浏览 ArkGraph,选择本次要执行的步骤。
The paper introduces KV-Invariant Transformer Expansion (KITE), a staged language-model scaling paradigm that adds capacity outside the key-value-producing path. Its Step Scale Transformer (SST) instantiation trains a source tower, adds a second KV-reading tower, and jointly continues training while retaining source-sized bulk prefill. The study compares a 66.959B-parameter SST model with 46.727B and 62.691B classic MoE Transformers under approximately matched theoretical cumulative training FLOPs. It reports lower final training loss for SST, higher scores on seven downstream tasks, slightly lower held-out ARXIV NLL, and lower analytical inference cost in prefill-heavy workload mixes, while emphasizing that serving speedups and causal contributions were not measured.
KITE suggests that a model can gain prediction capacity partway through training without making prompt-wide KV construction equally large. In the reported comparison, SST improves quality and an analytical prefill-heavy cost proxy at roughly matched estimated training compute. The evidence is limited to one selected SST configuration and theoretical cost accounting: it does not establish a scaling law, isolate which design ingredient causes the gains, identify an optimal connectivity pattern, or demonstrate real serving speedups.
论文中的结论可在「研究结论」中查看。