浏览 ArkGraph,选择本次要执行的步骤。
This paper separates two capabilities that value gating can add to softmax attention: abstention, which lets a head return no output, and noise filtering, which suppresses interfering content in value reads. Matched language models from roughly 10M to 350M non-embedding parameters compare learned phantom sink logits, norm and projection gates, routing controls, and combined variants. Paired validation-loss results indicate that explicit abstention matters most at small scale while filtering grows more important with scale; their benefits are largely additive. Evaluation-time interference injections support the filtering account and reveal magnitude and direction blind spots. Synthetic and pretrained-model analyses provide additional mechanism evidence.
In plain language, attention heads benefit from two separate escape hatches: one for saying “nothing here is worth reading,” and another for cleaning noisy information that is read. The experiments suggest that larger models become better at imitating the first behavior with attention sinks, while still needing the second. A sink logit plus a projection gate performs best in the tested range, but the gains are small, the largest models use one seed, only one training corpus is studied, and behavior beyond 350M remains unknown.
论文中的结论可在「研究结论」中查看。
历史研究报告
此前保存的报告仍可阅读,新报告生成已暂停。