Explore ArkGraph and select the steps to run.
This paper separates two capabilities that value gating can add to softmax attention: abstention, which lets a head return no output, and noise filtering, which suppresses interfering content in value reads. Matched language models from roughly 10M to 350M non-embedding parameters compare learned phantom sink logits, norm and projection gates, routing controls, and combined variants. Paired validation-loss results indicate that explicit abstention matters most at small scale while filtering grows more important with scale; their benefits are largely additive. Evaluation-time interference injections support the filtering account and reveal magnitude and direction blind spots. Synthetic and pretrained-model analyses provide additional mechanism evidence.
In plain language, attention heads benefit from two separate escape hatches: one for saying “nothing here is worth reading,” and another for cleaning noisy information that is read. The experiments suggest that larger models become better at imitating the first behavior with attention sinks, while still needing the second. A sink logit plus a projection gate performs best in the tested range, but the gains are small, the largest models use one seed, only one training corpus is studied, and behavior beyond 350M remains unknown.
The paper’s claims are available in Research claims.
Past research reports
Saved reports remain available. New report generation is paused.