Loading page…
On RULER-4k retrieval, attention-KL reordering outperformed the MSE ordering at 0.5 and 1 bpd for every 7–8B model, with its largest objective separation among the paper's benchmarks; at 2 bpd it came within 0.7–3.6 points of FP16. · CiteArk