正在加载页面…
On all 90 AIME 2024–2026 problems, per-request mean accepted length was broad for all three models and accepted length also fluctuated between verification rounds within requests. Reported spans/medians were 8.24–14.96/11.63 for Qwen3-1.7B, 9.24–16.46/12.03 for Qwen3-8B, and 8.05–17.64/12.22 for DS-Llama-8B; Figure 4 additionally prints distribution means of 11.60, 12.02, and 12.25. · CiteArk