Explore ArkGraph and select the steps to run.
This paper reframes practical token-level on-policy distillation as a temporally truncated approximation to sequence-level reverse-KL optimization. It introduces γOPD, which discounts future token log-ratios to trade long-range credit assignment against variance, and reward-compatible bounded mixing, which combines bounded teacher-derived credit with binary verifier rewards. Across mathematical reasoning, code generation, same-size, smaller-student, and multi-teacher settings, the paper reports generally stronger aggregate performance than compared OPD methods. Ablations attribute gains to complementary temporal discounting and bounded reward mixing, while training analyses report steadier optimization and negligible additional overhead. Experiments are confined to models below 10B parameters.
In plain language, the method lets earlier reasoning tokens receive feedback from what happens later while keeping that feedback from becoming unstable, and it lets objective correctness signals overrule pure imitation. The reported gains cover several math and code benchmarks and teacher–student arrangements, but they come from a specific family of sub-10B Qwen models and selected training data. Some headline improvement percentages conflict with the printed tables, and the token-credit examples are qualitative illustrations rather than broad quantitative evidence.
The paper’s claims are available in Research claims.
Past research reports
Saved reports remain available. New report generation is paused.