Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation
Shiqi Liu, Zeyu He, Letian Tao, Guojian Zhan, Jiaxin Gao, Feihong Zhang, Jingliang Duan, Wei Xiong, Kehua Sheng, Bo Zhang, Yang Guan, Shengbo Eben Li
Why it is worth reading
In plain language, the method lets earlier reasoning tokens receive feedback from what happens later while keeping that feedback from becoming unstable, and it lets objective correctness signals overrule pure imitation. The reported gains cover several math and code benchmarks and teacher–student arrangements, but they come from a specific family of sub-10B Qwen models and selected training data. Some headline improvement percentages conflict with the printed tables, and the token-credit examples are qualitative illustrations rather than broad quantitative evidence.
Core research claims
- Because of computational constraints, the experiments are limited to models smaller than 10B parameters; the paper leaves larger-scale evaluation and adaptive or entropy-aware discounting for future work.
- The experiments use veRL with vLLM rollouts and FSDP actor training; OPD-based methods use one response per prompt, batch and PPO minibatch size 1024, maximum prompt length 2048, maximum response length 16384, RADAR with constant 1×10^-5 learning rate, temperature/top-p 1.0/1.0, and tensor parallel size 4. Math uses a DAPO-style boxed-answer verifier and code uses execution tests; binary outcomes map to ±1.
- γOPD assigns each token a discounted sum of its current and future teacher–student log-ratios; γ=0 recovers local token-level OPD and γ=1 recovers sequence-level return-to-go. Reward-Compatible Bounded Mixing response-normalizes and softsign-bounds that teacher term before adding a binary verifier reward, so the verifier reward fixes the update sign while teacher credit modulates token-level strength.