Public research repositories on CiteArk, sorted by activity, update time, claims with supporting Assessments, and community reproduction requests.
Subject Programming Languages · 1 repository
Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation
This paper reframes practical token-level on-policy distillation as a temporally truncated approximation to sequence-level reverse-KL optimization. It introduces γOPD, which discounts future token log-ratios to trade long-range credit assignment against variance, and reward-compatible bounded mixing, which combines bounded teacher-derived credit with binary verifier rewards. Across mathematical reasoning, code generation, same-size, smaller-student, and multi-teacher settings, the paper reports generally stronger aggregate performance than compared OPD methods. Ablations attribute gains to complementary temporal discounting and bounded reward mixing, while training analyses report steadier optimization and negligible additional overhead. Experiments are confined to models below 10B parameters.