Loading page…
γOPD assigns each token a discounted sum of its current and future teacher–student log-ratios; γ=0 recovers local token-level OPD and γ=1 recovers sequence-level return-to-go. Reward-Compatible Bounded Mixing response-normalizes and softsign-bounds that teacher term before adding a binary verifier reward, so the verifier reward fixes the update sign while teacher credit modulates token-level strength. · CiteArk