γOPD assigns each token a discounted sum of its current and future teacher–student log-ratios; γ=0 recovers local token-level OPD and γ=1 recovers sequence-level return-to-go. Reward-Compatible Bounded Mixing response-normalizes and softsign-bounds that teacher term before adding a binary verifier reward, so the verifier reward fixes the update sign while teacher credit modulates token-level strength. · CiteArk
Not assessedNo independent reproduction scheduledMethodclaim-method-temporal-credit-rbm
γOPD assigns each token a discounted sum of its current and future teacher–student log-ratios; γ=0 recovers local token-level OPD and γ=1 recovers sequence-level return-to-go. Reward-Compatible Bounded Mixing response-normalizes and softsign-bounds that teacher term before adding a binary verifier reward, so the verifier reward fixes the update sign while teacher credit modulates token-level strength.
Source: paper:PDF pp. 4–5, Equations 7–9 and Sections 3.2–3.3
Reported and observed measurements
No structured measurement is attached to this Claim.
Assessments (0)
No immutable Assessment has been published for this Claim yet.