Without RBM in Qwen3-4B-Math→Qwen3-1.7B training, γ=1 produces much larger gradient norms, rapid policy-entropy collapse, and inferior AIME24 validation accuracy. Local OPD (γ=0) and discounted variants improve steadily, with γ=0.99 reported as best and better than γ=0.9. · CiteArk
Not assessedPlan blockedFindingclaim-gamma-sensitivity-figure4
Without RBM in Qwen3-4B-Math→Qwen3-1.7B training, γ=1 produces much larger gradient norms, rapid policy-entropy collapse, and inferior AIME24 validation accuracy. Local OPD (γ=0) and discounted variants improve steadily, with γ=0.99 reported as best and better than γ=0.9.
Source: paper:PDF pp. 8–9, Section 4.4 and Figure 4
Reported and observed measurements
No structured measurement is attached to this Claim.
Assessments (0)
No immutable Assessment has been published for this Claim yet.