正在加载页面…
Without RBM in Qwen3-4B-Math→Qwen3-1.7B training, γ=1 produces much larger gradient norms, rapid policy-entropy collapse, and inferior AIME24 validation accuracy. Local OPD (γ=0) and discounted variants improve steadily, with γ=0.99 reported as best and better than γ=0.9. · CiteArk