正在加载页面…
In vanilla-distillation training curves, the paper reports that γOPD maintains the highest verifiable reward, stable absolute OPD advantage, and the smallest gradient-norm fluctuations; its response length becomes shorter and more stable than most baselines. TOPD is the exception on length because it truncates the distillation signal, and its verifiable reward remains mostly below zero. · CiteArk