正在加载页面…
Across GSM8K, MATH-500, Countdown, and Sudoku, GDPO+IM reaches reward levels faster and finishes at a higher reward than GDPO. The paper says the higher plateau is especially pronounced on Countdown and Sudoku, while on MATH-500 IM matches the base method’s sustained rise in fewer steps. · CiteArk