Across GSM8K, MATH-500, Countdown, and Sudoku, GDPO+IM reaches reward levels faster and finishes at a higher reward than GDPO. The paper says the higher plateau is especially pronounced on Countdown and Sudoku, while on MATH-500 IM matches the base method’s sustained rise in fewer steps. · CiteArk
Across GSM8K, MATH-500, Countdown, and Sudoku, GDPO+IM reaches reward levels faster and finishes at a higher reward than GDPO. The paper says the higher plateau is especially pronounced on Countdown and Sudoku, while on MATH-500 IM matches the base method’s sustained rise in fewer steps.
来源:paper-fixed:PDF pp. 14–15, Appendix C Reward Dynamics and Figure 4