正在加载页面…
On AIME24 and AIME25, temporal discounting alone raises average accuracy by 2.04 points over vanilla OPD; naive reward mixing adds only 0.05 point beyond discounting, while the full discounting-plus-mixing-plus-normalization configuration reaches 57.56 average accuracy, a 3.91-point gain. Mixing plus normalization without discounting gains 2.00 points, supporting complementary contributions. · CiteArk