Additional pretraining or mid-training improves in-distribution math performance without yielding the best generalizable reasoning or the most robust alignment. The selected 4.5T checkpoint remains best in the additional checkpoint comparisons. · CiteArk
Not assessedNo independent reproduction scheduledLimitationlater-training-transfer-limit
Additional pretraining or mid-training improves in-distribution math performance without yielding the best generalizable reasoning or the most robust alignment. The selected 4.5T checkpoint remains best in the additional checkpoint comparisons.