Loading page…
After matched math SFT, the selected 4.5T OLMo3-32B checkpoint generalizes better to GPQA than the 4.9T checkpoint (36.3% versus 29.8%). The wider checkpoint comparison shows its strongest GPQA gain despite further pretraining/mid-training improving in-distribution performance. · CiteArk