Fine-tuning generalization varies abruptly across pretraining checkpoints: anonymized-function identification and city single-/multi-hop reasoning fluctuate. City multi-hop performance is weaker than city identification, with many near-zero checkpoints in the plotted small-model results. · CiteArk
Not assessedPlan blockedFindingfine-tuning-ooc-mode-hopping
Fine-tuning generalization varies abruptly across pretraining checkpoints: anonymized-function identification and city single-/multi-hop reasoning fluctuate. City multi-hop performance is weaker than city identification, with many near-zero checkpoints in the plotted small-model results.