Fine-tuning generalization varies abruptly across pretraining checkpoints: anonymized-function identification and city single-/multi-hop reasoning fluctuate. City multi-hop performance is weaker than city identification, with many near-zero checkpoints in the plotted small-model results. · CiteArk
Fine-tuning generalization varies abruptly across pretraining checkpoints: anonymized-function identification and city single-/multi-hop reasoning fluctuate. City multi-hop performance is weaker than city identification, with many near-zero checkpoints in the plotted small-model results.