Explore ArkGraph and select the steps to run.
This study tracks language-model generalization during pretraining using six behavioral evaluations and additional fine-tuning tests. Across OLMo3 and Apertus checkpoints, models repeatedly switch between following shallow patterns and producing answers consistent with task structure, even late in training. The reported fluctuations persist with probability-based metrics and alternative training-compute axes, while common benchmarks remain comparatively stable. Single-step updates and checkpoint averaging do not eliminate the phenomenon. Larger models generalize more often, but still fluctuate. Intermediate checkpoint selection improves post-training reasoning transfer and resistance to prefilling attacks in the reported comparisons. A preliminary data-selection experiment also steers generalization dynamics. Complexity proxies give inconsistent predictions across layers and datasets.
A model that performs well on familiar benchmarks may still alternate between solving a task and following a tempting shortcut. These results suggest that the final pretraining checkpoint need not transfer best to later reasoning or safety training. The evaluations are behavioral probes; they do not establish the proposed circuit-level explanation, and the data-selection demonstration is preliminary.
The paper’s claims are available in Research claims.