Larger parameter sizes and more pretraining tokens often increase how frequently models generalize, but neither prevents fallback to shallow patterns. Small models can remain locked in a shallow strategy; large models exhibit the same instability on harder tasks even long after Chinchilla-optimal training budgets. · CiteArk
Larger parameter sizes and more pretraining tokens often increase how frequently models generalize, but neither prevents fallback to shallow patterns. Small models can remain locked in a shallow strategy; large models exhibit the same instability on harder tasks even long after Chinchilla-optimal training budgets.