Larger parameter sizes and more pretraining tokens often increase how frequently models generalize, but neither prevents fallback to shallow patterns. Small models can remain locked in a shallow strategy; large models exhibit the same instability on harder tasks even long after Chinchilla-optimal training budgets. · CiteArk
Not assessedNo independent reproduction scheduledLimitationscaling-not-stable-cure
Larger parameter sizes and more pretraining tokens often increase how frequently models generalize, but neither prevents fallback to shallow patterns. Small models can remain locked in a shallow strategy; large models exhibit the same instability on harder tasks even long after Chinchilla-optimal training budgets.
Source: paper:§1, p.2; §6, p.9; Figures 2–8 and 18–25
Reported and observed measurements
No structured measurement is attached to this Claim.
Assessments (0)
No immutable Assessment has been published for this Claim yet.