The paper reports that larger models do not show an advantage in generalization at the same pretraining FLOPs, qualifying the advantage seen when comparing sizes at the same token count. · CiteArk
The paper reports that larger models do not show an advantage in generalization at the same pretraining FLOPs, qualifying the advantage seen when comparing sizes at the same token count.