正在加载页面…
The paper reports that larger models do not show an advantage in generalization at the same pretraining FLOPs, qualifying the advantage seen when comparing sizes at the same token count. · CiteArk