正在加载页面…
The main experiments train matched GPT-2-style decoder-only transformers from scratch on FineWeb-Edu at four non-embedding parameter tiers, sharing architecture, data order, optimizer, schedule, and paired seeds within each tier; three seeds are reported for most arms through 124M and one seed per arm at 350M. · CiteArk