正在加载页面…
In a 350M, 48-layer Pile model trained for 7B tokens with the GPT-2 tokenizer, adding a small number of attention blocks to Mamba-2 improves validation perplexity. The best reported result is around a 10% attention-layer ratio; the paper says SSD and attention are complementary and that exact spacing is not very important in these small-scale experiments. · CiteArk