Loading page…
At 2.7B scale, trained for 300B Pile tokens with 64 layers and matched parameters, adding six attention layers to Mamba-2 improves the reported Pile and downstream results over pure Mamba-2 and Transformer++. Adding MLP layers alone reduces the reported quality, while the combined SSD/MLP/attention model remains competitive. The paper reports that MLP layers may improve hardware efficiency or ease conversion to mixture-of-experts models. · CiteArk