Loading page…
At short sequence lengths such as 2K, the complete Mamba-2 model may be less efficient to train than a same-parameter Transformer because the Transformer uses L/2 MLP layers and L/2 attention layers while Mamba-2 uses L SSD layers. The paper notes that MLP layers are very hardware efficient and that SSD/MLP mixtures can speed training at short sequence lengths. · CiteArk