At short sequence lengths such as 2K, the complete Mamba-2 model may be less efficient to train than a same-parameter Transformer because the Transformer uses L/2 MLP layers and L/2 attention layers while Mamba-2 uses L SSD layers. The paper notes that MLP layers are very hardware efficient and that SSD/MLP mixtures can speed training at short sequence lengths. · CiteArk
Not assessedNo independent reproduction scheduledLimitationclaim-short-sequence-limit
At short sequence lengths such as 2K, the complete Mamba-2 model may be less efficient to train than a same-parameter Transformer because the Transformer uses L/2 MLP layers and L/2 attention layers while Mamba-2 uses L SSD layers. The paper notes that MLP layers are very hardware efficient and that SSD/MLP mixtures can speed training at short sequence lengths.
Source: paper_markdown:Section 9.3, PDF page 31
Reported and observed measurements
No structured measurement is attached to this Claim.
Assessments (0)
No immutable Assessment has been published for this Claim yet.