On the harder multi-query associative recall (MQAR) task, Mamba-1 struggles while Mamba-2 performs well across the tested settings. Mamba-2 is reported to be significantly better than Mamba-1 even when state size is controlled at N=16, and increasing Mamba-2 state size from N=16 to N=64 and N=256 consistently improves performance. Standard multi-head softmax attention and Based are included as baselines. · CiteArk
On the harder multi-query associative recall (MQAR) task, Mamba-1 struggles while Mamba-2 performs well across the tested settings. Mamba-2 is reported to be significantly better than Mamba-1 even when state size is controlled at N=16, and increasing Mamba-2 state size from N=16 to N=64 and N=256 consistently improves performance. Standard multi-head softmax attention and Based are included as baselines.
Source: paper_markdown:Figure 8 and Section 9.1, PDF pages 28 and 30
Reported and observed measurements
No structured measurement is attached to this Claim.
Assessments (0)
No immutable Assessment has been published for this Claim yet.