On the harder multi-query associative recall (MQAR) task, Mamba-1 struggles while Mamba-2 performs well across the tested settings. Mamba-2 is reported to be significantly better than Mamba-1 even when state size is controlled at N=16, and increasing Mamba-2 state size from N=16 to N=64 and N=256 consistently improves performance. Standard multi-head softmax attention and Based are included as baselines. · CiteArk
On the harder multi-query associative recall (MQAR) task, Mamba-1 struggles while Mamba-2 performs well across the tested settings. Mamba-2 is reported to be significantly better than Mamba-1 even when state size is controlled at N=16, and increasing Mamba-2 state size from N=16 to N=64 and N=256 consistently improves performance. Standard multi-head softmax attention and Based are included as baselines.
来源:paper_markdown:Figure 8 and Section 9.1, PDF pages 28 and 30