Loading page…
On the harder multi-query associative recall (MQAR) task, Mamba-1 struggles while Mamba-2 performs well across the tested settings. Mamba-2 is reported to be significantly better than Mamba-1 even when state size is controlled at N=16, and increasing Mamba-2 state size from N=16 to N=64 and N=256 consistently improves performance. Standard multi-head softmax attention and Based are included as baselines. · CiteArk