Loading page…
The paper reports that its SSD implementation is 2–8× faster than Mamba's fused associative scan for large state expansion, is faster than FlashAttention-2 at sequence lengths of 2K and above, and is 6× faster than FlashAttention-2 at sequence length 16K. At sequence length 4K, increasing state expansion slows the optimized Mamba scan approximately linearly, whereas SSD shows little slowdown. · CiteArk