正在加载页面…
Across the A100 attention benchmarks, the paper summarizes FlashAttention-2 as 1.7–3.0× faster than FlashAttention, 1.3–2.5× faster than the Triton implementation, and 3–10× faster than standard PyTorch attention. A more operation-specific summary says 1.3–1.5× over Triton in forward and around 2× in backward, and around 2× over FlashAttention and xFormers overall. · CiteArk