arXiv 预印本 2023
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
ganwumeng/flashattention-2-faster-attention-with-better-parallelism-and-work-partitioning--lpussw
FlashAttention-2 is an exact GPU attention algorithm and implementation that improves on FlashAttention by reducing non-matrix-multiplication work, parallelizing computation across sequence blocks, and repartitioning work among warps to reduce shared-memory communication. The paper benchmarks forward, backward, and combined attention on A100 GPUs across causal and non-causal settings, two head dimensions, and sequence lengths from 512 to 16k, and also reports unoptimized H100 results. It evaluates end-to-end GPT-style training for 1.3B- and 2.7B-parameter models at 2k and 8k context. Reported gains reach roughly twofold over FlashAttention and 225 TFLOPs/s/GPU in training.