FlashAttention-2 computes exact scaled-dot-product attention with no approximation, using tiled online softmax. It changes the earlier algorithm by delaying output normalization until the end, storing only row-wise logsumexp for backward, parallelizing over sequence blocks, and repartitioning warp work to avoid split-K communication where possible. · CiteArk
FlashAttention-2 computes exact scaled-dot-product attention with no approximation, using tiled online softmax. It changes the earlier algorithm by delaying output normalization until the end, storing only row-wise logsumexp for backward, parallelizing over sequence blocks, and repartitioning warp work to avoid split-K communication where possible.