FlashAttention-2 computes exact scaled-dot-product attention with no approximation, using tiled online softmax. It changes the earlier algorithm by delaying output normalization until the end, storing only row-wise logsumexp for backward, parallelizing over sequence blocks, and repartitioning warp work to avoid split-K communication where possible. · CiteArk
Not assessedNo independent reproduction scheduledMethodmethod-001
FlashAttention-2 computes exact scaled-dot-product attention with no approximation, using tiled online softmax. It changes the earlier algorithm by delaying output normalization until the end, storing only row-wise logsumexp for backward, parallelizing over sequence blocks, and repartitioning warp work to avoid split-K communication where possible.
Source: paper:PDF p. 5, §3.1 through PDF p. 9, §3.3
Reported and observed measurements
No structured measurement is attached to this Claim.
Assessments (0)
No immutable Assessment has been published for this Claim yet.