2 experiments
Execution agent’s result notes
Executed the complete assigned Figure 7 workload on one NVIDIA L4 (SM89, 23,034 MiB, driver 535.309.01) as an explicitly cross-hardware test. Every successful cell used FP16, five warmups, and 30 forward plus 30 backward CUDA-event samples. Fresh evidence provides 54 of 68 panel values plus both derived peaks: FA2 peaked at 58.000924 TFLOPs/s, and the non-causal d=128 16k FA2 cell was 55.108025 TFLOPs/s. FA2 beat PyTorch in all 18 available pairs but candidate FlashAttention v1.0.9 in only 10 of 12 available pairs, so the predeclared strict directional criterion was not met on L4. Twelve v1 d=128 cells were unsupported on SM89 and two PyTorch d=64 8k cells OOMed even in isolated expandable-segment retries.
A real two-configuration NVIDIA L4 probe ran with the exact fixed FlashAttention-2 source, two warmups, and five repetitions. FlashAttention-2 passed output/gradient checks and measured 49.927/30.172/34.018 TFLOPs/s forward/backward/combined at sequence 512 and 57.543/57.982/57.856 at sequence 8192. Both xFormers baselines ran; the historical Triton callable was unavailable, and PyTorch OOMed at 8192. The cumulative runner budget did not leave enough time for the full grid.
Figure 7a reports forward-plus-backward attention throughput on H100 80GB SXM5; PyTorch is marked OOM at sequence length 16k. The panel covers non-causal attention with head dimension 64.
Metric: attention throughput · TFLOPs/s
Data and sources · 16
| Source | Conditions | Metric | Value | Assessment and limits |
|---|---|---|---|---|
| Paperpaper:PDF p. 13, Figure 7afig7a/fig7a-512-pytorch/paper | method: PyTorch · sequenceLength: 512 | attention throughput | 62 TFLOPs/s | reported |
| Reproductionpaper:PDF p. 13, Figure 7afig7a/fig7a-512-pytorch/urn:citeark:assessment:c037e42a27da55578fd04b8296ae6e8745703dee93cbf78828d24a308ac61b3d | method: PyTorch · sequenceLength: 512 | attention throughput | 5.800733954913403 TFLOPs/s | inconclusive |
| Paperpaper:PDF p. 13, Figure 7afig7a/fig7a-512-flashattention/paper | method: FlashAttention · sequenceLength: 512 | attention throughput | 157 TFLOPs/s | reported |
| Reproductionpaper:PDF p. 13, Figure 7afig7a/fig7a-512-flashattention/urn:citeark:assessment:c037e42a27da55578fd04b8296ae6e8745703dee93cbf78828d24a308ac61b3d | method: FlashAttention · sequenceLength: 512 | attention throughput | 33.169356691259516 TFLOPs/s | inconclusive |
| Paperpaper:PDF p. 13, Figure 7afig7a/fig7a-512-flashattention-2/paper | method: FlashAttention-2 · sequenceLength: 512 | attention throughput | 215 TFLOPs/s | reported |
| Reproductionpaper:PDF p. 13, Figure 7afig7a/fig7a-512-flashattention-2/urn:citeark:assessment:c037e42a27da55578fd04b8296ae6e8745703dee93cbf78828d24a308ac61b3d | method: FlashAttention-2 · sequenceLength: 512 | attention throughput | 32.21329935387607 TFLOPs/s | inconclusive |
| Paperpaper:PDF p. 13, Figure 7afig7a/fig7a-1k-pytorch/paper | method: PyTorch · sequenceLength: 1k | attention throughput | 72 TFLOPs/s | reported |
| Reproductionpaper:PDF p. 13, Figure 7afig7a/fig7a-1k-pytorch/urn:citeark:assessment:c037e42a27da55578fd04b8296ae6e8745703dee93cbf78828d24a308ac61b3d | method: PyTorch · sequenceLength: 1k | attention throughput | 6.594632642560088 TFLOPs/s | inconclusive |
| Paperpaper:PDF p. 13, Figure 7afig7a/fig7a-1k-flashattention/paper | method: FlashAttention · sequenceLength: 1k | attention throughput | 159 TFLOPs/s | reported |
| Reproductionpaper:PDF p. 13, Figure 7afig7a/fig7a-1k-flashattention/urn:citeark:assessment:c037e42a27da55578fd04b8296ae6e8745703dee93cbf78828d24a308ac61b3d | method: FlashAttention · sequenceLength: 1k | attention throughput | 33.645996966704885 TFLOPs/s | inconclusive |
| Paperpaper:PDF p. 13, Figure 7afig7a/fig7a-1k-flashattention-2/paper | method: FlashAttention-2 · sequenceLength: 1k | attention throughput | 254 TFLOPs/s | reported |
| Reproductionpaper:PDF p. 13, Figure 7afig7a/fig7a-1k-flashattention-2/urn:citeark:assessment:c037e42a27da55578fd04b8296ae6e8745703dee93cbf78828d24a308ac61b3d | method: FlashAttention-2 · sequenceLength: 1k | attention throughput | 39.662784327600185 TFLOPs/s | inconclusive |
| Paperpaper:PDF p. 13, Figure 7afig7a/fig7a-2k-pytorch/paper | method: PyTorch · sequenceLength: 2k | attention throughput | 81 TFLOPs/s | reported |
| Reproductionpaper:PDF p. 13, Figure 7afig7a/fig7a-2k-pytorch/urn:citeark:assessment:c037e42a27da55578fd04b8296ae6e8745703dee93cbf78828d24a308ac61b3d | method: PyTorch · sequenceLength: 2k | attention throughput | 7.062177895084232 TFLOPs/s | inconclusive |
| Paperpaper:PDF p. 13, Figure 7afig7a/fig7a-2k-flashattention/paper | method: FlashAttention · sequenceLength: 2k | attention throughput | 161 TFLOPs/s | reported |
| Reproductionpaper:PDF p. 13, Figure 7afig7a/fig7a-2k-flashattention/urn:citeark:assessment:c037e42a27da55578fd04b8296ae6e8745703dee93cbf78828d24a308ac61b3d | method: FlashAttention · sequenceLength: 2k | attention throughput | 33.23882163846561 TFLOPs/s | inconclusive |
Comparison conditions
The paper used an H100 80GB SXM5; L4 measurements cannot verify the reported H100 335/338 TFLOPs/s values or hardware utilization.
The contract-fixed FlashAttention-2 source identifies as package 2.8.4 and postdates the paper. Its assigned CUDA kernels were unchanged, but an evaluation-only host dispatch was bounded to FP16 head dimensions 64/128 and rejected unassigned SplitKV paths.
The paper does not identify the original FlashAttention package version. The retained official v1.0.9 is only a source-grounded candidate and cannot execute d=128 backward on SM89, leaving 12 assigned comparison values unavailable.
PyTorch d=64 8k non-causal and causal backward exceeded the L4 memory capacity, leaving two assigned values unavailable; isolated fresh-process retries confirmed this was not allocator fragmentation.
The four PyTorch 16k figure positions have no numeric measurement IDs in the fixed contract and were retained as not-assigned rather than benchmarked.
Unresolved measurements: fig7a-8k-pytorch: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值; fig7b-512-flashattention: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值; fig7b-1k-flashattention: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值; fig7b-2k-flashattention: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值; fig7b-4k-flashattention: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值; fig7b-8k-flashattention: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值; fig7b-16k-flashattention: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值; fig7c-8k-pytorch: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值; fig7d-512-flashattention: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值; fig7d-1k-flashattention: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值; fig7d-2k-flashattention: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值; fig7d-4k-flashattention: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值; fig7d-8k-flashattention: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值; fig7d-16k-flashattention: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值
Execution used a roughly 23GB NVIDIA L4 instead of an H100 80GB SXM5, materially affecting throughput and memory behavior.
PyTorch at 16k—the claim's central OOM observation—was not run; it was skipped based on the paper label.
The listed measurements omit a numeric PyTorch 8k result because that L4 attempt OOMed; this does not establish H100 behavior.
The paper does not identify the exact FlashAttention baseline version; v1.0.9 was only a candidate.
The evaluated FlashAttention-2 source identifies as later package 2.8.4 and used a bounded evaluation-only host dispatch.
The inherited 16k-total-token and hidden-dimension-2048 grid remains an assumption because the H100 paragraph does not independently restate it.
The captured run freshly measured the listed Figure 7a cells with 30 forward and 30 backward CUDA-event samples, but on an NVIDIA L4 rather than the claimed H100 80GB SXM5. Since throughput is device-dependent, these absolute values are not comparable to Figure 7a. The central qualitative claim that PyTorch is OOM at 16k was not executed: the evaluator skipped PyTorch 16k and labeled it not_assigned based on the paper. The run therefore establishes only L4 portability behavior, not the H100 throughput or OOM claim. Some original claim measurements remain without evidence; available measurements are assessed individually.