Start with coverage and measured results; open a run only when you need evidence or technical details.
RunMatchRepeat
0/27
claims supported by evidence
0
Supported
0
Challenged or mixed
0
Contradicted
2
Inconclusive
25
Not assessed
Run history
Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.
ClaimResultFinishedStatus
Failure log
Failed paths grouped by cause — check them before reproducing.
Resource limits1 failures
Environment & dependencies1 failures
Claim–experiment reproduction matrix
See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions. There are also 2 claims with no independent reproduction scheduled in this plan.
20 targets·20 with evidence·0 active·0 need attention
A successful execution does not by itself validate a paper claim
Target state says whether the platform completed the work. The scientific conclusion is determined only by immutable evidence and Assessments. Resource shortages and platform failures are never presented as scientific contradictions.
Figure 6c reports backward attention throughput on A100 80GB SXM4; PyTorch is marked OOM at sequence length 16k. The panel covers causal attention with head dimension 64.
Official implementationfig6c1 plan0 runs
Scientific conclusionNot assessed
FlashAttention-2 A100 benchmark-grid portability check on L4
exp-a100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Figure 6a reports backward attention throughput on A100 80GB SXM4; PyTorch is marked OOM at sequence length 16k. The panel covers non-causal attention with head dimension 64.
Official implementationfig6a1 plan0 runs
Scientific conclusionNot assessed
FlashAttention-2 A100 benchmark-grid portability check on L4
exp-a100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Table 1 reports per-GPU training throughput for GPT3-2.7B at 2k context on 8×A100 GPUs for the baseline without FlashAttention, FlashAttention, and FlashAttention-2.
Information insufficienttable1-gpt3-2-7b-2k0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Figure 7c reports forward-plus-backward attention throughput on H100 80GB SXM5; PyTorch is marked OOM at sequence length 16k. The panel covers causal attention with head dimension 64.
Official implementationfig7c1 plan0 runs
Scientific conclusionNot assessed
Unoptimized H100 attention-panel portability check on L4
exp-h100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
The H100 measurements use the unmodified implementation and no H100-specific TMA or fourth-generation Tensor Core instructions. The authors forecast, rather than measure, a further 1.5–2× speedup from those features and leave that optimization to future work.
Information insufficientlimitation-0010 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Table 1 reports per-GPU training throughput for GPT3-1.3B at 2k context on 8×A100 GPUs for the baseline without FlashAttention, FlashAttention, and FlashAttention-2.
Information insufficienttable1-gpt3-1-3b-2k0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Figure 5c reports forward attention throughput on A100 80GB SXM4; PyTorch is marked OOM at sequence length 16k. The panel covers causal attention with head dimension 64.
Official implementationfig5c1 plan0 runs
Scientific conclusionNot assessed
FlashAttention-2 A100 benchmark-grid portability check on L4
exp-a100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Figure 7b reports forward-plus-backward attention throughput on H100 80GB SXM5; PyTorch is marked OOM at sequence length 16k. The panel covers non-causal attention with head dimension 128.
Official implementationfig7b1 plan0 runs
Scientific conclusionNot assessed
Unoptimized H100 attention-panel portability check on L4
exp-h100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Figure 4b reports forward-plus-backward attention throughput on A100 80GB SXM4; PyTorch is marked OOM at sequence length 16k. The panel covers non-causal attention with head dimension 128.
Official implementationfig4b1 plan0 runs
Scientific conclusionNot assessed
FlashAttention-2 A100 benchmark-grid portability check on L4
exp-a100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Figure 4d reports forward-plus-backward attention throughput on A100 80GB SXM4; PyTorch is marked OOM at sequence length 16k. The panel covers causal attention with head dimension 128.
Official implementationfig4d1 plan0 runs
Scientific conclusionNot assessed
FlashAttention-2 A100 benchmark-grid portability check on L4
exp-a100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Figure 5b reports forward attention throughput on A100 80GB SXM4; PyTorch is marked OOM at sequence length 16k. The panel covers non-causal attention with head dimension 128.
Official implementationfig5b1 plan0 runs
Scientific conclusionNot assessed
FlashAttention-2 A100 benchmark-grid portability check on L4
exp-a100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Table 1 reports per-GPU training throughput for GPT3-1.3B at 8k context on 8×A100 GPUs for the baseline without FlashAttention, FlashAttention, and FlashAttention-2.
Information insufficienttable1-gpt3-1-3b-8k0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Figure 4a reports forward-plus-backward attention throughput on A100 80GB SXM4; PyTorch is marked OOM at sequence length 16k. The panel covers non-causal attention with head dimension 64.
Official implementationfig4a1 plan0 runs
Scientific conclusionNot assessed
FlashAttention-2 A100 benchmark-grid portability check on L4
exp-a100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Figure 5d reports forward attention throughput on A100 80GB SXM4; PyTorch is marked OOM at sequence length 16k. The panel covers causal attention with head dimension 128.
Official implementationfig5d1 plan0 runs
Scientific conclusionNot assessed
FlashAttention-2 A100 benchmark-grid portability check on L4
exp-a100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Figure 6d reports backward attention throughput on A100 80GB SXM4; PyTorch is marked OOM at sequence length 16k. The panel covers causal attention with head dimension 128.
Official implementationfig6d1 plan0 runs
Scientific conclusionNot assessed
FlashAttention-2 A100 benchmark-grid portability check on L4
exp-a100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
The A100 results are summarized as reaching up to 230 TFLOPs/s, 73% of theoretical maximum in forward, and 63% of theoretical maximum in backward.
Official implementationfinding-0031 plan0 runs
Scientific conclusionNot assessed
FlashAttention-2 A100 benchmark-grid portability check on L4
exp-a100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Figure 5a reports forward attention throughput on A100 80GB SXM4; PyTorch is marked OOM at sequence length 16k. The panel covers non-causal attention with head dimension 64.
Official implementationfig5a1 plan0 runs
Scientific conclusionNot assessed
FlashAttention-2 A100 benchmark-grid portability check on L4
exp-a100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Running the same implementation on H100 without H100-specific instructions is described in the prose as reaching up to 335 TFLOPs/s.
Official implementationfinding-0041 plan0 runs
Scientific conclusionNot assessed
Unoptimized H100 attention-panel portability check on L4
exp-h100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Figure 4c reports forward-plus-backward attention throughput on A100 80GB SXM4; PyTorch is marked OOM at sequence length 16k. The panel covers causal attention with head dimension 64.
Official implementationfig4c1 plan0 runs
Scientific conclusionNot assessed
FlashAttention-2 A100 benchmark-grid portability check on L4
exp-a100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Across the A100 attention benchmarks, the paper summarizes FlashAttention-2 as 1.7–3.0× faster than FlashAttention, 1.3–2.5× faster than the Triton implementation, and 3–10× faster than standard PyTorch attention. A more operation-specific summary says 1.3–1.5× over Triton in forward and around 2× in backward, and around 2× over FlashAttention and xFormers overall.
Official implementationfinding-0021 plan0 runs
Scientific conclusionNot assessed
FlashAttention-2 A100 benchmark-grid portability check on L4
exp-a100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
The reported end-to-end TFLOPs/s uses 6 × sequence_length × parameter_count + 12 × layer_count × hidden_dimension × sequence_length². The paper notes that the attention term could be halved for causal masking, but deliberately does not halve it to remain consistent with Megatron-LM and prior reporting.
Information insufficientlimitation-0020 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Figure 7d reports forward-plus-backward attention throughput on H100 80GB SXM5; PyTorch is marked OOM at sequence length 16k. The panel covers causal attention with head dimension 128.
Official implementationfig7d1 plan0 runs
Scientific conclusionNot assessed
Unoptimized H100 attention-panel portability check on L4
exp-h100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
For 1.3B- and 2.7B-parameter GPT-style models at 2k and 8k context, FlashAttention-2 is summarized as up to 1.3× faster than FlashAttention and 2.8× faster than a baseline without FlashAttention, reaching 225 TFLOPs/s/GPU and 72% model-FLOPs utilization.
Information insufficientfinding-0050 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Figure 6b reports backward attention throughput on A100 80GB SXM4; PyTorch is marked OOM at sequence length 16k. The panel covers non-causal attention with head dimension 128.
Official implementationfig6b1 plan0 runs
Scientific conclusionNot assessed
FlashAttention-2 A100 benchmark-grid portability check on L4
exp-a100-grid-l4-cross-hardware1 attempt
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Table 1 reports per-GPU training throughput for GPT3-2.7B at 8k context on 8×A100 GPUs for the baseline without FlashAttention, FlashAttention, and FlashAttention-2.
Information insufficienttable1-gpt3-2-7b-8k0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Figure 7a reports forward-plus-backward attention throughput on H100 80GB SXM5; PyTorch is marked OOM at sequence length 16k. The panel covers non-causal attention with head dimension 64.
Official implementationfig7a1 plan1 run
Scientific conclusionInconclusive
Unoptimized H100 attention-panel portability check on L4
exp-h100-grid-l4-cross-hardware2 attempts
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
Assessed but inconclusive
Legacy task without target-level resource requirements
For causal attention, skipping blocks wholly above the causal diagonal is reported to make the kernel around 1.7–1.8× faster than attention without a causal mask; only one square block per row needs elementwise causal masking.
Official implementationfinding-0011 plan1 run
Scientific conclusionInconclusive
FlashAttention-2 A100 benchmark-grid portability check on L4
exp-a100-grid-l4-cross-hardware3 attempts
Evidence published
Next step
The platform will retry with the same resource profile.