Start with coverage and measured results; open a run only when you need evidence or technical details.
RunMatchRepeat
0/41
claims supported by evidence
0
Supported
0
Challenged or mixed
0
Contradicted
18
Inconclusive
23
Not assessed
Supporting Assessment
Measurement confirmed
Table 10 reports the full zero-shot evaluation row for GPT-Neo 2.7B.
Reported
5.63 perplexity
Observed
5.6244 perplexity
Difference -0.0056
Run history
Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.
ClaimResultFinishedStatus
Failure log
Failed paths grouped by cause — check them before reproducing.
Environment & dependencies6 failures
Data access5 failures
Other3 failures
Claim–experiment reproduction matrix
See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions. There are also 2 claims with no independent reproduction scheduled in this plan.
31 targets·18 with evidence·0 active·1 need attention
A successful execution does not by itself validate a paper claim
Target state says whether the platform completed the work. The scientific conclusion is determined only by immutable evidence and Assessments. Resource shortages and platform failures are never presented as scientific contradictions.
Table 10 reports the full zero-shot evaluation row for OPT-6.7B.
Official implementationclaim-table10-opt-6-7b1 plan0 runs
Scientific conclusionNot assessed
Table 10 zero-shot and Pile validation evaluation
exp-table10-official-checkpoint-evaluation
Cancelled
Next step
The task was cancelled and must be started again to continue.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Cancelled
Run attempt
No target-level attempt recorded
CAP evidence
No linked immutable Artifact yet
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
The Based and ReBased kernel-approximation variants do not improve over the corresponding simple activation or normalization baselines in the reported 130M and 380M ablations. The paper notes that the negative result may reflect the difference between SSD's 1-semiseparable mask and vanilla linear-attention methods, and that the expanded-feature variants use smaller B/C projections and adjusted layer counts.
Information insufficientclaim-based-rebased-ablation0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Table 5 reports the 360M, 7B-token head-structure ablation. The MIS/MVA pattern again has the lowest reported perplexity, while MCS/MQA and MES/MKA are worse despite the same stated total-state comparison.
Information insufficientclaim-head-structure-360m0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
At 2.7B scale, trained for 300B Pile tokens with 64 layers and matched parameters, adding six attention layers to Mamba-2 improves the reported Pile and downstream results over pure Mamba-2 and Transformer++. Adding MLP layers alone reduces the reported quality, while the combined SSD/MLP/attention model remains competitive. The paper reports that MLP layers may improve hardware efficiency or ease conversion to mixture-of-experts models.
Information insufficientclaim-hybrid-ssd-mlp-attention0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Table 5 reports the 125M, 2.5B-token head-structure ablation. Multi-input/Multi-value (MIS/MVA) has the lowest reported perplexity among the listed patterns; MQA/MCS and MKA/MES are substantially worse despite equal total state size. Parameter-matched multi-head patterns fall between these choices.
Information insufficientclaim-head-structure-125m0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Table 10 reports the full zero-shot evaluation row for Hybrid H3-2.7B.
Official implementationclaim-table10-hybrid-h3-2-7b1 plan0 runs
Legacy task without target-level resource requirements
Table 10 reports the full zero-shot evaluation row for RWKV4-3B.
Official implementationclaim-table10-rwkv4-3b1 plan0 runs
Scientific conclusionNot assessed
Table 10 zero-shot and Pile validation evaluation
exp-table10-official-checkpoint-evaluation
Cancelled
Next step
The task was cancelled and must be started again to continue.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Cancelled
Run attempt
No target-level attempt recorded
CAP evidence
No linked immutable Artifact yet
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Table 10 reports the full zero-shot evaluation row for Mamba-2-2.7B.
Official implementationclaim-table10-mamba2-2-7b1 plan0 runs
Scientific conclusionNot assessed
Table 10 zero-shot and Pile validation evaluation
exp-table10-official-checkpoint-evaluation
Cancelled
Next step
The task was cancelled and must be started again to continue.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Cancelled
Run attempt
No target-level attempt recorded
CAP evidence
No linked immutable Artifact yet
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Table 10 reports the full zero-shot evaluation row for Mamba-2.8B.
Official implementationclaim-table10-mamba-2-8b1 plan0 runs
Scientific conclusionNot assessed
Table 10 zero-shot and Pile validation evaluation
exp-table10-official-checkpoint-evaluation
Cancelled
Next step
The task was cancelled and must be started again to continue.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Cancelled
Run attempt
No target-level attempt recorded
CAP evidence
No linked immutable Artifact yet
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
The kernel-activation ablation finds no consistent improvement from the listed kernel approximations over simple pointwise activations. Random Feature Attention is the lowest-perplexity entry in the table, while Positive Random Features (Performer) is the highest; the paper retains Swish as the default and notes that removing the activation may be simpler but was not extensively tested.
Information insufficientclaim-kernel-activation-ablation0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Table 10 reports the full zero-shot evaluation row for Pythia-6.9B.
Official implementationclaim-table10-pythia-6-9b1 plan0 runs
Scientific conclusionNot assessed
Table 10 zero-shot and Pile validation evaluation
exp-table10-official-checkpoint-evaluation
Cancelled
Next step
The task was cancelled and must be started again to continue.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Cancelled
Run attempt
No target-level attempt recorded
CAP evidence
No linked immutable Artifact yet
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Table 10 reports the full zero-shot evaluation row for Pythia-2.8B.
Official implementationclaim-table10-pythia-2-8b1 plan0 runs
Scientific conclusionNot assessed
Table 10 zero-shot and Pile validation evaluation
exp-table10-official-checkpoint-evaluation
Cancelled
Next step
The task was cancelled and must be started again to continue.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Cancelled
Run attempt
No target-level attempt recorded
CAP evidence
No linked immutable Artifact yet
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
In the Chinchilla-style Pile scaling-law plots, Mamba-2 matches or exceeds Mamba and the Transformer++ recipe across approximately 125M to 1.3B parameters. Relative to the authors' Transformer baseline, the paper describes Mamba-2 as Pareto dominant in perplexity, theoretical FLOPs, and actual wall-clock time.
Information insufficientclaim-scaling-plot0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
In the Mamba-block ablation with SSD as the inner sequence-mixing layer, parallel ABCX projections use fewer parameters and slightly lower perplexity than sequential projections, while the extra normalization used in Mamba-2 slightly improves perplexity. The paper also reports preliminary larger-scale evidence that the extra normalization improves training stability.
Information insufficientclaim-block-ablation0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Table 10 reports the full zero-shot evaluation row for OPT-2.7B.
Official implementationclaim-table10-opt-2-7b1 plan0 runs
Scientific conclusionNot assessed
Table 10 zero-shot and Pile validation evaluation
exp-table10-official-checkpoint-evaluation
Cancelled
Next step
The task was cancelled and must be started again to continue.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Cancelled
Run attempt
No target-level attempt recorded
CAP evidence
No linked immutable Artifact yet
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Table 10 reports the full zero-shot evaluation row for Hybrid H3-1.3B.
Official implementationclaim-table10-hybrid-h3-1-3b1 plan0 runs
Legacy task without target-level resource requirements
Appendix Table 9 reports the model sizes, architectures, optimization settings, and token budgets used for the scaling-law experiments.
Information insufficientclaim-scaling-configurations0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
In a 350M, 48-layer Pile model trained for 7B tokens with the GPT-2 tokenizer, adding a small number of attention blocks to Mamba-2 improves validation perplexity. The best reported result is around a 10% attention-layer ratio; the paper says SSD and attention are complementary and that exact spacing is not very important in these small-scale experiments.
Information insufficientclaim-hybrid-ssd-attention0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Table 10 reports the full zero-shot evaluation row for GPT-J-6B.
Official implementationclaim-table10-gpt-j-6b1 plan0 runs
Scientific conclusionNot assessed
Table 10 zero-shot and Pile validation evaluation
exp-table10-official-checkpoint-evaluation
Cancelled
Next step
The task was cancelled and must be started again to continue.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Cancelled
Run attempt
No target-level attempt recorded
CAP evidence
No linked immutable Artifact yet
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
Table 10 reports the full zero-shot evaluation row for RWKV4-7.4B.
Official implementationclaim-table10-rwkv4-7-4b1 plan0 runs
Scientific conclusionNot assessed
Table 10 zero-shot and Pile validation evaluation
exp-table10-official-checkpoint-evaluation
Cancelled
Next step
The task was cancelled and must be started again to continue.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Cancelled
Run attempt
No target-level attempt recorded
CAP evidence
No linked immutable Artifact yet
Scientific judgment
No Assessment yet
Legacy task without target-level resource requirements
On the harder multi-query associative recall (MQAR) task, Mamba-1 struggles while Mamba-2 performs well across the tested settings. Mamba-2 is reported to be significantly better than Mamba-1 even when state size is controlled at N=16, and increasing Mamba-2 state size from N=16 to N=64 and N=256 consistently improves performance. Standard multi-head softmax attention and Based are included as baselines.
Information insufficientclaim-mqar0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
The paper reports that its SSD implementation is 2–8× faster than Mamba's fused associative scan for large state expansion, is faster than FlashAttention-2 at sequence lengths of 2K and above, and is 6× faster than FlashAttention-2 at sequence length 16K. At sequence length 4K, increasing state expansion slows the optimized Mamba scan approximately linearly, whereas SSD shows little slowdown.
Official implementationclaim-ssd-speed1 plan2 runs
Scientific conclusionInconclusive
SSD, fused Mamba scan, and FlashAttention-2 speed comparison
exp-ssd-speed-cross-hardware5 attempts
Evidence published
Next step
No automatic retry; a person must decide what to do next.
Unclassified execution failure·legacy.unknown
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
2 immutable Artifacts
Scientific judgment
Assessed but inconclusive
Legacy task without target-level resource requirements
The paper reports that the modified Mamba-2 block is tensor-parallel friendly and reduces synchronization points per block by half. It also describes sequence/context parallelism in which devices process contiguous sequence chunks and pass recurrent states between workers, with communication linear in the number of workers.
Official implementationclaim-systems-parallelism1 plan1 run
Scientific conclusionInconclusive
Tensor-parallel synchronization-point reduction from the fixed source
exp-tensor-parallel-sync-source-check3 attempts
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
Assessed but inconclusive
Legacy task without target-level resource requirements
Table 10 reports the full zero-shot evaluation row for Hybrid H3-130M.
Official implementationclaim-table10-hybrid-h3-130m1 plan1 run