3 experiments
Execution agent’s result notes
CPU preparation completed and handed off for GPU execution. The staged FlashAttention-2 artifact, task-local dependencies, fixed Mamba module import path, CUDA-guarded benchmark harness, and GPU setup/probe/full-run command are ready. No scientific timing result was collected because the assigned CPU allocation has no CUDA device.
Completed the synchronized GPU reproduction sweep on an NVIDIA L4 using the fixed Mamba-2 SSD and optimized Mamba fused-scan paths plus FlashAttention-2 2.8.3.post1. With batch 1, d_model 1024, float16, 10 warmups, 50 CUDA-event iterations, and lengths 1K/2K/4K/16K, the source-conditioned 2K/4K/16K SSD-over-Mamba range was 1.104x–1.840x; SSD first beat FlashAttention-2 at the tested 16K point; and the 16K SSD-over-FlashAttention-2 ratio was 1.279x. The 4K state sweep retained matched SSD and Mamba timings for state sizes 16, 32, 64, 128, and 256, showing lower SSD growth, but the paper defines no numerical minimal-slowdown threshold or scalar aggregation rule. These are L4 cross-hardware observations and do not establish the paper's NVIDIA A100 80GB PCIe numeric claims.
A representative probe and corrected full source analysis were run through the research-plan helper. The fixed paper source defines two baseline tensor-parallel all-reduces per block; the fixed Mamba-2 code exposes one final reduction on each mutually exclusive forward path (all_reduce normally, reduce_scatter with sequence parallelism). The retained computation is 1 / 2 = 0.5 fraction of baseline, matching the assigned half-reduction claim under the stated source-level forward TP assumptions.
Completed the assigned low-scope evaluation on all seven full task splits with the released danfu09/H3-125M checkpoint (revision 38030b27df2a9814aef9df42ac2d12db46f814af), using the official H3 source path, GPT-2 tokenizer, lm-eval 0.4.2, zero-shot evaluation, batch size 2, and CPU execution. The measured LAMBADA perplexity was 30.0757, LAMBADA accuracy 34.3683%, HellaSwag 29.3268%, PIQA 64.2546%, ARC-E 44.4024%, ARC-C 20.8191%, WinoGrande 50.5919%, OpenBookQA 17.8%, and the equal-weight seven-task accuracy average 37.3662%. These are fresh evaluator outputs bound to ordinary acc,none and perplexity,none fields; they differ from the paper's H3-130M row and therefore provide a completed but non-confirmatory reproduction under the recorded artifact and metric conditions.
Evaluated the released state-spaces/mamba-130m checkpoint with the repository lm-evaluation-harness wrapper on the complete LAMBADA, HellaSwag, PIQA, ARC-E, ARC-C, WinoGrande, and OpenBookQA task sets. CPU float32 execution with max length 2048 and batch size 8 produced LAMBADA perplexity 16.0657, LAMBADA accuracy 44.2461%, HellaSwag 30.8006%, PIQA 64.4723%, ARC-E 47.9798%, ARC-C 19.7099%, WinoGrande 52.0916%, OpenBookQA 16.6000%, and an unweighted seven-task average of 39.4143%. The LAMBADA perplexity and PIQA/ARC-E/WinoGrande values are close to the paper's displayed row, while HellaSwag, ARC-C, OpenBookQA, and the resulting average differ materially; these are observed evaluation results, not copied paper values. Exact Pile validation perplexity was not measured because the required validation input could not be acquired, so the assigned reproduction is partial.
Evaluated the released EleutherAI/pythia-160m checkpoint with the repository lm-eval 0.4.2 wrapper on CPU, using the GPT-NeoX tokenizer, zero-shot settings, batch size 32, one run, and the seven Table 10 test tasks. Fresh results were LAMBADA perplexity 38.09990492720415, LAMBADA accuracy 32.71880457985639%, HellaSwag 28.370842461661024%, PIQA 61.53427638737759%, ARC-E 43.686868686868685%, ARC-C 18.515358361774744%, WinoGrande 51.69692186266772%, and OpenBookQA 16%; the recomputed unweighted seven-task average was 36.07472462002945%. Exact Pile validation perplexity could not be evaluated because the verified EleutherAI/pile validation URL returned HTTP 404, so the reproduction is partial.
The official state-spaces/mamba2-130m checkpoint was evaluated with lm-eval 0.4.2 on the seven Table 10 downstream tasks using the repository's CPU reference path. Full-split results were LAMBADA ppl 16.7877, LAMBADA acc 43.9744%, HellaSwag 30.9799%, PIQA 64.9075%, ARC-E 47.3906%, ARC-C 20.9044%, WinoGrande 52.5651%, OpenBookQA 18.0000%, and equal-weight accuracy average 39.8174%. Pile validation was not run because the exact validation data/evaluator binding was unavailable.
Evaluated the released H3-355M checkpoint (SHA-256 e273ef09b6cb1984b5fc5f3e07a7acdeeab31641e277c60bf40e3ac86f247ea8; official H3 source commit 5c4d06b5795405170387c80998b58d76179a8a1a) on all seven Table 10 zero-shot tasks using GPT-2 tokenization, batch size 1, maximum length 2048, and one full evaluation per task on an NVIDIA L4. The observed results were LAMBADA perplexity 12.5790783 and accuracy 48.0303%, HellaSwag 41.4957%, PIQA 68.3351%, ARC-E 51.3468%, ARC-C 24.7440%, WinoGrande 54.2226%, and OpenBookQA 31.6000%. The equally weighted average recomputed from the seven retained selected task fields was 45.6821%. These are fresh numerical evaluations of the released checkpoint; the paper's Table 10 values were available before execution, and the observed values are close to the displayed row values but were obtained under the conditions below.
Evaluated the released EleutherAI/pythia-410m checkpoint with its NeoX tokenizer through the official lm-eval wrapper on CUDA/NVIDIA L4, using lm-eval 0.4.2, zero-shot evaluation, batch size 8, no limit, and the resolved source splits. Fresh selected ordinary-acc results were LAMBADA 51.446%, HellaSwag 33.718%, PIQA 67.084%, ARC-E 51.936%, ARC-C 21.246%, WinoGrande 53.433%, and OpenBookQA 18.000%, giving an equal-weight seven-task average of 42.409%; LAMBADA perplexity was 10.8461. The neighboring normalized accuracy fields were retained in the raw evidence but were not used because Table 10 labels the metric acc. Pile validation perplexity could not be independently evaluated, so the assigned reproduction is partial.
Evaluated the released state-spaces/mamba-370m checkpoint at Hub revision b6c47221dc4908532cc9773d469d6b8cbe3f0762 with the GPT-NeoX tokenizer, lm-eval 0.4.2, float16 CUDA inference on one NVIDIA L4, default seeds, and the complete seven-task Table 10 scope. LAMBADA perplexity was 8.1375; the source-resolved accuracy fields were LAMBADA 55.7151%, HellaSwag normalized 46.4748%, PIQA 69.4777%, ARC-E ordinary 54.8822%, ARC-C normalized 27.8157%, and WinoGrande 55.4065%; these round to the displayed source row values. Pile validation perplexity was not measured, and OpenBookQA plus its dependent average remain unresolved because the paper does not specify ordinary versus sequence-length-normalized accuracy.
Evaluated the released state-spaces/mamba2-370m checkpoint on the complete seven-task zero-shot workload with the NeoX tokenizer. LAMBADA perplexity (7.982040778406215), LAMBADA accuracy (55.753929749660394%), and WinoGrande accuracy (55.643251775848455%) are source-resolved. HellaSwag, PIQA, ARC-E, ARC-C, and OpenBookQA each have retained ordinary and length-normalized accuracy values, but the paper and fixed repository do not identify which field corresponds to the Table 10 acc columns; the dependent seven-task average is consequently unresolved. Canonical Pile validation data was unavailable.
Evaluated the released EleutherAI/pythia-1b checkpoint at revision main with its NeoX tokenizer through the repository's lm-eval==0.4.2 harness. The complete seven-task zero-shot workload ran once with no limit, batch size 8, CUDA on an NVIDIA L4, and PIQA custom-code loading enabled via HF_DATASETS_TRUST_REMOTE_CODE=1. The retained raw output contains ordinary and length-normalized neighboring fields. Using the paper-compatible task fields (LAMBADA acc, HellaSwag acc_norm, PIQA acc, ARC-E acc, ARC-C acc_norm, WinoGrande acc, OpenBookQA acc), the observed values were 56.2003%, 47.1221%, 70.7835%, 57.0286%, 26.8771%, 53.5122%, and 31.4%, respectively; their unweighted mean was 48.9891%. LAMBADA perplexity was 7.9167. These downstream observations round to the displayed Table 10 values, while the Pile validation perplexity was not evaluated.
A full zero-shot lm-eval 0.4.2 evaluation of the released state-spaces/mamba-790m checkpoint completed on an NVIDIA L4 with the NeoX tokenizer, batch size 1, and one repetition. Fresh results were LAMBADA PPL 6.0166589040205665, LAMBADA accuracy 61.45934407141471%, HellaSwag 42.30233021310496%, PIQA 72.19804134929271%, ARC-E 61.15319865319865%, ARC-C 26.535836177474405%, WinoGrande 55.406471981057614%, OpenBookQA 22.6%, and a recomputed seven-task average of 48.80788892079187%. Exact Pile validation perplexity was not established.
GPU evaluation completed on one NVIDIA L4 using the pinned state-spaces/mamba2-780m checkpoint (revision 2f1ce3195cbfbf0b2f0c01fdf8856b6dbe4ac10e; SHA-256 a210e4df8109cde42c578c51576600e828b4a61205df16edeff300c308fa568) and the EleutherAI/gpt-neox-20b tokenizer. The full zero-shot Table 10 task set ran on the configured splits with lm-eval 0.4.2: LAMBADA perplexity 5.8534176265, LAMBADA accuracy 61.5564%, HellaSwag 42.3820%, PIQA 72.0892%, ARC-E 61.0690%, ARC-C 26.4505%, WinoGrande 60.2210%, OpenBookQA 22.6000%, and an equal-weight seven-task average of 49.4812%. The raw run and derived operands are retained. Compared with the displayed Table 10 row, the fresh ordinary-acc results are close for LAMBADA, PIQA, ARC-E, and WinoGrande but materially lower for HellaSwag, ARC-C, OpenBookQA, and the average; neighboring acc_norm fields were retained but not selected by proximity. Pile validation perplexity was not reported because the paper and official repository do not define the required e
Completed the assigned OPT-1.3B Table 10 zero-shot evaluation on one NVIDIA L4 using the verified facebook/opt-1.3b checkpoint and OPT tokenizer, CUDA float16, batch size 8, zero shots, and complete harness splits for LAMBADA, HellaSwag, PIQA, ARC-E, ARC-C, WinoGrande, and OpenBookQA. The selected ordinary-acc results were LAMBADA perplexity 6.6403, LAMBADA accuracy 57.8692%, HellaSwag 41.5356%, PIQA 71.7084%, ARC-E 57.1128%, ARC-C 23.2935%, WinoGrande 59.4317%, OpenBookQA 23.4000%, and equal-weight average 47.7645%. Relative to the paper's OPT-1.3B row, LAMBADA and WinoGrande are close at displayed precision, while HellaSwag, ARC-C, OpenBookQA, and the derived average differ materially under the selected ordinary-acc interpretation.
The full GPU evaluation completed successfully on the released EleutherAI/pythia-1.4b checkpoint (revision fedc38a16eea3bd36a96b906d78d11d2ce18ed79; model.safetensors SHA-256 ad4388632e922d0c58c23cb315292d516deb70af0573f9c1543ce84158bff59c) using an NVIDIA L4, CUDA, float16, zero-shot lm-eval 0.4.2, and auto batch size resolved to 64. Fresh ordinary-acc results were LAMBADA 61.5952%, HellaSwag 40.3804%, PIQA 70.9467%, ARC-E 60.4377%, ARC-C 26.0239%, WinoGrande 57.6953%, and OpenBookQA 22.2%, with LAMBADA perplexity 6.08529 and an equal-weight seven-task average of 48.4685%. The paper's full Pile validation perplexity could not be established from the available evaluator and remains unresolved, so the complete Table 10 row is only partially reproduced.
Completed the assigned RWKV4-1.5B Table 10 GPU evaluation on an NVIDIA L4 using the pinned RWKV/rwkv-4-1b5-pile checkpoint, GPT-NeoX tokenizer, lm-evaluation-harness 0.4.2, seven zero-shot tasks, batch size 8, and seed 0,1234,1234. Fresh results were LAMBADA perplexity 7.081011983835502, LAMBADA accuracy 57.18998641568018%, HellaSwag 41.127265484963154%, PIQA 72.0348204570185%, ARC-E 60.81649831649831%, ARC-C 24.914675767918087%, WinoGrande 55.327545382794%, OpenBookQA 23.599999999999998%, and equal-weight average 47.85868454641032%. The seven-task row is evidenced from the completed 67,719-request run; Pile validation perplexity remains unresolved because the fixed sources do not specify the document-boundary versus packed-sequence aggregation policy.
On an NVIDIA L4, the released state-spaces/mamba-1.4b checkpoint with the GPT-NeoX tokenizer completed the seven zero-shot lm-eval 0.4.2 tasks at batch size 1. The retained evaluator output gives LAMBADA perplexity 5.04267677618444, LAMBADA accuracy 64.91364253832718%, HellaSwag 45.0408285202151%, PIQA 74.15669205658324%, ARC-E 65.40404040404042%, ARC-C 29.86348122866894%, WinoGrande 61.16811365430151%, OpenBookQA 26.200000000000003%, and a recomputed equal-weight seven-task accuracy mean of 52.39239977173377%. The exact whole-Pile validation perplexity remains unresolved.
Completed the assigned GPT-Neo 2.7B Table 10 evaluation on the exact released EleutherAI/gpt-neo-2.7B checkpoint at revision e24fa291132763e59f4a5422741b424fb5d59056. Using lm-eval 0.4.2's Hugging Face loader, GPT2-compatible tokenizer, CUDA float16, batch size 1, seed 0, zero-shot evaluation, and all seven requested tasks on one NVIDIA L4, the run exited 0 after about 44 minutes. The extracted values are LAMBADA perplexity 5.6244, LAMBADA accuracy 62.1192%, HellaSwag ordinary accuracy 42.6807%, PIQA 72.0892%, ARC-E 61.0269%, ARC-C 27.5597%, WinoGrande 58.1689%, OpenBookQA 23.4%, and an unweighted seven-task average of 49.5778%. The paper's GPT-Neo 2.7B row is 5.63, 62.2, 55.8, 72.1, 61.1, 30.2, 57.6, 33.2, and 53.2 respectively. The current plan intentionally binds the paper's acc-labeled columns to lm-eval's ordinary acc,none field; normalized alternatives remain in the raw output and were not selected merely to reduce the differences. Thus execution is complete, while exact claim eq
The paper reports that its SSD implementation is 2–8× faster than Mamba's fused associative scan for large state expansion, is faster than FlashAttention-2 at sequence lengths of 2K and above, and is 6× faster than FlashAttention-2 at sequence length 16K. At sequence length 4K, increasing state expansion slows the optimized Mamba scan approximately linearly, whereas SSD shows little slowdown.
Metric: speedup over Mamba fused associative scan (lower bound) · ×
Missing values are not zero.
Data and sources · 3
| Source | Conditions | Metric | Value | Assessment and limits |
|---|---|---|---|---|
| Paperpaper_markdown:Figure 10 caption, PDF page 29claim-ssd-speed/m-speed-mamba-min/paper | — | speedup over Mamba fused associative scan (lower bound) | 2 × | reported |
| Reproductionpaper_markdown:Figure 10 caption, PDF page 29claim-ssd-speed/m-speed-mamba-min/urn:citeark:assessment:5b3d6c49a444fc8cff7f2fa91c10f901d2b0eb38f621cf2e15529f6a6fdd9668 | — | speedup over Mamba fused associative scan (lower bound) | 1.1039711579139013 × | inconclusive |
| Reproductionpaper_markdown:Figure 10 caption, PDF page 29claim-ssd-speed/m-speed-mamba-min/urn:citeark:assessment:30fb2c163ecc8213f91368b1421e4aef87dd3ea7fa76f2698c3022c69a708ff6 | — | speedup over Mamba fused associative scan (lower bound) | — × | inconclusive |
Comparison conditions
The five assigned speed measurements remain unobserved until the GPU phase completes.
A clean package-root import is blocked by the unrelated Mamba-3 TileLang/TVM-FFI compatibility error; the harness bypasses that initializer with a local namespace shim and records the failure.
The paper's qualitative minimal-slowdown state-expansion phrase has no source-defined numeric threshold or aggregation rule, so it remains unresolved even after the GPU curve is collected.
Unresolved metric interpretation: m-speed-state-expansion: The paper says SSD supports 8x or larger state size with minimal slowdown, but does not define a numerical slowdown tolerance, baseline latency operand, or aggregation rule for the qualitative phrase. GPU execution must retain the complete N-versus-latency curve and report the tested multipliers; a scalar cannot be selected source-faithfully until the threshold is resolved.
Unresolved measurements: m-speed-mamba-min: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-mamba-max: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-fa2-crossover: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-fa2-16k: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-state-expansion: The submitted result did not promote parser evidence at ssd-speed-raw-results.json
The execution was partial during measurement because of metric_unavailable: CPU preparation completed and handed off for GPU execution. The staged FlashAttention-2 artifact, task-local dependencies, fixed Mamba module import path, CUDA-guarded benchmark harness, and GPU setup/probe/full-run command are ready. No scientific timing result was collected because the assigned CPU allocation has no CUDA device. The five assigned speed measurements remain unobserved until the GPU phase completes. A clean package-root import is blocked by the unrelated Mamba-3 TileLang/TVM-FFI compatibility error; the harness bypasses that initializer with a local namespace shim and records the failure. The paper's qualitative minimal-slowdown state-expansion phrase has no source-defined numeric threshold or aggregation rule, so it remains unresolved even after the GPU curve is collected. Unresolved metric interpretation: m-speed-state-expansion: The paper says SSD supports 8x or larger state size with minimal slowdown, but does not define a numerical slowdown tolerance, baseline latency operand, or aggregation rule for the qualitative phrase. GPU execution must retain the complete N-versus-latency curve and report the tested multipliers; a scalar cannot be selected source-faithfully until the threshold is resolved. Unresolved measurements: m-speed-mamba-min: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-mamba-max: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-fa2-crossover: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-fa2-16k: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-state-expansion: The submitted result did not promote parser evidence at ssd-speed-raw-results.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at s
The allocated device was NVIDIA L4 rather than the paper's NVIDIA A100 80GB PCIe, so the measured range, crossover, and 16K ratio are not a hardware-faithful confirmation of the reported A100 values.
The paper's qualitative state-expansion phrase does not define a numerical slowdown threshold, baseline operand, or aggregation rule; m-speed-state-expansion remains unresolved while the complete matched 4K curve is retained.
The image lacked nvcc/CUDA_HOME. The source-build path was replaced by official import-tested CUDA wheels matching the fixed Mamba source version and installed PyTorch ABI; the exact wheel identities and one-forward path checks are retained separately.
Unresolved metric interpretation: m-speed-state-expansion: The paper says SSD can handle much larger state expansion factors without much slowdown, but does not define a numerical slowdown tolerance or a scalar aggregation rule. The corrected GPU execution retains matched SSD and optimized-Mamba timings at the Figure 10 state sizes 16, 32, 64, 128, and 256 for the 4K input, together with the tested multipliers relative to the declared Mamba baseline. A scalar cannot be selected source-faithfully until the qualitative threshold and baseline rule are resolved.
Unresolved measurements: m-speed-state-expansion: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值
The executed accelerator was an NVIDIA L4, not the paper's NVIDIA A100 80GB PCIe.
The exact paper batch size, dtype, kernel versions, and precise FlashAttention-2 configuration were not tabulated; execution used batch 1, float16, Mamba 2.3.2.post1, and FlashAttention-2 2.8.3.post1.
Only 1K, 2K, 4K, and 16K were timed, so the broader sequence-length ordering was not exhaustively tested.
The retained raw evidence is partially truncated in the supplied context, although the captured execution records establish completion and the aggregate values.
The paper's qualitative state-expansion statement has no numerical slowdown threshold or scalar aggregation rule, so it cannot be converted into a comparable scalar judgment here.
The synchronized benchmark was completed on an NVIDIA L4, whereas the paper's reported results use an NVIDIA A100 80GB PCIe. The L4 results were 1.104–1.840× over Mamba, crossed FlashAttention-2 only at the tested 16K point, and measured 1.279× at 16K, differing substantially from the reported 2–8× range, 2K crossover, and 6× result. The recorded protocol explicitly identifies this hardware change as unable to verify the A100 claims. Thus the execution establishes L4 behavior but neither reproduces nor decisively contradicts the paper's reference-hardware results. The qualitative state-expansion clause also lacks a source-defined scalar threshold and is not adjudicated by the four requested measurements. Some original claim measurements remain without evidence; available measurements are assessed individually.