来源:paper_markdown:Figure 10 and Section 9.3, PDF pages 29–31; introduction, PDF pages 2–3
openai/gpt-5.6-luna-pro
9f680b3308
The allocated device was NVIDIA L4 rather than the paper's NVIDIA A100 80GB PCIe, so the measured range, crossover, and 16K ratio are not a hardware-faithful confirmation of the reported A100 values. · The paper's qualitative state-expansion phrase does not define a numerical slowdown threshold, baseline operand, or aggregation rule; m-speed-state-expansion remains unresolved while the complete matched 4K curve is retained. · The image lacked nvcc/CUDA_HOME. The source-build path was replaced by official import-tested CUDA wheels matching the fixed Mamba source version and installed PyTorch ABI; the exact wheel identities and one-forward path checks are retained separately. · Unresolved metric interpretation: m-speed-state-expansion: The paper says SSD can handle much larger state expansion factors without much slowdown, but does not define a numerical slowdown tolerance or a scalar aggregation rule. The corrected GPU execution retains matched SSD and optimized-Mamba timings at the Figure 10 state sizes 16, 32, 64, 128, and 256 for the 4K input, together with the tested multipliers relative to the declared Mamba baseline. A scalar cannot be selected source-faithfully until the qualitative threshold and baseline rule are resolved. · Unresolved measurements: m-speed-state-expansion: The preserved parser evidence did not contain a trustworthy metric: JSON 指针没有指向有限数值 · The executed accelerator was an NVIDIA L4, not the paper's NVIDIA A100 80GB PCIe. · The exact paper batch size, dtype, kernel versions, and precise FlashAttention-2 configuration were not tabulated; execution used batch 1, float16, Mamba 2.3.2.post1, and FlashAttention-2 2.8.3.post1. · Only 1K, 2K, 4K, and 16K were timed, so the broader sequence-length ordering was not exhaustively tested. · The retained raw evidence is partially truncated in the supplied context, although the captured execution records establish completion and the aggregate values. · The paper's qualitative state-expansion statement has no numerical slowdown threshold or scalar aggregation rule, so it cannot be converted into a comparable scalar judgment here. · The synchronized benchmark was completed on an NVIDIA L4, whereas the paper's reported results use an NVIDIA A100 80GB PCIe. The L4 results were 1.104–1.840× over Mamba, crossed FlashAttention-2 only at the tested 16K point, and measured 1.279× at 16K, differing substantially from the reported 2–8× range, 2K crossover, and 6× result. The recorded protocol explicitly identifies this hardware change as unable to verify the A100 claims. Thus the execution establishes L4 behavior but neither reproduces nor decisively contradicts the paper's reference-hardware results. The qualitative state-expansion clause also lacks a source-defined scalar threshold and is not adjudicated by the four requested measurements. Some original claim measurements remain without evidence; available measurements are assessed individually.
citeark-verification-engine
fb9bbdb5b0
The five assigned speed measurements remain unobserved until the GPU phase completes. · A clean package-root import is blocked by the unrelated Mamba-3 TileLang/TVM-FFI compatibility error; the harness bypasses that initializer with a local namespace shim and records the failure. · The paper's qualitative minimal-slowdown state-expansion phrase has no source-defined numeric threshold or aggregation rule, so it remains unresolved even after the GPU curve is collected. · Unresolved metric interpretation: m-speed-state-expansion: The paper says SSD supports 8x or larger state size with minimal slowdown, but does not define a numerical slowdown tolerance, baseline latency operand, or aggregation rule for the qualitative phrase. GPU execution must retain the complete N-versus-latency curve and report the tested multipliers; a scalar cannot be selected source-faithfully until the threshold is resolved. · Unresolved measurements: m-speed-mamba-min: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-mamba-max: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-fa2-crossover: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-fa2-16k: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-state-expansion: The submitted result did not promote parser evidence at ssd-speed-raw-results.json · The execution was partial during measurement because of metric_unavailable: CPU preparation completed and handed off for GPU execution. The staged FlashAttention-2 artifact, task-local dependencies, fixed Mamba module import path, CUDA-guarded benchmark harness, and GPU setup/probe/full-run command are ready. No scientific timing result was collected because the assigned CPU allocation has no CUDA device. The five assigned speed measurements remain unobserved until the GPU phase completes. A clean package-root import is blocked by the unrelated Mamba-3 TileLang/TVM-FFI compatibility error; the harness bypasses that initializer with a local namespace shim and records the failure. The paper's qualitative minimal-slowdown state-expansion phrase has no source-defined numeric threshold or aggregation rule, so it remains unresolved even after the GPU curve is collected. Unresolved metric interpretation: m-speed-state-expansion: The paper says SSD supports 8x or larger state size with minimal slowdown, but does not define a numerical slowdown tolerance, baseline latency operand, or aggregation rule for the qualitative phrase. GPU execution must retain the complete N-versus-latency curve and report the tested multipliers; a scalar cannot be selected source-faithfully until the threshold is resolved. Unresolved measurements: m-speed-mamba-min: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-mamba-max: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-fa2-crossover: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-fa2-16k: The submitted result did not promote parser evidence at ssd-speed-raw-results.json; m-speed-state-expansion: The submitted result did not promote parser evidence at ssd-speed-raw-results.json. This outcome is not retryable without changing the available inputs or conditions The submitted result did not promote parser evidence at s
已评估 5/5 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 2 × | 1.104 ×2 次评估 · 展开查看 | 无法判定 | |
| 8 × | 1.8397 ×2 次评估 · 展开查看 | 无法判定 | |
| 2,048 tokens | 16,384 tokens2 次评估 · 展开查看 | 无法判定 | |
| 6 × | 1.2789 ×2 次评估 · 展开查看 | 无法判定 | |
| 8 × baseline state size | —2 次评估 · 展开查看 | 无法判定 |