The latest reproduction results, one card per paper, ordered by the most recent Assessment. Skim the outcome here; open the paper's claims tab for the full evidence.
8 papers
The paper's Table 2 Citeseer claim states that on the public split SGC reaches 71.9% mean test accuracy, about one point above the authors' GCN result, alongside all listed baselines. One measurement of this composite claim was executed: the 100-epoch, public-split SGC arm gave 71.2% versus a reported 71.9% (0.7 percentage points lower). The 150-epoch sensitivity arm averaged 71.85% but was not substituted. The GCN comparison and all other Table 2 baselines remain unassessed, so the overall claim is inconclusive. Separately, the citation setup claim is supported: 100 epochs, learning rate 0.2, and 60 hyperopt iterations were each reproduced exactly. The 15 unassessed claims, including claim-citation-setup's sibling claims, are not run failures.
1/ 27 reproduced
The paper introduces SGVB and AEVB for variational inference with per-datapoint continuous latent variables, and reports experiments on MNIST and Frey Face, including Figure 3 marginal-likelihood estimates using the first 1,000 datapoints, 50 posterior samples per datapoint, 4 HMC leapfrog steps, and an MCEM baseline with 10 leapfrog steps, 5 weight updates, and a 90% acceptance target. Our reconstruction matched these operation counts: 1,000 datapoints, 50 samples, 4 steps, 10 steps, 5 updates, and 90% acceptance all matched. However, the reported estimator variance below 1 could not be comparably reproduced: observed values were 866.562 (Frey) and 0.033282 (MNIST) under a different variance definition, so the claim remains inconclusive.
2/ 13 reproduced
The paper reports an MCP census from a 2026-07-27 official-registry snapshot, retaining the latest version of each distinct server, querying tools/list anonymously on remote targets without tool calls, yielding 98,291 tools. Our offline reanalysis reproduced the artifact-derived counts: 59,625 entries, 18,688 distinct servers, 9,454 records without a remote endpoint, 9,234 captured remote-target records, 4,838 successful nonempty targets, 4,318 connection failures, 74 timeouts, 4 no-tool targets, 0 HTTP auth rejections, and a median of 11 tools per responding server. The overall claim remains inconclusive: measurement-tool-calls was unassessed, and tool-invocation absence, snapshot date, and registry completeness were not independently verified.
0/ 15 reproduced
FlashAttention-2 claims H100 80GB SXM5 forward-plus-backward attention throughput in Figure 7a, up to 296 TFLOPs/s, with PyTorch OOM at 16k. The run measured selected Figure 7a cells on an NVIDIA L4 instead of H100, so the reported values cannot be compared: e.g. FlashAttention-2 at 16k reported 296 TFLOPs/s versus observed 58.00092361049494, and PyTorch 512 reported 62 versus 5.800733954913403. The PyTorch 16k OOM claim was not executed. Matched throughput measurements are inconclusive; other claims were not assessed and remain unverified.
0/ 27 reproduced
This paper claims EGD, a gradient-whitening method, accelerates grokking across toy, parity, modular, MNIST, and transformer settings, with efficiency claims tied to exact-SVD and RSVD variants. Only the three efficiency claims were assessed, all inconclusive. For modular multiplication (c11) and addition (c10), most epoch counts reproduced closely, e.g. p=79 exact SVD 228 observed versus 221 reported and rank-128 RSVD 222 versus 233, reversing the reported ordering; p=97 exact SVD 118 versus 107. No wall-clock measurements were executed, so timing portions remain unverified. Sparse parity (c12) results are inconclusive because the rerun used squared loss while the paper describes hinge loss.
0/ 21 reproduced
The paper claims Table 10 reports full zero-shot evaluation rows for many models, and that its SSD implementation is 2–8× faster than Mamba's fused associative scan, beats FlashAttention-2 at 2K and above, and is 6× faster at 16K. Reproduction is partial: for the speed claim, observed values were 1.104–1.840× over Mamba, crossover at 16384 tokens (reported 2048), and 1.279× at 16K (reported 6×), measured on L4 rather than A100; the assessment is inconclusive. Table 10 rows and the remaining claims are not assessed here.
0/ 41 reproduced
The paper claims that teacher predictive entropy distinguishes procedural from knowledge-intensive domains with high ROC AUC across training stages and model sizes. Our reproduction evaluated only the 1B checkpoints on a reconstructed 240-document benchmark; the 7B and 13B measurements were omitted due to size/budget limits. For the evaluated 1B results, observed ROC AUC values (e.g., 0.8079 vs reported 0.815) were broadly consistent, but the claim remains inconclusive due to incomplete multi-model and multi-stage coverage.
0/ 76 reproduced
The paper introduces LongPIBench and reports key claims: prevention defenses show high attack success on long contexts, detection shows extreme trade-offs, GCG variants achieve high success, scope is limited to static workflows, and heuristic attacks, especially Authority spoof, substantially raise success. Our reproduction, currently inconclusive, allowed only a directional check of heuristics on document tasks: reported 1 vs observed 1 fraction attack success rate, with limitations from approximated model, corpus, and generation preventing strict matching. Other claims remain not assessed.
0/ 4 reproduced