Start with coverage and measured results; open a run only when you need evidence or technical details.
RunMatchRepeat
0/21
claims supported by evidence
0
Supported
0
Challenged or mixed
0
Contradicted
0
Inconclusive
21
Not assessed
Run history
Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.
No reproduction runs yet
Runs will appear here as the paper's experiment plans are executed.
Claim–experiment reproduction matrix
See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions. There are also 7 claims with no independent reproduction scheduled in this plan.
A successful execution does not by itself validate a paper claim
Target state says whether the platform completed the work. The scientific conclusion is determined only by immutable evidence and Assessments. Resource shortages and platform failures are never presented as scientific contradictions.
This task predates target-level records. The states below are conservative projections; no attempts, resource decisions, or evidence links are invented.
PaperDoctor findings were longer on average than referee or Stanford Agentic Reviewer feedback, with its average split into 33.9 evidence words and 16.1 suggestion words.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
All 30 authors gave PaperDoctor's overall review a positive score: 70% called it somewhat helpful and 30% very helpful, yielding a mean holistic score of +1.30.
CiteArk reconstructionclaim-study-0011 plan0 runs
Scientific conclusionNot assessed
Recompute the 30-paper author-rating study
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
In the same tumor-modeling paper, PaperDoctor reported that a novelty framing about DeepONet and Fourier Neural Operators being designed for homogeneous constant-coefficient PDEs was contradicted by the cited original papers, which include varying conditions or variable coefficients.
Information insufficientclaim-case-0050 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Authors accepted the evidence for 70.6% of all 1,299 issues; paper-level acceptance varied widely, and its positive correlation with the holistic helpfulness score was not statistically significant.
CiteArk reconstructionclaim-study-0021 plan0 runs
Scientific conclusionNot assessed
Recompute the 30-paper author-rating study
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
In an Agents4Science case study, the paper claimed SciBERT embeddings with 768 dimensions, but all three relevant released files configured all-MiniLM-L6-v2 and produced 384-dimensional vectors; SciBERT was not imported anywhere.
CiteArk reconstructionclaim-case-0031 plan0 runs
Scientific conclusionNot assessed
Audit the claimed and configured embedding model
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Experiment reproduction was bottlenecked before and after execution: only 53.0% of reproduction plans reached a command, and among plans that ran, 27.0% were warnings and 34.5% errors.
CiteArk reconstructionclaim-repro-0011 plan0 runs
Scientific conclusionNot assessed
Recompute the 40-paper reproduction-flow analysis
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
The 24 optional written comments praised breadth most often but also criticized it most often; proposed fixes, novelty/framing silence, and false positives were the only aspects with more criticism than praise.
CiteArk reconstructionclaim-study-0051 plan0 runs
Scientific conclusionNot assessed
Recompute the optional-comment aspect analysis
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
The four paper groups showed different reproduction failure locations: ICML plans usually reached execution but rarely passed, Agents4Science plans usually stopped before execution but often passed when they ran, and the two Nature groups were intermediate.
CiteArk reconstructionclaim-repro-0041 plan0 runs
Scientific conclusionNot assessed
Recompute the 40-paper reproduction-flow analysis
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
PaperDoctor's L1 check frequencies were broadly similar across paper groups, while L2 Code findings were highest for AI-written papers, Theory findings were highest for ICML, and Experiment Design findings separated the groups most strongly.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
PaperDoctor distributed its feedback more evenly across manuscript, literature, and code dimensions than referees or the Stanford Agentic Reviewer, both of which concentrated on experiments in the main body.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
PaperDoctor attached both evidence and a suggestion to every finding by construction; referees did so for 35.9% of findings and the Stanford Agentic Reviewer for 1.5%, whose dominant pattern was suggestion without evidence.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Authors were less uncertain but also less likely to agree with findings PaperDoctor labeled Error than with Warnings; acceptance of Evidence and Suggestion was also strongly coupled.
CiteArk reconstructionclaim-study-0041 plan0 runs
Scientific conclusionNot assessed
Recompute the 30-paper author-rating study
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Among plans that ran, high-priority plans passed more often and errored less often than medium- or low-priority plans; medium-priority pass rates were close to low-priority rates rather than intermediate.
CiteArk reconstructionclaim-repro-0021 plan0 runs
Scientific conclusionNot assessed
Recompute the 40-paper reproduction-flow analysis
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
In an Agents4Science case study, PaperDoctor found that a bibliography entry dated the GPT-4.5 research preview to 2024 even though the release was on February 27, 2025, making the citation metadata chronologically inconsistent with the late-2024 manuscript.
CiteArk reconstructionclaim-case-0021 plan0 runs
Scientific conclusionNot assessed
Verify the GPT-4.5 citation and release year
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Agreement differed sharply by check: L2 claim-verification items were rejected less often than L1 surface-screening items; Experiment Review had the highest Evidence acceptance and Figure Review the lowest.
CiteArk reconstructionclaim-study-0031 plan0 runs
Scientific conclusionNot assessed
Recompute the 30-paper author-rating study
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Across the same 40 papers, PaperDoctor generally produced more atomic findings per reviewer and paper than published referee reports or the Stanford Agentic Reviewer, with larger within-domain spread.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Among executed plans, training reruns passed least often and errored most often, inference/evaluation was intermediate, and statistical-test or analysis reruns passed most often and errored least often.
CiteArk reconstructionclaim-repro-0031 plan0 runs
Scientific conclusionNot assessed
Recompute the 40-paper reproduction-flow analysis
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
The interface example flags a partial paper-code mismatch: the paper says replay occurs every 10 steps for smaller models (and every 5 for FLAN-T5-3B), while the default configuration sets replay_freq to 1; replay_k does match the stated mini-batch size of 8.
CiteArk reconstructionclaim-case-0011 plan0 runs
Scientific conclusionNot assessed
Audit replay batch and interval configuration
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
In an Agents4Science theory case, PaperDoctor found an approximation bound stated without proof, assumptions, an appendix pointer, or definitions sufficient to interpret and verify it.
CiteArk reconstructionclaim-case-0041 plan0 runs
Scientific conclusionNot assessed
Audit definitions and support for the approximation bound
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
For a Nature Communications microscopy paper, rebuilding the shipped Figure 2d analysis yielded soft-focal contact-task recalls of 0.43 and 0.49, corresponding to a 1.14x best-pair ratio and 1.04x mean ratio rather than the prose claim of almost 2x.
CiteArk reconstructionclaim-case-0061 plan0 runs
Scientific conclusionNot assessed
Recompute the Figure 2d recall comparison
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Reasons plans never ran differed by paper group: missing code dominated SocialScience, wet-lab procedures dominated NatureScience, and restricted data plus shared blockers dominated Agents4Science; overall, missing code or model weights was the largest blocker class.
CiteArk reconstructionclaim-repro-0051 plan0 runs
Scientific conclusionNot assessed
Recompute the 40-paper reproduction-flow analysis
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Community reproductions
Executed locally and uploaded by users. The platform verifies file signatures; conclusions come from the uploaded runs.