PaperDoctor findings were longer on average than referee or Stanford Agentic Reviewer feedback, with its average split into 33.9 evidence words and 16.1 suggestion words.
Recompute the 40-paper feedback benchmark
实验方案已经存在,但当前处理任务没有对应的执行目标。
浏览 ArkGraph,选择本次要执行的步骤。
优先展示覆盖情况与实测结果;需要审计时再打开运行记录和技术细节。
0/21
条结论获得证据支持
0
获得支持
0
受到挑战或冲突
0
遭到反驳
0
无法判定
21
尚未评估
每一行是一条真实执行记录;命令、日志、哈希和签名都收纳在详情中。
尚无复现运行
论文中的实验方案开始执行后,运行记录会显示在这里。
逐个实验展示当前执行状态、阻塞原因、恢复动作和证据去向;技术执行与科研结论始终分开。 另有 7 条结论本次未安排独立复现。
执行成功不等于论文结论成立
目标状态回答平台有没有跑完;右侧科学结论只由不可变证据和 Assessment 决定。资源不足或平台故障不会被写成反驳论文的科研结论。
当前任务创建于目标级记录上线之前。下方状态来自历史任务的保守投影,不会伪造运行尝试、资源决策或证据关系。
PaperDoctor findings were longer on average than referee or Stanford Agentic Reviewer feedback, with its average split into 33.9 evidence words and 16.1 suggestion words.
Recompute the 40-paper feedback benchmark
实验方案已经存在,但当前处理任务没有对应的执行目标。
All 30 authors gave PaperDoctor's overall review a positive score: 70% called it somewhat helpful and 30% very helpful, yielding a mean holistic score of +1.30.
Recompute the 30-paper author-rating study
实验方案已经存在,但当前处理任务没有对应的执行目标。
In the same tumor-modeling paper, PaperDoctor reported that a novelty framing about DeepONet and Fourier Neural Operators being designed for homogeneous constant-coefficient PDEs was contradicted by the cited original papers, which include varying conditions or variable coefficients.
这条结论尚无可执行实验方案。
Authors accepted the evidence for 70.6% of all 1,299 issues; paper-level acceptance varied widely, and its positive correlation with the holistic helpfulness score was not statistically significant.
Recompute the 30-paper author-rating study
实验方案已经存在,但当前处理任务没有对应的执行目标。
In an Agents4Science case study, the paper claimed SciBERT embeddings with 768 dimensions, but all three relevant released files configured all-MiniLM-L6-v2 and produced 384-dimensional vectors; SciBERT was not imported anywhere.
Audit the claimed and configured embedding model
实验方案已经存在,但当前处理任务没有对应的执行目标。
Experiment reproduction was bottlenecked before and after execution: only 53.0% of reproduction plans reached a command, and among plans that ran, 27.0% were warnings and 34.5% errors.
Recompute the 40-paper reproduction-flow analysis
实验方案已经存在,但当前处理任务没有对应的执行目标。
The 24 optional written comments praised breadth most often but also criticized it most often; proposed fixes, novelty/framing silence, and false positives were the only aspects with more criticism than praise.
Recompute the optional-comment aspect analysis
实验方案已经存在,但当前处理任务没有对应的执行目标。
The four paper groups showed different reproduction failure locations: ICML plans usually reached execution but rarely passed, Agents4Science plans usually stopped before execution but often passed when they ran, and the two Nature groups were intermediate.
Recompute the 40-paper reproduction-flow analysis
实验方案已经存在,但当前处理任务没有对应的执行目标。
PaperDoctor's L1 check frequencies were broadly similar across paper groups, while L2 Code findings were highest for AI-written papers, Theory findings were highest for ICML, and Experiment Design findings separated the groups most strongly.
Recompute the 40-paper feedback benchmark
实验方案已经存在,但当前处理任务没有对应的执行目标。
PaperDoctor distributed its feedback more evenly across manuscript, literature, and code dimensions than referees or the Stanford Agentic Reviewer, both of which concentrated on experiments in the main body.
Recompute the 40-paper feedback benchmark
实验方案已经存在,但当前处理任务没有对应的执行目标。
PaperDoctor attached both evidence and a suggestion to every finding by construction; referees did so for 35.9% of findings and the Stanford Agentic Reviewer for 1.5%, whose dominant pattern was suggestion without evidence.
Recompute the 40-paper feedback benchmark
实验方案已经存在,但当前处理任务没有对应的执行目标。
Authors were less uncertain but also less likely to agree with findings PaperDoctor labeled Error than with Warnings; acceptance of Evidence and Suggestion was also strongly coupled.
Recompute the 30-paper author-rating study
实验方案已经存在,但当前处理任务没有对应的执行目标。
Among plans that ran, high-priority plans passed more often and errored less often than medium- or low-priority plans; medium-priority pass rates were close to low-priority rates rather than intermediate.
Recompute the 40-paper reproduction-flow analysis
实验方案已经存在,但当前处理任务没有对应的执行目标。
In an Agents4Science case study, PaperDoctor found that a bibliography entry dated the GPT-4.5 research preview to 2024 even though the release was on February 27, 2025, making the citation metadata chronologically inconsistent with the late-2024 manuscript.
Verify the GPT-4.5 citation and release year
实验方案已经存在,但当前处理任务没有对应的执行目标。
Agreement differed sharply by check: L2 claim-verification items were rejected less often than L1 surface-screening items; Experiment Review had the highest Evidence acceptance and Figure Review the lowest.
Recompute the 30-paper author-rating study
实验方案已经存在,但当前处理任务没有对应的执行目标。
Across the same 40 papers, PaperDoctor generally produced more atomic findings per reviewer and paper than published referee reports or the Stanford Agentic Reviewer, with larger within-domain spread.
Recompute the 40-paper feedback benchmark
实验方案已经存在,但当前处理任务没有对应的执行目标。
Among executed plans, training reruns passed least often and errored most often, inference/evaluation was intermediate, and statistical-test or analysis reruns passed most often and errored least often.
Recompute the 40-paper reproduction-flow analysis
实验方案已经存在,但当前处理任务没有对应的执行目标。
The interface example flags a partial paper-code mismatch: the paper says replay occurs every 10 steps for smaller models (and every 5 for FLAN-T5-3B), while the default configuration sets replay_freq to 1; replay_k does match the stated mini-batch size of 8.
Audit replay batch and interval configuration
实验方案已经存在,但当前处理任务没有对应的执行目标。
In an Agents4Science theory case, PaperDoctor found an approximation bound stated without proof, assumptions, an appendix pointer, or definitions sufficient to interpret and verify it.
Audit definitions and support for the approximation bound
实验方案已经存在,但当前处理任务没有对应的执行目标。
For a Nature Communications microscopy paper, rebuilding the shipped Figure 2d analysis yielded soft-focal contact-task recalls of 0.43 and 0.49, corresponding to a 1.14x best-pair ratio and 1.04x mean ratio rather than the prose claim of almost 2x.
Recompute the Figure 2d recall comparison
实验方案已经存在,但当前处理任务没有对应的执行目标。
Reasons plans never ran differed by paper group: missing code dominated SocialScience, wet-lab procedures dominated NatureScience, and restricted data plus shared blockers dominated Agents4Science; overall, missing code or model weights was the largest blocker class.
Recompute the 40-paper reproduction-flow analysis
实验方案已经存在,但当前处理任务没有对应的执行目标。
由用户本地执行并上传,平台已验证文件签名;结论由上传的运行提供。
还没有社区运行。