浏览 ArkGraph,选择本次要执行的步骤。
论文中的结论可在「研究结论」中查看。
0 / 28 条结论已通过验证
其余仍在验证中
实验方案已生成,还没有运行记录
已评估 0/5 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
source: Referee reports | 29.1 words/finding | — | 尚未评估 |
source: Stanford Agentic Reviewer | 27.1 words/finding | — | 尚未评估 |
source: PaperDoctor | 50 words/finding | — | 尚未评估 |
source: PaperDoctor · component: Evidence | 33.9 words/finding | — | 尚未评估 |
source: PaperDoctor · component: Suggestion | 16.1 words/finding | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/9 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 30 papers | — | 尚未评估 | |
| 25 students | — | 尚未评估 | |
| 1,299 items | — | 尚未评估 | |
rating: -1 harmful | 0 % of papers | — | 尚未评估 |
rating: 0 no help | 0 % of papers | — | 尚未评估 |
rating: +1 somewhat helpful | 70 % of papers | — | 尚未评估 |
rating: +2 very helpful | 30 % of papers | — | 尚未评估 |
| 30 papers | — | 尚未评估 | |
| 1.3 rating points | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/7 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 70.6 % of all items | — | 尚未评估 | |
| 68.5% | — | 尚未评估 | |
| 23.5% | — | 尚未评估 | |
| 97% | — | 尚未评估 | |
| 71.1% | — | 尚未评估 | |
| 0.29 Pearson r | — | 尚未评估 | |
| 0.12 p-value | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/3 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
source: paper claim · encoder: SciBERT | 768 dimensions | — | 尚未评估 |
source: released code · encoder: all-MiniLM-L6-v2 | 384 dimensions | — | 尚未评估 |
| 3 files | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/21 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 43.8 claims/paper | — | 尚未评估 | |
funnel stage: Planned | 14.5 plans/paper | — | 尚未评估 |
funnel stage: Planned | 100 % of plans | — | 尚未评估 |
| 29.3 claims/paper | — | 尚未评估 | |
initial feasibility: Ready | 2.6 plans/paper | — | 尚未评估 |
initial feasibility: Ready | 18 % of plans | — | 尚未评估 |
initial feasibility: Blocked | 11.9 plans/paper | — | 尚未评估 |
initial feasibility: Blocked | 82 % of plans | — | 尚未评估 |
execution: Ran | 7.7 plans/paper | — | 尚未评估 |
execution: Ran | 53 % of plans | — | 尚未评估 |
execution: Never ran | 6.8 plans/paper | — | 尚未评估 |
execution: Never ran | 47 % of plans | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/30 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 24 comments | — | 尚未评估 | |
| 79 responses | — | 尚未评估 | |
sentiment: praise | 58 responses | — | 尚未评估 |
sentiment: criticism | 21 responses | — | 尚未评估 |
aspect: Writing and language · sentiment: praise | 9 responses | — | 尚未评估 |
aspect: Writing and language · sentiment: criticism | 2 responses | — | 尚未评估 |
aspect: Presentation clarity · sentiment: praise | 6 responses | — | 尚未评估 |
aspect: Presentation clarity · sentiment: criticism | 0 responses | — | 尚未评估 |
aspect: Coverage of the review · sentiment: praise | 10 responses | — | 尚未评估 |
aspect: Coverage of the review · sentiment: criticism | 5 responses | — | 尚未评估 |
aspect: Cross-part consistency · sentiment: praise | 5 responses | — | 尚未评估 |
aspect: Cross-part consistency · sentiment: criticism | 0 responses | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/20 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
status: Total · paper group: Agents4Science | 86 plans | — | 尚未评估 |
status: Never ran · paper group: Agents4Science | 68 plans | — | 尚未评估 |
status: Ran & Not Pass · paper group: Agents4Science | 5 plans | — | 尚未评估 |
status: Ran & Passed · paper group: Agents4Science | 13 plans | — | 尚未评估 |
paper group: Agents4Science | 72.2% | — | 尚未评估 |
status: Total · paper group: NatureScience | 145 plans | — | 尚未评估 |
status: Never ran · paper group: NatureScience | 80 plans | — | 尚未评估 |
status: Ran & Not Pass · paper group: NatureScience | 34 plans | — | 尚未评估 |
status: Ran & Passed · paper group: NatureScience | 31 plans | — | 尚未评估 |
paper group: NatureScience | 47.7% | — | 尚未评估 |
status: Total · paper group: SocialScience | 181 plans | — | 尚未评估 |
status: Never ran · paper group: SocialScience | 89 plans | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/28 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
check: Writing · paper group: Agents4Science | 14.6 findings/paper | — | 尚未评估 |
check: Writing · paper group: ICML | 15.4 findings/paper | — | 尚未评估 |
check: Writing · paper group: NatureScience | 14.7 findings/paper | — | 尚未评估 |
check: Writing · paper group: SocialScience | 12.6 findings/paper | — | 尚未评估 |
check: Figure · paper group: Agents4Science | 7 findings/paper | — | 尚未评估 |
check: Figure · paper group: ICML | 6.1 findings/paper | — | 尚未评估 |
check: Figure · paper group: NatureScience | 5.6 findings/paper | — | 尚未评估 |
check: Figure · paper group: SocialScience | 5 findings/paper | — | 尚未评估 |
check: Citation · paper group: Agents4Science | 4 findings/paper | — | 尚未评估 |
check: Citation · paper group: ICML | 5.3 findings/paper | — | 尚未评估 |
check: Citation · paper group: NatureScience | 3.8 findings/paper | — | 尚未评估 |
check: Citation · paper group: SocialScience | 2.8 findings/paper | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/22 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
topic: Main body: experiments · source: Referee reports | 49.5 % of findings | — | 尚未评估 |
topic: Main body: methodology · source: Referee reports | 13.1 % of findings | — | 尚未评估 |
topic: Main body: writing · source: Referee reports | 16.7 % of findings | — | 尚未评估 |
topic: Main body: figures · source: Referee reports | 10.8 % of findings | — | 尚未评估 |
topic: External: literature · source: Referee reports | 7 % of findings | — | 尚未评估 |
topic: External: code · source: Referee reports | 3 % of findings | — | 尚未评估 |
topic: Main body: experiments · source: Stanford Agentic Reviewer | 72.6 % of findings | — | 尚未评估 |
topic: Main body: methodology · source: Stanford Agentic Reviewer | 12.8 % of findings | — | 尚未评估 |
topic: Main body: writing · source: Stanford Agentic Reviewer | 1.7 % of findings | — | 尚未评估 |
topic: Main body: figures · source: Stanford Agentic Reviewer | 2.4 % of findings | — | 尚未评估 |
topic: External: literature · source: Stanford Agentic Reviewer | 7.6 % of findings | — | 尚未评估 |
topic: External: code · source: Stanford Agentic Reviewer | 2.9 % of findings | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/16 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
source: Referee reports · pattern: Evidence + Suggestion | 35.9 % of findings | — | 尚未评估 |
source: Referee reports · pattern: Evidence only | 9.3 % of findings | — | 尚未评估 |
source: Referee reports · pattern: Suggestion only | 35.8 % of findings | — | 尚未评估 |
source: Referee reports · pattern: Neither | 19 % of findings | — | 尚未评估 |
source: Stanford Agentic Reviewer · pattern: Evidence + Suggestion | 1.5 % of findings | — | 尚未评估 |
source: Stanford Agentic Reviewer · pattern: Evidence only | 0.7 % of findings | — | 尚未评估 |
source: Stanford Agentic Reviewer · pattern: Suggestion only | 69.9 % of findings | — | 尚未评估 |
source: Stanford Agentic Reviewer · pattern: Neither | 27.8 % of findings | — | 尚未评估 |
source: PaperDoctor · pattern: Evidence + Suggestion | 100 % of findings | — | 尚未评估 |
source: PaperDoctor · pattern: Evidence only | 0 % of findings | — | 尚未评估 |
source: PaperDoctor · pattern: Suggestion only | 0 % of findings | — | 尚未评估 |
source: PaperDoctor · pattern: Neither | 0 % of findings | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/14 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
axis: Evidence · verdict: accepted · PaperDoctor severity: Warning | 71% | — | 尚未评估 |
axis: Evidence · verdict: uncertain · PaperDoctor severity: Warning | 11% | — | 尚未评估 |
axis: Evidence · verdict: rejected · PaperDoctor severity: Warning | 18% | — | 尚未评估 |
axis: Suggestion · verdict: accepted · PaperDoctor severity: Warning | 71% | — | 尚未评估 |
axis: Suggestion · verdict: uncertain · PaperDoctor severity: Warning | 12% | — | 尚未评估 |
axis: Suggestion · verdict: rejected · PaperDoctor severity: Warning | 17% | — | 尚未评估 |
axis: Evidence · verdict: accepted · PaperDoctor severity: Error | 67% | — | 尚未评估 |
axis: Evidence · verdict: uncertain · PaperDoctor severity: Error | 8% | — | 尚未评估 |
axis: Evidence · verdict: rejected · PaperDoctor severity: Error | 25% | — | 尚未评估 |
axis: Suggestion · verdict: accepted · PaperDoctor severity: Error | 66% | — | 尚未评估 |
axis: Suggestion · verdict: uncertain · PaperDoctor severity: Error | 8% | — | 尚未评估 |
axis: Suggestion · verdict: rejected · PaperDoctor severity: Error | 26% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/12 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
outcome: Pass · priority: High | 47.3% | — | 尚未评估 |
outcome: Warning · priority: High | 23.6% | — | 尚未评估 |
outcome: Error · priority: High | 29.1% | — | 尚未评估 |
priority: High | 110 plans | — | 尚未评估 |
outcome: Pass · priority: Medium | 33.3% | — | 尚未评估 |
outcome: Warning · priority: Medium | 32.1% | — | 尚未评估 |
outcome: Error · priority: Medium | 34.6% | — | 尚未评估 |
priority: Medium | 81 plans | — | 尚未评估 |
outcome: Pass · priority: Low | 33.6% | — | 尚未评估 |
outcome: Warning · priority: Low | 26.7% | — | 尚未评估 |
outcome: Error · priority: Low | 39.7% | — | 尚未评估 |
priority: Low | 116 plans | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/2 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
status: printed in bibliography | 2,024 year | — | 尚未评估 |
status: web-grounded release date | 2,025 year | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/66 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
axis: Evidence · scope: overall · verdict: accepted | 71% | — | 尚未评估 |
axis: Evidence · scope: overall · verdict: uncertain | 11% | — | 尚未评估 |
axis: Evidence · scope: overall · verdict: rejected | 19% | — | 尚未评估 |
axis: Suggestion · scope: overall · verdict: accepted | 70% | — | 尚未评估 |
axis: Suggestion · scope: overall · verdict: uncertain | 11% | — | 尚未评估 |
axis: Suggestion · scope: overall · verdict: rejected | 18% | — | 尚未评估 |
scope: overall | 43.3 items/paper | — | 尚未评估 |
axis: Evidence · scope: L1 paper-only screening · verdict: accepted | 62% | — | 尚未评估 |
axis: Evidence · scope: L1 paper-only screening · verdict: uncertain | 12% | — | 尚未评估 |
axis: Evidence · scope: L1 paper-only screening · verdict: rejected | 27% | — | 尚未评估 |
axis: Suggestion · scope: L1 paper-only screening · verdict: accepted | 63% | — | 尚未评估 |
axis: Suggestion · scope: L1 paper-only screening · verdict: uncertain | 12% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/16 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
source: Referee reports · paper group: Agents4Science | 6.9 findings/reviewer/paper | — | 尚未评估 |
source: Stanford Agentic Reviewer · paper group: Agents4Science | 44 findings/reviewer/paper | — | 尚未评估 |
source: PaperDoctor · paper group: Agents4Science | 56.4 findings/reviewer/paper | — | 尚未评估 |
source: Referee reports · paper group: ICML | 6.7 findings/reviewer/paper | — | 尚未评估 |
source: Stanford Agentic Reviewer · paper group: ICML | 25.2 findings/reviewer/paper | — | 尚未评估 |
source: PaperDoctor · paper group: ICML | 47.9 findings/reviewer/paper | — | 尚未评估 |
source: Referee reports · paper group: NatureScience | 16.2 findings/reviewer/paper | — | 尚未评估 |
source: Stanford Agentic Reviewer · paper group: NatureScience | 37 findings/reviewer/paper | — | 尚未评估 |
source: PaperDoctor · paper group: NatureScience | 47.7 findings/reviewer/paper | — | 尚未评估 |
source: Referee reports · paper group: SocialScience | 17.5 findings/reviewer/paper | — | 尚未评估 |
source: Stanford Agentic Reviewer · paper group: SocialScience | 36.2 findings/reviewer/paper | — | 尚未评估 |
source: PaperDoctor · paper group: SocialScience | 36.1 findings/reviewer/paper | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/12 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
outcome: Pass · rerun type: Train | 11.3% | — | 尚未评估 |
outcome: Warning · rerun type: Train | 30.2% | — | 尚未评估 |
outcome: Error · rerun type: Train | 58.5% | — | 尚未评估 |
rerun type: Train | 53 plans | — | 尚未评估 |
outcome: Pass · rerun type: Eval/inference | 39% | — | 尚未评估 |
outcome: Warning · rerun type: Eval/inference | 22.8% | — | 尚未评估 |
outcome: Error · rerun type: Eval/inference | 38.2% | — | 尚未评估 |
rerun type: Eval/inference | 123 plans | — | 尚未评估 |
outcome: Pass · rerun type: Analysis/statistical test | 48.9% | — | 尚未评估 |
outcome: Warning · rerun type: Analysis/statistical test | 29.8% | — | 尚未评估 |
outcome: Error · rerun type: Analysis/statistical test | 21.4% | — | 尚未评估 |
rerun type: Analysis/statistical test | 131 plans | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/6 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
configuration: smaller models | 8 examples | — | 尚未评估 |
configuration: smaller models | 10 training steps | — | 尚未评估 |
configuration: FLAN-T5-3B | 4 examples | — | 尚未评估 |
configuration: FLAN-T5-3B | 5 training steps | — | 尚未评估 |
| 1 config value | — | 尚未评估 | |
| 8 examples | — | 尚未评估 |
报告
5 groups
观测
—
实验方案已生成,还没有运行记录
已评估 0/1 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 5 groups | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/6 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 2 x | — | 尚未评估 | |
configuration: soft-focal configuration 1 | 0.43 recall score | — | 尚未评估 |
configuration: soft-focal configuration 2 | 0.49 recall score | — | 尚未评估 |
| 1.14 x | — | 尚未评估 | |
| 1.04 x | — | 尚未评估 | |
| 14% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/30 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 272 plans | — | 尚未评估 | |
reason: Incomplete runnable environment (missing code or model weights) | 33.1 % of all plans that never ran | — | 尚未评估 |
paper group: SocialScience | 89 plans | — | 尚未评估 |
reason: Missing code · paper group: SocialScience | 65.2% | — | 尚未评估 |
reason: Wet-lab · paper group: SocialScience | 0% | — | 尚未评估 |
reason: Restricted data · paper group: SocialScience | 6.7% | — | 尚未评估 |
reason: Same blocker as another · paper group: SocialScience | 0% | — | 尚未评估 |
reason: Resources · paper group: SocialScience | 0% | — | 尚未评估 |
reason: No reason given · paper group: SocialScience | 28.1% | — | 尚未评估 |
paper group: ICML | 35 plans | — | 尚未评估 |
reason: Missing code · paper group: ICML | 37.1% | — | 尚未评估 |
reason: Wet-lab · paper group: ICML | 0% | — | 尚未评估 |
The original FNO/DeepONet comparison remains scientifically meaningful source context in exp-case-theory-bound, but this handoff cannot bind it to an automatic numeric measurement without inventing a target.
The qualitative coverage matrix is preserved as source context in exp-feedback-benchmark, but no numeric surrogate is fabricated.
Method/architecture context with no independent reported measurement. The fixed repository is tree-sitter rather than PaperDoctor, so no implementation claim is inferred from it.
Interpretive caveat on the represented reproduction aggregates, not an additional experiment.
Interpretive limitation attached to the Figure Review acceptance result, not a separately measurable result.
Future-work and dependency context, not an independently reported measurement.
Normative scope statement about human judgment and unavailable advisor feedback, not an independently testable result.
Study-scope limitation, not an additional empirical result; it is preserved when interpreting exp-author-study-quantitative.