Explore ArkGraph and select the steps to run.
The paper’s claims are available in Research claims.
0 / 28 claims verified
the rest still being verified
Experiment plan ready; no runs yet.
0/5 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
source: Referee reports | 29.1 words/finding | — | Not assessed |
source: Stanford Agentic Reviewer | 27.1 words/finding | — | Not assessed |
source: PaperDoctor | 50 words/finding | — | Not assessed |
source: PaperDoctor · component: Evidence | 33.9 words/finding | — | Not assessed |
source: PaperDoctor · component: Suggestion | 16.1 words/finding | — | Not assessed |
Experiment plan ready; no runs yet.
0/9 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 30 papers | — | Not assessed | |
| 25 students | — | Not assessed | |
| 1,299 items | — | Not assessed | |
rating: -1 harmful | 0 % of papers | — | Not assessed |
rating: 0 no help | 0 % of papers | — | Not assessed |
rating: +1 somewhat helpful | 70 % of papers | — | Not assessed |
rating: +2 very helpful | 30 % of papers | — | Not assessed |
| 30 papers | — | Not assessed | |
| 1.3 rating points | — | Not assessed |
Experiment plan ready; no runs yet.
0/7 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 70.6 % of all items | — | Not assessed | |
| 68.5% | — | Not assessed | |
| 23.5% | — | Not assessed | |
| 97% | — | Not assessed | |
| 71.1% | — | Not assessed | |
| 0.29 Pearson r | — | Not assessed | |
| 0.12 p-value | — | Not assessed |
Experiment plan ready; no runs yet.
0/3 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
source: paper claim · encoder: SciBERT | 768 dimensions | — | Not assessed |
source: released code · encoder: all-MiniLM-L6-v2 | 384 dimensions | — | Not assessed |
| 3 files | — | Not assessed |
Experiment plan ready; no runs yet.
0/21 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 43.8 claims/paper | — | Not assessed | |
funnel stage: Planned | 14.5 plans/paper | — | Not assessed |
funnel stage: Planned | 100 % of plans | — | Not assessed |
| 29.3 claims/paper | — | Not assessed | |
initial feasibility: Ready | 2.6 plans/paper | — | Not assessed |
initial feasibility: Ready | 18 % of plans | — | Not assessed |
initial feasibility: Blocked | 11.9 plans/paper | — | Not assessed |
initial feasibility: Blocked | 82 % of plans | — | Not assessed |
execution: Ran | 7.7 plans/paper | — | Not assessed |
execution: Ran | 53 % of plans | — | Not assessed |
execution: Never ran | 6.8 plans/paper | — | Not assessed |
execution: Never ran | 47 % of plans | — | Not assessed |
Experiment plan ready; no runs yet.
0/30 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 24 comments | — | Not assessed | |
| 79 responses | — | Not assessed | |
sentiment: praise | 58 responses | — | Not assessed |
sentiment: criticism | 21 responses | — | Not assessed |
aspect: Writing and language · sentiment: praise | 9 responses | — | Not assessed |
aspect: Writing and language · sentiment: criticism | 2 responses | — | Not assessed |
aspect: Presentation clarity · sentiment: praise | 6 responses | — | Not assessed |
aspect: Presentation clarity · sentiment: criticism | 0 responses | — | Not assessed |
aspect: Coverage of the review · sentiment: praise | 10 responses | — | Not assessed |
aspect: Coverage of the review · sentiment: criticism | 5 responses | — | Not assessed |
aspect: Cross-part consistency · sentiment: praise | 5 responses | — | Not assessed |
aspect: Cross-part consistency · sentiment: criticism | 0 responses | — | Not assessed |
Experiment plan ready; no runs yet.
0/20 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
status: Total · paper group: Agents4Science | 86 plans | — | Not assessed |
status: Never ran · paper group: Agents4Science | 68 plans | — | Not assessed |
status: Ran & Not Pass · paper group: Agents4Science | 5 plans | — | Not assessed |
status: Ran & Passed · paper group: Agents4Science | 13 plans | — | Not assessed |
paper group: Agents4Science | 72.2% | — | Not assessed |
status: Total · paper group: NatureScience | 145 plans | — | Not assessed |
status: Never ran · paper group: NatureScience | 80 plans | — | Not assessed |
status: Ran & Not Pass · paper group: NatureScience | 34 plans | — | Not assessed |
status: Ran & Passed · paper group: NatureScience | 31 plans | — | Not assessed |
paper group: NatureScience | 47.7% | — | Not assessed |
status: Total · paper group: SocialScience | 181 plans | — | Not assessed |
status: Never ran · paper group: SocialScience | 89 plans | — | Not assessed |
Experiment plan ready; no runs yet.
0/28 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
check: Writing · paper group: Agents4Science | 14.6 findings/paper | — | Not assessed |
check: Writing · paper group: ICML | 15.4 findings/paper | — | Not assessed |
check: Writing · paper group: NatureScience | 14.7 findings/paper | — | Not assessed |
check: Writing · paper group: SocialScience | 12.6 findings/paper | — | Not assessed |
check: Figure · paper group: Agents4Science | 7 findings/paper | — | Not assessed |
check: Figure · paper group: ICML | 6.1 findings/paper | — | Not assessed |
check: Figure · paper group: NatureScience | 5.6 findings/paper | — | Not assessed |
check: Figure · paper group: SocialScience | 5 findings/paper | — | Not assessed |
check: Citation · paper group: Agents4Science | 4 findings/paper | — | Not assessed |
check: Citation · paper group: ICML | 5.3 findings/paper | — | Not assessed |
check: Citation · paper group: NatureScience | 3.8 findings/paper | — | Not assessed |
check: Citation · paper group: SocialScience | 2.8 findings/paper | — | Not assessed |
Experiment plan ready; no runs yet.
0/22 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
topic: Main body: experiments · source: Referee reports | 49.5 % of findings | — | Not assessed |
topic: Main body: methodology · source: Referee reports | 13.1 % of findings | — | Not assessed |
topic: Main body: writing · source: Referee reports | 16.7 % of findings | — | Not assessed |
topic: Main body: figures · source: Referee reports | 10.8 % of findings | — | Not assessed |
topic: External: literature · source: Referee reports | 7 % of findings | — | Not assessed |
topic: External: code · source: Referee reports | 3 % of findings | — | Not assessed |
topic: Main body: experiments · source: Stanford Agentic Reviewer | 72.6 % of findings | — | Not assessed |
topic: Main body: methodology · source: Stanford Agentic Reviewer | 12.8 % of findings | — | Not assessed |
topic: Main body: writing · source: Stanford Agentic Reviewer | 1.7 % of findings | — | Not assessed |
topic: Main body: figures · source: Stanford Agentic Reviewer | 2.4 % of findings | — | Not assessed |
topic: External: literature · source: Stanford Agentic Reviewer | 7.6 % of findings | — | Not assessed |
topic: External: code · source: Stanford Agentic Reviewer | 2.9 % of findings | — | Not assessed |
Experiment plan ready; no runs yet.
0/16 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
source: Referee reports · pattern: Evidence + Suggestion | 35.9 % of findings | — | Not assessed |
source: Referee reports · pattern: Evidence only | 9.3 % of findings | — | Not assessed |
source: Referee reports · pattern: Suggestion only | 35.8 % of findings | — | Not assessed |
source: Referee reports · pattern: Neither | 19 % of findings | — | Not assessed |
source: Stanford Agentic Reviewer · pattern: Evidence + Suggestion | 1.5 % of findings | — | Not assessed |
source: Stanford Agentic Reviewer · pattern: Evidence only | 0.7 % of findings | — | Not assessed |
source: Stanford Agentic Reviewer · pattern: Suggestion only | 69.9 % of findings | — | Not assessed |
source: Stanford Agentic Reviewer · pattern: Neither | 27.8 % of findings | — | Not assessed |
source: PaperDoctor · pattern: Evidence + Suggestion | 100 % of findings | — | Not assessed |
source: PaperDoctor · pattern: Evidence only | 0 % of findings | — | Not assessed |
source: PaperDoctor · pattern: Suggestion only | 0 % of findings | — | Not assessed |
source: PaperDoctor · pattern: Neither | 0 % of findings | — | Not assessed |
Experiment plan ready; no runs yet.
0/14 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
axis: Evidence · verdict: accepted · PaperDoctor severity: Warning | 71% | — | Not assessed |
axis: Evidence · verdict: uncertain · PaperDoctor severity: Warning | 11% | — | Not assessed |
axis: Evidence · verdict: rejected · PaperDoctor severity: Warning | 18% | — | Not assessed |
axis: Suggestion · verdict: accepted · PaperDoctor severity: Warning | 71% | — | Not assessed |
axis: Suggestion · verdict: uncertain · PaperDoctor severity: Warning | 12% | — | Not assessed |
axis: Suggestion · verdict: rejected · PaperDoctor severity: Warning | 17% | — | Not assessed |
axis: Evidence · verdict: accepted · PaperDoctor severity: Error | 67% | — | Not assessed |
axis: Evidence · verdict: uncertain · PaperDoctor severity: Error | 8% | — | Not assessed |
axis: Evidence · verdict: rejected · PaperDoctor severity: Error | 25% | — | Not assessed |
axis: Suggestion · verdict: accepted · PaperDoctor severity: Error | 66% | — | Not assessed |
axis: Suggestion · verdict: uncertain · PaperDoctor severity: Error | 8% | — | Not assessed |
axis: Suggestion · verdict: rejected · PaperDoctor severity: Error | 26% | — | Not assessed |
Experiment plan ready; no runs yet.
0/12 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
outcome: Pass · priority: High | 47.3% | — | Not assessed |
outcome: Warning · priority: High | 23.6% | — | Not assessed |
outcome: Error · priority: High | 29.1% | — | Not assessed |
priority: High | 110 plans | — | Not assessed |
outcome: Pass · priority: Medium | 33.3% | — | Not assessed |
outcome: Warning · priority: Medium | 32.1% | — | Not assessed |
outcome: Error · priority: Medium | 34.6% | — | Not assessed |
priority: Medium | 81 plans | — | Not assessed |
outcome: Pass · priority: Low | 33.6% | — | Not assessed |
outcome: Warning · priority: Low | 26.7% | — | Not assessed |
outcome: Error · priority: Low | 39.7% | — | Not assessed |
priority: Low | 116 plans | — | Not assessed |
Experiment plan ready; no runs yet.
0/2 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
status: printed in bibliography | 2,024 year | — | Not assessed |
status: web-grounded release date | 2,025 year | — | Not assessed |
Experiment plan ready; no runs yet.
0/66 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
axis: Evidence · scope: overall · verdict: accepted | 71% | — | Not assessed |
axis: Evidence · scope: overall · verdict: uncertain | 11% | — | Not assessed |
axis: Evidence · scope: overall · verdict: rejected | 19% | — | Not assessed |
axis: Suggestion · scope: overall · verdict: accepted | 70% | — | Not assessed |
axis: Suggestion · scope: overall · verdict: uncertain | 11% | — | Not assessed |
axis: Suggestion · scope: overall · verdict: rejected | 18% | — | Not assessed |
scope: overall | 43.3 items/paper | — | Not assessed |
axis: Evidence · scope: L1 paper-only screening · verdict: accepted | 62% | — | Not assessed |
axis: Evidence · scope: L1 paper-only screening · verdict: uncertain | 12% | — | Not assessed |
axis: Evidence · scope: L1 paper-only screening · verdict: rejected | 27% | — | Not assessed |
axis: Suggestion · scope: L1 paper-only screening · verdict: accepted | 63% | — | Not assessed |
axis: Suggestion · scope: L1 paper-only screening · verdict: uncertain | 12% | — | Not assessed |
Experiment plan ready; no runs yet.
0/16 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
source: Referee reports · paper group: Agents4Science | 6.9 findings/reviewer/paper | — | Not assessed |
source: Stanford Agentic Reviewer · paper group: Agents4Science | 44 findings/reviewer/paper | — | Not assessed |
source: PaperDoctor · paper group: Agents4Science | 56.4 findings/reviewer/paper | — | Not assessed |
source: Referee reports · paper group: ICML | 6.7 findings/reviewer/paper | — | Not assessed |
source: Stanford Agentic Reviewer · paper group: ICML | 25.2 findings/reviewer/paper | — | Not assessed |
source: PaperDoctor · paper group: ICML | 47.9 findings/reviewer/paper | — | Not assessed |
source: Referee reports · paper group: NatureScience | 16.2 findings/reviewer/paper | — | Not assessed |
source: Stanford Agentic Reviewer · paper group: NatureScience | 37 findings/reviewer/paper | — | Not assessed |
source: PaperDoctor · paper group: NatureScience | 47.7 findings/reviewer/paper | — | Not assessed |
source: Referee reports · paper group: SocialScience | 17.5 findings/reviewer/paper | — | Not assessed |
source: Stanford Agentic Reviewer · paper group: SocialScience | 36.2 findings/reviewer/paper | — | Not assessed |
source: PaperDoctor · paper group: SocialScience | 36.1 findings/reviewer/paper | — | Not assessed |
Experiment plan ready; no runs yet.
0/12 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
outcome: Pass · rerun type: Train | 11.3% | — | Not assessed |
outcome: Warning · rerun type: Train | 30.2% | — | Not assessed |
outcome: Error · rerun type: Train | 58.5% | — | Not assessed |
rerun type: Train | 53 plans | — | Not assessed |
outcome: Pass · rerun type: Eval/inference | 39% | — | Not assessed |
outcome: Warning · rerun type: Eval/inference | 22.8% | — | Not assessed |
outcome: Error · rerun type: Eval/inference | 38.2% | — | Not assessed |
rerun type: Eval/inference | 123 plans | — | Not assessed |
outcome: Pass · rerun type: Analysis/statistical test | 48.9% | — | Not assessed |
outcome: Warning · rerun type: Analysis/statistical test | 29.8% | — | Not assessed |
outcome: Error · rerun type: Analysis/statistical test | 21.4% | — | Not assessed |
rerun type: Analysis/statistical test | 131 plans | — | Not assessed |
Experiment plan ready; no runs yet.
0/6 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
configuration: smaller models | 8 examples | — | Not assessed |
configuration: smaller models | 10 training steps | — | Not assessed |
configuration: FLAN-T5-3B | 4 examples | — | Not assessed |
configuration: FLAN-T5-3B | 5 training steps | — | Not assessed |
| 1 config value | — | Not assessed | |
| 8 examples | — | Not assessed |
Reported
5 groups
Observed
—
Experiment plan ready; no runs yet.
0/1 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 5 groups | — | Not assessed |
Experiment plan ready; no runs yet.
0/6 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 2 x | — | Not assessed | |
configuration: soft-focal configuration 1 | 0.43 recall score | — | Not assessed |
configuration: soft-focal configuration 2 | 0.49 recall score | — | Not assessed |
| 1.14 x | — | Not assessed | |
| 1.04 x | — | Not assessed | |
| 14% | — | Not assessed |
Experiment plan ready; no runs yet.
0/30 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 272 plans | — | Not assessed | |
reason: Incomplete runnable environment (missing code or model weights) | 33.1 % of all plans that never ran | — | Not assessed |
paper group: SocialScience | 89 plans | — | Not assessed |
reason: Missing code · paper group: SocialScience | 65.2% | — | Not assessed |
reason: Wet-lab · paper group: SocialScience | 0% | — | Not assessed |
reason: Restricted data · paper group: SocialScience | 6.7% | — | Not assessed |
reason: Same blocker as another · paper group: SocialScience | 0% | — | Not assessed |
reason: Resources · paper group: SocialScience | 0% | — | Not assessed |
reason: No reason given · paper group: SocialScience | 28.1% | — | Not assessed |
paper group: ICML | 35 plans | — | Not assessed |
reason: Missing code · paper group: ICML | 37.1% | — | Not assessed |
reason: Wet-lab · paper group: ICML | 0% | — | Not assessed |
The original FNO/DeepONet comparison remains scientifically meaningful source context in exp-case-theory-bound, but this handoff cannot bind it to an automatic numeric measurement without inventing a target.
The qualitative coverage matrix is preserved as source context in exp-feedback-benchmark, but no numeric surrogate is fabricated.
Method/architecture context with no independent reported measurement. The fixed repository is tree-sitter rather than PaperDoctor, so no implementation claim is inferred from it.
Interpretive caveat on the represented reproduction aggregates, not an additional experiment.
Interpretive limitation attached to the Figure Review acceptance result, not a separately measurable result.
Future-work and dependency context, not an independently reported measurement.
Normative scope statement about human judgment and unavailable advisor feedback, not an independently testable result.
Study-scope limitation, not an additional empirical result; it is preserved when interpreting exp-author-study-quantitative.