Credit estimate for this reproduction
Expected usage
549
Maximum reservation
1683
Estimated compute time
About 210 min
Actual usage was 0 credits; 0 credits were charged. Compute ran for about 0 minutes and the model used 0 tokens.
Failed attempts also used about 14 minutes of compute and 7142207 model tokens. They produced no importable result and were not billed.
The paper’s claims are available in Research claims.
0 / 4 claims verified
the rest still being verified
Experiment plan ready; no runs yet.
0/3 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
trials: 5 · dataset: HotPotQA · metric_scope: task success rate · test_tasks_per_trial: 100 | 22 percentage_points | — | Not assessed |
trials: 5 · dataset: ChartQAPro · metric_scope: task success rate · test_tasks_per_trial: 100 | 26 percentage_points | — | Not assessed |
trials: 5 · dataset: Mind2Web · metric_scope: task success rate · test_tasks_per_trial: 100 | 27 percentage_points | — | Not assessed |
Required input unavailable · Execution will continue after the required code, data, or input is supplied.
Experiment plan ready; no runs yet.
0/3 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
split: held-in · method: ProFA · dataset: HotPotQA · value_scope: test set of constructed Who & When Pro dataset · accuracy_level: agent | 0.85 fraction | — | Not assessed |
split: held-out Algorithm-Generated · method: ProFA · dataset: Who & When · failure_logs: 184 · accuracy_level: step | 0.53 fraction | — | Not assessed |
split: held-out Hand-Crafted · method: ProFA · dataset: Who & When · failure_logs: 184 · accuracy_level: step | 0.2 fraction | — | Not assessed |
Experiment execution error · No automatic retry; a person must decide what to do next.
The claim is supported by plotted curves without machine-readable scalar values, and direct evaluation requires the unavailable datasets, model assets, external agent services, and exact implementation details.
This is an explicit limitation note, not an independent empirical measurement requiring reproduction.