Reproduction materials
Prepared 0/0 materials
资料与依赖由同一实验工作区准备,直接复用已有缓存,不另启准备 Agent。
This experiment has no separate content download, but dependencies, configuration, and loaders are still checked.
Prefetching only saves time. Unprepared material is handed to the complete reproduction Agent to download, adapt, and verify autonomously.
Credit estimate for this reproduction
Expected usage
509
Maximum reservation
1323
Estimated compute time
About 195 min
Actual usage was 156 credits; 0 credits were charged. Compute ran for about 9 minutes and the model used 6209368 tokens.
The paper’s claims are available in Research claims.
0 / 6 claims verified
the rest still being verified
The rollout ablation K=1 vs K=2 tests the core theoretical inflection point on Qwen3-0.6B-Base and can be partially evaluated alongside the main experiment.
0/10 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 6.67 percentage_points | — | Not assessed | |
| 6.67 percentage_points | — | Not assessed | |
| 22.8 percentage_points | — | Not assessed | |
| 15.15 percentage_points | — | Not assessed | |
| 6.67 percentage_points | — | Not assessed | |
| 10 percentage_points | — | Not assessed | |
| 24.2 percentage_points | — | Not assessed | |
| 20.2 percentage_points | — | Not assessed | |
| 20.71 percentage_points | — | Not assessed | |
| 20.71 percentage_points | — | Not assessed |
Training Qwen3-1.7B full-parameter fine-tuning for 1 epoch with 2 online rollouts exceeds the cumulative compute and runtime budget when executed alongside the primary 0.6B verification experiments.
0/6 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 5.42 percentage_points | — | Not assessed | |
| 26.67 percentage_points | — | Not assessed | |
| 3.96 percentage_points | — | Not assessed | |
| 23.34 percentage_points | — | Not assessed | |
| 57.4 percentage_points | — | Not assessed | |
| 26.27 percentage_points | — | Not assessed |
Experiment plan ready; no runs yet.
0/12 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 0.83 percentage_points | — | Not assessed | |
| 6.67 percentage_points | — | Not assessed | |
| 1.04 percentage_points | — | Not assessed | |
| 10 percentage_points | — | Not assessed | |
| 8.72 percentage_points | — | Not assessed | |
| 41.86 percentage_points | — | Not assessed | |
| 13.04 percentage_points | — | Not assessed | |
| 56.52 percentage_points | — | Not assessed | |
| 5.14 percentage_points | — | Not assessed | |
| 28.89 percentage_points | — | Not assessed | |
| 24.2 percentage_points | — | Not assessed | |
| 20.2 percentage_points | — | Not assessed |
Training Qwen3-4B full-parameter fine-tuning for 1 epoch with 2 online rollouts requires multi-GPU distributed setup and exceeds the single L4 GPU memory and budget limits.
0/6 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 13.33 percentage_points | — | Not assessed | |
| 36.67 percentage_points | — | Not assessed | |
| 16.25 percentage_points | — | Not assessed | |
| 36.67 percentage_points | — | Not assessed | |
| 79 percentage_points | — | Not assessed | |
| 32.32 percentage_points | — | Not assessed |
Running additional training runs at K=4 across multiple alternative weighting heuristics exceeds the cumulative paper budget allocation.
0/6 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 23.8 percentage_points | — | Not assessed | |
| 15.15 percentage_points | — | Not assessed | |
| 21.4 percentage_points | — | Not assessed | |
| 15.17 percentage_points | — | Not assessed | |
| 24 percentage_points | — | Not assessed | |
| 20.71 percentage_points | — | Not assessed |
This is a qualitative scope limitation acknowledged by the authors rather than an empirical numerical benchmark claim.