复现资料准备
已准备 0/0 项资料
资料与依赖由同一实验工作区准备,直接复用已有缓存,不另启准备 Agent。
本实验没有需要单独下载的正文资产,仍会核验依赖、配置和加载器。
预下载只是节省时间的优化;未准备的内容已交给完整复现 Agent 自主下载、调整和验证。
本次复现积分预估
预计消耗
509
最大预留
1323
预计算力时长
约 195 分钟
实际消耗 156 积分,扣除 0 积分;算力运行约 9 分钟,模型 Token 6209368。
论文中的结论可在「研究结论」中查看。
0 / 6 条结论已通过验证
其余仍在验证中
The rollout ablation K=1 vs K=2 tests the core theoretical inflection point on Qwen3-0.6B-Base and can be partially evaluated alongside the main experiment.
已评估 0/10 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 6.67 percentage_points | — | 尚未评估 | |
| 6.67 percentage_points | — | 尚未评估 | |
| 22.8 percentage_points | — | 尚未评估 | |
| 15.15 percentage_points | — | 尚未评估 | |
| 6.67 percentage_points | — | 尚未评估 | |
| 10 percentage_points | — | 尚未评估 | |
| 24.2 percentage_points | — | 尚未评估 | |
| 20.2 percentage_points | — | 尚未评估 | |
| 20.71 percentage_points | — | 尚未评估 | |
| 20.71 percentage_points | — | 尚未评估 |
Training Qwen3-1.7B full-parameter fine-tuning for 1 epoch with 2 online rollouts exceeds the cumulative compute and runtime budget when executed alongside the primary 0.6B verification experiments.
已评估 0/6 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 5.42 percentage_points | — | 尚未评估 | |
| 26.67 percentage_points | — | 尚未评估 | |
| 3.96 percentage_points | — | 尚未评估 | |
| 23.34 percentage_points | — | 尚未评估 | |
| 57.4 percentage_points | — | 尚未评估 | |
| 26.27 percentage_points | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/12 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 0.83 percentage_points | — | 尚未评估 | |
| 6.67 percentage_points | — | 尚未评估 | |
| 1.04 percentage_points | — | 尚未评估 | |
| 10 percentage_points | — | 尚未评估 | |
| 8.72 percentage_points | — | 尚未评估 | |
| 41.86 percentage_points | — | 尚未评估 | |
| 13.04 percentage_points | — | 尚未评估 | |
| 56.52 percentage_points | — | 尚未评估 | |
| 5.14 percentage_points | — | 尚未评估 | |
| 28.89 percentage_points | — | 尚未评估 | |
| 24.2 percentage_points | — | 尚未评估 | |
| 20.2 percentage_points | — | 尚未评估 |
Training Qwen3-4B full-parameter fine-tuning for 1 epoch with 2 online rollouts requires multi-GPU distributed setup and exceeds the single L4 GPU memory and budget limits.
已评估 0/6 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 13.33 percentage_points | — | 尚未评估 | |
| 36.67 percentage_points | — | 尚未评估 | |
| 16.25 percentage_points | — | 尚未评估 | |
| 36.67 percentage_points | — | 尚未评估 | |
| 79 percentage_points | — | 尚未评估 | |
| 32.32 percentage_points | — | 尚未评估 |
Running additional training runs at K=4 across multiple alternative weighting heuristics exceeds the cumulative paper budget allocation.
已评估 0/6 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 23.8 percentage_points | — | 尚未评估 | |
| 15.15 percentage_points | — | 尚未评估 | |
| 21.4 percentage_points | — | 尚未评估 | |
| 15.17 percentage_points | — | 尚未评估 | |
| 24 percentage_points | — | 尚未评估 | |
| 20.71 percentage_points | — | 尚未评估 |
This is a qualitative scope limitation acknowledged by the authors rather than an empirical numerical benchmark claim.