浏览 ArkGraph,选择本次要执行的步骤。
论文中的结论可在「研究结论」中查看。
0 / 13 条结论已通过验证
其余仍在验证中
实验方案已生成,还没有运行记录
已评估 0/102 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
method: Student · dataset: AIME24 | 22.6% | — | 尚未评估 |
method: Student · dataset: AIME24 | 60% | — | 尚未评估 |
method: Student · dataset: AIME25 | 20.83% | — | 尚未评估 |
method: Student · dataset: AIME25 | 33.33% | — | 尚未评估 |
method: Student · dataset: AMC23 | 60.16% | — | 尚未评估 |
method: Student · dataset: AMC23 | 92.5% | — | 尚未评估 |
method: Student · dataset: MATH500 | 67.4% | — | 尚未评估 |
method: Student · dataset: MATH500 | 72.6% | — | 尚未评估 |
method: Student · dataset: Average | 42.75% | — | 尚未评估 |
method: Student · dataset: Average | 64.61% | — | 尚未评估 |
method: Teacher · dataset: AIME24 | 57.4% | — | 尚未评估 |
method: Teacher · dataset: AIME24 | 76.67% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/12 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
operation: Discounted temporal credit | 609.85 milliseconds per step | — | 尚未评估 |
operation: Discounted temporal credit | 0.063 megabytes | — | 尚未评估 |
operation: Bounded advantage shaping | 0.42 milliseconds per step | — | 尚未评估 |
operation: Bounded advantage shaping | 0.125 megabytes | — | 尚未评估 |
operation: Verifier-reward mixing | 0.04 milliseconds per step | — | 尚未评估 |
operation: Verifier-reward mixing | 0.188 megabytes | — | 尚未评估 |
operation: Total additional overhead | 610.32 milliseconds per step | — | 尚未评估 |
operation: Total additional overhead | 0.375 megabytes | — | 尚未评估 |
operation: Vanilla OPD update (reference) | 573,669 milliseconds per step | — | 尚未评估 |
operation: Vanilla OPD update (reference) | 81,254.4 megabytes | — | 尚未评估 |
| 0.1064% | — | 尚未评估 | |
| 0.0005% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/65 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
method: Student · dataset: AIME24 | 22.6% | — | 尚未评估 |
method: Student · dataset: AIME25 | 20.83% | — | 尚未评估 |
method: Student · dataset: AMC23 | 60.16% | — | 尚未评估 |
method: Student · dataset: MATH500 | 67.4% | — | 尚未评估 |
method: Student · dataset: HumanEval+ | 75.61% | — | 尚未评估 |
method: Student · dataset: MBPP+ | 61.9% | — | 尚未评估 |
method: Student · dataset: LiveCodeBench v6 | 17.14% | — | 尚未评估 |
method: Student · dataset: TotalAvg | 47.15% | — | 尚未评估 |
method: Teacher · dataset: AIME24 | 57.4% | — | 尚未评估 |
method: Teacher · dataset: AIME25 | 51.67% | — | 尚未评估 |
method: Teacher · dataset: AMC23 | 93.52% | — | 尚未评估 |
method: Teacher · dataset: MATH500 | 69.8% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/102 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
method: Student · dataset: AIME24 | 11.77% | — | 尚未评估 |
method: Student · dataset: AIME24 | 46.67% | — | 尚未评估 |
method: Student · dataset: AIME25 | 8.33% | — | 尚未评估 |
method: Student · dataset: AIME25 | 20% | — | 尚未评估 |
method: Student · dataset: AMC23 | 39.84% | — | 尚未评估 |
method: Student · dataset: AMC23 | 90% | — | 尚未评估 |
method: Student · dataset: MATH500 | 59.2% | — | 尚未评估 |
method: Student · dataset: MATH500 | 72.8% | — | 尚未评估 |
method: Student · dataset: Average | 29.79% | — | 尚未评估 |
method: Student · dataset: Average | 57.37% | — | 尚未评估 |
method: Teacher · dataset: AIME24 | 57.4% | — | 尚未评估 |
method: Teacher · dataset: AIME24 | 76.67% | — | 尚未评估 |
实验方案已生成,还没有运行记录
已评估 0/19 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
dataset: AIME24 | 56.46% | — | 尚未评估 |
dataset: AIME25 | 50.83% | — | 尚未评估 |
dataset: AIME24 and AIME25 | 53.65% | — | 尚未评估 |
dataset: AIME24 | 58.29% | — | 尚未评估 |
dataset: AIME25 | 53.08% | — | 尚未评估 |
dataset: AIME24 and AIME25 | 55.69% | — | 尚未评估 |
| 2.04 percentage points | — | 尚未评估 | |
dataset: AIME24 | 58.81% | — | 尚未评估 |
dataset: AIME25 | 52.67% | — | 尚未评估 |
dataset: AIME24 and AIME25 | 55.74% | — | 尚未评估 |
| 2.09 percentage points | — | 尚未评估 | |
dataset: AIME24 | 57.55% | — | 尚未评估 |
The four gamma conditions require regenerated training trajectories, but the paper provides no numeric curve targets or safe complete categorical rule.
The exact illustrated response is retained in the ablation objective as qualitative source context; no scalar target is invented.
The full training package retains these curve claims as source context, but automatic verification cannot safely bind them.
The exact illustrated response is retained in the ablation objective as qualitative source context; no scalar target is invented.
This is protocol context rather than an independently testable result. Its reported settings are carried into each applicable training/evaluation objective, with missing revisions, seeds, stopping schedule, dtype, and multi-domain sampling ratio left explicit.
This is method definition and algebraic context rather than an independently reported empirical result. Its equations are operational prerequisites for the planned reconstruction packages; hand checks cannot substitute for the represented empirical comparisons.
The claim is a theorem under a stated moment assumption, not an empirical measurement. The fixed proof can be assessed mathematically; numerical examples would not reproduce it.
This is a stated scope limitation and future-work boundary, not an empirical result requiring a reproduction package.