Start with coverage and measured results; open a run only when you need evidence or technical details.
RunMatchRepeat
0/9
claims supported by evidence
0
Supported
0
Challenged or mixed
0
Contradicted
0
Inconclusive
9
Not assessed
Run history
Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.
No reproduction runs yet
Runs will appear here as the paper's experiment plans are executed.
Claim–experiment reproduction matrix
See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions. There are also 4 claims with no independent reproduction scheduled in this plan.
A successful execution does not by itself validate a paper claim
Target state says whether the platform completed the work. The scientific conclusion is determined only by immutable evidence and Assessments. Resource shortages and platform failures are never presented as scientific contradictions.
This task predates target-level records. The states below are conservative projections; no attempts, resource decisions, or evidence links are invented.
For Qwen3-4B-Math→Qwen3-4B mathematical distillation, Table 1 reports γOPD as the best method on both average accuracy (69.49) and average pass rate (82.41), exceeding the teacher averages (68.10 and 79.81) and the strongest OPD-baseline averages.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Vanilla 4B-to-4B distillation comparison
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Without RBM in Qwen3-4B-Math→Qwen3-1.7B training, γ=1 produces much larger gradient norms, rapid policy-entropy collapse, and inferior AIME24 validation accuracy. Local OPD (γ=0) and discounted variants improve steadily, with γ=0.99 reported as best and better than γ=0.9.
Information insufficientclaim-gamma-sensitivity-figure40 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Profiling on the main 4B student/teacher and 32-GPU setup reports total additional γOPD work of 610.32 ms/step and 0.375 MB, equal to 0.1064% time and 0.00046% memory relative to a vanilla OPD update.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Under joint math-and-code multi-teacher distillation into Qwen3-4B, γOPD records TotalAvg 66.00, the highest table value; it leads HumanEval+ and LiveCodeBench among code tasks and exceeds the teacher on the paper's average math and code comparisons, but not every individual benchmark or metric.
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Multi-teacher math and code reference controls
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
For the displayed incorrect response, local A_t^(0) penalizes tokens only weakly related to the actual reasoning error, whereas A_t^(0.99) concentrates stronger negative credit on the erroneous formula derivation without excessively penalizing the final end-of-sequence token.
Information insufficientclaim-token-credit-incorrect-example1 plan0 runs
Scientific conclusionNot assessed
Temporal-credit and reward-mixing component ablation
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
In vanilla-distillation training curves, the paper reports that γOPD maintains the highest verifiable reward, stable absolute OPD advantage, and the smallest gradient-norm fluctuations; its response length becomes shorter and more stable than most baselines. TOPD is the exception on length because it truncates the distillation signal, and its verifiable reward remains mostly below zero.
Information insufficientclaim-training-dynamics-vanilla1 plan0 runs
Scientific conclusionNot assessed
Vanilla 4B-to-4B distillation comparison
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
For Qwen3-4B-Math→Qwen3-1.7B mathematical distillation, every trained student remains below the teacher on the two reported averages, while γOPD gives the best trained-student average accuracy (56.51) and average pass rate (71.22).
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Math reference-model controls for Table 1
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
For the displayed correct response, local A_t^(0) gives sparse supervision and noticeable negative credit to only a few tokens, whereas normalized A_t^(0.99) propagates later information, smooths the signal, assigns stronger positive credit to the key mathematical derivation, and mildly penalizes redundant text.
Information insufficientclaim-token-credit-correct-example1 plan0 runs
Scientific conclusionNot assessed
Temporal-credit and reward-mixing component ablation
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
On AIME24 and AIME25, temporal discounting alone raises average accuracy by 2.04 points over vanilla OPD; naive reward mixing adds only 0.05 point beyond discounting, while the full discounting-plus-mixing-plus-normalization configuration reaches 57.56 average accuracy, a 3.91-point gain. Mixing plus normalization without discounting gains 2.00 points, supporting complementary contributions.