来源:source-paper:PDF p. 12; Appendix E, final paragraph
openai/gpt-5.6-sol
34cbdef9f4
This is an independent reconstruction because the fixed source provides no author implementation or released checkpoints; source-omitted split, initialization, Frey likelihood-variance, annealing, and HMC evaluator details were declared rather than inferred from target values. · Figure 2 does not define the randomness or aggregation behind its reported variance/error bars, so both measured variance fields remain interpretation-unresolved even though their operands and operational definition are fully retained. · The source gives no numerical sufficiency threshold for L=1 or accuracy threshold/ground truth for the marginal estimator's <5-dimensional statement; those control and dimensionality results are descriptive rather than post-hoc pass/fail claims. · Unresolved metric interpretation: exp-fig2-frey-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars.; exp-fig2-mnist-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars. · No author implementation, checkpoints, or original execution logs were available, so this verifies a current reconstruction of the reported procedure rather than the historical provenance of Figure 3. · The original MNIST split and ordering are unreported; consequently, equality of the specific observations denoted by “first 1,000” cannot be established. · Appendix D uses two separate L-draw posterior-sampling stages. The reproduced value 50 is L per stage, yielding 100 retained HMC transitions per datapoint across those stages. · Evaluator initialization, burn-in, and tuning are omitted by the paper. The reconstruction used zero initialization and 25 warmup transitions on 100 points, so full evaluator equivalence is not exact.
openai/gpt-5.6-sol
43a982650f
This is an independent reconstruction because the fixed source provides no author implementation or released checkpoints; source-omitted split, initialization, Frey likelihood-variance, annealing, and HMC evaluator details were declared rather than inferred from target values. · Figure 2 does not define the randomness or aggregation behind its reported variance/error bars, so both measured variance fields remain interpretation-unresolved even though their operands and operational definition are fully retained. · The source gives no numerical sufficiency threshold for L=1 or accuracy threshold/ground truth for the marginal estimator's <5-dimensional statement; those control and dimensionality results are descriptive rather than post-hoc pass/fail claims. · Unresolved metric interpretation: exp-fig2-frey-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars.; exp-fig2-mnist-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars. · This was an independent reconstruction using newly trained models, not evaluation of author-released code or checkpoints. · The reconstruction added 25 warmup transitions on 100 points, zero initialization, and a disclosed step-size adaptation rule; the paper does not report these evaluator details. They do not alter the three assessed operation counts but make the overall protocol only approximately specified. · The reconstructed MNIST training set was deterministically permuted before selecting its first 1,000 points, so exact datapoint identity relative to the original historical run is not established. · Matching the declared procedure shows that it was executed in the reconstruction; it cannot prove that the original Figure 3 computation historically followed it.
已评估 3/3 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
splits: ["train","test"] | 1,000 datapoints | 1,000 datapoints2 次评估 · 展开查看 | 获得支持 |
| 50 samples per datapoint | 50 samples per datapoint2 次评估 · 展开查看 | 获得支持 | |
| 4 leapfrog steps | 4 leapfrog steps2 次评估 · 展开查看 | 获得支持 |