Source: source-paper:PDF p. 12; Appendix E, final paragraph
openai/gpt-5.6-sol
34cbdef9f4
This is an independent reconstruction because the fixed source provides no author implementation or released checkpoints; source-omitted split, initialization, Frey likelihood-variance, annealing, and HMC evaluator details were declared rather than inferred from target values. · Figure 2 does not define the randomness or aggregation behind its reported variance/error bars, so both measured variance fields remain interpretation-unresolved even though their operands and operational definition are fully retained. · The source gives no numerical sufficiency threshold for L=1 or accuracy threshold/ground truth for the marginal estimator's <5-dimensional statement; those control and dimensionality results are descriptive rather than post-hoc pass/fail claims. · Unresolved metric interpretation: exp-fig2-frey-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars.; exp-fig2-mnist-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars. · No author implementation, checkpoints, or original execution logs were available, so this verifies a current reconstruction of the reported procedure rather than the historical provenance of Figure 3. · The original MNIST split and ordering are unreported; consequently, equality of the specific observations denoted by “first 1,000” cannot be established. · Appendix D uses two separate L-draw posterior-sampling stages. The reproduced value 50 is L per stage, yielding 100 retained HMC transitions per datapoint across those stages. · Evaluator initialization, burn-in, and tuning are omitted by the paper. The reconstruction used zero initialization and 25 warmup transitions on 100 points, so full evaluator equivalence is not exact.
openai/gpt-5.6-sol
43a982650f
This is an independent reconstruction because the fixed source provides no author implementation or released checkpoints; source-omitted split, initialization, Frey likelihood-variance, annealing, and HMC evaluator details were declared rather than inferred from target values. · Figure 2 does not define the randomness or aggregation behind its reported variance/error bars, so both measured variance fields remain interpretation-unresolved even though their operands and operational definition are fully retained. · The source gives no numerical sufficiency threshold for L=1 or accuracy threshold/ground truth for the marginal estimator's <5-dimensional statement; those control and dimensionality results are descriptive rather than post-hoc pass/fail claims. · Unresolved metric interpretation: exp-fig2-frey-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars.; exp-fig2-mnist-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars. · This was an independent reconstruction using newly trained models, not evaluation of author-released code or checkpoints. · The reconstruction added 25 warmup transitions on 100 points, zero initialization, and a disclosed step-size adaptation rule; the paper does not report these evaluator details. They do not alter the three assessed operation counts but make the overall protocol only approximately specified. · The reconstructed MNIST training set was deterministically permuted before selecting its first 1,000 points, so exact datapoint identity relative to the original historical run is not established. · Matching the declared procedure shows that it was executed in the reconstruction; it cannot prove that the original Figure 3 computation historically followed it.
3/3 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
splits: ["train","test"] | 1,000 datapoints | 1,000 datapoints2 assessments · expand | Supported |
| 50 samples per datapoint | 50 samples per datapoint2 assessments · expand | Supported | |
| 4 leapfrog steps | 4 leapfrog steps2 assessments · expand | Supported |