Source: source-paper:PDF p. 12; Appendix E, first paragraph
openai/gpt-5.6-sol
f00ff44720
This is an independent reconstruction because the fixed source provides no author implementation or released checkpoints; source-omitted split, initialization, Frey likelihood-variance, annealing, and HMC evaluator details were declared rather than inferred from target values. · Figure 2 does not define the randomness or aggregation behind its reported variance/error bars, so both measured variance fields remain interpretation-unresolved even though their operands and operational definition are fully retained. · The source gives no numerical sufficiency threshold for L=1 or accuracy threshold/ground truth for the marginal estimator's <5-dimensional statement; those control and dimensionality results are descriptive rather than post-hoc pass/fail claims. · Unresolved metric interpretation: exp-fig2-frey-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars.; exp-fig2-mnist-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars. · No author implementation, configuration, checkpoint, or original execution trace was available; this was an independent CiteArk reconstruction. · The paper specifies an automatically tuned 90% target and an annealing schedule but not their equations. The reconstruction chose clipped multiplicative HMC adaptation and eta_t=eta_0/sqrt(1+t/10000). · The revised plan added explicit adaptation, annealing, initialization, split, and device choices after the initial plan; these are disclosed reconstruction decisions rather than source-verified details. · Execution used one declared seed on an NVIDIA L4, not a documented original hardware/software environment. · The captured current execution shows that the independent reconstruction completed the 1,000-example AEVB, wake-sleep, and MCEM trajectories and evaluation. Its MCEM loop used 10 leapfrog steps, feedback toward a 90% acceptance target, and five decoder updates per cycle; after 60 million presented samples, cumulative acceptance was 89.99%. The reconstruction also used Adagrad with inverse-square-root annealing for all three methods. Thus the stated recipe and all supplied values were operationally reproduced. However, these settings were deliberately reconstructed from the paper without author code or checkpoints. The same matching execution could occur even if the historical baseline had not actually used them, so it cannot independently verify the paper's historical implementation claim.
openai/gpt-5.6-sol
1bbc2ae736
This is an independent reconstruction because the fixed source provides no author implementation or released checkpoints; source-omitted split, initialization, Frey likelihood-variance, annealing, and HMC evaluator details were declared rather than inferred from target values. · Figure 2 does not define the randomness or aggregation behind its reported variance/error bars, so both measured variance fields remain interpretation-unresolved even though their operands and operational definition are fully retained. · The source gives no numerical sufficiency threshold for L=1 or accuracy threshold/ground truth for the marginal estimator's <5-dimensional statement; those control and dimensionality results are descriptive rather than post-hoc pass/fail claims. · Unresolved metric interpretation: exp-fig2-frey-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars.; exp-fig2-mnist-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars. · This was an independent reconstruction, not execution of author-released code or checkpoints. · The revised plan introduced explicit multiplicative HMC adaptation and inverse-square-root Adagrad annealing formulas after source/environment audit. The paper specifies their purpose but not their equations, so exact historical schedule equivalence is unverifiable. · The supplied numeric measurements cover the MCEM counts and acceptance target; use of Adagrad and annealing across all three algorithms was established from inspected implementation and trajectory configurations rather than a separate supplied measurement.
3/3 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 10 leapfrog steps | 10 leapfrog steps2 assessments · expand | Supported | |
| 90% | 90%2 assessments · expand | Supported | |
| 5 update steps | 5 update steps2 assessments · expand | Supported |