来源:source-paper:PDF p. 12; Appendix E, first paragraph
openai/gpt-5.6-sol
f00ff44720
This is an independent reconstruction because the fixed source provides no author implementation or released checkpoints; source-omitted split, initialization, Frey likelihood-variance, annealing, and HMC evaluator details were declared rather than inferred from target values. · Figure 2 does not define the randomness or aggregation behind its reported variance/error bars, so both measured variance fields remain interpretation-unresolved even though their operands and operational definition are fully retained. · The source gives no numerical sufficiency threshold for L=1 or accuracy threshold/ground truth for the marginal estimator's <5-dimensional statement; those control and dimensionality results are descriptive rather than post-hoc pass/fail claims. · Unresolved metric interpretation: exp-fig2-frey-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars.; exp-fig2-mnist-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars. · No author implementation, configuration, checkpoint, or original execution trace was available; this was an independent CiteArk reconstruction. · The paper specifies an automatically tuned 90% target and an annealing schedule but not their equations. The reconstruction chose clipped multiplicative HMC adaptation and eta_t=eta_0/sqrt(1+t/10000). · The revised plan added explicit adaptation, annealing, initialization, split, and device choices after the initial plan; these are disclosed reconstruction decisions rather than source-verified details. · Execution used one declared seed on an NVIDIA L4, not a documented original hardware/software environment. · The captured current execution shows that the independent reconstruction completed the 1,000-example AEVB, wake-sleep, and MCEM trajectories and evaluation. Its MCEM loop used 10 leapfrog steps, feedback toward a 90% acceptance target, and five decoder updates per cycle; after 60 million presented samples, cumulative acceptance was 89.99%. The reconstruction also used Adagrad with inverse-square-root annealing for all three methods. Thus the stated recipe and all supplied values were operationally reproduced. However, these settings were deliberately reconstructed from the paper without author code or checkpoints. The same matching execution could occur even if the historical baseline had not actually used them, so it cannot independently verify the paper's historical implementation claim.
openai/gpt-5.6-sol
1bbc2ae736
This is an independent reconstruction because the fixed source provides no author implementation or released checkpoints; source-omitted split, initialization, Frey likelihood-variance, annealing, and HMC evaluator details were declared rather than inferred from target values. · Figure 2 does not define the randomness or aggregation behind its reported variance/error bars, so both measured variance fields remain interpretation-unresolved even though their operands and operational definition are fully retained. · The source gives no numerical sufficiency threshold for L=1 or accuracy threshold/ground truth for the marginal estimator's <5-dimensional statement; those control and dimensionality results are descriptive rather than post-hoc pass/fail claims. · Unresolved metric interpretation: exp-fig2-frey-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars.; exp-fig2-mnist-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars. · This was an independent reconstruction, not execution of author-released code or checkpoints. · The revised plan introduced explicit multiplicative HMC adaptation and inverse-square-root Adagrad annealing formulas after source/environment audit. The paper specifies their purpose but not their equations, so exact historical schedule equivalence is unverifiable. · The supplied numeric measurements cover the MCEM counts and acceptance target; use of Adagrad and annealing across all three algorithms was established from inspected implementation and trajectory configurations rather than a separate supplied measurement.
已评估 3/3 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 10 leapfrog steps | 10 leapfrog steps2 次评估 · 展开查看 | 获得支持 | |
| 90% | 90%2 次评估 · 展开查看 | 获得支持 | |
| 5 update steps | 5 update steps2 次评估 · 展开查看 | 获得支持 |