来源:source-paper:PDF p. 4; Algorithm 1 and paragraph following equation (8)
openai/gpt-5.6-sol
775d5cdb79
This is an independent reconstruction because the fixed source provides no author implementation or released checkpoints; source-omitted split, initialization, Frey likelihood-variance, annealing, and HMC evaluator details were declared rather than inferred from target values. · Figure 2 does not define the randomness or aggregation behind its reported variance/error bars, so both measured variance fields remain interpretation-unresolved even though their operands and operational definition are fully retained. · The source gives no numerical sufficiency threshold for L=1 or accuracy threshold/ground truth for the marginal estimator's <5-dimensional statement; those control and dimensionality results are descriptive rather than post-hoc pass/fail claims. · Unresolved metric interpretation: exp-fig2-frey-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars.; exp-fig2-mnist-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars. · Neither supplied measurement quantifies learning quality, gradient variance, or equivalence between L=1 and a higher-sample control. · The paper does not define “sufficient” numerically, so the paired trajectories cannot be converted into a source-comparable pass/fail result. · This was an independent approximate reconstruction with declared choices for source-omitted data splits, initialization, Gaussian likelihood details, and optimization behavior. · The plan was revised during the campaign; the relevant definitions of L and M remained source-grounded, and no post-result sufficiency threshold was introduced. · The rerun demonstrably used L=1 with minibatches of M=100, exactly reproducing the two reported configuration counts. However, those counts only show that the claimed condition was executed; they do not establish that one sample was scientifically “sufficient.” The paper supplies no sufficiency or equivalence criterion, and the reconstruction deliberately made its L=1 versus L=5 results descriptive. Thus the same observations could occur even if L=1 were not sufficient.
openai/gpt-5.6-sol
dd99efae28
This is an independent reconstruction because the fixed source provides no author implementation or released checkpoints; source-omitted split, initialization, Frey likelihood-variance, annealing, and HMC evaluator details were declared rather than inferred from target values. · Figure 2 does not define the randomness or aggregation behind its reported variance/error bars, so both measured variance fields remain interpretation-unresolved even though their operands and operational definition are fully retained. · The source gives no numerical sufficiency threshold for L=1 or accuracy threshold/ground truth for the marginal estimator's <5-dimensional statement; those control and dimensionality results are descriptive rather than post-hoc pass/fail claims. · Unresolved metric interpretation: exp-fig2-frey-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars.; exp-fig2-mnist-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars. · The paper supplies no quantitative definition, equivalence margin, or specified counterfactual for “sufficient”; L=5 is a disclosed reconstruction choice. · Independent learning-rate selection changed the optimizer condition between L=1 and L=5, despite the planned matched-control design. · Only one training seed per arm was retained, so retraining variability was not assessed. · Training-time gradient variability was not retained; repeated final ELBO evaluations measure estimator variability rather than gradient variability. · This was an independent reconstruction with declared dataset-split, initialization, regularization, and implementation choices rather than the unavailable original implementation. · The post-initial plan revision broadened and clarified the campaign after auditing; it did not add a source-grounded sufficiency criterion or repair the optimizer mismatch. · Current execution completed reconstructed MNIST AEVB runs with L=1 and L=5 at M=100 for 200 million presented samples. The recorded trajectories descriptively favored L=1 at the final checkpoint, but this does not verify the paper's undefined notion of “sufficient.” More importantly, the arms independently selected different learning rates (0.02 for L=1 and 0.01 for L=5), violating the protocol requirement to hold the optimizer fixed and materially confounding the comparison. The supplied measurements verify only that L=1 and M=100 were executed; those settings could be observed even if L=1 were insufficient.
已评估 2/2 项
保留同一结论下的完整结果组。已有评估排在前面;单项判断不替代整组结论。展开行查看条件、来源与历次证据。
| 指标与条件 | 论文报告 | 最近可用实测 | 证据判断 |
|---|---|---|---|
| 1 samples per datapoint | 1 samples per datapoint2 次评估 · 展开查看 | 获得支持 | |
| 100 datapoints | 100 datapoints2 次评估 · 展开查看 | 获得支持 |