7 experiments
Execution agent’s result notes
Completed all 28 distinct training states behind the 42-workload campaign on the allocated NVIDIA L4, reusing shared checkpoints and running the full assigned protocol/count validations, Figure 2 variance aggregates, Figure 3 HMC evaluations, L=1/L=5 controls, and z=3/5/10 dimensionality study. Under the declared 16-repeat full-split L=1 ELBO variance definition, MNIST's maximum was 0.033282 (<1) while Frey's was 866.562; exact equivalence to the paper's undefined error bars is unresolved. Final MNIST Figure 2 curves favored AEVB over wake-sleep at all five dimensions on both splits. The reconstructed Figure 3 HMC estimates favored wake-sleep for the 1k condition, AEVB on 50k test, and placed MCEM below both in both conditions.
For the Figure 2 lower-bound curves, the paper reports that estimator variance was below 1 and omits it from the plot.
Metric: upper bound on omitted lower-bound estimator variance · not specified
Data and sources · 3
| Source | Conditions | Metric | Value | Assessment and limits |
|---|---|---|---|---|
| Papersource-paper:PDF p. 7; Figure 2 captionclaim-004/measurement-003/paper | — | upper bound on omitted lower-bound estimator variance | 1 not specified | reported |
| Reproductionsource-paper:PDF p. 7; Figure 2 captionclaim-004/measurement-003/urn:citeark:assessment:9e0c7821f3be3472c8cf6d57e331c70ae98066ee02eee121e74753634f627bab | — | upper bound on omitted lower-bound estimator variance | 866.562336108187 not specified | inconclusive |
| Reproductionsource-paper:PDF p. 7; Figure 2 captionclaim-004/measurement-003/urn:citeark:assessment:7211ece5966ebdeb72161831b5d8b0414d88264d19fa5bcdd420e86f0cb656f4 | — | upper bound on omitted lower-bound estimator variance | 0.03328223825186328 not specified | inconclusive |
Comparison conditions
This is an independent reconstruction because the fixed source provides no author implementation or released checkpoints; source-omitted split, initialization, Frey likelihood-variance, annealing, and HMC evaluator details were declared rather than inferred from target values.
Figure 2 does not define the randomness or aggregation behind its reported variance/error bars, so both measured variance fields remain interpretation-unresolved even though their operands and operational definition are fully retained.
The source gives no numerical sufficiency threshold for L=1 or accuracy threshold/ground truth for the marginal estimator's <5-dimensional statement; those control and dimensionality results are descriptive rather than post-hoc pass/fail claims.
Unresolved metric interpretation: exp-fig2-frey-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars.; exp-fig2-mnist-lower-bound::measurement-003: Figure 2 defines the plotted quantity as estimated average variational lower bound per datapoint and equations (7)-(8) define its stochastic estimator, but the caption does not state whether its variance is across latent draws, minibatches, smoothing windows, or runs. The reconstruction therefore reports the maximum unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations without asserting exact equivalence to the omitted source error bars.
The paper does not define the randomness or aggregation underlying the Figure 2 variance statement.
The independent reconstruction uses source-unreported choices for the Frey split, Gaussian observation variance, initialization interpretation, weight decay, and other training details.
The revised variance definition was adopted after the source target had been read; although the retained maximum is unfavorable rather than selected to match, source equivalence remains unverifiable.
The experiment establishes only the variance of the reconstructed full-split ELBO evaluation, not the variance used for the paper’s omitted plot uncertainty.
The caption confirms that the paper reported estimator variance <1 and omitted it. The current reconstruction executed the full Frey workload and obtained a maximum variance of 866.562, but this was the unbiased variance across 16 repeated full-split mean analytic-KL ELBO evaluations, maximized over 64 method/dimension/checkpoint/split cells. The paper does not identify whether its variance was over latent draws, minibatches, runs, smoothing windows, or another quantity. Therefore the large discrepancy cannot validly support, challenge, or contradict the paper’s numerical claim. It only establishes high variance under the reconstruction’s disclosed definition.
The paper does not define the estimator-variance sampling axis or aggregation used for Figure 2.
The reconstruction used one training seed and independently chosen split, initialization interpretation, weight decay, and stopping details because the historical settings were incompletely reported.
The reported value 1 is a strict upper threshold, not a point estimate; absolute or relative difference from 1 is therefore not an appropriate reproduction criterion.
The paper’s omission of error bars prevents direct comparison with curve-specific historical variance values.
The paper caption explicitly states that estimator variance was “small (< 1) and omitted.” The reconstruction executed all five MNIST AEVB and wake-sleep trajectories through 200 million presented samples and computed a maximum variance of 0.033282 across 100 retained cells. However, it defined variance as the unbiased variance across 16 repeated full-split mean L=1 ELBO evaluations. The paper does not specify whether its variance was across latent draws, minibatches, smoothing windows, or runs, nor its aggregation across curves and checkpoints. Thus 0.033282 < 1 is not a comparable reproduction of the paper’s quantity, although it establishes low variance under the reconstruction’s declared definition.