Metric and scale shape the reconstruction
In the reconstructed MNIST curves, AEVB has higher final train and test ELBO than wake-sleep at all five latent dimensions. Yet the final test marginal-likelihood leader changes from wake-sleep with 1,000 training examples to AEVB with 50,000, while MCEM is lowest in both; finite probes at 3D, 5D, and 10D cannot validate the paper’s low-dimensional reliability boundary without ground truth.
View values and explanation
AEVB leads wake-sleep in final train and test ELBO at all five recorded dimensions; final test marginal likelihood is led by wake-sleep with 1,000 examples and by AEVB with 50,000, while the dimensionality probe cannot validate an accuracy boundary.
- Final test ELBO gap at 20D (AEVB minus wake-sleep)
- 4.337466
- Final test ELBO gap at 200D (AEVB minus wake-sleep)
- 4.443145
- Final test mean log marginal likelihood for AEVB, 50,000 examples
- -148.5871
- Final test mean log marginal likelihood for wake-sleep, 1,000 examples
- -159.2733
- Paper-reported strict dimension boundary (not validated by probe)
- 5