Figure 5 presents random samples from learned MNIST generative models at each of four latent-space dimensionalities: 2, 5, 10, and 20.
This claim has no executable experiment plan yet.
Start with coverage and measured results; open a run only when you need evidence or technical details.
2/13
claims supported by evidence
2
Supported
0
Challenged or mixed
0
Contradicted
3
Inconclusive
8
Not assessed
The Figure 3 marginal-likelihood estimates were computed on the first 1,000 datapoints of both the training and test sets, using 50 posterior samples per datapoint from HMC with 4 leapfrog steps.
Reported
1000 datapoints
Observed
1000 datapoints
Difference 0
Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.
Failed paths grouped by cause — check them before reproducing.
See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions. There are also 3 claims with no independent reproduction scheduled in this plan.
A successful execution does not by itself validate a paper claim
Target state says whether the platform completed the work. The scientific conclusion is determined only by immutable evidence and Assessments. Resource shortages and platform failures are never presented as scientific contradictions.
Figure 5 presents random samples from learned MNIST generative models at each of four latent-space dimensionalities: 2, 5, 10, and 20.
This claim has no executable experiment plan yet.
The authors report that the relative performance of the compared algorithms was not very sensitive to the chosen hidden-layer widths.
This claim has no executable experiment plan yet.
The paper states that a learned two-dimensional AEVB encoder can project MNIST and Frey Face observations into a low-dimensional latent space; Figure 4 itself shows grids of decoded model outputs p_theta(x|z) at inverse-Gaussian-CDF-transformed latent coordinates, not plotted encoded datapoint projections.
This claim has no executable experiment plan yet.
In Figure 3 with 1,000 MNIST training examples, the AEVB training marginal-likelihood curve rises faster than MCEM and the two approach similar displayed training endpoints; the displayed MCEM test curve finishes above AEVB, while wake-sleep has lower train and test curves than both near the end.
MNIST marginal-likelihood comparison with 1,000 training examples
Next step
No automatic retry; a person must decide what to do next.
Claim and experiment binding established
Execution failed
No target-level attempt recorded
No linked immutable Artifact yet
No Assessment yet
Legacy task without target-level resource requirements
Increasing the number of latent variables, including apparently superfluous variables, did not produce overfitting in the Figure 2 experiments; the paper attributes this to regularization by the variational lower bound.
MNIST AEVB versus wake-sleep lower-bound curves
Next step
No automatic retry; a person must decide what to do next.
Claim and experiment binding established
Execution failed
No target-level attempt recorded
No linked immutable Artifact yet
No Assessment yet
Legacy task without target-level resource requirements
Frey Face AEVB versus wake-sleep lower-bound curves
Next step
No automatic retry; a person must decide what to do next.
Claim and experiment binding established
Execution failed
No target-level attempt recorded
No linked immutable Artifact yet
No Assessment yet
Legacy task without target-level resource requirements
In Figure 3 with 50,000 MNIST training examples, AEVB has the highest displayed train and test marginal-likelihood curves, wake-sleep is lower, and MCEM improves much later and remains lower over the plotted training range.
MNIST marginal-likelihood comparison with 50,000 training examples
Next step
No automatic retry; a person must decide what to do next.
Claim and experiment binding established
Execution failed
No target-level attempt recorded
No linked immutable Artifact yet
No Assessment yet
Legacy task without target-level resource requirements
The Figure 2 computations took approximately 20-40 minutes per million training samples on an Intel Xeon CPU operating at an effective 40 GFLOPS.
Figure 2 runtime under the reported Intel Xeon reference condition
An experiment plan exists, but the current processing task has no matching execution target.
Across every lower-bound comparison shown in Figure 2, AEVB converged considerably faster than wake-sleep and reached a better variational-lower-bound solution on both training and test curves.
MNIST AEVB versus wake-sleep lower-bound curves
Next step
No automatic retry; a person must decide what to do next.
Claim and experiment binding established
Execution failed
No target-level attempt recorded
No linked immutable Artifact yet
No Assessment yet
Legacy task without target-level resource requirements
Frey Face AEVB versus wake-sleep lower-bound curves
Next step
No automatic retry; a person must decide what to do next.
Claim and experiment binding established
Execution failed
Experiment execution · Failed
No linked immutable Artifact yet
No Assessment yet
Legacy task without target-level resource requirements
For the Figure 2 lower-bound curves, the paper reports that estimator variance was below 1 and omits it from the plot.
MNIST AEVB versus wake-sleep lower-bound curves
Next step
The execution path is complete; inspect the scientific Assessment next.
Claim and experiment binding established
Evidence published
Evidence publication · Completed
1 immutable Artifact
Assessed but inconclusive
Legacy task without target-level resource requirements
Frey Face AEVB versus wake-sleep lower-bound curves
Next step
The execution path is complete; inspect the scientific Assessment next.
Claim and experiment binding established
Evidence published
Experiment execution · Running
1 immutable Artifact
Assessed but inconclusive
Legacy task without target-level resource requirements
The Figure 3 marginal-likelihood estimates were computed on the first 1,000 datapoints of both the training and test sets, using 50 posterior samples per datapoint from HMC with 4 leapfrog steps.
MNIST marginal-likelihood comparison with 1,000 training examples
Next step
The execution path is complete; inspect the scientific Assessment next.
Claim and experiment binding established
Evidence published
Evidence publication · Completed
1 immutable Artifact
Evidence supports the claim
Legacy task without target-level resource requirements
MNIST marginal-likelihood comparison with 50,000 training examples
Next step
The execution path is complete; inspect the scientific Assessment next.
Claim and experiment binding established
Evidence published
Evidence publication · Completed
1 immutable Artifact
Evidence supports the claim
Legacy task without target-level resource requirements
The MCEM baseline used 10 HMC leapfrog steps with automatically tuned step size targeting a 90% acceptance rate, followed by 5 parameter-weight update steps; all compared algorithms used Adagrad step sizes with an annealing schedule.
MNIST marginal-likelihood comparison with 1,000 training examples
Next step
The execution path is complete; inspect the scientific Assessment next.
Claim and experiment binding established
Evidence published
Evidence publication · Completed
1 immutable Artifact
Evidence supports the claim
Legacy task without target-level resource requirements
MNIST marginal-likelihood comparison with 50,000 training examples
Next step
The execution path is complete; inspect the scientific Assessment next.
Claim and experiment binding established
Evidence published
Evidence publication · Completed
1 immutable Artifact
Evidence supports the claim
Legacy task without target-level resource requirements
The marginal-likelihood estimator is reported to give good estimates only when the sampled latent space is very low-dimensional and enough samples are used; the experiment therefore used 3 latent variables, while estimates at higher dimensionality became unreliable.
MNIST marginal-likelihood comparison with 1,000 training examples
Next step
The execution path is complete; inspect the scientific Assessment next.
Claim and experiment binding established
Evidence published
Evidence publication · Completed
1 immutable Artifact
Assessed but inconclusive
Legacy task without target-level resource requirements
MNIST marginal-likelihood comparison with 50,000 training examples
Next step
The execution path is complete; inspect the scientific Assessment next.
Claim and experiment binding established
Evidence published
Evidence publication · Completed
1 immutable Artifact
Assessed but inconclusive
Legacy task without target-level resource requirements
Marginal-likelihood estimator sensitivity to latent dimensionality
Next step
The execution path is complete; inspect the scientific Assessment next.
Claim and experiment binding established
Evidence published
Evidence publication · Completed
1 immutable Artifact
Assessed but inconclusive
Legacy task without target-level resource requirements
In the reported experiments, one latent sample per datapoint was found sufficient when the minibatch contained about 100 datapoints.
MNIST L=1 latent-sample sufficiency at minibatch 100
Next step
The execution path is complete; inspect the scientific Assessment next.
Claim and experiment binding established
Evidence published
Evidence publication · Completed
1 immutable Artifact
Assessed but inconclusive
Legacy task without target-level resource requirements
Frey Face L=1 latent-sample sufficiency at minibatch 100
Next step
The execution path is complete; inspect the scientific Assessment next.
Claim and experiment binding established
Evidence published
Evidence publication · Completed
1 immutable Artifact
Assessed but inconclusive
Legacy task without target-level resource requirements
Executed locally and uploaded by users. The platform verifies file signatures; conclusions come from the uploaded runs.
No community runs yet.