Source: paper:PDF p.6, datasets and experimental setup
openai/gpt-5.6-sol
5f38f84e17
The paper does not report its historical seed sequence; this reproduction used the predeclared seeds 42-51, so dispersion and exact means need not match the authors' unknown runs. · The bundled fixed graph dictionaries materialize 5,278/4,676/44,327 unique undirected Cora/Citeseer/Pubmed edges, while paper Table 1 prints 5,429/4,732/44,338; node counts and public split sizes match and no substitute data were used. · The fixed official citation.py evaluates in process and does not persist model weights, so per-run observations and complete logs are retained but no checkpoint artifact exists. · The historical hyperopt search was intentionally not rerun under the signed protocol; 60 iterations were verified from paper and fixed tuning.py, and released tuned files were used. · Execution used the available modern CPU environment (PyTorch 2.8.0) rather than the historical PyTorch 1.x-era environment; the fixed code ran without metric failures, with expected isolated-row and sparse-constructor warnings retained in logs. · The historical hyperopt execution was not rerun, so the evidence verifies the paper and repository specification but not that all 60 historical evaluations actually completed. · Because the paper says 100 epochs and also says epoch count was tuned while the fixed README requests 150 for Citeseer, the actual epoch count underlying historical Table 2 remains unresolved; that unresolved condition is the claimed conflict. · The final plan revision followed a one-epoch Cora smoke probe but preceded the full Citeseer arms; it retained both predeclared epoch conditions, and the 150-epoch result was not substituted for the primary arm based on accuracy. · The supplied repository snapshot lacked Git metadata, so its declared commit identity could not be independently checked with git, although the inspected files were the runner-retained fixed snapshot.
openai/gpt-5.6-sol
998889086b
The paper does not report its historical seed sequence; this reproduction used the predeclared seeds 42-51, so dispersion and exact means need not match the authors' unknown runs. · The bundled fixed graph dictionaries materialize 5,278/4,676/44,327 unique undirected Cora/Citeseer/Pubmed edges, while paper Table 1 prints 5,429/4,732/44,338; node counts and public split sizes match and no substitute data were used. · The fixed official citation.py evaluates in process and does not persist model weights, so per-run observations and complete logs are retained but no checkpoint artifact exists. · The historical hyperopt search was intentionally not rerun under the signed protocol; 60 iterations were verified from paper and fixed tuning.py, and released tuned files were used. · Execution used the available modern CPU environment (PyTorch 2.8.0) rather than the historical PyTorch 1.x-era environment; the fixed code ran without metric failures, with expected isolated-row and sparse-constructor warnings retained in logs. · The historical 60-trial hyperparameter search was intentionally not rerun, so the evidence verifies its paper/repository specification but not that the authors completed those exact historical evaluations or how the released tuned files were generated. · The conflicting 100- and 150-epoch sources do not reveal which condition actually produced the historical Citeseer Table 2 value; they establish ambiguity rather than resolving it. · The repository snapshot lacked .git metadata, so its association with the declared commit could not be independently checked from repository history, although the exact retained README and implementation were inspectable. · Modern CPU/PyTorch execution verifies the device-independent configuration quantities assessed here, not full equivalence to the authors' historical software environment.
openai/gpt-5.6-sol
399d25238c
The paper does not report its historical seed sequence; this reproduction used the predeclared seeds 42-51, so dispersion and exact means need not match the authors' unknown runs. · The bundled fixed graph dictionaries materialize 5,278/4,676/44,327 unique undirected Cora/Citeseer/Pubmed edges, while paper Table 1 prints 5,429/4,732/44,338; node counts and public split sizes match and no substitute data were used. · The fixed official citation.py evaluates in process and does not persist model weights, so per-run observations and complete logs are retained but no checkpoint artifact exists. · The historical hyperopt search was intentionally not rerun under the signed protocol; 60 iterations were verified from paper and fixed tuning.py, and released tuned files were used. · Execution used the available modern CPU environment (PyTorch 2.8.0) rather than the historical PyTorch 1.x-era environment; the fixed code ran without metric failures, with expected isolated-row and sparse-constructor warnings retained in logs. · The historical hyperparameter searches were not rerun, so the evidence verifies the paper and repository configuration but not completion or provenance of the original searches. · Neither source resolves whether the historical Citeseer Table 2 result actually used 100 or 150 epochs; the evidence establishes the conflict, not one definitive historical condition. · The repository snapshot lacked .git metadata, although the retained files were identified as the fixed official implementation snapshot. · Execution used a modern CPU/PyTorch environment rather than the historical environment; this does not materially affect comparison of epoch counts, learning rate, or the static max_evals setting.
3/3 assessed
The complete result group stays with its claim. Assessed rows come first; their judgments do not replace the whole-claim conclusion. Expand a row for conditions, sources, and evidence history.
| Metric and conditions | Reported | Latest observation | Evidence judgment |
|---|---|---|---|
| 100 epochs | 100 epochs3 assessments · expand | Supported | |
| 0.2 unitless | 0.2 unitless3 assessments · expand | Supported | |
| 60 iterations | 60 iterations3 assessments · expand | Supported |