1 experiments
Execution agent’s result notes
Executed the fixed official CPU SGC entrypoint for seeds 42-51 on Cora, Citeseer, and Pubmed at the paper-bound 100 epochs, plus the separately declared Citeseer 150-epoch README sensitivity arm. The ten-seed primary mean test accuracies were 81.05%, 71.20%, and 78.96%, respectively; the Citeseer sensitivity mean was 71.85% and was not used to replace the primary result. Direct source/runtime evidence also resolved 100 epochs, learning rate 0.2, and 60 historical hyperopt iterations for every dataset while reusing the released tuned weight-decay files as required.
On the Citeseer public split, SGC reports 71.9% mean test accuracy, about one percentage point above the authors' GCN result; Table 2 reports all listed baselines.
Metric: test accuracy · percent
Data and sources · 2
| Source | Conditions | Metric | Value | Assessment and limits |
|---|---|---|---|---|
| Paperpaper:PDF p.6, Table 2claim-table2-citeseer/m-t2-cit-sgc/paper | reportedPlusMinus: 0.1 | test accuracy | 71.9 percent | reported |
| Reproductionpaper:PDF p.6, Table 2claim-table2-citeseer/m-t2-cit-sgc/urn:citeark:assessment:a9dc3cac47c652216efd64e89f2a553675555d7011d970f8682006a21620b8a4 | reportedPlusMinus: 0.1 | test accuracy | 71.2 percent | challenged |
Comparison conditions
The paper does not report its historical seed sequence; this reproduction used the predeclared seeds 42-51, so dispersion and exact means need not match the authors' unknown runs.
The bundled fixed graph dictionaries materialize 5,278/4,676/44,327 unique undirected Cora/Citeseer/Pubmed edges, while paper Table 1 prints 5,429/4,732/44,338; node counts and public split sizes match and no substitute data were used.
The fixed official citation.py evaluates in process and does not persist model weights, so per-run observations and complete logs are retained but no checkpoint artifact exists.
The historical hyperopt search was intentionally not rerun under the signed protocol; 60 iterations were verified from paper and fixed tuning.py, and released tuned files were used.
Execution used the available modern CPU environment (PyTorch 2.8.0) rather than the historical PyTorch 1.x-era environment; the fixed code ran without metric failures, with expected isolated-row and sparse-constructor warnings retained in logs.
Only one of 16 claim measurements was covered; the authors' GCN result, the claimed approximately one-point advantage, and the other listed baselines were not rerun.
The paper's historical seeds are unreported; the reproduction predeclared seeds 42–51.
Execution used CPU with PyTorch 2.8.0 rather than the historical software environment.
The bundled graph materialized 4,676 unique undirected Citeseer edges, whereas Table 1 reports 4,732; whether this is a counting convention or input difference was not resolved.
The official README requests 150 Citeseer epochs while the paper states 100, creating a source-level protocol ambiguity.
The plan was revised after a one-epoch Cora smoke probe but before the Citeseer campaign; the retained records show no outcome-driven selection between the 100- and 150-epoch arms.
Only 1 claim measurement(s) were executed; 15 remain unassessed.
This execution covers only part of the compound claim; unassessed measurements: m-t2-cit-gcn-lit, m-t2-cit-gat-lit, m-t2-cit-gln-lit, m-t2-cit-agnn-lit, m-t2-cit-lnet-lit, m-t2-cit-adalnet-lit, m-t2-cit-deepwalk-lit, m-t2-cit-dgi-lit, m-t2-cit-gcn-own, m-t2-cit-gat-own, m-t2-cit-fastgcn-own, m-t2-cit-gin-own, m-t2-cit-lnet-own, m-t2-cit-adalnet-own, m-t2-cit-dgi-own.
Only 1 claim measurement(s) were executed; 15 remain unassessed. Current execution challenges the covered SGC accuracy but cannot decide the composite claim. Ten predeclared 100-epoch Citeseer runs all produced 71.2%, 0.7 percentage points below the reported 71.9% and substantially larger than the paper's ±0.1 dispersion. The separately predeclared 150-epoch README sensitivity arm averaged 71.85%, but it materially changes the paper-stated 100-epoch condition and was appropriately not substituted. The GCN comparison and the other Table 2 baselines were not executed; with 15 claim measurements uncovered, the overall claim remains inconclusive.