Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking · CiteArk
Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking
PublicScope: Standard
Authors:Ali Saheb Pasand, Elvis Dohmatob
More
Copy repository link
Reproduction
Start with coverage and measured results; open a run only when you need evidence or technical details.
RunMatchRepeat
0/21
claims supported by evidence
0
Supported
0
Challenged or mixed
0
Contradicted
3
Inconclusive
18
Not assessed
Supporting Assessment
Measurement confirmed
On modular multiplication, exact-SVD EGD reaches 95% accuracy in the fewest epochs at every modulus and has the lowest wall-clock time at p=79. Column normalization has the lowest wall-clock time at p=97 and p=127; rank-128 RSVD is presented as the best overall balance of step count and per-epoch cost.
Reported
10116 epochs
Observed
10114 epochs
Difference -2
Run history
Each row is one recorded execution. Commands, logs, hashes, and signatures are available in its details.
ClaimResultFinishedStatus
Failure log
Failed paths grouped by cause — check them before reproducing.
Other1 failures
Claim–experiment reproduction matrix
See each experiment's execution state, blocker, recovery action, and evidence destination while keeping operations separate from scientific conclusions. There are also 4 claims with no independent reproduction scheduled in this plan.
3 targets·3 with evidence·0 active·0 need attention
A successful execution does not by itself validate a paper claim
Target state says whether the platform completed the work. The scientific conclusion is determined only by immutable evidence and Assessments. Resource shortages and platform failures are never presented as scientific contradictions.
When training and test data share the same conditioned Gaussian distribution, the paper's theory gives a vanilla-GD test-error plateau of order 1/(eta epsilon) only when both stated conditions hold, and no plateau if either fails; Figure 8 reports empirical agreement with the theoretical curve, while EGD has high test accuracy from the beginning of the plotted run.
Information insufficientc05-toy-equal-distribution-result0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
In CIFAR-10 label-shuffling experiments, adding exact-SVD EGD increases the area under the training-accuracy curve for Adam, RAdam, and RMSprop at every reported shuffle schedule. Without EGD, RMSprop is strongest in each schedule; the largest printed multiplicative gain is Adam with shuffling at 3 epochs.
Information insufficientc18-cifar10-nonstationarity0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
In the anisotropic two-feature toy problem, vanilla GD rapidly fits the training data but can retain chance-level or very poor test performance for a long plateau. With distribution shift, the reported plateau scales as order 1/(eta epsilon) for large initialization and order (1/eta) log(1/(tau epsilon)) for small initialization; with identical train and test distributions it again scales as order 1/(eta epsilon).
Information insufficientc03-toy-vanilla-gd-plateau0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
On MNIST with a three-layer ReLU MLP and AdamW, every EGD variant reaches 85% validation accuracy in fewer steps than Grokfast. RSVD variants also substantially reduce wall-clock time; rank 4 is fastest in time, ranks 8 and 4 tie for the fewest steps, and exact-SVD EGD has the highest printed final validation accuracy.
Information insufficientc14-mnist-threshold-and-final-accuracy0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
EGD transforms a layer gradient G as (GG^T)^(-1/2)G (using pseudoinverses when rank deficient), preserving its singular vectors while setting every retained singular value to 1; the paper describes this as a whitened form of natural gradient descent.
Information insufficientc01-egd-gradient-normalization0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
For modular multiplication at p=79, 97, and 127, EGD and vanilla SGD take diverging trajectories and converge to different parameter subspaces. Both overshoot their distance from initialization before settling in regions with the same fixed distance, with EGD moving little after its early large steps and vanilla SGD spanning a wider range.
Information insufficientc20-modular-multiplication-parameter-trajectories0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
On the distribution-shift toy experiment, EGD reaches perfect generalization after only a few iterations and appears insensitive to initialization scale, whereas large initialization lengthens the vanilla-GD plateau and small initialization attenuates it.
Information insufficientc04-toy-egd-immediate-grokking0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
The MNIST learning curves show EGD and its RSVD variants grokking earlier than Grokfast. Grokfast exhibits a sudden accuracy drop and recovery, and both Grokfast and vanilla AdamW have noisy training loss, whereas exact-SVD EGD and all plotted RSVD ranks have smoother training-loss curves; some RSVD ranks grok faster than exact SVD.
Information insufficientc15-mnist-learning-curve-stability0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
The Section 5 parity and modular-arithmetic experiments use the hyperparameters printed in Appendix Table 1, with ReLU activations in all cases.
Information insufficientc09-section5-hyperparameters0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Across Parity(n,k) settings (400,2), (100,3), and (50,4), EGD groks substantially earlier—typically after only a few epochs—while vanilla SGD has a long test-accuracy plateau before eventually grokking.
Information insufficientc06-main-sparse-parity-curves1 plan0 runs
Scientific conclusionNot assessed
Official sparse-parity EGD/RSVD matched-control rerun
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
For modular multiplication at p=97 with a two-layer transformer and Adam, Grokfast reaches 95% accuracy sooner than exact-SVD EGD in both steps and wall-clock time. RSVD rank 36 is closer to Grokfast but remains slower by the printed values; all three accelerated methods finish at 100% validation accuracy, versus 99.1% for vanilla Adam.
Information insufficientc16-transformer-modular-multiplication0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
For modular addition at each reported modulus (79, 97, and 127), EGD reaches high test accuracy after only a few epochs, vanilla SGD remains on a long plateau before eventually grokking, and column normalization also groks faster than vanilla SGD.
Information insufficientc07-main-modular-addition-curves1 plan0 runs
Scientific conclusionNot assessed
Official modular-addition EGD/RSVD matched-control rerun
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
Averaging the best setup of each method across the Appendix C tasks, the paper reports that tuned RSVD EGD has the largest average gains in both epochs and wall-clock time relative to vanilla SGD.
Information insufficientc13-cross-task-average-speedups0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
Randomized-SVD EGD is not hyperparameter-free: its rank and inner-loop iteration count must be chosen, and reducing rank can trade faster updates for delayed or absent grokking. The paper fixes two inner iterations in Appendix C, omits oversampling from its simplified RSVD implementation, leaves the effect of oversampling for future work, and states that rank tuning is needed for the best gain.
Official implementationc23-rsvd-rank-sensitivity1 plan0 runs
Scientific conclusionNot assessed
Official sparse-parity EGD/RSVD matched-control rerun
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
For addition modulo 97 with a two-layer ReLU network, the hidden-layer gradient remains poorly conditioned from the start through the end of vanilla optimization: its largest singular value is much larger than its smallest, which the paper links to stalled dynamics and delayed generalization; EGD equalizes these singular values.
Information insufficientc02-ill-conditioned-gradient-spectrum1 plan0 runs
Scientific conclusionNot assessed
Official modular-addition EGD/RSVD matched-control rerun
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
For modular addition at p=79, 97, and 127, EGD and vanilla SGD follow diverging trajectories and converge to different parameter subspaces. Both overshoot in distance from initialization and later settle in regions having the same fixed distance from initialization, while EGD moves little after its initial large steps and vanilla SGD traverses a wider range.
Information insufficientc19-modular-addition-parameter-trajectories0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
For modular multiplication at each reported modulus (79, 97, and 127), EGD groks after only a few epochs and earlier than all plotted alternatives, which remain on long plateaus before eventually grokking.
Information insufficientc08-main-modular-multiplication-curves1 plan0 runs
Scientific conclusionNot assessed
Official modular-multiplication EGD/RSVD matched-control rerun
An experiment plan exists, but the current processing task has no matching execution target.
Not in current task
For sparse parity at k=2, 3, and 4, EGD and vanilla SGD initially diverge, after which the paper's accompanying prose says vanilla SGD moves toward the vicinity where EGD began grokking. EGD overshoots its distance from initialization, while vanilla SGD approaches gradually; both ultimately maintain approximately the same fixed distance.
Information insufficientc21-sparse-parity-parameter-trajectories0 plans0 runs
Scientific conclusionNot assessed
This claim has no executable experiment plan yet.
On modular multiplication, exact-SVD EGD reaches 95% accuracy in the fewest epochs at every modulus and has the lowest wall-clock time at p=79. Column normalization has the lowest wall-clock time at p=97 and p=127; rank-128 RSVD is presented as the best overall balance of step count and per-epoch cost.
Official implementationc11-modular-multiplication-efficiency1 plan1 run
Scientific conclusionInconclusive
Official modular-multiplication EGD/RSVD matched-control rerun
exp-core-modular-multiplication2 attempts
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
Assessed but inconclusive
Legacy task without target-level resource requirements
On modular addition, exact-SVD EGD requires the fewest epochs to 95% accuracy at every modulus; column normalization has the lowest wall-clock time for p=97 and p=127, while RSVD rank 128 has the lowest wall-clock time for p=79. Reducing RSVD rank to 64 lowers per-epoch cost but delays threshold crossing relative to rank 128.
Official implementationc10-modular-addition-efficiency1 plan1 run
Scientific conclusionInconclusive
Official modular-addition EGD/RSVD matched-control rerun
exp-core-modular-addition2 attempts
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
Assessed but inconclusive
Legacy task without target-level resource requirements
For every sparse-parity k, at least one RSVD configuration reaches 95% accuracy in fewer epochs and less wall-clock time than exact-SVD EGD. The winning table rows are RSVD rank 40 for k=2 and rank 30 for k=3 and k=4. Column normalization and RSVD rank 40 do not reach 95% for k=4.
Official implementationc12-sparse-parity-efficiency1 plan1 run
Scientific conclusionInconclusive
Official sparse-parity EGD/RSVD matched-control rerun
exp-core-sparse-parity2 attempts
Evidence published
Next step
The execution path is complete; inspect the scientific Assessment next.
View local evidence pathExpandCollapse
Research plan
Claim and experiment binding established
Execution target
Evidence published
Run attempt
Evidence publication · Completed
CAP evidence
1 immutable Artifact
Scientific judgment
Assessed but inconclusive
Legacy task without target-level resource requirements
Community reproductions
Executed locally and uploaded by users. The platform verifies file signatures; conclusions come from the uploaded runs.