Loading page…
For addition modulo 97 with a two-layer ReLU network, the hidden-layer gradient remains poorly conditioned from the start through the end of vanilla optimization: its largest singular value is much larger than its smallest, which the paper links to stalled dynamics and delayed generalization; EGD equalizes these singular values. · CiteArk