For addition modulo 97 with a two-layer ReLU network, the hidden-layer gradient remains poorly conditioned from the start through the end of vanilla optimization: its largest singular value is much larger than its smallest, which the paper links to stalled dynamics and delayed generalization; EGD equalizes these singular values. · CiteArk
Not assessedPlan blockedFindingc02-ill-conditioned-gradient-spectrum
For addition modulo 97 with a two-layer ReLU network, the hidden-layer gradient remains poorly conditioned from the start through the end of vanilla optimization: its largest singular value is much larger than its smallest, which the paper links to stalled dynamics and delayed generalization; EGD equalizes these singular values.
Source: paper:PDF p. 3, Figure 4 and caption
Reported and observed measurements
No structured measurement is attached to this Claim.
Assessments (0)
No immutable Assessment has been published for this Claim yet.