Loading page…
For ResNet-110 on CIFAR-10, an initial learning rate of 0.1 was slightly too large for prompt convergence, so the reported run warmed up at 0.01 until training error fell below 80% (about 400 iterations) before returning to 0.1. The paper notes that starting directly at 0.1 eventually crossed 90% error after several epochs and reached similar final accuracy. · CiteArk