On CIFAR-10, deeper plain networks show higher training error as depth increases, including plain-56 versus plain-20, whereas residual networks overcome this optimization difficulty and gain accuracy from depths 20 through 56. The plain-110 error is reported as above 60% and is omitted from the plot. · CiteArk
On CIFAR-10, deeper plain networks show higher training error as depth increases, including plain-56 versus plain-20, whereas residual networks overcome this optimization difficulty and gain accuracy from depths 20 through 56. The plain-110 error is reported as above 60% and is omitted from the plot.
来源:paper:PDF pages 1, 7-8, Figures 1 and 6 and Section 4.2 discussion