正在加载页面…
The 1202-layer CIFAR-10 ResNet optimizes to below 0.1% training error but has 7.93% test error, worse than ResNet-110 despite similar training error. The authors attribute this generalization gap to overfitting by a 19.4M-parameter model on the small dataset and note that they did not use maxout or dropout. · CiteArk