For the ImageNet plain networks trained with batch normalization, the authors report non-zero forward-signal variances and healthy backward-gradient norms, so they argue that the observed 34-layer degradation is unlikely to be caused by vanishing gradients; they conjecture low convergence rates but leave the cause open. · CiteArk
For the ImageNet plain networks trained with batch normalization, the authors report non-zero forward-signal variances and healthy backward-gradient norms, so they argue that the observed 34-layer degradation is unlikely to be caused by vanishing gradients; they conjecture low convergence rates but leave the cause open.
来源:paper:PDF page 5, Section 4.1, paragraph following Figure 4 and Table 2