正在加载页面…
For the ImageNet plain networks trained with batch normalization, the authors report non-zero forward-signal variances and healthy backward-gradient norms, so they argue that the observed 34-layer degradation is unlikely to be caused by vanishing gradients; they conjecture low convergence rates but leave the cause open. · CiteArk