Residual networks (ResNets) have significantly better trainability and thus performance than feed-forward networks at large depth. Introducing skip connections facilitates signal propagation to deeper layers. In addition, previous works found that adding a scaling parameter for the residual branch further improves generalization performance. While they empirically identified a particularly beneficial range of values for this scaling parameter, the associated performance improvement and its universality across network hyperparameters yet need to be understood. For feed-forward networks (FFNets), finite-size theories have led to important insights with regard to signal propagation and hyperparameter tuning. We here derive a systematic finite-size theory for ResNets to study signal propagation and its dependence on the scaling for the residual branch. We derive analytical expressions for the response function, a measure for the network's sensitivity to inputs, and show that for deep networks the empirically found values for the scaling parameter lie within the range of maximal sensitivity. Furthermore, we obtain an analytical expression for the optimal scaling parameter that depends only weakly on other network hyperparameters, such as the weight variance, thereby explaining its universality across hyperparameters. Overall, this work provides a framework for theory-guided optimal scaling in ResNets and, more generally, provides the theoretical framework to study ResNets at finite widths.
翻译:残差网络(ResNets)在深度较大时具有显著优于前馈网络的训练性能。引入跳跃连接能够促进信号向更深层传播。此外,先前研究发现对残差分支添加缩放参数可进一步提升泛化性能。尽管已有研究通过实验确定了该缩放参数的特定有利取值范围,但相关性能提升及其在不同网络超参数间的普适性仍有待阐明。对于前馈网络(FFNets),有限尺寸理论在信号传播与超参数调优方面已取得重要进展。本文为研究残差网络信号传播及其对残差分支缩放参数的依赖关系,系统推导了面向ResNets的有限尺寸理论。我们推导了响应函数(衡量网络对输入敏感度的指标)的解析表达式,并证明在深度网络中,实验确定的缩放参数值恰处于最大敏感度区间。进一步,我们获得了最优缩放参数的解析表达式,该表达式对权重方差等其他网络超参数的依赖性很弱,从而解释了其跨超参数的普适性。总体而言,本研究为ResNets中的理论引导最优缩放提供了框架,并更广泛地建立了研究有限宽度ResNets的理论基础。