We bound the excess risk of interpolating deep linear networks trained using gradient flow. In a setting previously used to establish risk bounds for the minimum $\ell_2$-norm interpolant, we show that randomly initialized deep linear networks can closely approximate or even match known bounds for the minimum $\ell_2$-norm interpolant. Our analysis also reveals that interpolating deep linear models have exactly the same conditional variance as the minimum $\ell_2$-norm solution. Since the noise affects the excess risk only through the conditional variance, this implies that depth does not improve the algorithm's ability to "hide the noise". Our simulations verify that aspects of our bounds reflect typical behavior for simple data distributions. We also find that similar phenomena are seen in simulations with ReLU networks, although the situation there is more nuanced.
翻译:我们界定了使用梯度流训练的插值深度线性网络的超额风险。在先前用于建立最小ℓ2-范数插值器风险界的设定下,我们证明随机初始化的深度线性网络能够紧密逼近甚至匹配最小ℓ2-范数插值器的已知界限。分析还揭示,插值深度线性模型与最小ℓ2-范数解具有完全相同的条件方差。由于噪声仅通过条件方差影响超额风险,这意味着深度并未提升算法“隐藏噪声”的能力。仿真验证了我们的界限在简单数据分布上反映了典型行为。我们还发现,在ReLU网络的仿真中观察到类似现象,尽管其中情况更为复杂。