What makes waveform-based deep learning so hard? Despite numerous attempts at training convolutional neural networks (convnets) for filterbank design, they often fail to outperform hand-crafted baselines. These baselines are linear time-invariant systems: as such, they can be approximated by convnets with wide receptive fields. Yet, in practice, gradient-based optimization leads to suboptimal approximations. In our article, we approach this phenomenon from the perspective of initialization. We present a theory of large deviations for the energy response of FIR filterbanks with random Gaussian weights. We find that deviations worsen for large filters and locally periodic input signals, which are both typical for audio signal processing applications. Numerical simulations align with our theory and suggest that the condition number of a convolutional layer follows a logarithmic scaling law between the number and length of the filters, which is reminiscent of discrete wavelet bases.
翻译:基于波形深度学习为何如此困难?尽管在滤波器组设计中训练卷积神经网络(卷积网络)已进行多次尝试,但这些网络往往未能超越手工优化的基线模型。此类基线属于线性时不变系统:理论上可通过具有宽感受野的卷积网络近似实现。然而在实践中,基于梯度的优化会导致次优逼近。本文从初始化的角度探讨这一现象。我们提出关于随机高斯权重的FIR滤波器组能量响应的大偏差理论。研究发现,对于大规模滤波器及局部周期性输入信号(这恰是音频信号处理应用的典型特征),偏差会加剧。数值模拟与理论分析一致,表明卷积层的条件数在滤波器数量与长度之间遵循对数标度律,这一特性令人联想到离散小波基。