Deep neural networks (DNNs) trained for image denoising are able to generate high-quality samples with score-based reverse diffusion algorithms. These impressive capabilities seem to imply an escape from the curse of dimensionality, but recent reports of memorization of the training set raise the question of whether these networks are learning the "true" continuous density of the data. Here, we show that two DNNs trained on non-overlapping subsets of a dataset learn nearly the same score function, and thus the same density, when the number of training images is large enough. In this regime of strong generalization, diffusion-generated images are distinct from the training set, and are of high visual quality, suggesting that the inductive biases of the DNNs are well-aligned with the data density. We analyze the learned denoising functions and show that the inductive biases give rise to a shrinkage operation in a basis adapted to the underlying image. Examination of these bases reveals oscillating harmonic structures along contours and in homogeneous regions. We demonstrate that trained denoisers are inductively biased towards these geometry-adaptive harmonic bases since they arise not only when the network is trained on photographic images, but also when it is trained on image classes supported on low-dimensional manifolds for which the harmonic basis is suboptimal. Finally, we show that when trained on regular image classes for which the optimal basis is known to be geometry-adaptive and harmonic, the denoising performance of the networks is near-optimal.
翻译:深度神经网络(DNN)在图像去噪任务中经过训练后,能够通过基于分数的逆扩散算法生成高质量样本。这些令人瞩目的能力似乎暗示着对维度诅咒的规避,但近期关于训练集记忆化的报道引发疑问:这些网络是否在学习数据“真实”的连续密度?本文证明,当训练图像数量足够大时,在数据集非重叠子集上训练的两个DNN会学到几乎相同的分数函数,从而学习到相同的密度。在强泛化机制下,扩散生成的图像与训练集截然不同,且具有高视觉质量,表明DNN的归纳偏置与数据密度高度对齐。我们分析学习到的去噪函数后发现,归纳偏置在适应底层图像的基上产生了收缩操作。对这些基的检测揭示出轮廓边缘与均匀区域中振荡的谐波结构。我们证明训练后的去噪器天生偏向这些几何自适应谐波基——不仅当网络在摄影图像上训练时如此,当其训练于低维流形支撑的图像类(即使谐波基在此类情况下非最优)时亦然。最后,当网络在已知最优基为几何自适应谐波的规则图像类上训练时,其去噪性能接近最优。