Generative neural networks learn how to produce highly realistic images from a large, but finite number of examples - or do they simply memorise their training set? To settle this question, Kadkhodaie, Guth, Simoncelli and Mallat (ICLR '24) trained diffusion models independently on disjoint subsets of a dataset and showed that they converge to nearly the same density when the number of training images is large enough. This result raises two basic questions: how much data do you need for convergence, and what does convergence capture about learning the data distribution? Here, we address these questions by providing an exact analytical characterisation of the transition from memorisation to generalisation in linear generative models. We find that these models memorise at small load, while convergence emerges continuously when the number of samples is linear in the input dimension. Strikingly, we find that convergence is insensitive to recovery of the principal latent factors of the data, which are recovered in a sharp transition. After extending our approach to data with power-law spectra, we find the same distinction between convergence and latent recovery in our experiments with convolutional denoisers and in the data of Kadkhodaie et al. We thus show that generalisation in generative models decomposes into at least two distinct objectives: matching the bulk of the data distribution and recovering the principal latent factors. These objectives correspond to two different distances between true and learnt data distribution, and only the first one is captured by convergence.
翻译:生成神经网络通过大量但有限的样本学习生成高度逼真的图像——抑或它们仅仅是在记忆训练集?为解答此问题,Kadkhodaie、Guth、Simoncelli与Mallat(ICLR '24)在数据集的不相交子集上独立训练扩散模型,并发现当训练图像数量足够大时,这些模型会收敛至近乎相同的密度。这一结果引出了两个基本问题:需要多少数据才能实现收敛?收敛又揭示了关于学习数据分布的何种信息?本文通过在线性生成模型中精确刻画从记忆到泛化的转变来回答这些问题。我们发现,当样本负载较小时,模型呈现记忆行为;而当样本数量与输入维度呈线性关系时,收敛性逐渐涌现。引人注目的是,收敛性对数据主成分潜在因子的恢复并不敏感——而这些因子是在一个尖锐的转变中才被恢复的。在将方法拓展至具有幂律谱分布的数据后,我们在卷积去噪器实验及Kadkhodaie等人的数据中发现了同样的收敛性与潜在因子恢复之间的区别。因此,我们证明生成模型的泛化可分解为至少两个不同的目标:匹配数据分布的整体特征与恢复主成分潜在因子。这两个目标对应着真实数据分布与学习数据分布之间的两种不同距离,而收敛性仅能捕捉前者。