Memorization in diffusion models is often treated as a global property of the model or dataset. In practice, however, a single diffusion model can simultaneously generate both memorized and novel samples. Which training samples are most likely to be memorized? In this work, we show that memorization is governed by \emph{local data coverage}. Leveraging the connection between diffusion models and kernel density estimation (KDE), we derive a theoretical criterion that predicts whether a point is memorized based on the density of training data in its neighborhood and the size of the training dataset. In the high-dimensional limit, this leads to a sharp, local transition: regions of low coverage are dominated by isolated training samples, which are memorized, while dense regions support interpolation and generalization. We validate these predictions empirically, showing that memorization increases with local sparsity and that diffusion models exhibit a coexistence of memorized and novel samples within the same model. Extending this framework to multi-class settings, we further show that classes with higher intra-class sparsity (and thus lower local coverage) are more strongly memorized. Our results provide a local view of memorization in diffusion models, explaining when and where memorization occurs in terms of data geometry.
翻译:扩散模型中的记忆现象通常被视为模型或数据集的全局属性。然而在实践中,单个扩散模型能同时生成记忆样本与新颖样本。哪些训练样本最有可能被记忆?本文表明,记忆行为受局部数据覆盖度驱动。通过利用扩散模型与核密度估计(KDE)的内在联系,我们推导出理论判据:基于训练数据在样本邻域内的密度以及训练集规模,该判据可预测某数据点是否被记忆。在高维极限下,这会导致尖锐的局部相变:低覆盖区域被孤立训练样本主导(这些样本被记忆),而高密度区域则支持插值与泛化。我们通过实验验证了这些预测:记忆现象随局部稀疏性增强而加剧,且扩散模型在同一模型内呈现记忆样本与新样本共存的现象。将该框架扩展至多类场景后,我们进一步发现:类内稀疏性越高(即局部覆盖度越低)的类别会被更强烈地记忆。研究结果提供了扩散模型中记忆行为的局部视角,从数据几何角度阐释了记忆现象的发生条件与空间分布。