Score-based diffusion models are a highly effective method for generating samples from a distribution of images. We consider scenarios where the training data comes from a noisy version of the target distribution, and present an efficiently implementable modification of the inference procedure to generate noiseless samples. Our approach is motivated by the manifold hypothesis, according to which meaningful data is concentrated around some low-dimensional manifold of a high-dimensional ambient space. The central idea is that noise manifests as low magnitude variation in off-manifold directions in contrast to the relevant variation of the desired distribution which is mostly confined to on-manifold directions. We introduce the notion of an extended score and show that, in a simplified setting, it can be used to reduce small variations to zero, while leaving large variations mostly unchanged. We describe how its approximation can be computed efficiently from an approximation to the standard score and demonstrate its efficacy on toy problems, synthetic data, and real data.
翻译:基于分数的扩散模型是从图像分布中生成样本的一种高效方法。我们考虑训练数据来自目标分布的有噪声版本的情况,并提出一种可高效实现的推理过程修改方案,以生成无噪声样本。我们的方法受流形假设启发,该假设认为有意义的数据集中在高维环境空间中的某个低维流形附近。核心思想在于:噪声表现为流形外方向上的低幅度变化,而目标分布的相关变化则主要局限于流形内方向。我们引入扩展分数概念,并证明在简化设定下,该概念可将微小变化归零,同时基本保留较大变化。我们描述了如何通过标准分数的近似来高效计算扩展分数的近似值,并在玩具问题、合成数据和真实数据上验证了其有效性。