Flow matching models have shown great potential in image generation tasks among probabilistic generative models. However, most flow matching models in the literature do not explicitly utilize the underlying clustering structure in the target data when learning the flow from a simple source distribution like the standard Gaussian. This leads to inefficient learning, especially for many high-dimensional real-world datasets, which often reside in a low-dimensional manifold. To this end, we present $\texttt{Latent-CFM}$, which provides efficient training strategies by conditioning on the features extracted from data using pretrained deep latent variable models. Through experiments on synthetic data from multi-modal distributions and widely used image benchmark datasets, we show that $\texttt{Latent-CFM}$ exhibits improved generation quality with significantly less training and computation than state-of-the-art flow matching models by adopting pretrained lightweight latent variable models. Beyond natural images, we consider generative modeling of spatial fields stemming from physical processes. Using a 2d Darcy flow dataset, we demonstrate that our approach generates more physically accurate samples than competing approaches. In addition, through latent space analysis, we demonstrate that our approach can be used for conditional image generation conditioned on latent features, which adds interpretability to the generation process.
翻译:流匹配模型在概率生成模型的图像生成任务中展现出巨大潜力。然而,现有文献中的多数流匹配模型在学习从简单源分布(如标准高斯分布)到目标数据的流时,并未明确利用目标数据中的潜在聚类结构。这导致学习效率低下,尤其对于许多通常位于低维流形上的高维真实世界数据集而言。为此,我们提出$\texttt{Latent-CFM}$,通过利用预训练深度潜在变量模型从数据中提取的特征进行条件化处理,提供高效训练策略。通过对多模态分布合成数据及常用图像基准数据集的实验,我们表明:$\texttt{Latent-CFM}$通过采用预训练轻量级潜在变量模型,在显著减少训练与计算量的情况下,相比现有最先进的流匹配模型展现出更优的生成质量。除自然图像外,我们还考虑了源于物理过程的空间场生成建模。基于二维达西流数据集的实验表明,我们的方法相比竞争方法能够生成物理上更准确的样本。此外,通过潜在空间分析,我们证明该方法可基于潜在特征进行条件图像生成,从而为生成过程增加可解释性。