Diffusion models are often trained in low-dimensional latent spaces, which are then reused for related but shifted datasets. In this work, we study when such latent reuse remains reliable under distribution shift. We consider a source-target setting in which both datasets are approximately low-dimensional but may lie near different subspaces. We show that freezing and reusing a source latent space induces a target-domain score error governed by two quantities: the principal-angle misalignment between the source and target subspaces, and the target ambient noise amplified by the diffusion time scale. Motivated by these limits, we further study mixed source-target training and characterize how the required shared latent dimension depends on the relative geometry of the two distributions. Our results provide theoretical guidance on when latent reuse is reliable and when learning a shared representation may be necessary.
翻译:扩散模型通常训练于低维潜在空间,并进一步将其复用于相关但存在偏移的数据集。本研究探寻此类潜在空间复用在分布偏移条件下何时仍保持可靠性。我们考虑源-目标场景:两个数据集均近似低维,但可能位于不同子空间附近。研究表明,冻结并复用源域潜在空间会导致目标域得分误差,该误差由两个量决定:源子空间与目标子空间之间的主角度偏差,以及扩散时间尺度放大的目标环境噪声。基于这些边界,我们进一步研究混合源-目标训练,并刻画所需共享潜在维度如何取决于两个分布间的相对几何结构。研究结果为潜在空间复用在何时可靠,以及何时需要学习共享表示提供了理论指导。