This paper proposes ConsistDreamer - a novel framework that lifts 2D diffusion models with 3D awareness and 3D consistency, thus enabling high-fidelity instruction-guided scene editing. To overcome the fundamental limitation of missing 3D consistency in 2D diffusion models, our key insight is to introduce three synergetic strategies that augment the input of the 2D diffusion model to become 3D-aware and to explicitly enforce 3D consistency during the training process. Specifically, we design surrounding views as context-rich input for the 2D diffusion model, and generate 3D-consistent, structured noise instead of image-independent noise. Moreover, we introduce self-supervised consistency-enforcing training within the per-scene editing procedure. Extensive evaluation shows that our ConsistDreamer achieves state-of-the-art performance for instruction-guided scene editing across various scenes and editing instructions, particularly in complicated large-scale indoor scenes from ScanNet++, with significantly improved sharpness and fine-grained textures. Notably, ConsistDreamer stands as the first work capable of successfully editing complex (e.g., plaid/checkered) patterns. Our project page is at immortalco.github.io/ConsistDreamer.
翻译:本文提出ConsistDreamer——一种将二维扩散模型提升至具备三维感知与三维一致性的新型框架,从而实现高保真度的指令引导场景编辑。为克服二维扩散模型缺乏三维一致性的根本局限,我们的核心思路是引入三种协同策略:在训练过程中增强二维扩散模型的输入使其具备三维感知能力,并显式强化三维一致性。具体而言,我们设计了环绕视图作为二维扩散模型的上下文丰富输入,并生成具有三维一致性的结构化噪声而非图像独立噪声。此外,我们在逐场景编辑流程中引入了自监督的一致性强化训练机制。大量实验表明,ConsistDreamer在各类场景与编辑指令的引导编辑任务中均达到最先进性能,特别是在ScanNet++复杂大规模室内场景中,其锐利度与细粒度纹理表现显著提升。值得关注的是,ConsistDreamer是首个能成功编辑复杂图案(如格纹/棋盘格)的方法。项目页面详见immortalco.github.io/ConsistDreamer。