The success of image generative models has enabled us to build methods that can edit images based on text or other user input. However, these methods are bespoke, imprecise, require additional information, or are limited to only 2D image edits. We present GeoDiffuser, a zero-shot optimization-based method that unifies common 2D and 3D image-based object editing capabilities into a single method. Our key insight is to view image editing operations as geometric transformations. We show that these transformations can be directly incorporated into the attention layers in diffusion models to implicitly perform editing operations. Our training-free optimization method uses an objective function that seeks to preserve object style but generate plausible images, for instance with accurate lighting and shadows. It also inpaints disoccluded parts of the image where the object was originally located. Given a natural image and user input, we segment the foreground object using SAM and estimate a corresponding transform which is used by our optimization approach for editing. GeoDiffuser can perform common 2D and 3D edits like object translation, 3D rotation, and removal. We present quantitative results, including a perceptual study, that shows how our approach is better than existing methods. Visit https://ivl.cs.brown.edu/research/geodiffuser.html for more information.
翻译:图像生成模型的成功使我们能够构建基于文本或其他用户输入编辑图像的方法。然而,这些方法往往是特制的、不精确的,需要额外信息,或仅限于2D图像编辑。我们提出GeoDiffuser,一种基于零样本优化的方法,将常见的2D和3D图像对象编辑能力统一到单一方法中。我们的关键洞察是将图像编辑操作视为几何变换。我们证明,这些变换可以直接融入扩散模型的注意力层中,以隐式执行编辑操作。我们无需训练的优化方法使用一个目标函数,旨在保留对象风格的同时生成合理的图像(例如具有准确的照明和阴影)。它还能修复对象原始位置处被遮挡的图像部分。给定自然图像和用户输入,我们使用SAM分割前景对象并估计相应的变换,该变换用于我们的优化方法进行编辑。GeoDiffuser能够执行常见的2D和3D编辑,如对象平移、3D旋转和移除。我们展示了定量结果,包括一项感知研究,表明我们的方法优于现有方法。更多信息请访问https://ivl.cs.brown.edu/research/geodiffuser.html。