Existing image editing tools, while powerful, typically disregard the underlying 3D geometry from which the image is projected. As a result, edits made using these tools may become detached from the geometry and lighting conditions that are at the foundation of the image formation process. In this work, we formulate the newt ask of language-guided 3D-aware editing, where objects in an image should be edited according to a language instruction in context of the underlying 3D scene. To promote progress towards this goal, we release OBJECT: a dataset consisting of 400K editing examples created from procedurally generated 3D scenes. Each example consists of an input image, editing instruction in language, and the edited image. We also introduce 3DIT : single and multi-task models for four editing tasks. Our models show impressive abilities to understand the 3D composition of entire scenes, factoring in surrounding objects, surfaces, lighting conditions, shadows, and physically-plausible object configurations. Surprisingly, training on only synthetic scenes from OBJECT, editing capabilities of 3DIT generalize to real-world images.
翻译:现有图像编辑工具虽然功能强大,但通常忽略图像所投影的底层三维几何结构。因此,使用这些工具进行的编辑可能会脱离作为图像形成过程基础的几何形状与光照条件。本研究提出语言引导的三维感知编辑新任务,要求根据语言指令在底层三维场景语境中编辑图像中的物体。为推进该目标,我们发布了OBJECT数据集:包含由程序化生成的三维场景创建的40万个编辑示例,每个示例由输入图像、语言编辑指令及编辑后图像组成。同时提出3DIT模型——面向四项编辑任务的单任务与多任务模型。我们的模型展现出理解整个场景三维构成的卓越能力,能综合考量周围物体、表面、光照条件、阴影以及物理合理的物体配置。令人惊讶的是,仅使用OBJECT合成场景训练的3DIT模型,其编辑能力可泛化至真实世界图像。