Scene image editing is crucial for entertainment, photography, and advertising design. Existing methods solely focus on either 2D individual object or 3D global scene editing. This results in a lack of a unified approach to effectively control and manipulate scenes at the 3D level with different levels of granularity. In this work, we propose 3DitScene, a novel and unified scene editing framework leveraging language-guided disentangled Gaussian Splatting that enables seamless editing from 2D to 3D, allowing precise control over scene composition and individual objects. We first incorporate 3D Gaussians that are refined through generative priors and optimization techniques. Language features from CLIP then introduce semantics into 3D geometry for object disentanglement. With the disentangled Gaussians, 3DitScene allows for manipulation at both the global and individual levels, revolutionizing creative expression and empowering control over scenes and objects. Experimental results demonstrate the effectiveness and versatility of 3DitScene in scene image editing. Code and online demo can be found at our project homepage: https://zqh0253.github.io/3DitScene/.
翻译:场景图像编辑在娱乐、摄影和广告设计中至关重要。现有方法仅专注于二维单个物体或三维全局场景编辑,导致缺乏一种统一的方法来在不同粒度级别上有效控制和操作三维场景。在本工作中,我们提出3DitScene,一种新颖统一的场景编辑框架,利用语言引导解耦高斯溅射,实现从二维到三维的无缝编辑,允许对场景构图和单个物体进行精确控制。我们首先引入通过生成先验和优化技术精炼的三维高斯表示。随后,来自CLIP的语言特征将语义信息引入三维几何以实现物体解耦。借助解耦的高斯表示,3DitScene允许在全局和个体级别进行操作,彻底改变了创意表达方式,并增强了对场景和物体的控制能力。实验结果证明了3DitScene在场景图像编辑中的有效性和多功能性。代码和在线演示可在我们的项目主页找到:https://zqh0253.github.io/3DitScene/。