Editing objects within a scene is a critical functionality required across a broad spectrum of applications in computer vision and graphics. As 3D Gaussian Splatting (3DGS) emerges as a frontier in scene representation, the effective modification of 3D Gaussian scenes has become increasingly vital. This process entails accurately retrieve the target objects and subsequently performing modifications based on instructions. Though available in pieces, existing techniques mainly embed sparse semantics into Gaussians for retrieval, and rely on an iterative dataset update paradigm for editing, leading to over-smoothing or inconsistency issues. To this end, this paper proposes a systematic approach, namely TIGER, for coherent text-instructed 3D Gaussian retrieval and editing. In contrast to the top-down language grounding approach for 3D Gaussians, we adopt a bottom-up language aggregation strategy to generate a denser language embedded 3D Gaussians that supports open-vocabulary retrieval. To overcome the over-smoothing and inconsistency issues in editing, we propose a Coherent Score Distillation (CSD) that aggregates a 2D image editing diffusion model and a multi-view diffusion model for score distillation, producing multi-view consistent editing with much finer details. In various experiments, we demonstrate that our TIGER is able to accomplish more consistent and realistic edits than prior work.
翻译:在计算机视觉与图形学的广泛应用中,对场景内物体进行编辑是一项关键功能。随着3D高斯泼溅(3DGS)成为场景表征的前沿技术,对3D高斯场景进行有效修改变得日益重要。该过程需要精确检索目标物体,随后根据指令执行修改。尽管现有技术已实现部分功能,但其主要将稀疏语义嵌入高斯分布以进行检索,并依赖迭代式数据集更新范式进行编辑,导致过度平滑或不一致问题。为此,本文提出一种系统性方法TIGER,用于实现基于文本指令的连贯3D高斯检索与编辑。与针对3D高斯的自上而下语言定位方法不同,我们采用自下而上的语言聚合策略,生成支持开放词汇检索的稠密语言嵌入3D高斯分布。为克服编辑中的过度平滑与不一致问题,我们提出连贯分数蒸馏(CSD)方法,聚合2D图像编辑扩散模型与多视角扩散模型进行分数蒸馏,从而生成具有更精细细节的多视角一致编辑结果。在多项实验中,我们证明TIGER能够实现比先前工作更一致、更逼真的编辑效果。