Open-vocabulary querying in 3D space is challenging but essential for scene understanding tasks such as object localization and segmentation. Language-embedded scene representations have made progress by incorporating language features into 3D spaces. However, their efficacy heavily depends on neural networks that are resource-intensive in training and rendering. Although recent 3D Gaussians offer efficient and high-quality novel view synthesis, directly embedding language features in them leads to prohibitive memory usage and decreased performance. In this work, we introduce Language Embedded 3D Gaussians, a novel scene representation for open-vocabulary query tasks. Instead of embedding high-dimensional raw semantic features on 3D Gaussians, we propose a dedicated quantization scheme that drastically alleviates the memory requirement, and a novel embedding procedure that achieves smoother yet high accuracy query, countering the multi-view feature inconsistencies and the high-frequency inductive bias in point-based representations. Our comprehensive experiments show that our representation achieves the best visual quality and language querying accuracy across current language-embedded representations, while maintaining real-time rendering frame rates on a single desktop GPU.
翻译:开放词汇查询在三维空间中具有挑战性,但对于目标定位和分割等场景理解任务至关重要。语言嵌入场景表示通过将语言特征融入三维空间取得了进展,但其有效性严重依赖于训练和渲染时资源消耗巨大的神经网络。尽管近期3D高斯表示能够实现高效高质量的 novel view合成,但直接在其中嵌入语言特征会导致内存占用过高和性能下降。本文提出语言嵌入3D高斯表示——一种面向开放词汇查询任务的新型场景表示方法。我们摒弃在3D高斯上嵌入高维原始语义特征的做法,转而提出专用量化方案大幅缓解内存需求,并设计新型嵌入流程以应对点云表示中多视角特征不一致性与高频归纳偏置问题,从而实现更平滑且高精度的查询。综合实验表明,在现有语言嵌入表示中,我们的表示方法达到最优的视觉质量与语言查询精度,同时可在单桌面GPU上保持实时渲染帧率。