We present Uni-Fusion, a universal continuous mapping framework for surfaces, surface properties (color, infrared, etc.) and more (latent features in CLIP embedding space, etc.). We propose the first universal implicit encoding model that supports encoding of both geometry and different types of properties (RGB, infrared, features, etc.) without requiring any training. Based on this, our framework divides the point cloud into regular grid voxels and generates a latent feature in each voxel to form a Latent Implicit Map (LIM) for geometries and arbitrary properties. Then, by fusing a local LIM frame-wisely into a global LIM, an incremental reconstruction is achieved. Encoded with corresponding types of data, our Latent Implicit Map is capable of generating continuous surfaces, surface property fields, surface feature fields, and all other possible options. To demonstrate the capabilities of our model, we implement three applications: (1) incremental reconstruction for surfaces and color (2) 2D-to-3D transfer of fabricated properties (3) open-vocabulary scene understanding by creating a text CLIP feature field on surfaces. We evaluate Uni-Fusion by comparing it in corresponding applications, from which Uni-Fusion shows high-flexibility in various applications while performing best or being competitive. The project page of Uni-Fusion is available at https://jarrome.github.io/Uni-Fusion/ .
翻译:我们提出Uni-Fusion——一种面向表面、表面属性(颜色、红外等)及更多信息(如CLIP嵌入空间中的潜在特征)的通用连续映射框架。本研究首次提出无需任何训练即可同时支持几何与多种属性(RGB、红外、特征等)编码的通用隐式编码模型。基于此,本框架将点云划分为规则网格体素,在每个体素中生成潜在特征,从而构建几何与任意属性的潜在隐式映射(LIM)。通过逐帧将局部LIM融合为全局LIM,实现增量式重构。经对应类型数据编码后,我们的潜在隐式映射能够生成连续表面、表面属性场、表面特征场及其他所有可能的选项。为展示模型能力,我们实现三项应用:(1)表面与颜色的增量重构;(2)人造属性的二维到三维迁移;(3)通过在表面创建文本CLIP特征场实现开放词汇场景理解。我们通过相应应用对比评估Uni-Fusion,结果表明该框架在各类应用中展现出高度灵活性,同时达到最佳或具有竞争力的性能。Uni-Fusion的项目页面位于https://jarrome.github.io/Uni-Fusion/。