There is an emerging effort to combine the two popular technical paths, i.e., the multi-view stereo (MVS) and neural implicit surface (NIS), in scene reconstruction from sparse views. In this paper, we introduce a novel integration scheme that combines the multi-view stereo with neural signed distance function representations, which potentially overcomes the limitations of both methods. MVS uses per-view depth estimation and cross-view fusion to generate accurate surface, while NIS relies on a common coordinate volume. Based on this, we propose to construct per-view cost frustum for finer geometry estimation, and then fuse cross-view frustums and estimate the implicit signed distance functions to tackle noise and hole issues. We further apply a cascade frustum fusion strategy to effectively captures global-local information and structural consistency. Finally, we apply cascade sampling and a pseudo-geometric loss to foster stronger integration between the two architectures. Extensive experiments demonstrate that our method reconstructs robust surfaces and outperforms existing state-of-the-art methods.
翻译:近年来,学术界致力于结合多视角立体视觉(MVS)与神经隐式表面(NIS)这两条主流技术路径,以实现稀疏视角下的场景重建。本文提出一种新型集成方案,将多视角立体视觉与神经符号距离函数表征相结合,有望克服两种方法的固有局限。MVS通过逐视角深度估计与跨视角融合生成精确表面,而NIS则依赖公共坐标体积。基于此,我们提出构建逐视角代价截锥体以获取更精细的几何估计,进而融合跨视角截锥体并估计隐式符号距离函数,以解决噪声与空洞问题。进一步采用级联截锥体融合策略,有效捕获全局-局部信息与结构一致性。最后应用级联采样与伪几何损失函数,增强两种架构间的深度融合。大量实验表明,本方法可重建鲁棒表面,性能优于现有最先进方法。