We propose UpFusion, a system that can perform novel view synthesis and infer 3D representations for an object given a sparse set of reference images without corresponding pose information. Current sparse-view 3D inference methods typically rely on camera poses to geometrically aggregate information from input views, but are not robust in-the-wild when such information is unavailable/inaccurate. In contrast, UpFusion sidesteps this requirement by learning to implicitly leverage the available images as context in a conditional generative model for synthesizing novel views. We incorporate two complementary forms of conditioning into diffusion models for leveraging the input views: a) via inferring query-view aligned features using a scene-level transformer, b) via intermediate attentional layers that can directly observe the input image tokens. We show that this mechanism allows generating high-fidelity novel views while improving the synthesis quality given additional (unposed) images. We evaluate our approach on the Co3Dv2 and Google Scanned Objects datasets and demonstrate the benefits of our method over pose-reliant sparse-view methods as well as single-view methods that cannot leverage additional views. Finally, we also show that our learned model can generalize beyond the training categories and even allow reconstruction from self-captured images of generic objects in-the-wild.
翻译:我们提出UpFusion系统,能够在缺乏对应姿态信息的稀疏参考图像条件下,实现目标物体的新型视角合成与三维表征推断。当前基于稀疏视角的三维推断方法通常依赖相机姿态对输入视图进行几何信息聚合,但这类方法在真实场景中因姿态信息缺失或不准确而缺乏鲁棒性。相比之下,UpFusion通过隐式学习将可用图像作为条件生成模型的上下文来合成新视角,从而绕开这一需求。我们为扩散模型融合两种互补的条件化机制以利用输入视图:a) 通过场景级Transformer推断与查询视角对齐的特征;b) 通过可直接观测输入图像令牌的中间注意力层。实验表明,该机制不仅能生成高保真新视角,还能在增加(无姿态)图像数量时提升合成质量。我们在Co3Dv2与Google Scanned Objects数据集上验证了该方法,证明了其相比依赖姿态的稀疏视角方法及无法利用额外视图的单视角方法的优势。最后,我们的学习模型展现出超越训练类别的泛化能力,甚至能够基于真实场景中通用物体的自拍摄图像实现重建。