We study inferring 3D object-centric scene representations from a single image. While recent methods have shown potential in unsupervised 3D object discovery from simple synthetic images, they fail to generalize to real-world scenes with visually rich and diverse objects. This limitation stems from their object representations, which entangle objects' intrinsic attributes like shape and appearance with extrinsic, viewer-centric properties such as their 3D location. To address this bottleneck, we propose Unsupervised discovery of Object-Centric neural Fields (uOCF). uOCF focuses on learning the intrinsics of objects and models the extrinsics separately. Our approach significantly improves systematic generalization, thus enabling unsupervised learning of high-fidelity object-centric scene representations from sparse real-world images. To evaluate our approach, we collect three new datasets, including two real kitchen environments. Extensive experiments show that uOCF enables unsupervised discovery of visually rich objects from a single real image, allowing applications such as 3D object segmentation and scene manipulation. Notably, uOCF demonstrates zero-shot generalization to unseen objects from a single real image. Project page: https://red-fairy.github.io/uOCF/
翻译:我们研究从单张图像推断三维以对象为中心的场景表示。尽管近期方法已展现出从简单合成图像中无监督发现三维对象的潜力,但这些方法难以泛化至包含视觉丰富且多样对象的真实场景。这一限制源于其对象表示方式:将对象的固有属性(如形状和外观)与外在的、以观察者为中心的特性(如三维位置)相纠缠。为解决这一瓶颈,我们提出无监督发现以对象为中心的神经场(uOCF)。uOCF专注于学习对象的内在属性,并分别建模外在属性。我们的方法显著提升了系统性泛化能力,从而能够从稀疏的真实世界图像中无监督学习高保真度的以对象为中心的场景表示。为评估该方法,我们收集了三个新数据集,其中包括两个真实厨房环境。大量实验表明,uOCF能够从单张真实图像中无监督发现视觉丰富的对象,并可应用于三维对象分割和场景操作等任务。值得注意的是,uOCF展现出对单张真实图像中未见对象的零样本泛化能力。项目页面:https://red-fairy.github.io/uOCF/