This paper studies a new open-set problem, the open-vocabulary category-level object pose and size estimation. Given human text descriptions of arbitrary novel object categories, the robot agent seeks to predict the position, orientation, and size of the target object in the observed scene image. To enable such generalizability, we first introduce OO3D-9D, a large-scale photorealistic dataset for this task. Derived from OmniObject3D, OO3D-9D is the largest and most diverse dataset in the field of category-level object pose and size estimation. It includes additional annotations for the symmetry axis of each category, which help resolve symmetric ambiguity. Apart from the large-scale dataset, we find another key to enabling such generalizability is leveraging the strong prior knowledge in pre-trained visual-language foundation models. We then propose a framework built on pre-trained DinoV2 and text-to-image stable diffusion models to infer the normalized object coordinate space (NOCS) maps of the target instances. This framework fully leverages the visual semantic prior from DinoV2 and the aligned visual and language knowledge within the text-to-image diffusion model, which enables generalization to various text descriptions of novel categories. Comprehensive quantitative and qualitative experiments demonstrate that the proposed open-vocabulary method, trained on our large-scale synthesized data, significantly outperforms the baseline and can effectively generalize to real-world images of unseen categories. The project page is at https://ov9d.github.io.
翻译:本文研究一个新的开放集问题——开放词汇的类别级物体姿态与尺寸估计。给定任意新物体类别的人类文本描述,机器人智能体需预测观测场景图像中目标物体的位置、朝向和尺寸。为实现此类泛化能力,我们首先引入OO3D-9D——一个为此任务构建的大规模逼真数据集。该数据集基于OmniObject3D衍生而来,是类别级物体姿态与尺寸估计领域最大且最多样化的数据集,其额外包含了每个类别的对称轴标注,有助于解决对称性歧义。除大规模数据集外,我们发现实现此类泛化能力的另一关键在于利用预训练视觉-语言基础模型中的强大先验知识。为此,我们提出一个基于预训练DinoV2和文本到图像稳定扩散模型(Stable Diffusion)的框架,用于推断目标实例的归一化物体坐标空间(NOCS)图。该框架充分挖掘了DinoV2的视觉语义先验以及文本到图像扩散模型中对齐的视觉与语言知识,从而实现对各类新颖类别文本描述的泛化。全面的定量与定性实验表明,基于大规模合成数据训练的开放词汇方法显著优于基线,且能有效泛化至未见类别的真实世界图像。项目页面详见https://ov9d.github.io。