We introduce the task of localizing a flexible number of objects in real-world 3D scenes using natural language descriptions. Existing 3D visual grounding tasks focus on localizing a unique object given a text description. However, such a strict setting is unnatural as localizing potentially multiple objects is a common need in real-world scenarios and robotic tasks (e.g., visual navigation and object rearrangement). To address this setting we propose Multi3DRefer, generalizing the ScanRefer dataset and task. Our dataset contains 61926 descriptions of 11609 objects, where zero, single or multiple target objects are referenced by each description. We also introduce a new evaluation metric and benchmark methods from prior work to enable further investigation of multi-modal 3D scene understanding. Furthermore, we develop a better baseline leveraging 2D features from CLIP by rendering object proposals online with contrastive learning, which outperforms the state of the art on the ScanRefer benchmark.
翻译:[翻译摘要]
我们提出了一个任务:利用自然语言描述,在真实世界的3D场景中定位数量灵活的对象。现有的3D视觉定位任务主要关注于根据文本描述定位单个唯一对象。然而,这种严格设定并不自然,因为在现实场景和机器人任务(例如视觉导航和对象重排)中,潜在的多对象定位是常见需求。为解决此问题,我们提出了Multi3DRefer,该模型对ScanRefer数据集和任务进行了泛化。我们的数据集包含61926个描述,涵盖11609个对象,每个描述可能引用零个、单个或多个目标对象。我们还引入了新的评估指标和基准方法,这些方法基于先前工作,以促进多模态3D场景理解的进一步研究。此外,我们开发了一个更优的基线模型,通过在线渲染对象提议并结合对比学习利用CLIP的2D特征,在ScanRefer基准上超越了当前最佳性能。