3D visual grounding aims to identify the target object within a 3D point cloud scene referred to by a natural language description. While previous works attempt to exploit the verbo-visual relation with proposed cross-modal transformers, unstructured natural utterances and scattered objects might lead to undesirable performances. In this paper, we introduce DOrA, a novel 3D visual grounding framework with Order-Aware referring. DOrA is designed to leverage Large Language Models (LLMs) to parse language description, suggesting a referential order of anchor objects. Such ordered anchor objects allow DOrA to update visual features and locate the target object during the grounding process. Experimental results on the NR3D and ScanRefer datasets demonstrate our superiority in both low-resource and full-data scenarios. In particular, DOrA surpasses current state-of-the-art frameworks by 9.3% and 7.8% grounding accuracy under 1% data and 10% data settings, respectively.
翻译:三维视觉定位旨在根据自然语言描述,在三维点云场景中识别出所指代的目标物体。尽管先前的研究尝试利用所提出的跨模态变换器来挖掘语言与视觉之间的关系,但非结构化的自然语言表述与散乱的物体分布可能导致性能不佳。本文提出了一种新颖的具有顺序感知指代的三维视觉定位框架DOrA。DOrA利用大型语言模型(LLMs)解析语言描述,推断出锚定物体的指代顺序。这种有序的锚定物体使得DOrA能够在定位过程中更新视觉特征并确定目标物体。在NR3D与ScanRefer数据集上的实验结果表明,DOrA在低资源与全数据场景下均展现出优越性。特别是在1%数据与10%数据设定下,DOrA的定位准确率分别超越当前最先进框架9.3%与7.8%。