We propose a new task to benchmark scene understanding of embodied agents: Situated Question Answering in 3D Scenes (SQA3D). Given a scene context (e.g., 3D scan), SQA3D requires the tested agent to first understand its situation (position, orientation, etc.) in the 3D scene as described by text, then reason about its surrounding environment and answer a question under that situation. Based upon 650 scenes from ScanNet, we provide a dataset centered around 6.8k unique situations, along with 20.4k descriptions and 33.4k diverse reasoning questions for these situations. These questions examine a wide spectrum of reasoning capabilities for an intelligent agent, ranging from spatial relation comprehension to commonsense understanding, navigation, and multi-hop reasoning. SQA3D imposes a significant challenge to current multi-modal especially 3D reasoning models. We evaluate various state-of-the-art approaches and find that the best one only achieves an overall score of 47.20%, while amateur human participants can reach 90.06%. We believe SQA3D could facilitate future embodied AI research with stronger situation understanding and reasoning capability.
翻译:我们提出一项新任务以基准化具身智能体的场景理解能力:三维场景中的情境化问答(SQA3D)。给定一个场景上下文(例如三维扫描),SQA3D要求被测试的智能体首先理解其在三维场景中由文本描述的情境(位置、朝向等),然后推理其周围环境并回答该情境下的问题。基于ScanNet中的650个场景,我们提供了一个以约6800个独特情境为中心的数据集,包含2.04万条描述和3.34万个多样化推理问题。这些问题考察了智能体的广泛推理能力,从空间关系理解到常识推理、导航和多跳推理。SQA3D对当前多模态尤其是三维推理模型提出了重大挑战。我们评估了多种最先进方法,发现最佳方法仅达到47.20%的整体得分,而业余人类参与者可达到90.06%。我们相信SQA3D将推动未来具身人工智能研究具备更强的情境理解与推理能力。