Despite rapid progress, embodied agents still struggle with long-horizon manipulation that requires maintaining spatial consistency, causal dependencies, and goal constraints. A key limitation of existing approaches is that task reasoning is implicitly embedded in high-dimensional latent representations, making it challenging to separate task structure from perceptual variability. We introduce Grounded Scene-graph Reasoning (GSR), a structured reasoning paradigm that explicitly models world-state evolution as transitions over semantically grounded scene graphs. By reasoning step-wise over object states and spatial relations, rather than directly mapping perception to actions, GSR enables explicit reasoning about action preconditions, consequences, and goal satisfaction in a physically grounded space. To support learning such reasoning, we construct Manip-Cognition-1.6M, a large-scale dataset that jointly supervises world understanding, action planning, and goal interpretation. Extensive evaluations across RLBench, LIBERO, GSR-benchmark, and real-world robotic tasks show that GSR significantly improves zero-shot generalization and long-horizon task completion over prompting-based baselines. These results highlight explicit world-state representations as a key inductive bias for scalable embodied reasoning.


翻译:尽管发展迅速,具身智能体在执行需要保持空间一致性、因果依赖关系和目标约束的长时程操作任务时仍面临困难。现有方法的一个关键局限在于任务推理被隐式嵌入高维潜在表示中,这使得任务结构与感知可变性难以分离。我们提出基于场景图的接地推理范式,该结构化推理范式将世界状态演化显式建模为语义接地的场景图上的状态转移。通过逐步推理对象状态与空间关系,而非直接将感知映射为动作,GSR能够在物理接地空间中显式推理动作前提条件、执行结果及目标达成状态。为支持此类推理的学习,我们构建了Manip-Cognition-1.6M大规模数据集,该数据集联合监督世界理解、动作规划与目标解析。在RLBench、LIBERO、GSR基准测试及真实机器人任务上的广泛实验表明,相较于基于提示的基线方法,GSR在零样本泛化与长时程任务完成方面取得显著提升。这些结果凸显了显式世界状态表示作为可扩展具身推理的关键归纳偏置。

0
下载
关闭预览

相关内容

数据驱动的具身学习探索
专知会员服务
18+阅读 · 2025年2月26日
【博士论文】推理的表示学习:跨多样结构的泛化
专知会员服务
27+阅读 · 2024年10月20日
【牛津大学博士论文】深度具身智能体的空间推理与规划
【斯坦福博士论文】具身物体搜索的操作与推理方法
专知会员服务
39+阅读 · 2023年9月13日
【DTU博士论文】结构化表示学习的泛化
专知会员服务
54+阅读 · 2023年4月27日
「可解释知识图谱推理」最新方法综述
专知会员服务
89+阅读 · 2022年12月17日
可解释强化学习,Explainable Reinforcement Learning: A Survey
专知会员服务
133+阅读 · 2020年5月14日
理解人类推理的深度学习
论智
19+阅读 · 2018年11月7日
关系推理:基于表示学习和语义要素
计算机研究与发展
19+阅读 · 2017年8月22日
国家自然科学基金
44+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
16+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
VIP会员
最新内容
刚刚!Jev中文教程项目发布了
专知会员服务
0+阅读 · 10月4日
《人工智能赋能的适应性多功能电磁战》
专知会员服务
12+阅读 · 9月29日
俄乌战场实验室:全面战争如何重塑现代作战
专知会员服务
8+阅读 · 9月29日
2026年美空军协会会议上的无人机系统趋势
专知会员服务
11+阅读 · 9月28日
反制无人机:乌克兰提供的五点启示
专知会员服务
17+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
11+阅读 · 9月23日
相关VIP内容
数据驱动的具身学习探索
专知会员服务
18+阅读 · 2025年2月26日
【博士论文】推理的表示学习:跨多样结构的泛化
专知会员服务
27+阅读 · 2024年10月20日
【牛津大学博士论文】深度具身智能体的空间推理与规划
【斯坦福博士论文】具身物体搜索的操作与推理方法
专知会员服务
39+阅读 · 2023年9月13日
【DTU博士论文】结构化表示学习的泛化
专知会员服务
54+阅读 · 2023年4月27日
「可解释知识图谱推理」最新方法综述
专知会员服务
89+阅读 · 2022年12月17日
可解释强化学习,Explainable Reinforcement Learning: A Survey
专知会员服务
133+阅读 · 2020年5月14日
相关基金
国家自然科学基金
44+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
16+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员