Robotic mapping systems typically approach building metric-semantic scene representations from the robot's own sensors and cameras. However, these "first person" maps inherit the robot's own limitations due to its embodiment or skillset, which may leave many aspects of the environment unexplored. For example, the robot might not be able to open drawers or access wall cabinets. In this sense, the map representation is not as complete, and requires a more capable robot to fill in the gaps. We narrow these blind spots in current methods by leveraging egocentric data captured as a human naturally explores a scene wearing Project Aria glasses, giving a way to directly transfer knowledge about articulation from the human to any deployable robot. We demonstrate that, by using simple heuristics, we can leverage egocentric data to recover models of articulate object parts, with quality comparable to those of state-of-the-art methods based on other input modalities. We also show how to integrate these models into 3D scene graph representations, leading to a better understanding of object dynamics and object-container relationships. We finally demonstrate that these articulated 3D scene graphs enhance a robot's ability to perform mobile manipulation tasks, showcasing an application where a Boston Dynamics Spot is tasked with retrieving concealed target items, given only the 3D scene graph as input.
翻译:机器人建图系统通常依赖自身传感器与摄像头构建度量-语义场景表征。然而,这类"第一人称"地图受限于机器人本体的具身能力与技能集,可能导致环境中诸多区域无法被探索,例如机器人可能无法打开抽屉或触及壁柜。在此意义上,地图表征并不完备,需要能力更强的机器人填补这些空白。本文通过利用人类佩戴Project Aria眼镜自然探索场景时捕获的第一人称数据,缩小了现有方法中的盲区,为人机间铰接知识直接迁移提供了途径。实验证明,基于简单启发式规则,我们能够利用第一人称数据恢复可铰接物体部件的模型,其质量可与基于其他输入模态的先进方法相媲美。我们还展示了如何将这些模型集成至三维场景图表示中,从而更深入地理解物体动力学特性与容器间空间关联。最终验证表明,这些铰接式三维场景图能增强机器人执行移动操作任务的能力,并以波士顿动力公司Spot机器人为例,展示了其在仅输入三维场景图的情况下检索隐藏目标物品的应用场景。