Machine learning provides a powerful tool for building socially compliant robotic systems that go beyond simple predictive models of human behavior. By observing and understanding human interactions from past experiences, learning can enable effective social navigation behaviors directly from data. However, collecting navigation data in human-occupied environments may require teleoperation or continuous monitoring, making the process prohibitively expensive to scale. In this paper, we present a scalable data collection system for vision-based navigation, SACSoN, that can autonomously navigate around pedestrians in challenging real-world environments while encouraging rich interactions. SACSoN uses visual observations to observe and react to humans in its vicinity. It couples this visual understanding with continual learning and an autonomous collision recovery system that limits the involvement of a human operator, allowing for better dataset scaling. We use a this system to collect the SACSoN dataset, the largest-of-its-kind visual navigation dataset of autonomous robots operating in human-occupied spaces, spanning over 75 hours and 4000 rich interactions with humans. Our experiments show that collecting data with a novel objective that encourages interactions, leads to significant improvements in downstream tasks such as inferring pedestrian dynamics and learning socially compliant navigation behaviors. We make videos of our autonomous data collection system and the SACSoN dataset publicly available on our project page.
翻译:机器学习为构建超越简单人类行为预测模型的社交合规机器人系统提供了强大工具。通过从过往经验中观测和理解人类交互,学习过程能够直接从数据中获取有效的社交导航行为。然而,在有人类活动的环境中采集导航数据通常需要遥操作或持续监控,这使得该流程因成本过高而难以规模化。本文提出了一种基于视觉导航的可扩展数据采集系统SACSoN,该系统能在充满挑战的真实环境中自主绕行行人,同时促进丰富的人机交互。SACSoN通过视觉观测感知并响应周围人类,并将这种视觉理解与持续学习及自主碰撞恢复系统相结合,从而减少对人工操作员的依赖,实现更好的数据集扩展。我们利用该系统构建了SACSoN数据集——这是同类中规模最大的、在人类活动空间运行的自主机器人视觉导航数据集,包含超过75小时时长及4000次与人类的丰富交互。实验表明,采用鼓励交互的新型目标函数进行数据采集,能显著提升下游任务性能,包括行人动力学推断与社交合规导航行为的学习。我们已在项目页面公开发布自主数据采集系统演示视频及SACSoN数据集。