We investigate the feasibility of deploying reinforcement learning (RL) policies for constrained crowd navigation using a low-fidelity simulator. We introduce a representation of the dynamic environment, separating human and obstacle representations. Humans are represented through detected states, while obstacles are represented as computed point clouds based on maps and robot localization. This representation enables RL policies trained in a low-fidelity simulator to deploy in real world with a reduced sim2real gap. Additionally, we propose a spatio-temporal graph to model the interactions between agents and obstacles. Based on the graph, we use attention mechanisms to capture the robot-human, human-human, and human-obstacle interactions. Our method significantly improves navigation performance in both simulated and real-world environments. Video demonstrations can be found at https://sites.google.com/view/constrained-crowdnav/home.
翻译:本文研究了在低精度仿真器中部署强化学习(RL)策略以实现受限人群导航的可行性。我们提出了一种动态环境表征方法,将人与障碍物的表征分离。其中,行人通过检测状态进行表征,而障碍物则基于地图与机器人定位信息以计算点云的形式表征。该表征方式使得在低精度仿真器中训练的RL策略能够以较小的仿真到现实差距部署至真实环境。此外,我们提出了一种时空图模型来刻画智能体与障碍物之间的交互关系。基于该图结构,我们采用注意力机制来捕捉机器人-行人、行人-行人以及行人-障碍物之间的交互动态。我们的方法在仿真与真实环境中的导航性能均获得显著提升。视频演示可见:https://sites.google.com/view/constrained-crowdnav/home。