In recent years, simulations of pedestrians using the multi-agent reinforcement learning (MARL) have been studied. This study considered the roads on a grid-world environment, and implemented pedestrians as MARL agents using an echo-state network and the least squares policy iteration method. Under this environment, the ability of these agents to learn to move forward by avoiding other agents was investigated. Specifically, we considered two types of tasks: the choice between a narrow direct route and a broad detour, and the bidirectional pedestrian flow in a corridor. The simulations results indicated that the learning was successful when the density of the agents was not that high.
翻译:近年来,基于多智能体强化学习(MARL)的行人模拟研究已得到广泛开展。本研究考虑网格世界环境中的道路,利用回声状态网络与最小二乘策略迭代方法将行人实现为MARL智能体。在该环境下,探究了这些智能体通过学习避开其他智能体从而向前移动的能力。具体而言,我们考虑两类任务:较窄直行路径与较宽绕行路径之间的选择,以及走廊中的双向行人流。仿真结果表明,当智能体密度不高时,学习能够成功完成。