Crowd simulation is important for video-games design, since it enables to populate virtual worlds with autonomous avatars that navigate in a human-like manner. Reinforcement learning has shown great potential in simulating virtual crowds, but the design of the reward function is critical to achieving effective and efficient results. In this work, we explore the design of reward functions for reinforcement learning-based crowd simulation. We provide theoretical insights on the validity of certain reward functions according to their analytical properties, and evaluate them empirically using a range of scenarios, using the energy efficiency as the metric. Our experiments show that directly minimizing the energy usage is a viable strategy as long as it is paired with an appropriately scaled guiding potential, and enable us to study the impact of the different reward components on the behavior of the simulated crowd. Our findings can inform the development of new crowd simulation techniques, and contribute to the wider study of human-like navigation.
翻译:群体仿真对于电子游戏设计至关重要,因为它能够在虚拟世界中填充以类人方式导航的自主化身。强化学习在模拟虚拟群体方面展现出巨大潜力,但奖励函数的设计对于实现高效且有效的结果至关重要。本文探索了基于强化学习的群体仿真中奖励函数的设计。我们从理论上深入分析了某些奖励函数根据其解析特性所具备的有效性,并通过一系列场景以能量效率为指标进行实证评估。实验表明,直接最小化能量消耗是一种可行的策略,但需要与适当缩放的引导势能相结合;同时,这使我们能够研究不同奖励分量对模拟群体行为的影响。我们的发现可为新群体仿真技术的开发提供启示,并有助于拓展类人导航的广泛研究。