In this paper, we propose a new approach called Adaptive Behavioral Costs in Reinforcement Learning (ABC-RL) for training a human-like agent with competitive strength. While deep reinforcement learning agents have recently achieved superhuman performance in various video games, some of these unconstrained agents may exhibit actions, such as shaking and spinning, that are not typically observed in human behavior, resulting in peculiar gameplay experiences. To behave like humans and retain similar performance, ABC-RL augments behavioral limitations as cost signals in reinforcement learning with dynamically adjusted weights. Unlike traditional constrained policy optimization, we propose a new formulation that minimizes the behavioral costs subject to a constraint of the value function. By leveraging the augmented Lagrangian, our approach is an approximation of the Lagrangian adjustment, which handles the trade-off between the performance and the human-like behavior. Through experiments conducted on 3D games in DMLab-30 and Unity ML-Agents Toolkit, we demonstrate that ABC-RL achieves the same performance level while significantly reducing instances of shaking and spinning. These findings underscore the effectiveness of our proposed approach in promoting more natural and human-like behavior during gameplay.
翻译:本文提出了一种名为自适应行为代价强化学习(ABC-RL)的新方法,用于训练具有竞争强度的类人智能体。尽管深度强化学习智能体最近在各种视频游戏中取得了超人类表现,但其中一些无约束智能体可能表现出诸如抖动和旋转等人类行为中不常见的动作,导致怪异的游戏体验。为了像人类一样行动并保持类似的性能,ABC-RL将行为限制作为代价信号引入强化学习,并采用动态调整的权重。与传统的约束策略优化不同,我们提出了一种新公式,该公式在价值函数的约束下最小化行为代价。通过利用增广拉格朗日法,我们的方法是对拉格朗日调整的近似,用于处理性能与类人行为之间的权衡。通过在DMLab-30和Unity ML-Agents工具包的3D游戏上进行的实验,我们证明了ABC-RL在保持相同性能水平的同时,显著减少了抖动和旋转的实例。这些发现强调了所提方法在促进游戏过程中更自然和类人行为方面的有效性。