Existing game AI research mainly focuses on enhancing agents' abilities to win games, but this does not inherently make humans have a better experience when collaborating with these agents. For example, agents may dominate the collaboration and exhibit unintended or detrimental behaviors, leading to poor experiences for their human partners. In other words, most game AI agents are modeled in a "self-centered" manner. In this paper, we propose a "human-centered" modeling scheme for collaborative agents that aims to enhance the experience of humans. Specifically, we model the experience of humans as the goals they expect to achieve during the task. We expect that agents should learn to enhance the extent to which humans achieve these goals while maintaining agents' original abilities (e.g., winning games). To achieve this, we propose the Reinforcement Learning from Human Gain (RLHG) approach. The RLHG approach introduces a "baseline", which corresponds to the extent to which humans primitively achieve their goals, and encourages agents to learn behaviors that can effectively enhance humans in achieving their goals better. We evaluate the RLHG agent in the popular Multi-player Online Battle Arena (MOBA) game, Honor of Kings, by conducting real-world human-agent tests. Both objective performance and subjective preference results show that the RLHG agent provides participants better gaming experience.
翻译:现有游戏人工智能研究主要聚焦于提升智能体赢得游戏的能力,但这本质上并不能确保人类在与这些智能体协作时获得更好的体验。例如,智能体可能会主导协作过程并表现出非预期甚至有害的行为,导致人类合作伙伴体验不佳。换言之,大多数游戏人工智能智能体都是以"自我中心"方式建模的。本文提出了一种以人为中心的协作智能体建模方案,旨在提升人类的体验。具体而言,我们将人类体验建模为他们在任务执行过程中期望达成的目标。我们期望智能体在维持原有能力(如赢得游戏)的同时,学习如何提升人类实现这些目标的程度。为此,我们提出了基于人类增益的强化学习(RLHG)方法。该方法引入与人类原始目标达成程度对应的"基线",并鼓励智能体学习能有效提升人类更好实现目标的行为。我们在热门多人在线战术竞技(MOBA)游戏《王者荣耀》中进行了真实人机测试来评估RLHG智能体。客观表现与主观偏好结果均表明,RLHG智能体为参与者提供了更佳的游戏体验。