Modern Reinforcement Learning (RL) algorithms are able to outperform humans in a wide variety of tasks. Multi-agent reinforcement learning (MARL) settings present additional challenges, and successful cooperation in mixed-motive groups of agents depends on a delicate balancing act between individual and group objectives. Social conventions and norms, often inspired by human institutions, are used as tools for striking this balance. In this paper, we examine a fundamental, well-studied social convention that underlies cooperation in both animal and human societies: dominance hierarchies. We adapt the ethological theory of dominance hierarchies to artificial agents, borrowing the established terminology and definitions with as few amendments as possible. We demonstrate that populations of RL agents, operating without explicit programming or intrinsic rewards, can invent, learn, enforce, and transmit a dominance hierarchy to new populations. The dominance hierarchies that emerge have a similar structure to those studied in chickens, mice, fish, and other species.
翻译:现代强化学习算法在多种任务中已能超越人类表现。多智能体强化学习场景带来了额外挑战,而混合动机智能体群体中的成功合作取决于个体目标与群体目标间的精妙平衡。常受人类制度启发的社会惯例与规范,正是实现这一平衡的工具。本文研究了一个基础且经过深入探讨的社会惯例——主导等级,该惯例支撑着动物与人类社会中的合作行为。我们将生态学主导等级理论适配至人工智能体,在最小化改动的前提下沿用既有术语与定义。研究表明,未经显式编程或内在奖励设计的强化学习智能体群体,能够自主创造、学习、强制执行并跨群体传播主导等级。这些涌现出的主导等级结构与鸡、鼠、鱼等物种中研究的层级结构具有相似性。