Many settings of interest involving humans and machines -- from virtual personal assistants to autonomous vehicles -- can naturally be modelled as principals (humans) delegating to agents (machines), which then interact with each other on their principals' behalf. We refer to these multi-principal, multi-agent scenarios as delegation games. In such games, there are two important failure modes: problems of control (where an agent fails to act in line their principal's preferences) and problems of cooperation (where the agents fail to work well together). In this paper we formalise and analyse these problems, further breaking them down into issues of alignment (do the players have similar preferences?) and capabilities (how competent are the players at satisfying those preferences?). We show -- theoretically and empirically -- how these measures determine the principals' welfare, how they can be estimated using limited observations, and thus how they might be used to help us design more aligned and cooperative AI systems.
翻译:许多涉及人类与机器的有趣场景——从虚拟个人助手到自动驾驶汽车——都可以自然地建模为委托人(人类)将任务委托给代理人(机器),代理人随后代表其委托人进行交互。我们将这种多委托人、多代理人的场景称为委托博弈。在此类博弈中,存在两种重要的失效模式:控制问题(代理人未能按照委托人的偏好行事)与合作问题(代理人之间未能良好协作)。本文对这些问题进行形式化分析与解构,进一步将其划分为对齐问题(参与者是否具有相似偏好?)与能力问题(参与者在满足这些偏好方面的胜任程度如何?)。我们通过理论与实证研究表明:这些度量如何决定委托人的福祉,如何利用有限观测对其进行估计,以及如何借此指导我们设计更具对齐性与协作性的AI系统。