Instilling virtuous behavior in artificial intelligence has seen increasing interest. One of the techniques proposed is known as affinity-based reinforcement learning, which uses policy regularization on the objective function to incentivize virtuous actions without being fully dependent on the reward function design. Thus far, this technique has been demonstrated to be effective in grid worlds and toy-problem environments with minimal state and action spaces. To expand this research to more sophisticated environments, we introduce a two-player multi-agent environment based on the role-playing board game known as Fog of Love. In this environment, two agents compete to fulfill their individual virtues, while also cooperating to satisfy their relationship. Given the multi-agent nature, this is a complex problem where multi-agent deep deterministic policy gradient agents neither compete nor cooperate successfully. We present evidence that localized affinities enhance agent performance in achieving both competitive and cooperative objectives, resulting from superior overall scores in both domains. This not only results in virtuous choices but also clarifies an agent's teleology and makes its behavior human-level interpretable.
翻译:人工智能注入美德行为的研究日益受到关注。其中一种提出的技术称为基于亲和力的强化学习,该方法通过对目标函数进行策略正则化,在不完全依赖奖励函数设计的情况下激励美德行为。迄今为止,该技术已在具有极小状态和动作空间的网格世界及简单问题环境中被证明有效。为将研究拓展至更复杂的环境,我们引入了一个基于角色扮演桌游《爱之迷雾》的双智能体多智能体环境。在该环境中,两个代理在合作维系关系的同时,竞争追求各自的个体美德。由于多智能体的特性,这是一个复杂问题,其中多智能体深度确定性策略梯度代理既无法成功竞争也无法实现有效合作。我们提供的证据表明,局部亲和力增强了代理在竞争与合作双重目标上的表现,体现在两个领域均获得更优的总分。这不仅产生美德选择,还明晰了代理的目的论,并使其行为达到可被人类理解的解释水平。