In general-sum games, the interaction of self-interested learning agents commonly leads to socially worse outcomes, such as defect-defect in the iterated stag hunt (ISH). Previous works address this challenge by sharing rewards or shaping their opponents' learning process, which require too strong assumptions. In this paper, we demonstrate that agents trained to optimize expected returns are more likely to choose a safe action that leads to guaranteed but lower rewards. However, there typically exists a risky action that leads to higher rewards in the long run only if agents cooperate, e.g., cooperate-cooperate in ISH. To overcome this, we propose using action value distribution to characterize the decision's risk and corresponding potential payoffs. Specifically, we present Adaptable Risk-Sensitive Policy (ARSP). ARSP learns the distributions over agent's return and estimates a dynamic risk-seeking bonus to discover risky coordination strategies. Furthermore, to avoid overfitting training opponents, ARSP learns an auxiliary opponent modeling task to infer opponents' types and dynamically alter corresponding strategies during execution. Empirically, agents trained via ARSP can achieve stable coordination during training without accessing opponent's rewards or learning process, and can adapt to non-cooperative opponents during execution. To the best of our knowledge, it is the first method to learn coordination strategies between agents both in iterated prisoner's dilemma (IPD) and iterated stag hunt (ISH) without shaping opponents or rewards, and can adapt to opponents with distinct strategies during execution. Furthermore, we show that ARSP can be scaled to high-dimensional settings.
翻译:在一般和博弈中,自利学习智能体的互动通常会导致社会性较差的结果,例如在迭代猎鹿博弈(ISH)中的背叛-背叛行为。先前的研究通过共享奖励或塑造对手的学习过程来应对这一挑战,但这些方法需要过于严格的假设。本文中,我们表明,优化期望回报的智能体更倾向于选择能带来有保障但较低奖励的安全动作。然而,通常存在一种风险动作,只有在智能体合作时才能长期获得更高奖励,例如在ISH中的合作-合作行为。为克服这一问题,我们提出利用动作价值分布来表征决策的风险及相应潜在收益。具体而言,我们提出了可适应性风险敏感策略(ARSP)。ARSP学习智能体回报的分布,并估计动态风险寻求奖励,以发现风险性协调策略。此外,为避免过度适应训练对手,ARSP学习一个辅助的对手建模任务来推断对手类型,并在执行过程中动态调整相应策略。实验表明,通过ARSP训练的智能体在训练过程中无需获取对手奖励或学习过程即可实现稳定协调,并在执行过程中能适应非合作对手。据我们所知,这是首个无需塑造对手或奖励即可在迭代囚徒困境(IPD)和迭代猎鹿博弈(ISH)中学习智能体间协调策略的方法,且能在执行过程中适应具有不同策略的对手。此外,我们证明ARSP可扩展到高维设置。