In this paper, we study regret minimization in repeated games with \emph{adaptive} opponents who can respond based on histories of play. The standard metric of \emph{external regret} in online learning is known to fail to capture such adaptivity. To account for players' counterfactual reasoning, we introduce {\tt Repeated Policy Regret (RP-Regret)}, a game-theoretic metric that measures the difference between the \emph{realized} and the \emph{best-in-hindsight} accumulated utility when all players can \emph{respond} to the history of play. Compared to existing regret notions in this setting, ours is native to repeated game playing, enabling stronger comparators and opponents with fewer constraints, while maintaining the possibility of finding better equilibria when all players minimize it. We first identify necessary conditions for obtaining {\tt RP-Regret} sublinear in time, on the variation of the player's comparator strategies in the regret definition and on the memories of both the comparator and opponents' strategies. We then study additional conditions and provable algorithms to minimize {\tt RP-Regret}, which is by definition \emph{non-convex} in the strategy space. To address this challenge, we propose three algorithms: (i) one based on an optimization oracle, as assumed in some prior work in online non-convex learning; (ii) one that minimizes a convex and \emph{linearized} surrogate of {\tt RP-Regret} at each iteration; (iii) one that directly minimizes {\tt RP-Regret} when opponents change strategies slowly. Furthermore, when all players can run algorithms to minimize the {\tt RP-Regret} (or its linearized variant), certain subgame perfect equilibria of the repeated game can be learned. We also provide experiments showing that minimizing our regret notions can lead to more cooperative solutions with higher utility in games such as Stag-Hunt.


翻译:本文研究重复博弈中面对**自适应**对手(即能依据历史博弈过程做出响应的对手)时的遗憾最小化问题。在线学习中的标准**外部遗憾**指标已被证明无法刻画这种自适应性。为体现玩家的反事实推理,我们提出**重复策略遗憾(RP-Regret)**——一种博弈论度量指标,衡量当所有玩家均可对历史博弈过程**做出响应**时,其**实际累积效用**与**事后最优累积效用**之间的差异。相较于该领域的现有遗憾概念,本文提出的度量天然适用于重复博弈情境,既能支持更强的比较对象与约束更少的对手,又能在所有玩家均最小化该遗憾时保留发现更优均衡的可能性。我们首先确定了实现时间次线性的**RP-Regret**所需的必要条件:涉及遗憾定义中玩家比较策略的变化幅度,以及比较策略与对手策略的记忆长度。随后,我们研究额外条件并设计可证明的算法以最小化**RP-Regret**——该目标在策略空间中天然具有**非凸性**。为应对这一挑战,我们提出三种算法:(一)基于优化预言机(部分在线非凸学习研究中的假设)的算法;(二)每轮迭代最小化**RP-Regret**的凸**线性化**代理项的算法;(三)当对手策略缓慢变化时直接最小化**RP-Regret**的算法。进一步地,当所有玩家均可运行最小化**RP-Regret**(或其线性化变体)的算法时,重复博弈的某些子博弈完美均衡可被学习。实验表明,在诸如猎鹿博弈中,最小化本文提出的遗憾概念可引导出具有更高效用的合作解。

0
下载
关闭预览

相关内容

智能博弈对抗方法:博弈论与强化学习综合视角对比分析
专知会员服务
199+阅读 · 2022年8月28日
【CVPR2020-北京大学】自适应间隔损失的提升小样本学习
专知会员服务
85+阅读 · 2020年6月9日
面向多智能体博弈对抗的对手建模框架
专知
18+阅读 · 2022年9月28日
机器学习中的最优化算法总结
人工智能前沿讲习班
22+阅读 · 2019年3月22日
自定义损失函数Gradient Boosting
AI研习社
14+阅读 · 2018年10月16日
推荐算法:Match与Rank模型的交织配合
从0到1
15+阅读 · 2017年12月18日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Arxiv
0+阅读 · 6月2日
Arxiv
0+阅读 · 5月13日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员