We study sequential decision-making in partially observable environments against strategic, adaptive opponents, modeled as partially observable Markov games (POMGs). The central challenge is to learn latent dynamics from partial observations while facing an adversary whose behavior depends on the learner's strategy, making standard regret notions inadequate. We prove that an epoch-based optimistic maximum-likelihood algorithm achieves $\tilde{O}(\sqrt{T})$ policy regret for fixed problem parameters, with explicit dependence on the horizon, adversary memory, confidence radius, and the aggregate Eluder dimension of the observable-operator class. The algorithm selects one policy per geometrically growing epoch using confidence sets built cumulatively from past data, which keeps the cost of comparing adversary responses across policies logarithmic in $T$. We also prove a lower bound matching the $\sqrt{T}$ and aggregate-Eluder-dimension dependence, up to problem-dependent and logarithmic factors. Finally, we extend the framework to horizon-adaptive guarantees and adversaries with geometric fading memory.


翻译:我们研究部分可观测环境下面对策略性自适应对手的序贯决策问题,该问题被建模为部分可观测马尔可夫博弈(POMG)。核心挑战在于:在面临行为依赖于学习者策略的对手时,需从部分观测中学习潜在动态,这使得标准遗憾定义不再适用。我们证明,基于轮次的乐观最大似然算法在固定问题参数下可实现$\tilde{O}(\sqrt{T})$的策略遗憾,其显式依赖项包括时域长度、对手记忆长度、置信半径以及可观测算子类的聚合埃尔德维度。该算法通过基于历史数据累积构建的置信集,在每个几何增长的轮次中选择单一策略,从而将跨策略比较对手响应的代价控制在$T$的对数阶内。我们还证明了与$\sqrt{T}$及聚合埃尔德维度依赖匹配的下界(除问题依赖因子与对数因子外)。最后,我们将该框架扩展至时域自适应保证和几何衰减记忆型对手场景。

0
下载
关闭预览

相关内容

《分布式多智能体强化学习策略的可解释性研究》
专知会员服务
31+阅读 · 2025年11月17日
《主观概率约束下寻找可行系统及其军事应用》69页
专知会员服务
29+阅读 · 2025年9月27日
《战略智能体与有限反馈下的序贯决策》211页
专知会员服务
38+阅读 · 2025年5月7日
【伯克利博士论文】在部分可观察性下的对齐问题
专知会员服务
20+阅读 · 2025年1月9日
深入理解BERT Transformer ,不仅仅是注意力机制
大数据文摘
22+阅读 · 2019年3月19日
不用数学讲清马尔可夫链蒙特卡洛方法?
算法与数学之美
16+阅读 · 2018年8月8日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 5月31日
VIP会员
最新内容
致命七类无人机:无人机时代的演进型合成兵种
专知会员服务
1+阅读 · 13分钟前
《异构无人水面艇集群作战自主制导算法》130页
《人工智能能通过美国陆军战争学院吗?》报告
军事域人工智能驱动系统的治理
专知会员服务
4+阅读 · 9月14日
相关基金
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员