We study sequential decision-making in partially observable environments against strategic, adaptive opponents, modeled as partially observable Markov games (POMGs). The central challenge is to learn latent dynamics from partial observations while facing an adversary whose behavior depends on the learner's strategy, making standard regret notions inadequate. We prove that an epoch-based optimistic maximum-likelihood algorithm achieves $\tilde{O}(\sqrt{T})$ policy regret for fixed problem parameters, with explicit dependence on the horizon, adversary memory, confidence radius, and the aggregate Eluder dimension of the observable-operator class. The algorithm selects one policy per geometrically growing epoch using confidence sets built cumulatively from past data, which keeps the cost of comparing adversary responses across policies logarithmic in $T$. We also prove a lower bound matching the $\sqrt{T}$ and aggregate-Eluder-dimension dependence, up to problem-dependent and logarithmic factors. Finally, we extend the framework to horizon-adaptive guarantees and adversaries with geometric fading memory.
翻译:我们研究部分可观测环境下面对策略性自适应对手的序贯决策问题,该问题被建模为部分可观测马尔可夫博弈(POMG)。核心挑战在于:在面临行为依赖于学习者策略的对手时,需从部分观测中学习潜在动态,这使得标准遗憾定义不再适用。我们证明,基于轮次的乐观最大似然算法在固定问题参数下可实现$\tilde{O}(\sqrt{T})$的策略遗憾,其显式依赖项包括时域长度、对手记忆长度、置信半径以及可观测算子类的聚合埃尔德维度。该算法通过基于历史数据累积构建的置信集,在每个几何增长的轮次中选择单一策略,从而将跨策略比较对手响应的代价控制在$T$的对数阶内。我们还证明了与$\sqrt{T}$及聚合埃尔德维度依赖匹配的下界(除问题依赖因子与对数因子外)。最后,我们将该框架扩展至时域自适应保证和几何衰减记忆型对手场景。