We study infinite-horizon average-reward Markov decision processes (AMDPs) in the context of general function approximation. Specifically, we propose a novel algorithmic framework named Local-fitted Optimization with OPtimism (LOOP), which incorporates both model-based and value-based incarnations. In particular, LOOP features a novel construction of confidence sets and a low-switching policy updating scheme, which are tailored to the average-reward and function approximation setting. Moreover, for AMDPs, we propose a novel complexity measure -- average-reward generalized eluder coefficient (AGEC) -- which captures the challenge of exploration in AMDPs with general function approximation. Such a complexity measure encompasses almost all previously known tractable AMDP models, such as linear AMDPs and linear mixture AMDPs, and also includes newly identified cases such as kernel AMDPs and AMDPs with Bellman eluder dimensions. Using AGEC, we prove that LOOP achieves a sublinear $\tilde{\mathcal{O}}(\mathrm{poly}(d, \mathrm{sp}(V^*)) \sqrt{T\beta} )$ regret, where $d$ and $\beta$ correspond to AGEC and log-covering number of the hypothesis class respectively, $\mathrm{sp}(V^*)$ is the span of the optimal state bias function, $T$ denotes the number of steps, and $\tilde{\mathcal{O}} (\cdot) $ omits logarithmic factors. When specialized to concrete AMDP models, our regret bounds are comparable to those established by the existing algorithms designed specifically for these special cases. To the best of our knowledge, this paper presents the first comprehensive theoretical framework capable of handling nearly all AMDPs.
翻译:我们研究通用函数逼近背景下无限时域平均奖励马尔可夫决策过程(AMDPs)。具体而言,我们提出一种名为局部拟合优化置信上界(LOOP)的新型算法框架,该框架同时包含基于模型和基于价值的形式。特别地,LOOP采用专为平均奖励与函数逼近设置设计的置信集构建新方案与低切换策略更新机制。此外,针对AMDPs,我们提出一种新型复杂度度量——平均奖励广义埃尔得系数(AGEC)——该度量刻画了通用函数逼近框架下AMDPs中探索的挑战性。这种复杂度度量几乎涵盖所有先前已知的可处理AMDP模型(如线性AMDPs与线性混合AMDPs),同时也包含新识别的案例(如核AMDPs及具有贝尔曼埃尔得维数的AMDPs)。利用AGEC,我们证明LOOP可实现次线性遗憾界$\tilde{\mathcal{O}}(\mathrm{poly}(d, \mathrm{sp}(V^*)) \sqrt{T\beta} )$,其中$d$与$\beta$分别对应AGEC与假设类的对数覆盖数,$\mathrm{sp}(V^*)$为最优状态偏置函数的跨度,$T$表示步数,$\tilde{\mathcal{O}} (\cdot) $省略对数因子。当具体化到特定AMDP模型时,我们的遗憾界与专为这些特例设计的现有算法所建立的界相当。据我们所知,本文首次提出能够处理几乎所有AMDPs的统一理论框架。