We consider the problem of learning a control policy that is robust against the parameter mismatches between the training environment and testing environment. We formulate this as a distributionally robust reinforcement learning (DR-RL) problem where the objective is to learn the policy which maximizes the value function against the worst possible stochastic model of the environment in an uncertainty set. We focus on the tabular episodic learning setting where the algorithm has access to a generative model of the nominal (training) environment around which the uncertainty set is defined. We propose the Robust Phased Value Learning (RPVL) algorithm to solve this problem for the uncertainty sets specified by four different divergences: total variation, chi-square, Kullback-Leibler, and Wasserstein. We show that our algorithm achieves $\tilde{\mathcal{O}}(|\mathcal{S}||\mathcal{A}| H^{5})$ sample complexity, which is uniformly better than the existing results by a factor of $|\mathcal{S}|$, where $|\mathcal{S}|$ is number of states, $|\mathcal{A}|$ is the number of actions, and $H$ is the horizon length. We also provide the first-ever sample complexity result for the Wasserstein uncertainty set. Finally, we demonstrate the performance of our algorithm using simulation experiments.
翻译:我们研究在训练环境与测试环境的参数不匹配下学习鲁棒控制策略的问题。我们将此问题建模为分布鲁棒强化学习(DR-RL)问题,其目标是在一个不确定性集合中,针对环境最坏可能的随机模型,学习能最大化值函数的策略。我们重点研究表格型情节式学习场景,其中算法能够访问名义(训练)环境(围绕其定义不确定性集合)的生成模型。我们提出鲁棒分段值学习(RPVL)算法,用于解决由四种不同散度(全变差、卡方、KL散度和Wasserstein)指定不确定性集合下的该问题。我们证明该算法达到$\tilde{\mathcal{O}}(|\mathcal{S}||\mathcal{A}| H^{5})$的样本复杂度,相比于现有结果在系数$|\mathcal{S}|$上一致更优,其中$|\mathcal{S}|$是状态数,$|\mathcal{A}|$是动作数,$H$是情节长度。我们还首次给出了Wasserstein不确定性集合的样本复杂度结果。最后,通过仿真实验展示了算法的性能。