Recently, reinforcement learning has gained prominence in modern statistics, with policy evaluation being a key component. Unlike traditional machine learning literature on this topic, our work places emphasis on statistical inference for the parameter estimates computed using reinforcement learning algorithms. While most existing analyses assume random rewards to follow standard distributions, limiting their applicability, we embrace the concept of robust statistics in reinforcement learning by simultaneously addressing issues of outlier contamination and heavy-tailed rewards within a unified framework. In this paper, we develop an online robust policy evaluation procedure, and establish the limiting distribution of our estimator, based on its Bahadur representation. Furthermore, we develop a fully-online procedure to efficiently conduct statistical inference based on the asymptotic distribution. This paper bridges the gap between robust statistics and statistical inference in reinforcement learning, offering a more versatile and reliable approach to policy evaluation. Finally, we validate the efficacy of our algorithm through numerical experiments conducted in real-world reinforcement learning experiments.
翻译:近年来,强化学习在现代统计学中日益凸显其重要性,其中策略评估是关键组成部分。与现有机器学习文献不同,本研究聚焦于强化学习算法参数估计的统计推断。现有分析大多假设随机奖励服从标准分布,这限制了其适用性;而我们通过统一框架同时解决异常值污染和重尾奖励问题,将稳健统计概念引入强化学习。本文提出一种在线稳健策略评估方法,并基于其巴哈杜尔表示建立估计量的极限分布。此外,我们开发了完全在线程序,利用渐近分布高效进行统计推断。本研究弥合了强化学习中稳健统计与统计推断之间的鸿沟,为策略评估提供更通用可靠的方法。最后,通过实际强化学习实验的数值研究验证了我们算法的有效性。