We study the problem of estimating the distribution of the return of a policy using an offline dataset that is not generated from the policy, i.e., distributional offline policy evaluation (OPE). We propose an algorithm called Fitted Likelihood Estimation (FLE), which conducts a sequence of Maximum Likelihood Estimation (MLE) and has the flexibility of integrating any state-of-the-art probabilistic generative models as long as it can be trained via MLE. FLE can be used for both finite-horizon and infinite-horizon discounted settings where rewards can be multi-dimensional vectors. Our theoretical results show that for both finite-horizon and infinite-horizon discounted settings, FLE can learn distributions that are close to the ground truth under total variation distance and Wasserstein distance, respectively. Our theoretical results hold under the conditions that the offline data covers the test policy's traces and that the supervised learning MLE procedures succeed. Experimentally, we demonstrate the performance of FLE with two generative models, Gaussian mixture models and diffusion models. For the multi-dimensional reward setting, FLE with diffusion models is capable of estimating the complicated distribution of the return of a test policy.
翻译:我们研究了利用非目标策略生成的离线数据集来估计策略回报分布的问题,即分布离线策略评估(OPE)。我们提出了一种名为拟合似然估计(FLE)的算法,该算法执行一系列最大似然估计(MLE),并具有整合任意最先进概率生成模型的灵活性,前提是该模型可通过MLE训练。FLE可同时应用于有限时域和无限时域折扣奖励场景,且支持多维奖励向量。理论结果表明,在有限时域和无限时域折扣场景下,FLE能够分别在全变差距离和Wasserstein距离度量下学习到接近真实值的分布。我们的理论结果在以下条件下成立:离线数据覆盖了测试策略的轨迹,且监督学习MLE过程成功实现。实验方面,我们通过两种生成模型(高斯混合模型和扩散模型)验证了FLE的性能。在多维奖励场景中,结合扩散模型的FLE能够有效估计测试策略复杂回报分布。