We study the problem of offline policy optimization in stochastic contextual bandit problems, where the goal is to learn a near-optimal policy based on a dataset of decision data collected by a suboptimal behavior policy. Rather than making any structural assumptions on the reward function, we assume access to a given policy class and aim to compete with the best comparator policy within this class. In this setting, a standard approach is to compute importance-weighted estimators of the value of each policy, and select a policy that minimizes the estimated value up to a "pessimistic" adjustment subtracted from the estimates to reduce their random fluctuations. In this paper, we show that a simple alternative approach based on the "implicit exploration" estimator of \citet{Neu2015} yields performance guarantees that are superior in nearly all possible terms to all previous results. Most notably, we remove an extremely restrictive "uniform coverage" assumption made in all previous works. These improvements are made possible by the observation that the upper and lower tails importance-weighted estimators behave very differently from each other, and their careful control can massively improve on previous results that were all based on symmetric two-sided concentration inequalities. We also extend our results to infinite policy classes in a PAC-Bayesian fashion, and showcase the robustness of our algorithm to the choice of hyper-parameters by means of numerical simulations.
翻译:我们研究了随机上下文赌博机问题中的离线策略优化问题,其目标是基于由次优行为策略收集的决策数据集学习近优策略。不同于对奖励函数施加任何结构性假设,我们假设可以访问给定的策略类,并旨在与该类中的最优比较策略竞争。在此设置下,标准方法是计算每个策略值的重要性加权估计量,并选择使估计值最小化(减去“悲观”调整项以减少随机波动)的策略。本文表明,基于\citet{Neu2015}的“隐式探索”估计量的简单替代方法,在几乎所有方面均能提供优于先前结果的性能保证。最显著的是,我们消除了以往所有工作中极为严格的“均匀覆盖”假设。这些改进源于观察到重要性加权估计量的上尾与下尾行为存在显著差异,通过精细控制它们,可以大幅改进以往基于对称双边集中不等式的结果。我们还将结果以PAC-Bayesian方式扩展到无限策略类,并通过数值模拟展示了算法对超参数选择的鲁棒性。