We study Off-Policy Evaluation (OPE) in contextual bandit settings with large action spaces. The benchmark estimators suffer from severe bias and variance tradeoffs. Parametric approaches suffer from bias due to difficulty specifying the correct model, whereas ones with importance weight suffer from variance. To overcome these limitations, Marginalized Inverse Propensity Scoring (MIPS) was proposed to mitigate the estimator's variance via embeddings of an action. To make the estimator more accurate, we propose the doubly robust estimator of MIPS called the Marginalized Doubly Robust (MDR) estimator. Theoretical analysis shows that the proposed estimator is unbiased under weaker assumptions than MIPS while maintaining variance reduction against IPS, which was the main advantage of MIPS. The empirical experiment verifies the supremacy of MDR against existing estimators.
翻译:我们研究大动作空间上下文赌博机设置下的离线策略评估(OPE)。基准估计器存在严重的偏差与方差权衡问题。参数化方法因难以指定正确模型而产生偏差,而基于重要性权重的估计器则面临方差问题。为克服这些限制,研究者提出了边际逆概率评分(MIPS)方法,通过动作嵌入来降低估计器方差。为提升估计精度,我们提出MIPS的双稳健估计器——边际双稳健(MDR)估计器。理论分析表明,在保持IPS方差缩减优势(即MIPS的主要优势)的前提下,所提估计器在比MIPS更弱的假设下无偏。实证实验验证了MDR相较现有估计器的优越性。