Contextual bandit algorithms often estimate reward models to inform decision-making. However, true rewards can contain action-independent redundancies that are not relevant for decision-making. We show it is more data-efficient to estimate any function that explains the reward differences between actions, that is, the treatment effects. Motivated by this observation, building on recent work on oracle-based bandit algorithms, we provide the first reduction of contextual bandits to general-purpose heterogeneous treatment effect estimation, and we design a simple and computationally efficient algorithm based on this reduction. Our theoretical and experimental results demonstrate that heterogeneous treatment effect estimation in contextual bandits offers practical advantages over reward estimation, including more efficient model estimation and greater flexibility to model misspecification.
翻译:上下文赌博机算法常通过估计奖励模型来指导决策。然而,真实奖励可能包含与决策无关的、独立于动作的冗余信息。我们证明,估计能够解释动作间奖励差异的函数(即处理效应)更具数据效率。受此观察启发,并基于近期关于预言机辅助赌博机算法的研究,我们首次将上下文赌博机问题约化为通用异质性处理效应估计,并设计了一种基于此约化的简单且计算高效的算法。理论和实验结果表明,在上下文赌博机中采用异质性处理效应估计相比奖励估计具有实践优势,包括更高效的模型估计及对模型误设的更强鲁棒性。