We study off-policy evaluation (OPE) in the problem of slate contextual bandits where a policy selects multi-dimensional actions known as slates. This problem is widespread in recommender systems, search engines, marketing, to medical applications, however, the typical Inverse Propensity Scoring (IPS) estimator suffers from substantial variance due to large action spaces, making effective OPE a significant challenge. The PseudoInverse (PI) estimator has been introduced to mitigate the variance issue by assuming linearity in the reward function, but this can result in significant bias as this assumption is hard-to-verify from observed data and is often substantially violated. To address the limitations of previous estimators, we develop a novel estimator for OPE of slate bandits, called Latent IPS (LIPS), which defines importance weights in a low-dimensional slate abstraction space where we optimize slate abstractions to minimize the bias and variance of LIPS in a data-driven way. By doing so, LIPS can substantially reduce the variance of IPS without imposing restrictive assumptions on the reward function structure like linearity. Through empirical evaluation, we demonstrate that LIPS substantially outperforms existing estimators, particularly in scenarios with non-linear rewards and large slate spaces.
翻译:我们研究版面情境赌博机问题中的离线策略评估(OPE),其中策略选择称为版面的多维动作。该问题广泛存在于推荐系统、搜索引擎、市场营销乃至医疗应用中,然而典型的逆倾向得分(IPS)估计器由于动作空间巨大而产生显著方差,使得有效的OPE成为重大挑战。伪逆(PI)估计器通过假设奖励函数的线性来缓解方差问题,但这一假设难以从观测数据验证且常被严重违反,导致显著偏差。为解决先前估计器的局限性,我们提出一种新颖的版面赌博机OPE估计器——潜在IPS(LIPS),该方法在低维版面抽象空间中定义重要性权重,并通过数据驱动方式优化版面抽象以最小化LIPS的偏差和方差。通过该设计,LIPS能在不施加线性等限制性奖励函数结构假设的情况下显著降低IPS的方差。实证评估表明,LIPS在非线性奖励和大版面空间场景中显著优于现有估计器。