Off-policy learning methods are intended to learn a policy from logged data, which includes context, action, and feedback (cost or reward) for each sample point. In this work, we build on the counterfactual risk minimization framework, which also assumes access to propensity scores. We propose learning methods for problems where feedback is missing for some samples, so there are samples with feedback and samples missing-feedback in the logged data. We refer to this type of learning as semi-supervised batch learning from logged data, which arises in a wide range of application domains. We derive a novel upper bound for the true risk under the inverse propensity score estimator to address this kind of learning problem. Using this bound, we propose a regularized semi-supervised batch learning method with logged data where the regularization term is feedback-independent and, as a result, can be evaluated using the logged missing-feedback data. Consequently, even though feedback is only present for some samples, a learning policy can be learned by leveraging the missing-feedback samples. The results of experiments derived from benchmark datasets indicate that these algorithms achieve policies with better performance in comparison with logging policies.
翻译:离线策略学习方法的目的是从包含上下文、动作和每个样本点反馈(成本或奖励)的日志数据中学习策略。本研究基于反事实风险最小化框架,该框架同样假设可以获取倾向性得分。我们针对部分样本缺失反馈的问题提出了学习方法,因此在日志数据中既有包含反馈的样本,也有缺失反馈的样本。我们将此类学习称为基于日志数据的半监督批量学习,其在广泛的应用领域中均有出现。我们基于逆倾向性得分估计器推导出了真实风险的一个新上界,以解决此类学习问题。利用这一上界,我们提出了一种基于日志数据的正则化半监督批量学习方法,其中正则化项与反馈无关,因此可利用日志中缺失反馈的数据进行评估。这样一来,即使只有部分样本包含反馈,也能通过利用缺失反馈的样本来学习策略。基于基准数据集进行的实验结果表明,与日志策略相比,这些算法能够学习到性能更优的策略。