While there now exists a large literature on policy evaluation and learning, much of prior work assumes that the treatment assignment of one unit does not affect the outcome of another unit. Unfortunately, ignoring interference may lead to biased policy evaluation and ineffective learned policies. For example, treating influential individuals who have many friends can generate positive spillover effects, thereby improving the overall performance of an individualized treatment rule (ITR). We consider the problem of evaluating and learning an optimal ITR under clustered network interference (also known as partial interference) where clusters of units are sampled from a population and units may influence one another within each cluster. Unlike previous methods that impose strong restrictions on spillover effects, the proposed methodology only assumes a semiparametric structural model where each unit's outcome is an additive function of individual treatments within the cluster. Under this model, we propose an estimator that can be used to evaluate the empirical performance of an ITR. We show that this estimator is substantially more efficient than the standard inverse probability weighting estimator, which does not impose any assumption about spillover effects. We derive the finite-sample regret bound for a learned ITR, showing that the use of our efficient evaluation estimator leads to the improved performance of learned policies. Finally, we conduct simulation and empirical studies to illustrate the advantages of the proposed methodology.
翻译:尽管目前已有大量关于政策评估与学习的文献,但多数先前研究假设某个单元的干预分配不会影响另一单元的结果。遗憾的是,忽略干扰效应可能导致有偏的政策评估和低效的学习策略。例如,对拥有众多朋友的影响力个体进行干预可产生正向溢出效应,从而提升个性化治疗方案规则的总体表现。本文探讨在集群网络干扰(亦称部分干扰)情境下评估与学习最优个性化治疗方案规则的问题,其中集群单元从总体中抽样,且同一集群内的单元可能相互影响。与以往对溢出效应施加严格限制的方法不同,本文提出的方法仅假设一种半参数结构模型,其中每个单元的结果是该集群内个体干预的加性函数。在该模型框架下,我们提出一种可用于评估个性化治疗方案规则经验表现的估计量。研究表明,该估计量的效率显著高于不施加任何溢出效应假设的标准逆概率加权估计量。我们推导了学习所得个性化治疗方案规则的有限样本遗憾界,证明使用高效评估估计量可提升学习策略的表现。最后,通过仿真与实证研究验证了所提方法的优势。