We consider the problem of learning personalized treatment policies that are externally valid or generalizable: they perform well in other target populations besides the experimental (or training) population from which data are sampled. We first show that welfare-maximizing policies for the experimental population are robust to shifts in the distribution of outcomes (but not characteristics) between the experimental and target populations. We then develop new methods for learning policies that are robust to shifts in outcomes and characteristics. In doing so, we highlight how treatment effect heterogeneity within the experimental population affects the generalizability of policies. Our methods may be used with experimental or observational data (where treatment is endogenous). Many of our methods can be implemented with linear programming.
翻译:我们研究了外部有效或可泛化的个性化治疗策略的学习问题:这些策略除了在实验(或训练)群体(数据来源于此)中表现良好外,在其他目标群体中也能保持良好性能。我们首先证明,针对实验群体的福利最大化策略对实验群体与目标群体之间结果分布(而非特征分布)的变动具有稳健性。随后,我们开发了能够同时应对结果分布和特征分布变动的新策略学习方法。在此过程中,我们阐明了实验群体内的治疗效果异质性如何影响策略的泛化能力。我们的方法可应用于实验数据或观察数据(其中治疗具有内生性),且多数方法可通过线性规划实现。