We consider the problem of learning personalized treatment policies that are externally valid or generalizable: they perform well in other target populations besides the experimental (or training) population from which data are sampled. We first show that welfare-maximizing policies for the experimental population are robust to shifts in the distribution of outcomes (but not characteristics) between the experimental and target populations. We then develop new methods for learning policies that are robust to shifts in outcomes and characteristics. In doing so, we highlight how treatment effect heterogeneity within the experimental population affects the generalizability of policies. Our methods may be used with experimental or observational data (where treatment is endogenous). Many of our methods can be implemented with linear programming.
翻译:我们考虑学习具有外部有效性或可泛化性的个性化治疗策略问题:这些策略在除实验(或训练)群体(数据采样来源)之外的其他目标群体中同样表现良好。我们首先证明,针对实验群体的福利最大化策略在实验群体与目标群体之间的结果分布(而非特征分布)发生偏移时具有稳健性。随后,我们开发了能够同时应对结果和特征分布偏移的新方法。在此过程中,我们揭示了实验群体内治疗效果异质性如何影响策略的可泛化性。我们的方法可应用于实验数据或观察性数据(其中治疗具有内生性)。其中多数方法可通过线性规划实现。