Methods for learning optimal policies use causal machine learning models to create human-interpretable rules for making choices around the allocation of different policy interventions. However, in realistic policy-making contexts, decision-makers often care about trade-offs between outcomes, not just single-mindedly maximising utility for one outcome. This paper proposes an approach termed Multi-Objective Policy Learning (MOPoL) which combines optimal decision trees for policy learning with a multi-objective Bayesian optimisation approach to explore the trade-off between multiple outcomes. It does this by building a Pareto frontier of non-dominated models for different hyperparameter settings which govern outcome weighting. The key here is that a low-cost greedy tree can be an accurate proxy for the very computationally costly optimal tree for the purposes of making decisions which means models can be repeatedly fit to learn a Pareto frontier. The method is applied to a real-world case-study of non-price rationing of anti-malarial medication in Kenya.
翻译:学习最优策略的方法利用因果机器学习模型生成可解释规则,用于分配不同政策干预措施。然而,在现实政策制定场景中,决策者通常关注目标之间的权衡,而非单一地追求某一目标的效用最大化。本文提出一种名为“多目标策略学习”(MOPoL)的方法,该方法将策略学习中的最优决策树与多目标贝叶斯优化相结合,以探索多个目标之间的权衡关系。具体而言,通过构建不同超参数设置(用于控制目标权重)下的非支配模型帕累托前沿来实现这一目标。关键在于,低成本的贪婪树可以作为计算成本极高的最优树在决策过程中的精确代理,从而能够通过重复拟合模型来学习帕累托前沿。该方法应用于肯尼亚抗疟药物非价格配给的实际案例研究。