Across domains such as medicine, employment, and criminal justice, predictive models often target labels that imperfectly reflect the outcomes of interest to experts and policymakers. For example, clinical risk assessments deployed to inform physician decision-making often predict measures of healthcare utilization (e.g., costs, hospitalization) as a proxy for patient medical need. These proxies can be subject to outcome measurement error when they systematically differ from the target outcome they are intended to measure. However, prior modeling efforts to characterize and mitigate outcome measurement error overlook the fact that the decision being informed by a model often serves as a risk-mitigating intervention that impacts the target outcome of interest and its recorded proxy. Thus, in these settings, addressing measurement error requires counterfactual modeling of treatment effects on outcomes. In this work, we study intersectional threats to model reliability introduced by outcome measurement error, treatment effects, and selection bias from historical decision-making policies. We develop an unbiased risk minimization method which, given knowledge of proxy measurement error properties, corrects for the combined effects of these challenges. We also develop a method for estimating treatment-dependent measurement error parameters when these are unknown in advance. We demonstrate the utility of our approach theoretically and via experiments on real-world data from randomized controlled trials conducted in healthcare and employment domains. As importantly, we demonstrate that models correcting for outcome measurement error or treatment effects alone suffer from considerable reliability limitations. Our work underscores the importance of considering intersectional threats to model validity during the design and evaluation of predictive models for decision support.
翻译:在医疗、就业和刑事司法等领域,预测模型往往针对无法完美反映专家与政策制定者所关注目标结果的标签进行建模。例如,用于辅助医生决策的临床风险评估模型,常以医疗资源使用指标(如费用、住院率)作为患者医疗需求的代理变量。当这些代理变量与预期测量的目标结果存在系统性差异时,便可能产生结果测量误差。然而,现有针对结果测量误差的特征描述与缓解方法忽略了关键事实:作为决策依据的模型本身常构成一种风险缓解干预,既影响目标结果,也影响其记录代理变量。因此,在此类场景中,纠正测量误差需要建立关于处理对结果影响的反事实模型。本文系统研究了由结果测量误差、处理效应以及历史决策政策引发的选择性偏差对模型可靠性造成的交叉威胁。我们提出了一种无偏风险最小化方法,在已知代理变量测量误差特性的前提下,能够同步纠正这些挑战的复合影响。同时,针对处理相关测量误差参数未知的场景,我们还开发了对应的参数估计方法。我们通过理论推导与医疗及就业领域的随机对照试验真实数据实验,验证了该方法的有效性。更重要的是,研究表明单纯纠正结果测量误差或处理效应的模型存在严重可靠性局限。本研究强调了在设计评估决策支持预测模型时,必须考虑模型有效性的多重交叉威胁。