We study how to learn treatment policies from multimodal electronic health records (EHRs) that consist of tabular data and clinical text. These policies can help physicians make better treatment decisions and allocate healthcare resources more efficiently. Causal policy learning methods prioritize patients with the largest expected treatment benefit. Yet, existing estimators are designed for tabular covariates under causal assumptions that may be hard to justify in the multimodal setting. A pragmatic alternative is to apply causal estimators directly to multimodal representations, but this can produce biased treatment effect estimates when the representations do not preserve the relevant confounding information. As a result, predictive models of baseline risk are commonly used in practice to guide treatment decisions, although they are not designed to identify which patients benefit most from treatment. We propose AACE (Annotation-Assisted Coarsened Effects), an annotation-assisted approach to causal policy learning for multimodal EHRs. The method uses expert-provided annotations during training to support confounding adjustment, and then predicts treatment benefit from only multimodal representations at inference. We show that the proposed method achieves strong empirical performance across synthetic, semi-synthetic, and real-world EHR datasets, outperforming risk-based and representation-based causal baselines, and offering practical insights for applying causal machine learning in clinical practice.
翻译:我们研究如何从包含表格数据和临床文本的多模态电子健康记录中学习治疗策略。这些策略可帮助医生优化治疗决策并更高效地分配医疗资源。因果策略学习方法优先考虑具有最大预期治疗获益的患者,然而现有估计器基于因果假设针对表格协变量设计,在多模态环境下这些假设难以验证。一种实用替代方案是直接对多模态表征应用因果估计器,但当表征未能保留相关混杂信息时,这种做法会产生有偏的治疗效果估计。因此,临床实践中常用基线风险预测模型指导治疗决策,尽管这些模型并非为识别最能从治疗中获益的患者而设计。我们提出AACE(注释辅助粗化效应)方法,这是一种面向多模态电子健康记录的因果策略学习注释辅助方法。该方法在训练阶段利用专家提供的注释进行混杂调整,推理阶段仅通过多模态表征预测治疗获益。实验结果表明,所提方法在合成、半合成及真实电子健康记录数据集上均取得优异实证表现,超越基于风险与基于表征的因果基线方法,为因果机器学习在临床实践中的应用提供了实用见解。