Electronic health records (EHR) data have considerable variability in data completeness across sites and patients. Lack of "EHR data-continuity" or "EHR data-discontinuity", defined as "having medical information recorded outside the reach of an EHR system" can lead to a substantial amount of information bias. The objective of this study was to comprehensively evaluate (1) how EHR data-discontinuity introduces data bias, (2) case finding algorithms affect downstream prediction models, and (3) how algorithmic fairness is associated with racial-ethnic disparities. We leveraged our EHRs linked with Medicaid and Medicare claims data in the OneFlorida+ network and used a validated measure (i.e., Mean Proportions of Encounters Captured [MPEC]) to estimate patients' EHR data continuity. We developed a machine learning model for predicting type 2 diabetes (T2D) diagnosis as the use case for this work. We found that using cohorts selected by different levels of EHR data-continuity affects utilities in disease prediction tasks. The prediction models trained on high continuity data will have a worse fit on low continuity data. We also found variations in racial and ethnic disparities in model performances and model fairness in models developed using different degrees of data continuity. Our results suggest that careful evaluation of data continuity is critical to improving the validity of real-world evidence generated by EHR data and health equity.
翻译:电子健康档案(EHR)数据在不同医疗中心和患者间存在显著的数据完整性差异。"EHR数据连续性缺失"或"EHR数据不连续性"(定义为"超出EHR系统覆盖范围的医疗信息记录")可能导致大量信息偏倚。本研究旨在全面评估:(1)EHR数据不连续性如何引入数据偏倚;(2)病例发现算法对下游预测模型的影响;以及(3)算法公平性与种族-民族差异的关联。我们利用OneFlorida+网络中与医疗补助及医疗保险索赔数据关联的EHR,采用已验证的测量指标(即捕获就诊平均比例[MPEC])估算患者的EHR数据连续性。以2型糖尿病(T2D)诊断预测为应用场景,构建机器学习模型。研究发现,基于不同EHR数据连续性水平选择的队列会影响疾病预测任务的效用。基于高连续性数据训练的预测模型,在低连续性数据上的拟合效果较差。此外,不同数据连续性程度开发的模型在模型性能及公平性方面存在显著的种族与民族差异。研究结果表明,审慎评估数据连续性对于提升基于EHR数据生成的真实世界证据的有效性与健康公平性至关重要。