Electronic health records (EHRs) are increasingly recognized as a cost-effective resource for patient recruitment in clinical research. However, how to optimally select a cohort from millions of individuals to answer a scientific question of interest remains unclear. Consider a study to estimate the mean or mean difference of an expensive outcome. Inexpensive auxiliary covariates predictive of the outcome may often be available in patients' health records, presenting an opportunity to recruit patients selectively which may improve efficiency in downstream analyses. In this paper, we propose a two-phase sampling design that leverages available information on auxiliary covariates in EHR data. A key challenge in using EHR data for multi-phase sampling is the potential selection bias, because EHR data are not necessarily representative of the target population. Extending existing literature on two-phase sampling design, we derive an optimal two-phase sampling method that improves efficiency over random sampling while accounting for the potential selection bias in EHR data. We demonstrate the efficiency gain from our sampling design via simulation studies and an application to evaluating the prevalence of hypertension among US adults leveraging data from the Michigan Genomics Initiative, a longitudinal biorepository in Michigan Medicine.
翻译:电子健康档案(EHRs)正日益被视为临床研究中经济高效的患者招募资源。然而,如何从数百万个体中优化选择队列以回答相关的科学问题仍不明确。考虑一项估计昂贵结局均值或均值差异的研究。患者健康档案中通常可获取与结局相关的廉价辅助协变量,这为选择性招募患者提供了机会,从而可能提升下游分析效率。本文提出一种两阶段抽样设计,利用EHR数据中辅助协变量的可用信息。使用EHR数据进行多阶段抽样的关键挑战在于潜在的选择偏差——EHR数据未必能代表目标人群。通过拓展现有两阶段抽样设计文献,我们推导出一种最优两阶段抽样方法,在权衡EHR数据选择偏差的同时,较随机抽样提升了效率。通过模拟研究及一项利用密歇根基因组计划(密歇根医学纵向生物样本库)数据评估美国成年人高血压患病率的应用案例,我们验证了该抽样设计的效率增益。