Early prediction of Alzheimer's disease (AD) is crucial for timely intervention and treatment. This study aims to use machine learning approaches to analyze longitudinal electronic health records (EHRs) of patients with AD and identify signs and symptoms that can predict AD onset earlier. We used a case-control design with longitudinal EHRs from the U.S. Department of Veterans Affairs Veterans Health Administration (VHA) from 2004 to 2021. Cases were VHA patients with AD diagnosed after 1/1/2016 based on ICD-10-CM codes, matched 1:9 with controls by age, sex and clinical utilization with replacement. We used a panel of AD-related keywords and their occurrences over time in a patient's longitudinal EHRs as predictors for AD prediction with four machine learning models. We performed subgroup analyses by age, sex, and race/ethnicity, and validated the model in a hold-out and "unseen" VHA stations group. Model discrimination, calibration, and other relevant metrics were reported for predictions up to ten years before ICD-based diagnosis. The study population included 16,701 cases and 39,097 matched controls. The average number of AD-related keywords (e.g., "concentration", "speaking") per year increased rapidly for cases as diagnosis approached, from around 10 to over 40, while remaining flat at 10 for controls. The best model achieved high discriminative accuracy (ROCAUC 0.997) for predictions using data from at least ten years before ICD-based diagnoses. The model was well-calibrated (Hosmer-Lemeshow goodness-of-fit p-value = 0.99) and consistent across subgroups of age, sex and race/ethnicity, except for patients younger than 65 (ROCAUC 0.746). Machine learning models using AD-related keywords identified from EHR notes can predict future AD diagnoses, suggesting its potential use for identifying AD risk using EHR notes, offering an affordable way for early screening on large population.
翻译:阿尔茨海默病(AD)的早期预测对于及时干预和治疗至关重要。本研究旨在利用机器学习方法分析AD患者的纵向电子健康记录(EHRs),并识别可提前预测AD发病的体征和症状。我们采用病例对照设计,使用了美国退伍军人事务部退伍军人健康管理局(VHA)2004年至2021年的纵向EHRs数据。病例组为2016年1月1日后根据ICD-10-CM编码确诊为AD的VHA患者,按年龄、性别和临床使用情况以1:9的比例与对照组进行匹配(可重复匹配)。我们以患者纵向EHRs中一组AD相关关键词及其随时间出现的频率作为预测因子,采用四种机器学习模型进行AD预测。我们按年龄、性别和种族/族裔进行了亚组分析,并在留出集和“未见”VHA站点的数据上验证了模型。报告了在基于ICD诊断前长达十年的预测中,模型的区分度、校准度及其他相关指标。研究队列包括16,701例病例和39,097例匹配对照。病例组每年AD相关关键词(如“注意力不集中”、“言语不清”)的平均数量随诊断临近而快速增加,从约10个增至40个以上,而对照组则稳定在10个左右。最佳模型在基于ICD诊断前至少十年的数据预测中实现了高区分度(ROCAUC 0.997),模型校准良好(Hosmer-Lemeshow拟合优度p值=0.99),且在除年龄小于65岁患者(ROCAUC 0.746)外的各年龄、性别及种族/族裔亚组中表现一致。利用从EHR记录中识别的AD相关关键词的机器学习模型可预测未来AD诊断,表明其具有通过EHR记录识别AD风险的潜力,为大规模人群的早期筛查提供了经济可行的方法。