Background. Atrial fibrillation (AF) is the most prevalent cardiac arrhythmia and a major determinant of prognosis. Established AF risk scores rely on factors (older age, hypertension) nearly ubiquitous among patients with cardiovascular disease (CVD), offering limited stratification in this high-risk group. Most target long-term (5-10 year) rather than medium-term prediction. We developed interpretable ML models predicting AF risk over a 24-month and entire follow-up horizon in CVD patients using routinely collected hospital data. Methods. Single-center retrospective study of electronic health records from the National Research Cardiology Center (Russia) for patients aged >=18 with CVD but without pre-existing AF, hospitalized more than once between January 2012 and May 2019. A custom NLP pipeline transformed unstructured discharge reports into 73 structured features, combining a rule-based parser with transformer-based NER. Using LightAutoML we built a full model (73 features), a simple model (reduced subset), and a linear model for a bedside risk score. Performance was assessed by ROC AUC, compared with CHARGE-AF, C2HEST, MHS, and HAVOC, and interpreted via SHAP. Results. Of 80,576 records from 45,000 patients, 17,562 met inclusion criteria; 1,438 (8.19%) developed AF. The full model reached ROC AUC 0.735 (24-month) and 0.696 (entire follow-up); the simple model was nearly identical (0.725, 0.696). All non-linear models outperformed the four clinical risk scores (ROC AUC 0.53-0.64). The simple model uses 13 features and is named Pre-AF 13. SHAP identified age and left atrial volume as dominant predictors. A linear risk score (Pre-AF 9) stratified observed 24-month AF incidence from ~7% to 36%. Conclusion. Interpretable ML models built from routinely collected EHR data identify high-AF-risk CVD patients, outperforming established clinical risk scores.
翻译:背景:房颤(AF)是最常见的心律失常,也是预后的主要决定因素。现有的房颤风险评分依赖于心血管疾病(CVD)患者中几乎普遍存在的因素(如高龄、高血压),因此在该高风险群体中提供的分层能力有限。大多数评分针对长期(5-10年)而非中期预测。我们利用常规收集的医院数据,开发了可解释的机器学习模型,用于预测心血管疾病患者在24个月及整个随访期间内的房颤风险。方法:这是一项单中心回顾性研究,数据来自俄罗斯国家心脏病学研究中心2012年1月至2019年5月期间多次住院、年龄≥18岁、患有心血管疾病但无房颤病史患者的电子健康记录。我们通过定制的自然语言处理流程,将非结构化的出院报告转化为73个结构化特征,该流程结合了基于规则的解析器和基于Transformer的命名实体识别。利用LightAutoML,我们构建了包含所有特征的完整模型、精简特征集的简单模型以及用于床旁风险评分的线性模型。模型性能通过ROC AUC评估,并与CHARGE-AF、C2HEST、MHS和HAVOC评分进行比较,同时利用SHAP进行解释。结果:在45,000名患者共80,576份记录中,有17,562份符合纳入标准;其中1,438名(8.19%)发生了房颤。完整模型在24个月和整个随访期间的ROC AUC分别达到0.735和0.696;简单模型的表现几乎相同(分别为0.725和0.696)。所有非线性模型的性能均优于四种临床风险评分(ROC AUC为0.53-0.64)。简单模型使用13个特征,命名为Pre-AF 13。SHAP分析显示年龄和左心房容积是主要预测因子。线性风险评分(Pre-AF 9)可将观察到的24个月房颤发生率从约7%分层至36%。结论:基于常规收集的电子健康记录数据构建的可解释机器学习模型能够识别高风险房颤的心血管疾病患者,其性能优于现有的临床风险评分。