Rapid and accurate identification of Venous thromboembolism (VTE), a severe cardiovascular condition including deep vein thrombosis (DVT) and pulmonary embolism (PE), is important for effective treatment. Leveraging Natural Language Processing (NLP) on radiology reports, automated methods have shown promising advancements in identifying VTE events from retrospective data cohorts or aiding clinical experts in identifying VTE events from radiology reports. However, effectively training Deep Learning (DL) and the NLP models is challenging due to limited labeled medical text data, the complexity and heterogeneity of radiology reports, and data imbalance. This study proposes novel method combinations of DL methods, along with data augmentation, adaptive pre-trained NLP model selection, and a clinical expert NLP rule-based classifier, to improve the accuracy of VTE identification in unstructured (free-text) radiology reports. Our experimental results demonstrate the model's efficacy, achieving an impressive 97\% accuracy and 97\% F1 score in predicting DVT, and an outstanding 98.3\% accuracy and 98.4\% F1 score in predicting PE. These findings emphasize the model's robustness and its potential to significantly contribute to VTE research.
翻译:静脉血栓栓塞(VTE)是一种包括深静脉血栓(DVT)和肺栓塞(PE)在内的严重心血管疾病,快速准确识别其对有效治疗至关重要。利用自然语言处理(NLP)技术分析放射学报告,自动化方法在从回顾性数据队列中识别VTE事件或辅助临床专家从放射学报告中识别VTE事件方面已展现出显著进展。然而,由于标注医学文本数据有限、放射学报告的复杂性与异质性以及数据不平衡,高效训练深度学习(DL)和NLP模型面临挑战。本研究提出结合DL方法、数据增强、自适应预训练NLP模型选择及临床专家NLP规则分类器的新型方法组合,以提高对非结构化(自由文本)放射学报告中VTE的识别准确率。实验结果表明,该模型在DVT预测中实现了97%的准确率和97%的F1分数,在PE预测中达到了98.3%的准确率和98.4%的F1分数,展现了优异的性能。这些发现强调了该模型的稳健性及其对VTE研究的潜在重要贡献。