Monitoring the health status of patients in the ICU is crucial for providing them with better care and treatment. Massive raw electronic health records (EHR) give machine learning models more clinical texts and vital signs to make accurate predictions. Currently, many advanced NLP models have emerged for clinical note analysis. However, due to the complicated textual structure and noise in raw clinical data, coarse embedding approaches without domain-specific refining limit the accuracy improvement. To address this issue, we propose FINEEHR, a system adopting two representation learning techniques, including metric learning and fine-tuning, to refine clinical note embeddings, utilizing the inner correlation among different health statuses and note categories. We evaluate the performance of FINEEHR using two metrics, AUC and AUC-PR, on a real-world MIMIC III dataset. Our experimental results demonstrate that both refining approaches can improve prediction accuracy, and their combination presents the best results. It outperforms previous works, achieving an AUC improvement of over 10%, with an average AUC of 96.04% and an average AUC-PR of 96.48% across various classifiers.
翻译:监测ICU患者的健康状况对于提供更优质的护理与治疗至关重要。海量的原始电子健康记录(EHR)为机器学习模型提供了丰富的临床文本和生命体征数据,从而实现精准预测。当前,许多先进的NLP模型已涌现用于临床笔记分析。然而,由于原始临床数据中复杂的文本结构和噪声,缺乏领域特定优化的粗粒度嵌入方法限制了准确率的提升。针对这一问题,我们提出FINEEHR系统,采用度量学习和微调两种表示学习技术,利用不同健康状况与笔记类别间的内在关联,优化临床笔记嵌入。我们使用AUC和AUC-PR两个指标,在真实世界的MIMIC III数据集上评估FINEEHR性能。实验结果表明,两种优化方法均能提升预测准确率,且二者结合效果最佳。该方法超越以往研究成果,AUC提升超过10%,在不同分类器上平均AUC达96.04%,平均AUC-PR达96.48%。