Identifying disease phenotypes from electronic health records (EHRs) is critical for numerous secondary uses. Manually encoding physician knowledge into rules is particularly challenging for rare diseases due to inadequate EHR coding, necessitating review of clinical notes. Large language models (LLMs) offer promise in text understanding but may not efficiently handle real-world clinical documentation. We propose a zero-shot LLM-based method enriched by retrieval-augmented generation and MapReduce, which pre-identifies disease-related text snippets to be used in parallel as queries for the LLM to establish diagnosis. We show that this method as applied to pulmonary hypertension (PH), a rare disease characterized by elevated arterial pressures in the lungs, significantly outperforms physician logic rules ($F_1$ score of 0.62 vs. 0.75). This method has the potential to enhance rare disease cohort identification, expanding the scope of robust clinical research and care gap identification.
翻译:从电子健康记录(EHR)中识别疾病表型对于众多二次应用至关重要。将医生知识手动编码为规则对于罕见疾病而言尤为困难,这是因为电子健康记录编码不充分,需要查阅临床笔记。大语言模型在文本理解方面前景广阔,但可能无法高效处理现实世界的临床文档。我们提出一种基于大语言模型的零样本方法,该方法通过检索增强生成和MapReduce得到增强,能预先识别与疾病相关的文本片段,并作为查询并行用于大语言模型以确立诊断。我们表明,将该方法应用于肺动脉高压(一种以肺部动脉压力升高为特征的罕见疾病)时,其性能显著优于医生逻辑规则($F_1$分为0.62对比0.75)。该方法有望增强罕见疾病队列识别,从而拓展稳健临床研究的范围并助力识别护理差距。