Machine reasoning has made great progress in recent years owing to large language models (LLMs). In the clinical domain, however, most NLP-driven projects mainly focus on clinical classification or reading comprehension, and under-explore clinical reasoning for disease diagnosis due to the expensive rationale annotation with clinicians. In this work, we present a ``reasoning-aware'' diagnosis framework that rationalizes the diagnostic process via prompt-based learning in a time- and labor-efficient manner, and learns to reason over the prompt-generated rationales. Specifically, we address the clinical reasoning for disease diagnosis, where the LLM generates diagnostic rationales providing its insight on presented patient data and the reasoning path towards the diagnosis, namely Clinical Chain-of-Thought (Clinical CoT). We empirically demonstrate LLMs/LMs' ability of clinical reasoning via extensive experiments and analyses on both rationale generation and disease diagnosis in various settings. We further propose a novel set of criteria for evaluating machine-generated rationales' potential for real-world clinical settings, facilitating and benefiting future research in this area.
翻译:近年来,得益于大型语言模型(LLMs),机器推理取得了显著进展。然而在临床领域,大多数自然语言处理驱动的项目主要聚焦于临床分类或阅读理解,由于需要临床医生进行昂贵的人工标注,针对疾病诊断的临床推理研究仍不充分。本文提出了一种“推理感知”诊断框架,通过基于提示学习以省时省力的方式合理化诊断过程,并学会对提示生成的理由进行推理。具体而言,我们聚焦疾病诊断中的临床推理问题,其中LLM生成诊断理由,阐述其对患者数据的见解以及通往诊断的推理路径,即临床思维链(Clinical CoT)。我们通过在不同设置下对理由生成和疾病诊断进行广泛的实验与分析,实证证明了LLMs/LMs的临床推理能力。此外,我们提出了一套新颖的评估标准,用于评价机器生成理由在真实临床环境中的潜力,从而促进并有益于该领域的未来研究。