Large language models (LLMs) have recently demonstrated their potential in clinical applications, providing valuable medical knowledge and advice. For example, a large dialog LLM like ChatGPT has successfully passed part of the US medical licensing exam. However, LLMs currently have difficulty processing images, making it challenging to interpret information from medical images, which are rich in information that supports clinical decisions. On the other hand, computer-aided diagnosis (CAD) networks for medical images have seen significant success in the medical field by using advanced deep-learning algorithms to support clinical decision-making. This paper presents a method for integrating LLMs into medical-image CAD networks. The proposed framework uses LLMs to enhance the output of multiple CAD networks, such as diagnosis networks, lesion segmentation networks, and report generation networks, by summarizing and reorganizing the information presented in natural language text format. The goal is to merge the strengths of LLMs' medical domain knowledge and logical reasoning with the vision understanding capability of existing medical-image CAD models to create a more user-friendly and understandable system for patients compared to conventional CAD systems. In the future, LLM's medical knowledge can be also used to improve the performance of vision-based medical-image CAD models.
翻译:大型语言模型(LLMs)近期在临床应用中展现出巨大潜力,能够提供有价值的医学知识与建议。例如,以ChatGPT为代表的大规模对话型LLM已成功通过部分美国执业医师资格考试。然而,当前LLMs难以处理图像信息,这导致其无法解读富含临床决策关键信息的医学影像。另一方面,基于深度学习算法的医学图像计算机辅助诊断(CAD)网络通过先进算法支持临床决策,已在医学领域取得显著成功。本文提出一种将LLMs整合至医学图像CAD网络的方法。该框架利用LLMs对诊断网络、病灶分割网络及报告生成网络等多个CAD网络的输出进行总结与重组,以自然语言文本形式呈现增强信息。其核心目标是将LLMs的医学领域知识与逻辑推理能力,与现有医学图像CAD模型的视觉理解能力相融合,构建比传统CAD系统更易用、更易于患者理解的交互系统。未来,LLMs的医学知识还可用于提升基于视觉的医学图像CAD模型的性能。