As advances in large language models (LLMs) and multimodal techniques continue to mature, the development of general-purpose multimodal large language models (MLLMs) has surged, with significant applications in natural image interpretation. However, the field of pathology has largely remained untapped in this regard, despite the growing need for accurate, timely, and personalized diagnostics. To bridge the gap in pathology MLLMs, we present the PathAsst in this study, which is a generative foundation AI assistant to revolutionize diagnostic and predictive analytics in pathology. To develop PathAsst, we collect over 142K high-quality pathology image-text pairs from a variety of reliable sources, including PubMed, comprehensive pathology textbooks, reputable pathology websites, and private data annotated by pathologists. Leveraging the advanced capabilities of ChatGPT/GPT-4, we generate over 180K instruction-following samples. Furthermore, we devise additional instruction-following data, specifically tailored for the invocation of the pathology-specific models, allowing the PathAsst to effectively interact with these models based on the input image and user intent, consequently enhancing the model's diagnostic capabilities. Subsequently, our PathAsst is trained based on Vicuna-13B language model in coordination with the CLIP vision encoder. The results of PathAsst show the potential of harnessing the AI-powered generative foundation model to improve pathology diagnosis and treatment processes. We are committed to open-sourcing our meticulously curated dataset, as well as a comprehensive toolkit designed to aid researchers in the extensive collection and preprocessing of their own datasets. Resources can be obtained at https://github.com/superjamessyx/Generative-Foundation-AI-Assistant-for-Pathology.
翻译:随着大型语言模型(LLMs)及多模态技术的持续成熟,通用型多模态大语言模型(MLLMs)的开发蓬勃发展,并在自然图像理解领域展现出重要应用。然而,尽管对精准、及时且个性化诊断的需求日益增长,病理学领域在此方面仍基本处于未开发状态。为填补病理学MLLMs的空白,本研究提出PathAsst——一种生成式基础AI助手,旨在革新病理学中的诊断与预测分析。为开发PathAsst,我们从多种可靠来源收集了超过14.2万对高质量病理学图像-文本对,包括PubMed、综合性病理学教科书、权威病理学网站以及由病理学家标注的私有数据。利用ChatGPT/GPT-4的先进能力,我们生成了超过18万个指令跟随样本。此外,我们还设计了专门用于调用病理学特定模型的额外指令跟随数据,使PathAsst能够基于输入图像和用户意图高效与这些模型交互,从而增强模型的诊断能力。随后,PathAsst基于Vicuna-13B语言模型并与CLIP视觉编码器协同训练。PathAsst的结果表明,利用基于AI的生成式基础模型来改进病理学诊断与治疗流程具有巨大潜力。我们承诺开源精心整理的数据集以及一套综合性工具包,旨在帮助研究者广泛收集与预处理其自有数据集。资源可从https://github.com/superjamessyx/Generative-Foundation-AI-Assistant-for-Pathology获取。