JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability

With the explosive growth of medical data and the rapid development of artificial intelligence technology, precision medicine has emerged as a key to enhancing the quality and efficiency of healthcare services. In this context, Large Language Models (LLMs) play an increasingly vital role in medical knowledge acquisition and question-answering systems. To further improve the performance of these systems in the medical domain, we introduce an innovative method that jointly trains an Information Retrieval (IR) system and an LLM during the fine-tuning phase. This approach, which we call Joint Medical LLM and Retrieval Training (JMLR), is designed to overcome the challenges faced by traditional models in handling medical question-answering tasks. By employing a synchronized training mechanism, JMLR reduces the demand for computational resources and enhances the model's ability to leverage medical knowledge for reasoning and answering questions. Our experimental results demonstrate that JMLR-13B (81.2% on Amboos, 61.3% on MedQA) outperforms models using conventional pre-training and fine-tuning Meditron-70B (76.4% on AMBOSS, 60.3% on MedQA). For models of the same 7B scale, JMLR-7B(68.7% on Amboos, 51.7% on MedQA) significantly outperforms other public models (Meditron-7B: 50.1%, 47.9%), proving its superiority in terms of cost (our training time: 37 hours, traditional method: 144 hours), efficiency, and effectiveness in medical question-answering tasks. Through this work, we provide a new and efficient knowledge enhancement tool for healthcare, demonstrating the great potential of integrating IR and LLM training in precision medical information retrieval and question-answering systems.

翻译：随着医学数据的爆炸性增长和人工智能技术的快速发展，精准医学已成为提升医疗服务质量和效率的关键。在此背景下，大语言模型在医学知识获取和问答系统中发挥着日益重要的作用。为进一步提升这些系统在医学领域的性能，我们提出了一种创新方法，在微调阶段联合训练信息检索系统与大语言模型。这种被称为联合医学大语言模型与检索训练的方法，旨在解决传统模型处理医学问答任务时面临的挑战。通过采用同步训练机制，JMLR降低了计算资源需求，并增强了模型运用医学知识进行推理和回答问题的能力。实验结果表明，JMLR-13B（在Amboos上达81.2%，MedQA上达61.3%）优于采用传统预训练和微调的Meditron-70B（AMBOSS上76.4%，MedQA上60.3%）。在同等7B规模模型中，JMLR-7B（Amboos上68.7%，MedQA上51.7%）显著超越其他公开模型（Meditron-7B：50.1%，47.9%），证明了其在医学问答任务中的成本优势（训练时间：传统方法144小时，本方法37小时）、效率及有效性。通过本研究，我们为医疗领域提供了一种新型高效知识增强工具，展示了将信息检索与大语言模型训练相结合在精准医学信息检索与问答系统中的巨大潜力。