This paper introduces MedExQA, a novel benchmark in medical question-answering, to evaluate large language models' (LLMs) understanding of medical knowledge through explanations. By constructing datasets across five distinct medical specialties that are underrepresented in current datasets and further incorporating multiple explanations for each question-answer pair, we address a major gap in current medical QA benchmarks which is the absence of comprehensive assessments of LLMs' ability to generate nuanced medical explanations. Our work highlights the importance of explainability in medical LLMs, proposes an effective methodology for evaluating models beyond classification accuracy, and sheds light on one specific domain, speech language pathology, where current LLMs including GPT4 lack good understanding. Our results show generation evaluation with multiple explanations aligns better with human assessment, highlighting an opportunity for a more robust automated comprehension assessment for LLMs. To diversify open-source medical LLMs (currently mostly based on Llama2), this work also proposes a new medical model, MedPhi-2, based on Phi-2 (2.7B). The model outperformed medical LLMs based on Llama2-70B in generating explanations, showing its effectiveness in the resource-constrained medical domain. We will share our benchmark datasets and the trained model.
翻译:本文提出MedExQA,一种新颖的医学问答基准数据集,旨在通过解释性评估大语言模型对医学知识的理解能力。通过构建五个当前数据集覆盖不足的医学专科数据集,并为每个问答对引入多重解释,我们填补了当前医学问答基准中缺乏对大语言模型生成细微医学解释能力的全面评估这一重要空白。本研究强调了可解释性在医学大语言模型中的重要性,提出了一种超越分类准确率的模型评估方法论,并揭示了语言病理学这一特定领域——包括GPT4在内的当前大语言模型对此缺乏深入理解。实验结果表明,基于多重解释的生成式评估与人工评估结果具有更高一致性,这为大语言模型提供了更鲁棒的自动化理解评估契机。为丰富开源医学大语言模型(当前主要基于Llama2),本研究还提出基于Phi-2(2.7B参数)的新型医学模型MedPhi-2。该模型在生成解释任务上优于基于Llama2-70B的医学大语言模型,展示了其在资源受限医学领域的有效性。我们将公开共享基准数据集与训练模型。