Recent studies indicate that Generative Pre-trained Transformer 4 with Vision (GPT-4V) outperforms human physicians in medical challenge tasks. However, these evaluations primarily focused on the accuracy of multi-choice questions alone. Our study extends the current scope by conducting a comprehensive analysis of GPT-4V's rationales of image comprehension, recall of medical knowledge, and step-by-step multimodal reasoning when solving New England Journal of Medicine (NEJM) Image Challenges - an imaging quiz designed to test the knowledge and diagnostic capabilities of medical professionals. Evaluation results confirmed that GPT-4V outperforms human physicians regarding multi-choice accuracy (88.0% vs. 77.0%, p=0.034). GPT-4V also performs well in cases where physicians incorrectly answer, with over 80% accuracy. However, we discovered that GPT-4V frequently presents flawed rationales in cases where it makes the correct final choices (27.3%), most prominent in image comprehension (21.6%). Regardless of GPT-4V's high accuracy in multi-choice questions, our findings emphasize the necessity for further in-depth evaluations of its rationales before integrating such models into clinical workflows.
翻译:近期研究表明,具备视觉能力的生成式预训练Transformer 4(GPT-4V)在医学挑战任务中表现优于人类医生。然而,这些评估主要聚焦于多项选择题的准确性。本研究通过全面分析GPT-4V在解决《新英格兰医学杂志》(NEJM)图像挑战——一项旨在测试医学专业人员知识与诊断能力的影像学测验——时的图像理解推理、医学知识回忆以及逐步多模态推理过程,拓展了当前研究范围。评估结果证实,在多选题准确性方面(88.0% vs. 77.0%,p=0.034),GPT-4V优于人类医生。在医生回答错误的病例中,GPT-4V同样表现良好,准确率超过80%。但我们发现,在做出正确最终选择的病例中(27.3%),GPT-4V频繁呈现有缺陷的推理,其中在图像理解方面最为突出(21.6%)。尽管GPT-4V在多选题中具有高准确性,我们的研究结果强调,在将该类模型整合到临床工作流程之前,有必要进一步深入评估其推理过程。