This paper presents a comprehensive evaluation of GPT-4V's capabilities across diverse medical imaging tasks, including Radiology Report Generation, Medical Visual Question Answering (VQA), and Visual Grounding. While prior efforts have explored GPT-4V's performance in medical image analysis, to the best of our knowledge, our study represents the first quantitative evaluation on publicly available benchmarks. Our findings highlight GPT-4V's potential in generating descriptive reports for chest X-ray images, particularly when guided by well-structured prompts. Meanwhile, its performance on the MIMIC-CXR dataset benchmark reveals areas for improvement in certain evaluation metrics, such as CIDEr. In the domain of Medical VQA, GPT-4V demonstrates proficiency in distinguishing between question types but falls short of the VQA-RAD benchmark in terms of accuracy. Furthermore, our analysis finds the limitations of conventional evaluation metrics like the BLEU scores, advocating for the development of more semantically robust assessment methods. In the field of Visual Grounding, GPT-4V exhibits preliminary promise in recognizing bounding boxes, but its precision is lacking, especially in identifying specific medical organs and signs. Our evaluation underscores the significant potential of GPT-4V in the medical imaging domain, while also emphasizing the need for targeted refinements to fully unlock its capabilities.
翻译:本文系统评估了GPT-4V在多样化医学影像任务中的能力,涵盖放射报告生成、医学视觉问答(VQA)及视觉定位领域。尽管已有研究探索了GPT-4V在医学图像分析中的表现,但据我们所知,本研究首次在公开基准数据集上进行了定量评估。实验结果表明,GPT-4V在生成胸部X光图像描述性报告方面展现出潜力,尤其在结构良好的提示词引导下表现突出。然而其在MIMIC-CXR数据集基准上的性能显示,部分评估指标(如CIDEr)仍有改进空间。在医学VQA领域,GPT-4V虽能有效区分问题类型,但准确率未达VQA-RAD基准水平。此外,我们的分析揭示了BLEU评分等传统评估指标的局限性,倡导开发语义更稳健的评估方法。在视觉定位任务中,GPT-4V在识别边界框方面展现出初步潜力,但精确度仍有欠缺,尤其难以精准定位特定医学器官与征象。本研究既肯定了GPT-4V在医学影像领域的巨大潜力,也强调需针对性优化以充分释放其能力。