Medical Visual Question Answering~(VQA) is a combination of medical artificial intelligence and popular VQA challenges. Given a medical image and a clinically relevant question in natural language, the medical VQA system is expected to predict a plausible and convincing answer. Although the general-domain VQA has been extensively studied, the medical VQA still needs specific investigation and exploration due to its task features. In the first part of this survey, we collect and discuss the publicly available medical VQA datasets up-to-date about the data source, data quantity, and task feature. In the second part, we review the approaches used in medical VQA tasks. We summarize and discuss their techniques, innovations, and potential improvements. In the last part, we analyze some medical-specific challenges for the field and discuss future research directions. Our goal is to provide comprehensive and helpful information for researchers interested in the medical visual question answering field and encourage them to conduct further research in this field.
翻译:医学视觉问答(VQA)是医学人工智能与广泛VQA挑战的结合。给定一张医学图像和一个用自然语言表述的临床相关问题,医学VQA系统需预测出合理且有说服力的答案。尽管通用领域的VQA已被广泛研究,但医学VQA因其任务特性仍需专门的探究与探索。本综述第一部分收集并讨论了截至当前公开的医学VQA数据集,涵盖数据来源、数据量及任务特征。第二部分回顾了医学VQA任务中采用的方法,总结并讨论了其技术、创新点及潜在改进方向。最后一部分分析了该领域特有的医学挑战,并探讨了未来研究方向。我们的目标是为对医学视觉问答领域感兴趣的研究者提供全面且有参考价值的信息,并鼓励他们在该领域开展进一步研究。