Recent advances in machine learning models have greatly increased the performance of automated methods in medical image analysis. However, the internal functioning of such models is largely hidden, which hinders their integration in clinical practice. Explainability and trust are viewed as important aspects of modern methods, for the latter's widespread use in clinical communities. As such, validation of machine learning models represents an important aspect and yet, most methods are only validated in a limited way. In this work, we focus on providing a richer and more appropriate validation approach for highly powerful Visual Question Answering (VQA) algorithms. To better understand the performance of these methods, which answer arbitrary questions related to images, this work focuses on an automatic visual Turing test (VTT). That is, we propose an automatic adaptive questioning method, that aims to expose the reasoning behavior of a VQA algorithm. Specifically, we introduce a reinforcement learning (RL) agent that observes the history of previously asked questions, and uses it to select the next question to pose. We demonstrate our approach in the context of evaluating algorithms that automatically answer questions related to diabetic macular edema (DME) grading. The experiments show that such an agent has similar behavior to a clinician, whereby asking questions that are relevant to key clinical concepts.
翻译:近期机器学习模型的进步极大提升了医学图像分析中自动化方法的性能。然而,这类模型的内部运作机制在很大程度上仍不透明,这阻碍了其在临床实践中的整合。可解释性与可信度被视为现代方法在临床社区广泛应用的关键因素。因此,机器学习模型的验证是一个重要环节,但目前大多数方法仅以有限方式进行验证。本研究聚焦于为高性能视觉问答(VQA)算法提供更丰富、更恰当的验证方法。为了更深入理解这些能回答图像相关任意问题的方法的性能,本文聚焦于自动化视觉图灵测试(VTT)。具体而言,我们提出一种自适应问答方法,旨在揭示VQA算法的推理行为:引入一个强化学习(RL)智能体,该智能体通过观测先前提出问题的历史记录,自适应选择下一个需提出的问题。我们以糖尿病黄斑水肿(DME)分级自动问答算法的评估为背景验证该方法。实验表明,该智能体表现出与临床医生相似的提问行为,能够聚焦与关键临床概念相关的问题。