In the dynamic landscape of medical artificial intelligence, this study explores the vulnerabilities of the Pathology Language-Image Pretraining (PLIP) model, a Vision Language Foundation model, under targeted adversarial conditions. Leveraging the Kather Colon dataset with 7,180 H&E images across nine tissue types, our investigation employs Projected Gradient Descent (PGD) adversarial attacks to intentionally induce misclassifications. The outcomes reveal a 100% success rate in manipulating PLIP's predictions, underscoring its susceptibility to adversarial perturbations. The qualitative analysis of adversarial examples delves into the interpretability challenges, shedding light on nuanced changes in predictions induced by adversarial manipulations. These findings contribute crucial insights into the interpretability, domain adaptation, and trustworthiness of Vision Language Models in medical imaging. The study emphasizes the pressing need for robust defenses to ensure the reliability of AI models.
翻译:在医学人工智能的动态格局中,本研究探讨了病理语言-图像预训练(PLIP)模型——一种视觉语言基础模型——在针对性对抗条件下的脆弱性。利用包含来自九种组织类型的7180张H&E图像的Kather结肠数据集,我们采用投影梯度下降(PGD)对抗攻击故意诱导误分类。结果显示,PLIP预测的成功操控率达100%,凸显其对对抗扰动的敏感性。对抗性样本的定性分析深入探究了可解释性挑战,揭示了由对抗操控引起的预测细微变化。这些发现为医学成像中视觉语言模型的可解释性、领域适应性和可信度提供了关键见解。研究强调了迫切需要鲁棒防御机制,以确保人工智能模型的可靠性。