In the dynamic landscape of medical artificial intelligence, this study explores the vulnerabilities of the Pathology Language-Image Pretraining (PLIP) model, a Vision Language Foundation model, under targeted adversarial conditions. Leveraging the Kather Colon dataset with 7,180 H&E images across nine tissue types, our investigation employs Projected Gradient Descent (PGD) adversarial attacks to intentionally induce misclassifications. The outcomes reveal a 100% success rate in manipulating PLIP's predictions, underscoring its susceptibility to adversarial perturbations. The qualitative analysis of adversarial examples delves into the interpretability challenges, shedding light on nuanced changes in predictions induced by adversarial manipulations. These findings contribute crucial insights into the interpretability, domain adaptation, and trustworthiness of Vision Language Models in medical imaging. The study emphasizes the pressing need for robust defenses to ensure the reliability of AI models.
翻译:在医学人工智能的动态格局中,本研究探讨了病理语言-图像预训练(PLIP)模型——一种视觉语言基础模型——在定向对抗条件下的脆弱性。利用包含9种组织类型、共计7180张H&E染色图像的Kather结肠数据集,我们采用投影梯度下降(PGD)对抗攻击方法刻意诱导分类错误。结果显示,PLIP模型预测的操纵成功率高达100%,凸显其对抗扰动的易感性。对抗样本的定性分析深入探讨了可解释性挑战,揭示了由对抗性操纵引起的预测中细微变化。这些发现为医学影像中视觉语言模型的可解释性、领域适应性和可信度提供了关键见解。本研究强调了建立稳健防御机制以确保人工智能模型可靠性的迫切需求。