Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image. That inference is unsafe: a model exploiting finding-name priors scores like one that reads the scan, and no standard benchmark separates them. We introduce a causal audit that intervenes on the image, occluding the relevant region, occluding an irrelevant one, and swapping in another patient's same-label scan, and combines three behavioral metrics to test whether a correct answer depends on the image. Across nine systems, a text-only model with no image access reaches within 5.7 accuracy points of the best multimodal one, and a 119-billion-parameter multimodal model is statistically indistinguishable from a 7-billion text-only baseline. The audit splits the cohort into three models that ignore the image, one that is unstable, and five that use it selectively, for a subset of findings; the categories hold across a second dataset, resolution, and prompt phrasing. Against board-certified radiologists, a text-only model is statistically indistinguishable from a radiologist's accuracy while grounding at zero, whereas the image-using models ground at radiologist-comparable rates. Reported confidence flags ungrounded answers only when a model uses the image. Grounding audits, not accuracy, should gate clinical deployment.
翻译:医学视觉语言模型在胸部X光片上报告了高准确率,这日益被解读为模型利用了图像信息的证据。然而这种推断并不安全:利用疾病名称先验知识的模型所获分数与读取扫描图像的模型相当,而现有标准基准无法区分二者。我们提出一种因果审计方法:通过遮挡图像相关区域、遮挡无关区域以及替换为另一患者同标签扫描图像来干预图像,并结合三项行为指标检测正确回答是否依赖图像。对九种系统进行测试表明:无图像访问权限的纯文本模型与最佳多模态模型准确率差距在5.7个百分点以内,而1190亿参数的多模态模型与70亿参数的纯文本基线模型在统计上无显著差异。该审计方法将系统群组分为三类:忽略图像的模型(3个)、不稳定的模型(1个)、选择性利用有限疾病图像信息的模型(5个);该分类在第二个数据集、不同分辨率和提示措辞下保持稳定。与委员会认证放射科医生相比:纯文本模型准确率在统计上与放射科医生无差异但基础归因率为零,而使用图像的模型达到了与放射科医生可比的基础归因率。报告置信度仅当模型使用图像时才可标记无依据的回答。临床部署应以归因审计而非准确率为准入门槛。