A Benchmark for Hallucination Detection in VLMs for Gastrointestinal Endoscopy - 专知论文

会员服务 ·

0

AUC · MoDELS · 黑盒 · 白盒 · 视觉问答 ·

A Benchmark for Hallucination Detection in VLMs for Gastrointestinal Endoscopy

翻译：暂无翻译

Aminu Lawal,Niyoj Oli,Sachin Acharya,Prashnna Gyawali,Maria Carmen Romano,Binod Bhattarai

from arxiv, Accepted at the Medical Image Understanding and Analysis (MIUA) 2026 conference

Vision-language models (VLMs) are prone to hallucination, which remains a major barrier to their safe deployment in clinical practice. To date, most hallucination detection methods have been evaluated on radiology benchmarks such as MIMIC-CXR and VQA-RAD, while gastrointestinal (GI) endoscopy remains largely underexplored. In this paper, we benchmark nine hallucination detection methods on the Gut-VLM dataset, a GI diagnostic Visual Question Answering (VQA) dataset with 4,392 test VQA pairs, across five VLMs (MedGemma-4B, MedGemma-27B, LLaVA-Med-7B, LLaVA-v1.6-7B, and Lingshu-32B). The methods span three categories: black-box methods (RadFlag, SelfCheckGPT-NLI), gray-box methods (AvgProb, AvgEnt, MaxProb, MaxEnt, Semantic Entropy, and VASE), and a white-box method (ReXTrust). Our results show that ReXTrust, a white-box method, achieves the highest AUC across all five models, outperforming the strongest alternative method on each VLM by a statistically significant margin (paired permutation test, p < 0.001 in all cases), reaching a peak AUC of 93.0 on MedGemma-4B. White-box hidden-state access provides a consistent advantage of 19.5 AUC points on average (range: 9.5--33.5), with ReXTrust maintaining strong performance even on LLaVA-v1.6-7B (AUC 79.9), where black-box methods and clustering-based gray-box methods collapse to near-chance performance. Among non-white-box methods, token-level gray-box statistics (MaxEnt, MaxProb) are the strongest alternatives, outperforming both clustering-based gray-box methods (Semantic Entropy, VASE) and black-box approaches on average. We further identify confident confabulation, a failure mode in which models hallucinate with high inter-sample consistency or high token-level probability, as a systemic failure for both consistency and uncertainty-based methods.

翻译：暂无翻译

0

相关内容

AUC

在无标注条件下适配视觉—语言模型：全面综述

在无标注条件下适配视觉—语言模型：全面综述

专知会员服务

13+阅读 · 2025年8月9日

面向视觉语言模型的持续学习：遗忘之外的综述与分类体系

面向视觉语言模型的持续学习：遗忘之外的综述与分类体系

专知会员服务

21+阅读 · 2025年8月9日

视觉语言模型泛化到新领域：全面综述

视觉语言模型泛化到新领域：全面综述

专知会员服务

38+阅读 · 2025年6月27日

【ICML2025】使用树搜索重新排序推理上下文，使大型视觉语言模型更强大

【ICML2025】使用树搜索重新排序推理上下文，使大型视觉语言模型更强大

专知会员服务

7+阅读 · 2025年6月10日

高效视觉语言模型研究综述

高效视觉语言模型研究综述

专知会员服务

14+阅读 · 2025年4月18日

【CVPR2025】Mamba 作为桥梁：连接视觉基础模型与视觉语言模型以实现领域泛化语义分割

【CVPR2025】Mamba 作为桥梁：连接视觉基础模型与视觉语言模型以实现领域泛化语义分割

专知会员服务

14+阅读 · 2025年4月12日

【CVPR2025】Mamba 作为桥梁：连接视觉基础模型与视觉语言模型以实现跨领域的语义分割

【CVPR2025】Mamba 作为桥梁：连接视觉基础模型与视觉语言模型以实现跨领域的语义分割

专知会员服务

17+阅读 · 2025年4月7日

《Med3DVLM：面向三维医学图像分析的高效视觉-语言模型》

《Med3DVLM：面向三维医学图像分析的高效视觉-语言模型》

专知会员服务

9+阅读 · 2025年3月27日

【NeurlPS2024】一种适用于跨模态和任务的视觉-语言模型的统一去偏方法

【NeurlPS2024】一种适用于跨模态和任务的视觉-语言模型的统一去偏方法

专知会员服务

22+阅读 · 2024年10月11日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

NLP领域最近比较火的Prompt，能否借鉴到多模态领域？一文跟进最新进展

NLP领域最近比较火的Prompt，能否借鉴到多模态领域？一文跟进最新进展

PaperWeekly

17+阅读 · 2022年3月8日

图卷积神经网络蒸馏知识，Distillating Knowledge from GCN

图卷积神经网络蒸馏知识，Distillating Knowledge from GCN

专知

41+阅读 · 2020年3月25日

预训练语言模型关系图+必读论文列表，清华荣誉出品

预训练语言模型关系图+必读论文列表，清华荣誉出品

机器之心

18+阅读 · 2019年10月11日

ACL 2019论文分享：ARNOR增强模型注意力，降低远监督学习中的噪声

ACL 2019论文分享：ARNOR增强模型注意力，降低远监督学习中的噪声

AINLP

53+阅读 · 2019年8月15日

【泡泡图灵智库】VITAMIN-E:极密集特征点的视觉跟踪和建图（CVPR）

【泡泡图灵智库】VITAMIN-E:极密集特征点的视觉跟踪和建图（CVPR）

泡泡机器人SLAM

10+阅读 · 2019年6月14日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

Focal Loss for Dense Object Detection

Focal Loss for Dense Object Detection

统计学习与视觉计算组

12+阅读 · 2018年3月15日

论文浅尝 | Improved Neural Relation Detection for KBQA

论文浅尝 | Improved Neural Relation Detection for KBQA

开放知识图谱

13+阅读 · 2018年1月21日

新型双组份Camassa-Holm方程的等谱问题及适定性研究

国家自然科学基金

0+阅读 · 2015年12月31日

Sigma 1受体对血管性痴呆小鼠血脑屏障的调节作用及机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

外源性外泌体对mTOR通路参与柯萨奇病毒B3诱导细胞凋亡的调控作用及机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

Polysulfides介导Caveolin-1硫巯基化修饰在LDL跨内皮细胞穿胞及早期AS中的作用

国家自然科学基金

0+阅读 · 2015年12月31日

肿瘤细胞生长抑制剂Gemmacolides的靶标研究

国家自然科学基金

0+阅读 · 2015年12月31日

Waardenburg综合征的拷贝数变异检测及其致病机制的研究

国家自然科学基金

0+阅读 · 2015年12月31日

阿尔茨海默病生物标志物的电化学发光成像分析

国家自然科学基金

0+阅读 · 2015年12月31日

HOXA5通过CHOP介导的凋亡途径抑制胆管癌的增殖作用研究

国家自然科学基金

0+阅读 · 2015年12月31日

骨髓间充质干细胞在缺血性卒中后抑制水通道蛋白4保护血脑屏障完整性的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

腹侧被盖区GABA通路对缺血海马新生神经发生的影响及机制

国家自然科学基金

0+阅读 · 2014年12月31日

REALM: A Unified Red-Teaming Benchmark for Physical-World VLMs

Arxiv

0+阅读 · 6月22日

Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation

Arxiv

0+阅读 · 6月22日

Wasserstein-Aligned Localisation for VLM-Based Distributional OOD Detection in Medical Imaging

Arxiv

0+阅读 · 6月20日

PROTON: Prototype-Based Test-Time Online OOD Detection for Medical VLMs

Arxiv

0+阅读 · 6月18日

Spectral Query-Key Product Weight Steering for Training-Free VLM Hallucination Mitigation

Arxiv

0+阅读 · 6月18日

Beyond 'One Language, One Script': Quantifying Orthographic Bias in Multilingual VLMs with PuMVR

Arxiv

0+阅读 · 6月18日

Beyond the Linear Separability Ceiling: Aligning Representations in VLMs

Arxiv

0+阅读 · 6月17日

Moving Beyond Diversity: Visual Token Pruning as Subspace Reconstruction for Efficient VLMs

Arxiv

0+阅读 · 6月17日

Hallucination Detection and Correction in Medical VLMs via Counter-Evidence Verification

Arxiv

0+阅读 · 6月17日

CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework

Arxiv

0+阅读 · 6月16日

VIP会员

文章信息

相关主题

最新内容

无人机自主控制与人工智能：系统性综述

无人机自主控制与人工智能：系统性综述

专知会员服务

1+阅读 · 8分钟前

巡飞弹与反无人机系统——现代战场的两大支柱

巡飞弹与反无人机系统——现代战场的两大支柱

专知会员服务

1+阅读 · 39分钟前

《打造“黄金舰队”》57页报告

《打造“黄金舰队”》57页报告

专知会员服务

0+阅读 · 41分钟前

《北约数字教官网络发展路径》128页报告

《北约数字教官网络发展路径》128页报告

专知会员服务

1+阅读 · 今天6:33

ECCV 2026 | MIMFlow：MIM与归一化流统一图像生成

ECCV 2026 | MIMFlow：MIM与归一化流统一图像生成

专知会员服务

6+阅读 · 6月25日

超越自回归边界：扩散模型、世界模型与SSM如何重塑代码智能

超越自回归边界：扩散模型、世界模型与SSM如何重塑代码智能

专知会员服务

5+阅读 · 6月25日

重塑决策优势：美军作战艺术与多域作战中联盟联合全域指挥控制（CJADC2）体系的融合

重塑决策优势：美军作战艺术与多域作战中联盟联合全域指挥控制（CJADC2）体系的融合

专知会员服务

7+阅读 · 6月25日

网状网络及其在军事领域的运用

网状网络及其在军事领域的运用

专知会员服务

7+阅读 · 6月25日

《意识即战场——全球安全体系中认知战的演进：乌克兰构建认知作战体系的展望》

《意识即战场——全球安全体系中认知战的演进：乌克兰构建认知作战体系的展望》

专知会员服务

7+阅读 · 6月25日

无美国参与的欧洲战争方式（万字长文）

无美国参与的欧洲战争方式（万字长文）

专知会员服务

8+阅读 · 6月25日

重构“下一场战争”的制胜理论：超越兰彻斯特方程与现代系统

重构“下一场战争”的制胜理论：超越兰彻斯特方程与现代系统

专知会员服务

9+阅读 · 6月25日

《国防工业中基于模型定义的实施：产品定义数字化转型的战略路径》90页

《国防工业中基于模型定义的实施：产品定义数字化转型的战略路径》90页

专知会员服务

9+阅读 · 6月25日

《国防领域敏感性分析白皮书》

《国防领域敏感性分析白皮书》

专知会员服务

8+阅读 · 6月25日

综述 | 从问答到任务完成：Agent系统与Harness设计

综述 | 从问答到任务完成：Agent系统与Harness设计

专知会员服务

9+阅读 · 6月24日

Agentic RL：框架、实践与长程智能体训练

Agentic RL：框架、实践与长程智能体训练

专知会员服务

10+阅读 · 6月24日

相关VIP内容

在无标注条件下适配视觉—语言模型：全面综述

在无标注条件下适配视觉—语言模型：全面综述

专知会员服务

13+阅读 · 2025年8月9日

面向视觉语言模型的持续学习：遗忘之外的综述与分类体系

面向视觉语言模型的持续学习：遗忘之外的综述与分类体系

专知会员服务

21+阅读 · 2025年8月9日

视觉语言模型泛化到新领域：全面综述

视觉语言模型泛化到新领域：全面综述

专知会员服务

38+阅读 · 2025年6月27日

【ICML2025】使用树搜索重新排序推理上下文，使大型视觉语言模型更强大

【ICML2025】使用树搜索重新排序推理上下文，使大型视觉语言模型更强大

专知会员服务

7+阅读 · 2025年6月10日

高效视觉语言模型研究综述

高效视觉语言模型研究综述

专知会员服务

14+阅读 · 2025年4月18日

【CVPR2025】Mamba 作为桥梁：连接视觉基础模型与视觉语言模型以实现领域泛化语义分割

【CVPR2025】Mamba 作为桥梁：连接视觉基础模型与视觉语言模型以实现领域泛化语义分割

专知会员服务

14+阅读 · 2025年4月12日

【CVPR2025】Mamba 作为桥梁：连接视觉基础模型与视觉语言模型以实现跨领域的语义分割

【CVPR2025】Mamba 作为桥梁：连接视觉基础模型与视觉语言模型以实现跨领域的语义分割

专知会员服务

17+阅读 · 2025年4月7日

《Med3DVLM：面向三维医学图像分析的高效视觉-语言模型》

《Med3DVLM：面向三维医学图像分析的高效视觉-语言模型》

专知会员服务

9+阅读 · 2025年3月27日

【NeurlPS2024】一种适用于跨模态和任务的视觉-语言模型的统一去偏方法

【NeurlPS2024】一种适用于跨模态和任务的视觉-语言模型的统一去偏方法

专知会员服务

22+阅读 · 2024年10月11日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

热门VIP内容

开通专知VIP会员享更多权益服务

《打造“黄金舰队”》57页报告

ECCV 2026 | MIMFlow：MIM与归一化流统一图像生成

巡飞弹与反无人机系统——现代战场的两大支柱

《北约数字教官网络发展路径》128页报告

相关资讯

NLP领域最近比较火的Prompt，能否借鉴到多模态领域？一文跟进最新进展

NLP领域最近比较火的Prompt，能否借鉴到多模态领域？一文跟进最新进展

PaperWeekly

17+阅读 · 2022年3月8日

图卷积神经网络蒸馏知识，Distillating Knowledge from GCN

图卷积神经网络蒸馏知识，Distillating Knowledge from GCN

专知

41+阅读 · 2020年3月25日

预训练语言模型关系图+必读论文列表，清华荣誉出品

预训练语言模型关系图+必读论文列表，清华荣誉出品

机器之心

18+阅读 · 2019年10月11日

ACL 2019论文分享：ARNOR增强模型注意力，降低远监督学习中的噪声

ACL 2019论文分享：ARNOR增强模型注意力，降低远监督学习中的噪声

AINLP

53+阅读 · 2019年8月15日

【泡泡图灵智库】VITAMIN-E:极密集特征点的视觉跟踪和建图（CVPR）

【泡泡图灵智库】VITAMIN-E:极密集特征点的视觉跟踪和建图（CVPR）

泡泡机器人SLAM

10+阅读 · 2019年6月14日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

44+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

Focal Loss for Dense Object Detection

Focal Loss for Dense Object Detection

统计学习与视觉计算组

12+阅读 · 2018年3月15日

论文浅尝 | Improved Neural Relation Detection for KBQA

论文浅尝 | Improved Neural Relation Detection for KBQA

开放知识图谱

13+阅读 · 2018年1月21日

相关论文

REALM: A Unified Red-Teaming Benchmark for Physical-World VLMs

Arxiv

0+阅读 · 6月22日

Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation

Arxiv

0+阅读 · 6月22日

Wasserstein-Aligned Localisation for VLM-Based Distributional OOD Detection in Medical Imaging

Arxiv

0+阅读 · 6月20日

PROTON: Prototype-Based Test-Time Online OOD Detection for Medical VLMs

Arxiv

0+阅读 · 6月18日

Spectral Query-Key Product Weight Steering for Training-Free VLM Hallucination Mitigation

Arxiv

0+阅读 · 6月18日

Beyond 'One Language, One Script': Quantifying Orthographic Bias in Multilingual VLMs with PuMVR

Arxiv

0+阅读 · 6月18日

Beyond the Linear Separability Ceiling: Aligning Representations in VLMs

Arxiv

0+阅读 · 6月17日

Moving Beyond Diversity: Visual Token Pruning as Subspace Reconstruction for Efficient VLMs

Arxiv

0+阅读 · 6月17日

Hallucination Detection and Correction in Medical VLMs via Counter-Evidence Verification

Arxiv

0+阅读 · 6月17日

CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework

Arxiv

0+阅读 · 6月16日

相关基金

新型双组份Camassa-Holm方程的等谱问题及适定性研究

国家自然科学基金

0+阅读 · 2015年12月31日

Sigma 1受体对血管性痴呆小鼠血脑屏障的调节作用及机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

外源性外泌体对mTOR通路参与柯萨奇病毒B3诱导细胞凋亡的调控作用及机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

Polysulfides介导Caveolin-1硫巯基化修饰在LDL跨内皮细胞穿胞及早期AS中的作用

国家自然科学基金

0+阅读 · 2015年12月31日

肿瘤细胞生长抑制剂Gemmacolides的靶标研究

国家自然科学基金

0+阅读 · 2015年12月31日

Waardenburg综合征的拷贝数变异检测及其致病机制的研究

国家自然科学基金

0+阅读 · 2015年12月31日

阿尔茨海默病生物标志物的电化学发光成像分析

国家自然科学基金

0+阅读 · 2015年12月31日

HOXA5通过CHOP介导的凋亡途径抑制胆管癌的增殖作用研究

国家自然科学基金

0+阅读 · 2015年12月31日

骨髓间充质干细胞在缺血性卒中后抑制水通道蛋白4保护血脑屏障完整性的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

腹侧被盖区GABA通路对缺血海马新生神经发生的影响及机制

国家自然科学基金

0+阅读 · 2014年12月31日

微信扫码咨询专知VIP会员