Accurately detecting and localizing hallucinations is a critical task for ensuring high reliability of image captions. In the era of Multimodal Large Language Models (MLLMs), captions have evolved from brief sentences into comprehensive narratives, often spanning hundreds of words. This shift exponentially increases the challenge: models must now pinpoint specific erroneous spans or words within extensive contexts, rather than merely flag response-level inconsistencies. However, existing benchmarks lack the fine granularity and domain diversity required to evaluate this capability. To bridge this gap, we introduce DetailVerifyBench, a rigorous benchmark comprising 1,000 high-quality images across five distinct domains. With an average caption length of over 200 words and dense, token-level annotations of multiple hallucination types, it stands as the most challenging benchmark for precise hallucination localization in the field of long image captioning to date. Our benchmark is available at https://zyx-hhnkh.github.io/DetailVerifyBench/.
翻译:准确定位与检测幻觉是确保图像描述高可靠性的关键任务。在多模态大语言模型时代,图像描述已从简短语句演进为涵盖数百词的综合叙事。这一转变指数级提升了挑战难度:模型需在长篇语境中精准定位错误片段或词汇,而非仅标记整体响应层面的不一致。然而现有基准缺乏评估该能力的细粒度与领域多样性。为填补这一空白,我们提出DetailVerifyBench——包含横跨五个不同领域的1,000张高质量图像的严格基准。其平均描述长度超过200词,并配有密集的标记级多类型幻觉标注,成为目前长图像描述领域最具挑战性的精准幻觉定位基准。我们的基准库访问地址为https://zyx-hhnkh.github.io/DetailVerifyBench/。