Fine-grained perception of multimodal information is critical for advancing human-AI interaction. With recent progress in audio-visual technologies, Omni Language Models (OLMs), capable of processing audio and video signals in parallel, have emerged as a promising paradigm for achieving richer understanding and reasoning. However, their capacity to capture and describe fine-grained details remains limited explored. In this work, we present a systematic and comprehensive investigation of omni detailed perception from the perspectives of the data pipeline, models, and benchmark. We first identify an inherent "co-growth" between detail and hallucination in current OLMs. To address this, we propose Omni-Detective, an agentic data generation pipeline integrating tool-calling, to autonomously produce highly detailed yet minimally hallucinatory multimodal data. Based on the data generated with Omni-Detective, we train two captioning models: Audio-Captioner for audio-only detailed perception, and Omni-Captioner for audio-visual detailed perception. Under the cascade evaluation protocol, Audio-Captioner achieves the best performance on MMAU and MMAR among all open-source models, surpassing Gemini 2.5 Flash and delivering performance comparable to Gemini 2.5 Pro. On existing detailed captioning benchmarks, Omni-Captioner sets a new state-of-the-art on VDC and achieves the best trade-off between detail and hallucination on the video-SALMONN 2 testset. Given the absence of a dedicated benchmark for omni detailed perception, we design Omni-Cloze, a novel cloze-style evaluation for detailed audio, visual, and audio-visual captioning that ensures stable, efficient, and reliable assessment. Experimental results and analysis demonstrate the effectiveness of Omni-Detective in generating high-quality detailed captions, as well as the superiority of Omni-Cloze in evaluating such detailed captions.


翻译:多模态信息的细粒度感知对于推进人机交互至关重要。随着视听技术的近期进展,能够并行处理音频与视频信号的Omni语言模型(OLMs)已成为实现更丰富理解与推理的前沿范例。然而,它们在捕捉与描述细粒度细节方面的能力仍缺乏深入探索。本工作从数据流程、模型和基准三个维度对全维精细感知展开了系统而全面的研究。我们首先发现当前OLMs中细节与幻觉之间存在固有的“协同增长”现象。为解决此问题,我们提出Omni-Detective——一种集成工具调用的智能体数据生成流程,可自主产生高细节度且低幻觉的多模态数据。基于Omni-Detective生成的数据,我们训练了两个描述模型:面向纯音频精细感知的Audio-Captioner,以及面向视听精细感知的Omni-Captioner。在级联评估协议下,Audio-Captioner在所有开源模型中于MMAU和MMAR基准上取得最佳性能,超越Gemini 2.5 Flash,并与Gemini 2.5 Pro性能相当。在现有精细描述基准上,Omni-Captioner在VDC上创下新最优结果,并在video-SALMONN 2测试集上实现了细节与幻觉的最佳权衡。鉴于全维精细感知专用基准的缺失,我们设计了Omni-Cloze——一种面向精细音频、视觉及视听描述的新型完形填空式评估方法,可确保稳定、高效且可靠的评测。实验结果与分析证明了Omni-Detective在生成高质量精细描述方面的有效性,以及Omni-Cloze在评估此类精细描述上的优越性。

0
下载
关闭预览

相关内容

大规模视觉-语言模型的基准、评估、应用与挑战
专知会员服务
18+阅读 · 2025年2月10日
视觉里程计:起源、优势、对比、应用
计算机视觉life
18+阅读 · 2017年7月17日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
面向2027年及未来的海军情报改革
专知会员服务
2+阅读 · 8月5日
《无人机蜂群:释放人类-蜂群编队的潜能》
专知会员服务
4+阅读 · 8月5日
《战略战术化:一项综合性述评》
专知会员服务
2+阅读 · 8月5日
相关VIP内容
大规模视觉-语言模型的基准、评估、应用与挑战
专知会员服务
18+阅读 · 2025年2月10日
相关资讯
视觉里程计:起源、优势、对比、应用
计算机视觉life
18+阅读 · 2017年7月17日
相关基金
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员