Large Audio Language Models (LALMs) still struggle in complex acoustic scenes because they often fail to preserve task-relevant acoustic evidence before reasoning begins. We identify this error pattern as the evidence bottleneck: state-of-the-art systems show larger deficits in acoustic evidence extraction than in downstream reasoning, suggesting that upstream perception is often the limiting factor. To address this problem, we propose EvA (Evidence-First Audio), a dual-path architecture that enhances acoustic evidence preservation through hierarchical aggregation and non-compressive, time-aligned fusion. We also build EvA-Perception, a large-scale training set with about 54K event-ordered captions and 500K evidence-grounded QA pairs. Under a unified zero-shot protocol, EvA achieves the best open-source \emph{Perception} results on MMAU, MMAR, and MMSU, with the largest gains on perception-heavy splits. Human evaluation on open-ended captioning further shows improved fine-grained acoustic coverage and caption quality. These results support the evidence-first hypothesis: stronger audio understanding depends on preserving acoustic evidence before reasoning. Project can be found at https://satsuki2486441738.github.io/EvA/.
翻译:大型音频语言模型(LALMs)在复杂声学场景中仍面临挑战,因其在推理开始前往往无法保留与任务相关的声学证据。我们将此错误模式定义为"证据瓶颈":现有系统在声学证据提取环节的缺陷远大于下游推理环节,表明感知层往往是制约整体性能的关键因素。针对该问题,我们提出EvA(证据优先音频)双路径架构,通过层次化聚合与非压缩式时域对齐融合增强声学证据的保留能力。同时构建包含约54K条事件有序描述和500K组基于证据的问答对的EvA-Perception大规模训练集。在统一零样本协议下,EvA在MMAU、MMAR和MMSU数据集上取得开源模型最优感知性能,尤其在感知密集型任务子集中提升最为显著。针对开放描述任务的人工评估进一步证实,该模型在细粒度声学覆盖度与描述质量方面均有提升。实验结果支持证据优先假说:更强的音频理解依赖于在推理前保留声学证据。项目主页见https://satsuki2486441738.github.io/EvA/。