Audio agents extend large audio-language models (LALMs) by decomposing audio questions into tool calls, intermediate evidence, and iterative reasoning steps. However, as LALMs become stronger, the key challenge shifts from enabling tool use to determining when agentic evidence acquisition genuinely benefits audio understanding. We propose Audio-Mind, an auditable and pluggable framework for conditional evidence acquisition in audio understanding. Audio-Mind dynamically combines a strong frontend with planner-guided tool use, preserving frontend judgment when initial evidence is sufficient while acquiring bounded external evidence for questions with unresolved evidence gaps. Experiments on MMAR and MSU-Bench show that Audio-Mind outperforms prior audio-agent baselines, reaching 80.4% accuracy on MMAR and 82.8% accuracy on MSU-Bench. A matched-backbone comparison highlights why this design matters: under strong audio frontends, agentic decomposition can become an orchestration bottleneck when the workflow does not preserve the frontend's holistic audio-grounded judgment. Beyond accuracy, Audio-Mind produces higher-quality, auditable reasoning traces that expose uncertainty, tool evidence, and answer rationales, offering a potential basis for more reliable audio-QA annotation and error analysis.
翻译:音频智能体通过将音频问题分解为工具调用、中间证据和迭代推理步骤,扩展了大型音频语言模型(LALMs)。然而,随着LALMs能力的增强,核心挑战已从启用工具使用转变为确定何时通过智能体证据获取真正有益于音频理解。我们提出Audio-Mind,一种可审计且可插拔的音频理解条件性证据获取框架。Audio-Mind动态结合了强大的前端系统与规划器引导的工具使用:当初始证据充分时保留前端判断,同时针对仍存在证据缺口的问题获取有限的外部证据。在MMAR和MSU-Bench上的实验表明,Audio-Mind超越了先前的音频智能体基线,在MMAR上达到80.4%的准确率,在MSU-Bench上达到82.8%的准确率。匹配骨干网络的对比进一步揭示了该设计的必要性:在强音频前端条件下,若工作流程未能保留前端基于整体音频的全局判断能力,智能体分解反而可能成为编排瓶颈。除准确率提升外,Audio-Mind还生成了更高质量、可审计的推理轨迹,清晰呈现不确定性、工具证据和答案推理过程,为更可靠的音频问答标注和错误分析提供了潜在基础。