Reference Audio-Visual Segmentation (Ref-AVS) aims to segment objects in audible videos based on multimodal cues in reference expressions. Previous methods overlook the explicit recognition of expression difficulty and dominant modality in multimodal cues, over-rely on the quality of the instruction-tuning dataset for object reasoning, and lack reflective validation of segmentation results, leading to erroneous mask predictions. To address these issues, in this paper, we propose a novel training-free Multi-Agent Recognition, Reasoning, and Reflection framework to achieve high-quality Reference Audio-Visual Segmentation, termed MAR3. Incorporating the sociological Delphi theory to achieve robust analysis, a Consensus Multimodal Recognition mechanism is proposed that enables LLM agents to explicitly recognize the difficulty of reference expressions and the dominant modality of multimodal cues. Based on our modality-dominant difficulty rule, we propose an adaptive Collaborative Object Reasoning strategy to reliably reason about the referred object. To further ensure precise mask prediction, we develop a Reflective Learning Segmentation mechanism, in which a check agent examines intermediate segmentation results and iteratively corrects the object text prompt of the segment agent. Experiments demonstrate that MAR3 achieves superior performance (69.2% in J&F) on the Ref-AVSBench dataset, outperforming SOTA by 3.4% absolutely.
翻译:参考音频-视觉分割(Ref-AVS)旨在根据参考表达式中的多模态线索,对可听视频中的对象进行分割。以往方法忽视了显式识别表达式难度和多模态线索中的主导模态,过度依赖指令调优数据集的质量进行对象推理,且缺乏对分割结果的反思性验证,导致掩膜预测出现错误。针对这些问题,本文提出了一种新颖的免训练多智能体识别、推理与反思框架,以实现高质量参考音频-视觉分割,简称MAR3。通过引入社会学德尔菲理论实现稳健分析,提出了共识多模态识别机制,使大语言模型智能体能够显式识别参考表达式的难度以及多模态线索中的主导模态。基于本文提出的模态主导难度规则,进一步设计了自适应协同对象推理策略,以可靠地推理所指对象。为确保掩膜预测的精确性,开发了反思式学习分割机制,其中核查智能体对中间分割结果进行审查,并迭代修正分割智能体的对象文本提示。实验表明,MAR3在Ref-AVSBench数据集上取得了优越性能(J&F指标69.2%),绝对性能超过现有最优方法3.4%。