While vision-language pretrained models (VLMs) excel in various multimodal understanding tasks, their potential in fine-grained audio-visual reasoning, particularly for audio-visual question answering (AVQA), remains largely unexplored. AVQA presents specific challenges for VLMs due to the requirement of visual understanding at the region level and seamless integration with audio modality. Previous VLM-based AVQA methods merely used CLIP as a feature encoder but underutilized its knowledge, and mistreated audio and video as separate entities in a dual-stream framework as most AVQA methods. This paper proposes a new CLIP-powered target-aware single-stream (TASS) network for AVQA using the image-text matching knowledge of the pretrained model through the audio-visual matching characteristic of nature. It consists of two key components: the target-aware spatial grounding module (TSG+) and the single-stream joint temporal grounding module (JTG). Specifically, we propose a TSG+ module to transfer the image-text matching knowledge from CLIP models to our region-text matching process without corresponding ground-truth labels. Moreover, unlike previous separate dual-stream networks that still required an additional audio-visual fusion module, JTG unifies audio-visual fusion and question-aware temporal grounding in a simplified single-stream architecture. It treats audio and video as a cohesive entity and further extends the pretrained image-text knowledge to audio-text matching by preserving their temporal correlation with our proposed cross-modal synchrony (CMS) loss. Extensive experiments conducted on the MUSIC-AVQA benchmark verified the effectiveness of our proposed method over existing state-of-the-art methods.
翻译:尽管视觉-语言预训练模型(VLMs)在各种多模态理解任务中表现出色,但其在细粒度视听推理中的潜力,特别是针对视听问答(AVQA)任务,仍未得到充分探索。AVQA给VLMs带来了特定挑战,因为它要求对区域级别的视觉理解以及与音频模态的无缝集成。以往基于VLM的AVQA方法仅将CLIP用作特征编码器,未能充分利用其知识,并且像大多数AVQA方法一样,在双流框架中将音频和视频视为独立实体。本文提出了一种新的CLIP驱动的目标感知单流(TASS)网络用于AVQA,利用预训练模型的图像-文本匹配知识以及自然界的视听匹配特性。它由两个关键组件组成:目标感知空间定位模块(TSG+)和单流联合时间定位模块(JTG)。具体而言,我们提出了TSG+模块,将CLIP模型的图像-文本匹配知识迁移到区域-文本匹配过程,而无需对应的真实标签。此外,与仍需要额外视听融合模块的以往独立双流网络不同,JTG在一个简化的单流架构中统一了视听融合和基于问题的时间定位。它将音频和视频视为一个整体实体,并通过我们提出的跨模态同步(CMS)损失保留它们的时间相关性,将预训练的图像-文本知识进一步扩展到音频-文本匹配。在MUSIC-AVQA基准上进行的大量实验验证了所提方法相对于现有最先进方法的有效性。