The aim of audio-visual segmentation (AVS) is to precisely differentiate audible objects within videos down to the pixel level. Traditional approaches often tackle this challenge by combining information from various modalities, where the contribution of each modality is implicitly or explicitly modeled. Nevertheless, the interconnections between different modalities tend to be overlooked in audio-visual modeling. In this paper, inspired by the human ability to mentally simulate the sound of an object and its visual appearance, we introduce a bidirectional generation framework. This framework establishes robust correlations between an object's visual characteristics and its associated sound, thereby enhancing the performance of AVS. To achieve this, we employ a visual-to-audio projection component that reconstructs audio features from object segmentation masks and minimizes reconstruction errors. Moreover, recognizing that many sounds are linked to object movements, we introduce an implicit volumetric motion estimation module to handle temporal dynamics that may be challenging to capture using conventional optical flow methods. To showcase the effectiveness of our approach, we conduct comprehensive experiments and analyses on the widely recognized AVSBench benchmark. As a result, we establish a new state-of-the-art performance level in the AVS benchmark, particularly excelling in the challenging MS3 subset which involves segmenting multiple sound sources. To facilitate reproducibility, we plan to release both the source code and the pre-trained model.
翻译:音频-视觉分割(AVS)的目标是在像素级别精确区分视频中的可听对象。传统方法通常通过融合多模态信息来解决这一挑战,其中每种模态的贡献被显式或隐式建模。然而,不同模态之间的相互关联在音频-视觉建模中往往被忽视。本文受人类在脑海中模拟物体声音及其视觉表象能力的启发,提出了一种双向生成框架。该框架建立了物体视觉特征与其关联声音之间的稳健关联,从而提升了AVS的性能。为实现这一目标,我们采用了一个视觉到音频的投影组件,从对象分割掩码中重建音频特征并最小化重建误差。此外,考虑到许多声音与物体运动相关,我们引入了一个隐式体积运动估计模块,用于处理传统光流方法难以捕捉的时间动态。为展示我们方法的有效性,我们在广泛认可的AVSBench基准上进行了全面实验与分析。最终,我们在AVS基准上实现了新的最优性能,尤其在涉及多声源分割的具有挑战性的MS3子集上表现突出。为促进可重复性,我们计划公开源代码和预训练模型。