Audio and visual signals typically occur simultaneously, and humans possess an innate ability to correlate and synchronize information from these two modalities. Recently, a challenging problem known as Audio-Visual Segmentation (AVS) has emerged, intending to produce segmentation maps for sounding objects within a scene. However, the methods proposed so far have not sufficiently integrated audio and visual information, and the computational costs have been extremely high. Additionally, the outputs of different stages have not been fully utilized. To facilitate this research, we introduce a novel Progressive Confident Masking Attention Network (PMCANet). It leverages attention mechanisms to uncover the intrinsic correlations between audio signals and visual frames. Furthermore, we design an efficient and effective cross-attention module to enhance semantic perception by selecting query tokens. This selection is determined through confidence-driven units based on the network's multi-stage predictive outputs. Experiments demonstrate that our network outperforms other AVS methods while requiring less computational resources.
翻译:音频与视觉信号通常同时出现,人类天生具备关联和同步这两种模态信息的能力。近期,一个被称为视听分割(AVS)的挑战性问题应运而生,其目标是为场景中的发声物体生成分割图。然而,现有方法未能充分整合音频与视觉信息,且计算成本极高。此外,不同阶段的输出尚未得到充分利用。为推进该领域研究,本文提出一种新颖的渐进式置信掩码注意力网络(PMCANet)。该网络利用注意力机制揭示音频信号与视觉帧之间的内在关联。进一步地,我们设计了一个高效且有效的交叉注意力模块,通过选择查询令牌来增强语义感知能力。该选择过程由基于网络多阶段预测输出的置信度驱动单元决定。实验表明,我们的网络在减少计算资源需求的同时,性能优于其他AVS方法。