Weakly supervised audio-visual video parsing (AVVP) methods aim to detect audible-only, visible-only, and audible-visible events using only video-level labels. Existing approaches tackle this by leveraging unimodal and cross-modal contexts. However, we argue that while cross-modal learning is beneficial for detecting audible-visible events, in the weakly supervised scenario, it negatively impacts unaligned audible or visible events by introducing irrelevant modality information. In this paper, we propose CoLeaF, a novel learning framework that optimizes the integration of cross-modal context in the embedding space such that the network explicitly learns to combine cross-modal information for audible-visible events while filtering them out for unaligned events. Additionally, as videos often involve complex class relationships, modelling them improves performance. However, this introduces extra computational costs into the network. Our framework is designed to leverage cross-class relationships during training without incurring additional computations at inference. Furthermore, we propose new metrics to better evaluate a method's capabilities in performing AVVP. Our extensive experiments demonstrate that CoLeaF significantly improves the state-of-the-art results by an average of 1.9% and 2.4% F-score on the LLP and UnAV-100 datasets, respectively.
翻译:弱监督音视频解析(AVVP)方法旨在仅利用视频级别标签检测仅音频事件、仅视频事件以及音视频共存事件。现有方法通过利用单模态和跨模态上下文来解决该问题。然而,我们认为,尽管跨模态学习有助于检测音视频共存事件,但在弱监督场景下,引入不相关的模态信息会对未对齐的音频或视频事件产生负面影响。本文提出CoLeaF,一种新颖的学习框架,通过在嵌入空间中优化跨模态上下文的整合,使网络能够显式学习为音视频共存事件组合跨模态信息,同时为未对齐事件过滤跨模态信息。此外,由于视频常涉及复杂的类别关系,建模这些关系可提升性能,但这会引入额外计算成本。我们的框架可在训练期间利用跨类别关系,而无需在推理时增加计算量。我们还提出了新的评估指标,以更全面地衡量方法的AVVP能力。大量实验表明,CoLeaF在LLP和UnAV-100数据集上分别将最先进结果的平均F值显著提升了1.9%和2.4%。