Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them. However, this setting is unrealistic as natural videos often contain numerous audio-visual events with different categories. To better adapt to real-life applications, in this paper we focus on the task of dense-localizing audio-visual events, which aims to jointly localize and recognize all audio-visual events occurring in an untrimmed video. The problem is challenging as it requires fine-grained audio-visual scene and context understanding. To tackle this problem, we introduce the first Untrimmed Audio-Visual (UnAV-100) dataset, which contains 10K untrimmed videos with over 30K audio-visual events. Each video has 2.8 audio-visual events on average, and the events are usually related to each other and might co-occur as in real-life scenes. Next, we formulate the task using a new learning-based framework, which is capable of fully integrating audio and visual modalities to localize audio-visual events with various lengths and capture dependencies between them in a single pass. Extensive experiments demonstrate the effectiveness of our method as well as the significance of multi-scale cross-modal perception and dependency modeling for this task.
翻译:现有的音视频事件定位(AVE)方法仅处理经过人工裁剪且包含单个实例的视频。然而,这种设定并不符合现实——自然视频往往包含多个不同类别的音视频事件。为更好地适应实际应用场景,本文聚焦于密集定位音视频事件任务,旨在联合定位并识别未裁剪视频中发生的所有音视频事件。该问题具有挑战性,因其需要细粒度的音视频场景与上下文理解。为此,我们首次提出未裁剪音视频数据集(UnAV-100),包含10K个未裁剪视频及超过30K个音视频事件。每个视频平均包含2.8个音视频事件,这些事件通常相互关联,并可能像真实场景中那样同时发生。随后,我们提出一种新的学习框架来形式化该任务,该框架能够充分融合音频与视觉模态,在单次处理中定位不同长度的音视频事件并捕获其依赖关系。大量实验证明了我们方法的有效性,以及多尺度跨模态感知和依赖建模对该任务的重要意义。