We introduce EPIC-SOUNDS, a large-scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos. We propose an annotation pipeline where annotators temporally label distinguishable audio segments and describe the action that could have caused this sound. We identify actions that can be discriminated purely from audio, through grouping these free-form descriptions of audio into classes. For actions that involve objects colliding, we collect human annotations of the materials of these objects (e.g. a glass object being placed on a wooden surface), which we verify from visual labels, discarding ambiguities. Overall, EPIC-SOUNDS includes 78.4k categorised segments of audible events and actions, distributed across 44 classes as well as 39.2k non-categorised segments. We train and evaluate two state-of-the-art audio recognition models on our dataset, highlighting the importance of audio-only labels and the limitations of current models to recognise actions that sound.
翻译:我们提出了EPIC-SOUNDS,这是一个大规模音频注释数据集,捕捉了第一人称视角视频中音频流的时间范围与类别标签。我们设计了一套注释流程,标注员可对可区分的音频片段进行时间标注,并描述可能引发该声音的动作。通过将这些自由形式的音频描述进行分组归类,我们识别出可仅凭音频进行区分的动作。针对涉及物体碰撞的动作,我们收集了这些物体材质的人类注释(例如玻璃物体放置在木质表面上),并通过视觉标签进行验证,排除了歧义。总体而言,EPIC-SOUNDS包含78.4k个已分类的可闻事件与动作片段,分布于44个类别,以及39.2k个未分类片段。我们训练并评估了两个当前最先进的音频识别模型,突显了仅凭音频标签的重要性以及现有模型在识别“声音动作”方面的局限性。