In recent years, the task of weakly supervised audio-visual violence detection has gained considerable attention. The goal of this task is to identify violent segments within multimodal data based on video-level labels. Despite advances in this field, traditional Euclidean neural networks, which have been used in prior research, encounter difficulties in capturing highly discriminative representations due to limitations of the feature space. To overcome this, we propose HyperVD, a novel framework that learns snippet embeddings in hyperbolic space to improve model discrimination. Our framework comprises a detour fusion module for multimodal fusion, effectively alleviating modality inconsistency between audio and visual signals. Additionally, we contribute two branches of fully hyperbolic graph convolutional networks that excavate feature similarities and temporal relationships among snippets in hyperbolic space. By learning snippet representations in this space, the framework effectively learns semantic discrepancies between violent and normal events. Extensive experiments on the XD-Violence benchmark demonstrate that our method outperforms state-of-the-art methods by a sizable margin.
翻译:近年来,弱监督音视频暴力检测任务受到广泛关注。该任务旨在基于视频级标签识别多模态数据中的暴力片段。尽管该领域已取得进展,但先前研究采用的欧式神经网络因特征空间局限性,难以捕获高度判别性表征。为此,我们提出HyperVD框架——一种在双曲空间中学习片段嵌入以增强模型判别能力的新框架。该框架包含绕行融合模块用于多模态融合,有效缓解音频与视觉信号间的模态不一致性。此外,我们贡献了两个全双曲图卷积网络分支,用于挖掘双曲空间中片段间的特征相似性与时序关系。通过在该空间中学习片段表征,框架可有效学习暴力事件与正常事件间的语义差异。在XD-Violence基准上的大量实验表明,我们的方法以显著优势超越现有最优方法。