The growing importance of multi-modal humor detection within affective computing correlates with the expanding influence of short-form video sharing on social media platforms. In this paper, we propose a novel two-branch hierarchical model for short-form video humor detection (SVHD), named Comment-aided Video-Language Alignment (CVLA) via data-augmented multi-modal contrastive pre-training. Notably, our CVLA not only operates on raw signals across various modal channels but also yields an appropriate multi-modal representation by aligning the video and language components within a consistent semantic space. The experimental results on two humor detection datasets, including DY11k and UR-FUNNY, demonstrate that CVLA dramatically outperforms state-of-the-art and several competitive baseline approaches. Our dataset, code and model release at https://github.com/yliu-cs/CVLA.
翻译:多模态幽默检测在情感计算中的重要性日益凸显,这与短视频分享在社交媒体平台上的影响力不断扩大密切相关。本文提出了一种新颖的双分支层次化模型用于短视频幽默检测(SVHD),命名为基于数据增强多模态对比预训练的评论辅助视频-语言对齐(CVLA)。值得注意的是,我们的CVLA不仅能在多个模态通道上处理原始信号,还能通过将视频和语言组件对齐到一致的语义空间中,生成恰当的多模态表示。在两个幽默检测数据集DY11k和UR-FUNNY上的实验结果表明,CVLA显著优于现有最先进方法及多个具有竞争力的基线方法。我们的数据集、代码和模型已在https://github.com/yliu-cs/CVLA发布。