As short-form funny videos on social networks are gaining popularity, it becomes demanding for AI models to understand them for better communication with humans. Unfortunately, previous video humor datasets target specific domains, such as speeches or sitcoms, and mostly focus on verbal cues. We curate a user-generated dataset of 10K multimodal funny videos from YouTube, called ExFunTube. Using a video filtering pipeline with GPT-3.5, we verify both verbal and visual elements contributing to humor. After filtering, we annotate each video with timestamps and text explanations for funny moments. Our ExFunTube is unique over existing datasets in that our videos cover a wide range of domains with various types of humor that necessitate a multimodal understanding of the content. Also, we develop a zero-shot video-to-text prompting to maximize video humor understanding of large language models (LLMs). With three different evaluation methods using automatic scores, rationale quality experiments, and human evaluations, we show that our prompting significantly improves LLMs' ability for humor explanation.
翻译:随着社交网络上搞笑短视频日益流行,人工智能模型理解这些内容以更好地与人类沟通变得愈发重要。然而,现有的视频幽默数据集大多局限于特定领域(如演讲或情景喜剧),且主要关注语言线索。我们构建了一个名为ExFunTube的用户生成数据集,包含来自YouTube的1万个多模态搞笑视频。通过采用基于GPT-3.5的视频过滤流水线,我们验证了语言和视觉元素对幽默的贡献。在过滤后,我们为每个视频标注了搞笑时刻的时间戳和文字解释。与现有数据集相比,ExFunTube的独特性在于其视频覆盖广泛领域,包含多种幽默类型,需要多模态内容理解。此外,我们提出了一种零样本视频到文本提示方法,以最大化大型语言模型(LLMs)对视频幽默的理解能力。通过自动评分、推理质量实验和人工评估三种评价方法,我们证明该提示方法显著提升了LLMs的幽默解释能力。