Despite the recent advances of the artificial intelligence, building social intelligence remains a challenge. Among social signals, laughter is one of the distinctive expressions that occurs during social interactions between humans. In this work, we tackle a new challenge for machines to understand the rationale behind laughter in video, Video Laugh Reasoning. We introduce this new task to explain why people laugh in a particular video and a dataset for this task. Our proposed dataset, SMILE, comprises video clips and language descriptions of why people laugh. We propose a baseline by leveraging the reasoning capacity of large language models (LLMs) with textual video representation. Experiments show that our baseline can generate plausible explanations for laughter. We further investigate the scalability of our baseline by probing other video understanding tasks and in-the-wild videos. We release our dataset, code, and model checkpoints on https://github.com/SMILE-data/SMILE.
翻译:尽管人工智能近年来取得显著进展,构建社会智能仍是一项挑战。在众多社交信号中,笑声是人类社会互动中最具特色的表达方式之一。本研究提出全新的"视频笑声推理"任务,旨在让机器理解视频中笑声产生的原因。我们为这一任务构建了数据集SMILE,该数据集包含视频片段及其对应笑声成因的语言描述。我们通过利用大语言模型(LLMs)的推理能力结合文本化视频表征,提出了基线方法。实验表明,该基线能够生成具有合理性的笑声解释。我们进一步通过探究其他视频理解任务及真实场景视频,验证了该方法的可扩展性。数据集、代码及模型检查点已发布于https://github.com/SMILE-data/SMILE。