Publishing open-source academic video recordings is an emergent and prevalent approach to sharing knowledge online. Such videos carry rich multimodal information including speech, the facial and body movements of the speakers, as well as the texts and pictures in the slides and possibly even the papers. Although multiple academic video datasets have been constructed and released, few of them support both multimodal content recognition and understanding tasks, which is partially due to the lack of high-quality human annotations. In this paper, we propose a novel multimodal, multigenre, and multipurpose audio-visual academic lecture dataset (M$^3$AV), which has almost 367 hours of videos from five sources covering computer science, mathematics, and medical and biology topics. With high-quality human annotations of the spoken and written words, in particular high-valued name entities, the dataset can be used for multiple audio-visual recognition and understanding tasks. Evaluations performed on contextual speech recognition, speech synthesis, and slide and script generation tasks demonstrate that the diversity of M$^3$AV makes it a challenging dataset.
翻译:发布开源学术视频记录是一种新兴且流行的在线知识分享方式。此类视频承载着丰富的多模态信息,包括语音、演讲者的面部及身体动作,以及幻灯片乃至论文中的文字和图片。尽管已有多个学术视频数据集被构建并发布,但其中鲜有能同时支持多模态内容识别与理解任务的数据集,这在一定程度上是由于缺乏高质量的人工标注。本文提出了一个新颖的多模态、多体裁、多用途学术讲座音视频数据集(M$^3$AV),该数据集包含来自五个来源的近367小时视频,涵盖计算机科学、数学、医学及生物学等主题。凭借对口语和书面文字(尤其是高价值命名实体)的高质量人工标注,该数据集可用于多种音视频识别与理解任务。在上下文语音识别、语音合成以及幻灯片与脚本生成任务上的评估表明,M$^3$AV的多样性使其成为一个具有挑战性的数据集。