This paper proposes a practical multimodal video summarization task setting and a dataset to train and evaluate the task. The target task involves summarizing a given video into a predefined number of keyframe-caption pairs and displaying them in a listable format to grasp the video content quickly. This task aims to extract crucial scenes from the video in the form of images (keyframes) and generate corresponding captions explaining each keyframe's situation. This task is useful as a practical application and presents a highly challenging problem worthy of study. Specifically, achieving simultaneous optimization of the keyframe selection performance and caption quality necessitates careful consideration of the mutual dependence on both preceding and subsequent keyframes and captions. To facilitate subsequent research in this field, we also construct a dataset by expanding upon existing datasets and propose an evaluation framework. Furthermore, we develop two baseline systems and report their respective performance.
翻译:本文提出了一种实用的多模态视频摘要任务设定,以及用于训练和评估该任务的数据集。目标任务是将给定视频摘要为预设数量的关键帧-描述对,并以列表形式呈现,以便快速掌握视频内容。该任务旨在以图像(关键帧)形式提取视频中的关键场景,并生成相应描述来解释每个关键帧的情境。该任务不仅具有实用价值,还构成了值得研究的高度挑战性问题。具体而言,要实现关键帧选择性能与描述质量的同步优化,必须充分考虑前后关键帧及描述之间的相互依赖关系。为促进该领域的后续研究,我们基于现有数据集进行了扩展构建,并提出了评估框架。此外,我们开发了两个基准系统,并报告了其各自的表现。