There is growing interest in searching for information from large video corpora. Prior works have studied relevant tasks, such as text-based video retrieval, moment retrieval, video summarization, and video captioning in isolation, without an end-to-end setup that can jointly search from video corpora and generate summaries. Such an end-to-end setup would allow for many interesting applications, e.g., a text-based search that finds a relevant video from a video corpus, extracts the most relevant moment from that video, and segments the moment into important steps with captions. To address this, we present the HiREST (HIerarchical REtrieval and STep-captioning) dataset and propose a new benchmark that covers hierarchical information retrieval and visual/textual stepwise summarization from an instructional video corpus. HiREST consists of 3.4K text-video pairs from an instructional video dataset, where 1.1K videos have annotations of moment spans relevant to text query and breakdown of each moment into key instruction steps with caption and timestamps (totaling 8.6K step captions). Our hierarchical benchmark consists of video retrieval, moment retrieval, and two novel moment segmentation and step captioning tasks. In moment segmentation, models break down a video moment into instruction steps and identify start-end boundaries. In step captioning, models generate a textual summary for each step. We also present starting point task-specific and end-to-end joint baseline models for our new benchmark. While the baseline models show some promising results, there still exists large room for future improvement by the community. Project website: https://hirest-cvpr2023.github.io
翻译:从大规模视频语料库中搜索信息的需求日益增长。现有工作孤立地研究了相关任务,如基于文本的视频检索、时刻检索、视频摘要和视频描述,缺乏能够联合实现从视频语料库搜索并生成摘要的端到端框架。这种端到端框架可支持诸多有趣应用,例如:通过文本搜索从视频语料库中定位相关视频,提取该视频中最相关的时刻,并将该时刻细分为带描述的关键步骤。为此,我们提出了HiREST(层级检索与步骤描述)数据集,并构建了一个涵盖层级信息检索及从教学视频语料库生成视觉/文本逐步摘要的新基准。HiREST包含来自教学视频数据集的3400个文本-视频对,其中1100个视频标注了与文本查询相关的时刻跨度,并将每个时刻分解为带描述和时间戳的关键教学步骤(总计8600条步骤描述)。我们的层级基准包括视频检索、时刻检索,以及两项新颖的时刻分割与步骤描述任务。在时刻分割中,模型需将视频时刻分解为教学步骤并识别起止边界;在步骤描述中,模型需为每个步骤生成文本摘要。我们还为这一新基准提供了任务特定基线模型与端到端联合基线模型。尽管基线模型展示出一定潜力,但该领域仍有大量改进空间。项目网站:https://hirest-cvpr2023.github.io