StepFormer: Self-supervised Step Discovery and Localization in Instructional Videos

Instructional videos are an important resource to learn procedural tasks from human demonstrations. However, the instruction steps in such videos are typically short and sparse, with most of the video being irrelevant to the procedure. This motivates the need to temporally localize the instruction steps in such videos, i.e. the task called key-step localization. Traditional methods for key-step localization require video-level human annotations and thus do not scale to large datasets. In this work, we tackle the problem with no human supervision and introduce StepFormer, a self-supervised model that discovers and localizes instruction steps in a video. StepFormer is a transformer decoder that attends to the video with learnable queries, and produces a sequence of slots capturing the key-steps in the video. We train our system on a large dataset of instructional videos, using their automatically-generated subtitles as the only source of supervision. In particular, we supervise our system with a sequence of text narrations using an order-aware loss function that filters out irrelevant phrases. We show that our model outperforms all previous unsupervised and weakly-supervised approaches on step detection and localization by a large margin on three challenging benchmarks. Moreover, our model demonstrates an emergent property to solve zero-shot multi-step localization and outperforms all relevant baselines at this task.

翻译：教学视频是从人类示范中学习程序性任务的重要资源。然而，此类视频中的指令步骤通常短且稀疏，大部分视频内容与流程无关。这促使我们需要对这些视频中的指令步骤进行时间定位，即被称为关键步骤定位的任务。传统的关键步骤定位方法需要视频级人工标注，因此无法扩展到大规模数据集。在本文中，我们无需任何人工监督来解决该问题，并提出了StepFormer——一种自监督模型，能够发现并定位视频中的指令步骤。StepFormer是一种Transformer解码器，通过可学习查询关注视频内容，并生成捕获视频中关键步骤的槽位序列。我们在大规模教学视频数据集上训练该系统，仅使用自动生成的字幕作为监督来源。具体而言，我们通过文本叙述序列并结合顺序感知损失函数对系统进行监督，该损失函数可过滤掉无关短语。实验表明，我们的模型在三个具有挑战性的基准测试中，其步骤检测与定位性能大幅超越了所有先前的无监督和弱监督方法。此外，该模型展现出解决零样本多步骤定位的新兴能力，并在该任务上优于所有相关基线方法。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

Into the Metaverse，93页ppt介绍元宇宙概念、应用、趋势

专知会员服务

49+阅读 · 2022年2月19日

最新《自监督表示学习》报告，70页ppt

专知会员服务

86+阅读 · 2020年12月22日

最新《Transformers模型》教程，64页ppt

专知会员服务

326+阅读 · 2020年11月26日

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

专知会员服务

74+阅读 · 2020年8月2日