Determining when people are struggling from video enables a finer-grained understanding of actions and opens opportunities for building intelligent support visual interfaces. In this paper, we present a new dataset with three assembly activities and corresponding performance baselines for the determination of struggle from video. Three real-world problem-solving activities including assembling plumbing pipes (Pipes-Struggle), pitching camping tents (Tent-Struggle) and solving the Tower of Hanoi puzzle (Tower-Struggle) are introduced. Video segments were scored w.r.t. the level of struggle as perceived by annotators using a forced choice 4-point scale. Each video segment was annotated by a single expert annotator in addition to crowd-sourced annotations. The dataset is the first struggle annotation dataset and contains 5.1 hours of video and 725,100 frames from 73 participants in total. We evaluate three decision-making tasks: struggle classification, struggle level regression, and struggle label distribution learning. We provide baseline results for each of the tasks utilising several mainstream deep neural networks, along with an ablation study and visualisation of results. Our work is motivated toward assistive systems that analyze struggle, support users during manual activities and encourage learning, as well as other video understanding competencies.
翻译:从视频中判定人们何时遇到困难,有助于实现更细粒度的动作理解,并为构建智能支持视觉界面开辟了可能。本文提出了一个包含三种装配活动的新数据集,以及相应的视频困难程度判定性能基线基准。我们引入了三项真实世界的问题解决活动,包括组装水管(水管装配困难)、搭建露营帐篷(帐篷搭建困难)以及解决汉诺塔谜题(汉诺塔困难)。视频片段根据标注者感知到的困难程度,采用强迫选择四点量表进行评分。每个视频片段均有一位专家标注员进行标注,同时辅以众包标注。该数据集是首个困难程度标注数据集,包含来自73名参与者的总计5.1小时视频及725,100帧画面。我们评估了三个决策任务:困难分类、困难程度回归以及困难标签分布学习。我们利用多种主流深度神经网络为每个任务提供了基线结果,并进行了消融研究与结果可视化。本研究旨在推动能够分析困难程度、在手动活动中为用户提供支持并鼓励学习,以及提升其他视频理解能力的辅助系统的发展。