Spatio-temporal action detection (STAD) aims to classify the actions present in a video and localize them in space and time. It has become a particularly active area of research in computer vision because of its explosively emerging real-world applications, such as autonomous driving, visual surveillance, entertainment, etc. Many efforts have been devoted in recent years to building a robust and effective framework for STAD. This paper provides a comprehensive review of the state-of-the-art deep learning-based methods for STAD. Firstly, a taxonomy is developed to organize these methods. Next, the linking algorithms, which aim to associate the frame- or clip-level detection results together to form action tubes, are reviewed. Then, the commonly used benchmark datasets and evaluation metrics are introduced, and the performance of state-of-the-art models is compared. At last, this paper is concluded, and a set of potential research directions of STAD are discussed.
翻译:时空动作检测(STAD)旨在识别视频中存在的动作类别,并在空间和时间维度上对其进行定位。由于其在自动驾驶、视觉监控、娱乐等新兴实际应用中的蓬勃发展,该领域已成为计算机视觉中特别活跃的研究方向。近年来,大量研究致力于构建鲁棒且有效的STAD框架。本文系统综述了基于深度学习的最先进STAD方法。首先,建立了对这些方法进行分类的体系结构。其次,对用于将帧级或片段级检测结果关联形成动作管的连接算法进行了梳理。随后,介绍了常用的基准数据集和评估指标,并比较了最先进模型的性能。最后,对全文进行总结,并探讨了STAD的若干潜在研究方向。