Video segmentation encompasses a wide range of categories of problem formulation, e.g., object, scene, actor-action and multimodal video segmentation, for delineating task-specific scene components with pixel-level masks. Recently, approaches in this research area shifted from concentrating on ConvNet-based to transformer-based models. In addition, various interpretability approaches have appeared for transformer models and video temporal dynamics, motivated by the growing interest in basic scientific understanding, model diagnostics and societal implications of real-world deployment. Previous surveys mainly focused on ConvNet models on a subset of video segmentation tasks or transformers for classification tasks. Moreover, component-wise discussion of transformer-based video segmentation models has not yet received due focus. In addition, previous reviews of interpretability methods focused on transformers for classification, while analysis of video temporal dynamics modelling capabilities of video models received less attention. In this survey, we address the above with a thorough discussion of various categories of video segmentation, a component-wise discussion of the state-of-the-art transformer-based models, and a review of related interpretability methods. We first present an introduction to the different video segmentation task categories, their objectives, specific challenges and benchmark datasets. Next, we provide a component-wise review of recent transformer-based models and document the state of the art on different video segmentation tasks. Subsequently, we discuss post-hoc and ante-hoc interpretability methods for transformer models and interpretability methods for understanding the role of the temporal dimension in video models. Finally, we conclude our discussion with future research directions.
翻译:视频分割涵盖广泛的问题形式化类别,例如物体、场景、演员-动作以及多模态视频分割,旨在通过像素级掩膜描绘任务特定的场景组件。近年来,该研究领域的方法从基于卷积神经网络模型转向基于Transformer模型。此外,受基础科学理解、模型诊断以及现实世界部署的社会影响等日益增长的兴趣驱动,针对Transformer模型和视频时间动态性的各种可解释性方法已经出现。以往的综述主要聚焦于卷积神经网络模型在部分视频分割任务上的应用,或Transformer在分类任务上的应用。此外,针对基于Transformer的视频分割模型的组件级讨论尚未得到应有的关注。同时,先前关于可解释性方法的综述侧重于用于分类的Transformer,而对视频模型的时间动态建模能力的分析关注较少。在本综述中,我们通过深入讨论各类视频分割任务、基于最新Transformer模型的组件级分析以及相关可解释性方法的回顾,来应对上述问题。我们首先介绍不同视频分割任务类别、其目标、特定挑战和基准数据集。接着,我们提供基于Transformer的最新模型的组件级回顾,并记录在不同视频分割任务上的最新进展。随后,我们讨论Transformer模型的事后和事前可解释性方法,以及用于理解视频模型中时间维度作用的可解释性方法。最后,我们以未来研究方向作为总结。