Recently, video-based action recognition methods using convolutional neural networks (CNNs) achieve remarkable recognition performance. However, there is still lack of understanding about the generalization mechanism of action recognition models. In this paper, we suggest that action recognition models rely on the motion information less than expected, and thus they are robust to randomization of frame orders. Furthermore, we find that motion monotonicity remaining after randomization also contributes to such robustness. Based on this observation, we develop a novel defense method using temporal shuffling of input videos against adversarial attacks for action recognition models. Another observation enabling our defense method is that adversarial perturbations on videos are sensitive to temporal destruction. To the best of our knowledge, this is the first attempt to design a defense method without additional training for 3D CNN-based video action recognition models.
翻译:最近,基于卷积神经网络(CNN)的视频动作识别方法取得了显著的识别性能。然而,目前对动作识别模型泛化机制的理解仍存在不足。本文提出,动作识别模型对运动信息的依赖程度低于预期,因此对帧顺序的随机化具有鲁棒性。此外,我们发现随机化后保留的运动单调性也促进了这种鲁棒性。基于这一观察,我们开发了一种新颖的防御方法,通过输入视频的时间混洗来应对针对动作识别模型的对抗攻击。另一个支持我们防御方法的观察是,视频上的对抗扰动对时间破坏敏感。据我们所知,这是首次尝试设计一种无需额外训练的、基于3D CNN的视频动作识别模型防御方法。