Graph convolutional networks have been widely used in skeleton-based action recognition. However, existing approaches are limited in fine-grained action recognition due to the similarity of inter-class data. Moreover, the noisy data from pose extraction increases the challenge of fine-grained recognition. In this work, we propose a flexible attention block called Channel-Variable Spatial-Temporal Attention (CVSTA) to enhance the discriminative power of spatial-temporal joints and obtain a more compact intra-class feature distribution. Based on CVSTA, we construct a Multi-Dimensional Refinement Graph Convolutional Network (MDR-GCN), which can improve the discrimination among channel-, joint- and frame-level features for fine-grained actions. Furthermore, we propose a Robust Decouple Loss (RDL), which significantly boosts the effect of the CVSTA and reduces the impact of noise. The proposed method combining MDR-GCN with RDL outperforms the known state-of-the-art skeleton-based approaches on fine-grained datasets, FineGym99 and FSD-10, and also on the coarse dataset NTU-RGB+D X-view version.
翻译:图卷积网络已广泛用于基于骨骼的动作识别。然而,现有方法在细粒度动作识别中受限于类间数据的相似性。此外,姿态提取产生的噪声数据进一步增加了细粒度识别的挑战。本文提出一种灵活注意力模块——通道可变时空注意力(CVSTA),以增强时空关节的判别能力并获取更紧凑的类内特征分布。基于CVSTA,我们构建了多维细化图卷积网络(MDR-GCN),该网络能够提升通道级、关节级和帧级特征在细粒度动作中的区分度。此外,提出鲁棒解耦损失(RDL),该损失显著增强CVSTA的效果并降低噪声影响。将MDR-GCN与RDL相结合的方法在细粒度数据集FineGym99和FSD-10以及粗粒度数据集NTU-RGB+D X-view版本上均优于已知最先进的基于骨骼的方法。