Human action recognition aims at classifying the category of human action from a segment of a video. Recently, people dive into designing GCN-based models to extract features from skeletons for performing this task, because skeleton representations are much efficient and robust than other modalities such as RGB frames. However, when employing the skeleton data, some important clues like related items are also dismissed. It results in some ambiguous actions that are hard to be distinguished and tend to be misclassified. To alleviate this problem, we propose an auxiliary feature refinement head (FR Head), which consists of spatial-temporal decoupling and contrastive feature refinement, to obtain discriminative representations of skeletons. Ambiguous samples are dynamically discovered and calibrated in the feature space. Furthermore, FR Head could be imposed on different stages of GCNs to build a multi-level refinement for stronger supervision. Extensive experiments are conducted on NTU RGB+D, NTU RGB+D 120, and NW-UCLA datasets. Our proposed models obtain competitive results from state-of-the-art methods and can help to discriminate those ambiguous samples.
翻译:人体动作识别旨在从视频片段中对人体动作类别进行分类。近年来,由于骨架表示相比RGB帧等其他模态更高效、鲁棒,研究者深入设计基于GCN的模型来提取骨架特征以完成该任务。然而,在使用骨架数据时,相关物体等重要线索也被丢弃,导致一些模糊动作难以区分且易于误分类。为解决该问题,我们提出一种辅助特征精炼头(FR Head),它由时空解耦与对比特征精炼组成,用于获得骨架的判别表示。在特征空间中对模糊样本进行动态发现与校正。此外,FR Head可施加于GCN的不同阶段,构建多层级精炼以提供更强的监督。我们在NTU RGB+D、NTU RGB+D 120和NW-UCLA数据集上进行了大量实验。我们提出的模型取得了与最先进方法相竞争的结果,并能有效区分那些模糊样本。