Semi-supervised action recognition aims to improve spatio-temporal reasoning ability with a few labeled data in conjunction with a large amount of unlabeled data. Albeit recent advancements, existing powerful methods are still prone to making ambiguous predictions under scarce labeled data, embodied as the limitation of distinguishing different actions with similar spatio-temporal information. In this paper, we approach this problem by empowering the model two aspects of capability, namely discriminative spatial modeling and temporal structure modeling for learning discriminative spatio-temporal representations. Specifically, we propose an Adaptive Contrastive Learning~(ACL) strategy. It assesses the confidence of all unlabeled samples by the class prototypes of the labeled data, and adaptively selects positive-negative samples from a pseudo-labeled sample bank to construct contrastive learning. Additionally, we introduce a Multi-scale Temporal Learning~(MTL) strategy. It could highlight informative semantics from long-term clips and integrate them into the short-term clip while suppressing noisy information. Afterwards, both of these two new techniques are integrated in a unified framework to encourage the model to make accurate predictions. Extensive experiments on UCF101, HMDB51 and Kinetics400 show the superiority of our method over prior state-of-the-art approaches.
翻译:半监督动作识别旨在利用少量标注数据结合大量未标注数据,提升时空推理能力。尽管近期取得进展,现有强大方法在标注数据稀缺时仍易产生模糊预测,具体表现为难以区分时空信息相似的不同动作。本文通过增强模型两方面的能力来解决该问题,即判别性空间建模和时间结构建模,以学习判别性时空表示。具体而言,我们提出了一种自适应对比学习(Adaptive Contrastive Learning, ACL)策略,该策略利用标注数据的类别原型评估所有未标注样本的置信度,并从伪标签样本库中自适应选取正负样本以构建对比学习。此外,我们引入多尺度时间学习(Multi-scale Temporal Learning, MTL)策略,该策略能从长片段中突出信息性语义,并将其整合到短视频片段中,同时抑制噪声信息。随后,这两种新技术被集成到一个统一框架中,以促使模型做出准确预测。在UCF101、HMDB51和Kinetics400上的大量实验表明,我们的方法优于以往最先进的方法。