Due to the scarcity of manually annotated data required for fine-grained video understanding, few-shot fine-grained (FS-FG) action recognition has gained significant attention, with the aim of classifying novel fine-grained action categories with only a few labeled instances. Despite the progress made in FS coarse-grained action recognition, current approaches encounter two challenges when dealing with the fine-grained action categories: the inability to capture subtle action details and the insufficiency of learning from limited data that exhibit high intra-class variance and inter-class similarity. To address these limitations, we propose M$^3$Net, a matching-based framework for FS-FG action recognition, which incorporates \textit{multi-view encoding}, \textit{multi-view matching}, and \textit{multi-view fusion} to facilitate embedding encoding, similarity matching, and decision making across multiple viewpoints. \textit{Multi-view encoding} captures rich contextual details from the intra-frame, intra-video, and intra-episode perspectives, generating customized higher-order embeddings for fine-grained data. \textit{Multi-view matching} integrates various matching functions enabling flexible relation modeling within limited samples to handle multi-scale spatio-temporal variations by leveraging the instance-specific, category-specific, and task-specific perspectives. \textit{Multi-view fusion} consists of matching-predictions fusion and matching-losses fusion over the above views, where the former promotes mutual complementarity and the latter enhances embedding generalizability by employing multi-task collaborative learning. Explainable visualizations and experimental results on three challenging benchmarks demonstrate the superiority of M$^3$Net in capturing fine-grained action details and achieving state-of-the-art performance for FS-FG action recognition.
翻译:由于细粒度视频理解所需的人工标注数据稀缺,少样本细粒度动作识别(FS-FG)受到广泛关注,其目标是在仅有少量标注样本的条件下对新颖的细粒度动作类别进行分类。尽管少样本粗粒度动作识别已取得进展,但现有方法在处理细粒度动作类别时仍面临两大挑战:无法捕捉细微动作细节,以及在类内方差大、类间相似度高的有限数据中学习不足。为解决上述问题,我们提出M$^3$Net——一种基于匹配的FS-FG动作识别框架,该框架融合了多视角编码、多视角匹配与多视角融合,以促进来自多视角的嵌入编码、相似度匹配及决策制定。多视角编码从帧内、视频内及情节内视角捕捉丰富的上下文细节,为细粒度数据生成定制化的高阶嵌入。多视角匹配整合多种匹配函数,通过利用实例特定、类别特定及任务特定的视角,在有限样本内实现灵活的关系建模以应对多尺度时空变化。多视角融合包含上述视角上的匹配-预测融合与匹配-损失融合,前者促进相互互补性,后者通过多任务协同学习增强嵌入泛化能力。在三个具有挑战性的基准数据集上的可解释可视化结果与实验表明,M$^3$Net在捕捉细粒度动作细节方面具有优越性,并在FS-FG动作识别中达到了最先进性能。