Wide-angle videos in few-shot action recognition (FSAR) effectively express actions within specific scenarios. However, without a global understanding of both subjects and background, recognizing actions in such samples remains challenging because of the background distractions. Receptance Weighted Key Value (RWKV), which learns interaction between various dimensions, shows promise for global modeling. While directly applying RWKV to wide-angle FSAR may fail to highlight subjects due to excessive background information. Additionally, temporal relation degraded by frames with similar backgrounds is difficult to reconstruct, further impacting performance. Therefore, we design the CompOund SegmenTation and Temporal REconstructing RWKV (Otter). Specifically, the Compound Segmentation Module~(CSM) is devised to segment and emphasize key patches in each frame, effectively highlighting subjects against background information. The Temporal Reconstruction Module (TRM) is incorporated into the temporal-enhanced prototype construction to enable bidirectional scanning, allowing better reconstruct temporal relation. Furthermore, a regular prototype is combined with the temporal-enhanced prototype to simultaneously enhance subject emphasis and temporal modeling, improving wide-angle FSAR performance. Extensive experiments on benchmarks such as SSv2, Kinetics, UCF101, and HMDB51 demonstrate that Otter achieves state-of-the-art performance. Extra evaluation on the VideoBadminton dataset further validates the superiority of Otter in wide-angle FSAR.
翻译:广角视频在少样本动作识别(FSAR)中能有效表达特定场景下的动作。然而,由于缺乏对主体和背景的全局理解,此类样本中的动作识别因背景干扰而仍具挑战性。接受权重键值(RWKV)通过学习不同维度间的交互,展现出全局建模的潜力。但直接应用RWKV至广角FSAR可能因背景信息过多而无法突出主体。此外,由背景相似帧导致的时序关系退化难以重建,进一步影响性能。为此,我们设计了复合分割与时序重建RWKV(Otter)。具体而言,复合分割模块(CSM)被设计用于分割并突出每帧中的关键区域,有效从背景信息中凸显主体。时序重建模块(TRM)被集成至时序增强原型构建中,实现双向扫描,从而更好地重建时序关系。同时,将常规原型与时序增强原型结合,以同步增强主体聚焦与时序建模,提升广角FSAR性能。在SSv2、Kinetics、UCF101和HMDB51等基准数据集上的广泛实验表明,Otter达到了最先进性能。在VideoBadminton数据集上的额外评估进一步验证了Otter在广角FSAR中的优越性。