基于多尺度特征融合的骨架片段对比学习用于动作定位 (Skeleton-Snippet Contrastive Learning with Multiscale Feature Fusion for Action Localization) - 专知论文

会员服务 ·

0

骨架 · 片段 · 对比学习 · 融合 · 动作识别 ·

2025 年 12 月 22 日

Skeleton-Snippet Contrastive Learning with Multiscale Feature Fusion for Action Localization

翻译：基于多尺度特征融合的骨架片段对比学习用于动作定位

Qiushuo Cheng,Jingjing Liu,Catherine Morgan,Alan Whone,Majid Mirmehdi

The self-supervised pretraining paradigm has achieved great success in learning 3D action representations for skeleton-based action recognition using contrastive learning. However, learning effective representations for skeleton-based temporal action localization remains challenging and underexplored. Unlike video-level {action} recognition, detecting action boundaries requires temporally sensitive features that capture subtle differences between adjacent frames where labels change. To this end, we formulate a snippet discrimination pretext task for self-supervised pretraining, which densely projects skeleton sequences into non-overlapping segments and promotes features that distinguish them across videos via contrastive learning. Additionally, we build on strong backbones of skeleton-based action recognition models by fusing intermediate features with a U-shaped module to enhance feature resolution for frame-level localization. Our approach consistently improves existing skeleton-based contrastive learning methods for action localization on BABEL across diverse subsets and evaluation protocols. We also achieve state-of-the-art transfer learning performance on PKUMMD with pretraining on NTU RGB+D and BABEL.

翻译：自监督预训练范式通过对比学习，在基于骨架的动作识别领域学习三维动作表示方面取得了巨大成功。然而，针对基于骨架的时序动作定位学习有效表示仍然具有挑战性且研究不足。与视频级别的动作识别不同，检测动作边界需要具有时间敏感性的特征，以捕捉标签发生变化的相邻帧之间的细微差异。为此，我们构建了一个用于自监督预训练的片段判别前置任务，该任务将骨架序列密集地投影到非重叠的片段中，并通过对比学习增强能够区分不同视频间这些片段的特征。此外，我们在基于骨架的动作识别模型的强大骨干网络基础上，通过一个U形模块融合中间特征，以增强用于帧级定位的特征分辨率。我们的方法在BABEL数据集的各种子集和评估协议上，持续改进了现有基于骨架的对比学习方法在动作定位任务上的性能。通过在NTU RGB+D和BABEL上进行预训练，我们在PKUMMD数据集上也实现了最先进的迁移学习性能。

0

相关内容

【NeurIPS2024】SAFE: 慢速与快速参数高效调优用于基于预训练模型的持续学习

【NeurIPS2024】SAFE: 慢速与快速参数高效调优用于基于预训练模型的持续学习

专知会员服务

18+阅读 · 2024年11月5日

【CVPR2024】DiffusionMTL: 从部分标注数据学习多任务去噪扩散模型

【CVPR2024】DiffusionMTL: 从部分标注数据学习多任务去噪扩散模型

专知会员服务

34+阅读 · 2024年3月25日

【NeurIPS2023】半监督端到端对比学习用于时间序列分类

【NeurIPS2023】半监督端到端对比学习用于时间序列分类

专知会员服务

36+阅读 · 2023年10月17日

【ICML2022】DRIBO:基于多视图信息瓶颈的鲁棒深度强化学习

【ICML2022】DRIBO:基于多视图信息瓶颈的鲁棒深度强化学习

专知会员服务

17+阅读 · 2022年8月13日

强化学习在机器人中的应用，附视频与Slides，Animesh Garg, UoT

强化学习在机器人中的应用，附视频与Slides，Animesh Garg, UoT

专知会员服务

37+阅读 · 2022年7月12日

【超越消息传递:图神经网络的物理启发范式】Beyond Message Passing: a Physics-Inspired Paradigm for Graph Neural Networks

【超越消息传递:图神经网络的物理启发范式】Beyond Message Passing: a Physics-Inspired Paradigm for Graph Neural Networks

专知会员服务

17+阅读 · 2022年5月10日

【CVPR 2022】基于双噪声标签的可见光-红外人再识别学习，Learning with Twin Noisy Labels for Visible-Infrared Person Re-Identification

【CVPR 2022】基于双噪声标签的可见光-红外人再识别学习，Learning with Twin Noisy Labels for Visible-Infrared Person Re-Identification

专知会员服务

14+阅读 · 2022年3月28日

【CVPR 2022】基于实例深度估计的统一深度感知全景分割 PanopticDepth: Per-Instance Depth Estimation for Unified Depth-Aware Panoptic Segmentation

【CVPR 2022】基于实例深度估计的统一深度感知全景分割 PanopticDepth: Per-Instance Depth Estimation for Unified Depth-Aware Panoptic Segmentation

专知会员服务

18+阅读 · 2022年3月19日

【ICCV2021】基于对比视频表示学习的长短视图特征分解

专知会员服务

10+阅读 · 2021年10月6日

【ICML2021】图对比学习自动化

专知会员服务

41+阅读 · 2021年6月19日

深度学习图像检索(CBIR): 十年之大综述

深度学习图像检索(CBIR): 十年之大综述

专知

66+阅读 · 2020年12月5日

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

专知

11+阅读 · 2020年8月28日

【ACMMM2020-北航】KBGN:用于视觉对话中自适应视觉-文本推理的知识桥图网络

【ACMMM2020-北航】KBGN:用于视觉对话中自适应视觉-文本推理的知识桥图网络

专知

10+阅读 · 2020年8月12日

Python图像处理，366页pdf，Image Operators Image Processing in Python

Python图像处理，366页pdf，Image Operators Image Processing in Python

专知

15+阅读 · 2020年7月23日

【CVPR 2020 Oral】小样本类增量学习

【CVPR 2020 Oral】小样本类增量学习

专知

20+阅读 · 2020年6月26日

【CVPR2020-旷视】DPGN：分布传播图网络的小样本学习

【CVPR2020-旷视】DPGN：分布传播图网络的小样本学习

专知

13+阅读 · 2020年4月1日

【阿里巴巴-WWW2020】对抗性多模态表示学习的点击率预测，Adversarial Multimodal RL

【阿里巴巴-WWW2020】对抗性多模态表示学习的点击率预测，Adversarial Multimodal RL

专知

11+阅读 · 2020年3月17日

如何用机器学习精准辨别“背景”和“目标”

如何用机器学习精准辨别“背景”和“目标”

论智

10+阅读 · 2018年10月22日

资源 | GitHub新项目：轻松使用多种预训练卷积网络抽取图像特征

资源 | GitHub新项目：轻松使用多种预训练卷积网络抽取图像特征

机器之心

12+阅读 · 2018年4月16日

语义分割中的深度学习方法全解：从FCN、SegNet到DeepLab

语义分割中的深度学习方法全解：从FCN、SegNet到DeepLab

炼数成金订阅号

26+阅读 · 2017年7月10日

间接优化的高效Monte Carlo声传播研究

国家自然科学基金

0+阅读 · 2017年12月31日

基于DASH的交互式三维视频系统建模

国家自然科学基金

1+阅读 · 2015年12月31日

视觉识别中的实用鲁棒回归技术研究

国家自然科学基金

3+阅读 · 2015年12月31日

基于MEMS加速度传感器的智能终端手势识别及三维交互模型

国家自然科学基金

6+阅读 · 2015年12月31日

基于随机有限集理论的复杂背景视频多目标跟踪研究

国家自然科学基金

2+阅读 · 2015年12月31日

不确定知识图谱中面向结构查询的众包清洗研究

国家自然科学基金

4+阅读 · 2015年12月31日

基于支撑函数的不规则形态扩展目标建模和估计研究

国家自然科学基金

0+阅读 · 2015年12月31日

基于上下文感知和异质特征集成的SAR图像分割与评价

国家自然科学基金

2+阅读 · 2015年12月31日

自由视点三维视频中纹理-深度图像联合建模及应用

国家自然科学基金

0+阅读 · 2015年12月31日

基于字典学习的小样本高光谱遥感图像稀疏表示分类精度研究与应用

国家自然科学基金

3+阅读 · 2014年12月31日

LocationAgent: A Hierarchical Agent for Image Geolocation via Decoupling Strategy and Evidence from Parametric Knowledge

Arxiv

0+阅读 · 1月27日

Efficient Rehearsal for Continual Learning in ASR via Singular Value Tuning

Arxiv

0+阅读 · 1月26日

Learning Domain Knowledge in Multimodal Large Language Models through Reinforcement Fine-Tuning

Arxiv

0+阅读 · 1月23日

BayesianVLA: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries

Arxiv

0+阅读 · 1月22日

BayesianVLA: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries

Arxiv

0+阅读 · 1月21日

Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey

Arxiv

0+阅读 · 1月19日

Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval

Arxiv

0+阅读 · 1月19日

Language-Based Swarm Perception: Decentralized Person Re-Identification via Natural Language Descriptions

Arxiv

0+阅读 · 1月18日

Video Joint-Embedding Predictive Architectures for Facial Expression Recognition

Arxiv

0+阅读 · 1月14日

Safe Heterogeneous Multi-Agent RL with Communication Regularization for Coordinated Target Acquisition

Arxiv

0+阅读 · 1月13日

VIP会员

文章信息

相关主题

相关VIP内容

【NeurIPS2024】SAFE: 慢速与快速参数高效调优用于基于预训练模型的持续学习

【NeurIPS2024】SAFE: 慢速与快速参数高效调优用于基于预训练模型的持续学习

专知会员服务

18+阅读 · 2024年11月5日

【CVPR2024】DiffusionMTL: 从部分标注数据学习多任务去噪扩散模型

【CVPR2024】DiffusionMTL: 从部分标注数据学习多任务去噪扩散模型

专知会员服务

34+阅读 · 2024年3月25日

【NeurIPS2023】半监督端到端对比学习用于时间序列分类

【NeurIPS2023】半监督端到端对比学习用于时间序列分类

专知会员服务

36+阅读 · 2023年10月17日

【ICML2022】DRIBO:基于多视图信息瓶颈的鲁棒深度强化学习

【ICML2022】DRIBO:基于多视图信息瓶颈的鲁棒深度强化学习

专知会员服务

17+阅读 · 2022年8月13日

强化学习在机器人中的应用，附视频与Slides，Animesh Garg, UoT

强化学习在机器人中的应用，附视频与Slides，Animesh Garg, UoT

专知会员服务

37+阅读 · 2022年7月12日

【超越消息传递:图神经网络的物理启发范式】Beyond Message Passing: a Physics-Inspired Paradigm for Graph Neural Networks

【超越消息传递:图神经网络的物理启发范式】Beyond Message Passing: a Physics-Inspired Paradigm for Graph Neural Networks

专知会员服务

17+阅读 · 2022年5月10日

【CVPR 2022】基于双噪声标签的可见光-红外人再识别学习，Learning with Twin Noisy Labels for Visible-Infrared Person Re-Identification

【CVPR 2022】基于双噪声标签的可见光-红外人再识别学习，Learning with Twin Noisy Labels for Visible-Infrared Person Re-Identification

专知会员服务

14+阅读 · 2022年3月28日

【CVPR 2022】基于实例深度估计的统一深度感知全景分割 PanopticDepth: Per-Instance Depth Estimation for Unified Depth-Aware Panoptic Segmentation

【CVPR 2022】基于实例深度估计的统一深度感知全景分割 PanopticDepth: Per-Instance Depth Estimation for Unified Depth-Aware Panoptic Segmentation

专知会员服务

18+阅读 · 2022年3月19日

【ICCV2021】基于对比视频表示学习的长短视图特征分解

专知会员服务

10+阅读 · 2021年10月6日

【ICML2021】图对比学习自动化

专知会员服务

41+阅读 · 2021年6月19日

热门VIP内容

开通专知VIP会员享更多权益服务

《无人机与战争：被忽视的环境影响及无人机保护潜力》

俄罗斯规划未来无人机驱动军队

《整合杀伤链：一个用于边缘目标验证与战术推理的零样本框架》最新资料

《人工智能、武器与影响力：前沿模型在模拟核危机中展现复杂推理》2026最新46页报告

相关资讯

深度学习图像检索(CBIR): 十年之大综述

深度学习图像检索(CBIR): 十年之大综述

专知

66+阅读 · 2020年12月5日

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

【KDD2020-Tutorial】因果推理与稳定学习，Causal Inference and Stable Learning

专知

11+阅读 · 2020年8月28日

【ACMMM2020-北航】KBGN:用于视觉对话中自适应视觉-文本推理的知识桥图网络

【ACMMM2020-北航】KBGN:用于视觉对话中自适应视觉-文本推理的知识桥图网络

专知

10+阅读 · 2020年8月12日

Python图像处理，366页pdf，Image Operators Image Processing in Python

Python图像处理，366页pdf，Image Operators Image Processing in Python

专知

15+阅读 · 2020年7月23日

【CVPR 2020 Oral】小样本类增量学习

【CVPR 2020 Oral】小样本类增量学习

专知

20+阅读 · 2020年6月26日

【CVPR2020-旷视】DPGN：分布传播图网络的小样本学习

【CVPR2020-旷视】DPGN：分布传播图网络的小样本学习

专知

13+阅读 · 2020年4月1日

【阿里巴巴-WWW2020】对抗性多模态表示学习的点击率预测，Adversarial Multimodal RL

【阿里巴巴-WWW2020】对抗性多模态表示学习的点击率预测，Adversarial Multimodal RL

专知

11+阅读 · 2020年3月17日

如何用机器学习精准辨别“背景”和“目标”

如何用机器学习精准辨别“背景”和“目标”

论智

10+阅读 · 2018年10月22日

资源 | GitHub新项目：轻松使用多种预训练卷积网络抽取图像特征

资源 | GitHub新项目：轻松使用多种预训练卷积网络抽取图像特征

机器之心

12+阅读 · 2018年4月16日

语义分割中的深度学习方法全解：从FCN、SegNet到DeepLab

语义分割中的深度学习方法全解：从FCN、SegNet到DeepLab

炼数成金订阅号

26+阅读 · 2017年7月10日

相关论文

LocationAgent: A Hierarchical Agent for Image Geolocation via Decoupling Strategy and Evidence from Parametric Knowledge

Arxiv

0+阅读 · 1月27日

Efficient Rehearsal for Continual Learning in ASR via Singular Value Tuning

Arxiv

0+阅读 · 1月26日

Learning Domain Knowledge in Multimodal Large Language Models through Reinforcement Fine-Tuning

Arxiv

0+阅读 · 1月23日

BayesianVLA: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries

Arxiv

0+阅读 · 1月22日

BayesianVLA: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries

Arxiv

0+阅读 · 1月21日

Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey

Arxiv

0+阅读 · 1月19日

Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval

Arxiv

0+阅读 · 1月19日

Language-Based Swarm Perception: Decentralized Person Re-Identification via Natural Language Descriptions

Arxiv

0+阅读 · 1月18日

Video Joint-Embedding Predictive Architectures for Facial Expression Recognition

Arxiv

0+阅读 · 1月14日

Safe Heterogeneous Multi-Agent RL with Communication Regularization for Coordinated Target Acquisition

Arxiv

0+阅读 · 1月13日

相关基金

间接优化的高效Monte Carlo声传播研究

国家自然科学基金

0+阅读 · 2017年12月31日

基于DASH的交互式三维视频系统建模

国家自然科学基金

1+阅读 · 2015年12月31日

视觉识别中的实用鲁棒回归技术研究

国家自然科学基金

3+阅读 · 2015年12月31日

基于MEMS加速度传感器的智能终端手势识别及三维交互模型

国家自然科学基金

6+阅读 · 2015年12月31日

基于随机有限集理论的复杂背景视频多目标跟踪研究

国家自然科学基金

2+阅读 · 2015年12月31日

不确定知识图谱中面向结构查询的众包清洗研究

国家自然科学基金

4+阅读 · 2015年12月31日

基于支撑函数的不规则形态扩展目标建模和估计研究

国家自然科学基金

0+阅读 · 2015年12月31日

基于上下文感知和异质特征集成的SAR图像分割与评价

国家自然科学基金

2+阅读 · 2015年12月31日

自由视点三维视频中纹理-深度图像联合建模及应用

国家自然科学基金

0+阅读 · 2015年12月31日

基于字典学习的小样本高光谱遥感图像稀疏表示分类精度研究与应用

国家自然科学基金

3+阅读 · 2014年12月31日

微信扫码咨询专知VIP会员