Manually annotating accurate 3D hand poses is extremely time-consuming and labor-intensive. Existing self-supervised hand pose estimation methods leverage the discrepancy between input images and rendered outputs, or multi-view consistency constraints, as the driving force to optimize networks and progressively refine pose accuracy. However, these methods are highly susceptible to noisy pseudo-labels and overlook the importance of fully exploiting fine-grained spatial correlations, which undermines the stability of model training. To address these issues, we propose UST-Hand, a self-supervised learning framework that estimates uncertainty distribution of hand pose and constructs a probabilistic point cloud feature space, which enables the complex spatiotemporal relationship modeling. UST-Hand employs a conditional normalizing flow model to capture hand pose distributions and samples diverse hypotheses, facilitating robust learning under noisy pseudo-labels supervision with enhanced stability. These multi-hypothesis are mapped to a unified probabilistic 3D point cloud space for multi-view and temporal feature interaction, comprehensively exploring hand motion patterns and fine-grained spatial correlations. Extensive experiments on three challenging datasets demonstrate that UST-Hand achieves state-of-the-art performance, outperforming existing self-supervised methods by up to 37.8% in Mean Per Vertex Position Error (MPVPE).


翻译:手动标注精确的三维手部姿态极为耗时费力。现有自监督手部姿态估计方法利用输入图像与渲染输出之间的差异或多视角一致性约束作为驱动力,优化网络并逐步提升姿态精度。然而,这些方法极易受噪声伪标签影响,且忽视充分挖掘细粒度空间相关性的重要性,从而削弱了模型训练的稳定性。为解决上述问题,我们提出UST-Hand——一种自监督学习框架,通过估计手部姿态的不确定性分布并构建概率点云特征空间,实现复杂的时空关系建模。UST-Hand采用条件归一化流模型捕获手部姿态分布并采样多样假设,从而在噪声伪标签监督下实现鲁棒学习并增强稳定性。这些多假设被映射至统一的概率三维点云空间,用于多视角与时序特征交互,全面探索手部运动模式及细粒度空间相关性。在三个具有挑战性的数据集上的大量实验表明,UST-Hand取得了最先进性能,在平均每顶点位置误差(MPVPE)指标上较现有自监督方法提升高达37.8%。

0
下载
关闭预览

相关内容

根据激光测量原理得到的点云,包括三维坐标(XYZ)和激光反射强度(Intensity)。 根据摄影测量原理得到的点云,包括三维坐标(XYZ)和颜色信息(RGB)。 结合激光测量和摄影测量原理得到点云,包括三维坐标(XYZ)、激光反射强度(Intensity)和颜色信息(RGB)。 在获取物体表面每个采样点的空间坐标后,得到的是一个点的集合,称之为“点云”(Point Cloud)
基于深度学习的物体姿态估计综述
专知会员服务
27+阅读 · 2024年5月15日
专知会员服务
34+阅读 · 2021年10月11日
专知会员服务
65+阅读 · 2021年4月11日
最新《深度学习人体姿态估计》综述论文,26页pdf
专知会员服务
41+阅读 · 2020年12月29日
计算机视觉方向简介 | 人体姿态估计
计算机视觉life
28+阅读 · 2019年6月6日
深度学习人体姿态估计算法综述
AI前线
25+阅读 · 2019年5月19日
SkeletonNet:完整的人体三维位姿重建方法
计算机视觉life
21+阅读 · 2019年1月21日
【前沿】凌空手势识别综述
科技导报
12+阅读 · 2017年8月17日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
最新内容
《多域冲突比较支持模型》60页
专知会员服务
3+阅读 · 今天4:35
面向2027年及未来的海军情报改革
专知会员服务
3+阅读 · 8月5日
相关VIP内容
基于深度学习的物体姿态估计综述
专知会员服务
27+阅读 · 2024年5月15日
专知会员服务
34+阅读 · 2021年10月11日
专知会员服务
65+阅读 · 2021年4月11日
最新《深度学习人体姿态估计》综述论文,26页pdf
专知会员服务
41+阅读 · 2020年12月29日
相关基金
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员