Manually annotating accurate 3D hand poses is extremely time-consuming and labor-intensive. Existing self-supervised hand pose estimation methods leverage the discrepancy between input images and rendered outputs, or multi-view consistency constraints, as the driving force to optimize networks and progressively refine pose accuracy. However, these methods are highly susceptible to noisy pseudo-labels and overlook the importance of fully exploiting fine-grained spatial correlations, which undermines the stability of model training. To address these issues, we propose UST-Hand, a self-supervised learning framework that estimates uncertainty distribution of hand pose and constructs a probabilistic point cloud feature space, which enables the complex spatiotemporal relationship modeling. UST-Hand employs a conditional normalizing flow model to capture hand pose distributions and samples diverse hypotheses, facilitating robust learning under noisy pseudo-labels supervision with enhanced stability. These multi-hypothesis are mapped to a unified probabilistic 3D point cloud space for multi-view and temporal feature interaction, comprehensively exploring hand motion patterns and fine-grained spatial correlations. Extensive experiments on three challenging datasets demonstrate that UST-Hand achieves state-of-the-art performance, outperforming existing self-supervised methods by up to 37.8% in Mean Per Vertex Position Error (MPVPE).
翻译:手动标注精确的三维手部姿态极为耗时费力。现有自监督手部姿态估计方法利用输入图像与渲染输出之间的差异或多视角一致性约束作为驱动力,优化网络并逐步提升姿态精度。然而,这些方法极易受噪声伪标签影响,且忽视充分挖掘细粒度空间相关性的重要性,从而削弱了模型训练的稳定性。为解决上述问题,我们提出UST-Hand——一种自监督学习框架,通过估计手部姿态的不确定性分布并构建概率点云特征空间,实现复杂的时空关系建模。UST-Hand采用条件归一化流模型捕获手部姿态分布并采样多样假设,从而在噪声伪标签监督下实现鲁棒学习并增强稳定性。这些多假设被映射至统一的概率三维点云空间,用于多视角与时序特征交互,全面探索手部运动模式及细粒度空间相关性。在三个具有挑战性的数据集上的大量实验表明,UST-Hand取得了最先进性能,在平均每顶点位置误差(MPVPE)指标上较现有自监督方法提升高达37.8%。