Learning-based methods have dominated the 3D human pose estimation (HPE) tasks with significantly better performance in most benchmarks than traditional optimization-based methods. Nonetheless, 3D HPE in the wild is still the biggest challenge of learning-based models, whether with 2D-3D lifting, image-to-3D, or diffusion-based methods, since the trained networks implicitly learn camera intrinsic parameters and domain-based 3D human pose distributions and estimate poses by statistical average. On the other hand, the optimization-based methods estimate results case-by-case, which can predict more diverse and sophisticated human poses in the wild. By combining the advantages of optimization-based and learning-based methods, we propose the Zero-shot Diffusion-based Optimization (ZeDO) pipeline for 3D HPE to solve the problem of cross-domain and in-the-wild 3D HPE. Our multi-hypothesis ZeDO achieves state-of-the-art (SOTA) performance on Human3.6M as minMPJPE $51.4$mm without training with any 2D-3D or image-3D pairs. Moreover, our single-hypothesis ZeDO achieves SOTA performance on 3DPW dataset with PA-MPJPE $42.6$mm on cross-dataset evaluation, which even outperforms learning-based methods trained on 3DPW.
翻译:学习型方法在三维人体姿态估计任务中占据主导地位,在大多数基准测试中表现显著优于传统优化型方法。然而,野外三维人体姿态估计仍是学习型模型面临的最大挑战——无论是通过2D-3D提升、图像到3D映射,还是扩散型方法——因为训练后的网络隐式学习了相机内参和基于领域的三维人体姿态分布,并通过统计平均来估计姿态。相反,优化型方法逐案例估计结果,能够预测野外更多样且复杂的人体姿态。通过结合优化型与学习型方法的优势,我们提出零样本扩散优化管道用于三维人体姿态估计,以解决跨域和野外三维人体姿态估计问题。我们的多假设ZeDO在Human3.6M数据集上以最小平均每关节位置误差51.4毫米达到当前最优性能,且无需使用任何2D-3D或图像-3D配对进行训练。此外,我们的单假设ZeDO在3DPW数据集上通过跨数据集评估实现PA-MPJPE 42.6毫米的当前最优性能,甚至超越了在3DPW上训练的学习型方法。