Continuous diffusion models have demonstrated their effectiveness in addressing the inherent uncertainty and indeterminacy in monocular 3D human pose estimation (HPE). Despite their strengths, the need for large search spaces and the corresponding demand for substantial training data make these models prone to generating biomechanically unrealistic poses. This challenge is particularly noticeable in occlusion scenarios, where the complexity of inferring 3D structures from 2D images intensifies. In response to these limitations, we introduce the Discrete Diffusion Pose ($\text{Di}^2\text{Pose}$), a novel framework designed for occluded 3D HPE that capitalizes on the benefits of a discrete diffusion model. Specifically, $\text{Di}^2\text{Pose}$ employs a two-stage process: it first converts 3D poses into a discrete representation through a \emph{pose quantization step}, which is subsequently modeled in latent space through a \emph{discrete diffusion process}. This methodological innovation restrictively confines the search space towards physically viable configurations and enhances the model's capability to comprehend how occlusions affect human pose within the latent space. Extensive evaluations conducted on various benchmarks (e.g., Human3.6M, 3DPW, and 3DPW-Occ) have demonstrated its effectiveness.
翻译:连续扩散模型已证明其在解决单目三维人体姿态估计中固有的不确定性与模糊性方面的有效性。尽管具备这些优势,但此类模型需要较大的搜索空间及相应的海量训练数据,容易生成生物力学上不真实的姿态。这一挑战在遮挡场景中尤为明显,因为从二维图像推断三维结构的复杂性在此类场景中加剧。针对这些局限性,我们提出了离散扩散姿态估计模型,这是一种专为遮挡三维人体姿态估计设计的新型框架,充分利用了离散扩散模型的优势。具体而言,该模型采用两阶段流程:首先通过姿态量化步骤将三维姿态转换为离散表示,随后通过离散扩散过程在隐空间中对这些表示进行建模。这一方法创新将搜索空间严格限制在物理可行的构型范围内,并增强了模型在隐空间中理解遮挡如何影响人体姿态的能力。在多个基准数据集上进行的广泛评估已验证了其有效性。