Current unsupervised 2D-3D human pose estimation (HPE) methods do not work in multi-person scenarios due to perspective ambiguity in monocular images. Therefore, we present one of the first studies investigating the feasibility of unsupervised multi-person 2D-3D HPE from just 2D poses alone, focusing on reconstructing human interactions. To address the issue of perspective ambiguity, we expand upon prior work by predicting the cameras' elevation angle relative to the subjects' pelvis. This allows us to rotate the predicted poses to be level with the ground plane, while obtaining an estimate for the vertical offset in 3D between individuals. Our method involves independently lifting each subject's 2D pose to 3D, before combining them in a shared 3D coordinate system. The poses are then rotated and offset by the predicted elevation angle before being scaled. This by itself enables us to retrieve an accurate 3D reconstruction of their poses. We present our results on the CHI3D dataset, introducing its use for unsupervised 2D-3D pose estimation with three new quantitative metrics, and establishing a benchmark for future research.
翻译:当前的无监督二维到三维人体姿态估计方法在多人场景中因单目图像的透视歧义而失效。为此,我们首次系统探究了仅从二维姿态出发实现无监督多人三维人体姿态估计的可行性,重点聚焦人体交互重建。为克服透视歧义问题,我们在前人工作基础上引入相机相对于受试者骨盆的仰角预测机制。通过该机制,可将预测姿态旋转至与地平面齐平,同时获取个体间在三维空间中的垂直偏移量估计值。我们采用独立提升各受试者二维姿态至三维空间后,再将其合并至共享三维坐标系的方案。最终,姿态经预测仰角旋转与偏移处理后完成尺度缩放,仅此流程即可实现姿态的精确三维重建。我们在CHI3D数据集上展开实验,首次引入三种新型量化指标用于无监督二维到三维姿态估计任务,并为后续研究建立基准。