We present Ego3DPose, a highly accurate binocular egocentric 3D pose reconstruction system. The binocular egocentric setup offers practicality and usefulness in various applications, however, it remains largely under-explored. It has been suffering from low pose estimation accuracy due to viewing distortion, severe self-occlusion, and limited field-of-view of the joints in egocentric 2D images. Here, we notice that two important 3D cues, stereo correspondences, and perspective, contained in the egocentric binocular input are neglected. Current methods heavily rely on 2D image features, implicitly learning 3D information, which introduces biases towards commonly observed motions and leads to low overall accuracy. We observe that they not only fail in challenging occlusion cases but also in estimating visible joint positions. To address these challenges, we propose two novel approaches. First, we design a two-path network architecture with a path that estimates pose per limb independently with its binocular heatmaps. Without full-body information provided, it alleviates bias toward trained full-body distribution. Second, we leverage the egocentric view of body limbs, which exhibits strong perspective variance (e.g., a significantly large-size hand when it is close to the camera). We propose a new perspective-aware representation using trigonometry, enabling the network to estimate the 3D orientation of limbs. Finally, we develop an end-to-end pose reconstruction network that synergizes both techniques. Our comprehensive evaluations demonstrate that Ego3DPose outperforms state-of-the-art models by a pose estimation error (i.e., MPJPE) reduction of 23.1% in the UnrealEgo dataset. Our qualitative results highlight the superiority of our approach across a range of scenarios and challenges.
翻译:我们提出了Ego3DPose,一种高精度的双目自我中心3D姿态重建系统。双目自我中心设置在实际应用中兼具实用性与有效性,但仍未得到充分探索。由于观看畸变、严重自遮挡以及自我中心2D图像中关节视野受限,该系统长期面临姿态估计精度低的问题。在此,我们注意到自我中心双目输入中包含的两类重要3D线索——立体对应关系与透视——被忽略了。当前方法过度依赖2D图像特征,隐式学习3D信息,这引入了对常见运动模式的偏见,导致整体精度降低。我们观察到这些方法不仅在具有挑战性的遮挡情况下表现不佳,甚至在估计可见关节位置时也存在缺陷。为解决这些挑战,我们提出两种创新方法。首先,我们设计了一种双路径网络架构,其中一条路径通过双目热力图独立估计每个肢体的姿态。由于未提供全身信息,该方法缓解了对训练过的全身分布的偏见。其次,我们利用身体肢体的自我中心视角,该视角表现出强烈的透视差异(例如,当手部靠近相机时呈现显著的大尺寸)。我们提出一种基于三角学的新的透视感知表示,使网络能够估计肢体的3D方向。最后,我们开发了一个端到端的姿态重建网络,协同整合这两种技术。全面评估表明,Ego3DPose在UnrealEgo数据集上以23.1%的姿态估计误差(即MPJPE)降低超越了现有最优模型。我们的定性结果凸显了该方法在多种场景与挑战下的优越性。