Full-body ego-pose estimation from head and hand poses alone has become an active area of research to power articulate avatar representation on headset-based platforms. However, existing methods over-rely on the confines of the motion-capture spaces in which datasets were recorded, while simultaneously assuming continuous capture of joint motions and uniform body dimensions. In this paper, we propose EgoPoser, which overcomes these limitations by 1) rethinking the input representation for headset-based ego-pose estimation and introducing a novel motion decomposition method that predicts full-body pose independent of global positions, 2) robustly modeling body pose from intermittent hand position and orientation tracking only when inside a headset's field of view, and 3) generalizing across various body sizes for different users. Our experiments show that EgoPoser outperforms state-of-the-art methods both qualitatively and quantitatively, while maintaining a high inference speed of over 600 fps. EgoPoser establishes a robust baseline for future work, where full-body pose estimation needs no longer rely on outside-in capture and can scale to large-scene environments.
翻译:仅通过头部和手部姿态进行全身自我姿态估计,已成为一项活跃的研究领域,旨在为头戴式平台上的具身化身表示提供支持。然而,现有方法过度依赖数据集录制时所处的动作捕捉空间的限制,同时假设关节运动连续捕捉且身体尺寸统一。在本文中,我们提出EgoPoser,它通过以下方式克服了这些限制:1)重新思考基于头戴设备的自我姿态估计的输入表示,并引入一种新颖的运动分解方法,该方法独立于全局位置预测全身姿态;2)仅在头戴设备视野范围内跟踪手部位置和方向时,鲁棒地建模身体姿态;3)针对不同用户实现跨各种身体尺寸的泛化。我们的实验表明,EgoPoser在定性和定量上均优于现有最优方法,同时保持超过600帧/秒的高推理速度。EgoPoser为未来工作建立了鲁棒的基线,其中全身姿态估计不再依赖于外部向内捕获,并能够扩展到大规模场景环境。