We present SimXR, a method for controlling a simulated avatar from information (headset pose and cameras) obtained from AR / VR headsets. Due to the challenging viewpoint of head-mounted cameras, the human body is often clipped out of view, making traditional image-based egocentric pose estimation challenging. On the other hand, headset poses provide valuable information about overall body motion, but lack fine-grained details about the hands and feet. To synergize headset poses with cameras, we control a humanoid to track headset movement while analyzing input images to decide body movement. When body parts are seen, the movements of hands and feet will be guided by the images; when unseen, the laws of physics guide the controller to generate plausible motion. We design an end-to-end method that does not rely on any intermediate representations and learns to directly map from images and headset poses to humanoid control signals. To train our method, we also propose a large-scale synthetic dataset created using camera configurations compatible with a commercially available VR headset (Quest 2) and show promising results on real-world captures. To demonstrate the applicability of our framework, we also test it on an AR headset with a forward-facing camera.
翻译:我们提出SimXR方法,用于从AR/VR头显获取的信息(头显姿态与摄像头数据)控制虚拟化身。由于头戴摄像头的视角具有挑战性,人体常被截断在视野之外,使得传统基于图像的自我中心姿态估计变得困难。另一方面,头显姿态虽能提供全身运动的关键信息,但缺乏手部和脚部的精细细节。为协同利用头显姿态与摄像头数据,我们让虚拟人追踪头显运动,同时分析输入图像以决定肢体运动模式:当身体部位可见时,手部和脚部运动由图像引导;当不可见时,物理规律驱动控制器生成合理动作。我们设计了不依赖中间表征的端到端方法,学习直接建立从图像与头显姿态到虚拟人控制信号的映射。为训练该方法,我们提出了一种大型合成数据集,该数据集采用与商用VR头显(Quest 2)兼容的摄像头配置创建,并在真实世界采集数据上展现了显著效果。为验证框架的适用性,我们还在配备前向摄像头的AR头显上进行了测试。