We show, for the first time, that neural networks trained only on synthetic data achieve state-of-the-art accuracy on the problem of 3D human pose and shape (HPS) estimation from real images. Previous synthetic datasets have been small, unrealistic, or lacked realistic clothing. Achieving sufficient realism is non-trivial and we show how to do this for full bodies in motion. Specifically, our BEDLAM dataset contains monocular RGB videos with ground-truth 3D bodies in SMPL-X format. It includes a diversity of body shapes, motions, skin tones, hair, and clothing. The clothing is realistically simulated on the moving bodies using commercial clothing physics simulation. We render varying numbers of people in realistic scenes with varied lighting and camera motions. We then train various HPS regressors using BEDLAM and achieve state-of-the-art accuracy on real-image benchmarks despite training with synthetic data. We use BEDLAM to gain insights into what model design choices are important for accuracy. With good synthetic training data, we find that a basic method like HMR approaches the accuracy of the current SOTA method (CLIFF). BEDLAM is useful for a variety of tasks and all images, ground truth bodies, 3D clothing, support code, and more are available for research purposes. Additionally, we provide detailed information about our synthetic data generation pipeline, enabling others to generate their own datasets. See the project page: https://bedlam.is.tue.mpg.de/.
翻译:我们首次证明,仅使用合成数据训练的神经网络能够在真实图像的三维人体姿态与形状(HPS)估计任务中达到最先进的精度。此前的合成数据集存在规模小、缺乏真实性或缺少逼真衣物等问题。实现足够高的真实性并非易事,我们展示了如何为全身运动场景做到这一点。具体而言,我们的BEDLAM数据集包含以SMPL-X格式标注真实三维人体的单目RGB视频,涵盖多样的体型、动作、肤色、发型和衣物。衣物通过商业物理模拟技术对运动中的身体进行逼真仿真。我们渲染了不同人数在真实场景中的图像,并设置多样化的光照和相机运动。随后,我们使用BEDLAM训练多种HPS回归器,尽管仅使用合成数据,仍在真实图像基准上达到了最先进的精度。借助BEDLAM,我们深入分析了模型设计选择对精度的影响。结果表明,在优质合成训练数据支持下,基础方法(如HMR)的精度可接近当前最先进方法(CLIFF)。BEDLAM适用于多项任务,所有图像、真实人体标注、三维衣物数据、支持代码等均开放供研究使用。此外,我们提供了合成数据生成管线的详细信息,便于他人自行生成数据集。项目页面详见:https://bedlam.is.tue.mpg.de/。