This study addresses the problem of 3D human mesh reconstruction from multi-view images. Recently, approaches that directly estimate the skinned multi-person linear model (SMPL)-based human mesh vertices based on volumetric heatmap representation from input images have shown good performance. We show that representation learning of vertex heatmaps using an autoencoder helps improve the performance of such approaches. Vertex heatmap autoencoder (VHA) learns the manifold of plausible human meshes in the form of latent codes using AMASS, which is a large-scale motion capture dataset. Body code predictor (BCP) utilizes the learned body prior from VHA for human mesh reconstruction from multi-view images through latent code-based supervision and transfer of pretrained weights. According to experiments on Human3.6M and LightStage datasets, the proposed method outperforms previous methods and achieves state-of-the-art human mesh reconstruction performance.
翻译:本研究针对多视角图像中的三维人体网格重建问题展开。近年来,基于输入图像中体素热力图表示直接估计蒙皮多人线性模型(SMPL)人体网格顶点的方法展现出良好性能。本文证明,利用自编码器进行顶点热力图表示学习能够提升此类方法的性能。顶点热力图自编码器(VHA)通过大规模动作捕捉数据集AMASS以潜在编码形式学习合理人体网格的流形结构。人体编码预测器(BCP)利用从VHA中习得的身体先验知识,通过基于潜在编码的监督与预训练权重迁移,实现多视角图像的人体网格重建。在Human3.6M和LightStage数据集上的实验表明,所提方法优于现有方法,达到了最先进的人体网格重建性能。