There has been significant progress in generating an animatable 3D human avatar from a single image. However, recovering texture for the 3D human avatar from a single image has been relatively less addressed. Because the generated 3D human avatar reveals the occluded texture of the given image as it moves, it is critical to synthesize the occluded texture pattern that is unseen from the source image. To generate a plausible texture map for 3D human avatars, the occluded texture pattern needs to be synthesized with respect to the visible texture from the given image. Moreover, the generated texture should align with the surface of the target 3D mesh. In this paper, we propose a texture synthesis method for a 3D human avatar that incorporates geometry information. The proposed method consists of two convolutional networks for the sampling and refining process. The sampler network fills in the occluded regions of the source image and aligns the texture with the surface of the target 3D mesh using the geometry information. The sampled texture is further refined and adjusted by the refiner network. To maintain the clear details in the given image, both sampled and refined texture is blended to produce the final texture map. To effectively guide the sampler network to achieve its goal, we designed a curriculum learning scheme that starts from a simple sampling task and gradually progresses to the task where the alignment needs to be considered. We conducted experiments to show that our method outperforms previous methods qualitatively and quantitatively.
翻译:单张图像生成可动画三维人体化身的研究已取得显著进展,但如何从单张图像恢复三维人体化身的纹理仍相对较少被探讨。由于生成的三维人体化身在运动时会暴露原始图像中被遮挡的纹理区域,因此合成未见之于源图像的遮挡纹理模式至关重要。为生成合理的三维人体化身纹理图,需根据给定图像中的可见纹理合成遮挡纹理模式,同时确保生成的纹理与目标三维网格表面对齐。本文提出一种融合几何信息的三维人体化身纹理合成方法。该方法包含两个卷积网络,分别用于采样与优化过程:采样网络利用几何信息填充源图像的遮挡区域,并将纹理与目标三维网格表面对齐;优化网络进一步细化与调整采样纹理。为保持给定图像的清晰细节,最终纹理图由采样纹理与优化纹理混合生成。为有效引导采样网络达成目标,我们设计了一种课程学习方案,从简单采样任务开始,逐步过渡至需考虑对齐的任务。实验结果表明,本方法在定性与定量指标上均优于现有方法。