Human generation has achieved significant progress. Nonetheless, existing methods still struggle to synthesize specific regions such as faces and hands. We argue that the main reason is rooted in the training data. A holistic human dataset inevitably has insufficient and low-resolution information on local parts. Therefore, we propose to use multi-source datasets with various resolution images to jointly learn a high-resolution human generative model. However, multi-source data inherently a) contains different parts that do not spatially align into a coherent human, and b) comes with different scales. To tackle these challenges, we propose an end-to-end framework, UnitedHuman, that empowers continuous GAN with the ability to effectively utilize multi-source data for high-resolution human generation. Specifically, 1) we design a Multi-Source Spatial Transformer that spatially aligns multi-source images to full-body space with a human parametric model. 2) Next, a continuous GAN is proposed with global-structural guidance and CutMix consistency. Patches from different datasets are then sampled and transformed to supervise the training of this scale-invariant generative model. Extensive experiments demonstrate that our model jointly learned from multi-source data achieves superior quality than those learned from a holistic dataset.
翻译:人体生成技术已取得显著进展。然而,现有方法在合成人脸、手部等特定区域时仍存在困难。我们认为根本原因在于训练数据——完整人体数据集不可避免地在局部区域存在信息不足和分辨率偏低的问题。为此,我们提出利用包含不同分辨率图像的多源数据集联合训练高分辨率人体生成模型。但多源数据存在两个固有挑战:a) 不同来源的数据包含的空间位置不对齐的局部区域,无法直接构成连贯人体;b) 数据具有不同尺度。针对这些问题,我们提出端到端框架UnitedHuman,赋予连续生成对抗网络有效利用多源数据生成高分辨率人体的能力。具体地:1) 设计多源空间变换器,通过人体参数模型将多源图像空间对齐至全身空间;2) 提出具有全局结构引导和CutMix一致性的连续生成对抗网络,通过对不同数据集的图像块进行采样变换,监督该尺度不变生成模型的训练。大量实验表明,联合多源数据训练的模型在生成质量上显著优于基于单一完整数据集的模型。