Existing one-shot 4D head synthesis methods usually learn from monocular videos with the aid of 3DMM reconstruction, yet the latter is evenly challenging which restricts them from reasonable 4D head synthesis. We present a method to learn one-shot 4D head synthesis via large-scale synthetic data. The key is to first learn a part-wise 4D generative model from monocular images via adversarial learning, to synthesize multi-view images of diverse identities and full motions as training data; then leverage a transformer-based animatable triplane reconstructor to learn 4D head reconstruction using the synthetic data. A novel learning strategy is enforced to enhance the generalizability to real images by disentangling the learning process of 3D reconstruction and reenactment. Experiments demonstrate our superiority over the prior art.
翻译:现有单样本4D头部生成方法通常借助3DMM重建从单目视频中学习,但后者本身具有较高难度,限制了这些方法实现合理的4D头部生成。本文提出一种通过大规模合成数据学习单样本4D头部生成的方法。其核心在于:首先通过对抗学习从单目图像中学习部件化4D生成模型,以合成多视角、多身份及完整运动序列的训练数据;随后利用基于Transformer的可驱动三平面重建器,通过合成数据学习4D头部重建。我们提出一种新颖的学习策略,通过解耦三维重建与重演的学习过程,增强模型对真实图像的泛化能力。实验证明本方法优于现有先进技术。