Existing person image generative models can do either image generation or pose transfer but not both. We propose a unified diffusion model, UPGPT to provide a universal solution to perform all the person image tasks - generative, pose transfer, and editing. With fine-grained multimodality and disentanglement capabilities, our approach offers fine-grained control over the generation and the editing process of images using a combination of pose, text, and image, all without needing a semantic segmentation mask which can be challenging to obtain or edit. We also pioneer the parameterized body SMPL model in pose-guided person image generation to demonstrate new capability - simultaneous pose and camera view interpolation while maintaining a person's appearance. Results on the benchmark DeepFashion dataset show that UPGPT is the new state-of-the-art while simultaneously pioneering new capabilities of edit and pose transfer in human image generation.
翻译:现有的人物图像生成模型要么只能执行图像生成,要么只能进行姿态迁移,而无法同时完成这两类任务。本文提出了一种统一扩散模型UPGPT,为所有人物图像任务(包括生成、姿态迁移和编辑)提供通用解决方案。凭借细粒度的多模态与解耦能力,该方法能够通过结合姿态、文本和图像对图像的生成与编辑过程进行精细控制,且无需使用难以获取或编辑的语义分割掩码。此外,我们首次在姿态引导的人物图像生成中引入参数化的SMPL人体模型,展示了新的能力——在保持人物外观的同时实现姿态与摄像机视角的同步插值。在基准数据集DeepFashion上的结果表明,UPGPT达到了新的最先进水平,同时开创了人物图像生成中编辑与姿态迁移的新能力。