We propose a data-driven approach for context-aware person image generation. Specifically, we attempt to generate a person image such that the synthesized instance can blend into a complex scene. In our method, the position, scale, and appearance of the generated person are semantically conditioned on the existing persons in the scene. The proposed technique is divided into three sequential steps. At first, we employ a Pix2PixHD model to infer a coarse semantic mask that represents the new person's spatial location, scale, and potential pose. Next, we use a data-centric approach to select the closest representation from a precomputed cluster of fine semantic masks. Finally, we adopt a multi-scale, attention-guided architecture to transfer the appearance attributes from an exemplar image. The proposed strategy enables us to synthesize semantically coherent realistic persons that can blend into an existing scene without altering the global context. We conclude our findings with relevant qualitative and quantitative evaluations.
翻译:我们提出了一种数据驱动的上下文感知人物图像生成方法。具体而言,我们尝试生成人物图像,使得合成实例能够融入复杂场景。在该方法中,生成人物的位置、尺度与外观在语义上受场景中已有人员的条件约束。所提出的技术分为三个连续步骤。首先,我们采用Pix2PixHD模型推断出表示新人物空间位置、尺度及潜在姿态的粗语义掩码。其次,通过数据驱动的策略从预计算的精细语义掩码聚类中选择最接近的表示。最后,我们采用多尺度注意力引导架构,从示例图像中迁移外观特征。该策略使我们能够合成语义一致且逼真的人物,在不改变全局上下文的前提下融入现有场景。我们通过相关的定性与定量评估对结论进行了验证。