We study the problem of inferring scene affordances by presenting a method for realistically inserting people into scenes. Given a scene image with a marked region and an image of a person, we insert the person into the scene while respecting the scene affordances. Our model can infer the set of realistic poses given the scene context, re-pose the reference person, and harmonize the composition. We set up the task in a self-supervised fashion by learning to re-pose humans in video clips. We train a large-scale diffusion model on a dataset of 2.4M video clips that produces diverse plausible poses while respecting the scene context. Given the learned human-scene composition, our model can also hallucinate realistic people and scenes when prompted without conditioning and also enables interactive editing. A quantitative evaluation shows that our method synthesizes more realistic human appearance and more natural human-scene interactions than prior work.
翻译:我们研究场景功能推断问题,提出一种将人物真实感地插入场景的方法。给定包含标注区域的场景图像与人物图像,该方法能在尊重场景功能的前提下将人物插入场景。模型可根据场景上下文推断合理姿态集合、调整参考人物姿态并协调整体构图。我们通过视频片段中的人物姿态重定向学习,以自监督方式构建任务。基于包含240万视频片段的数据集训练大规模扩散模型,生成符合场景上下文的多样化合理姿态。利用学习到的人-场景构图关系,模型在无需条件约束时也能生成逼真的人物与场景,并支持交互式编辑。定量评估表明,本方法合成的角色外观与人物-场景交互自然性均优于现有方案。