Text-conditioned image editing is a recently emerged and highly practical task, and its potential is immeasurable. However, most of the concurrent methods are unable to perform action editing, i.e. they can not produce results that conform to the action semantics of the editing prompt and preserve the content of the original image. To solve the problem of action editing, we propose KV Inversion, a method that can achieve satisfactory reconstruction performance and action editing, which can solve two major problems: 1) the edited result can match the corresponding action, and 2) the edited object can retain the texture and identity of the original real image. In addition, our method does not require training the Stable Diffusion model itself, nor does it require scanning a large-scale dataset to perform time-consuming training.
翻译:文本条件图像编辑是一项新近兴起且极具实用价值的任务,其潜力不可估量。然而,当前大多数方法无法实现动作编辑,即无法生成既符合编辑提示动作语义、又保留原始图像内容的结果。针对动作编辑问题,我们提出KV反演——一种能够实现出色重建性能与动作编辑的方法,该方法可解决两大难题:1) 编辑结果能与相应动作匹配;2) 编辑对象能保持原始真实图像的纹理与身份特征。此外,我们的方法无需训练Stable Diffusion模型本身,也无需扫描大规模数据集进行耗时训练。