Hair editing has made tremendous progress in recent years. Early hair editing methods use well-drawn sketches or masks to specify the editing conditions. Even though they can enable very fine-grained local control, such interaction modes are inefficient for the editing conditions that can be easily specified by language descriptions or reference images. Thanks to the recent breakthrough of cross-modal models (e.g., CLIP), HairCLIP is the first work that enables hair editing based on text descriptions or reference images. However, such text-driven and reference-driven interaction modes make HairCLIP unable to support fine-grained controls specified by sketch or mask. In this paper, we propose HairCLIPv2, aiming to support all the aforementioned interactions with one unified framework. Simultaneously, it improves upon HairCLIP with better irrelevant attributes (e.g., identity, background) preservation and unseen text descriptions support. The key idea is to convert all the hair editing tasks into hair transfer tasks, with editing conditions converted into different proxies accordingly. The editing effects are added upon the input image by blending the corresponding proxy features within the hairstyle or hair color feature spaces. Besides the unprecedented user interaction mode support, quantitative and qualitative experiments demonstrate the superiority of HairCLIPv2 in terms of editing effects, irrelevant attribute preservation and visual naturalness. Our code is available at \url{https://github.com/wty-ustc/HairCLIPv2}.
翻译:近年来,发型编辑领域取得了巨大进展。早期的发型编辑方法使用精心绘制的草图或遮罩来指定编辑条件。尽管这些方法能够实现非常精细的局部控制,但对于可通过语言描述或参考图像轻松指定的编辑条件而言,这种交互方式效率较低。得益于跨模态模型(如CLIP)的最新突破,HairCLIP成为首个支持基于文本描述或参考图像进行发型编辑的工作。然而,此类文本驱动与参考图像驱动的交互模式使得HairCLIP无法支持由草图或遮罩指定的精细控制。本文提出HairCLIPv2,旨在通过统一框架支持上述所有交互模式。同时,它在保留无关属性(如身份、背景)和应对未见文本描述方面均优于HairCLIP。核心思路是将所有发型编辑任务转换为发型迁移任务,并相应地将编辑条件转换为不同代理特征。通过在发型或发色特征空间中融合对应代理特征,将编辑效果施加于输入图像。除了支持前所未有的用户交互模式外,定量与定性实验证明了HairCLIPv2在编辑效果、无关属性保留及视觉自然度方面的优越性。我们的代码已开源:\url{https://github.com/wty-ustc/HairCLIPv2}。