In the realm of multi-modality, text-guided image retouching techniques emerged with the advent of deep learning. Most currently available text-guided methods, however, rely on object-level supervision to constrain the region that may be modified. This not only makes it more challenging to develop these algorithms, but it also limits how widely deep learning can be used for image retouching. In this paper, we offer a text-guided mask-free image retouching approach that yields consistent results to address this concern. In order to perform image retouching without mask supervision, our technique can construct plausible and edge-sharp masks based on the text for each object in the image. Extensive experiments have shown that our method can produce high-quality, accurate images based on spoken language. The source code will be released soon.
翻译:在多模态领域,随着深度学习的兴起,文本引导的图像修描技术应运而生。然而,目前大多数可用的文本引导方法依赖于对象级监督来约束可能被修改的区域。这不仅增加了开发这些算法的难度,也限制了深度学习在图像修描中的广泛应用。针对这一问题,本文提出了一种文本引导的无遮罩图像修描方法,能够生成一致的结果。我们的技术可以根据文本为图像中的每个对象构建合理且边缘锐利的遮罩,从而在没有遮罩监督的情况下执行图像修描。大量实验表明,我们的方法能够基于自然语言生成高质量、准确的图像。源代码即将发布。