Text-guided image editing is widely needed in daily life, ranging from personal use to professional applications such as Photoshop. However, existing methods are either zero-shot or trained on an automatically synthesized dataset, which contains a high volume of noise. Thus, they still require lots of manual tuning to produce desirable outcomes in practice. To address this issue, we introduce MagicBrush (https://osu-nlp-group.github.io/MagicBrush/), the first large-scale, manually annotated dataset for instruction-guided real image editing that covers diverse scenarios: single-turn, multi-turn, mask-provided, and mask-free editing. MagicBrush comprises over 10K manually annotated triples (source image, instruction, target image), which supports trainining large-scale text-guided image editing models. We fine-tune InstructPix2Pix on MagicBrush and show that the new model can produce much better images according to human evaluation. We further conduct extensive experiments to evaluate current image editing baselines from multiple dimensions including quantitative, qualitative, and human evaluations. The results reveal the challenging nature of our dataset and the gap between current baselines and real-world editing needs.
翻译:文本引导的图像编辑在日常生活中需求广泛,涵盖从个人使用到Photoshop等专业应用场景。然而,现有方法要么采用零样本学习方式,要么在自动合成的高噪声数据集上进行训练,导致实际应用中仍需大量人工调参才能获得理想效果。为解决这一问题,我们提出MagicBrush(https://osu-nlp-group.github.io/MagicBrush/)——首个面向指令引导的真实图像编辑的大规模人工标注数据集,覆盖单轮编辑、多轮编辑、提供遮罩编辑和无遮罩编辑等多样化场景。该数据集包含超过1万组人工标注三元组(源图像、指令、目标图像),能够支持大规模文本引导图像编辑模型的训练。我们在MagicBrush上微调InstructPix2Pix,实验表明新模型在人工评估中能生成质量显著提升的图像。此外,我们通过定量、定性和人工评估等多维度实验,全面评测了当前图像编辑基线方法。实验结果揭示了本数据集的挑战性,以及现有基线方法与真实编辑需求之间存在的差距。