Recent advances in diffusion models (DMs) have achieved exceptional visual quality in image editing tasks. However, the global denoising dynamics of DMs inherently conflate local editing targets with the full-image context, leading to unintended modifications in non-target regions. In this paper, we shift our attention beyond DMs and turn to Masked Generative Transformers (MGTs) as an alternative approach to tackle this challenge. By predicting multiple masked tokens rather than holistic refinement, MGTs exhibit a localized decoding paradigm that endows them with the inherent capacity to explicitly preserve non-relevant regions during the editing process. Building upon this insight, we introduce the first MGT-based image editing framework, termed EditMGT. We first demonstrate that MGT's cross-attention maps provide informative localization signals for localizing edit-relevant regions and devise a multi-layer attention consolidation scheme that refines these maps to achieve fine-grained and precise localization. On top of these adaptive localization results, we introduce region-hold sampling, which restricts token flipping within low-attention areas to suppress spurious edits, thereby confining modifications to the intended target regions and preserving the integrity of surrounding non-target areas. To train EditMGT, we construct CrispEdit-2M, a high-resolution dataset spanning seven diverse editing categories. Without introducing additional parameters, we adapt a pre-trained text-to-image MGT into an image editing model through attention injection. Extensive experiments across four standard benchmarks demonstrate that, with fewer than 1B parameters, our model achieves similarity performance while enabling 6 times faster editing. Moreover, it delivers comparable or superior editing quality, with improvements of 3.6% and 17.6% on style change and style transfer tasks, respectively.


翻译:近年来,扩散模型(DMs)在图像编辑任务中取得了卓越的视觉质量。然而,DM的全局去噪动力学本质上会将局部编辑目标与全图像上下文混为一谈,导致非目标区域发生意外修改。本文跳出DM的范式,将目光转向掩码生成式Transformer(MGT),将其作为应对这一挑战的替代方案。通过预测多个掩码标记而非整体精化,MGT展现出局部化解码范式,使其天然具备在编辑过程中显式保留无关区域的能力。基于这一洞见,我们提出了首个基于MGT的图像编辑框架,命名为EditMGT。我们首先证明MGT的交叉注意力图能够为定位编辑相关区域提供信息丰富的定位信号,并设计了一种多层注意力整合方案,以精化这些注意力图实现细粒度精准定位。在这些自适应定位结果的基础上,我们引入区域保持采样,通过限制低注意力区域内的标记翻转来抑制虚假编辑,从而将修改约束在预期目标区域内,并保持周边非目标区域的完整性。为训练EditMGT,我们构建了涵盖七个不同编辑类别的高分辨率数据集CrispEdit-2M。在不引入额外参数的前提下,我们通过注意力注入将预训练的文本到图像MGT适配为图像编辑模型。在四个标准基准上的大量实验表明,我们的模型在参数少于10亿的情况下实现了相似性能,同时编辑速度提升6倍。此外,它在风格变换和风格迁移任务上分别取得了3.6%和17.6%的改进,提供了相当或更优的编辑质量。

0
下载
关闭预览

相关内容

【CVPR2025】基于组合表示移植的图像编辑方法
专知会员服务
8+阅读 · 2025年4月5日
Sora的幕后功臣?详解大火的DiT:拥抱Transformer的扩散模型
《扩散模型图像编辑》综述
专知会员服务
28+阅读 · 2024年2月28日
Graph Transformer近期进展
专知会员服务
65+阅读 · 2023年1月5日
【Tutorial】计算机视觉中的Transformer,98页ppt
专知
21+阅读 · 2021年10月25日
百闻不如一码!手把手教你用Python搭一个Transformer
大数据文摘
18+阅读 · 2019年4月22日
多图带你读懂 Transformers 的工作原理
AI研习社
10+阅读 · 2019年3月18日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关VIP内容
【CVPR2025】基于组合表示移植的图像编辑方法
专知会员服务
8+阅读 · 2025年4月5日
Sora的幕后功臣?详解大火的DiT:拥抱Transformer的扩散模型
《扩散模型图像编辑》综述
专知会员服务
28+阅读 · 2024年2月28日
Graph Transformer近期进展
专知会员服务
65+阅读 · 2023年1月5日
相关基金
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员