InstructAttribute：基于指令的细粒度物体属性编辑 (InstructAttribute: Fine-grained Object Attributes editing with Instruction)

Text-to-image (T2I) diffusion models are widely used in image editing due to their powerful generative capabilities. However, achieving fine-grained control over specific object attributes, such as color and material, remains a considerable challenge. Existing methods often fail to accurately modify these attributes or compromise structural integrity and overall image consistency. To fill this gap, we introduce Structure Preservation and Attribute Amplification (SPAA), a novel training-free framework that enables precise generation of color and material attributes for the same object by intelligently manipulating self-attention maps and cross-attention values within diffusion models. Building on SPAA, we integrate multi-modal large language models (MLLMs) to automate data curation and instruction generation. Leveraging this object attribute data collection engine, we construct the Attribute Dataset, encompassing a comprehensive range of colors and materials across diverse object categories. Using this generated dataset, we propose InstructAttribute, an instruction-tuned model that enables fine-grained and object-level attribute editing through natural language prompts. This capability holds significant practical implications for diverse fields, from accelerating product design and e-commerce visualization to enhancing virtual try-on experiences. Extensive experiments demonstrate that InstructAttribute outperforms existing instruction-based baselines, achieving a superior balance between attribute modification accuracy and structural preservation.

翻译：文本到图像（T2I）扩散模型因其强大的生成能力而被广泛应用于图像编辑。然而，实现对特定物体属性（如颜色和材质）的细粒度控制仍然是一个相当大的挑战。现有方法通常无法准确修改这些属性，或者会损害结构完整性和整体图像一致性。为填补这一空白，我们提出了结构保持与属性增强（SPAA），这是一种无需训练的新颖框架，通过智能操控扩散模型中的自注意力图与交叉注意力值，能够为同一物体精确生成颜色和材质属性。基于SPAA，我们整合了多模态大语言模型（MLLMs）以实现数据自动整理与指令生成。利用这一物体属性数据收集引擎，我们构建了属性数据集，涵盖了多样化物体类别中全面的颜色与材质范围。使用此生成的数据集，我们提出了InstructAttribute，这是一个经过指令微调的模型，能够通过自然语言提示实现细粒度的、物体级别的属性编辑。该能力对于从加速产品设计、电子商务可视化到增强虚拟试穿体验等多个领域具有重要的实际意义。大量实验表明，InstructAttribute优于现有的基于指令的基线方法，在属性修改准确性与结构保持之间实现了更优的平衡。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

【CVPR 2022】一个完全无监督的框架，从噪声和部分测量中学习图像，Robust Equivariant Imaging: a fully unsupervised framework for learning to image

专知会员服务

25+阅读 · 2022年3月3日

【NeurIPS2021】用于文本图表示学习的 GNN 嵌套 Transformer 模型：GraphFormers

专知会员服务

46+阅读 · 2021年11月24日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日