Modern Text-to-Image (T2I) diffusion models have achieved remarkable semantic alignment, yet they often suffer from a significant lack of variety, converging on a narrow set of visual solutions for any given prompt. This typicality bias presents a challenge for creative applications that require a wide range of generative outcomes. We identify a fundamental trade-off in current approaches to diversity: modifying model inputs requires costly optimization to incorporate feedback from the generative path. In contrast, acting on spatially-committed intermediate latents tends to disrupt the forming visual structure, leading to artifacts. In this work, we propose to apply repulsion in the Contextual Space as a novel framework for achieving rich diversity in Diffusion Transformers. By intervening in the multimodal attention channels, we apply on-the-fly repulsion during the transformer's forward pass, injecting the intervention between blocks where text conditioning is enriched with emergent image structure. This allows for redirecting the guidance trajectory after it is structurally informed but before the composition is fixed. Our results demonstrate that repulsion in the Contextual Space produces significantly richer diversity without sacrificing visual fidelity or semantic adherence. Furthermore, our method is uniquely efficient, imposing a small computational overhead while remaining effective even in modern "Turbo" and distilled models where traditional trajectory-based interventions typically fail.
翻译:现代文本到图像扩散模型在语义对齐方面取得了显著进展,但其多样性严重不足——给定同一提示词时,模型往往会收敛到少量视觉解决方案上。这种典型性偏差给需要广泛生成结果的创意应用带来了挑战。我们指出现有多样性方法中存在的基本权衡:修改模型输入需要昂贵的优化来整合生成路径的反馈;而对空间已定型的中间隐变量施加操作则容易破坏正在形成的视觉结构,导致伪影。本文提出在扩散Transformer的上下文空间中施加排斥机制,作为一种实现丰富多样性的新框架。通过干预多模态注意力通道,我们在Transformer前向传播过程中实时施加排斥操作,将干预注入在文本条件与新兴图像结构融合的模块之间。这使得我们能在引导轨迹获得结构信息后、构图定型前重新定向其路径。实验表明,上下文空间的排斥机制在不牺牲视觉保真度和语义一致性的前提下,显著提升了生成多样性。此外,本方法具有独特的高效性:仅带来微小计算开销,即使在传统轨迹干预方法通常会失效的现代"Turbo"模型和蒸馏模型中仍能保持有效性。