Numerous studies on text-to-image (T2I) generative models have utilized cross-attention maps to boost application performance and interpret model behavior. However, the distinct characteristics of attention maps from different attention heads remain relatively underexplored. In this study, we show that selectively aggregating cross-attention maps from heads most relevant to a target concept can improve visual interpretability. Compared to the diffusion-based segmentation method DAAM, our approach achieves higher mean IoU scores. We also find that the most relevant heads capture concept-specific features more accurately than the least relevant ones, and that selective aggregation helps diagnose prompt misinterpretations. These findings suggest that attention head selection offers a promising direction for improving the interpretability and controllability of T2I generation.
翻译:关于文本到图像(T2I)生成模型的多项研究已利用交叉注意力图来提升应用性能并解释模型行为。然而,不同注意力头产生的注意力图的独特特征仍相对未被充分探索。本研究表明,选择性地聚合与目标概念最相关的注意力头的交叉注意力图,能够提升视觉可解释性。与基于扩散的分割方法DAAM相比,我们的方法获得了更高的平均交并比(mean IoU)分数。我们还发现,最相关的注意力头比最不相关的注意力头能更准确地捕获概念特定特征,且选择性聚合有助于诊断提示误解。这些发现表明,注意力头选择为改进T2I生成的可解释性和可控性提供了一个有前景的方向。