We present AttentionBender, a tool that manipulates cross-attention in Video Diffusion Transformers to help artists probe the internal mechanics of black-box video generation. While generative outputs are increasingly realistic, prompt-only control limits artists' ability to build intuition for the model's material process or to work beyond its default tendencies. Using an autobiographical research-through-design approach, we built on Network Bending to design AttentionBender, which applies 2D transforms (rotation, scaling, translation, etc.) to cross-attention maps to modulate generation. We assess AttentionBender by visualizing 4,500+ video generations across prompts, operations, and layer targets. Our results suggest that cross-attention is highly entangled: targeted manipulations often resist clean, localized control, producing distributed distortions and glitch aesthetics over linear edits. AttentionBender contributes a tool that functions both as an Explainable AI style probe of transformer attention mechanisms, and as a creative technique for producing novel aesthetics beyond the model's learned representational space.
翻译:我们提出AttentionBender,一种操控视频扩散Transformer中交叉注意力的工具,旨在帮助艺术家探索黑盒视频生成的内在机制。尽管生成输出日益逼真,但仅依赖提示词的控制方式限制了艺术家建立对模型材料化过程的直觉,或突破其默认生成倾向的能力。通过自传式研究型设计方法,我们基于网络弯曲(Network Bending)构建了AttentionBender,该工具对交叉注意力图施加二维变换(旋转、缩放、平移等)以调节生成过程。我们通过可视化4500余段视频生成结果(覆盖不同提示词、操作类型与层目标)评估AttentionBender。实验表明,交叉注意力具有高度纠缠性:定向操控往往难以实现纯净的局部化控制,反而在线性编辑中产生分布性失真与故障美学。AttentionBender提供了一种工具,既可作为面向Transformer注意力机制的可解释人工智能(XAI)式探针,又是一种创造性技术,能够生成超出模型习得表征空间的新颖美学效果。