Despite recent progress towards scaling up multimodal vision-language models, these models are still known to struggle on compositional generalization benchmarks such as Winoground. We find that a critical component lacking from current vision-language models is relation-level alignment: the ability to match directional semantic relations in text (e.g., "mug in grass") with spatial relationships in the image (e.g., the position of the mug relative to the grass). To tackle this problem, we show that relation alignment can be enforced by encouraging the directed language attention from 'mug' to 'grass' (capturing the semantic relation 'in') to match the directed visual attention from the mug to the grass. Tokens and their corresponding objects are softly identified using the cross-modal attention. We prove that this notion of soft relation alignment is equivalent to enforcing congruence between vision and language attention matrices under a 'change of basis' provided by the cross-modal attention matrix. Intuitively, our approach projects visual attention into the language attention space to calculate its divergence from the actual language attention, and vice versa. We apply our Cross-modal Attention Congruence Regularization (CACR) loss to UNITER and improve on the state-of-the-art approach to Winoground.
翻译:尽管多模态视觉-语言模型在规模化方面取得了近期进展,但这些模型在组合泛化基准测试(如Winoground)上仍显得力不从心。我们发现当前视觉-语言模型缺失的关键组件是关系级对齐:即匹配文本中方向性语义关系(例如“草地里的杯子”)与图像中空间关系(例如杯子相对于草地的位置)的能力。为解决这一问题,我们证明通过鼓励从“杯子”到“草地”的有向语言注意力(捕获语义关系“在……里”)匹配从杯子到草地的有向视觉注意力,可以实现关系对齐。令牌及其对应物体通过跨模态注意力被软性识别。我们证明这种软关系对齐的概念等价于在跨模态注意力矩阵提供的“基变换”下,强制视觉与语言注意力矩阵之间的一致性。直观地,我们的方法将视觉注意力投影到语言注意力空间以计算其与实际语言注意力的差异,反之亦然。我们将跨模态注意力一致性正则化(CACR)损失应用于UNITER,并改进了Winoground上的最先进方法。