Scene-Text Visual Question Answering (ST-VQA) aims to understand scene text in images and answer questions related to the text content. Most existing methods heavily rely on the accuracy of Optical Character Recognition (OCR) systems, and aggressive fine-tuning based on limited spatial location information and erroneous OCR text information often leads to inevitable overfitting. In this paper, we propose a multimodal adversarial training architecture with spatial awareness capabilities. Specifically, we introduce an Adversarial OCR Enhancement (AOE) module, which leverages adversarial training in the embedding space of OCR modality to enhance fault-tolerant representation of OCR texts, thereby reducing noise caused by OCR errors. Simultaneously, We add a Spatial-Aware Self-Attention (SASA) mechanism to help the model better capture the spatial relationships among OCR tokens. Various experiments demonstrate that our method achieves significant performance improvements on both the ST-VQA and TextVQA datasets and provides a novel paradigm for multimodal adversarial training.
翻译:场景文本视觉问答(ST-VQA)旨在理解图像中的场景文本,并回答与文本内容相关的问题。现有方法大多严重依赖光学字符识别(OCR)系统的准确性,且基于有限的空间位置信息和错误的OCR文本信息进行激进微调,往往会导致不可避免的过拟合。本文提出了一种具有空间感知能力的多模态对抗训练架构。具体来说,我们引入了一个对抗性OCR增强(AOE)模块,该模块利用OCR模态嵌入空间中的对抗训练来增强OCR文本的容错表示,从而减少由OCR错误引起的噪声。同时,我们添加了空间感知自注意力(SASA)机制,以帮助模型更好地捕捉OCR标记之间的空间关系。大量实验表明,我们的方法在ST-VQA和TextVQA数据集上均取得了显著的性能提升,并为多模态对抗训练提供了一种新的范式。