While vision transformers have been highly successful in improving the performance in image-based tasks, not much work has been reported on applying transformers to multilingual scene text recognition due to the complexities in the visual appearance of multilingual texts. To fill the gap, this paper proposes an augmented transformer architecture with n-grams embedding and cross-language rectification (TANGER). TANGER consists of a primary transformer with single patch embeddings of visual images, and a supplementary transformer with adaptive n-grams embeddings that aims to flexibly explore the potential correlations between neighbouring visual patches, which is essential for feature extraction from multilingual scene texts. Cross-language rectification is achieved with a loss function that takes into account both language identification and contextual coherence scoring. Extensive comparative studies are conducted on four widely used benchmark datasets as well as a new multilingual scene text dataset containing Indonesian, English, and Chinese collected from tourism scenes in Indonesia. Our experimental results demonstrate that TANGER is considerably better compared to the state-of-the-art, especially in handling complex multilingual scene texts.
翻译:尽管视觉Transformer在提升图像任务性能方面取得了巨大成功,但由于多语言文本视觉外观的复杂性,目前鲜有研究将Transformer应用于多语言场景文本识别。为填补这一空白,本文提出了一种增强型Transformer架构,该架构融合了n-gram嵌入和跨语言校正模块(TANGER)。TANGER由主Transformer(采用视觉图像的单块嵌入)和辅助Transformer(采用自适应n-gram嵌入)组成,旨在灵活探索相邻视觉块之间的潜在关联性,这对多语言场景文本的特征提取至关重要。跨语言校正通过结合语言识别和上下文连贯性评分的损失函数实现。我们基于四个广泛使用的基准数据集以及一个新收集的包含印尼语、英语和中文的多语言场景文本数据集(取自印尼旅游场景)进行了大量对比研究。实验结果表明,TANGER的性能显著优于现有最先进方法,尤其在处理复杂多语言场景文本时表现突出。