We present in this paper a novel scheme for multimodal learning named the Parallel Attention mechanism. In addition, to take into account the advantages of grammar and context in Vietnamese, we propose the Hierarchical Linguistic Features Extractor instead of using an LSTM network to extract linguistic features. Based on these two novel modules, we introduce the Parallel Attention Transformer (PAT), achieving the best accuracy compared to all baselines on the benchmark ViVQA dataset and other SOTA methods including SAAA and MCAN.
翻译:本文提出了一种名为并行注意力机制的多模态学习新方案。此外,为充分利用越南语的语法与语境优势,我们提出层级化语言特征提取器(Hierarchical Linguistic Features Extractor)替代LSTM网络来提取语言特征。基于这两个创新模块,我们构建了并行注意力Transformer(PAT),在ViVQA基准数据集上相比所有基线模型以及包括SAAA和MCAN在内的现有最优方法,均取得了最佳准确率。