Multimodality Representation Learning, as a technique of learning to embed information from different modalities and their correlations, has achieved remarkable success on a variety of applications, such as Visual Question Answering (VQA), Natural Language for Visual Reasoning (NLVR), and Vision Language Retrieval (VLR). Among these applications, cross-modal interaction and complementary information from different modalities are crucial for advanced models to perform any multimodal task, e.g., understand, recognize, retrieve, or generate optimally. Researchers have proposed diverse methods to address these tasks. The different variants of transformer-based architectures performed extraordinarily on multiple modalities. This survey presents the comprehensive literature on the evolution and enhancement of deep learning multimodal architectures to deal with textual, visual and audio features for diverse cross-modal and modern multimodal tasks. This study summarizes the (i) recent task-specific deep learning methodologies, (ii) the pretraining types and multimodal pretraining objectives, (iii) from state-of-the-art pretrained multimodal approaches to unifying architectures, and (iv) multimodal task categories and possible future improvements that can be devised for better multimodal learning. Moreover, we prepare a dataset section for new researchers that covers most of the benchmarks for pretraining and finetuning. Finally, major challenges, gaps, and potential research topics are explored. A constantly-updated paperlist related to our survey is maintained at https://github.com/marslanm/multimodality-representation-learning.
翻译:多模态表示学习作为一种将来自不同模态的信息及其关联进行嵌入学习的技术,已在视觉问答(VQA)、自然语言视觉推理(NLVR)和视觉语言检索(VLR)等多种应用中取得了显著成功。在这些应用中,跨模态交互以及不同模态间的互补信息对于先进模型执行任何多模态任务(如理解、识别、检索或最优生成)至关重要。研究者们已提出多种方法来解决这些任务。基于Transformer架构的不同变体在多模态处理上表现尤为突出。本综述全面梳理了深度学习多模态架构的演变与增强历程,以处理文本、视觉和音频特征,从而应对多样化的跨模态与现代多模态任务。本研究归纳了:(i) 近期面向特定任务的深度学习方法论,(ii) 预训练类型与多模态预训练目标,(iii) 从最先进的预训练多模态方法到统一架构的演进,以及 (iv) 多模态任务类别及其潜在未来改进方向,以促进更优的多模态学习。此外,我们为新手研究者整理了一个数据集章节,覆盖了预训练与微调所需的大多数基准。最后,我们探讨了主要挑战、现有空白及潜在研究课题。与本综述相关的持续更新论文列表可在 https://github.com/marslanm/multimodality-representation-learning 获取。