Amidst the evolving landscape of artificial intelligence, the convergence of visual and textual information has surfaced as a crucial frontier, leading to the advent of image-text multimodal models. This paper provides a comprehensive review of the evolution and current state of image-text multimodal models, exploring their application value, challenges, and potential research trajectories. Initially, we revisit the basic concepts and developmental milestones of these models, introducing a novel classification that segments their evolution into three distinct phases, based on their time of introduction and subsequent impact on the discipline. Furthermore, based on the tasks' significance and prevalence in the academic landscape, we propose a categorization of the tasks associated with image-text multimodal models into five major types, elucidating the recent progress and key technologies within each category. Despite the remarkable accomplishments of these models, numerous challenges and issues persist. This paper delves into the inherent challenges and limitations of image-text multimodal models, fostering the exploration of prospective research directions. Our objective is to offer an exhaustive overview of the present research landscape of image-text multimodal models and to serve as a valuable reference for future scholarly endeavors. We extend an invitation to the broader community to collaborate in enhancing the image-text multimodal model community, accessible at: \href{https://github.com/i2vec/A-survey-on-image-text-multimodal-models}{https://github.com/i2vec/A-survey-on-image-text-multimodal-models}.
翻译:在人工智能不断发展的背景下,视觉与文本信息的融合已成为一个关键前沿领域,催生了图像-文本多模态模型的出现。本文对图像-文本多模态模型的演变历程与当前状态进行了全面综述,探讨了其应用价值、面临的挑战及潜在的研究方向。首先,我们回顾了这些模型的基本概念与发展里程碑,并根据其引入时间及对学科领域的后续影响,提出了一种将其演变划分为三个不同阶段的新颖分类方法。此外,基于任务在学术领域中的重要性和普遍性,我们将与图像-文本多模态模型相关的任务归纳为五大类型,并阐释了每类任务的最新进展与关键技术。尽管这些模型取得了显著成就,但诸多挑战与问题依然存在。本文深入探讨了图像-文本多模态模型固有的挑战与局限,以促进对前瞻性研究方向的探索。我们的目标是全面概述图像-文本多模态模型当前的研究格局,并为未来的学术研究提供有价值的参考。我们诚挚邀请更广泛的学术社区参与合作,共同推动图像-文本多模态模型社区的发展,相关资源可访问:\href{https://github.com/i2vec/A-survey-on-image-text-multimodal-models}{https://github.com/i2vec/A-survey-on-image-text-multimodal-models}。