Recent advances in Large Language Models (LLMs) have stimulated a surge of research aimed at extending their applications to the visual domain. While these models exhibit promise in generating abstract image captions and facilitating natural conversations, their performance on text-rich images still requires improvement. In this paper, we introduce Contrastive Reading Model (Cream), a novel neural architecture designed to enhance the language-image understanding capability of LLMs by capturing intricate details that are often overlooked in existing methods. Cream combines vision and auxiliary encoders, fortified by a contrastive feature alignment technique, to achieve a more effective comprehension of language information in visually situated contexts within the images. Our approach bridges the gap between vision and language understanding, paving the way for the development of more sophisticated Document Intelligence Assistants. Through rigorous evaluations across diverse visually-situated language understanding tasks that demand reasoning capabilities, we demonstrate the compelling performance of Cream, positioning it as a prominent model in the field of visual document understanding. We provide our codebase and newly-generated datasets at https://github.com/naver-ai/cream .
翻译:大语言模型(LLMs)的最新进展引发了一系列旨在将其应用扩展至视觉领域的研究热潮。尽管这类模型在生成抽象图像描述和促进自然对话方面展现出潜力,但在处理富含文本的图像时仍有改进空间。本文提出对比阅读模型(Cream),这是一种新颖的神经架构,通过捕捉现有方法中常被忽视的复杂细节,增强大语言模型的图文理解能力。Cream结合视觉编码器与辅助编码器,并辅以对比特征对齐技术,从而在图像的视觉情境中实现对语言信息的更有效理解。我们的方法弥合了视觉与语言理解之间的鸿沟,为开发更复杂的文档智能助手铺平了道路。通过在各类需要推理能力的视觉情境语言理解任务上的严格评估,我们证明了Cream模型的卓越性能,使其成为视觉文档理解领域的杰出模型。我们提供的代码库和新生成的数据集见https://github.com/naver-ai/cream。