We propose InternLM-XComposer, a vision-language large model that enables advanced image-text comprehension and composition. The innovative nature of our model is highlighted by three appealing properties: 1) Interleaved Text-Image Composition: InternLM-XComposer can effortlessly generate coherent and contextual articles that seamlessly integrate images, providing a more engaging and immersive reading experience. Simply provide a title, and our system will generate the corresponding manuscript. It can intelligently identify the areas in the text where images would enhance the content and automatically insert the most appropriate visual candidates. 2) Comprehension with Rich Multilingual Knowledge: The text-image comprehension is empowered by training on extensive multi-modal multilingual concepts with carefully crafted strategies, resulting in a deep understanding of visual content. 3) State-of-the-art Performance: Our model consistently achieves state-of-the-art results across various mainstream benchmarks for vision-language foundational models, including MME Benchmark, MMBench, MMBench-CN, Seed-Bench, and CCBench (Chinese Cultural Benchmark). Collectively, InternLM-XComposer seamlessly blends advanced text-image comprehension and composition, revolutionizing vision-language interaction and offering new insights and opportunities. The InternLM-XComposer model series with 7B parameters are publicly available at https://github.com/InternLM/InternLM-XComposer.
翻译:我们提出了InternLM-XComposer,这是一个支持高级图像-文本理解与合成的视觉语言大模型。该模型的创新性体现在三个引人注目的特性上:1)交错式文本-图像合成:InternLM-XComposer能够轻松生成连贯且富有语境的文章,并自然融入图像,提供更具吸引力和沉浸感的阅读体验。仅需提供标题,系统即可生成对应的稿件。它能智能识别文本中需用图像增强内容的位置,并自动插入最合适的视觉候选元素。2)基于丰富多语言知识的理解能力:通过在大规模多模态多语言概念上采用精心设计的策略进行训练,模型对视觉内容的理解得以增强,从而实现对视觉内容的深层认知。3)最先进的性能:本模型在多个视觉语言基础模型的主流基准测试中持续取得最先进成果,包括MME Benchmark、MMBench、MMBench-CN、Seed-Bench及CCBench(中国文化基准)。总的来说,InternLM-XComposer无缝融合了高级图文理解与合成技术,革新了视觉语言交互方式,并提供了全新的见解与机遇。含70亿参数的InternLM-XComposer模型系列已开源至https://github.com/InternLM/InternLM-XComposer。