We propose InternLM-XComposer, a vision-language large model that enables advanced image-text comprehension and composition. The innovative nature of our model is highlighted by three appealing properties: 1) Interleaved Text-Image Composition: InternLM-XComposer can effortlessly generate coherent and contextual articles that seamlessly integrate images, providing a more engaging and immersive reading experience. Simply provide a writing instruction, and our system will generate the corresponding manuscript. It can intelligently identify the areas in the text where images would enhance the content and automatically insert the most appropriate visual candidates. 2) Comprehension with Rich Multilingual Knowledge: The text-image comprehension is empowered by training on an extensive multi-modal multilingual database with carefully crafted strategies, resulting in a deep understanding of visual content. 3) State-of-the-art Performance: Our model consistently achieves state-of-the-art results across various mainstream benchmarks for vision-language foundational models, including MME Benchmark, MMBench, MMBench-CN, Seed-Bench, CCBench (Chinese Cultural Benchmark), QBench and Tiny LVLM. Owing to the absence of established metrics for quantitatively assessing text-image composition, we have devised a robust evaluation procedure that comprises both human and GPT4-Vision (GPT4-V) to ensure reliability. Notably, our InternLM-XComposer achieves competitive text-image composition scores compared to public solutions, including GPT4-V and GPT3.5. Collectively, InternLM-XComposer seamlessly blends advanced text-image comprehension and composition, revolutionizing vision-language interaction and offering new insights and opportunities. The InternLM-XComposer model series are publicly available at https://github.com/InternLM/InternLM-XComposer.
翻译:我们提出InternLM-XComposer,一种支持高级图文理解与合成的视觉语言大模型。其创新性体现在三个引人注目的特性:1)交错文本-图像合成:InternLM-XComposer能够轻松生成连贯且上下文丰富的文章,并自然嵌入图像,提供更具吸引力和沉浸感的阅读体验。仅需提供写作指令,系统即可生成对应稿件,智能识别文本中需增强内容的图像位置,并自动插入最合适的视觉候选。2)富含多语言知识的理解:通过在大规模多模态多语言数据库上采用精心设计的策略进行训练,模型实现了对视觉内容的深度理解。3)最先进的性能:该模型在多个视觉语言基础模型主流基准(包括MME Benchmark、MMBench、MMBench-CN、Seed-Bench、CCBench(中文文化基准)、QBench及Tiny LVLM)上持续取得最佳结果。鉴于缺乏评估图文合成的成熟量化指标,我们设计了一套包含人工与GPT4-Vision(GPT4-V)双重验证的稳健评估流程。值得注意的是,与GPT4-V和GPT3.5等公开方案相比,InternLM-XComposer在图文合成评分上具有竞争力。综上所述,InternLM-XComposer将高级图文理解与合成无缝融合,革新了视觉语言交互方式,并提供了新的见解与机遇。InternLM-XComposer模型系列已公开于https://github.com/InternLM/InternLM-XComposer。