Large Vision Language Models (LVLMs) have demonstrated impressive zero-shot capabilities in various vision-language dialogue scenarios. However, the absence of fine-grained visual object detection hinders the model from understanding the details of images, leading to irreparable visual hallucinations and factual errors. In this paper, we propose Lyrics, a novel multi-modal pre-training and instruction fine-tuning paradigm that bootstraps vision-language alignment from fine-grained cross-modal collaboration. Building on the foundation of BLIP-2, Lyrics infuses local visual features extracted from a visual refiner that includes image tagging, object detection and semantic segmentation modules into the Querying Transformer, while on the text side, the language inputs equip the boundary boxes and tags derived from the visual refiner. We further introduce a two-stage training scheme, in which the pre-training stage bridges the modality gap through explicit and comprehensive vision-language alignment targets. During the instruction fine-tuning stage, we introduce semantic-aware visual feature extraction, a crucial method that enables the model to extract informative features from concrete visual objects. Our approach achieves robust performance on 13 datasets across various vision-language tasks, and demonstrates promising multi-modal understanding, perception and conversation capabilities in 11 scenario-based benchmark toolkits.
翻译:大型视觉语言模型(LVLMs)在多种视觉语言对话场景中展现出令人印象深刻的零样本能力。然而,细粒度视觉目标检测的缺失阻碍了模型理解图像细节,导致无法修复的视觉幻觉和事实错误。本文提出歌词(Lyrics),一种新颖的多模态预训练与指令微调范式,通过细粒度跨模态协作引导视觉语言对齐。基于BLIP-2框架,歌词将视觉精炼器(包含图像标注、目标检测和语义分割模块)提取的局部视觉特征注入查询变换器(Querying Transformer),同时在文本侧,语言输入配备从视觉精炼器导出的边界框与标签。我们进一步引入两阶段训练方案:预训练阶段通过显式且全面的视觉语言对齐目标弥合模态差距;在指令微调阶段,我们提出语义感知视觉特征提取这一关键方法,使模型能够从具体视觉对象中提取信息特征。我们的方法在涵盖多种视觉语言任务的13个数据集上取得稳健性能,并在11个场景化基准工具包中展现出卓越的多模态理解、感知与对话能力。