Image-text training like CLIP has dominated the pretraining of vision foundation models in recent years. Subsequent efforts have been made to introduce region-level visual learning into CLIP's pretraining but face scalability challenges due to the lack of large-scale region-level datasets. Drawing inspiration from supervised fine-tuning (SFT) in natural language processing such as instruction tuning, we explore the potential of fine-grained SFT in enhancing the generation of vision foundation models after their pretraining. Thus a two-stage method ViSFT (Vision SFT) is proposed to unleash the fine-grained knowledge of vision foundation models. In ViSFT, the vision foundation model is enhanced by performing visual joint learning on some in-domain tasks and then tested on out-of-domain benchmarks. With updating using ViSFT on 8 V100 GPUs in less than 2 days, a vision transformer with over 4.4B parameters shows improvements across various out-of-domain benchmarks including vision and vision-linguistic scenarios.
翻译:图像-文本训练(如CLIP)近年来主导了视觉基础模型的预训练。后续研究尝试将区域级视觉学习引入CLIP的预训练过程,但受限于缺乏大规模区域级数据集而面临可扩展性挑战。受自然语言处理中监督微调(如指令微调)的启发,我们探索了细粒度监督微调在提升预训练后视觉基础模型生成能力方面的潜力。为此,提出两阶段方法ViSFT(Vision SFT),旨在释放视觉基础模型的细粒度知识。ViSFT通过在某些领域内任务上进行视觉联合学习来增强视觉基础模型,随后在领域外基准测试上进行评估。在8块V100 GPU上使用ViSFT更新不到2天后,一个具有超过44亿参数的视觉Transformer在包括视觉和视觉-语言场景在内的多个领域外基准测试中展现出性能提升。