Large vision--language models (VLMs) often use a frozen vision backbone, whose image features are mapped into a large language model through a lightweight connector. While transformer-based encoders are the standard visual backbone, we ask whether state space model (SSM) vision backbones can be a strong alternative. We systematically evaluate SSM vision backbones for VLMs in a controlled setting. Under matched ImageNet-1K initialization, the SSM backbone achieves the strongest overall performance across both VQA and grounding/localization. We further adapt both SSM and ViT-family backbones with detection or segmentation training and find that dense-task tuning generally improves performance across families; after this adaptation, the SSM backbone remains competitive while operating at a substantially smaller model scale. We further observe that (i) higher ImageNet accuracy or larger backbones do not reliably translate into better VLM performance, and (ii) some visual backbones are unstable in localization. Based on these findings, we propose stabilization strategies that improve robustness for both backbone families and highlight SSM backbones as a strong alternative to transformer-based vision encoders in VLMs.
翻译:大型视觉-语言模型(VLM)通常采用冻结的视觉主干网络,通过轻量级连接器将图像特征映射到大语言模型中。尽管基于Transformer的编码器是标准的视觉主干,我们探究状态空间模型(SSM)视觉主干能否成为有力的替代方案。我们在受控条件下系统评估了VLM中的SSM视觉主干。在ImageNet-1K初始化一致的情况下,SSM主干在视觉问答(VQA)和定位/ grounding任务中均取得了最佳整体性能。我们进一步对SSM和ViT系列主干进行检测或分割训练的适配,发现密集任务微调通常能提升各系列的性能;经过此适配后,SSM主干在显著较小的模型规模下仍保持竞争力。我们还观察到:(i)更高的ImageNet精度或更大的主干网络并不稳定提升VLM性能,(ii)部分视觉主干在定位任务中不稳定。基于这些发现,我们提出能提升两种主干系列鲁棒性的稳定化策略,并强调SSM主干是VLM中基于Transformer的视觉编码器的强有力替代方案。