The contrastive vision-language pre-training, known as CLIP, demonstrates remarkable potential in perceiving open-world visual concepts, enabling effective zero-shot image recognition. Nevertheless, few-shot learning methods based on CLIP typically require offline fine-tuning of the parameters on few-shot samples, resulting in longer inference time and the risk of over-fitting in certain domains. To tackle these challenges, we propose the Meta-Adapter, a lightweight residual-style adapter, to refine the CLIP features guided by the few-shot samples in an online manner. With a few training samples, our method can enable effective few-shot learning capabilities and generalize to unseen data or tasks without additional fine-tuning, achieving competitive performance and high efficiency. Without bells and whistles, our approach outperforms the state-of-the-art online few-shot learning method by an average of 3.6\% on eight image classification datasets with higher inference speed. Furthermore, our model is simple and flexible, serving as a plug-and-play module directly applicable to downstream tasks. Without further fine-tuning, Meta-Adapter obtains notable performance improvements in open-vocabulary object detection and segmentation tasks.
翻译:基于对比学习的视觉-语言预训练模型CLIP在感知开放世界视觉概念方面展现出显著潜力,能够实现有效的零样本图像识别。然而,基于CLIP的小样本学习方法通常需要在少量样本上进行参数的离线微调,导致推理时间延长,并存在特定领域过拟合的风险。为应对这些挑战,我们提出Meta-Adapter——一种轻量级残差式适配器,能够通过少量样本以在线方式优化CLIP特征。基于少量训练样本,本方法即可实现高效的小样本学习能力,并无需额外微调即可泛化至未见数据或任务,实现竞争力性能与高时效性。无需繁复设计,我们的方法在八个图像分类数据集上以更高推理速度平均超越现有最优在线小样本学习方法3.6%。此外,本模型简洁灵活,可作为即插即用模块直接应用于下游任务。在无需进一步微调的情况下,Meta-Adapter在开放词汇目标检测与分割任务中均取得显著性能提升。