The development of large vision-language models, notably CLIP, has catalyzed research into effective adaptation techniques, with a particular focus on soft prompt tuning. Conjointly, test-time augmentation, which utilizes multiple augmented views of a single image to enhance zero-shot generalization, is emerging as a significant area of interest. This has predominantly directed research efforts toward test-time prompt tuning. In contrast, we introduce a robust MeanShift for Test-time Augmentation (MTA), which surpasses prompt-based methods without requiring this intensive training procedure. This positions MTA as an ideal solution for both standalone and API-based applications. Additionally, our method does not rely on ad hoc rules (e.g., confidence threshold) used in some previous test-time augmentation techniques to filter the augmented views. Instead, MTA incorporates a quality assessment variable for each view directly into its optimization process, termed as the inlierness score. This score is jointly optimized with a density mode seeking process, leading to an efficient training- and hyperparameter-free approach. We extensively benchmark our method on 15 datasets and demonstrate MTA's superiority and computational efficiency. Deployed easily as plug-and-play module on top of zero-shot models and state-of-the-art few-shot methods, MTA shows systematic and consistent improvements.
翻译:大型视觉语言模型(尤其是CLIP)的发展推动了高效适应技术的研究,其中软提示调优成为焦点。同时,测试时增强技术利用单张图像的多视角增强来提升零样本泛化能力,正成为一个重要的研究方向,这促使研究重点转向测试时提示调优。与此不同,我们提出了一种鲁棒的测试时均值漂移增强(MTA)方法,该方法无需训练过程即可超越基于提示的方法。这使得MTA成为独立和基于API应用的理想解决方案。此外,本方法不依赖先前某些测试时增强技术中用于筛选增强视图的特定规则(如置信度阈值)。相反,MTA将每个视图的质量评估变量(称为内点得分)直接纳入优化过程。该得分与密度模式搜索过程联合优化,从而形成一种无需训练和超参数的高效方法。我们在15个数据集上进行了全面基准测试,证明MTA的优越性和计算效率。作为零样本模型和最新少样本方法上的即插即用模块,MTA展现出系统且一致的性能提升。