Recently, efficient fine-tuning of large-scale pre-trained models has attracted increasing research interests, where linear probing (LP) as a fundamental module is involved in exploiting the final representations for task-dependent classification. However, most of the existing methods focus on how to effectively introduce a few of learnable parameters, and little work pays attention to the commonly used LP module. In this paper, we propose a novel Moment Probing (MP) method to further explore the potential of LP. Distinguished from LP which builds a linear classification head based on the mean of final features (e.g., word tokens for ViT) or classification tokens, our MP performs a linear classifier on feature distribution, which provides the stronger representation ability by exploiting richer statistical information inherent in features. Specifically, we represent feature distribution by its characteristic function, which is efficiently approximated by using first- and second-order moments of features. Furthermore, we propose a multi-head convolutional cross-covariance (MHC$^3$) to compute second-order moments in an efficient and effective manner. By considering that MP could affect feature learning, we introduce a partially shared module to learn two recalibrating parameters (PSRP) for backbones based on MP, namely MP$_{+}$. Extensive experiments on ten benchmarks using various models show that our MP significantly outperforms LP and is competitive with counterparts at less training cost, while our MP$_{+}$ achieves state-of-the-art performance.
翻译:近期,大规模预训练模型的高效微调引起了广泛研究兴趣,其中线性探测(LP)作为基础模块,被用于利用最终表示进行任务依赖的分类。然而,现有方法大多聚焦于如何有效引入少量可学习参数,鲜有关注常用的LP模块本身。本文提出了一种新颖的矩探测(MP)方法,以进一步挖掘LP的潜力。不同于基于最终特征均值(例如ViT的词令牌)或分类令牌构建线性分类头的LP,我们的MP在特征分布上执行线性分类器,通过利用特征中更丰富的统计信息来提供更强的表示能力。具体而言,我们使用特征的特征函数来表示特征分布,该函数通过一阶矩和二阶矩高效近似。此外,我们提出了多头卷积互协方差(MHC$^3$)以高效且有效地计算二阶矩。考虑到MP可能影响特征学习,我们引入部分共享模块,为基于MP的主干网络学习两个重校准参数(PSRP),即MP$_{+}$。使用多种模型在十个基准上的大量实验表明,我们的MP以更低的训练成本显著优于LP,并与同类方法性能相当;而MP$_{+}$则取得了最先进性能。