Understanding which concepts models can and cannot represent has been fundamental to many tasks: from effective and responsible use of models to detecting out of distribution data. We introduce Gaussian process probes (GPP), a unified and simple framework for probing and measuring uncertainty about concepts represented by models. As a Bayesian extension of linear probing methods, GPP asks what kind of distribution over classifiers (of concepts) is induced by the model. This distribution can be used to measure both what the model represents and how confident the probe is about what the model represents. GPP can be applied to any pre-trained model with vector representations of inputs (e.g., activations). It does not require access to training data, gradients, or the architecture. We validate GPP on datasets containing both synthetic and real images. Our experiments show it can (1) probe a model's representations of concepts even with a very small number of examples, (2) accurately measure both epistemic uncertainty (how confident the probe is) and aleatory uncertainty (how fuzzy the concepts are to the model), and (3) detect out of distribution data using those uncertainty measures as well as classic methods do. By using Gaussian processes to expand what probing can offer, GPP provides a data-efficient, versatile and uncertainty-aware tool for understanding and evaluating the capabilities of machine learning models.
翻译:理解模型能够表征与无法表征哪些概念,是许多任务(从模型的有效与负责任使用,到检测分布外数据)的基础。我们提出高斯过程探针(GPP)——一种用于探测模型所表征概念并测量其不确定性的统一且简洁的框架。作为线性探测方法的贝叶斯扩展,GPP关注模型在概念分类器上诱导出何种分布。该分布既可用来衡量模型表征的内容,也可用来评估探针对模型表征结果的置信程度。GPP可应用于任何具有输入向量表示(如激活值)的预训练模型,且无需访问训练数据、梯度或模型架构。我们在包含合成图像与真实图像的数据集上验证了GPP。实验表明,GPP能够:(1) 即使在样本数量极少的情况下探测模型对概念的表征;(2) 准确测量认知不确定性(探针的置信度)与偶然不确定性(概念对模型的模糊程度);(3) 利用这些不确定性度量,与经典方法一样有效地检测分布外数据。通过利用高斯过程扩展探测的能力,GPP为理解与评估机器学习模型的能力提供了一种数据高效、通用且具备不确定性感知能力的工具。