By now there is substantial evidence that deep learning models learn certain human-interpretable features as part of their internal representations of data. As having the right (or wrong) concepts is critical to trustworthy machine learning systems, it is natural to ask which inputs from the model's original training set were most important for learning a concept at a given layer. To answer this, we combine data attribution methods with methods for probing the concepts learned by a model. Training network and probe ensembles for two concept datasets on a range of network layers, we use the recently developed TRAK method for large-scale data attribution. We find some evidence for convergence, where removing the 10,000 top attributing images for a concept and retraining the model does not change the location of the concept in the network nor the probing sparsity of the concept. This suggests that rather than being highly dependent on a few specific examples, the features that inform the development of a concept are spread in a more diffuse manner across its exemplars, implying robustness in concept formation.
翻译:目前已有大量证据表明,深度学习模型在其内部数据表征中习得了某些人类可解释的特征。由于拥有正确(或错误)的概念对于可信赖的机器学习系统至关重要,因此自然要问:模型原始训练集中哪些输入对特定层学习某个概念最为关键?为解答这一问题,我们将数据归因方法与探测模型所学概念的方法相结合。通过在多个网络层上训练针对两个概念数据集的网络与探测集成,我们采用了近期开发的TRAK方法进行大规模数据归因。研究发现存在一定收敛证据:移除概念中前10,000张贡献最大的图像并重新训练模型后,概念在网络中的定位及其探测稀疏性均未改变。这表明支撑概念发展的特征并非高度依赖少数特定样本,而是以更弥散的方式分布于其范例中,体现出概念形成的鲁棒性。