Recent work has compared neural network representations via similarity-based analyses to improve model interpretation. The quality of a similarity measure is typically evaluated by its success in assigning a high score to representations that are expected to be matched. However, existing similarity measures perform mediocrely on standard benchmarks. In this work, we develop a new similarity measure, dubbed ContraSim, based on contrastive learning. In contrast to common closed-form similarity measures, ContraSim learns a parameterized measure by using both similar and dissimilar examples. We perform an extensive experimental evaluation of our method, with both language and vision models, on the standard layer prediction benchmark and two new benchmarks that we introduce: the multilingual benchmark and the image-caption benchmark. In all cases, ContraSim achieves much higher accuracy than previous similarity measures, even when presented with challenging examples. Finally, ContraSim is more suitable for the analysis of neural networks, revealing new insights not captured by previous measures.
翻译:近期研究通过基于相似度的分析比较神经网络表征,以提升模型可解释性。相似度度量质量的评估标准通常在于其能否对预期匹配的表征赋予较高分数。然而,现有相似度度量在标准基准测试中表现平平。本研究基于对比学习提出了一种名为ContraSim的新型相似度度量。与常见的闭式相似度度量不同,ContraSim通过同时利用相似与不相似样本学习参数化度量。我们在标准层次预测基准测试以及我们引入的两个新基准测试(多语言基准测试与图像-文本基准测试)上,对语言模型和视觉模型开展了系统性实验评估。在所有案例中,即使面对具有挑战性的样本,ContraSim的准确度也显著优于先前相似度度量。此外,ContraSim更适用于神经网络分析,能够揭示先前度量方法未能捕捉的新洞察。