Motivated by the goals of dataset pruning and defect identification, a growing body of methods have been developed to score individual examples within a dataset. These methods, which we call "example difficulty scores", are typically used to rank or categorize examples, but the consistency of rankings between different training runs, scoring methods, and model architectures is generally unknown. To determine how example rankings vary due to these random and controlled effects, we systematically compare different formulations of scores over a range of runs and model architectures. We find that scores largely share the following traits: they are noisy over individual runs of a model, strongly correlated with a single notion of difficulty, and reveal examples that range from being highly sensitive to insensitive to the inductive biases of certain model architectures. Drawing from statistical genetics, we develop a simple method for fingerprinting model architectures using a few sensitive examples. These findings guide practitioners in maximizing the consistency of their scores (e.g. by choosing appropriate scoring methods, number of runs, and subsets of examples), and establishes comprehensive baselines for evaluating scores in the future.
翻译:受数据集剪枝与缺陷识别目标的驱动,研究者们开发了越来越多用于对数据集中单个样本进行评分的方法。这些被称为"样本难度分数"的方法通常用于对样本进行排序或分类,但不同训练轮次、评分方法及模型架构之间排序结果的一致性通常未知。为探究样本排序如何因这些随机与可控因素而变化,我们系统比较了不同评分公式在多种训练轮次与模型架构下的表现。研究发现,这些分数大多具有以下共同特征:在模型的单次运行中存在噪声、与单一难度概念高度相关,并揭示出样本对某些模型架构的归纳偏差具有高度敏感至不敏感的变化范围。借鉴统计遗传学方法,我们开发了一种利用少数敏感样本对模型架构进行指纹识别的简单方法。这些发现可指导实践者最大化分数的一致性(例如通过选择适当的评分方法、运行次数及样本子集),并为未来评估分数建立全面的基准。