Machine learning algorithms have become ubiquitous in a number of applications (e.g. image classification). However, due to the insufficient measurement of traditional metrics (e.g. the coarse-grained Accuracy of each classifier), substantial gaps are usually observed between the real-world performance of these algorithms and their scores in standardized evaluations. In this paper, inspired by the psychometric theories from human measurement, we propose a task-agnostic evaluation framework Camilla, where a multi-dimensional diagnostic metric Ability is defined for collaboratively measuring the multifaceted strength of each machine learning algorithm. Specifically, given the response logs from different algorithms to data samples, we leverage cognitive diagnosis assumptions and neural networks to learn the complex interactions among algorithms, samples and the skills (explicitly or implicitly pre-defined) of each sample. In this way, both the abilities of each algorithm on multiple skills and some of the sample factors (e.g. sample difficulty) can be simultaneously quantified. We conduct extensive experiments with hundreds of machine learning algorithms on four public datasets, and our experimental results demonstrate that Camilla not only can capture the pros and cons of each algorithm more precisely, but also outperforms state-of-the-art baselines on the metric reliability, rank consistency and rank stability.
翻译:机器学习算法在诸多应用(如图像分类)中已变得无处不在。然而,由于传统指标(例如每个分类器的粗粒度准确率)存在测量不足的问题,这些算法在实际场景中的表现与其在标准化评估中的得分之间通常存在显著差距。本文受人类测量中的心理测量学理论启发,提出了一种任务无关的评估框架Camilla,其中定义了一种多维诊断指标“能力”(Ability),用于协同衡量每个机器学习算法的多维度优势。具体而言,给定不同算法对数据样本的响应日志,我们利用认知诊断假设与神经网络来学习算法、样本及各样本(显式或隐式预定义的)技能之间的复杂交互。通过这种方式,每个算法在多项技能上的能力以及部分样本因子(如样本难度)可被同时量化。我们在四个公开数据集上使用数百种机器学习算法进行了广泛实验,结果表明Camilla不仅能更精确地捕捉每种算法的优缺点,还在指标可靠性、排名一致性与排名稳定性上优于现有最优基线方法。