Models that can actively seek out the best quality training data hold the promise of more accurate, adaptable, and efficient machine learning. Active learning techniques often tend to prefer examples that are the most difficult to classify. While this works well on homogeneous datasets, we find that it can lead to catastrophic failures when performed on multiple distributions with different degrees of label noise or heteroskedasticity. These active learning algorithms strongly prefer to draw from the distribution with more noise, even if their examples have no informative structure (such as solid color images with random labels). To this end, we demonstrate the catastrophic failure of these active learning algorithms on heteroskedastic distributions and propose a fine-tuning-based approach to mitigate these failures. Further, we propose a new algorithm that incorporates a model difference scoring function for each data point to filter out the noisy examples and sample clean examples that maximize accuracy, outperforming the existing active learning techniques on the heteroskedastic datasets. We hope these observations and techniques are immediately helpful to practitioners and can help to challenge common assumptions in the design of active learning algorithms.
翻译:能够主动寻找高质量训练数据的模型有望实现更准确、更适应且更高效的机器学习。主动学习技术通常倾向于选择最难分类的样本。虽然这种方法在均匀数据集上表现良好,但我们发现在存在不同标签噪声程度或异方差性的多个分布上运行时,它可能导致灾难性失败。这些主动学习算法强烈倾向于从噪声更大的分布中采样,即使这些样本不具备任何信息性结构(例如带有随机标签的纯色图像)。为此,我们展示了主动学习算法在异方差分布上的灾难性失败,并提出了一种基于微调的方法来减轻此类失败。此外,我们提出了一种新算法,该算法为每个数据点引入模型差异评分函数,以滤除噪声样本并采样能最大化准确率的干净样本,在异方差数据集上优于现有的主动学习技术。我们希望这些观察结果和技术能立即为实践者提供帮助,并有助于挑战主动学习算法设计中的常见假设。