Deep neural classifiers tend to rely on spurious correlations between spurious attributes of inputs and targets to make predictions, which could jeopardize their generalization capability. Training classifiers robust to spurious correlations typically relies on annotations of spurious correlations in data, which are often expensive to get. In this paper, we tackle an annotation-free setting and propose a self-guided spurious correlation mitigation framework. Our framework automatically constructs fine-grained training labels tailored for a classifier obtained with empirical risk minimization to improve its robustness against spurious correlations. The fine-grained training labels are formulated with different prediction behaviors of the classifier identified in a novel spuriousness embedding space. We construct the space with automatically detected conceptual attributes and a novel spuriousness metric which measures how likely a class-attribute correlation is exploited for predictions. We demonstrate that training the classifier to distinguish different prediction behaviors reduces its reliance on spurious correlations without knowing them a priori and outperforms prior methods on five real-world datasets.
翻译:深度神经分类器倾向于利用输入与目标之间虚假属性的虚假相关性进行预测,这可能损害其泛化能力。训练对虚假相关性鲁棒的分类器通常依赖于数据中虚假相关性的标注,而此类标注往往成本高昂。本文针对无标注场景,提出了一种自导虚假相关性缓解框架。该框架自动为基于经验风险最小化得到的分类器构建细粒度训练标签,以提升其对虚假相关性的鲁棒性。细粒度训练标签通过分类器在新型虚假性嵌入空间中的不同预测行为来构建。我们利用自动检测的概念属性和新型虚假性度量(用于衡量类别-属性相关性被用于预测的可能性)构建该空间。实验表明,训练分类器区分不同预测行为可在未知先验知识的情况下降低其对虚假相关性的依赖,并在五个真实数据集上优于现有方法。