We present a simple but effective method to measure and mitigate model biases caused by reliance on spurious cues. Instead of requiring costly changes to one's data or model training, our method better utilizes the data one already has by sorting them. Specifically, we rank images within their classes based on spuriosity (the degree to which common spurious cues are present), proxied via deep neural features of an interpretable network. With spuriosity rankings, it is easy to identify minority subpopulations (i.e. low spuriosity images) and assess model bias as the gap in accuracy between high and low spuriosity images. One can even efficiently remove a model's bias at little cost to accuracy by finetuning its classification head on low spuriosity images, resulting in fairer treatment of samples regardless of spuriosity. We demonstrate our method on ImageNet, annotating $5000$ class-feature dependencies ($630$ of which we find to be spurious) and generating a dataset of $325k$ soft segmentations for these features along the way. Having computed spuriosity rankings via the identified spurious neural features, we assess biases for $89$ diverse models and find that class-wise biases are highly correlated across models. Our results suggest that model bias due to spurious feature reliance is influenced far more by what the model is trained on than how it is trained.
翻译:我们提出一种简单有效的方法,用于测量并缓解模型因依赖偶发性线索而产生的偏差。该方法无需对数据或模型训练进行高成本改动,而是通过排序现有数据实现更优利用。具体而言,我们基于可解释网络的深层神经特征代理偶发性(常见偶发性线索的呈现程度),对类别内图像进行排序。通过偶发性排序,可轻松识别少数子群体(即低偶发性图像),并以高/低偶发性图像间的准确率差距评估模型偏差。甚至可通过在低偶发性图像上微调分类头,以极小准确率代价高效消除模型偏差,使样本获得更公平的偶发性无关处理。我们在ImageNet上验证该方法,标注了5000个类别-特征依赖关系(其中630个被判定为偶发性),并生成包含32.5万张对应特征软分割掩码的数据集。通过已识别的偶发性神经特征计算偶发性排序后,我们评估了89种不同模型的偏差,发现类别级偏差在各模型间高度相关。实验结果表明,因偶发性特征依赖导致的模型偏差更多取决于训练数据本身,而非训练方式。