Neural networks employ spurious correlations in their predictions, resulting in decreased performance when these correlations do not hold. Recent works suggest fixing pretrained representations and training a classification head that does not use spurious features. We investigate how spurious features are represented in pretrained representations and explore strategies for removing information about spurious features. Considering the Waterbirds dataset and a few pretrained representations, we find that even with full knowledge of spurious features, their removal is not straightforward due to entangled representation. To address this, we propose a linear autoencoder training method to separate the representation into core, spurious, and other features. We propose two effective spurious feature removal approaches that are applied to the encoding and significantly improve classification performance measured by worst group accuracy.
翻译:神经网络在预测中依赖虚假相关性,当这些相关性不成立时会导致性能下降。近期研究建议固定预训练表征并训练不利用虚假特征的分类头。我们探究了虚假特征在预训练表征中的表示方式,并探讨了消除虚假特征信息的策略。以Waterbirds数据集及若干预训练表征为例,我们发现即使完全知晓虚假特征,由于表征的纠缠性,其去除过程仍非直接。为解决该问题,我们提出了一种线性自编码器训练方法,将表征分离为核心特征、虚假特征及其他特征。我们提出了两种有效的虚假特征消除方法,应用于编码过程,显著提升了以最差组准确率衡量的分类性能。