Distribution shifts are common in real-world datasets and can affect the performance and reliability of deep learning models. In this paper, we study two types of distribution shifts: diversity shifts, which occur when test samples exhibit patterns unseen during training, and correlation shifts, which occur when test data present a different correlation between seen invariant and spurious features. We propose an integrated protocol to analyze both types of shifts using datasets where they co-exist in a controllable manner. Finally, we apply our approach to a real-world classification problem of skin cancer analysis, using out-of-distribution datasets and specialized bias annotations. Our protocol reveals three findings: 1) Models learn and propagate correlation shifts even with low-bias training; this poses a risk of accumulating and combining unaccountable weak biases; 2) Models learn robust features in high- and low-bias scenarios but use spurious ones if test samples have them; this suggests that spurious correlations do not impair the learning of robust features; 3) Diversity shift can reduce the reliance on spurious correlations; this is counter intuitive since we expect biased models to depend more on biases when invariant features are missing. Our work has implications for distribution shift research and practice, providing new insights into how models learn and rely on spurious correlations under different types of shifts.
翻译:现实世界的数据集普遍存在分布偏移,这会影响深度学习模型的性能和可靠性。本文研究了两种类型的分布偏移:多样性偏移(测试样本呈现训练中未见过的模式)和相关偏移(测试数据中已见不变特征与虚假特征之间的相关性发生变化)。我们提出了一种集成协议,利用可控方式下同时存在这两种偏移的数据集对其进行分析。最后,我们将该方法应用于皮肤癌分析这一真实世界分类问题,使用分布外数据集和专门的偏差标注。我们的协议揭示了三个发现:1)即使在低偏差训练中,模型也会学习并传播相关偏移;这带来了积累和组合无法解释的弱偏差的风险;2)模型在高偏差和低偏差场景下都能学习鲁棒特征,但如果测试样本中存在虚假特征,模型会使用它们;这表明虚假相关性不会损害鲁棒特征的学习;3)多样性偏移可以减少对虚假相关性的依赖;这有悖直觉,因为我们期望在缺乏不变特征时,有偏模型会更依赖于偏差。我们的工作对分布偏移研究和实践具有启示意义,提供了关于模型在不同类型的偏移下如何学习并依赖虚假相关性的新见解。