Prediction models can perform poorly when deployed to target distributions different from the training distribution. To understand these operational failure modes, we develop a method, called DIstribution Shift DEcomposition (DISDE), to attribute a drop in performance to different types of distribution shifts. Our approach decomposes the performance drop into terms for 1) an increase in harder but frequently seen examples from training, 2) changes in the relationship between features and outcomes, and 3) poor performance on examples infrequent or unseen during training. These terms are defined by fixing a distribution on $X$ while varying the conditional distribution of $Y \mid X$ between training and target, or by fixing the conditional distribution of $Y \mid X$ while varying the distribution on $X$. In order to do this, we define a hypothetical distribution on $X$ consisting of values common in both training and target, over which it is easy to compare $Y \mid X$ and thus predictive performance. We estimate performance on this hypothetical distribution via reweighting methods. Empirically, we show how our method can 1) inform potential modeling improvements across distribution shifts for employment prediction on tabular census data, and 2) help to explain why certain domain adaptation methods fail to improve model performance for satellite image classification.
翻译:预测模型在部署到与训练分布不同的目标分布时,其性能可能显著下降。为理解这些运行失效模式,我们提出了一种名为"分布偏移分解"(DISDE)的方法,用于将性能下降归因于不同类型的分布偏移。该方法将性能下降分解为三个分量:1)训练集中常见但难度增大样本的增加;2)特征与结果间关系的变化;3)训练中罕见或未出现样本上的性能下降。这些分量通过固定输入分布$X$并比较训练与目标条件下$Y \mid X$的条件分布,或固定条件分布$Y \mid X$并改变输入分布$X$来定义。为此,我们定义了一个由训练集与目标集中常见值构成的假设输入分布$X$,在该分布上易于比较$Y \mid X$及预测性能,并通过重加权方法估计该假设分布下的性能。实验表明,该方法可:1)针对表格人口普查数据的就业预测,指导跨分布偏移的建模改进策略;2)解释为何某些域适应方法未能提升卫星图像分类模型的性能。