In real-world machine learning systems, labels are often derived from user behaviors that the system wishes to encourage. Over time, new models must be trained as new training examples and features become available. However, feedback loops between users and models can bias future user behavior, inducing a presentation bias in the labels that compromises the ability to train new models. In this paper, we propose counterfactual augmentation, a novel causal method for correcting presentation bias using generated counterfactual labels. Our empirical evaluations demonstrate that counterfactual augmentation yields better downstream performance compared to both uncorrected models and existing bias-correction methods. Model analyses further indicate that the generated counterfactuals align closely with true counterfactuals in an oracle setting.
翻译:在现实世界的机器学习系统中,标签通常来源于系统希望鼓励的用户行为。随着时间的推移,新的训练样本和特征出现时,必须训练新模型。然而,用户与模型之间的反馈循环会扭曲未来的用户行为,导致标签中出现呈现偏差,从而损害新模型的训练能力。本文提出反事实增强——一种利用生成反事实标签来校正呈现偏差的新型因果方法。实证评估表明,与未校正模型及现有偏差校正方法相比,反事实增强能取得更优的下游任务性能。模型分析进一步揭示,在理想条件(oracle)设定下,生成的反事实与真实反事实高度一致。