Statistical fairness stipulates equivalent outcomes for every protected group, whereas causal fairness prescribes that a model makes the same prediction for an individual regardless of their protected characteristics. Counterfactual data augmentation (CDA) is effective for reducing bias in NLP models, yet models trained with CDA are often evaluated only on metrics that are closely tied to the causal fairness notion; similarly, sampling-based methods designed to promote statistical fairness are rarely evaluated for causal fairness. In this work, we evaluate both statistical and causal debiasing methods for gender bias in NLP models, and find that while such methods are effective at reducing bias as measured by the targeted metric, they do not necessarily improve results on other bias metrics. We demonstrate that combinations of statistical and causal debiasing techniques are able to reduce bias measured through both types of metrics.
翻译:摘要:统计公平性要求每个受保护群体获得同等结果,而因果公平性则规定模型对个体的预测应独立于其受保护特征。反事实数据增强(CDA)能有效减少NLP模型中的偏见,但经CDA训练的模型通常仅在紧密关联因果公平性概念的指标上被评估;类似地,旨在促进统计公平性的基于采样的方法也很少接受因果公平性评估。本研究对NLP模型中性别偏见的统计与因果去偏方法进行了评估,发现虽然这些方法能有效减少目标度量指标所衡量的偏见,但未必能在其他偏见指标上改善结果。我们证明,通过结合统计与因果去偏技术,可同时降低这两类指标衡量的偏见。