Societal biases are reflected in large pre-trained language models and their fine-tuned versions on downstream tasks. Common in-processing bias mitigation approaches, such as adversarial training and mutual information removal, introduce additional optimization criteria, and update the model to reach a new debiased state. However, in practice, end-users and practitioners might prefer to switch back to the original model, or apply debiasing only on a specific subset of protected attributes. To enable this, we propose a novel modular bias mitigation approach, consisting of stand-alone highly sparse debiasing subnetworks, where each debiasing module can be integrated into the core model on-demand at inference time. Our approach draws from the concept of \emph{diff} pruning, and proposes a novel training regime adaptable to various representation disentanglement optimizations. We conduct experiments on three classification tasks with gender, race, and age as protected attributes. The results show that our modular approach, while maintaining task performance, improves (or at least remains on-par with) the effectiveness of bias mitigation in comparison with baseline finetuning. Particularly on a two-attribute dataset, our approach with separately learned debiasing subnetworks shows effective utilization of either or both the subnetworks for selective bias mitigation.
翻译:社会偏见反映在大规模预训练语言模型及其下游任务的微调版本中。常见的处理中偏差缓解方法(如对抗训练和互信息移除)会引入额外的优化准则,并更新模型以达到新的去偏状态。然而在实践中,最终用户和从业者可能更希望切换回原始模型,或仅对特定受保护属性子集应用去偏处理。为实现这一目标,我们提出一种新颖的模块化偏差缓解方法,该方法由独立的高度稀疏去偏子网络构成,每个去偏模块可在推理时按需集成到核心模型中。我们的方法借鉴了"差异剪枝"的概念,并提出了一种可适配多种表征解耦优化策略的新型训练机制。我们在性别、种族和年龄作为受保护属性的三个分类任务上进行了实验。结果表明,我们的模块化方法在保持任务性能的同时,相较基线微调其偏差缓解效果得到提升(或至少持平)。特别地,在双属性数据集上,通过分别学习的去偏子网络,我们的方法能够有效利用其中一个或两个子网络实现选择性偏差缓解。