Pretrained Language Models (PLMs) are widely used in NLP for various tasks. Recent studies have identified various biases that such models exhibit and have proposed methods to correct these biases. However, most of the works address a limited set of bias dimensions independently such as gender, race, or religion. Moreover, the methods typically involve finetuning the full model to maintain the performance on the downstream task. In this work, we aim to modularly debias a pretrained language model across multiple dimensions. Previous works extensively explored debiasing PLMs using limited US-centric counterfactual data augmentation (CDA). We use structured knowledge and a large generative model to build a diverse CDA across multiple bias dimensions in a semi-automated way. We highlight how existing debiasing methods do not consider interactions between multiple societal biases and propose a debiasing model that exploits the synergy amongst various societal biases and enables multi-bias debiasing simultaneously. An extensive evaluation on multiple tasks and languages demonstrates the efficacy of our approach.
翻译:预训练语言模型(PLMs)广泛应用于自然语言处理领域的多种任务。近期研究揭示了这些模型存在的多种偏见,并提出了相应的矫正方法。然而,大多数研究仅独立处理有限的偏见维度,如性别、种族或宗教。此外,这些方法通常需要微调完整模型以保持在下游任务上的性能。本研究旨在以模块化方式对预训练语言模型进行多维度去偏。先前的工作广泛探索了利用有限的以美国为中心的反事实数据增强(CDA)进行PLM去偏。我们利用结构化知识与大型生成模型,以半自动化方式构建跨多个偏见维度的多样化CDA。我们指出现有去偏方法未能考虑多种社会偏见之间的交互作用,并提出一种能够利用多种社会偏见协同效应、实现多偏见同步去偏的模型。在多任务和多语言上的广泛评估证明了我们方法的有效性。