Recent instruction fine-tuned models can solve multiple NLP tasks when prompted to do so, with machine translation (MT) being a prominent use case. However, current research often focuses on standard performance benchmarks, leaving compelling fairness and ethical considerations behind. In MT, this might lead to misgendered translations, resulting, among other harms, in the perpetuation of stereotypes and prejudices. In this work, we address this gap by investigating whether and to what extent such models exhibit gender bias in machine translation and how we can mitigate it. Concretely, we compute established gender bias metrics on the WinoMT corpus from English to German and Spanish. We discover that IFT models default to male-inflected translations, even disregarding female occupational stereotypes. Next, using interpretability methods, we unveil that models systematically overlook the pronoun indicating the gender of a target occupation in misgendered translations. Finally, based on this finding, we propose an easy-to-implement and effective bias mitigation solution based on few-shot learning that leads to significantly fairer translations.
翻译:近期,指令微调模型能够在提示引导下解决多项自然语言处理任务,其中机器翻译是一个典型应用场景。然而,当前研究通常聚焦于标准性能基准测试,忽视了紧迫的公平性与伦理考量。在机器翻译领域,这种忽视可能导致性别误译,进而造成刻板印象与偏见的强化等危害。本研究通过探究此类模型在机器翻译中是否存在性别偏见及其程度,以及如何缓解该偏见,来弥合这一研究空白。具体而言,我们在WinoMT语料库上(涵盖英语到德语及西班牙语的翻译任务)计算了既定的性别偏见度量指标。研究发现,指令微调模型默认生成阳性屈折变化的译文,甚至无视女性职业刻板印象,进而产生性别误译。随后,我们借助可解释性方法揭示,在发生性别误译的案例中,模型系统性地忽略了指示目标职业性别的代词。基于这一发现,我们提出了一种基于少样本学习的简便易行且有效的偏见缓解方案,可实现显著更公平的译文。