Discriminatory social biases, including gender biases, have been found in Pre-trained Language Models (PLMs). In Natural Language Inference (NLI), recent bias evaluation methods have observed biased inferences from the outputs of a particular label such as neutral or entailment. However, since different biased inferences can be associated with different output labels, it is inaccurate for a method to rely on one label. In this work, we propose an evaluation method that considers all labels in the NLI task. We create evaluation data and assign them into groups based on their expected biased output labels. Then, we define a bias measure based on the corresponding label output of each data group. In the experiment, we propose a meta-evaluation method for NLI bias measures, and then use it to confirm that our measure can evaluate bias more accurately than the baseline. Moreover, we show that our evaluation method is applicable to multiple languages by conducting the meta-evaluation on PLMs in three different languages: English, Japanese, and Chinese. Finally, we evaluate PLMs of each language to confirm their bias tendency. To our knowledge, we are the first to build evaluation datasets and measure the bias of PLMs from the NLI task in Japanese and Chinese.
翻译:预训练语言模型(PLMs)中存在包括性别偏见在内的歧视性社会偏见。在自然语言推理(NLI)任务中,近期偏见评估方法发现,模型在特定标签(如中性或蕴含)的输出中表现出有偏推理。然而,由于不同有偏推理可能对应不同输出标签,仅依赖单一标签的评估方法存在不准确性。本研究提出一种考虑NLI任务中所有标签的评估方法:我们构建评估数据,并根据其预期有偏输出标签进行分组,进而基于每组数据对应的标签输出定义偏见度量指标。实验部分,我们提出NLI偏见度量的元评估方法,并验证所提度量能比基线方法更准确地评估偏见。此外,通过对英语、日语和汉语三种语言的PLMs进行元评估,证明本方法适用于多语言场景。最终,我们评估各语言PLMs的偏见倾向性。据我们所知,本研究首次构建日语和汉语NLI任务的评估数据集,并测量其PLMs的偏见水平。