Language models (LMs) are pretrained on diverse data sources, including news, discussion forums, books, and online encyclopedias. A significant portion of this data includes opinions and perspectives which, on one hand, celebrate democracy and diversity of ideas, and on the other hand are inherently socially biased. Our work develops new methods to (1) measure political biases in LMs trained on such corpora, along social and economic axes, and (2) measure the fairness of downstream NLP models trained on top of politically biased LMs. We focus on hate speech and misinformation detection, aiming to empirically quantify the effects of political (social, economic) biases in pretraining data on the fairness of high-stakes social-oriented tasks. Our findings reveal that pretrained LMs do have political leanings that reinforce the polarization present in pretraining corpora, propagating social biases into hate speech predictions and misinformation detectors. We discuss the implications of our findings for NLP research and propose future directions to mitigate unfairness.
翻译:语言模型(LMs)在包括新闻、论坛、百科全书和书籍等多种数据源上进行预训练。这些数据中有相当一部分包含观点与立场——一方面推崇民主与思想多样性,另一方面又内在地带有社会偏见。我们的工作开发了新方法用于:(1) 测量基于此类语料库训练的语言模型在社会与经济维度上的政治偏见;(2) 测量基于带有政治偏见的语言模型构建的下游NLP模型的公平性。我们聚焦于仇恨言论和虚假信息检测,旨在通过实证方式量化预训练数据中的政治(社会、经济)偏见对高风险社会导向任务公平性的影响。研究发现预训练语言模型确实存在政治倾向,这种倾向强化了预训练语料库中的两极分化现象,并将社会偏见传播至仇恨言论预测与虚假信息检测器中。我们讨论了这些发现对NLP研究的意义,并提出了缓解不公平性的未来研究方向。