Natural Language Processing (NLP) plays an important role in our daily lives, particularly due to the enormous progress of Large Language Models (LLM). However, NLP has many fairness-critical use cases, e.g., as an expert system in recruitment or as an LLM-based tutor in education. Since NLP is based on human language, potentially harmful biases can diffuse into NLP systems and produce unfair results, discriminate against minorities or generate legal issues. Hence, it is important to develop a fairness certification for NLP approaches. We follow a qualitative research approach towards a fairness certification for NLP. In particular, we have reviewed a large body of literature on algorithmic fairness, and we have conducted semi-structured expert interviews with a wide range of experts from that area. We have systematically devised six fairness criteria for NLP, which can be further refined into 18 sub-categories. Our criteria offer a foundation for operationalizing and testing processes to certify fairness, both from the perspective of the auditor and the audited organization.
翻译:自然语言处理(NLP)在我们的日常生活中扮演着重要角色,这尤其得益于大型语言模型(LLM)的巨大进步。然而,NLP 拥有许多公平性至关重要的应用场景,例如作为招聘系统中的专家系统,或作为基于LLM的教育辅导工具。由于NLP基于人类语言,潜在的有害偏见可能渗透进NLP系统,导致不公平的结果、歧视少数群体或引发法律问题。因此,为NLP方法制定公平性认证至关重要。我们遵循定性研究方法,探索NLP的公平性认证。具体而言,我们系统回顾了关于算法公平性的大量文献,并对该领域的广泛专家进行了半结构化访谈。我们系统地提出了NLP的六项公平性标准,这些标准可进一步细分为18个子类别。我们的标准为从审计者和被审计组织两个视角出发,实施和测试公平性认证流程奠定了基础。