Reinforcing or even exacerbating societal biases and inequalities will increase significantly as generative AI increasingly produces useful artifacts, from text to images and beyond, for the real world. We address these issues by formally characterizing the notion of fairness for generative AI as a basis for monitoring and enforcing fairness. We define two levels of fairness using the notion of infinite sequences of abstractions of AI-generated artifacts such as text or images. The first is the fairness demonstrated on the generated sequences, which is evaluated only on the outputs while agnostic to the prompts and models used. The second is the inherent fairness of the generative AI model, which requires that fairness be manifested when input prompts are neutral, that is, they do not explicitly instruct the generative AI to produce a particular type of output. We also study relative intersectional fairness to counteract the combinatorial explosion of fairness when considering multiple categories together with lazy fairness enforcement. Finally, fairness monitoring and enforcement are tested against some current generative AI models.
翻译:随着生成式人工智能日益高效地为现实世界产出从文本到图像等多样化人造制品,其强化甚至加剧社会偏见与不平等的风险将显著提升。我们通过形式化刻画生成式人工智能的公平性概念来应对这些问题,以此作为监测与执行公平性的基础。利用人工智能生成制品(如文本或图像)的无限抽象序列概念,我们定义了两种公平性层级:其一是生成序列展现的公平性,该评估仅基于输出结果,与所用提示词及模型无关;其二是生成式人工智能模型的内在公平性,要求当输入提示词为中性——即未明确指令生成式人工智能产生特定类型输出时——必须体现公平性。我们还研究了相对交叉公平性,以应对同时考虑多重类别时公平性面临的组合爆炸问题,并引入惰性公平性执行机制。最后,针对当前部分生成式人工智能模型进行了公平性监测与执行的测试。