Text-to-image generation models have recently achieved astonishing results in image quality, flexibility, and text alignment and are consequently employed in a fast-growing number of applications. Through improvements in multilingual abilities, a larger community now has access to this kind of technology. Yet, as we will show, multilingual models suffer similarly from (gender) biases as monolingual models. Furthermore, the natural expectation is that these models will provide similar results across languages, but this is not the case and there are important differences between languages. Thus, we propose a novel benchmark MAGBIG intending to foster research in multilingual models without gender bias. We investigate whether multilingual T2I models magnify gender bias with MAGBIG. To this end, we use multilingual prompts requesting portrait images of persons of a certain occupation or trait (using adjectives). Our results show not only that models deviate from the normative assumption that each gender should be equally likely to be generated, but that there are also big differences across languages. Furthermore, we investigate prompt engineering strategies, i.e. the use of indirect, neutral formulations, as a possible remedy for these biases. Unfortunately, they help only to a limited extent and result in worse text-to-image alignment. Consequently, this work calls for more research into diverse representations across languages in image generators.
翻译:文本到图像生成模型近期在图像质量、灵活性与文本对齐方面取得了惊人成果,因而被广泛应用于快速增长的应用场景中。随着多语言能力的提升,更多用户群体得以接触此类技术。然而,我们将证明多语言模型与单语言模型同样存在(性别)偏见问题。更令人意外的是,尽管用户自然期望这些模型在不同语言中输出相似结果,实际情况却非如此——语言间存在显著差异。为此,我们提出新型基准数据集MAGBIG,旨在推动无性别偏见的 multilingual 模型研究。通过MAGBIG,我们探究多语言T2I模型是否放大性别偏见:使用多语言提示词请求生成特定职业或特质(以形容词描述)的人物肖像图像。研究结果不仅显示模型偏离了"各性别生成概率应均等"的规范性假设,更暴露出跨语言间的巨大差异。此外,我们考察了提示工程策略(即使用间接中性表述)作为偏见补救措施的可能性。遗憾的是,这些策略效果有限,反而导致文本-图像对齐质量下降。因此,本研究呼吁加强对图像生成器中跨语言多样化表征的研究。