Recent studies have highlighted the issue of stereotypical depictions for people of different identity groups in Text-to-Image (T2I) model generations. However, these existing approaches have several key limitations, including a noticeable lack of coverage of global identity groups in their evaluation, and the range of their associated stereotypes. Additionally, they often lack a critical distinction between inherently visual stereotypes, such as `underweight' or `sombrero', and culturally dependent stereotypes like `attractive' or `terrorist'. In this work, we address these limitations with a multifaceted approach that leverages existing textual resources to ground our evaluation of geo-cultural stereotypes in the generated images from T2I models. We employ existing stereotype benchmarks to identify and evaluate visual stereotypes at a global scale, spanning 135 nationality-based identity groups. We demonstrate that stereotypical attributes are thrice as likely to be present in images of these identities as compared to other attributes. We further investigate how disparately offensive the depictions of generated images are for different nationalities. Finally, through a detailed case study, we reveal how the 'default' representations of all identity groups have a stereotypical appearance. Moreover, for the Global South, images across different attributes are visually similar, even when explicitly prompted otherwise. CONTENT WARNING: Some examples may contain offensive stereotypes.
翻译:近期研究揭示了文本到图像(T2I)模型生成结果中对不同身份群体存在刻板印象描绘的问题。然而,现有方法存在若干关键局限,包括评估对象未能覆盖全球身份群体及其相关刻板印象范围。此外,这些研究常缺乏对固有视觉刻板印象(如"体重过轻"或"宽檐帽")与文化依赖型刻板印象(如"有魅力"或"恐怖分子")的关键性区分。本研究通过多维度方法解决上述局限,利用现有文本资源构建评估框架,系统分析T2I模型生成图像中的地理文化刻板印象。我们采用现有刻板印象基准测试,对覆盖135个国籍身份群体的全球范围视觉刻板印象进行识别与评估。研究证实,相较其他属性,刻板印象属性出现在这些身份群体图像中的概率高出三倍。我们进一步探究不同国籍群体的生成图像在冒犯性程度上的差异。最终通过详细案例分析,揭示所有身份群体的"默认"表征均具有刻板化外观。更值得注意的是,对于全球南方群体,即使明确提示其他属性,其不同属性下的图像在视觉上仍呈现高度相似性。内容警告:部分示例可能包含具有冒犯性的刻板印象。