We introduce VisoGender, a novel dataset for benchmarking gender bias in vision-language models. We focus on occupation-related biases within a hegemonic system of binary gender, inspired by Winograd and Winogender schemas, where each image is associated with a caption containing a pronoun relationship of subjects and objects in the scene. VisoGender is balanced by gender representation in professional roles, supporting bias evaluation in two ways: i) resolution bias, where we evaluate the difference between pronoun resolution accuracies for image subjects with gender presentations perceived as masculine versus feminine by human annotators and ii) retrieval bias, where we compare ratios of professionals perceived to have masculine and feminine gender presentations retrieved for a gender-neutral search query. We benchmark several state-of-the-art vision-language models and find that they demonstrate bias in resolving binary gender in complex scenes. While the direction and magnitude of gender bias depends on the task and the model being evaluated, captioning models are generally less biased than Vision-Language Encoders. Dataset and code are available at https://github.com/oxai/visogender
翻译:我们提出了VisoGender,一个用于基准测试视觉语言模型中性别偏见的新型数据集。受Winograd和Winogender模式的启发,我们在一个二元性别的霸权体系内聚焦于职业相关的偏见,其中每张图像配有一段包含场景中主体与客体代词关系的标题。VisoGender通过专业角色中的性别表征实现平衡,并支持两种方式的偏见评估:i) 指代消解偏见——评估人类标注者感知为男性化与女性化性别呈现的图像主体之间代词指代消解准确率的差异;ii) 检索偏见——比较针对中性性别搜索查询时,被感知为男性化与女性化性别呈现的专业人士的检索比率。我们对多个最先进的视觉语言模型进行了基准测试,发现它们在复杂场景中消解二元性别时表现出偏见。尽管性别偏见的指向和幅度取决于被评估的任务和模型,但图像描述模型通常比视觉语言编码器偏见更少。数据集和代码可在 https://github.com/oxai/visogender 获取。