We generalize the notion of social biases from language embeddings to grounded vision and language embeddings. Biases are present in grounded embeddings, and indeed seem to be equally or more significant than for ungrounded embeddings. This is despite the fact that vision and language can suffer from different biases, which one might hope could attenuate the biases in both. Multiple ways exist to generalize metrics measuring bias in word embeddings to this new setting. We introduce the space of generalizations (Grounded-WEAT and Grounded-SEAT) and demonstrate that three generalizations answer different yet important questions about how biases, language, and vision interact. These metrics are used on a new dataset, the first for grounded bias, created by augmenting extending standard linguistic bias benchmarks with 10,228 images from COCO, Conceptual Captions, and Google Images. Dataset construction is challenging because vision datasets are themselves very biased. The presence of these biases in systems will begin to have real-world consequences as they are deployed, making carefully measuring bias and then mitigating it critical to building a fair society.
翻译:我们将语言嵌入中的社会偏见概念推广到有根视觉与语言嵌入。有根嵌入中存在偏见,且其显著程度似乎不低于甚至高于无根嵌入。尽管视觉与语言可能承受不同的偏见(人们可能希望这种差异能削弱两者各自的偏见),但事实依然如此。存在多种将词嵌入偏见度量指标推广至这一新场景的方法。我们引入泛化空间(有根-WEAT和有根-SEAT),并论证三种泛化方法对回答偏见、语言与视觉如何交互的不同重要问题各有侧重。这些度量方法应用于首个专为有根偏见构建的新数据集——通过用COCO、概念描述和谷歌图像中的10,228张图像扩展标准语言偏见基准而构建。数据集构建具有挑战性,因为视觉数据集本身存在严重偏见。随着这些系统的部署,其中存在的偏见将开始产生现实影响,因此精心测量并缓解偏见对于建设公平社会至关重要。