Image captioning models are known to perpetuate and amplify harmful societal bias in the training set. In this work, we aim to mitigate such gender bias in image captioning models. While prior work has addressed this problem by forcing models to focus on people to reduce gender misclassification, it conversely generates gender-stereotypical words at the expense of predicting the correct gender. From this observation, we hypothesize that there are two types of gender bias affecting image captioning models: 1) bias that exploits context to predict gender, and 2) bias in the probability of generating certain (often stereotypical) words because of gender. To mitigate both types of gender biases, we propose a framework, called LIBRA, that learns from synthetically biased samples to decrease both types of biases, correcting gender misclassification and changing gender-stereotypical words to more neutral ones.
翻译:图像描述模型已知会延续并放大训练集中有害的社会偏见。本研究旨在缓解图像描述模型中的此类性别偏见。虽然先前的工作通过强制模型关注人物以减少性别误分类来解决该问题,但这反而以牺牲正确性别预测为代价产生了性别刻板印象词汇。基于这一观察,我们提出假设:影响图像描述模型的性别偏见存在两种类型:1)利用上下文预测性别的偏见,以及2)因性别而产生的特定(通常为刻板印象)词汇生成概率的偏见。为缓解这两种性别偏见,我们提出一个名为LIBRA的框架,该框架通过从合成偏置样本中学习来降低两种偏见,从而修正性别误分类并将性别刻板印象词汇转换为更中性的表达。