Due to their similarity-based learning objectives, pretrained sentence encoders often internalize stereotypical assumptions that reflect the social biases that exist within their training corpora. In this paper, we describe several kinds of stereotypes concerning different communities that are present in popular sentence representation models, including pretrained next sentence prediction and contrastive sentence representation models. We compare such models to textual entailment models that learn language logic for a variety of downstream language understanding tasks. By comparing strong pretrained models based on text similarity with textual entailment learning, we conclude that the explicit logic learning with textual entailment can significantly reduce bias and improve the recognition of social communities, without an explicit de-biasing process
翻译:由于基于相似性的学习目标,预训练句子编码器常常内化反映其训练语料中社会偏见的刻板假设。本文描述了流行句子表征模型中存在的关于不同社群的多种刻板偏见类型,包括预训练下一句预测模型和对比句子表征模型。我们将此类模型与学习语言逻辑以处理多种下游语言理解任务的文本蕴含模型进行比较。通过比较基于文本相似性的强预训练模型与文本蕴含学习方法,我们得出结论:无需显式去偏过程,文本蕴含带来的显式逻辑学习即可显著减少偏见并提升对社会社群的识别能力。