Growing literature has shown that NLP systems may encode social biases; however, the political bias of summarization models remains relatively unknown. In this work, we use an entity replacement method to investigate the portrayal of politicians in automatically generated summaries of news articles. We develop an entity-based computational framework to assess the sensitivities of several extractive and abstractive summarizers to the politicians Donald Trump and Joe Biden. We find consistent differences in these summaries upon entity replacement, such as reduced emphasis of Trump's presence in the context of the same article and a more individualistic representation of Trump with respect to the collective US government (i.e., administration). These summary dissimilarities are most prominent when the entity is heavily featured in the source article. Our characterization provides a foundation for future studies of bias in summarization and for normative discussions on the ideal qualities of automatic summaries.
翻译:日益增长的文献表明,自然语言处理系统可能编码社会偏见;然而,摘要模型的政治偏见仍相对未知。本研究采用实体替换方法,考察新闻文章自动生成摘要中政治人物的呈现方式。我们开发了一个基于实体的计算框架,评估几种抽取式和生成式摘要生成器对政客唐纳德·特朗普和乔·拜登的敏感度。研究发现,进行实体替换后,摘要中存在一致差异,例如同一文章背景下对特朗普存在的强调程度降低,且特朗普相对于美国集体政府(即行政机构)表现为更个人化的呈现方式。当源文章中该实体被大量提及时,这些摘要差异最为显著。本研究特征描绘为未来摘要偏见研究及自动摘要理想特质的规范性讨论奠定了基础。