Opinion summarisation aims to summarise the salient information and opinions presented in documents such as product reviews, discussion forums, and social media texts into short summaries that enable users to effectively understand the opinions therein. Generating biased summaries has the risk of potentially swaying public opinion. Previous studies focused on studying bias in opinion summarisation using extractive models, but limited research has paid attention to abstractive summarisation models. In this study, using political bias as a case study, we first establish a methodology to quantify bias in abstractive models, then trace it from the pre-trained models to the task of summarising social media opinions using different models and adaptation methods. We find that most models exhibit intrinsic bias. Using a social media text summarisation dataset and contrasting various adaptation methods, we find that tuning a smaller number of parameters is less biased compared to standard fine-tuning; however, the diversity of topics in training data used for fine-tuning is critical.
翻译:意见摘要旨在将产品评论、论坛讨论和社交媒体文本等文档中的关键信息和观点提炼为简短摘要,使用户能够有效理解其中的内容。生成带有偏见的摘要可能带来影响公众舆论的风险。以往研究主要聚焦于使用提取式模型探究意见摘要中的偏见,但针对抽象式摘要模型的关注相对有限。本研究以政治偏见为例,首先建立了一种量化抽象式模型偏见的分析方法,继而追溯偏见从预训练模型到不同模型及适配方法下的社交媒体意见摘要任务中的表现。研究发现,大多数模型存在内在偏见。通过使用社交媒体文本摘要数据集并对比多种适配方法,我们发现:与标准微调相比,调整较少参数的方法产生的偏见更少;然而,微调训练数据主题的多样性对偏见程度具有关键影响。