Unsupervised extractive summarization aims to extract salient sentences from a document as the summary without labeled data. Recent literatures mostly research how to leverage sentence similarity to rank sentences in the order of salience. However, sentence similarity estimation using pre-trained language models mostly takes little account of document-level information and has a weak correlation with sentence salience ranking. In this paper, we proposed two novel strategies to improve sentence similarity estimation for unsupervised extractive summarization. We use contrastive learning to optimize a document-level objective that sentences from the same document are more similar than those from different documents. Moreover, we use mutual learning to enhance the relationship between sentence similarity estimation and sentence salience ranking, where an extra signal amplifier is used to refine the pivotal information. Experimental results demonstrate the effectiveness of our strategies.
翻译:无监督抽取式摘要旨在无标注数据的情况下,从文档中提取关键句子作为摘要。近期文献大多研究如何利用句子相似度按显著性对句子进行排序。然而,使用预训练语言模型进行句子相似度评估时,通常较少考虑文档级信息,且与句子显著性排序的相关性较弱。本文提出了两种新策略以改进无监督抽取式摘要中的句子相似度评估。我们采用对比学习优化文档级目标,使同一文档的句子比不同文档的句子更为相似。此外,我们利用互学习增强句子相似度评估与句子显著性排序之间的关系,通过额外信号放大器精炼关键信息。实验结果表明了所提策略的有效性。