Pre-trained sentence representations are crucial for identifying significant sentences in unsupervised document extractive summarization. However, the traditional two-step paradigm of pre-training and sentence-ranking, creates a gap due to differing optimization objectives. To address this issue, we argue that utilizing pre-trained embeddings derived from a process specifically designed to optimize cohensive and distinctive sentence representations helps rank significant sentences. To do so, we propose a novel graph pre-training auto-encoder to obtain sentence embeddings by explicitly modelling intra-sentential distinctive features and inter-sentential cohesive features through sentence-word bipartite graphs. These pre-trained sentence representations are then utilized in a graph-based ranking algorithm for unsupervised summarization. Our method produces predominant performance for unsupervised summarization frameworks by providing summary-worthy sentence representations. It surpasses heavy BERT- or RoBERTa-based sentence representations in downstream tasks.
翻译:预训练句子表示对于无监督文档抽取式摘要中识别关键句子至关重要。然而,传统预训练与句子排序的两阶段范式因优化目标不同而产生语义鸿沟。为解决该问题,我们提出:利用专门优化共性与个性句子表示的过程生成的预训练嵌入,有助于关键句子排序。为此,我们提出一种新颖的图预训练自编码器,通过句子-词二分图显式建模句内区分性特征与句间连贯性特征,从而获取句子嵌入。这些预训练句子表示随后被用于基于图的排序算法实现无监督摘要。通过提供符合摘要质量的句子表示,本方法在无监督摘要框架中取得显著性能优势,并超越了基于BERT或RoBERTa的繁重句子表示在下游任务中的表现。