Citation graphs can be helpful in generating high-quality summaries of scientific papers, where references of a scientific paper and their correlations can provide additional knowledge for contextualising its background and main contributions. Despite the promising contributions of citation graphs, it is still challenging to incorporate them into summarization tasks. This is due to the difficulty of accurately identifying and leveraging relevant content in references for a source paper, as well as capturing their correlations of different intensities. Existing methods either ignore references or utilize only abstracts indiscriminately from them, failing to tackle the challenge mentioned above. To fill that gap, we propose a novel citation-aware scientific paper summarization framework based on citation graphs, able to accurately locate and incorporate the salient contents from references, as well as capture varying relevance between source papers and their references. Specifically, we first build a domain-specific dataset PubMedCite with about 192K biomedical scientific papers and a large citation graph preserving 917K citation relationships between them. It is characterized by preserving the salient contents extracted from full texts of references, and the weighted correlation between the salient contents of references and the source paper. Based on it, we design a self-supervised citation-aware summarization framework (CitationSum) with graph contrastive learning, which boosts the summarization generation by efficiently fusing the salient information in references with source paper contents under the guidance of their correlations. Experimental results show that our model outperforms the state-of-the-art methods, due to efficiently leveraging the information of references and citation correlations.
翻译:摘要:引用图有助于生成高质量的科学论文摘要,其中参考文献及其关联性可为论文的背景与主要贡献提供额外知识。尽管引用图具有显著潜力,但将其整合到摘要生成任务中仍具挑战性。这主要源于难以准确识别并利用参考文献中与目标论文相关的内容,同时捕捉不同强度的引用关联。现有方法或完全忽略参考文献,或仅不加区分地使用其摘要,未能解决上述难题。为弥补这一空白,我们提出了一种基于引用图的新型引用感知科学论文摘要生成框架,能够准确定位并融合参考文献中的显著内容,同时捕捉目标论文与其参考文献间的动态相关性。具体而言,我们首先构建了领域专用数据集PubMedCite,包含约19.2万篇生物医学论文及其引用关系(共91.7万条)。该数据集的特点在于保留了从参考文献全文提取的显著内容,以及这些内容与目标论文之间的加权关联。基于此,我们设计了自监督的引用感知摘要生成框架CitationSum,通过图对比学习,在引用关系引导下高效融合参考文献的显著信息与目标论文内容,从而提升摘要质量。实验结果表明,由于充分利用了参考文献信息及引用关联,我们的模型性能优于现有最先进方法。