Summarisation of research results in plain language is crucial for promoting public understanding of research findings. The use of Natural Language Processing to generate lay summaries has the potential to relieve researchers' workload and bridge the gap between science and society. The aim of this narrative literature review is to describe and compare the different text summarisation approaches used to generate lay summaries. We searched the databases Web of Science, Google Scholar, IEEE Xplore, Association for Computing Machinery Digital Library and arXiv for articles published until 6 May 2022. We included original studies on automatic text summarisation methods to generate lay summaries. We screened 82 articles and included eight relevant papers published between 2020 and 2021, all using the same dataset. The results show that transformer-based methods such as Bidirectional Encoder Representations from Transformers (BERT) and Pre-training with Extracted Gap-sentences for Abstractive Summarization (PEGASUS) dominate the landscape of lay text summarisation, with all but one study using these methods. A combination of extractive and abstractive summarisation methods in a hybrid approach was found to be most effective. Furthermore, pre-processing approaches to input text (e.g. applying extractive summarisation) or determining which sections of a text to include, appear critical. Evaluation metrics such as Recall-Oriented Understudy for Gisting Evaluation (ROUGE) were used, which do not consider readability. To conclude, automatic lay text summarisation is under-explored. Future research should consider long document lay text summarisation, including clinical trial reports, and the development of evaluation metrics that consider readability of the lay summary.
翻译:以通俗语言总结研究成果对于促进公众理解研究发现至关重要。利用自然语言处理生成通俗摘要具有减轻研究人员工作负担并弥合科学与公众鸿沟的潜力。本叙述性文献综述旨在描述并比较用于生成通俗摘要的不同文本摘要方法。我们检索了Web of Science、Google Scholar、IEEE Xplore、美国计算机协会数字图书馆和arXiv数据库中截至2022年5月6日发表的文章,纳入采用自动文本摘要方法生成通俗摘要的原创研究。经筛选82篇文章后,最终纳入8篇发表于2020至2021年间的相关论文,这些论文均使用同一数据集。结果表明,基于Transformer的方法(如BERT和PEGASUS)主导了通俗文本摘要领域,除一项研究外均采用此类方法。混合式抽取式与生成式摘要方法的组合被证实最为有效。此外,输入文本的预处理策略(例如应用抽取式摘要)或确定文本中应包含的章节显得至关重要。现有评估指标(如ROUGE)未考虑可读性。总之,自动通俗文本摘要的研究尚不充分。未来研究应关注长文档的通俗文本摘要(包括临床试验报告),并开发能评估通俗摘要可读性的评价指标。