Text summarization has been a crucial problem in natural language processing (NLP) for several decades. It aims to condense lengthy documents into shorter versions while retaining the most critical information. Various methods have been proposed for text summarization, including extractive and abstractive summarization. The emergence of large language models (LLMs) like GPT3 and ChatGPT has recently created significant interest in using these models for text summarization tasks. Recent studies \cite{goyal2022news, zhang2023benchmarking} have shown that LLMs-generated news summaries are already on par with humans. However, the performance of LLMs for more practical applications like aspect or query-based summaries is underexplored. To fill this gap, we conducted an evaluation of ChatGPT's performance on four widely used benchmark datasets, encompassing diverse summaries from Reddit posts, news articles, dialogue meetings, and stories. Our experiments reveal that ChatGPT's performance is comparable to traditional fine-tuning methods in terms of Rouge scores. Moreover, we highlight some unique differences between ChatGPT-generated summaries and human references, providing valuable insights into the superpower of ChatGPT for diverse text summarization tasks. Our findings call for new directions in this area, and we plan to conduct further research to systematically examine the characteristics of ChatGPT-generated summaries through extensive human evaluation.
翻译:文本摘要数十年来一直是自然语言处理(NLP)中的一个关键问题,其目标是将冗长文档压缩为简短版本,同时保留最关键的信息。目前已提出多种文本摘要方法,包括抽取式摘要和生成式摘要。近年来,随着GPT3和ChatGPT等大型语言模型(LLMs)的出现,利用这些模型进行文本摘要任务引起了广泛关注。近期研究\cite{goyal2022news, zhang2023benchmarking}表明,LLMs生成的新闻摘要已能与人类相媲美。然而,LLMs在方面或查询导向摘要等更实际应用中的表现仍待探索。为填补这一空白,我们在四个广泛使用的基准数据集(涵盖来自Reddit帖子、新闻文章、对话会议和故事的不同类型摘要)上评估了ChatGPT的性能。实验发现,在Rouge分数方面,ChatGPT的性能与传统微调方法相当。此外,我们揭示了ChatGPT生成摘要与人类参考摘要之间的一些独特差异,为理解ChatGPT在多样化文本摘要任务中的强大能力提供了宝贵见解。我们的发现为该领域指出了新方向,并计划进一步通过大规模人工评估系统性地探究ChatGPT生成摘要的特性。