Large Language Models (LLMs) have achieved significant success across various natural language processing (NLP) tasks, encompassing question-answering, summarization, and machine translation, among others. While LLMs excel in general tasks, their efficacy in domain-specific applications remains under exploration. Additionally, LLM-generated text sometimes exhibits issues like hallucination and disinformation. In this study, we assess LLMs' capability of producing concise survey articles within the computer science-NLP domain, focusing on 20 chosen topics. Automated evaluations indicate that GPT-4 outperforms GPT-3.5 when benchmarked against the ground truth. Furthermore, four human evaluators provide insights from six perspectives across four model configurations. Through case studies, we demonstrate that while GPT often yields commendable results, there are instances of shortcomings, such as incomplete information and the exhibition of lapses in factual accuracy.
翻译:大型语言模型(LLMs)已在多种自然语言处理任务中取得显著成功,涵盖问答、摘要及机器翻译等。尽管LLMs在通用任务中表现优异,但其在特定领域应用中的有效性仍需探索。此外,LLM生成的文本有时会出现幻觉和虚假信息等问题。在本研究中,我们评估了LLMs在计算机科学-自然语言处理领域生成简洁综述文章的能力,重点关注20个选定主题。自动评估表明,GPT-4在以真实数据为基准测试时优于GPT-3.5。此外,四名人类评估员从六个不同角度对四种模型配置提供了见解。通过案例研究,我们发现虽然GPT通常能产生值得赞许的结果,但也存在信息不完整、事实准确性缺失等不足之处。