Large language models (LLMs) have shown promise for automatic summarization but the reasons behind their successes are poorly understood. By conducting a human evaluation on ten LLMs across different pretraining methods, prompts, and model scales, we make two important observations. First, we find instruction tuning, and not model size, is the key to the LLM's zero-shot summarization capability. Second, existing studies have been limited by low-quality references, leading to underestimates of human performance and lower few-shot and finetuning performance. To better evaluate LLMs, we perform human evaluation over high-quality summaries we collect from freelance writers. Despite major stylistic differences such as the amount of paraphrasing, we find that LMM summaries are judged to be on par with human written summaries.
翻译:大型语言模型(LLMs)在自动摘要方面展现出潜力,但其成功的原因尚不明确。通过对十种不同预训练方法、提示词和模型规模的大型语言模型进行人工评估,我们得出两项重要发现。首先,我们发现指令微调(而非模型规模)是LLMs零样本摘要能力的关键。其次,现有研究因受限于低质量参考摘要,导致对人类性能的低估以及少样本和微调性能的下降。为更准确评估LLMs,我们对从自由撰稿人处收集的高质量摘要进行了人工评估。尽管在改写程度等风格特征上存在显著差异,但我们发现LLM生成的摘要与人工撰写的摘要质量相当。