Large Language Models (LLMs) exhibit powerful summarization abilities. However, their capabilities on conversational summarization remains under explored. In this work we evaluate LLMs (approx. 10 billion parameters) on conversational summarization and showcase their performance on various prompts. We show that the summaries generated by models depend on the instructions and the performance of LLMs vary with different instructions sometimes resulting steep drop in ROUGE scores if prompts are not selected carefully. We also evaluate the models with human evaluations and discuss the limitations of the models on conversational summarization
翻译:大型语言模型展现出强大的摘要生成能力。然而,它们在对话摘要方面的能力仍未得到充分探索。在本研究中,我们评估了大型语言模型(约100亿参数)在对话摘要任务上的表现,并展示了它们在不同提示下的性能。研究表明,模型生成的摘要依赖于指令,且大型语言模型的性能会随指令变化而波动,若提示选择不当,ROUGE分数可能显著下降。我们还通过人工评估对模型进行了评价,并讨论了模型在对话摘要任务上的局限性。