This study explores the capabilities of prompt-driven Large Language Models (LLMs) like ChatGPT and GPT-4 in adhering to human guidelines for dialogue summarization. Experiments employed DialogSum (English social conversations) and DECODA (French call center interactions), testing various prompts: including prompts from existing literature and those from human summarization guidelines, as well as a two-step prompt approach. Our findings indicate that GPT models often produce lengthy summaries and deviate from human summarization guidelines. However, using human guidelines as an intermediate step shows promise, outperforming direct word-length constraint prompts in some cases. The results reveal that GPT models exhibit unique stylistic tendencies in their summaries. While BERTScores did not dramatically decrease for GPT outputs suggesting semantic similarity to human references and specialised pre-trained models, ROUGE scores reveal grammatical and lexical disparities between GPT-generated and human-written summaries. These findings shed light on the capabilities and limitations of GPT models in following human instructions for dialogue summarization.
翻译:本研究旨在探索指令驱动型大语言模型(如ChatGPT和GPT-4)在遵循人类对话摘要指南方面的能力。实验采用DialogSum(英文社交对话)和DECODA(法语音视频中心交互)数据集,测试了多种提示策略:包括来自现有文献的提示、人类摘要指南提示以及两步式提示方法。研究结果表明,GPT模型生成的摘要通常较长,且偏离人类摘要指南。然而,将人类指南作为中间步骤的方法展现出潜力,在某些情况下优于直接的词长约束提示。结果揭示GPT模型在摘要生成中表现出独特的风格倾向。尽管GPT输出的BERTScores并未显著下降(表明其与人类参考摘要及专用预训练模型存在语义相似性),但ROUGE评分显示GPT生成摘要与人工撰写摘要之间存在语法和词汇层面的差异。这些发现揭示了GPT模型在遵循人类指令进行对话摘要时的能力与局限性。