This study focuses on emotion-sensitive spoken dialogue in human-machine speech interaction. With the advancement of Large Language Models (LLMs), dialogue systems can handle multimodal data, including audio. Recent models have enhanced the understanding of complex audio signals through the integration of various audio events. However, they are unable to generate appropriate responses based on emotional speech. To address this, we introduce the Emotional chat Model (E-chat), a novel spoken dialogue system capable of comprehending and responding to emotions conveyed from speech. This model leverages an emotion embedding extracted by a speech encoder, combined with LLMs, enabling it to respond according to different emotional contexts. Additionally, we introduce the E-chat200 dataset, designed explicitly for emotion-sensitive spoken dialogue. In various evaluation metrics, E-chat consistently outperforms baseline LLMs, demonstrating its potential in emotional comprehension and human-machine interaction.
翻译:本研究聚焦于人机语音交互中的情感敏感型口语对话。随着大语言模型(LLMs)的发展,对话系统已能处理包括音频在内的多模态数据。近期模型通过整合多种音频事件,增强了对复杂音频信号的理解能力。然而,这些模型仍无法基于情感语音生成恰当响应。为解决这一问题,我们提出了情感对话模型(E-chat),这是一种能够理解并回应语音中传达的情感的新型语音对话系统。该模型利用语音编码器提取的情感嵌入,结合大语言模型,使其能够根据不同的情感语境进行响应。此外,我们专门为情感敏感型口语对话构建了E-chat200数据集。在多项评估指标中,E-chat始终优于基线大语言模型,展示了其在情感理解与人机交互方面的潜力。