Large Language Model (LLM)-enhanced agents become increasingly prevalent in Human-AI communication, offering vast potential from entertainment to professional domains. However, current multi-modal dialogue systems overlook the acoustic information present in speech, which is crucial for understanding human communication nuances. This oversight can lead to misinterpretations of speakers' intentions, resulting in inconsistent or even contradictory responses within dialogues. To bridge this gap, in this paper, we propose PerceptiveAgent, an empathetic multi-modal dialogue system designed to discern deeper or more subtle meanings beyond the literal interpretations of words through the integration of speech modality perception. Employing LLMs as a cognitive core, PerceptiveAgent perceives acoustic information from input speech and generates empathetic responses based on speaking styles described in natural language. Experimental results indicate that PerceptiveAgent excels in contextual understanding by accurately discerning the speakers' true intentions in scenarios where the linguistic meaning is either contrary to or inconsistent with the speaker's true feelings, producing more nuanced and expressive spoken dialogues. Code is publicly available at: \url{https://github.com/Haoqiu-Yan/PerceptiveAgent}.
翻译:大型语言模型(LLM)增强的智能体在人机交互中日益普及,展现出从娱乐到专业领域的巨大潜力。然而,当前的多模态对话系统忽视了语音中蕴含的听觉信息,而这对理解人类交流的细微差别至关重要。这种忽视可能导致对说话者意图的误解,从而在对话中产生不一致甚至矛盾的回答。为弥补这一不足,本文提出PerceptiveAgent——一个共情的多模态对话系统,旨在通过整合语音模态感知,辨别超越字面含义的更深层或更微妙的语义。该系统以LLM作为认知核心,从输入语音中感知听觉信息,并基于自然语言描述的说话风格生成共情回应。实验结果表明,在语言含义与说话者真实情感相悖或不一致的场景中,PerceptiveAgent能够通过准确辨别说话者的真实意图,在上下文理解方面表现优异,从而生成更细腻且富有表现力的口语对话。代码公开于:\url{https://github.com/Haoqiu-Yan/PerceptiveAgent}。