Current Spoken Dialogue Systems (SDSs) often serve as passive listeners that respond only after receiving user speech. To achieve human-like dialogue, we propose a novel future prediction architecture that allows an SDS to anticipate future affective reactions based on its current behaviors before the user speaks. In this work, we investigate two scenarios: speech and laughter. In speech, we propose to predict the user's future emotion based on its temporal relationship with the system's current emotion and its causal relationship with the system's current Dialogue Act (DA). In laughter, we propose to predict the occurrence and type of the user's laughter using the system's laughter behaviors in the current turn. Preliminary analysis of human-robot dialogue demonstrated synchronicity in the emotions and laughter displayed by the human and robot, as well as DA-emotion causality in their dialogue. This verifies that our architecture can contribute to the development of an anticipatory SDS.
翻译:当前的口语对话系统(SDS)往往扮演被动倾听者的角色,仅在接收到用户语音后才做出响应。为了实现类人对话,我们提出了一种新颖的未来预测架构,使口语对话系统能够在用户发言前,根据自身当前行为预判用户未来的情感反应。本研究探讨了两种场景:言语和笑声。在言语场景中,我们基于系统当前情感与用户未来情感的时间关联性,以及系统当前对话行为(DA)与用户情感之间的因果关系,预测用户未来的情感状态。在笑声场景中,我们利用系统在当前轮次中的笑声行为,预测用户笑声是否发生及其类型。对人机对话的初步分析显示,人类与机器人在情感和笑声表达上存在同步性,且其对话中对话行为与情感之间存在因果关系。这一发现验证了我们的架构有助于开发具有预见能力的口语对话系统。