Full-duplex spoken dialogue models (SDMs) can listen and speak simultaneously, enabling interaction dynamics closer to human conversation than turn-based systems. Inspired by neural coupling in human communication, we study how such models coordinate their internal representations during interaction. We simulate full-duplex dialogues between two instances of the pretrained \textit{Moshi} model under controlled conditions, manipulating channel noise and decoding bias. Synchronization is measured using Centered Kernel Alignment (CKA) across temporal lags, while anticipatory turn-taking cues are probed from delayed internal activations using causal LSTM models, from both speaker and listener perspectives. We find strong representational synchronization under no noise conditions, peaking near zero lag and degrading with noise, and we show that internal states encode anticipatory information that supports turn-taking prediction ahead of time.
翻译:全双工口语对话模型能够同时进行听与说,实现比基于话轮切换机制更接近人类对话的交互动态。受人类交流中神经耦合现象的启发,本研究探讨了此类模型在交互过程中如何协调其内部表征。我们在可控条件下模拟了两个预训练\textit{Moshi}模型实例间的全双工对话,并操控信道噪声与解码偏置。通过中心核对齐方法测量不同时间延迟下的表征同步性,并利用因果长短期记忆模型从说话者和听者两个视角探测延迟内部激活中蕴含的预期性话轮交替线索。研究发现:在无噪声条件下表征同步性最强,在零延迟附近达到峰值,且随噪声增加而退化;同时,内部状态编码了支持话轮提前预测的预期性信息。