Despite the success of end-to-end (E2E) spoken dialogue systems, maintaining strict context adherence in multi-round conversations remains a challenge. While prior works attribute these failures to models forgetting dialogue history, we highlight an equally critical but overlooked bottleneck: a gap between latent context awareness and active adherence. Although models internally recognize relevant past utterances, strong parametric priors often overshadow these signals during decoding. To bridge this gap, we propose an audio-adapted Context-Aware Decoding (CAD) approach. By leveraging internal attention mechanisms to isolate key historical rounds, our approach contrasts output distributions with and without this key context during inference, directly amplifying multimodal contextual signals. Evaluations on the Audio MultiChallenge benchmark demonstrate significant improvements in Semantic Memory and Self Coherence subtasks, successfully enforcing strict, context-faithful adherence.
翻译:尽管端到端(E2E)口语对话系统取得了成功,但在多轮对话中保持严格的上下文遵循仍然是一个挑战。虽然先前的研究将此类失败归因于模型遗忘对话历史,我们强调了一个同样关键但被忽视的瓶颈:潜在上下文感知与主动遵循之间的鸿沟。尽管模型内部能够识别相关的过往话语,但强大的参数先验常在解码过程中遮蔽这些信号。为弥合这一鸿沟,我们提出了一种音频适配的上下文感知解码(CAD)方法。通过利用内部注意力机制隔离关键历史轮次,我们的方法在推理时对比包含与不包含该关键上下文的输出分布,从而直接放大多模态上下文信号。在Audio MultiChallenge基准上的评估表明,该方法在语义记忆和自我一致性子任务上取得了显著改进,成功实现了严格的、符合上下文的遵循。