A huge number of multi-participant dialogues happen online every day, which leads to difficulty in understanding the nature of dialogue dynamics for both humans and machines. Dialogue disentanglement aims at separating an entangled dialogue into detached sessions, thus increasing the readability of long disordered dialogue. Previous studies mainly focus on message-pair classification and clustering in two-step methods, which cannot guarantee the whole clustering performance in a dialogue. To address this challenge, we propose a simple yet effective model named CluCDD, which aggregates utterances by contrastive learning. More specifically, our model pulls utterances in the same session together and pushes away utterances in different ones. Then a clustering method is adopted to generate predicted clustering labels. Comprehensive experiments conducted on the Movie Dialogue dataset and IRC dataset demonstrate that our model achieves a new state-of-the-art result.
翻译:摘要:每天在线发生大量多参与者对话,这导致人类和机器都难以理解对话动态的本质。对话解缠旨在将纠缠的对话分离为独立的会话片段,从而提高长时无序对话的可读性。以往研究主要关注两步方法中的消息对分类与聚类,但此类方法无法保证对话中整体聚类性能。为解决这一挑战,我们提出一个简单而有效的模型——CluCDD,通过对比学习聚合话语。具体而言,我们的模型将同一会话中的话语拉近,同时推远不同会话中的话语。随后采用聚类方法生成预测聚类标签。在电影对话数据集和IRC数据集上进行综合实验表明,我们的模型取得了新的最先进结果。