Speaker diarization has gained considerable attention within speech processing research community. Mainstream speaker diarization rely primarily on speakers' voice characteristics extracted from acoustic signals and often overlook the potential of semantic information. Considering the fact that speech signals can efficiently convey the content of a speech, it is of our interest to fully exploit these semantic cues utilizing language models. In this work we propose a novel approach to effectively leverage semantic information in clustering-based speaker diarization systems. Firstly, we introduce spoken language understanding modules to extract speaker-related semantic information and utilize these information to construct pairwise constraints. Secondly, we present a novel framework to integrate these constraints into the speaker diarization pipeline, enhancing the performance of the entire system. Extensive experiments conducted on the public dataset demonstrate the consistent superiority of our proposed approach over acoustic-only speaker diarization systems.
翻译:说话人日志在语音处理研究领域中引起了广泛关注。主流说话人日志方法主要依赖于从声学信号中提取的说话人声音特征,往往忽视了语义信息的潜力。考虑到语音信号能够有效传达讲话内容,我们感兴趣的是如何充分利用语言模型来挖掘这些语义线索。在本工作中,我们提出了一种新颖的方法,以有效利用基于聚类的说话人日志系统中的语义信息。首先,我们引入口语理解模块来提取与说话人相关的语义信息,并利用这些信息构建成对约束。其次,我们提出了一种新的框架,将这些约束集成到说话人日志流程中,从而提升整个系统的性能。在公开数据集上进行的大量实验表明,我们提出的方法相较于仅基于声学的说话人日志系统具有持续的优势。