Speaker diarization, the task of segmenting an audio recording based on speaker identity, constitutes an important speech pre-processing step for several downstream applications. The conventional approach to diarization involves multiple steps of embedding extraction and clustering, which are often optimized in an isolated fashion. While end-to-end diarization systems attempt to learn a single model for the task, they are often cumbersome to train and require large supervised datasets. In this paper, we propose an end-to-end supervised hierarchical clustering algorithm based on graph neural networks (GNN), called End-to-end Supervised HierARchical Clustering (E-SHARC). The E-SHARC approach uses front-end mel-filterbank features as input and jointly learns an embedding extractor and the GNN clustering module, performing representation learning, metric learning, and clustering with end-to-end optimization. Further, with additional inputs from an external overlap detector, the E-SHARC approach is capable of predicting the speakers in the overlapping speech regions. The experimental evaluation on several benchmark datasets like AMI, VoxConverse and DISPLACE, illustrates that the proposed E-SHARC framework improves significantly over the state-of-art diarization systems.
翻译:说话人日志作为一项基于说话人身份对音频录音进行分割的任务,是多个下游应用中重要的语音预处理步骤。传统日志方法涉及嵌入提取和聚类的多步骤流程,这些步骤通常以孤立方式进行优化。尽管端到端日志系统试图为这一任务学习单一模型,但它们往往训练繁琐且需要大规模监督数据集。本文提出了一种基于图神经网络(GNN)的端到端监督分层聚类算法,称为端到端监督分层聚类(E-SHARC)。E-SHARC方法以前端梅尔滤波器组特征为输入,联合学习嵌入提取器和GNN聚类模块,通过端到端优化实现表示学习、度量学习和聚类。此外,通过引入外部重叠检测器的额外输入,E-SHARC方法能够预测重叠语音区域中的说话人。在AMI、VoxConverse和DISPLACE等多个基准数据集上的实验评估表明,所提出的E-SHARC框架显著优于当前最先进的日志系统。