We performed an experimental review of current diarization systems for the conversational telephone speech (CTS) domain. In detail, we considered a total of eight different algorithms belonging to clustering-based, end-to-end neural diarization (EEND), and speech separation guided diarization (SSGD) paradigms. We studied the inference-time computational requirements and diarization accuracy on four CTS datasets with different characteristics and languages. We found that, among all methods considered, EEND-vector clustering (EEND-VC) offers the best trade-off in terms of computing requirements and performance. More in general, EEND models have been found to be lighter and faster in inference compared to clustering-based methods. However, they also require a large amount of diarization-oriented annotated data. In particular EEND-VC performance in our experiments degraded when the dataset size was reduced, whereas self-attentive EEND (SA-EEND) was less affected. We also found that SA-EEND gives less consistent results among all the datasets compared to EEND-VC, with its performance degrading on long conversations with high speech sparsity. Clustering-based diarization systems, and in particular VBx, instead have more consistent performance compared to SA-EEND but are outperformed by EEND-VC. The gap with respect to this latter is reduced when overlap-aware clustering methods are considered. SSGD is the most computationally demanding method, but it could be convenient if speech recognition has to be performed. Its performance is close to SA-EEND but degrades significantly when the training and inference data characteristics are less matched.
翻译:我们对当前面向会话电话语音领域的说话人分割系统进行了实验性综述。具体而言,我们考虑了分属基于聚类的、端到端神经说话人分割(EEND)以及语音分离引导的说话人分割(SSGD)三类范式的总计八种不同算法。我们在四个具有不同特征和语言的会话电话语音数据集上研究了这些方法的推理时计算需求与分割准确性。结果发现,在所有被考虑的方法中,EEND向量聚类(EEND-VC)在计算需求与性能之间实现了最佳权衡。更普遍地讲,相较于基于聚类的方法,EEND模型在推理时更为轻量且快速,但其也需要大量面向说话人分割的标注数据。特别地,当数据集规模减小时,EEND-VC在我们的实验中的性能有所下降,而自注意力EEND(SA-EEND)则受此影响较小。我们还发现,与EEND-VC相比,SA-EEND在所有数据集上给出的结果一致性较差,其在语音稀疏度高的长对话中性能会下降。基于聚类的说话人分割系统(特别是VBx)相较于SA-EEND具有更一致的表现,但性能不如EEND-VC。当采用重叠感知聚类方法时,与后者的差距会缩小。SSGD对计算资源的需求最高,但若需要执行语音识别,则其可能更为便利。其性能与SA-EEND相近,但当训练与推理数据的特征匹配度较低时,其性能会显著下降。