While recent research advances in speaker diarization mostly focus on improving the quality of diarization results, there is also an increasing interest in improving the efficiency of diarization systems. In this paper, we demonstrate that a multi-stage clustering strategy that uses different clustering algorithms for input of different lengths can address multi-faceted challenges of on-device speaker diarization applications. Specifically, a fallback clusterer is used to handle short-form inputs; a main clusterer is used to handle medium-length inputs; and a pre-clusterer is used to compress long-form inputs before they are processed by the main clusterer. Both the main clusterer and the pre-clusterer can be configured with an upper bound of the computational complexity to adapt to devices with different resource constraints. This multi-stage clustering strategy is critical for streaming on-device speaker diarization systems, where the budgets of CPU, memory and battery are tight.
翻译:尽管近期说话人日志研究主要集中在提升日志质量上,但对提高系统效率的关注也与日俱增。本文证明,采用对不同长度输入应用不同聚类算法的多阶段聚类策略,可应对设备端说话人日志应用的多重挑战。具体而言,使用回退聚类器处理短时输入,主聚类器处理中等长度输入,预聚类器则在长时输入被主聚类器处理前进行压缩。主聚类器与预聚类器均可通过设置计算复杂度上界,适配不同资源约束的设备。这种多阶段聚类策略对CPU、内存及电池资源受限的流式设备端说话人日志系统至关重要。