Voice activity and overlapped speech detection (respectively VAD and OSD) are key pre-processing tasks for speaker diarization. The final segmentation performance highly relies on the robustness of these sub-tasks. Recent studies have shown VAD and OSD can be trained jointly using a multi-class classification model. However, these works are often restricted to a specific speech domain, lacking information about the generalization capacities of the systems. This paper proposes a complete and new benchmark of different VAD and OSD models, on multiple audio setups (single/multi-channel) and speech domains (e.g. media, meeting...). Our 2/3-class systems, which combine a Temporal Convolutional Network with speech representations adapted to the setup, outperform state-of-the-art results. We show that the joint training of these two tasks offers similar performances in terms of F1-score to two dedicated VAD and OSD systems while reducing the training cost. This unique architecture can also be used for single and multichannel speech processing.
翻译:语音活动检测(VAD)和重叠语音检测(OSD)是说话人日志化的关键预处理任务。最终分割性能高度依赖这些子任务的鲁棒性。近期研究表明,VAD和OSD可通过多类分类模型实现联合训练,但这类研究常局限于特定语音领域,缺乏对系统泛化能力的评估。本文提出全新的完整基准,涵盖多种音频配置(单/多通道)和语音领域(如媒体、会议等),对不同VAD与OSD模型进行评测。我们构建的2/3类系统结合时序卷积网络与适配音频配置的语音表征,在性能上超越了当前最优结果。研究表明,联合训练这两项任务在F1分数上可与专用VAD+OSD双系统相当,同时降低训练成本。这种统一架构还可同时应用于单通道与多通道语音处理。