In a speech recognition system, voice activity detection (VAD) is a crucial frontend module. Addressing the issues of poor noise robustness in traditional binary VAD systems based on DFSMN, the paper further proposes semantic VAD based on multi-task learning with improved models for real-time and offline systems, to meet specific application requirements. Evaluations on internal datasets show that, compared to the real-time VAD system based on DFSMN, the real-time semantic VAD system based on RWKV achieves relative decreases in CER of 7.0\%, DCF of 26.1\% and relative improvement in NRR of 19.2\%. Similarly, when compared to the offline VAD system based on DFSMN, the offline VAD system based on SAN-M demonstrates relative decreases in CER of 4.4\%, DCF of 18.6\% and relative improvement in NRR of 3.5\%.
翻译:在语音识别系统中,语音激活检测(VAD)是至关重要的前端模块。针对传统基于DFSMN的二值VAD系统抗噪能力不足的问题,本文进一步提出基于多任务学习的语义VAD,通过改进模型分别应用于实时与离线系统,以满足特定应用需求。在内部数据集上的评估显示,相较于基于DFSMN的实时VAD系统,基于RWKV的实时语义VAD系统实现了字符错误率(CER)相对降低7.0%、检测代价函数(DCF)相对降低26.1%,以及噪声抑制比(NRR)相对提升19.2%。类似地,相较于基于DFSMN的离线VAD系统,基于SAN-M的离线VAD系统实现了CER相对降低4.4%、DCF相对降低18.6%,以及NRR相对提升3.5%。