In this paper, we propose MM-KWS, a novel approach to user-defined keyword spotting leveraging multi-modal enrollments of text and speech templates. Unlike previous methods that focus solely on either text or speech features, MM-KWS extracts phoneme, text, and speech embeddings from both modalities. These embeddings are then compared with the query speech embedding to detect the target keywords. To ensure the applicability of MM-KWS across diverse languages, we utilize a feature extractor incorporating several multilingual pre-trained models. Subsequently, we validate its effectiveness on Mandarin and English tasks. In addition, we have integrated advanced data augmentation tools for hard case mining to enhance MM-KWS in distinguishing confusable words. Experimental results on the LibriPhrase and WenetPhrase datasets demonstrate that MM-KWS outperforms prior methods significantly.
翻译:本文提出MM-KWS,一种利用文本与语音模板进行多模态注册的用户自定义关键词检测新方法。与以往仅关注文本或语音特征的方法不同,MM-KWS从两种模态中提取音素、文本及语音嵌入表示,随后通过比对查询语音嵌入以检测目标关键词。为确保MM-KWS在不同语言间的适用性,我们采用融合多种多语言预训练模型的特征提取器,并分别在汉语和英语任务上验证其有效性。此外,我们集成了先进的数据增强工具进行困难样本挖掘,以提升MM-KWS对易混淆词汇的区分能力。在LibriPhrase与WenetPhrase数据集上的实验结果表明,MM-KWS显著优于现有方法。