User-defined keyword spotting (KWS) is crucial for personalized voice interaction, yet existing methods face several challenges: (1) insufficient discriminability among confusable words, (2) performance inconsistency across speakers with varying pronunciations, and (3) high data cost to ensure reliable wake-word performance. In this paper, we introduce DMA-KWS, an efficient and robust framework for user-defined keyword spotting. First, it adopts a dual-stage matching pipeline: CTC decoding with streaming phoneme search to locate candidate segments, followed by QbyT with a phoneme matcher for fine-grained verification, enabling it to better distinguish confusable words. Next, multi-modal enrollment fuses user-specific speech with text embeddings to further improve accuracy for registered users. Finally, a parameter-efficient continual adaptation mechanism performs lightweight updates using synthetic and real data. Extensive experiments demonstrate the superior performance of DMA-KWS. On the LibriPhrase Hard subset, it achieves 97.85% AUC and 6.13% EER, reaching state-of-the-art performance. In speaker-dependent settings, DMA-KWS consistently outperforms text-only enrollment, demonstrating significant performance gains. Moreover, the proposed parameter-efficient fine-tuning mechanism adapts DMA-KWS with only 187k updated parameters, further enhancing KWS performance while ensuring suitability for on-device deployment.
翻译:用户自定义关键词检测对于个性化语音交互至关重要,然而现有方法面临若干挑战:(1) 易混淆词之间区分性不足,(2) 不同发音习惯的说话者间性能不一致,(3) 高数据成本难以保障可靠的唤醒词性能。本文提出DMA-KWS,一种高效稳健的用户自定义关键词检测框架。首先,它采用双阶段匹配流水线:基于CTC解码的流式音素搜索定位候选片段,接着通过带音素匹配器的QbyT进行细粒度验证,从而更好区分混淆词。其次,多模态注册融合用户特定语音与文本嵌入,进一步提升注册用户准确性。最后,参数高效的持续自适应机制利用合成数据和真实数据进行轻量级更新。大量实验证明了DMA-KWS的卓越性能。在LibriPhrase Hard子集上,该方法达到97.85%的AUC和6.13%的EER,实现当前最优性能。在说话者相关设置中,DMA-KWS持续优于纯文本注册方案,展现出显著性能提升。此外,所提出的参数高效微调机制仅需更新18.7万个参数即可实现DMA-KWS自适应,在确保设备端部署适用性的同时进一步提升KWS性能。