Most existing keyword spotting research focuses on conditions with slight or moderate noise. In this paper, we try to tackle a more challenging task: detecting keywords buried under strong interfering speech (10 times higher than the keyword in amplitude), and even worse, mixed with other keywords. We propose a novel Mix Training (MT) strategy that encourages the model to discover low-energy keywords from noisy and mixed speech. Experiments were conducted with a vanilla CNN and two EfficientNet (B0/B2) architectures. The results evaluated with the Google Speech Command dataset demonstrated that the proposed mix training approach is highly effective and outperforms standard data augmentation and mixup training.
翻译:现有的大多数关键词定位研究仅针对轻微或中等噪声条件。本文尝试解决更具挑战性的任务:在强干扰语音(振幅为关键词10倍以上)甚至与其他关键词混杂的恶劣环境中检测被掩盖的关键词。我们提出了一种新颖的混合训练策略,能够促使模型从含噪混合语音中发现低能量关键词。实验采用标准卷积神经网络及两种EfficientNet(B0/B2)架构进行验证,基于Google Speech Command数据集的评估结果表明,所提出的混合训练方法具有极高有效性,且性能显著优于标准数据增强与混合训练方法。