Modern noise-cancelling headphones have significantly improved users' auditory experiences by removing unwanted background noise, but they can also block out sounds that matter to users. Machine learning (ML) models for sound event detection (SED) and speaker identification (SID) can enable headphones to selectively pass through important sounds; however, implementing these models for a user-centric experience presents several unique challenges. First, most people spend limited time customizing their headphones, so the sound detection should work reasonably well out of the box. Second, the models should be able to learn over time the specific sounds that are important to users based on their implicit and explicit interactions. Finally, such models should have a small memory footprint to run on low-power headphones with limited on-chip memory. In this paper, we propose addressing these challenges using HiSSNet (Hierarchical SED and SID Network). HiSSNet is an SEID (SED and SID) model that uses a hierarchical prototypical network to detect both general and specific sounds of interest and characterize both alarm-like and speech sounds. We show that HiSSNet outperforms an SEID model trained using non-hierarchical prototypical networks by 6.9 - 8.6 percent. When compared to state-of-the-art (SOTA) models trained specifically for SED or SID alone, HiSSNet achieves similar or better performance while reducing the memory footprint required to support multiple capabilities on-device.
翻译:现代降噪耳机通过消除不必要的背景噪声显著提升了用户的听觉体验,但它们也可能屏蔽对用户重要的声音。用于声音事件检测(SED)和说话人识别(SID)的机器学习模型可使耳机选择性通过重要声音;然而,为提供以用户为中心的体验而实现这些模型面临若干独特的挑战。首先,大多数人很少花时间定制耳机,因此声音检测应在开箱即用时具有合理表现。其次,模型应能根据用户的隐式和显式交互,随时间学习用户关注的具体声音。最后,此类模型应具有较小的内存占用,以便在片上内存受限的低功耗耳机上运行。本文提出利用HiSSNet(层次化SED与SID网络)应对这些挑战。HiSSNet是一种SEID(SED与SID)模型,通过层次原型网络同时检测通用与特定关注声音,并表征警报类声音和语音。研究表明,相较于使用非层次原型网络训练的SEID模型,HiSSNet的性能提升了6.9%至8.6%。与专门为单独SED或SID任务训练的最新(SOTA)模型相比,HiSSNet在实现相似或更优性能的同时,减少了支持设备上多能力所需的内存占用。