Automatic Speaker Diarization (ASD) is an enabling technology with numerous applications, which deals with recordings of multiple speakers, raising special concerns in terms of privacy. In fact, in remote settings, where recordings are shared with a server, clients relinquish not only the privacy of their conversation, but also of all the information that can be inferred from their voices. However, to the best of our knowledge, the development of privacy-preserving ASD systems has been overlooked thus far. In this work, we tackle this problem using a combination of two cryptographic techniques, Secure Multiparty Computation (SMC) and Secure Modular Hashing, and apply them to the two main steps of a cascaded ASD system: speaker embedding extraction and agglomerative hierarchical clustering. Our system is able to achieve a reasonable trade-off between performance and efficiency, presenting real-time factors of 1.1 and 1.6, for two different SMC security settings.
翻译:自动说话人日志(ASD)是一项具有众多应用的关键技术,涉及多说话人录音,在隐私方面引发特别关切。实际上,在远程场景中,当录音被共享至服务器时,客户端不仅交出其对话内容的隐私,还放弃了可从其语音中推断的所有信息的隐私。然而,据我们所知,迄今为止,隐私保护ASD系统的开发一直被忽视。在本工作中,我们通过结合两种密码学技术——安全多方计算(SMC)和安全模块化哈希——来解决这一难题,并将其应用于级联ASD系统的两个主要步骤:说话人嵌入提取和凝聚层次聚类。我们的系统能够在性能与效率之间实现合理权衡,在两种不同的SMC安全设置下,分别呈现1.1和1.6的实时因子。