Speaker extraction (SE) aims to segregate the speech of a target speaker from a mixture of interfering speakers with the help of auxiliary information. Several forms of auxiliary information have been employed in single-channel SE, such as a speech snippet enrolled from the target speaker or visual information corresponding to the spoken utterance. The effectiveness of the auxiliary information in SE is typically evaluated by comparing the extraction performance of SE with uninformed speaker separation (SS) methods. Following this evaluation protocol, many SE studies have reported performance improvement compared to SS, attributing this to the auxiliary information. However, such studies have been conducted on a few datasets and have not considered recent deep neural network architectures for SS that have shown impressive separation performance. In this paper, we examine the role of the auxiliary information in SE for different input scenarios and over multiple datasets. Specifically, we compare the performance of two SE systems (audio-based and video-based) with SS using a common framework that utilizes the recently proposed dual-path recurrent neural network as the main learning machine. Experimental evaluation on various datasets demonstrates that the use of auxiliary information in the considered SE systems does not always lead to better extraction performance compared to the uninformed SS system. Furthermore, we offer insights into the behavior of the SE systems when provided with different and distorted auxiliary information given the same mixture input.
翻译:说话人提取(SE)旨在借助辅助信息,从干扰说话人的混合语音中分离出目标说话人的语音。在单通道SE中已采用多种形式的辅助信息,例如从目标说话人处录制的语音片段、或与口语句子相对应的视觉信息。通常通过将SE的提取性能与无信息引导的说话人分离(SS)方法进行比较,来评估SE中辅助信息的有效性。遵循这一评估协议,许多SE研究报道了相较于SS的性能提升,并将其归因于辅助信息。然而,这些研究仅基于少量数据集,且未考虑近期在SS中展现出卓越分离性能的深度神经网络架构。本文针对不同输入场景及多个数据集,考察了SE中辅助信息的作用。具体而言,我们采用基于最近提出的双路径递归神经网络作为主要学习机器的通用框架,比较了两种SE系统(基于音频和基于视频)与SS的性能。多组数据集的实验评估表明,相较于无信息引导的SS系统,所考虑的SE系统中辅助信息的引入并非总能带来更优的提取性能。此外,我们还针对给定相同混合输入时,不同失真程度的辅助信息下SE系统的行为提供了相关见解。