Speaker anonymization systems continue to improve their ability to obfuscate the original speaker characteristics in a speech signal, but often create processing artifacts and unnatural sounding voices as a tradeoff. Many of those systems stem from the VoicePrivacy Challenge (VPC) Baseline B1, using a neural vocoder to synthesize speech from an F0, x-vectors and bottleneck features-based speech representation. Inspired by this, we investigate the reproduction capabilities of the aforementioned baseline, to assess how successful the shared methodology is in synthesizing human-like speech. We use four objective metrics to measure speech quality, waveform similarity, and F0 similarity. Our findings indicate that both the speech representation and the vocoder introduces artifacts, causing an unnatural perception. A MUSHRA-like listening test on 18 subjects corroborate our findings, motivating further research on the analysis and synthesis components of the VPC Baseline B1.
翻译:说话人匿名化系统在模糊语音信号中原始说话人特征的能力上持续进步,但常常以产生处理伪影和不自然的声音为代价。其中许多系统源自语音隐私挑战(VPC)基线B1,该基线使用神经声码器从基于F0、x向量和瓶颈特征的语音表征中合成语音。受此启发,我们研究了上述基线的复现能力,以评估该共享方法在合成类人语音方面的成功程度。我们采用四个客观指标来测量语音质量、波形相似度和F0相似度。研究结果表明,语音表征和声码器均引入了伪影,导致不自然的听觉感知。基于18名受试者的MUSHRA听力测试验证了我们的发现,这激励了对VPC基线B1分析与合成组件的进一步研究。