Speech restoration (SR) is a task of converting degraded speech signals into high-quality ones. In this study, we propose a robust SR model called Miipher, and apply Miipher to a new SR application: increasing the amount of high-quality training data for speech generation by converting speech samples collected from the Web to studio-quality. To make our SR model robust against various degradation, we use (i) a speech representation extracted from w2v-BERT for the input feature, and (ii) a text representation extracted from transcripts via PnG-BERT as a linguistic conditioning feature. Experiments show that Miipher (i) is robust against various audio degradation and (ii) enable us to train a high-quality text-to-speech (TTS) model from restored speech samples collected from the Web. Audio samples are available at our demo page: google.github.io/df-conformer/miipher/
翻译:摘要:语音修复(Speech Restoration, SR)旨在将退化的语音信号转换为高质量信号。本研究提出了一种名为Miipher的鲁棒语音修复模型,并将其应用于一项新的SR任务:通过将网络收集的语音样本转换为录音室级质量,增加语音生成任务的高质量训练数据量。为提升SR模型对各种退化现象的鲁棒性,我们采用了(i)从w2v-BERT提取的语音表示作为输入特征,以及(ii)通过PnG-BERT从转录文本中提取的文本表示作为语言条件特征。实验表明,Miipher(i)对多种音频退化具有鲁棒性,且(ii)能够利用网络收集的修复语音样本训练出高质量文本转语音(TTS)模型。音频样本可在我们的演示页面获取:google.github.io/df-conformer/miipher/