We present SRC4VC, a new corpus containing 11 hours of speech recorded on smartphones by 100 Japanese speakers. Although high-quality multi-speaker corpora can advance voice conversion (VC) technologies, they are not always suitable for testing VC when low-quality speech recording is given as the input. To this end, we first asked 100 crowdworkers to record their voice samples using smartphones. Then, we annotated the recorded samples with speaker-wise recording-quality scores and utterance-wise perceived emotion labels. We also benchmark SRC4VC on any-to-any VC, in which we trained a multi-speaker VC model on high-quality speech and used the SRC4VC speakers' voice samples as the source in VC. The results show that the recording quality mismatch between the training and evaluation data significantly degrades the VC performance, which can be improved by applying speech enhancement to the low-quality source speech samples.
翻译:本文介绍了SRC4VC,这是一个包含100名日语说话人通过智能手机录制的总计11小时语音的新语料库。尽管高质量的多说话人语料库能够推动语音转换技术的发展,但在输入为低质量语音录音时,它们并不总是适用于测试语音转换系统。为此,我们首先邀请了100名众包工作人员使用智能手机录制其语音样本。随后,我们对录制的样本进行了标注,包括按说话人划分的录音质量评分和按话语划分的感知情感标签。我们还在任意对任意语音转换任务上对SRC4VC进行了基准测试,其中我们在高质量语音上训练了一个多说话人语音转换模型,并将SRC4VC说话人的语音样本用作语音转换的源语音。结果表明,训练数据与评估数据之间的录音质量不匹配会显著降低语音转换性能,而通过对低质量源语音样本应用语音增强技术可以改善此问题。