This paper introduces CraBERT, a pre-trained phoneme encoder (PPEnc) designed for efficient pre-training in text-to-speech (TTS). CraBERT employs a cascade-fusion architecture and a subword-phoneme alignment algorithm to integrate representations from a pre-trained subword-level BERT into a phoneme-level BERT. This design provides prior word- and sentence-level information, reducing the amount of pre-training required by the phoneme encoder. Subjective listening evaluations show that CraBERT achieves MOS values comparable to existing PPEncs after approximately one epoch of pre-training, whereas the baselines in our comparison are pre-trained for approximately ten epochs. These results demonstrate that CraBERT can efficiently learn representations suitable for improving the perceived naturalness and prosody of synthesized speech.
翻译:摘要:本文提出CraBERT——一种专为文本到语音(TTS)高效预训练设计的预训练音素编码器(PPEnc)。CraBERT采用级联融合架构与子词-音素对齐算法,将预训练子词级BERT的表示集成至音素级BERT中。该设计提供了先验的词级与句级信息,降低了音素编码器所需的预训练数据量。主观听感评估表明,CraBERT在约一个训练周期(epoch)的预训练后可达到与现有PPEnc相当的MOS值,而对比基线需预训练约十个周期。这些结果证明,CraBERT能够高效学习适用于提升合成语音自然度与韵律感知效果的表示。