Discrete audio tokens derived from self-supervised learning models have gained widespread usage in speech generation. However, current practice of directly utilizing audio tokens poses challenges for sequence modeling due to the length of the token sequence. Additionally, this approach places the burden on the model to establish correlations between tokens, further complicating the modeling process. To address this issue, we propose acoustic BPE which encodes frequent audio token patterns by utilizing byte-pair encoding. Acoustic BPE effectively reduces the sequence length and leverages the prior morphological information present in token sequence, which alleviates the modeling challenges of token correlation. Through comprehensive investigations on a speech language model trained with acoustic BPE, we confirm the notable advantages it offers, including faster inference and improved syntax capturing capabilities. In addition, we propose a novel rescore method to select the optimal synthetic speech among multiple candidates generated by rich-diversity TTS system. Experiments prove that rescore selection aligns closely with human preference, which highlights acoustic BPE's potential to other speech generation tasks.
翻译:从自监督学习模型中派生出的离散音频令牌在语音生成领域得到了广泛应用。然而,当前直接使用音频令牌的做法因令牌序列长度过长而给序列建模带来挑战。此外,这一方法迫使模型自行建立令牌间的关联,进一步增加了建模的复杂性。为解决此问题,我们提出声学BPE方法,通过字节对编码对频繁出现的音频令牌模式进行编码。声学BPE有效缩短了序列长度,并利用了令牌序列中蕴含的先验形态信息,从而缓解了令牌关联建模的困难。通过对基于声学BPE训练的语音语言模型进行系统研究,我们确认了其显著优势,包括更快的推理速度和更强的句法捕获能力。此外,我们提出了一种新的重打分方法,用于从高多样性TTS系统生成的多个候选语音中筛选最优合成结果。实验证明,该重打分选择与人类偏好高度一致,凸显了声学BPE在其他语音生成任务中的应用潜力。