This paper integrates graph-to-sequence into an end-to-end text-to-speech framework for syntax-aware modelling with syntactic information of input text. Specifically, the input text is parsed by a dependency parsing module to form a syntactic graph. The syntactic graph is then encoded by a graph encoder to extract the syntactic hidden information, which is concatenated with phoneme embedding and input to the alignment and flow-based decoding modules to generate the raw audio waveform. The model is experimented on two languages, English and Mandarin, using single-speaker, few samples of target speakers, and multi-speaker datasets, respectively. Experimental results show better prosodic consistency performance between input text and generated audio, and also get higher scores in the subjective prosodic evaluation, and show the ability of voice conversion. Besides, the efficiency of the model is largely boosted through the design of the AI chip operator with 5x acceleration.
翻译:本文将图到序列整合到端到端文本到语音框架中,利用输入文本的句法信息实现语法感知建模。具体而言,输入文本通过依存句法分析模块解析为句法图,随后由图编码器编码以提取句法隐藏信息,该信息与音素嵌入拼接后输入对齐和基于流的解码模块,最终生成原始音频波形。该模型在英语和普通话两种语言上进行了实验,分别使用了单说话人、目标说话人少量样本以及多说话人数据集。实验结果表明,输入文本与生成音频之间的韵律一致性性能更优,且在主观韵律评价中获得了更高分数,同时展示了语音转换能力。此外,通过AI芯片算子的设计,模型效率大幅提升,实现了5倍加速。