Text-based speech editing (TSE) techniques are designed to enable users to edit the output audio by modifying the input text transcript instead of the audio itself. Despite much progress in neural network-based TSE techniques, the current techniques have focused on reducing the difference between the generated speech segment and the reference target in the editing region, ignoring its local and global fluency in the context and original utterance. To maintain the speech fluency, we propose a fluency speech editing model, termed \textit{FluentEditor}, by considering fluency-aware training criterion in the TSE training. Specifically, the \textit{acoustic consistency constraint} aims to smooth the transition between the edited region and its neighboring acoustic segments consistent with the ground truth, while the \textit{prosody consistency constraint} seeks to ensure that the prosody attributes within the edited regions remain consistent with the overall style of the original utterance. The subjective and objective experimental results on VCTK demonstrate that our \textit{FluentEditor} outperforms all advanced baselines in terms of naturalness and fluency. The audio samples and code are available at \url{https://github.com/Ai-S2-Lab/FluentEditor}.
翻译:基于文本的语音编辑(TSE)技术旨在使用户能够通过修改输入文本转录而非音频本身来编辑输出音频。尽管基于神经网络的TSE技术取得了诸多进展,但当前技术主要集中在减少编辑区域内生成语音片段与参考目标之间的差异,忽视了其在语境和原始语句中的局部与整体流畅性。为保持语音流畅性,我们提出了一种流畅性语音编辑模型,称为FluentEditor,通过在TSE训练中引入流畅性感知训练准则。具体而言,“声学一致性约束”旨在平滑编辑区域及其邻近声学片段之间的过渡,使其与真实情况一致;而“韵律一致性约束”则力求确保编辑区域内的韵律属性与原始语句的整体风格保持一致。在VCTK上的主观与客观实验结果表明,我们的FluentEditor在自然度和流畅性方面优于所有先进基线方法。音频样本和代码可在\url{https://github.com/Ai-S2-Lab/FluentEditor}获取。