With the fast development of zero-shot text-to-speech technologies, it is possible to generate high-quality speech signals that are indistinguishable from the real ones. Speech editing, including speech insertion and replacement, appeals to researchers due to its potential applications. However, existing studies only considered clean speech scenarios. In real-world applications, the existence of environmental noise could significantly degrade the quality of generation. In this study, we propose a noise-resilient speech editing framework, SeamlessEdit, for noisy speech editing. SeamlessEdit adopts a frequency-band-aware noise suppression module and an in-content refinement strategy. It can well address the scenario where the frequency bands of voice and background noise are not separated. The proposed SeamlessEdit framework outperforms state-of-the-art approaches in multiple quantitative and qualitative evaluations.
翻译:随着零样本文本转语音技术的快速发展,生成与真实语音无法区分的高质量语音信号已成为可能。语音编辑(包括语音插入和替换)因其潜在应用而备受研究者关注。然而,现有研究仅考虑无噪声场景。在实际应用中,环境噪声的存在会显著降低生成质量。本研究提出一种抗噪声的语音编辑框架SeamlessEdit,用于噪声环境下的语音编辑。SeamlessEdit采用频带感知噪声抑制模块和上下文精炼策略,能有效处理语音与背景噪声频带未分离的情况。实验表明,所提出的SeamlessEdit框架在多项定量与定性评估中均优于现有最优方法。