In this paper, we explore audio-editing with non-rigid text edits. We show that the proposed editing pipeline is able to create audio edits that remain faithful to the input audio. We explore text prompts that perform addition, style transfer, and in-painting. We quantitatively and qualitatively show that the edits are able to obtain results which outperform Audio-LDM, a recently released text-prompted audio generation model. Qualitative inspection of the results points out that the edits given by our approach remain more faithful to the input audio in terms of keeping the original onsets and offsets of the audio events.
翻译:本文探讨了基于非刚性文本编辑的音频编辑技术。我们提出的编辑流程能够生成忠实于输入音频的编辑结果。实验考察了执行添加、风格迁移和修复功能的文本提示。通过定量与定性分析,证实本方法获得的编辑效果优于近期发布的文本提示音频生成模型Audio-LDM。定性评估结果表明,该方法在保持原始音频事件的起止时间方面,能更忠实地还原输入音频特征。