A song is a combination of singing voice and accompaniment. However, existing works focus on singing voice synthesis and music generation independently. Little attention was paid to explore song synthesis. In this work, we propose a novel task called text-to-song synthesis which incorporating both vocals and accompaniments generation. We develop Melodist, a two-stage text-to-song method that consists of singing voice synthesis (SVS) and vocal-to-accompaniment (V2A) synthesis. Melodist leverages tri-tower contrastive pretraining to learn more effective text representation for controllable V2A synthesis. A Chinese song dataset mined from a music website is built up to alleviate data scarcity for our research. The evaluation results on our dataset demonstrate that Melodist can synthesize songs with comparable quality and style consistency. Audio samples can be found in https://text2songMelodist.github.io/Sample/.
翻译:歌曲是歌声与伴奏的结合体。然而,现有研究分别独立聚焦于歌声合成与音乐生成,鲜有工作探索歌曲合成。本文提出一项名为文本到歌曲合成的新任务,该任务同时包含人声与伴奏的生成。我们开发了Melodist方法,一种两阶段的文本到歌曲方法,由歌声合成和歌声到伴奏合成两个阶段构成。Melodist利用三塔对比预训练学习更有效的文本表示,以实现可控的歌声到伴奏合成。为缓解研究中的数据稀缺问题,我们从音乐网站挖掘构建了一个中文歌曲数据集。在该数据集上的评估结果表明,Melodist能够合成具有相当质量与风格一致性的歌曲。音频样本可在https://text2songMelodist.github.io/Sample/获取。