In recent years, the burgeoning interest in diffusion models has led to significant advances in image and speech generation. Nevertheless, the direct synthesis of music waveforms from unrestricted textual prompts remains a relatively underexplored domain. In response to this lacuna, this paper introduces a pioneering contribution in the form of a text-to-waveform music generation model, underpinned by the utilization of diffusion models. Our methodology hinges on the innovative incorporation of free-form textual prompts as conditional factors to guide the waveform generation process within the diffusion model framework. Addressing the challenge of limited text-music parallel data, we undertake the creation of a dataset by harnessing web resources, a task facilitated by weak supervision techniques. Furthermore, a rigorous empirical inquiry is undertaken to contrast the efficacy of two distinct prompt formats for text conditioning, namely, music tags and unconstrained textual descriptions. The outcomes of this comparative analysis affirm the superior performance of our proposed model in terms of enhancing text-music relevance. Finally, our work culminates in a demonstrative exhibition of the excellent capabilities of our model in text-to-music generation. We further demonstrate that our generated music in the waveform domain outperforms previous works by a large margin in terms of diversity, quality, and text-music relevance.
翻译:近年来,扩散模型日益引起关注,推动了图像和语音生成的重大进展。然而,从无限制文本提示直接合成音乐波形仍是一个相对未充分探索的领域。针对这一空白,本文提出了一项开创性贡献——一种基于扩散模型的文本到波形音乐生成模型。我们的方法核心在于创新性地将自由形式的文本提示作为条件因素,用于引导扩散模型框架内的波形生成过程。为应对文本-音乐配对数据有限的挑战,我们利用网络资源构建数据集,并通过弱监督技术辅助完成。此外,我们开展了严格的实证研究,以对比两种不同提示格式(即音乐标签与无约束文本描述)在文本条件化中的效果。比较分析结果证实,所提模型在提升文本-音乐相关性方面表现更优。最后,本文通过示范性展示证明了模型在文本到音乐生成中的卓越能力。进一步表明,我们生成的波形域音乐在多样性、质量及文本-音乐相关性上大幅超越了先前工作。