While recent generative models can produce engaging music, their utility is limited. The variation in the music is often left to chance, resulting in compositions that lack structure. Pieces extending beyond a minute can become incoherent or repetitive. This paper introduces an approach for generating structured, arbitrarily long musical pieces. Central to this approach is the creation of musical segments using a conditional generative model, with transitions between these segments. The generation of prompts that determine the high-level composition is distinct from the creation of finer, lower-level details. A large language model is then used to suggest the musical form.
翻译:尽管近年来的生成模型能够产生引人入胜的音乐,但其实用性仍存在局限。音乐中的变化往往依赖于随机性,导致作品缺乏结构。持续时间超过一分钟的乐曲可能变得不连贯或重复。本文提出了一种生成结构化、任意长度音乐作品的方法。该方法的核心在于使用条件生成模型创建音乐片段,并在这些片段之间实现过渡。决定高层级作曲的提示生成过程,与创作更精细的低层级细节相互独立。随后,利用大型语言模型来建议音乐形式。