Recent deep music generation studies have put much emphasis on long-term generation with structures. However, we are yet to see high-quality, well-structured whole-song generation. In this paper, we make the first attempt to model a full music piece under the realization of compositional hierarchy. With a focus on symbolic representations of pop songs, we define a hierarchical language, in which each level of hierarchy focuses on the semantics and context dependency at a certain music scope. The high-level languages reveal whole-song form, phrase, and cadence, whereas the low-level languages focus on notes, chords, and their local patterns. A cascaded diffusion model is trained to model the hierarchical language, where each level is conditioned on its upper levels. Experiments and analysis show that our model is capable of generating full-piece music with recognizable global verse-chorus structure and cadences, and the music quality is higher than the baselines. Additionally, we show that the proposed model is controllable in a flexible way. By sampling from the interpretable hierarchical languages or adjusting pre-trained external representations, users can control the music flow via various features such as phrase harmonic structures, rhythmic patterns, and accompaniment texture.
翻译:近期深度音乐生成研究多聚焦于具有结构的长时段生成,但尚未实现高质量、结构完整的整首歌曲生成。本文首次在作曲层级结构框架下探索完整音乐作品的建模。聚焦流行歌曲的符号表征,我们定义了一种层级语言体系,其中每个层级专注于特定音乐范围内的语义与上下文依赖关系:高层语言揭示整首歌曲的曲式结构、乐句与终止式,而低层语言聚焦音符、和弦及其局部模式。通过训练级联扩散模型对层级语言进行建模,使每个层级以上层信息为条件。实验与分析表明,本模型能够生成具有可辨识全局主歌-副歌结构与终止式的完整乐曲,且音乐质量优于基线方法。此外,我们证明了该模型具备灵活的可控性:通过从可解释的层级语言中采样或调整预训练的外部表征,用户可借助乐句和声结构、节奏模式、伴奏织体等多种特征控制音乐流向。