Although deep learning has revolutionized music generation, existing methods for structured melody generation follow an end-to-end left-to-right note-by-note generative paradigm and treat each note equally. Here, we present WuYun, a knowledge-enhanced deep learning architecture for improving the structure of generated melodies, which first generates the most structurally important notes to construct a melodic skeleton and subsequently infills it with dynamically decorative notes into a full-fledged melody. Specifically, we use music domain knowledge to extract melodic skeletons and employ sequence learning to reconstruct them, which serve as additional knowledge to provide auxiliary guidance for the melody generation process. We demonstrate that WuYun can generate melodies with better long-term structure and musicality and outperforms other state-of-the-art methods by 0.51 on average on all subjective evaluation metrics. Our study provides a multidisciplinary lens to design melodic hierarchical structures and bridge the gap between data-driven and knowledge-based approaches for numerous music generation tasks.
翻译:尽管深度学习已革新音乐生成领域,现有有结构旋律生成方法仍遵循从左至右逐音符生成的端到端范式,且对每个音符等量齐观。本文提出WuYun——一种知识增强的深度学习架构,通过先生成结构上最重要的音符构建旋律骨架,再以动态装饰音将其填充为完整旋律,从而优化生成旋律的结构性。具体而言,我们利用音乐领域知识提取旋律骨架,并采用序列学习重构旋律骨架,这些骨架作为辅助知识为旋律生成过程提供引导。实验表明,WuYun能生成具有更优长程结构与音乐性的旋律,在所有主观评估指标上平均超越其他最先进方法0.51分。本研究为设计旋律层级结构提供了跨学科视角,并为众多音乐生成任务弥合了数据驱动与知识驱动方法之间的鸿沟。