Automatic song writing is a topic of significant practical interest. However, its research is largely hindered by the lack of training data due to copyright concerns and challenged by its creative nature. Most noticeably, prior works often fall short of modeling the cross-modal correlation between melody and lyrics due to limited parallel data, hence generating lyrics that are less singable. Existing works also lack effective mechanisms for content control, a much desired feature for democratizing song creation for people with limited music background. In this work, we propose to generate pleasantly listenable lyrics without training on melody-lyric aligned data. Instead, we design a hierarchical lyric generation framework that disentangles training (based purely on text) from inference (melody-guided text generation). At inference time, we leverage the crucial alignments between melody and lyrics and compile the given melody into constraints to guide the generation process. Evaluation results show that our model can generate high-quality lyrics that are more singable, intelligible, coherent, and in rhyme than strong baselines including those supervised on parallel data.
翻译:自动作词是一个具有显著实际意义的话题。然而,由于版权问题导致训练数据匮乏,加之其创作本质带来的挑战,相关研究在很大程度上受到阻碍。最值得注意的是,由于平行数据有限,以往的研究往往难以建模旋律与歌词之间的跨模态关联,从而生成可唱性较差的歌词。现有工作也缺乏有效的内容控制机制,而这一功能对于让音乐背景有限的人群普及歌曲创作而言极为重要。在本工作中,我们提出在不依赖旋律-歌词对齐数据进行训练的情况下,生成悦耳动听的歌词。为此,我们设计了一种分层歌词生成框架,将训练过程(完全基于文本)与推理过程(旋律引导的文本生成)解耦。在推理阶段,我们利用旋律与歌词之间的关键对齐关系,将给定旋律编译为约束条件以引导生成过程。评估结果表明,我们的模型能够生成高质量歌词,与包括基于平行数据监督训练的强基线模型相比,其可唱性、可理解性、连贯性和押韵性均更优。