Despite previous efforts in melody-to-lyric generation research, there is still a significant compatibility gap between generated lyrics and melodies, negatively impacting the singability of the outputs. This paper bridges the singability gap with a novel approach to generating singable lyrics by jointly Learning wOrding And Formatting during Melody-to-Lyric training (LOAF-M2L). After general-domain pretraining, our proposed model acquires length awareness first from a large text-only lyric corpus. Then, we introduce a new objective informed by musicological research on the relationship between melody and lyrics during melody-to-lyric training, which enables the model to learn the fine-grained format requirements of the melody. Our model achieves 3.75% and 21.44% absolute accuracy gains in the outputs' number-of-line and syllable-per-line requirements compared to naive fine-tuning, without sacrificing text fluency. Furthermore, our model demonstrates a 63.92% and 74.18% relative improvement of music-lyric compatibility and overall quality in the subjective evaluation, compared to the state-of-the-art melody-to-lyric generation model, highlighting the significance of formatting learning.
翻译:尽管先前在旋律到歌词生成研究中已有诸多尝试,但生成的歌词与旋律之间仍存在显著的兼容性缺口,对输出的可演唱性产生负面影响。本文通过提出在旋律到歌词训练中联合学习措辞与格式(LOAF-M2L)的新方法,桥接了可演唱性缺口。经过通用领域预训练后,我们的模型首先从大规模纯文本歌词语料库中获取长度感知能力。随后,我们根据音乐学对旋律与歌词关系的研究引入新目标,使模型能在旋律到歌词训练中学习旋律的细粒度格式要求。与朴素微调相比,我们的模型在输出的行数要求和每行音节数要求上分别实现了3.75%和21.44%的绝对准确率提升,且未牺牲文本流畅性。此外,在主观评估中,与当前最优的旋律到歌词生成模型相比,我们的模型在歌词与音乐兼容性和整体质量上分别实现了63.92%和74.18%的相对提升,凸显了格式学习的重要性。