We are interested in a challenging task, Realistic-Music-Score based Singing Voice Synthesis (RMS-SVS). RMS-SVS aims to generate high-quality singing voices given realistic music scores with different note types (grace, slur, rest, etc.). Though significant progress has been achieved, recent singing voice synthesis (SVS) methods are limited to fine-grained music scores, which require a complicated data collection pipeline with time-consuming manual annotation to align music notes with phonemes. Furthermore, these manual annotation destroys the regularity of note durations in music scores, making fine-grained music scores inconvenient for composing. To tackle these challenges, we propose RMSSinger, the first RMS-SVS method, which takes realistic music scores as input, eliminating most of the tedious manual annotation and avoiding the aforementioned inconvenience. Note that music scores are based on words rather than phonemes, in RMSSinger, we introduce word-level modeling to avoid the time-consuming phoneme duration annotation and the complicated phoneme-level mel-note alignment. Furthermore, we propose the first diffusion-based pitch modeling method, which ameliorates the naturalness of existing pitch-modeling methods. To achieve these, we collect a new dataset containing realistic music scores and singing voices according to these realistic music scores from professional singers. Extensive experiments on the dataset demonstrate the effectiveness of our methods. Audio samples are available at https://rmssinger.github.io/.
翻译:我们对一项具有挑战性的任务——基于真实乐谱的歌唱声音合成(RMS-SVS)——感兴趣。RMS-SVS旨在根据包含不同音符类型(如装饰音、连音、休止符等)的真实乐谱生成高质量的歌唱声音。尽管已取得显著进展,但近年来的歌唱声音合成(SVS)方法仍局限于细粒度乐谱,这些乐谱需要复杂的数据采集流程以及耗时的人工标注来对齐音乐音符与音素。此外,此类人工标注会破坏乐谱中音符时长的规律性,使得细粒度乐谱不便于作曲。为应对这些挑战,我们提出了RMSSinger——首个RMS-SVS方法,它以真实乐谱为输入,消除了大部分繁琐的人工标注,并避免了上述不便。值得注意的是,乐谱基于单词而非音素,因此在RMSSinger中,我们引入了单词级建模,以避免耗时的音素时长标注以及复杂的音素级旋律-音符对齐。此外,我们提出了首个基于扩散模型的音高建模方法,改善了现有音高建模方法的自然度。为实现这些目标,我们收集了一个新数据集,其中包含专业歌手根据真实乐谱演唱的歌唱声音。在该数据集上的大量实验证明了我们方法的有效性。音频样本可访问https://rmssinger.github.io/。