Singing voice synthesis (SVS), as a specific task for generating the vocal singing voice from a music score, has drawn much attention in recent years. SVS faces the challenge that the singing has various pronunciation flexibility conditioned on the same music score. Most of the previous works of SVS can not well handle the misalignment between the music score and actual singing. In this paper, we propose an acoustic feature processing strategy, named PHONEix, with a phoneme distribution predictor, to alleviate the gap between the music score and the singing voice, which can be easily adopted in different SVS systems. Extensive experiments in various settings demonstrate the effectiveness of our PHONEix in both objective and subjective evaluations.
翻译:歌唱声音合成(SVS)作为根据乐谱生成人声歌唱声音的具体任务,近年来备受关注。SVS面临的挑战在于,在相同乐谱条件下歌唱具有多样的发音灵活性。以往的大多数SVS方法难以妥善处理乐谱与实际演唱之间的错位。本文提出名为PHONEix的声学特征处理策略,该策略结合音素分布预测器,旨在缓解乐谱与歌唱声音之间的差异,且可便捷地应用于不同SVS系统中。多种设置下的广泛实验表明,我们的PHONEix在客观与主观评价中均具有有效性。