Singing voice transcription converts recorded singing audio to musical notation. Sound contamination (such as accompaniment) and lack of annotated data make singing voice transcription an extremely difficult task. We take two approaches to tackle the above challenges: 1) introducing multimodal learning for singing voice transcription together with a new multimodal singing dataset, N20EMv2, enhancing noise robustness by utilizing video information (lip movements to predict the onset/offset of notes), and 2) adapting self-supervised learning models from the speech domain to the singing voice transcription task, significantly reducing annotated data requirements while preserving pretrained features. We build a self-supervised learning based audio-only singing voice transcription system, which not only outperforms current state-of-the-art technologies as a strong baseline, but also generalizes well to out-of-domain singing data. We then develop a self-supervised learning based video-only singing voice transcription system that detects note onsets and offsets with an accuracy of about 80\%. Finally, based on the powerful acoustic and visual representations extracted by the above two systems as well as the feature fusion design, we create an audio-visual singing voice transcription system that improves the noise robustness significantly under different acoustic environments compared to the audio-only systems.
翻译:歌唱语音转录将录制的歌唱音频转换为乐谱符号。声音污染(如伴奏)和标注数据匮乏使得歌唱语音转录成为一项极具挑战性的任务。我们采用两种方法应对上述挑战:1)引入多模态学习进行歌唱语音转录,并创建新的多模态歌唱数据集N20EMv2,通过利用视频信息(用唇部运动预测音符的起止)增强噪声鲁棒性;2)将语音领域的自监督学习模型迁移至歌唱语音转录任务,在保留预训练特征的同时显著降低标注数据需求。我们构建了基于自监督学习的纯音频歌唱语音转录系统,该系统不仅作为强基线超越当前最先进技术,还能良好泛化至域外歌唱数据。随后开发基于自监督学习的纯视频歌唱语音转录系统,能以约80%的准确率检测音符起止。最终,基于上述两个系统提取的强大声学与视觉表征以及特征融合设计,我们创建了视听融合的歌唱语音转录系统,在不同声学环境下相比纯音频系统显著提升了噪声鲁棒性。