We introduce LyricWhiz, a robust, multilingual, and zero-shot automatic lyrics transcription method achieving state-of-the-art performance on various lyrics transcription datasets, even in challenging genres such as rock and metal. Our novel, training-free approach utilizes Whisper, a weakly supervised robust speech recognition model, and GPT-4, today's most performant chat-based large language model. In the proposed method, Whisper functions as the "ear" by transcribing the audio, while GPT-4 serves as the "brain," acting as an annotator with a strong performance for contextualized output selection and correction. Our experiments show that LyricWhiz significantly reduces Word Error Rate compared to existing methods in English and can effectively transcribe lyrics across multiple languages. Furthermore, we use LyricWhiz to create the first publicly available, large-scale, multilingual lyrics transcription dataset with a CC-BY-NC-SA copyright license, based on MTG-Jamendo, and offer a human-annotated subset for noise level estimation and evaluation. We anticipate that our proposed method and dataset will advance the development of multilingual lyrics transcription, a challenging and emerging task.
翻译:我们提出LyricWhiz,一种鲁棒的多语言零样本自动歌词转录方法,在各类歌词转录数据集上实现了最先进的性能,即使在摇滚和金属等具有挑战性的音乐类型中也是如此。这种新颖的无训练方法利用了Whisper(一种弱监督的鲁棒语音识别模型)和GPT-4(当今性能最佳的基于聊天的大型语言模型)。在该方法中,Whisper作为“耳朵”负责转录音频,而GPT-4作为“大脑”充当注释器,在上下文感知的输出选择和校正方面表现优异。实验表明,LyricWhiz能够显著降低英语单词错误率,并有效转录多种语言的歌词。此外,我们利用LyricWhiz基于MTG-Jamendo创建了首个公开可用、大规模、多语言的歌词转录数据集,采用CC-BY-NC-SA版权许可,并提供人工标注的子集用于噪声水平估计和评估。我们预期,所提出的方法和数据集将推动多语言歌词转录这一具有挑战性的新兴任务的发展。