Extraction of the predominant pitch from polyphonic audio is one of the fundamental tasks in the field of music information retrieval and computational musicology. To accomplish this task using machine learning, a large amount of labeled audio data is required to train the model that predicts the pitch contour. But a classical model pre-trained on data from one domain (source), e.g, songs of a particular singer or genre, may not perform comparatively well in extracting melody from other domains (target). The performance of such models can be boosted by adapting the model using some annotated data in the target domain. In this work, we study various adaptation techniques applied to machine learning models for polyphonic melody extraction. Experimental results show that meta-learning-based adaptation performs better than simple fine-tuning. In addition to this, we find that this method outperforms the existing state-of-the-art non-adaptive polyphonic melody extraction algorithms.
翻译:在多音音频中提取主旋律音高是音乐信息检索与计算音乐学领域的基础任务之一。为通过机器学习完成该任务,需要大量带标注的音频数据来训练预测音高轮廓的模型。然而,在某一领域(源域)数据上预训练的经典模型(例如针对特定歌手或音乐类型的歌曲),在从其他领域(目标域)提取旋律时可能表现不佳。通过利用目标域中的部分标注数据进行模型自适应,可提升此类模型的性能。本研究系统分析了应用于多音旋律提取机器学习模型的各种自适应技术。实验结果表明,基于元学习的自适应方法优于简单的微调方法。此外,我们发现该方法在性能上超越了现有最先进的非自适应多音旋律提取算法。