Self-supervision methods learn representations by solving pretext tasks that do not require human-generated labels, alleviating the need for time-consuming annotations. These methods have been applied in computer vision, natural language processing, environmental sound analysis, and recently in music information retrieval, e.g. for pitch estimation. Particularly in the context of music, there are few insights about the fragility of these models regarding different distributions of data, and how they could be mitigated. In this paper, we explore these questions by dissecting a self-supervised model for pitch estimation adapted for tempo estimation via rigorous experimentation with synthetic data. Specifically, we study the relationship between the input representation and data distribution for self-supervised tempo estimation.
翻译:自监督方法通过解决无需人工标注的预文任务来学习表征,从而缓解了对耗时标注的需求。这些方法已应用于计算机视觉、自然语言处理、环境声音分析,以及近年来在音乐信息检索领域(如音高估计)。尤其在音乐场景下,关于这些模型对不同数据分布的脆弱性及其缓解措施,目前尚缺乏深入见解。本文通过剖析一个为节拍估计而改编的自监督音高估计模型,并基于合成数据进行严谨实验来探讨这些问题。具体而言,我们研究了输入表征与数据分布对自监督节拍估计的影响。