Symbolic music is widely used in various deep learning tasks, including generation, transcription, synthesis, and Music Information Retrieval (MIR). It is mostly employed with discrete models like Transformers, which require music to be tokenized, i.e., formatted into sequences of distinct elements called tokens. Tokenization can be performed in different ways. As Transformer can struggle at reasoning, but capture more easily explicit information, it is important to study how the way the information is represented for such model impact their performances. In this work, we analyze the common tokenization methods and experiment with time and note duration representations. We compare the performances of these two impactful criteria on several tasks, including composer and emotion classification, music generation, and sequence representation learning. We demonstrate that explicit information leads to better results depending on the task.
翻译:符号音乐广泛应用于各类深度学习任务,包括生成、转录、合成与音乐信息检索(MIR)。该类任务主要采用Transformer等离散模型,因此需要将音乐进行标记化处理,即格式化为由不同元素(称为标记)构成的序列。标记化可通过多种方式实现。鉴于Transformer在推理方面存在局限性,但能更轻松地捕捉显式信息,因而研究信息表征方式对该模型性能的影响至关重要。本研究分析了常见的标记化方法,并对时间与音符时长表征进行了实验。我们针对作曲者与情感分类、音乐生成及序列表征学习等多项任务,比较了这两个关键因素的表现。结果表明,显式信息在不同任务中均能带来更优的效果。