Existing approaches for generating multitrack music with transformer models have been limited in terms of the number of instruments, the length of the music segments and slow inference. This is partly due to the memory requirements of the lengthy input sequences necessitated by existing representations. In this work, we propose a new multitrack music representation that allows a diverse set of instruments while keeping a short sequence length. Our proposed Multitrack Music Transformer (MMT) achieves comparable performance with state-of-the-art systems, landing in between two recently proposed models in a subjective listening test, while achieving substantial speedups and memory reductions over both, making the method attractive for real time improvisation or near real time creative applications. Further, we propose a new measure for analyzing musical self-attention and show that the trained model attends more to notes that form a consonant interval with the current note and to notes that are 4N beats away from the current step.
翻译:现有基于Transformer模型生成多轨音乐的方法在乐器数量、音乐片段长度及推理速度方面均存在局限。这在一定程度上源于现有表示方法所需的长输入序列带来的内存消耗。本文提出一种新型多轨音乐表示方法,能够在保持较短序列长度的同时涵盖多种乐器。我们提出的多轨音乐Transformer(MMT)在主观听测实验中取得了与现有最优系统相当的性能——介于近期提出的两个模型之间——同时相比两者在速度提升与内存缩减方面均实现显著突破,使其成为适合实时即兴创作或近实时创意应用的理想方案。此外,我们提出一种新的音乐自注意力分析度量方法,并证明训练后的模型更关注与当前音符形成协和音程的节拍,以及距离当前步长4N拍远的音符。