In this paper, we introduce Jointist, an instrument-aware multi-instrument framework that is capable of transcribing, recognizing, and separating multiple musical instruments from an audio clip. Jointist consists of an instrument recognition module that conditions the other two modules: a transcription module that outputs instrument-specific piano rolls, and a source separation module that utilizes instrument information and transcription results. The joint training of the transcription and source separation modules serves to improve the performance of both tasks. The instrument module is optional and can be directly controlled by human users. This makes Jointist a flexible user-controllable framework. Our challenging problem formulation makes the model highly useful in the real world given that modern popular music typically consists of multiple instruments. Its novelty, however, necessitates a new perspective on how to evaluate such a model. In our experiments, we assess the proposed model from various aspects, providing a new evaluation perspective for multi-instrument transcription. Our subjective listening study shows that Jointist achieves state-of-the-art performance on popular music, outperforming existing multi-instrument transcription models such as MT3. We conducted experiments on several downstream tasks and found that the proposed method improved transcription by more than 1 percentage points (ppt.), source separation by 5 SDR, downbeat detection by 1.8 ppt., chord recognition by 1.4 ppt., and key estimation by 1.4 ppt., when utilizing transcription results obtained from Jointist. Demo available at \url{https://jointist.github.io/Demo}.
翻译:摘要:本文提出Jointist——一个乐器感知的多乐器框架,能够从音频片段中转录、识别并分离多种乐器。Jointist包含一个乐器识别模块,该模块为其他两个模块提供条件:一个输出乐器特定钢琴卷帘的转录模块,以及一个利用乐器信息和转录结果的源分离模块。转录与源分离模块的联合训练有助于提升两项任务的性能。乐器模块为可选项,可直接由人类用户控制,这使得Jointist成为一个灵活的用户可控框架。我们的挑战性问题定义使该模型在现实世界中具有高度实用性——鉴于现代流行音乐通常由多种乐器组成。然而其新颖性要求我们从新角度评估此类模型。在实验中,我们从多个维度评估所提模型,为多乐器转录提供了新的评估视角。主观听力研究表明,Jointist在流行音乐上达到了最先进性能,优于MT3等现有模型。我们在多个下游任务上进行实验发现,当使用Jointist生成的转录结果时,所提方法使转录准确率提升超过1个百分点、源分离SDR提升5、强拍检测提升1.8个百分点、和弦识别提升1.4个百分点、调性估计提升1.4个百分点。演示地址:\url{https://jointist.github.io/Demo}。