This paper explores the modeling method of polyphonic music sequence. Due to the great potential of Transformer models in music generation, controllable music generation is receiving more attention. In the task of polyphonic music, current controllable generation research focuses on controlling the generation of chords, but lacks precise adjustment for the controllable generation of choral music textures. This paper proposed Condition Choir Transformer (CoCoFormer) which controls the output of the model by controlling the chord and rhythm inputs at a fine-grained level. In this paper, the self-supervised method improves the loss function and performs joint training through conditional control input and unconditional input training. In order to alleviate the lack of diversity on generated samples caused by the teacher forcing training, this paper added an adversarial training method. CoCoFormer enhances model performance with explicit and implicit inputs to chords and rhythms. In this paper, the experiments proves that CoCoFormer has reached the current better level than current models. On the premise of specifying the polyphonic music texture, the same melody can also be generated in a variety of ways.
翻译:本文探索了复调音乐序列的建模方法。鉴于Transformer模型在音乐生成中的巨大潜力,可控音乐生成日益受到关注。在复调音乐任务中,当前可控生成研究主要集中于和弦生成的控制,但对于合唱织体的可控生成缺乏精细调节。本文提出了Condition Choir Transformer(CoCoFormer),通过细粒度控制和弦与节奏输入来调控模型输出。本文采用自监督方法改进损失函数,并通过条件控制输入与无条件输入训练进行联合训练。为缓解教师强制训练导致的生成样本多样性不足问题,本文引入了对抗训练方法。CoCoFormer通过显式与隐式输入和弦及节奏信息增强了模型性能。实验证明,CoCoFormer已达到当前模型中的先进水平。在指定复调音乐织体的前提下,相同旋律亦可生成多种变体。