Deep learning has recently empowered and democratized generative modeling of images and text, with additional concurrent works exploring the possibility of generating more complex forms of data, such as audio. However, the high dimensionality, long-range dependencies, and lack of standardized datasets currently makes generative modeling of audio and music very challenging. We propose to model music as a series of discrete notes upon which we can use autoregressive natural language processing techniques for successful generative modeling. While previous works used similar pipelines on data such as sheet music and MIDI, we aim to extend such approaches to the under-studied medium of guitar tablature. Specifically, we develop the first work to our knowledge that models one specific genre as guitar tablature: heavy rock. Unlike other works in guitar tablature generation, we have a freely available public demo at https://huggingface.co/spaces/josuelmet/Metal_Music_Interpolator
翻译:深度学习近年来已推动并普及了图像与文本的生成式建模,同时也有并行研究探索生成更复杂数据形式(如音频)的可能性。然而,音频与音乐的高维度、长程依赖特性以及标准化数据集的缺乏,使得其生成式建模仍面临重大挑战。我们提出将音乐建模为一系列离散音符,从而可运用自回归自然语言处理技术实现有效的生成式建模。尽管此前已有研究采用类似流水线处理乐谱与MIDI等数据,本工作首次将此类方法扩展至研究尚不充分的吉他指法谱领域。具体而言,我们首次针对特定音乐体裁——硬摇滚,开发出吉他指法谱生成模型。与同类吉他指法谱生成研究不同,本工作提供可免费访问的公开演示链接:https://huggingface.co/spaces/josuelmet/Metal_Music_Interpolator