Text-to-music generation (T2M-Gen) faces a major obstacle due to the scarcity of large-scale publicly available music datasets with natural language captions. To address this, we propose the Music Understanding LLaMA (MU-LLaMA), capable of answering music-related questions and generating captions for music files. Our model utilizes audio representations from a pretrained MERT model to extract music features. However, obtaining a suitable dataset for training the MU-LLaMA model remains challenging, as existing publicly accessible audio question answering datasets lack the necessary depth for open-ended music question answering. To fill this gap, we present a methodology for generating question-answer pairs from existing audio captioning datasets and introduce the MusicQA Dataset designed for answering open-ended music-related questions. The experiments demonstrate that the proposed MU-LLaMA model, trained on our designed MusicQA dataset, achieves outstanding performance in both music question answering and music caption generation across various metrics, outperforming current state-of-the-art (SOTA) models in both fields and offering a promising advancement in the T2M-Gen research field.
翻译:文本到音乐生成(T2M-Gen)面临一个主要障碍,即缺乏大规模公开可用且带有自然语言字幕的音乐数据集。为解决这一问题,我们提出了音乐理解LLaMA(MU-LLaMA)模型,能够回答音乐相关问题并为音乐文件生成字幕。该模型利用预训练的MERT模型提取音频表示,进而获取音乐特征。然而,获取适合训练MU-LLaMA模型的数据集仍然具有挑战性,因为现有公开可用的音频问答数据集缺乏开放式音乐问答所需的深度。为填补这一空白,我们提出了一种从现有音频字幕数据集生成问答对的方法,并引入了专为回答开放式音乐问题而设计的MusicQA数据集。实验表明,基于我们设计的MusicQA数据集训练的MU-LLaMA模型,在音乐问答和音乐字幕生成两项任务上,各项指标均表现出色,超越了当前这两个领域的现有最优(SOTA)模型,为T2M-Gen研究领域提供了有前景的进展。