Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio. In this paper, we (1) provide empirical evidence that assessment of LALMs using the popular MusicQA dataset fails to measure whether a model's responses about music are factually correct, and (2) develop a new protocol for assessing the music comprehension capabilities of LALMs. Specifically, we propose an evaluation protocol that prompts a LALM for factually verifiable information, and parses its open-ended response into a structured format that can be objectively assessed using Precision, Recall, and F1 scores. Using this protocol, we define a benchmark consisting of six factual information retrieval tasks defined on three diverse datasets: MusicNet, the Free Music Archive, and OverClocked ReMix. We benchmark nine recent LALMs, including frontier models like Gemini and the latest open models like Music Flamingo, and release the suite of evaluation scripts at https://github.com/DCL2004/LALM-Eval to facilitate benchmarking of new LALMs.


翻译:大型音频语言模型(LALMs)利用多模态表示对关于音频的自然语言查询生成开放式回答。本文中,我们(1)提供实证证据表明,使用流行的MusicQA数据集评估LALMs无法衡量模型对音乐的回答是否符合事实正确性,(2)开发了一种评估LALMs音乐理解能力的新协议。具体而言,我们提出了一种评估协议,该协议提示LALM提供可事实验证的信息,并将其开放式的回答解析为结构化格式,从而能够使用精确率、召回率和F1分数进行客观评估。利用该协议,我们定义了一个基准测试,包含定义在三个多样化数据集(MusicNet、自由音乐档案馆和OverClocked ReMix)上的六个事实信息检索任务。我们对九个近期LALMs进行了基准测试,包括像Gemini这样的前沿模型和像Music Flamingo这样的最新开放模型,并在https://github.com/DCL2004/LALM-Eval发布了全套评估脚本,以促进新LALMs的基准测试。

0
下载
关闭预览

相关内容

音乐,广义而言,指精心组织声音,并将其排布在时间和空间上的艺术类型。
多模态大型语言模型:综述
专知会员服务
47+阅读 · 2025年6月14日
《语音大语言模型》最新进展综述
专知会员服务
58+阅读 · 2024年10月8日
《多模态大语言模型评估综述》
专知会员服务
41+阅读 · 2024年8月29日
多模态大规模语言模型基准的综述
专知会员服务
41+阅读 · 2024年8月25日
「大型语言模型评测」综述
专知会员服务
70+阅读 · 2024年3月30日
《大型语言模型视频理解》综述
专知会员服务
59+阅读 · 2024年1月2日
天大最新《大型语言模型评估》全面综述,111页pdf
专知会员服务
89+阅读 · 2023年10月31日
近期语音类前沿论文
深度学习每日摘要
14+阅读 · 2019年3月17日
自然语言处理中的语言模型预训练方法
PaperWeekly
14+阅读 · 2018年10月21日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Arxiv
0+阅读 · 6月3日
VIP会员
最新内容
刚刚!Jev中文教程项目发布了
专知会员服务
0+阅读 · 10月4日
《人工智能赋能的适应性多功能电磁战》
专知会员服务
12+阅读 · 9月29日
俄乌战场实验室:全面战争如何重塑现代作战
专知会员服务
8+阅读 · 9月29日
2026年美空军协会会议上的无人机系统趋势
专知会员服务
11+阅读 · 9月28日
反制无人机:乌克兰提供的五点启示
专知会员服务
17+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
11+阅读 · 9月23日
相关VIP内容
多模态大型语言模型:综述
专知会员服务
47+阅读 · 2025年6月14日
《语音大语言模型》最新进展综述
专知会员服务
58+阅读 · 2024年10月8日
《多模态大语言模型评估综述》
专知会员服务
41+阅读 · 2024年8月29日
多模态大规模语言模型基准的综述
专知会员服务
41+阅读 · 2024年8月25日
「大型语言模型评测」综述
专知会员服务
70+阅读 · 2024年3月30日
《大型语言模型视频理解》综述
专知会员服务
59+阅读 · 2024年1月2日
天大最新《大型语言模型评估》全面综述,111页pdf
专知会员服务
89+阅读 · 2023年10月31日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员