Improving Perceptual Quality of Drum Transcription with the Expanded Groove MIDI Dataset

We introduce the Expanded Groove MIDI dataset (E-GMD), an automatic drum transcription (ADT) dataset that contains 444 hours of audio from 43 drum kits, making it an order of magnitude larger than similar datasets, and the first with human-performed velocity annotations. We use E-GMD to optimize classifiers for use in downstream generation by predicting expressive dynamics (velocity) and show with listening tests that they produce outputs with improved perceptual quality, despite similar results on classification metrics. Via the listening tests, we argue that standard classifier metrics, such as accuracy and F-measure score, are insufficient proxies of performance in downstream tasks because they do not fully align with the perceptual quality of generated outputs.

翻译：我们引入了扩大的Groove MIDI数据集(E-GMD ), 这是一个自动的桶转录(ADT)数据集,包含来自43个桶包的444小时音频,使它成为比类似数据集更大的数量级,也是第一个具有人性化速度说明的。我们使用E-GMD 优化分类器,用于下游生成,方法是预测表达动态(速度 ), 并通过监听测试显示,尽管在分类标准上取得了类似结果,它们仍能以更高的感知质量产生产出。通过监听测试,我们认为,标准分类指标,如准确度和F计量分,不足以替代下游任务,因为它们不完全符合生成产出的感知质量。

相关内容

数据集

关注 88

数据集，又称为资料集、数据集合或资料集合，是一种由数据所组成的集合。
Data set（或dataset）是一个数据的集合，通常以表格形式出现。每一列代表一个特定变量。每一行都对应于某一成员的数据集的问题。它列出的价值观为每一个变量，如身高和体重的一个物体或价值的随机数。每个数值被称为数据资料。对应于行数，该数据集的数据可能包括一个或多个成员。

还在修改博士论文？这份《博士论文写作技巧》为你指南

专知会员服务

167+阅读 · 2020年6月9日

【IJCAI2020】从语言图谱到常识图谱，TransOMCS: From Linguistic Graphs to Commonsense Knowledge

专知会员服务

26+阅读 · 2020年5月6日

【Google大脑】进化正则激活层，Evolving Normalization-Activation Layers

专知会员服务

19+阅读 · 2020年4月9日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日