Machine learning from training data with a skewed distribution of examples per class can lead to models that favor performance on common classes at the expense of performance on rare ones. AudioSet has a very wide range of priors over its 527 sound event classes. Classification performance on AudioSet is usually evaluated by a simple average over per-class metrics, meaning that performance on rare classes is equal in importance to the performance on common ones. Several recent papers have used dataset balancing techniques to improve performance on AudioSet. We find, however, that while balancing improves performance on the public AudioSet evaluation data it simultaneously hurts performance on an unpublished evaluation set collected under the same conditions. By varying the degree of balancing, we show that its benefits are fragile and depend on the evaluation set. We also do not find evidence indicating that balancing improves rare class performance relative to common classes. We therefore caution against blind application of balancing, as well as against paying too much attention to small improvements on a public evaluation set.
翻译:从每类示例分布偏斜的训练数据中进行机器学习,可能导致模型偏向于在常见类别上表现优异,同时牺牲稀有类别的性能。AudioSet在其527个声音事件类别中具有极宽的先验分布范围。AudioSet上的分类性能通常通过各类别指标的简单平均值来评估,这意味着稀有类别的性能与常见类别的性能同等重要。近期多篇论文采用数据集平衡技术来提升AudioSet的性能。然而,我们发现平衡虽然提升了公开AudioSet评估数据上的性能,但同时在相同条件下收集的未公开评估集上损害了性能。通过改变平衡程度,我们证明其优势是脆弱的,并且依赖于评估集。我们也没有发现证据表明平衡相比常见类别能提升稀有类别的性能。因此,我们警告不要盲目应用平衡技术,同时也要警惕过度关注公开评估集上的微小改进。