This article describes the Data-Efficient Low-Complexity Acoustic Scene Classification Task in the DCASE 2024 Challenge and the corresponding baseline system. The task setup is a continuation of previous editions (2022 and 2023), which focused on recording device mismatches and low-complexity constraints. This year's edition introduces an additional real-world problem: participants must develop data-efficient systems for five scenarios, which progressively limit the available training data. The provided baseline system is based on an efficient, factorized CNN architecture constructed from inverted residual blocks and uses Freq-MixStyle to tackle the device mismatch problem. The baseline system's accuracy ranges from 42.40% on the smallest to 56.99% on the largest training set.
翻译:本文介绍了DCASE 2024挑战赛中的数据高效低复杂度声学场景分类任务及其对应的基线系统。该任务框架延续了前两届(2022年和2023年)的设置,重点关注录音设备不匹配和低复杂度约束。本届挑战赛引入了额外的实际问题:参赛者需针对五种场景开发数据高效的系统,这些场景逐渐限制可用训练数据。提供的基线系统基于由倒置残差块构建的高效因式分解CNN架构,并采用Freq-MixStyle方法解决设备不匹配问题。该基线系统的准确率在最小训练集上为42.40%,在最大训练集上达到56.99%。