Dataset condensation, a concept within data-centric learning, efficiently transfers critical attributes from an original dataset to a synthetic version, maintaining both diversity and realism. This approach significantly improves model training efficiency and is adaptable across multiple application areas. Previous methods in dataset condensation have faced challenges: some incur high computational costs which limit scalability to larger datasets (e.g., MTT, DREAM, and TESLA), while others are restricted to less optimal design spaces, which could hinder potential improvements, especially in smaller datasets (e.g., SRe2L, G-VBSM, and RDED). To address these limitations, we propose a comprehensive design framework that includes specific, effective strategies like implementing soft category-aware matching and adjusting the learning rate schedule. These strategies are grounded in empirical evidence and theoretical backing. Our resulting approach, Elucidate Dataset Condensation (EDC), establishes a benchmark for both small and large-scale dataset condensation. In our testing, EDC achieves state-of-the-art accuracy, reaching 48.6% on ImageNet-1k with a ResNet-18 model at an IPC of 10, which corresponds to a compression ratio of 0.78%. This performance exceeds those of SRe2L, G-VBSM, and RDED by margins of 27.3%, 17.2%, and 6.6%, respectively.
翻译:数据集凝练是数据驱动学习中的关键概念,它能高效地将原始数据集的关键属性迁移至合成版本,同时保持多样性与真实性。该方法显著提升模型训练效率,并适用于多个应用领域。此前数据集凝练方法面临挑战:部分方法计算成本高,限制了大规模数据集的可扩展性(如MTT、DREAM和TESLA),而其他方法则受限于次优设计空间,可能阻碍性能提升(尤其在小型数据集上,如SRe2L、G-VBSM和RDED)。为解决这些问题,我们提出综合设计框架,包含软类别感知匹配和学习率调度调整等具体有效策略,这些策略均有实证与理论支撑。最终方法Elucidate Dataset Condensation(EDC)为小规模与大规模数据集凝练确立了基准。实验结果显示,EDC在ImageNet-1k数据集上以ResNet-18模型、IPC=10(压缩比0.78%)的条件下达到48.6%的准确率,分别超越SRe2L、G-VBSM和RDED达27.3%、17.2%和6.6%。