Coreset selection is powerful in reducing computational costs and accelerating data processing for deep learning algorithms. It strives to identify a small subset from large-scale data, so that training only on the subset practically performs on par with full data. Practitioners regularly desire to identify the smallest possible coreset in realistic scenes while maintaining comparable model performance, to minimize costs and maximize acceleration. Motivated by this desideratum, for the first time, we pose the problem of refined coreset selection, in which the minimal coreset size under model performance constraints is explored. Moreover, to address this problem, we propose an innovative method, which maintains optimization priority order over the model performance and coreset size, and efficiently optimizes them in the coreset selection procedure. Theoretically, we provide the convergence guarantee of the proposed method. Empirically, extensive experiments confirm its superiority compared with previous strategies, often yielding better model performance with smaller coreset sizes.
翻译:核心集选择在降低深度学习算法计算成本、加速数据处理方面具有重要作用。该方法旨在从大规模数据中识别出小型子集,使得仅在该子集上训练能达到与全量数据训练相当的性能。在实际场景中,研究者通常希望在使用最少核心集的前提下维持模型性能,以最小化成本并最大化加速效果。受此需求驱动,我们首次提出精炼核心集选择问题,探索在模型性能约束下的核心集最小规模。针对该问题,我们提出一种创新方法,该方法在模型性能与核心集规模之间建立优化优先级顺序,并在核心集选择过程中对二者进行高效优化。理论上,我们给出了所提方法的收敛性保证。实验表明,该方法相较现有策略具有显著优势,常能以更小的核心集规模实现更优的模型性能。