Data pruning algorithms are commonly used to reduce the memory and computational cost of the optimization process. Recent empirical results reveal that random data pruning remains a strong baseline and outperforms most existing data pruning methods in the high compression regime, i.e., where a fraction of $30\%$ or less of the data is kept. This regime has recently attracted a lot of interest as a result of the role of data pruning in improving the so-called neural scaling laws; in [Sorscher et al.], the authors showed the need for high-quality data pruning algorithms in order to beat the sample power law. In this work, we focus on score-based data pruning algorithms and show theoretically and empirically why such algorithms fail in the high compression regime. We demonstrate ``No Free Lunch" theorems for data pruning and present calibration protocols that enhance the performance of existing pruning algorithms in this high compression regime using randomization.
翻译:数据剪枝算法常用于降低优化过程中的内存与计算成本。近期实证结果表明,随机数据剪枝仍是强基准方法,且在高压缩率场景下(即仅保留不超过30%数据时)优于大多数现有数据剪枝方法。由于数据剪枝在改善所谓神经缩放定律中的作用,该场景近来引起了广泛关注;在[Sorscher等]的研究中,作者论证了为突破样本幂律限制而需要高质量数据剪枝算法的必要性。本研究聚焦于基于评分的数据剪枝算法,从理论与实证层面阐释此类算法在高压缩率场景失效的根本原因。我们提出数据剪枝的“无免费午餐”定理,并建立校准协议——该协议通过引入随机化机制,可增强现有剪枝算法在此高压缩率场景下的性能表现。