The bootstrap is a popular data-driven method to quantify statistical uncertainty, but for modern high-dimensional problems, it could suffer from huge computational costs due to the need to repeatedly generate resamples and refit models. We study the use of bootstraps in high-dimensional environments with a small number of resamples. In particular, we show that by using sample-resample independence from a recent "cheap" bootstrap perspective, running a number of resamples as small as one could attain valid coverage even when the dimension grows closely with the sample size, thus supporting the implementability of the bootstrap for large-scale problems. We validate our theoretical results and compare the performance of our approach with other benchmarks via a range of experiments.
翻译:Bootstrap是一种流行的数据驱动方法,用于量化统计不确定性,但在现代高维问题中,由于需要反复生成重采样并重新拟合模型,该方法可能面临巨大的计算成本。我们研究了在重采样次数较少的情况下,Bootstrap在高维环境中的应用。特别地,通过利用近期提出的"廉价"Bootstrap视角中的样本-重样本独立性,我们证明了即使重采样次数少至一次,且当维度随样本量快速增长时,仍能获得有效的覆盖概率,从而支持Bootstrap在大规模问题中的可实现性。我们通过一系列实验验证了理论结果,并将所提方法与其它基准方法的性能进行了比较。