The additive model is a popular nonparametric regression method due to its ability to retain modeling flexibility while avoiding the curse of dimensionality. The backfitting algorithm is an intuitive and widely used numerical approach for fitting additive models. However, its application to large datasets may incur a high computational cost and is thus infeasible in practice. To address this problem, we propose a novel approach called independence-encouraging subsampling (IES) to select a subsample from big data for training additive models. Inspired by the minimax optimality of an orthogonal array (OA) due to its pairwise independent predictors and uniform coverage for the range of each predictor, the IES approach selects a subsample that approximates an OA to achieve the minimax optimality. Our asymptotic analyses demonstrate that an IES subsample converges to an OA and that the backfitting algorithm over the subsample converges to a unique solution even if the predictors are highly dependent in the original big data. The proposed IES method is also shown to be numerically appealing via simulations and a real data application.
翻译:可加模型是一种流行的非参数回归方法,因其既能保持建模灵活性又能避免维度灾难而受到青睐。回拟合算法是拟合可加模型的一种直观且广泛使用的数值方法。然而,将其应用于大规模数据集可能会产生高昂的计算成本,因此在实际中不可行。为解决这一问题,我们提出了一种名为“独立性鼓励子抽样”(IES)的新方法,用于从大数据中选择子样本以训练可加模型。受正交阵列(OA)因其成对独立预测变量和对每个预测变量范围的均匀覆盖而达到极小化极大最优性的启发,IES方法选择的子样本近似于OA,从而实现了极小化极大最优性。我们的渐近分析表明,IES子样本收敛到OA,并且即使在原始大数据中预测变量高度相关的情况下,子样本上的回拟合算法也能收敛到唯一解。通过模拟实验和实际数据应用,所提出的IES方法在数值上也表现出良好的效果。