In this paper, we study a one-shot distributed learning algorithm via refitting Bootstrap samples, which we refer to as ReBoot. Given the local models that are fit on multiple independent subsamples, ReBoot refits a new model on the union of the Bootstrap samples drawn from these local models. The whole procedure requires only one round of communication of model parameters. Theoretically, we analyze the statistical rate of ReBoot for generalized linear models (GLM) and noisy phase retrieval, which represent convex and non-convex problems respectively. In both cases, ReBoot provably achieves the full-sample statistical rate whenever the subsample size is not too small. In particular, we show that the systematic bias of ReBoot, the error that is independent of the number of subsamples, is $O(n ^ {-2})$ in GLM, where n is the subsample size. This rate is sharper than that of model parameter averaging and its variants, implying the higher tolerance of ReBoot with respect to data splits to maintain the full-sample rate. Simulation study exhibits the statistical advantage of ReBoot over competing methods including averaging and CSL (Communication-efficient Surrogate Likelihood) with up to two rounds of gradient communication. Finally, we propose FedReBoot, an iterative version of ReBoot, to aggregate convolutional neural networks for image classification, which exhibits substantial superiority over FedAve within early rounds of communication.
翻译:本文研究了一种通过重抽Bootstrap样本的单轮分布式学习算法,我们将其称为ReBoot。给定基于多个独立子样本拟合的局部模型,ReBoot从这些局部模型中抽取Bootstrap样本,并在其并集上重新拟合新模型。整个过程仅需一轮模型参数通信。理论上,我们分析了ReBoot在广义线性模型(GLM)和有噪相位恢复中的统计速率,这两个问题分别代表凸优化与非凸优化场景。在两种情形下,只要子样本量足够大,ReBoot均可达到全样本统计速率。特别地,我们证明ReBoot的系统性偏差(即独立于子样本数量的误差)在GLM中为$O(n^{-2})$,其中n为子样本大小。该速率优于模型参数平均及其变体的结果,表明ReBoot在维持全样本速率时对数据划分具有更高容忍度。仿真实验显示,相比平均法和基于至多两轮梯度通信的CSL(通信高效代理似然)等竞争方法,ReBoot具有统计优势。最后,我们提出ReBoot的迭代版本FedReBoot,用于聚合卷积神经网络进行图像分类,在早期通信轮次中显著优于FedAve。