This paper investigates efficient distributed training of a Federated Learning~(FL) model over a wireless network of wireless devices. The communication iterations of the distributed training algorithm may be substantially deteriorated or even blocked by the effects of the devices' background traffic, packet losses, congestion, or latency. We abstract the communication-computation impacts as an `iteration cost' and propose a cost-aware causal FL algorithm~(FedCau) to tackle this problem. We propose an iteration-termination method that trade-offs the training performance and networking costs. We apply our approach when clients use the slotted-ALOHA, the carrier-sense multiple access with collision avoidance~(CSMA/CA), and the orthogonal frequency-division multiple access~(OFDMA) protocols. We show that, given a total cost budget, the training performance degrades as either the background communication traffic or the dimension of the training problem increases. Our results demonstrate the importance of proactively designing optimal cost-efficient stopping criteria to avoid unnecessary communication-computation costs to achieve only a marginal FL training improvement. We validate our method by training and testing FL over the MNIST dataset. Finally, we apply our approach to existing communication efficient FL methods from the literature, achieving further efficiency. We conclude that cost-efficient stopping criteria are essential for the success of practical FL over wireless networks.
翻译:本文研究了在由无线设备构成的无线网络中对联邦学习(FL)模型进行高效分布式训练的问题。分布式训练算法的通信迭代可能会因设备的背景流量、数据包丢失、拥塞或延迟等因素而显著恶化甚至阻塞。我们将通信与计算的影响抽象为“迭代成本”,并提出一种成本感知的因果联邦学习算法(FedCau)来解决该问题。我们提出了一种迭代终止方法,以权衡训练性能与网络成本。当客户端采用时隙ALOHA、带碰撞避免的载波侦听多路访问(CSMA/CA)以及正交频分多址(OFDMA)协议时,我们应用了该方法。研究表明,在给定总成本预算的情况下,训练性能会随着背景通信流量或训练问题维度的增加而下降。我们的结果凸显了主动设计最优成本效益停止准则的重要性,以避免因追求仅边际提升的联邦学习训练而付出不必要的通信与计算成本。我们通过在MNIST数据集上对联邦学习进行训练与测试,验证了该方法的有效性。最后,我们将方法应用于现有文献中的通信高效联邦学习方法中,进一步提升了效率。我们得出结论:成本效益停止准则对于在无线网络环境中成功实现实际联邦学习至关重要。