The impact of local averaging on the performance of federated learning (FL) systems is studied in the presence of communication delay between the clients and the parameter server. To minimize the effect of delay, clients are assigned into different groups, each having its own local parameter server (LPS) that aggregates its clients' models. The groups' models are then aggregated at a global parameter server (GPS) that only communicates with the LPSs. Such setting is known as hierarchical FL (HFL). Different from most works in the literature, the number of local and global communication rounds in our work is randomly determined by the (different) delays experienced by each group of clients. Specifically, the number of local averaging rounds are tied to a wall-clock time period coined the sync time $S$, after which the LPSs synchronize their models by sharing them with the GPS. Such sync time $S$ is then reapplied until a global wall-clock time is exhausted.
翻译:在客户端与参数服务器之间存在通信延迟的情况下,研究了局部平均对联邦学习(FL)系统性能的影响。为最小化延迟的影响,客户端被划分为不同的组,每组拥有自己的局部参数服务器(LPS),负责聚合其组内客户端的模型。随后,各组模型在全局参数服务器(GPS)处聚合,该服务器仅与LPS通信。这种设置被称为分层联邦学习(HFL)。与文献中的大多数工作不同,我们的工作中局部和全局通信轮数由每组客户端经历的不同延迟随机决定。具体而言,局部平均轮数与一个壁钟时间周期(称为同步时间$S$)相关联,在该时间后,LPS通过将其模型共享给GPS来同步。此同步时间$S$随后重复应用,直至全局壁钟时间耗尽。