Distributed machine learning is becoming increasingly popular for geo-distributed data analytics, facilitating the collaborative analysis of data scattered across data centers in different regions. This paradigm eliminates the need for centralizing sensitive raw data in one location but faces the significant challenge of high parameter synchronization delays, which stems from the constraints of bandwidth-limited, heterogeneous, and fluctuating wide-area networks. Prior research has focused on optimizing the synchronization topology, evolving from starlike to tree-based structures. However, these solutions typically depend on regular tree structures and lack an adequate topology metric, resulting in limited improvements. This paper proposes NetStorm, an adaptive and highly efficient communication scheduler designed to speed up parameter synchronization across geo-distributed data centers. First, it establishes an effective metric for optimizing a multi-root FAPT synchronization topology. Second, a network awareness module is developed to acquire network knowledge, aiding in topology decisions. Third, a multipath auxiliary transmission mechanism is introduced to enhance network awareness and facilitate multipath transmissions. Lastly, we design policy consistency protocols to guarantee seamless updates of transmission policies. Empirical results demonstrate that NetStorm significantly outperforms distributed training systems like MXNET, MLNET, and TSEngine, with a speedup of 6.5~9.2 times over MXNET.
翻译:分布式机器学习在地理分布式数据分析中日益普及,促进了跨区域数据中心间数据的协同分析。该范式消除了将敏感原始数据集中存储于单一地点的需求,但面临参数同步延迟高的重大挑战,其根源在于带宽受限、异构且波动的广域网环境。先前研究聚焦于同步拓扑优化,已从星型结构演进至树型结构。然而,这些方案通常依赖规则化树结构,且缺乏有效的拓扑度量指标,导致性能提升有限。本文提出NetStorm——一种旨在加速跨地理分布式数据中心参数同步的自适应高效通信调度器。首先,建立了优化多根FAPT同步拓扑的有效度量指标。其次,开发网络感知模块以获取网络知识,辅助拓扑决策。第三,引入多路径辅助传输机制以增强网络感知能力并支持多路径传输。最后,设计策略一致性协议以保障传输策略的无缝更新。实验结果表明,NetStorm显著优于MXNET、MLNET和TSEngine等分布式训练系统,相较MXNET可实现6.5~9.2倍的加速比。