Recently, researchers have explored ML-based Traffic Engineering (TE), leveraging neural networks to solve TE problems traditionally addressed by optimization. However, existing ML-based TE schemes remain impractical: they either fail to handle topology changes or suffer from poor scalability due to excessive computational and memory overhead. To overcome these limitations, we propose Geminet, a lightweight and scalable ML-based TE framework that can handle changing topologies. Geminet is built upon two key insights: (i) a methodology that decouples neural networks from topology by learning an iterative gradient-descent-based adjustment process, as the update rule of gradient descent is topology-agnostic, relying only on a few gradient-related quantities; (ii) shifting optimization from path-level routing weights to edge-level dual variables, reducing memory consumption by leveraging the fact that edges are far fewer than paths. Evaluations on WAN and data center datasets show that Geminet significantly improves scalability. Its neural network size is only 0.04% to 7% of existing schemes, while handling topology variations as effectively as HARP, a state-of-the-art ML-based TE approach, without performance degradation. When trained on large-scale topologies, Geminet consumes under 10 GiB of memory, more than eight times less than the 80-plus GiB required by HARP, while achieving 5.45 times faster convergence speed, demonstrating its potential for large-scale deployment.
翻译:近年来,研究者们探索了基于机器学习的流量工程(TE),利用神经网络解决传统上由优化方法处理的TE问题。然而,现有基于ML的TE方案仍不实用:它们要么无法处理拓扑变化,要么因计算和内存开销过大而导致可扩展性差。为克服这些局限,我们提出Geminet——一种轻量级、可扩展且能处理拓扑变化的ML驱动TE框架。Geminet基于两个关键见解构建:(i)一种通过学习基于梯度下降的迭代调整过程将神经网络与拓扑解耦的方法论,因为梯度下降的更新规则与拓扑无关,仅依赖少数梯度相关量;(ii)将优化从路径级路由权重转移到边级对偶变量,利用边数量远少于路径的事实降低内存消耗。在广域网和数据中心数据集上的评估表明,Geminet显著提升了可扩展性。其神经网络规模仅为现有方案的0.04%至7%,同时在处理拓扑变化时性能与当前最先进的基于ML的TE方法HARP相当且无退化。当在大规模拓扑上训练时,Geminet消耗的内存低于10 GiB,比HARP所需的80 GiB以上内存减少八倍多,同时收敛速度提升5.45倍,展现了其大规模部署的潜力。