This paper introduces a full solution for decentralized routing in Low Earth Orbit satellite constellations based on continual Deep Reinforcement Learning (DRL). This requires addressing multiple challenges, including the partial knowledge at the satellites and their continuous movement, and the time-varying sources of uncertainty in the system, such as traffic, communication links, or communication buffers. We follow a multi-agent approach, where each satellite acts as an independent decision-making agent, while acquiring a limited knowledge of the environment based on the feedback received from the nearby agents. The solution is divided into two phases. First, an offline learning phase relies on decentralized decisions and a global Deep Neural Network (DNN) trained with global experiences. Then, the online phase with local, on-board, and pre-trained DNNs requires continual learning to evolve with the environment, which can be done in two different ways: (1) Model anticipation, where the predictable conditions of the constellation are exploited by each satellite sharing local model with the next satellite; and (2) Federated Learning (FL), where each agent's model is merged first at the cluster level and then aggregated in a global Parameter Server. The results show that, without high congestion, the proposed Multi-Agent DRL framework achieves the same E2E performance as a shortest-path solution, but the latter assumes intensive communication overhead for real-time network-wise knowledge of the system at a centralized node, whereas ours only requires limited feedback exchange among first neighbour satellites. Importantly, our solution adapts well to congestion conditions and exploits less loaded paths. Moreover, the divergence of models over time is easily tackled by the synergy between anticipation, applied in short-term alignment, and FL, utilized for long-term alignment.
翻译:本文提出了一种基于连续深度强化学习(DRL)的低地球轨道卫星星座分散式路由完整解决方案。该方案需应对多重挑战,包括卫星局部知识受限与持续运动、系统中通信流量、通信链路及通信缓冲区等时变不确定性源。我们采用多智能体方法,每颗卫星作为独立决策智能体,通过邻域卫星的反馈获取有限环境知识。解决方案分为两个阶段:第一,离线学习阶段依靠分散式决策与全局深度神经网络(DNN),该网络基于全局经验进行训练;第二,在线阶段利用本地机载预训练DNN,通过持续学习与环境共同演进。该阶段可采用两种方式实现:(1)模型预测,即每颗卫星利用星座可预测条件,将本地模型与下一颗卫星共享;(2)联邦学习(FL),即各智能体模型先在集群层级合并,再在全局参数服务器中聚合。结果表明,在非高拥塞场景下,所提多智能体DRL框架可达到与最短路径方案同等的端到端性能,但后者需集中式节点承担高通信开销以实现实时全网知识同步,而本方案仅需一阶邻域卫星间有限反馈交换。更重要的是,本方案能良好适应拥塞条件并利用低负载路径。此外,通过短期对齐中应用的模型预测与长期对齐中应用的联邦学习协同作用,模型随时间发散的问题可被轻松解决。