The on-orbit intelligent planning of satellites swarm has attracted increasing attention from scholars. Especially in tasks such as the pursuit and attachment of non-cooperative satellites, satellites swarm must achieve coordinated cooperation with limited resources. The study proposes a reinforcement learning framework that integrates the transformer and expert networks. Firstly, under the constraints of incomplete information about non-cooperative satellites, an implicit multi-satellites cooperation strategy was designed using a communication sharing mechanism. Subsequently, for the characteristics of the pursuit-attachment tasks, the multi-agent reinforcement learning framework is improved by introducing transformers and expert networks inspired by transfer learning ideas. To address the issue of satellites swarm scalability, sequence modelling based on transformers is utilized to craft memory-augmented policy networks, meanwhile increasing the scalability of the swarm. By comparing the convergence curves with other algorithms, it is shown that the proposed method is qualified for pursuit-attachment tasks of satellites swarm. Additionally, simulations under different maneuvering strategies of non-cooperative satellites respectively demonstrate the robustness of the algorithm and the task efficiency of the swarm system. The success rate of pursuit-attachment tasks is analyzed through Monte Carlo simulations.
翻译:星上智能规划已成为学者日益关注的研究热点。在非合作卫星追捕附着等任务中,卫星集群需在有限资源下实现协同合作。本研究提出一种融合Transformer与专家网络的强化学习框架。首先,在非合作卫星信息不完备约束下,通过通信共享机制设计了隐式多星协同策略。随后,针对追捕附着任务特性,借鉴迁移学习思想引入Transformer与专家网络改进多智能体强化学习框架。为解决卫星集群可扩展性问题,采用基于Transformer的序列建模构建记忆增强策略网络,同时提升集群规模适应性。通过与其他算法的收敛曲线对比,表明所提方法适用于卫星集群追捕附着任务。此外,针对非合作卫星不同机动策略的仿真分别验证了算法的鲁棒性与集群系统任务效能。通过蒙特卡洛仿真分析了追捕附着任务的成功率。