This paper presents an expert-guided active-inference-inspired framework for adaptive UAV swarm trajectory planning. The proposed method converts multi-UAV trajectory design from a repeated combinatorial optimization problem into a hierarchical probabilistic inference problem. In the offline phase, a genetic-algorithm planner with repulsive-force collision avoidance (GA--RF) generates expert demonstrations, which are abstracted into Mission, Route, and Motion dictionaries. These dictionaries are used to learn a probabilistic world model that captures how expert mission allocations induce route orders and how route orders induce motion-level behaviors. During online operation, the UAV swarm evaluates candidate actions by forming posterior beliefs over symbolic states and minimizing KL-divergence-based abnormality indicators with respect to expert-derived reference distributions. This enables mission allocation, route insertion, motion adaptation, and collision-aware replanning without rerunning the offline optimizer. Bayesian state estimators, including EKF and PF modules, are integrated at the motion level to improve trajectory correction under uncertainty. Simulation results show that the proposed framework preserves expert-like planning structure while producing smoother and more stable behavior than modified Q-learning. Additional validation using real-flight UAV trajectory data demonstrates that the learned world model can correct symbolic predictions under noisy and non-smooth observations, supporting its applicability to adaptive UAV swarm autonomy.
翻译:本文提出了一种基于专家引导的主动推理框架,用于自适应无人机集群的轨迹规划。该方法将多无人机轨迹设计从重复的组合优化问题转化为分层概率推理问题。离线阶段,采用具有排斥力避碰机制的遗传算法规划器生成专家演示,并将其抽象为任务字典、航线字典和运动字典。通过这三个字典学习概率世界模型,捕捉专家任务分配如何诱导航线顺序、航线顺序如何诱导运动层行为。在线运行阶段,无人机集群通过形成对符号状态的后验信念,并最小化与专家参考分布之间的KL散度异常指标来评估候选动作。这使得在不重新运行离线优化器的情况下,即可实现任务分配、航线插入、运动适应与考虑碰撞的重新规划。运动层集成了包含扩展卡尔曼滤波器和粒子滤波模块的贝叶斯状态估计器,以提升不确定性下的轨迹修正能力。仿真结果表明,该框架在保持类似专家规划结构的同时,相比改进的Q学习算法产生了更平滑稳定的行为。利用真实飞行无人机轨迹数据的额外验证表明,所学世界模型能在噪声和非平滑观测条件下修正符号预测,支持其在自适应无人机集群自主性中的适用性。