Unmanned Aerial Vehicle (UAV) swarms play an effective role in timely data collection from ground sensors in remote and hostile areas. Optimizing the collective behavior of swarms can improve data collection performance. This paper puts forth a new mean field flight resource allocation optimization to minimize age of information (AoI) of sensory data, where balancing the trade-off between the UAVs movements and AoI is formulated as a mean field game (MFG). The MFG optimization yields an expansive solution space encompassing continuous state and action, resulting in significant computational complexity. To address practical situations, we propose, a new mean field hybrid proximal policy optimization (MF-HPPO) scheme to minimize the average AoI by optimizing the UAV's trajectories and data collection scheduling of the ground sensors given mixed continuous and discrete actions. Furthermore, a long short term memory (LSTM) is leveraged in MF-HPPO to predict the time-varying network state and stabilize the training. Numerical results demonstrate that the proposed MF-HPPO reduces the average AoI by up to 45 percent and 57 percent in the considered simulation setting, as compared to multi-agent deep Q-learning (MADQN) method and non-learning random algorithm, respectively.
翻译:无人机集群在远程和恶劣区域的地面传感器数据及时收集中发挥着重要作用。优化集群的集体行为能够提升数据收集性能。本文提出了一种新的平均场飞行资源分配优化方法,以最小化感知数据的信息年龄(AoI),其中无人机移动与AoI之间的权衡被建模为平均场博弈(MFG)。该MFG优化产生包含连续状态与动作的广阔解空间,导致显著的计算复杂度。为应对实际场景,我们提出了一种新的平均场混合近端策略优化(MF-HPPO)方案,通过在混合连续与离散动作下优化无人机轨迹和地面传感器数据采集调度,以最小化平均AoI。此外,MF-HPPO利用长短期记忆(LSTM)网络预测时变网络状态并稳定训练过程。数值结果表明,与多智能体深度Q学习(MADQN)方法及非学习随机算法相比,所提出的MF-HPPO在设定的仿真场景中分别将平均AoI降低了45%和57%。