Autonomous aerial vehicles (AAVs) empower sixth-generation (6G) Internet-of-Things (IoT) networks through mobility-driven data collection. However, conventional reward-driven reinforcement learning for AAV trajectory planning suffers from severe credit assignment issues and training instability, because sparse scalar rewards fail to capture the long-term and nonlinear effects of sequential movements. To address these challenges, this paper proposes Learn for Variation (L4V), a gradient-informed trajectory learning framework that replaces high-variance scalar reward signals with dense and analytically grounded policy gradients. Particularly, the coupled evolution of AAV kinematics, distance-dependent channel gains, and per-user data-collection progress is first unrolled into an end-to-end differentiable computational graph. Backpropagation through time then serves as a discrete adjoint solver, which propagates exact sensitivities from the cumulative mission objective to every control action and policy parameter. These structured gradients are used to train a deterministic neural policy with temporal smoothness regularization and gradient clipping. Extensive simulations demonstrate that L4V consistently outperforms representative baselines, including a genetic algorithm, DQN, A2C, and DDPG, in mission completion time, average transmission rate, and training cost
翻译:自主飞行器(AAV)通过运动驱动的数据采集赋能第六代(6G)物联网(IoT)网络。然而,针对AAV轨迹规划的传统奖励驱动强化学习存在严重的信用分配问题和训练不稳定现象,这是因为稀疏的标量奖励无法捕捉序列运动的长期非线性效应。针对这些挑战,本文提出“为变异而学习”(L4V)——一种梯度引导的轨迹学习框架,用密集且具有解析基础的策略梯度替代高方差标量奖励信号。具体而言,首先将AAV运动学、距离相关信道增益和每用户数据采集进度的耦合演化展开为端到端可微计算图。随后,通过时间反向传播作为离散伴随求解器,将精确灵敏度从累积任务目标传播至每个控制动作和策略参数。这些结构化梯度用于训练具有时间平滑正则化和梯度裁剪的确定性神经策略。大量仿真表明,L4V在任务完成时间、平均传输速率和训练成本方面始终优于代表性基线方法(包括遗传算法、DQN、A2C和DDPG)。