We study the problem of long-term (multiple days) mapping of a river plume using multiple autonomous underwater vehicles (AUVs), focusing on the Douro river representative use-case. We propose an energy - and communication - efficient multi-agent reinforcement learning approach in which a central coordinator intermittently communicates with the AUVs, collecting measurements and issuing commands. Our approach integrates spatiotemporal Gaussian process regression (GPR) with a multi-head Q-network controller that regulates direction and speed for each AUV. Simulations using the Delft3D ocean model demonstrate that our method consistently outperforms both single- and multi-agent benchmarks, with scaling the number of agents both improving mean squared error (MSE) and operational endurance. In some instances, our algorithm demonstrates that doubling the number of AUVs can more than double endurance while maintaining or improving accuracy, underscoring the benefits of multi-agent coordination. Our learned policies generalize across unseen seasonal regimes over different months and years, demonstrating promise for future developments of data-driven long-term monitoring of dynamic plume environments.
翻译:我们研究了使用多台自主水下航行器(AUV)对河流羽流进行长期(多日)映射的问题,重点聚焦于杜罗河的代表性应用场景。我们提出了一种节能且通信高效的多智能体强化学习方法,中央协调器间歇性地与AUV通信,收集测量数据并下达指令。该方法将时空高斯过程回归(GPR)与多头Q网络控制器相结合,用于调控每台AUV的航行方向与速度。基于Delft3D海洋模型的仿真表明,我们的方法在性能上持续优于单智能体和多智能体基准方法,而增加智能体数量不仅能降低均方误差(MSE),还能提升作业续航能力。在某些案例中,我们的算法表明,将AUV数量加倍时,续航能力可提升一倍以上,同时保持或提高测量精度,这凸显了多智能体协调的优势。所学策略可泛化至不同月份和年份中未见过的季节性场景,为未来开发动态羽流环境数据驱动的长期监测技术展现了应用前景。