Catch Me If You Can: Deep Meta-RL for Search-and-Rescue using LoRa UAV Networks

Long range (LoRa) wireless networks have been widely proposed as a efficient wireless access networks for the battery-constrained Internet of Things (IoT) devices. In many practical search-and-rescue (SAR) operations, one challenging problem is finding the location of devices carried by a lost person. However, using a LoRa-based IoT network for SAR operations will have a limited coverage caused by high signal attenuation due to the terrestrial blockages especially in highly remote areas. To overcome this challenge, the use of unmanned aerial vehicles (UAVs) as a flying LoRa gateway to transfer messages from ground LoRa nodes to the ground rescue station can be a promising solution. In this paper, the problem of the flying LoRa (FL) gateway control in the search-and-rescue system using the UAV-assisted LoRa network is modeled as a partially observable Markov decision process. Then, a deep meta-RL-based policy is proposed to control the FL gateway trajectory during SAR operation. For initialization of proposed deep meta-RL-based policy, first, a deep RL-based policy is designed to determine the adaptive FL gateway trajectory in a fixed search environment including a fixed radio geometry. Then, as a general solution, a deep meta-RL framework is used for SAR in any new and unknown environments to integrate the prior FL gateway experience with information collected from the other search environments and rapidly adapt the SAR policy model for SAR operation in a new environment. The proposed UAV-assisted LoRa network is then experimentally designed and implemented. Practical evaluation results show that if the deep meta-RL based control policy is applied instead of the deep RL-based one, the number of SAR time slots decreases from 141 to 50.

翻译：长距离（LoRa）无线网络被广泛提出作为电池受限的物联网（IoT）设备的高效无线接入网络。在许多实际搜救（SAR）行动中，一个具有挑战性的问题是寻找失联人员携带设备的位置。然而，将基于LoRa的物联网网络用于搜救行动时，由于地面障碍物（尤其是在偏远地区）导致的高信号衰减，其覆盖范围会受到限制。为解决这一挑战，将无人机（UAV）作为空中LoRa网关，用于将地面LoRa节点的消息传输至地面救援站，是一种有前景的方案。本文中，采用无人机辅助LoRa网络的搜救系统中，空中LoRa（FL）网关的控制问题被建模为部分可观测马尔可夫决策过程。随后，提出了一种基于深度元强化学习的策略，以控制搜救行动中空中LoRa网关的轨迹。为初始化所提出的深度元强化学习策略，首先设计了一个基于深度强化学习的策略，用于在包含固定无线电几何结构的固定搜索环境中确定自适应空中LoRa网关轨迹。然后，作为一种通用解决方案，采用深度元强化学习框架实现未知新环境中的搜救，以整合先前空中LoRa网关经验与其他搜索环境收集的信息，并快速适应新环境中的搜救策略模型。随后，对所提出的无人机辅助LoRa网络进行了实验设计与实现。实际评估结果表明，若采用基于深度元强化学习的控制策略替代基于深度强化学习的策略，搜救时隙数将从141次减少至50次。