This paper presents a hierarchical decision-making framework for unmanned aerial vehicle (UAV) missions motivated by search-and-rescue (SAR) scenarios under limited simulation training. The framework combines a fixed rule-based high-level advisor with an online goal-conditioned low-level reinforcement learning (RL) controller. To stress-test early adaptation, we also consider a strict no-pretraining deployment regime. The high-level advisor is defined offline from a structured task specification and compiled into deterministic rules. It provides interpretable mission- and safety-aware guidance through recommended actions, avoided actions, and regime-dependent arbitration weights. The low-level controller learns online from task-defined dense rewards and reuses experience through a mode-aware prioritized replay mechanism augmented with rule-derived metadata. We evaluate the framework on two tasks: battery-aware multi-goal delivery and moving-target delivery in obstacle-rich environments. Across both tasks, the proposed method improves early safety and sample efficiency primarily by reducing collision terminations, while preserving the ability to adapt online to scenario-specific dynamics.
翻译:本文提出了一种面向搜索救援场景的无人机任务层次化决策框架,该框架适用于有限模拟训练环境。该框架将固定规则的高层决策器与在线目标条件低层强化学习控制器相结合。为检验早期适应能力,我们同时考虑了严格的零预训练部署模式。高层决策器根据结构化任务规范离线定义,并编译为确定性规则集,通过推荐动作、禁止动作以及依赖状态的任务仲裁权重,提供可解释的任务与安全导向指导。低层控制器从任务定义的密集奖励中在线学习,并通过融入规则衍生元数据的模式感知优先重放机制复用经验。我们在两个任务中评估该框架:障碍物丰富环境下的电池感知多目标运输任务与移动目标运输任务。实验结果表明,所提方法主要通过减少碰撞终止次数提升了早期安全性与样本效率,同时保持了对特定场景动力学的在线适应能力。