Ensuring safety and meeting temporal specifications are critical challenges for long-term robotic tasks. Signal temporal logic (STL) has been widely used to systematically and rigorously specify these requirements. However, traditional methods of finding the control policy under those STL requirements are computationally complex and not scalable to high-dimensional or systems with complex nonlinear dynamics. Reinforcement learning (RL) methods can learn the policy to satisfy the STL specifications via hand-crafted or STL-inspired rewards, but might encounter unexpected behaviors due to ambiguity and sparsity in the reward. In this paper, we propose a method to directly learn a neural network controller to satisfy the requirements specified in STL. Our controller learns to roll out trajectories to maximize the STL robustness score in training. In testing, similar to Model Predictive Control (MPC), the learned controller predicts a trajectory within a planning horizon to ensure the satisfaction of the STL requirement in deployment. A backup policy is designed to ensure safety when our controller fails. Our approach can adapt to various initial conditions and environmental parameters. We conduct experiments on six tasks, where our method with the backup policy outperforms the classical methods (MPC, STL-solver), model-free and model-based RL methods in STL satisfaction rate, especially on tasks with complex STL specifications while being 10X-100X faster than the classical methods.
翻译:确保安全性并满足时序规范是长期机器人任务的关键挑战。信号时序逻辑(STL)已被广泛用于系统化且严格地指定这些要求。然而,在STL约束下寻找控制策略的传统方法计算复杂,难以扩展到高维系统或具有复杂非线性动力学的系统。强化学习(RL)方法可以通过手工设计或STL启发的奖励函数来学习满足STL规范的政策,但由于奖励的模糊性和稀疏性,可能会遭遇意外行为。本文提出了一种直接学习神经网络控制器以满足STL指定要求的方法。我们的控制器在训练中学习生成轨迹,以最大化STL鲁棒性得分。在测试中,类似于模型预测控制(MPC),学习到的控制器在规划时域内预测轨迹,以确保部署中满足STL要求。当控制器失效时,设计了一个备用策略来保障安全。我们的方法能够适应多种初始条件和环境参数。我们在六项任务上进行了实验,结果表明,我们的方法结合备用策略在STL满足率上优于经典方法(MPC、STL求解器)、无模型和基于模型的强化学习方法,尤其是在具有复杂STL规范的任务上,同时比经典方法快10倍至100倍。