Linear Temporal Logic (LTL) is widely used to specify high-level objectives for system policies, and it is highly desirable for autonomous systems to learn the optimal policy with respect to such specifications. However, learning the optimal policy from LTL specifications is not trivial. We present a model-free Reinforcement Learning (RL) approach that efficiently learns an optimal policy for an unknown stochastic system, modelled using Markov Decision Processes (MDPs). We propose a novel and more general product MDP, reward structure and discounting mechanism that, when applied in conjunction with off-the-shelf model-free RL algorithms, efficiently learn the optimal policy that maximizes the probability of satisfying a given LTL specification with optimality guarantees. We also provide improved theoretical results on choosing the key parameters in RL to ensure optimality. To directly evaluate the learned policy, we adopt probabilistic model checker PRISM to compute the probability of the policy satisfying such specifications. Several experiments on various tabular MDP environments across different LTL tasks demonstrate the improved sample efficiency and optimal policy convergence.
翻译:线性时态逻辑(LTL)被广泛用于描述系统策略的高层目标,自主系统非常需要学习满足此类规范的最优策略。然而,从LTL规范中学习最优策略并非易事。我们提出了一种无模型强化学习方法,能够高效学习未知随机系统(以马尔可夫决策过程建模)的最优策略。我们设计了一种新颖且更通用的乘积MDP、奖励结构和折扣机制,当与现成的无模型强化学习算法结合使用时,能以最优性保证高效地学习最大化满足给定LTL规范概率的最优策略。我们还提供了关于选择强化学习关键参数以确保最优性的改进理论结果。为直接评估所学策略,我们采用概率模型检验器PRISM计算策略满足此类规范的概率。在多个表格MDP环境上针对不同LTL任务的实验表明,该方法具有更好的样本效率和最优策略收敛性。