The difficulty of manually specifying reward functions has led to an interest in using linear temporal logic (LTL) to express objectives for reinforcement learning (RL). However, LTL has the downside that it is sensitive to small perturbations in the transition probabilities, which prevents probably approximately correct (PAC) learning without additional assumptions. Time discounting provides a way of removing this sensitivity, while retaining the high expressivity of the logic. We study the use of discounted LTL for policy synthesis in Markov decision processes with unknown transition probabilities, and show how to reduce discounted LTL to discounted-sum reward via a reward machine when all discount factors are identical.
翻译:手动指定奖励函数的难度促使人们利用线性时序逻辑(LTL)来表达强化学习(RL)的目标。然而,LTL的缺点在于它对转移概率的微小扰动敏感,这导致在没有额外假设的情况下无法实现概率近似正确(PAC)学习。时间折扣提供了一种消除这种敏感性的方法,同时保留了该逻辑的高表达性。我们研究了在转移概率未知的马尔可夫决策过程中,使用折扣LTL进行策略合成的问题,并展示了当所有折扣因子相同时,如何通过奖励机将折扣LTL简化为折扣和奖励。