Standard Markov decision process (MDP) and reinforcement learning algorithms optimize the policy with respect to the expected gain. We propose an algorithm which enables to optimize an alternative objective: the probability that the gain is greater than a given value. The algorithm can be seen as an extension of the value iteration algorithm. We also show how the proposed algorithm could be generalized to use neural networks, similarly to the deep Q learning extension of Q learning.
翻译:标准马尔可夫决策过程(MDP)与强化学习算法基于期望收益优化策略。我们提出一种算法,能够优化替代性目标:收益大于给定值的概率。该算法可视为价值迭代算法的延伸。我们还展示了如何将所提算法推广至神经网络,类似于深度Q学习对Q学习的扩展方式。