Standard Markov decision process (MDP) and reinforcement learning algorithms optimize the policy with respect to the expected gain. We propose an algorithm which enables to optimize an alternative objective: the probability that the gain is greater than a given value. The algorithm can be seen as an extension of the value iteration algorithm. We also show how the proposed algorithm could be generalized to use neural networks, similarly to the deep Q learning extension of Q learning.
翻译:标准马尔可夫决策过程(MDP)与强化学习算法以期望收益为优化目标来优化策略。本文提出一种能够优化另一目标的算法:收益大于给定值的概率。该算法可视为值迭代算法的扩展。我们进一步展示了如何将该算法泛化至使用神经网络,类似于Q学习向深度Q学习的扩展方式。