Human preferences are not always represented via complete linear orders: It is natural to employ partially-ordered preferences for expressing incomparable outcomes. In this work, we consider decision-making and probabilistic planning in stochastic systems modeled as Markov decision processes (MDPs), given a partially ordered preference over a set of temporally extended goals. Specifically, each temporally extended goal is expressed using a formula in Linear Temporal Logic on Finite Traces (LTL$_f$). To plan with the partially ordered preference, we introduce order theory to map a preference over temporal goals to a preference over policies for the MDP. Accordingly, a most preferred policy under a stochastic ordering induces a stochastic nondominated probability distribution over the finite paths in the MDP. To synthesize a most preferred policy, our technical approach includes two key steps. In the first step, we develop a procedure to transform a partially ordered preference over temporal goals into a computational model, called preference automaton, which is a semi-automaton with a partial order over acceptance conditions. In the second step, we prove that finding a most preferred policy is equivalent to computing a Pareto-optimal policy in a multi-objective MDP that is constructed from the original MDP, the preference automaton, and the chosen stochastic ordering relation. Throughout the paper, we employ running examples to illustrate the proposed preference specification and solution approaches. We demonstrate the efficacy of our algorithm using these examples, providing detailed analysis, and then discuss several potential future directions.
翻译:人类偏好并非总是通过完全线性序来表示:采用偏序偏好来表达不可比较的结果是自然的方式。本文针对随机系统(建模为马尔可夫决策过程MDP)中的决策与概率规划问题,考虑在有限迹线性时序逻辑(LTL$_f$)公式表达的一组时序扩展目标上存在偏序偏好的场景。为处理偏序偏好下的规划问题,我们引入序理论将时序目标的偏好映射为MDP策略的偏好。在此框架下,随机序下的最优策略会在MDP有限路径上诱导出随机非支配概率分布。本文技术方案包含两个关键步骤:首先,我们开发了一种将时序目标上的偏序偏好转化为偏好自动机的计算方法,该自动机是一个带有接受条件偏序关系的半自动机;其次,我们证明寻找最优策略等价于计算多目标MDP中的帕累托最优策略,该多目标MDP由原始MDP、偏好自动机及选定的随机序关系共同构建。全文通过运行实例阐明所提出的偏好规范与求解方法,并利用这些实例验证算法的有效性,在提供详细分析后探讨了若干潜在未来研究方向。