Configurable Markov Decision Processes (Conf-MDPs) have recently been introduced as an extension of the traditional Markov Decision Processes (MDPs) to model the real-world scenarios in which there is the possibility to intervene in the environment in order to configure some of its parameters. In this paper, we focus on a particular subclass of Conf-MDP that satisfies regularity conditions, namely Lipschitz continuity. We start by providing a bound on the Wasserstein distance between $\gamma$-discounted stationary distributions induced by changing policy and configuration. This result generalizes the already existing bounds both for Conf-MDPs and traditional MDPs. Then, we derive a novel performance improvement lower bound.
翻译:可配置马尔可夫决策过程(Conf-MDPs)近期被引入作为传统马尔可夫决策过程(MDPs)的扩展,以模拟那些能够干预环境并配置其部分参数的真实场景。本文聚焦于满足正则性条件(即利普希茨连续性)的Conf-MDP特定子类。首先,我们给出了由策略与配置改变所诱导的$\gamma$-折扣平稳分布之间的Wasserstein距离界。该结果推广了Conf-MDPs和传统MDPs的现有界。随后,我们推导了一个新颖的性能提升下界。