Optimizing static risk-averse objectives in Markov decision processes is challenging because they do not readily admit dynamic programming decompositions. Prior work has proposed to use a dynamic decomposition of risk measures that help to formulate dynamic programs on an augmented state space. This paper shows that several existing decompositions are inherently inexact, contradicting several claims in the literature. In particular, we give examples that show that popular decompositions for CVaR and EVaR risk measures are strict overestimates of the true risk values. However, an exact decomposition is possible for VaR, and we give a simple proof that illustrates the fundamental difference between VaR and CVaR dynamic programming properties.
翻译:在马尔可夫决策过程中优化静态风险规避目标具有挑战性,因其难以直接采用动态规划分解方法。已有研究提出利用风险测度的动态分解,在扩展状态空间中构建动态规划模型。本文证明现有若干分解方法本质存在不精确性,这与文献中的多项结论相悖。具体而言,我们通过实例表明,针对CVaR和EVaR风险测度的常用分解方法会严格高估真实风险值。然而,VaR测度可实现精确分解,我们通过简洁证明阐明了VaR与CVaR动态规划性质的根本差异。