We consider the problem of learning models for risk-sensitive reinforcement learning. We theoretically demonstrate that proper value equivalence, a method of learning models which can be used to plan optimally in the risk-neutral setting, is not sufficient to plan optimally in the risk-sensitive setting. We leverage distributional reinforcement learning to introduce two new notions of model equivalence, one which is general and can be used to plan for any risk measure, but is intractable; and a practical variation which allows one to choose which risk measures they may plan optimally for. We demonstrate how our framework can be used to augment any model-free risk-sensitive algorithm, and provide both tabular and large-scale experiments to demonstrate its ability.
翻译:我们研究了风险敏感强化学习中的模型学习问题。我们从理论上证明,在风险中性设置中可用于最优规划的适当值等价方法,在风险敏感设置中不足以实现最优规划。我们利用分布强化学习引入了两种新的模型等价概念:一种是一般性的、可针对任意风险度量进行规划但难以处理的等价概念;另一种是实用变体,允许用户选择能够实现最优规划的风险度量。我们展示了该框架如何增强任何无模型风险敏感算法,并通过表格实验和大规模实验验证了其有效性。