We consider off-policy evaluation (OPE) in Partially Observable Markov Decision Processes, where the evaluation policy depends only on observable variables but the behavior policy depends on latent states (Tennenholtz et al. (2020a)). Prior work on this problem uses a causal identification strategy based on one-step observable proxies of the hidden state, which relies on the invertibility of certain one-step moment matrices. In this work, we relax this requirement by using spectral methods and extending one-step proxies both into the past and future. We empirically compare our OPE methods to existing ones and demonstrate their improved prediction accuracy and greater generality. Lastly, we derive a separate Importance Sampling (IS) algorithm which relies on rank, distinctness, and positivity conditions, and not on the strict sufficiency conditions of observable trajectories with respect to the reward and hidden-state structure required by Tennenholtz et al. (2020a).


翻译:在部分可观察的Markov决策程序中,我们考虑政策外评价,因为评价政策仅取决于可观察的变量,而行为政策则取决于潜在状态(Tennnholtz等人(2020年a) ) 。 先前关于该问题的工作采用了基于隐藏状态的一步可观察的近似值的因果关系识别战略,这一战略依赖于某些一分一秒的矩阵的可视性。在这项工作中,我们通过使用光谱方法,将一步的代理人延伸到过去和将来,放松了这一要求。我们实证地比较了我们的OPE方法与现有的方法,并表明其预测的准确性和更加笼统性。 最后,我们得出了一种独立的重要性抽样算法,该算法依赖于等级、区别性和假设性条件,而不是Tenenholtz等人(2020年a)要求的奖赏和隐藏状态结构方面的可观察轨迹的严格充分条件。

0
下载
关闭预览

相关内容

Fariz Darari简明《博弈论Game Theory》介绍,35页ppt
专知会员服务
113+阅读 · 2020年5月15日
Stabilizing Transformers for Reinforcement Learning
专知会员服务
61+阅读 · 2019年10月17日
强化学习最新教程,17页pdf
专知会员服务
182+阅读 · 2019年10月11日
Hierarchically Structured Meta-learning
CreateAMind
27+阅读 · 2019年5月22日
Transferring Knowledge across Learning Processes
CreateAMind
29+阅读 · 2019年5月18日
已删除
将门创投
4+阅读 · 2019年4月1日
逆强化学习-学习人先验的动机
CreateAMind
16+阅读 · 2019年1月18日
【SIGIR2018】五篇对抗训练文章
专知
12+阅读 · 2018年7月9日
Hierarchical Disentangled Representations
CreateAMind
4+阅读 · 2018年4月15日
强化学习族谱
CreateAMind
26+阅读 · 2017年8月2日
强化学习 cartpole_a3c
CreateAMind
9+阅读 · 2017年7月21日
Arxiv
0+阅读 · 2021年11月15日
Arxiv
3+阅读 · 2018年1月31日
VIP会员
最新内容
《美军水下战与海床战概述及本地实施》
专知会员服务
0+阅读 · 6分钟前
面向未来冲突推进陆军情报体制改革
专知会员服务
0+阅读 · 24分钟前
乌克兰纵深打击如何重塑俄罗斯的战略选择
专知会员服务
2+阅读 · 7月24日
俄乌战争中关于中程打击无人机部署的经验启示
《基于强化学习的自动化红队测试》
专知会员服务
4+阅读 · 7月23日
伊朗不对称防空战略的演进
专知会员服务
4+阅读 · 7月23日
对抗环境下超视距目标打击的情报支援
专知会员服务
11+阅读 · 7月22日
相关资讯
Hierarchically Structured Meta-learning
CreateAMind
27+阅读 · 2019年5月22日
Transferring Knowledge across Learning Processes
CreateAMind
29+阅读 · 2019年5月18日
已删除
将门创投
4+阅读 · 2019年4月1日
逆强化学习-学习人先验的动机
CreateAMind
16+阅读 · 2019年1月18日
【SIGIR2018】五篇对抗训练文章
专知
12+阅读 · 2018年7月9日
Hierarchical Disentangled Representations
CreateAMind
4+阅读 · 2018年4月15日
强化学习族谱
CreateAMind
26+阅读 · 2017年8月2日
强化学习 cartpole_a3c
CreateAMind
9+阅读 · 2017年7月21日
Top
微信扫码咨询专知VIP会员