We formulate episodic Markov decision process (MDP) planning as Bayesian inference over policies. The primary contribution is conceptual: the policy itself is treated as the latent variable, and expected return defines an unnormalized posterior density over policies. This preserves the standard expected-return objective, in contrast to trajectory-centric planning-as-inference formulations that introduce auxiliary optimality variables and to entropy-regularized policy optimization methods that solve a different objective. In the exact formulation, the posterior over deterministic policies induces what we define here as an optimal stochastic policy under preference uncertainty, namely the stochastic policy induced by that posterior. For discrete MDPs with stochastic transitions, we study variational sequential Monte Carlo (VSMC) as one approximate inference method for this posterior, introducing policy consistency under state revisitation and coupled transition randomness across particles. Experiments on grid worlds, Blackjack, Triangle Tireworld, and Academic Advising examine the consequences of inference over policies and compare its induced behavior with entropy-regularized policy optimization. The results support the view that MDP planning can be naturally cast as Bayesian inference over policies.


翻译:暂无翻译

0
下载
关闭预览

相关内容

《生成可解释军事行动方案(COA)》
专知会员服务
77+阅读 · 2025年2月23日
《海军规划过程概述》15页slides
专知会员服务
32+阅读 · 2025年2月21日
《警力保护 (FP) 的战术规划考虑因素》45页slides
专知会员服务
11+阅读 · 2025年1月14日
基于KG+LLM的联合作战计划智能生成方法
专知会员服务
45+阅读 · 2025年1月9日
军事决策过程 (MDMP):综合指南
专知会员服务
42+阅读 · 2024年7月26日
《军事决策过程:组织和实施规划》美陆军最新156页
专知会员服务
82+阅读 · 2024年2月5日
《多域作战环境下的军事决策过程》
专知会员服务
204+阅读 · 2023年4月11日
《多域作战环境下的军事决策过程》
专知
114+阅读 · 2023年4月12日
跟踪 | 美国防部DARPA技术研发项目立项流程概述
走向智能论坛
39+阅读 · 2019年6月18日
利用动态深度学习预测金融时间序列基于Python
量化投资与机器学习
18+阅读 · 2018年10月30日
disentangled-representation-papers
CreateAMind
26+阅读 · 2018年9月12日
Focal Loss for Dense Object Detection
统计学习与视觉计算组
12+阅读 · 2018年3月15日
国家自然科学基金
122+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Arxiv
0+阅读 · 7月21日
VIP会员
最新内容
分层反无人机系统发展新趋势
专知会员服务
9+阅读 · 9月3日
何为协作武器?
专知会员服务
10+阅读 · 9月1日
《理解认知战:超越信息》
专知会员服务
14+阅读 · 9月1日
美国战争部在GenAI.mil上推出OpenAI的ChatGPT Mil
专知会员服务
10+阅读 · 8月31日
人工智能赋能军事维护:重新定义国防战备
专知会员服务
5+阅读 · 8月31日
相关VIP内容
《生成可解释军事行动方案(COA)》
专知会员服务
77+阅读 · 2025年2月23日
《海军规划过程概述》15页slides
专知会员服务
32+阅读 · 2025年2月21日
《警力保护 (FP) 的战术规划考虑因素》45页slides
专知会员服务
11+阅读 · 2025年1月14日
基于KG+LLM的联合作战计划智能生成方法
专知会员服务
45+阅读 · 2025年1月9日
军事决策过程 (MDMP):综合指南
专知会员服务
42+阅读 · 2024年7月26日
《军事决策过程:组织和实施规划》美陆军最新156页
专知会员服务
82+阅读 · 2024年2月5日
《多域作战环境下的军事决策过程》
专知会员服务
204+阅读 · 2023年4月11日
相关基金
国家自然科学基金
122+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
11+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员