Certification methods for stochastic systems provide sufficient proof rules, based on real-valued supermartingale certificates, to determine the almost-sure satisfaction of $ω$-regular properties (and therefore of linear temporal logic) over general state spaces, encompassing both countably infinite and continuous state spaces. Conversely, reinforcement learning (RL) methods for $ω$-regular tasks have received considerable attention, but they typically lack formal guarantees that the learned policy satisfies the specification, except possibly for finite state and action spaces. We bridge these two lines of research by establishing a novel theoretical connection: under an appropriate reward, the value function associated to a policy that almost surely satisfies an $ω$-regular property encodes a Streett supermartingale certificate for that specification. Our results, validated experimentally on finite Markov decision processes, hold for finite, countably infinite, and continuous state spaces, suggesting a principled route to certificate synthesis via RL.


翻译:随机系统的认证方法提供了基于实值超鞅凭证的充分证明规则,用于确定在包括可数无穷和连续状态空间的一般状态空间上,$ω$-正则属性(因此也包括线性时序逻辑)的几乎必然满足性。反之,针对$ω$-正则任务的强化学习方法虽受到广泛关注,但除有限状态和动作空间外,通常缺乏对所学策略满足规范的形式化保证。我们通过建立一项新颖的理论联系来弥合这两个研究方向:在适当奖励下,与几乎必然满足$ω$-正则属性的策略相关联的值函数编码了该规范的一个Streett超鞅凭证。我们的结果在有限马尔可夫决策过程上经过实验验证,适用于有限、可数无穷和连续状态空间,为通过强化学习进行凭证合成提供了一条原理性路径。

0
下载
关闭预览

相关内容

【博士论文】通过秩的概念理解深度学习,206页pdf
专知会员服务
50+阅读 · 2024年8月7日
因果关联学习,Causal Relational Learning
专知会员服务
185+阅读 · 2020年4月21日
使用 Keras Tuner 调节超参数
TensorFlow
15+阅读 · 2020年2月6日
激活函数还是有一点意思的!
计算机视觉战队
12+阅读 · 2019年6月28日
从信息论的角度来理解损失函数
深度学习每日摘要
17+阅读 · 2019年4月7日
换个角度看GAN:另一种损失函数
机器之心
16+阅读 · 2019年1月1日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
Arxiv
0+阅读 · 5月29日
VIP会员
最新内容
致命七类无人机:无人机时代的演进型合成兵种
《异构无人水面艇集群作战自主制导算法》130页
《人工智能能通过美国陆军战争学院吗?》报告
军事域人工智能驱动系统的治理
专知会员服务
4+阅读 · 9月14日
相关VIP内容
【博士论文】通过秩的概念理解深度学习,206页pdf
专知会员服务
50+阅读 · 2024年8月7日
因果关联学习,Causal Relational Learning
专知会员服务
185+阅读 · 2020年4月21日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员