Reinforcement learning algorithms assume that observations satisfy the Markov property, yet real-world sensors frequently violate this assumption through correlated noise, latency, or partial observability. Standard performance metrics conflate Markov breakdowns with other sources of suboptimality, leaving practitioners without diagnostic tools for such violations. This paper introduces a prediction-based scoring method that quantifies non-Markovian structure in observation trajectories. A random forest first removes nonlinear Markov-compliant dynamics; ridge regression then tests whether historical observations reduce prediction error on the residuals beyond what the current observation provides. The resulting score is bounded in [0, 1] and requires no causal graph construction. Evaluation spans six environments (CartPole, Pendulum, Acrobot, HalfCheetah, Hopper, Walker2d), three algorithms (PPO, A2C, SAC), controlled AR(1) noise at six intensity levels, and 10 seeds per condition. In post-hoc detection, 7 of 16 environment-algorithm pairs, primarily high-dimensional locomotion tasks, show significant positive monotonicity between noise intensity and the violation score (Spearman rho up to 0.78, confirmed under repeated-measures analysis); under training-time noise, 13 of 16 pairs exhibit statistically significant reward degradation. An inversion phenomenon is documented in low-dimensional environments where the random forest absorbs the noise signal, causing the score to decrease as true violations grow, a failure mode analyzed in detail. A practical utility experiment demonstrates that the proposed score correctly identifies partial observability and guides architecture selection, fully recovering performance lost to non-Markovian observations. Source code to reproduce all results is provided at https://github.com/NAVEENMN/Markovianes.


翻译:强化学习算法假设观测满足马尔可夫性质,但现实世界中的传感器常通过相关噪声、延迟或部分可观测性违反这一假设。标准性能指标将马尔可夫性失效与其他次优性来源混为一谈,使从业者缺乏诊断此类违背的工具。本文提出一种基于预测的评分方法,用于量化观测轨迹中的非马尔可夫结构。该方法首先利用随机森林去除非线性马尔可夫合规动力学;随后通过岭回归检验历史观测能否在残差上进一步降低当前观测未能消除的预测误差。所得评分区间为[0, 1],且无需构建因果图。评估覆盖六类环境(CartPole、Pendulum、Acrobot、HalfCheetah、Hopper、Walker2d)、三种算法(PPO、A2C、SAC)、六种强度的受控AR(1)噪声以及每种条件下10个随机种子。在事后检测中,16个环境-算法配对里有7个(主要为高维运动控制任务)显示噪声强度与违背评分之间存在显著正单调性(Spearman ρ最高达0.78,经重复测量分析验证);在训练阶段注入噪声时,16个配对中有13个出现统计显著的奖励衰减现象。研究记录了低维环境中的反转现象:随机森林吸收了噪声信号,导致评分在真实违背加剧时反而下降,本文深入分析了该失效模式。一项实用效用实验表明,所提评分能正确识别部分可观测性并指导架构选择,完全恢复因非马尔可夫观测损失的性能。重现所有结果的源代码已发布于https://github.com/NAVEENMN/Markovianes。

0
下载
关闭预览

相关内容

【CMU博士论文】强化学习中策略评估的统计推断
专知会员服务
27+阅读 · 2024年9月15日
【NeurIPS2023】强化学习中的概率推理:正确的方法
专知会员服务
28+阅读 · 2023年11月25日
【ICML2023】在受限逆强化学习中的可识别性和泛化能力
专知会员服务
26+阅读 · 2023年6月5日
基于模型的强化学习综述
专知会员服务
48+阅读 · 2023年1月9日
强化学习可解释性基础问题探索和方法综述
专知会员服务
93+阅读 · 2022年1月16日
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
「强化学习可解释性」最新2022综述
专知
12+阅读 · 2022年1月16日
Distributional Soft Actor-Critic (DSAC)强化学习算法的设计与验证
深度强化学习实验室
20+阅读 · 2020年8月11日
多智能体强化学习(MARL)近年研究概览
PaperWeekly
38+阅读 · 2020年3月15日
【学界】虚拟对抗训练:一种新颖的半监督学习正则化方法
GAN生成式对抗网络
10+阅读 · 2019年6月9日
你的算法可靠吗? 神经网络不确定性度量
专知
40+阅读 · 2019年4月27日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
VIP会员
相关主题
最新内容
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
1+阅读 · 今天13:43
博士论文 | 大动作空间中的在线与离线策略学习
专知会员服务
0+阅读 · 今天13:36
综述 | Autonomous Research Agents:AI 科学家与验证缺口
《多域冲突比较支持模型》60页
专知会员服务
9+阅读 · 8月7日
面向2027年及未来的海军情报改革
专知会员服务
6+阅读 · 8月5日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
Top
微信扫码咨询专知VIP会员