Recent benchmarks for memory-augmented reinforcement learning (RL) have introduced partially observable Markov decision process (POMDP) environments in which agents must use historical observations to make decisions. However, these benchmarks often lack fine-grained control over the challenges posed to memory models. Synthetic environments offer a solution, enabling precise manipulation of environment dynamics for rigorous and interpretable evaluation of memory-augmented RL. This paper advances the design of such customizable POMDPs with three key contributions: (1) a theoretical framework for analyzing POMDPs based on Memory Demand Structure (MDS) and related concepts; (2) a methodology using linear dynamics, state aggregation, and reward redistribution to construct POMDPs with predefined MDS; and (3) a suite of lightweight, scalable POMDP environments with tunable difficulty, grounded in our theoretical insights. Overall, our work clarifies core challenges in partially observable RL, offers principled guidelines for POMDP design, and aids in selecting and developing suitable memory architectures for RL tasks.
翻译:近期针对记忆增强强化学习(RL)的基准测试引入了部分可观测马尔可夫决策过程(POMDP)环境,其中智能体必须利用历史观测信息进行决策。然而,这些基准测试通常缺乏对记忆模型所面临挑战的精细控制。合成环境提供了一种解决方案,能够精确调控环境动态,从而对记忆增强强化学习进行严格且可解释的评估。本文通过三项关键贡献推进了此类可定制POMDP的设计:(1)提出了基于记忆需求结构(MDS)及相关概念的POMDP理论分析框架;(2)建立了一种利用线性动力学、状态聚合与奖励再分配来构建具有预定义MDS的POMDP的方法论;(3)基于理论洞见开发了一套轻量级、可扩展且难度可调的POMDP环境套件。总体而言,我们的工作厘清了部分可观测强化学习的核心挑战,为POMDP设计提供了规范化指导原则,并有助于为强化学习任务选择与开发合适的记忆架构。