Neural ordinary differential equations (ODEs) are widely recognized as the standard for modeling physical mechanisms, which help to perform approximate inference in unknown physical or biological environments. In partially observable (PO) environments, how to infer unseen information from raw observations puzzled the agents. By using a recurrent policy with a compact context, context-based reinforcement learning provides a flexible way to extract unobservable information from historical transitions. To help the agent extract more dynamics-related information, we present a novel ODE-based recurrent model combines with model-free reinforcement learning (RL) framework to solve partially observable Markov decision processes (POMDPs). We experimentally demonstrate the efficacy of our methods across various PO continuous control and meta-RL tasks. Furthermore, our experiments illustrate that our method is robust against irregular observations, owing to the ability of ODEs to model irregularly-sampled time series.
翻译:神经常微分方程(ODEs)被广泛认为是物理机制建模的标准工具,有助于在未知物理或生物环境中进行近似推断。在部分可观测(PO)环境中,如何从原始观测中推断未知信息一直困扰着智能体。通过使用具有紧凑上下文的循环策略,基于上下文的强化学习提供了一种从历史转换中提取不可观测信息的灵活方式。为了帮助智能体提取更多与动力学相关的信息,我们提出了一种新颖的基于ODE的循环模型,并将其与无模型强化学习(RL)框架相结合,以解决部分可观测马尔可夫决策过程(POMDPs)。我们通过实验展示了该方法在多种部分可观测连续控制任务和元强化学习任务中的有效性。此外,我们的实验表明,由于ODE具有建模不规则采样时间序列的能力,该方法对不规则观测具有鲁棒性。