Stochastic optimal control of dynamical systems is a crucial challenge in sequential decision-making. Recently, control-as-inference approaches have had considerable success, providing a viable risk-sensitive framework to address the exploration-exploitation dilemma. Nonetheless, a majority of these techniques only invoke the inference-control duality to derive a modified risk objective that is then addressed within a reinforcement learning framework. This paper introduces a novel perspective by framing risk-sensitive stochastic control as Markovian score climbing under samples drawn from a conditional particle filter. Our approach, while purely inference-centric, provides asymptotically unbiased estimates for gradient-based policy optimization with optimal importance weighting and no explicit value function learning. To validate our methodology, we apply it to the task of learning neural non-Gaussian feedback policies, showcasing its efficacy on numerical benchmarks of stochastic dynamical systems.
翻译:动态系统的随机最优控制是序贯决策中的关键挑战。近年来,“控制即推理”方法取得了显著成功,为解决探索-利用困境提供了一个可行的风险敏感性框架。然而,这些技术中的大多数仅利用推理-控制对偶性推导出修正的风险目标,随后在强化学习框架内加以处理。本文提出了一种全新视角,将风险敏感性随机控制建模为基于条件粒子滤波采样下的马尔可夫评分攀爬。我们的方法纯以推理为中心,通过最优重要性加权为基于梯度的策略优化提供渐近无偏估计,且无需显式学习价值函数。为验证该方法,我们将其应用于学习非高斯神经反馈策略任务,并通过随机动态系统的数值基准实验展示了其有效性。