In this paper, we present a Deep Reinforcement Learning (RL)-driven Adaptive Stochastic Nonlinear Model Predictive Control (SNMPC) to optimize uncertainty handling, constraints robustification, feasibility, and closed-loop performance. To this end, we conceive an RL agent to proactively anticipate upcoming control tasks and to dynamically determine the most suitable combination of key SNMPC parameters - foremost the robustification factor $\kappa$ and the Uncertainty Propagation Horizon (UPH) $T_u$. We analyze the trained RL agent's decision-making process and highlight its ability to learn context-dependent optimal parameters. One key finding is that adapting the constraints robustification factor with the learned policy reduces conservatism and improves closed-loop performance while adapting UPH renders previously infeasible SNMPC problems feasible when faced with severe disturbances. We showcase the enhanced robustness and feasibility of our Adaptive SNMPC (aSNMPC) through the real-time motion control task of an autonomous passenger vehicle to follow an optimal race line when confronted with significant time-variant disturbances. Experimental findings demonstrate that our look-ahead RL-driven aSNMPC outperforms its Static SNMPC (sSNMPC) counterpart in minimizing the lateral deviation both with accurate and inaccurate disturbance assumptions and even when driving in previously unexplored environments.
翻译:本文提出一种深度强化学习驱动的自适应随机非线性模型预测控制(SNMPC)方法,用于优化不确定性处理、约束鲁棒化、可行性及闭环性能。为此,我们设计了一个强化学习(RL)智能体,以主动预判即将到来的控制任务,并动态确定关键的SNMPC参数——主要是鲁棒化因子$\kappa$和不确定性传播水平(UPH)$T_u$——的最优组合。我们分析了训练后RL智能体的决策过程,并强调了其学习上下文相关最优参数的能力。一个关键发现是:通过所学策略自适应调整约束鲁棒化因子可降低保守性并提升闭环性能,而调整UPH则能在面临严重扰动时使原本不可行的SNMPC问题变得可行。我们通过一个实时运动控制任务验证了所提出的自适应SNMPC(aSNMPC)增强的鲁棒性与可行性:在显著时变扰动下,控制自动驾驶乘用车沿最优赛道线行驶。实验结果表明,无论是在精确或非精确扰动假设下,甚至面对未见过的驾驶环境,所提出的前向RL驱动的aSNMPC在最小化横向偏差方面均优于其静态SNMPC(sSNMPC)对应方法。