Differential Dynamic Programming (DDP) is an efficient computational tool for solving nonlinear optimal control problems. It was originally designed as a single shooting method and thus is sensitive to the initial guess supplied. This work considers the extension of DDP to multiple shooting (MS), improving its robustness to initial guesses. A novel derivation is proposed that accounts for the defect between shooting segments during the DDP backward pass, while still maintaining quadratic convergence locally. The derivation enables unifying multiple previous MS algorithms, and opens the door to many smaller algorithmic improvements. A penalty method is introduced to strategically control the step size, further improving the convergence performance. An adaptive merit function and a more reliable acceptance condition are employed for globalization. The effects of these improvements are benchmarked for trajectory optimization with a quadrotor, an acrobot, and a manipulator. MS-DDP is also demonstrated for use in Model Predictive Control (MPC) for dynamic jumping with a quadruped robot, showing its benefits over a single shooting approach.
翻译:微分动态规划(DDP)是求解非线性最优控制问题的高效计算工具。该方法最初被设计为单点打靶法,因此对初始解的选取较为敏感。本研究探讨将DDP扩展至多点打靶(MS),以提升其对初始解的鲁棒性。本文提出一种新颖的推导方法,在DDP反向传递过程中处理打靶段间的缺陷变量,同时保持局部二次收敛性。该推导统一了多种现有的多点打靶算法,并为众多微小的算法改进开辟了道路。引入一种惩罚方法以策略性地控制步长,进一步改善收敛性能。采用自适应价值函数与更可靠的接受条件实现全局化。通过四旋翼飞行器、Acrobot机械臂与操作臂的轨迹优化问题,对这些改进效果进行基准测试。MS-DDP还被验证可用于四足机器人动态跳跃的模型预测控制(MPC),展现出相较于单点打靶法的优势。