This paper presents a novel approach to learning free terminal time closed-loop control for robotic manipulation tasks, enabling dynamic adjustment of task duration and control inputs to enhance performance. We extend the supervised learning approach, namely solving selected optimal open-loop problems and utilizing them as training data for a policy network, to the free terminal time scenario. Three main challenges are addressed in this extension. First, we introduce a marching scheme that enhances the solution quality and increases the success rate of the open-loop solver by gradually refining time discretization. Second, we extend the QRnet in Nakamura-Zimmerer et al. (2021b) to the free terminal time setting to address discontinuity and improve stability at the terminal state. Third, we present a more automated version of the initial value problem (IVP) enhanced sampling method from previous work (Zhang et al., 2022) to adaptively update the training dataset, significantly improving its quality. By integrating these techniques, we develop a closed-loop policy that operates effectively over a broad domain with varying optimal time durations, achieving near globally optimal total costs.
翻译:本文提出了一种新颖的方法,用于学习自由终端时间下的机器人操作任务闭环控制,能够动态调整任务持续时间和控制输入以提升性能。我们将监督学习方法——即求解选定的最优开环问题并将其用作策略网络的训练数据——扩展到自由终端时间场景。该扩展面临三个主要挑战:首先,我们引入了一种渐进式方案,通过逐步细化时间离散化来提高开环求解器的解质量和成功率;其次,将Nakamura-Zimmerer等(2021b)中的QRnet扩展到自由终端时间设置,以解决终端状态处的非连续性问题并提升稳定性;第三,我们提出了一个更自动化的初值问题(IVP)增强采样方法(源自Zhang等(2022)的前期工作),用于自适应更新训练数据集,显著提升其质量。通过整合这些技术,我们开发出一种能在广泛域内有效运行的闭环策略,可适应不同最优时间持续时间,并实现接近全局最优的总成本。