Many autonomous systems face safety challenges, requiring robust closed-loop control to handle physical limitations and safety constraints. Real-world systems, like autonomous ships, encounter nonlinear dynamics and environmental disturbances. Reinforcement learning is increasingly used to adapt to complex scenarios, but standard frameworks ensuring safety and stability are lacking. Predictive Safety Filters (PSF) offer a promising solution, ensuring constraint satisfaction in learning-based control without explicit constraint handling. This modular approach allows using arbitrary control policies, with the safety filter optimizing proposed actions to meet physical and safety constraints. We apply this approach to marine navigation, combining RL with PSF on a simulated Cybership II model. The RL agent is trained on path following and collision avpodance, while the PSF monitors and modifies control actions for safety. Results demonstrate the PSF's effectiveness in maintaining safety without hindering the RL agent's learning rate and performance, evaluated against a standard RL agent without PSF.
翻译:许多自主系统面临安全挑战,需要鲁棒的闭环控制以应对物理限制和安全约束。诸如自主船舶等真实系统会遭遇非线性动力学和环境扰动。强化学习日益被用于适应复杂场景,但缺乏能确保安全性与稳定性的标准框架。预测性安全滤波器(Predictive Safety Filters, PSF)提供了一种有前景的解决方案,可在基于学习的控制中确保约束满足,而无需显式处理约束。这种模块化方法允许使用任意控制策略,通过安全滤波器优化所提出的动作以满足物理和安全约束。我们将该方法应用于海洋导航,在模拟的Cybership II模型上将强化学习与PSF相结合。RL智能体针对路径跟踪和避碰任务进行训练,而PSF则监测并修正控制动作以保障安全。结果表明,与未使用PSF的标准RL智能体相比,PSF在维持安全性的同时未阻碍RL智能体的学习速率和性能。