We introduce a novel Proximal Policy Optimization (PPO) algorithm aimed at addressing the challenge of maintaining a uniform proton beam intensity delivery in the Muon to Electron Conversion Experiment (Mu2e) at Fermi National Accelerator Laboratory (Fermilab). Our primary objective is to regulate the spill process to ensure a consistent intensity profile, with the ultimate goal of creating an automated controller capable of providing real-time feedback and calibration of the Spill Regulation System (SRS) parameters on a millisecond timescale. We treat the Mu2e accelerator system as a Markov Decision Process suitable for Reinforcement Learning (RL), utilizing PPO to reduce bias and enhance training stability. A key innovation in our approach is the integration of a neuralized Proportional-Integral-Derivative (PID) controller into the policy function, resulting in a significant improvement in the Spill Duty Factor (SDF) by 13.6%, surpassing the performance of the current PID controller baseline by an additional 1.6%. This paper presents the preliminary offline results based on a differentiable simulator of the Mu2e accelerator. It paves the groundwork for real-time implementations and applications, representing a crucial step towards automated proton beam intensity control for the Mu2e experiment.
翻译:我们提出了一种新颖的近端策略优化(PPO)算法,旨在解决费米国家加速器实验室(Fermilab)Muon到Electron转换实验(Mu2e)中保持均匀质子束强度传输的挑战。主要目标是调控引出过程以确保一致的强度分布,最终目标是开发一种自动化控制器,能够在毫秒时间尺度上为引出调节系统(SRS)参数提供实时反馈和校准。我们将Mu2e加速器系统视为适用于强化学习(RL)的马尔可夫决策过程,利用PPO来减少偏差并提升训练稳定性。本方法的关键创新在于将神经化比例-积分-微分(PID)控制器集成到策略函数中,使得引出占空因子(SDF)显著提升13.6%,相比当前PID控制器基线额外提高了1.6%。本文基于Mu2e加速器的可微分模拟器展示了初步离线结果,为实时实现和应用奠定了基础,标志着向Mu2e实验自动化质子束强度控制迈出的关键一步。