We propose a model-based reinforcement learning (RL) approach for noisy time-dependent gate optimization with improved sample complexity over model-free RL. Sample complexity is the number of controller interactions with the physical system. Leveraging an inductive bias, inspired by recent advances in neural ordinary differential equations (ODEs), we use an auto-differentiable ODE parametrised by a learnable Hamiltonian ansatz to represent the model approximating the environment whose time-dependent part, including the control, is fully known. Control alongside Hamiltonian learning of continuous time-independent parameters is addressed through interactions with the system. We demonstrate an order of magnitude advantage in the sample complexity of our method over standard model-free RL in preparing some standard unitary gates with closed and open system dynamics, in realistic numerical experiments incorporating single shot measurements, arbitrary Hilbert space truncations and uncertainty in Hamiltonian parameters. Also, the learned Hamiltonian can be leveraged by existing control methods like GRAPE for further gradient-based optimization with the controllers found by RL as initializations. Our algorithm that we apply on nitrogen vacancy (NV) centers and transmons in this paper is well suited for controlling partially characterised one and two qubit systems.
翻译:我们提出了一种基于模型的强化学习方法,用于含噪时变量子门优化,相较于无模型强化学习,该方法具有更优的样本复杂度。样本复杂度指控制器与物理系统相互作用的次数。受神经常微分方程最新进展的启发,我们通过引入归纳偏置,利用可微分的常微分方程(参数化于可学习的哈密顿量假设形式)来表征近似环境的模型。该模型中除控制项外的时变部分完全已知,而控制项与连续时间无关的哈密顿参数学习均通过系统交互实现。通过包含单次测量、任意希尔伯特空间截断及哈密顿参数不确定性的实际数值实验,我们证明了该方法在闭系统和开放系统动力学中制备标准酉门时,其样本复杂度比标准无模型强化学习具有一个数量级的优势。此外,学习到的哈密顿量可被现有控制方法(如GRAPE)利用,以强化学习发现的控制器为初始值进行进一步梯度优化。本文中应用于氮空位中心与超导量子比特的算法,特别适用于控制部分表征的单量子比特与双量子比特系统。