Reinforcement learning is a growing field in AI with a lot of potential. Intelligent behavior is learned automatically through trial and error in interaction with the environment. However, this learning process is often costly. Using variational quantum circuits as function approximators potentially can reduce this cost. In order to implement this, we propose the quantum natural policy gradient (QNPG) algorithm -- a second-order gradient-based routine that takes advantage of an efficient approximation of the quantum Fisher information matrix. We experimentally demonstrate that QNPG outperforms first-order based training on Contextual Bandits environments regarding convergence speed and stability and moreover reduces the sample complexity. Furthermore, we provide evidence for the practical feasibility of our approach by training on a 12-qubit hardware device.
翻译:强化学习是人工智能中一个具有巨大潜力的新兴领域。智能行为通过与环境的交互试错自动学习,然而这一学习过程往往成本高昂。利用变分量子电路作为函数逼近器可能有助于降低这一成本。为实现此目标,我们提出了量子自然策略梯度(QNPG)算法——一种基于二阶梯度的方案,该方案利用量子Fisher信息矩阵的高效近似。实验表明,在上下文赌博机环境中,QNPG在收敛速度与稳定性方面优于基于一阶梯度的训练方法,并进一步降低了样本复杂度。此外,通过在12量子比特硬件设备上的训练,我们为该方法在实际应用中的可行性提供了证据。