We propose an approximate Thompson sampling algorithm that learns linear quadratic regulators (LQR) with an improved Bayesian regret bound of $O(\sqrt{T})$. Our method leverages Langevin dynamics with a meticulously designed preconditioner as well as a simple excitation mechanism. We show that the excitation signal induces the minimum eigenvalue of the preconditioner to grow over time, thereby accelerating the approximate posterior sampling process. Moreover, we identify nontrivial concentration properties of the approximate posteriors generated by our algorithm. These properties enable us to bound the moments of the system state and attain an $O(\sqrt{T})$ regret bound without the unrealistic restrictive assumptions on parameter sets that are often used in the literature.
翻译:我们提出一种近似Thompson采样算法,用于学习线性二次调节器(LQR),其贝叶斯遗憾界改进为$O(\sqrt{T})$。该方法利用带有精心设计预处理器的朗之万动力学以及简单的激励机制。我们证明激励信号会使预处理器的极小特征值随时间增长,从而加速近似后验采样过程。此外,我们识别出该算法生成的近似后验分布具有非平凡集中性质。这些性质使我们能够界定系统状态的矩,并在无需文献中常采用的对参数集不切实际的限制性假设条件下,获得$O(\sqrt{T})$遗憾界。