We study the problem of learning decentralized linear quadratic regulator when the system model is unknown a priori. We propose an online learning algorithm that adaptively designs a control policy as new data samples from a single system trajectory become available. Our algorithm design uses a disturbance-feedback representation of state-feedback controllers coupled with online convex optimization with memory and delayed feedback. We show that our controller enjoys an expected regret that scales as $\sqrt{T}$ with the time horizon $T$ for the case of partially nested information pattern. For more general information patterns, the optimal controller is unknown even if the system model is known. In this case, the regret of our controller is shown with respect to a linear sub-optimal controller. We validate our theoretical findings using numerical experiments.
翻译:本文研究了系统模型先验未知时的分散式线性二次型调节器学习问题。我们提出了一种在线学习算法,该算法能够根据单条系统轨迹采集的新数据样本自适应地设计控制策略。算法设计采用状态反馈控制器的扰动反馈表示,并结合具有记忆和延迟反馈的在线凸优化方法。研究表明,在部分嵌套信息模式下,所提控制器随时间范围 $T$ 的期望遗憾量级为 $\sqrt{T}$。对于更一般的信息模式,即使系统模型已知,最优控制器也无法确定。此时,我们证明该控制器相对于线性次优控制器具有遗憾界。通过数值实验验证了理论结果。