Given the success of model-free methods for control design in many problem settings, it is natural to ask how things will change if realistic communication channels are utilized for the transmission of gradients or policies. While the resulting problem has analogies with the formulations studied under the rubric of networked control systems, the rich literature in that area has typically assumed that the model of the system is known. As a step towards bridging the fields of model-free control design and networked control systems, we ask: \textit{Is it possible to solve basic control problems - such as the linear quadratic regulator (LQR) problem - in a model-free manner over a rate-limited channel?} Toward answering this question, we study a setting where a worker agent transmits quantized policy gradients (of the LQR cost) to a server over a noiseless channel with a finite bit-rate. We propose a new algorithm titled Adaptively Quantized Gradient Descent (\texttt{AQGD}), and prove that above a certain finite threshold bit-rate, \texttt{AQGD} guarantees exponentially fast convergence to the globally optimal policy, with \textit{no deterioration of the exponent relative to the unquantized setting}. More generally, our approach reveals the benefits of adaptive quantization in preserving fast linear convergence rates, and, as such, may be of independent interest to the literature on compressed optimization.
翻译:鉴于无模型方法在多种控制设计问题中的成功应用,自然引发思考:若通过实际通信信道传输梯度或策略,情况将如何改变?尽管由此产生的问题与网络化控制系统中的已有研究框架存在相似性,但该领域现有文献通常假设系统模型已知。为弥合无模型控制设计与网络化控制系统两个领域的差距,我们提出以下问题:能否在速率受限信道上以无模型方式解决线性二次型调节器(LQR)等基本控制问题?针对该问题,我们研究了一种场景:工作智能体通过有限比特率的无噪声信道向服务器传输量化后的策略梯度(关于LQR代价函数)。我们提出了一种名为"自适应量化梯度下降(\texttt{AQGD})"的新算法,并证明当比特率超过某一有限阈值时,\texttt{AQGD}可保证以指数级速度收敛到全局最优策略,且收敛速率相对于未量化情形无退化。更广泛而言,我们的方法揭示了自适应量化在保持快速线性收敛速率方面的优势,因此可能对压缩优化领域的文献具有独立参考价值。