We study the problem of Online Convex Optimization (OCO) with memory, which allows loss functions to depend on past decisions and thus captures temporal effects of learning problems. In this paper, we introduce dynamic policy regret as the performance measure to design algorithms robust to non-stationary environments, which competes algorithms' decisions with a sequence of changing comparators. We propose a novel algorithm for OCO with memory that provably enjoys an optimal dynamic policy regret in terms of time horizon, non-stationarity measure, and memory length. The key technical challenge is how to control the switching cost, the cumulative movements of player's decisions, which is neatly addressed by a novel switching-cost-aware online ensemble approach equipped with a new meta-base decomposition of dynamic policy regret and a careful design of meta-learner and base-learner that explicitly regularizes the switching cost. The results are further applied to tackle non-stationarity in online non-stochastic control (Agarwal et al., 2019), i.e., controlling a linear dynamical system with adversarial disturbance and convex cost functions. We derive a novel gradient-based controller with dynamic policy regret guarantees, which is the first controller provably competitive to a sequence of changing policies for online non-stochastic control.
翻译:我们研究了带记忆的在线凸优化(OCO)问题,该问题允许损失函数依赖于历史决策,从而捕捉学习问题中的时间效应。本文引入动态策略遗憾作为性能度量,以设计对非平稳环境具有鲁棒性的算法,该算法将决策与一系列变化的比较器进行竞争。我们针对带记忆的OCO提出了一种新算法,该算法在时间范围、非平稳性度量和记忆长度方面具有最优的动态策略遗憾保证。关键技术挑战在于如何控制切换成本(即决策者的累积移动量),我们通过一种新颖的感知切换成本的在线集成方法巧妙解决了这一问题:该方法包含动态策略遗憾的新型元基分解,以及通过显式正则化切换成本精心设计的元学习器和基学习器。进一步地,我们将结果应用于解决在线非随机控制(Agarwal等人,2019)中的非平稳性问题(即控制具有对抗性扰动和凸代价函数的线性动力学系统)。我们推导出一种具有动态策略遗憾保证的梯度型控制器,这是首个被证明能够与在线非随机控制中一系列变化策略竞争的控制器。