We study a variant of online convex optimization where the player is permitted to switch decisions at most $S$ times in expectation throughout $T$ rounds. Similar problems have been addressed in prior work for the discrete decision set setting, and more recently in the continuous setting but only with an adaptive adversary. In this work, we aim to fill the gap and present computationally efficient algorithms in the more prevalent oblivious setting, establishing a regret bound of $O(T/S)$ for general convex losses and $\widetilde O(T/S^2)$ for strongly convex losses. In addition, for stochastic i.i.d.~losses, we present a simple algorithm that performs $\log T$ switches with only a multiplicative $\log T$ factor overhead in its regret in both the general and strongly convex settings. Finally, we complement our algorithms with lower bounds that match our upper bounds in some of the cases we consider.
翻译:我们研究在线凸优化的一种变体,其中玩家在$T$轮决策中期望最多允许$S$次决策切换。类似问题已在离散决策集设置下的先前工作中得到研究,最近在连续设置中也有针对自适应对手的探讨。本文旨在填补这一空白,提出在更常见的 oblivious 设置下计算高效的算法,为一般凸损失建立$O(T/S)$的遗憾界,并为强凸损失建立$\widetilde O(T/S^2)$的遗憾界。此外,对于随机独立同分布损失,我们提出一个简单算法,仅需$\log T$次切换,且在一般凸和强凸设置中遗憾仅增加$\log T$倍乘因子。最后,我们通过下界补充算法分析,在某些考虑情形中这些下界与我们的上界相匹配。