We study the problems of distributed online and bandit convex optimization against an adaptive adversary. We aim to minimize the average regret on $M$ machines working in parallel over $T$ rounds with $R$ intermittent communications. Assuming the underlying cost functions are convex and can be generated adaptively, our results show that collaboration is not beneficial when the machines have access to the first-order gradient information at the queried points. This is in contrast to the case for stochastic functions, where each machine samples the cost functions from a fixed distribution. Furthermore, we delve into the more challenging setting of federated online optimization with bandit (zeroth-order) feedback, where the machines can only access values of the cost functions at the queried points. The key finding here is identifying the high-dimensional regime where collaboration is beneficial and may even lead to a linear speedup in the number of machines. We further illustrate our findings through federated adversarial linear bandits by developing novel distributed single and two-point feedback algorithms. Our work is the first attempt towards a systematic understanding of federated online optimization with limited feedback, and it attains tight regret bounds in the intermittent communication setting for both first and zeroth-order feedback. Our results thus bridge the gap between stochastic and adaptive settings in federated online optimization.
翻译:我们研究了针对自适应对手的分布式在线与赌博机凸优化问题。目标是最小化在$M$台并行工作的机器上,经过$T$轮次、$R$次间歇通信的平均遗憾。假设底层代价函数是凸的且可自适应生成,我们的结果表明,当机器能获取查询点处的一阶梯度信息时,协作并非有益。这与随机函数情况(每台机器从固定分布中采样代价函数)形成对比。此外,我们深入研究了更具挑战性的具有赌博机(零阶)反馈的联邦在线优化场景,其中机器只能访问查询点处的代价函数值。关键发现是识别出了高维机制,在该机制下协作是有益的,甚至可能带来机器数量的线性加速。我们进一步通过开发新颖的分布式单点与两点反馈算法,利用联邦对抗线性赌博机阐明了我们的发现。我们的工作是首次系统理解具有有限反馈的联邦在线优化的尝试,并且在间歇通信设置下为一阶和零阶反馈都获得了紧致的遗憾界。因此,我们的结果弥合了联邦在线优化中随机设置与自适应设置之间的差距。