When the data used for reinforcement learning (RL) are collected by multiple agents in a distributed manner, federated versions of RL algorithms allow collaborative learning without the need for agents to share their local data. In this paper, we consider federated Q-learning, which aims to learn an optimal Q-function by periodically aggregating local Q-estimates trained on local data alone. Focusing on infinite-horizon tabular Markov decision processes, we provide sample complexity guarantees for both the synchronous and asynchronous variants of federated Q-learning. In both cases, our bounds exhibit a linear speedup with respect to the number of agents and near-optimal dependencies on other salient problem parameters. In the asynchronous setting, existing analyses of federated Q-learning, which adopt an equally weighted averaging of local Q-estimates, require that every agent covers the entire state-action space. In contrast, our improved sample complexity scales inverse proportionally to the minimum entry of the average stationary state-action occupancy distribution of all agents, thus only requiring the agents to collectively cover the entire state-action space, unveiling the blessing of heterogeneity in enabling collaborative learning by relaxing the coverage requirement of the single-agent case. However, its sample complexity still suffers when the local trajectories are highly heterogeneous. In response, we propose a novel federated Q-learning algorithm with importance averaging, giving larger weights to more frequently visited state-action pairs, which achieves a robust linear speedup as if all trajectories are centrally processed, regardless of the heterogeneity of local behavior policies.
翻译:当用于强化学习的数据由多个智能体以分布式方式收集时,联邦版本的强化学习算法允许协作学习,而无需智能体共享其本地数据。在本文中,我们考虑联邦Q学习,其目标是通过定期聚合仅基于本地数据训练的局部Q估计来学习最优Q函数。聚焦于无限时域的表格型马尔可夫决策过程,我们为联邦Q学习的同步和异步变体提供了样本复杂度保证。在这两种情形下,我们的界均展现出关于智能体数量的线性加速,以及在其他关键问题参数上的接近最优依赖性。在异步设置中,现有联邦Q学习的分析采用对局部Q估计的等权重平均化,要求每个智能体覆盖整个状态-动作空间。相比之下,我们改进的样本复杂度与所有智能体平均稳态状态-动作占据分布的最小条目成反比,因此仅要求智能体集体覆盖整个状态-动作空间,揭示了异构性在通过放宽单智能体情况下的覆盖需求来实现协作学习中的裨益。然而,当局部轨迹高度异构时,其样本复杂度仍然受到影响。为此,我们提出了一种具有重要性平均化的新型联邦Q学习算法,对更频繁访问的状态-动作对赋予更大权重,从而实现了鲁棒的线性加速,仿佛所有轨迹被集中处理,而与局部行为策略的异构性无关。