In many real-world multi-agent cooperative tasks, due to high cost and risk, agents cannot continuously interact with the environment and collect experiences during learning, but have to learn from offline datasets. However, the transition dynamics in the dataset of each agent can be much different from the ones induced by the learned policies of other agents in execution, creating large errors in value estimates. Consequently, agents learn uncoordinated low-performing policies. In this paper, we propose a framework for offline decentralized multi-agent reinforcement learning, which exploits value deviation and transition normalization to deliberately modify the transition probabilities. Value deviation optimistically increases the transition probabilities of high-value next states, and transition normalization normalizes the transition probabilities of next states. They together enable agents to learn high-performing and coordinated policies. Theoretically, we prove the convergence of Q-learning under the altered non-stationary transition dynamics. Empirically, we show that the framework can be easily built on many existing offline reinforcement learning algorithms and achieve substantial improvement in a variety of multi-agent tasks.
翻译:在许多现实世界的多智能体协作任务中,由于高昂的成本和风险,智能体无法在学习过程中持续与环境交互并收集经验,而只能从离线数据集中学习。然而,每个智能体数据集中的转移动态可能与执行过程中其他智能体学习策略所诱导的转移动态存在显著差异,从而导致价值估计出现较大误差,进而使智能体学习到不协调的低性能策略。本文提出了一种离线去中心化多智能体强化学习框架,该框架利用价值偏差和转移归一化来有意修改转移概率。价值偏差乐观地提高高价值下一状态的转移概率,而转移归一化则对下一状态的转移概率进行归一化处理。二者共同使智能体能够学习到高性能且协调的策略。理论上,我们证明了在改变的非平稳转移动态下Q学习的收敛性。实验上,我们表明该框架可轻松构建于许多现有离线强化学习算法之上,并在多种多智能体任务中实现显著改进。