We introduce networked communication to the mean-field game framework, in particular to oracle-free settings where $N$ decentralised agents learn along a single, non-episodic evolution path of the empirical system. We prove that our architecture, with only a few reasonable assumptions about network structure, has sample guarantees bounded between those of the centralised- and independent-learning cases. We discuss how the sample guarantees of the three theoretical algorithms do not actually result in practical convergence. Accordingly, we show that in practical settings where the theoretical parameters are not observed (leading to poor estimation of the Q-function), our communication scheme significantly accelerates convergence over the independent case, without relying on the undesirable assumption of a centralised controller. We contribute several further practical enhancements to all three theoretical algorithms, allowing us to showcase their first empirical demonstrations. Our experiments confirm that we can remove several of the key theoretical assumptions of the algorithms, and display the empirical convergence benefits brought by our new networked communication. We additionally show that the networked approach has significant advantages, over both the centralised and independent alternatives, in terms of robustness to unexpected learning failures and to changes in population size.
翻译:我们将网络化通信引入均场博弈框架,特别关注无先知环境——其中N个去中心化智能体沿着经验系统的单一非周期性演化路径进行学习。我们证明,在仅对网络结构做出少数合理假设的前提下,所提出的架构具有介于集中式学习与独立学习之间的样本保证界限。我们讨论了三种理论算法的样本保证在实际中并未导致收敛的原因。为此,我们表明在理论参数不可观测(导致Q函数估计不佳)的实际环境中,我们的通信方案能显著加速收敛,且无需依赖集中控制器的不可取假设。我们对所有三种理论算法进行了多项实际增强,从而首次实现其实验验证。实验证实,我们可去除算法的若干关键理论假设,并展示新型网络化通信带来的经验收敛优势。此外,我们还表明网络化方法在应对意外学习故障和群体规模变化的鲁棒性方面,显著优于集中式与独立式替代方案。