Multi-agent reinforcement learning (MARL) models multiple agents that interact and learn within a shared environment. This paradigm is applicable to various industrial scenarios such as autonomous driving, quantitative trading, and inventory management. However, applying MARL to these real-world scenarios is impeded by many challenges such as scaling up, complex agent interactions, and non-stationary dynamics. To incentivize the research of MARL on these challenges, we develop MABIM (Multi-Agent Benchmark for Inventory Management) which is a multi-echelon, multi-commodity inventory management simulator that can generate versatile tasks with these different challenging properties. Based on MABIM, we evaluate the performance of classic operations research (OR) methods and popular MARL algorithms on these challenging tasks to highlight their weaknesses and potential.
翻译:多智能体强化学习(MARL)通过建模在共享环境中交互与学习的多个智能体,可应用于自动驾驶、量化交易及库存管理等工业场景。然而,将MARL应用于此类真实场景面临扩展性、复杂智能体交互及非平稳动态等挑战。为激励MARL应对上述难题的研究,我们开发了MABIM(库存管理多智能体基准平台)——一种多层级、多品类的库存管理模拟器,可生成具有不同挑战特性的多样化任务。基于MABIM,我们评估了经典运筹学(OR)方法与主流MARL算法在这些挑战性任务上的表现,以揭示其不足与潜在改进方向。