This paper introduces Local Learner (2L), an algorithm for providing a set of reference strategies to guide the search for programmatic strategies in two-player zero-sum games. Previous learning algorithms, such as Iterated Best Response (IBR), Fictitious Play (FP), and Double-Oracle (DO), can be computationally expensive or miss important information for guiding search algorithms. 2L actively selects a set of reference strategies to improve the search signal. We empirically demonstrate the advantages of our approach while guiding a local search algorithm for synthesizing strategies in three games, including MicroRTS, a challenging real-time strategy game. Results show that 2L learns reference strategies that provide a stronger search signal than IBR, FP, and DO. We also simulate a tournament of MicroRTS, where a synthesizer using 2L outperformed the winners of the two latest MicroRTS competitions, which were programmatic strategies written by human programmers.
翻译:本文介绍了一种名为局部学习者(2L)的算法,旨在为两人零和博弈中的程序化策略搜索提供一组基准策略。以往的学习算法,例如迭代最佳响应(IBR)、虚拟博弈(FP)和双 oracle(DO),要么计算成本高昂,要么会遗漏引导搜索算法所需的重要信息。2L 通过主动选择一组基准策略来提升搜索信号。我们通过引导一种局部搜索算法在三款游戏中合成策略,实证展示了我们方法的优势,这三款游戏包括具有挑战性的即时策略游戏 MicroRTS。结果表明,2L 学习的基准策略相比 IBR、FP 和 DO 提供了更强的搜索信号。我们还模拟了 MicroRTS 锦标赛,其中使用 2L 的合成器性能优于最近两届 MicroRTS 比赛的获胜者——这些获胜者是由人类程序员编写的程序化策略。