In the first-order query model for zero-sum $K\times K$ matrix games, players observe the expected pay-offs for all their possible actions under the randomized action played by their opponent. This classical model has received renewed interest after the discovery by Rakhlin and Sridharan that $\epsilon$-approximate Nash equilibria can be computed efficiently from $O(\frac{\ln K}{\epsilon})$ instead of $O(\frac{\ln K}{\epsilon^2})$ queries. Surprisingly, the optimal number of such queries, as a function of both $\epsilon$ and $K$, is not known. We make progress on this question on two fronts. First, we fully characterise the query complexity of learning exact equilibria ($\epsilon=0$), by showing that they require a number of queries that is linear in $K$, which means that it is essentially as hard as querying the whole matrix, which can also be done with $K$ queries. Second, for $\epsilon > 0$, the current query complexity upper bound stands at $O(\min(\frac{\ln(K)}{\epsilon} , K))$. We argue that, unfortunately, obtaining a matching lower bound is not possible with existing techniques: we prove that no lower bound can be derived by constructing hard matrices whose entries take values in a known countable set, because such matrices can be fully identified by a single query. This rules out, for instance, reducing to an optimization problem over the hypercube by encoding it as a binary payoff matrix. We then introduce a new technique for lower bounds, which allows us to obtain lower bounds of order $\tilde\Omega(\log(\frac{1}{K\epsilon})$ for any $\epsilon \leq 1 / (cK^4)$, where $c$ is a constant independent of $K$. We further discuss possible future directions to improve on our techniques in order to close the gap with the upper bounds.
翻译:在一阶查询模型下进行零和$K \times K$矩阵博弈时,玩家根据对手的随机化行动观察其所有可能行动的期望收益。自Rakhlin与Sridharan发现可从$O(\frac{\ln K}{\epsilon})$(而非$O(\frac{\ln K}{\epsilon^2})$)次查询高效计算$\epsilon$-近似纳什均衡以来,这一经典模型重新引起关注。令人惊讶的是,作为$\epsilon$和$K$的函数,此类查询的最优数量尚不明确。我们从两个方面推进了该问题的研究。首先,我们完整刻画了学习精确均衡($\epsilon=0$)的查询复杂度,证明其所需查询次数与$K$呈线性关系,意味着其难度本质上等同于查询整个矩阵(同样可通过$K$次查询完成)。其次,对于$\epsilon > 0$,当前的查询复杂度上界为$O(\min(\frac{\ln(K)}{\epsilon} , K))$。我们认为,遗憾的是,现有技术无法获得匹配的下界:我们证明,无法通过构造条目取值于已知可数集的困难矩阵来推导下界,因为此类矩阵可通过单次查询完全识别。这排除了例如将超立方体上的优化问题编码为二元收益矩阵的归约途径。随后,我们引入了一种新的下界技术,使得对于任意$\epsilon \leq 1 / (cK^4)$(其中$c$为与$K$无关的常数),可获得$\tilde\Omega(\log(\frac{1}{K\epsilon})$量级的下界。我们进一步讨论了未来可能改进该技术以缩小与上界差距的方向。