We study repeated multi-player vector-valued games in which a player observes a payoff vector each round and evaluates outcomes through linear scalarizations of those vectors. Different from most prior works, the choice of scalarization is treated as an online decision variable rather than a fixed modeling decision. We propose a bi-level learning framework in which an outer learner chooses a scalarization from a finite candidate class on a slow timescale, while a faster inner bandit no-regret learner selects actions using the scalar feedback induced by the chosen scalarization. Performance of this approach is defined with respect to a certain true weight vector, and the deployed scalarizations act as control signals that shape the induced payoff trajectory. We provide implementable algorithms based on bandit online mirror descent with stabilized importance weighting, and we derive finite-time performance guarantees in the form of sublinear regret bounds. Experiments on a vector-valued extension of a canonical game show that convergence to the preferred equilibrium rises from roughly $50\%$ under non-adaptive scalarization to about $80\%$ under our proposed method.
翻译:我们研究重复多参与者向量值博弈问题,其中每轮博弈参与者观察到收益向量,并通过这些向量的线性标量化来评估结果。与以往多数研究不同,本文将标量化选择视为在线决策变量而非固定建模参数。我们提出一种双层学习框架:外层学习者以较慢时间尺度从有限候选类别中选择标量化,而内层采用更快速度的无遗憾强盗学习器,基于所选标量化产生的标量反馈选择动作。该方法的性能参照某个真实权重向量定义,实际部署的标量化作为控制信号塑造诱导收益轨迹。我们提出基于带稳定重要性加权的强盗在线镜像下降的可实现算法,并推导出具有次线性遗憾界的有限时间性能保证。在规范博弈的向量值扩展实验表明,在非自适应标量化下收敛至偏好均衡的概率约为50%,而本文方法可提升至约80%。