Reinforcement learning with multinomial logistic (MNL) function approximation has become an important framework due to its flexibility and broad applicability. While existing studies have established regret guarantees under worst-case analysis, they do not capture how performance depends on the variability of the interaction between the learner and the environment. In this paper, we develop a new theoretical analysis for MNL-based Markov decision processes that yields explicit variance-adaptive regret bounds. Our algorithm is computationally efficient and achieves the instance-wise optimal rate of regret, narrowing the gap between upper and lower bounds. Our numerical experiments validate that our method learns optimal policies more efficiently than conventional approaches.
翻译:多项式逻辑(MNL)函数逼近下的强化学习因其灵活性和广泛适用性已成为重要研究框架。现有研究虽在最坏情况分析下建立了遗憾界,但未能刻画性能如何依赖于学习器与交互环境之间的变异性。本文针对基于MNL的马尔可夫决策过程提出了新的理论分析框架,得出了显式的方差自适应遗憾界。所提算法计算高效且实现了实例最优遗憾率,收窄了上下界之间的差距。数值实验验证了该方法比传统方法能更高效地学习最优策略。