We present BLINQ, a new model-based algorithm that learns the Whittle indices of an indexable, communicating and unichain Markov Decision Process (MDP). Our approach relies on building an empirical estimate of the MDP and then computing its Whittle indices using an extended version of a state-of-the-art existing algorithm. We provide a proof of convergence to the Whittle indices we want to learn as well as a bound on the time needed to learn them with arbitrary precision. Moreover, we investigate its computational complexity. Our numerical experiments suggest that BLINQ significantly outperforms existing Q-learning approaches in terms of the number of samples needed to get an accurate approximation. In addition, it has a total computational cost even lower than Q-learning for any reasonably high number of samples. These observations persist even when the Q-learning algorithms are speeded up using neural networks to predict Q-values.
翻译:我们提出BLINQ,这是一种新的基于模型的算法,用于学习可索引、通信且单链马尔可夫决策过程(MDP)的惠特尔指数。该方法首先构建MDP的经验估计,然后利用现有最优算法的一个扩展版本计算其惠特尔指数。我们给出了收敛到目标惠特尔指数的证明,以及以任意精度学习这些指数所需时间的上界。此外,我们还研究了其计算复杂度。数值实验表明,在获得精确近似所需的样本数量方面,BLINQ显著优于现有的Q学习方法。同时,对于任何合理高数量的样本,其总计算成本甚至低于Q学习。即使使用神经网络加速Q值的预测,这一观察结果依然成立。