The knowledge gradient (KG) algorithm is a popular policy for the best arm identification (BAI) problem. It is built on the simple idea of always choosing the measurement that yields the greatest expected one-step improvement in the estimate of the best mean of the arms. In this research, we show that this policy has limitations, causing the algorithm not asymptotically optimal. We next provide a remedy for it, by following the manner of one-step look ahead of KG, but instead choosing the measurement that yields the greatest one-step improvement in the probability of selecting the best arm. The new policy is called improved knowledge gradient (iKG). iKG can be shown to be asymptotically optimal. In addition, we show that compared to KG, it is easier to extend iKG to variant problems of BAI, with the $\epsilon$-good arm identification and feasible arm identification as two examples. The superior performances of iKG on these problems are further demonstrated using numerical examples.
翻译:知识梯度(KG)算法是解决最佳臂识别(BAI)问题的一种常用策略。其核心思想十分简单:始终选择能够使臂均值最优估计的期望单步改进量最大的测量方式。本研究表明,该策略存在局限性,导致算法无法达到渐近最优性。为此,我们遵循KG算法单步前瞻的思想提出改进方案,转而选择能够使选择最优臂概率的单步改进量最大的测量方式。新策略被称为改进知识梯度(iKG)算法。可以证明iKG具有渐近最优性。此外,研究表明相较于KG,iKG更容易扩展至BAI的变体问题,并以$\epsilon$-优臂识别和可行臂识别为例进行了验证。数值实验进一步证明了iKG在这些问题上的优越性能。