This paper formulates model selection as an infinite-armed bandit problem. The models are arms, and picking an arm corresponds to a partial training of the model (resource allocation). The reward is the accuracy of the selected model after its partial training. In this best arm identification problem, regret is the gap between the expected accuracy of the optimal model and that of the model finally chosen. We first consider a straightforward generalization of UCB-E to the stochastic infinite-armed bandit problem and show that, under basic assumptions, the expected regret order is $T^{-\alpha}$ for some $\alpha \in (0,1/5)$ and $T$ the number of resources to allocate. From this vanilla algorithm, we introduce the algorithm Mutant-UCB that incorporates operators from evolutionary algorithms. Tests carried out on three open source image classification data sets attest to the relevance of this novel combining approach, which outperforms the state-of-the-art for a fixed budget.
翻译:本文将模型选择问题形式化为无穷臂赌博机问题。其中模型被视为臂,选择臂对应于对模型进行部分训练(资源分配),奖励为所选模型经部分训练后的准确率。在该最优臂识别问题中,遗憾定义为最优模型期望准确率与最终选定模型期望准确率之间的差距。我们首先考虑将UCB-E算法直接推广至随机无穷臂赌博机问题,并证明在基本假设下,期望遗憾阶数为$T^{-\alpha}$(其中$\alpha \in (0,1/5)$,$T$为可分配资源数)。基于该基础算法,我们提出了融合进化算法算子的Mutant-UCB算法。在三个开源图像分类数据集上的测试结果表明,这种新型混合方法具有实际价值,在固定预算下其性能优于当前最优方法。