Recently multi-armed bandit problem arises in many real-life scenarios where arms must be sampled in batches, due to limited time the agent can wait for the feedback. Such applications include biological experimentation and online marketing. The problem is further complicated when the number of arms is large and the number of batches is small. We consider pure exploration in a batched multi-armed bandit problem. We introduce a general linear programming framework that can incorporate objectives of different theoretical settings in best arm identification. The linear program leads to a two-stage algorithm that can achieve good theoretical properties. We demonstrate by numerical studies that the algorithm also has good performance compared to certain UCB-type or Thompson sampling methods.
翻译:近年来,多臂老虎机问题出现在许多实际场景中,由于智能体等待反馈的时间有限,臂必须按批次进行采样。此类应用包括生物实验和在线营销。当臂的数量较多而批次数较少时,该问题进一步复杂化。本文研究了批处理多臂老虎机问题中的纯探索。我们引入了一个通用线性规划框架,该框架能够整合最优臂识别中不同理论设定的目标。该线性规划导出一个两阶段算法,可实现良好的理论性质。通过数值研究,我们证明该算法相比某些基于UCB或汤普森采样的方法也具有良好性能。