Multi-armed bandits a simple but very powerful framework for algorithms that make decisions over time under uncertainty. An enormous body of work has accumulated over the years, covered in several books and surveys. This book provides a more introductory, textbook-like treatment of the subject. Each chapter tackles a particular line of work, providing a self-contained, teachable technical introduction and a brief review of the further developments; many of the chapters conclude with exercises. The book is structured as follows. The first four chapters are on IID rewards, from the basic model to impossibility results to Bayesian priors to Lipschitz rewards. The next three chapters cover adversarial rewards, from the full-feedback version to adversarial bandits to extensions with linear rewards and combinatorially structured actions. Chapter 8 is on contextual bandits, a middle ground between IID and adversarial bandits in which the change in reward distributions is completely explained by observable contexts. The last three chapters cover connections to economics, from learning in repeated games to bandits with supply/budget constraints to exploration in the presence of incentives. The appendix provides sufficient background on concentration and KL-divergence. The chapters on "bandits with similarity information", "bandits with knapsacks" and "bandits and agents" can also be consumed as standalone surveys on the respective topics.
翻译:多臂老虎机是一种简单但极为强大的算法框架,用于在不确定性环境下逐步做出决策。多年来,该领域积累了大量的研究成果,多部专著与综述论文已对其进行了系统总结。本书以教科书式的入门方式对该主题进行讲解。每一章聚焦某一特定研究方向,提供自洽且适于教学的技术性导论,并简要回顾后续进展;多数章节末尾附有习题。全书结构如下:前四章讨论独立同分布奖励,涵盖基础模型、不可能性结果、贝叶斯先验及Lipschitz奖励;随后三章介绍对抗性奖励,涉及全反馈版本、对抗式老虎机及其在线性奖励与组合结构动作上的扩展;第八章讨论上下文老虎机,其奖励分布变化完全由可观测上下文解释,是独立同分布与对抗性老虎机之间的过渡模型;最后三章探讨与经济学的联系,包括重复博弈中的学习、带供应/预算约束的老虎机,以及激励存在下的探索行为。附录对浓度不等式及KL散度进行了充分背景介绍。关于"带相似性信息的老虎机"、"带背包约束的老虎机"及"老虎机与智能体"的章节亦可作为独立综述阅读。