In this monograph, I introduce the basic concepts of Online Learning through a modern view of Online Convex Optimization. Here, online learning refers to the framework of regret minimization under worst-case assumptions. I present first-order and second-order algorithms for online learning with convex losses, in Euclidean and non-Euclidean settings. All the algorithms are clearly presented as instantiation of Online Mirror Descent or Follow-The-Regularized-Leader and their variants. Particular attention is given to the issue of tuning the parameters of the algorithms and learning in unbounded domains, through adaptive and parameter-free online learning algorithms. Non-convex losses are dealt through convex surrogate losses and through randomization. The bandit setting is also briefly discussed, touching on the problem of adversarial and stochastic multi-armed bandits. These notes do not require prior knowledge of convex analysis and all the required mathematical tools are rigorously explained. Moreover, all the included proofs have been carefully chosen to be as simple and as short as possible.
翻译:在本专著中,我通过在线凸优化的现代视角介绍在线学习的基本概念。这里的在线学习是指在最坏情况假设下进行遗憾最小化的框架。我介绍了在欧几里得和非欧几里得环境中处理凸损失的在线学习的一阶和二阶算法。所有算法均以在线镜像下降或跟随正则化领导者及其变体的实例形式清晰地呈现。特别关注算法参数的调整问题以及通过自适应和无参数在线学习算法在无界域中进行学习的问题。非凸损失通过凸替代损失和随机化处理。此外,简要讨论了赌博机设置,涉及对抗性和随机多臂赌博机问题。本讲义不要求具备凸分析先验知识,所有必需的数学工具均予以严格解释。此外,所有包含的证明均经过精心选择,力求尽可能简洁。