We study the adversarial online learning problem and create a completely online algorithmic framework that has data dependent regret guarantees in both full expert feedback and bandit feedback settings. We study the expected performance of our algorithm against general comparators, which makes it applicable for a wide variety of problem scenarios. Our algorithm works from a universal prediction perspective and the performance measure used is the expected regret against arbitrary comparator sequences, which is the difference between our losses and a competing loss sequence. The competition class can be designed to include fixed arm selections, switching bandits, contextual bandits, periodic bandits or any other competition of interest. The sequences in the competition class are generally determined by the specific application at hand and should be designed accordingly. Our algorithm neither uses nor needs any preliminary information about the loss sequences and is completely online. Its performance bounds are data dependent, where any affine transform of the losses has no effect on the normalized regret.
翻译:我们研究了对抗性在线学习问题,并提出了一种完全在线的算法框架,在完整专家反馈和带状反馈设置下均具有基于数据的遗憾保证。我们分析了该算法相对于通用比较器的预期性能,使其适用于多种问题场景。该算法基于通用预测视角,采用的性能指标是相对于任意比较器序列的预期遗憾,即我们的损失与竞争损失序列之间的差值。竞争类别可设计为包含固定臂选择、切换赌博机、上下文赌博机、周期赌博机或任何其他感兴趣的竞争形式。竞争类别中的序列通常由具体应用场景决定,需相应设计。我们的算法既不使用也不依赖于损失序列的任何先验信息,完全在线运行。其性能界限基于数据,损失的任何仿射变换均不会影响归一化遗憾。