In many online sequential decision-making scenarios, a learner's choices affect not just their current costs but also the future ones. In this work, we look at one particular case of such a situation where the costs depend on the time average of past decisions over a history horizon. We first recast this problem with history dependent costs as a problem of decision making under stage-wise constraints. To tackle this, we then propose the novel Follow-The-Adaptively-Regularized-Leader (FTARL) algorithm. Our innovative algorithm incorporates adaptive regularizers that depend explicitly on past decisions, allowing us to enforce stage-wise constraints while simultaneously enabling us to establish tight regret bounds. We also discuss the implications of the length of history horizon on design of no-regret algorithms for our problem and present impossibility results when it is the full learning horizon.
翻译:在许多在线序贯决策场景中,学习者的选择不仅影响当前成本,也会影响未来成本。本研究聚焦于一种特定情形:成本取决于历史时间窗口内过去决策的时间平均值。我们首先将此类历史依赖成本问题重新表述为具有阶段约束的决策问题。为解决该问题,我们提出了新型的"自适应正则化领导者跟随"(FTARL)算法。该创新算法采用显式依赖于过去决策的自适应正则化项,使我们能在施加阶段约束的同时,推导出紧凑的遗憾界。我们还探讨了历史窗口长度对设计无遗憾算法的影响,并展示了当历史窗口为完整学习周期时的不可行性结果。