Logistic bandit is a ubiquitous framework of modeling users' choices, e.g., click vs. no click for advertisement recommender system. We observe that the prior works overlook or neglect dependencies in $S \geq \lVert \theta_\star \rVert_2$, where $\theta_\star \in \mathbb{R}^d$ is the unknown parameter vector, which is particularly problematic when $S$ is large, e.g., $S \geq d$. In this work, we improve the dependency on $S$ via a novel approach called {\it regret-to-confidence set conversion (R2CS)}, which allows us to construct a convex confidence set based on only the \textit{existence} of an online learning algorithm with a regret guarantee. Using R2CS, we obtain a strict improvement in the regret bound w.r.t. $S$ in logistic bandits while retaining computational feasibility and the dependence on other factors such as $d$ and $T$. We apply our new confidence set to the regret analyses of logistic bandits with a new martingale concentration step that circumvents an additional factor of $S$. We then extend this analysis to multinomial logistic bandits and obtain similar improvements in the regret, showing the efficacy of R2CS. While we applied R2CS to the (multinomial) logistic model, R2CS is a generic approach for developing confidence sets that can be used for various models, which can be of independent interest.
翻译:逻辑斯蒂老虎机是建模用户选择(如广告推荐系统中点击与未点击行为)的通用框架。我们注意到,现有工作忽视或忽略了$S \geq \lVert \theta_\star \rVert_2$中的依赖关系(其中$\theta_\star \in \mathbb{R}^d$为未知参数向量),当$S$较大(如$S \geq d$)时这一问题尤为突出。本研究通过提出一种名为"遗憾-置信集转换(R2CS)"的新方法,改进了对$S$的依赖关系。该方法仅基于存在遗憾保证的在线学习算法,即可构造凸置信集。采用R2CS后,我们在逻辑斯蒂老虎机中实现了关于$S$的遗憾界严格改进,同时保持计算可行性及对$d$、$T$等其他因素的依赖关系。我们将新置信集应用于逻辑斯蒂老虎机的遗憾分析,通过引入避免额外$S$因子的新鞅集中步骤进一步优化。进而将分析扩展到多项逻辑斯蒂老虎机,获得类似的遗憾改进,验证了R2CS的有效性。虽以(多项)逻辑斯蒂模型为例进行应用,但R2CS作为构建置信集的通用方法,可适用于多种模型,具有独立的研究价值。