We propose a novel combinatorial stochastic-greedy bandit (SGB) algorithm for combinatorial multi-armed bandit problems when no extra information other than the joint reward of the selected set of $n$ arms at each time step $t\in [T]$ is observed. SGB adopts an optimized stochastic-explore-then-commit approach and is specifically designed for scenarios with a large set of base arms. Unlike existing methods that explore the entire set of unselected base arms during each selection step, our SGB algorithm samples only an optimized proportion of unselected arms and selects actions from this subset. We prove that our algorithm achieves a $(1-1/e)$-regret bound of $\mathcal{O}(n^{\frac{1}{3}} k^{\frac{2}{3}} T^{\frac{2}{3}} \log(T)^{\frac{2}{3}})$ for monotone stochastic submodular rewards, which outperforms the state-of-the-art in terms of the cardinality constraint $k$. Furthermore, we empirically evaluate the performance of our algorithm in the context of online constrained social influence maximization. Our results demonstrate that our proposed approach consistently outperforms the other algorithms, increasing the performance gap as $k$ grows.
翻译:我们提出一种新颖的组合式随机贪婪赌博机算法,用于解决组合多臂赌博机问题,该问题中每个时间步$t\in [T]$仅能观测到所选$n$个臂的联合奖励,而无其他额外信息。SGB采用优化的随机探索-然后-提交策略,专门针对基臂数量庞大的场景设计。与现有方法在每个选择步骤中探索全部未选择基臂不同,我们的SGB算法仅采样优化比例的未选择臂,并从中选取动作。我们证明,该算法在单调随机子模奖励下可达到$(1-1/e)$-遗憾界$\mathcal{O}(n^{\frac{1}{3}} k^{\frac{2}{3}} T^{\frac{2}{3}} \log(T)^{\frac{2}{3}})$,在基数约束$k$上优于现有最优方法。此外,我们通过在线受限社交影响力最大化问题对算法性能进行实证评估。结果表明,我们提出的方法持续优于其他算法,且随着$k$增大性能差距进一步扩大。