Two-sided online matching platforms have been employed in various markets. However, agents' preferences in present market are usually implicit and unknown and must be learned from data. With the growing availability of side information involved in the decision process, modern online matching methodology demands the capability to track preference dynamics for agents based on their contextual information. This motivates us to consider a novel Contextual Online Matching Bandit prOblem (COMBO), which allows dynamic preferences in matching decisions. Existing works focus on multi-armed bandit with static preference, but this is insufficient: the two-sided preference changes as along as one-side's contextual information updates, resulting in non-static matching. In this paper, we propose a Centralized Contextual - Explore Then Commit (CC-ETC) algorithm to adapt to the COMBO. CC-ETC solves online matching with dynamic preference. In theory, we show that CC-ETC achieves a sublinear regret upper bound O(log(T)) and is a rate-optimal algorithm by proving a matching lower bound. In the experiments, we demonstrate that CC-ETC is robust to variant preference schemes, dimensions of contexts, reward noise levels, and contexts variation levels.
翻译:双边在线匹配平台已应用于多种市场。然而,当前市场中参与者的偏好通常是隐式且未知的,必须从数据中学习。随着决策过程中涉及的辅助信息日益丰富,现代在线匹配方法要求具备基于上下文信息追踪参与者偏好动态的能力。这促使我们提出一个新颖的上下文在线匹配赌博机问题,该问题允许匹配决策中的动态偏好。现有研究集中于静态偏好的多臂赌博机,但这并不充分:只要一方的上下文信息发生更新,双边偏好就会随之改变,从而导致非静态匹配。本文提出了一种集中式上下文-先探索后提交算法来适应上述问题,该算法能够解决动态偏好下的在线匹配。理论分析表明,算法实现了次线性的累积遗憾上界,并通过证明匹配下界证明了其是速率最优算法。实验证明,算法对不同偏好方案、上下文维度、奖励噪声水平及上下文变化程度均具有鲁棒性。