We study social learning dynamics where the agents collectively follow a simple multi-armed bandit protocol. Agents arrive sequentially, choose arms and receive associated rewards. Each agent observes the full history (arms and rewards) of the previous agents, and there are no private signals. While collectively the agents face exploration-exploitation tradeoff, each agent acts myopically, without regards to exploration. Motivating scenarios concern reviews and ratings on online platforms. We allow a wide range of myopic behaviors that are consistent with (parameterized) confidence intervals, including the "unbiased" behavior as well as various behaviorial biases. While extreme versions of these behaviors correspond to well-known bandit algorithms, we prove that more moderate versions lead to stark exploration failures, and consequently to regret rates that are linear in the number of agents. We provide matching upper bounds on regret by analyzing "moderately optimistic" agents. As a special case of independent interest, we obtain a general result on failure of the greedy algorithm in multi-armed bandits. This is the first such result in the literature, to the best of our knowledge
翻译:我们研究了一种社交学习动态过程,其中所有主体共同遵循简单的多臂赌博机协议。主体依次到达,选择臂并获取相应奖励。每个主体都能观察到之前所有主体的完整历史(包括臂选择和奖励),且不存在私有信号。虽然所有主体共同面临探索与利用的权衡,但每个主体均表现出短视行为,不考虑探索。典型应用场景涉及在线平台上的评论和评分。我们允许广泛符合(参数化)置信区间的短视行为,包括“无偏”行为以及各种行为偏差。虽然这些行为的极端版本对应于著名的赌博机算法,但我们证明更加适度的版本会导致严重的探索失败,进而产生与主体数量成线性关系的遗憾率。通过分析“适度乐观”的主体,我们给出了遗憾率的上界匹配。作为一个独立兴趣的特殊情况,我们得到了多臂赌博机中贪婪算法失败的一般性结论。据我们所知,这是文献中首次出现此类结果。