This paper investigates whether online learning algorithms in pricing produce competitive outcomes or tacit collusion. This issue has drawn considerable attention from competition regulators as algorithmic pricing becomes more common in digital markets. Understanding when such algorithms lead to equilibrium or supra-competitive prices is critical for buyers, sellers, and policymakers. We study the behavior of multi-armed bandit (MAB) online learning algorithms in repeated price competition. These algorithms require little information to learn, making them realistic models of automated pricing. Our analysis shows that mean-based algorithms, a special variant of online learning algorithms, converge to correlated rationalizable actions. In the Bertrand environments considered, this implies convergence to the Nash equilibrium or adjacent prices. Numerical experiments reveal that most MAB algorithms, including those that are not mean-based, also converge. We observe supra-competitive prices only in specific cases where all sellers implement the same symmetric version of certain algorithms, such as UCB. This effect diminishes as the number of competitors increases. Our results suggest that, even in a stylized repeated Bertrand competition, sustained supra-competitive prices may be less of a concern when independent agents use different online learning algorithms. Our insights are relevant for regulators and managers considering the use of algorithmic pricing algorithms.
翻译:本文研究定价中的在线学习算法是否会产生竞争性结果或隐性合谋。随着算法定价在数字市场中日益普遍,这一问题已引起竞争监管机构的极大关注。理解这些算法何时导致均衡价格或超竞争价格,对买方、卖方和政策制定者至关重要。我们研究了重复价格竞争中多臂老虎机(MAB)在线学习算法的行为。这些算法所需信息极少即可学习,使其成为自动化定价的现实模型。我们的分析表明,基于均值的算法(在线学习算法的一种特殊变体)收敛于相关理性化行动。在所考虑的伯特兰竞争环境中,这意味着收敛于纳什均衡或邻近价格。数值实验揭示,大多数MAB算法(包括非基于均值的算法)也会收敛。仅在所有卖方实施某些算法(如UCB)的相同对称版本时,我们才观察到超竞争价格。随着竞争者数量增加,这种效应逐渐减弱。我们的结果表明,即使在简化的重复伯特兰竞争中,当独立主体使用不同的在线学习算法时,持续的超竞争价格可能并非主要担忧。我们的见解对考虑使用算法定价算法的监管者和管理者具有参考意义。