We consider the setting of repeated fair division between two players, denoted Alice and Bob, with private valuations over a cake. In each round, a new cake arrives, which is identical to the ones in previous rounds. Alice cuts the cake at a point of her choice, while Bob chooses the left piece or the right piece, leaving the remainder for Alice. We consider two versions: sequential, where Bob observes Alice's cut point before choosing left/right, and simultaneous, where he only observes her cut point after making his choice. The simultaneous version was first considered by Aumann and Maschler (1995). We observe that if Bob is almost myopic and chooses his favorite piece too often, then he can be systematically exploited by Alice through a strategy akin to a binary search. This strategy allows Alice to approximate Bob's preferences with increasing precision, thereby securing a disproportionate share of the resource over time. We analyze the limits of how much a player can exploit the other one and show that fair utility profiles are in fact achievable. Specifically, the players can enforce the equitable utility profile of $(1/2, 1/2)$ in the limit on every trajectory of play, by keeping the other player's utility to approximately $1/2$ on average while guaranteeing they themselves get at least approximately $1/2$ on average. We show this theorem using a connection with Blackwell approachability. Finally, we analyze a natural dynamic known as fictitious play, where players best respond to the empirical distribution of the other player. We show that fictitious play converges to the equitable utility profile of $(1/2, 1/2)$ at a rate of $O(1/\sqrt{T})$.
翻译:我们考虑两个玩家(爱丽丝和鲍勃)在私有蛋糕估值下的重复公平分配场景。每轮都会出现与先前轮次相同的新蛋糕。爱丽丝选择切点切蛋糕,鲍勃则选择左边或右边的一块,剩余部分归爱丽丝。我们研究两种版本:顺序版本,鲍勃在观察到爱丽丝的切点后再选择左右;同时版本,鲍勃在做出选择后才观察到切点。同时版本最初由Aumann和Maschler(1995)提出。我们发现,若鲍勃近乎短视且频繁选择偏爱的一块,爱丽丝可通过类似二分搜索的策略系统性利用他。该策略使爱丽丝以递增精度逼近鲍勃的偏好,从而长期获取不成比例的份额。我们分析玩家利用对方的理论极限,并证明公平效用配置文件实际可达。具体而言,玩家可通过在每轮博弈轨迹中限制对方平均效用约为1/2,同时保证自身平均效用至少约为1/2,逐步强制执行(1/2, 1/2)的公平效用配置。我们通过Blackwell可逼近性定理证明该结论。最后,分析名为“虚拟博弈”的自然动态过程,其中玩家针对对方经验分布做出最优反应。我们证明虚拟博弈以O(1/√T)的速率收敛至(1/2, 1/2)的公平效用配置。