The objective of the KPR agents are to learn themselves in the minimum (learning) time to have maximum success or utilization probability ($f$). A dictator can easily solve the problem with $f = 1$ in no time, by asking every one to form a queue and go to the respective restaurant, resulting in no fluctuation and full utilization from the first day (convergence time $\tau = 0$). It has already been shown that if each agent chooses randomly the restaurants, $f = 1 - e^{-1} \simeq 0.63$ (where $e \simeq 2.718$ denotes the Euler number) in zero time ($\tau = 0$). With the only available information about yesterday's crowd size in the restaurant visited by the agent (as assumed for the rest of the strategies studied here), the crowd avoiding (CA) strategies can give higher values of $f$ but also of $\tau$. Several numerical studies of modified learning strategies actually indicated increased value of $f = 1 - \alpha$ for $\alpha \to 0$, with $\tau \sim 1/\alpha$. We show here using Monte Carlo technique, a modified Greedy Crowd Avoiding (GCA) Strategy can assure full utilization ($f = 1$) in convergence time $\tau \simeq eN$, with of course non-zero probability for an even larger convergence time. All these observations suggest that the strategies with single step memory of the individuals can never collectively achieve full utilization ($f = 1$) in finite convergence time and perhaps the maximum possible utilization that can be achieved is about eighty percent ($f \simeq 0.80$) in an optimal time $\tau$ of order ten, even when $N$ the number of customers or of the restaurants goes to infinity.
翻译:科莱帕耶斯餐馆问题(KPR)中,智能体的目标是在最小(学习)时间内实现最大成功或利用率概率($f$)。一个独裁者能在瞬间轻松解决该问题,通过要求所有人排队并前往各自餐馆,从而从第一天起实现无波动且完全利用(收敛时间$\tau = 0$)。已有研究表明,若每个智能体随机选择餐馆,则$f = 1 - e^{-1} \simeq 0.63$(其中$e \simeq 2.718$为欧拉数),且收敛时间$\tau = 0$。仅凭智能体所访餐馆昨日人流量信息(本文后续策略均基于此假设),避众(CA)策略虽能提升$f$值,但也会增加$\tau$值。多组改进学习策略的数值研究实际上表明,当$\alpha \to 0$时,$f = 1 - \alpha$增大,且$\tau \sim 1/\alpha$。本文通过蒙特卡洛方法证明,改进的贪婪避众(GCA)策略可在收敛时间$\tau \simeq eN$内确保完全利用($f = 1$),但需注意存在更大收敛时间的非零概率。所有观测结果表明,具有单步记忆的个体策略永远无法在有限收敛时间内集体实现完全利用($f = 1$),而即便当顾客数或餐馆数$N$趋于无穷时,最优时间内可实现的最大利用率可能约为百分之八十($f \simeq 0.80$),对应收敛时间$\tau$数量级为十。