We study two reproducible failure modes of deep multi-agent reinforcement learning in continuous-time pricing markets: (i) tacit cartel formation between competing DDPG agents, and (ii) actor--critic instability at high event rates. We instantiate both inside a single CT-MARL benchmark (Poisson-clocked price updates, observation latency $δ$, interior-optimum logit demand), show that synchronous DDPG agents reliably trigger Failure Mode 1 with collusion index $Δ= 0.69 \pm 0.11$, and quantify a partial microstructure fix: asynchrony alone cuts collusion by 48\% and adding latency drives it to a minimum of $Δ= 0.28$. The fix has clearly documented costs: it is partial ($Δ$ remains supra-Bertrand), it is non-monotone in $δ$, and it does not survive Failure Mode 2, which emerges as DDPG critic divergence at $λ= 5$ and corrupts the phase-diagram cell at $(λ{=}5, δ{=}1)$. We accompany the scalar collusion index with trajectory-level trace diagnostics that expose the within-episode signalling collapse and the post-shock non-recovery.
翻译:我们研究了连续时间定价市场中深度多智能体强化学习的两种可复现失效模式:(i) 竞争性DDPG智能体之间形成隐性卡特尔,以及(ii)高事件率下的演员-评论家不稳定性。我们在单一CT-MARL基准测试(泊松时钟定价更新、观测延迟δ、内部最优对数需求)中实例化两者,证明同步DDPG智能体以共谋指数Δ=0.69±0.11可靠触发失效模式1,并量化了一种部分微观结构修复方案:仅异步性即可削减48%的共谋,而增加延迟则将其降至最低Δ=0.28。该修复方案具有明确记录的成本:它是部分的(Δ仍高于伯特兰水平)、在δ上非单调,且无法在失效模式2下存活——该模式在λ=5时表现为DDPG评论家发散,并污染了(λ=5, δ=1)处的相图单元格。我们伴随标量共谋指数提供轨迹级诊断,揭示回合内信号崩溃和冲击后不可恢复现象。