Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain. We introduce RetailBench, a data-grounded simulation benchmark for evaluating tool-using LLM agents in single-store supermarket operation. RetailBench models retail management as a partially observable decision process and is designed to support thousand-day-scale simulations. In this environment, agents must manage pricing, replenishment, supplier selection, shelf assortment, inventory aging, customer feedback, external events, and cash-flow constraints. We evaluate seven contemporary LLMs under representative agent frameworks over a 180-day evaluation horizon and compare them with a privileged oracle policy. Results show substantial variation across models: only a small subset survives the full evaluation horizon, and even the strongest LLM runs remain substantially behind the oracle policy in final net worth and sales outcomes. Behavioral analysis attributes these gaps to incomplete evidence acquisition, surface-level decision making, and the lack of a consistent long-horizon policy. RetailBench provides a controlled testbed for studying reliable autonomy in economically grounded long-horizon decision-making.
翻译:摘要:大语言模型(LLM)代理在短期、范围明确的任务上取得了快速进展,但其在动态长期环境中维持连贯决策的能力仍不确定。我们提出RetailBench,一个基于数据驱动的仿真基准,用于评估单店超市运营场景中使用工具的LLM代理。RetailBench将零售管理建模为部分可观测的决策过程,并支持千天尺度仿真。在该环境中,代理需要管理定价、补货、供应商选择、货架组合、库存老化、客户反馈、外部事件及现金流约束。我们在180天评估周期内,基于代表性代理框架评估七种当代LLM,并将其与具有特权信息的参考策略进行比较。结果表明模型间存在显著差异:仅少数子集能完整存活于评估周期,即使最强的LLM运行在最终净资产和销售业绩上仍远落后于参考策略。行为分析将这些差距归因于证据获取不完整、表层决策模式以及缺乏一致的长期策略。RetailBench为研究经济导向长期决策中的可靠自主性提供了受控测试平台。