Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain. We introduce RetailBench, a data-grounded simulation benchmark for evaluating tool-using LLM agents in single-store supermarket operation. RetailBench models retail management as a partially observable decision process and is designed to support thousand-day-scale simulations. In this environment, agents must manage pricing, replenishment, supplier selection, shelf assortment, inventory aging, customer feedback, external events, and cash-flow constraints. We evaluate seven contemporary LLMs under representative agent frameworks over a 180-day evaluation horizon and compare them with a privileged oracle policy. Results show substantial variation across models: only a small subset survives the full evaluation horizon, and even the strongest LLM runs remain substantially behind the oracle policy in final net worth and sales outcomes. Behavioral analysis attributes these gaps to incomplete evidence acquisition, surface-level decision making, and the lack of a consistent long-horizon policy. RetailBench provides a controlled testbed for studying reliable autonomy in economically grounded long-horizon decision-making.


翻译:摘要:大语言模型(LLM)代理在短期、范围明确的任务上取得了快速进展,但其在动态长期环境中维持连贯决策的能力仍不确定。我们提出RetailBench,一个基于数据驱动的仿真基准,用于评估单店超市运营场景中使用工具的LLM代理。RetailBench将零售管理建模为部分可观测的决策过程,并支持千天尺度仿真。在该环境中,代理需要管理定价、补货、供应商选择、货架组合、库存老化、客户反馈、外部事件及现金流约束。我们在180天评估周期内,基于代表性代理框架评估七种当代LLM,并将其与具有特权信息的参考策略进行比较。结果表明模型间存在显著差异:仅少数子集能完整存活于评估周期,即使最强的LLM运行在最终净资产和销售业绩上仍远落后于参考策略。行为分析将这些差距归因于证据获取不完整、表层决策模式以及缺乏一致的长期策略。RetailBench为研究经济导向长期决策中的可靠自主性提供了受控测试平台。

0
下载
关闭预览

相关内容

大语言模型中的隐式推理:综合综述
专知会员服务
34+阅读 · 2025年9月4日
重新思考不确定性:大语言模型时代的关键综述与分析
专知会员服务
39+阅读 · 2024年11月20日
大型语言模型代理的安全与隐私综述
专知会员服务
30+阅读 · 2024年8月5日
大语言模型的终身学习综述
专知会员服务
77+阅读 · 2024年6月15日
《大型语言模型持续学习》综述
专知会员服务
94+阅读 · 2024年4月26日
大型语言模型高效推理综述
专知会员服务
65+阅读 · 2024年4月23日
「大型语言模型评测」综述
专知会员服务
70+阅读 · 2024年3月30日
自然语言处理中的语言模型预训练方法
PaperWeekly
14+阅读 · 2018年10月21日
推荐算法:Match与Rank模型的交织配合
从0到1
15+阅读 · 2017年12月18日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Arxiv
14+阅读 · 2023年8月7日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
10+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
8+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
8+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
10+阅读 · 8月1日
相关VIP内容
大语言模型中的隐式推理:综合综述
专知会员服务
34+阅读 · 2025年9月4日
重新思考不确定性:大语言模型时代的关键综述与分析
专知会员服务
39+阅读 · 2024年11月20日
大型语言模型代理的安全与隐私综述
专知会员服务
30+阅读 · 2024年8月5日
大语言模型的终身学习综述
专知会员服务
77+阅读 · 2024年6月15日
《大型语言模型持续学习》综述
专知会员服务
94+阅读 · 2024年4月26日
大型语言模型高效推理综述
专知会员服务
65+阅读 · 2024年4月23日
「大型语言模型评测」综述
专知会员服务
70+阅读 · 2024年3月30日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员