As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or revealed only when the right question is asked. Benchmarks that expose full intent upfront and grade only the final choice can neither pose this long-horizon challenge nor explain which requirement an agent missed. To address this gap, we introduce EComAgentBench, a benchmark of 662 tasks grounded in real Amazon products and reviews. Each task scatters these requirements across a visible query, a tool-gated profile, and scripted clarification; an agent must uncover hidden intent, verify candidates against attributes and review evidence, and commit to a single product within 100 tool calls. Moreover, typed, source-tagged rubrics grade every task, attributing each failure to a requirement and its source. Construction is automated yet reliable, with every answer fixed in code before any text is generated and every sample validated. Our evaluation of seven models reveals that even the strongest attains only 57.1% overall accuracy, and rubric satisfaction degrades from visible to hidden sources. Overall, we believe EComAgentBench will serve as a reproducible foundation for moving shopping agents from single-query search toward dependable assistance over long horizons.


翻译:随着基于大语言模型的购物代理进入实际应用,现有基准无法捕捉购物者需求的呈现方式:这些需求或隐含在查询语句中,或记录在用户画像中,或仅在提出恰当问题时才会显式。那些事先完全展示意图且仅对最终选择评分的基准,既无法构成长时域挑战,也无法解释代理遗漏了哪项需求。为填补这一空白,我们提出EComAgentBench——一个基于真实亚马逊商品及评论的662项任务评测基准。每项任务将需求分散在可见查询、工具管控的用户画像及脚本化的澄清对话中;代理必须揭示隐藏意图,对照属性及评论证据验证候选商品,并在100次工具调用内确定唯一商品。此外,带类型与来源标签的评分细则对每项任务进行评分,将每次失败归因于具体需求及其来源。构建过程自动化且可靠:所有答案在生成任何文本前就已编码固定,所有样本均经过验证。我们对七种模型的评估显示,即使最强模型也仅达到57.1%的整体准确率,且评分细则满足度从可见来源到隐藏来源逐步下降。总体而言,我们相信EComAgentBench将为推动购物代理从单次查询搜索向长时域可靠辅助过渡奠定可复现的基础。

0
下载
关闭预览

相关内容

【Facebook】人工智能基准(Benchmarking)测试再思考,55页ppt
专知会员服务
32+阅读 · 2020年12月20日
搜索query意图识别的演进
DataFunTalk
13+阅读 · 2020年11月15日
一文读懂目标检测:R-CNN、Fast R-CNN、Faster R-CNN、YOLO、SSD
七月在线实验室
11+阅读 · 2018年7月18日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
9+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
7+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
7+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
10+阅读 · 8月1日
相关VIP内容
【Facebook】人工智能基准(Benchmarking)测试再思考,55页ppt
专知会员服务
32+阅读 · 2020年12月20日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员