As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchmarks are typically built on popular applications with relatively simple tasks and focus on a narrow set of capabilities while overlooking broader dimensions, resulting in saturated performance on modern agents and failing to probe their limitations. To this end, we introduce GauntletBench, a web-based benchmark for evaluating agent generalisation in challenging scenarios, focusing on three underexplored capabilities (temporal perception, graphical understanding, and 3D reasoning), across five less-covered professional applications (Video Editor, Workflow Builder, 3D Modeller, Flight Analyser, and Circuit Designer), each with 20 vision-intensive tasks (100 in total). Our benchmark provides a modular pipeline that comprises an environment compatible with both open- and closed-source agent frameworks, a controlled web-based application, a well-structured task suite, and an automated evaluation engine with diverse metrics. Contrary to widespread expectations, our empirical results reveal that frontier agentic systems remain far from achieving human-level performance. Even the state-of-the-art agent achieves only a 19.1% success rate on our GauntletBench, highlighting the limitations in these overlooked capabilities and generalisation. By comparison, non-expert human annotators achieve over 80% success on our challenging yet feasible tasks, revealing the substantial gap between current agent capabilities and those required for complex real-world scenarios.


翻译:随着智能体系统的不断演进并广泛部署于现实场景,对其能力进行忠实评估的需求日益增长。然而,当前的基准测试通常构建于任务相对简单的流行应用之上,且仅关注狭窄的能力维度而忽视更广泛的领域,导致现代智能体在该类测试中表现饱和,无法探测其局限性。为此,我们提出GauntletBench——一个基于网页的基准测试,专门用于评估智能体在挑战性场景中的泛化能力,聚焦于三项未被充分探索的能力(时间感知、图形理解和三维推理),覆盖五个较少涉及的专业应用领域(视频编辑器、工作流构建器、三维建模器、航程分析器和电路设计器),每个应用包含20项视觉密集型任务(总计100项)。我们的基准测试提供模块化流水线,包含兼容开源与闭源智能体框架的环境、可操控的网页应用程序、结构完善的任务套件以及配备多样化指标的自动化评估引擎。与普遍预期相反,实验结果表明前沿智能体系统仍远未达到人类水平。即使最先进的智能体在GauntletBench上的成功率也仅达19.1%,凸显了这些被忽视的能力及泛化性的局限。相比之下,非专家人类标注员在我们具有挑战性但可行的任务上实现了超过80%的成功率,揭示了当前智能体能力与复杂现实场景需求之间的显著差距。

0
下载
关闭预览

相关内容

通用智能体评估的逻辑架构
专知会员服务
22+阅读 · 2月28日
智能体适应
专知会员服务
28+阅读 · 2025年12月11日
Agent AI:多模态交互的新地平线
专知会员服务
22+阅读 · 2025年5月26日
【NUS博士论文】面向交互的多智能体行为预测,156页pdf
专知会员服务
32+阅读 · 2024年11月17日
专知会员服务
98+阅读 · 2021年1月24日
【Facebook】人工智能基准(Benchmarking)测试再思考,55页ppt
专知会员服务
32+阅读 · 2020年12月20日
浅谈群体智能——新一代AI的重要方向
中国科学院自动化研究所
44+阅读 · 2019年10月16日
人工智能训练师的再定义
竹间智能Emotibot
11+阅读 · 2019年5月15日
群体智能:新一代人工智能的重要方向
走向智能论坛
12+阅读 · 2017年8月16日
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
28+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
4+阅读 · 今天7:07
《无人机空中监控:通信实验洞察》
专知会员服务
3+阅读 · 今天7:05
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
6+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
13+阅读 · 7月31日
相关VIP内容
通用智能体评估的逻辑架构
专知会员服务
22+阅读 · 2月28日
智能体适应
专知会员服务
28+阅读 · 2025年12月11日
Agent AI:多模态交互的新地平线
专知会员服务
22+阅读 · 2025年5月26日
【NUS博士论文】面向交互的多智能体行为预测,156页pdf
专知会员服务
32+阅读 · 2024年11月17日
专知会员服务
98+阅读 · 2021年1月24日
【Facebook】人工智能基准(Benchmarking)测试再思考,55页ppt
专知会员服务
32+阅读 · 2020年12月20日
相关基金
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
28+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
21+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员