As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchmarks are typically built on popular applications with relatively simple tasks and focus on a narrow set of capabilities while overlooking broader dimensions, resulting in saturated performance on modern agents and failing to probe their limitations. To this end, we introduce GauntletBench, a web-based benchmark for evaluating agent generalisation in challenging scenarios, focusing on three underexplored capabilities (temporal perception, graphical understanding, and 3D reasoning), across five less-covered professional applications (Video Editor, Workflow Builder, 3D Modeller, Flight Analyser, and Circuit Designer), each with 20 vision-intensive tasks (100 in total). Our benchmark provides a modular pipeline that comprises an environment compatible with both open- and closed-source agent frameworks, a controlled web-based application, a well-structured task suite, and an automated evaluation engine with diverse metrics. Contrary to widespread expectations, our empirical results reveal that frontier agentic systems remain far from achieving human-level performance. Even the state-of-the-art agent achieves only a 19.1% success rate on our GauntletBench, highlighting the limitations in these overlooked capabilities and generalisation. By comparison, non-expert human annotators achieve over 80% success on our challenging yet feasible tasks, revealing the substantial gap between current agent capabilities and those required for complex real-world scenarios.
翻译:随着智能体系统的不断演进并广泛部署于现实场景,对其能力进行忠实评估的需求日益增长。然而,当前的基准测试通常构建于任务相对简单的流行应用之上,且仅关注狭窄的能力维度而忽视更广泛的领域,导致现代智能体在该类测试中表现饱和,无法探测其局限性。为此,我们提出GauntletBench——一个基于网页的基准测试,专门用于评估智能体在挑战性场景中的泛化能力,聚焦于三项未被充分探索的能力(时间感知、图形理解和三维推理),覆盖五个较少涉及的专业应用领域(视频编辑器、工作流构建器、三维建模器、航程分析器和电路设计器),每个应用包含20项视觉密集型任务(总计100项)。我们的基准测试提供模块化流水线,包含兼容开源与闭源智能体框架的环境、可操控的网页应用程序、结构完善的任务套件以及配备多样化指标的自动化评估引擎。与普遍预期相反,实验结果表明前沿智能体系统仍远未达到人类水平。即使最先进的智能体在GauntletBench上的成功率也仅达19.1%,凸显了这些被忽视的能力及泛化性的局限。相比之下,非专家人类标注员在我们具有挑战性但可行的任务上实现了超过80%的成功率,揭示了当前智能体能力与复杂现实场景需求之间的显著差距。