Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

Mykola Vysotskyi,Runqi Lin,Grzegorz Biziel,Michal Zakrzewski,Sebastian Montagna,Damian Rynczak,Shreyansh Padarha,Kumail Alhamoud,Zihao Fu,William Lugoloobi,Kai Rawal,Hanna Yershova,Xander Davies,Taras Rumezhak,Guohao Li,Fazl Barez,Baoyuan Wu,Arkadiusz Drohomirecki,Yarin Gal,Chris Russell,Christopher Summerfield,Adam Mahdi,Volodymyr Karpiv,Philip Torr,Adel Bibi

As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchmarks are typically built on popular applications with relatively simple tasks and focus on a narrow set of capabilities while overlooking broader dimensions, resulting in saturated performance on modern agents and failing to probe their limitations. To this end, we introduce GauntletBench, a web-based benchmark for evaluating agent generalisation in challenging scenarios, focusing on three underexplored capabilities (temporal perception, graphical understanding, and 3D reasoning), across five less-covered professional applications (Video Editor, Workflow Builder, 3D Modeller, Flight Analyser, and Circuit Designer), each with 20 vision-intensive tasks (100 in total). Our benchmark provides a modular pipeline that comprises an environment compatible with both open- and closed-source agent frameworks, a controlled web-based application, a well-structured task suite, and an automated evaluation engine with diverse metrics. Contrary to widespread expectations, our empirical results reveal that frontier agentic systems remain far from achieving human-level performance. Even the state-of-the-art agent achieves only a 19.1% success rate on our GauntletBench, highlighting the limitations in these overlooked capabilities and generalisation. By comparison, non-expert human annotators achieve over 80% success on our challenging yet feasible tasks, revealing the substantial gap between current agent capabilities and those required for complex real-world scenarios.

翻译：随着智能体系统的不断演进并广泛部署于现实场景，对其能力进行忠实评估的需求日益增长。然而，当前的基准测试通常构建于任务相对简单的流行应用之上，且仅关注狭窄的能力维度而忽视更广泛的领域，导致现代智能体在该类测试中表现饱和，无法探测其局限性。为此，我们提出GauntletBench——一个基于网页的基准测试，专门用于评估智能体在挑战性场景中的泛化能力，聚焦于三项未被充分探索的能力（时间感知、图形理解和三维推理），覆盖五个较少涉及的专业应用领域（视频编辑器、工作流构建器、三维建模器、航程分析器和电路设计器），每个应用包含20项视觉密集型任务（总计100项）。我们的基准测试提供模块化流水线，包含兼容开源与闭源智能体框架的环境、可操控的网页应用程序、结构完善的任务套件以及配备多样化指标的自动化评估引擎。与普遍预期相反，实验结果表明前沿智能体系统仍远未达到人类水平。即使最先进的智能体在GauntletBench上的成功率也仅达19.1%，凸显了这些被忽视的能力及泛化性的局限。相比之下，非专家人类标注员在我们具有挑战性但可行的任务上实现了超过80%的成功率，揭示了当前智能体能力与复杂现实场景需求之间的显著差距。