Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon, high-value professional workflows across diverse domains. Current GUI benchmarks still predominantly focus on general-purpose software, relatively simple applications, and short-horizon tasks, leaving it largely unknown whether modern agents can follow user instructions to autonomously operate domain-specific professional software and accomplish economically valuable work in an end-to-end manner. To bridge this gap, we introduce Workflow-GYM, a benchmark for long-horizon GUI tasks centered on professional domains and specialized software environments. Through extensive experiments on state-of-the-art models, we find that even the strongest models achieve only slightly above 30% success rates, highlighting that professional long-horizon GUI workflows remain highly challenging for current GUI agents. Further analysis reveals that current agents struggle to maintain long-horizon workflow consistency, frequently exhibiting workflow stage omission, error propagation, objective drift, and insufficient understanding of professional software environments. Our findings provide important insights into the limitations of current agent systems and suggest key directions for the next generation of GUI-agent research.


翻译:近年来,人工智能代理向处理日益复杂、真实世界任务的快速演进已为人所瞩目。然而,现有基准测试极少评估代理能否操纵图形用户界面以跨领域完成长周期、高价值的专业工作流程。当前GUI基准测试仍主要聚焦于通用软件、相对简单的应用及短周期任务,致使现代代理能否遵循用户指令、自主操作专业领域特定软件并端到端完成具有经济价值的工作基本未知。为弥合这一差距,我们提出Workflow-GYM——一个聚焦专业领域与专用软件环境的长期GUI任务基准。通过对最先进模型的广泛实验,我们发现即便最强模型也仅能达成略高于30%的成功率,凸显出专业长周期GUI工作流程对当前GUI代理而言仍极具挑战性。进一步分析表明,当前代理难以维持长周期工作流程的一致性,频繁出现工作流阶段缺失、错误传播、目标偏移及对专业软件环境理解不足等问题。我们的发现为揭示当前代理系统的局限性提供了重要见解,并为下一代GUI代理研究指明了关键方向。

0
下载
关闭预览

相关内容

通用智能体评估的逻辑架构
专知会员服务
22+阅读 · 2月28日
【Facebook】人工智能基准(Benchmarking)测试再思考,55页ppt
专知会员服务
32+阅读 · 2020年12月20日
人工智能训练师的再定义
竹间智能Emotibot
11+阅读 · 2019年5月15日
爱奇艺基于AI的移动端自动化测试框架的设计
前端之巅
18+阅读 · 2019年2月27日
深度学习在推荐系统上的应用
架构文摘
13+阅读 · 2018年2月22日
尽早跑通深度学习的实践代码,是入门深度学习的最快途径
算法与数据结构
22+阅读 · 2017年12月13日
tensorflow项目学习路径
北京思腾合力科技有限公司
10+阅读 · 2017年11月23日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
人工智能与未来空战管理
专知会员服务
6+阅读 · 9月9日
机器的崛起:美海军陆战队组建机器人营思考
专知会员服务
9+阅读 · 9月8日
无面之战:人工智能如何重绘权力版图
专知会员服务
6+阅读 · 9月8日
《最强大的军事网状网络》
专知会员服务
9+阅读 · 9月7日
《预测陆军征兵任务分配》110页
专知会员服务
7+阅读 · 9月7日
相关VIP内容
通用智能体评估的逻辑架构
专知会员服务
22+阅读 · 2月28日
【Facebook】人工智能基准(Benchmarking)测试再思考,55页ppt
专知会员服务
32+阅读 · 2020年12月20日
相关资讯
人工智能训练师的再定义
竹间智能Emotibot
11+阅读 · 2019年5月15日
爱奇艺基于AI的移动端自动化测试框架的设计
前端之巅
18+阅读 · 2019年2月27日
深度学习在推荐系统上的应用
架构文摘
13+阅读 · 2018年2月22日
尽早跑通深度学习的实践代码,是入门深度学习的最快途径
算法与数据结构
22+阅读 · 2017年12月13日
tensorflow项目学习路径
北京思腾合力科技有限公司
10+阅读 · 2017年11月23日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员