Graphical User Interface (GUI) agents are increasingly used to automate complex computer tasks across applications, websites, and operating systems. To improve their reliability, recent work has introduced experiential memory, where agents retrieve prior trajectories to guide decision-making in similar states. More recent approaches further extend this idea to visual memory by storing and retrieving screenshots from past interactions, providing agents with richer contextual information than text-only memories. However, the effect of visual memory in GUI agents remains insufficiently understood: it is unclear which failures visual memory mitigates, or which failures it exacerbates. To systematically analyze the effect of visual memory, we introduce a taxonomy of four GUI agent failures (i.e., cognitive failure, visual state misunderstanding, hidden operation blindness, and grounding error) that map to distinct stages of the perception-reasoning-action pipeline. We find that prepending full-image memory has a divergent effect on the failure distribution: it reduces state-level failures but worsens action-level ones, and increases hidden operation blindness and grounding error. Motivated by this finding, we propose Action-Grounded Visual Memory (AGMem), an action-grounded memory framework for GUI agents. The core idea of AGMem is to store image crops that capture the local GUI region closely related to a successful action or a recovery, rather than storing full screenshots. Experiments on OSWorld show that AGMem improves task success rates by 33.3 % over full-image memory. These results demonstrate that AGMem is an effective representation for visual memory in GUI agents.


翻译:图形用户界面(GUI)代理越来越多地被用于自动化跨应用、网站和操作系统的复杂计算机任务。为了提高其可靠性,近期研究引入了经验记忆机制,使代理能够检索先前的轨迹以指导在相似状态下的决策。更先进的方法进一步将这一思想扩展到视觉记忆,通过存储和检索来自过去交互的截图,为代理提供比纯文本记忆更丰富的上下文信息。然而,视觉记忆在GUI代理中的影响尚未得到充分理解:目前尚不清楚视觉记忆能够缓解哪些失败,或者加剧哪些失败。为了系统分析视觉记忆的影响,我们提出了一种包含四种GUI代理失败类型(即认知失败、视觉状态误解、隐藏操作盲视和基础错误)的分类法,这些类型对应感知-推理-动作流程的不同阶段。我们发现,预置全图记忆对失败分布产生了分歧性影响:它减少了状态级失败,但加剧了动作级失败,并增加了隐藏操作盲视和基础错误。基于这一发现,我们提出了动作基础视觉记忆(AGMem),一种面向GUI代理的动作基础记忆框架。AGMem的核心思想是存储与成功动作或恢复操作密切相关的局部GUI区域图像裁剪,而非存储完整截图。在OSWorld上的实验表明,AGMem将任务成功率相较于全图记忆提升了33.3%。这些结果证明,AGMem是GUI代理中视觉记忆的有效表示方法。

0
下载
关闭预览

相关内容

《软件定义网络元素与机器代码的形式化验证》
专知会员服务
14+阅读 · 2025年11月18日
视觉弱监督学习研究进展
专知会员服务
32+阅读 · 2022年6月28日
【博士论文】视觉语言交互中的视觉推理研究
专知会员服务
65+阅读 · 2021年12月1日
专知会员服务
48+阅读 · 2021年7月2日
基于关系网络的视觉建模:有望替代卷积神经网络
微软研究院AI头条
10+阅读 · 2019年7月12日
【综述】计算机视觉简介:历史、现状和发展趋势【可下载】
机器学习算法与Python学习
15+阅读 · 2018年9月21日
计算机视觉简介:历史、现状和发展趋势
机器学习研究会
22+阅读 · 2017年11月21日
【观点】计算机视觉:历史、现状和发展趋势|胡占义研究员
中国科学院自动化研究所
14+阅读 · 2017年11月21日
国家自然科学基金
7+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 5月27日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
5+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
8+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
13+阅读 · 8月7日
相关基金
国家自然科学基金
7+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员