Large language model (LLM) web agents are usually deployed as tool callers: each turn, the model reads a fresh page observation and emits one structured tool action. When every action is a low-level primitive, horizons grow quickly and so do policy-facing LLM completions, dominating latency and cost on benchmarks such as Mind2Web and WebArena. Recent systems therefore wrap repeated interaction fragments as web skills: callable tools built from successful trajectories or induced programs, so one call can replace several primitives. However, prior skill libraries are still triggered mainly by instruction similarity or coarse site metadata, which yields low skill reuse on held-out sites and leaves much of the potential step and token reduction on the table. We present SkillMigrator, an agent that learns reusable web skills and transfers them across sites by matching layout structure rather than specific element references. Each induced skill is stored as a transferable interaction pattern (TIP): the skill paired with a structural sketch of the snapshot at induction time. At test time, SkillMigrator retrieves TIPs by layout similarity and grounds their references on the live page. The rest of the stack is standard: accessibility-snapshot observations with stable references, and fixed tool calling over primitives plus skill invocations. Compared with the state-of-the-art approaches, SkillMigrator reduces the average LLM-action count on successful trajectories by 8-10% across both WebArena and Mind2Web at matched success rate.


翻译:大型语言模型(LLM)网络代理通常以工具调用方式部署:每轮交互中,模型读取新页面观察并生成结构化工具动作。当每个动作均为低级原语时,决策链迅速增长,导致面向策略的LLM补全次数激增,在Mind2Web和WebArena等基准测试中造成显著延迟与成本。为此,近期系统将重复交互片段封装为"网络技能"——基于成功轨迹或诱导程序构建的可调用工具,单次调用即可替代多个原语。然而,现有技能库仍主要依赖指令相似性或粗略站点元数据进行触发,导致在未见站点上技能复用率低下,潜在步骤与令牌缩减能力未能充分发挥。本文提出SkillMigrator代理,通过学习可复用网络技能并基于布局结构匹配(而非特定元素引用)实现跨站点迁移。每个诱导技能存储为可迁移交互模式(TIP):技能与诱导时刻快照的结构化草图配对。测试阶段,SkillMigrator通过布局相似性检索TIP,并将其引用锚定于实时页面。其余架构保持标准:采用稳定引用的无障碍快照观察,以及基于原语与技能调用的固定工具调用机制。与现有最优方法相比,在同等成功率下,SkillMigrator在WebArena和Mind2Bench成功轨迹上的平均LLM动作次数分别降低8-10%。

0
下载
关闭预览

相关内容

OpenAI 32页《智能体》指南,如何构建首个智能体系统
专知会员服务
51+阅读 · 2025年4月18日
2019年新书推荐-《神经网络与深度学习》-Michael Nielsen
深度学习与NLP
14+阅读 · 2019年2月21日
国家自然科学基金
18+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
最新内容
非对称防御中的自组织临界性:俄乌战争
专知会员服务
1+阅读 · 今天14:36
《战争中的大语言模型监管》
专知会员服务
2+阅读 · 今天14:26
边缘计算的军事应用
专知会员服务
8+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
9+阅读 · 8月8日
相关VIP内容
OpenAI 32页《智能体》指南,如何构建首个智能体系统
专知会员服务
51+阅读 · 2025年4月18日
相关基金
国家自然科学基金
18+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员