Training computer-use agents (CUAs) -- models that interact with graphical desktops through screenshots and keyboard/mouse actions -- requires large-scale, diverse trajectory data collected in full desktop environments. The largest public resource, AgentNet (22.5K human trajectories), leads to negative transfer when used for supervised fine-tuning (SFT): continuing training UI-TARS 7B on AgentNet causes OSWorld success rate to fall from 26.3% to 8-10%. We present ProCUA-SFT, a dataset of 3.1M step-level SFT samples distilled from 93K synthetic trajectories across 2,484 application combinations. The dataset is produced by a fully automated pipeline that (i) synthesizes grounded tasks on live desktops seeded with real-world content -- 912 spreadsheets from SpreadsheetBench, approximately 10K permissively-licensed presentations from Zenodo10K, and multi-application OSWorld configs -- and (ii) verifies each task's feasibility through binary precondition checking before rollout. A single VLM (Kimi-K2.5) serves as goal generator, precondition judge, and trajectory executor, eliminating planner-actor capability gaps. Each trajectory is expanded into step-prefix samples that exactly reproduce the context layout seen at inference time. Fine-tuning UI-TARS 7B on ProCUA-SFT for one epoch yields 45.0% on OSWorld -- an 18.7 percentage-point improvement over the base model and over 35% above AgentNet-trained counterparts. A subset of ProCUA was incorporated into the training data for the Nemotron 3 Nano Omni model, contributing to its computer-use capabilities.


翻译:训练计算机使用智能体——即通过截屏和键盘/鼠标操作与图形桌面交互的模型——需要在完整桌面环境中收集大规模、多样化的轨迹数据。目前最大的公开资源AgentNet(含2.25万条人类轨迹)在用于监督微调时会导致负迁移:在UI-TARS 7B模型上继续训练AgentNet数据,将使其在OSWorld任务上的成功率从26.3%降至8-10%。本文提出ProCUA-SFT数据集,包含从涵盖2484种应用组合的9.3万条合成轨迹中蒸馏得到的310万步级SFT样本。该数据集通过全自动流水线生成,其流程包括:(i)在注入真实世界内容的活跃桌面上合成接地任务——这些内容包括来自SpreadsheetBench的912个电子表格、从Zenodo10K获取的约1万个开放许可演示文稿,以及多应用OSWorld配置;(ii)在轨迹执行前通过二元前置条件检查验证每项任务的可行性。单个视觉语言模型(Kimi-K2.5)同时担任目标生成器、前置条件判断器和轨迹执行器,从而消除了规划器与执行器之间的能力鸿沟。每条轨迹被扩展为步前缀样本,精确复现推理时的上下文布局。在ProCUA-SFT上对UI-TARS 7B模型进行单周期微调后,其在OSWorld上的成功率提升至45.0%:较基础模型提高18.7个百分点,较使用AgentNet训练的模型高出35%以上。ProCUA的子集已被纳入Nemotron 3 Nano Omni模型的训练数据,为其计算机使用能力的提升做出了贡献。

0
下载
关闭预览

相关内容

CVPR 2022 将于2022年 6 月 21-24 日在美国的新奥尔良举行。CVPR是IEEE Conference on Computer Vision and Pattern Recognition的缩写,即IEEE国际计算机视觉与模式识别会议。该会议是由IEEE举办的计算机视觉和模式识别领域的顶级会议,会议的主要内容是计算机视觉与模式识别技术。

知识荟萃

精品入门和进阶教程、论文和代码整理等

更多

查看相关VIP内容、论文、资讯等
《用于蜂群相对定位的射频测距技术》64页技术报告
专知会员服务
28+阅读 · 2024年4月27日
《大模型驱动的汽车行业群体智能技术白皮书》,176页pdf
《TextCycleGAN 技术报告》
专知会员服务
34+阅读 · 2023年5月4日
《建立智能体-仿真物技术关系 (ASTR)》美国陆军55页报告
专知会员服务
40+阅读 · 2023年3月28日
最新《弱监督预训练语言模型微调》报告,52页ppt
专知会员服务
38+阅读 · 2020年12月26日
技术动态 | TechKG:一个面向中文学术领域的大型知识图谱
开放知识图谱
25+阅读 · 2018年12月20日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
0+阅读 · 6月16日
Arxiv
0+阅读 · 6月14日
Arxiv
0+阅读 · 6月5日
Arxiv
0+阅读 · 6月4日
Arxiv
0+阅读 · 5月28日
Arxiv
0+阅读 · 5月18日
Arxiv
0+阅读 · 5月6日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
9+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
6+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
9+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
14+阅读 · 7月31日
相关论文
Arxiv
0+阅读 · 6月16日
Arxiv
0+阅读 · 6月14日
Arxiv
0+阅读 · 6月5日
Arxiv
0+阅读 · 6月4日
Arxiv
0+阅读 · 5月28日
Arxiv
0+阅读 · 5月18日
Arxiv
0+阅读 · 5月6日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员