The progress of autonomous web navigation has been hindered by the dependence on billions of exploratory interactions via online reinforcement learning, and domain-specific model designs that make it difficult to leverage generalization from rich out-of-domain data. In this work, we study data-driven offline training for web agents with vision-language foundation models. We propose an instruction-following multimodal agent, WebGUM, that observes both webpage screenshots and HTML pages and outputs web navigation actions, such as click and type. WebGUM is trained by jointly finetuning an instruction-finetuned language model and a vision encoder with temporal and local perception on a large corpus of demonstrations. We empirically demonstrate this recipe improves the agent's ability of grounded multimodal perception, HTML comprehension, and multi-step reasoning, outperforming prior works by a significant margin. On the MiniWoB, we improve over the previous best offline methods by more than 45.8%, even outperforming online-finetuned SoTA, humans, and GPT-4-based agent. On the WebShop benchmark, our 3-billion-parameter model achieves superior performance to the existing SoTA, PaLM-540B. Furthermore, WebGUM exhibits strong positive transfer to the real-world planning tasks on the Mind2Web. We also collect 347K high-quality demonstrations using our trained models, 38 times larger than prior work, and make them available to promote future research in this direction.
翻译:自主网页导航的进展一直受限于对在线强化学习中数十亿次探索性交互的依赖,以及难以利用丰富域外数据泛化能力的领域特定模型设计。本研究探索了基于视觉-语言基础模型的网页智能体数据驱动离线训练方法。我们提出了一种遵循指令的多模态智能体WebGUM,该模型同时观察网页截图与HTML页面,输出点击、键入等网页导航动作。通过联合微调指令微调语言模型与配备时序及局部感知能力的视觉编码器,我们在大规模演示数据集上对WebGUM进行训练。实验证明,该方法显著提升了智能体在具身多模态感知、HTML理解及多步推理方面的能力,性能大幅超越以往工作。在MiniWoB基准上,我们较先前最优离线方法提升超过45.8%,甚至超越了在线微调的最先进模型(SoTA)、人类及基于GPT-4的智能体。在WebShop基准测试中,我们30亿参数的模型性能超越现有SoTA模型PaLM-540B。此外,WebGUM在Mind2Web的真实世界规划任务中展现出强大的正向迁移能力。我们利用训练模型收集了347K条高质量演示数据(规模为先前工作的38倍),并已公开以推动该方向的未来研究。