Training models to act as agents that can effectively navigate and perform actions in a complex environment, such as a web browser, has typically been challenging due to lack of training data. Large language models (LLMs) have recently demonstrated some capability to navigate novel environments as agents in a zero-shot or few-shot fashion, purely guided by natural language instructions as prompts. Recent research has also demonstrated LLMs have the capability to exceed their base performance through self-improvement, i.e. fine-tuning on data generated by the model itself. In this work, we explore the extent to which LLMs can self-improve their performance as agents in long-horizon tasks in a complex environment using the WebArena benchmark. In WebArena, an agent must autonomously navigate and perform actions on web pages to achieve a specified objective. We explore fine-tuning on three distinct synthetic training data mixtures and achieve a 31\% improvement in task completion rate over the base model on the WebArena benchmark through a self-improvement procedure. We additionally contribute novel evaluation metrics for assessing the performance, robustness, capabilities, and quality of trajectories of our fine-tuned agent models to a greater degree than simple, aggregate-level benchmark scores currently used to measure self-improvement.
翻译:训练模型作为智能体在复杂环境(如网络浏览器)中有效导航并执行操作通常因缺乏训练数据而具有挑战性。近期研究表明,大型语言模型(LLMs)已展现出在零样本或少样本条件下作为智能体导航新环境的能力,其行为完全由自然语言指令提示引导。最新研究还表明,LLMs能够通过自我改进(即对模型自身生成的数据进行微调)来超越其基础性能。本研究基于WebArena基准测试,深入探索LLMs作为智能体在复杂环境中执行长周期任务时通过自我改进提升性能的潜力。在WebArena中,智能体必须自主导航网页并执行操作以实现特定目标。我们通过三种不同的合成训练数据混合策略进行微调,在自我改进过程中使WebArena基准测试的任务完成率较基础模型提升31%。此外,我们提出了一套新颖的评估指标,相较于当前用于衡量自我改进的简单聚合级基准分数,这些指标能更全面地评估微调后智能体模型的性能、鲁棒性、能力及轨迹质量。