Artificial intelligence research increasingly depends on prolonged cycles of reproduction, debugging, and iterative refinement to achieve State-Of-The-Art (SOTA) performance, creating a growing need for systems that can accelerate the full pipeline of empirical model optimization. In this work, we introduce AutoSOTA, an end-to-end automated research system that advances the latest SOTA models published in top-tier AI papers to reproducible and empirically improved new SOTA models. We formulate this problem through three tightly coupled stages: resource preparation and goal setting; experiment evaluation; and reflection and ideation. To tackle this problem, AutoSOTA adopts a multi-agent architecture with eight specialized agents that collaboratively ground papers to code and dependencies, initialize and repair execution environments, track long-horizon experiments, generate and schedule optimization ideas, and supervise validity to avoid spurious gains. We evaluate AutoSOTA on recent research papers collected from eight top-tier AI conferences under filters for code availability and execution cost. Across these papers, AutoSOTA achieves strong end-to-end performance in both automated replication and subsequent optimization. Specifically, it successfully discovers 105 new SOTA models that surpass the original reported methods, averaging approximately five hours per paper. Case studies spanning LLM, NLP, computer vision, time series, and optimization further show that the system can move beyond routine hyperparameter tuning to identify architectural innovation, algorithmic redesigns, and workflow-level improvements. These results suggest that end-to-end research automation can serve not only as a performance optimizer, but also as a new form of research infrastructure that reduces repetitive experimental burden and helps redirect human attention toward higher-level scientific creativity.
翻译:人工智能研究日益依赖于长时间的复现、调试和迭代优化才能达到最优性能(State-Of-The-Art, SOTA),这使得对加速整个经验性模型优化流程的系统需求不断增长。本文介绍AutoSOTA——一个端到端自动化研究系统,能够将顶级AI论文中发布的最新SOTA模型推进为可复现且经过实证改进的新SOTA模型。我们通过三个紧密耦合的阶段来形式化这一问题:资源准备与目标设定;实验评估;以及反思与构思。为解决该问题,AutoSOTA采用多智能体架构,包含八个专门智能体,它们协同工作以将论文与代码及依赖项关联、初始化并修复执行环境、追踪长期实验、生成并调度优化思路,以及监督结果有效性以避免虚假提升。我们在来自八场顶级AI会议的最新研究论文上,依据代码可用性和执行成本筛选后,对AutoSOTA进行了评估。在这些论文中,AutoSOTA在自动复现和后续优化两方面均实现了强大的端到端性能。具体而言,它成功发现了105个超越原始报告方法的新SOTA模型,平均每篇论文耗时约五小时。覆盖大语言模型(LLM)、自然语言处理(NLP)、计算机视觉、时间序列和优化领域的案例研究进一步表明,该系统能够超越常规超参数调优,识别出架构创新、算法重新设计以及工作流层面的改进。这些结果表明,端到端研究自动化不仅可作为性能优化器,更可成为一种新型研究基础设施,减少重复性实验负担,并帮助将人类注意力重新引导至更高层次的科学创造力。