Artificial intelligence research increasingly depends on prolonged cycles of reproduction, debugging, and iterative refinement to achieve State-Of-The-Art (SOTA) performance, creating a growing need for systems that can accelerate the full pipeline of empirical model optimization. In this work, we introduce AutoSOTA, an end-to-end automated research system that advances the latest SOTA models published in top-tier AI papers to reproducible and empirically improved new SOTA models. We formulate this problem through three tightly coupled stages: resource preparation and goal setting; experiment evaluation; and reflection and ideation. To tackle this problem, AutoSOTA adopts a multi-agent architecture with eight specialized agents that collaboratively ground papers to code and dependencies, initialize and repair execution environments, track long-horizon experiments, generate and schedule optimization ideas, and supervise validity to avoid spurious gains. We evaluate AutoSOTA on recent research papers collected from eight top-tier AI conferences under filters for code availability and execution cost. Across these papers, AutoSOTA achieves strong end-to-end performance in both automated replication and subsequent optimization. Specifically, it successfully discovers 105 new SOTA models that surpass the original reported methods, averaging approximately five hours per paper. Case studies spanning LLM, NLP, computer vision, time series, and optimization further show that the system can move beyond routine hyperparameter tuning to identify architectural innovation, algorithmic redesigns, and workflow-level improvements. These results suggest that end-to-end research automation can serve not only as a performance optimizer, but also as a new form of research infrastructure that reduces repetitive experimental burden and helps redirect human attention toward higher-level scientific creativity.


翻译:人工智能研究日益依赖于长时间的复现、调试和迭代优化才能达到最优性能(State-Of-The-Art, SOTA),这使得对加速整个经验性模型优化流程的系统需求不断增长。本文介绍AutoSOTA——一个端到端自动化研究系统,能够将顶级AI论文中发布的最新SOTA模型推进为可复现且经过实证改进的新SOTA模型。我们通过三个紧密耦合的阶段来形式化这一问题:资源准备与目标设定;实验评估;以及反思与构思。为解决该问题,AutoSOTA采用多智能体架构,包含八个专门智能体,它们协同工作以将论文与代码及依赖项关联、初始化并修复执行环境、追踪长期实验、生成并调度优化思路,以及监督结果有效性以避免虚假提升。我们在来自八场顶级AI会议的最新研究论文上,依据代码可用性和执行成本筛选后,对AutoSOTA进行了评估。在这些论文中,AutoSOTA在自动复现和后续优化两方面均实现了强大的端到端性能。具体而言,它成功发现了105个超越原始报告方法的新SOTA模型,平均每篇论文耗时约五小时。覆盖大语言模型(LLM)、自然语言处理(NLP)、计算机视觉、时间序列和优化领域的案例研究进一步表明,该系统能够超越常规超参数调优,识别出架构创新、算法重新设计以及工作流层面的改进。这些结果表明,端到端研究自动化不仅可作为性能优化器,更可成为一种新型研究基础设施,减少重复性实验负担,并帮助将人类注意力重新引导至更高层次的科学创造力。

0
下载
关闭预览

相关内容

AutoResearch AI综述:迈向AI驱动的科学发现自动化
专知会员服务
18+阅读 · 5月26日
端到端自动驾驶系统研究综述
专知会员服务
31+阅读 · 2024年11月29日
概述自动机器学习(AutoML)
人工智能学家
19+阅读 · 2019年8月11日
深度解读:小米AI实验室AutoML团队最新成果FairNAS
PaperWeekly
32+阅读 · 2019年7月11日
AutoML研究综述:让AI学习设计AI
机器之心
15+阅读 · 2019年5月7日
【综述】自动机器学习AutoML最新65页综述,带你了解最新进展
中国人工智能学会
48+阅读 · 2019年5月3日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
VIP会员
最新内容
从采集到决策:美军视角下的战术情报范式重构
专知会员服务
0+阅读 · 今天2:42
《履带式无人地面战车技术发展现状》
专知会员服务
2+阅读 · 今天1:46
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员