AI scientist systems are beginning to automate the production, evaluation, and iteration of scientific hypotheses. Their promise is speed; their risk is that speed also scales errors embedded in the scientific record. We argue that a near-term risk is corpus failure: AI scientist systems are trained on and grounded in a literature that over-represents positive results and under-represents null findings. We formalise this distortion as the null result gap, estimate it across three domains (drug discovery ~0.60, psychology ~0.56, cancer biology ~0.35), and introduce an amplification index for reasoning about how retrieval, generation, and automated evaluation can compound the raw gap. Using first-order estimates, we argue that a standard three-stage pipeline can amplify corpus distortion by a factor of 2.18x, with the conclusion unchanged under more conservative multipliers. We identify four governance failure modes: confident rediscovery, ghost evidence accumulation, replication laundering, and confidence miscalibration. We then propose three interventions: null-result databases as training infrastructure, retraction-aware evaluation metrics, and mandatory training corpus disclosure. The central takeaway is that AI scientists will not only accelerate science. Without governance, they will accelerate science's blind spots before they accelerate its discoveries.


翻译:AI科学家系统正开始自动化地生成、评估和迭代科学假设。其优势在于速度;其风险在于速度同样会放大科学文献中嵌入的错误。我们认为近期面临的风险是语料库失效:AI科学家系统所依赖的训练和基础文献过度呈现阳性结果,而阴性发现则被严重低估。我们将这种扭曲形式化为零结果差距,并在三个领域(药物发现约0.60、心理学约0.56、癌症生物学约0.35)进行了估算,同时引入一个放大指数来推演检索、生成和自动评估如何加剧原始差距。通过一阶估计,我们证明标准的三阶段流水线可将语料库扭曲放大2.18倍,即便采用更保守的乘数因子,结论依然成立。我们识别出四种治理失效模式:自信重现、幽灵证据积累、复制洗白以及置信度校准偏差。随后提出三项干预措施:将零结果数据库作为训练基础设施、采用感知撤回的评估指标,以及实施强制训练语料库披露。核心结论是:AI科学家不仅会加速科学进程。若缺乏治理,它们在加速科学发现之前,将先加速放大科学中的盲区。

0
下载
关闭预览

相关内容

人工智能杂志AI(Artificial Intelligence)是目前公认的发表该领域最新研究成果的主要国际论坛。该期刊欢迎有关AI广泛方面的论文,这些论文构成了整个领域的进步,也欢迎介绍人工智能应用的论文,但重点应该放在新的和新颖的人工智能方法如何提高应用领域的性能,而不是介绍传统人工智能方法的另一个应用。关于应用的论文应该描述一个原则性的解决方案,强调其新颖性,并对正在开发的人工智能技术进行深入的评估。 官网地址:http://dblp.uni-trier.de/db/journals/ai/
AutoResearch AI综述:迈向AI驱动的科学发现自动化
专知会员服务
18+阅读 · 5月26日
迈向医学人工智能科学家
专知会员服务
21+阅读 · 4月1日
【新书】生成式人工智能傻瓜书入门
专知会员服务
61+阅读 · 2024年9月24日
生成式人工智能,40页pdf
专知会员服务
105+阅读 · 2023年8月19日
【AI4Science】《人工智能科学:深度学习革命》2023新书,
专知会员服务
214+阅读 · 2023年6月15日
8月最新-《可解释机器学习-Christoph Molnar》-新书分享
深度学习与NLP
10+阅读 · 2019年8月12日
完备的 AI 学习路线,最详细的资源整理!
新智元
18+阅读 · 2019年5月4日
类脑计算的前沿论文,看我们推荐的这7篇
人工智能前沿讲习班
21+阅读 · 2019年1月7日
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Arxiv
0+阅读 · 6月11日
VIP会员
最新内容
非对称防御中的自组织临界性:俄乌战争
专知会员服务
8+阅读 · 8月10日
《战争中的大语言模型监管》
专知会员服务
8+阅读 · 8月10日
《边缘计算关键技术分析及美军作战实践应用》
边缘计算的军事应用
专知会员服务
11+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
12+阅读 · 8月8日
相关VIP内容
AutoResearch AI综述:迈向AI驱动的科学发现自动化
专知会员服务
18+阅读 · 5月26日
迈向医学人工智能科学家
专知会员服务
21+阅读 · 4月1日
【新书】生成式人工智能傻瓜书入门
专知会员服务
61+阅读 · 2024年9月24日
生成式人工智能,40页pdf
专知会员服务
105+阅读 · 2023年8月19日
【AI4Science】《人工智能科学:深度学习革命》2023新书,
专知会员服务
214+阅读 · 2023年6月15日
相关基金
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
12+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2012年12月31日
Top
微信扫码咨询专知VIP会员