AI scientist systems are beginning to automate the production, evaluation, and iteration of scientific hypotheses. Their promise is speed; their risk is that speed also scales errors embedded in the scientific record. We argue that a near-term risk is corpus failure: AI scientist systems are trained on and grounded in a literature that over-represents positive results and under-represents null findings. We formalise this distortion as the null result gap, estimate it across three domains (drug discovery ~0.60, psychology ~0.56, cancer biology ~0.35), and introduce an amplification index for reasoning about how retrieval, generation, and automated evaluation can compound the raw gap. Using first-order estimates, we argue that a standard three-stage pipeline can amplify corpus distortion by a factor of 2.18x, with the conclusion unchanged under more conservative multipliers. We identify four governance failure modes: confident rediscovery, ghost evidence accumulation, replication laundering, and confidence miscalibration. We then propose three interventions: null-result databases as training infrastructure, retraction-aware evaluation metrics, and mandatory training corpus disclosure. The central takeaway is that AI scientists will not only accelerate science. Without governance, they will accelerate science's blind spots before they accelerate its discoveries.
翻译:AI科学家系统正开始自动化地生成、评估和迭代科学假设。其优势在于速度;其风险在于速度同样会放大科学文献中嵌入的错误。我们认为近期面临的风险是语料库失效:AI科学家系统所依赖的训练和基础文献过度呈现阳性结果,而阴性发现则被严重低估。我们将这种扭曲形式化为零结果差距,并在三个领域(药物发现约0.60、心理学约0.56、癌症生物学约0.35)进行了估算,同时引入一个放大指数来推演检索、生成和自动评估如何加剧原始差距。通过一阶估计,我们证明标准的三阶段流水线可将语料库扭曲放大2.18倍,即便采用更保守的乘数因子,结论依然成立。我们识别出四种治理失效模式:自信重现、幽灵证据积累、复制洗白以及置信度校准偏差。随后提出三项干预措施:将零结果数据库作为训练基础设施、采用感知撤回的评估指标,以及实施强制训练语料库披露。核心结论是:AI科学家不仅会加速科学进程。若缺乏治理,它们在加速科学发现之前,将先加速放大科学中的盲区。