Tackling the most pressing problems for humanity, such as the climate crisis and the threat of global pandemics, requires accelerating the pace of scientific discovery. While science has traditionally relied on trial and error and even serendipity to a large extent, the last few decades have seen a surge of data-driven scientific discoveries. However, in order to truly leverage large-scale data sets and high-throughput experimental setups, machine learning methods will need to be further improved and better integrated in the scientific discovery pipeline. A key challenge for current machine learning methods in this context is the efficient exploration of very large search spaces, which requires techniques for estimating reducible (epistemic) uncertainty and generating sets of diverse and informative experiments to perform. This motivated a new probabilistic machine learning framework called GFlowNets, which can be applied in the modeling, hypotheses generation and experimental design stages of the experimental science loop. GFlowNets learn to sample from a distribution given indirectly by a reward function corresponding to an unnormalized probability, which enables sampling diverse, high-reward candidates. GFlowNets can also be used to form efficient and amortized Bayesian posterior estimators for causal models conditioned on the already acquired experimental data. Having such posterior models can then provide estimators of epistemic uncertainty and information gain that can drive an experimental design policy. Altogether, here we will argue that GFlowNets can become a valuable tool for AI-driven scientific discovery, especially in scenarios of very large candidate spaces where we have access to cheap but inaccurate measurements or to expensive but accurate measurements. This is a common setting in the context of drug and material discovery, which we use as examples throughout the paper.
翻译:应对人类面临的最紧迫问题,如气候危机和全球大流行病的威胁,需要加速科学发现的进程。尽管传统科学在很大程度上依赖试错甚至偶然发现,但过去几十年中,数据驱动的科学发现激增。然而,为了真正利用大规模数据集和高通量实验装置,机器学习方法需要进一步改进,并更好地融入科学发现流程。在此背景下,当前机器学习方法面临的一个关键挑战是对极大搜索空间的高效探索,这需要估计可约简(认知)不确定性并生成多样且有信息量的实验方案的技术。这一需求催生了一种名为GFlowNets的新型概率机器学习框架,该框架可应用于实验科学循环中的建模、假设生成和实验设计阶段。GFlowNets学习从由未归一化概率对应的奖励函数间接定义的概率分布中进行采样,从而能够采样多样且高奖励的候选方案。GFlowNets还可用于构建因果模型的高效且摊销化的贝叶斯后验估计器,该模型以已获取的实验数据为条件。拥有此类后验模型后,即可提供认知不确定性和信息增益的估计量,进而驱动实验设计策略。综上所述,我们认为GFlowNets有望成为人工智能驱动科学发现的有力工具,尤其是在候选空间极大、且我们可访问廉价但不精确或昂贵但精确的测量数据的场景中。这是药物和材料发现中的常见设定,本文也以此类场景作为贯穿全文的示例。