Generative Flow Networks (GFlowNets) are amortized samplers that learn stochastic policies to sequentially generate compositional objects from a given unnormalized reward distribution. They can generate diverse sets of high-reward objects, which is an important consideration in scientific discovery tasks. However, as they are typically trained from a given extrinsic reward function, it remains an important open challenge about how to leverage the power of pre-training and train GFlowNets in an unsupervised fashion for efficient adaptation to downstream tasks. Inspired by recent successes of unsupervised pre-training in various domains, we introduce a novel approach for reward-free pre-training of GFlowNets. By framing the training as a self-supervised problem, we propose an outcome-conditioned GFlowNet (OC-GFN) that learns to explore the candidate space. Specifically, OC-GFN learns to reach any targeted outcomes, akin to goal-conditioned policies in reinforcement learning. We show that the pre-trained OC-GFN model can allow for a direct extraction of a policy capable of sampling from any new reward functions in downstream tasks. Nonetheless, adapting OC-GFN on a downstream task-specific reward involves an intractable marginalization over possible outcomes. We propose a novel way to approximate this marginalization by learning an amortized predictor enabling efficient fine-tuning. Extensive experimental results validate the efficacy of our approach, demonstrating the effectiveness of pre-training the OC-GFN, and its ability to swiftly adapt to downstream tasks and discover modes more efficiently. This work may serve as a foundation for further exploration of pre-training strategies in the context of GFlowNets.
翻译:生成流网络(GFlowNets)是一种摊销采样器,通过学习随机策略从给定的非归一化奖励分布中顺序生成组合对象。它们能够生成多样化的高奖励对象集,这在科学发现任务中至关重要。然而,由于这些网络通常基于给定的外部奖励函数进行训练,如何利用预训练的强大能力以无监督方式训练GFlowNets并使其高效适应下游任务,仍是一个重要的开放挑战。受近年来无监督预训练在各领域取得成功的启发,我们提出了一种新颖的无奖励GFlowNet预训练方法。通过将训练问题构建为自监督任务,我们提出了结果条件生成流网络(OC-GFN),使其能够探索候选空间。具体而言,OC-GFN学习达到任意目标结果,类似于强化学习中的目标条件策略。研究表明,预训练的OC-GFN模型可直接提取策略,用于在下游任务中从任意新奖励函数中采样。然而,在适应下游任务特定奖励时,对OC-GFN的微调涉及对可能结果的难以处理边际化。我们提出了一种新颖的近似方法,通过学习摊销预测器实现高效微调。大量实验结果验证了我们方法的有效性,证明了OC-GFN预训练的有效性及其快速适应下游任务、更高效发现模式的能力。本工作可为GFlowNets中预训练策略的进一步探索奠定基础。