We present a general strategy for turning generative models into candidate solution samplers for batch Bayesian optimization (BO). The use of generative models for BO enables large batch scaling as generative sampling, optimization of non-continuous design spaces, and high-dimensional and combinatorial design. Inspired by the success of direct preference optimization (DPO), we show that one can train a generative model with noisy, simple utility values directly computed from observations to then form proposal distributions whose densities are proportional to the expected utility, i.e., BO's acquisition function values. Furthermore, this approach is generalizable beyond preference-based feedback to general types of reward signals and loss functions. This perspective avoids the construction of surrogate (regression or classification) models, common in previous methods that have used generative models for black-box optimization. Theoretically, we show that the generative models within the BO process follow a sequence of distributions which asymptotically approximate an optimal target under certain conditions. We also evaluate the performance through experiments on challenging optimization problems involving large batches in high dimensions.
翻译:我们提出了一种通用策略,将生成模型转化为批量贝叶斯优化(BO)中的候选解采样器。生成模型在BO中的应用能够实现大规模批量采样、非连续设计空间的优化,以及高维和组合设计。受直接偏好优化(DPO)成功经验的启发,我们展示了如何利用直接根据观测计算出的含噪声简单效用值来训练生成模型,从而形成其密度与期望效用(即BO的采集函数值)成正比的提议分布。此外,该方法可推广至偏好反馈之外的通用奖励信号和损失函数类型。这种视角避免了构建以往使用生成模型进行黑箱优化方法中常见的代理(回归或分类)模型。理论上,我们证明BO过程中的生成模型遵循一系列分布,这些分布在特定条件下渐近逼近最优目标。我们还通过涉及高维大规模批量的挑战性优化问题的实验评估了该方法的性能。