Preferential Bayesian optimization (PBO) is a sample-efficient framework for learning human preferences between candidate designs. PBO classically relies on homoscedastic noise models to represent human aleatoric uncertainty. Yet, such noise fails to accurately capture the varying levels of human aleatoric uncertainty, particularly when the user possesses partial knowledge among different pairs of candidates. For instance, a chemist with solid expertise in glucose-related molecules may easily compare two compounds from that family while struggling to compare alcohol-related molecules. Currently, PBO overlooks this uncertainty during the search for a new candidate through the maximization of the acquisition function, consequently underestimating the risk associated with human uncertainty. To address this issue, we propose a heteroscedastic noise model to capture human aleatoric uncertainty. This model adaptively assigns noise levels based on the distance of a specific input to a predefined set of reliable inputs known as anchors provided by the human. Anchors encapsulate partial knowledge and offer insight into the comparative difficulty of evaluating different candidate pairs. Such a model can be seamlessly integrated into the acquisition function, thus leading to candidate design pairs that elegantly trade informativeness and ease of comparison for the human expert. We perform an extensive empirical evaluation of the proposed approach, demonstrating a consistent improvement over homoscedastic PBO.
翻译:偏好贝叶斯优化(PBO)是一种用于学习候选设计之间人类偏好的样本高效框架。经典的PBO依赖于同方差噪声模型来表示人类的偶然不确定性。然而,此类噪声无法准确捕捉人类偶然不确定性的变化水平,特别是当用户在不同候选对之间具备部分知识时。例如,一位在葡萄糖相关分子领域具有扎实专业知识的化学家可能轻松比较该家族中的两种化合物,却在比较醇类相关分子时遇到困难。目前,PBO在通过最大化采集函数搜索新候选方案时忽略了这种不确定性,从而低估了与人类不确定性相关的风险。为解决这一问题,我们提出了一种异方差噪声模型来捕捉人类的偶然不确定性。该模型根据特定输入到人类提供的预定义可靠输入集(称为锚点)的距离自适应地分配噪声水平。锚点封装了部分知识,并为评估不同候选对的相对难度提供了洞见。此类模型可以无缝集成到采集函数中,从而为人类专家产生在信息量与易比较性之间取得优雅平衡的候选设计对。我们对所提方法进行了广泛的实证评估,结果表明其相较于同方差PBO取得了持续性的改进。