Scalarization is a general technique that can be deployed in any multiobjective setting to reduce multiple objectives into one, such as recently in RLHF for training reward models that align human preferences. Yet some have dismissed this classical approach because linear scalarizations are known to miss concave regions of the Pareto frontier. To that end, we aim to find simple non-linear scalarizations that can explore a diverse set of $k$ objectives on the Pareto frontier, as measured by the dominated hypervolume. We show that hypervolume scalarizations with uniformly random weights are surprisingly optimal for provably minimizing the hypervolume regret, achieving an optimal sublinear regret bound of $O(T^{-1/k})$, with matching lower bounds that preclude any algorithm from doing better asymptotically. As a theoretical case study, we consider the multiobjective stochastic linear bandits problem and demonstrate that by exploiting the sublinear regret bounds of the hypervolume scalarizations, we can derive a novel non-Euclidean analysis that produces improved hypervolume regret bounds of $\tilde{O}( d T^{-1/2} + T^{-1/k})$. We support our theory with strong empirical performance of using simple hypervolume scalarizations that consistently outperforms both the linear and Chebyshev scalarizations, as well as standard multiobjective algorithms in bayesian optimization, such as EHVI.
翻译:标量化是一种通用技术,可在任意多目标场景中应用,以将多个目标简化为单一目标,例如近期在基于人类反馈的强化学习(RLHF)中用于训练对齐人类偏好的奖励模型。然而,部分研究者曾摒弃这一经典方法,因为线性标量化已知会遗漏帕累托前沿的凹区域。为此,我们旨在寻找简单的非线性标量化方法,使其能够探索帕累托前沿上由$k$个目标组成的多样化集合(以支配超体积衡量)。我们证明,采用均匀随机权重的超体积标量化在理论上对于最小化超体积遗憾具有最优性,可达到$O(T^{-1/k})$的最优次线性遗憾界,并匹配了下界,表明任何算法在渐近意义上都无法取得更优性能。作为理论案例研究,我们考虑多目标随机线性臂问题,并展示通过利用超体积标量化的次线性遗憾界,可导出一类新颖的非欧几里得分析方法,从而得到改进的超体积遗憾界$\tilde{O}( d T^{-1/2} + T^{-1/k})$。我们通过使用简单超体积标量化的强实证性能支持理论——该标量化方法在贝叶斯优化中始终优于线性标量化、切比雪夫标量化以及标准多目标算法(如EHVI)。