Optimization problems find widespread use in both single-objective and multi-objective scenarios. In practical applications, users aspire for solutions that converge to the region of interest (ROI) along the Pareto front (PF). While the conventional approach involves approximating a fitness function or an objective function to reflect user preferences, this paper explores an alternative avenue. Specifically, we aim to discover a method that sidesteps the need for calculating the fitness function, relying solely on human feedback. Our proposed approach entails conducting direct preference learning facilitated by an active dueling bandit algorithm. The experimental phase is structured into three sessions. Firstly, we assess the performance of our active dueling bandit algorithm. Secondly, we implement our proposed method within the context of Multi-objective Evolutionary Algorithms (MOEAs). Finally, we deploy our method in a practical problem, specifically in protein structure prediction (PSP). This research presents a novel interactive preference-based MOEA framework that not only addresses the limitations of traditional techniques but also unveils new possibilities for optimization problems.
翻译:优化问题在单目标和多目标场景中有着广泛应用。在实际应用中,用户期望获得沿帕累托前沿收敛至感兴趣区域的解。传统方法通过近似适应度函数或目标函数来反映用户偏好,而本文探索了另一种途径。具体而言,我们旨在发现一种无需计算适应度函数、仅依赖人类反馈的方法。所提方法通过主动对决赌博机算法实现直接偏好学习。实验阶段分为三个部分:首先评估主动对决赌博机算法的性能,其次在多目标进化算法框架中实现该方法,最后将其应用于蛋白质结构预测这一实际问题。本研究提出了一种新颖的交互式偏好型多目标进化算法框架,不仅突破了传统技术的局限性,更为优化问题开辟了新的可能性。