Recently, large multimodal models, such as CLIP and Stable Diffusion have experimented tremendous successes in both foundations and applications. However, as these models increase in parameter size and computational requirements, it becomes more challenging for users to personalize them for specific tasks or preferences. In this work, we address the problem of adapting the previous models towards sets of particular human preferences, aligning the retrieved or generated images with the preferences of the user. We leverage the Bradley-Terry preference model to develop a fast adaptation method that efficiently fine-tunes the original model, with few examples and with minimal computing resources. Extensive evidence of the capabilities of this framework is provided through experiments in different domains related to multimodal text and image understanding, including preference prediction as a reward model, and generation tasks.
翻译:近年来,诸如CLIP和Stable Diffusion等大型多模态模型在基础研究与应用实践中均取得了显著成功。然而,随着这些模型参数规模与计算需求的增长,用户为特定任务或偏好对其进行个性化适配的难度持续增加。本研究针对现有模型向特定人类偏好集合适配的问题展开探讨,旨在使检索或生成的图像与用户偏好相契合。我们利用Bradley-Terry偏好模型开发了一种快速适应方法,该方法能够以少量样本和最低计算资源高效微调原始模型。通过跨多模态文本-图像理解相关领域(包括作为奖励模型的偏好预测任务与生成任务)的大量实验,充分验证了该框架的效能。