Incorporating human feedback has been shown to be crucial to align text generated by large language models to human preferences. We hypothesize that state-of-the-art instructional image editing models, where outputs are generated based on an input image and an editing instruction, could similarly benefit from human feedback, as their outputs may not adhere to the correct instructions and preferences of users. In this paper, we present a novel framework to harness human feedback for instructional visual editing (HIVE). Specifically, we collect human feedback on the edited images and learn a reward function to capture the underlying user preferences. We then introduce scalable diffusion model fine-tuning methods that can incorporate human preferences based on the estimated reward. Besides, to mitigate the bias brought by the limitation of data, we contribute a new 1M training dataset, a 3.6K reward dataset for rewards learning, and a 1K evaluation dataset to boost the performance of instructional image editing. We conduct extensive empirical experiments quantitatively and qualitatively, showing that HIVE is favored over previous state-of-the-art instructional image editing approaches by a large margin.
翻译:整合人类反馈已被证明对于将大语言模型生成的文本与人类偏好对齐至关重要。我们假设,当前最先进的指令式图像编辑模型(其输出基于输入图像和编辑指令生成)同样能从人类反馈中受益,因为其输出可能无法遵循正确的指令和用户偏好。在本文中,我们提出了一种利用人类反馈进行指令式视觉编辑(HIVE)的新框架。具体而言,我们收集编辑图像上的人类反馈,并学习一个奖励函数来捕捉潜在的底层用户偏好。随后,我们引入了可扩展的扩散模型微调方法,能够基于估计的奖励值将人类偏好融入其中。此外,为缓解数据局限性带来的偏差,我们贡献了一个新的100万训练数据集、一个用于奖励学习的3600奖励数据集,以及一个1000评估数据集,以提升指令式图像编辑的性能。我们进行了大量定性与定量实证实验,结果表明HIVE在性能上大幅优于此前最先进的指令式图像编辑方法。