Vision Language Models (VLMs) such as CLIP are powerful models; however they can exhibit unwanted biases, making them less safe when deployed directly in applications such as text-to-image, text-to-video retrievals, reverse search, or classification tasks. In this work, we propose a novel framework to generate synthetic counterfactual images to create a diverse and balanced dataset that can be used to fine-tune CLIP. Given a set of diverse synthetic base images from text-to-image models, we leverage off-the-shelf segmentation and inpainting models to place humans with diverse visual appearances in context. We show that CLIP trained on such datasets learns to disentangle the human appearance from the context of an image, i.e., what makes a doctor is not correlated to the person's visual appearance, like skin color or body type, but to the context, such as background, the attire they are wearing, or the objects they are holding. We demonstrate that our fine-tuned CLIP model, $CF_\alpha$, improves key fairness metrics such as MaxSkew, MinSkew, and NDKL by 40-66\% for image retrieval tasks, while still achieving similar levels of performance in downstream tasks. We show that, by design, our model retains maximal compatibility with the original CLIP models, and can be easily controlled to support different accuracy versus fairness trade-offs in a plug-n-play fashion.
翻译:视觉语言模型(如CLIP)是强大的模型;然而,它们可能表现出不良偏见,在直接部署于文本到图像、文本到视频检索、反向搜索或分类任务等应用时安全性较低。在本研究中,我们提出了一种新颖的框架,用于生成合成反事实图像,以创建多样化和平衡的数据集,可用于微调CLIP。给定一组来自文本到图像模型的多样化合成基础图像,我们利用现成的分割和修复模型,将具有多样化视觉外观的人物置于特定情境中。我们证明,在此类数据集上训练的CLIP能够学会将人物外观与图像情境解耦,即决定“医生”身份的因素并非与个人的视觉外观(如肤色或体型)相关,而是与情境(如背景、所穿着的服装或手持的物品)相关。我们展示了经过微调的CLIP模型$CF_\alpha$在图像检索任务中,将关键公平性指标(如MaxSkew、MinSkew和NDKL)提升了40-66%,同时在下游任务中仍能保持相近的性能水平。我们表明,通过设计,我们的模型保持了与原始CLIP模型的最大兼容性,并可通过即插即用的方式轻松调控,以支持不同的准确性与公平性权衡。