We study cultural and socioeconomic diversity in contrastive vision-language models (VLMs). Using a broad range of benchmark datasets and evaluation metrics, we bring to attention several important findings. First, the common filtering of training data to English image-text pairs disadvantages communities of lower socioeconomic status and negatively impacts cultural understanding. Notably, this performance gap is not captured by - and even at odds with - the currently popular evaluation metrics derived from the Western-centric ImageNet and COCO datasets. Second, pretraining with global, unfiltered data before fine-tuning on English content can improve cultural understanding without sacrificing performance on said popular benchmarks. Third, we introduce the task of geo-localization as a novel evaluation metric to assess cultural diversity in VLMs. Our work underscores the value of using diverse data to create more inclusive multimodal systems and lays the groundwork for developing VLMs that better represent global perspectives.
翻译:本研究探讨对比视觉-语言模型中的文化与社会经济多样性。通过使用广泛的基准数据集与评估指标,我们揭示了若干重要发现。首先,将训练数据过滤为英文图文对的常见做法,对较低社会经济地位的群体造成不利影响,并损害文化理解能力。值得注意的是,这种性能差距未被当前基于西方中心主义的ImageNet和COCO数据集衍生的流行评估指标所捕捉——甚至与之相悖。其次,在针对英文内容进行微调前,使用未经筛选的全球数据进行预训练,可在不牺牲上述流行基准性能的前提下提升文化理解能力。第三,我们提出地理定位任务作为评估视觉-语言模型文化多样性的新型评估指标。本研究强调了使用多样化数据构建更具包容性的多模态系统的重要性,并为开发更能代表全球视角的视觉-语言模型奠定了基础。