For more than a decade, researchers have measured progress in object recognition on ImageNet-based generalization benchmarks such as ImageNet-A, -C, and -R. Recent advances in foundation models, trained on orders of magnitude more data, have begun to saturate these standard benchmarks, but remain brittle in practice. This suggests standard benchmarks, which tend to focus on predefined or synthetic changes, may not be sufficient for measuring real world generalization. Consequently, we propose studying generalization across geography as a more realistic measure of progress using two datasets of objects from households across the globe. We conduct an extensive empirical evaluation of progress across nearly 100 vision models up to most recent foundation models. We first identify a progress gap between standard benchmarks and real-world, geographical shifts: progress on ImageNet results in up to 2.5x more progress on standard generalization benchmarks than real-world distribution shifts. Second, we study model generalization across geographies by measuring the disparities in performance across regions, a more fine-grained measure of real world generalization. We observe all models have large geographic disparities, even foundation CLIP models, with differences of 7-20% in accuracy between regions. Counter to modern intuition, we discover progress on standard benchmarks fails to improve geographic disparities and often exacerbates them: geographic disparities between the least performant models and today's best models have more than tripled. Our results suggest scaling alone is insufficient for consistent robustness to real-world distribution shifts. Finally, we highlight in early experiments how simple last layer retraining on more representative, curated data can complement scaling as a promising direction of future work, reducing geographic disparity on both benchmarks by over two-thirds.
翻译:十余年来,研究者们通过基于ImageNet的泛化基准(如ImageNet-A、-C、-R)来衡量目标识别任务的进展。近年来,基于数量级更大数据训练的基座模型虽已开始饱和这些标准基准,但在实际应用中仍表现出脆弱性。这表明,侧重预定义或合成变化的标准基准可能不足以衡量现实世界的泛化能力。因此,我们提出以跨地理区域的泛化能力作为更现实的进展度量标准,采用两组来自全球各家庭物品的数据集。我们对近100个视觉模型(涵盖最新基座模型)开展了广泛的实证评估。首先,我们发现标准基准与现实世界地理迁移之间存在进展鸿沟:ImageNet上的进展在标准泛化基准上的提升幅度,是现实世界分布迁移下的2.5倍。其次,我们通过测量模型在不同地区的性能差异(更精细的现实世界泛化指标)来研究跨地理区域的泛化能力。观察显示所有模型均存在显著的地理差异,即使是基座CLIP模型,不同地区间的准确率差异也达7%-20%。与直观认识相悖的是,我们发现标准基准上的进展未能缩小地理差异,反而往往加剧了这种差异:性能最差模型与当前最佳模型间的地理差异已扩大三倍以上。我们的结果表明,仅靠规模扩展不足以实现针对现实世界分布迁移的持续鲁棒性。最后,我们通过初步实验强调,在更具代表性、经过筛选的数据上进行简单的最后一层重训练,可作为扩展策略的有力补充研究方向,此举将两个基准上的地理差异均降低超过三分之二。