Analyzing model performance in various unseen environments is a critical research problem in the machine learning community. To study this problem, it is important to construct a testbed with out-of-distribution test sets that have broad coverage of environmental discrepancies. However, existing testbeds typically either have a small number of domains or are synthesized by image corruptions, hindering algorithm design that demonstrates real-world effectiveness. In this paper, we introduce CIFAR-10-Warehouse, consisting of 180 datasets collected by prompting image search engines and diffusion models in various ways. Generally sized between 300 and 8,000 images, the datasets contain natural images, cartoons, certain colors, or objects that do not naturally appear. With CIFAR-10-W, we aim to enhance the evaluation and deepen the understanding of two generalization tasks: domain generalization and model accuracy prediction in various out-of-distribution environments. We conduct extensive benchmarking and comparison experiments and show that CIFAR-10-W offers new and interesting insights inherent to these tasks. We also discuss other fields that would benefit from CIFAR-10-W.
翻译:分析模型在各种未见环境中的性能是机器学习领域的一个关键研究问题。为研究该问题,构建一个包含分布外测试集且具有广泛环境差异覆盖的测试平台至关重要。然而,现有测试平台通常要么领域数量较少,要么通过图像损坏合成,这阻碍了能体现真实世界有效性的算法设计。本文引入CIFAR-10-Warehouse,它由通过不同方式提示图像搜索引擎和扩散模型收集的180个数据集组成。这些数据集大小通常在300到8000张图像之间,包含自然图像、卡通图像、特定颜色图像或非自然存在的物体。借助CIFAR-10-W,我们旨在增强两类泛化任务的评估并深化对其理解:即领域泛化以及在不同分布外环境中的模型准确率预测。我们进行了广泛的基准测试和对比实验,结果表明CIFAR-10-W为这些任务提供了新的有趣见解。我们还讨论了其他能够从CIFAR-10-W中受益的研究领域。