Active learning aims to enhance model performance by strategically labeling informative data points. While extensively studied, its effectiveness on large-scale, real-world datasets remains underexplored. Existing research primarily focuses on single-source data, ignoring the multi-domain nature of real-world data. We introduce a multi-domain active learning benchmark to bridge this gap. Our benchmark demonstrates that traditional single-domain active learning strategies are often less effective than random selection in multi-domain scenarios. We also introduce CLIP-GeoYFCC, a novel large-scale image dataset built around geographical domains, in contrast to existing genre-based domain datasets. Analysis on our benchmark shows that all multi-domain strategies exhibit significant tradeoffs, with no strategy outperforming across all datasets or all metrics, emphasizing the need for future research.
翻译:主动学习旨在通过策略性地标注信息量大的数据点来提升模型性能。尽管已有广泛研究,但其在大规模真实世界数据集上的有效性仍待深入探索。现有研究主要关注单一来源数据,忽视了真实数据多领域的特性。为弥补这一空白,我们引入了多领域主动学习基准测试。该基准测试表明,在多领域场景下,传统单领域主动学习策略往往不如随机选择有效。我们还提出了CLIP-GeoYFCC——一个基于地理领域构建的新型大规模图像数据集,以区别于现有基于体裁的领域数据集。基准分析显示,所有多领域策略均存在显著的权衡问题,没有一种策略能在所有数据集或所有指标上表现卓越,这凸显了未来研究的必要性。