Long-tailed data distributions are prevalent in a variety of domains, including finance, e-commerce, biomedical science, and cyber security. In such scenarios, the performance of machine learning models is often dominated by the head categories, while the learning of tail categories is significantly inadequate. Given abundant studies conducted to alleviate the issue, this work aims to provide a systematic view of long-tailed learning with regard to three pivotal angles: (A1) the characterization of data long-tailedness, (A2) the data complexity of various domains, and (A3) the heterogeneity of emerging tasks. To achieve this, we develop the most comprehensive (to the best of our knowledge) long-tailed learning benchmark named HeroLT, which integrates 13 state-of-the-art algorithms and 6 evaluation metrics on 14 real-world benchmark datasets across 4 tasks from 3 domains. HeroLT with novel angles and extensive experiments (264 in total) enables researchers and practitioners to effectively and fairly evaluate newly proposed methods compared with existing baselines on varying types of datasets. Finally, we conclude by highlighting the significant applications of long-tailed learning and identifying several promising future directions. For accessibility and reproducibility, we open-source our benchmark HeroLT and corresponding results at https://github.com/SSSKJ/HeroLT.
翻译:长尾数据分布在金融、电子商务、生物医学科学和网络安全等多个领域中普遍存在。在此类场景下,机器学习模型的性能往往由头部类别主导,而尾部类别的学习则显著不足。鉴于已有大量研究致力于缓解该问题,本文旨在从三个关键视角系统审视长尾学习:(A1)数据长尾性表征、(A2)不同领域的数据复杂性,以及(A3)新兴任务的异构性。为此,我们开发了目前(据我们所知)最全面的长尾学习基准测试平台HeroLT,该平台集成了13种最先进算法和6种评估指标,涵盖来自3个领域、4项任务中的14个真实世界基准数据集。HeroLT凭借新颖的视角和大量实验(共计264项),使研究人员和实践者能够针对不同类型的数据集,有效且公平地评估新提出的方法与现有基线方法。最后,我们总结了长尾学习的重大应用,并指出了若干有前景的未来方向。为确保可访问性和可复现性,我们在https://github.com/SSSKJ/HeroLT 开源了HeroLT基准测试平台及相应结果。