类别分布偏移下的文本分类研究综述 (A Survey of Text Classification Under Class Distribution Shift) - 专知论文

会员服务 ·

0

分布偏移 · 文本分类 · 类别 · 类别分布 · ML ·

A Survey of Text Classification Under Class Distribution Shift

翻译：类别分布偏移下的文本分类研究综述

Adriana Valentina Costache,Silviu Florin Gheorghe,Eduard Gabriel Poesina,Paul Irofti,Radu Tudor Ionescu

from arxiv, Accepted at EACL 2026 (main)

The basic underlying assumption of machine learning (ML) models is that the training and test data are sampled from the same distribution. However, in daily practice, this assumption is often broken, i.e.~the distribution of the test data changes over time, which hinders the application of conventional ML models. One domain where the distribution shift naturally occurs is text classification, since people always find new topics to discuss. To this end, we survey research articles studying open-set text classification and related tasks. We divide the methods in this area based on the constraints that define the kind of distribution shift and the corresponding problem formulation, i.e.~learning with the Universum, zero-shot learning, and open-set learning. We next discuss the predominant mitigation approaches for each problem setup. Finally, we identify several future work directions, aiming to push the boundaries beyond the state of the art. Interestingly, we find that continual learning can solve many of the issues caused by the shifting class distribution. We maintain a list of relevant papers at https://github.com/Eduard6421/Open-Set-Survey.

翻译：机器学习（ML）模型的基本前提假设是训练数据与测试数据来自同一分布。然而，在日常实践中，这一假设常被打破，即测试数据的分布随时间发生变化，这阻碍了传统ML模型的应用。文本分类是分布偏移自然发生的领域之一，因为人们总会发现新的讨论主题。为此，本文综述了研究开放集文本分类及相关任务的学术文献。我们根据定义分布偏移类型及相应问题表述的约束条件，将该领域方法分为三类：通用集学习、零样本学习与开放集学习。随后，我们讨论了每种问题设置下的主流缓解方法。最后，我们指出了若干未来研究方向，旨在推动该领域超越现有技术水平。值得注意的是，我们发现持续学习能够解决由类别分布偏移引发的诸多问题。相关论文列表维护于 https://github.com/Eduard6421/Open-Set-Survey。

0

相关内容

分布偏移

深度图学习在分布偏移下的综述：从图的分布外泛化到自适应

深度图学习在分布偏移下的综述：从图的分布外泛化到自适应

专知会员服务

18+阅读 · 2024年10月28日

文本分类算法及其应用场景研究

文本分类算法及其应用场景研究

专知会员服务

19+阅读 · 2024年7月31日

【牛津大学博士论文】学习分布不确定性估计的语义分割，191页pdf

【牛津大学博士论文】学习分布不确定性估计的语义分割，191页pdf

专知会员服务

30+阅读 · 2024年7月31日

文本分类算法及其应用场景研究综述

文本分类算法及其应用场景研究综述

专知会员服务

29+阅读 · 2024年6月18日

《分布外泛化评估》综述

《分布外泛化评估》综述

专知会员服务

43+阅读 · 2024年3月6日

基于图卷积神经网络的文本分类方法研究综述

基于图卷积神经网络的文本分类方法研究综述

专知会员服务

40+阅读 · 2022年8月26日

NTU最新《广义分布外OOD检测》综述论文，20页pdf阐述离群/异常/新类/开集/分布外检测的异同

NTU最新《广义分布外OOD检测》综述论文，20页pdf阐述离群/异常/新类/开集/分布外检测的异同

专知会员服务

29+阅读 · 2021年10月26日

分布外泛化(Out-Of-Distribution Generalization) 综述论文，22页pdf240篇文献

专知会员服务

64+阅读 · 2021年9月2日

多标签文本分类研究进展

专知会员服务

40+阅读 · 2021年5月18日

【Snapchat-谷歌-微软】最新《深度学习文本分类》2020综述论文大全，150+DL分类模型，42页pdf215篇参考文献

【Snapchat-谷歌-微软】最新《深度学习文本分类》2020综述论文大全，150+DL分类模型，42页pdf215篇参考文献

专知会员服务

84+阅读 · 2020年4月9日

《文本分类大综述：从浅层到深度学习》最新2020版35页pdf

《文本分类大综述：从浅层到深度学习》最新2020版35页pdf

专知

59+阅读 · 2020年8月6日

最新《迁移学习:域自适应理论》综述论文，128页ppt讲解迁移学习与最优传输

最新《迁移学习:域自适应理论》综述论文，128页ppt讲解迁移学习与最优传输

专知

16+阅读 · 2020年4月27日

五年12篇顶会论文综述！一文读懂深度学习文本分类方法

五年12篇顶会论文综述！一文读懂深度学习文本分类方法

AI100

10+阅读 · 2019年6月5日

NLP基础任务:文本分类近年发展汇总,68页超详细解析

NLP基础任务:文本分类近年发展汇总,68页超详细解析

专知

167+阅读 · 2019年4月18日

《小样本学习(Few-shot learning)》最新41页综述论文，来自港科大和第四范式

《小样本学习(Few-shot learning)》最新41页综述论文，来自港科大和第四范式

专知

363+阅读 · 2019年4月12日

里昂大学博士学位论文-图像分类中的迁移学习

里昂大学博士学位论文-图像分类中的迁移学习

专知

12+阅读 · 2019年4月10日

【GitHub项目推荐】文本分类最好的几个深度学习方法 TensorFlow 实践

【GitHub项目推荐】文本分类最好的几个深度学习方法 TensorFlow 实践

专知

39+阅读 · 2018年11月27日

深度学习文本分类方法综述（代码）

深度学习文本分类方法综述（代码）

中国人工智能学会

28+阅读 · 2018年6月16日

深度学习在文本分类中的应用

深度学习在文本分类中的应用

AI研习社

13+阅读 · 2018年1月7日

fastText、TextCNN、TextRNN…这套NLP文本分类深度学习方法库供你选择

fastText、TextCNN、TextRNN…这套NLP文本分类深度学习方法库供你选择

数据派THU

29+阅读 · 2017年8月2日

图文混合跨媒体知识单元的模糊分类方法研究

国家自然科学基金

1+阅读 · 2015年12月31日

有效融合多源异构数据的集成分类器研究

国家自然科学基金

5+阅读 · 2015年12月31日

分布式有监督学习的学习理论

国家自然科学基金

17+阅读 · 2015年12月31日

多标记文本数据流分类方法研究

国家自然科学基金

3+阅读 · 2015年12月31日

方差正则化的分类模型选择方法研究

国家自然科学基金

1+阅读 · 2015年12月31日

基于生态演替的文本大数据特征学习研究

国家自然科学基金

1+阅读 · 2015年12月31日

数据内在结构和稀疏保持的大间隔分类方法研究

国家自然科学基金

2+阅读 · 2015年12月31日

面向异分布数据的主动学习方法

国家自然科学基金

12+阅读 · 2015年12月31日

基于异构信息网络的分类算法推荐方法研究

国家自然科学基金

7+阅读 · 2015年12月31日

上市公司文本信息分析研究：基于大数据的视角

国家自然科学基金

8+阅读 · 2014年12月31日

Sensitivity analysis of image classification models using generalized polynomial chaos

Arxiv

0+阅读 · 2月3日

Simple-Sampling and Hard-Mixup with Prototypes to Rebalance Contrastive Learning for Text Classification

Arxiv

0+阅读 · 1月23日

Bridging the Gap Between Simulated and Real Network Data Using Transfer Learning

Arxiv

0+阅读 · 1月21日

Statistical Learning Theory for Distributional Classification

Arxiv

0+阅读 · 1月21日

A survey on Clustered Federated Learning: Taxonomy, Analysis and Applications

Arxiv

0+阅读 · 1月20日

Classification Imbalance as Transfer Learning

Arxiv

0+阅读 · 1月15日

Text Classification Under Class Distribution Shift: A Survey

Arxiv

0+阅读 · 1月15日

An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift

Arxiv

0+阅读 · 1月9日

A survey on Clustered Federated Learning: Taxonomy, Analysis and Applications

Arxiv

0+阅读 · 1月7日

Measures of classification bias derived from sample size analysis

Arxiv

0+阅读 · 1月6日

VIP会员

文章信息

相关主题

相关VIP内容

深度图学习在分布偏移下的综述：从图的分布外泛化到自适应

深度图学习在分布偏移下的综述：从图的分布外泛化到自适应

专知会员服务

18+阅读 · 2024年10月28日

文本分类算法及其应用场景研究

文本分类算法及其应用场景研究

专知会员服务

19+阅读 · 2024年7月31日

【牛津大学博士论文】学习分布不确定性估计的语义分割，191页pdf

【牛津大学博士论文】学习分布不确定性估计的语义分割，191页pdf

专知会员服务

30+阅读 · 2024年7月31日

文本分类算法及其应用场景研究综述

文本分类算法及其应用场景研究综述

专知会员服务

29+阅读 · 2024年6月18日

《分布外泛化评估》综述

《分布外泛化评估》综述

专知会员服务

43+阅读 · 2024年3月6日

基于图卷积神经网络的文本分类方法研究综述

基于图卷积神经网络的文本分类方法研究综述

专知会员服务

40+阅读 · 2022年8月26日

NTU最新《广义分布外OOD检测》综述论文，20页pdf阐述离群/异常/新类/开集/分布外检测的异同

NTU最新《广义分布外OOD检测》综述论文，20页pdf阐述离群/异常/新类/开集/分布外检测的异同

专知会员服务

29+阅读 · 2021年10月26日

分布外泛化(Out-Of-Distribution Generalization) 综述论文，22页pdf240篇文献

专知会员服务

64+阅读 · 2021年9月2日

多标签文本分类研究进展

专知会员服务

40+阅读 · 2021年5月18日

【Snapchat-谷歌-微软】最新《深度学习文本分类》2020综述论文大全，150+DL分类模型，42页pdf215篇参考文献

【Snapchat-谷歌-微软】最新《深度学习文本分类》2020综述论文大全，150+DL分类模型，42页pdf215篇参考文献

专知会员服务

84+阅读 · 2020年4月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【CMU博士论文】基于自适应表征的高效视觉建模

《多域作战中融合网络、电子战与动能机动》

AI智能体时代大模型安全风险与攻防新挑战

迈向个性化大语言模型驱动的智能体：基础、评估与未来方向

相关资讯

《文本分类大综述：从浅层到深度学习》最新2020版35页pdf

《文本分类大综述：从浅层到深度学习》最新2020版35页pdf

专知

59+阅读 · 2020年8月6日

最新《迁移学习:域自适应理论》综述论文，128页ppt讲解迁移学习与最优传输

最新《迁移学习:域自适应理论》综述论文，128页ppt讲解迁移学习与最优传输

专知

16+阅读 · 2020年4月27日

五年12篇顶会论文综述！一文读懂深度学习文本分类方法

五年12篇顶会论文综述！一文读懂深度学习文本分类方法

AI100

10+阅读 · 2019年6月5日

NLP基础任务:文本分类近年发展汇总,68页超详细解析

NLP基础任务:文本分类近年发展汇总,68页超详细解析

专知

167+阅读 · 2019年4月18日

《小样本学习(Few-shot learning)》最新41页综述论文，来自港科大和第四范式

《小样本学习(Few-shot learning)》最新41页综述论文，来自港科大和第四范式

专知

363+阅读 · 2019年4月12日

里昂大学博士学位论文-图像分类中的迁移学习

里昂大学博士学位论文-图像分类中的迁移学习

专知

12+阅读 · 2019年4月10日

【GitHub项目推荐】文本分类最好的几个深度学习方法 TensorFlow 实践

【GitHub项目推荐】文本分类最好的几个深度学习方法 TensorFlow 实践

专知

39+阅读 · 2018年11月27日

深度学习文本分类方法综述（代码）

深度学习文本分类方法综述（代码）

中国人工智能学会

28+阅读 · 2018年6月16日

深度学习在文本分类中的应用

深度学习在文本分类中的应用

AI研习社

13+阅读 · 2018年1月7日

fastText、TextCNN、TextRNN…这套NLP文本分类深度学习方法库供你选择

fastText、TextCNN、TextRNN…这套NLP文本分类深度学习方法库供你选择

数据派THU

29+阅读 · 2017年8月2日

相关论文

Sensitivity analysis of image classification models using generalized polynomial chaos

Arxiv

0+阅读 · 2月3日

Simple-Sampling and Hard-Mixup with Prototypes to Rebalance Contrastive Learning for Text Classification

Arxiv

0+阅读 · 1月23日

Bridging the Gap Between Simulated and Real Network Data Using Transfer Learning

Arxiv

0+阅读 · 1月21日

Statistical Learning Theory for Distributional Classification

Arxiv

0+阅读 · 1月21日

A survey on Clustered Federated Learning: Taxonomy, Analysis and Applications

Arxiv

0+阅读 · 1月20日

Classification Imbalance as Transfer Learning

Arxiv

0+阅读 · 1月15日

Text Classification Under Class Distribution Shift: A Survey

Arxiv

0+阅读 · 1月15日

An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift

Arxiv

0+阅读 · 1月9日

A survey on Clustered Federated Learning: Taxonomy, Analysis and Applications

Arxiv

0+阅读 · 1月7日

Measures of classification bias derived from sample size analysis

Arxiv

0+阅读 · 1月6日

相关基金

图文混合跨媒体知识单元的模糊分类方法研究

国家自然科学基金

1+阅读 · 2015年12月31日

有效融合多源异构数据的集成分类器研究

国家自然科学基金

5+阅读 · 2015年12月31日

分布式有监督学习的学习理论

国家自然科学基金

17+阅读 · 2015年12月31日

多标记文本数据流分类方法研究

国家自然科学基金

3+阅读 · 2015年12月31日

方差正则化的分类模型选择方法研究

国家自然科学基金

1+阅读 · 2015年12月31日

基于生态演替的文本大数据特征学习研究

国家自然科学基金

1+阅读 · 2015年12月31日

数据内在结构和稀疏保持的大间隔分类方法研究

国家自然科学基金

2+阅读 · 2015年12月31日

面向异分布数据的主动学习方法

国家自然科学基金

12+阅读 · 2015年12月31日

基于异构信息网络的分类算法推荐方法研究

国家自然科学基金

7+阅读 · 2015年12月31日

上市公司文本信息分析研究：基于大数据的视角

国家自然科学基金

8+阅读 · 2014年12月31日

微信扫码咨询专知VIP会员