Recently, many improved naive Bayes methods have been developed with enhanced discrimination capabilities. Among them, regularized naive Bayes (RNB) produces excellent performance by balancing the discrimination power and generalization capability. Data discretization is important in naive Bayes. By grouping similar values into one interval, the data distribution could be better estimated. However, existing methods including RNB often discretize the data into too few intervals, which may result in a significant information loss. To address this problem, we propose a semi-supervised adaptive discriminative discretization framework for naive Bayes, which could better estimate the data distribution by utilizing both labeled data and unlabeled data through pseudo-labeling techniques. The proposed method also significantly reduces the information loss during discretization by utilizing an adaptive discriminative discretization scheme, and hence greatly improves the discrimination power of classifiers. The proposed RNB+, i.e., regularized naive Bayes utilizing the proposed discretization framework, is systematically evaluated on a wide range of machine-learning datasets. It significantly and consistently outperforms state-of-the-art NB classifiers.


翻译:近年来,许多改进的朴素贝叶斯方法在判别能力方面取得了显著提升。其中,正则化朴素贝叶斯(RNB)通过平衡判别能力与泛化能力取得了优异的表现。数据离散化在朴素贝叶斯中至关重要。通过将相似值归并为一个区间,可以更准确地估计数据分布。然而,现有方法(包括RNB)通常将数据离散化为过少的区间,这可能导致显著的信息损失。为解决这一问题,我们提出了一种面向朴素贝叶斯的半监督自适应判别离散化框架,该框架通过伪标签技术同时利用标记数据与未标记数据,从而更准确地估计数据分布。所提方法还通过采用自适应判别离散化策略显著减少了离散化过程中的信息损失,进而大幅提升了分类器的判别能力。我们提出的RNB+(即采用所提离散化框架的正则化朴素贝叶斯)在广泛的机器学习数据集上进行了系统性评估。实验结果表明,RNB+显著且稳定地优于当前最先进的朴素贝叶斯分类器。

0
下载
关闭预览

相关内容

朴素贝叶斯法是基于贝叶斯定理与特征条件独立假设的分类方法。对于给定的训练数据集,首先基于“特征条件独立”的假设学习输入/输出的联合概率分布。然后基于此模型,对给定输入x,利用贝叶斯定理求后验概率最大的y。 朴素贝叶斯实现简单,学习与预测的效率都很高,是一种常用的方法。
专知会员服务
32+阅读 · 2021年7月2日
专知会员服务
30+阅读 · 2021年5月20日
【CVPR2021】现实世界域泛化的自适应方法
专知会员服务
58+阅读 · 2021年3月31日
专知会员服务
33+阅读 · 2021年3月7日
【ICLR2021】对未标记数据进行深度网络自训练的理论分析
Transferring Knowledge across Learning Processes
CreateAMind
29+阅读 · 2019年5月18日
强化学习的Unsupervised Meta-Learning
CreateAMind
18+阅读 · 2019年1月7日
无监督元学习表示学习
CreateAMind
27+阅读 · 2019年1月4日
Unsupervised Learning via Meta-Learning
CreateAMind
44+阅读 · 2019年1月3日
迁移学习之Domain Adaptation
全球人工智能
18+阅读 · 2018年4月11日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
2+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2008年12月31日
Arxiv
0+阅读 · 2023年5月24日
Arxiv
0+阅读 · 2023年5月21日
Arxiv
13+阅读 · 2021年3月29日
VIP会员
最新内容
《国防与国际安全中的量子-人工智能融合》报告
专知会员服务
1+阅读 · 今天14:44
致命七类无人机:无人机时代的演进型合成兵种
《异构无人水面艇集群作战自主制导算法》130页
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
2+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员