The cost of manual data labeling can be a significant obstacle in supervised learning. Data programming (DP) offers a weakly supervised solution for training dataset creation, wherein the outputs of user-defined programmatic labeling functions (LFs) are reconciled through unsupervised learning. However, DP can fail to outperform an unweighted majority vote in some scenarios, including low-data contexts. This work introduces a Bayesian extension of classical DP that mitigates failures of unsupervised learning by augmenting the DP objective with regularization terms. Regularized learning is achieved through maximum a posteriori estimation with informative priors. Majority vote is proposed as a proxy signal for automated prior parameter selection. Results suggest that regularized DP improves performance relative to maximum likelihood and majority voting, confers greater interpretability, and bolsters performance in low-data regimes.
翻译:手动数据标注的成本可能成为监督学习中的重大障碍。数据编程(DP)提供了一种弱监督解决方案,用于创建训练数据集,其中用户定义的编程标注函数(LF)的输出通过无监督学习进行协调。然而,在某些场景下(包括低数据量环境),DP可能无法超越未加权多数投票的性能。本研究提出了一种经典DP的贝叶斯扩展方法,通过向DP目标函数引入正则化项来缓解无监督学习的失效问题。正则化学习通过基于信息先验的最大后验估计实现。我们提出将多数投票作为自动化先验参数选择的代理信号。结果表明,与最大似然估计和多数投票相比,正则化DP能提升性能、增强可解释性,并在低数据量场景下强化模型表现。