The availability of large annotated data can be a critical bottleneck in training machine learning algorithms successfully, especially when applied to diverse domains. Weak supervision offers a promising alternative by accelerating the creation of labeled training data using domain-specific rules. However, it requires users to write a diverse set of high-quality rules to assign labels to the unlabeled data. Automatic Rule Induction (ARI) approaches circumvent this problem by automatically creating rules from features on a small labeled set and filtering a final set of rules from them. In the ARI approach, the crucial step is to filter out a set of a high-quality useful subset of rules from the large set of automatically created rules. In this paper, we propose an algorithm (Filtering of Automatically Induced Rules) to filter rules from a large number of automatically induced rules using submodular objective functions that account for the collective precision, coverage, and conflicts of the rule set. We experiment with three ARI approaches and five text classification datasets to validate the superior performance of our algorithm with respect to several semi-supervised label aggregation approaches. Further, we show that achieves statistically significant results in comparison to existing rule-filtering approaches.
翻译:大规模标注数据的可用性可能是成功训练机器学习算法的关键瓶颈,尤其在应用于多样化领域时。弱监督学习通过使用特定领域规则加速标注训练数据的创建,提供了一种有前景的替代方案。然而,该方法要求用户编写一组高质量且多样化的规则来为未标注数据分配标签。自动规则归纳方法通过仅基于少量标注数据的特征自动生成规则,并从中筛选出最终规则集,从而规避了这一问题。在自动规则归纳方法中,关键步骤是从大量自动生成的规则中筛选出高质量且有用的子集。本文提出了一种算法(自动归纳规则的过滤方法),该算法利用子模目标函数(综合考虑规则集的集体精确率、覆盖率和冲突情况)从大量自动归纳的规则中过滤规则。我们通过三种自动规则归纳方法和五个文本分类数据集进行实验,验证了该算法相对于若干半监督标签聚合方法的优越性能。此外,我们证明该算法相比现有规则过滤方法取得了统计显著的结果。