Weak supervision (WS) is a powerful method to build labeled datasets for training supervised models in the face of little-to-no labeled data. It replaces hand-labeling data with aggregating multiple noisy-but-cheap label estimates expressed by labeling functions (LFs). While it has been used successfully in many domains, weak supervision's application scope is limited by the difficulty of constructing labeling functions for domains with complex or high-dimensional features. To address this, a handful of methods have proposed automating the LF design process using a small set of ground truth labels. In this work, we introduce AutoWS-Bench-101: a framework for evaluating automated WS (AutoWS) techniques in challenging WS settings -- a set of diverse application domains on which it has been previously difficult or impossible to apply traditional WS techniques. While AutoWS is a promising direction toward expanding the application-scope of WS, the emergence of powerful methods such as zero-shot foundation models reveals the need to understand how AutoWS techniques compare or cooperate with modern zero-shot or few-shot learners. This informs the central question of AutoWS-Bench-101: given an initial set of 100 labels for each task, we ask whether a practitioner should use an AutoWS method to generate additional labels or use some simpler baseline, such as zero-shot predictions from a foundation model or supervised learning. We observe that in many settings, it is necessary for AutoWS methods to incorporate signal from foundation models if they are to outperform simple few-shot baselines, and AutoWS-Bench-101 promotes future research in this direction. We conclude with a thorough ablation study of AutoWS methods.
翻译:弱监督(WS)是一种在缺乏标注数据时构建带标签数据集以训练监督模型的有效方法。它通过聚合多个由标注函数(LF)表示的含噪但低成本的标签估计值,替代了人工标注数据。尽管弱监督已在众多领域成功应用,但其应用范围受限于在特征复杂或高维领域构建标注函数的困难。为此,少量方法提出利用小规模真实标签集实现标注函数设计过程的自动化。本文提出AutoWS-Bench-101框架,用于评估自动弱监督(AutoWS)技术在具有挑战性的弱监督场景——即传统弱监督技术此前难以或无法应用的多样化应用领域——中的表现。尽管自动弱监督是拓展弱监督应用范围的前沿方向,但零样本基座模型等强大方法的涌现揭示出亟需理解自动弱监督技术如何与现代零样本或少样本学习器进行比较或协同。这构成了AutoWS-Bench-101的核心问题:给定每个任务初始的100个标签,实践者应使用自动弱监督方法生成额外标签,还是采用更简单的基线方法(如基座模型的零样本预测或监督学习)?我们观察到,在许多场景中,自动弱监督方法必须整合基座模型的信号才能超越简单的少样本基线,而AutoWS-Bench-101将推动该方向的未来研究。最后,我们对自动弱监督方法进行了全面的消融实验。