Semi-supervised learning (SSL) has witnessed great progress with various improvements in the self-training framework with pseudo labeling. The main challenge is how to distinguish high-quality pseudo labels against the confirmation bias. However, existing pseudo-label selection strategies are limited to pre-defined schemes or complex hand-crafted policies specially designed for classification, failing to achieve high-quality labels, fast convergence, and task versatility simultaneously. To these ends, we propose a Semi-supervised Reward framework (SemiReward) that predicts reward scores to evaluate and filter out high-quality pseudo labels, which is pluggable to mainstream SSL methods in wide task types and scenarios. To mitigate confirmation bias, SemiReward is trained online in two stages with a generator model and subsampling strategy. With classification and regression tasks on 13 standard SSL benchmarks of three modalities, extensive experiments verify that SemiReward achieves significant performance gains and faster convergence speeds upon Pseudo Label, FlexMatch, and Free/SoftMatch.
翻译:半监督学习在基于伪标签的自训练框架下取得了显著进展,其核心挑战是如何在避免确认偏差的前提下筛选出高质量的伪标签。然而,现有伪标签选择策略局限于预定义方案或专门为分类任务设计的复杂手工规则,难以同时实现高质量标签、快速收敛及任务通用性。为此,我们提出半监督奖励框架(SemiReward),通过预测奖励分数评估并筛选高质量伪标签,该框架可即插即用于主流半监督方法,覆盖多种任务类型与场景。为缓解确认偏差,SemiReward采用两阶段在线训练策略,结合生成器模型与子采样方法。在涵盖三种模态的13个标准半监督基准数据集上的分类与回归实验表明,SemiReward在Pseudo Label、FlexMatch及Free/SoftMatch方法基础上,实现了显著性能提升与更快的收敛速度。