Self-training (ST) is a simple and standard approach in semi-supervised learning that has been applied to many machine learning problems. Despite its widespread acceptance and practical effectiveness, it is still not well understood why and how ST improves performance by fitting the model to potentially erroneous pseudo-labels. To investigate the properties of ST, in this study, we derive and analyze a sharp characterization of the behavior of iterative ST when training a linear classifier by minimizing the ridge-regularized convex loss for binary Gaussian mixtures, in the asymptotic limit where input dimension and data size diverge proportionally. The derivation is based on the replica method of statistical mechanics. The result indicates that, when the total number of iterations is large, ST may find a classification plane with the optimal direction regardless of the label imbalance by accumulating small parameter updates over long iterations. It is argued that this is because the small update of ST can accumulate information of the data in an almost noiseless way. However, when a label imbalance is present in true labels, the performance of the ST is significantly lower than that of supervised learning with true labels, because the ratio between the norm of the weight and the magnitude of the bias can become significantly large. To overcome the problems in label imbalanced cases, several heuristics are introduced. By numerically analyzing the asymptotic formula, it is demonstrated that with the proposed heuristics, ST can find a classifier whose performance is nearly compatible with supervised learning using true labels even in the presence of significant label imbalance.
翻译:自训练是半监督学习中一种简单而标准的方法,已应用于许多机器学习问题。尽管其被广泛接受且在实践中有效,但对于自训练为何及如何通过将模型拟合至可能错误的伪标签来提升性能,目前仍缺乏深入理解。为探究自训练的性质,本研究推导并分析了迭代自训练在线性分类器训练中的精确特征描述:在高维极限下(即输入维度和数据规模同比发散时),通过最小化带脊正则化的凸损失函数对二元高斯混合进行分类。推导基于统计力学的副本方法。结果表明,当总迭代次数较大时,自训练可通过长迭代过程中积累微小参数更新,找到具有最优方向(无论标签是否平衡)的分类平面。论证指出,这是因为自训练的微小更新能以近似无噪声的方式积累数据信息。然而,当真实标签存在不平衡时,自训练的性能显著低于使用真实标签的监督学习,原因在于权重范数与偏置幅度的比值可能变得极大。为解决标签不平衡情况下的问题,本文引入了若干启发式策略。通过数值分析渐近公式证明,采用所提出的启发式策略后,即使存在显著标签不平衡,自训练也能找到性能与使用真实标签的监督学习几乎相当的分类器。