In semi-supervised learning, student-teacher distribution matching has been successful in improving performance of models using unlabeled data in conjunction with few labeled samples. In this paper, we aim to replicate that success in the self-supervised setup where we do not have access to any labeled data during pre-training. We introduce our algorithm, Q-Match, and show it is possible to induce the student-teacher distributions without any knowledge of downstream classes by using a queue of embeddings of samples from the unlabeled dataset. We focus our study on tabular datasets and show that Q-Match outperforms previous self-supervised learning techniques when measuring downstream classification performance. Furthermore, we show that our method is sample efficient--in terms of both the labels required for downstream training and the amount of unlabeled data required for pre-training--and scales well to the sizes of both the labeled and unlabeled data.
翻译:在半监督学习中,师生分布匹配方法已成功利用少量标注样本结合未标注数据提升模型性能。本文旨在将这一成功经验迁移至自监督学习场景,即在预训练阶段无需任何标注数据。我们提出Q-Match算法,并证明通过利用未标注数据集中样本嵌入的队列,无需知晓下游类别即可构建师生分布。我们以表格数据集为重点研究对象,实验表明,在下游分类性能评估中,Q-Match优于以往的自监督学习技术。此外,该方法在样本效率上具有优势——无论是对下游训练所需的标注数据量,还是预训练所需的未标注数据量——且能较好地适应标注与未标注数据的规模变化。