This paper proposes an unsupervised method that leverages topological characteristics of data manifolds to estimate class separability of the data without requiring labels. Experiments conducted in this paper on several datasets demonstrate a clear correlation and consistency between the class separability estimated by the proposed method with supervised metrics like Fisher Discriminant Ratio~(FDR) and cross-validation of a classifier, which both require labels. This can enable implementing learning paradigms aimed at learning from both labeled and unlabeled data, like semi-supervised and transductive learning. This would be particularly useful when we have limited labeled data and a relatively large unlabeled dataset that can be used to enhance the learning process. The proposed method is implemented for language model fine-tuning with automated stopping criterion by monitoring class separability of the embedding-space manifold in an unsupervised setting. The proposed methodology has been first validated on synthetic data, where the results show a clear consistency between class separability estimated by the proposed method and class separability computed by FDR. The method has been also implemented on both public and internal data. The results show that the proposed method can effectively aid -- without the need for labels -- a decision on when to stop or continue the fine-tuning of a language model and which fine-tuning iteration is expected to achieve a maximum classification performance through quantification of the class separability of the embedding manifold.
翻译:本文提出了一种无监督方法,通过利用数据流形的拓扑特征来估计数据的类别可分性,而无需依赖标签。在多个数据集上的实验表明,该方法估计的类别可分性与需要标签的监督指标(如Fisher判别比FDR及分类器交叉验证)之间存在显著的相关性和一致性。这有助于实现从有标签和无标签数据中共同学习的范式(如半监督学习和直推式学习),尤其在仅有少量标签数据但拥有大量未标注数据可增强学习过程时具有重要价值。该方法被应用于语言模型微调,通过无监督环境下监测嵌入空间流形的类别可分性,实现自动停止准则的设定。首先在合成数据上验证了该方法,结果显示其估计的类别可分性与FDR计算结果高度一致。随后在公开和内部数据集上实施的实验表明,该方法无需标签即可有效辅助判断语言模型微调何时停止或继续,并通过量化嵌入流形的类别可分性,确定能实现最大分类性能的微调迭代次数。