Scientific observations generate large quantities of unlabeled data which is laborious to hand-label, making unsupervised learning techniques valuable for processing datasets. Among these approaches, contrastive learning provides a convenient mechanism for extracting structural representations from unannotated datasets. For natural imagery, the general approach is to use a variety of data-space augmentation methods in order to generate synthetic samples; however, for scientific observations data-space perturbations can fundamentally alter the underlying data. Our proposed method is to generate contrastive samples by perturbing the network weights rather than the underlying data, thus more closely preserving the structure of the data. We demonstrate this technique using a SimCLR-based pipeline applied over radar observations of meteors, and show performance gains under matched protocols.
翻译:科学观测产生大量未标注数据,而人工标注工作繁重,这使得无监督学习技术在处理数据集方面具有重要价值。在这些方法中,对比学习提供了一种从无标注数据集中提取结构表征的便捷机制。对于自然图像,通用做法是采用多种数据空间增强方法生成合成样本;然而针对科学观测数据,数据空间扰动会从根本上改变数据底层结构。我们提出的方法是通过扰动网络权重而非原始数据来生成对比样本,从而更紧密地保留数据结构。我们采用基于SimCLR的框架对雷达流星观测数据进行实验验证,在匹配协议下展现了性能提升。