Metric learning aims at finding a suitable distance metric over the input space, to improve the performance of distance-based learning algorithms. In high-dimensional settings, metric learning can also play the role of dimensionality reduction, by imposing a low-rank restriction to the learnt metric. In this paper, instead of training a low-rank metric on high-dimensional data, we consider a randomly compressed version of the data, and train a full-rank metric there. We give theoretical guarantees on the error of distance-based metric learning, with respect to the random compression, which do not depend on the ambient dimension. Our bounds do not make any explicit assumptions, aside from i.i.d. data from a bounded support, and automatically tighten when benign geometrical structures are present. Experimental results on both synthetic and real data sets support our theoretical findings in high-dimensional settings.
翻译:度量学习旨在寻找输入空间上的合适距离度量,以提升基于距离的学习算法的性能。在高维场景中,通过对待学习度量施加低秩约束,度量学习还可起到降维作用。本文考虑对数据采用随机压缩版本,并在该压缩空间上训练全秩度量,而非直接在高维数据上训练低秩度量。我们给出了基于距离的度量学习在随机压缩下的误差理论保证,该保证不依赖于环境维度。除数据来自有界支撑的独立同分布假设外,我们的界无需额外显式假设,且在存在良性几何结构时会自动收紧。对合成数据集和真实数据集的实验结果均支持我们在高维场景中的理论发现。