To overcome difficulties in classifying large dimensionality data with a large number of classes, we propose a novel approach called JLSPCADL. This paper uses the Johnson-Lindenstrauss (JL) Lemma to select the dimensionality of a transformed space in which a discriminative dictionary can be learned for signal classification. Rather than reducing dimensionality via random projections, as is often done with JL, we use a projection transformation matrix derived from Modified Supervised PC Analysis (M-SPCA) with the JL-prescribed dimension. JLSPCADL provides a heuristic to deduce suitable distortion levels and the corresponding Suitable Description Length (SDL) of dictionary atoms to derive an optimal feature space and thus the SDL of dictionary atoms for better classification. Unlike state-of-the-art dimensionality reduction-based dictionary learning methods, a projection transformation matrix derived in a single step from M-SPCA provides maximum feature-label consistency of the transformed space while preserving the cluster structure of the original data. Despite confusing pairs, the dictionary for the transformed space generates discriminative sparse coefficients, with fewer training samples. Experimentation demonstrates that JLSPCADL scales well with an increasing number of classes and dimensionality. Improved label consistency of features due to M-SPCA helps to classify better. Further, the complexity of training a discriminative dictionary is significantly reduced by using SDL. Experimentation on OCR and face recognition datasets shows relatively better classification performance than other supervised dictionary learning algorithms.
翻译:针对高维大数据且类别数量众多时的分类难题,我们提出一种名为JLSPCADL的新方法。本文运用Johnson-Lindenstrauss (JL)引理选择变换空间维度,在该空间中可学习用于信号分类的判别性字典。与JL方法通常采用的随机投影降维不同,我们采用基于修正监督主成分分析(M-SPCA)且符合JL规定维度的投影变换矩阵。JLSPCADL提供启发式准则,用于推导合适的失真水平及对应的字典原子适宜描述长度(SDL),从而获得最优特征空间及字典原子的SDL以实现更优分类。与当前基于降维的字典学习方法不同,通过M-SPCA单步推导的投影变换矩阵,能在保持原始数据聚类结构的同时最大化变换空间的特征-标签一致性。即便存在混淆类别对,变换空间中的字典也能用更少训练样本生成判别性稀疏系数。实验表明,JLSPCADL在类别数量和维度增加时具有良好的可扩展性。由于M-SPCA提升了特征标签一致性,分类性能得到改善。此外,通过使用SDL显著降低了判别性字典的训练复杂度。在OCR和人脸识别数据集上的实验表明,其分类性能优于其他监督字典学习算法。