Data-driven neighborhood definitions and graph constructions are often used in machine learning and signal processing applications. k-nearest neighbor~(kNN) and $\epsilon$-neighborhood methods are among the most common methods used for neighborhood selection, due to their computational simplicity. However, the choice of parameters associated with these methods, such as k and $\epsilon$, is still ad hoc. We make two main contributions in this paper. First, we present an alternative view of neighborhood selection, where we show that neighborhood construction is equivalent to a sparse signal approximation problem. Second, we propose an algorithm, non-negative kernel regression~(NNK), for obtaining neighborhoods that lead to better sparse representation. NNK draws similarities to the orthogonal matching pursuit approach to signal representation and possesses desirable geometric and theoretical properties. Experiments demonstrate (i) the robustness of the NNK algorithm for neighborhood and graph construction, (ii) its ability to adapt the number of neighbors to the data properties, and (iii) its superior performance in local neighborhood and graph-based machine learning tasks.
翻译:数据驱动的邻域定义和图构建方法广泛应用于机器学习与信号处理领域。k近邻(kNN)和$\epsilon$-邻域方法因其计算简便性而成为最常用的邻域选择方法,但这些方法中参数(如k和$\epsilon$)的选取仍依赖于经验。本文主要有两项贡献:首先,我们提出了邻域选择的新视角,证明邻域构建等价于稀疏信号逼近问题;其次,我们提出了一种非负核回归(NNK)算法,该算法可构建具有更优稀疏表示能力的邻域。NNK算法与信号表示中的正交匹配追踪方法具有相似性,并具备理想的几何与理论性质。实验表明:(i)NNK算法在邻域与图构建中具有鲁棒性;(ii)其能根据数据特性自适应调整邻域数量;(iii)在基于局部邻域和图的机器学习任务中表现更优。