In some practical learning tasks, such as traffic video analysis, the number of available training samples is restricted by different factors, such as limited communication bandwidth and computation power. Determinantal Point Process (DPP) is a common method for selecting the most diverse samples to enhance learning quality. However, the number of selected samples is restricted to the rank of the kernel matrix implied by the dimensionality of data samples. Secondly, it is not easily customizable to different learning tasks. In this paper, we propose a new way of measuring task-oriented diversity based on the Rate-Distortion (RD) theory, appropriate for multi-level classification. To this end, we establish a fundamental relationship between DPP and RD theory. We observe that the upper bound of the diversity of data selected by DPP has a universal trend of $\textit{phase transition}$, which suggests that DPP is beneficial only at the beginning of sample accumulation. This led to the design of a bi-modal method, where RD-DPP is used in the first mode to select initial data samples, then classification inconsistency (as an uncertainty measure) is used to select the subsequent samples in the second mode. This phase transition solves the limitation to the rank of the similarity matrix. Applying our method to six different datasets and five benchmark models suggests that our method consistently outperforms random selection, DPP-based methods, and alternatives like uncertainty-based and coreset methods under all sampling budgets, while exhibiting high generalizability to different learning tasks.
翻译:在某些实际学习任务(如交通视频分析)中,可用训练样本数量受限于通信带宽和计算能力等因素。行列式点过程(DPP)是选择最具多样性样本以提升学习质量的常用方法,但其存在两大局限:首先,所选样本数量受限于数据样本维度所隐含的核矩阵秩;其次,该方法难以针对不同学习任务进行定制化调整。本文提出一种基于率失真(RD)理论的新型任务导向多样性度量方法,适用于多级分类场景。为此,我们建立了DPP与率失真理论的基本联系,发现DPP所选择数据多样性的上界存在普适的相变趋势,表明DPP仅在样本积累初期有效。基于此发现,我们设计了双模态方法:第一模态采用RD-DPP选择初始数据样本,第二模态利用分类不一致性(作为不确定性度量)选择后续样本。这种相变策略有效解决了相似矩阵秩的限制问题。在六个数据集和五个基准模型上的实验表明,本方法在所有采样预算下始终优于随机采样、基于DPP的方法以及其他备选方案(如基于不确定性和核心集方法),同时展现出对不同学习任务的强泛化能力。