Active learning reduces labeling costs by selecting samples that maximize information gain. A dominant framework, Query-by-Committee (QBC), typically relies on perturbation-based diversity by inducing model disagreement through random feature subsetting or data blinding. While this approximates one notion of epistemic uncertainty, it sacrifices direct characterization of the plausible hypothesis space. We propose the complementary approach: Rashomon Ensembled Active Learning (REAL) which constructs a committee by exhaustively enumerating the Rashomon Set of all near-optimal models. To address functional redundancy within this set, we adopt a PAC-Bayesian framework using a Gibbs posterior to weight committee members by their empirical risk. Leveraging recent algorithmic advances, we exactly enumerate this set for the class of sparse decision trees. Across synthetic and established active learning baselines, REAL outperforms randomized ensembles, particularly in moderately noisy environments where it strategically leverages expanded model multiplicity to achieve faster convergence.
翻译:主动学习通过选择最大化信息增益的样本来降低标注成本。主流框架“委员会查询”(QBC)通常依赖基于扰动的多样性,通过随机特征子集选择或数据遮蔽诱发模型分歧。虽然这近似了某种认知不确定性概念,但牺牲了对合理假设空间的直接刻画。我们提出互补方法:拉什蒙集成主动学习(REAL),该算法通过穷举枚举所有近似最优模型的拉什蒙集来构建委员会。针对该集合内的功能冗余问题,我们采用基于吉布斯后验的PAC-Bayesian框架,通过经验风险对委员会成员进行加权。借助最新的算法进展,我们针对稀疏决策树类别实现了该集合的精确枚举。在合成数据集和已建立的主动学习基准测试中,REAL均优于随机集成方法,尤其在中等噪声环境中,该方法通过策略性地利用扩展的模型多样性实现更快的收敛。