Modern datasets in biology and chemistry are often characterized by the presence of a large number of variables and outlying samples due to measurement errors or rare biological and chemical profiles. To handle the characteristics of such datasets we introduce a method to learn a robust ensemble comprised of a small number of sparse, diverse and robust models, the first of its kind in the literature. The degree to which the models are sparse, diverse and resistant to data contamination is driven directly by the data based on a cross-validation criterion. We establish the finite-sample breakdown of the ensembles and the models that comprise them, and we develop a tailored computing algorithm to learn the ensembles by leveraging recent developments in l0 optimization. Our extensive numerical experiments on synthetic and artificially contaminated real datasets from genomics and cheminformatics demonstrate the competitive advantage of our method over state-of-the-art sparse and robust methods. We also demonstrate the applicability of our proposal on a cardiac allograft vasculopathy dataset.
翻译:现代生物学和化学数据集常因测量误差或罕见生物化学特征而呈现变量众多且含异常样本的特点。为应对此类数据的特性,我们提出一种学习鲁棒集成模型的方法,该集成由少量稀疏、多样且鲁棒的子模型构成,这在文献中尚属首创。模型稀疏性、多样性及抗数据污染的程度由基于交叉验证准则的数据直接驱动。我们建立了集成模型及其组成子模型的有限样本崩溃点理论,并基于l0优化领域最新进展开发了定制的计算算法来学习该集成。通过合成数据集与基因组学、化学信息学领域人工污染真实数据集的广泛数值实验,验证了本方法相较现有顶级稀疏与鲁棒方法的竞争优势。我们还通过心脏同种异体移植血管病变数据集展示了所提方法的实用价值。