Ensemble methods are commonly used in classification due to their remarkable performance. Achieving high accuracy in a data stream environment is a challenging task considering disruptive changes in the data distribution, also known as concept drift. A greater diversity of ensemble components is known to enhance prediction accuracy in such settings. Despite the diversity of components within an ensemble, not all contribute as expected to its overall performance. This necessitates a method for selecting components that exhibit high performance and diversity. We present a novel ensemble construction and maintenance approach based on MMR (Maximal Marginal Relevance) that dynamically combines the diversity and prediction accuracy of components during the process of structuring an ensemble. The experimental results on both four real and 11 synthetic datasets demonstrate that the proposed approach (DynED) provides a higher average mean accuracy compared to the five state-of-the-art baselines.
翻译:集成方法因其卓越性能而广泛应用于分类任务。在数据流环境中,数据分布发生颠覆性变化(即概念漂移)会显著增加实现高精度的难度。研究表明,在数据流场景下,集成组件的高度多样性有助于提升预测准确率。然而,集成中各组件的多样性程度不同,并非所有组件都能如预期般对整体性能做出贡献,因此需要选择兼具高性能与多样性的组件。本文提出一种基于MMR(最大边际相关性)的新型集成构建与维护方法,该方法在集成构建过程中动态融合组件的多样性与预测准确性。在四个真实数据集与十一个合成数据集上的实验结果表明,与五种最新基线方法相比,所提方法(DynED)实现了更高的平均准确率。