Feature selection is popular for obtaining small, interpretable, yet highly accurate prediction models. Conventional feature-selection methods typically yield one feature set only, which might not suffice in some scenarios. For example, users might be interested in finding alternative feature sets with similar prediction quality, offering different explanations of the data. In this article, we introduce alternative feature selection and formalize it as an optimization problem. In particular, we define alternatives via constraints and enable users to control the number and dissimilarity of alternatives. We consider sequential as well as simultaneous search for alternatives. Next, we discuss how to integrate conventional feature-selection methods as objectives. In particular, we describe solver-based search methods to tackle the optimization problem. Further, we analyze the complexity of this optimization problem and prove NP-hardness. Additionally, we show that a constant-factor approximation exists under certain conditions and propose corresponding heuristic search methods. Finally, we evaluate alternative feature selection in comprehensive experiments with 30 binary-classification datasets. We observe that alternative feature sets may indeed have high prediction quality, and we analyze factors influencing this outcome.
翻译:特征选择是获得小型、可解释且预测精度高的预测模型的常用方法。传统特征选择方法通常仅产生一个特征集,这在某些场景下可能不足。例如,用户可能希望找到具有相似预测质量但能提供不同数据解释的替代特征集。本文提出替代特征选择,并将其形式化为一个优化问题。具体而言,我们通过约束定义替代方案,使用户能够控制替代方案的数量和相异性。我们考虑顺序搜索和同时搜索替代方案。接下来,我们讨论如何将传统特征选择方法作为目标函数进行集成,并描述基于求解器的搜索方法来处理该优化问题。此外,我们分析该优化问题的复杂度并证明其NP困难性。同时,我们证明在特定条件下存在常数因子近似,并提出相应的启发式搜索方法。最后,我们在包含30个二分类数据集的综合实验中评估替代特征选择。实验表明,替代特征集确实可能具有高预测质量,我们分析了影响这一结果的因素。