Selecting from or ranking a set of candidates variables in terms of their capacity for predicting an outcome of interest is an important task in many scientific fields. A variety of methods for variable selection and ranking have been proposed in the literature. In practice, it can be challenging to know which method is most appropriate for a given dataset. In this article, we propose methods of comparing variable selection and ranking algorithms. We first introduce measures of the quality of variable selection and ranking algorithms. We then define estimators of our proposed measures, and establish asymptotic results for our estimators in the regime where the dimension of the covariates is fixed as the sample size grows. We use our results to conduct large-sample inference for our measures, and we propose a computationally efficient partial bootstrap procedure to potentially improve finite-sample inference. We assess the properties of our proposed methods using numerical studies, and we illustrate our methods with an analysis of data for predicting wine quality from its physicochemical properties.
翻译:从一组候选变量中根据其预测感兴趣结果的能力进行选择或排序,是许多科学领域的重要任务。文献中已提出多种变量选择与排序方法。在实践中,如何针对给定数据集选择最合适的方法往往具有挑战性。本文提出了比较变量选择与排序算法的方法。我们首先引入衡量变量选择与排序算法质量的指标,随后定义所提出指标的估计量,并在协变量维度固定而样本量增长的渐近框架下建立估计量的渐近性质。基于这些结果,我们对指标进行大样本推断,并提出一种计算高效的偏自助法以潜在改进有限样本推断效果。通过数值研究评估所提方法的性质,并基于葡萄酒品质与其理化性质关系的预测数据分析对方法进行实例验证。