In many categorical response regression applications, the response categories admit a multiresolution structure. That is, subsets of the response categories may naturally be combined into coarser response categories. In such applications, practitioners are often interested in estimating the resolution at which a predictor affects the response category probabilities. In this article, we propose a method for fitting the multinomial logistic regression model in high dimensions that addresses this problem in a unified and data-driven way. In particular, our method allows practitioners to identify which predictors distinguish between coarse categories but not fine categories, which predictors distinguish between fine categories, and which predictors are irrelevant. For model fitting, we propose a scalable algorithm that can be applied when the coarse categories are defined by either overlapping or nonoverlapping sets of fine categories. Statistical properties of our method reveal that it can take advantage of this multiresolution structure in a way existing estimators cannot. We use our method to model cell type probabilities as a function of a cell's gene expression profile (i.e., cell type annotation). Our fitted model provides novel biological insights which may be useful for future automated and manual cell type annotation methodology.
翻译:在许多分类响应回归应用中,响应类别具有多分辨率结构。即,响应类别的子集可以自然地合并为更粗粒度的响应类别。在此类应用中,研究者通常关注预测变量影响响应类别概率的分辨率估计。本文提出一种高维多项式逻辑回归模型的拟合方法,以统一且数据驱动的方式解决该问题。具体而言,我们的方法允许研究者识别以下三类预测变量:区分粗粒度类别但无法区分细粒度类别的变量、区分细粒度类别的变量,以及无关变量。在模型拟合方面,我们提出了一种可扩展算法,适用于粗粒度类别由重叠或非重叠细粒度类别集合定义的情形。统计性质表明,该方法能够利用多分辨率结构,而现有估计量无法实现。我们将该方法应用于基于细胞基因表达谱建模细胞类型概率(即细胞类型注释)。拟合模型提供了新颖的生物学见解,可能对未来自动化和人工细胞类型注释方法具有参考价值。