Ordinal data are quite common in applied statistics. Although some model selection and regularization techniques for categorical predictors and ordinal response models have been developed over the past few years, less work has been done concerning ordinal-on-ordinal regression. Motivated by survey datasets on food products consisting of Likert-type items, we propose a strategy for smoothing and selection of ordinally scaled predictors in the cumulative logit model. First, the original group lasso is modified by use of difference penalties on neighbouring dummy coefficients, thus taking into account the predictors' ordinal structure. Second, a fused lasso type penalty is presented for fusion of predictor categories and factor selection. The performance of both approaches is evaluated in simulation studies, while our primary case study is a survey on the willingness to pay for luxury food products.
翻译:序数数据在应用统计学中十分常见。尽管近年来已发展出针对分类预测变量和序数响应变量的模型选择与正则化技术,但关于序数变量对序数变量的回归研究仍较为有限。受食品调查数据中李克特量表型项目的启发,我们提出一种在累积logit模型中对序数尺度预测变量进行平滑与选择的策略。首先,通过对相邻虚拟变量系数施加差分惩罚来修改原始组套索(group lasso),从而纳入预测变量的序数结构。其次,引入融合套索型惩罚以实现预测变量类别的融合与因子选择。模拟研究评估了两种方法的性能,而主要案例研究则聚焦于奢侈品食品支付意愿的调查。