The application of kernel-based Machine Learning (ML) techniques to discrete choice modelling using large datasets often faces challenges due to memory requirements and the considerable number of parameters involved in these models. This complexity hampers the efficient training of large-scale models. This paper addresses these problems of scalability by introducing the Nystr\"om approximation for Kernel Logistic Regression (KLR) on large datasets. The study begins by presenting a theoretical analysis in which: i) the set of KLR solutions is characterised, ii) an upper bound to the solution of KLR with Nystr\"om approximation is provided, and finally iii) a specialisation of the optimisation algorithms to Nystr\"om KLR is described. After this, the Nystr\"om KLR is computationally validated. Four landmark selection methods are tested, including basic uniform sampling, a k-means sampling strategy, and two non-uniform methods grounded in leverage scores. The performance of these strategies is evaluated using large-scale transport mode choice datasets and is compared with traditional methods such as Multinomial Logit (MNL) and contemporary ML techniques. The study also assesses the efficiency of various optimisation techniques for the proposed Nystr\"om KLR model. The performance of gradient descent, Momentum, Adam, and L-BFGS-B optimisation methods is examined on these datasets. Among these strategies, the k-means Nystr\"om KLR approach emerges as a successful solution for applying KLR to large datasets, particularly when combined with the L-BFGS-B and Adam optimisation methods. The results highlight the ability of this strategy to handle datasets exceeding 200,000 observations while maintaining robust performance.
翻译:基于核的机器学习技术应用于离散选择建模时,常因内存需求及模型参数数量庞大而面临挑战,这种复杂性阻碍了大规模模型的高效训练。本文通过引入Nyström近似,解决了大规模数据集上核逻辑回归的可扩展性问题。研究首先进行理论分析,包括:i) 刻画核逻辑回归解集的特征,ii) 给出Nyström近似核逻辑回归解的上界,iii) 描述针对Nyström核逻辑回归的优化算法特化。随后对Nyström核逻辑回归进行计算验证。测试了四种地标点选择方法:基础均匀采样、k-means采样策略、以及两种基于杠杆分数的非均匀方法。利用大规模出行方式选择数据集评估这些策略的性能,并与传统方法(如多项Logit模型)及现代机器学习技术进行对比。研究还评估了多种优化技术对所提出的Nyström核逻辑回归模型的效率,在数据集上检验了梯度下降、动量法、Adam和L-BFGS-B优化方法的性能。其中,k-means Nyström核逻辑回归方法结合L-BFGS-B和Adam优化算法,成为将核逻辑回归应用于大规模数据集的有效方案。结果表明该策略能处理超过20万观测值的数据集,同时保持稳健性能。