The emergence of a variety of Machine Learning (ML) approaches for travel mode choice prediction poses an interesting question to transport modellers: which models should be used for which applications? The answer to this question goes beyond simple predictive performance, and is instead a balance of many factors, including behavioural interpretability and explainability, computational complexity, and data efficiency. There is a growing body of research which attempts to compare the predictive performance of different ML classifiers with classical random utility models. However, existing studies typically analyse only the disaggregate predictive performance, ignoring other aspects affecting model choice. Furthermore, many studies are affected by technical limitations, such as the use of inappropriate validation schemes, incorrect sampling for hierarchical data, lack of external validation, and the exclusive use of discrete metrics. We address these limitations by conducting a systematic comparison of different modelling approaches, across multiple modelling problems, in terms of the key factors likely to affect model choice (out-of-sample predictive performance, accuracy of predicted market shares, extraction of behavioural indicators, and computational efficiency). We combine several real world datasets with synthetic datasets, where the data generation function is known. The results indicate that the models with the highest disaggregate predictive performance (namely extreme gradient boosting and random forests) provide poorer estimates of behavioural indicators and aggregate mode shares, and are more expensive to estimate, than other models, including deep neural networks and Multinomial Logit (MNL). It is further observed that the MNL model performs robustly in a variety of situations, though ML techniques can improve the estimates of behavioural indices such as Willingness to Pay.
翻译:各类机器学习方法在出行方式选择预测中的涌现,为交通建模者提出了一个有趣的问题:不同应用场景应选用何种模型?该问题的答案不仅关乎简单的预测性能,而是需要权衡多重因素,包括行为可解释性、计算复杂度和数据效率。目前已有越来越多研究尝试比较不同机器学习分类器与经典随机效用模型的预测性能。然而,现有研究通常仅分析个体层面的预测表现,忽略了影响模型选择的其他方面。此外,许多研究存在技术局限,例如采用不恰当的验证方案、对分层数据进行错误抽样、缺乏外部验证,以及仅使用离散评价指标。我们通过系统比较不同建模方法在多个建模问题中的关键影响因素(样本外预测性能、市场份额预测精度、行为指标提取能力及计算效率),解决了上述局限。研究结合了多个真实世界数据集与已知数据生成函数的合成数据集。结果表明,个体层面预测性能最高的模型(即极端梯度提升和随机森林)在行为指标和总体方式分担率估算方面表现较差,且其估算成本高于深度神经网络和多元Logit等其他模型。同时观察发现,多元Logit模型在各种情境下均保持稳健表现,但机器学习技术可提升支付意愿等行为指标的估算精度。