The methodology discussed in this paper aims to enhance choice models' comprehensiveness and explanatory power for forecasting choice outcomes. To achieve these, we have developed a data-driven method that leverages machine learning procedures for identifying the most effective representation of variables in mode choice empirical probability specifications. The methodology will show its significance, particularly in the face of big data and an abundance of variables where it can search through many candidate models. Furthermore, this study will have potential applications in transportation planning and policy-making, which will be achieved by introducing a sparse identification method that looks for the sparsest specification ( parsimonious model ) in the domain of candidate functions. Finally, this paper applies the method to synthetic choice data as a proof of concept. We perform two experiments and show that if the functional form used to generate the synthetic data lies in the domain of base functions, the methodology can recover that. Otherwise, the method will raise a red flag by outputting small coefficients ( near zero ) for base functions.
翻译:本文探讨的方法旨在提升选择模型在预测选择结果时的全面性与解释力。为实现该目标,我们开发了一种数据驱动方法,利用机器学习流程识别模式选择经验概率规范中变量的最有效表征。该方法在面临大数据及变量丰富场景时尤为重要,因其可在众多候选模型中高效搜索。此外,本研究通过引入一种稀疏识别方法,在候选函数域中寻找最稀疏规范(简约模型),从而在交通规划与政策制定领域具有潜在应用价值。最后,本文将该方法应用于合成选择数据作为概念验证。我们开展了两项实验,结果表明:若生成合成数据的函数形式属于基函数域,该方法可恢复该函数形式;反之,方法将通过输出接近零的小系数对基函数发出警示信号。