The Empirical Bayes (EB) procedure of Hauer et al. (2002) is the workhorse of highway safety analysis: it combines a Safety Performance Function with observed crash counts to produce shrinkage estimates of segment-level crash rates. EB delivers practicality by holding several quantities fixed at calibration: SPF coefficients, per-type overdispersion, observed ADT, and a fixed exposure exponent. These assumptions strain when ADT is missing on a majority of segments. We present a fully Bayesian hierarchical model that moves beyond EB by relaxing each of these assumptions in a single joint inference. Fit on Ohio's road inventory (408,304 segments, 2.9 million crashes, 2013-2025), the model jointly imputes missing ADT and estimates per-segment crash rates with uncertainty. Posterior predictive checks of an initial fixed-exposure model expose a tail misfit; relaxing the exposure structure to a per-functional-class exposure exponent and an estimated length exponent, in place of a single scalar and a fixed offset, resolves it and improves out-of-sample predictive accuracy (PSIS-LOO $Δ\mathrm{elpd}$ = 9,394, SE 238). Crash count is sublinear in traffic in every class (exposure exponents 0.49-0.70, all $<1$, the safety-in-numbers effect) and sublinear in segment length ($β_{\mathrm{len}} = 0.69$). Partial pooling substantially improves out-of-sample predictive accuracy over complete pooling (PSIS-LOO $Δ\mathrm{elpd}$ = 4,780, SE 225). The Bayesian ADT submodel attains $R^2_{\log} = 0.756$ by encoding county and functional class as hierarchical priors, versus $0.653$ for a LightGBM restricted to the same continuous predictors. The output is a posterior crash rate distribution per segment, replacing the median-by-type point estimates used in our prior risk-aware routing framework.
翻译:Hauer等人(2002)提出的经验贝叶斯(EB)方法是公路安全分析的核心工作:它通过将安全性能函数与观测事故计数相结合,对路段级事故率产生收缩估计。EB方法通过将若干参数在标定阶段固定为常数来实现实用性:SPF系数、各类型过度离散参数、观测平均日交通量(ADT)以及固定暴露度指数。当大部分路段缺失ADT数据时,这些假设会产生偏差。我们提出一个完全贝叶斯分层模型,通过将上述每个假设纳入单一联合推断中,超越了EB方法。该模型在俄亥俄州道路数据库(408,304个路段,290万起事故,2013-2025年)上拟合,能够联合插补缺失ADT数据并估计具有不确定性的路段级事故率。初始固定暴露度模型的后验预测检验暴露出尾部拟合不良问题;将暴露度结构放宽为每个功能类别的暴露度指数和估计的长度指数(代替单一标量值和固定偏移)解决了这一问题,并提升了样本外预测精度(PSIS-LOO $Δ\mathrm{elpd}$ = 9,394,标准误 238)。每个类别的事故次数与交通量呈亚线性关系(暴露度指数0.49-0.70,全部$<1$,体现"数字中安全"效应),同时与路段长度呈亚线性关系($β_{\mathrm{len}} = 0.69$)。部分池化显著优于完全池化的样本外预测精度(PSIS-LOO $Δ\mathrm{elpd}$ = 4,780,标准误 225)。贝叶斯ADT子模型通过将县和功能类别编码为分层先验,达到$R^2_{\log} = 0.756$,而受限于相同连续预测变量的LightGBM仅为$0.653$。模型输出每个路段的后验事故率分布,取代了我们先前风险感知路径规划框架中使用的按类型中位数点估计。