Non-parametric machine learning models, such as random forests and gradient boosted trees, are frequently used to estimate house prices due to their predictive accuracy, but such methods are often limited in their ability to quantify prediction uncertainty. Conformal Prediction (CP) is a model-agnostic framework for constructing confidence sets around machine learning prediction models with minimal assumptions. However, due to the spatial dependencies observed in house prices, direct application of CP leads to confidence sets that are not calibrated everywhere, i.e., too large of confidence sets in certain geographical regions and too small in others. We survey various approaches to adjust the CP confidence set to account for this and demonstrate their performance on a data set from the housing market in Oslo, Norway. Our findings indicate that calibrating the confidence sets on a \textit{locally weighted} version of the non-conformity scores makes the coverage more consistently calibrated in different geographical regions. We also perform a simulation study on synthetically generated sale prices to empirically explore the performance of CP on housing market data under idealized conditions with known data-generating mechanisms.
翻译:非参数机器学习模型(如随机森林和梯度提升树)因其预测精度常被用于房价估算,但这类方法在量化预测不确定性方面存在局限性。共形预测(CP)是一种与模型无关的框架,可在最小假设条件下为机器学习预测模型构建置信区间。然而,由于房价观测值存在空间依赖性,直接应用CP会导致置信区间在各区域校准不一致,即某些地理区域置信区间过大,而另一些区域则过小。我们系统研究了多种调整CP置信区间以应对该问题的方法,并在挪威奥斯陆住房市场数据集上验证其性能。研究结果表明,基于非一致性分数的局部加权版本校准置信区间,可使不同地理区域的覆盖率保持更一致的校准水平。我们还在已知数据生成机制的理想条件下,利用合成销售价格数据进行仿真研究,以实证探讨CP在住房市场数据中的表现。