Due to rapidly rising healthcare costs worldwide, there is significant interest in controlling them. An important aspect concerns price transparency, as preliminary efforts have demonstrated that patients will shop for lower costs, driving efficiency. This requires the data to be made available, and models that can predict healthcare costs for a wide range of patient demographics and conditions. We present an approach to this problem by developing a predictive model using machine-learning techniques. We analyzed de-identified patient data from New York State SPARCS (statewide planning and research cooperative system), consisting of 2.3 million records in 2016. We built models to predict costs from patient diagnoses and demographics. We investigated two model classes consisting of sparse regression and decision trees. We obtained the best performance by using a decision tree with depth 10. We obtained an R-square value of 0.76 which is better than the values reported in the literature for similar problems.
翻译:全球医疗成本迅速上升引发了对其控制的高度关注。一个重要方面涉及价格透明度,因为初步实践表明,患者会主动寻求更低成本的医疗服务,从而提升效率。这要求数据公开可用,并建立能够预测广泛患者人口统计学特征及疾病状况下医疗成本的模型。我们提出了一种基于机器学习技术构建预测模型的方法。我们分析了纽约州SPARCS(全州规划与合作研究系统)的去标识化患者数据,包含2016年230万条记录。我们构建了基于患者诊断信息和人口统计学特征预测成本的模型,并研究了稀疏回归与决策树两类模型框架。采用深度为10的决策树模型获得了最佳性能,其R平方值达到0.76,优于文献中同类问题的报告结果。