Flexible modeling of the entire distribution as a function of covariates is an important generalization of mean-based regression that has seen growing interest over the past decades in both the statistics and machine learning literature. This review outlines selected state-of-the-art statistical approaches to distributional regression, complemented with alternatives from machine learning. Topics covered include the similarities and differences between these approaches, extensions, properties and limitations, estimation procedures, and the availability of software. In view of the increasing complexity and availability of large-scale data, this review also discusses the scalability of traditional estimation methods, current trends, and open challenges. Illustrations are provided using data on childhood malnutrition in Nigeria and Australian electricity prices.
翻译:以协变量函数形式对整个分布进行灵活建模,是均值回归的重要推广。近几十年来,这一方法在统计学与机器学习领域均受到越来越多关注。本综述系统梳理了分布回归领域中若干前沿统计方法,并辅以机器学习领域的替代方案。涵盖内容包括:不同方法间的异同点、扩展形式、性质与局限性、估计过程以及软件实现情况。鉴于大规模数据的复杂性与可获得性日益提升,本文亦讨论了传统估计方法在可扩展性方面面临的挑战、当前研究趋势及未解难题。通过尼日利亚儿童营养不良数据与澳大利亚电价数据,本文提供了具体应用例证。