We consider the question of estimating multi-dimensional Gaussian mixtures (GM) with compactly supported or subgaussian mixing distributions. Minimax estimation rate for this class (under Hellinger, TV and KL divergences) is a long-standing open question, even in one dimension. In this paper we characterize this rate (for all constant dimensions) in terms of the metric entropy of the class. Such characterizations originate from seminal works of Le Cam (1973); Birge (1983); Haussler and Opper (1997); Yang and Barron (1999). However, for GMs a key ingredient missing from earlier work (and widely sought-after) is a comparison result showing that the KL and the squared Hellinger distance are within a constant multiple of each other uniformly over the class. Our main technical contribution is in showing this fact, from which we derive entropy characterization for estimation rate under Hellinger and KL. Interestingly, the sequential (online learning) estimation rate is characterized by the global entropy, while the single-step (batch) rate corresponds to local entropy, paralleling a similar result for the Gaussian sequence model recently discovered by Neykov (2022) and Mourtada (2023). Additionally, since Hellinger is a proper metric, our comparison shows that GMs under KL satisfy the triangle inequality within multiplicative constants, implying that proper and improper estimation rates coincide.
翻译:本文考虑估计具有紧支撑或次高斯混合分布的多维高斯混合模型的问题。此类模型在Hellinger、TV和KL散度下的极小极大估计率是一个长期悬而未决的问题,即便在单维情形下也是如此。本文利用该类模型的度量熵来刻画这个速率(对所有常数维度而言)。这种刻画方法源于Le Cam(1973)、Birge(1983)、Haussler和Opper(1997)、Yang和Barron(1999)的开创性工作。然而,对于高斯混合模型,先前研究中缺失的一个关键要素(也是广泛寻求的)是比较结果,表明KL散度与平方Hellinger距离在该类模型上均匀地相差一个常数倍数。我们的主要技术贡献在于证明了这一事实,并由此推导出在Hellinger和KL散度下估计率的熵特征。有趣的是,序贯(在线学习)估计率由全局熵刻画,而单步(批处理)估计率则对应于局部熵,这与Neykov(2022)和Mourtada(2023)最近为高斯序列模型发现的类似结果相呼应。此外,由于Hellinger是一种适当的度量,我们的比较表明,在KL散度下的高斯混合模型满足乘法常数意义下的三角不等式,这意味着适当估计率与不适当估计率是一致的。