We study the problem of learning mixtures of Gaussians with censored data. Statistical learning with censored data is a classical problem, with numerous practical applications, however, finite-sample guarantees for even simple latent variable models such as Gaussian mixtures are missing. Formally, we are given censored data from a mixture of univariate Gaussians $$ \sum_{i=1}^k w_i \mathcal{N}(\mu_i,\sigma^2), $$ i.e. the sample is observed only if it lies inside a set $S$. The goal is to learn the weights $w_i$ and the means $\mu_i$. We propose an algorithm that takes only $\frac{1}{\varepsilon^{O(k)}}$ samples to estimate the weights $w_i$ and the means $\mu_i$ within $\varepsilon$ error.
翻译:我们研究了基于删失数据学习高斯混合模型的问题。删失数据下的统计学习是一个经典问题,具有众多实际应用,然而,即使是高斯混合这类简单潜变量模型,目前仍缺乏有限样本下的理论保证。具体而言,给定来自单变量高斯混合模型的删失数据 $$ \sum_{i=1}^k w_i \mathcal{N}(\mu_i,\sigma^2), $$ 即样本仅当其位于集合 $S$ 内部时才能被观测到。目标是学习权重 $w_i$ 和均值 $\mu_i$。我们提出了一种算法,仅需 $\frac{1}{\varepsilon^{O(k)}}$ 个样本即可在 $\varepsilon$ 误差范围内估计出权重 $w_i$ 和均值 $\mu_i$。