Empirical Bayes provides a powerful framework for learning and adapting to latent structure in data. In sequence models where an independent observation is associated to each latent parameter, theory and methods around empirical Bayes are well-developed. However, in models where latent parameters and observed data interact through more complex designs, many statistical and algorithmic questions remain unanswered. In this work, we study a canonical setting of empirical Bayes estimation for the distribution of regression coefficients in a high-dimensional Bayesian or random effects linear model. Computationally, we propose a new system of gradient flow equations for computing a nonparametric maximum likelihood estimator (NPMLE), which jointly optimizes over the prior and posterior distributions of the regression coefficients in a Gibbs variational representation of the marginal log-likelihood. A diffusion-based implementation yields an adaptive Langevin dynamics algorithm in which the prior evolves continuously to optimize a sequence model log-likelihood defined by the coordinates of the Langevin sample. Theoretically, we show polynomial-time convergence of the proposed gradient flow to a near-NPMLE from any initialization within a convex sub-level set of the marginal log-likelihood, by developing a high-temperature log-Sobolev inequality for the posterior law. We establish the statistical consistency of any near-NPMLE under deterministic conditions for the regression design as $n,p\rightarrow\infty$.
翻译:暂无翻译