We study the Gibbs posterior distribution for sparse deep neural nets in a nonparametric regression setting. The posterior can be accessed via Metropolis-adjusted Langevin algorithms. Using a mixture over uniform priors on sparse sets of network weights, we prove an oracle inequality which shows that the method adapts to the unknown regularity and hierarchical structure of the regression function. The estimator achieves the minimax-optimal rate of convergence (up to a logarithmic factor).
翻译:我们研究非参数回归背景下稀疏深度神经网络的吉布斯后验分布。该后验分布可通过Metropolis调整Langevin算法获取。利用网络权重的稀疏集上均匀先验的混合分布,我们证明了一个oracle不等式,表明该方法能够自适应回归函数的未知正则性与层级结构。该估计量达到了极小化最优收敛速度(仅含对数因子)。