Is Stochastic Gradient Descent (SGD) substantially different from Glauber dynamics? This is a fundamental question at the time of understanding the most used training algorithm in the field of Machine Learning, but it received no answer until now. Here we show that in discrete optimization and inference problems, the dynamics of an SGD-like algorithm resemble very closely that of Metropolis Monte Carlo with a properly chosen temperature, which depends on the mini-batch size. This quantitative matching holds both at equilibrium and in the out-of-equilibrium regime, despite the two algorithms having fundamental differences (e.g.\ SGD does not satisfy detailed balance). Such equivalence allows us to use results about performances and limits of Monte Carlo algorithms to optimize the mini-batch size in the SGD-like algorithm and make it efficient at recovering the signal in hard inference problems.
翻译:随机梯度下降(SGD)与格劳伯动力学是否存在本质区别?这是理解机器学习领域最常用训练算法时的根本性问题,但至今未得到解答。本文证明,在离散优化与推理问题中,类SGD算法的动力学过程与采用适当温度(该温度取决于小批量样本规模)的梅特罗波利斯蒙特卡洛方法高度相似。尽管两种算法存在根本性差异(例如SGD不满足细致平衡条件),这种定量匹配在平衡态与非平衡态区间均成立。该等价性使我们能够利用蒙特卡洛算法性能与极限的相关研究成果,优化类SGD算法中的小批量样本规模,从而在困难推理问题中实现高效信号恢复。