Stochastic differential equations of Langevin-diffusion form have received significant attention, thanks to their foundational role in both Bayesian sampling algorithms and optimization in machine learning. In the latter, they serve as a conceptual model of the stochastic gradient flow in training over-parameterized models. However, the literature typically assumes smoothness of the potential, whose gradient is the drift term. Nevertheless, there are many problems for which the potential function is not continuously differentiable, and hence the drift is not Lipschitz continuous everywhere. This is exemplified by robust losses and Rectified Linear Units in regression problems. In this paper, we show some foundational results regarding the flow and asymptotic properties of Langevin-type Stochastic Differential Inclusions under assumptions appropriate to the machine-learning settings. In particular, we show strong existence of the solution, as well as an asymptotic minimization of the canonical free-energy functional.
翻译:Langevin扩散形式的随机微分方程因其在贝叶斯采样算法和机器学习优化中的基础性作用而受到广泛关注。在优化领域,它们被用作训练过参数化模型时随机梯度流的概念模型。然而,现有文献通常假设势函数光滑,其梯度构成漂移项。但许多问题中的势函数并非连续可微,因此漂移项并非处处Lipschitz连续,例如回归问题中的鲁棒损失函数和修正线性单元。本文在适合机器学习场景的假设下,展示了Langevin型随机微分包含的流与渐近性质的一些基础性结果。特别地,我们证明了解的强存在性,以及典型自由能泛函的渐近最小化性质。