We propose a categorical semantics for machine learning algorithms in terms of lenses, parametric maps, and reverse derivative categories. This foundation provides a powerful explanatory and unifying framework: it encompasses a variety of gradient descent algorithms such as ADAM, AdaGrad, and Nesterov momentum, as well as a variety of loss functions such as MSE and Softmax cross-entropy, and different architectures, shedding new light on their similarities and differences. Furthermore, our approach to learning has examples generalising beyond the familiar continuous domains (modelled in categories of smooth maps) and can be realised in the discrete setting of Boolean and polynomial circuits. We demonstrate the practical significance of our framework with an implementation in Python.
翻译:摘要:我们提出了一种基于透镜、参数映射和反向导数范畴的机器学习算法范畴语义。该基础框架提供了强大的解释与统一机制:它涵盖了ADAM、AdaGrad和Nesterov动量等多种梯度下降算法,MSE和Softmax交叉熵等多种损失函数,以及不同架构,揭示了它们之间的相似性与差异。此外,我们的学习方法实例不仅超越了常见的连续域(在光滑映射范畴中建模),还可在布尔电路和多项式电路等离散场景中实现。我们通过Python实现展示了该框架的实际意义。