Most mathematical distortions used in ML are fundamentally integral in nature: $f$-divergences, Bregman divergences, (regularized) optimal transport distances, integral probability metrics, geodesic distances, etc. In this paper, we unveil a grounded theory and tools which can help improve these distortions to better cope with ML requirements. We start with a generalization of Riemann integration that also encapsulates functions that are not strictly additive but are, more generally, $t$-additive, as in nonextensive statistical mechanics. Notably, this recovers Volterra's product integral as a special case. We then generalize the Fundamental Theorem of calculus using an extension of the (Euclidean) derivative. This, along with a series of more specific Theorems, serves as a basis for results showing how one can specifically design, alter, or change fundamental properties of distortion measures in a simple way, with a special emphasis on geometric- and ML-related properties that are the metricity, hyperbolicity, and encoding. We show how to apply it to a problem that has recently gained traction in ML: hyperbolic embeddings with a "cheap" and accurate encoding along the hyperbolic vs Euclidean scale. We unveil a new application for which the Poincar\'e disk model has very appealing features, and our theory comes in handy: \textit{model} embeddings for boosted combinations of decision trees, trained using the log-loss (trees) and logistic loss (combinations).
翻译:机器学习中使用的绝大多数数学失真本质上都是积分形式的:$f$散度、Bregman散度、(正则化的)最优传输距离、积分概率度量、测地距离等。本文揭示了一套基础理论及工具,可帮助改进这些失真度量以更好地满足机器学习需求。我们首先推广了黎曼积分,使其能处理非严格可加性函数——即如同非广延统计力学中更一般的$t$-可加函数。值得注意的是,该方法将沃尔泰拉乘积积分作为特例加以恢复。随后,我们利用欧几里得导数的推广形式对微积分基本定理进行泛化。该定理与一系列更具体的定理共同构成基础,以证明如何通过简单方式专门设计、改变或修正失真度量的基本性质,重点聚焦几何与机器学习相关性质:度量性、双曲性和编码性。我们展示了如何将该理论应用于一个近期在机器学习领域备受关注的问题:通过沿双曲与欧几里得尺度进行"廉价"且精准的编码实现双曲嵌入。本研究揭示了一个新应用场景——庞加莱圆盘模型在此具有极具吸引力的特性,且我们的理论恰好适用:即针对基于对数损失(树模型)和逻辑损失(组合模型)训练的增强决策树组合的模型嵌入。