Given a source and a target probability measure supported on $\mathbb{R}^d$, the Monge problem asks to find the most efficient way to map one distribution to the other. This efficiency is quantified by defining a \textit{cost} function between source and target data. Such a cost is often set by default in the machine learning literature to the squared-Euclidean distance, $\ell^2_2(\mathbf{x},\mathbf{y})=\tfrac12|\mathbf{x}-\mathbf{y}|_2^2$. Recently, Cuturi et. al '23 highlighted the benefits of using elastic costs, defined through a regularizer $\tau$ as $c(\mathbf{x},\mathbf{y})=\ell^2_2(\mathbf{x},\mathbf{y})+\tau(\mathbf{x}-\mathbf{y})$. Such costs shape the \textit{displacements} of Monge maps $T$, i.e., the difference between a source point and its image $T(\mathbf{x})-\mathbf{x})$, by giving them a structure that matches that of the proximal operator of $\tau$. In this work, we make two important contributions to the study of elastic costs: (i) For any elastic cost, we propose a numerical method to compute Monge maps that are provably optimal. This provides a much-needed routine to create synthetic problems where the ground truth OT map is known, by analogy to the Brenier theorem, which states that the gradient of any convex potential is always a valid Monge map for the $\ell_2^2$ cost; (ii) We propose a loss to \textit{learn} the parameter $\theta$ of a parameterized regularizer $\tau_\theta$, and apply it in the case where $\tau_{A}(\mathbf{z})=|A^\perp \mathbf{z}|^2_2$. This regularizer promotes displacements that lie on a low dimensional subspace of $\mathbb{R}^d$, spanned by the $p$ rows of $A\in\mathbb{R}^{p\times d}$.
翻译:给定支撑在$\mathbb{R}^d$上的源概率测度和目标概率测度,Monge问题旨在寻找将一个分布映射到另一个分布的最有效方式。这种效率通过定义源数据与目标数据之间的\textit{成本}函数来量化。在机器学习文献中,此类成本通常默认设置为平方欧氏距离,即$\ell^2_2(\mathbf{x},\mathbf{y})=\tfrac12|\mathbf{x}-\mathbf{y}|_2^2$。最近,Cuturi等人'23的研究强调了使用弹性成本的优势,该成本通过正则化项$\tau$定义为$c(\mathbf{x},\mathbf{y})=\ell^2_2(\mathbf{x},\mathbf{y})+\tau(\mathbf{x}-\mathbf{y})$。此类成本通过赋予Monge映射$T$的\textit{位移}(即源点与其像$T(\mathbf{x})-\mathbf{x}$之间的差值)与$\tau$的近端算子结构相匹配的特性,从而塑造了这些位移的形态。在本工作中,我们对弹性成本的研究做出了两项重要贡献:(i) 针对任意弹性成本,我们提出了一种数值方法来计算可证明最优的Monge映射。这提供了一个亟需的流程,通过类比Brenier定理(该定理指出任意凸势能的梯度总是$\ell_2^2$成本的有效Monge映射),可以创建已知真实最优传输映射的合成问题;(ii) 我们提出了一种损失函数来\textit{学习}参数化正则化项$\tau_\theta$的参数$\theta$,并将其应用于$\tau_{A}(\mathbf{z})=|A^\perp \mathbf{z}|^2_2$的情况。该正则化项促使位移位于$\mathbb{R}^d$的一个低维子空间上,该子空间由$A\in\mathbb{R}^{p\times d}$的$p$个行向量张成。