We propose Log-Averaged Mirror Prox (LAMP), a linear-space primal-dual method for large-scale optimal transport. LAMP implements primal mirror prox updates by tracking an averaged dual sequence, reducing storage complexity from ${O}(nm)$ to $O(n+m)$ while preserving dense, GPU-friendly reductions. Consequently, LAMP preserves the last-iterate $\widetilde{O}( nm\varepsilon^{-1})$ arithmetic complexity of conservatively parameterized primal-dual mirror prox. We further analyze LAMP as a direct optimal transport solver in a more performant parameter regime, providing a last-iterate sub-optimality certificate dependent on infeasibility and an explicit $O(1/t)$ term. Moreover, we give a computable sufficient condition for best-iterate convergence to a saddle-point. Numerical experiments with an optimized CUDA implementation show that LAMP outperforms first-order baselines in several high-accuracy (entropic) optimal transport problems. LAMP is further shown to scale up to problems with $n=m=2^{18}$ marginal supports, which were previously beyond the reach of primal-dual first-order methods.
翻译:我们提出对数平均镜像近端法(LAMP),一种用于大规模最优输运的线性空间原始-对偶方法。LAMP通过追踪平均对偶序列实现原始镜像近端更新,将存储复杂度从${O}(nm)$降低至$O(n+m)$,同时保留密集、GPU友好的归约操作。因此,LAMP保持了保守参数化原始-对偶镜像近端的最后迭代$\widetilde{O}( nm\varepsilon^{-1})$算术复杂度。我们进一步将LAMP分析为一种在更高性能参数区间内的直接最优输运求解器,提供了依赖于不可行性的最后迭代次优性认证以及显式的$O(1/t)$项。此外,我们给出了一个可计算的充分条件以确保最佳迭代收敛到鞍点。通过优化的CUDA实现进行的数值实验表明,LAMP在多个高精度(熵正则化)最优输运问题中优于一阶基线方法。进一步证明,LAMP可扩展至边际支撑点数$n=m=2^{18}$的问题,此类问题此前超出原始-对偶一阶方法的求解范围。