The problem of optimization on Stiefel manifold, i.e., minimizing functions of (not necessarily square) matrices that satisfy orthogonality constraints, has been extensively studied. Yet, a new approach is proposed based on, for the first time, an interplay between thoughtfully designed continuous and discrete dynamics. It leads to a gradient-based optimizer with intrinsically added momentum. This method exactly preserves the manifold structure but does not require additional operation to keep momentum in the changing (co)tangent space, and thus has low computational cost and pleasant accuracy. Its generalization to adaptive learning rates is also demonstrated. Notable performances are observed in practical tasks. For instance, we found that placing orthogonal constraints on attention heads of trained-from-scratch Vision Transformer [Dosovitskiy et al. 2022] could markedly improve its performance, when our optimizer is used, and it is better that each head is made orthogonal within itself but not necessarily to other heads. This optimizer also makes the useful notion of Projection Robust Wasserstein Distance [Paty & Cuturi 2019; Lin et al. 2020] for high-dim. optimal transport even more effective.
翻译:施蒂费尔流形上的优化问题,即最小化满足正交约束的(未必为方阵)矩阵函数,已被广泛研究。然而,本文首次提出了一种基于精心设计的连续与离散动力学相互作用的优化方法,由此导出了具有内禀动量机制的梯度优化器。该方法精确保持流形结构,但无需额外操作来维持动量在变化的(余)切空间中的适应性,因而具有较低的计算成本与良好的精度。其自适应学习率的推广形式亦被论证。在实际任务中观察到显著性能提升。例如,我们发现当使用该优化器时,对从头训练的视觉Transformer [Dosovitskiy et al. 2022] 的注意力头施加正交约束可显著提升其性能,且每个注意力头内部正交化优于头间正交化。该优化器还使投影鲁棒Wasserstein距离 [Paty & Cuturi 2019; Lin et al. 2020] ——这一用于高维最优输运的有效概念——更具效力。