This paper studies delayed stochastic algorithms for weakly convex optimization in a distributed network with workers connected to a master node. Recently, Xu et al. 2022 showed that an inertial stochastic subgradient method converges at a rate of $\mathcal{O}(\tau_{\text{max}}/\sqrt{K})$ which depends on the maximum information delay $\tau_{\text{max}}$. In this work, we show that the delayed stochastic subgradient method ($\texttt{DSGD}$) obtains a tighter convergence rate which depends on the expected delay $\bar{\tau}$. Furthermore, for an important class of composition weakly convex problems, we develop a new delayed stochastic prox-linear ($\texttt{DSPL}$) method in which the delays only affect the high-order term in the complexity rate and hence, are negligible after a certain number of $\texttt{DSPL}$ iterations. In addition, we demonstrate the robustness of our proposed algorithms against arbitrary delays. By incorporating a simple safeguarding step in both methods, we achieve convergence rates that depend solely on the number of workers, eliminating the effect of the delay. Our numerical experiments further confirm the empirical superiority of our proposed methods.
翻译:本文研究了在连接至主节点的工作节点分布式网络中,针对弱凸优化的延迟随机算法。近期,Xu等人(2022)表明,惯性随机次梯度方法以 $\mathcal{O}(\tau_{\text{max}}/\sqrt{K})$ 的速率收敛,该速率依赖于最大信息延迟 $\tau_{\text{max}}$。本文证明,延迟随机次梯度方法 ($\texttt{DSGD}$) 可获得更紧的收敛速率,该速率依赖于期望延迟 $\bar{\tau}$。此外,针对一类重要的组合弱凸问题,我们开发了一种新的延迟随机近端线性 ($\texttt{DSPL}$) 方法,其中延迟仅影响复杂度速率中的高阶项,因此在经过一定数量的 $\texttt{DSPL}$ 迭代后可忽略。同时,我们展示了所提出算法对任意延迟的鲁棒性。通过在两种方法中引入简单的保护步骤,我们实现了仅依赖于工作节点数量的收敛速率,从而消除了延迟的影响。数值实验进一步证实了所提出方法的经验优势。