This paper addresses stochastic optimization in a streaming setting with time-dependent and biased gradient estimates. We analyze several first-order methods, including Stochastic Gradient Descent (SGD), mini-batch SGD, and time-varying mini-batch SGD, along with their Polyak-Ruppert averages. Our non-asymptotic analysis establishes novel heuristics that link dependence, biases, and convexity levels, enabling accelerated convergence. Specifically, our findings demonstrate that (i) time-varying mini-batch SGD methods have the capability to break long- and short-range dependence structures, (ii) biased SGD methods can achieve comparable performance to their unbiased counterparts, and (iii) incorporating Polyak-Ruppert averaging can accelerate the convergence of the stochastic optimization algorithms. To validate our theoretical findings, we conduct a series of experiments using both simulated and real-life time-dependent data.
翻译:本文研究了在流式数据环境下,面对时变且有偏梯度估计的随机优化问题。我们分析了多种一阶方法,包括随机梯度下降(SGD)、小批量SGD以及时变小批量SGD,并探讨了它们的Polyak-Ruppert平均形式。我们的非渐近分析建立了新的启发式准则,将依赖性、偏差和凸性水平联系起来,从而加速收敛。具体而言,研究发现:(i)时变小批量SGD方法能够打破长程和短程依赖结构;(ii)有偏SGD方法可达到与无偏方法相当的性能;(iii)引入Polyak-Ruppert平均能够加速随机优化算法的收敛。为验证理论结果,我们使用模拟数据和真实时变数据开展了一系列实验。