We study stochastic approximation algorithms with Markovian noise and constant step-size $\alpha$. We develop a method based on infinitesimal generator comparisons to study the bias of the algorithm, which is the expected difference between $\theta_n$ -- the value at iteration $n$ -- and $\theta^*$ -- the unique equilibrium of the corresponding ODE. We show that, under some smoothness conditions, this bias is of order $O(\alpha)$. Furthermore, we show that the time-averaged bias is equal to $\alpha V + O(\alpha^2)$, where $V$ is a constant characterized by a Lyapunov equation, showing that $\esp{\bar{\theta}_n} \approx \theta^*+V\alpha + O(\alpha^2)$, where $\bar{\theta}_n=(1/n)\sum_{k=1}^n\theta_k$ is the Polyak-Ruppert average. We also show that $\bar{\theta}_n$ converges with high probability around $\theta^*+\alpha V$. We illustrate how to combine this with Richardson-Romberg extrapolation to derive an iterative scheme with a bias of order $O(\alpha^2)$.
翻译:我们研究了具有马尔可夫噪声和定步长 $\alpha$ 的随机逼近算法。我们发展了一种基于无穷小生成元比较的方法来研究算法的偏差,即 $\theta_n$(第 $n$ 次迭代的值)与 $\theta^*$(对应常微分方程的唯一平衡点)的期望差值。我们证明,在某些光滑性条件下,该偏差的阶为 $O(\alpha)$。此外,我们证明了时间平均偏差等于 $\alpha V + O(\alpha^2)$,其中 $V$ 是一个由李雅普诺夫方程刻画的常数,这表明 $\esp{\bar{\theta}_n} \approx \theta^*+V\alpha + O(\alpha^2)$,其中 $\bar{\theta}_n=(1/n)\sum_{k=1}^n\theta_k$ 为 Polyak-Ruppert 平均。我们还证明了 $\bar{\theta}_n$ 以高概率收敛于 $\theta^*+\alpha V$ 附近。我们阐述了如何将此方法与 Richardson-Romberg 外推法结合,以推导出一个偏差阶为 $O(\alpha^2)$ 的迭代方案。