Pith. sign in

REVIEW 1 cited by

Computing the Bias of Constant-step Stochastic Approximation with Markovian Noise

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.14285 v2 pith:4TWACHJL submitted 2024-05-23 stat.ML cs.LGmath.OC

classification stat.MLcs.LGmath.OC
keywords alphathetabiasapproximationconstantmarkoviannoiseorder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study stochastic approximation algorithms with Markovian noise and constant step-size $\alpha$. We develop a method based on infinitesimal generator comparisons to study the bias of the algorithm, which is the expected difference between $\theta_n$ -- the value at iteration $n$ -- and $\theta^*$ -- the unique equilibrium of the corresponding ODE. We show that, under some smoothness conditions, this bias is of order $O(\alpha)$. Furthermore, we show that the time-averaged bias is equal to $\alpha V + O(\alpha^2)$, where $V$ is a constant characterized by a Lyapunov equation, showing that $\mathbb{E}[\bar{\theta}_n] \approx \theta^*+V\alpha + O(\alpha^2)$, where $\bar{\theta}_n=(1/n)\sum_{k=1}^n\theta_k$ is the Polyak-Ruppert average. We also show that $\bar{\theta}_n$ converges with high probability around $\theta^*+\alpha V$. We illustrate how to combine this with Richardson-Romberg extrapolation to derive an iterative scheme with a bias of order $O(\alpha^2)$.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Homogenization of Multi-agent Learning Dynamics in Finite-state Markov Games

    stat.ML 2025-06 conditional novelty 5.0 of 10

    Under uniform ergodicity and Lipschitz assumptions, the rescaled parameter process of multi-agent RL learners in a finite-state Markov game converges weakly to the ODE that averages each update against the stationary ...

Pith tools