Scaling Limits of Constant-Stepsize SGD at Flat Minima


Abstract

For stochastic gradient descent (SGD) with a constant stepsize \(\alpha\), the invariant law of the iterates, centered at a minimizer, describes the behavior of the algorithm over long time horizons. In the strongly convex case, this invariant law has the familiar \(\sqrt{\alpha}\) scaling and a Gaussian limit as \(\alpha\downarrow 0\). We show that this behavior changes fundamentally for convex objectives \(H\) with flat minima and (sub)quadratic tails.

More specifically, we study SGD with Markovian noise generated by a contractive driving chain. For every sufficiently small constant stepsize \(\alpha\), we prove existence, uniqueness, and geometric convergence to an augmented invariant law in a Wasserstein distance induced by an \(\alpha\)-dependent metric. When the minimizer \(x_\star\) has local flatness exponent \(m\ge2\), meaning that \(\nabla^2H(x)\asymp \left\lVert x-x_\star \right\rVert^{m-2}I_d\) as \(x\to x_\star\), we obtain a contraction bound with factor \(1-c\alpha^{m-1}\), where \(c>0\) is a constant. This recovers the factor \(1-c\alpha\) in the quadratic case \(m=2\). We then analyze the small-stepsize scaling limit. We show that the invariant law concentrates on the scale \(\alpha^{1/m}\) and that the rescaled iterates converge weakly to the stationary distribution of the stochastic differential equation \[dY_t=-h_0(Y_t)\,dt+\Sigma^{1/2}\,dB_t ,\] where \(h_0\) is the limiting drift at the minimizer and \(\Sigma\) denotes the asymptotic covariance. This recovers the Gaussian limit when \(m=2\) and gives generally non-Gaussian stationary limits in the flat case \(m>2\). Finally, we give corresponding results for coordinate-separable objectives with unequal flatness exponents.

1

2

3

1 Introduction↩︎

Stochastic approximation originated in recursive methods for root-finding and optimization based on noisy observations [1][5]. Classical theory uses decreasing stepsizes and studies almost-sure convergence to a target point. In large-scale optimization, by contrast, constant or piecewise-constant stepsizes are often used for long time intervals. With a constant stepsize and nondegenerate gradient noise, the last stochastic iterate generally does not converge to the minimizer. It is instead natural to study the invariant law of the Markov chain generated by the algorithm, which describes the stationary error of the last iterate.

For smooth strongly convex objectives, the stationary error of stochastic gradient descent (SGD) is well understood. It is typically of order \(\sqrt\alpha\), and after the \(\alpha^{-1/2}\)-rescaling the invariant law converges to the stationary distribution of an Ornstein–Uhlenbeck diffusion, hence Gaussian. This picture appears in small-stepsize asymptotic laws [6], in diffusion approximations for constant-step SGD [7], and in stationary small-stepsize characterizations for SGD-type algorithms [8]. Complementary Markov chain analyses establish invariant measure expansions, convergence, and concentration in strongly convex settings [9], [10]. However, the quadratic case is not representative of all convex objectives. When the deterministic drift vanishes to order higher than one at the minimizer, the normalization of the invariant law and its scaling limit both change.

This paper studies convex objectives whose minimizer may be flatter than quadratic. Near the minimizer, the deterministic drift may satisfy \[\nabla^2H(x)\asymp \left\lVert x \right\rVert^{m-2}I_d, \qquad m\ge2,\] in dimension \(d\), so the restoring force is of order \(\left\lVert x \right\rVert^{m-1}\). At the same time, the objective may have subquadratic tails: for \(1\le \beta<2\), the drift for large \(\left\lVert x \right\rVert\) may grow only as \(\left\lVert x \right\rVert^{\beta-1}\). These features occur in standard statistical objectives. Quantile estimation when the distribution function crosses the target probability level at higher order provides a canonical example of local flatness. For median and \(L_1\) regression, this type of nonregular local behavior was studied by Knight [11]; their regularly varying mechanism applies to general quantile loss introduced by Koenker and Bassett [12]. The same local geometry appears in the Rockafellar–Uryasev variational representation of conditional value-at-risk (CVaR) [13]. Subquadratic tails arise in robust losses with bounded or sublinear scores, such as Huber-type and generalized Charbonnier losses [14][16], and in logistic losses [17]. Recent work on subquadratic SGD treats the locally strongly convex case \(m=2\) with \(\beta<2\) [18]; the case \(m>2\) requires a different local scale and a different contraction argument.

We emphasize the different roles played by the two key exponents used throughout the paper: The local flatness exponent \(m\) determines the order of the restoring drift near the minimizer and therefore the stationary scaling. The tail exponent \(\beta\) determines the drift for large \(\left\lVert x \right\rVert\) and the weight function needed to control excursions.

Moreover, we consider SGD with Markovian noise. A driving Markov chain \((\xi_n)_{n \ge 0}\) evolves according to a contractive recursion, and the SGD iterate follows \[X_{n+1}=X_n-\alpha\{h(X_n)+g(X_n,\xi_{n+1})\}, \qquad h=\nabla H.\] The augmented process \((X_n,\xi_n)\) is Markov. This formulation covers not only statistical settings with i.i.d.noise but also temporally dependent data streams.

1.0.0.1 Challenges and main contributions.

There are three difficulties for analyzing SGD in the above setting. First, ordinary Euclidean contraction degenerates near the minimizer when \(m>2\), which is also the region where the invariant laws concentrate as \(\alpha\downarrow0\). Second, for \(\beta<2\), quadratic Lyapunov functions are not well matched to the drift for large \(\left\lVert x \right\rVert\). Third, under Markovian noise, the current iterate is correlated with future values of the driving chain, which prevents a direct one-step martingale argument. Our main results below address these challenges.

First, for the convergence of SGD with a small fixed stepsize \(\alpha\), we prove that the SGD iterates converge to a unique invariant law geometrically in a suitably chosen Wasserstein distance, with contraction factor \(1-c\alpha^{m-1}\) for a constant \(c>0\). The proof uses the metric of Qu, Blanchet, and Glynn [19] induced by a Lyapunov weight function \(V_\alpha\) that we carefully construct. The weight function consists of two terms, one for controlling the subquadratic tail of the objective and the other for controlling the near-minimizer region. Moreover, directly applying the result of [19] does not work in our case, and we show contraction of SGD iterates with a new argument.

Second, with the fixed-\(\alpha\) invariant law in hand, we then identify its scaling limit as \(\alpha \downarrow 0\). The local \(m\)-flat geometry determines the normalization: the SGD iterates concentrate on the scale \(\alpha^{1/m}\). Then, the scaling limit is given by the stationary distribution of the stochastic differential equation \[dY_t=-h_0(Y_t)\,dt+\Sigma^{1/2}\,dB_t ,\] where \(h_0\) denotes the limiting drift, and \(\Sigma\) denotes the asymptotic variance which is identified via a Poisson equation argument. When \(m=2\), the above equation recovers the Ornstein–Uhlenbeck diffusion, and the invariant law is Gaussian. For \(m>2\), the drift \(h_0\) is nonlinear, and the stationary distribution is generally non-Gaussian. We note that Chen, Mou, and Maguluri [8] numerically exhibited the non-Gaussian quartic scaling. Wang et al. [20] conjectured a one-dimensional flat-minimum Gibbs approximation with the corresponding nonstandard scaling and supported it numerically. Our result establishes the scaling limit rigorously in a multidimensional setting with Markovian noise, confirming and extending the phenomenon suggested by prior work.

Finally, for coordinate-separable objective functions, we provide the corresponding constant-stepsize and scaling limit results. Notably, if the coordinates have unequal local flatness exponents, only the flattest coordinates can remain nonzero under the common scale associated with the largest exponent.

Table 1 summarizes the cases appearing in the main results.

Table 1: The primary features covered by the main results. The exponent \(m\) is local and determines the normalization of the invariant law and the limiting drift \(h_0\). The exponent \(\beta\) is global and determines the tail part of the induced metric.
Feature Representative condition Effect on the invariant law
Quadratic \(h(x)\sim Ax\) \(\alpha^{1/2}\) scale, Gaussian limit
Flat minimum \(h(x)\sim \norm{x}^{m-2}x\) \(\alpha^{1/m}\) scale, nonlinear diffusion
Markovian noise correlated \(g(0,\xi_n)\) asymptotic covariance \(\Sigma\)
Subquadratic tail \(\norm{h(x)}\asymp \norm{x}^{\beta-1}\) \(\norm{x}^{2-\beta}\) term in the Lyapunov weight

1.0.0.2 Related literature.

The classical stochastic approximation literature studies decreasing stepsizes and averaging of iterates, establishing almost-sure convergence and asymptotic normality [1][5], [21][24]. Modern nonasymptotic analyses of decreasing-step stochastic optimization include [25][27]. These results concern convergence of the iterates to an optimizer. The object here is different: for a constant stepsize, the last iterate has a nontrivial invariant law, and the limit \(\alpha\downarrow0\) is taken at stationarity.

For constant stepsizes, the algorithm is analyzed as a Markov chain. Pflug [6] studied small-stepsize asymptotic laws for stochastic optimization. Mandt, Hoffman, and Blei [7] developed quadratic diffusion approximation for constant-stepsize SGD. Dieuleveut, Durmus, and Bach [9] developed invariant measure and bias expansions for smooth strongly convex SGD. Merad and Gaïffas [10] proved Wasserstein convergence and concentration in related strongly convex settings. In a different asymptotic regime, Yu, Balasubramanian, Volgushev, and Erdogdu [28] established a central limit theorem for averages of iterates at a fixed stepsize, together with a characterization of the invariant bias.

Besides the classical work of Pflug [6], the recent works most closely related to our scaling limit of invariant laws are Chen, Mou, and Maguluri [8], Zhang et al. [29], and Wang et al. [20]. Pflug [6] and Chen, Mou, and Maguluri [8] both first pass to constant-stepsize invariant laws and then let \(\alpha\downarrow0\) to obtain a Gaussian limit. The latter work obtains Gaussian limits in smooth strongly convex, linear, and contractive settings, and numerically demonstrates the \(\alpha^{1/4}\) scaling and a non-Gaussian limit for a quartic objective. Zhang et al. [29] establish steady-state convergence for nonsmooth contractive stochastic approximation: their scale remains \(\sqrt\alpha\), while nonsmooth local dynamics can produce a non-Gaussian limit. Wang et al. [20] give a one-dimensional Gibbs approximation for flat convex objectives with the correct nonstandard scaling, conditional on two conjectures concerning stationary moments and Stein equation regularity. We establish the flat-minimum limit unconditionally and allow a multidimensional setting with Markovian noise.

The tail behavior we study in this work is closest to recent work on subquadratic SGD for robust and quantile regression [18]. In the notation below, that setting corresponds to \(m=2\) and \(\beta<2\): the objective is locally strongly convex, but its tail for large \(\left\lVert x \right\rVert\) is subquadratic. The present paper allows this tail behavior to coexist with local flatness \(m>2\).

Markovian noise introduces a separate issue because the stationary iterate is correlated with future values of the driving chain. Poisson equations are the standard tool for separating this temporal dependence in stochastic approximation. For linear stochastic approximation, Huo, Chen, and Xie [30] proved convergence to a unique stationary distribution and developed a stepsize expansion of its bias. More recent analyses quantify how Markovian memory and nonlinear updates affect stationary bias and weak convergence [31], [32]. Here the Poisson equation facilitates identifying the diffusion covariance \(\Sigma\) as the asymptotic covariance of the stationary sequence \(g(0,\xi_n)\).

Our proof of constant-stepsize ergodicity builds on recent work by Qu, Blanchet, and Glynn [19]. Their work proves Wasserstein convergence from contractive drift conditions. We do not verify their conditions directly as they may not always hold in our case. Thus, while we adopt the induced metric of [19], our constant-stepsize ergodicity result requires a new directional kernel estimate adapted to the SGD dynamics.

For the scaling limit as \(\alpha\downarrow0\), a related line of work addresses quantitative stationary approximations, often through generator comparisons or Stein-type arguments. As a methodological precursor, Gast [33] used generator comparison to obtain rates for mean-field models. Allmeier and Gast [34] used related generator expansions to compute bias corrections for constant-step stochastic approximation with Markovian noise. Wang et al. [20] proved nonasymptotic Wasserstein and tail bounds in smooth settings, including Markovian noise through Poisson equations. Wei et al. [35] obtained finite-sample Gaussian approximations for constant-stepsize SGD through linearization. Our result instead identifies the scaling limit of invariant laws for multidimensional flat objectives, where the limiting drift is nonlinear and the limit is generally non-Gaussian. While we do not obtain a finite-\(\alpha\) convergence rate, to the best of our knowledge this is the first rigorous invariant-law scaling limit for flat minima, even in dimension one.

1.0.0.3 Organization.

The paper is organized around the two limits involved in the analysis: first the limit of the SGD iterates at constant stepsize, and then the small-stepsize limit at stationarity. Section 2 states all the main results. Section 3 studies statistical examples that motivate local flatness and subquadratic tails. Section 4 provides the proof strategy and proves the main theorems. Additional technical proofs are deferred to the appendices.

2 Main results↩︎

To simplify the statement of the main results, we assume that zero is the minimizer of the objective function without loss of generality. Indeed, with \(x_\star\) denoting the minimizer of the original objective, replacing the optimization variable by \(x-x_\star\) translates the minimizer to the origin.

Let \(H:\mathbb{R}^d\to\mathbb{R}\) be the objective function, and let \(h:=\nabla H\). Consider SGD with Markovian noise defined as follows. Let \((\xi_n)\) be the driving Markov chain. It takes values in a closed convex set \(\Xi\subseteq\mathbb{R}^{d_\Xi}\) and evolves according to \[\xi_{n+1}=\Phi(\xi_n,U_{n+1}),\] where \((U_n)_{n\ge1}\) are i.i.d.innovations taking values in a measurable space \(\mathcal{U}\) and \(\Phi:\Xi\times\mathcal{U}\to\Xi\) is Borel measurable. For a constant stepsize \(\alpha\in(0,1)\), the SGD is defined by \[\label{eq:markov-sgd-main} X_{n+1} = X_n-\alpha\{h(X_n)+g(X_n,\xi_{n+1})\}.\tag{1}\] Equivalently, with \(Z_n=(X_n,\xi_n)\in\mathsf Z:=\mathbb{R}^d\times\Xi\), \[Z_{n+1}=F_{\alpha,U_{n+1}}(Z_n),\] where \[\label{eq:augmented-random-map} F_{\alpha,u}(x,\xi) := \left( x-\alpha\{h(x)+g(x,\Phi(\xi,u))\}, \Phi(\xi,u) \right).\tag{2}\] We suppress the dependence on \(\alpha\) when it is clear from context. Note that the augmented process \(Z_n\) is a Markov chain while the \(X\)-coordinate alone need not be Markov.

Remark 1. The i.i.d.noise setting is a special case of the preceding Markovian formulation. Indeed, to construct i.i.d.\(\xi_1,\xi_2,\ldots\) with common law \(\mu\), it suffices to take \[\mathcal{U}=\Xi,\qquad U_1,U_2,\ldots\stackrel{\rm i.i.d.}{\sim}\mu, \qquad \Phi(\xi,u)=u.\] Then \[\xi_{n+1}=\Phi(\xi_n,U_{n+1})=U_{n+1}.\] Consequently, the results below apply to this i.i.d.setting.

2.1 Geometric ergodicity in Wasserstein distance↩︎

Our first result is regarding the convergence of the constant-stepsize SGD to an invariant law. More precisely, we establish the geometric ergodicity of the SGD iterates in a Wasserstein distance. We first introduce the metrics used in this work. These metrics come from the work of Qu, Blanchet, and Glynn [19], specialized to the augmented finite-dimensional state space.

Definition 1. Let \(E\) be a closed convex subset of a finite-dimensional normed space, with base norm \(\vert\cdot\vert_\star\). For a continuous weight function \(V:E\to[1,\infty)\), define the metric induced by \(V\) \[d_{V,\star}(z,z') := \inf_{\gamma\in\operatorname{AC}_E(z,z')} \int_0^1 V(\gamma(t))\,\vert\dot{\gamma}(t)\vert_\star\,dt,\] where \(\operatorname{AC}_E(z,z')\) is the set of absolutely continuous curves \(\gamma:[0,1]\to E\) with \(\gamma(0)=z\) and \(\gamma(1)=z'\). For probability measures \(\mu,\nu\) on \(E\), define \[W_{V,\star}(\mu,\nu) := \inf_{\gamma\in\Pi(\mu,\nu)} \int_{E\times E} d_{V,\star}(z,z')\,\gamma(dz,dz'),\] where \(\Pi(\mu,\nu)\) is the set of couplings of \(\mu\) and \(\nu\).

When \(V\equiv1\) on \(\mathbb{R}^d\) with the Euclidean norm, we simply write \(W_1\) for the Wasserstein distance. For the augmented chain, the base norm is defined to be \[\label{eq:aug-base-norm} \vert (x,\xi)\vert_\alpha := \left\lVert x \right\rVert+\alpha^{-1}\left\lVert \xi \right\rVert, \qquad (x,\xi)\in\mathbb{R}^d\times\mathbb{R}^{d_\Xi},\tag{3}\] where the choice of the balancing factor \(\alpha^{-1}\) is suggested by the analysis. The corresponding induced metric and Wasserstein distance are denoted by \(d_{V,\alpha}\) and \(W_{V,\alpha}\).

Next, we introduce the assumptions on the objective function \(H\).

Assumption 1 (Objective function). Fix \(m\ge2\) and \(\beta\in[1,2]\). The objective \(H:\mathbb{R}^d\to\mathbb{R}\) satisfies the following conditions.

  1. Regularity and convexity. \(H\in C^2(\mathbb{R}^d)\), \(H\) is convex, and \(0\in\arg\min H\).

  2. Local flatness. There exists \(R_H>0\) and constants \(c_{\rm in},C_{\rm in}>0\) such that, for all \(0<\left\lVert x \right\rVert\le R_H\), \[\nabla^2H(x)\succeq c_{\rm in}\left\lVert x \right\rVert^{m-2}I_d, \qquad \left\lVert h(x) \right\rVert\le C_{\rm in}\left\lVert x \right\rVert^{m-1}.\] When \(m=2\), this condition is understood to extend to \(x=0\) by continuity of \(\nabla^2H\).

  3. Tail condition. One of the following two tail conditions holds.

    1. Quadratic tail. For \(\beta=2\), there exists \(c_{\rm out}>0\) such that, for all \(\left\lVert x \right\rVert\ge R_H\), \[\nabla^2H(x)\succeq c_{\rm out}I_d.\]

    2. Subquadratic tail. For \(1\le\beta<2\), there exist \(c_{\rm out},C_{\rm out}>0\) such that, for all \(\left\lVert x \right\rVert\ge R_H\), \[\left\langle x,h(x)\right\rangle\ge c_{\rm out}\left\lVert x \right\rVert^{\beta},\qquad \left\lVert h(x) \right\rVert\le C_{\rm out}\left\lVert x \right\rVert^{\beta-1}.\]

The tail condition ([[itm:H3]](#itm:H3){reference-type=“ref” reference=“itm:H3”}) supplies the restoring drift towards the minimizer for the SGD, while the local flatness ([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}) determines the scale of fluctuations near the minimizer and the scaling limit.

Remark 2. To exemplify ([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}), we note that it includes locally defined functions:

  • powers of norms such as \(H(x)=\|x\|^m/m\) and \(H(x)=m^{-1}(x^\top A x)^{m/2}\) with \(A\succ0\);

  • polynomials such as \(H(x)=\|x\|^4+c\|x\|_4^4\) and \(H(x)=\|x\|^4 + c\|x\|^6\) where \(c>0\) for \(m = 4\);

  • non-polynomial objectives such as \(H(x)=\sin^4(\|x\|)/4\).

Examples with a degenerate Hessian or unequal flatness along different coordinates, such as \(x_1^4+x_2^4\) and \(x_1^4+x_2^6\), fail ([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}), but they are covered by the coordinate-separable results in Section 2.3. Exponentially flat functions, such as \(H(x)=e^{-1/x^2}\) near \(0\) with \(H(0)=0\), vanish to infinite order at the minimizer and therefore do not satisfy ([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}) for any finite \(m\).

The next assumptions concern the noise in the SGD and the one-step iterate in 12 .

Assumption 2 (Stochastic update). The driving Markov chain and stochastic update satisfy the following conditions.

  1. Driving-chain contraction. There exists a measurable function \(L_\Phi:\mathcal{U}\to[0,1]\) such that, for all \(\xi,\eta\in\Xi\) and all \(u\in\mathcal{U}\), \[\left\lVert \Phi(\xi,u)-\Phi(\eta,u) \right\rVert \le L_\Phi(u)\left\lVert \xi-\eta \right\rVert, \qquad \mathbb{E}L_\Phi(U_1)<1.\]

  2. Lipschitzness of noise. There exists \(L_{g,\Phi}>0\) such that, for all \(x\in\mathbb{R}^d\), all \(\xi,\eta\in\Xi\), and all \(u\in\mathcal{U}\), \[\left\lVert g(x,\Phi(\xi,u))-g(x,\Phi(\eta,u)) \right\rVert \le L_{g,\Phi}\left\lVert \xi-\eta \right\rVert.\]

  3. Reference-point integrability. There exists \(\xi_\star\in\Xi\) such that \[\label{eq:markov-reference-state} \mathbb{E}\left\lVert \Phi(\xi_\star,U_1)-\xi_\star \right\rVert<\infty.\tag{4}\] If \(\beta=2\), assume also \[\mathbb{E}\left\lVert g(0,\Phi(\xi_\star,U_1)) \right\rVert<\infty.\] When \(1\le\beta<2\), this condition follows from ([[itm:N5]](#itm:N5){reference-type=“ref” reference=“itm:N5”}) at \(x=0\).

  4. Subquadratic tail dissipativity. Define \[\label{eq:markov-conditional-mean-g} \bar g(x,\xi) :=\mathbb{E}_{U_1}[g(x,\Phi(\xi,U_1))].\tag{5}\] When \(\beta=2\), this condition is not imposed. When \(1\le\beta<2\), there exists \(c_{\rm diss}>0\) such that, for all \(\left\lVert x \right\rVert\ge R_H\) and all \(\xi\in\Xi\), \[\left\langle x,h(x)+\bar g(x,\xi)\right\rangle \ge c_{\rm diss}\left\lVert x \right\rVert^{\beta}.\]

  5. Subquadratic exponential integrability. When \(\beta=2\), this condition is not imposed. When \(1\le\beta<2\), there exists \(\lambda_0>0\) such that \[\sup_{x\in\mathbb{R}^d}\sup_{\xi\in\Xi} \mathbb{E}\exp\!\left( \lambda_0\frac{\left\lVert g(x,\Phi(\xi,U_1)) \right\rVert}{1+\left\lVert x \right\rVert^{\beta-1}} \right)<\infty.\] Here and below, \(\left\lVert x \right\rVert^0=1\) by convention.

  6. Nondegeneracy of noise at minimizer. If \(m=2\), this condition is not imposed. If \(m>2\), assume that there exist constants \(\varepsilon_g>0\) and \(p_g>0\) such that \[\inf_{\xi\in\Xi} \mathbb{P}\!\left(\left\lVert g(0,\Phi(\xi,U_1)) \right\rVert\ge\varepsilon_g\right) \ge p_g.\]

  7. Co-coercivity and noise perturbation. For \((x,\xi)\in\mathbb{R}^d\times\Xi\), set \[\mathsf G(x,\xi):=h(x)+g(x,\xi).\] For every \(\xi\in\Xi\), the map \(x\mapsto\mathsf G(x,\xi)\) is \(C^1\). There exist constants \(L_{\mathsf G}\ge1\) and \(\theta\in[0,1)\) such that, for every \(x,y\in\mathbb{R}^d\) and every \(\xi\in\Xi\), \[\label{eq:sample-cocoercivity} \left\langle x-y,\mathsf G(x,\xi)-\mathsf G(y,\xi)\right\rangle \ge L_{\mathsf G}^{-1} \left\lVert \mathsf G(x,\xi)-\mathsf G(y,\xi) \right\rVert^2,\tag{6}\] and \(\bar g\) defined in 5 satisfies \[\label{eq:mean-perturbation-control} \left\langle x-y,\bar g(x,\xi)-\bar g(y,\xi)\right\rangle \ge -\theta\left\langle x-y,h(x)-h(y)\right\rangle.\tag{7}\]

Remark 3. We now make a few clarifications about the above assumptions.

Before stating the first result, recall that the augmented space \(\mathsf Z=\mathbb{R}^d\times\Xi\) is equipped with the base norm \(\vert\cdot\vert_\alpha\) in 3 . The metrics \(d_{V,\alpha}\) and \(W_{V,\alpha}\) are those of Definition 1. For a metric \(d\) on a space \(E\), define \[\mathcal{P}_1(E,d) := \left\{\mu:\int_E d(z,z_0)\,\mu(dz)<\infty \text{ for some } z_0\in E\right\}.\]

Theorem 4 (Geometric ergodicity of constant-stepsize SGD). Assume Assumptions 1 and 2. Then there exist \(\alpha_0>0\) and \(c\in(0,1)\) such that, for every \(0<\alpha\le\alpha_0\), one can choose a continuous weight function \(V_\alpha:\mathsf Z\to[1,\infty)\) for which the augmented chain \(Z_n=(X_n,\xi_n)\) admits a unique invariant law \[\pi_\alpha\in\mathcal{P}_1(\mathsf Z,d_{V_\alpha,\alpha}).\] Moreover, \[\label{eq:markov-contraction-laws} W_{V_\alpha,\alpha}(\mu P_\alpha^n,\nu P_\alpha^n) \le (1-c\alpha^{m-1})^n W_{V_\alpha,\alpha}(\mu,\nu), \qquad n\ge0,\tag{8}\] for all initial laws \(\mu,\nu\) with finite \(W_{V_\alpha,\alpha}(\mu,\nu)\), where \(P_\alpha\) is the transition kernel of the augmented chain.

Remark 5. The contraction factor \(1-c\alpha^{m-1}\) reflects the local flatness of the objective. In the quadratic case \(m=2\), the deterministic curvature at the minimizer gives the factor \(1-c\alpha\), matching the usual strongly convex behavior. When \(m>2\), the curvature vanishes exactly where the invariant law concentrates. Nevertheless, the locally flat objective and the noise still ensure strict contraction for each fixed small stepsize, with the contraction factor approaching one at order \(\alpha^{m-1}\).

The proof of Theorem 4 is given in Section [sec:proof-fixed-markov]. Projecting the contraction of the augmented chain to the \(X\)-coordinate gives the following ordinary \(W_1\) consequence, whose proof is given in Appendix 6.7.

Corollary 1. Assume the hypotheses of Theorem 4, and fix \(\alpha\in(0,\alpha_0]\). Let \((\pi_\alpha)_X\) denote the \(X\)-marginal of the invariant law \(\pi_\alpha\) from Theorem 4. Let \(z_\star=(0,\xi_\star)\), with \(\xi_\star\) as in Assumption 2([[itm:N3]](#itm:N3){reference-type=“ref” reference=“itm:N3”}). There exists a sufficiently small constant \(\kappa>0\), depending only on the constants in Assumptions 1 and 2 and independent of \(\alpha\), such that the following holds. Define \[\Gamma_\beta(r):= (1+r)^{\beta-1} \exp\!\left\{ \kappa\bigl((1+r)^{2-\beta}-1\bigr) \right\}, \qquad 1\le\beta\le2.\] Then \[\mathcal{P}_1 \bigl(\mathsf Z,d_{V_\alpha,\alpha}\bigr) = \left\{ \mu: \int_{\mathsf Z} \Bigl[ \Gamma_\beta(\left\lVert x \right\rVert) +\alpha^{-1}\left\lVert \xi-\xi_\star \right\rVert \Bigr]\,\mu(dx,d\xi)<\infty \right\}.\] For every \(\mu\in\mathcal{P}_1(\mathsf Z,d_{V_\alpha,\alpha})\) and every \(n\ge0\), \[W_1\bigl((\mu P_\alpha^n)_X,(\pi_\alpha)_X\bigr) \le (1-c\alpha^{m-1})^n W_{V_\alpha,\alpha}(\mu,\pi_\alpha).\]

The set \(\mathcal{P}_1 \bigl(\mathsf Z,d_{V_\alpha,\alpha}\bigr)\) is explicitly characterized in the above corollary. In particular, every deterministic initial condition belongs to \(\mathcal{P}_1(\mathsf Z,d_{V_\alpha,\alpha})\). When \(1\le\beta<2\), membership in this class requires an exponential moment of order \(2-\beta\) in the \(X\)-coordinate and a first moment in the \(\xi\)-coordinate.

2.2 Small-stepsize scaling limit↩︎

The invariant law \(\pi_\alpha\) of the SGD iterates in the last subsection is implicit and generally difficult to characterize. Therefore, we turn to studying its scaling limit as \(\alpha \downarrow 0\). To this end, we introduce two additional conditions.

Assumption 3 (Objective function). Assume Assumption 1 together with the following.

  1. Local gradient expansion. There exists \(H_0\in C^2(\mathbb{R}^d)\), convex and homogeneous of degree \(m\), such that \(H_0(0)=0\), \(\nabla H_0(0)=0\), and \[\nabla H(x)=\nabla H_0(x)+o(\left\lVert x \right\rVert^{m-1}) \qquad\text{as }x\to0.\]

Assumption 4 (Stochastic update). Assume Assumption 2 together with the following. Let \(\pi_\Xi\) denote the invariant law of the driving chain \((\xi_n)_{n \ge 0}\), whose existence and uniqueness are proved in Lemma 20. The Markovian noise satisfies the following condition.

  1. Stationary centering and second moments. The gradient error \(g(x,\xi)\) satisfies \[\label{eq:stationary-centering-gkx} \int_\Xi g(x,\xi)\,\pi_\Xi(d\xi)=0, \qquad x\in\mathbb{R}^d,\tag{9}\] and \[\label{eq:g0-moment-gk} \int_\Xi \bigl(\left\lVert g(0,\xi) \right\rVert^2+\left\lVert \xi-\xi_\star \right\rVert^2\bigr)\, \pi_\Xi(d\xi)<\infty.\tag{10}\]

Let \((\xi_n)_{n\ge0}\) be the driving chain initialized according to its invariant law \(\pi_\Xi\), so that \((\xi_n)_{n\ge0}\) is stationary. By Assumption 4([[itm:N8]](#itm:N8){reference-type=“ref” reference=“itm:N8”}), \((g(0,\xi_n))_{n\ge0}\) is centered and stationary. Define \[\label{eq:GK-lag-series} \Sigma:=\Gamma_0+\sum_{k=1}^{\infty}(\Gamma_k+\Gamma_k^\top), \qquad \Gamma_k:=\mathbb{E}[g(0,\xi_0)g(0,\xi_k)^\top].\tag{11}\] We will show that \(\Sigma\) is the asymptotic covariance appearing in the scaling limit. To see that \(\Sigma\) is a well-defined covariance matrix, Lemma 5 shows that the series converges absolutely and that \(\Sigma\) is a finite positive semidefinite matrix.

Remark 6. The identity 11 is the Green–Kubo representation of the asymptotic covariance; see, for example, [3], [4]. For Markovian noise, the correlations between \(\xi_k\)’s accumulate over time. Lemma 5 uses the Poisson equation to decompose the correlated sequence into a conditionally centered increment plus a telescoping term. The telescoping part is lower order in the generator calculation, while the quadratic variation of the centered increment gives \(\Sigma\) as in 11 .

The following theorem, proved in Section [sec:proof-clt-markov], identifies both the scale \(\alpha^{1/m}\) of the invariant law and its limiting distribution as the solution of a stochastic differential equation.

Theorem 7 (Scaling limit of invariant law). Assume Assumptions 3 and 4. Let \(H_0\) be as in Assumption 3([[itm:H4]](#itm:H4){reference-type=“ref” reference=“itm:H4”}) and set \(h_0:=\nabla H_0\). For each sufficiently small \(\alpha>0\), let \((X_\infty^{(\alpha)},\xi_\infty^{(\alpha)})\) have law \(\pi_\alpha\) from Theorem 4, and set \[Y_\alpha:=\alpha^{-1/m}X_\infty^{(\alpha)}.\] Then \(Y_\alpha\Rightarrow Y_\infty\) as \(\alpha\downarrow0\), where \(Y_\infty\) is the unique invariant distribution of \[\label{eq:limitSDE-markov} dY_t=-h_0(Y_t)\,dt+\Sigma^{1/2}\,dB_t,\tag{12}\] where \(\Sigma\) is defined in 11 and \(B=(B_t)_{t\ge0}\) is a standard \(d\)-dimensional Brownian motion.

We note that the covariance matrix \(\Sigma\) is allowed to be degenerate, in which case the scaling limit is singular.

Remark 8. If \(m=2\) and \(h_0(y)=Ay\), then 12 defines an Ornstein–Uhlenbeck process. Its invariant law is Gaussian with covariance \(C\) solving \[AC+CA^\top=\Sigma.\] Thus Theorem 7 recovers the classical \(\sqrt\alpha\) scaling and Gaussian limit. For \(m>2\), the drift \(h_0\) is nonlinear, the scale becomes \(\alpha^{1/m}\), and the limiting invariant law is non-Gaussian.

2.3 Coordinate-separable objectives↩︎

The preceding results impose a common local flatness exponent in all directions. We now consider coordinate-separable objectives, for which the flatness exponent, and hence the natural scale of the invariant law, may vary across coordinates. Suppose that \[\label{eq:separable-model} H(x)=\sum_{i=1}^d H_i(x_i), \qquad g(x,\xi)=\bigl(g_1(x_1,\xi),\ldots,g_d(x_d,\xi)\bigr).\tag{13}\] Writing \(h_i=H_i'\), the SGD recursion becomes \[\label{eq:separable-sgd} X_{n+1,i} =X_{n,i}-\alpha\{h_i(X_{n,i})+g_i(X_{n,i},\xi_{n+1})\}, \qquad 1\le i\le d,\tag{14}\] where \(\xi_{n+1}=\Phi(\xi_n,U_{n+1})\) is the driving-chain recursion from the beginning of this section. Thus the coordinates may remain dependent through the common driving chain.

For each \(1\le i\le d\), assume that the one-dimensional pair \((H_i,g_i)\) satisfies the one-dimensional versions of Assumptions 1 and 2, with parameters \((m_i,\beta_i)\). The constants in these assumptions may depend on \(i\). Applying the construction of Section 4.1 with \(d=1\) and \((H,g,m,\beta)\) replaced by \((H_i,g_i,m_i,\beta_i)\), on \(\mathbb{R}\times\Xi\) equipped with the common base norm \[\lvert(x_i,\xi)\rvert_\alpha :=|x_i|+\alpha^{-1}\left\lVert \xi \right\rVert,\] gives a weight \(V_{\alpha,i}\) and an induced metric \(d_{i,\alpha}:=d_{V_{\alpha,i},\alpha}^{(i)}\). Define the additive metric on \(\mathbb{R}^d\times\Xi\) by \[\label{eq:separable-metric} d_{{\rm sep},\alpha}((x,\xi),(y,\eta)) :=\sum_{i=1}^d d_{i,\alpha}((x_i,\xi),(y_i,\eta)),\tag{15}\] and let \(W_{{\rm sep},\alpha}\) denote the Wasserstein–1 distance with cost \(d_{{\rm sep},\alpha}\). Set \[m_{\max}:=\max_{1\le i\le d}m_i.\]

The following corollary extends Theorem 4 to this separable setting.

Corollary 2 (Geometric ergodicity for separable objectives). Suppose that 13 holds and that, for every \(1\le i\le d\), the pair \((H_i,g_i)\) satisfies the one-dimensional versions of Assumptions 1 and 2. Then there exist \(\alpha_0\in(0,1)\) and \(c_0\in(0,1)\) such that, for every \(\alpha\in(0,\alpha_0]\), the augmented chain \(Z_n=(X_n,\xi_n)\) admits a unique invariant law \[\pi_\alpha^{\rm sep} \in\mathcal{P}_1(\mathbb{R}^d\times\Xi,d_{{\rm sep},\alpha}).\] Moreover, if \(P_\alpha^{\rm sep}\) denotes its transition kernel, then \[\label{eq:sep-markov-contraction} W_{{\rm sep},\alpha} \bigl(\mu(P_\alpha^{\rm sep})^n,\nu(P_\alpha^{\rm sep})^n\bigr) \le (1-c_0\alpha^{m_{\max}-1})^n W_{{\rm sep},\alpha}(\mu,\nu), \qquad n\ge0,\tag{16}\] for all probability laws \(\mu,\nu\) for which the right-hand side is finite.

The contraction rate is determined by \(m_{\max}\): the coordinates with the largest flatness exponent have the largest contraction factor and hence govern the convergence of the full chain. Corollary 2 is proved in Appendix 8.1.

Remark 9. The setup above also includes independent coordinate driving chains as a special case. More precisely, suppose that \(\Xi=\prod_{i=1}^d\Xi_i\), the driving map and innovations act independently across coordinates, and each \(g_i\) depends only on the \(i\)th coordinate of \(\xi\). Then the transition kernel is a product kernel. Hence the product of the one-dimensional invariant laws is invariant and, by the uniqueness in Corollary 2, equals \(\pi_\alpha^{\rm sep}\).

We next turn to the scaling limit. For each coordinate, assume in addition the one-dimensional versions of Assumptions 3 and 4. Let \(H_{i,0}\) be the homogeneous function appearing in the one-dimensional version of Assumption 3([[itm:H4]](#itm:H4){reference-type=“ref” reference=“itm:H4”}), and set \(h_{i,0}:=H_{i,0}'\). Initialize the driving chain in its invariant law \(\pi_\Xi\), and define \[g_0^{\rm sep}(\xi) :=\bigl(g_1(0,\xi),\ldots,g_d(0,\xi)\bigr), \qquad \xi\in\Xi.\] Then \((g_0^{\rm sep}(\xi_n))_{n\ge0}\) is centered and stationary. Its asymptotic covariance is \[\label{eq:separable-GK-series} \Sigma_{\rm sep} :=\Gamma_0^{\rm sep} +\sum_{k=1}^{\infty} \bigl(\Gamma_k^{\rm sep}+(\Gamma_k^{\rm sep})^\top\bigr), \qquad \Gamma_k^{\rm sep} :=\mathbb{E}\bigl[g_0^{\rm sep}(\xi_0)g_0^{\rm sep}(\xi_k)^\top\bigr].\tag{17}\] The coordinatewise Poisson construction in Lemma 5 shows that the series converges absolutely and that \(\Sigma_{\rm sep}\) is finite and positive semidefinite. Its off-diagonal entries retain the temporal cross-covariances created by a common driving chain. Let \[\mathcal{I}_*:=\{i:m_i=m_{\max}\}, \qquad P_*:=\operatorname{diag}\bigl(\mathbf{1}_{\{i\in\mathcal{I}_*\}}\bigr), \qquad \Sigma_*:=P_*\Sigma_{\rm sep}P_*. \label{eq:def-Sigma-star}\tag{18}\] Both \(\Sigma_{\rm sep}\) and \(\Sigma_*\) are allowed to be degenerate.

For each \(1\le i\le d\), the \((x_i,\xi)\)-marginal of \(\pi_\alpha^{\rm sep}\) is invariant for the one-dimensional augmented chain associated with \((H_i,g_i)\). Therefore Theorem 7 applied to this chain gives \[\alpha^{-1/m_i}X_{\infty,i}^{(\alpha),\rm sep} \Rightarrow Y_{\infty,i},\] where \(Y_{\infty,i}\) has the unique invariant law of the one-dimensional diffusion \[dY_{i,t} =-h_{i,0}(Y_{i,t})\,dt +(\Sigma_{\rm sep})_{ii}^{1/2}\,dB_{i,t},\] and \(B_i\) is a standard one-dimensional Brownian motion. Thus coordinate \(i\) has natural scale \(\alpha^{1/m_i}\).

Corollary 3 (Scaling limit for separable objectives). Suppose that 13 holds and that, for every \(1\le i\le d\), the pair \((H_i,g_i)\) satisfies the one-dimensional versions of Assumptions 3 and 4. Let \((X_\infty^{(\alpha),\rm sep},\xi_\infty^{(\alpha),\rm sep})\) have law \(\pi_\alpha^{\rm sep}\), and set \[h_0^{\rm sep}(y) :=\bigl(h_{1,0}(y_1),\ldots,h_{d,0}(y_d)\bigr).\] Then, under the common normalization determined by \(m_{\max}\), \[\alpha^{-1/m_{\max}}X_\infty^{(\alpha),\rm sep} \Rightarrow Y_\infty^{\rm sep},\] where \(Y_\infty^{\rm sep}\) has the unique invariant law of the \(d\)-dimensional diffusion \[\label{eq:sep-limitSDE-d} dY_t =-h_0^{\rm sep}(Y_t)\,dt +\Sigma_*^{1/2}\,dB_t.\tag{19}\] Here, \(\Sigma_*\) is defined in 18 and \(B\) is a standard \(d\)-dimensional Brownian motion.

For \(i\notin\mathcal{I}_*\), the \(i\)th component of 19 has no Brownian forcing and follows the deterministic flow \(dY_{i,t}=-h_{i,0}(Y_{i,t})\,dt\), whose unique invariant law is \(\delta_0\). Thus, under the common normalization, only the coordinates with the largest flatness exponent can have a nonzero limit. If \(m_i=m\) for every \(i\), then \(P_*=I_d\) and \(\Sigma_*=\Sigma_{\rm sep}\). Corollary 3 then gives \(\alpha^{-1/m}X_\infty^{(\alpha),\rm sep} \Rightarrow Y_\infty^{\rm sep},\) where 19 becomes \(dY_t =-h_0^{\rm sep}(Y_t)\,dt +\Sigma_{\rm sep}^{1/2}\,dB_t.\) The proof of Corollary 3 is given in Appendix 8.2.

Remark 10. Under the independent-coordinate conditions of Remark 9, the naturally rescaled vector satisfies \[\left( \frac{X_{\infty,1}^{(\alpha),\rm sep}}{\alpha^{1/m_1}},\ldots, \frac{X_{\infty,d}^{(\alpha),\rm sep}}{\alpha^{1/m_d}} \right) \Rightarrow (Y_{\infty,1},\ldots,Y_{\infty,d}),\] where the limiting coordinates are independent. See Appendix 8.3 for more details.

3 Applications and numerical experiments↩︎

This section illustrates the roles of the local exponent \(m\) and the tail exponent \(\beta\), and reports numerical experiments for the invariant-law predictions. Section 3.1 uses quantile and tail-risk estimation to interpret \(m\) as a local flatness exponent and illustrates its consequences for local scaling. Section 3.2 uses robust and logistic losses to interpret \(\beta\) as a global tail-growth exponent; these examples are locally quadratic with \(m=2\), emphasizing that \(\beta\) controls large excursions rather than the small-stepsize normalization.

3.1 Local flatness in quantile and tail-risk estimation↩︎

Quantile and tail-risk estimation problems provide a statistical interpretation of the local flatness exponent. The quantile-loss formulation goes back to Koenker and Bassett [12], while Knight [11] studied its asymptotic behavior under general local conditions on the distribution function, including the higher-order crossings considered below.

3.1.0.1 Background.

Let \(Y\) be a scalar response, let \(F\) denote its distribution function, and let \(t_\star\) be a \(\tau\)-quantile, so that \(F(t_\star)=\tau\). The population quantile objective is \[H_\tau(t):=\mathbb{E}\rho_\tau(Y-t), \qquad \rho_\tau(u):=u(\tau-\mathbf{1}_{\{u<0\}}).\] At continuity points of \(F\), the population gradient is \(F(t)-\tau\). Consequently, \[t_\star\in\operatorname*{argmin}_{t\in\mathbb{R}}H_\tau(t).\] The local objective behavior is exactly the local crossing behavior of \(F\) at the target quantile. Suppose that the following higher-order crossing condition holds: for some \(m\ge2\) and constants \(c_+,c_->0\), \[F(t_\star+u)-\tau =c_+u_+^{m-1}-c_-u_-^{m-1}+o(|u|^{m-1}), \qquad u\to0, \label{eq:weak-quantile-crossing-v3}\tag{20}\] where \(u_+=\max\{u,0\}\) and \(u_-=\max\{-u,0\}\). Integrating the population gradient gives \[H_\tau(t_\star+u)-H_\tau(t_\star) =\frac{c_+}{m}u_+^m+\frac{c_-}{m}u_-^m+o(|u|^m).\] The positive-density case corresponds to \(m=2\). If the density vanishes at the target quantile like \(|u|^{m-2}\), then the equation \(F(t)=\tau\) still has the unique solution \(t=t_\star\), but \(F(t)-\tau\) vanishes at order \(m-1\) near \(t_\star\). As a result, the population risk is \(m\)-flat. This setup is closely related to the regularly varying local regime studied by Knight [11].

The Rockafellar–Uryasev variational representation of CVaR [13] gives an analogous population objective. For a scalar loss \(L\) and confidence level \(\tau\in(0,1)\), this representation is \[\operatorname{CVaR}_\tau(L) =\min_{t\in\mathbb{R}}C_\tau(t), \qquad C_\tau(t):=t+\frac{1}{1-\tau}\mathbb{E}(L-t)_+.\] At continuity points of \(F_L\), \[C_\tau'(t) =1-\frac{1}{1-\tau}\mathbb{P}(L>t) =\frac{F_L(t)-\tau}{1-\tau}.\] Hence, if \(t_\star\) is the unique \(\tau\)-quantile of \(L\), then \[\operatorname*{argmin}_{t\in\mathbb{R}}C_\tau(t) =\{t_\star\}.\] The local expansion in 20 then implies \[C_\tau(t_\star+u)-C_\tau(t_\star) =\frac{c_+}{m(1-\tau)}u_+^m +\frac{c_-}{m(1-\tau)}u_-^m+o(|u|^m).\] Thus, combining the variational representation with the local expansion in 20 shows that the quantile objective and the variational objective for CVaR have the same local exponent \(m\).

3.1.0.2 Invariant law for median estimation.

To illustrate the local effect of this flatness, consider median estimation with distribution function \[F_m(x)=\frac{1}{2}+\frac{1}{2}\operatorname{sgn}(x)|x|^{m-1}, \qquad |x|\le1. \label{eq:weak-median-family-v3}\tag{21}\] Thus, \(\tau=1/2\) and the median is zero. The stochastic gradient at the minimizer has nonzero variance, while the restoring force is of order \(|x|^{m-1}\). Their balance gives the scale \(\alpha^{1/m}\) and a nonlinear diffusion when \(m>2\). The raw pinball score is nonsmooth, so the finite-state calculation below illustrates this mechanism rather than directly applying Theorem 7.

The finite-state example in Figure 1 uses a two-sample lazy version of the recursion \[X_{n+1}=X_n- \alpha\left\{\frac{1}{2}\bigl(\mathbf{1}_{\{Y_{n+1,1}\le X_n\}} +\mathbf{1}_{\{Y_{n+1,2}\le X_n\}}\bigr)-\frac{1}{2}\right\}. \label{eq:lazy-weak-median-v3}\tag{22}\] For \(\alpha=2^{-j}\), the chain lives on the grid \(\{-1,-1+\alpha/2,\ldots,1\}\). If \(x_k=-1+k\alpha/2\), then its right and left transition probabilities are \[p_k=(1-F_m(x_k))^2, \qquad q_k=F_m(x_k)^2,\] and its invariant law is computed exactly from the detailed balance equation \[\frac{\pi_\alpha(k+1)}{\pi_\alpha(k)} =\frac{p_k}{q_{k+1}}.\] The variance of the noise at the minimizer in 22 is \(\Sigma=1/8\), and the limiting diffusion is \[dZ_t=-\frac{1}{2} Z_t|Z_t|^{m-2}\,dt+\sqrt{\Sigma}\,dB_t.\] Let \(Z_\infty\) have the invariant law of this diffusion. Its density is \[p_m(z)=\frac{m(1/(m\Sigma))^{1/m}}{2\Gamma(1/m)} \exp\!\left(-\frac{|z|^m}{m\Sigma}\right).\] In particular, for \(m=4\), \(p_4(z)\propto \exp(-2z^4)\), which is not Gaussian. Figure 1 confirms all these theoretical predictions.

Figure 1: SGD for median estimation. Panel (a) shows the L_2 error for several crossing orders m; the fitted slopes match the theoretical exponent 1/m. Panel (b) checks the predicted local moment scale: in this model \alpha^{-1}\mathbb{E}|X_\infty^{(\alpha)}|^m\to\Sigma=1/8. Panel (c) shows the convergence of \alpha^{-1/4}X_\infty^{(\alpha)} for m=4 to the quartic invariant density of the limiting diffusion, together with a Gaussian of the same variance for comparison.

3.1.0.3 Markovian covariance and unequal exponents.

A related finite-state experiment in Figure 2 (a) verifies the effect of temporal dependence on the asymptotic covariance. Let \(\rho\in[0,1)\), and let \((S_n)\) be a stationary two-state Markov chain on \(\{-1,1\}\) with transition probabilities \[\mathbb{P}(S_{n+1}=s\mid S_n=s)=\frac{1+\rho}{2}, \qquad \mathbb{P}(S_{n+1}=-s\mid S_n=s)=\frac{1-\rho}{2}, \qquad s\in\{-1,1\}.\] Its stationary distribution is uniform on \(\{-1,1\}\), and \(\mathbb{E}[S_nS_{n+k}]=\rho^k\). Independently, let \((R_n)\) be i.i.d.with \[\mathbb{P}(R_n\le r)=r^{m-1},\quad 0\le r\le1,\] and set \(Y_n=S_nR_n\). The two-state chain can be realized on \([-1,1]\) by taking \(\Phi(s,U)=s\) with probability \(\rho\), and \(\Phi(s,U)=1\) or \(-1\), each with probability \((1-\rho)/2\). Hence \(\mathbb{E}L_\Phi(U)=\rho<1\), where \(L_\Phi(U)\) is the random Lipschitz coefficient from Assumption 2([[itm:N1]](#itm:N1){reference-type=“ref” reference=“itm:N1”}). The marginal law of \(Y_n\) is still 21 , but at the median \(\mathbf{1}_{\{Y_n\le0\}}-\frac{1}{2}=-\frac{1}{2}S_n.\) Thus the asymptotic covariance in the Green–Kubo representation 11 is \[\Sigma =\frac{1}{4}\left(1+2\sum_{k\ge1}\rho^k\right) =\frac{1}{4} \cdot \frac{1+\rho}{1-\rho}.\] For \(m=4\), the invariant law satisfies \(\mathbb{E}Z_\infty^4=\Sigma\), and hence \[\alpha^{-1}\mathbb{E}|X_\infty^{(\alpha)}|^4 \longrightarrow \Sigma.\] Figure 2 (a) shows that the rescaled moment approaches the predicted value \(\Sigma = \Sigma_{\rm GK}\) as \(\alpha \downarrow 0\).

Moreover, we consider the separable case by taking the product of two scalar recursions of the form 22 , with \(m_1=2\) and \(m_2=4\). Under the common scale associated with \(m_2\), only the second coordinate has a nonzero limit. Figure 2 (b) confirms this conclusion.

Figure 2: Two consequences of the scaling limit of invariant laws. Panel (a) uses a Markovian sign stream and verifies that the fourth moment constant in the m=4 quantile example is governed by the asymptotic covariance rather than by the marginal variance. Panel (b) considers a separable two-coordinate model with exponents m_1=2 and m_2=4; under the common \alpha^{-1/4} scale, the m_1=2 coordinate converges to zero while the m_2=4 coordinate converges to a nondegenerate limit.

3.2 Subquadratic tails for robust and logistic losses↩︎

The tail exponent \(\beta\) describes the growth of the objective at infinity. The generalized Charbonnier family below realizes every \(\beta\in[1,2]\) [15], [16]; the classical Huber loss is a related piecewise-defined robust loss with linear tails [14]. Under the bounded, nondegenerate, nonseparable model verified below, the population logistic risk has linear coercive growth corresponding to \(\beta=1\) [17].

3.2.0.1 Robust location estimation.

Consider a scalar location model. For \(1\le\beta\le2\), define the generalized Charbonnier loss \[\rho_\beta(u):=\frac{(1+u^2)^{\beta/2}-1}{\beta}, \qquad \psi_\beta(u):=\rho_\beta'(u) =u(1+u^2)^{\beta/2-1}.\] The endpoint \(\beta=1\) is the pseudo-Huber/Charbonnier loss, while \(\beta=2\) is ordinary least squares. In the scalar location model \[Y=\theta_\star+\varepsilon, \qquad H_\beta(\theta):=\mathbb{E}\rho_\beta(\theta-Y),\] recenter \(x=\theta-\theta_\star\). If the error distribution is symmetric, then the constant-stepsize SGD becomes \[X_{n+1}=X_n-\alpha\psi_\beta(X_n-\varepsilon_{n+1}).\] The population drift is \[h_\beta(x)=\mathbb{E}\psi_\beta(x-\varepsilon).\] Since \[\psi_\beta'(u) = (1+u^2)^{\beta/2-2}\{1+(\beta-1)u^2\},\] we have \[h_\beta'(0)=\mathbb{E}\psi_\beta'(-\varepsilon)>0\] for every nondegenerate bounded symmetric error. Thus the minimizer is locally quadratic with \(m=2\). Consequently the invariant law remains the classical Gaussian limit, \[\alpha^{-1/2}X_\infty^{(\alpha)} \Rightarrow N\left(0,\frac{\mathop{\mathrm{Var}}(\psi_\beta(-\varepsilon))}{2h_\beta'(0)}\right).\]

The difference from least squares is not local but global. As \(|u|\to\infty\), \[\rho_\beta(u)\sim \frac{|u|^\beta}{\beta}, \qquad \psi_\beta(u)\sim \operatorname{sgn}(u)|u|^{\beta-1}.\] Therefore, for bounded errors, \[H_\beta(x)\sim\frac{|x|^\beta}{\beta}, \qquad h_\beta(x)\sim \operatorname{sgn}(x)|x|^{\beta-1}, \qquad |x|\to\infty.\] For \(1\le\beta<2\), the behavior at infinity of the deterministic flow \(\dot{x}_t=-h_\beta(x_t)\) is described by \[\frac{d}{dt}|x_t|^{2-\beta} \sim -(2-\beta), \qquad |x_t|\to\infty. \label{eq:robust-tail-coordinate-v3}\tag{23}\] This calculation explains why the weight function for subquadratic tails involves \(|x|^{2-\beta}\).

For the numerical illustration, take \[\varepsilon\sim0.9\,\operatorname{Unif}[-1,1] +0.1\,\operatorname{Unif}[-10,10].\] This bounded symmetric mixture satisfies the assumptions used above. Figure 3 reports the corresponding numerical results.

Figure 3: Robust location estimation with generalized Charbonnier losses. Panel (a) verifies the population tail H_\beta(x)\asymp |x|^\beta. Panel (b) verifies the Gaussian scaling limit of X_\infty^{(\alpha)}/\sqrt\alpha for the pseudo-Huber loss with \beta=1 because it remains locally quadratic. Panel (c) shows that the decay of the normalized state \mathbb{E} X_t/x_0, starting from a large x_0>0, depends on the tail exponent. Panel (d) normalizes the same trajectories by the tail scaling |x|^{2-\beta}, showing that they follow the linear prediction in 23 .

3.2.0.2 Logistic regression.

Finally, we verify theoretically that our results apply to SGD for logistic regression under a correctly specified model with bounded, nondegenerate design and nonvanishing label noise. Let \((Y_n,Z_n)_{n\ge1}\) be i.i.d.data with covariates satisfying \[\left\lVert Z_n \right\rVert \le R, \qquad \mathbb{E}[Z_n Z_n^\top]\succeq\lambda I_d\] for some \(R,\lambda>0\). With \(\sigma(t):=(1+e^{-t})^{-1}\), suppose that \[\mathbb{P}(Y_n=1\mid Z_n)=\sigma(Z_n^\top\theta_\star), \qquad Y_n\in\{-1,1\}.\] Set \(S_n:=Y_nZ_n\). For the logistic loss \[\ell(\theta;S):=\log(1+e^{-S^\top\theta}),\] the unique population minimizer is \(\theta_\star\). After recentering \(x=\theta-\theta_\star\), the sample and population gradients are \[\mathsf G(x,S) :=-\frac{S}{1+e^{S^\top(\theta_\star+x)}}, \qquad h(x):=\mathbb{E}\mathsf G(x,S).\] The centered population objective is \(C^2\) and convex, with \(h(0)=0\), and hence satisfies Assumption 1([[itm:H1]](#itm:H1){reference-type=“ref” reference=“itm:H1”}); moreover, \(\left\lVert \mathsf G(x,S) \right\rVert\le R\).

Since \(\lvert Z^\top\theta_\star\rvert\le R\left\lVert \theta_\star \right\rVert\), setting \(q:=\sigma(-R\left\lVert \theta_\star \right\rVert)\) gives \(q\le \mathbb{P}(Y=1\mid Z)\le 1-q .\) Hence, for every \(v\in\mathbb{S}^{d-1}\), \[\gamma(v):=\mathbb{E}[(-S^\top v)_+] \ge q\mathbb{E}|Z^\top v| \ge \frac{q}{R}\mathbb{E}[(Z^\top v)^2] \ge \frac{q\lambda}{R}.\] Moreover, writing \(B:=R\left\lVert \theta_\star \right\rVert\), one has the uniform bound \[\sup_{v\in\mathbb{S}^{d-1}} \left|\frac{1}{r}\left\langle rv,h(rv)\right\rangle-\gamma(v)\right| \le \frac{e^B}{er}.\] Thus, for all sufficiently large \(r\), \[\left\langle rv,h(rv)\right\rangle\ge\frac{q\lambda}{2R}r, \qquad \left\lVert h(rv) \right\rVert\le R,\] which verifies Assumption 1([[itm:H3b]](#itm:H3b){reference-type=“ref” reference=“itm:H3b”}) with \(\beta=1\). The above condition and the nondegenerate design also imply, by a standard covering and concentration argument, that an i.i.d.sample is not linearly separable with probability tending to one as its size grows, for fixed \(d\).

The population Hessian at \(\theta_\star\) satisfies \[A:=\mathbb{E}\!\left[ \sigma(Z^\top\theta_\star){1-\sigma(Z^\top\theta_\star)}ZZ^\top \right] \succeq q(1-q)\lambda I_d.\] Continuity of the population Hessian verifies Assumption 1([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}) with \(m=2\), while \[h(x)=Ax+o(\left\lVert x \right\rVert), \qquad x\to0,\] verifies Assumption 3([[itm:H4]](#itm:H4){reference-type=“ref” reference=“itm:H4”}) with \(H_0(x)=\tfrac12x^\top A x\).

For each observation, \[\nabla_x\mathsf G(x,S) = \frac{e^{S^\top(\theta_\star+x)}}{(1+e^{S^\top(\theta_\star+x)})^2}SS^\top \preceq \frac{R^2}{4}I_d.\] Thus each sample loss is convex with \((R^2/4)\)-Lipschitz gradient, and the Baillon–Haddad theorem [36] verifies the co-coercivity part of Assumption 2([[itm:N7]](#itm:N7){reference-type=“ref” reference=“itm:N7”}) with any \(L_{\mathsf G}\ge\max\{1,R^2/4\}\). Although the displayed Hessian has rank at most one, no samplewise strict curvature is required. With \(g(x,S):=\mathsf G(x,S)-h(x)\), i.i.d.sampling gives \(\mathbb{E}g(x,S)=0\) and \(\left\lVert g(x,S) \right\rVert\le2R\). Thus the mean-perturbation part of ([[itm:N7]](#itm:N7){reference-type=“ref” reference=“itm:N7”}) holds with \(\theta=0\), and the tail estimate above verifies Assumption 2([[itm:N4]](#itm:N4){reference-type=“ref” reference=“itm:N4”}). The uniform bound on \(g\) verifies ([[itm:N5]](#itm:N5){reference-type=“ref” reference=“itm:N5”}), while centering and boundedness verify Assumption 4([[itm:N8]](#itm:N8){reference-type=“ref” reference=“itm:N8”}); ([[itm:N6]](#itm:N6){reference-type=“ref” reference=“itm:N6”}) is not imposed because \(m=2\). Finally, the i.i.d.embedding in Remark 1 verifies ([[itm:N1]](#itm:N1){reference-type=“ref” reference=“itm:N1”})([[itm:N3]](#itm:N3){reference-type=“ref” reference=“itm:N3”}). Thus Assumptions 3 and 4 hold with \(m=2\) and \(\beta=1\), so our main results apply.

4 Proofs of the main results↩︎

We prove the main results in this section. In Section 4.1, we show that for any sufficiently small constant stepsize \(\alpha\), the augmented Markov chain admits a unique invariant law and contracts geometrically in the Wasserstein distance, thereby establishing Theorem 4. In Section 4.2, with the fixed-\(\alpha\) invariant law in hand, we derive its scaling limit as \(\alpha \downarrow 0\) and prove Theorem 7. Additional technical proofs are deferred to Appendices 6, 7, and 8.

4.1 Analysis of constant-stepsize SGD↩︎

At fixed \(\alpha\), we use the framework of Qu, Blanchet, and Glynn [19], where the main task is to prove a one-step contraction for a certain weight-induced metric. More specifically, to prove Theorem 4, it suffices to construct a metric \(d_{V_\alpha,\alpha}\) induced by a weight function \(V_\alpha\) such that \[\mathbb{E}d_{V_\alpha,\alpha}(F_{U_1}(z),F_{U_1}(z')) \le (1-c_0 \alpha^{m-1})d_{V_\alpha,\alpha}(z,z') \label{eq:main-contraction-aug}\tag{24}\] for some constant \(c_0>0\) and all \(z,z'\in\mathsf Z\). Proposition 11 below then converts this into the invariant law, uniqueness, and geometric convergence.

We remark that the technique for establishing the contraction in [19] cannot be applied directly here. In short, it controls each realized random map through its local Lipschitz modulus before averaging. For the SGD maps considered here, a stochastic-gradient Jacobian may be rank deficient, so the corresponding update map can have local Lipschitz modulus one and therefore may not be a strict contraction. Instead, we retain the induced metric but prove the desired contraction in expectation via a direction kernel (cf.Lemma 2).

For a metric \(d\) on a space \(E\), we write \(W_1^d\) for the Wasserstein–1 distance induced by \(d\), namely \[W_1^d(\mu,\nu) := \inf_{\gamma\in\Pi(\mu,\nu)} \int_{E\times E} d(z,z')\,\gamma(dz,dz').\]

Proposition 11. Let \((E,d)\) be a Polish space and let \(Z_{n+1}=F_{U_{n+1}}(Z_n)\) be a random-map Markov chain with transition kernel \(P\). Suppose that there exist \(r\in(0,1)\) and \(z_\star\in E\) such that \[\label{eq:metric-one-step-contraction} \mathbb{E}d(F_U(z),F_U(z'))\le r d(z,z'), \qquad z,z'\in E,\qquad{(1)}\] and \[\label{eq:metric-reference-integrability} \mathbb{E}d(F_U(z_\star),z_\star)<\infty.\qquad{(2)}\] Then the chain admits a unique invariant law \(\pi\in\mathcal{P}_1(E,d)\). Moreover, for all probability laws \(\mu,\nu\) for which the right-hand side is finite, \[\label{eq:metric-W1-contraction} W_1^d(\mu P^n,\nu P^n) \le r^n W_1^d(\mu,\nu), \qquad n\ge0.\qquad{(3)}\] For every bounded function \(\varphi:E\to\mathbb{R}\) that is Lipschitz with respect to the metric \(d\), and every fixed \(z\in E\), one has \(P^n\varphi(z)\to\pi(\varphi)\).

The proof is a Banach fixed-point argument for the Markov operator in the metric \(W_1^d\); the details are given in Appendix 6.1.

Before applying Proposition 11 to the augmented Markovian map \(F_U\), we first record a lemma which converts Assumption 2([[itm:N7]](#itm:N7){reference-type=“ref” reference=“itm:N7”}) into samplewise nonexpansiveness and averaged contraction for the stochastic-gradient Jacobian. For \((x,\xi)\in\mathbb{R}^d\times\Xi\), write \[A(x,\xi):=\nabla_x\mathsf G(x,\xi) =\nabla^2H(x)+\nabla_xg(x,\xi).\]

Lemma 1. Under Assumptions 1([[itm:H1]](#itm:H1){reference-type=“ref” reference=“itm:H1”}) and 2([[itm:N7]](#itm:N7){reference-type=“ref” reference=“itm:N7”}), \[\label{eq:global-jacobian-bounds} \left\lVert A(x,\xi) \right\rVert_{\mathrm{op}} \le L_{\mathsf G},\qquad \left\lVert \nabla_xg(x,\xi) \right\rVert_{\mathrm{op}} \le\frac{(2-\theta)L_{\mathsf G}}{1-\theta}, \qquad (x,\xi)\in\mathbb{R}^d\times\Xi.\tag{25}\] Consequently, \(h\) is globally \(\frac{L_{\mathsf G}}{1-\theta}\)-Lipschitz, \(g(\cdot,\xi)\) is globally \(\frac{2-\theta}{1-\theta}L_{\mathsf G}\)-Lipschitz uniformly in \(\xi\), and \[\label{eq:population-cocoercivity} \left\lVert h(x) \right\rVert^2 \le\frac{L_{\mathsf G}}{1-\theta}\left\langle x,h(x)\right\rangle, \qquad x\in\mathbb{R}^d.\tag{26}\] For every \(0<\alpha\le L_{\mathsf G}^{-1}\), \[\label{eq:samplewise-nonexpansive} \left\lVert (I_d-\alpha A(x,\xi))v \right\rVert\le \left\lVert v \right\rVert, \qquad x\in\mathbb{R}^d,\;\xi\in\Xi,\;v\in\mathbb{R}^d.\tag{27}\] Furthermore, for every \(x\in\mathbb{R}^d\), \(\xi\in\Xi\), and \(v\ne0\), \[\label{eq:averaged-directional-contraction} \mathbb{E}\!\left[ \left\lVert (I_d-\alpha A(x,\Phi(\xi,U_1)))v \right\rVert \right] \le \left( 1-\frac{\alpha(1-\theta)}{2} \frac{\left\langle v,\nabla^2H(x)v\right\rangle}{\left\lVert v \right\rVert^2} \right)\left\lVert v \right\rVert.\tag{28}\]

To apply Proposition 11, the key is to establish 24 . Towards this end, we need to construct a suitable weight function \(V_\alpha\). Let \(R_H\) be as in Assumption 1([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}). For the tail region, define \[\mathcal{T}_\beta(x):= \begin{cases} 1, & \beta=2,\\[0.3em] \exp\!\left\{\kappa\bigl(\left\lVert x \right\rVert^{2-\beta}-(\tfrac{3R_H}{4})^{2-\beta}\bigr)_+\right\}, & 1\le\beta<2, \end{cases}\] where \(\kappa>0\) will be chosen sufficiently small in the subquadratic case. Near the minimizer, define \[\begin{align} \omega(x) &:= \begin{cases} 0, & m=2,\\ \left(1-\left(\dfrac{2\left\lVert x \right\rVert}{R_H}\right)^{m-2}\right)_+, & 2<m\le3,\\[0.4em] \left(1-\dfrac{2\left\lVert x \right\rVert}{R_H}\right)_+, & m>3, \end{cases} &\qquad \delta_\alpha &:= \begin{cases} 0, & m=2,\\ \kappa_0\alpha, & 2<m\le3,\\ \kappa_0\alpha^{m-2}, & m>3, \end{cases} \end{align}\] where the constant \(\kappa_0>0\) will be chosen in the technical estimates below. Finally set \[V_\alpha(x,\xi):=V_\alpha(x):=\mathcal{T}_\beta(x)+\delta_\alpha\omega(x).\] Define \(d_{V_\alpha}^{X}\) to be the metric from Definition 1 with the Euclidean base norm and weight \(V_\alpha\).

Regarding the two summands in \(V_\alpha\), the term \(\mathcal{T}_\beta\) controls the tail region: it is constant in the quadratic-tail case \(\beta=2\), and is exponential in \(\left\lVert x \right\rVert^{2-\beta}\) when \(1\le\beta<2\). The compactly supported function \(\omega\) is used when \(m>2\): it provides contraction in a neighborhood of the minimizer. In the flat cases \(m>2\), the coefficient \(\delta_\alpha\) is the amplitude of this near-minimizer correction and is chosen so that its decrease has order \(\alpha^{m-1}\).

For fixed \(\xi\in\Xi\), write \[\xi^+:=\Phi(\xi,U_1), \qquad f_{\xi^+}(x):=x-\alpha\{h(x)+g(x,\xi^+)\}.\] For \(e\in\mathbb{S}^{d-1}\), define the directional contraction factor \[\mathcal{D}_{\xi,\alpha}(x,e) := \left\lVert (I_d-\alpha A(x,\Phi(\xi,U_1)))e \right\rVert.\] For a nonnegative function \(\phi(x)\), define the directional kernel \[(\mathcal{K}_{\xi,\alpha}\phi)(x;e) := \mathbb{E}\!\left[ \mathcal{D}_{\xi,\alpha}(x,e)\, \phi\bigl(f_{\Phi(\xi,U_1)}(x)\bigr) \right], \qquad e\in\mathbb{S}^{d-1}.\]

The next lemma is the main estimate that shows the contraction of the weight \(V_\alpha\) under the kernel \(\mathcal{K}_{\xi,\alpha}\).

Lemma 2. Assume Assumptions 1 and 2. There exist constants \[c_0\in(0,1),\qquad C_V > 0,\qquad \alpha_0\in(0,1],\] such that, for every \(0<\alpha\le\alpha_0\), \(\xi\in\Xi\), and \(x\in\mathbb{R}^d\), \[\label{eq:conditional-directional-drift-main} \sup_{e\in\mathbb{S}^{d-1}} (\mathcal{K}_{\xi,\alpha}V_\alpha)(x;e) \le (1-c_0\alpha^{m-1})V_\alpha(x),\tag{29}\] and \[\label{eq:conditional-growth-main} \mathbb{E}\!\left[ V_\alpha\bigl(f_{\Phi(\xi,U_1)}(x)\bigr) \right] \le (1+C_V\alpha)V_\alpha(x).\tag{30}\]

The next lemma lifts the directional contraction to the induced metric on the augmented space, proving 24 .

Lemma 3. Assume Assumptions 1 and 2, and let \(c_0\) be as in Lemma 2. There exists \(\alpha_0\in(0,1]\) such that 24 holds for every \(0<\alpha\le\alpha_0\).

The final estimate verifies the reference-point integrability required by Proposition 11.

Lemma 4. There exists \(\alpha_0\in(0,1]\) such that, for every \(0<\alpha\le\alpha_0\), with \(z_\star=(0,\xi_\star)\) and \(\xi_\star\) as in Assumption 2([[itm:N3]](#itm:N3){reference-type=“ref” reference=“itm:N3”}), \[\label{eq:fixed-reference-integrability} \mathbb{E}\,d_{V_\alpha,\alpha}\bigl(F_{U_1}(z_\star),z_\star\bigr)<\infty.\tag{31}\]

The proofs of Lemmas 1, 2, 3, and 4 are given in Appendices 6.2, 6.4, 6.5, and 6.6, respectively.

Proof of Theorem 4. We first check that \((\mathsf Z,d_{V_\alpha,\alpha})\) is Polish. Since \(\mathsf Z\) is closed in a finite-dimensional Euclidean space, \(V_\alpha\) is continuous, and \(V_\alpha\ge1\), the metric \(d_{V_\alpha,\alpha}\) dominates the augmented base norm 3 . Conversely, on every set \(K_R:=\{z\in\mathsf Z:\vert z\vert_\alpha\le R\},\) continuity of \(V_\alpha\) gives \(\sup_{K_R}V_\alpha<\infty\), so \(d_{V_\alpha,\alpha}\) and the augmented base metric are locally comparable on \(K_R\). Hence \((\mathsf Z,d_{V_\alpha,\alpha})\) is Polish by elementary general topology.

Let \(c_0\) be as in Lemma 2, and choose \(\alpha_0 \in (0, L_{\mathsf G}^{-1}]\) small enough that the conclusions of Lemmas 23, and 4 hold. Fix \(0<\alpha\le\alpha_0\). Equations 24 and 31 verify respectively the two conditions ?? and ?? of Proposition 11, with \((E,d)=(\mathsf Z,d_{V_\alpha,\alpha})\), \(r=1-c_0\alpha^{m-1}\), and \(z_\star\) as in Lemma 4. Proposition 11 therefore yields a unique invariant law \(\pi_\alpha\in\mathcal{P}_1(\mathsf Z,d_{V_\alpha,\alpha})\) and gives 8 with \(c = c_0\). ◻

4.2 Analysis of scaling limit↩︎

Next, we turn to the scaling limit of the invariant law \(\pi_\alpha\) given by Theorem 4 as \(\alpha \downarrow 0\). For every sufficiently small \(\alpha>0\), let \[(X_\alpha,\xi_\alpha)\sim\pi_\alpha, \qquad \xi_\alpha^+:=\Phi(\xi_\alpha,U_1), \qquad X_\alpha^+:=X_\alpha-\alpha\{h(X_\alpha)+g(X_\alpha,\xi_\alpha^+)\},\] where \(U_1\) is independent of \((X_\alpha,\xi_\alpha)\). Set \[Y_\alpha:=\alpha^{-1/m}X_\alpha, \qquad Y_\alpha^+:=\alpha^{-1/m}X_\alpha^+, \qquad h_\alpha(y):=\alpha^{-(1-1/m)}h(\alpha^{1/m}y).\] By stationarity of \(\pi_\alpha\), \[(Y_\alpha^+,\xi_\alpha^+)\stackrel d=(Y_\alpha,\xi_\alpha),\] and \[Y_\alpha^+-Y_\alpha = -\alpha^{2-2/m}h_\alpha(Y_\alpha) -\alpha^{1-1/m}g(\alpha^{1/m}Y_\alpha,\xi_\alpha^+).\] Thus the drift and noise have orders \(\alpha^{2-2/m}\) and \(\alpha^{1-1/m}\), respectively (recovering the familiar \(\alpha\) and \(\sqrt{\alpha}\) scalings for \(m=2\) in particular). We prove Theorem 7 by showing tightness of \(Y_\alpha\), identifying every weak subsequential limit through the stationary generator equation, and using uniqueness of the invariant law of the limiting diffusion.

The following key lemma gives the Poisson decomposition of the Markovian noise and identifies the covariance matrix \(\Sigma\). Let \(Q\) be the transition kernel of the driving chain \((\xi_n)_{n \ge 0}\), \[Q\varphi(\xi):=\mathbb{E}[\varphi(\Phi(\xi,U_1))].\]

Lemma 5. Under Assumptions 1 and 4, for \(x\in\mathbb{R}^d\), write \(g_x(\xi):=g(x,\xi)\). Then the series \[\chi_x(\xi):=\sum_{k=1}^{\infty}Q^k g_x(\xi), \qquad x\in\mathbb{R}^d, \;\xi\in\Xi,\] converges absolutely and defines a centered solution of the Poisson equation \[\label{eq:chi-poisson} \chi_x-Q\chi_x=Qg_x.\tag{32}\] Define \[D_x(\xi,u):= g_x(\Phi(\xi,u))+\chi_x(\Phi(\xi,u))-\chi_x(\xi).\] Then \[\label{eq:Dx-martingale-difference} \mathbb{E}[D_x(\xi,U_1)\mid \xi]=0,\tag{33}\] and, with \(\xi^+=\Phi(\xi,u)\), \[\label{eq:martingale-poisson-decomposition} g_x(\xi^+)=D_x(\xi,u)+\chi_x(\xi)-\chi_x(\xi^+).\tag{34}\] The series in 11 converges absolutely, and the resulting matrix satisfies \[\Sigma = \mathbb{E}_{\xi\sim\pi_\Xi}\mathbb{E}\!\bigl[D_0(\xi,U_1)D_0(\xi,U_1)^\top\bigr] \succeq0, \qquad D_0(\xi,u):=D_x(\xi,u)\big|_{x=0}.\]

The next lemma gives moment estimates for \(X_\alpha\) which will imply tightness of \(Y_\alpha=\alpha^{-1/m}X_\alpha\).

Lemma 6. Assume Assumptions 3 and 4. There exist constants \(C,\alpha_0>0\) such that, for every \(0<\alpha\le\alpha_0\), \[\begin{align} \mathbb{E}\left\langle X_\alpha,h(X_\alpha)\right\rangle &\le C\alpha, \tag{35}\\ \mathbb{E}\bigl[\left\lVert X_\alpha \right\rVert^m\mathbf{1}_{\{\left\lVert X_\alpha \right\rVert\le1\}}\bigr] &\le C\alpha, \tag{36}\\ \mathbb{P}(\left\lVert X_\alpha \right\rVert\ge1) &\le C\alpha, \tag{37}\\ \mathbb{E}\left\lVert h(X_\alpha)+g(X_\alpha,\xi_\alpha^+) \right\rVert^2 &\le C. \tag{38} \end{align}\]

The following lemma controls the dependence between \(Y_\alpha\) and \(\xi_\alpha\).

Lemma 7. Assume Assumptions 3 and 4. Let \(\psi:\mathbb{R}^d\to\mathbb{R}\) be bounded and globally Lipschitz, and let \(f:\Xi\to\mathbb{R}\) be bounded, Lipschitz, and centered under \(\pi_\Xi\). Then \[\left\lvert \mathbb{E}[\psi(Y_\alpha)f(\xi_\alpha)] \right\rvert \le C_{\psi,f}\alpha^{1-1/m},\] where \(C_{\psi,f}>0\) is independent of \(\alpha\). In particular, \[\mathbb{E}[\psi(Y_\alpha)f(\xi_\alpha)]\to0.\]

The next lemma identifies every weak subsequential limit of \(Y_\alpha\).

Lemma 8. Assume Assumptions 3 and 4. Let \(\alpha_k\downarrow0\) and suppose \[Y_{\alpha_k}\Rightarrow \nu.\] Then, for every \(\varphi\in C_c^3(\mathbb{R}^d)\), \[\label{eq:stationary-generator-limit-main} \int_{\mathbb{R}^d} \left[ -\left\langle h_0(y),\nabla\varphi(y)\right\rangle + \frac{1}{2}\mathop{\mathrm{tr}}\bigl(\Sigma\nabla^2\varphi(y)\bigr) \right]\,\nu(dy) = 0.\tag{39}\]

Finally, Equation 39 identifies a unique limiting law.

Lemma 9. Assume Assumption 3, and let \(\Sigma\succeq0\). The (possibly degenerate) diffusion \[dY_t=-h_0(Y_t)\,dt+\Sigma^{1/2}\,dB_t\] has a unique invariant law, denoted by \(\nu_\infty\). Moreover, any probability law \(\nu\) satisfying 39 for every \(\varphi\in C_c^3(\mathbb{R}^d)\) equals \(\nu_\infty\).

The proofs of Lemmas 5, 6, 7, 8, and 9 are given in Appendices 7.3, 7.4, 7.5, 7.6, and 7.7, respectively.

Proof of Theorem 7. Fix \(R\ge1\). For all sufficiently small \(\alpha\), we have \(R\alpha^{1/m}\le1\), and Lemma 6 gives \[\begin{align} \mathbb{P}(\left\lVert Y_\alpha \right\rVert>R) &= \mathbb{P}(\left\lVert X_\alpha \right\rVert>R\alpha^{1/m})\\ &\le \mathbb{P}(R\alpha^{1/m}<\left\lVert X_\alpha \right\rVert\le1) +\mathbb{P}(\left\lVert X_\alpha \right\rVert>1)\\ &\le \frac{ \mathbb{E}[\left\lVert X_\alpha \right\rVert^m \mathbf{1}_{\{\left\lVert X_\alpha \right\rVert\le1\}}] }{R^m\alpha} +\mathbb{P}(\left\lVert X_\alpha \right\rVert>1)\\ &\le CR^{-m}+C\alpha, \end{align}\] where the last line uses 36 and 37 . Hence \[\limsup_{\alpha\downarrow0} \mathbb{P}(\left\lVert Y_\alpha \right\rVert>R) \le CR^{-m}.\] Letting \(R\to\infty\) shows that the family \(\{Y_\alpha\}_{\alpha>0}\) is tight as \(\alpha\downarrow0\).

We next identify all possible subsequential limits. Let \((\alpha_k)_{k\ge1}\) be an arbitrary sequence satisfying \(\alpha_k\downarrow0\). By tightness, there exist integers \(1\le k_1<k_2<\cdots\) and a probability law \(\nu\) on \(\mathbb{R}^d\) such that \(Y_{\alpha_{k_j}}\Rightarrow\nu\) as \(j\to\infty\). Since \(\alpha_{k_j}\downarrow0\), we may apply Lemma 8 to the sequence \((\alpha_{k_j})_{j\ge1}\). It follows that \(\nu\) satisfies the stationary generator equation 39 . Therefore, Lemma 9 implies that \(\nu=\nu_\infty,\) where \(\nu_\infty\) is the unique invariant law of the limiting diffusion 12 . Thus, \(Y_{\alpha_{k_j}}\Rightarrow\nu_\infty.\)

The same argument applies to every subsequence of the original sequence \((\alpha_k)_{k\ge1}\): every such subsequence has a further subsequence along which the corresponding random variables converge weakly to \(\nu_\infty\). As a result, the entire sequence satisfies \(Y_{\alpha_k}\Rightarrow\nu_\infty\) as \(k\to\infty\). Because the sequence \((\alpha_k)_{k\ge1}\) was arbitrary, we conclude that \(Y_\alpha=\alpha^{-1/m}X_\alpha \Rightarrow Y_\infty\) as \(\alpha\downarrow0\) and \(Y_\infty\sim\nu_\infty\). ◻

Remark 12. The generator argument above identifies weak subsequential limits rather than solving a Stein equation for the limiting invariant law. This is particularly useful when \(m>2\), since the limiting generator has nonlinear \(m\)-flat drift, a generally non-Gaussian invariant law, and possibly degenerate covariance. Obtaining uniform estimates for solutions of the associated Stein equations that are needed for quantitative convergence rates would require additional analysis of Poisson equations for diffusions with nonlinear drift.

5 Conclusion↩︎

This paper develops an invariant-law theory for constant-stepsize SGD beyond the strongly convex setting. For each sufficiently small constant stepsize, we prove geometric convergence to a unique invariant law in a Wasserstein distance induced by a metric adapted to two geometric features of the objective: an \(m\)-flat near-minimizer region and a \(\beta\)-subquadratic tail. We then identify the small-stepsize scaling limit of the invariant law as the solution of 12 . In particular, the classical Gaussian limit is recovered when \(m=2\), while flatter objectives can produce larger and generally non-Gaussian stationary fluctuations. Moreover, Markovian dependence in gradient noise enters through the asymptotic covariance \(\Sigma\).

The main conceptual point is that local and global features play different roles. The exponent \(m\) determines the scale of the invariant law and the scaling limit, whereas \(\beta\) determines the weight function needed to control excursions. This separation is reflected in the examples: quantile and tail-risk estimation illustrate nonquadratic local flatness, while smooth robust and logistic losses illustrate subquadratic tail growth.

Several directions remain open. First, the separable extension developed here allows different coordinate flatness exponents, but the genuinely nonseparable anisotropic case remains to be understood. If different coordinate groups have natural scales \(s_i(\alpha)=\alpha^{1/m_i}\), the anisotropically rescaled generator contains several effective time scales; a natural approach is multi-time-scale stochastic averaging [4], [37][39]. Second, quantitative convergence rates for the scaling limit of invariant laws may be accessible through Stein’s method for the limiting diffusion, based on the solution of the Poisson equation [40][44]. Other important extensions include replacing the one-step contraction in geometric ergodicity by a multi-step contraction, which would cover Markovian data streams for which contraction only emerges over several transitions, and relaxing the globally contractive driving-chain dynamics toward Harris-type or mixing-based Markovian noise [45], [46].

Acknowledgments↩︎

Cheng Mao was supported in part by NSF CAREER Award 2338062. Debankur Mukherjee was partially supported by the NSF grant CPS-2240982.

6 Constant-stepsize contraction↩︎

This appendix proves the metric fixed-point principle and the estimates used in Theorem 4, followed by the projected Wasserstein–1 corollary. Constants denoted by \(C,c\) may change from line to line but are independent of \(\alpha\), \(x\), and the current value \(\xi\) of the driving chain. We decrease the small-stepsize threshold when needed.

6.1 Fixed-point principle↩︎

Proof of Proposition 11. The assumptions imply that \(P\) maps \(\mathcal{P}_1(E,d)\) into itself, since \(\mathbb{E}d(F_U(z),z_\star)\le r d(z,z_\star)+\mathbb{E}d(F_U(z_\star),z_\star)\). The one-step estimate ?? iterates under the synchronous coupling and gives ?? . Let \(\mu_0=\delta_{z_\star}\) and \(\mu_n=\mu_0P^n\). By ?? , \(W_1^d(\mu_1,\mu_0)<\infty\), and hence \(W_1^d(\mu_{k+1},\mu_k)\le r^kW_1^d(\mu_1,\mu_0)\). Thus \((\mu_n)\) is Cauchy in \((\mathcal{P}_1(E,d),W_1^d)\). Let its limit be \(\pi\). Applying the contraction to \(\mu_n\) and \(\pi\) shows \(\mu_{n+1}\to\pi P\), while \(\mu_{n+1}\to\pi\), so \(\pi P=\pi\).

For fixed \(z\in E\), taking \(\mu=\delta_z\) and \(\nu=\pi\) in ?? gives \(W_1^d(\delta_zP^n,\pi)\le r^nW_1^d(\delta_z,\pi)\to0\); the right-hand side is finite because \(\pi\in\mathcal{P}_1(E,d)\). Consequently, every bounded \(d\)-Lipschitz \(\varphi\) satisfies \(|P^n\varphi(z)-\pi(\varphi)|\le \mathop{\mathrm{Lip}}_d(\varphi)W_1^d(\delta_zP^n,\pi)\to0\).

Let \(\widetilde{\pi}\) be any Borel invariant probability law. By the pointwise convergence, invariance, and dominated convergence, \(\widetilde{\pi}(\varphi)=\lim_n\widetilde{\pi}(P^n\varphi)=\pi(\varphi)\). Bounded Lipschitz functions determine Borel probability measures on the Polish space \((E,d)\), so \(\widetilde{\pi}=\pi\). ◻

6.2 Proof of Lemma 1↩︎

Proof of Lemma 1. Put \(A_{\rm s}(x,\xi):=(A(x,\xi)+A(x,\xi)^\top)/2\). Fix \(\xi\in\Xi\). Applying Cauchy–Schwarz to 6 gives \(\left\lVert \mathsf G(x,\xi)-\mathsf G(y,\xi) \right\rVert \le L_{\mathsf G}\left\lVert x-y \right\rVert\) for \(x,y\in\mathbb{R}^d\); the conclusion is immediate when the left-hand difference vanishes and otherwise follows by division by its norm. Since \(A(x,\xi)=\nabla_x\mathsf G(x,\xi)\), this proves the first bound in 25 .

Next fix \(x,v\in\mathbb{R}^d\). Apply 6 to \(x+tv\) and \(x\), divide by \(t^2\), and let \(t\to0\). Then \[\left\lVert A(x,\xi)v \right\rVert^2 \le L_{\mathsf G}\left\langle v,A(x,\xi)v\right\rangle =L_{\mathsf G}\left\langle v,A_{\rm s}(x,\xi)v\right\rangle, \qquad 0\preceq A_{\rm s}(x,\xi)\preceq L_{\mathsf G}I_d.\] Here the first inequality implies \(A_{\rm s}(x,\xi)\succeq0\), and the second also uses \(\left\lVert A_{\rm s} \right\rVert_{\mathrm{op}}\le\left\lVert A \right\rVert_{\mathrm{op}}\).

Since \(\mathsf G=h+g\), 7 implies \[\left\langle x-y,\mathbb{E}[\mathsf G(x,\Phi(\xi,U_1))- \mathsf G(y,\Phi(\xi,U_1))]\right\rangle \ge(1-\theta)\left\langle x-y,h(x)-h(y)\right\rangle.\] Apply this inequality to \(x+tv\) and \(x\), divide by \(t^2\), and let \(t\to0\). For \(t\ne0\), the preceding Lipschitz bound gives \(\left\lVert [\mathsf G(x+tv,\Phi(\xi,U_1))- \mathsf G(x,\Phi(\xi,U_1))]/t \right\rVert\le L_{\mathsf G}\left\lVert v \right\rVert\), so dominated convergence yields \[\mathbb{E}[A_{\rm s}(x,\Phi(\xi,U_1))]\succeq(1-\theta)\nabla^2H(x), \qquad 0\preceq\nabla^2H(x)\preceq\frac{L_{\mathsf G}}{1-\theta}I_d.\] Since \(\nabla_xg(x,\xi)=A(x,\xi)-\nabla^2H(x)\), the triangle inequality proves the second bound in 25 . Integrating the two Jacobian bounds along line segments yields the stated global Lipschitz estimates. Finally, 26 is the Baillon–Haddad inequality for the convex \(L_{\mathsf G}/(1-\theta)\)-smooth function \(H\), applied with \(h(0)=0\).

It remains to prove the two contraction bounds. Write \(A=A(x,\xi)\) and \(A_{\rm s}=A_{\rm s}(x,\xi)\). Since \(\alpha L_{\mathsf G}\le1\), \[\begin{align} \left\lVert (I_d-\alpha A)v \right\rVert^2 &=\left\lVert v \right\rVert^2-2\alpha\left\langle v,A_{\rm s}v\right\rangle+\alpha^2\left\lVert Av \right\rVert^2\\ &\le\left\lVert v \right\rVert^2-(2\alpha-\alpha^2L_{\mathsf G})\left\langle v,A_{\rm s}v\right\rangle \le\left\lVert v \right\rVert^2-\alpha\left\langle v,A_{\rm s}v\right\rangle\le\left\lVert v \right\rVert^2. \end{align}\] This proves 27 . If \(v\ne0\), then \[\begin{align} 0&\le\alpha\frac{\left\langle v,A_{\rm s}v\right\rangle}{\left\lVert v \right\rVert^2} \le\alpha\left\lVert A_{\rm s} \right\rVert_{\mathrm{op}}\le\alpha\left\lVert A \right\rVert_{\mathrm{op}}\le1,\\ \left\lVert (I_d-\alpha A)v \right\rVert &\le\left(1-\frac{\alpha}{2} \frac{\left\langle v,A_{\rm s}v\right\rangle}{\left\lVert v \right\rVert^2}\right)\left\lVert v \right\rVert, \end{align}\] where the second inequality uses \(\sqrt{1-t}\le1-t/2\). Replacing \(\xi\) by \(\Phi(\xi,U_1)\), taking expectation with \(v\) fixed, and using the averaged matrix inequality above proves 28 . ◻

6.3 Preparatory estimates↩︎

Lemma 10. Under Assumption 1, there exists \(C_h>0\) such that \(\left\lVert h(x) \right\rVert\le C_h\left\lVert x \right\rVert^{m-1}\) whenever \(\left\lVert x \right\rVert\le R_H/2\).

Proof. This is the second inequality in Assumption 1([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}). ◻

Lemma 11. Assume \(m>2\), and define \[\begin{align} G_\xi&:=\left\lVert g(0,\Phi(\xi,U_1)) \right\rVert,& \bar a_\alpha&:=\inf_{\xi\in\Xi}\mathbb{E}\min\{\alpha G_\xi,1\},\\ \bar a_{\alpha,p} &:=\inf_{\xi\in\Xi}\mathbb{E}\bigl[\min\{(\alpha G_\xi)^p,1\}\bigr],& &p\in(0,1]. \end{align}\] Under Assumption 2([[itm:N6]](#itm:N6){reference-type=“ref” reference=“itm:N6”}), there exist constants \(c_1,c_2>0\) and \(\alpha_g\in(0,1]\) such that, for \(0<\alpha\le\alpha_g\), \[\begin{align} \bar a_\alpha&\ge c_1\alpha, \tag{40}\\ \bar a_{\alpha,p}&\ge c_2\bar a_\alpha^p. \tag{41} \end{align}\] Moreover, under Assumption 2([[itm:N3]](#itm:N3){reference-type=“ref” reference=“itm:N3”}) if \(\beta=2\), and under ([[itm:N5]](#itm:N5){reference-type=“ref” reference=“itm:N5”}) if \(\beta<2\), there exists \(C_g>0\) such that \[\label{eq:fixed-bara-upper} \bar a_\alpha\le C_g\alpha.\tag{42}\] Consequently \(\bar a_\alpha=\Theta(\alpha)\).

Proof. Assumption 2([[itm:N6]](#itm:N6){reference-type=“ref” reference=“itm:N6”}) gives \(\inf_{\xi\in\Xi}\mathbb{P}(G_\xi\ge\varepsilon_g)\ge p_g\). For \(0<\alpha\le\varepsilon_g^{-1}\), \(\min\{\alpha G_\xi,1\}\ge \alpha\varepsilon_g\mathbf{1}_{\{G_\xi\ge\varepsilon_g\}}\), and therefore \(\bar a_\alpha\ge\alpha\varepsilon_gp_g\), proving 40 . Similarly, \(\min\{(\alpha G_\xi)^p,1\}\ge (\alpha\varepsilon_g)^p\mathbf{1}_{\{G_\xi\ge\varepsilon_g\}}\), so \(\bar a_{\alpha,p}\ge\alpha^p\varepsilon_g^pp_g\).

It remains to compare \(\alpha^p\) with \(\bar a_\alpha^p\). Let \(\xi_\star\) be the reference point in Assumption 2([[itm:N3]](#itm:N3){reference-type=“ref” reference=“itm:N3”}). If \(\beta=2\), then \(\mathbb{E}G_{\xi_\star}<\infty\) by ([[itm:N3]](#itm:N3){reference-type=“ref” reference=“itm:N3”}); if \(\beta<2\), Assumption 2([[itm:N5]](#itm:N5){reference-type=“ref” reference=“itm:N5”}) at \(x=0\) and \(\xi=\xi_\star\) gives the same conclusion. Thus \(\bar a_\alpha\le\mathbb{E}\min(\alpha G_{\xi_\star},1) \le\alpha\mathbb{E}G_{\xi_\star}\). Combining this with the previous lower bound on \(\bar a_{\alpha,p}\) proves 41 ; the same inequality gives 42 . Together with 40 , this proves \(\bar a_\alpha=\Theta(\alpha)\). ◻

Lemma 12. The following estimates hold.

  1. Let \(Z\) be an \(\mathbb{R}^d\)-valued random vector with \(\mathbb{E}Z=0\) and \(\mathbb{E}e^{\lambda\left\lVert Z \right\rVert}\le M\) for some \(\lambda>0\). Then there exist \(t_0\in(0,\lambda/2]\) and \(C>0\) such that \(\mathbb{E}e^{t\langle u,Z\rangle}\le e^{Ct^2}\) for every unit vector \(u\) and every \(|t|\le t_0\).

  2. If \(Y\ge0\) and \(\mathbb{E}e^{\lambda Y}\le M\), then \(\mathbb{E}e^{tY}\le1+(2M/\lambda)t\) for \(0\le t\le\lambda/2\).

  3. Suppose \(1\le\beta<2\). There exists \(C_\beta>0\) such that, whenever \(y\ne0\) and \(\left\lVert z \right\rVert\le\left\lVert y \right\rVert/2\), \[\left\lVert y-z \right\rVert^{2-\beta} \le \left\lVert y \right\rVert^{2-\beta}-(2-\beta)\left\lVert y \right\rVert^{1-\beta} \left\langle\frac{y}{\left\lVert y \right\rVert},z\right\rangle +C_\beta\left\lVert y \right\rVert^{-\beta}\left\lVert z \right\rVert^2.\]

Proof. For (i), set \(X=\langle u,Z\rangle\), so \(\mathbb{E}X=0\) and \(|X|\le\left\lVert Z \right\rVert\). For \(|t|\le\lambda/2\), \[\begin{align} e^{tX}&\le 1+tX+\frac{t^2\left\lVert Z \right\rVert^2}{2}e^{|t|\left\lVert Z \right\rVert},& r^2e^{(\lambda/2)r}&\le \frac{16}{\lambda^2}e^{\lambda r},\\ \mathbb{E}e^{tX}&\le 1+\frac{8M}{\lambda^2}t^2 \le \exp\!\left(\frac{8M}{\lambda^2}t^2\right). \end{align}\] For (ii), \(e^{tY}-1\le tYe^{tY}\le tYe^{(\lambda/2)Y}\) and \(re^{(\lambda/2)r}\le(2/\lambda)e^{\lambda r}\), so \(\mathbb{E}e^{tY}\le1+(2M/\lambda)t\). For (iii), apply Taylor’s theorem to \(\psi(w)=\left\lVert w \right\rVert^{2-\beta}\) on \(\mathbb{R}^d\setminus\{0\}\). Along \(y-sz\), \(s\in[0,1]\), the condition \(\left\lVert z \right\rVert\le\left\lVert y \right\rVert/2\) gives \(\left\lVert y-sz \right\rVert\ge\left\lVert y \right\rVert/2\); since \(\left\lVert \nabla^2\psi(w) \right\rVert_{\mathrm{op}}\le C_\beta\left\lVert w \right\rVert^{-\beta}\), the claimed inequality follows. ◻

6.4 Proof of Lemma 2↩︎

For fixed \(\xi\), write \[\xi^+:=\Phi(\xi,U_1), \qquad f_{\xi^+}(x)=x-\alpha\{h(x)+g(x,\xi^+)\}.\] For \(e\in\mathbb{S}^{d-1}\), set \(\mathcal{D}_{\xi,\alpha}(x,e) := \left\lVert (I_d-\alpha A(x,\xi^+))e \right\rVert.\) By 27 , \[\label{eq:fixed-directional-nonexpansive} \mathcal{D}_{\xi,\alpha}(x,e)\le1, \qquad e\in\mathbb{S}^{d-1}.\tag{43}\] Moreover, by 28 , \[\label{eq:fixed-directional-basic} \mathbb{E}\mathcal{D}_{\xi,\alpha}(x,e) \le 1-\frac{\alpha(1-\theta)}{2} \left\langle e,\nabla^2H(x)e\right\rangle, \qquad e\in\mathbb{S}^{d-1}.\tag{44}\]

We first record two weight estimates used below.

Lemma 13. There exist constants \(C,c>0\), independent of \(\alpha,x,\xi,\kappa,\kappa_0\), such that, for all sufficiently small \(\alpha\), the following hold. If \(\left\lVert x \right\rVert\le \frac{R_H}{2}\), then \[\label{eq:core-overshoot-restored} \mathbb{E}\bigl[(\mathcal{T}_\beta(f_{\xi^+}(x))-1)_+\bigr] \le C e^{-c/\alpha},\tag{45}\] with the left-hand side equal to zero when \(\beta=2\). If \(\left\lVert x \right\rVert\le R_H\), then, with the convention that \(\kappa=0\) when \(\beta=2\), \[\label{eq:compact-positive-increment} \mathbb{E}\bigl[(V_\alpha(f_{\xi^+}(x))-V_\alpha(x))_+\bigr] \le C(\kappa+\kappa_0)\alpha V_\alpha(x),\tag{46}\] and consequently \[\label{eq:compact-growth} \mathbb{E}V_\alpha(f_{\xi^+}(x)) \le \bigl(1+C(\kappa+\kappa_0)\alpha\bigr)V_\alpha(x).\tag{47}\]

Proof. The case \(\beta=2\) has \(\mathcal{T}_\beta\equiv1\), so 45 is immediate. Suppose throughout the rest of the proof that \(1\le\beta<2\). Since \(h\) is continuous, it is bounded on the compact set \(\{\left\lVert x \right\rVert\le R_H\}\). Also \(1+\left\lVert x \right\rVert^{\beta-1}\) is bounded on this compact set, so Assumption 2([[itm:N5]](#itm:N5){reference-type=“ref” reference=“itm:N5”}) implies that, for some \(\lambda_K,M_K>0\), \[\label{eq:compact-g-exp-moment} \sup_{\left\lVert x \right\rVert\le R_H}\sup_{\xi\in\Xi} \mathbb{E}\exp\{\lambda_K\left\lVert g(x,\xi^+) \right\rVert\}\le M_K.\tag{48}\]

First assume \(\left\lVert x \right\rVert\le \frac{R_H}{2}\). If \(\mathcal{T}_\beta(f_{\xi^+}(x))>1\), then \(\left\lVert f_{\xi^+}(x) \right\rVert>\frac{3R_H}{4}\). Since \(\frac{3R_H}{4}>\frac{R_H}{2}\) and \(\left\lVert h(x) \right\rVert\le C\), this implies, after decreasing the stepsize threshold, \(\alpha\left\lVert g(x,\xi^+) \right\rVert\ge c\). Moreover, \[\bigl(\left\lVert f_{\xi^+}(x) \right\rVert^{2-\beta}-(\tfrac{3R_H}{4})^{2-\beta}\bigr)_+ \le C\{1+\alpha\left\lVert g(x,\xi^+) \right\rVert\},\] because \(2-\beta\le1\). Taking \(\kappa\) small enough that \(C\kappa\alpha\le \lambda_K/2\), 48 gives \[\begin{align} \mathbb{E}\bigl[(\mathcal{T}_\beta(f_{\xi^+}(x))-1)_+\bigr] &\le C\mathbb{E}\!\bigl[ e^{C\kappa\alpha\left\lVert g(x,\xi^+) \right\rVert} \mathbf{1}_{\{\left\lVert g(x,\xi^+) \right\rVert\ge c/\alpha\}} \bigr] \\ &\le C e^{-c/\alpha}, \end{align}\] which proves 45 .

It remains to prove 46 . Fix \(\left\lVert x \right\rVert\le R_H\) and set \(G=\left\lVert g(x,\xi^+) \right\rVert\). Split according to \(B_\alpha:=\{G\le \alpha^{-1/2}\}\). On \(B_\alpha\), \(\left\lVert f_{\xi^+}(x)-x \right\rVert\le C\alpha(1+G) \le C\alpha^{1/2},\) so both \(x\) and \(f_{\xi^+}(x)\) lie in a fixed compact enlargement of \(\{\left\lVert y \right\rVert\le R_H\}\). On this enlargement, \(\mathcal{T}_\beta\) has Lipschitz constant at most \(C\kappa\), for \(0<\kappa\le1\). Hence \[\begin{align} \mathbb{E}\bigl[(\mathcal{T}_\beta(f_{\xi^+}(x))-\mathcal{T}_\beta(x))_+\mathbf{1}_{B_\alpha}\bigr] &\le C\kappa\alpha\mathbb{E}(1+G) \le C\kappa\alpha. \end{align}\] On \(B_\alpha^c\), use \(e^a-e^b\le (a-b)_+e^a\) for \(a\ge b\), the bound \(\bigl(\left\lVert f_{\xi^+}(x) \right\rVert^{2-\beta}-(\tfrac{3R_H}{4})^{2-\beta}\bigr)_+\le C\{1+(\alpha G)^{2-\beta}\}\), and 48 . Decreasing \(\alpha\) if necessary, \[\begin{align} &\mathbb{E}\bigl[(\mathcal{T}_\beta(f_{\xi^+}(x))-\mathcal{T}_\beta(x))_+\mathbf{1}_{B_\alpha^c}\bigr] \\ &\qquad\le C\kappa\mathbb{E}\!\bigl[(1+(\alpha G)^{2-\beta}) e^{C\kappa(1+(\alpha G)^{2-\beta})} \mathbf{1}_{\{G>\alpha^{-1/2}\}}\bigr] \le C\kappa e^{-c\alpha^{-1/2}} \le C\kappa\alpha. \end{align}\] Therefore \[\label{eq:compact-U-positive-increment-expanded} \mathbb{E}\bigl[(\mathcal{T}_\beta(f_{\xi^+}(x))-\mathcal{T}_\beta(x))_+\bigr] \le C\kappa\alpha.\tag{49}\] The \(\omega\)-part satisfies \(0\le\omega\le1\), and hence \[\delta_\alpha(\omega(f_{\xi^+}(x))-\omega(x))_+ \le \delta_\alpha\le C\kappa_0\alpha.\] Since \(V_\alpha(x)\ge1\), 46 follows from 49 . Finally, 47 follows from \(V_\alpha(f)\le V_\alpha(x)+(V_\alpha(f)-V_\alpha(x))_+\). ◻

Lemma 14. Assume \(1\le\beta<2\). There exist \(c>0\) and \(\alpha_{\rm tail}\in(0,1]\) such that, for all \(0<\alpha\le\alpha_{\rm tail}\), all \(\xi\in\Xi\), and all \(\left\lVert x \right\rVert\ge R_H\), \[\label{eq:far-U-beta-decrease-new} \mathbb{E}\mathcal{T}_\beta(f_{\xi^+}(x)) \le (1-c\alpha)\mathcal{T}_\beta(x).\tag{50}\]

Proof. Let \(r=\left\lVert x \right\rVert\), \(G=g(x,\xi^+)\), and \(\zeta=G-\bar g(x,\xi)\). Then \(\mathbb{E}[\zeta\mid\xi]=0\) and \(f_{\xi^+}(x)=x-\alpha\{h(x)+\bar g(x,\xi)+\zeta\}.\) By Assumption 2([[itm:N5]](#itm:N5){reference-type=“ref” reference=“itm:N5”}), \(\left\lVert \bar g(x,\xi) \right\rVert\le C(1+\left\lVert x \right\rVert^{\beta-1})\), uniformly in \((x,\xi)\). Fix \(\eta_0\in(0,1/2)\), put \[N_x:=\frac{\left\lVert g(x,\xi^+) \right\rVert}{1+\left\lVert x \right\rVert^{\beta-1}}, \qquad E_x:=\{N_x\le\alpha^{-\eta_0}\}.\] On \(E_x\), for sufficiently small \(\alpha\), \[\alpha\left\lVert h(x)+G \right\rVert \le C\alpha^{1-\eta_0}r^{\beta-1} \le \varepsilon_* r, \qquad r\ge R_H,\] with \(\varepsilon_*>0\) chosen so that the segment from \(x\) to \(f_{\xi^+}(x)\) remains in the tail region where the cutoff in \(\mathcal{T}_\beta\) is inactive. Lemma 12(iii), applied to \(z=\alpha\{h(x)+G\}\), gives on \(E_x\) \[\begin{align} \left\lVert f_{\xi^+}(x) \right\rVert^{2-\beta} &\le r^{2-\beta} -(2-\beta)\alpha r^{-\beta}\left\langle x,h(x)+\bar g(x,\xi)\right\rangle -(2-\beta)\alpha r^{-\beta}\left\langle x,\zeta\right\rangle +C\alpha^{2-2\eta_0}. \end{align}\] By the tail dissipativity condition ([[itm:N4]](#itm:N4){reference-type=“ref” reference=“itm:N4”}), \(r^{-\beta}\left\langle x,h(x)+\bar g(x,\xi)\right\rangle\ge c_{\rm diss}\). Therefore, conditionally on \(\xi\), \[\mathcal{T}_\beta(f_{\xi^+}(x))\mathbf{1}_{E_x} \le \mathcal{T}_\beta(x) \exp\{-\kappa(2-\beta)c_{\rm diss}\alpha -\kappa(2-\beta)\alpha r^{-\beta}\left\langle x,\zeta\right\rangle +C\alpha^{2-2\eta_0}\}.\] Writing \(e=x/r\), the centered term is \[r^{-\beta}\left\langle x,\zeta\right\rangle =r^{1-\beta}(1+\left\lVert x \right\rVert^{\beta-1}) \left\langle e, \frac{g(x,\xi^+)}{1+\left\lVert x \right\rVert^{\beta-1}} -\mathbb{E}\!\left[\frac{g(x,\xi^+)}{1+\left\lVert x \right\rVert^{\beta-1}}\,\middle|\,\xi\right]\right\rangle.\] Since \(r^{1-\beta}(1+\left\lVert x \right\rVert^{\beta-1})\) is uniformly bounded for \(r\ge R_H\), the coefficient of the centered normalized noise is \(O(\alpha)\). The conditional exponential moment in ([[itm:N5]](#itm:N5){reference-type=“ref” reference=“itm:N5”}) also gives a uniform exponential moment for this centered normalized noise. Lemma 12(i) therefore implies \[\mathbb{E}\left[ \exp\{-\kappa(2-\beta)\alpha r^{-\beta}\left\langle x,\zeta\right\rangle\}\mid \xi \right] \le e^{C\alpha^2}.\] Thus \[\label{eq:tail-good-event-bound} \mathbb{E}[\mathcal{T}_\beta(f_{\xi^+}(x))\mathbf{1}_{E_x}] \le e^{-c\alpha+C\alpha^{2-2\eta_0}}\mathcal{T}_\beta(x).\tag{51}\]

It remains to control \(E_x^c\). Because \(2-\beta\in(0,1]\) and \(s\mapsto s^{2-\beta}\) is concave, \[\begin{align} \left\lVert x-\alpha(h(x)+G) \right\rVert^{2-\beta}-r^{2-\beta} &\le (r+\alpha\left\lVert h(x)+G \right\rVert)^{2-\beta}-r^{2-\beta} \\ &\le (2-\beta)r^{1-\beta}\alpha\left\lVert h(x)+G \right\rVert. \end{align}\] For \(r\ge R_H\), Assumption 1([[itm:H3b]](#itm:H3b){reference-type=“ref” reference=“itm:H3b”}) and the definition of \(N_x\) give \(\left\lVert h(x)+G \right\rVert\le C r^{\beta-1}(1+N_x).\) The powers of \(r\) therefore cancel, and we obtain the deterministic bound \(\left\lVert f_{\xi^+}(x) \right\rVert^{2-\beta}-r^{2-\beta} \le C\alpha(1+N_x).\) Consequently, \[\mathcal{T}_\beta(f_{\xi^+}(x)) \le \mathcal{T}_\beta(x)\exp\{C\kappa\alpha(1+N_x)\}.\] Choosing \(\alpha\) so that \(C\kappa\alpha\le\lambda_0/2\), Assumption ([[itm:N5]](#itm:N5){reference-type=“ref” reference=“itm:N5”}) yields \[\label{eq:tail-bad-event-bound} \mathbb{E}[\mathcal{T}_\beta(f_{\xi^+}(x))\mathbf{1}_{E_x^c}] \le Ce^{-c\alpha^{-\eta_0}}\mathcal{T}_\beta(x).\tag{52}\] Combining 51 and 52 , using \(2-2\eta_0>1\) and the fact that \(e^{-c\alpha^{-\eta_0}}=o(\alpha)\), gives 50 after reducing the stepsize threshold. ◻

Lemma 15. There exist \(C>0\) and \(\alpha_0\in(0,1]\) such that, for every \(0<\alpha\le\alpha_0\), \(x\in\mathbb{R}^d\), and \(\xi\in\Xi\), \[\label{eq:global-weight-positive-increment} \mathbb{E}\left[ \bigl(V_\alpha(f_{\Phi(\xi,U_1)}(x))-V_\alpha(x)\bigr)_+ \right] \le C\alpha V_\alpha(x).\tag{53}\] Consequently, \[\label{eq:mean-contractive-weight-gate} \mathbb{E}\left[ L_\Phi(U_1)V_\alpha(f_{\Phi(\xi,U_1)}(x)) \right] \le(\rho_\Xi+C\alpha)V_\alpha(x).\tag{54}\]

Proof. For \(\left\lVert x \right\rVert\le R_H\), 53 is 46 . Suppose \(\left\lVert x \right\rVert\ge R_H\). If \(\beta=2\), then \(\mathcal{T}_\beta\equiv1\) and \(\omega(x)=0\), so \[\bigl(V_\alpha(f_{\Phi(\xi,U_1)}(x))-V_\alpha(x)\bigr)_+ \le\delta_\alpha\le C\alpha V_\alpha(x).\] If \(1\le\beta<2\), use the normalized noise \(N_x:=\frac{\left\lVert g(x,\Phi(\xi,U_1)) \right\rVert}{1+\left\lVert x \right\rVert^{\beta-1}}.\) The deterministic tail comparison used in the proof of Lemma 14 gives \[\mathcal{T}_\beta(f_{\Phi(\xi,U_1)}(x)) \le\mathcal{T}_\beta(x)e^{C\kappa\alpha(1+N_x)}.\] Since \(e^t-1\le te^t\) for \(t\ge0\), Assumption 2 ([[itm:N5]](#itm:N5){reference-type=“ref” reference=“itm:N5”}) yields, uniformly in \(x,\xi\), \[\begin{align} \mathbb{E}\bigl[ (\mathcal{T}_\beta(f_{\Phi(\xi,U_1)}(x)) -\mathcal{T}_\beta(x))_+\bigr] &\le C\alpha\mathcal{T}_\beta(x) \mathbb{E}[(1+N_x)e^{C\kappa\alpha(1+N_x)}]\\ &\le C\alpha V_\alpha(x). \end{align}\] The positive increment of \(\delta_\alpha\omega\) is at most \(\delta_\alpha\le C\alpha V_\alpha(x)\), which proves 53 . Finally, \(0\le L_\Phi\le1\) and \(\mathbb{E}L_\Phi(U_1)=\rho_\Xi\) give \[\begin{align} \mathbb{E}[L_\Phi(U_1)V_\alpha(f_{\Phi(\xi,U_1)}(x))] &\le \rho_\Xi V_\alpha(x) +\mathbb{E}[(V_\alpha(f_{\Phi(\xi,U_1)}(x))-V_\alpha(x))_+], \end{align}\] and 54 follows. ◻

Lemma 16. There exist \(c_{\rm core}>0\) and \(\alpha_{\rm core}\in(0,1]\) such that, for every \(0<\alpha\le\alpha_{\rm core}\), every \(\xi\in\Xi\), and every \(\left\lVert x \right\rVert\le \frac{R_H}{2}\), \[\sup_{e\in\mathbb{S}^{d-1}} (\mathcal{K}_{\xi,\alpha}V_\alpha)(x;e) \le (1-c_{\rm core}\alpha^{m-1})V_\alpha(x).\]

Proof. Fix \(x\) with \(r=\left\lVert x \right\rVert\le \frac{R_H}{2}\), and fix \(e\in\mathbb{S}^{d-1}\). Since \(\frac{3R_H}{4}>\frac{R_H}{2}\), \(\mathcal{T}_\beta(x)=1\). Put \(\mathcal{R}(u)=\mathcal{T}_\beta(u)-1\) when \(\beta<2\), and \(\mathcal{R}\equiv0\) when \(\beta=2\). Then \[V_\alpha(u)=1+\delta_\alpha\omega(u)+\mathcal{R}(u), \qquad V_\alpha(x)=1+\delta_\alpha\omega(x).\] Using 43 , \[\label{eq:fixed-core-pre} \begin{align} (\mathcal{K}_{\xi,\alpha}V_\alpha)(x;e)-V_\alpha(x) &\le \mathbb{E}[\mathcal{D}_{\xi,\alpha}(x,e)]-1 +\delta_\alpha\{\mathbb{E}\omega(f_{\xi^+}(x))-\omega(x)\} \\ &\qquad +\mathbb{E}\mathcal{R}(f_{\xi^+}(x)). \end{align}\tag{55}\]

Put \(a_{\rm in}:=(1-\theta)c_{\rm in}/2\). If \(m=2\), Assumption 1([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}) gives \(\left\langle e,\nabla^2H(x)e\right\rangle\ge c_{\rm in}\) for \(\left\lVert x \right\rVert\le \frac{R_H}{2}\). Hence \[\mathbb{E}\mathcal{D}_{\xi,\alpha}(x,e) \le 1-a_{\rm in}\alpha.\] Since \(\omega\equiv0\) when \(m=2\), 55 and 45 give the claim after absorbing the exponentially small term into \(\alpha\).

Assume now \(m>2\), and put \(s_m:=(m-2)\wedge1\). Then \[\omega(y) = \left(1-\left(\frac{2\left\lVert y \right\rVert}{R_H}\right)^{s_m}\right)_+, \qquad y\in\mathbb{R}^d.\] By Assumption 1([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}) and 44 , \(\mathbb{E}\mathcal{D}_{\xi,\alpha}(x,e) \le 1-a_{\rm in}\alpha r^{m-2}.\) We claim that \[\label{eq:fixed-core-omega-drop-new} \mathbb{E}\omega(f_{\xi^+}(x))-\omega(x) \le -c_\omega\mathbb{E}\min\{(\alpha\left\lVert g(0,\xi^+) \right\rVert)^{s_m},1\} +C_\omega r^{s_m}.\tag{56}\] Indeed, write \[f_{\xi^+}(x)=-\alpha g(0,\xi^+)+v, \qquad v=x-\alpha h(x)-\alpha\{g(x,\xi^+)-g(0,\xi^+)\}.\] Lemma 10 gives \(\left\lVert h \right\rVert(x)\le Cr^{m-1}\) for \(\left\lVert x \right\rVert\le \frac{R_H}{2}\), while 25 gives \(\left\lVert g(x,\xi^+)-g(0,\xi^+) \right\rVert \le (2-\theta)L_{\mathsf G}r/(1-\theta)\). Hence, after reducing the stepsize threshold, \(\left\lVert v \right\rVert\le Cr\). The function \(y\mapsto\omega(y)\) is \(s_m\)-Hölder, because \((1-s^{s_m})_+\) is \(s_m\)-Hölder on \([0,\infty)\). Therefore \[\omega(f_{\xi^+}(x)) \le \omega(-\alpha g(0,\xi^+))+C\left\lVert v \right\rVert^{s_m} \le \left(1-\left(\frac{2\alpha\left\lVert g(0,\xi^+) \right\rVert}{R_H}\right)^{s_m}\right)_+ +Cr^{s_m}.\] Since \[\left(1-\left(\frac{2z}{R_H}\right)^{s_m}\right)_+ \le 1-c\min\{z^{s_m},1\}, \qquad z\ge0,\] and \(\omega(x)=1-(\frac{2r}{R_H})^{s_m}\) for \(r\le \frac{R_H}{2}\), 56 follows.

Combining 55 , 56 , 45 , and Lemma 11, we get \[\label{eq:core-combined-before-absorption} (\mathcal{K}_{\xi,\alpha}V_\alpha)(x;e)-V_\alpha(x) \le -a_{\rm in}\alpha r^{m-2} -c\delta_\alpha\bar a_{\alpha,s_m} +C\delta_\alpha r^{s_m} +Ce^{-c/\alpha}.\tag{57}\] If \(2<m\le3\), then \(s_m=m-2\) and \(\delta_\alpha=\kappa_0\alpha\). Using \(\bar a_{\alpha,s_m}\ge c\bar a_\alpha^{m-2}\), 57 becomes \[(\mathcal{K}_{\xi,\alpha}V_\alpha)(x;e)-V_\alpha(x) \le -(a_{\rm in}-C\kappa_0)\alpha r^{m-2} -c\kappa_0\alpha\bar a_\alpha^{m-2} +Ce^{-c/\alpha}.\] Choose \(\kappa_0\) so small that \(a_{\rm in}-C\kappa_0\ge a_{\rm in}/2\). Then \(\bar a_\alpha\ge c_1\alpha\), and the exponentially small term is negligible compared with \(\alpha^{m-1}\). Hence \[\label{eq:core-additive-drop-flat-case} (\mathcal{K}_{\xi,\alpha}V_\alpha)(x;e)-V_\alpha(x) \le -c\alpha^{m-1}.\tag{58}\]

If \(m>3\), then \(s_m=1\), \(\delta_\alpha=\kappa_0\alpha^{m-2}\), and \(\bar a_{\alpha,1}=\bar a_\alpha\). Factoring out \(\alpha\) in 57 gives \[\begin{align} (\mathcal{K}_{\xi,\alpha}V_\alpha)(x;e)-V_\alpha(x) \le \alpha\bigl[ -a_{\rm in}r^{m-2} -c\kappa_0\alpha^{m-3}\bar a_\alpha +C\kappa_0\alpha^{m-3}r \bigr] +Ce^{-c/\alpha}. \end{align}\] Put \(q:=(m-2)/(m-3)>1\). Young’s inequality yields \[C\kappa_0\alpha^{m-3}r \le \frac{a_{\rm in}}{2}r^{m-2} +C_*\kappa_0^q\alpha^{m-2}.\] Since \(\bar a_\alpha\ge c_1\alpha\), the negative noise term inside the brackets is at most \(-cc_1\kappa_0\alpha^{m-2}\). Choose \(\kappa_0\) so small that \(C_*\kappa_0^{q-1}\le cc_1/2\). After restoring the outer factor \(\alpha\), the Young remainder is absorbed by the negative noise term. Finally, \(e^{-c/\alpha}=o(\alpha^{m-1})\), so 58 also holds for \(m>3\).

On the near-minimizer region, \(1\le V_\alpha(x)\le C\) uniformly in \(\alpha\). Since 58 gives an additive decrease of at least \(c\alpha^{m-1}\), it implies \[(\mathcal{K}_{\xi,\alpha}V_\alpha)(x;e) \le (1-c_{\rm core}\alpha^{m-1})V_\alpha(x),\] after decreasing \(c_{\rm core}\). Taking the supremum over \(e\) completes the proof. ◻

Lemma 17. There exist \(c_{\rm br}>0\) and \(\alpha_{\rm br}\in(0,1]\) such that, for every \(0<\alpha\le\alpha_{\rm br}\), every \(\xi\in\Xi\), and every \(\frac{R_H}{2}\le\left\lVert x \right\rVert\le R_H\), \[\sup_{e\in\mathbb{S}^{d-1}} (\mathcal{K}_{\xi,\alpha}V_\alpha)(x;e) \le (1-c_{\rm br}\alpha^{m-1})V_\alpha(x).\]

Proof. Fix \(e\in\mathbb{S}^{d-1}\). By Assumption 1([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}), for \(\frac{R_H}{2}\le\left\lVert x \right\rVert\le R_H\), \(\left\langle e,\nabla^2H(x)e\right\rangle \ge c_{\rm in}\left(\frac{R_H}{2}\right)^{m-2}.\) Together with 44 , this gives \(\mathbb{E}\mathcal{D}_{\xi,\alpha}(x,e) \le 1-c\alpha\) for some \(c>0\). Since \(\mathcal{D}_{\xi,\alpha}(x,e)\le1\), \[\begin{align} (\mathcal{K}_{\xi,\alpha}V_\alpha)(x;e) &\le V_\alpha(x)\mathbb{E}\mathcal{D}_{\xi,\alpha}(x,e) +\mathbb{E}[(V_\alpha(f_{\xi^+}(x))-V_\alpha(x))_+] \\ &\le (1-c\alpha+C(\kappa+\kappa_0)\alpha)V_\alpha(x), \end{align}\] where the last line uses 46 . Choose \(\kappa\) and \(\kappa_0\) sufficiently small to obtain a factor \(1-c\alpha\). Since \(0<\alpha\le1\) and \(m\ge2\), \(\alpha\ge\alpha^{m-1}\), which gives the stated estimate after decreasing the constant. ◻

Lemma 18. There exist \(c_{\rm far}>0\) and \(\alpha_{\rm far}\in(0,1]\) such that, for every \(0<\alpha\le\alpha_{\rm far}\), every \(\xi\in\Xi\), and every \(\left\lVert x \right\rVert\ge R_H\), \[\sup_{e\in\mathbb{S}^{d-1}} (\mathcal{K}_{\xi,\alpha}V_\alpha)(x;e) \le (1-c_{\rm far}\alpha^{m-1})V_\alpha(x).\]

Proof. First suppose \(\beta=2\). Then \(\mathcal{T}_\beta\equiv1\), and \(\omega(x)=0\) for \(\left\lVert x \right\rVert\ge R_H\), so \(V_\alpha(x)=1\). By Assumption 1([[itm:H3a]](#itm:H3a){reference-type=“ref” reference=“itm:H3a”}) and 44 , \(\mathbb{E}\mathcal{D}_{\xi,\alpha}(x,e) \le 1-c\alpha\) for some \(c>0\). Since \(0\le\omega\le1\), \((\mathcal{K}_{\xi,\alpha}V_\alpha)(x;e) \le 1-c\alpha+\delta_\alpha.\) Since \(\delta_\alpha\le\kappa_0\alpha\), choosing \(\kappa_0\) small gives a factor \(1-c\alpha\).

Now suppose \(1\le\beta<2\). On \(\left\lVert x \right\rVert\ge R_H\), \(\omega(x)=0\), so \(V_\alpha(x)=\mathcal{T}_\beta(x)\). By 43 and Lemma 14, \[(\mathcal{K}_{\xi,\alpha}V_\alpha)(x;e) \le \mathbb{E}\mathcal{T}_\beta(f_{\xi^+}(x))+\delta_\alpha \le (1-c\alpha)\mathcal{T}_\beta(x)+\delta_\alpha.\] Since \(\mathcal{T}_\beta(x)\ge \exp\{\kappa(R_H^{2-\beta}-(\tfrac{3R_H}{4})^{2-\beta})\}>1\) on the far region and \(\delta_\alpha\le C\kappa_0\alpha\), choosing \(\kappa_0\) small absorbs the last term and again gives a factor \(1-c\alpha\). In both tail cases, \(\alpha\ge\alpha^{m-1}\); taking the supremum over \(e\) and decreasing the constant proves the stated estimate. ◻

Proof of Lemma 2. The estimates in Lemmas 16, 17, and 18 give 29 after taking \(c_0:=\min\{c_{\rm core},c_{\rm br},c_{\rm far}\}\) and reducing the common stepsize threshold.

It remains to prove the growth estimate 30 . On \(\left\lVert x \right\rVert\le R_H\), it is exactly 47 . If \(\left\lVert x \right\rVert\ge R_H\) and \(\beta=2\), then \(V_\alpha(x)=1\) and \(V_\alpha(f_{\xi^+}(x))\le1+\delta_\alpha\le(1+C\alpha)V_\alpha(x)\). If \(\left\lVert x \right\rVert\ge R_H\) and \(1\le\beta<2\), then Lemma 14 and \(0\le\omega\le1\) give \[\mathbb{E}V_\alpha(f_{\xi^+}(x)) \le (1-c\alpha)\mathcal{T}_\beta(x)+\delta_\alpha \le (1+C\alpha)V_\alpha(x),\] again after reducing the threshold and increasing \(C\). This proves Lemma 2. ◻

6.5 Proof of Lemma 3↩︎

Set \(\rho_\Xi:=\mathbb{E}L_\Phi(U_1)<1,\) where the inequality follows from Assumption 2([[itm:N1]](#itm:N1){reference-type=“ref” reference=“itm:N1”}).

Proof of Lemma 3. Let \(C_V\) be as in Lemma 2. Let \(C_+\) be as in Lemma 15. Choose \(\alpha_0\in(0,1]\) small enough that the conclusions of Lemmas 2 and 15 hold for \(0<\alpha\le\alpha_0\) and \[\label{eq:augmented-xi-dominance-new} \rho_\Xi+C_+\alpha +\alpha^2L_{g,\Phi}(1+C_V\alpha) \le 1-c_0\alpha^{m-1}, \qquad 0<\alpha\le\alpha_0.\tag{59}\] This is possible because \(\rho_\Xi<1\) and \(\alpha^{m-1}\to0\).

Fix \(z=(x,\xi)\), \(z'=(y,\eta)\), and an absolutely continuous path \(\gamma(t)=(x(t),\xi(t))\) from \(z\) to \(z'\). For fixed \(u\), write \[\xi^+(t):=\Phi(\xi(t),u), \qquad X^+(t):=x(t)-\alpha\{h(x(t))+g(x(t),\xi^+(t))\}.\] The curve \(F_u\circ\gamma=(X^+(\cdot),\xi^+(\cdot))\) is absolutely continuous. Indeed, \(\Phi(\cdot,u)\) is Lipschitz by Assumption 2([[itm:N1]](#itm:N1){reference-type=“ref” reference=“itm:N1”}), the composed dependence \(g(x,\Phi(\xi,u))\) is Lipschitz in \(\xi\) by Assumption 2([[itm:N2]](#itm:N2){reference-type=“ref” reference=“itm:N2”}), and the map \(x\mapsto h(x)+g(x,\zeta)\) has derivative \(A(x,\zeta)\), which is uniformly bounded by 25 .

Let \(\operatorname{md}_\alpha\) denote the metric derivative with respect to the augmented base norm \(\vert\cdot\vert_\alpha\). For a.e. \(t\), \[\label{eq:augmented-metric-derivative-bound} \begin{align} \operatorname{md}_\alpha(F_u\circ\gamma)(t) &\le \left\lVert (I_d-\alpha A(x(t),\xi^+(t)))\dot{x}(t) \right\rVert \\ &\qquad +(L_\Phi(u)+\alpha^2L_{g,\Phi})\alpha^{-1}\left\lVert \dot{\xi}(t) \right\rVert. \end{align}\tag{60}\] To prove 60 , compare \(F_u(\gamma(s))\) and \(F_u(\gamma(t))\). In the \(x\)-coordinate, first freeze the noise argument at \(\xi^+(t)\); division by \(|s-t|\) and passage to the a.e.differentiability point gives the derivative \((I_d-\alpha A(x(t),\xi^+(t)))\dot{x}(t)\). The remaining change in the \(x\)-coordinate is bounded by \(\alpha L_{g,\Phi}\left\lVert \xi(s)-\xi(t) \right\rVert\). In the noise coordinate, Assumption 2([[itm:N1]](#itm:N1){reference-type=“ref” reference=“itm:N1”}) gives \(\left\lVert \Phi(\xi(s),u)-\Phi(\xi(t),u) \right\rVert \le L_\Phi(u)\left\lVert \xi(s)-\xi(t) \right\rVert\). Dividing by \(|s-t|\) and using the \(\alpha^{-1}\)-weight in the augmented norm gives 60 .

If \(\dot{x}(t)\ne0\), set \(e(t)=\dot{x}(t)/\left\lVert \dot{x}(t) \right\rVert\); otherwise choose any measurable unit vector \(e(t)\). By the definition of induced length, 60 , Lemmas 2 and 15, and 59 , \[\begin{align} &\mathbb{E}\, d_{V_\alpha,\alpha}\bigl(F_{U_1}(z),F_{U_1}(z')\bigr) \\ &\quad\le \int_0^1 \mathbb{E}\!\left[ V_\alpha(X^+(t)) \left\lVert (I_d-\alpha A(x(t),\xi^+(t)))\dot{x}(t) \right\rVert \right]dt \\ &\qquad + \int_0^1 \mathbb{E}\!\left[ (L_\Phi(U_1)+\alpha^2L_{g,\Phi})V_\alpha(X^+(t)) \right] \alpha^{-1}\left\lVert \dot{\xi}(t) \right\rVert\,dt \\ &\quad\le (1-c_0\alpha^{m-1}) \int_0^1 V_\alpha(x(t)) \left(\left\lVert \dot{x}(t) \right\rVert+\alpha^{-1}\left\lVert \dot{\xi}(t) \right\rVert\right)dt. \end{align}\] Taking the infimum over all admissible paths \(\gamma\) gives 24 . ◻

6.6 Proof of Lemma 4↩︎

Proof of Lemma 4. Let \(z_\star=(0,\xi_\star)\), where \(\xi_\star\) is the point from Assumption 2([[itm:N3]](#itm:N3){reference-type=“ref” reference=“itm:N3”}). Write \[\xi_\star^+:=\Phi(\xi_\star,U_1), \qquad x_{\rm ref}^+:=-\alpha g(0,\xi_\star^+).\] Move from \(z_\star=(0,\xi_\star)\) to \((0,\xi_\star^+)\), and then from \((0,\xi_\star^+)\) to \((x_{\rm ref}^+,\xi_\star^+)\). Since \(V_\alpha\) depends only on the \(x\)-coordinate, \[d_{V_\alpha,\alpha}\bigl(F_{U_1}(z_\star),z_\star\bigr) \le V_\alpha(0)\alpha^{-1}\left\lVert \xi_\star^+-\xi_\star \right\rVert + d_{V_\alpha}^{X}(x_{\rm ref}^+,0).\] The first term is integrable by Assumption 2([[itm:N3]](#itm:N3){reference-type=“ref” reference=“itm:N3”}).

If \(\beta=2\), then \(\mathcal{T}_\beta\equiv1\) and \(V_\alpha\le1+\delta_\alpha\). Hence \[d_{V_\alpha}^{X}(x_{\rm ref}^+,0) \le (1+\delta_\alpha)\left\lVert x_{\rm ref}^+ \right\rVert = (1+\delta_\alpha)\alpha\left\lVert g(0,\xi_\star^+) \right\rVert,\] whose expectation is finite by Assumption 2([[itm:N3]](#itm:N3){reference-type=“ref” reference=“itm:N3”}).

If \(\beta<2\), Assumption 2([[itm:N5]](#itm:N5){reference-type=“ref” reference=“itm:N5”}) at \(x=0\) and \(\xi=\xi_\star\) gives an exponential moment of \(g(0,\xi_\star^+)\). Along the straight line from \(0\) to \(x_{\rm ref}^+\), \[d_{V_\alpha}^{X}(x_{\rm ref}^+,0) \le \alpha\left\lVert g(0,\xi_\star^+) \right\rVert \left[ 1+\delta_\alpha+ \exp\!\left\{\kappa\bigl(\alpha^{2-\beta}\left\lVert g(0,\xi_\star^+) \right\rVert^{2-\beta}-(\tfrac{3R_H}{4})^{2-\beta}\bigr)_+\right\} \right].\] Since \(2-\beta\in(0,1]\), \(r^{2-\beta}\le1+r\) for \(r\ge0\), we may choose \(\alpha_0\in(0,1]\) such that, for \(0<\alpha\le\alpha_0\), the exponential factor above is dominated by \(C e^{\lambda\left\lVert g(0,\xi_\star^+) \right\rVert}\) for some \(\lambda>0\) below the exponential-moment threshold supplied by ([[itm:N5]](#itm:N5){reference-type=“ref” reference=“itm:N5”}). Thus \(\mathbb{E}d_{V_\alpha}^{X}(x_{\rm ref}^+,0)<\infty.\) This proves 31 . ◻

6.7 Proof of Corollary 1↩︎

Proof of Corollary 1. Since \(V_\alpha\ge1\), the augmented induced metric dominates the Euclidean distance in the \(X\)-coordinate: \(d_{V_\alpha,\alpha}\bigl((x,\xi),(y,\eta)\bigr) \ge \left\lVert x-y \right\rVert.\) Hence, for probability laws \(\mu,\nu\) on \(\mathsf Z\), \[W_1(\mu_X,\nu_X)\le W_{V_\alpha,\alpha}(\mu,\nu).\] Applying this with \(\nu=\pi_\alpha\), and then using Theorem 4, gives \[W_1\bigl((\mu P_\alpha^n)_X,(\pi_\alpha)_X\bigr) \le W_{V_\alpha,\alpha}(\mu P_\alpha^n,\pi_\alpha) \le (1-c\alpha^{m-1})^n W_{V_\alpha,\alpha}(\mu,\pi_\alpha).\] Since \(\pi_\alpha\in\mathcal{P}_1(\mathsf Z,d_{V_\alpha,\alpha})\), the final quantity is finite if and only if \(\mu\in\mathcal{P}_1(\mathsf Z,d_{V_\alpha,\alpha})\).

It remains to characterize \(\mathcal{P}_1(\mathsf Z,d_{V_\alpha,\alpha})\). Since \(\vert(x,\xi)\vert_\alpha = \left\lVert x \right\rVert+\alpha^{-1}\left\lVert \xi \right\rVert,\) and since \(V_\alpha\ge1\), projection of paths gives the lower bounds \[d_{V_\alpha,\alpha}\bigl((x,\xi),(0,\xi_\star)\bigr) \ge d_{V_\alpha}^{X}(x,0), \qquad d_{V_\alpha,\alpha}\bigl((x,\xi),(0,\xi_\star)\bigr) \ge \alpha^{-1}\left\lVert \xi-\xi_\star \right\rVert.\] Conversely, moving first in the \(x\)-coordinate and then in the \(\xi\)-coordinate gives \[d_{V_\alpha,\alpha}\bigl((x,\xi),(0,\xi_\star)\bigr) \le d_{V_\alpha}^{X}(x,0) + V_\alpha(0)\alpha^{-1}\left\lVert \xi-\xi_\star \right\rVert.\] Thus \(d_{V_\alpha,\alpha}((x,\xi),(0,\xi_\star))\)-integrability is equivalent to integrability of \(d_{V_\alpha}^{X}(x,0) + \alpha^{-1}\left\lVert \xi-\xi_\star \right\rVert.\)

If \(\beta=2\), then \(\mathcal{T}_\beta\equiv1\) and \(1\le V_\alpha\le1+\delta_\alpha\). Hence \(\left\lVert x-y \right\rVert \le d_{V_\alpha}^{X}(x,y) \le (1+\delta_\alpha)\left\lVert x-y \right\rVert,\) so \(d_{V_\alpha}^{X}(x,0)\)-integrability is equivalent to \(\left\lVert x \right\rVert\)-integrability.

If \(1\le\beta<2\), let \(R_H\) be as in Assumption 1([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}) and define \[\Psi_\kappa(r) := \int_0^r \exp\{\kappa\bigl(s^{2-\beta}-(\tfrac{3R_H}{4})^{2-\beta}\bigr)_+\}\,ds.\] Because the correction \(\delta_\alpha\omega\) is compactly supported and \(0\le\omega\le1\), \[\Psi_\kappa(\left\lVert x \right\rVert) \le d_{V_\alpha}^{X}(x,0) \le \Psi_\kappa(\left\lVert x \right\rVert)+\frac{\delta_\alpha R_H}{2}.\] Moreover, \[\Psi_\kappa(r) \sim \frac{e^{-\kappa(\frac{3R_H}{4})^{2-\beta}}}{\kappa(2-\beta)} r^{\beta-1} e^{\kappa r^{2-\beta}}, \qquad r\to\infty.\] Therefore \(d_{V_\alpha}^{X}(x,0)\)-integrability is equivalent to \[\int (1+\left\lVert x \right\rVert)^{\beta-1} \exp\{\kappa\left\lVert x \right\rVert^{2-\beta}\}\,\mu(dx,d\xi) <\infty.\] Since \(2-\beta\in(0,1]\), for \(r\ge0\), \(r^{2-\beta}\le(1+r)^{2-\beta}\le1+r^{2-\beta},\) and hence \[e^{-\kappa}e^{\kappa r^{2-\beta}} \le \exp\!\left\{\kappa\bigl((1+r)^{2-\beta}-1\bigr)\right\} \le e^{\kappa r^{2-\beta}}.\] Thus the preceding integrability condition is equivalent to integrability of \(\Gamma_\beta(\left\lVert x \right\rVert)\). Combining this with the \(\xi\)-coordinate term gives the displayed description of \(\mathcal{P}_1(\mathsf Z,d_{V_\alpha,\alpha})\). ◻

7 Scaling-limit estimates↩︎

This appendix proves the auxiliary results used for the scaling limit of invariant laws. Constants denoted by \(C,c\) may change from line to line but are independent of \(\alpha\), unless stated otherwise.

7.1 Deterministic scaling estimates↩︎

Lemma 19. Under Assumption 1, the following bounds hold.

  1. If \(\left\lVert x \right\rVert\le \frac{R_H}{2}\), then \(\left\langle x,h(x)\right\rangle \ge \frac{c_{\rm in}}{m-1}\left\lVert x \right\rVert^m.\)

  2. There exists \(C_{\rm tail}>0\) such that \[\left\lVert x \right\rVert^\beta \le C_{\rm tail}\bigl(1+\left\langle x,h(x)\right\rangle\bigr), \qquad 1\le\beta<2,\] and \[\left\lVert x \right\rVert^2 \le C_{\rm tail}\bigl(1+\left\langle x,h(x)\right\rangle\bigr), \qquad \beta=2.\]

  3. With \(r_0:=\min\{1,\frac{R_H}{2}\}\), there exists \(c_*>0\) such that \[\left\langle x,h(x)\right\rangle\ge c_*, \qquad \left\lVert x \right\rVert\ge r_0.\]

  4. There exists \(C_a>0\) such that, for all \(x\in\mathbb{R}^d\), \((1+\left\lVert x \right\rVert)(1+\left\lVert x \right\rVert^{\beta-1}) \le C_a\bigl(1+\left\langle x,h(x)\right\rangle\bigr).\)

Proof. For the lower bound near the minimizer, use \(h(x)=\int_0^1\nabla^2H(tx)x\,dt.\) Assumption 1([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}) gives \[\left\langle x,h(x)\right\rangle = \int_0^1\left\langle x,\nabla^2H(tx)x\right\rangle\,dt \ge c_{\rm in}\left\lVert x \right\rVert^m\int_0^1t^{m-2}\,dt = \frac{c_{\rm in}}{m-1}\left\lVert x \right\rVert^m.\]

If \(1\le\beta<2\), the tail condition \(\left\langle x,h(x)\right\rangle\ge c_{\rm out}\left\lVert x \right\rVert^\beta\) on \(\{\left\lVert x \right\rVert\ge R_H\}\), together with boundedness on \(\{\left\lVert x \right\rVert<R_H\}\), gives \(\left\lVert x \right\rVert^{\beta} \le C\bigl(1+\left\langle x,h(x)\right\rangle\bigr).\) If \(\beta=2\), then for \(\left\lVert x \right\rVert\ge2R_H\), convexity and minimality at zero imply \(\left\langle x,h(x/2)\right\rangle\ge0.\) Using Assumption 1([[itm:H3a]](#itm:H3a){reference-type=“ref” reference=“itm:H3a”}) on the segment \(\{tx:1/2\le t\le1\}\), \[\begin{align} \left\langle x,h(x)\right\rangle &\ge \left\langle x,h(x)-h(x/2)\right\rangle \\ &= \int_{1/2}^1\left\langle x,\nabla^2H(tx)x\right\rangle\,dt \ge \frac{c_{\rm out}}{2}\left\lVert x \right\rVert^2. \end{align}\] The region \(\{\left\lVert x \right\rVert<2R_H\}\) is absorbed into the additive constant.

Finally, write \(x=ru\), where \(r=\left\lVert x \right\rVert\) and \(u\in\mathbb{S}^{d-1}\). Convexity of \(H\) and minimality at zero imply that \(s\mapsto\left\langle u,h(su)\right\rangle\) is nondecreasing on \([0,\infty)\). For \(r\ge r_0\), \[\left\langle x,h(x)\right\rangle = r\left\langle u,h(ru)\right\rangle \ge r_0\left\langle u,h(r_0u)\right\rangle = \left\langle r_0u,h(r_0u)\right\rangle \ge \frac{c_{\rm in}}{m-1}r_0^m.\] This proves item (iii). For item (iv), if \(\beta=2\), then \((1+\left\lVert x \right\rVert)(1+\left\lVert x \right\rVert^{\beta-1}) \le C(1+\left\lVert x \right\rVert^2),\) and the quadratic-tail estimate in item (ii) gives the result. If \(1\le\beta<2\), then \((1+\left\lVert x \right\rVert)(1+\left\lVert x \right\rVert^{\beta-1}) \le C(1+\left\lVert x \right\rVert^\beta),\) and the subquadratic-tail estimate in item (ii) gives the result. This proves the lemma. ◻

7.2 Driving-chain ergodicity and Poisson equation↩︎

Lemma 20. Assume Assumption 2([[itm:N1]](#itm:N1){reference-type=“ref” reference=“itm:N1”}) and the reference-point condition 4 . Let \(Q\) be the transition kernel of \(\xi_{n+1}=\Phi(\xi_n,U_{n+1}).\) Then \(Q\) admits an invariant law \(\pi_\Xi\in\mathcal{P}_1(\Xi)\), unique among all Borel invariant probability laws on \(\Xi\). Moreover, \(\mathbb{E}\left\lVert \xi_n^\xi-\xi_n^\eta \right\rVert \le\rho_\Xi^n\left\lVert \xi-\eta \right\rVert,\) for synchronously coupled chains started from \(\xi\) and \(\eta\), and \[W_1(\mu Q,\nu Q)\le \rho_\Xi W_1(\mu,\nu), \qquad \mu,\nu\in\mathcal{P}_1(\Xi),\] and if \(\pi_\alpha\) is the invariant law of the augmented chain, then its \(\Xi\)-marginal is \(\pi_\Xi\).

Proof. Let \(\xi_\star\) be as in Assumption 2([[itm:N3]](#itm:N3){reference-type=“ref” reference=“itm:N3”}). If \(\eta\sim\mu\in\mathcal{P}_1(\Xi)\), then Assumption 2([[itm:N1]](#itm:N1){reference-type=“ref” reference=“itm:N1”}) gives \[\begin{align} \mathbb{E}\left\lVert \Phi(\eta,U_1)-\xi_\star \right\rVert &\le \mathbb{E}\left\lVert \Phi(\eta,U_1)-\Phi(\xi_\star,U_1) \right\rVert + \mathbb{E}\left\lVert \Phi(\xi_\star,U_1)-\xi_\star \right\rVert \\ &\le \rho_\Xi\mathbb{E}\left\lVert \eta-\xi_\star \right\rVert + \mathbb{E}\left\lVert \Phi(\xi_\star,U_1)-\xi_\star \right\rVert. \end{align}\] Thus \(Q\) maps \(\mathcal{P}_1(\Xi)\) into itself.

For any coupling \((\eta,\widetilde{\eta})\) of \((\mu,\nu)\), drive both chains by the same innovation \(U_1\). Then \[\mathbb{E}\left\lVert \Phi(\eta,U_1)-\Phi(\widetilde{\eta},U_1) \right\rVert \le \rho_\Xi\mathbb{E}\left\lVert \eta-\widetilde{\eta} \right\rVert.\] Iteration with independent innovations gives the stated synchronous \(n\)-step estimate. Taking the infimum over couplings yields \(W_1(\mu Q,\nu Q)\le\rho_\Xi W_1(\mu,\nu).\) Since \(\Xi\) is closed in a finite-dimensional Euclidean space, \((\mathcal{P}_1(\Xi),W_1)\) is complete. Banach’s fixed-point theorem gives a unique invariant law \(\pi_\Xi\in\mathcal{P}_1(\Xi)\).

To prove uniqueness among all Borel invariant laws, let \(\rho\) be such a law on \(\Xi\). For each fixed \(\xi\), the contraction estimate gives \(\delta_\xi Q^n\to\pi_\Xi\) in \(W_1\). Hence, for every bounded Lipschitz \(\varphi\), \(Q^n\varphi(\xi)\to\pi_\Xi(\varphi).\) By invariance and dominated convergence, \(\rho(\varphi) = \rho(Q^n\varphi) \to \pi_\Xi(\varphi).\) Bounded Lipschitz functions determine Borel probability laws on the Polish space \(\Xi\), so \(\rho=\pi_\Xi\).

Finally, if \((X,\xi)\sim\pi_\alpha\), invariance of the augmented chain implies that \(\Phi(\xi,U_1)\) has the same law as \(\xi\). Thus the \(\Xi\)-marginal of \(\pi_\alpha\) is invariant for \(Q\) and equals \(\pi_\Xi\). ◻

Lemma 21. Assume the hypotheses of Lemma 20. Let \(f:\Xi\to\mathbb{R}\) be bounded, Lipschitz, and centered under \(\pi_\Xi\). Define \(u(\xi):=\sum_{k=0}^{\infty}Q^kf(\xi).\) Then the series converges absolutely for every \(\xi\), \(u\) is Lipschitz with \(\mathop{\mathrm{Lip}}(u)\le\frac{\mathop{\mathrm{Lip}}(f)}{1-\rho_\Xi},\) \(u\in L^2(\pi_\Xi)\), and \(u-Qu=f.\)

Proof. Let \(\xi_k^\xi\) be the driving chain started from \(\xi\). Let \(\eta_0\sim\pi_\Xi\) and drive the stationary chain \((\eta_k)\) by the same innovations. Since \(f\) is centered under \(\pi_\Xi\), \(Q^kf(\xi) = \mathbb{E}[f(\xi_k^\xi)-f(\eta_k)].\) Therefore \(\left\lvert Q^kf(\xi) \right\rvert \le \min\left\{ 2\left\lVert f \right\rVert_\infty,\, \mathop{\mathrm{Lip}}(f)\rho_\Xi^k\bigl(\left\lVert \xi \right\rVert+\mathbb{E}_{\pi_\Xi}\left\lVert \eta \right\rVert\bigr) \right\}.\) This bound proves absolute convergence. Similarly, synchronous coupling of two chains started from \(\xi\) and \(\eta\) gives \(\left\lvert Q^kf(\xi)-Q^kf(\eta) \right\rvert \le \mathop{\mathrm{Lip}}(f)\rho_\Xi^k\left\lVert \xi-\eta \right\rVert,\) and summing over \(k\) proves the Lipschitz bound for \(u\).

The first bound also implies logarithmic growth: \(\left\lvert u(\xi) \right\rvert \le C_f\bigl(1+\log(1+\left\lVert \xi \right\rVert)\bigr).\) Since \(\pi_\Xi\in\mathcal{P}_1(\Xi)\) and \(\log^2(1+r)\le C(1+r)\), we get \(u\in L^2(\pi_\Xi)\). Finally, absolute convergence allows termwise application of \(Q\), giving \[Qu=\sum_{k=1}^{\infty}Q^kf, \qquad u-Qu=f.\] ◻

7.3 Proof of Lemma 5↩︎

Proof of Lemma 5. Let \(\xi_k^\xi\) and \(\xi_k^\eta\) be two copies of the driving chain, started from \(\xi\) and \(\eta\) and coupled through the same innovations. By Assumption 2([[itm:N1]](#itm:N1){reference-type=“ref” reference=“itm:N1”}), \(\mathbb{E}\left\lVert \xi_k^\xi-\xi_k^\eta \right\rVert \le \rho_\Xi^k\left\lVert \xi-\eta \right\rVert.\) For \(k\ge1\), Assumption 2([[itm:N2]](#itm:N2){reference-type=“ref” reference=“itm:N2”}) gives \[\left\lVert Q^kg_x(\xi)-Q^kg_x(\eta) \right\rVert \le L_{g,\Phi}\rho_\Xi^{k-1}\left\lVert \xi-\eta \right\rVert.\] Let \(\eta\sim\pi_\Xi\), independent of the driving innovations. Since \(\pi_\Xi g_x=0\) by 9 , \[Q^kg_x(\xi)=\mathbb{E}[g_x(\xi_k^\xi)-g_x(\xi_k^\eta)], \qquad k\ge1.\] Consequently \[\left\lVert Q^kg_x(\xi) \right\rVert \le L_{g,\Phi}\rho_\Xi^{k-1} \bigl(\left\lVert \xi-\xi_\star \right\rVert +\mathbb{E}_{\pi_\Xi}\left\lVert \eta-\xi_\star \right\rVert\bigr),\] and the series defining \(\chi_x\) converges absolutely for every \(\xi\). Applying \(Q\) termwise gives \(\chi_x-Q\chi_x=Qg_x\). The same estimate gives \[\label{eq:poisson-solution-growth} \left\lVert \chi_x(\xi) \right\rVert \le C_\chi R_\chi(\xi), \qquad R_\chi(\xi):=1+\left\lVert \xi-\xi_\star \right\rVert,\tag{61}\] with \(R_\chi\in L^2(\pi_\Xi)\) by 10 . Centering of \(\chi_x\) follows from invariance of \(\pi_\Xi\): \(\int_\Xi \chi_x(\xi)\,\pi_\Xi(d\xi)=0,\) because \(\int Q^kg_x\,d\pi_\Xi=\int g_x\,d\pi_\Xi=0\) for every \(k\ge1\).

It remains to control the dependence on \(x\). The uniform derivative bound 25 shows that, for some \(L_x>0\), \[\left\lVert g_x(\zeta)-g_y(\zeta) \right\rVert \le L_x\left\lVert x-y \right\rVert, \qquad x,y\in\mathbb{R}^d,\;\zeta\in\Xi.\] Fix \(p\in(0,1)\), and put \(f_{x,y}:=g_x-g_y\). By 9 , \(\pi_\Xi f_{x,y}=0\). On the one hand, \[\left\lVert Q^kf_{x,y}(\xi) \right\rVert \le L_x\left\lVert x-y \right\rVert, \qquad k\ge1.\] On the other hand, comparison with the stationary chain used above and Assumption 2([[itm:N2]](#itm:N2){reference-type=“ref” reference=“itm:N2”}), applied separately to \(g_x\) and \(g_y\), give \[\begin{align} \left\lVert Q^kf_{x,y}(\xi) \right\rVert &\le 2L_{g,\Phi}\rho_\Xi^{k-1} \bigl(\left\lVert \xi-\xi_\star \right\rVert +\mathbb{E}_{\pi_\Xi}\left\lVert \eta-\xi_\star \right\rVert\bigr)\\ &\le C\rho_\Xi^{k-1}R_\chi(\xi). \end{align}\] Since \(\min\{a,b\}\le a^pb^{1-p}\) for \(a,b\ge0\), summing these two bounds over \(k\ge1\) yields the Hölder estimate \[\label{eq:poisson-solution-x-holder} \left\lVert \chi_x(\xi)-\chi_y(\xi) \right\rVert \le C_p\left\lVert x-y \right\rVert^pR_\chi(\xi)^{1-p}, \qquad x,y\in\mathbb{R}^d,\;\xi\in\Xi.\tag{62}\] The martingale-difference property 33 is immediate from 32 .

Finally, let \(f(\xi)=g(0,\xi)\). The geometric estimate used to prove 61 , with \(x=0\), gives \[\left\lVert Q^kf(\xi) \right\rVert \le C\rho_\Xi^{k-1}(1+\left\lVert \xi-\xi_\star \right\rVert), \qquad k\ge1.\] Together with 10 , this implies absolute summability of the lag covariances. Write the stationary chain as \((\xi_n)_{n\in\mathbb{Z}}\), set \(f_n=f(\xi_n)\) and \(\chi_n=\chi_0(\xi_n)\), and define \(D_n:=f_{n+1}+\chi_{n+1}-\chi_n.\) Then \((D_n)\) is a martingale-difference sequence and \(f_{n+1}=D_n+\chi_n-\chi_{n+1}.\) Therefore, for \(N\ge1\), \(\sum_{n=1}^N f_n = \sum_{n=0}^{N-1}D_n+\chi_0-\chi_N.\) Since \(\chi_0\in L^2(\pi_\Xi)\), the telescoping term is negligible after division by \(N\) in the covariance of the partial sums. Hence \[\lim_{N\to\infty} \frac{1}{N}\mathop{\mathrm{Var}}\!\left(\sum_{n=1}^N f_n\right) = \mathbb{E}[D_0D_0^\top] = \Sigma.\] On the other hand, stationarity and absolute summability of the lag covariances give \[\lim_{N\to\infty} \frac{1}{N}\mathop{\mathrm{Var}}\!\left(\sum_{n=1}^N f_n\right) = \Gamma_0+\sum_{k\ge1}(\Gamma_k+\Gamma_k^\top).\] This proves 11 and completes the proof of the lemma. ◻

7.4 Proof of Lemma 6↩︎

Lemma 22. Assume Assumptions 3 and 4. Then there exist constants \(A,B>0\), independent of \(\alpha\), such that \(\mathbb{E}\left\lVert g(X_\alpha,\xi_\alpha^+) \right\rVert^2 \le A+B\,\mathbb{E}\left\langle X_\alpha,h(X_\alpha)\right\rangle.\)

Proof. By Lemma 20, \(\xi_\alpha^+\sim\pi_\Xi\).

If \(\beta=2\), 25 gives a uniform \(x\)-Lipschitz constant \((2-\theta)L_{\mathsf G}/(1-\theta)\) for \(g\). Hence \[\left\lVert g(x,\zeta) \right\rVert^2 \le 2\left\lVert g(0,\zeta) \right\rVert^2 +2\left(\frac{(2-\theta)L_{\mathsf G}}{1-\theta}\right)^2\left\lVert x \right\rVert^2.\] The stationary second-moment condition 10 gives \(\mathbb{E}\left\lVert g(0,\xi_\alpha^+) \right\rVert^2<\infty,\) and Lemma 19 gives \(\left\lVert x \right\rVert^2\le C(1+\left\langle x,h(x)\right\rangle).\) Taking expectations proves the claim in the quadratic-tail case.

If \(1\le\beta<2\), Assumption 2([[itm:N5]](#itm:N5){reference-type=“ref” reference=“itm:N5”}) gives a uniform second moment for \(\frac{g(x,\Phi(\xi,U_1))}{1+\left\lVert x \right\rVert^{\beta-1}}.\) Therefore, conditionally on \((X_\alpha,\xi_\alpha)\), \[\mathbb{E}\bigl[\left\lVert g(X_\alpha,\xi_\alpha^+) \right\rVert^2\mid X_\alpha,\xi_\alpha\bigr] \le C(1+\left\lVert X_\alpha \right\rVert^{\beta-1})^2 \le C(1+\left\lVert X_\alpha \right\rVert^{\beta}).\] Using again Lemma 19, \(\left\lVert x \right\rVert^{\beta}\le C(1+\left\langle x,h(x)\right\rangle),\) and taking expectations proves the claim. ◻

Lemma 23. Assume Assumptions 3 and 4. For all sufficiently small \(\alpha\), \[\mathbb{E}\left\lVert X_\alpha \right\rVert^2<\infty, \qquad \mathbb{E}\left\lVert h(X_\alpha)+g(X_\alpha,\xi_\alpha^+) \right\rVert^2<\infty.\] Moreover, \[\label{eq:stationary-cross-poisson-bound} \mathbb{E}\left\langle X_\alpha,g(X_\alpha,\xi_\alpha^+)\right\rangle \ge -\theta\mathbb{E}\left\langle X_\alpha,h(X_\alpha)\right\rangle -C\alpha\left(1+\mathbb{E}\left\langle X_\alpha,h(X_\alpha)\right\rangle\right).\tag{63}\]

Proof. We first justify the second moments. If \(1\le\beta<2\), the induced first-moment bound in Theorem 4 implies that \(X_\alpha\) has finite moments of every polynomial order. Lemma 22 and the global Lipschitz bound on \(h\) from Lemma 1 then give \(\mathbb{E}\left\lVert h(X_\alpha)+g(X_\alpha,\xi_\alpha^+) \right\rVert^2<\infty.\)

For \(\beta=2\), Lemma 19 gives the global coercivity bound \[\label{eq:h-quadratic-coercive-for-L2} \left\langle x,h(x)\right\rangle\ge c\left\lVert x \right\rVert^2-C, \qquad x\in\mathbb{R}^d.\tag{64}\] Indeed, outside a sufficiently large ball this follows from the quadratic-tail lower Hessian bound, while the remaining bounded region is absorbed into the additive constant.

Start the augmented chain from \(X_0=0\) and \(\xi_0\sim\pi_\Xi\). Then \(\xi_n\sim\pi_\Xi\) for all \(n\). Define the martingale-corrected quadratic Lyapunov function \(J_\alpha^0(x,\xi) :=\left\lVert x \right\rVert^2-2\alpha\left\langle x,\chi_0(\xi)\right\rangle.\) For any law whose \(\Xi\)-marginal is \(\pi_\Xi\), 61 , Young’s inequality, and \(R_\chi\in L^2(\pi_\Xi)\) give \[\label{eq:Lalpha-comparable-L2} \frac{1}{2}\mathbb{E}\left\lVert X \right\rVert^2-C\alpha^2 \le \mathbb{E}J_\alpha^0(X,\xi) \le \frac{3}{2}\mathbb{E}\left\lVert X \right\rVert^2+C\alpha^2.\tag{65}\] For finite \(n\), the second moments below are finite by induction from \(X_0=0\), the finite variance of \(g(0,\xi)\) under \(\pi_\Xi\), and the global Lipschitz bounds in Lemma 1. Write \[\mathsf G_n:=\mathsf G(X_n,\xi_{n+1}), \qquad X_{n+1}=X_n-\alpha\mathsf G_n.\] The mean-perturbation condition 7 , with \(y=0\), implies \[\begin{align} &\mathbb{E}\!\left[ \left\langle X_n,\mathsf G(X_n,\xi_{n+1})-\mathsf G(0,\xi_{n+1})\right\rangle \,\middle|\,X_n,\xi_n\right]\\ &\qquad= \left\langle X_n, h(X_n)+\bar g(X_n,\xi_n)-\bar g(0,\xi_n)\right\rangle \ge (1-\theta)\left\langle X_n,h(X_n)\right\rangle. \end{align}\] The Poisson identity at the minimizer gives \(\mathbb{E}\left\langle X_n,g(0,\xi_{n+1})\right\rangle = \mathbb{E}\left\langle X_n,\chi_0(\xi_n)-\chi_0(\xi_{n+1})\right\rangle.\) Expanding \(J_\alpha^0\), using \(g(X_n,\xi_{n+1})=g(0,\xi_{n+1})+ \mathsf G(X_n,\xi_{n+1})-\mathsf G(0,\xi_{n+1})-h(X_n)\), and applying the preceding cancellation yield \[\begin{align} &\mathbb{E}[J_\alpha^0(X_{n+1},\xi_{n+1})- J_\alpha^0(X_n,\xi_n)]\\ &\quad=-2\alpha\mathbb{E}\left\langle X_n, \mathsf G(X_n,\xi_{n+1})-\mathsf G(0,\xi_{n+1})\right\rangle \\ &\quad+\alpha^2\mathbb{E}\left\lVert \mathsf G_n \right\rVert^2 +2\alpha^2\mathbb{E}\left\langle \mathsf G_n,\chi_0(\xi_{n+1})\right\rangle. \end{align}\] The global Lipschitz bounds in Lemma 1, together with \(\mathbb{E}_{\pi_\Xi}\left\lVert g(0,\xi) \right\rVert^2<\infty\), imply \(\mathbb{E}\left\lVert \mathsf G_n \right\rVert^2\le C\bigl(1+\mathbb{E}\left\lVert X_n \right\rVert^2\bigr).\) Moreover, Cauchy–Schwarz and \(\chi_0\in L^2(\pi_\Xi)\) give \[\left\lvert \mathbb{E}\left\langle \mathsf G_n,\chi_0(\xi_{n+1})\right\rangle \right\rvert \le \bigl(\mathbb{E}\left\lVert \mathsf G_n \right\rVert^2\bigr)^{1/2} \bigl(\mathbb{E}_{\pi_\Xi}\left\lVert \chi_0 \right\rVert^2\bigr)^{1/2} \le C\bigl(1+\mathbb{E}\left\lVert X_n \right\rVert^2\bigr).\] Combining these bounds with 64 yields \[\mathbb{E}J_\alpha^0(X_{n+1},\xi_{n+1}) \le \mathbb{E}J_\alpha^0(X_n,\xi_n)-c\alpha\mathbb{E}\left\lVert X_n \right\rVert^2+C\alpha +C\alpha^2\bigl(1+\mathbb{E}\left\lVert X_n \right\rVert^2\bigr).\] For sufficiently small \(\alpha\), 65 absorbs the \(C\alpha^2\mathbb{E}\left\lVert X_n \right\rVert^2\) term and gives \[\label{eq:Lalpha-geometric-bound} \mathbb{E}J_\alpha^0(X_{n+1},\xi_{n+1}) \le (1-c\alpha)\mathbb{E}J_\alpha^0(X_n,\xi_n)+C\alpha.\tag{66}\] Iterating 66 and using \(X_0=0\) gives \(\sup_n\mathbb{E}\left\lVert X_n \right\rVert^2<\infty\). By Theorem 4, the finite-time laws converge weakly to \(\pi_\alpha\). Lower semicontinuity of \(x\mapsto\left\lVert x \right\rVert^2\) then gives \(\mathbb{E}_{\pi_\alpha}\left\lVert X_\alpha \right\rVert^2<\infty\). Lemma 22 and the global Lipschitz bound in Lemma 1 give the second moment of \(h(X_\alpha)+g(X_\alpha,\xi_\alpha^+)\).

We now prove 63 . Lemma 22 and 26 imply \[\mathbb{E}\left\lVert h(X_\alpha)+g(X_\alpha,\xi_\alpha^+) \right\rVert^2 \le C\left(1+\mathbb{E}\left\langle X_\alpha,h(X_\alpha)\right\rangle\right).\] By stationarity and the Poisson identity at the minimizer, \[\begin{align} \mathbb{E}\left\langle X_\alpha,g(0,\xi_\alpha^+)\right\rangle &= \mathbb{E}\left\langle X_\alpha,\chi_0(\xi_\alpha)-\chi_0(\xi_\alpha^+)\right\rangle\\ &= \mathbb{E}\left\langle X_\alpha^+-X_\alpha,\chi_0(\xi_\alpha^+)\right\rangle\\ &= -\alpha\mathbb{E}\left\langle h(X_\alpha)+g(X_\alpha,\xi_\alpha^+),\chi_0(\xi_\alpha^+)\right\rangle. \end{align}\] Thus Cauchy–Schwarz gives \[\mathbb{E}\left\langle X_\alpha,g(0,\xi_\alpha^+)\right\rangle \ge -C\alpha \left(\mathbb{E}\left\lVert h(X_\alpha)+g(X_\alpha,\xi_\alpha^+) \right\rVert^2\right)^{1/2} \ge -C\alpha\left(1+\mathbb{E}\left\langle X_\alpha,h(X_\alpha)\right\rangle\right).\] Finally, the mean-perturbation condition with \(y=0\) gives \[\begin{align} &\mathbb{E}\left\langle X_\alpha, h(X_\alpha)+g(X_\alpha,\xi_\alpha^+)-g(0,\xi_\alpha^+)\right\rangle\\ &\qquad\ge(1-\theta)\mathbb{E}\left\langle X_\alpha,h(X_\alpha)\right\rangle. \end{align}\] Combining the last two displays and subtracting \(\mathbb{E}\left\langle X_\alpha,h(X_\alpha)\right\rangle\) proves 63 . This proves the lemma. ◻

Proof of Lemma 6. By Lemma 23, the stationary square identity is justified. Since \(X_\alpha^+\stackrel d=X_\alpha\), \[0 = -2\alpha \mathbb{E}\left\langle X_\alpha,h(X_\alpha)+g(X_\alpha,\xi_\alpha^+)\right\rangle + \alpha^2 \mathbb{E}\left\lVert h(X_\alpha)+g(X_\alpha,\xi_\alpha^+) \right\rVert^2.\] Using 63 , Lemma 22, and 26 , we obtain \[(1-\theta)\mathbb{E}\left\langle X_\alpha,h(X_\alpha)\right\rangle \le C\alpha +C\alpha\mathbb{E}\left\langle X_\alpha,h(X_\alpha)\right\rangle.\] For small \(\alpha\), the last term is absorbed into the left-hand side, giving \(\mathbb{E}\left\langle X_\alpha,h(X_\alpha)\right\rangle\le C\alpha.\) This proves 35 , and Lemma 22 then gives 38 .

The local moment and outside-probability estimates now follow from the same drift lower bounds. Let \(r_0:=\min\{1,\frac{R_H}{2}\}\). Lemma 19 gives \[\left\lVert x \right\rVert^m\le C\left\langle x,h(x)\right\rangle, \qquad \left\lVert x \right\rVert\le r_0, \qquad \left\langle x,h(x)\right\rangle\ge c_*, \qquad \left\lVert x \right\rVert\ge r_0.\] Therefore \[\begin{align} \mathbb{E}[\left\lVert X_\alpha \right\rVert^m\mathbf{1}_{\{\left\lVert X_\alpha \right\rVert\le1\}}] &\le C\mathbb{E}\left\langle X_\alpha,h(X_\alpha)\right\rangle +\mathbb{P}(\left\lVert X_\alpha \right\rVert\ge r_0) \\ &\le C\alpha, \end{align}\] and \[\mathbb{P}(\left\lVert X_\alpha \right\rVert\ge1) \le \mathbb{P}(\left\lVert X_\alpha \right\rVert\ge r_0) \le c_*^{-1}\mathbb{E}\left\langle X_\alpha,h(X_\alpha)\right\rangle \le C\alpha.\] This proves all estimates in Lemma 6. ◻

Lemma 24. Under the assumptions of Lemma 6, for every \(\eta>0\), \[\label{eq:scaling-gradient-lindeberg-main} \mathbb{E}\left[ \frac{\left\lVert Y_\alpha^+-Y_\alpha \right\rVert^2}{\alpha^{2-2/m}} \mathbf{1}_{\{\left\lVert Y_\alpha^+-Y_\alpha \right\rVert>\eta\}} \right]\longrightarrow0.\tag{67}\] Moreover, for every fixed \(R>0\), the family \[\left\{ \left\lVert g(\alpha^{1/m}Y_\alpha,\xi_\alpha^+) \right\rVert^2 \mathbf{1}_{\{\left\lVert Y_\alpha \right\rVert\le R\}}:0<\alpha\le\alpha_0 \right\}\] is uniformly integrable.

Proof. We first prove the local uniform integrability statement. On the event \(\{\left\lVert Y_\alpha \right\rVert\le R\}\), the point \(x=\alpha^{1/m}Y_\alpha\) remains in a fixed compact set. If \(\beta=2\), the uniform \(x\)-Lipschitz bound 25 gives \(\left\lVert g(x,\xi_\alpha^+) \right\rVert^2 \le C_R\{1+\left\lVert g(0,\xi_\alpha^+) \right\rVert^2\},\) and \(\xi_\alpha^+\sim\pi_\Xi\), so 10 gives uniform integrability. If \(1\le\beta<2\), Assumption 2([[itm:N5]](#itm:N5){reference-type=“ref” reference=“itm:N5”}) gives a uniform exponential moment for the ratio \(\frac{\left\lVert g(x,\xi_\alpha^+) \right\rVert}{1+\left\lVert x \right\rVert^{\beta-1}},\) while \(1+\left\lVert x \right\rVert^{\beta-1}\) is bounded on the same compact set. This again gives uniform integrability.

For 67 , write \[Y_\alpha^+-Y_\alpha =-\alpha^{1-1/m} \{h(X_\alpha)+g(X_\alpha,\xi_\alpha^+)\}.\] It is enough to prove \[\label{eq:scaled-Falpha-lindeberg-proof} \mathbb{E}\left[ \left\lVert \mathsf G_\alpha \right\rVert^2 \mathbf{1}_{\{\alpha^{1-1/m}\left\lVert \mathsf G_\alpha \right\rVert>\eta\}} \right]\to0, \qquad \mathsf G_\alpha:=\mathsf G(X_\alpha,\xi_\alpha^+).\tag{68}\]

Consider first \(\beta=2\). The global Lipschitz bounds in Lemma 1 imply \(\left\lVert \mathsf G_\alpha \right\rVert \le C\left\lVert X_\alpha \right\rVert+\left\lVert g(0,\xi_\alpha^+) \right\rVert.\) Using \((a+b)^2\mathbf{1}_{\{a+b>t\}}\le4a^2\mathbf{1}_{\{a>t/2\}}+4b^2\mathbf{1}_{\{b>t/2\}}\), the contribution of \(g(0,\xi_\alpha^+)\) to 68 tends to zero by dominated convergence under \(\pi_\Xi\). For the \(X_\alpha\)-term, the threshold \(c\eta\alpha^{1/m-1}\) tends to infinity. For all sufficiently small \(\alpha\), it lies in the quadratic-tail region, where \(\left\lVert x \right\rVert^2\le C\left\langle x,h(x)\right\rangle\). Therefore \[\mathbb{E}\!\bigl[\left\lVert X_\alpha \right\rVert^2 \mathbf{1}_{\{\alpha^{1-1/m}C\left\lVert X_\alpha \right\rVert>\eta\}}\bigr] \le C\mathbb{E}\left\langle X_\alpha,h(X_\alpha)\right\rangle\to0\] by 35 . This verifies the Lindeberg condition in the quadratic-tail case.

Now assume \(1\le\beta<2\), and set \(w_\alpha:=1+\left\lVert X_\alpha \right\rVert^{\beta-1}.\) Since \(\left\lVert h(x) \right\rVert\le C(1+\left\lVert x \right\rVert^{\beta-1})\) globally, conditionally on \((X_\alpha,\xi_\alpha)\) we have \[\left\lVert \mathsf G_\alpha \right\rVert \le C w_\alpha\left\{1+ \frac{\left\lVert g(X_\alpha,\xi_\alpha^+) \right\rVert}{w_\alpha}\right\}.\] The exponential moment in Assumption 2([[itm:N5]](#itm:N5){reference-type=“ref” reference=“itm:N5”}) gives a uniform conditional second moment and a uniform conditional exponential tail for the normalized factor. Split according to \(w_\alpha\le \alpha^{-(m-1)/(2m)}\). On this event, the normalized factor must exceed a multiple of \(\alpha^{-(m-1)/(2m)}\), so the contribution is bounded by \(C\mathbb{E}w_\alpha^2 e^{-c\alpha^{-(m-1)/(2m)}}\to0.\) On the complementary event, the conditional second-moment bound gives a contribution at most \(C\mathbb{E}\bigl[w_\alpha^2 \mathbf{1}_{\{w_\alpha>\alpha^{-(m-1)/(2m)}\}}\bigr].\) To show that this expectation vanishes, fix a sufficiently large \(R_0\). The event \(\{w_\alpha>\alpha^{-(m-1)/(2m)}\}\) is eventually contained in \(\{\left\lVert X_\alpha \right\rVert\ge R_0\}\). On this region, Lemma 19 gives \((1+\left\lVert x \right\rVert^{\beta-1})^2\le C(1+\left\lVert x \right\rVert^\beta)\le C\left\langle x,h(x)\right\rangle\), after increasing \(R_0\) if necessary. Hence \[\mathbb{E}\bigl[w_\alpha^2 \mathbf{1}_{\{w_\alpha>\alpha^{-(m-1)/(2m)}\}}\bigr] \le C\mathbb{E}\left\langle X_\alpha,h(X_\alpha)\right\rangle\to0.\] This proves 68 , and hence 67 . ◻

7.5 Proof of Lemma 7↩︎

Proof of Lemma 7. Let \(u\) be the Poisson solution from Lemma 21; thus \[u-Qu=f, \qquad u\in L^2(\pi_\Xi).\] Since \(\psi(Y_\alpha)\) is measurable with respect to \((X_\alpha,\xi_\alpha)\), \[\begin{align} \mathbb{E}[\psi(Y_\alpha)f(\xi_\alpha)] &= \mathbb{E}[\psi(Y_\alpha)u(\xi_\alpha)] - \mathbb{E}[\psi(Y_\alpha)Qu(\xi_\alpha)] \\ &= \mathbb{E}[\psi(Y_\alpha)u(\xi_\alpha)] - \mathbb{E}[\psi(Y_\alpha)u(\xi_\alpha^+)]. \end{align}\] Add and subtract \(\psi(Y_\alpha^+)u(\xi_\alpha^+)\). Since \((Y_\alpha^+,\xi_\alpha^+)\stackrel d=(Y_\alpha,\xi_\alpha),\) the first and third terms cancel, giving \[\mathbb{E}[\psi(Y_\alpha)f(\xi_\alpha)] = \mathbb{E}[(\psi(Y_\alpha^+)-\psi(Y_\alpha))u(\xi_\alpha^+)].\] Therefore, by Cauchy–Schwarz, \[\left\lvert \mathbb{E}[\psi(Y_\alpha)f(\xi_\alpha)] \right\rvert \le \mathop{\mathrm{Lip}}(\psi) \bigl(\mathbb{E}\left\lVert Y_\alpha^+-Y_\alpha \right\rVert^2\bigr)^{1/2} \bigl(\mathbb{E}\left\lvert u(\xi_\alpha^+) \right\rvert^2\bigr)^{1/2}.\] The second factor is finite and independent of \(\alpha\), because \(\xi_\alpha^+\sim\pi_\Xi\). For the first factor, \[Y_\alpha^+-Y_\alpha = -\alpha^{1-1/m} \{h(X_\alpha)+g(X_\alpha,\xi_\alpha^+)\}.\] Lemma 6 gives \(\mathbb{E}\left\lVert Y_\alpha^+-Y_\alpha \right\rVert^2 \le C\alpha^{2-2/m}.\) This proves the desired bound. ◻

7.6 Proof of Lemma 8↩︎

By 26 , 35 , and Lemma 22, \[\label{eq:scaling-generator-moment-bounds} \mathbb{E}\left\lVert h(X_\alpha) \right\rVert^2\le C\alpha, \qquad \sup_{0<\alpha\le\alpha_0} \mathbb{E}\left\lVert g(X_\alpha,\xi_\alpha^+) \right\rVert^2<\infty, \qquad \mathbb{E}\bigl[\left\lVert h(X_\alpha) \right\rVert \left\lVert g(X_\alpha,\xi_\alpha^+) \right\rVert\bigr]\longrightarrow0.\tag{69}\] The last conclusion follows from the first two by Cauchy–Schwarz.

For the covariance calculation below, define, for \(x\in\mathbb{R}^d\), \[M_x(\zeta) := g_x(\zeta)g_x(\zeta)^\top +g_x(\zeta)\chi_x(\zeta)^\top +\chi_x(\zeta)g_x(\zeta)^\top, \qquad \zeta\in\Xi.\] The next two lemmas justify replacing \(M_{\alpha^{1/m}Y_\alpha}\) by \(M_0\) on compact sets and handle integrable functions of the driving chain.

Lemma 25. Under Assumptions 3 and 4, for every compact set \(K\subset\mathbb{R}^d\), one has \(QM_0\in L^1(\pi_\Xi)\) and \[\label{eq:Q-Mx-M0-L1-compact} \mathbb{E}\left[ \left\lVert QM_{\alpha^{1/m}Y_\alpha}(\xi_\alpha)-QM_0(\xi_\alpha) \right\rVert \mathbf{1}_{\{Y_\alpha\in K\}} \right] \longrightarrow0.\tag{70}\]

Proof. The uniform \(x\)-Lipschitz bound 25 gives \(\left\lVert g_x(\zeta)-g_0(\zeta) \right\rVert\le C\left\lVert x \right\rVert.\) Fix \(p\in(0,1)\). The estimates for the solution of the Poisson equation in 61 and 62 give \[\left\lVert \chi_x(\zeta) \right\rVert+\left\lVert \chi_0(\zeta) \right\rVert\le C R_\chi(\zeta), \qquad \left\lVert \chi_x(\zeta)-\chi_0(\zeta) \right\rVert \le C_p\left\lVert x \right\rVert^pR_\chi(\zeta)^{1-p}.\] Expanding \(M_x-M_0\), for \(\left\lVert x \right\rVert\le1\) we obtain \(\left\lVert M_x(\zeta)-M_0(\zeta) \right\rVert \le C_p(\left\lVert x \right\rVert+\left\lVert x \right\rVert^p)F(\zeta),\) where \[F(\zeta):= 1+\left\lVert g_0(\zeta) \right\rVert+R_\chi(\zeta) +\left\lVert g_0(\zeta) \right\rVert R_\chi(\zeta).\] The stationary second moments in 10 and Cauchy–Schwarz imply \(F\in L^1(\pi_\Xi)\). In particular, \(M_0\in L^1(\pi_\Xi)\), and invariance gives \(QM_0\in L^1(\pi_\Xi)\). If \(K\) is compact, then, for all sufficiently small \(\alpha\), \(\alpha^{1/m}\sup_{y\in K}\left\lVert y \right\rVert\le1\). On \(\{Y_\alpha\in K\}\), the preceding pointwise bound, invariance of \(\pi_\Xi\), and \(\xi_\alpha\sim\pi_\Xi\) give \[\begin{align} &\mathbb{E}\left[ \left\lVert QM_{\alpha^{1/m}Y_\alpha}(\xi_\alpha)-QM_0(\xi_\alpha) \right\rVert \mathbf{1}_{\{Y_\alpha\in K\}}\right]\\ &\qquad\le C_{p,K}(\alpha^{1/m}+\alpha^{p/m})\mathbb{E}QF(\xi_\alpha) = C_{p,K}(\alpha^{1/m}+\alpha^{p/m})\int_\Xi F\,d\pi_\Xi \longrightarrow0. \end{align}\] This proves 70 . ◻

Lemma 26. Assume Assumptions 3 and 4. Let \(\alpha_k\downarrow0\) and suppose \(Y_{\alpha_k}\Rightarrow\nu\). If \(\psi:\mathbb{R}^d\to\mathbb{R}\) is bounded and globally Lipschitz, and if \(\Psi\in L^1(\pi_\Xi)\), then \[\mathbb{E}[\psi(Y_{\alpha_k})\Psi(\xi_{\alpha_k})] \longrightarrow \left(\int_{\mathbb{R}^d}\psi(y)\,\nu(dy)\right) \left(\int_\Xi\Psi(\xi)\,\pi_\Xi(d\xi)\right).\]

Proof. The constant part of \(\Psi\) is handled by weak convergence of \(Y_{\alpha_k}\), because the \(\Xi\)-marginal of \(\pi_{\alpha_k}\) is \(\pi_\Xi\) by Lemma 20. It remains to consider centered \(\Psi\), i.e. \(\pi_\Xi(\Psi)=0\). Bounded Lipschitz functions are dense in \(L^1(\pi_\Xi)\) on the closed Euclidean set \(\Xi\). Thus, for every \(\varepsilon>0\), choose a bounded Lipschitz \(f\) such that \(\left\lVert \Psi-f \right\rVert_{L^1(\pi_\Xi)}<\varepsilon\). Replacing \(f\) by \(f-\pi_\Xi(f)\) increases this error by at most another \(\varepsilon\), so we may assume \(\pi_\Xi(f)=0\). Since \(\xi_{\alpha_k}\sim\pi_\Xi\), \[\left\lvert \mathbb{E}[\psi(Y_{\alpha_k})(\Psi-f)(\xi_{\alpha_k})] \right\rvert \le 2\left\lVert \psi \right\rVert_\infty\varepsilon.\] Lemma 7 gives \(\mathbb{E}[\psi(Y_{\alpha_k})f(\xi_{\alpha_k})]\to0\). Letting \(\varepsilon\downarrow0\) proves the centered case and hence the lemma. ◻

Lemma 27. Under Assumptions 3 and 4, for every \(\varphi\in C_c^3(\mathbb{R}^d)\), \[\label{eq:prelimit-generator-identity} 0= -\mathbb{E}\left\langle h_\alpha(Y_\alpha),\nabla\varphi(Y_\alpha)\right\rangle +\frac{1}{2}\mathbb{E}\mathop{\mathrm{tr}}\!\left( QM_{\alpha^{1/m}Y_\alpha}(\xi_\alpha)\nabla^2\varphi(Y_\alpha) \right)+o(1),\tag{71}\] where the \(o(1)\) is deterministic and tends to zero as \(\alpha\downarrow0\).

Proof. Write \[G_\alpha:=g_{X_\alpha}(\xi_\alpha^+), \qquad \mathsf G_\alpha:=\mathsf G(X_\alpha,\xi_\alpha^+).\] Then \[Y_\alpha^+-Y_\alpha =-\alpha^{2-2/m}h_\alpha(Y_\alpha)-\alpha^{1-1/m}G_\alpha =-\alpha^{1-1/m}\mathsf G_\alpha.\] Let \(K_0:=\mathop{\mathrm{supp}}(\nabla\varphi)\), and let \(K_1:=\{y\in\mathbb{R}^d:\operatorname{dist}(y,K_0)\le1\}.\) If \(K_0=\varnothing\), the conclusion is immediate, so assume otherwise. Define \[\widetilde{\chi}_{\alpha,y}(\xi) := \begin{cases} \chi_{\alpha^{1/m}y}(\xi),&y\in K_1,\\ 0,&y\notin K_1. \end{cases}\] Set \(A_\alpha(y,\xi) :=\left\langle \nabla\varphi(y),\widetilde{\chi}_{\alpha,y}(\xi)\right\rangle.\) This is well defined because \(\nabla\varphi=0\) outside \(K_0\), and \(\left\lvert A_\alpha(y,\xi) \right\rvert\le C\left\lVert \nabla\varphi \right\rVert_\infty R_\chi(\xi)\). Introduce the localized perturbed test function \[\label{eq:perturbed-test-function-GK} \varphi_\alpha(y,\xi) := \varphi(y)-\alpha^{1-1/m}A_\alpha(y,\xi).\tag{72}\] The integrability required below follows from 38 , \(R_\chi\in L^2(\pi_\Xi)\), and 61 . Stationarity gives \[\label{eq:perturbed-stationary-identity} \mathbb{E}[\varphi_\alpha(Y_\alpha^+,\xi_\alpha^+)-\varphi_\alpha(Y_\alpha,\xi_\alpha)]=0.\tag{73}\]

First consider the uncorrected part. Taylor’s formula and \(Y_\alpha^+-Y_\alpha=-\alpha^{2-2/m}h_\alpha(Y_\alpha)-\alpha^{1-1/m}G_\alpha\) give \[\label{eq:uncorrected-taylor-generator} \begin{align} \frac{\varphi(Y_\alpha^+)-\varphi(Y_\alpha)}{\alpha^{2-2/m}} &=-\left\langle h_\alpha(Y_\alpha),\nabla\varphi(Y_\alpha)\right\rangle -\alpha^{1/m-1}\left\langle \nabla\varphi(Y_\alpha),G_\alpha\right\rangle \\ &\quad+ \frac{1}{2}\mathop{\mathrm{tr}}\!\left(G_\alpha G_\alpha^\top\nabla^2\varphi(Y_\alpha)\right) +r_\alpha^{(1)}, \end{align}\tag{74}\] with \(\mathbb{E}|r_\alpha^{(1)}|\to0\). The terms in the quadratic expansion containing \(h_\alpha\) are negligible because \[\alpha^{2-2/m}\left\lVert h_\alpha(Y_\alpha) \right\rVert^2=\left\lVert h(X_\alpha) \right\rVert^2, \qquad \alpha^{1-1/m}\left\lVert h_\alpha(Y_\alpha) \right\rVert\left\lVert G_\alpha \right\rVert =\left\lVert h(X_\alpha) \right\rVert\left\lVert G_\alpha \right\rVert,\] and 69 makes both expectations tend to zero. The third-order Taylor remainder is controlled by the uniform continuity of \(\nabla^2\varphi\): for every \(\eta>0\), after division by \(\alpha^{2-2/m}\) its absolute value is bounded by \[\eta\frac{\left\lVert Y_\alpha^+-Y_\alpha \right\rVert^2}{\alpha^{2-2/m}} +C_\eta \frac{\left\lVert Y_\alpha^+-Y_\alpha \right\rVert^2}{\alpha^{2-2/m}} \mathbf{1}_{\{\left\lVert Y_\alpha^+-Y_\alpha \right\rVert>\eta\}}.\] The first term is bounded in expectation uniformly in \(\alpha\) by 38 ; the second tends to zero by 67 . Sending \(\eta\downarrow0\) proves \(\mathbb{E}|r_\alpha^{(1)}|\to0\).

Now expand the term involving the solution of the Poisson equation in 72 . In \[-\alpha^{1/m-1}\mathbb{E}\left[ A_\alpha(Y_\alpha^+,\xi_\alpha^+) -A_\alpha(Y_\alpha,\xi_\alpha) \right],\] add and subtract \(A_\alpha(Y_\alpha,\xi_\alpha^+)\). Since \(A_\alpha(Y_\alpha,\cdot)\) can be nonzero only when \(Y_\alpha\in K_0\), the first resulting increment is \[\begin{align} -\alpha^{1/m-1} \mathbb{E}\left[A_\alpha(Y_\alpha,\xi_\alpha^+) -A_\alpha(Y_\alpha,\xi_\alpha)\right] &=\alpha^{1/m-1} \mathbb{E}\left\langle \nabla\varphi(Y_\alpha),Qg_{X_\alpha}(\xi_\alpha)\right\rangle. \end{align}\] Here the solution of the Poisson equation is evaluated at \(X_\alpha=\alpha^{1/m}Y_\alpha\) with \(Y_\alpha\in K_0\). The last display follows from \(\chi_x-Q\chi_x=Qg_x\) and cancels the conditional mean of the singular term in 74 , since \(\mathbb{E}[G_\alpha\mid Y_\alpha,\xi_\alpha]=Qg_{X_\alpha}(\xi_\alpha)\).

It remains to control \[-\alpha^{1/m-1}\mathbb{E}\left[ A_\alpha(Y_\alpha^+,\xi_\alpha^+) -A_\alpha(Y_\alpha,\xi_\alpha^+)\right].\] Let \(E_\alpha:=\{\left\lVert Y_\alpha^+-Y_\alpha \right\rVert\le1/2\}\). By 67 , \[\mathbb{P}(E_\alpha^c) \le4\mathbb{E}\left[\left\lVert Y_\alpha^+-Y_\alpha \right\rVert^2\mathbf{1}_{E_\alpha^c}\right] =o(\alpha^{2-2/m}).\] The bound on \(A_\alpha\), Cauchy–Schwarz, and \(\xi_\alpha^+\sim\pi_\Xi\) therefore give \[\label{eq:localized-poisson-large-jump} \begin{align} &\alpha^{1/m-1}\mathbb{E}\left[ \left|A_\alpha(Y_\alpha^+,\xi_\alpha^+) -A_\alpha(Y_\alpha,\xi_\alpha^+)\right| \mathbf{1}_{E_\alpha^c}\right]\\ &\qquad\le C\alpha^{1/m-1} \bigl(\mathbb{E}_{\pi_\Xi}R_\chi^2\bigr)^{1/2} \mathbb{P}(E_\alpha^c)^{1/2}=o(1). \end{align}\tag{75}\]

On \(E_\alpha\), if either term involving \(A_\alpha\) is nonzero, then one of \(Y_\alpha,Y_\alpha^+\) belongs to \(K_0\), and both belong to \(K_1\). Hence \[\begin{align} &A_\alpha(Y_\alpha^+,\xi_\alpha^+) -A_\alpha(Y_\alpha,\xi_\alpha^+)\\ &\quad= \left\langle \nabla\varphi(Y_\alpha^+)-\nabla\varphi(Y_\alpha),\widetilde{\chi}_{\alpha,Y_\alpha}(\xi_\alpha^+)\right\rangle +\left\langle \nabla\varphi(Y_\alpha^+),\widetilde{\chi}_{\alpha,Y_\alpha^+}(\xi_\alpha^+) -\widetilde{\chi}_{\alpha,Y_\alpha}(\xi_\alpha^+)\right\rangle. \end{align}\]

For the first term on \(E_\alpha\), Taylor’s formula gives \[\begin{align} &-\alpha^{1/m-1} \mathbb{E}\left\langle \nabla\varphi(Y_\alpha^+)-\nabla\varphi(Y_\alpha),\widetilde{\chi}_{\alpha,Y_\alpha}(\xi_\alpha^+)\right\rangle\mathbf{1}_{E_\alpha} \\ &\qquad= \mathbb{E}\mathop{\mathrm{tr}}\!\left(G_\alpha \widetilde{\chi}_{\alpha,Y_\alpha}(\xi_\alpha^+)^\top \nabla^2\varphi(Y_\alpha)\right)+o(1). \end{align}\] The contribution of \(E_\alpha^c\) to the displayed leading term tends to zero. Indeed, \[\mathbb{E}[\left\lVert G_\alpha \right\rVert^2\mathbf{1}_{E_\alpha^c}] \le2\mathbb{E}[\left\lVert \mathsf G_\alpha \right\rVert^2\mathbf{1}_{E_\alpha^c}] +2\mathbb{E}\left\lVert h(X_\alpha) \right\rVert^2\longrightarrow0\] by 67 and 69 ; Cauchy–Schwarz with \(R_\chi\in L^2(\pi_\Xi)\) then applies. Also, \(\nabla^2\varphi(Y_\alpha)\ne0\) implies \(Y_\alpha\in K_0\), so the solution appearing in the leading term is well defined.

The contribution of \(-\alpha^{2-2/m}h_\alpha\) is \(O(\alpha^{1-1/m})\): the factor \(\nabla^2\varphi(Y_\alpha)\) restricts this term to a fixed compact set, where \(h_\alpha\) is locally uniformly bounded by the expansion in Assumption 3([[itm:H4]](#itm:H4){reference-type=“ref” reference=“itm:H4”}), while \(\mathbb{E}R_\chi(\xi_\alpha^+)<\infty\).

To control the Taylor remainder, let \(\omega_\varphi\) be a bounded modulus of continuity of \(\nabla^2\varphi\). Since \(Y_\alpha^+-Y_\alpha=-\alpha^{1-1/m}\mathsf G_\alpha\), the absolute expectation of the remainder after division by \(\alpha^{1-1/m}\) is at most \[C\mathbb{E}\!\left[ \omega_\varphi(\alpha^{1-1/m}\left\lVert \mathsf G_\alpha \right\rVert) \left\lVert \mathsf G_\alpha \right\rVert R_\chi(\xi_\alpha^+)\right].\] For every \(\eta>0\), the contribution of \(\{\alpha^{1-1/m}\left\lVert \mathsf G_\alpha \right\rVert\le\eta\}\) is bounded by \[C\omega_\varphi(\eta) \bigl(\mathbb{E}\left\lVert \mathsf G_\alpha \right\rVert^2\bigr)^{1/2} \bigl(\mathbb{E}_{\pi_\Xi}R_\chi^2\bigr)^{1/2} \le C\omega_\varphi(\eta).\] On the complementary event, Cauchy–Schwarz bounds the contribution by \[C\left( \mathbb{E}\!\left[\left\lVert \mathsf G_\alpha \right\rVert^2 \mathbf{1}_{\{\alpha^{1-1/m}\left\lVert \mathsf G_\alpha \right\rVert>\eta\}}\right] \right)^{1/2} \bigl(\mathbb{E}_{\pi_\Xi}R_\chi^2\bigr)^{1/2},\] which tends to zero by 67 . Sending \(\eta\downarrow0\) proves the \(o(1)\) remainder in the gradient increment.

For the second term on \(E_\alpha\), which changes the solution of the Poisson equation from \(\alpha^{1/m}Y_\alpha\) to \(\alpha^{1/m}Y_\alpha^+\), fix the exponent \(p\in(1-1/m,1)\) in 62 . That estimate gives \[\begin{align} &\alpha^{1/m-1} \mathbb{E}\left|\left\langle \nabla\varphi(Y_\alpha^+),\widetilde{\chi}_{\alpha,Y_\alpha^+}(\xi_\alpha^+) -\widetilde{\chi}_{\alpha,Y_\alpha}(\xi_\alpha^+)\right\rangle\right| \mathbf{1}_{E_\alpha} \\ &\qquad\le C_p\alpha^{p-(1-1/m)} \mathbb{E}[\left\lVert \mathsf G_\alpha \right\rVert^p R_\chi(\xi_\alpha^+)^{1-p}]\\ &\qquad\le C_p\alpha^{p-(1-1/m)} =o(1). \end{align}\] Here the second inequality follows from weighted AM–GM, \(a^pb^{1-p}\le pa+(1-p)b\), the uniform second moment of \(\mathsf G_\alpha\), and \(R_\chi\in L^2(\pi_\Xi)\).

Combining the preceding identities with 73 gives \[\begin{align} 0&=-\mathbb{E}\left\langle h_\alpha(Y_\alpha),\nabla\varphi(Y_\alpha)\right\rangle \\ &\quad+\frac{1}{2}\mathbb{E}\mathop{\mathrm{tr}}\!\left( \Bigl[G_\alpha G_\alpha^\top +G_\alpha\widetilde{\chi}_{\alpha,Y_\alpha}(\xi_\alpha^+)^\top +\widetilde{\chi}_{\alpha,Y_\alpha}(\xi_\alpha^+)G_\alpha^\top\Bigr] \nabla^2\varphi(Y_\alpha) \right)+o(1). \end{align}\] After multiplication by \(\nabla^2\varphi(Y_\alpha)\), conditioning on \((Y_\alpha,\xi_\alpha)\) identifies the conditional mean with \(QM_{\alpha^{1/m}Y_\alpha}(\xi_\alpha)\nabla^2\varphi(Y_\alpha)\). This proves 71 . ◻

Proof of Lemma 8. Let \(\varphi\in C_c^3(\mathbb{R}^d)\) and let \(K:=\mathop{\mathrm{supp}}(\nabla^2\varphi)\). Lemma 27 gives 71 . Along the subsequence \(\alpha_k\downarrow0\), local uniform convergence of \(h_\alpha\) and weak convergence give \[\label{eq:drift-limit-markov-new} \mathbb{E}\left\langle h_{\alpha_k}(Y_{\alpha_k}),\nabla\varphi(Y_{\alpha_k})\right\rangle \longrightarrow \int_{\mathbb{R}^d}\left\langle h_0(y),\nabla\varphi(y)\right\rangle\,\nu(dy).\tag{76}\] Here \(\nabla\varphi\) is compactly supported and \(h_\alpha\to h_0\) locally uniformly by Assumption 3([[itm:H4]](#itm:H4){reference-type=“ref” reference=“itm:H4”}).

For the covariance term, 70 replaces \(QM_{\alpha_k^{1/m}Y_{\alpha_k}}(\xi_{\alpha_k})\) by \(QM_0(\xi_{\alpha_k})\) on \(K\). Each entry of \(QM_0\) belongs to \(L^1(\pi_\Xi)\) by Lemma 25. Applying Lemma 26 entrywise, with the bounded Lipschitz entries of \(\nabla^2\varphi\), yields \[\label{eq:QM0-decorrelation-limit} \mathbb{E}\mathop{\mathrm{tr}}\!\left(QM_0(\xi_{\alpha_k})\nabla^2\varphi(Y_{\alpha_k})\right) \longrightarrow \int_{\mathbb{R}^d}\mathop{\mathrm{tr}}\!\left(\bar M\nabla^2\varphi(y)\right)\nu(dy),\tag{77}\] where \(\bar M:=\int_\Xi M_0(\xi)\,\pi_\Xi(d\xi).\) It remains only to identify \(\bar M\) with \(\Sigma\). Let \((\xi_0,\xi_1)\) be two successive values of the stationary driving chain and write \(f=g_0\), \(\chi=\chi_0\). Since \(\chi-Q\chi=Qf\), \(\mathbb{E}[f(\xi_1)+\chi(\xi_1)\mid\xi_0]=Qf(\xi_0)+Q\chi(\xi_0)=\chi(\xi_0).\) With \(D_0=f(\xi_1)+\chi(\xi_1)-\chi(\xi_0)\), expansion and the preceding conditional identity give \[\begin{align} \Sigma &=\mathbb{E}[D_0D_0^\top] \\ &=\mathbb{E}\bigl[f(\xi_1)f(\xi_1)^\top +f(\xi_1)\chi(\xi_1)^\top +\chi(\xi_1)f(\xi_1)^\top\bigr] \\ &=\int_\Xi M_0(\xi)\,\pi_\Xi(d\xi)=\bar M. \end{align}\] Combining 71 , 76 , and 77 proves 39 . ◻

7.7 Proof of Lemma 9↩︎

Lemma 28. Under Assumption 3 (specifically, ([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}) and ([[itm:H4]](#itm:H4){reference-type=“ref” reference=“itm:H4”})), \[\left\langle y,h_0(y)\right\rangle \ge \frac{c_{\rm in}}{m-1}\left\lVert y \right\rVert^m, \qquad y\in\mathbb{R}^d.\]

Proof. The claim is trivial at \(y=0\). For \(y\ne0\), set \(x=\alpha^{1/m}y\). For all sufficiently small \(\alpha\), \(\left\lVert x \right\rVert\le \frac{R_H}{2}\). The identity \(h(x)=\int_0^1\nabla^2H(tx)x\,dt\) and Assumption 1([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}) give \(\left\langle x,h(x)\right\rangle \ge \frac{c_{\rm in}}{m-1}\left\lVert x \right\rVert^m.\) Equivalently, \[\left\langle y,\alpha^{-(m-1)/m}h(\alpha^{1/m}y)\right\rangle = \alpha^{-1}\left\langle x,h(x)\right\rangle \ge \frac{c_{\rm in}}{m-1}\left\lVert y \right\rVert^m.\] Letting \(\alpha\downarrow0\) and using the expansion in Assumption 3([[itm:H4]](#itm:H4){reference-type=“ref” reference=“itm:H4”}) together with the homogeneity of \(h_0\) gives the result. ◻

Lemma 29. Under Assumption 3 (specifically, ([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}) and ([[itm:H4]](#itm:H4){reference-type=“ref” reference=“itm:H4”})), there exists \(c_m>0\) such that, for all \(y,z\in\mathbb{R}^d\), \(\left\langle y-z,h_0(y)-h_0(z)\right\rangle \ge c_m\left\lVert y-z \right\rVert^m.\)

Proof. The claim is trivial when \(y=z\). Set \(e:=y-z\ne0\). For sufficiently small \(\alpha\), the segment between \(\alpha^{1/m}z\) and \(\alpha^{1/m}y\) is contained in \(\{\left\lVert x \right\rVert\le \frac{R_H}{2}\}\). It can pass through the origin for at most one parameter value, which does not affect the integral below. Using Assumption 1([[itm:H2]](#itm:H2){reference-type=“ref” reference=“itm:H2”}) along the segment gives \[\begin{align} &\left\langle e,\alpha^{-(m-1)/m}\{h(\alpha^{1/m}y)-h(\alpha^{1/m}z)\}\right\rangle \\ &\qquad = \alpha^{-1}\left\langle \alpha^{1/m}e,h(\alpha^{1/m}y)-h(\alpha^{1/m}z)\right\rangle \\ &\qquad = \alpha^{-1}\int_0^1 \left\langle \alpha^{1/m}e,\nabla^2H(\alpha^{1/m}(z+te))\alpha^{1/m}e\right\rangle\,dt \\ &\qquad \ge c_{\rm in}\left\lVert e \right\rVert^2 \int_0^1\left\lVert z+te \right\rVert^{m-2}\,dt. \end{align}\] We use the elementary segment bound \[\label{eq:elementary-segment-bound-power} \inf_{\left\lVert v \right\rVert=1}\inf_{w\in\mathbb{R}^d} \int_0^1\left\lVert w+tv \right\rVert^{m-2}\,dt>0.\tag{78}\] To prove 78 , the case \(m=2\) is immediate. For \(m>2\), decompose \(w=a v+w_\perp\), with \(w_\perp\perp v\). Then \[\int_0^1\left\lVert w+tv \right\rVert^{m-2}\,dt \ge \int_0^1 |a+t|^{m-2}\,dt.\] The last integral is minimized over \(a\in\mathbb{R}\) when the interval \([a,a+1]\) is centered at the origin, and its minimum is \(2\int_0^{1/2}t^{m-2}\,dt>0\). Thus 78 holds and gives \(\int_0^1\left\lVert z+te \right\rVert^{m-2}\,dt \ge c\left\lVert e \right\rVert^{m-2}.\) Letting \(\alpha\downarrow0\) and using the expansion in Assumption 3([[itm:H4]](#itm:H4){reference-type=“ref” reference=“itm:H4”}) together with the homogeneity of \(h_0\) gives the result. ◻

Proof of Lemma 9. The drift \(-h_0\) is locally Lipschitz because \(H_0\in C^2\), and it has polynomial growth by homogeneity. Thus the SDE has a pathwise unique maximal strong solution up to its explosion time. Let \[W(y):=1+\left\lVert y \right\rVert^2, \qquad \mathcal{L}\varphi(y) :=-\left\langle h_0(y),\nabla\varphi(y)\right\rangle +\frac{1}{2}\mathop{\mathrm{tr}}(\Sigma\nabla^2\varphi(y)).\] By Lemma 28, \[\label{eq:limit-diffusion-LW} \mathcal{L}W(y) =-2\left\langle y,h_0(y)\right\rangle+\mathop{\mathrm{tr}}(\Sigma) \le C-c\left\lVert y \right\rVert^m.\tag{79}\] If \(\tau_R:=\inf\{t:\left\lVert Y_t \right\rVert\ge R\}\), Itô’s formula applied to \(W(Y_{t\wedge\tau_R})\) gives \(\mathbb{E}_y W(Y_{t\wedge\tau_R})\le W(y)+Ct.\) On \(\{\tau_R\le t\}\), the left-hand side is at least \(1+R^2\). Hence \(\mathbb{P}_y(\tau_R\le t)\le (W(y)+Ct)/(1+R^2)\to0\), which proves nonexplosion. Applying Itô’s formula without stopping and using 79 gives \(\frac{1}{T}\int_0^T \mathbb{E}_y\left\lVert Y_t \right\rVert^m\,dt \le C+\frac{W(y)}{cT}.\) The time-averaged laws \(T^{-1}\int_0^T P_t(y,\cdot)\,dt\) are therefore tight. Since the nonexplosive locally Lipschitz SDE is Feller, the Krylov–Bogoliubov argument gives at least one invariant probability law.

We next show that every invariant probability law has finite \(m\)-moment. Let \(\theta_n\in C^2([0,\infty))\) be nondecreasing and concave, with \(0\le\theta_n'\le1\), \(\theta_n'(s)=1\) for \(s\le n\), \(\theta_n'(s)=0\) for \(s\ge2n\), \(\theta_n''\le0\), and \(\theta_n'(s)\uparrow1\) for every fixed \(s\). Set \(W_n(y):=\theta_n(W(y))\). Then \(W_n\) is bounded with bounded first and second derivatives. If \(\mu\) is invariant, then \(\int \mathcal{L}W_n\,d\mu=0\). Moreover, \[\mathcal{L}W_n(y) =\theta_n'(W(y))\mathcal{L}W(y) +\frac{1}{2}\theta_n''(W(y)) \left\lVert \Sigma^{1/2}\nabla W(y) \right\rVert^2 \le \theta_n'(W(y))(C-c\left\lVert y \right\rVert^m),\] because \(\theta_n''\le0\) and \(\Sigma\succeq0\). Hence \[c\int \theta_n'(W(y))\left\lVert y \right\rVert^m\,\mu(dy) \le C\int \theta_n'(W(y))\,\mu(dy) \le C.\] Letting \(n\to\infty\) and using monotone convergence gives \(\int \left\lVert y \right\rVert^m\,\mu(dy)<\infty.\) In particular every invariant law has finite second moment.

Let \(\mu\) and \(\nu\) be invariant laws. Choose a coupling \((Y_0,\widetilde{Y}_0)\) with finite second moment and drive both solutions by the same Brownian motion. Since the diffusion coefficient is constant, \(\Delta_t:=Y_t-\widetilde{Y}_t\) satisfies, up to localization, \[\frac{d}{dt}\left\lVert \Delta_t \right\rVert^2 =-2\left\langle \Delta_t,h_0(Y_t)-h_0(\widetilde{Y}_t)\right\rangle \le -2c_m\left\lVert \Delta_t \right\rVert^m\] by Lemma 29. With \(D(t):=\mathbb{E}\left\lVert \Delta_t \right\rVert^2\), Fatou’s lemma and Jensen’s inequality give \(D'(t)\le -2c_m\mathbb{E}\left\lVert \Delta_t \right\rVert^m \le -2c_mD(t)^{m/2}.\) Thus \(D(t)\to0\), exponentially if \(m=2\) and polynomially if \(m>2\). The marginals remain \(\mu\) and \(\nu\), so \(W_2(\mu,\nu)^2\le D(t)\to0\). Hence \(\mu=\nu\). Denote the unique invariant law by \(\nu_\infty\).

Finally, suppose \(\rho\) satisfies the stationary weak equation for every \(\varphi\in C_c^3(\mathbb{R}^d)\). It then holds on \(C_c^\infty(\mathbb{R}^d)\). The martingale problem for \(\mathcal{L}\) is well posed because the SDE is pathwise unique and nonexplosive. The Echeverria criterion therefore implies that \(\rho\) is invariant; see [47]. By uniqueness, \(\rho=\nu_\infty\). ◻

8 Separable extensions↩︎

This appendix proves the separable results from Section 2.3. We sum the one-dimensional contractions for the constant-stepsize result and apply the generator argument to the coordinates that remain nonzero under the common normalization.

8.1 Proof of Corollary 2↩︎

Proof of Corollary 2. Write \(d_{i,\alpha}:=d_{V_{\alpha,i},\alpha}^{(i)}\). For each coordinate \(i\), the one-dimensional proof gives a constant \(c_i>0\) and, after taking the minimum of the finitely many stepsize thresholds, the synchronous one-step estimate \[\mathbb{E}\Bigl[d_{i,\alpha}\bigl(F_{U_1}^{(i)}(x_i,\xi),F_{U_1}^{(i)}(y_i,\eta)\bigr)\Bigr] \le (1-c_i\alpha^{m_i-1})\,d_{i,\alpha}\bigl((x_i,\xi),(y_i,\eta)\bigr)\] for the coordinate map \[F_u^{(i)}(x_i,\xi) := \Bigl(x_i-\alpha\bigl(h_i(x_i)+g_i(x_i,\Phi(\xi,u))\bigr),\Phi(\xi,u)\Bigr).\] The \(i\)th coordinate projection of the full separable map is \((x,\xi)\mapsto F_u^{(i)}(x_i,\xi).\) Since the same innovation \(U_1\) drives every coordinate, summing the coordinatewise estimates gives \[\mathbb{E}\Bigl[d_{{\rm sep},\alpha}\bigl(F_{U_1}(x,\xi),F_{U_1}(y,\eta)\bigr)\Bigr] \le (1-c_0\alpha^{m_{\max}-1})\,d_{{\rm sep},\alpha}\bigl((x,\xi),(y,\eta)\bigr),\] where \(c_0:=\min_{1\le i\le d}c_i>0\). Indeed, for \(0<\alpha\le1\), \[c_i\alpha^{m_i-1} \ge c_0\alpha^{m_{\max}-1}, \qquad 1\le i\le d.\] This is the required one-step contraction in the additive separable metric.

It remains to check the reference-point integrability. The one-dimensional hypotheses may use different reference points. Let \(\xi_\star^{(i)}\) be one for coordinate \(i\), and fix \(\xi^\circ\in\Xi\). By Assumption 2([[itm:N1]](#itm:N1){reference-type=“ref” reference=“itm:N1”}), \[\mathbb{E}\left\lVert \Phi(\xi^\circ,U_1)-\xi^\circ \right\rVert \le (1+\rho_\Xi)\left\lVert \xi^\circ-\xi_\star^{(i)} \right\rVert + \mathbb{E}\left\lVert \Phi(\xi_\star^{(i)},U_1)-\xi_\star^{(i)} \right\rVert<\infty.\] If \(\beta_i=2\), then Assumption 2([[itm:N2]](#itm:N2){reference-type=“ref” reference=“itm:N2”}) gives \[\mathbb{E}\left\lvert g_i(0,\Phi(\xi^\circ,U_1)) \right\rvert \le \mathbb{E}\left\lvert g_i(0,\Phi(\xi_\star^{(i)},U_1)) \right\rvert + L_{g,\Phi}\left\lVert \xi^\circ-\xi_\star^{(i)} \right\rVert<\infty.\] Thus \(\xi^\circ\) works for every coordinate. Each \(d_{i,\alpha}\) dominates the base norm \(\left\lvert x_i-y_i \right\rvert+\alpha^{-1}\left\lVert \xi-\eta \right\rVert\) and is locally comparable to it on bounded sets. Hence \(d_{{\rm sep},\alpha}\) is complete and separable on \(\mathbb{R}^d\times\Xi\) and has the usual Borel \(\sigma\)-field. Summing the coordinatewise reference-point costs at \((0,\xi^\circ)\) and applying Proposition 11 gives the invariant law, uniqueness, and 16 . ◻

8.2 Proof of Corollary 3↩︎

Proof of Corollary 3. Write \(q:=\lvert\mathcal{I}_*\rvert\), use a subscript \(*\) to denote restriction to the coordinates in \(\mathcal{I}_*\), and set \(Y_{\alpha,*} :=\alpha^{-1/m_{\max}} X_{\infty,*}^{(\alpha),\rm sep}.\) The coordinatewise moment bounds from the proof of Theorem 7 imply that \((Y_{\alpha,*})_\alpha\) is tight in \(\mathbb{R}^q\). We first identify its subsequential limits. By separability, the update of the active block does not depend on the remaining iterate coordinates. On the scale \(\alpha^{1/m_{\max}}\), its stationary one-step increment is \[Y_{\alpha,*}^+-Y_{\alpha,*} =-\alpha^{2-2/m_{\max}}h_{\alpha,*}(Y_{\alpha,*}) -\alpha^{1-1/m_{\max}} g_*(\alpha^{1/m_{\max}}Y_{\alpha,*},\xi_\alpha^+),\] where \[h_{\alpha,*}(y) :=\bigl( \alpha^{-(m_{\max}-1)/m_{\max}} h_i(\alpha^{1/m_{\max}}y_i) \bigr)_{i\in\mathcal{I}_*},\] and \(g_*(x,\xi):=(g_i(x_i,\xi))_{i\in\mathcal{I}_*}\). The coordinatewise tangent assumptions give \[h_{\alpha,*}\longrightarrow h_{0,*}, \qquad h_{0,*}(y):=(h_{i,0}(y_i))_{i\in\mathcal{I}_*}\] locally uniformly.

For \(x\in\mathbb{R}^q\), let \(\chi_{*,x}(\xi) :=(\chi_{i,x_i}(\xi))_{i\in\mathcal{I}_*},\) where \(\chi_{i,x_i}\) is the one-dimensional solution of the Poisson equation for coordinate \(i\). At the minimizer, define \[D_i(\xi,U_1) :=g_i(0,\Phi(\xi,U_1)) +\chi_{i,0}(\Phi(\xi,U_1))-\chi_{i,0}(\xi), \qquad i\in\mathcal{I}_*.\] The vector \(D_*=(D_i)_{i\in\mathcal{I}_*}\) is conditionally centered and, by Lemma 5, \[\Sigma_{\rm act}:=\mathbb{E}[D_*D_*^\top] =(\Sigma_{\rm sep})_{\mathcal{I}_*\times\mathcal{I}_*}.\] Thus the off-diagonal asymptotic covariances created by the common driving chain are retained within the active block.

For \(\varphi\in C_c^3(\mathbb{R}^q)\), use the perturbed test function \[\varphi_\alpha(y,\xi) :=\varphi(y) -\alpha^{1-1/m_{\max}} \left\langle \nabla\varphi(y),\chi_{*,\alpha^{1/m_{\max}}y}(\xi)\right\rangle.\] The calculation in Lemma 8 applies componentwise to this block. For \[M_{*,x}(\xi) :=g_*(x,\xi)g_*(x,\xi)^\top +g_*(x,\xi)\chi_{*,x}(\xi)^\top +\chi_{*,x}(\xi)g_*(x,\xi)^\top,\] each entry is a finite sum of products of one-dimensional terms. The coordinatewise Lipschitz estimates for \(g_i\), the Hölder estimates 62 for \(\chi_i\), and the coordinatewise stationary second moments therefore give, as in Lemma 25, \[\mathbb{E}\!\left[ \left\lVert QM_{*,\alpha^{1/m_{\max}}Y_{\alpha,*}}(\xi_\alpha) -QM_{*,0}(\xi_\alpha) \right\rVert \mathbf{1}_{\{Y_{\alpha,*}\in K\}} \right]\longrightarrow0\] for every compact \(K\subset\mathbb{R}^q\). The entries of \(QM_{*,0}\) belong to \(L^1(\pi_\Xi)\). Applying Lemma 26 entrywise to \(\nabla^2\varphi(Y_{\alpha,*})\) gives the required product limits. Consequently, every subsequential weak limit \(\nu_*\) of \(Y_{\alpha,*}\) satisfies \[\int_{\mathbb{R}^q} \left[ -\left\langle \nabla\varphi(y),h_{0,*}(y)\right\rangle +\frac{1}{2}\mathop{\mathrm{tr}}\bigl(\Sigma_{\rm act}\nabla^2\varphi(y)\bigr) \right]\nu_*(dy)=0, \qquad \varphi\in C_c^3(\mathbb{R}^q).\] Because every coordinate in \(\mathcal{I}_*\) has exponent \(m_{\max}\), the active drift satisfies, for constants \(c,c'>0\), \[\left\langle y,h_{0,*}(y)\right\rangle \ge c\sum_{i\in\mathcal{I}_*}|y_i|^{m_{\max}} \ge c'\left\lVert y \right\rVert^{m_{\max}},\] and the same argument gives \(\left\langle y-z,h_{0,*}(y)-h_{0,*}(z)\right\rangle \ge c'\left\lVert y-z \right\rVert^{m_{\max}}.\) Lemma 9 therefore identifies a unique invariant law for the active block of 19 . Hence the entire family \(Y_{\alpha,*}\) converges to this law.

It remains to consider coordinates outside \(\mathcal{I}_*\). If \(i\notin\mathcal{I}_*\), then \[\frac{X_{\infty,i}^{(\alpha),\rm sep}}{\alpha^{1/m_{\max}}} =\alpha^{1/m_i-1/m_{\max}} \left( \frac{X_{\infty,i}^{(\alpha),\rm sep}}{\alpha^{1/m_i}} \right) \longrightarrow0 \qquad\text{in probability},\] because the bracketed family is tight by the coordinatewise limit stated before the corollary. In 19 , these inactive coordinates have no Brownian forcing and solve \(dY_{i,t}=-h_{i,0}(Y_{i,t})\,dt.\) The one-dimensional coercivity gives \(y_i h_{i,0}(y_i)\ge c|y_i|^{m_i}\), so every such deterministic flow converges to zero and has \(\delta_0\) as its unique invariant law. The active block has the unique invariant law identified above. Therefore 19 has a unique invariant law, with inactive coordinates equal to zero. Combining the active-block convergence with the inactive-coordinate convergence proves the asserted convergence to \(Y_\infty^{\rm sep}\). ◻

8.3 Proof of Remark 10↩︎

Proof of Remark 10. For each coordinate \(i\), the one-dimensional version of Theorem 7 gives \(\frac{X_{\infty,i}^{(\alpha)}}{\alpha^{1/m_i}}\Rightarrow Y_{\infty,i}.\) For each fixed \(\alpha\), the stationary law factorizes by Remark 9, so the rescaled coordinates are independent. Hence the joint characteristic function factorizes: \[\mathbb{E}\exp\left(i\sum_{i=1}^d t_i\frac{X_{\infty,i}^{(\alpha)}}{\alpha^{1/m_i}}\right) = \prod_{i=1}^d \mathbb{E}\exp\left(i t_i\frac{X_{\infty,i}^{(\alpha)}}{\alpha^{1/m_i}}\right) \longrightarrow \prod_{i=1}^d \mathbb{E}e^{it_iY_{\infty,i}}.\] Lévy’s continuity theorem gives the claimed joint convergence. ◻

References↩︎

[1]
H. Robbins and S. Monro, “A stochastic approximation method,” Ann. Math. Statist., vol. 22, no. 3, pp. 400–407, 1951, doi: 10.1214/aoms/1177729586.
[2]
J. Kiefer and J. Wolfowitz, “Stochastic estimation of the maximum of a regression function,” Ann. Math. Statist., vol. 23, no. 3, pp. 462–466, 1952, doi: 10.1214/aoms/1177729392.
[3]
A. Benveniste, M. Métivier, and P. Priouret, Adaptive algorithms and stochastic approximations, vol. 22. Berlin: Springer-Verlag, 1990.
[4]
H. J. Kushner and G. G. Yin, Stochastic approximation and recursive algorithms and applications, Second., vol. 35. New York: Springer, 2003.
[5]
V. S. Borkar, Stochastic approximation: A dynamical systems viewpoint, vol. 48. Gurgaon: Hindustan Book Agency, 2008.
[6]
G. Ch. Pflug, “Stochastic minimization with constant step-size: Asymptotic laws,” SIAM J. Control Optim., vol. 24, no. 4, pp. 655–666, 1986, doi: 10.1137/0324039.
[7]
S. Mandt, M. D. Hoffman, and D. M. Blei, “Stochastic gradient descent as approximate Bayesian inference,” J. Mach. Learn. Res., vol. 18, no. 134, pp. 1–35, 2017.
[8]
Z. Chen, S. Mou, and S. T. Maguluri, “Stationary behavior of constant stepsize SGD type algorithms: An asymptotic characterization,” Proc. ACM Meas. Anal. Comput. Syst., vol. 6, no. 1, pp. 19:1–19:24, 2022, doi: 10.1145/3508039.
[9]
A. Dieuleveut, A. Durmus, and F. Bach, “Bridging the gap between constant step size stochastic gradient descent and Markov chains,” Ann. Statist., vol. 48, no. 3, pp. 1348–1382, 2020, doi: 10.1214/19-AOS1850.
[10]
I. Merad and S. Gaïffas, “Convergence and concentration properties of constant step-size SGD through Markov chains,” Electron. J. Stat., vol. 19, no. 2, pp. 5843–5894, 2025, doi: 10.1214/25-EJS2471.
[11]
K. Knight, “Limiting distributions for \(L_1\) regression estimators under general conditions,” Ann. Statist., vol. 26, no. 2, pp. 755–770, 1998, doi: 10.1214/aos/1028144858.
[12]
R. Koenker and G. Bassett Jr., “Regression quantiles,” Econometrica, vol. 46, no. 1, pp. 33–50, 1978, doi: 10.2307/1913643.
[13]
R. T. Rockafellar and S. Uryasev, “Optimization of Conditional Value-at-Risk,” J. Risk, vol. 2, no. 3, pp. 21–41, 2000, doi: 10.21314/JOR.2000.038.
[14]
P. J. Huber, “Robust estimation of a location parameter,” Ann. Math. Statist., vol. 35, no. 1, pp. 73–101, 1964, doi: 10.1214/aoms/1177703732.
[15]
P. Charbonnier, L. Blanc-Féraud, G. Aubert, and M. Barlaud, “Deterministic edge-preserving regularization in computed imaging,” IEEE Trans. Image Process., vol. 6, no. 2, pp. 298–311, 1997, doi: 10.1109/83.551699.
[16]
J. T. Barron, “A general and adaptive robust loss function,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 4331–4339, doi: 10.1109/CVPR.2019.00446.
[17]
P. L. Bartlett, M. I. Jordan, and J. D. McAuliffe, “Convexity, classification, and risk bounds,” J. Amer. Statist. Assoc., vol. 101, no. 473, pp. 138–156, 2006, doi: 10.1198/016214505000000907.
[18]
Y. Zhang, D. L. Huo, Y. Chen, and Q. Xie, Full version: arXiv:2504.08178“A piecewise Lyapunov analysis of sub-quadratic SGD: Applications to robust and quantile regression,” ACM SIGMETRICS Perform. Eval. Rev., vol. 53, no. 1, pp. 85–87, 2025, doi: 10.1145/3744970.3727269.
[19]
Y. Qu, J. Blanchet, and P. W. Glynn, “Computable bounds on convergence of Markov chains in Wasserstein distance via contractive drift,” Ann. Appl. Probab., vol. 35, no. 4, pp. 2678–2715, 2025, doi: 10.1214/25-AAP2184.
[20]
Z. Wang, Y. Wang, I. Narang, F. Wang, Y. Wang, and S. T. Maguluri, arXiv:2602.13960“Steady-state behavior of constant-stepsize stochastic approximation: Gaussian approximation and tail bounds.” 2026, [Online]. Available: https://arxiv.org/abs/2602.13960.
[21]
V. Fabian, “On asymptotic normality in stochastic approximation,” Ann. Math. Statist., vol. 39, no. 4, pp. 1327–1332, 1968, doi: 10.1214/aoms/1177698258.
[22]
L. Ljung, “Analysis of recursive stochastic algorithms,” IEEE Trans. Automat. Control, vol. 22, no. 4, pp. 551–575, 1977, doi: 10.1109/TAC.1977.1101561.
[23]
D. Ruppert, “Efficient estimations from a slowly convergent Robbins–Monro process,” School of Operations Research; Industrial Engineering, Cornell University, Ithaca, NY, Technical Report 781, 1988.
[24]
B. T. Polyak and A. B. Juditsky, “Acceleration of stochastic approximation by averaging,” SIAM J. Control Optim., vol. 30, no. 4, pp. 838–855, 1992, doi: 10.1137/0330046.
[25]
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro, “Robust stochastic approximation approach to stochastic programming,” SIAM J. Optim., vol. 19, no. 4, pp. 1574–1609, 2009, doi: 10.1137/070704277.
[26]
É. Moulines and F. R. Bach, “Non-asymptotic analysis of stochastic approximation algorithms for machine learning,” in Advances in neural information processing systems, 2011, vol. 24, pp. 451–459.
[27]
F. Bach and É. Moulines, “Non-strongly-convex smooth stochastic approximation with convergence rate O(1/n),” in Advances in neural information processing systems, 2013, vol. 26, pp. 773–781.
[28]
L. Yu, K. Balasubramanian, S. Volgushev, and M. A. Erdogdu, “An analysis of constant step size SGD in the non-convex regime: Asymptotic normality and bias,” in Advances in neural information processing systems, 2021, vol. 34, pp. 4234–4248.
[29]
Y. Zhang, D. L. Huo, Y. Chen, and Q. Xie, “Prelimit coupling and steady-state convergence of constant-stepsize nonsmooth contractive SA,” ACM SIGMETRICS Perform. Eval. Rev., vol. 52, no. 1, pp. 35–36, 2024, doi: 10.1145/3673660.3655076.
[30]
D. Huo, Y. Chen, and Q. Xie, Articles in Advance“Bias and extrapolation in Markovian linear stochastic approximation with constant stepsizes,” Math. Oper. Res., 2026, doi: 10.1287/moor.2024.0471.
[31]
D. L. Huo, Y. Zhang, Y. Chen, and Q. Xie, “The collusion of memory and nonlinearity in stochastic approximation with constant stepsize,” in Advances in neural information processing systems, 2024, vol. 37, pp. 21699–21762, doi: 10.52202/079017-0684.
[32]
H. Hadavi, W. Mou, S. Samsonov, and H.-T. Wai, arXiv:2604.13378“Revisiting the constant stepsize stochastic approximation with decision-dependent Markovian noise.” 2026, [Online]. Available: https://arxiv.org/abs/2604.13378.
[33]
N. Gast, “Expected values estimated via mean-field approximation are \(1/N\)-accurate,” Proc. ACM Meas. Anal. Comput. Syst., vol. 1, no. 1, pp. 17:1–17:26, 2017, doi: 10.1145/3084454.
[34]
S. Allmeier and N. Gast, “Computing the bias of constant-step stochastic approximation with Markovian noise,” in Advances in neural information processing systems, 2024, vol. 37, pp. 137873–137902, doi: 10.52202/079017-4379.
[35]
Z. Wei, J. Li, Z. Lou, and W. B. Wu, “Gaussian approximation and concentration of constant learning-rate stochastic gradient descent,” in Advances in neural information processing systems, 2025, vol. 38.
[36]
H. H. Bauschke and P. L. Combettes, “The Baillon–Haddad theorem revisited,” J. Convex Anal., vol. 17, no. 3–4, pp. 781–787, 2010.
[37]
R. Z. Khasminskii, “On the principle of averaging the Itô’s stochastic differential equations,” Kybernetika, vol. 4, no. 3, pp. 260–279, 1968.
[38]
V. S. Borkar, “Stochastic approximation with two time scales,” Systems Control Lett., vol. 29, no. 5, pp. 291–294, 1997, doi: 10.1016/S0167-6911(97)90015-3.
[39]
H.-W. Kang and T. G. Kurtz, “Separation of time-scales and model reduction for stochastic reaction networks,” Ann. Appl. Probab., vol. 23, no. 2, pp. 529–583, 2013, doi: 10.1214/12-AAP841.
[40]
A. D. Barbour, “Stein’s method for diffusion approximations,” Probab. Theory Related Fields, vol. 84, no. 3, pp. 297–322, 1990, doi: 10.1007/BF01197887.
[41]
A. Braverman, J. G. Dai, and J. Feng, “Stein’s method for steady-state diffusion approximations: An introduction through the Erlang-A and Erlang-C models,” Stoch. Syst., vol. 6, no. 2, pp. 301–366, 2016, doi: 10.1214/15-SSY212.
[42]
É. Pardoux and A. Yu. Veretennikov, “On the Poisson equation and diffusion approximation. I,” Ann. Probab., vol. 29, no. 3, pp. 1061–1085, 2001, doi: 10.1214/aop/1015345596.
[43]
É. Pardoux and A. Yu. Veretennikov, “On Poisson equation and diffusion approximation. II,” Ann. Probab., vol. 31, no. 3, pp. 1166–1192, 2003, doi: 10.1214/aop/1055425774.
[44]
É. Pardoux and A. Yu. Veretennikov, “On the Poisson equation and diffusion approximation. III,” Ann. Probab., vol. 33, no. 3, pp. 1111–1133, 2005, doi: 10.1214/009117905000000062.
[45]
S. P. Meyn and R. L. Tweedie, Markov chains and stochastic stability, Second. Cambridge: Cambridge University Press, 2009.
[46]
A. Eberle, A. Guillin, and R. Zimmer, “Quantitative Harris-type theorems for diffusions and McKean–Vlasov processes,” Trans. Amer. Math. Soc., vol. 371, no. 10, pp. 7135–7173, 2019, doi: 10.1090/tran/7576.
[47]
S. N. Ethier and T. G. Kurtz, Markov processes: Characterization and convergence. New York: John Wiley & Sons, 1986.

  1. \(^1\)Georgia Institute of Technology, Email: ↩︎

  2. \(^2\)Georgia Institute of Technology, Email: ↩︎

  3. \(^3\)Georgia Institute of Technology, Email: ↩︎