Potential Hessian Ascent: The Sherrington-Kirkpatrick Model


Abstract

We present the first iterative spectral algorithm to find near-optimal solutions for the Sherrington-Kirkpatrick model, which is a random quadratic objective over the discrete hypercube, resolving a conjecture of Subag [1]. The algorithm is a randomized Hessian ascent in the solid cube, with the objective modified by subtracting an instance-independent potential function [2], [3].

Using tools from free probability theory, we construct an approximate projector into the top eigenspaces of the Hessian, which serves as the covariance matrix for the random increments. With high probability, the iterates’ empirical distribution approximates the solution to the primal version of the Auffinger-Chen SDE [4]. The per-iterate change in the modified objective is bounded via a Taylor expansion, where the derivatives are controlled through Gaussian concentration bounds and smoothness properties of a semiconcave regularization of the Fenchel-Legendre dual to the Parisi PDE [5].

These results lay the groundwork for (possibly) demonstrating low-degree sum-of-squares certificates over high-entropy step distributions for a relaxed version of the Parisi formula [6].

1 Introduction↩︎

1.1 Overview↩︎

Subag [1] introduced a marvellously simple algorithm for maximizing the Hamiltonian of a spherical spin glass: start at the origin and repeatedly take small steps along the top eigenspaces of the Hessian of the Hamiltonian until eventually reaching a near-maximizer on the boundary. This can be understood as an analog of gradient ascent in a situation where the gradient is zero along the path we want to follow. The algorithm finds a near-optimum configuration whenever the spherical spin glass problem satisfies full replica symmetry breaking (fRSB)1, and the conjectured optimum achievable by efficient algorithms otherwise [7], [8]. When working with the same Hamiltonian on the discrete hypercube rather than the sphere, efficient algorithms using approximate message passing (AMP) [9], [10] have been developed, but it has since remained unclear how to provably extend Subag’s conceptually straightforward approach which relied heavily on the highly symmetric geometry of the sphere; see [1], [9] and [11].

To this end, Subag conjectured that it should be possible to follow a path of local maximizers of the objective corrected with a so-called generalized TAP correction under fRSB on the cube [1]. In this work, we positively resolve this conjecture by developing a Hessian ascent algorithm for the Sherrington-Kirkpatrick (SK) model that works by locally maximizing this TAP-corrected objective.

The conceptual framework for the algorithm can be termed potential Hessian ascent (PHA). It subtracts a regularizing penalty function from the underlying objective function which encourages coordinates far from a corner to move more and those nearby to move less, and then iteratively optimizes the resulting objective function in the interior of the solution domain. The PHA algorithm takes small steps, each in a randomized direction that follows the critical points of a modified objective. Solving the stationary conditions introduced by this formulation yields an iterative spectral algorithm which follows the top eigenspace of a TAP-corrected Hessian.

In the special case of the SK model, the Hamiltonian is given by \(H(\sigma) = \angles{\sigma, A \sigma}\) for \(\sigma \in \{\pm 1\}^n\), where \(A\) is a real Ginibre random matrix (a matrix whose entries are i.i.d.Gaussian random variables with mean \(0\) and variance \(1/n\)), and our goal is to find a near-maximizer of \(H(\sigma)\) over \(\{\pm 1\}^n\). Given an input \(n\) and an inverse temperature \(\beta\) (related to the desired accuracy \(\varepsilon\) of the output), our algorithm performs a Hessian ascent on the modified objective function \[\label{eq:32first32introduction32of32objective} \mathop{\mathrm{obj}}_{\beta,\gamma,n}\left(t, \sigma\right) = \beta\left\langle\sigma, A \sigma\right\rangle - V_{\beta,\gamma,n}\left(t, \sigma\right),\tag{1}\] where the potential \(V_{\beta,\gamma,n}: \mathbb{R}^n \to \mathbb{R}\) is a combination of per-coordinate terms and a global energy term: \[\label{eq:32first32introduction32of32potential} V_{\beta,\gamma,n}\left(t, \sigma\right) := \sum_{j=1}^n \tilde{\Lambda}_{\beta,\gamma}\left(t, \sigma_j\right) + n r_{\beta,\mu_\beta}\left(t\right).\tag{2}\] Here \(\tilde{\Lambda}_{\beta,\gamma}\) and \(r_{\beta,\mu_\beta}\) are both defined in terms of the Parisi formula for the SK model [12], which we explain in §1.2. Briefly, there is a probability measure \(\mu_\beta\) on \([0,1]\) and a function \(\Phi_{\beta,\mu_\beta}: [0,1] \times \mathbb{R}\to \mathbb{R}\) solving a certain optimization problem depending on \(\beta\). Our \(\tilde{\Lambda}_{\beta,\gamma}\) is the Fenchel–Legendre dual in the spatial variable of \(\Phi_{\beta,\mu_\beta}(t,x) + \gamma x^2 / 2\), where the \(\gamma\) term is added to prevent \(\tilde{\Lambda}_{\beta,\gamma}\) from taking infinite values outside \([-1,1]\). The energy term \(r_{\beta,\mu_\beta}\left(t\right)\) is given by \(\beta^2 \int_t^1 s\, \mu_\beta([0,s])\,ds\).

Assuming a strong form of full replica symmetry breaking, the total amount we expect to gain in the function \(\angles{\sigma, A\sigma} -\sum_{j=1}^n \tilde{\Lambda}_{\beta,\gamma}\left(t, \sigma_j\right)\) on the time interval \([t,1]\) is described by \(nr_{\beta}\left(t,\mu_\beta\right)\), and so by subtracting this term, we arrange that the modified objective function should have approximately a constant value over the entire path given by our algorithm. The necessary fRSB assumption is widely believed to hold for the SK model at low temperature, and Auffinger, Chen & Zeng [13] have made significant partial progress toward establishing this rigorously.

The Hessian ascent algorithm produces iterates \((\sigma_k)_{k=1}^K\) in \(\mathbb{R}^n\) (staying close to \([-1,1]^n\)) whose increments \(\sigma_{k+1} - \sigma_k\) are given by following the top part of the bulk of the spectrum of the Hessian of the modified objective \(\mathop{\mathrm{obj}}_{\beta,\gamma,n}(t,\sigma)\), analogously to Subag’s approach on the sphere. To analyze the PHA algorithm on the cube, we control the Taylor expansion of the objective function at each iterate \(\sigma_k\) using tools from random matrix theory and stochastic analysis. Conditioned on the values of \(A\) and the preceding iterate \(\sigma_k\), each increment \(\sigma_{k+1} - \sigma_k\) in the algorithm is a Gaussian random vector with a covariance matrix designed to concentrate on the top part of the bulk of the spectrum of the Hessian at \(\sigma_k\). Since the potential on the cube is a sum of functions of the individual coordinates, it is essential to analyze the averaged behavior of the individual coordinates of \(\sigma_k\) (i.e.the empirical distribution of \(\sigma_k\)), which in turn requires precise control over the diagonal entries of the covariance matrix used to generate the increment \(\sigma_{k+1} - \sigma_k\). Our argument proceeds in several stages, each of which has some independent interest:

  1. Analysis of the Hessian of the modified objective using random matrix theory (). Prior work of Gufler, Schertzer and Schmidt [14] in the RS regime used free probability to study the limiting spectral distribution of the modified Hessian, and the use of free probability to study the diagonal entries of a resolvent was suggested in [15] as well. We characterize various properties about the spectral distribution as well as bulk statistics of the diagonal entries, in the lower-temperature fRSB regime.

  2. A proof that the empirical distribution of the PHA algorithm’s iterates converge to the (primal) Auffinger-Chen SDE () .

  3. Concentration bounds for the derivatives of an infimum-convolved Fenchel-Legendre dual of the solution to the Parisi PDE (§6). These bounds are combined with the previous two results via a Taylor-expansion based approach to control the fluctuations of the modified objective function () .

1.2 Background↩︎

To set the stage for a more detailed statement of our results, we now recall background on the Sherrington-Kirkpatrick model and previous work on finding the associated maximum and maximizers.

1.2.1 The Sherrington-Kirkpatrick model↩︎

The Sherrington-Kirkpatrick model [16] was introduced as a “mean-field" simplification of Ising spin glass models on \(d\)-dimensional lattices (\(\mathbb{Z}^d\)). The configuration space is described by assigning \(\pm 1\) to each vertex of a complete graph on \(n\) vertices. The graph has Gaussian (signed) weights assigned to each edge. Thus, the configuration space is \(\{\pm 1\}^n\), and the Hamiltonian is a random quadratic polynomial on the hypercube \(\{\pm 1\}^n\). Specifically, \[\label{def:sk-model} H_n(\sigma) = \sum_{i,j} A_{i,j}\sigma_i\sigma_j\,,\tag{3}\] where \(A\) is a Gaussian random matrix with i.i.d.entries. Moreover, we assume that \(A\) is normalized so that \(A_{i,j} \overset{i.i.d.}{\sim}\mathcal{N}(0,1/n)\) for every \(i, j \in [n]\). Under this normalization \(A_{\mathop{\mathrm{sym}}} := (A + A^{\mathsf{T}})/2\) is \(1/\sqrt{2}\) times a GOE random matrix, and hence the spectral distribution converges almost surely as \(n \to \infty\) to the Wigner semicircle law (see §4.1 below).

One of the main problems of interest is to find the energy of the ground state; here the ground state would be the maximizer of \(H(\sigma)\) on \(\{\pm 1\}^n\) and its energy is the maximum value. As \[H_n(\sigma) = \sum_{i,j} A_{i,j} - 2 \sum_{\sigma_i\sigma_j = -1} A_{i,j},\] maximizing \(H(\sigma)\) is equivalent to minimizing \(\sum_{\sigma_i\sigma_j = -1} A_{i,j}\), i.e.minimizing the weighted “cut” between the vertex sets \(\{i: \sigma_i = 1\}\) and \(\{j: \sigma_j = -1\}\).

In particular, we want to analyze the behavior of the ground state as the dimension \(n\) tends to infinity. Because of standard concentration of measure results for Lipschitz functions of Gaussian variables, the value of the maximum is close to its expectation with high probability [17], and hence the asymptotic behavior is described by the ground state energy density given by \[\begin{align} \lim_{n \to \infty} \frac{1}{n} \mathop{{}\mathbb{E}}_{A}\left[\max_{\sigma \in \{\pm 1\}^n}H_n(\sigma)\right] := \lim_{n \to \infty} \frac{1}{n} \mathop{{}\mathbb{E}}_A\left[\max_{\sigma \in \{\pm 1\}^n} \langle \sigma, A\sigma\rangle\right]. \end{align}\]

1.2.2 Parisi formula and Auffinger-Chen representation↩︎

The free energy of the Sherrington-Kirkpatrick model is given by a variational principle first proposed by Parisi [12] and then rigorously proved nearly twenty-five years later by Guerra [18], [19], Talagrand [20] and Panchenko [21], [22]. The variational principle transforms the static optimization problem of the expression of the free energy into a variational optimization problem over the space of probability distributions on \([0,1]\), and simultaneously minimizes a free-entropy term that comes from solving the so-called Parisi PDE and a variance-like quantity.

Theorem 1 (Parisi Variational Principle, [12], [20]).
For \(n \in \mathbb{N}\) and an inverse temperature \(\beta = \frac{1}{T} > 0\), let \[\mathcal{F}_{n, \beta} := \log \sum_{\sigma \in \{\pm 1\}^n} e^{\beta H_n(\sigma)}\] be the free energy of the \(n\)-dimensional Sherrington-Kirkpatrick Hamiltonian \(H_n\).

For each probability measure \(\mu \in \mathcal{P}([0,1])\), let \(F_\mu(t) = \mu([0,t])\) be its cumulative distribution function, and let \(\Phi_{\beta,\mu}\) be the solution to the parabolic PDE \[\label{eq:32Parisi32PDE} \partial_t \Phi(t,x) = -\beta^2\left(\partial_{x,x}\Phi(t, x) + F_{\mu}(t)(\partial_x\Phi(t,x))^2\right)\, ,\tag{4}\] with the terminal condition \[\Phi(1, x) = \log(2\cosh(x))\, .\] Let \(\mathcal{P}_\beta\) be the optimal free entropy \[\label{eq:parisi-formula-optimization} \mathcal{P}_\beta := \inf_{\mu \in \mathcal{P}[0, 1]} \left[ \Phi_{\beta,\mu}(0, 0) - \beta^2\int_0^1tF_{\mu}(t)dt \right].\tag{5}\] Then \[\label{eq:parisi-formula} \lim_{n \to \infty}\frac{1}{n}\mathcal{F}_{n, \beta} \overset{a.s.}{=} \mathcal{P}_\beta.\tag{6}\]

The parabolic PDE 4 is solved “backwards" in time, and if \(\mu\) is atomic then the solution can be computed explicitly using the Hopf-Cole transformations [4]; see . Crucially, note that the free-entropy term \(\Phi_{\beta,\mu}(0,0)\) that comes from the PDE itself depends on the probability distribution \(\mu\) to be optimized over.

in particular allows one to evaluate the ground state energy of the Sherrington-Kirkpatrick model as the zero temperature limit of the underlying Parisi formula [23]. We also mention that Auffinger and Chen have given a version of the Parisi PDE at \(\beta = \infty\), although it has less regularity than the finite \(\beta\) versions [24].

Corollary 1 (Ground State Energy of the Sherrington-Kirkpatrick Model [23]). \[\lim_{n \to \infty}\frac{1}{n}\mathop{{}\mathbb{E}}_A\left[\max_{\sigma \in \{\pm 1\}^n}H_n(\sigma)\right] = \lim_{\beta \to \infty}\frac{1}{\beta}\,\mathcal{P}_\beta\, .\]

The study of the ground state energy as well as the configurations that realize the optimum depends crucially on the regularity of the solution \(\Phi_\mu\) and the measure \(\mu_\beta\) itself. To obtain a unique minimizer \(\mu\) in the Parisi formula [4], Auffinger and Chen showed that the associated functional is convex, by reformulating the Parisi formula as a stochastic control problem.

Theorem 2 (Auffinger-Chen Principle, [4], [9]).
For a fixed \(\beta\) and \(\mu\), the free entropy term \(\Phi_\mu(t, x)\) at \(t = 0\) can be equivalently rewritten as a maximization over a set \(\mathcal{D}\) of Brownian-motion-adapted stochastic processes that are stopped when the process exits \([-1,1]\), \[\label{eq:auffinger-chen} \Phi_{\beta,\mu}(0, x) = \max_{u \in \mathcal{D}}\left(\mathop{{}\mathbb{E}}\left[\Phi_{\beta,\mu}\left(1, x + 2\beta^2\int_0^1F_{\mu}(s)u_s\,ds + \sqrt{2}\beta\int_0^1dW_s\right) \right] - \beta^2\int_0^1F_{\mu}(s)\mathop{{}\mathbb{E}}[u_s^2]ds \right)\, ,\tag{7}\] where \(\{W_s\}\) is standard Brownian motion, \(\mathcal{F}_t\) is the filtration generated \(\{W_t\}\), and \[\mathcal{D}:= \left\{(u_t)_{0 \leqslant t \leqslant 1}\, |\, u_t\text{ is progressively measurable w.r.t. }\mathcal{F}_t,\, |u_t| \leqslant 1\right\}\, .\] The maximizer of 7 is unique and is given by, \[u_s = \partial_x\Phi_{\beta,\mu}(s, x + X_s)\, ,\] and \(X_s\) is the strong solution following the Stochastic Differential Equation, \[\label{eq:auffinger-chen-sde} dX_s = 2\beta^2F_{\mu}(s)\partial_x\Phi_{\beta,\mu}(s,x + X_s)ds + \sqrt{2}\beta dW_s,\, X_0 = 0\, .\tag{8}\] with initial condition \(X_0 = 0\).

1.2.3 Algorithms to find the maximizers↩︎

The Auffinger-Chen representation is critical not only for showing regularity of the Parisi optimal measure \(\mu_\beta\), but also for describing the maximizers of \(H_n(\sigma)\). Montanari used the Auffinger-Chen SDE to design an approximate message passing (AMP) algorithm for finding near maximizers [9]. The successive iterates of the AMP algorithm for every coordinate \(i \in [n]\) are given by multiplying the current value of \(i\) with a non-linear function evaluated at a point chosen by discretizing the increments of 8 and propagating the updated value via the Gaussian matrix \(A\) that specifies the input. Various results using the AMP framework have given efficient algorithms to optimize average-case problems such as mean-field spin glasses on the hypercube [9], [10] and sparse random Max-CSPs [25], [26]. The AMP framework has also been applied with great success to problems in estimation [27][29], sampling [30], [31] and compressed sensing [32]. In the case of spin-glasses on the sphere, Subag gave a conceptually simpler algorithm called Hessian ascent [1] which we describe further in §1.3.1 as it forms the conceptual basis for our work.

For our algorithm to attain the actual maximum in the Parisi formula, one requires certain assumptions on the measure \(\mu_\beta\), namely full Replica Symmetry Breaking (fRSB) which means that the measure \(\mu_\beta\) has full support asymptotically as \(\beta \to \infty\). Auffinger and Chen [5] showed that if \(\mu\) has full support, then \(\Phi_{\beta,\mu_\beta}\) is a sufficiently smooth function of \((t,x)\) for our purposes. The fRSB condition is widely believed to hold in the SK model for sufficiently large \(\beta\), but the state of the art for rigorous results on large \(\beta\) is that \(\mu\) has infinite support [13].

Without assuming fRSB, efficient algorithms are expected to not perform better than an explicit threshold given by a relaxed Parisi formula introduced by Huang and Sellke [7] (which agrees with the original Parisi formula in the fRSB case). The AMP algorithms produce solutions that approximately achieve this relaxed optimum in general, while there are many known hardness results for going beyond the relaxed optimum. Variants of the so-called overlap-gap property (OGP), introduced by Gamarnik and Sudan [33], obstruct various natural families of algorithms, including local classical & quantum algorithms [33][35], low-degree polynomials [36] and AMP [37]. The unifying implications of this work are found in a result by Huang and Sellke [7] via the introduction of the most general variant of the OGP, which holds for mean-field spin glasses on the hypercube with even degree \(\geqslant 4\). These obstructions are transferred to the setting of sparse random Max-CSPs due to a result of Jones, Marwaha, Sandhu and Shi [8].

1.3 Main Results↩︎

In this section we state the main result of this paper, which is a polynomial time approximation scheme (PTAS) for the SK model that extends the algorithmic program of Subag [1] to the cube. Note that here “polynomial time” means polynomial in \(n\) with constants that can depend on \(\beta\) or \(\varepsilon\). Like Subag’s work, the algorithm relies on spectral analysis and does not use the AMP framework.

1.3.1 Concept: Potential Hessian Ascent↩︎

Suppose that we are trying to maximize some real-valued objective function \(H\) on some domain \(D\). In particular, for some given precision \(\varepsilon\), we want to output in polynomial time some \(\sigma \in D\) such that \[H(\sigma) \geqslant(1 - \varepsilon) \sup_{\sigma' \in D} H(\sigma').\] Akin to the celebrated work of Subag [1] the PHA algorithm considers an extension \(\tilde{H}: \tilde{D}\to \mathbb{R}\) of the original objective function \(H\)—which was originally defined on the domain \(D\). For instance, if \(D\) is the cube \(\{\pm 1\}^n\), then \(\tilde{D}\) will be the \([-1,1]^n\), and more generally one could consider a bounded set \(D\) and let \(\tilde{D}\) be its convex hull.

We want to choose \(\tilde{H}\) so that its maximum value is equal to that of \(H\), and this maximum is attained along an entire path from a given point \(\sigma_0 \in \tilde{D}\) to a maximizer of \(H\) in \(D\). We then define iterates inductively that approximately follow such a path. Since the gradient of \(\tilde{H}\) is zero along the path, maintaining the optimum value requires us to move only in directions in the kernel of \(\nabla^2 \tilde{H}\). We therefore define iterates inductively starting from the given point \(\sigma_0\) by \(\sigma_{k+1} = \sigma_k + Q_k Z_k\), where \(Z_k\) is a standard Gaussian vector and \(Q_k^2\) is a certain covariance matrix, which should approximately project into the kernel of \(\nabla^2 \tilde{H}\). While Subag used matrix power iterations to approximately choose an eigenvector with largest eigenvalue, we can straightforwardly generalize this to a random vector from the top part of the spectrum [6].

In the case of the SK model, \(D = \{\pm 1\}^n\) and \(\tilde{D}= [-1,1]^n\). We will choose the modified objective function \(\tilde{H} = \mathop{\mathrm{obj}}\) given by 1 and 2 . The potential \(V_{\beta,\gamma,n}\) in 2 is a function \([0,1] \times [-1,1]^n \to \mathbb{R}\) given by \(\sum_j \tilde{\Lambda}_{\beta,\gamma}(t, \sigma_j) + n r_{\beta,\mu_\beta}(t)\), where \(\tilde{\Lambda}_{\beta,\gamma}: [0,1] \times \mathbb{R}\to \mathbb{R}\cup \{\pm \infty\}\) is a regularized version of the Fenchel-Legendre conjugate in space of the solution \(\Phi_{\beta,\mu_\beta}\) to the Parisi PDE, and \(r_{\beta,\mu_\beta}\) is a certain correction depending only on time. Note that \(V_{\beta,\gamma,n}\) is independent of the particular instance of \(A\), and thus only requires access to the Parisi solution \(\Phi_{\beta,\mu_\beta}\) and the measure \(\mu_\beta\), which is a \(2\)-dimensional rather than an \(n\)-dimensional problem. Our goal is to show that Hessian ascent succeeds in finding a near optimizer using this modified objective function.

Our modified objective function is motivated by the concept of generalized TAP free energy [2], [3] in Subag’s algorithm. We imagine that we start with a certain total budget of energy, and at each step, we balance increasing the objective \(\langle\sigma,A\sigma \rangle\) with spending a limited amount of energy (i.e. not increasing \(V_\beta\) too much). Subag’s algorithm maximized the TAP free energy on the ball \(\tilde{D}\) by following the top eigenvector of the Hessian of this energy. Furthermore, the Parisi-like formula on the sphere [3], [38] simplifies in the fRSB regime to yield an energy that matches the contribution from the Hessian term. Similarly, our PHA algorithm for the SK model follows the top-eigenspace of the Hessian of the generalized TAP free energy of the SK model on the cube. The top-eigenvalue of this Hessian corresponds to a term that is, roughly, the rate at which an “entropy” term changes in a direction orthogonal to the current iterate. This term ends up matching the energy given by the Parisi formula in the fRSB regime, giving the final value achieved by the algorithm.

1.3.2 Main results on the algorithm↩︎

The PHA algorithm for the SK model has the following inputs:

  1. Oracle access to the solution to the Parisi PDE \(\Phi_{\beta,\mu_\beta}(t,x)\) and its derivatives,

  2. Oracle access to the cumulative distribution function \(F_{\mu_\beta}: [0,1] \to [0,1]\) of the Parisi measure \(\mu_\beta\).

  3. An instance of the real Ginibre matrix \(A\) that specifies the Sherrington–Kirkpatrick Hamiltonian.

We also make the following assumption.

Assumption 1 (The SK model is fRSB).
The unique minimizer \(\mu_\beta\) in the Parisi formula has support equal to \([0,q_\beta^*]\) where \(q^*_\beta \to 1\) as \(\beta \to \infty\).

Recall from [5], if the support of \(\mu_{\beta}\) contains \([0,q_\beta^*]\), then \(\Phi_{\beta,\mu_\beta}\) will have sufficient smoothness as a function of \((t,x)\).

We provide a simplified and informal version of the PHA algorithm below (); see for a complete description. The main result is that under  and the choice of potential function \(V_{\beta,\gamma,n}(t,\sigma)\) given above, the PHA algorithm is a PTAS for the SK model. More precisely, for \(\varepsilon > 0\) and \(\beta = 10 / \varepsilon\) and \(\delta \in (0,1/19]\), with high probability, the algorithm achieves an error of \(O(\varepsilon) + o_n(1)\) in runtime \(\widetilde{O}\left(n^{2+4\delta} \exp(\text{constant}/\varepsilon^2)) \right)\), where \(\widetilde{O}\) hides factors that are polylog in \(n\). Below, the parameter \(\eta(\varepsilon) = \mathsf{exp}\left(-\mathsf{poly}\left(\frac{1}{\varepsilon}\right)\right)\) controls the step-size of the algorithm, and the number of steps is \(K = O\left(\frac{q^*}{\eta}\right) = O\left(\mathsf{exp}\left(\mathsf{poly}\left(\frac{1}{\varepsilon}\right)\right)\right)\).

Theorem 3 (PHA is an asymptotic PTAS for the SK model).
Let \(\varepsilon> 0\) be sufficiently small and \(\delta \in \left(0, \frac{1}{19}\right]\). Then, fix \(\beta = \frac{10}{\varepsilon}\), \(\eta = e^{-C\beta^2}\) where \(C\) is a sufficiently large absolute constant, and \(\gamma = \eta^{1/8}\) and let \(n \geqslant\eta^{-90}\).
Sample an instance \(\{A\}_{i,j \in [n]\times [n]}\) of the real Ginibre matrix used in the Hamiltonian as in .
Then, under , the PHA algorithm () takes as input \(A\) and \(\varepsilon\), as well as oracle access to \(\Phi\), its derivatives and \(\mu_{\beta}\), and outputs a configuration \(\sigma^* \in \{\pm 1\}^n\) such that \[\frac{1}{n} H_n(\sigma^*) \geqslant \frac{\mathcal{P}_\beta}{\beta} -\frac{\varepsilon}{5}-O(\varepsilon^2)-O_{\varepsilon}(n^{-\alpha}),\] where \(\alpha = \min\left(\frac{1}{24}, \frac{1}{4}\delta\right)\), with probability \(\geqslant 1 - n\eta^{-2}e^{-\Omega(n^{2/9 - 4\delta})}\) (in the matrix \(A\) and the randomness of the algorithm itself). Furthermore, the algorithm runs in time \(\widetilde{O}\left(n^{2 + 4\delta}\exp\left(O(1/\varepsilon^{2})\right)\right)\).

Consequently, for every fixed \(\varepsilon>0\), with suitable parameters depending only on \(\varepsilon\), PHA runs in time polynomial in \(n\) and outputs \(\sigma^*\in\{\pm1\}^n\) satisfying \[H_n(\sigma^*)\geqslant \left(1-\frac{\varepsilon}{2}-O(\varepsilon^2)-o_n(1)\right) \max_{\sigma\in\{\pm1\}^n}H_n(\sigma)\] with probability \(1-\exp(-n^{\Omega(1)})\) as \(n\to\infty\).

This theorem is proven in and .

Figure 1: Potential Hessian Ascent (Informal)

1.3.3 Ingredients in the analysis↩︎

The proof of convergence of our algorithm proceeds in three stages.

  1. Spectral analysis of the modified Hessian: We use free probability theory to locate the top part of the spectrum of \(2 \beta A_{\mathop{\mathrm{sym}}} - D\) where \(\beta\) is the inverse temperature, \(A_{\mathop{\mathrm{sym}}}\) is the symmetrization of the Ginibre random matrix, and \(D\) is a diagonal matrix which is the Hessian of \(\sum_j \Lambda\left(t, \sigma_j\right)\). We construct a positive operator \(Q\) whose weight is concentrated on the top part of the spectrum, and then we show that the diagonal of \(Q^2\) is close to \(D^{-2}\) using non-commutative conditional expectation formulas for free sums together with high-dimensional concentration of measure for the Gaussian matrix (see §4). The proof works for arbitrary nonnegative diagonal matrices and the same tools could be applied to a variety of combinations of a Gaussian random matrix and another matrix depending on the choice of potential.

  2. Convergence of empirical distribution of \(\sigma_k\) to Auffinger-Chen SDE: Using this control over the diagonal entries of \(Q^2\) and concentration of measure for the Gaussian vector, we show that on average the \(j\)th entry of \(\sigma_{k+1} - \sigma_k\) behaves like an independent Gaussian of variance \(v(t, \sigma_{k,j})^2\) where \(v(t,y) = \sqrt{2}\beta / \partial_{y,y} \Lambda(t,y)\). We deduce that with high probability the empirical distribution of \(\sigma_k\) is well approximated by the distribution of \(Y_t\), where \(Y_t\) is the solution to the SDE \(dY_t = v(t,Y_t)\,dW_t\) driven by a Brownian motion \(W_t\) (see §5). This is the SDE studied by Auffinger and Chen that closely relates to Parisi PDE for \(\Phi\), and thus our argument gives some hint about how PHA can generate solutions to an SDE appropriate to the geometry of the underlying region.

  3. Taylor expansion and concentration estimates: Finally, we analyze the change in the modified objective function at each iteration \(\sigma_k\) using a mixture of Taylor expansions and concentration estimates. To understand the derivatives of the objective function, we express Parisi’s PDE for \(\Phi\) (in the dual coordinate space) in terms of the inf-convolved Fenchel-Legendre conjugate \(\Lambda\) (in the primal coordinate space) and prove various regularity properties about it (see §2). The terms in the Taylor expansion can then be controlled using the ingredients from the first two steps as well as concentration bounds for norms of minorly correlated Gaussians (see §6). This argument provides a template to analyze modified objective functions under different potentials: to do so, systematically combine spectral estimates about the covariance matrix used by the PHA algorithm and moment estimates for the empirical distribution of its iterates with the regularity properties of the potential function and an appropriate set of concentration inequalities. We remark that the Taylor expansion procedure has interest beyond simply showing convergence since this methodology would be critical to obtain sum-of-squares proofs of near-optimality under high-entropy step distributions (see §7.2).

An intriguing feature of the proof is that we effectively compartmentalize the randomness of the Gaussian matrix \(A\) in the objective function and the randomness of the Gaussian vectors used in the algorithm. Indeed, §4 guarantees properties of the random matrix \(A\) in relation to all diagonal matrices uniformly, and the argument of §5 and §6 would apply to any matrix \(A\) with these properties; this leaves open the possibility to generalize the random matrix used as input for the optimization problem.

1.3.4 Significance↩︎

Our algorithm implements Subag’s Hessian ascent idea [1] in the case of the cube rather than the sphere. It is consequently the first spectral algorithm for analyzing spin glasses on \(\{\pm 1\}^n\) that is a PTAS. The workings of this algorithm are very different from the AMP algorithm of Montanari [9]. Both the PHA algorithm and the AMP algorithm require access to the function \(\Phi_{\beta,\mu_\beta}(t,x): [0,1] \times \mathbb{R}\to \mathbb{R}\) which solves the Parisi PDE (see ), in addition to its derivatives and the Parisi order parameter \(F_{\mu_\beta}: [0,1] \to [0,1]\). These are reasonable assumptions, as the \(\Phi_{\beta,\mu_\beta}\) does not depend on \(n\), and it is known how to solve the Parisi formula efficiently [4], [39] and the interested reader may consult [9] for details.

While AMP is already known to optimize the SK model under fRSB, the contribution of PHA is to build from the basic concepts of quadratic optimization and spectral analysis of random matrices, and to demonstrate convergence to the Auffinger-Chen SDE—and therefore near-optimal solutions—as a consequence of these foundations. The resulting algorithm is the first to directly leverage a geometric understanding of the solution landscape of the SK model, and in this way makes progress toward de-mystifying the algorithmic implications of the Parisi formula.

Our algorithm has the same time complexity as AMP up to a factor of \(n^{o(1)}\). We believe that the dependence on \(\varepsilon\) needs to be of leading order that is \(\exp\left(\mathrm{poly}(\frac{1}{\varepsilon})\right)\) in our algorithm, but that it may be possible to bring the dependence down with use of better-than-worst-case Lipschitz estimates implied by the regularity bounds developed in §2.

In future work, we hope that the PHA framework may be applied to optimize more general random polynomials on the cube, and other domains as well. In particular, we propose as a problem for future work to investigate the applications to sums-of-squares (SoS) certificates for the optimizers. It is well established that no low-degree SoS certificates over the standard SoS hierarchy exist for the Parisi formula [40][42]. This is largely due to the inability of low-degree SoS to capture concentration-of-measure [43]. However, it has recently been proved that there do exist low-degree SoS certificates that points generated by high entropy step (HES) distributions cannot have value larger than (some constant times) the relaxed value of the Parisi formula on the sphere [6]. Our work could allow analogous results to be proved on the cube in the fRSB regime as proposed in [6]. For further discussion, see §7.

1.4 Notation↩︎

Here we briefly overview some notations and conventions that will be used throughout the paper, including in particular the notation for matrices and vectors. Note that we will often write \([n]\) for the index set \(\{1,\dots,n\}\) for \(n \in \mathbb{N}\).

Notation 4 (Matrices).  

  • \(M_n(\mathbb{R})\) and \(M_n(\mathbb{C})\) denote the \(n \times n\) real and complex matrices respectively.

  • We use \(*\) for the adjoint on \(M_n(\mathbb{C})\) and \(\mathsf{T}\) for the transpose on \(M_n(\mathbb{R})\) or \(M_n(\mathbb{C})\).

  • We write \(\mathsf{Tr}_n\) for the usual trace \(\mathsf{Tr}_n(A) = \sum_{j=1}^n A_{j,j}\) and \(\operatorname{tr}_n\) for the normalized trace \(\operatorname{tr}_n(A) = (1/n) \mathsf{Tr}_n(A)\).

  • For \(\lambda \in \mathbb{C}\), we denote the multiple \(\lambda I\) of the identity matrix simply by \(\lambda\) (for instance, \(\lambda + A\) denotes \(\lambda I + A\)).

We generally follow the convention from free probability (see §4.1) that traces and Schatten norms of operators are normalized so that they are equal to the moments of the spectral distribution of the operator; this convention makes it easy to compare the operators with the idealized limiting objects.

Notation 5 (Norms for matrices). The normalized Schatten norms are given for \(p \in [1,\infty)\) by \[\left\lVert{A}\right\rVert_p = (\operatorname{tr}_n((A^*A)^{p/2}))^{1/p}.\] Hence, if \(A\) has singular values \(a_1\), …, \(a_n\), then \[\left\lVert{A}\right\rVert_p = \left( \frac{1}{n}\sum_{j=1}^n a_j^p \right)^{1/p}.\] The Schatten-\(\infty\) norm is by definition equal to the operator norm, which we denote simply by \(\left\lVert{A}\right\rVert := \left\lVert{A}\right\rVert_{\infty} := \max_{i\in[n]} |a_i|\).

By contrast, the \(\ell^p\) norms for our vectors are not normalized. To visually distinguish these from the operator norms, we use single bars \(|\cdot|\) to denote vector norms.

Notation 6 (Norms for vectors). For an \(n\)-dimensional vector \(v\) with components \(v_1, \dots, v_n\), we have \(|v|_p := \sqrt[p]{\sum_{i \in [n]} |v_i|^p}\) for \(p \in [1,\infty)\), and \(|v|_\infty = \max_i |v_i|\).

Notation 7 (Operations with diagonal matrices). Given a vector \(v = (v_1,\dots,v_n)\), we denote by \(\operatorname{diag}(v)\) or \(\operatorname{diag}(v_j)\) the diagonal matrix with diagonal entries \(v_1\), …, \(v_n\). Meanwhile, we denote by \(\mathcal{D}_n\) the subalgebra of diagonal matrices in \(M_n(\mathbb{C})\), and \(E_{\mathcal{D}_n}: M_n(\mathbb{C}) \to \mathcal{D}_n\) denotes the projection onto the diagonal matrices, which zeroes out the off-diagonal entries of the input matrix (this is understood as a non-commutative conditional expectation in §4.1).

Definition 1. A (normalized) real Ginibre matrix is a random matrix \(A\) in \(M_n(\mathbb{R})\) such that the entries of \(A\) are i.i.d.Gaussian random variables with mean \(0\) and variance \(1/n\).

Definition 2. A (normalized) Gaussian orthogonal ensemble (GOE) matrix is a random self-adjoint matrix \(A\) in \(M_n(\mathbb{R})\) such that the entries \(\{A_{i,j}: i \leqslant j\}\) are independent, the entries \(A_{i,i}\) are Gaussian with mean zero and variance \(2/n\), and the entries \(A_{i,j}\) for \(i \neq j\) are normal with mean zero and variance \(1/n\).

Lemma 1. If \(A\) is a real Ginibre random matrix, then \((A + A^{\mathsf{T}})/\sqrt{2}\) is a GOE random matrix.

Finally, while functions such \(V_{\beta,\gamma,n}\) depend on many different parameters, we will sometimes suppress some of the dependencies in the course of the technical arguments, to prevent the equations from becoming unwieldy. Indeed, \(\beta\), \(\gamma\), and \(n\) will be fixed for most of our analysis, especially in §4 and §5. The dependence on these parameters will be included only when it is relevant to the argument, such as in §6.

2 Primal Parisi PDE↩︎

The PHA algorithm considers an extension of the original objective function into the convex hull of the original domain. The main component in the potential is the Fenchel-Legendre dual, or convex conjugate, of the Parisi solution \(\Phi = \Phi_{\beta,\mu_\beta}\), given by \[\label{eq:32Lambda32def} \Lambda_\beta(t,y) = \sup_{x \in \mathbb{R}} \left( xy - \Phi_{\beta,\mu_\beta}(t,x) \right),\tag{9}\] which turns out to be continuous on \([-1,1]\) and infinite outside this interval. Note that this sign convention for \(\Lambda\) is the opposite of the one in [2]. Moreover, as mentioned in §1.4, we suppress the dependence on \(\beta\) in the notation. The point \(x\) where the supremum is achieved satisfies \(y = \partial_x \Phi(t,x)\) and \(x = \partial_y \Lambda(t,y)\), so that \(\partial_x \Phi(t,\cdot)\) and \(\partial_y \Lambda(t,\cdot)\) describe a change of coordinates between \((-1,1)\) and \(\mathbb{R}\), and hence also from the cube \((-1,1)^n\) to \(\mathbb{R}^n\). Montanari’s algorithm works with the dual coordinate \(x\) (in the same way as the Parisi formula itself) and expresses the near optimizer as \(\partial_x \Phi(q_\beta^*,x)\), but we want to work directly in the primal space with the coordinate \(y\).

Conceptually, \(\Lambda_\beta\) represents the entropy of the current point in the cube, and at each step a certain amount of entropy will be spent as we move closer to a corner. In fact, \(\Lambda_\beta(1,y)\) is equal to the Shannon entropy function of a Bernoulli distribution with probabilities \((1 + y)/2\) and \((1 - y)/2\) (compare ). Our modified objective function can thus be understood as the original objective \(\angles{ \sigma, A \sigma}\) minus the entropy budget that we have to spend (given by \(\sum_j \Lambda_\beta(t,\sigma_j)\)) and the energy gain that we hope to achieve (given by \(r_{\beta,\mu_\beta}\)), which will balance out to zero if the algorithm achieves its goal. This idea is inspired by the Gibbs variational principle that the free energy is the expected internal energy plus the Shannon entropy times temperature, which has been in the background of the Parisi formula from the beginning.

Though working with the point \(y\) in the original domain, rather than the point \(x\) in the dual domain, is intuitive, significant challenges arise from \(\partial_y \Lambda_\beta(t,\cdot)\) blowing up at \(\pm 1\). Therefore, we use a regularized version of \(\Lambda_\beta\); for \(\gamma > 0\), let \[\begin{align} \tilde{\Lambda}_{\beta,\gamma}(t,y) &= \sup_{x \in \mathbb{R}} \left( xy - \Phi_{\beta,\mu_\beta}(t,x) - \frac{\gamma}{2} x^2 \right) \tag{10} \\ &= \inf_{y' \in \mathbb{R}} \left( \Lambda_\beta(t,y') + \frac{1}{2\gamma} (y' - y)^2 \right). \tag{11} \end{align}\] This function \(\tilde{\Lambda}_{\beta,\gamma}\) will be smooth on all of \(\mathbb{R}^n\) with its second derivative bounded by \(1/\gamma\). Thus, \(\tilde{\Lambda}_{\beta,\gamma}\) serves to modify our objective in the primal coordinate space, while also having similar smoothness properties as \(\Phi\). While using the function \(\Lambda_\beta\) would theoretically prevent iterates from leaving the cube, this would require the step size to be made sufficiently small when the point is near the boundary; using the function \(\tilde{\Lambda}_{\beta,\gamma}\) will allow the coordinates of the iterates to leave the cube, but with high probability they will not be too far away, as we will see later on.

The goal of this section is to establish analytic properties for \(\Phi = \Phi_{\beta,\mu_\beta}\) and \(\tilde{\Lambda} = \tilde{\Lambda}_{\beta,\gamma}\) as groundwork for our algorithm. As noted in §1.4, the dependence on \(\beta\) and \(\gamma\) will sometimes be suppressed in the notation. Moreover, many of the properties hold for a general probability measure \(\mu\) and not only the optimizer \(\mu_\beta\). Thus, for the bulk of the section, we will study the functions \(\Phi\), \(\Lambda\), and \(\tilde{\Lambda}_\gamma\) associated to a fixed \(\mu\) and \(\beta\). In particular, we will:

  • Recall regularity properties for the Parisi solution \(\Phi\) from the literature.

  • Prove regularity properties for \(\tilde{\Lambda}_\gamma\) from Fenchel-Legendre duality.

  • Write down a differential equation for \(\tilde{\Lambda}_\gamma\).

  • Translate the Auffinger-Chen SDE into the primal coordinate space, writing equivalent SDEs associated to \(\tilde{\Lambda}_\gamma\).

  • Obtain uniform continuity estimates for \(\Lambda\) as a function \(t\) and \(y\), as well as convergence estimates for \(\tilde{\Lambda}_\gamma\) as \(\gamma \to 0\).

2.1 Smoothness of the Parisi solution↩︎

First, we gather some results on the regularity of solutions to the Parisi PDE; see also [4], [5], [39].

Proposition 8 (Convexity and smoothness for \(\Phi\) [4] and [5]). Let \(\Phi\) be a strong solution to the Parisi PDE for some measure \(\mu\) with initial condition \(\Phi(1,x) = \log (2\cosh(x))\). Then

  1. Strict convexity: \(\Phi\) is strictly convex in \(x\); see [2].

  2. Smoothness: The spatial derivatives \(\partial_x^k \Phi\) exist for all \(k \geqslant 0\), and each derivative is bounded and continuous on \([0,1] \times \mathbb{R}\) [39].

  3. Smoothness in time: The weak derivatives \(\partial_t^{\pm} \partial_x^k \Phi\) exist and are bounded [5], [39]. Moreover, the left and right-hand time derivatives \(\partial_t^{\pm} \partial_x^k \Phi\) exist and they agree whenever \(t\) is a point of continuity of \(F_\mu\) [39]. See also [4] and [5].

As mentioned above, since the maximizer in 9 satisfies \(y = \partial_x \Phi(t,x)\), it will be important for us to get more refined control over \(\partial_x \Phi(t,x)\). From Hopf-Cole computations, it is not hard to see that, when \(\mu\) is atomic, the derivative \(\partial_x \Phi(t,x)\) goes to \(1\) as \(x \to \infty\) and \(-1\) as \(x \to -\infty\). In fact, one can give a uniform bound as follows.

Proposition 9 (Refined bounds for \(\partial_x \Phi\) and \(\partial_{x,x}\Phi\)). Let \(\Phi\) be a strong solution to the Parisi PDE for some measure \(\mu\) with initial condition \(\Phi(1,x) = \log (2\cosh(x))\). Then \[1 - 2 \exp(8 \beta^2(1 - t)) \exp(-2x) \leqslant\mathop{\mathrm{sgn}}(x) \partial_x \Phi(t,x) \leqslant 1.\] and \[e^{-6 \beta^2(1 - t)} \mathop{\mathrm{sech}}(x)^2 \leqslant\partial_{x,x} \Phi(t,x) \leqslant 1.\]

A similar bound for \(\partial_{x,x} \Phi\) is given in [44], [4]. Rather than approximating by atomic measures as in those works, we use a stochastic expression for \(\partial_x \Phi\) due to Jagannath and Tabasco [39].

Lemma 2 (See [39]). Let \(x_0 \in \mathbb{R}\) and \(t_0 \in [0,1]\), and for \(t \in [t_0,1]\) let \(X_t\) be the solution to the Auffinger-Chen SDE \[dX_t = \sqrt{2} \beta \,dW_t + 2 \beta^2 F_\mu(t)\partial_x \Phi(t,X_t)\,dt\, ,\] with initial condition \(X_{t_0} = x_0\).Then, \[\begin{align} \partial_x \Phi(t_0,x_0) &= \mathop{{}\mathbb{E}}[\tanh(X_1)]\,, \\ \partial_{x,x} \Phi(t_0,x_0) &= \mathop{{}\mathbb{E}}\left[ \mathop{\mathrm{sech}}(X_1)^2 + 2 \beta^2 \int_{t_0}^1 F_\mu(s) \partial_{x,x} \Phi(s,X_s)^2\,ds \right]\,,\\ \partial_{x,x,x}\Phi(t_0,x_0) &= -2\mathop{{}\mathbb{E}}\left[\tanh(X_1)\mathop{\mathrm{sech}}^2(X_1)\right] + 6\beta^2\mathop{{}\mathbb{E}}\left[\int_{t_0}^1 F_\mu(s)\partial_{x,x,x}\Phi(s,X_s)\partial_{x,x}\Phi(s,X_s)ds\right]\,. \end{align}\]

The other ingredient that we need is the following moment generating function bound for \(X_t\) proved by Itô calculus.

Lemma 3 (MGF estimate for AC solution). Let \(X_t\) be the solution to the Auffinger-Chen SDE with initial condition \(X_{t_0} = x_0\). Then for \(\lambda \in \mathbb{R}\), \[\mathop{{}\mathbb{E}}[\exp(\lambda X_t)] \leqslant\exp(\beta^2 (2|\lambda| + \lambda^2) (t - t_0)) \exp(\lambda x_0)\]

Proof. By Itô calculus (see ), \[\begin{align} d \exp(\lambda X_t) &= \lambda \exp(\lambda X_t) \,dX_t + \frac{\lambda^2}{2} \exp(\lambda X_t)(dX_t)^2 \\ &= \lambda \exp(\lambda X_t) \sqrt{2} \beta \,dW_t + \lambda \exp(\lambda X_t) 2 \beta^2 F_\mu(t) \partial_x \Phi(t,X_t)\,dt + 2 \beta^2 \frac{\lambda^2}{2} \exp(\lambda X_t)\,dt. \end{align}\] Taking expectations (which makes the \(dW_t\) term vanish) and then differentiating yields, \[\begin{align} \frac{d}{dt} \mathop{{}\mathbb{E}}[\exp(\lambda X_t)] &= 2 \beta^2 |\lambda| \mathop{{}\mathbb{E}}[\exp(\lambda X_t) F_\mu(t) \partial_x \Phi(t,X_t)] + \beta^2 \lambda^2 \mathop{{}\mathbb{E}}[\exp(\lambda X_t)] \\ &\leqslant\beta^2 (2 |\lambda| + \lambda^2) \mathop{{}\mathbb{E}}[\exp(\lambda X_t)]\, , \end{align}\] where we use the fact that \(X_t\) is finite almost surely and that \(|\partial_x \Phi(t,x)| \leqslant 1\) (by ). Then, by Grönwall’s inequality, \[\mathop{{}\mathbb{E}}[\exp(\lambda X_t)] \leqslant\exp(\beta^2(2|\lambda| + \lambda^2)(t-t_0)) \mathop{{}\mathbb{E}}[\exp(\lambda X_{t_0})] = \exp(\beta^2 (2|\lambda| +\lambda^2)(t-t_0)) \exp(\lambda x_0). \qedhere\] ◻

Proof of . Write \((t_0,x_0)\) for the point where we want to bound \(\partial_x \Phi\). Let \(X_t\) be as in . The upper point on \(\partial_x \Phi\) follows from the fact that \(\tanh(X_1) \leqslant 1\). For the lower bound, note that \[1 - \tanh(X_1) = \frac{2 \exp(-X_1)}{\exp(X_1) + \exp(-X_1)} \leqslant 2 \exp(-2X_1)\,.\] Hence, taking the expectation \[1 - \partial_x \Phi(t_0,x_0) \leqslant 2 \mathop{{}\mathbb{E}}[\exp(-2X_1)] \leqslant 2 \exp(8 \beta^2(1 - t_0)) \exp(-2x_0). \qedhere\] The upper bound \(\partial_{x,x} \Phi \leqslant 1\) is given in [39]. For the lower bound, note by that \[\partial_{x,x} \Phi(t_0,x_0) \geqslant\mathop{{}\mathbb{E}}[\mathop{\mathrm{sech}}(X_1)^2] \geqslant[\mathop{{}\mathbb{E}}\mathop{\mathrm{sech}}(X_1)]^2 \geqslant\frac{1}{\mathop{{}\mathbb{E}}[\cosh(X_1)]^2},\] where the last line follows because \[1 = \mathop{{}\mathbb{E}}[ \cosh(X_1)^{1/2} \mathop{\mathrm{sech}}(X_1)^{1/2} ] \leqslant(\mathop{{}\mathbb{E}}[\cosh(X_1)] \mathop{{}\mathbb{E}}[\mathop{\mathrm{sech}}(X_1)])^{1/2}.\] Note that by , \[\begin{align} \mathop{{}\mathbb{E}}[\cosh(X_1)] &= \frac{1}{2} \left( \mathop{{}\mathbb{E}}[e^{X_1}] + \mathop{{}\mathbb{E}}[e^{-X_1}] \right) \\ &\leqslant\frac{1}{2} \exp(3 \beta^2(1-t_0))[e^{x_0} + e^{-x_0}]. \end{align}\] Hence, \[\partial_{x,x} \Phi(t_0,x_0) \geqslant[\exp(3 \beta^2(1 - t_0)) \cosh(x_0)]^{-2},\] which is the desired lower bound. ◻

2.2 Smoothness of the convex conjugates↩︎

Now we turn to the properties of the convex conjugate \(\Lambda\). Many of the claims in this proposition are standard facts about Fenchel-Legendre duality, but the proofs are short in this case so we include them for completeness.

Proposition 10 (Basic properties of \(\Lambda\)). Let \(\Phi\) be a strong solution to the Parisi PDE for some measure \(\mu\) with initial condition \(\Phi(1,x) = \log (2\cosh x)\). Let \[\Lambda(t,y) = \sup_{x \in \mathbb{R}} \left( xy - \Phi(t,x) \right).\]

  1. Domain: We have \(\Lambda(t,y) = +\infty\) if and only if \(|y| > 1\).

  2. Gradient and unique maximizer: If \(|y| < 1\), then there is a unique maximizer \(x\) in the formula above, and \(x\) is the unique solution to \(\partial_x \Phi(t,x) = y\). Furthermore, the maximizer is \(x = \partial_y \Lambda(t,y)\).

  3. Smoothness: \(\Lambda(t,y)\) is a \(C^\infty\) function of \(y\) on \((-1,1)\).

Proof. (1) Since \(|\partial_x \Phi| \leqslant 1\), the function \(\Phi(t,x)\) is \(1\)-Lipschitz in \(x\). Hence, if \(|y| > 1\), then \(xy\) grows faster than \(\Phi(t,x)\) as \(x \to \pm \infty\), and so the supremum will be infinite. For the case \(|y| \leqslant 1\), note that \[\int_0^\infty |1 - \partial_x \Phi(t,x)| \,dx < \infty\] by , and therefore \(\Phi(t,x) - x\) is bounded for \(x > 0\). Thus, by evenness of \(\Phi(t,\cdot)\), we have that \(|x| - \Phi(t,x)\) is bounded. In particular for \(y \in [-1,1]\), \(xy - \Phi(t,x)\) is bounded above, so \(\Lambda(t,y) < \infty\).

(2) By the strict convexity of \(\Phi\), the function \(\partial_x \Phi(t,\cdot)\) is strictly increasing. Also, by the limits at \(\pm \infty\) are \(\pm 1\). Hence, \(\partial_x \Phi(t,\cdot)\) defines a bijection \(\mathbb{R}\to (-1,1)\) which has a smooth inverse since \(\partial_{x,x} \Phi > 0\). Hence, for every \(y \in (-1,1)\), there is a unique \(x = [\partial_x \Phi(t,\cdot)]^{-1}(y)\) such that \(\partial_x \Phi(t,x) = y\). This relation implies that \(\partial_x[ xy - \Phi(t,x)] = 0\). Since the function \(xy - \Phi(t,x)\) is concave in \(x\), any critical point must be a maximizer. Writing \(x(y) = [\partial_x \Phi(t,\cdot)]^{-1}(y)\), so by the chain rule \[\partial_y \Lambda(t,y) = \frac{\partial}{\partial y} \left[ xy - \Phi(t,x) \right] = x + y \frac{dx}{dy} - \partial_x \Phi(t,x) \frac{dx}{dy} = x.\] (3) Therefore, since the maximizer is unique, integrating the previous equality gets \[\Lambda(t,y) = y [\partial_x \Phi(t,\cdot)]^{-1}(y) - \Phi(t, [\partial_x \Phi(t,\cdot)]^{-1}(y)).\] By the inverse function theorem, \([\partial_x \Phi(t,\cdot)]^{-1}\) is \(C^\infty\) since \(\Phi \in C^\infty\) by and \(\partial_{x,x}\Phi > 0\). It follows that \(\Lambda(t,\cdot)\) is a smooth function of \(y\) because the \(C^{\infty}\) property is closed under finite sums and products and compositions. ◻

Now we estimate \(\partial_y \Lambda\) and hence obtain uniform continuity of \(\Lambda\) on \([-1,1]\).

Lemma 4 (Continuity estimates for \(\Lambda\)). Continue the same setup as the previous proposition.

  1. Estimate for first derivative of \(\Lambda\): \[|\partial_y \Lambda(t,y)| \leqslant\frac{1}{2} \log \frac{2}{1-|y|} + 4 \beta^2(1-t).\]

  2. Uniform continuity: For \(y, y' \in [-1,1]\), we have \[|\Lambda(t,y') - \Lambda(t,y)| \leqslant\frac{1}{2} |y' - y| \left(\log \frac{2}{|y'-y|} + 1 + 8 \beta^2(1 - t) \right) \leqslant(1 + 8 \beta^2) \left| \frac{y'-y}{2} \right|^{\frac{8\beta^2}{1 + 8 \beta^2}}.\]

Proof. (1) Fix \(y\), and let \(x = \partial_y \Lambda(t,y)\), so that \(y = \partial_x \Phi(t,x)\). Then we have \[1 - y = 1 - \partial_x \Phi(t,x) \leqslant 2 \exp(8\beta^2(1-t)) \exp(-2x).\] Thus, \[\log(1 - y) \leqslant\log(2) + 8 \beta^2(1 - t) - 2x.\] and \[2x \leqslant-\log(1 - y) + \log(2) + 8 \beta^2(1-t)\,,\] and so \[\partial_y \Lambda(t,y) \leqslant\frac{1}{2} \log \frac{2}{1-y} + 4 \beta^2(1-t).\] This is the bound we want in the case that \(y \geqslant 0\), and the bound follows in the negative case too since \(\partial_y \Lambda\) is an odd function of \(y\).

(2) Since \(\Lambda\) is an even function of \(y\), we have \[|\Lambda(t,y') - \Lambda(t,y)| = |\Lambda(t,|y'|) - \Lambda(t,|y|)|.\] Therefore, since also \(||y'| - |y|| \leqslant|y' - y|\), it suffices to prove the claim when \(y', y \geqslant 0\). Furthermore, assume without loss of generality that \(y' \geqslant y \geqslant 0\). Thus, \[\begin{align} 0 \leqslant\Lambda(t,y') - \Lambda(t,y) &= \int_y^{y'} \partial_y \Lambda(t,u)\,du \\ &\leqslant\int_y^{y'} \left( \frac{1}{2} \log \frac{2}{1-u} + 4 \beta^2 (1 - t) \right)\,du. \end{align}\] Now since right-hand side is an increasing function of \(u\), shifting the interval of integration from \([y,y']\) to \([1-(y'-y),1]\) can only increase the value. Hence, \[\begin{align} \Lambda(t,y') - \Lambda(t,y) &\leqslant\int_{1-(y'-y)}^1 \left( \frac{1}{2} \log \frac{2}{1-u} + 4 \beta^2 (1 - t) \right)\,du \\ &= \int_0^{(y'-y)/2} \log \frac{1}{v}\,dv + 4 \beta^2 (1 - t)(y'-y), \end{align}\] where we made the substitution \(1 - u = 2v\) in the integral of the first term. By integration by parts, \[\begin{align} \int_0^w \log \frac{1}{v}\,dv &= -\int_0^w \log v\,dv \\ &= \left[ -v \log v \right]_0^w + \int_0^w v(1/v)\,dv \\ &= -w \log w + w \\ &= w\left(\log \frac{1}{w} + 1 \right). \end{align}\] Thus, we get \[\Lambda(t,y') - \Lambda(t,y) \leqslant\frac{y' - y}{2} \left(\log \frac{2}{y'-y} + 1 + 8 \beta^2(1 - t) \right).\] This is the desired bound in the case where \(y' \geqslant y \geqslant 0\), which implies the general case.

Finally, to show the last claim, observe that \[\begin{align} \left(\log \frac{2}{|y'-y|} + 1 + 8 \beta^2(1-t) \right) &\leqslant\left(\log \frac{2}{|y'-y|} + 1 + 8 \beta^2 \right) \\ &= (1 + 8 \beta^2) \left( 1 + \frac{1}{1 + 8 \beta^2} \log \frac{2}{|y'-y|} \right) \\ &\leqslant(1 + 8 \beta^2) \exp \left( \frac{1}{1 + 8 \beta^2} \log \frac{2}{|y'-y|} \right) \\ &= (1 + 8 \beta^2) \left| \frac{y' - y}{2} \right|^{-\frac{1}{1 + 8\beta^2}}. \end{align}\] Therefore, \[\frac{|y' - y|}{2} \left(\log \frac{2}{|y'-y|} + 1 + 8 \beta^2(1 - t) \right) \leqslant(1 + 8 \beta^2) \left| \frac{y' - y}{2} \right|^{1 - \frac{1}{1 + 8\beta^2}},\] which is the desired estimate. ◻

The functions \(\tilde{\Lambda}_{\gamma}\) have similar properties as in that can be described as follows.

Proposition 11 (Properties of the regularized conjugate \(\tilde{\Lambda}_{\gamma}\)). Let \(\Phi\) be a strong solution to the Parisi PDE for some measure \(\mu\) with initial condition \(\Phi(1,x) = \log (2\cosh x)\). Let \[\tilde{\Lambda}_{\gamma}(t,y) = \sup_{x \in \mathbb{R}} \left( xy - \Phi(t,x) - \frac{\gamma}{2} x^2 \right).\] Then

  1. Domain: \(\tilde{\Lambda}_{\gamma}(t,y) < \infty\) for all \(y \in \mathbb{R}\).

  2. Unique maximizer: For each \(y \in \mathbb{R}\), there is a unique maximizer \(x\) in the above formula, which is the unique solution to \(y = \partial_x \Phi(t,x) + \gamma x\).

  3. Smoothness: \(\tilde{\Lambda}_{\gamma}\) is a \(C^\infty\) function of \(y\).

  4. Gradient and maximizer: The maximizer \(x\) is given by \(x = \partial_y \tilde{\Lambda}_{\gamma}(t,y)\).

Proof. (1) Note that since \(|x| - \Phi(t,x)\) is bounded, the function \(xy - \Phi(t,x) - \frac{\gamma}{2} x^2\) is bounded above by a quadratic with negative leading term and hence the supremum is finite.

(2) Note that \(x \mapsto \partial_x \Phi(t,x) + \gamma x\) is a smooth strictly increasing function with derivative bounded below, and therefore it defines a bijection \(\mathbb{R}\to \mathbb{R}\) with smooth inverse. The rest of the argument for (2), as well as (3) and (4), is the same as . ◻

Proposition 12 (Relating \(\Lambda\) and \(\tilde{\Lambda}_{\gamma}\)). Continue the same setup from the previous propositions. Then

  1. Inf-convolution formula: \[\tilde{\Lambda}_{\gamma}(t,y) = \inf_{y' \in [-1,1]} \Lambda(t,y') + \frac{1}{2 \gamma} (y' - y)^2.\]

  2. Unique minimizer: The formula above has a unique minimizer \(y'\) that satisfies \(y' + \gamma \partial_y \Lambda(t,y') = y\) and \(y' = y - \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,y)\).

  3. Uniform convergence as \(\gamma \to 0\): For \(y \in [-1,1]\), we have \[\Lambda(t,y) \geqslant\tilde{\Lambda}_{\gamma}(t,y) \geqslant\Lambda(t,y) - (1 + 4 \beta^2) (2 \beta^2 \gamma)^{4\beta^2/(1 + 4 \beta^2)}.\]

Proof. We prove (1) and (2) together. From the proof of , we know that \(\partial_y \Lambda(t,\cdot) = [\partial_x \Phi(t,\cdot)]^{-1}\) is an increasing diffeomorphism from \((-1,1)\) to \(\mathbb{R}\), and hence so is \(\mathop{\mathrm{id}}+ \gamma \partial_y \Lambda(t, \cdot)\). Thus, there is a unique point \(y'\) with \(y' + \gamma \partial_y \Lambda(t,y') = y\). Direct computation shows that \(y'\) is a critical point of \(\Lambda(t,y') + \frac{1}{2 \gamma} (y' - y)^2\). Since that function is strictly convex on \([-1,1]\), the critical point must be the minimizer.

To solve for the minimum value, first write \(x = \partial_y \Lambda(t,y')\), so that \(y' = \partial_x \Phi(t,x)\), and thus \(y = y' + \gamma \partial_y \Lambda(t,y') = y' + \gamma x = \partial_x \Phi(t,x) + \gamma x\). Hence, \(x\) is the maximizer in the definition of \(\tilde{\Lambda}_{\gamma}(t,y)\) and also the maximizer in the definition of \(\Lambda(t,y')\), so that \[\begin{align} \tilde{\Lambda}_{\gamma}(t,y) &= xy - \Phi(t,x) - \frac{\gamma}{2} x^2 \\ &= x(y' + \gamma x) - \Phi(t,x) - \frac{\gamma}{2} x^2 \\ &= xy' - \Phi(t,x) + \frac{\gamma}{2} x^2 \\ &= \Lambda(t,y') + \frac{1}{2 \gamma} (y' - y)^2 \\ &= \inf_{y'' \in [-1,1]} \Lambda(t,y'') + \frac{1}{2 \gamma} (y'' - y)^2, \end{align}\] which finishes the proof of (1). Furthermore, since \(x\) is the maximizer for \(\tilde{\Lambda}_{\gamma}(t,y)\), we have \(\partial_y \tilde{\Lambda}_{\gamma}(t,y) = x = \partial_y \Lambda(t,y')\), and so \[y' = y - \gamma \partial_y \Lambda(t,y') = y - \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,y),\] which finishes the proof of (2).

(3) The upper bound on \(\tilde{\Lambda}_{\gamma}\) follows by taking \(y' = y\) as a candidate for the infimum in (1). For the lower bound, note by that \[\Lambda(t,y') + \frac{1}{2 \gamma} (y' - y)^2 \geqslant\Lambda(t,y) - (1 + 8 \beta^2) \left( \frac{|y'-y|}{2} \right)^{1 - 1/(1 + 8 \beta^2)} + \frac{2}{\gamma} \left( \frac{|y'-y|}{2} \right)^2.\] Let us find the minimum value of the function \(f: [0,\infty) \to \mathbb{R}\) given by \[f(s) = -(1 + 8 \beta^2) s^{8\beta^2/(1 + 8 \beta^2)} + \frac{2}{\gamma} s^2.\] Note \[f'(s) = -8 \beta^2 s^{-1/(1 + 8 \beta^2)} + \frac{4}{\gamma} s,\] so critical points occur when \[\begin{align} 2 \beta^2 s^{-1/(1+8\beta^2)} &= \frac{1}{\gamma} s \\ 2 \beta^2 \gamma &= s^{1 + 1/(1 + 8 \beta^2)} = s^{(2 + 8\beta^2)/(1 + 8 \beta^2)} \\ s &= (2 \beta^2 \gamma)^{(1 + 8 \beta^2)/(2 + 8 \beta^2)}, \end{align}\] and here \[\begin{align} f(s) &= -(1 + 8 \beta^2) (2 \beta^2 \gamma)^{8 \beta^2/(2 + 8 \beta^2)} + \frac{2}{\gamma} (2 \beta^2 \gamma)^{2(1+8\beta^2)/(2 + 8 \beta^2)} \\ &= -(1 + 8 \beta^2) (2 \beta^2 \gamma)^{8 \beta^2/(2 + 8 \beta^2)} + \frac{2}{\gamma}(2 \beta^2 \gamma) (2 \beta^2 \gamma)^{8\beta^2/(2 + 8 \beta^2)} \\ &= -(1 + 4 \beta^2) (2 \beta^2 \gamma)^{4\beta^2/(1 + 4 \beta^2)}. \end{align}\] Since \(f(0) = 0\) and \(\lim_{s \to \infty} f(s) = \infty\), this is the minimum and hence gives a lower bound for \(f(|y'-y|/2)\), which proves the desired lower bound for \(\tilde{\Lambda}_{\gamma}(t,y)\). ◻

2.3 Primal Parisi PDE and primal Auffinger-Chen SDE↩︎

Proposition 13 (Primal Parisi PDE and regularity properties). Let \(\mu\) be a probability measure on \([0,1]\), and let \(\Phi\) be the corresponding solution to the Parisi PDE with initial condition \(\Phi(1,x) = \log( 2 \cosh(x))\). Let \(\gamma \geqslant 0\).

  1. Lipschitz estimates: \(\partial_y^k \tilde{\Lambda}_{\gamma}(t,y)\) is Lipschitz in \(t\), uniformly for \(y\) in each compact subinterval \([-a,a]\) of \(\mathop{\mathrm{dom}}(\tilde{\Lambda}_{\gamma})\) (which is \((-1,1)\) if \(\gamma = 0\) and \(\mathbb{R}\) otherwise).

  2. Differentiability in \(t\): Let \(t_0\) be a point where \(F_\mu\) is continuous and let \(y_0 \in \mathop{\mathrm{dom}}(\tilde{\Lambda}_{\gamma})\). Then for \(k \in \mathbb{N}\), \(\partial_t \partial^k \tilde{\Lambda}_{\gamma}(t,y)\) exists at \((t_0,y_0)\).

  3. Partial differential equation: Let \(t_0\) a continuity point for \(F_\mu\), and \(y_0 \in \mathop{\mathrm{dom}}(\tilde{\Lambda}_{\gamma})\), we have \[\partial_t \tilde{\Lambda}_{\gamma}(t_0,y_0) = \beta^2 \left( \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_0,y_0)} - \gamma + F_\mu(t)(y_0 - \gamma \partial_y \tilde{\Lambda}_{\gamma}(t_0,y_0))^2 \right).\]

  4. Boundary conditions: We have \[\tilde{\Lambda}_{\gamma}(1,y) = \inf_{y'}\left(\Lambda(1,y') + \frac{1}{2\gamma}(y - y')^2\right)\] where \[\Lambda(1,y') = \frac{(1-y')\log(1-y') + (1+y')\log(1+y')}{2} - \log 2\,.\]

Proof. (1) Note that there is some interval \([-b,b]\) such that \(\partial_y \tilde{\Lambda}_{\gamma}(t, \cdot)[-a,a] \subseteq [-b,b]\) for all \(t \in [0,1]\). This follows in the case \(\gamma = 0\) from and in the case \(\gamma > 0\) from . Next, writing \(\tilde{\Phi}_{\gamma}(t,x) = \Phi(t,x) + \frac{\gamma}{2} x^2\), we see that \(\partial_{x,x} \tilde{\Phi}_{\gamma}\) is bounded below on \([-b,b]\) by some constant \(c\) using , and therefore, for \(x, x' \in [-b,b]\), we have \[c|x - x'| \leqslant|\partial_x \tilde{\Phi}_{\gamma}(t,x) - \partial_x \tilde{\Phi}_{\gamma}(t,x')|.\] Furthermore, from , since \(\partial_x^k \Phi\) is weakly differentiable in \(t\) and the derivative is bounded, we know that \(\partial_x^k \Phi\) is Lipschitz in \(t\), uniformly in \(y\).

Now fix \(y \in [-a,a]\) and let \(t, t' \in [0,1]\). Write \[\begin{align} 0 &= y - y \\ &= \partial_x \tilde{\Phi}_{\gamma}(t, \partial_y \tilde{\Lambda}_{\gamma}(t,y)) - \partial_x \tilde{\Phi}_{\gamma}(t',\partial_y \tilde{\Lambda}_{\gamma}(t',y)) \\ &= \partial_x \tilde{\Phi}_{\gamma}(t, \partial_y \tilde{\Lambda}_{\gamma}(t,y)) - \partial_x \tilde{\Phi}_{\gamma}(t, \partial_y \tilde{\Lambda}_{\gamma}(t',y)) + \partial_x \tilde{\Phi}_{\gamma}(t, \partial_y \tilde{\Lambda}_{\gamma}(t',y)) - \partial_x \tilde{\Phi}_{\gamma}(t', \partial_y \tilde{\Lambda}_{\gamma}(t',y)) \end{align}\] We have that \(\partial_x \tilde{\Phi}_{\gamma}(t,y)\) is Lipschitz in \(t\) uniformly in \(y\), and therefore, \(|\partial_x \tilde{\Phi}_{\gamma}(t,x) - \partial_x \tilde{\Phi}_{\gamma}(t',x)| \leqslant L\,|t - t'|\). Therefore \[\begin{align} c\,|\partial_y \tilde{\Lambda}_{\gamma}(t,y) - \partial_y \tilde{\Lambda}_{\gamma}(t',y)| &\leqslant|\partial_x \tilde{\Phi}_{\gamma}(t, \partial_y \tilde{\Lambda}_{\gamma}(t,y)) - \partial_x \tilde{\Phi}_{\gamma}(t, \partial_y \tilde{\Lambda}_{\gamma}(t',y))| \\ &= |\partial_x \tilde{\Phi}_{\gamma}(t, \partial_y \tilde{\Lambda}_{\gamma}(t',y)) - \partial_x \tilde{\Phi}_{\gamma}(t', \partial_y \tilde{\Lambda}_{\gamma}(t',y))| \\ &\leqslant L|t - t'|. \end{align}\] Hence, \(|\partial_y \tilde{\Lambda}_{\gamma}(t,y) - \partial_y \tilde{\Lambda}_{\gamma}(t',y)| \leqslant(L/c) |t - t'|\) as desired.

Lipschitzness and boundedness of \(\partial_y \tilde{\Lambda}_{\gamma}(t,y)\) in \(t\) uniformly for \(y \in [-a,a]\) implies Lipschitzness of \(\tilde{\Lambda}_{\gamma}\) in \(t\) uniformly for \(y \in [-a,a]\) because \(\tilde{\Lambda}_{\gamma}(t,y) = y \partial_y \tilde{\Lambda}_{\gamma}(t,y) - \tilde{\Phi}_{\gamma}(t, \partial_y \tilde{\Lambda}_{\gamma}(t,y))\). To obtain Lipschitzness of the higher-order derivatives, we use the formula \[\partial_{y,y} \tilde{\Lambda}_{\gamma} = \frac{1}{\partial_{x,x} \tilde{\Phi}_{\gamma}(t,\partial_y \tilde{\Lambda}_{\gamma}(t,y))}\, ,\] which is obtained by the observation that \(x = \partial_y\tilde{\Lambda}_{\gamma}(t,y) = \left(\partial_x \tilde{\Phi}_{\gamma}(t,x)\right)^{-1}[y]\) is the unique maximizer, followed by an invocation of the inverse function theorem as \(\partial_x \tilde{\Phi}_{\gamma}(t,x)\) is continuously differentiable. Taking spatial derivatives of this formula expresses \(\partial_y^k \tilde{\Lambda}_{\gamma}(t,y)\) as a composition of spatial derivatives of \(\Phi\) (which are Lipschitz in \(t\) and \(x\)) with \(\partial_y \tilde{\Lambda}_{\gamma}(t,y)\) (which is Lipschitz in \(t\)), where the only term in the denominator is \(\partial_{x,x} \tilde{\Phi}_{\gamma}\) which we know is bounded below on \([-b,b]\). This implies Lipschitzness of the higher derivatives of \(\tilde{\Lambda}_{\gamma}\).

(2) Let \(t_0\) be a point of continuity of \(F_\mu\), so that \(\partial_x^k\Phi\) is differentiable there, and hence so is \(\partial_x^k \tilde{\Phi}_{\gamma}\) since \(\tilde{\Phi}_{\gamma}(t,x) = \Phi(t,x) + (\gamma/2) x^2\). Let \(y \in [-a,a]\) as in the previous argument. Recall \(\partial^k \Phi(t,x)\) is bounded for \(x \in [-b,b]\). Using the same relations as in the previous argument, write \[\begin{align} 0 &= \partial_x \tilde{\Phi}_{\gamma}(t, \partial_y \tilde{\Lambda}_{\gamma}(t,y)) - \partial_x \tilde{\Phi}_{\gamma}(t, \partial_y \tilde{\Lambda}_{\gamma}(t_0,y)) + \partial_x \tilde{\Phi}_{\gamma}(t, \partial_y \tilde{\Lambda}_{\gamma}(t_0,y)) - \partial_x \tilde{\Phi}_{\gamma}(t_0, \partial_y \tilde{\Lambda}_{\gamma}(t_0,y)) \\ &= \partial_{x,x} \tilde{\Phi}_{\gamma}(t,\partial_y \tilde{\Lambda}_{\gamma}(t_0,y)) ( \partial_y \tilde{\Lambda}_{\gamma}(t,y) - \partial_y \tilde{\Lambda}_{\gamma}(t_0,y)) + O(| \partial_y \tilde{\Lambda}_{\gamma}(t,y) - \partial_y \tilde{\Lambda}_{\gamma}(t_0,y)|^2) \\ &\quad + \partial_t\partial_x \tilde{\Phi}_{\gamma}(t_0, \partial_y \tilde{\Lambda}_{\gamma}(t_0,y))(t- t_0) + o(t-t_0). \end{align}\] Since \(\partial_y \tilde{\Lambda}_{\gamma}\) is Lipschitz in \(t\) uniformly for \(y \in [-a,a]\), the term \(O(| \partial_y \tilde{\Lambda}_{\gamma}(t,y) - \partial_y \tilde{\Lambda}_{\gamma}(t_0,y)|^2)\) is \(O(|t - t_0|^2)\), and therefore, \[\partial_y \tilde{\Lambda}_{\gamma}(t,y) - \partial_y \tilde{\Lambda}_{\gamma}(t_0,y) = -\frac{1}{\partial_{x,x} \tilde{\Phi}_{\gamma}(t,\partial_y \tilde{\Lambda}_{\gamma}(t_0,y))} \partial_t \tilde{\Phi}_{\gamma}(t_0, \partial_y \tilde{\Lambda}_{\gamma}(t_0,y))(t - t_0) + o(t - t_0),\] which shows differentiability \[\label{eq:32partial32t32y32Lambda} \partial_t \partial_y \tilde{\Lambda}_{\gamma}(t_0,y) = -\frac{\partial_t \partial_x \tilde{\Phi}_{\gamma}(t_0, \partial_y \tilde{\Lambda}_{\gamma}(t_0,y))}{\partial_{x,x} \tilde{\Phi}_{\gamma}(t_0,\partial_y \tilde{\Lambda}_{\gamma}(t_0,y))}.\tag{12}\] For differentiability of \(\Lambda\) in \(t\) at the point \((t_0,y)\), we again express \(\Lambda\) in terms of \(\partial_y \Lambda\) and \(\Phi\). Similarly, the higher order derivatives of \(\Lambda\) are shown to be differentiable in \(t\) at \(t = t_0\) by expressing them as compositions of derivatives \(\partial_x^k \Phi\) (which are differentiable in \(t\) at \(t = t_0\) by (2)) and \(\partial_y \Lambda\) and using the chain rule.

(3) Let \(t\) be a point of continuity for \(F_\mu\). Note that \[\begin{align} \partial_t \tilde{\Phi}_{\gamma}(t,x) &= \partial_t \Phi(t,x) \\ &= -\beta^2 \left( \partial_{x,x} \Phi(t,x) + F_\mu(t) \partial_x \Phi(t,x)^2 \right) \\ &= -\beta^2 \left( \partial_{x,x} \tilde{\Phi}_{\gamma}(t,x) - \gamma + F_\mu(t) (\partial_x \tilde{\Phi}_{\gamma}(t,x) - \gamma x)^2 \right). \end{align}\] Now differentiate the relation \[\tilde{\Lambda}_{\gamma}(t,y) = y \partial_y \tilde{\Lambda}_{\gamma}(t,y) - \tilde{\Phi}_{\gamma}(t, \partial_y \tilde{\Lambda}_{\gamma}(t,y))\] in \(t\) to obtain \[\partial_t \tilde{\Lambda}_{\gamma}(t,y) = y \partial_t \partial_y \tilde{\Lambda}_{\gamma}(t,y) - \partial_t \tilde{\Phi}_{\gamma}(t,\partial_y \tilde{\Lambda}_{\gamma}(t,y)) - \partial_x \tilde{\Phi}_{\gamma}(t,\partial_y \tilde{\Lambda}_{\gamma}(t,y)) \partial_t \partial_y\tilde{\Lambda}_{\gamma}(t,y).\] Since \(y = \partial_x \tilde{\Phi}_{\gamma}(t, \partial_y \tilde{\Lambda}_{\gamma}(t,y))\), the first and third terms on the right-hand side cancel and \[\begin{align} \partial_t \tilde{\Lambda}_{\gamma}(t,y) &= -\partial_t \tilde{\Phi}_{\gamma}(t, \partial_y \tilde{\Lambda}_{\gamma}(t,y)) \\ &= \beta^2 \left( \partial_{x,x} \tilde{\Phi}_{\gamma}(t, \partial_y \tilde{\Lambda}_{\gamma}(t,y)) - \gamma + F_\mu(t)( \partial_x \tilde{\Phi}_{\gamma}(t,\partial_y \tilde{\Lambda}_{\gamma}(t,y)) - \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,y))^2 \right) \\ &= \beta^2 \left( \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,y)} -\gamma + F_\mu(t)(y - \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,y))^2 \right). \end{align}\]

(4) In the case \(\gamma = 0\), the initial condition arises due to \(\log \left(2\cosh\right)\) being the Fenchel-Legendre conjugate of the binary entropy function rescaled to the domain \([-1,1]\), see for example [45]. The case for \(\gamma > 0\) then follows by . ◻

A recent result of Mourrat [46] gives another rewrite of the Parisi formula using Legendre-Fenchel duality, stated in terms of an expectation over a martingale process. We give a related Itô computation below that describes the transformation of the solution to the Auffinger-Chen SDE from the dual coordinates to the primal coordinates. In light of the proposition below, we refer to \[\label{eq:primal-auffinger-chen} dY_t = \frac{\sqrt{2} \beta}{\partial_{y,y} \Lambda(t,Y_t)}\,dW_t\,,\tag{13}\] as the primal Auffinger-Chen SDE. Later in below, we will study a similar SDE associated to \(\tilde{\Lambda}_{\gamma}\). To justify existence and uniqueness of solutions, note that (3) and (4) below gives Lipschitz estimates for \(1 / \partial_{y,y} \Lambda\), and clearly \(1/\partial_{y,y} \Lambda\) vanishes at \(y = \pm 1\) so that it can be extended by \(0\) on the complement of \((-1,1)\). Finally, we remark that [39] shows that Itô calculus can be applied in this situation, even though the function has limited regularity especially in the time variable.

Proposition 14 (Equivalence between primal and dual AC SDE). Fix \(\beta \in (0,\infty)\). Let \(\mu\) be a probability measure on \([0,1]\), and let \(\Phi\) be the solution to the Parisi PDE. Let \(W_t\) be a standard Brownian motion. Let \(X_t\) and \(Y_t\) be stochastic processes with \(Y_t = \partial_x \Phi(t,X_t)\), or equivalently \(X_t = \partial_y \Lambda(t,Y_t)\). Then, we have \[dY_t = \frac{\sqrt{2} \beta}{\partial_{y,y} \Lambda(t,Y_t)}\,dW_t \iff dX_t = \sqrt{2} \beta \,dW_t + 2 \beta^2 F_{\mu}(t) \partial_x \Phi(t,X_t)\,dt.\]

Proof. Assuming the equation for \(X_t\), we have by Itô calculus that \[\begin{align} dY_t &= d \partial_x \Phi(t,X_t) \\ &= \partial_{x,x} \Phi(t,X_t) \,dX_t + \partial_{t,x} \Phi(t,X_t)\,dt + \frac{1}{2} \partial_{x,x,x} \Phi(t,X_t) (dX_t)^2 \\ &= \sqrt{2} \beta \partial_{x,x} \Phi(t,X_t)\,dW_t + 2 \beta^2 F_\mu(t)\partial_{x,x} \Phi(t,X_t) \partial_x \Phi(t,X_t)\,dt + \partial_{t,x} \Phi(t,X_t)\,dt + \beta^2 \partial_{x,x,x} \Phi(t,X_t)\,dt \\ &= \sqrt{2} \beta \partial_{x,x} \Phi(t,X_t)\,dW_t + \partial_x[\partial_t \Phi + \beta^2 \partial_{x,x} \Phi + \beta^2 F_{\mu}(t) (\partial_x \Phi)^2](t,X_t)\,dt \\ &= \frac{\sqrt{2} \beta}{\partial_{y,y}\Lambda(t,Y_t)}\,dW_t + 0. \end{align}\] Conversely, suppose that \(Y_t\) satisfies this equation and note \(X_t = \partial_y \Lambda(t,Y_t)\). Then by Itô calculus, \[\begin{align} dX_t &= d \partial_y\Lambda(t,Y_t) \\ &= \partial_{y,y} \Lambda(t,Y_t)\,dY_t + \partial_{t,y} \Lambda(t,Y_t)\,dt + \frac{1}{2} \partial_{y,y,y} \Lambda(t,Y_t) (dY_t)^2 \\ &= \sqrt{2} \beta \,dW_t + \partial_{t,y} \Lambda(t,Y_t)\,dt + \beta^2 \frac{\partial_{y,y,y} \Lambda(t,Y_t)}{\left(\partial_{y,y} \Lambda(t,Y_t)\right)^2}\,dt \\ &= \sqrt{2} \beta\,dW_t + \partial_y\left[\partial_t \Lambda - \frac{\beta^2}{\partial_{y,y} \Lambda} \right](t,Y_t)\,dt \\ &= \sqrt{2} \beta\,dW_t + \partial_y\left[\beta^2 F_{\mu}(t) y^2 \right]_{y=Y_t}\,dt \\ &= \sqrt{2} \beta\,dW_t + 2 \beta^2 F_{\mu}(t) Y_t\,dt \\ &= \sqrt{2} \beta\,dW_t + 2 \beta^2 F_{\mu}(t) \partial_x \Phi(t,X_t)\,dt. \qedhere \end{align}\] ◻

2.4 Estimates for the primal solutions↩︎

Recall \(\tilde{\Lambda}_{\gamma}\) is one of the main ingredients in our objective function, which we will seek to estimate at the iterates of our algorithm by using Taylor expansion at each iterate (see ). Hence, to prove the validity of our algorithm, we need as good of bounds on higher derivatives of \(\tilde{\Lambda}_{\gamma}\) as we can find. Moreover, since we study the SDE driven by Brownian motion with the coefficient function \(1 / \partial_{y,y} \tilde{\Lambda}_{\gamma}\), we also want Lipschitz bounds for \(1 / \partial_{y,y} \tilde{\Lambda}_{\gamma}\) in both time and space; this will be crucial for our convergence argument in and energy analysis in . Specifically, we will show the following result.

Proposition 15 (Estimates for the derivatives of \(\tilde{\Lambda}_{\gamma}\)). Let \(\Phi\) and \(\tilde{\Lambda}_{\gamma}\) be as above. Then we have the following estimates for \(\gamma \geqslant 0\) and \(y \in \mathop{\mathrm{dom}}(\tilde{\Lambda}_{\gamma})\).

  1. Bounds for second derivative of \(\tilde{\Lambda}_{\gamma}\): \[\frac{1}{1 + \gamma} \leqslant\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,y) \leqslant\frac{1}{\gamma}.\]

  2. Bound for third derivative of \(\tilde{\Lambda}_{\gamma}\): \[\left|\partial_{y,y,y} \tilde{\Lambda}_{\gamma}(t,y)\right| \leqslant\frac{2}{\gamma^2}.\]

  3. Spatial Lipschitz estimate for \(1 / \partial_{y,y} \tilde{\Lambda}_{\gamma}\): \[\left| \partial_y \left( \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,y)} \right) \right| \leqslant 2.\]

  4. Temporal Lipschitz estimate for \(1 / \partial_{y,y} \tilde{\Lambda}_{\gamma}\): \[\left| \partial_t \left( \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,y)} \right) \right| \leqslant 18 \beta^2.\]

The proof proceeds in several steps:

  1. For atomic \(\mu\), we express the solution \(\Phi(t,x)\) using Ruelle probability cascades.

  2. We obtain estimates comparing various derivatives of \(\Phi\) for atomic \(\mu\), which we then extend to arbitrary measures by density.

  3. We express derivatives of \(\tilde{\Lambda}_{\gamma}\) in terms of derivatives of \(\Phi\) and estimate them.

The description of the Parisi formula in terms of RPCs is standard in spin-glass theory, but we provide some explanation in . In particular, we use the explicit RPC-based representation for \(\Phi\) over atomic measures to estimate uniform bounds on the spatial derivatives of \(\Phi\).

Lemma 5 (Polynomial expressions for \(\partial^{(j)}_{x}\Phi\) via the RPC representation). Let \(0 \leqslant t_0 < \dots < t_r = 1\), and fix a finitely supported probability measure \(\mu\) with \(\mathrm{supp}(\mu) = [0,t_0] \cup \{t_1,\dots,t_r\}\). Then there exists a random variable \(T(x)\) depending on \(x \in \mathbb{R}\) such that

  1. \(|T(x)| \leqslant 1\).

  2. \(T(x)\) is differentiable in \(x\) and \(T'(x) = 1 - T(x)^2\).

  3. Let \(F_j(\tau)\) be the polynomial given recursively by \[F_1(\tau) = \tau, \qquad F_{j+1}(\tau) = F_j'(\tau)(1 - \tau^2).\] Then for all \(x\), \[\partial_x^j \Phi(t_0,x) = \mathop{{}\mathbb{E}}[F_j(T(x))].\] In particular, \[\begin{align} \partial_x \Phi(t,x) &= \mathop{{}\mathbb{E}}[T(x)] \\ \partial_{x,x} \Phi(t,x) &= \mathop{{}\mathbb{E}}[1 - T(x)^2] \\ \partial_{x,x,x} \Phi(t,x) &= \mathop{{}\mathbb{E}}[-2T(x)(1 - T(x)^2)] \\ \partial_{x,x,x,x} \Phi(t,x) &= \mathop{{}\mathbb{E}}[-2(1 - 3 T(x)^2)(1 - T(x)^2)]. \end{align}\]

Proof. As in the proof of [44], we use the Ruelle Probability Cascade construction, which we include for the reader’s convenience in . By , there exist nonnegative random variables \((v_\alpha)_{\alpha \in \mathbb{N}^r}\) such that \(\sum_{\alpha \in \mathbb{N}^r} v_\alpha = 1\), and random variables \((Z_\alpha)_{\alpha \in \mathbb{N}^r}\) such that for all \(x\), \[\Phi(t_0,x) = \mathop{{}\mathbb{E}}\log \sum_{\alpha \in \mathbb{N}^r} 2 v_\alpha \cosh(x + Z_\alpha).\] Now let \[T(x) = \frac{d}{dx} \log \sum_{\alpha \in \mathbb{N}^r} 2 v_\alpha \cosh(x + Z_\alpha) = \frac{\sum_{\alpha \in \mathbb{N}^r} v_\alpha \sinh(x + Z_\alpha)}{\sum_{\alpha \in \mathbb{N}^r} v_\alpha \cosh(x + Z_\alpha)}.\] Since \(|\sinh| \leqslant|\cosh|\), we have \(|T(x)| \leqslant 1\). Also, by direct computation \[\begin{align} T'(x) &= \frac{d}{dx} \frac{\sum_{\alpha \in \mathbb{N}^r} v_\alpha \sinh(x + Z_\alpha)}{\sum_{\alpha \in \mathbb{N}^r} v_\alpha \cosh(x + Z_\alpha)} \\ &= \frac{\frac{d}{dx} [\sum_{\alpha \in \mathbb{N}^r} v_\alpha \sinh(x + Z_\alpha)]}{\sum_{\alpha \in \mathbb{N}^r} v_\alpha \cosh(x + Z_\alpha)} - \frac{[\sum_{\alpha \in \mathbb{N}^r} v_\alpha \sinh(x + Z_\alpha)] \frac{d}{dx} [\sum_{\alpha \in \mathbb{N}^r} v_\alpha \cosh(x + Z_\alpha)]}{[\sum_{\alpha \in \mathbb{N}^r} v_\alpha \cosh(x + Z_\alpha)]^2} \\ &= \frac{\sum_{\alpha \in \mathbb{N}^r} v_\alpha \cosh(x + Z_\alpha)}{\sum_{\alpha \in \mathbb{N}^r} v_\alpha \cosh(x + Z_\alpha)} - \frac{[\sum_{\alpha \in \mathbb{N}^r} v_\alpha \sinh(x + Z_\alpha)] [\sum_{\alpha \in \mathbb{N}^r} v_\alpha \sinh(x + Z_\alpha)]}{[\sum_{\alpha \in \mathbb{N}^r} v_\alpha \cosh(x + Z_\alpha)]^2} \\ &= 1 - T(x)^2. \end{align}\] Note that \(\log \sum_{\alpha \in \mathbb{N}^r} 2 v_\alpha \cosh(x + Z_\alpha)\) is integrable over the probability space. This random variable above is also \(1\)-Lipschitz in \(x\) since \(|T(x)| \leqslant 1\). Hence, we may apply the bounded convergence theorem to the difference quotients to conclude that \[\partial_x \Phi(t_0,x) = \frac{d}{dx} \mathop{{}\mathbb{E}}\log \sum_{\alpha \in \mathbb{N}^r} 2 v_\alpha \cosh(x + Z_\alpha) = \mathop{{}\mathbb{E}}[T(x)].\] By similar reasoning, for any polynomial \(F(\tau)\), since \((d/dx) F(T(x)) = F'(T(x)) (1 - T(x)^2)\) is bounded, we have \[\frac{d}{dx} \mathop{{}\mathbb{E}}[F(T(x))] = \mathop{{}\mathbb{E}}[F'(T(x))(1 - T(x)^2)].\] We apply this procedure inductively starting with \(F_1(\tau) = \tau\) and this results in claim (3). ◻

Now we are ready to begin estimating the derivatives of \(\Phi\).

Lemma 6 (Bounds for \(\partial^{(j)}_{x}\Phi\) in terms of \(\partial_{x,x}\Phi\)). Let \(\Phi\) be a solution to the Parisi equation for some measure \(\mu\) on \([0,1]\). Then for \(j \geqslant 3\), there exists a constant \(C_j\) independent of \(\beta\) such that \[|\partial_x^j \Phi(t,x)| \leqslant C_j \partial_{x,x} \Phi(t,x).\] In particular, we can take \(C_3 = 2\) and \(C_4 = 4\).

Proof. Fix \(t_0\) and assume first that \(\mu\) is atomic. Let \(T(x)\) be as in . Then \[\partial_x^j \Phi(t_0,x) = \mathop{{}\mathbb{E}}[F_j(T(x))] = \mathop{{}\mathbb{E}}[F_{j-1}'(T(x))(1 - T(x)^2)].\] Let \(C_j = \max_{\tau \in [-1,1]} |F_{j-1}'(\tau)|\). Then \[|\partial_x^j \Phi(t_0,x)| \leqslant\mathop{{}\mathbb{E}}[|F_{j-1}'(T(x))|(1 - T(x)^2)] \leqslant C_j \mathop{{}\mathbb{E}}[1 - T(x)^2] = C_j \partial_x^2 \Phi(t_0,x).\] In particular, since \(F_2(\tau) = 1 - \tau^2\) and \(F_2'(\tau) = -2 \tau\), we have \(C_3 = 2\). Moreover, \(F_3(\tau) = -2 \tau(1 - \tau^2)\) and \(F_3'(\tau) = -2(1 - 3 \tau^2)\). Clearly, \(-2 \leqslant 1 - 3 \tau^2 \leqslant 1\), and so \(C_4 = 4\).

It remains to extend the claim from atomic measures to general measures. By [5], if \(\mu_k\) is a sequence of measures with \(\mu_k \to \mu\) in the weak-\(*\) topology, then \(\Phi_{\mu_k} \to \Phi_\mu\) uniformly, and hence our estimates also hold for \(\mu\). ◻

Finally, we deduce the asserted estimates for \(\Lambda\).

Proof of . (1) Fix \(t\), and let \(x(y) = \partial_y \tilde{\Lambda}_{\gamma}(t,y)\), so that \(y = \partial_x \Phi(t, x(y)) + \gamma x(y)\). Recall also by the formula for derivatives of inverse function that \[\frac{dx}{dy} = \partial_{y,y} \tilde{\Lambda}_{\gamma}(t,y) = \frac{1}{\partial_{x,x} \Phi(t,x(y)) + \gamma}.\] Then because \(0 \leqslant\partial_{x,x} \Phi(t,x) \leqslant 1\) by , we get \[\frac{1}{1 + \gamma} \leqslant\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,y) \leqslant\frac{1}{\gamma},\] which is the asserted estimate.

(2) By the chain rule, \[\partial_{y,y,y} \tilde{\Lambda}_{\gamma}(t,y) = \partial_y \left( \frac{1}{\partial_{x,x} \Phi(t,x(y)) + \gamma} \right) = -\frac{\partial_{x,x,x} \Phi(t,x)}{(\partial_{x,x} \Phi(t,x) + \gamma)^2} \frac{dx}{dy} = -\frac{\partial_{x,x,x} \Phi(t,x)}{(\partial_{x,x} \Phi(t,x) + \gamma)^3}.\] Then by , \[\frac{|\partial_{x,x,x} \Phi(t,x)|}{(\partial_{x,x} \Phi(t,x) + \gamma)^3} \leqslant\frac{|\partial_{x,x,x} \Phi(t,x)|}{\gamma^2 \partial_{x,x} \Phi(t,x)} \leqslant\frac{2}{\gamma^2}.\]

(3) By the chain rule, \[\partial_y \left( \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma} (t,y)} \right) = \partial_y \left( \partial_{x,x} \Phi(t,x) + \gamma \right) = \partial_{x,x,x} \Phi(t,x) \frac{dx}{dy} = \frac{\partial_{x,x,x} \Phi(t,x)}{\partial_{x,x} \Phi(t,x) + \gamma}.\] By , this is bounded in absolute value by \(2\).

(4) Note that \[\begin{align} \partial_t \left( \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma} (t,y)} \right) &= \partial_t \left( \partial_{x,x} \Phi(t, \partial_y \tilde{\Lambda}_{\gamma}(t,y)) \right) \\ &= \partial_t \partial_{x,x} \Phi(t, x) + \partial_{x,x,x} \Phi(t,x) \partial_t \partial_y\tilde{\Lambda}_{\gamma}(t,y). \end{align}\] From 12 , \[\begin{align} \partial_t \partial_y \tilde{\Lambda}_{\gamma}(t,y) &= -\frac{\partial_t \partial_x \tilde{\Phi}_{\gamma}(t,x)}{\partial_{x,x} \tilde{\Phi}_{\gamma}(t,x)} \\ &= -\frac{\partial_t \partial_x \Phi(t,x)}{\partial_{x,x} \Phi(t,x) + \gamma}, \end{align}\] so that \[\label{eq:32formula32for32dt32dy32dy} \partial_t \left( \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma} (t,y)} \right) = \partial_t \partial_{x,x} \Phi(t, x) -\partial_{x,x,x} \Phi(t,x)\frac{\partial_t \partial_x \Phi(t,x)}{\partial_{x,x} \Phi(t,x) + \gamma};\tag{14}\] we now proceed to estimate both terms on the right-hand side. By the Parisi PDE , we have \[\begin{align} \partial_t \partial_x \Phi &= -\beta^2 \partial_x \left( \partial_{x,x} \Phi + F_\mu(t) \partial_x (\partial_x \Phi)^2 \right) \\ &= -\beta^2 \left( \partial_{x,x,x} \Phi + 2 F_\mu(t) \partial_x \Phi \partial_{x,x} \Phi \right) \\ \partial_t \partial_{x,x} \Phi &= -\beta^2 \partial_x \left( \partial_{x,x,x} \Phi + 2 F_\mu(t) \partial_x \Phi \partial_{x,x} \Phi \right) \\ &= -\beta^2 \left( \partial_{x,x,x,x} \Phi + 2 F_\mu(t) (\partial_{x,x} \Phi)^2 + 2 F_\mu(t) \partial_x \Phi \partial_{x,x,x} \Phi \right). \end{align}\] In particular, by , \[\begin{align} |\partial_{x,x,x} \Phi(t,x) \cdot \partial_t \partial_x \Phi(t, x)| &\leqslant\beta^2 |\partial_{x,x,x} \Phi|^2 + 2 \beta^2 F_\mu(t) |\partial_x \Phi\partial_{x,x,x} \Phi| \partial_{x,x} \Phi \label{eq:32estimate32for32dt32dx} \\ &\leqslant 4 \beta^2 (\partial_{x,x} \Phi)^2 + 4 \beta^2 (\partial_{x,x} \Phi)^2. \nonumber \end{align}\tag{15}\] and \[\begin{align} |\partial_t \partial_{x,x} \Phi(t, x)| &\leqslant\beta^2 |\partial_{x,x,x,x} \Phi| + 2 \beta^2 F_\mu(t) (\partial_{x,x} \Phi)^2 + 2 \beta^2 F_\mu(t) |\partial_x \Phi \partial_{x,x,x} \Phi| \label{eq:32estimate32for32dt32dx32dx} \\ &\leqslant 4 \beta^2 \partial_{x,x} \Phi + 2 \beta^2 (\partial_{x,x} \Phi)^2 + 4 \beta^2 |\partial_x \Phi| \partial_{x,x} \Phi \nonumber \\ &\leqslant 10 \beta^2. \nonumber \end{align}\tag{16}\] Substituting our estimates 16 for \(\partial_t \partial_{x,x} \Phi\) and 15 for \(\partial_t \partial_x \Phi\) into 14 yields \[\begin{align} \left| \partial_t \left( \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma} (t,y)} \right) \right| &\leqslant|\partial_t \partial_{x,x} \Phi| + \left| \frac{\partial_t \partial_x \Phi(t,x)}{\partial_{x,x} \Phi(t,x) + \gamma} \right| \\ &\leqslant 10 \beta^2 + \frac{8 \beta^2 (\partial_{x,x} \Phi(t,x))^2}{\partial_{x,x} \Phi(t,x) + \gamma} \\ &\leqslant 18 \beta^2. \qedhere \end{align}\] ◻

Remark 16 (Differential equation for \(1 / \partial_{y,y} \Lambda\)). Let \(v(t,y) = 1 / \partial_{y,y} \Lambda(t,y) = \partial_{x,x} \Phi(t,\partial_y \Lambda(t,y))\) for \(|y| < 1\) and set \(v(t,y) = 0\) for \(|y| > 1\). Then \(v\) satisfies the differential equation, \[\partial_t v = -\beta^2 v^2 \left(\partial_{y,y} v + 2 F_\mu(t)\right)\,,\] for \(|y| < 1\) since \[\partial_t v = -\frac{\partial_t \partial_{y,y} \Lambda}{(\partial_{y,y} \Lambda)^2} = -v^2 \partial_{y,y} \partial_t \Lambda = -\beta^2 v^2 \partial_{y,y} \left( \frac{1}{\partial_{y,y} \Lambda} + F_\mu(t) y^2 \right) = -\beta^2 v^2 \left( \partial_{y,y} v + 2 F_\mu(t) \right).\] Moreover, it trivially satisfies the equation when \(|y| > 1\) since both sides are zero. The terminal condition at \(t=1\) can easily be computed from (4) as \(v(1,y) = 1 - y^2\). This equation somewhat resembles the equation for \(\Phi\) itself, but now with the worse nonlinearity \(v^2 \partial_{y,y} v\) instead of the additive nonlinear \((\partial_x \Phi)^2\) term from the Parisi PDE. This equation should be investigated further using the tools of viscosity solutions and free boundary problems.

2.5 Energy simplifications under fRSB↩︎

Previous approaches to optimizing spin glass Hamiltonians required the fRSB regularity property for the measure \(\mu\) that minimizes the Parisi formula [1], [9]. Under fRSB, the Auffinger-Chen SDE has quadratic variation \(\mathop{{}\mathbb{E}}Y_t^2 = t\) for \(t\) up to a certain time \(q_\beta^*\). An equivalent condition is that \[\mathop{{}\mathbb{E}}\frac{2\beta^2}{(\partial_{y,y}\Lambda(t,Y_t))^2} = 1,\] which in our paper is exactly the condition needed to satisfy a self-consistency equation that arises from computing the Cauchy-Stieltjes transform of the spectrum of \(\nabla^2 \mathop{\mathrm{obj}}\) while requiring that the largest eigenvalue of \(\nabla^2 \mathop{\mathrm{obj}}\) is near zero. The precise statement of this fRSB simplification is as follows.

Lemma 7 (fRSB Property of \(Y_t\) [9], [47]).
Recall \(q^*\) from . Assume \(\beta > 0\) is large enough and that holds. Let \(\mu\) be the optimizer in the Parisi formula for a given \(\beta\), and let \(\Phi\) and \(\Lambda\) be the corresponding solutions. Let \(Y_t\) be the solution to the primal AC SDE. Then, for every \(t \in [0, q^*_\beta]\), we have \[\mathop{{}\mathbb{E}}_{Y_t} \left[\frac{2\beta^2}{(\partial_{y,y}\Lambda(t,Y_t))^2}\right] = 1\,.\] Equivalently, \[\mathop{{}\mathbb{E}}_{Y_t} Y_t^2 = t\,.\] Moreover, \[\mathop{{}\mathbb{E}}\left[\frac{1}{\partial_{y,y} \Lambda(t,Y_t)}\right] = \int_t^1\mu([0,s])\,ds = \int_t^1 F_\mu(s)\,ds.\]

A direct consequence is the following formula for \(\mathop{{}\mathbb{E}}\Lambda(t,Y_t)\) at \(t = q_\beta^*\) under fRSB. Montanari [9] shows that it agrees with \(\beta\) times a certain energy functional \(\mathcal{E}_\beta\), which agrees with \(\mathcal{P}_\beta\) in the large \(\beta\) limit.

Corollary 2 (Value of entropy along the AC process). Under Assumption 1, with the same setup as in the previous lemma, we have, \[\mathop{{}\mathbb{E}}\Lambda(q_\beta^*,Y_{q_\beta^*}) - \Lambda(0,0) - \beta^2 \int_0^{q^*_\beta} s F_\mu(s)\,ds = 2 \beta^2 \int_0^{q^*_\beta} \int_s^1 F_\mu(u)\,du\,ds.\]

Proof. To estimate this value, we use  and  to simplify the expected value under the primal Auffinger-Chen dynamics , \[\begin{align} d \Lambda(t,Y_t) &= \partial_y \Lambda(t,Y_t) \,dY_t + \frac{1}{2} \partial_{y,y} \Lambda(t,Y_t) (dY_t)^2 + \partial_t \Lambda(t,Y_t)\,dt \\ &= \frac{\sqrt{2} \beta \partial_y \Lambda(t,Y_t)}{\partial_{y,y} \Lambda(t,Y_t)}\,dW_t + \left( \frac{\beta^2}{\partial_{y,y} \Lambda(t,Y_t)} + \partial_t \Lambda(t,Y_t) \right)\,dt \\ &= \frac{\sqrt{2} \beta \partial_y \Lambda(t,Y_t)}{\partial_{y,y} \Lambda(t,Y_t)}\,dW_t + \left( \frac{2 \beta^2}{\partial_{y,y} \Lambda(t,Y_t)} + \beta^2 F_\mu(t) Y_t^2 \right)\,dt\,. \end{align}\] Hence, after taking expectations and using , \[\begin{align} \frac{d}{dt} \mathop{{}\mathbb{E}}\Lambda(t,Y_t) &= 2 \beta^2 \mathop{{}\mathbb{E}}\frac{1}{\partial_{y,y} \Lambda(t,Y_t)} + \beta^2 F_\mu(t) \mathop{{}\mathbb{E}}[Y_t^2] \\ &= \beta^2\left( 2 \int_t^1 F_\mu(s)\,ds + t F_\mu(t) \right). \end{align}\] We obtain the statement asserted in the lemma by integrating this and noting that \(Y_0 = 0\). ◻

Since our algorithm uses \(\tilde{\Lambda}_{\gamma}\) rather than \(\Lambda\) itself, we want a version of for \(\tilde{\Lambda}_{\gamma}\). For this purpose we compare in Wasserstein distance two processes associated to \(\Lambda\) and \(\tilde{\Lambda}_{\gamma}\). This lemma does not itself invoke fRSB.

Lemma 8 (Wasserstein-\(2\) distance between the non-convolved & convolved primal SDE).
Let \(\Phi\) be the solution to the Parisi equation for some \(\mu\), and let \(\tilde{\Lambda}_{\gamma}\) be as above. Let \(Y_t\) and \(Y_t^\gamma\) solve the equations \[\label{eq:primal-ac-sde} dY_t = \frac{\sqrt{2} \beta}{\partial_{y,y} \Lambda(t, Y_t)}\,dW_t\,,\tag{17}\] and \[\label{eq:regularized-primal-ac-sde} dY_t^\gamma = \frac{\sqrt{2} \beta}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t, Y_t^\gamma)}\,dW_t\,,\tag{18}\] with \(Y^\gamma_0 = 0\) and \(Y_0 = 0\). Then \[\label{eq:32sde-closeness321} \mathop{{}\mathbb{E}}\left|Y_t^\gamma - \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,Y_t^\gamma) - Y_t\right|^2 \leqslant\mathop{{}\mathbb{E}}\left|Y_t^\gamma - (Y_t + \gamma \partial_y \Lambda(t,Y_t))\right|^2 \leqslant\frac{1}{5} \gamma^2 (e^{10\beta^2 t} - 1)\tag{19}\] and \[\label{eq:32sde-closeness322} \left\lVert{Y_t^\gamma - Y_t}\right\rVert_{L^2}^2 \leqslant 2 \gamma^2 (e^{10 \beta^2 t} - 1).\tag{20}\]

Proof. First, let \(Z_t^\gamma = Y_t + \gamma \partial_y \Lambda(t,Y_t)\). Proposition 12 (2) implies that \(\mathop{\mathrm{id}}- \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,\cdot)\) is the inverse function of \(\mathop{\mathrm{id}}+ \gamma \partial_y \Lambda(t,\cdot)\). Thus we have \(Y_t = Z_t^\gamma - \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,Z_t^\gamma)\). Moreover, the function \(\mathop{\mathrm{id}}- \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,\cdot)\) is \(1\)-Lipschitz since its derivative is \(1 - \gamma \partial_{y,y} \tilde{\Lambda}_{\gamma}(t,\cdot)\), which is bounded below by \(1 - \gamma / \gamma = 0\) and bounded above by \(1 - \gamma / (1 + \gamma) \leqslant 1\), using Proposition 15 (1). Hence, we have \[\mathop{{}\mathbb{E}}|Y_t^\gamma - \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,Y_t^\gamma) - Y_t|^2 \leqslant\mathop{{}\mathbb{E}}|Y_t^\gamma - (Y_t + \gamma \partial_y \Lambda(t,Y_t))|^2,\] which is the first inequality in 19 . It remains to estimate \(\mathop{{}\mathbb{E}}|Y_t^\gamma - (Y_t + \gamma \partial_y \Lambda(t,Y_t))|^2\).

It follows from Itô calculus that \[\begin{align} dZ_t^\gamma &= dY_t + \gamma \partial_{y,y}\Lambda(t,Y_t)\,dY_t + \frac{1}{2} \gamma \partial_{y,y,y} \Lambda(t,Y_t) (dY_t)^2 + \gamma \partial_{t,y} \Lambda(t,Y_t)\,dt \\ &= \sqrt{2} \beta \left( \frac{1}{\partial_{y,y} \Lambda(t,Y_t)} + \gamma \right)\,dW_t + \gamma \left( \beta^2 \frac{\partial_{y,y,y} \Lambda(t,Y_t)}{\partial_{y,y} \Lambda(t,Y_t)^2} + \partial_{t,y} \Lambda(t,Y_t) \right)\,dt \\ &= \frac{\sqrt{2}\beta}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Z_t^\gamma)}\,dW_t + 2 \gamma \beta^2 F_\mu(t) Y_t\,dt. \end{align}\] Here we simplified the \(dt\) term using the same computations as converting from the dual to primal PDE. To simplify the \(dW_t\) time, we used the following fact: Recall from the proof of that \(\partial_y \tilde{\Lambda}_{\gamma} = \partial_y \Lambda \circ (\mathop{\mathrm{id}}+ \gamma \partial_y \Lambda)^{-1}\), and hence \[\partial_{y,y} \tilde{\Lambda}_{\gamma} = \frac{\partial_{y,y} \Lambda}{1 + \gamma \partial_{y,y} \Lambda} \circ (\mathop{\mathrm{id}}+ \gamma \partial_y \Lambda)^{-1},\] and hence \[\frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}} \circ (\mathop{\mathrm{id}}+ \gamma \partial_y \Lambda) = \frac{1}{\partial_{y,y} \Lambda} + \gamma.\] Thus in particular, \[\label{eq:32relating32two32drivers} \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Z_t^\gamma)} = \frac{1}{\partial_{y,y} \Lambda(t,Y_t)} + \gamma.\tag{21}\] Finally, we remark that some further justification is required for the Itô computation above since \(\partial_{y,y} \Lambda\) blows up at \(\pm 1\). One can proceed rigorously by expressing \(Z_t^\gamma\) in terms of the dual process \(X_t\) as \[Z_t^\gamma = \partial_x \tilde{\Phi}_{\gamma}(t,X_t) = \partial_x \Phi(t,X_t) + \gamma X_t.\] Computing \(dZ_t^\gamma\) by the Itô formula is justified by [39] and results in the same expression as above, as in the proof of .

With \(dZ_t^\gamma\) in hand, we now compute \[\begin{align} d[(Z_t^\gamma - Y_t^\gamma)^2] &= 2(Z_t^\gamma - Y_t^\gamma)(dZ_t^\gamma - dY_t^\gamma) + (dZ_t^\gamma - dY_t^\gamma)^2 \\ &= 2(Z_t^\gamma - Y_t^\gamma) \left( \frac{\sqrt{2} \beta}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Z_t^\gamma)} - \frac{\sqrt{2} \beta}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Y_t^\gamma)} \right)\,dW_t + 2(Z_t^\gamma - Y_t^\gamma) \cdot 2 \gamma \beta^2 F_\mu(t) Y_t\,dt \\ & \quad + 2 \beta^2 \left( \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Z_t^\gamma)} - \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Y_t^\gamma)} \right)^2\,dt. \end{align}\] Taking the expectation and using the fact that \(1/ \partial_{y,y} \tilde{\Lambda}_{\gamma}\) is \(2\)-Lipschitz by (3), we get that \[\begin{align} \frac{d}{dt} \mathop{{}\mathbb{E}}[(Z_t^\gamma - Y_t^\gamma)^2] &= 4 \beta^2 \mathop{{}\mathbb{E}}[(Z_t^\gamma - Y_t^\gamma) \cdot \gamma F_\mu(t) Y_t] + 2 \beta^2 \mathop{{}\mathbb{E}}\left( \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Z_t^\gamma)} - \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Y_t^\gamma)} \right)^2 \\ &\leqslant 2 \beta^2 \mathop{{}\mathbb{E}}[(Z_t^\gamma - Y_t^\gamma)^2] + 2 \beta^2 \gamma^2 \mathop{{}\mathbb{E}}[Y_t^2] + 8 \beta^2 \mathop{{}\mathbb{E}}[(Z_t^\gamma - Y_t^\gamma)^2] \\ &\leqslant 10 \beta^2 \mathop{{}\mathbb{E}}[(Z_t^\gamma - Y_t^\gamma)^2] + 2 \beta^2 \gamma^2. \end{align}\] since \(\mathop{{}\mathbb{E}}[Y_t^2] \leqslant 1\). Therefore, since \(Z_0^\gamma = 0 = Y_0^\gamma\), Grönwall’s inequality implies that \[\label{eq:32Y32minus32Z32estimate} \mathop{{}\mathbb{E}}[(Z_t^\gamma - Y_t^\gamma)^2] \leqslant\int_0^t 2 \beta^2 \gamma^2 e^{10\beta^2 (t-s)} \,ds = \frac{1}{5} \gamma^2 (e^{10\beta^2 t} - 1).\tag{22}\] This proves the second inequality in 19 . For the last claim 20 , note that \(\partial_y\Lambda(t,Y_t) = X_t\) and \(|\partial_x\Phi(x,t)| \leqslant 1\) by . Moreover, \[\begin{align} d(X_t^2) &= 2 X_t \,dX_t + (dX_t)^2 \\ &= 2X_t \left( \sqrt{2} \beta\,dW_t + 2 \beta^2 F_\mu(t) \partial_x \Phi(t,X_t)\,dt \right) + 2 \beta^2 \,dt \\ &\leqslant 2\sqrt{2}\beta X_t \,dW_t + 2 \beta^2 (2 + X_t^2) \,dt, \end{align}\] where we have used the inequality \(2 X_t \leqslant 1 + X_t^2\). Hence, \[\frac{d}{dt} \mathop{{}\mathbb{E}}[X_t^2] \leqslant 2 \beta^2(2 + \mathop{{}\mathbb{E}}[X_t^2]) \leqslant 10 \beta^2 (2/5 + \mathop{{}\mathbb{E}}[X_t^2])\] Therefore, since \(X_0 = 0\), we have \[\mathop{{}\mathbb{E}}[X_t^2] \leqslant\frac{2}{5} (e^{10 \beta^2 t} - 1).\] Hence, by this inequality and , \[\begin{align} \left\lVert{Y_t^\gamma - Y_t}\right\rVert_{L^2} &\leqslant\left\lVert{Y_t^\gamma - Y_t - \gamma X_t}\right\rVert_{L^2} + \gamma \left\lVert{X_t}\right\rVert_{L^2} \\ &\leqslant\frac{\gamma}{\sqrt{5}} (e^{10 \beta^2 t} - 1)^{1/2} + \frac{\sqrt{2}}{\sqrt{5}} \gamma (e^{10 \beta^2 t} - 1)^{1/2} \\ &\leqslant\sqrt{2} \gamma (e^{10 \beta^2 t} - 1)^{1/2}, \end{align}\] which proves the asserted estimate 20 . ◻

Corollary 3. Suppose holds, and consider the functions \(\Phi\), \(\tilde{\Lambda}_{\gamma}\) associated to the optimizing measure \(\mu\) in the Parisi formula. Let \(Y_t\) and \(Y_t^\gamma\) be as in the previous lemma. Then we have for \(t \in [0,q_\beta^*]\) that \[\label{eq:32first32L232norm32estimate} \frac{1}{\sqrt{2} \beta} - e^{5 \beta^2 t} \gamma \leqslant\left\lVert{\frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Y_t^\gamma)}}\right\rVert_{L^2} \leqslant\frac{1}{\sqrt{2} \beta} + (1 + e^{5 \beta^2 t})\gamma,\tag{23}\] and hence \[\label{eq:32second32L232norm32estimate} t^{1/2} \left(1 - \sqrt{2} \beta e^{5 \beta^2 t} \gamma \right) \leqslant\left\lVert{Y_t^\gamma}\right\rVert_{L^2} \leqslant t^{1/2} \left(1 + \sqrt{2} \beta (1 + e^{5 \beta^2 t})\gamma \right).\tag{24}\]

Proof. Consider \(Z_t^\gamma\) as in the previous proof. From 21 , we observe that \[\label{eq:32compare32L232norm32for32Zt32and32Yt} \left\lVert{ \frac{1}{\partial_{y,y} \Lambda(t,Y_t)} }\right\rVert_{L^2} \leqslant\left\lVert{ \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Z_t^\gamma)} }\right\rVert_{L^2} \leqslant\left\lVert{ \frac{1}{\partial_{y,y} \Lambda(t,Y_t)} }\right\rVert_{L^2} + \gamma.\tag{25}\] Using the fact that \(1 / \partial_{y,y} \tilde{\Lambda}_{\gamma}\) is \(2\)-Lipschitz ( (3)), together with 22 , \[\left\lVert{\frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Z_t^\gamma)} - \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Y_t^\gamma)}}\right\rVert_{L^2} \leqslant 2 \left\lVert{Z_t^\gamma - Y_t^\gamma}\right\rVert_{L^2} \leqslant 2 \cdot 5^{-1/2} \gamma (e^{10 \beta^2 t} - 1)^{1/2} \leqslant e^{5 \beta^2 t} \gamma.\] Hence, \[\label{eq:32compare32L232norm32for32Zt32and32Yt322} \left\lVert{\frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Z_t^\gamma)}}\right\rVert_{L^2} - e^{5 \beta^2 t} \gamma \leqslant\left\lVert{\frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Y_t^\gamma)}}\right\rVert_{L^2} \leqslant\left\lVert{\frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Z_t^\gamma)}}\right\rVert_{L^2} + e^{5 \beta^2 t} \gamma.\tag{26}\] By 25 and 26 , we have \[\left\lVert{\frac{1}{\partial_{y,y} \Lambda(t,Y_t)} }\right\rVert_{L^2} - e^{5 \beta^2 t}\gamma \leqslant\left\lVert{\frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Y_t^\gamma)}}\right\rVert_{L^2} \leqslant\left\lVert{\frac{1}{\partial_{y,y} \Lambda(t,Y_t)} }\right\rVert_{L^2} + ( 1 + e^{5 \beta^2 t}) \gamma\] We finally substitute in the fact that under fRSB \(\left\lVert{\frac{1}{\partial_{y,y} \Lambda(t,Y_t)}}\right\rVert_{L^2} = \frac{1}{\sqrt{2} \beta}\) () to obtain the first asserted estimate 23 .

To prove the second estimate 24 , observe that by the Itô isometry, \[\left\lVert{Y_t^\gamma}\right\rVert_{L^2}^2 = \int_0^t 2 \beta^2 \left\lVert{\frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(s,Y_s^\gamma)}}\right\rVert_{L^2}^2\,ds.\] Using 23 , the latter can be upper-bounded by \[\int_0^t 2\beta^2\left(\frac{1}{\sqrt{2}\beta} + \left(1 + e^{5\beta^2 t}\gamma\right)\right)^2\,ds \leqslant t\left(1 + \sqrt{2} \beta (1 + e^{5 \beta^2 t})\gamma\right)^2,\] and the lower bound is proved similarly. ◻

3 Extended Objective Function & Quadratic Optimization↩︎

In this section we will introduce the potential function that is derived directly from the generalized TAP free energy [2]. The generalized TAP free energy gives the extension of the original objective function into \([-1,1]^n\). While technically the modified objective is of interest only in the convex hull, as the increments are Gaussian, it will be the case that some small fraction of the coordinates of the final iterate will escape the cube. Consequently, the entropy term is regularized to deal with these outliers, ensuring that the modified objective is over \(\mathbb{R}^n\) while continuing to be a uniformly near-faithful representation on \([-1,1]^n\) (see ).

3.1 Extended objective via the generalized TAP representation↩︎

We define the potential function for the PHA algorithm which consists of a sum of “entropy-like” functions for every coordinate, corrected by an average radial term which depends on time. For conceptual and interpretative reasons that make the analogy of the use of the generalized TAP free energy in  similar to the structure of Subag’s algorithm [1], it would be preferable to have the potential function be purely dependent on geometry (and not time). While we believe it is perfectly plausible to work with the radial term evaluated at the normalized distance \(\left(\frac{1}{n}\left|\sigma\right|^2_2\right)\), we substitute in an “external clock” \(t \in [0,1]\) since it simplifies certain technical details in the analysis.

Definition 3 (Potential function & objective function). Let the potential \(V_{\beta}: [0,1] \times \mathbb{R}^n \to \mathbb{R}\cup \{\pm \infty\}\) be given by \[V_{\beta}(t,\sigma) := \left(\sum_{i \in [n]} \tilde{\Lambda}_{\gamma}(t, \sigma_i)\right) + \beta^2n\int_{t}^1F_{\mu_{\beta}}(s)\,s\,ds\,.\] Note that \(\tilde{\Lambda}_{\gamma}\) carries an implicit dependence on \(\beta\) and \(\mu_{\beta}\), and that \(\mu_{\beta}\) is from .

Given \(H(\sigma) = \sigma^{\mathsf{T}}A\sigma\), let the objective function \(\mathrm{obj}: [0,1] \times \mathbb{R}^n \to \mathbb{R}\cup \{\pm \infty\}\) be \[\mathop{\mathrm{obj}}(t,\sigma) := \beta\, H(\sigma) - V_{\beta}(t,\sigma) \,.\]

We will evaluate \(t\) as given by the step-size \(\eta\) scaled by the iterate number \(j \in [K]\). Below, the spatial derivatives of the potential function are introduced, which are important during the Taylor expansion analysis calculation (see ) to track the change in the modified objective function.

Corollary 4 (Spatial derivatives of time-dependent \(V\)).
Let \(V_\beta(t,\sigma_t)\) be the time-dependent potential function defined in . Then, the following expressions hold for the partial derivatives (with respect to space) of \(V_{\beta}(t,\sigma_t)\), where \(e_i\) is the \(i\)th standard basis vector of \(\mathbb{R}^n\): \[\begin{align} \nabla V_{\beta}(t,\sigma_t) &= \sum_{i \in [n]}\partial_y \tilde{\Lambda}_{\gamma}(t,(\sigma_t)_i)e_i\, , \\ \nabla^2 V_{\beta}(t,\sigma_t) &= \sum_{i \in [n]}\partial_{y,y}\tilde{\Lambda}_{\gamma}(t,(\sigma_t)_i)e_ie_i^\mathsf{T}\, , \\ \nabla^3 V_{\beta}(t,\sigma_t) &= \sum_{i \in [n]}\partial_{y,y,y}\tilde{\Lambda}_{\gamma}(t,(\sigma_t)_i) e_i^{\otimes 3}\,. \end{align}\]

3.2 Potential Hessian ascent: quadratic optimization with entropic corrections↩︎

The algorithm generates a sequence of iterates \(\sigma_k \in \mathbb{R}^n\) starting at \(\sigma_0 := (0, \dots, 0)\).

Let \(\delta\) be a precision parameter controlling how concentrated the random step will be in the top part of the spectrum of the Hessian and let \(\beta := 10/\varepsilon\). We will take \(K := \lceil\frac{q^*_\beta}{\eta}\rceil\) iterations in total, where \(\eta\) is chosen to depend on the approximation parameter \(\varepsilon\) (see ).

At each iterate \(\sigma_k\), we locally optimize a quadratic approximation to the objective function \(\mathop{\mathrm{obj}}(t,\sigma)\) defined in . Rather than setting the next iterate deterministically, we sample from a distribution of near-optimizers. Specifically, we sample a Gaussian vector \(u_k \sim \mathcal{N}(0, Q_k^2)\) with the following covariance2: \[Q_k^2 := 2\beta\tilde{b}n^{\delta}\Pi_{\sigma_k^{\perp}}\left(\tilde{b}^2 + \left(\tilde{a} - \nabla^2\mathop{\mathrm{obj}}(k\eta, \sigma_k)\right)^2\right)^{-1}\Pi_{\sigma_k^{\perp}} \,.\] where \(\tilde{a}\) and \(\tilde{b}\) are specific near-zero quantities controlling the sharpness of the distribution, and \(\Pi_{\sigma_k^{\perp }}\) projects to the subspace orthogonal to \(\sigma_k\). To elucidate what this does, the spectral theorem in conjunction with  shows us that \(Q_k^2\) has the same eigenvectors as \(\nabla^2\mathop{\mathrm{obj}}(k\eta, \sigma_k)\), with transformed eigenvalues: eigenvalues of the Hessian that are close to \(0\) become large eigenvalues (of order \(2\beta^2\)) in \(Q_k^2\), while eigenvalues much smaller than \(0\) in the Hessian become vanishing in \(Q_k^2\).

We then set the update to decide the next iterate as, \[\sigma_{k+1} := \sigma_k + \eta^{1/2} u_k\,.\]

In the end, to obtain a point in \(\{-1,+1\}^n\), we first truncate the coordinates to be in \([-1,1]\) individually and then sample the \(i\)th coordinate to be \(+1\) with probability \((1+(\sigma_K)_i)/2\) and \(-1\) with probability \((1-(\sigma_K)_i)/2\).

Figure 2: Potential Hessian Ascent

Lemma 9 (Fast matrix square root [48]). For every \(J, m \in \mathbb{N}\), given \(M \in \mathbb{R}^{n\times n}\) satisfying \(M \succeq 0\) and \(v \in \mathbb{R}^{n}\), it is possible to compute \(u_J\) satisfying \[\left|u_J - M^{-1/2}v\right|_2 \leqslant O\left(\frac{1}{\lambda_{\min}}\exp\left(-\frac{\Omega(m)}{\log \kappa + 3}\right)\right) + O\left(\frac{m\kappa\log(\kappa)}{\sqrt{\lambda_{\min}}}\right)\left(\frac{\sqrt{\kappa}-1}{\sqrt{\kappa}+1}\right)^{J-1}\left|v\right|_2\] in time \(O(Jn^2)\), where the runtime bottleneck comes from exactly \(J\) matrix-vector product computations with the matrix \(M\), \(\lambda_{\min}\) is the smallest eigenvalue of \(M\), and \(\kappa\) is the condition number of \(M\), defined as the ratio between the largest and smallest eigenvalues of \(M\).

Corollary 5. Given \(M\in\mathbb{R}^{n\times n}\) satisfying \(M \succeq 0\), \(v \in \mathbb{R}^{n}\), and \(\varepsilon' > 0\), it is possible to compute a vector \(u\) satisfying \[\left|u - M^{-1/2}v\right|_2 \leqslant\varepsilon'\left|v\right|_2\] in time \(O(\sqrt{\kappa} n^2\operatorname{polylog}(\kappa,1/\varepsilon', 1/\left|v\right|_2,\lambda_{\min}))\), where \(\kappa\) is the condition number of \(M\) and \(\lambda_{\min}\) is the least eigenvalue of \(M\).

As a result, if \(g \sim \mathcal{N}(0,I_n)\) so that \(h = M^{-1/2}g\) is a sample from \(\mathcal{N}(0, M^{-1})\), when \(g\) and \(M\) are given explicitly, it is possible to compute \(h'\) satisfying \(|h' - h|_2 \leqslant\varepsilon'|h|_2\) in time \(O(\sqrt{\kappa} n^2\operatorname{polylog}(\kappa,1/\varepsilon', 1/\left|g\right|_2,\lambda_{\min}))\).

Proof. Invoke with \(J = \sqrt{\kappa}\log\left (\frac{m\kappa\log\kappa}{\varepsilon'\sqrt{\lambda_{\min}}}\right)\) and with \(m = (\log \kappa + 3)\log(1/(\varepsilon'\lambda_{\min}\left|v\right|_2))\). ◻

Proposition 17 (Time complexity of ).
Assume that \(A\) satisfies the high probability event stated in , which happens with probability \(1 - C_1 \exp(-C_2 n)\). Then, there is an implementation of that runs in time \(O\left(n^{2+4\delta}e^{O(1/\varepsilon^{2})}\operatorname{polylog}(n,1/\varepsilon)\right)\) where generates \(u_k\) satisfying \(|u_k - h_k|_2 \leqslant O_{\varepsilon}(n^{-100})\) with overwhelming probability, where \(h_k \sim \mathcal{N}(0,Q_k^2)\). This implementation maintains the same algorithmic guarantees with only an \(o(1)\) loss in the value of \(H(\sigma^*)/n\), where \(\sigma^*\) is the output of the algorithm.

Proof. The runtime bottleneck will be sampling \(u_k\) approximating \(h_k \sim \mathcal{N}(0, Q_k^2)\), which may be characterized as \(h_k := L_kg_k\) for \(g_k \in \mathbb{R}^n\) a vector with i.i.d. standard Gaussian entries, where \[\begin{align} L_k := \sqrt{2\beta \tilde{b}n^{\delta}}\;\Pi_{\sigma_k^{\perp}}\left(\tilde{b}^2 + (\tilde{a} - \nabla^2\operatorname{obj}(k\eta,\sigma_k))^2\right)^{-1/2} \end{align}\] so that \(L_kL_k^{\mathsf{T}} = Q_k^2\). To apply , we will note that since \(\nabla^2\mathop{\mathrm{obj}}(k\eta, \sigma_k) = \beta(A + A^{\mathsf{T}}) - \nabla^2V(k\eta, \sigma_k)\), matrix-vector multiplication by the matrix \(K_k := \tilde{b}^2 + (\tilde{a} - \nabla^2\mathop{\mathrm{obj}}(k\eta, \sigma_k))^2\) can be computed in \(O(n^2)\) time. Furthermore, the spectrum of \(K_k\) is lower-bounded by \(\tilde{b}^2\) and it is upper-bounded by \(\tilde{b}^2 + (|\tilde{a}| + \left\lVert \nabla^2\mathop{\mathrm{obj}}(k\eta, \sigma_k) \right\rVert_{_\mathsf{op}})^2 \leqslant\exp(O(\beta^2))\) since \(\left\lVert \nabla^2\mathop{\mathrm{obj}}(k\eta, \sigma_k) \right\rVert_{_\mathsf{op}} \leqslant 3\beta + \gamma^{-1} \leqslant 3\beta + e^{O(\beta^2)}\).

Therefore, by , there is a procedure that computes a vector \(u_k\in \mathbb{R}^n\) satisfying \(|u_k - h_k|_2 \leqslant\sqrt{2\beta\tilde{b} n^\delta}\,\varepsilon' |K_k^{-1/2}g_k|_2\) where \(h_k \sim \mathcal{N}(0, Q_k^2)\) and \(|K_k^{-1/2}g_k|_2\) is upper-bounded by \(O_{\varepsilon}(n)\) with overwhelming probability in time \(O(\tilde{b}^{-1}e^{O(\beta^2)}n^{2}\operatorname{polylog}(n, 1/\varepsilon'))\). By , \(\tilde{b} \geqslant\Omega(b^3/\beta^2) \geqslant\Omega(\beta n^{-3\delta})\) when \(b = \beta n^{-\delta}\). We may take \(\varepsilon'\) to be \(O(n^{-102})\) to obtain \(|u_k - h_k|_2 \leqslant O(n^{-100})\) with high probability while incurring only an additional logarithmic dependence on \(n\) in runtime.

The value of \(a\) in can be found by binary search in \(O(n\operatorname{polylog}(n,1/\varepsilon))\) time, since the step of computing the trace of \(D^{-2}\) is linear time, with all matrices involved being diagonal. Similarly, to find \(\tilde{a}\) and \(\tilde{b}\), it is only necessary to compute the traces and inverses of diagonal matrices.

The error between \(u_k\) and \(h_k\) accumulates linearly over the iterates in the analysis of \(H(\sigma)\) and its rounding \(H(\sigma^*)\) in . ◻

The next proposition deals with the rounding scheme, which given a point \(\sigma\) in the solid cube finds a corner \(\sigma^*\) of the cube with close to the same value for the function we want to maximize. The value of each coordinate \(\sigma_j^*\) is a Bernoulli distribution on \(\{\pm 1\}\) with mean \(\sigma_j\), and we control the error with high probability using Hoeffding’s inequality. While this proposition assumes \(\sigma \in [-1,1]^n\), if the vector is merely close to \([-1,1]^n\), one can truncate it first and then apply the proposition. This is the chosen approach in .

Proposition 18 (Small fluctuations to the energy under rounding).
Let \(M\) be a real symmetric matrix. Let \(\sigma \in [-1,1]^n\) be a fixed vector. Let \(\sigma^*\) be a random vector such that the coordinates \(\sigma_j^*\) are independent for \(j = 1\), …, \(n\) and \[\mathop{{}\mathbb{P}}_{\sigma^*}(\sigma_j^* = 1) = \frac{1 + \sigma_j}{2}, \qquad \mathop{{}\mathbb{P}}_{\sigma^*}(\sigma_j^* = -1) = \frac{1 - \sigma_j}{2}.\] Then there is an absolute constant \(c\) such that, for \(\alpha \in [0,1/2]\), we have \[\mathop{{}\mathbb{P}}_{\sigma^*}\left(\angles{\sigma^*, M \sigma^*} \geqslant\angles{\sigma, M \sigma} - 4 \left\lVert{M}\right\rVert n^{1-\alpha} - \left\lVert{M}\right\rVert_2n^{1-\alpha}/\sqrt{c} - \left\lVert{M}\right\rVert n^{1-2\alpha}/c - \sum_{i \in [n]}|M_{i,i}| \right) \geqslant 1 - 2\exp(-n^{1-2\alpha}).\]

If \(A\) and \(H\) as given in satisfy the high-probability conditions stated in or , and if \(A\) also satisfies \(|A_{i,i}| < 3n^{-\alpha}/2\) for all \(i\), which occurs with probability at least \(1-ne^{-n^{2(1-\alpha)}}\), then \[\mathop{{}\mathbb{P}}_{\sigma^*}\left(H(\sigma^*) \geqslant H(\sigma) - \frac{3}{2}n^{1-\alpha}\left(5 + 1/\sqrt{c} + n^{-\alpha}/c \right)\right) \geqslant 1 - 2\exp(-n^{1-2\alpha}).\]

Proof. Note that \[\mathop{{}\mathbb{E}}[\sigma_j^*] = \frac{1+\sigma_j}{2} - \frac{1 - \sigma_j}{2} = \sigma_j.\] In other words, \(\mathop{{}\mathbb{E}}[\sigma^*] = \sigma\). Let \(\tilde{\sigma} = \sigma^* - \mathop{{}\mathbb{E}}\sigma^* = \sigma^* - \sigma\). Write \[\angles{\sigma^*, M \sigma^*} = \angles{\tilde{\sigma}, M \tilde{\sigma}} + 2 \angles{\tilde{\sigma}, M \sigma} + \angles{\sigma, M \sigma}.\] Since each component of \(\tilde{\sigma}\) is independent, zero-mean, and has sub-Gaussian norm at most 1, by the Hanson-Wright inequality [49] there is a constant \(c > 0\) such that \[\mathop{{}\mathbb{P}}\left(\angles{\tilde{\sigma}, M \tilde{\sigma}} - \mathop{{}\mathbb{E}}[\angles{\tilde{\sigma}, M \tilde{\sigma}}] < -t_{\mathrm{HW}}\right) \leqslant\exp\left(-c\min\left( \frac{t_{\mathrm{HW}}^2}{n\left\lVert{M}\right\rVert_2^2},\frac{t_{\mathrm{HW}}}{\left\lVert{M}\right\rVert}\right)\right)\,.\] So take \(t_{\mathrm{HW}} = \max(\left\lVert{M}\right\rVert_2n^{1-\alpha}/\sqrt{c},\left\lVert{M}\right\rVert n^{1-2\alpha}/c)\). Then we have

\[\mathop{{}\mathbb{E}}[\angles{\tilde{\sigma}, M \tilde{\sigma}}] = \angles{M, \mathop{{}\mathbb{E}}[\tilde{\sigma}\tilde{\sigma}^{\mathsf{T}}]} = \sum_{i \in [n]} M_{i,i}(1-\sigma_i^2) \geqslant-\sum_{i \in [n]}|M_{i,i}|\,.\]

Furthermore, \[\angles{\tilde{\sigma}, M \sigma} = \sum_{j=1}^n (M\sigma)_j \tilde{\sigma}_j.\] The random variables \((M \sigma)_j \tilde{\sigma}_j\) are independent with mean zero and each one is supported in \(\{(M \sigma)_j(-1 + \sigma_j), (M \sigma)_j(1 + \sigma_j)\}\). Therefore, by Hoeffding’s inequality, for \(t > 0\), \[\mathop{{}\mathbb{P}}(\angles{\tilde{\sigma}, M \sigma} \leqslant-t) \leqslant\exp\left( -\frac{t^2}{\sum_{j=1}^n (2(M \sigma)_j)^2} \right) = \exp\left(- \frac{t^2}{4 |M\sigma|_2^2} \right) \leqslant\exp\left(- \frac{t^2}{4 n\left\lVert{M}\right\rVert^2} \right),\] since \(|\sigma|_2^2 \leqslant n\). Now take \(t = 2 \left\lVert{M}\right\rVert n^{1-\alpha}\). Thus, \(\angles{\tilde{\sigma}, M \sigma} \geqslant-t\) implies that \[\angles{\sigma^*, M \sigma^*} \geqslant\angles{\sigma, M \sigma} - 4 \left\lVert{M}\right\rVert n^{1-\alpha} - \max\left(\left\lVert{M}\right\rVert_2n^{1-\alpha}/\sqrt{c},\left\lVert{M}\right\rVert n^{1-2\alpha}/c\right) -\sum_{i \in [n]}|M_{i,i}|,\] and the probability of violating this bound simplifies to \(\exp(-n^{1 - 2 \alpha})\). ◻

4 Random Matrix Analysis of the Modified Hessian↩︎

Our goal in this section is to define an operator that will serve as an approximate spectral projection for the entropy-corrected Hessian. Recall the entropy-shifted Hessian at a point \(\sigma \in [-1,1]^n\) and time \(t \in [0,1]\) has the form \(\beta(A + A^{\mathsf{T}}) - \operatorname{diag}(\partial_{y,y} \Lambda(t,\sigma_j))\). Here \(A\) is a non-symmetric real Gaussian matrix where the entries are independent with mean zero and variance \(1/n\), and \(A + A^{\mathsf{T}}\) is \(\sqrt{2}\) times a standard GOE matrix. We set \(A_{\mathop{\mathrm{sym}}} = (A + A^{\mathsf{T}})/2\). In this section, the specific form of the entropy correction will not play much role, and so we proceed more generally to study \(2 \beta A_{\mathop{\mathrm{sym}}} - D\) for any positive diagonal matrix \(D\). We will, however, assume that \(2 \beta^2 \operatorname{tr}_n(D^{-2}) = 1\). Later in §5 we will show that under our fRSB assumption this will be approximately true for the Hessian at the points in our algorithm, and we will be able to make it exactly true by a multiplicative renormalization of the diagonal matrix.

The large-\(n\) behavior of the spectrum of \(2 \beta A_{\mathop{\mathrm{sym}}} - D\) for any deterministic matrix with operator norm \(O(1)\) can be described using the tools of free probability, which have already been used for the spectral analysis of the TAP equations in [14]; the authors in [15] also suggested to use random matrix theory to describe the diagonal entries of the resolvent of \(2 \beta A_{\mathop{\mathrm{sym}}} - D\). Voiculescu’s asymptotic freeness theory [50] tells us that the matrices \(2 \beta A_{\mathop{\mathrm{sym}}}\) and \(D\) will be asymptotically freely independent, and their spectral measure will be well-approximated by the free convolution of the semicircular distribution with the spectral distribution of \(D\). Furthermore, in many situations, with high probability, there are no eigenvalues outside the bulk spectrum (see e.g.[51]), and so we expect that computations done in the free limit accurately reflect the maximum of the spectrum. Free probabilistic tools have already been used in [52] for the study of Kraichnan equations related to the Sherrington-Kirkpatrick model.

Concretely, our goal is the following:

Construct an operator \(P\) to approximately project onto the maximum part of the spectrum of \(2 \beta A_{\mathop{\mathrm{sym}}} - D\).

Letting \(\lambda\) (which is \(\tilde{a}(D)\) below) be the maximum of the spectrum of a free semicircular minus \(D\) (the idealized spectrum for large \(n\)), and then for small \(\rho > 0\) consider the operator \(P(D)\) given by \[P(D)^2 = -\mathop{\mathrm{Im}}(\lambda + i \rho - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1} = \rho (\rho^2 + (\lambda - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^2)^{-1}.\] The function \(t \mapsto \rho / (\rho^2 + (\lambda - t)^2)\) will as \(\rho \to 0\) be highly spiked around the point \(\lambda\). Hence, applying this function to the operator \(2 \beta A_{\mathop{\mathrm{sym}}} - D\) should output approximate eigenvectors for the top part of the spectrum. However, one challenge is that we expect the spectral density for \(2 \beta A_{\mathop{\mathrm{sym}}} - D\) to vanish at the boundary. This means that \(\operatorname{tr}(P(D)^2)\) will approach zero as \(\rho \to 0\). It is thus essential to arrange the weight \(P(D)^2\) assigns away from \(\lambda\) to go to zero faster than the trace of \(P(D)^2\) as \(\rho \to 0\). This requires some careful estimates. Furthermore, \(\rho\) will depend on \(n\); more precisely, we will set \(\rho \approx n^{-3\delta}\) for a fixed small positive \(\delta\).

In order to obtain convergence of our algorithm to the primal Auffinger-Chen SDE, we will additionally need to show that the diagonal entries of \(P(D)^2\) behave like a scalar multiple of \(D^{-2}\). We will study the projection of \((\lambda + i \rho - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}\) onto the diagonal using the idea that \(A_{\mathop{\mathrm{sym}}}\) is approximately freely independent of the diagonal subalgebra as well as a magic formula for conditional expectations of the resolvent of the sum of two freely independent operators from free probability (see below).

Another technical detail is that we want all our estimates to hold with high probability in \(A\), simultaneously for all choices of diagonal matrix \(D\) that we input; this is because the same sample of the Gaussian matrix \(A\) is being used through the whole algorithm, while the point \(\sigma \in [-1,1]\) is being updated at each step. Thus, we want uniform asymptotic freeness results that hold for all diagonal matrices simultaneously. This can be achieved using union bounds and concentration of measure for the Gaussian matrix \(A\), which is a standard technique in random matrix theory and probability more generally. Since the space of diagonal matrices is \(n\)-dimensional, the number of diagonal matrices in a sufficiently dense subset will only be exponential in \(n\), but the concentration bounds will be exponential in \(-n^2\) and hence overwhelm the number of test points that we union-bound. The same idea of “uniform asymptotic freeness from the diagonal” was applied by the first author and Srivatsav Kunnawalkam Elayavalli in another context in [53].

The following theorem is a summary of the results of this section which will be used in the rest of the paper. Background on free independence will be given in §4.1. Recall also the notation pertaining to matrix algebras in §1.4.

Theorem 19. Let \(\beta \geqslant 1\) and \(\delta \in (0,1/17]\) and \(n \in \mathbb{N}\). Let \(b = \beta n^{-\delta}\). Write \(g_{-D}(z) = \operatorname{tr}_n((z+D)^{-1})\), let \[\tilde{z} := ib + 2 \beta^2 g_{-D}(ib),\] and let \(\tilde{a}\) and \(\tilde{b}\) be the real and imaginary parts of \(\tilde{z}\). For diagonal matrices \(D \geqslant 0\) with \(2 \beta^2 \operatorname{tr}(D^{-2}) = 1\), let \(P = P(D)\) be the matrix \[P = \sqrt{\tilde{b}(\tilde{b}^2 + (\tilde{a} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^2)^{-1}}.\] Then for some absolute constants \(C_1\), \(C_2\), \(\dots > 0\), we have the following properties.

  1. Maximum of idealized spectrum: \(2 \beta^2 \operatorname{tr}_n(D^{-1})\) is the maximum of the spectrum of \(\sqrt{2} \beta S - D\) where \(S\) is a semicircular operator free from \(D\). Moreover, \[|2 \beta^2 \operatorname{tr}_n(D^{-1}) - \tilde{a}| \leqslant 2 \beta^4 n^{-2\delta} \operatorname{tr}_n(D^{-3}).\]

  2. Approximate eigenvector condition: \[\operatorname{tr}_n(P^2 (\tilde{a} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^2 ) \leqslant 2 \beta^5 n^{-3\delta} \operatorname{tr}_n(D^{-4}) \leqslant 2 \beta^3 n^{-3\delta} \left\lVert{D^{-1}}\right\rVert^2.\]

  3. Approximation for diagonal: With probability at least \[1 - C_1 \exp(-C_2 n)\] in the Gaussian matrix \(A\), we have that \(\left\lVert{2A_{\mathop{\mathrm{sym}}}}\right\rVert \leqslant 3\) and for all nonnegative diagonal matrices with \(2 \beta^2 \operatorname{tr}(D^{-2}) = 1\), \[\left\lVert{E_{\mathcal{D}_n}[P^2] - b(b^2 + D^2)^{-1}}\right\rVert_2 \leqslant C_3 \beta n^{-2\delta}\]

  4. Lower bound for trace: We have \[\operatorname{tr}_n[b(b^2 + D^2)^{-1}] \geqslant\frac{1}{2} \beta^{-1} n^{-\delta} - \beta \left\lVert{D^{-1}}\right\rVert^2 n^{-3\delta}\] so in particular \[\operatorname{tr}_n[P^2] \geqslant C_4 \beta^{-1} n^{-\delta} - C_3 \beta n^{-2\delta} - \frac{1}{2} \beta \left\lVert{D^{-1}}\right\rVert^2 n^{-3\delta}.\]

  5. Upper bound for operator norm: \[\left\lVert{P^2}\right\rVert \leqslant C_5 \beta^{-1} n^{3\delta}.\]

The theorem is based on replacing the Gaussian matrix \(\sqrt{2} A_{\mathop{\mathrm{sym}}}\) with a semicircular operator \(S\) freely independent of \(D\), an object that describes the idealized large-\(n\) limit of a GOE matrix. Point (1) of is based on spectral analysis for this idealized situation, carried out in §4.2, while point (2) follows quickly from the choice of \(P\).

The crux of the proof is point (3). We rely on the fact that \(P^2\) is chosen as the operator imaginary part of the resolvent \((\tilde{a} + i\tilde{b} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}\) (see ), and resolvents relate in a natural way with non-commutative conditional expectations (see ). The proof is broken into several parts:

  • In §4.2, we analyze the projection onto the diagonal of the resolvent of \(\sqrt{2} \beta S - D\).

  • In §4.3, we compare the diagonal of the resolvent of \(2 \beta A_{\mathop{\mathrm{sym}}} - D\) with its expectation (over \(A\)) using Gaussian concentration inequalities.

  • In §4.4, we compare the expectation of the resolvent to the idealized semicircular version studied in §4.2.

These results are combined using the triangle inequality to obtain (3) of the theorem, and then (4) follows quickly from (3).

We remark that although we rely on \(P^2\) being the imaginary part of some resolvent at a point on the complex plane, the matrix \(P\) itself always has real entries since \(\sqrt{\tilde{b}(\tilde{b}^2 + (\tilde{a} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^2)^{-1}}\) where \(A_{\mathop{\mathrm{sym}}}\) and \(D\) have real entries. Thus, while we use complex matrices in the proof of Theorem 19, which is natural for the free probability setting, we obtain in the end a real matrix to use as the covariance matrix for the real Gaussian vectors in our algorithm.

We do not care at present about the specific values of the numerical constants in the theorem, but something like the Big \(O\) notation is insufficient for the logical and algebraic manipulations. We will therefore call the numerical constants in the statements \(C_1\), \(C_2\), …. However, each statement will have its own “scope” for numbering constants, thus for instance the \(C_1\) in one lemma is not the same as the \(C_1\) in another lemma. Since in the course of the proofs we will need to define various constants as well, we will call the constants in the proofs \(M_1\), \(M_2\), …, and the constants \(C_j\) in the statements will be defined in terms of these \(M_j\)’s. Thus, each proof will also have its own scope for numbering constants.

4.1 Free probability background↩︎

Before going into the proof of , we review some background in free probability.

We consider the setting of tracial non-commutative probability spaces, a non-commutative analog of a probability space. As a motivation, recall that for a probability measure space \((\Omega,\mathcal{F},P)\), the bounded random variables form the space \(L^\infty(\Omega,\mathcal{F},P)\), and the expectation \(E = \int (\cdot)\,dP\) defines a linear map \(L^\infty(\Omega,\mathcal{F},P) \to \mathbb{C}\). We can replace \(L^\infty(\Omega,\mathcal{F},P)\) with a non-commutative algebra that has similar properties. For instance, we could consider the \(n \times n\) complex matrices \(M_n(\mathbb{C})\) as a space of “random variables” and the normalized trace \(\operatorname{tr}_n\) as the “expectation.” The distribution of a self-adjoint random variable \(X\) would then be interpreted as its spectral measure \((1/n) \sum_{j=1}^n \delta_{\lambda_j}\). More generally, \(M_n(\mathbb{C})\) can be replaced by an infinite-dimensional object with similar properties, namely a von Neumann algebra. Since the subtleties of the theory of von Neumann algebras are not needed in this paper, we present the definitions here and refer the reader to the introductory texts [54][56] and the reference books [57][59] for further information.

Von Neumann algebras:↩︎

A von Neumann algebra \(\mathcal{M} \subseteq B(H)\) is a subset of bounded operators on a Hilbert space \(H\) with the following properties:

  1. \(\mathcal{M}\) is a unital \(*\)-subalgebra, i.e., it is closed under addition, multiplication, scaling, and adjoints, and it contains the identity.

  2. \(\mathcal{M}\) is closed in the weak operator topology, i.e., if \(T_i\) is a net in \(\mathcal{M}\) and \(\langle\xi, T_i \eta \rangle\to \langle\xi, T\eta \rangle\) for all \(\xi, \eta \in H\), then \(T \in \mathcal{M}\).

The analog of the expectation (or integration with respect to a probability measure) will be tracial state on \(\mathcal{M}\). A linear functional \(\tau: \mathcal{M} \to \mathbb{C}\) is said to be

  1. normal if it is continuous with respect to the weak operator topology,

  2. unital if \(\tau(1) = 1\),

  3. tracial if \(\tau(XY) = \tau(YX)\) for \(X, Y \in \mathcal{M}\),

  4. positive if \(\tau(X^*X) \geqslant 0\) for \(X \in \mathcal{M}\),

  5. faithful if \(\tau(X^*X) = 0\) implies that \(X = 0\).

In particular, \(\tau\) is called a state if it is positive and unital. Note that the normalized trace \(\operatorname{tr}_n\) is an example of a faithful normal tracial state on \(M_n(\mathbb{C})\).

Non-commutative probability spaces:↩︎

A tracial non-commutative probability space is a pair \((\mathcal{M},\tau)\) where \(\mathcal{M}\) is a von Neumann algebra and \(\tau\) is a faithful normal tracial state. This should be viewed as a non-commutative analog of \(L^\infty(\Omega,\mathcal{F},P)\) with its expectation and as an infinite-dimensional analog of \(M_n(\mathbb{C})\) with its trace. Just as in those examples, self-adjoint elements in \(\mathcal{M}\) have a well-defined distribution. If \(X\) in \(\mathcal{M}\) is self-adjoint, then there is a unique compactly supported measure \(\mu\) such that \[\tau[f(X)] = \int_{\mathbb{R}} f\,d\mu,\] for polynomials \(f\), and we call \(\mu\) the (spectral) distribution of \(X\).

Non-commutative \(L^p\) spaces:↩︎

Given a tracial non-commutative probability space \((\mathcal{M},\tau)\) and \(p \in [1,\infty)\), we define the \(L^p\) norm \(\left\lVert{X}\right\rVert_p = \tau(|X|^p)^{1/p}\) where \(|X| = (X^*X)^{1/2}\). Also, set \(\left\lVert{X}\right\rVert = \left\lVert{X}\right\rVert_\infty\) to be the operator norm. For instance, in the case of a matrix algebra with \(\operatorname{tr}_n\), \(p = 1\) gives the normalized trace norm and \(p = 2\) gives the normalized Hilbert-Schmidt norm. These norms satisfy the non-commutative Hölder’s inequality: If \(1/p = 1/p_1 + \dots + 1/p_k\), then \[\left\lVert{x_1 \dots x_k}\right\rVert_p \leqslant\left\lVert{x_1}\right\rVert_{p_1} \dots \left\lVert{x_k}\right\rVert_{p_k}.\] We will frequently use this in the cases where \(p_j\)’s are \(1\), \(2\), or \(\infty\). We also have \(\left\lVert{x}\right\rVert_\infty = \lim_{p \to \infty} \left\lVert{x}\right\rVert_p\).

We denote by \(L^p(\mathcal{M},\tau)\) the completion of \(\mathcal{M}\) with respect to \(\left\lVert{\cdot}\right\rVert_p\). In particular, \(L^2(\mathcal{M},\tau)\) is a Hilbert space with inner product \(\langle x,y \rangle_\tau = \tau(x^*y)\). Moreover, left multiplication by \(x\) defines a representation of \(\mathcal{M}\) on \(L^2(\mathcal{M},\tau)\), i.e.a \(*\)-homomorphism \(\mathcal{M} \to B(L^2(\mathcal{M},\tau))\). For further background on non-commutative \(L^p\) spaces, see [60][62].

Conditional expectation:↩︎

Let \((\mathcal{M},\tau)\) be a tracial non-commutative probability space and let \(\mathcal{A} \subseteq \mathcal{M}\) be a von Neumann subalgebra of \(\mathcal{M}\). Then we can view \(L^2(\mathcal{A},\tau|_{\mathcal{A}})\) as a subspace of \(L^2(\mathcal{M},\tau)\). The orthogonal projection \(E_{\mathcal{A}}\) onto \(L^2(\mathcal{A},\tau)\) restricts to a mapping \(\mathcal{M} \to \mathcal{A}\), and in fact it is a contraction in all of the \(L^p\) norms. We call \(E_{\mathcal{A}}\) the canonical conditional expectation onto \(\mathcal{A}\). One example we will use in the sequel is \(\mathcal{M} = M_n(\mathbb{C})\) and \(\mathcal{A}\) the diagonal subalgebra \(\mathcal{D}_n\). Then \(E_{\mathcal{D}_n}(X)\) is simply the diagonal matrix obtained by zeroing out the off-diagonal entries of \(X\).

Free independence:↩︎

Tracial non-commutative probability space \((\mathcal{M},\tau)\) admit a non-commutative version of independence, known as free independence, studied in [63], [64]. For introductory textbooks on the topic, see [65], [51], [66]. Let \((\mathcal{M},\tau)\) be a tracial non-commutative probability space. Let \(\mathcal{A}_1\), …, \(\mathcal{A}_d\) be \(*\)-subalgebras. We say that \(\mathcal{A}_1\), …, \(\mathcal{A}_d\) are freely independent if whenever \(i_1\), …, \(i_k \in [d]\) with \(i_1 \neq i_2 \neq i_3 \neq \dots \neq i_k\) and \(X_j \in \mathcal{A}_{i_j}\), we have \[\tau \left[ (X_1 - \tau(X_1)) \dots (X_k - \tau(X_k)) \right] = 0.\] Similarly, elements, tuples, or sets in \(\mathcal{M}\) are said to be freely independent if the \(*\)-subalgebras that they generate are freely independent. It is always possible to construct freely independent copies of any given families of non-commutative random variables; given any tracial non-commutative probability spaces \((\mathcal{A},\tau_{\mathcal{A}})\) and \((\mathcal{B},\tau_{\mathcal{B}})\), one forms the free product \((\mathcal{A} * \mathcal{B}, \tau_{\mathcal{A}} * \tau_{\mathcal{B}})\); see e.g.[65], [51].

Asymptotic free independence:↩︎

We next recall Voiculescu’s results on asymptotic freeness [50], [67], which shows that GOE and GUE matrices are asymptotically free from deterministic matrices in the large-\(n\) limit; the same applies to random matrices whose distributions are invariant under unitary or orthogonal conjugations, and to many matrices with independent entries. The limiting spectral distribution of each GOE or GUE matrix is given by Wigner’s semicircle law. Let’s state an instance of this theorem relevant to our case; for proof see e.g.[51].

Theorem 20 (Asymptotic freeness for GOE and deterministic matrix). Let \(Z_n\) be a GOE random matrix, and let \(D_n\) be a deterministic matrix with \(\left\lVert{D_n}\right\rVert \leqslant M\) for some constant \(M\). Assume that the empirical spectral distribution \(\mu_{D_n}\) converges to some \(\mu\) as \(n \to \infty\). Consider a tracial non-commutative probability space \(\mathcal{M}\) generated by two freely independent self-adjoint elements \(S\) and \(D\), with spectral distributions \((1/2\pi) \mathbb{1}_{[-2,2]}(t) \sqrt{4 - t^2}\,dt\) and \(\mu\) respectively. Then we have almost surely that for non-commutative polynomials \(f\) in two variables, \[\lim_{n \to \infty} \operatorname{tr}_n[f(Z_n,D_n)] = \tau[f(S,D)].\]

Sum of freely independent variables:↩︎

An extensive theory has been developed around the sum of two variables \(X\) and \(Y\) that are freely independent, which in this paper we will apply with \(X\) being a semicircular variable \(\sqrt{2} \beta S\) and \(Y\) being \(-D\) for a positive diagonal matrix. If \(X\) and \(Y\) have distributions \(\mu\) and \(\nu\) respectively, then the distribution of \(X+Y\) (which is uniquely determined by \(\mu\) and \(\nu\)) is called the free convolution of \(\mu\) and \(\nu\) and is denoted \(\mu \boxplus \nu\). Free convolution is computed using several complex-analytic functions related to the measure \(\mu\).

Definition 4 (Cauchy-Stieltjes Transform). The Cauchy-Stieltjes transform of a probability measure \(\mu\) is the function \[g_\mu(z) = \int_{\mathbb{R}} \frac{1}{z - x}\,d\mu(x),\] defined for \(z \in \mathbb{C} \setminus \mathrm{supp}(\mu)\). Similarly, for a self-adjoint \(X\) in \(\mathcal{M}\), we write \[g_X(z) = \tau[(z - X)^{-1}] \text{ for } z \in \mathbb{C} \setminus \mathop{\mathrm{Spec}}(X),\] which is easily seen to be the Cauchy-Stieltjes transform of the spectral distribution of \(X\). Here, \(\mathop{\mathrm{Spec}}(X)\) is the set of eigenvalues of \(X\).

The following properties of the Cauchy-Stieltjes transform are well known and can be found in most textbooks that prove the spectral theorem for self-adjoint operators on Hilbert space, for instance [68].

Proposition 21 (Properties of Cauchy-Stieltjes transform). Let \(\mu \in \mathcal{P}(\mathbb{R})\). Let \(\mathrm{supp}(\mu)\) denote the closed support of \(\mu\). For \(S \subseteq \mathbb{C}\) and \(z \in \mathbb{C}\), recall that \(d(z,S) := \inf_{w \in S} |z - w|\).

  1. \(g_\mu(z)\) maps the upper half-plane into the lower half-plane and vice versa.

  2. \(\displaystyle |g_\mu(z)| \leqslant\frac{1}{d(z,\mathrm{supp}(\mu))} \leqslant\frac{1}{|\mathop{\mathrm{Im}}z|}\).

  3. A point \(a \in \mathbb{R}\) is not* in the support of \(\mu\) if and only if there is an analytic function \(f\) defined in a neighborhood \(O\) of \(a\) that agrees with \(g_\mu\) on \(O \setminus \mathbb{R}\).*

  4. If \(\mu\) is compactly supported, then \[\lim_{z \to \infty} z g_\mu(z) = 1.\] In particular, \(g_\mu^{-1}\) is defined in a neighborhood of \(0\) and \(g_\mu^{-1}(z) - 1/z\) is analytic near \(0\).

Definition 5 (\(R\)-transform). Let \(\mu \in \mathcal{P}(\mathbb{R})\). Then \(r_\mu(z) = g_\mu^{-1}(z) - 1/z\) where defined. In particular, if \(\mu\) is compactly supported, then \(r_\mu\) is defined in a neighborhood of \(0\). We also write \(r_X(z)\) for the \(R\)-transform of the spectral measure of \(X\).

Proposition 22 (Additivity of \(R\)-transform under free convolution). Let \(X\) and \(Y\) be freely independent self-adjoint elements in a tracial non-commutative probability space. Then \(r_{X+Y}(z) = r_X(z) + r_Y(z)\) in a neighborhood of \(0\). See e.g.[65], [51].

Analytic subordination:↩︎

The complex-analytic framework also allows us to understand the conditional expectation \(E_{\mathcal{A}}[(z - X - Y)^{-1}]\) when \(X \in \mathcal{A}\) and \(\mathcal{A}\) is freely independent of \(Y\) and \(z\) is in the upper half-plane \(\mathbb{H} := \{z: \mathop{\mathrm{Im}}(z) > 0\}\). In fact, \(E_{\mathcal{A}}[(z - X - Y)^{-1}]\) is given by \((f(z) - X)^{-1}\) where \(f\) is an analytic function \(\mathbb{H} \to \mathbb{H}\). This will be the key to arrange the desired behavior of the diagonal entries in (3). The following theorem is due mostly to Biane [69]; for further history and generalizations, see [69][74], Voiculescu2004free-analysis?.

Theorem 23 (See [69]). Let \(\mathcal{A}\) and \(\mathcal{B}\) be freely independent subalgebras of \(\mathcal{M}\). Let \(X \in \mathcal{A}\) and \(Y \in \mathcal{B}\) be freely independent self-adjoint operators in a tracial von Neumann algebra \((\mathcal{M},\tau)\).

  1. There exists a unique analytic function \(f: \mathbb{H} \to \mathbb{H}\) such that \(f(z) = z + O(1)\) for sufficiently large \(z\) and \(g_{X+Y}(z) = g_X(f(z))\).

  2. Let \(E_{\mathcal{A}}\) denote the canonical conditional expectation onto \(\mathcal{A}\). Then for \(z \in \mathbb{H}\), \[E_{\mathcal{A}}[(z - X - Y)^{-1}] = (f(z) - X)^{-1}.\]

Consequences of the resolvent identity:↩︎

In light of the central role played by resolvents in , we close the section with a few estimates for resolvents that will be used many times in our analysis. In the following lemmas, recall that the plain norm \(\left\lVert{x}\right\rVert\) of an element of a von Neumann algebra refers to its operator norm in \(B(H)\), which by definition is the same as \(\left\lVert{x}\right\rVert_\infty\) in the sense of the non-commutative \(L^p\) spaces.

Lemma 10. Let \(z \in \mathbb{C}\setminus \mathbb{R}\). Let \((\mathcal{M},\tau)\) be a tracial von Neumann algebra.

  1. For \(A \in \mathcal{M}_{\operatorname{sa}}\) where \(\mathcal{M}_{\operatorname{sa}}\) is the subalgebra of self-adjoint operators, we have \(\left\lVert{(z - A)^{-1}}\right\rVert_\infty \leqslant 1/|\mathop{\mathrm{Im}}z|\).

  2. Let \(A, B \in \mathcal{M}_{\operatorname{sa}}\) and \(p \in [1,\infty]\). Then \[\left\lVert{(z - A)^{-1} - (z - B)^{-1}}\right\rVert_p \leqslant\frac{1}{|\mathop{\mathrm{Im}}z|^2} \left\lVert{A - B}\right\rVert_p.\]

  3. Let \(A(t)\) be a self-adjoint element of \(\mathcal{M}\) depending in a \(C^1\) manner on a parameter \(t\) with respect \(\left\lVert{\cdot}\right\rVert_p\). Then \[\frac{d}{dt} [(z - A(t))^{-1}] = (z - A(t))^{-1} \,A'(t)\, (z - A(t))^{-1}.\]

Proof. (1) For vectors \(\xi \in H\), we have \[|\angles{\xi, (z - A) \xi}| \geqslant|\mathop{\mathrm{Im}}\angles {\xi, (z - A) \xi}| = |\mathop{\mathrm{Im}}z|\, |\xi|^2.\] Thus, \(|(z - A) \xi| \geqslant|\mathop{\mathrm{Im}}z| \,|\xi|\), which implies that \(z - A\) is injective and has closed range. Meanwhile, the same inequality holds for the adjoint \(\overline{z} I - A\), and hence \(z I - A\) is surjective, hence invertible. Then \(|(z - A) \xi| \geqslant|\mathop{\mathrm{Im}}z| \,|\xi|\) implies that \(\left\lVert{(z - A)^{-1}}\right\rVert \leqslant 1 / |\mathop{\mathrm{Im}}z|\).

(2) By the resolvent identity and the non-commutative Hölder inequality, \[\begin{align} \left\lVert{(z - A)^{-1} - (z - B)^{-1}}\right\rVert_p &= \left\lVert{(z - A)^{-1}(A - B)(z - B)^{-1}}\right\rVert_p \\ &\leqslant\left\lVert{(z - A)^{-1}}\right\rVert \left\lVert{A - B}\right\rVert_p \left\lVert{(z - B)^{-1}}\right\rVert \\ &\leqslant\frac{1}{|\mathop{\mathrm{Im}}z|^2} \left\lVert{A - B}\right\rVert_p. \end{align}\]

(3) By the resolvent identity, \[\begin{align} &(z - A(t+\epsilon))^{-1} - (z - A(t))^{-1} = \\ &(z - A(t+\epsilon))^{-1}(z - A(t))(z - A(t))^{-1} - (z - A(t+\epsilon))^{-1} (z - A(t+\epsilon))(z - A(t))^{-1} \\ &= (z - A(t+\epsilon))^{-1}(A(t+\epsilon) - A(t))(z - A(t))^{-1}. \end{align}\] Dividing by \(\epsilon\) and taking the limit as \(\epsilon \to 0\) completes the proof. ◻

We also recall the following definition: Given an operator \(T\) on a complex Hilbert space, the real and imaginary parts are defined as \(\mathop{\mathrm{Re}}(T) = (T+T^*)/2\) and \(\mathop{\mathrm{Im}}(T) = (T - T^*)/2i\). The following facts about operator real and imaginary parts are standard and easy to check.

Lemma 11. Let \(T\) be an operator in a tracial von Neumann algebra \((\mathcal{M},\tau)\).

  1. \(\mathop{\mathrm{Re}}(T)\) and \(\mathop{\mathrm{Im}}(T)\) are self-adjoint and \(T = \mathop{\mathrm{Re}}(T) + i \mathop{\mathrm{Im}}(T)\).

  2. \(\tau(\mathop{\mathrm{Re}}(T)) = \mathop{\mathrm{Re}}(\tau(T))\) and \(\tau(\mathop{\mathrm{Im}}(T)) = \mathop{\mathrm{Im}}(\tau(T))\).

  3. For \(p \in [1,\infty]\), we have \(\left\lVert{\mathop{\mathrm{Re}}(T)}\right\rVert_p \leqslant\left\lVert{T}\right\rVert_p\) and \(\left\lVert{\mathop{\mathrm{Im}}(T)}\right\rVert_p \leqslant\left\lVert{T}\right\rVert_p\).

  4. If \(T\) is invertible, then \(\mathop{\mathrm{Re}}(T^{-1}) = (T^*)^{-1} \mathop{\mathrm{Re}}(T) T^{-1}\) and \(\mathop{\mathrm{Im}}(T^{-1}) = -(T^*)^{-1} \mathop{\mathrm{Im}}(T) T^{-1}\).

4.2 Analysis of the free semicircular operator↩︎

Our argument is based on comparing \(2 \beta A_{\mathop{\mathrm{sym}}} - D\) with \(\sqrt{2} \beta S - D\) where \(S\) is a semicircular operator freely independent of \(D\), and in this section, we first perform the analysis with the semicircular operator, and we thus set up the following notation.

Notation 24. Let \(\mathcal{B} = L^\infty[-2,2]\), and let \(\tau_{\mathcal{B}}(f) = \frac{1}{2\pi} \int_{-2}^2 f(x)\sqrt{4 - x^2}\,dx\), and let \(S \in \mathcal{B}\) be the identity function, so that \(S\) is a standard semicircular element. Let \((\mathcal{M}_n,\tau_n)\) be the free product of \((\mathbb{M}_n(\mathbb{C}),\operatorname{tr}_n)\) with \((\mathcal{B},\tau_{\mathcal{B}})\). We view \(\mathbb{M}_n(\mathbb{C})\) as a subalgebra of \(\mathcal{M}_n\), so that \(\mathcal{M}_n\) is generated by \(\mathbb{M}_n(\mathbb{C})\) and a freely independent semicircular element \(S\).

Recall that the Cauchy transform of \(\sqrt{2} \beta S - D\) is given by \[g_{\sqrt{2} \beta S - D}(z) = \tau_n[(z - \sqrt{2} \beta S + D)^{-1}]\] and \[g_{-D}(z) = \operatorname{tr}_n[(z + D)^{-1}].\] Our first step is to describe the subordination function \(f\) for the free sum \(\sqrt{2} \beta S + (-D)\) more explicitly using the special form of the \(R\)-transform of a semicircular variable.

Proposition 25. Consider the setup of Notation 24, let \(D \in M_n(\mathbb{C})\). Let \(f\) be the subordination function such that \(g_{\sqrt{2} \beta S - D}(z) = g_{-D}(f(z))\) as in . Then \(f(z)\) is injective and \[f^{-1}(z) = z + 2 \beta^2 g_{-D}(z) \text{ for } z \in \operatorname{dom}(f^{-1}).\] Moreover, for \(z \in \mathbb{H}\), we have \(z \in \operatorname{dom}(f^{-1})\) if and only if \(z + 2 \beta^2 g_{-D}(z)\) has positive imaginary part if and only if \(2 \beta^2 \operatorname{tr}[((\mathop{\mathrm{Im}}z)^2 + (\mathop{\mathrm{Re}}z + D)^2)^{-1}] < 1\).

Proof. Recall that the \(R\)-transform of a semicircular element of variance \(2 \beta^2\) is \(r_{\sqrt{2} \beta S}(z) = 2 \beta^2 z\) (see e.g.[51]). By definition of the \(R\)-transform and by free independence, we have for \(z\) in a neighborhood of \(0\) that \[\begin{align} z &= g_{\sqrt{2} \beta S - D}(1/z + r_{\sqrt{2} \beta S - D}(z)) \\ &= g_{\sqrt{2} \beta S - D}(1/z + r_{-D}(z) + r_{\sqrt{2} \beta S}(z)) \\ &= g_{\sqrt{2} \beta S - D}( g_{-D}^{-1}(z) + 2 \beta^2 z). \end{align}\] Now we substitute \(w = g_{-D}^{-1}(z)\) (so \(w\) will range in a neighborhood of \(\infty\)) and obtain \[g_{-D}(w) = g_{\sqrt{2} \beta S - D}(w + 2 \beta^2 g_{-D}(w)).\] Then substituting \(w = f(z)\) for \(z\) large, we obtain \[g_{\sqrt{2} \beta S - D}(z) = g_{-D}(f(z)) = g_{\sqrt{2} \beta S - D}(f(z) + 2 \beta^2 g_{-D}(f(z))).\] Since \(g_{\sqrt{2} \beta S - D}\) is injective on a neighborhood of \(\infty\), this implies \[z = f(z) + 2 \beta^2 g_{-D}(f(z))\] holds in a neighborhood of infinity. But note that both sides are analytic on the upper half-plane and hence the identity extends to the entire upper half-plane by the identity theorem. Thus, \(f\) is injective and its inverse is given on its domain by \(z + 2 \beta^2 g_{-D}(z)\).

Next, let us prove the claim about the domain of \(f^{-1}\). First, if \(z \in \mathop{\mathrm{dom}}(f^{-1})\), then \(z + 2 \beta^2 g_{-D}(z)\) being equal to \(f^{-1}(z)\) must have positive imaginary part. Second, suppose that \(z + 2 \beta^2 g_{-D}(z)\) has positive imaginary part. Note that \[\begin{align} \mathop{\mathrm{Im}}[z + 2 \beta^2 g_{-D}(z)] &= \mathop{\mathrm{Im}}(z) + 2 \beta^2 \mathop{\mathrm{Im}}\operatorname{tr}[(z + D)^{-1}] \\ &= \mathop{\mathrm{Im}}(z) + 2 \beta^2 \operatorname{tr}( \mathop{\mathrm{Im}}((i \mathop{\mathrm{Im}}z + \mathop{\mathrm{Re}}z + D)^{-1})) \\ &= \mathop{\mathrm{Im}}(z) - 2 \beta^2 (\mathop{\mathrm{Im}}z) \operatorname{tr}( ((\mathop{\mathrm{Im}}z)^2 + (\mathop{\mathrm{Re}}z + D)^2)^{-1} ) \\ &= \mathop{\mathrm{Im}}(z) \left( 1 - 2 \beta^2 \operatorname{tr}( ((\mathop{\mathrm{Im}}z)^2 + (\mathop{\mathrm{Re}}z + D)^2)^{-1} )\right). \end{align}\] Hence, \(z + 2 \beta^2 g_{-D}(z)\) has positive imaginary part if and only if \(2 \beta^2 \operatorname{tr}( ((\mathop{\mathrm{Im}}z)^2 + (\mathop{\mathrm{Re}}z + D)^2)^{-1} ) < 1\). Furthermore, \(2 \beta^2 \operatorname{tr}( ((\mathop{\mathrm{Im}}z)^2 + (\mathop{\mathrm{Re}}z + D)^2)^{-1} )\) is a decreasing function of \(\mathop{\mathrm{Im}}z\), when \(\mathop{\mathrm{Re}}z\) is fixed. Hence, \(2 \beta^2 \operatorname{tr}( ((\mathop{\mathrm{Im}}z)^2 + (\mathop{\mathrm{Re}}z + D)^2)^{-1} ) < 1\) holds for \(z\), it also holds for \(z + iy\) for all \(y > 0\). Now for sufficiently large \(y\), we have that \(z + iy \in \mathop{\mathrm{dom}}(f^{-1})\) and so \[f(z + iy + 2 \beta^2 g_{-D}(z + iy)) = z + iy.\] Now \(z + iy + 2 \beta^2 g_{-D}(z + iy)\) has positive imaginary part and hence is in the domain of \(f\) for all \(y \geqslant 0\). Thus, by analytic continuation the above identity extends to all \(y \geqslant 0\), and in particular, \(f(z + 2 \beta^2 g_{-D}(z)) = z\), so that \(z \in \mathop{\mathrm{dom}}(f^{-1})\). ◻

The next lemma locates the maximum of the spectrum of \(\sqrt{2} \beta S - D\). Note that when we assume \(2 \beta^2 \operatorname{tr}(D^{-2}) = 1\), this means that the point \(a\) in the lemma below will be zero. The approach will be to use property 3 of to reduce statements about the spectrum of \(\sqrt{2} \beta S - D\) to complex-analytic properties of its Cauchy-Stieltjes transform.

Lemma 12.  

  1. For each real diagonal matrix \(D \in M_n(\mathbb{C})\), there exists a unique real number \(a > -\min \mathop{\mathrm{Spec}}(D)\) such that \(2 \beta^2 \operatorname{tr}_n((a + D)^{-2}) = 1\).

  2. \(\max \mathop{\mathrm{Spec}}(\sqrt{2} \beta S - D) = a + 2\beta^2\operatorname{tr}_n((a+D)^{-1})\).

Proof. (1) Note that \(2 \beta^2 \operatorname{tr}_n((a + D)^{-2})\) is a strictly decreasing function of \(a\) on \([-\min \mathop{\mathrm{Spec}}(D),\infty)\). As \(a \to \infty\), it converges to zero and as \(a \to -\min \mathop{\mathrm{Spec}}(D)\), it converges to \(+\infty\). Hence, by the intermediate value theorem and strict monotonicity, there is a unique point where it equals \(1\).

(2) Let \(h(z) = z + 2\beta^2 g_{-D}(z) = z + 2\beta^2 \operatorname{tr}((z+D)^{-1})\), which is the inverse function for \(f\) on the appropriate domain. Our first goal is to show that \(h\) is invertible on \((a,\infty) + i\mathbb{R}\) and in particular this region is contained in the image of \(f\). Suppose that \(x \in (a,\infty)\) and \(y \in \mathbb{R}\). We claim first that \(\mathop{\mathrm{Im}}h(x+iy)\) has the same sign as \(y\). Note that \[\mathop{\mathrm{Im}}h(x+iy) = y - 2\beta^2\operatorname{tr}(y(y^2 + (x+D)^2)^{-1}) = y\left(1 - 2\beta^2\operatorname{tr}((y^2 + (x + D)^2)^{-1})\right).\] Now \(2\beta^2\operatorname{tr}((y^2 + (x+D)^2)^{-1}) \leqslant 2\beta^2\operatorname{tr}((x+D)^{-2}) < 2\beta^2\operatorname{tr}((a+D)^{-2}) = 1\), and so \(\mathop{\mathrm{Im}}h(x+iy)\) is \(y\) times a strictly positive number, and so has the same sign as \(y\). By , we see that if \(y > 0\), then \(x+iy \in \mathop{\mathrm{dom}}(f)\). Similarly, if \(y < 0\), then \(x+iy\) is in the domain of the mirror version of \(f\) in the lower half-plane. In particular, \(h\) is injective on these two regions. We note the identity that \(h'(z) = 1 - 2\beta^2\operatorname{tr}((z+D)^{-2})\), and hence since \(\operatorname{tr}\mathop{\mathrm{Re}}X \leqslant\operatorname{tr}|X|\) for all \(X\), \[\begin{align} \mathop{\mathrm{Re}}h'(x+iy) &\geqslant 1 - 2 \beta^2 \operatorname{tr}(|(x + iy +D)^{-2}|) \\ &= 1 - 2 \beta^2 \operatorname{tr}((y^2 + (x+D)^2)^{-1}) \\ &> 1 - 2 \beta^2 \operatorname{tr}((a+D)^{-2}) \\ &= 0, \end{align}\] In particular, we see that \(h\) is strictly increasing on \((a,\infty)\) and thus injective there as well. Overall \(h\) is injective on \((a,\infty) + i\mathbb{R}\).

Also, since \(a > \min \mathop{\mathrm{Spec}}(D)\), we know \(g_{-D}\) is analytic on \((a,\infty) + i\mathbb{R}\). Therefore, \(g_{-D} \circ h^{-1}\) is analytic on \(h((a,\infty) + i\mathbb{R})\). Furthermore, we know that \(h^{-1}\) agrees with \(f\) on \(((a,\infty) + i\mathbb{R}) \cap \mathbb{H}\), so that \(g_{-D} \circ h^{-1}\) agrees with \(g_{-D} \circ f = g_{\sqrt{2}\beta S - D}\) on this region; the symmetrical statement holds for the lower half-plane. Hence, \(g_{\sqrt{2} \beta S - D}\) is analytic on \(h((a,\infty) + i\mathbb{R})\), and it follows by that \(h((a,\infty) + i\mathbb{R})\) is in the complement of the spectrum of \(\sqrt{2} \beta S - D\). In particular, \(h((a,\infty)) = (a + 2\beta^2 g_{-D}(a),\infty)\) is in the complement of the spectrum, and so \[\max \mathop{\mathrm{Spec}}(\sqrt{2} \beta S - D) \leqslant a + 2\beta^2 g_{-D}(a) = a + 2\beta^2\operatorname{tr}((a+D)^{-1}).\]

To see the reverse inequality, consider the behavior of \(h\) near the point \(a\). By construction, \(h'(a) = 0\) and \(h''(a) = 4 \beta^2 \operatorname{tr}((a+D)^{-3}) > 0\). Thus, \[h(z) - h(a) = \frac{1}{2}h''(a)(z - a)^2 + O((z-a)^3),\] hence, \(h(z) - h(a) = (z-a)^2 p(z)\) for some nonzero analytic function \(p(z)\), and in particular, \(h(z) - h(a) = q(z)^2\) for some analytic function \(q\) with \(q(a) = 0\) and \(q'(a) > 0\). Now \(q^{-1}\) is defined on some neighborhood \((a-\delta,a+\delta) + i(-\delta,\delta)\). We have \(h^{-1}(z) = q^{-1}(\sqrt{z - h(a)})\) on \((a,a+\delta) + i(-\delta,\delta)\), where \(\sqrt{z - h(a)}\) is defined by taking the slit in the negative real direction. This implies also that \[g_{\sqrt{2} \beta S - D}(z) = g_{-D}\left(q^{-1}\left(\sqrt{z - h(a)}\right) \right).\] Since \(g_{-D}\) is analytic on a neighborhood of \(a\), this identity extends to a neighborhood of \(h(a)\) say \((h(a)-\delta',h(a)+\delta') + i (-\delta',\delta')\). Our choice of square root \(\sqrt{z - h(a)}\) cuts the plane along \((-\infty,h(a)]\), mapping the “upper side” onto the positive imaginary axis and the “lower side” onto the negative imaginary axis. Since \((q^{-1})'(0) > 0\) and \[g_{-D}'(a) = \frac{d}{dz} \Bigr|_{z=a} \operatorname{tr}_n((z + D)^{-1}) = -\operatorname{tr}_n((a+D)^{-2}) = -\frac{1}{2 \beta^2} < 0,\] we get that for \(x \in (h(a)-\delta',h(a))\), we have \[\lim_{y \to 0^+} g_{\sqrt{2} \beta S - D}(x+iy) = \lim_{y \to 0^+} g_{-D}\left(q^{-1}\left(\sqrt{x+iy - h(a)}\right)\right) > 0,\] and the behavior as \(y \to 0^-\) is described by the complex conjugate. Thus, \(g_{\sqrt{2} \beta S - D}\) cannot extend to be analytic in a neighborhood of \(h(a)\) since the values from the upper side and the lower side disagree on \((h(a)-\delta,h(a))\). Therefore by , \(h(a) \in \mathop{\mathrm{Spec}}(\sqrt{2} \beta S - D)\). ◻

As mentioned before, in Theorem 19, we assume \(2 \beta^2 \operatorname{tr}_n(D^{-2}) = 1\), take \(a=0\), and fix \(b = \beta n^{-\delta}\) for some small positive \(\delta\). Letting \(\tilde{z} = \tilde{a} + i \tilde{b} = ib + 2 \beta^2 g_{-D}(ib)\), we have \(\tilde{a}\) is the maximum of the spectrum of \(\sqrt{2} \beta S - D\). To obtain an operator that concentrates on the upper part of the spectrum of \(\sqrt{2} \beta S - D\), we will use the imaginary part of the resolvent \([\tilde{a} + i \tilde{b} - (\sqrt{2} \beta S - D)]^{-1}\). Our goal is then to study the conditional expectation of this operator onto the diagonal subalgebra \(\mathcal{D}_n \subseteq \mathbb{M}_n(\mathbb{C}) \subseteq \mathcal{M}_n\), which is described by \((f(\tilde{a} + i\tilde{b}) + D)^{-1} = (ib + D)^{-1}\), thanks to .

We remark that our estimates need to take the renormalization into account. Indeed, the normalized \(\operatorname{tr}_n[\mathop{\mathrm{Im}}[\tilde{z} - (\sqrt{2} \beta S - D)]]^{-1}\) will vanish as \(b \to 0\), and so this operator will later be renormalized in §5. The reason it vanishes is that density of the spectral measure of \(\sqrt{2} \beta S - D\) will vanish like a square root function at the edge of the spectrum (as one can see from the proof of (2) and Stieltjes inversion formula). The operator \(-\mathop{\mathrm{Im}}[\tilde{z} - (\sqrt{2} \beta S - D)]^{-1}\) is obtained by applying the “spiked” function \(x \mapsto (\tilde{a} + i\tilde{b} - x)^{-1}\) to the operator \((\sqrt{2} \beta S - D)\); the spiked function is concentrated near the edge of the spectrum at \(\tilde{a}\) where the density is small, and so its integral with respect to \(\mu_{\sqrt{2} \beta S - D}\) will vanish as \(b \to 0\). Thus, the error terms for our approximation of the diagonal will need to be small relative to \(\operatorname{tr}_n[\mathop{\mathrm{Im}}[\tilde{z} - (\sqrt{2} \beta S - D)]]^{-1}\). The trace turns out to be on the order of \(b\) while the error estimates for the diagonal terms are controlled in terms of \(\tilde{b}\). Thus, a key point is that we prove in the next lemma that \(\tilde{b}\) is on the order \(b^3\) because we chose exactly the point \(\tilde{a}\) at the edge of the spectrum (this is related to the fact that \(h'(a) = 0\) in the proof of (2) above). It is quite necessary for our argument that \(\tilde{a}\) is at the top of the spectrum (and hence \(a = 0\)); indeed, if \(a\) and hence \(\tilde{a}\) were larger, then the “spike” of the function \(x \mapsto -\mathop{\mathrm{Im}}(\tilde{a} + i\tilde{b} - x)^{-1}\) fall outside the spectrum of \(\sqrt{2} \beta S - D\). On the other hand, if we took \(a < 0\), then \(a+ib\) would end up outside the domain of \(f^{-1}\) when \(b\) is sufficiently small, which would prevent us from using the subordination function to estimate the resolvent.

Lemma 13. Fix \(b \in (0,\beta)\). Let \(D \geqslant 0\) be a diagonal matrix with \(2\beta^2\operatorname{tr}(D^{-2}) = 1\). Let \(\tilde{z} = ib + 2 \beta^2 g_{-D}(ib)\), and let \(\tilde{a}\) and \(\tilde{b}\) be the real and imaginary parts of \(\tilde{z}\) (which also depend on \(D\)). Then

  1. \(b^3 / (3 \beta^2) \leqslant\tilde{b} \leqslant 2 \beta^2 b^3 \operatorname{tr}_n(D^{-4}) \leqslant b^3 \left\lVert{D^{-1}}\right\rVert^2\).

  2. \(-2 \beta^2 \mathop{\mathrm{Im}}g_{-D}(ib) \geqslant b - 2 \beta^2 b^3 \operatorname{tr}_n(D^{-4}) \geqslant b - b^3 \left\lVert{D^{-1}}\right\rVert^2\).

Proof. (1) Note that using and resolvent identities, \[\begin{align} \tilde{b} &= b + 2 \beta^2 \mathop{\mathrm{Im}}\operatorname{tr}_n[(ib + D)^{-1}] \\ &= b - 2 \beta^2 b \operatorname{tr}_n[(b^2 + D^2)^{-1}] \\ &= b(1 - 2 \beta^2 \operatorname{tr}_n[D^{-2}] - 2 \beta^2 \operatorname{tr}_n[(b^2 + D^2)^{-1} - D^{-2}]) \\ &= b(0 + 2 \beta^2 b^2 \operatorname{tr}_n[(b^2 + D^2)^{-1}D^{-2}]) \\ &= 2 \beta^2 b^3 \operatorname{tr}_n[(b^2 + D^2)^{-1}D^{-2}]). \end{align}\] Then we note that \[\operatorname{tr}_n[(b^2 + D^2)^{-1}D^{-2}]) \leqslant\operatorname{tr}_n[D^{-4}],\] and \[2 \beta^2 \operatorname{tr}_n(D^{-4}) \leqslant 2 \beta^2 \operatorname{tr}_n(D^{-2}) \left\lVert{D^{-1}}\right\rVert^2 = \left\lVert{D^{-1}}\right\rVert^2\] which implies the upper bounds from (1).

On the other hand, by Jensen’s inequality applied to the convex function \(t \mapsto (b^2t + 1)^{-1} t^2\) on \([0,\infty)\), \[\begin{align} \operatorname{tr}_n[(b^2 + D^2)^{-1}D^{-2}]) &= \operatorname{tr}_n[(b^2 D^{-2}+1)^{-1} D^{-4}] \\ &\geqslant(b^2 \operatorname{tr}_n(D^{-2}) + 1)^{-1} \operatorname{tr}_n(D^{-2})^2 \\ &= (b^2/2\beta^2 + 1)^{-1} \frac{1}{4 \beta^4} \\ &= (b^2 + 2 \beta^2)^{-1} (2 \beta^2)^{-1}. \end{align}\] Therefore, \[\tilde{b} \geqslant b^3(b^2 + 2 \beta^2)^{-1} \geqslant b^3(\beta^2 + 2 \beta^2)^{-1} = \frac{b^3}{3 \beta^2},\] which gives the lower bound from (1).

(2) Observe that \[-2 \beta^2 \mathop{\mathrm{Im}}g_{-D}(ib) = b - \tilde{b} \geqslant b - 2 \beta^2 b^3 \operatorname{tr}_n(D^{-4}),\] and again note \(2 \beta^2 \operatorname{tr}_n(D^{-4}) \leqslant\left\lVert{D^{-1}}\right\rVert^2\). ◻

4.3 Concentration estimates for the resolvent↩︎

The goal of this subsection is to estimate the difference between the diagonal of the resolvent \((\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}\) and its expectation with respect to the randomness of \(A\) for all fixed \(D\). We argue using concentration inequalities, which give estimates on the probability of some random variable deviating from its expectation by a certain amount. In particular, we use Herbst’s concentration inequality for Lipschitz functions in this section and in the next section we use the Poincaré inequality. These inequalities are closely connected to the log-Sobolev inequality, Talagrand entropy-cost inequality, and other aspects of information geometry. For general background on concentration inequalities, see [75][77].

Lemma 14 (Herbst concentration estimate for Gaussian matrix). Let \(A\) be a normalized real Ginibre matrix (see Definition 1). Let \(f: M_n(\mathbb{R}) \to \mathbb{R}\) be a Lipschitz function with respect to \(\left\lVert{\cdot}\right\rVert_2\). Then for all \(\delta \geqslant 0\), \[\mathop{{}\mathbb{P}}( |f(A) - \mathop{{}\mathbb{E}}[f(A)]| \geqslant\delta) \leqslant 2 e^{-n^2 \delta^2 / (2 \left\lVert{f}\right\rVert_{\operatorname{Lip}}^2)}\]

For proof of this lemma, we refer to [51]. A related inequality gives a bound on the variance of a Lipschitz function of a Gaussian variable.

Lemma 15 (Poincaré inequality for Gaussian). Let \(f: M_n(\mathbb{R}) \to \mathbb{R}\) be Lipschitz with respect to \(\left\lVert{\cdot}\right\rVert_2\). Let \(A\) be a normalized real Gaussian random matrix. Then \[\mathop{{}\boldsymbol{\mathrm{Var}}}(f(A)) \leqslant\frac{1}{n^2} \mathop{{}\mathbb{E}}\left\lVert{\nabla f(A)}\right\rVert_2^2 \leqslant\frac{1}{n^2} \left\lVert{f}\right\rVert_{\operatorname{Lip}}^2.\]

Corollary 6. Let \(W\) be a real or complex inner product space, and let \(f, g: M_n(\mathbb{R}) \to W\) be Lipschitz. \[\mathop{{}\mathbb{E}}|\angles{ f(A) - \mathop{{}\mathbb{E}}[f(A)], g(A) - \mathop{{}\mathbb{E}}[g(A)] }| \leqslant\frac{\dim_{\mathbb{R}} W}{n^2} \left\lVert{f}\right\rVert_{\operatorname{Lip}} \left\lVert{g}\right\rVert_{\operatorname{Lip}}.\]

Proof. First, consider the case where \(f = g\) and the inner product space is real. Fix an orthonormal basis \(w_1\), …, \(w_d\) for \(W\). Apply the previous lemma to \(\langle f(A),w_j \rangle\) for each \(j\), and then sum up the results over the basis. In the complex case, we apply the previous lemma to the real and imaginary parts of \(\angles{f(A), w_j}\).

In the case where \(f \ne g\), the left-hand side can be estimated using the Cauchy-Schwarz inequality, and then we apply the case of \(f = g\) proved above. ◻

The other ingredient we will need is an estimate on the probability of large operator norm for the GOE matrix [51]. Here recall \(A + A^{\mathsf{T}}\) is \(\sqrt{2}\) times a GOE matrix, and so asymptotically its operator norm will converge to \(2 \sqrt{2} < 3\).

Lemma 16. Let \(A\) be a normalized real Gaussian matrix. For some universal constants \(C_1\) and \(C_2\), we have \[\mathop{{}\mathbb{P}}(\left\lVert{A + A^{\mathsf{T}}}\right\rVert \geqslant 3) \leqslant C_1 e^{-C_2 n}.\]

The expectation of the operator norm can be estimated by looking at high moments [51], and then we can estimate the probability using concentration. A short argument for this type of bound is given by Ledoux [78] in the GUE case, and this can be adapted to the GOE as well using the joint distribution of eigenvalues in [51].

Now we are ready to apply these concentration estimates to our particular choice of resolvent operator. In the following lemma we do not use the precise form of \(\tilde{z}\) from , and so we state it more generally for any \(\tilde{z}\) with positive imaginary part.

Lemma 17. Let \(A\) be a normalized real Gaussian matrix. Let \(E_{\mathcal{D}_n}: M_n(\mathbb{C}) \to \mathcal{D}_n\) be the orthogonal projection (or non-commutative conditional expectation) onto the diagonal matrices. Let \(\tilde{z}\) with \(\mathop{\mathrm{Im}}\tilde{z} > 0\), and let \(\delta > 0\). Fix a real diagonal matrix \(D\). Then \[\left\lVert{E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1} - \mathop{{}\mathbb{E}}E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}]]}\right\rVert_2 \leqslant\frac{4 \beta}{n^{8\delta} |\mathop{\mathrm{Im}}\tilde{z}|^2}\] with probability at least \[1 - 2\cdot 3^{2n} e^{-n^{2-16\delta} / 2}.\]

Proof. Here we use a classic \(\epsilon\)-net argument. Let \(\Omega = \{D': \operatorname{tr}(D'(D')^*) = 1\}\) be a set of diagonal matrices and let \(\Omega_0\) be a maximal \(1/2\)-separated subset of \(\Omega\) with respect to \(\left\lVert{\cdot}\right\rVert_2\). Thus, every element of \(\Omega\) is within a distance of \(1/2\) from some element of \(\Omega_0\). The balls of radius \(1/2\) centered at points in \(\Omega_0\) are disjoint and contained in the ball of radius \(3/2\) and hence \(|\Omega_0| \leqslant 3^{2n}\) since we used complex entries. Thus, for any diagonal matrix \(B\), we have \[\left\lVert{B}\right\rVert_2 = \sup_{D' \in \Omega} \mathop{\mathrm{Re}}\angles{B,D'} \leqslant\max_{D' \in \Omega_0} |\mathop{\mathrm{Re}}\operatorname{tr}(BD')| + \frac{1}{2} \left\lVert{B}\right\rVert_2,\] so that \[\left\lVert{B}\right\rVert_2 \leqslant 2 \max_{D' \in \Omega_0} |\mathop{\mathrm{Re}}\operatorname{tr}(BD')|.\] If \(B\) is not necessarily diagonal, then \(\operatorname{tr}(BD') = \operatorname{tr}(E_{\mathcal{D}_n}[B]D')\) and hence \[\left\lVert{E_{\mathcal{D}_n}[B]}\right\rVert_2 \leqslant 2 \max_{D' \in \Omega_0} |\mathop{\mathrm{Re}}\operatorname{tr}(BD')|.\] We apply this with \(B = (\tilde{z} - 2\beta A_{\mathop{\mathrm{sym}}} + D)^{-1} - \mathop{{}\mathbb{E}}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}]\), so \[\begin{align} \bigl \lVert E_{\mathcal{D}_n}[(\tilde{z} &- 2\beta A_{\mathop{\mathrm{sym}}} + D)^{-1} - \mathop{{}\mathbb{E}}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}]] \bigr \rVert_2 \\ &\leqslant 2 \max_{D' \in \Omega_0} \bigl|\mathop{\mathrm{Re}}\operatorname{tr}_n[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}D' - \mathop{{}\mathbb{E}}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}]D'] \bigr|. \end{align}\] Moreover, for every such \(D'\), we have that \(\mathop{\mathrm{Re}}\operatorname{tr}_n[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1} D']\) is \(2 \beta / |\mathop{\mathrm{Im}}\tilde{z}|^2\)-Lipschitz as a function of \(A\) with respect to \(\left\lVert{\cdot}\right\rVert_2\). Therefore, by the concentration inequality of , \[\mathop{{}\mathbb{P}}\left( |(\mathop{\mathrm{id}}- \mathop{{}\mathbb{E}})\mathop{\mathrm{Re}}\operatorname{tr}_n[(\tilde{z} - 2\beta A_{\mathop{\mathrm{sym}}} + D)^{-1} D']| \geqslant\frac{2 \beta}{n^{8\delta} |\mathop{\mathrm{Im}}\tilde{z}|^2} \right) \leqslant 2 e^{-n^{2 - 16\delta} / 2}.\] By taking a union bound, we can arrange that \(|(\mathop{\mathrm{id}}- \mathop{{}\mathbb{E}})\operatorname{tr}_n[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1} D']| \leqslant\frac{2 \beta}{n^{8\delta} |\mathop{\mathrm{Im}}\tilde{z}|^2}\) for all \(D' \in \Omega_0\) with probability at least \(1 - 2\cdot 3^{2n} e^{-n^{2-16\delta} / 2}\). ◻

Lemma 18. Let \(\beta \geqslant\sqrt{3}\). Let \(\delta \in [0,1/17]\). Suppose \(n^\delta \geqslant\sqrt{3}\). Fix \(b = \beta n^{-\delta}\). For nonnegative diagonal matrices \(D\) with \(2 \beta^2 \operatorname{tr}(D^{-2}) = 1\), let \(\tilde{z} = \tilde{a} + i \tilde{b} = ib + 2\beta^2g_{-D}(ib)\), which depends on \(D\), \(n\), and \(\delta\). Then with probability \(1 - C_1 \exp(-C_2 n)\) in the Gaussian matrix \(A\), we have \[\left\lVert{E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1} - \mathop{{}\mathbb{E}}E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}]]}\right\rVert_2 \leqslant C_3 \beta n^{-2\delta}\] uniformly for all real diagonal \(D\) with \(D \geqslant 0\) and \(2 \beta^2 \operatorname{tr}(D^{-2}) = 1\).

Proof. We want to view \(E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}]\) as a function of \(D^{-1}\) in the sphere of radius \(1/(\sqrt{2} \beta)\) with respect to \(\left\lVert{\cdot}\right\rVert_2\), show that it is Lipschitz, and hence deduce an estimate for all \(D^{-1}\) from an estimate on a sufficiently dense subset.

For the Lipschitz estimate, fix two nonnegative diagonal matrices \(D\) and \(D'\) with \(\operatorname{tr}(D^{-2}) = \operatorname{tr}((D')^{-2}) = 1/(2\beta^2)\), and let \(\tilde{z} = ib + 2 \beta^2 \operatorname{tr}((ib + D)^{-1})\) and \(\tilde{z}' = ib + 2 \beta^2 \operatorname{tr}((ib + D')^{-1})\). Note by the resolvent identity that \[\begin{align} (ib + D)^{-1} - (ib + D')^{-1} &= (ib + D)^{-1} (D' - D) (ib + D')^{-1} \\ &= (ib + D)^{-1} D (D^{-1} - (D')^{-1}) D'(ib + D')^{-1} \\ &= (1 - ib(ib + D)^{-1}) (D^{-1} - (D')^{-1}) (1 - ib(ib + D')^{-1}). \end{align}\] Furthermore, \(\left\lVert{(ib + D)^{-1}}\right\rVert \leqslant 1/b\) so that \(\left\lVert{1 - ib(ib + D)^{-1}}\right\rVert \leqslant 2\), and similarly for \(D'\). Hence, by the non-commutative Hölder’s inequality, \[\left\lVert{(ib + D)^{-1} - (ib + D')^{-1}}\right\rVert_2 \leqslant\left\lVert{1 - ib(ib + D)^{-1}}\right\rVert \left\lVert{D^{-1} - (D')^{-1}}\right\rVert_2 \left\lVert{1 - ib(ib + D')^{-1}}\right\rVert \leqslant 4 \left\lVert{D^{-1} - (D')^{-1}}\right\rVert_2.\] Thus, \[|\tilde{z} - \tilde{z}'| \leqslant 2 \beta^2 \left\lVert{(ib + D)^{-1} - (ib + D')^{-1}}\right\rVert_2 \leqslant 8 \beta^2 \left\lVert{D^{-1} - (D')^{-1}}\right\rVert_2.\] Another application of the resolvent identity yields that \[\begin{align} (\tilde{z} - 2\beta A_{\mathop{\mathrm{sym}}} + D)^{-1} & - (\tilde{z}' - 2 \beta A_{\mathop{\mathrm{sym}}} + D')^{-1} \\ &= (\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}(\tilde{z}' - \tilde{z} + D' - D)(\tilde{z}' - 2 \beta A_{\mathop{\mathrm{sym}}} + D')^{-1} \tag{27} \\ &= (\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}(\tilde{z}' - \tilde{z})(\tilde{z}' - 2 \beta A_{\mathop{\mathrm{sym}}} + D')^{-1} \tag{28} \\ &\quad + (\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}(D' - D)(\tilde{z}' - 2 \beta A_{\mathop{\mathrm{sym}}} + D')^{-1} \tag{29} \end{align}\] Now by and (1), \[\left\lVert{(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}}\right\rVert \leqslant\frac{1}{\tilde{b}} \leqslant\frac{3 \beta^2}{b^3},\] and similarly for \(D'\). Hence, 28 can be bounded by \[\begin{align} \left\lVert{(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}(\tilde{z}' - \tilde{z})(\tilde{z}' - 2\beta A_{\mathop{\mathrm{sym}}} + D')^{-1}}\right\rVert_2 &\leqslant\frac{M_1}{\beta^4 b^6} |\tilde{z} - \tilde{z}'| \\ &\leqslant\frac{M_2 \beta^6}{b^6} \left\lVert{D^{-1} - (D')^{-1}}\right\rVert_2 \\ &= M_2 n^{6\delta} \left\lVert{D^{-1} - (D')^{-1}}\right\rVert_2, \end{align}\] where \(M_1\) and \(M_2\) are universal constants. Then to estimate 29 , we write \[\begin{gather} \label{eq:32resolvent32to32estimate322} (\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}(D' - D)(\tilde{z}' - 2 \beta A_{\mathop{\mathrm{sym}}} + D')^{-1} = (\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}D(D^{-1} - (D')^{-1})D'(\tilde{z}' - 2 \beta A_{\mathop{\mathrm{sym}}} + D')^{-1}. \end{gather}\tag{30}\] We write \[\begin{align} (\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}D &= 1 - (\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}}) \nonumber \\ &= 1 - (\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}i\tilde{b} - (\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1} (\tilde{a} - 2 \beta A_{\mathop{\mathrm{sym}}}). \label{eq:32split32of32terms32to32estimate32resolvent} \end{align}\tag{31}\] As before, \[\left\lVert{(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1} \tilde{b}}\right\rVert \leqslant 1.\] Meanwhile, \[\tilde{a} = \mathop{\mathrm{Re}}2\beta^2\operatorname{tr}((ib + D)^{-1}) = 2\beta^2\operatorname{tr}(D(b^2 + D^2)^{-1}) \leqslant 2\beta^2\operatorname{tr}(D^{-1}) \leqslant 2\beta^2\operatorname{tr}(D^{-2})^{1/2} = \sqrt{2}\beta.\] Furthermore, recall by , we have that \(\left\lVert{2A_{\mathop{\mathrm{sym}}}}\right\rVert \leqslant 3\) with probability \(1 - M_3 \exp(-M_4 n)\), and so \(\left\lVert{\tilde{a} - 2 \beta A_{\mathop{\mathrm{sym}}}}\right\rVert \leqslant(\sqrt{2} + 3) \beta\). Thus, \[\left\lVert{(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}(\tilde{a} - 2 \beta A_{\mathop{\mathrm{sym}}})}\right\rVert \leqslant\frac{3 \beta^2}{b^3} (\sqrt{2} + 3) \beta \leqslant M_5 n^{3\delta}.\] Substituting all these estimates into 31 yields \[\left\lVert{(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}D}\right\rVert \leqslant M_6 n^{3 \delta},\] and of course a similar estimate holds for \(D'(\tilde{z}' - 2 \beta A_{\mathop{\mathrm{sym}}} + D')^{-1}\). Plugging these back into 30 , we get \[\left\lVert{(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}(D' - D)(\tilde{z}' - 2 \beta A_{\mathop{\mathrm{sym}}} + D')^{-1}}\right\rVert_2 \leqslant M_6^2 n^{6\delta} \left\lVert{D^{-1} - (D')^{-1}}\right\rVert_2.\] Therefore, overall \[\left\lVert{(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1} - (\tilde{z}' - 2 \beta A_{\mathop{\mathrm{sym}}} + D')^{-1}}\right\rVert_2 \leqslant M_7 n^{6 \delta} \left\lVert{D^{-1} - (D')^{-1}}\right\rVert_2.\] Finally, since \(E_{\mathcal{D}_n}\) is contractive in the \(2\)-norm, the map \[D \mapsto E_{\mathcal{D}_n}\big[(\tilde{z}-2\beta A_{\mathop{\mathrm{sym}}}+D)^{-1}\big]\] is \(M_7 n^{6\delta}\)-Lipschitz in \(D^{-1}\) on the event \(\{\left\lVert{2A_{\mathop{\mathrm{sym}}}}\right\rVert\leqslant 3\}\). Hence, since this event does not depend on \(D\), the centered truncated map \[D^{-1}\mapsto E_{\mathcal{D}_n}\big[(\tilde{z}-2\beta A_{\mathop{\mathrm{sym}}}+D)^{-1}\big]\mathbb{1}_{\{\left\lVert{2A_{\mathop{\mathrm{sym}}}}\right\rVert\leqslant 3\}}-\mathop{{}\mathbb{E}}[E_{\mathcal{D}_n}\big[(\tilde{z}-2\beta A_{\mathop{\mathrm{sym}}}+D)^{-1}\big]\mathbb{1}_{\{\left\lVert{2A_{\mathop{\mathrm{sym}}}}\right\rVert\leqslant 3\}}]\] is \(2M_7 n^{6\delta}\)-Lipschitz. Moreover, uniformly in \(D\), \[\left\lVert{E_{\mathcal{D}_n}\big[(\tilde{z}-2\beta A_{\mathop{\mathrm{sym}}}+D)^{-1}\big]}\right\rVert_2 \leqslant\frac{1}{\tilde{b}}\leqslant\frac{3\beta^2}{b^3}=3 \beta^2 \beta^{-3} n^{3\delta},\] and therefore \[\label{eq:A-diagonal-expected-truncation-ok} \left\lVert{\mathop{{}\mathbb{E}}E_{\mathcal{D}_n}\big[(\tilde{z}-2\beta A_{\mathop{\mathrm{sym}}}+D)^{-1}\big]-\mathop{{}\mathbb{E}}[E_{\mathcal{D}_n}\big[(\tilde{z}-2\beta A_{\mathop{\mathrm{sym}}}+D)^{-1}\big]\mathbb{1}_{\{\left\lVert{2A_{\mathop{\mathrm{sym}}}}\right\rVert\leqslant 3\}}]}\right\rVert_2 \leqslant 3M_3\beta^{-1} n^{3\delta}e^{-M_4 n}.\tag{32}\]

Next, fix a maximal subset \(\Omega_0\) of \(\Omega = \{D^{-1}: 2 \beta^2 \operatorname{tr}(D^{-2}) = 1\}\) that is \(\beta n^{-8\delta}\)-separated with respect to \(\left\lVert{\cdot}\right\rVert_2\). Thus, every element of \(\Omega\) is within a distance of \(\beta/n^{8\delta}\) from \(\Omega_0\). The balls centered at the elements of \(\Omega_0\) of radius \(\beta/(2n^{8\delta})\) are disjoint and contained in the ball of radius \(3\beta /2\) centered at \(0\), and so \(|\Omega_0| \leqslant(3 n^{8 \delta})^n\).

By the union bound and , with probability at least \[1 - 2(3n^{8\delta})^n 3^{2n} e^{-n^{2-16\delta}/2},\] over the Gaussian matrix, we have for all \(D^{-1} \in \Omega_0\), \[\left\lVert{E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1} - \mathop{{}\mathbb{E}}E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}]]}\right\rVert_2 \leqslant\frac{4 \beta}{n^{8\delta} |\mathop{\mathrm{Im}}\tilde{z}|^2} \leqslant M_9 \beta^{-1} n^{-2\delta}\] where again we have applied (1) to estimate \(1 / |\mathop{\mathrm{Im}}\tilde{z}|^2 \leqslant 9 \beta^4 / (\beta n^{-\delta})^6\). By our choice of \(\Omega_0\), every \(D^{-1} \in \Omega\) is within a distance of \(\beta / n^{8\delta}\) of some \((D')^{-1} \in \Omega_0\). Therefore, when \(\left\lVert{2A_{\mathop{\mathrm{sym}}}}\right\rVert \leqslant 3\), by , \[\begin{align} &\left\lVert{E_{\mathcal{D}_n}\big[(\tilde{z}-2\beta A_{\mathop{\mathrm{sym}}}+D)^{-1}\big]-\mathop{{}\mathbb{E}}E_{\mathcal{D}_n}\big[(\tilde{z}-2\beta A_{\mathop{\mathrm{sym}}}+D)^{-1}\big]}\right\rVert_2 \\ &\qquad\qquad - \left\lVert{ E_{\mathcal{D}_n}\big[(\tilde{z}'-2\beta A_{\mathop{\mathrm{sym}}}+D')^{-1}\big]-\mathop{{}\mathbb{E}}E_{\mathcal{D}_n}\big[(\tilde{z}'-2\beta A_{\mathop{\mathrm{sym}}}+D')^{-1}\big]}\right\rVert_2\\ &\leqslant 2M_7 n^{6\delta}\left\lVert{D^{-1}-(D')^{-1}}\right\rVert_2 + 3M_3\beta^{-1} n^{3\delta}e^{-M_4 n} \\ &\leqslant 2M_7 \beta n^{-2\delta} + 3M_3\beta^{-1} n^{3\delta}e^{-M_4 n}. \end{align}\] The last term is exponentially small in \(n\), so it may be absorbed into the constants. Therefore, \[\forall D \in \Omega, \quad \left\lVert{E_{\mathcal{D}_n}[(\tilde{z} - 2\beta A_{\mathop{\mathrm{sym}}} + D)^{-1} - \mathop{{}\mathbb{E}}E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}]]}\right\rVert_2 \leqslant M_{11} \beta n^{-2\delta},\] with probability at least \[1 - M_3 \exp(-M_4 n) - 2(3n^{8\delta})^n 3^{2n} e^{-n^{2-16\delta}/2}.\] Note since \(\delta \leqslant 1/17\), \[2(3n^{8\delta})^n 3^{2n} e^{-n^{2-16\delta}/2} = 2\exp(-\tfrac{1}{2}n^{2-16\delta} + 8\delta n \log n + 3 n \log 3) \leqslant M_{12} e^{-M_{13} n}.\] so the overall probability of error is bounded by \(M_{14} \exp(-M_{15} n)\). In the statement, we take \(C_1 = M_{14}\), \(C_2 = M_{15}\), and \(C_3 = M_{11}\). ◻

4.4 Evaluation of the expected resolvent by interpolation↩︎

We have now studied the behavior of the resolvent \((\tilde{z} - \sqrt{2} \beta S + D)^{-1}\), and also estimated how far \((\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}\) is from its expectation. It remains to compare \((\tilde{z} - \sqrt{2} \beta S + D)^{-1}\) with the expectation of \((\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}\). This we will do using the following proposition. Here we consider an arbitrary point \(z\) in the upper half-plane for ease of notation, but ultimately, we will plug in \(\tilde{z} = ib + 2 \beta^2 g_{-D}(ib) = i\beta n^{-\delta} + 2 \beta^2 g_{-D}(in^{\delta})\) instead of \(z\).

Proposition 26. Fix \(z\) with positive imaginary part, let \(D\) be a real diagonal matrix, and let \(f\) be the subordination function with \(g_{\sqrt{2} \beta S - D} = g_{-D} \circ f\). Then \[\label{eq:32expectation32estimate} \left\lVert{\mathop{{}\mathbb{E}}\left[ E_{\mathcal{D}_n}[(z - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}]\right] - (f(z) + D)^{-1}}\right\rVert_2 \leqslant\frac{32 \beta^4}{n^{3/2} |\mathop{\mathrm{Im}}z|^5} + \frac{2 \beta^2}{n |\mathop{\mathrm{Im}}z|^3}.\qquad{(1)}\]

We prove this proposition using several common techniques in random matrix theory: interpolation, integration by parts, and concentration estimates. In particular, our argument is inspired by the work of Collins, Guionnet, and Parraud [79] in the setting of the GUE. It is well-known that for a polynomial \(p\), the expectation of \(\operatorname{tr}_n[f(X^{(n)})]\), where \(X^{(n)}\) is a GUE matrix, is equal to the trace of \(f(S)\), where \(S\) is standard semicircular operator, plus a correction of order \(1/n^2\), and in fact the expectation has an asymptotic expansion in powers of \(1/n^2\), known as the genus expansion (see [80] for an introduction and historical survey). Since it is unclear how the genus expansion would help with smooth functions beyond polynomials, Collins, Guionnet, and Parraud [79] and Parraud [81] developed another asymptotic expansion formula where the terms are expressed through non-commutative derivatives of the input function \(f\), which are amenable to analytic techniques. The idea of the proof is to put the GUE matrix \(X^{(n)}\) and the semicircular operator in the same space and study the interpolation \((1-t)^{1/2} X^{(n)} + t^{1/2} S\) inspired by the free Ornstein-Uhlenbeck process. After differentiating in \(t\) the expected trace of \((1-t)^{1/2} X^{(n)} + t^{1/2} S\), they use integration by parts and other tricks to get a tractable expression for the \(1/n^2\) correction. A similar interpolation method was used in [82] to obtain sharp non-asymptotic bounds for the operator norms of general Gaussian random matrices.

We remark that the technique of Gaussian interpolation has formed the backbone of several rigorous results in mathematical spin-glass theory [44], [83][85], in particular in the proofs of the existence of the limit of the free energy density [18] as well as the upper bounds in (both) the replica-symmetric and replica-symmetry breaking regimes [19]. The use of a non-commutative version of Gaussian interpolation to bound the resolvent of the “shifted” Hessian, which ultimately provides a bound on the free energy achieved by the algorithmic process, is a notable parallel in technique.

We will follow a similar strategy with the GOE rather than GUE matrix, which is slightly more complicated because the expansion will have \(O(1/n)\) terms. However, since our goal is only to get a concrete bound rather than a full asymptotic expansion, we will be content to integrate by parts once and then estimate the result using the Poincaré inequality (compare [81]). In fact, rather than directly arguing with the semicircular operator \(\sqrt{2} S\), we will approximate it by a Gaussian matrix \(B\) of size \(nk\) much larger than \(n\), which we may regard as fixed throughout this discussion (a larger GUE matrix of size \(nk\) is similarly used in [79] and [81]).

Thus, fix \(n\) and \(k\), and let \(B\) be an \(nk \times nk\) Gaussian matrix where again the entries are standard normal with variance \(1/(nk)\). Let \(B_{\mathop{\mathrm{sym}}} = \frac{1}{2}(B + B^{\mathsf{T}})\), so that \(2B_{\mathop{\mathrm{sym}}}\) is \(\sqrt{2}\) times a GOE matrix. Then we consider \(A \otimes I_k\) and \(B\) as elements of \(M_{nk}(\mathbb{R})\). For \(t \in [0,1]\) and \(z \in \mathbb{C}\setminus \mathbb{R}\), let \[G = G(z,t) = (z I_{nk} - 2 \beta [(1-t)^{1/2} A_{\mathop{\mathrm{sym}}} \otimes I_k + t^{1/2} B_{\mathop{\mathrm{sym}}}] + D \otimes I_k)^{-1};\] we will often abbreviate \(G(z,t)\) to \(G\) for the sake of readability in our equations. Then fix another diagonal matrix \(D'\) and let \[h(t) = \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[G(z,t) (D' \otimes I_k)].\] Note that \[\label{eq:32h32of32zero} h(0) = \mathop{{}\mathbb{E}}\operatorname{tr}_n[(z I_n - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1} D'],\tag{33}\] while \[\label{eq:32h32of32one} h(1) = \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[(z I_{nk} - \beta(B + B^{\mathsf{T}}) + D \otimes I_k)^{-1} (D' \otimes I_k)],\tag{34}\] which will involve the freely independent semicircular operator after we take \(k \to \infty\) at the very end of the argument.

We will estimate \(h(1) - h(0)\) by studying \(h'(t)\). Here we rely on which implies that \[\frac{d}{dt} G(z,t) = \beta t^{-1/2} G(z,t) B_{\mathop{\mathrm{sym}}} G(z,t) - \beta (1 - t)^{-1/2} G(z,t)(A_{\mathop{\mathrm{sym}}} \otimes I_k) G(z,t).\] Thus, \[\label{eq:32formula32for32h32prime} h'(t) = \beta t^{-1/2} \mathbb{E} \operatorname{tr}_{nk}[G B_{\mathop{\mathrm{sym}}} G(D' \otimes I_k)] - \beta (1-t)^{-1/2} \mathbb{E} \operatorname{tr}_{nk}[G(z,t)(A_{\mathop{\mathrm{sym}}} \otimes I_k) G (D' \otimes I_k)].\tag{35}\] We will study these terms using Gaussian integration by parts.

Lemma 19. Consider the setup above, and abbreviate \(G(z,t)\) to \(G\) for convenience. Then \[\label{eq:32IBP32132formula} \beta t^{-1/2} \mathbb{E} \operatorname{tr}_{nk}[G B_{\mathop{\mathrm{sym}}} G (D' \otimes I_k)] = 2 \beta^2 \mathop{{}\mathbb{E}}\bigg[ \operatorname{tr}_{nk}[G] \operatorname{tr}_{nk}[G(D' \otimes I_k)G] + \frac{1}{nk} \operatorname{tr}_{nk}[G^3 (D' \otimes I_k)] \bigg].\tag{36}\]

Proof. First, by cyclic symmetry of the trace, \[\operatorname{tr}_{nk}[G (B+B^{\mathsf{T}}) G (D' \otimes I_k)] = \operatorname{tr}_{nk}[B G (D' \otimes I_k) G] + \operatorname{tr}_{nk}[G (D' \otimes I_k) G B^{\mathsf{T}}]\] Now since \(G(z,t)\) is the inverse of a symmetric matrix, it is symmetric (here we mean symmetric with respect to transposition rather than Hermitian conjugation). Also, \(D'\) is symmetric. Thus, since transposition preserves the trace, \[\operatorname{tr}_{nk}[G (D' \otimes I_k) G B^{\mathsf{T}}] = \operatorname{tr}_{nk}[B G (D' \otimes I_k) G].\] Thus, we get \[\operatorname{tr}_{nk}[G B_{\mathop{\mathrm{sym}}} G (D' \otimes I_k)] = \operatorname{tr}_{nk}[B G (D' \otimes I_k) G] = \frac{1}{kn} \sum_{i,j=1}^{kn} B_{i,j} [G (D' \otimes I_k) G]_{j,i}.\] We take expectations on both sides and then use Gaussian integration by parts: \[\frac{1}{kn} \sum_{i,j=1}^{kn} \mathop{{}\mathbb{E}}\left[ B_{i,j} [G (D' \otimes I_k) G]_{j,i} \right] = \frac{1}{k^2n^2} \sum_{i,j=1}^{kn} \mathop{{}\mathbb{E}}\left[ \frac{\partial}{\partial B_{i,j}} \left[ G (D' \otimes I_k) G \right]_{j,i} \right]\] Using , we get \[\begin{align} \frac{\partial}{\partial B_{i,j}} G &= G \frac{\partial}{\partial B_{i,j}}[\beta t^{1/2}(B + B^{\mathsf{T}})] G \\ &= \beta t^{1/2} G (E_{i,j} + E_{j,i}) G \\ &= \beta t^{1/2} G E_{i,j} G + \beta t^{1/2} G E_{j,i} G, \end{align}\] where \(E_{i,j}\) are the standard matrix units in \(M_{nk}(\mathbb{C})\). Thus, \[\begin{align} \beta t^{-1/2} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[G B_{\mathop{\mathrm{sym}}} G (D' \otimes I_k)] &= \frac{\beta^2}{k^2n^2} \sum_{i,j=1}^{nk} \mathop{{}\mathbb{E}}[G E_{i,j} G (D' \otimes I_k) G]_{j,i} \tag{37} \\ &\quad + \frac{\beta^2}{k^2n^2} \sum_{i, j=1}^{nk} \mathop{{}\mathbb{E}}[G E_{j,i} G (D' \otimes I_k) G]_{j,i} \tag{38} \\ &\quad + \frac{\beta^2}{k^2n^2} \sum_{i,j=1}^{nk} \mathop{{}\mathbb{E}}[G(D' \otimes I_k)G E_{i,j} G]_{j,i} \tag{39} \\ &\quad + \frac{\beta^2}{k^2n^2} \sum_{i,j=1}^{nk} \mathop{{}\mathbb{E}}[G (D' \otimes I_k)G E_{j,i} G]_{j,i} \tag{40}, \end{align}\] Taking first the terms 38 and 40 that contain \(E_{j,i}\), note that \[\eqref{eq:32matrix32resolvent32computation32term322} = \frac{\beta^2}{k^2n^2} \sum_{i, j=1}^{nk} \mathop{{}\mathbb{E}}[G_{j,j} [G (D' \otimes I_k) G]_{i,i}] = \beta^2 \operatorname{tr}_{nk}[G] \operatorname{tr}_{nk}[G (D' \otimes I_k) G].\] Likewise, 40 evaluates to \(\operatorname{tr}_{nk}[G(D' \otimes I_k)G] \operatorname{tr}_{nk}[G]\), which also equals 38 . Therefore, 38 and 40 become the term \(2 \beta^2 \mathop{{}\mathbb{E}}\bigg[ \operatorname{tr}_{nk}[G] \operatorname{tr}_{nk}[G(D' \otimes I_k)G]\) on the right-hand side of 36 in the statement of the lemma.

Next, consider the terms 37 and 39 that contain \(E_{i,j}\). Recall that \(G^{\mathsf{T}} = G\), or \(G_{i,j} = G_{j,i}\). Hence, \[\begin{align} \eqref{eq:32matrix32resolvent32computation32term321} &= \frac{\beta^2}{k^2n^2} \sum_{i,j=1}^{nk} \mathop{{}\mathbb{E}}[G_{j,i} [G (D' \otimes I_k) G]_{j,i}] = \frac{\beta^2}{k^2n^2} \sum_{i,j=1}^{nk} \mathop{{}\mathbb{E}}[G_{i,j} [G (D' \otimes I_k) G]_{j,i}] \\ &= \frac{\beta^2}{nk} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[G[G(D' \otimes I_k)G]] = \frac{\beta^2}{nk} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[G^3(D' \otimes I_k)]. \end{align}\] Similarly, 39 also evaluates to \(\frac{\beta^2}{nk} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[G^3(D' \otimes I_k)]\). Therefore, 37 and 39 become the term \(\frac{2 \beta^2}{nk} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[G^3(D' \otimes I_k)]\) on the right-hand side of 36 in the statement of the lemma. ◻

We now carry out a parallel computation for the matrix \(A_{\mathop{\mathrm{sym}}} \otimes I_k\) rather than \(B_{\mathop{\mathrm{sym}}}\). In the following, recall that we identify \(M_{nk}(\mathbb{C})\) with \(M_n(\mathbb{C}) \otimes M_k(\mathbb{C})\). Let \(\operatorname{tr}_n \otimes \mathop{\mathrm{id}}: M_{nk}(\mathbb{C}) \to M_k(\mathbb{C})\) be the partial trace map given on simple tensors by \(X \otimes Y \mapsto \operatorname{tr}_n(X)Y\), and let \(\mathsf{T}\otimes \mathop{\mathrm{id}}: M_{nk}(\mathbb{C}) \to M_{nk}(\mathbb{C})\) be the partial transpose map given by \(X \otimes Y \mapsto X^{\mathsf{T}} \otimes Y\).

Lemma 20. Continuing with the same setup as above and again abbreviating \(G(z,t)\) to \(G\), we have \[\begin{gather} \label{eq:32IBP32232formula} \beta (1-t)^{-1/2} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[G (A_{\mathop{\mathrm{sym}}} \otimes I_k) G (D' \otimes I_k)] = 2 \beta^2 \mathop{{}\mathbb{E}}\operatorname{tr}_k[(\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G] (\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G(D' \otimes I_k) G]] \\ + \frac{2\beta^2}{n} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[(\mathsf{T}\otimes \mathop{\mathrm{id}})[G] G (D' \otimes I_k) G]. \end{gather}\tag{41}\]

Proof. The argument is similar to the previous lemma. First, we observe that \((A \otimes I_k)^{\mathsf{T}} = A^{\mathsf{T}} \otimes I_k\) and \(G^{\mathsf{T}} = G\), and apply symmetry of the trace and transposes to get \[\operatorname{tr}_{nk}[G (A^{\mathsf{T}} \otimes I_k) G (D' \otimes I_k)] = \operatorname{tr}_{nk}[(D' \otimes I_k) G (A \otimes I_k) G] = \operatorname{tr}_{nk}[(A \otimes I_k)G(D' \otimes I_k) G].\] Therefore, \[\begin{align} \operatorname{tr}_{nk}[G (A_{\mathop{\mathrm{sym}}} \otimes I_k) G (D' \otimes I_k)] &= \operatorname{tr}_{nk}[(A \otimes I_k) G (D' \otimes I_k) G] \\ &= \sum_{i,j=1}^n A_{i,j} \operatorname{tr}_{nk}[(E_{i,j} \otimes I_k) G (D' \otimes I_k) G] \\ &= \frac{1}{n} \sum_{i,j=1}^n \mathop{{}\mathbb{E}}\left[ \frac{\partial}{\partial A_{i,j}} \operatorname{tr}_{nk}[(E_{i,j} \otimes I_k) G (D' \otimes I_k) G] \right], \end{align}\] where the last line follows from Gaussian integration by parts. Note that \[\frac{\partial}{\partial A_{i,j}} G = \beta (1 - t)^{1/2} G((E_{i,j} + E_{j,i}) \otimes I_k) G,\] and so \[\frac{\partial}{\partial A_{i,j}} [(E_{i,j} \otimes I_k) G (D' \otimes I_k) G] = (E_{i,j} \otimes I_k) G ((E_{i,j} + E_{j,i}) \otimes I_k) G (D' \otimes I_k) + (E_{i,j} \otimes I_k) G (D' \otimes I_k) ((E_{i,j} + E_{j,i}) \otimes I_k) G.\] Thus, we get \[\begin{align} \beta (1-t)^{-1/2} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[G (A_{\mathop{\mathrm{sym}}} \otimes I_k) G (D' \otimes I_k)] &= \frac{\beta^2}{n} \sum_{i,j=1}^{n} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[(E_{i,j} \otimes I_k)G(E_{i,j} \otimes I_k) G (D' \otimes I_k) G] \tag{42} \\ &\quad + \frac{\beta^2}{n} \sum_{i,j=1}^{n} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk} [(E_{i,j} \otimes I_k) G (E_{j,i} \otimes I_k) G (D' \otimes I_k) G] \tag{43} \\ &\quad + \frac{\beta^2}{n} \sum_{i,j=1}^{n} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[(E_{i,j} \otimes I_k)G(D' \otimes I_k)G (E_{i,j} \otimes I_k) G] \tag{44} \\ &\quad + \frac{\beta^2}{n} \sum_{i,j=1}^{n} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[(E_{i,j} \otimes I_k)G (D' \otimes I_k)G (E_{j,i} \otimes I_k)G] \tag{45}. \end{align}\] Note that 43 and 45 are the same up to cyclic symmetry of the trace and switching the indices \(i\) and \(j\). Each of these two terms evaluates to \[\begin{align} \frac{\beta^2}{n} \sum_{i,j=1}^{n} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[(E_{i,j} \otimes I_k)G(E_{i,j} \otimes I_k) G (D' \otimes I_k) G] &= \beta^2 \operatorname{tr}_{nk}[((\operatorname{tr}_n \otimes \mathop{\mathrm{id}})(G) \otimes I_k) G (D' \otimes I_k) G] \\ &= \beta^2 \mathbb{E} \operatorname{tr}_k[(\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G] (\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G(D' \otimes I_k) G]. \end{align}\] Hence, 43 and 45 together produce the term \(2\beta^2 \mathbb{E} \operatorname{tr}_k[(\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G] (\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G(D' \otimes I_k) G]\) on the right-hand side of 41 in the lemma statement. Meanwhile, 42 and 44 are equal to each other using cyclic symmetry of the trace, and each of these terms evaluates to \[\begin{align} \frac{\beta^2}{n} \sum_{i,j=1}^{n} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[(E_{i,j} \otimes I_k)G(E_{i,j} \otimes I_k) G (D' \otimes I_k) G] &= \frac{\beta^2}{n} \sum_{i,j=1}^{n} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[(E_{i,i} \otimes I_k)(\mathsf{T}\otimes \mathop{\mathrm{id}})[G](E_{j,j} \otimes I_k) G (D' \otimes I_k) G] \\ &= \frac{\beta^2}{n} \operatorname{tr}_{nk}[(\mathsf{T}\otimes \mathop{\mathrm{id}})[G] G (D' \otimes I_k) G]. \end{align}\] Hence, 42 and 44 together produce the term \(\frac{2\beta^2}{n} \operatorname{tr}_{nk}[(\mathsf{T}\otimes \mathop{\mathrm{id}})[G] G (D' \otimes I_k) G]\) on the right-hand side of 41 in the lemma statement. ◻

To summarize, the last two lemmas together with 35 show that \(h'(t)\) is the difference of 36 and 41 , which we want to estimate in order to prove Proposition 26. We first give an estimate for the rightmost terms in 36 and 41 respectively, which are the easier terms to deal with since they already have factors of \(1/n\).

Lemma 21. With the setup above, we have \[\label{eq:32estimate32transpose32trace32term321} \left|\frac{2\beta^2}{nk} \mathbb{E} \operatorname{tr}_{nk}[G^3 (D' \otimes I_k)]| \right| \leqslant\frac{2\beta^2}{nk |\mathop{\mathrm{Im}}z|^3} \left\lVert{D'}\right\rVert_2\tag{46}\] and \[\label{eq:32estimate32transpose32trace32term322} \left| \frac{2\beta^2}{n} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[(\mathsf{T}\otimes \mathop{\mathrm{id}})[G] G (D' \otimes I_k) G] \right| \leqslant\frac{2\beta^2}{n |\mathop{\mathrm{Im}}z|^3} \left\lVert{D'}\right\rVert_2.\tag{47}\]

Proof. We first note that \(\left\lVert{G(z,t)}\right\rVert \leqslant\frac{1}{|\mathop{\mathrm{Im}}z|}\) as a direct application of (1). Hence, \[|\operatorname{tr}_{nk}[G^3 (D' \otimes I_k)]| \leqslant\left\lVert{G^3}\right\rVert_2 \left\lVert{D' \otimes I_k}\right\rVert_2 \leqslant\left\lVert{G}\right\rVert^3 \left\lVert{D'}\right\rVert_2,\] where we use that \(\left\lVert{D' \otimes I_k}\right\rVert_2^2 = \operatorname{tr}_{nk}(((D')^*D' \otimes I_k)) = \operatorname{tr}_n((D')^*D') = \left\lVert{D'}\right\rVert_2^2\). This easily implies 46 . For the other estimate, note that \[\begin{align} |\operatorname{tr}_{nk}[(\mathsf{T}\otimes \mathop{\mathrm{id}})[G] G (D' \otimes I_k) G]| &\leqslant\left\lVert{(\mathsf{T}\otimes \mathop{\mathrm{id}})[G]}\right\rVert_2 \left\lVert{G(D' \otimes I_k)G}\right\rVert_2 \\ &\leqslant\left\lVert{G}\right\rVert_2 \left\lVert{G}\right\rVert \left\lVert{D' \otimes I_k}\right\rVert_2 \left\lVert{G}\right\rVert \\ &\leqslant\left\lVert{G}\right\rVert^3 \left\lVert{D'}\right\rVert_2. \end{align}\] Here we use the fact that \((\mathsf{T}\otimes \mathop{\mathrm{id}})\) is isometric with respect to \(\left\lVert{\cdot}\right\rVert_2\) since it performs a permutation of the entries of the matrix.3 We also use the non-commutative Hölder’s inequality. ◻

In order to estimate 41 minus 36 , it remains to estimate \[\mathop{{}\mathbb{E}}\operatorname{tr}_k[(\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G] (\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G(D' \otimes I_k) G]] - \mathop{{}\mathbb{E}}[\operatorname{tr}_{nk}[G] \operatorname{tr}_{nk}[G(D' \otimes I_k) G]].\] This is the more challenging part, and the key idea is to understand this as a covariance and use the Poincaré inequality.

Lemma 22. \[\begin{gather} \left| \mathop{{}\mathbb{E}}\operatorname{tr}_k\Bigl[(\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G] (\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G(D' \otimes I_k) G] \Bigr] - \mathop{{}\mathbb{E}}\Bigl[\operatorname{tr}_{nk}[G] \operatorname{tr}_{nk}[G(D' \otimes I_k) G] \Bigr] \right| \\ \leqslant\frac{16 \beta^2 \left\lVert{D'}\right\rVert_2}{n^{3/2} |\mathop{\mathrm{Im}}z|^5} + \frac{16 \beta^2 \left\lVert{D'}\right\rVert_2}{n^{3/2} k^2 |\mathop{\mathrm{Im}}z|^5}. \end{gather}\]

Proof. Consider the mappings \(M_{nk}(\mathbb{R}) \to M_k(\mathbb{C})\) given by \(f(B) = (\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G^*]\) and \(g(B) = (\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G(D' \otimes I) G]\), and note that \[\mathop{{}\mathbb{E}}\operatorname{tr}_k\Bigl[(\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G] (\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G(D' \otimes I_k) G]\Bigr] = \mathop{{}\mathbb{E}}\angles{f(B),g(B)}_{\operatorname{tr}_k}.\] To apply the Poincaré inequality, we first want to verify these are Lipschitz. Using (2), we see that \(G\) is \(\frac{2 \beta}{|\mathop{\mathrm{Im}}z|^2}\)-Lipschitz as a function of \(B\) with respect to \(\left\lVert{\cdot}\right\rVert_2\), which implies that \[f \text{ is } \frac{2 \beta}{|\mathop{\mathrm{Im}}z|^2} \text{-Lipschitz in } B.\] For \(g\), we use Lipschitzness of \(G\) in \(B\) together with the behavior of Lipschitz functions under products; more precisely, given two inputs \(B\) and \(\tilde{B}\), let \(G\) and \(\tilde{G}\) be the corresponding values of the resolvent \(G\). Then \[\begin{align} \left\lVert{G(D' \otimes I_k) G - \tilde{G}(D' \otimes I_k)\tilde{G}}\right\rVert_2 &\leqslant\left\lVert{(G - \tilde{G})(D' \otimes I_k)G}\right\rVert_2 + \left\lVert{\tilde{G}(D' \otimes I_k)(G - \tilde{G})}\right\rVert_2 \\ &\leqslant\left\lVert{G - \tilde{G}}\right\rVert_2 \left\lVert{D'}\right\rVert \left\lVert{G}\right\rVert + \left\lVert{\tilde{G}}\right\rVert \left\lVert{D'}\right\rVert \left\lVert{G - \tilde{G}}\right\rVert_2 \\ &\leqslant 2 \frac{\beta}{|\mathop{\mathrm{Im}}z|^2} \frac{1}{|\mathop{\mathrm{Im}}z|} \left\lVert{D'}\right\rVert. \end{align}\] Note here that we have the operator norm rather than the \(2\)-norm of \(D'\). Then since the \((\operatorname{tr}_n \otimes \mathop{\mathrm{id}})\) is a contraction with respect to \(2\)-norm, we see that \[g \text{ is } \frac{4 \beta}{|\mathop{\mathrm{Im}}z|^3} \left\lVert{D'}\right\rVert \text{-Lipschitz in } B.\] Now let \(\mathop{{}\mathbb{E}}_B\) denote the expectation with respect to \(B\) while holding \(A\) fixed. Note that \(B\) is invariant in distribution with respect to conjugation by \(I_n \otimes O\) for any matrix \(O\) in the \(k \times k\) orthogonal group, and of course \(I_n \otimes O\) leaves \(A \otimes I_k\) and \(D' \otimes I_k\) literally invariant. It follows that \(\mathop{{}\mathbb{E}}_B[(\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G]]\) is invariant under conjugation by the orthogonal group and hence is equal to a multiple of the identity; thus, \[\mathop{{}\mathbb{E}}_B[(\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G]] = \mathop{{}\mathbb{E}}_B \operatorname{tr}_{nk}[G].\] Similarly, \[\mathop{{}\mathbb{E}}_B[(\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G(D' \otimes I_k) G]] = \mathop{{}\mathbb{E}}_B \operatorname{tr}_{nk}[G(D' \otimes I_k) G].\] By the Poincaré inequality (), \[\begin{align} \Bigl|\mathop{{}\mathbb{E}}_B \operatorname{tr}_k[(\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G](\operatorname{tr}_n \otimes \mathop{\mathrm{id}})&[G(D' \otimes I_k) G]] - \mathop{{}\mathbb{E}}_B \operatorname{tr}_{nk}[G] \mathop{{}\mathbb{E}}_B \operatorname{tr}_{nk}[G(D' \otimes I_k) G] \Bigr| \nonumber \\ &= |\mathop{{}\mathbb{E}}_B \angles{f(B),g(B)}_{\operatorname{tr}_k} - \angles{\mathop{{}\mathbb{E}}_B[f(B)], \mathop{{}\mathbb{E}}_B[g(B)]}_{\operatorname{tr}_k}| \nonumber \\ &= |\mathop{{}\mathbb{E}}_B \angles{f(B) - \mathop{{}\mathbb{E}}_B f(B), g(B) - \mathop{{}\mathbb{E}}_B g(B)}| \nonumber \\ &\leqslant\frac{\dim_{\mathbb{R}} M_k(\mathbb{C})}{(nk)^2} \left\lVert{f}\right\rVert_{\operatorname{Lip}} \left\lVert{g}\right\rVert_{\operatorname{Lip}} \nonumber \\ &\leqslant\frac{2}{n^2} \frac{2 \beta}{|\mathop{\mathrm{Im}}z|^2} \frac{4 \beta \left\lVert{D'}\right\rVert}{|\mathop{\mathrm{Im}}z|^3} \nonumber \\ &= \frac{16 \beta^2 \left\lVert{D'}\right\rVert}{n^2 |\mathop{\mathrm{Im}}z|^5}. \label{eq:32first32application32of32Poincare} \end{align}\tag{48}\] We can also apply the same reasoning to the scalar-valued functions \(B \mapsto \operatorname{tr}_{nk}[G]\) and \(B \mapsto \operatorname{tr}_{nk}[G(D' \otimes I_k) G]\), and thus we get \[\label{eq:32second32application32of32Poincare} |\mathop{{}\mathbb{E}}_B[\operatorname{tr}_{nk}[G] \operatorname{tr}_{nk}[G(D' \otimes I_k) G]] - \mathop{{}\mathbb{E}}_B \operatorname{tr}_{nk}[G] \mathop{{}\mathbb{E}}_B \operatorname{tr}_{nk}[G(D' \otimes I_k) G]| \leqslant\frac{16 \beta^2 \left\lVert{D'}\right\rVert}{(nk)^2 |\mathop{\mathrm{Im}}z|^5}.\tag{49}\] Combining 48 and 49 with the triangle inequality shows that \[\begin{gather} \Bigl|\mathop{{}\mathbb{E}}_B \operatorname{tr}_k[(\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G(z,t)](\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G(z,t)(D' \otimes I_k) G(z,t)]] - \mathop{{}\mathbb{E}}_B[\operatorname{tr}_{nk}[G(z,t)] \operatorname{tr}_{nk}[G(z,t)(D' \otimes I_k) G(z,t)]] \Bigr| \\ \leqslant\frac{16 \beta^2 \left\lVert{D'}\right\rVert}{n^2 |\mathop{\mathrm{Im}}z|^5} + \frac{16 \beta^2 \left\lVert{D'}\right\rVert}{(nk)^2 |\mathop{\mathrm{Im}}z|^5}. \end{gather}\] By integrating over \(A\), we can obtain the same inequality with \(\mathop{{}\mathbb{E}}\) rather than \(\mathop{{}\mathbb{E}}_B\). Finally, we recall that \(\left\lVert{D'}\right\rVert \leqslant n^{1/2} \left\lVert{D'}\right\rVert_2\) to complete the proof. ◻

Lemma 23. \[|h'(t)| \leqslant\frac{32 \beta^4 \left\lVert{D'}\right\rVert_2}{n^{3/2} |\mathop{\mathrm{Im}}z|^5} + \frac{32 \beta^4 \left\lVert{D'}\right\rVert_2}{n^{3/2}k^2 |\mathop{\mathrm{Im}}z|^5} + \frac{2 \beta^2 \left\lVert{D'}\right\rVert_2}{n |\mathop{\mathrm{Im}}z|^3} + \frac{2 \beta^2 \left\lVert{D'}\right\rVert_2}{nk |\mathop{\mathrm{Im}}z|^3}.\]

Proof. By substituting 36 and 41 into 35 , we have \[\begin{align} h'(t) &= 2 \beta^2 \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[G] \operatorname{tr}_{nk}[G(D' \otimes I_k)G] - 2 \beta^2 \mathop{{}\mathbb{E}}\operatorname{tr}_k[(\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G] (\operatorname{tr}_n \otimes \mathop{\mathrm{id}})[G(D' \otimes I_k) G]] \tag{50} \\ &\quad + \frac{\beta^2}{nk} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[G^3 (D' \otimes I_k)] \tag{51} \\ &\quad -\frac{2\beta^2}{n} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[(\mathsf{T}\otimes \mathop{\mathrm{id}})[G] G (D' \otimes I_k) G]. \tag{52} \end{align}\] We then estimate 50 by Lemma 22, estimate 51 by 46 , and estimate 52 by 47 , which yields the bound asserted in this lemma. ◻

Proof of . Recall the values of \(h(0)\) and \(h(1)\) from 33 and 34 , note that \(|h(0) - h(1)| \leqslant\sup_{t \in [0,1]} |h'(t)|\), and apply our estimate for \(h'\) from Lemma 23 to obtain \[\begin{gather} \label{eq:32final32expectation32error32estimate32finite32k} |\mathop{{}\mathbb{E}}\operatorname{tr}_n[(z I_n - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1} D'] - \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[(z I_{nk} - 2\beta B_{\mathop{\mathrm{sym}}} + D \otimes I_k)^{-1} (D' \otimes I_k)]| \\ \leqslant\frac{32 \beta^4 \left\lVert{D'}\right\rVert_2}{n^{3/2} |\mathop{\mathrm{Im}}z|^5} + \frac{32 \beta^4 \left\lVert{D'}\right\rVert_2}{n^{3/2} k^2 |\mathop{\mathrm{Im}}z|^5} + \frac{2 \beta^2 \left\lVert{D'}\right\rVert_2}{n |\mathop{\mathrm{Im}}z|^3} + \frac{2 \beta^2 \left\lVert{D'}\right\rVert_2}{nk |\mathop{\mathrm{Im}}z|^3}. \end{gather}\tag{53}\] Note that \(2B_{\mathop{\mathrm{sym}}}\) is \(\sqrt{2}\) times a standard \(nk \times nk\) GOE matrix. Moreover, \(D \otimes I_k\) and \(D' \otimes I_k\) are deterministic matrices. Thus, the asymptotic freeness theorem () implies that, for every non-commutative polynomial \(p\), the traces \(\operatorname{tr}_n[p(D \otimes I_k, D' \otimes I_k, 2 B_{\mathop{\mathrm{sym}}})]\) converge almost surely as \(k \to \infty\) to \(\tau_n[p(D,D',\sqrt{2} S)]\). The same also applies to the resolvent since this can be approximated by polynomials uniformly on compact subsets of the upper half-plane. Hence, \[\lim_{k \to \infty} \mathop{{}\mathbb{E}}\operatorname{tr}_{nk}[(z I_{nk} - 2 \beta B_{\mathop{\mathrm{sym}}} + D \otimes I_k)^{-1} (D' \otimes I_k)] = \tau_n[(z - \sqrt{2} \beta S + D)^{-1} D'].\] Thus, when we take \(k \to \infty\) in 53 , we obtain \[|\mathop{{}\mathbb{E}}\operatorname{tr}_n[(z - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}D'] - \tau_n[(z - \sqrt{2} \beta S + D)^{-1}D']| \leqslant\frac{32 \beta^4 \left\lVert{D'}\right\rVert_2}{n^{3/2} |\mathop{\mathrm{Im}}z|^5} + \frac{2 \beta^2 \left\lVert{D'}\right\rVert_2}{n |\mathop{\mathrm{Im}}z|^3}.\] By , letting \(f(z)\) be the subordination function, we obtain \[\tau_n[(z - \sqrt{2} \beta S + D)^{-1}D'] = \tau_n[(f(z) + D)^{-1} D'] = \operatorname{tr}_n[(f(z) + D)^{-1}D'],\] so that \[\left|\operatorname{tr}_n[(\mathop{{}\mathbb{E}}[(z - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}] - (f(z) + D)^{-1})D'] \right| \leqslant\left( \frac{32 \beta^4}{n^{3/2} |\mathop{\mathrm{Im}}z|^5} + \frac{2 \beta^2}{n |\mathop{\mathrm{Im}}z|^3} \right) \left\lVert{D'}\right\rVert_2.\] Taking the supremum over \(\left\lVert{D'}\right\rVert_2 \leqslant 1\), we obtain the \(\left\lVert{\cdot}\right\rVert_2\)-norm of \(\mathop{{}\mathbb{E}}\circ E_{\mathcal{D}_n}[(z - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}] - (f(z) + D)^{-1}\), and so ?? follows. ◻

4.5 Conclusion of the random matrix argument↩︎

Proof of . Let \(b\), \(\tilde{a}\), \(\tilde{b}\), \(\tilde{z}\) be as in the statement of the theorem, namely, \(b = \beta n^{-\delta}\), and \(\tilde{z} := ib + 2 \beta^2 g_{-D}(ib)\) where \(g_{-D}(z) = \operatorname{tr}_n((z+D)^{-1})\), and \(\tilde{a}\) and \(\tilde{b}\) are the real and imaginary parts of \(\tilde{z}\).

stating that \(2 \beta^2 \operatorname{tr}_n(D^{-1})\) is the maximum of the spectrum of \(\sqrt{2} \beta S - D\) where \(S\) is a semicircular operator freely independent of \(D\) follows from (2). Now recall \[\tilde{a} = 2 \beta^2 \mathop{\mathrm{Re}}\operatorname{tr}_n((ib+D)^{-1}),\] and so \[|2 \beta^2 \operatorname{tr}_n(D^{-1}) - \tilde{a}| \leqslant 2 \beta^2 |\mathop{\mathrm{Re}}\operatorname{tr}_n[D^{-1} - (ib + D)^{-1}]|.\] We have by and resolvent identities that \[\begin{align} \mathop{\mathrm{Re}}\operatorname{tr}_n[D^{-1} - (ib + D)^{-1}] &= \operatorname{tr}_n[D^{-1} - D(b^2 + D^2)^{-1}] \\ &= \operatorname{tr}_n[D[D^{-2} - (b^2 + D^2)^{-1}]] \\ &= \operatorname{tr}_n[DD^{-2}b^2(b^2+D^2)^{-1}] \\ &\leqslant b^2 \operatorname{tr}_n[D^{-3}]. \end{align}\] This proves the asserted estimate on \(a + 2 \beta^2 \operatorname{tr}_n[D^{-1}] - \tilde{a}\) since \(b^2 = \beta^2 n^{-2\delta}\).

For , note that \[\operatorname{tr}_n[P^2(\tilde{a} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^2] = \operatorname{tr}_n[\tilde{b}(\tilde{b}^2+(\tilde{a} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^2)^{-1}(\tilde{a} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^2] \leqslant\tilde{b},\] and by (1), we have \(\tilde{b} \leqslant 2 \beta^2 b^3 \operatorname{tr}_n(D^{-4})\) and then we substitute \(b = \beta n^{-\delta}\).

For , note that \[P^2 = -\mathop{\mathrm{Im}}[(\tilde{a} + i\tilde{b} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}],\] and observe that \(\mathop{\mathrm{Im}}\circ E_{\mathcal{D}_n} = E_{\mathcal{D}_n} \circ \mathop{\mathrm{Im}}\). Letting \(z := ib\), we estimate \[\begin{align} \left\lVert{E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}] - (z + D)^{-1}}\right\rVert_2 &\leqslant\left\lVert{E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}] - \mathop{{}\mathbb{E}}E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}]}\right\rVert_2 \tag{54} \\ &\qquad\qquad + \left\lVert{\mathop{{}\mathbb{E}}E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}] - (z + D)^{-1}}\right\rVert_2 \tag{55} \end{align}\] By , with probability at least \(1 - M_1 \exp(-M_2 n)\) in the Gaussian matrix \(A\), we have for all choices of \(D\) that \[\label{eq:32RMT32main32splitting32solution} \left\lVert{E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}] - \mathop{{}\mathbb{E}}E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}]}\right\rVert_2 \leqslant M_3 \beta n^{-2\delta},\tag{56}\] which takes care of 54 .

To estimate 55 , let \(f\) be the subordination function as in ; by construction \(\tilde{z} = z + 2 \beta^2 g_{-D}(z)\), and so \(f(\tilde{z}) = z\). Hence, by , we have \[\left\lVert{\mathop{{}\mathbb{E}}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}] - (z + D)^{-1}}\right\rVert_2 \leqslant\frac{32 \beta^4}{n^{3/2} \tilde{b}^5} + \frac{2 \beta^2}{n \tilde{b}^3}.\] Then recall by (1) that \(\tilde{b}^{-1} \leqslant 3\beta^2 b^{-3} = 3 \beta^{-1} n^{3\delta}\). Also, \(E_{\mathcal{D}_n}\) is contractive in \(\left\lVert{\cdot}\right\rVert_2\), so that \[\left\lVert{\mathop{{}\mathbb{E}}E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}] - (z + D)^{-1}}\right\rVert_2 \leqslant\frac{M_4}{\beta n^{3/2-15\delta}} + \frac{M_5}{\beta n^{1-9\delta}}.\] Since \(\delta \leqslant 1/17 < 3/34 < 1/11\), we have that \(3/2 - 15 \delta \geqslant 2 \delta\) and \(1 - 9 \delta \geqslant 2 \delta\). Therefore, \[\label{eq:32RMT32main32splitting32solution322} \left\lVert{\mathop{{}\mathbb{E}}E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}] - (z + D)^{-1}}\right\rVert_2 \leqslant M_6 \beta^{-1} n^{-2\delta}.\tag{57}\] Estimating 54 by 56 and 55 by 57 and bounding \(\beta^{-1}\) by a constant, we get \[\left\lVert{E_{\mathcal{D}_n}[(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}] - (z + D)^{-1}}\right\rVert_2 \leqslant M_7 \beta n^{-2\delta}.\] Then taking the negated imaginary parts of the operators, we get \[\left\lVert{E_{\mathcal{D}_n}[\tilde{b}(\tilde{b}^2 + (\tilde{a} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^2)^{-1}] - b(b^2 + D^2)^{-1}}\right\rVert_2 \leqslant M_7 \beta n^{-2\delta},\] which is the desired estimate with \(C_3 = M_7\).

For , note by (2), \[\operatorname{tr}_n[b(b^2+D^2)^{-1}] = -\mathop{\mathrm{Im}}g_{-D}(ib) \geqslant\frac{1}{2 \beta^2} (b - b^3 \left\lVert{D^{-1}}\right\rVert^2) \geqslant\frac{1}{2} \beta^{-1} n^{-\delta} (1 - \beta^2 n^{-2\delta} \left\lVert{D^{-1}}\right\rVert^2).\] Then using , \[\begin{align} \operatorname{tr}_n[P^2] = \operatorname{tr}[E_{\mathcal{D}_n}(P^2)] &\geqslant\operatorname{tr}_n[b(b^2+D^2)^{-1}] - \left\lVert{E_{\mathcal{D}_n}(P^2) - b(b^2+D^2)^{-1}}\right\rVert_2 \\ &\geqslant\frac{1}{2} \beta^{-1} n^{-\delta} - M_7 \beta n^{-2\delta} - \frac{1}{2} \beta \left\lVert{D^{-1}}\right\rVert^2 n^{-3\delta}. \end{align}\]

For , observe that \[\left\lVert{P^2}\right\rVert = \left\lVert{\mathop{\mathrm{Im}}(\tilde{z} - 2 \beta A_{\mathop{\mathrm{sym}}} + D)^{-1}}\right\rVert \leqslant\frac{1}{\mathop{\mathrm{Im}}\tilde{z}} \leqslant\frac{3 \beta^2}{b^3} = 3 \beta^{-1} n^{3\delta}. \qedhere\] ◻

5 Convergence to the Primal Auffinger-Chen SDE under fRSB↩︎

In this section, we analyze the steps of our algorithm and describe how it converges in empirical distribution to the solution to the primal Auffinger-Chen SDE. To this end, for a vector \(v\in \mathbb{R}^n\), we define \(\mathop{\mathrm{emp}}(v) := \frac{1}{n} \sum_{j=1}^n \delta_{v_j}\) to refer to the empirical distribution of coordinates of that vector and for \(Y\) a random variable in \(\mathbb{R}\), we let \(\operatorname{dist}(Y) \in \mathcal{P}(\mathbb{R})\) denote the probability distribution of \(Y\).

In , with high probability over \(A\), we constructed, for any diagonal matrix \(D\) with \(2 \beta^2 \operatorname{tr}(D^{-2}) = 1\), a positive semi-definite matrix \(P(D)\) whose mass was concentrated on the highest part of the bulk of the spectrum of \(2 \beta A_{\mathop{\mathrm{sym}}} - D\). While the normalization of \(P(D)\) was convenient for the analysis of resolvents in the last section, it will be helpful now to renormalize it so that the covariance matrix we use in the algorithm has normalized trace approximately \(1\). In addition, to make the Taylor expansion analysis easier in the next section, we will arrange that our step is exactly orthogonal to the current position. Thus, we introduce the following quantities: For any \(\sigma \in \mathbb{R}^n\), let \[D(t,\sigma) := \left( \frac{2 \beta^2}{n} \sum_{j=1}^n \partial_{y,y} \tilde{\Lambda}_{\gamma} \left(t, \sigma_j\right)^{-2} \right)^{1/2} \operatorname{diag}\left[\partial_{y,y} \tilde{\Lambda}_{\gamma} \left(t, \sigma_j\right)\right].\] Note that by construction \(2 \beta^2 \operatorname{tr}_n(D(t,\sigma)^{-2}) = 1\).

Then, let \(Q(t,\sigma)\) be the positive semi-definite matrix given by \[\begin{align} Q(t,\sigma)^2 &:= 2 \beta n^{\delta} \Pi_{\sigma^\perp} P(D(t,\sigma))^2 \Pi_{\sigma^\perp}\\ &= 2 \beta n^\delta \Pi_{\sigma^\perp} \tilde{b}(\tilde{b}^2 + (\tilde{a} - 2 \beta A_{\mathop{\mathrm{sym}}} + D(t,\sigma))^2)^{-1} \Pi_{\sigma^\perp}, \end{align}\] where we use the notation from . As we will see below, \(\operatorname{tr}_n(Q(t,\sigma)^2)\) tends to \(1\) as \(n \to \infty\), and the diagonal entries of \(Q(t,\sigma)^2\) are well-approximated by \(D(t,\sigma)^{-2}\).

For our algorithm, suppose holds, fix a step size \(\eta\) and number of steps \(K\) such that \(K \eta \leqslant q_\beta^*\). For \(k = 0\), …, \(K\), write \(t_k = k \eta\). Define a random vector \(\sigma_k\) in \(\mathbb{R}^n\) inductively as follows: Let \(\sigma_0 := 0\), and for \(k \in \{0, \dots, K-1\}\), let \[\label{eq:32sigma32step} \sigma_{k+1} := \sigma_k + \eta^{1/2} Q(t_k,\sigma_k) Z_k,\tag{58}\] where \(Z_0, Z_1\), \(Z_2\), …, \(Z_{K-1}\) are independent standard Gaussian random vectors in \(\mathbb{R}^n\).

As in , let \(Y_t^{\gamma}\) be the solution to the SDE: \[dY_t^{\gamma} = \frac{\sqrt{2} \beta}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,Y^{\gamma}_t)} \,dW_t,\;\; Y_0 = 0.\] where \(W_t\) is a standard Brownian motion (see ). By (3), \(1 / \partial_{y,y} \tilde{\Lambda}_{\gamma}(t,y)\) is \(2\)-Lipschitz in space, which implies that the solution to the SDE is well-defined. For convenience, we write \[v(t,y) := \frac{1}{ \partial_{y,y} \tilde{\Lambda}_{\gamma}(t,y)}\,.\]

For the remainder of this section, we use the following abbreviated notations: \[\begin{align} D_k &:= D(t_k, \sigma_k) \\ P_k &:= P(D_k) \\ Q_k &:= Q(t_k, \sigma_k) \\ v_k(y) &:= v(t_k, y) \end{align}\] We also understand \(v_k(\sigma_k)\) to be the component-wise application of \(v_k\) to the vector \(\sigma_k\), i.e. the \(i\)th component of \(v_k(\sigma_k)\) is \(v_k((\sigma_k)_i)\).

Theorem 27 (Convergence of empirical distribution to primal SDE). Assume fRSB. Let \(\delta \in (0,1/19]\). For some universal constants \(C_1\), \(C_2\), …, the following statement holds. Suppose that the matrix \(A\) satisfies the conclusions of , which happens with probability \(1 - C_1 \exp(-C_2 n)\). Suppose that \(\eta \in (0,1)\) and \(\beta \in (1,\infty)\), \(\gamma > 0\), and \(n \in \mathbb{N}\) such that \(\eta n^{1/90} \geqslant 1\) and \[\label{eq:32induction32hypothesis32for32all32thm} C_4^{-1} (\exp( C_4 \beta^2) - 1) \left[ C_5 \beta^4 \eta + C_6 n^{-\delta} + C_7 e^{10 \beta^2} \gamma^2 + C_8 \beta^{-2} \eta^{-1} n^{-1/18} \right] \leqslant\frac{1}{8 \beta^2}.\tag{59}\] Let \(k' \in \mathbb{N}\) with \(k' \leqslant K \leqslant 1/\eta\), and let \(\sigma_k\) be defined as in for \(k = 0\), …, \(k'\). Then with probability \[1 - C_9 n \eta^{-2} \exp(-C_{10} n^{2/9-4\delta}),\] in the Gaussian vectors \(Z_0\), …, \(Z_{k'-1}\) from , we have that for \(k = 1\), …, \(k'\), \[d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k), \operatorname{dist}(Y_{t_k}^\gamma))^2 \leqslant C_4^{-1} (\exp( C_4 \beta^2 t_k) - 1) \left[ C_5 \beta^4 \eta + C_6 n^{-\delta} + C_7 e^{10 \beta^2} \gamma^2 + C_8 \beta^{-2} \eta^{-1} n^{-1/18} \right]\] and \[2d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_{k}), \operatorname{dist}(Y_{t_{k}}^\gamma)\right) + (1 + e^{5 \beta^2}) \gamma \leqslant\frac{1}{2 \sqrt{2} \beta} .\]

We argue inductively that with high probability \(d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k), \operatorname{dist}(Y^{\gamma}_{t_k}))\) is small. The bulk of the argument is for the inductive step, and during this argument we will assume that \(\sigma_k\) has already been chosen, and hence we treat it as deterministic. Then we go from \(\mathop{\mathrm{emp}}(\sigma_k)\) to \(\operatorname{dist}(Y_{t_k}^\gamma)\) in several steps. We compare the random empirical distribution of the next step \(\mathop{\mathrm{emp}}(\sigma_{k+1})\) with its expectation using concentration of measure (see §5.3). We compare the expectation with the distribution of \(Y_{t_{k+1}}^\gamma\) using the inductive hypothesis and Lipschitz estimates on \(v\) (see §5.2), for which we also need to use the control of the diagonal entries of the matrix \(Q_k\) from (see §5.1).

We will continue to use \(C_1\), \(C_2\), …to denote absolute constants in each statement and \(M_1\), \(M_2\), …for constants in the proofs. We continue to assume \(\beta \geqslant 1\).

5.1 Controlling the diagonal of the covariance↩︎

Here we will estimate the diagonal of the covariance matrix \(Q_k^2\). From , we can approximate the diagonal entries of \(P(D_k)^2\) by \(b(b^2 + D_k^2)^{-1}\) where \(b = \beta n^{-\delta}\). We want to show this is close to \(b D_k^{-2}\), the diagonal matrix whose entries are given by \(v_k(\sigma_k)\), so that the increments of \(\sigma_k\) will resemble those in the SDE. We first estimate the error from the normalization we performed on \(D_k\) to make \(2 \beta^2 \operatorname{tr}_n(D_k^{-2}) = 1\).

Lemma 24. We have \[\begin{align} \left\lVert{D_k^{-1} - \operatorname{diag}(v_k(\sigma_{k}))}\right\rVert_2 &= \left| \left( \frac{1}{n} \sum_{j=1}^n \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j})^2} \right)^{1/2} - \frac{1}{\sqrt{2} \beta} \right| \\ &\leqslant 2 d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k), \operatorname{dist}(Y_{t_k}^\gamma)) + (1 + e^{5 \beta^2}) \gamma. \end{align}\]

Proof. By definition, \[\sqrt{2} \beta D_k^{-1} = \frac{\sqrt{2} \beta \operatorname{diag}(v_k(\sigma_{k}))}{\left\lVert{\sqrt{2} \beta \operatorname{diag}(v_k(\sigma_{k}))}\right\rVert_2}.\] In general, for a nonzero vector \(h\) in a Hilbert space, \[\left\lVert{\frac{h}{\left\lVert{h}\right\rVert} - h}\right\rVert = \left| \frac{1}{\left\lVert{h}\right\rVert} - 1 \right| \left\lVert{h}\right\rVert = \left| \left\lVert{h}\right\rVert - 1 \right|.\] Applying this to \(\sqrt{2} \beta \operatorname{diag}(v_k(\sigma_{k,j}))\) yields the first equality. For the second estimate, let \(\Sigma\) be a random variable on the same probability space as \(Y_{t_k}^\gamma\) such that the marginal of \(\Sigma\) is \(\mathop{\mathrm{emp}}(\sigma_k)\) and \(\Sigma\) and \(Y_{t_k}^\gamma\) are optimally coupled, i.e. \(\left\lVert{\Sigma - Y_{t_k}^\gamma}\right\rVert_{L^2} = d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k),\operatorname{dist}(Y_t^\gamma))\). Since \(v\) is \(2\)-Lipschitz by , we have \[\left|\frac{1}{\sqrt{n}} \left|v_k(\sigma_{k})\right|_2 - \left\lVert{v_k(Y_{t_k}^\gamma)}\right\rVert_{L^2} \right| \leqslant\left\lVert{v_k(\Sigma) - v_k(Y_{t_k}^\gamma)}\right\rVert_{L^2} \leqslant 2 d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k), \operatorname{dist}(Y_{t_k}^\gamma)).\] We then recall from that \[\left| \left\lVert{v_k(Y_{t_k}^\gamma)}\right\rVert_{L^2} - \frac{1}{\sqrt{2} \beta} \right| \leqslant(1 + e^{5 \beta^2 t_k}) \gamma \leqslant(1 + e^{5 \beta^2}) \gamma.\] Combining these estimates yields the asserted statement. ◻

Now we are ready to control the diagonal of \(Q_k^2\). At the same time, we will update the properties of for the new normalization of \(Q\) for later use in §6.

Proposition 28. Suppose that \(A\) is chosen from the high probability event in . Suppose that \(\gamma < 2\) and \[2 d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_k), \operatorname{dist}(Y_{t_k}^\gamma)\right) + (1 + e^{5 \beta^2}) \gamma \leqslant\frac{1}{2 \sqrt{2} \beta}.\] Then

  1. Approximation for diagonal: \(\left\lVert{E_{\mathcal{D}_n}[Q_k^2] - 2 \beta^2 D_k^{-2} }\right\rVert_2 \leqslant C_1 \beta^2 n^{-\delta}\).

  2. Approximation for diagonal in square root: \(\left\lVert{(E_{\mathcal{D}_n}[Q_k^2])^{1/2} - \sqrt{2} \beta D_k^{-1}}\right\rVert_2 \leqslant C_3 \beta n^{-\delta/2}\).

  3. Approximation for trace: \(|\operatorname{tr}_n[Q_k^2] - 1| \leqslant C_1 \beta^2 n^{-\delta}\).

  4. Hilbert-Schmidt norm of diagonal: \(\left\lVert{E_{\mathcal{D}_n}[Q_k^2]}\right\rVert_2 \leqslant C_2 \beta^2\).

  5. Operator norm bound: \(\left\lVert{Q_k^2}\right\rVert \leqslant C_4 n^{4 \delta}\).

  6. Approximate eigenvector condition: \[\left\lVert{\left(2 \beta A_{\mathop{\mathrm{sym}}} - D_k - 2 \beta^2 \operatorname{tr}_n(D_k^{-1})\right) Q_k}\right\rVert_2 \leqslant C_5 \beta^2 n^{-\delta} + C_6(\beta+\gamma^{-1})n^{-1/4+2\delta}.\]

Proof. (1) By , writing \(b = \beta n^{-\delta}\), \[\label{eq:5463461-first-bound} \left\lVert{E_{\mathcal{D}_n}[P_k^2] - b (b^2 + D_k^2)^{-1}}\right\rVert_1 \leqslant\left\lVert{E_{\mathcal{D}_n}[P_k^2] - b (b^2 + D_k^2)^{-1}}\right\rVert_2 \leqslant M_1 \beta n^{-2\delta}.\tag{60}\] We define an intermediate term \(\Xi\) and find a common denominator \[\Xi := D_k^{-2} - (b^2 + D_k^2)^{-1} = b^2 (b^2 + D_k^2)^{-1} D_k^{-2}.\] Hence, \[\begin{align} \left\lVert{ \Xi }\right\rVert_1 &\leqslant b^2 \left\lVert{(b^2 + D_k^2)^{-1}}\right\rVert \left\lVert{D_k^{-2}}\right\rVert_1 \notag \\ &\leqslant b^2 \left\lVert{D_k^{-1}}\right\rVert^2 \frac{1}{2 \beta^2} \label{eq:Xi-bound} \end{align}\tag{61}\] Now recall that \(D_k^{-1} = \operatorname{diag}(v_k(\sigma_{k})) / \left\lVert{\sqrt{2} \beta \operatorname{diag}(v_k(\sigma_{k}))}\right\rVert_2\), and \(v \leqslant(1 + \gamma)\) by so that \[\left\lVert{D_k^{-1}}\right\rVert \leqslant\frac{1+\gamma}{\left\lVert{\sqrt{2} \beta \operatorname{diag}(v_k(\sigma_{k}))}\right\rVert_2}\] By , if \(2 d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k), \operatorname{dist}(Y_{t_k}^\gamma)) + (1 + e^{5 \beta^2}) \gamma \leqslant\frac{1}{2 \sqrt{2} \beta}\), then we obtain \[\left\lVert{\operatorname{diag}(v_k(\sigma_{k}))}\right\rVert_2 \geqslant\frac{1}{\sqrt{2} \beta} - \frac{1}{2\sqrt{2} \beta} = \frac{1}{2 \sqrt{2} \beta}.\] Hence, \(\left\lVert{D_k^{-1}}\right\rVert \leqslant 2(1+\gamma) \leqslant 6\), and so, substituting the last two inequalities into , \[\left\lVert{\Xi }\right\rVert_1 \leqslant\frac{4(1+\gamma)^2}{2 \beta^2} b^2 = 2(1+\gamma)^2n^{-2\delta}.\] Therefore, by the triangle inequality with the preceding bound and and because \(\gamma < 2\), \[\left\lVert{E_{\mathcal{D}_n}[P_k^2] - b D_k^{-2}}\right\rVert_1 \leqslant M_1 \beta n^{-2\delta} + 2 b n^{-2\delta} \leqslant M_2 \beta n^{-2\delta}.\] Furthermore, by , \[\left\lVert{P_k^2 - \Pi_{\sigma_k^\perp} P_k^2 \Pi_{\sigma_k^\perp}}\right\rVert_2 \leqslant 2 \left\lVert{P_k^2}\right\rVert \, \left\lVert{1 - \Pi_{\sigma_k^\perp}}\right\rVert_2 \leqslant\frac{C}{n^{1/2} \tilde{b}} \leqslant M_3 \beta^2 b^{-3} n^{-1/2} = M_4 \beta^{-1} n^{-1/2+3\delta}.\] We can bound \(n^{-1/2+3\delta}\) by \(n^{-2\delta}\) as well since \(\delta \leqslant 1/17\). Hence, by triangle inequality with the two preceding bounds and the contractivity of \(E_{\mathcal{D}_n}\), \[\left\lVert{E_{\mathcal{D}_n}[\Pi_{\sigma_k^\perp} P_k^2 \Pi_{\sigma_k^\perp}] - b D_k^{-2}}\right\rVert_2 \leqslant M_4 \beta n^{-2\delta}.\] Then multiplying by \(2 \beta^2 / b = 2 \beta n^\delta\), \[\left\lVert{E_{\mathcal{D}_n}[Q_k^2] - 2 \beta^2 D_k^{-2}}\right\rVert_2 \leqslant 2M_4 \beta^2 n^{-\delta}.\] This proves the first claim.

(2) By the Powers-Størmer inequality, since \(E_{\mathcal{D}_n}[Q_k^2]^{1/2}\) and \(D_k\) are nonnegative matrices, we have \[\begin{align} \left\lVert{E_{\mathcal{D}_n}[Q_k^2]^{1/2} - \sqrt{2} \beta D_k^{-1}}\right\rVert_2 &\leqslant\left\lVert{E_{\mathcal{D}_n}[Q_k^2] - 2 \beta^2 D_k^{-2}}\right\rVert_1^{1/2} \\ &\leqslant\left\lVert{E_{\mathcal{D}_n}[Q_k^2] - 2 \beta^2 D_k^{-2}}\right\rVert_2^{1/2} \\ &\leqslant(2M_4)^{1/2} \beta n^{-\delta/2}. \end{align}\]

(3) The third claim is immediate from the first claim since \(2 \beta^2 \operatorname{tr}_n(D_k^{-2}) = 1\) by construction.

(4) Note that \[\left\lVert{E_{\mathcal{D}_n}[Q_k^2]}\right\rVert_2 \leqslant 72 \beta^2 \left\lVert{D_k^{-2}}\right\rVert_2 + 2M_4 \beta^2 n^{-\delta} \leqslant 8 \beta^2 + 2M_4 \beta^2\] since we have \(\left\lVert{D_k^{-1}}\right\rVert \leqslant 6\).

(5) Note that by , \[\left\lVert{Q_k^2}\right\rVert \leqslant 2 \beta n^\delta \left\lVert{P_k^2}\right\rVert \leqslant C n^{4\delta}.\]

(6) Using , letting \(\overline{Q}_k^2 := 2 \beta n^\delta P_k^2\), we get \[\left\lVert{\left(2 \beta A_{\mathop{\mathrm{sym}}} - D_k - \tilde{a}\right) \overline{Q}_k}\right\rVert_2^2 = \operatorname{tr}_n\left(\overline{Q}_k^2 (\tilde{a} - 2 \beta A_{\mathop{\mathrm{sym}}} + D_k)^2 \right) \leqslant 2 \beta^4 n^{-2\delta} \left\lVert{D_k^{-1}}\right\rVert^2 \leqslant 72 \beta^4 n^{-2 \delta}\] since \(\left\lVert{D_k^{-1}}\right\rVert \leqslant 6\). By the Powers-Størmer inequality, \[\left\lVert{Q_k - \overline{Q}_k}\right\rVert_2^2 \leqslant\left\lVert{Q_k^2 - \overline{Q}_k^2}\right\rVert_1 \leqslant 2\beta n^{\delta}\left(2\left\lVert{(1 - \Pi_{\sigma_k^\perp})P_k^2}\right\rVert_1 + \left\lVert{(1 - \Pi_{\sigma_k^\perp})P_k^2(1 - \Pi_{\sigma_k^\perp})}\right\rVert_1\right),\] which is bounded above by \(6\beta n^{\delta-1/2}\left\lVert{P_k^2}\right\rVert\), since \(\left\lVert{1 - \Pi_{\sigma_k^\perp}}\right\rVert_2 = n^{-1/2}\). So by triangle inequality, \[\label{eq:5463466-third-bound} \left\lVert{\left(2 \beta A_{\mathop{\mathrm{sym}}} - D_k - \tilde{a}\right) Q_k}\right\rVert_2 \leqslant 6 \sqrt{2} \beta^2 n^{-\delta} + \sqrt{6\beta} n^{\delta/2-1/4}\left\lVert{P_k}\right\rVert\left\lVert{2 \beta A_{\mathop{\mathrm{sym}}} - D_k - \tilde{a}}\right\rVert.\tag{62}\] Then recall from that \[|\tilde{a} - 2 \beta^2 \operatorname{tr}_n(D_k^{-1})| \leqslant 2 \beta^4 n^{-2\delta} \operatorname{tr}_n(D_k^{-3}) \leqslant n^{-2 \delta} \left\lVert{D_k^{-1}}\right\rVert \cdot 2 \beta^4 \operatorname{tr}_n(D_k^{-2}) \leqslant 12 \beta^2 n^{-2\delta}.\] Hence, by , \[\left\lVert{\left(\tilde{a} - 2 \beta^2 \operatorname{tr}_n(D_k^{-1})\right) Q_k}\right\rVert_2 \leqslant M_5 \beta \cdot 2\beta^2 n^{-2\delta}.\] Therefore, applying triangle inequality with and the preceding bound, \[\left\lVert{\left(2 \beta A_{\mathop{\mathrm{sym}}} - D_k - 2 \beta^2 \operatorname{tr}_n(D_k^{-1})\right) Q_k}\right\rVert_2 \leqslant 6 \sqrt{2} \beta^2 n^{-\delta} + M_6n^{-1/4+2\delta}(\beta+\gamma^{-1}) + M_5 \beta^3 \cdot 2 n^{-2\delta} ,\] where we used (1) and the high-probability event from to obtain \(\left\lVert{\tilde{a} - 2 \beta A_{\mathop{\mathrm{sym}}} + D_k}\right\rVert \leqslant M_7(\beta + \gamma^{-1})\) and gives \(\left\lVert{P_k}\right\rVert \leqslant M_8 \beta^{-1/2} n^{3\delta/2}\). ◻

5.2 The expected empirical distribution of the next iterate↩︎

Our next goal is, having fixed \(\sigma_k\), to estimate the distance between the expected empirical distribution \(\mathop{{}\mathbb{E}}[\mathop{\mathrm{emp}}(\sigma_{k+1}) \mid A, \sigma_k]\) and \(\operatorname{dist}(Y_{t_{k+1}}^\gamma \mid Y_{t_{k}}^\gamma)\). First, to make solution of the SDE easier to compare with the result of our algorithm, we want to approximate \(Y_{t_{k+1}}^\gamma\) by \(Y_{t_{k}}^\gamma + \sqrt{2} \beta v_k(Y_{t_k}^\gamma) (W_{t_{k+1}} - W_{t_k})\). For this purpose, we use the following standard estimate for solutions to SDEs.

Lemma 25. For all \(t \in [0,1-\eta]\), we have \[\left\lVert{Y_{t+\eta}^\gamma - \left[Y_t^\gamma + \sqrt{2} \beta v(t,Y_t^\gamma) (W_{t+\eta} - W_t)\right]}\right\rVert_{L^2}^2 \leqslant C \beta^6 \eta^2.\]

Proof. Recall \[Y_{t + \eta}^\gamma - \left(Y_t^\gamma + \sqrt{2} \beta v(t,Y_t^\gamma)(W_{t+\eta} - W_t)\right) = \int_{0 \leqslant h \leqslant\eta} \sqrt{2} \beta \left[v(t+h,Y_{t+h}^\gamma) - v(t,Y_t^\gamma)\right] dW_{t+h}.\] By the Itô isometry, the Lipschitz bounds for \(v\) (), and the arithmetic-geometric mean inequality, \[\begin{align} \mathop{{}\mathbb{E}}\left|Y_{t+\eta}^\gamma - Y_t^\gamma - \sqrt{2} \beta v(t,Y_t^\gamma) (W_{t+\eta} - W_t)\right|^2 &= 2 \beta^2 \int_{0\leqslant h \leqslant\eta} \mathop{{}\mathbb{E}}\left|v(t+h,Y_{t+h}^\gamma) - v(t,Y_t^\gamma)\right|^2\,dh \\ &\leqslant 2 \beta^2 \int_0^\eta \left(M_1 \beta^2 h + 2|Y_{t+h}^\gamma - Y_t^\gamma|\right)^2 \,dh \\ &\leqslant\int_0^\eta \left(4M_1^2 \beta^6 h^2 + 8 \beta^2 |Y^{\gamma}_{t+h} - Y^{\gamma}_t|^2\right)\,dh. \end{align}\] Now note that since \(v\) is bounded by \(1+\gamma\), we also have \[\mathop{{}\mathbb{E}}\left|Y_{t+h}^\gamma - Y_t^\gamma\right|^2 = 2 \beta^2 \int_0^h \left|v(t+h',Y_{t+h'}^\gamma)\right|^2\,dh' \leqslant 2 \beta^2 h(1+\gamma).\] Therefore, we get \[\mathop{{}\mathbb{E}}\left|Y_{t+\eta}^\gamma - Y_t^\gamma - \sqrt{2} \beta v(t,Y^{\gamma}_t) (W_{t+\eta} - W_t)\right|^2 \leqslant\int_0^\eta (4M_1^2 \beta^6 h^2 + 16 \beta^4 h(1+\gamma)) \,dh = M_2 \beta^6 \eta^3 + M_3 \beta^4 \eta^2 \leqslant M_4 \beta^6 \eta^2.\] ◻

Lemma 26. Assume \(A\) satisfies the conclusions of . Let \(\sigma_k\) and \(\operatorname{dist}(Y_{t_k}^\gamma)\) be fixed with \[2 d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_k), \operatorname{dist}(Y_{t_k}^\gamma)\right) + (1 + e^{5 \beta^2 t_k}) \gamma \leqslant\frac{1}{2 \sqrt{2} \beta}.\] Then, \[\begin{gather} d_{W,2}\left(\mathop{{}\mathbb{E}}\left[\mathop{\mathrm{emp}}(\sigma_{k+1}) | A, \sigma_1, \dots, \sigma_k\right], \operatorname{dist}(Y_{t_{k+1}}^\gamma \mid Y_{t_{k}}^\gamma)\right)^2 - d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_k), \operatorname{dist}(Y_{t_k}^\gamma)\right)^2 \\ \leqslant\eta \beta^2\cdot \left[ C_1 \beta^4 \eta + C_2 n^{-\delta} + C_3 d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_k),\operatorname{dist}(Y_{t_k}^\gamma)\right)^2 + C_4 e^{10\beta^2} \gamma^2 \right]. \end{gather}\]

Proof. Consider the discrete probability space \([n]\) with normalized counting measure, and let \(\Sigma\) be a random variable that takes value \(\sigma_{k,j}\) on point \(j\), so that \(\Sigma\) has distribution \(\mathop{\mathrm{emp}}(\sigma_k)\). Also, let \(\Xi\) be the random variable that takes value \(((Q_k^2)_{j,j})^{1/2}\) at point \(j\). Let \(W\) be a normal random variable of variance \(\eta\) independent of \((\Sigma,\Xi)\). We first observe that \(\Sigma + \Xi W\) has the distribution \(\mathop{{}\mathbb{E}}\mathop{\mathrm{emp}}(\sigma_{k+1})\). Indeed, for smooth functions \(f\), \[\mathop{{}\mathbb{E}}\frac{1}{n} \sum_{j=1}^n f(\sigma_{k+1,j}) = \mathop{{}\mathbb{E}}\frac{1}{n} \sum_{j=1}^n f\left(\sigma_{k,j} + \eta^{1/2} (Q_k Z_k)_j\right).\] Since \(\eta^{1/2}(Q_k Z_k)_j\) is a normal random variable of variance \(\eta (Q_k^2)_{j,j}\), the expectation is the same if we replace \(\eta^{1/2}(Q_k Z_k)_j\) with \(((Q_k^2)_{j,j})^{1/2} W\). Thus, \[\mathop{{}\mathbb{E}}\frac{1}{n} \sum_{j=1}^n f\left(\sigma_{k,j} + \eta^{1/2} (Q_k Z_k)_j\right) = \mathop{{}\mathbb{E}}\frac{1}{n} \sum_{j=1}^n f\left(\sigma_{k,j} + ((Q_k^2)_{j,j})^{1/2} W\right) = \mathop{{}\mathbb{E}}f(\Sigma + \Xi W).\]

By choosing an appropriate larger probability space, we can arrange that \(\Sigma\), \(\Xi\), and \((W_t)_{0 \leqslant t \leqslant t_k}\) be in the same probability space (hence also \(Y_t^\gamma\) is on the same probability space since \(Y_t^\gamma\) is obtained by solving the SDE with Brownian motion \(W_t\)) such that \(\Sigma\) and \(Y_{t_{k}}^\gamma\) are optimally coupled, that is, \[\left\lVert{\Sigma - Y_{t_k}^\gamma}\right\rVert_{L^2} = d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_k), \operatorname{dist}(Y^{\gamma}_{t_k})\right).\] Then take the rest of the Brownian motion \((W_t - W_{t_k})_{t \geqslant t_k}\) to be independent of \((\Sigma,\Xi,(W_t)_{0 \leqslant t \leqslant t_{k+1}}\)). Since \(W_{t_{k+1}} - W_{t_k}\) has variance \(\eta\), we can identify it with the \(W\) from the previous paragraph.

Making the conditioning in \(\mathop{{}\mathbb{E}}\left[\mathop{\mathrm{emp}}(\sigma_{k+1}) | A, \sigma_1, \dots, \sigma_k\right]\) notationally implicit for brevity, we want to estimate \[d_{W,2}\left(\mathop{{}\mathbb{E}}\mathop{\mathrm{emp}}(\sigma_{k+1}),\operatorname{dist}(Y_{t_{k+1}}^\gamma \mid Y_{t_{k}}^\gamma)\right) \leqslant\left\lVert{(\Sigma + \Xi W) - Y_{t_{k+1}}^\gamma}\right\rVert_{L^2} = \left\lVert{(\Sigma - Y_{t_k}^\gamma) + (\Xi W - Y_{t_{k+1}}^\gamma + Y_{t_k}^\gamma)}\right\rVert_{L^2}.\] Observe that the conditional expectation of \(\Xi W\) given \((Y_{t_k}^\gamma,\Sigma)\) is zero because \(W\) is independent of these variables. Moreover, since \(Y_{t_{k+1}}^\gamma - Y_{t_k}^\gamma = \int_{t_k}^{t_{k+1}} \sqrt{2} \beta v(t,Y_t^\gamma)\,dW_t\), this also has conditional expectation zero given \((Y_{t_k}^\gamma,\Sigma)\). Thus, \(\Xi W - Y_{t_{k+1}}^\gamma + Y_{t_k}^\gamma\) and \(\Sigma - Y_{t_k}^\gamma\) are orthogonal in \(L^2\) and so \[\left\lVert{\Sigma + \Xi W - Y^{\gamma}_{t_{k+1}}}\right\rVert_{L^2}^2 = \left\lVert{\Sigma - Y_{t_k}^\gamma}\right\rVert_{L^2}^2 + \left\lVert{\Xi W - Y_{t_{k+1}}^\gamma + Y_{t_k}^\gamma}\right\rVert_{L^2}^2.\] Thus, \[\label{eq:lemma-5465-pythagorean} d_{W,2}\left(\mathop{{}\mathbb{E}}\mathop{\mathrm{emp}}(\sigma_{k+1}), \operatorname{dist}(Y^{\gamma}_{t_{k+1}} \mid Y_{t_{k}}^\gamma)\right)^2 \leqslant d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_k),\operatorname{dist}(Y_{t_k}^\gamma)\right)^2 + \left\lVert{\Xi W - Y_{t_{k+1}}^\gamma + Y_{t_k}^\gamma}\right\rVert_{L^2}^2.\tag{63}\] Next, we estimate by triangle inequality \[\label{eq:lemma-5465-first-triangle} \left\lVert{\Xi W - Y_{t_{k+1}}^\gamma + Y_{t_k}^\gamma}\right\rVert_{L^2} \leqslant\left\lVert{\Xi W - \sqrt{2} \beta v_k(Y_{t_k}^\gamma) W}\right\rVert_{L^2} + \left\lVert{Y_{t_{k+1}}^\gamma - Y_{t_k}^\gamma - \sqrt{2} \beta v_k(Y_{t_k}^\gamma) W}\right\rVert_{L^2}.\tag{64}\] By , the second term of satisfies \[\label{eq:lemma-5465-first-triangle-second-term} \left\lVert{Y_{t_{k+1}}^\gamma - Y_{t_k}^\gamma - \sqrt{2} \beta v_k(Y_{t_k}^\gamma) W}\right\rVert_{L^2} \leqslant M_1 \beta^3 \eta.\tag{65}\] Meanwhile, by another triangle inequality, the first term satisfies \[\begin{align} \left\lVert{\Xi W - \sqrt{2} \beta v_k(Y_{t_k}^\gamma) W}\right\rVert_{L^2} &= \eta^{1/2} \left\lVert{\Xi - \sqrt{2} \beta v_k(Y_{t_k}^\gamma)}\right\rVert_{L^2} \notag \\ \label{eq:lemma-5465-second-triangle} &\leqslant\eta^{1/2} \left\lVert{\Xi - \sqrt{2} \beta v_k( \Sigma)}\right\rVert_{L^2} + \eta^{1/2} \left\lVert{\sqrt{2} \beta v_k(\Sigma) - \sqrt{2} \beta v_k(Y_{t_k}^\gamma)}\right\rVert_{L^2}. \end{align}\tag{66}\] Furthermore, and , along with the fact that \(1+e^{5\beta^2} \leqslant 2e^{5\beta^2}\), show that the first term of satisfies \[\begin{align} \left\lVert{\Xi - \sqrt{2} \beta v_k(\Sigma)}\right\rVert_{L^2} &\leqslant\left\lVert{E_{\mathcal{D}_n}[Q_k^2]^{1/2} - \sqrt{2} \beta D_k^{-1}}\right\rVert_2 + \sqrt{2} \beta \left\lVert{D_k^{-1} - \operatorname{diag}(v_k(\sigma_{k}))}\right\rVert_2 \\ &\leqslant M_2 \beta n^{-\delta/2} + 2 \sqrt{2} \beta d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_k),\operatorname{dist}(Y_{t_k}^\gamma)\right) + 2 \sqrt{2} \beta e^{5 \beta^2} \gamma. \end{align}\] Moreover, since \(v\) is \(2\)-Lipschitz in space by , we get for the second term \[\left\lVert{\sqrt{2} \beta v_k(\Sigma) - \sqrt{2} \beta v_{k}(Y^{\gamma}_{t_{k}})}\right\rVert_{L^2} \leqslant M_3 \beta d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_k),\operatorname{dist}(Y^{\gamma}_{t_k})\right).\] Thus, combining the last two inequalities with , \[\left\lVert{\Xi W - \sqrt{2} \beta v_{k}(Y^{\gamma}_{t_k})W}\right\rVert_{L^2} \leqslant\eta^{1/2}\left[M_2 \beta n^{-\delta/2} + M_4 \beta d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_k),\operatorname{dist}(Y^{\gamma}_{t_k})\right) + M_5 \beta e^{5\beta^2}\gamma\right],\] which then combines with and to yield \[\left\lVert{\Xi W - Y^{\gamma}_{t_{k+1}} + Y^{\gamma}_{t_k}}\right\rVert_{L^2} \leqslant M_1 \beta^3 \eta + \eta^{1/2} \beta \left[ M_2 n^{-\delta/2} + M_4 d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_k),\operatorname{dist}(Y^{\gamma}_{t_k})\right) + M_5 e^{5\beta^2}\gamma \right].\] Finally, using the arithmetic-geometric mean inequality, \[\left\lVert{\Xi W - Y^{\gamma}_{t_{k+1}} + Y^{\gamma}_{t_{k}}}\right\rVert_{L^2}^2 \leqslant\eta \beta^2 \cdot \left[ M_6 \beta^4 \eta + M_7 n^{-\delta} + M_8 d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_k),\operatorname{dist}(Y^{\gamma}_{t_k})\right)^2 + M_9 e^{10 \beta^2} \gamma^2 \right].\] This with shows the desired conclusion with \(C_1 = M_6\), \(C_2 = M_7\), \(C_3 = M_8\), \(C_4 = M_9\). ◻

5.3 Concentration for the empirical distribution↩︎

Finally, we will estimate the difference between the empirical distribution of \(\sigma_{k+1}\) and its expectation using concentration of measure. The estimate that we obtain below is derived from concentration of measure alone, and is certainly not optimal. Obtaining optimal estimates for the Wasserstein distances for empirical distributions is a challenging problem, and readers well-versed in probability theory may try to obtain better estimates using more refined tools. Here it will be convenient to estimate the \(L^1\)-Wasserstein distance first, rather than directly attacking the \(L^2\)-Wasserstein distance. Recall that the \(L^1\)-Wasserstein distance is \[d_{W,1}(\mu,\nu) = \inf \big\{ \left\lVert{X - Y}\right\rVert_{L^1}: X \sim \mu, Y \sim \nu \big\} = \sup_{f \in \mathsf{Lip}_1(\mathbb{R}), f(0) = 0}\left|\int f\,d(\mu - \nu)\right|,\] where the last equality is by Monge-Kantorovich-Rubinstein duality; see .

We also use the following construction of a dense family of \(1\)-Lipschitz functions which is well known and easy to verify.

Lemma 27. Fix \(\ell,m \in \mathbb{N}\). Let \(\mathsf{Lip}_1([-2^\ell,2^\ell])\) be the set of \(1\)-Lipschitz functions on \([-2^\ell,2^\ell]\) that vanish at \(0\). Then there exist \(2^{2^{\ell+ m+1}}\) functions in \(\mathsf{Lip}_1([-2^\ell,2^\ell])\) that are \(1/2^m\)-dense with respect to the uniform norm. Specifically take the Lipschitz functions that vanish at zero and are linear with slopes \(\pm 1\) on each interval of the form \([(j-1)/2^m,j/2^m]\), for \(j = -2^{\ell+m}+1,\dots, 2^{\ell+m}\). The number of such functions is \(2^{2^{\ell+m+1}}\) since they are uniquely described by the choice of slope \(\pm 1\) on each of the \(2\cdot 2^{\ell+m}\) intervals.

This allows us to estimate the \(L^1\) Wasserstein distance using concentration if the measures are supported in a bounded interval. For this purpose, we estimate probabilities of the maximum coordinate of \(\sigma_k\) being large.

Lemma 28. Let \(\alpha > 0\). We have \[\max_j |(Q_k Z_k)_j| \leqslant n^\alpha\] with probability at least \[1 - C_1 n \exp(-C_2 n^{2\alpha - 4 \delta}).\]

Proof. Recall from that \(\left\lVert{Q_k^2}\right\rVert \leqslant M_1 n^{4\delta}\). Therefore, \((Q_k Z_k)_j\) is a normal random variable with mean zero and variance bounded by \(M_1 n^{4\delta}\). Thus, \[\mathop{{}\mathbb{P}}\left(\left|\left(Q_k Z_k\right)_j\right| \geqslant n^\alpha\right) \leqslant M_2 \exp(-M_3 n^{2 \alpha-4\delta}).\] The asserted estimate therefore follows from a union bound over the coordinates. ◻

Lemma 29. Let \(Z\) be a real Gaussian random variable with mean \(a\) and variance \(b\). Then \[\mathbb{E}\left[\mathbb{1}_{|Z - a| \geqslant c} |Z - a|^2\right] \leqslant\frac{2b}{\sqrt{2 \pi}}(b^{-1/2}c + b^{1/2}c^{-1}) e^{-b^{-1}c^2/2}.\]

Proof. Without loss of generality, assume that \(a = 0\). Let \(\tilde{Z} = b^{-1/2} Z\) which is a standard normal. Then \[\begin{align} \mathbb{E}\left[\mathbb{1}_{|Z| \geqslant c} |Z|^2\right] &= b \mathbb{E}\left[\mathbb{1}_{|\tilde{Z}| \geqslant b^{-1/2}c} |\tilde{Z}|^2\right] \\ &= \frac{2b}{\sqrt{2 \pi}} \int_{b^{-1/2}c}^\infty z^2 e^{-z^2/2}\,dz \\ &= \frac{2b}{\sqrt{2 \pi}} \left(\left[-ze^{-z^2/2}\right]_{b^{-1/2}c}^\infty + \int_{b^{-1/2}c}^\infty e^{-z^2/2}\,dz\right) \\ &\leqslant\frac{2b}{\sqrt{2 \pi}} \left[b^{-1/2}c e^{-b^{-1}c^2/2} + b^{1/2} c^{-1} \int_{b^{-1/2}c}^\infty ze^{-z^2/2}\,dz\right] \\ &= \frac{2b}{\sqrt{2 \pi}}(b^{-1/2}c + b^{1/2}c^{-1}) e^{-b^{-1}c^2/2}. \qedhere \end{align}\] ◻

Lemma 30. Let \(\delta \in (0,1/19]\) and let \(\alpha \in (2\delta,1/7-4\delta/7)\). Assume that the parameters of satisfy \(q^* \leqslant 1\) with \(\eta\) and \(K\) held constant. Assume that \(\sigma_k\) has been chosen (and so is deterministic for the purposes of this lemma) and assume that \[\max_j |\sigma_{k,j}| \leqslant k \eta^{1/2} n^\alpha.\] Then, conditioned on \(A, Z_0, \dots, Z_{k-1}\), with probability at least \[1 - C_1 \eta^{-1} n \exp(-C_2 n^{2 \alpha-4\delta}) - C_3 \exp \left( -C_4 \beta^2n^{1-4\delta-4 \alpha} + C_5\beta^{-1}\eta^{-1}n^{3\alpha} \right).\] in the Gaussian vector \(Z_k\), we have \[\max_j |\sigma_{k+1,j}| \leqslant(k+1) \eta^{1/2} n^\alpha\] and \[d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_{k+1}), \mathop{{}\mathbb{E}}\left[\mathop{\mathrm{emp}}(\sigma_{k+1}) | A, \sigma_1, \dots, \sigma_k\right]\right) \leqslant C_7 n^{-\alpha/2}.\]

Remark 29. For concreteness, one may take \(\alpha = 1/9\) and then \(n^{2\alpha - 4 \delta} = n^{2/9-4\delta} = n^{1-4 \delta - 7 \alpha}\).

Proof. By , \(\max_j |\sigma_{k+1,j} - \sigma_{k,j}| \leqslant\eta^{1/2} n^\alpha\) with probability at least \(1 - M_1 n \exp(-M_2 n^{2 \alpha-4\delta})\) in the Gaussian vector \(Z_k\) conditioned on \(A, Z_0, \dots, Z_{k-1}\). This in particular, together with \(K\eta \leqslant q^* \leqslant 1\), implies that \[\label{eq:max-coordinate-event} \max_j |\sigma_{k+1,j}| \leqslant(k+1) \eta^{1/2} n^\alpha \leqslant K \eta^{1/2} n^{\alpha} \leqslant\eta^{-1/2} n^\alpha\tag{67}\] with probability at least \(1 - M_1 Kn \exp(-M_2 n^{2 \alpha-4\delta})\). Fix \(\ell\) such that \(2^{\ell-1} \leqslant 2 \eta^{-1/2} n^\alpha \leqslant 2^\ell\). We truncate the expected empirical distribution as follows: Let \(\operatorname{proj}_{[-2^\ell,2^\ell]}\) be the truncation map \[\operatorname{proj}_{[-2^\ell,2^\ell]}(s) = \begin{cases} -2^\ell, & s \in (-\infty, -2^\ell] \\ s, & s \in [-2^\ell, 2^\ell] \\ 2^\ell, & s \in [2^\ell,\infty). \end{cases}\] and let \(\tau_0\) be the pushforward of \(\tau = \mathop{{}\mathbb{E}}\left[\mathop{\mathrm{emp}}(\sigma_{k+1}) | A, \sigma_1, \dots, \sigma_k\right]\) under \(\operatorname{proj}_{[-2^\ell,2^\ell]}\). Since \(|\sigma_{k,j}| \leqslant k\eta^{1/2} n^{\alpha} \leqslant\eta^{-1/2} n^\alpha\) by assumption, we have \[|\sigma_{k+1,j} - \sigma_{k,j}| \leqslant\eta^{-1/2} n^\alpha \implies |\sigma_{k+1,j}| \leqslant 2 \eta^{-1/2} n^\alpha \leqslant 2^\ell.\] and thus \[|\sigma_{k+1,j} - \mathop{\mathrm{proj}}_{[-2^\ell,2^\ell]}(\sigma_{k+1,j})| \leqslant\mathbb{1}_{|\sigma_{k+1,j} - \sigma_{k,j}| \geqslant\eta^{-1/2} n^\alpha} |\sigma_{k+1,j} - \sigma_{k,j}|.\] Note that \(\sigma_{k+1,j}\) is Gaussian with mean \(\sigma_{k,j}\) and variance bounded by \(M \eta n^{4 \delta}\) by . Therefore, by Lemma 29, \[\begin{align} \mathbb{E}[\mathbb{1}_{|\sigma_{k+1,j} - \sigma_{k,j}| \geqslant\eta^{-1/2} n^\alpha} (\sigma_{k+1,j} - \sigma_{k,j})^2] &\leqslant\frac{2M \eta n^{4 \delta}}{\sqrt{2\pi}}(M^{-1/2} \eta^{-1} n^{\alpha-\delta} + M^{1/2} \eta n^{\delta-\alpha}) e^{-M^{-1} \eta^{-1} n^{-2\delta} \eta^{-1} n^{2\alpha}} \\ &\leqslant 2 M^{3/2} n^{\alpha+3\delta}e^{-M^{-1} \eta^{-2} n^{2\alpha-2\delta}} . \end{align}\] Therefore, \[\begin{align} d_{W,2}(\tau,\tau_0)^2 &\leqslant\int_{|x| \geqslant\eta^{-1/2} n^\alpha} \left(x - \operatorname{proj}_{[-2^\ell,2^\ell]}(x)\right)^2\,d\tau(x) \notag\\ &\leqslant\frac{1}{n} \sum_{j=1}^n \mathbb{E}\left[|\sigma_{k+1,j} - \mathop{\mathrm{proj}}_{[-2^\ell,2^\ell]}(\sigma_{k+1,j})|^2\right] \notag\\ &\leqslant\frac{1}{n} \sum_{j=1}^n \mathbb{E}\left[ \mathbb{1}_{|\sigma_{k+1,j} - \sigma_{k,j}| \geqslant\eta^{-1/2} n^\alpha} (\sigma_{k+1,j} - \sigma_{k,j})^2\right] \notag\\ &\leqslant 2M^{3/2} n^{\alpha+3\delta} e^{-M^{-1} \eta^{-2} n^{2\alpha-2\delta}} \notag\\ &\leqslant M' n^{-\alpha},\label{eq:wass-tau-tau0} \end{align}\tag{68}\] where we choose \(M'\) such that \(M' \geqslant 2M^{3/2}n^{2\alpha+3\delta}e^{-M^{-1}\eta^{-2}n^{2\alpha-2\delta}}\) by choosing \(M' = 2M^{3/2}n_0e^{-M^{-1}n_0^{1/9}}\) for \(n_0\) solving \(0 = \frac{d}{dn_0}\left(n_0e^{-M^{-1}n_0^{1/9}}\right)\).

Now we proceed to estimate \(d_{W,1}(\mathop{\mathrm{emp}}(\sigma_{k+1}), \tau_0)\). Take \(f \in \mathsf{Lip}_1([-2^\ell,2^\ell])\), and extend \(f\) to an \(1\)-Lipschitz function on \(\mathbb{R}\) by setting \(f\) to be constant on \((-\infty,-2^\ell]\) and \([2^\ell,\infty)\), so that \(f \circ \mathop{\mathrm{proj}}_{[-2^\ell,2^\ell]} = f\). Consider \[F_f(Z_k) = \frac{1}{n} \sum_{j=1}^n f\left(\sigma_{k,j} + \eta^{1/2} (Q_k Z_k)_j \right)\] where \(Z_k\) is a standard Gaussian vector and recall from the proof of that the distribution of \(\sigma_{k,j}+\eta^{1/2}(Q_kZ_k)\) over the randomness of \(j\) and \(Z_k\) is equal to \(\mathop{{}\mathbb{E}}\left[\mathop{\mathrm{emp}}(\sigma_{k+1}) | A, \sigma_1, \dots, \sigma_k\right]\), so \(\mathop{{}\mathbb{E}}_{Z_k} [F_f(Z_k)] = \int f\,d\tau = \int (f \circ \mathop{\mathrm{proj}}_{[-2^\ell,2^\ell]})d\tau = \int f\,d\tau_0\). Moreover, \(F_f(Z_k)\) is an \(\eta^{1/2} \left\lVert{Q_k}\right\rVert /\sqrt{n}\)-Lipschitz function of \(Z_k\), since \[\begin{align} \frac{1}{n} \sum_{j=1}^n \left|f\left(\sigma_{k,j} + \eta^{1/2} (Q_k Z_k)_j \right) - f\left(\sigma_{k,j} + \eta^{1/2} (Q_k Z_k')_j \right)\right| &\leqslant\frac{1}{n} \eta^{1/2}\sum_{j=1}^n \left| (Q_k (Z_k-Z_k'))_j \right| \\&= \frac{\eta^{1/2}}{n} |Q_k (Z_k-Z_k')|_1 \leqslant\frac{\eta^{1/2}}{\sqrt{n}} \left\lVert{Q_k}\right\rVert\,|Z_k-Z_k'|_2 \,. \end{align}\] Now we apply the concentration estimate to the Lipschitz function \(F_f(Z_k)\) above, while noting that \(\left\lVert{Q_k}\right\rVert n^{-1/2} \leqslant M_5 n^{2\delta-1/2}\) by . By concentration of Lipschitz functions of Gaussians (i.e. with \(n=1\)), we have \[\left| F_f(Z_k) - \mathop{{}\mathbb{E}}_{Z_k} \left[F_f(Z_k)\right]\right| \leqslant\frac{1}{2^m}\] with probability at least \[1 - M_6 \exp \left( -M_7 n^{1-4\delta}2^{-2m}\eta^{-1} \right).\]

Choose \(m\) so that \(2^{-m} \leqslant\beta \eta^{1/2} n^{-2 \alpha} \leqslant 2^{-m+1}\). By and , we can estimate the \(L^1\) Wasserstein distance between measures supported in \([-2^\ell,2^\ell]\) up to an error of size \(1/2^m\) by testing only the \(1\)-Lipschitz functions \(f\) in some set \(\mathcal{F}\) satisfying \(|\mathcal{F}| = 2^{2^{\ell+m+1}}\). We use a union bound for the probability of error for each of these Lipschitz functions \(f\). This yields \[\begin{align} d_{W,1}(\mathop{\mathrm{emp}}(\sigma_{k+1}), \tau_0) &= \sup_{\left\lVert{f}\right\rVert_{\mathrm{Lip}}\leqslant 1}\left|\int f(d\mathop{\mathrm{emp}}(\sigma_{k+1}) - d\tau_0)\right| \\&\leqslant \frac{1}{2^m} + \sup_{f \in \mathcal{F}}\left|\int f(d\mathop{\mathrm{emp}}(\sigma_{k+1}) - d\tau_0)\right| \\&= \frac{1}{2^m} + \sup_{f \in \mathcal{F}}\left| F_f(Z_k) - \mathop{{}\mathbb{E}}_{Z_k} \left[F_f(Z_k)\right]\right| \\&\leqslant\frac{2}{2^m} \leqslant 2 \eta^{1/2} n^{-2\alpha} \end{align}\] with probability at least \[\label{eq:prob-lipschitz-function-union-bound} 1 - M_6 2^{2^{\ell+ m+1}} \exp \left( -M_7 n^{1-4\delta} 2^{-2m} \eta^{-1} \right),\tag{69}\] provided that \(\max_j |\sigma_{k+1,j}| \leqslant\eta^{-1/2} n^\alpha \leqslant 2^\ell\) which by happens with probability at least \[\label{eq:max-coordinate-event-prob} 1 - M_1 Kn \exp(-M_2 n^{2 \alpha-4\delta}).\tag{70}\] Hence, in this high-probability event, by Hölder’s inequality, \[d_{W,2}(\mathop{\mathrm{emp}}(\sigma_{k+1}),\tau_0) \leqslant 2^{\ell/2} d_{W,1}(\mathop{\mathrm{emp}}(\sigma_{k+1}),\tau_0)^{1/2} \leqslant M_8 \eta^{-1/4} n^{\alpha/2} \eta^{1/4} n^{-\alpha} = M_8 n^{-\alpha/2},\] which combines with by triangle inequality to obtain \[\label{eq:event-emp-close-to-exp} d_{W,2}(\mathop{\mathrm{emp}}(\sigma_{k+1}),\tau) \leqslant M_9 n^{-\alpha/2}.\tag{71}\]

It remains to evaluate the probability coming from concentration. By construction, \(2^{\ell} \leqslant 4 \eta^{-1/2} n^\alpha\) and \(2^{m} \leqslant 2\beta^{-1} \eta^{-1/2} n^{2 \alpha}\). Thus, \[2^{2^{\ell+m+1}} \leqslant M_{9} \exp\left(M_{10}\beta^{-1}\eta^{-1}n^{3\alpha}\right).\] Hence, \[\begin{align} M_6 2^{2^{\ell+ m+1}}&\exp \left( -M_7 n^{1-4\delta} 2^{-2m} \eta^{-1} \right) \\ &\leqslant M_6 M_{9} \exp\bigg(M_{10}\beta^{-1}\eta^{-1}n^{3\alpha} \bigg) \exp \left( -M_7 n^{1-4\delta} 2^{-2m}\eta^{-1} \right) \\ &\leqslant M_{11} \exp \left( -M_{12} \beta^2n^{1-4\delta-4 \alpha} +M_{10}\beta^{-1}\eta^{-1}n^{3\alpha} \right), \end{align}\] which, when combined with and , yields the final probability of the bounds and . ◻

5.4 Conclusion of the convergence argument↩︎

We are ready to finish proving . As preparation, we record a small computation for the inductive step.

Corollary 7. Let \(\delta \in (0,1/19]\), and let \(\alpha \in (2\delta,1/7-4\delta/7)\). Assume that the parameters \(\eta\) and \(n\) satisfy \(n^{1/90} \eta \geqslant 1\). Let \(A\) satisfy the conclusions of . Assume that \(\max_j |\sigma_{k,j}| \leqslant k \eta^{1/2} n^\alpha\), that \(2 d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k), \operatorname{dist}(Y_{t_k}^\gamma)) + (1 + e^{5 \beta^2 t_k}) \gamma \leqslant\frac{1}{2 \sqrt{2} \beta}\), and moreover that \[d_{W,2}(\mathop{{}\mathbb{E}}\left[\mathop{\mathrm{emp}}(\sigma_{k+1}) | A, \sigma_1, \dots, \sigma_k\right],\operatorname{dist}(Y_{t_{k+1}}^\gamma)) \leqslant 1.\] Then, conditioned on \(A, Z_0, \dots, Z_{k-1}\), with probability in the Gaussian vector \(Z_k\) of at least \[1 - C_1 n\eta^{-1}\exp \left( -C_2 n^{2/9-4\delta} \right),\] we have both \[\begin{gather} d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_{k+1}), \operatorname{dist}(Y^{\gamma}_{t_{k+1}})\right)^2 - d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_k), \operatorname{dist}(Y^{\gamma}_{t_k})\right)^2 \\ \leqslant\eta \cdot \left[ C_4 \beta^2 d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_k),\operatorname{dist}(Y^{\gamma}_{t_k})\right)^2 + C_5 \beta^6 \eta + C_6 \beta^2 n^{-\delta} + C_7 \beta^2 e^{10\beta^2} \gamma^2 + C_8 \eta^{-1} n^{-1/18} \right]. \end{gather}\] and \[\max_j |\sigma_{k+1,j}| \leqslant(k+1)\eta^{1/2} n^{\alpha}.\]

Proof. By triangle inequality and making the conditioning in \(\mathop{{}\mathbb{E}}\left[\mathop{\mathrm{emp}}(\sigma_{k+1}) | A, \sigma_1, \dots, \sigma_k\right]\) notationally implicit for brevity, \[\begin{align} d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_{k+1}), \operatorname{dist}(Y^{\gamma}_{t_{k+1}})\right)^2 &\leqslant\left( d_{W,2}\left(\mathop{{}\mathbb{E}}\mathop{\mathrm{emp}}(\sigma_{k+1}), \operatorname{dist}(Y^{\gamma}_{t_{k+1}})\right) + d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_{k+1}), \mathop{{}\mathbb{E}}\mathop{\mathrm{emp}}(\sigma_{k+1})\right)\right)^2 \\ &= d_{W,2}\left(\mathop{{}\mathbb{E}}\mathop{\mathrm{emp}}(\sigma_{k+1}), \operatorname{dist}(Y^{\gamma}_{t_{k+1}})\right)^2 \\ & \quad + 2 d_{W,2}\left(\mathop{{}\mathbb{E}}\mathop{\mathrm{emp}}(\sigma_{k+1}), \operatorname{dist}(Y^{\gamma}_{t_{k+1}})\right) d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_{k+1}), \mathop{{}\mathbb{E}}\mathop{\mathrm{emp}}(\sigma_{k+1})\right) \\ &\quad + d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_{k+1}), \mathop{{}\mathbb{E}}\mathop{\mathrm{emp}}(\sigma_{k+1})\right)^2. \end{align}\] Assuming that \(d_{W,2}(\mathop{{}\mathbb{E}}\mathop{\mathrm{emp}}(\sigma_{k+1}),\operatorname{dist}(Y^{\gamma}_{t_{k+1}})) \leqslant 1\), we can bound \[\begin{gather} 2 d_{W,2}\left(\mathop{{}\mathbb{E}}\mathop{\mathrm{emp}}(\sigma_{k+1}), \operatorname{dist}(Y^{\gamma}_{t_{k+1}})\right) d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_{k+1}), \mathop{{}\mathbb{E}}\mathop{\mathrm{emp}}(\sigma_{k+1})\right) + d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_{k+1}), \mathop{{}\mathbb{E}}\mathop{\mathrm{emp}}(\sigma_{k+1})\right)^2 \\ \leqslant M_1 n^{-\alpha/2} + M_2 n^{-\alpha}, \end{gather}\] using with \(\alpha = 1/9\). Combining this with using triangle inequality, we obtain the asserted estimate with probability \[1 - M_3\eta^{-1}n \exp \left( -M_4 n^{2/9-4\delta}\right) - M_5\exp\left(-M_6 n^{5/9-4\delta} + M_7 \beta^{-1} \eta^{-1} n^{3/9} \right).\] By our assumptions \(\beta \geqslant 1\) and \(\eta n^{1/90} \geqslant 1\), and hence \(\beta^{-1} \eta^{-1} n^{3/9} \leqslant n^{31/90}\). Since \(-M_6 n^{5/9-4\delta} + M_7 n^{31/90} - M_8 \leqslant-M_9n^{2/9-4\delta}\), we can absorb the \(M_5\exp\left(-M_6 n^{5/9-4\delta} + M_7 n^{31/90} \right)\) term into the \(M_3\eta^{-1}n\exp(-M_4 n^{2/9 - 4 \delta})\) term. ◻

Now we are ready to finish the argument.

Proof of . We induct over \(k = 0, \dots, k'-1\). Write \(\zeta_{n,\beta,\eta,\gamma} := M_2 \beta^4 \eta + M_3 n^{-\delta} + M_4 e^{10 \beta^2} \gamma^2 + M_5 \beta^{-2}\eta^{-1} n^{-1/18}\) and \(\Delta_k := d_{W,2}(\mathop{\mathrm{emp}}(\sigma_{k}), \operatorname{dist}(Y_{t_{k}}^\gamma))\). The inductive hypothesis will be that with probability at least \(1 - M_6 n(k-1)\eta^{-1} \exp \left( -M_7 n^{2/9-4\delta} \right)\), \[\label{eq:inductive-delta} \Delta_k^2 \leqslant\sum_{j=0}^{k-1} (1 + M_1 \eta \beta^2)^j \eta\beta^2 \cdot \zeta_{n,\beta,\eta,\gamma}\tag{72}\] and that the hypotheses of will hold for this value of \(k\) with \(\alpha = 1/9\).

Note that \(\Delta_0 = 0\) since \(\sigma_0 = 0\) and \(Y_0 = 0\), and furthermore the hypotheses of are vacuous at \(k=0\).

For the inductive step, combining the probability from the inductive hypothesis with the probability from puts us in an event with probability at least \(1 - M_6 nk\eta^{-1} \exp \left( -M_7 n^{2/9-4\delta} \right)\), so we can assume both bounds of hold. The first bound says that \[\Delta_{k+1}^2 - \Delta_k^2 \leqslant\eta \cdot \left[ M_1 \beta^2 \Delta_k^2 + M_2 \beta^6 \eta + M_3 \beta^2 n^{-\delta} + M_4 \beta^2 e^{10 \beta^2} \gamma^2 + M_5 \eta^{-1} n^{-1/18} \right],\] and hence by , \[\Delta_{k+1}^2 \leqslant(1 + M_1 \eta \beta^2) \Delta_k^2 + \eta\beta^2 \zeta_{n,\beta,\eta,\gamma} \leqslant\sum_{j=0}^{k} (1 + M_1 \eta \beta^2)^j \eta\beta^2 \cdot \zeta_{n,\beta,\eta,\gamma}\,,\] as required by for the next step of the induction.

The second bound of gives the first hypothesis of for the next step.

For the second hypothesis, note that \[\begin{align} \Delta_{k+1}^2 &\leqslant\sum_{j=0}^{k} (1 + M_1 \eta \beta^2)^j \eta\beta^2 \cdot \zeta_{n,\beta,\eta,\gamma} \notag\\ &= \frac{(1 + M_1 \eta \beta^2)^{k+1} - 1}{M_1 \eta \beta^2} \eta\beta^2 \cdot \zeta_{n,\beta,\eta,\gamma} \notag\\ \label{eq:SDE-inductive-hypothesis-exponential} &\leqslant M_1^{-1} \left(\exp( M_1 \beta^2 (k+1) \eta) - 1\right) \zeta_{n,\beta,\eta,\gamma}\,. \end{align}\tag{73}\] By the assumption in the theorem statement, an application of the arithmetic-geometric mean inequality, upper-bounding \((k+1) \eta\) by 1, and upper-bounding \((1 + e^{5\beta^2})^2\gamma^2\) by \(C_4^{-1}C_7 (\exp( C_4 \beta^2) - 1)e^{10\beta^2}\gamma^2\), \[\label{eq:32induction32hypothesis32for32all} 2\left( M_1^{-1} \left(\exp( M_1 \beta^2 (k+1) \eta) - 1\right) \zeta_{n,\beta,\eta,\gamma} \right)^{1/2} + (1 + e^{5\beta^2}) \gamma \leqslant\frac{1}{2 \sqrt{2} \beta}\,.\tag{74}\] The last two inequalities combine to give \[\label{eq:final-wasserstein-2-bound-loose} 2d_{W,2}\left(\mathop{\mathrm{emp}}(\sigma_{k+1}), \operatorname{dist}(Y_{t_{k+1}}^\gamma)\right) + (1 + e^{5 \beta^2}) \gamma \leqslant\frac{1}{2 \sqrt{2} \beta},\tag{75}\] which implies the second hypothesis of for the next step.

Finally, observe that the same induction that leads to also works when \(\Delta_k\) is taken instead to be \(d_{W,2}(\mathop{{}\mathbb{E}}\left[\mathop{\mathrm{emp}}(\sigma_{k+1}) | A, \sigma_1, \dots, \sigma_k\right], \operatorname{dist}(Y_{t_{k+1}}^\gamma))\), with the only difference being fewer applications of the triangle inequality and consequently fewer terms to upper-bound (alternatively, we may apply with ), so, invoking \(\beta \geqslant 1\), \[d_{W,2}\left(\mathop{{}\mathbb{E}}\left[\mathop{\mathrm{emp}}(\sigma_{k+1}) | A, \sigma_1, \dots, \sigma_k\right], \operatorname{dist}(Y^{\gamma}_{t_{k+1}})\right) \leqslant\frac{1}{2 \sqrt{2} \beta} \leqslant 1,\] which is the final assumption needed for , finishing the proof of the inductive step.

By , we obtain that for all \(k \in \{0,\dots,k'\}\), \[\begin{align} \Delta_k^2 &\leqslant M_1^{-1} \left(\exp( M_1 \beta^2 k' \eta) - 1\right) \zeta_{n,\beta,\eta,\gamma} \end{align}\] with probability at least \(1 - M_6 nk\eta^{-1} \exp \left( -M_7 n^{2/9-4\delta} \right)\), as required by the first conclusion of the theorem statement. The second conclusion is immediately implied by . ◻

6 Energy Analysis↩︎

In this section we finally obtain guarantees on the quality of the output of by estimating the changes in our modified objective function at each step using Taylor expansion. We will estimate the terms in the Taylor expansion using the bounds on the spectral properties of the Hessian (\(\S\) 4), the convergence of the iterates in empirical distribution (\(\S\) 5), and concentration inequalities. The Taylor expansion is taken to the second degree in space (with third-order remainder) due to the Brownian-motion-like nature of the process and the need to use the spectral properties of the Hessian.

Theorem 30. Assume . Let the Gaussian matrix \(A\) be chosen from the high probability event in . Consider the objective function \[\mathop{\mathrm{obj}}(t,\sigma) = \beta \angles{\sigma, A \sigma} - \sum_{i=1}^n\tilde{\Lambda}_{\gamma}(t,\sigma_i) - \beta^2 n \int_t^1 s F_\mu(s)\,ds\,,\] and fix \(\beta = \frac{10}{\varepsilon}\), \(\gamma = e^{-C_1 \beta^2}\) and \(\eta = \gamma^{8}\) and \(n \geqslant\eta^{-90}\) for some sufficiently large constant \(C\). Let \(z \in \{-1,1\}^n\) be the final solution output by  on input \(A\) with access to \(\mu_\beta\) and \(\Phi\) (and its derivatives). Then, conditioned on \(A\), with probability at least \(1 - C_1 \exp(-C_2 n^{1/3 - 4\delta} + C_3 (\log \eta^{-1})^2)\), \[\frac{1}{n}H(z) \geqslant\frac{\mathcal{P}_{\beta}}{\beta} - \frac{\varepsilon}{5} - O(\varepsilon^2) - O(n^{-\alpha})\, ,\] where \(\alpha = \min\left(\frac{\delta}{4}, \frac{1}{24}\right)\).

Here we use the following notation: \[\begin{align} t_k &= k \eta \\ Q_k &= Q(t_k,\sigma_k) \\ \Delta \sigma_k &= \sigma_{k+1} - \sigma_k = \eta^{1/2} Q_k Z_k. \end{align}\] and \[v(t,\sigma) = \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,\sigma)}.\] For a function \(f(t,y)\) (for instance \(\partial_y \tilde{\Lambda}_{\gamma}\) or \(\partial_{y,y} \tilde{\Lambda}_{\gamma}\)), we write \(f(t,\cdot)\) for the coordinate-wise application of \(f\) to a vector in space, i.e. \[f(t,\sigma_k) = (f(t,\sigma_{k,j}))_{j=1}^n\,.\]

6.1 Overview of Taylor expansion↩︎

We will express the total change in the objective function as a telescoping sum of the changes at each step. Thus, most of the section will be devoted to estimating the increment \(\mathop{\mathrm{obj}}(t_{k+1},\sigma_{k+1}) - \mathop{\mathrm{obj}}(t_k,\sigma_k)\). Here we will describe the terms that arise from the Taylor expansion, then in the following sections we will estimate each one of them, and in the conclusion we will sum up the steps and deduce that the value of \(\angles{\sigma, A \sigma}\) is close to the energy functional that describes the theoretical maximum, which is computed in [9].

To evaluate the increment, let us first separate the space update and the time update: \[\operatorname{obj}(t_{k+1}, \sigma_{k+1}) - \operatorname{obj}(t_k,\sigma_k) = [\operatorname{obj}(t_k,\sigma_{k+1}) - \operatorname{obj}(t_k,\sigma_k)] + [\operatorname{obj}(t_{k+1},\sigma_{k+1}) - \operatorname{obj}(t_k,\sigma_{k+1})]\] The time update is further broken down as follows, using the integral form of Taylor expansion: \[\begin{align} & \quad \operatorname{obj}(t_{k+1},\sigma_{k+1}) - \operatorname{obj}(t_k, \sigma_{k+1}) \\ &= -\sum_{j=1}^n \int_{t_k}^{t_{k+1}} \partial_t \tilde{\Lambda}_{\gamma}(t,\sigma_{k+1,j})\,dt + n \beta^2 \int_{t_k}^{t_{k+1}} F_\mu(t) t \,dt \\ &=_{\text{\prettyref{prop:primal-pde-lipschitz}}} -\beta^2 \sum_{j=1}^n \int_{t_k}^{t_{k+1}} \left( \frac{1}{\partial_{y,y} \tilde{\Lambda}_{\gamma}(t,\sigma_{k+1,j})} - \gamma + F_{\mu}(t)\left((\sigma_{k+1})_j - \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,(\sigma_{k+1,j}))\right)^2 \right) \,dt \\ &\qquad \qquad \qquad \qquad \qquad + n \beta^2 \int_{t_k}^{t_{k+1}} F_\mu(t) t \,dt \\ &= -\beta^2 \sum_{j=1}^n \int_{t_k}^{t_{k+1}} v(t,\sigma_{k+1,j})\,dt + n \beta^2 \gamma \eta - \beta^2 \int_{t_k}^{t_{k+1}} F_\mu(t) \left( \sum_{j=1}^n (\mathop{\mathrm{id}}- \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,\cdot))(\sigma_{k+1,j})^2 - nt \right)\,dt \end{align}\] The space update can be broken down as follows: \[\begin{align} & \quad \operatorname{obj}(t_k, \sigma_{k+1}) - \operatorname{obj}(t_k,\sigma_k) \\ &= \beta \left( \angles{\sigma_{k+1}, A_{\mathop{\mathrm{sym}}} \sigma_{k+1}} - \angles{\sigma_k, A_{\mathop{\mathrm{sym}}} \sigma_k} \right) - \sum_{j=1}^n \left( \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k+1,j}) - \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}) \right). \end{align}\] Now we write \[\angles{\sigma_{k+1}, A_{\mathop{\mathrm{sym}}} \sigma_{k+1}} - \angles{\sigma_k, A_{\mathop{\mathrm{sym}}} \sigma_k} = 2 \angles{\Delta \sigma_k, A_{\mathop{\mathrm{sym}}} \sigma_k} + \angles{\Delta \sigma_k, A_{\mathop{\mathrm{sym}}} \Delta \sigma_k}.\] Furthermore, by the Taylor expansion with Lagrange remainder, there is some \(\xi_{k,j}\) between \(\sigma_{k,j}\) and \(\sigma_{k+1,j}\) such that \[\tilde{\Lambda}_{\gamma}(t_k,\sigma_{k+1,j}) - \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}) = \partial_y \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}) \Delta \sigma_{k,j} + \frac{1}{2} \partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}) (\Delta \sigma_{k,j})^2 + \frac{1}{6} \partial_{y,y,y} \tilde{\Lambda}_{\gamma}(t_k, \xi_{k,j}) (\Delta \sigma_{k,j})^3.\]

Finally, we combine all these terms. For reasons that will be apparent later, we will add and subtract \(\beta^2 \eta \sum_j v(t_k,\sigma_{k,j})\). Thus, \[\begin{align} \operatorname{obj}(t_{k+1}, \sigma_{k+1}) &- \operatorname{obj}(t_k,\sigma_k) \\ =&\,\angles{2 \beta A_{\mathop{\mathrm{sym}}} \sigma_k - \partial_y \tilde{\Lambda}_{\gamma}(t_k,\sigma_k), \Delta \sigma_k} \tag{76} \\ &+ \frac{1}{2} \angles{\Delta \sigma_k, (2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}))) \Delta \sigma_k} - \beta^2 \eta \sum_{j=1}^n v(t_k,\sigma_{k,j}) \tag{77} \\ &+ \frac{1}{6} \sum_{j=1}^n \partial_{y,y,y} \tilde{\Lambda}_{\gamma}(t_k, \xi_{k,j}) (\Delta \sigma_{k,j})^3 \tag{78} \\ &- \beta^2 \int_{t_k}^{t_{k+1}} F_\mu(t) \left( \sum_{j=1}^n \left(\mathop{\mathrm{id}}- \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,\cdot)\right)(\sigma_{k+1,j})^2 - nt \right)\,dt \tag{79} \\ & + \beta^2 \sum_{j=1}^n \eta v(t_k,\sigma_{k,j}) - \beta^2 \sum_{j=1}^n \int_{t_k}^{t_{k+1}} v(t,\sigma_{k+1,j})\,dt \tag{80} \\ &+ n\beta^2 \gamma \eta \tag{81}\,. \end{align}\] In the subsequent sections, we will show that each of these terms is small with high probability over the Gaussian vectors \(Z_k\), where “small” means bounded by \(\eta n\) times something which vanishes in the limit as \(n \to \infty\), then \(\eta \to 0\), then \(\gamma \to 0\), then \(\beta \to \infty\).

  • The gradient term 76 will have expectation zero in \(Z_k\) and the fluctuations will be controlled by concentration of measure arguments in conjunction with regularity properties of \(\tilde{\Lambda}_{\gamma}\).

  • The Hessian term 77 will be estimated using the approximate eigenvector condition for the covariance matrix that we chose, together with concentration arguments on the norms of a single update vector.

  • The third-order remainder term 78 will be estimated directly using concentration arguments that apply to the \(\ell^3\) norm of the update \(\Delta \sigma_k\).

  • The time term 79 will be estimated using convergence in Wasserstein distance to the SDE as well as the relationship between \(Y_t^\gamma\) and \(Y_t\).

  • The remaining error term 80 will be handled by time continuity estimates for \(v\) and by telescoping, and 81 does not require any further comment.

Summing up the estimates for the increments \(\mathop{\mathrm{obj}}(t_{k+1},\sigma_{k+1}) - \mathop{\mathrm{obj}}(t_k,\sigma_k)\), we arrive at a lower bound for \(\mathop{\mathrm{obj}}(t_K,\sigma_K)\), which then translates into a lower bound for the value of the Hamiltonian \(H(\sigma_K)\) at the last iterate. To analyze the potential function \(V(t_K,\sigma_K)\) as well as the telescoping term 80 , we use convergence of the empirical distributions of \(\sigma_k\) to the distribution of \(Y_t\) established in §5. Under the form of fRSB that we have assumed, the energy in the large \(n\) and large \(\beta\) limit matches the true optimum, using standard arguments already presented in [9]. In other words, the convergence results and fRSB enable us to show that the energy gain produced by the algorithm matches that predicted by the Parisi solution. We conclude by analyzing the effect of rounding and making appropriate choices for all the parameters.

6.2 The gradient term↩︎

Here we estimate the gradient term 76 using concentration of measure.

Lemma 31 (Bounded gradient of modified objective under primal process).
Let \(\sigma_k\) be generated as in and . Assume that the vectors \(Z_0\), …, \(Z_{k-1}\) from are in the high probability event from and that the matrix \(A\) satisfies the conclusions of . Then, conditioned on \(A, Z_0, \dots, Z_{k-1}\), with probability at least \(1 - C_5 \exp(-C_6 n^{1/3 - 4 \delta})\) in the Gaussian vector \(Z_k\), we have \[\left| \angles{2 \beta A_{\mathop{\mathrm{sym}}} \sigma_k - \partial_y \tilde{\Lambda}_{\gamma}(t_k,\sigma_k), \Delta \sigma_k} \right| \leqslant C_3 \left( 3 \beta + \frac{1}{\gamma} \right) \eta^{1/2} n^{2/3}.\]

Proof. Recall \(\Delta \sigma_k = \eta^{1/2} Q_k Z_k\). Note \[\mathop{{}\mathbb{E}}[\angles{2 \beta A_{\mathop{\mathrm{sym}}} \sigma_k - \partial_y \tilde{\Lambda}_{\gamma}(t_k,\sigma_k), \eta^{1/2} Q_k Z_k} \mid A, Z_0, \dots, Z_{k-1}] = 0\,,\] and the quantity in the expectation is a Lipschitz function of \(Z_k\) with Lipschitz norm bounded by \[\left( 2 \beta \left\lVert{A_{\mathop{\mathrm{sym}}}}\right\rVert + \max_{i \in [n]}|\partial_y \tilde{\Lambda}_{\gamma}(t_k,(\sigma_k)_i)| \right)|\sigma_k|_2 \eta^{1/2} \left\lVert{Q_k}\right\rVert.\] Recall \(2 \left\lVert{A_{\mathop{\mathrm{sym}}}}\right\rVert \leqslant 3\) since we are in the high probability event from . Furthermore, \(\left\lVert{Q_k}\right\rVert \leqslant M_1 n^{2\delta}\), which follows from  and . Lastly, it is apparent from that \(\Lambda(t,y)\) is an even function of \(y\), so \(\partial_y \Lambda(t,y)\) is an odd function, and this immediately yields that \(\partial_y \tilde{\Lambda}_{\gamma}(t,0) = 0\). Consequently, since \(\partial_{y,y} \tilde{\Lambda}_{\gamma} \leqslant 1/\gamma\) by , we have \[|\partial_y \tilde{\Lambda}_{\gamma}(t,y)| \leqslant\frac{1}{\gamma} |y|.\]

Hence, the Lipschitz norm can be bounded by \[M_1 \left( 3 \beta + \frac{1}{\gamma} \right) |\sigma_k|_2 \eta^{1/2} n^{2\delta}.\] Moreover, using , \[\begin{align} n^{-1/2} |\sigma_k|_2 = \left( \frac{1}{n}\sum_{j=1}^n \sigma_{k,j}^2\right)^{1/2}\!\! &{}\leqslant d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k),\mathop{\mathrm{dist}}(Y_{t_k}^\gamma)) + \left\lVert{Y_{t_k}^\gamma}\right\rVert_{L^2} \\&{}\leqslant d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k),\mathop{\mathrm{dist}}(Y_{t_k}^\gamma)) + t_k^{1/2} (1 + \sqrt{2} \beta (1 + e^{5 \beta^2 t_k}) \gamma). \end{align}\] Therefore, the Lipschitz norm can be bounded by \[\left( 3 \beta + \frac{1}{\gamma} \right) \left( M_2 + M_3\, d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k),\mathop{\mathrm{dist}}(Y_{t_k}^\gamma)) + M_4 \beta e^{5 \beta^2} \gamma \right) \eta^{1/2} n^{1/2 + 2\delta}.\] Since we are in the high probability event from , we can assume that \[\left( M_2 + M_3\, d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k),\mathop{\mathrm{dist}}(Y_{t_k}^\gamma)) + M_4 \beta e^{5 \beta^2} \gamma \right) \leqslant M_5.\] By the Gaussian concentration inequality for Lipschitz functions [49], we obtain the asserted error bound; here note that \((n^{2/3})^2 / (n^{1/2 + 2 \delta})^2 = n^{1/3 - 4 \delta}\). ◻

6.3 The Hessian term↩︎

In order to control the Hessian term 77 , we will first compare the quadratic expression \[\angles{\Delta \sigma_k, (2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}))) \Delta \sigma_k}\] with its expectation by a concentration inequality due to Hanson and Wright. Then to control the expectation, we will use the approximate eigenvector condition from which shows that the range of \(Q(t_k,\sigma_k)^2\) is approximately in the kernel of \(2 \beta A_{\mathop{\mathrm{sym}}} - D(t_k,\sigma_k) - 2 \beta^2 \operatorname{tr}(D(t_k,\sigma_k)^{-1})\). However, since \(D(t_k,\sigma_k)\) has a multiplicative normalization compared to \(\operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}))\), we also need to control a few other approximation errors.

Lemma 32 (Concentration of the Hessian term). Suppose that the matrix \(A\) satisfies the conclusions of . Suppose that the parameters \(\beta\), \(K\), \(\eta\), \(n\) satisfy \(n^{1/90} \eta \geqslant 1\) and 59 and that the vectors \(Z_0\), …, \(Z_{k-1}\) are chosen in the high probability event in . Then conditioned on \(A, Z_0, \dots, Z_{k-1}\), with probability at least \(1 - \exp(-n^{1/3-4\delta})\) in the Gaussian vector \(Z_k\), we have \[\begin{gather} \left| \angles{\Delta \sigma_k, (2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}))) \Delta \sigma_k} - n \eta \operatorname{tr}_n\big[Q_k\big(2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}))\big) Q_k\big] \right| \\ \leqslant C \left( 3 \beta + \frac{1}{\gamma} \right) \beta^2 \eta n^{2/3}. \end{gather}\]

Proof. Recall that \[\angles{\Delta \sigma_k, \left(2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}))\right) \Delta \sigma_k} = \eta \angles{Z_k, Q_k \big(2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}))\big) Q_k Z_k}.\] Since \(Z_k\) is a standard Gaussian vector, \[\mathop{{}\mathbb{E}}\angles{\Delta \sigma_k, \left(2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}))\right) \Delta \sigma_k} = n \eta \operatorname{tr}_n\left[Q_k \left(2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}))\right) Q_k\right].\] For convenience, let \[R_k := \eta Q_k \left(2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}))\right) Q_k.\] By the Hanson-Wright inequality, \[\mathop{{}\mathbb{P}}\left(|\angles{Z_k, R_k Z_k} - \mathop{{}\mathbb{E}}\angles{Z_k, R_k Z_k}| > \zeta \right) \leqslant 2 \exp\left(-\min\left(\frac{\zeta^2}{16 n\left\lVert{R_k}\right\rVert_{2}^2}, \frac{\zeta}{4 \left\lVert{R_k}\right\rVert}\right) \right).\] Here we compute \(\left\lVert{R_k}\right\rVert_{2} \leqslant\left\lVert{R_k}\right\rVert\) and \[\begin{align} \left\lVert{R_k}\right\rVert &= \eta \left\lVert{Q_k \left(2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j})) \right)Q_k}\right\rVert \\ &\leqslant\eta \left\lVert{2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}))}\right\rVert \left\lVert{Q_k^2}\right\rVert \\ &\leqslant\eta \left( 3 \beta + \frac{1}{\gamma} \right) M_1 n^{4\delta}, \end{align}\] where we applied and the facts that \(\left\lVert{2 A_{\mathop{\mathrm{sym}}}}\right\rVert \leqslant 3\) and \(\partial_{y,y} \tilde{\Lambda}_{\gamma} \leqslant 1/\gamma\). Now we substitute \[\zeta = 4\eta \left( 3 \beta + \frac{1}{\gamma} \right) M_1 n^{2/3},\] so that \[\min\left(\frac{\zeta^2}{16n \left\lVert{R_k}\right\rVert_{2}^2}, \frac{\zeta}{4 \left\lVert{R_k}\right\rVert}\right) = \min \left( \frac{ n^{4/3}}{n^{1+4\delta}}, \frac{n^{2/3}}{n^{4 \delta}} \right) \geqslant n^{1/3-4\delta}. \qedhere\] ◻

Lemma 33 (Approximation errors in the Hessian term). Suppose that the matrix \(A\) satisfies the conclusions of . Suppose that the parameters \(\beta\), \(K\), \(\eta\), \(n\) satisfy \(n^{1/90} \eta \geqslant 1\) and 59 and that the vectors \(Z_0\), …, \(Z_{k-1}\) are chosen in the high probability event in . Then \[\begin{gather} \left| \operatorname{tr}_n[Q_k(2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}))) Q_k] - \frac{2 \beta^2}{n} \sum_{j=1}^n v(t_k,\sigma_{k,j}) \right| \\ \leqslant C_1 \beta^3 n^{-\delta} + C_2\beta^2(\beta+\gamma^{-1}) n^{-1/4+2\delta} + C_3 \beta^4 \left( 2 d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k), \mathop{\mathrm{dist}}(Y_{t_k}^\gamma)) + (1 + e^{5 \beta^2}) \gamma \right). \end{gather}\]

Proof. Recall from that \[D_k = \sqrt{2} \beta \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2 \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}))\] Therefore, we can write \[\begin{align} 2 \beta A_{\mathop{\mathrm{sym}}} &- \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j})) - 2 \beta^2 \operatorname{tr}_n(\operatorname{diag}(v(t_k,\sigma_k))) \\ &= \frac{1}{\sqrt{2} \beta \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2} \left( 2 \beta A_{\mathop{\mathrm{sym}}} - D_k - 2 \beta^2 \operatorname{tr}_n(D_k^{-1}) \right) \\ & \quad + \left( 1 - \frac{1}{\sqrt{2} \beta \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2} \right) 2 \beta A_{\mathop{\mathrm{sym}}} \\ &\quad + \left(1 - \frac{1}{2 \beta^2 \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2^2} \right) 2 \beta^2 \operatorname{tr}_n(\operatorname{diag}(v(t_k,\sigma_k))) \end{align}\] Thus, by the triangle inequality, \[\begin{align} \bigg \lVert \big(2 \beta A_{\mathop{\mathrm{sym}}} &- \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j})) - 2 \beta^2 \operatorname{tr}_n(\operatorname{diag}(v(t_k,\sigma_k)))\big) Q_k \bigg \rVert_2 \tag{82} \\ &\leqslant\frac{1}{\sqrt{2} \beta \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2} \left\lVert{ (2 \beta A_{\mathop{\mathrm{sym}}} - D_k - 2 \beta^2 \operatorname{tr}(D_k^{-1})) Q_k }\right\rVert_2 \tag{83} \\ & \quad + \left| 1 - \frac{1}{\sqrt{2} \beta \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2} \right| 2 \beta \left\lVert{A_{\mathop{\mathrm{sym}}} Q_k}\right\rVert_2 \tag{84} \\ &\quad + \left| \sqrt{2} \beta \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2 - \frac{1}{\sqrt{2} \beta \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2} \right| \sqrt{2} \beta \left\lVert{Q_k}\right\rVert_2, \tag{85} \end{align}\] where we used that \(\operatorname{tr}(\operatorname{diag}(v(t_k,\sigma_{k,j}))) \leqslant\left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2\). To estimate 83 , recall from that \[\left\lVert{ (2 \beta A_{\mathop{\mathrm{sym}}} - D_k - 2 \beta^2 \operatorname{tr}(D_k^{-1})) Q_k }\right\rVert_2 \leqslant M_1 \beta^2 n^{-\delta} + M_0(\beta+\gamma^{-1})n^{-1/4+2\delta}.\] To handle 84 and 85 , we recall from that \[\left\lVert{Q_k}\right\rVert_2 \leqslant M_2 \beta,\] and in addition, \(\left\lVert{2 A_{\mathop{\mathrm{sym}}}}\right\rVert \leqslant 3\). In order to bound 84 and 85 , we also need to estimate \(\sqrt{2} \beta \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2 - 1\); recall that in the proof of , and in particular in , we arranged that \[\sqrt{2} \beta \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2 \geqslant 1/2\] and \[\label{eq:32another32estimate32for32Hessian32error} \left| \sqrt{2} \beta \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2 - 1 \right| \leqslant\sqrt{2} \beta \left( 2 d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k), \mathop{\mathrm{dist}}(Y_{t_k}^\gamma)) + (1 + e^{5 \beta^2}) \gamma \right).\tag{86}\] We therefore have that \[\label{eq:32another32estimate32for32Hessian32error322} \left| 1 - \frac{1}{(\sqrt{2} \beta \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2)} \right| \leqslant 2 \cdot \sqrt{2} \beta \left( 2 d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k), \mathop{\mathrm{dist}}(Y_{t_k}^\gamma)) + (1 + e^{5 \beta^2}) \gamma \right),\tag{87}\] which we will substitute into 84 . Similarly, to estimate 85 , we use that \[\left| \sqrt{2} \beta \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2 - \frac{1}{\sqrt{2} \beta \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2} \right| \leqslant\left| \sqrt{2} \beta \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2 - 1 \right| + \left| 1 - \frac{1}{\sqrt{2} \beta \left\lVert{\operatorname{diag}(v(t_k,\sigma_{k,j}))}\right\rVert_2} \right|,\] which can then be estimated by 86 and 87 . Substituting all these estimates into 83 - 85 , we obtain the following bound for 82 : \[\begin{align} \bigg \lVert \big(2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j})) - 2 \beta^2 \operatorname{tr}_n(\operatorname{diag}(v(t_k,\sigma_k)))\big) Q_k \bigg \rVert_2 &\leqslant M_3 \beta^2 n^{-\delta} + M_3\beta(\beta+\gamma^{-1}) n^{-1/4+2\delta} +\\ &M_4 \beta^3 \left( 2 d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k), \mathop{\mathrm{dist}}(Y_{t_k}^\gamma)) + (1 + e^{5 \beta^2}) \gamma \right). \end{align}\] This in turn implies that \[\begin{align} \operatorname{tr}_n\big[Q_k\big(2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j})) - &2 \beta^2 \operatorname{tr}_n(\operatorname{diag}(v(t_k,\sigma_k)))\big) Q_k\big] \leqslant M_5\beta^3 n^{-\delta} + M_5\beta^2(\beta+\gamma^{-1}) n^{-1/4+2\delta} + \\ &\qquad\qquad\qquad M_6 \beta^4 \left( 2 d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k), \mathop{\mathrm{dist}}(Y_{t_k}^\gamma)) + (1 + e^{5 \beta^2}) \gamma \right). \end{align}\] Finally, we write \[\begin{align} &\operatorname{tr}_n[Q_k(2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}))) Q_k] - 2\beta^2\operatorname{tr}_n[\operatorname{diag}(v(t_k,\sigma_{k,j}))] = \\ &\operatorname{tr}_n\big[Q_k\big(2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j})) - 2 \beta^2 \operatorname{tr}_n(\operatorname{diag}(v(t_k,\sigma_k)))\big) Q_k\big] + 2\beta^2(\operatorname{tr}_n[Q_k^2] - 1) \operatorname{tr}_n[\operatorname{diag}(v(t_k,\sigma_{k,j}))], \end{align}\] and recall from that \(|\operatorname{tr}_n[Q_k^2] - 1| \leqslant M_7 \beta^2 n^{-\delta}\). Since we already have the error term \(M_5 \beta^3 n^{-\delta}\), this new error can be combined with the previous one up to changing constants. ◻

By combining and , we obtain the following result.

Lemma 34 (Bound for the Hessian term). Suppose that \(Z_0\), …, \(Z_{k-1}\) are chosen in the high probability event from , which occurs with probability at least \(1 - n\eta^{-2}C_1 \exp(-C_2 n^{2/9 - 4 \delta})\), and that the matrix \(A\) satisfies the conclusions of , which happens with probability \(1 - C_3 \exp(-C_4 n)\). Then, conditioned on \(A, Z_0, \dots, Z_{k-1}\), with probability \(1 - 2 \exp(-n^{1/3-4\delta})\) in the Gaussian vector \(Z_k\), we have \[\begin{gather} \left| \frac{1}{2} \angles{\Delta \sigma_k, \big(2 \beta A_{\mathop{\mathrm{sym}}} - \operatorname{diag}(\partial_{y,y} \tilde{\Lambda}_{\gamma}(t_k,\sigma_{k,j}))\big) \Delta \sigma_k} - \beta^2 \eta \sum_{j=1}^n v(t_k,\sigma_{k,j}) \right| \leqslant\eta n \cdot \biggl( C_1(3 \beta + \gamma^{-1}) \beta^2 n^{-1/3}\\ + C_2 \beta^3 n^{-\delta} + C_3\beta^2(\beta+\gamma^{-1}) n^{-1/4+2\delta} + C_4 \beta^4 d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k),\mathop{\mathrm{dist}}(Y_{t_k}^\gamma)) + C_5 e^{5 \beta^2} \beta^4 \gamma \biggr). \end{gather}\]

6.4 The remainder term↩︎

This section proves the following bound for the third-order remainder.

Lemma 35 (Bounded third-order Taylor remainder). Assume that \(Z_0\), …, \(Z_{k-1}\) are in the high probability event from and that the matrix \(A\) satisfies the conclusions of . Then, conditioned on \(A, Z_0, \dots, Z_{k-1}\), with probability at least \(1 - C_1 \exp(-C_2 n^{1/3-4\delta})\), we have \[\sum_{j=1}^n |\partial_{y,y,y} \tilde{\Lambda}_{\gamma}(t_k, \xi_{k,j})| |(\Delta \sigma_k)_j|^3 \leqslant\frac{C_3}{\gamma^2} \beta^{3/2} \eta^{3/2} n.\]

Because \(|\partial_{y,y,y} \tilde{\Lambda}_{\gamma}| \leqslant 2 / \gamma^2\), this lemma follows immediately once we establish the following bounds for the \(3\)-norm of the updates.

Lemma 36 (Bounded \(3\)-norm of the updates).
Suppose that the matrix \(A\) satisfies the conclusions of . Suppose that the parameters \(\beta\), \(K\), \(\eta\), \(n\) satisfy \(n^{1/90} \eta \geqslant 1\) and 59 and that the vectors \(Z_0\), …, \(Z_{k-1}\) are chosen in the high probability event in . Then, conditioned on \(A, Z_0, \dots, Z_{k-1}\), with probability at least \(1 - C_1 \exp(-C_2 n^{2/3-4\delta})\) in the Gaussian vector \(Z_k\), we have \[\sum_{i=1}^n |(\Delta \sigma_k)_i|^3 \leqslant C_3 \beta^{3/2} \eta^{3/2} n.\]

Proof. Recall that \(|x|_r \leqslant|x|_p\) for every \(1 \leqslant p\leqslant r\). Therefore, \[|x|_3 = \left( \sum_{j=1}^n |x_j|^3 \right)^{1/3} \leqslant\left( \sum_{j=1}^n |x_j|^2 \right)^{1/2} = |x|_2\,.\] Consequently, \(f(\Delta \sigma_k) = \left( \sum_{j=1}^n |(\Delta \sigma_k)_j|^3 \right)^{1/3}\) is a \(1\)-Lipschitz function of \(\Delta \sigma_k\). Moreover, conditioned on \(A\) and \(Z_0\), …, \(Z_{k-1}\), the variable \(\Delta \sigma_k\) is Gaussian with covariance \(\eta Q_k^2\). By the concentration inequality for Lipschitz functions of Gaussian random variables [49], \[\label{eq:3-norm-remainder} \mathop{{}\mathbb{P}}\left(\left|f(\Delta \sigma_k) - \mathop{{}\mathbb{E}}\left[f(\Delta \sigma_k)\right]\right| \geqslant\eta^{1/2} n^{1/3}\right) \leqslant\exp\left(-\frac{n^{2/3}}{2 \left\lVert{Q_k^2}\right\rVert}\right) \leqslant\exp(-M_1 n^{2/3-4\delta}).\tag{88}\] As \(y = z^{1/3}\) is concave for \(z \in \mathbb{R}_{\geqslant 0}\), Jensen’s inequality implies \[\begin{align} \mathop{{}\mathbb{E}}f(\Delta \sigma_k) &= \mathop{{}\mathbb{E}}\left( \sum_{j=1}^n |(\Delta \sigma_k)_j|^3 \right)^{1/3} \\ &\leqslant\left( \mathop{{}\mathbb{E}}\sum_{j=1}^n |(\Delta \sigma_k)_j|^3 \right)^{1/3}\,. \end{align}\] Since \((\Delta \sigma_k)_i \sim \mathcal{N}\left(0, \eta(Q_k^2)_{i,i}\right)\), \[\mathop{{}\mathbb{E}}|\Delta \sigma_i|^3 = M_2 (Q_k^2)_{i,i}^{3/2}\eta^{3/2}\] for some constant \(M_2\). Thus, \[\begin{align} \mathop{{}\mathbb{E}}\left[f(\Delta \sigma_k)\right] &\leqslant M^{1/3}_2 \mathsf{Tr}_n(E_{\mathcal{D}_n}[\eta Q_k^2]^{3/2})^{1/3}\\ &= M_2^{1/3} n^{1/3} \eta^{1/2} \operatorname{tr}_n(E_{\mathcal{D}_n}[Q_k^2]^{3/2})^{1/3} \\ &\leqslant M_2^{1/3} n^{1/3} \eta^{1/2} \operatorname{tr}_n(E_{\mathcal{D}_n}[Q_k^2]^2)^{1/4} \\ &\leqslant M_3 n^{1/3} \eta^{1/2} \beta^{1/2}, \end{align}\] where the last inequality follows from and the assumed application of to \(Z_0, \dots, Z_{k-1}\). Therefore, by , \[\mathop{{}\mathbb{P}}\left(f(\Delta \sigma_k) \geqslant M_3 n^{1/3} \beta^{1/2}\eta^{1/2}\right) \leqslant\exp(-M_1 n^{2/3 - 4 \delta}).\] It follows that \[\sum_{i=1}^n |\Delta \sigma_i|^3 \leqslant M_4 n \beta^{3/2}\eta^{3/2}\,.\] with probability at least \[1 - \exp(-M_1 n^{2/3 - 4 \delta}).\] in the Gaussian vector \(Z_k\), conditioning on the vectors \(Z_0\), …, \(Z_{k-1}\) and assuming that they fall under the high probability event in . ◻

6.5 The time terms↩︎

We will now estimate 79 by controlling how close the regularization-corrected self-overlap \(\frac{1}{n} \sum_{j=1}^n (\sigma_{k+1,j} - \gamma\partial_y \tilde{\Lambda}_{\gamma}(t, \sigma_{k+1,j}))^2\) is to \(t\) through the convergence to the primal Auffinger-Chen SDE .

Lemma 37 (Bounding the \(2\)-norm under the \(\gamma\)-convolved SDE).
Assume that the hypotheses and conclusions of  are satisfied and that \(\beta \eta^{1/2} < 1\). Then for some constants \(C_j\), we have \[\left| \frac{1}{n} \sum_{j=1}^n (\sigma_{k+1,j} - \gamma\partial_y \tilde{\Lambda}_{\gamma}(t, \sigma_{k+1,j}))^2 - t \right| \leqslant C_1 d_{W,2}(\mathop{\mathrm{emp}}(\sigma_{k+1}), \mathop{\mathrm{dist}}(Y_{t_{k+1}}^\gamma)) + C_2 e^{5 \beta^2} \gamma + C_3 \beta^3 \eta^{1/2}\,,\] where \(t \in [t_k,t_{k+1}]\).

Proof. Let \(\Sigma\) be a random variable with distribution \(\mathop{\mathrm{emp}}(\sigma_{k+1})\) and let \(Y_t^\gamma\) be the stochastic process defined earlier. Consider an optimal coupling of \(\Sigma\) and \(Y_{t_{k+1}}^\gamma\) on some probability space. Recall that \(\mathop{\mathrm{id}}- \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,\cdot)\) is \(1\)-Lipschitz. Therefore, \[\begin{align} \left\lVert{[\Sigma - \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,\Sigma)] - [Y_{t_{k+1}}^\gamma - \gamma \partial_y \tilde{\Lambda}_{\gamma}(t, Y_{t_{k+1}}^\gamma)]}\right\rVert_{L^2} &\leqslant\left\lVert{\Sigma - Y_{t_{k+1}}^\gamma}\right\rVert_{L^2} \\ &= d_{W,2}(\operatorname{emp}(\sigma_{k+1}), \mathop{\mathrm{dist}}(Y_{t_{k+1}}^\gamma)). \end{align}\] Meanwhile, by , \[\left\lVert{[Y_{t_{k+1}}^\gamma - \gamma \partial_y \tilde{\Lambda}_{\gamma}(t, Y_{t_{k+1}}^\gamma)] - Y_{t_{k+1}}}\right\rVert_{L^2} \leqslant 5^{-1/2} e^{5 \beta^2} \gamma.\] Furthermore, from the proof of  in the case that \(\gamma = 0\), \[\left\lVert{Y_{t_{k+1}} - Y_t}\right\rVert_{L^2} \leqslant 2\beta\sqrt{\eta}\,.\] Finally, for any \(t \in [0,q^*_\beta)\)[9] implies that \(\left\lVert{Y_t}\right\rVert_{L^2} = t^{1/2}\). Thus, combining these estimates with the triangle inequality yields, \[\begin{align} \left| \left( \frac{1}{n} \sum_{j=1}^n (\sigma_{k+1,j} - \gamma\partial_y \tilde{\Lambda}_{\gamma}(t, \sigma_{k+1,j}))^2 \right)^{1/2} - t^{1/2} \right| &\leqslant\left\lVert{\Sigma - Y_t}\right\rVert_{L^2} \\ &\leqslant d_{W,2}(\operatorname{emp}(\sigma_{k+1}), \mathop{\mathrm{dist}}(Y_{t_{k+1}}^\gamma)) + 5^{-1/2} e^{5 \beta^2} \gamma + 2\beta \eta^{1/2}. \end{align}\] Now writing temporarily \(s = \frac{1}{n} \sum_{j=1}^n (\sigma_{k+1,j} - \gamma\partial_y \tilde{\Lambda}_{\gamma}(t, \sigma_{k+1,j}))^2\), we have \[\begin{align} |s - t| &= |s^{1/2} - t^{1/2}|(s^{1/2} + t^{1/2}) \\ &\leqslant|s^{1/2} - t^{1/2}|(t^{1/2} + |s^{1/2} - t^{1/2}| + t^{1/2}) \\ &\leqslant|s^{1/2} - t^{1/2}|(2 + |s^{1/2} - t^{1/2}|) \end{align}\] since \(t \leqslant 1\). Furthermore, if the assumptions of the theorem are satisfied, then \(d_{W,2}(\operatorname{emp}(\sigma_{k+1}), \mathop{\mathrm{dist}}(Y_{t_{k+1}}^\gamma)) + 5^{-1/2} e^{5 \beta^2} \gamma + 2 \beta \eta^{1/2}\) is bounded by a constant \(M\), so we can simply estimate this by \(|s - t| \leqslant M |s^{1/2} - t^{1/2}|\). ◻

We thus obtain the following estimate for the final distortion term.

Lemma 38 (Distortion of drift term in time error).
Assume the hypotheses and conclusions of . Then \[\begin{align} &\left| \beta^2\! \int_{t_k}^{t_{k+1}}\! F_\mu(t) \!\left( \sum_{j=1}^n ((\mathop{\mathrm{id}}- \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,\cdot))(\sigma_{k+1,j}))^2 - nt \!\right)\!dt \right| \\&\qquad\qquad\qquad{}\leqslant n \beta^2 \eta \left[ C_1 d_{W,2}(\operatorname{emp}(\sigma_{k+1}), \mathop{\mathrm{dist}}(Y_{t_{k+1}}^\gamma)) + C_2 e^{5 \beta^2} \gamma + 2\beta \eta^{1/2} \right]. \end{align}\]

Proof. Note that \(F_\mu(t) \leqslant 1\) and apply the previous lemma to estimate \(|\frac{1}{n} \sum_{j=1}^n ((\mathop{\mathrm{id}}- \gamma \partial_y \tilde{\Lambda}_{\gamma}(t,\cdot))(\sigma_{k+1,j}))^2 - t|\). ◻

We finish this section with an estimate to help with 80 . Note that because we use telescoping, we include the summation over \(k\) in the statement here.

Lemma 39 (Bound for telescoping term). \[\left| \sum_{k=0}^{K-1} \left( \beta^2 \sum_{j=1}^n \eta v(t_k,(\sigma_k)_j) - \beta^2 \sum_{j=1}^n \int_{t_k}^{t_{k+1}} v(t,(\sigma_{k+1})_j)\,dt \right) \right| \leqslant 9 \beta^4 \eta n.\]

Proof. We write \[\eta v(t_k,\sigma_k) - \int_{t_k}^{t_{k+1}} v(t,\sigma_{k+1})\,dt = \left( \eta v(t_k,\sigma_k) - \eta v(t_{k+1},\sigma_{k+1}) \right) + \left( \eta v(t_{k+1},\sigma_{k+1}) - \int_{t_k}^{t_{k+1}} v(t,\sigma_{k+1})\,dt \right)\] The first term telescopes when we sum over \(k\), that is, \[\left| \sum_{k=0}^{K-1} (v(t_k,\sigma_k) - v(t_{k+1},\sigma_{k+1})) \right| = |v(0,0) - v(t_K, \sigma_K)| \leqslant 1\] since \(\gamma \leqslant v \leqslant 1 + \gamma\) by . After multiplying by \(\beta^2 \eta\) and summing over \(j \in [n]\), the final contribution is \(\beta^2 \eta n \leqslant\beta^4 \eta n\).

To estimate the second term, recall that by (4), \[|v(t_{k+1},\sigma_{k+1}) - v(t,\sigma_{k+1})| \leqslant 18 \beta^2 (t_{k+1} - t).\] Integrating from \(t_k\) to \(t_{k+1}\) yields \[\left| \eta v(t_{k+1},\sigma_{k+1}) - \int_{t_k}^{t_{k+1}} v(t,\sigma_{k+1})\,dt \right| \leqslant 9 \beta^2 \eta^2.\] After summing over \(j\) and multiplying by \(\beta^2\), we obtain \[\left| \beta^2 \sum_{j=1}^n \eta v(t_{k+1},(\sigma_{k+1})_j) - \beta^2 \sum_{j=1}^n \int_{t_k}^{t_{k+1}} v(t,(\sigma_{k+1})_j)\,dt \right| \leqslant 9 \beta^4 \eta^2 n.\] Finally, summing over \(k\) gives a final contribution of \(9 \beta^4 (K \eta) \eta n \leqslant 9 \beta^4 \eta n\). ◻

6.6 Conclusion of the energy argument↩︎

Now we can put together all the estimates we have made in the individual terms in the Taylor expansion and sum over \(k\).

Theorem 31 (Estimates for change in the objective function). Suppose . Suppose that the matrix \(A\) satisfies the conclusions of . Suppose that the parameters \(\beta\), \(K\), \(\eta\), \(n\) satisfy \(n^{1/90} \eta \geqslant 1\) and 59 and that the vectors \(Z_0\), …, \(Z_{K-1}\) are chosen in the high probability event in . Then with probability at least \(1 - C_1n\eta^{-2} \exp(-C_2 n^{2/9 - 4\delta})\) in the Gaussian vectors \(Z_0\), …, \(Z_{K-1}\), the conclusions of hold and in addition \[\operatorname{obj}(t_K, \sigma_K) - \operatorname{obj}(t_0,\sigma_0) \geqslant-n \cdot e^{C_4 \beta^2} \left( C_5 \gamma^{-1} \eta^{-1/2} n^{-1/36} + C_6 \gamma^{-2} \eta^{1/2} + C_7 n^{-\delta/2} + C_8 \gamma \right).\]

Proof. We take the Taylor expansion in 76 - 81 and apply the bounds for each individual terms that we have proved, namely:

  • 76 is estimated by ,

  • 77 is estimated by ,

  • 78 is estimated by ,

  • 79 is estimated by .

Hence, assuming we are in the high probability events described in Theorem 27 and in each of those lemmas, we have \[\begin{align} \operatorname{obj}(t_{k+1}, \sigma_{k+1}) &- \operatorname{obj}(t_k,\sigma_k) \\ \geqslant&-\,n\eta \cdot M_1 \left( 3 \beta + \gamma^{-1} \right) \eta^{-1/2} n^{-1/3} \tag{89} \\ &- n\eta \cdot \biggl( M_2(3 \beta + \gamma^{-1}) \beta^2 n^{-1/3} + M_3 \beta^3 n^{-\delta} + M_4 \beta^4 d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k),\mathop{\mathrm{dist}}(Y_{t_k}^\gamma)) + M_5 e^{5 \beta^2} \beta^4 \gamma \biggr) \tag{90} \\ &- n\eta \cdot M_2\beta^2(\beta + \gamma^{-1})n^{-1/4+2\delta} \nonumber\\ &- n\eta \cdot M_6 \beta^{3/2} \gamma^{-2}\eta^{1/2} \tag{91} \\ &- n \eta \cdot \left[ M_7 \beta^2 d_{W,2}(\mathop{\mathrm{emp}}(\sigma_{k+1}), \mathop{\mathrm{dist}}(Y_{t_{k+1}}^\gamma)) + M_8 \beta^2 e^{5 \beta^2} \gamma + M_9 \beta^3 \eta^{1/2} \right] \tag{92} \\ & + \beta^2 \sum_{j=1}^n \eta v(t_k,\sigma_{k,j}) - \beta^2 \sum_{j=1}^n \int_{t_k}^{t_{k+1}} v(t,\sigma_{k+1,j})\,dt \tag{93} \\ &- n\eta \cdot \beta^2 \gamma \tag{94}\,, \end{align}\] where we have left 93 alone in order to apply telescoping later. We combine the remaining errors 89 - 92 and 81 and throw away terms that are dominated by other terms in the expansion, and we thus obtain: \[\begin{align} n \eta \cdot \biggl( M_{10} \left( 3 \beta + \gamma^{-1} \right) \eta^{-1/2} (n^{-1/3}+n^{-1/4+2\delta})& + M_{11} \beta^3 \gamma^{-2} \eta^{1/2} + M_{12} \beta^3 n^{-\delta} \\ &+ M_{13} \beta^2 d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k),\mathop{\mathrm{dist}}(Y_{t_k}^\gamma)) + M_{13} \beta^4 e^{5 \beta^2} \gamma \biggr). \end{align}\] Then we recall from that \[\begin{align} d_{W,2}(\mathop{\mathrm{emp}}(\sigma_k),\mathop{\mathrm{dist}}(Y_{t_k}^\gamma)) &\leqslant\exp( M_{14} \beta^2) \left[ M_{15} \beta^4 \eta + M_{16} n^{-\delta} + M_{17} e^{10 \beta^2} \gamma^2 + M_{18} \beta^{-2} \eta^{-1} n^{-1/18} \right]^{1/2} \\ &\leqslant\exp( M_{14} \beta^2) \left( M_{19} \beta^2 \eta^{1/2} + M_{20} n^{-\delta/2} + M_{21} e^{5 \beta^2} \gamma + M_{21} \beta^{-1} \eta^{-1/2} n^{-1/36} \right). \end{align}\] Hence, the errors from 89 - 92 and 81 can be bounded by \[\label{eq:32most32Taylor32terms} n \eta \cdot e^{M_{22} \beta^2} \left( M_{23} \gamma^{-1} \eta^{-1/2} n^{-1/36} + M_{24} \gamma^{-2} \eta^{1/2} + M_{25} n^{-\delta/2} + M_{26} \gamma \right).\tag{95}\] Finally, we sum the error for \(k = 0\), …, \(K-1\). Hence, the \(n \eta\) in front of 95 becomes \(n K \eta \leqslant n\). Meanwhile, the summation over \(k\) of 93 is estimated by as smaller than \(9 \beta^4 \eta n\). Since we already have terms of the form \(n \exp(M \beta^2) \gamma^{-2} \eta^{1/2}\), this additional error can be absorbed by changing the constants. Therefore, \[\operatorname{obj}(t_K, \sigma_K) - \operatorname{obj}(t_0,\sigma_0) \geqslant-n \cdot e^{M_{22} \beta^2} \left( M_{23} \gamma^{-1} \eta^{-1/2} n^{-1/36} + M_{27} \gamma^{-2} \eta^{1/2} + M_{25} n^{-\delta/2} + M_{26} \gamma \right).\]

It remains to estimate the probability of failure. We take a union bound over the errors in each step that arise both from the proof of and from the arguments in this section. In , , , and , the probabilities are all bounded by \(M_{27} \exp(-M_{28} n^{1/3 - 4 \delta})\), assuming that \(Z_0\), …, \(Z_{k-1}\) are chosen already from the high probability events in . We take a union bound over the steps, where the number of terms is bounded by \(\eta^{-1}\), for a total probability of \(M_{27}\eta^{-1}\exp(-M_{28} n^{1/3 - 4 \delta})\). Meanwhile, the probabilities of error from are bounded by \(M_{29} n\eta^{-2}\exp(-M_{30} n^{2/9-4\delta})\), and so can combine these error estimates together at the cost of changing the constants. ◻

Combining , which shows that the per step deviation of the objective function is sufficiently small, with high probability, with the uniform estimates for the moduli of continuity of \(\tilde{\Lambda}_{\gamma}\) provided in §2 allows one to conclude that, under fRSB, the desired energy increments come arbitrarily close to the value given by the Parisi formula. The last step is to use the fact that the norm of the iterates is sufficiently large, and that they are close to the cube to guarantee that the rounding scheme outlined in  doesn’t lose too much energy.

Theorem 32 (Final energy estimate for truncated \(\sigma_K\)). Fix \(\varepsilon> 0\). Assume , assume \(n^{1/90} \eta \geqslant 1\) and 59 , and assume we are in the high probability events of , , . Let \(\sigma_K\) be the vector at the last step of the algorithm and let \(\tilde{\sigma}\) be the truncation \(\tilde{\sigma}_j = \mathop{\mathrm{sgn}}(\sigma_{K,j}) \min(1, |\sigma_{K,j}|)\). Finally, write \[\mathcal{E}(\beta) = 2 \beta \int_0^{q_\beta^*} \int_t^1 F_\mu(s)\,ds\,.\] Then \[\begin{align} \frac{1}{n} H(\tilde{\sigma}) &\geqslant\mathcal{E}(\beta) -\frac{1}{\beta} \exp( C_2 \beta^2) \left[ C_5 \eta^{-1/2} n^{-1/36} + C_6 \eta^{1/2} + C_7 n^{-\delta/2} \right]^{8\beta^2 / (1 + 8 \beta^2)} \\ & \quad\quad\quad - \frac{1}{\beta} \exp(C_8 \beta^2) [C_9 \gamma^{-1} \eta^{-1/2} n^{-1/36} + C_{10} \gamma^{-2} \eta^{1/2} + C_{11} \gamma^{4 \beta^2/(1+4\beta^2)}]. \end{align}\]

Proof.  

Unwinding the objective function.↩︎

By the definition of \(\mathop{\mathrm{obj}}\), \[\beta \angles{\sigma_K, A_{\mathop{\mathrm{sym}}} \sigma_K} = \mathop{\mathrm{obj}}(q_\beta^*,\sigma_K) - \mathop{\mathrm{obj}}(0,0) + \sum_{j=1}^n \tilde{\Lambda}_{\gamma}(q_\beta^*,\sigma_{K,j}) - n \Lambda(0,0) - n \beta^2 \int_0^{q_\beta^*} t F_\mu(t)\,dt.\] Thus, substituting \(\mathcal{E}(\beta)\) in , we get \[\beta \angles{\sigma_K, A_{\mathop{\mathrm{sym}}} \sigma_K} = \beta \mathcal{E}(\beta)n + \mathop{\mathrm{obj}}(q_\beta^*,\sigma_K) - \mathop{\mathrm{obj}}(0,0) + \sum_{j=1}^n \tilde{\Lambda}_{\gamma}(q_\beta^*,\sigma_{K,j}) - n \mathop{{}\mathbb{E}}\Lambda(q_\beta^*,Y_{q_\beta^*}).\] Recall that \(t_0 = 0\), \(\sigma_0 = 0\), and \(t_K = q_\beta^*\), and hence gives us a lower bound for \(\mathop{\mathrm{obj}}(q_\beta^*,\sigma_K) - \mathop{\mathrm{obj}}(0,0)\). It thus remains to estimate the approximation error \(\sum_{j=1}^n \tilde{\Lambda}_{\gamma}(q_\beta^*,\sigma_{K,j}) - n \mathop{{}\mathbb{E}}\Lambda(q_\beta^*,Y_{q_\beta^*})\) for the entropy function, and then to perform a truncation to obtain a vector \(\tilde{\sigma} \in \{-1,1\}^n\) from \(\sigma_K\).

Approximation error for entropy.↩︎

We now use the SDE and regularity analysis from §2 to estimate \(\sum_{j=1}^n \tilde{\Lambda}_{\gamma}(q_\beta^*, (\sigma_K)_j)\). Let \(\Sigma\) be a random variable with distribution \(\mathop{\mathrm{emp}}(\sigma_K)\), so that \[\frac{1}{n} \sum_{j=1}^n \tilde{\Lambda}_{\gamma}(q_\beta^*, (\sigma_K)_j) = \mathop{{}\mathbb{E}}\tilde{\Lambda}_{\gamma}(q_\beta^*, \Sigma).\] Assume that \(\Sigma\) is on the same probability space as the processes \(Y_t\) and \(Y_t^\gamma\) such that \(\Sigma\) and \(Y_{q_\beta^*}\) are optimally coupled, that is, \(\left\lVert{\Sigma - Y_{q_\beta^*}}\right\rVert_{L^2} = d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K),\mathop{\mathrm{dist}}(Y_{q_\beta^*}))\). Let \(f: \mathbb{R}\to [-1,1]\) be the truncation \(f(y) = \mathop{\mathrm{sgn}}(y) \min(1, |y|)\). Since \(\tilde{\Lambda}_{\gamma}(t,y)\) is an increasing function of \(|y|\), we have \[\tilde{\Lambda}_{\gamma}(q_\beta^*, \Sigma) \geqslant\tilde{\Lambda}_{\gamma}(q_\beta^*, f(\Sigma)) \geqslant\Lambda(q_\beta^*, f(\Sigma)) - (1 + 4 \beta^2) (2 \beta^2 \gamma)^{4 \beta^2/(1 + 4 \beta^2)},\] where the second estimate follows from . Furthermore, by , \[\begin{align} \mathop{{}\mathbb{E}}|\Lambda(q_\beta^*, f(\Sigma)) - \Lambda(q_\beta^*, Y_{q_\beta^*})| &\leqslant(1 + 8 \beta^2) \mathop{{}\mathbb{E}}\left| \frac{f(\Sigma) - Y_{q_\beta^*}}{2} \right|^{8\beta^2 / (1 + 8 \beta^2)} \\ &\leqslant(1 + 8 \beta^2) \left( \frac{\left\lVert{f(\Sigma) - Y_{q_\beta^*}}\right\rVert_{L^1}}{2} \right)^{8\beta^2 / (1 + 8 \beta^2)}, \end{align}\] where the last estimate follows from Hölder’s inequality. Finally, note that since \(Y_{q_\beta^*} \in [-1,1]\), \[\begin{align} \left\lVert{f(\Sigma) - Y_{q_\beta^*}}\right\rVert_{L^1} &\leqslant\left\lVert{\Sigma - Y_{q_\beta^*}}\right\rVert_{L^1} \leqslant\left\lVert{\Sigma - Y_{q_\beta^*}}\right\rVert_{L^2} \\ &= d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K),\mathop{\mathrm{dist}}(Y_{q_\beta^*})) \\ &\leqslant d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K), \mathop{\mathrm{dist}}(Y_{q_\beta^*}^\gamma)) + \left\lVert{Y_{q_\beta^*}^\gamma - Y_{q_\beta^*}}\right\rVert_{L^2} \\ &\leqslant_{\text{\prettyref{lem:sde-closeness}}} d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K), \mathop{\mathrm{dist}}(Y_{q_\beta^*}^\gamma)) + \sqrt{2} e^{5 \beta^2} \gamma. \end{align}\] Thus, \[\begin{gather} \label{eq:32entropy32approximation32error} \sum_{j=1}^n \tilde{\Lambda}_{\gamma}(q_\beta^*, (\sigma_K)_j) - n \mathop{{}\mathbb{E}}\Lambda(q_\beta^*, Y_{q^*_\beta}) \\ \geqslant-n \left( (1 + 8\beta^2) \left( d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K), \mathop{\mathrm{dist}}(Y_{q_\beta^*}^\gamma)) + \sqrt{2} e^{5 \beta^2} \gamma \right)^{8\beta^2 / (1 + 8 \beta^2)} + (1 + 4 \beta^2) (2 \beta^2 \gamma)^{4 \beta^2/(1 + 4 \beta^2)} \right). \end{gather}\tag{96}\]

Truncation error.↩︎

Recall we set \(\tilde{\sigma}_j = f(\sigma_{K,j})\) where \(f(y) = \mathop{\mathrm{sgn}}(y) \min(|y|,1)\). We will now estimate \(|\sigma_K - \tilde{\sigma}|_2\). Note that \[\frac{1}{n} |\sigma_K - \tilde{\sigma}|_2^2 = \mathop{{}\mathbb{E}}|\Sigma - f(\Sigma)|^2,\] where \(\Sigma\) is a random variable with distribution \(\mathop{\mathrm{emp}}(\sigma_K)\) as above. Since \(Y_{q_\beta^*} \in [-1,1]\), we have \(|\Sigma - f(\Sigma)| \leqslant|\Sigma - Y_{q_\beta^*}|\), so that, by , \[\begin{align} n^{-1/2} |\sigma_K - \tilde{\sigma}|_2 &= \left\lVert{\Sigma - f(\Sigma)}\right\rVert_{L^2} \\ &\leqslant\left\lVert{\Sigma - Y_{q_\beta^*}}\right\rVert_{L^2} \\ &\leqslant d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K), \mathop{\mathrm{dist}}(Y_{q_\beta^*})) \\ &\leqslant d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K), \mathop{\mathrm{dist}}(Y_{q_\beta^*}^\gamma)) + \sqrt{2} e^{5 \beta^2} \gamma. \end{align}\] Since \(\left\lVert{A_{\mathop{\mathrm{sym}}}}\right\rVert \leqslant 3\) in our given sample, we conclude that \[\begin{align} \angles{\tilde{\sigma}, A_{\mathop{\mathrm{sym}}} \tilde{\sigma}} &\geqslant\angles{\sigma_K, A_{\mathop{\mathrm{sym}}} \sigma_K} + 2 \angles{\tilde{\sigma}, A_{\mathop{\mathrm{sym}}} (\tilde{\sigma} - \sigma_K)} - \angles{(\tilde{\sigma} - \sigma_K), A_{\mathop{\mathrm{sym}}}(\tilde{\sigma} - \sigma_K)} \\ &\geqslant\angles{\sigma_K, A_{\mathop{\mathrm{sym}}} \sigma_K} - 6 |\tilde{\sigma}|_2 |\tilde{\sigma} - \sigma_K|_2 - 3 |\tilde{\sigma} - \sigma_K|_2^2 \\ &\geqslant\angles{\sigma_K, A_{\mathop{\mathrm{sym}}} \sigma_K} - 6n \left\lVert{\Sigma - f(\Sigma)}\right\rVert_{L^2} - 3n \left\lVert{\Sigma - f(\Sigma)}\right\rVert_{L^2}^2. \end{align}\] In the high probability event of , \(d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K), \mathop{\mathrm{dist}}(Y_{q_\beta^*}^\gamma)) + \sqrt{2} e^{5 \beta^2} \gamma\) will be bounded by a constant, and hence \(\left\lVert{\Sigma - f(\Sigma)}\right\rVert_{L^2}^2 \leqslant M_1 \left\lVert{\Sigma - f(\Sigma)}\right\rVert_{L^2}\). Thus, \[\label{eq:32rounding32approximation32error} \angles{\tilde{\sigma}, A_{\mathop{\mathrm{sym}}} \tilde{\sigma}} \geqslant\angles{\sigma_K, A_{\mathop{\mathrm{sym}}} \sigma_K} - nM_2 \left(d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K), \mathop{\mathrm{dist}}(Y_{q_\beta^*}^\gamma)) + \sqrt{2} e^{5 \beta^2} \gamma \right).\tag{97}\]

Putting the estimates together.↩︎

We write \[\begin{align} \beta \angles{\tilde{\sigma}, A_{\mathop{\mathrm{sym}}} \tilde{\sigma}} - \beta \mathcal{E}(\beta)n &= \beta \angles{\tilde{\sigma}, A_{\mathop{\mathrm{sym}}} \tilde{\sigma}} - \beta \angles{\sigma_K, A_{\mathop{\mathrm{sym}}} \sigma_K} \\ &\quad + \sum_{j=1}^n \tilde{\Lambda}_{\gamma}(q_\beta^*, \sigma_{K,j}) - n \mathop{{}\mathbb{E}}[\Lambda(q_\beta^*,Y_{q_\beta^*})] \\ &\quad +\mathop{\mathrm{obj}}(q_\beta^*,\sigma_K) - \mathop{\mathrm{obj}}(0,0). \end{align}\] On the right-hand side, the first line is estimated by 97 , the second line is estimated by 96 , and the third line is estimated by , which yields \[\begin{align} \beta \angles{\tilde{\sigma}, A_{\mathop{\mathrm{sym}}} \tilde{\sigma}} - \beta \mathcal{E}(\beta)n &\geqslant-n M_2 \beta \left( d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K), \mathop{\mathrm{dist}}(Y_{q_\beta^*}^\gamma)) + \sqrt{2} e^{5 \beta^2} \gamma \right) \\ &\quad -n M_3 \beta^2 \left( \left( d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K), \mathop{\mathrm{dist}}(Y_{q_\beta^*}^\gamma)) + \sqrt{2} e^{5 \beta^2} \gamma \right)^{8\beta^2 / (1 + 8 \beta^2)} + (2 \beta^2 \gamma)^{4 \beta^2/(1 + 4 \beta^2)} \right) \\ &\quad -n \cdot e^{C_4 \beta^2} \left( M_4 \gamma^{-1} \eta^{-1/2} n^{-1/36} + M_5 \gamma^{-2} \eta^{1/2} + M_6 n^{-\delta/2} + M_7 \gamma \right), \end{align}\] where we have estimated \((1 + \beta^2)^{8\beta^2 / (1 + 8 \beta^2)}\) by a constant times \(\beta^2\). To combine these estimates conveniently, note that the \(d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K),\mathop{\mathrm{dist}}(Y_{q_\beta^*}^{\gamma}))\) and \(\gamma\) terms on the first line can be absorbed into the estimate on the third line from , in the same way as we did in the proof of that theorem. For the second line, we write \[\left( d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K), \mathop{\mathrm{dist}}(Y_{q_\beta^*}^\gamma)) + \sqrt{2} e^{5 \beta^2} \gamma \right)^{8\beta^2 / (1 + 8 \beta^2)} \leqslant M_8 d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K), \mathop{\mathrm{dist}}(Y_{q_\beta^*}^\gamma))^{8\beta^2 / (1 + 8 \beta^2)} + M_9 (e^{5 \beta^2} \gamma)^{8\beta^2 / (1 + 8 \beta^2)}.\] Since \(\gamma \leqslant 1\), and since \(4 \beta^2 / (1 + 4 \beta^2) \leqslant 8\beta^2/(1 + 8 \beta^2)\), we can thus write \[M_9 (e^{5 \beta^2} \gamma)^{8\beta^2 / (1 + 8 \beta^2)} + (2 \beta^2 \gamma)^{4 \beta^2/(1 + 4 \beta^2)} \leqslant e^{M_{10} \beta^2} \gamma^{4 \beta^2/(1 + 4 \beta^2)}.\] We finally combine this with the \(\gamma\) terms on the third line, and altogether the error is bounded by \[\begin{gather} n \biggl( M_{11} d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K),\mathop{\mathrm{dist}}(Y_{q_\beta^*}^\gamma))^{8\beta^2/(1 + 8 \beta^2)} \\ + e^{C_4 \beta^2} \left( M_4 \gamma^{-1} \eta^{-1/2} n^{-1/36} + M_5 \gamma^{-2} \eta^{1/2} + M_6 n^{-\delta/2} + M_7 \gamma^{4 \beta^2/(1+4\beta^2)} \right) \biggr). \end{gather}\] Then we estimate \(d_{W,2}(\mathop{\mathrm{emp}}(\sigma_K),\mathop{\mathrm{dist}}(Y_{q_\beta^*}^\gamma))\) by and combine terms by similar reasoning as did before. This yields the estimate asserted in the theorem after dividing by \(\beta\). ◻

It remains to choose appropriate values of \(\beta\), \(\gamma\), \(\eta\), and \(n\) to the make the final error smaller than \(\varepsilon\).

Corollary 8 (Quantitative lower bound on energy approximation).
Suppose . Let \(\sigma_K\) be the final iterate of , with \(\tilde{\sigma} = \mathsf{trunc}\left(\sigma_K\right)\) as defined in . Furthermore, assume we are in the high-probability events of , , , and . Then, with \(\mathcal{P}_{\beta}\) as in 5 , \[\frac{1}{n}H(\sigma^*) \geqslant\frac{\mathcal{P}_\beta}{\beta} - \frac{\varepsilon}{5} - O(\varepsilon^2) - O_{\varepsilon}(n^{-\alpha})\,,\] provided that \(\alpha = \min\left(\delta/4, 1/24\right)\), \(\beta = \frac{10}{\varepsilon}\), \(\eta = e^{-C\beta^2}\), \(\gamma = \eta^{1/8} = e^{-C\beta^2/8}\), and \(n\geqslant\eta^{-90}\) for some sufficiently large absolute constant \(C > 0\), where \(\sigma^* = \mathsf{round}(\tilde{\sigma})\) with the rounding procedure described in .

Consequently, \[H(\sigma^*)\geqslant \left(1-\frac{\varepsilon}{2}-O(\varepsilon^2)-o_n(1)\right) \max_{\sigma\in\{\pm1\}^n}H(\sigma)\] with probability \(1-\exp(-n^{\Omega(1)})\) as \(n\to\infty\).

Proof. We now invoke  in conjunction with a precise set of parameterizations for \(\beta\), \(\eta\) and \(\gamma\) as functions of \(\varepsilon\in (0,1/2)\) that give the desired approximation to the energy, with high probability under the choice of input and the algorithm’s internal randomness. Direct computation shows that for large enough \(C\), our assumptions on \(\beta\), \(\eta\), \(\gamma\), and \(n\) imply the hypothesis 59 in Theorem 27, so that we can apply Theorem 31.

By  with our choices of \(\beta\), \(\eta\), and \(\gamma\), \[\frac{1}{n}H(\tilde{\sigma}) \geqslant\mathcal{E}(\beta) - O(\varepsilon^{2}) - O_{\varepsilon}(n^{-\alpha}).\] Set \(\sigma^* = \mathsf{round}(\tilde{\sigma})\) with \(\mathsf{round}(\cdot)\) defined as in . Then \[\frac{1}{n}H(\sigma^*) \geqslant_{\text{w.h.p. as in \prettyref{prop:rounding-energy}}} \frac{1}{n}H(\tilde{\sigma}) - \frac{O(1)}{n^\alpha}\,.\]

Putting the previous two estimates together with \(\alpha = \min\left(\delta/4,1/24\right)\) gives \[\begin{align} \frac{1}{n}H(\sigma^*) \geqslant\mathcal{E}(\beta) - \frac{O(1)}{n^\alpha} - O(\varepsilon^{2}) - O_{\varepsilon}(n^{-\alpha})\,. \end{align}\] Using the elementary log-sum-exp bound \[\frac{\mathcal{P}_{\beta}}{\beta} \leqslant\lim_{n \to \infty}\mathop{{}\mathbb{E}}\frac{1}{n}\max_{\sigma\in\{\pm1\}^n}H(\sigma) + \frac{\log 2}{\beta}\] together with [9], a simple integration-by-parts in conjunction with [9] (after adjusting for our normalization of the Hamiltonian) and \(\beta = \frac{10}{\varepsilon}\) so that \((2 \log 2 + 1/2)\varepsilon/10 < \varepsilon/5\) imply that, under fRSB, \[\mathcal{E}(\beta) = 2\beta\left(\int_0^{q^*_\beta}\left(\int_t^1 F_\mu(s)ds\right)dt\right) \geqslant\frac{\mathcal{P}_\beta}{\beta} - \frac{\varepsilon}{5}.\] Combining this with the previous inequality yields our first conclusion: \[\frac{1}{n}H(\sigma^*) \geqslant\frac{\mathcal{P}_\beta}{\beta} - \frac{\varepsilon}{5} - O(\varepsilon^2) - O_{\varepsilon}(n^{-\alpha}).\]

Then with the elementary log-sum-exp bound \[\frac{\mathcal{P}_{\beta}}{\beta} \geqslant\lim_{n \to \infty}\mathop{{}\mathbb{E}}\frac{1}{n}\max_{\sigma\in\{\pm1\}^n}H(\sigma) = \lim_{\beta \to \infty} \frac{\mathcal{P}_{\beta}}{\beta}.\] combined with  and the concentration of the ground state energy [49], we ultimately obtain with high probability \[\frac{1}{n}H(\sigma^*) \geqslant\frac{1}{n}\max_{\sigma \in \{-1,1\}^n} H(\sigma) - \frac{\varepsilon}{5} - O(\varepsilon^2) - o_n(1).\] By using the simple bound \(\lim_{n\to\infty}\mathop{{}\mathbb{E}}\frac{1}{n}\max_{\sigma \in \{-1,1\}^n} H(\sigma) \geqslant 1/2\) from  [9] (again, after adjusting for our normalization of the Hamiltonian), we obtain the multiplicative form \[H(\sigma^*)\geqslant \left(1-\frac{\varepsilon}{2}-O(\varepsilon^2)-o_n(1)\right) \max_{\sigma\in\{\pm1\}^n}H(\sigma). \qedhere\] ◻

7 Discussion & Open Problems↩︎

We present a few natural directions for further inquiry given the successful initiation of Subag’s algorithmic program to the domain of the hypercube.

7.1 Extending to higher-degree polynomials↩︎

It is natural to apply the PHA algorithmic framework on higher-degree random polynomials over the hypercube, such as mixed \(p\)-spin models. A mixed \(p\)-spin model is a non-homogeneous random degree-\(p\) polynomial defined as, \[H_p(\sigma) := \sum_{k \geqslant 2}^{p}\frac{\gamma_k}{n^{(k-1)/2}}\left\langle A^{(k)}, \sigma^{\otimes k}\right\rangle\, ,\] where \(\{A^{(k)}\}_{k \in [p]}\) is a family of independent order-\(k\) Gaussian tensors and \(\{\gamma_k \geqslant 0\}_{k \in[p]}\) are non-negative constants that weight the model. The goal is to compute, \[\max_{\sigma \in \{\pm 1\}^n} H_p(\sigma)\, .\] This maximum value is captured, almost surely, by a modest generalization of the Parisi formula for the SK model [23], [44], [85]. There is an existing AMP algorithm [10] that achieves an approximation ratio given by a natural relaxation of the Parisi formula. We suspect that some more technical work extending the analysis in §5 and §6 to work with a potential function for the more general model [2] can be combined with estimates of spectral convergence for the Hessian of \(H_p\) provided by Subag [1] to straightforwardly generalize the PHA algorithm.

7.2 Relaxed Parisi formula and high-entropy step sum-of-squares hierarchy↩︎

High-entropy step (HES) distributions are defined as Euler-Maruyama discretizations of a family of sufficiently regular Itô processes without drift. In a recent result, it was demonstrated that if \(H^{\mathrm{sp}}(\sigma)\) is a spherical spin glass Hamiltonian in the fRSB regime, then the maximum expected value of \(H^{\mathrm{sp}}(\sigma)\) attained by any HES process over \(\sigma\) can be certified by a low-degree SoS proof (up to constant factors) [6].

The constant factors the certificates are off by arise due to imprecise estimates on the higher-order derivatives of the Hamiltonian. An approach is suggested by the authors to strengthen the bound [6]. Roughly, the approach would involve removing a wasteful \(\ell^{\infty}\) bound to account for the Fourier mass restricted to fixed levels when bounding the nuclear norm of the cumulants of the HES distribution [6]. The feasibility argument of [6] critically uses the fact that Subag’s Hessian ascent can be randomized over a HES distribution, though, the challenging part is demonstrating that no HES process can do better.

Note that the PHA algorithm is quite similar to a HES distribution (over the cube), as it is generated by a specific discretized Itô process. To instantiate a HES SoS hierarchy over the hypercube [6] it is crucial to have a HES distribution to demonstrate feasibility for the same problem over the hypercube. The PHA algorithm provides this HES distribution conditioned on there existing a low-degree matrix polynomial in the prior iterates that approximates \(Q^2(t_k,\sigma_k)\) for every \(k \in [K]\). The systematic control over the maximum of the bulk spectrum and Frobenius norm of \(Q^2(t_k,\sigma_k)\) provided by  suggests that this is likely the case with a large constant choice of degree and an appropriate choice of \(\delta\) to curtail the operator norm (see [6]).

Additionally, the fact that the objective function’s fluctuations can, when expressed via Taylor expansion-type arguments, be certified via concentration of measure and regularity arguments over the high-entropy process induced by the PHA algorithm is yet more evidence that the technique of proofs in [6] can possibly be adapted to work on the cube. Therefore, provided that a low-degree matrix polynomial approximation for \(Q^2(t_k,\sigma_k)\) exists, it is likely possible to instantiate the SoS HES hierarchy on the hypercube to provide low-degree SoS certificates for the relaxed Parisi formula by combining the techniques in [6] with combinatorial techniques from free probability theory [86].

8 Itô Calculus↩︎

In the energy analysis of the algorithm, convergence to a primal version of the Auffinger-Chen representation () plays a crucial role. The representation is a stochastic restatement of the Parisi Variational-Principle () in terms of functions of a drifted Brownian motion. We briefly state elementary results in Itô calculus that we utilize; for a detailed review see [87].

Itô processes are closed under twice-differentiable functions.

Definition 6 (Itô formula). Let \(\{X_t\}_{t \in \mathbb{R}_{\geqslant 0}}\) be an Itô process that satisfies the SDE, \[dX_t = f(t, X_t)dt + g(t, X_t)dW_t\, .\] Further, suppose that \(h \in C^{1,2}(\mathbb{R}_{\geqslant 0},\mathbb{R})\). Then, \[dh(t,X_t) = \left(\partial_t h(t,X_t) + f(t,X_t)\partial_x h(t,X_t) + \frac{g(t,X_t)^2}{2}\partial_{x,x}h(t,X_t)\right)dt + g(t,X_t) \partial_x h(t,X_t) dW_t\,.\]

An important fact is that the set of self-driven Brownian motions follow a Plancherel-like identity known as the Itô isometry.

Lemma 40 (Itô isometry). Given an Itô process \(\{X_t\}_{t \in \mathbb{R}_{\geqslant 0}}\), the Itô integral forms an isometry of normed vector spaces induced by the functional inner product over the vector space of square-integrable functions. Consequently, \[\mathop{{}\mathbb{E}}\left[\left(\int_0^t X_s dW_s\right)^2\right] = \mathop{{}\mathbb{E}}\left[\int_0^t X^2_s ds\right]\,.\]

9 Wasserstein Metrics↩︎

Denote by \(\operatorname{Lip}_b(M)\) the set of bounded Lipschitz functions over a metric space \((M,d)\). Let, \[\left\lVert{f}\right\rVert_{\operatorname{Lip}} := \sup_{x \neq y}\frac{\|f(x) - f(y)\|}{d(x,y)}\,.\] Below we review some elementary statements about Wasserstein metrics from the theory of optimal transport; see [88] for a detailed treatment.

Definition 7 (Wasserstein metric [88]).
Fix \(p \in [1,\infty]\). Given a metric space \((M,d)\) with probability measures \(\mu\) and \(\nu\) on \(M\) with \(p\)-finite moments, the Wasserstein \(p\) metric is given as, \[d_{W,p}(\mu,\nu) = \inf_{\pi \in \Gamma(\mu,\nu)}\left(\mathop{{}\mathbb{E}}_{(x,y) \sim \pi} d(x,y)^p\right)^{1/p}\,,\] where \(\Gamma(\mu,\nu)\) denotes the set of all probability measures \(\pi\) on \(M \times M\) that satisfy \[\pi(S \times M) = \mu(S)\,,\text{ and, } \pi(M \times S) = \nu(S)\, ,\] for every measurable set \(S\).

The Kantorovich-Rubinstein duality gives a dual characterization of \(d_{W,1}\) in terms of a variational problem over Lipschitz pushforwards of the underlying measures.

Theorem 33 (Kantorovich-Rubinstein Duality, special case of [88]).
Let \((M,d)\) be a metric space, and \(\mu\) and \(\nu\) be measures over \(M\) with finite first moment. Then, \[d_{W,1}\left(\mu,\nu\right) = \inf_{\pi \in \Gamma(\mu,\nu)}\left(\mathop{{}\mathbb{E}}_{(x,y) \sim \pi} d(x,y)\right) = \sup_{(f,g)\,\in I_d}\left(\mathop{{}\mathbb{E}}_{x\sim\mu}\left[f(x)\right] + \mathop{{}\mathbb{E}}_{y\sim\nu}\left[g(y)\right]\right)\,,\] where, \[I_d := \left\{(f,g)\,\in \operatorname{Lip}(M) \times \operatorname{Lip}(M)\mid f(x) + g(y) \leqslant d(x,y)\right\}\,.\]

By fixing a base point \(x_0 \in M\) and choosing \(k(x) = \inf_{y \in M}\left(d(x,y) - g(y)\right) - \inf_{z \in M}\left(d(x_0,z) - g(z)\right)\) we see that finite first moments imply that \(k\in L^1(\mu)\cap L^1(\nu)\) and that it satisfies \(k(x_0) = 0\). Hence, an easy corollary is that \((f,g)\) can be replaced by \((k,-k)\) without decreasing the dual value, and so the supremum can be taken over all \(1\)-Lipschitz functions over \((M,d)\) that evaluate to 0 at the base point \(x_0\).

Corollary 9 (Monge-Kantorovich-Rubinstein duality (simplified)).
Let \(\mu\) and \(\nu\) be as in . Then, \[d_{W,1}\left(\mu,\nu\right) = \sup_{f(x_0) = 0,\, \left\lVert{f}\right\rVert_{\operatorname{Lip}} \leqslant 1} \left| \int f\,d\mu - \int f\,d\nu \right|\]

We will also use the comparison between Wasserstein distances, especially the \(L^1\) and \(L^2\) Wasserstein distances.

Lemma 41 (Comparison of Wasserstein distances). If \(p \leqslant q\), then \(d_{W,p}(\mu,\nu) \leqslant d_{W,q}(\mu,\nu)\). On the other hand, suppose that \(\mu\) and \(\nu\) are supported in a set \(S\) of diameter at most \(K\). Then \(d_{W,q}(\mu,\nu) \leqslant K^{1-p/q} d_{W,p}(\mu,\nu)^{p/q}\).

Proof. Fix a transport plan \(\pi\) between \(\mu\) and \(\nu\). Then we have by Hölder’s inequality that \[\left( \mathbb{E}_{(x,y) \sim \pi} d(x,y)^p \right)^{1/p} \leqslant\left( \mathbb{E}_{(x,y) \sim \pi} d(x,y)^q \right)^{1/q}.\] Hence, taking the infimum over \(\pi\) proves the first claim. Similarly, if \(\mu\) and \(\nu\) are supported in a set of diameter \(K\), then \[\mathbb{E}_{(x,y) \sim \pi} d(x,y)^q \leqslant K^{q-p} \mathbb{E}_{(x,y) \sim \pi} d(x,y)^p.\] Taking the \(1/q\) power and taking the infimum over \(\pi\) completes the proof. ◻

10 Ruelle Probability Cascades↩︎

To aid the reader in translating the Ruelle Probability Cascades into an expression for the solution \(\Phi\) of the Parisi PDE, we briefly recall a critical part of the construction of the Ruelle Probability Cascades, as well as the Hopf-Cole transformation which provides an explicit formula for \(\Phi\) when the measure \(\mu\) is atomic; see [23] and [44] for a detailed overview. We focus on expressing the Parisi solution at a fixed time \(t_0\) when \(\mu\) is a finitely supported measure. The measure \(\mu\) here is not necessarily the Parisi minimizer, but simply an arbitrary input probability measure for the Parisi differential equation. Since the Parisi equation is solved backwards in time, the solution at time \(t_0\) only depends on \(\mu|_{(t_0,1]}\), and hence we set up the following notation.

Notation 34 (Atomic measure \(\mu\)).
Let \(0 \leqslant t_0 < \dots < t_r = 1\), and fix a finitely supported probability measure \(\mu\) with \(\mathrm{supp}(\mu) \subseteq [0,t_0] \cup \{t_1,\dots,t_r\}\). Let \(\zeta_j = \mu([0,t_j])\), and write \(\mu|_{(t_0,1]}\) as \(\sum_{j=1}^r (\zeta_j - \zeta_{j-1}) \delta_{t_j}\).

Below, we detail a part of the construction of the RPCs that assigns a random measure to the leaves of a certain \(\infty\)-ary tree.

Definition 8 (\(\operatorname{RPC}(\mu)\) [23]).
Let \(\mu\) be an atomic measure as defined in to parameterize the RPC tree. Fix \(\mathcal{A} = \bigsqcup_{j=0}^r \mathbb{N}^j\) to represent an \(\infty\)-ary tree of depth \(r+1\) rooted at \(\emptyset\), with every node \(\pi \in \mathbb{N}^j\) uniquely specified by a path \(p(\pi) = (\emptyset, \alpha_1, (\alpha_1,\alpha_2),\dots,(\alpha_1,\dots,\alpha_j))\). Now, to every vertex \(\pi \in \mathbb{N}^j\) for \(0 \leqslant j < r\) associate an independent Poisson point process with mean measure \(\rho_j(dx) = \zeta_jx^{-1-\zeta_j}dx\) and arrange all arriving points in decreasing order \(u_{\pi 1},u_{\pi 2},\dots\) to associate with the children of \(\pi\). Finally, define a random measure on the leaves of the tree \(\{\alpha\}_{\alpha \in \mathbb{N}^r}\) as follows, \[v_\alpha = \frac{\prod_{\beta \in p(\alpha)}u_\beta}{\sum_{\gamma \in \mathbb{N}^r}\prod_{\beta \in p(\gamma)}u_\beta}\,.\]

Standard arguments dictate that the probability mass \(v_\alpha\) assigned to every leaf \(\alpha \in \mathbb{N}^r\) is almost-surely finite (see [23]). Asymptotically, when the parameter \(\mu\) is taken to be the Parisi measure \(\mu_\beta\) from , the random measure on the leaves of \(\operatorname{RPC}(\mu_{\beta})\) approaches the Gibbs measure up to orthogonal transformation. Each vertex in the RPC tree has a geometric location constructed via a weighted linear combination over a complete orthonormal basis for the underlying separable Hilbert space, with the weighting decided by the points \(\{t_j\}_{j=0}^r\) in the support of the atomic measure \(\mu\). However, this does not affect the proof of the following lemmata, and is therefore omitted; see [23] for more details.

Lemma 42 (Hopf-Cole transformation).
Let \(\mu\) be a finitely supported measure and use . Let \(\Phi\) be the solution to the Parisi PDE associated to \(\mu\). Let \(Z_1\), …, \(Z_r\) be independent normal random variables with \(E[Z_{j+1}] = 0\) and \(\mathop{{}\boldsymbol{\mathrm{Var}}}(Z_{j+1}) = 2 \beta^2 (t_{j+1} - t_j)\) for every \(0 \leqslant j \leqslant r-1\). Then \[\Phi(t_j,x) = \begin{cases} \frac{1}{\zeta_j} \log \mathop{{}\mathbb{E}}\exp(\zeta_j \Phi(t_{j+1},x+Z_{j+1}) ), & \zeta_j > 0 \\ \mathbb{E} \Phi(t_{j+1},x+Z_{j+1}), & \zeta_j = 0 \end{cases}\]

Proof. On the interval \([t_j, t_{j+1}]\), we want \[\partial_t \Phi(t,x) = -\beta^2 (\partial_{x,x} \Phi(t,x) + \zeta_j \partial_x \Phi(t,x)^2).\] Hence, in the case where \(\zeta_j > 0\), \[\begin{align} \partial_t \exp( \zeta_j \Phi(t,x) ) &= \zeta_j \partial_t \Phi(t,x) \exp(\zeta_j \Phi(t,x)) \\ &= -\zeta_j \beta^2 (\partial_{x,x} \Phi(t,x) + \zeta_j \partial_x \Phi(t,x)^2) \exp(\zeta_j \Phi(t,x)) \\ &= -\beta^2 [\zeta_j \partial_{x,x} \Phi(t,x) + \zeta_j^2 \partial_x \Phi(t,x)^2]\exp\left(\zeta_j\Phi(t,x)\right) \\ &= -\beta^2 \partial_{x,x} [\exp(\zeta_j \Phi(t,x))]. \end{align}\] Recall from that \(|\partial_x \Phi(t,x)| \leqslant 1\) so \(\Phi\) grows linearly, and it is not hard to see that the solution to the heat equation is unique for a function that grows at most exponentially. Therefore, for \(t \in [t_j,t_{j+1}]\), we have \[\exp( \zeta_j \Phi(t,x) ) = \frac{1}{\sqrt{4\pi \beta^2 (t_{j+1} - t)}} \int_{\mathbb{R}} \exp(\zeta_j \Phi(t_{j+1},x+z)) \exp\left(-\frac{z^2}{4 \beta^2(t_{j+1} - t)} \right) \,dz.\] Hence, if \(Z_{j+1}\) is a normal random variable of mean zero and variance \(2 \beta^2(t_{j+1} - t_j)\), then we have \[\exp( \zeta_j \Phi(t_j,x) ) = \mathop{{}\mathbb{E}}[\exp(\zeta_j \Phi(t_{j+1},x+Z_{j+1}))],\] which is equivalent to the asserted formula. In the case when \(\zeta_j = 0\), the formula reduces to the standard probabilistic formula for solutions to the heat equation. ◻

Below, we briefly sketch how the Ruelle Probability Cascades can be used to construct a solution to provide a specific representation to the Hopf-Cole transformed expression for \(\Phi(t_0,x)\).

Lemma 43 (RPC based representation of \(\Phi(t_0,x)\)).
Let \(\mu\) be a finitely supported measure on \([0,1]\), let \(t_0 \in [0,1)\) and use . Let \(\mathcal{A} = \bigsqcup_{j = 0}^{r} \mathbb{N}^j\). Then there exist nonnegative random variables \((v_\alpha)_{\alpha \in \mathbb{N}^r}\) such that \(\sum_{\alpha \in \mathbb{N}^r} v_\alpha = 1\), and there exist Gaussian random variables \((Z_\alpha)_{\alpha \in \mathbb{N}^r}\) with \(\mathop{{}\mathbb{E}}Z_\alpha = 0\) and \(\mathop{{}\boldsymbol{\mathrm{Var}}}(Z_\alpha) = 2 \beta^2(1 - t_0)\) such that for all \(x\), \[\Phi(t_0,x) = \mathop{{}\mathbb{E}}\log \sum_{\alpha \in \mathbb{N}^r} 2 v_\alpha \cosh(x + Z_\alpha).\]

Proof. First consider the case where \(\mu([0,t_0]) > 0\) so that all the \(\zeta_j\)’s are strictly positive. Let \(Z_1\), …, \(Z_r\) be independent normal random variables with \(\mathop{{}\mathbb{E}}[Z_{j+1}] = 0\) and \(\mathop{{}\boldsymbol{\mathrm{Var}}}(Z_{j+1}) = 2 \beta^2 (t_{j+1} - t_j)\). For compatibility with the notation of [23], recall that we can express \(Z_j\) as a function of a uniform random variable \(\omega_j\) on \([0,1]\) so that \(\omega_1\), …, \(\omega_r\) are independent. Let \(X_j\) be the random variable \[X_j(\omega_1,\dots,\omega_j) = \Phi(t_j, x + Z_1(\omega_1) + \dots + Z_j(\omega_j)),\] so that \[X_j(\omega_1,\dots,\omega_j) = \frac{1}{\zeta_j} \log \mathop{{}\mathbb{E}}\left[ \exp(\zeta_j X_{j+1}(\omega_1,\dots,\omega_{j+1})) \mid \omega_1, \dots, \omega_j \right]\] when \(\zeta_j > 0\) and \(X_j(\omega_1,\dots,\omega_j) =\mathbb{E} [X_{j+1}(\omega_1,\dots,\omega_{j+1}) \mid \omega_1, \dots, \omega_j]\) if \(\zeta_j = 0\). Now consider independent uniform random variables \(\omega_\alpha\) in \([0,1]\) indexed by strings \(\alpha = (n_1,\dots,n_j)\) of natural numbers, i.e.the index set is \(\mathcal{A} = \bigsqcup_{j=0}^\infty \mathbb{N}^j\). For \(\alpha = (n_1,\dots,n_r) \in \mathbb{N}^r\), write \[\Omega_\alpha = (\omega_{(n_1)}, \omega_{(n_1,n_2)},\dots,\omega_{(n_1,\dots,n_r)}).\] Let \((v_\alpha)_{\alpha \in \mathcal{A}}\) be the random weights constructed using the RPCs (). In particular, \(\sum_{\alpha \in \mathbb{N}^r} v_\alpha = 1\). Then [23] shows that \[X_0 = \mathop{{}\mathbb{E}}\log \sum_{\alpha \in \mathbb{N}^r} v_\alpha \exp X_r(\Omega_\alpha).\] In other words, \[\Phi(t_0,x) = \mathop{{}\mathbb{E}}\log \sum_{(n_1,\dots,n_r) \in \mathbb{N}^r} v_{(n_1,\dots,n_r)} \exp\left( \Phi(1,x+Z_1(\omega_{n_1}) + \dots Z_r(\omega_{(n_1,\dots,n_r)})) \right).\] Writing \[Z_\alpha(\Omega_\alpha) = Z_1(\omega_{n_1}) + \dots + Z_r(\omega_{(n_1,\dots,n_r)}),\] we see that \(Z_\alpha\) is normal with mean zero and variance \(\sum_{j=0}^{r-1} 2\beta^2 (t_{j+1} - t_j) = 2\beta^2 (1 - t_0)\). Finally, note \(\exp(\Phi(1,\cdot)) = 2 \cosh(\cdot)\), which proves the asserted formula.

In the case where \(\mu([0,t_0]) = 0\), note that \(t_1\) is the infimum of the support of \(\mu\). Then by the preceding argument, we can express \[\Phi(t_1,x) = \mathop{{}\mathbb{E}}\log \sum_{\alpha \in \mathbb{N}^r} 2 v_\alpha \cosh(x + Z_\alpha),\] where \(Z_\alpha\) is a Gaussian random variable of variance \(2 \beta^2(1 - t_1)\). Recall \(\Phi\) satisfies \(\partial_t \Phi = -\beta^2 \partial_{x,x} \Phi\) on \([t_0,t_1]\). Let \(Z\) be a normal random variable of mean zero and variance \(2 \beta^2 (t_1 - t_0)\) independent of all the other random variables. Then \[\begin{align} \Phi(t_0,x) &= \mathop{{}\mathbb{E}}\Phi(t_1,x+Z) \\ &= \mathop{{}\mathbb{E}}\log \sum_{\alpha \in \mathbb{N}^r} 2 v_\alpha \cosh(x + Z + Z_\alpha). \end{align}\] Hence, we have the asserted formula at \(t_0\) using \(\tilde{Z}_\alpha = Z + Z_\alpha\) which is normal of mean zero and variance \(2 \beta^2(1 - t_1) + 2 \beta^2(t_1 - t_0) = 2 \beta^2(1 - t_0)\). ◻

Acknowledgements↩︎

DJ & JS thank the union of postdocs and academic researchers at University of California for introducing them to each other.

JS & JSS thank Chris Jones for a stimulating discussion on the algorithmic applications of Brownian motion. JS & JSS are grateful to Todd Kemp for an instructive and helpful discussion in Summer-2023, and DJ is grateful to Kemp for his advice as postdoc supervisor in 2020-2023. JSS is indebted to Brice Huang for a particularly inspiring & informative discussion at Harvard University in Fall-2022 which motivated the formulation of a Hessian ascent framework on the hypercube, and for comments on a draft of this paper.

The initial development of this work was conducted in Summer-2023 jointly at the Halıcıoğlu Data Science Institute and the department of mathematics at the University of California San Diego during a set of particularly fruitful discussions between all the authors.

JSS is extremely grateful to Alexandra Kolla for her advice and support as postdoc supervisor. JSS did this work in part as a visiting scholar at the Flatiron Institute hosted by Prof. SueYeon Chung.

The authors are grateful to the anonymous referees for their careful review and detailed comments.

Funding↩︎

DJ acknowledges funding from the National Science Foundation (US), grant DMS-2002826; the National Sciences and Engineering Research Council (Canada), grant RGPIN-2017-05650; Denmark’s Independent Research Fund, grant 1026-00371B; and the Horizon EU Marie Skłodowska-Curie Action FREEINFOGEOM, grant 101209517. JS acknowledges support from National Science Foundation (US) awards #2217058 and #2112665. JSS acknowledges partial support from Defense Advanced Research Projects Agency ONISQ program award HR001120C0068 during the early stages of this work.

Competing interests↩︎

The authors have no relevant financial or non-financial interests to disclose.

Data availability↩︎

No datasets were generated or analysed during the current study.

References↩︎

[1]
E. Subag, “Following the ground states of full-RSB spherical spin glasses,” Communications on Pure and Applied Mathematics, vol. 74, no. 5, pp. 1021–1044, 2021.
[2]
W.-K. Chen, D. Panchenko, and E. Subag, “Generalized TAP free energy,” Communications on Pure and Applied Mathematics, 2018.
[3]
E. Subag, “Free energy landscapes in spherical spin glasses,” Duke Mathematical Journal, vol. 173, no. 7, pp. 1291–1357, May 2024, doi: 10.1215/00127094-2023-0033.
[4]
A. Auffinger and W.-K. Chen, “The Parisi formula has a unique minimizer,” Communications in Mathematical Physics, vol. 335, no. 3, pp. 1429–1444, 2015.
[5]
A. Auffinger and W.-K. Chen, “On properties of Parisi measures,” Probability Theory and Related Fields, vol. 161, no. 3, pp. 817–850, 2015.
[6]
J. S. Sandhu and J. Shi, “Sum-of-squares & Gaussian processes I: Certification,” arXiv preprint arXiv:2401.14383, 2024.
[7]
B. Huang and M. Sellke, “Tight Lipschitz hardness for optimizing mean field spin glasses,” in 2022 IEEE 63rd annual symposium on foundations of computer science (FOCS), 2022, pp. 312–322.
[8]
C. Jones, K. Marwaha, J. S. Sandhu, and J. Shi, “Random Max-CSPs inherit algorithmic hardness from spin glasses,” in 14th innovations in theoretical computer science conference (ITCS 2023), 2023.
[9]
A. Montanari, “Optimization of the Sherrington-Kirkpatrick Hamiltonian,” in 2019 IEEE 60th annual symposium on foundations of computer science (FOCS), 2019, pp. 1417–1433, doi: 10.1109/FOCS.2019.00087.
[10]
A. El Alaoui, A. Montanari, and M. Sellke, “Optimization of mean-field spin glasses,” The Annals of Probability, vol. 49, no. 6, pp. 2922–2960, 2021.
[11]
A. Montanari, “Optimization of the SK model, LIDS student seminar, MIT.” Youtube; https://www.youtube.com/watch?v=kPMYRHyy8V0, 2019.
[12]
G. Parisi, “A sequence of approximated solutions to the SK model for spin glasses,” Journal of Physics A: Mathematical and General, vol. 13, no. 4, p. L115, 1980.
[13]
A. Auffinger, W.-K. Chen, and Q. Zeng, “The SK model is infinite step replica symmetry breaking at zero temperature,” Communications on Pure and Applied Mathematics, vol. 73, no. 5, 2020.
[14]
S. Gufler, A. Schertzer, and M. A. Schmidt, “On the concavity of the TAP free energy in the SK model,” Stochastic Processes and their Applications, vol. 164, pp. 160–182, 2023.
[15]
A. Adhikari, C. Brennecke, P. von Soosten, and H.-T. Yau, “Dynamical approach to the TAP equations for the Sherrington–Kirkpatrick model,” Journal of Statistical Physics, vol. 183, no. 3, p. 35, 2021.
[16]
D. Sherrington and S. Kirkpatrick, “Solvable model of a spin-glass,” Physical Review Letters, vol. 35, no. 26, p. 1792, 1975.
[17]
R. J. Adler and J. E. Taylor, Random fields and geometry. New York, NY, United States: Springer, 2007.
[18]
F. Guerra and F. L. Toninelli, “The thermodynamic limit in mean field spin glass models,” Communications in Mathematical Physics, vol. 230, no. 1, pp. 71–79, 2002.
[19]
F. Guerra, “Broken replica symmetry bounds in the mean field spin glass model,” Communications in mathematical physics, vol. 233, pp. 1–12, 2003.
[20]
M. Talagrand, “The Parisi formula,” Annals of mathematics, pp. 221–263, 2006.
[21]
D. Panchenko, “The Parisi ultrametricity conjecture,” Annals of Mathematics, pp. 383–393, 2013.
[22]
D. Panchenko, “The Parisi formula for mixed \(p\)-spin models,” The Annals of Probability, vol. 42, no. 3, pp. 946–958, 2014.
[23]
D. Panchenko, The Sherrington-Kirkpatrick model. New York: Springer Science & Business Media, 2013.
[24]
A. Auffinger and W.-K. Chen, Parisi formula for the ground state energy in the mixed \(p\)-spin model,” The Annals of Probability, vol. 45, no. 6B, pp. 4617–4631, 2017, doi: 10.1214/16-AOP1173.
[25]
A. El Alaoui, A. Montanari, and M. Sellke, “Local algorithms for maximum cut and minimum bisection on locally treelike regular graphs of large degree,” Random Structures & Algorithms, vol. 63, no. 3, pp. 689–715, 2023.
[26]
A. Chen, N. Huang, and K. Marwaha, “Local algorithms and the failure of log-depth quantum advantage on sparse random CSPs,” arXiv preprint arXiv:2310.01563, 2023.
[27]
M. Bayati and A. Montanari, “The LASSO risk for Gaussian matrices,” IEEE Transactions on Information Theory, vol. 58, no. 4, pp. 1997–2017, 2011.
[28]
A. Montanari and R. Venkataramanan, Estimation of low-rank matrices via approximate message passing,” The Annals of Statistics, vol. 49, no. 1, pp. 321–345, 2021, doi: 10.1214/20-AOS1958.
[29]
M. Celentano and A. Montanari, “Fundamental barriers to high-dimensional regression with convex penalties,” The Annals of Statistics, vol. 50, no. 1, pp. 170–196, 2022.
[30]
A. El Alaoui, A. Montanari, and M. Sellke, “Sampling from the Sherrington-Kirkpatrick Gibbs measure via algorithmic stochastic localization,” in 2022 IEEE 63rd annual symposium on foundations of computer science (FOCS), 2022, pp. 323–334.
[31]
B. Huang, A. Montanari, and H. T. Pham, “Sampling from spherical spin glasses in total variation via algorithmic stochastic localization,” arXiv preprint arXiv:2404.15651, 2024.
[32]
D. L. Donoho, A. Maleki, and A. Montanari, “Message-passing algorithms for compressed sensing,” Proceedings of the National Academy of Sciences, vol. 106, no. 45, pp. 18914–18919, 2009.
[33]
D. Gamarnik and M. Sudan, “Limits of local algorithms over sparse random graphs,” in Proceedings of the 5th conference on innovations in theoretical computer science, 2014, pp. 369–376.
[34]
W.-K. Chen, D. Gamarnik, D. Panchenko, M. Rahman, et al., “Suboptimality of local algorithms for a class of max-cut problems,” Annals of Probability, vol. 47, no. 3, pp. 1587–1618, 2019.
[35]
C.-N. Chou, P. J. Love, J. S. Sandhu, and J. Shi, “Limitations of local quantum algorithms on random max-k-XOR and beyond,” in 49th international colloquium on automata, languages, and programming (ICALP 2022), 2022.
[36]
D. Gamarnik, A. Jagannath, and A. S. Wein, “Low-degree hardness of random optimization problems,” in 2020 IEEE 61st annual symposium on foundations of computer science (FOCS), 2020, pp. 131–140.
[37]
D. Gamarnik and A. Jagannath, The overlap gap property and approximate message passing algorithms for \(p\)-spin models,” The Annals of Probability, vol. 49, no. 1, pp. 180–205, 2021, doi: 10.1214/20-AOP1448.
[38]
W.-K. Chen and A. Sen, “Parisi formula, disorder chaos and fluctuation for the ground state energy in the spherical mixed p-spin models,” Communications in Mathematical Physics, vol. 350, no. 1, pp. 129–173, 2017.
[39]
A. Jagannath and I. Tobasco, “A dynamic programming approach to the Parisi functional,” Proceedings of the American Mathematical Society, vol. 144, no. 7, pp. 3135–3150, 2016.
[40]
V. Bhattiprolu, V. Guruswami, and E. Lee, “Sum-of-squares certificates for maxima of random tensors on the sphere,” in Approximation, randomization, and combinatorial optimization. Algorithms and techniques (APPROX/RANDOM 2017), 2017.
[41]
S. B. Hopkins, P. K. Kothari, A. Potechin, P. Raghavendra, T. Schramm, and D. Steurer, “The power of sum-of-squares for detecting hidden structures,” in 2017 IEEE 58th annual symposium on foundations of computer science (FOCS), 2017, pp. 720–731.
[42]
M. Ghosh, F. G. Jeronimo, C. Jones, A. Potechin, and G. Rajendran, “Sum-of-squares lower bounds for Sherrington-Kirkpatrick via planted affine planes,” in 2020 IEEE 61st annual symposium on foundations of computer science (FOCS), 2020, pp. 954–965.
[43]
B. Barak and D. Steurer, “Proofs, beliefs, and algorithms through the lens of sum-of-squares,” Course notes: http://www.sumofsquares.org/public/index.html, vol. 1, 2016.
[44]
M. Talagrand, Mean field models for spin glasses: Advanced replica-symmetry and low temperature. Heidelberg: Springer, 2011.
[45]
Y.-P. Hsieh, A. Kavis, P. Rolland, and V. Cevher, “Mirrored Langevin dynamics,” Advances in Neural Information Processing Systems, vol. 31, 2018.
[46]
J.-C. Mourrat, “Un-inverting the Parisi formula,” Ann. Inst. H. Poincaré Probab. Statist., vol. 61, no. 4, pp. 2709–2720, 2025, doi: 10.1214/24-AIHP1487.
[47]
W. K. Chen, “Variational representations for the Parisi functional and the two-dimensional Guerra-Talagrand bound,” Annals of Probability, vol. 45, no. 6, pp. 3929–3966, 2017.
[48]
G. Pleiss, M. Jankowiak, D. Eriksson, A. Damle, and J. Gardner, “Fast matrix square roots with applications to Gaussian processes and Bayesian optimization,” Advances in neural information processing systems, vol. 33, pp. 22268–22281, 2020.
[49]
R. Vershynin, High-dimensional probability: An introduction with applications in data science, vol. 47. Cambridge, United Kingdom: Cambridge University Press, 2018.
[50]
D.-V. Voiculescu, “Limit laws for random matrices and free products,” Inventiones mathematicae, vol. 104, no. 1, pp. 201–220, Dec. 1991, doi: 10.1007/BF01245072.
[51]
G. W. Anderson, A. Guionnet, and O. Zeitouni, An introduction to random matrices. Cambridge, United Kingdom: Cambridge University Press, 2010.
[52]
A. Guionnet and C. Mazza, “Long time behaviour of the solution to non-linear Kraichnan equations,” Probab. Theory Relat. Fields, vol. 131, pp. 493–518, 2005, doi: 10.1007/s00440-004-0382-7.
[53]
D. Jekel and S. Kunnawalkam Elayavalli, “Upgraded free independence phenomena for random unitaries,” Trans. Amer. Math. Soc. Ser. B, vol. 13, pp. 1–29, 2026, doi: 10.1090/btran/244.
[54]
K. Zhu, An introduction to operator algebras. Ann Arbor: CRC Press, 1993.
[55]
V. Jones and V. S. Sunder, Introduction to subfactors. Cambridge: Cambridge University Press, 1997.
[56]
C. Anantharaman-Delaroche and S. Popa, Preprint available at https://www.math.ucla.edu/~popa/Books/IIunV15.pdf“An introduction to \(\mathrm{II}_1\) factors,” 2021.
[57]
M. Takesaki, Theory of operator algebras i, vol. 124. Berlin Heidelberg: Springer-Verlag, 2002.
[58]
R. V. Kadison and J. R. Ringrose, Fundamentals of the theory of operator algebras i, vol. 15. Providence: American Mathematical Society, 1983.
[59]
B. Blackadar, Operator algebras: Theory of \({C}^*\)-algebras and von Neumann algebras, vol. 122. Berlin, Heidelberg: Springer-Verlag, 2006.
[60]
G. Pisier and Q. Xu, “Non-commutative \(L^p\)-spaces,” in Handbook of the geometry of Banach spaces, vol. 2, W. B. Johnson and J. Lindenstrauss, Eds. Amsterdam: Elsevier, 2003, pp. 1459–1517.
[61]
J. Dixmier, “Formes linéaires sur un anneau d’opérateurs,” Bulletin de la Société Mathématique de France, vol. 81, pp. 9–39, 1953.
[62]
R. C. da Silva, arXiv:1803.02390“Lecture notes on non-commutative \(L_p\)-spaces,” 2018.
[63]
D.-V. Voiculescu, Symmetries of some reduced free product \({C}^*\)-algebras,” in Operator algebras and their connections with topology and ergodic theory, H. Araki, C. C. Moore, Ş.-V. Stratila, and D.-V. Voiculescu, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 1985, pp. 556–588.
[64]
D.-V. Voiculescu, “Addition of certain non-commuting random variables,” Journal of Functional Analysis, vol. 66, no. 3, pp. 323–346, 1986, doi: 10.1016/0022-1236(86)90062-5.
[65]
D.-V. Voiculescu, K. J. Dykema, and A. Nica, Free random variables, vol. 1. Providence: American Mathematical Society, 1992.
[66]
J. A. Mingo and R. Speicher, Free probability and random matrices, vol. 35. New York: Springer, 2017.
[67]
D.-V. Voiculescu, “A strengthened asymptotic freeness result for random matrices with applications to free entropy,” International Mathematics Research Notices, vol. 1998, no. 1, pp. 41–63, 1998.
[68]
N. I. Akhiezer and I. M. Glazman, Theory of linear operators in hilbert space. New York: Ungar, 1963.
[69]
P. Biane, “Processes with free increments,” Math. Z., vol. 227, pp. 143–174, 1998.
[70]
D.-V. Voiculescu, “The analogues of entropy and of Fisher’s information measure in free probability i.” Comm. Math. Phys., vol. 155, pp. 71–92, 1993.
[71]
D.-V. Voiculescu, “Analytic subordination consequences of free Markovianity,” Indiana University Mathematics Journal, vol. 51, pp. 1161–1166, 2002, doi: 10.1512/iumj.2002.51.2252.
[72]
S. T. Belinschi, “Complex analysis methods in non-commutative probability,” PhD thesis, Indiana University, 2005.
[73]
S. T. Belinschi and H. Bercovici, “A new approach to subordination results in free probability,” J Anal Math, vol. 101, pp. 357–365, 2007, doi: 10.1007/s11854-007-0013-1.
[74]
K. J. Dykema, “Multilinear function series and transforms in free probability theory,” Advances in Mathematics, vol. 208, no. 1, pp. 351–407, 2007.
[75]
M. Ledoux, “A heat semigroup approach to concentration on the sphere and on a compact Riemannian manifold,” Geometric & Functional Analysis, vol. 2, no. 2, pp. 221–224, Jun. 1992, doi: 10.1007/BF01896974.
[76]
S. G. Bobkov and M. Ledoux, “From Brunn-Minkowski to Braskamp-Lieb and to logarithmic Sobolev inequalities,” Geom. Funct. Anal., vol. 10, pp. 1028–1052, 2000.
[77]
M. Ledoux, The concentration of measure phenomenon, vol. 89. Providence, RI: American Mathematical Society, 2001.
[78]
M. Ledoux, A remark on hypercontractivity and tail inequalities for the largest eigenvalues of random matrices,” in Séminaire de probabilités XXXVII, Berlin, Heidelberg: Springer Berlin Heidelberg, 2003, pp. 360–369.
[79]
B. Collins, A. Guionnet, and F. Parraud, “On the operator norm of non-commutative polynomials in deterministic matrices and iid GUE matrices,” Cambridge Journal of Mathematics, vol. 10, no. 1, pp. 195–260, 2022.
[80]
A. Zvonkin, “Matrix integrals and map enumeration: An accessible introduction,” Mathematical and Computer Modelling, vol. 26, no. 8, pp. 281–304, 1997, doi: 10.1016/S0895-7177(97)00210-0.
[81]
F. Parraud, “Asymptotic expansion of smooth functions in polynomials in deterministic matrices and iid GUE matrices,” Communications in Mathematical Physics, vol. 399, no. 1, pp. 249–294, 2023, doi: 10.1007/s00220-022-04551-2.
[82]
A. Bandeira, M. Boedihardjo, and R. van Handel, “Matrix concentration inequalities and free probability,” Invent. math., vol. 234, pp. 419–487, 2023, doi: 10.1007/s00222-023-01204-6.
[83]
F. Guerra, “Sum rules for the free energy in the mean field spin glass model,” Fields Institute Communications, vol. 30, no. 11, 2001.
[84]
M. Aizenman, R. Sims, and S. L. Starr, “Extended variational principle for the Sherrington-Kirkpatrick spin-glass model,” Physical Review B, vol. 68, no. 21, p. 214403, 2003.
[85]
M. Talagrand, Mean field models for spin glasses: Volume I: Basic examples, vol. 54. Heidelberg: Springer Science & Business Media, 2010.
[86]
A. Nica and R. Speicher, Lectures on the combinatorics of free probability, vol. 13. Cambridge: Cambridge University Press, 2006.
[87]
B. Oksendal, Stochastic differential equations: An introduction with applications. Heidelberg: Springer Science & Business Media, 2013.
[88]
L. Ambrosio, E. Brué, and D. Semola, Lectures on optimal transport, vol. 130. Cham: Springer, 2021.

  1.  See for a definition.↩︎

  2. In the language of resolvents, this is equivalent to \[Q_k^2 := -2\beta\tilde{b}n^{\delta}\Pi_{\sigma_k^{\perp}}\left(\Im R(\tilde{a} - i\tilde{b};\, \nabla^2\mathop{\mathrm{obj}}(k\eta, \sigma_k))\right)\Pi_{\sigma_k^{\perp}} \,.\]↩︎

  3. We need to use the \(2\)-norm here since \((\mathsf{T}\otimes \mathop{\mathrm{id}})\) is not isometric with respect to operator norm.↩︎