Rapid mixing for Gibbs measures in Riemannian manifolds


Abstract

Langevin dynamics on Riemannian manifolds is analyzed. Conditions ensuring the existence of a suitable logarithmic Sobolev inequality (rapid mixing to the Gibbs measure) are identified. These conditions involve the curvature of the manifold, the inverse temperature, escaping directions from saddle points, and exclude barren plateaus and spurious local minima. We show that when these conditions are met, mixing times polynomial in the dimension of the manifold are achievable. This result is obtained through a relation between Langevin processes in the domain and in the image of a Riemannian submersion. Such a relation can be of independent interest.

= [line width=0.8pt]

1 Introduction↩︎

Given some sufficiently smooth function \(F : M \to \mathbb{R}\) on a Riemannian manifold \((M, g)\), and some fixed positive constant \(\beta > 0\)—usually referred to as the inverse temperature—the problem of sampling with respect to the associated Gibbs distribution \[\nu(x) := \frac{1}{Z} e^{-\beta F(x)},\] where \(Z\) is the normalizing constant of \(\nu\), \(Z := \int_M e^{-\beta F} \, \mathrm{dVol}_g\), has played an essential role in several areas. For example, in lattice gauge theory, where \(M\) represents the set of states a system can be in, and \(F\) its energy, solutions to this problem are of paramount importance. Without the ability to sample distributions of the above form we would not be able to predict masses of elementary particles or the nature of phase transitions in high-energy physics [1]. The problem of sampling \(\nu\) is also closely related to that of minimizing functions with constraints, which is crucial in numerical analysis [2], [3]. Gibbs sampling also has applications in differential privacy, the gold standard of privacy in machine learning [4][6].

More often than not, the distribution \(\nu\) cannot be sampled directly. For this reason, it is common to aim for approximate sampling through the use of some Markovian process whose fixed point is the target distribution \(\nu\), see [7][10]. This work focuses on one such process, known as the Langevin diffusion or Langevin dynamics, a continuous-time stochastic process \(X_t\) on \(M\) that combines gradient descent and Brownian motion2. This process is known to have \(\nu\) as its stationary distribution—i.e. for any initial condition, its probability density at time \(t\), \(\rho_t\), converges to \(\nu\) as \(t\) tends to infinity [12].

But convergence alone is of very limited interest. Equally important, and much more challenging, is the issue of controlling convergence rates. In the simple case where \(M\) is the Euclidean space, this issue has been studied extensively [13][16]. In that simplified scenario, the relation between convergence rates, the dimension of \(M\), the properties of \(F\) and the value of \(\beta\) is nowadays better understood. The picture is much less clear for general Riemannian manifolds.

In order to prove rapid convergence for the density of \(X_t\) toward its stationary density, a classic approach is to derive a suitable logarithmic Sobolev inequality (log-Sobolev for short), which in turn provides exponentially-decaying upper bounds for the distance between the densities \(\rho_t\) and \(\nu\) [12]. In our work, we will first prove a weaker version of the log-Sobolev inequality, known as a Poincaré inequality, which will be then tightened into a log-Sobolev inequality.

In a recent work by M. B. Li and M. A. Erdogdu [17], the Langevin diffusion process \(X_t\) on the product of \(n\)-dimensional spheres was analyzed. In that setting, they proved the existence of a useful log-Sobolev inequality, provided the function \(F\) associated with \(X_t\) verifies suitable assumptions. In this work, we will show that qualitatively similar results also hold for a much broader class of Riemannian manifolds, including those studied in [17]. In particular, our results will be valid for quotient manifolds. Our own construction of a log-Sobolev inequality borrows elements from theirs, but also involves new ideas3.

Note that one may distinguish two regimes of the Langevin diffusion, in which \(X_t\) exhibits different behaviors; namely when \(\frac{1}{\beta} \ll 1\) and when \(\frac{1}{\beta} \gg 1\)—which can be understood as the low-temperature and high-temperature regimes. In the former, the evolution of \(X_t\) is mainly dictated by the gradient of \(F\), whereas in the latter, it is mainly noise. In this work, we will aim at proving Poincaré and log-Sobolev inequalities valid in the low-temperature regime. As we shall see, the major obstacle to overcome is ensuring that \(X_t\) is able to escape the saddle points of \(F\) rapidly.

The high-temperature regime is interesting too [19], but obtaining a log-Sobolev inequality is then simple whenever the Ricci curvature of \((M, g)\) is positive. In this case, for sufficiently large values of \(\frac{1}{\beta}\), there exists some constant \(\kappa > 0\) such that \[\nabla^2F +\frac{1}{\beta} \mathrm{Ric}_g \geq \kappa g.\] This inequality, which is known as a dimension-curvature condition (cf. 10) implies that \(M\) satisfies a log-Sobolev inequality with the same constant by classic Bakry-Émery theory [12]. To be more precise, the condition \[\frac{1}{\beta} \in \Omega\Big(\frac{\lambda_{\min}(\nabla^2 F)}{\mathrm{Ric}_g}\Big)\] is sufficient to obtain the desired log-Sobolev inequality.

Lastly, note that the curvature-dimension condition also holds in the low-temperature regime whenever \(F\) is assumed to be strongly convex.

1.1 Main results↩︎

As mentioned earlier, this work is a study of rapid mixing for the Langevin diffusion on a Riemannian manifold in the low-temperature regime. It turns out that several technical difficulties arise when the global minimum of \(F\) is not unique. We will see, however, that when the multiplicity of global minima is only caused by a symmetry of \(F\), the situation improves. Such symmetries are very common in physics, as can be appreciated in the literature on lattice gauge theory or tensor networks [20], [21]. The study of lattice gauge theories and sampling and optimizing tensor networks are two potential fields of application of the results of this paper; first steps in this direction are taken in [22].

Informally, the possible symmetries of \(F\) will be factored out of the original manifold \(M\). On the resulting quotient manifold, the minimum will be unique. More precisely, let \((M, g)\) be a compact Riemannian manifold, and let \(F: M \to \mathbb{R}\) be some smooth function on \(M\). We assume that the symmetries of \(F\) can be described via a suitable action of some Lie group \(G\) on \(M\), i.e. there exists some compact and connected Lie group \(G\) acting on \(M\) and such that \[F(gx) = F(x),\] for every \(x \in M\) and every \(g \in G\). If we consider the orbit space \(M/G\), and the projection \(\pi: M \to M/G\), we can define a function \(\tilde{F}\) on \(M/G\) such that \(F = \tilde{F} \circ \pi\). Assuming that \(\tilde{F}\) has a unique global minimum, our strategy will consist in studying Langevin dynamics on \(M/G\), obtain a Poincaré inequality, and use it to derive a log-Sobolev inequality for \(\nu\) in \(M\). A key aspect of our derivation will be to endow \(M/G\) with a Riemannian metric \(h\) so as to make \(\pi\) a Riemannian submersion (cf. 8.4). We will study in parallel the Langevin dynamics on \((M, g)\) and \((M/G, h)\) associated with \(F\) and \(\tilde{F}\) respectively, i.e. those satisfying the formal4 SDEs \[\label{eq:defprocesos} dX_t = -\mathrm{grad}_g\, F(X_t) dt + \sqrt{\frac{2}{\beta}} dW_t, \quad \mathrm{and} \quad d\tilde{X}_t = -\mathrm{grad}_h\, \tilde{F}(\tilde{X}_t) dt + \sqrt{\frac{2}{\beta}} d\tilde{W}_t,\tag{1}\] where \(W_t\) and \(\tilde{W}_t\) are Brownian motions in \(M\) and \(M/G\). Such pairs of processes have attracted considerable attention in recent years [24][27].

To illustrate the foregoing discussion with an example, let us consider the function \[\require{physics} F(X) = \frac{\Tr(X^\dagger A X)}{\Tr(X^\dagger BX)},\] where \(X\) is an \(n \times m\) complex matrix, with \(n \geq m\) fixed, such that \(X^\dagger X = \mathbb{1}_m\), \(A, B\) square Hermitian matrices, and \(B\) positive-definite. Finding the minimum of this function is a well-known problem, known as trace-ratio minimization, relevant in principal component analysis [28] or in the issue of graph embedding [29]. \(F\) has a symmetry: \(F(XU) = F(X)\) for every unitary matrix \(U\) of dimension \(m\). Nevertheless, considering the action of the unitary group \(\mathrm{U}(m)\) on the space of \(n \times m\) complex matrices satisfying \(X^\dagger X = \mathbb{1}_m\) given by \[(U, X) \mapsto XU,\] we can define a projected version of \(F\) which has a unique global minimum [30]. The trace ratio function will be discussed in further details in 7.

The main results of this work are obtained when the following conditions are met.

Assumption 1 (Symmetry). There exists a compact and connected Lie group \(G\) acting freely, isometrically and smoothly on \((M, g)\).

Under 1, it is known (cf. 66) that the quotient space \(M/G\) can be endowed with a manifold structure and a Riemannian metric \(h\) for which the projection \[\pi: (M, g) \to (M/G, h),\] is a Riemannian submersion.

Next, we need two assumptions regarding the curvature of \((M, g)\) and \((M/G, h)\). The first is needed for the lifting process of the Poincaré inequality (cf. 3), the next to obtain explicit curvature-dimension condition constants.

Assumption 2 (Curvature). The fibers of \(\pi\) have non-negative Ricci curvature.

Assumption 3 (Curvature). There exists a constant \(\mathbf{K} \geq 1\) such that for every \(p\in M\), in normal coordinates centered at \(p\), \(|R_{ijk}^l(p)| \leq \mathbf{K}\) for every \(i, j, k, l \in \{1, \dotsc, \dim(M)\}\), where \(R\) denotes the Riemann curvature tensor of \((M, g)\). There also exist constants \(R_M, R_{M/G} \geq 0\) that lower bound the Ricci curvatures of \((M, g)\) and \((M/G, h)\), respectively, \[\mathrm{Ric}_h \geq -R_{M/G},\quad \mathrm{\textit{and}}\quad \mathrm{Ric}_g \geq -R_M.\]

Note that whenever the sectional curvature of \((M, g)\) is lower bounded by some constant \(\kappa\), the Ricci curvature—which is the sum of the sectional curvature associated with all the planes containing \(v\)—is lower bounded by \((\dim(M)-1)\kappa\). Furthermore, by O’Neill’s curvature formulas (cf. 21), we know that the sectional curvature of \((M/G, h)\) is also lower bounded by \(\kappa\), and so the Ricci curvature of \((M/G, h)\) is lower bounded by \((\dim(M/G)-1)\kappa\).

Next, we need to assume that the group action of \(G\) on \(M\) represents the symmetries of \(F\).

Assumption 4 (Symmetry). \(F\) is constant in the fibers of \(\pi\)—i.e. \(F(x) = F(y)\) for every \(x, y \in M\) such that \(\pi(x) = \pi(y)\).

As we mentioned earlier, whenever 4 holds, we can define a unique smooth function \(\tilde{F} : M/G \to \mathbb{R}\) such that \(\tilde{F} \circ \pi = F\).

Remark 1. Since \(M\) is compact and \(F\) is smooth, there exist constants \(A_1, A_2, A_3 \geq 1\) such that \(F\) is \(A_1\)-Lipschitz, \(\mathrm{grad}_{g}\,F\) is \(A_2\)-Lipschitz, and \(\nabla^2 F\) is \(A_3\)-Lipschitz. The same Lipschitz constants are valid for \(\tilde{F}\), as it will be shown in 81.

The following assumption is made for the sake of simplicity, and without loss of generality.

Assumption 5. \(\min_{x \in M} F(x) = 0.\)

Our derivation of a Poincaré inequality on \(M/G\) requires that \(\tilde{F}\) has a unique (global) minimum. We believe this limitation is not fundamental, but an artifact of our construction; results qualitatively similar to ours should hold when \(\tilde{F}\) has multiple minima. For \(M = \mathbb{R}^n\), it has been shown that this difficulty can be overcome. In this context, the Langevin dynamics converge fast to a mixture of Gibbs measures concentrated around the various minima of \(\tilde{F}\) [13]. One avenue to deal with local minima is to turn them into saddle points by overparametrization [31], [32]. Studying functions \(\tilde{F}: M/G \to \mathbb{R}\) with multiple minima is the subject of ongoing work and will be discussed elsewhere.

Assumption 6. \(\tilde{F}\) has a unique minimum.

This assumption implies that \(F\) has no local minima, and all of the global minima of \(F\) lay in a single fiber of \(\pi\). For simplicity, we will use the term saddle point for any critical point that is not the global minimum of \(\tilde{F}\). With a slight abuse of terminology, local maxima will also be called saddle points in this work, as their treatment in the proofs is the same.

Let \(\tilde{\mathcal{C}}\) and \(\tilde{\mathcal{S}}\) denote the set of critical and saddle points of \(\tilde{F}\), respectively. We need to assume that they are isolated and sufficiently spaced.

Assumption 7 (Critical point spacing). The critical points of \(\tilde{F}\) are isolated points separated by some distance \(D\), i.e. \[d_h(x_1, x_2) \geq D, \quad \forall x_1, x_2 \in \tilde{\mathcal{C}},\;x_1 \neq x_2,\] where \(d_h\) denotes the geodesic distance on \((M/G, h)\).

This assumption is made for technical reasons but we believe it should eventually be possible to remove it. It is meant to control the behavior of \(X_t\) when it escapes from saddle points.

One could argue that 7 is generically satisfied. In fact, Morse functions—those for which the Hessian’s kernel is zero at every critical point—defined on a compact manifold have a finite number of critical points [33]. Moreover, Morse functions form a dense open subset of \(C^\infty(M)\) when endowed with the strong topology (cf. [34]).

We believe the next assumption on \(\tilde{F}\) is necessary to guarantee rapid mixing.

Assumption 8. For every saddle point \(y \in \tilde{\mathcal{S}}\), there exists an escape direction. That is, there exists some constant \(\lambda_* \in (0, 1]\) such that for every saddle point \(y \in \tilde{\mathcal{S}}\) it holds that \[\lambda_{\min}\big(\nabla^2 \tilde{F}(y)\big) \leq -\lambda_*.\]

Assumption 9. The global minimum is an attractor for the process \(\tilde{X}_t\). In other words, the Hessian of \(\tilde{F}\) at the global minimum is non-degenerate. More precisely, at the unique minimum \(x^*\) of \(\tilde{F}\), we have \[\lambda_{\min}\big(\nabla^2 \tilde{F}(x^*)\big) \geq \lambda_*.\] where \(\lambda_*\) is the same constant as in 8.

Remark 2. Assumptions 8 and 9 hold if

  1. For every saddle point \(y\) of \(F\), \(\lambda_{\min}\big(\nabla^2 F(y)\big) \leq -\lambda_*\).

  2. For every \(x \in \pi^{-1}(x^*)\), and every \(v \in \ker\big(\nabla^2 F(x)\big)\), \(v \in \ker(d\pi|_x)\). Furthermore, for every \(v \in \ker\big(\nabla^2 F(x)\big)^\bot\), \(\nabla^2 F(x)[v] \geq \lambda_* v\).

As we will show in 76, the only non-zero eigenvalues of \(\nabla^2 F\) are equal to those of \(\nabla^2 \tilde{F}\).

While the above assumption on the Hessian of \(\tilde{F}\) allows us to escape saddle points, we also need to ensure that the process \(\tilde{X}_t\) moves sufficiently fast whenever it is sufficiently far from the critical points. For that, we would like the evolution of \(\tilde{X}_t\) (1 ) to be dominated by the gradient term. This requires the norm of the latter to be large enough away from \(\tilde{\mathcal{C}}\).

Assumption 10 (No barren plateaus). There exists a constant \(0 < C_{\tilde{F}} \leq 1\) such that \[\label{eq:boundgradientdistance} |\mathrm{grad}_h\, \tilde{F}(x)|_{h} \geq C_{\tilde{F}} d_{h}(x, \tilde{\mathcal{C}}),\qquad{(1)}\] for every \(x \in M/G\).

Remark 3. This assumption holds whenever \[|\mathrm{grad}_g\, F(x)|_g \geq C_{\tilde{F}} d_g(x, \mathcal{C}),\] for every \(x \in M\), where \(\mathcal{C}\) denotes the set of critical points of \(F\). Indeed, the norms of \(\mathrm{grad}_{g}\, F\) and \(\mathrm{grad}_{h}\, \tilde{F}\) are equal (cf. 74), and distances between points do not increase under Riemannian submersions.

We are now in a position to state the two main results of this work. For \(\beta\) large enough, we have

  • Rapid convergence of the distributions of \(X_t\) and \(\tilde{X}_t\) to their respective Gibbs measures.

  • Concentration of the Gibbs measures around the minima of \(F\) and \(\tilde{F}\).

Theorem 4 (Main result 1 - informal version). Let \((M, g)\) be a compact Riemannian manifold of dimension \(\dim(M) \geq 2\) which is furthermore a symmetric space. Assume there exists a group \(G\) acting on \(M\) and satisfying 1. Let \(\pi : (M, g) \to (M/G, h)\) be the associated Riemannian submersion, where \(M\) and \(M/G\) satisfy 2 3 with constants \(\mathbf{K}\), \(R_{M}\) and \(R_{M/G}\). Let \(F: M \to \mathbb{R}\) be a smooth function satisfying and let \(\tilde{F}: M/G \to \mathbb{R}\) be the unique function such that \(F = \tilde{F} \circ \pi\). Further assume that \(\tilde{F}\) satisfies . Let \(\beta > 0\) be such that \[\beta \in \Omega\Big(\mathrm{poly}\Big(R_{M/G}, \dim(M), A_2, A_3, \mathbf{K}, \frac{1}{i(M/G)}, \frac{1}{i(M)}, \frac{1}{\mathit{conv}(M/G)}, \frac{1}{\lambda_*}, \frac{1}{D}, \frac{1}{C_{\tilde{F}}}\Big)\Big),\] where \(C_{\tilde{F}}\) is the constant from ?? , \(\mathit{conv}(M/G)\) denotes the convexity radius of \((M/G, h)\), \(i(M)\) and \(i(M/G)\) denote the injectivity radii of \((M, g)\) and \((M/G, h)\) respectively, \(D\) is the lower bound on the distance between any two critical points of \(\tilde{F}\), \(\lambda_*\) is the bound on the eigenvalues of \(\nabla^2 \tilde{F}\) at the critical points, and \(A_2, A_3\) are the Lipschitz constants of the gradient and the Hessian of \(\tilde{F}\), respectively.

Then the distributions \(\rho_t\) and \(\tilde{\rho}_t\) of the processes \(X_t\) and \(\tilde{X}_t\) solving the SDEs \[dX_t = -\mathrm{grad}_g\, F(X_t) dt + \sqrt{\frac{2}{\beta}} dW_t, \quad \mathrm{and} \quad d\tilde{X}_t = -\mathrm{grad}_h\, \tilde{F}(\tilde{X}_t) dt + \sqrt{\frac{2}{\beta}} d\tilde{W}_t,\] with initial uniform distribution on \(M\) and \(M/G\), respectively, converge exponentially fast to their respective Gibbs measures \(\nu\), and \(\tilde{\nu}\), i.e. \[\begin{align} \norm{\nu - \rho_t}^2_{\mathrm{TV}} &\leq \beta\, e^{-2\alpha t} \max_{y \in M} F(y), \quad \forall t \geq 0,\\ \norm{\tilde{\nu} - \tilde{\rho}_t}^2_{\mathrm{TV}} &\leq \beta\, e^{-2\tilde{\alpha} t} \max_{y \in M/G} \tilde{F}(y) = \beta\, e^{-2\tilde{\alpha} t} \max_{y \in M} F(y) , \quad \forall t \geq 0, \end{align}\] where the constants \(\alpha\), and \(\tilde{\alpha}\) correspond to the log-Sobolev inequality constants associated with the manifolds \(M\) and \(M/G\), which in particular scale as \[\begin{align} \frac{1}{\alpha} &\in O\Big(\mathrm{poly}\Big(\beta, A_2, R_M, \mathrm{diam}(M), \mathrm{diam}(G), \frac{1}{\lambda_*}\Big)\Big),\\ \frac{1}{\tilde{\alpha}} &\in O\Big(\mathrm{poly}\Big(\beta, A_2, R_{M/G}, \mathrm{diam}(M/G), \frac{1}{\lambda_*}\Big)\Big). \end{align}\] where \(\mathrm{diam}(G)\) denotes the diameter of \(G\) as a fiber of \(\pi\) and \(\mathrm{diam}(M)\), \(\mathrm{diam}(M/G)\) denote the diameter of \((M, g)\) and \((M/G, h)\).

Theorem 5 (Main result 2 - informal version). Let \((M, g)\) be a compact Riemannian manifold of dimension \(\dim(M) \geq 2\). Let \(F: M \to \mathbb{R}\) be twice differentiable with an \(A_2\)-Lipschitz gradient. Let \[\varepsilon_{\max} \in O(i(M)^2 A_2),\] where \(i(M)\) denotes the injectivity radius of \(M\). For every \(\varepsilon \in (0, \varepsilon_{\max}]\) and \(\delta \in (0, 1)\), if \[\begin{align} \beta &\in \Omega\bigg(\frac{\dim(M)^2\log (A_2, \mathrm{Vol}(M),\dim(M), \varepsilon^{-1}, \delta^{-1})}{\varepsilon}\bigg), \end{align}\] the Gibbs distribution \(\nu = \frac{1}{Z}e^{-\beta F}\) satisfies \[\nu\left(F - \min_{y \in M} F(y) \geq \varepsilon\right) \leq \delta.\]

The above results can be of interest when considering a family of minimization problems defined on a sequence of manifolds with growing dimension. It is then natural to wonder how \(\beta\) and the convergence rates \(\alpha, \tilde{\alpha}\) scale with the dimension of the manifold on which \(F\) is defined.

Remark 6. As a consequence of 4, if \[\beta \in \Omega(\mathrm{poly}(\dim(M))),\] i.e. if the constants that lower bound \(\beta\) grow at most polynomially with respect to the dimension of \(M\), assuming furthermore that the diameter of \(M, M/G\) and \(G\), as well as the lower bound on the Ricci curvature of \(M\) grow at most polynomially with respect to the dimension of \(M\), then the log-Sobolev constants satisfy \[\frac{1}{\alpha} \in O(\mathrm{poly}(\dim(M))),\quad \mathrm{and} \quad \frac{1}{\tilde{\alpha}} \in O(\mathrm{poly}(\dim(M))).\]

In particular, whenever \(\max_{y \in M} F\) scales as \(\mathrm{poly}(\dim(M))\), the distance between the distribution of \(X_t\) and \(\nu\) decays exponentially at a rate \(\Omega\Big(\frac{1}{\mathrm{poly}(\dim(M))}\Big)\). Similarly, the distance between the distribution of \(\tilde{X}_t\) and \(\tilde{\nu}\) decays exponentially at a rate \(\Omega\Big(\frac{1}{\mathrm{poly}(\dim(M/G))}\Big)\).

1.2 Structure of the text↩︎

The structure of the next five sections reflects our proof strategy.

  • 2. We will prove that, under slightly milder assumptions than those considered for 4, the base space \(M/G\) satisfies a Poincaré inequality with a constant that only depends on the bound on the eigenvalues of \(\nabla^2 \tilde{F}\) at its critical points, \(\lambda_*\) (cf. 8 9).

  • 3. We will show how to lift a Poincaré inequality from the base space to the total space of a Riemannian submersion \(\pi\), when the Gibbs measure on \(M\) is defined with respect to a function \(F\) that is constant in the fibers of \(\pi\), and the Gibbs measure on \(M/G\) is defined with respect to the projected version of \(F\). This result will allow us to obtain a Poincaré inequality for \(M\).

  • 4. After a short review on classic results which relate the existence of a log-Sobolev inequality to mixing time estimates, we will show how to upgrade a Poincaré inequality to a log-Sobolev inequality under the existence of a curvature-dimension condition. This will ultimately allow us to prove the existence of a logarithmic Sobolev inequality for both \(M/G\) and \(M\), which will in turn translate to upper bounds for the distance between the distribution of the processes \(X_t\) and \(\tilde{X}_t\) (cf. 1 ) and their associated Gibbs measures.

  • 5. We will formalize and prove 4, regarding convergence rates of the processes \(X_t\) and \(\tilde{X}_t\).

  • 6. We will formalize and prove 5, regarding the concentration of the Gibbs measures \(\nu\) and \(\tilde{\nu}\).

7 is an analysis of two scenarios in which most of the assumptions of 4 5 can be easily verified, namely the trace quotient minimization problem, and the mean-field energy minimization problem associated with the two-dimensional ferromagnetic Ising model.

In an effort to make our work as self-contained (and accessible) as possible, background is provided on Riemannian geometry in 8, Sobolev spaces in 9, Bakry-Émery theory in 10, stochastic differential equations in 11, and partial differential equations on manifolds in 12. 13 is a discussion of a technical aspect of 2.

2 Poincaré inequality under unique minimum↩︎

This section is devoted to the main element in the proof of the rapid mixing result stated in 4: obtaining a suitable Poincaré inequality in the base space of a Riemannian submersion.

Definition 1. Let \((M, g)\) be a Riemannian manifold, \(F \in C^\infty(M)\), and \(\beta > 0\). The generator of the Langevin dynamics \(\operatorname{L}\) is the operator acting on5 \(C^2(M)\) as \[\operatorname{L}f := \langle -\mathrm{grad}_{g}\, F, \mathrm{grad}_{g}\,f\rangle_g + \frac{1}{\beta} \Delta_g f.\]

We will follow the notation and definitions of [36]. For an in-depth study of stochastic processes on Riemannian manifolds, see [23].

Definition 2. Let \((\Omega, \mathcal{F}, \mathbb{P})\) be a probability space and \(M\) be a smooth manifold. A continuous-time stochastic process in \(M\) is a family of random variables \[\{Y_t: t \in \mathbb{R}^+\},\] where \(Y_t : \Omega \to M\) for every \(t \in \mathbb{R}^+\).

Definition 3. Given some filtered probability space \((\Omega, \mathcal{F})\), a compact Riemannian manifold \((M,g)\), the Langevin diffusion process \(X_t\) is the unique \(\mathcal{F}\)-adapted semimartingale on \(M\) such that \[\label{martingaleproblem} M^f_t := f(X_t) - f(X_0) - \int_0^t \operatorname{L}f(X_s)ds,\qquad{(2)}\] is an \(\mathcal{F}\)-adapted local martingale for all \(f \in C^2(M)\).

The existence and uniqueness of such a process is proven in [23]. We often say that \(X_t\) is generated by \(\operatorname{L}\) or that \(X_t\) solves the martingale problem associated with \(\operatorname{L}\). The formal SDE associated with this process is \[\label{LangevinDiffusionEq} dX_t = -\mathrm{grad}_{g}\,F(X_t)\,dt + \sqrt{\frac{2}{\beta}}\,dW_t,\tag{2}\] where \(W_t\) is a Brownian motion on \(M\) and \(X_0\) is initialized with a distribution \(\rho_0\) supported on \(M\).

Poincaré inequalities involve the notions of carré du champ operator and Markov triple.

Definition 4 (Carré du champ operator). Let \((M, g)\) be a compact manifold, and let \(\operatorname{L}\) be as given in 1, we define the carré du champ operator \(\Gamma\) on \(C^2(M) \times C^2(M)\) as \[\Gamma(f_1,f_2) := \frac{1}{2}(\operatorname{L}(f_1f_2) - f_1\operatorname{L}f_2 - f_2\operatorname{L}f_1).\]

Notation 7. We will often use \(\Gamma(f)\) to denote \(\Gamma(f,f)\).

Remark 8. The carré du champ operator associated with \(\operatorname{L}\) as given in 1 can be written as \[\Gamma(f_1, f_2) = \frac{1}{\beta}\langle \mathrm{grad}_{g}\,f_1, \mathrm{grad}_{g}\,f_2\rangle_g,\] for any \(f_1, f_2 \in C^2(M)\), and so \[\Gamma(f) = \frac{1}{\beta}|\mathrm{grad}_{g}\,f|^2_g,\] for any \(f \in C^2(M)\).

Proof. The result automatically follows, using the well-known identities \[\Delta_g(f_1f_2) = f_1\Delta_g f_2 + f_2\Delta_g f_1 + 2 \langle \mathrm{grad}_{g}\,f_1, \mathrm{grad}_{g}\,f_2\rangle_g,\] and \[\mathrm{grad}(f_1f_2) = f_1\,\mathrm{grad}_{g}\,f_2 + f_2\, \mathrm{grad}_{g}\,f_1.\] ◻

Definition 5 (Markov triple). Let \((M, g)\) be a compact manifold, \(F \in C^\infty(M)\), \({\beta >0}\), and \(\operatorname{L}\) be the associated Langevin diffusion generator. The Markov triple \((M, \nu, \Gamma)\) associated with \(\operatorname{L}\) consists of the Gibbs measure \[d\nu := \frac{1}{Z}e^{-\beta F} \mathrm{dVol}_g,\] where \(Z\) is the normalizing constant, \(Z := \int_M e^{-\beta F} \mathrm{dVol}_g\), and \(\Gamma\) is the carré du champ operator associated with \(\operatorname{L}\).

Definition 6 (Poincaré inequality). Let \((M, g)\) be a compact Riemannian manifold, \(F \in C^\infty(M)\), and \(\beta >0\). The Markov triple \((M, \nu, \Gamma)\) satisfies a Poincaré inequality with constant \(\kappa > 0\), denoted \(\mathrm{PI}(\kappa)\), if for every \(f \in C^2(M)\), \[\int_M f^2 \,d\nu - \Big(\int_M f \,d\nu\Big)^2 \leq \frac{1}{\kappa} \int_M \Gamma(f) \, d\nu.\]

The Poincaré inequality constant \(\kappa\) controls the rate of convergence of the Langevin diffusion process [12]. Bounding this constant is also the first step toward obtaining a log-Sobolev inequality.

Let us consider two Riemannian manifolds \((M, g)\) and \((B, h)\), and let us assume that \((M, g)\) is a compact symmetric space. Further assume that there exists a—surjective—Riemmanian submersion with totally geodesic fibers, \(\pi: (M, g) \to (B, h)\). Note that, as we are assuming that \(M\) is compact, \(B\) must be compact as well. Now let \(F: M \to \mathbb{R}\) be a smooth function that is constant in the fibers of \(\pi\). The existence of such a function \(F\) is equivalent to the existence of a function \(\tilde{F} : B \to \mathbb{R}\) satisfying \(F = \tilde{F} \circ \pi\). It is known that \(\tilde{F}\) is smooth if and only if \(F\) is smooth [2]. Therefore, for such fiber-invariant function \(F\), the Markov triple \((M, \nu, \Gamma)\) naturally induces another \((B, \tilde{\nu}, \tilde{\Gamma})\) for \(\tilde{F}\) and vice-versa, where of course \[\label{eq:defMarkovB} \tilde{\operatorname{L}} := -\mathrm{grad}_h \tilde{F} + \frac{1}{\beta} \Delta_h,\quad \tilde{\Gamma}(f) = \frac{1}{\beta}|\mathrm{grad}_{h}\,f|^2_h,\quad \mathrm{and}\quad d\tilde{\nu} = \frac{1}{\tilde{Z}} e^{-\beta \tilde{F}}\mathrm{dVol}_h,\tag{3}\] where \(\tilde{Z}\) is the normalizing constant for \(\tilde{\nu}\) on \((B, h)\).

As alluded in the previous section, when attempting to prove rapid mixing for a Langevin process, the multiplicity of minima turns out to be a serious technical obstruction. If a submersion allows to remove this multiplicity, it turns out that deriving a Poincaré inequality for the Markov triple \((B, \tilde{\nu}, \tilde{\Gamma})\) is much simpler than for the original Markov triple \((M, \nu, \Gamma)\), hence our interest in the former. The problem of lifting such a Poincaré inequality to the latter Markov triple, and of upgrading it to a log-Sobolev inequality will be dealt with in [SectionLift] [SectionPItoLSI], respectively.

To prove a Poincaré inequality for the Markov triple \((B, \tilde{\nu}, \tilde{\Gamma})\), we will follow the steps of [17]. We will first define and prove the existence of two suitable Lyapunov functions. Next, we will show how these Lyapunov functions can be used to extend a local Poincaré inequality, which is easy to obtain, to the whole manifold \(B\).

2.1 A first Poincaré inequality↩︎

Let us restate for completeness.

Assumption 11. There exists a constant \(\mathbf{K} \geq 1\) such that for every \(p\in M\), in normal coordinates centered at \(p\), \(|R_{ijk}^l(p)| \leq \mathbf{K}\) for every \(i, j, k, l \in \{1, \dotsc, \dim(M)\}\), where \(R\) denotes the Riemann curvature tensor of \((M, g)\). There also exist some constants \(R_M, R_B \geq 0\) that lower bound the Ricci curvatures of \((M, g)\) and \((B, h)\), respectively, \[\mathrm{Ric}_h \geq -R_{B},\quad\mathrm{\textit{and}}\quad \mathrm{Ric}_g \geq -R_M.\]

Assumption 12. \(F\) is constant in the fibers of \(\pi\)—i.e. \(F(x) = F(y)\) for every \(x, y \in M\) such that \(\pi(x) = \pi(y)\).

Remark 9. Since \(M\) is compact and \(F\) is smooth, we know that there exist constants \(A_1, A_2, A_3 \geq 1\) such that \(F\) is \(A_1\)-Lipschitz, \(\mathrm{grad}_{g}\,F\) is \(A_2\)-Lipschitz, and \(\nabla^2 F\) is \(A_3\)-Lipschitz. The same Lipschitz constants are valid for \(\tilde{F}\), as it will be shown in 81.

Assumption 13. \(\min_{x \in M} F(x) = 0\).

Assumption 14. \(\tilde{F}\) has a unique minimum.

Assumption 15. The critical points of \(\tilde{F}\) are isolated points separated by some distance \(D\), i.e. \[d_h(x_1, x_2) \geq D, \quad \forall x_1, x_2 \in \tilde{\mathcal{C}},\;x_1 \neq x_2.\]

Assumption 16. For every saddle point \(y \in \tilde{\mathcal{S}}\), there exists an escape direction. That is, there exists some constant \(\lambda_* \in (0, 1]\) such that for every saddle point \(y \in \tilde{\mathcal{S}}\) it holds that \[\lambda_{\min}\big(\nabla^2 \tilde{F}(y)\big) \leq -\lambda_* < 0.\]

Assumption 17. The global minimum is an attractor for the process \(\tilde{X}_t\). In other words, the Hessian of \(\tilde{F}\) at the global minimum is non-degenerate. More precisely, at the unique minimum \(x^*\) of \(\tilde{F}\), we have \[\lambda_{\min}\big(\nabla^2 \tilde{F}(x^*)\big) \geq \lambda_*.\] where \(\lambda_*\) is the same constant as in 8.

Assumption 18. There exists a constant \(0 < C_{\tilde{F}} \leq 1\) such that \[\label{eq94641bis} |\mathrm{grad}_{h}\,\tilde{F}(x)|_{h} \geq C_{\tilde{F}} d_{h}(x, \tilde{\mathcal{C}}),\qquad{(3)}\] for every \(x \in B\).

Theorem 10 (Poincaré Inequality on \(B\)). Let \((M, g)\) and \((B, h)\) be two Riemannian manifolds satisfying 11 with constants \(\mathbf{K}\), \(R_M\) and \(R_B\). Assume that \((M, g)\) is furthermore a compact symmetric space. Let \(\pi: (M, g) \to (B, h)\) be a surjective Riemannian submersion with totally geodesic fibers. Let \(F: M \to \mathbb{R}\) be a function satisfying and let \(\tilde{F}: B \to \mathbb{R}\) be the unique function such that \(F = \tilde{F} \circ \pi\). Assume that \(\tilde{F}\) satisfies . Let \(a, \beta > 0\) be such that \[\begin{align} a^2 &\geq \max\Big\{\frac{24 A_2 \dim(B)}{C^2_{\tilde{F}}}, \frac{544}{\lambda_*}\Big\}, \\\beta &\geq \max\Bigg\{\frac{72^2 \dim(M)^5 A_2 A^2_3 \mathbf{K}^2 a^6}{\lambda_*^2}, \frac{9a^2}{D^2}, \frac{a^2}{i(B)^2}, \frac{a^2}{i(M)^2}, \frac{4R_{B}}{\lambda_*}, \frac{a^2}{\mathit{conv}(B)^2}\Bigg\}, \end{align}\] where \(C_{\tilde{F}}\) is the constant from ?? , \(\mathit{conv}(B)\) denotes the convexity radius of \(B\), \(i(M)\) and \(i(B)\) denote the injectivity radii of \((M, g)\) and \((B, h)\) respectively, \(D\) is the lower bound on the distance between any two critical points of \(\tilde{F}\), \(\lambda_*\) is the bound on the eigenvalues of \(\nabla^2 \tilde{F}\) at the critical points, and \(A_2, A_3\) are the Lipschitz constants of the gradient and the Hessian of \(\tilde{F}\), respectively.

Then the Markov triple \((B, \tilde{\nu}, \tilde{\Gamma})\) satisfies a \(\mathrm{PI}(\kappa)\), where \[\kappa = \frac{\lambda_*}{184}.\]

Informally, we will see that, whenever the amount of noise \(\frac{1}{\beta}\) is sufficiently small, we obtain a Poincaré inequality with a constant that only depends on the value of \(\lambda_*\), which in turn controls how fast \(\tilde{X}_t\) escapes from the saddle points of \(\tilde{F}\). The constant \(a\) sets a lower bound on the values of \(\beta\) for which we can claim that the Poincaré inequality holds.

The rest of the section will be devoted to the proof of this result, which will be given in 2.3.

2.2 Lyapunov functions↩︎

Definition 7 (Lyapunov function). Let \((M, g)\) be a Riemannian manifold, and let \(\operatorname{L}\) denote the Langevin generator from 1 for some \(F \in C^\infty(M)\) and \(\beta > 0\). A function \(W : M \rightarrow \mathbb{R}\), \(W \in C^2(M)\) is said to be a Lyapunov function on a measurable subset \(U \subset M\) with parameters \(\theta > 0\) and \(b \geq 0\) if \(W \geq 1\) and \[\frac{\operatorname{L}W(x)}{W(x)} \leq -\theta + b\mathbb{1}_U(x), \quad \forall x \in M.\]

Definition 8 (Quasi-Lyapunov function). Let \((M, g)\) be a Riemannian manifold, and \(\operatorname{L}\) be the Langevin generator. A function \(W: M \rightarrow \mathbb{R}\) is said to be a quasi-Lyapunov function on an open subset \(U \subset M\) with parameter \(\theta > 0\) if \(W \in C^2(\overline{U}) \cap C^2(U^c)\), \(W \geq 1\) and \[\operatorname{L}W(x) \leq -\theta W(x), \quad \forall x \in U.\]

As we mentioned earlier, we are interested in obtaining Lyapunov functions6 defined on (submanifolds of) \((B, h)\) and associated with the operator \[\tilde{\operatorname{L}} = - \mathrm{grad}_{h}\,\tilde{F} + \frac{1}{\beta} \Delta_h,\] where \(\tilde{F}\) is the function from 10 satisfying .

2.2.1 First Lyapunov function↩︎

To construct our first Lyapunov function, we will use the following result.

Lemma 1 ([18]). Let \((B, h)\) be a compact Riemannian manifold. Let \(\tilde{F}: B \to \mathbb{R}\) be smooth, and suppose that it satisfies . Let \(\tilde{\mathcal{C}}\) be the set of critical points of \(\tilde{F}\). Then, for every \(\beta > 0\) and every \(a > 0\) satisfying \[a^2 \geq \frac{24A_2 \dim(B)}{C^2_{\tilde{F}}},\] where \(A_2\) is the Lipschitz constant of the gradient of \(\tilde{F}\), it holds that \[\frac{\Delta_h \tilde{F}(x)}{2} - \frac{\beta}{4}|\mathrm{grad}_{h}\,\tilde{F}(x)|^2_h \leq -A_2 \dim(B),\quad \forall x \in B : d_h(x, \tilde{\mathcal{C}})^2 \geq \frac{a^2}{4\beta}.\]

Proof. Using 18 it holds that \[|\mathrm{grad}_{h}\,\tilde{F}(x)|_h \geq C_{\tilde{F}} d_h(x, \tilde{\mathcal{C}}).\] Thus, for every \(x \in B\) such that \(d_h(x, \tilde{\mathcal{C}}) \geq \frac{a^2}{4\beta}\) we obtain that \[\begin{align} \frac{\Delta_h \tilde{F}(x)}{2} - \frac{\beta}{4}|\mathrm{grad}_{h}\,\tilde{F}(x)|^2_h &\leq \frac{\Delta_h \tilde{F}(x)}{2} - \frac{\beta}{4}C_{\tilde{F}}^2 d_h(x, \tilde{\mathcal{C}})^2 \\&\leq \frac{\Delta_h \tilde{F}(x)}{2} - \frac{a^2}{16}C_{\tilde{F}}^2 \\&\leq \frac{A_2\dim(B)}{2} - \frac{1}{16}C_{\tilde{F}}^2 \frac{24A_2 \dim(B)}{C^2_{\tilde{F}}} \\&\leq -A_2 \dim(B), \end{align}\] where the second-last inequality follows by the assumption \(a^2 \geq \frac{24A_2 \dim(B)}{C^2_{\tilde{F}}}\). ◻

In fact, 18 is only used in the proof of 1. Therefore, this assumption can be relaxed; it need only hold for every \(x \in B\) satisfying \(d_h(x, \tilde{\mathcal{C}})^2 \geq \frac{a^2}{4\beta}\), where \(a^2\) is defined as above and \(\beta\) is some positive constant.

Notation 11. For all \(r > 0\) and \(X \subset B\), we will denote the set of points within distance \(r\) to \(X\) as \(U(r,X)\), i.e. \[U(r,X) := \{x \in B : d_h(x, X) < r\}.\]

Let \(\tilde{\mathcal{S}}\) denote the saddle points of \(\tilde{F}\). We are going to prove that \[\label{eq:defW1} W_1: \overline{U}\Big(\frac{a}{2\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)^c \to \mathbb{R},\quad x \mapsto \exp\left(\frac{\beta}{2}\tilde{F}(x)\right)\tag{4}\] is a Lyapunov function on the geodesic ball \(\mathcal{B}(\frac{a}{\sqrt{\beta}}, x^*)\), where \(x^*\) is the unique minimum of \(\tilde{F}\), \(\overline{U}\) denotes the closure of \(U\) and \(\cdot^c\) denotes the complement.

Proposition 12. Under the assumptions of 1, the function \(W_1\) from 4 is a Lyapunov function on \(\mathcal{B}(\frac{a}{\sqrt{\beta}}, x^*)\) with respect to the Langevin generator \[\tilde{\operatorname{L}} = - \mathrm{grad}_{h}\,\tilde{F} + \frac{1}{\beta} \Delta_h,\] whenever \[\frac{a}{\sqrt{\beta}} \leq \frac{D}{3},\] where \(D\) denotes the least distance between any two critical points of \(\tilde{F}\).

Proof. From the choice of \(\frac{a}{\sqrt{\beta}}\), we can decompose \[\overline{U}\Big(\frac{a}{2\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)^c = \mathcal{B}\Big(\frac{a}{\sqrt{\beta}}, x^*\Big) \cup \overline{U}\Big(\frac{a}{2\sqrt{\beta}}, \tilde{\mathcal{C}}\Big)^c.\] Now, let us write \[\begin{align} \tilde{\operatorname{L}} W_1(x) &= \Big\langle -\mathrm{grad}_{h}\,\tilde{F}(x), \mathrm{grad}_{h}\,\exp\Big(\frac{\beta}{2}\tilde{F}(x)\Big)\Big\rangle_h + \frac{1}{\beta} \Delta_h \exp\left(\frac{\beta}{2}\tilde{F}(x)\right) \\&= -\frac{\beta}{2} W_1(x) |\mathrm{grad}_{h}\,\tilde{F}(x)|^2_h + \frac{1}{\beta} \left(\frac{\beta}{2}W_1(x) \Delta_h \tilde{F}(x) + \frac{\beta^2}{4} W_1(x) |\mathrm{grad}_{h}\,\tilde{F}(x)|^2_h\right) \\&= W_1(x) \Big(\frac{1}{2} \Delta_h \tilde{F}(x) - \frac{\beta}{4} |\mathrm{grad}_{h}\,\tilde{F}(x)|^2_h\Big). \end{align}\]

We can now apply 1 and the trivial bound involving the Lipschitz constant \(A_2\), \[\frac{1}{2} \Delta_h \tilde{F}(x) - \frac{\beta}{4} |\mathrm{grad}_{h}\,\tilde{F}(x)|^2_h \leq \frac{1}{2}\Delta_h \tilde{F}(x) \leq \frac{\dim(B) A_2}{2},\] which holds for every \(x \in B\), to conclude that \[\label{lyapunovW1} \frac{\tilde{\operatorname{L}} W_1(x)}{W_1(x)} \leq - A_2 \dim(B) + \frac{3}{2} A_2 \dim(B) \mathbb{1}_{\mathcal{B}\big(\frac{a}{\sqrt{\beta}}, x^*\big)}(x),\tag{5}\] for every \(x \in \overline{U}\Big(\frac{a}{2\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)^c\). 5 proves that \(W_1\) is a Lyapunov function on \(\mathcal{B}\Big(\frac{a}{\sqrt{\beta}}, x^*\Big)\), as desired. ◻

2.2.2 Second Lyapunov function: local escape time bounds↩︎

We now show how to obtain a quasi-Lyapunov function on the set \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\). To this end, we will first study the PDE \[\tilde{\operatorname{L}}W_2 = -\theta W_2, \quad \forall x \in U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big),\] where \(\theta >0\) is some fixed constant. The next proposition provides an explicit expression for the solution of this PDE, under the assumption that it exists and under an additional finiteness assumption.

Proposition 13. Let \(\mathscr{L}\) be the generator of a Markov process \(Y_t\) on a manifold \(M\) and assume that the associated martingale problem has a unique solution, given by \(Y_t\). Let \(V \subset M\) be an open set. Assume that \(f \in \mathcal{D}(\mathscr{L})\) is a solution of the following Dirichlet problem: \[\begin{cases} (-\mathscr{L}f)(x) -\theta f(x) = 0,&\quad \forall x \in V, \\ f(x) = 1,&\quad \forall x \in V^c, \end{cases}\] where \(\theta\) is some positive constant and \(\mathcal{D}(\mathscr{L})\) denotes the domain of \(\mathscr{L}\). Let \[\tau_{V^c} := \inf\{t > 0 : Y_t \notin V\}\] denote the first escaping time of \(Y_t\) from \(V\). If \[\label{eq:finitenessassumption} \mathbb{E}\left[\exp(\theta\tau_{V^c}) | Y_0 = x\right] < \infty, \quad \forall x \in V,\qquad{(4)}\] then \(f\) can be written as \[f(x) = \mathbb{E}[\exp(\theta \tau_{V^c}) | Y_0 = x].\]

This proposition is a variation of [38] where both our claim and our assumptions are weaker. The resulting proof is analogous and outlined here.

Proof. Let \(f\) be the solution of the above Dirichlet problem. Note that \[(-\mathscr{L}f)(x) -\theta f(x) = 0,\] for every \(x \in V\), and \(Y_t\) solves the martingale problem associated with \(\mathscr{L}\), which implies that \[Z_t := e^{\theta t}f(Y_t)\] is a martingale [38]. Moreover, for every \(x \in V\) it holds that \[\mathbb{E}[|\exp(\theta \tau_{V^c})f(Y_{\tau_{V^c}})| \,|\, Y_0=x ] = \mathbb{E}[|\exp(\theta \tau_{V^c})|\,|\, Y_0=x] < \infty,\] where the equality follows from \(f(x)\) being 1 in \(V^c\), and finiteness follows from ?? . Therefore, it holds that \[\mathbb{E}[\exp(\theta \tau_{V^c}) | Y_0=x] = \mathbb{E}[Z_{\tau_{V^c}} | Y_0 = x]= Z_0 = f(x),\] where the second equality follows from [38] with \(T = \tau_{V^c}\) and \(S = 0\). ◻

Let us apply 13 to our setting. Under the same finiteness assumption, we will be able to ensure that the PDE associated with our desired quasi-Lyapunov condition has a unique solution. Thus, using the previous lemma, we will obtain our candidate for quasi-Lyapunov function.

Corollary 1. Let \((B, h)\) be a compact Riemannian manifold. Suppose that \(\tilde{F} : B \to \mathbb{R}\) is smooth and satisfies . Let \(a, \beta > 0\), \[\tilde{\operatorname{L}} = -\mathrm{grad}_{h}\,\tilde{F} + \frac{1}{\beta} \Delta_h,\] and let \(\tilde{X}_t\) be the Langevin diffusion process generated by \(\tilde{L}\). Consider \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\), and assume that \(\partial U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\) is \(C^\infty\) (cf. 64). Furthermore, assume that there exists some constant \(\theta > 0\) such that \(-\theta\) is not in the spectrum of \(\tilde{\operatorname{L}}\)—denoted as \(\sigma(\tilde{\operatorname{L}})\)—and such that7 \[\label{eqfiniteness} \mathbb{E}[\exp(\theta \tau^*) | \tilde{X}_0 = x] < \infty,\quad \forall x \in U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big),\qquad{(5)}\] where \(\tau^*\) is the first escaping time of \(\tilde{X}_t\) from \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\). Then the function \(W_2\) defined as \[W_2 : B \to \mathbb{R},\quad x \mapsto \mathbb{E}[\exp(\theta \tau^*) | \tilde{X}_0 = x]\] is a quasi-Lyapunov function with respect to \(\tilde{\operatorname{L}}\) on \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\).

Remark 14. The regularity assumption of the boundary of \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\) is made to guarantee that \(W_2\) is smooth up to the boundary of \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\). It can be satisfied imposing that \(\frac{a}{\sqrt{\beta}}\) is upper bounded by the injectivity radius of \(B\), \(i(B)\). Indeed, \(\partial U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big) = d^{-1}_h(\frac{a}{\sqrt{\beta}})\), and the distance function \(d_h(\cdot, \tilde{\mathcal{S}})\) is \(C^\infty\) in \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big) \setminus \tilde{\mathcal{S}}\) whenever \(\frac{a}{\sqrt{\beta}} \leq \min\{i(B), D/3\}\) [39], where \(D\) is the least distance between two critical points (cf. 15). Thus, \(\partial U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\) will be \(C^\infty\) whenever \(\frac{a}{\sqrt{\beta}}\) is a regular value (cf. 65) of the distance function to the saddle points. Again, this can be assumed without loss of generality, as the set of regular values for a smooth function is dense (cf. 116).

In order to guarantee that the solution to the PDE exists, we need to assume that \(-\theta\) is not in the spectrum of \(\tilde{\operatorname{L}}\). Again, this can be done without loss of generality, by simply changing \(\theta\) in ?? to \(\theta - \varepsilon\) for some \(\varepsilon > 0\) sufficiently small. In fact, if the finiteness assumption holds for some \(\theta > 0\), it holds for every \(\theta' < \theta\). After this change, the resulting \(\theta\) will not be in the spectrum of \(\tilde{\operatorname{L}}\), since its spectrum is a non-increasing sequence tending to minus infinity by 114.

Proof of 1. Let us first define the PDE \[\begin{cases} \tilde{\operatorname{L}}f = -\theta f, \quad &x \in U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big),\\ f = 1,\quad &x \in U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)^c. \end{cases}\] Under the assumption that \(-\theta\) is not in the spectrum of \(\tilde{\operatorname{L}}\), the PDE has a unique solution. Indeed the boundary problem has a solution (cf. 12), which can be extended to \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)^c\) by simply setting it to \(1\).

Thus, we can apply 13 to conclude that the solution \(f\) is of the form \[f(x) = \mathbb{E}[\exp(\theta \tau^*) | \tilde{X}_0 = x],\] which matches the definition of \(W_2\). This implies that \(W_2 \geq 1\), and, using the assumption on the regularity of the boundary of \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\), we can conclude that \[W_2 \in C^{\infty}\Big(\overline{U}\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\Big) \cap C^\infty\Big(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)^c\Big)\] (cf. 12.3). ◻

In order to guarantee that \(W_2\) is our desired quasi-Lyapunov function, it only remains to ensure that there exists some \(\theta > 0\) such that \[\label{eq:finitenesscond2} \mathbb{E}[\exp(\theta \tau^*) | \tilde{X}_0=x] < \infty,\quad \forall x \in U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big),\tag{6}\] which will in turn allow us to apply 1. A sufficient condition for 6 to hold is that the distribution for \(\tau^*\) has exponentially decaying tails.

Lemma 2 ([40]). Let \(X\) be a random variable, and assume that there exist two constants \(c_1, c_2 > 0\) such that \[\label{eq94616} \mathbb{P}[|X| \geq t] \leq c_1 e^{-c_2 t},\quad \forall t > 0.\qquad{(6)}\] Then taking \(c_0 = c_2/2\), it holds that \[\mathbb{E}[e^{cX}] < \infty,\quad \forall|c|\leq c_0.\]

Proof. Let \(a > 0\) and \(T > 0\). Then \[\mathbb{E}[e^{a|X|}\mathbb{1}_{\{e^{a|X|} \leq e^{aT}\}}] \leq \int_0^{e^{aT}} \mathbb{P}[e^{a|X|} \geq t] dt \leq 1 + \int_1^{e^{aT}} \mathbb{P}\left[|X| \geq \frac{\log t}{a}\right]dt.\] Since we are assuming that ?? holds, we can conclude that \[\mathbb{E}[e^{a|X|}\mathbb{1}_{\{e^{a|X|} \leq e^{aT}\}}] \leq 1 + c_1 \int_1^{e^{aT}} e^{-\frac{c_2 \log t}{a}} dt = 1 + c_1 \int_1^{e^{aT}} t^{-c_2/a} dt.\] Thus, for any \(a \in [0, \frac{c_2}{2}]\), it holds that \[\mathbb{E}[e^{a|X|}\mathbb{1}_{\{e^{a|X|} \leq e^{aT}\}}] \leq 1 + c_1(1 - e^{-aT}) \leq 1 + c_1.\] Taking the limit as \(T \rightarrow \infty\), we conclude that \(\mathbb{E}[e^{a|X|}]\) is finite for all \(a \in [0, \frac{c_2}{2}]\). Since both \(e^{aX}\) and \(e^{-aX}\) are upper bounded by \(e^{|a||X|}\), it follows that \(\mathbb{E}[e^{aX}]\) is finite for all \(|a|\leq \frac{c_2}{2}\). ◻

This lemma will be used to prove the claim of finiteness of 6 . For that, for every \(x \in U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\) we will seek two constants \(c_1, c_2 > 0\) such that \[\label{eq:desiredfiniteness} \mathbb{P}[\tau^* \geq t | \tilde{X}_0 = x] \leq c_1 e^{-c_2 t}, \quad \forall t \geq 0,\tag{7}\] and such that \(c_2\) does not depend on \(x\).

To do this, we will study how fast \(\tilde{X}_t\), escapes from \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\). The result will be proven in 15. It requires a few auxiliary results and definitions. We start with a version of the mean value theorem in Riemannian manifolds.

Lemma 3 ([2]). Let \((M, g)\) be a Riemannian manifold and let \(f \in C^2(M)\). Let \(x \in M\), \(v \in T_xM\), and let \(\gamma(t) = \exp_x(tv)\) be defined on \([0, 1]\), where \(\exp_x: T_xM \to M\) denotes the exponential map at \(x\). Also assume that there exists some \(\kappa \geq 0\) such that, for all \(t \in [0, 1]\), \[\norm{\mathrm{P}^{-1}_{tv} \circ \nabla^2 f(\gamma(t)) \circ \mathrm{P}_{tv} - \nabla^2 f(x)} \leq \kappa|tv|_g.\] Then \[|\mathrm{P}^{-1}_v \mathrm{grad}_{g}\,f(\exp_x(v)) - \mathrm{grad}_{g}\,f(x) - \nabla^2f(x)[v]|_g \leq \frac{\kappa}{2}|v|^2_g,\] where \(\mathrm{P}_{tv}\) denotes the parallel transport from \(x\) to \(\gamma(t)\) along \(\gamma\).

To simplify the notation, let us define the following auxiliary function.

Definition 9. Let \(\tilde{F}\) be the function on \(B\) associated with the Markov triple \((B, \tilde{\nu}, \tilde{\Gamma})\). Let \(y \in B\) be some fixed critical point of \(\tilde{F}\). Using normal coordinates centered at \(y\), let us define \(H_y(x)\) for every \(x\) outside the cut locus of \(y\) as \[H_y(x) := \delta^{ij} \partial_{jk} \tilde{F}(y)x^k \partial_i = \nabla^2 \tilde{F}(y)[\log_y x] \in T_y B,\] where \(\log_y\) is the inverse of the exponential map \(\exp_y\) and \[\nabla^2 \tilde{F}(y)[\log_y x] = \nabla_{\log_y x} \, \mathrm{grad}_{h}\,\tilde{F}|_y.\]

Lemma 4. Let \(y \in \tilde{\mathcal{S}}\) be some fixed saddle point of \(\tilde{F}\). Then for every \(x\) outside the cut locus of \(y\) it holds that \[|\mathrm{P}^{-1}_{\log_y x}\, \mathrm{grad}_{h}\,\tilde{F}(x) - H_y(x)|_h \leq \frac{A_3}{2} d_h(x, y)^2,\] where \(A_3\) is the Lipschitz constant of the Hessian of \(\tilde{F}\).

Proof. The proof follows automatically by applying 3, using the fact that \(\mathrm{grad}_{h}\,\tilde{F}(y) = 0\) and that \(\tilde{F}\) has a \(K_3\)-Lipschitz Hessian. ◻

When analysing how \(\tilde{X}_t\) escapes the saddle points of \(\tilde{F}\), instead of considering the distance between \(\tilde{X}_t\) and \(\tilde{\mathcal{S}}\), we will study an auxiliary function.

Definition 10. Let \(y \in B\) be fixed, and let \(v \in T_y B\) be a fixed unit vector. For every \(x\) outside the cut locus of \(y\), we define, in normal coordinates \[\tilde{r}_{y, v}(x) := \langle v, \log_y x\rangle.\]

The main reason why we should look at \(\tilde{r}_{y, v}\) instead of the distance function is because the gradient of \(\tilde{r}_{y, v}\) is always in the direction of \(v\), which is key in the proof of 15 (see also 33). This is clearly not the case for the gradient of the distance function \(x \mapsto d_h(x, \tilde{\mathcal{S}})\).

We observe that, by Cauchy-Schwarz inequality \[\label{eq:ineqtilderd} |\tilde{r}_{y, v}(x)| \leq |\log_y x|_h = d_h(x, y),\tag{8}\] for all \(y \in B\) and any \(x \in B\) outside the cut locus of \(y\). This implies that for every \(y \in \tilde{\mathcal{S}}\), any lower bound on \(|\tilde{r}_{y, v}(\tilde{X}_t)|\) becomes a lower bound on \(d_h(\tilde{X}_t, \tilde{\mathcal{S}})\).

We are now in the position to provide the proof that \(\tau^*\) has exponentially decaying tails (cf. 7 ). This proof follows the same lines as that of [18].

Proposition 15 (Local Escape Time Bound). Let \((M, g)\) be a compact manifold which is furthermore a symmetric space. Let \(\pi: (M, g) \to (B, h)\) be a surjective Riemannian submersion with totally geodesic fibers. Assume that there exists a constant \(\mathbf{K} \geq 1\) such that for every \(p\in M\), in normal coordinates centered at \(p\), \(|R_{ijk}^l(p)| \leq \mathbf{K}\) for every \(i, j, k, l \in \{1, \dotsc, \dim(M)\}\), where \(R\) denotes the Riemann curvature tensor of \((M, g)\). Let \(F\) be a smooth function on \(M\) verifying 12 13, and let \(\tilde{F}\) be the unique function on \(B\) satisfying \(F = \tilde{F} \circ \pi\). Suppose that \(\tilde{F}\) satisfies .

Let \(\tilde{X}_t\) be the Langevin diffusion process on \(B\) generated by \[\tilde{\operatorname{L}} = -\mathrm{grad}_{h}\,\tilde{F} + \frac{1}{\beta} \Delta_h.\] Let \(a \geq 1\), \[\begin{align} \beta &\geq \max\Bigg\{72^2 \dim(M)^5 A_2 A^2_3 \mathbf{K}^2 a^6, \frac{9a^2}{D^2}, \frac{a^2}{i(B)^2}, \frac{a^2}{i(M)^2}\Bigg\}, \end{align}\] where \(D\) is such that \(d_h(x_i, x_j) \geq D\), for every \(x_i, x_j \in \tilde{\mathcal{C}}\), \(A_2, A_3 \geq 1\) are the Lipschitz constants of the gradient and the Hessian of \(\tilde{F}\), and \(i(B)\), \(i(M)\) denote the injectivity radii of \((B, h)\) and \((M, g)\), respectively. For every \(x \in U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\), there exists a constant \(C > 0\) such that \[\mathbb{P}[\tau^* \geq t | \tilde{X}_0 = x] \leq C e^{-\lambda_* t / 8},\quad \forall\, t \geq 0,\] where \(C > 0\) is a constant independent of \(t\), \(\lambda_*\) is the bound on the eigenvalues of \(\nabla^2 \tilde{F}\) at the critical points, and \(\tau^*\) is the first escaping time of \(\tilde{X}_t\) from \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\), i.e. \[\tau^* = \inf\Big\{t > 0 : \tilde{X}_t \notin U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\Big\}.\]

15 allows us to guarantee that, whenever we consider a sufficiently small neighborhood around the saddle points of \(\tilde{F}\) and the amount of noise \(\frac{1}{\beta}\) is sufficiently small, \(\tilde{X}_t\) is able to escape from the saddle points sufficiently fast. More precisely, the probability distribution of the first escaping time of \(\tilde{X}_t\) from \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\) has exponentially decaying tails.

Let us now briefly discuss the structure of the proof. As we mentioned earlier, in order to analyze the first escaping time of \(\tilde{X}_t\) from \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\), instead of considering the distance between \(\tilde{X}_t\) and the saddle points, we will use \(\tilde{r}_{y, v}(\tilde{X}_t)\), where \(y\) is some fixed saddle point, and we will prove that this quantity becomes sufficiently large with high probability, which in turn allows us to conclude that \(\tilde{X}_t\) escapes the saddle points with high probability.

First, we will obtain the SDE corresponding to \(\frac{1}{2}\tilde{r}_{y, v}(\tilde{X}_t)^2\). Instead of directly studying this process, we will consider an auxiliary process \(Y_t\) which lower bounds \(\frac{1}{2}\tilde{r}_{y, v}(\tilde{X}_t)^2\), i.e. \[\mathbb{P}\Big[Y_t \leq \frac{1}{2}\tilde{r}_{y, v}(\tilde{X}_t)^2\Big] = 1,\] before \(\tilde{X}_t\) escapes the saddle points. For this, we will use a simple comparison result for SDEs (see 11.3). Finally, we will prove that \(Y_t\) becomes sufficiently large with high probability, which will in turn allow us to obtain an upper bound on \[\mathbb{P}[\tau^* \geq t | \tilde{X}_0 = x].\]

Proof of 15. First of all, note that we are assuming that \(\frac{a}{\sqrt{\beta}} \leq \frac{D}{3}\). Thus, as the critical points of \(\tilde{F}\) are isolated by 15 we can write \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\) as a disjoint union \[U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big) = \bigsqcup_{y_i \in \tilde{\mathcal{S}}} B\big(y_i, \frac{a}{\sqrt{\beta}}\big),\] where each ball \(B\big(y_i, \frac{a}{\sqrt{\beta}}\big)\) is a connected component of \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\). Now, let \(x \in U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\). This decomposition implies that there exists one single \(y_i \in \tilde{\mathcal{S}}\) such that \(x \in \mathcal{B}\Big(\frac{a}{\sqrt{\beta}}, y_i\Big)\). Let us denote \(y := y_i\) for simplicity.

Moreover, as we are also assuming that \(\frac{a}{\sqrt{\beta}} \leq i(B)\), it holds that \[\tau^* = \tau_{\mathcal{B}(\frac{a}{\sqrt{\beta}}, y)^c} \leq \tau_{\mathcal{B}(i(B), y)^c}\] and so we can always study \(\tilde{X}_t\) in normal coordinates centered at \(y\).

Let \(v \in T_y B\) be the unitary eigenvector of \(\nabla^2 \tilde{F}(y)\) that corresponds to the minimum eigenvalue of \(\nabla^2 \tilde{F}(y)\). By 16 we know that \(\nabla^2 \tilde{F}(y)[v, v] \leq -\lambda_*\).

As mentioned earlier, we will study the process \(\tilde{r}_{y, v}(\tilde{X}_t)\) instead of \(d_h(y, \tilde{X}_t)\). We wish to show that with high probability, \(\tilde{r}_{y, v}(\tilde{X}_t)\) is larger than \(\frac{a}{\sqrt{\beta}}\). By 8 , escape of \(\tilde{X}_t\) from \(B\big(y, \frac{a}{\sqrt{\beta}}\big)\) thus occurs with high probability.

For the sake of clarity, we will use \(\tilde{r}\) as a shorthand for \(\tilde{r}_{y, v}\). Using Itô’s formula (see 98) and the standard chain rule, we know that before \(\tilde{X}_t\) leaves \(\mathcal{B}(i(B), y)\), \(\frac{1}{2}\tilde{r}(\tilde{X}_t)^2\) solves the following SDE; \[\begin{align} \label{originalSDE} \begin{aligned} d\left[\frac{1}{2}\tilde{r}(\tilde{X}_t)^2\right] &= \tilde{\operatorname{L}}\left[\frac{1}{2}\tilde{r}(\tilde{X}_t)^2\right]dt + \sqrt{\frac{2}{\beta}} \Big|\mathrm{grad}_{h}\,\frac{1}{2}\tilde{r}(\tilde{X}_t)^2\Big|_hdB_t \\&= \left(\langle -\mathrm{grad}_{h}\,\tilde{F}(\tilde{X}_t), \frac{1}{2} \mathrm{grad}_{h}\,\tilde{r}(\tilde{X}_t)^2\rangle_h + \frac{1}{2\beta} \Delta_h \tilde{r}(\tilde{X}_t)^2\right)dt \\&\quad + \sqrt{\frac{2}{\beta}} \Big|\mathrm{grad}_{h}\,\frac{1}{2}\tilde{r}(\tilde{X}_t)^2\Big|_hdB_t \\&= \left(\langle -\mathrm{grad}_{h}\,\tilde{F}(\tilde{X}_t), \tilde{r}(\tilde{X}_t)\mathrm{grad}_{h}\,\tilde{r}(\tilde{X}_t) \rangle_h \right. \\&\quad\quad \left. + \frac{1}{\beta}\left(\tilde{r}(\tilde{X}_t)\Delta_h \tilde{r}(\tilde{X}_t) + |\mathrm{grad}_{h}\,\tilde{r}(\tilde{X}_t)|^2_h\right)\right)dt \\&\quad + \sqrt{\frac{2}{\beta}} |\tilde{r}(\tilde{X}_t)||\mathrm{grad}_{h}\,\tilde{r}(\tilde{X}_t)|_h dB_t, \end{aligned} \end{align}\tag{9}\] where \(B_t\) is a one-dimensional Brownian motion in \(\mathbb{R}\).

With the aim of controlling the first escaping time of \(\tilde{X}_t\) from \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\), it would be useful to obtain an upper bound on \[\label{eq:probtilderpeq} \mathbb{P}\Big[\frac{1}{2}\tilde{r}(\tilde{X}_t)^2 < \frac{a^2}{2\beta}\Big],\tag{10}\] since whenever \(\frac{1}{2}\tilde{r}(\tilde{X}_t)^2 \geq \frac{a^2}{2\beta}\), it holds that \(\frac{1}{2}d_h(\tilde{X}_t, \tilde{\mathcal{S}}) \geq \frac{a^2}{2\beta}\), and so it must be the case that \(\tau^* \leq t\). Therefore, having an upper bound for the probability shown in 10 which decays exponentially with respect to \(t\) would allow us to finish the proof, as we will see later.

Nevertheless, to the best of our knowledge, with the SDE shown in 9 it is far from obvious to obtain upper bounds for 10 . To circumvent this problem, we will obtain a stochastic process which lower bounds \(\frac{1}{2}\tilde{r}(\tilde{X}_t)^2\). To find such a process, we will use standard comparison results for SDEs (cf. 11.3). Indeed, if we are able to lower bound the drift term of 9 , we will be able to define a new SDE for a process which lower bounds \(\frac{1}{2}\tilde{r}(\tilde{X}_t)^2\).

The drift term of the SDE shown in 9 is given by \[\label{eq:originaldrift} \mathcal{D}(\tilde{X}_t) := \langle -\mathrm{grad}_{h}\,\tilde{F}(\tilde{X}_t), \tilde{r}(\tilde{X}_t)\mathrm{grad}_{h}\,\tilde{r}(\tilde{X}_t) \rangle_h + \frac{1}{\beta}\left(\tilde{r}(\tilde{X}_t)\Delta_h \tilde{r}(\tilde{X}_t) + |\mathrm{grad}_{h}\,\tilde{r}(\tilde{X}_t)|^2_h\right).\tag{11}\] Let us first lower bound the first summand of this expression. Because parallel transport is an isometry, we can write \[\begin{align} \label{eq:auxbound1} \begin{aligned} \langle -\mathrm{grad}_{h}\,\tilde{F}(x), \mathrm{grad}_{h}\,\tilde{r}(x)\rangle_h &= \langle -\mathrm{P}_{\log_y x}\,H(x), \mathrm{grad}_{h}\,\tilde{r}(x)\rangle_h \\&\quad -\langle \mathrm{grad}_{h}\,\tilde{F}(x) - \mathrm{P}_{\log_y x}\, H(x), \mathrm{grad}_{h}\,\tilde{r}(x)\rangle_h \\&\geq \langle -H(x), \mathrm{P}^{-1}_{\log_y x}\, \mathrm{grad}_{h}\,\tilde{r} (x)\rangle_h \\&\quad - |\mathrm{P}^{-1}_{\log_y x}\,\mathrm{grad}_{h}\,\tilde{F}(x) - H(x)|_h |\mathrm{grad}_{h}\,\tilde{r}(x)|_h \\&\geq \langle -H(x), \mathrm{P}^{-1}_{\log_y x}\,\mathrm{grad}_{h}\,\tilde{r}(x)\rangle_h - \frac{A_3 a^2}{2\beta} |\mathrm{grad}_{h}\,\tilde{r}(x)|_h, \end{aligned} \end{align}\tag{12}\] for every \(x \in \mathcal{B}(\frac{a}{\sqrt{\beta}}, y)\), where \(\mathrm{P}_{\log_y x}\) denotes the parallel transport from \(y\) to \(x\) along the unique geodesic joining the two points. The last inequality follows from 4.

Moreover, since \(\frac{a}{\sqrt{\beta}} \leq \min\{i(M), i(B), \frac{1}{72 \dim(M)^{5/2} \mathbf{K}}\}\) by assumption, we can apply 32 to conclude that8 \[\begin{align} \label{eq:auxbound2} \begin{aligned} |\mathrm{grad}_{h}\,\tilde{r}(x)|^2_h &\leq 2,\\ |\mathrm{grad}_{h}\,\tilde{r}(x)|^2_h + \tilde{r}(x)\Delta_h\tilde{r}(x) &\geq \frac{1}{2}, \end{aligned} \end{align}\tag{13}\] for every \(x \in \mathcal{B}(\frac{a}{\sqrt{\beta}}, y)\), and so, inserting 12 13 into 11 shows that \[\mathcal{D}(x) \geq \tilde{r}(x) \left(\langle -H(x), \mathrm{P}^{-1}_{\log_y x}\,\mathrm{grad}_{h}\,\tilde{r}(x)\rangle_h - \frac{A_3 a^2}{\beta}\right) + \frac{1}{2\beta},\] for every \(x \in \mathcal{B}(\frac{a}{\sqrt{\beta}}, y)\).

Lastly, as \(\frac{a}{\sqrt{\beta}} \leq \frac{1}{\sqrt{2}\dim(M)^{3/2} \mathbf{K}}\), using 33 we can further bound \[\langle -H(x), \mathrm{P}^{-1}_{\log_y x}\, \mathrm{grad}_{h}\,\tilde{r}(x)\rangle_h \geq \lambda_* \tilde{r}(x) - 4\dim(M)^5 A_2 \mathbf{K} \frac{a^3}{\beta\sqrt{\beta}},\] for every \(x \in \mathcal{B}(\frac{a}{\sqrt{\beta}}, y)\), and conclude that \[\begin{align} \mathcal{D}(x) &\geq \tilde{r}(x) \left(\lambda_* \tilde{r}(x) - 4\dim(M)^5 A_2 \mathbf{K} \frac{a^3}{\beta\sqrt{\beta}}- \frac{A_3 a^2}{\beta} \right) + \frac{1}{2\beta} \\&= \lambda_* \tilde{r}(x)^2 - \tilde{r}(x)\left(4\dim(M)^5 A_2 \mathbf{K} \frac{a^3}{\beta\sqrt{\beta}}+ \frac{A_3 a^2}{\beta}\right) + \frac{1}{2\beta}, \end{align}\] for every \(x \in \mathcal{B}(\frac{a}{\sqrt{\beta}},y)\).

At this point, we can use the fact that \(\tilde{r}(x) < \frac{a}{\sqrt{\beta}}\) for every \(x \in \mathcal{B}(\frac{a}{\sqrt{\beta}},y)\) to reach \[\frac{1}{2\beta}- \tilde{r}(x)\left(4\dim(M)^5 A_2 \mathbf{K} \frac{a^3}{\beta\sqrt{\beta}}+ \frac{A_3 a^2 }{\beta}\right) > \frac{1}{\beta} \left(\frac{1}{2}- \left(4\dim(M)^5 A_2 \mathbf{K} \frac{a^4}{\beta}+ \frac{A_3 a^3 }{\sqrt{\beta}}\right)\right).\]

As we are also choosing \[\beta \geq 72^2 \dim(M)^5 A_2 A^2_3 \mathbf{K}^2 a^6 \geq \max\{32 \dim(M)^5 A_2 \mathbf{K}a^4, 64 A_3^2a^6\},\] we find that \[\left(\frac{1}{2}- \left(4\dim(M)^5 A_2 \mathbf{K} \frac{a^4}{\beta}+ \frac{A_3 a^3 }{\sqrt{\beta}}\right)\right) \geq \frac{1}{2} - \Big(\frac{1}{8} + \frac{1}{8}\Big) = \frac{1}{4}.\] This bound implies that, before \(\tilde{X}_t\) leaves \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\), it holds that \[\label{ineqdttermsSDE} \mathcal{D}(\tilde{X}_t) \geq \lambda_* \tilde{r}(\tilde{X}_t)^2 + \frac{1}{4\beta}.\tag{14}\]

Using the comparison result for SDEs (see 106), we can obtain a process \(z_t\) such that its square lower bounds \(\tilde{r}(\tilde{X}_t)^2\) before \(X_t\) leaves \(\mathcal{B}(\frac{a}{\sqrt{\beta}}, y)\), i.e. \[\label{eq:comparisonprocesses} \mathbb{P}\Big[\tilde{r}^2(\tilde{X}_t) \geq z_t^2 \, |\, t < \tau_{\mathcal{B}(\frac{a}{\sqrt{\beta}},y)^c}\Big] = 1,\quad \forall t \geq 0.\tag{15}\] The process \(z_t\) is such that \(\frac{1}{2}(z_t)^2\) verifies the SDE \[d\left[\frac{1}{2}(z_t)^2\right] = \left[2 \lambda_* \frac{1}{2} (z_t)^2 + \frac{1}{4\beta} \right]dt + \sqrt{\frac{2}{\beta}} |z_t| \theta_t\, dB_t,\] where \(\theta_t = |\mathrm{grad}_{h}\,\tilde{r}(\tilde{X}_t)|_h\). Choosing \(Y_t = \frac{1}{2}(z_t)^2\), we can rewrite this SDE as \[dY_t = \Big(2\lambda_* Y_t + \frac{1}{4\beta}\Big)dt + 2\sqrt{\frac{Y_t}{\beta}} \theta_t\, dB_t.\]

Recall that in 13 we showed that \(\theta^2_t \leq 2\). Thus, by 102 we know that for every \(A > 0\) and every \(0 < \theta \leq \frac{\lambda_* \beta}{2}\), it holds that \[P[Y_t \leq A] \leq e^{\theta (A - Y_0)} e^{-\frac{\theta}{4\beta}t},\quad \forall t \geq 0.\] Choosing \(\theta = \frac{\lambda_* \beta}{2}\) and \(A = \frac{a^2}{2\beta}\) we can conclude that \[P\Big[Y_t \leq \frac{a^2}{2\beta}\Big] \leq C e^{-\frac{\lambda_*}{8}t},\] with \(C := \exp\Big(\frac{\lambda_* \beta}{2} (\frac{a^2}{2\beta} - Y_0)\Big)\). This bound allows us to conclude the proof: since \[\tau^* \geq t \iff \sup_{s \in [0, t]} \frac{1}{2}d_h(\tilde{X}_s, \tilde{\mathcal{S}})^2 < \frac{a^2}{2\beta},\] we know that \[\mathbb{P}[\tau^* \geq t | \tilde{X}_0 = x] = \mathbb{P}\left[\sup_{s \in [0, t]} \frac{1}{2}d_h(\tilde{X}_s, \tilde{\mathcal{S}})^2 < \frac{a^2}{2\beta} \bigg| \tilde{X}_0 = x\right].\] Now, \[\mathbb{P}\left[\sup_{s \in [0, t]} \frac{1}{2}d_h(\tilde{X}_s, \tilde{\mathcal{S}})^2 < \frac{a^2}{2\beta} \bigg| \tilde{X}_0 = x\right] \leq \mathbb{P}\left[\frac{1}{2}d_h(\tilde{X}_t, \tilde{\mathcal{S}})^2 < \frac{a^2}{2\beta} \bigg| \tilde{X}_0 = x\right].\] Therefore, since with probability one \[Y_t \leq \frac{1}{2}\tilde{r}(\tilde{X}_t)^2 \leq \frac{1}{2}d_h(\tilde{X}_t, S)^2,\] we can conclude that \[\begin{align} \mathbb{P}\left[ \frac{1}{2}d_h(\tilde{X}_t, \tilde{\mathcal{S}})^2 < \frac{a^2}{2\beta} \bigg| \tilde{X}_0 = x\right] &\leq \mathbb{P}\left[\frac{1}{2}\tilde{r}_{y}(\tilde{X}_t)^2 < \frac{a^2}{2\beta} \bigg| \tilde{X}_0 = x \right] \\&\leq \mathbb{P}\left[\left. Y_t < \frac{a^2}{2 \beta} \right| Y_0 = \frac{1}{2}\tilde{r}^2(x) \right] \\&\leq e^{-\frac{\lambda_*}{8} t}C, \end{align}\] where the second inequality follows from the bound of 15 , and \(\lambda_*\) is the bound on the eigenvalues of \(\nabla^2 F\) from 14. ◻

Corollary 2. Let \((M, g)\) be a compact and symmetric manifold. Let \(\pi: (M, g) \to (B, h)\) be a surjective Riemannian submersion with totally geodesic fibers. Assume that there exists a constant \(\mathbf{K} \geq 1\) such that for every \(p\in M\), in normal coordinates centered at \(p\), \(|R_{ijk}^l(p)| \leq \mathbf{K}\) for every \(i, j, k, l \in \{1, \dotsc, \dim(M)\}\), where \(R\) denotes the Riemann curvature tensor of \((M, g)\). Let \(F\) be a smooth function on \(M\) verifying 12 13, and let \(\tilde{F}\) be the unique function on \(B\) satisfying \(F = \tilde{F} \circ \pi\). Assume that \(\tilde{F}\) satisfies .

Let \(a \geq 1\) and let \[\begin{align} \beta &\geq \max\Bigg\{72^2 \dim(M)^5 A_2 A^2_3 \mathbf{K}^2 a^6, \frac{9a^2}{D^2}, \frac{a^2}{i(B)^2}, \frac{a^2}{i(M)^2}\Bigg\}. \end{align}\] Then \[W_2(x) = \mathbb{E}[\exp(\lambda_* \tau^*/16) | \tilde{X}_0 = x],\] defined on \(B\) is a quasi-Lyapunov function on \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\) with parameter \(\theta = \lambda_*/16\).

Proof. In 15 we proved that for every \(x \in U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\), there exists some \(C >0\) such that \[\mathbb{P}[\tau^* \geq t | \tilde{X}_0 = x] \leq C e^{-\lambda_* t / 8},\quad \forall\, t \geq 0\] and so by 2 we know that \[\mathbb{E}\Big[\exp\Big(\frac{\lambda_*}{16} \tau^*\Big) \Big| \tilde{X}_0 = x\Big] < \infty,\] for every \(x \in U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\).

Furthermore, by the choice of \(\beta\), it holds that \(\frac{a}{\sqrt{\beta}} \leq i(B)\) and so, as we saw in 14 it holds without loss of generality that \(\partial U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\) is smooth.

Finally, it also holds without loss of generality that \(-\lambda_*/16 \notin \sigma(\tilde{\operatorname{L}})\), and so all of the conditions of 1 are satisfied, which allows us to conclude the proof. ◻

2.3 Proof of Theorem 10↩︎

Lastly, in this subsection we will prove that the Markov triple \((B, \tilde{\nu}, \tilde{\Gamma})\) satisfies a Poincaré inequality with a—positive—constant that only depends on the value of \(\lambda_*\), which is the positive constant controlling the escape directions from the saddle points (cf. 16), i.e. \[\lambda_{\min}\big(\nabla^2 \tilde{F}(y)\big) \leq -\lambda_* < 0,\quad \forall y \in \tilde{\mathcal{S}}.\]

Let \(a, \beta > 0\) be fixed. We define the following sets in order to lighten the notation. \[\begin{align} \label{eq:subsetsB} \begin{aligned} G_1 &:= \overline{U}\Big(\frac{a}{2\sqrt{\beta}}, \tilde{\mathcal{S}}\Big) = \Big\{x \in B\, |\, d_h(x, \tilde{\mathcal{S}})^2 \leq \frac{a^2}{4\beta}\Big\},\\ G_2 &:= U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big) = \Big\{x \in B \, |\, d_h(x, \tilde{\mathcal{S}})^2 < \frac{a^2}{\beta}\Big\},\\ G_3 &:= \mathcal{B}(\frac{a}{\sqrt{\beta}},x^*) = \Big\{x \in B \, |\, d_h(x, x^*)^2 < \frac{a^2}{\beta}\Big\},\\ G_4 &:= B \setminus (G_1 \cup G_3), \end{aligned} \end{align}\tag{16}\] where \(x^*\) is the unique minimum of \(\tilde{F}\). See 1 for a visual representation.

Figure 1: Sketch of the subsets defined in 16 . We have used stripped filling to denote that some region is included in two subsets. The critical points are isolated by 15, after choosing a sufficiently small value for \frac{a}{\sqrt{\beta}}.

2.3.1 Local Poincaré inequality near the global minimum↩︎

Let us derive a Poincaré inequality for \((G_3, \left.\tilde{\nu}\right|_{G_3}, \tilde{\Gamma}|_{G_3})\), where \(\left.\tilde{\nu}\right|_{G_3}\) denotes the Gibbs measure restricted to \(G_3\), \[\tilde{\nu}|_{G_3}(x) := \frac{\tilde{\nu}(x)}{\int_{G_3} \tilde{\nu}(y)dy} \mathbb{1}_{G_3}(x),\] and \(\tilde{\Gamma}|_{G_3}\) denotes the operator \(\tilde{\Gamma}\) restricted to act on functions \(f \in C^2(G_3)\). This inequality will follow from the geodesic convexity of \(G_3\) along with the so-called curvature-dimension condition, \[\nabla^2\tilde{F} + \frac{1}{\beta}\mathrm{Ric}_h \geq \kappa h,\] with \(\kappa > 0\). In fact, whenever this inequality holds on a geodesically convex subset \(A \subset B\), the triple \((A, \left.\tilde{\nu}\right|_A, \tilde{\Gamma}|_A)\) satisfies a \(\mathrm{PI}(\kappa/2)\). We refer the reader to 10 for more details. The geodesic convexity of the subset \(A\) is crucial if one wants to directly obtain a Poincaré inequality from the curvature-dimension condition. For an in-depth study on how the convexity of a subset affects the interplay between the curvature-dimension condition and the Poincaré inequality, we refer to [35], [41], [42].

In the following statement, which is analogous to [18], we will obtain a Poincaré inequality for \((G_3, \left.\tilde{\nu}\right|_{G_3}, \tilde{\Gamma}|_{G_3})\) by using the curvature-dimension condition, as \(G_3\) can always be assumed to be convex by taking \(\frac{a}{\sqrt{\beta}}\) sufficiently small. Note that, unlike in [18], we are assuming that \(\tilde{F}\) only has one global minimum in \(B\) which simplifies the proof.

Lemma 5 (Poincaré Inequality on \(G_3\)). Suppose that \(\tilde{F}\) satisfies and assume that the Ricci curvature of \((B, h)\) is lower-bounded by some non-positive constant, i.e. \(\mathrm{Ric}_{h} \geq -R_{B}\) for some \(R_{B} \geq 0\). Then for all choices of \(\beta > 0\) such that \[\beta \geq \max \left\{a^2\frac{4A^2_3}{\lambda_*^2}, \frac{4R_{B}}{\lambda_*}, \frac{a^2}{\mathit{conv}(B)^2}\right\},\] where \(\lambda_*\) is defined in 14, and \(\mathit{conv}(B)\) denotes the convexity radius of \((B, h)\), it holds that the Markov triple \((G_3, \left.\tilde{\nu}\right|_{G_3}, \tilde{\Gamma}|_{G_3})\), satisfies a \(\mathrm{PI}(\frac{\lambda_*}{8})\), i.e. \[\int_{G_3} f^2\,d\nu|_{G_3} + \Big(\int_{G_3} f \, d\nu|_{G_3}\Big)^2 \leq \frac{8}{\lambda_* \beta} \int_{G_3} |\mathrm{grad}_{h}\,f|_h^2\, d\nu|_{G_3},\] for every \(f \in C^2(G_3) \cap H^1(G_3)\), where \(H^1\) denotes the Sobolev space of \(L^2\) functions on \(G_3\) with \(L^2\) gradient (see 9).

Proof. Let \(x^*\) be the unique minimum of \(\tilde{F}\). From the definition of \(\tilde{F}\) having an \(A_3\)-Lipschitz Hessian (cf. 43) we know that for every \(x \in G_3\) \[\max_{\substack{s \in T_{x^*}B\\|s|_h=1}} |\mathrm{P}^{-1}_{\log_{x^*}x} \circ \nabla^2 \tilde{F}(x) \circ \mathrm{P}_{\log_{x^*}x}[s] - \nabla^2 \tilde{F}(x^*)[s]|_h \leq A_3 d_h(x,x^*),\] where \(\mathrm{P}_{\log_{x^*}x}\) denotes the parallel transport from \(x^*\) to \(x\) along the unique geodesic joining the two points. Since the parallel transport is an isometry, we can use Weyl’s inequality to conclude that for every \(x \in G_3\), \[\begin{align} |\lambda_1(\nabla^2 \tilde{F}(x)) - \lambda_1(\nabla^2 \tilde{F}(x^*))| &\leq \max_{\substack{s \in T_{x^*}M\\|s|_h=1}} |\mathrm{P}^{-1}_{\log_{x^*}x} \circ \nabla^2 \tilde{F}(x) \circ \mathrm{P}_{\log_{x^*}x}[s] - \nabla^2 \tilde{F}(x^*)[s]|_h \\&\leq A_3 d_h(x,x^*), \end{align}\] where \(\lambda_1(\nabla^2\tilde{F}(x))\) denotes the first eigenvalue of the Hessian of \(\tilde{F}\) at \(x\). This way, we can apply 17 to conclude that \[\lambda_1(\nabla^2 \tilde{F}(x)) \geq \lambda_1(\nabla^2 \tilde{F}(x^*)) - A_3 d_h(x, x^*) \geq \frac{\lambda_*}{2},\quad \mathrm{whenever } d_h(x, x^*) \leq \frac{\lambda_*}{2A_3}.\] The above inequality is fulfilled on \(G_3 = \{x \in B : d_h(x, x^*)^2 < \frac{a^2}{\beta}\}\), as by assumption \(\frac{a^2}{\beta} \leq \frac{\lambda^2_*}{4 A^2_3}\).

For this reason, and the fact that we assumed that \(\mathrm{Ric}_h \geq -R_B\) where \(R_B \geq 0\), then, given our choice of \(\beta\), \[\nabla^2 \tilde{F}(x) + \frac{1}{\beta}\mathrm{Ric}_h \geq \left(\frac{\lambda_*}{2} - \frac{1}{\beta}R_B\right)h \geq \frac{\lambda_*}{4}h.\]

Since we are ensuring that \(\frac{a}{\sqrt{\beta}} \leq \mathit{conv}(B)\), we know that \((G_3, \left.\tilde{\nu}\right|_{G_3}, \tilde{\Gamma}|_{G_3})\) satisfies a \(\mathrm{PI}(\lambda_*/8)\). See 97 for more details. ◻

2.3.2 Extending the Poincaré inequality↩︎

Having obtained a local Poincaré inequality on \(G_3\), in this subsection we will extend it to obtain a Poincaré inequality for \((B, \tilde{\nu}, \tilde{\Gamma})\). This will be done using the Lyapunov functions constructed in 2.2, \[W_1(x) = \exp\Big(\frac{\beta}{2}\tilde{F}(x)\Big),\quad \mathrm{and}\quad W_2(x) = \mathbb{E}[\exp(\lambda_*\, \tau^*/16) | \tilde{X}_0 = x].\]

In order to extend our Poincaré inequality from \(G_3\) to \(B\), smooth bump functions will come handy. The following two lemmas allow to construct one that vanishes on \(G_1\), is equal to one on \(G_2^c\), and such that its derivative is upper bounded by a suitable constant.

Lemma 6 ([39]). Let \(y \in B\) be some fixed point. The distance function \(x \mapsto d_h(x, y)\) is a smooth function on \(\mathcal{B}(i(B),y) \setminus \{y\}\) with unit gradient.

Lemma 7. Let \[r := \frac{a}{\sqrt{\beta}} \leq \min\Big\{i(B), \frac{D}{3}\Big\},\] where \(D\) is the lower bound on the distance between any two saddle points of \(\tilde{F}\) (cf. 15). Then for every \(\varepsilon > 0\) there exists a smooth partition function \(\chi: B \to [0, 1]\) such that \(\chi \equiv 1\) on \(G_2^c\), \(\chi \equiv 0\) on \(G_1\) and such that \[\norm{\tilde{\Gamma}(\chi)}_\infty \leq \frac{1}{\beta}\Big(\frac{2}{r} + \varepsilon\Big)^2.\]

Proof. First, we can define a smooth bump function \(\psi: \mathbb{R} \to [0, 1]\) such that \[\psi(x) = \begin{cases} 0, \quad & x \leq r/2,\\ \mathit{increasing}, \quad &x\in (r/2, r),\\ 1,\quad & x\geq r, \end{cases}\] with \(\norm{\psi'}_{\infty} \leq \frac{2}{r} + \varepsilon\) following standard methods (cf. [18]). Now, we define \(\chi: B \to [0, 1]\) as \[\chi(x) := \psi(d_h(x,\tilde{\mathcal{S}})).\] Note that \(\psi(x)\) is constant in the interior of \(G_1\) and in the complement of \(G_2\). Furthermore, since we are assuming that \(r \leq i(B)\), it holds by 6 that the distance function is smooth in \(G_2 \setminus \tilde{\mathcal{S}}\), and so \(\chi\) is smooth in \(B\) and \[\mathrm{grad}_{h}\,\chi = \mathrm{grad}_{h}\,\psi(d_h(x, \tilde{\mathcal{S}})) = \psi'(d_h(x, S)) \mathrm{grad}_{h}\,d_h(x, \tilde{\mathcal{S}}),\] for every \(x \in B\) such that \(r/2 < d_h(x, \tilde{\mathcal{S}}) < r\) and \(\mathrm{grad}_{h}\,\chi \equiv 0\) elsewhere. Again, by 6, the distance function has a gradient with constant norm one, and so it holds that \[|\mathrm{grad}_{h}\,\chi(x)|_h = |\psi'(d_h(x, \tilde{\mathcal{S}}))|_h\] on \(B\).

For this reason, for every \(x \in B\) such that \(r/2 < d_h(x, \tilde{\mathcal{S}}) < r\), \[\lVert\tilde{\Gamma}(\chi(x))\rVert_{\infty} = \frac{1}{\beta} \lVert\mathrm{grad}_{h}\,\chi(x)\rVert^2_\infty = \frac{1}{\beta}\lVert\psi'(d_h(x, \tilde{\mathcal{S}}))\rVert^2_\infty \leq \frac{1}{\beta}\Big(\frac{2}{r} + \varepsilon\Big)^2.\] ◻

Our final tool to derive a Poincaré inequality is an integration by parts formula on subsets of \(B\).

Lemma 8 (Neumann boundary condition). Let \(\tilde{\operatorname{L}}\) and \(\tilde{\nu}\) be as defined in 3 , and let \(N\) be a submanifold of \(B\) with boundary. For every pair of functions \(W\), and \(\phi\) in \(C^2(N)\) such that \(\phi\) vanishes at \(\partial N\), it holds that \[\int_N \phi \tilde{\operatorname{L}}W \,d\tilde{\nu} = - \frac{1}{\beta} \int_N \langle \mathrm{grad}_{h}\,\phi,\mathrm{grad}_{h}\,W\rangle_h \,d\tilde{\nu} .\]

Proof. By the definition of \(\tilde{\nu}\) and \(\tilde{\operatorname{L}}\), \[\begin{align} \int_N \phi \tilde{\operatorname{L}}W \,d\tilde{\nu} &= \int_N \phi\langle-\mathrm{grad}_{h}\,\tilde{F}, \mathrm{grad}_{h}\,W\rangle_h \,d\tilde{\nu} + \frac{1}{\beta} \int_N \phi \Delta_h W \,d\tilde{\nu} \\&= \int_N \phi\langle-\mathrm{grad}_{h}\,\tilde{F}, \mathrm{grad}_{h}\,W\rangle_h \,d\tilde{\nu} + \frac{1}{\beta} \int_N \phi e^{-\beta \tilde{F}}\Delta_h W\; \mathrm{dVol}_h. \end{align}\] Using Green’s identity, since we are assuming that \(\phi\) vanishes at \(\partial N\), it holds that \[\begin{align} \int_N \phi \tilde{\operatorname{L}}W \,d\tilde{\nu} &= \int_N \phi\langle-\mathrm{grad}_{h}\,\tilde{F}, \mathrm{grad}_{h}\,W\rangle_h \,d\tilde{\nu} - \frac{1}{\beta} \int_N \langle \mathrm{grad}_{h}\,(\phi e^{-\beta \tilde{F}}),\mathrm{grad}_{h}\,W\rangle_h\; \mathrm{dVol}_h \\&= \int_N \phi\langle-\mathrm{grad}_{h}\,\tilde{F}, \mathrm{grad}_{h}\,W\rangle_h \,d\tilde{\nu} - \frac{1}{\beta} \int_N \langle \mathrm{grad}_{h}\,\phi,\mathrm{grad}_{h}\,W\rangle_h \,d\tilde{\nu} \\&\quad + \int_N \phi \langle \mathrm{grad}_{h}\,\tilde{F},\mathrm{grad}_{h}\,W\rangle_h \,d\tilde{\nu} \\&= - \frac{1}{\beta} \int_N \langle \mathrm{grad}_{h}\,\phi,\mathrm{grad}_{h}\,W\rangle_h \,d\tilde{\nu}. \end{align}\] ◻

We are now ready to prove 10. The proof strategy adopted in the setting considered in [17] remains valid in ours. We give here the full proof for the sake of completeness.

Proof of 10. Let \(f \in C^2(B)\). Without loss of generality, we can assume that \(\int_{G_3} f \, d\tilde{\nu} = 0\), as \[\int_B f^2 \, d\tilde{\nu} - \Big(\int_B f \,d\tilde{\nu} \Big)^2 \leq \int_B (f + c)^2 \, d\tilde{\nu},\] for every \(c \in \mathbb{R}\). Our goal will be to upper bound the right-hand side of this inequality by some multiple of \(\int_B |\mathrm{grad}_{h}\,f|^2_h \,d\tilde{\nu}\) in order to obtain a Poincaré inequality.

Let \(\chi\) be the bump function from 7. Since \(\chi \equiv 0\) in \(G_1\) and \(\chi \equiv 1\) in \(G_2^c\), we know that \[\begin{align} \label{firstinequality} \begin{aligned} \int_B f^2 \,d\tilde{\nu} &= \int_B (f(1 - \chi) + f \chi)^2 \,d\tilde{\nu} \\&\leq 2\int_B f^2(1-\chi)^2 \, d\tilde{\nu} + 2\int_B f^2\chi^2 \, d\tilde{\nu} \\&= 2\int_{G_2} f^2(1-\chi)^2 \, d\tilde{\nu} + 2\int_{G_1^c} f^2\chi^2 \, d\tilde{\nu}. \end{aligned} \end{align}\tag{17}\]

In 2 we proved that \(W_2 \in C^\infty(\overline{G_2}) \cap C^\infty(G_2^c)\), \(W_2 \geq 1\) and \[\label{eq:lyapunovW2} \frac{\tilde{\operatorname{L}}W_2(x)}{W_2(x)} = -\frac{\lambda_*}{16},\quad \forall x \in G_2.\tag{18}\] Using 18 and the Neumann boundary condition from 8 we can bound the first summand of 17 as follows: \[\begin{align} \int_{G_2} f^2(1-\chi)^2 \,d\tilde{\nu} &= -\frac{16}{\lambda_*}\int_{G_2} \frac{\tilde{\operatorname{L}}W_2}{W_2}f^2(1-\chi)^2 \,d\tilde{\nu} \\&= \frac{16}{\lambda_* \beta} \int_{G_2} \Big\langle \mathrm{grad}_{h}\,\frac{f^2(1-\chi)^2}{W_2}, \mathrm{grad}_{h}\,W_2\Big\rangle_h \,d\tilde{\nu}, \end{align}\] since \((1-\chi)\) vanishes at \(\partial G_2\). Using the chain rule we can further bound this expression, \[\begin{align} \label{eq:ineqchainrule} \begin{aligned} \frac{16}{\lambda_* \beta} \int_{G_2} \Big\langle \mathrm{grad}_{h}\,\frac{f^2(1-\chi)^2}{W_2}, \mathrm{grad}_{h}\,&W_2\Big\rangle_h \,d\tilde{\nu} \\&= \frac{32}{\lambda_*\beta} \int_{G_2} \frac{f(1-\chi)}{W_2}\langle \mathrm{grad}_{h}\,f(1-\chi), \mathrm{grad}_{h}\,W_2\rangle_h \,d\tilde{\nu} \\&\quad - \frac{16}{\lambda_*\beta}\int_{G_2} \frac{f^2(1-\chi)^2}{W_2^2} |\mathrm{grad}_{h}\,W_2|^2_h \,d\tilde{\nu} \\&= \frac{16}{\lambda_* \beta} \int_{G_2} |\mathrm{grad}_{h}\,f(1-\chi)|^2_h \,d\tilde{\nu} \\&\quad- \frac{16}{\lambda_* \beta} \int_{G_2} \Big|\mathrm{grad}_{h}\,f(1-\chi)- \frac{f(1-\chi)}{W_2}\mathrm{grad}_{h}\,W_2\Big|^2_h \,d\tilde{\nu} \\&\leq \frac{16}{\lambda_* \beta} \int_{G_2} |\mathrm{grad}_{h}\,f(1-\chi)|^2_h \,d\tilde{\nu} \\&= \frac{16}{\lambda_*}\int_{G_2} \tilde{\Gamma}(f(1-\chi)) \,d\tilde{\nu}. \end{aligned} \end{align}\tag{19}\] Lastly, using the fact that \[\label{productGamma} \tilde{\Gamma}(fg) \leq 2(f^2 \tilde{\Gamma}(g) + g^2 \tilde{\Gamma}(f)),\quad \forall f,g \in C^2(B),\tag{20}\] it holds that \[\begin{align} \frac{16}{\lambda_*}\int_{G_2} \tilde{\Gamma}(f(1-\chi)) \,d\tilde{\nu} &\leq \frac{32}{\lambda_*} \left(\int_{G_2} f^2 \tilde{\Gamma}(\chi) \,d\tilde{\nu} + \int_{G_2} (1-\chi)^2 \tilde{\Gamma}(f) \,d\tilde{\nu}\right) \\&\leq \frac{32\lVert\tilde{\Gamma}(\chi)\rVert_\infty}{\lambda_*} \int_{G_2} f^2 \,d\tilde{\nu} + \frac{32}{\lambda_*}\int_{G_2} \tilde{\Gamma}(f) \,d\tilde{\nu} \\&\leq \frac{32\lVert\tilde{\Gamma}(\chi)\rVert_\infty}{\lambda_*} \int_B f^2 \,d\tilde{\nu} + \frac{32}{\lambda_*}\int_B \tilde{\Gamma}(f) \,d\tilde{\nu}. \end{align}\]

Thus, we have concluded that \[\label{conclusion1} 2\int_{G_2} f^2(1-\chi)^2 d\tilde{\nu} \leq \frac{64\lVert\tilde{\Gamma}(\chi)\rVert_\infty}{\lambda_*} \int_B f^2 d\tilde{\nu} + \frac{64}{\lambda_*}\int_B \tilde{\Gamma}(f)d\tilde{\nu}.\tag{21}\]

Moreover, in 12 we proved that \(W_1\) is smooth, \(W_1 \geq 1\) and \[\label{eq:lyapunovW1} \frac{\tilde{\operatorname{L}}W_1(x)}{W_1(x)} \leq -A_2 \dim(B) + \frac{3}{2}A_2 \dim(B) \mathbb{1}_{G_3}(x), \quad \forall x \in G_1^c \cup G_3.\tag{22}\] Thus, the second summand of 17 can be bounded similarly, using 22 and 8, since \(\chi\) vanishes on \(\partial G_1\). Recall that \(G_1^c = G_3 \cup G_4\). Thus, \[\begin{align} \int_{G_1^c} f^2 \chi^2 \,d\tilde{\nu} &\leq \frac{1}{A_2 \dim(B)} \int_{G_1^c} \frac{-\tilde{\operatorname{L}}W_1}{W_1} f^2 \chi^2 \,d\tilde{\nu} + \frac{3}{2} \int_{G_3} f^2 \,d\tilde{\nu} \\&= \frac{1}{A_2 \beta \dim(B)} \int_{G_1^c} \Big\langle \mathrm{grad}_{h}\,\frac{f^2 \chi^2}{W_1}, \mathrm{grad}_{h}\,W_1\Big\rangle_h \,d\tilde{\nu} + \frac{3}{2} \int_{G_3} f^2 \,d\tilde{\nu} \\&\leq \frac{1}{A_2 \dim(B)} \int_{G_1^c} \tilde{\Gamma}(f\chi) \,d\tilde{\nu} + \frac{3}{2}\int_{G_3} f^2 \,d\tilde{\nu}, \end{align}\] where the last inequality is obtained following an analogous procedure as in 19 . Now, since we already proved that \((G_3, \left.\tilde{\nu}\right|_{G_3}, \tilde{\Gamma}|_{G_3})\) satisfies a \(\mathrm{PI}(\frac{\lambda_*}{8})\), it holds that \[\int_{G_3} f^2 \,d\tilde{\nu} = \frac{Z_{G_3}}{Z_B} \int_{G_3} f^2 \,\left.d\tilde{\nu}\right|_{G_3} \leq \frac{Z_{G_3}}{Z_B} \frac{8}{\lambda_*}\int_{G_3} \tilde{\Gamma}(f) \, \left.d\tilde{\nu}\right|_{G_3} = \frac{8}{\lambda_*}\int_{G_3} \tilde{\Gamma}(f) \,d\tilde{\nu}.\] Thus, \[\begin{align} \int_{G_1^c} f^2 \chi^2 \,d\tilde{\nu} &= \frac{1}{A_2 \dim(B)} \int_{G_1^c} \tilde{\Gamma}(f\chi) \,d\tilde{\nu} + \frac{3}{2}\int_{G_3} f^2 \,d\tilde{\nu} \\&\leq \frac{1}{A_2 \dim(B)} \int_{G_1^c} \tilde{\Gamma}(f\chi) \,d\tilde{\nu} + \frac{12}{\lambda_*}\int_{G_3} \tilde{\Gamma}(f) \,d\tilde{\nu} \\&\leq \frac{1}{A_2 \dim(B)} \int_{B} \tilde{\Gamma}(f\chi) \,d\tilde{\nu} + \frac{12}{\lambda_*}\int_B \tilde{\Gamma}(f) \,d\tilde{\nu}. \end{align}\] Using 20 we can upper bound this expression, \[\begin{align} \frac{1}{A_2 \dim(B)} \int_{B} \tilde{\Gamma}(f\chi) \,d\tilde{\nu} + \frac{12}{\lambda_*}\int_B \tilde{\Gamma}(f) \,d\tilde{\nu} &\leq \frac{2\lVert\tilde{\Gamma}(\chi)\rVert_\infty}{A_2 \dim(B)} \int_{B} f^2 \,d\tilde{\nu} \\&\quad + \left(\frac{2}{A_2 \dim(B)}+ \frac{12}{\lambda_*}\right) \int_{B} \tilde{\Gamma}(f) \,d\tilde{\nu}, \end{align}\] which ultimately leads to \[\label{conclusion2} 2\int_{G_1^c} f^2 \chi^2 \,d\tilde{\nu} \leq \frac{4\lVert\tilde{\Gamma}(\chi)\rVert_\infty}{A_2 \dim(B)} \int_{B} f^2 \,d\tilde{\nu} + \left(\frac{4}{A_2 \dim(B)} + \frac{24}{\lambda_*}\right) \int_{B} \tilde{\Gamma}(f) \,d\tilde{\nu}.\tag{23}\]

Inserting 21 23 into 17 , we can conclude that \[\begin{align} \int_B f^2 \,d\tilde{\nu} \leq 4\lVert\tilde{\Gamma}(\chi)\rVert_\infty\left(\frac{16}{\lambda_*} + \frac{1}{A_2 \dim(B)} \right)\int_{B} f^2 \,d\tilde{\nu} + \left(\frac{4}{A_2 \dim(B)} + \frac{88}{\lambda_*}\right) \int_{B} \tilde{\Gamma}(f) \,d\tilde{\nu}. \end{align}\]

Recall that \(\chi\) is such that \[\lVert\tilde{\Gamma}(\chi)\rVert_{\infty} \leq \frac{1}{\beta}\Big(\frac{2}{r} + \varepsilon\Big)^2,\] for every \(\varepsilon > 0\), where \(r = \frac{a}{\sqrt{\beta}}\). Taking \(\varepsilon \to 0\), we get \[\lVert\Gamma(\chi)\rVert_{\infty} \leq \frac{4}{a^2}.\] Therefore, since \(A_2 \dim(B) \geq 1 \geq \lambda_*\), it holds that \[\begin{align} 1 - 4\norm{\Gamma(\chi)}_{\infty} \left(\frac{16}{\lambda_*} + \frac{1}{A_2 \dim(B)} \right) &\geq 1 - 4\norm{\Gamma(\chi)}_{\infty} \frac{17}{\lambda_*} \\&\geq 1 - \frac{272}{a^2 \lambda_*} \\&\geq \frac{1}{2}, \end{align}\] where the last inequality follows from the assumption that \(a^2 \geq \frac{544}{\lambda_*}\). Therefore, we know that \[\int_B f^2 \,d\tilde{\nu} \leq \Big(\frac{8}{A_2 \dim(B)} + \frac{176}{\lambda_*}\Big) \int_{B} \tilde{\Gamma}(f) \,d\tilde{\nu}.\] Lastly, using again that \(A_2 \dim(B) \geq 1 \geq \lambda_*\) we conclude that \[\int_B f^2 \,d\tilde{\nu} \leq \frac{184}{\lambda_*} \int_{B} \tilde{\Gamma}(f) \,d\tilde{\nu},\] and so the claim follows. ◻

3 Poincaré inequalities through Riemannian submersions↩︎

In this section, we will study the Poincaré inequality in the context of Riemannian submersions9. In particular, we will show how, given a Riemannian submersion \(\pi: (M, g) \to (B, h)\), the existence of a Poincaré inequality on the base space \(B\) can be lifted to obtain a Poincaré inequality on the total space \(M\) and vice versa.

Let \((M, g)\) be a complete Riemannian manifold—not necessarily compact10—and let \(G\) be some compact Lie group acting smoothly, freely and isometrically on \(M\). Let \(h\) be the metric on \(M/G\) for which the projection \(\pi: (M, g) \to (M/G, h)\) is a Riemannian submersion. As in [sec:intro] [SectionPI], we consider a smooth function \(F : M \to \mathbb{R}\) which is constant in the fibers of \(\pi\)—i.e. \(F(gx) = F(x)\) for every \(x \in M\) and every \(g \in G\)—and let \(\tilde{F}\) be the unique function on \(B\) satisfying \(F = \tilde{F} \circ \pi\). We again define the Gibbs measures on \(M\) and \(M/G\) as \[\nu = \frac{1}{Z} e^{-\beta F},\quad \tilde{\nu} = \frac{1}{\tilde{Z}} e^{-\beta \tilde{F}}.\]

Let us assume that a Poincaré inequality holds for either \(M\) or \(M/G\). That is, there exists some constant \(\kappa > 0\) such that for every \(f \in C^2(M) \cap H^1(M, \nu)\), \[\int_M f^2 \, d\nu - \Bigg(\int_M f \, d\nu\Bigg)^2 \leq \frac{1}{\kappa\beta} \int_M |\mathrm{grad}_g \;f|^2_g \, d\nu,\] or there exists some constant \(\tilde{\kappa} > 0\) such that for every \(f \in C^2(M/G) \cap H^1(M/G, \tilde{\nu})\), \[\int_{M/G} f^2 \, d\tilde{\nu} - \Bigg(\int_{M/G} f \, d\tilde{\nu}\Bigg)^2 \leq \frac{1}{\tilde{\kappa}\beta} \int_{M/G} |\mathrm{grad}_{h} \;f|^2_{h} \, d\tilde{\nu},\] where \(H^1(M, \nu), H^1(M/G, \tilde{\nu})\) denote the Sobolev spaces11 of square-integrable functions with square-integrable gradient on \(M\) and \(M/G\) with respect to \(\nu\) and \(\tilde{\nu}\), respectively. In either case, assuming that one of the inequalities holds, we will prove that the other holds as well. In order to do so, we will make use of two main results. The first one, which will be discussed in 3.1, can be understood as a version of Fubini’s theorem for Riemannian submersions. The second, which will be shown in 3.2, guarantees that compact manifolds with non-negative Ricci curvature always verify a Poincaré inequality with respect to their normalized Riemannian measure. We will then prove the equivalence between a Poincaré inequality with respect to \(\nu\) in \(M\) and one with respect to \(\tilde{\nu}\) in \(M/G\).

Note that, when \(M\)—and thus \(M/G\)—is compact, \(C^2(M) \cap H^1(M, \nu) = C^2(M)\) and \(C^2(M/G) \cap H^1(M/G, \tilde{\nu}) = C^2(M/G)\), and so the above Poincaré inequalities correspond to the Poincaré inequalities introduced in 6 for the Markov triples \((M, \nu, \Gamma)\) and \((M/G, \tilde{\nu}, \tilde{\Gamma})\), where \[\Gamma(f) = \frac{1}{\beta} |\mathrm{grad}_{g}\,f|^2_g, \quad \tilde{\Gamma}(\tilde{f}) = \frac{1}{\beta} |\mathrm{grad}_{h}\,\tilde{f}|^2_{h},\] for every \(f \in C^2(M)\) and every \(\tilde{f} \in C^2(M/G)\).

3.1 Integrals and Riemannian submersions↩︎

The following theorem will be used to derive a variation of Fubini’s theorem in the context of Riemannian submersions.

Theorem 16 ([43]). Let \(M\) and \(N\) be two smooth manifolds of dimensions \(m\) and \(n\), respectively. Let \(\varphi : M \to N\) be a \(C^1\) function. Assume that the set of critical values of \(\varphi\) has measure zero. Let \(f : M \to \mathbb{R}\) be a measurable function. Let \(\omega \in \Omega^{m-n}(M)\) be an \((m-n)\)-form on \(M\), and let \(\eta \in \Omega^n(N)\) be an \(n\)-form on \(N\). If \(f\) is integrable with respect to \(\omega \wedge \varphi^*\eta\) on \(M\)—where \(\varphi^* \eta\) denotes the pullback of \(\eta\) by \(\varphi\)—then, for almost every \(x \in N\), \[F(x) := \begin{cases}\int_{\varphi^{-1}(x)} f\, \omega, \quad& \mathrm{if } \varphi^{-1}(x) \neq \emptyset,\\ 0\quad& \text{otherwise,}\end{cases}\] is well-defined, measurable, integrable with respect to \(\eta\) and \[\int_M f\; \omega \wedge \varphi^* \eta = \int_N F\; \eta.\] Moreover, if \[\hat{F}(x) := \begin{cases} \int_{\varphi^{-1}(x)} |f|\, |\omega|,\quad& \mathrm{if } \varphi^{-1}(x) \neq \emptyset,\\ 0\quad& \text{otherwise,} \end{cases}\] is integrable with respect to \(\eta\) on \(N\), then \(f\) is integrable with respect to \(\omega \wedge \varphi^* \eta\) on \(M\).

In particular, if we take \(\varphi\) in 16 to be a Riemannian submersion between two Riemannian manifolds, and define \(\omega\) as the volume form of the fibers of \(\varphi\), and \(\eta\) as the volume form of the base space with respect to the induced metric, we obtain the following result.

Corollary 3. Let \(\pi : (M, g) \rightarrow (N, h)\) be a Riemannian submersion. Let us denote by \(\hat{g}(x) := g|_{\pi^{-1}(x)}\) the induced metric on the fiber \(\pi^{-1}(x)\), for every \(x \in N\). For any integrable function \(f\) on \(M\) it holds that \(\int_{\pi^{-1}(x)} f(y)\; \mathrm{dVol}_{\hat{g}(x)}(y)\) is integrable on \(N\) and \[\int_M f(x)\; \mathrm{dVol}_g(x) = \int_N \left(\int_{\pi^{-1}(x)} f(y)\; \mathrm{dVol}_{\hat{g}(x)}(y) \right)\; \mathrm{dVol}_h(x).\] Moreover, if \[\hat{F}(x) := \int_{\pi^{-1}(x)} |f(y)|\,\mathrm{dVol}_{\hat{g}(x)}(y)\] is integrable on \(N\), then \(f\) is integrable on \(M\).

Proof. Let \(\eta = \mathrm{dVol}_{h(\pi(x))}\), and \(\omega = \mathrm{dVol}_{\hat{g}(x)}\). Since \(\pi\) is a Riemannian submersion, it has no critical points and \[\mathrm{dVol}_{\hat{g}(x)} \wedge \pi^* \mathrm{dVol}_{h(\pi(x))} = \mathrm{dVol}_{g(x)},\] which follows from the decomposition of \(g\) into its horizontal and vertical components (cf. 8.4). ◻

Two straightforward consequences of 3 arise when considering \(G\)-bundles with compact isometric fibers. Firstly, we can strengthen the result obtained in 16 to ensure regularity of \[x \mapsto \int_{\pi^{-1}(x)} f(y)\, \mathrm{dVol}_{\hat{g}(x)}(y),\] provided that \(f\) is sufficiently regular. Secondly, given a function \(\tilde{F}\) on \(B\), and \(F\) defined on \(M\) verifying \(F = \tilde{F} \circ \pi\), we can relate the normalizing constant of the Gibbs measure associated with \(F\) with that of the Gibbs measure associated with \(\tilde{F}\) and the volume of \(G\).

Proposition 17 (Upgraded regularity). Consider a principal \(G\)-bundle which is also a Riemannian submersion, \(\pi: (M, g) \to (M/G, h)\). Further assume that the fibers of \(\pi\) are all isometric to some fixed compact Lie group \((G, \hat{g})\). For every integrable function \(f \in C^k(M)\), and for almost every \(x \in M/G\), the function \[\hat{F}(x) := \int_{\pi^{-1}(x)} f(y)\, \mathrm{dVol}_{\hat{g}}(y)\] is well-defined, integrable on \(M/G\) and \(C^k(M/G)\).

Proof. It follows from 16 that \(\hat{F}\) is well-defined and integrable, so it only remains to prove that \(\hat{F} \in C^k(M/G)\). For every \(x_0 \in M/G\) we can consider an open neighborhood \(U\) around \(x_0\) and a local trivialization, so that in local coordinates \[\hat{F}(x) = \int_{\pi^{-1}(x)} f(y) \,\mathrm{dVol}_{\hat{g}}(y) = \int_{G} f(x, y)\, \mathrm{dVol}_{\hat{g}}(y),\] for every \(x \in U\). Now, since \(G\) is compact, if \(f\) is continuous on \(M\), we can apply the dominated convergence theorem to conclude that \(\hat{F}\) is continuous at \(x_0\). Thus, \(\hat{F}\) is continuous in \(M/G\). Furthermore, if \(f\) is assumed to be in \(C^k(M)\), then, using the compactness of \(G\) and the regularity of \(f\), we can differentiate under the integral sign to conclude that \(\hat{F}\) is \(C^k(M/G)\). ◻

Lastly, let us state the variation of Fubini’s theorem in \(G\)-bundles which are also Riemannian submersions, when integrating with respect to the Gibbs measure associated with a function which is constant in the fibers of the submersion \(\pi\). As we mentioned previously, this will allow us to obtain a relation between the normalizing constants of the Gibbs measures and the volume of \(G\).

Proposition 18 (Fubini’s theorem for Gibbs measures). Consider a principal \(G\)-bundle which is also a Riemannian submersion, \(\pi: (M, g) \to (M/G, h)\). Further assume that the fibers of \(\pi\) are all isometric to some compact Lie group \((G, \hat{g})\). Let \(F\) be some function on \(M\) that is constant on the fibers of \(\pi\), and \(\tilde{F}\) be such that \(\tilde{F} \circ \pi = F\). Consider the Gibbs measures \(d\nu(x) = \frac{1}{Z}e^{-\beta F(x)}\mathrm{dVol}_g(x)\) on \(M\) and \(d\tilde{\nu}(x) = \frac{1}{\tilde{Z}}e^{-\beta \tilde{F}(x)}\mathrm{dVol}_{h}(x)\) on \(M/G\). Provided both \(\nu\) and \(\tilde{\nu}\) are probability measures, given any integrable function \(f\) on \(M\) with respect to \(\nu\), the function \[\tilde{f}: x \mapsto \tilde{f}(x) = \int_{\pi^{-1}(x)} f(y)\, \mathrm{dVol}_{\hat{g}}(y)\] is integrable on \(M/G\) with respect to \(\tilde{\nu}\) and \[\int_M f(x) d\nu(x) = \frac{\tilde{Z}}{Z}\int_{M/G} \tilde{f}(x)\, d\tilde{\nu}(x).\]

Proof. The result simply follows by rewriting \[\begin{align} \int_M f(x) d\nu(x) &= \frac{1}{Z}\int_M f(x)e^{-\beta F(x)}\; \mathrm{dVol}_{g}(x) \\&= \frac{1}{Z}\int_{M/G} \left(\int_{\pi^{-1}(x)} f(y)e^{-\beta F(y)}\; \mathrm{dVol}_{\hat{g}}(y) \right)\; \mathrm{dVol}_{h}(x) \\&= \frac{1}{Z}\int_{M/G} \left(\int_{\pi^{-1}(x)} f(y)\; \mathrm{dVol}_{\hat{g}}(y) \right)\; e^{-\beta \tilde{F}(x)}\; \mathrm{dVol}_{h}(x) \\&= \frac{\tilde{Z}}{Z}\int_{M/G} \left(\int_{\pi^{-1}(x)} f(y)\; \mathrm{dVol}_{\hat{g}}(y) \right)\; d\tilde{\nu}(x) \\&= \frac{\tilde{Z}}{Z}\int_{M/G} \tilde{f}(x)\, d\tilde{\nu}(x). \end{align}\] ◻

From this expression, we obtain the following relation between the volume of \(G\) and the normalizing constants \(Z\) and \(\tilde{Z}\) when considering \(f \equiv 1\). \[\label{volumenzetas} 1 = \int_M d\nu(x) = \frac{\tilde{Z}}{Z}\int_{M/G} \mathrm{Vol}_{\hat{g}}(G)\; d\tilde{\nu}(x) = \frac{\tilde{Z}}{Z} \mathrm{Vol}_{\hat{g}}(G).\tag{24}\]

3.2 Lifting and lowering a Poincaré inequality↩︎

The following result allows us to guarantee that every compact manifold satisfies a Poincaré inequality with respect to its normalized Riemannian measure, i.e. the probability measure associated with its volume form, which corresponds to the Gibbs measure at infinite temperature. Note that this result on its own is not useful in our setting; indeed, we are interested in proving the existence of a Poincaré inequality with respect to the Gibbs measure associated with a given function \(F\) at finite temperature, which in general is different from the Riemannian measure. Nevertheless, this result is crucial for the lifting process of the Poincaré inequality, as will be shown in the proof of 20.

Theorem 19 ([44]). Let \((N,h)\) be a compact, connected Riemannian manifold of finite volume \(\mathrm{Vol}(N)\). Assume that \(N\) has non-negative Ricci curvature. Let \(d\mu = \frac{\mathrm{dVol}_h}{\mathrm{Vol}(N)}\) be the normalized Riemannian measure on \((N, h)\) and let \(D\) be its diameter. Then the Markov triple \((N, \mu, |\mathrm{grad}_{h}\,\cdot |_h^2)\) satisfies a \(\mathrm{PI}(\frac{\pi^2}{D^2})\), i.e. for every \(f \in C^2(N)\), \[\int_N f^2\; d\mu - \left(\int_N f\; d\mu \right)^2 \leq \frac{D^2}{\pi^2}\int_N |\mathrm{grad}_{h}\,f|_h^2 \; d\mu,\] which can be rewritten as \[\int_N f^2\; \mathrm{dVol}_h - \frac{1}{\mathrm{Vol}(N)}\left(\int_N f\; \mathrm{dVol}_h \right)^2 \leq \frac{D^2}{\pi^2}\int_N |\mathrm{grad}_{h}\, f|_h^2 \; \mathrm{dVol}_h .\]

We now proceed to state and prove the main two results of the section, namely 20 21. The former will allow us to obtain a Poincaré inequality in the total space of a Riemannian submersion, provided the base space satisfies a Poincaré inequality, while the latter will allows us to do the opposite. Recall that, as we have mentioned previously, these Poincaré inequalities are always considered with respect to the Gibbs measure associated with a function which is constant in the fibers of the Riemannian submersion.

Theorem 20. Let \((M, g)\) be a complete manifold and let \(\pi: (M, g) \to (M/G, h)\) be a principal \(G\)-bundle that is also a Riemannian submersion. Assume that the fibers are all isometric to a compact Lie group \((G, \hat{g})\). Further assume that \((G, \hat{g})\) has non-negative Ricci curvature. Let \(\mathrm{diam}(G)\) be the diameter of \(G\). Consider a \(C^2\) function \(F: M \to \mathbb{R}\) which is constant on the fibers of \(\pi\). Let \(\tilde{F}: M/G \to \mathbb{R}\) be the (unique) function on \(M/G\) such that \(F = \tilde{F} \circ \pi\). Let \[d\tilde{\nu} = \frac{1}{\tilde{Z}} e^{-\beta \tilde{F}(x)}\; \mathrm{dVol}_{h},\quad d\nu = \frac{1}{Z}e^{-\beta F(x)}\mathrm{dVol}_g,\] be the Gibbs probability measures on \(M/G\) and \(M\), respectively. If \((M/G, \tilde{\nu}, \tilde{\Gamma})\) satisfies a \(\mathrm{PI}(\tilde{\kappa})\), then \((M, \nu, \Gamma)\) also satisfies a \(\mathrm{PI}(\kappa)\) with \[\frac{1}{\kappa} = \max\Big\{\frac{1}{\tilde{\kappa}}, \frac{\mathrm{diam}(G)^2 \beta }{\pi^2}\Big\},\] where \(\Gamma(\cdot) = \frac{1}{\beta}|\mathrm{grad}_g\, \cdot|^2_{g}\), and \(\tilde{\Gamma}(\cdot) = \frac{1}{\beta} |\mathrm{grad}_{h} \,\cdot|^2_{h}\).

Theorem 21. Under the conditions of 20, where \((G, \hat{g})\) may not have non-negative Ricci curvature, if \((M, \nu, \Gamma)\) satisfies a \(\mathrm{PI}(\kappa)\) then \((M/G, \tilde{\nu}, \tilde{\Gamma})\) also satisfies a \(\mathrm{PI}(\kappa)\).

For the proofs of 20 21, we will use two auxiliary lemmas that relate the horizontal and the vertical components of the gradient of a function in the total space of a Riemannian submersion to the gradient with respect to the metrics of the base space and the fibers. These components are denoted by \(\cdot^\mathcal{H}\) and \(\cdot^\mathcal{V}\), respectively (see 63).

Lemma 9. Let \(\pi: (M, g) \to (M/G, h)\) be a principal \(G\)-bundle which is also a Riemannian submersion, where \(G\) is a compact Lie group. Assume that the fibers are all isometric to \((G, \hat{g})\). Then for every \(x \in M/G\) and for every \(f \in C^2(M)\), the following holds: \[\left|\frac{1}{\mathrm{Vol}_{\hat{g}}(G)}\mathrm{grad}_{h}\int_{\pi^{-1}(x)} f(y)\; \mathrm{dVol}_{\hat{g}}(y)\right|_{h}^2 \leq \frac{1}{\mathrm{Vol}_{\hat{g}}(G)} \int_{\pi^{-1}(x)} \left|\left(\mathrm{grad}_{g} \, f(y)\right)^{\mathcal{H}}\right|_{g}^2 \mathrm{dVol}_{\hat{g}}(y).\]

Proof. First, recall that as \((G, \hat{g})\) has finite volume and \(f \in C^2(M)\), it holds that \(|f|\) and \(|(\mathrm{grad}_g\, f)^{\mathcal{H}}|^2_g\) are both integrable in \(\pi^{-1}(x)\) for every \(x \in M/G\), so the above expression is well-defined.

Let us now fix \(x \in M/G\). We choose coordinates \((z^1, \dotsc, z^k)\) in \(M/G\) centered at \(x = (0, \dotsc, 0)\) and such that \(\{\frac{\partial}{\partial z^1}, \dotsc, \frac{\partial}{\partial z^k}\}\) is an orthonormal basis for \(T_x M/G\), i.e. \(\left.h_{ij}\right|_x = \delta_{ij}\).

Furthermore, we can extend these coordinates to \(\pi^{-1}(x) \setminus \mathcal{X} (\simeq G \setminus \mathcal{X})\), where \(\mathcal{X}\) is a measure-zero set. By a slight abuse of notation, we denote such coordinates as \[(z^1, \dotsc, z^k, y^1, \dotsc, y^{n-k}).\] This way, every point \(y \in \pi^{-1}(x) \setminus \mathcal{X}\) can be written as \((0, \dotsc, 0, y^1, \dotsc, y^{n-k})\). Furthermore, in these coordinates, the local frame \[\left\{\frac{\partial}{\partial z^1}, \dotsc, \frac{\partial}{\partial z^k}, \frac{\partial}{\partial y^1}, \dotsc, \frac{\partial}{\partial y^{n-k}}\right\}\] is such that \(\left\{\frac{\partial}{\partial z^1}, \dotsc, \frac{\partial}{\partial z^k}\right\}\) are orthonormal with respect to \(g\) and horizontal with respect to \(\pi\) (cf. 26), \(\left\{\frac{\partial}{\partial y^1}, \dotsc, \frac{\partial}{\partial y^{n-k}}\right\}\) are vertical and \(g(\frac{\partial}{\partial z^i}, \frac{\partial}{\partial y^j}) = 0\) for every \(i \in \{1, \dotsc, k\}\), \(j \in \{1, \dotsc, n-k\}\).

From our choice of coordinates, we know that on \(\pi^{-1}(x) \setminus \mathcal{X}\), \(g\) is of the form \[\label{matrixmetric} \left(\begin{array}{ c | c } \mathbb{1} & 0 \\ \hline 0 & A \end{array}\right),\tag{25}\] where \(\mathbb{1}\) is the identity matrix of size \(k \times k\) and \(A\) is an \((n-k) \times (n-k)\) matrix.

In such coordinates, since the metric \(h\) at \(x\) is simply the Euclidean metric, \[\left.\mathrm{grad}_{h}( \cdot)\right|_x = h^{ij}(x)\frac{\partial }{\partial z^i}(\cdot)\frac{\partial }{\partial z^j} = \delta^{ij}\frac{\partial }{\partial z^i}(\cdot)\frac{\partial }{\partial z^j},\] and so \[\begin{gather} \left|\frac{1}{\mathrm{Vol}_{\hat{g}}(G)} \mathrm{grad}_{h} \int_{\pi^{-1}(x)} f(y)\; \mathrm{dVol}_{\hat{g}}(y)\right|_{h}^2 = \frac{1}{\mathrm{Vol}_{\hat{g}}(G)^2} \sum_{i = 1}^k \left(\frac{\partial}{\partial z^i} \int_{\pi^{-1}(x)} f(y)\; \mathrm{dVol}_{\hat{g}}(y)\right)^2 \\= \frac{1}{\mathrm{Vol}_{\hat{g}}(G)^2} \sum_{i = 1}^k \left(\frac{\partial}{\partial z^i} \int_{G \setminus \mathcal{X}} f(0, \dotsc, 0, y)\; \mathrm{dVol}_{\hat{g}}(y)\right)^2 \\= \frac{1}{\mathrm{Vol}_{\hat{g}}(G)^2} \sum_{i = 1}^k \left(\int_{G \setminus \mathcal{X}} \frac{\partial}{\partial z^i} f(0, \dotsc, 0, y)\; \mathrm{dVol}_{\hat{g}}(y)\right)^2 \\\leq \frac{1}{\mathrm{Vol}_{\hat{g}}(G)} \sum_{i = 1}^k \int_{G \setminus \mathcal{X}} \left(\frac{\partial}{\partial z^i} f(0, \dotsc, 0, y)\right)^2 \mathrm{dVol}_{\hat{g}}(y), \end{gather}\] where the last inequality is Cauchy-Schwarz’s inequality: \[\left(\int_G g(z)\; \mathrm{dVol}_{\hat{g}}(z)\right)^2 \leq \int_G g(z)^2\; \mathrm{dVol}_{\hat{g}}(z) \int_{G} \mathrm{dVol}_{\hat{g}}(z).\]

On the other hand, let us rewrite \[\label{intofgrad} \frac{1}{\mathrm{Vol}_{\hat{g}}(G)} \int_{\pi^{-1}(x)} |(\mathrm{grad}_g\, f(y))^{\mathcal{H}}|^2_g\; \mathrm{dVol}_{\hat{g}}(y)\tag{26}\] using the coordinates from the beginning of this proof, for some point in \(\pi^{-1}(x) \setminus \mathcal{X}\), \[(\mathrm{grad}_g\, f)^{\mathcal{H}} = g^{ij}\frac{\partial f}{\partial x^i}\frac{\partial }{\partial z^j},\] where \(x^i = z^i\) for \(i \in \{1, \dotsc, k\}\) and \(x^i = y^{i-k}\) for \(i \in \{k+1, \dotsc, n\}\). Moreover, since \(g\) in \(\pi^{-1}(x) \setminus \mathcal{X}\) is of the form shown in 25 we can conclude that \[(\mathrm{grad}_g\, f)^{\mathcal{H}} = \delta^{ij}\frac{\partial f}{\partial z^i}\frac{\partial }{\partial z^j}.\] Therefore, its norm with respect to \(g\) is simply \[|(\mathrm{grad}_g\, f(y))^{\mathcal{H}}|^2_g = g\left(\delta^{ij}\frac{\partial f}{\partial z^i}\frac{\partial }{\partial z^j}, \delta^{ij}\frac{\partial f}{\partial z^i}\frac{\partial }{\partial z^j}\right) = \delta_{ij}\frac{\partial f}{\partial z^i}\frac{\partial f}{\partial z^j}.\] This allows us to write 26 as \[\frac{1}{\mathrm{Vol}_{\hat{g}}(G)} \int_{G \setminus \mathcal{X}} \sum_{i = 1}^k \left(\frac{\partial}{\partial z^i}f(0,\dotsc,0, y)\right)^2 \mathrm{dVol}_{\hat{g}}(y),\] finishing the proof. ◻

Lemma 10. Let \(\pi: (M, g) \to (M/G, h)\) be a principal \(G\)-bundle, for some compact Lie group \(G\), that is also a Riemannian submersion. Assume that the fibers are all isometric to \((G, \hat{g})\). Then for every \(x \in M/G\) and for every \(f \in C^1(M)\), the following holds: \[\left|\mathrm{grad}_{\hat{g}}\, f(y)\right|_{\hat{g}}^2 = \left|\left(\mathrm{grad}_{g} \,f(y)\right)^{\mathcal{V}}\right|_{g}^2\] for almost every \(y \in \pi^{-1}(x)\).

Proof. Fix \(x \in M/G\). Working in the same coordinates as in the proof of 9 we know that on \(\pi^{-1}(x) \setminus \mathcal{X}\), \(g\) is of the form \[\left(\begin{array}{ c | c } \mathbb{1} & 0 \\ \hline 0 & A \end{array}\right).\] Using the same argument as before, we obtain \[\left(\mathrm{grad}_{g}\, f\right)^{\mathcal{V}} = g^{ij} \frac{\partial f}{\partial y^i}\frac{\partial }{\partial y^j},\] finishing the proof. ◻

Finally, we prove the following auxiliary technical lemma before turning to the main results of this section; it is stated separately for the sake of clarity.

Lemma 11. Under the conditions of 20, let \(f \in C^2(M) \cap H^1(M, \nu)\). Then \[\hat{F}(x) := \int_{\pi^{-1}(x)} f(y)\,\mathrm{dVol}_{\hat{g}}(y)\] is a function in \(C^2(M/G) \cap H^1(M/G, \tilde{\nu})\),

Proof. The fact that \(\hat{F}\) belongs to \(C^2(M/G)\) follows from the regularity of \(f\) by 17. Now, using Cauchy-Schwarz’s inequality \[\hat{F}(x)^2 = \left(\int_{\pi^{-1}(x)} f(y)\, \mathrm{dVol}_{\hat{g}}(y)\right)^2 \leq \mathrm{Vol}_{\hat{g}}(G)\int_{\pi^{-1}(x)} f^2(y)\, \mathrm{dVol}_{\hat{g}}(y),\] and so by 18 \[\begin{align} \int_{M/G} \hat{F}(x)^2\, d\tilde{\nu}(x) &\leq \mathrm{Vol}_{\hat{g}}(G) \int_{M/G} \left(\int_{\pi^{-1}(x)} f^2(y)\, \mathrm{dVol}_{\hat{g}}(y)\right) d\tilde{\nu} (x) \\&= \frac{Z}{\tilde{Z}}\mathrm{Vol}_{\hat{g}}(G) \int_M f^2(x) \,d\nu(x) < \infty, \end{align}\] as \(f \in H^1(M, \nu)\), which implies that \(\hat{F} \in L^2(M/G, \tilde{\nu})\). Moreover, from 9, we know that for every \(x \in M/G\) \[|\mathrm{grad}_{h}\, \hat{F}(x)|^2_{h} \leq \mathrm{Vol}_{\hat{g}}(G) \int_{\pi^{-1}(x)} \left|\left(\mathrm{grad}_{g}\, f(y)\right)^{\mathcal{H}}\right|_{g}^2 \,\mathrm{dVol}_{\hat{g}}(y).\] Using again 18 we can conclude that \[\begin{align} \int_{M/G} |\mathrm{grad}_{h}\, \hat{F}(x)|^2_{h}\, d\tilde{\nu}(x) &\leq \mathrm{Vol}_{\hat{g}}(G) \int_{M/G} \left( \int_{\pi^{-1}(x)} \left|\left(\mathrm{grad}_{g} \,f(y)\right)^{\mathcal{H}}\right|_{g}^2 \,\mathrm{dVol}_{\hat{g}}(y)\right) d\tilde{\nu}(x) \\&= \mathrm{Vol}_{\hat{g}}(G) \frac{Z}{\tilde{Z}} \int_M \left|\left(\mathrm{grad}_{g} \,f(x)\right)^{\mathcal{H}}\right|_{g}^2\, d\nu(x) \\&\leq \mathrm{Vol}_{\hat{g}}(G) \frac{Z}{\tilde{Z}} \int_M \left|\mathrm{grad}_{g} \,f(x)\right|_{g}^2 \,d\nu(x) < \infty, \end{align}\] as \(f \in H^1(M, \nu)\), which allows us to conclude that \(\hat{F} \in H^1(M/G, \tilde{\nu})\). ◻

With the above lemmas in mind, we will now proceed to the proof of the lifting result for Poincaré inequalities.

Proof of 20. First, recall that \((M/G, \tilde{\nu}, \tilde{\Gamma})\) satisfies a \(\mathrm{PI}(\tilde{\kappa})\) iff for every \(f \in C^2(M/G) \cap H^1(M/G, \tilde{\nu})\) \[\int_{M/G} f^2\; d\tilde{\nu} - \left(\int_{M/G} f \;d\tilde{\nu} \right)^2 \leq \frac{1}{\tilde{\kappa}\beta} \int_{M/G} |\mathrm{grad}_{h}\, f|_{h}^2\, d\tilde{\nu}.\] Let \(f \in C^2(M) \cap H^1(M, \nu)\). Without loss of generality, we can assume that \(\int_M f d\nu = 0\). Using 18 we can write \[\label{eq:initialeqPoincare} \int_M f^2\; d\nu = \frac{\tilde{Z}}{Z} \int_{M/G} \left(\int_{\pi^{-1}(x)} f^2(y) \,\mathrm{dVol}_{\hat{g}}(y)\right) d\tilde{\nu}(x).\tag{27}\] As \(f \in C^2(\pi^{-1}(x))\) for every \(x \in M/G\), we can now use the fact that the fibers satisfy a Poincaré inequality; we apply 19 to 27 and obtain \[\begin{align} \label{eq:firstauxineqpoincare} \begin{aligned} \int_M f^2\; d\nu &\leq \frac{\tilde{Z}}{Z}\left\{\frac{1}{\mathrm{Vol}_{\hat{g}}(G)}\int_{M/G} \left(\int_{\pi^{-1}(x)} f(y)\, \mathrm{dVol}_{\hat{g}}(y) \right)^2 d\tilde{\nu}(x)\right. \\&\quad\left. + C \int_{M/G} \left(\int_{\pi^{-1}(x)} \left|\mathrm{grad}_{\hat{g}} \,f(y)\right|_{\hat{g}}^2 \,\mathrm{dVol}_{\hat{g}}(y)\right) d\tilde{\nu}(x) \right\}, \end{aligned} \end{align}\tag{28}\] where \(C := \frac{\mathrm{diam}(G)^2}{\pi^2}\).

Using 10, we know that \[\left|\mathrm{grad}_{\hat{g}}\, f\right|_{\hat{g}}^2 = \left|\left(\mathrm{grad}_{g} \,f\right)^{\mathcal{V}}\right|_{g}^2,\] and so, using the fact that \(f \in H^1(M, \nu)\), we can apply 18 to the second term on the right-hand side of 28 to conclude that \[\int_M f^2\; d\nu = \frac{\tilde{Z}}{Z\mathrm{Vol}_{\hat{g}}(G)}\int_{M/G} \left(\int_{\pi^{-1}(x)} f(y)\, \mathrm{dVol}_{\hat{g}}(y) \right)^2 d\tilde{\nu}(x) + C \int_{M} \left|\left(\mathrm{grad}_{g}\, f\right)^{\mathcal{V}}\right|_{g}^2 d\nu.\]

Now, it follows from 11 that \[x \mapsto \int_{\pi^{-1}(x)} f(y) \,\mathrm{dVol}_{\hat{g}}(y)\] is a function in \(C^2(M/G) \cap H^1(M/G, \tilde{\nu})\). Since we are assuming that \((M/G, \tilde{\nu}, \tilde{\Gamma})\) satisfies a Poincaré inequality and that \(\int_M f d\nu = 0\) we can further bound \[\begin{align} \int_M f^2\; d\nu &\leq \frac{\tilde{Z}}{Z\mathrm{Vol}_{\hat{g}}(G)}\Biggl\{\left( \int_{M/G} \left(\int_{\pi^{-1}(x)} f(y)\, \mathrm{dVol}_{\hat{g}}(y)\right)\; d\tilde{\nu}(x)\right)^2 \\&\quad + \frac{1}{\tilde{\kappa}\beta} \int_{M/G} \left|\mathrm{grad}_{h} \int_{\pi^{-1}(x)} f(y) \,\mathrm{dVol}_{\hat{g}}(y) \right|^2_{h} d\tilde{\nu}(x)\Biggr\} \\&\quad + C \int_{M} \left|\left(\mathrm{grad}_{g} \,f\right)^{\mathcal{V}}\right|_{g}^2 d\nu \\&= \frac{\tilde{Z}}{Z\mathrm{Vol}_{\hat{g}}(G)}\Biggl\{\frac{Z^2}{\tilde{Z}^2}\left(\int_M f\, d\nu\right)^2 + \frac{1}{\tilde{\kappa}\beta}\int_{M/G} \left|\mathrm{grad}_{h} \int_{\pi^{-1}(x)} f(y) \,\mathrm{dVol}_{\hat{g}}(y) \right|^2_{h} d\tilde{\nu}(x)\Biggr\} \\&\quad + C \int_{M} \left|\left(\mathrm{grad}_{g} \,f\right)^{\mathcal{V}}\right|_{g}^2 d\nu \\&= \frac{\tilde{Z}}{Z}\frac{1}{\mathrm{Vol}_{\hat{g}}(G)\tilde{\kappa}\beta}\int_{M/G} \left|\mathrm{grad}_{h} \int_{\pi^{-1}(x)} f(y) \, \mathrm{dVol}_{\hat{g}}(y) \right|^2_{h} d\tilde{\nu}(x) \\&\quad + C \int_{M} \left|\left(\mathrm{grad}_{g} \, f\right)^{\mathcal{V}}\right|_{g}^2 d\nu \\&= \frac{1}{\mathrm{Vol}_{\hat{g}}(G)^2 \tilde{\kappa}\beta}\int_{M/G} \left|\mathrm{grad}_{h} \int_{\pi^{-1}(x)} f(y)\, \mathrm{dVol}_{\hat{g}}(y) \right|^2_{h} d\tilde{\nu}(x) \\&\quad + C \int_{M} \left|\left(\mathrm{grad}_{g} \, f\right)^{\mathcal{V}}\right|_{g}^2 d\nu, \end{align}\] where the last equality follows from 24 . Lastly, we apply 9 and 18 to conclude that \[\begin{align} \int_M f^2\; d\nu &\leq \frac{1}{\mathrm{Vol}_{\hat{g}}(G)\tilde{\kappa}\beta }\int_{M/G} \left(\int_{\pi^{-1}(x)} |\left(\mathrm{grad}_{g} \,f(y)\right)^{\mathcal{H}}|_{g}^2 \,\mathrm{dVol}_{\hat{g}}(y)\right) d\tilde{\nu}(x) \\&\quad + C \int_{M} \left|\left(\mathrm{grad}_{g}\, f\right)^{\mathcal{V}}\right|_{g}^2 d\nu \\&= \frac{\tilde{Z}}{Z}\frac{1}{\tilde{\kappa}\beta}\int_{M/G} \left(\int_{\pi^{-1}(x)} |\left(\mathrm{grad}_{g}\, f(y)\right)^{\mathcal{H}}|_{g}^2 \,\mathrm{dVol}_{\hat{g}}(y)\right) d\tilde{\nu}(x) \\&\quad+ C \int_{M} \left|\left(\mathrm{grad}_{g} \,f\right)^{\mathcal{V}}\right|_{g}^2 d\nu \\&= \frac{1}{\tilde{\kappa}\beta}\int_{M} \left|(\mathrm{grad}_{g}\, f)^{\mathcal{H}} \right|_{g}^2 d\nu + C \int_{M} \left|\left(\mathrm{grad}_{g}\, f\right)^{\mathcal{V}}\right|_{g}^2 d\nu. \end{align}\]

Therefore, if \((M/G, \tilde{\nu}, \tilde{\Gamma})\) satisfies a \(\mathrm{PI}(\tilde{\kappa})\), \((M, \nu, \Gamma)\) satisfies a \(\mathrm{PI}(\kappa)\) verifying \[\frac{1}{\kappa} = \max \Big\{\frac{1}{\tilde{\kappa}}, \beta C\Big\}= \max\Big\{\frac{1}{\tilde{\kappa}}, \frac{\mathrm{diam}(G)^2 \beta }{\pi^2}\Big\},\] where \(\mathrm{diam}(G)\) is the diameter of the fibers with the restricted metric \((G, \hat{g})\). ◻

Lastly, let us prove that Poincaré inequalities can be lowered from the total space to the base space of a Riemannian submersion. Unlike for the previous result, we now do not need to use the fact that the fibers satisfy a Poincaré inequality, and the Poincaré constant obtained in the base space is the same as that of the total space.

Proof of 21. Let \(\tilde{f} \in C^2(M/G)\cap H^1(M/G, \tilde{\nu})\), and let \(f: M \to \mathbb{R}\) be the unique function such that \(\tilde{f} \circ \pi = f\). In particular, \(f\) is constant in the fibers of \(\pi\) and belongs to \(C^2(M)\). Using 3 and the fact that \(\tilde{f} \in H^1(M/G, \tilde{\nu})\) we know that \[x \mapsto \int_{\pi^{-1}(x)} f^2(y) e^{-\beta F(y)}\,\mathrm{dVol}_{\hat{g}}(y) = \tilde{f}^2(x) e^{-\beta \tilde{F}(x)}\, \mathrm{Vol(G)}\] is integrable on \(M/G\). Thus \(f^2 e^{-\beta F}\) is integrable on \(M\)—i.e. \(f \in L^2(M, \nu)\). Substituting \(f^2\) by \(|\mathrm{grad}_g\, f|^2_g\) in the above expression we can similarly conclude that \(|\mathrm{grad}_g\, f|^2_g\, e^{-\beta F}\) is integrable on \(M\), which implies that \(f \in H^1(M, \nu)\).

Without loss of generality, assume that \(\int_{M/G} \tilde{f}\; d\tilde{\nu} = 0\). Thus, by 18 we know that \[\begin{align} \int_M f(x)\, d\nu(x) &= \frac{\tilde{Z}}{Z} \int_{M/G} \left(\int_{\pi^{-1}(x)} f(y)\, \mathrm{dVol}_{\hat{g}}(y)\right) d\tilde{\nu}(x) \\&= \frac{\tilde{Z}}{Z} \int_{M/G} \left(\tilde{f}(x) \int_{\pi^{-1}(x)} \mathrm{dVol}_{\hat{g}}(y)\right) d\tilde{\nu}(x) \\&= \frac{\tilde{Z}}{Z} \mathrm{Vol}(G) \int_{M/G} \tilde{f}(x)\, d\tilde{\nu}(x) = 0. \end{align}\]

Furthermore, using again 18, we know that \[\begin{align} \int_M f^2(x) \; d\nu(x) = \frac{\tilde{Z}}{Z} \mathrm{Vol}(G) \int_{M/G} \tilde{f}^2(x) \;d\tilde{\nu}(x). \end{align}\]

Thus, since \((M, \nu, \Gamma)\) satisfies a \(\mathrm{PI}(\kappa)\) and \(f \in C^2(M) \cap H^1(M, \nu)\), we know that \[\begin{align} \int_{M/G} \tilde{f}^2(x)\; d\tilde{\nu}(x) &= \frac{Z}{\tilde{Z} \mathrm{Vol}(G)} \int_M f^2(x)\; d\nu(x) \\&\leq \frac{Z}{\tilde{Z} \mathrm{Vol}(G)} \frac{1}{\kappa \beta} \int_M |\mathrm{grad}_g \,f(x)|^2_g\; d\nu(x) \\&= \frac{1}{\mathrm{Vol}(G)}\frac{1}{\kappa \beta} \int_{M/G} \left(\int_{\pi^{-1}(x)} |\mathrm{grad}_g\, f(y)|^2_g \,\mathrm{dVol}_{\hat{g}}(y)\right)d\tilde{\nu}(x). \end{align}\]

Moreover, by 74 we can conclude that for every \(x \in M\), it holds that \[|\mathrm{grad}_g\, f(x)|^2_g = |\mathrm{grad}_{h}\, \tilde{f}(\pi(x))|^2_{h},\] and so \[\begin{align} \int_{M/G} \tilde{f}^2(x)\; d\tilde{\nu}(x) &\leq \frac{1}{\mathrm{Vol}(G)}\frac{1}{\kappa \beta} \int_{M/G} \left(\int_{\pi^{-1}(x)} |\mathrm{grad}_{h}\, \tilde{f}(x)|^2_{h}\, \mathrm{dVol}_{\hat{g}}(y)\right)d\tilde{\nu}(x) \\&= \frac{1}{\kappa \beta} \int_{M/G} |\mathrm{grad}_{h}\, \tilde{f}(x)|^2_{h}\; d\tilde{\nu}(x), \end{align}\] finishing the proof. ◻

Although the lowering result for the Poincaré inequality is not required for the proof of our main result or in the examples studied, we include it here as it may be of independent interest in this and other related settings.

4 From a Poincaré to a log-Sobolev inequality and rapid mixing↩︎

In this section, we show how to derive a log-Sobolev inequality from a Poincaré inequality together with the curvature-dimension condition. The key tool is a weaker version of the log-Sobolev inequality, known as a defective log-Sobolev inequality. The use of such weakened inequalities (including the related weak log-Sobolev inequalities) is common practice in this type of argument (cf. [45], [46]).

Definition 11 (Logarithmic Sobolev inequalities). We say that the Markov triple \((M, \nu, \Gamma)\) satisfies a defective logarithmic Sobolev inequality with constants \(\alpha\) and \(A > 0\), LSI\((\alpha, A)\) if, for every probability measure \(\mu\) such that12 \(\mu \ll \nu\) with \(f := \frac{d\mu}{d\nu} \in C^1(M)\), \[\label{defectiveLSI} H(\mu | \nu) \leq \frac{1}{2\alpha} I(\mu | \nu) + A,\qquad{(7)}\] where \(H(\mu | \nu) := \int_M \log \frac{d\mu}{d\nu} d\mu\) is known as the relative entropy (or Kullback-Leibler divergence) of the probability measure \(\mu\) with respect to \(\nu\), and \(I(\mu | \nu) := \int_M \frac{\Gamma(f)}{f} d\nu\) is known as the Fisher information of \(\mu\) with respect to \(\nu\).

We say the Markov triple \((M, \nu, \Gamma)\) satisfies a tight logarithmic Sobolev inequality—or simply log-Sobolev inequality—with constant \(\alpha > 0\), LSI(\(\alpha\)), if ?? is satisfied with \(A = 0\), i.e. \[H(\mu | \nu) \leq \frac{1}{2\alpha}I(\mu | \nu).\]

The log-Sobolev inequality allows us to study the rate at which the distribution of the Langevin diffusion \(X_t\), \(\rho_t\), tends to the Gibbs distribution \(\nu\). In order to have a notion of distance between two probability distributions, we can make use of the so-called total variation distance.

Definition 12 (Total variation distance). Let \(\mu\) and \(\nu\) be two probability measures on a measure space \((E, \mathcal{F})\). Then we define their total variation distance as \[\norm{\mu - \nu}_{\text{TV}} := \sup\{|\mu(A) - \nu(A)| : A \in \mathcal{F}\}.\]

We are interested in obtaining an upper bound for the total variation distance between \(\rho_t\) and \(\nu\), which decreases fast with \(t\). In fact, one of the implications of having a log-Sobolev inequality is that the total variation distance between the distribution of \(X_t\), \(\rho_t\) and the Gibbs distribution \(\nu\) decreases exponentially fast with respect to \(t\).

Proposition 22. If the Markov triple \((M, \nu, \Gamma)\) satisfies an \(\mathrm{LSI}(\alpha)\), then \[\norm{\rho_t - \nu}^2_{\mathrm{TV}} \leq \frac{1}{2} e^{-2\alpha t} H(\rho_0 | \nu),\quad \forall t \geq 0.\]

Proof. Indeed, by [12] we know that \[\label{eq:boundKL} H(\rho_t | \nu) \leq e^{-2\alpha t} H(\rho_0 | \nu), \quad \forall t \geq 0.\tag{29}\] Using the Pinsker-Csizsár-Kullback inequality, we get \[\norm{\rho_t - \nu}^2_{\mathrm{TV}} \leq \frac{1}{2} H(\rho_t | \nu) \leq \frac{1}{2} e^{-2\alpha t} H(\rho_0 | \nu).\] ◻

A log-Sobolev inequality shows that the total variation distance between \(\rho_t\) and \(\nu\) decreases exponentially fast with time. In many situations, it is also crucial to obtain bounds on \(H(\rho_0 |\nu)\).

Proposition 23. Let \(\mu\) be the uniform distribution on a compact Riemannian manifold \((M, g)\), let \(F : M \to \mathbb{R}\) be a smooth function satisfying \(\min_{x \in M} F(x) \geq 0\), and let \(\nu\) be the Gibbs distribution \(\nu = \frac{1}{Z}e^{-\beta F}\). Then \[H(\mu | \nu) \leq \beta \max_{x \in M} F(x).\]

Proof. The fact that \(\mu \ll \nu\) follows immediately. Now, since \(d\mu = \frac{1}{\mathrm{Vol}(M)} \mathrm{dVol}_g(x)\) and \(d\nu = \frac{1}{Z}e^{-\beta F(x)} \mathrm{dVol}_g(x)\), \[\frac{d\mu}{d\nu} = \frac{Z}{\mathrm{Vol}(M)} e^{\beta F(x)},\] which implies that \[\log \frac{d\mu}{d\nu} = \log (\frac{Z}{\mathrm{Vol}(M)}) + \beta F(x).\] This expression allows us to write the relative entropy of \(\mu\) with respect to \(\nu\) as \[H(\mu | \nu) = \frac{1}{\mathrm{Vol}(M)} \int_M \log \frac{d\mu}{d\nu} \mathrm{dVol}_g(x) = \log (\frac{Z}{\mathrm{Vol}(M)}) + \frac{\beta}{\mathrm{Vol}(M)} \int_M F(x) \mathrm{dVol}_g(x).\] Now, note that since \(F \geq 0\), \[Z = \int_M e^{-\beta F(x)} \mathrm{dVol}_g(x) \leq \int_M \mathrm{dVol}_g(x) = \mathrm{Vol}_g(M),\] and so we can upper bound the relative entropy as \[H(\mu | \nu) \leq \log(1) + \frac{\beta}{\mathrm{Vol}(M)} \int_M \max_{x \in M} F(x) \mathrm{dVol}_g(x) \leq \beta \max_{x \in M} F(x).\] ◻

In particular, whenever the initial distribution of the Langevin diffusion \(X_t\) is uniform on \(M\), \(H(\rho_0| \nu)\) can be upper bounded by the product of \(\beta\) and the maximum of \(F\).

The Poincaré inequality yields an analogue of 29 with the relative entropy replaced by the \(\chi^2\) divergence (cf. [12]), which in turn yields the total variation distance bound \(\norm{\rho_t - \nu}^2_{\mathrm{TV}} \leq \frac{1}{2} e^{-2\alpha t} \chi^2(\rho_0 | \nu)\) (cf. [47]). However, the relative entropy \(H\) is upper bounded by the logarithm of \(\chi^2\) ([47]). For this reason, while both the Poincaré and log-Sobolev inequalities provide exponentially decaying bounds on the total variation distance between \(\rho_t\) and \(\nu\), the prefactor obtained via the latter is exponentially smaller.

4.1 From Poincaré to log-Sobolev↩︎

Theorem 24 (Poincaré inequality and curvature-dimension condition imply log-Sobolev inequality). Consider the Markov triple \((M, \nu, \Gamma)\) and let \(\beta \geq 1\). Suppose that the following two conditions hold:

  1. \((M, \nu, \Gamma)\) satisfies a curvature-dimension condition with constant \(-\kappa_1 < 0\), denoted \(\mathrm{CD}(-\kappa_1)\), i.e. \[\label{CD3} \nabla^2 F + \frac{1}{\beta} \mathrm{Ric}_g \geq -\kappa_1 g.\qquad{(8)}\]

  2. \((M, \nu, \Gamma)\) satisfies a \(\mathrm{PI}(\kappa_2)\) with \(\kappa_2 > 0\).

Then \((M, \nu, \Gamma)\) satisfies an \(\mathrm{LSI}(\alpha)\) with \[\frac{1}{\alpha} = \frac{4\beta \kappa_1 \mathrm{diam}(M)^2}{\kappa_2},\] where \(\mathrm{diam}(M)\) is the diameter of \(M\).

Remark 25. Note that we can assume that \(\kappa_1 > 1\) since \(\mathrm{CD}(-\kappa_1)\) with \(0 < \kappa_1 \leq 1\) implies \(\mathrm{CD}(-\kappa')\) for any \(\kappa' > 1\). Similarly, we can assume that \(1 \geq \kappa_2 > 0\), since \((M, \nu, \Gamma)\) satisfying a \(\mathrm{PI}(\kappa_2)\) with \(\kappa_2 > 1\) implies a \(\mathrm{PI}(1)\).

Remark 26. 24 is a generalization of [18]. Note that we are allowing for the curvature-dimension condition to hold with respect to the negative constant \(-\kappa_1\), as opposed to when we obtained a Poincaré inequality from the curvature-dimension condition, in which case we needed the constant to be positive (cf. 5 97).

To prove 24, we will proceed as in [18]. We will first obtain an upper bound on the relative entropy of \(\rho_t\)—the distribution of the Langevin diffusion \(X_t\)—with respect to \(\nu\), under the assumption that the curvature-dimension condition holds. This upper bound will be written in terms of the Fisher information and the Wasserstein distance between \(\rho_t\) and \(\nu\). We will next get a defective log-Sobolev inequality from it. Finally, this inequality will be tightened using standard techniques.

Definition 13 (Wasserstein distance). Let \((\mathcal{X}, d)\) be a separable and complete metric space, and let \(p \in [1, \infty)\). Let \(\mu\) and \(\eta\) be two probability measures of \(\mathcal{X}\). We define the Wasserstein distance of order \(p\) between \(\mu\) and \(\eta\) as \[W_p(\mu, \eta) := \left(\inf_{\pi \in \Pi(\mu, \eta)} \int_{\mathcal{X}} d(x, y)^p \, d\pi(x, y)\right)^{1/p},\] where \(\Pi(\mu, \eta)\) is the set of joint probability measures on \(\mathcal{X} \times \mathcal{X}\) whose marginals are \(\mu\) and \(\eta\).

As we mentioned earlier, the Wasserstein distance allows us to upper bound the relative entropy between two probability distributions in the presence of a curvature-dimension condition.

Proposition 27 (HWI Inequality, [48]). Let \((M, g)\) be a compact Riemannian manifold equipped with a measure \(\mu = e^{-V}\), \(V \in C^2(M)\), satisfying the following curvature-dimension bound: \[\nabla^2 V + \mathrm{Ric}_g \geq \kappa g,\] for some \(\kappa \in \mathbb{R}\). Then if \(\mu \in P_2(M)\), i.e. \(\mu\) is a probability measure and \[\int_M d_g(x, x_0)^2 d\mu < \infty,\quad \forall x_0 \in M,\] for any \(\eta \in P_2(M)\), it holds that \[\label{eq:prelimineqH} H(\eta | \mu) \leq W_2(\eta, \mu)\sqrt{\mathbf{I}(\eta | \mu)} - \frac{\kappa}{2}W_2(\eta, \mu)^2,\qquad{(9)}\] where \(\mathbf{I}(\eta | \mu)\) is the unadjusted* Fisher information, i.e. \(\mathbf{I}(\eta | \mu) := \int_M \frac{|\mathrm{grad}_{g}\, f|_g^2}{f} d\mu\), with \(f = \frac{d\eta}{d\mu}\).*

The following lemma provides an upper bound on the Wasserstein distance that will be useful to derive a defective log-Sobolev inequality from ?? .

Lemma 12 ([48]). Let \(\mu, \eta\) be two probability measures on a separable and complete metric space \((\mathcal{X}, d)\). Then for every \(p \geq 1\) and every \(x_0 \in \mathcal{X}\), it holds that \[W_p(\mu, \eta) \leq 2^{\frac{1}{p'}}\left\{\int_{\mathcal{X}} d(x_0, x)^p \, d|\mu - \eta|(x)\right\}^{\frac{1}{p}},\] where \(\frac{1}{p} + \frac{1}{p'} = 1\).

Using this result we can easily upper bound the Wasserstein distance by the diameter of the manifold. This result is analogous to [48].

Remark 28. Let \((M, g)\) be a compact Riemannian manifold with diameter \(\mathrm{diam}(M)\). Then for any two probability measures \(\mu, \eta \in P_2(M)\) it holds that \[W_2(\mu, \eta)^2 \leq 4\, \mathrm{diam}(M)^2.\]

Proof. Setting \(p = 2\) in 12 we obtain \[W_2(\mu, \eta)^2 \leq 2\, \mathrm{diam}(M)^2 \int_{M} d|\mu - \eta|(x) = 2\,\mathrm{diam}(M)^2 \norm{\mu - \eta}_{\text{TV}}.\] Now \[\norm{\mu - \eta}_{\text{TV}} = \sup\{|\mu(A) - \eta(A)| : A \in \mathcal{F}\} \leq 2,\] yielding \[W_2(\mu, \eta)^2 \leq 4\, \mathrm{diam}(M)^2.\] ◻

Lastly, as we mentioned earlier, we will end the proof of 24 with a tightening argument, which will allow us to obtain a tight log-Sobolev inequality from the existence of a defective log-Sobolev inequality and a Poincaré inequality. To this end, we will use the following result, which we include without proof.

Proposition 29 (Tightening of a log-Sobolev inequality with a Poincaré inequality - [12]). Assume that \((M, \nu, \Gamma)\) satisfies an \(\mathrm{LSI}(C, D)\) and a \(\mathrm{PI}(C')\). Then the Markov triple satisfies a log-Sobolev inequality with constant \[\frac{1}{\frac{1}{C} + \frac{1}{C'}(\frac{D}{2} + 1)}.\]

With the above results in mind, we are now ready to prove 24.

Proof of 24. Let \(\nu\) be the Gibbs distribution, and let \(\mu\) be any measure on \(M\) satisfying \(\mu \ll \nu\). Consider the relative entropy between \(\mu\) and \(\nu\), \(H(\mu|\nu)\). In order to apply 27, we have to rewrite \(\nu\) as \(\nu = e^{-V}\) for a suitable \(C^2\) function \(V\). It suffices to consider \(V = \beta F + \log(Z)\). For this choice of \(V\), the curvature-dimension condition of 27 is fulfilled. Indeed, \[\begin{align} \nabla^2 V + \mathrm{Ric}_g = \nabla^2(\beta F + \log(Z)) + \mathrm{Ric}_g = \beta \nabla^2 F + \mathrm{Ric}_g \geq -\beta \kappa_1 g, \end{align}\] where the inequality follows by multiplying the curvature-dimension condition written in ?? by \(\beta\). Therefore, applying 27 we conclude that \[\label{ineqrelativeentropy} H(\mu | \nu) \leq W_2(\mu, \nu) \sqrt{\mathbf{I}(\mu | \nu)} + \frac{\beta \kappa_1}{2} W_2(\mu, \nu)^2,\tag{30}\] where \[\mathbf{I}(\mu | \nu) = \beta I(\mu | \nu),\] and \(I(\mu | \nu)\) is the Fisher information. Furthermore, we can apply Young’s inequality \[ab \leq \frac{\varepsilon}{2}a^2 + \frac{1}{2\varepsilon}b^2, \quad \forall \varepsilon > 0,\;\forall a, b \geq 0.\] to 30 with \(a = \sqrt{\mathbf{I}(\mu | \nu)}\) and \(b = W_2(\mu, \nu)\) to obtain \[\begin{align} H(\mu | \nu) &\leq \frac{\varepsilon}{2}\mathbf{I}(\mu | \nu) + \left(\frac{1}{2\varepsilon} + \frac{\beta \kappa_1}{2}\right)W_2(\mu, \nu)^2\\ &\leq \frac{\varepsilon}{2}\mathbf{I}(\mu | \nu) + \left(\frac{1}{2\varepsilon}+ \frac{\beta \kappa_1}{2}\right)4\,\mathrm{diam}(M)^2\\ &= \frac{\varepsilon\beta}{2}I(\mu | \nu) + \left(\frac{1}{2\varepsilon}+ \frac{\beta \kappa_1}{2}\right)4\, \mathrm{diam}(M)^2, \end{align}\] where the second inequality follows from 28.

The above inequality is a defective logarithmic Sobolev inequality (cf. ?? ) with constants \(\alpha = \frac{1}{\varepsilon \beta}\) and \(A = \left(\frac{1}{2\varepsilon}+ \frac{\beta \kappa_1}{2}\right)4\, \mathrm{diam}(M)^2\). Thus, we can apply the tightening result from 29 to conclude that \((M, \nu, \Gamma)\) satisfies a tight logarithmic Sobolev inequality with constant \[\frac{1}{\varepsilon \beta + \frac{1}{\kappa_2}\left(\left(\frac{1}{2\varepsilon}+ \frac{\beta \kappa_1}{2}\right)2\,\mathrm{diam}(M)^2 + 1\right)},\] i.e. \[H(\mu | \nu) \leq \frac{1}{2}\left(\varepsilon\beta + \frac{1}{\kappa_2}\left\{\left(\frac{1}{2\varepsilon}+ \frac{\beta \kappa_1}{2}\right)2\,\mathrm{diam}(M)^2 + 1\right\}\right)I(\mu | \nu).\] Since this inequality works for all \(\varepsilon > 0\) we can choose \(\varepsilon = \frac{\mathrm{diam}(M)}{\sqrt{\beta \kappa_2}}\) to obtain the least upper bound, i.e. \[\frac{1}{\tilde{\alpha}} := 2\,\mathrm{diam}(M)\sqrt{\frac{\beta}{\kappa_2}} + \frac{\beta \kappa_1 \mathrm{diam}(M)^2 + 1}{\kappa_2}.\]

Let us conclude by obtaining a more readable log-Sobolev inequality constant. Indeed as we are assuming without loss of generality that \(\beta \geq 1\), \(0 < \kappa_2 \leq 1\), \(\mathrm{diam}(M) \geq 1\) and \(\kappa_1 > 1\) we can bound \[\begin{align} 2\,\mathrm{diam}(M)\sqrt{\frac{\beta}{\kappa_2}} + \frac{\beta \kappa_1 \mathrm{diam}(M)^2 + 1}{\kappa} &\leq \frac{2\beta \,\mathrm{diam}(M) + \beta \kappa_1 \mathrm{diam}(M)^2 + 1}{\kappa_2}\\ &\leq \frac{4\beta \kappa_1 \mathrm{diam}(M)^2}{\kappa_2}, \end{align}\] finishing the proof. ◻

5 Rapid mixing for Gibbs measures in Riemannian manifolds↩︎

We are now in a position to state and prove our rapid mixing result for Langevin dynamics on Riemannian manifolds. An informal presentation thereof was given in 4. To this end, we will put together the main results of [SectionPI] [SectionLift] [SectionPItoLSI] [SectionSuboptimality]. We will first prove that \((M, \nu, \Gamma)\) and \((M/G, \tilde{\nu}, \tilde{\Gamma})\) both satisfy a log-Sobolev inequality. Next, we will see how the rapid mixing result follows.

Theorem 30 (4 - formal version). Let \((M, g)\) be a compact Riemannian manifold of dimension \(\dim(M) \geq 2\), which is furthermore a symmetric space. Let \(G\) be a compact and connected Lie group, and let \(\pi: (M, g) \to (M/G, h)\) be a principal \(G\)-bundle that is also a Riemannian submersion with totally geodesic fibers. Further assume that the fibers of \(\pi\)—which are isometric to \((G, \hat{g})\), where \(\hat{g}\) corresponds to their induced metric—satisfy 2. Assume that \(M\) and \(M/G\) satisfy 3 with constants \(\mathbf{K}\), \(R_M\) and \(R_{M/G}\).

Let \(F: M \to \mathbb{R}\) be a function satisfying and let \(\tilde{F}: M/G \to \mathbb{R}\) be the unique function such that \(F = \tilde{F} \circ \pi\). Assume that \(\tilde{F}\) satisfies . Let \(a, \beta > 0\) be such that \[\begin{align} a^2 &\geq \max\Big\{\frac{24 A_2 \dim(M/G)}{C^2_{\tilde{F}}}, \frac{544}{\lambda_*}\Big\}, \\\beta &\geq \max\Bigg\{\frac{72^2 \dim(M)^5 A_2 A^2_3 \mathbf{K}^2 a^6}{\lambda_*^2}, \frac{9a^2}{D^2}, \frac{a^2}{i(M/G)^2}, \frac{a^2}{i(M)^2}, \frac{4R_{M/G}}{\lambda_*}, \frac{a^2}{\mathit{conv}(M/G)^2}\Bigg\}, \end{align}\] where \(C_{\tilde{F}}\) is the constant from ?? , \(\mathit{conv}(M/G)\) denotes the convexity radius of \((M/G, h)\), \(i(M)\) and \(i(M/G)\) denote the injectivity radii of \((M, g)\) and \((M/G, h)\) respectively, \(D\) is the lower bound on the distance between any two critical points of \(\tilde{F}\), \(\lambda_*\) is the lower bound on the eigenvalues of \(\nabla^2 \tilde{F}\) at the critical points, and \(A_2, A_3\) are the Lipschitz constants of the gradient and the Hessian of \(\tilde{F}\), respectively.

Then the Markov triple \((M, \nu, \Gamma)\) satisfies an \(\mathrm{LSI}(\alpha_M)\), with \[\label{eq:constantalphaM} \frac{1}{\alpha_M} = 4\beta (A_2 + R_M) \mathrm{diam}(M)^2\max \Big\{\frac{184}{\lambda_*}, \frac{\mathrm{diam}(G)^2 \beta}{\pi^2}\Big\},\qquad{(10)}\] where \(\mathrm{diam}(G)\) denotes the diameter of \(G\) as a fiber of \(\pi\) and \(\mathrm{diam}(M)\) denotes the diameter of \((M, g)\). Furthermore, the Markov triple \((M/G, \tilde{\nu}, \tilde{\Gamma})\) satisfies an \(\mathrm{LSI}(\alpha_{M/G})\), with \[\label{eq:constantalphaMG} \frac{1}{\alpha_{M/G}} = \frac{736\beta (A_2 + R_{M/G}) \mathrm{diam}(M/G)^2}{\lambda_*}.\qquad{(11)}\]

Note that there are two extra assumptions needed to prove the existence of a log-Sobolev inequality for the Markov triples \((M, \nu, \Gamma)\) and \((M/G, \tilde{\nu}, \tilde{\Gamma})\), compared to those needed to prove the existence of a Poincaré inequality for \((M/G, \tilde{\nu}, \tilde{\Gamma})\). The first is that the Ricci curvature of the fibers of \(\pi\) is non-negative, in order to be able to lift the Poincaré inequality from \(M/G\) to \(M\). The second is the existence of a lower bound on the Ricci curvature of \(M\), which allows to tighten the Poincaré inequality lifted to \(M\) to a log-Sobolev inequality.

Also observe that in the above theorem we claim that the fibers of \(\pi\) are isometric. This is always the case whenever the Riemannian submersion \(\pi\) has totally geodesic fibers [49].

As we mentioned before, the proof of 30 only consists in putting together the results obtained in the previous sections. In particular, we will first use the Poincaré inequality theorem—which corresponds to 10—to conclude that \((B, \tilde{\nu}, \tilde{\Gamma})\) satisfies a Poincaré inequality. Then using the lifting result shown in 20 we will obtain a Poincaré inequality on \((M, \nu, \Gamma)\). Lastly, using the tightening arguments from 24 we will upgrade both Poincaré inequalities to obtain logarithmic Sobolev inequalities.

Proof of 30. First, recall that under the assumptions of the theorem, we can apply 10 and conclude that the Markov triple \((M/G, \tilde{\nu}, \tilde{\Gamma})\) satisfies a \(\mathrm{PI}(\kappa_{M/G})\) with \[\kappa_{M/G} = \frac{\lambda_*}{184}.\] Now, we can lift this Poincaré inequality from \(M/G\) to \(M\) by simply applying 20, and claim that \((M, \nu, \Gamma)\) satisfies a \(\mathrm{PI}(\kappa_M)\) with \[\begin{align} \frac{1}{\kappa_{M}} &= \max \Big\{\frac{1}{\kappa_{M/G}}, \frac{\mathrm{diam}(G)^2 \beta}{\pi^2}\Big\}\\ &= \max \Big\{\frac{184}{\lambda_*}, \frac{\mathrm{diam}(G)^2 \beta}{\pi^2}\Big\}. \end{align}\]

After obtaining a Poincaré inequality for \((M, \nu, \Gamma)\) and \((M/G, \tilde{\nu}, \tilde{\Gamma})\), it only remains to tighten them to logarithmic Sobolev inequalities. This can be done using 24. Indeed, note that by 4 it holds that \[\begin{align} \nabla^2_g F &\geq -A_2 g,\\ \nabla^2_h \tilde{F} &\geq -A_2 h, \end{align}\] and so, as we are assuming that \(\mathrm{Ric}_g \geq -R_M\), \(\mathrm{Ric}_h \geq -R_{M/G}\) for some constants \(R_M, R_{M/G} \geq 0\), we can conclude that \[\begin{align} \nabla_g^2 F + \frac{1}{\beta}\mathrm{Ric}_g &\geq -(A_2 + R_M),\\ \nabla_h^2 \tilde{F} + \frac{1}{\beta}\mathrm{Ric}_h &\geq -(A_2 + R_{M/G}). \end{align}\] Thus, by 24 we can conclude that \((M, \nu, \Gamma)\) satisfies an \(\mathrm{LSI}(\alpha_M)\) with \[\begin{align} \frac{1}{\alpha_M} &= \frac{4\beta (A_2 + R_M) \mathrm{diam}(M)^2}{\kappa_M}\\ &= 4\beta (A_2 + R_M) \mathrm{diam}(M)^2\max \Big\{\frac{184}{\lambda_*}, \frac{\mathrm{diam}(G)^2 \beta}{\pi^2}\Big\}. \end{align}\] and \((M/G, \tilde{\nu}, \tilde{\Gamma})\) satisfies an \(\mathrm{LSI}(\alpha_{M/G})\) with \[\begin{align} \frac{1}{\alpha_{M/G}} = \frac{4\beta (A_2 + R_{M/G}) \mathrm{diam}(M/G)^2}{\kappa_{M/G}} = \frac{736\beta (A_2 + R_{M/G}) \mathrm{diam}(M/G)^2}{\lambda_*}. \end{align}\] ◻

As we mentioned previously, the rapid mixing result shown in 4 becomes a corollary of 30.

Corollary 4. Under the assumptions of 30, the Langevin diffusion process \(X_t\) defined in 2 with initial uniform distribution on \(M\) converges exponentially fast to the Gibbs measure \(\nu\), i.e. \[\norm{\nu - \rho_t}^2_{\mathrm{TV}} \leq \beta\, e^{-2\alpha_M t} \max_{y \in M} F(y), \quad \forall t \geq 0,\] where \(\rho_t\) denotes the distribution of \(X_t\) and \(\alpha_M\) denotes the log-Sobolev inequality constant of \((M, \nu, \Gamma)\). Similarly, the Langevin diffusion process \(\tilde{X}_t\) defined analogously with initial uniform distribution on \(M/G\) also converges exponentially fast to the Gibbs measure \(\tilde{\nu}\), i.e. \[\norm{\tilde{\nu} - \tilde{\rho}_t}^2_{\mathrm{TV}} \leq \beta\, e^{-2\alpha_{M/G} t} \max_{y \in M/G} \tilde{F}(y) = \beta\, e^{-2\alpha_{M/G} t} \max_{y \in M} F(y) , \quad \forall t \geq 0,\] where \(\tilde{\rho}_t\) denotes the distribution of \(\tilde{X}_t\) and \(\alpha_{M/G}\) denotes the log-Sobolev inequality constant of \((M/G, \tilde{\nu}, \tilde{\Gamma})\).

Proof. In 30 we proved that \((M, \nu, \Gamma)\) satisfies an \(\mathrm{LSI}(\alpha_M)\) (see ?? ) and the Markov triple \((M/G, \tilde{\nu}, \tilde{\Gamma})\) satisfies an \(\mathrm{LSI}(\alpha_{M/G})\) (see ?? ). Using 22 we can obtain an upper bound on the total variation distance between the densities of \(X_t\) and \(\tilde{X}_t\) and the Gibbs measures \(\nu\) and \(\tilde{\nu}\), \[||\rho_t - \nu||_{\mathrm{TV}}^2 \leq \frac{1}{2} e^{-2\alpha_M t} H(\rho_0|\nu),\quad \mathrm{and}\quad ||\tilde{\rho}_t - \tilde{\nu}||_{\mathrm{TV}}^2 \leq \frac{1}{2} e^{-2\alpha_{M/G}\, t} H(\tilde{\rho}_0|\tilde{\nu}),\] for every \(t \geq 0\). Moreover, by 23, we can bound the relative entropy between the uniform distribution and the Gibbs distribution on \(M\) and \(M/G\) as \[H(\rho_0 | \nu) \leq \beta \max_{y \in M} F(y),\quad \mathrm{and}\quad H(\tilde{\rho}_0 | \tilde{\nu}) \leq \beta \max_{y \in M/G} \tilde{F}(y) = \beta \max_{y \in M} F(y),\] concluding the proof. ◻

6 Suboptimality of the Gibbs distribution↩︎

The aim of this section is to prove that, given a compact Riemannian manifold \((M, g)\) and a sufficiently smooth function \(F: M \to \mathbb{R}\), the Gibbs distribution associated with \(F\), \(\nu = \frac{1}{Z}e^{-\beta F}\), finds the minima of \(F\). This result was stated informally in 5.

More formally, we will show that, by choosing appropriate values for \(\beta\), which scale polynomially with the dimension of the manifold, and logarithmically with respect to the volume of the manifold and the Lipschitz constant of the gradient of \(F\), the distribution of \(\nu\) gets concentrated around the minima of the function \(F\).

Theorem 31 (5 - formal version). Let \((M, g)\) be a compact Riemannian manifold of dimension \(d \geq 2\), and let \(F : M \to \mathbb{R}\) be twice differentiable with an \(A_2\)-Lipschitz gradient. Let \[\varepsilon_{\max} := \min \Big\{\frac{i(M)^2 A_2}{8}, 1\Big\},\] where \(i(M)\) denotes the injectivity radius of \(M\). For every \(\varepsilon \in (0, \varepsilon_{\max}]\) and \(\delta \in (0, 1)\), if \[\beta \geq \frac{2}{\varepsilon}\left( \frac{1}{2} +d^2\log\frac{d^3A_2 \mathrm{Vol}(M)\sqrt{2\pi}}{\delta\varepsilon} \right),\] then the Gibbs distribution \(\nu(x) = \frac{1}{Z} e^{-\beta F(x)}\) satisfies \[\nu\left(F - \min_{y \in M} F(y) \geq \varepsilon\right) \leq \delta.\]

In the above theorem one can assume, without loss of generality, that \[\min_{x\in M} F(x) = 0.\]

This section is based on [18]. Note that the statement of the theorem—which corresponds to [17]—has remained almost identical to the original, as well as the auxiliary lemmas needed to prove it. Nevertheless, some of the proofs have been adapted to generalize the result to an arbitrary compact manifold \(M\).

Lemma 13 ([18]). Let \((M, g)\), \(\varepsilon_{\max}\) and \(F\) be as in 31. Let \(x^* \in M\) be any global minimum of \(F\). Then for every \(\varepsilon \in (0, \varepsilon_{\max}]\) the following bound holds: \[\nu(F \geq \varepsilon) \leq \frac{e^{-\beta \varepsilon} \mathrm{Vol}(M)}{\int_{\mathcal{B}(R, x^*)} e^{-\beta \frac{A_2}{2} d_g(x^*, x)^2}\; \mathrm{dVol}_g(x)},\] where \(R := \sqrt{\frac{2\varepsilon}{A_2}}\), \(\mathcal{B}(R, x^*)\) denotes the geodesic ball of radius \(R\) centered at \(x^*\), and \(d_g(\cdot, \cdot)\) is the geodesic distance of \((M, g)\).

Note that we require the geodesic ball of radius \(R\) centered at \(x^*\) to be contained in the geodesic ball with radius equal to the injectivity radius of \(M\). Under this extra assumption, the original proof of 13 remains valid. For this reason, we do not include a proof for it in this work.

For the next auxiliary lemma, which is an adapted version of [18], we will incorporate a proof, as the techniques used are different.

Lemma 14. Let \((M, g)\) be a compact Riemannian manifold of dimension \(d\). Let \[\varepsilon_{\max} = \min\Big\{\frac{i(M)^2 A_2}{8}, 1\Big\}\] and let \(\varepsilon \in (0, \varepsilon_{\max}]\), \(R = \frac{2\varepsilon}{A_2}\). Then for every \[\beta \geq \frac{1+2\log2}{2\varepsilon},\] and every \(y \in M\) it holds that \[\int_{\mathcal{B}(R, y)} e^{-\beta \frac{A_2}{2} d_g(y, x)^2} \; \mathrm{dVol}_g(x) \geq 2^{d-1} \pi^{1/2} \frac{\Gamma(\frac{d+1}{2})^{d-1}}{\Gamma(\frac{d}{2})^d d^{d-1}} \left(\frac{1}{\beta A_2}\right)^{d/2}e^{-1/2}.\]

Proof. By the Gauss lemma, since \(R \leq i(M)/2\), \[\label{eq:rewrittenintegral} \int_{\mathcal{B}(R, y)} e^{-\beta \frac{A_2}{2} d_g(y, x)^2} \; \mathrm{dVol}_g(x) = \int_0^R e^{-\beta\frac{A_2}{2}\rho^2} \left(\int_{S_{y}(\rho)} d\sigma \right)\; d\rho,\tag{31}\] where \[S_{y}(\rho) := \{\text{exp}_{y}(v)\, : \, |v|_g = \rho\}\] is the geodesic sphere of radius \(\rho\) centered at \(y\), \(\exp_{y}\) denotes the exponential map centered at \(y\), and \(d\sigma\) is the induced volume form on \(S_{y}(\rho)\). Now, since \(M\) is compact, by [50], it holds that \[\mathrm{Vol}(S_{y}(\rho)) \geq 2^{d-1} \frac{\mathrm{Vol}_{g_{\mathit{round}}}(\mathbb{S}^{d-1})^d}{\mathrm{Vol}_{g_{\mathit{round}}}(\mathbb{S}^{d})^{d-1} d^{d-1}} \rho^{d-1},\] for all \(\rho \leq i(M)/2\), where \(\mathrm{Vol}_{g_{\mathit{round}}}(\mathbb{S}^{d})\) denotes the volume of the \(d\)-dimensional round sphere of radius \(1\), which is given by (cf. [51]) \[\mathrm{Vol}_{g_{\mathit{round}}}(\mathbb{S}^{d}) = \frac{2 \pi^{(d+1)/2}}{\Gamma(\frac{d+1}{2})}.\] Thus, \[\begin{align} \mathrm{Vol}(S_{y}(\rho)) &\geq 2^{d-1} \frac{\mathrm{Vol}_{g_{\mathit{round}}}(\mathbb{S}^{d-1})^d}{\mathrm{Vol}_{g_{\mathit{round}}}(\mathbb{S}^{d})^{d-1} d^{d-1}} \rho^{d-1} \\&= 2^{d-1} \frac{\left(\frac{2 \pi^{d/2}}{\Gamma(\frac{d}{2})}\right)^d}{\left(\frac{2 \pi^{(d+1)/2}}{\Gamma(\frac{d+1}{2})}\right)^{d-1} d^{d-1}} \rho^{d-1}\\ &=2^{d} \pi^{1/2} \frac{\Gamma(\frac{d+1}{2})^{d-1}}{\Gamma(\frac{d}{2})^d d^{d-1}} \rho^{d-1} \\&= \mathcal{A}_d\, \rho^{d-1}, \end{align}\] where \[\mathcal{A}_d := 2^{d} \pi^{1/2} \frac{\Gamma(\frac{d+1}{2})^{d-1}}{\Gamma(\frac{d}{2})^d d^{d-1}}.\]

This way, we can lower-bound the right-hand side of 31 as \[\int_0^R e^{-\beta\frac{A_2}{2}\rho^2} \left(\int_{S_{y}(\rho)}\; d\sigma \right)\; d\rho \geq \mathcal{A}_d \int_0^R e^{-\beta\frac{A_2}{2}\rho^2} \rho^{d-1}\; d\rho.\] applying the change of variable \(r = \sqrt{\beta A_2}\rho\) we get \[\begin{align} \label{eq1} \nonumber \int_0^R e^{-\beta\frac{A_2}{2}\rho^2} \rho^{d-1}\; d\rho &= \int_0^{\sqrt{\beta A_2}R} e^{-\frac{r^2}{2}} \left(\frac{1}{\beta A_2}\right)^{d/2} r^{d-1}\; dr \\&= \left(\frac{1}{\beta A_2}\right)^{d/2} \int_0^{\sqrt{\beta A_2}R} e^{-\frac{r^2}{2}} r^{d-1}\; dr. \end{align}\tag{32}\] Now, since \(d \geq 2\) and \(R = \sqrt{\frac{2\varepsilon}{A_2}}\), if we choose \(\beta\) such that \[\beta \geq \frac{1}{A_2 R^2} = \frac{1}{2\varepsilon},\] it holds that \[\sqrt{\beta A_2}R \geq 1,\] and so we can lower bound 32 as follows: \[\begin{align} \left(\frac{1}{\beta A_2}\right)^{d/2} \int_0^{\sqrt{\beta A_2}R} e^{-\frac{r^2}{2}} r^{d-1}\; dr &\geq \left(\frac{1}{\beta A_2}\right)^{d/2} \int_1^{\sqrt{\beta A_2}R} e^{-\frac{r^2}{2}} r^{d-1}\; dr\\ &\geq \left(\frac{1}{\beta A_2}\right)^{d/2} \int_1^{\sqrt{\beta A_2}R} e^{-\frac{r^2}{2}} r \; dr \\ &= \left(\frac{1}{\beta A_2}\right)^{d/2} \left.\left(-e^{-\frac{r^2}{2}}\right)\right|_{1}^{\sqrt{\beta A_2} R}\\ &= \left(\frac{1}{\beta A_2}\right)^{d/2} \left(e^{-1/2} - e^{-\frac{\beta A_2 R^2}{2}}\right)\\ &= \left(\frac{1}{\beta A_2}\right)^{d/2} \left(e^{-1/2} - e^{-\beta \varepsilon}\right). \end{align}\]

Furthermore, \[e^{-1/2} - e^{-\beta \varepsilon} \geq \frac{1}{2}e^{-1/2} \iff \log\frac{1}{2} - \frac{1}{2} \geq -\beta \varepsilon \iff \beta \geq \frac{1}{2\varepsilon} + \frac{\log 2}{\varepsilon} = \frac{1+2\log 2}{2\varepsilon}.\] Therefore, for such values of \(\beta\) \[\left(\frac{1}{\beta A_2}\right)^{d/2} \int_0^{\sqrt{\beta A_2}R} e^{-\frac{r^2}{2}} r^{d-1}\; dr \geq \frac{1}{2} \left(\frac{1}{\beta A_2}\right)^{d/2}e^{-1/2},\] which allows us to conclude that \[\int_{\mathcal{B}(R, y)} e^{-\beta \frac{A_2}{2} d_g(y, x)^2} \; \mathrm{dVol}_g(x) \geq 2^{d-1} \pi^{1/2} \frac{\Gamma(\frac{d+1}{2})^{d-1}}{\Gamma(\frac{d}{2})^d d^{d-1}} \left(\frac{1}{\beta A_2}\right)^{d/2}e^{-1/2} .\] ◻

With the above two auxiliary lemmas in mind, we will now write the proof of 31. It has the same structure as in [18]. Nevertheless, we have decided to incorporate it in our work for completeness, as the theorem is now stated for any general compact manifold \(M\).

Proof of 31. From the previous lemmas we obtained the following bound: \[\nu(F \geq \varepsilon) \leq \frac{e^{-\beta \varepsilon} \mathrm{Vol}(M)}{2^{d-1} \pi^{1/2} \frac{\Gamma(\frac{d+1}{2})^{d-1}}{\Gamma(\frac{d}{2})^d d^{d-1}} \left(\frac{1}{\beta A_2}\right)^{d/2}e^{-1/2}},\] for \(\beta \geq \frac{1+2\log 2}{2\varepsilon}\). If we now define \[\mathscr{B}_d := \frac{e^{1/2}\Gamma(\frac{d}{2})^d d^{d-1}A_2^{d/2} \mathrm{Vol}(M)}{2^{d-1} \pi^{1/2} \Gamma(\frac{d+1}{2})^{d-1} },\] it only remains to find a value of \(\beta\) such that \[\label{goalineq} \beta^{d/2} e^{-\beta \varepsilon} \mathscr{B}_d \leq \delta.\tag{33}\] Taking logarithms in this expression we obtain \[\frac{d}{2}\log \beta - \beta\varepsilon + \log \mathscr{B}_d \leq \log \delta,\] which can be rewritten as \[\label{eq:inequalitybeta} \beta\varepsilon - \frac{d}{2} \log \beta \geq \log \frac{1}{\delta} + \log \mathscr{B}_d.\tag{34}\] Now, note that \[\frac{\beta\varepsilon}{2} - \frac{d}{2} \log \beta\] is an increasing function of \(\beta\) for \(\beta \geq \frac{d}{\varepsilon}\). Therefore, for such values \[\frac{\beta\varepsilon}{2} - \frac{d}{2}\log \beta \geq \frac{d}{2}\Big(1 - \log \frac{d}{\varepsilon}\Big).\] Inserting this inequality into 34 , it is easy to see that we only need to find the values of \(\beta\) for which \[\frac{\beta\varepsilon}{2} + \frac{d}{2}\Big(1 - \log \frac{d}{\varepsilon}\Big) \geq \log \frac{1}{\delta} + \log \mathscr{B}_d,\] that is, \[\beta \geq \frac{2}{\varepsilon}\left(\log \frac{1}{\delta} + \log \mathscr{B}_d + \frac{d}{2}\Big(\log \frac{d}{\varepsilon} - 1\Big)\right).\]

Let us express the logarithm of \(\mathscr{B}_d\) explicitly: \[\begin{align} \log \mathscr{B}_d &= \log \left(\frac{e^{1/2}\Gamma(\frac{d}{2})^d d^{d-1}A_2^{d/2} \mathrm{Vol}(M)}{2^{d-1} \pi^{1/2} \Gamma(\frac{d+1}{2})^{d-1} }\right) \\&= \frac{1}{2} + d \log\Gamma\left(\frac{d}{2}\right) + (d-1)\log d + \frac{d}{2}\log A_2 + \log \mathrm{Vol}(M) \\&\quad - (d-1)\log 2 -\frac{1}{2}\log\pi - (d-1)\log\Gamma\left(\frac{d+1}{2}\right) \\&\leq \frac{1}{2} + d \log\Gamma\left(\frac{d}{2}\right) + (d-1)\log d + \frac{d}{2}\log A_2 + \log \mathrm{Vol}(M) \\&\leq \frac{1}{2} + d \log\Gamma\left(\frac{d}{2}\right) + d\log (dA_2 \mathrm{Vol}(M)). \end{align}\] Therefore, we can choose \[\begin{align} \beta &\geq \frac{2}{\varepsilon}\left(\frac{1}{2} + d \log\Gamma\left(\frac{d}{2}\right) + d\log\frac{d^2A_2 \mathrm{Vol}(M)}{\delta\varepsilon}\right)\\ &\geq \frac{2}{\varepsilon}\left(\log \frac{1}{\delta} + \log \mathscr{B}_d + \frac{d}{2}(\log \frac{d}{\varepsilon} - 1)\right). \end{align}\]

Now, for \(d > 2\) we can use the following upper bound on the Gamma function (see [52]) \[\Gamma\left(\frac{d}{2}\right) < \sqrt{2\pi} \left(\frac{d - 1}{2e}\right)^{(d-1)/2},\] and so \[\log \Gamma\left(\frac{d}{2}\right) < \log \sqrt{2\pi} + \frac{d-1}{2}\log \frac{d-1}{2e} < d\log(d\sqrt{2\pi}).\] Thus, a suitable lower bound for \(\beta\) is \[\beta \geq \frac{2}{\varepsilon}\left( \frac{1}{2} + d^2 \log(d\sqrt{2\pi}) +d\log\frac{d^2A_2 \mathrm{Vol}(M)}{\delta\varepsilon} \right).\] Lastly, to simplify this, we note that whenever \[\beta \geq \frac{2}{\varepsilon}\left( \frac{1}{2} +d^2\log\frac{d^3A_2 \mathrm{Vol}(M)\sqrt{2\pi}}{\delta\varepsilon} \right),\] 33 holds, finishing the proof.

In particular, if \(d = 2\), since \(\Gamma(1) = 1\), we can bound \[\log \mathscr{B}_d \leq \frac{1}{2} + \log 2 + \log A_2 + \log \mathrm{Vol}(M) = \frac{1}{2} + \log 2 + \log(A_2 \mathrm{Vol}(M)),\] and so it suffices to choose \[\beta \geq \frac{2}{\varepsilon}\left(2\log \frac{2A_2\mathrm{Vol}(M)}{\varepsilon\delta} + \frac{1}{2} + \log 2\right).\] In particular, the general bound found for \(d > 2\) is also sufficient. ◻

Remark 32. The choice of \(\beta\) in 31 scales as \[\Omega\left(\frac{d^2}{\varepsilon}\log\left(\frac{dA_2 \mathrm{Vol}(M)}{\delta \varepsilon}\right)\right).\]

In particular, in order for \(\beta\) to grow polynomially with the dimension \(d\), the value of \(\mathrm{Vol}(M)\) and \(A_2\) can grow exponentially with the dimension of \(M\).

7 Examples↩︎

In this section we will consider two scenarios in which most of the assumptions of 30 can be easily verified—in particular, the existence of a unique minimum for the projected function, and those regarding the eigenvalues of its Hessian at the critical points. We will study the functions considered in the trace ratio minimization problem, and in the problem of minimizing the energy of a two-dimensional classical ferromagnetic Ising spin model.

7.1 Trace ratio minimization↩︎

The trace ratio minimization problem consists in finding the minimum of the function \[\require{physics} \begin{align} \label{eq:traceratiofunction} \begin{aligned} f: \mathrm{V}_k(\mathbb{C}^n) \to \mathbb{R},\quad X \mapsto \frac{\Tr(X^\dagger A X)}{\Tr(X^\dagger B X)}, \end{aligned} \end{align}\tag{35}\] where \(\mathrm{V}_k(\mathbb{C}^n)\) denotes the Stiefel manifold of linear isometries from \(\mathbb{C}^k\) to \(\mathbb{C}^n\), with \(k \leq n\) (cf. 8.4.2), and the matrices \(A, B \in \mathbb{C}^{n \times n}\) are Hermitian with \(B\) positive definite.

Finding the minimum of this function—or the maximum, which corresponds to the minimization problem for \((-A, B)\)—has attracted sustained attention in the literature. This problem is related to principal component analysis (PCA) [28] and more generally to dimensionality reduction techniques, such as graph embedding [29]. Furthermore, very similar problems are studied in the context of quantum information [53], [54].

In order to apply our results to the function studied in this problem, we want to find a projected version of \(f\) which has a unique minimum. To that end, let us first lift it to \(\mathrm{U}(n)\) by defining \[\require{physics} \label{eq:traceratioUn} F: \mathrm{U}(n) \to \mathbb{R},\quad U := (X|Y) \mapsto \frac{\Tr(X^\dagger A X)}{\Tr(X^\dagger B X)},\tag{36}\] where \(X\) corresponds to the first \(k\) columns of \(U\) and \(Y\) corresponds to its last \(n-k\) columns. If we consider the projection map \[\pi_1: \mathrm{U}(n) \to \mathrm{V}_k(\mathbb{C}^n),\quad (X | Y) \mapsto X\] studied in 8.4.2, it is clear that \[F = f \circ \pi_1.\]

If we now consider the actions of the groups \(\mathrm{U}(k) \times \mathrm{U}(n-k)\) and \(\mathrm{U}(k)\) on \(\mathrm{U}(n)\) and \(\mathrm{V}_k(\mathbb{C}^n)\), respectively, defined as

Figure 2: image.

Figure 3: image.

we showed that they induce the projection maps onto the Grassmann manifold \[\pi_2: \mathrm{V}_k(\mathbb{C}^n) \to \mathrm{Gr}_k(\mathbb{C}^n)\quad \mathrm{and}\quad \pi_3: \mathrm{U}(n) \to \mathrm{Gr}_k(\mathbb{C}^n),\] as seen in 8.4.3. Furthermore, \(F\) is invariant under the above action on \(\mathrm{U}(n)\): given any \(U = (X | Y ) \in \mathrm{U}(n)\) and any \((A, B) \in \mathrm{U}(k) \times \mathrm{U}(n-k)\), it holds that \[\begin{align} F\bigg(U \begin{pmatrix} A& 0\\ 0 & B \end{pmatrix}\bigg) = F((XA|YB)) = F((X|Y)) = F(U). \end{align}\] Therefore, we can project \(F\) onto \(\mathrm{Gr}_k(\mathbb{C}^n)\), obtaining the function \[\require{physics} \begin{align} \label{eq:traceratiograssmann} \begin{aligned} \tilde{f}: \mathrm{Gr}_k(\mathbb{C}^n) \to \mathbb{R},\quad P \mapsto \frac{\Tr(P A)}{\Tr(P B)}, \end{aligned} \end{align}\tag{37}\] where we are using the description of the Grassmann manifold given in 8.4.3, i.e. \[\require{physics} \mathrm{Gr}_k(\mathbb{C}^n) = \{P \in \mathbb{C}^{n \times n} : P^\dagger = P, P^2 = P, \Tr(P) = k\}.\] The function \(\tilde{f}\) can also be seen as a projected version of \(f\) via \(\pi_2\), by simply rewriting \[\require{physics} \frac{\Tr(X^\dagger A X)}{\Tr(X^\dagger B X)} = \frac{\Tr(XX^\dagger A)}{\Tr(XX^\dagger B)},\] and denoting \(P = XX^\dagger\).

Recall that, if we endow \(\mathrm{U}(n)\) with its bi-invariant metric \(g\), there exists a metric \(h_1\) on \(\mathrm{V}_k(\mathbb{C}^n)\) and a metric \(h_2\) on \(\mathrm{Gr}_k(\mathbb{C}^n)\) for which the projections \[\pi_1: (\mathrm{U}(n), g) \to (\mathrm{V}_k(\mathbb{C}^n), h_1),\quad \mathrm{and}\quad \pi_3: (\mathrm{U}(n), g) \to (\mathrm{Gr}_k(\mathbb{C}^n), h_2),\] are Riemannian submersions with totally geodesic fibers. Therefore, by 70 we know that \(\pi_2: (\mathrm{V}_k(\mathbb{C}^n), h_1) \to (\mathrm{Gr}_k(\mathbb{C}^n), h_2)\) is also a Riemannian submersion with totally geodesic fibers.

This way, we have defined three functions—namely \(F\), \(f\) and \(\tilde{f}\)—verifying \[F = f \circ \pi_1,\quad \mathrm{and}\quad F = \tilde{f} \circ \pi_3,\] (see 4). Having defined the trace minimization problem on \(\mathrm{U}(n)\), \(\mathrm{V}_k(\mathbb{C}^n)\) and \(\mathrm{Gr}_k(\mathbb{C}^n)\), our aim will be to obtain log-Sobolev inequalities for the Markov triples \[(\mathrm{U}(n), \nu_F, \Gamma_g),\quad (\mathrm{V}_k(\mathbb{C}^n), \nu_f, \Gamma_{h_1}), \quad \mathrm{and}\quad (\mathrm{Gr}_k(\mathbb{C}^n), \nu_{\tilde{f}}, \Gamma_{h_2}),\] induced by the operators \[\label{eq:diferentesoperators} \operatorname{L}_F := -\mathrm{grad}_g\, F + \frac{1}{\beta}\Delta_g,\;\operatorname{L}_f := -\mathrm{grad}_{h_1}\, f + \frac{1}{\beta}\Delta_{h_1},\;\mathrm{and}\;\operatorname{L}_{\tilde{f}} := -\mathrm{grad}_{h_2} \, \tilde{f} + \frac{1}{\beta}\Delta_{h_2},\tag{38}\] where \(\beta > 0\) is some fixed constant. Recall that, when considering these operators, the associated Gibbs measures are given by \[\nu_F = \frac{1}{Z_F} e^{-\beta F},\quad \nu_f = \frac{1}{Z_f}e^{-\beta f},\quad \text{and}\quad \nu_{\tilde{f}} = \frac{1}{Z_{\tilde{f}}} e^{-\beta \tilde{f}},\] where \(Z_F, Z_f\) and \(Z_{\tilde{f}}\) are their respective normalizing constants, and the carré du champ operators are given by \[\Gamma_g(\phi_1) = \frac{1}{\beta}|\mathrm{grad}_g\, \phi_1|^2_g,\quad \Gamma_{h_1}(\phi_2) = \frac{1}{\beta}|\mathrm{grad}_{h_1}\, \phi_2|^2_{h_1},\quad \Gamma_{h_2}(\phi_3) = \frac{1}{\beta}|\mathrm{grad}_{h_2}\, \phi_3|^2_{h_2},\] for any \(C^2\) functions \(\phi_1, \phi_2, \phi_3\) in \(\mathrm{U}(n)\), \(\mathrm{V}_k(\mathbb{C}^n)\) and \(\mathrm{Gr}_k(\mathbb{C}^n)\), respectively.

Figure 4: Sketch describing the Markov triples, functions and Riemannian submersions considered in this subsection, where f, F and \tilde{f} are as defined in 36 35 37 , respectively.

The function \(\tilde{f}\) has been studied in depth13 in [30]. Two of the main results proven in the aforementioned reference concern the study of local minima and global minima of \(\tilde{f}\). In particular, they prove that \(\tilde{f}\) has no spurious local minima, and that under generic assumptions, its global minimum is unique and the Hessian of \(\tilde{f}\) at the global minimum is non-degenerate.

Proposition 33 ([30]). Every local minimum of \(\tilde{f}\) is a global minimum.

Proposition 34 ([30]). Let \(P\) be a global minimum of \(\tilde{f}\). Assume that there is a gap between the \(k\)-th and the \(k+1\)-th largest eigenvalue of \(\Phi(P) := A - \tilde{f}(P) B\). Then the global maximum is unique and its Hessian is non-degenerate at \(P\).

Let us now characterize the critical points of \(\tilde{f}\). We want to ensure that they are isolated, separated by a sufficiently large distance, and that there always exists an escape direction for every saddle point. This will be done in the following subsection.

7.1.1 Critical points of \(\tilde{f}\)↩︎

Although it is difficult to characterize the critical points of \(\tilde{f}\) in the general case, we will consider two settings in which the critical points of \(\tilde{f}\) can be described explicitly, namely when \(\tilde{f}\) is defined on \(\mathrm{Gr}_1(\mathbb{C}^n)\) and when \(B = \mathbb{1}\). We believe that similar results can be derived in more general scenarios.

First, note that in [30] it is proven that the critical points of \(\tilde{f}\) correspond to those \(P \in \mathrm{Gr}_k(\mathbb{C}^n)\) for which there exists a common orthonormal eigenbasis for \(\Phi(P) = A - \tilde{f}(P) B\) and \(P\). This way, the question of finding the set of critical points of \(\tilde{f}\) reduces to characterizing the set \[\require{physics} \bigg\{P \in \mathrm{Gr}_k(\mathbb{C}^n) : \bigg[P, A - \frac{\Tr(PA)}{\Tr(PB)}B\bigg] = 0\bigg\}.\] To do so, we will make use of the following simultaneous diagonalization result.

Proposition 35 ([55]). Let \(A, B\) be two Hermitian matrices. If \(B\) is positive definite, then there exists a non-singular matrix \(S\) such that \(B = SS^\dagger\) and \(A = S \Lambda S^{\dagger}\), where \(\Lambda\) is real-diagonal. Furthermore, the diagonal entries of \(\Lambda\) correspond to the eigenvalues of \(B^{-1}A\).

Lastly, to study the escape directions of \(\tilde{f}\) at its saddle points, we will analyze its Hessian. It is shown in [30] that the Hessian of \(\tilde{f}\) at a critical point \(P\) in the direction of \(X = \begin{pmatrix} X_{11} & X_{12}\\ X_{12}^\dagger & X_{22} \end{pmatrix} \in T_P\mathrm{Gr}_k(\mathbb{C}^n)\)—where \(X_{11}, X_{12}, X_{22}\) are matrices of size \(k \times k\), \(k \times (n-k)\) and \((n-k) \times (n-k)\), respectively—is given by \[\require{physics} \nabla^2 \tilde{f}(P)[X, X] = \frac{2}{\Tr(BP)} \sum_{i = 1}^k \sum_{j = 1}^{n-k} |z_{ij}|^2 (\sigma_j - \lambda_i),\] where \((z_{ij})\) are the coordinates of \(X_{12}\), \(\{\sigma_j\}_{j = 1}^{n-k}\) correspond to the eigenvalues of \(\Phi(P)\) associated with the zero eigenvectors of \(P\), and \(\{\lambda_i\}_{i = 1}^{k}\) correspond to the eigenvalues of \(\Phi(P)\) associated with the non-zero eigenvectors of \(P\).

7.1.1.1 First case: rank-one \(P\).


Let us first consider the case when \(\tilde{f}\) is defined on \(\mathrm{Gr}_1(\mathbb{C}^n)\). Here, we can rewrite 35 as \[\require{physics} \begin{align} \label{eq:traceratioGr1} \begin{aligned} \tilde{f}: \mathrm{Gr}_1(\mathbb{C}^n) \to \mathbb{R},\quad \ketbra{\varphi} \mapsto \frac{\Tr(\ketbra{\varphi} A)}{\Tr(\ketbra{\varphi} B)}, \end{aligned} \end{align}\tag{39}\] where \(\ket{\varphi}\) is some unit vector in \(\mathbb{C}^n\), and \(\bra{\varphi} := \ket{\varphi}^\dagger\). Moreover, the set of critical points of \(\tilde{f}\) can be easily described.

Proposition 36. Let \(\tilde{f}\) be as defined in 39 , and assume that \(B^{-1}A\) has \(n\) distinct eigenvalues. Then \(\tilde{f}\) has exactly \(n\) critical points.

Proof. As we saw earlier, the critical points of \(\tilde{f}\) correspond to the set of rank-one projectors \(\ketbra{\varphi}\) that satisfy \[[\ketbra{\varphi}, \Phi(\ketbra{\varphi})] = 0.\] This condition can be rewritten as \[\Phi(\ketbra{\varphi}) \ketbra{\varphi} = \ketbra{\varphi} \Phi(\ketbra{\varphi}).\] Now, since \(\braket{\varphi} = 1\), it holds that \[\begin{align} \Phi(\ketbra{\varphi}) \ket{\varphi} &=\ketbra{\varphi} \Phi(\ketbra{\varphi}) \ket{\varphi}\\ &= \bra{\varphi}(A - \tilde{f}(\ketbra{\varphi})B)\ket{\varphi}\ket{\varphi} \\&= (\bra{\varphi}A\ket{\varphi} - \tilde{f}(\ketbra{\varphi})\bra{\varphi}B\ket{\varphi})\ket{\varphi} \\&= 0, \end{align}\] and so \[\label{eq:psieig} A\ket{\varphi} = \tilde{f}(\ketbra{\varphi}) B \ket{\varphi}.\tag{40}\]

Using 35 we know that there exists some non-singular matrix \(S\) such that \(A = S\Lambda S^\dagger\) and \(B = SS^\dagger\), where \(\Lambda\) is a diagonal matrix with diagonal entries equal to the eigenvalues of \(A^{-1}B\). Therefore, 40 can be rewritten as \[S\Lambda S^\dagger \ket{\varphi} = \tilde{f}(\ketbra{\varphi}) S S^\dagger \ket{\varphi},\] and so \[\Lambda S^\dagger \ket{\varphi} = \tilde{f}(\ketbra{\varphi}) S^\dagger \ket{\varphi},\] which implies that \(S^\dagger \ket{\varphi}\) is an eigenvector of \(\Lambda\). Since we are assuming that \(\Lambda\) has \(n\) distinct eigenvalues, then, up to a complex phase, there are only \(n\) possible choices for \(S^\dagger \ket{\varphi}\), which in turn implies that \(\tilde{f}\) has \(n\) critical points. ◻

We have proven that the set of critical points of \(\tilde{f}\) corresponds to \[\{\ketbra{\varphi} \in \mathrm{Gr}_1(\mathbb{C}^n): S^\dagger \ket{\varphi} \mathrm{ is an eigenvector of } \Lambda\},\] where \(S\) and \(\Lambda\) are as defined in the previous proof. Now, using the expression for the distance in the Grassmann manifold [56] we know that the minimum distance between any two critical points of \(\tilde{f}\) is given by \[D := \min\Big\{\arccos(|\bra{\varphi_1}\ket{\varphi_2}|) \in [0, \frac{\pi}{2}] : \ketbra{\varphi_1},\ketbra{\varphi_2} \mathrm{ critical points of } \tilde{f}\Big\}.\]

To conclude, we analyze the Hessian of \(\tilde{f}\) at the critical points. As we mentioned earlier, at any critical point \(\ketbra{\varphi}\) of \(\tilde{f}\), its Hessian in the direction of \(X \in T_P \mathrm{Gr}_1(\mathbb{C}^n)\) can be written as \[\require{physics} \nabla^2 \tilde{f}(\ketbra{\varphi})[X, X] = \frac{2}{\Tr(B\ketbra{\varphi})} \sum_{j = 1}^{n-1} |z_{1j}|^2 (\sigma_j - \lambda_1).\] Moreover, recall that \(\lambda_1\) is the eigenvalue of \(\Phi(\ketbra{\varphi})\) associated with \(\ket{\varphi}\), which we showed in the proof of 36 to be zero. Thus, we can rewrite this expression as \[\require{physics} \nabla^2 \tilde{f}(\ketbra{\varphi})[X, X] = \frac{2}{\Tr(B\ketbra{\varphi})} \sum_{j = 1}^{n-1} |z_{1j}|^2 \sigma_j,\] where the values \(\sigma_j\) correspond to the eigenvalues of \[A - \tilde{f}(\ketbra{\varphi}) B\] not associated with \(\ket{\varphi}\). Using 33 we know that there is only one critical point \(\ketbra{\varphi}\) for which all of the \(\sigma_j\) are non-negative, and only one critical point for which all of the \(\sigma_j\) are non-positive. Therefore, denoting by \(\tilde{\mathcal{C}}\) the set of critical points of \(\tilde{f}\), we know that for every \(P \in \tilde{\mathcal{C}}\), \(\lambda_{\min}(\nabla^2\tilde{f}(P)) \leq - \lambda_* < 0\), with \[-\lambda_* := \max_{P \in \tilde{\mathcal{C}}} \min \{\sigma_i(P) : i \in \{1, \dotsc, n-1\}\},\] where \(\sigma_i(P)\) denotes the \(i\)-th eigenvalue \(\sigma_i\) at the critical point \(P\).

7.1.1.2 Second case: \(B = \mathbb{1}\).


We can rewrite 35 as \[\require{physics} \begin{align} \label{eq:traceratiograssmannBid} \begin{aligned} \tilde{f}: \mathrm{Gr}_k(\mathbb{C}^n) \to \mathbb{R},\quad P \mapsto \frac{\Tr(P A)}{k}. \end{aligned} \end{align}\tag{41}\] In this case, any two critical points of \(\tilde{f}\) lay in each other’s cut locus, which in particular implies that \[d(P, Q) \geq i(\mathrm{Gr}_k(\mathbb{C}^n)),\] for any two critical points \(P, Q\) of \(\tilde{f}\) (cf. 8.3).

Describing the cut locus of a point in a manifold is not an easy task in general. Nevertheless, for the case of the Grassmann manifold with the induced metric, the cut locus of a point \(P = U U^\dagger\) can be described easily [56]. Indeed, for any \(P = UU^\dagger \in \mathrm{Gr}_k(\mathbb{C}^n)\), its cut locus \(C(P)\) is given by \[\label{eq:cutlocusdescription} C(P) = \{ Q = YY^\dagger \in \mathrm{Gr}_k(\mathbb{C}^n) : \rank(U^\dagger Y) < k\}.\tag{42}\] Therefore, the cut locus of a projector \(P\) in \(\mathrm{Gr}_k(\mathbb{C}^n)\) corresponds to those projectors whose associated subspace contains at least one orthogonal direction to the subspace on which \(P\) projects.

Proposition 37. Let \(\tilde{f}\) be defined as in 41 , and assume that \(A\) has \(n\) distinct eigenvalues. Then \(\tilde{f}\) has \(\binom{n}{k}\) distinct critical points, which are separated by a distance of at least \(\frac{\sqrt{2}\pi}{2}\).

Proof. When \(B = \mathbb{1}\), any critical point \(P\) of \(\tilde{f}\) satisfies \[[P, A - \tilde{f}(P)] = 0,\] which is equivalent to the condition that \[[P, A] = 0.\] This way, for every critical point \(P\) of \(\tilde{f}\), we know that \(P\) and \(A\) can be diagonalized in the same basis. Assuming again that \(A\) has \(n\) distinct eigenvalues with associated eigenvectors \(\{\ket{\varphi_i}\}_{i = 1}^n\), it must be the case that we can write \(P\) as \[P = \sum_{j = i_1}^{i_k} \ketbra{\varphi_{j}},\] where \(\{i_1, \dotsc, i_k\} \subset \{1, \dotsc, n\}\). Thus, there are \(\binom{n}{k}\) choices of \(P\).

Using the above description of the critical points, we know that every two critical points will differ in at least one orthogonal direction. Thus, by 42 , any two critical points \(P\) and \(Q\) will lay in each other’s cut locus. Therefore, their distance is lower bounded by the injectivity radius of \(\mathrm{Gr}_k(\mathbb{C}^n)\) which in turn is greater than \(\frac{\sqrt{2}\pi}{2}\) (cf. 7). ◻

Let us now analyze the Hessian of \(\tilde{f}\) (cf. 41 ) at any critical point \(P\). Recall from the previous result that we can write \(A\) and \(P\) as \[A = \sum_{i = 1}^k \lambda_i \ketbra{\varphi_i} + \sum_{i = k+1}^{n} \sigma_i \ketbra{\varphi_i},\quad \mathrm{and}\quad P = \sum_{i = i}^k \ketbra{\varphi_i},\] where \(\{\ket{\varphi_i}\}_{i = 1}^n\) is a common eigenbasis for both matrices, and \(\{\lambda_i\}_{i = 1}^k\), \(\{\sigma_i\}_{k+1}^n\) are the associated eigenvalues for \(A\).

Furthermore, from the expression of the Hessian of \(\tilde{f}\) at some critical point \(P\), and using the fact that \(\Phi(P) = A - \tilde{f}(P)\), we know that \[\require{physics} \nabla^2 \tilde{f}(P)[X, X] = \frac{2}{\Tr(BP)} \sum_{i = 1}^k \sum_{j = 1}^{n-k} |z_{ij}|^2 (\sigma_j - \lambda_i).\]

This way, denoting by \(\tilde{\mathcal{C}}\) the set of critical points of \(\tilde{f}\), we know that for every \(P \in \tilde{\mathcal{C}}\), \(\lambda_{\min}(\nabla^2\tilde{f}(P)) \leq - \lambda_* < 0\), where \[-\lambda_* := \min\{ E_i - E_j : i,j \in \{1, \dotsc, n\}, E_i, E_j \mathrm{ eigenvalues of } A\}.\]

7.1.2 Rapid mixing and suboptimality for the Gibbs distributions↩︎

Lastly, let us apply 30 31 4 to the trace ratio problem. Recall that we are interested in the functions \(F, f\) and \(\tilde{f}\) defined on \((\mathrm{U}(n), g)\), \((\mathrm{V}_k(\mathbb{C}^n), h_1)\), \((\mathrm{Gr}_k(\mathbb{C}^n), h_2)\) as presented in 35 36 37 .

The following result allows us to guarantee that, for sufficiently large values of \(\beta\), the Langevin dynamics associated to the trace ratio problem in \(\mathrm{U}(n)\), \(\mathrm{V}_k(\mathbb{C}^n)\) and \(\mathrm{Gr}_k(\mathbb{C}^n)\) mix rapidly to the associated Gibbs distribution, which in turn finds the global minima.

Proposition 38. Let \(F\), \(f\) and \(\tilde{f}\) be defined as in 35 36 37 and let \(A_2\) be the Lipschitz constant of the gradient of \(\tilde{f}\). Let \(\operatorname{L}_F\), \(\operatorname{L}_f\) and \(\operatorname{L}_{\tilde{f}}\) be the operators defined as in 38 . For every \(\varepsilon \in (0, 1]\) and \(\delta \in (0, 1)\), if \(\beta\) is such that \[\beta \geq \frac{2}{\varepsilon}\bigg(\frac{1}{2} + 6n^4\log n + \frac{n^5(n+1)}{2}\log(2\pi) + n^4 \log \frac{A_2 \sqrt{2\pi}}{\delta\varepsilon}\bigg),\] then the Gibbs distribution \(\nu_F(x) = \frac{1}{Z_F} e^{-\beta F(x)}\) on \(\mathrm{U}(n)\) is such that \[\nu_F\left(F - \min_{y \in \mathrm{U}(n)} F(y) \geq \varepsilon\right) \leq \delta.\] If we define \[\varepsilon^M_{\max} := \min \Big\{\frac{\pi^2 A_2}{16}, 1\Big\},\] then, for every \(\varepsilon \in (0, \varepsilon^M_{\max}]\) and \(\delta \in (0, 1)\), if \(\beta\) is such that \[\beta \geq \frac{2}{\varepsilon}\bigg(\frac{1}{2} + k^2(2n-k)^2\bigg(3\log(k(2n-k)) +\log \frac{A_2 \sqrt{2\pi}}{\delta\varepsilon}\bigg) + k^3(2n-k)^2\frac{2n-k+1}{2}\log(2\pi) \bigg),\] the Gibbs distribution \(\nu_f(x) = \frac{1}{Z_f} e^{-\beta f(x)}\) on \(\mathrm{V}_k(\mathbb{C}^n)\) is such that \[\nu_f\left(f - \min_{y \in \mathrm{V}_k(\mathbb{C}^n)} f(y) \geq \varepsilon\right) \leq \delta.\] and if \(\beta\) is such that \[\beta \geq \frac{2}{\varepsilon}\bigg(\frac{1}{2} + 4k^2(n-k)^2\bigg(3\log (2k(n-k)) + \log \frac{A_2 \sqrt{2\pi}}{\delta\varepsilon} + k(n-k)\log(2\pi)\bigg) \bigg),\] the Gibbs distribution \(\nu_{\tilde{f}}(x) = \frac{1}{Z_{\tilde{f}}} e^{-\beta \tilde{f}(x)}\) on \(\mathrm{Gr}_k(\mathbb{C}^n)\) is such that \[\nu_{\tilde{f}}\left(\tilde{f} - \min_{y \in \mathrm{Gr}_k(\mathbb{C}^n)} \tilde{f}(y) \geq \varepsilon\right) \leq \delta.\]

Proof. In order to prove this result, it suffices to bound the injectivity radius and the volume of \(\mathrm{U}(n)\), \(\mathrm{V}_k(\mathbb{C}^n)\) and \(\mathrm{Gr}_k(\mathbb{C}^n)\), and apply 31. The bounds obtained in 7 for the injectivity radii of these manifolds are \[i(\mathrm{U}(n)) \geq \pi, \quad i(\mathrm{V}_k(\mathbb{C}^n)) \geq \frac{\sqrt{2}\pi}{2}, \quad \mathrm{and}\quad i(\mathrm{Gr}_k(\mathbb{C}^n)) \geq \frac{\sqrt{2}\pi}{2},\] and the bounds on their volume can be deduced from the expressions shown in 8.6.3, i.e. \[\mathrm{Vol}(\mathrm{U}(n)) \leq (2\pi)^{n(n+1)/2}, \quad \mathrm{Vol}(\mathrm{V}_k(\mathbb{C}^n)) \leq (2\pi)^{nk - \frac{k(k-1)}{2}},\quad \mathrm{Vol}(\mathrm{Gr}_k(\mathbb{C}^n)) \leq (2\pi)^{nk -k^2}.\] Lastly, the dimensions of \(\mathrm{U}(n)\), \(\mathrm{V}_k(\mathbb{C}^n)\) and \(\mathrm{Gr}_k(\mathbb{C}^n)\) can be substituted by the explicit expressions found in 8.6.3, i.e. \[\dim(\mathrm{U}(n)) = n^2,\quad \dim(\mathrm{V}_k(\mathbb{C}^n)) =2nk-k^2, \quad \dim(\mathrm{Gr}_k(\mathbb{C}^n)) = 2k(n-k).\] ◻

Next, let us first restate some of the assumptions needed in order for 30 to hold.

Assumption 19. The critical points of \(\tilde{f}\) are isolated points separated by some distance \(D\), i.e. \[d_{h_2}(x_1, x_2) \geq D, \quad \forall x_1, x_2 \in \tilde{\mathcal{C}},\;x_1 \neq x_2,\] where \(\tilde{\mathcal{C}}\) denotes the set of critical points of \(\tilde{f}\).

Assumption 20. For every saddle point \(y \in \tilde{\mathcal{S}}\), there exists an escape direction. That is, there exists some constant \(\lambda_* \in (0, 1]\) such that for every saddle point \(y \in \tilde{\mathcal{S}}\) it holds that \[\lambda_{\min}\big(\nabla^2 \tilde{F}(y)\big) \leq -\lambda_* < 0.\]

Assumption 21. The global minimum is an attractor for the process \(\tilde{X}_t\). In other words, the Hessian of \(\tilde{F}\) at the global minimum is non-degenerate. More precisely, at the unique minimum \(x^*\) of \(\tilde{F}\), we have \[\lambda_{\min}\big(\nabla^2 \tilde{F}(x^*)\big) \geq \lambda_*.\] where \(\lambda_*\) is the same constant as in 8.

Assumption 22. There exists a constant \(0 < C_{\tilde{f}} \leq 1\) such that \[|\mathrm{grad}_{h_2}\,\tilde{f}(x)|_{h_2} \geq C_{\tilde{f}} d_{h_2}(x, \tilde{\mathcal{C}}),\] for every \(x \in \mathrm{Gr}_k(\mathbb{C}^n)\).

Proposition 39. Let \(F\), \(f\) and \(\tilde{f}\) be defined as in 35 36 37 . Assume that there is a gap between the \(k\)-th and the \(k+1\)-th largest eigenvalue of \[\Phi(P^*)= A - \tilde{f}(P^*)B,\] where \(P^*\) is the global minimum of \(\tilde{f}\). Furthermore, assume that \(\tilde{f}\) satisfies —where in particular, \(\lambda_*\) depends on the gap of \(\Phi(P^*)\). Let \(\operatorname{L}_F\), \(\operatorname{L}_f\) and \(\operatorname{L}_{\tilde{f}}\) be the operators defined as in 38 , and let \(a, \beta > 0\) be such that \[\begin{align} a^2 &\geq \max\Big\{\frac{48 A_2 k(n-k)}{C^2_{\tilde{f}}}, \frac{544}{\lambda_*}\Big\}, \\\beta &\geq \max\Bigg\{\frac{72^2 n^{10} A_2 A^2_3 a^6}{\lambda_*^2}, \frac{9a^2}{D^2}, \frac{8a^2}{\pi^2}\Bigg\}, \end{align}\] where \(A_2\) and \(A_3\) are the Lipschitz constants of the gradient and the Hessian of \(\tilde{f}\), respectively. Then the Markov triple \((\mathrm{U}(n), \nu_F, \Gamma_g)\) associated with the operator \(\operatorname{L}_F\) satisfies an \(\mathrm{LSI}(\alpha)\) verifying \[\frac{1}{\alpha} = 4\beta A_2 \pi^2 n (1 + \sqrt{2})^2\max \Big\{\frac{184}{\lambda_*}, n (1 + \sqrt{2})^2 \beta\Big\},\] the Markov triple \((\mathrm{V}_k(\mathbb{C}^n), \nu_{f}, \Gamma_{h_1})\) associated with the operator \(\operatorname{L}_f\) satisfies an \(\mathrm{LSI}(\hat{\alpha})\) verifying \[\frac{1}{\hat{\alpha}} = 4\beta A_2 \pi^2 n (1 + \sqrt{2})^2\max \Big\{\frac{184}{\lambda_*}, k (1 + \sqrt{2})^2 \beta\Big\},\] and the Markov triple \((\mathrm{Gr}_k(\mathbb{C}^n), \nu_{\tilde{f}}, \Gamma_{h_2})\) associated with the operator \(\operatorname{L}_{\tilde{f}}\) satisfies an \(\mathrm{LSI}(\tilde{\alpha})\) verifying \[\frac{1}{\tilde{\alpha}} = \frac{736\beta A_2 \pi^2 n (1 + \sqrt{2})^2}{\lambda_*}.\]

Proof. We begin by applying 30 to the Markov triples \((\mathrm{U}(n), \nu_F, \Gamma_g)\) and
\((\mathrm{Gr}_k(\mathbb{C}^n), \nu_{\tilde{f}}, \Gamma_{h_2})\). We showed in 57 that the coefficients of the Riemann curvature tensor of \(\mathrm{U}(n)\) in normal coordinates when endowed with the bi-invariant metric are upper bounded by \(\frac{1}{2}\)—and so we can take \(\mathbf{K} = 1\) in 30. Furthermore, we also showed in 55 73 that the Ricci curvatures of the unitary group and the Grassmann manifold are non-negative, and so the constants \(R_{\mathrm{U}(n)}\) and \(R_{\mathrm{Gr}_k(\mathbb{C}^n)}\) from 30 can be taken to be \(0\).

Also, the Riemannian submersion \(\pi_1 : (\mathrm{U}(n), g) \to (\mathrm{Gr}_k(\mathbb{C}^n), h_2)\) has totally geodesic fibers, which implies by 21 22 that the sectional curvature of its fibers coincides with that of the total space. Since \((\mathrm{U}(n), g)\) has non-negative sectional curvature (cf. 55), we conclude that the sectional curvature of the fibers is also non-negative and so their Ricci curvature is non-negative as well.

The result now follows by substituting the manifold-dependent constants in the statement of 30 with the known bounds derived in 8. We use the lower bounds for the injectivity radii of \((\mathrm{U}(n), g)\) and \((\mathrm{Gr}_k(\mathbb{C}^n), h_2)\), along with the lower bound on the convexity radius of \((\mathrm{Gr}_k(\mathbb{C}^n), h_2)\) obtained in 7, i.e. \[\mathit{conv}(\mathrm{Gr}_k(\mathbb{C}^n)) \geq \frac{\sqrt{2}\pi}{4}.\] We also use the upper bound on the diameter of \((\mathrm{U}(n), g)\), and \((\mathrm{Gr}_k(\mathbb{C}^n), h_2)\) found in 9 56 , i.e. \[\mathrm{diam}(\mathrm{Gr}_k(\mathbb{C}^n)) \leq \mathrm{diam}(\mathrm{U}(n)) \leq \pi \sqrt{n}(1 + \sqrt{2}),\] and substitute the dimensions of \(\mathrm{U}(n)\) and \(\mathrm{Gr}_k(\mathbb{C}^n)\).

Lastly, we obtain a bound on the diameter of the fibers \(\mathrm{U}(k) \times \mathrm{U}(n-k)\) with the induced metric, which can be derived via 86 and 9, i.e. \[\mathrm{diam}(\mathrm{U}(k) \times \mathrm{U}(n-k)) \leq \sqrt{\pi^2 k (1+ \sqrt{2})^2 + \pi^2 (n-k) (1+\sqrt{2})^2} = \pi (1 + \sqrt{2}) \sqrt{n}.\]

All of the above bounds allow us to obtain the log-Sobolev inequality constants for the Markov triples \((\mathrm{U}(n), \nu_F, \Gamma_g)\) and \((\mathrm{Gr}_k(\mathbb{C}^n), \nu_{\tilde{f}}, \Gamma_{h_2})\).

To obtain the log-Sobolev inequality of the triple \((\mathrm{V}_k(\mathbb{C}^n), \nu_f, \Gamma_{h_1})\), we proceed as in the proof of 30: we first apply 10 to \((\mathrm{Gr}_k(\mathbb{C}^n), \nu_{\tilde{f}}, \Gamma_{h_2})\) to find that it satisfies a \(\mathrm{PI}(\kappa)\) with \[\kappa = \frac{\lambda_*}{184}.\] Then using the lifting theorem for Poincaré inequalities shown in 20, we lift the Poincaré inequality via \(\pi_2\) to ensure that \((\mathrm{V}_k(\mathbb{C}^n), \nu_f, \Gamma_{h_1})\) satisfies a \(\mathrm{PI}(\tilde{\kappa})\) where \[\frac{1}{\tilde{\kappa}} = \max\Big\{\frac{184}{\lambda_*}, (1+\sqrt{2})^2 k \beta\Big\},\] which can be then tightened using 24 to obtain an \(\mathrm{LSI}(\hat{\alpha})\), where \[\frac{1}{\hat{\alpha}} = \frac{4\beta A_2 \mathrm{diam}(\mathrm{V}_k(\mathbb{C}^n))^2}{\tilde{\kappa}}.\] Lastly, using the bound on the diameter of the Stiefel manifold shown in 56 we see that \[\mathrm{diam}(\mathrm{V}_k(\mathbb{C}^n)) \leq \pi \sqrt{n}(1 + \sqrt{2}),\] and so the result follows. ◻

Recall that in 7.1.1 we proved that the critical points of \(\tilde{f}\) are isolated when it is defined on \(\mathrm{Gr}_1(\mathbb{C}^n)\) or when \(B = \mathbb{1}\). Furthermore, we also gave an expression for the minimum distance between any two critical points, and the constant \(\lambda_*\) from 20.

Remark 40. In the case when \(\tilde{f}\) is defined on \(\mathrm{Gr}_1(\mathbb{C}^n)\) or \(B = \mathbb{1}\), 19 holds, i.e. the critical points of \(\tilde{f}\) are isolated, and the constant \(D\)—which lower bounds the distance between any two critical points—is given by \[D = \frac{\sqrt{2}\pi}{2},\] when \(B = \mathbb{1}\)—which in particular does not depend on \(k\) or \(n\)—and by \[D = \min\Big\{\arccos(|\bra{\varphi_1}\ket{\varphi_2}|) \in [0, \frac{\pi}{2}] : \ketbra{\varphi_1},\ketbra{\varphi_2} \mathrm{ critical points of } \tilde{f}\Big\},\] when \(\tilde{f}\) is defined on \(\mathrm{Gr}_1(\mathbb{C}^n)\).

Moreover, 20 21 hold, and the constant \(\lambda_*\) can be written as \[-\lambda_* := \min\{ E_i - E_j : i,j \in \{1, \dotsc, n\}, E_i, E_j \mathrm{ eigenvalues of } A\},\] when \(B = \mathbb{1}\), and as \[-\lambda_* := \max_{P \in \tilde{\mathcal{C}}} \min \{\sigma_i(P) : i \in \{1, \dotsc, n-1\}\}.\] when \(\tilde{f}\) is defined on \(\mathrm{Gr}_1(\mathbb{C}^n)\).

Lastly, let us restate 4.

Corollary 5. Let \(X_t\), \(\hat{X}_t\) and \(\tilde{X}_t\) be the Langevin diffusion processes generated by \(\operatorname{L}_F\), \(\operatorname{L}_f\), and \(\operatorname{L}_{\tilde{f}}\), with initial uniform distribution on \(\mathrm{U}(n)\), \(\mathrm{V}_k(\mathbb{C}^n)\) and \(\mathrm{Gr}_k(\mathbb{C}^n)\), respectively. Under the assumptions of 39, they converge exponentially fast to the Gibbs measures \(\nu_F\), \(\nu_f\), and \(\nu_{\tilde{f}}\), respectively, i.e. \[\begin{align} \norm{\nu_F - \rho_t}^2_{\mathrm{TV}} &\leq \beta\, e^{-2\alpha t} \max_{y \in M} F(y), \\ \norm{\nu_f - \hat{\rho}_t}^2_{\mathrm{TV}} &\leq \beta\, e^{-2\hat{\alpha} t} \max_{y \in \mathrm{V}_k(\mathbb{C}^n)} f(y) = \beta\, e^{-2\hat{\alpha} t} \max_{y \in \mathrm{U}(n)} F(y) ,\\ ||\nu_{\tilde{f}} - \tilde{\rho}_t||^2_{\mathrm{TV}} &\leq \beta\, e^{-2\tilde{\alpha} t} \max_{y \in \mathrm{Gr}_k(\mathbb{C}^n)} \tilde{f}(y) = \beta\, e^{-2\tilde{\alpha} t} \max_{y \in \mathrm{U}(n)} F(y), \end{align}\] for every \(t \geq 0\), where \(\rho_t, \hat{\rho}_t\) and \(\tilde{\rho}_t\) denote the distributions of \(X_t, \hat{X}_t\) and \(\tilde{X}_t\), and \(\alpha, \hat{\alpha}, \tilde{\alpha}\) denote the log-Sobolev inequality constants obtained in 39.

7.2 Two-dimensional Ising model↩︎

The second function that we consider is defined as \[\begin{align} \label{eq:isingfunction} \begin{aligned} F : \mathrm{SU}(2)^{\times n^2} &\rightarrow \mathbb{R}\\ (U_1, \dotsc, U_{n^2}) &\mapsto \bra{0}^{\otimes n^2} \big(U_1^\dagger \otimes \dotsm \otimes U_{n^2}^\dagger \big)H \big(U_1 \otimes \dotsm \otimes U_{n^2}\big) \ket{0}^{\otimes n^2}, \end{aligned} \end{align}\tag{43}\] where \(H\) corresponds to a ferromagnetic Ising model Hamiltonian, i.e. \[H = -\sum_{\langle I, J\rangle \in \Lambda} \sigma_z^I \otimes \sigma_z^J - \sum_{I\in V} h_I \sigma_z^I,\] over a square lattice \((V, \Lambda)\) with nearest-neighbor interactions, \(h_{(i, j)}\in \mathbb{R}\) are defined as \[\label{eq:definitionmagneticfield} h_{(i, j)} = \begin{cases} 4 + \varepsilon_{(i, j)} &\quad \text{if } 1 < i < n \text{ and } 1 < j < n,\\ 3 + \varepsilon_{(i, j)} &\quad \text{if } 1 < i < n\text{ xor } 1 < j < n,\\ 2 + \varepsilon_{(i, j)} &\quad \text{otherwise}, \end{cases}\tag{44}\] with \(\varepsilon_{(i, j)} > 0\) for every \((i, j) \in V\), and \(\sigma_z\) denotes the \(Z\)-Pauli matrix, i.e. \(\sigma_z = \begin{pmatrix} 1 & 0 \\ 0 & -1 \end{pmatrix}\). We denote by \(\ket{0}\) and \(\ket{1}\) the eigenvectors of \(\sigma_z\) with eigenvalue \(1\) and \(-1\), respectively. \(F\) is commonly regarded as the expectation value of \(H\) in a given state. Previous results for performing efficient Gibbs sampling for this model are known [57].

As in the previous example, the function \(F\) defined in 43 has a clear symmetry: let \(i \in \{1, \dotsc, n^2\}\) and let \(U_i, \tilde{U}_i \in \mathrm{SU}(2)\) be such that \[\tilde{U}_i\ket{0} = e^{i\theta} U_i\ket{0},\] for some \(\theta \in [0, 2\pi)\). Then \[F(U_1, \dotsc, U_i, \dotsc, U_{n^2}) = F(U_1, \dotsc, \tilde{U}_i, \dotsc, U_{n^2}).\]

The symmetry of \(F\) can be studied in terms of a group action on \(\mathrm{SU}(2)\). Indeed, if we consider the action of \(\mathrm{U}(1)\) on \(\mathrm{SU}(2)\) defined as \[\begin{align} \label{eq:actiononSU2} \begin{aligned} \mathrm{SU}(2) \times \mathrm{U}(1) &\to \mathrm{SU}(2)\\ (X, e^{i\theta}) &\mapsto X\begin{pmatrix} e^{i\theta} & 0\\ 0 & e^{-i\theta} \end{pmatrix}, \end{aligned} \end{align}\tag{45}\] it is clear that \(F\) is invariant under this action in any of its coordinates \((U_1, \dotsc, U_{n^2})\).

Note that the action shown in 45 is free: for every \(U \in \mathrm{SU}(2)\) and every \(e^{i\theta} \in \mathrm{U}(1)\), it holds that \[U\begin{pmatrix} e^{i\theta} & 0\\ 0 & e^{-i\theta} \end{pmatrix} = U \iff \begin{pmatrix} e^{i\theta} & 0\\ 0 & e^{-i\theta} \end{pmatrix} = \mathbb{1}\] which implies that \(e^{i\theta} = 1\).

Moreover, the action is smooth and isometric whenever \(\mathrm{SU}(2)\) is endowed with its bi-invariant metric \(g_{\mathit{bi}}\). Therefore, by 66, there exists a Riemannian metric \(h\) on \(\mathrm{U}(2)/\mathrm{U}(1)\) such that the projection \(\pi: (\mathrm{U}(2), g_{\mathit{bi}}) \to (\mathrm{U}(2)/\mathrm{U}(1), h)\) is a Riemannian submersion with totally geodesic fibers.

In fact, \(\pi\) corresponds to a well-studied Riemannian submersion, known as the Hopf fibration between \(\mathbb{S}^3\) and \(\mathbb{S}^2\). To illustrate this, let us introduce the following auxiliary result.

Proposition 41. Let \(\mathrm{SU}(2)\) be endowed with the bi-invariant metric \(g_{\mathit{bi}}\). Then \[(\mathrm{SU}(2), g_{\mathit{bi}}) \simeq (\mathbb{S}^3, 2g_{\textit{round}}),\] where \(g_{\textit{round}}\) is the standard round metric on the sphere \(\mathbb{S}^3\).

Proof. First, note that \[\mathrm{SU}(2) = \left\{\begin{pmatrix} a & -\overline{b} \\ b & \overline{a} \end{pmatrix}\;:\;a, b \in \mathbb{C},\; |a|^2+|b|^2=1\right\} \simeq \mathbb{S}^3.\] Recall from 48 that its lie algebra \(\mathfrak{su}(2)\) consists of \(2 \times 2\) anti-Hermitian matrices with zero trace, \[\mathfrak{su}(2) = T_{\mathbb{1}} \mathrm{SU}(2) = \left\{ \begin{pmatrix} ix & z\\ -\overline{z} & -ix \end{pmatrix}\;:\;x \in \mathbb{R},\;z \in \mathbb{C}\right\}.\]

Moreover, the bi-invariant of \(\mathrm{SU}(2)\) at the identity is given by \[\require{physics} \Tr \Bigg(\begin{pmatrix} -ix_2 & -z_2\\ \overline{z_2} & ix_2 \end{pmatrix} \begin{pmatrix} ix_1 & z_1\\ -\overline{z_1} & -ix_1 \end{pmatrix}\Bigg) = x_1x_2 + z_2 \overline{z_1} + z_1 \overline{z_2} + x_1 x_2\] for any two tangent vectors \(\begin{pmatrix} ix_2 & z_2\\ -\overline{z_2} & -ix_2 \end{pmatrix}\), \(\begin{pmatrix} ix_1 & z_1\\ -\overline{z_1} & -ix_1 \end{pmatrix} \in T_{\mathbb{1}}\mathrm{SU}(2)\), which corresponds to \(2g_{\mathit{Eucl}}\) in \(\mathbb{R}^3\).

Now, the metric \(g_{\mathit{bi}}\) is extended from \(T_\mathbb{1} \mathrm{SU}(2)\) to \(T\mathrm{SU}(2)\) as \[\require{physics} g_{\mathit{bi}}(AG, BG) := \Tr(G^\dagger B^\dagger A G) = \Tr (B^\dagger A).\]

Note that the map \[\begin{align} \mathbb{R}^4 \rightarrow \mathbb{R}^4,\quad \begin{pmatrix} x & y\\ -\overline{y} & \overline{x} \end{pmatrix} \mapsto \begin{pmatrix} x & y\\ -\overline{y} & \overline{x} \end{pmatrix}G \end{align}\] where \(x, y \in \mathbb{C}\), is an isometry in \(\mathbb{R}^4\), and so the extension of the of the scalar product coincides with twice the usual scalar product in the round sphere \(\mathbb{S}^3\), i.e. \(2g_{\mathit{round}}\). ◻

With this result in mind, it is clear that the action defined in 45 is equivalent to the action of \(\mathrm{U}(1)\) on \(\mathbb{S}^3\) given by \[\begin{align} \mathrm{U}(1) \times \mathbb{S}^3 \to \mathbb{S}^3,\quad (e^{i\theta}, x) \mapsto e^{i\theta}x. \end{align}\] This action induces a Riemannian submersion, known as the Hopf fibration between \((\mathbb{S}^3, \lambda g_{\mathit{round}})\) and \((\mathbb{S}^2, \frac{\lambda}{4}g_{\mathit{round}})\), for any \(\lambda > 0\) [58]. Using 41 it holds that the Riemannian submersion with totally geodesic fibers \[\pi:(\mathrm{SU}(2), g_{\mathit{bi}}) \to (\mathrm{SU}(2)/\mathrm{U}(1), h),\] corresponds to the Hopf fibration \[\pi: (\mathbb{S}^3, 2 g_{\mathit{round}}) \to (\mathbb{S}^2, \frac{1}{2}g_{\mathit{round}}).\]

Let us now study the function \(F\) on \(\mathrm{SU}(2)^{\times n^2}\) and its projected version \(\tilde{F}\) on \((\mathbb{S}^2)^{\times n^2}\), which can be understood as the function \(F\) defined on the product of Bloch spheres.

7.2.1 Critical points↩︎

Let us characterize the critical points of \(F\). To do so, we will first show that every eigenvector of \(H\) is a critical point of \(F\), and then prove that only the eigenvectors of \(H\) can be critical points.

Let us introduce an auxiliary result that will allow us to conclude that the eigenvectors of \(H\) are critical points of \(F\).

Lemma 15. Let \(A\) be an \(n\)-square complex matrix. Then \[\mathrm{Tr}(AX) = 0\] for every Hermitian matrix \(X\) if and only if \(A = 0\).

With this result in mind, let us study the critical points of an auxiliary function \(f\).

Proposition 42. Let \(\hat{H}\) be some Hermitian matrix and let \(f\) be defined as \[\begin{align} f: &\;\mathrm{SU}(n) \rightarrow \mathbb{R}\\ &U \mapsto \bra{0}U^\dagger \hat{H} U \ket{0}, \end{align}\] where \(\ket{0}\) denotes some fixed vector in \(\mathbb{C}^{n}\). Then \(U\ket{0}\) is an eigenvector of \(\hat{H}\) if and only if \(U\) is a critical point of \(f\).

Proof. Let us give an expression for the differential of \(f\) at any point \(U\) in the direction of a tangent vector \(iXU\), where \(X\) is some Hermitian matrix. To do so, we parametrize \(f\) along a geodesic starting at \(U\) in the direction of \(iXU\), \[\gamma(t) = e^{itX} U,\] yielding \[f(t) := f(\gamma(t)) = \bra{0} U^\dagger e^{-itX} \hat{H} e^{itX} U \ket{0}.\] Computing its derivative, we obtain \[\begin{align} f'(t) &= -i \bra{0} U^\dagger e^{-itX}X\hat{H}e^{itX}U\ket{0} + i\bra{0} U^\dagger e^{-itX}\hat{H}Xe^{itX}U\ket{0} \\&= i\bra{0} U^\dagger e^{-itX}[\hat{H}, X]e^{itX}U\ket{0}. \end{align}\] Using the notation \(\ket{U} := U\ket{0}\) we can rewrite this expression as \[f'(t) = i\bra{U} e^{-itX}[\hat{H}, X]e^{itX}\ket{U},\] which allows us to conclude that \[df|_U(iXU) = f'(0) = i\bra{U}[\hat{H}, X]\ket{U}.\]

Assume now that \(U\ket{0}\) is an eigenvector of \(\hat{H}\), i.e. \(\hat{H}\ket{U} = E \ket{U}\) for some \(E \in \mathbb{R}\). Thus, for every \(iX \in \mathfrak{su}(n)\) \[df|_U (iXU) = \bra{U}[\hat{H}, X]\ket{U} = \bra{U}\hat{H}X - X\hat{H}\ket{U} = E \bra{U}X - X\ket{U} = 0.\] To prove the converse, let \(U \in \mathrm{SU}(n)\) and suppose that \[\label{eq:eq1} \bra{U}[\hat{H}, X]\ket{U} = 0,\quad \forall iX \in \mathfrak{su}(n).\tag{46}\] We can write \(\ket{U}\) in the (orthonormal) eigenbasis of \(\hat{H}\); \(\ket{U} = \sum_{i = 1}^n \alpha_i \ket{\varphi_i}\), where \(\hat{H} \ket{\varphi_i} = E_i \ket{\varphi_i}\). With this expression, 46 becomes \[\left(\sum_{i = 1}^n \overline{\alpha_i} \bra{\varphi_i}\right)\hat{H}X\left(\sum_{i = 1}^n \alpha_i \ket{\varphi_i}\right) = \left(\sum_{i = 1}^n \overline{\alpha_i} \bra{\varphi_i}\right)X\hat{H}\left(\sum_{i = 1}^n \alpha_i \ket{\varphi_i}\right),\] which implies that \[\left(\sum_{i = 1}^n \overline{\alpha_i} E_i \bra{\varphi_i}\right)X\left(\sum_{i = 1}^n \alpha_i \ket{\varphi_i}\right) = \left(\sum_{i = 1}^n \overline{\alpha_i} \bra{\varphi_i}\right)X\left(\sum_{i = 1}^n \alpha_i E_i\ket{\varphi_i}\right),\] and so \[\sum_{i, j = 1}^n \overline{\alpha_i} \alpha_j E_i \bra{\varphi_i}X\ket{\varphi_j} = \sum_{i, j = 1}^n \overline{\alpha_i} \alpha_j E_j \bra{\varphi_i}X\ket{\varphi_j}.\]

Grouping all the terms on the left-hand side we obtain \[\sum_{i, j= 1}^n \overline{\alpha_i} \alpha_j (E_i-E_j) \bra{\varphi_i}X\ket{\varphi_j} = 0.\] If we define the matrix \(A = \left(\overline{\alpha_j}\alpha_i (E_j - E_i)\right)_{ij}\) we can rewrite the expression as \[\sum_{i, j = 1}^n \overline{\alpha_i} \alpha_j (E_i-E_j) \bra{\varphi_i}X\ket{\varphi_j} = \sum_{i, j = 1}^n A_{ji} X_{ij} = \mathrm{Tr}(AX) = 0\] for every \(iX \in \mathfrak{su}(n)\).

In order to apply 15, take \(Y\) to be any Hermitian matrix. Therefore, we can write \(Y = \Lambda + X\) where \(\Lambda\) is diagonal and \(X\) is such that \(iX \in \mathfrak{su}(n)\), thus \[\mathrm{Tr}(AY) = \mathrm{Tr}(A\Lambda) + \mathrm{Tr}(AX) = \mathrm{Tr}(A\Lambda) = 0,\] where the last equality follows from the fact that \(\Lambda\) is diagonal. Indeed \[\mathrm{Tr}(A\Lambda) = \sum_{i = 1}^n A_{ii}\Lambda_{ii} = \sum_{i = 1}^n |\alpha_i|^2 (E_i-E_i)\Lambda_{ii} = 0.\] Using 15 we can conclude that \[\overline{\alpha_j}\alpha_i (E_j - E_i) = 0,\quad \forall i, j \in \{1, \dotsc, n\}.\] This implies that, if all of the eigenvalues are different, at most one \(\alpha_i\) is non-zero, implying that \(\ket{U}\) is an eigenvector of \(\hat{H}\). In the case where some eigenvalues are equal, we can conclude that \(\ket{U}\) has to be a linear combination of eigenvectors associated with the same eigenvalue, implying that \(\ket{U}\) is too an eigenvector of \(\hat{H}\). ◻

In the next auxiliary result, we will show how to compute the gradient of a function defined on the product of two manifolds, when endowed with the product metric.

Lemma 16. Let \((M_1, g_1)\) and \((M_2, g_2)\) be two Riemannian manifolds and let \(M_1 \times M_2\) be the product manifold endowed with the product metric \(g\). Let us define \[\begin{align} \hat{f}: &\;M_1 \times M_2 \rightarrow \mathbb{R}\\ &\quad(x, y) \mapsto f_1(x)f_2(y), \end{align}\] where \(f_i: M_i \rightarrow \mathbb{R}\) are smooth functions. Then \[\mathrm{grad}_{g}\, \hat{f}(x,y) = (f_2(y) \mathrm{grad}_{g_1}\, f_1(x),\, f_1(x) \mathrm{grad}_{g_2}\, f_2(y)), \quad \forall (x, y) \in M_1 \times M_2.\]

Proof. Let \((x, y) \in M_1 \times M_2\), let \(W \in T_{(x,y)} (M_1 \times M_2)\) and let \(\gamma(t)\) be the geodesic such that \(\gamma(0) = x\) and \(\gamma'(0) = W\). Writing both \(W\) and \(\gamma(t)\) in coordinates, we know that \(W = (U, V)\) with \(U \in T_x M_1\) and \(V \in T_y M_2\), and \(\gamma(t) = (\gamma_1(t), \gamma_2(t))\) with \(\gamma_1(0) = x\), \(\gamma_1'(0) = U\), \(\gamma_2(0) = y\), and \(\gamma_2'(0) = V\).

Using the definition of the differential of \(\hat{f}\) and the product rule in \(\mathbb{R}\) we conclude that \[\begin{align} d\hat{f}|_{(x,y)}(U, V) = \left.\frac{d}{dt}\right|_{t = 0} \hat{f}(\gamma(t)) &= \left.\frac{d}{dt}\right|_{t = 0} \hat{f}(\gamma_1(t), \gamma_2(t)) \\&= \left.\frac{d}{dt}\right|_{t = 0} f_1(\gamma_1(t))f_2(\gamma_2(t))\\ &= f_2(y)\left.\frac{d}{dt}\right|_{t = 0} f_1(\gamma_1(t)) + f_1(y)\left.\frac{d}{dt}\right|_{t = 0} f_2(\gamma_2(t)) \\&=f_2(y)df_1|_{x}(U) + f_1(x)df_2|_{y}(V). \end{align}\] Therefore, \[\begin{align} g(\mathrm{grad}_g\, \hat{f}, W) = d\hat{f}(W) &= f_2(y)df_1(U) + f_1(x)df_2(V) \\&= f_2(y)g_1(\mathrm{grad}_{g_1}\,f_1, U) + f_1(x)g_{2}(\mathrm{grad}_{g_2}\,f_2, V) \\&= g\left((f_2(y)\mathrm{grad}_{g_1}\,f_1, f_1(x)\mathrm{grad}_{g_2}\,f_2), (U, V)\right), \end{align}\] and the claim follows. ◻

With the above results in mind, we will now study the critical points of the function \(F\) defined in 43 . In particular, we will prove the following result.

Proposition 43. Let \(F\) be defined as in 43 . Then every critical point of \(F\) corresponds to an eigenvector of \(H\).

Let us first introduce some notation; using \((x,y)\) coordinates on the square lattice \(V\) we can express \(F\) explicitly as \[\begin{align} F(U_{(1,1)}, \dotsc, U_{(n,n)}) &= -\sum_{\langle (i_1, i_2), (j_1, j_2)\rangle \in \Lambda} \bra{0}U_{(i_1, i_2)}^\dagger \sigma_z U_{(i_1, i_2)}\ket{0}\bra{0}U_{(j_1, j_2)}^\dagger \sigma_z U_{(j_1, j_2)}\ket{0} \\&\quad- \sum_{(i_1, i_2) \in V}h_{(i_1, i_2)} \bra{0}U_{(i_1, i_2)}^\dagger \sigma_z U_{(i_1, i_2)}\ket{0}, \end{align}\] where \(h_{(i_1, i_2)}\) is as defined in 44 . To simplify the notation, we will denote \[f_{(i, j)} := \bra{0}U_{(i, j)}^\dagger \sigma_z U_{(i, j)}\ket{0},\] and we will denote the gradient as \(\nabla\) for simplicity. This way, we can rewrite \(F\) as \[F = -\sum_{\langle (i_1, i_2), (j_1, j_2)\rangle \in \Lambda} f_{(i_1, i_2)}f_{(j_1, j_2)}- \sum_{(i_1, i_2) \in V}h_{(i_1, i_1)} f_{(i_1, i_2)}.\]

Note that, using 16, we know that the gradient of \(F\) with respect to \(U_{(i, j)}\) is \[\begin{align} \label{localgradient} \begin{aligned} \nabla_{(i, j)}F &:= -\left(f_{(i-1, j)}+f_{(i+1, j)}\right. \\&\quad\quad\left. +f_{(i, j-1)}+f_{(i, j+1)}+h_{(i, j)}\right)\nabla f_{(i,j)}, \end{aligned} \end{align}\tag{47}\] where some of the summands may be omitted (i.e. when \(i-1=0, j-1=0, i+1>n\) or \(j+1>n\)).

Also recall from 42, that \(\nabla f_{(i, j)}\) only vanishes when \(f_{(i, j)} = \pm 1\), that is, when \(U_{i, j}\ket{0}\) is—up to a phase—\(\ket{0}\) or \(\ket{1}\), respectively.

We are interested in characterizing the critical points of \(F\), i.e, when \[\nabla_{(i, j)} F = 0\quad \forall i,j \in \{1, \dotsc, n\}.\] First of all, note that when \(\ket{U}\) in an eigenvector of \(H\), then it holds that \(\nabla f_{(i, j)} = 0\) for every \((i, j) \in V\) and so it is a critical point of \(F\).

Proof of 43. Assume that we are working on an \(n \times n\) lattice, with \(n \geq 2\). We will prove that every critical point corresponds to an eigenvector of \(H\). To do so, we will consider a critical point and assume that at least one of the \(\nabla f_{(i, j)}\) is non-zero. This will be done in three steps:

  1. We will first assume that \(\nabla f_{(1,1)}\) is non-zero, and reach a contradiction. Therefore, by the symmetry of the problem we can conclude that every critical point must have \[\nabla f_{(1, 1)} = \nabla f_{(1, n)} = \nabla f_{(n, 1)} = \nabla f_{(n, n)} = 0.\]

  2. For the next step, assuming that \(\nabla f_{(1,1)} = \nabla f_{(1, n)} = 0\) we will prove that if \(\nabla f_{(1, k)} \neq 0\) then we cannot be at a critical point.

    Again, by the symmetry of the problem, for the third step we may assume that \(\nabla f_{(1,i)} = \nabla f_{(i, 1)} = 0\) for every \(i = 1, \dotsc, n\).

  3. For the last step of the proof, we will assume that \(\nabla f_{(k, l)} \neq 0\) but \(\nabla f_{(k-1, l)} = \nabla f_{(k, l-1)} = 0\) for some \((k, l)\) with \(k \in \{2, \dots, n-1\}, l \in \{2, n-1\}\) and prove that we cannot be at a critical point.

7.2.1.1 First step: corners.

Assume that \(\nabla f_{(1,1)} \neq 0\). Therefore, by 47 \[f_{(1,2)} + f_{(2,1)} + \varepsilon_{(1,1)} + 2 = 0,\] which cannot happen, as \(-1 \leq f_{(i, j)} \leq 1\) for every \((i, j) \in V\). Therefore, we may now assume that the gradients \(\nabla f_{(i, j)}\) of the corners vanish, and continue with the second step of our proof.

7.2.1.2 Second step: first row.

Assume that, for some \(1 < k < n\), \(\nabla f_{(1,k)} \neq 0\), and \(\nabla f_{(1, k-1)} = 0\). This implies that \[f_{(1,k-1)} + f_{(1,k+1)} + f_{(2,k)} + 3 + \varepsilon_{(1, k)} = 0.\] Since \(\nabla f_{(1,k-1)} = 0\) by assumption, it has to be the case that \(f_{(1, k-1)} = \pm1\), and so \[f_{(1,k+1)} + f_{(2,k)} + 2 + \varepsilon_{(1,k)} = 0,\quad \mathrm{or}\quad f_{(1,k+1)} + f_{(2,k)} + 4 + \varepsilon_{(1,k)} = 0,\] which in either case is a contradiction.

Again, by the symmetry of the problem, we can assume that all of the \(f_{(i,j)}\) laying on the outer layer of our lattice have vanishing gradients.

7.2.1.3 Third step: inner spins.

It only remains to prove the case when \(\nabla f_{(i, j)} \neq 0\) for some \(1 < i, j < n\), and \[\nabla f_{(i-1, j)} = \nabla f_{(i, j-1)} = 0.\] In this case \[f_{(i-1, j)} + f_{(i, j-1)} + f_{(i+1, j)} + f_{(i, j+1)} + 4 + \varepsilon_{(i, j)} = 0,\] which cannot happen, as \[f_{(i-1, j)} + f_{(i, j-1)} \in \{-2, 0, 2\}.\] ◻

Remark 44. Note that the proof of 43 only relies on the magnetic field being greater than the interaction degree at each site. For this reason, by considering a suitable magnetic field for each spin, the proof can be easily adapted to hold for any dimension and any interaction graph.

43 allows us to conclude that the set of critical points of \(F\) can be written as \[\{ (U_1, \dotsc, U_{n^2}) \in \mathrm{SU}(2)^{\times n} : (U_1 \otimes \cdots \otimes U_{n^2})\ket{0}^{\otimes n^2} = e^{i\theta}\ket{x} : x \in \{0, 1\}^{n^2}, \theta \in [0, 2\pi)\}.\] Therefore, if we consider the Riemannian submersion with totally geodesic fibers \(\pi\) from \(\mathrm{SU}(2)^{\times n^2}\) to \((\mathbb{S}^2)^{\times n^2}\), and consider \(\tilde{F}\), the unique function on \((\mathbb{S}^2)^{\times n^2}\) such that \(F = \tilde{F} \circ \pi\), we can conclude that \(\tilde{F}\) has exactly \(2^{n^2}\) critical points, which correspond—via a slight abuse of notation where \(\ket{\psi} \equiv e^{i\theta}\ket{\psi}\) for every vector \(\ket{\psi}\) and any \(\theta \in [0, 2\pi)\)—to the points in \((\mathbb{S}^2)^{\times n^2}\) given by \[\{\ket{x} : x \in \{0, 1\}^{n^2}\}.\] Thus, we can lower bound the distance between any two critical points of \(\tilde{F}\).

Lemma 17. Let \(\tilde{F}\) be the unique function on \(((\mathbb{S}^2)^{\times n^2}, \frac{1}{2}g_{\mathit{round}})\) such that \(F = \tilde{F} \circ \pi\). Then for any two critical points \(x, y\) of \(\tilde{F}\), it holds that \[d_{g_{\mathit{round}}}(x, y) \geq \frac{\sqrt{2}\pi}{2}.\]

Proof. For any two eigenvectors \(\ket{x} = \ket{x_1, \dotsc, x_{n^2}}\), \(\ket{x'} = \ket{x'_1, \dotsc, x'_{n^2}}\) of \(H\) there exists some \(j \in \{1, \dotsc, n^2\}\) such that \(x_j = 1\) and \(x'_j = 0\). Since any point in \(\mathbb{S}^2\) can be written as \[\cos\frac{\theta}{2} \ket{0} + \sin \frac{\theta}{2} e^{i \varphi} \ket{1},\] it is clear that the points in the Bloch sphere \(\mathbb{S}^2\) corresponding to \(x_j\) and \(x'_j\) are antipodal. Therefore, the distance between any two distinct critical points, \(\ket{x}\) and \(\ket{x'}\), when considering the round metric in \(\mathbb{S}^2\) scaled by \(\frac{1}{2}\), is lower bounded by \(\frac{\sqrt{2}}{2}\pi\). Indeed, the distance between two antipodal points in \((\mathbb{S}^2,g_{\mathit{round}})\) is \(\pi\), and when scaling the metric by \(\frac{1}{2}\), the distances get scaled by a factor of \(\frac{1}{\sqrt{2}}\). ◻

Lastly, let us study the Hessian of \(F\) at the critical points.

Lemma 18. Let \(F\) be as defined in 43 . Then \(F\) has a unique minimum, a unique maximum and the rest of the critical points are saddle points. Moreover, for every critical point \(y\) of \(F\), \(\lambda_{\min}\big(\nabla^2 \tilde{F}(y)\big) \leq -\lambda_* < 0\), where \[\lambda_* := \min_{(i, j) \in V} \varepsilon_{(i, j)}.\]

Proof. Let us parametrize \(F\) along a geodesic \[\gamma(t) = (e^{itX_1}U_1, \dotsc, e^{itX_{n^2}}U_{n^2}).\] Taking derivatives, we may conclude that \[\begin{align} \nabla^2 F|_{\overrightarrow{U}}(i\overrightarrow{XU},i\overrightarrow{XU}) &= - \sum_{i, j = 1}^{n^2} \bra{U} (X_i \otimes X_j) H \ket{U} - \sum_{i, j = 1}^{n^2} \bra{U} H (X_i \otimes X_j) \ket{U} \\&\quad + 2\sum_{i, j = 1}^{n^2} \bra{U} X_i H X_j \ket{U}, \end{align}\] where we denote \[\overrightarrow{U} := (U_1, \dotsc, U_{n^2}),\;\ket{U} := U_1 \otimes \dotsm \otimes U_{n^2}\ket{0}^{\otimes n^2}\;\mathrm{and}\;\overrightarrow{XU} := (X_1U_1, X_2U_2, \dotsc, X_{n^2}U_{n^2}),\] for simplicity. Furthermore, if \(\ket{U}\) is an eigenvector of \(H\) with eigenvalue \(E\), then \[\label{eqHessianMultiGate} \left.\nabla^2 F\right|_{\overrightarrow{U}}(i\overrightarrow{XU},i\overrightarrow{XU}) = 2\left(\sum_{i, j = 1}^d \bra{U} X_i H X_j \ket{U} - E\sum_{i, j = 1}^d \bra{U} (X_i \otimes X_j) \ket{U}\right).\tag{48}\]

Writing each \(X_j\) in the eigenbasis \(\{\ket{0}, \ket{1}\}\), we obtain \[X_j = \mathbb{1} \otimes \dotsm \otimes \mathbb{1} \otimes \left(\sum_{k,l = 0}^1 \alpha_{k, l}^{(j)} \ket{k}\bra{l}\right)\otimes \mathbb{1} \otimes \dotsm \otimes \mathbb{1}.\] Moreover, writing \(\ket{U}\) as \[\ket{U} = \ket{l_1} \otimes \ket{l_2} \otimes \dotsm \otimes \ket{l_{n^2}},\] where \(l_i \in \{0, 1\}\) for every \(i \in \{1, \dotsc, n^2\}\), we can conclude that \[\begin{align} X_j \ket{U} &= \ket{l_1} \otimes \overset{j)}{\dotsm} \otimes \left(\sum_{k = 0}^1 \alpha_{k, l_j}^{(j)} \ket{k}\right)\otimes \dotsm \otimes \ket{l_{n^2}} \\&= \sum_{k = 0}^1 \alpha_{k, l_j}^{(j)} \ket{l_1} \otimes \overset{j)}{\dotsm} \otimes \ket{k}\otimes \dotsm \otimes \ket{l_{n^2}}. \end{align}\] If \(i \neq j\) we can write the second term on the right-hand side of 48 as \[\begin{align} \bra{U}\left(X_i \otimes X_j\right) \ket{U}&= \bra{U}X_i X_j \ket{U} \\&= \left(\sum_{k = 0}^1 \overline{\alpha}_{k, l_i}^{(i)} \bra{l_1} \otimes \overset{i)}{\dotsm} \otimes \bra{k}\otimes \dotsm \otimes \bra{l_{n^2}}\right) \\&\quad \times \left(\sum_{k = 0}^1 \alpha_{k, l_j}^{(j)} \ket{l_1} \otimes \overset{j)}{\dotsm} \otimes \ket{k}\otimes \dotsm \otimes \ket{l_{n^2}}\right) \\&= \left(\sum_{k = 0}^1 \overline{\alpha}_{k, l_i}^{(i)}\bra{k}\right)\ket{l_i}\bra{l_j} \left(\sum_{k = 0}^1 \alpha_{k, l_j}^{(j)} \ket{k}\right) \\&= \overline{\alpha}^{(i)}_{l_i, l_i}\alpha^{(j)}_{l_j, l_j}. \end{align}\] If \(i = j\) we can write \[\bra{U} X_i^2 \ket{U} = \sum_{k = 0}^1 |\alpha^{(i)}_{k, l_i}|^2.\]

On the other hand, if \(i \neq j\) we can write the first term on the right-hand side of 48 as \[\begin{align} \bra{U}X_i H X_j\ket{U}&= \bra{U} X_i H \left(\sum_{k = 0}^1 \alpha_{k, l_j}^{(j)} \ket{l_1} \otimes \overset{j)}{\dotsm} \otimes \ket{k}\otimes \dotsm \otimes \ket{l_d}\right) \\&= \bra{U} X_i \left(\sum_{k = 0}^1 \alpha_{k, l_j}^{(j)}E_{l_1, \overset{j)}{\dotsc}, k, \dotsc, l_{n^2}} \ket{l_1} \otimes \overset{j)}{\dotsm} \otimes \ket{k}\otimes \dotsm \otimes \ket{l_d}\right) \\&= \overline{\alpha}^{(i)}_{l_i, l_i}\alpha_{l_j, l_j}^{(j)}E_{l_1, \dotsc, l_j, \dotsc, l_{n^2}}, \end{align}\] where \(E_{l_1, \dotsc, l_{n^2}}\) denotes the eigenvalue of \(H\) associated with the eigenstate \(\ket{l_1, \dotsc, l_{n^2}}\).

If \(i = j\) then \[\begin{align} \bra{U}X_i H X_i\ket{U}&= \bra{U} X_i H \left(\sum_{k = 0}^1 \alpha_{k, l_i}^{(i)} \ket{l_1} \otimes \overset{i)}{\dotsm} \otimes \ket{k}\otimes \dotsm \otimes \ket{l_{n^2}}\right) \\&= \bra{U} X_i \left(\sum_{k = 0}^1 \alpha_{k, l_i}^{(i)}E_{l_1, \overset{i)}{\dotsc}, k, \dotsc, l_{n^2}} \ket{l_1} \otimes \overset{i)}{\dotsm} \otimes \ket{k}\otimes \dotsm \otimes \ket{l_{n^2}}\right) \\&= \sum_{k = 1}^n |\alpha^{(i)}_{k, l_i}|^2 E_{l_1, \overset{i)}{\dotsc}, k, \dotsc, l_{n^2}}. \end{align}\]

Substituting these expressions into 48 we obtain \[\begin{align} \label{expressionhessiano} \begin{aligned} \frac{1}{2}\left.\nabla^2 F\right|_{\overrightarrow{U}}(i\overrightarrow{XU},i\overrightarrow{XU}) &= \sum_{i, j = 1}^{n^2} \bra{U} X_i H X_j \ket{U} - E\sum_{i, j = 1}^{n^2} \bra{U} (X_i \otimes X_j) \ket{U} \\&= \sum_{\substack{i, j = 1\\i \neq j}}^{n^2} \left(\overline{\alpha}^{(i)}_{l_i, l_i}\alpha_{l_j, l_j}^{(j)}E_{l_1, \dotsc, l_j, \dotsc, l_d} - E \overline{\alpha}^{(i)}_{l_i, l_i}\alpha^{(j)}_{l_j, l_j}\right) \\&\quad + \sum_{i = 1}^{n^2} \left(\sum_{k = 1}^n |\alpha^{(i)}_{k, l_i}|^2 E_{l_1, \dotsc, k, \dotsc, l_d} - E \sum_{k = 1}^n |\alpha^{(i)}_{k, l_i}|^2\right) \\&= \sum_{i = 1}^{n^2} \sum_{k = 0}^1 \left(E_{l_1, \overset{i)}{\dotsc}, k, \dotsc, l_d} - E\right) |\alpha^{(i)}_{k, l_i}|^2 \\&= \sum_{i = 1}^{n^2} \left(E_{l_1, \overset{i)}{\dotsc}, l_i + 1 \mathrm{ mod } 2, \dotsc, l_d} - E\right) |\alpha^{(i)}_{l_i + 1 \mathrm{ mod } 2, l_i}|^2. \end{aligned} \end{align}\tag{49}\] This way, if \(E\) corresponds to the lowest (resp. highest) eigenvalue of \(H\), the Hessian of \(F\) is positive-semidefinite (resp. negative-semidefinite).

Observe that in 49 \(E_{l_1, \overset{i)}{\dotsc}, (l_i + 1 \mod 2), \dotsc, l_d}\) corresponds to the eigenvalue obtained when flipping a single spin—in the \(i\)-th position. Thus, if the eigenvector is not that with minimal eigenvalue of \(H\) we can flip at least one spin from \(\ket{1}\) to \(\ket{0}\), in which case the associated eigenvalue decreases at least by \(\min_{I \in V}\varepsilon_I\). ◻

Lastly, we would like to see that the Hessian of \(\tilde{F}\) at the global minimum is positive definite. To do this, we will see which directions \(\overrightarrow{XU}\) lie in the kernel of \(\nabla^2 F\).

Proposition 45. The Hessian of \(\tilde{F}\) at its global minimum is definite-positive.

Proof. We want to conclude that the kernel of \(\nabla^2 F\) at its global minimum is included in the vertical tangent space of \(\mathrm{SU}(2)^{\times n^2}\) with respect to the submersion \(\pi: \mathrm{SU}(2)^{\times n^2} \to (\mathbb{S}^2)^{\times n^2}\). This way, using 78, we will conclude that \(\tilde{F}\) has a positive-definite Hessian at the global minimum.

Let us use the expression for the Hessian shown in 49 and let us denote by \(E\) the smallest (non-degenerate) eigenvalue of \(H\), which corresponds to the eigenvector \(\ket{0}^{\otimes n^2}\). Then at any \(U\) for which \(U\ket{0}^{\otimes n^2} = e^{i\theta} \ket{0}^{\otimes n^2}\), it holds that \[\frac{1}{2}\left.\nabla^2 F\right|_{\overrightarrow{U}}(i\overrightarrow{XU},i\overrightarrow{XU}) = \sum_{i = 1}^{n^2} \left(E_{0, \overset{i)}{\dotsc}, 1, \dotsc, 0} - E\right) |\alpha^{(i)}_{1, 0}|^2,\] where \(E_{0, \overset{i)}{\dotsc}, 1, \dotsc, 0}\) denotes the eigenvalue associated with \(|0, \overset{i)}{\dotsc}, 1, \dotsc, 0\rangle\) and the terms \(\alpha^{(i)}_{j,k}\) are the coordinates of the \(i\)-th component of the vector \(\overrightarrow{X}\).

Since \(E\) is the unique lowest eigenvalue of \(H\), the only directions \(i\overrightarrow{XU}\) that lie in the kernel of \(\left.\nabla^2 F\right|_{\overrightarrow{U}}\) correspond to those such that \(\alpha^{(i)}_{k, li} = 0\) for every \(i\) and every \(k \neq l_i\). Otherwise, we would have a non-zero term in the sum: \(E_{l_1, \overset{i)}{\dotsc}, k, \dotsc, l_n} - E > 0\).

For this reason, every \(X_j\) is of the form \[X_j = \mathbb{1} \otimes \dotsm \otimes \mathbb{1} \otimes \left(\alpha_{0, 0}^{(j)} \ketbra{0} - \alpha_{1, 1}^{(j)} \ketbra{1}\right)\otimes \mathbb{1} \otimes \dotsm \otimes \mathbb{1},\] and so it can be written as \[X_j = \mathbb{1} \otimes \dotsm \otimes \mathbb{1} \otimes \begin{pmatrix} ia & 0\\ 0 & -ia \end{pmatrix}\otimes \mathbb{1} \otimes \dotsm \otimes \mathbb{1},\] for some \(a\). These directions are vertical with respect to the action shown in 45 , and so the Hessian of \(\tilde{F}\) at its global minimum is positive definite. ◻

7.2.2 Rapid mixing and suboptimality for the Gibbs distributions↩︎

Let us conclude this section by applying 30 31 4 to our setting. In particular, we will study the function \(F\) defined on the manifold \(\mathrm{SU}(2)^{\times n^2}\) endowed with the metric \(g\) which corresponds to the product of the bi-invariant metric of \(\mathrm{SU}(2)\) and its projection \(\tilde{F}\) on \((\mathbb{S}^2)^{\times n^2}\).

Proposition 46. Let \(F: \mathrm{SU}(2)^{\times n^2} \to \mathbb{R}\) be defined as in 43 and let \(\tilde{F}: (\mathbb{S}^2)^{\times n^2} \to \mathbb{R}\) be the unique function such that \(F = \tilde{F} \circ \pi\), where \(\pi\) is the Riemannian submersion \[\pi:(\mathrm{SU}(2)^{\times n^2}, g) \to ((\mathbb{S}^2)^{\times n^2}, h),\] \(g\) corresponds to the product of the bi-invariant metric of \(\mathrm{SU}(2)\), and \(h\) corresponds to the product of the metric \(\frac{1}{2}g_{\mathit{round}}\), where \(g_{\mathit{round}}\) denotes the usual round metric on the sphere. For every \(\varepsilon \in (0, 1]\) and \(\delta \in (0, 1)\), if \(\beta\) is such that \[\beta \geq \frac{2}{\varepsilon}\bigg(\frac{1}{2} + 3^3n^4 \log (3n^2) + 9 n^6 \log(2^{5/2} \pi^2) + 9n^4 \log \frac{A_2\sqrt{2\pi}}{\delta\varepsilon}\bigg),\] the Gibbs distribution \(\nu_F(x) = \frac{1}{Z_F} e^{-\beta F(x)}\) on \(\mathrm{SU}(2)^{\times n^2}\) is such that \[\nu_F\Big(F - \min_{y \in \mathrm{SU}(2)^{\times n^2}} F(y) \geq \varepsilon\Big) \leq \delta.\] Now, if we define \[\varepsilon_{\max} := \min \Big\{\frac{\pi^2 A_2}{16}, 1\Big\},\] then, for every \(\varepsilon \in (0, \varepsilon_{\max}]\) and \(\delta \in (0, 1)\), if \(\beta\) satisfies \[\beta \geq \frac{2}{\varepsilon}\bigg(\frac{1}{2} + 12n^4 \log (2n^2) + 4n^6 \log \frac{\pi^{3/2}}{\Gamma(3/2)} + 4n^4 \log \frac{A_2\sqrt{2\pi}}{\delta\varepsilon}\bigg),\] the Gibbs distribution \(\nu_{\tilde{F}}(x) = \frac{1}{Z_{\tilde{F}}} e^{-\beta \tilde{F}(x)}\) on \((\mathbb{S}^{2})^{\times n^2}\) is such that \[\nu_{\tilde{F}}\Big(\tilde{F} - \min_{y \in (\mathbb{S}^{2})^{\times n^2}} \tilde{F}(y) \geq \varepsilon\Big) \leq \delta.\]

Before we present the proof of 46, note that \(((\mathbb{S}^2)^{\times n^2}, h)\) is assumed to be endowed with the product of the round metric scaled by \(\frac{1}{2}\), i.e. \[((\mathbb{S}^2)^{\times n^2}, h) \simeq (\mathbb{S}^2, \frac{1}{2}g_{\mathit{round}})^{\times n^2}.\] In contrast, in 8.6 we studied the sphere endowed with the round metric. Nevertheless, the results derived in the aforementioned section can be easily adapted to hold for the scaled sphere. Indeed, given some \(n\)-dimensional Riemannian manifold \((M, g)\), if we consider the same manifold \(M\) endowed with the scaled metric \(\lambda g\), for some \(\lambda > 0\), the Ricci curvature of the scaled manifold stays the same as the original
([39]), while the distances get scaled by \(\sqrt{\lambda}\) and the volume form gets scaled by \(\lambda^{n/2}\).

Proof of 46. To prove the suboptimality result for the Gibbs measures \(\nu_F\) and \(\nu_{\tilde{F}}\), we need to bound the injectivity radii and the volume of \((\mathrm{SU}(2)^{\times n^2}, g)\) and \(((\mathbb{S}^{2})^{\times n^2}, h)\). First, as we saw in 41 that \[(\mathrm{SU}(2), g_{\mathit{bi}}) \simeq (\mathbb{S}^3, 2g_{\mathit{round}}),\] where \(g_{\mathit{bi}}\) denotes the bi-invariant metric of \(\mathrm{SU}(2)\), we know that \[\mathrm{Vol}(\mathrm{SU}(2)) = 2^{3/2} \mathrm{Vol}(\mathbb{S}^3) = 2^{3/2} \frac{2\pi^2}{\Gamma(2)} = 2^{5/2} \pi^2,\] where we used the expression for the volume of the sphere endowed with the round metric seen in 8.6.3, and so \[\mathrm{Vol}(\mathrm{SU}(2)^{\times n^2}) = (2^{5/2} \pi^2)^{n^2}.\] Also \[\mathrm{Vol}\Big(\Big(\mathbb{S}^2, \frac{1}{2}g_{\mathit{round}}\Big)\Big) = \frac{1}{2}\mathrm{Vol}(\mathbb{S}^2) = \frac{\pi^{3/2}}{\Gamma(3/2)},\] and so \[\mathrm{Vol}(((\mathbb{S}^2)^{\times n^2}, h))) = \Big(\frac{\pi^{3/2}}{\Gamma(3/2)}\Big)^{n^2}.\] For the injectivity radii, we simply use the scaling property of the distance, along with the bounds obtained in 7 to conclude that \[i(\mathrm{SU}(2)^{\times n^2}, g) = i(\mathrm{SU}(2), g_{\mathit{bi}}) = i(\mathbb{S}^3, 2g_{\mathit{round}}) = \sqrt{2}i(\mathbb{S}^3) \geq \sqrt{2}\pi.\] \[i(((\mathbb{S}^2)^{\times n^2}, h)) = i(\mathbb{S}^2, \frac{1}{2}g_{\mathit{round}}) = \frac{\sqrt{2}}{2}i(\mathbb{S}^2) \geq \frac{\sqrt{2}\pi}{2}.\] Substituting these bounds into 31, the suboptimality for the Gibbs measures follows. ◻

Now, let us restate the assumption needed for 30 to hold in this case.

Assumption 23. There exists a constant \(0 < C_{\tilde{F}} \leq 1\) such that \[|\mathrm{grad}_{h}\, \tilde{F}(x)|_h \geq C_{\tilde{F}} d_h(x, \tilde{\mathcal{C}}),\] for every \(x \in (\mathbb{S}^2)^{\times n^2}\), where \(\tilde{\mathcal{C}}\) denotes the set of critical points of \(\tilde{F}\).

Proposition 47. Let \(F: \mathrm{SU}(2)^{\times n^2} \to \mathbb{R}\) be defined as in 43 and let \(\tilde{F}: (\mathbb{S}^2)^{\times n^2} \to \mathbb{R}\) be the unique function such that \(F = \tilde{F} \circ \pi\), where \(\pi\) is the Riemannian submersion \[\pi:(\mathrm{SU}(2)^{\times n^2}, g) \to ((\mathbb{S}^2)^{\times n^2}, h),\] \(g\) corresponds to the product of the bi-invariant metric of \(\mathrm{SU}(2)\), and \(h\) corresponds to the product of the metric \(\frac{1}{2}g_{\mathit{round}}\), where \(g_{\mathit{round}}\) denotes the usual round metric on the sphere. Assume that \(\tilde{F}\) satisfies 23, and let \[\lambda_* := \min_{I \in V} \varepsilon_I,\] (cf. 44 ). Furthermore, let \(a, \beta > 0\) be such that \[\begin{align} a^2 &\geq \max\Big\{\frac{48 A_2 n^2}{C^2_{\tilde{F}}}, \frac{544}{\lambda_{*}}\Big\}, \\\beta &\geq \max\Bigg\{\frac{72^2 3^5 n^{10} A_2 A^2_3 a^6}{\lambda_{*}^2}, \frac{8a^2}{\pi^2}\Bigg\}, \end{align}\] where \(A_2\) and \(A_3\) are the Lipschitz constants of the gradient and the Hessian of \(\tilde{F}\), respectively. Let \(\operatorname{L}_F\) and \(\operatorname{L}_{\tilde{F}}\) be defined as \[\operatorname{L}_F := -\mathrm{grad}_g\, F + \frac{1}{\beta}\Delta_g,\quad \mathrm{and}\quad \operatorname{L}_{\tilde{F}} := -\mathrm{grad}_h\, \tilde{F} + \frac{1}{\beta} \Delta_h.\] Then the Markov triple \((\mathrm{SU}(2)^{\times n^2}, \nu_F, \Gamma_g)\) associated with \(\operatorname{L}_F\) satisfies an \(\mathrm{LSI}(\alpha)\), where \[\frac{1}{\alpha} = 16\beta A_2 \pi^2 n^2\max \Big\{\frac{184}{\lambda_*}, n^2 (1 + \sqrt{2})^2 \beta\Big\},\] and the Markov triple \(((\mathbb{S}^2)^{\times n^2}, \nu_{\tilde{F}}, \Gamma_h)\) associated with \(\operatorname{L}_{\tilde{F}}\) satisfies an \(\mathrm{LSI}(\tilde{\alpha})\), where \[\frac{1}{\tilde{\alpha}} = \frac{368\beta A_2 n^2 \pi^2}{\lambda_*}.\]

Proof. In order to obtain the log-Sobolev inequalities, recall from 43 and 17 that the critical points of \(\tilde{F}\) are isolated and separated by a distance of at least \(\frac{\sqrt{2}\pi}{2}\). Moreover, we proved in 18 45 that 8 9 hold with \(\lambda_*\) given by \[\min_{I \in V} \varepsilon_I.\]

The Ricci curvatures of both \((\mathrm{SU}(2)^{\times n^2}, g)\) and \(((\mathbb{S}^2)^{\times n^2}, h)\) are non-negative (cf. 58 56 82). Furthermore, it follows by 57 58 that the coefficients of the Riemann curvature tensor of \((\mathrm{SU}(2)^{\times n^2}, g)\) in normal coordinates when endowed with the product of bi-invariant metrics are bounded by \(\frac{1}{2}\).

Also, it holds that the Riemannian submersion \(\pi_1 : (\mathrm{SU}(2)^{\times n^2}, g) \to ((\mathbb{S}^2)^{\times n^2}, h)\) has totally geodesic fibers, which means by 21 22 that the sectional curvature of the fibers coincides with that of the total space. Since \((\mathrm{SU}(2)^{\times n^2}, g)\) has non-negative sectional curvature by 55 [curvatureproduct], we conclude that the sectional curvature of the fibers is also non-negative and so their Ricci curvature is non-negative as well.

Now, the result follows by substituting the manifold-dependent constants in the statement of 30 with the known bounds found in 8. In particular, the dimensions of \(\mathrm{SU}(2)^{\times n^2}\) and \((\mathbb{S}^2)^{\times n^2}\) are known to be \(3n^2\) and \(2n^2\), respectively. Also, using 86, 8 and the scaling property of the distances, we know that \[\begin{align} \mathrm{diam}(((\mathbb{S}^2)^{\times n^2}, h)) &= n\, \mathrm{diam}(\mathbb{S}^2, \frac{1}{2}g_{\mathit{round}}) = n\frac{\sqrt{2}\pi}{2},\\ \mathrm{diam}(\mathrm{SU}(2)^{\times n^2}, g) &= n\, \mathrm{diam}(\mathrm{SU}(2), g_{\mathit{bi}}) \leq 2\pi n. \end{align}\] Lastly, the diameter of the fibers with the induced metric can be bounded using 86 and 9 as \[\mathrm{diam}(\mathrm{U}(1)^{\times n^2}) = n\, \mathrm{diam}(\mathrm{U}(1)) \leq n\pi(1 + \sqrt{2}),\] and by 7 and the scaling of the metric, the convexity radius of \(((\mathbb{S}^2)^{\times n^2}, g)\) is lower bounded by \[\mathit{conv}(((\mathbb{S}^2)^{\times n^2}, g)) = \mathit{conv}((\mathbb{S}^2, \frac{1}{2}g_{\mathit{round}})) = \frac{\sqrt{2}}{2}\mathit{conv}((\mathbb{S}^2, g_{\mathit{round}})) \geq \frac{\pi\sqrt{2}}{4},\] so the result follows. ◻

Lastly, let us restate 4.

Corollary 6. Let \(X_t\) and \(\tilde{X}_t\) be the Langevin diffusion processes generated by \(\operatorname{L}_F\) and \(\operatorname{L}_{\tilde{F}}\), with initial uniform distribution on \(\mathrm{SU}(2)^{\times n^2}\) and \((\mathbb{S}^2)^{\times n^2}\), respectively. Under the assumptions of 47, they converge exponentially fast to their associated Gibbs measures \(\nu_F\), and \(\nu_{\tilde{F}}\) i.e. \[\begin{align} \norm{\nu_F - \rho_t}^2_{\mathrm{TV}} &\leq \beta\, e^{-2\alpha t} \max_{y \in \mathrm{SU}(2)^{\times n^2}} F(y),\\ \norm{\nu_{\tilde{F}} - \tilde{\rho}_t}^2_{\mathrm{TV}} &\leq \beta\, e^{-2\tilde{\alpha} t} \max_{y \in (\mathbb{S}^2)^{\times n^2}} f(y) = \beta\, e^{-2\tilde{\alpha} t} \max_{y \in \mathrm{SU}(2)^{\times n^2}} F(y), \end{align}\] for every \(t \geq 0\), where \(\rho_t\) and \(\tilde{\rho}_t\) denote the distribution of \(X_t\) and \(\tilde{X}_t\), and \(\alpha\) and \(\tilde{\alpha}\) denote the log-Sobolev inequality constants of \((\mathrm{SU}(2)^{\times n^2}, \nu_F, \Gamma_g)\), and \(((\mathbb{S}^2)^{\times n^2}, \nu_{\tilde{F}}, \Gamma_h)\).

Acknowledgements↩︎

P. P. V. would like to thank Pablo Hidalgo Palencia for many fruitful conversations throughout the development of this work, and Tomasz Kania for generously providing the argument for the escaping time estimates of the generalized CIR process (see 102). A. C. acknowledges the support of the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - Project-ID 470903074 - TRR 352. This project was funded within the QuantERA II Programme which has received funding from the EU’s H2020 research and innovation programme under the GA No 101017733. M. C. L. was supported by Project PID2024-156578NB-I00 funded by MICIU /AEI /10.13039/501100011033 / FEDER, EU. S. I. acknowledges financial support from the “Center of Excellence Maria de Maeztu 2025-2029” award to the Institute of Cosmos Sciences, grant CEX2024-001451-M, funded by MICIU/AEI/10.13039/501100011033. A. L. acknowledges support from the Italian Ministry of University and Research (MUR), through “Programma per Giovani Ricercatori Rita Levi Montalcini”, the grant “Dipartimento di Eccellenza 2023-2027” of Dipartimento di Matematica, Politecnico di Milano, as well as the National Group of Mathematical Physics (GNFM) of the Italian Institute for High Mathematics (INdAM). P. P.  V. acknowledges support of the Spanish Ministry of Science and Innovation MCIN/AEI/10.13039/501100011033 (CEX2023-001347-S,CEX2019-000904-S, CEX2019-000904-S-21-2). This work has been funded by the Spanish Ministry of Science, Innovation and Universities MICIU/AEI/10.13039/501100011033 (CEX2023-001347-S, PID2023-146758NB- I00), Comunidad de Madrid (TEC-2024/COM-84-QUITEMAD-CM), Universidad Complutense de Madrid (FEI-EU-22-06), and the Ministry for Digital Transformation and of Civil Service of the Spanish Government through the QUANTUM ENIA project call - Quantum Spain project, and by the European Union through the Recovery, Transformation and Resilience Plan - NextGenerationEU within the framework of the Digital Spain 2026 Agenda.

8 Background in Riemannian geometry↩︎

This section gathers most of the results from Riemannian geometry that are needed in this work. In particular, we will give a general introduction to Lie groups, and analyze the curvature of the unitary and special unitary groups. We will also introduce Riemannian submersions, and obtain results regarding the connection between the curvature of their domain and their image, focusing on the particular case of Stiefel and Grassmann manifolds. We will conclude the section with a study of fiber-invariant functions in the context of Riemannian submersions, and a set of auxiliary results regarding some properties about the sphere and the aforementioned manifolds, that are used throughout the text.

8.1 Lie groups↩︎

We begin by introducing some of the most elementary definitions and properties of Lie groups. Lie groups can be both studied from the point of view of algebra and differential/Riemannian geometry. We will focus on the latter approach—see [59] for the algebraic viewpoint.

Definition 14 (Lie group). A Lie group is a group \(G\) that can be endowed with a smooth (real) manifold structure in such a way that both the group operation and the inverse map are smooth.

In this work, we will restrict our attention to a specific kind of Lie groups, namely linear Lie groups, which are closed subsets of \(\mathrm{GL}(n)\). In fact, the only Lie groups that we consider in this text are the unitary group \(\mathrm{U}(n)\) and special unitary group \(\mathrm{SU}(n)\). As we mentioned earlier, this section is only intended to be a short summary of the main definitions and results used in 7. For this reason, we have decided not to include any proofs. For a more general and detailed introduction to Lie groups, see for example [60].

To each Lie group \(G\) we can associate a real vector space, known as its Lie algebra \(\mathfrak{g}\), which in the case of linear Lie groups is given by \[\mathfrak{g} := \{X \in \mathrm{M}_n(\mathbb{C}) : e^{tX} \in G \mathrm{ \textit{for all} } t \in \mathbb{R}\}.\]

Proposition 48. For the Lie groups \(\mathrm{U}(n)\) and \(\mathrm{SU}(n)\), their associated Lie algebras \(\mathfrak{u}(n)\) and \(\mathfrak{su}(n)\) are given by the following sets: \[\mathfrak{u}(n) = \{ X \in \mathrm{M}_n(\mathbb{C}) : X^\dagger = - X\},\] and \[\require{physics} \mathfrak{su}(n) = \{ X \in \mathrm{M}_n(\mathbb{C}) : X^\dagger = - X,\;\Tr(X) = 0\}.\]

The vector space \(\mathfrak{g}\) is called a Lie algebra because it is endowed with a bilinear map, known as the Lie bracket.

Definition 15 (Lie bracket). Given a Lie algebra \(\mathfrak{g}\), a Lie bracket is a bilinear map \[[\cdot, \cdot] : \mathfrak{g} \times \mathfrak{g} \to \mathfrak{g},\] satisfying

  • \([u, u] = 0\) for every \(u \in \mathfrak{g}\).

  • \([u, [v, w]] + [w, [u, v]] + [v, [w,u ]] = 0\), for every \(u, v, w \in \mathfrak{g}\). This is known as the Jacobi identity, and in particular implies that \([u, v] = -[v, u]\) for every \(u, v \in \mathfrak{g}\).

In the case of linear Lie groups, the Lie bracket associated with their Lie algebra is given by the matrix commutator operator, i.e. \([u, v] := uv - vu\), for every \(u, v \in \mathfrak{g}\).

Bear in mind that the notation \([\cdot, \cdot]\) is also used for the Lie bracket of vector fields on a general manifold \(M\), not necessarily a Lie group. Both uses of the same notation are instances of the same structure; indeed, the space of vector fields on \(M\) forms a Lie algebra—the algebra of the diffeomorphism group of \(M\)—endowed with the Lie bracket of vector fields. Nevertheless, the two can be distinguished by the context: while the Lie bracket of \(\mathfrak{g}\) is a map defined on \(\mathfrak{g} \times \mathfrak{g}\), the Lie bracket of vector fields is defined on \(\mathfrak{X}(M) \times \mathfrak{X}(M)\). The Lie bracket of vector fields is also known as the Lie derivative.

One essential property of Lie groups is that their Lie algebra \(\mathfrak{g}\) can be identified with the tangent space at the identity, \(\mathfrak{g} \equiv T_e G\) (cf. [60]). Furthermore, for each element \(p \in G\), \[T_p G = p\mathfrak{g} = \mathfrak{g}p.\]

Lie groups also give us a very convenient way to define vector fields on them, given a tangent vector at some fixed point. Indeed, given a Lie group \(G\) and a vector \(v \in \mathfrak{g}\), we can define the following vector fields \(v^L, v^R \in \mathfrak{X}(G)\) in the following way: \[v^L(a) = av,\quad v^R(a) = va,\quad \forall a \in G.\] Such vector fields are known as left and right-invariant vector fields, respectively.

Invariant vector fields have various useful properties; one of which appears in the following proposition (cf. [60]).

Proposition 49. The Lie derivative of two left-invariant (resp. right-invariant) vector fields is again left-invariant (resp. right-invariant).

In fact, the Lie derivative of two invariant vector fields is related to the Lie bracket of the Lie algebra as follows:

Proposition 50 ([60]). Let \(G\) be a linear Lie group with unit \(e\), and let \(u, v \in \mathfrak{g}\). Then \[[u^L, v^L](e) = uv - vu\] and \[[u^R, v^R](e) = vu - uv.\]

Since Lie groups are manifolds, they can be endowed with a Riemannian metric. In particular, some Lie groups can be endowed with what is known as a bi-invariant metric.

Definition 16 (Bi-invariant metric). Given a linear Lie group \(G\) endowed with a Riemannian metric \(g\), we say that \(g\) is bi-invariant if, for every \(a, b \in G\) and every \(X, Y \in T_b G\) \[g(X, Y) = g(aX, aY),\] and \[g(X, Y) = g(Xa, Ya).\]

Note that, in the above definition, if \(X, Y \in T_b G\), then \(aX, aY \in T_{ab} G\) and \(Xa, Ya \in T_{ba} G\).

Remark 51. We can define a bi-invariant metric on \(G = \mathrm{U}(n)\) or \(\mathrm{SU}(n)\) as follows; for every \(X, Y \in \mathfrak{g}\), let \[\require{physics} g(X, Y) = \Tr(Y^\dagger X).\] Now, we can extend the metric to \(TG\); let \(pX, pY \in T_p G\), with \(X, Y \in \mathfrak{g}\), we define \[\require{physics} g(pX, pY) := \Tr(Y^\dagger p^\dagger p X) = \Tr(Y^\dagger X) = g(X, Y),\] where the equality follows from \(p\) being unitary.

We will refer to \(\require{physics} g(X, Y) = \Tr (Y^\dagger X)\) as the bi-invariant metric of \(\mathrm{SU}(n)\) and \(\mathrm{U}(n)\). Working with this metric allows us to characterize all the geodesics of \(\mathrm{SU}(n)\) and \(\mathrm{U}(n)\) (cf. [60]).

Proposition 52. Consider \(G = \mathrm{SU}(n)\) or \(\mathrm{U}(n)\) endowed with the bi-invariant metric shown in 51. For every \(p \in G\), the geodesic passing through \(p\) in the direction of \(Xp \in T_p G\), \(X \in \mathfrak{g}\) is of the form \[\gamma(t) = e^{tX}p.\]

One key property of Lie groups endowed with a bi-invariant metric is that they are symmetric spaces. This fact is used several times in [SectionPI] [sec:LaplacianAndGradientBound] and allows us to obtain an explicit expression for the Taylor series of the metric in normal coordinates.

Definition 17. A connected Riemannian manifold \((M, g)\) is a symmetric space if for each point \(p \in M\) there exists a point reflexion, i.e. an isometry \(\varphi : M \to M\) that fixes \(p\) and such that \(d\varphi|_p : T_p M \to T_p M\) is equal to \(-\mathrm{Id}\).

As noted previously, it is known that every connected Lie group endowed with a bi-invariant metric is a symmetric space. Furthermore, the round sphere, and Grassmann manifolds endowed with their induced metric are symmetric spaces too (c.f. [39]).

8.1.1 Curvature of the unitary and special unitary groups↩︎

Let us now study the curvature of \(\mathrm{U}(n)\) and \(\mathrm{SU}(n)\) when endowed with the bi-invariant metric. Since \(\mathrm{U}(n)\) and \(\mathrm{SU}(n)\) are linear Lie groups, their curvature can be studied by inspecting what is known as the Killing form.

Definition 18 (Killing form). Given a Lie algebra \(\mathfrak{g}\) of a linear Lie group \(G\), the Killing form of \(G\) is a symmetric bilinear form \(\mathfrak{B}: \mathfrak{g} \times \mathfrak{g} \to \mathbb{C}\) given by \[\require{physics} \mathfrak{B}(u, v) := \Tr([u, [v, \cdot]]),\] where \([\cdot, \cdot]\) denotes the Lie bracket associated with the Lie algebra of \(G\), which is given by the commutator.

The Killing form can in fact be defined for any Lie algebra \(\mathfrak{g}\), not just for Lie algebras associated with linear Lie groups. For further details, see [60].

Proposition 53 ([60]). For any Lie group \(G\) equipped with a bi-invariant metric \(g\), the following hold:

  1. For all pairs of orthonormal vectors \(u, v \in \mathfrak{g}\), \[K_g(u,v) = \frac{1}{4}|[u,v]|^2_g.\]

  2. For all \(u, v, w \in \mathfrak{g}\), \[R(u, v)w = \frac{1}{4}[[u, v], w].\]

  3. For all \(u, v \in \mathfrak{g}\), \[\mathrm{Ric}_g(u,v) = -\frac{1}{4}\mathfrak{B}(u,v),\] where \(\mathfrak{B}\) is the Killing form of \(G\), and \(K_g\), \(R\) and \(\mathrm{Ric}_g\) denote the sectional curvature, and the Riemann and Ricci curvature tensors of \((G, g)\), respectively.

This proposition is very useful in our context, as there exists an explicit formula for the Killing form of \(\mathrm{U}(n)\) and \(\mathrm{SU}(n)\).

Remark 54 ([60]). The Killing form of \(\mathrm{U}(n)\) is \[\require{physics} \mathfrak{B}(X,Y) = 2n\Tr(XY) - 2\Tr(X) \Tr(Y), \quad X, Y \in \mathfrak{u}(n).\] Furthermore, the Killing form of \(\mathrm{SU}(n)\) is \[\require{physics} \mathfrak{B}(X, Y)= 2n\Tr(XY),\quad X, Y \in \mathfrak{su}(n).\]

The expressions shown in 54 allow us to bound the curvatures of \(\mathrm{U}(n)\) and \(\mathrm{SU}(n)\) when endowed with their bi-invariant metric.

Proposition 55. Let \(\mathrm{U}(n)\) be endowed with the bi-invariant metric \(g\) from 51. Then \[\mathrm{Ric}_{g} \geq 0,\quad \mathrm{\textit{and}}\quad 0 \leq K_{g} \leq \frac{1}{2}.\]

Proof. By 53 54 it holds that \[\require{physics} \mathrm{Ric}_{g}(X, X) = -\frac{1}{2}\left(n\Tr(X^2) - \Tr(X)^2\right) \geq 0,\] for any \(p \in \mathrm{U}(n)\) and any \(X \in T_p \mathrm{U}(n)\).

Now, given two orthonormal vectors \(X,Y \in T_p \mathrm{U}(n)\), using 53 we obtain \[\begin{align} 0\leq K_{g}(X, Y) &= \frac{1}{4} \left|[X, Y]\right|_g^2. \end{align}\] Furthermore, it is known [61] that \[\label{ineqliebracket} \left|[X, Y]\right|_g \leq 2 |X|_g^2 |Y|_g^2,\tag{50}\] which allows us to conclude that \[0\leq K_{g}(X, Y) \leq \frac{2}{4} \left|X\right|_g^2 \left|Y\right|_g^2 \leq \frac{1}{2}.\] ◻

Proposition 56. Consider \(\mathrm{SU}(n)\) endowed with the bi-invariant metric \(g\) from 51. Then \[\mathrm{Ric}_{g} = \frac{n}{2}g,\quad \mathrm{\textit{and}}\quad 0 \leq K_{g} \leq \frac{1}{2}.\]

Proof. The bound of the sectional curvature can be proven following the same steps as in 55.

In order to obtain the equation for the Ricci curvature, note that by 53 54 \[\require{physics} \mathrm{Ric}_{g}(X, Y) = -\frac{2}{4}n\Tr(XY) = \frac{n}{2}g(X, Y),\] for every \(p \in \mathrm{SU}(n)\) and every pair \(X, Y \in T_p \mathrm{SU}(n)\). ◻

Lastly, let us show that, when endowing \(\mathrm{U}(n)\) or \(\mathrm{SU}(n)\) with the bi-invariant metric, the Riemannn curvature tensor can be easily bounded.

Proposition 57. Let \(G = \mathrm{SU}(n)\) or \(\mathrm{U}(n)\) be endowed with the bi-invariant metric \(g\). Then, for every \(x \in G\), considering normal coordinates centered at \(x\) it follows that \[|R_{ijk}^l(x)| = |R_{ijkl} (x)| \leq \frac{1}{2},\] for every \(i, j, k, l \in \{1, \dotsc, d\}\), where \(d = n^2\) if \(G = \mathrm{U}(n)\) and \(d = n^2-1\) otherwise.

Proof. Let \(x \in G\). Taking normal coordinates centered at \(x\) (cf. 60), we know that \[|R_{ijk}^l(x)| = |\delta^{la}R_{ijka}(x)|,\] for every \(i, j, k, l \in \{1, \dotsc, d\}\). Now, using 53 we know that \[R(u, v)w = \frac{1}{4}[[u, v],w],\] for every \(u, v, w \in \mathfrak{g}\). Let \(\{e_n\}_{n = 1}^d\) be an orthonormal basis of \(T_x G\), and consider four vectors \(e_i, e_j, e_k, e_l \in \{e_n\}_{n = 1}^d\). Then \[R_{ijkl}(x) = \frac{1}{4}g([[e_i, e_j],e_k], e_l) \leq \frac{1}{4} |[[e_i, e_j],e_k]|_g |e_l|_g.\] Now, we can apply 50 along with the fact that the vectors have norm one, to conclude that \[|R_{ijk}^l(x)| = |\delta^{la}R_{ijka}(x)| \leq \frac{\sqrt{2}}{4} |[e_i, e_j]|_g |e_k|_g \leq \frac{1}{2}.\] ◻

8.2 Curvature of product manifolds↩︎

Let us now study the curvature of product manifolds and how it is related to that of each factor. The usual metric considered for the product of two Riemannian manifolds is known as the product metric, which is given as the direct sum of the metrics each manifold is endowed with. This is the metric considered in the manifolds appearing in 7.

Before we show how to bound the curvature of the product metric, let us present the following two auxiliary lemmas.

Lemma 19. Let \((M, g)\) and \((N, h)\) be two Riemannian manifolds and let \(M \times N\) be the product manifold endowed with the product metric \(g \oplus h\). Let \(X, Y \in \mathfrak{X}(M\times N)\) be two decomposable vector fields, i.e. \(X = (X_1, X_2)\), and \(Y = (Y_1, Y_2)\), with \(X_1, Y_1 \in \mathfrak{X}(M)\), \(X_2, Y_2 \in \mathfrak{X}(N)\). Then \((g \oplus h)(X, Y)\) verifies \[(g \oplus h)(X, Y) = g(X_1, Y_1) + h(X_2, Y_2),\] and \[[X, Y]_{M \times N} = [X_1, Y_1]_M + [X_2, Y_2]_N,\] where \([\cdot, \cdot]_{M \times N}\), \([\cdot, \cdot]_{M}\), and \([\cdot, \cdot]_{N}\) denote the Lie derivative in \(M\times N\), \(M\) and \(N\), respectively, and the sum is written in the sense of \(T_{(x_1, x_2)} (M\times N) = T_{x_1} M \oplus T_{x_2} N\).

Furthermore, (c.f. [62]), using the same sum notation for \(T_{(x_1, x_2)} (M\times N)\), \[\nabla_X Y = \nabla^M_{X_1} Y_1 + \nabla^N_{X_2} Y_2,\] where \(\nabla\), \(\nabla^M\) and \(\nabla^N\) denote the Levi-Civita connection of \(M \times N\), \(M\) and \(N\), respectively.

Using 19 it is straightforward to see how the Riemann curvature tensor of a product is related to those of its factors.

Lemma 20. Let \((M, g)\) and \((N, h)\) be two Riemannian manifolds and let \(M \times N\) be the product manifold endowed with the product metric. Then for every point \((x, y) \in M \times N\) and every four tangent vectors \(X, Y, Z, W \in T_{(x, y)} M \times N\), \[R_{g \oplus h}(X, Y, Z, W) = R_{g}(X_1, Y_1, Z_1, W_1) + R_{h}(X_2, Y_2, Z_2, W_2),\] where \(R_{g \oplus h}\) denotes the Riemann curvature tensor of \(M \times N\), and \(R_{g}\), \(R_{h}\) denote the Riemann curvature tensors of \(M\) and \(N\) respectively. We have used the same notation as in 19 for the decomposition of the tangent vectors.

With the two above lemmas in mind, it is now straightforward to prove the following statement.

Proposition 58 (Curvature of the product). Let \((M, g)\), \((N, h)\) be two Riemannian manifolds of dimension \(d_M\) and \(d_N\), respectively, verifying \[K_g \leq D_1,\quad K_h \leq D_2,\] \[\mathrm{Ric}_g \geq \delta_1,\quad \mathrm{Ric}_h \geq \delta_2,\] for some constants \(D_1, D_2, \delta_1, \delta_2 \in \mathbb{R}\), where \(K_g\) (resp. \(K_h\)) denotes the sectional curvature of \(M\) (resp. \(N\)) and \(\mathrm{Ric}_g\) (resp. \(\mathrm{Ric}_h\)) denotes the Ricci curvature of \(M\) (resp. \(N\)). Assume that for every \(x \in M\) and every \(y \in N\), the coefficients of the Riemann curvature tensors of \(M\) and \(N\), denoted as \(R_g\) and \(R_h\) respectively, in normal coordinates centered at \(x\) and \(y\) are such that \[|(R_g)_{ijkl}(x)| \leq \Lambda_1, \quad |(R_h)_{mnpq}(y)| \leq \Lambda_2,\] for every \(i,j,k,l \in \{1, \dotsc, d_M\}\) and every \(m,n,p,q \in \{1, \dotsc, d_N\}\), for some constants \(\Lambda_1, \Lambda_2 \in \mathbb{R}\). Then, considering \(M \times N\) endowed with the product metric, it holds that \[K_{g \oplus h} \leq \max\{D_1, D_2\},\quad \mathrm{Ric}_{g \oplus h} \geq \min\{\delta_1, \delta_2 \},\] and, for every \((x,y) \in M \times N\), in normal coordinates centered at \((x,y)\), \[|(R_{g \oplus h})_{ijkl}(x,y)| \leq \max\{\Lambda_1, \Lambda_2\},\] for every \(i, j, k, l \in \{1, \dotsc, d_M + d_N\}\).

Proof. From the previous Lemma, we know that for any point \((x, y) \in M \times N\) and any four tangent vectors \(X, Y, Z, W \in T_{(x, y)} M \times N\), it holds that \[R_{g \oplus h}(X, Y, Z, W) = R_g(X_1, Y_1, Z_1, W_1) + R_h(X_2, Y_2, Z_2, W_2),\] where we are using the same notation as in 19 for the decomposition of the tangent vectors. This identity allows us to conclude that, in normal coordinates centered at any point \((x, y) \in M \times N\), \[|(R_{g \oplus h})_{ijkl}(x,y)| \leq \max\{\Lambda_1, \Lambda_2\},\] for every \(i, j, k, l \in \{1, \dotsc, d_M + d_N\}\), by simply considering an orthonormal basis of \(T_{(x, y)} M \times N\). Thus, \[|(R_{g \oplus h})_{ijkl}(x,y)| = |R_{g \oplus h}(e_i, e_j, e_k, e_l)(x,y)| \neq 0\] if and only if \(e_i, e_k, e_k\) and \(e_l\) all belong to either \(T_x M\) or \(T_y N\).

Now, let \(\Pi\) be a plane in \(T_{(x, y)} M \times N\). If \(\Pi \subset T_x M\), then \(\Pi = \mathit{span}\{X, Y\}\), for some orthonormal vectors \(X, Y \in T_x M\). Thus \[K_{g \oplus h}(\Pi) = K_{g \oplus h}(X, Y) = R_{g \oplus h}(X, Y, Y, X) = R_g(X, Y, Y, X) = K_g(X, Y).\] If \(\Pi \subset T_y N\) it follows that \(K_{g \oplus h}(\Pi) = K_h(\tilde{X}, \tilde{Y})\) for some orthonormal vectors \(\tilde{X}, \tilde{Y} \in T_y N\). If \(\Pi \not\subset T_x M\), \(\Pi \not\subset T_y N\) then we can take \(X \in T_x M\), \(Y \in T_y N\) such that \(\Pi = \mathit{span}\{X, Y\}\). In this case, by 20 \[K_{g \oplus h}(X, Y) = R_{g \oplus h}(X, Y, Y, X) = R_g(X, 0, 0, X) + R_h(0, Y, Y, 0) = 0.\]

For the last inequality, let \(X \in T_{(x, y)} M \times N\). Then \[\begin{align} \mathrm{Ric}_{g \oplus h}(X, X) &= \sum_{i = 1}^{d_M+d_N} R_{g \oplus h}(e_i, X, X, e_i) \\&= \sum_{i = 1}^{d_M} R_g(e_i, X_1, X_1, e_i) + \sum_{i = d_M+1}^{d_M+d_N} R_h(e_i, X_2, X_2, e_i) \\&= \mathrm{Ric}_g(X_1, X_1) + \mathrm{Ric}_h(X_2, X_2), \end{align}\] where \(\{e_i\}_{i = 1}^{d_M+d_N}\) is an orthonormal basis for \(T_{(x, y)} M \times N\), which can be divided into an orthonormal basis for \(T_x M\), and an orthonormal basis for \(T_y N\), namely \(\{e_i\}_{i = 1}^{d_M}\) and \(\{e_i\}_{i = d_M+1}^{d_M+d_N}\), respectively. ◻

There is one particular family of manifolds which have convenient curvature properties. These manifolds are known as Einstein manifolds.

Definition 19 (Einstein manifold). We say a manifold \(M\) is Einstein if \[\mathrm{Ric}_g = \lambda g,\] for some constant \(\lambda \in \mathbb{R}\).

Proposition 59 ([63]). The product of two Riemannian manifolds which are Einstein with the same constant \(\lambda\) is again Einstein with the same constant when endowed with the product metric.

8.3 Injectivity and convexity radii↩︎

Before we study Riemannian submersions, let us first define the concept of injectivity and convexity radii of a Riemannian manifold. In fact, such quantities play a crucial role in the proof of the main results from [SectionSuboptimality] [SectionPI]. Having lower bounds for such quantities ultimately allows us to control the Poincaré and the log-Sobolev inequality constants studied in [SectionPI] [SectionPItoLSI].

In order to define the injectivity radius of a Riemannian manifold, we need to introduce the exponential map and its inverse.

Definition 20 (Exponential map, logarithm). Let \((M, g)\) be a compact14 Riemannian manifold. For every \(x \in M\), the exponential map at \(p\), \(\exp_p : T_p M \to M\), is defined as \[\exp_p(v) := \gamma_{p, v}(1),\] where \(\gamma_{p, v}\) is the unique curve in \(M\) such that \(\gamma(0) = p\), \(\gamma'(0) = v\).

Whenever the exponential map is invertible, one can define its inverse, known as the logarithm* at \(p\), denoted as \(\log_p : M \to T_pM\).*

Remark 60 (Normal coordinates). On any connected neighborhood of \(0 \in T_xM\) on which \(\exp_x\) is a diffeomorphism onto its image, the exponential map induces local coordinates in a natural way. These coordinates are known as normal coordinates centered at \(x\) (see [39]).

Definition 21 (Injectivity Radius). Let \((M, g)\) be a Riemannian manifold. The injectivity radius of \(M\), denoted by \(i(M)\), is the greatest radius \(r > 0\) for which the exponential map defined on the ball \(\mathcal{B}(r, x)\) is a diffeomorphism onto its image for all \(x \in M\); \[i(M) := \inf_{\vphantom{\tilde{M}}x \in M} \sup_{r > 0} \{\exp_x |_{\mathcal{B}(r,x)} \text{ is a diffeomorphism onto its image}\}.\]

Bounding the injectivity radius is not an easy task in general. Nevertheless, for compact manifolds with upper bounded sectional curvature, it becomes easier.

Proposition 61 ([64]). Let \((M, g)\) be a compact Riemannian manifold, and let \(K_g\) denote its sectional curvature. Assume that there exists a constant \(D > 0\) such that \(K_g \leq D\). Then \[i(M) \geq \min \left\{\frac{\pi}{\sqrt{D}}, \frac{1}{2}l(M)\right\},\] where \(l(M)\) is the length of the shortest non-trivial periodic geodesic in \(M\).

Before defining the convexity radius of a Riemannian manifold, let us define the cut locus of a point in a Riemannian manifold, which is closely related to its injectivity radius.

Definition 22 (Cut locus). Let \((M, g)\) be a complete and connected Riemannian manifold, and let \(p \in M\), \(v \in T_p M\). We define the cut time associated with \(p\) and \(v\) as the maximum time \(t\) for which the geodesic \(\gamma : [0, t] \to M\), \(\gamma_{(p, v)}(s) := \exp_p(sv)\) minimizes the distance between \(p\) and \(\exp_p(sv)\) for every \(s \in [0, t]\), \[t_{\mathit{cut}}(p, v) := \sup \{t > 0 : \left.\gamma_{(p, v)}\right|_{[0, t]} \mathrm{ is minimizing}\}.\] When \(t_{\mathit{cut}}(p, v) < \infty\), we denote \(\gamma_{(p, v)}(t_{\mathit{cut}}(p, v))\) as the cut point of \(p\) along \(\gamma_{(p, v)}\).

This way, we define the cut locus of \(p \in M\) as the set of cut points of \(p\) along all possible geodesics, \[C(p) := \{\gamma_{(p, v)}(t_{\mathit{cut}}(p, v)) : t_{\mathit{cut}}(p, v) < \infty, \;v \in T_p M, |v|_g = 1\}.\]

From the definition of the cut locus, given \(p \in M\) and \(q \in C(p)\), it must be the case that \(d_g(p, q) \geq i(M)\). Otherwise, there would exist a unique distance-minimizing geodesic joining \(p\) and \(q\).

The next quantity of interest is the convexity radius, which is defined as the greatest radius for which any ball of such radius centered anywhere in the manifold is geodesically convex.

Definition 23. Given a manifold \(M\) a subset \(U \subset M\) is said to be geodesically convex—or simply convex—if, for any two points \(x\) and \(y\) in \(U\), there exists a unique length-minimizing geodesic in \(M\) connecting \(x\) and \(y\) which lies entirely in \(U\).

Definition 24 (Convexity radius). The convexity radius of a Riemannian manifold, denoted by \(\mathit{conv}(M)\), is the greatest radius \(r > 0\) for which the ball \(\mathcal{B}(r,x)\) is geodesically convex for all \(x \in M\); \[\mathit{conv}(M) := \inf_{\vphantom{\tilde{M}} x \in M}\sup_{r > 0}\{\mathcal{B}(r,x) \text{ is geodesically convex}\}.\]

Again, as for the injectivity radius, controlling the convexity radius is in general a complicated task, which becomes easier for compact manifolds with upper bounded sectional curvature.

Proposition 62 ([65]). Let \((M, g)\) be a compact Riemannian manifold, and let \(K_g\) denote its sectional curvature. Assume that there exists a constant \(D > 0\) such that \(K_g \leq D\), then \[\mathit{conv}(M) \geq \frac{1}{2} \min\left\{\frac{\pi}{\sqrt{D}}, \frac{1}{2}l(M)\right\},\] where again \(l(M)\) is the length of the shortest non-trivial periodic geodesic in \(M\).

61 62 show that, in order to obtain lower bounds on the injectivity and convexity radii of a compact manifold, it is convenient to bound its sectional curvature and to control the length of its shortest non-trivial periodic geodesic.

8.4 Riemannian submersions↩︎

In this subsection we will gather some of the most basic definitions and results in the field of Riemannian submersions. For a more thorough analysis we refer the reader to classic references such as [39] or O’Neill’s articles [66], [67].

Definition 25 (Submersion). Let \(M\) and \(B\) be two smooth manifolds, and let \(F : M \rightarrow B\) be a smooth map. We say that \(F\) is a submersion if its differential \(F_*|_p\) is surjective for all \(p \in M\). \(M\) and \(B\) are often referred to as the total and base spaces, respectively.

Given a submersion, the tangent space of the total space at each point can be split into its vertical and horizontal components.

Definition 26 (Vertical and horizontal tangent spaces). Given a submersion \(\pi: M \rightarrow B\), we define the vertical tangent space at \(p \in M\) as \(V_p M := \ker \pi_*|_p\). Note that \(V_p M\) is a vector space of dimension \(\dim(M) - \dim(B)\). This way, if \(M\) is a Riemannian manifold, we can define the horizontal tangent space as \(H_p M = (V_p M)^{\bot}\). Thus, \[T_p M = V_p M \oplus H_p M,\] for every \(p \in M\).

Notation 63. If \(\pi: M \to B\) is a submersion and \(X\) is a tangent vector at \(x \in M\), \(X \in T_x M\), we will denote the horizontal and vertical components of \(X\) by \(X^\mathcal{H}\) and \(X^\mathcal{V}\), respectively. This way, \[X = X^\mathcal{V} + X^\mathcal{H} \in V_x M \oplus H_x M.\]

Next, we will introduce Riemannian submersions, which are submersions between manifolds that preserve the metric between the total and the base space.

Definition 27 (Riemannian Submersion). Let \((M, g)\) and \((B, h)\) be two Riemannian manifolds, and let \(\pi: M \rightarrow B\) be a submersion. We say \(\pi\) is a Riemannian submersion if for every \(p \in M\), \(\pi_*|_p\) maps \(H_p M\) isometrically onto \(T_{\pi(p)} B\). That is, for each \(p \in M\) and \(X, Y \in H_p M\), \[\langle X, Y \rangle_g = \langle \pi_*|_p X,\pi_*|_p Y \rangle_h.\]

In particular, given a Riemannian submersion \(\pi: (M, g) \to (B, h)\) the metric \(g\) splits as \[g = g^\mathcal{H} + g^\mathcal{V},\] where \(g^\mathcal{H}\) and \(g^\mathcal{V}\) correspond to the restrictions of \(g\) on the horizontal and the vertical tangent spaces, respectively.

When given a Riemannian submersion between two Riemannian manifolds, one can relate the geodesics of the base and the total space using the following result.

Proposition 64 ([67]). Let \(\pi: (M, g) \rightarrow (B, h)\) be a Riemannian submersion. If \(\gamma\) is a geodesic on \(M\) that is horizontal at some point \(\gamma(t_0)\) ( i.e. \(\gamma'(t_0) \in H_{\gamma(t_0)} M\)), then \(\gamma(t)\) is horizontal for every \(t\) in the domain of \(\gamma\), and so \(\pi \circ \gamma\) is a geodesic of \(B\). If \(M\) is complete, every geodesic \(\tilde{\gamma}\) in \(B\) is the projection of a horizontal geodesic \(\gamma\) in \(M\), i.e. \(\tilde{\gamma} = \pi \circ \gamma\).

One interesting property of Riemannian submersions is that they map sufficiently small balls in the total space to balls in the base space of the same radius.

Proposition 65. Let \((M, g)\) be a complete Riemannian manifold, let \(x \in M\) and let \(\pi: (M, g) \to (B, h)\) be a Riemannian submersion, then \[\pi(\mathcal{B}(r,x)) = \mathcal{B}(r,\pi(x)),\] for every \(r < \min\{i(M), i(B)\}\), where \(i(M)\) and \(i(B)\) denote the injectivity radii of \(M\) and \(B\), respectively.

Proof. First, let \(y \in \mathcal{B}(r,x)\). Let \(\gamma\) be the unique—as \(d_g(x, y) < i(M)\)—geodesic in \(\mathcal{B}(r, x)\) such that \(\gamma(0) = x\) and \(\gamma(1) = y\). Then \[r > d_g(x, y) = \int_0^1 |\gamma'(t)|_g \,dt \geq \int_0^1 |\gamma'(t)^\mathcal{H}|_g \,dt = \int_{0}^1 |(\pi \circ \gamma)'(t)|_h\, dt \geq d_h(\pi(x), \pi(y)),\] and so \(\pi(y) \in \mathcal{B}(r, \pi(x))\).

Now, let \(y \in \mathcal{B}(r, \pi(x))\), and let \(\tilde{\gamma}\) be the unique geodesic in \(\mathcal{B}(r, \pi(x))\) such that \(\tilde{\gamma}(0) = y\) and \(\tilde{\gamma}(1) = \pi(x)\). Since \(M\) is complete, by 64 we know that there exists a horizontal geodesic \(\gamma\) starting at \(x\) and such that \(\tilde{\gamma} = \pi \circ \gamma\), \(l(\tilde{\gamma}) = l(\gamma) < r\). Therefore, \(\gamma(1) \in \mathcal{B}(r, x)\) and so \(y = \tilde{\gamma}(1) = \pi(\gamma(1))\), where \(\gamma(1) \in \mathcal{B}(r, x)\). ◻

One simple way to construct surjective Riemannian submersions given a Riemannian manifold is by considering quotient maps with respect to group actions. Therefore, we will first introduce some basic definitions about group actions on manifolds and then proceed to state the main results, which guarantee that certain group actions induce Riemannian submersions.

Definition 28 (Group action on a manifold). Let \(G\) be a group and \(M\) be a manifold. A left action of \(G\) on \(M\) is a map \(G \times M \to M\), usually denoted as \((x, p) \mapsto x\cdot p\) verifying the following:

  • \(x_2 \cdot (x_1 \cdot p) = (x_2x_1)\cdot p\) for every \(x_1, x_2 \in G\), \(p \in M\).

  • \(e \cdot p = p\) for every \(p \in M\), \(e\) being the unit of \(G\).

Right actions are defined analogously.

Although group actions often give rise to Riemannian submersions, not every group action does. Nevertheless, if the group action fulfills certain properties, this is the case.

Definition 29 (Free action). A group action of \(G\) on a manifold \(M\) is free if for every \(p \in M\), \[\{x \in G : x p = p\} = \{e\},\] where \(e\) denotes the unit of \(G\).

Definition 30 (Smooth action). The action of a Lie group \(G\) on a manifold \(M\) is smooth if the map \(G \times M \rightarrow M\) given by \((x, p)\mapsto x \cdot p\) is smooth.

Definition 31 (Proper action). A group action of \(G\) on a manifold \(M\) is said to be proper if the map \[\begin{align} G \times M &\rightarrow M \times M\\ (x, p) &\mapsto (x\cdot p, p) \end{align}\] is proper, i.e., the preimage of every compact set is compact.

In particular, every smooth action by a compact Lie group is proper (cf. [39]).

Definition 32 (Isometric action). A group action of \(G\) on a Riemannian manifold \((M, g)\) is said to be isometric if, for every \(x \in G\), the map \[\begin{align} \alpha_x : M &\to M\\ p &\mapsto x \cdot p \end{align}\] is an isometry on \((M, g)\) (i.e. a diffeomorphism such that for every point \(p \in M\), its differential satisfies that \(g(u, v) = g(d\alpha_x(u), d\alpha_x(v))\) for every \(u, v \in T_p M\)).

When a group action on a manifold satisfies all of the above conditions, the quotient map from the manifold to the orbit space is a Riemannian submersion.

Proposition 66 ([39]). Suppose \((M, g)\) is a Riemannian manifold, and let \(G\) be a Lie group acting smoothly, freely, properly and isometrically on \(M\). Then the orbit space \(M/G\) has a unique smooth manifold structure and Riemannian metric such that \(\pi : M \rightarrow M/G\) is a Riemannian submersion.

The Riemannian submersions that arise as the orbit space of a suitable group action are known as \(G\)-bundles. Let us briefly introduce them, as they are discussed in 3. For a more in-depth study of fiber bundles or \(G\)-bundles, we refer to [68].

Definition 33. Let \(M\) and \(B\) be two connected manifolds. We say that \(\pi: M \to B\) is a locally trivial fibration, or fiber bundle with fiber \(F\) if it satisfies all of the following conditions:

  • \(\pi^{-1}(x) \simeq F\) for every \(x \in B\).

  • \(\pi\) is surjective.

  • For every \(x \in B\) there exists an open neighborhood \(U \subset B\) and a homeomorphism—also known as a local trivialization—\(\Phi: \pi^{-1}(U) \to U \times F\) such that the following diagram commutes: \[\begin{tikzcd} {\pi^{-1}(U)} && {U \times F} \\ \\ & U \arrow["\Phi", from=1-1, to=1-3] \arrow["\pi"', from=1-1, to=3-2] \arrow["{\pi_1}", from=1-3, to=3-2] \end{tikzcd},\] where \(\pi_1\) denotes the projection onto the first component.

Definition 34. Let \(G\) be a Lie group. A principal \(G\)-bundle is a fiber bundle \(\pi: M \to B\) with fiber \(G\) satisfying all of the following conditions:

  • There exists a free action \(\mu\) of \(G\) on \(M\) such that \[\begin{tikzcd} {M \times G} &&& M \\ \\ {B \times\{e\}} &&& B \arrow["\mu", from=1-1, to=1-4] \arrow["{\pi \times \epsilon}"', from=1-1, to=3-1] \arrow["\pi", from=1-4, to=3-4] \arrow["{=}"{description}, draw=none, from=3-1, to=3-4] \end{tikzcd}\] commutes, where \(\epsilon(g) = e\) for every \(g \in G\), and \(e\) denotes the unit of \(G\).

  • The induced action on the fibers \[\mu|_{\pi^{-1}(x)}: \pi^{-1}(x) \times G \to \pi^{-1}(x),\] is free and transitive—i.e. for any two points \(x_1, x_2 \in \pi^{-1}(x)\) there exists some \(g \in G\) such that \(x_1 = \mu(x_2,g)\).

  • There exist local trivializations \(\Phi: \pi^{-1}(U) \to U \times G\) such that the following diagrams commute: \[\begin{tikzcd} {\pi^{-1}(U) \times G} &&& {U \times G \times G} \\ \\ {\pi^{-1}(U)} &&& {U \times G} \arrow["{\Phi \times \mathrm{Id}}", from=1-1, to=1-4] \arrow["\mu"', from=1-1, to=3-1] \arrow["{\mathrm{Id} \times \cdot}", from=1-4, to=3-4] \arrow["\Phi", from=3-1, to=3-4] \end{tikzcd},\] where \(\cdot\) denotes the usual group multiplication of \(G\).

As we mentioned earlier, whenever the assumptions described in 66 hold, the Riemannian submersion is also a principal \(G\)-bundle.

Proposition 67 ([68]). Let \(M\) be a manifold and let \(G\) be a compact Lie group acting freely and smoothly on \(M\). Then \[\pi: M \to M/G\] is a principal \(G\)-bundle.

The Riemannian submersions that we consider in this text satisfy one more property; all of the fibers of \(\pi\) are totally geodesic.

Definition 35 (Totally geodesic submanifold). A Riemannian submanifold \((\tilde{M}, \tilde{g})\) of \((M, g)\) is said to be totally geodesic if every \(g\)-geodesic that is tangent to \(\tilde{M}\) at some point \(t_0\) stays in \(\tilde{M}\) for all \(t\) in some interval \((t_0-\varepsilon, t_0+\varepsilon)\), \(\varepsilon > 0\).

The definition of a totally geodesic submanifold can also be understood in the terms of the following proposition.

Proposition 68 ([39]). Let \((\tilde{M}, \tilde{g})\) be an embedded Riemannian submanifold of a Riemannian manifold \((M, g)\). If \(\tilde{M}\) is totally geodesic in \(M\), then every \(\tilde{g}\)-geodesic in \(\tilde{M}\) is also a \(g\)-geodesic in \(M\).

In particular, given a Lie group \(G\) endowed with a bi-invariant metric \(g\) and some closed subgroup \(H \subset G\), it holds that there exists a metric \(h\) on \(G/H\) such that \(\pi: (G, g) \mapsto (G/H, h)\) is a Riemannian submersion with totally geodesic fibers [66].

8.4.1 Curvature of Riemannian submersions↩︎

Let us now examine the relationship between the curvature of the total and the base space of a Riemannian submersion. For the purpose of this work, we will only define the so-called O’Neill tensor \(T\). For an in-depth study of the O’Neill tensors, see [66].

Definition 36 (O’Neill tensor). Let \(\pi : M \to B\) be a Riemannian submersion. For any two vector fields \(E, F \in \mathfrak{X}(M)\), we define the O’Neill tensor \(T\) as \[T_E F := (\nabla_{E^\mathcal{V}} (F^\mathcal{V}))^\mathcal{H} + (\nabla_{E^\mathcal{V}} (F^\mathcal{H}))^\mathcal{V},\] where \(\cdot ^\mathcal{V}\) and \(\cdot^{H}\) denote the vertical and horizontal parts of a vector field, respectively.

The tensor \(T\) allows us to relate the sectional curvature of the base space and the fibers of a Riemannian submersion with that of the total space.

Lemma 21 ([66]). Let \(\pi : (M, g) \rightarrow (B, h)\) be a Riemannian submersion. Given two horizontal orthonormal vectors \(X, Y \in H_xM\) and two vertical orthonormal vectors \(V, W \in V_xM\) at some point \(x \in M\), it holds that \[K_g(X, Y) = K^*_h(\pi_* X, \pi_* Y) - \frac{3}{4}\left|[X, Y]^\mathcal{V}\right|_g^2,\] and \[K_g(V, W) = \hat{K}(V, W) - \langle T_V V, T_W W\rangle_g + |T_V W|_g^2,\] where \(K_g\) is the sectional curvature of \(M\), \(\hat{K}\) is the sectional curvature of the fibers of \(\pi\), \(K^*_h\) is the sectional curvature of \(B\), and \(T\) is the O’Neill tensor.

Notice that, while the sectional curvature takes a pair of tangent vectors as arguments, both the Lie derivative and the O’Neill tensor take two vector fields as arguments. Nevertheless, via a slight abuse of notation, the expressions are well-defined.

Remark 69. The expressions shown in 21 are well-defined, as for any two horizontal vector fields \(X, Y \in \mathfrak{X}(M)\), and \(p \in M\) \[\left.\left([X, Y]^\mathcal{V}\right)\right|_p\] only depends on \(X_p\) and \(Y_p\).

Similarly, for any two vertical vector fields \(U, V \in \mathfrak{X}(M)\), and \(p \in M\) \[\left.\left(T_U V\right)\right|_p\] only depends on \(U_p\) and \(V_p\).

Proof. Taking coordinates on a neighborhood of \(p\), we can write \[X = X^i(x) \left.\frac{\partial}{\partial h^i}\right|_x,\quad Y = Y^i(x) \left.\frac{\partial}{\partial h^i}\right|_x,\] where \(\{\left.\frac{\partial}{\partial h^i}\right|_x\}\) form an orthonormal basis of \(H_x M\), for every \(x\) in a neighborhood of \(p\). In these coordinates, \[\left.[X, Y]\right|_p = \sum_{i, j} X^i(p) Y^j(p)\Big[\frac{\partial}{\partial h^i}\bigg|_p,\, \frac{\partial}{\partial h^j}\bigg|_p\Big] + \sum_{j} X(Y^j(p))\frac{\partial}{\partial h^j}\bigg|_p - \sum_{i} Y(X^i(p)) \frac{\partial}{\partial h^i}\bigg|_p.\] Note that the two last summands in this expression are horizontal, thus \[\left(\left.[X, Y]\right|_p\right)^\mathcal{V} = \left(\sum_{i, j} X^i(p) Y^j(p)\Big[\frac{\partial}{\partial h^i}\bigg|_p, \frac{\partial}{\partial h^j}\bigg|_p\Big]\right)^\mathcal{V},\] which only depends on the point \(p\).

Working in the same coordinates, and using the definition of \(T\) \[(T_U V)|_p = \left.(\nabla_{U} V)^\mathcal{H}\right|_p = (U(V^k(p)) + U^i(p) V^j(p) \Gamma^k_{ij}(p)) \left.\frac{\partial}{\partial h^k}\right|_p.\] Note that, since \(V\) is vertical, \(V^k(x) \equiv 0\) for every \(k\) (we are only summing over the horizontal components). Thus \[(T_U V)|_p = (U^i(p) V^j(p) \Gamma^k_{ij}(p)) \left.\frac{\partial}{\partial h^k}\right|_p,\] which only depends on \(p\). ◻

Let us conclude with a useful result, which can be seen in [49], and relates the O’Neill tensor \(T\) to the fibers of a Riemannian submersion being totally geodesic.

Lemma 22. Let \(\pi: (M, g) \to (B, h)\) be a Riemannian submersion. \(\pi\) has totally geodesic fibers if and only if the O’Neill tensor \(T\) from 36 vanishes identically.

This result allows us to conclude that, whenever a Riemannian submersion has totally geodesic fibers, the sectional curvature of the fibers is the same as that of the total space.

Moreover, using 22, we can guarantee that certain nested submersions inherit totally geodesic fibers.

Proposition 70. Let \((M_1, g_1)\), \((M_2, g_2)\) and \((M_3, g_3)\) be three Riemannian manifolds, and let \(\pi_1, \pi_2\) and \(\pi_3\) be three Riemannian submersions, \[\begin{align} \pi_1 : (M_1, g_1) \to (M_2, g_2),\\ \pi_2: (M_2, g_2) \to (M_3, g_3),\\ \pi_3: (M_1, g_1) \to (M_3, g_3), \end{align}\] such that \(\pi_3 = \pi_2 \circ \pi_1\). If \(\pi_3\) has totally geodesic fibers, then \(\pi_2\) has totally geodesic fibers.

Proof. We want to show that the O’Neill tensor \(T\) associated with \(\pi_2\) vanishes identically. To do so, we will show that for any two vertical vector fields with respect to \(\pi_2\), \(U, V \in \mathfrak{X}(M_2)\), it holds that \[\label{eq:vanishingverticalcon} (\nabla^{M_2}_U V)^{\mathcal{H}} = 0.\tag{51}\] Should this hold, the result would follow: using the compatibility of the Levi-Civita connection \(\nabla^{M_2}\), we know that for every two vertical vector fields with respect to \(\pi_2\), \(U, V \in \mathfrak{X}(M_2)\) and any horizontal vector field with respect to \(\pi_2\), \(X \in \mathfrak{X}(M_2)\), \[0 = U g_2(V, X) = g_2(\nabla^{M_2}_U V, X) + g_2(V, \nabla^{M_2}_U X) = g_2(V, \nabla^{M_2}_U X),\] where the first equality follows from \(V\) begin vertical and \(X\) horizontal, and the last equality follows from the assumption that 51 holds. This allows us to conclude that \(\nabla^{M_2}_U X\) is always horizontal, and so \(T\) vanishes (cf. 36).

Let us then prove that 51 holds. Given \(U, V \in \mathfrak{X}(M_2)\)—which are vertical with respect to \(\pi_2\) by assumption—we can lift them to \(M_1\), and obtain \(\tilde{U}, \tilde{V} \in \mathfrak{X}(M_1)\) which are such that \(d\pi_1(\tilde{U}) = U\) and \(d\pi_1(\tilde{V}) = V\). Note that \(\tilde{U}\) and \(\tilde{V}\) are vertical with respect to \(\pi_3\), \[\begin{align} d\pi_3(\tilde{U}) = d\pi_2(d\pi_1(\tilde{U})) = d\pi_2(U) = 0,\\ d\pi_3(\tilde{V}) = d\pi_2(d\pi_1(\tilde{V})) = d\pi_2(V) = 0, \end{align}\] where the last equality follows from the assumption that \(U\) and \(V\) are vertical with respect to \(\pi_2\). Now, from \(\pi_3\) having totally geodesic fibers, we know that \(\nabla^{M_1}_{\tilde{U}} \tilde{V}\) is vertical with respect to \(\pi_3\). Moreover, it is known [66] that \(\nabla^{M_1}_{\tilde{U}} \tilde{V}\) is such that \(\pi_1(\nabla^{M_1}_{\tilde{U}} \tilde{V}) = \nabla^{M_2}_U V\), and so \[0 = d\pi_3(\nabla^{M_1}_{\tilde{U}} \tilde{V}) = d\pi_2(d\pi_1(\nabla^{M_1}_{\tilde{U}} \tilde{V})) = d\pi_2(\nabla^{M_2}_U V),\] finishing the proof. ◻

8.4.2 Curvature of complex Stiefel manifolds↩︎

Definition 37 (Stiefel manifold). Given \(k, n \in \mathbb{N}\) with \(1 \leq k \leq n\), the set of linear isometries from \(\mathbb{C}^k\) to \(\mathbb{C}^n\) is a smooth manifold known as the complex Stiefel manifold (cf. [69]), \[\mathrm{V}_k(\mathbb{C}^n) := \{X \in \mathbb{C}^{n \times k} : X^\dagger X = \mathbb{1}\}.\]

We will restrict our attention to the case when \(1 < k < n\), as \(\mathrm{V}_n(\mathbb{C}^n) = \mathrm{U}(n)\), and \(\mathrm{V}_1(\mathbb{C}^n) = \mathbb{S}^{2n-1}\). These manifolds can be seen as the quotient space \(\mathrm{U}(n)/\mathrm{U}(n-k)\), where \(\mathrm{U}(n-k)\) acts on \(\mathrm{U}(n)\) as \[\begin{align} \label{eq:defactionstiefel} \begin{aligned} \mathrm{U}(n-k) \times \mathrm{U}(n) &\to \mathrm{U}(n)\\ (X, U) &\mapsto U\begin{pmatrix} \mathbb{1}_k & 0\\ 0 & X \end{pmatrix}. \end{aligned} \end{align}\tag{52}\] It is easy to show that this action is free: let \(U \in \mathrm{U}(n)\) and \(X \in \mathrm{U}(n-k)\). If \[U\begin{pmatrix} \mathbb{1}_k & 0\\ 0 & X \end{pmatrix} = U,\] then \[\begin{pmatrix} \mathbb{1}_k & 0\\ 0 & X \end{pmatrix} = \mathbb{1}_n,\] and so it must be the case that \(X = \mathbb{1}_{n-k}\). Moreover, the action is smooth and isometric whenever \(\mathrm{U}(n)\) is endowed with its bi-invariant metric \(g\). Therefore, by 66 we know that there exists a metric \(h\) on \(\mathrm{V}_k(\mathbb{C}^n)\) for which the quotient map \[\pi: (\mathrm{U}(n), g) \to (\mathrm{V}_k(\mathbb{C}^n), h)\] is a Riemannian submersion. Furthermore, as we discussed in 8.4, the submersion has totally geodesic fibers.

In order to control the curvature of the complex Stiefel manifolds, we will need an auxiliary result, which will allow us to lift vector fields in a convenient way.

Proposition 71. Let \(G\) be a Lie group endowed with a bi-invariant metric and let \(H\) be a closed subgroup of \(G\). Consider the action of \(H\) on \(G\) from the right: \[\begin{align} H \times G \to G\\ (h, g) \mapsto gh. \end{align}\] Let \(X \in T_e G\) be a horizontal vector with respect to the fibration \(G \to G/H\). Then its corresponding left-invariant vector field \[X^L(p) := pX\] is horizontal for every \(p \in G\).

Proof. Note that \[\begin{align} L_g: G \to G\\ x \mapsto gx \end{align}\] is an isometry for every \(g \in G\). Furthermore, \(L_g\) maps fibers to fibers: \(L_g(xH) = (gx)H\). Therefore, if \(X\) is horizontal at \(e\)—i.e. perpendicular to the fiber \(H\)—it follows that \((L_g)_* X\) is perpendicular to \((dL_g)_* H\), where \((L_g)_* X = gX\) and \((dL_g)_* H\) is the tangent space of the fiber \(gH\). ◻

71 allows us to bound the curvature of the Stiefel and Grassmann manifolds.

Proposition 72. Let \(\mathrm{U}(n)\) be endowed with its bi-invariant metric \(g\) (cf. 51). Let \(k < n\) and consider the induced metric \(h\) on \(\mathrm{V}_k(\mathbb{C}^{n})\) for which \(\pi : (\mathrm{U}(n), g) \to (\mathrm{V}_k(\mathbb{C}^{n}), h)\) is a Riemannian submersion. Then \[0 \leq K_{h} \leq 2,\quad \mathrm{\textit{and}}\quad \mathrm{Ric}_{h} \geq 0.\]

Proof. Recall that \(\mathrm{V}_k(\mathbb{C}^{n})\) is the quotient space associated with the action shown in 52 . Now, let \(p \in \mathrm{U}(n)\). We can use 21 71 to conclude that, given two horizontal orthonormal vectors \(X, Y \in T_p \mathrm{U}(n)\), we can extend them as left-invariant vector fields \(\tilde{X}, \tilde{Y}\) which remain horizontal. Thus, \[\begin{align} 0 \leq K_{h}(\pi_* \tilde{X}(p), \pi_* \tilde{Y}(p)) &= K_{g}(\tilde{X}(p), \tilde{Y}(p)) + \frac{3}{4}|[\tilde{X}, \tilde{Y}]^\mathcal{V}|^2_g \\&\leq K_{g}(X, Y) + \frac{3}{4}|[\tilde{X}, \tilde{Y}]|^2_g \\&= |[X, Y]|^2_g \\&\leq 2. \end{align}\] where the last equality follows from 49 50 53, and the last inequality follows from 50 .

Furthermore, since the sectional curvature is non-negative, it holds that its Ricci curvature is non-negative as well. Indeed, let \(V \in T_p \mathrm{V}_k(\mathbb{C}^n)\) be a unitary vector, we can extend \(V\) to a orthonormal basis of \(T_p \mathrm{V}_k(\mathbb{C}^n)\), \(\{V, e_2, \dotsc, e_n\}\). This way, writing \(\mathrm{Ric}_{h}(V, V)\) as \[\mathrm{Ric}_{h}(V, V) = \sum_{i = 2}^n K_{h}(V, e_k) \geq 0,\] the result follows. ◻

8.4.3 Curvature of complex Grassmann manifolds↩︎

Lastly, we will bound the curvature of the manifold of linear subspaces of complex dimension \(k\) in \(\mathbb{C}^n\), which is known a the Grassmann manifold \(\mathrm{Gr}_k(\mathbb{C}^n)\). For an in-depth study of Grassmann manifolds, see [56] and the references therein.

Definition 38 (Grassmann manifold). We define the complex Grassmann manifold \(\mathrm{Gr}_k(\mathbb{C}^n)\) as \[\require{physics} \mathrm{Gr}_k(\mathbb{C}^n) = \{P \in \mathbb{C}^{n \times n} : P^\dagger = P,\;P^2 = P,\;\Tr(P) = k\}.\]

In particular, each point \(P \in \mathrm{Gr}_k(\mathbb{C}^n)\) can be seen as a rank-\(k\) projector in \(\mathbb{C}^n\). Furthermore, \(\mathrm{Gr}_k(\mathbb{C}^n)\) can also be seen as the quotient space \[\mathrm{U}(n) / (\mathrm{U}(n-k) \times \mathrm{U}(k)),\] where the quotient is considered with respect to the right action on \(\mathrm{U}(n)\) defined as \[\begin{align} \mathrm{U}(k) \times \mathrm{U}(n-k) \times \mathrm{U}(n) &\to \mathrm{U}(n)\\ ((U, V),X) &\mapsto X \begin{pmatrix} U & 0\\ 0 & V \end{pmatrix}. \end{align}\] This action is again free, smooth and isometric whenever \(\mathrm{U}(n)\) is endowed with the bi-invariant metric \(g\). Therefore, the quotient description of \(\mathrm{Gr}_k(\mathbb{C}^n)\) allows us to conclude that it can be endowed with a Riemannian metric \(h\) for which \[\pi: (\mathrm{U}(n), g) \to (\mathrm{Gr}_k(\mathbb{C}^n), h)\] is a Riemannian submersion with totally geodesic fibers.

This way, following the same reasoning as in the proof of 72, we obtain the same bounds on the sectional and Ricci curvatures of \(\mathrm{Gr}_k(\mathbb{C}^n)\).

Proposition 73. Let \(\mathrm{U}(n)\) be endowed with its bi-invariant metric \(g\) (cf. 51). Consider the induced metric \(h\) on \(\mathrm{Gr}_k(\mathbb{C}^{n})\) for which \(\pi : (\mathrm{U}(n), g) \to (\mathrm{Gr}_k(\mathbb{C}^{n}), h)\) is a Riemannian submersion. Then \[0 \leq K_{h} \leq 2,\quad \mathrm{\textit{and}}\quad \mathrm{Ric}_{h} \geq 0.\]

Note that we can also define Grassmann manifolds via the action of \(\mathrm{U}(k)\) on the Stiefel manifold \(\mathrm{V}_k(\mathbb{C}^n)\) given as \[\begin{align} \mathrm{U}(k) \times \mathrm{V}_k(\mathbb{C}^n) &\to \mathrm{V}_k(\mathbb{C}^n) \\ (U, X) &\mapsto XU. \end{align}\] The action is smooth, free—\(XU = X\) implies that \(U = \mathbb{1}\) by simply multiplying by \(X^\dagger\) on both sides—and isometric whenever \(\mathrm{V}_k(\mathbb{C}^n)\) is endowed with the induced metric \(\tilde{g}\) making the projection map \(\pi_1 : (\mathrm{U}(n), g) \to (\mathrm{V}_k(\mathbb{C}^n), \tilde{g})\) a Riemannian submersion, where \(g\) is the bi-invariant metric of \(\mathrm{U}(n)\). This way, by 66 we know that there exists a smooth manifold structure and a Riemannian metric \(\tilde{h}\) on \(\mathrm{Gr}_k(\mathbb{C}^n)\) for which the projection \(\pi_2 : (\mathrm{V}_k(\mathbb{C}^n), \tilde{g}) \to (\mathrm{Gr}_k(\mathbb{C}^n), \tilde{h})\) is a Riemannian submersion. Furthermore, since \(\pi = \pi_2 \circ \pi_1\) is also a Riemannian submersion, by the uniqueness of the construction (cf. 66), we know that the manifold structures induced by \(\pi\) and \(\pi_2\) are the same, and that \(h = \tilde{h}\). Lastly, by 70 we know that \(\pi_2\) has totally geodesic fibers.

8.5 Riemannian submersions and fiber-invariant functions↩︎

In this subsection, we present several results concerning functions that are constant along the fibers of a surjective Riemannian submersion. The majority of the results presented here can be found in [2].

We will always be working with a surjective Riemmanian submersion \(\pi: (M, g) \to (B, h)\) and a sufficiently smooth function \(f: M \to \mathbb{R}\) which is constant in the fibers of \(\pi\), or a sufficiently smooth function \(\tilde{f} : B \to \mathbb{R}\). The invariance of \(f\) implies that there exists a unique function \(\tilde{f}\) satisfying \(f = \tilde{f} \circ \pi\) and vice-versa. Furthermore, both functions have the same smoothness properties [2].

We will conclude that having information about the gradient, the Hessian or the Laplacian of either \(f\) or \(\tilde{f}\) easily translates to having information about that of its projected or lifted version. This will be very convenient for us, as it will allow us to work in either the total or the base space of a Riemannian submersion, and then translate the properties obtained using the results that we will present in this section.

Let us first define the horizontal lift of a tangent vector in a Riemannian submersion.

Definition 39. Let \(\pi: (M, g) \to (B, h)\) be a Riemannian submersion, let \(x \in M\), and let \(v \in T_{\pi(x)} B\). The horizontal lift of \(v\) at \(x\), denoted as \(\mathrm{lift}_x\, v\), is the unique horizontal vector \(u \in H_x(M)\) such that \(\pi_*|_x(u) = v\).

The first result allows us to relate the gradient of a function in the base space \(\tilde{f}\) with the gradient of the lifted function \(f\) in the total space. In fact, as the following proposition claims, any function that is constant in the fibers of a Riemannian submersion has a horizontal gradient which is the lift of the gradient of \(\tilde{f}\).

Definition 40 (Gradient). Let \((M, g)\) be a Riemannian manifold. Given a differentiable function \(f: M \rightarrow \mathbb{R}\), the gradient of \(f\), denoted by \(\mathrm{grad}_{g}\, f\) it the unique vector field on \(M\) such that \[\langle\mathrm{grad}_{g}\,f, X\rangle_g = df(X),\] for any vector field \(X \in \mathfrak{X}(M)\), where \(df\) denotes the differential of \(f\), also written as \(f_*\)

Proposition 74 ([2]). Let \(\pi: (M, g) \to (B, h)\) be a surjective Riemannian submersion. Let \(\tilde{f} \in C^1(B)\), and let \(f : M \to \mathbb{R}\) be the unique function such that \(f = \tilde{f} \circ \pi\). Then for every \(x \in M\), \[\mathrm{grad}_g \, f(x) = \mathrm{lift}_x\, \mathrm{grad}_{h}\, \tilde{f}(\pi(x)).\]

74 allows us to relate the critical points of a fiber-invariant function \(f\) and its projection \(\tilde{f}\). Indeed, for any \(p \in B\) which is a critical point of \(\tilde{f}\), and any \(x \in \pi^{-1}(p)\), it holds that \(x\) is a critical point of \(f\). Furthermore, as \(\pi\) is a Riemannian submersion, we can also relate the norm of both gradients at any given point.

Remark 75. Since \(\pi\) preserves the norm of horizontal vectors, it holds that \[|\mathrm{grad}_g\, f(x)|_g = |\mathrm{grad}_{h}\, \tilde{f}(\pi(x))|_{h},\] for every \(x \in M\).

The following propositions allow us to relate the Hessian of the fiber-invariant function in the total space with that of the function in the base space.

Definition 41 (Hessian). Let \((M, g)\) be a Riemannian manifold. Let \(f : M \rightarrow \mathbb{R}\) be a twice differentiable function on \(M\). The Hessian of \(f\) is the tensor defined as \[\nabla^2 f := \nabla \nabla f = \nabla df,\] where \(\nabla\) denotes the Levi-Civita connection of \((M, g)\), acting as \[\nabla^2 f(X, Y) = \langle \nabla_X \mathrm{grad}_{g}\,f, Y \rangle_g,\] for any vector fields \(X, Y \in \mathfrak{X}(M)\). Moreover, \(\nabla^2 f\) is symmetric.

We will also denote \[\nabla^2 f(x)[X] := \nabla_X\, \mathrm{grad}_{g}\,f|_x.\] for any \(x \in M\) and any \(X \in T_x M\).

Proposition 76 ([2]). Let \((M, g)\) and \((B, h)\) be two Riemannian manifolds. Let \(\pi: (M, g) \to (B, h)\) be a surjective Riemannian submersion. Let \(\tilde{f} \in C^2(B)\), and let \(f = \tilde{f} \circ \pi\). Then for every \(x \in M\) and every horizontal vector \(u \in H_x M\), \[(\nabla^2_g\, f(x)[u])^{\mathcal{H}} = \mathrm{lift}_x \nabla^2_h\, \tilde{f}(\pi(x))[\pi_*u],\] where we used \(\nabla^2_g\) and \(\nabla^2_h\) to distinguish between the Hessian on \(M\) and \(B\).

In particular, if \(x \in M\) is a critical point of \(f\), then the eigenvalues of \(\nabla^2_g\, f(x)\) and \(\nabla^2_h\, \tilde{f}(\pi(x))\), are the same, together with a set of \(\dim(M) - \dim(B)\) additional eigenvalues equal to zero.

Proposition 77 ([2]). Let \(\pi: (M, g) \to (B, h)\) be a surjective Riemannian submersion. If \(x \in M\) is a critical point of a function \(f \in C^2(M)\) which is constant in the fibers of \(\pi\), then \(\nabla^2_g\, f(x)[v] = 0\) for every \(v \in V_x M\).

The following proposition allows us to relate the Hessian of a fiber-invariant function \(f\) on the total space of a Riemannian submersion to that of the projected function \(\tilde{f}\) in the base space.

Proposition 78. Let \(\pi: (M, g) \to (B, h)\) be a surjective Riemannian submersion. Let \(\tilde{f} \in C^2(B)\), and let \(f: M \to \mathbb{R}\) be such that \(f = \tilde{f} \circ \pi\). Let \(x \in M\) be a critical point of \(f\) in \(M\). Then for every \(v \in T_xM\) it holds that \[\nabla^2_g\,f(x)[v, v] = \nabla^2_g\,f(x)[v^{\mathcal{H}}, v^{\mathcal{H}}]= \nabla^2_h\, \tilde{f}(\pi(x))[\pi_* v, \pi_* v].\]

In particular, if \(p \in M\) is a critical point of \(\tilde{f}\) in \(B\), then for every \(v \in T_p B\) and every \(x \in \pi^{-1}(p)\) it holds that \(\nabla^2_h\, \tilde{f}(p)[v, v] = \nabla^2_g\, f(x)[\mathrm{lift}_x v, \mathrm{lift}_x v]\).

Proof. Assume that \(x \in M\) is some critical point of \(f\), and let \(v \in T_x M\). Using 77 and the linearity of \(\nabla\) we conclude that \[\nabla_v \, \mathrm{grad}_g\, f|_x = \nabla_{v^{\mathcal{V}}} \,\mathrm{grad}_g\, f|_x + \nabla_{v^{\mathcal{H}}} \,\mathrm{grad}_g\, f|_x = \nabla_{v^{\mathcal{H}}} \,\mathrm{grad}_g\, f|_x.\]

Therefore, \[\begin{align} \nabla^2_g\, f(x)[v, v] = \langle \nabla_v \,\mathrm{grad}_g\, f, v\rangle_g &= \langle \nabla_{v^{\mathcal{H}}} \,\mathrm{grad}_g\, f, v\rangle_g \\&= \langle \nabla_{v^{\mathcal{H}}} \,\mathrm{grad}_g\, f, v^{\mathcal{H}}\rangle_g + \langle \nabla_{v^{\mathcal{H}}} \,\mathrm{grad}_g\, f, v^{\mathcal{V}}\rangle_g \\&= \langle \nabla_{v^{\mathcal{H}}} \,\mathrm{grad}_g\, f, v^{\mathcal{H}}\rangle_g + \langle \nabla_{v^{\mathcal{V}}} \,\mathrm{grad}_g\, f, v^{\mathcal{H}}\rangle_g \\&= \langle \nabla_{v^{\mathcal{H}}} \,\mathrm{grad}_g\, f, v^{\mathcal{H}}\rangle_g \\&= \nabla^2_g\, f(x)[v^{\mathcal{H}}, v^{\mathcal{H}}], \end{align}\] where the third-last equality follows from the symmetry of the Hessian. Thus, \[\begin{align} \nabla^2_g\, f(x)[v, v] = \nabla^2_g \, f(x)[v^{\mathcal{H}}, v^{\mathcal{H}}] = \langle \nabla^2_g\, f(x)[v^H], v^{\mathcal{H}}\rangle_g &= \langle (\nabla^2_g\, f(x)[v^H])^{\mathcal{H}}, v^{\mathcal{H}}\rangle_g \\ &= \langle \nabla^2_h\, \tilde{f}(\pi(x))[\pi_* v^H], \pi_* v^{\mathcal{H}}\rangle_h \\&= \nabla^2_h\, \tilde{f}(\pi(x))[\pi_* v, \pi_* v], \end{align}\] where the second-last equality follows from 76 and from \(\pi\) being a Riemannian submersion. This expression also allows us to conclude that, for every critical point \(p \in B\) of \(\tilde{f}\), every \(x \in \pi^{-1}(p)\), and every \(v \in T_p B\), \[\nabla^2_h\, \tilde{f}(p)[v, v] = \nabla^2_g\, f(x)[\mathrm{lift}_x\, v, \mathrm{lift}_x\, v].\] ◻

Lastly, under the assumption that the Riemannian submersion \(\pi\) has totally geodesic fibers, we can conclude that the Laplacian of a fiber-invariant function in the total space is the same as that of the projected function in the base space.

Definition 42 (Laplace-Beltrami operator). Given a Riemannian manifold \((M, g)\), let \(f : M \to \mathbb{R}\) be twice differentiable. We define the Laplacian of \(f\) as \[\Delta_g f = \mathrm{div}(\mathrm{grad}_{g}\,f),\] where \[\require{physics} \mathrm{div} X:= \Tr(\nabla X),\] for any vector field \(X \in \mathfrak{X}(M)\). The operator \(\Delta_g\) is the Laplace-Beltrami operator.

Proposition 79. Let \(\pi: (M, g) \to (B, h)\) be a surjective Riemannian submersion with totally geodesic fibers. Let \(\tilde{f} \in C^2(B)\), and let \(f = \tilde{f} \circ \pi\). Then for every \(x \in M\) \[\Delta_g\, f(x) = \Delta_h\, f(\pi(x)).\]

Proof. For every \(x \in M\), we pick an orthonormal basis of \(T_x M = H_x M \oplus V_xM\), given by \(\{H_i\} \cup \{V_i\}\). Then \[\begin{align} \Delta_g\, f(x) = \sum_{V_i \in V_x M} \langle \nabla_{V_i}\, \mathrm{grad}_g\, f, V_i \rangle_g + \sum_{H_i \in H_x M} \langle \nabla_{H_i}\, \mathrm{grad}_g\, f, H_i \rangle_g. \end{align}\] Note that the gradient of \(f\) is horizontal (cf. 74) and so for every \(V_i \in V_xM\) it holds that \[\langle \nabla_{V_i}\, \mathrm{grad}_g\, f, V_i \rangle_g = \langle (\nabla_{V_i}\, \mathrm{grad}_g\, f)^{\mathcal{V}}, V_i \rangle_g = \langle T_{V_i} \,\mathrm{grad}_g\, f, V_i\rangle_g = 0,\] since \(T \equiv 0\) by 22. Therefore, \[\begin{align} \Delta_g\, f(x) = \sum_{H_i \in H_x M} \langle \nabla_{H_i}\, \mathrm{grad}_g \,f, H_i \rangle_g &= \sum_{H_i \in H_x M} \langle \nabla^2_g\, f(x)[H_i], H_i \rangle_g \\&= \sum_{H_i \in H_x M} \langle (\nabla^2_g\, f(x)[H_i])^{\mathcal{H}}, H_i \rangle_g \\&= \sum_{H_i \in H_x M} \langle \mathrm{lift}_x \nabla^2_h\, \tilde{f}(\pi(x))[\pi_* H_i], H_i \rangle_g \\&= \sum_{H_i \in H_x M} \langle \nabla^2_h\, \tilde{f}(\pi(x))[\pi_* H_i], \pi_* H_i \rangle_h \\&= \Delta_h\, \tilde{f}(\pi(x)), \end{align}\] where the third-last equality follows from 76 and the second-last equality follows from \(\pi\) being a Riemannian submersion, which in turn implies that \(\{\pi_* H_i\}_{H_i \in H_x M}\) form an orthonormal basis of \(T_{\pi(x)} B\). ◻

To conclude the section, we briefly discuss how the Lipschitz constants of a function \(f\) that is constant in the fibers of a Riemannian submersion are valid for its projected version \(\tilde{f}\). Before we do so, let us first discuss parallel transport.

Remark 80 (parallel transport). Given some curve \(\gamma\) in \(M\) and two points \(\gamma(t_1), \gamma(t_2)\) along the curve, the parallel transport along \(\gamma\) is a vector space isomorphism between \(T_{\gamma(t_1)} M\) and \(T_{\gamma(t_2)} M\). For every \(x \in M\) and every \(v \in T_xM\), we denote the parallel transport from \(x\) to \(\exp_x(v)\) along the curve \(t \mapsto \exp_x (tv)\) as \(\mathrm{P}_v\) . For an in-depth study of the parallel transport, see for example [39].

Definition 43 (Lipschitz function). Let \((M, g)\) be a compact Riemannian manifold. We say that a function \(f: M \rightarrow \mathbb{R}\) is \(A_1\)-Lipschitz if \[|f(x) - f(y)| \leq A_1 d_g(x,y),\quad \forall x, y \in M.\] Similarly, we say that \(\mathrm{grad}_{g}\,f\) is \(A_2\)-Lipschitz if \[|\mathrm{P}^{-1}_{v}\,\mathrm{grad}_{g}\,f(\exp_x(v)) - \mathrm{grad}_{g}\,f(x)|_g \leq A_2 |v|_g,\quad \forall x \in M,\;\forall v \in T_xM,\] where \(\mathrm{P}^{-1}_v\) denotes the inverse of \(\mathrm{P}_v\).

Lastly, we say that the Hessian \(\nabla^2 f\) is \(A_3\)-Lipschitz if \[\max_{\substack{s \in T_xM\\|s|_g=1}} |\mathrm{P}^{-1}_{v} \circ \nabla^2 f(\exp_x(v)) \circ \mathrm{P}_{v}[s] - \nabla^2 f(x)[s]|_g \leq A_3 |v|_g,\] for every \(x \in M\) and every \(v \in T_xM\).

To relate the Lipschitz constants of a fiber-invariant function \(f\) to those of its projected version, we will use the fact that the parallel transport and the exponential map commute with the differential of a Riemannian submersion.

Lemma 23 ([60]). Let \(\pi:(M, g) \to (B, h)\) be a Riemannian submersion. Then for every \(x \in M\) it holds that \[\pi \circ \exp_x^M = \exp^B_{\pi(x)} \circ\, \pi_*|_{x},\] whenever both sides are defined. Moreover, for every \(x, y \in M\) and every curve \(\gamma\) connecting \(x\) and \(y\), it holds that \[\pi_*|_y \circ \mathrm{P}^{\gamma}_{y \leftarrow x} = \mathrm{P}^{\pi \circ \gamma}_{\pi(y) \leftarrow \pi(x)} \circ \pi_*|_x,\] where \(\mathrm{P}^{\gamma}_{y \leftarrow x}\) denotes the parallel transport in \((M, g)\) from \(x\) to \(y\) along the curve \(\gamma\) and \(\mathrm{P}^{\pi \circ \gamma}_{\pi(y) \leftarrow \pi(x)}\) denotes the parallel transport in \((B, h)\) from \(\pi(x)\) to \(\pi(y)\) along the curve \(\pi \circ \gamma\).

Proposition 81. Let \((M, g)\) be a compact Riemannian manifold, and let \(\pi: (M, g) \to (B, h)\) be a surjective Riemannian submersion. Let \(f \in C^2(M)\) be constant in the fibers of \(\pi\) and assume that it is \(A_1\)-Lipschitz, with an \(A_2\)-Lipschitz gradient and an \(A_3\)-Lipschitz Hessian. Let \(\tilde{f} : B \to \mathbb{R}\) be such that \(f = \tilde{f} \circ \pi\). Then it follows that \(\tilde{f}\) is \(A_1\)-Lipschitz, with an \(A_2\)-Lipschitz gradient and an \(A_3\)-Lipschitz Hessian.

Proof. Let \(p, q \in B\), and let \(x \in \pi^{-1}(p)\), \(y \in \pi^{-1}(q)\) be such that \(d_h(p, q) = d_g(x, y)\). Then \[|\tilde{f}(p) - \tilde{f}(q)| = |f(x) - f(y)| \leq A_1d_g(x, y) = A_1 d_h(p, q).\]

Now, recall that the gradient of \(f\) is \(A_2\)-Lipschitz if \[|\mathrm{P}^{-1}_{v}\,\mathrm{grad}_g\, f(\exp_x(v)) - \mathrm{grad}_g\, f(x)|_g \leq A_2 |v|_g,\] for every \(x \in M\) and every \(v \in T_xM\), where \(\mathrm{P}^{-1}_{v}\) is the inverse of the parallel transport in \(M\) from \(x\) to \(\exp_x(v)\) along \(\gamma(t) = \exp^M_x(t v)\).

We will denote by \(\exp^M\) and \(\exp^B\) the exponential maps in \(M\) and \(B\), respectively.

Let \(p \in B\), \(x \in \pi^{-1}(p)\), \(w \in T_p B\), and let us denote by \(\tilde{\mathrm{P}}^{-1}_{w}\) the inverse of the parallel transport in \(B\) from \(p\) to \(\exp_p^B(v)\) along \(\tilde{\gamma}(t) = \exp^B_p(t w)\). Then relating the gradients of \(f\) and \(\tilde{f}\) via 74 and 64 we can write \[\begin{align} |\tilde{\mathrm{P}}^{-1}_{w}\,\mathrm{grad}_h\, \tilde{f}(\exp^B_p(w)) - \mathrm{grad}_h\, &\tilde{f}(p)|_h \\&= |\tilde{\mathrm{P}}^{-1}_{w}\, \pi_*(\mathrm{grad}_g\, f(\exp^M_x(\mathrm{lift}_x w))) - \pi_*(\mathrm{grad}_g\, f(x))|_h \\&= |\pi_*\, \mathrm{P}^{-1}_{\mathrm{lift}_x w}\, \mathrm{grad}_g\, f(\exp_x^M(\mathrm{lift}_x w)) - \pi_*(\mathrm{grad}_g\, f(x))|_h, \end{align}\] where the second equality follows from 23. Furthermore, as \(\pi\) is a Riemannian submersion, it holds that \[\begin{align} |\pi_*\, \mathrm{P}^{-1}_{\mathrm{lift}_x w}\, \mathrm{grad}_g\, f(\exp_x^M(\mathrm{lift}_x w)) - &\pi_*(\mathrm{grad}_g\, f(x))|_h \\&=|(\mathrm{P}^{-1}_{\mathrm{lift}_x w}\, \mathrm{grad}_g\, f(\exp_x^M(\mathrm{lift}_x w)))^{\mathcal{H}} - (\mathrm{grad}_g\, f(x))^{\mathcal{H}}|_g \\&\leq |\mathrm{P}^{-1}_{\mathrm{lift}_x w}\, \mathrm{grad}_g\, f(\exp_x^M(\mathrm{lift}_x w)) - \mathrm{grad}_g\, f(x)|_g \\&\leq A_2 |\mathrm{lift}_x w|_g \\&= A_2 |w|_h. \end{align}\]

Lastly, the Hessian of \(f\) is \(A_3\) Lipschitz if \[\label{eq:HessianLipschitzcond} \max_{\substack{s \in T_xM\\|s|_g=1}} |\mathrm{P}^{-1}_{v} \circ \nabla^2_g\, f(\exp^M_x(v)) \circ \mathrm{P}_{v}[s] - \nabla^2_g\, f(x)[s]|_g \leq A_3 |v|_g,\tag{53}\] for every \(x \in M\) and every \(v \in T_xM\).

Let \(p \in B\), \(x\in \pi^{-1}(p)\), and \(v \in H_x M\) be some horizontal tangent vector at \(x\). We can lower bound 53 by \[\begin{align} A_3 |\pi_* v|_h = A_3 |v|_g &\geq \max_{\substack{s \in T_xM\\|s|_g=1}} |\mathrm{P}^{-1}_{v} \circ \nabla^2_g\, f(\exp^M_x(v)) \circ \mathrm{P}_{v}[s] - \nabla^2_g\, f(x)[s]|_g \\&\geq \max_{\substack{s \in T_xM\\|s|_g=1}} |(\mathrm{P}^{-1}_{v} \circ \nabla^2_g\, f(\exp^M_x(v)) \circ \mathrm{P}_{v}[s])^{\mathcal{H}} - (\nabla^2_g\, f(x)[s])^{\mathcal{H}}|_g \\&= \max_{\substack{s \in T_xM\\|s|_g=1}} |\pi_*(\mathrm{P}^{-1}_{v} \circ \nabla^2_g\, f(\exp^M_x(v)) \circ \mathrm{P}_{v}[s]) - \pi_* (\nabla^2_g\, f(x)[s])|_h, \end{align}\] where the last equality follows from the fact that \(\pi\) is a Riemannian submersion. We can use 76 and 23 to conclude that \[\begin{align} \max_{\substack{s \in T_xM\\|s|_g=1}} &|\pi_*(\mathrm{P}^{-1}_{v} \circ \nabla^2_g\, f(\exp^M_x(v)) \circ \mathrm{P}_{v}[s]) - \pi_* (\nabla^2_g f(x)[s])|_h \\&= \max_{\substack{s \in T_xM\\|s|_g=1}} |\tilde{\mathrm{P}}^{-1}_{\pi_* v} \circ \pi_* \circ \nabla^2_g\, f(\exp^M_x(v)) \circ \mathrm{P}_{v}[s] - \nabla^2_h\, \tilde{f}(\pi(x))[\pi_* s])|_h \\&= \max_{\substack{s \in T_xM\\|s|_g=1}} |\tilde{\mathrm{P}}^{-1}_{\pi_* v} \circ \nabla^2_h\, \tilde{f}(\exp^B_p(\pi_* v))[\pi_* \mathrm{P}_{v}(s)] - \nabla^2_h\, \tilde{f}(\pi(x))[\pi_* s])|_h \\&= \max_{\substack{s \in T_xM\\|s|_g=1}} |\tilde{\mathrm{P}}^{-1}_{\pi_* v} \circ \nabla^2_h\, \tilde{f}(\exp^B_p(\pi_* v)) \circ \tilde{\mathrm{P}}_{\pi_* v}[\pi_* s] - \nabla^2_h \tilde{f}(\pi(x))[\pi_* s])|_h. \end{align}\]

Putting everything together, we obtain that \[\max_{\substack{s \in T_xM\\|s|_g=1}} |\tilde{\mathrm{P}}^{-1}_{\pi_* v} \circ \nabla^2_h\, \tilde{f}(\exp^B_p(\pi_* v)) \circ \tilde{\mathrm{P}}_{\pi_* v}[\pi_* s] - \nabla^2_h\, \tilde{f}(p)[\pi_* s])|_h \leq A_3 |\pi_* v|_h,\] for every \(p \in B\), \(x \in \pi^{-1}(p)\), and \(v \in H_x M\). Since there is a one-to-one correspondence between horizontal vectors in \(T_xM\) and tangent vectors in \(T_p B\), the result follows. ◻

8.6 Periodic geodesics, diameter, dimension and volume↩︎

In this subsection, we will first give upper bounds for the curvature of the round sphere. Then we will give lower bounds for the injectivity and convexity radii of the sphere, the unitary group, as well as the Stiefel and Grassmann manifolds via the study of their shortest non-trivial periodic geodesic and the upper bounds on their sectional curvature. Lastly, we will bound their diameter and volume and obtain an expression for their dimension.

First, the sectional and Ricci curvatures of the sphere endowed with the round metric are widely known. Let us therefore state the following result without proof.

Proposition 82 ([39]). The sphere \(\mathbb{S}^n\) endowed with the round metric has constant sectional curvature \(1\). Furthermore, it is an Einstein manifold with constant \((n-1)\).

This result, together with the bounds found for the sectional and Ricci curvatures of the unitary and special unitary groups—see 55 56—and the Stiefel and Grassmann manifolds—see 72 73—are summarized in 1.

Table 1: Summary of the various bounds obtained for the sectional and Ricci curvatures of the manifolds studied.
Manifold Metric Sectional curvature Ricci curvature
\(\textup{U}(n)\) Bi-invariant \(\leq 1/2\) \(\geq 0\)
\(\textup{SU}(n)\) Bi-invariant \(\leq 1/2\) \(= n/2\)
\(\bb{S}^n\) Round \(= 1\) \(= n-1\)
\(\textup{V}_k(\bb{C}^n)\) Induced \(\leq 2\) \(\geq 0\)
\(\textup{Gr}_k(\bb{C}^n)\) Induced \(\leq 2\) \(\geq 0\)

8.6.1 Bounding the injectivity and convexity radii↩︎

In order to obtain lower bounds on the injectivity and convexity radii in the manifolds of our interest, we saw in 8.3 that it is useful to have lower bounds on the length of their shortest non-trivial periodic geodesic.

Let us first briefly discuss the length of the shortest non-trivial periodic geodesic in a product manifold. This is in fact a rather easy task: given two Riemannian manifolds, \((M, g)\) and \((N, h)\), every geodesic of \(M \times N\) endowed with the product metric can be written as \[\gamma(t) = (\gamma_1(t), \gamma_2(t)),\] where \(\gamma_1(t)\) and \(\gamma_2(t)\) are geodesics in \(M\) and \(N\), respectively. For this reason, the length of the shortest non-trivial periodic geodesic in \(M \times N\), is \[l(M\times N) = \min \{l(M), l(N)\}.\] Certainly, the shortest non-trivial periodic geodesic of the product manifold corresponds to fixing a point for either \(M\) or \(N\), and considering the shortest non-trivial periodic geodesic of \(N\) or \(M\), respectively.

Let us now lower bound the length of the shortest non-trivial periodic geodesics in the manifolds of our interest. The lower bounds that will be derived are summarized in 2 along with the upper bounds for the diameter of the manifolds, which will be discussed in 8.6.2.

Table 2: Summary of various bounds for the diameter and length of the shortest non-trivial periodic geodesic of the manifolds that are studied in [sec:sec:traceratio].
Manifold Metric \(l(M)\) Diameter
\(\textup{U}(n)\) Bi-invariant \(\geq 2\pi\) \(\leq \pi \sqrt{n} (\sqrt{2}+1)\)
\(\bb{S}^n\) Round \(= 2\pi\) \(\leq \pi\)
\(\textup{V}_k(\bb{C}^n)\) Induced \(\geq 2\sqrt{2}\pi\) \(\leq \pi \sqrt{n} (\sqrt{2}+1)\)
\(\textup{Gr}_k(\bb{C}^n)\) Induced \(\geq \sqrt{2}\pi\) \(\leq \pi \sqrt{n} (\sqrt{2}+1)\)

The length of the shortest non-trivial periodic geodesic on the sphere when endowed with the round metric is known to be \(2\pi\) [39], as geodesics on the sphere correspond to great circles. Let us now compute the length of the shortest non-trivial periodic geodesic in \(\mathrm{U}(n)\) endowed with the bi-invariant metric.

Proposition 83. The length of the shortest non-trivial periodic geodesic of \(\mathrm{U}(n)\) endowed with the bi-invariant metric is bounded from below by \(2\pi\).

Proof. Let us fix \(X \in \mathfrak{u}(n)\) with norm one, i.e. \(\require{physics} -\Tr(X^2) = 1\), and a geodesic starting in \(U \in \mathrm{U}(n)\) in the direction of \(X\), i.e. \(\gamma(t) = e^{tX}U\). We want to find \[\min\{t > 0 : e^{tX}U = U\} = \min\{t > 0 : e^{tX} = \mathbb{1}\}.\]

Since \(X \in \mathfrak{u}(n)\) we know that \(X^\dagger = -X\), therefore all its eigenvalues are imaginary. Furthermore, \(X\) is normal (i.e. \(X^\dagger X = X X^\dagger\)). Therefore, without loss of generality, we can consider \(X\) to be in its diagonal form \[\Sigma = \begin{bmatrix} i\lambda_1 & & & \\ & i\lambda_2 & & \\ & & \ddots & \\ & & & i\lambda_n \end{bmatrix},\] where \(\lambda_j \in \mathbb{R}\) for every \(j \in \{1, \dotsc, n\}\). Otherwise, there would exist a unitary matrix \(V\) such that \(X = V \Sigma V^\dagger\) and \(e^{tX} = V e^{t\Sigma} V^\dagger\), which yields \[e^{tX} = \mathbb{1} \iff V e^{t\Sigma} V^\dagger = \mathbb{1} \iff e^{t\Sigma} = \mathbb{1}.\] Now, \[e^{tX} = \begin{bmatrix} e^{it\lambda_1} & & \\ & e^{it\lambda_2} & & \\ & & \ddots &\\ & & & e^{it\lambda_n} \end{bmatrix},\] where all the off-diagonal elements are zero. In order for \(e^{tX}\) to be the identity it has to hold that \[e^{it\lambda_j} = 1,\quad t > 0, \text{ for each } j \in \{1, \dotsc, n\},\] which implies that \[t\lambda_j = 2\pi k_j\quad \text{for some } k_j \in \mathbb{Z},\quad \text{for every }j \in \{1,\dotsc,n\}.\] Since for every eigenvalue \(i\lambda_j\) it holds that \(|\lambda_j|^2 \leq 1\), we can conclude that \(t\) has to be greater than \(2\pi\) and so \(l(\mathrm{U}(n)) \geq 2\pi\). ◻

Let us now finish by obtaining a lower bound on the length of the shortest non-trivial periodic geodesic on the Stiefel and Grassmann manifolds when endowed with their induced metrics.

Let \(g\) be the bi-invariant metric of \(\mathrm{U}(n)\). The length of the shortest non-trivial periodic geodesic on the Stiefel manifold \(\mathrm{V}_k(\mathbb{C}^n)\) when endowed with the metric \(h\) making \(\pi: (\mathrm{U}(n), g) \to (\mathrm{V}_k(\mathbb{C}^n), h)\) a Riemannian submersion is known to be lower bounded by \(2\sqrt{2}\pi\) [70], [71].

Finally, let us obtain a lower bound for the shortest non-trivial periodic geodesic of Grassmann manifolds.

Proposition 84. Consider \(\mathrm{U}(n)\) endowed with the bi-invariant metric, and let \(\mathrm{Gr}_k(\mathbb{C}^n)\) be endowed with the metric making the quotient map \(\pi: \mathrm{U}(n) \to \mathrm{Gr}_k(\mathbb{C}^n)\) a Riemannian submersion. The length of its shortest non-trivial periodic geodesic is lower bounded by \(\sqrt{2}\pi\).

Proof. Let us fix some \(P \in \mathrm{Gr}_k(\mathbb{C}^n)\), and some unitary tangent vector \(X \in T_P \mathrm{Gr}_k(\mathbb{C}^n)\). It is known (cf. [56]) that \(X\) corresponds to a Hermitian matrix such that \(X = PX + XP\) and that the geodesic starting at \(P\) in the direction \(X\) is given by \[\gamma_{P, X}(t) := e^{t[X, P]}P e^{-t[X, P]}.\]

Without loss of generality, we may assume that \(P = \begin{pmatrix} \mathbb{1}_k & 0 \\ 0 & 0_{n-k} \end{pmatrix}\), and so the condition \(X = XP + PX\) implies that \(X\) is of the form \[X = \begin{pmatrix} 0_k & A\\ A^\dagger & 0_{n-k} \end{pmatrix},\] where \(A\) is an \(k \times (n-k)\)-complex matrix.

Standard calculations allow us to conclude that \[\begin{align} [X, P]^{2k} &= \begin{pmatrix} (-1)^k (AA^\dagger)^k & 0\\ 0 & (-1)^k (A^\dagger A)^k \end{pmatrix},\\ [X, P]^{2k+1} &= \begin{pmatrix} 0 & (-1)^{k+1}(A A^\dagger)^k A\\ (-1)^{k}(A^\dagger A)^k A^\dagger & 0 \end{pmatrix}, \end{align}\] for every \(k \in \mathbb{N}\cup\{0\}\). Thus, \[\begin{align} e^{t[X, P]} &= \begin{pmatrix} \cos(t\sqrt{AA^\dagger}) & -\frac{\sin(t\sqrt{AA^\dagger})}{\sqrt{AA^\dagger}} A\\ -\frac{\sin(t\sqrt{A^\dagger A})}{\sqrt{A^\dagger A}} A^\dagger & \cos(t\sqrt{A^\dagger A}) \end{pmatrix}. \end{align}\]

This way, in order for \(P = e^{t[X, P]} P e^{-t[X, P]}\) to hold, it must be the case that \[\label{eq:identitycosines} \cos^2(t\sqrt{AA^\dagger}) = \mathbb{1}_k.\tag{54}\]

Let us assume without loss of generality that \(A A^\dagger\) is diagonal with eigenvalues \(\{\lambda_i\}_{i = 1}^k\). If 54 holds, every \(\lambda_i\) must be non-negative and \[\label{eq:conditionclosedgeodGr} t\sqrt{\lambda_i} = \pi k,\quad k \in \mathbb{Z},\tag{55}\] for every \(i \in \{1, \dotsc, k\}\).

Finally, note that, as \(X\) has norm one, \[\require{physics} 1 = \Tr(X^\dagger X) = \Tr(AA^\dagger) + \Tr(A^\dagger A) = 2\sum_{i = 1}^k \lambda_i,\] and so, as the eigenvalues are assumed to be non-negative, it holds that \(\lambda_i \leq \frac{1}{2}\) for every \(i \in \{1, \dotsc, k\}\), and \(\sqrt{\lambda_i} \leq \frac{\sqrt{2}}{2}\). Thus, in order for 55 to hold, it must be the case that \[t \geq \frac{2\pi}{\sqrt{2}} = \sqrt{2}\pi,\] finishing the proof. ◻

The lower bounds on the length of the shortest non-trivial periodic geodesic and the upper bounds on the sectional curvature of the manifolds considered allow us to obtain bounds on their injectivity and convexity radii.

Corollary 7. Consider the manifolds \(\mathrm{U}(n)\), \(\mathbb{S}^n\), \(\mathrm{V}_k(\mathbb{C}^n)\) and \(\mathrm{Gr}_k(\mathbb{C}^n)\) endowed with the metrics shown in 2. Then their injectivity and convexity radii can be lower bounded as follows; \[\begin{align} i(\mathrm{U}(n)) &\geq \pi,&\quad \mathit{conv}(\mathrm{U}(n)) &\geq \frac{\pi}{2},\\ i(\mathbb{S}^n) &\geq \pi, &\quad \mathit{conv}(\mathbb{S}^n) &\geq \frac{\pi}{2},\\ i(\mathrm{V}_k(\mathbb{C}^n)) &\geq \frac{\sqrt{2}\pi}{2}, &\quad \mathit{conv}(\mathrm{V}_k(\mathbb{C}^n)) &\geq \frac{\sqrt{2}\pi}{4}, \\ i(\mathrm{Gr}_k(\mathbb{C}^n)) &\geq \frac{\sqrt{2}\pi}{2},&\quad \mathit{conv}(\mathrm{Gr}_k(\mathbb{C}^n)) &\geq \frac{\sqrt{2}\pi}{4}.\\ \end{align}\]

Proof. To obtain the bounds, we can use 61 62, along with the bounds on the sectional curvature and the length of their shortest non-trivial periodic geodesic shown in 1 2.

For the unitary group endowed with the bi-invariant metric, we obtained that \[K \leq \frac{1}{2},\quad \text{and}\quad l(\mathrm{U}(n)) \geq 2\pi.\] Thus \[i(\mathrm{U}(n)) \geq \min\{\sqrt{2}\pi, \pi\} = \pi,\quad \mathrm{and}\quad \mathit{conv}(\mathrm{U}(n)) \geq \frac{\pi}{2}.\]

For the sphere endowed with the round metric, we showed that \[K \leq 1,\quad \mathrm{and}\quad l(\mathbb{S}^n) \leq 2\pi,\] so \[i(\mathbb{S}^n) \geq \min\{\pi, \pi\} = \pi,\quad \mathrm{and}\quad \mathit{conv}(\mathbb{S}^n) \geq \frac{\pi}{2}.\]

For the Stiefel manifold \(\mathrm{V}_k(\mathbb{C}^n)\) endowed with the metric \(h\) making \(\pi: (\mathrm{U}(n), g) \to (\mathrm{V}_k(\mathbb{C}^n), h)\) a Riemannian submersion, where \(g\) is the bi-invariant metric of \(\mathrm{U}(n)\), we proved that \[K \leq 2,\quad \mathrm{and}\quad l(\mathrm{V}_k(\mathbb{C}^n)) \geq 2\sqrt{2}\pi,\] which implies that \[i(\mathrm{V}_k(\mathbb{C}^n)) \geq \min\{\frac{\sqrt{2}\pi}{2}, \sqrt{2}\pi\} = \frac{\sqrt{2}\pi}{2},\quad \mathrm{and}\quad \mathit{conv}(\mathrm{V}_k(\mathbb{C}^n)) \geq \frac{\sqrt{2}\pi}{4}.\]

Lastly, for the Grassmann manifold also endowed with its induced metric, we showed that \[K \leq 2,\quad \mathrm{and}\quad l(\mathrm{Gr}_k(\mathbb{C}^n)) \geq \sqrt{2}\pi,\] yielding \[i(\mathrm{Gr}_k(\mathbb{C}^n)) \geq \min\{\frac{\sqrt{2}\pi}{2}, \frac{\sqrt{2}\pi}{2}\} = \frac{\sqrt{2}\pi}{2},\quad \mathrm{and}\quad \mathit{conv}(\mathrm{Gr}_k(\mathbb{C}^n)) \geq \frac{\sqrt{2}\pi}{4}.\] ◻

8.6.2 Controlling the diameter↩︎

We will now derive the diameter bounds shown in 2. In this subsection we will also show how the diameter of a product manifold is related to that of its components, and how the diameters of the total and the base space of a Riemannian submersion are related.

Proposition 85. Let \(M\) be a compact Riemannian manifold and let \(\pi : (M, g) \to (B, h)\) be a Riemannian submersion. Denote by \(d_g\) and \(d_h\) the geodesic distances on \(M\) and \(B\), respectively. Then for every \(x, y \in M\) \[d_g(x, y) \geq d_h(\pi(x), \pi(y)).\] In particular, if \(\pi\) is surjective, \(\mathrm{diam}(B) \leq \mathrm{diam}(M)\).

Proof. Since \(M\) is compact, the distance between \(x\) and \(y\) is attained by a geodesic. We denote such a geodesic by \(\gamma\), and assume that \(\gamma(0) = x\), \(\gamma(1) = y\). Then \[\begin{align} d_g(x, y) = \int_0^1 |\gamma'(s)|_g\, ds \geq \int_0^1 |(\gamma')(s)^\mathcal{H}|_g\, ds = \int_0^1 |\pi_* (\gamma')(s)^\mathcal{H}|_h\, ds \geq d_h(\pi(x), \pi(y)), \end{align}\] concluding the proof. ◻

Lastly, the diameter of a product manifold endowed with the product metric can be easily studied, by simply obtaining an expression for the distance function in the product manifold.

Proposition 86 (Diameter of the product). Given two Riemannian manifolds \((M, g)\) and \((N, h)\), if we consider the product manifold \((M \times N, g\oplus h)\) endowed with the product metric, the distance induced by this metric is such that for every \((x_1,y_1), (x_2, y_2) \in M \times N\), \[d_{g\oplus h}((x_1,y_1), (x_2, y_2)) = \sqrt{d^2_g(x_1, x_2) + d^2_h(y_1,y_2)},\] and so \[\mathrm{diam}(M\times N) = \sqrt{\mathrm{diam}(M)^2 + \mathrm{diam}(N)^2}.\]

With 86 85 in mind, we can now bound the diameter of the manifolds of our interest. We start by establishing upper bounds on the diameter of both the unitary group and the round sphere. To this end, we will use the following result.

Proposition 87 (Bonnet-Myers). Let \((M, g)\) be a compact, connected Riemmanian manifold of dimension \(n\), and suppose that there is a positive constant \(c = \frac{1}{R^2}\) such that the Ricci curvature of \(M\) satisfies \(\mathrm{Ric}_g(v,v) \geq (n-1)c\) for all unit vectors \(v\). Then the diameter of \(M\) is less than or equal to \(\pi R\).

This result allows us to upper bound the diameter of the round sphere and \(\mathrm{SU}(n)\).

Corollary 8. Let \(\mathbb{S}^n\) be endowed with the round metric, and let \(\mathrm{SU}(n)\) be endowed with the bi-invariant metric. Then \[\mathrm{diam}(\mathbb{S}^n) \leq \pi,\quad \mathrm{\textit{and}}\quad \mathrm{diam}(\mathrm{SU}(n)) \leq \pi \sqrt{2n}.\]

Proof. First, using 82 we know that the round sphere \(\mathbb{S}^n\) is an Einstein manifold with constant \(n-1\), and so it satisfies the Bonnet-Myers theorem with constant \(c = 1\).

For the special unitary group, recall from 56 that \(\mathrm{SU}(n)\) is an Einstein manifold with constant \(\frac{n}{2}\). Moreover, \(\mathrm{SU}(n)\) is connected, compact and has dimension \(n^2-1\). Therefore \(\mathrm{Ric}(v,v) = ((n^2-1) -1)c\) with \[c = \frac{n}{2(n^2 -2)}.\] This way, \(R = \sqrt{\frac{2(n^2-2)}{n}}\) and so we can conclude that \[\begin{align} \mathrm{diam}(\mathrm{SU}(n)) &\leq \pi \sqrt{\frac{2(n^2-2)}{n}} \leq \pi \sqrt{2n}. \end{align}\] ◻

Having obtained a bound on the diameter of \(\mathrm{SU}(n)\), we can now upper bound the diameter of \(\mathrm{U}(n)\) when endowed with its bi-invariant metric.

Corollary 9. Let \(\mathrm{U}(n)\) be endowed with the bi-invariant metric. Then \[\mathrm{diam}(\mathrm{U}(n)) \leq \pi\sqrt{n}(1+ \sqrt{2}).\]

Proof. Let \(A, B \in \mathrm{U}(n)\). We can rewrite \(A\) and \(B\) as \(A = n_A h_A\), \(B = n_B h_B\), where \(n_A, n_B \in \mathrm{U}(1)\) and \(h_A, h_B \in \mathrm{SU}(n)\). Using this decomposition we can write \[d_{\mathrm{U}(n)}(A, B) = d_{\mathrm{U}(n)}(n_A h_A, n_B h_B).\] Moreover, since the metric of \(\mathrm{U}(n)\) is bi-invariant, it follows that \[d_{\mathrm{U}(n)}(n_A h_A, n_B h_B) = d_{\mathrm{U}(n)}(h_A, \overline{n}_A n_B h_B).\] Now, applying the triangular inequality \[d_{\mathrm{U}(n)}(h_A, \overline{n}_A n_B h_B) \leq d_{\mathrm{U}(n)}(h_A, h_B) + d_{\mathrm{U}(n)}(h_B, \overline{n}_A n_B h_B).\]

Using the fact that \(d_{\mathrm{U}(n)}(h_A, h_B) \leq d_{\mathrm{SU}(n)}(h_A, h_B)\) and the bi-invariance of the metric again, \[\begin{align} d_{\mathrm{U}(n)}(h_A, h_B) + d_{\mathrm{U}(n)}(h_B, \overline{n}_A n_B h_B) &\leq d_{\mathrm{SU}(n)}(h_A, h_B) + d_{\mathrm{U}(n)} (h_B, \overline{n}_A n_B h_B) \\&\leq \mathrm{diam}(\mathrm{SU}(n)) + d_{\mathrm{U}(n)}(\mathbb{1}, \overline{n}_A n_B). \end{align}\]

Finally, since \(n_A, n_B \in \mathrm{U}(1)\) and \(d_{\mathrm{U}(n)}(\mathbb{1}, \overline{n}_A n_B) = \sqrt{n}d_{\mathrm{U}(1)}(1, \overline{n}_A n_B)\), \[\begin{align} \mathrm{diam}(\mathrm{SU}(n)) + d_{\mathrm{U}(n)}(\mathbb{1}, \overline{n}_A n_B) &\leq \mathrm{diam}(\mathrm{SU}(n)) + d_{\mathrm{U}(1)}(1, \overline{n}_A n_B) \\&\leq \mathrm{diam}(\mathrm{SU}(n)) + \pi\sqrt{n}, \end{align}\] where the last inequality follows from \(\mathrm{U}(1) \simeq \mathbb{S}^1\) having a diameter upper bounded by \(\pi\). ◻

Lastly, we can apply 85 to obtain a bound on the diameter of \(\mathrm{V}_k(\mathbb{C}^n)\) and \(\mathrm{Gr}_k(\mathbb{C}^n)\) when endowed with their induced metrics. Indeed, it follows automatically that 15 \[\begin{align} \label{eq:diamaterboundstiefelgrassmann} \begin{aligned} \mathrm{diam}(\mathrm{V}_k(\mathbb{C}^n)) &\leq \mathrm{diam}(\mathrm{U}(n)) \leq \pi\sqrt{n}(1+ \sqrt{2}),\\ \mathrm{diam}(\mathrm{Gr}_k(\mathbb{C}^n)) &\leq \mathrm{diam}(\mathrm{U}(n)) \leq \pi\sqrt{n}(1+ \sqrt{2}). \end{aligned} \end{align}\tag{56}\]

8.6.3 Dimension and volume↩︎

Lastly, let us briefly discuss the dimension and volume of the manifolds studied in this section, when endowed with the metrics shown in 2.

8.6.3.1 Sphere.

The dimension of the \(n\)-sphere \(\mathbb{S}^n\) is \(n\). Its volume when endowed with the round metric is known to be \[\mathrm{Vol}_{g_{\textit{round}}}(\mathbb{S}^n) = \frac{2\pi^{(n+1)/2}}{\Gamma(\frac{n+1}{2})},\] (cf. [39]).

8.6.3.2 Unitary group.

The unitary group \(\mathrm{U}(n)\) is a Lie group of dimension \(n^2\). Furthermore, it is known (cf. [72]) that when endowed with the bi-invariant metric, its volume is \[\mathrm{Vol}(\mathrm{U}(n)) = \frac{(2\pi)^{n(n+1)/2}}{\prod_{k = 1}^{n-1} k!}.\]

Now, in order to study Stiefel and Grassmann manifolds, we will make use of two formulae which allow us to relate their dimension and volumes to those of the total space and the fibers. Given a Riemannian submersion \(\pi: (M, g) \to (M/G, h)\), it holds that \[\begin{align} \dim(M/G) &= \dim(M) - \dim(G),\\ \mathrm{Vol}(M) &= \mathrm{Vol(G)} \mathrm{Vol}(M/G). \end{align}\] Note that the volume formula can be inferred using the Fubini formula from 3.1.

8.6.3.3 Stiefel manifold.

Using the formulae for the volume and the dimension of quotient spaces, it is easy to see that the dimension of the Stiefel manifold \(\mathrm{V}_{k}(\mathbb{C}^{n})\) is \[n^2 - (n-k)^2 = 2nk - k^2.\] Moreover, when endowed with the metric \(h\) which makes \(\pi: (\mathrm{U}(n), g) \to (\mathrm{U}(n)/\mathrm{U}(n-k), h)\) a Riemannian submersion—where \(g\) is the bi-invariant metric of \(\mathrm{U}(n)\)—its volume is given by \[\begin{align} \mathrm{Vol}(\mathrm{V}_{k}(\mathbb{C}^{n})) &= \mathrm{Vol}(\mathrm{U}(n))/ \mathrm{Vol}(\mathrm{U}(n-k)) \\&= \frac{(2\pi)^{nk - \frac{k(k-1)}{2}}}{\prod_{a = n-k}^{n-1}a!}. \end{align}\]

8.6.3.4 Grassmann manifold.

Similarly, the dimension of the Grassmann manifold \(\mathrm{Gr}_k (\mathbb{C}^n)\) is \[n^2 - (n-k)^2 - k^2 = 2k(n - k),\] and its volume, when endowed with its corresponding induced metric is given by \[\begin{align} \mathrm{Vol}(\mathrm{Gr}_{k}(\mathbb{C}^{n})) &= \mathrm{Vol}(\mathrm{U}(n))/ \mathrm{Vol}(\mathrm{U}(n-k) \times U(k)) \\&= \frac{(2\pi)^{nk - k^2} \prod_{a = 1}^{k-1} a!}{\prod_{a = n-k}^{n-1}a!}. \end{align}\]

9 Sobolev spaces on manifolds↩︎

In this section, we will introduce some basic notions about Sobolev spaces on manifolds, which provide the natural setting for the Poincaré and log-Sobolev inequalities. We will follow the structure of Alexander Grigor’yan’s notes [73]. Let us first define weighted manifolds.

Definition 44 (Weighted manifold). Let \((M, g)\) be a Riemannian manifold, and let \(D(x)\) be a smooth and positive function on \(M\), also known as a density function. Let us define the measure \(\mu\) on \(M\) as \[d\mu(x) := D(x)\mathrm{dVol}_g(x).\] We say \((M, g, \mu)\) is a weighted manifold.

Since \((M, \mu)\) is a measure space, it is straightforward to define the space of square-integrable functions on \(M\) with respect to \(\mu\), denoted as \(L^2(M, \mu)\). We also denote by \(L^2_{\mathit{loc}}(M, \mu)\) the space of measurable functions on \(M\) such that \(f \in L^2(V, \mu)\) for every \(V \subset M\) compact.

It is a well-known result that, whenever we consider a measure space endowed with a finite measure, the spaces \(L^p\) are nested.

Lemma 24. Let \((E, \mu)\) be a measure space with finite measure \(\mu\). Let \(p, q \in [1, \infty]\), be such that \(p > q\). Then \(L^p(E, \mu) \subset L^q(E, \mu)\).

Let us now introduce an analogous of the space \(L^2(M, \mu)\) for vector fields on a weighted manifold \((M, g, \mu)\).

Definition 45 (Spaces \(\overrightarrow{L}^2(M, \mu)\) and \(\overrightarrow{L}^2_{\mathit{loc}}(M, \mu)\)). Given a weighted manifold \((M, g, \mu)\), we denote by \(\overrightarrow{L}^2(M, \mu)\) the space of all measurable vector fields \(X \in \mathfrak{X}(M)\) such that \(|X|_g \in L^2(M, \mu)\).

We can define an inner product in \(\overrightarrow{L}^2(M, \mu)\) given as \[(v, w)_{\overrightarrow{L}^2(M, \mu)} := \int_M \langle v, w\rangle_g\; d\mu,\] for every pair of vector fields \(v, w \in \mathfrak{X}(M)\). Its corresponding norm is given as \[\norm{v}^2_{\overrightarrow{L}^2(M, \mu)} = \int_M |v|^2_g\; d\mu.\]

The space \(\overrightarrow{L}^2_{\mathit{loc}}(M, \mu)\) is defined as the space of all vector fields \(X \in \mathfrak{X}(M)\) such that \(|X|_g \in L^2_{\mathit{loc}}(M, \mu)\).

Before we define Sobolev spaces on a weighted manifold \((M, g, \mu)\), we need to define the notion of weak gradient.

Definition 46 (Weak gradient). Let \(u : M \to \mathbb{R}\) be such that \(u \in L^2_{\mathit{loc}}(M, \mu)\). A weak gradient of \(u\) is a vector field \(v \in \overrightarrow{L}^2_{\mathit{loc}}(M, \mu)\)—also denoted \(\mathrm{grad}_{g}\, u\)—such that, for any smooth vector field \(\psi\) on \(M\) with compact support, it holds that \[\label{defweakgradient} \int_M u\, \mathrm{div}_{g, \mu} \psi\; d\mu = -\int_M \langle v, \psi\rangle_g\; d\mu,\qquad{(12)}\] where \[\label{divergencemeasure} \mathrm{div}_{g, \mu} X := \frac{1}{D} \mathrm{div} (DX),\quad \forall X \in \mathfrak{X}(M).\qquad{(13)}\]

The weak gradient of a given function \(u\) is uniquely defined in \(\overrightarrow{L}^2_{\mathit{loc}}(M)\). Moreover, if \(u\) is assumed to be smooth, the usual gradient \(\mathrm{grad}_{g}\,u\) coincides with the weak gradient. Indeed, the left hand side of ?? can be rewritten using the following proposition:

Proposition 88. Given a smooth function \(f: M \rightarrow \mathbb{R}\) and a smooth vector field \(X \in \mathfrak{X}(M)\), the following identity holds \[\mathrm{div}(fX) = f\mathrm{div}(X) + X(f).\]

Applying the divergence theorem—which we state below—?? holds.

Proposition 89 (Divergence theorem). Let \((M, g)\) be an oriented Riemannian manifold with boundary and let \(X \in \mathfrak{X}(M)\) be a compactly supported smooth vector field, then \[\int_{M} \mathrm{div}_g(X)\, \mathrm{dVol}_g = \int_{\partial M} g(X,\hat{n})\, \mathrm{dVol}_{\tilde{g}},\] where \(\hat{n}\) is the unit normal vector field to \(\partial M\) and \(\tilde{g}\) is the induced metric on the boundary.

Note that the weak gradient of a function \(u\) is independent of the measure considered. Indeed, given a weighted manifold \((M, g, \mu)\), recall that for a given locally square-integrable function \(u\) on \((M, g, \mu)\), its weak gradient is defined as the unique locally square-integrable vector field such that, for any smooth compactly supported vector field \(\psi\) on \(M\), it holds that \[\int_M u\, \frac{1}{D}\mathrm{div}(D \psi) D\; \mathrm{dVol}_g = -\int_M \langle v, \psi\rangle_g\, D\; \mathrm{dVol}_g,\] which can be rewritten as \[\int_M u\, \mathrm{div}(D \psi) \mathrm{dVol}_g = -\int_M \langle v, D\psi\rangle_g\; \mathrm{dVol}_g.\] Since this identity is satisfied by every smooth and compactly supported vector field on \(M\), it suffices to consider \(\tilde{\psi} := D \psi\) to conclude that the definition is independent of the measure \(\mu\).

Having defined the weak gradient of a locally square-integrable function, we can define the Sobolev space \(H^1\).

Definition 47 (Sobolev space \(H^1\)). We denote by \(H^1(M, \mu)\) the Sobolev space \[H^1(M, \mu) := \left\{u \in L^2(M, \mu) : \mathrm{grad}_{g}\,u \in \overrightarrow{L}^2(M, \mu) \right\},\] which can be endowed with the following inner product \[(u, v)_{H^1(M, \mu)} := (u, v)_{L^2(M, \mu)} + (\mathrm{grad}_{g}\,u, \mathrm{grad}_{g}\,v)_{\overrightarrow{L}^2(M, \mu)},\quad \forall u, v \in H^1(M, \mu).\] The associated norm is given by \[\norm{u}^2_{H^1(M, \mu)} = \norm{u}^2_{L^2(M, \mu)} + \norm{\mathrm{grad}_{g}\,u}^2_{\overrightarrow{L}^2(M, \mu)}.\]

The Sobolev space \(H^1(M, \mu)\) endowed with the inner product described above is in fact a Hilbert space [73]. Moreover, whenever \(\mu\) is a finite measure on \(M\), it follows from 24 that \(H^1(M, \mu) \subset L^1(M, \mu)\).

Note that both the Sobolev and \(L^2\) spaces have been defined for a general weighted manifold \((M, g, \mu)\). In particular, given any open set \(U \subset M\) and some measure \(\tilde{\mu}\) on \(U\), we can define \(L^2(U, \tilde{\mu})\) and \(H^1(U, \tilde{\mu})\).

Notation 90. When the measure \(\mu\) considered is the one given by the volume form associated with \(g\)—i.e. \(D(x) \equiv 1\)—we will simply write \(L^2(M)\), \(L^2_{\mathit{loc}}(M)\), \(\overrightarrow{L}^2(M)\), \(\overrightarrow{L}^2_{\mathit{loc}}(M)\) and \(H^1(M)\) for simplicity.

Whenever \(M\) is compact, the following holds.

Remark 91. Let \((M, g, \mu)\) be a compact weighted manifold, and let \(U\) be a submanifold of \(M\) endowed with the restricted measure \(\mu\) on \(U\). Then

  1. \(L^2(U, \mu|_U) = L^2(U),\)

  2. \(L^2_{\mathit{loc}}(U, \mu|_U) = L^2_{\mathit{loc}}(U),\)

  3. \(\overrightarrow{L}^2(U, \mu|_U) = \overrightarrow{L}^2(U),\)

  4. \(\overrightarrow{L}^2_{\mathit{loc}}(U, \mu|_U) = \overrightarrow{L}^2_{\mathit{loc}}(U).\)

Proof. Let us only prove \(1\), as it implies the rest of the equalities. Recall that \(d\mu(x) = D(x) \mathrm{dVol}_g(x)\), where \(D(x)\) is assumed to be smooth and positive. Thus, since \(M\) is compact, there exist two constants \(k_1\) and \(k_2\) such that \[0 < k_1 \leq D(x) \leq k_2 < \infty,\] and so \[k_1 \int_U |f|^2 \mathrm{dVol}_g \leq \int_U |f|^2 d\mu|_U \leq k_2\int_U |f|^2 \mathrm{dVol}_g,\] which finishes the proof. ◻

In fact, the same holds for Sobolev spaces.

Remark 92. Let \((M, g, \mu)\) be a compact weighted manifold, and let \(U\) be a submanifold of \(M\) endowed with the restricted measure \(\mu\) on \(U\). Then \[H^1(U) = H^1(U, \mu|_U).\]

10 Bakry-Émery theory in domains with a convex boundary↩︎

The goal of this section is to introduce the curvature-dimension condition, formulated by Bakry and Émery [12], and to explore its connection to the Poincaré inequality in a domain with convex boundary. In particular, given a Riemannian manifold \((M, g)\) and a smooth function \(F\) on \(M\), the curvature-dimension condition relates the Hessian of \(F\), the Ricci curvature of \(M\) and the metric \(g\). Moreover, when the curvature-dimension condition holds on a domain of \(M\) with a convex boundary, we can obtain a lower bound on the first non-zero eigenvalue of the Langevin dynamics generator \[\operatorname{L} = -\mathrm{grad}_{g}\, F + \frac{1}{\beta} \Delta_g,\] which in turn yields a Poincaré inequality.

Definition 48 (Domain). Let \(M\) be a manifold, and let \(U \subset M\) be a subset. We say that \(U\) is a domain if the closure of \(U\) coincides with that of its interior.

Throughout the section we will only consider domains which are connected and have a smooth boundary.

Let us now give two definitions of convexity.

Definition 49 (Weakly geodesically convex set). Let \((M, g)\) be a Riemannian manifold. A subset \(U \subset M\) is said to be weakly geodesically convex if, for each \(p, q \in U\) there exists a distance-minimizing geodesic in \(M\) joining them and staying inside \(U\).

For the second notion of convexity, we first need to define the second fundamental form of a manifold with boundary.

Definition 50 (Second fundamental form of a manifold with boundary). Let \((M, g)\) be a Riemannian manifold with boundary, and let \(N\) be the inward pointing unit normal vector field of \(\partial M\). We define the second fundamental form of \(\partial M\) as \[\mathbb{I}(X, Y) := - \langle \nabla_X N, Y\rangle_g,\] for every \(X, Y \in \mathfrak{X}(\partial M)\).

Definition 51 (Convex boundary). Let \((M, g)\) be a Riemmanian manifold and let \(U \subset M\) be a domain. We say that \(\partial U\) is convex if \(\mathbb{I}(X,X) \geq 0\) for every \(X \in \mathfrak{X}(\partial U)\).

Although the two notions of convexity may seem quite different, it is known that if a closed domain is weakly geodesically convex, then its boundary is convex [35].

In order to derive a Poincaré inequality on a convex domain, let us consider a slightly more general setting than in the main text. This way, let \((M, g)\) be a compact Riemannian manifold. Let us define the operator \[\label{eq:generaloperatorL} \mathcal{L} := \alpha(\mathrm{grad}_{g}\, V + \Delta_g),\tag{57}\] where \(\alpha > 0\) is some constant and \(V \in C^2(M)\). In particular, choosing \(\alpha = \frac{1}{\beta}\), \(V = -\beta F\), for some smooth function \(F\) on \(M\), we recover the Langevin dynamics generator \(\operatorname{L}\).

As we mentioned in 2, \(\mathcal{L}\) is naturally defined on \(C^2(M)\), and it can be extended via its closure in \(L^2(M)\). Let us also denote this extension by \(\mathcal{L}\). To each operator \(\mathcal{L}\), we can associate a Markov triple \((M, \mu, \mathbf{\Gamma})\).

Definition 52 (Carré du champ operator). Let \((M, g)\) be a compact manifold, given some operator \(\mathcal{L}\) as in 57 , we define the carré du champ operator \(\mathbf{\Gamma}\) on \(C^2(M) \times C^2(M)\) as \[\mathbf{\Gamma}(f_1, f_2) := \frac{1}{2}(\mathcal{L}(f_1f_2) - f_1\mathcal{L}f_2 - f_2\mathcal{L}f_1),\] and denote \(\mathbf{\Gamma}(f,f)\) by \(\mathbf{\Gamma}(f)\).

Following an analogous reasoning to that followed in 8, we can obtain an explicit expression for the carré du champ operator associated with \(\mathcal{L}\).

Remark 93. The carré du champ operator associated with \(\mathcal{L}\) as defined in 57 can be written as \[\mathbf{\Gamma}(f_1, f_2) = \alpha\langle \mathrm{grad}_{g}\, f_1, \mathrm{grad}_{g}\, f_2\rangle_g,\] for any \(f_1, f_2 \in C^2(M)\), and so \[\mathbf{\Gamma}(f) = \alpha |\mathrm{grad}_{g}\, f|^2_g,\] for any \(f \in C^2(M)\).

Definition 53 (Markov triple). Let \((M, g)\) be a compact manifold, and let \(\mathcal{L}\) be defined as \[\mathcal{L} = \alpha(\mathrm{grad}_{g}\, V + \Delta_g),\] where \(\alpha > 0\) is some constant and \(V \in C^2(M)\). The Markov triple \((M, \mu, \mathbf{\Gamma})\) associated with \(\mathcal{L}\) consists of \[d\mu := e^{V} \mathrm{dVol}_g,\] which can be assumed to be a probability constant by simply shifting \(V\) by a suitable constant (\(\mathcal{L}\) is invariant under this shift), and the associated carré du champ operator given by \[\mathbf{\Gamma}(f_1, f_2) = \alpha \langle \mathrm{grad}_{g}\, f_1, \mathrm{grad}_{g}\, f_2\rangle_g,\] for any \(f_1, f_2 \in C^2(M)\).

To define the curvature-dimension condition, we first need to define the second order carré du champ operator.

Definition 54 (Second order carré du champ operator). Given a compact manifold \((M, g)\) and an operator \(\mathcal{L}\) defined as above, we define the second order carré du champ operator as \[\mathbf{\Gamma}_2(f, h) := \frac{1}{2}(\mathcal{L}\mathbf{\Gamma}(f, h) - \mathbf{\Gamma}(f,\mathcal{L}h) - \mathbf{\Gamma}(\mathcal{L}f, h)),\quad \forall f, g \in C^2(M),\] where \(\mathbf{\Gamma}\) is the carré du champ operator associated with \(\mathcal{L}\).

Definition 55 (Curvature-dimension condition). Let \((M, g)\) be a compact Riemannian manifold, let \(\mathcal{L}\) be as above and let \((M, \mu, \mathbf{\Gamma})\) be the associated Markov triple. Given some constant \(\kappa \in \mathbb{R}\), we say that \((M, \mu, \mathbf{\Gamma})\) satisfies a (tight) curvature-dimension condition \(\mathrm{CD}(\kappa)\)16 if, for all \(f \in C^2(M)\), \[\label{eq:CDOriginal} \mathbf{\Gamma}_2(f, f) \geq \kappa \mathbf{\Gamma}(f, f).\qquad{(14)}\]

While this definition of the curvature-dimension condition is given with respect to a general Markov triple, we are interested in curvature-dimension condition associated with the Markov triple induced by the Langevin diffusion generator \[\operatorname{L} = -\mathrm{grad}_{g}\, F + \frac{1}{\beta} \Delta_g,\] where \(F\) is some smooth function on \(M\) and \(\beta\) is some positive constant. To obtain such an expression, let us first present the second order carré du champ operator associated with \(\operatorname{L}\).

Lemma 25 ([18]). Let \((M, g)\) be a compact Riemannian manifold. Let \(\beta >0\), let \(F\) be some smooth function on \(M\) and consider the Langevin diffusion generator \(\operatorname{L}\). Then its associated second order carré du champ operator can be written as \[\Gamma_2(f, f) = \frac{1}{\beta^2}|\nabla^2 f|^2_g + \frac{1}{\beta^2} \mathrm{Ric}_g(\mathrm{grad}_{g}\, f, \mathrm{grad}_{g}\, f) + \frac{1}{\beta} \nabla^2F(\mathrm{grad}_{g}\, f, \mathrm{grad}_{g}\, f),\] for any \(f \in C^2(M)\).

With this lemma and the explicit expression for the carré du champ operator \(\Gamma\) associated with \(\operatorname{L}\) in mind, we can rewrite the curvature-dimension condition associated with \(\operatorname{L}\) in terms of the Ricci curvature of \(M\), the metric \(g\) and the Hessian of \(F\).

Remark 94. Let \((M, g)\) be a compact Riemannian manifold, let \(\beta > 0\), let \(F\) be some smooth function on \(M\) and let \(\operatorname{L}\) be the Langevin diffusion generator. Then the curvature-dimension condition \(\mathrm{CD}(\kappa)\) from ?? is equivalent to \[\label{eq:CDineqaux} \nabla^2 F(X,X) + \frac{1}{\beta} \mathrm{Ric}_g(X,X) \geq \kappa g(X,X),\qquad{(15)}\] for every \(X \in \mathfrak{X}(M)\).

Proof. Assume first that ?? holds. Using the explicit expression for the carré du champ operator \(\Gamma\) associated with \(\operatorname{L}\), which is given by \[\Gamma(f, h) = \frac{1}{\beta} \langle \mathrm{grad}_{g}\,f, \mathrm{grad}_{g}\,h\rangle_g,\] and using 25 we know that \[\begin{align} \Gamma_2(f, f) &= \frac{1}{\beta^2}|\nabla^2 f|^2_g + \frac{1}{\beta^2}\mathrm{Ric}_g(\mathrm{grad}_{g}\,f,\mathrm{grad}_{g}\,f) + \frac{1}{\beta}\nabla^2 F(\mathrm{grad}_{g}\,f, \mathrm{grad}_{g}\,f) \\&\geq \frac{1}{\beta^2}\mathrm{Ric}_g(\mathrm{grad}_{g}\,f,\mathrm{grad}_{g}\,f) + \frac{1}{\beta}\nabla^2 F(\mathrm{grad}_{g}\,f, \mathrm{grad}_{g}\,f) \\&\geq \frac{\kappa}{\beta}\langle\mathrm{grad}_{g}\,f, \mathrm{grad}_{g}\,f \rangle_g \\&= \kappa \Gamma(f, f). \end{align}\]

To prove the converse result, let \(p \in M\) and \(X \in T_pM\). Let \(f\) be a function such that \(\mathrm{grad}_{g}\,f(p) = X\) and \(\nabla^2 f(p) = 0\). Then \[\begin{align} \Gamma_2(f, f) &= \frac{1}{\beta^2}|\nabla^2 f|^2_g + \frac{1}{\beta^2} \mathrm{Ric}_g(\mathrm{grad}_{g}\,f, \mathrm{grad}_{g}\,f) + \frac{1}{\beta} \nabla^2F(\mathrm{grad}_{g}\,f, \mathrm{grad}_{g}\,f) \\&= \frac{1}{\beta^2} \mathrm{Ric}_g(X, X) + \frac{1}{\beta} \nabla^2F(X, X), \end{align}\] and \[\Gamma(f, f) = \frac{1}{\beta} |\mathrm{grad}_{g}\,f|^2_g = \frac{1}{\beta} |X|^2_g.\] Lastly, since \[\Gamma_2(f, f) \geq \kappa \Gamma(f, f),\] holds by assumption, we obtain that \[\frac{1}{\beta^2} \mathrm{Ric}_g(X, X) + \frac{1}{\beta} \nabla^2F(X, X) \geq \frac{\kappa}{\beta} |X|^2_g,\] finishing the proof. ◻

Let us finish the section by showing the connection between the curvature-dimension condition and the Poincaré inequality on a domain of a compact Riemannian manifold. As we mentioned before, we will be working on \(U \subset M\), a connected domain with smooth boundary \(\partial U\) of a compact Riemannian manifold \((M, g)\). We will consider the operator \[\label{eq:operatorsinalpha} \hat{\mathcal{L}} := \mathrm{grad}_{g}\,V + \Delta_g,\tag{58}\] where \(V \in C^2(M)\), and the Neumann eigenvalue problem \[\begin{cases} \hat{\mathcal{L}}f = -\lambda f,\\ Nf|_{\partial U} = 0, \end{cases}\] where \(N\) denotes the inward pointing unit normal vector field of \(\partial U\).

More specifically, we will study the minimum value \(\lambda > 0\) for which there exists some \(f \in C^2_N(U)\)—the space of \(C^2\) functions on \(U\) with Neumann boundary condition—such that \(\hat{\mathcal{L}} f = - \lambda f\). This quantity is known as the first Neumann eigenvalue of \(\hat{\mathcal{L}}\), and is denoted by \(\lambda_{1, \hat{\mathcal{L}}}(U)\).

The first Neumann eigenvalue of \(\hat{\mathcal{L}}\) corresponds to the Poincaré inequality associated with the Markov triple \((U, \mu|_U, |\mathrm{grad}_{g}\,\cdot|_g^2)\), where \(\mu|_U = \mathbb{1}_U e^{V(x)}\mathrm{dVol}_g(x)\) is assumed to be a probability measure. Before we prove this result in 95, let us present a simple characterization of \(\lambda_{1, \hat{\mathcal{L}}}\) (cf. [35]). In the following Lemma, given an operator \(\mathscr{L}\), we will use the notation \((\mathscr{L}, D(\mathscr{L}))\) to highlight that its domain is \(D(\mathscr{L})\).

Lemma 26. Let \((M, g)\) be a compact Riemannian manifold and let \(U\) be a connected bounded open domain of \(M\) with smooth boundary. Let \(\hat{\mathcal{L}}\) be defined as in 58 and let \((\hat{\mathcal{L}}, D(\mathcal{\hat{L}}))\) be the unique self-adjoint extension (cf. [35]) of \((\hat{\mathcal{L}}, C^2_N(U))\) on \(L^2(U, \mu)\). Then \(\lambda_{1, \hat{\mathcal{L}}}(U)\) is the spectral gap of \((\hat{\mathcal{L}}, D(\hat{\mathcal{L}}))\) and so \[\label{firstNeumanneig} \lambda_{1, \hat{\mathcal{L}}}(U) = \inf \{\mu|_U(|\mathrm{grad}_{g}\,f|_g^2) : f \in C^2(U),\;\left.\mu\right|_U(f) = 0,\;\left.\mu\right|_U(f^2) = 1\},\qquad{(16)}\] where \(d\mu|_U := \mathbb{1}_U e^{V(x)}\,\mathrm{dVol}_g(x)\).

Let us now show that the first Neumann eigenvalue \(\lambda_{1, \hat{\mathcal{L}}}(U)\) corresponds to the Poincaré inequality constant of \((U, \mu|_U, |\mathrm{grad}_{g}\,\cdot|_g^2)\).

Remark 95. Let \((M, g)\) be a compact Riemannian manifold and let \(U\) be a connected bounded open domain of \(M\) with smooth boundary. Let \(\hat{\mathcal{L}}\) be defined as in 58 , and consider its associated Markov triple \((U, \mu|_U, |\mathrm{grad}_{g}\,\cdot|_g^2)\), where \(\mu|_U\) is assumed to be a probability measure. Let \(\lambda_{1, \hat{\mathcal{L}}}(U)\) be the first Neumann eigenvalue of \(\hat{\mathcal{L}}\) on \(U\). Then \((U, \mu|_U, |\mathrm{grad}_{g}\,\cdot|_g^2)\) satisfies a Poincaré inequality with constant \(\lambda_{1, \hat{\mathcal{L}}}(U)\), i.e. \[\int_U f^2\, d\mu|_U - \Big(\int_U f\, d\mu|_U\Big)^2 \leq \frac{1}{\lambda_{1, \hat{\mathcal{L}}}(U)} \int_U |\mathrm{grad}_{g}\,f|_g^2 \, d\mu|_U,\] for every \(f \in C^2(U) \cap H^1(U, \mu|_U)\).

Proof. Let \(f \in C^2(U) \cap H^1(U,\mu|_U)\) and consider \[\tilde{f} := \frac{f - \left.\mu\right|_U(f)}{\left.\mu\right|_U((f - \left.\mu\right|_U(f))^2)^{\frac{1}{2}}}.\] Since \(f \in L^2(U, \mu|_U)\) and \(\mu|_U\) is a probability measure on \(U\), 24 ensures that \(f \in L^1(U, \mu|_U)\), and so \(\tilde{f}\) is well-defined and belongs to \(C^2(U) \cap H^1(U, \mu|_U)\).

It is clear that \(\left.\mu\right|_U(\tilde{f}) = 0\) and \(\left.\mu\right|_U(\tilde{f}^2) = 1\). Thus by ?? it holds that \[\lambda_{1, \hat{\mathcal{L}}}(U) \leq \left.\mu\right|_U(|\mathrm{grad}_{g}\,\tilde{f}|_g^2) = \frac{\left.\mu\right|_U(|\mathrm{grad}_{g}\,f|_g^2)}{\left.\mu\right|_U((f - \left.\mu\right|_U(f))^2)},\] i.e. \[\label{eq:preliminaryineqpoincare} \left.\mu\right|_U((f - \left.\mu\right|_U(f))^2) \leq \frac{1}{\lambda_{1, \hat{\mathcal{L}}}(U)}\left.\mu\right|_U(|\mathrm{grad}_{g}\,f|_g^2).\tag{59}\]

Moreover, since \(\mu|_U\) is a probability measure on \(U\), it holds that \[\left.\mu\right|_U((f - \left.\mu\right|_U(f))^2) = \left.\mu\right|_U(f^2) - 2\left.\mu\right|_U(f)^2 + \left.\mu\right|_U(f)^2 = \left.\mu\right|_U(f^2) - \left.\mu\right|_U(f)^2,\] and so we can rewrite 59 as \[\left.\mu\right|_U(f^2) - \left.\mu\right|_U(f)^2 \leq \frac{1}{\lambda_{1, \hat{\mathcal{L}}}(U)}\left.\mu\right|_U(|\mathrm{grad}_{g}\,f|_g^2)\] for every \(f \in C^2(U) \cap H^1(U, \mu|_U)\). ◻

The following proposition allows us to relate the curvature-dimension condition with the existence of a lower bound on the first eigenvalue of \(\hat{\mathcal{L}}\). To do this, it is necessary to ask for \(\partial U\) to be convex, as otherwise, the curvature-dimension condition does not necessarily imply a Poincaré inequality. See [74] for more details.

Proposition 96 ([35]). Let \((M, g)\) be a compact Riemannian manifold and let \(U\) be a bounded open domain of \(M\) with smooth boundary \(\partial U\). Consider the second-order elliptic differential operator \[\hat{\mathcal{L}} = \mathrm{grad}_{g}\,V + \Delta_g,\] where \(V \in C^2(M)\). Assume that there exists some constant \(\require{upgreek} \upkappa \in \mathbb{R}\) such that \[\require{upgreek} \label{modifiedCD} \mathrm{Ric}_g(X,X) - \langle \nabla_X\, \mathrm{grad}_{g}\,V, X\rangle_g \geq -\upkappa |X|^2_g,\qquad{(17)}\] for every \(X \in \mathfrak{X}(U)\). If \(\partial U\) is convex, then the first Neumann Eigenvalue of \(\hat{\mathcal{L}}\) satisfies \[\require{upgreek} \lambda_{1, \hat{\mathcal{L}}}(U) \geq \frac{8}{D^2} - \frac{1}{2} \upkappa,\] where \(D\) is the diameter of \(U\).

As we mentioned above, this result allows us to relate the curvature-dimension condition with the Poincaré inequality of the Markov triple \((U, \nu|_U, \Gamma)\) associated with the operator \[\operatorname{L} f = \langle -\mathrm{grad}_{g}\,F, \mathrm{grad}_{g}\,f\rangle_g + \frac{1}{\beta} \Delta_g\, f,\] where \(F\) is some smooth function on \(M\) and \(\beta\) is some positive constant.

Proposition 97. Let \((M, g)\) be a compact Riemannian manifold and let \(U\) be a bounded open domain with smooth convex boundary \(\partial U\). Let \(F\) be some smooth function on \(M\) and assume that there exists some constant \(\kappa\) for which \[\label{eq:CDconditionremark} \nabla^2 F(X, X) + \frac{1}{\beta} \mathrm{Ric}_g(X, X) \geq \kappa g(X, X),\qquad{(18)}\] for every \(X \in \mathfrak{X}(U)\). Then it holds that \[\int_U f^2\, d\nu|_U - \Big(\int_U f\, d\nu|_U\Big)^2 \leq \frac{2}{\kappa\beta} \int_U |\mathrm{grad}_{g}\,f|_g^2 \, d\nu|_U,\] for every \(f \in C^2(U) \cap H^1(U, \nu|_U)\).

Proof. First, let \(V\) be defined as \[\label{eq:defofV} V(x) := -\beta F(x) - \log(Z),\tag{60}\] where \(Z = \int_U e^{-\beta F(x)} dx\), and define \[\mathrm{\boldsymbol{L}} := \beta \langle - \mathrm{grad}_{g}\,F, \cdot\rangle_g + \Delta_g.\]

Let us obtain a lower bound on the first Neumann eigenvalue \(\lambda_{1, \mathrm{\boldsymbol{L}}}(U)\). In order to apply 96, let us begin by rewriting the curvature-dimension condition.

The curvature-dimension condition from ?? can be rewritten as \[\label{eq:compare1} \beta \nabla^2 F +\mathrm{Ric}_g \geq \beta \kappa g.\tag{61}\]

If we now rewrite the left-hand side of ?? for \(V(x)\) as defined in 60 , using the fact that \[\nabla^2 F(X, Y) = \langle \nabla_X \mathrm{grad}_{g}\,F, Y\rangle_g,\] we arrive at \[\require{upgreek} \label{eq:compare2} \mathrm{Ric}_g(X,X) +\beta \nabla^2 F(X,X) \geq -\upkappa |X|^2_g.\tag{62}\]

Comparing 61 62 , we conclude that 62 holds in our case for \(\require{upgreek} -\upkappa := \beta\kappa\). Thus, we can apply 96 to find that \[\lambda_{1, \mathrm{\boldsymbol{L}}}(U) \geq \frac{8}{D^2} + \frac{1}{2}\beta \kappa \geq \frac{\beta \kappa}{2},\] where \(D\) is the diameter of \(U\). And so, using 95, since \(\nu(x)|_U := \mathbb{1}_U\frac{1}{Z} e^{-\beta F}\) is a probability measure on \(U\), it holds that \[\int_U f^2\, d\nu|_U - \Big(\int_U f \, d\nu|_U \Big)^2 \leq \frac{2}{\kappa \beta} \int_U |\mathrm{grad}_{g}\,f|_g^2 \, d\nu|_U,\] for every \(f \in C^2(U) \cap H^1(U, \nu|_U)\). ◻

11 Miscellaneous auxiliary results for SDEs↩︎

In this section, we will briefly discuss three auxiliary results regarding stochastic differential equations that are used in 2.2.2. We will first provide a version of Itô’s formula for processes solving a martingale problem on a manifold. Next, we will obtain a bound on the escaping time of a generalized Cox-Ingersoll-Ross (CIR) process. Lastly, we will briefly discuss a comparison result for SDEs. We adopt the notation and definitions used in [36] and will often not restate them here; the interested reader may consult this reference for details. For a short introduction to stochastic calculus, see also [12]

11.1 Itô’s formula↩︎

We aim to prove a version of Itô’s formula, which is used to obtain the SDE of the process \(\tilde{r}_{p, v}(\tilde{X}_t)\), where \(\tilde{X}_t\) is the Langevin diffusion process generated by \(\tilde{\operatorname{L}}\) as defined in 3 , and \(\tilde{r}_{p,v}\) is defined for every \(x\) outside the cut locus of \(p\) as \[\tilde{r}_{p, v}(x)= \langle v, \log_p x\rangle,\] where \(v \in T_p M\) is fixed—cf. 10. Studying \(\tilde{r}_{p, v}(\tilde{X}_t)\) is key in the analysis of the escaping time of the process \(\tilde{X}_t\) from the saddle points of \(\tilde{F}\).

Proposition 98. Let \((M, g)\) be a compact manifold and let \(F \in C^\infty (M)\). Consider the Langevin diffusion generator \[\operatorname{L}\phi = \langle -\mathrm{grad}_{g}\,F, \mathrm{grad}_{g}\,\phi\rangle_g + \frac{1}{\beta}\Delta_g \phi,\] and let \(X_t\) be the Langevin diffusion process. Then for every \(f \in C^\infty(M)\) it holds that \[df(X_t) = \operatorname{L}f(X_t)dt + \sqrt{\frac{2}{\beta}} |\mathrm{grad}_{g}\,f(X_t)|_g dB_t,\] where \(B_t\) is a one-dimensional Brownian motion.

In order to prove 98, we will use the fact that \(X_t\) solves the martingale problem associated with \(\operatorname{L}\); namely, for every \(f \in C^\infty(M)\), there exists a local martingale \(M^f_t\) such that \[\label{eq:localmartingale} f(X_t) = f(X_0) + \int_0^t \operatorname{L}f(X_s) ds + M^f_t\tag{63}\] Thus, it suffices to identify the local martingale \(M^f_t\).

Before we give a proof for the above statement, let us introduce some auxiliary results.

Definition 56. A continuous semimartingale is a continuous process \(Y_t\) which can be (uniquely) written as \(Y_t = M_t + A_t\), where \(M_t\) is a continuous local martingale (cf. [36]) and \(A_t\) is a continuous adapted process of finite variation (cf. [36]).

Proposition 99 ([36]). A continuous real semimartingale \(Y_t = M_t + A_t\) has a finite quadratic variation and \(\langle Y, Y\rangle_t = \langle M , M\rangle_t\).

Lastly, let us identify the quadratic variation associated with the martingale \(M^f_t\), for any given function \(f \in C^\infty (M)\). To this end, we will adapt the statement and the proof of [75].

Proposition 100. Let \((M, g)\) be a compact manifold and let \(F \in C^\infty (M)\). Consider the Langevin diffusion generator \(\operatorname{L}\) and let \(X_t\) be the Langevin diffusion process. Given \(f \in C^\infty(M)\), let \(M^f_t\) be the local martingale such that 63 holds. Then \[\langle M^f, M^f\rangle_t = \frac{2}{\beta}\int_0^t |\mathrm{grad}_{g}\,f(X_s)|_g^2 ds.\]

Proof. First, it is known (cf. [36]) that the square of any continuous real semi-martingale \(Y_t\) satisfies the SDE given by \[Y^2_t = Y_0^2 + 2\int_0^t Y_s dY_s + \langle Y, Y \rangle_t.\] Substituting \(f(X_t)\) in this equation we obtain that \[\label{eq:squareIto} f^2(X_t) = f^2(X_0) + 2\int_0^t f(X_s)df(X_s) + \langle f(X), f(X)\rangle_t.\tag{64}\]

Now, the quadratic variation of \(f(X_t)\) can be rewritten in terms of the quadratic variation of \(M^f\). Indeed, since \(f(X_t)\) satisfies the SDE from 63 , we can use 99 to conclude that \[\label{eq:quadraticvariation} \langle f(X), f(X)\rangle_t = \langle M^f, M^f\rangle_t,\tag{65}\] and so, substituting 65 and the expression for \(df(X_t)\) shown in 63 into 64 , we conclude that \[\label{eq:firstidentityfsquared} f^2(X_t) = f^2(X_0) + 2\int_0^t f(X_s)dM^f_s + 2\int_0^t f(X_s) \operatorname{L}f(X_s) ds + \langle M^f, M^f\rangle_t.\tag{66}\]

On the other hand, we can apply 63 to \(f^2\) to conclude that there exists a local martingale \(M^{f^2}_t\) such that \[\label{eq:secondidentityfsquared} f^2(X_t) = f^2(X_0) + M^{f^2}_t + \int_0^t \operatorname{L}f^2(X_s) ds.\tag{67}\]

Since the decomposition of a semimartingale into a local martingale and a process with finite variation is unique, we can equate the bounded variation parts in 66 67 to conclude that \[2\int_0^t f(X_s) \operatorname{L}f(X_s) ds + \langle M^f, M^f\rangle_t = \int_0^t \operatorname{L}f^2(X_s) ds,\] which allows us to rewrite \[\begin{align} \langle M^f, M^f\rangle_t = \int_0^t \operatorname{L}f^2(X_s) ds - 2 f(X_s) \operatorname{L}f(X_s) ds &= 2\int_0^t \Gamma(f(X_s)) ds \\&= \frac{2}{\beta} \int_0^t |\mathrm{grad}_{g}\,f(X_s)|^2_g\; ds, \end{align}\] finishing the proof. ◻

Having obtained an expression for the quadratic variation of the local martingale \(M^f_t\), it is straightforward to obtain its associated SDE. To this end, we will use the following result.

Proposition 101 ([36]). If \(M_t\) is a continuous local martingale such that the measure \(d\langle M,M\rangle_t\) is almost surely equivalent17 to the Lebesgue measure, there exist a \((\mathcal{F}^M_t)\)-predictable process \(f_t\) (cf. [36]) which is strictly positive almost surely and an \((\mathcal{F}^M_t)\)-Brownian motion \(B\) such that \[d\langle M, M\rangle_t = f_t dt \quad \text{and}\quad M_t = M_0 + \int_0^t f^{1/2}_s dB_s.\]

We can now prove 98.

Proof of 98. Recall that we know that for any \(f \in C^\infty(M)\), there exists a local martingale satisfying \[f(X_t) = f(X_0) + M^f_t + \int_0^t \operatorname{L}f(X_s) ds,\] where by 100, the quadratic variation of \(M^f\) can be written as \[\langle M^f, M^f\rangle_t = \frac{2}{\beta}\int_0^t |\mathrm{grad}_{g}\,f(X_s)|_g^2 ds.\] Since \(M\) is compact, it is clear that \[d\langle M^f, M^f\rangle_t = \frac{2}{\beta} |\mathrm{grad}_{g}\,f(X_t)|_g^2 dt\] is equivalent to the Lebesgue measure \(dt\). Thus, we apply 101 to conclude that \[dM_t = \sqrt{\frac{2}{\beta}} |\mathrm{grad}_{g}\,f(X_t)|_gdB_s,\] finishing the proof. ◻

11.2 Escaping time of a generalized CIR process↩︎

Next, we will focus on the first escaping time of a process which can be seen as a slight generalization of the Cox-Ingersoll-Ross process [76]. In particular, we will study the real-valued process given by the SDE \[\label{eq:SDEmodifiedCIR} Y_t = \xi + \int_0^t (aY_s + b) ds + \int_0^t c \sigma_s \sqrt{Y_s} dB_s,\quad t\geq 0,\tag{68}\] where \(a, b, c \in \mathbb{R}\) are some positive constants, \(B_t\) is a one-dimensional Brownian motion and \(\sigma_t\) is some continuous process for which there exists some constant \(\sigma\) such that \(0 < \sigma^2_t \leq \sigma\), for all \(t \geq 0\).

We will prove the following result, which provides an explicit exponentially decaying upper bound on the probability of \(Y_t\) having a value upper bounded by some constant \(A\). The proof and some further discussion on this result can also be found in [77].

Proposition 102. Let \(Y_t\) be the solution to 68 with initial condition \(Y_0 = \xi \geq 0\). Then for any two constants \(A> 0\) and \(0 < \theta \leq \frac{2a}{c^2 \sigma}\), it holds that \[\mathbb{P}[Y_t \leq A | Y_0 = \xi] \leq e^{\theta(A - \xi)} e^{-\theta b t},\quad \forall t\geq 0.\]

Let us first ensure that there actually exists a unique solution to the SDE given in 68 .

Proposition 103. Let \((\Omega, \mathcal{F}, (\mathcal{F}_t)_{t \geq 0}, \mathbb{P})\) carrying a one-dimensional Brownian motion \((B_t)_{t \geq 0}\). Let \(\sigma_t\) be a \((\mathcal{F}_t)\)-adapted continuous process (cf. [36]) for which there exists some constant \(\sigma\) such that \(0 <\sigma^2_t \leq \sigma\) for all \(t \geq 0\), and let \(\xi\) be \(\mathcal{F}_0\)-measurable with \(\mathbb{E}[\xi^2] < \infty\). Then \[\label{eq:SDEforYt} Y_t = \xi + \int_0^t (aY_s + b) ds + \int_0^t c \sigma_s \sqrt{|Y_s|} dB_s,\quad t\geq 0,\qquad{(19)}\] where \(a, b, c > 0\) has a unique strong solution. Furthermore, whenever \(\xi \geq 0\), it holds that \(Y_t \geq 0\) for every \(t \geq 0\) and so the SDE can be written as \[Y_t = \xi + \int_0^t (aY_s + b) ds + \int_0^t c \sigma_s \sqrt{Y_s} dB_s,\quad t\geq 0.\]

Proof. As the SDE is defined in \(\mathbb{R}\), it suffices to ensure that the diffusion coefficient \[c \sigma_t \sqrt{|Y_t|}\] is Hölder and the drift coefficient \[aY_t + b\] is Lipschitz. To this end, we define \[\begin{align} F(x) &:= ax + b,\\ G(t, \omega, x) &:= c \sigma_t(\omega) \sqrt{x}. \end{align}\] Then for any \(T > 0\), \(t \in [0, T]\), \(\omega \in \Omega\) and \(x, y \geq 0\), it holds that \[\begin{align} |F(x) - F(y)| &\leq a|x-y|,\\ |G(t, \omega, x) - G(t, \omega, y)| &\leq c |\sigma_t(\omega)| |\sqrt{x} - \sqrt{y}| \leq c \sqrt{\sigma} |x - y|^\frac{1}{2}, \end{align}\] as \(x \mapsto \sqrt{x}\) is Hölder \(1/2\). Thus, we can apply [78] to ensure that the SDE has a unique strong solution.

Furthermore, note that whenever \(\xi \geq 0\), \(Y_t\) will be non-negative. Indeed, by continuity of the process \(Y_t\), before becoming negative it has to have value \(0\). In this case, whenever \(Y_s(\omega) = 0\) for some \(s > 0\) and \(\omega \in \Omega\), ?? implies that the dynamics of \(Y_t(\omega)\) at time \(s\) are deterministic and given by \(dY_t = b\,dt\). Now, since \(b > 0\) this forces \(Y_t\) to grow, bouncing it back to the positive values. ◻

Now that we have proven that the SDE from 68 has a unique solution, we can prove 102. Before we do this, let us first state two auxiliary results, namely Itô’s formula for a real-valued process, and Markov’s inequality.

Proposition 104 ([79]). If \(f \in C^{1,2}(\mathbb{R}^+ \times \mathbb{R})\) and \(Y_t\) is a standard process satisfying \[Y_t = a(t, Y_t) dt + b(t, Y_t) dB_t,\quad t \geq 0,\] then the process \(Z_t := f(t, Y_t)\) satisfies \[dZ_t = \Big(\frac{\partial f}{\partial t}(t, Y_t) + a(t, Y_t) \frac{\partial f}{\partial x}(t, Y_t) + \frac{1}{2} b(t, Y_t)^2\frac{\partial^2 f}{\partial x^2}(t, Y_t)\Big) dt + b(t, Y_t) \frac{\partial f}{\partial x}(t, Y_t) dB_t.\]

Proposition 105 (Markov’s inequality). Let \(X\) be a non-negative real variable. Then for every positive constant \(\lambda > 0\), it holds that \[\mathbb{P}[X \geq \lambda] \leq \frac{1}{\lambda}\mathbb{E}[X].\]

We can now provide a proof for 102.

Proof of 102. First of all, recall from 103 that the unique solution of 68 is such that \(Y_t \geq 0\) for every \(t \geq 0\) whenever \(Y_0 \geq 0\). Now, let \(0 < \theta \leq \frac{2a}{c^2 \sigma}\) be fixed, and let us define \[M_t := f(t, Y_t) = \exp(\theta b t - \theta Y_t),\] where \(f(t, x) := \exp(\theta b t - \theta x)\). Thus, we can apply Itô’s formula from 104, using the identities \[\begin{align} \frac{\partial f}{\partial t}(t, Y_t) &= \theta b M_t,\\ \frac{\partial f}{\partial x}(t, Y_t) &= -\theta M_t,\\ \frac{\partial^2 f}{\partial x^2}(t, Y_t) &= \theta^2 M_t, \end{align}\] to conclude that \(M_t\) satisfies the SDE \[\begin{align} dM_t &= M_t\Big( \big(\theta b -\theta (a Y_t + b) + \frac{1}{2} c^2 \sigma_t^2 Y_t \theta^2\big)dt - c\sqrt{Y_t} \sigma_t dB_t\Big)\\ &= M_t\Big( \big(-\theta a + \frac{1}{2} c^2 \sigma_t^2 \theta^2\big)Y_t dt - c\sqrt{Y_t} \sigma_t dB_t\Big). \end{align}\]

Note that, as \(Y_t \geq 0\) and \(\sigma_t^2 \leq \sigma\) for every \(t \geq 0\), it holds that \[\label{eq:negativedrift} \big(-\theta a + \frac{1}{2} c^2 \sigma_t^2 \theta^2\big)Y_t \leq \big(-\theta a + \frac{1}{2} c^2 \sigma \theta^2\big)Y_t \leq 0,\quad \forall t\geq 0,\tag{69}\] since by definition, \(\theta\) is such that \[0< \theta \leq \frac{2a}{c^2 \sigma}.\]

69 implies that \(M_t\) is a supermartingale: \[\begin{align} \mathbb{E}[M_t] &= \mathbb{E}[M_0] + \mathbb{E}\Big[\int_0^t M_s \big(-\theta a + \frac{1}{2} c^2 \sigma_s^2 \theta^2\big)Y_s ds\Big] - \mathbb{E}\Big[\int_0^t M_s c\sqrt{Y_s} \sigma_s dB_s\Big]\\ &= \mathbb{E}[M_0] + \mathbb{E}\Big[\int_0^t M_s \big(-\theta a + \frac{1}{2} c^2 \sigma_s^2 \theta^2\big)Y_s ds\Big]\\ &\leq \mathbb{E}[M_0], \end{align}\] where the second identity follows from the fact that \[\int_0^t M_s c\sqrt{Y_s} \sigma_s dB_s\] is a martingale by [80]—and so its expected value vanishes—and the inequality follows from 69 and the fact that \(M_t\) is non-negative by definition. Therefore, for every \(t \geq 0\), \[\mathbb{E}[M_t | Y_0 = \xi] \leq M_0 = e^{-\theta \xi},\] which can be rewritten as \[\mathbb{E}[\exp(\theta b t - \theta Y_t) | Y_0 = \xi] = e^{\theta b t} \mathbb{E}[\exp(- \theta Y_t) | Y_0 = \xi] \leq e^{-\theta \xi},\] and so \[\mathbb{E}[\exp(- \theta Y_t) | Y_0 = \xi] \leq e^{-\theta \xi}e^{-\theta b t},\quad \forall t\geq 0.\]

Lastly, we can use Markov’s inequality from 105 to conclude that for every \(t \geq 0\), \[\mathbb{P}[Y_t \leq A | Y_0 = \xi] = \mathbb{P}[e^{-\theta Y_t} \geq e^{-\theta A} | Y_0 = \xi ] \leq e^{\theta A} \mathbb{E}[e^{-\theta Y_t} | Y_0 = \xi] \leq e^{\theta(A - \xi)} e^{-\theta b t},\] finishing the proof. ◻

11.3 Comparison result for SDEs↩︎

Lastly, let us provide a comparison result for the processes \(\tilde{r}(\tilde{X}_t)^2\) and \(z_t^2\) as defined in 70 71 . For more general comparison results, and a more detailed discussion, see for example [81].

Recall that in 2.2.2 we obtained two processes in \(\mathbb{R}\), denoted by \(\tilde{r}^2(\tilde{X}_t)\) and \(z_t\). In particular, given a Riemannian manifold \((B, h)\) and some fixed point \(y \in B\), we showed that \(\frac{1}{2}\tilde{r}^2(\tilde{X}_t)\) satisfies \[\begin{align} \label{sde1} \begin{aligned} d\left[\frac{1}{2}\tilde{r}(\tilde{X}_t)^2\right] &= \Big[\langle -\mathrm{grad}_{h}\,\tilde{F}(\tilde{X}_t), \tilde{r}(\tilde{X}_t)\mathrm{grad}_{h}\,\tilde{r}(\tilde{X}_t) \rangle_h \\&\quad\quad + \frac{1}{\beta}\left(\tilde{r}(\tilde{X}_t)\Delta_h\tilde{r}(\tilde{X}_t) + |\mathrm{grad}_{h}\,\tilde{r}(\tilde{X}_t)|_h^2\right)\Big]dt \\&\quad + \sqrt{\frac{2}{\beta}}|\tilde{r}(\tilde{X}_t)||\mathrm{grad}_{h}\,\tilde{r}(\tilde{X}_t)|_h\, dB_t, \end{aligned} \end{align}\tag{70}\] before \(\tilde{X}_t\) escapes from the ball \(\mathcal{B}(i(B), y)\), and \(\frac{1}{2}z^2_t\) satisfies \[\label{sde2} d\left[\frac{1}{2}(z_t)^2\right] = \left[2 \lambda_* \frac{1}{2} (z_t)^2 + \frac{1}{4\beta} \right]dt + \sqrt{\frac{2}{\beta}} |z_t| |\mathrm{grad}_{h}\,\tilde{r}(\tilde{X}_t)|_h\,dB_t.\tag{71}\]

Proposition 106. The processes \(\tilde{r}(\tilde{X}_t)\) and \(z_t\) are such that, if \(\tilde{r}(X_0) = z_0\). Then \[\mathbb{P}[\tilde{r}^2(\tilde{X}_t) \geq z_t^2 | t < \tau^c_{\mathcal{B}(i(B), p)}] = 1,\quad \forall t \geq 0,\] where \(\tau^c_{\mathcal{B}(i(B), y)}\) is the first escaping time of \(\tilde{X}_t\) from \(\mathcal{B}(i(B), y)\).

Proof. In order to see this, let us first define two auxiliary functions, \[\begin{align} \Phi(t, x) &:= \langle -\mathrm{grad}_{h}\,\tilde{F}(x), \tilde{r}(x)\mathrm{grad}_{h}\,\tilde{r}(x) \rangle_h + \frac{1}{\beta}\left(\tilde{r}(x)\Delta_h\tilde{r}(x) + |\mathrm{grad}_{h}\,\tilde{r}(x)|_h^2\right),\\ \tilde{\Phi}(t, z) &:= 2 \lambda_* \frac{1}{2} z^2 + \frac{1}{4\beta}, \end{align}\] where \(x \in B\) and \(z \in \mathbb{R}\). In 14 we showed that \[\label{eq:ineqPhis} \Phi(t, \tilde{X}_t) \geq \tilde{\Phi}(t, \tilde{r}(\tilde{X}_t)),\tag{72}\] for every \(t \geq 0\) before \(\tilde{X}_t\) escapes from \(\mathcal{B}(i(B), y)\).

Let us now define the following auxiliary processes; \[\begin{align} U_t &:= \frac{1}{2}\left(\tilde{r}(\tilde{X}_t)^2 - (z_t)^2\right),\\ a_t &:= \begin{cases} \frac{1}{U_t}(\tilde{\Phi}(t, \tilde{r}(\tilde{X}_t)) - \tilde{\Phi}(t, z_t)), \quad &\text{if } U_t \neq 0,\\ 0,\quad &\text{if } U_t = 0, \end{cases}\\ b_t &:= \Phi(t, \tilde{X}_t) - \tilde{\Phi}(t, \tilde{r}(\tilde{X}_t)),\\ c_t &:= \begin{cases} \frac{1}{U_t} \sqrt{\frac{2}{\beta}}(|\tilde{r}(\tilde{X}_t)| - |z_t|)|\mathrm{grad}_{h}\,\tilde{r}(\tilde{X}_t)|_h, \quad &\text{if } U_t \neq 0,\\ 0,\quad &\text{if } U_t = 0. \end{cases} \end{align}\] These processes satisfy \[\begin{align} b_t + a_t U_t &= \Phi(t, \tilde{X}_t) - \tilde{\Phi}(t, z_t),\\ U_t c_t &= \sqrt{\frac{2}{\beta}}(|\tilde{r}(\tilde{X}_t)| - |z_t|)|\mathrm{grad}_{h}\,\tilde{r}(\tilde{X}_t)|_h. \end{align}\] Then, subtracting 71 from 70 we obtain that \(U_t\) satisfies \[U_t = U_0 + \int_0^t (a_s U_s + b_s)ds + \int_0^t U_s c_s dB_s.\]

The process \(U_t\) is continuous—as \(\tilde{r}(\tilde{X}_t)\) and \(z_t\) are continuous. Furthermore, \(U_t\) is non-negative. This follows from the same reasoning used to prove that the solution of the modified CIR process from 68 was non-negative. Indeed, \(b_t \geq 0\) by 72 and so whenever \(U_t\) is zero, its associated SDE at that time is of the form \[dU_t = b_t dt,\] which does not allow for \(U_t\) to attain negative values. ◻

12 PDEs on manifolds↩︎

The goal of this appendix is to guarantee that there exists a solution to the PDE that appears in the construction of the second Lyapunov function (cf. 1). Indeed, in that construction we need to ensure that there exists some function \(f: \overline{U} \rightarrow \mathbb{R}\) verifying \[\label{EDPoriginal} \begin{cases} \tilde{\operatorname{L}}f = -\theta f&\quad x \in U,\\ f = 1&\quad x \in \partial U, \end{cases}\tag{73}\] where \(U\) is some open bounded subset of a compact Riemannian manifold \((B, h)\), \(\theta > 0\) and \(\tilde{\operatorname{L}}\) is defined as \[\tilde{\operatorname{L}} \phi = \langle -\mathrm{grad}_{h}\, \tilde{F}, \mathrm{grad}_{h}\, \phi \rangle + \frac{1}{\beta} \Delta_h \phi,\] for some smooth function \(\tilde{F}\) on \(B\) (cf. 3 ).

One could be tempted to directly apply classic results of existence of weak solutions to the above equation, such as those presented in Evans’ work [82]. Nevertheless, to the best of our knowledge, most of the existing literature on this topic focuses on PDEs defined on (subsets of) \(\mathbb{R}^n\), which clearly is not case for 73 .

Although the literature does not usually focus on PDEs on manifolds, some of the existence results can be adapted in a natural way. This will be done in the following sections. We will adapt [82], where the author proves several existence results for PDEs defined in terms of a general elliptic operator \(\mathfrak{L}\)—a second order differential operator for which the matrix of its second order coefficients is negative-definite.

The reason why such existence results can be adapted is that the existence of weak solutions is guaranteed using results from functional analysis, such as the Lax-Milgram theorem (cf. 110), the Rellich-Kondrachov theorem (cf. 107) and the Fredholm alternative theorem (cf. 108), which apply to operators defined on Hilbert spaces, and in particular function spaces. For this reason, since we can construct analogous Hilbert spaces on Riemannian manifolds (cf. 9), which retain the key properties of their Euclidean counterparts, such as completeness and compact embeddings, the results can be applied without essential modification.

With this in mind, in 12.1 we will introduce the main definitions and tools which will allow us to adapt the classic theory to smooth manifolds. In 12.2 we will include the adapted versions of the theorems from [82]. Finally, in 12.3, we will discuss the regularity of the weak solution of 73 .

12.1 Tools from functional analysis↩︎

Definition 57. Let \(M\) be a manifold, we say a subset \(U\) of \(M\) is precompact if its closure \(\overline{U}\) is compact in \(M\).

Definition 58. Let \(M\) be a manifold and let \(U\) be an open subset of \(M\). Define \(H^1_0(U)\) as the closure of \(C^\infty_c(M)\), the space of smooth functions with compact support, in \(H^1(U)\)—cf. 47.

One nice property of \(H^1_0(U)\), is that it can be compactly embedded into \(L^2(U)\).

Theorem 107 (Compact embedding theorem - Rellich–Kondrachov - [73]). Let \(M\) be a manifold and let \(U\) be a precompact open subset of \(M\). Then the identical embedding \[H^1_0(U) \hookrightarrow L^2(U)\] is a compact operator.

Theorem 108 (Fredholm alternative, [82]). Let \(H\) be a Hilbert space and \(\mathcal{K}: H \rightarrow H\) be a compact linear operator. Then either for each \(f \in H\), the equation \[u - \mathcal{K}u = f\] has a unique solution or the homogeneous equation \[u - \mathcal{K}u = 0\] has solutions \(u \neq 0\).

In addition, should the homogeneous equation have non-zero solutions, the space of solutions of the homogeneous problem is finite dimensional. Furthermore, the non-homogeneous equation \[u - \mathcal{K}u = f\] has a solution if and only if \(f \in \ker(I-\mathcal{K}^*)^\bot\).

We now continue with a result regarding the spectrum of a bounded, compact linear operator on a Banach space. In order to state the result, we need to introduce some definitions and notation.

Definition 59 (Spectrum of an operator). Given a bounded linear operator \(\mathcal{K} : H \to H\) on a real Banach space \(H\), the resolvent of \(\mathcal{K}\) is \[\rho(\mathcal{K}) := \{x \in \mathbb{R} : (\mathcal{K} - xI) \text{ is bijective}\}.\] The spectrum of \(\mathcal{K}\) is then defined as \[\sigma(\mathcal{K}) = \mathbb{R} - \rho(\mathcal{K}).\] Finally, the point spectrum of \(\mathcal{K}\), denoted by \(\sigma_p(\mathcal{K})\) is defined as \[\sigma_p(\mathcal{K}) := \{x \in \sigma(\mathcal{K}) : \ker(\mathcal{K} - xI) \neq \{0\}\}.\]

Theorem 109 ([82]). Let \(H\) be a Hilbert space, assume \(\dim(H) = \infty\), and let \(\mathcal{K}: H \rightarrow H\) be a compact linear operator. Then the following hold:

  1. \(0 \in \sigma(\mathcal{K})\),

  2. \(\sigma(\mathcal{K}) - \{0\} = \sigma_p(\mathcal{K}) - \{0\}\),

  3. Either \(\sigma(\mathcal{K}) - \{0\}\) is finite, or else \(\sigma(\mathcal{K}) - \{0\}\) is a sequence tending to 0.

It only remains to state a well-known theorem known as Lax-Milgram theorem, which is a key ingredient when ensuring the existence of weak solutions to the PDEs that we will study in the next section. Before we state the theorem, let us define two properties of a bilinear form \(\mathfrak{B}\) on a Hilbert space.

Definition 60 (Bounded operator). A bilinear form \(\mathfrak{B}\) on a Hilbert space \(H\) is said to be bounded if there exists a constant \(\kappa>0\) such that \[|\mathfrak{B}(x,y)|\leq \kappa\norm{x}\norm{y}\quad \text{for all } x,y \in H.\]

Definition 61 (Coercive). A bilinear form \(\mathfrak{B}\) on a Hilbert space \(H\) is said to be coercive if there exists a number \(k > 0\) such that \[\mathfrak{B}(x,x) \geq k \norm{x}^2\quad \text{for all }x \in H.\]

Theorem 110 (Lax-Milgram, [82]). Let \(\mathfrak{B}\) be a bounded, coercive and bilinear form on a Hilbert space \(H\). Then for every bounded linear functional \(\psi \in H^*\), there exists a unique element \(f \in H\) such that \[\mathfrak{B}(x,f) = \psi(x)\quad \text{for all }x \in H.\]

12.2 Existence of weak solutions to boundary problems↩︎

In this section, we will prove an existence result for weak solutions of the boundary problem presented in 73 . Let us start with the definition of a weak solution of a PDE and the bilinear form associated with an operator.

Definition 62 (Weak solution and bilinear form associated with \(\mathfrak{L}\)). Let \(M\) be a compact manifold and let \(U \subset M\) be open. Given some second-order differential operator \(\mathfrak{L}\) on \(M\), some constant \(\lambda \in \mathbb{R}\) and some \(f \in L^2(U)\) we say that \(u\) is a weak solution of the PDE \[\mathfrak{L}u = \lambda u + f,\quad x \in U,\] if, for every \(\phi \in C^\infty_0(U)\), it holds that \[\int_U (\mathfrak{L}u)\phi = \lambda(u, \phi)_{{L^2(U)}} + (f, \phi)_{L^2(U)}.\]

In this setting, \[\mathfrak{B}[u, \phi] := \lambda(u, \phi)_{{L^2(U)}} + (f, \phi)_{L^2(U)}\] is known as the bilinear form associated with the operator \(\mathfrak{L}\).

Theorem 111. Let \((M, g)\) be a compact Riemannian manifold, let \(U \subset M\) be open, and let \(F\) be some smooth function on \(M\). Let \(\operatorname{L}\) be the Langevin diffusion operator given by \[\operatorname{L}f = \langle -\mathrm{grad}_{g}\, F, \mathrm{grad}_{g}\, f\rangle_g + \frac{1}{\beta}\Delta_g f .\] Let \(\theta > 0\) be such that \(-\theta \notin \sigma(\operatorname{L})\). Then there exists a unique weak solution of the problem \[\begin{cases} \operatorname{L}f = -\theta f&\quad x \in U,\\ f = 1&\quad x \in \partial U. \end{cases}\]

Before showing the proof of this result, we will state some intermediate theorems needed to ensure the existence of a weak solution, following the same structure as in Evans’ book [82].

We will not incorporate the proofs of the intermediate theorems from Evans. Instead, we will sketch them, highlighting the key steps. We believe that doing so allows the reader understand how the proofs can be adapted without actually rewriting anything, simply by recalling that \(H^1_0(U)\) is a Hilbert space for any open subset of a compact manifold \(M\) and using the results of 12.1. To see the original proofs in detail, refer to [82].

Theorem 112 (First existence theorem for weak solutions, [82]). Let \(\mathfrak{L}\) be an elliptic operator. Let \(M\) be a compact manifold and let \(U\) be some open subset of \(M\). There exists a number \(\gamma \geq 0\) such that, for all \(\mu \geq \gamma\), and each function \(f \in L^2(U)\) there exists a unique weak solution \(u \in H^1_0(U)\) of the boundary-value problem \[\begin{cases} \mathfrak{L}u + \mu u= + f&\quad x \in U,\\ f = 0&\quad x \in \partial U. \end{cases}\]

Sketch of the proof. The proof relies on 88 and 89 to define a bilinear form \(\mathfrak{B}_\mu\) associated with the operator \(\mathfrak{L} + \mu\). The operator \(\mathfrak{B}_\mu\) is defined on \(H^1_0(U)\) and is proven to be bounded and coercive. Then using 110 the result follows. ◻

Definition 63 (Weak solution of adjoint problem). Let \(M\) be a compact manifold and let \(U\) be some open subset of \(M\). Given some \(f \in L^2(U)\), we say that \(v \in H^1_0(U)\) is a weak solution of the adjoint problem \[\begin{cases} \mathfrak{L}^* v = f&\quad \text{in } U,\\ v = 0&\quad \text{on }\partial U, \end{cases}\] provided \[\mathfrak{B}^*[v,u] = (f, u)_{L^2(U)}\] for all \(u \in H^1_0(U)\), where \(\mathfrak{B}^*\) is the adjoint bilinear form of \(\mathfrak{B}\)—the form associated with the operator \(\mathfrak{L}\)—i.e. \[\mathfrak{B}^*[v,u] := \mathfrak{B}[u, v],\] for all \(u, v \in H^1_0(U)\).

Theorem 113 (Second existence theorem for weak solutions, [82]). Let \(\mathfrak{L}\) be an elliptic operator. Let \(M\) be a compact manifold and let \(U\) be some open subset of \(M\). Then one and only one of the following statements hold:

  1. For each \(f \in L^2(U)\) there exists a unique weak solution \(u\) of the boundary-value problem \[\label{alpha2} \begin{cases} \mathfrak{L}u = f&\quad \text{in } U,\\ u = 0&\quad \text{on } \partial U.\\ \end{cases}\qquad{(20)}\]

  2. There exists a weak solution \(u \not\equiv 0\) of the homogeneous problem \[\label{beta2} \begin{cases} \mathfrak{L}u = 0&\quad \text{in } U,\\ u = 0&\quad \text{on } \partial U.\\ \end{cases}\qquad{(21)}\]

Furthermore, should the second assertion hold, the dimension of the subspace \(N \subset H^1_0(U)\) of weak solutions of the problem in ?? is finite and equals the dimension of the subspace \(N^* \subset H^1_0(U)\) of weak solutions of \[\begin{cases} \mathfrak{L}^* v = 0&\quad \text{in } U,\\ v = 0&\quad \text{on }\partial U. \end{cases}\]

Finally, the boundary-value problem in ?? has a weak solution if and only if \[(f,v)_{L^2(U)} = 0\text{ for all } v \in N^*.\]

Sketch of the proof. The proof relies on 112. A bounded and linear operator on \(L^2(U)\) is defined, which is proven to be compact by the Rellich-Kondrachov theorem (107). Finally, the Fredholm alternative (108) is applied to the operator, and its spectrum is studied using 109. ◻

Theorem 114 (Third Existence Theorem for weak solutions, [82]). Let \(\mathfrak{L}\) be an elliptic operator. Let \(M\) be a compact manifold and let \(U\) be some open subset of \(M\).

There exists an at most countable set \(\Sigma \subset \mathbb{R}\)—the spectrum of \(\mathfrak{L}\)—such that the boundary-value problem \[\begin{cases} \mathfrak{L}u = \lambda u + f\quad &\text{in } U,\\ u = 0\quad &\text{on } \partial U, \end{cases}\] has a unique weak solution for each \(f \in L^2(U)\) if and only if \(\lambda \notin \Sigma\).

If \(\Sigma\) is infinite, then the sequence \(\Sigma = \{\lambda_k\}_{k = 1}^\infty\) is non-decreasing and \[\lambda_k \rightarrow +\infty.\]

Sketch of the proof. The proof relies on 112 113, and uses the same linear, bounded and compact operator from the proof of 113. To conclude, its eigenvalues are studied using 109. ◻

Finally, we can present the proof of 111, which regards the existence of weak solutions to the PDE shown in 73 .

Proof of 111. Note that the operator considered in the statement of the theorem is not elliptic but rather \(-\operatorname{L}\) is. Therefore, instead of considering the problem from 73 , we can consider \(\tilde{f} = f - 1\) and the associated PDE \[\begin{cases} -\operatorname{L}\tilde{f} = -\theta \tilde{f} - \theta&\quad x \in U,\\ \tilde{f} = 0&\quad x \in \partial U. \end{cases}\]

Now, we know that this problem has a solution for almost every \(\theta \in \mathbb{R}\)—except for at most a countable subset of \(\mathbb{R}\)—by 114. Thus, the problem from the statement has a unique solution for almost every \(\theta \in \mathbb{R}\). ◻

12.3 Regularity↩︎

In this subsection we will briefly discuss the regularity of the weak solution of the PDE from 73 up to the boundary of \(U\).

Note that, unlike in the previous section, there is no need to adapt the existing definitions and results from [82], as the regularity of a function is a local notion. See [83] for more details.

In order to ensure that the weak solution of a PDE is regular up to its boundary, we need to assume some regularity of the boundary of the subset \(U\). Thus, let us define what it means for a subset of a manifold \(M\) to have a smooth boundary. Again, since the definition is local, we may consider \(M\) to be \(\mathbb{R}^n\) for simplicity.

Definition 64 (Boundary regularity). Let \(U\) be a subset of \(\mathbb{R}^n\). We say that its boundary \(\partial U\) is \(C^k\) if, for every \(x \in \partial U\), there exists some radius \(r > 0\) and a \(C^k\) function \(\gamma: \mathbb{R}^{n-1} \rightarrow \mathbb{R}\) such that—possibly after relabelling and reorienting the coordinate axes if necessary—it holds that \[U \cap \mathcal{B}(r, x) = \{x \in \mathcal{B}(r, x) : x_n > \gamma(x_1,\dotsc, x_{n-1}\},\] where \(\mathcal{B}(r, x)\) denotes the Euclidean ball centered in \(x\) with radius \(r\).

We say \(\partial U\) is \(C^\infty\) if \(\partial U\) is \(C^k\) for all \(k \in \mathbb{N}\).

With the above definition in mind, we now present the result concerning the regularity of the weak solution obtained in 111.

Theorem 115 ([82]). Let \(u\) be the weak solution of the boundary-value problem from 111. Assume that the coefficients of \(\operatorname{L}\) are of class \(C^\infty\) in \(\overline{U}\). If \(\partial U\) is \(C^\infty\), then \[u \in C^\infty(\overline{U}).\]

Let us conclude with an auxiliary result that is key when studying the smoothness of the boundary of certain subsets (cf. 14). Before stating the result, let us define regular points and regular values.

Definition 65 (Regular value). Let \(M\) be an \(n\)-dimensional manifold, and let \(N\) be a \(k\)-dimensional manifold. Let \(F: X \to Y\) be a smooth function. We say \(x \in M\) is a regular point if \(dF\) is surjective at \(x\). A point \(y \in N\) is said to be a regular value if every \(x \in F^{-1}(y)\) is a regular point.

As we mentioned earlier, if the boundary of a set can be described—at least locally—as \(f^{-1}(y)\) for some smooth function \(f\), where \(y\) is a regular value of \(f\), we can apply the implicit function theorem to conclude that the boundary is smooth. The following result guarantees that, given a smooth function \(f\), almost every point is a regular value of \(f\) (cf. [84]).

Theorem 116. Let \(M\) and \(N\) be two smooth manifolds, and let \(f: M \to N\) be a smooth function. The set of regular values of \(f\) is everywhere dense in \(N\).

13 Studying the pseudo-radial process↩︎

The goal of this section is to analyze in detail the function \(\tilde{r}_{p, v}\) defined in 2.2 and used in 15 to study the escaping time of the Langevin dynamics from the saddle points. Given a manifold \((M, g)\), some fixed point \(p \in M\) and a fixed tangent vector \(v \in T_p M\) with norm one, \(\tilde{r}_{p, v}\) is defined outside the cut locus of \(p\) as \[\label{eq:deftilder} \tilde{r}_{p, v}(x) = \langle v, \log_p x\rangle,\tag{74}\] where \(\log_p x\) denotes the inverse of the exponential mapping, and \(\langle \cdot , \cdot \rangle\) is the usual Euclidean inner product.

In the proof of 15 both the norm of the gradient of \(\tilde{r}_{p, v}\) and its Laplacian need to be bounded. A natural strategy would be to try to show that its Laplacian is non-negative and its gradient has constant norm one. However, in this section we will first show that the Laplacian of \(\tilde{r}_{p, v}\) cannot be assumed to be non-negative, nor the norm of its gradient can be assumed to be constant and equal to one. Then we will show how, under the assumption of \((M, g)\) being a symmetric space (cf. 17), the absolute value of its Laplacian and the norm of its gradient stay close to zero and one, respectively, when restricting \(\tilde{r}_{p, v}\) to a geodesic ball centered at \(p\) of sufficiently small radius. Lastly, we will state and prove the auxiliary lemmas 33 and 32 used in 2.2.2.

13.1 No-go results for the Laplacian and the gradient of \(\tilde{r}_{p,v}\)↩︎

Consider a very simple Riemannian manifold; the two-dimensional round sphere.

Remark 117. Let \((\mathbb{S}^2, g_{\mathit{round}})\) be the two-sphere endowed with the usual round metric, let \(p \in \mathbb{S}^2\) be a fixed point and let \(v \in T_p \mathbb{S}^2\) with \(|v|_{g_{\mathit{round}}} = 1\). There is no radius \(R\) such that \(\Delta_{g_{\mathit{round}}} \tilde{r}_{p, v}\) is non-negative when restricted to \(\mathcal{B}(R, p)\).

Proof. Without loss of generality, we can assume that \(p \in \mathbb{S}^2\) has coordinates \((0, 0, 1)\), and therefore \(v = (v_1, v_2, 0)\) with \(v_1^2 + v_2^2 = 1\). In these coordinates, \[\log_p(x) = (\theta \cos\varphi, \theta\sin\varphi, 0)\] where \(\theta \in (0, \pi)\) and \(\varphi \in [0, 2\pi)\). Therefore, \[\tilde{r}_{p, v}(x) = \theta(v_1\cos\varphi + v_2 \sin\varphi).\]

It is a standard fact that the round metric of \(\mathbb{S}^2\) in normal coordinates is of the form \[g_{\mathit{round}}(\theta, \varphi) = d\theta^2 + \sin^2 \theta d\varphi^2,\] which implies that \[\Gamma^\theta_{\varphi \varphi} = - \sin\theta\cos\theta,\quad \Gamma^\varphi_{\varphi \theta} = \Gamma^\varphi_{\theta\varphi} = \cot\theta,\] while the rest of the Christoffel symbols vanish identically.

Since the Laplacian can be written in local coordinates as \[\label{eq:Laplacianincoords} \Delta_g = g^{ij}(x) \frac{\partial^2}{\partial x^i \partial x^j} - g^{kl}(x) \Gamma^n_{kl}(x) \frac{\partial}{\partial x^n},\tag{75}\] we can use the expressions for the Christoffel symbols, the metric and its inverse to conclude that \[\begin{align} \Delta_{g_\mathit{round}} &= g^{\theta \theta}_{\mathit{round}}(x) \frac{\partial^2}{\partial \theta^2} + g^{\varphi \varphi}_{\mathit{round}}(x) \frac{\partial^2}{\partial \varphi^2} - g^{\varphi\varphi}_{\mathit{round}}(x) \Gamma^\theta_{\varphi\varphi}(x) \frac{\partial}{\partial \theta} \\&= \frac{\partial^2}{\partial \theta^2} + \frac{1}{\sin^2 \theta} \frac{\partial^2}{\partial \varphi^2} + \cot \theta \frac{\partial}{\partial \theta}, \end{align}\] and so \[\begin{align} \Delta_{g_\mathit{round}} \tilde{r}_{p, v}(x) &= \Big(\frac{\partial^2}{\partial \theta^2} + \frac{1}{\sin^2 \theta} \frac{\partial^2}{\partial \varphi^2} + \cot \theta \frac{\partial}{\partial \theta}\Big)\theta(v_1\cos\varphi + v_2 \sin\varphi) \\&= -\frac{\theta}{\sin^2 \theta}(v_1\cos\varphi+v_2\sin\varphi) + \cot\theta (v_1\cos\varphi + v_2 \sin\varphi) \\&= (v_1\cos\varphi + v_2 \sin\varphi)\Big(\cot\theta -\frac{\theta}{\sin^2 \theta}\Big) . \end{align}\] Observing this identity we can conclude that, even if the Laplacian of \(\tilde{r}_{p, v}\) tends to zero as \(\theta\) becomes small, this does not imply that \(\Delta_{g_\mathit{round}}\, \tilde{r}_{p, v}(x) \geq 0\) near \(p\), as the expression on the right-hand side has opposite signs for \(\varphi = 0\) and \(\varphi = \pi\) when \(\theta > 0\). ◻

Now, working again with the two-sphere, let us show that the norm of the gradient of \(\tilde{r}_{p, v}\) cannot be assumed to be constant.

Remark 118. There is no radius \(R\) such that \(|\mathrm{grad}_{g_{\mathit{round}}}\, \tilde{r}_{p, v}|_{g_{\mathit{round}}} \equiv 1\) then restricted to the ball \(\mathcal{B}(R, p)\).

Proof. Using the expressions for \(\tilde{r}_{p, v}\) and the round metric in normal coordinates obtained in the previous remark, along with the expression of \(\mathrm{grad}_{g_{\mathit{round}}}\, \tilde{r}_{p, v}\) in local coordinates, namely \[\mathrm{grad}_{g_{\mathit{round}}}\, \tilde{r}_{p, v}(x) = g_{\mathit{round}}^{ij}(x) \frac{\partial \tilde{r}_{p, v}}{\partial x^i} \partial_j,\] we conclude that \[\mathrm{grad}_{g_{\mathit{round}}}\, \tilde{r}_{p, v}(x) = (v_1 \cos \varphi + v_2 \sin \varphi) \partial_{\theta} + \frac{1}{\sin^2\theta} \theta (v_2 \cos\varphi - v_1 \sin \varphi)\partial_{\varphi}.\] Therefore \[|\mathrm{grad}_{g_{\mathit{round}}}\, \tilde{r}_{p, v}(x)|^2_{g_{\mathit{round}}} = (v_1 \cos \varphi + v_2 \sin \varphi)^2 + \frac{1}{\sin^2\theta} \theta^2 (v_2 \cos\varphi - v_1 \sin \varphi)^2.\] Although it is easy to see that the norm of the gradient tends to \(1\) as \(\theta\) approaches zero, it is clear that the expression is not constant in \(\mathcal{B}(R, p)\) for any value of \(R\). ◻

13.2 Bounding the Laplacian in symmetric spaces↩︎

Although \(\Delta \tilde{r}_{p, v}\) cannot be assumed to be non-negative, here we will prove that if \(\tilde{r}_{p, v}\) is defined on a symmetric space, its absolute value can be bounded when restricting \(\tilde{r}_{p, v}\) to a sufficiently small geodesic ball. Throughout this section, \((M, g)\) will be assumed to be a complete and connected symmetric space of dimension \(d\).

Let \(p \in M\) be fixed. Using normal coordinates centered at \(p\) we can rewrite 74 as \[\tilde{r}_{p, v}(x) = \delta_{ij} v^i x^i.\] Therefore, using the expression of the Laplacian in local coordinates (cf. 75 ), and assuming without loss of generality that \(v = \frac{\partial }{\partial x^1}\), we can rewrite the Laplacian of \(\tilde{r}_{p, v}\) as \[\label{eq:exprlaplacianr} \Delta_g \tilde{r}_{p, v}(x) = -g^{kl}(x) \Gamma^{1}_{kl}(x).\tag{76}\] Thus, if we are able to bound the coefficients of the inverse metric and the Christoffel symbols near \(p\), we will be able to bound \(\Delta_g \tilde{r}_{p, v}(x)\), as desired. With this end, we will first obtain an explicit expression for the Taylor series of the coefficients of the metric \(g\) in normal coordinates. This expression could be of independent interest; although there are many references in which the Taylor expansion of the metric is computed up to a certain order—to the best of our knowledge, always lower than 7th—the full Taylor series of \(g\) is not so easy to find in the literature. This will be done below. Then, in 13.2.2 we will obtain bounds for the absolute value of \(g^{ij}(x)\) and \(\partial_a g^{ij}(x)\), for every \(a, i, j \in \{1, \dotsc, d\}\) using the Taylor expansion of \(g\). These bounds will allow us to relate the Laplacian of \(\tilde{r}_{p, v}\) to the coefficients of the Riemann curvature tensor at \(p\) and the radius of the ball to which \(\tilde{r}_{p, v}\) is restricted.

13.2.1 Taylor expansion of the metric coefficients↩︎

As we mentioned earlier, we will give an explicit expression for the Taylor series of the coefficients of \(g\) in normal coordinates, when \((M, g)\) is a complete and connected symmetric space. This subsection is based in the proof of [85].

Proposition 119. Let \((M, g)\) be a complete and connected \(d\)-dimensional Riemannian manifold which is furthermore a symmetric space, and let \(p \in M\) be some fixed point in \(M\). For every \(x \in M\) outside the cut locus of \(p\), the coefficients \(g_{ij}(x)\) can be written in normal coordinates centered in \(p\) as \[g_{ij}(x) = \delta_{ij} + \sum_{n = 1}^{\infty} \frac{1}{2n!}g^{(2n)}_{ij}(0),\] where \[g^{(k)}_{ij}(t) := (\partial_t)^k g_{ij}(\gamma(t)),\] and \(\gamma\) is defined as \[\begin{align} \gamma: \mathbb{R} &\to M\\ t &\mapsto \exp_p(t\log_p x). \end{align}\]

Before we proceed to the proof of this result, let us prove some auxiliary lemmas, which will be fundamental when obtaining an explicit expression for the Taylor expansion of \(g\).

Definition 66. Let \((M, g)\) be a complete and connected Riemannian manifold. For every fixed \(p \in M\) and any \(x \in M\) outside the cut locus of \(p\), we define \(\varphi\) as \[\begin{align} \varphi: \mathbb{R}^2 &\to M\\ (s, t) &\mapsto \exp_p(tV(s)), \end{align}\] where \(V: \mathbb{R} \to T_p M\) is some smooth function satisfying \(V(0) = \log_p x\).

Notation 120. Let us denote the pullback connection by \(\varphi\) as \({}^{\varphi}\nabla\), which is defined as \[{}^{\varphi}\nabla_X Y|_{(s, t)} := \nabla_{\varphi_* X} Y|_{\varphi(s, t)},\] for every \(X \in \mathfrak{X}(\mathbb{R}^2)\) and any \(Y \in \mathfrak{X}(M)\).

Lemma 27. Let \((M, g)\) be a complete and connected Riemannian manifold. Let \(p \in M\) be fixed, let \(x \in M\) be outside the cut locus of \(p\) and let \(\varphi\) be as defined previously. Consider \((\partial_s, \partial_t)\) to be the (global) frame of \(\mathbb{R}^2\) associated with the coordinates \((s, t)\). Let us denote \(u := V(0), w:= V'(0)\). Then it holds that \[\begin{align} \label{eq:firstcovder} \begin{aligned} \varphi_* \partial_s|_{(0,0)} &= 0,\\ \varphi_* \partial_t|_{(0, 0)} &= u,\\ {}^{\varphi}\nabla_{\partial_t}\, \varphi_* \partial_t|_{(0,0)} & = 0,\\ {}^{\varphi}\nabla_{\partial_t}\,\varphi_* \partial_s|_{(0,0)} &= w. \end{aligned} \end{align}\qquad{(22)}\] Moreover, it follows that \({}^{\varphi}\nabla_{\partial_t}\, \varphi_* \partial_t \equiv 0\).

Proof. Writing \(\varphi\) in normal coordinates centered at \(p\) one gets \[\varphi(s, t) = (tV^1(s), \dotsc, tV^d(s)),\] so it is easy to see that the pushforward of the vector fields \(\partial_s\) and \(\partial_t\) by \(\varphi\) are given by \[\varphi_* \partial_s|_{(s,t)} = d\varphi(\partial_s)|_{(s, t)} = \partial_s \varphi(s, t) = t\,\partial_s V^i(s) \partial_i|_{\varphi(s, t)},\] and \[\varphi_* \partial_t|_{(s, t)} = V^{i}(s) \partial_i|_{\varphi(s, t)}.\]

Now, since in normal coordinates the Christoffel symbols all vanish at \(p = \varphi(0, 0)\), we can conclude that \[\nabla_X Y|_p = X(Y^k)\frac{\partial }{\partial x^k}, \quad \forall X, Y \in \mathfrak{X}(M).\] Therefore, \[\begin{align} {}^{\varphi}\nabla_{\partial_t}\, \varphi_* \partial_t|_{(0,0)} &= \nabla_{\varphi_* \partial_t} \,\varphi_* \partial_t|_{\varphi(0,0)} = (\partial_t V^i(s)) \partial_i|_{\varphi(0,0)} = 0,\\ {}^{\varphi}\nabla_{\partial_t}\,\varphi_* \partial_s|_{(0,0)} &= \nabla_{\varphi_* \partial_t} \,\varphi_* \partial_s|_{\varphi(0,0)} = \partial_t (t \partial_s V^i(s)) \partial_i|_{\varphi(0,0)} = \partial_s V^i(0) \partial_i|_{\varphi(0,0)} = w, \end{align}\] where we have used the fact that \((\varphi_* \partial_t) f = \partial_t (f \circ \varphi)\).

Moreover, since \(\gamma_{V(s)}(\cdot) := \varphi(s, \cdot)\) is a geodesic for every \(s\), it holds that \({}^{\varphi}\nabla_{\partial_t}\, \varphi_* \partial_t \equiv 0\), finishing the proof. ◻

Lemma 28. For every \(k \in \mathbb{N}\), it holds that \[{}^{\varphi}\nabla^{(2k)}_{\partial_t} \varphi_* \partial_s|_{(0,0)} = 0.\]

Proof. First, note that for each \(k \geq 2\) (cf. [85]) \[\label{eq:originalcovderkth} {}^{\varphi}\nabla^{(k)}_{\partial_t} \varphi_* \partial_s = \sum_{l = 0}^{k-2} {\binom{k-2}{l}} (\nabla^{(k-2-l)} R)_{\varphi} (\underbrace{\varphi_* \partial_t, \dotsc, \varphi_*\partial_t}_{k-2-l \text{ times}}, {}^{\varphi}\nabla^{(l)}_{\partial t} (\varphi_* \partial s), \varphi_* \partial_t)\varphi_* \partial_t,\tag{77}\] where \(R\) denotes the Riemann curvature tensor of \(M\).

Since \((M, g)\) is a complete and connected symmetric space by assumption, its Riemann curvature tensor is paralell, i.e. \(\nabla R \equiv 0\) (cf. [39]). Therefore, 77 simplifies to \[\label{eq:kthcovdertivative} {}^{\varphi}\nabla^{(k)}_{\partial_t} \varphi_* \partial_s = R_{\varphi} ({}^{\varphi}\nabla^{(k-2)}_{\partial t} (\varphi_* \partial s), \varphi_* \partial_t)\varphi_* \partial_t.\tag{78}\] Moreover, recall that \(\varphi_* \partial_s|_{\varphi(0,0)} = 0\) and so \[{}^{\varphi}\nabla^{(2)}_{\partial_t} \varphi_* \partial_s|_{(0,0)} = R_{\varphi(0,0)} (\varphi_* \partial s, \varphi_* \partial_t)\varphi_* \partial_t = R_{\varphi(0,0)}(0, u, u) = 0.\] The result now follows by induction on \(k\). ◻

With the two previous lemmas in mind, let us now prove 119.

Proof of 119. Let \(x \in M\) be outside the cut locus of \(p\), and consider the curve \(\gamma(t) = \varphi(0, t) = \exp_p(t\log_p x)\). We can write \[g_{ij}(\gamma(t)) = \delta_{ij} + \sum_{n = 1}^{\infty} \frac{1}{n!}g^{(n)}_{ij}(0) t^n,\] where \(g^{(n)}_{ij}\) is as defined in 119. Choosing \(t = 1\) we obtain that \[g_{ij}(x) = \delta_{ij} + \sum_{n = 1}^{\infty} \frac{1}{n!}g^{(n)}_{ij}(0).\]

In order to obtain an explicit formula for \(g^{(n)}_{ij}(0)\), we start by showing that, using the expression for \(\varphi_* \partial_s\) shown in 27, we have \[\label{eq:pullbackmetric} g_{\varphi(s, t)}(\varphi_* \partial_s, \varphi_* \partial_s) = t^2 \partial_s V^i(s) \partial_s V^j(s) g_{ij}(\varphi(s, t)).\tag{79}\] Thus, using the general Leibniz rule on 79 we find that for every \(k \geq 2\) \[\begin{align} (\partial_t)^{k} g_{\varphi(s, t)}(\varphi_* \partial_s, \varphi_* \partial_s) &= \big\{k(k-1) (\partial_t)^{k-2} g_{ij}(\varphi(s, t))\partial_s + 2k t (\partial_t)^{k-1} g_{ij}(\varphi(s, t)) \\&\quad + t^2 (\partial_t)^{k} g_{ij}(\varphi(s, t))\big\} \partial_s V^i(s) \partial_s V^j(s). \end{align}\] In particular, for \((s, t) = (0, 0)\), \[\label{eq:derivativemetric2} (\partial_t)^k g_{\varphi(0,0)}(\varphi_* \partial_s, \varphi_* \partial_s) = k(k-1) g^{(k-2)}_{ij}(0) \partial_s V^i(0) \partial_s V^j(0) = k (k-1) g^{(k-2)}_{ij}(0) w^i w^j,\tag{80}\] where \(w = V'(0)\).

On the other hand, since \({}^{\varphi}\nabla\) is compatible with \(g_{\varphi}\) (cf. [85]), it follows that \[\partial_t g_{\varphi(s, t)}(\varphi_* \partial_s, \varphi_* \partial_s) = g_{\varphi(s, t)}({}^{\varphi}\nabla_{\partial_t} \varphi_* \partial_s, \varphi_* \partial_s) + g_{\varphi(s, t)}(\varphi_* \partial_s, {}^{\varphi}\nabla_{\partial_t} \varphi_* \partial_s),\] and so, for every \(k \geq 0\) it holds that \[\label{eq:derivativemetric1} (\partial_t)^k g_{\varphi(s, t)}(\varphi_* \partial_s, \varphi_* \partial_s) = \sum_{l = 0}^k \binom{k}{l} g_{\varphi(s, t)}({}^{\varphi}\nabla^{(k-l)}_{\partial_t} \varphi_* \partial_s, {}^{\varphi}\nabla^{(l)}_{\partial_t} \varphi_* \partial_s).\tag{81}\]

Putting together 81 80 we conclude that for every \(k \geq 0\) \[\label{eq:expressionderivativemetric} (k+2)(k+1)g^{(k)}_{ij}(0) w^i w^j = \sum_{l = 0}^{k+2} \binom{k+2}{l} g_{\gamma(0)}({}^{\varphi}\nabla^{(k+2-l)}_{\partial t} \varphi_* \partial s, {}^{\varphi}\nabla^{(l)}_{\partial t} \varphi_* \partial s ).\tag{82}\] In particular, this identity implies that if \(k\) is odd, \[\label{eq:oddderivativesmetriczero} (k+2)(k+1)g^{(k)}_{ij}(0) w^i w^j = 0,\tag{83}\] since for every value of \(l\) either \(k+2-l\) or \(l\) is even, and so by 28 the right-hand side of 82 is zero.

If \(k\) is even, following an analogous reasoning we conclude that \[\label{eq:identitywithderivativesofg} (k+2)(k+1)g^{(k)}_{ij}(0) w^i w^j = \sum_{\substack{l = 1\\l \text{ odd}}}^{k+1} \binom{k+2}{l} g_{\gamma(0)}({}^{\varphi}\nabla^{(k+2-l)}_{\partial t} \varphi_* \partial s, {}^{\varphi}\nabla^{(l)}_{\partial t} \varphi_* \partial s ).\tag{84}\] Going back to the Taylor expansion of \(g_{ij}(x)\), we get that \[g_{ij}(x) = \delta_{ij} + \sum_{n = 1}^{\infty} \frac{1}{2n!}g^{(2n)}_{ij}(0),\] finishing the proof. ◻

Remark 121. Although we did not give an explicit formula for the terms in the Taylor series of \(g_{ij}(x)\) in the statement of 119, we did find such an explicit formula in the proof. In particular, we proved that for every \(n \in \mathbb{N}\), the coefficients \(g_{ij}^{2n}(0)\) satisfy 84 .

Before we obtain a bound for the Laplacian of the function \(\tilde{r}_{p, v}\), let us first give a simpler and more explicit expression for the even derivatives of the metric coefficients of \((M, g)\).

Proposition 122. Let \((M, g)\) be a complete and connected symmetric space of dimension \(d\). Let us denote by \(R\) its Riemann curvature tensor. Let \(p \in M\) be fixed and let \(x \in M\) be outside the cut locus of \(p\). Let \(\gamma(t) := \exp_p(t \log_p x)\). Then, in normal coordinates centered at \(p\), \[g^{(2n)}_{ij}(0) = \frac{2^{2n+1} }{(2n+2)(2n+1)} \delta_{\alpha j} T^\alpha_{i, i_1, \dotsc, i_{2n}} x^{i_1}\dotsm x^{i_{2n}}, \quad \forall n \geq 0, \forall i, j \in \{1, \dotsc, d\},\] where \[g^{(k)}_{ij}(t) = (\partial_t)^k g_{ij}(\gamma(t)),\] and \[\begin{align} T^\alpha_\beta &:= \delta^\alpha_\beta,\\ T^\alpha_{\beta i_1i_2} &:= R^\alpha_{\beta i_1 i_2},\\ T^\alpha_{\beta i_1, \dotsc, i_{2n}} &:= R^\alpha_{\alpha_1 i_1 i_2} \dotsm R^{\alpha_{n-2}}_{\alpha_{n-1} i_{2n-3} i_{2n-2}} R^{\alpha_{n-1}}_{\beta i_{2n-1} i_{2n}},\quad n \geq 2, \end{align}\] where the coefficients of \(R\) are considered in normal coordinates centered at \(p\) and evaluated at \(p\).

Proof. Let us define \(\varphi\) as in 27. We will obtain expressions for \({}^{\varphi}\nabla^{(n)}_{\partial_t} \varphi_* \partial_s|_{(0, 0)}\) in normal coordinates, for odd values of \(n\)—the even case was studied in 28. Using ?? 78 we see that, in normal coordinates centered at \(p\), \[\begin{align} {}^{\varphi}\nabla_{\partial_t}\, \varphi_* \partial_s|_{(0, 0)} &= w^b \partial_b,\\ {}^{\varphi}\nabla^{(3)}_{\partial_t} \varphi_* \partial_s|_{(0, 0)} &= R_{\varphi(0,0)} ({}^{\varphi}\nabla^{}_{\partial t} (\varphi_* \partial s), u)u = R_{\varphi(0,0)} (w, u)u = R_{ai_1 i_2}^b w^a x^{i_1} x^{i_2} \partial_b,\\ {}^{\varphi}\nabla^{(2n+1)}_{\partial_t} \varphi_* \partial_s|_{(0, 0)} &= R^b_{\alpha_1 i_1 i_2} \dotsm R^{\alpha_{n-2}}_{\alpha_{n-1} i_{2n-3} i_{2n-2}} R^{\alpha_{n-1}}_{a i_{2n-1} i_{2n}} w^a x^{i_1} \dotsm x^{i_{2n}} \partial_b, \quad n \geq 2, \end{align}\] where again \(w = V'(0)\) and \(u = \log_p x\). To simplify the notation, we can write these expressions using the notation from the statement of the proposition, \[{}^{\varphi}\nabla^{(2n+1)}_{\partial_t} \varphi_* \partial_s|_{(0, 0)} = T^\alpha_{\beta i_1, \dotsc, i_{2n}} w^\beta x^{i_1}\dotsm x^{i_{2n}}\partial_\alpha,\] for every \(n \geq 0\).

This way, since in normal coordinates \(g_{\varphi(0,0)}\) is the Euclidean metric, for every \(n \in \mathbb{N}\) and every odd value \(2l+1\), \(l \geq 0\) we can write \[\begin{align} g_{\gamma(0)}({}^{\varphi}\nabla^{(2n+2-(2l+1))}_{\partial t} \varphi_* \partial s, {}^{\varphi}\nabla^{(2l+1)}_{\partial t} \varphi_* \partial s ) &= \delta_{\alpha \beta}T^\alpha_{i, i_1, \dotsc, i_{2(n-l)}} T^\beta_{j, j_1, \dotsc, j_{2l}} \\&\quad \times w^i w^j x^{i_1}\dotsm x^{i_{2(n-l)}} x^{j_1}\dotsm x^{j_{2l}}\\ &= \delta_{\alpha \beta}T^\alpha_{i, i_1, \dotsc, i_{2(n-l)}} T^\beta_{j, i_{2(n-l)+1}, \dotsc, i_{2n}} w^i w^j x^{i_1}\dotsm x^{i_{2n}} \\&= \delta_{\alpha j} T^\alpha_{i, i_1, \dotsc, i_{2n}} w^i w^j x^{i_1}\dotsm x^{i_{2n}}, \end{align}\] where the last equality follows simply by re-ordering the indices. Inserting this identity into 84 , since \(w\) was arbitrary, we obtain that for every \(i, j \in \{1, \dotsc, d\}\) \[\begin{align} \label{eq:expansiontaylorterms} \begin{aligned} (2n+2)(2n+1)g^{(2n)}_{ij}(0) &= \sum_{\substack{l = 1\\l \text{ odd}}}^{2n+1} \binom{2n+2}{l}\delta_{\alpha j} T^\alpha_{i, i_1, \dotsc, i_{2n}} x^{i_1}\dotsm x^{i_{2n}} \\&= 2^{2n+1} \delta_{\alpha j} T^\alpha_{i, i_1, \dotsc, i_{2n}} x^{i_1}\dotsm x^{i_{2n}}. \end{aligned} \end{align}\tag{85}\]  ◻

In the following proposition, we will obtain a similar expression for the derivatives of the coefficients of the inverse metric in normal coordinates. Although this will not be used to study the Laplacian or the gradient of \(\tilde{r}_{p, v}\), it will play an essential role in 13.4.2.

Proposition 123. Let \((M, g)\) be a complete and connected \(d\)-dimensional Riemannian manifold which is furthermore a symmetric space, and let \(p \in M\) be some fixed point in \(M\). For any \(x \in M\) outside the cut locus of \(p\), taking normal coordinates centered at \(p\), we denote the coefficients of the inverse metric \(g^{-1}\) at \(x\) as \(g^{ij}(x)\).

Let \((g^{(k)}(0))^{ij}\) be the analogous of \(g^{(k)}_{ij}(0)\), i.e. \[(g^{(k)}(0))^{ij} := \left.\frac{d^k}{dt^k}\right|_{t = 0} g^{ij}(\gamma(t)),\] where \(\gamma(t) = \exp_p(t \log_p x)\). Then for every \(n \geq 0\) and every \(i, j \in \{1, \dotsc, d\}\) it holds that \[\begin{align} (g^{(2n+1)}(0))^{ij} &= 0,\\ \big(g^{(2n)}(0)\big)^{ij} &= C_{2n} \delta^{i\beta} T^j_{\beta, i_1, \dotsc, i_{2n}} x^{i_1}\cdots x^{i_{2n}}, \end{align}\] where \[\begin{align} C_0 &:= 1,\\ C_{2n} &:= -\sum_{m = 1}^n {\binom{2n}{2m}} C_{2(n-m)} \frac{2^{2m+1}}{(2m+2)(2m+3)}. \end{align}\]

Proof. By definition of the inverse metric, it holds that \[g_{ik}(\gamma(t)) g^{kj}(\gamma(t)) = \delta_i^j, \quad \forall i, j \in \{1, \dotsc, d\}.\] Taking the differential with respect to \(t\), at \(t = 0\), we conclude that \[\label{eq:firstderivativegginverse} g^{(1)}_{ik}(0) g^{kj}(\gamma(0)) + g_{ik}(\gamma(0)) (g^{(1)}(0))^{kj} = 0, \quad \forall i, j \in \{1, \dotsc, d\}.\tag{86}\] Now, recall that \(g_{ik}(\gamma(0)) = \delta_{ik}\) and \(g^{(1)}_{ik}(0) = 0\) for every \(i, k \in \{1, \dotsc, d\}\), which allows us to rewrite 86 as \[\delta_{ik} (g^{(1)}(0))^{kj} = 0, \quad \forall i, j \in \{1, \dotsc, d\}.\] Therefore, we conclude that \[(g^{(1)}(0))^{ij} = 0, \quad \forall i, j \in \{1, \dotsc, d\}.\]

We can now prove that \[(g^{(2n+1)}(0))^{ij} = 0,\quad \forall n \geq 0,\; \forall i, j \in \{1, \dotsc, d\}\] by induction. Indeed, for every \(i, j \in \{1, \dotsc, d\}\) using the general Leibniz rule it holds that \[\label{eq:oddderivativecontractionmetric} 0 = \left.\frac{d^{2n-1}}{dt^{2n-1}}\right|_{t = 0} \Big(g_{ik}(\gamma(t)) g^{kj}(\gamma(t))\Big) = \sum_{m = 0}^{2n-1} \binom{2n-1}{m} g^{(2n-1-m)}_{ik}(0) (g^{(m)}(0))^{kj}.\tag{87}\] In every summand on the right-hand side of this identity either \(m\) or \(2n-1-m\) is odd, and so by induction hypothesis, using the fact that \(g^{(2n-1)}_{ij}(0) = 0\), for every \(n \geq 1\) (cf. 83 ) we can rewrite 87 as \[0 = \delta_{ik} (g^{(2n+1)}(0))^{kj},\] and the claim follows.

Let us now consider the even derivatives of the inverse metric coefficients. We proceed by induction. For \(n = 0\), it holds that \[\delta^i_j = \big(g^{}(0)\big)^{ik}\delta_{kj}, \quad \forall i, j \in \{1, \dotsc, d\},\] which allows us to conclude that \[\big(g^{}(0)\big)^{ij}= \delta^{ij}, \quad \forall i, j \in \{1, \dotsc, d\}.\] On the other hand, \[C_{0} \delta^{i\beta} T^j_{\beta} = \delta^{i\beta}\delta^j_\beta = \delta^{ij}, \quad \forall i, j \in \{1, \dotsc, d\},\] which proves the base case.

Now, for \(n \geq 1\), using the general Leibniz rule, we know that for every \(i, j \in \{1, \dotsc, d\}\), \[\begin{align} (g^{(2n)}(0))^{ik}\delta_{kj} &= - \sum_{m = 1}^n {\binom{2n}{2m}} (g^{2(n-m)}(0))^{ik} g_{kj}^{(2m)}(0) \\&= - \sum_{m = 1}^n {\binom{2n}{2m}} (g^{2(n-m)}(0))^{ik} \frac{2^{2m+1}}{(2m+2)(2m+1)} \delta_{\alpha j} T^\alpha_{k, i_1, \dotsc, i_{2m}} x^{i_1}\dotsm x^{i_{2m}} \\&= - \left(\sum_{m = 1}^n {\binom{2n}{2m}} C_{2(n-m)} \frac{2^{2m+1}}{(2m+2)(2m+1)} \delta^{i\beta} \delta_{\alpha j} T^k_{\beta, j_1, \dotsc, j_{2(n-m)}} T^\alpha_{k, i_1, \dotsc, i_{2m}}\right) \\&\quad \times x^{j_1}\cdots x^{j_{2(n-m)}} x^{i_1}\dotsm x^{i_{2n}}, \end{align}\] where the second equality follows from 122 and the third equality follows from the induction hypothesis.

By simply reordering the indices, it holds that \[T^k_{\beta, j_1, \dotsc, j_{2(n-m)}} T^\alpha_{k, i_1, \dotsc, i_{2m}} x^{j_1}\cdots x^{j_{2(n-m)}} x^{i_1}\dotsm x^{i_{2m}} = T^\alpha_{\beta, i_1, \dotsc, i_{2n}}x^{i_1}\dotsm x^{i_{2n}},\] and so we can write \[\begin{align} (g^{(2n)}(0))^{ik}\delta_{kj} = - \sum_{m = 1}^n {\binom{2n}{2m}} C_{2(n-m)} \frac{2^{2m+1}}{(2m+2)(2m+1)} \delta^{i\beta} \delta_{\alpha j} T^\alpha_{\beta, i_1, \dotsc, i_{2n}}x^{i_1}\dotsm x^{i_{2n}}, \end{align}\] for every \(i, j \in \{1, \dotsc, d\}\). The result now follows by rewriting this expression as \[(g^{(2n)}(0))^{ij} = - \sum_{m = 1}^n {\binom{2n}{2m}} C_{2(n-m)} \frac{2^{2m+1}}{(2m+2)(2m+1)} \delta^{i\beta} T^j_{\beta, i_1, \dotsc, i_{2n}}x^{i_1}\dotsm x^{i_{2n}}.\] and defining \[C_{2n} := - \sum_{m = 1}^n {\binom{2n}{2m}} C_{2(n-m)} \frac{2^{2m+1}}{(2m+2)(2m+1)}.\] ◻

13.2.2 Bounding the Laplacian↩︎

Having obtained an explicit expression for the Taylor series of the coefficients of the metric \(g\) i normal coordinates, we now proceed to upper bound the absolute value of \(g_{ij}(x)\) and \(\partial_a g_{ij}(x)\) for every \(a, i, j \in \{1, \dotsc, d\}\). By bounding the former, we will be able to upper bound the absolute value of the coefficients of the inverse metric, which will allow us to upper bound the absolute value of the Christoffel symbols, and ultimately \(|\Delta_g \tilde{r}_{p, v}(x)|\).

Proposition 124. Let \((M, g)\) be a complete and connected \(d\)-dimensional Riemannian manifold which is also a symmetric space. Let \(p \in M\) be a fixed point and let us consider normal coordinates centered at \(p\). Let \(0 < r < i(M)\) and let \(x \in M\) be such that \(d_g(x, p) \leq r\). Assume that, at \(p\), there exists some constant \(\mathbf{K} \geq 1\) such that the coefficients of the Riemann curvature tensor in these coordinates satisfy \(|R_{ijk}^l(p)| \leq \mathbf{K}\) for every \(i, j, k, l \in \{1, \dotsc, d\}\). Then \[\left|\sum_{n = 1}^{\infty} \frac{1}{2n!}g^{(2n)}_{ij}(0)\right| \leq\frac{2}{d} \sum_{n = 1}^{\infty}\frac{d^{3n} \mathbf{K}^n (2r)^{2n} }{(2n+2)!},\] for every \(i, j \in \{1, \dotsc, d\}\).

Proof. In 122 we showed that for every \(n \in \mathbb{N}\) \[g^{(2n)}_{ij}(0) = \frac{2^{2n+1} }{(2n+2)(2n+1)} \delta_{\alpha j} T^\alpha_{i, i_1, \dotsc, i_{2n}} x^{i_1}\dotsm x^{i_{2n}}, \quad \forall i, j \in \{1, \dotsc, d\}.\] It is now straightforward to obtain a bound for this expression. As in normal coordinates centered at \(p\) it holds that \(|R_{ijk}^l(p)| \leq \mathbf{K}\) for every \(i, j, k, l \in \{1, \dotsc, d\}\), it follows that for each \(n \in \mathbb{N}\) \[\label{eq:tensorTbound} |T^\alpha_{\beta i_1, \dotsc, i_{2n}} | \leq d^{n-1} \mathbf{K}^{n}, \quad \forall \alpha, \beta, i_1, \dotsc, i_{2n} \in \{1, \dotsc, d\}.\tag{88}\] Since we are also assuming that \(d_g(p, x) = |\log_p x|_g \leq r\), it follows that for every \(n \in \mathbb{N}\) \[\begin{align} (2n+2)(2n+1)|g^{(2n)}_{ij}(0)| &\leq 2^{2n+1} d^{2n} d^{n-1} \mathbf{K}^{n} r^{2n} \\&= 2^{2n+1} d^{3n-1} \mathbf{K}^n r^{2n}, \end{align}\] for every \(i, j \in \{1, \dotsc, d\}\).

This way, we proved that for every \(i, j \in \{1, \dotsc, d\}\) \[\left|\sum_{n = 1}^{\infty} \frac{1}{2n!}g^{(2n)}_{ij}(0)\right| \leq \frac{2}{d} \sum_{n = 1}^{\infty}\frac{d^{3n} \mathbf{K}^n (2r)^{2n} }{(2n+2)!},\] finishing the proof. ◻

124 allows us to obtain bounds which can be as small as desired, simply by requiring \(r\) to be small enough.

Remark 125. Under the conditions of 124, let \[r \leq \frac{1}{d^{3/2} \mathbf{K}^{1/2}}.\] Then \[\left|\sum_{n = 1}^{\infty} \frac{1}{2n!}g^{(2n)}_{ij}(0)\right| < \frac{1}{d}.\]

Proof. By simply inserting the bound on \(r\) into the expression found in 124 we see that \[\begin{align} \left|\sum_{n = 1}^{\infty} \frac{1}{2n!}g^{(2n)}_{ij}(0)\right| &\leq \frac{2}{d} \sum_{n = 1}^{\infty}\frac{d^{3n} \mathbf{K}^n (2r)^{2n} }{(2n+2)!} \\&\leq \frac{2}{d} \sum_{n = 1}^{\infty}\frac{d^{3n} \mathbf{K}^n (2 \frac{1}{d^{3/2} \mathbf{K}^{1/2}})^{2n} }{(2n+2)!} \\&= \frac{2}{d} \sum_{n = 1}^{\infty}\frac{4^n}{(2n+2)!} \\&\leq \frac{1}{d}. \end{align}\] ◻

As we mentioned earlier, we need to find an upper bound on the absolute value of the coefficients of \(g^{-1}\) at \(x\) in normal coordinates, as they appear both in the expression of the Christoffel symbols and in the Laplacian of \(\tilde{r}_{p, v}\). In order to do so, we will use two auxiliary lemmas.

Lemma 29. For every matrix \(A\) of size \(n \times m\) it holds that \[\max_{ij} |a_{ij}| \leq \norm{A}_F \leq \sqrt{nm} \max_{ij} |a_{ij}|,\] where \(\norm{A}_F\) denotes the Frobenius norm of \(A\), namely \[\require{physics} \norm{A}_F := \Tr(A^T A) = \Big(\sum_{ij} |a_{ij}|^2\Big)^{1/2}.\]

The next lemma will allow us to obtain a Taylor series expansion for the coefficients of \(g^{-1}\) in normal coordinates using what is known as a Neumann series approximation. We refer the reader to [86] for more details.

Lemma 30. Let \(X\) be some matrix such that \(\norm{X} < 1\) for some matrix norm \(\norm{\cdot}\). Then \((\mathbb{1}-X)\) is non-singular and \[(\mathbb{1} - X)^{-1} = \sum_{k = 0}^{\infty} X^k.\]

As we mentioned previously, these lemmas allow us to prove the following result, which gives us a bound on the absolute value of the coefficients of \(g^{-1}\) in normal coordinates.

Proposition 126. Let \((M, g)\) be a complete and connected \(d\)-dimensional Riemannian manifold which is also a symmetric space. Let \(p \in M\) be a fixed point and let us consider normal coordinates centered at \(p\). Let \(0 < r < i(M)\) and let \(x \in M\) be such that \(d_g(x, p) \leq r\). Assume that, at \(p\), there exists some constant \(\mathbf{K} \geq 1\) such that the coefficients of the Riemann curvature tensor in these coordinates satisfy \(|R_{ijk}^l(p)| \leq \mathbf{K}\) for every \(i, j, k, l \in \{1, \dotsc, d\}\). Further assume that \[r \leq \frac{1}{d^{3/2} \mathbf{K}^{1/2}}.\] Then for every \(i, j \in \{1, \dotsc, d\}\), \[|g^{ij}(x) | \leq \frac{1}{1-L},\] where \[\label{eq:defconstanteL} L := 2\sum_{n = 1}^{\infty}\frac{d^{3n} \mathbf{K}^n (2r)^{2n} }{(2n+2)!}.\qquad{(23)}\]

Proof. First, we can write \(g\) as \[g(x) = \mathbb{1} + X,\] where \[X_{ij} = g_{ij} - \delta_{ij} = \sum_{n = 1}^\infty \frac{g^{(2n)}_{ij}(0)}{2n!},\] for every \(i, j \in \{1, \dotsc, d\}\).

From 124 we can conclude that for every \(i, j \in \{1, \dotsc, d\}\), \[|X_{ij}| = |g_{ij} - \delta_{ij}| \leq \frac{L}{d},\] where \(L\) is as defined in ?? . Therefore, by 125 29 we can conclude that \[\label{eq:normboundX} \norm{X}_F \leq d \max_{ij}|X_{ij}| \leq L < 1,\tag{89}\] and so by 30 the metric \(g\) is non-singular and \[g^{-1} = (\mathbb{1} + X)^{-1} = \sum_{k = 0}^{\infty} (-1)^k X^k,\] which implies that \[\label{eq:normboundinverseg} \norm{g^{-1}}_F \leq \sum_{k= 0}^{\infty} \norm{X}^k_F \leq \frac{1}{1-L}.\tag{90}\]

Lastly, we can use 29 again to guarantee that for every \(i, j \in \{1, \dotsc, d\}\) \[|g^{ij}(x)| \leq \max_{ij}|g^{ij}(x)| \leq \norm{g^{-1}}_F \leq \frac{1}{1-L}.\] ◻

To finish, we will follow the same steps as in 124 to give an upper bound on \(|\partial_a g_{ij}(x)|\) in normal coordinates centered at \(p\), depending on \(d_g(x, p)\) and the bound on the Riemann curvature tensor.

Proposition 127. Let \((M, g)\) be a complete and connected \(d\)-dimensional Riemannian manifold which is also a symmetric space. Let \(p \in M\) be a fixed point and let us consider normal coordinates centered at \(p\). Let \(0 < r < i(M)\) and let \(x \in M\) be such that \(d_g(x, p) \leq r\). Assume that, at \(p\), there exists some constant \(\mathbf{K} \geq 1\) such that the coefficients of the Riemann curvature tensor in these coordinates satisfy \(|R_{ijk}^l(p)| \leq \mathbf{K}\) for every \(i, j, k, l \in \{1, \dotsc, d\}\). Then \[|\partial_a g_{ij}(x)| \leq \sum_{n = 1}^\infty \frac{n 2^{2n+2} d^{3n-2} \mathbf{K}^{n} r^{2n-1}}{(2n+2)!},\] for every \(a, i, j \in \{1, \dotsc, d\}\).

Proof. For every \(a, i \in \{1, \dotsc, d\}\), we can write \[\partial_a x^i = \delta^i_a.\] Therefore, if we take the derivative with respect to the \(a\)-th coordinate in 85 we obtain that for every \(n \in \mathbb{N}\) and every \(i, j \in \{1, \dotsc, d\}\) \[(2n+2)(2n+1)\partial_a g^{(2n)}_{ij}(0) = 2^{2n+1} \delta_{\alpha j} T^\alpha_{i, i_1, \dotsc, i_{2n}} x^{i_1, \dotsc, i_{2n}}_a\] where \[x^{i_1,\dotsc,i_{2n}}_a := \delta^{i_1}_a x^{i_2} \dotsm x^{i_{2n}} + x^{i_1}\delta^{i_2}_a x^{i_3} \dotsm x^{i_{2n}} + \dotsm + x^{i_1} \dotsm x^{i_{2n-1}} \delta^{i_{2n}}_a.\] The norm of these elements can be bounded as \[\begin{align} |x^{i_1,\dotsc,i_{2n}}_a| &\leq |\delta^{i_1}_a x^{i_2} \dotsm x^{i_{2n}}| + |x^{i_1}\delta^{i_2}_a x^{i_3} \dotsm x^{i_{2n}}| + \dotsm + |x^{i_1} \dotsm x^{i_{2n-1}} \delta^{i_{2n}}_a|\\ &\leq 2n r^{2n-1}. \end{align}\] Therefore, using this inequality together with 88 we conclude that for every \(n \geq 1\) and every \(a, i, j \in \{1, \dotsc, d\}\), \[\begin{align} (2n+2)(2n+1)|\partial_a g^{(2n)}_{ij}(0)| &= |2^{2n+1} \delta_{\alpha j} T^\alpha_{i, i_1, \dotsc, i_{2n}} x^{i_1, \dotsc, i_{2n}}_a| \\&\leq 2^{2n+1} |T^j_{i, i_1, \dotsc, i_{2n}}| |x^{i_1, \dotsc, i_{2n}}_a| \\&\leq 2^{2n+1} d^{2n-1} d^{n-1} \mathbf{K}^n 2nr^{2n-1} \\&= 2^{2n+2} n d^{3n-2} \mathbf{K}^n r^{2n-1}, \end{align}\] and so \[\begin{align} |\partial_a g_{ij}(x)| &\leq \sum_{n = 1}^\infty \frac{1}{2n!} |\partial_ag^{(2n)}_{ij}(0)| \\&\leq \sum_{n = 1}^\infty \frac{n 2^{2n+2} d^{3n-2} \mathbf{K}^{n} r^{2n-1}}{(2n+2)!}, \end{align}\] for every \(a, i, j \in \{1, \dotsc, d\}\). ◻

Having obtained a bound on \(|\partial_a g_{ij}(x)|\) and \(|g^{ij}(x)|\) for every \(a, i, j \in \{1, \dotsc, d\}\), the bound on the Laplacian of \(\tilde{r}_{p, v}\) follows almost immediately, as we will show in the next theorem. Note that the bound for the absolute value of the Laplacian can be made arbitrarily small by shrinking \(r\). Nevertheless, the bound we obtain below is sufficiently good for the purpose of this work.

Proposition 128 (Upper bound on the absolute value of the Laplacian of \(\tilde{r}_{p, v}\)). Let \((M, g)\) be a complete and connected \(d\)-dimensional Riemannian manifold, which is also a symmetric space. Let \(p \in M\) be a fixed point and let us consider normal coordinates centered at \(p\). Let \(0 < r < i(M)\) and let \(x \in M\) be such that \(d_g(x, p) \leq r\). Assume that, at \(p\), there exists some constant \(\mathbf{K} \geq 1\) such that the coefficients of the Riemann curvature tensor in these coordinates satisfy \(|R_{ijk}^l(p)| \leq \mathbf{K}\) for every \(i, j, k, l \in \{1, \dotsc, d\}\). Then if \[r \leq \frac{1}{d^{3/2} \mathbf{K}},\] it holds that \[|\Delta_g \tilde{r}_{p, v}(x)| \leq 24 d^{5/2}.\]

Proof. In normal coordinates, the Christoffel symbols can be written as \[\Gamma^1_{kl}(x) = \frac{1}{2} g^{1m}(x) \Big(\partial_l g_{mk}(x) + \partial_k g_{ml}(x) - \partial_m g_{kl}(x)\Big),\] for every \(k, l \in \{1, \dotsc, d\}\). Thus, we can rewrite 76 as \[\label{eq:explicitLaplacianr} \Delta_g \tilde{r}_{p, v}(x) = -g^{ij}(x) \Gamma^1_{ij}(x) = -\frac{1}{2}g^{ij}(x)g^{1k}(x)(\partial_j g_{ki}(x) + \partial_i g_{kj}(x) - \partial_k g_{ij}(x)).\tag{91}\]

By 127, we know that for every \(a, i, j \in \{1, \dotsc, d\}\), \[|\partial_a g_{ij}(x)| \leq \sum_{n = 1}^\infty \frac{n 2^{2n+2} d^{3n-2} \mathbf{K}^{n} r^{2n-1}}{(2n+2)!}.\] Therefore, since \[r \leq \frac{1}{d^{3/2} \mathbf{K}},\] and since we are assuming without loss of generality that \(\mathbf{K} \geq 1\), we see that \[\begin{align} \begin{aligned} \label{eq:ineqvaluederivativemetric} |\partial_a g_{ij}(x)| &\leq \sum_{n = 1}^\infty \frac{n 2^{2n+2} d^{3n-2} \mathbf{K}^{n} (\frac{1}{d^{3/2} \mathbf{K}})^{2n-1}}{(2n+2)!} \\&= \sum_{n = 1}^\infty \frac{n 2^{2n+2} d^{-1/2} \mathbf{K}^{-n+1}}{(2n+2)!} \\&\leq \frac{4}{\sqrt{d}} \sum_{n = 1}^\infty \frac{n 4^{n}}{(2n+2)!} \\&\leq \frac{4}{\sqrt{d}}. \end{aligned} \end{align}\tag{92}\] Now, we can use 126 to conclude that for every \(i, j \in \{1, \dotsc, d\}\) \[|g^{ij}(x)| \leq \frac{1}{1-L},\] where \(L\) is as defined in ?? .

Again, using that \[r \leq \frac{1}{d^{3/2} \mathbf{K}} \leq \frac{1}{d^{3/2} \mathbf{K}^{1/2}},\] we can conclude that \[L \leq 2 \sum_{n = 1}^{\infty}\frac{4^{n}}{(2n+2)!} < \frac{1}{2}.\] and so \[\label{eq:ineqvalueinversemetric} |g^{ij}(x)| \leq \frac{1}{1-\frac{1}{2}} = 2.\tag{93}\]

Inserting 92 93 into 91 it follows that \[|\Delta_g \tilde{r}_{p, v}(x)| \leq 24 d^{5/2},\] concluding the proof. ◻

13.3 Bounding the norm of the gradient in symmetric spaces↩︎

To finish, let us focus on the gradient of \(\tilde{r}_{p, v}(x)\), for which we will obtain a lower and an upper bound on its norm. To this end, we will use the bounds on the metric derived in the previous section.

Lemma 31. Let \((M, g)\) be a complete and connected Riemannian manifold which is also a symmetric space of dimension \(d\). Let \(p \in M\) be a fixed point and let us consider normal coordinates centered at \(p\). Let \(0 < r < i(M)\) and let \(x \in M\) be such that \(d_g(x, p) \leq r\). In these coordinates, the gradient of \(\tilde{r}_{p, v}\) can be written as \[\mathrm{grad}_{g}\, \tilde{r}_{p, v}(x) = g^{ij}(x) \delta_{ik} v^k \partial_j,\] where \(v = v^i \partial_i \in T_p M\) is the vector with respect to which \(\tilde{r}_{p, v}\) is defined.

Proof. In normal coordinates centered at \(p\), \(\tilde{r}_{p, v}(x) = \langle v, x\rangle = \delta_{ij} v^i x^j\). Recall that, in these coordinates, \[\mathrm{grad}_{g}\, \tilde{r}_{p, v}(x) = g^{ij}(x) \frac{\partial \tilde{r}(x)}{\partial x^i} \partial_j.\] Therefore, using that \(\partial_a x^i = \delta_a^i\) for every \(a, i \in \{1, \dotsc, d\}\) it holds that \[\partial_a \tilde{r}_{p, v}(x) = \delta_{ai} v^i,\] and so \[\mathrm{grad}_{g}\, \tilde{r}_{p, v}(x) = g^{ij}(x) \delta_{ik} v^k \partial_j.\] ◻

Proposition 129. Under the conditions of 31, assume that there exists a constant \(\mathbf{K} \geq 1\) such that for every \(x \in M\), in normal coordinates centered at \(x\), the coefficients of the Riemann curvature tensor are such that \(|R^l_{ijk}(x)| \leq \mathbf{K}\) for every \(i, j, k, l \in \{1, \dotsc, d\}\). If \[r \leq \frac{1}{d^{3/2} \mathbf{K}^{1/2}},\] then \[\frac{1}{1 + L} \leq |\mathrm{grad}_{g}\, \tilde{r}_{p, v}(x)|^2_g \leq \frac{1}{1 - L},\] where \[L = 2\sum_{n = 1}^{\infty}\frac{d^{3n} \mathbf{K}^n (2r)^{2n} }{(2n+2)!}.\]

Proof. In 31 we showed that \[\mathrm{grad}_{g}\, \tilde{r}_{p, v}(x) = g^{ij}(x) \delta_{ik} v^k \partial_j,\] and so \[|\mathrm{grad}_{g}\,\tilde{r}_{p, v}(x)|^2_g = g_{ij}(x) g^{ai}(x) \delta_{ab}v^b g^{cj}(x) \delta_{cd}v^d = \delta_{jb} \delta_{cd} g^{cj}(x) v^b v^d = \sum_{ij} g^{ij}(x) v^i v^j.\] This expression can be easily lower bounded, \[\begin{align} |\mathrm{grad}_{g}\,\tilde{r}_{p, v}(x)|^2_g = \sum_{ij} g^{ij}(x) v^i v^j \geq \lambda_{\min}(g^{-1}(x)) |v|_g^2 = \frac{1}{\lambda_{\max}(g(x))}, \end{align}\] as \(|v|_g^2 = 1\).

Note that, as \(g\) is symmetric and positive-definite, its singular values correspond to its eigenvalues, and so \[\norm{g}_2 := \sigma_{\max}(g) = \lambda_{\max}(g),\] where \(\sigma_{\max}(g)\) denotes the greatest singular value of \(g\) and \(\lambda_{\max}(g)\) denotes its greatest eigenvalue. Moreover, this norm is upper bounded by the Frobenius norm from 29, which can be written as \[\norm{g}_F = \sqrt{\sigma_1(g)^2 + \dotsc + \sigma_d(g)^2 } = \sqrt{\lambda_1(g)^2 + \dotsc + \lambda_d(g)^2},\] and so \(\norm{g}_2 \leq \norm{g}_F\). Thus, using 89 we can conclude that \[\lambda_{\max}(g(x)) = \norm{g}_2 = \norm{\mathbb{1} + X}_2 \leq \norm{1}_2 + \norm{X}_2 \leq 1 + \norm{X}_F \leq 1 + L,\] and so \[|\mathrm{grad}_{g}\,\tilde{r}_{p, v}(x)|^2_g \geq \frac{1}{1 + L}.\]

Similarly, \[|\mathrm{grad}_{g}\,\tilde{r}_{p, v}(x)|^2_g \leq \lambda_{\max}(g^{-1}(x)) = \norm{g^{-1}(x)}_2 \leq \norm{g^{-1}(x)}_F \leq \frac{1}{1-L},\] where the last inequality follows from 90 . ◻

13.4 Two auxiliary lemmas↩︎

We now state and prove the auxiliary lemmas 33 and 32 used in 2.2.2. We begin with 32, as its proof is simpler.

Let us recall the framework in which the lemmas are formulated. Let \(\pi:(M, g) \to (B, h)\) be a Riemannian submersion with totally geodesic fibers, where \((M, g)\) is furthermore a complete and connected symmetric space. Let us fix some \(p \in B\), and define the following two functions for every \(q\in B\) outside the cut locus of \(p\), \[\tilde{r}_{p, v}(q) = \langle v, \log_p q\rangle,\] and \[H_p(q) = \nabla^2 \tilde{F}(p)[\log_p q],\] where \(\tilde{F}\) is some suitable smooth function on \(B\), and \(v \in T_p B\) is some fixed tangent vector.

To analyze the above functions, we will lift them from \(B\) to \(M\), study their lifted versions, and then transfer the properties of the lifted functions to the original ones using the results from 8.5.

13.4.1 First auxiliary lemma↩︎

Lemma 32. Let \(\pi: (M, g) \to (B, h)\) be a surjective Riemannian submersion with totally geodesic fibers, where \((M, g)\) is a compact and connected symmetric space. Assume that there exists some constant \(\mathbf{K} \geq 1\) such that, for every \(x \in M\), the coefficients of the Riemann curvature tensor in normal coordinates centered at \(x\) satisfy \(|R_{ijk}^l(x)| \leq \mathbf{K}\) for every \(i, j, k, l \in \{1, \dotsc, \dim(M)\}\). Let \[r \leq \min\{i(M), i(B), \frac{1}{72 \dim(M)^{5/2} \mathbf{K}}\}.\] Let \(p \in B\) be some fixed point of \(B\), and let \(v \in T_p B\) be a unitary tangent vector at \(p\). Let us consider \[\tilde{r}_{p, v}(q) = \langle v, \log_p q\rangle,\] for every \(q \in B\) outside the cut locus of \(p\). Then it holds that \[\frac{10}{11} \leq |\mathrm{grad}_h \tilde{r}_{p, v}(q)|^2_h \leq \frac{10}{9},\] and \[|\mathrm{grad}_h \tilde{r}_{p, v}(q)|^2_h + \tilde{r}_{p, v}(q)\Delta_g \tilde{r}_{p, v}(q) \geq \frac{1}{2},\] for every \(q \in \mathcal{B}(r, p)\).

As we mentioned earlier, to obtain the above bounds, we will first obtain a lifted version of \(\tilde{r}_{p, v}\), which we will denote \(\tilde{r}^M_{z, v}\). We will then use the results from [sec:boundingLaplaciansymmetricspaces] [sec:boundinggradientsymmetric] to obtain bounds for the norm of the gradient and the Laplacian of \(\tilde{r}^M_{z, v}\), using the fact that \(\tilde{r}^M_{z, v}\) is defined on a symmetric space. Lastly, we will obtain the desired bounds for \(\tilde{r}_{p, v}\) using 8.5.

In the following proposition we explicitly construct the function which is the lifted version of \(\tilde{r}_{p, v}\).

To clarify the notation, we will use \(\log^B\) and \(\log^M\) throughout the subsection for the inverses of the exponential maps in \(B\) and \(M\), respectively.

Proposition 130. Let \(\pi: (M, g) \to (B, h)\) be a surjective Riemannian submersion. Let \(p \in B\) and let \(z \in \pi^{-1}(p)\) be fixed. Let \(\alpha \leq \min\{i(M), i(B)\}\). Then for every \(q \in \mathcal{B}(\alpha, p)\) and every \(x, y \in \pi^{-1}(q) \cap \mathcal{B}(\alpha, z)\), it holds that \[(\log^M_z x)^{\mathcal{H}} = (\log^M_z y)^{\mathcal{H}},\] and \[\pi_* ((\log^M_z x)^{\mathcal{H}}) = \log_p^B q.\]

In particular, given a fixed point \(p \in B\), it holds that for every \(v \in T_p B\) and every \(z \in \pi^{-1}(p)\) the functions \[\begin{align} \tilde{r}_{p, v}: \mathcal{B}(\alpha, p) &\to \mathbb{R} &\quad \tilde{r}^M_{z, v} : \mathcal{B}(\alpha, z) &\to \mathbb{R}\\ q &\mapsto \langle v, \log^B_p q\rangle &\quad x &\mapsto \langle \mathrm{lift}_z v, \log^M_z x\rangle, \end{align}\] are such that \(\tilde{r}_{p, v}(\pi(x)) = \tilde{r}^M_{z, v}(x)\) for every \(x \in \mathcal{B}(\alpha, z)\).

Proof. Given \(p \in B\) and \(z \in \pi^{-1}(p)\), let us fix some \(q \in \mathcal{B}(\alpha, p)\) and let us consider two points \(x, y \in \pi^{-1}(q) \cap \mathcal{B}(\alpha, z )\). First of all, recall that \(\pi \circ \exp^M_z = \exp^B_p \circ \,\pi_*|_z\) (cf. 23), and so \[\exp^B_p(\pi_*((\log^M_z x)^{\mathcal{H}})) = \exp^B_p(\pi_*(\log^M_z x)) = \pi(\exp_z^M(\log^M_z x)) = \pi(x) = q.\] Following the same reasoning, we also obtain that \[\exp^B_p(\pi_*((\log^M_z y)^{\mathcal{H}})) = q,\] so both \(\pi_*((\log^M_z x)^{\mathcal{H}})\) and \(\pi_*((\log^M_z y)^{\mathcal{H}})\) are tangent vectors at \(p\) which get mapped to \(q\) via the exponential map.

Moreover, from \(\pi\) being a Riemannian submersion, it holds that \[\begin{align} |\pi_*((\log^M_z x)^{\mathcal{H}})|_h &= |(\log^M_z x)^{\mathcal{H}}|_g \leq |\log^M_z x|_g \leq \alpha \leq i(B),\\ |\pi_*((\log^M_z y)^{\mathcal{H}})|_h &= |(\log^M_z y)^{\mathcal{H}}|_g \leq |\log^M_z y|_g \leq \alpha \leq i(B), \end{align}\] and so from \(\exp^B_p\) being a local diffeomorphism in the ball of radius \(i(B)\), it has to be the case that \[\pi_*((\log^M_z x)^{\mathcal{H}}) = \pi_*((\log^M_z y)^{\mathcal{H}}),\] which in turn allows us to conclude that \[(\log^M_z x)^{\mathcal{H}} = (\log^M_z y)^{\mathcal{H}},\] and \[\pi_*((\log^M_z x)^{\mathcal{H}}) = \log^B_p q.\]

Given some fixed \(p \in B\), and any \(z \in \pi^{-1}(p)\) it holds that \[\pi(\mathcal{B}(\alpha, z)) = \mathcal{B}(\alpha, x),\] as we saw in 65. Moreover, \[\begin{align} \tilde{r}_{p, v}(\pi(x)) = \langle v, \log^B_p \pi(x)\rangle = \langle v, \pi_*((\log^M_z x)^{\mathcal{H}})\rangle &= \langle \mathrm{lift}_z v, (\log^M_z x)^{\mathcal{H}}\rangle \\&= \langle \mathrm{lift}_z v, \log^M_z x\rangle \\&= \tilde{r}^M_{z, v}(x), \end{align}\] so the result follows. ◻

With the previous lemma in mind, we are ready to prove 32.

Proof of 32. Given some \(p \in B\) fixed, let us fix some \(z \in \pi^{-1}(p)\). As we are assuming that \(r \leq \min \{i(M), i(B)\}\), and using 130 we obtain a function \(\tilde{r}^M_{z, v}\) defined on \(\mathcal{B}(r, z)\) satisfying \(\tilde{r}^M_{z, v} = \tilde{r}_{p, v} \circ \pi\).

Moreover, by 74 79 we know that their Laplacians and their gradients are related by the following identities; \[\begin{align} \label{eq:igualdadesrtildelifted} \begin{aligned} \Delta_g \, \tilde{r}^M_{z, v} (x) &= \Delta_h \,\tilde{r}_{p, v} (\pi(x)),\\ |\mathrm{grad}_g\, \tilde{r}^M_{z, v} (x)|_g &= |\mathrm{grad}_h\, \tilde{r}_{p, v}(\pi(x))|_h, \end{aligned} \end{align}\tag{94}\] for every \(x \in \mathcal{B}(r, z)\).

As we are assuming that \[r \leq \frac{1}{72 \dim(M)^{5/2} \mathbf{K}},\] we can apply 128 129 to \(\tilde{r}^M_{z, v}\) to conclude that for every \(x \in \mathcal{B}(r, z)\) it holds that \[\begin{align} |\Delta_g\, \tilde{r}^M_{z, v}(x)| &\leq 24 \dim(M)^{5/2},\\ \frac{1}{1+L} \leq |\mathrm{grad}_g \, \tilde{r}^M_{z, v}(x)|^2_g &\leq \frac{1}{1-L}. \end{align}\]

Thus, inserting these bounds on 94 we can conclude that \[\begin{align} |\Delta_h \tilde{r}_{p, v}(q)| &\leq 24 \text{dim}(M)^{5/2},\\ \frac{1}{1+L} \leq |\mathrm{grad}_h\, \tilde{r}_{p, v}(q)|^2_h &\leq \frac{1}{1-L}. \end{align}\] for every \(q \in \mathcal{B}(r, p)\).

Since \(r \leq \frac{1}{2\text{dim}(M)^{3/2} \mathbf{K}}\), and \(\mathbf{K} \geq 1\), we can bound \(L\) as \[L \leq 2\sum_{n = 1}^\infty \frac{\mathbf{K}^{-n}}{(2n+2)!} \leq 2\sum_{n = 1}^\infty \frac{1}{(2n+2)!} \leq \frac{1}{10},\] and so \[\frac{10}{11} \leq \frac{1}{1+L} \leq |\mathrm{grad}_h \,\tilde{r}_{p, v}(q)|^2_h \leq \frac{1}{1-L} \leq \frac{10}{9},\] for every \(q \in \mathcal{B}(r, p)\).

Lastly, note that \(\tilde{r}_{p, v}(q) \leq d_g(p,q)\) for every \(q \in \mathcal{B}(r, p)\). Thus, since \(r \leq \frac{1}{72\dim(M)^{5/2}}\) by assumption, it holds that \[|\mathrm{grad}_h \,\tilde{r}_{p, v}(q)|^2_h + \tilde{r}_{p, v}(q) \Delta_h \tilde{r}_{p, v}(q) \geq \frac{10}{11} - \frac{1}{3} \geq \frac{1}{2},\] for every \(q \in \mathcal{B}(r, p)\). ◻

13.4.2 Second auxiliary Lemma↩︎

Lemma 33. Let \(\pi: (M, g) \to (B, h)\) be a surjective Riemannian submersion, where \((M, g)\) is a compact and connected symmetric space. Assume that there exists some constant \(\mathbf{K} \geq 1\) such that, for every \(x \in M\), the coefficients of the Riemann curvature tensor in normal coordinates centered at \(x\) satisfy \(|R_{ijk}^l(x)| \leq \mathbf{K}\) for every \(i, j, k, l \in \{1, \dotsc, \dim(M)\}\).

Let \(F\) be a smooth function on \(M\) verifying 12 13, and let \(\tilde{F}\) be the unique function on \(B\) satisfying \(F = \tilde{F} \circ \pi\). Assume that \(\tilde{F}\) satisfies . Let \[r \leq \min\Bigg\{i(M), i(B), \frac{1}{\sqrt{2}\dim(M)^{3/2}\mathbf{K}}\Bigg\}.\]

Let \(p \in B\) be a fixed saddle point of \(\tilde{F}\) and let \(v\) be the unitary eigenvector of \(\nabla^2 \tilde{F}(p)\) that corresponds to its minimum eigenvalue. Defining \(\tilde{r}_{p, v}(q)\) with respect to \(v\), i.e. \[\tilde{r}_{p, v}(q) = \langle v, \log^B_p q\rangle,\] it holds that \[\langle -H_p(q), \mathrm{P}^{-1}_{\log_p q}\, \mathrm{grad}_h \, \tilde{r}_{p, v}(q)\rangle_h \geq \lambda_* \tilde{r}_{p, v}(q) - 4\dim(M)^5 A_2 \mathbf{K} r^3,\] for every \(q \in B\big(y, r\big)\), where \[H_p(q) := \nabla^2_h\, \tilde{F}(p)[\log^B_p q],\] and \(A_2\) is the Lipschitz constant of the gradient of \(\tilde{F}\).

We proceed in the same manner as in the previous section; to prove 33 we will lift the expression \[\label{eq:expressionofinterest} \langle -H_p(q), \mathrm{P}^{-1}_{\log_p q}\, \mathrm{grad}_h \, \tilde{r}_{p, v}(q)\rangle_h\tag{95}\] to the total space \((M, g)\). We will obtain bounds for the lifted expression, using the fact that \((M, g)\) is assumed to be a symmetric space, and then transfer the bounds to the original expression.

In the previous subsection we saw that, given some fixed \(p \in B\), we can define a function \(\tilde{r}^M_{z, v}\) on \(\mathcal{B}(r, z)\), for any \(z \in \pi^{-1}(p)\), when \(r\) is sufficiently small, satisfying \(\tilde{r}^M_{z, v} = \tilde{r}_{p, v} \circ \pi\). In the same spirit, we can also lift the function \(H_p(q)\).

Lemma 34. In the same setting as for 33, let \(p \in B\) be a fixed saddle point of \(\tilde{F}\), and let \(z \in \pi^{-1}(p)\). Let \[H^M_z(x) := \nabla^2_g\, F(z)[\log^M_z x]\] be defined in \(\mathcal{B}(\alpha, z)\), where \(\alpha = \min\{i(M), i(B)\}\). Then \(H^M_z\) is invariant in the fibers of \(\pi\), i.e. for every \(q \in \mathcal{B}(\alpha, p)\) and every \(x_1, x_2 \in \pi^{-1}(q) \cap \mathcal{B}(\alpha, z)\), it holds that \[H^M_z(x_1) = H^M_z(x_2),\] and \[\pi_* H^M_z(x) = H_p(\pi(x)),\] for every \(x \in \mathcal{B}(\alpha, z)\).

Proof. Note that \(z \in \pi^{-1}(p)\) is a critical point of \(F\) by 74. Thus, we can apply 77 to conclude that \[H^M_z(x) = \nabla^2_g\, F(z)[\log^M_z x] = \nabla^2_g\;F(z)[(\log^M_z x)^\mathcal{H}], \quad \forall x \in \mathcal{B}(\alpha, z).\]

Now, for every \(q \in \mathcal{B}(\alpha, p)\) and every \(x_1, x_2 \in \pi^{-1}(q) \cap \mathcal{B}(\alpha, z)\), we proved in 130 that \[(\log^M_z x_1)^{\mathcal{H}} = (\log^M_z x_2)^{\mathcal{H}},\] and so \(H^M_z(x_1) = H^M_z(x_2)\).

Moreover, \[\begin{align} (H^M_z(x))^{\mathcal{H}} = (\nabla^2_g\, F(z)[\log^M_z x])^{\mathcal{H}} = (\nabla^2_g\, F(z)[(\log^M_z x)^\mathcal{H}])^{\mathcal{H}} &= \text{lift}_z \nabla^2_h\, \tilde{F}(p)[\pi_* \log^M_z x] \\&= \text{lift}_z \nabla^2_h\, \tilde{F}(p)[\log^B_p q] \\&= \text{lift}_z\, H_p(q), \end{align}\] where the third-last identity follows from 76 and the second-last identity follows from 130. Thus, it holds that \[\pi_* H^M_z(x) = H_p(\pi(x)),\] for every \(x \in \mathcal{B}(\alpha, z)\), finishing the proof. ◻

34 allows us to relate the expression shown in 95 to its lifted version.

Lemma 35. In the same setting as for 33, let \(p \in B\) be a fixed saddle point of \(\tilde{F}\), and let \(z \in \pi^{-1}(p)\) be fixed. Then for every \(x \in \mathcal{B}(\alpha, z)\) it holds that \[\langle H_p(q), \mathrm{P}^{-1}_{\log^B_p q} \, \mathrm{grad}_h\, \tilde{r}_{p, v}(q)\rangle_h = \langle H^M_z(x)^{\mathcal{H}}, (\mathrm{P}^{-1}_{\log^M_z x} \, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x))^{\mathcal{H}}\rangle_g,\] where \(q = \pi(x)\).

Proof. Given \(p \in B\) and \(z \in \pi^{-1}(p)\), let \(x \in \mathcal{B}(\alpha, z)\) and let us denote its projection as \(q := \pi(x)\). First, we can use 74 along with 130 34 to rewrite \[\langle H_p(q), \mathrm{P}^{-1}_{\log^B_p q} \, \mathrm{grad}_h\, \tilde{r}_{p, v}(q)\rangle_h = \langle \pi_*\, H^M_z(x), \mathrm{P}^{-1}_{\log^B_p q} \circ \pi_* (\, \mathrm{grad}_g \,\tilde{r}^M_{z, v}(x))\rangle_h.\] Now, we will make use of 23, which guarantees that \[\pi_*|_z \circ \mathrm{P}^{-1}_{\log^M_z x} = \mathrm{P}^{-1}_{\log^B_{p}{q}} \circ \pi_* |_x.\] Therefore, we can commute the parallel transport and the pushforward, to obtain \[\begin{align} \langle \pi_*\, H^M_z(x), \mathrm{P}^{-1}_{\log^B_p q} \circ \pi_* (\, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x))\rangle_h &= \langle \pi_*\, H^M_z(x), \pi_*\, \mathrm{P}^{-1}_{\log^M_z x} \, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x)\rangle_h \\&= \langle H^M_z(x)^{\mathcal{H}}, (\mathrm{P}^{-1}_{\log^M_z x} \, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x))^{\mathcal{H}}\rangle_g, \end{align}\] where the last equality follows from \(\pi\) being a Riemannian submersion. ◻

Now, in order to study \(\mathrm{P}^{-1}_{\log_z^M x}\, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x)\), we can use the following Taylor expansion for the parallel transport (cf. [87]), \[\begin{align} \mathrm{P}^{-1}_{\log^M_z x}\, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x) &= \sum_{n = 0}^\infty \frac{1}{n!} \left.\nabla^{(n)}_{\log^M_z x} \,\mathrm{grad}_g\, \tilde{r}^M_{z, v}(\gamma(t))\right|_{t = 0} \\&= \mathrm{grad}_g\, \tilde{r}^M_{z, v}(z) + \sum_{n = 1}^\infty \frac{1}{n!} \left.\frac{d^n}{dt^n}\right|_{t = 0} \mathrm{grad}_g\, \tilde{r}^M_{z, v}(\gamma(t)), \end{align}\] where \(\gamma(t) = \exp_z(t \log^M_z x)\). Using this expression for the parallel transport of the gradient of \(\tilde{r}^M_{z,v}\) and 35, we can obtain a lower bound on 95 .

Lemma 36. In the same setting as for 33, let \(p \in B\) be a fixed saddle point of \(\tilde{F}\), and let \(z \in \pi^{-1}(p)\) be fixed. Then for every \(x \in \mathcal{B}(\alpha, z)\) it holds that \[\langle H_p(q), \mathrm{P}^{-1}_{\log^B_p q} \, \mathrm{grad}_h\, \tilde{r}_{p, v}(q)\rangle_h \geq \langle H^M_z(x), \mathrm{grad}_g\, \tilde{r}^M_{z, v}(z))\rangle_g - 2|H^M_z(x)|_g |u|_g,\] where \(q = \pi(x)\) and \(u := \mathrm{P}^{-1}_{\log^M_z x} (\, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x)) - \mathrm{grad}_g\, \tilde{r}^M_{z, v}(z)\).

Proof. First, we can use 35 to rewrite \[\label{eq:firstidentityHgrad} \langle H_p(q), \mathrm{P}^{-1}_{\log^B_p q} \, \mathrm{grad}_h\, \tilde{r}_{p, v}(q)\rangle_h = \langle H^M_z(x)^{\mathcal{H}}, (\mathrm{P}^{-1}_{\log^M_z x} \, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x))^{\mathcal{H}}\rangle_g.\tag{96}\] Now, using the decomposition of the tangent space \(T_z M\) into the vertical and horizontal subspaces, we can rewrite the right-hand side of 96 as \[\begin{align} \langle H^M_z(x)^{\mathcal{H}}, (\mathrm{P}^{-1}_{\log^M_z x} \,\mathrm{grad}_g\, \tilde{r}^M_{z, v}(x))^{\mathcal{H}}\rangle_g &= \langle H^M_z(x), \mathrm{P}^{-1}_{\log^M_z x} \, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x)\rangle_g \\&\quad - \langle H^M_z(x)^{\mathcal{V}}, (\mathrm{P}^{-1}_{\log^M_z x} \, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x))^{\mathcal{V}}\rangle_g \\&\geq \langle H^M_z(x), \mathrm{P}^{-1}_{\log^M_z x} \, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x)\rangle_g \\&\quad - |H^M_z(x)^{\mathcal{V}}|_g |\mathrm{P}^{-1}_{\log^M_z x} \, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x))^{\mathcal{V}}|_g \\&\geq \langle H^M_z(x), \mathrm{P}^{-1}_{\log^M_z x} \, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x)\rangle_g - |H^M_z(x)|_g |u|_g, \end{align}\] where the last inequality follows since the first term in the Taylor series of the parallel transport corresponds to \(\mathrm{grad}_g\, \tilde{r}^M_{z, v}(z)\), which is horizontal by 74.

Lastly, we can again use the fact that \[\mathrm{P}^{-1}_{\log^M_z x}\, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x) = \mathrm{grad}_g\, \tilde{r}^M_{z, v}(z) + u,\] to conclude that \[\begin{align} \langle H^M_z(x)^{\mathcal{H}}, (\mathrm{P}^{-1}_{\log^M_z x} \, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x))^{\mathcal{H}}\rangle_g &\geq \langle H^M_z(x), \mathrm{P}^{-1}_{\log^M_z x} \, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x)\rangle_g - |H^M_z(x)|_g |u|_g \\&= \langle H^M_z(x), \mathrm{grad}_g\, \tilde{r}^M_{z, v}(z))\rangle_g + \langle H^M_z(x), u\rangle_g \\&\quad - |H^M_z(x)|_g |u|_g \\&\geq \langle H^M_z(x), \mathrm{grad}_g\, \tilde{r}^M_{z, v}(z))\rangle_g - 2|H^M_z(x)|_g |u|_g. \end{align}\] ◻

To obtain a more explicit bound, we need to upper bound the norm of \(u\). To do this, we will study the derivatives of \(\mathrm{grad}_g\, \tilde{r}^M_{z, v}\) along the geodesic \(\gamma(t) = \exp_z(t\log^M_z x)\).

From the definition of \(\tilde{r}^M_{z, v}\), and the expression for its gradient given in 31, we know that in normal coordinates centered at \(p\), \[\mathrm{grad}_g\, \tilde{r}^M_{z, v}(\gamma(t)) = g^{ij}(\gamma(t)) \delta_{ik} (\text{lift}_z v)^k \partial_j,\] and so \[\left.\frac{d^n}{dt^n}\right|_{t = 0} \mathrm{grad}_g\, \tilde{r}^M_{z, v}(\gamma(t)) = (g^{(n)}(0))^{ij} \delta_{ik} (\text{lift}_z v)^k \partial_j,\] where \((g^{(n)}(0))^{ij}\) is as defined in 123, i.e. \[(g^{(n)}(0))^{ij} = \left.\frac{d^n}{dt^n}\right|_{t = 0} g^{ij}(\gamma(t)).\] Therefore, the derivatives of the gradient of \(\tilde{r}^M_{z, v}\) along \(\gamma(t)\), can be studied by considering the derivatives of the coefficients of the inverse metric along the same geodesic. Recall from 123 that \[\begin{align} (g^{(2n+1)}(0))^{ij} &= 0,\\ \big(g^{(2n)}(0)\big)^{ij} &= C_{2n} \delta^{i\beta} T^j_{\beta, i_1, \dotsc, i_{2n}} x^{i_1}\cdots x^{i_{2n}}, \end{align}\] for every \(n \geq 0\) and every \(i, j \in \{1, \dotsc, d\}\), where \[\begin{align} T^\alpha_\beta &= \delta^\alpha_\beta,\\ T^\alpha_{\beta i_1i_2} &= R^\alpha_{\beta i_1 i_2},\\ T^\alpha_{\beta i_1, \dotsc, i_{2n}} &= R^\alpha_{\alpha_1 i_1 i_2} \dotsm R^{\alpha_{n-2}}_{\alpha_{n-1} i_{2n-3} i_{2n-2}} R^{\alpha_{n-1}}_{\beta i_{2n-1} i_{2n}},\quad n \geq 2, \end{align}\] (cf. 122) and \[\begin{align} C_0 &= 1,\\ C_{2n} &= -\sum_{m = 1}^n {\binom{2n}{2m}} C_{2(n-m)} \frac{2^{2m+1}}{(2m+2)(2m+3)}, \quad n \geq 1. \end{align}\]

In the following lemma, we obtain an upper bound for the absolute value of the coefficients \(C_{2n}\).

Lemma 37. Let \(C_{2n}\) be defined as above. Then \[|C_{2n}| \leq (2n)!\] for every \(n \geq 0\).

Proof. The result holds trivially for \(C_0\). Let now \(n \geq 1\) and assume that \(|C_{2k}| \leq (2k)!\) for every \(k < n\). Then \[\begin{align} |C_{2n}| &\leq \sum_{m = 1}^n {\binom{2n}{2m}} |C_{2(n-m)}| \frac{2^{2m+1}}{(2m+2)(2m+3)} \\&\leq \sum_{m = 1}^\infty {\binom{2n}{2m}} |C_{2(n-m)}| \frac{2^{2m+1}}{(2m+2)(2m+3)} \\&\leq \sum_{m = 1}^\infty {\binom{2n}{2m}} (2(n-m))! \frac{2^{2m+1}}{(2m+2)(2m+3)} \\&= 2n! \sum_{m = 1}^\infty \frac{(2m+1)2^{2m+1}}{(2m+3)!} \\&\leq 2n! \end{align}\] ◻

Using this bound on the coefficients \(C_{2n}\) and the bound on the tensor \(T\) derived in 88 , we can bound the norm of the derivatives of the gradient of \(\tilde{r}^M_{z, v}\) along \(\gamma(t)\).

Lemma 38. Let \((M, g)\) be a Riemannian manifold of dimension \(d\) which is also a complete and connected symmetric space. Assume that there exists some constant \(\mathbf{K} \geq 1\) such that, for every \(x \in M\), the coefficients of the Riemann curvature tensor in normal coordinates centered at \(x\) satisfy \(|R_{ijk}^l(x)| \leq \mathbf{K}\) for every \(i, j, k, l \in \{1, \dotsc, d\}\). Let \(z \in M\) be a fixed point, and let \(x \in M\) be outside the cut locus of \(z\). Let \(\gamma(t) = \exp_z(t \log^M_z x)\). Then it holds that \[\begin{align} \left.\frac{d^{2n+1}}{dt^{2n+1}}\right|_{t = 0} \mathrm{grad}_g\, \tilde{r}^M_{z, v}(\gamma(t)) &= 0,\\ \Bigg|\left.\frac{d^{2n}}{dt^{2n}}\right|_{t = 0} \mathrm{grad}_g\, \tilde{r}^M_{z, v}(\gamma(t))\Bigg|_g &\leq d^{3n+1} \mathbf{K}^n (2n)! r^{2n}, \end{align}\] for every \(n \geq 0\), where \(r = d_g(x, z)\).

Proof. Taking normal coordinates centered at \(z\), \[\left.\frac{d^n}{dt^n}\right|_{t = 0} \mathrm{grad}_g\, \tilde{r}^M_{z, v}(\gamma(t)) = (g^{(n)}(0))^{ij} \delta_{ik} (\text{lift}_z v)^k \partial_j.\] Using the expression for the derivatives of the coefficients of the inverse metric in normal coordinates found in 123 we know that \[\left.\frac{d^{2n+1}}{dt^{2n+1}}\right|_{t = 0} \mathrm{grad}_g\, \tilde{r}^M_{z, v}(\gamma(t)) = (g^{(2n+1)}(0))^{ij} \delta_{ik} (\text{lift}_z v)^k \partial_j = 0.\]

Also, from the bound on the coefficients \(C_{2n}\) shown in 37 and the bound on the tensor \(T\) obtained in 88 , we know that \[\begin{align} \label{eq:inequalityabsvalueinverseterms} \begin{aligned} \big|\big(g^{(2n)}(0)\big)^{ij}\big| = |C_{2n} \delta^{i\beta} T^j_{\beta, i_1, \dotsc, i_{2n}} x^{i_1}\cdots x^{i_{2n}}| &\leq (2n)! d^{2n} d^{n-1} \mathbf{K}^n r^{2n} \\&= (2n)! d^{3n-1} \mathbf{K}^n r^{2n}, \end{aligned} \end{align}\tag{97}\] for every \(n \geq 1\), and every \(i, j \in \{1, \dotsc, d\}\).

Therefore, since \(|\mathrm{lift}_z v|_g = 1\), we can use 97 to conclude that for every \(n \geq 1\) \[\begin{align} \Bigg|\left.\frac{d^{2n}}{dt^{2n}}\right|_{t = 0} \mathrm{grad}_g\, \tilde{r}^M_{z, v}(\gamma(t))\Bigg|_g &= \big|(g^{(2n)}(0))^{ij} \delta_{ik} (\mathrm{lift}_z\, v)^k \partial_j\big|_g \\&= \Big(\delta_{\alpha\beta} (g^{(2n)}(0))^{i\alpha} \delta_{ik} (\mathrm{lift}_z\, v)^k (g^{(2n)}(0))^{j\beta} \delta_{jl} (\mathrm{lift}_z\, v)^l \Big)^{1/2} \\&\leq \Big(d (2n)!^2 d^{6n} \mathbf{K}^{2n} r^{4n} \Big)^{1/2} \\&\leq d^{3n+1} \mathbf{K}^n (2n)! r^{2n}. \end{align}\]

Lastly, for \(n = 0\) note that \(g^{ij}(\gamma(0)) = \delta^{ij}\), and so \[|\mathrm{grad}_g\, \tilde{r}^M_{z, v}(\gamma(0))|_ g \leq 1 \leq d.\] ◻

With all of the above lemmas in mind, we can finally give a proof for 33.

Proof of 33. Given \(p \in B\) and \(z \in \pi^{-1}(p)\), recall from 36 that \[\label{eq:auxinequalityinnerproductH} \langle H_p(q), \mathrm{P}^{-1}_{\log^B_p q} \, \mathrm{grad}_h\, \tilde{r}_{p, v}(q)\rangle_h \geq \langle H^M_z(x), \mathrm{grad}_g\, \tilde{r}^M_{z, v}(z))\rangle_g - 2|H^M_z(x)|_g |u|_g,\tag{98}\] for every \(x \in \mathcal{B}(\alpha, z)\), where \(q = \pi(x)\).

Recall that \[\begin{align} u &= \mathrm{P}^{-1}_{\log^M_z x} \, \mathrm{grad}_g\, \tilde{r}^M_{z, v}(x) - \mathrm{grad}_g\, \tilde{r}^M_{z, v}(z) \\&= \sum_{n = 1}^\infty \frac{1}{n!} \left.\frac{d^n}{dt^n}\right|_{t = 0} \mathrm{grad}_g\, \tilde{r}^M_{z, v}(\gamma(t)) \\&= \sum_{n = 1}^\infty \frac{1}{2n!} \left.\frac{d^{2n}}{dt^{2n}}\right|_{t = 0} \mathrm{grad}_g\, \tilde{r}^M_{z, v}(\gamma(t)), \end{align}\] where \(\gamma(t) = \exp_z(t\log^M_z x)\), and where the last inequality follows from 38, as the odd derivatives vanish. Using the same lemma, we can bound the norm of \(u\), \[\begin{align} |u|_g &\leq \sum_{n = 1}^\infty \frac{1}{2n!} \Bigg|\left.\frac{d^{2n}}{dt^{2n}}\right|_{t = 0} \mathrm{grad}_g\, \tilde{r}^M_{z, v}(\gamma(t))\Bigg|_g \\&\leq \sum_{n = 1}^\infty \dim(M)^{3n+1} \mathbf{K}^n r^{2n}. \end{align}\] Lastly, recall that \(r \leq \frac{1}{\sqrt{2} \dim(M)^{3/2} \mathbf{K}^{1/2}}\) by assumption and so it holds that \[|u|_g \leq \dim(M)\sum_{n = 1}^\infty (\dim(M)^3 \mathbf{K} r^2)^n = \frac{\dim(M)^4 \mathbf{K} r^2}{1 - \dim(M)^3\mathbf{K} r^2} \leq 2 \dim(M)^4 \mathbf{K} r^2,\] since \(1 - \dim(M)^3\mathbf{K}r^2 \geq \frac{1}{2}\).

Thus, we can lower bound the right-hand side of 98 to conclude that \[\label{eq:anotherbound} \langle H_p(q), \mathrm{P}^{-1}_{\log^B_p q} \, \mathrm{grad}_h\, \tilde{r}_{p, v}(q)\rangle_h \geq \langle H^M_z(x), \mathrm{grad}_g\, \tilde{r}^M_{z, v}(z))\rangle_g - 4|H^M_z(x)|_g \dim(M)^4 \mathbf{K} r^2.\tag{99}\]

Let us now lower bound the first term on the right-hand side of 99 . Considering normal coordinates centered at \(z\), recall from the proof of 38 that \[\mathrm{grad}_g\, \tilde{r}^M_{z, v}(x) = g^{ij}(x) \delta_{ik} (\mathrm{lift}_z v)^k \partial_j,\] and so in particular \[\mathrm{grad}_g\, \tilde{r}^M_{z, v}(z) = (\mathrm{lift}_z v)^i \partial_i.\]

Moreover, recall that \[H^M_z(x) = \nabla^2_g\, F(z)[\log^M_z x] = \left.\nabla_{\log^M_z x} \,\mathrm{grad}_g \, F\right|_z,\] and so writing \(H^M_z(x)\) in normal coordinates, since the Christoffel symbols and the derivatives of the coefficients of the inverse metric vanish at \(z = 0\), we obtain \[\begin{align} H^M_z(x) =\left.\nabla_{\log^M_z x} \,\mathrm{grad}_g \, F\right|_z &= x^l \partial_l (\mathrm{grad}_g\, F)^i \partial_i \\&= x^l g^{ik}(0) \partial_l\partial_k F(0) \partial_i \\&= x^l \delta^{ik} \partial_{lk} F(0) \partial_i. \end{align}\]

Using the expressions for the gradient of \(\tilde{r}^M_{z, v}\) and \(H^M_z\) in normal coordinates, we can rewrite \[\begin{align} \label{eq:expressionatz} \begin{aligned} \langle -H^M_z(x), \mathrm{grad}_g\,\tilde{r}^M_{z, v}(z)\rangle_g &= -(H^M_z(x))^i\left(\mathrm{grad}_g\, \tilde{r}^M_{z, v}(z)\right)^j \delta_{ij} \\&= -\delta^{ik}\partial_{kl} F(0) x^l (\mathrm{lift}_z\, v)^j \delta_{ij} \\&= -\delta^{k}_j \partial_{kl} F(0) x^l (\mathrm{lift}_z\, v)^j \\&= - \partial_{kl} F(0) x^l (\mathrm{lift}_z\, v)^k. \end{aligned} \end{align}\tag{100}\]

Now, since \(v\) is the eigenvector associated with the minimum eigenvalue of \(\nabla^2 \tilde{F}(p)\), it holds that \(\mathrm{lift}_z\, v\) is an eigenvector associated with the minimum eigenvalue of \(\nabla^2 F(z)\). Denoting the eigenvalue associated with \(v\) as \(\lambda_v\), it holds that \[\nabla^2_h\, \tilde{F}(p)[v] = \lambda_v v,\] so we can apply 76 to conclude that \[(\nabla^2_g\, F(z)[\mathrm{lift}_z\, v])^{\mathcal{H}} = \mathrm{lift}_z \nabla^2_h \,\tilde{F}(p)[v] = \lambda_v\, \mathrm{lift}_z\, v.\] Furthermore, \[(\nabla^2_g\, F(z)[\mathrm{lift}_z\, v])^{\mathcal{H}} = \nabla^2_g\, F(z)[\mathrm{lift}_z\, v].\] Using the symmetry of the Hessian and 77, we know that \[(\nabla^2_g \,F(z)[\mathrm{lift}_z\, v])^{\mathcal{V}} = \sum_{w \in V_z M} \langle \nabla^2_g \,F(z)[\mathrm{lift}_z\, v], w\rangle_g\, w = \sum_{w \in V_z M} \langle \nabla^2_g \,F(z)[w], \mathrm{lift}_z v\rangle_g\, w =0,\] where \(V_z M\) denotes the vertical tangent space of \(M\) at \(z\). Thus, \[\nabla^2_g\, F(z)[\mathrm{lift}_z\, v] = \lambda_v \,\mathrm{lift}_z\, v,\] and so, since \(\lambda_v \leq -\lambda_*\) by 16, we know that \[\partial_{kl} F(0) (\mathrm{lift}_z v)^k \leq -\lambda_{*} \delta_{il} (\mathrm{lift}_z v)^i,\] so, we can lower bound 100 as \[\begin{align} \langle -H^M_z(x), \mathrm{grad}_g\,\tilde{r}^M_{z, v}(z)\rangle_g = -\partial_{kl} F(0) x^l (\mathrm{lift}_z v)^k \geq \lambda_{*} \delta_{il} (\mathrm{lift}_z v)^i x^l &= \lambda_* \tilde{r}^M_{z, v}(x) \\&= \lambda_* \tilde{r}_{p, v}(q), \end{align}\] where the last equality follows from 130. Inserting this inequality into 99 we conclude that \[\langle H_p(q), \mathrm{P}^{-1}_{\log^B_p q} \, \mathrm{grad}_h\, \tilde{r}_{p, v}(q)\rangle_h \geq \lambda_* \tilde{r}_{p, v}(q) - 4|H^M_z(x)|_g \dim(M)^4 \mathbf{K} r^2.\] Lastly, note that \[|H^M_z(x)|_g = |\nabla^2 F(z)[\log_z x]|_g \leq \dim(M) A_2 r,\] where \(A_2\) is the Lipschitz constant of the gradient of \(\tilde{F}\)—which coincides with that of the gradient of \(F\).

This allows us to conclude that \[\langle H_p(q), \mathrm{P}^{-1}_{\log^B_p q} \, \mathrm{grad}_h\, \tilde{r}_{p, v}(q)\rangle_h \geq \lambda_* \tilde{r}_{p, v}(q) - 4 \dim(M)^5 \mathbf{K} A_2 r^3,\] as desired. ◻

References↩︎

[1]
M. Creutz, Quarks, gluons and lattices. Cambridge University Press, 2023.
[2]
N. Boumal, “An introduction to optimization on smooth manifolds.” To appear with Cambridge University Press, 2022, [Online]. Available: https://www.nicolasboumal.net/book.
[3]
P.-A. Absil, R. Mahony, and R. Sepulchre, Optimization algorithms on matrix manifolds. Princeton University Press, 2008.
[4]
R. Chourasia, J. Ye, and R. Shokri, “Differential privacy dynamics of langevin diffusion and noisy gradient descent,” Advances in Neural Information Processing Systems, vol. 34, pp. 14771–14781, 2021.
[5]
G. Heller and E. Fetaya, “Can stochastic gradient langevin dynamics provide differential privacy for deep learning?” in 2023 IEEE conference on secure and trustworthy machine learning (SaTML), 2023, pp. 68–106.
[6]
T. Ryffel, F. Bach, and D. Pointcheval, “Differential privacy guarantees for stochastic gradient langevin dynamics.” 2022, [Online]. Available: https://arxiv.org/abs/2201.11980.
[7]
S. Brooks, A. Gelman, G. Jones, and X.-L. Meng, Handbook of markov chain monte carlo. CRC press, 2011.
[8]
A. Gelman, J. B. Carlin, H. S. Stern, and D. B. Rubin, Bayesian data analysis. Chapman; Hall/CRC, 1995.
[9]
D. J. MacKay, Information theory, inference and learning algorithms. Cambridge university press, 2003.
[10]
C. P. Robert, G. Casella, and G. Casella, Monte carlo statistical methods, vol. 2. Springer, 2004.
[11]
X. Cheng, J. Zhang, and S. Sra, “Efficient sampling on riemannian manifolds via langevin MCMC,” Advances in Neural Information Processing Systems, vol. 35, pp. 5995–6006, 2022.
[12]
D. Bakry, I. Gentil, and M. Ledoux, Analysis and geometry of markov diffusion operators. Springer International Publishing, 2013.
[13]
G. Menz and A. Schlichting, Poincaré and logarithmic Sobolev inequalities by decomposition of the energy landscape,” The Annals of Probability, vol. 42, no. 5, pp. 1809–1884, 2014, doi: 10.1214/14-AOP908.
[14]
S. Chewi, M. A. Erdogdu, M. Li, R. Shen, and M. S. Zhang, “Analysis of langevin monte carlo from poincare to log-sobolev,” Foundations of Computational Mathematics, vol. 25, no. 4, pp. 1345–1395, 2025.
[15]
S. Vempala and A. Wibisono, “Rapid convergence of the unadjusted langevin algorithm: Isoperimetry suffices,” Advances in neural information processing systems, vol. 32, 2019.
[16]
A. Wibisono, “Proximal langevin algorithm: Rapid convergence under isoperimetry.” 2019, [Online]. Available: https://arxiv.org/abs/1911.01469.
[17]
M. Li and M. A. Erdogdu, “Riemannian langevin algorithm for solving semidefinite programs,” Bernoulli, vol. 29, no. 4, pp. 3093–3113, 2023.
[18]
M. Li and M. A. Erdogdu, Supplement to "Riemannian Langevin algorithm for solving semidefinite programs".” 2023, doi: 10.3150/22-BEJ1576SUPP.
[19]
S. Cao, R. Nissim, and S. Sheffield, “Dynamical approach to area law for lattice yang-mills.” 2025, [Online]. Available: https://arxiv.org/abs/2509.04688.
[20]
J. I. Cirac, D. Pérez-Garcı́a, N. Schuch, and F. Verstraete, “Matrix product states and projected entangled pair states: Concepts, symmetries, theorems,” Rev. Mod. Phys., vol. 93, p. 045003, 2021.
[21]
I. Montvay and G. Münster, Quantum fields on a lattice. Cambridge University Press, 1994.
[22]
P. Páez Velasco, unpublished“Tensor network manifolds and riemannian fundamental theorem for tensor networks.”
[23]
E. P. Hsu, Stochastic analysis on manifolds. American Mathematical Soc., 2002.
[24]
Y. Xu et al., Submitted to The Fourteenth International Conference on Learning Representations. Under review“Quotient-space diffusion model,” 2026.
[25]
G. Menon, “The geometry of the deep linear network.” 2024, [Online]. Available: https://arxiv.org/abs/2411.09004.
[26]
G. Menon, A. J. Stromme, and A. Vacher, “On the implicit regularization of langevin dynamics with projected noise.” 2026, [Online]. Available: https://arxiv.org/abs/2602.12257.
[27]
C.-P. Huang, D. Inauen, and G. Menon, Motion by mean curvature and Dyson Brownian Motion,” Electronic Communications in Probability, vol. 28, pp. 1–10, 2023, doi: 10.1214/23-ECP540.
[28]
I. Jolliffe, Principal component analysis. Springer, 2011, pp. 1094–1096.
[29]
S. Yan, D. Xu, B. Zhang, and H.-J. Zhang, “Graph embedding: A general framework for dimensionality reduction,” in 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05), 2005, vol. 2, pp. 830–837.
[30]
H. Shen, K. Diepold, and K. Hüper, “A geometric revisit to the trace quotient problem,” in Proceedings of the 19th international symposium of mathematical theory of networks and systems (MTNS 2010), 2010, p. 1.
[31]
I. M. Safran, G. Yehudai, and O. Shamir, “The effects of mild over-parameterization on the optimization landscape of shallow relu neural networks,” in Conference on learning theory, 2021, pp. 3889–3934.
[32]
K. Fukumizu, S. Yamaguchi, Y. Mototake, and M. Tanaka, “Semi-flat minima and saddle points by embedding neural networks to overparameterization,” Advances in neural information processing systems, vol. 32, 2019.
[33]
J. Cheeger, D. G. Ebin, and D. G. Ebin, Comparison theorems in riemannian geometry, vol. 9. North-Holland Amsterdam, 1975.
[34]
M. W. Hirsch, Differential topology. Springer Science & Business Media, 2012.
[35]
F. Wang, Functional inequalities markov semigroups and spectral theory. Elsevier, 2006.
[36]
D. Revuz and M. Yor, Continuous martingales and brownian motion, vol. 293. Springer Science & Business Media, 2013.
[37]
D. Bakry, F. Barthe, P. Cattiaux, and A. Guillin, A simple proof of the Poincaré inequality for a large class of probability measures,” Electronic Communications in Probability, vol. 13, pp. 60–66, 2008, doi: 10.1214/ECP.v13-1352.
[38]
A. Bovier and F. den Hollander, Metastability: A potential-theoretic approach. Springer International Publishing, 2015.
[39]
J. M. Lee, Introduction to riemannian manifolds, vol. 2. Springer, 2018.
[40]
M. J. Wainwright, High-dimensional statistics: A non-asymptotic viewpoint. Cambridge University Press, 2019.
[41]
[42]
Z. Qian, “A gradient estimate on a manifold with convex boundary,” Proceedings of the Royal Society of Edinburgh: Section A Mathematics, vol. 127, no. 1, pp. 171–179, 1997, doi: 10.1017/S0308210500023568.
[43]
R. Sulanke and P. Wintgen, Differentialgeometrie und faserbündel, vol. 48. Springer, 1972.
[44]
J. Q. Zhong, “On the estimate of first eigenvalue of a compact riemannian manifold,” Sci. Sinica Ser. A, vol. 27, pp. 1265–1273, 1984.
[45]
P. Cattiaux, I. Gentil, and A. Guillin, “Weak logarithmic sobolev inequalities and entropic convergence,” Probability theory and related fields, vol. 139, no. 3, pp. 563–603, 2007.
[46]
E. Carlen and M. Loss, “Logarithmic sobolev inequalities and spectral gaps,” Contemporary Mathematics, vol. 353, pp. 53–60, 2004.
[47]
A. B. Tsybakov, Introduction to nonparametric estimation. Springer New York, NY, 2008.
[48]
C. Villani, Optimal transport: Old and new. Springer Berlin Heidelberg, 2008.
[49]
R. H. Escobales Jr, “Riemannian submersions with totally geodesic fibers,” Journal of Differential Geometry, vol. 10, no. 2, pp. 253–276, 1975.
[50]
C. B. Croke, “Some isoperimetric inequalities and eigenvalue estimates,” Annales scientifiques de l’École Normale Supérieure, vol. Ser. 4, 13, no. 4, pp. 419–435, 1980, doi: 10.24033/asens.1390.
[51]
A. Gray and L. Vanhecke, Riemannian geometry as determined by the volumes of small geodesic balls,” Acta Mathematica, vol. 142, pp. 157–198, 1979, doi: 10.1007/BF02395060.
[52]
N. Batir, “Inequalities for the gamma function,” Archiv der Mathematik, vol. 91, no. 6, pp. 554–563, 2008.
[53]
Á. Yángüez, T. A. Hahn, and J. Kochanowski, “Efficient quantum measurements: Computational max- and measured rényi divergences and applications.” 2025, [Online]. Available: https://arxiv.org/abs/2509.21308.
[54]
F. Verstraete, D. Porras, and J. I. Cirac, “Density matrix renormalization group and periodic boundary conditions: A quantum information perspective,” Phys. Rev. Lett., vol. 93, p. 227205, 2004, doi: 10.1103/PhysRevLett.93.227205.
[55]
R. A. Horn and C. R. Johnson, Matrix analysis. Cambridge university press, 2012.
[56]
T. Bendokat, R. Zimmermann, and P.-A. Absil, “A grassmann manifold handbook: Basic geometry and computational aspects,” Advances in Computational Mathematics, vol. 50, no. 1, p. 6, 2024.
[57]
M. Jerrum and A. Sinclair, “Polynomial-time approximation algorithms for the ising model,” SIAM Journal on computing, vol. 22, no. 5, pp. 1087–1116, 1993.
[58]
P. Petersen, Riemannian geometry. Springer New York, 2006.
[59]
W. Fulton and J. Harris, Representation theory: A first course. Springer Science & Business Media, 2013.
[60]
J. Q. Gallier and J. Quaintance, Differential geometry and lie groups, vol. 12. Springer, 2020.
[61]
A. Böttcher and D. Wenzel, “The frobenius norm and the commutator,” Linear Algebra and its Applications, vol. 429, no. 8–9, pp. 1864–1885, Oct. 2008, doi: 10.1016/j.laa.2008.05.020.
[62]
M. P. Do Carmo and J. Flaherty Francis, Riemannian geometry, vol. 6. Springer, 1992.
[63]
A. L. Besse, Einstein manifolds. Springer Berlin Heidelberg, 2007.
[64]
W. Klingenberg, “Contributions to riemannian geometry in the large,” Annals of Mathematics, vol. 69, no. 3, pp. 654–666, 1959, Accessed: Jun. 25, 2024. [Online]. Available: http://www.jstor.org/stable/1970029.
[65]
M. Berger, A panoramic view of riemannian geometry. Springer Berlin Heidelberg, 2007.
[66]
B. O’Neill, “The fundamental equations of a submersion.” Michigan Mathematical Journal, vol. 13, no. 4, pp. 459–469, 1966.
[67]
B. O’Neill, “Submersions and geodesics,” Duke Math. J., vol. 34, no. 1, pp. 363–373, 1967.
[68]
R. L. Cohen, “The topology of fiber bundles lecture notes,” Standford University, 1998.
[69]
C. Autenried and I. Markina, “Sub-riemannian geometry of stiefel manifolds,” SIAM Journal on Control and Optimization, vol. 52, no. 2, pp. 939–959, 2014.
[70]
Q. Rentmeesters et al., “Algorithms for data fitting on some common homogeneous spaces,” PhD thesis, Ph. D. thesis, Université Catholique de Louvain, Louvain, Belgium, 2013.
[71]
P.-A. Absil and S. Mataigne, “The ultimate upper bound on the injectivity radius of the stiefel manifold,” SIAM Journal on Matrix Analysis and Applications, vol. 46, no. 2, pp. 1145–1167, 2025.
[72]
S. K. Lando and A. K. Zvonkin, Graphs on surfaces and their applications, vol. 141. Springer Science & Business Media, 2013.
[73]
A. Grigor’yan, “Lecture notes on analysis on manifolds.” Bielefeld University, 2024, [Online]. Available: https://www.math.uni-bielefeld.de/~grigor/anman2.pdf.
[74]
F.-Y. Wang, “Log-sobolev inequality on non-convex riemannian manifolds,” Advances in Mathematics, vol. 222, no. 5, pp. 1503–1520, 2009.
[75]
E. P. Hsu, “A brief introduction to brownian motion on a riemannian manifold,” lecture notes, 2008.
[76]
J. C. Cox, J. E. Ingersoll, S. A. Ross, et al., “A theory of the term structure of interest rates,” Econometrica, vol. 53, no. 2, pp. 385–407, 1985.
[77]
mathusername, “Escaping time of a modified CIR process.” Mathematics Stack Exchange, 2025, [Online]. Available: https://math.stackexchange.com/q/5117199.
[78]
B. Akhtari and H. Li, “The cox-ingersoll-ross process under volatility uncertainty,” Journal of Mathematical Analysis and Applications, vol. 531, no. 1, p. 127867, 2024.
[79]
J. M. Steele, Stochastic calculus and financial applications. Springer New York, 2012.
[80]
B. Oksendal, Stochastic differential equations: An introduction with applications. Springer Science & Business Media, 2013.
[81]
E. Pardoux and A. Răşcanu, Stochastic differential equations, backward SDEs, partial differential equations. Springer International Publishing, 2014.
[82]
L. C. Evans, Partial differential equations. American Mathematical Society, 2010.
[83]
D. Gilbarg and N. S. Trudinger, Elliptic partial differential equations of second order. Springer, 1977.
[84]
J. W. Milnor and D. W. Weaver, Topology from the differentiable viewpoint, vol. 21. Princeton university press, 1997.
[85]
B. Andrews and C. Hopper, The ricci flow in riemannian geometry: A complete proof of the differentiable 1/4-pinching sphere theorem. Springer, 2010.
[86]
G. W. Stewart, Matrix algorithms: Volume 1: Basic decompositions. SIAM, 1998.
[87]
B. F. Schutz, Geometrical methods of mathematical physics. Cambridge university press, 1980.

  1. pablopaez@ucm.es↩︎

  2. Certainly, discrete-time stochastic evolutions are more relevant to computer simulations, but to be able to mobilize the conceptual tools that will lead to a proof of our main results, we have chosen to work in continuous time. Deriving a discrete-time algorithm from a given stochastic differential equation in a way that preserves convergence properties is a non-trivial problem, which has been solved in [11] for a setting similar to ours.↩︎

  3. Furthermore, in our approach, we avoid some non-verified assumptions used in some of the proofs of [18].↩︎

  4. Although the expressions of the SDEs are non-rigorous, they provide a clear insight to the dynamics of the processes \(X_t\) and \(\tilde{X}_t\). Furthermore, they show the connection to the usual gradient descent. Nevertheless, these SDEs will not be used in this work. We will only use the definition provided in ?? . See [23].↩︎

  5. In fact, \(\operatorname{L}\) can be extended to act on a broader subset of functions of \(L^2(M)\) via its closure [35].↩︎

  6. Our definition of Lyapunov functions is not standard, as they are often defined with slight modifications—see, for example [37].↩︎

  7. In order to avoid cumbersome expressions, we will always denote by \(\tau^*\) the first escaping time of \(\tilde{X}_t\) from \(U\Big(\frac{a}{\sqrt{\beta}}, \tilde{\mathcal{S}}\Big)\).↩︎

  8. Despite \(\Delta_h \tilde{r}\) being zero in the case of flat space and its gradient having constant norm one, bounding the Laplacian or the norm of the gradient of \(\tilde{r}\) becomes non-trivial even in very simple settings such as the two-dimensional round sphere (cf. 13.1). Nevertheless, when restricting \(x\) to a sufficiently small neighborhood around \(y\), we are able to obtain bounds for such quantities.↩︎

  9. For a short introduction to Riemannian submersions, see 8.↩︎

  10. Although in the rest of the sections of this work \(M\) is always assumed to be compact, we have decided to write this section for any Riemannian manifold \((M, g)\), not necessarily compact, as we believe that the result shown here is relevant on its own.↩︎

  11. For a short introduction to Sobolev spaces on manifolds, we refer the reader to 9.↩︎

  12. Recall that \(\mu \ll \nu\) is used to denote that \(\mu\) is absolutely continuous* with respect to \(\nu\), i.e. for every measurable set \(A\), \(\nu(A) = 0\) implies that \(\mu(A) = 0\).*↩︎

  13. Although their work focuses on real Grassmann manifolds, their results automatically apply to the complex case after replacing symmetric matrices by Hermitian matrices, and the transpose with the conjugate transpose.↩︎

  14. We require \(M\) to be compact for simplicity, in order to be able to define \(\exp_x\) on \(T_xM\), as in the general case the exponential map is only defined on a subset of \(T_x M\), unless \(M\) is complete.↩︎

  15. In fact, these bounds can be tightened, obtaining the upper bound of \(\pi\sqrt{2n}\), by considering the Riemannian submersions \[\begin{align} \mathrm{SU}(n) &\to \mathrm{SU}(n) /\mathrm{SU}(n-k) \simeq \mathrm{V}_k(\mathbb{C}^n),\\ \mathrm{SU}(n) &\to \mathrm{SU}(n) /\mathrm{S}(\mathrm{U}(n-k) \times \mathrm{U}(k)) \simeq \mathrm{Gr}_k(\mathbb{C}^n), \end{align}\] which are analogous to the ones studied in this work. Nevertheless, we will keep the above non-optimal upper bounds for simplicity.↩︎

  16. Typically, the curvature-dimension condition as stated here is denoted as \(\mathrm{CD}(\kappa, \infty)\). Nevertheless, we will use \(\mathrm{CD}(\kappa)\) for readability.↩︎

  17. Two measures \(\mu\), \(\nu\) are said to be equivalent if and only if \(\mu \ll \nu\) and \(\nu \ll \mu\).↩︎