May 24, 2026
We study hypercontractivity for the underdamped Langevin dynamics with a convex confining potential. Unlike in the overdamped case, the noise acts only on the velocity variable, so the usual argument based on the logarithmic Sobolev inequality (LSI) does not apply. Nevertheless, we prove that, when the spatial marginal satisfies an LSI with constant \(\rho\) and the friction parameter is of order \(\sqrt{\rho}\), the underdamped Langevin semigroup improves integrability over the kinetic time scale \(t \sim \rho^{-1/2}\).
The proof rests on a novel space-time logarithmic Sobolev inequality, derived from a controlled version of the entropy decay estimate, which captures how dissipation in the velocity variable is transferred to the position variable. Combining this space-time LSI with a duality argument based on a forward/backward interpolation of the underdamped Langevin semigroup yields the desired hypocoercive hypercontractivity estimate. As a corollary, we obtain decay of the Rényi divergence at the sharp hypocoercive rate \(\mathcal{O}(\sqrt{\rho})\).
The notion of hypercontractivity originates in the seminal work of Nelson Nelson1973? on \(L^p\)-to-\(L^q\) estimates for the Ornstein–Uhlenbeck semigroup. Gross Gross1975? identified the logarithmic Sobolev inequality (LSI) as the underlying mechanism, establishing that for reversible Markov semigroups, an LSI is in fact equivalent to hypercontractivity. Concretely, if \((\mathcal{P}_t)_{t\ge 0}\) is a Markov semigroup with invariant law \(\mu\), then \(\mu\) satisfies an LSI with constant \(\rho\) if and only if \[\label{eq:Gross} \lVert\mathcal{P}_t f\rVert_{L^{q(t)}(\mu)} \le \lVert f\rVert_{L^p(\mu)}, \qquad q(t)-1 = \mathrm{e}^{2\rho t}(p-1), \qquad \forall\, p>1.\tag{1}\] That is, the semigroup improves integrability over time, with the admissible exponent growing exponentially at a rate set by the LSI constant. Bakry and Émery BakryEmery1985? subsequently developed the \(\Gamma_2\) calculus approach to hypercontractivity for diffusion semigroups, based on a symmetric dual formulation Neveu1976?; we refer to BakryGentilLedoux2014? for a systematic exposition. Beyond their intrinsic interest, LSI and the associated hypercontractive estimates are now standard tools with deep connections to many areas, including concentration of measure, transportation cost inequalities, isoperimetric inequalities, and mixing-time bounds for Markov chains guionnet2004lectures?, ledoux2006concentration?, BCR2006?.
In this work, we extend such hypercontractivity to the underdamped Langevin dynamics \[\label{eq:underdamped-Langevin} \begin{align} \,\mathrm{d}X_t&=V_t\,\mathrm{d}t,\\ \,\mathrm{d}V_t&=-\nabla U(X_t)\,\mathrm{d}t-\gamma V_t\,\mathrm{d}t+\sqrt{2\gamma}\,\mathrm{d}B_t, \end{align}\tag{2}\] where \((X_t,V_t)\in\mathbb{R}^d\times\mathbb{R}^d\) are the position and velocity, \(U:\mathbb{R}^d\to\mathbb{R}\) is a convex confining potential, \(\gamma>0\) is the friction parameter, and \((B_t)_{t\ge0}\) is a standard Brownian motion in \(\mathbb{R}^d\). The invariant law is the Gibbs measure \(\mu(\,\mathrm{d}x\,\mathrm{d}v)\propto\mathrm{e}^{-U(x)-\lvert v\rvert^2/2}\,\mathrm{d}x\,\mathrm{d}v\). The underdamped Langevin dynamics is more delicate to analyze than the overdamped case since the diffusion is degenerate and acts only on the velocity variable. The regularization must therefore be transferred through the Hamiltonian transport to the spatial variable. The understanding of this phenomenon dates back to Kolmogorov Kolmogorov1934? and Hörmander Hormander1967? in the framework of hypoellipticity; see, e.g., DesvillettesVillani2001?, HerauNier2004?, HelfferNier2005?, Herau2007?, GolseImbertMouhotVasseur2019? for more recent developments. The hypocoercivity framework was developed by Villani Villani2009? to quantify the indirect dissipation and to obtain more explicit decay estimates. The quantitative hypocoercivity analysis has been further developed in recent years; see, e.g., DolbeaultMouhotSchmeiser2015?, Baudoin2017?, BernardFathiLevittStoltz2022?, BrigatiStoltz2025?, CaoLuWang2023?, AlbrittonArmstrongMourratNovack2024?, FanLiLu2026?, Lu2026?. In particular, recent works have established sharp hypocoercive decay estimates for the underdamped Langevin equation CaoLuWang2023?, Lu2026?, FanLiLu2026?.
Hypercontractivity for underdamped Langevin dynamics was previously investigated by Feng-Yu Wang Wang2017?, whose analysis is based on a coupling and dimension-free Harnack inequality argument (see also Wang1997?). This approach has also been refined and applied to sampling algorithms by Altschuler, Chewi, and Zhang AltschulerChewiZhang2025?, ZhangAltschulerChewi2026?. While the approach of Wang2017? works for underdamped Langevin dynamics, it does not yield the kinetic hypercontractivity rate: the analysis is based on a contractive synchronous coupling and thus only applies in the high-friction regime.
Our main result gives a kinetic analogue of the Gross estimate 1 for the underdamped Langevin semigroup \(\mathcal{P}_t=\mathrm{e}^{t(-\mathcal{L}_a+\gamma\mathcal{L}_s)}\) with friction parameter \(\gamma=\Gamma\sqrt\rho\) in the low-friction regime. For each fixed \(\tau\ge\tau_{\rm ST}(\Gamma)\), there exists \(\alpha_{\Gamma,\tau}>1\), independent of \(d\) and \(\rho\), such that \[\lVert\mathcal{P}_t f\rVert_{L^{q_{\Gamma,\tau}(t)}(\mu)} \le \lVert f\rVert_{L^p(\mu)}, \qquad q_{\Gamma,\tau}(t)-1 = \alpha_{\Gamma,\tau}^{\lfloor \sqrt\rho\,t/\tau\rfloor}(p-1), \qquad \forall\, p > 1.\] Thus the exponent improves on the ballistic time scale \(t\sim\rho^{-1/2}\), consistent with the sharp hypocoercivity decay estimates of Lu2026?.
Before describing the main idea behind the proof, let us briefly recall the hypocoercive entropy decay established by one of us in Lu2026?: for the underdamped Langevin semigroup \(\mathcal{P}_t=\mathrm{e}^{t(-\mathcal{L}_a+\gamma\mathcal{L}_s)}\) with friction \(\gamma=\Gamma\sqrt\rho\) in the low-friction regime, \[\mathop{\mathrm{Ent}}_\mu(\mathcal{P}_tf) \le C_\Gamma\,\mathrm{e}^{-\lambda_\Gamma\sqrt\rho\,t}\mathop{\mathrm{Ent}}_\mu(f), \qquad t\ge 0,\] for every probability density \(f\), with constants \(C_\Gamma,\lambda_\Gamma>0\) depending only on \(\Gamma\). The decay rate \(\sqrt\rho\) occurs on the ballistic (kinetic) time scale and is sharp. In the overdamped case, such exponential entropy decay follows directly from LSI through the gradient-flow viewpoint of Otto calculus JordanKinderlehrerOtto1998?, Otto2001?, OttoVillani2000?. However, no LSI of the same type holds for underdamped Langevin dynamics due to the degeneracy of the diffusion. The analysis of Lu2026? instead uses an entropic hypocoercive Lyapunov function with a Wasserstein entropy-current corrector \[\mathcal{H}_\theta(g) = \mathop{\mathrm{Ent}}_\mu(g)+\theta\sqrt\rho\,\mathcal{C}_{\rm OT}(g),\] where \(\mathcal{C}_{\rm OT}(g)\) couples the spatial marginal and the velocity current through optimal transport. This modified entropy satisfies a differential inequality strong enough to control the full entropy.
In the reversible diffusion case, LSI has two important consequences: entropy decay and hypercontractivity. While Lu2026? establishes the entropy decay, it does so without proving an LSI, so the classical Gross argument for hypercontractivity Gross1975? cannot be applied directly.
The key insight in this work is to establish a version of the logarithmic Sobolev inequality that holds for underdamped Langevin dynamics. The idea originates from another thread of quantitative hypocoercivity, in which one lifts the functional inequality from phase space to space-time. In the \(L^2\) theory, space-time Poincaré inequalities were introduced by Albritton, Armstrong, Mourrat, and Novack AlbrittonArmstrongMourratNovack2024? and used quantitatively by Cao, Lu, and Wang CaoLuWang2023? to obtain a sharp hypocoercive \(L^2\) estimate for the underdamped Langevin semigroup.
Our analysis of hypercontractivity rests on an entropic version of the space-time functional inequality: for a density \(g\) on \([0,T]\times\mathbb{R}^{2d}\), we prove a space-time logarithmic Sobolev inequality of the form (see 1 for the precise statement) \[\mathop{\mathrm{Ent}}_{\bar\mu_T}(g) \lesssim I_{v,T}(g) +\frac{1}{\rho} \int_0^T \lVert(\partial_t+\mathcal{L}_a)g_t\rVert_{-1,g_t}^2\,\frac{\,\mathrm{d}t}{T}.\] This estimate replaces the missing instantaneous LSI, and is established in 3. The central argument is a controlled version of the entropy-decay estimate of Lu2026?.
Equipped with the space-time LSI, we prove hypocoercive hypercontractivity by a dual forward/backward time interpolation, in the spirit of the Bakry–Émery semigroup method BakryEmery1985?. Since the underdamped Langevin semigroup is non-reversible, the interpolation pairs the forward semigroup with its \(L^2(\mu)\)-adjoint. The details are given in [sec:interpolation] [sec:hypercontractivity].
Finally, before closing the introduction, let us mention that a closely related Rényi form of kinetic regularization, termed hyperequilibration by Altschuler and Chewi AltschulerChewi2024?, was used in the analysis of sampling algorithms. If \(\nu_t\) denotes the law at time \(t\), one asks for exponential relaxation of the Rényi divergence \[\operatorname R_q(\nu_t\mid\mu)\lesssim \mathrm{e}^{-\lambda t}\operatorname R_q(\nu_0\mid\mu), \qquad q>1.\] For overdamped Langevin dynamics, this type of Rényi decay was established earlier by Cao, Lu, and Lu CaoLuLu2019? and by Vempala and Wibisono VempalaWibisono2019?. In the analysis of sampling algorithms, in particular the Hamiltonian Monte Carlo (HMC) algorithm, such Rényi estimates are important for establishing warm starts and obtaining optimal complexity bounds ChenGatmiry2023?. In particular, Zhang, Altschuler, and Chewi ZhangAltschulerChewi2026? prove that a non-Metropolized HMC scheme, which is essentially a time-splitting discretization of the underdamped Langevin dynamics, can produce the required Rényi warm start using a Harnack inequality and shifted-composition estimates. As noted earlier, the coupling-based technique of Wang2017?, ZhangAltschulerChewi2026? only treats the high-friction regime and thus does not yield the kinetic hypercontractive rate.
As a consequence of the hypocoercive hypercontractivity, 2 gives the corresponding hyperequilibration estimate \[\operatorname R_q(\nu_t\mid\mu) \lesssim \mathrm{e}^{-\lambda_{\Gamma,\tau}\sqrt\rho\,t} \operatorname R_q(\nu_0\mid\mu), \qquad q>1,\] with the sharp hypocoercive rate (matching the entropy decay rate in Lu2026?). This improves the estimate used in ZhangAltschulerChewi2026? and may lead to sharper complexity estimates for the HMC algorithm; we leave this to future work.
Throughout, we work on the phase space \(\mathbb{R}^{2d}\) equipped with the reference Gibbs measure: \[\mu(\,\mathrm{d}x\,\mathrm{d}v) = \mu_x(\,\mathrm{d}x)\,\kappa(\,\mathrm{d}v),\] where \[\mu_x(\,\mathrm{d}x) = Z_x^{-1}\mathrm{e}^{-U(x)}\,\mathrm{d}x, \qquad \kappa(\,\mathrm{d}v) = (2\pi)^{-d/2}\mathrm{e}^{-\lvert v\rvert^2/2}\,\mathrm{d}v,\] and \(Z_x = \int_{\mathbb{R}^d} \mathrm{e}^{-U(x)}\,\mathrm{d}x\) is the spatial normalizing constant. We consider the underdamped Langevin semigroup \(\mathcal{P}_t = \mathrm{e}^{t(-\mathcal{L}_a+ \gamma \mathcal{L}_s)}\), \(t \ge 0\), with friction parameter \(\gamma > 0\), acting on probability densities relative to \(\mu\), where the Hamiltonian transport and velocity Ornstein–Uhlenbeck generators are \[\mathcal{L}_a= v\cdot\nabla_x - \nabla U\cdot\nabla_v, \qquad \mathcal{L}_s= \Delta_v - v\cdot\nabla_v,\] respectively. In \(L^2(\mu)\), the transport generator \(\mathcal{L}_a\) is skew-adjoint, while \(\mathcal{L}_s= -\nabla_v^{*}\nabla_v\) is symmetric and non-positive; here \(\nabla_v^{*} F = -\operatorname{div}_v F + v\cdot F\) denotes the adjoint of \(\nabla_v\) in \(L^2(\kappa)\).
We adopt the standing hypotheses of Lu2026?*Assumptions 2.1 and 2.2: the potential \(U \in C^\infty(\mathbb{R}^d)\) is convex, and the spatial marginal \(\mu_x\) satisfies the logarithmic Sobolev inequality \[\label{eq:lsi} \mathop{\mathrm{Ent}}_{\mu_x}(f) \le \frac{1}{2\rho}\int \frac{\lvert\nabla f\rvert^2}{f}\,\mathrm{d}\mu_x \qquad \text{for all } f \ge 0 \text{ with } \int f\,\mathrm{d}\mu_x= 1,\tag{3}\] with LSI constant \(\rho > 0\). In addition, \(U\) satisfies the tame Hérau–Nier confining bounds HerauNier2004?; we use them only through the regularization and semigroup-smoothing arguments of Lu2026?*Lemmas 5.3 and 7.1.
Let \(\Pi_v\) denote integration along the velocity fiber against \(\kappa\), i.e., \[(\Pi_v\varphi)(x) = \int_{\mathbb{R}^d} \varphi(x, v)\,\kappa(\,\mathrm{d}v).\] For a probability density \(g\) on \(\mathbb{R}^{2d}\) with respect to \(\mu\), the associated spatial marginal density and velocity current (i.e., the first-order \(v\)-moment) are \[q = \Pi_vg, \qquad j = \Pi_v(v g).\] On \(\{q > 0\}\) we define the conditional density of \(g\) given \(x\), \[h_x(v) = \frac{g(x, v)}{q(x)},\] extended by the convention \(h_x \equiv 1\) on \(\{q = 0\}\). With this notation, on \(\{q > 0\}\) the conditional velocity field \[m(x) := \frac{j(x)}{q(x)} = \int_{\mathbb{R}^d} v\, h_x(v)\,\kappa(\,\mathrm{d}v) = \mathbb{E}_{g\mu}[V \mid X = x]\] is the conditional mean of \(V\) given \(X\) under the joint law \((X, V) \sim g\mu\). We adopt the convention \(j = m = 0\) on \(\{q = 0\}\).
The velocity Fisher information of \(g\) is \[I_v(g) = \int \frac{\lvert\nabla_v g\rvert^2}{g}\,\mathrm{d}\mu\] with the standard convention \(I_v(g) = +\infty\) if the right-hand side is not well-defined. The relative entropy admits the chain-rule decomposition, as in Lu2026?*(2.14)–(2.15), \[\label{eq:ent-decomp} \mathop{\mathrm{Ent}}(g) := \mathop{\mathrm{Ent}}_\mu(g) = \mathop{\mathrm{Ent}}_x(q) + \mathop{\mathrm{Ent}}_v(g),\tag{4}\] where \[\mathop{\mathrm{Ent}}_x(q) := \mathop{\mathrm{Ent}}_{\mu_x}(q) = \int q \log q\,\mathrm{d}\mu_x\] is the relative entropy of the spatial marginal \(q\mu_x\) with respect to \(\mu_x\), and \[\mathop{\mathrm{Ent}}_v(g) := \int_{\{q > 0\}} q(x)\, \mathop{\mathrm{Ent}}_\kappa(h_x)\,\mathrm{d}\mu_x(x), \qquad \mathop{\mathrm{Ent}}_\kappa(h_x) := \int h_x \log h_x\,\mathrm{d}\kappa,\] is the conditional velocity entropy under \(g\mu\), relative to \(\kappa\) and averaged over the spatial marginal \(q\mu_x\). Both summands are non-negative, and the identity 4 is understood in the extended sense \([0, +\infty]\).
A central ingredient in the modified entropy method of Lu2026? for sharp hypocoercive entropy decay is the Wasserstein current corrector \(\mathcal{C}_{\rm OT}\), built from the Brenier transport map between \(q\mu_x\) and \(\mu_x\). The LSI 3 implies, via Talagrand’s inequality, that \(\mu_x\) has finite second moment. Hence, for every probability density \(q\) on \(\mathbb{R}^d\) such that \(q\mu_x\) has finite second moment, Brenier’s theorem yields a unique \(q\mu_x\)-a.e. optimal transport map \(T_q\colon\mathbb{R}^d\to\mathbb{R}^d\) pushing \(q\mu_x\) onto \(\mu_x\) brenier1991polar?, villani2021topics?. With the associated Brenier displacement \(\xi_q(x) := x - T_q(x)\), we set \[\label{eq:COT} \mathcal{C}_{\rm OT}(g) := \int j(x)\cdot\xi_q(x)\,\mathrm{d}\mu_x,\tag{5}\] whenever the integral is well-defined; a sufficient condition is recorded in 8 below.
Applying the Gaussian logarithmic Sobolev inequality of Gross Gross1975? to the conditional density \(h_x(v)\) and integrating against \(q\,\mathrm{d}\mu_x\) yields \[\label{eq:gausslsi} \mathop{\mathrm{Ent}}_v(g) \le \tfrac12I_v(g),\tag{6}\] which will be used repeatedly. The following lemma collects the estimates of Lu2026?*Section 3 that control the kinetic energy of the velocity current and the size of the Wasserstein corrector.
Lemma 1 (Lu2026?*Lemmas 3.1–3.2). For every probability density \(g\) on \(\mathbb{R}^{2d}\), \[\label{eq:imported-size} J(g) := \int \frac{\lvert j\rvert^2}{q}\,\mathrm{d}\mu_x\le 2\mathop{\mathrm{Ent}}_v(g),\tag{7}\] with the convention \(\lvert j\rvert^2/q := 0\) on \(\{q = 0\}\). If, in addition, \(\mathop{\mathrm{Ent}}(g) < \infty\), then \[\label{est:corrector} \lvert\mathcal{C}_{\rm OT}(g)\rvert \le \left(\frac{2}{\rho}\,\mathop{\mathrm{Ent}}_x(q)\,J(g)\right)^{1/2} \le \rho^{-1/2}\,\mathop{\mathrm{Ent}}(g),\tag{8}\] and, in particular, if \(I_v(g) < \infty\), \[\label{eq:imported-cot-product} \lvert\mathcal{C}_{\rm OT}(g)\rvert \le \left(\frac{2}{\rho}\,\mathop{\mathrm{Ent}}_x(q)\,I_v(g)\right)^{1/2},\tag{9}\] where \(\rho > 0\) is the LSI constant in 3 .
Two further structural inequalities from Lu2026? will be used to control the time evolution of the corrector \(\mathcal{C}_{\rm OT}(g_t)\). The first is the Wasserstein acceleration inequality of Lu2026?*Lemma 4.1, which depends only on the continuity equation for the spatial marginal \(q_t = \Pi_vg_t\). The second is the Brenier stress estimate of Lu2026?*Lemma 5.3, a static bound on the centered stress tensor, independent of the underlying evolution of \(g_t\). To formulate the estimates, we use a regularity class of time-dependent density-current pairs adapted from Lu2026?*Definition 2.5, retaining only the properties needed for the present argument.
Definition 1 (Regular density-current pair). A pair \(\{(q_t,j_t)\}_{t\in(0,T)}\) on \(\mathbb{R}^d\) is called a regular density-current pair if \[q\in C^1((0,T);C^\infty(\mathbb{R}^d)), \qquad j\in C^1((0,T);C^\infty(\mathbb{R}^d;\mathbb{R}^d)),\] and the following conditions hold.
For every \(t\in(0,T)\), \(q_t\) is a probability density with respect to \(\mu_x\), and \[\inf_{(t,x)\in K\times\mathbb{R}^d} q_t(x)>0 \qquad \text{for every compact } K\subset(0,T).\] Moreover, \(q_t\mu_x\) has a finite second moment, \(\mathop{\mathrm{Ent}}_x(q_t)<\infty\), and \(\int\frac{\lvert j_t\rvert^2}{q_t}\,\mathrm{d}\mu_x< \infty\).
The continuity equation \[\label{eq:continue} \partial_t q_t-\nabla_x^*j_t=0\tag{10}\] holds pointwise on \((0,T)\times\mathbb{R}^d\). Letting \[m_t=\frac{j_t}{q_t}, \qquad a_t=\partial_t m_t+(m_t\cdot\nabla_x)m_t,\] we assume that \(m_t,a_t\in L^2(q_t\mu_x)\) locally uniformly in \(t\). Moreover, when 10 is written as \(\partial_t(q_t\mu_x)+\operatorname{div}(m_tq_t\mu_x)=0\), its characteristics satisfy Lu2026?*Definition 2.5 (R3).
There exists a sequence of cutoff functions \(\chi_R\in C_c^\infty(\mathbb{R}^d)\), \(0\le\chi_R\le 1\), \(\chi_R\uparrow 1\), and \(\lvert\nabla\chi_R\rvert\le C/R\), such that, locally uniformly for \(t\in(0,T)\), \(\xi_{q_t}\cdot\nabla_x q_t\in L^1(\mu_x)\) and, as \(R\to\infty\), \[\begin{align} \int \chi_R\,\nabla_x q_t\cdot\xi_{q_t}\,\mathrm{d}\mu_x &\;\longrightarrow\; \int \nabla_x q_t\cdot\xi_{q_t}\,\mathrm{d}\mu_x,\\ \int q_t\,\nabla\chi_R\cdot\xi_{q_t}\,\mathrm{d}\mu_x &\;\longrightarrow\;0. \end{align}\]
Remark 1. Condition (A1) ensures that the Brenier map \(T_{q_t}\) is defined and that \(\xi_{q_t}=x-T_{q_t}(x)\) belongs to \(L^2(q_t\mu_x)\), by Talagrand’s inequality following from 3 . Condition (A2) justifies the second-order Wasserstein expansion used in the proof of 11 . Condition (A3), together with the stress cutoff in 3 below, is used to pass to the limit at spatial infinity in the localized integration-by-parts identities.
Lemma 2 (Lu2026?*Lemma 4.1). Let \(\{(q_t,j_t)\}_{t\in(0,T)}\) be a regular density-current pair in the sense of 1, and let \(\xi_{q_t}(x)=x-T_{q_t}(x)\) be the Brenier displacement from \(q_t\mu_x\) to \(\mu_x\). Then, in the sense of distributions on \((0,T)\), \[\label{eq:wasserstein-acceleration} \frac{\,\mathrm{d}}{\,\mathrm{d}t}\int j_t\cdot\xi_{q_t}\,\mathrm{d}\mu_x \le \int\frac{\lvert j_t\rvert^2}{q_t}\,\mathrm{d}\mu_x +\int \left( \partial_t j_t -\nabla_x^*\!\left(\frac{j_t\otimes j_t}{q_t}\right) \right)\cdot\xi_{q_t}\,\mathrm{d}\mu_x.\tag{11}\]
Lemma 3 (Lu2026?*Lemma 5.3). Let \(g\) be a probability density on \(\mathbb{R}^{2d}\) with respect to \(\mu\), and set \(q=\Pi_vg\) and \(j=\Pi_v(vg)\). Assume that \(q\) is strictly positive and belongs to \(C^\infty(\mathbb{R}^d)\), and write \(g(x,v)=q(x)h_x(v)\). Assume that the conditional mean and covariance \[m(x)=\int v\,h_x(v)\,\mathrm{d}\kappa, \qquad \Sigma_g(x)=\int (v-m(x))\otimes(v-m(x))h_x(v)\,\mathrm{d}\kappa\] exist for a.e. \(x\). Define the centered stress tensor by \[\label{def:tensor} \Theta_g := \Pi_v(v\otimes v\,g)-\frac{j\otimes j}{q}-qI_d = q(\Sigma_g-I_d).\tag{12}\] Assume that \(\mathop{\mathrm{Ent}}_x(q)\) and \(\mathop{\mathrm{Ent}}_v(g)\) are finite, and that \(\Theta_g\in C^1_{\mathrm{loc}}(\mathbb{R}^d;\mathbb{R}^{d\times d})\). Let \(\xi_q=x-T_q\) be the Brenier displacement from \(q\mu_x\) to \(\mu_x\). Assume moreover that there exists a cutoff sequence \(\{\chi_R\}\) such that condition (A3) in 1 holds for the marginal \(q\), and \[\int \xi_q\cdot\nabla_x^*(\chi_R\Theta_g)\,\mathrm{d}\mu_x \longrightarrow \int \xi_q\cdot\nabla_x^*\Theta_g\,\mathrm{d}\mu_x,\] with \(\xi_q\cdot\nabla_x q,\xi_q\cdot\nabla_x^*\Theta_g\in L^1(\mu_x)\). Then, for every \(0<\beta<1/2\), \[\label{eq:brenier-stress} -\int\nabla_x q\cdot\xi_q\,\mathrm{d}\mu_x + \int\xi_q\cdot\nabla_x^*\Theta_g\,\mathrm{d}\mu_x \le \beta^{-1}\mathop{\mathrm{Ent}}_v(g)-\mathop{\mathrm{Ent}}_x(q).\tag{13}\]
This section is devoted to the space-time logarithmic Sobolev inequality for the underdamped Langevin dynamics. To this end, we first introduce the time-augmented reference measure: \[\,\mathrm{d}\bar\mu_T= \frac{1}{T}\,\mathrm{d}t\,\mathrm{d}\mu\] on \([0,T]\times\mathbb{R}^{2d}\). For a positive space-time density \(g\) on \([0,T]\times\mathbb{R}^{2d}\), the time-averaged entropy and time-averaged velocity Fisher information are defined analogously by \[\mathop{\mathrm{Ent}}_{\bar\mu_T}(g) = \int g \log g \,\mathrm{d}\bar\mu_T, \qquad I_{v,T}(g) = \int \frac{\lvert\nabla_v g\rvert^2}{g}\,\mathrm{d}\bar\mu_T.\] For a signed distribution \(r_t\) on \(\mathbb{R}^{2d}\) in the dual of \(C_c^\infty(\mathbb{R}^{2d})\), we define the weighted velocity negative norm by \[\label{eq:negative-norm} \lVert r_t\rVert_{-1, g_t}^2 = \sup_{\zeta \in C_c^\infty(\mathbb{R}^{2d})} \left\{ 2 \langle r_t, \zeta \rangle - \int g_t \lvert\nabla_v \zeta\rvert^2\,\mathrm{d}\mu \right\},\tag{14}\] where \(\langle\cdot,\cdot\rangle\) denotes the distributional pairing (which reduces to \(\int r_t \zeta\,\mathrm{d}\mu\) when \(r_t\) is a function). Equivalently, by Legendre duality, \[\label{eq:negative-representation} \lVert r_t\rVert_{-1, g_t}^2 = \inf \left\{ \int g_t \lvert A_t\rvert^2\,\mathrm{d}\mu \;:\; A_t \in L^2(g_t\,\mathrm{d}\mu;\, \mathbb{R}^d),\;r_t = \nabla_v^{*}(g_t A_t) \right\},\tag{15}\] with the convention that the infimum is \(+\infty\) if no such velocity field \(A_t\) exists. In particular, \(\langle \nabla_v^{*}(g_t A_t), \varphi \rangle = 0\) for every test function \(\varphi = \varphi(x)\) depending only on \(x\); hence, if \(\lVert r_t\rVert_{-1, g_t} < \infty\), then \(\Pi_vr_t = 0\), i.e., the \(x\)-marginal of \(r_t\) vanishes.
Remark 2. This norm is the velocity-fiber analogue of the Wasserstein tangent metric ambrosio2005gradient?*Section 8. To see this, write \(g_t(x,v)=q_t(x)h_{t,x}(v)\); then on \(\{q_t>0\}\) we have \[\dot{h}_{t,x}(v) = r_t(x,v)/q_t(x) \qquad \text{with } \int \dot{h}_{t,x}\,\mathrm{d}\kappa=0, \;\text{for \mu_x-a.e. x,}\] when \(\lVert r_t\rVert_{-1, g_t} < \infty\). For regular data, the representation 15 disintegrates as \[\lVert r_t\rVert_{-1,g_t}^2 = \int_{\mathbb{R}^d} q_t(x)\, \lVert\dot{h}_{t,x}\rVert_{-1,h_{t,x}}^2 \,\mathrm{d}\mu_x(x),\] where \[\lVert\dot{h}\rVert_{-1,h}^2 := \inf_{\dot{h}=\nabla_v^*(ha)} \int h\,\lvert a\rvert^2\,\mathrm{d}\kappa = \sup_{\phi\in C_c^\infty(\mathbb{R}^d)} \left\{ 2\int \dot{h}\,\phi\,\mathrm{d}\kappa - \int h\,\lvert\nabla_v\phi\rvert^2\,\mathrm{d}\kappa \right\}.\] Thus \(\lVert r_t\rVert_{-1,g_t}^2\) is nothing else than the average over \(q_t\mu_x\) of the Wasserstein tangent norm at the conditional law \(h_{t,x}\kappa\).
The remainder of this section establishes the space-time LSI, 1, in three steps. The first step, 4, is a controlled version of the entropy-current calculation in Lu2026?*Section 6: it bounds the evolution of the corrector \(\mathcal{C}_{\rm OT}(g_t)\) along a path driven by an additional velocity forcing. The second step, 1, combines this corrector inequality with the entropy identity for the controlled equation to yield a differential inequality for a modified entropy, at the cost of a quadratic penalty in the forcing. The third step integrates this differential inequality in time and applies an elementary endpoint estimate to deliver the space-time LSI 36 in 1. Finally, we discuss how to remove the regularity assumption used along the way.
Lemma 4 (Controlled corrector differential inequality). Let \(\{g_t\}_{t\in(0,T)}\) be a path of probability densities with respect to \(\mu\) satisfying \[\label{eq:reg-gt-smooth} g\in C^1\bigl((0,T);\,C^\infty(\mathbb{R}^{2d})\bigr) \quad\text{and}\quad \inf_{(t,x,v)\in K\times\mathbb{R}^{2d}} g_t(x,v)>0 \quad\text{for every compact }K\subset(0,T),\tag{16}\] and let \(u:(0,T)\times\mathbb{R}^{2d}\to\mathbb{R}^d\) be a measurable velocity-control field with finite action \[\label{eq:finite-action-u} \int_0^T\int g_t\lvert u_t\rvert^2\,\mathrm{d}\mu\,\mathrm{d}t<\infty.\tag{17}\] Assume that \(g_t\) satisfies the controlled forward Kolmogorov equation \[\label{eq:controlled-kolmogorov} (\partial_t+\mathcal{L}_a)g_t = \gamma\mathcal{L}_sg_t+\nabla_v^{*}(g_tu_t)\tag{18}\] pointwise on \((0,T)\times\mathbb{R}^{2d}\). Assume that \(\{(q_t,j_t)= (\Pi_vg_t, \Pi_v(vg_t))\}_{t\in(0,T)}\) is a regular density-current pair in the sense of 1, and that, for a.e.\(t\in(0,T)\), the centered stress tensor \(\Theta_t\), defined as in 12 , satisfies the hypotheses of 3.
Define the spatial source induced by the velocity control: \[\label{eq:Bt-def} B_t(x) := \Pi_v(g_tu_t)(x) = \int g_t(x,v)u_t(x,v)\,\kappa(\,\mathrm{d}v).\tag{19}\] Then, for a.e.\(t\in(0,T)\), \(B_t\in L^2(q_t^{-1}\mu_x;\mathbb{R}^d)\) with \[\int\frac{\lvert B_t\rvert^2}{q_t}\,\mathrm{d}\mu_x \le \int g_t\lvert u_t\rvert^2\,\mathrm{d}\mu,\] and hence \[\label{eq:Bt-xi-bound} \left\lvert\int B_t\cdot \xi_{q_t}\,\mathrm{d}\mu_x\right\rvert \le \left(\frac{2}{\rho}\,\mathop{\mathrm{Ent}}_x(q_t)\right)^{1/2} \left(\int g_t\lvert u_t\rvert^2\,\mathrm{d}\mu\right)^{1/2}.\tag{20}\] Moreover, the controlled corrector differential inequality \[\label{eq:forced-cot-diff} \frac{\,\mathrm{d}}{\,\mathrm{d}t}\mathcal{C}_{\rm OT}(g_t) \le -\mathop{\mathrm{Ent}}_x(q_t) -\gamma\mathcal{C}_{\rm OT}(g_t) +3I_v(g_t) +\int B_t\cdot\xi_{q_t}\,\mathrm{d}\mu_x\tag{21}\] holds in the sense of distributions on \((0,T)\).
Proof. All identities below are understood in the weak sense in the spatial variables and in the sense of distributions in time.
We first record an elementary bound on \(B_t(x)\) defined in 19 . A direct application of Jensen’s inequality gives, for a.e.\(t\), \[\label{eq:Bt-action-proof} |B_t(x)|^2 = q_t(x)^2 \left| \int u_t(x,v) h_{t,x}(v)\,\kappa(\,\mathrm{d}v) \right|^2 \le q_t(x) \int g_t(x,v)|u_t(x,v)|^2\,\kappa(\,\mathrm{d}v).\tag{22}\] We then have \[\label{eq:Bt-L2q-proof} \int \frac{|B_t|^2}{q_t}\,\mathrm{d}\mu_x \le \int g_t |u_t|^2\,\mathrm{d}\mu < \infty.\tag{23}\] In particular, the pairing with the Brenier displacement \(\xi_{q_t}\) is well-defined by \[\left|\int B_t\cdot \xi_{q_t}\,\mathrm{d}\mu_x\right| \le \left(\int \frac{|B_t|^2}{q_t}\,\mathrm{d}\mu_x\right)^{1/2} \left(\int q_t |\xi_{q_t}|^2\,\mathrm{d}\mu_x\right)^{1/2},\] where the second term is bounded by Talagrand’s inequality Otto2001?: \[\int q_t\lvert\xi_{q_t}\rvert^2\,\mathrm{d}\mu_x =W_2^2(q_t\mu_x,\mu_x) \le \frac{2}{\rho}\,\mathop{\mathrm{Ent}}_x(q_t).\] Combining the last two displays with 23 gives 20 .
We next show that the velocity control does not contribute to the equation for the \(x\)-marginal \(q_t\). Indeed, for any smooth test function \(b=b(x)\), integration by parts in \(v\) gives \[\int b(x)\,\nabla_v^{*}(g_tu_t)\,\mathrm{d}\mu = \int g_tu_t\cdot \nabla_v b(x)\,\mathrm{d}\mu = 0,\] because \(\nabla_v b\equiv 0\). It follows that \[\Pi_v\bigl(\nabla_v^{*}(g_t u_t)\bigr)=0 .\] Applying \(\Pi_v\) to 18 , and using \(\Pi_v(\mathcal{L}_sg_t)=0\), we obtain the same spatial continuity equation as in the uncontrolled case (i.e., \(u_t = 0\)): \[\label{eq:qt-eqn-controlled} \partial_t q_t=\nabla_x^{*}j_t.\tag{24}\]
We now compute the equation for the first velocity moment \(j_t=\Pi_v(vg_t)\). For the control term, testing against a smooth vector field \(a=a(x)\) gives \[\label{eq:sourcefirst} \begin{align} \int a(x)\cdot \Pi_v\bigl(v\,\nabla_v^{*}(g_tu_t)\bigr)\,\mathrm{d}\mu_x &= \int (v\cdot a(x))\,\nabla_v^{*}(g_tu_t)\,\mathrm{d}\mu \\ &= \int g_tu_t\cdot a(x)\,\mathrm{d}\mu = \int a(x)\cdot B_t(x)\,\mathrm{d}\mu_x. \end{align}\tag{25}\] Thus, the control contributes a source term \(B_t\) to the first-moment equation. The remaining terms follow from the same computation as in the uncontrolled case Lu2026?*(4.3). It follows that \[\label{eq:jt-eqn-controlled} \partial_t j_t = -\nabla_x q_t +\nabla_x^{*}\!\left(\frac{j_t\otimes j_t}{q_t}\right) +\nabla_x^{*}\Theta_t -\gamma j_t +B_t.\tag{26}\]
Since the marginal equation 24 is unchanged by the velocity control, the Wasserstein acceleration inequality 11 applies to \((q_t,j_t)\). Substituting 26 into it and recalling \(\mathcal{C}_{\rm OT}(g_t)=\int j_t\cdot \xi_{q_t}\,\mathrm{d}\mu_x\), we obtain \[\begin{align} \frac{\,\mathrm{d}}{\,\mathrm{d}t}\mathcal{C}_{\rm OT}(g_t) &\le \int\frac{\lvert j_t\rvert^2}{q_t}\,\mathrm{d}\mu_x -\int\nabla_x q_t\cdot\xi_{q_t}\,\mathrm{d}\mu_x +\int\xi_{q_t}\cdot\nabla_x^*\Theta_t\,\mathrm{d}\mu_x -\gamma\int j_t\cdot\xi_{q_t}\,\mathrm{d}\mu_x +\int B_t\cdot\xi_{q_t}\,\mathrm{d}\mu_x. \end{align}\] The Brenier stress estimate 13 , applied with \(\beta=1/4\), gives \[-\int\nabla_x q_t\cdot\xi_{q_t}\,\mathrm{d}\mu_x +\int\xi_{q_t}\cdot\nabla_x^*\Theta_t\,\mathrm{d}\mu_x \le 4\mathop{\mathrm{Ent}}_v(g_t)-\mathop{\mathrm{Ent}}_x(q_t).\] The remaining kinetic-current term satisfies \[\int\frac{\lvert j_t\rvert^2}{q_t}\,\mathrm{d}\mu_x+4\mathop{\mathrm{Ent}}_v(g_t) \le 6\mathop{\mathrm{Ent}}_v(g_t) \le 3I_v(g_t),\] by 7 and the Gaussian LSI 6 , while the friction term is exactly the corrector: \[-\gamma \int j_t\cdot \xi_{q_t}\,\mathrm{d}\mu_x = -\gamma \mathcal{C}_{\rm OT}(g_t).\] Consequently, the controlled corrector inequality is \[\frac{\,\mathrm{d}}{\,\mathrm{d}t}\mathcal{C}_{\rm OT}(g_t) \le -\mathop{\mathrm{Ent}}_x(q_t) -\gamma \mathcal{C}_{\rm OT}(g_t) +3I_v(g_t) +\int B_t\cdot \xi_{q_t}\,\mathrm{d}\mu_x,\] in the sense of distributions on \((0,T)\). ◻
The estimate 21 in 4 differs from the uncontrolled corrector inequality only through the source pairing \(\int B_t\cdot\xi_{q_t}\,\mathrm{d}\mu_x\) with \(B_t=\Pi_v(g_tu_t)\). The next proposition combines this inequality 21 with the entropy dissipation identity for the controlled equation to derive a differential inequality, along the controlled path, for the modified entropy functional of Lu2026?, \[\label{def:modifiedentropy} \mathcal{H}_\theta(g_t):=\mathop{\mathrm{Ent}}(g_t)+\theta\sqrt\rho\,\mathcal{C}_{\rm OT}(g_t).\tag{27}\] With the hypocoercive weight \(\theta\sqrt\rho\) and friction parameter \(\gamma=\Gamma\sqrt\rho\) chosen appropriately, the spatial entropy \(\mathop{\mathrm{Ent}}_x(q_t)\) and velocity Fisher information \(I_v(g_t)\) jointly absorb the corrector term \(-\theta\sqrt\rho\,\gamma\,\mathcal{C}_{\rm OT}(g_t)\) from 21 , while the control-dependent terms are bounded by a single quadratic action cost \(\rho^{-1/2}\int g_t\lvert u_t\rvert^2\,\mathrm{d}\mu\).
Proposition 1 (Controlled modified entropy estimate). Fix \(\Gamma>0\) and set \[\label{eq:gammaprop} \gamma=\Gamma\sqrt\rho, \qquad \theta=\frac{1}{8}\min\bigl\{\Gamma,\Gamma^{-1}\bigr\}.\qquad{(1)}\] Let \((g_t,u_t)\) satisfy the assumptions of 4, and \(\mathcal{H}_\theta(g_t)\) be the modified entropy 27 . Then, in the sense of distributions on \((0,T)\), \[\label{eq:controlled-diff} \frac{\,\mathrm{d}}{\,\mathrm{d}t}\mathcal{H}_\theta(g_t) \le -\frac{\theta}{2}\sqrt\rho\,\mathop{\mathrm{Ent}}(g_t) +C_\Gamma\,\rho^{-1/2}\int g_t\lvert u_t\rvert^2\,\mathrm{d}\mu,\qquad{(2)}\] where the constant \(C_\Gamma:=\Gamma^{-1}+2\theta\) depends only on \(\Gamma\).
Proof. By the smoothness assumption 16 on \(g_t\), all identities below hold pointwise in \((t,x,v)\); the displayed differential inequalities are then understood in the sense of distributions in \(t\), i.e., tested against nonnegative \(\varphi\in C_c^\infty((0,T))\).
First, 4 applied with \(\gamma=\Gamma\sqrt\rho\) from ?? gives the controlled corrector inequality \[\label{eq:controlled-cot-diff} \frac{\,\mathrm{d}}{\,\mathrm{d}t}\mathcal{C}_{\rm OT}(g_t) \le -\mathop{\mathrm{Ent}}_x(q_t)-\gamma\mathcal{C}_{\rm OT}(g_t)+3I_v(g_t) +\int B_t\cdot\xi_{q_t}\,\mathrm{d}\mu_x.\tag{28}\] Moreover, the skew-adjointness of \(\mathcal{L}_a\) and the identity \(\mathcal{L}_s=-\nabla_v^{*}\nabla_v\) in \(L^2(\mu)\), together with the mass conservation \(\int g_t\,\mathrm{d}\mu=1\), yield the entropy dissipation identity for the controlled equation: \[\label{eq:controlled-entropy-diff} \frac{\,\mathrm{d}}{\,\mathrm{d}t}\mathop{\mathrm{Ent}}(g_t) =-\gamma I_v(g_t) +\int g_tu_t\cdot\nabla_v\log g_t\,\mathrm{d}\mu.\tag{29}\]
We next estimate the two \(u_t\)-linear terms in 28 and 29 arising from the control. For the entropy cross term in 29 , Young’s inequality with weight \(\gamma/2\) gives \[\label{eq:entropy-control-bound} \int g_tu_t\cdot\nabla_v\log g_t\,\mathrm{d}\mu \le \frac{\gamma}{4}I_v(g_t) +\frac{1}{\gamma}\int g_t\lvert u_t\rvert^2\,\mathrm{d}\mu.\tag{30}\] On the other hand, for the source pairing in 28 , combining the bound 20 from 4 with Young’s inequality yields \[\label{eq:cot-control-bound} \left\lvert\int B_t\cdot\xi_{q_t}\,\mathrm{d}\mu_x\right\rvert \le \frac{1}{4}\mathop{\mathrm{Ent}}_x(q_t) +\frac{2}{\rho}\int g_t\lvert u_t\rvert^2\,\mathrm{d}\mu.\tag{31}\]
We now bound the time derivative of the modified entropy functional \(\mathcal{H}_\theta(g_t)\) in 27 . For this, it suffices to combine 28 and 29 , and substitute the estimates 30 and 31 . A straightforward computation, using \(\gamma=\Gamma\sqrt\rho\), leads to \[\begin{gather} \label{eq:pre-absorb-controlled} \frac{\,\mathrm{d}}{\,\mathrm{d}t}\mathcal{H}_\theta(g_t) \le -\frac{3\theta}{4}\sqrt\rho\,\mathop{\mathrm{Ent}}_x(q_t) -\theta\Gamma\rho\,\mathcal{C}_{\rm OT}(g_t) -\left(\frac{3\Gamma}{4}-3\theta\right)\sqrt\rho\,I_v(g_t) \\ +C_\Gamma\,\rho^{-1/2}\int g_t\lvert u_t\rvert^2\,\mathrm{d}\mu, \end{gather}\tag{32}\] with \(C_\Gamma=\Gamma^{-1}+2\theta\).
It remains to estimate the indefinite-sign term \(-\theta\Gamma\rho\,\mathcal{C}_{\rm OT}(g_t)\) in 32 . By 9 in 1 and Young’s inequality again, we have \[\label{eq:cot-absorption} \theta\Gamma\rho\,\lvert\mathcal{C}_{\rm OT}(g_t)\rvert \le \theta\Gamma\sqrt{2\rho\,\mathop{\mathrm{Ent}}_x(q_t)\,I_v(g_t)} \le \frac{\theta}{4}\sqrt\rho\,\mathop{\mathrm{Ent}}_x(q_t) +2\theta\Gamma^2\sqrt\rho\,I_v(g_t).\tag{33}\] We take \(\theta=\frac{1}{8}\min\{\Gamma,\Gamma^{-1}\}\) as in ?? , ensuring \(\theta\le\Gamma/8\) and \(\theta\Gamma^2\le\Gamma/8\), so that \[\label{eq:constantest} \frac{3\Gamma}{4}-3\theta-2\theta\Gamma^2 \ge \frac{3\Gamma}{4}-\frac{3\Gamma}{8}-\frac{2\Gamma}{8} = \frac{\Gamma}{8} \ge \frac{\theta}{4}.\tag{34}\] Substituting 33 into 32 and using 34 , we arrive at \[\frac{\,\mathrm{d}}{\,\mathrm{d}t}\mathcal{H}_\theta(g_t) \le -\frac{\theta}{2}\sqrt\rho\,\mathop{\mathrm{Ent}}_x(q_t) -\frac{\theta}{4}\sqrt\rho\,I_v(g_t) +C_\Gamma\,\rho^{-1/2}\int g_t\lvert u_t\rvert^2\,\mathrm{d}\mu.\] Finally, the Gaussian LSI 6 gives \(\mathop{\mathrm{Ent}}_v(g_t)\le\frac{1}{2}I_v(g_t)\), and the entropy decomposition 4 then yields \[-\frac{\theta}{2}\sqrt\rho\,\mathop{\mathrm{Ent}}_x(q_t)-\frac{\theta}{4}\sqrt\rho\,I_v(g_t) \le -\frac{\theta}{2}\sqrt\rho\,\bigl(\mathop{\mathrm{Ent}}_x(q_t)+\mathop{\mathrm{Ent}}_v(g_t)\bigr) =-\frac{\theta}{2}\sqrt\rho\,\mathop{\mathrm{Ent}}(g_t),\] which establishes ?? . ◻
The space-time LSI now follows by applying ?? to a controlled equation built from a near-minimizer in the negative norm. More precisely, if \(r_t=(\partial_t+\mathcal{L}_a)g_t\), then 15 allows us, for every \(\varepsilon>0\), to choose a space-time measurable velocity field \(A_t(x,v)\), defined \(g_t\,\mathrm{d}\bar\mu_T\)-almost everywhere, such that \(A\in L^2(g_t\,\mathrm{d}\bar\mu_T)\), \[r_t=\nabla_v^{*}(g_tA_t) \quad\text{in }\mathcal{D}'(\mathbb{R}^{2d})\text{ for a.e. }t, \qquad \int g_t\lvert A_t\rvert^2\,\mathrm{d}\bar\mu_T \le \int_0^T\lVert(\partial_t+\mathcal{L}_a)g_t\rVert_{-1,g_t}^2\,\frac{\,\mathrm{d}t}{T} +\varepsilon.\] We then set \[u_t=A_t+\gamma\nabla_v\log g_t,\] so that, because \(\mathcal{L}_s=-\nabla_v^{*}\nabla_v\), the equation above becomes the controlled equation in 1. The quadratic control cost is then bounded by the negative-norm action plus a multiple of the velocity Fisher information. A short endpoint estimate controls \(\mathop{\mathrm{Ent}}(g_0)\) by the time average and the action, and the hypothesis \(T\ge\tau_{\rm ST}\rho^{-1/2}\) leaves a positive coefficient in front of the resulting time-averaged entropy.
Theorem 1 (Space-time logarithmic Sobolev inequality). Fix \(\Gamma>0\), and set \[\label{eq:constantsl} \gamma:=\Gamma\sqrt\rho, \qquad M:=\max\{\Gamma,\Gamma^{-1}\}\ge 1.\tag{35}\] There exist constants \[\tau_{\rm ST}(\Gamma) = 32M+4, \qquad C_{\rm ST}(\Gamma) = 80M^2+16M+2,\] depending only on \(\Gamma\), such that the following holds. Let \(T\ge\tau_{\rm ST}\,\rho^{-1/2}\), and let \(\{g_t\}_{t\in(0,T)}\) be a path of probability densities on \(\mathbb{R}^{2d}\) with respect to \(\mu\) satisfying the regularity assumptions on \(\{g_t\}\) in 4. Then the space-time logarithmic Sobolev inequality holds: \[\label{eq:stlsi} \mathop{\mathrm{Ent}}_{\bar\mu_T}(g) \le C_{\rm ST} \left[ I_{v,T}(g) +\frac{1}{\rho} \int_0^T\lVert(\partial_t+\mathcal{L}_a)g_t\rVert_{-1,g_t}^2\,\frac{\,\mathrm{d}t}{T} \right].\tag{36}\]
Proof. If the right-hand side of 36 is infinite, the inequality holds trivially; we therefore assume it is finite. From the constant choice 35 , we have \(\theta=1/(8M)\) in 1.
First, by the definition 15 of \(\|\cdot\|_{-1,g_t}\), for every \(\varepsilon>0\), one can choose a measurable velocity field \(A_t\) such that \[\label{eq:st-current} (\partial_t+\mathcal{L}_a)g_t=\nabla_v^{*}(g_tA_t)\tag{37}\] holds in \(\mathcal{D}'(\mathbb{R}^{2d})\) for a.e.\(t\in[0,T]\), with \[\label{eq:st-current-cost} \int_0^T\int g_t\lvert A_t\rvert^2\,\mathrm{d}\mu\,\frac{\,\mathrm{d}t}{T} \le \int_0^T\lVert(\partial_t+\mathcal{L}_a)g_t\rVert_{-1,g_t}^2\,\frac{\,\mathrm{d}t}{T}+\varepsilon.\tag{38}\] Recalling that \(\int g_t\,\mathrm{d}\mu=1\) for every \(t\in[0,T]\), we have \[\label{eq:space-time-entropy-average} \mathop{\mathrm{Ent}}_{\bar\mu_T}(g)=\frac{1}{T}\int_0^T\mathop{\mathrm{Ent}}(g_t)\,\mathrm{d}t,\tag{39}\] due to \[\mathop{\mathrm{Ent}}(g_t) = \int g_t \log g_t \,\mathrm{d}\mu - \left(\int g_t \,\mathrm{d}\mu\right) \log \left(\int g_t \,\mathrm{d}\mu\right) = \int g_t \log g_t \,\mathrm{d}\mu\,.\]
We next introduce \[u_t:=A_t+\gamma\nabla_v\log g_t.\] Since \(\mathcal{L}_s=-\nabla_v^{*}\nabla_v\), the divergence equation 37 can be rewritten as the controlled forward Kolmogorov equation: \[(\partial_t+\mathcal{L}_a)g_t=\gamma\mathcal{L}_sg_t+\nabla_v^{*}(g_tu_t),\] to which 1 applies with the explicit constant \(C_\Gamma=\Gamma^{-1}+2\theta\). Integrating ?? over \([0,T]\) gives \[\label{est:modifyentropy} \frac{\theta}{2}\sqrt\rho\int_0^T\mathop{\mathrm{Ent}}(g_t)\,\mathrm{d}t \le \mathcal{H}_\theta(g_0)-\mathcal{H}_\theta(g_T) +(\Gamma^{-1}+2\theta)\,\rho^{-1/2}\int_0^T\int g_t\lvert u_t\rvert^2\,\mathrm{d}\mu\,\mathrm{d}t.\tag{40}\] From the corrector bound 8 , we readily have \[(1-\theta)\mathop{\mathrm{Ent}}(g_t)\le\mathcal{H}_\theta(g_t)\le(1+\theta)\mathop{\mathrm{Ent}}(g_t),\] and thus \[\mathcal{H}_\theta(g_0)-\mathcal{H}_\theta(g_T) \le \mathcal{H}_\theta(g_0) \le (1+\theta)\mathop{\mathrm{Ent}}(g_0).\] Using the elementary bound \[\lvert u_t\rvert^2\le 2\lvert A_t\rvert^2+2\gamma^2\lvert\nabla_v\log g_t\rvert^2\] and \(\gamma^2=\Gamma^2\rho\), one can estimate the \(u_t\)-action cost term in 40 as follows: \[\rho^{-1/2}\int g_t\lvert u_t\rvert^2\,\mathrm{d}\mu \le 2\rho^{-1/2}\int g_t\lvert A_t\rvert^2\,\mathrm{d}\mu +2\Gamma^2\sqrt\rho\,I_v(g_t).\] Introducing constants \[\label{eq:K1K2-def} K_1 := 2(\Gamma^{-1}+2\theta)\Gamma^2 = 2\Gamma+4\theta\Gamma^2, \qquad K_2 := 2(\Gamma^{-1}+2\theta) = 2\Gamma^{-1}+4\theta,\tag{41}\] by the above estimates, we obtain \[\label{eq:integrated-controlled} \frac{\theta}{2}\sqrt\rho\int_0^T\mathop{\mathrm{Ent}}(g_t)\,\mathrm{d}t \le (1+\theta)\mathop{\mathrm{Ent}}(g_0) +K_1\sqrt\rho\int_0^TI_v(g_t)\,\mathrm{d}t +\frac{K_2}{\sqrt\rho}\int_0^T\int g_t\lvert A_t\rvert^2\,\mathrm{d}\mu\,\mathrm{d}t.\tag{42}\]
We now estimate the initial entropy \(\mathop{\mathrm{Ent}}(g_0)\) in 42 . For this, we compute the entropy dissipation for the solution \(g_t\) to 37 : \[\frac{\,\mathrm{d}}{\,\mathrm{d}t}\mathop{\mathrm{Ent}}(g_t) =\int g_tA_t\cdot\nabla_v\log g_t\,\mathrm{d}\mu,\] which by Young’s inequality with weight \(\sqrt\rho\) gives \[\label{eq:ent-derivative-bound} \left\lvert\frac{\,\mathrm{d}}{\,\mathrm{d}t}\mathop{\mathrm{Ent}}(g_t)\right\rvert \le\frac{\sqrt\rho}{2}I_v(g_t) +\frac{1}{2\sqrt\rho}\int g_t\lvert A_t\rvert^2\,\mathrm{d}\mu.\tag{43}\] Since \(t\mapsto\mathop{\mathrm{Ent}}(g_t)\) is absolutely continuous, for every \(s\in[0,T]\), \[\label{eq:estentropy} \mathop{\mathrm{Ent}}(g_0) =\mathop{\mathrm{Ent}}(g_s)-\int_0^s\frac{\,\mathrm{d}}{\,\mathrm{d}\ell}\mathop{\mathrm{Ent}}(g_\ell)\,\mathrm{d}\ell \le\mathop{\mathrm{Ent}}(g_s)+\int_0^T\left\lvert\frac{\,\mathrm{d}}{\,\mathrm{d}\ell}\mathop{\mathrm{Ent}}(g_\ell)\right\rvert\,\mathrm{d}\ell.\tag{44}\] Averaging 44 over \(s\in[0,T]\), we obtain \[\label{eq:endpoint-control} \mathop{\mathrm{Ent}}(g_0) \le\mathop{\mathrm{Ent}}_{\bar\mu_T}(g) +\frac{\sqrt\rho}{2}\int_0^TI_v(g_t)\,\mathrm{d}t +\frac{1}{2\sqrt\rho}\int_0^T\int g_t\lvert A_t\rvert^2\,\mathrm{d}\mu\,\mathrm{d}t,\tag{45}\] by using 39 and substituting 43 .
Finally, substituting 45 into 42 , and rearranging the terms, we can conclude \[\begin{gather} \label{eq:estintentropy} \left(\frac{\theta}{2}\sqrt\rho-\frac{1+\theta}{T}\right) \int_0^T\mathop{\mathrm{Ent}}(g_t)\,\mathrm{d}t \\ \le \left(K_1+\frac{1+\theta}{2}\right)\sqrt\rho\int_0^TI_v(g_t)\,\mathrm{d}t +\frac{K_2+(1+\theta)/2}{\sqrt\rho}\int_0^T\int g_t\lvert A_t\rvert^2\,\mathrm{d}\mu\,\mathrm{d}t. \end{gather}\tag{46}\]
Define \[\label{eq:tau-ST-def} \tau_{\rm ST}:=\frac{4(1+\theta)}{\theta}=32M+4.\tag{47}\] The choice of \(T\ge\tau_{\rm ST}\,\rho^{-1/2}\) then yields \[\frac{\theta}{2}\sqrt\rho-\frac{1+\theta}{T} \ge \frac{\theta}{4}\sqrt\rho\] for the prefactor in 46 . Therefore, dividing 46 by \(\frac{\theta}{4}\sqrt\rho\cdot T\) and using 39 again gives \[\label{eq:estintentropy2} \mathop{\mathrm{Ent}}_{\bar\mu_T}(g) \le \frac{4}{\theta}\!\left(K_1+\frac{1+\theta}{2}\right)I_{v,T}(g) +\frac{4}{\theta}\!\left(K_2+\frac{1+\theta}{2}\right)\frac{1}{\rho}\int_0^T\int g_t\lvert A_t\rvert^2\,\mathrm{d}\mu\,\frac{\,\mathrm{d}t}{T}.\tag{48}\] A direct computation from 41 with \(\theta=1/(8M)\) gives \(K_1,K_2\le 5M/2\), and then the two coefficients on the right-hand side of 48 are both bounded by \[\label{eq:CST-explicit} C_{\rm ST} :=\frac{4}{\theta}\!\left(\frac{5M}{2}+\frac{1+\theta}{2}\right) =\frac{10M}{\theta}+\frac{2(1+\theta)}{\theta} =80M^2+16M+2.\tag{49}\] Therefore \[\mathop{\mathrm{Ent}}_{\bar\mu_T}(g) \le C_{\rm ST} \left[ I_{v,T}(g) +\frac{1}{\rho}\int_0^T\int g_t\lvert A_t\rvert^2\,\mathrm{d}\mu\,\frac{\,\mathrm{d}t}{T} \right].\] Recalling 38 , letting \(\varepsilon\to 0\) yields 36 . The proof is complete. ◻
Remark 3 (Optimal constants). With \(M:=\max\{\Gamma,\Gamma^{-1}\}\ge 1\), the constants in 1 read \[\tau_{\rm ST}(\Gamma)=32M+4, \qquad C_{\rm ST}(\Gamma)=80M^2+16M+2\le 98M^2,\] and both attain their minimum at the friction \(\Gamma=1\), where \(\tau_{\rm ST}(1)=36\) and \(C_{\rm ST}(1)=98\). Moreover, as \(\Gamma\to 0\) or \(\Gamma\to\infty\), we have \[\tau_{\rm ST}(\Gamma)\asymp\max\{\Gamma,\Gamma^{-1}\}, \qquad C_{\rm ST}(\Gamma)\asymp\max\{\Gamma^2,\Gamma^{-2}\}.\] Hence, the space-time LSI constant degrades only polynomially in \(\max\{\Gamma,\Gamma^{-1}\}\), identifying \(\Gamma\sim 1\) (equivalently, \(\gamma\sim\sqrt\rho\)) as the optimal friction-scaling regime.
The argument above takes \(g_t\) to be regular enough to differentiate \(\mathop{\mathrm{Ent}}(g_t)\) and \(\mathcal{C}_{\rm OT}(g_t)\) along the controlled flow. To deduce 36 for a general smooth positive density with finite right-hand side but no a priori regularity of the velocity current, we approximate it by regular controlled pairs whose action, entropy, and velocity Fisher information all converge to those of the original, and then pass to the limit.
Lemma 5 (Finite-action regularization). Let \(g\) be a smooth positive space-time density with slice mass one, and suppose that \[(\partial_t+\mathcal{L}_a)g_t=\nabla_v^{*}(g_tA_t)\] for a velocity current with finite action, entropy, and velocity Fisher information. Then there are regular positive pairs \((g^{\varepsilon},A^{\varepsilon})\), with the same slice mass and satisfying \[(\partial_t+\mathcal{L}_a)g_t^{\varepsilon} = \nabla_v^{*}(g_t^{\varepsilon}A_t^{\varepsilon}),\] to which the entropy-current identities above apply, such that, as \(\varepsilon\downarrow0\), \[g^{\varepsilon}\to g \quad\text{in } L^1(\,\mathrm{d}\bar\mu_T),\] and \[\mathop{\mathrm{Ent}}_{\bar\mu_T}(g^{\varepsilon})\to\mathop{\mathrm{Ent}}_{\bar\mu_T}(g),\qquad I_{v,T}(g^{\varepsilon})\to I_{v,T}(g),\] \[\int_0^T\int g_t^{\varepsilon}\lvert A_t^{\varepsilon}\rvert^2\,\mathrm{d}\mu\,\frac{\,\mathrm{d}t}{T} \to \int_0^T\int g_t\lvert A_t\rvert^2\,\mathrm{d}\mu\,\frac{\,\mathrm{d}t}{T}.\]
Proof. This is the regularization step from Lu2026?*Definition 2.5 and Section 7, applied to the linear finite-action continuity equation with velocity current. One first adds a small equilibrium density, then uses the convolution, cutoff, and moment truncation procedure of that section, transporting the current along the same regularization. Convexity of \((h,W)\mapsto \int \frac{\lvert W\rvert^2}{h}\,\mathrm{d}\mu\) gives the action convergence, while the entropy and Fisher-information convergences are the same as those used in Lu2026?*Section 7. The approximation preserves the controlled equations and the time-slice mass because both \(\mathcal{L}_a\) and \(\nabla_v^{*}\) are treated in divergence form in the regularization. Applying the regular inequality 36 to \((g^\varepsilon,A^\varepsilon)\) and letting \(\varepsilon\downarrow0\) then yields 36 for \(g\); the Brenier maps, moment equations, stress pairings, and lower-semicontinuity inputs invoked are precisely those of Lu2026?*Definition 2.5 and Sections 4–7, so no additional Brenier regularity argument is needed beyond the one already carried out there. ◻
Throughout this section and the next, we fix a parameter \(\Gamma>0\), set the friction parameter \(\gamma:=\Gamma\sqrt\rho\), and choose a time period \(T\ge\tau_{\rm ST}(\Gamma)\,\rho^{-1/2}\) such that the space-time LSI 36 of 1 holds with constant \(C_{\rm ST}=C_{\rm ST}(\Gamma)\). For ease of exposition, we write \[\mathcal{P}_t:=\mathrm{e}^{t(-\mathcal{L}_a+\gamma\mathcal{L}_s)}, \qquad \mathcal{P}_t^\ast:=\mathrm{e}^{t(\mathcal{L}_a+\gamma\mathcal{L}_s)}\] for the density semigroup of the underdamped Langevin dynamics 2 and its \(L^2(\mu)\)-adjoint, respectively. Both are positivity-preserving Markov contractions on \(L^p(\mu)\) for any \(1\le p\le\infty\) BakryGentilLedoux2014?.
The goal of this section is to control the bilinear pairing \(\int\mathcal{P}_T\varphi\cdot\psi\,\mathrm{d}\mu\), which, by duality with \(\psi\in L^{q'}(\mu)\), recovers the \(L^q(\mu)\)-norm of \(\mathcal{P}_T\varphi\). Here and in what follows, \(q'\) denotes the conjugate exponent of \(q\), \(1/q+1/q'=1\).
To this end, we construct a forward/backward time interpolation between \(\varphi\) and \(\psi\) whose associated path satisfies the assumptions of 1, allowing the use of the space-time LSI. The construction is in the spirit of the semigroup-interpolation technique in the Bakry–Émery \(\Gamma_2\) calculus BakryEmery1985? and the earlier work Neveu1976?, with one essential adaptation: the underdamped generator \(-\mathcal{L}_a+\gamma\mathcal{L}_s\) has a skew-adjoint part \(-\mathcal{L}_a\), so the density semigroup \(\mathcal{P}_t\) and its \(L^2(\mu)\)-adjoint \(\mathcal{P}^{*}_t\) no longer coincide as they would in the reversible case, and we therefore propagate \(\varphi\) and \(\psi\) under \(\mathcal{P}_t\) and \(\mathcal{P}_t^{*}\) simultaneously.
Specifically, given smooth strictly positive functions \(\varphi,\psi\), we define the forward/backward interpolation density \(G_s\) by \[\label{def:interpolationdensity} G_s=\frac{\varphi_s\psi_s}{Z}, \qquad Z=\int\varphi_s\psi_s\,\mathrm{d}\mu, \qquad \text{with} \qquad \varphi_s=\mathcal{P}_s\varphi,\qquad \psi_s=\mathcal{P}^\ast_{T-s}\psi.\tag{50}\] By the skew-adjointness \(\mathcal{L}_a^{*}=-\mathcal{L}_a\) and the symmetry \(\mathcal{L}_s^{*}=\mathcal{L}_s\) in \(L^2(\mu)\), \[\frac{\,\mathrm{d}}{\,\mathrm{d}s}\int\varphi_s\psi_s\,\mathrm{d}\mu =\int(-\mathcal{L}_a+\gamma\mathcal{L}_s)\varphi_s\cdot\psi_s\,\mathrm{d}\mu -\int\varphi_s\cdot(\mathcal{L}_a+\gamma\mathcal{L}_s)\psi_s\,\mathrm{d}\mu=0,\] so the normalization \(Z\) is independent of \(s\), and each \(G_s\) is indeed a probability density with respect to \(\mu\). In particular, at \(s=T\), we have \(\varphi_T=\mathcal{P}_T\varphi\) and \(\psi_T=\psi\), so that \(Z\) recovers the bilinear pairing of interest: \[Z=\int\mathcal{P}_T\varphi\cdot\psi\,\mathrm{d}\mu.\] Thus, the goal of this section is to bound \(Z\) from above, which is the main step for proving hypocoercive hypercontractivity in 5. To this end, we apply the space-time LSI to the path \(\{G_s\}_{s\in[0,T]}\) and propagate the resulting bound on \(\mathop{\mathrm{Ent}}_{\bar\mu_T}(G)\) to the endpoint entropies \(\mathop{\mathrm{Ent}}_\mu(G_0)\) and \(\mathop{\mathrm{Ent}}_\mu(G_T)\); see 2. We then translate these endpoint estimates into an upper bound on \(Z\) in 7.
We first establish the following lemma, which identifies the continuity equation satisfied by \(G_s\) and bounds the dissipation entering the space-time LSI (i.e., the right-hand side of 36 ).
Lemma 6. The forward/backward interpolation density \(G_s\) defined in 50 satisfies the velocity continuity equation \[\label{eq:interpolation-current} (\partial_s+\mathcal{L}_a)G_s=\nabla_v^{*}(G_sA_s), \qquad A_s:=-\gamma\bigl(\nabla_v\log\varphi_s-\nabla_v\log\psi_s\bigr).\tag{51}\] Moreover, the two quantities on the right-hand side of the space-time LSI 36 along \(G_s\) are both bounded by the two-sided Fisher information: \[\label{eq:interpolation-energy} \mathcal{E}_T(\varphi,\psi) :=\frac{1}{T}\int_0^T\int \left( \lvert\nabla_v\log\varphi_s\rvert^2+\lvert\nabla_v\log\psi_s\rvert^2 \right) G_s \,\mathrm{d}\mu\,\mathrm{d}s,\tag{52}\] namely, \[\label{eq:interpolation-norms} I_{v,T}(G)\le 2\,\mathcal{E}_T(\varphi,\psi), \qquad \int_0^T\lVert(\partial_s+\mathcal{L}_a)G_s\rVert_{-1,G_s}^2\,\frac{\,\mathrm{d}s}{T} \le 2\gamma^2\,\mathcal{E}_T(\varphi,\psi).\tag{53}\]
Proof. Since \(\mathcal{L}_a\) is a first-order differential operator, a direct computation gives \[\mathcal{L}_a(\varphi_s\psi_s)=(\mathcal{L}_a\varphi_s)\psi_s+\varphi_s\mathcal{L}_a\psi_s.\] Combining this with the forward and backward equations \(\partial_s\varphi_s=(-\mathcal{L}_a+\gamma\mathcal{L}_s)\varphi_s\) and \(\partial_s\psi_s=-(\mathcal{L}_a+\gamma\mathcal{L}_s)\psi_s\), the \(\mathcal{L}_a\)-contributions cancel and we obtain \[\label{eq:partial-La-product} (\partial_s+\mathcal{L}_a)(\varphi_s\psi_s) =\gamma\bigl(\psi_s\mathcal{L}_s\varphi_s-\varphi_s\mathcal{L}_s\psi_s\bigr).\tag{54}\] Using \(\mathcal{L}_s=-\nabla_v^{*}\nabla_v\) and the Leibniz rule \(\nabla_v^{*}(fF)=f\nabla_v^{*}F-\nabla_v f\cdot F\) on each of \(\psi_s\nabla_v^{*}\nabla_v\varphi_s\) and \(\varphi_s\nabla_v^{*}\nabla_v\psi_s\), the cross-terms \(\nabla_v\psi_s\cdot\nabla_v\varphi_s\) and \(\nabla_v\varphi_s\cdot\nabla_v\psi_s\) cancel by symmetry of the dot product, leaving \[\psi_s\mathcal{L}_s\varphi_s-\varphi_s\mathcal{L}_s\psi_s =-\nabla_v^{*}\bigl(\psi_s\nabla_v\varphi_s-\varphi_s\nabla_v\psi_s\bigr) =-\nabla_v^{*}\bigl(\varphi_s\psi_s\,(\nabla_v\log\varphi_s-\nabla_v\log\psi_s)\bigr),\] where the last identity factors out \(\varphi_s\psi_s = ZG_s\). Substituting into 54 and dividing by the time-independent normalization \(Z\) yields 51 .
We next bound the velocity Fisher information \(I_v(G_s)\). Recalling that \(G_s = \varphi_s \psi_s / Z\) and applying the product rule \[\nabla_v \log G_s = \nabla_v \log \varphi_s + \nabla_v \log \psi_s,\] the elementary inequality \(\lvert X + Y\rvert^2 \le 2\lvert X\rvert^2 + 2\lvert Y\rvert^2\) gives \[I_v(G_s) = \int G_s \lvert\nabla_v \log G_s\rvert^2\,\mathrm{d}\mu \le 2\int G_s\bigl(\lvert\nabla_v \log \varphi_s\rvert^2 + \lvert\nabla_v \log \psi_s\rvert^2\bigr)\,\mathrm{d}\mu.\] Time-averaging over \([0, T]\) yields the first bound in 53 .
For the second bound in 53 , we apply the variational characterization 15 with the explicit current \(A_s\) given in 51 , which yields \[\begin{align} \lVert(\partial_s + \mathcal{L}_a)G_s\rVert_{-1, G_s}^2 &\le \int G_s \lvert A_s\rvert^2\,\mathrm{d}\mu \\ &= \gamma^2 \int G_s \lvert\nabla_v \log \varphi_s - \nabla_v \log \psi_s\rvert^2\,\mathrm{d}\mu \\ &\le 2\gamma^2 \int G_s \bigl(\lvert\nabla_v \log \varphi_s\rvert^2 + \lvert\nabla_v \log \psi_s\rvert^2\bigr)\,\mathrm{d}\mu. \end{align}\] Again, time-averaging over \([0, T]\) completes the proof. ◻
We now apply the space-time logarithmic Sobolev inequality to the interpolation \(G_s\). The space-time LSI 36 , combined with 6, bounds the time-averaged entropy \(\mathop{\mathrm{Ent}}_{\bar\mu_T}(G)\) in terms of the two-sided Fisher information \(\mathcal{E}_T(\varphi, \psi)\). A direct estimate of the entropy derivative \(\frac{\,\mathrm{d}}{\,\mathrm{d}s}\mathop{\mathrm{Ent}}_\mu(G_s)\) then transfers this time-averaged bound to the two endpoints \(\mathop{\mathrm{Ent}}_\mu(G_0)\) and \(\mathop{\mathrm{Ent}}_\mu(G_T)\), which is the estimate needed for the duality argument in 5.
Proposition 2. The forward/backward interpolation \(G_s\) satisfies the time-averaged entropy bound: \[\label{eq:interpolation-average} \mathop{\mathrm{Ent}}_{\bar\mu_T}(G) \le 2C_{\mathrm{ST}}\!\left(1 + \frac{\gamma^2}{\rho}\right)\mathcal{E}_T(\varphi, \psi),\qquad{(3)}\] and the endpoint entropy bound: \[\label{eq:interpolation-endpoint} \mathop{\mathrm{Ent}}_\mu(G_0) + \mathop{\mathrm{Ent}}_\mu(G_T) \le K_T\,\mathcal{E}_T(\varphi, \psi),\qquad{(4)}\] where \[\label{eq:KT} K_T := 4C_{\mathrm{ST}}\!\left(1 + \frac{\gamma^2}{\rho}\right) + 2\gamma T.\qquad{(5)}\]
Proof. The time-averaged bound ?? follows directly from the space-time LSI of 1 applied to \(G_s\), combined with the negative-norm estimate 53 of 6.
For the endpoint bound ?? , we estimate the entropy derivative along the interpolation path. From 51 , the entropy evolves as \[\frac{\,\mathrm{d}}{\,\mathrm{d}s}\mathop{\mathrm{Ent}}_\mu(G_s) = \int G_s A_s \cdot \nabla_v \log G_s\,\mathrm{d}\mu.\] Substituting \(A_s = -\gamma(\nabla_v \log \varphi_s - \nabla_v \log \psi_s)\) and \(\nabla_v \log G_s = \nabla_v \log \varphi_s + \nabla_v \log \psi_s\), the integrand simplifies to \[G_s A_s \cdot \nabla_v \log G_s = -\gamma G_s\bigl(\lvert\nabla_v \log \varphi_s\rvert^2 - \lvert\nabla_v \log \psi_s\rvert^2\bigr),\] and hence \[\label{eq:entropy-derivative-bound} \left\lvert\frac{\,\mathrm{d}}{\,\mathrm{d}s}\mathop{\mathrm{Ent}}_\mu(G_s)\right\rvert \le \gamma \int G_s\bigl(\lvert\nabla_v \log \varphi_s\rvert^2 + \lvert\nabla_v \log \psi_s\rvert^2\bigr)\,\mathrm{d}\mu.\tag{55}\] Integrating over \([0, T]\) gives \[\int_0^T \left\lvert\frac{\,\mathrm{d}}{\,\mathrm{d}s}\mathop{\mathrm{Ent}}_\mu(G_s)\right\rvert\,\mathrm{d}s \le \gamma T\,\mathcal{E}_T(\varphi, \psi).\] For each endpoint \(a \in \{0, T\}\) and every \(r \in [0, T]\), \[\mathop{\mathrm{Ent}}_\mu(G_a) \le \mathop{\mathrm{Ent}}_\mu(G_r) + \int_0^T \left\lvert\frac{\,\mathrm{d}}{\,\mathrm{d}s}\mathop{\mathrm{Ent}}_\mu(G_s)\right\rvert\,\mathrm{d}s \le \mathop{\mathrm{Ent}}_\mu(G_r) + \gamma T\,\mathcal{E}_T(\varphi, \psi).\] Averaging in \(r\) over \([0, T]\) yields \[\mathop{\mathrm{Ent}}_\mu(G_a) \le \mathop{\mathrm{Ent}}_{\bar\mu_T}(G) + \gamma T\,\mathcal{E}_T(\varphi, \psi).\] Adding the two endpoint inequalities (\(a = 0\) and \(a = T\)) and applying ?? gives ?? , with \(K_T\) as in ?? . ◻
The remaining ingredient for 5 is a bound on \(\log Z\) in terms of the endpoint entropies and the two-sided Fisher information \(\mathcal{E}_T(\varphi, \psi)\). Once \(\varphi\) and \(\psi\) are normalized in \(L^p(\mu)\) and \(L^{q'}(\mu)\) respectively, \(\log Z\) is bounded above by a weighted combination of \(\mathop{\mathrm{Ent}}_\mu(G_0)\) and \(\mathop{\mathrm{Ent}}_\mu(G_T)\), minus a dissipation term \(\gamma T\,\mathcal{E}_T(\varphi,\psi)\). In 5, this bound will be combined with 2 to conclude \(Z \le 1\) for an appropriate choice of exponents \(p\) and \(q\) (see 3).
Lemma 7. Let \(1 < p, q < \infty\), \(\lVert\varphi\rVert_{L^p(\mu)} = 1\), and \(\lVert\psi\rVert_{L^{q'}(\mu)} = 1\), where \(q' = q/(q-1)\). Then \[\label{eq:closure} 2\log Z \le \left(\frac{2}{p} - 1\right)\mathop{\mathrm{Ent}}_\mu(G_0) + \left(1 - \frac{2}{q}\right)\mathop{\mathrm{Ent}}_\mu(G_T) - \gamma T\,\mathcal{E}_T(\varphi, \psi).\tag{56}\]
Proof. Using \(\partial_sG_s=-\mathcal{L}_aG_s+\nabla_v^{*}(G_sA_s)\) and \(\partial_s\varphi_s=-\mathcal{L}_a\varphi_s+\gamma\mathcal{L}_s\varphi_s\), we compute \[\label{eq:timederivative} \begin{align} \frac{\,\mathrm{d}}{\,\mathrm{d}s}\int G_s\log\varphi_s\,\mathrm{d}\mu &= \int(-\mathcal{L}_aG_s)\log\varphi_s\,\mathrm{d}\mu +\int\nabla_v^{*}(G_sA_s)\log\varphi_s\,\mathrm{d}\mu \\ &\quad +\int G_s\frac{-\mathcal{L}_a\varphi_s+\gamma\mathcal{L}_s\varphi_s}{\varphi_s}\,\mathrm{d}\mu. \end{align}\tag{57}\] From the chain rule: \[\mathcal{L}_a(\log f) = \frac{\mathcal{L}_af}{f}\,,\] and the skew-adjointness of \(\mathcal{L}_a\), the two \(\mathcal{L}_a\) contributions cancel: \[\begin{align} \int(-\mathcal{L}_aG_s)\log\varphi_s\,\mathrm{d}\mu + \int G_s\frac{-\mathcal{L}_a\varphi_s}{\varphi_s}\,\mathrm{d}\mu = \int G_s\frac{\mathcal{L}_a\varphi_s}{\varphi_s}\,\mathrm{d}\mu + \int G_s\frac{-\mathcal{L}_a\varphi_s}{\varphi_s}\,\mathrm{d}\mu = 0\,. \end{align}\] The identity 57 thus simplifies to \[\frac{\,\mathrm{d}}{\,\mathrm{d}s}\int G_s\log\varphi_s\,\mathrm{d}\mu = \int G_sA_s\cdot\nabla_v\log\varphi_s\,\mathrm{d}\mu +\gamma\int G_s\frac{\mathcal{L}_s\varphi_s}{\varphi_s}\,\mathrm{d}\mu.\]
Next, using \(\mathcal{L}_s=-\nabla_v^{*}\nabla_v\) and \(G_s/\varphi_s=\psi_s/Z\), we compute \[\int G_s\frac{\mathcal{L}_s\varphi_s}{\varphi_s}\,\mathrm{d}\mu =-\int\nabla_v\!\left(\frac{\psi_s}{Z}\right)\cdot\nabla_v\varphi_s\,\mathrm{d}\mu =-\int G_s\nabla_v\log\varphi_s\cdot\nabla_v\log\psi_s\,\mathrm{d}\mu,\] and since \(A_s=-\gamma(\nabla_v\log\varphi_s-\nabla_v\log\psi_s)\), we obtain \[\label{eq:gsvp} \frac{\,\mathrm{d}}{\,\mathrm{d}s}\int G_s\log\varphi_s\,\mathrm{d}\mu =-\gamma\int G_s\lvert\nabla_v\log\varphi_s\rvert^2\,\mathrm{d}\mu.\tag{58}\] The analogous computation with \(\partial_s\psi_s=-\mathcal{L}_a\psi_s-\gamma\mathcal{L}_s\psi_s\) gives \[\label{eq:gspsi} \frac{\,\mathrm{d}}{\,\mathrm{d}s}\int G_s\log\psi_s\,\mathrm{d}\mu = \gamma\int G_s\lvert\nabla_v\log\psi_s\rvert^2\,\mathrm{d}\mu.\tag{59}\] Integrating these two identities 58 59 over \([0,T]\) and then subtracting gives \[\label{eq:closure-energy-identity} \gamma T\mathcal{E}_T(\varphi,\psi) = \int G_0\log\varphi\,\mathrm{d}\mu -\int G_T\log\varphi_T\,\mathrm{d}\mu +\int G_T\log\psi\,\mathrm{d}\mu -\int G_0\log\psi_0\,\mathrm{d}\mu,\tag{60}\] due to the definition 52 of \(\mathcal{E}_T(\varphi,\psi)\) and endpoints \(\varphi_0=\varphi\) and \(\psi_T=\psi\). On the other hand, by the definition of \(G_s\), we find \[\mathop{\mathrm{Ent}}_\mu(G_s) = \int G_s\log\varphi_s\,\mathrm{d}\mu +\int G_s\log\psi_s\,\mathrm{d}\mu -\log Z,\] and then 60 can be rewritten as \[\label{eq:energy95zgvp} \gamma T\mathcal{E}_T(\varphi,\psi) = 2\int G_0\log\varphi\,\mathrm{d}\mu +2\int G_T\log\psi\,\mathrm{d}\mu -\mathop{\mathrm{Ent}}_\mu(G_0)-\mathop{\mathrm{Ent}}_\mu(G_T)-2\log Z.\tag{61}\]
Finally, we recall the variational inequality for the entropy: for any probability density \(g\) with respect to \(\mu\) and any measurable function \(\phi\), \[\label{eq:variational} \int \phi\, g\,\,\mathrm{d}\mu \le \mathop{\mathrm{Ent}}_\mu(g) + \log\int \mathrm{e}^{\phi}\,\,\mathrm{d}\mu.\tag{62}\] Since \(\lVert\varphi\rVert_{L^p(\mu)} = \lVert\psi\rVert_{L^{q'}(\mu)} = 1\), applying 62 to \(\phi = p\log\varphi\) and \(g = G_0\) gives \[p\int G_0\log\varphi\,\,\mathrm{d}\mu \le \mathop{\mathrm{Ent}}_\mu(G_0) + \log\int \varphi^p\,\,\mathrm{d}\mu = \mathop{\mathrm{Ent}}_\mu(G_0),\] which yields, by dividing by \(p\), \[\int G_0\log\varphi\,\,\mathrm{d}\mu \le \frac{1}{p}\mathop{\mathrm{Ent}}_\mu(G_0).\] The analogous application to \(\phi = q'\log\psi\) and \(g = G_T\), combined with \(\lVert\psi\rVert_{L^{q'}(\mu)}^{q'} = 1\), gives \[\int G_T\log\psi\,\,\mathrm{d}\mu \le \frac{1}{q'}\mathop{\mathrm{Ent}}_\mu(G_T).\] Substituting both bounds into 61 and using \(1/q' = 1 - 1/q\) gives 56 . ◻
We now exploit the endpoint entropy bound ?? and the normalization bound 56 to establish the hypercontractivity property for \(\mathcal{P}_T\). We first prove the contraction at the critical pair \((p_{\rm c},q_{\rm c})\), at which the two bounds ?? and 56 balance exactly and force the normalization constant to satisfy \(Z\le1\). The remaining \((p,q)\) exponents then follow by Riesz–Thorin interpolation with the \(L^1\to L^1\) and \(L^\infty\to L^\infty\) Markov contraction bounds for \(\mathcal{P}_T\).
Proposition 3 (Critical hypercontractive estimate). Let \(K_T\) be the constant ?? for the endpoint entropy estimate ?? , and introduce the constant \[\eta_T=\frac{\gamma T}{K_T}.\] Then \(0<\eta_T<1/2\), and, with \[\label{eq:criexponents} p_{\rm c}=\frac{2}{1+\eta_T}, \qquad q_{\rm c}=\frac{2}{1-\eta_T},\qquad{(6)}\] one has \[\label{eq:critest} \lVert\mathcal{P}_Tf\rVert_{L^{q_{\rm c}}(\mu)} \le \lVert f\rVert_{L^{p_{\rm c}}(\mu)}\qquad{(7)}\] for every \(f\in L^{p_{\rm c}}(\mu)\).
Proof. The inequality \(0<\eta_T<1/2\) is immediate from the definition \(K_T=4C_{\mathrm{ST}}(1+\gamma^2/\rho)+2\gamma T\), since \(4C_{\mathrm{ST}}(1+\gamma^2/\rho)>0\). Since the mapping \(\mathcal{P}_T\) is positive, i.e., \[\lvert\mathcal{P}_Tf\rvert\le \mathcal{P}_T\lvert f\rvert,\] it suffices, after a standard density argument, to prove the estimate \[\label{eq:target} \lVert\mathcal{P}_T \varphi\rVert_{L^{q_{\rm c}}(\mu)} \le \lVert\varphi\rVert_{L^{p_{\rm c}}(\mu)},\tag{63}\] for smooth strictly positive \(\varphi\) with \(\lVert\varphi\rVert_{L^{p_{\rm c}}(\mu)}=1\). Indeed, if 63 holds, replacing a smooth non-negative \(\varphi\) by \(\varphi+\varepsilon\) and letting \(\varepsilon\downarrow0\) gives the same bound for smooth non-negative \(\varphi\). If \(f_n\to f\) in \(L^{p_{\rm c}}(\mu)\) with \(f_n\) smooth and non-negative, the preceding estimate shows that \(\mathcal{P}_Tf_n\) is Cauchy in \(L^{q_{\rm c}}(\mu)\); its limit agrees with the \(L^{p_{\rm c}}\)-semigroup value \(\mathcal{P}_Tf\) by the \(L^{p_{\rm c}}\) Markov contraction and a subsequence argument. Thus the estimate 63 extends to all non-negative \(f\in L^{p_{\rm c}}(\mu)\). The general signed case follows from positivity, since \(\lvert\mathcal{P}_Tf\rvert\le\mathcal{P}_T\lvert f\rvert\).
We now prove 63 . Let \(q_{\rm c}'=q_{\rm c}/(q_{\rm c}-1)\) be the conjugate exponent of \(q_{\rm c}\). By duality in \(L^{q_{\rm c}}(\mu)\) and density in \(L^{q_{\rm c}'}(\mu)\), the hypercontraction 63 is equivalent to the assertion that, for every smooth strictly positive \(\psi\) with \(\lVert\psi\rVert_{L^{q_{\rm c}'}(\mu)}=1\), \[Z=\int \mathcal{P}_T\varphi\cdot\psi\,\mathrm{d}\mu\le1.\] For the critical exponents ?? , applying 7 with \(p=p_{\rm c}\) and \(q=q_{\rm c}\), and then using the endpoint entropy estimate ?? , gives \[2\log Z \le \eta_T\bigl[\mathop{\mathrm{Ent}}_\mu(G_0)+\mathop{\mathrm{Ent}}_\mu(G_T)\bigr] -\gamma T\mathcal{E}_T(\varphi,\psi) \le \eta_TK_T\mathcal{E}_T(\varphi,\psi) -\gamma T\mathcal{E}_T(\varphi,\psi) =0,\] where the last equality follows from the choice \(\eta_T=\gamma T/K_T\). Hence \(Z\le1\) holds as desired. ◻
Theorem 2 (One-step hypocoercive hypercontractivity). With \(\eta_T\) as in 3, set \[\label{def:alpha} \alpha_T=\frac{1+\eta_T}{1-\eta_T} > 1.\tag{64}\] Then for any \(p>1\) and \(f\in L^p(\mu)\), \[\label{eq:one-slab} \lVert\mathcal{P}_Tf\rVert_{L^{1+\alpha_T(p-1)}(\mu)} \le \lVert f\rVert_{L^p(\mu)}.\tag{65}\] Consequently, \[\label{eq:one-slab2} \lVert\mathcal{P}_Tf\rVert_{L^q(\mu)} \le \lVert f\rVert_{L^p(\mu)}\tag{66}\] whenever \(1\le q\le 1+\alpha_T(p-1)\).
Proof. The proof is based on the critical estimate of 3 together with the \(L^1\to L^1\) and \(L^\infty\to L^\infty\) Markov contraction bounds for \(\mathcal{P}_T\).
First let \(1<p\le p_{\rm c}\). By the Markov contraction property and 3, we have \[\lVert\mathcal{P}_Tf\rVert_{L^1(\mu)} \le \lVert f\rVert_{L^1(\mu)} \qquad\text{and}\qquad \lVert\mathcal{P}_Tf\rVert_{L^{q_{\rm c}}(\mu)} \le \lVert f\rVert_{L^{p_{\rm c}}(\mu)}.\] Hence, the Riesz–Thorin theorem applied to \(\mathcal{P}_T\) with the endpoints \(L^1 \to L^1\) and \(L^{p_{\rm c}} \to L^{q_{\rm c}}\) gives, for \(0\le\beta\le1\), \[\label{eq:interpolation} \lVert\mathcal{P}_Tf\rVert_{L^{q_\beta}(\mu)} \le \lVert f\rVert_{L^{p_\beta}(\mu)}, \qquad \frac{1}{p_\beta}=1-\beta+\frac{\beta}{p_{\rm c}}, \quad \frac{1}{q_\beta}=1-\beta+\frac{\beta}{q_{\rm c}}.\tag{67}\] Since \(p_{\rm c}>1\), the map \(\beta\mapsto p_\beta\) is increasing from \(1\) to \(p_{\rm c}\). Thus, for a given \(p\in(1,p_{\rm c}]\), there is a unique \(\beta = \tfrac{2(p-1)}{p(1-\eta_T)}\in(0,1]\) such that \(p=p_\beta\). Moreover, using ?? , or equivalently \(1/p_{\rm c}=(1+\eta_T)/2\) and \(1/q_{\rm c}=(1-\eta_T)/2\), one can directly compute the interpolated exponents as \[p_\beta-1 = \frac{\beta(1-\eta_T)}{2-\beta(1-\eta_T)}, \qquad q_\beta-1 = \frac{\beta(1+\eta_T)}{2-\beta(1+\eta_T)}.\] With the choice of 64 , we then have \[\alpha_T(p_\beta-1) = \frac{\beta(1+\eta_T)}{2-\beta(1-\eta_T)} \le \frac{\beta(1+\eta_T)}{2-\beta(1+\eta_T)} = q_\beta-1.\] Therefore, at \(p=p_\beta\), by the interpolation 67 and the monotonicity of \(\lVert\cdot\rVert_{L^r(\mu)}\), \[\lVert\mathcal{P}_Tf\rVert_{L^{1+\alpha_T(p-1)}(\mu)} \le \lVert\mathcal{P}_Tf\rVert_{L^{q_\beta}(\mu)} \le \lVert f\rVert_{L^p(\mu)}.\] This proves 65 for \(1<p\le p_{\rm c}\).
We now consider the case \(p\ge p_{\rm c}\); the analysis is similar. Interpolating the critical estimate ?? with the \(L^\infty\to L^\infty\) contraction gives, for \(0<\beta\le1\), \[\lVert\mathcal{P}_Tf\rVert_{L^{q_\beta}(\mu)} \le \lVert f\rVert_{L^{p_\beta}(\mu)}, \qquad \frac{1}{p_\beta}=\frac{\beta}{p_{\rm c}}, \quad \frac{1}{q_\beta}=\frac{\beta}{q_{\rm c}}.\] Now \(p_\beta\) runs from \(p_{\rm c}\) to \(\infty\) as \(\beta\) runs from \(1\) down to \(0\), so we choose \(\beta\) with \(p=p_\beta\). Since \(q_{\rm c}/p_{\rm c}=\alpha_T\), we have \(q_\beta=\alpha_T p_\beta\), and hence \[q_\beta\ge 1+\alpha_T(p_\beta-1).\] Monotonicity of \(L^q(\mu)\) norms again gives 65 .
The final claim 66 follows directly from 65 and the norm monotonicity: if \(1\le q\le 1+\alpha_T(p-1)\), then \[\lVert\mathcal{P}_Tf\rVert_{L^q(\mu)} \le \lVert\mathcal{P}_Tf\rVert_{L^{1+\alpha_T(p-1)}(\mu)} \le \lVert f\rVert_{L^p(\mu)}. \qedhere\] ◻
Corollary 1 (Iterated hypercontractivity). Let \(p>1\) and, for \(n=0,1,2,\ldots\), set \[\label{def:pn} p_n=1+\alpha_T^n(p-1).\tag{68}\] Then, for any \(n \ge 0\) and \(f\in L^p(\mu)\), \[\lVert\mathcal{P}_{nT}f\rVert_{L^{p_n}(\mu)} \le \lVert f\rVert_{L^p(\mu)}.\] More generally, if \(t\ge0\) and \(n=\lfloor t/T\rfloor\), then \[\label{eq:targecontract} \lVert\mathcal{P}_tf\rVert_{L^q(\mu)} \le \lVert f\rVert_{L^p(\mu)}\tag{69}\] for every \(1\le q\le p_n\).
Proof. We consider the recursion \[p_{k+1}=1+\alpha_T(p_k-1)\] from the definition 68 . Applying 65 with input exponent \(p_k\) to the function \(\mathcal{P}_{kT}f\) gives \[\lVert\mathcal{P}_{(k+1)T}f\rVert_{L^{p_{k+1}}(\mu)} = \lVert\mathcal{P}_T(\mathcal{P}_{kT}f)\rVert_{L^{p_{k+1}}(\mu)} \le \lVert\mathcal{P}_{kT}f\rVert_{L^{p_k}(\mu)}.\] Iterating this inequality for \(k=0,\ldots,n-1\) yields \[\lVert\mathcal{P}_{nT}f\rVert_{L^{p_n}(\mu)} \le \lVert f\rVert_{L^p(\mu)}.\] For arbitrary \(t\ge0\), write \(t=nT+r\) with \(n=\lfloor t/T\rfloor\) and \(0\le r<T\). The semigroup property gives \(\mathcal{P}_t=\mathcal{P}_r\mathcal{P}_{nT}\), and the Markov contraction property gives \[\lVert\mathcal{P}_tf\rVert_{L^{p_n}(\mu)} = \lVert\mathcal{P}_r(\mathcal{P}_{nT}f)\rVert_{L^{p_n}(\mu)} \le \lVert\mathcal{P}_{nT}f\rVert_{L^{p_n}(\mu)} \le \lVert f\rVert_{L^p(\mu)}.\] Finally, since \(L^r(\mu)\) norms are monotone increasing in \(r\), the same estimate holds for any output exponent \(1\le q\le p_n\), i.e., the desired 69 holds. ◻
Remark 4 (Kinetic scaling). Fix \(\Gamma > 0\) and choose \(\tau \ge \tau_{\mathrm{ST}}(\Gamma)\). Set \[\gamma = \Gamma\sqrt{\rho}, \qquad T = \tau\rho^{-1/2}.\] This is the kinetic time scale on which 2 applies. Under this scaling, the two dimensionless combinations entering \(K_T\) in ?? are \[\gamma T = \Gamma\tau, \qquad \frac{\gamma^2}{\rho} = \Gamma^2,\] and hence the constants \[K_T = K_{\Gamma,\tau} := 4C_{\mathrm{ST}}(\Gamma)(1 + \Gamma^2) + 2\Gamma\tau, \qquad \eta_T = \eta_{\Gamma,\tau} := \frac{\Gamma\tau}{K_{\Gamma,\tau}},\] as well as \[\alpha_T = \alpha_{\Gamma,\tau} := \frac{1 + \eta_{\Gamma,\tau}}{1 - \eta_{\Gamma,\tau}} > 1, \qquad \lambda_{\Gamma,\tau} := \frac{\log\alpha_{\Gamma,\tau}}{\tau},\] depend only on \(\Gamma\) and \(\tau\), and in particular are independent of \(\rho\).
Consequently, the LSI constant \(\rho\) enters the iteration only through the kinetic clock \(\sqrt{\rho}\,t\). Indeed, for \(t \ge 0\) and \(n = \lfloor t/T \rfloor = \lfloor\sqrt{\rho}\,t/\tau\rfloor\), 1 gives the admissible output exponent \[q_*(t) = 1 + (p - 1)\alpha_{\Gamma,\tau}^{\,n}.\] Since \(x - 1 \le \lfloor x \rfloor \le x\), this implies \[(p - 1)\alpha_{\Gamma,\tau}^{-1} \exp\!\left(\lambda_{\Gamma,\tau}\sqrt{\rho}\,t\right) \le q_*(t) - 1 \le (p - 1)\exp\!\left(\lambda_{\Gamma,\tau}\sqrt{\rho}\,t\right),\] so the admissible \(L^q\) exponent grows exponentially at the rate \(\lambda_{\Gamma,\tau}\sqrt{\rho}\) (for comparison, in the reversible case 1 , the admissible exponent \(q\) grows as \(\mathrm{e}^{2\rho t}\)). The integrability gain therefore occurs on the ballistic time scale \(t \sim \rho^{-1/2}\), consistent with the sharp hypocoercive entropy-decay rate of Lu2026?, and improving on the diffusive scale \(\rho^{-1}\) in the overdamped Langevin case suggested by the LSI 3 .
We now translate the above hypocoercive hypercontractivity into statements about the Rényi divergence decay. Recall that, for \(r>1\), the Rényi divergence of order \(r\) of a probability measure \(\nu\) with respect to \(\mu\) is \[\operatorname R_r(\nu\mid\mu) = \frac{1}{r-1}\log\int\left(\frac{\,\mathrm{d}\nu}{\,\mathrm{d}\mu}\right)^r\,\mathrm{d}\mu = \frac{r}{r-1}\log\lVert\,\mathrm{d}\nu/\,\mathrm{d}\mu\rVert_{L^r(\mu)}\] if \(\nu\ll\mu\), and \(+\infty\) otherwise. The next corollary is essentially a reformulation of our hypercontractivity result in 1, simultaneously yielding the order-improvement (hypercontractive) bound and the long-time decay estimate for \(\operatorname R_r(\nu_t\mid\mu)\).
Corollary 2 (Hypocoercive Rényi hypercontractivity and decay). Let \(h\) be a probability density, set \(\nu_t=(\mathcal{P}_t h)\,\mu\), and let \(n=\lfloor t/T\rfloor\) with \(p_n=1+\alpha_T^n(p-1)\).
(Rényi hypercontractivity.) For every \(p>1\) and every \(q\) with \(1<p\le q\le p_n\), \[\label{eq:renyi} \operatorname R_q(\nu_t\mid\mu) \le \frac{q(p-1)}{p(q-1)} \operatorname R_p(\nu_0\mid\mu).\tag{70}\]
(Rényi divergence decay.) For every \(q>1\) and every \(t\ge0\), \[\label{eq:renyi-long-time} \operatorname R_q(\nu_t\mid\mu) \le q\alpha_T \mathrm{e}^{-(\log\alpha_T)t/T} \operatorname R_q(\nu_0\mid\mu).\tag{71}\] In particular, if \(T=\tau\rho^{-1/2}\) with \(\tau\ge\tau_{\rm ST}(\Gamma)\), then \(\alpha_T=\alpha_{\Gamma,\tau}\) depends only on \(\Gamma\) and \(\tau\), and \[\label{eq:renyi-kinetic-decay} \operatorname R_q(\nu_t\mid\mu) \le q\alpha_{\Gamma,\tau}\, \mathrm{e}^{-\lambda_{\Gamma,\tau}\sqrt\rho\,t} \operatorname R_q(\nu_0\mid\mu), \qquad \lambda_{\Gamma,\tau}=\frac{\log\alpha_{\Gamma,\tau}}{\tau}.\tag{72}\]
Proof. (i). By 1 and monotonicity of \(L^q\) norms on the probability space, \[\lVert\mathcal{P}_t h\rVert_{L^q(\mu)}\le\lVert h\rVert_{L^p(\mu)}\,, \quad \text{whenever q\le p_n}.\] Using \(\operatorname R_r(\nu\mid\mu)=\tfrac r{r-1}\log\lVert\,\mathrm{d}\nu/\,\mathrm{d}\mu\rVert_{L^r(\mu)}\), we obtain 70 : \[\operatorname R_q(\nu_t\mid\mu) \le \frac{q}{q-1}\log\lVert h\rVert_{L^p(\mu)} = \frac{q(p-1)}{p(q-1)} \operatorname R_p(\nu_0\mid\mu).\]
(ii). Given \(q>1\), choose \(p=1+\alpha_T^{-n}(q-1)\), so that \(q=1+\alpha_T^n(p-1)\le p_n\). Applying (i) at this \((p,q)\) gives \[\operatorname R_q(\nu_t\mid\mu) \le \frac{q\alpha_T^{-n}}{1+(q-1)\alpha_T^{-n}} \operatorname R_p(\nu_0\mid\mu).\] Since \(p\le q\), monotonicity of Rényi divergence in the order gives \(\operatorname R_p(\nu_0\mid\mu)\le\operatorname R_q(\nu_0\mid\mu)\). Moreover, \(\alpha_T^{-n}\le\alpha_T\mathrm{e}^{-(\log\alpha_T)t/T}\), proving 71 . Under the kinetic scaling \(T=\tau\rho^{-1/2}\), the identities \(\gamma T=\Gamma\tau\) and \(\gamma^2/\rho=\Gamma^2\) make \(\alpha_T\) a function only of \(\Gamma\) and \(\tau\), yielding 72 . ◻
Remark 5 (Rate degradation at extreme friction). We trace the asymptotic behavior of the hypocoercive Rényi-decay rate \(\lambda_{\Gamma,\tau}\sqrt\rho\) at extreme friction, taking \(\tau\) at its minimal admissible value \(\tau=\tau_{\rm ST}(\Gamma)\). With \(M:=\max\{\Gamma,\Gamma^{-1}\}\), substituting the asymptotics \(\tau_{\rm ST}(\Gamma)\asymp M\) and \(C_{\rm ST}(\Gamma)\asymp M^2\) from 3 into the constants \(K_T=4C_{\rm ST}(\Gamma)(1+\Gamma^2)+2\Gamma\tau\) and \(\eta_T=\Gamma\tau/K_T\) of 4 yields \[K_T \asymp \begin{cases} M^4 & \text{as }\Gamma\to\infty\;(M=\Gamma),\quad\text{with }1+\Gamma^2\asymp M^2,\;\Gamma\tau\asymp M^2,\\ M^2 & \text{as }\Gamma\to 0\;(M=\Gamma^{-1}),\quad\text{with }1+\Gamma^2\asymp 1,\;\Gamma\tau\asymp 1. \end{cases}\] While \(K_T\) scales asymmetrically in the two extremes, in both cases \[\eta_T=\frac{\Gamma\tau}{K_T}\asymp M^{-2}, \qquad \log\alpha_T=\log\frac{1+\eta_T}{1-\eta_T}=2\eta_T+O(\eta_T^3)\asymp M^{-2}, \qquad \lambda_{\Gamma,\tau_{\rm ST}(\Gamma)} =\frac{\log\alpha_T}{\tau_{\rm ST}(\Gamma)} \asymp M^{-3}.\] The hypocoercive Rényi-decay rate \(\lambda_{\Gamma,\tau}\sqrt\rho\) therefore degrades cubically in \(\max\{\Gamma,\Gamma^{-1}\}\) at extreme friction, identifying \(\Gamma\sim 1\) as the optimal friction-scaling regime for the decay rate as well; the minimum of \(\tau_{\rm ST}\) and \(C_{\rm ST}\) at \(\Gamma=1\) for the space-time LSI (see 3) thus translates into the qualitatively optimal Rényi-decay rate.