Riesz-Kernel Stein Variational Gradient Descent:
Renormalized Entropy and Long-Time Particle Limits

Trevor Teolis
Rice University

,

Maarten V. de Hoop
Rice University


Abstract

Stein variational gradient descent (SVGD) transports interacting particles toward a target distribution through deterministic kernelized dynamics. Singular Riesz kernels are attractive because they can provide quantitative population-level convergence, but at the finite-particle level the corresponding Stein energy has infinite self-interaction. We study periodic Riesz SVGD with self-interaction removed and prove a many-particle, long-time sampling theorem. Throughout the range in which the singular Stein energy is locally integrable, under a uniform bound on the initial relative entropy per particle, the time-averaged empirical-measure law converges weakly to the point mass \(\delta_\pi\) at the target as the particle number and any diverging averaging horizon tend to infinity. We also show that the empirical-measure laws induced by invariant particle laws of finite relative entropy converge weakly to \(\delta_\pi\), without a uniform entropy bound. Below the logarithmic singularity threshold, we obtain an explicit algebraic finite-particle error bound. These results extend the joint-entropy approach for smooth-kernel SVGD to singular interactions.

1 Introduction↩︎

Sampling from a target probability density \(\pi\), known only up to normalization, is a basic problem in Bayesian inference and computational statistics. At the population level, overdamped Langevin dynamics realizes the Wasserstein gradient flow of relative entropy through independent score-driven diffusions. The corresponding deterministic continuity equation has velocity \(-\nabla\log(\rho/\pi)\), which depends on the unknown evolving density \(\rho\) and therefore cannot be evaluated directly on an empirical measure.

Stein variational gradient descent (SVGD), introduced by [1], takes a different route. It projects this velocity onto a kernel space and uses integration by parts to remove \(\nabla\log\rho\). Write \(\pi=Z^{-1}e^{-V}\) on a \(d\)-dimensional state space, where \(Z\) is the normalizing constant, and let \(b_\pi:=\nabla\log\pi=-\nabla V\). For a smooth symmetric positive-definite kernel \(k\), the standard continuous-time particle system is \[\label{eq:standard-svgd-intro} \dot{X}_i(t) =\frac{1}{N}\sum_{j=1}^N \left\{\nabla_1 k(X_j(t),X_i(t)) +k(X_j(t),X_i(t))b_\pi(X_j(t))\right\}.\tag{1}\] Here \(\nabla_1\) differentiates the first kernel argument, and the sum includes self-interaction. The projected velocity depends on the target only through its score, so the normalizing constant is not needed. The score-weighted term transports the ensemble toward the target, while the kernel derivative supplies a repulsive interaction. Thus SVGD moves the empirical measure \(\mu_t^N:=N^{-1}\sum_{i=1}^N\delta_{X_i(t)}\) without estimating its density or choosing a parametric approximating family. It trades independent stochastic trajectories for a coordinated deterministic approximation of the target law.

Our concern is the long-time behavior of this coordinated particle transport. Unlike Langevin dynamics, the finite system has no diffusive mechanism enforcing ergodicity, and the particles cannot be analyzed as independent samples. The resulting question is how to obtain a many-particle, long-time sampling statement from deterministic interacting dynamics whose fixed-\(N\) trajectories need not converge to the target. The quantity connecting the dynamics to convergence is the kernel Stein discrepancy. If \(\mathcal{H}_k\) is the reproducing-kernel Hilbert space of \(k\), then \[\label{eq:standard-ksd-intro} \begin{align} \operatorname{KSD}_k^2(\mu\mid\pi) &:= \left\|\int \bigl\{\nabla_x k(x,\cdot)+k(x,\cdot)b_\pi(x)\bigr\} \,\,\mathrm d\mu(x)\right\|_{\mathcal{H}_k^d}^{2}. \end{align}\tag{2}\] It vanishes at the target and measures the size of the kernel-projected transport velocity. Along the population SVGD flow, it is exactly the rate at which relative entropy decreases. For \(P\ll Q\), write \[\label{eq:relative-entropy} \operatorname{Ent}(P\mid Q):=\int\log(\,\mathrm dP/\,\mathrm dQ)\,\,\mathrm dP.\tag{3}\] If \(\rho_t^k\) is the smooth-kernel population SVGD flow, then \[\label{eq:standard-population-dissipation-intro} \frac{\,\mathrm d}{\,\mathrm dt}\operatorname{Ent}(\rho_t^k\mid\pi) =-\operatorname{KSD}_k^2(\rho_t^k\mid\pi).\tag{4}\] The population gradient-flow theory was developed by [2] and [3], while [4] established the mean-field limit. For smooth kernels, [5] subsequently obtained longer-time stability estimates for particle systems initialized near the target. The KSD is also used by [6] to test whether a finite sample is consistent with a specified target distribution and by [7] to quantify how well samples approximate the target.

This perspective has led both to new analyses and to modifications of the algorithm. [8] introduced matrix-valued kernels that incorporate problem geometry through preconditioning, while [9] used random batches to reduce the cost of the all-to-all interaction. Kernel design can also change the underlying dissipation: [10] recast SVGD as a kernelized Wasserstein gradient flow of the chi-squared divergence and introduced Laplacian Adjusted Wasserstein Gradient Descent (LAWGD), whose target-adapted kernel represents the inverse Langevin generator and yields strong population convergence under a Poincaré inequality. The required spectral decomposition, however, makes LAWGD a different computational regime from the translation-invariant scalar kernels considered here.

For smooth kernels, the KSD can be evaluated directly on empirical measures, including the self-interaction terms. [11] established non-asymptotic descent estimates for population SVGD and quantitative bounds on its time-averaged KSD, [12] developed a variational description of the many-particle and long-time regimes, and [13] obtained the first explicit finite-particle rate for driving the KSD toward zero. Most closely related to our argument, [14] introduced the normalized joint-law entropy \[H_N^{\mathrm{sm}}(t) :=\frac{1}{N} \operatorname{Ent}\!\left(P_t^{N,\mathrm{sm}}\mid\pi^{\otimes N}\right),\] where \(P_t^{N,\mathrm{sm}}\) is the joint law of the smooth-kernel particle system. Their entropy calculation yields \[\frac{1}{T}\int_0^T \mathbb{E}\!\left[\operatorname{KSD}_k^2(\mu_t^N\mid\pi)\right]\,\,\mathrm dt \le \frac{H_N^{\mathrm{sm}}(0)}{T}+\frac{C_k}{N}.\] Consequently, for uniformly bounded initial normalized entropy, \[\lim_{T\to\infty}\limsup_{N\to\infty} \frac{1}{T}\int_0^T \mathbb{E}\!\left[\operatorname{KSD}_k^2(\mu_t^N\mid\pi)\right]\,\,\mathrm dt =0.\] Under their additional assumptions, this also yields convergence of time-averaged particle marginals. More recently, [15] proved propagation-of-chaos bounds, uniform over the averaging horizon, for time-averaged smooth-kernel SVGD; their uniform-in-physical-time parametric rates concern special finite-rank systems.

Smooth-kernel results show that time-averaged many-particle convergence can be obtained despite the absence of fixed-\(N\) ergodicity. At the population level, general results for standard smooth scalar kernels give qualitative last-iterate convergence and quantitative time-averaged or best-iterate KSD bounds, but no general quantitative rate is known for the last iterate in a strong topology [11], [16]. This contrast motivates kernels with finite-order Fourier decay.

On \(\mathbb{T}^d\), the mean-zero periodic Riesz kernel \(\kappa_a\) is characterized by \[\label{eq:kappa-fourier} \widehat\kappa_a(0)=0, \qquad \widehat\kappa_a(\ell)=(2\pi|\ell|)^{-2a} \quad(\ell\in\mathbb{Z}^d\setminus\{0\}),\tag{5}\] or equivalently by \((-\Delta)^a\kappa_a=\delta_0-1\). Its Fourier multiplier makes the associated Stein discrepancy coercive in a negative Sobolev norm, while the Riesz force provides short-range repulsion. [16] show that this reduced smoothing can yield quantitative population last-iterate convergence. Under their Sobolev regularity and small-initial-entropy assumptions, for every \(a>1\) and a regularity index \(\gamma>\max\{d/2,a-1\}\), \[\label{eq:population-polynomial-rate} \operatorname{Ent}(\rho_t\mid\pi) \le \left( \operatorname{Ent}(\rho_0\mid\pi)^{-(a-1)/\gamma}+\frac{t}{C} \right)^{-\gamma/(a-1)},\tag{6}\] together with corresponding decay in strong Sobolev norms; in particular, the representative entropy and squared \(L^2\) rates are \(t^{-\gamma/(a-1)}\). At the inverse-Laplacian endpoint \(a=1\), [17] establish global exponential convergence for a class of kernels with quadratic Fourier decay, while [16] recover the corresponding Coulomb result on the torus. In a different singular limit, [18] show that concentrating a regular kernel leads to a local quadratic-mobility flow.

The same finite-order Fourier structure creates the particle-level obstruction. Applying the Stein operator in both variables produces a scalar energy kernel with an infinite diagonal, whereas the singular particle dynamics excludes self-interaction. Thus the nonnegative population dissipation cannot simply be evaluated on an empirical measure. The issue is not singularity by itself: reduced smoothing strengthens population coercivity while simultaneously exposing the diagonal at finite \(N\).

1.1 The particle system and its Stein energy↩︎

We now take the state space to be the flat torus \(\mathbb{T}^d\) and specialize to \(k(y,x)=\kappa_a(y-x)\). The first-order Stein operator is \[\label{eq:A-intro} \mathcal{A}_{\pi,y}k(y,x) :=\nabla_y k(y,x)+b_\pi(y)k(y,x).\tag{7}\] The Riesz interaction is not smooth on the diagonal, and its self-interaction is not consistently defined across the range considered below. We therefore study the self-interaction-free system \[\label{eq:particle-system} \dot{X}_i = \frac{1}{N}\sum_{j\ne i}\mathcal{A}_{\pi,X_j}k(X_j,X_i) = \frac{1}{N}\sum_{j\ne i} \left\{\nabla\kappa_a(X_j-X_i) +\kappa_a(X_j-X_i)b_\pi(X_j)\right\}.\tag{8}\] Write \(\boldsymbol{X}(t):=(X_1(t),\ldots,X_N(t))\). For a configuration \(\boldsymbol{x}=(x_1,\ldots,x_N)\), let \[\mu_{\boldsymbol{x}}^N:=\frac{1}{N}\sum_{i=1}^N\delta_{x_i}\] be its empirical measure, and let \[\mathcal{D}_N:=\{\boldsymbol{x}\in(\mathbb{T}^d)^N:x_i\ne x_j\text{ for }i\ne j\}\] be the collision-free configuration space. The first term is the repulsive kernel force; the second transports that interaction through the target score. With \(\kappa_a\) defined by 5 , we work in the finite-energy regime \[\label{eq:parameter-range-intro} 1<a<1+\frac{d}{2}, \qquad \sigma:=d+2-2a\in(0,d).\tag{9}\]

To connect the particle velocity with entropy dissipation, we use divergence relative to the target measure. For a vector field \(u\), this is \[\mathcal{S}_{\pi,x}u :=\frac{1}{\pi(x)}\nabla_x\cdot\bigl(\pi(x)u\bigr) =\nabla_x\cdot u+b_\pi(x)\cdot u.\] Applying this operator to the velocity generated at \(x\) by a source point \(y\) defines the scalar Stein kernel \[\label{eq:G-intro} G_\pi(x,y):=\mathcal{S}_{\pi,x}\mathcal{A}_{\pi,y}k(y,x).\tag{10}\] Thus \(G_\pi(x,y)\) measures the contribution of the interaction between \(x\) and \(y\) to compression of the flow relative to the target. For \(x\ne y\), a direct calculation gives \[\label{eq:G-expansion} \begin{align} G_\pi(x,y) ={}&-\Delta\kappa_a(y-x) +\bigl(b_\pi(x)-b_\pi(y)\bigr)\cdot\nabla\kappa_a(y-x)\\ &+b_\pi(x)\cdot b_\pi(y)\,\kappa_a(y-x). \end{align}\tag{11}\]

At the population level, the quadratic form generated by \(G_\pi\) is nonnegative. Indeed, for a smooth probability measure \(\mu\), \[\label{eq:F-cont} \begin{align} \operatorname{KSD}_{\kappa_a}^2(\mu\mid\pi) &:= \left\| \int_{\mathbb{T}^d}\mathcal{A}_{\pi,y}k(y,\cdot)\,\,\mathrm d\mu(y) \right\|_{\mathcal{H}_{\kappa_a}^d}^{2}\\ &=\iint_{\mathbb{T}^d\times\mathbb{T}^d}G_\pi(x,y) \,\,\mathrm d(\mu-\pi)(x)\,\,\mathrm d(\mu-\pi)(y) =:2\mathcal{F}^\pi(\mu). \end{align}\tag{12}\] We call \(\mathcal{F}^\pi\) the Stein energy. The space \(\mathcal{H}_{\kappa_a}\) is the native Fourier space of \(\kappa_a\), defined in 22 . Along the population SVGD flow, this energy is exactly the entropy dissipation: \[\label{eq:population-dissipation-intro} \frac{\,\mathrm d}{\,\mathrm dt}\operatorname{Ent}(\rho_t\mid\pi) =-\operatorname{KSD}_{\kappa_a}^2(\rho_t\mid\pi) =-2\mathcal{F}^\pi(\rho_t).\tag{13}\] Identity 12 is proved for smooth measures in 5, and the population entropy-dissipation identity 13 is proved in 11. Both arguments use Fourier duality and therefore do not require pointwise reproduction for the singular kernel.

The role of \(G_\pi\) at the particle level is analogous. When the Liouville equation for the joint particle law is tested against relative entropy with respect to \(\pi^{\otimes N}\), each ordered pair contributes \(G_\pi(x_i,x_j)\). The leading term \(-\Delta\kappa_a(y-x)\) in 11 diverges on the diagonal, so the full empirical Stein energy is undefined. Since the particle system omits self-interaction, its entropy production instead involves the off-diagonal empirical Stein energy \[\label{eq:F-discrete} \mathcal{F}_N^\pi(\boldsymbol{x}) :=\frac{1}{2N^2}\sum_{i\ne j}G_\pi(x_i,x_j).\tag{14}\]

1.2 The singular entropy argument↩︎

The smooth-kernel entropy method suggests applying the same normalized joint-law entropy to the off-diagonal Riesz system. Let \(P_t^N\) be its joint particle law and write \[H_N(t):=\frac{1}{N}\operatorname{Ent}(P_t^N\mid\pi^{\otimes N}).\] Removing self-interaction gives the off-diagonal identity \[\label{eq:entropy-intro} \begin{align} \frac{\,\mathrm d}{\,\mathrm dt}H_N(t) &=-\frac{1}{N^2}\mathbb{E}_{P_t^N}\!\left[\sum_{i\ne j} G_\pi(X_i(t),X_j(t))\right]\\ &=-2\mathbb{E}_{P_t^N}\!\left[\mathcal{F}_N^\pi(\boldsymbol{X}(t))\right]. \end{align}\tag{15}\] The problem is to recover the nonnegativity lost by deleting an infinite diagonal.

For comparison, modulated-energy methods compare an empirical measure with a continuum density without evaluating the singular energy on the particle diagonal. The framework was developed for Coulomb-type flows by [19] and extended to translation-invariant Riesz flows by [20], using renormalization and smearing ideas related to [21]. The resulting corrected modulated energy and commutator estimate give finite-time mean-field convergence for Riesz flows. In SVGD, however, the velocity decomposes as \[\label{eq:velocity-intro} v_\mu=-\nabla\kappa_a*\mu+\kappa_a*(b_\pi\mu).\tag{16}\] The first term has the commutator structure treated by that theory within its admissible range; the target-dependent second term creates a discrete empirical trace for which the same estimate is unavailable. This prevents a direct application of existing Riesz-flow theory to a last-iterate particle limit. The time-averaged argument below avoids such a particle-to-population comparison.

Our argument combines the joint-law entropy identity from smooth-kernel SVGD with a direct capped-diagonal passage from particle energies to the continuum KSD. Making this work for the target-dependent, self-interaction-free Riesz system requires three steps:

  1. Singular entropy identity. We first prove global collision avoidance and then justify the joint entropy calculation for the singular flow.

  2. Capped-diagonal empirical liminf. Capping the full Stein kernel below its singular diagonal gives a direct lower bound from the off-diagonal particle energy to the continuum KSD. This proves that the negative part is bounded by a correction \(c_N\to0\) and permits simultaneous averaging horizons.

  3. Quantitative correction below the logarithmic threshold. When \(0<\sigma<2\), an operator factorization and a positive Bessel minorant give the explicit rate \(c_N\lesssim N^{-1+\sigma/d}\).

The resulting finite-\(N\) renormalized entropy method is the central contribution: it converts the singular off-diagonal entropy production into a nonnegative continuum discrepancy up to a vanishing \(N\)-dependent correction. Because the argument works directly with the joint particle law, it does not require a propagation-of-chaos comparison with the population flow.

1.3 Main result and proof mechanism↩︎

The exact off-diagonal identity 15 does not yet provide dissipation because \(\mathcal{F}_N^\pi\) need not be nonnegative. Set \[\label{eq:onsager-intro} c_N:=\left(-\inf_{\boldsymbol{x}\in\mathcal{D}_N}\mathcal{F}_N^\pi(\boldsymbol{x})\right)_+.\tag{17}\] The capped-diagonal liminf proves, throughout \(0<\sigma<d\), that \(c_N\to0\). Therefore the renormalized energy \(\mathcal{E}_N^\pi:=\mathcal{F}_N^\pi+c_N\) is nonnegative. Integrating 15 and using \(H_N(T)\ge0\) gives \[\label{eq:rd-intro} \frac{1}{T}\int_0^T \mathbb{E}_{P_t^N}\!\left[\mathcal{E}_N^\pi(\boldsymbol{X}(t))\right]\,\,\mathrm dt \le \frac{H_N(0)}{2T}+c_N.\tag{18}\] For \(0<\sigma<2\), the quantitative refinement gives \[\label{eq:quantitative-onsager-intro} c_N\le C_{\rm ons}N^{-1+\sigma/d}.\tag{19}\] Thus 18 is the singular replacement for the finite-diagonal correction in the smooth-kernel entropy method, with an explicit rate for \(0<\sigma<2\).

The entropy identity is therefore naturally suited to time-averaged convergence. Let \(P_{t,i}^N\) be the \(i\)-th marginal of \(P_t^N\) and set \[\label{eq:tagged-average-intro} \overline{P}_{\mathrm{av},T}^{N} :=\frac{1}{T}\int_0^T\frac{1}{N}\sum_{i=1}^N P_{t,i}^N\,\,\mathrm dt.\tag{20}\] Writing \(d_{\mathrm{BL}}\) for the bounded-Lipschitz distance, we prove that whenever \(T_N\to\infty\) and \(H_N(0)/T_N\to0\), the time-averaged empirical-measure law converges to \(\delta_\pi\), the Dirac mass at the point \(\pi\in\mathcal{P}(\mathbb{T}^d)\): \[\label{eq:ordered-result-intro} \frac{1}{T_N}\int_0^{T_N} (\boldsymbol{x}\mapsto\mu_{\boldsymbol{x}}^N)_\#P_t^N\,\,\mathrm dt \rightharpoonup\delta_\pi, \qquad d_{\mathrm{BL}}\!\left(\overline{P}_{\mathrm{av},T_N}^{N},\pi\right)\to0.\tag{21}\] The first convergence therefore takes place in \(\mathcal{P}(\mathcal{P}(\mathbb{T}^d))\). In particular, every sequence \(T_N\to\infty\) is admissible under a uniform initial normalized-entropy bound. No exchangeability is required. We also prove that invariant laws of finite relative entropy have empirical-measure laws converging to \(\delta_\pi\), without any uniform entropy bound. Exchangeability is needed only to conclude that their first marginals converge to \(\pi\).

The proof of 21 has four conceptual steps. First, the Riesz repulsion prevents collisions, so the singular flow and its Liouville equation are globally defined. Second, the joint-entropy identity gives an upper bound on the averaged off-diagonal energy. Third, a bounded continuous cap of \(G_\pi\) gives a lower bound for every sequence of empirical-measure laws, including laws averaged over \(T_N\). Finally, compactness and the coercive identity \(\mathcal{F}^\pi(\mu)=0\) if and only if \(\mu=\pi\) identify the limit. The Bessel-minorant argument is an additional quantitative refinement and is not needed for 21 .

2 Setting and theorem statements↩︎

2.1 The torus, the target, and the functional setting↩︎

We identify \(\mathbb{T}^d\) with \([-\tfrac12,\tfrac12)^d\) and use normalized Lebesgue measure. Fourier coefficients are \[\widehat f(\ell)=\int_{\mathbb{T}^d}f(x)e^{-2\pi i\ell\cdot x}\,\,\mathrm dx.\] We use the bounded-Lipschitz metric \[d_{\mathrm{BL}}(\mu,\nu) := \sup_{\|f\|_\infty+\operatorname{Lip}(f)\le1} \left|\int_{\mathbb{T}^d}f\,\,\mathrm d(\mu-\nu)\right|,\] which metrizes weak convergence on \(\mathcal{P}(\mathbb{T}^d)\).

Assumption 1 (Target). The density \(\pi=Z^{-1}e^{-V}\) is strictly positive. For some \[m>\frac{d}{2}+2, \qquad m\ge a+1,\] we have \(\pi,V\in H^m(\mathbb{T}^d)\) and \(V\in W^{2,\infty}(\mathbb{T}^d)\).

The Sobolev assumptions in 1 are used in the continuum identification of the Riesz KSD with a negative Sobolev norm and in the operator estimates for \(0<\sigma<2\) in 2. Collision avoidance and the entropy calculation use only the bounded derivatives of the target score displayed in their proofs.

The Riesz kernel \(\kappa_a\) was defined in 5 . Its native space is the mean-zero Fourier space \[\label{eq:native-space} \mathcal{H}_{\kappa_a} :=\left\{f:\widehat f(0)=0,\; \|f\|_{\mathcal{H}_{\kappa_a}}^2 :=\sum_{\ell\ne0}\widehat\kappa_a(\ell)^{-1} |\widehat f(\ell)|^2<\infty\right\} =\dot{H}^a_0(\mathbb{T}^d).\tag{22}\] If \(a>d/2\), point evaluation is continuous and this is the usual mean-zero reproducing-kernel Hilbert space generated by \(\kappa_a\). The range considered here may also contain \(a\le d/2\), where point evaluation on \(\dot{H}^a\) is not continuous and \(\kappa_a\) is singular on the diagonal. The literal pointwise reproducing property is then unavailable. We therefore use 22 through Fourier duality; no point evaluation of the singular kernel is needed.

The main finite-particle theorems assume 9 . In this range the kernel is smooth away from the origin, its force is less singular than the Coulomb force \((a=1,\sigma=d)\), and the scalar kernel \(G_\pi\) has locally integrable singularity \(\operatorname{dist}(x,y)^{-\sigma}\).

The identity 12 initially defines \(\mathcal{F}^\pi\) for smooth measures. Later we apply it to weak limits of empirical measures, which need not have densities. We therefore use \(\mathcal{F}^\pi\) for the lower-semicontinuous relaxation of the smooth Gram energy to \(\mathcal{P}(\mathbb{T}^d)\), allowing the value \(+\infty\). The diagonal-cap construction in 6 gives a concrete representation of this relaxation.

2.2 Entropy dissipation and long-time sampling↩︎

Recall from the introduction that \(\mathcal{D}_N\) denotes the collision-free configuration space on which the particle vector field is defined.

The first theorem establishes the analytic foundation for the entropy method. Its three conclusions have distinct roles: collision avoidance makes the singular flow global; the Liouville calculation identifies its exact entropy production; and a qualitative correction turns that sign-indefinite off-diagonal production into a nonnegative energy up to a vanishing finite-\(N\) error.

Theorem 1 (Global dynamics and renormalized entropy). Suppose 1 and 9 hold.

  1. Every solution of 8 starting in \(\mathcal{D}_N\) exists globally and remains in \(\mathcal{D}_N\).

  2. Let \(P_0^N\ll\pi^{\otimes N}\) have finite relative entropy, and let \(P_t^N\) be its pushforward by the flow. Then, for every \(t\ge0\), \[\label{eq:entropy-identity} H_N(t)+2\int_0^t\mathbb{E}_{P_s^N}\mathcal{F}_N^\pi\,\,\mathrm ds=H_N(0), \qquad H_N(t):=\frac{1}{N}\operatorname{Ent}(P_t^N\mid\pi^{\otimes N}).\tag{23}\]

  3. If \[\label{eq:qualitative-correction} m_N:=\inf_{\boldsymbol{x}\in\mathcal{D}_N}\mathcal{F}_N^\pi(\boldsymbol{x}), \qquad c_N:=(-m_N)_+, \qquad \mathcal{E}_N^\pi:=\mathcal{F}_N^\pi+c_N,\tag{24}\] then \(c_N\to0\), \(\mathcal{E}_N^\pi\ge0\) on \(\mathcal{D}_N\), and \[\label{eq:renormalized-dissipation} \frac{1}{T}\int_0^T\mathbb{E}_{P_t^N}\mathcal{E}_N^\pi\,\,\mathrm dt \le \frac{H_N(0)}{2T}+c_N \qquad(T>0).\tag{25}\]

For a smooth kernel, the analogue of the third conclusion follows by adding the finite diagonal of the Gram matrix. Here that diagonal is infinite. The correction in 24 is shown to vanish by the direct capped-diagonal liminf of 9; no smearing estimate is needed for this qualitative conclusion.

Below the logarithmic threshold \(\sigma=2\), the correction has an explicit natural-scale bound.

Theorem 2 (Quantitative Onsager bound for \(0<\sigma<2\)). Under the assumptions of 1, assume in addition that \(0<\sigma<2\). There is \(C_{\rm ons}<\infty\), depending only on \(d,a,\pi\), such that \[\label{eq:onsager} \mathcal{F}_N^\pi(\boldsymbol{x})\ge-C_{\rm ons}N^{-1+\sigma/d} \qquad(N\ge2,\;\boldsymbol{x}\in\mathcal{D}_N).\tag{26}\] Consequently, \[\label{eq:quantitative-correction} c_N\le C_{\rm ons}N^{-1+\sigma/d}.\tag{27}\] Moreover, for the particle laws in 1, \[\label{eq:quantitative-dissipation} \frac{1}{T}\int_0^T\mathbb{E}_{P_t^N}\!\left[ \mathcal{F}_N^\pi+C_{\rm ons}N^{-1+\sigma/d}\right]\,\,\mathrm dt \le \frac{H_N(0)}{2T}+C_{\rm ons}N^{-1+\sigma/d}.\tag{28}\]

The next theorem converts the averaged energy bound into concentration of the time-averaged empirical-measure law at the target, and hence convergence of its label-averaged marginal.

Theorem 3 (Simultaneous many-particle and long-time limit). Under the assumptions of 1, let \(T_N\to\infty\) and assume \[\label{eq:entropy-time-condition} \frac{H_N(0)}{T_N}\longrightarrow0.\tag{29}\] Define the time-averaged empirical-measure law \[\label{eq:Lambda-simultaneous} \Lambda_{N,T_N} :=\frac{1}{T_N}\int_0^{T_N} (\boldsymbol{x}\mapsto\mu_{\boldsymbol{x}}^N)_\#P_t^N\,\,\mathrm dt.\tag{30}\] Then \[\label{eq:ordered-main} \Lambda_{N,T_N}\rightharpoonup\delta_\pi, \qquad d_{\mathrm{BL}}\!\left(\overline{P}_{\mathrm{av},T_N}^{N},\pi\right) \longrightarrow0.\tag{31}\] In particular, 31 holds for every \(T_N\to\infty\) if \(\sup_NH_N(0)<\infty\).

The theorem concerns the full law of the empirical measure and therefore also the label-averaged marginal. It does not require the joint law to be exchangeable. Condition 29 shows that a uniform entropy bound is sufficient but not necessary.

At stationarity, the entropy identity forces the mean off-diagonal dissipation to vanish. Passing this fact to the continuum yields the following consequence.

Corollary 1 (No macroscopic spurious stationary law). For each \(N\), let \(Q^N\ll\pi^{\otimes N}\) be invariant under the flow 8 and have finite relative entropy. Then \[\label{eq:stationary-empirical-law} (\boldsymbol{x}\mapsto\mu_{\boldsymbol{x}}^N)_\#Q^N\rightharpoonup\delta_\pi, \qquad \frac{1}{N}\sum_{i=1}^NQ_i^N\rightharpoonup\pi.\tag{32}\] Here \(Q_i^N\) denotes the \(i\)-th marginal of \(Q^N\). If the laws are exchangeable, then \(Q_1^N\rightharpoonup\pi\).

3 Stein–Riesz structure↩︎

This section establishes the two properties of \(G_\pi\) used throughout the proof. Locally, it has a positive Riesz singularity; this drives collision avoidance and permits a monotone cap of the diagonal. Globally, it is a centered Gram kernel; this makes the continuum energy nonnegative and identifies it as a KSD. The finite-particle difficulty is precisely that these global properties do not make sense on the infinite diagonal.

3.1 Local singularity↩︎

Proposition 4 (Local expansion). There are \(r_0,c_0,C>0\) such that, for \(0<r:=\operatorname{dist}(x,y)<r_0\), \[\label{eq:local-G} G_\pi(x,y)=c_0r^{-\sigma}+R_\pi(x,y),\qquad{(1)}\] where \[\label{eq:R-bound} |R_\pi(x,y)| \le C \begin{cases} 1,&0<\sigma<2,\\ 1+|\log r|,&\sigma=2,\\ 1+r^{-(\sigma-2)},&2<\sigma<d. \end{cases}\qquad{(2)}\] In particular, after decreasing \(r_0\) if needed, \[\label{eq:G-positive-near} G_\pi(x,y)\ge\frac{c_0}{2}\operatorname{dist}(x,y)^{-\sigma}>0 \qquad(0<\operatorname{dist}(x,y)<r_0).\qquad{(3)}\]

Proof. The periodic Riesz asymptotics in 8 give \[-\Delta\kappa_a(z)=c_0|z|^{-\sigma}+O(1)\] and \[|\kappa_a(z)|\lesssim \begin{cases} 1,&\sigma<2,\\ 1+|\log|z||,&\sigma=2,\\ |z|^{-(\sigma-2)},&\sigma>2, \end{cases} \qquad |\nabla\kappa_a(z)|\lesssim 1+|z|^{1-\sigma}.\] Since \(b_\pi\) is Lipschitz, \[|(b_\pi(x)-b_\pi(y))\cdot\nabla\kappa_a(y-x)| \lesssim |x-y|\bigl(1+|x-y|^{1-\sigma}\bigr).\] Substitution in 11 proves ?? –?? . Every remainder is lower order than \(r^{-\sigma}\), which yields ?? . ◻

The proposition shows that the target dependence does not alter the leading singularity. The positivity in ?? will be used twice: first to repel collapsing clusters, and later to cap the singular diagonal from below without losing an empirical lower bound.

3.2 Stein centering and positivity↩︎

Local positivity controls the singularity but does not yet explain why the continuum energy measures discrepancy from \(\pi\). That information comes from the global Stein identities.

Proposition 5 (Centering and Gram representation). The kernel is centered in each variable: \[\label{eq:centering} \int_{\mathbb{T}^d}G_\pi(x,y)\,\,\mathrm d\pi(y)=0, \qquad \int_{\mathbb{T}^d}G_\pi(x,y)\,\,\mathrm d\pi(x)=0\qquad{(4)}\] in the distributional sense. If \(\mu\) is smooth, then \[\label{eq:gram} 2\mathcal{F}^\pi(\mu) = \left\| \int_{\mathbb{T}^d}\mathcal{A}_{\pi,y}k(y,\cdot)\,\,\mathrm d\mu(y) \right\|_{\mathcal{H}_{\kappa_a}^d}^{\,2}\ge0.\qquad{(5)}\] Consequently \(\mathcal{F}^\pi\) is convex.

Proof. We first record the two integrations by parts separately. For every smooth periodic vector field \(u\), \[\int_{\mathbb{T}^d}\mathcal{S}_{\pi}u\,\,\mathrm d\pi =\int_{\mathbb{T}^d}\nabla\cdot(u\pi)\,\,\mathrm dx=0.\] For fixed \(x\), the feature is also centered: \[\int_{\mathbb{T}^d}\mathcal{A}_{\pi,y}k(y,x)\,\,\mathrm d\pi(y) =\int_{\mathbb{T}^d}\nabla_y\!\bigl(k(y,x)\pi(y)\bigr)\,\,\mathrm dy=0.\] These formulas are understood after testing in the remaining variable, so they remain valid for the singular kernel. Apply the first formula to \(u(x)=\mathcal{A}_{\pi,y}k(y,x)\), and apply the second before \(\mathcal{S}_{\pi,x}\), to obtain the two identities in ?? .

We next make the Gram calculation explicit without invoking pointwise reproduction. Write \(\,\mathrm d\mu=\rho\,\,\mathrm dx\) and set \[q_\rho:=b_\pi\rho-\nabla\rho.\] Integration by parts in the feature variable gives \[\Phi_\rho :=\int_{\mathbb{T}^d}\mathcal{A}_{\pi,y}k(y,\cdot)\rho(y)\,\,\mathrm dy =\kappa_a*q_\rho.\] For each fixed \(y\), integration by parts in \(x\) gives \[\int_{\mathbb{T}^d}\rho(x)\, \mathcal{S}_{\pi,x}\mathcal{A}_{\pi,y}k(y,x)\,\,\mathrm dx =\int_{\mathbb{T}^d}q_\rho(x)\cdot \mathcal{A}_{\pi,y}k(y,x)\,\,\mathrm dx.\] Integrating against \(\rho(y)\,\,\mathrm dy\) and using the definition of \(\Phi_\rho\), we find \[\label{eq:gram-intermediate} \iint G_\pi(x,y)\rho(x)\rho(y)\,\,\mathrm dx\,\,\mathrm dy =\int_{\mathbb{T}^d}q_\rho(x)\cdot\Phi_\rho(x)\,\,\mathrm dx =\int_{\mathbb{T}^d}q_\rho\cdot(\kappa_a*q_\rho)\,\,\mathrm dx.\tag{33}\] This is the quadratic form associated with the scalar Stein kernel. Since \(\widehat\kappa_a(\ell)>0\) for \(\ell\ne0\), Parseval’s identity yields \[\int q_\rho\cdot(\kappa_a*q_\rho) =\sum_{\ell\ne0}\widehat\kappa_a(\ell) |\widehat q_\rho(\ell)|^2 =\sum_{\ell\ne0}\widehat\kappa_a(\ell)^{-1} |\widehat\Phi_\rho(\ell)|^2 =\|\Phi_\rho\|_{\mathcal{H}_{\kappa_a}^d}^{2},\] where the middle equality uses \(\widehat\Phi_\rho=\widehat\kappa_a\widehat q_\rho\), and the last is exactly 22 . Finally, the feature generated by \(\pi\) vanishes because \(b_\pi\pi-\nabla\pi=0\). Thus the same identity holds with \(\rho\) replaced by \(\rho-\pi\). Together with ?? and 33 , this proves ?? . Convexity follows because the right-hand side is the squared norm of a linear function of \(\mu\); it passes to the lower-semicontinuous extension by approximation. ◻

Remark 6. Equations ?? and ?? explain why the centered form in 12 equals \(\frac{1}{2}\iint G_\pi\,\,\mathrm d\mu\,\,\mathrm d\mu\) whenever either expression is finite. They do not justify inserting an atom on the diagonal; that is precisely the finite-\(N\) renormalization problem.

4 Collision-free particle dynamics and Liouville structure↩︎

Before differentiating the joint entropy, we must know that the singular ODE is globally well-posed. The only possible finite-time obstruction is a collision. Below the logarithmic threshold, the radius of a collapsing cluster has a strictly positive derivative at each sufficiently small running minimum. At and above the threshold, the Riesz pair energy diverges at collision and is decreasing near the collision set. Once collisions are excluded, the Liouville calculation can be performed on the collision-free configuration space.

4.1 Cluster repulsion below the logarithmic threshold↩︎

Lemma 1 (Cluster differential inequality). Assume \(0<\sigma<2\). In local Euclidean lifts, let \[\bar X_I=\frac{1}{|I|}\sum_{i\in I}X_i, \qquad R_I^2=\sum_{i\in I}|X_i-\bar X_I|^2\] for a set \(I\) with \(|I|\ge2\). If, at the time under consideration, \(R_I\) is sufficiently small and the particles in \(I\) are separated from those outside \(I\) by a fixed positive distance, then \[\label{eq:cluster-ineq} \frac{1}{2}\frac{\,\mathrm d}{\,\mathrm dt}R_I^2 \ge c\bigl(R_I^2\bigr)^{1-\sigma/2}-CR_I^2,\tag{34}\] where \(c,C>0\) may depend on the separation and on \(I\).

Proof. Write \(q_i=X_i-\bar X_I\). Symmetrizing the internal Riesz force gives \[\frac{1}{N}\sum_{i\in I}q_i\cdot \sum_{\substack{j\in I\\j\ne i}}\nabla\kappa_a(X_j-X_i) = \frac{1}{N}\sum_{\substack{i<j\\i,j\in I}} -(X_j-X_i)\cdot\nabla\kappa_a(X_j-X_i).\] The asymptotic bound \[\label{eq:radial-repulsion} -z\cdot\nabla\kappa_a(z)\ge c|z|^{2-\sigma}\tag{35}\] and the identity \(\sum_{i<j}|X_i-X_j|^2=|I|R_I^2\) show that this is bounded below by \(c(R_I^2)^{1-\sigma/2}\).

For the target-weighted internal force, subtract \((|I|-1)\kappa_a(0)b_\pi(\bar X_I)\) from each inner sum; its total contribution vanishes because \(\sum_iq_i=0\). The remaining summands are \[\bigl(\kappa_a(X_j-X_i)-\kappa_a(0)\bigr)b_\pi(X_j) +\kappa_a(0)\bigl(b_\pi(X_j)-b_\pi(\bar X_I)\bigr).\] Since \[|\kappa_a(z)-\kappa_a(0)|\lesssim |z|^{2-\sigma}, \qquad |b_\pi(X_j)-b_\pi(\bar X_I)|\lesssim R_I,\] their contribution is \(O(R_I^{3-\sigma})+O(R_I^2)\), which is lower order than \(R_I^{2-\sigma}\). Finally, the velocity generated by particles outside \(I\) is Lipschitz across the separated cluster. Subtracting its value at \(\bar X_I\) leaves \(O(R_I^2)\). These estimates give 34 . ◻

4.2 A pair-energy barrier at and above the logarithmic threshold↩︎

Define \[\label{eq:pair-energy} \mathcal{W}_N(\boldsymbol{x}):=\sum_{1\le i<j\le N}\kappa_a(x_j-x_i)\tag{36}\] and let \(\boldsymbol{C}^N=(C_1^N,\ldots,C_N^N)\), where \[C_i^N(\boldsymbol{x}) :=\frac{1}{N}\sum_{j\ne i}\kappa_a(x_j-x_i)b_\pi(x_j).\] Then the particle system can be written as \[\label{eq:pair-energy-flow} \dot{\boldsymbol{X}}=-\frac{1}{N}\nabla\mathcal{W}_N(\boldsymbol{X})+\boldsymbol{C}^N(\boldsymbol{X}).\tag{37}\]

Lemma 2 (Pair-energy barrier). Assume \(2\le\sigma<d\). As a collision-free configuration approaches the collision set, \[\label{eq:pair-barrier-limits} \mathcal{W}_N(\boldsymbol{x})\longrightarrow+\infty, \qquad \frac{|\boldsymbol{C}^N(\boldsymbol{x})|}{|\nabla\mathcal{W}_N(\boldsymbol{x})|} \longrightarrow0.\tag{38}\]

Proof. Let \(r(\boldsymbol{x}):=\min_{i\ne j}\operatorname{dist}(x_i,x_j)\). The first limit follows from the logarithmic divergence of \(\kappa_a\) when \(\sigma=2\), its \(r^{2-\sigma}\) divergence when \(\sigma>2\), and its boundedness from below away from the diagonal. Moreover, \[|\boldsymbol{C}^N(\boldsymbol{x})| \le C \begin{cases} 1+|\log r(\boldsymbol{x})|,&\sigma=2,\\ r(\boldsymbol{x})^{2-\sigma},&\sigma>2. \end{cases}\] The closest-scale cluster estimate in 9 gives \[|\nabla\mathcal{W}_N(\boldsymbol{x})| \ge c\,r(\boldsymbol{x})^{1-\sigma}\] for all sufficiently small \(r(\boldsymbol{x})\). The quotient is therefore bounded by \(Cr(1+|\log r|)\) at \(\sigma=2\) and by \(Cr\) when \(\sigma>2\), proving 38 . ◻

Proposition 7 (No collision). Every solution of 8 with initial condition in \(\mathcal{D}_N\) is global and collision-free.

Proof. The vector field is locally Lipschitz on \(\mathcal{D}_N\), so a maximal solution can fail to continue only by approaching the collision set. Suppose first that \(0<\sigma<2\) and a first collision occurs at \(\tau<\infty\). For \(|I|\ge2\), set \[Y_I(t):=\frac{1}{|I|}\sum_{\substack{i<j\\i,j\in I}} \operatorname{dist}(X_i(t),X_j(t))^2.\] On a sufficiently small sublevel, the particles in \(I\) lie in a common coordinate chart and \(Y_I=R_I^2\). A collision sequence supplies at least one set with \(\liminf_{t\uparrow\tau}Y_I(t)=0\); among all such sets, choose one that is maximal under inclusion. The elementary running-minimum selection then gives times \(t_n\uparrow\tau\) such that \[Y_I(t_n)\longrightarrow0, \qquad Y_I'(t_n)\le0.\] Maximality implies, after passing to a subsequence, that the particles in \(I\) remain a fixed positive distance from every exterior particle at these times; otherwise \(I\) could be enlarged. Applying 1 at \(t_n\) gives a strictly positive derivative for all large \(n\), a contradiction.

Now let \(2\le\sigma<d\). By 2, sufficiently near the collision set, \[|\boldsymbol{C}^N|\le\frac{1}{2N}|\nabla\mathcal{W}_N|.\] Along 37 , \[\frac{\,\mathrm d}{\,\mathrm dt}\mathcal{W}_N(\boldsymbol{X}(t)) =-\frac{1}{N}|\nabla\mathcal{W}_N|^2 +\nabla\mathcal{W}_N\cdot\boldsymbol{C}^N \le-\frac{1}{2N}|\nabla\mathcal{W}_N|^2.\] To account for trajectories that enter and leave this neighborhood, choose \(\delta>0\) so that the displayed inequality holds whenever \(r(\boldsymbol{X})<\delta\). On every connected component of \(\{t:r(\boldsymbol{X}(t))<\delta\}\), the pair energy is nonincreasing. At an entry time it is uniformly bounded because \(\{\boldsymbol{x}:r(\boldsymbol{x})\ge\delta\}\) is compact; on a component containing the initial time it is bounded by its initial value. Thus \(\mathcal{W}_N\) stays uniformly bounded along every collision sequence, contradicting 38 . Hence no finite-time collision occurs in either range. ◻

Having ruled out collisions, we pass from individual trajectories to the law of the full particle configuration. The appropriate equation is the Liouville equation, the continuity equation for a probability density transported by a deterministic vector field. This is useful here because differentiating relative entropy along a continuity equation produces the divergence of that vector field relative to the reference measure \(\pi^{\otimes N}\). The next identity recognizes this relative divergence as the off-diagonal Stein energy.

4.3 Liouville equation↩︎

From this point on, denote the right-hand side of 8 by \[B_i^N(\boldsymbol{X}) :=\frac{1}{N}\sum_{j\ne i}\mathcal{A}_{\pi,X_j}k(X_j,X_i), \qquad \dot{X}_i=B_i^N(\boldsymbol{X}).\]

Lemma 3 (Divergence identity). On \(\mathcal{D}_N\), the vector field \(\boldsymbol{B}^N=(B_1^N,\ldots,B_N^N)\) satisfies \[\label{eq:divergence} \operatorname{div}_{\boldsymbol{x}}\boldsymbol{B}^N(\boldsymbol{x}) +\sum_{i=1}^Nb_\pi(x_i)\cdot B_i^N(\boldsymbol{x}) =2N\mathcal{F}_N^\pi(\boldsymbol{x}).\tag{39}\]

Proof. By 10 , \[\nabla_{x_i}\cdot\mathcal{A}_{\pi,x_j}k(x_j,x_i) +b_\pi(x_i)\cdot\mathcal{A}_{\pi,x_j}k(x_j,x_i) =G_\pi(x_i,x_j).\] Sum over \(i\ne j\) and divide by \(N\). ◻

Thus \(G_\pi\) is exactly the pairwise contribution to the compressibility of the joint flow relative to \(\pi^{\otimes N}\). Identity 39 is what turns the Liouville transport equation into the entropy-dissipation identity in 8.

If \(P_t^N\) has density \(p_t^N\), it solves \[\label{eq:liouville} \partial_t p_t^N+\operatorname{div}_{\boldsymbol{x}}(p_t^N\boldsymbol{B}^N)=0\tag{40}\] on \(\mathcal{D}_N\). No boundary condition is imposed at collisions: by 7, the flow never reaches that boundary.

5 Joint entropy and Onsager renormalization↩︎

We now turn the deterministic flow into a dissipative estimate for its joint law. The calculation has two parts. The Liouville equation gives an exact identity with \(\mathcal{F}_N^\pi\) as entropy production. Since this off-diagonal quantity can be negative, the direct empirical liminf in 9 supplies a qualitative vanishing correction. For \(0<\sigma<2\), the Bessel-minorant argument of 8 gives its explicit rate.

5.1 Exact entropy production↩︎

The entropy identity below is the structural ingredient inherited from the smooth-kernel joint-law method. The work here is to justify it despite the singular collision set; the no-collision theorem and the approximation in 12 provide that justification.

Proposition 8 (Entropy identity). Let \(P_0^N\ll\pi^{\otimes N}\) have finite relative entropy, and let \(P_t^N=(\Phi_t^N)_\#P_0^N\), where \(\Phi_t^N\) is the collision-free flow of 8 . Then 23 holds for every \(t\ge0\).

Proof. Assume first that \(p_t^N\) is smooth and that its support stays a positive distance from the collision set on the time interval under consideration. Mass conservation removes the derivative of the linear part of the entropy, and integration by parts in 40 gives \[\begin{align} \frac{\,\mathrm d}{\,\mathrm dt}\operatorname{Ent}(P_t^N\mid\pi^{\otimes N}) &=\int p_t^N\boldsymbol{B}^N\cdot \nabla_{\boldsymbol{x}}\log\frac{p_t^N}{\pi^{\otimes N}}\,\,\mathrm d\boldsymbol{x}\\ &=-\int p_t^N\left( \operatorname{div}_{\boldsymbol{x}}\boldsymbol{B}^N+ \sum_ib_\pi(x_i)\cdot B_i^N\right)\,\mathrm d\boldsymbol{x}\\ &=-2N\,\mathbb{E}_{P_t^N}\mathcal{F}_N^\pi, \end{align}\] where the last equality is 39 . Divide by \(N\) and integrate in time.

For a general finite-entropy initial law, localize to sets on which the density ratio is bounded and trajectories remain uniformly separated from collisions. The same calculation applies on each localized flow tube. The exhaustion argument, including passage of the shifted energy and both endpoint entropies to the limit, is given in 12. ◻

5.2 Renormalized dissipation↩︎

The entropy identity alone is insufficient because its production term is the off-diagonal energy, not the nonnegative continuum KSD. The empirical liminf in 6 supplies the vanishing correction throughout \(0<\sigma<d\); 8 later gives its explicit rate for \(0<\sigma<2\).

Proof of 1. Global existence is 7, and the exact identity is 8. The qualitative correction \(c_N\to0\) and \(\mathcal{E}_N^\pi\ge0\) are proved in 10. Since relative entropy is nonnegative, 23 then gives \[\frac{1}{T}\int_0^T\mathbb{E}\mathcal{E}_N^\pi\,\,\mathrm dt =\frac{H_N(0)-H_N(T)}{2T}+c_N \le\frac{H_N(0)}{2T}+c_N.\] ◻

6 Passage to continuum KSD↩︎

Estimate 25 controls a scalar particle energy. To obtain convergence of measures, we must show that this energy cannot disappear in the many-particle limit. We consider laws of the empirical measure \(\mu_{\boldsymbol{x}}^N\) defined in the introduction. Compactness of \(\mathcal{P}(\mathbb{T}^d)\) gives subsequential limits automatically. The work is to pass the singular off-diagonal energy to those limits. We do this by capping \(G_\pi\) at a finite height, passing the resulting bounded continuous kernel to the limit, and then raising the cap. The finite diagonal introduced by the cap costs only \(L/(2N)\). Since this comparison holds for every configuration law, it also applies when the averaging horizon depends on \(N\).

6.1 A capped-diagonal representation↩︎

By 4, \(G_\pi(x,y)\to+\infty\) uniformly as \(\operatorname{dist}(x,y)\to0\). Since it is continuous away from the diagonal, it is also bounded below on \(\mathbb{T}^d\times\mathbb{T}^d\setminus\operatorname{diag}\), where \(\operatorname{diag}:=\{(x,x):x\in\mathbb{T}^d\}\); fix \(C_0\) such that \(G_\pi\ge-C_0\). For \(L\ge1\), define the capped kernel \[G_\pi^{(L)}(x,y) := \begin{cases} \min\{G_\pi(x,y),L\},&x\ne y,\\ L,&x=y, \end{cases}\] and \[\label{eq:F-cap} \mathcal{F}_L^\pi(\mu) :=\frac{1}{2}\iint G_\pi^{(L)}(x,y)\,\,\mathrm d\mu(x)\,\,\mathrm d\mu(y).\tag{41}\] The value on the diagonal makes \(G_\pi^{(L)}\) bounded and continuous on the full product space. Thus \(\mathcal{F}_L^\pi\) is continuous under weak convergence. The family increases with \(L\), and \(\mathcal{F}_L^\pi\ge-C_0/2\).

Lemma 4 (Capped-diagonal characterization). The lower-semicontinuous Gram energy satisfies \[\label{eq:F-cutoff-limit} \mathcal{F}^\pi(\mu)=\lim_{L\to\infty}\mathcal{F}_L^\pi(\mu) \quad\text{for every }\mu\in\mathcal{P}(\mathbb{T}^d).\tag{42}\] In particular, an atom has infinite energy, and \(\mathcal{F}^\pi\) is lower semicontinuous.

Proof. Set \(I(\mu):=\sup_L\mathcal{F}_L^\pi(\mu)\). For a smooth density, \(G_\pi\) is locally integrable because \(\sigma<d\), and monotone convergence after adding \(C_0\), together with ?? , gives \(I(\mu)=\mathcal{F}^\pi(\mu)\). Moreover, \(\mathcal{F}_L^\pi\le\mathcal{F}^\pi\) on smooth measures. Since \(\mathcal{F}_L^\pi\) is continuous, the definition of \(\mathcal{F}^\pi\) as the lower-semicontinuous relaxation of the smooth Gram form yields \[\label{eq:cap-below-gram} I(\mu)\le\mathcal{F}^\pi(\mu) \qquad\text{for every }\mu\in\mathcal{P}(\mathbb{T}^d).\tag{43}\]

It remains to prove the reverse inequality when \(I(\mu)<\infty\). The local expansion and the lower bound away from the diagonal imply \[G_\pi(x,y) \ge c\,\operatorname{dist}(x,y)^{-\sigma} \mathbf{1}_{\{\operatorname{dist}(x,y)<r_0\}}-C.\] Thus \(\mu\) has finite Riesz \(\sigma\)-energy. Let \(\mu_\varepsilon=p_\varepsilon*\mu\), where \(p_\varepsilon\) is the periodic heat kernel. The nonnegative Fourier series of \(G_0=-\Delta\kappa_a\) shows both that the principal energies of \(\mu_\varepsilon\) are uniformly bounded and that they converge to the principal energy of \(\mu\).

The target-dependent remainder has strictly smaller singular order. More precisely, by ?? , for some \(\tau<\sigma\), \[|R_\pi(x,y)|\le C\bigl(1+\operatorname{dist}(x,y)^{-\tau}\bigr),\] where a logarithm is dominated by any positive power. For \(\delta<r_0\), \[\operatorname{dist}(x,y)^{-\tau}\mathbf{1}_{\{\operatorname{dist}(x,y)<\delta\}} \le \delta^{\sigma-\tau}\operatorname{dist}(x,y)^{-\sigma}.\] The uniform principal-energy bound therefore makes the remainder uniformly integrable under \(\mu_\varepsilon\otimes\mu_\varepsilon\). Away from the diagonal it is continuous, so weak convergence and then \(\delta\downarrow0\) give convergence of the remainder integrals. Hence \[\mathcal{F}^\pi(\mu_\varepsilon) \longrightarrow \frac{1}{2}\iint G_\pi(x,y)\,\,\mathrm d\mu(x)\,\,\mathrm d\mu(y) =I(\mu),\] where the last identity is monotone convergence for \(G_\pi+C_0\). The definition of the relaxation now gives \(\mathcal{F}^\pi(\mu)\le\liminf_{\varepsilon\downarrow0} \mathcal{F}^\pi(\mu_\varepsilon)=I(\mu)\). Together with 43 , this proves 42 when \(I(\mu)<\infty\). If \(I(\mu)=\infty\), then 43 already forces \(\mathcal{F}^\pi(\mu)=\infty\).

If \(\mu\{x\}=m>0\), then \(\mathcal{F}_L^\pi(\mu)\ge \frac{1}{2}Lm^2-C_0/2\), which proves the assertion about atoms. Finally, \(\mathcal{F}^\pi\) is a monotone supremum of the continuous functionals \(\mathcal{F}_L^\pi\), after the fixed shift \(C_0/2\), and is therefore lower semicontinuous. ◻

6.2 Microscopic-to-continuum lower bound↩︎

The next proposition is the bridge from particle dissipation to continuum dissipation. Its lower-bound direction is especially important: it prevents microscopic concentration near the singular diagonal from creating an artificially small limiting energy.

Proposition 9 (Empirical liminf). Let \(Q^N\in\mathcal{P}(\mathcal{D}_N)\) and let \[\Lambda_N:=(\boldsymbol{x}\mapsto\mu_{\boldsymbol{x}}^N)_\#Q^N \in\mathcal{P}(\mathcal{P}(\mathbb{T}^d)).\] If \(\Lambda_N\rightharpoonup\Lambda\), then \[\label{eq:empirical-liminf} \int_{\mathcal{P}(\mathbb{T}^d)}\mathcal{F}^\pi(\mu)\,\,\mathrm d\Lambda(\mu) \le \liminf_{N\to\infty} \int_{\mathcal{D}_N}\mathcal{F}_N^\pi(\boldsymbol{x})\,\,\mathrm dQ^N(\boldsymbol{x}).\qquad{(6)}\]

Proof. Fix \(L\ge1\). Since \(G_\pi^{(L)}\le G_\pi\) off the diagonal, \[\begin{align} \mathcal{F}_N^\pi(\boldsymbol{x}) &\ge \frac{1}{2N^2}\sum_{i\ne j}G_\pi^{(L)}(x_i,x_j)\\ &=\mathcal{F}_L^\pi(\mu_{\boldsymbol{x}}^N)-\frac{L}{2N}. \end{align}\] The last term is exactly the diagonal contribution of the capped kernel. Therefore \[\liminf_{N\to\infty}\int\mathcal{F}_N^\pi\,\,\mathrm dQ^N \ge \lim_{N\to\infty}\int\mathcal{F}_L^\pi(\mu)\,\,\mathrm d\Lambda_N(\mu) = \int\mathcal{F}_L^\pi(\mu)\,\,\mathrm d\Lambda(\mu),\] where \(L/(2N)\to0\), and the last equality uses continuity and boundedness of \(\mathcal{F}_L^\pi\). Let \(L\to\infty\) and apply monotone convergence after adding \(C_0/2\). ◻

Proposition 10 (Qualitative correction). For every \(0<\sigma<d\), the numbers \(m_N,c_N\) defined in 24 satisfy \[-\infty<m_N\le0, \qquad c_N\longrightarrow0.\] Consequently, \(\mathcal{E}_N^\pi=\mathcal{F}_N^\pi+c_N\ge0\).

Proof. The global off-diagonal lower bound \(G_\pi\ge-C_0\) gives \(m_N\ge-C_0/2\). On the other hand, if \(X_1,\ldots,X_N\) are independent with law \(\pi\), local integrability and Stein centering give \[\mathbb{E}\mathcal{F}_N^\pi(X_1,\ldots,X_N) =\frac{N(N-1)}{2N^2} \iint G_\pi(x,y)\,\,\mathrm d\pi(x)\,\,\mathrm d\pi(y) =0.\] Hence \(m_N\le0\).

If \(c_N\not\to0\), then along a subsequence \(c_N\ge2\varepsilon\) for some \(\varepsilon>0\). Since \(m_N=-c_N\), choose approximate minimizers \(\boldsymbol{x}^N\in\mathcal{D}_N\) such that \[\mathcal{F}_N^\pi(\boldsymbol{x}^N)\le m_N+\varepsilon\le-\varepsilon.\] After extracting once more, \(\mu_{\boldsymbol{x}^N}^N\rightharpoonup\mu\). Apply 9 with \(Q^N=\delta_{\boldsymbol{x}^N}\): \[0\le\mathcal{F}^\pi(\mu) \le\liminf_{N\to\infty}\mathcal{F}_N^\pi(\boldsymbol{x}^N) \le-\varepsilon,\] a contradiction. Thus \(c_N\to0\), and nonnegativity follows from its definition. ◻

Remark 11 (Why the comparison is uniform in the horizon). The lower-bound direction replaces \(G_\pi\) by a pointwise smaller capped kernel and pays only its explicit diagonal term. The lower-order, possibly sign-indefinite part of \(G_\pi\) is never separated from the full kernel. No time horizon enters ?? ; this is precisely what permits the simultaneous limit in 3.

6.3 Time-averaged empirical laws↩︎

For \(T>0\), define \[\label{eq:Lambda-NT} \Lambda_{N,T} :=\frac{1}{T}\int_0^T (\boldsymbol{x}\mapsto\mu_{\boldsymbol{x}}^N)_\#P_t^N\,\,\mathrm dt.\tag{44}\] Its barycenter is exactly the measure in 20 : \[\label{eq:barycenter} \int_{\mathcal{P}(\mathbb{T}^d)}\mu\,\,\mathrm d\Lambda_{N,T}(\mu) =\overline{P}_{\mathrm{av},T}^{N}.\tag{45}\]

7 Target identification and long-time sampling↩︎

The preceding section passes the off-diagonal particle energy to every weak limit of empirical-measure laws. It remains to identify the zero set of the continuum energy. The Riesz Fourier multiplier identifies \(\mathcal{F}^\pi\) with a negative Sobolev norm of \(\mu-\pi\), so \(\pi\) is its unique zero.

7.1 The zero set of the Riesz KSD↩︎

The following coercive identification is the point at which the Riesz Fourier structure matters globally.

Proposition 12 (Negative-Sobolev equivalence). Under 1, there are constants \(0<c\le C<\infty\) such that for every \(\mu\in\mathcal{P}(\mathbb{T}^d)\), \[\label{eq:sobolev-equivalence} c\|\mu-\pi\|_{\dot{H}^{1-a}}^2 \le 2\mathcal{F}^\pi(\mu) \le C\|\mu-\pi\|_{\dot{H}^{1-a}}^2,\qquad{(7)}\] with the convention that both sides may be infinite. Consequently \[\label{eq:unique-zero} \mathcal{F}^\pi(\mu)=0\quad\Longleftrightarrow\quad\mu=\pi.\qquad{(8)}\]

Proof. Let \(\Pi f:=f-\int_{\mathbb{T}^d}f\) denote removal of the zero Fourier mode, acting componentwise on vector distributions, and, for a zero-mass distribution \(\nu\), define \[P_\pi\nu:=\Pi(\nabla\nu-b_\pi\nu).\] For smooth zero-mass \(\nu\), the Fourier form of ?? gives \[\label{eq:Ppi-energy} 2\mathcal{F}^\pi(\pi+\nu)=\|P_\pi\nu\|_{\dot{H}^{-a}}^2,\tag{46}\] where the left-hand side denotes the natural quadratic extension to signed perturbations. The upper estimate in ?? follows because differentiation maps \(\dot{H}^{1-a}\) to \(\dot{H}^{-a}\), while multiplication by \(b_\pi\) is bounded on \(H^{1-a}\) under 1.

For the lower estimate, it suffices to prove that \[\|\nu\|_{\dot{H}^{1-a}} \le C\|P_\pi\nu\|_{\dot{H}^{-a}} \qquad\left(\int_{\mathbb{T}^d}\nu=0\right).\] Otherwise there are smooth zero-mass \(\nu_n\) with \(\|\nu_n\|_{\dot{H}^{1-a}}=1\) and \(P_\pi\nu_n\to0\) in \(\dot{H}^{-a}\). After passing to a subsequence, \(\nu_n\rightharpoonup\nu\) in \(H^{1-a}\). Multiplication by \(b_\pi\), viewed from \(H^{1-a}\) to \(H^{-a}\), is compact: it is bounded at the first regularity and is followed by the compact one-order Sobolev embedding. Hence \[\nabla\nu_n =P_\pi\nu_n+b_\pi\nu_n-\int_{\mathbb{T}^d}b_\pi\nu_n\] converges in \(H^{-a}\). Indeed, \(b_\pi\nu_n\to b_\pi\nu\) strongly there, its zero Fourier mode also converges, and distributional convergence identifies the resulting strong limit with \(\nabla\nu\). The zero-mass Poincaré equivalence \(\|\nu_n-\nu\|_{\dot{H}^{1-a}} \asymp\|\nabla(\nu_n-\nu)\|_{\dot{H}^{-a}}\) then gives strong convergence in \(\dot{H}^{1-a}\), so \(\|\nu\|_{\dot{H}^{1-a}}=1\) and \(P_\pi\nu=0\).

The last identity means that \[\nabla\nu-b_\pi\nu=c\] for a constant vector \(c\). Since \(b_\pi=\nabla\log\pi\), \[\pi\nabla\left(\frac{\nu}{\pi}\right)=c.\] Integrating each coordinate derivative on the torus gives \[0=c_j\int_{\mathbb{T}^d}\pi^{-1}\,\,\mathrm dx,\] and therefore \(c=0\). Thus \(\nu/\pi\) is constant; the zero-mass condition forces \(\nu=0\), a contradiction. This proves the lower estimate.

It remains to pass from smooth densities to the relaxed energy. Set \(\nu=\mu-\pi\). If \(\nu\in H^{1-a}\), then \(\mu_\varepsilon=p_\varepsilon*\mu\) is a smooth probability measure and \(\mu_\varepsilon-\pi\to\nu\) strongly in \(H^{1-a}\). Boundedness of \(P_\pi\) and 46 provide a recovery sequence whose Gram energies converge to \(\frac{1}{2}\|P_\pi\nu\|_{\dot{H}^{-a}}^2\). Conversely, if smooth probabilities \(\mu_n\rightharpoonup\mu\) have bounded Gram energy, the lower estimate just proved bounds \(\mu_n-\pi\) in \(H^{1-a}\). Every weak Sobolev limit agrees distributionally with \(\mu-\pi\), and weak lower semicontinuity gives the matching lower bound. Thus the relaxed energy equals \(\frac{1}{2}\|P_\pi(\mu-\pi)\|_{\dot{H}^{-a}}^2\) when \(\mu-\pi\in H^{1-a}\), and is \(+\infty\) otherwise. The same estimate appears in [16]. Finally, the lower bound shows that \(\mathcal{F}^\pi(\mu)=0\) only when \(\mu-\pi=0\), proving ?? . ◻

7.2 Proofs of the long-time conclusions↩︎

Proof of 3. Consider an arbitrary subsequence in \(N\). Compactness of \(\mathcal{P}(\mathcal{P}(\mathbb{T}^d))\) gives a further subsequence such that \(\Lambda_{N,T_N}\rightharpoonup\Lambda\). Apply 9 to \[Q^{N,T_N}:=\frac{1}{T_N}\int_0^{T_N}P_t^N\,\,\mathrm dt.\] The entropy identity gives \[A_N:=\int\mathcal{F}_N^\pi\,\,\mathrm dQ^{N,T_N} =\frac{H_N(0)-H_N(T_N)}{2T_N} \le\frac{H_N(0)}{2T_N}.\] Consequently, \[0\le\int\mathcal{F}^\pi(\mu)\,\,\mathrm d\Lambda(\mu) \le\liminf_NA_N \le\limsup_NA_N \le0.\] By 12, \(\mathcal{F}^\pi\) is nonnegative and vanishes only at \(\pi\), so \(\Lambda=\delta_\pi\). Every subsequential limit is the same; hence \(\Lambda_{N,T_N}\rightharpoonup\delta_\pi\).

The barycenter map on \(\mathcal{P}(\mathcal{P}(\mathbb{T}^d))\) is weakly continuous, and 45 identifies the barycenter of \(\Lambda_{N,T_N}\) with \(\overline{P}_{\mathrm{av},T_N}^{N}\). This proves 31 . ◻

Remark 13 (Fixed-horizon consequence). The same compactness argument also gives a fixed-horizon estimate. If \(\sup_NH_N(0)\le H_*\), define \[\omega_\pi(r):= \sup\left\{d_{\mathrm{BL}}(\mu,\pi): \mu\in\mathcal{P}(\mathbb{T}^d),\;\mathcal{F}^\pi(\mu)\le r\right\}.\] Lower semicontinuity, compactness, and ?? imply \(\omega_\pi(r)\downarrow0\) as \(r\downarrow0\). Applying 9 at fixed \(T\), followed by convexity of \(\mathcal{F}^\pi\) under barycenters, gives \[\limsup_{N\to\infty} d_{\mathrm{BL}}\!\left(\overline{P}_{\mathrm{av},T}^{N},\pi\right) \le\omega_\pi\!\left(\frac{H_*}{2T}\right).\]

Proof of 1. Since \(Q^N\ll\pi^{\otimes N}\) and the collision set is \(\pi^{\otimes N}\)-null, \(Q^N\) may be regarded as a law on \(\mathcal{D}_N\). Stationarity and 23 imply \[\int_{\mathcal{D}_N}\mathcal{F}_N^\pi\,\,\mathrm dQ^N=0.\] Let \(\Lambda_N\) be the law of the empirical measure under \(Q^N\). Along any convergent subsequence, 9 yields \[\int\mathcal{F}^\pi(\mu)\,\,\mathrm d\Lambda(\mu) \le\liminf_{N\to\infty}\int\mathcal{F}_N^\pi\,\,\mathrm dQ^N=0.\] Nonnegativity and ?? force \(\Lambda=\delta_\pi\). Thus every subsequential limit is \(\delta_\pi\). The barycenter of \(\Lambda_N\) is \(N^{-1}\sum_iQ_i^N\), so this label-averaged marginal converges to \(\pi\). If \(Q^N\) is exchangeable, the barycenter equals \(Q_1^N\). ◻

8 Quantitative renormalization below the logarithmic threshold↩︎

The qualitative correction \(c_N\to0\) follows from compactness and applies throughout \(0<\sigma<d\). For \(0<\sigma<2\), the target-dependent part of the Stein kernel can also be separated from a positive translation-invariant minorant. This gives the explicit rate in 2. Throughout this section, assume \[0<\sigma<2, \qquad s:=d-\sigma=2a-2\in(0,d).\] The operator factorization and Bessel minorant below apply throughout this range.

8.1 Normalized Stein operator↩︎

Let \(\mathsf L=(-\Delta)^{1/2}\) on mean-zero periodic distributions, with multiplier \(\lambda_\ell:=2\pi|\ell|\) for \(\ell\ne0\), and let \(\Pi\) denote the mean-zero projection, componentwise for vector fields. Set \[D_{b_\pi}:=\nabla-b_\pi, \qquad T_\pi:=D_{b_\pi}^*\mathsf L^{-2a}D_{b_\pi}.\] The inverse power is the mean-zero pseudoinverse. The Gram calculation in 5 can equivalently be written, for smooth signed \(f\), as \[\label{eq:operator-form} \iint G_\pi(x,y)f(x)f(y)\,\,\mathrm dx\,\,\mathrm dy =\left\langle D_{b_\pi}f,\mathsf L^{-2a}D_{b_\pi}f\right\rangle.\tag{47}\] For \(b_\pi=0\), the corresponding operator is \[T_0=\mathsf L^{2-2a}=\mathsf L^{-s}.\]

We normalize the principal part to the identity. On the mean-zero subspace \(L_0^2(\mathbb{T}^d)\), define \[\mathcal{R}:=\nabla\mathsf L^{-1}, \qquad \mathcal{C}:=\mathsf L^{-a}\Pi M_{b_\pi}\mathsf L^{a-1}.\] Here \(M_{b_\pi}\) denotes multiplication by the vector field \(b_\pi\). Since \(\mathcal{R}^*\mathcal{R}=I\) on \(L_0^2\), \[\label{eq:normalized-B} \mathcal{B} :=\mathsf L^{s/2}T_\pi\mathsf L^{s/2} =(\mathcal{R}-\mathcal{C})^*(\mathcal{R}-\mathcal{C}) =I+\mathcal{K}.\tag{48}\] The next lemma shows that the normalized target perturbation gains two derivatives as a quadratic form.

Lemma 5 (Two-order target perturbation). There is \(C_\pi<\infty\) such that \[\label{eq:two-order-form} |\langle u,\mathcal{K} v\rangle| \le C_\pi\|u\|_{H^{-1}}\|v\|_{H^{-1}} \qquad(u,v\in L_0^2(\mathbb{T}^d)).\tag{49}\]

Proof. Put \(\alpha=m-1\). The parameter assumptions give \[\alpha>\frac{d}{2}+1, \qquad \alpha>a,\] the second inequality following from \(a<1+d/2\). We first record the periodic commutator estimate \[\label{eq:commutator-estimate} \|[\mathsf L^a,b]g\|_{H^{1-a}} \le C\|b\|_{H^\alpha}\|g\|_2\tag{50}\] for each component \(b\) of \(b_\pi\). By skew-adjointness of the commutator, it is enough to prove \[\|[\mathsf L^a,b]h\|_2 \le C\|b\|_{H^\alpha}\|h\|_{H^{a-1}}.\] The periodic Kato–Ponce inequality, obtained from the Euclidean estimate by standard transference, gives \[\|[\mathsf L^a,b]h\|_2 \lesssim \|\nabla b\|_\infty\|h\|_{H^{a-1}} +\|\mathsf L^ab\|_{L^p}\|h\|_{L^q}, \qquad \frac{1}{p}+\frac{1}{q}=\frac{1}{2};\] see, for example, [22]. Write \(\delta=\alpha-a>0\). If \(0<\delta<d/2\), take \[\frac{1}{p}=\frac{1}{2}-\frac{\delta}{d}, \qquad \frac{1}{q}=\frac{\delta}{d}.\] If \(\delta>d/2\), take \(p=\infty\), \(q=2\). At the borderline \(\delta=d/2\), choose \(0<\varepsilon<a-1\), use \(H^{d/2}\hookrightarrow H^{d/2-\varepsilon}\hookrightarrow L^p\) with \(1/p=\varepsilon/d\), and take \(1/q=1/2-\varepsilon/d\). In every case, \(\alpha-1>d/2\) supplies the required embedding of \(H^{a-1}\) into \(L^q\), while \(H^\alpha\hookrightarrow W^{1,\infty}\). This proves 50 .

The mean-zero projection is important in the following exact decomposition: \[\label{eq:C-decomposition} \mathcal{C} =\Pi M_{b_\pi}\mathsf L^{-1}+\mathcal{E}, \qquad \mathcal{E} :=-\mathsf L^{-a}\Pi[\mathsf L^a,M_{b_\pi}]\mathsf L^{-1}.\tag{51}\] Indeed, this is just \(M_{b_\pi}\mathsf L^a =\mathsf L^aM_{b_\pi}-[\mathsf L^a,M_{b_\pi}]\), with the homogeneous powers interpreted as mean-zero pseudoinverses. Estimate 50 gives \[\mathcal{E}:H^{-1}\longrightarrow H^1.\] Sobolev multiplication by \(b_\pi\in H^\alpha\) also gives \[\mathcal{C}:H^{-1}\longrightarrow L^2.\]

Expanding 48 , \[\mathcal{K} =-\mathcal{R}^*\mathcal{C}-\mathcal{C}^*\mathcal{R} +\mathcal{C}^*\mathcal{C}.\] The nominal order-\(-1\) contributions in 51 cancel exactly: \[\label{eq:order-one-cancellation} -\mathcal{R}^*\Pi M_{b_\pi}\mathsf L^{-1} -(\Pi M_{b_\pi}\mathsf L^{-1})^*\mathcal{R} = \mathsf L^{-1}M_{\operatorname{div}b_\pi}\mathsf L^{-1}.\tag{52}\] The projections may be dropped in this calculation because the range of \(\mathcal{R}\) is mean zero and \(\mathcal{R}^*\) annihilates constants. Identity 52 follows componentwise from \[\partial_jM_{(b_\pi)_j}-M_{(b_\pi)_j}\partial_j =M_{\partial_j(b_\pi)_j}.\] Since \(\operatorname{div}b_\pi\in L^\infty\), the right-hand side maps \(H^{-1}\) to \(H^1\). Every remaining term contains \(\mathcal{E}\) or \(\mathcal{C}^*\mathcal{C}\), and the two mapping properties above give the same \(H^{-1}\to H^1\) bound. Duality yields 49 . ◻

Lemma 6 (Strict positivity of the normalized operator). There is \(\beta>0\), depending only on \(d,a,\pi\), such that \[\label{eq:B-positive} \mathcal{B}\ge\beta I \qquad\text{on }L_0^2(\mathbb{T}^d).\tag{53}\]

Proof. By 48 , \(\mathcal{B}\ge0\). The preceding lemma shows that \(\mathcal{K}\) maps \(H^{-1}\) to \(H^1\), hence is compact on \(L_0^2\). It remains to exclude a kernel.

Suppose \(\mathcal{B}u=0\) and set \(f=\mathsf L^{a-1}u\). Then \(\mathsf L^{-a}\Pi D_{b_\pi}f=0\), so \[D_{b_\pi}f=c\] for a constant vector \(c\). Since \[D_{b_\pi}f =\nabla f-b_\pi f =\pi\nabla\left(\frac{f}{\pi}\right),\] we have \(\nabla(f/\pi)=c/\pi\). Integrating each coordinate derivative on the torus gives \[0=c_j\int_{\mathbb{T}^d}\pi^{-1}\,\,\mathrm dx,\] and therefore \(c=0\). Thus \(f=c_0\pi\). But \(f\) has zero Lebesgue mean, whereas \(\int\pi=1\), so \(c_0=0\) and then \(u=0\). Hence \(\mathcal{B}=I+\mathcal{K}\) is injective. Its spectrum can accumulate only at \(1\), and 53 follows. ◻

8.2 A positive Bessel minorant↩︎

For \(M\ge1\), let \(J_M\) be the periodic Bessel kernel with Fourier coefficients \[\label{eq:Bessel-multiplier} \widehat J_M(\ell) =\bigl(M^2+\lambda_\ell^2\bigr)^{-s/2}, \qquad \ell\in\mathbb{Z}^d,\tag{54}\] where \(\lambda_0=0\). Its heat representation is \[\label{eq:Bessel-heat} J_M(z) =\frac{1}{\Gamma(s/2)} \int_0^\infty t^{s/2-1}e^{-M^2t}p_t(z)\,\,\mathrm dt.\tag{55}\] The periodic heat kernel is nonnegative, so \(J_M\ge0\). Moreover, \[J_M\in L^1(\mathbb{T}^d), \qquad \int_{\mathbb{T}^d}J_M\,\,\mathrm dx=M^{-s}.\]

Center this kernel with respect to \(\pi\): \[\label{eq:centered-Bessel} \overline{J}_M^\pi(x,y) :=J_M(x-y)-u_M(x)-u_M(y)+\gamma_M,\tag{56}\] where \[u_M:=J_M*\pi, \qquad \gamma_M:=\int_{\mathbb{T}^d}u_M\,\,\mathrm d\pi.\] Then \(\overline{J}_M^\pi\) is centered in each variable, \[0\le u_M(x)\le\|\pi\|_\infty M^{-s}, \qquad \gamma_M\ge0,\] and \(u_M\) is continuous because \(J_M\in L^1\) and \(\pi\) is continuous.

Proposition 14 (Bessel minorant). There is \(M_0<\infty\) such that, for every \(M\ge M_0\), \[\label{eq:Bessel-minorant} T_\pi-J_M*\ge0\qquad{(9)}\] as a quadratic form on smooth zero-mass functions.

Proof. Conjugating by \(\mathsf L^{s/2}\), the Bessel convolution operator becomes the multiplier \(\mathcal{J}_M\) with \[\widehat{\mathcal{J}_M}(\ell) =\chi_M(\ell) :=\left(\frac{\lambda_\ell^2}{M^2+\lambda_\ell^2}\right)^{s/2}, \qquad \ell\ne0.\] Thus \[\mathsf L^{s/2}(T_\pi-J_M*)\mathsf L^{s/2} =\mathcal{B}-\mathcal{J}_M =\mathcal{D}_M+\mathcal{K}, \qquad \mathcal{D}_M:=I-\mathcal{J}_M.\] The scalar inequality \[1-(1+x)^{-s/2}\ge c_s\frac{x}{1+x}, \qquad x\ge0,\] gives \[\label{eq:D-lower-symbol} 1-\chi_M(\ell) \ge c_s\frac{M^2}{M^2+\lambda_\ell^2}.\tag{57}\]

Let \(P_R\) project onto \(0<\lambda_\ell\le R\), let \(Q_R=I-P_R\), and write \(u=u_{\rm lo}+u_{\rm hi}\), where \(u_{\rm lo}=P_Ru\) and \(u_{\rm hi}=Q_Ru\). If \(M\ge R\ge1\), 57 implies \[\label{eq:D-high} \langle u_{\rm hi},\mathcal{D}_Mu_{\rm hi}\rangle \ge c_s'R^2\|u_{\rm hi}\|_{H^{-1}}^2.\tag{58}\] Indeed, \[\frac{M^2(1+\lambda^2)}{M^2+\lambda^2} \ge\frac{1+R^2}{2} \qquad(\lambda\ge R,\;M\ge R).\]

Choose \(R\) so large that \[\frac{C_\pi}{c_s'R^2}\le\frac{1}{4}, \qquad \frac{4C_\pi^2}{\beta c_s'R^2}\le\frac{1}{4}.\] Then 49 and 58 give \[\label{eq:Bessel-high} \langle u_{\rm hi},(\mathcal{D}_M+\mathcal{K})u_{\rm hi}\rangle \ge\frac{3}{4}\langle u_{\rm hi},\mathcal{D}_Mu_{\rm hi}\rangle.\tag{59}\] On the finite-dimensional low block, \(\mathcal{B}\ge\beta I\), while \[\|\mathcal{J}_MP_R\|_{2\to2} \le\left(\frac{R^2}{M^2+R^2}\right)^{s/2}\longrightarrow0.\] After increasing \(M_0=M_0(R)\), \[\label{eq:Bessel-low} \langle u_{\rm lo},(\mathcal{B}-\mathcal{J}_M)u_{\rm lo}\rangle \ge\frac{3\beta}{4}\|u_{\rm lo}\|_2^2.\tag{60}\] Because \(\mathcal{D}_M\) is diagonal in Fourier variables, the low–high cross term comes only from \(\mathcal{K}\). A final use of 49 , 58 , and Young’s inequality gives \[\label{eq:Bessel-cross} 2|\langle u_{\rm lo},\mathcal{K}u_{\rm hi}\rangle| \le\frac{\beta}{4}\|u_{\rm lo}\|_2^2 +\frac{1}{4}\langle u_{\rm hi},\mathcal{D}_Mu_{\rm hi}\rangle.\tag{61}\] Combining 5961 , \[\langle u,(\mathcal{B}-\mathcal{J}_M)u\rangle \ge\frac{\beta}{2}\|u_{\rm lo}\|_2^2 +\frac{1}{2}\langle u_{\rm hi},\mathcal{D}_Mu_{\rm hi}\rangle \ge0.\] Conjugating back proves ?? . ◻

8.3 A continuous positive-semidefinite remainder↩︎

In this subsection and the next, assume \(0<\sigma<2\).

For \(x\ne y\), define \[\label{eq:KM-definition} K_M^\pi(x,y):=G_\pi(x,y)-\overline{J}_M^\pi(x,y).\tag{62}\] Both terms are singular on the diagonal; the next lemma shows that their difference has a continuous extension.

Lemma 7 (Continuous remainder). For every \(M\ge M_0\), the kernel \(K_M^\pi\) extends continuously to \(\mathbb{T}^d\times\mathbb{T}^d\), is positive semidefinite, and satisfies \[\label{eq:KM-diagonal} \sup_{x\in\mathbb{T}^d}K_M^\pi(x,x)\le CM^\sigma.\tag{63}\]

Proof. Write \[G_\pi(x,y)=G_0(y-x)+R_\pi(x,y), \qquad G_0=-\Delta\kappa_a.\] Because \(\sigma<2\), \(\kappa_a\) is continuous at the origin. Moreover, \[|\nabla\kappa_a(z)|\lesssim1+|z|^{1-\sigma},\] and \(b_\pi\) is Lipschitz. Hence \[|(b_\pi(x)-b_\pi(y))\cdot\nabla\kappa_a(y-x)| \lesssim |x-y|+|x-y|^{2-\sigma}\longrightarrow0.\] The remaining term \(b_\pi(x)\cdot b_\pi(y)\kappa_a(y-x)\) is also continuous across the diagonal. Thus \(R_\pi\) extends continuously.

The Fourier coefficients of \(G_0-J_M\) are \(-M^{-s}\) at zero and, for \(\ell\ne0\), \[\lambda_\ell^{-s} -\bigl(M^2+\lambda_\ell^2\bigr)^{-s/2}.\] Uniformly in \(M\ge1\) and \(\ell\ne0\), \[\label{eq:Bessel-coefficient-bound} 0\le \lambda_\ell^{-s} -\bigl(M^2+\lambda_\ell^2\bigr)^{-s/2} \le C_s\min\{\lambda_\ell^{-s}, M^2\lambda_\ell^{-s-2}\}.\tag{64}\] Since \(s+2>d\) is equivalent to \(\sigma<2\), this sequence is absolutely summable. Therefore \(G_0-J_M\) is continuous. Splitting the sum at \(\lambda_\ell=M\) gives \[\label{eq:G0-J-diagonal} (G_0-J_M)(0) \le C\sum_{0<\lambda_\ell\le M}\lambda_\ell^{-s} +CM^2\sum_{\lambda_\ell>M}\lambda_\ell^{-s-2} \le CM^{d-s}=CM^\sigma.\tag{65}\]

Using 56 , the off-diagonal expression 62 is \[(G_0-J_M)(y-x)+R_\pi(x,y)+u_M(x)+u_M(y)-\gamma_M.\] Every term now has a continuous extension. The bounds on \(u_M,\gamma_M\), 65 , and boundedness of the extended \(R_\pi\) prove 63 .

It remains to prove positive semidefiniteness. Both \(G_\pi\) and \(\overline{J}_M^\pi\) are centered with respect to \(\pi\). For any finite signed measure \(\nu\), put \(\nu_0=\nu-\nu(\mathbb{T}^d)\pi\). Centering gives \[\iint K_M^\pi\,\,\mathrm d\nu\,\,\mathrm d\nu =\iint K_M^\pi\,\,\mathrm d\nu_0\,\,\mathrm d\nu_0.\] For smooth \(\nu_0\) of zero mass, the right-hand side is nonnegative by 14. Periodic mollification and continuity of \(K_M^\pi\) extend this inequality to every finite signed measure. Hence \(K_M^\pi\) is an ordinary continuous positive-semidefinite kernel. ◻

8.4 Completion of the quantitative estimate↩︎

Proof of 2 for \(0<\sigma<2\). Fix \(M\ge M_0\). Positive semidefiniteness and 63 give \[\sum_{i\ne j}K_M^\pi(x_i,x_j) \ge-\sum_iK_M^\pi(x_i,x_i) \ge-CNM^\sigma.\] Therefore \[\label{eq:KM-offdiagonal} \frac{1}{2N^2}\sum_{i\ne j}K_M^\pi(x_i,x_j) \ge-C\frac{M^\sigma}{N}.\tag{66}\]

For the centered Bessel kernel, \[\begin{align} \sum_{i\ne j}\overline{J}_M^\pi(x_i,x_j) ={}& \sum_{i\ne j}J_M(x_i-x_j) -2(N-1)\sum_i u_M(x_i)\\ &+N(N-1)\gamma_M. \end{align}\] The first and third terms are nonnegative, while \(u_M\le\|\pi\|_\infty M^{-s}\). Thus \[\label{eq:Bessel-offdiagonal} \frac{1}{2N^2}\sum_{i\ne j}\overline{J}_M^\pi(x_i,x_j) \ge-CM^{-s}.\tag{67}\] Since \(G_\pi=K_M^\pi+\overline{J}_M^\pi\) off the diagonal, 6667 yield \[\mathcal{F}_N^\pi(\boldsymbol{x}) \ge-C\left(\frac{M^\sigma}{N}+M^{-s}\right).\] Choose \(M=N^{1/d}\). Since \(s+\sigma=d\), \[\frac{M^\sigma}{N}=M^{-s}=N^{-1+\sigma/d}.\] For the finitely many \(N\) with \(N^{1/d}<M_0\), enlarge the constant. This proves 26 , and 27 follows from the definition of \(c_N\). ◻

Remark 15 (The endpoint obstruction). The cancellation above gains exactly two orders. At \(\sigma=2\), 64 has a high-frequency \(|\ell|^{-d}\) tail, so the residual diagonal diverges logarithmically and \(K_M^\pi\) is no longer a continuous finite-diagonal kernel. For \(\sigma>2\), the residual singularity is stronger. A further renormalization is therefore needed for \(2\le\sigma<d\).

9 Open directions↩︎

Three extensions remain particularly natural.

  1. The qualitative correction \(c_N\to0\) holds for the full range \(0<\sigma<d\), but an explicit rate in the range \(2\le\sigma<d\) remains open. At and above the endpoint, the post-cancellation remainder no longer has a finite diagonal. The ball-growth and smearing methods of [20], [21] control the translation-invariant principal Riesz interaction, but the target-dependent remainder is variable-coefficient and sign-indefinite. An explicit rate would require a corresponding commutator or pseudodifferential truncation estimate.

  2. A last-iterate result on logarithmic time scales would follow from a singular propagation-of-chaos estimate with controlled time dependence. Existing Riesz commutator estimates address the translation-invariant term \(-\nabla\kappa_a*(\mu^N-\rho)\) in their admissible range, but not the discrete trace generated by the target-weighted term \(\kappa_a*(b_\pi(\mu^N-\rho))\). Controlling that trace would connect the population decay of [16] to a particle last-iterate limit.

  3. It would be useful to admit deterministic atomic initialization. Such initial laws have infinite entropy relative to \(\pi^{\otimes N}\), so the joint-entropy argument would need an initial-layer estimate or a replacement that measures a finite defect.

An extension to \(\mathbb{R}^d\) is also natural, but it brings separate questions of tightness, confinement, and long-range control.

Acknowledgments↩︎

The authors used OpenAI’s ChatGPT during the preparation of this manuscript to assist with mathematical exploration, proof checking, and improvements to the language and readability. The authors reviewed and verified all resulting content and take full responsibility for the manuscript.

10 Periodic Riesz asymptotics↩︎

Lemma 8. Let \(1<a<1+d/2\) and \(\sigma=d+2-2a\). Near \(z=0\), the periodic kernel \(\kappa_a\) has the same singular part as the Euclidean Riesz potential of order \(2a\). In particular, \[\begin{align} -\Delta\kappa_a(z) &=c_{d,a}|z|^{-\sigma}+O(1),\tag{68}\\ -z\cdot\nabla\kappa_a(z) &\ge c_{d,a}'|z|^{2-\sigma}\tag{69} \end{align}\] for sufficiently small \(z\ne0\). Moreover, \[\kappa_a(z)= \begin{cases} c|z|^{2-\sigma}+O(1),&\sigma>2,\\ c\log(1/|z|)+h(z),\quad h\in C^\infty\text{ near }0,&\sigma=2,\\ \kappa_a(0)-c|z|^{2-\sigma}+o(|z|^{2-\sigma}),&\sigma<2, \end{cases}\] with constants whose signs are consistent with 68 . In addition, \[|\nabla^2\kappa_a(z)| \le C\bigl(1+\operatorname{dist}(z,0)^{-\sigma}\bigr) \qquad(z\ne0).\]

Proof. Choose a cutoff supported in a coordinate chart around the origin. Poisson summation, or equivalently the standard parametrix for the elliptic pseudodifferential operator \((-\Delta)^a\), decomposes the periodic fundamental solution into the Euclidean homogeneous fundamental solution plus a smooth periodic remainder. The Euclidean homogeneity is \(2a-d=2-\sigma\), with the logarithmic replacement at zero homogeneity. Applying \(-\Delta\) lowers the degree by two and gives degree \(-\sigma\). Radial differentiation gives 69 ; the coefficient is positive because \((-\Delta)^a\kappa_a=\delta_0-1\) and \(\widehat\kappa_a(\ell)>0\). Differentiating the homogeneous singular part twice, and using smoothness of the periodic remainder, gives the Hessian bound. ◻

10.1 The pair-energy gradient near collisions↩︎

The next estimate rules out cancellation of the singular pair forces inside a multiscale cluster. It is the input used in 2.

Lemma 9 (Closest-scale virial estimate). Fix \(N\ge2\) and \(2\le\sigma<d\). There are \(c,r_0>0\) such that \[|\nabla\mathcal{W}_N(\boldsymbol{x})| \ge c\,r(\boldsymbol{x})^{1-\sigma}, \qquad r(\boldsymbol{x}):=\min_{i\ne j}\operatorname{dist}(x_i,x_j),\] whenever \(0<r(\boldsymbol{x})<r_0\).

Proof. Argue by contradiction. Suppose that for a sequence \(\boldsymbol{x}^n\in\mathcal{D}_N\), with \(r_n:=r(\boldsymbol{x}^n)\to0\), \[r_n^{\sigma-1}|\nabla\mathcal{W}_N(\boldsymbol{x}^n)|\longrightarrow0.\] After passing to a subsequence, the pair realizing \(r_n\) is fixed and, for every pair \(i,j\), the ratio \(\operatorname{dist}(x_i^n,x_j^n)/r_n\) converges in \([1,\infty]\). Declare \(i\sim j\) when this limit is finite. The triangle inequality makes this an equivalence relation. Let \(I\) be the class containing the closest pair. Then \(|I|\ge2\), every internal distance is comparable to \(r_n\), and \[\label{eq:terminal-cluster-separation} \frac{\operatorname{dist}(x_i^n,x_k^n)}{r_n}\longrightarrow\infty \qquad(i\in I,\;k\notin I).\tag{70}\]

Choose local lifts of the points in \(I\), and write \[\bar x_I^n=\frac{1}{|I|}\sum_{i\in I}x_i^n, \qquad q_i^n=x_i^n-\bar x_I^n, \qquad R_n^2=\sum_{i\in I}|q_i^n|^2.\] The internal distance comparison gives \(R_n\asymp r_n\). Pairwise symmetrization yields \[\begin{align} \sum_{i\in I}q_i^n\cdot\nabla_{x_i}\mathcal{W}_N(\boldsymbol{x}^n) ={}& \sum_{\substack{i<j\\i,j\in I}} (x_j^n-x_i^n)\cdot \nabla\kappa_a(x_j^n-x_i^n) +E_n, \label{eq:pair-virial} \end{align}\tag{71}\] where \(E_n\) contains the interactions with particles outside \(I\). By 8, the internal sum is bounded above by \[-c\Phi_\sigma(r_n), \qquad \Phi_\sigma(r):= \begin{cases} 1,&\sigma=2,\\ r^{2-\sigma},&\sigma>2. \end{cases}\]

It remains to check that the exterior terms cannot cancel this contribution. The same local asymptotics give \[|\nabla^2\kappa_a(z)| \le C\bigl(1+\operatorname{dist}(z,0)^{-\sigma}\bigr).\] For \(k\notin I\), let \(D_{k,n}:=\min_{i\in I}\operatorname{dist}(x_i^n,x_k^n)\). Subtracting the value of the exterior force at \(\bar x_I^n\), whose contribution vanishes because \(\sum_iq_i^n=0\), gives \[|E_n| \le C R_n^2\sum_{k\notin I}\bigl(1+D_{k,n}^{-\sigma}\bigr).\] Here \(D_{k,n}/R_n\to\infty\), so for large \(n\) the interpolation segment from each \(x_i^n\) to \(\bar x_I^n\) remains at least \(D_{k,n}/2\) from \(x_k^n\), which justifies the displayed Hessian bound throughout the mean-value estimate. By 70 , this is \(o(\Phi_\sigma(r_n))\): for each exterior particle, the singular part divided by \(\Phi_\sigma(r_n)\) is \(O((r_n/D_{k,n})^\sigma)\), while the bounded part is \(O(r_n^\sigma)\).

Consequently, the absolute value of the left-hand side of 71 is at least \(c\Phi_\sigma(r_n)\) for large \(n\). Cauchy–Schwarz and \(R_n\asymp r_n\) now give \[|\nabla\mathcal{W}_N(\boldsymbol{x}^n)| \ge \frac{c\Phi_\sigma(r_n)}{R_n} \ge c r_n^{1-\sigma},\] contradicting the choice of the sequence. ◻

11 Population entropy dissipation for the singular kernel↩︎

We justify 13 without using pointwise reproduction or evaluating the singular kernel on the diagonal. Let \((\rho_t)_{t\in[0,T]}\) be a smooth, strictly positive solution of \[\label{eq:population-flow} \partial_t\rho_t+\nabla\cdot(\rho_t v_t)=0, \qquad v_t(x):= \int_{\mathbb{T}^d}\mathcal{A}_{\pi,y}k(y,x)\rho_t(y)\,\,\mathrm dy.\tag{72}\] Set \[q_t:=b_\pi\rho_t-\nabla\rho_t =-\rho_t\nabla\log\frac{\rho_t}{\pi}.\] Because \(\rho_t\) is smooth, the convolution of the distribution \(\kappa_a\) with \(q_t\) is smooth. Distributional integration by parts in the source variable gives \[v_t =\int_{\mathbb{T}^d}\mathcal{A}_{\pi,y}k(y,\cdot)\rho_t(y)\,\,\mathrm dy =\kappa_a*q_t.\] Differentiating the relative entropy along 72 and integrating by parts on the torus therefore yields \[\begin{align} \frac{\,\mathrm d}{\,\mathrm dt}\operatorname{Ent}(\rho_t\mid\pi) &=\int_{\mathbb{T}^d}\log\frac{\rho_t}{\pi}\,\partial_t\rho_t\,\,\mathrm dx\\ &=\int_{\mathbb{T}^d}\rho_t\nabla\log\frac{\rho_t}{\pi}\cdot v_t\,\,\mathrm dx\\ &=-\int_{\mathbb{T}^d}q_t\cdot(\kappa_a*q_t)\,\,\mathrm dx. \end{align}\] The Fourier identity established in the proof of 5 gives \[\int_{\mathbb{T}^d}q_t\cdot(\kappa_a*q_t)\,\,\mathrm dx =\left\| \int_{\mathbb{T}^d}\mathcal{A}_{\pi,y}k(y,\cdot)\,\,\mathrm d\rho_t(y) \right\|_{\mathcal{H}_{\kappa_a}^d}^{2} =\operatorname{KSD}_{\kappa_a}^2(\rho_t\mid\pi) =2\mathcal{F}^\pi(\rho_t).\] This proves 13 .

12 Entropy approximation↩︎

We give the localization argument behind 8. Fix \(N\) and a terminal time \(t>0\), let \(\Phi_s^N\) denote the particle flow, and write \[h_0:=\frac{\,\mathrm dP_0^N}{\,\mathrm d\pi^{\otimes N}}.\] For \(m\ge1\), set \[A_m:=\left\{\boldsymbol{x}: m^{-1}\le h_0(\boldsymbol{x})\le m, \quad \min_{0\le s\le t}\min_{i\ne j} \operatorname{dist}\bigl(\Phi_s^N(\boldsymbol{x})_i,\Phi_s^N(\boldsymbol{x})_j\bigr) \ge m^{-1} \right\}.\] The sets \(A_m\) increase. Moreover, \(P_0^N(\bigcup_mA_m)=1\): the density ratio is finite and positive \(P_0^N\)-almost surely, and 7 implies that the minimum pair distance along each trajectory is positive on the compact interval \([0,t]\). Put \(\alpha_m:=P_0^N(A_m)\) and, after discarding finitely many empty sets, define \[P_0^{N,m}:=\frac{P_0^N(\,\cdot\cap A_m)}{\alpha_m}, \qquad P_s^{N,m}:=(\Phi_s^N)_\#P_0^{N,m}.\]

On the trajectory tube generated by \(A_m\), the vector field and its first derivatives are bounded. The flow is a \(C^1\) one-to-one map there, so the area formula applies to the (not necessarily smooth) bounded density of \(P_0^{N,m}\). Equivalently, one may first smooth inside this tube and then remove the smoothing. The Jacobian satisfies \[\frac{\,\mathrm d}{\,\mathrm ds}\log\det D\Phi_s^N(\boldsymbol{x}) =\operatorname{div}_{\boldsymbol{x}}\boldsymbol{B}^N(\Phi_s^N(\boldsymbol{x})),\] whereas \[\frac{\,\mathrm d}{\,\mathrm ds}\log\pi^{\otimes N}(\Phi_s^N(\boldsymbol{x})) =\sum_{i=1}^Nb_\pi(\Phi_s^N(\boldsymbol{x})_i) \cdot B_i^N(\Phi_s^N(\boldsymbol{x})).\] Combining these two formulas with 39 gives \[\label{eq:localized-entropy} \operatorname{Ent}(P_t^{N,m}\mid\pi^{\otimes N}) +2N\int_0^t\mathbb{E}_{P_s^{N,m}}\mathcal{F}_N^\pi\,\,\mathrm ds =\operatorname{Ent}(P_0^{N,m}\mid\pi^{\otimes N}).\tag{73}\] This is the same calculation as in the smooth proof, now with every integration performed on a region uniformly separated from the singular set.

It remains to pass \(m\to\infty\). The local positivity of \(G_\pi\) and compactness away from the diagonal give a constant \(C_G\) such that \(G_\pi(x,y)\ge-C_G\) for \(x\ne y\). Hence the elementary shift \[\widehat{\mathcal{E}}_N^\pi:=\mathcal{F}_N^\pi+\frac{C_G}{2}\] is nonnegative for every configuration; the sharper Onsager estimate is not needed for this fixed-\(N\) approximation. Equation 73 becomes \[\label{eq:localized-renormalized-entropy} \operatorname{Ent}(P_t^{N,m}\mid\pi^{\otimes N}) +2N\int_0^t\mathbb{E}_{P_s^{N,m}}\widehat{\mathcal{E}}_N^\pi\,\,\mathrm ds =\operatorname{Ent}(P_0^{N,m}\mid\pi^{\otimes N}) +NC_Gt.\tag{74}\] Because \(\alpha_m\uparrow1\), conditioning on \(A_m\) gives \[\operatorname{Ent}(P_0^{N,m}\mid\pi^{\otimes N}) =\frac{1}{\alpha_m}\int_{A_m}h_0\log h_0\,\,\mathrm d\pi^{\otimes N} -\log\alpha_m \longrightarrow \operatorname{Ent}(P_0^N\mid\pi^{\otimes N}).\] The right-hand side of 74 is therefore bounded. Lower semicontinuity first shows that \(P_t^N=(\Phi_t^N)_\#P_0^N\) has finite entropy. Since the flow is injective, \(P_t^{N,m}\) is precisely \(P_t^N\) conditioned on \(\Phi_t^N(A_m)\). The same conditional-entropy formula at time \(t\) now gives \[\operatorname{Ent}(P_t^{N,m}\mid\pi^{\otimes N}) \longrightarrow\operatorname{Ent}(P_t^N\mid\pi^{\otimes N}).\] Finally, Tonelli’s theorem and monotone convergence apply to the energy: \[\int_0^t\mathbb{E}_{P_s^{N,m}}\widehat{\mathcal{E}}_N^\pi\,\,\mathrm ds =\frac{1}{\alpha_m}\int_{A_m}\int_0^t \widehat{\mathcal{E}}_N^\pi(\Phi_s^N(\boldsymbol{x}))\,\,\mathrm ds\,\,\mathrm dP_0^N(\boldsymbol{x}) \longrightarrow \int_0^t\mathbb{E}_{P_s^N}\widehat{\mathcal{E}}_N^\pi\,\,\mathrm ds.\] Passing to the limit in 74 and subtracting the deterministic correction \(NC_Gt\) proves the exact identity 23 .

References↩︎

[1]
Q. Liu and D. Wang, Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm,” in Advances in neural information processing systems, 2016, vol. 29, pp. 2378–2386.
[2]
Q. Liu, Stein Variational Gradient Descent as Gradient Flow,” in Advances in neural information processing systems, 2017, vol. 30.
[3]
A. Duncan, N. Nüsken, and Ł. Szpruch, On the Geometry of Stein Variational Gradient Descent,” Journal of Machine Learning Research, vol. 24, no. 56, pp. 1–39, 2023.
[4]
J. Lu, Y. Lu, and J. Nolen, Scaling Limit of the Stein Variational Gradient Descent: The Mean-Field Regime,” SIAM Journal on Mathematical Analysis, vol. 51, no. 2, pp. 648–671, 2019, doi: 10.1137/18M1187611.
[5]
J. A. Carrillo and J. Skrzeczkowski, arXiv:2312.16344Convergence and Stability Results for the Particle System in the Stein Gradient Descent Method.” 2023.
[6]
K. Chwialkowski, H. Strathmann, and A. Gretton, A Kernel Test of Goodness of Fit,” in Proceedings of the 33rd international conference on machine learning, 2016, vol. 48, pp. 2606–2615.
[7]
J. Gorham and L. Mackey, Measuring Sample Quality with Kernels,” in Proceedings of the 34th international conference on machine learning, 2017, vol. 70, pp. 1292–1301.
[8]
D. Wang, Z. Tang, C. Bajaj, and Q. Liu, Stein Variational Gradient Descent with Matrix-Valued Kernels,” in Advances in neural information processing systems, 2019, vol. 32, pp. 7836–7846.
[9]
L. Li, Y. Li, J.-G. Liu, Z. Liu, and J. Lu, A Stochastic Version of Stein Variational Gradient Descent for Efficient Sampling,” Communications in Applied Mathematics and Computational Science, vol. 15, no. 1, pp. 37–63, 2020, doi: 10.2140/camcos.2020.15.37.
[10]
S. Chewi, T. Le Gouic, C. Lu, T. Maunu, and P. Rigollet, arXiv:2006.02509SVGD as a Kernelized Wasserstein Gradient Flow of the Chi-Squared Divergence,” in Advances in neural information processing systems, 2020, vol. 33.
[11]
A. Korba, A. Salim, M. Arbel, G. Luise, and A. Gretton, A Non-Asymptotic Analysis for Stein Variational Gradient Descent,” in Advances in neural information processing systems, 2020, vol. 33, pp. 4672–4682.
[12]
N. Nüsken and D. R. M. Renger, Stein Variational Gradient Descent: Many-Particle and Long-Time Asymptotics,” Foundations of Data Science, 2023, doi: 10.3934/fods.2022023.
[13]
J. Shi and L. Mackey, A Finite-Particle Convergence Rate for Stein Variational Gradient Descent,” in Advances in neural information processing systems, 2023, vol. 36, pp. 26831–26844.
[14]
K. Balasubramanian, S. Banerjee, and P. Ghosal, Revised 2026; arXiv:2409.08469Improved Finite-Particle Convergence Rates for Stein Variational Gradient Descent.” 2024.
[15]
K. Balasubramanian, S. Banerjee, and A. Korba, arXiv:2607.00149Uniform-in-Time Propagation-of-Chaos for Stein Variational Gradient Descent.” 2026.
[16]
L. Chizat, M. Colombo, R. Colombo, and X. Fernández-Real, arXiv:2605.09456Quantitative Local Convergence of Mean-Field Stein Variational Gradient Flow.” 2026.
[17]
J. A. Carrillo, J. Skrzeczkowski, and J. Warnett, arXiv:2412.10295The Stein–Log-Sobolev Inequality and the Exponential Rate of Convergence for the Continuous Stein Variational Gradient Descent Method.” 2024.
[18]
J. A. Carrillo, J. Skrzeczkowski, and J. Warnett, arXiv:2605.03627Stein Variational Gradient Descent Dynamics for Highly Concentrated Kernels.” 2026.
[19]
S. Serfaty, Mean Field Limit for Coulomb-Type Flows,” Duke Mathematical Journal, vol. 169, no. 15, pp. 2887–2935, 2020, doi: 10.1215/00127094-2020-0019.
[20]
Q.-H. Nguyen, M. Rosenzweig, and S. Serfaty, Mean-Field Limits of Riesz-Type Singular Flows,” Ars Inveniendi Analytica, pp. Paper No. 4, 45, 2022, doi: 10.15781/nvv7-jy87.
[21]
M. Petrache and S. Serfaty, Next Order Asymptotics and Renormalized Energy for Riesz Interactions,” Journal of the Institute of Mathematics of Jussieu, vol. 16, no. 3, pp. 501–569, 2017, doi: 10.1017/S1474748015000201.
[22]
T. Kato and G. Ponce, Commutator Estimates and the Euler and Navier–Stokes Equations,” Communications on Pure and Applied Mathematics, vol. 41, no. 7, pp. 891–907, 1988, doi: 10.1002/cpa.3160410704.