No title


Abstract

Diffusion Flow Matching (DFM) has recently emerged as a versatile framework for generative modeling, yet its theoretical convergence properties remain only partially understood. In this work, we provide refined and novel convergence guarantees for Brownian motion based DFMs, focusing on the discretization error. Our analysis is conducted under the Kullback–Leibler (KL) divergence and the 2-Wasserstein distance. Under finite-moment conditions and a mild score integrability assumption, we derive KL convergence bounds with improved dimensional dependence compared to prior work, achieving, up to our knowledge, state-of-the-art scaling under minimal conditions. We further extend the analysis to the 2-Wasserstein distance: under an additional first-order score integrability assumption and a weak log-concavity condition, we obtain convergence guarantees with dimensional dependence consistent with the KL case.

1 Introduction↩︎

A central problem in machine learning and statistics is the generation of new samples from a target distribution that is only accessible through finite data. Generative modeling addresses this challenge by learning a mechanism that transforms a simple and tractable base distribution into the data distribution of interest. Deep generative models have been successfully applied in a wide range of domains such as image and video synthesis [1], [2], speech and audio generation [3], molecular and material design [4][6], and medical imaging [7], owing to their ability to learn and simulate complex data-generating processes from observations alone.

Among the broad class of generative approaches, continuous-time generative models have recently emerged as a particularly powerful and conceptually elegant paradigm. These models describe the transformation between probability distributions through differential equations or stochastic dynamics evolving over time, enabling fine-grained control over the generation process and providing a natural connection to tools from stochastic analysis and optimal transport. Examples include score-based generative models (SGMs) that leverage stochastic differential equations [1], [8], [9], and probability flow ordinary differential equations [10][13] that enable deterministic sampling.

Within this landscape, Flow Matching (FM) models have gained significant attention as a unifying framework for constructing continuous-time transports between probability distributions [14][20]. The key idea underlying FM is to define a coupling \(\pi\) between a base distribution \(\mu\) and a target distribution \(\nu^{\star}\), and to introduce an interpolating process, referred to as an interpolant, that bridges the two distributions over a finite time horizon. The dynamics of this interpolant induce a time-dependent velocity field that can be learned from samples, thereby specifying a transport map from \(\mu\) to \(\nu^{\star}\). When the interpolant follows deterministic dynamics, the resulting model corresponds to a deterministic FM formulation. Allowing for stochasticity in the interpolant leads instead to Diffusion FM (DFM) models. While FM and its stochastic formulation offer greater flexibility and robustness, they also introduce substantial technical challenges. In general, the interpolant does not define a Markov diffusion process and therefore cannot be directly characterized by a stochastic differential equation. To address this issue, DFM relies on the construction of a Markovian projection: a diffusion process that matches the marginal distributions of the interpolant at every time. The drift of this mimicking diffusion satisfies a regression identity and can be efficiently approximated using a neural network trained on samples from the interpolant. Once learned, the diffusion process can be simulated using standard numerical schemes, such as the Euler–Maruyama method, to generate samples from the learned distribution.

Despite promising empirical results [17] and its conceptual appeal, the theoretical foundations of Diffusion Flow Matching remain comparatively underdeveloped. Existing analyses are largely confined to deterministic FM models, while rigorous convergence guarantees for DFMs, particularly in terms of quantitative error bounds and distributional metrics, are still scarce. This gap motivates a deeper theoretical investigation of DFM models and their convergence properties.

Our contribution

In this work, we analyze DFM models based on \(d\)-dimensional Brownian bridge. We establish theoretical guarantees for convergence to the target distribution, both in Kullback–Leibler (\(\mathrm{KL}\)) divergence and in Wasserstein-2 (\(\mathscr{W}_2\)) distance, under standard and mild assumptions on the data. Our results extend and improve upon prior work by sharpening the dimensional dependence of the \(\mathrm{KL}\) guarantees and by establishing \(\mathscr{W}_2\) bounds accounting for all the sources of error. Our main contributions are summarized as follows:

  1. \(\mathrm{KL}\) convergence bounds.

    1. Without early stopping and constant step-size. We derive in 1 an improved explicit upper bound on the \(\mathrm{KL}\) divergence between the target distribution \(\nu^{\star}\) and the DFM output without early stopping, under standard assumptions. In particular, we suppose that the two marginals \(\mu\) and \(\nu^{\star}\) admit finite \(8\)-th moment (2) that the coupling \(\pi\) admits a score with finite \(8\)-th moment (3), and standard \(\mathrm{L}^2\) drift-approximation accuracy (1) under the Markovian projection. The resulting bound achieves \(\mathcal{O}(h)\) dependence on the time step size, while improving the dimensional scaling from \(\mathcal{O}(d^4)\) to \(\mathcal{O}(d^3)\) compared to prior works.

    2. With early stopping and constant step-size. By assuming the same moment conditions on \(\mu\) and \(\nu^{\star}\) and approximation error of the drift, replacing the condition 3 on \(\pi\) by the mild condition that the score of the conditional coupling \(\pi_{0|1}\) admits finite \(8\)-th order moment 4, we obtain in 2 an explicit bound on the \(\mathrm{KL}\) divergence between a smoothed target distribution and the early-stopped DFM output. Our bound preserves the \(\mathcal{O}(h)\) dependence and the improved \(\mathcal{O}(d^3)\) dimensional scaling of the non-early-stopped case. From this result, we deduce in 1 that choosing \(\delta = \mathcal{O}( \varepsilon^2/d)\) and \(h= \mathcal{O}(\varepsilon^{10}/d^{7})\) yields a \(\mathscr{W}_{2,\text{FM}}^2\)-error of order \(\mathcal{O}(\varepsilon^2)\).

    3. With early stopping and novel step-size schedule. In 3, we establish faster \(\mathrm{KL}\) convergence rates in the early stopping regime via a novel step-size schedule, while preserving the \(\mathcal{O}(d^3)\) and \(\mathcal{O}(h)\) dependences. The result relies only on the moment and drift approximation assumptions in [ass_moment,ass_drift_approx], together with the integrability condition on the score of the conditional coupling \(\pi_{0|1}\) (4). As a consequence, 2 yields improved complexity bounds in the Fortet–Mourier metric: choosing \(\delta = \mathcal{O}( \varepsilon^2/d)\), it suffices to take \(h= \tilde{\mathcal{O}}(\varepsilon^{2}/d^{3})\), where \(\tilde{\mathcal{O}}\) hides logarithmic factors in \(d\) and \(1/\varepsilon\), to guarantee a \(\mathscr{W}_{2,\text{FM}}^2\)-error of order \(\mathcal{O}(\varepsilon^2)\).

  2. \(\mathscr{W}_2\) convergence bounds.

    1. Without early stopping and constant step-size. We establish \(\mathscr{W}_2\) convergence bounds for DFMs in the non-early-stopped regime in 4, applicable to a broad class of distributions. Under appropriate \(\mathrm{L}^2\)-approximation error for the drift 5, moment conditions [ass_moment,ass_score], a weak log-concavity assumption on the coupling \(\pi\) (6) and an integrability condition on the Jacobian of the score associated with \(\pi\) (7), we derive bounds scaling as \(\mathcal{O}(\sqrt{h})\) in the time step and \(\mathcal{O}(\sqrt{d^3})\) in the dimension. These rates are consistent with the corresponding \(\mathrm{KL}\) guarantees. Furthermore, in 3, we show that our results apply in the case where \(\pi\) is the independent coupling, provided that the scores of the marginals \(\mu\) and \(\nu^{\star}\), together with their Jacobians, are integrable, and weakly concave (8).

1.0.0.1 Notation

Given a measurable space \((\mathsf{E}, \mathcal{E})\), we denote by \(\mathcal{P}(\mathsf{E})\) the set of probability measures of \(\mathsf{E}\). Also, given a topological space \((\mathsf{E}, \tau)\), we use \(\mathcal{B}(\mathsf{E})\) to denote the Borel \(\sigma\)-algebra on \(\mathsf{E}\). We denote by \(\mathbb{W}= \mathrm{C}([0,1],\mathbb{R}^d)\) the set of continuous functions from \([0,1]\) to \(\mathbb{R}^d\) and we refer to it as the Wiener space. We denote by \(\mathrm{Leb}^d\) the Lebesgue measure on \(\mathbb{R}^d\). Given two probability measures \(\mu, \nu \in \mathcal{P}(\mathbb{R}^d)\), we denote by \(\Pi(\mu,\nu)\) the set of couplings between \(\mu\) and \(\nu\), i.e., \(\xi\in\Pi(\mu,\nu)\) if and only if \(\xi\) is a probability measure on \(\mathbb{R}^d\times\mathbb{R}^d\) and \(\xi(\mathsf{A}\times \mathbb{R}^d) = \mu(\mathsf{A})\) and \(\xi(\mathbb{R}^d\times \mathsf{A}) = \nu(\mathsf{A})\) for all measurable \(\mathsf{A}\subseteq \mathbb{R}^d\). The relative entropy (or \(\mathrm{KL}\)-divergence) of \(\mu\) with respect to \(\nu\) is defined by \(\mathrm{KL}(\mu |\nu) := \int \log (\mathrm{d}\mu /\mathrm{d}\nu) \mathrm{d}\mu\) if \(\mu\) is absolutely continuous with respect to \(\nu\), and \(\mathrm{KL}(\mu |\nu) := +\infty\) otherwise. If \(\mu\) and \(\nu\) have finite second moment, the \(2-\)Wasserstein distance is defined by \(\mathscr{W}_2^2 (\mu, \nu) := \inf_{\xi \in \Pi(\mu,\nu)} \int \norm{x-y}^2 \mathrm{d}\xi(x,y)\) and the Fortet-Mourier distance of order \(2\) is defined by \(\mathscr{W}_{2,\text{FM}}^2(\mu,\nu)=\inf_{\pi\in\Pi(\mu,\nu)}\int \min\{\|x-y\|^2, 1\} \mathrm{d}\pi (x,y).\) Given \(\xi\in\mathcal{P}(\mathbb{R}^d)\) and \(p\ge 1\), we denote by \(\|\cdot\|_{\mathrm{L}^p(\xi)}:=(\int\|\cdot\|^p \mathrm{d}\xi)^{1/p}\) the \(\mathrm{L}^p\)-norm with respect to \(\xi\). Given two real numbers \(u,v\in \mathbb{R}\), we write \(u\lesssim v\) (resp. \(u \gtrsim v\)) to mean \(u\le C v\) (resp. \(u\ge C v\)) for a universal constant \(C>0\). Also, we denote by \(\norm{x}\) the Euclidean norm of \(x\in\mathbb{R}^d\), by \(\langle x, y\rangle\) the scalar product between \(x,y \in\mathbb{R}^d\), and by \(x^{\operatorname{T}}\) the transpose of \(x\). Last, we use standard Big-\(\mathcal{O}\) notation.

2 Diffusion Flow Matching↩︎

In this section, we provide a concise yet self-contained overview of Brownian motion based Diffusion Flow Matching (DFM) models, following the probabilistic formulation introduced in recent works. Recall that \(\nu^{\star}\in \mathcal{P}(\mathbb{R}^d)\) denote the target (or data) distribution and \(\mu \in \mathcal{P}(\mathbb{R}^d)\) denote the base (or prior) distribution.

As already highlighted, DFM is a procedure that designs a generative model able to produce new samples approximately distributed according to \(\nu^{\star}\), by learning a stochastic transport from \(\mu\) to \(\nu^{\star}\) over a finite time horizon. To this end, DFM proceeds as follows:

  1. Markovian projection. First, DFM constructs the Markovian projection \((X^{\mathrm M}_t)_{t\in\left[0,1\right]}\) of an interpolated process \((X^{\mathrm I}_t)_{t\in\left[0,1\right]}\), such that \(X^{\mathrm M}_0\) and \(X^{\mathrm M}_1\) have distributions \(\mu\) and \(\nu^\star\), respectively. The process \((X^{\mathrm M}_t)_{t\in\left[0,1\right]}\) is defined as the solution to a stochastic differential equation (SDE).

  2. Approximation of the Markovian projection. While \((X^{\mathrm M}_t)_{t\in\left[0,1\right]}\) would yield an exact generative model mapping samples from \(\mu\) to samples from \(\nu^\star\) if the associated SDE could be solved exactly, this is generally infeasible in practice. Consequently, suitable numerical or modeling approximations must be introduced.

We detail these two crucial stages in what follows.

2.0.0.1 Markovian projection.

The core objects at the basis of this construction is a coupling \(\pi\) between \(\mu\) and \(\nu^{\star}\), i.e., a probability measure on the product space \(\mathbb{R}^{2d}\) such that, for any \(\mathsf{A}\in\mathcal{B}(\mathbb{R}^d)\), \(\pi(\mathsf{A}\times \mathbb{R}^d) = \mu(\mathsf{A})\) and \(\pi(\mathbb{R}^d\times \mathsf{A})= \nu^{\star}(\mathsf{A})\), and the transition of the Brownian bridge process that we now recall. It is well-known that if \((B_t)_{t \in [0,1]}\) is a Brownian motion starting from a distribution \(\mu\in\mathcal{P}(\mathbb{R}^d)\), then it induces a Markov kernel \(\mathrm{b}\mathbb{B}\) which corresponds to a regular version of the conditional distribution of the path \((B_t)_{t\in[0,1]}\) given \(B_0\) and \(B_1\), i.e., \[\mathbb{P}[(B_t)_{t\in[0,1]} \in \mathsf{A}| (B_0,B_1)] = \mathrm{b}\mathbb{B}((B_0,B_1),\mathsf{A}) \;,\] for any \(\mathsf{A}\in \mathcal{B}(\mathbb{W})\). The existence of this kernel is guaranteed, for instance, by Theorem 8.37 in [21]. Then, by integrating these bridges against the coupling \(\pi\), one defines the stochastic interpolant \[\begin{align} \label{def:stochastic95interpolant} \inter{\pi,\mathbb{B}}(\mathsf{A}) = \int \mathrm{b}\mathbb{B}((x_0,x_1),\mathsf{A})\, \pi(\mathrm{d}x_0,\mathrm{d}x_1)\;, \end{align}\tag{1}\] for any \(\mathsf{A}\in \mathcal{B}(\mathbb{W})\). This construction corresponds to the law of a stochastic process \((X_t^{\rm I})_{t\in[0,1]}\), which defines a random path connecting the two marginals: by construction, the interpolant satisfies \(X^{\rm I}_0 \sim \mu\) and \(X^{\rm I}_1 \sim \nu^{\star}\).

While \((X_t^{\rm I})_{t\in[0,1]}\) provides a conceptually meaningful interpolation between \(\mu\) and \(\nu^{\star}\), it is generally non-Markovian. Indeed, its dynamics depend on the terminal value \(X^{\rm I}_1\) and satisfies \[\label{eq:sde95stochastic95interpolant} \mathrm{d}X^{\rm I}_t = 2 \nabla_x \log p_{1-t}(X^{\rm I}_1 \mid X^{\rm I}_t)\mathrm{d}t + \sqrt{2}\mathrm{d}W_t\;,\quad t\in [0,1]\;,\tag{2}\] with \(X^{\rm I}_0\sim \mu.\) Here \((W_t)_{t\in [0,1]}\) denotes a \(d\)-dimensional Brownian motion and \((y,x) \mapsto p_{s}(y|x)\) denotes the heat kernel, i.e., for any \(x,y\in\mathbb{R}^d\) and \(0<s\le 1\) \[\label{def95heat95kernel} p_{s}(y|x) = \frac{1}{(4\pi s)^{d/2}} \exp \Big(-\frac{\norm{y-x}^2}{4s}\Big)\;\;.\tag{3}\] We refer to Section 2.1 of [22] for details.

To obtain a tractable generative model, one constructs a Markov diffusion \((X_t^{\rm M})_{t\in[0,1]}\) that shares the same time marginals as the interpolant. This is achieved via the Markovian projection [23], [24]: \(X^{\rm M}_0 \sim \mu\) and \[\label{sde:markovian95proj} \mathrm{d}X_t^{\rm M} = \tilde{\beta}_t(X_t^{\rm M})\mathrm{d}t + \sqrt{2}\mathrm{d}W_t\;, \quad t \in [0,1]\;,\tag{4}\] where the mimicking drift is given by the conditional expectation \[\begin{align} \label{drift95as95conditional95expectation} \tilde{\beta}_t(x) &= \mathbb{E}\big[ 2 \nabla_x \log p_{1-t}(X^{\rm I}_1 \mid X^{\rm I}_t) \,\big|\, X^{\rm I}_t = x \big]\\ &=\mathbb{E}\left[ \frac{X^\mathrm{I}_1-X^\mathrm{I}_t}{1-t} \Big| X_t^{\rm I}=x\right]\;. \end{align}\tag{5}\] The resulting process satisfies \[\begin{align} \label{equality95in95law95of95marginals} X^{\rm I}_t \stackrel{\text{dist}}{=}X^{\rm M}_t \;, \end{align}\tag{6}\] for all \(t \in [0,1]\) [22], [25], and therefore constitutes an ideal diffusion model transporting \(\mu\) to \(\nu^{\star}\).

While the Markovian projection provides an ideal generative model, it is not directly accessible. First, the mimicking drift defined in 5 is intractable and even using some approximation, the continuous time SDE 4 cannot be simulated exactly. As a result, turning this construction into a workable model requires overcoming these practical limitations.

2.0.0.2 Approximation of the Markovian projection.

The first challenge in approximating the Markovian projection is that of approximating the mimicking drift \(\tilde{\beta}_t\). Since, for any \(t\in [0,1]\), \(\tilde{\beta}_t\) writes as a conditional expectation 5 , then, by Corollary 8.17 in [21] and 6 , \(\tilde{\beta}_t\) solves the regression problem: \[\arg\min_{f}\mathbb{E}\Big[ \norm{ f(t,X^{\rm M}_{t}) - \tilde{\beta}_{t}(X^{\rm M}_{t}) }^2 \Big]\;.\] Therefore, for any \(t\in [0,1]\), \(\tilde{\beta}_t\) can be estimated using a family of neural networks \(\{x \mapsto s_\theta(t,x)\}_{\theta \in \Theta}\) and minimizing the \(\mathrm{L}^2\) regression loss \[\label{minimization95problem} \theta \mapsto \mathbb{E}\Big[ \norm{ s_\theta(t,X^{\rm M}_{t}) - \tilde{\beta}_{t}(X^{\rm M}_{t}) }^2 \Big]\;.\tag{7}\]

The trained network \(s_{\theta^\star}(t,\cdot)\) then serves as proxy of \(\tilde{\beta}_{t}\). However, the resulting diffusion which reads: \[\begin{align} \mathrm{d}X^{\mathrm{NN}}_t=s_{\theta^\star}(t, X^{\mathrm{NN}}_t)\mathrm{d}t+ \sqrt{2} \mathrm{d}W_t\;, \quad t\in [0,1]\;, \end{align}\]with \(X^{\mathrm{NN}}_0\sim \mu\), cannot be solved in general and we have to rely on numerical discretization. We make use of the Euler-Maruyama scheme, i.e., for a choice of sequence of step sizes \(\{h_k\}_{k=1}^N\), \(N \geqslant 1\), and the corresponding time discretization \(t_k= \sum_{i=1}^{k} h_i\), such that \(t_0=0\) and \(t_N = 1\), we define the continuous process \((X^{\theta^{\star}}_t)_{t \in [0,1]}\), recursively on the intervals \([t_k, t_{k+1}]\) by \[\label{model} \mathrm{d}X^{\theta^{\star}}_t=s_{\theta^\star}(t_k, X^{\theta^{\star}}_{t_k})\mathrm{d}t +\sqrt{2} \mathrm{d}W_t\;, \quad t\in [t_k,t_{k+1}]\;,\tag{8}\]

with \(X^{\theta^{\star}}_0\sim \mu\). Finally, this dynamics is used to define the generative model associated with the DFM procedure. In particular, approximate samples from \(\nu^{\star}\) are obtained by simulating trajectories of \((X^{\theta^{\star}}_t)_{t\le 1-\delta}\) with \(\delta \geqslant 0\). When \(\delta=0\), the dynamics is run until its terminal time and no early stopping is performed. In contrast, choosing \(\delta>0\) results in an early stopping of the dynamics, which may be beneficial in practice to mitigate numerical instabilities near the terminal time.

In short, DFM generates samples by simulating an Euler–Maruyama discretization of a neural-network approximation of the Markovian projection associated with the stochastic interpolant 8 . The simulation is run up to time \(t = 1 - \delta\), for some \(\delta \ge 0\).

3 Main Results↩︎

In this section, we present our main theoretical contributions: non-asymptotic convergence guarantees for Brownian-motion-based Diffusion DFM 8 . We analyze convergence in two metrics: Kullback–Leibler divergence and Wasserstein-2 distance. Our results are established under mild moment and regularity conditions on the distributions \(\mu\), \(\nu^{\star}\), and the coupling \(\pi\), together with standard \(\mathrm{L}^2\) drift-approximation assumptions common in the literature on diffusion and score-based generative models [22], [26][28].

3.1 Convergence in Kullback–Leibler Divergence↩︎

We first focus on convergence in \(\mathrm{KL}\) divergence, which provides a strong notion of similarity between distributions and is widely used in theoretical analyses of diffusion models [25], [27], [29]. We state next two assumptions that we will be made in all of our results.

As is standard, we first assume that the learned drift approximates the true drift with \(\varepsilon^2\) accuracy in \(\mathrm{L}^2\) norm.

Assumption 1 (Drift approximation). There exists \(\theta^\star\in\Theta\) and \(\varepsilon^2>0\) such that \[\sum_{k=0}^{N-1} h_{k+1} \mathbb{E}\big[\|s_{\theta^\star}(t_k, X^{\mathrm{M}}_{t_k}) - \tilde{\beta}_{t_k}(X^{\mathrm{M}}_{t_k})\|^2 \big] \le \varepsilon^2 \;.\]

Furthermore, we impose the following moment conditions. For \(m\in\mathbb{N}\), \(\zeta\in\mathcal{P}(\mathbb{R}^m)\), and \(p\ge1\), we denote by \(\mathbf{m}_{p}[{\zeta}] = \int \|x\|^p \mathrm{d}\zeta(x)\) the \(p\)-th moment. When \(\zeta\) has a density with respect to \(\mathrm{Leb}^m\), we identify it with the density itself.

Assumption 2 (Moment condition). It holds \(\mathbf{m}_{8}[{\mu}] + \mathbf{m}_{8}[{\nu^{\star}}]<+\infty\).

3.1.0.1 KL convergence without early stopping and constant step-size.

In the full regime setting, we also impose standard integrability conditions on the score of the coupling, which are routinely assumed in analyses of diffusion-based generative models [22], [29].

Assumption 3 (Score integrability of the coupling). The coupling \(\pi\) is absolutely continuous with respect to \(\mathrm{Leb}^{2d}\) with positive density, \(\log\pi\) is \(\mathrm{C}^1\), and \(\|\nabla \log \pi \|_{\mathrm{L}^8(\pi)}<+\infty.\)

Under these assumptions, we can quantify how close the law of the generated samples is to the target \(\nu^{\star}\) in \(\mathrm{KL}\) divergence.

Theorem 1. Let \(\{t_k\}_{k=0}^{N_h}\) be a uniform partition of \([0,1]\) with step size \(h=1/N_h>0\). Under [ass_drift_approx,ass_moment,ass_score], if \(\mathbf{m}_{8}[{\mu}], \mathbf{m}_{8}[{\nu^{\star}}]\lesssim d^4\), denoting by \(\nu^{\theta^\star}_1\) the law of \(X^{\theta^{\star}}_1\), we have \[\begin{align} \label{convergence95bound95no95early} \mathrm{KL}(\nu^{\star}| \nu^{\theta^\star}_1)&\lesssim \varepsilon^2 +h(h^{1/8}+1) \left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right)d\;. \end{align}\tag{9}\]

The bound provided in 1 has two additive contributions: the first term, \(\varepsilon^2\), corresponds to the drift-approximation error; while the second term accounts for the discretization error and scales as \(\mathcal{O}(h)\). Importantly, the bound scales cubically with the dimension \(d\), improving upon previous \(d^4\)-scaling results [22]. We refer to 4 for a detailed literature comparison. Note that, to obtain a \(\mathrm{KL}\) error of order \(\mathcal{O}(\varepsilon^2)\) and a computational complexity \(\mathcal{O}(\varepsilon^{-2})\), one needs \[\begin{align} N_h& = \frac{1}{\varepsilon^2}\left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right)d \;, \end{align}\] discretization points.

Remark 1. For simplicity, we set the time horizon to \(T=1\). For an arbitrary \(T>0\), the bound in 1 remains valid with a Brownian reference on \([0,T]\), up to a multiplicative factor \(\max\{1,T^8\}\) in the discretization error.

3.1.0.2 KL convergence with early stopping and constant step-size.

If the process is stopped before reaching \(t=1\), we can relax 3 and refine the bound.

More precisely, let \(\pi_{0|1}\) denote a regular conditional distribution of \(X^{\mathrm{I}}_0\) given \(X^{\mathrm{I}}_1\). We assume the following integrability condition.

Assumption 4 (Score integrability of the conditional coupling). The conditional coupling \(\pi_{0|1}\) is absolutely continuous with respect to \(\mathrm{Leb}^d\), with strictly positive density. Moreover, \(\log \pi_{0|1}\) is \(\mathrm{C}^1\) and satisfies \(\|\nabla \log \pi_{0|1} \|_{\mathrm{L}^8(\pi_{0|1} )}<+\infty.\)

We remark that 4 is substantially weaker than 3. For instance, when the coupling is chosen to be independent, that is \(\pi=\mu\otimes\nu^{\star}\), the conditional distribution \(\pi_{0|1}\) coincides with the prior distribution \(\mu\), and 4 reduces to an integrability condition on the score of \(\mu\): \(\|\nabla \log \mu \|_{\mathrm{L}^8(\mu)} < +\infty.\) Consequently, the only assumption involving the target distribution \(\nu^{\star}\) is the moment condition 2.

Under this milder assumption, we are able to derive the following result.

Theorem 2. Fix \(0<\delta<1/2\). Let \(\{t_k\}_{k=0}^{N_h}\) be a uniform partition of \([0,1]\) with step size \(h=1/N_h>0\). Under [ass_drift_approx,ass_moment,ass_score_conditioned], if \(\mathbf{m}_{8}[{\mu}], \mathbf{m}_{8}[{\nu^{\star}}]\lesssim d^4\), then, for \(\nu^{\star}_{1-\delta}\) and \(\nu^{\theta^\star}_{1-\delta}\) denoting the laws of \(X^{\mathrm{M}}_{1-\delta}\) and \(X^{\theta^{\star}}_{1-\delta}\), we have \[\begin{align} \label{convergence95bound95early} \mathrm{KL}(\nu^{\star}_{1-\delta} | \nu^{\theta^\star}_{1-\delta})&\lesssim \varepsilon^2+ h (h^{1/8}+1) \\ &\quad \cdot\Bigg(\frac{d^2}{\delta^4} + \|\nabla\log \pi_{0|1}\|_{\mathrm{L}^8(\pi_{0|1})}^4 \Bigg)d\;. \end{align}\tag{10}\]

Consistently with Theorem 1, the dependence on the dimension remains cubic in \(d\). As a consequence, our result improves upon Theorem 3 in [22]. In this setting, achieving a \(\mathrm{KL}\) error of order \(\mathcal{O}(\varepsilon^2)\) with computational complexity \(\mathcal{O}(\varepsilon^{-2})\) is ensured by choosing \[\begin{align} N_h&=\frac{1}{\varepsilon^2}\Bigg(\frac{d^2}{\delta^4} + \|\nabla\log \pi_{0|1}\|_{\mathrm{L}^8(\pi_{0|1})}^4 \Bigg)d\;. \end{align}\] An extensive comparison with the related literature is provided in 4.

Corollary 1. Fix \(\delta = \mathcal{O}( \varepsilon^2/d)\) and \(h= \mathcal{O}(\varepsilon^{10}/d^{7})\). Let \(\{t_k\}_{k=0}^{N_h}\) be a uniform partition of \([0,1]\) with step size \(h=1/N_h>0\). Under [ass_drift_approx,ass_moment,ass_score_conditioned], if \(\mathbf{m}_{8}[{\mu}], \mathbf{m}_{8}[{\nu^{\star}}]\lesssim d^4\), then, for \(\nu^{\theta^\star}_{1-\delta}\) denoting the law of \(X^{\theta^{\star}}_{1-\delta}\), we have \[\mathscr{W}_{2,\text{FM}}^2(\nu^{\star}|\nu^{\theta^{\star}}_{1-\delta} )\lesssim \mathcal{O}(\varepsilon^2)\;.\]

This result ensures a computational complexity \(\mathcal{O}(\varepsilon^{-10}).\)

3.1.0.3 KL convergence with early stopping and novel step-size schedule.

In the early stopping regime, we show that an appropriately designed step-size schedule yields an accelerated rate of convergence in \(\mathrm{KL}\). Our result holds under [ass_moment,ass_drift_approx,ass_score_conditioned].

Theorem 3. Fix \(0<\delta<1/2\). Let \(\{t_k\}_{k=0}^{M_h+N}\) be a partition of \([0,1]\) with step sizes \(\{h_k\}_{k=1}^{M_h+N}\) such that \(h_k=h=1/(2M_h)\) for \(k\le M_h\) and \(h_{k}=h\min\{t_k, 1-t_k\}\) for \(M_h<k\le M_h+N\). Under [ass_moment,ass_drift_approx,ass_score_conditioned], denoting by \(\nu^{\star}_{1-\delta}\) and \(\nu^{\theta^\star}_{1-\delta}\) the laws of \(X^{\mathrm{M}}_{1-\delta}\) and \(X^{\theta^{\star}}_{1-\delta}\), we have \[\begin{align} \label{czqfkmbs} \mathrm{KL}(\nu^{\star}_{1-\delta} | \nu^{\theta^\star}_{1-\delta})&\lesssim \varepsilon^2+ h d^3\log\frac{1}{\delta}+h(h^{1/8}+1)\\ &\quad \cdot\Big(d^2+ \|\nabla\log \pi_{0|1}\|_{\mathrm{L}^8(\pi_{0|1})}^4 \Big)d\;. \end{align}\tag{11}\]

Note that, to obtain a \(\mathrm{KL}\) error of order \(\mathcal{O}(\varepsilon^2)\) and a computational complexity \(\mathcal{O}(\varepsilon^{-2})\), one needs \[\begin{align} M_h& = \frac{1}{2\varepsilon^2}\left[\Big(d^2+ \|\nabla\log \pi_{0|1}\|_{\mathrm{L}^8(\pi_{0|1})}^4 \Big)d +d^3\log\frac{1}{\delta}\right], \end{align}\] and \(N=2 M_h \log(1/\delta)\).

Corollary 2. Fix \(\delta = \mathcal{O}( \varepsilon^2/d)\) and \(h= \tilde{\mathcal{O}}(\varepsilon^{2}/d^{3})\), where \(\tilde{\mathcal{O}}\) hides logarithmic factors in \(d\) and \(1/\varepsilon\). Let \(\{t_k\}_{k=0}^{M_h+N}\) be a partition of \([0,1]\) with step sizes \(\{h_k\}_{k=1}^{M_h+N}\) such that \(h_k=h\) for \(k\le M_h\) and \(h_{k}=h\min\{t_k, 1-t_k\}\) for \(M_h<k\le M_h+N\). Under [ass_moment,ass_drift_approx,ass_score_conditioned], if \(\mathbf{m}_{8}[{\mu}], \mathbf{m}_{8}[{\nu^{\star}}]\lesssim d^4\), then for \(\nu^{\theta^\star}_{1-\delta}\) denoting the law of \(X^{\theta^{\star}}_{1-\delta}\), we have \[\mathscr{W}_{2,\text{FM}}^2(\nu^{\star},\nu^{\theta^{\star}}_{1-\delta} )\lesssim \mathcal{O}(\varepsilon^2)\;.\]

This result ensures a computational complexity \(\mathcal{O}(\varepsilon^{-2}).\)

3.2 Convergence in Wasserstein-2 Distance↩︎

We now turn our attention to convergence in Wasserstein-2 distance, which is a natural metric for evaluating the quality of generated samples in generative modeling [30][32].

In this section, we introduce a modified \(\mathrm{L}^2\) drift approximation condition, tailored to the Wasserstein setting:

Assumption 5 (Drift approximation for Wasserstein analysis). There exist \(\theta^\star\) and \(\varepsilon>0\) such that \[\sum_{k=0}^{N-1} h_{k+1} \mathbb{E}\Big[\|s_{\theta^\star}(t_k,X^{\theta^{\star}}_{t_k})-\tilde{\beta}_{t_k}(X^{\theta^{\star}}_{t_k})\|^2 \Big]^{1/2} \le \varepsilon\;.\]

Assumptions of this form are standard in the analysis of SGMs under Wasserstein metrics (see, e.g., [28], [30][33]). An important feature of 5 is that the expectation is taken over the trajectory of the generative process \((X^{\theta^{\star}}_{t_k})_{k=0}^{N}\).

In addition, we assume that \(\pi\) satisfies a weak log-concavity condition in the sense of [34], together with suitable integrability assumptions on the Jacobian of its associated score function.

Let \(\beta:\mathbb{R}^d\to \mathbb{R}\) be a smooth potential function. Its weak convexity profile is defined as \[\begin{align} \label{eq:def:weak95convexity95profile} \kappa_\beta(r) = \inf\left\{ \frac{\langle \nabla \beta(x) - \nabla \beta(y), x-y\rangle}{\|x-y\|^2} : \|x-y\|=r \right\}\;, \end{align}\tag{12}\] for any \(r>0\). This function quantifies non-uniform convexity lower bounds, allowing one to weaken classical convexity conditions. Furthermore, for any \(M\ge 0\) and \(r>0\), define \[f_M(r) = 2\sqrt{M}\,\tanh\big(r\sqrt{M}/2\big)\;\;.\]

The weak convexity profile 12 is widely used in the analyses of diffusion-based generative models. Recent results establish state-of-the-art \(\mathrm{KL}\) and \(\mathscr{W}_2\) convergence guarantees for these models under such weak regularity conditions [25], [30], [35], [36].

We assume that \(\pi\) is weakly log-concave:

Assumption 6 (Weak log-concavity of the coupling). \(\pi\in\mathcal{P}(\mathbb{R}^d)\) is absolutely continuous with respect to \(\mathrm{Leb}^{2d}\), \(\log\zeta\) is \(\mathcal{C}^1\), and satisfies for some \(\alpha_{\pi}>0,M_{\pi}\ge0\) and any \(r>0,\) \[\begin{align} \kappa_{-\log\pi}(r)\ge \alpha_{\pi}-r^{-1}f_{M_{\pi}}(r)\;. \end{align}\]

Strongly log-concave distributions satisfy 6 as a special case choosing \(M_{\zeta} =0\). Furthermore, perturbations of strongly log-concave distributions are weakly log-concave: for instance, if \(-\log \xi = V + W\) with \(V\) strongly convex and \(W\) smooth with Lipschitz gradient, then \(\xi\) is weakly log-concave. This includes classical double-well potentials, which are widely used in physics and statistics. Finally, it is well-known that Gaussian mixtures satisfy 6.
Moreover, we assume that the Jacobian of the score function associated to \(\pi\) is integrable:

Assumption 7 (First-order score integrability of the coupling). The coupling \(\pi\) is absolutely continuous with respect to \(\mathrm{Leb}^{2d}\) with positive density, \(\log\pi\) is \(\mathrm{C}^2\), and \[\begin{align} \|\nabla^2 \log \pi \|_{\mathrm{L}^2(\pi)}&:=\left(\int_{\mathbb{R}^{2d}}\left\| \nabla^2 \log \frac{\mathrm{d}\pi}{\mathrm{d}\mathrm{Leb}^{2d}}\right\|^2_{\mathrm{op}} \,\mathrm{d}\pi \right)^{1/2}\\ &<+\infty\;. \end{align}\]

Note that a Lipschitz score function implies 7.

3.2.0.1 \(\mathscr{W}_2\) convergence bound without early stopping and constant step-size.

Under the conditions above, we can establish precise quantitative bounds on the \(\mathscr{W}_2\) distance between the target distribution \(\nu^{\star}\) and the distribution of the generated samples.

Theorem 4. Let \(\{t_k\}_{k=0}^{N_h}\) be a uniform partition of \([0,1]\) with step size \(h=1/N_h>0\). Under [ass_drift_approx_wass,ass_moment,ass_score,ass_regularity,ass_hessian], if \(\mathbf{m}_{8}[{\mu}], \mathbf{m}_{8}[{\nu^{\star}}]\lesssim d^4\), then, denoting by \(\nu^{\theta^\star}_1\) the law of \(X^{\theta^{\star}}_1\), we have \[\begin{align} \label{wass95weak95convergence95bound} &\mathscr{W}_2(\nu^{\star},\nu^{\theta^\star}_1)\\ &\lesssim \mathrm{C}\left(\varepsilon+\sqrt{h}(h^\frac{1}{16}+1) \sqrt{\left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right)d}\right)\;. \end{align}\tag{13}\] where \[\begin{align} \label{def:constant95in95weak95bound} \mathrm{C} &=\exp\left( \frac{8\sqrt{2}}{\sqrt{\alpha_{\pi}}}\exp\left(\frac{M_\pi}{\alpha_\pi}\right)\|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}\right) \;. \end{align}\tag{14}\]

The bound provided in 4 naturally decomposes into two additive contributions: one coming from the drift-approximation error, \(\varepsilon\), and one from the discretization error, which, consistently with the \(\mathrm{KL}\) setting, exhibits a \(\mathcal{O}(\sqrt{h})\) dependence on the time step and a \(\mathcal{O}(\sqrt{d^3})\) dependence on the space dimension. Furthermore, an implication of 4 is that choosing the number of discretization steps \(N_h = d^3/\varepsilon^2\) guarantees a \(\mathscr{W}_2\) error of order \(\mathcal{O}(\varepsilon)\), with overall computational complexity scaling as \(\mathcal{O}(\varepsilon^{-1})\).
We can specify our results to the case where \(\pi = \mu \otimes \nu^{\star}\). In this case, it is sufficient that the marginals \(\mu\) and \(\nu^{\star}\) are weakly log-concave in the sense of [34], and that their scores, together with the corresponding Jacobians, are integrable, to guarantee convergence.

Assumption 8. The marginals \(\mu\) and \(\nu^{\star}\) are absolutely continuous with respect to \(\mathrm{Leb}^{d}\) with positive densities, \(\log\mu, \log\nu^{\star}\) are \(\mathrm{C}^2\), and \[\begin{align} &\|\nabla \log \mu \|_{\mathrm{L}^8(\mu)}+ \|\nabla\log\nu^{\star}\|_{\mathrm{L}^8(\nu^{\star})}< +\infty\;,\\ &\|\nabla^2 \log \mu \|_{\mathrm{L}^2(\mu)}+ \|\nabla^2\log\nu^{\star}\|_{\mathrm{L}^2(\nu^{\star})}< +\infty\;. \end{align}\] Moreover, they satisfy for some \(\alpha_\mu, \alpha_{\nu^{\star}}>0\), \(M_\mu, M_{\nu^{\star}}\ge0\) and any \(r\ge 0\), \[\begin{align} &\kappa_{-\log\mu}(r)\ge \alpha_{\mu}-r^{-1}f_{M_{\mu}}(r)\;,\\ &\kappa_{-\log\nu^{\star}}(r)\ge \alpha_{\nu^{\star}}-r^{-1}f_{M_{\nu^{\star}}}(r)\;. \end{align}\]

Under the above conditions, [ass_score,ass_hessian,ass_regularity] hold:

Lemma 1. Under 8, [ass_score,ass_regularity,ass_hessian] hold true for \(\pi = \mu \otimes \nu^{\star}\). In particular, \[\begin{align} &\|\nabla\log\pi\|_{\mathrm{L}^8(\pi)}\lesssim \|\nabla\log\mu\|_{\mathrm{L}^8(\mu)}+\|\nabla\log\nu^{\star}\|_{\mathrm{L}^8(\nu^{\star})}\\ &\|\nabla^2\log\pi\|_{\mathrm{L}^2(\pi)}\lesssim \|\nabla^2\log\mu\|_{\mathrm{L}^2(\mu)}+\|\nabla^2\log\nu^{\star}\|_{\mathrm{L}^2(\nu^{\star})}\;, \end{align}\]and \(\pi\) is weakly log-concave with parameters \[\begin{align} \alpha_\pi= \min\{\alpha_\mu, \alpha_{\nu^{\star}}\}\;, \quad M_\pi=2\max\{M_\mu, M_{\nu^{\star}}\}. \end{align}\]

We then automatically deduce the following corollary.

Corollary 3. Let \(\{t_k\}_{k=0}^{N_h}\) be a uniform partition of \([0,1]\) with step size \(h=1/N_h>0\). Under [ass_drift_approx_wass,ass_moment,ass_indep_coupling], choosing \(\pi = \mu \otimes \nu^{\star}\) and assuming \(\mathbf{m}_{8}[{\mu}], \mathbf{m}_{8}[{\nu^{\star}}]\lesssim d^4\), the bound in 4 holds with \(\|\nabla\log\pi\|_{\mathrm{L}^8(\pi)}\), \(\|\nabla^2\log\pi\|_{\mathrm{L}^2(\pi)}\), \(\alpha_\pi\) and \(M_\pi\) as in 1.

4 Related works and comparison with existing literature↩︎

The theoretical understanding of DFMs remains limited, despite encouraging empirical results [17]. While SGMs have been extensively studied in terms of both practical performance and theoretical properties [27], [29], [30], [37], rigorous convergence guarantees for DFM models are relatively recent and still leave room for improvement.

4.0.0.1 KL Convergence Bounds.

Non-early stopping. In this setting, DFM has been analyzed in [22], where the authors derived non-asymptotic KL convergence bounds (Theorem 2). Their analysis accounts for all practical sources of error, including drift estimation and time discretization. The bounds require moment conditions on the marginals (Assumption H1), their scores (Assumption H2(i)), and the score of the coupling (Assumption H2(ii)). The drift estimator is assumed to approximate the true drift in \(\mathrm{L}^2\) norm, without any further smoothness assumptions. The resulting rates scale with the dimension as \(d^4\). These results provide the first explicit KL guarantees for non-early-stopped DFM that incorporate all sources of error. In 1, we improve upon Theorem 2 in [22] by relaxing the integrability conditions on the scores of the marginals (Assumption H2(i)) and by reducing the dimension dependence to \(d^3\).
Early stopping. The same work [22] also considers KL convergence under early stopping. In Theorem 3, the authors establish explicit bounds on the KL divergence between a smoothed target distribution and the output of an early-stopped DFM, assuming an independent coupling \(\pi = \mu \otimes \nu^{\star}\), moment conditions on \(\mu\) and \(\nu^{\star}\) (Assumption H1), and moment assumptions on the score of \(\mu\). The resulting bounds again scale as \(d^4\) and, to the best of our knowledge, represent the only existing KL guarantees for early-stopped DFM. In 2, we consider a more general set of assumptions, which reduces to those of Theorem 3 in [22] when \(\pi\) is chosen as the independent coupling, and we obtain a \(\mathcal{O}(d^3)\) dimensional scaling. Consequently, 2 improves upon Theorem 3 of [22] both in terms of assumptions and dimensional dependence. Moreover, using a tailored step-size schedule and under the same assumptions as 2, we obtain in 3 accelerated \(\mathrm{KL}\) convergence in the early-stopping regime, while maintaining \(\mathcal{O}(d^3)\) complexity. From this result, we derive in 2 improved complexity bounds in the Fortet–Mourier metric: \(\delta = \mathcal{O}(\varepsilon^2/d)\) and \(h=\tilde{\mathcal{O}}(\varepsilon^2/d^3)\) yield a \(\mathcal{W}_{2,\mathrm{FM}}^2\)-error of order \(O(\varepsilon^2)\).
Two-sided early stopping. More recently, [38] studied \(\mathrm{KL}\) convergence for DFMs built on a more general class of stochastic interpolants, allowing for bridge processes beyond the classical Brownian bridge. Their analysis assumes finite eighth-order moments for \(\mu\) and \(\nu^\star\), as well as bounded variation of the time derivative of the interpolation function (Assumption 4.1), together with a standard \(\mathrm{L}^2\) approximation error for the learned drift. However, their convergence guarantees (Theorem 4.3) depend crucially on an early stopping procedure applied not only at the terminal time \(t=1\), but also at the initial time \(t=0\). Consequently, the \(\mathrm{KL}\) bound decomposes into three additive terms. The first term, of order \(\varepsilon^2\), corresponds to the drift estimation error. The second term, \(\mathrm{KL}(\mathcal{L}(X^\mathrm{M}_{\delta}) \,|\, \hat{\mu})\), with \(\delta>0\) and \(\hat{\mu}\in\mathcal{P}(\mathbb{R}^d)\) denoting the initialization point of the model 8 , captures the initialization error: it is strictly positive, analytically intractable, and cannot be controlled in practice. The third term accounts for the discretization error and scales as \(\mathcal{O}(d^3)\). Notably, the discretization error diverges as the process approaches the prior distribution. This divergence necessitates truncating the dynamics away from \(t=0\), which in turn induces the non-vanishing initialization error. When specializing their bounds to the Brownian bridge setting (Section 5), the same structural limitations persist: the initialization error remains non-zero and the discretization error diverges when approaching the prior. As a result, their method forfeits one of the main advantages of DFMs over SGMs, namely the ability to generate samples without any intrinsic initialization bias.

4.0.0.2 Wasserstein-2 Convergence Bounds.

Non-early stopping. Convergence in the \(2\)-Wasserstein distance has been investigated in [39] and [40]. In [39], the authors establish \(2\)-Wasserstein convergence guarantees for pre-trained diffusion models via Lagrangian and Eulerian distillation losses, thereby controlling the Wasserstein gap between the teacher and the student models. However, their analysis requires the velocity field to satisfy a one-sided Lipschitz condition in space with a time-dependent Lipschitz constant, and it does not account for time-discretization errors. Theorem 3.15 in [40] provides bounds on the \(2\)-Wasserstein distance between the generative and target distributions that explicitly incorporate all sources of error and scale as \(\mathcal{O}(\sqrt{d})\). It is important to emphasize that these guarantees are obtained under the assumption of a Gaussian prior distribution, a setting in which such strong results are to be expected and which is therefore rather restrictive. Moreover, their analysis relies on Gaussian tail assumptions on the target distribution (Assumption 3.7) and on regularity conditions on the learned velocity field (Assumption 3.13). In contrast, in 4 we establish non-asymptotic and fully quantitative convergence guarantees in \(\mathscr{W}_2\) distance that simultaneously capture both drift-approximation and time-discretization errors. These results hold for arbitrary prior distributions and rely only on mild moment and integrability conditions ([ass_moment,ass_score,ass_hessian]) on the coupling \(\pi\), its associated score, and the Jacobian of the score, as well as weak convexity assumptions on the log-density of \(\pi\) (6). Furthermore, our analysis does not require any a priori smoothness assumptions on either the drift or its estimator; instead, the necessary regularity is derived directly from [ass_regularity,ass_hessian]. In 3, we further specialize our bounds to the case of the independent coupling, showing that the same guarantees are obtained by imposing analogous assumptions on the marginals, rather than on the joint coupling (8).

Early stopping. The same work [40] also establishes \(\mathscr{W}_2\)-convergence guarantees under an early stopping procedure (Theorem 3.19). While these bounds account for all sources of error and scale as \(\mathcal{O}(\sqrt{d})\), they remain valid only for Gaussian priors and compactly supported target distributions (Assumption 3.17). Moreover, the analysis continues to rely on regularity conditions on the estimator (Assumption 3.13). Therefore, the resulting \(\mathscr{W}_2\) bounds exhibit the same structural limitations as in the non-early-stopped setting.

5 Conclusion↩︎

In this work, we conduct a thorough theoretical analysis of a DFM model built upon a \(d\)-dimensional Brownian bridge. We establish dimension-improved \(\mathrm{KL}\) convergence bounds and provide \(\mathscr{W}_2\) convergence guarantees, under mild and standard assumptions. Despite these advances, several avenues remain open for improvement. In particular, it would be highly valuable to relax the integrability requirements on the score functions further; to achieve sharper dependence on the space dimension and to undertake a statistical analysis of DFMs.

Acknowledgments and Disclosure of Funding↩︎

The work of Marta Gentiloni-Silveri has been supported by the Paris Ile-de-France Région in the framework of DIM AI4IDF. Alain Durmus has received funding from the Fondation de l’École polytechnique as part of its “Servir la science” campaign. The work of Alain Durmus is supported by the France 2030 program with the reference ANR-25-PEIA-0001 (THEOREM project); by Hi! Paris and Agence Nationale de la Recherche (Grant 11-LABX-0047) and by the European Union (ERC-2022-SYG-OCEAN-101071601). Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Council Executive Agency. Neither the European Union nor the granting authority can be held responsible for them.

Impact Statement↩︎

This paper presents work whose goal is to advance the field of machine learning. There are many potential societal consequences of our work, none of which we feel must be specifically highlighted here.

6 Preliminaries↩︎

6.0.0.1 Extra Notation

Given a matrix \(\mathbf{A}\in \mathbb{R}^{n\times s}\), we denote by \(\norm{\mathbf{A}}_{\mathrm{op}}\) the operator norm of \(\mathbf{A}\). For \(f :\left[0,1\right]\times \mathbb{R}^d \to \mathbb{R}\) regular enough, we denote by \(\nabla_x f(t,x), \nabla^2_x f(t,x)\) and \(\Delta_x f(t,x)\) respectively the gradient, hessian and laplacian of \(f\), defined for \(t,x\in [0,1]\times \mathbb{R}^d\) by \(\nabla_x f(t,x):= (\partial_{x_i}f(t,x))_i\), \(\nabla^2_x f(t,x):=(\partial_{x_i}\partial_{x_j} f(t,x))_{i,j}\), \(\Delta f(t,x):=\sum_{i=1}^{d} \partial^2_{x_i}f(t,x)\), where \(\partial_{x_j}\) denotes the partial derivative with respect to the \(j\)-th variable. For \(F: \left[0,1\right]\times \mathbb{R}^d\to \mathbb{R}^d\) regular enough, we denote by \(D_x F\), \(\operatorname{div}_x F\) and \(\Delta_x F\) respectively, the Jacobian matrix, the divergence and the vectorial laplacian of \(F\), defined for \(t,x \in \left[0,1\right] \times \mathbb{R}^d\) by \(D_x F(t,x) = (\partial_{x_j}F_i(t,x))_{i,j}\), \(\operatorname{div}_x F(t,x):= \sum_{j=1}^{d} \partial_{x_j} F_j(t,x)\), \(\Delta_x F(t,x)= (\Delta_x F_1 (t,x),..., \Delta_x F_d (t,x))\).
We remark that the stochastic interpolant admits the simple closed form \[\label{interpolant95in95linear95form} X^{\mathrm{I}}_t \stackrel{\text{dist}}{=}(1-t)X^{\mathrm{I}}_0+tX^{\mathrm{I}}_1+\sqrt{2t(1-t)}\mathrm{Z}\;, \quad \mathrm{Z}\sim \mathrm{N}(0, \operatorname{Id})\;,\tag{15}\] Denote by \((p_t^{\mathrm{I}})_{t\in\left[0,1\right]}\) the time marginal densities of \((X_t^{\mathrm{I}})_{t\in\left[0,1\right]}\) with respect to the Lebesgue measure. Note that, as a consequence of the very definition 1 of the stochastic interpolant, they write as \[\label{marginals95interpolant} p^\mathrm{I}_t(x)=\int_{\mathbb{R}^{2d}} p_{t}(x|x_0) p_{1-t}( x_1|x) \tilde{\pi}(\mathrm{d}x_0, \mathrm{d}x_1)\;,\tag{16}\] with \[\begin{align} \label{def:tilde95pi} \tilde{\pi}(\mathrm{d}x_0, \mathrm{d}x_1)=\frac{ \pi(\mathrm{d}x_0, \mathrm{d}x_1)}{p_1(x_1|x_0)}\;. \end{align}\tag{17}\]

Also, denote by \(\mathrm{K}_{t}\) the regular kernel associated with the conditional distribution of \((X^\mathrm{I}_0, X^\mathrm{I}_1)\) given \(X^\mathrm{I}_t\), i.e., the map \(\mathrm{K}_{t}:\mathbb{R}^d\times\mathcal{B}(\mathbb{R}^{2d})\rightarrow [0,1]\) such that

  1. \(y\mapsto \mathrm{K}_{t}(y, \mathsf{A})\) is measurable for any \(\mathsf{A}\in \mathcal{B}(\mathbb{R}^{2d})\);

  2. \(\mathsf{A}\mapsto \mathrm{K}_{t}(y, \mathsf{A})\) is a probability measure for any \(y\in \mathbb{R}^d\);

  3. almost surely it holds \[\begin{align} \mathrm{K}_{t}(X^\mathrm{I}_t, \mathsf{A}):=\mathbb{E}\left[\mathbb{1}_{\mathsf{A}}(X^\mathrm{I}_0, X^\mathrm{I}_T)|X^\mathrm{I}_t\right]\;, \quad \mathsf{A}\in \mathcal{B}(\mathbb{R}^{2d})\;. \end{align}\]

It follows again from 1 that \(\mathrm{K}_{t}\) admits a transition density with respect to \(\mathrm{Leb}^{2d}\) given by \[\begin{align} \label{def:pi95t95x} \mathrm{k}_{t}(x_0,x_1|x_t)=p_{t|(0,1)}(x_t|x_0,x_1)(p^\mathrm{I}_{t}(x_t))^{-1}\pi(x_0,x_1)\;, \end{align}\tag{18}\] with \(p_{t|(0,1)}\) denoting the conditional density of \(B_t\) given \((B_0, B_1)\), i.e., \[\begin{align} p_{t|(0,1)}(x_0,x_1|x_t)=\frac{p_{1-t}(x_1|x_t)p_t(x_t|x_0)}{p_1(x_1|x_0)}=\exp\left(-\frac{\|x_t-(1-t)x_0-t x_1\|^2}{4t(1-t)}\right)\;. \end{align}\] Denote also by \(\mathrm{K}_{t}^{\#0}\) and \(\mathrm{K}_{t}^{\#1}\) its first and second marginal respectively and by \(\mathrm{k}_{t}^{\#0}:=\mathrm{d}\mathrm{K}_{t}^{\#0}/\mathrm{d}\mathrm{Leb}^d,\) and \(\mathrm{k}_{t}^{\#1}:=\mathrm{d}\mathrm{K}_{t}^{\#1}/\mathrm{d}\mathrm{Leb}^d\) the associated densities. We observe that, with the introduced notation, for any \(t\in [0,1)\) and \(y_t\in\mathbb{R}^d\), we have that \[\begin{align} \label{eq:representation95of95drift95as95conditional95expectation95of95linear95function} \tilde{\beta}_t(x_t)=\int \frac{x_1-x_t}{1-t}\mathrm{k}_t(\mathrm{d}x_0, \mathrm{d}x_1|x_t)=\int \frac{x_1-x_t}{1-t}\mathrm{k}^{\#1}_t(\mathrm{d}x_1|x_t)\;. \end{align}\tag{19}\]

6.1 On the heat kernel↩︎

It is well-known that \((s,x,y)\mapsto p_s(y|x)\) defined in 3 is twice continuously differentiable in the space variables \(x\) and \(y\) and satisfies for \(x,y\in \mathbb{R}^d, s\in (0,1],\) \[\label{sec:simmetry95nabla95heat95kernel} \nabla_x p_s(y|x)= -\frac{x-y}{2s} p_s(y|x)=-\nabla_y p_s(y|x)\;,\tag{20}\] \[\label{simmetry95hessian95heat95kernel} \nabla^2_x p_s(y|x)= -\frac{1}{2s} p_s(y|x)\operatorname{Id}+ \frac{(x-y)(x-y)^{\operatorname{T}}}{4s^2} p_s(y|x)=\nabla^2_y p_s(y|x)\;,\tag{21}\] \[\label{simmetry95delta95heat95kernel} \Delta_x p_s(y|x)= -\frac{d}{2s} p_s(y|x)+ \norm{\frac{x-y}{2s}}^2 p_s(y|x)=\Delta_y p_s(y|x)\;.\tag{22}\] Moreover 3 satisfies the heat equation, i.e., \[\label{fp95eq95heat95kernel} \partial_s p_s(y|x)=\Delta_x p_s(y|x)\;, \quad s\in (0,1]\;, \;x,y\in \mathbb{R}^d\;.\tag{23}\] Thus, in particular, \((s,x,y)\mapsto p_s(y|x)\) is continuously differentiable in the time variable \(s\).

6.2 On score’s integrability↩︎

Lemma 2. Under 3, for any \(p\in \{2,4,8\},\) it holds \[\begin{align} \label{eq:bound95score95marginals} \|\nabla \log\mu\|_{\mathrm{L}^p(\mu)}^p\;,\|\nabla \log\nu^{\star}\|_{\mathrm{L}^p(\nu^{\star})}^p\lesssim \|\nabla \log\pi\|_{\mathrm{L}^p(\pi)}^p\;. \end{align}\tag{24}\] Moreover, if we assume also 2, it holds \[\begin{align} \label{eq:bound95score95tilde95pi} \|\nabla \log\tilde{\pi}\|_{\mathrm{L}^p(\pi)}\le \|\nabla \log\pi\|_{\mathrm{L}^p(\pi)}+\sqrt[p]{\mathbf{m}_{p}[{\mu}]}+\sqrt[p]{\mathbf{m}_{p}[{\nu^{\star}}]}\;. \end{align}\tag{25}\]

Proof of 2:. We start with 24 . We only deal with \(\mu\), as the argument for \(\nu^{\star}\) is the very same. We have that \[\begin{align} \nabla \log\mu(x_0)&= \frac{\nabla \mu(x_0)}{\mu(x_0)}= \frac{ \int \nabla_{x_0} \pi(x_0,x_1)\mathrm{d}x_1}{\int \pi(x_0,x_1)\mathrm{d}x_1} = \frac{\int\pi(x_0,x_1)\nabla_{x_0}\log\pi(x_0,x_1)\mathrm{d}x_1}{\int \pi(x_0,x_1)\mathrm{d}x_1}\\ &=\mathbb{E}\left[\nabla_{x_0}\log\pi(X^\mathrm{I}_0,X^\mathrm{I}_1)|X^\mathrm{I}_0=x_0\right]\;. \end{align}\] Jensen inequality yields that \[\begin{align} \|\nabla \log\mu(x_0)\|^p\le \mathbb{E}\left[\left\|\nabla_{x_0}\log\pi(X^\mathrm{I}_0,X^\mathrm{I}_1)\right\|^p|X^\mathrm{I}_0=x_0\right]\;. \end{align}\] Taking expectation and using the properties of conditional expectation, we get that \[\begin{align} \|\nabla \log\mu\|_{\mathrm{L}^p(\mu)}^p&=\mathbb{E}\left[\|\nabla \log\mu(X^\mathrm{I}_0)\|^p\right]\le \mathbb{E}\left[ \mathbb{E}\left[\left\|\nabla_{x_0}\log\pi(X^\mathrm{I}_0,X^\mathrm{I}_1)\right\|^p|X^\mathrm{I}_0\right]\right]\\ &= \mathbb{E}\left[ \left\|\nabla_{x_0}\log\pi(X^\mathrm{I}_0,X^\mathrm{I}_1)\right\|^p\right]=\|\nabla_{x_0} \log\pi\|_{\mathrm{L}^p(\pi)}^p\le \|\nabla \log\pi\|_{\mathrm{L}^p(\pi)}^p\;. \end{align}\] The bound in 25 is a direct consequence of the definitions 17 of \(\tilde{\pi}\) and 3 of \(p_1(x_1|x_0)\). ◻

6.3 On the Stochastic Interpolant↩︎

We collect here several results on Stochastic Interpolants that will be used later. 3 reports the result of Lemma 1 in [22] and is a direct consequence of 15 . [lemma:lemma2_back_in_dfm,lemma:lemma2_for_in_dfm] are taken from Lemma 2 in [22], while [lemma:lemma3_back_in_dfm,lemma:lemma3_for_in_dfm] follow from Lemma 3 in [22] and our 2. We refer to [22] for proofs.

Lemma 3. For any \(p\ge 1\), they hold \[\mathbb{E}[\norm{X^{\mathrm{I}}_s-X^{\mathrm{I}}_0}^{2p}] \lesssim s^{2p} \mathbf{m}_{2p}[{\mu}]+s^{2p}\mathbf{m}_{2p}[{\nu^{\star}}]+d^{p}s^p(1-s)^p\;,\]and \[\mathbb{E}[\norm{X^{\mathrm{I}}_1-X^{\mathrm{I}}_s}^{2p}]\mathrm{d}s \lesssim (1-s)^{2p}\mathbf{m}_{2p}[{\mu}]+(1-s)^{2p}\mathbf{m}_{2p}[{\nu^{\star}}]+d^{p}s^p(1-s)^p\;.\]

Proposition 1. The time reversal of the stochastic interpolant \(((X^\mathrm{I}_t)_{t\in [0,1]})^\mathrm{R}\) solves weakly \[\label{def:X95backward} \mathrm{d}\overleftarrow{X}_t= 2\overleftarrow{b}_t(\overleftarrow{X}_0, \overleftarrow{X}_t)\mathrm{d}t + \sqrt{2}\mathrm{d}\overleftarrow{B}_t\;, \quad t\in [0,1]\;, \quad \overleftarrow{X}_0\sim \nu^{\star}\;.\qquad{(1)}\] with \((\overleftarrow{B}_t)_{t\in [0,1]}\) \(d\)-dimensional Brownian motion independent of \(\overleftarrow{X}_0\) and \((t,x)\mapsto \overleftarrow{b}_t(x)\) as in [22], Lemma 2. In particular, they hold \[\label{eq:law95eq95XI95Xba} (X^\mathrm{I}_t)_{t\in [0,1]} \stackrel{\text{dist}}{=}((\overleftarrow{X}_t)_{t\in [0,1]})^{\mathrm{R}}\;,\qquad{(2)}\] and, for any \(u\in [0,1]\) and \(t\in [0, 1-u]\), \[\label{eq:XI95toXba} X^{\mathrm{I}}_{1-(t+u)}-X^{\mathrm{I}}_{1-u}\stackrel{\text{dist}}{=}\overleftarrow{f}_u^t+\overleftarrow{g}_u^t\;,\quad \overleftarrow{f}_u^t=2\int_{u}^{t+u} \overleftarrow{b}_r (\overleftarrow{X}_0, \overleftarrow{X}_r)\mathrm{d}r\;, \quad \overleftarrow{g}_u^t=\overleftarrow{B}_{t+u}-\overleftarrow{B}_{u}\;.\qquad{(3)}\] Moreover, \((\overleftarrow{X}_t)_{t\in [0,1]}\) is \((\overleftarrow{\mathcal{F}}_t)_{t\in [0,1]}\)-adapted, with \(\overleftarrow{\mathcal{F}}_t=\sigma(\overleftarrow{X}_0, (\overleftarrow{B}_u)_{u\le t})\).

Lemma 4. Assume [ass_moment,ass_score]. For any \(u\in [0,1]\), \(t\in [0, 1-u]\), \(p\in \{2,4,8\}\), we have that \[\mathbb{E}\Big[\norm{\overleftarrow{f}_u^t}^{p}\Big]\lesssim t^p \left( \norm{\nabla\log \pi}_{\mathrm{L}^p(\pi)}^p +\mathbf{m}_{p}[{\mu}]+\mathbf{m}_{p}[{\nu^{\star}}]\right)\;.\]

Proposition 2. The stochastic interpolant \((X^\mathrm{I}_t)_{t\in [0,1]}\) solves weakly \[\label{def:X95forward} \mathrm{d}\overrightarrow{X}_t= 2\overrightarrow{b}_t(\overrightarrow{X}_0, \overrightarrow{X}_t)\mathrm{d}t + \sqrt{2}\mathrm{d}\overrightarrow{B}_t\;, \quad t\in [0,1]\;, \quad \overrightarrow{X}_0\sim \mu\;.\qquad{(4)}\] with \((\overrightarrow{B}_t)_{t\in [0,1]}\) \(d\)-dimensional Brownian motion independent of \(\overrightarrow{X}_0\) and \((t,x)\mapsto \overrightarrow{b}_t(x)\) as in [22], Lemma 2. In particular, they hold \[\begin{align} (X^\mathrm{I}_t)_{t\in [0,1]} \stackrel{\text{dist}}{=}(\overrightarrow{X}_t)_{t\in [0,1]}\;, \end{align}\]and for any \(u\in [0,1]\) and \(t\in [0, 1-u]\), \[X^{\mathrm{I}}_{t+u}-X^{\mathrm{I}}_{u}\stackrel{\text{dist}}{=}\overrightarrow{f}_u^t+\overrightarrow{g}_u^t\;, \quad \overrightarrow{f}_u^t=2\int_{u}^{t+u} \overrightarrow{b}_r (\overrightarrow{X}_0, \overrightarrow{X}_r)\mathrm{d}r\;, \quad \overrightarrow{g}_u^t=\overrightarrow{B}_{t+u}-\overrightarrow{B}_{u}\;.\]Moreover, \((\overrightarrow{X}_t)_{t\in [0,1]}\) is \((\overrightarrow{\mathcal{F}}_t)_{t\in [0,1]}\)-adapted, with \(\overrightarrow{\mathcal{F}}_t=\sigma(\overrightarrow{X}_0, (\overrightarrow{B}_u)_{u\le t})\).

Lemma 5. Assume [ass_moment,ass_score]. For any \(u\in [0,1]\), \(t\in [0, 1-u]\), \(p\in \{2,4,8\}\), we have that \[\mathbb{E}\Big[\norm{\overrightarrow{f}_u^t}^{p}\Big]\lesssim t^p \left(\norm{\nabla\log \pi}_{\mathrm{L}^p(\pi)}^p +\mathbf{m}_{p}[{\mu}]+\mathbf{m}_{p}[{\nu^{\star}}]\right)\;,\]

6.4 On the conditional Markov kernel↩︎

6.4.1 Strongly log-concave case↩︎

Strongly log-concave couplings satisfy 6. More specifically, let

Assumption 9 (Strong log-concavity of the coupling). \(\pi\in\mathcal{P}(\mathbb{R}^{2d})\) is absolutely continuous with respect to \(\mathrm{Leb}^{2d}\), \(\log\pi\) is \(\mathrm{C}^2\), and satisfies \[\begin{align} \nabla^2(-\log\pi)\succeq \alpha_{\pi}\operatorname{Id}\;, \end{align}\]

then \(\pi\) satisfies 6 with \(M_\pi=0\). For sake of clarity, we first consider the case where \(\pi\in\Pi(\mu,\nu^{\star})\) satisfies 9.

Lemma 6. Assume that \(\pi\in\Pi(\mu,\nu^{\star})\) satisfies 9. Fix \(t\in (0,1)\) and \(y_t\in \mathbb{R}^d\). Then, \(\mathrm{K}_t(y_t, \cdot)\) satisfies 9 with the same parameter \(\alpha_{\pi}\). Moreover, \(\mathrm{K}^{\#0}_t(y_t, \cdot)\) and \(\mathrm{K}^{\#1}_t(y_t, \cdot)\) satisfy 9 as well with better parameters. Specifically, \[\begin{align} \label{eq:strong95regularity95conditioned95kernel} \nabla^2(-\log\mathrm{k}^{\#0}_t(y_t|\cdot))\succeq \left(\alpha_{\pi} +\frac{1-t}{2t} \right) \operatorname{Id}\;, \quad \nabla^2(-\log\mathrm{k}^{\#1}_t(y_t|\cdot))\succeq \left(\alpha_{\pi} +\frac{t}{2(1-t)} \right) \operatorname{Id}\;. \end{align}\tag{26}\]

Proof of 6:. Fix \(t\in [0,1)\) and \(y_t\in \mathbb{R}^d\) and denote by \[\begin{align} V(y_0,y_1):= -\log \pi(y_0,y_1) \;. \end{align}\]Then, because of 18 , it holds \[\begin{align} \mathrm{k}_{t}(y_0,y_1|y_t)\propto \exp\left(-\frac{\|y_t-(1-t)y_0-t y_1\|^2}{4t(1-t)}-V(y_0,y_1)\right)\;, \end{align}\]hence \[\begin{align} \label{eq:log95conditional95density} -\log \mathrm{k}_{t}(y_0,y_1|y_t)= \frac{\|y_t-(1-t)y_0-t y_1\|^2}{4t(1-t)}+V(y_0,y_1)\;, \end{align}\tag{27}\] and \[\begin{align} \label{eq:hessians95conditional95marginals} \nabla^2_{y_0} (-\log \mathrm{k}_{t}(y_0,y_1|y_t))= \frac{1-t}{2t}+\nabla^2_{y_0}V(y_0,y_1)\;,\;\nabla^2_{y_1} (-\log \mathrm{k}_{t}(y_0,y_1|y_t))= \frac{t}{2(1-t)}+\nabla^2_{y_1}V(y_0,y_1)\;. \end{align}\tag{28}\] At this pont, 26 follows directly from 9. ◻

6.4.2 Weakly log-concave case↩︎

We now consider the general case where \(\pi\in\Pi(\mu,\nu^{\star})\) satisfies 6.

Lemma 7. Assume that \(\pi\in \Pi(\mu,\nu^{\star})\) satisfies 6. Fix \(t\in (0,1)\) and \(y_t\in \mathbb{R}^d\). Then, \(\mathrm{K}_t(y_t, \cdot)\) satisfies 6 with the same parameters \(\alpha_{\pi}, M_{\pi}\). Moreover, \(\mathrm{K}^{\#0}_t(y_t, \cdot)\) and \(\mathrm{K}^{\#1}_t(y_t, \cdot)\) satisfy 6 with better parameters. Specifically, \[\begin{align} \label{def:param95weak95conv95improved} k_{-\log\mathrm{k}^{\#0}_t}(r)\ge \left(\alpha_\pi+\frac{1-t}{2t}\right)-r^{-1}f_{M_\pi}(r)\;, \quad k_{-\log\mathrm{k}^{\#1}_t}(r)\ge \left(\alpha_\pi+\frac{t}{2(1-t)}\right)-r^{-1}f_{M_\pi}(r)\;. \end{align}\tag{29}\]

Proof of 7:. Fix \(t\in (0,1)\) and \(y_t\in \mathbb{R}^d\). Proceed as in the proof of 6 to get 27 . Using 6 and 27 , we get that \(\mathrm{K}_t(y_t, \cdot)\) satisfies 6 with the same parameters \(\alpha_{\pi}, M_{\pi}\) and that the marginals of \(\mathrm{K}_t(y_t, \cdot)\) satisfy 6 with parameters given by 29 . ◻

6.5 On the mimicking drift’s regularity↩︎

6.5.1 Strongly log concave case↩︎

We preface this section with an auxiliary technical lemma.

Lemma 8. Let \[\begin{align} Y = \begin{pmatrix} Y^{(1)} \\[2mm] Y^{(2)} \end{pmatrix}\;, \quad Y^{(1)}, Y^{(2)} \in \mathbb{R}^d\;, \end{align}\]be a random vector with finite second order moments. Then \[\begin{align} \left\|\mathrm{Cov}(Y^{(1)}, Y^{(2)})\right\|_{\mathrm{op}}\le \sqrt{ \left\|\mathrm{Cov}(Y^{(1)})\right\|_{\mathrm{op}} \left\|\mathrm{Cov}( Y^{(2)})\right\|_{\mathrm{op}}}\;. \end{align}\]

Proof of 8:. Let \(\Sigma:=\mathrm{Cov}(Y)\). Then \(\Sigma\) can be written in block form as \[\begin{align} \Sigma = \begin{pmatrix} \Sigma_{11} & \Sigma_{12} \\[1mm] \Sigma_{21} & \Sigma_{22} \end{pmatrix}\;, \end{align}\] where \[\begin{align} \Sigma_{11} = \mathrm{Cov}(Y^{(1)})\;, \quad \Sigma_{22} = \mathrm{Cov}(Y^{(2)})\;, \quad \Sigma_{12} = \mathrm{Cov}(Y^{(1)}, Y^{(2)})\;, \quad \Sigma_{21} = \Sigma_{12}^T\;. \end{align}\] Recall that, for any matrix \(\mathrm{A}\) or order \(d\), we have \[\begin{align} \label{def95op95norm} \|\mathrm{A}\|_{\mathrm{op}}=\sup_{u,v\in\mathbb{R}^d, \|u\|,\|v\|=1} |u^\top \mathrm{A} v|\;. \end{align}\tag{30}\] Let \(u,v \in \mathbb{R}^d\) be arbitrary unit vectors. Being \(\Sigma\) positive semi-definite, for any \(t\in \mathbb{R}\), we have that \[\begin{align} \label{eq95in95lemma95discriminant} \begin{pmatrix} u \\ tv \end{pmatrix}^T \Sigma \begin{pmatrix} u \\ tv \end{pmatrix} = u^T \Sigma_{11} u + \left(2 u^T \Sigma_{12} v\right)t +\left( v^T \Sigma_{22} v\right)t^2 \ge 0\;. \end{align}\tag{31}\] This implies that the discriminant of the above second order polynomial in \(t\) is always non-positive, i.e., \[\begin{align} \big(2\, u^T \Sigma_{12} v\big)^2 - 4 \,(u^T \Sigma_{11} u) (v^T \Sigma_{22} v)\le 0\;, \end{align}\] or equivalently, \[\begin{align} (u^T \Sigma_{12} v)^2 \le (u^T \Sigma_{11} u)( v^T \Sigma_{22} v )\;. \end{align}\] It follows from 30 that \[\begin{align} |u^T \Sigma_{12} v| \le \sqrt{\|\Sigma_{11}\|_{\mathrm{op}} \|\Sigma_{22}\|_{\mathrm{op}}}\;. \end{align}\] The thesis then follows from the arbitrary of the unit vectors \(u,v\in\mathbb{R}^d.\) ◻

Proposition 3. Assume that \(\pi\in \Pi(\mu,\nu^{\star})\) satisfies [ass_strong_regularity,ass_hessian]. Then, for any \(s\in (0,1)\) and \(x,y\in\mathbb{R}^d\), it holds \[\begin{align} \label{eq:strong95lipschitz95mimicking95drift} \|\tilde{\beta}_s(x)-\tilde{\beta}_s(y)\|\le \frac{4\sqrt{2}}{\sqrt{\alpha_{\pi}}} \|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}\frac{1}{\sqrt{s(1-s)}}\|x-y\|\;. \end{align}\qquad{(5)}\]

Proof of 3:. Few computations lead to \[\label{jacobian95mimicking95drift} \begin{align} & D_x \tilde{\beta}_s(x)\\ &= 2 \frac{\int_{\mathbb{R}^{2d}} \nabla_x p_{1-s}(x_1|x)(\nabla_x p_s(x|x_0))^{\operatorname{T}} \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\\ &+ 2 \frac{\int_{\mathbb{R}^{2d}} p_{s}(x|x_0) \nabla^2_x p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)} \\ &-2 \Big(\frac{\int_{\mathbb{R}^{2d}} p_{s}(x|x_0) \nabla_x p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\Big)\Big(\frac{ \int_{\mathbb{R}^{2d}} \nabla_x p_{s}(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\Big)^{\operatorname{T}} \\ &- 2\Big(\frac{\int_{\mathbb{R}^{2d}} p_{s}(x|x_0) \nabla_x p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\Big) \Big(\frac{\int_{\mathbb{R}^{2d}} p_{s}(x|x_0) \nabla_x p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\Big)^{\operatorname{T}}\;. \end{align}\tag{32}\]

Using 20 and integration by parts formula, we get \[\begin{align} & D_x \tilde{\beta}_s(x)\\ &= \frac{\int_{\mathbb{R}^{2d}} \frac{\nabla_{x_1}\tilde{\pi}}{\tilde{\pi}}(x_0,x_1)(\frac{x_0-x}{s})^{\operatorname{T}} p_{s}(x|x_0)p_{1-s}(x_1|x)\tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\\ &+ \frac{\int_{\mathbb{R}^{2d}} \frac{x_1-x}{1-s}(\frac{\nabla_{x_1}\tilde{\pi}}{\tilde{\pi}}(x_0,x_1))^{\operatorname{T}} p_{s}(x|x_0)p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\\ &- \Big(\frac{\int_{\mathbb{R}^{2d}} \frac{\nabla_{x_1}\tilde{\pi}}{\tilde{\pi}}(x_0,x_1)p_{s}(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\Big)\\ &\quad \cdot\Big(\frac{ \int_{\mathbb{R}^{2d}}\frac{x_0-x}{s} p_{s}(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\Big)^{\operatorname{T}} \\ &- \Big(\frac{\int_{\mathbb{R}^{2d}} \frac{x_1-x}{1-s}p_{s}(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\Big) \\ &\quad \cdot\Big(\frac{\int_{\mathbb{R}^{2d}} \frac{\nabla_{x_1}\tilde{\pi}}{\tilde{\pi}}(x_0,x_1)p_{s}(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\Big)^{\operatorname{T}}\\ &\label{eq:jacobian95as95covariance}=\text{Cov}_{(X^{\mathrm{I}}_0,X^\mathrm{I}_1)\sim \mathrm{K}_s(x,\cdot)}\left(\frac{X^\mathrm{I}_0-x}{s}, \frac{\nabla_{x_1}\tilde{\pi}}{\tilde{\pi}}(X^{\mathrm{I}}_0,X^\mathrm{I}_1)\right)+\text{Cov}_{(X^{\mathrm{I}}_0,X^\mathrm{I}_1)\sim \mathrm{K}_s(x,\cdot)}\left(\frac{X^\mathrm{I}_1-x}{1-s}, \frac{\nabla_{x_1}\tilde{\pi}}{\tilde{\pi}}(X^{\mathrm{I}}_0,X^\mathrm{I}_1)\right)\;. \end{align}\tag{33}\] Therefore, using 8, [ass_strong_regularity,ass_hessian], 6, Brascamp-Lieb inequality (see Lemma 2 in [41]) and 2, we get \[\begin{align} \| D_x \tilde{\beta}_s(x)\|_{\mathrm{op}}&\le \left(\sqrt{\left\|\text{Cov}_{X^\mathrm{I}_0\sim \mathrm{K}^{\# 0}_s(x,\cdot)}\left(\frac{X^\mathrm{I}_0-x}{s}\right)\right\|_{\mathrm{op}}}+\sqrt{\left\|\text{Cov}_{X^\mathrm{I}_1\sim \mathrm{K}^{\# 1}_s(x,\cdot)}\left(\frac{X^\mathrm{I}_1-x}{1-s}\right)\right\|_{\mathrm{op}}}\right)\\ &\quad \cdot\sqrt{\left\|\text{Cov}_{(X^{\mathrm{I}}_0,X^\mathrm{I}_1)\sim \mathrm{K}_s(x,\cdot)}\left(\nabla\log\tilde{\pi}(X^{\mathrm{I}}_0,X^\mathrm{I}_1)\right)\right\|_{\mathrm{op}}}\\ &\preceq \left(\frac{1}{s}\sqrt{\frac{2s}{2\alpha_{\pi}s+ 1-s} }+ \frac{1}{1-s}\sqrt{\frac{2(1-s)}{s+2\alpha_\pi(1-s)}}\right)\sqrt{\frac{\|\nabla \log \tilde{\pi}\|_{\mathrm{L}^2(\pi)}^2}{\alpha_{\pi}}}\\ &\le \frac{4\sqrt{2}}{\sqrt{\alpha_{\pi}}} \|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}\frac{1}{\sqrt{s(1-s)}}\;. \end{align}\] Hence, we have that \[\begin{align} \| D_x \tilde{\beta}_s(x)\|_{\mathrm{op},\infty}&\le \frac{4\sqrt{2}}{\sqrt{\alpha_{\pi}}} \|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}\frac{1}{\sqrt{s(1-s)}}\;. \end{align}\]The thesis follows directly. ◻

6.5.2 Weakly log concave case↩︎

Proposition 4. Assume that \(\pi\in \Pi(\mu,\nu^{\star})\) satisfies [ass_regularity,ass_hessian]. Then, for any \(s\in (0,1)\) and \(x,y\in\mathbb{R}^d\), it holds \[\begin{align} \label{eq:extremes95lipschitz95mimicking95drift} \|\tilde{\beta}_s(x)-\tilde{\beta}_s(y)\|\le \frac{4\sqrt{2}}{\sqrt{\alpha_{\pi}}}\exp\left(\frac{M_\pi}{\alpha_\pi}\right) \|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}\frac{1}{\sqrt{s(1-s)}}\;. \end{align}\qquad{(6)}\]

Proof of 4:. Proceed as in the proof of 3, to get 33 and therefore \[\begin{align} \label{eq:upper95bound95jacobian} \|D_x \tilde{\beta}_s(x)\|_{\mathrm{op}}&\le \left(\sqrt{\left\|\text{Cov}_{X^\mathrm{I}_0\sim \mathrm{K}^{\# 0}_s(x,\cdot)}\left(\frac{X^\mathrm{I}_0-x}{s}\right)\right\|_{\mathrm{op}}}+\sqrt{\left\|\text{Cov}_{X^\mathrm{I}_1\sim \mathrm{K}^{\# 1}_s(x,\cdot)}\left(\frac{X^\mathrm{I}_1-x}{1-s}\right)\right\|_{\mathrm{op}}}\right)\\ &\quad \cdot\sqrt{\left\|\text{Cov}_{(X^{\mathrm{I}}_0,X^\mathrm{I}_1)\sim \mathrm{K}_s(x,\cdot)}\left(\nabla\log\tilde{\pi}(X^{\mathrm{I}}_0,X^\mathrm{I}_1)\right)\right\|_{\mathrm{op}}}\;. \end{align}\tag{34}\] For any \(\alpha>0, M\ge 0\), denote by \[\begin{align} \label{eq:poincare95constant} \xi(\alpha,M):=\frac{\alpha}{\exp\left(M/\alpha\right)}\;. \end{align}\tag{35}\] Also, fix \(s\in (0,1)\) and \(y_s\in \mathbb{R}^d\). The byproduct of 7 and Theorem 5.7 in [42] implies that \(\mathrm{K}_s(y_s, \cdot)\), \(\mathrm{K}^{\#0}_s(y_s, \cdot)\) and \(\mathrm{K}^{\#1}_s(y_s, \cdot)\) satisfy log-Sobolev inequality with parameters \(\xi^{-1}(\alpha_{\pi},M_{\pi})\), \(\xi^{-1}(\alpha_{\pi}+(1-s)/(2s),M_{\pi})\) and \(\xi^{-1}(\alpha_{\pi}+s/(2(1-s)),M_{\pi})\) respectively, with \(\xi(\alpha,M)\) for \(\alpha>0, M\ge 0\) defined in 35 . Consequently, as shown in [43], \(\mathrm{K}_s(y_t, \cdot)\), \(\mathrm{K}^{\#0}_s(y_s, \cdot)\) and \(\mathrm{K}^{\#1}_s(y_s, \cdot)\) satisfy Poincaré inequality with constants \(2\xi(\alpha_{\pi},M_{\pi})\), \(2\xi^{-1}(\alpha_{\pi}+(1-s)/(2s),M_{\pi})\) and \(2\xi^{-1}(\alpha_{\pi}+s/(2(1-s)),M_{\pi})\) respectively. We therefore get that

\[\begin{align} &\|D_x \tilde{\beta}_s(x) \|_{\mathrm{op}}\\ &\le \left( \frac{1}{s}\sqrt{2\xi^{-1}\left(\alpha_{\pi}+\frac{1-s}{2s},M_{\pi}\right)}+ \frac{1}{1-s}\sqrt{2\xi^{-1}\left(\alpha_{\pi}+\frac{s}{2(1-s)},M_{\pi}\right)}\right) \|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}\sqrt{2\xi^{-1}(\alpha_{\pi},M_{\pi})}\\ &=\left(\frac{2\sqrt{2}}{\sqrt{s}}\frac{1}{\sqrt{2s\alpha_\pi+1-s}}\exp\left(\frac{M_\pi}{\alpha_\pi+(1-s)/(2s)}\right)+\frac{2\sqrt{2}}{\sqrt{1-s}}\frac{1}{\sqrt{2(1-s)\alpha_\pi+s}}\exp\left(\frac{M_\pi}{\alpha_\pi+s/(2(1-s))}\right) \right)\|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}\\ &\le \frac{4\sqrt{2}}{\sqrt{\alpha_{\pi}}}\exp\left(\frac{M_\pi}{\alpha_\pi}\right) \|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}\frac{1}{\sqrt{s(1-s)}}\;. \end{align}\] ◻

7 Convergence Bounds in Kullback–Leibler Divergence↩︎

7.1 Non-early-stopping regime with constant step-size↩︎

Proof of 1:. For any \(t\in [0,1]\), we denote by \(\nu^{\star}_t=\text{Law}(X^{\mathrm{M}}_t)\) and \(\nu^{\theta^{\star}}_t=\text{Law}(X^{\theta^{\star}}_t).\) Also, we denote by \(\mathcal{L}^\mathrm{M}\) the generator of \((X^{\mathrm{M}}_t)_{t\in [0,1]}\), which is defined for any \(t\in [0,1]\) and \(\rho \in C^2(\mathbb{R}^d)\) as \[\mathcal{L}^\mathrm{M}_t \rho :=\langle\nabla_x \rho ,\tilde{\beta}_t\rangle + \Delta_x \rho \;.\] We fix \(0<\epsilon_1<\min\{h, 1/2\}\). First, using the data processing inequality (see, \(e.g.\), Lemma 1.6 in [44]), the standard decomposition of the KL divergence [26], [27], [29] based on Girsanov theorem, triangle inequality and 1, we bound the KL divergence between \(\nu^{\star}_{1-\epsilon_1}\) and \(\nu^{\theta^{\star}}_{1-\epsilon_1}\) as follows

\[\begin{align} \label{Decomposition95KL95Girsanov} &\mathrm{KL}(\nu^{\star}_{1-\epsilon_1}|\nu^{\theta^{\star}}_{1-\epsilon_1})\\ & \le \mathrm{KL}(\text{Law}((X^{\mathrm{M}}_t)_{t\in [0, 1-\epsilon_1]}) | \mathrm{Law}((X^{\theta^{\star}}_t)_{t\in [0,1-\epsilon_1]}))\\ &\lesssim \sum_{k=0}^{N-2}\int_{t_k}^{t_{k+1}} \mathbb{E}\Big[\norm{ s_{\theta^\star}(t_k, X^{\mathrm{M}}_{t_k})- \tilde{\beta}_{t}(X^{\mathrm{M}}_{t})}^2\Big]\mathrm{d}t + \int_{1-h}^{1-\epsilon_1} \mathbb{E}\Big[\norm{ s_{\theta^\star}(1-h, X^{\mathrm{M}}_{1-h})- \tilde{\beta}_{t}(X^{\mathrm{M}}_{t})}^2\Big]\mathrm{d}t\\ &\lesssim \varepsilon^2 +\sum_{k=0}^{N-2}\int_{t_k}^{t_{k+1}} \mathbb{E}\Big[\norm{ \tilde{\beta}_{t_k}(X^{\mathrm{M}}_{t_k})- \tilde{\beta}_{t}(X^{\mathrm{M}}_{t})}^2\Big]\mathrm{d}t + \int_{1-h}^{1-\epsilon_1} \mathbb{E}\Big[\norm{ \tilde{\beta}_{1-h}(X^{\mathrm{M}}_{1-h})- \tilde{\beta}_{t}(X^{\mathrm{M}}_{t})}^2\Big]\mathrm{d}t\;. \end{align}\tag{36}\] Second, we aim at bounding the RHS of 36 uniformly in \(\epsilon_1\). Indeed, if we assume to be able to bound it with a constant \(A\) independent of \(\epsilon_1\), then, using the weak convergence of \(X^{\mathrm{M}}_{1-\epsilon_1}\) to \(X^\mathrm{I}_1\) (whose law is given by \(\nu^{\star}\)) as \(\epsilon_1 \to 0\), the continuity of \((X^{\theta^{\star}}_t)_{t\in [0,1]}\), hence the weak convergence of \(X^{\theta^{\star}}_{1-\epsilon_1}\) to \(X^{\theta^{\star}}_1\) (whose law is given by \(\nu^{\theta^\star}_1\)) as \(\epsilon_1 \to 0\), and the lower semi-continuity of the \(\mathrm{KL}\)-divergence with respect to the weak convergence (see, \(e.g.\), Theorem 19 in [45]), we will get \[\label{lsc95of95kl} \begin{align} \mathrm{KL}(\nu^{\star}|\nu^{\theta^{\star}}_{1}) \le \liminf_{\epsilon_1\to 0} \mathrm{KL}(\nu^{\star}_{1-\epsilon_1}|\nu^{\theta^{\star}}_{1-\epsilon_1})\lesssim \liminf_{\epsilon_1\to 0} A=A \;. \end{align}\tag{37}\] Let us therefore bound the RHS of 36 . We will do so by using stochastic calculus tools, and, more precisely, Ito’s formula. This formula yields \[\label{ito95dec95drift} \mathrm{d}\tilde{\beta}_t(X^{\mathrm{M}}_t)= (\partial_t +\mathcal{L}^\mathrm{M}_t)\tilde{\beta}_t(X^{\mathrm{M}}_t) \mathrm{d}t + \sqrt{2} D_x \tilde{\beta}_t(X^{\mathrm{M}}_t) \mathrm{d}W_t\;,\quad t\in [0,1-\epsilon_1]\;.\tag{38}\] So, applying Young inequality and Ito’s isometry, we have that, for any \(k=0,\cdots,N-1\) \[\label{bound95via95ito95of95difference95of95b95t} \begin{align} &\mathbb{E}\Big[\norm{ \tilde{\beta}_{t_k}(X^{\mathrm{M}}_{t_k})- \tilde{\beta}_{t}(X^{\mathrm{M}}_{t})}^2\Big]\\ &=\mathbb{E}\Bigg[\norm{\int_{t_k}^t(\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s(X^{\mathrm{M}}_s) \mathrm{d}s +\sqrt{2} \int_{t_k}^t D_x \tilde{\beta}_s(X^{\mathrm{M}}_s) \mathrm{d}W_s\;}^2\Bigg]\\ &\lesssim \mathbb{E}\Bigg[\norm{\int_{t_k}^t (\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s(X^{\mathrm{M}}_s) \mathrm{d}s }^2\Bigg] + 2 \int_{t_k}^{t_{k+1}} \mathbb{E}\Big[\norm{D_x \tilde{\beta}_s(X^{\mathrm{M}}_s)}^2\Big] \mathrm{d}s\;. \end{align}\tag{39}\] We now bound separately the two upper addends. To do so, we introduce the auxiliary measures \(\lambda_k^h(\mathrm{d}s)\in \mathcal{P}([t_k, t_{k+1}])\) for \(k=0,..., N-2\) and \(\lambda_{N-1}^h(\mathrm{d}s)\in \mathcal{P}([1-h, 1-\epsilon_1])\) which will help us, via a double change of measure argument, to mitigate the bad behaviour at \(t=0\) and \(t=1\) of the reciprocal characteristic of the mimicking drift (i.e., \(\partial_s +\mathcal{L}^\mathrm{M}_s\)), which is the trickiest addend. Namely, for \(k=0,..., N-2\) we consider the measures \(\lambda_k^h(\mathrm{d}s)\in \mathcal{P}([t_k, t_{k+1}])\) defined as \[\require{upgreek} \label{Auxiliary95measure95on95time} \lambda_k^h(\mathrm{d}s)= \frac{\uprho(s)^{-1}}{Z_k}\mathbb{1}_{[t_k,t_{k+1}]} \mathrm{d}s\;,\tag{40}\] with \[\require{upgreek} \label{def:uprho} \uprho(s)^{-1}= s^{-7/8}\mathbb{1}_{\{s\le 1/2\}}+(1-s)^{-7/8}\mathbb{1}_{\{s> 1/2\}}\;,\tag{41}\] and \[Z_k = \int_{\min\{t_k, 1/2\}}^{\min\{t_{k+1}, 1/2\}} r^{-7/8} \mathrm{d}r+ \int_{\max\{t_k,1/2\}}^{\max\{t_{k+1}, 1/2\}} (1-r)^{-7/8} \mathrm{d}r\;.\] For \(k=N-1\), we consider the measure \(\lambda_{N-1}^h(\mathrm{d}s)\in \mathcal{P}([1-h, 1-\epsilon_1])\) defined as \[\require{upgreek} \label{jadgiyrl} \lambda_{N-1}^h(\mathrm{d}s)= \frac{\uprho(s)^{-1}}{Z_{N-1}}\mathbb{1}_{[1-h,1-\epsilon_1]} \mathrm{d}s\;,\tag{42}\] with \(\require{upgreek} \uprho(s)^{-1}\) as in 41 and \[Z_{N-1} = \int_{\min\{1-h, 1/2\}}^{\min\{1-\epsilon_1, 1/2\}} r^{-7/8} \mathrm{d}r+ \int_{\max\{1-h,1/2\}}^{\max\{1-\epsilon_1, 1/2\}} (1-r)^{-7/8} \mathrm{d}r\;.\] Note that, for any \(s\in [0, 1-\epsilon_1]\) and for any \(k=0,..., N-1\), they hold \[\require{upgreek} \label{properties95of95auxiliary95measure95on95time} \uprho(s)\lesssim 1\;, \quad Z_k\lesssim h^{1/8} \;, \quad \int_{0}^{1-\epsilon_1} \uprho^{-1}(s) \mathrm{d}s\lesssim 1\;.\tag{43}\] We start by bounding the first addend, that is the one that involves the reciprocal characteristic of the mimicking drift. With a first change of measure argument, we get for any \(k=0,..., N-1\), \[\require{upgreek} \mathbb{E}\Bigg[\norm{\int_{t_k}^t (\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s(X^{\mathrm{M}}_s) \mathrm{d}s }^2\Bigg] = Z_k^2 \mathbb{E}\Bigg[\norm{\int_{t_k}^t (\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s(X^{\mathrm{M}}_s) \uprho(s)\lambda_k^h(\mathrm{d}s) }^2\Bigg]\;,\]where, in the last inequality, we used 43 . But then, if we apply Jensen inequality and use an other change of measure argument, we get

\[\require{upgreek} \begin{align} \label{bd:first95addend} \mathbb{E}\Bigg[\norm{\int_{t_k}^t (\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s(X^{\mathrm{M}}_s) \mathrm{d}s }^2\Bigg]&\le Z_k^2 \mathbb{E}\Bigg[\int_{t_k}^t \norm{(\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s(X^{\mathrm{M}}_s)}^2 \uprho(s)^{2} \lambda_k^h(\mathrm{d}s)\Bigg]\\ &\le Z_k \int_{t_k}^{t}\mathbb{E}\Big[\norm{(\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s(X^{\mathrm{M}}_s)}^2 \Big]\uprho(s) \mathrm{d}s\\ &\lesssim h^{1/8}\int_{t_k}^{t_{k+1}} \mathbb{E}\Big[\norm{(\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s(X^{\mathrm{M}}_s)}^2 \Big]\uprho(s) \mathrm{d}s\;. \end{align}\tag{44}\] Let us now focus on the second addend. Remarkably, this addend can be bounded via the reciprocal characteristic of \(\tilde{\beta}\): because of Ito’s formula, for \(t\in [0,1-\epsilon_1]\), it holds true \[\mathrm{d}\norm{\tilde{\beta}_t(X^{\mathrm{M}}_t)}^2 = \Big\{2 \langle \tilde{\beta}_t, (\partial_t +\mathcal{L}^\mathrm{M}_t)\tilde{\beta}_t\rangle + 2\norm{D_x\tilde{\beta}_t}^2 \Big\}(X^{\mathrm{M}}_t)\mathrm{d}t + 2\sqrt{2} \langle \tilde{\beta}_t, D_x\tilde{\beta}_t\rangle(X^{\mathrm{M}}_t)\mathrm{d}W_t\;.\]Note that, under 2 for \(\mu\) and \(\nu^{\star}\) and 3 for \(\tilde{\pi}\), Lemma 4 in [22] ensures that the process \((\int_0^s \langle \tilde{\beta}_t, D_x\tilde{\beta}_t\rangle(X^{\mathrm{M}}_t)\mathrm{d}W_t)_{s\in [0,1-\epsilon_1]}\) is a true martingale. Consequently, we have that \[\label{RHS952} \begin{align} &2 \int_{t_k}^{t_{k+1}} \mathbb{E}\Big[\norm{D_x\tilde{\beta}_s(X^{\mathrm{M}}_s)}^2\Big]\mathrm{d}s \\ &\le \mathbb{E}\Big[\norm{\tilde{\beta}_{t_{k+1}}(X^{\mathrm{M}}_{t_{k+1}})}^2\Big]- \mathbb{E}\Big[\norm{\tilde{\beta}_{t_{k}}(X^{\mathrm{M}}_{t_{k}})}^2\Big] +2\Big| \int_{t_k}^{t_{k+1}} \mathbb{E}[ \langle \tilde{\beta}_s (X^{\mathrm{M}}_s),(\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s (X^{\mathrm{M}}_s)\rangle] \mathrm{d}s \Big|\;. \end{align}\tag{45}\] But then, using, as before, a double change of measure argument and applying Cauchy-Schwartz inequality, we can bound the above expression as follows \[\require{upgreek} \begin{align} &2 \int_{t_k}^{t_{k+1}} \mathbb{E}\Big[\norm{D_x\tilde{\beta}_s(X^{\mathrm{M}}_s)}^2\Big]\mathrm{d}s \\ &\le \mathbb{E}\Big[\norm{\tilde{\beta}_{t_{k+1}}(X^{\mathrm{M}}_{t_{k+1}})}^2\Big]- \mathbb{E}\Big[\norm{\tilde{\beta}_{t_{k}}(X^{\mathrm{M}}_{t_{k}})}^2\Big] +2 Z_k\Big| \int_{t_k}^{t_{k+1}} \mathbb{E}[ \langle \tilde{\beta}_s (X^{\mathrm{M}}_s),(\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s (X^{\mathrm{M}}_s)\rangle] \uprho(s) \lambda_k^h(\mathrm{d}s) \Big|\\ &\le\mathbb{E}\Big[\norm{\tilde{\beta}_{t_{k+1}}(X^{\mathrm{M}}_{t_{k+1}})}^2\Big]- \mathbb{E}\Big[\norm{\tilde{\beta}_{t_{k}}(X^{\mathrm{M}}_{t_{k}})}^2\Big] + Z_k\int_{t_k}^{t_{k+1}} \mathbb{E}\Big[\norm{\tilde{\beta}_s(X^{\mathrm{M}}_s)}^2\Big] \lambda_k^h(\mathrm{d}s)\\ &\quad + Z_k\int_{t_k}^{t_{k+1}} \mathbb{E}\Big[\norm{(\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s (X^{\mathrm{M}}_s)}^2 \Big] \uprho(s)^{2} \lambda_k^h(\mathrm{d}s)\\ &= \mathbb{E}\Big[\norm{\tilde{\beta}_{t_{k+1}}(X^{\mathrm{M}}_{t_{k+1}})}^2\Big]- \mathbb{E}\Big[\norm{\tilde{\beta}_{t_{k}}(X^{\mathrm{M}}_{t_{k}})}^2\Big] + \int_{t_k}^{t_{k+1}} \mathbb{E}\Big[\norm{\tilde{\beta}_s(X^{\mathrm{M}}_s)}^2\Big] \uprho(s)^{-1} \mathrm{d}s\\ &\quad + \int_{t_k}^{t_{k+1}} \mathbb{E}\Big[\norm{(\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s (X^{\mathrm{M}}_s)}^2 \Big] \uprho(s) \mathrm{d}s\;. \end{align}\]Plugging this bound and 44 in 36 , we get \[\require{upgreek} \begin{align} \label{Decomposition95KL95via95generator} &\mathrm{KL}(\nu^{\star}_{1-\epsilon_1}|\nu^{\theta^{\star}}_{1-\epsilon_1})\\ &\lesssim \varepsilon^2 + h\mathbb{E}\Big[\norm{\tilde{\beta}_{1-\epsilon_1}(X^{\mathrm{M}}_{1-\epsilon_1})}^2\Big]+h \int_{0}^{{1-\epsilon_1}}\mathbb{E}\Big[ \norm{\tilde{\beta}_s (X^{\mathrm{M}}_s)}^2\Big]\uprho(s)^{-1}\mathrm{d}s\\ &\quad +h (h^{1/8}+1)\int_{0}^{1-\epsilon_1}\mathbb{E}\Big[ \norm{(\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s (X^{\mathrm{M}}_s)}^2\Big] \uprho(s)\mathrm{d}s\;. \end{align}\tag{46}\] We now compute explicitly and upper bound each term appearing in the RHS of 46 , recalling that, because of Theorem 1 in [22], for any \(s\in [0, 1]\), \(\nu^{\star}_s=\text{Law}(X^\mathrm{I}_s)\). We start with the second term, that is \[\begin{align} h\mathbb{E}\Big[ \norm{\tilde{\beta}_{1-\epsilon_1} (X^{\mathrm{M}}_{1-\epsilon_1})}^2\Big] \;. \end{align}\] Using 20 , integration by part, Jensen inequality, 2 and \(\mathbf{m}_{8}[{\mu}],\mathbf{m}_{8}[{\nu^{\star}}]\lesssim d^4\), we get \[\label{term95295dec95KL95gen} \begin{align} &\mathbb{E}\Big[ \norm{\tilde{\beta}_{1-\epsilon_1} (X^{\mathrm{M}}_{1-\epsilon_1})}^2\Big]\\ &\int_{\mathbb{R}^d}\norm{\frac{\int_{\mathbb{R}^{2d}} p_{1-\epsilon_1} (x|x_0)\nabla_x p_{\epsilon_1}(x_1|x)\tilde{\pi}(x_0,x_1) \mathrm{d}x_1 \mathrm{d}x_0}{p^{\mathrm{I}}_{1-\epsilon_1}(x)}}^2 p^{\mathrm{I}}_{1-\epsilon_1}(x)\mathrm{d}x\\ &\lesssim\int_{\mathbb{R}^d}\norm{\frac{\int_{\mathbb{R}^{2d}} p_{1-\epsilon_1}(x|x_0) p_{\epsilon_1}(x_1|x) (\nabla_{x_1}\tilde{\pi}(x_0,x_1)/\tilde{\pi}(x_0,x_1))\tilde{\pi}(x_0,x_1)\mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_{1-\epsilon_1}(x)}}^2 p^{\mathrm{I}}_{1-\epsilon_1}(x)\mathrm{d}x \\ &= \mathbb{E}\Bigg[\norm{\mathbb{E}\Bigg[\frac{\nabla_{x_1}\tilde{\pi}}{\tilde{\pi}}(X_0^{\mathrm{I}},X_1^{\mathrm{I}})\Bigg| X^{\mathrm{I}}_{1-\epsilon_1}\Bigg]}^2\Bigg] \le \mathbb{E}\Bigg[\mathbb{E}\Bigg[\norm{\frac{\nabla_{x_1}\tilde{\pi}}{\tilde{\pi}}(X_0^{\mathrm{I}},X_1^{\mathrm{I}})}^2\Bigg|X^{\mathrm{I}}_{1-\epsilon_1}\Bigg]\Bigg] \\ &= \mathbb{E}\Bigg[\norm{\frac{\nabla_{x_1}\tilde{\pi}}{\tilde{\pi}}(X_0^{\mathrm{I}},X_1^{\mathrm{I}})}^2\Bigg] =\norm{\frac{\nabla_{x_1}\tilde{\pi}}{\tilde{\pi}}}^2_{\mathrm{L}^2(\pi)}\le \norm{\nabla \log \tilde{\pi}}^2_{\mathrm{L}^2(\pi)}\lesssim \norm{\nabla \log \pi}^2_{\mathrm{L}^2(\pi)}+\mathbf{m}_{2}[{\mu}]+\mathbf{m}_{2}[{\nu^{\star}}]\\ &\lesssim \norm{\nabla \log \pi}^2_{\mathrm{L}^2(\pi)}+d\;. \end{align}\tag{47}\] Hence, we can bound the second term in 46 as follows \[\begin{align} \label{eq:second95term} h\mathbb{E}\Big[ \norm{\tilde{\beta}_{1-\epsilon_1} (X^{\mathrm{M}}_{1-\epsilon_1})}^2\Big] \lesssim h \left(\norm{\nabla \log \pi}^2_{\mathrm{L}^2(\pi)}+d\right)\;. \end{align}\tag{48}\] We now deal with the third term of the RHS of 46 , i.e., \[\require{upgreek} \begin{align} h \int_{0}^{{1-\epsilon_1}}\mathbb{E}\Big[ \norm{\tilde{\beta}_s (X^{\mathrm{M}}_s)}^2\Big]\uprho(s)^{-1}\mathrm{d}s\;. \end{align}\] Note that, proceeding as before and using 43 , we have that \[\require{upgreek} \label{term95395dec95KL95gen} \begin{align} &\int_{0}^{1-\epsilon_1}\mathbb{E}\Big[ \norm{\tilde{\beta}_s (X^{\mathrm{M}}_s)}^2\Big]\uprho(s)^{-1}\mathrm{d}s\\ &\lesssim \int_{0}^{1-\epsilon_1} \uprho(s)^{-1}\int_{\mathbb{R}^d}\norm{\frac{\int_{\mathbb{R}^{2d}} p_s(x|x_0) \nabla_x p_{1-s} (x_1|x)\tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}}^2 p^{\mathrm{I}}_s(x)\mathrm{d}x \mathrm{d}s\\ &= \int_{0}^{1-\epsilon_1} \uprho(s)^{-1}\int_{\mathbb{R}^d}\norm{\frac{1}{p^\mathrm{I}_s(x)}\int_{\mathbb{R}^{2d}} p_s(x|x_0) p_{1-s} (x_1| x)\frac{\nabla_{x_1}\tilde{\pi}}{\tilde{\pi}}(x_0, x_1)\tilde{\pi}(x_0, x_1) \mathrm{d}x_0, \mathrm{d}x_1}^2 p^{\mathrm{I}}_s(x)\mathrm{d}x\\ &= \int_{0}^{1-\epsilon_1} \uprho(s)^{-1}\mathbb{E}\Bigg[\norm{\mathbb{E}\Bigg[\frac{\nabla_{x_1}\tilde{\pi}}{\tilde{\pi}}(X_0^{\mathrm{I}},X_1^{\mathrm{I}})\Bigg| X^{\mathrm{I}}_s \Bigg]}^2\Bigg] \mathrm{d}s\\ &\le \int_{0}^{1-\epsilon_1} \uprho(s)^{-1}\mathbb{E}\Bigg[\mathbb{E}\Bigg[\norm{\frac{\nabla_{x_1}\tilde{\pi}}{\tilde{\pi}}(X_0^{\mathrm{I}},X_1^{\mathrm{I}})}^2\Bigg| X^{\mathrm{I}}_s \Bigg]\Bigg] \mathrm{d}s \\ &=\int_{0}^{1-\epsilon_1} \uprho(s)^{-1}\mathbb{E}\Bigg[\norm{\frac{\nabla_{x_1}\tilde{\pi}}{\tilde{\pi}}(X_0^{\mathrm{I}},X_1^{\mathrm{I}})}^2\Bigg] \mathrm{d}s\\ &= \int_{0}^{1-\epsilon_1} \uprho(s)^{-1} \mathrm{d}s\:\norm{\nabla \log \tilde{\pi}}^2_{\mathrm{L}^2(\pi)}\lesssim \norm{\nabla \log \tilde{\pi}}^2_{\mathrm{L}^2(\pi)}\lesssim \norm{\nabla \log \pi}^2_{\mathrm{L}^2(\pi)}+\mathbf{m}_{2}[{\mu}]+\mathbf{m}_{2}[{\nu^{\star}}] \\ &\lesssim \norm{\nabla \log \pi}^2_{\mathrm{L}^2(\pi)}+d\;. \end{align}\tag{49}\] Therefore, we can bound the third term in 46 as follows \[\require{upgreek} \begin{align} \label{eq:third95term} h \int_{0}^{{1-\epsilon_1}}\mathbb{E}\Big[ \norm{\tilde{\beta}_s (X^{\mathrm{M}}_s)}^2\Big]\uprho(s)^{-1}\mathrm{d}s\lesssim h\left(\norm{\nabla \log \pi}^2_{\mathrm{L}^2(\pi)}+d\right) \;. \end{align}\tag{50}\] We now turn to the last term appearing in the r.h.s. of 46 , i.e., \[\require{upgreek} \begin{align} &h (h^{1/8}+1) \int_{0}^{1-\epsilon_1}\mathbb{E}\Big[ \norm{(\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s (X^{\mathrm{M}}_s)}^2\Big] \uprho(s)\mathrm{d}s\\ &= h (h^{1/8}+1) \int_{0}^{1-\epsilon_1}\int_{\mathbb{R}^{d}} \norm{(\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s (x)}^2 \uprho(s) p^{\mathrm{I}}_s(x) \mathrm{d}x \;\mathrm{d}s\;. \end{align}\] Some computations and 23 (we refer to Appendix A4 in [22] for more details) lead to \[\label{ito95drift951} \begin{align} &(\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s(x)=\sum_{k=1}^{6} A_s^k(x)\;, \end{align}\tag{51}\] where we have defined \[\begin{gather} A_s^1(x)=4\frac{\int_{\mathbb{R}^{2d}} \Delta_x p_s(x|x_0) \nabla_x p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)} \\ - 4\frac{\int_{\mathbb{R}^{2d}} \Delta_x p_s(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1 \int_{\mathbb{R}^{2d}} p_s(x|x_0) \nabla_x p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{(p^{\mathrm{I}}_s(x))^2}\;, \end{gather}\] \[\begin{gather} A_s^2(x)= 4\frac{\int_{\mathbb{R}^{2d}} \nabla^2_x p_{1-s}(x_1|x) \nabla_x p_s(x|x_0) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\\ - 4\frac{\int_{\mathbb{R}^{2d}} p_s(x|x_0) \nabla^2_x p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1 \int_{\mathbb{R}^{2d}} \nabla_x p_s(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{(p^{\mathrm{I}}_s(x))^2}\;. \end{gather}\] \[\begin{gather} A^3_s(x)=-4\frac{\int_{\mathbb{R}^{2d}} \nabla_x p_{1-s}(x_1|x)(\nabla_x p_s(x|x_0))^{\operatorname{T}} \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\\ \cdot\frac{ \int_{\mathbb{R}^{2d}} \nabla_x p_s(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\;, \end{gather}\] \[\begin{gather} A^4_s(x)= -4\frac{\int_{\mathbb{R}^{2d}} \langle \nabla_x p_{1-s}(x_1|x), \nabla_x p_s(x|x_0)\rangle \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\\ \cdot\frac{\int_{\mathbb{R}^{2d}} p_s(x|x_0) \nabla_x p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\;, \end{gather}\] \[\begin{gather} A_s^5(x)= 4\frac{\norm{\int_{\mathbb{R}^{2d}} \nabla_x p_s(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}^2 }{(p^{\mathrm{I}}_s(x))^2}\\ \cdot\frac{\int_{\mathbb{R}^{2d}} p_s(x|x_0) \nabla_x p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\;, \end{gather}\] \[\begin{gather} A_s^6(x)=4\frac{\int_{\mathbb{R}^{2d}} p_s(x|x_0) \nabla_x p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)} \\ \cdot\Bigg(\frac{\int_{\mathbb{R}^{2d}} \nabla_x p_s(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\Bigg)^{\operatorname{T}}\frac{\int_{\mathbb{R}^{2d}} p_s(x|x_0) \nabla_x p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)} \;. \end{gather}\] Therefore \[\require{upgreek} \begin{align} h (h^{1/8}+1)\int_{0}^{1-\epsilon_1}\mathbb{E}\Big[ \norm{(\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s (X^{\mathrm{M}}_s)}^2\Big]\uprho(s)\mathrm{d}s\lesssim h (h^{1/8}+1)\sum_{k=1}^6 \int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\norm{A_s^k(x)}^2 \uprho(s) p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\;. \end{align}\] Plugging the above inequality, 48 and 50 in 46 , we obtain \[\require{upgreek} \begin{align} \label{eq:mid95decomposition95kl} \mathrm{KL}(\nu^{\star}_{1-\epsilon_1}|\nu^{\theta^{\star}}_{1-\epsilon_1})\lesssim \varepsilon^2 +h\norm{\nabla \log \tilde{\pi}}^2_{\mathrm{L}^2(\pi)}+h (h^{1/8}+1)\sum_{k=1}^6 \int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\norm{A_s^k(x)}^2 \uprho(s) p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\;. \end{align}\tag{52}\]

We now bound each term \(A_s^k\) in the sum, starting with \(A_s^1\). Integrating by parts and using 20 and 22 , we get \[\require{upgreek} \begin{align} &\int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\norm{A_s^1(x)}^2 \uprho(s)p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\; \\ &\lesssim\int_0^{1-\epsilon_1} \int_{\mathbb{R}^d} \left\|\frac{\int_{\mathbb{R}^{2d}} \Delta_x p_s(x|x_0) \nabla_x p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\right.\\ &\left.\quad - \frac{\int_{\mathbb{R}^{2d}} \Delta_x p_s(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\right.\\ &\left. \quad \cdot\frac{\int_{\mathbb{R}^{2d}} p_s(x|x_0) \nabla_x p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\right\|^2 \uprho(s)p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\\ &=\int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d} \left\|\frac{\int_{\mathbb{R}^{2d}} \left\langle\nabla_x p_s(x|x_0),\nabla_{x_0}\tilde{\pi}(x_0,x_1) \right\rangle \nabla_x p_{1-s}(x_1|x) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\right.\\ &\left.\quad - \frac{\int_{\mathbb{R}^{2d}} \left\langle\nabla_x p_s(x|x_0),\nabla_{x_0}\tilde{\pi}(x_0,x_1) \right\rangle p_{1-s}(x_1|x) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\right.\\ &\left. \quad \cdot\frac{\int_{\mathbb{R}^{2d}} p_s(x|x_0) \nabla_x p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\right\|^2 \uprho(s) p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\\ &=\int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d} \left\|\frac{\int_{\mathbb{R}^{2d}} \left\langle\frac{x-x_0}{s},\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}} \right\rangle(x_0,x_1)\frac{x_1-x}{1-s} p_s(x|x_0) p_{1-s}(x_1|x)\tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\right.\\ &\left.\quad - \frac{\int_{\mathbb{R}^{2d}} \left\langle\frac{x-x_0}{s},\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(x_0,x_1) \right\rangle p_s(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\right.\\ &\left. \quad \cdot\frac{\int_{\mathbb{R}^{2d}} \frac{x_1-x}{1-s} p_s(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\right\|^2 \uprho(s) p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\\ &=\int_{0}^{1-\epsilon_1} \mathbb{E}\left[\left\|\mathbb{E}\left[\left\langle\frac{X^\mathrm{I}_s-X^\mathrm{I}_0}{s},\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1) \right\rangle\frac{X^\mathrm{I}_1-X^\mathrm{I}_s}{1-s}\Big|X^\mathrm{I}_s\right]\right.\right.\\ &\left.\left. \quad -\mathbb{E}\left[\left\langle\frac{X^\mathrm{I}_s-X^\mathrm{I}_0}{s},\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1) \right\rangle\Big|X^\mathrm{I}_s\right]\mathbb{E}\left[\frac{X^\mathrm{I}_1-X^\mathrm{I}_s}{1-s}\Big|X^\mathrm{I}_s\right]\right\|^2\right]\uprho(s)\mathrm{d}s\\ &\lesssim\label{eq:where95to95plug95measurability95property95in95conditional95expectation} \int_{0}^{1-\epsilon_1} \left\{\mathbb{E}\left[\left\|\mathbb{E}\left[\left\langle\frac{X^\mathrm{I}_s-X^\mathrm{I}_0}{s},\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1) \right\rangle\frac{X^\mathrm{I}_1-X^\mathrm{I}_s}{1-s}\Big|X^\mathrm{I}_s\right]\right\|^2\right]\right.\\ &\quad \left.+\mathbb{E}\left[\left\|\mathbb{E}\left[\left\langle\frac{X^\mathrm{I}_s-X^\mathrm{I}_0}{s},\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1) \right\rangle\Big|X^\mathrm{I}_s\right]\mathbb{E}\left[\frac{X^\mathrm{I}_1-X^\mathrm{I}_s}{1-s}\Big|X^\mathrm{I}_s\right]\right\|^2\right]\right\}\uprho(s)\mathrm{d}s\;. \end{align}\tag{53}\]

At this point, we split the time interval \([0,1-\epsilon_1]\) in two, \([0,1/2]\) and \([1/2, 1-\epsilon_1]\) and we rewrite the RHS as \[\require{upgreek} \begin{align} &\int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\norm{A_s^1(x)}^2 \uprho(s) p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\; \\ &\lesssim\int_{0}^{1/2} \left\{\mathbb{E}\left[\left\|\left\langle\frac{X^\mathrm{I}_s-X^\mathrm{I}_0}{s},\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1) \right\rangle X^\mathrm{I}_1-X^\mathrm{I}_s\right\|^2\right]\right.\\ &\quad \left.+\mathbb{E}\left[\left\|\mathbb{E}\left[\left\langle\frac{X^\mathrm{I}_s-X^\mathrm{I}_0}{s},\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1) \right\rangle\Big|X^\mathrm{I}_s\right]\mathbb{E}\left[X^\mathrm{I}_1-X^\mathrm{I}_s\Big|X^\mathrm{I}_s\right]\right\|^2\right]\right\}\uprho(s)\mathrm{d}s\\ &\quad + \int_{1/2}^{1-\epsilon_1} \left\{\mathbb{E}\left[\left\|\left\langle X^\mathrm{I}_s-X^\mathrm{I}_0,\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1) \right\rangle\frac{X^\mathrm{I}_1-X^\mathrm{I}_s}{1-s}\right\|^2\right]\right.\\ &\quad \left. +\mathbb{E}\left[\left\|\mathbb{E}\left[\left\langle X^\mathrm{I}_s-X^\mathrm{I}_0,\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1) \right\rangle\Big|X^\mathrm{I}_s\right]\mathbb{E}\left[\frac{X^\mathrm{I}_1-X^\mathrm{I}_s}{1-s}\Big|X^\mathrm{I}_s\right]\right\|^2\right]\right\}\uprho(s)\mathrm{d}s\;. \end{align}\]We first focus on the time sub-interval \(s\in[0,1/2]\). To this aim, we fix \(s\in [0,1/2]\). Holder and Jensen inequalities yield

\[\begin{align} & \mathbb{E}\left[\left\|\left\langle\frac{X^\mathrm{I}_s-X^\mathrm{I}_0}{s},\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1) \right\rangle X^\mathrm{I}_1-X^\mathrm{I}_s\right\|^2\right]\\ &\quad +\mathbb{E}\left[\left\|\mathbb{E}\left[\left\langle\frac{X^\mathrm{I}_s-X^\mathrm{I}_0}{s},\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1) \right\rangle\Big|X^\mathrm{I}_s\right]\mathbb{E}\left[X^\mathrm{I}_1-X^\mathrm{I}_s\Big|X^\mathrm{I}_s\right]\right\|^2\right]\\ &\le \mathbb{E}\left[\left\|\left\langle\frac{X^\mathrm{I}_s-X^\mathrm{I}_0}{s},\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1) \right\rangle\right\|^4\right]^{\frac{1}{2}}\mathbb{E}\left[\left\|X^\mathrm{I}_1-X^\mathrm{I}_s\right\|^{4}\right]^{\frac{1}{2}}\\ &\quad +\mathbb{E}\left[\left\|\mathbb{E}\left[\left\langle\frac{X^\mathrm{I}_s-X^\mathrm{I}_0}{s},\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1) \right\rangle\Big|X^\mathrm{I}_s\right]\right\|^4\right]^{\frac{1}{2}}\mathbb{E}\left[\left\|\mathbb{E}\left[X^\mathrm{I}_1-X^\mathrm{I}_s\Big|X^\mathrm{I}_s\right]\right\|^4\right]^{\frac{1}{2}}\\ &\lesssim \mathbb{E}\left[\left\|\left\langle\frac{X^\mathrm{I}_s-X^\mathrm{I}_0}{s},\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1) \right\rangle\right\|^4\right]^{\frac{1}{2}}\mathbb{E}\left[\left\|X^\mathrm{I}_1-X^\mathrm{I}_s\right\|^{4}\right]^{\frac{1}{2}}\\ &\le \mathbb{E}\left[\left\|\frac{X^\mathrm{I}_s-X^\mathrm{I}_0}{s}\right\|^{8}\right]^{\frac{1}{4}}\mathbb{E}\left[\left\|\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1)\right\|^{8}\right]^{\frac{1}{4}}\mathbb{E}\left[\left\|X^\mathrm{I}_1-X^\mathrm{I}_s\right\|^{4}\right]^{\frac{1}{2}}\;. \end{align}\] Therefore, we have that \[\require{upgreek} \begin{align} \int_{0}^{1/2} \int_{\mathbb{R}^d}\norm{A_s^1(x)}^2 \uprho(s) p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\; \lesssim \int_{0}^{1/2}\mathbb{E}\left[\left\|\frac{X^\mathrm{I}_s-X^\mathrm{I}_0}{s}\right\|^{8}\right]^{\frac{1}{4}}\mathbb{E}\left[\left\|\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1)\right\|^{8}\right]^{\frac{1}{4}}\mathbb{E}\left[\left\|X^\mathrm{I}_1-X^\mathrm{I}_s\right\|^{4}\right]^{\frac{1}{2}}\uprho(s) \mathrm{d}s\;. \end{align}\] Using 3, 1, 4, Equation 43 , 2 and the fact that \(\mathbf{m}_{8}[{\mu}], \mathbf{m}_{8}[{\nu^{\star}}]\lesssim d^4\), we get that \[\require{upgreek} \begin{align} &\int_{0}^{1/2} \int_{\mathbb{R}^d}\norm{A_s^1(x)}^2 \uprho(s) p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\; \\ &\lesssim \int_{0}^{1/2}\mathbb{E}\left[\left\|\frac{\overleftarrow{X}_{1}-\overleftarrow{X}_{1-s}}{s}\right\|^{8}\right]^{\frac{1}{4}}\mathbb{E}\left[\left\|\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(X^\mathrm{I}_0,X^\mathrm{I}_1)\right\|^{8}\right]^{\frac{1}{4}}\mathbb{E}\left[\left\|X^\mathrm{I}_1-X^\mathrm{I}_s\right\|^{4}\right]^{\frac{1}{2}}\uprho(s) \mathrm{d}s\\ &\lesssim \|\nabla\log \tilde{\pi}\|^2_{\mathrm{L}^{8}(\pi)} (d^2+\mathbf{m}_{4}[{\mu}]+\mathbf{m}_{4}[{\nu^{\star}}])^{\frac{1}{2}}\int_{0}^{1/2}\mathbb{E}\left[\left\|\frac{\overleftarrow{f}_{1-s}^s+\overleftarrow{g}_{1-s}^s}{s}\right\|^{8} \right]^{\frac{1}{4}}\uprho(s) \mathrm{d}s\\ &\lesssim \|\nabla\log \tilde{\pi}\|^2_{\mathrm{L}^{8}(\pi)} \left(d+\sqrt{\mathbf{m}_{4}[{\mu}]}+\sqrt{\mathbf{m}_{4}[{\nu^{\star}}]}\right)\int_{0}^{1/2}\left\{\mathbb{E}\left[\left\|\frac{\overleftarrow{f}_{1-s}^s}{s}\right\|^{8}\right]^{\frac{1}{4}}+ \mathbb{E}\left[\left\|\frac{\overleftarrow{g}_{1-s}^s}{s}\right\|^{8} \right]^{\frac{1}{4}}\right\}\uprho(s) \mathrm{d}s\\ &\lesssim \|\nabla\log \tilde{\pi}\|^2_{\mathrm{L}^{8}(\pi)} \left(d+\sqrt{\mathbf{m}_{4}[{\mu}]}+\sqrt{\mathbf{m}_{4}[{\nu^{\star}}]}\right) \left(\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^2+\sqrt[4]{\mathbf{m}_{8}[{\mu}]}+\sqrt[4]{\mathbf{m}_{8}[{\nu^{\star}}]}+\int_{0}^{1/2}\{d^4 s^{4-8}\}^{\frac{1}{4}}s^{\frac{7}{8}} \mathrm{d}s\right)\\ &=\|\nabla\log \tilde{\pi}\|^2_{\mathrm{L}^{8}(\pi)} \left(d+\sqrt{\mathbf{m}_{4}[{\mu}]}+\sqrt{\mathbf{m}_{4}[{\nu^{\star}}]}\right) \left(\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^2+\sqrt[4]{\mathbf{m}_{8}[{\mu}]}+\sqrt[4]{\mathbf{m}_{8}[{\nu^{\star}}]}+d \int_{0}^{1/2} s^{-\frac{1}{8}}\mathrm{d}s \right)\\ &\lesssim \|\nabla\log \tilde{\pi}\|^2_{\mathrm{L}^{8}(\pi)} \left(d+\sqrt{\mathbf{m}_{4}[{\mu}]}+\sqrt{\mathbf{m}_{4}[{\nu^{\star}}]}\right) \left(\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^2+\sqrt[4]{\mathbf{m}_{8}[{\mu}]}+\sqrt[4]{\mathbf{m}_{8}[{\nu^{\star}}]}+d\right)\\ &\lesssim \left(\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^2+\sqrt[4]{\mathbf{m}_{8}[{\mu}]}+\sqrt[4]{\mathbf{m}_{8}[{\nu^{\star}}]}\right) \left(d+\sqrt{\mathbf{m}_{4}[{\mu}]}+\sqrt{\mathbf{m}_{4}[{\nu^{\star}}]}\right) \left(\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^2+\sqrt[4]{\mathbf{m}_{8}[{\mu}]}+\sqrt[4]{\mathbf{m}_{8}[{\nu^{\star}}]}+d\right)\\ &\lesssim\left(\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^2+d\right) d\left(\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^2+d\right)\lesssim d \left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right)\;. \end{align}\] To handle with the time sub-interval \(s\in[1/2, 1-\epsilon_1]\), we proceed in a similar way, but this time using [lemma:lemma2_for_in_dfm,lemma:lemma3_for_in_dfm] rather than [lemma:lemma2_back_in_dfm,lemma:lemma3_back_in_dfm]. By doing so we obtain that \[\require{upgreek} \begin{align} \int_{1/2}^{1-\epsilon_1} \norm{A_s^1(x)}^2 \uprho(s) p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\lesssim d \left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right) \;. \end{align}\] Putting together the two bounds we get \[\require{upgreek} \begin{align} \int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\norm{A_s^1(x)}^2 \uprho(s) p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\lesssim d \left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right) \;. \end{align}\] Proceeding in a similar way (we omit the argument, as it is almost a duplication of the previous one), we get \[\require{upgreek} \begin{align} \int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\norm{A_s^2(x)}^2 \uprho(s) p_s^{\mathrm{I}} (x)\mathrm{d}x\; \mathrm{d}s\lesssim d \left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right) \;. \end{align}\] We now switch to \(A_s^3\). Using 22 , integration by parts, we have that \[\require{upgreek} \begin{align} &\int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\norm{A_s^3(x)}^2 \uprho(s)p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\; \\ &\lesssim \int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\left\|\frac{\int_{\mathbb{R}^{2d}} \nabla_x p_{1-s}(x_1|x)(\nabla_x p_s(x|x_0))^{\operatorname{T}} \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\right.\\ &\left. \quad \cdot \frac{ \int_{\mathbb{R}^{2d}} \nabla_x p_s(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\right\|^2\uprho(s)p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\;\\ & \lesssim\int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\left\|\frac{\int_{\mathbb{R}^{2d}} \frac{x_1-x}{1-s}\left(\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(x_0,x_1) \right)^{\operatorname{T}} p_{1-s}(x_1|x) p_s(x|x_0) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\right.\\ &\left. \quad \cdot\frac{ \int_{\mathbb{R}^{2d}} \frac{x-x_0}{s} p_s(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)}\right\|^2\uprho(s)p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\;\\ &=\int_{0}^{1-\epsilon_1} \mathbb{E}\left[\left\|\mathbb{E}\left[\frac{X^\mathrm{I}_1-X^\mathrm{I}_s}{1-s}\left(\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}\left(X^\mathrm{I}_0, X^\mathrm{I}_1\right)\right)^{\operatorname{T}}\Big| X^\mathrm{I}_s\right]\mathbb{E}\left[\frac{X^\mathrm{I}_s-X^\mathrm{I}_0}{s}\Big| X^\mathrm{I}_s\right]\right\|^2\right]\uprho(s) \mathrm{d}s \end{align}\]Mimicking the argument used to bound \(A^1_s\) (we omit the details, as they are almost a duplication of the previous one), we eventually get that

\[\require{upgreek} \begin{align} \int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\norm{A_s^3(x)}^2 \uprho(s)p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\; \lesssim d \left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right) \;. \end{align}\] In the very same way, we get \[\require{upgreek} \begin{align} &\int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\norm{A_s^4(x)}^2 \uprho(s)p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\; \lesssim d \left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right) \;. \end{align}\] We now turn to \(A^5_s.\) Using 22 and integration by parts, we obtain \[\require{upgreek} \begin{align} &\int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\norm{A_s^5(x)}^2 \uprho(s)p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\; \\ &\lesssim\int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\left\| \frac{\norm{\int_{\mathbb{R}^{2d}} \nabla_x p_s(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}^2 }{(p^{\mathrm{I}}_s(x))^2}\right.\\ &\left. \quad \cdot\frac{\int_{\mathbb{R}^{2d}} p_s(x|x_0) \nabla_x p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)} \right\|^2 \uprho(s)p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\\ &=\int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\left\| \left\langle \frac{\int_{\mathbb{R}^{2d}} \frac{x-x_0}{s} p_s(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1 }{p^{\mathrm{I}}_s(x)},\right.\right.\\ &\left.\left. \cdot\frac{\int_{\mathbb{R}^{2d}} \frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}(x_0,x_1) p_s(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1 }{p^{\mathrm{I}}_s(x)}\right\rangle \right.\\ &\left. \quad \cdot\frac{\int_{\mathbb{R}^{2d}} \frac{x_1-x}{1-s}p_s(x|x_0) p_{1-s}(x_1|x) \tilde{\pi}(x_0,x_1) \mathrm{d}x_0 \mathrm{d}x_1}{p^{\mathrm{I}}_s(x)} \right\|^2 \uprho(s)p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\\ &=\int_{0}^{1-\epsilon_1} \mathbb{E}\left[\left\|\left\langle\mathbb{E}\left[\frac{X^\mathrm{I}_s-X^\mathrm{I}_0}{s}\Big|X^\mathrm{I}_s\right],\mathbb{E}\left[\frac{\nabla_{x_0}\tilde{\pi}}{\tilde{\pi}}\left(X^\mathrm{I}_0, X^\mathrm{I}_1\right)\Big|X^\mathrm{I}_s\right] \right\rangle\mathbb{E}\left[\frac{X^\mathrm{I}_1-X^\mathrm{I}_s}{1-s}\Big|X^\mathrm{I}_s\right]\right\|^2\right] \uprho(s) \mathrm{d}s\;. \end{align}\] Proceeding as for \(A^1_s\), we eventually get \[\require{upgreek} \begin{align} \int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\norm{A_s^5(x)}^2 \uprho(s)p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\; \lesssim d \left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right) \;. \end{align}\] The argument to bound \(A^6_s\) is almost the same (hence omitted) and leads to \[\require{upgreek} \begin{align} \int_{0}^{1-\epsilon_1} \int_{\mathbb{R}^d}\norm{A_s^6(x)}^2 \uprho(s)p_s^{\mathrm{I}} (x)\mathrm{d}x \mathrm{d}s\; \lesssim d \left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right) \;. \end{align}\] Putting together the bounds on the \(\{A_s^k\}_{k=1}^{6}\) derived so far, we eventually obtain

\[\require{upgreek} \begin{align} \label{term95495dec95KL95gen} h(h^{1/8}+1)\int_{0}^{1-\epsilon_1}\mathbb{E}\Big[ \norm{(\partial_s +\mathcal{L}^\mathrm{M}_s)\tilde{\beta}_s (\overrightarrow{X}_s)}^2\Big]\uprho(s) \mathrm{d}s\lesssim h(h^{1/8}+1)d \left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right)\;. \end{align}\tag{54}\] Plugging 54 , 47 and 49 into 46 , we get \[\begin{align} \mathrm{KL}(\nu^{\star}_{1-\epsilon_1}|\nu^{\theta^{\star}}_{1-\epsilon_1})\lesssim \varepsilon^2 +h(h^{1/8}+1)d \left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right)\;. \end{align}\] The byproduct of the above estimate and 37 leads to \[\begin{align} \mathrm{KL}(\nu^{\star}| \nu^{\theta^\star}_1)\lesssim \varepsilon^2 + h(h^{1/8}+1) d \left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right)\;. \end{align}\] ◻

7.2 Early-stopping regime with constant step-size↩︎

Proof of 2:. Fix \(0<\delta<1/2\). We want to apply 1 to \(\mu\), \(\nu^{\star}_{1-\delta}\) and the coupling \(\pi_{1-\delta}\in \Pi(\mu, \nu^{\star}_{1-\delta})\) defined as \[\begin{align} \pi_{1-\delta}(x_0, x_{1-\delta})=\int_{\mathbb{R}^d} p^{\mathrm{I}}_{1-\delta|0,1}(x_{1-\delta}|x_0, x_1) \pi_{0|1}(x_0| x_1)\nu^{\star}(\mathrm{d}x_1)\;, \end{align}\]where \((x_0,x_1,x_{1-\delta}) \mapsto p^{\mathrm{I}}_{1-\delta|0,1}(x_{1-\delta}|x_0, x_1)\) denotes the density of \(X^\mathrm{I}_{1-\delta}\) given \((X^\mathrm{I}_0, X^\mathrm{I}_1)\) with respect to the Lebesgue measure. To this aim, we need to prove that \(\nu^{\star}_{1-\delta}\) has finite 8th order moment and that \(\pi_{1-\delta}\) satisfies \[\begin{align} \|\nabla\log\pi_{1-\delta}\|_{\mathrm{L}^8(\pi_{1-\delta})}<+\infty\;. \end{align}\] We start with \(\nu^{\star}_{1-\delta}\). Note that, as a very consequence of 15 , we have that \[\begin{align} \label{bound95moment95hat} \mathbf{m}_{8}[{\nu^{\star}_{1-\delta}}]= \mathbb{E}\left[\norm{X^{\mathrm{I}}_{1-\delta}}^8\right]\lesssim \delta^8 \mathbf{m}_{8}[{\mu}]+(1-\delta)^8\mathbf{m}_{8}[{\nu^{\star}}]+ d^4 \delta^4(1-\delta)^4<+\infty\;. \end{align}\tag{55}\]

We now switch to \(\pi_{1-\delta}\). It follows from the very definition of the stochastic interpolant that, \[p^{\mathrm{I}}_{1-\delta|0,1}(x_{1-\delta}|x_0, x_1)= \frac{1}{(4\pi \delta(1-\delta))^{d/2}} \exp\Bigg(-\frac{\norm{x_{1-\delta}-\delta x_0 -(1-\delta)x_1}^2}{4\delta(1-\delta)}\Bigg)\;.\]Therefore, \[\nabla_{x_0} p^{\mathrm{I}}_{1-\delta|0,1}(x_{1-\delta}|x_0, x_1) = \frac{x_{1-\delta}-\delta x_0 -(1-\delta)x_1}{2(1-\delta)}p^{\mathrm{I}}_{1-\delta|0,1}(x_{1-\delta}|x_0, x_1)\;,\]and \[\nabla_{x_{1-\delta}} p^{\mathrm{I}}_{1-\delta|0,1}(x_{1-\delta}|x_0, x_1)= -\frac{x_{1-\delta}-\delta x_0 -(1-\delta)x_1}{2\delta(1-\delta)}p^{\mathrm{I}}_{1-\delta|0,1}(x_{1-\delta}|x_0, x_1)\;.\] Furthermore, \[\begin{align} \frac{p^{\mathrm{I}}_{1-\delta|0,1}(x_{1-\delta}|x_0, x_1)\pi_{0|1}(x_0| x_1)\nu^{\star}(\mathrm{d}x_1)}{\int_{\mathbb{R}^d}p^{\mathrm{I}}_{1-\delta|0,1}(x_{1-\delta}|x_0, \tilde{x}_1)\pi_{0|1}(x_0| \tilde{x}_1)\nu^{\star}(\mathrm{d}\tilde{x}_1)}&= p^{\mathrm{I}}_{1|0, 1-\delta}(\mathrm{d}x_1|x_0, x_{1-\delta})\;. \end{align}\]Consequently, we have that \[\begin{align} &\frac{\nabla_{x_0}\pi_{1-\delta}}{\pi_{1-\delta}}(x_0, x_{1-\delta})\\ &= \frac{\int_{\mathbb{R}^d} p^{\mathrm{I}}_{1-\delta|0,1}(x_{1-\delta}|x_0, x_1)\nabla_{x_0}\pi_{0|1}(x_0| x_1)\nu^{\star}(\mathrm{d}x_1) + \int_{\mathbb{R}^d} \nabla_{x_0} p^{\mathrm{I}}_{1-\delta|0,1}(x_{1-\delta}|x_0, x_1) \pi_{0|1}(x_0| x_1)\nu^{\star}(\mathrm{d}x_1)}{\int_{\mathbb{R}^d} p^{\mathrm{I}}_{1-\delta|0,1}(x_{1-\delta}|x_0, \tilde{x}_1) \pi_{0|1}(x_0| \tilde{x}_1)\nu^{\star}(\mathrm{d}\tilde{x}_1)}\\ &= \int_{\mathbb{R}^d}\nabla_{x_0}\log\pi_{0|1}(x_0| x_1) p^{\mathrm{I}}_{1|0, 1-\delta}(\mathrm{d}x_1|x_0, x_{1-\delta}) + \int_{\mathbb{R}^d}\frac{x_{1-\delta}-\delta x_0 -(1-\delta)x_1}{2(1-\delta)} p^{\mathrm{I}}_{1|0, 1-\delta}(\mathrm{d}x_1|x_0, x_{1-\delta}) \;, \end{align}\] and that \[\begin{align} \frac{\nabla_{x_{1-\delta}}\pi_{1-\delta}}{\pi_{1-\delta}}(x_0, x_{1-\delta})&= \frac{ \int_{\mathbb{R}^d} \nabla_{x_{1-\delta}} p^{\mathrm{I}}_{1-\delta|0,1}(x_{1-\delta}|x_0, x_1) \pi_{0|1}(x_0| x_1)\nu^{\star}(\mathrm{d}x_1)}{\int_{\mathbb{R}^d} p^{\mathrm{I}}_{1-\delta|0,1}(x_{1-\delta}|x_0, \tilde{x}_1) \pi_{0|1}(x_0| \tilde{x}_1)\nu^{\star}(\mathrm{d}\tilde{x}_1)}\\ &=-\int_{\mathbb{R}^{d}}\frac{x_{1-\delta}-\delta x_0 -(1-\delta)x_1}{2\delta(1-\delta)}p^{\mathrm{I}}_{1|0, 1-\delta}(\mathrm{d}x_1|x_0, x_{1-\delta}) \;. \end{align}\]But then, if we use Jensen inequality and 55 , we get \[\begin{align} \int_{\mathbb{R}^{2d}} \norm{\nabla_{x_0}\log \frac{\mathrm{d}\pi_{1-\delta}}{\mathrm{d}\mathrm{Leb}^{2d}}}^8 \mathrm{d}\pi_{1-\delta}& \lesssim \mathbb{E}\Bigg[\norm{\mathbb{E}\Bigg[\nabla_{x_0}\log\pi_{0|1}(X^{\mathrm{I}}_0| X^{\mathrm{I}}_{1})\Bigg|(X^{\mathrm{I}}_0, X^{\mathrm{I}}_{1-\delta})\Bigg]}^8\Bigg] \\ &\quad + \mathbb{E}\Bigg[\norm{\mathbb{E}\Bigg[\frac{X^{\mathrm{I}}_{1-\delta}-\delta X^{\mathrm{I}}_0 -(1-\delta)X^{\mathrm{I}}_1}{1-\delta}\Bigg|(X^{\mathrm{I}}_0, X^{\mathrm{I}}_{1-\delta})\Bigg]}^8\Bigg] \\ &\lesssim \norm{\nabla\log\pi_{0|1}}^8_{\mathrm{L}^8(\pi_{0|1})}+ \mathbb{E}\Bigg[\norm{\frac{X^{\mathrm{I}}_{1-\delta}-\delta X^{\mathrm{I}}_0 -(1-\delta)X^{\mathrm{I}}_1}{1-\delta}}^8\Bigg] \\ &\lesssim \norm{\nabla\log\pi_{0|1}}^8_{\mathrm{L}^8(\pi_{0|1})} +\mathbf{m}_{8}[{\nu^{\star}_{1-\delta}}]\frac{1}{(1-\delta)^8}+ \mathbf{m}_{8}[{\mu}]\frac{\delta^8}{(1-\delta)^8}+\mathbf{m}_{8}[{\nu^{\star}}]\\ &\lesssim \norm{\nabla\log\pi_{0|1}}^8_{\mathrm{L}^8(\pi_{0|1})}+ \mathbf{m}_{8}[{\mu}]\frac{\delta^8}{(1-\delta)^8}+ \mathbf{m}_{8}[{\nu^{\star}}] +d^4\frac{\delta^4}{(1-\delta)^4}\\ &\lesssim \norm{\nabla\log\pi_{0|1}}^8_{\mathrm{L}^8(\pi_{0|1})}+\mathbf{m}_{8}[{\mu}]\frac{1}{(1-\delta)^8}+ \mathbf{m}_{8}[{\nu^{\star}}]\frac{1}{\delta^8} +d^4\frac{1}{\delta^4(1-\delta)^4}\;. \end{align}\] and (similarly) \[\begin{align} \int_{\mathbb{R}^{2d}} \norm{\nabla_{x_{1-\delta}}\log \frac{\mathrm{d}\pi_{1-\delta}}{\mathrm{d}\mathrm{Leb}^{2d}}}^8 \mathrm{d}\pi_{1-\delta}& \lesssim \mathbb{E}\Bigg[\norm{\mathbb{E}\Bigg[\frac{X^{\mathrm{I}}_{1-\delta}-\delta X^{\mathrm{I}}_0 -(1-\delta)X^{\mathrm{I}}_1}{\delta(1-\delta)}\Bigg| (X^{\mathrm{I}}_0, X^{\mathrm{I}}_{1-\delta})\Bigg]}^8\Bigg] \\ &\lesssim \mathbf{m}_{8}[{\mu}]\frac{1}{(1-\delta)^8}+ \mathbf{m}_{8}[{\nu^{\star}}]\frac{1}{\delta^8} +d^4\frac{1}{\delta^4(1-\delta)^4}\;. \end{align}\] Plugging these estimates in the convergence bound provided in 1 and using the fact that \(\mathbf{m}_{8}[{\mu}], \mathbf{m}_{8}[{\nu^{\star}}]\lesssim d^4\) allows to conclude. ◻

Proof of 1:. It follows from the very definition of Fortet- Mourier distance, triangle inequality and Pinsker’s inequality that \[\begin{align} \label{bound:first95cor} \mathscr{W}_{2,\text{FM}}^2(\nu^{\star},\nu^{\theta^{\star}}_{1-\delta} )\lesssim \mathscr{W}_{2}^2(\nu^{\star},\nu^{\star}_{1-\delta} )+ \mathrm{TV}^2(\nu^{\star}_{1-\delta}, \nu^{\theta^\star}_{1-\delta})\lesssim\mathscr{W}_{2}^2(\nu^{\star},\nu^{\star}_{1-\delta} )+\mathrm{KL}(\nu^{\star}_{1-\delta} | \nu^{\theta^\star}_{1-\delta})\;. \end{align}\tag{56}\] with \(\mathrm{TV}\) denoting the total variation distance. It follows from 15 , that \[\begin{align} \mathscr{W}_{2}^2(\nu^{\star},\nu^{\star}_{1-\delta} )\lesssim \delta^2\mathbf{m}_{2}[{\mu}]+\delta^2\mathbf{m}_{2}[{\nu^{\star}}]+\delta(1-\delta)d\;. \end{align}\] Plugging this and the bound provided in 2 into 56 leads to \[\begin{align} \mathscr{W}_{2,\text{FM}}^2(\nu^{\star},\nu^{\theta^{\star}}_{1-\delta} )\lesssim \delta^2\mathbf{m}_{2}[{\mu}]+\delta^2\mathbf{m}_{2}[{\nu^{\star}}]+\delta d+ \varepsilon^2 + h (h+1) \Bigg(\frac{d^2}{\delta^4} +\norm{\nabla\log\pi_{0|1}}^4_{\mathrm{L}^8(\pi_{0|1})} \Bigg)d\;. \end{align}\] Therefore, if \(\delta = \mathcal{O}( \varepsilon^2/d)\), when choosing \(h= \mathcal{O}(\varepsilon^{10}/d^{7})\), we have that \[\begin{align} \mathscr{W}_{2,\text{FM}}^2(\nu^{\star},\nu^{\theta^{\star}}_{1-\delta} )\lesssim \mathcal{O}(\varepsilon^2)\;, \end{align}\]which proves the desired bound. ◻

7.3 Early-stopping regime with novel step-size schedule↩︎

Proof of 3:. Our aim is to apply 1 on the sub-partition \(\{t_k\}_{k=0}^{M_h}\) of \([0,1/2]\) with constant step size \(h_k=h\) for \(k\le M_h\), and Proposition 5.1 in [38] on the sub-partition \(\{t_k\}_{k=M_h}^{M_h+N}\) of \([1/2,1]\) with exponentially decreasing step sizes \(h_k=h\min\{t_k, 1-t_k\}\) for \(M_h<k\le M_h+N\). Since the hypothesis of Proposition 5.1 in [38] are satisfied in our setting, we immediately get that \[\begin{align} \label{eq:first95part95expo} \mathrm{KL}(\nu^{\star}|\nu^{\theta^\star}_{1})&\lesssim \mathrm{KL}(\nu^{\star}_{1/2}|\nu^{\theta^\star}_{1/2}) + \varepsilon^2 +h^2d^3+h^2d^3\log\frac{1}{\delta} +hd^2+hd^2 \log\frac{1}{\delta}\\ &=\mathrm{KL}(\nu^{\star}_{1/2}|\nu^{\theta^\star}_{1/2}) + \varepsilon^2 +h^2d^3\left(1+\log\frac{1}{\delta}\right)+ hd^2\left(1+\log\frac{1}{\delta}\right)\\ &=\mathrm{KL}(\nu^{\star}_{1/2}|\nu^{\theta^\star}_{1/2}) + \varepsilon^2 +\left(1+\log\frac{1}{\delta}\right)\left(h^2d^3+ hd^2\right) \\ &\lesssim \mathrm{KL}(\nu^{\star}_{1/2}|\nu^{\theta^\star}_{1/2}) + \varepsilon^2 +hd^3\left(1+\log\frac{1}{\delta}\right)\\ &\lesssim \mathrm{KL}(\nu^{\star}_{1/2}|\nu^{\theta^\star}_{1/2}) + \varepsilon^2 + h d^3\log\frac{1}{\delta}\;, \end{align}\tag{57}\] with \(\nu^{\theta^\star}_{1/2}\) denoting the law of \(X^{\theta^{\star}}_{1/2}.\) In order to apply 1, we first need to prove that the probability distribution \(\nu^{\star}_{1/2}=\mathcal{L}\left(X^{\mathrm{I}}_{1/2}\right)\) and the coupling \[\begin{align} \pi_{0, 1/2}(x_0, x_{1/2})=\int_{\mathbb{R}^{d}} p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1) \pi_{0|1}(x_0| x_1)\nu^{\star}(\mathrm{d}x_1)\in\Pi(\mu, \nu^{\star}_{1/2})\;, \end{align}\]with \((x_0,x_1,x_{1/2}) \mapsto p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1)\) denoting the density of \(X^\mathrm{I}_{1/2}\) given \((X^\mathrm{I}_0, X^\mathrm{I}_1)\) with respect to the Lebesgue measure, satisfy \(\mathbf{m}_{8}[{\nu^{\star}}]<+\infty,\) and \[\begin{align} \|\nabla\log \pi_{0, 1/2}\|_{\mathrm{L}^8(\pi_{0, 1/2})}<+\infty\;, \end{align}\]respectively. On the one side, as a very consequence of 15 , we have that \[\begin{align} \label{zwbxfduh} \mathbf{m}_{8}[{\nu^{\star}_{1/2}}]= \mathbb{E}\left[\norm{X^{\mathrm{I}}_{1/2}}^8\right]\lesssim \mathbf{m}_{8}[{\mu}]+\mathbf{m}_{8}[{\nu^{\star}}]+ d^4<+\infty\;. \end{align}\tag{58}\] On the other side, being \[\begin{align} &\nabla_{x_0} \pi_{0, 1/2}(x_0, x_{1/2})\\ &=\int_{\mathbb{R}^{d}} \nabla_{x_0} p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1) \pi_{0|1}(x_0| x_1)\nu^{\star}(\mathrm{d}x_1)+ \int_{\mathbb{R}^{d}} p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1) \nabla_{x_0}\pi_{0|1}(x_0| x_1)\nu^{\star}(\mathrm{d}x_1)\;, \end{align}\]we have that \[\begin{align} &\nabla_{x_0} \log \pi_{0, 1/2}(x_0, x_{1/2})\\ &= \frac{\int_{\mathbb{R}^{d}} \nabla_{x_0}\log p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1) p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1)\pi_{0|1}(x_0| x_1)\nu^{\star}(\mathrm{d}x_1)}{\int_{\mathbb{R}^{d}} p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, \tilde{x}_1) \pi_{0|1}(x_0| \tilde{x}_1)\nu^{\star}(\mathrm{d}\tilde{x}_1)}\\ &\quad +\frac{\int_{\mathbb{R}^{d}} \nabla_{x_0}\log \pi_{0|1}(x_0| x_1) p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1) \pi_{0|1}(x_0| x_1)\nu^{\star}(\mathrm{d}x_1)}{\int_{\mathbb{R}^{d}} p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, \tilde{x}_1) \pi_{0|1}(x_0| \tilde{x}_1)\nu^{\star}(\mathrm{d}\tilde{x}_1)}\;. \end{align}\] Since \[\begin{align} \frac{p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1)\pi_{0|1}(x_0| x_1)\nu^{\star}(\mathrm{d}x_1)}{\int_{\mathbb{R}^{d}} p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, \tilde{x}_1) \pi_{0|1}(x_0| \tilde{x}_1)\nu^{\star}(\mathrm{d}\tilde{x}_1)}=p^{\mathrm{I}}_{1|0, 1/2}(\mathrm{d}x_1|x_0, x_{1/2})\;, \end{align}\] we can rewrite \(\nabla_{x_0}\log\pi_{0,1/2}\) as \[\begin{align} \label{eq:to95plug95in} \nabla_{x_0} \log \pi_{0, 1/2}(x_0, x_{1/2})&= \int_{\mathbb{R}^{d}} \nabla_{x_0}\log p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1) p^{\mathrm{I}}_{1|0, 1/2}(\mathrm{d}x_1|x_0, x_{1/2})\\ &\quad +\int_{\mathbb{R}^{d}} \nabla_{x_0}\log \pi_{0|1}(x_0| x_1) p^{\mathrm{I}}_{1|0, 1/2}(\mathrm{d}x_1|x_0, x_{1/2})\;. \end{align}\tag{59}\] It follows from the very definition of the stochastic interpolant that, \[\label{eq:conditional95density95interpolant95mid95interval} p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1)= \frac{1}{\pi ^{d/2}} \exp\Bigg(-\norm{x_{1/2}-\frac{1}{2} x_0 -\frac{1}{2}x_1}^2\Bigg)\;.\tag{60}\] Hence \[\begin{align} \nabla_{x_0} \log p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1)=-\frac{1}{2} x_0-\frac{1}{2} x_1+x_{1/2}\;. \end{align}\] Plugging this equality in 59 , we get \[\begin{align} \nabla_{x_0} \log \pi_{0, 1/2}(x_0, x_{1/2})&= \int_{\mathbb{R}^{d}} \left(-\frac{1}{2} x_0-\frac{1}{2} x_1+x_{1/2} \right) p^{\mathrm{I}}_{1|0, 1/2}(\mathrm{d}x_1|x_0, x_{1/2})\\ &\quad +\int_{\mathbb{R}^{d}} \nabla_{x_0}\log \pi_{0|1}(x_0| x_1) p^{\mathrm{I}}_{1|0, 1/2}(\mathrm{d}x_1|x_0, x_{1/2})\;. \end{align}\] Therefore, using Jensen inequality we get that \[\begin{align} & \|\nabla_{x_0}\log \pi_{0, 1/2}\|_{\mathrm{L}^8(\pi_{0, 1/2})}^8\\ &\lesssim \mathbb{E}\left[\left\|\mathbb{E}\left[-\frac{1}{2} X^\mathrm{I}_0-\frac{1}{2} X^\mathrm{I}_1+X^\mathrm{I}_{1/2}\Bigg| (X^\mathrm{I}_0, X^\mathrm{I}_{1/2})\right]\right\|^8\right]+\mathbb{E}\left[\left\|\mathbb{E}\left[\nabla_{x_0}\log \pi_{0|1}(X^\mathrm{I}_0|X^\mathrm{I}_1)\Bigg| (X^\mathrm{I}_0, X^\mathrm{I}_{1/2})\right]\right\|^8\right]\\ &\lesssim \mathbb{E}\left[\left\|-\frac{1}{2} X^\mathrm{I}_0-\frac{1}{2} X^\mathrm{I}_1+X^\mathrm{I}_{1/2}\right\|^8\right]+\mathbb{E}\left[\left\|\nabla_{x_0}\log \pi_{0|1}(X^\mathrm{I}_0|X^\mathrm{I}_1)\right\|^8\right]\;. \end{align}\] But then, leveraging 55 , we obtain \[\begin{align} \|\nabla_{x_0}\log \pi_{0, 1/2}\|_{\mathrm{L}^8(\pi_{0, 1/2})}\lesssim \sqrt[8]{\mathbf{m}_{8}[{\mu}]}+\sqrt[8]{\mathbf{m}_{8}[{\nu^{\star}}]}+ \|\nabla\log \pi_{0|1}\|_{\mathrm{L}^8(\pi_{0|1})}\lesssim \sqrt{d}+\|\nabla\log \pi_{0|1}\|_{\mathrm{L}^8(\pi_{0|1})}\;. \end{align}\] Similarly, being \[\begin{align} \nabla_{x_{1/2}} \pi_{0, 1/2}(x_0, x_{1/2})=\int_{\mathbb{R}^{d}} \nabla_{x_{1/2}} p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1) \pi_{0|1}(x_0| x_1)\nu^{\star}(\mathrm{d}x_1)\;, \end{align}\] we have that \[\begin{align} &\nabla_{x_{1/2}} \log \pi_{0, 1/2}(x_0, x_{1/2})\\ &=\frac{\int_{\mathbb{R}^{d}} \nabla_{x_{1/2}}\log p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1) p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1)\pi_{0|1}(x_0| x_1)\nu^{\star}(\mathrm{d}x_1)}{\int_{\mathbb{R}^{d}} \nabla_{x_{1/2}} p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, \tilde{x}_1) \pi_{0|1}(x_0| \tilde{x}_1)\nu^{\star}(\mathrm{d}\tilde{x}_1)}\\ &= \int_{\mathbb{R}^{d}} \nabla_{x_{1/2}}\log p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1) p^{\mathrm{I}}_{1|0,1/2}(\mathrm{d}x_1|x_0, x_{1/2})\;. \end{align}\] It follows from 60 that \[\begin{align} \nabla_{x_{1/2}}\log p^{\mathrm{I}}_{1/2|0,1}(x_{1/2}|x_0, x_1) = -2x_{1/2}+x_0+x_1\;. \end{align}\] At this point, proceeding as before, we get \[\begin{align} \|\nabla_{x_{1/2}}\log \pi_{0, 1/2}\|_{\mathrm{L}^8(\pi_{0, 1/2})}\lesssim \sqrt[8]{\mathbf{m}_{8}[{\mu}]}+\sqrt[8]{\mathbf{m}_{8}[{\nu^{\star}}]}\lesssim \sqrt{d}\;. \end{align}\] Therefore, we have that \[\begin{align} \|\nabla\log \pi_{0, 1/2}\|_{\mathrm{L}^8(\pi_{0, 1/2})}\lesssim \sqrt{d}+ \|\nabla\log \pi_{0|1}\|_{\mathrm{L}^8(\pi_{0|1})}\;. \end{align}\] Recalling 1 and applying 1 on the time interval \([0,1/2]\), with prior \(\mu\), target \(\nu^{\star}_{1/2}\) and coupling \(\pi_{0, 1/2}\), we get \[\begin{align} \mathrm{KL}(\nu^{\star}_{1/2}|\nu^{\theta^\star}_{1/2})\lesssim \varepsilon^2 +h(h^{1/8}+1)\Big(d^2+ \|\nabla\log \pi_{0|1}\|_{\mathrm{L}^8(\pi_{0|1})}^4 \Big)d\;. \end{align}\] Combining the above bound with 57 yields \[\begin{align} \mathrm{KL}(\nu^{\star}|\nu^{\theta^\star}_{1})\lesssim \varepsilon^2 +h(h^{1/8}+1)\Big(d^2+ \|\nabla\log \pi_{0|1}\|_{\mathrm{L}^8(\pi_{0|1})}^4 \Big)d+h d^3\log\frac{1}{\delta}\;. \end{align}\] ◻

Proof of 2:. It follows from the very definition of Fortet- Mourier distance, triangle inequality and Pinsker’s inequality that \[\begin{align} \label{bound:first95cor95expo} \mathscr{W}_{2,\text{FM}}^2(\nu^{\star},\nu^{\theta^{\star}}_{1-\delta} )\lesssim \mathscr{W}_{2}^2(\nu^{\star},\nu^{\star}_{1-\delta} )+ \mathrm{TV}^2(\nu^{\star}_{1-\delta}, \nu^{\theta^\star}_{1-\delta})\lesssim\mathscr{W}_{2}^2(\nu^{\star},\nu^{\star}_{1-\delta} )+\mathrm{KL}(\nu^{\star}_{1-\delta} | \nu^{\theta^\star}_{1-\delta})\;. \end{align}\tag{61}\] with \(\mathrm{TV}\) denoting the total variation distance. It follows from 15 , that \[\begin{align} \mathscr{W}_{2}^2(\nu^{\star},\nu^{\star}_{1-\delta} )\lesssim \delta^2\mathbf{m}_{2}[{\mu}]+\delta^2\mathbf{m}_{2}[{\nu^{\star}}]+\delta(1-\delta)d\;. \end{align}\] Plugging this and the bound provided in 3 into 61 leads to \[\begin{align} \mathscr{W}_{2,\text{FM}}^2(\nu^{\star},\nu^{\theta^{\star}}_{1-\delta} )\lesssim \delta^2\mathbf{m}_{2}[{\mu}]+\delta^2\mathbf{m}_{2}[{\nu^{\star}}]+\delta d+ \varepsilon^2 +h d^3\log\frac{1}{\delta} +h(h^{1/8}+1)\Big(d^2+ \|\nabla\log \pi_{0|1}\|_{\mathrm{L}^8(\pi_{0|1})}^4 \Big)d\;. \end{align}\] Therefore, if \(\delta = \mathcal{O}( \varepsilon^2/d)\), when choosing \(h= \tilde{\mathcal{O}}(\varepsilon^{2}/d^{3})\), we have that \[\begin{align} \mathscr{W}_{2,\text{FM}}^2(\nu^{\star},\nu^{\theta^{\star}}_{1-\delta} )\lesssim \mathcal{O}(\varepsilon^2)\;, \end{align}\]which proves the desired bound. ◻

8 Convergence Bounds in Wasserstein-2 Distance↩︎

8.1 Strong Log-Concave and Full Log-Lipschitz Distributions↩︎

For sake of clarity, we first prove 13 under 9, i.e., we prove the following result.

Theorem 5. Let \(\{t_k\}_{k=0}^{N_h}\) be a uniform partition of \([0,1]\) with step size \(h=1/N_h>0\). Under [ass_drift_approx_wass,ass_moment,ass_score,ass_strong_regularity,ass_hessian], denoting by \(\nu^{\theta^\star}_1\) the law of \(X^{\theta^{\star}}_1\), we have that

\[\begin{align} \label{wass95strong95convergence95bound} \mathscr{W}_2(\nu^{\star},\nu^{\theta^\star}_1) \lesssim \exp\left( \frac{8\sqrt{2} \|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}}{\sqrt{\alpha_{\pi}}}\right) \left(\varepsilon +\sqrt{h}(h^{1/16}+1) \sqrt{\left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right)d}\right)\;. \end{align}\tag{62}\]

Proof of 5:. Consider the synchronous coupling between \((X^\mathrm{M}_t)_{t\in [0,1]}\) and the continuous time interpolation of \((X^{\theta^{\star}}_t)_{t\in [0,1]}\) with the same initialization, i.e., use the same Brownian motion to drive the two processes and set \(X^{\theta^{\star}}_0=X^\mathrm{M}_0\). Then, it holds \[\begin{align} \label{eq:wass95bounded95by95l2} \mathscr{W}_2(\nu^{\star}, \nu^{\theta^{\star}})\le \left\|X^\mathrm{M}_{T}-X^{\theta^{\star}}_{T}\right\|_{\mathrm{L}^2}=\left\|X^\mathrm{M}_{t_N}-X^{\theta^{\star}}_{t_N}\right\|_{\mathrm{L}^2}\;, \end{align}\tag{63}\] where, with abuse of notation, we denoted by \((X^{\theta^{\star}}_t)_{t\in [0,1]}\) either the process 8 and its time continuous interpolation. To upper bound the r.h.s. of the above expression, we estimate \(\left\|X^\mathrm{M}_{t_{k+1}}-X^{\theta^{\star}}_{t_{k+1}}\right\|_{\mathrm{L}^2}\) by means of \(\left\|X^\mathrm{M}_{t_{k}}-X^{\theta^{\star}}_{t_{k}}\right\|_{\mathrm{L}^2}\) and develop the recursion. Fix \(k\in \{0,...,N-1\}\). As we considered the synchronous coupling, we have that \[\begin{align} \label{eq:big95one95theo95wass} &\left\|X^\mathrm{M}_{t_{k+1}}-X^{\theta^{\star}}_{t_{k+1}}\right\|_{\mathrm{L}^2}\\ &= \left\|X^\mathrm{M}_{t_{k}}-X^{\theta^{\star}}_{t_{k}}+\int_{t_k}^{t_{k+1}} \left\{\tilde{\beta}_t(X^\mathrm{M}_t)-s_{\theta^\star}(t_k, X^{\theta^{\star}}_{t_k})\right\}\mathrm{d}t\right\|_{\mathrm{L}^2}\\ &\le \left\|X^\mathrm{M}_{t_{k}}-X^{\theta^{\star}}_{t_{k}}\right\|_{\mathrm{L}^2}+\sqrt{h}\left(\int_{t_k}^{t_{k+1}}\mathbb{E}\left[\left\| \left\{\tilde{\beta}_t(X^\mathrm{M}_t)-\tilde{\beta}_{t_k}(X^\mathrm{M}_{t_k})\right\}\right\|^2\right]\mathrm{d}t\right)^{1/2}+h\left\|\tilde{\beta}_{t_k}(X^\mathrm{M}_{t_k})-\tilde{\beta}_{t_k}(X^{\theta^{\star}}_{t_k})\right\|_{\mathrm{L}^2}\\ &\quad + h\left\|\tilde{\beta}_{t_k}(X^{\theta^{\star}}_{t_k})-s_{\theta^\star}(t_k, X^{\theta^{\star}}_{t_k})\right\|_{\mathrm{L}^2}\\ &\le \left\|X^\mathrm{M}_{t_{k}}-X^{\theta^{\star}}_{t_{k}}\right\|_{\mathrm{L}^2}+h\left\|\tilde{\beta}_{t_k}(X^\mathrm{M}_{t_k})-\tilde{\beta}_{t_k}(X^{\theta^{\star}}_{t_k})\right\|_{\mathrm{L}^2}+h\varepsilon\\ &\quad +h\left(\sum_{k=0}^{N-1}\int_{t_k}^{t_{k+1}}\mathbb{E}\left[\left\| \left\{\tilde{\beta}_t(X^\mathrm{M}_t)-\tilde{\beta}_{t_k}(X^\mathrm{M}_{t_k})\right\}\right\|^2\right]\mathrm{d}t\right)^{1/2}\;, \end{align}\tag{64}\] where, in the last inequality, we have used 1. We now focus on the second term appearing in the r.h.s. of the above expression. First note that if \(k=0\) then this term is null. Therefore, we can assume \(0<k<N\), so that \(0<t_k<1\), use 9 and apply 3. By doing so, we get \[\begin{align} h\left\|\tilde{\beta}_{t_k}(X^\mathrm{M}_{t_k})-\tilde{\beta}_{t_k}(X^{\theta^{\star}}_{t_k})\right\|_{\mathrm{L}^2}&\le h \frac{4\sqrt{2}}{\sqrt{\alpha_{\pi}}} \|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}\frac{1}{\sqrt{t_k(1-t_k)}}\left\|X^\mathrm{M}_{t_{k}}-X^{\theta^{\star}}_{t_{k}}\right\|_{\mathrm{L}^2}\;. \end{align}\]Plugging this into 64 , we get

\[\begin{align} \label{eq:big95second95theo95wass} \left\|X^\mathrm{M}_{t_{k+1}}-X^{\theta^{\star}}_{t_{k+1}}\right\|_{\mathrm{L}^2}&\le \left\|X^\mathrm{M}_{t_{k}}-X^{\theta^{\star}}_{t_{k}}\right\|_{\mathrm{L}^2}+h \frac{4\sqrt{2}}{\sqrt{\alpha_{\pi}}} \|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}\frac{1}{\sqrt{t_k(1-t_k)}}\left\|X^\mathrm{M}_{t_{k}}-X^{\theta^{\star}}_{t_{k}}\right\|_{\mathrm{L}^2}\\ &\quad +h\varepsilon+h\left(\sum_{k=0}^{N-1}\int_{t_k}^{t_{k+1}}\mathbb{E}\left[\left\| \left\{\tilde{\beta}_t(X^\mathrm{M}_t)-\tilde{\beta}_{t_k}(X^\mathrm{M}_{t_k})\right\}\right\|^2\right]\mathrm{d}t\right)^{1/2}\;. \end{align}\tag{65}\]

We are left with bounding the last term appearing in the r.h.s.. To this aim, we proceed as in the proof of 1 and obtain \[\begin{align} &h\left(\sum_{k=0}^{N-1}\int_{t_k}^{t_{k+1}}\mathbb{E}\left[\left\| \left\{\tilde{\beta}_t(X^\mathrm{M}_t)-\tilde{\beta}_{t_k}(X^\mathrm{M}_{t_k})\right\}\right\|^2\right]\mathrm{d}t\right)^{1/2}\\ &\le \mathbf{c} h\sqrt{ h(h^{1/8}+1) \left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right)d}\\ &\le \mathbf{c} h\sqrt{h}(h^{1/16}+1)\sqrt{\left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right)d}\;, \end{align}\]for some universal constant \(\mathbf{c}>0\), which may change from line to line. Plugging this bound in 65 , we get \[\begin{align} \left\|X^\mathrm{M}_{t_{k+1}}-X^{\theta^{\star}}_{t_{k+1}}\right\|_{\mathrm{L}^2}\le \gamma_k\left\|X^\mathrm{M}_{t_{k}}-X^{\theta^{\star}}_{t_{k}}\right\|_{\mathrm{L}^2}+ h\sqrt{h}\mathrm{D}+h\varepsilon\;,\quad k\in\{0,...,N-2\}\;, \end{align}\]with \[\begin{align} \label{def:l2} \gamma_k=1+h \frac{4\sqrt{2}}{\sqrt{\alpha_{\pi}}} \|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}\frac{1}{\sqrt{t_k(1-t_k)}}\;, \end{align}\tag{66}\] and \[\begin{align} \label{def:strong95constant95D} \mathrm{D}=\mathbf{c}(h^{1/16}+1)\sqrt{\left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right)d}\;. \end{align}\tag{67}\] Therefore, if we develop the recursion and use the fact that we set \(X^\mathrm{M}_{0}=X^{\theta^{\star}}_{0}\), we get that \[\begin{align} \label{eq:development95of95recursion} \left\|X^\mathrm{M}_{T}-X^{\theta^{\star}}_{T}\right\|_{\mathrm{L}^2}&\le \left\|X^\mathrm{M}_{0}-X^{\theta^{\star}}_{0}\right\|_{\mathrm{L}^2}\prod_{l=0}^{N-2}\gamma_l+ (h\sqrt{h}\mathrm{D}+h\varepsilon)\sum_{k=0}^{N-1}\prod_{l=k}^{N-2}\gamma_l\\ &=( \sqrt{h}\mathrm{D}+\varepsilon)\left(h\;\sum_{k=0}^{N-1}\prod_{l=k}^{N-1}\gamma_l\right)\;. \end{align}\tag{68}\] Note that \[\begin{align} h\; \sum_{k=0}^{N-1}\prod_{l=k}^{N-1}\gamma_l&= h\;\sum_{k=0}^{\lfloor N/2\rfloor}\prod_{l=k}^{\lfloor N/2\rfloor}\gamma_l+h\;\sum_{k=\lfloor N/2\rfloor}^{N-1}\prod_{l=k}^{N-1}\gamma_l\\ &\le h\;\sum_{k=0}^{\lfloor N/2\rfloor}\prod_{l=k}^{\lfloor N/2\rfloor}\left(1+\sqrt{h} \frac{4\sqrt{2}\;\|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}}{\sqrt{\alpha_{\pi}}}\frac{1}{\sqrt{l}}\right)+h\;\sum_{k=\lfloor N/2\rfloor}^{N-1}\prod_{l=k}^{N-1}\left(1+\sqrt{h} \frac{4\sqrt{2}\;\|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}}{\sqrt{\alpha_{\pi}}}\frac{1}{\sqrt{l}}\right)\\ &= h\;\sum_{k=0}^{N-1}\prod_{l=k}^{N-1}\left(1+\sqrt{h} \frac{4\sqrt{2}\;\|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}}{\sqrt{\alpha_{\pi}}}\frac{1}{\sqrt{l}}\right)\le h\;\sum_{k=0}^{N-1}\prod_{l=k}^{N-1} \exp\left(\sqrt{h} \frac{4\sqrt{2}\;\|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}}{\sqrt{\alpha_{\pi}}}\frac{1}{\sqrt{l}}\right)\\ &=h\;\sum_{k=0}^{N-1}\exp\left(\sqrt{h} \frac{4\sqrt{2}\;\|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}}{\sqrt{\alpha_{\pi}}}\sum_{l=k}^{N-1}\frac{1}{\sqrt{l}}\right)\le h\;\sum_{k=0}^{N-1}\exp\left(\sqrt{hN} \frac{8\sqrt{2}\;\|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}}{\sqrt{\alpha_{\pi}}}\right)\\ &\le Nh\exp\left( \frac{8\sqrt{2}\;\|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}}{\sqrt{\alpha_{\pi}}}\right)= \exp\left( \frac{8\sqrt{2}\;\|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}}{\sqrt{\alpha_{\pi}}}\right)\;. \end{align}\] Therefore, we have that \[\begin{align} \left\|X^\mathrm{M}_{T}-X^{\theta^{\star}}_{T}\right\|_{\mathrm{L}^2}\le \exp\left( \frac{8\sqrt{2}}{\sqrt{\alpha_{\pi}}}\|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}\right)( \sqrt{h}\mathrm{D}+\varepsilon)\;. \end{align}\] Recalling the definition 67 of \(\mathrm{D}\), and plugging this bound in 63 , we obtain 62 . ◻

8.2 Weakly Log-Concave and One-Sided Log-Lipschitz Distributions↩︎

Proof of 4:. We proceed as in the proof of 5, to obtain 64 , i.e., \[\begin{align} \left\|X^\mathrm{M}_{t_{k+1}}-X^{\theta^{\star}}_{t_{k+1}}\right\|_{\mathrm{L}^2}&\le \left\|X^\mathrm{M}_{t_{k}}-X^{\theta^{\star}}_{t_{k}}\right\|_{\mathrm{L}^2}+h\left\|\tilde{\beta}_{t_k}(X^\mathrm{M}_{t_k})-\tilde{\beta}_{t_k}(X^{\theta^{\star}}_{t_k})\right\|_{\mathrm{L}^2}+h\varepsilon\\ &\quad +h\left(\sum_{k=0}^{N-1}\int_{t_k}^{t_{k+1}}\mathbb{E}\left[\left\| \left\{\tilde{\beta}_t(X^\mathrm{M}_t)-\tilde{\beta}_{t_k}(X^\mathrm{M}_{t_k})\right\}\right\|^2\right]\mathrm{d}t\right)^{1/2}\;. \end{align}\]To bound the second term of the r.h.s. of the above expression, we use 4 instead of 3. By doing so, we get \[\begin{align} h\left\|\tilde{\beta}_{t_k}(X^\mathrm{M}_{t_k})-\tilde{\beta}_{t_k}(X^{\theta^{\star}}_{t_k})\right\|_{\mathrm{L}^2}\le (\gamma_k-1) \left\|X^\mathrm{M}_{t_{k}}-X^{\theta^{\star}}_{t_{k}}\right\|_{\mathrm{L}^2}\;, \end{align}\]with \[\begin{align} \label{eq:gamma95weak} \gamma_k=1+ h\frac{4\sqrt{2}}{\sqrt{\alpha_{\pi}}}\exp\left(\frac{M_\pi}{\alpha_\pi}\right) \|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}\frac{1}{\sqrt{t_k(1-t_k)}}\;. \end{align}\tag{69}\] Therefore, we have that \[\begin{align} \left\|X^\mathrm{M}_{t_{k+1}}-X^{\theta^{\star}}_{t_{k+1}}\right\|_{\mathrm{L}^2}&\le \gamma_k\left\|X^\mathrm{M}_{t_{k}}-X^{\theta^{\star}}_{t_{k}}\right\|_{\mathrm{L}^2}+h\varepsilon +h\left(\sum_{k=0}^{N-1}\int_{t_k}^{t_{k+1}}\mathbb{E}\left[\left\| \left\{\tilde{\beta}_t(X^\mathrm{M}_t)-\tilde{\beta}_{t_k}(X^\mathrm{M}_{t_k})\right\}\right\|^2\right]\mathrm{d}t\right)^{1/2}\;. \end{align}\] To bound the last term of the r.h.s. of the above equation, as before, we first proceed as in the proof of 1 and obtain

\[\begin{align} h\left(\sum_{k=0}^{N-1}\int_{t_k}^{t_{k+1}}\mathbb{E}\left[\left\| \left\{\tilde{\beta}_t(X^\mathrm{M}_t)-\tilde{\beta}_{t_k}(X^\mathrm{M}_{t_k})\right\}\right\|^2\right]\mathrm{d}t\right)^{1/2}\le \mathbf{c}h\sqrt{h}(h^{1/16}+1)\sqrt{\left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right)d}\;, \end{align}\]for some universal constant \(\mathbf{c}>0.\)

Therefore, this time we get a recursive formula of the type \[\begin{align} \left\|X^\mathrm{M}_{t_{k+1}}-X^{\theta^{\star}}_{t_{k+1}}\right\|_{\mathrm{L}^2}\le \gamma_k\left\|X^\mathrm{M}_{t_{k}}-X^{\theta^{\star}}_{t_{k}}\right\|_{\mathrm{L}^2}+ h\sqrt{h}\mathrm{D}+h\varepsilon\;,\quad k\in\{0,...,N-2\}\;, \end{align}\]with \(\gamma_k=(\gamma_k-1)+1\) and \((\gamma_k-1)\) as in 69 and \[\begin{align} \label{def:constant95D} \mathrm{D}=\mathbf{c}(h^{1/16}+1)\sqrt{\left(d^2+\norm{\nabla\log \pi}_{\mathrm{L}^8(\pi)}^4\right)d}\;. \end{align}\tag{70}\]

Developing the recursion as before and using 63 , we obtain that \[\begin{align} \label{eq:weak95development95of95recursion} \mathscr{W}_2(\nu^{\star}, \nu^{\theta^{\star}})&\le \left\|X^\mathrm{M}_{T}-X^{\theta^{\star}}_{T}\right\|_{\mathrm{L}^2}\le \left\|X^\mathrm{M}_{0}-X^{\theta^{\star}}_{0}\right\|_{\mathrm{L}^2}\prod_{l=0}^{N-2}\gamma_l+ (h\sqrt{h}\mathrm{D}+h\varepsilon)\sum_{k=0}^{N-1}\prod_{l=k}^{N-2}\gamma_l\\ &=( \sqrt{h}\mathrm{D}+\varepsilon)\left(h\;\sum_{k=0}^{N-1}\prod_{l=k}^{N-1}\gamma_l\right)\;. \end{align}\tag{71}\]

Proceeding as before, we get \[\begin{align} h\;\sum_{k=0}^{N-1}\prod_{l=k}^{N-1}\gamma_l\le \exp\left( \frac{8\sqrt{2}}{\sqrt{\alpha_{\pi}}}\exp\left(\frac{M_\pi}{\alpha_\pi}\right)\|\nabla^2 \log \pi\|_{\mathrm{L}^2(\pi)}\right) \end{align}\] Plugging this bound in 71 , and recalling the definition 70 of \(\mathrm{D}\) respectively, we obtain 13 . ◻

References↩︎

[1]
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models.” 2020, [Online]. Available: https://arxiv.org/abs/2006.11239.
[2]
A. Q. Nichol and P. Dhariwal, “Improved denoising diffusion probabilistic models,” in Proceedings of the 38th international conference on machine learning, 2021, pp. 8162–8171.
[3]
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, arXiv:2010.11119“DiffWave: A versatile diffusion model for audio synthesis,” 2020.
[4]
C. Liu and et al., arXiv:2302.02591“Generative diffusion models on graphs: Methods and applications,” 2023.
[5]
C. Zang and F. Wang, arXiv:2006.10137“MoFlow: An invertible flow model for generating molecular graphs,” 2020.
[6]
I. Dunn and D. R. Koes, arXiv:2508.12629“FlowMol3: Flow matching for 3D de novo small-molecule generation,” 2025.
[7]
J. Luo, L. Yang, Y. Liu, et al., “Review of diffusion models and its applications in biomedical informatics,” BMC Medical Informatics and Decision Making, 2025, doi: 10.1186/s12911-025-03210-5.
[8]
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution.” 2020, [Online]. Available: https://arxiv.org/abs/1907.05600.
[9]
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” Advances in neural information processing systems, vol. 32, 2019.
[10]
S. Chen, S. Chewi, H. Lee, Y. Li, J. Lu, and A. Salim, “The probability flow ODE is provably fast,” in Thirty-seventh conference on neural information processing systems, 2023.
[11]
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10684–10695.
[12]
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen, “Hierarchical text-conditional image generation with clip latents,” arXiv preprint arXiv:2204.06125, vol. 1, no. 2, p. 3, 2022.
[13]
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov, “Grad-tts: A diffusion probabilistic model for text-to-speech,” in International conference on machine learning, 2021, pp. 8599–8608.
[14]
S. Peluchetti, “Non-denoising forward-time diffusions.” 2022.
[15]
Y. Lipman and et al., online resource“Flow matching guide and code.” 2024.
[16]
M. S. Albergo and E. Vanden-Eijnden, “Building normalizing flows with stochastic interpolants,” arXiv preprint arXiv:2209.15571, 2022.
[17]
M. S. Albergo, N. M. Boffi, and E. Vanden-Eijnden, “Stochastic interpolants: A unifying framework for flows and diffusions,” arXiv preprint arXiv:2303.08797, 2023.
[18]
Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le, “Flow matching for generative modeling,” in The eleventh international conference on learning representations, 2023.
[19]
Q. Liu, “Rectified flow: A marginal preserving approach to optimal transport,” arXiv preprint arXiv:2209.14577, 2022.
[20]
X. Liu, C. Gong, and Q. Liu, “Flow straight and fast: Learning to generate and transfer data with rectified flow,” International Conference on Learning Representations (ICLR), 2023.
[21]
A. Klenke, Probability theory: A comprehensive course. Springer Science & Business Media, 2013.
[22]
M. Gentiloni Silveri, A. Durmus, and G. Conforti, “Theoretical guarantees in kl for diffusion flow matching,” Advances in Neural Information Processing Systems, vol. 37, pp. 138432–138473, 2024.
[23]
I. Gyöngy, “Mimicking the one-dimensional marginal distributions of processes having an itô differential,” Probability theory and related fields, vol. 71, no. 4, pp. 501–516, 1986.
[24]
N. V. Krylov, “On the relation between differential operators of second order and the solutions of stochastic differential equations,” Steklov Seminar.
[25]
M. Gentiloni Silveri, G. Conforti, and A. Oliviero Durmus, “Exponential convergence guarantees for iterative markovian fitting,” in The thirty-ninth annual conference on neural information processing systems, 2025.
[26]
S. Chen, S. Chewi, J. Li, Y. Li, A. Salim, and A. R. Zhang, “Sampling is as easy as learning the score: Theory for diffusion models with minimal data assumptions,” International Conference on Learning Representations, 2023.
[27]
H. Chen, H. Lee, and J. Lu, “Improved analysis of score-based generative modeling: User-friendly bounds under minimal smoothness assumptions,” in International conference on machine learning, 2023, pp. 4735–4763.
[28]
X. Gao, H. M. Nguyen, and L. Zhu, “Wasserstein convergence guarantees for a general class of score-based generative models,” arXiv preprint arXiv:2311.11003, 2023.
[29]
G. Conforti, A. Durmus, and M. Gentiloni Silveri, “KL convergence guarantees for score diffusion models under minimal data assumptions,” SIAM Journal on Mathematics of Data Science, vol. 7, no. 1, pp. 86–109, 2025.
[30]
M. Gentiloni Silveri and A. Ocello, “Beyond log-concavity and score regularity: Improved convergence bounds for score-based generative models in W2-distance,” in Forty-second international conference on machine learning, 2025.
[31]
S. Strasman, A. Ocello, C. Boyer, S. L. Corff, and V. Lemaire, “An analysis of the noise schedule for score-based generative models,” arXiv preprint arXiv:2402.04650, 2024.
[32]
S. Strasman, S. Surendran, C. Boyer, S. Le Corff, V. Lemaire, and A. Ocello, “Wasserstein convergence of critically damped langevin diffusions,” in The thirty-ninth annual conference on neural information processing systems.
[33]
S. Bruno, Y. Zhang, D.-Y. Lim, Ö. D. Akyildiz, and S. Sabanis, “On diffusion-based generative models and their error bounds: The log-concave case with full convergence estimates,” arXiv preprint arXiv:2311.13584, 2023.
[34]
G. Conforti, “Weak semiconvexity estimates for schr\(\backslash\)" odinger potentials and logarithmic sobolev inequality for schr\(\backslash\)" odinger bridges,” arXiv preprint arXiv:2301.00083, 2022.
[35]
S. Bruno and S. Sabanis, “Wasserstein convergence of score-based generative models under semiconvexity and discontinuous gradients,” arXiv preprint arXiv:2505.03432, 2025.
[36]
G. Kremling, F. Iafrate, M. Taheri, and J. Lederer, “Non-asymptotic error bounds for probability flow ODEs under weak log-concavity,” arXiv preprint arXiv:2510.17608, 2025.
[37]
V. Arsenyan, E. Vardanyan, and A. Dalalyan, “Assessing the quality of denoising diffusion models in wasserstein distance: Noisy score and optimal bounds,” arXiv preprint arXiv:2506.09681, 2025.
[38]
Y. Liu, Y. Chen, R. Hu, and L. Huang, “Finite-time analysis of discrete-time stochastic interpolants,” arXiv preprint arXiv:2502.09130, 2025.
[39]
N. M. Boffi, M. S. Albergo, and E. Vanden-Eijnden, “Flow map matching with stochastic interpolants: A mathematical framework for consistency models,” Transactions on Machine Learning Research, 2025.
[40]
M. Xiangjun and n. W. Zhongjia, “Pathway to $o(\sqrt{d})$ complexity bound under wasserstein metric of flow-based models,” 2025.
[41]
S. Chewi and A.-A. Pooladian, “An entropic generalization of caffarelli’s contraction theorem via covariance inequalities,” Comptes Rendus. Mathématique, vol. 361, no. G9, pp. 1471–1482, 2023.
[42]
G. Conforti, D. Lacker, and S. Pal, “Projected langevin dynamics and a gradient flow for entropic optimal transport,” arXiv preprint arXiv:2309.08598, 2023.
[43]
C. Villani et al., Optimal transport: Old and new, vol. 338. Springer, 2008.
[44]
M. Nutz, “Introduction to entropic optimal transport,” Lecture notes, Columbia University, 2021.
[45]
T. Van Erven and P. Harremos, “Rényi divergence and kullback-leibler divergence,” IEEE Transactions on Information Theory, vol. 60, no. 7, pp. 3797–3820, 2014.