Bridging Schrödinger and Bass: A Semimartingale Optimal Transport Problem with Diffusion Control


Abstract

We study a semimartingale optimal transport problem interpolating between the Schrödinger bridge and the stretched Brownian motion associated with the Bass solution of the Skorokhod embedding problem. The cost combines an entropy term on the drift with a quadratic penalization of the diffusion coefficient, leading to a stochastic control problem over drift and volatility.

We establish a complete duality theory for this problem, despite the lack of coercivity in the diffusion component. In particular, we prove strong duality and dual attainment, and derive an equivalent reduced dual formulation in terms of a variational problem over terminal potentials.

Optimal solutions are characterized by a coupled Schrödinger–Bass bridge system, involving a backward heat potential and a transport map given by the gradient of a convex function. This system interpolates between the classical Schrödinger system and the Bass martingale transport. Our results provide a unified framework encompassing entropic and martingale optimal transport, and yield a variational foundation for data-driven diffusion models.

: Optimal transport along Itô processes, Schrödinger bridge, Bass martingale, dual potentials, Schrödinger–Bass bridge system.

: Primary: 49Q22, 60H30. Secondary: 93E20, 60J60, 60G44.

1 Introduction↩︎

Classical Schrödinger bridges and martingale optimal transport provide two fundamental constructions of optimal couplings of probability measures via continuous-time stochastic processes. In the Schrödinger problem, one fixes a reference Brownian motion and prescribes initial and terminal distributions, then minimises the relative entropy of the law of the process with respect to the reference path measure. This leads to a stochastic control problem with controlled drift and fixed diffusion coefficient, admitting a well-posed dual formulation and a characterisation in terms of the Schrödinger system, see [1], [2], [3]. In contrast, martingale optimal transport imposes a martingale constraint and allows for non-trivial control of the diffusion coefficient. The Bass problem, whose optimiser is the stretched Brownian motion, consists in prescribing marginal laws of a continuous martingale while minimising a deviation from Brownian motion; see [4][6]. This formulation is closely related to martingale Benamou–Brenier transport [7] and plays a central role in model-independent finance. In contrast with the Schrödinger problem, it exhibits degeneracies and lacks a regular dual structure.

The aim of this paper is to introduce and analyse a class of optimal semimartingale transport problems that interpolate between these two regimes. Given \(\mu_0,\mu_T\in{\cal P}_2(\mathbb{R}^d)\), we consider Itô processes \(X=(X_t)_{0\le t\le T}\) with \(X_0\sim\mu_0\), \(X_T\sim\mu_T\), and we minimise a functional which penalises both the drift and the volatility of \(X\), weighted by a parameter \(\beta\) \(>\) \(0\). We denote the resulting value by \({\rm SBB}(\mu_0,\mu_T)\) and refer to this optimisation problem as the Schrödinger–Bass bridge (SBB) problem. As \(\beta\to\infty\), one recovers the Schrödinger bridge, while as \(\beta\to 0\), the problem reduces to a Bass-type martingale transport. The SBB problem therefore provides a unified stochastic control formulation of entropic and martingale transport.

The main difficulty in the analysis stems from the lack of coercivity in the diffusion variable. In contrast with the framework of controlled diffusions studied in [8], [9], the cost functional does not penalise large diffusion coefficients in a coercive manner, and standard compactness and duality arguments do not apply directly.

Our first main result establishes strong duality for the SBB problem. We prove that the primal stochastic control problem admits a minimiser and that its value coincides with that of a dual problem over functions \(v\in C^{1,2}\) solving a fully nonlinear Hamilton–Jacobi–Bellman equation under a curvature constraint \(D^2 v<\beta I_d\). Moreover, the HJB equation admits an explicit representation in terms of a terminal potential \(\phi\) through a quadratic inf–convolution operator, leading to an equivalent static dual formulation of Donsker–Varadhan type. We show that the static dual problem attains its supremum over a suitable relaxed space.

Our second main result is a complete characterisation of the solution to SBB. We show that the optimal SBB bridge is encoded by a triplet \((h,\nu,\mathscr{Y})\) solving a coupled system, which we call the Schrödinger–Bass bridge system. The functions \(h\) and \(\nu\) solve backward and forward heat equations, while the map \(\mathscr{Y}\) is the gradient of a quadratic inf–convolution transform. The marginal laws \((\mu_t)_{t}\) of the optimal process \((X_t)_t\) satisfy \[\begin{align} \mathscr{Y}_t \#\mu_t & = \; h_t \nu_t, \quadi.e.\quad \mu_t \; = \: \mathscr{X}_t\#(h_t \nu_t), \quad 0 \leq t \leq T, \end{align}\] where \(\#\) denotes the push-forward operation, \(\mathscr{X}_t\) \(=\) \(\mathscr{Y}_t^{-1}\) is the inverse of the homeomorphism \(\mathscr{Y}_t\), and is given by the gradient of a convex function: \[\begin{align} \mathscr{X}_t(y) &= \; \nabla_y \Big( \frac{|y|^2}{2} + \frac{1}{\beta} \log h_t(y) \Big) \; = \; y + \frac{1}{\beta} \nabla_y \log h_t(y), \end{align}\] thus \(\mathscr{X}_t\) is expressed in terms of the score of the potential density \(h_t\). This system reduces to the classical Schrödinger system as \(\beta\to\infty\) and to the Bass martingale transport system as \(\beta\to 0\).

Finally, we identify the structure of the optimal semimartingale. The optimal drift and diffusion coefficients admit explicit feedback representations in terms of \((h,\mathscr{Y})\). Moreover, after the change of variables \(Y_t\) \(=\) \(\mathscr{Y}_t(X_t)\), \(0\leq t\leq T\), the process \(Y\) is a Schrödinger bridge diffusion and becomes a Brownian motion under an equivalent change of measure. This yields a stretched Brownian representation of the optimal SBB bridge, extending the Bass construction from martingales to general semimartingales.

The paper is organised as follows. Section 2 introduces the SBB problem and derives a first dual formulation. Section 3 states the main results. Sections 4 and 5 establish duality and analyse the reduced dual problem. Sections 6 and 7 prove dual and primal attainment and derive the structure of optimal solutions. 6

2 Schrödinger–Bass bridge optimal transport↩︎

2.1 Problem formulation↩︎

Let \(T>0\) be some finite maturity and let \(\Omega=C([0,T],\mathbb{R}^d)\) be the canonical space equipped with its canonical filtration \(\mathbb{F}\) \(=\) \(({\cal F}_t)_t\) and canonical process \(X\), i.e. \(X_t(\omega)=\omega(t)\) for all \(\omega\in\Omega\), \(t\in [0,T]\). We denote by \({\cal P}\) the set of probability measures \(\mathbb{P}\) on \(\Omega\) under which \(X\) has the diffusion decomposition \[\begin{align} \label{decdiff} X_t &= \; X_0 + \int_0^t \alpha_s^\mathbb{P}\mathrm{d}s + \int_0^t \sigma_s^\mathbb{P}\mathrm{d}W_s^\mathbb{P}, \qquad t \in [0,T], \; \mathbb{P}-a.s. \end{align}\tag{1}\] for some \(d\)-dimensional \(\mathbb{P}-\)Brownian motion \(W^\mathbb{P}\), and some characteristics \(\gamma^\mathbb{P}=(\alpha^\mathbb{P},\sigma^\mathbb{P})\) that are \(\mathbb{F}\)-progressively measurable processes valued in \(\mathbb{R}^d\times\mathbb{S}_+^d\), satisfying \(\int_0^T |\alpha_t^\mathbb{P}|\mathrm{d}t+\int_0^T |\sigma_t^\mathbb{P}|^2 \mathrm{d}t<\infty\), \(\mathbb{P}\)-a.s.

Given two distributions \(\mu_0,\mu_T\in{\cal P}_2(\mathbb{R}^d)\), we introduce the subset of transport plans \[{\cal P}(\mu_0,\mu_T) \;=\; \big\{ \mathbb{P}\in {\cal P}: \mathbb{P}\circ X_0^{-1} = \mu_0, \;and\; \mathbb{P}\circ X_T^{-1} = \mu_T \big\}.\] Given a parameter \(\beta>0\), and denoting by \(I_d\) the identity matrix in \(\mathbb{R}^{d\times d}\), we introduce the cost function \[\label{def:c} c(a,b):=\frac{1}{2}|a|^2+\frac{\beta}{2}|b - I_d |^2, ~a\in\mathbb{R}^d,~b\in\mathbb{S}_+^d.\tag{2}\] Our objective in this paper is to derive a full characterization of a critical transport plan obtained through the following optimal transport problem on the canonical path space: \[\label{SBB} {\rm SBB}(\mu_0,\mu_T) := \inf_{\mathbb{P}\in{\cal P}(\mu_0,\mu_T)} J^0(\mathbb{P}), ~with~ J^0(\mathbb{P}) := \mathbb{E}^{\mathbb{P}} \Big[ \int_0^Tc(\gamma_t^\mathbb{P})\mathrm{d}t \Big].\tag{3}\] Here the SBB acronym stands for Schrödinger–Bass bridge, and is justified by the two extreme cases \(\beta\to 0\) and \(\beta\to\infty\):

  • Formally, when \(\beta\) goes to infinity, we constrain to those transport plans in the subset \({\cal P}_0(\mu_0,\mu_T):=\{\mathbb{P}\!\in\!{\cal P}(\mu_0,\mu_T):\sigma^\mathbb{P}=I_d~on~[0,T]\}\); the problem SBB reduces to the classical Schrödinger bridge problem which aims to find the closest transport plan \(\mathbb{P}\in{\cal P}_0(\mu_0,\mu_T)\) to the Brownian path measure with initial law \(\mu_0\) in the sense of the relative entropy (Kullback-Leibler) distance, see, e.g. [1], \[{\rm SB}(\mu_0,\mu_T) := \inf_{\mathbb{P}\in{\cal P}_0(\mu_0,\mu_T)} \mathbb{E}^{\mathbb{P}} \Big[ \frac{1}{2}\int_0^T|\alpha_t^\mathbb{P}|^2\mathrm{d}t \Big].\]

  • In the other extreme case, by dividing the criterion \(J^0\) by \(\beta\), and sending \(\beta\) to zero, we formally constrain the drift coefficient to be zero, and then we are looking for a continuous martingale, namely the stretched Brownian motion/Bass martingale in the appropriate setting, which is closest to Brownian motion according to the quadratic volatility norm, under marginal constraints. This martingale transport problem was studied in Backhoff-Veraguas, Beiglböck, Huesmann & Källblad [7], Backhoff-Veraguas, Schachermayer & Tschilderer [5], who called its solution the Stretched Brownian motion and which is intimately related to the Bass martingale with degenerate initial law. The latter was motivated by calibration problems in financial engineering, and was first studied by Conze & Henry-Labordère [4] and Acciaio, Pammer & Marini [6].

The solution (when it exists) \(\hat{\mathbb{P}}\) of \({\rm SBB}(\mu_0,\mu_T)\) is called the Schrödinger–Bass bridge (SBB) transport plan. We start with some initial considerations which will underpin the subsequent analysis of the problem SBB.

Remark 1. The linear coupling \(\mathbb{P}^\pi_{\rm lin}\in{\cal P}(\mu_0,\mu_T)\) is defined for all static coupling \(\pi\in\Pi(\mu_0,\mu_T)\) as follows. Let \((Y_0,Y_T)\) be a random vector on some probability space, with law \(\pi\), and let \(\mathbb{P}^\pi_{\mathrm{lin}} := \text{Law}(Y)\) be the law of the linear interpolation \(Y_t :=\frac{T-t}{T}Y_0+\frac{t}{T}Y_T,\) \(t\in[0,T]\). Back to our notations on the canonical space, we observe that \(\mathrm{d}X_t=\alpha^{\mathbb{P}^\pi_{\mathrm{lin}}}_t\mathrm{d}t+0\;\mathrm{d}W_t,\) \(\mathbb{P}^\pi_{\mathrm{lin}}-\)a.s. with corresponding \(\alpha^{\mathbb{P}^\pi_{\mathrm{lin}}}_t:=\frac{X_T-X_0}{T},\; 0\le t\le T\), and, for \(t>0\), equivalently \(\alpha^{\mathbb{P}^\pi_{\mathrm{lin}}}_t=\frac{X_t-X_0}{t}\) satisfying \(J^0(\mathbb{P}^\pi_{\mathrm{lin}})=\frac{1}{2}\mathbb{E}\int_0^T|\alpha^{\mathbb{P}^\pi_{\mathrm{lin}}}_t|^2dt + \frac{\beta d T}{2}=\frac{1}{2}T^{-1}\mathbb{E}|X_T-X_0|^2+ \frac{\beta d T}{2}<\infty\). Consequently \(\mathbb{P}^\pi_{\mathrm{lin}}\in{\cal P}(\mu_0,\mu_T)\).

Lemma 1. \({\rm SBB}(\mu_0,\mu_T)\in[0,\infty)\), and \(\mathbb{E}^\mathbb{P}[\sup_{t\le T}|X_t|^2]<\infty\) for every \(\mathbb{P}\in{\cal P}(\mu_0,\mu_T)\) with \(J^0(\mathbb{P})<\infty\).

By Remark 1, we have \({\rm SBB}(\mu_0,\mu_T)\le J^0(\mathbb{P}_{\mathrm{lin}})<\infty\). Next, for \(\mathbb{P}\in{\cal P}(\mu_0,\mu_T)\) with \(J^0(\mathbb{P})<\infty\), it follows from the BDG inequality that \(\mathbb{E}^\mathbb{P}[\sup_{t\le T}|X_t|^2]\le C(1+\mathbb{E}^\mathbb{P}[|X_0|^2]+\mathbb{E}^\mathbb{P}\!\int_0^T\big(|\alpha_t^\mathbb{P}|^2+|\sigma_t^\mathbb{P}|^2\big)\mathrm{d}t)\le C'(1+\mathbb{E}^\mathbb{P}[|X_0|^2]+J^0(\mathbb{P}))<\infty\), for some constants \(C,C'>0\).   \({\cal t}\)  \({\cal u}\)

2.2 A first dual problem↩︎

The problem SBB is a semimartingale optimal transport problem in the sense of [8], see also [9]. Notice however that our cost function does not satisfy the coercivity condition in \(b\) that is required in [8]. Despite this, we shall justify that the linear programming duality still holds in our setting and we may then express the constrained control problem SBB in terms of the penalized unconstrained control problem.

Let \(w:=1+|.|^2\) and let \({\cal C}_w:=\{\psi\in C^0(\mathbb{R}^d,\mathbb{R}):|\frac{\psi}{w}|_\infty<\infty\}\) be the collection of all continuous functions with quadratic growth on \(\mathbb{R}^d\). As \(\mu_T\in{\cal P}_2(\mathbb{R}^d)\), we have \[{\cal P}(\mu_0,\mu_T) = \big\{\mathbb{P}\in{\cal P}(\mu_0): \mathbb{E}^\mathbb{P}[\psi(X_T)]=\mu_T(\psi) ~for all~\psi\in {\cal C}_w \big\},\] with \({\cal P}(\mu_0):=\cup_{\nu\in{\cal P}_2(\mathbb{R}^d)}{\cal P}(\mu_0,\nu)\). We may then rewrite our problem as the penalized control problem \[\label{Jpsi} {\rm SBB}(\mu_0,\mu_T) = \inf_{\mathbb{P}\in{\cal P}(\mu_0)} \sup_{\psi\in{\cal C}_w} \mu_T(\psi) + J^\psi(\mathbb{P}), ~J^\psi(\mathbb{P}):=\mathbb{E}^{\mathbb{P}} \Big[\!-\psi(X_T)+\!\!\int_0^T\!\!\!c(\gamma_t^\mathbb{P})\mathrm{d}t \Big].\tag{4}\] Following the standard terminology in optimal transport theory, we call the penalization maps \(\psi\) potential functions. By the trivial inequality \(\inf_{\mathbb{P}\in{\cal P}(\mu_0)} \sup_{\psi\in{\cal C}_w}\ge \sup_{\psi\in{\cal C}_w}\inf_{\mathbb{P}\in{\cal P}(\mu_0)}\), this provides the weak duality inequality \[\label{weakduality1} {\rm SBB}(\mu_0,\mu_T) \ge \sup_{\psi\in{\cal C}_w} \mu_T(\psi) +\inf_{\mathbb{P}\in{\cal P}(\mu_0)}J^\psi(\mathbb{P}).\tag{5}\] We rewrite the minimization on the right-hand side as a standard stochastic control problem by expressing it as the integration of the value function of a standard stochastic control problem with respect to \(\mu_0\) . To do this, we denote for all \(\mathbb{P}\in{\cal P}(\mu_0)\) by \(\{\mathbb{P}^x\}_{x\in\mathbb{R}^d}\) a regular conditional law of \(\mathbb{P}\) given \(X_0\), and we write by the tower property that \(J^\psi(\mathbb{P})=\int J^\psi(\mathbb{P}^x)\mu_0(\mathrm{d}x)\ge\int \inf_{\mathbb{Q}\in{\cal P}(\delta_x)}J^\psi(Q)\mu_0(\mathrm{d}x)\). Combining with 5 , we see that \[\label{weakduality} {\rm SBB}(\mu_0,\mu_T) \;\ge\; {\cal V}(\mu_0,\mu_T) \;:=\; \sup_{\psi\in{\cal C}_w} \mu_T(\psi) -\mu_0(V^\psi_0),\tag{6}\] where \(V^\psi_0\) is the value function of a standard stochastic control problem: \[\label{Vpsi} V^\psi_0(x) := \sup_{\mathbb{P}\in{\cal P}(\delta_x)}\!-J^\psi(\mathbb{P}) = \sup_{\mathbb{P}\in{\cal P}(\delta_x)} \mathbb{E}^\mathbb{P}\Big[\psi(X_T)-\int_0^Tc(\gamma^\mathbb{P}_t)\;\mathrm{d}t\Big], ~~x\in\mathbb{R}^d.\tag{7}\] The penalized optimisation problem \({\cal V}(\mu_0,\mu_T)\) is our first dual problem. Our primary objective in this paper is to prove that there is no duality gap in 6 . In addition, we shall prove existence for both primal and dual problems, together with a complete characterization of the corresponding solution which exhibits a perfect interpolation between the well-known Schrödinger bridge system and the Bass-Brenier transport map.

2.3 HJB equation↩︎

We denote by \(Q_T:=[0,T)\times\mathbb{R}^d\) the time-space parabolic domain, and \(\overline{Q}_T:=[0,T]\times\mathbb{R}^d\) its closure. By standard optimal control theory, the HJB equation corresponding to the problem \(V^\psi_0\) is: \[\label{HJB} \partial_tv+H(Dv,D^2v) =0~on~Q_T, ~and~v(T,\cdot) =\psi~on~\mathbb{R}^d,\tag{8}\] where the Hamiltonian \(H\) is given by \[\begin{align} H(p,A) &:= \sup_{a\in\mathbb{R}^d,\;b\in\mathbb{S}_+^d}\Big\{a\!\cdot\!p+\frac{1}{2}bb^\intercal\!:\!A-c(a,b)\Big\} \\ &= \frac{1}{2}|p|^2 + \frac{\beta}{2}\; I_d\!:\![\beta(\beta I_d - A)^{-1}-I_d] ~on~{\rm dom}(H):=\{(p,A):A<\beta I_d\}, \end{align}\] and \(H=\infty\) outside \({\rm dom}(H)\). We record that the maximizers of the Hamiltonian are: \[\widehat{a}(p)=p,~\widehat{b}(A)=\beta(\beta I_d-A)^{-1}, ~(p,A)\in{\rm dom}(H).\] Notice that finiteness of \(H(Dv,D^2 v)\) formally forces \(D^2v<\beta I_d\), and we then expect that the solution of the HJB equation 8 exhibits a boundary layer in the sense that \(\lim_{t\nearrow T}v(t,\cdot)-\frac{\beta}{2}|.|^2\) is concave. For this reason, it turns out that the set of potential maps can be reduced to the subset of all such \(\beta-\)concave maps: \[{\cal C}_w^{\rm conc} :={\cal C}_w\cap{\cal C}^{\rm conc} ~where~ {\cal C}^{\rm conc}:=\big\{\phi\in C^0(\mathbb{R}^d):\phi ~is \beta-concave \big\}.\] where \(\phi\) is \(\beta-\)concave if \(\phi -\frac{\beta}{2}|\cdot|^2\) is concave. Similarly, we say that \(\phi\) is \(\beta-\)convex if the map \(\phi+\frac{\beta{2}}{|}\cdot |^2\) is convex, and we denote \[{\cal C}_w^{\rm conv} :={\cal C}_w\cap{\cal C}^{\rm conv}, ~with~ {\cal C}^{\rm conv}:=\big\{\phi\!\in\! C^0(\mathbb{R}^d):\phi ~is \beta-convex \big\} =-{\cal C}^{\rm conc}.\]

2.4 Dual potential maps and reduced dual problem↩︎

Due to the particular choice 2 of the cost function, the HJB equation 8 turns out to be amenable to interesting manipulation leading to more explicit solution structure. We shall use extensively the Moreau transform, also called inf–convolution: \[\begin{align} \mathbf{T}_\beta^+[\phi](x) := \inf_{y\in\mathbb{R}^d} \phi(y)+\frac{\beta{2}}{|}x-y|^2, &x\in\mathbb{R}^d.& \end{align}\] Notice that \(\psi:=\mathbf{T}_\beta^+[\phi]\) is \(\beta-\)concave as \(\mathbf{T}_\beta^+[\phi]-\frac{\beta{2}}{|}\cdot |^2\) is an infimum of affine functions. We also introduce the dual operator \(\mathbf{T}_\beta^-[\psi]:=-\mathbf{T}_\beta^+[-\psi]\). We shall use the standard Moreau biconjugation property: if \(\phi\) is proper, lower semicontinuous and \(\beta\)-convex, then \[\label{TbetaTbeta} \mathbf{T}_\beta^-\circ\mathbf{T}_\beta^+[\phi]=\phi.\tag{9}\] More generally, for arbitrary \(\phi\) such that \(\mathbf{T}_\beta^+[\phi]\) is finite-valued, the function \(\phi^\beta:=\mathbf{T}_\beta^-\circ\mathbf{T}_\beta^+[\phi]\) is \(\beta\)-convex, satisfies \(\phi^\beta\le \phi\), and \(\mathbf{T}_\beta^+[\phi^\beta]=\mathbf{T}_\beta^+[\phi]\). When \(\phi=\mathbf{T}_\beta^-[\psi]\) for some potential \(\psi\), we call \(\phi\) a dual potential map.

Our main characterization result of the SBB problem reduces the dual problem \({\cal V}(\mu_0,\mu_T)\) to the following optimisation problem over the set of dual potential maps: \[\begin{align} \label{D} {\cal V}_{\rm red}(\mu_0,\mu_T) := \sup_{\phi\in {\cal C}^{\rm conv}_{w}} \mathfrak{J}(\phi), ~~with~~& \mathfrak{J}(\phi) := \mu_T\big(\mathbf{T}_\beta^+[\phi]\big) - \mu_0\big(\mathbf{T}_\beta^+[u_T^\phi]\big) \\ & e^{u_s^\phi(y)}:={\cal N}_{s}*e^{\phi}(y), ~(s,y)\in\overline{Q}_T, \nonumber \end{align}\tag{10}\] where \({\cal N}_s:=(2\pi s)^{-\frac{d}{2}}e^{{-\frac{|.|^2}{2s}}}\) is the heat kernel in \(\mathbb{R}^d\), with the convention \({\cal N}_0*e^{\phi} = e^{\phi}\). The justification of this expression for the dual is clarified by the subsequent Lemma 4 (3-b) below, which connects \(\mathbf{T}_\beta^+[u_T^\phi]\) to the HJB equation corresponding to the control problem 7 .

Remark 2. The restriction of the dual functions to be \(\beta-\)convex in 10 can be relaxed. To see this, set \(\tilde{\phi}:=\mathbf{T}_\beta^-\circ\mathbf{T}_\beta^+[\phi]\). By the above Moreau envelope property, \(\tilde{\phi}\) is \(\beta\)-convex, \(\tilde{\phi}\le \phi\), and \(\mathbf{T}_\beta^+[\tilde{\phi}]=\mathbf{T}_\beta^+[\phi]\). Hence \(u_T^{\tilde{\phi}}\le u_T^\phi\), and therefore \(\mathfrak{J}(\tilde{\phi})\ge\mathfrak{J}(\phi)\). Thus the \(\beta\)-convexity restriction on the dual potential does not affect the dual maximization problem.

2.5 A relaxed dual formulation↩︎

It turns out that the formulation 6 of the dual problem \({\cal V}\) does not guarantee existence of an optimal potential map. For this reason, we need to introduce an appropriate relaxation so as to allow for existence while preserving the same value. First, from the previous considerations, we already know that it is sufficient to restrict the dual maps to those \(\beta-\)concave maps. As \(\beta-\)concave maps are bounded from above by a quadratic function, it remains to specify the asymptotic behavior of \(\psi^-\) at infinity. Instead, we introduce the following relaxed dual space by enforcing an appropriate integrability condition: \[\label{barCconc} \bar{\cal C}^{\rm conc} := \Big\{\psi~is finite-valued and \beta-concave:\;\mu_0\Big( \mathbf{T}_\beta^+\Big[u_T^{\mathbf{T}_\beta^-[\psi]}\Big]\Big)<\infty \Big\},\tag{11}\] where \(V^\psi_0\) for \(\psi\in\bar{\cal C}^{\rm conc}\) is defined as in 7 with the expectation understood in the extended sense. The quantity \(\mu_T(\psi)-\mu_0(V^\psi_0)\) is also understood in the extended sense, with the convention that it is equal to \(-\infty\) whenever the positive and negative parts are not both well-defined.

Our relaxed dual formulation is defined as \[\label{barVc} \bar{\cal V}(\mu_0,\mu_T) := \sup_{\psi\in\bar{\cal C}^{\rm conc}} \big\{\mu_T(\psi)-\mu_0(V^\psi_0)\big\}.\tag{12}\]

3 Main results↩︎

The main objective of this paper is to establish strong duality between the primal problem \({\rm SBB}\) in 3 and the dual problem \({\cal V}\) in 6 . To ensure existence, the latter is relaxed to 12 ; moreover, it can be expressed in the reduced form \({\cal V}_{\rm red}\) defined in 10 .

Theorem 3. Let \(\mu_0,\mu_T\in{\cal P}_2(\mathbb{R}^d)\). Then

\({\rm SBB}={\cal V}=\bar{\cal V}={\cal V}_{\rm red}\in[0,\infty)\) at \((\mu_0,\mu_T)\);

Under the additional condition \(\beta T>1\):

  • The reduced dual problem \({\cal V}_{\rm red}\) has a solution \(\hat{\phi}\in {\cal C}^{\rm conv}_{w}\), which induces a solution \(\hat{\psi}:=\mathbf{T}_\beta^+[\hat{\phi}]\in\bar{\cal C}^{\rm conc}\) of the dual problem \(\bar{\cal V}\);

  • The primal problem SBB has a solution \(\hat{\mathbb{P}}\in{\cal P}(\mu_0,\mu_T)\);

  • For all \(t<T\), the map \(\mathscr{Y}_t:={\rm id}-\frac{1}{\beta}\nabla \mathbf{T}_\beta^+[u^{\hat{\phi}}_{T-t}]\) is a well-defined one-to-one map, and we denote its inverse by \(\mathscr{X}_t:=\mathscr{Y}_t^{-1}\). At terminal time \(t=T\), \(\mathscr{Y}_T\) is understood as the corresponding set-valued argmin map. Moreover, there exists a probability measure \(\hat{\mathbb{Q}}\) under which \(Y\) is a Schrödinger bridge diffusion,

    \[Y_t = Y_0+\int_0^t \nabla u^{\hat{\phi}}_{T-s}(Y_s)\,\mathrm{d}s +W_t^{\hat{\mathbb{Q}}}, \qquad 0\le t<T,\]

    and the process \(X_t:=\mathscr{X}_t(Y_t)\), \(0\le t<T\), admits a continuous extension to \([0,T]\) satisfying \(\hat{\mathbb{P}}=\text{Law}_{\hat{\mathbb{Q}}}(X)\). Equivalently, if \(\mathbf{R}\) denotes the Brownian law with initial distribution \(m_0:=\mathscr{Y}_0\#\mu_0\), then:

    \[\frac{d\hat{\mathbb{Q}}}{d\mathbf{R}}\Big|_{{\cal F}_t} = e^{u^{\hat{\phi}}_{T-t}(Y_t)-u^{\hat{\phi}}_T(Y_0)}, \qquad 0\le t<T.\]

  • There exists a non-zero \(\sigma\)-finite measure \(\nu_0\) such that, setting \(\nu_t={\cal N}_t\!*\!\nu_0\) and \(m_T:=e^{\hat{\phi}}\nu_T\), the dual optimiser \(\hat{\phi}\) is characterized by the SBB system

    \[\frac{\mathrm{d}\mathscr{Y}_0\#\mu_0}{\mathrm{d}\nu_0} = e^{u^{\hat{\phi}}_{T}}, \qquad \mu_T=\mathscr{X}_T\#m_T.\]

We report the proof of the duality (i) in Sections 4 and 5, and the remaining part (ii) in Sections 6 and 7. To better understand the last characterization of our solution, recall that the Schrödinger bridge corresponds to the formal limit \(\beta\to\infty\), where \(\mathscr{Y}\) formally reduces to \(\mathscr{Y}_t={\rm id}\) for all \(t\in[0,T]\), while the Bass martingale corresponds to the formal limit \(\beta\to0\), where \(\hat{\phi}=0\). We observe here that the characterization (ii-c) in the last statement combines features of the Schrödinger bridge and the Bass martingale:

  • Similar to the structure of the Schrödinger bridge, the flow of probability measures \((\nu_t)_{t\le T}\) satisfies the forward Fokker-Planck equation, while the map \((t,y)\) \(\mapsto\) \(e^{u_{T-t}^{\hat{\phi}}(y)}\) is a solution of the backward heat equation;

  • The map \(\mathscr{Y}\) is the gradient of the convex map \(\frac{1}{\beta}(\mathbf{T}_\beta^+[u^{\hat{\phi}}_{T-t}]-\frac{\beta{2}}{|}.|^2)\) is our substitute for the Brenier map in the Bass martingale component of our characterization. Its inverse \(\mathscr{X}\) is expressed as \(\mathscr{X}_t\) \(=\) \({\rm id} + \frac{1}{\beta}\nabla u_{T-t}^{\hat{\phi}}\).

We represent the SBB system in the standard graphical form adopted in the Schrödinger bridge literature as: \[\begin{array}{ccccccl} \mu_T & \!\!\!\!-\!\!\!-\!\!\!-\!\!\!\longrightarrow &\!\!\!\!m_T:=\mathscr{Y}_T \# \mu_T & \!\!\!\!-\!\!\!-\!\!\!-\!\!\!\longrightarrow & \!\!\!\!{ \frac{dm_T}{d\nu_T} = e^{\hat{\phi}}} & \!\!\!\!-\!\!\!-\!\!\!-\!\!\!\longrightarrow & \nu_T \\ \Big\uparrow & & & & & &\Big\uparrow \\ \Big| & &\colorbox{gray}{\color{white} ~~~~Bass~~~~}& && \colorbox{gray}{\color{white} Schr\"odinger} & \Big| \\ \Big| & & & & & & \Big| \\ \mu_0 & \!\!\!\!-\!\!\!-\!\!\!-\!\!\!\longrightarrow & \!\!\!\!m_0:=\mathscr{Y}_0 \# \mu_0 & \!\!\!\!-\!\!\!-\!\!\!-\!\!\!\longrightarrow & \!\!\!\!\frac{dm_0}{d\nu_0} = e^{u^{\hat{\phi}}_T} & \!\!\!\!-\!\!\!-\!\!\!-\!\!\!\longrightarrow & \nu_0 \end{array}\] At terminal time, \(\mathscr{Y}_T\) is understood as the argmin correspondence, and the rigorous terminal relation is \(\frac{dm_T}{d\nu_T}=e^{\hat{\phi}}, \;\mu_T=(\mathscr{X}_T)_\#m_T\), rather than a deterministic push-forward identity \(m_T=\mathscr{Y}_T\#\mu_T\).

One can view the SBB solution as a stretched Schrödinger diffusion, i.e., a Bass transport of a Schrödinger bridge.

4 Proof of the first duality results↩︎

In order to prove the first duality result SBB \(={\cal V}\) in Theorem 3 (i), we focus on the dependence of the value function SBB on its second argument. We thus fix \(\mu_0\in{\cal P}_2(\mathbb{R}^d)\) and we analyse in this section the map \(F:{\cal P}_2(\mathbb{R}^d)\longrightarrow\mathbb{R}\) defined by: \[F(m):={\rm SBB}(\mu_0,m),~m\in{\cal P}_2(\mathbb{R}^d).\] Let \({\cal M}_2\) denote the vector space of finite signed measures with finite second moment, endowed with the topology \(\sigma({\cal M}_2,{\cal C}_w)\).

Lemma 2. The map \(F\) is convex and continuous on \({\cal P}_2(\mathbb{R}^d)\).

1. We first show that \(F\) is convex on \({\cal P}_2(\mathbb{R}^d)\). Let \(m_0,m_1\in{\cal P}_2(\mathbb{R}^d)\) and \(\lambda\in(0,1)\). Fix \(\eta>0\) and choose \(\mathbb{P}_i\in{\cal P}(\mu_0,m_i)\) with \(\mathbb{E}^{\mathbb{P}_i}\big[\int_0^T c(\gamma_t^{\mathbb{P}_i})\mathrm{d}t \big] \le F(m_i)+\eta\), \(i=0,1\).

Consider the enlarged space \(\tilde{\Omega}:=\Omega\times\{0,1\}\) with canonical process \((\tilde{X},U)\) and canonical filtration \(\tilde{\mathbb{F}}=\mathbb{F}^{\tilde{X}}\vee\mathbb{F}^U\), \(\mathbb{F}^{\tilde{X}}=\{{\cal F}^{\tilde{X}}_t=\sigma(\widetilde{X}_s,s\le t)\}_{t\le T}\) and \(\mathbb{F}^{U}=\{{\cal F}^{U}_t=\sigma(U)\}_{t\le T}\). Let \(\tilde{\mathbb{P}}\) be the probability measure under which \(U\) is a Bernoulli\((\lambda)\) variable independent of \(\tilde{X}_0\), \(\widetilde{X}\) has conditional law \(\mathbb{P}_i\) given \(\{U=i\}\), \(i=0,1\), and thus its characteristics \(\tilde{\gamma}\) are defined by \(\tilde{\gamma}^i:=(\tilde{\alpha}^i,\tilde{\sigma}^i)\) conditionally on \(\{U=i\}\).

By the Markovian projection argument of Gyöngi [11], we see that \(\mathbb{P}:=\tilde{\mathbb{P}}\circ \widetilde{X}^{-1}\in{\cal P}(\mu_0,m_\lambda)\), \(m_\lambda:=(1-\lambda)m_0+\lambda m_1\), with characteristics \(\gamma_t^\mathbb{P}=(\alpha_t^\mathbb{P},\sigma_t^\mathbb{P})\) defined by \((\alpha_t^\mathbb{P},\sigma_t^\mathbb{P}(\sigma_t^\mathbb{P})^\intercal)=\mathbb{E}^{\widetilde{\mathbb{P}}}[(\widetilde{\alpha}_t,\widetilde{\sigma}_t(\widetilde{\sigma}_t)^\intercal)| {\cal F}^{\tilde{X}}_t]\), and by the convexity of the cost function \(c(a,b)\) in \((a,bb^\intercal)\), we see that \[\begin{align} F(m_\lambda) \le \mathbb{E}^{\mathbb{P}}\!\int_0^T \!\!c(\gamma_t^{\mathbb{P}})\;dt & \le (1-\lambda) \mathbb{E}^{\mathbb{P}_0}\!\int_0^T \!\! c(\gamma_t^{\mathbb{P}_0})\;\mathrm{d}t +\lambda\mathbb{E}^{\mathbb{P}_1}\!\int_0^T \!\! c(\gamma_t^{\mathbb{P}_1})\;\mathrm{d}t \\ &\le (1-\lambda) F(m_0) +\lambda F(m_1)+2\eta. \end{align}\] The required convexity of \(F\) now follows from the arbitrariness of \(\eta>0\).

We next show the continuity of \(F\). We shall prove in Step 3 below that \[\label{ineq:step3} F(m) \le (1+C_0\kappa)F(m')+d(1+\beta)\kappa, ~whenever~ \kappa:=W_2(m,m')\le\frac{T}{2}.\tag{13}\] Let \(m_n\in{\cal P}_2(\mathbb{R}^d)\) with \(m_n\to m\) in \(\sigma({\cal M}_2,{\cal C}_w)\). Then \(m_n\rightharpoonup m\) and, since \(|.|^2\in{\cal C}_w\), \(\int|.|^2\mathrm{d}m_n\to\int|.|^2\mathrm{d}m\), which implies \(W_2(m_n,m)\to0\). Then, as \(\kappa_n:=W_2(m_n,m)\le\frac{T}{2}\) for large \(n\), it follows from 13 that: \[F(m) \le \liminf_{n\to\infty}(1+C_0\kappa_n)F(m_n)+d(1+\beta)\kappa_n =\liminf_{n\to\infty}F(m_n).\] Applying 13 again with \(m\) and \(m_n\) exchanged, we obtain the reverse inequality. Hence \[\lim_{n\to\infty}F(m_n)=F(m),\] which proves the continuity of \(F\).

It remains to prove 13 . Fix \(\eta>0\) and pick \(\mathbb{P}\in{\cal P}(\mu_0,m')\) such that \[\label{eq:nearmin} \mathbb{E}^{\mathbb{P}}\!\int_0^T c(\gamma_t^\mathbb{P})\mathrm{d}t \le F(m')+\eta.\tag{14}\] 3-a. Time change on \([0,T-\kappa]\). Set \(\theta:=\frac{T}{T-\kappa}\) and define on the same space \(\widetilde{X}_u:=X_{\theta u}\) for \(u\in[0,T-\kappa]\). Let \(\widetilde{\mathbb{P}}\) be the law of \(\widetilde{X}\) on \(C([0,T-\kappa];\mathbb{R}^d)\). Then \(\widetilde{X}_0\sim\mu_0\) and \(\widetilde{X}_{T-\kappa}=X_T\sim m'\). A direct change-of-variables computation in the martingale problem shows that \(\widetilde{X}\) has characteristics \(\widetilde{\alpha}_u=\theta\alpha_{\theta u}^\mathbb{P}\) and \(\widetilde{\sigma}_u=\sqrt\theta\, \sigma_{\theta u}^\mathbb{P}\), and therefore: \[\mathbb{E}^{\widetilde{\mathbb{P}}}\!\int_0^{T-\kappa} c(\widetilde{\alpha}_t,\widetilde{\sigma}_t)\;\mathrm{d}t = \frac{1}{2}\theta \; \mathbb{E}^{\mathbb{P}}\!\int_0^T\big(|\alpha_t^\mathbb{P}|^2 +\frac{\beta}{\theta^2} |\sqrt\theta(\sigma_t^\mathbb{P}-I_d)+(\sqrt{\theta}-1)I_d|^2 \big)\;\mathrm{d}t.\] Note that, for \(b\in\mathbb{R}^{d\times d}\), \(|\sqrt{\theta}b+(\sqrt{\theta}-1)I_d|^2 = \theta |b|^2 +2\sqrt{\theta}(\sqrt{\theta}-1)b:I_d +(\sqrt{\theta}-1)^2 d\). By Young’s inequality, \(2\sqrt{\theta}(\sqrt{\theta}-1)b:I_d \le \theta(\theta-1)|b|^2 + \frac{(\sqrt{\theta}-1)^2}{\theta-1}d\). Hence \(|\sqrt{\theta}b+(\sqrt{\theta}-1)I_d|^2 \le \theta^2|b|^2+\frac{\theta(\sqrt{\theta}-1)}{\sqrt{\theta}+1}d\). We deduce from 14 that \[\begin{align} \mathbb{E}^{\widetilde{\mathbb{P}}}\!\int_0^{T-\kappa} c(\widetilde{\alpha}_t,\widetilde{\sigma}_t)\;\mathrm{d}t &\le& \theta \; \mathbb{E}^{\mathbb{P}}\!\int_0^Tc(\alpha_t^\mathbb{P},\sigma^\mathbb{P}_t)\; \mathrm{d}t +\frac{T\beta d(\sqrt\theta-1)}{2(\sqrt\theta+1)} \nonumber\\ &\le& \theta (F(m')+\eta) +\frac{T\beta d(\sqrt\theta-1)}{2(\sqrt\theta+1)}. \label{Flsc1} \end{align}\tag{15}\]

3-b. Terminal correction on \([T-\kappa,T]\). Let \(\pi\) be an optimal \(W_2\)-coupling between \(m'\) and \(m\), disintegrated as \(\pi(dy| x)m'(dx)\). Enlarge the space and sample \(Y\sim \pi(\cdot\mid \widetilde{X}_{T-\kappa})\). Then \(\text{\rm Law}(Y)=m\) and \(\mathbb{E}|Y-\widetilde{X}_{T-\kappa}|^2=W_2(m',m)^2=\kappa^2\). Define \[\widehat X_t := \mathbf{1}_{\{t\in[0,T-\kappa]\}} \widetilde{X}_t +\mathbf{1}_{\{t\in[T-\kappa,T]\}} \Big\{\widetilde{X}_{T-\kappa}+\dfrac{t-(T-\kappa)}{\kappa}\bigl(Y-\widetilde{X}_{T-\kappa}\bigr)\Big\}.\] On \([T-\kappa,T]\), \(\widehat X\) has drift \(\widehat\alpha_t=(Y-\widetilde{X}_{T-\kappa})/\kappa\) and diffusion \(\widehat\sigma_t\equiv0\), hence \[\label{eq:bridge} \mathbb{E}\!\int_{T-\kappa}^T c(\widehat\alpha_t,\widehat\sigma_t)\;dt =\frac{W_2(m',m)^2}{2\kappa} + \frac{\beta}{2} d\;\kappa =\frac{\kappa}{2} + \frac{\beta}{2} d\;\kappa.\tag{16}\] Let \(\widehat \mathbb{P}\) be the law of \(\widehat X\) on the canonical space. By a Markovian projection argument as in Step 1 of the current proof, we see that \(\widehat \mathbb{P}\in{\cal P}(\mu_0,m)\) and \[\begin{align} \label{eq:jensenbridge} F(m) \le \mathbb{E}^{\widehat \mathbb{P}}\!\int_0^T c(\alpha_t^{\widehat \mathbb{P}},\sigma_t^{\widehat \mathbb{P}})\;dt &\le \mathbb{E}\!\int_0^T c(\widehat\alpha_t,\widehat\sigma_t)\;dt \\ &\le \theta (F(m')+\eta) +\frac{T\beta d(\sqrt\theta-1)}{2(\sqrt\theta+1)} +\frac{\kappa}{2} + \frac{\beta}{2} d\;\kappa, \end{align}\tag{17}\] where the last inequality follows from 15 and 16 . Substituting \(\theta:=\frac{T}{T-\kappa}\) and using \(\kappa\le T/2\), we have \(\theta=\frac{T}{T-\kappa}\le 1+\frac{2\kappa}{T}, \; \frac{\sqrt\theta-1}{\sqrt\theta+1}\le C_T\kappa\) for some constant \(C_T<\infty\) depending only on \(T\). Letting \(\eta\downarrow0\) and enlarging the constant if necessary gives 13 .   \({\cal t}\)  \({\cal u}\)

We are now ready for the first duality result in Theorem 3 (i).

Proposition 4. For \(\mu_0,\mu_T\in{\cal P}_2(\mathbb{R}^d)\), we have \({\rm SBB}(\mu_0,\mu_T)={\cal V}(\mu_0,\mu_T)\).

Define \(\bar F:{\cal M}_2\to \mathbb{R}\cup\{+\infty\}\) by \[\bar F(m):= \left\{ \begin{array}{ll} F(m), & ifm\in{\cal P}_2(\mathbb{R}^d),\\ +\infty, & ifm\in{\cal M}_2\setminus{\cal P}_2(\mathbb{R}^d). \end{array} \right.\] Since \(F\) is convex on \({\cal P}_2(\mathbb{R}^d)\), the map \(\bar F\) is convex on \({\cal M}_2\). Moreover, \(\bar F\) is \(\sigma({\cal M}_2,{\cal C}_w)-\)lower semicontinuous on \({\cal M}_2\). Then, \(F(\mu_T)=\bar F(\mu_T)=\bar F^{**}(\mu_T)\) by the Fenchel-Moreau theorem, where: \[\bar F^{**}(m)=\sup_{\psi\in{\cal C}_w}m(\psi)-\bar F^*(\psi) ~and~ \bar F^{*}(\psi)=\sup_{m\in{\cal M}_2}m(\psi)-\bar F(m) =\sup_{m\in{\cal P}_2}m(\psi)-F(m)\] Hence \(F(\mu_T)=\bar F(\mu_T)=\bar F^{**}(\mu_T).\)

We now complete the required dual representation by identifying the map \(\bar F^{**}\) to the value \({\cal V}(\mu_0,\mu_T)\) introduced in 6 . First, \[\begin{align} \bar F^*(\psi) &= \sup_{m\in{\cal P}_2}m(\psi)-F(m) = \sup_{m\in{\cal P}_2}m(\psi)-\inf_{\mathbb{P}\in{\cal P}(\mu_0,m)}J^0(\mathbb{P}) = \sup_{\mathbb{P}\in{\cal P}(\mu_0)}-J^\psi(\mathbb{P}), \end{align}\] and by Step 1, \[F(\mu_T) =\bar F^{**}(\mu_T) = \sup_{\psi\in{\cal C}_w}\mu_T(\psi)-\bar F^*(\psi) = \sup_{\psi\in{\cal C}_w}\mu_T(\psi)-\sup_{\mathbb{P}\in{\cal P}(\mu_0)}-J^\psi(\mathbb{P}).\] We now prove that \(\bar F^{**}(\mu_T)={\cal V}(\mu_0,\mu_T)\) as defined in 6 . This is similar to Lemma 3.5 in [8] which was established under their coercivity condition on the cost function.

2-a. We first show that it suffices to prove the required result for potential maps \(\psi\) with \(|\psi^+|_{\infty}<\infty\). Indeed, for \(\psi\in{\cal C}_w\), define \(\psi_k:=\psi\wedge k\), and notice that as \(\psi_k\uparrow\psi\), \(c\ge0\), and \(\psi^-(X_T)\in \mathbb{L}^1(\mathbb{P})\), it follows from the monotone convergence theorem that \(J^{\psi_k}(\mathbb{P})\downarrow J^\psi(\mathbb{P})\). By the standard approximate optimiser, we also obtain the monotone convergence \(\inf_{\mathbb{P}\in{\cal P}(\mu_0)}J^{\psi_k}(\mathbb{P}) \downarrow\inf_{\mathbb{P}\in{\cal P}(\mu_0)}J^{\psi}(\mathbb{P})\).

The inequality \(\inf_{\mathbb{P}\in{\cal P}(\mu_0)}J^\psi(\mathbb{P})\ge-\mu_0(V_0^\psi)\) was already established right before introducing the dual problem \({\cal V}(\mu_0,\mu_T)\) in 6 . We now prove the reverse inequality by means of a measurable selection argument. Since \(|\psi^+|_\infty<\infty\) and \(c\ge0\), we have \(-J^\psi(\mathbb{P}) = \mathbb{E}^\mathbb{P}\Big[\psi(X_T)-\int_0^T c(\gamma_t^\mathbb{P})\mathrm{d}t\Big] \le |\psi^+|_\infty, \; \mathbb{P}\in{\cal P}(\mu_0)\), and thus \(V^\psi_0<\infty\). For fixed \(\eta>0\), it follows from the measurable selection result, see e.g. [8], that there exists a \(\mu_0\)-measurable map \(x\mapsto\mathbb{P}^{x,\eta}\in{\cal P}(\delta_x)\) such that \(J^{\psi}(\mathbb{P}^{x,\eta}) \le -V^\psi_0(x)+\eta ~\text{for }\mu_0\text{-a.e. }x.\) Define \(\mathbb{P}^\eta:=\int_{\mathbb{R}^d} \mathbb{P}^{x,\eta}(\cdot)\,\mu_0(dx)\) on \(\Omega\). As \(\mathbb{P}^\eta[\cdot\,|\,X_0=x]=\mathbb{P}^{x,\eta}\), \(\mu_0\)-a.s., we have \(\mathbb{P}^\eta\in{\cal P}(\mu_0)\), and it follows from Fubini’s theorem that \[J^{\psi}(\mathbb{P}^\eta) =\int_{\mathbb{R}^d} J^\psi(\mathbb{P}^{x,\eta})\,\mu_0(dx) \le -\int_{\mathbb{R}^d} V^\psi_0(x)\,\mu_0(dx)+\eta.\] This shows that \(\mathbb{P}^\eta\) is an \(\eta-\)optimiser, and we deduce the required result by the arbitrariness of \(\eta>0\).   \({\cal t}\)  \({\cal u}\)

In preparation for the proof of the reduced dual formulation in the next subsection, we now show that we may restrict the dual maximization to upper bounded \(\beta\)–concave potentials \(\psi\): \[{\cal C}_{w,\uparrow}^{\text{conc}}:={\cal C}^{\rm conc}\cap{\cal C}_{w,\uparrow} ~with~ {\cal C}_{w,\uparrow}:=\{\psi\in {\cal C}_w:\;|\psi^+|_\infty<\infty\}.\] We also set \({\cal C}_{w,\uparrow}^{\mathrm{conv}} := \big\{\mathbf{T}_\beta^-[\psi]:\psi\in{\cal C}_{w,\uparrow}^{\text{conc}}\big\}\).

For \(t<T\) and \(x\in\mathbb{R}^d\), we use the shifted notation \({\cal P}_t(\delta_x)\) for the admissible laws on \([t,T]\) starting from \(x\) at time \(t\), and \[V_t^\psi(x) := \sup_{\mathbb{P}\in{\cal P}_t(\delta_x)} \mathbb{E}^\mathbb{P}\Big[ \psi(X_T)-\int_t^T c(\gamma_r^\mathbb{P})\,\mathrm{d}r \Big].\]

Lemma 3. We have \[{\cal V}(\mu_0,\mu_T) = \underline{{\cal V}}(\mu_0,\mu_T) := \sup_{\psi\in{\cal C}_{w,\uparrow}^{\text{conc}}} \{\mu_T(\psi)-\mu_0(V_0^\psi)\}.\] Moreover, \[\label{eq:dual-enlarged-conc} \mu_T(\psi)-\mu_0(V_0^\psi) \le {\cal V}(\mu_0,\mu_T), ~~for all~~ \psi\in{\cal C}^{\text{conc}},\qquad{(1)}\] where the objective is understood in the extended sense.

(a) We first restrict the dual maximization to \({\cal C}_{w,\uparrow}\). For arbitrary \(\psi\in{\cal C}_w\), we have \(\psi^n:=\psi\wedge n\in{\cal C}_{w,\uparrow}\), \(n\ge1\). As \(\psi^n\uparrow\psi\) and \(\psi^n\ge\psi^1\in\mathbb{L}^1(\mu_T)\), it follows by monotone convergence that \(\mu_T(\psi^n)\;\uparrow\;\mu_T(\psi).\)

On the other hand, monotonicity in the terminal payoff gives \(V_0^{\psi^n}\uparrow V_0^\psi\) pointwise. As \(A:=|\frac{\psi{w}}{|}_\infty<\infty\), and the Wiener measure \(\mathbb{Q}^{0,x}\in{\cal P}(\delta_x)\) started from \(\delta_x\) has characteristics \(\gamma^{\mathbb{Q}^{0,x}}=(0,I_d)\) inducing the cost \(c(\gamma^{\mathbb{Q}^{0,x}})=0\), we have \[V_0^{\psi^n}(x)\ge E^{\mathbb{Q}^{0,x}}[\psi^n(X_T)] \ge -A(1+\mathbb{E}^{\mathbb{Q}^{0,x}}|X_T|^2) = -A(1+|x|^2+dT), ~for all~n\ge 1.\] Then \(V_0^{\psi^n}\ge V_0^{\psi^1}\in \mathbb{L}^1(\mu_0)\), and we get by monotone convergence \(\mu_0(V_0^{\psi^n})\;\uparrow\;\mu_0(V_0^\psi).\)

Hence \(\mu_T(\psi^n)-\mu_0(V_0^{\psi^n})\longrightarrow \mu_T(\psi)-\mu_0(V_0^\psi)\), and we deduce from the arbitrariness of \(\psi\in{\cal C}_w\) that \({\cal V}(\mu_0,\mu_T) = \sup_{\psi\in{\cal C}_{w,\uparrow}} \{\mu_T(\psi)-\mu_0(V_0^\psi)\}\).

(b) We next show that we may restrict further the dual maximization to \(\beta\)–concave potentials. Fix now \(\psi\in {\cal C}_{w,\uparrow}\) and define its \(\beta\)–concave envelope \[\hat{\psi}:={\mathbf{T}}_\beta^+\circ {\mathbf{T}}_\beta^-[\psi]\in {\cal C}_w^{\text{conc}}, ~\text{so that}~\psi\le \hat{\psi}\le |\psi^+|_\infty.\] On the other hand, notice that \(V^\psi\) is l.s.c. as the supremum of continuous maps, and it follows from the easy part of dynamic programming principle (DPP) that \(V^\psi\) is a l.s.c.viscosity supersolution of the HJB equation 8 on \([0,T)\times\mathbb{R}^d\). Since the Hamiltonian \(H\) is infinite outside its domain \({\rm dom}(H):=\{(p,A):A<\beta I_d\}\), we deduce that \(V_t^\psi\) is \(\beta\)–concave for every \(t<T\). As \(V_t^\psi\ge \psi-\frac{\beta d}{2}(T-t)\) (take the control \(\gamma\equiv0\)), we have \[\label{eq:terminal-layer-beta-envelope} \liminf_{\substack{s\uparrow T\\ y\to x}} V_s^\psi(y)\ge \hat{\psi}(x), \qquad x\in\mathbb{R}^d .\tag{18}\] We now propagate this terminal estimate backward by dynamic programming. Fix \((t,x)\in[0,T)\times\mathbb{R}^d\) and let \(\mathbb{P}\in{\cal P}_t(\delta_x)\) be such that \(\mathbb{E}^\mathbb{P}\int_t^T c(\gamma_r^\mathbb{P})\,\mathrm{d}r<\infty\). For \(s\in(t,T)\), the dynamic programming principle gives \(V_t^\psi(x) \ge \mathbb{E}^\mathbb{P}\left[ V_s^\psi(X_s)-\int_t^s c(\gamma_r^\mathbb{P})\,\mathrm{d}r \right]\). Let \(s\uparrow T\). Since \(X_s\to X_T\), the boundary layer estimate 18 gives \(\liminf_{s\uparrow T} V_s^\psi(X_s)\ge \hat{\psi}(X_T)\), \(\mathbb{P}\)-a.s. To apply Fatou’s lemma, we use the zero control only as a lower bound. Starting from \(z\) at time \(s\), the control \((\alpha,\sigma)=(0,0)\) gives \(V_s^\psi(z) \ge \psi(z)-\frac{\beta d}{2}(T-s)\). Since \(\psi\in{\cal C}_w\), there exists \(C>0\) such that \(\psi(z)\ge -C(1+|z|^2)\). Hence, for \(s\) close to \(T\), \[V_s^\psi(X_s) \ge -C\Big(1+\sup_{r\in[t,T]}|X_r|^2\Big)-\frac{\beta d}{2}T.\] The right-hand side is \(\mathbb{P}\)-integrable by the same BDG estimate as in Lemma 1. Fatou’s lemma therefore yields \(\liminf_{s\uparrow T}\mathbb{E}^\mathbb{P}[V_s^\psi(X_s)] \ge \mathbb{E}^\mathbb{P}[\hat{\psi}(X_T)]\). Moreover, \[\int_t^s c(\gamma_r^\mathbb{P})\,\mathrm{d}r \uparrow \int_t^T c(\gamma_r^\mathbb{P})\,\mathrm{d}r. \quad \text{Consequently,} \quad V_t^\psi(x) \ge \mathbb{E}^\mathbb{P}\left[ \hat{\psi}(X_T)-\int_t^T c(\gamma_r^\mathbb{P})\,\mathrm{d}r \right].\] Taking the supremum over all such \(\mathbb{P}\in{\cal P}_t(\delta_x)\) gives \(V_t^\psi(x)\ge V_t^{\hat{\psi}}(x)\). Together with the opposite inequality \(V_t^\psi\le V_t^{\hat{\psi}}\), we obtain \(V_t^\psi=V_t^{\hat{\psi}} \quadon [0,T)\times\mathbb{R}^d\). In particular \(V_0^\psi=V_0^{\hat{\psi}}\), and therefore \(\mu_T(\psi)-\mu_0(V_0^\psi) \le \mu_T(\hat{\psi})-\mu_0(V_0^{\hat{\psi}})\), which yields the desired restriction to \(\beta\)–concave potentials.

(c) Fix \(\psi\in {\cal C}^{\rm conc}\) and \((t,x)\in[0,T)\times\mathbb{R}^d\). Fix any \(y\in\mathbb{R}^d\) and consider on \([t,T]\) the deterministic admissible characteristics \(\alpha_s\equiv \frac{y-x}{T-t}, \; \sigma_s\equiv 0\). Then \(X_T=y\) a.s.and \(\int_t^T c(\alpha_s,\sigma_s)\;ds = \frac{|y-x|^2}{2(T-t)}+\frac{\beta d}{2}(T-t)\), since \(| I_d|^2=\text{tr}( I_d)=d\). Hence \[\label{eq:V-lower-stepc} V_t^\psi(x) \ge \psi(y)-\frac{|y-x|^2}{2(T-t)}-\frac{\beta d}{2}(T-t) \ge -C_{t,\psi}(1+|x|^2).\tag{19}\]

For \(n\ge 1\) set \(\psi_n:=\psi\wedge n\). Then \(\psi_n\uparrow\psi\) and \(\psi_n(y)\ge \psi_1(y)\), so 19 yields the uniform-in-\(n\) bound \[V_t^{\psi_n}(x)\ge \psi_1(y)-\frac{|y-x|^2}{2(T-t)}-\frac{\beta d}{2}(T-t)\ge -C_t(1+|x|^2), \qquad n\ge 1,\] with \(C_t<\infty\) depending on \(t,\beta,d,y,\psi_1(y)\) but not on \(n\).

(d) Fix \(\psi\in{\cal C}^{\rm conc}\) and set \(\psi_n:=\psi\wedge n\).

(d1) \(\mu_T(\psi_n)\uparrow \mu_T(\psi)\) in the extended sense. Since \(\psi_n=\psi\wedge n\), we have \((\psi_n)^-=\psi^-\) and \((\psi_n)^+\uparrow \psi^+\). Because \(\psi(x)\le C(1+|x|^2)\) and \(\mu_T\in{\cal P}_2(\mathbb{R}^d)\), we have \(\psi^+\in \mathbb{L}^1(\mu_T)\). If \(\int_{\mathbb{R}^d}\psi^-\;d\mu_T=\infty\), then \(\int_{\mathbb{R}^d}\psi_n\;d\mu_T = \int_{\mathbb{R}^d}(\psi_n)^+\;d\mu_T-\int_{\mathbb{R}^d}\psi^-\;d\mu_T = -\infty \;for all n,\) so \(\mu_T(\psi_n)\uparrow \mu_T(\psi)=-\infty\). If instead \(\int_{\mathbb{R}^d}\psi^-\;d\mu_T<\infty\), then \(\psi^-\in \mathbb{L}^1(\mu_T)\), and we obtain \(\mu_T(\psi_n)\uparrow\mu_T(\psi)\) by monotone convergence.

(d2) Weak duality on the enlarged class. Fix \(\mathbb{P}\in{\cal P}(\mu_0,\mu_T)\). If \(\mathbb{E}^\mathbb{P}\int_0^T c(\gamma_s^\mathbb{P})\;ds=+\infty\), then there is nothing to prove. We may thus assume that \(\mathbb{E}^\mathbb{P}\int_0^T c(\gamma_s^\mathbb{P})\;ds<\infty\). Let \(\{\mathbb{P}_x\}_{x\in\mathbb{R}^d}\) be a regular conditional law of \(\mathbb{P}\) given \(X_0\). Then \(\mathbb{P}_x\in{\cal P}(\delta_x)\) for \(\mu_0\)–a.e.\(x\), and by definition \(V_0^{\psi_n}(x) \ge \mathbb{E}^{\mathbb{P}_x}\!\left[ \psi_n(X_T)-\int_0^T c(\gamma_s^{\mathbb{P}_x})\,\mathrm{d}s \right] = -J^{\psi_n}(\mathbb{P}_x)\). Integrating with respect to \(\mu_0\) and using disintegration, \(\mu_T(\psi_n)-\mu_0(V_0^{\psi_n}) \le \mathbb{E}^\mathbb{P}\int_0^T c(\gamma_s^\mathbb{P})\;ds.\) Taking the infimum over \(\mathbb{P}\in{\cal P}(\mu_0,\mu_T)\) and using Proposition 4, we obtain \[\label{eq:weak-dual-n-inf-stepd} \mu_T(\psi_n)-\mu_0(V_0^{\psi_n}) \le {\rm SBB}(\mu_0,\mu_T) = {\cal V}(\mu_0,\mu_T), \qquad n\ge 1.\tag{20}\]

(d3) Sending \(n\to\infty\). Since \(\psi_n\uparrow\psi\), monotonicity in the terminal payoff gives \(V_0^{\psi_n}\uparrow V_0^\psi\) pointwise, hence \(V_0^{\psi_n}\le V_0^\psi\). Moreover, by Step (c) with \(t=0\), there exists \(C<\infty\) such that \(V_0^{\psi_n}(x)\ge -C(1+|x|^2), \; n\ge 1,\) so that \(\mu_0(V_0^{\psi_n})\) and \(\mu_0(V_0^\psi)\) are well-defined in \((-\infty,+\infty]\) and \(\mu_0(V_0^{\psi_n})\le \mu_0(V_0^\psi)\). Thus from 20 , \(\mu_T(\psi_n)-\mu_0(V_0^\psi) \le \mu_T(\psi_n)-\mu_0(V_0^{\psi_n}) \le {\cal V}(\mu_0,\mu_T).\) Letting \(n\to\infty\) and using \(\mu_T(\psi_n)\uparrow \mu_T(\psi)\) from (d1), we obtain \[\label{eq:weak-dual-extended-stepd} \mu_T(\psi)-\mu_0(V_0^\psi)\le {\cal V}(\mu_0,\mu_T), \qquad \psi\in{\cal C}^{\rm conc}.\tag{21}\] This proves ?? .   \({\cal t}\)  \({\cal u}\)

5 Reduced dual verification↩︎

The first duality result \({\rm SBB}={\cal V}\) was established in the last section, and we now complement it by justifying equality with the reduced dual value \({\cal V}_{\rm red}\) as claimed in Theorem 3 (i). We start with the analysis of the regularity of the maps \(u^\phi\).

Lemma 4. Fix \(T,\beta>0\) and \(\psi\in\bar{\cal C}^{\text{conc}}\). Set \(\phi:=\mathbf{T}_\beta^-[\psi]\). Then:

  1. \(h^\phi_t:={\cal N}_{T-t}*e^\phi >0\) for all \(t\le T\), and \(h^\phi\in C^{1,3}(Q_T)\) satisfies \[\partial_t h^\phi+\tfrac12\Delta h^\phi=0~\text{on } Q_T.\]

  2. \(\tilde{u}^\phi:=\log{h^\phi}\!\in\! C^{1,3}(Q_T)\), with \(D^2\tilde{u}^\phi_t\!+\!\kappa(t)I\succeq 0,\) \(\kappa(t)\!:=\!\frac{\beta}{1+\beta(T-t)}\), \(t\!<\!T\), and \[\partial_t \tilde{u}^\phi+\tfrac12\big(\Delta \tilde{u}^\phi+|\nabla \tilde{u}^\phi|^2\big)=0~\text{on } Q_T.\]

  3. \(v^\phi:=\mathbf{T}_\beta^+[\tilde{u}^\phi]\in C^{1,2}(Q_T)\), and the minimum in \(\mathbf{T}_\beta^+[\tilde{u}^\phi_t](x)\) is uniquely attained at \(\mathscr{Y}_t(x)\), for all \((t,x)\in Q_T\). Moreover,
    (3-a) \(\nabla v^\phi=\beta({\rm id}-\mathscr{Y})\), \(\partial_t v^\phi=\partial_t \tilde{u}^\phi(\mathscr{Y})\), and \(D^2 v^\phi-\beta I=-\beta^2\big(\beta I+D^2\tilde{u}^\phi(\mathscr{Y})\big)^{-1}\). In particular, \(-\frac{1}{T-t}I\preceq D^2 v^\phi_t\prec \beta I\), for all \(t<T\).
    (3-b) \(v^\phi\) is a solution of the HJB equation 8 .
    (3-c) If, in addition, \(\psi\in{\cal C}_{w,\uparrow}^{\text{conc}}\), then \(v^\phi_t\to \psi\) locally uniformly as \(t\uparrow T\), and \(v^\phi\) has quadratic growth uniformly in \(t\), i.e. there exists a constant \(C<\infty\) such that \[|v^\phi_t(x)|\le C(1+|x|^2), \qquad (t,x)\in Q_T.\]

By definition of \(\bar{\cal C}^{\text{conc}}\), we have \(\mu_0\Big(\mathbf{T}_\beta^+\Big[u_T^{\mathbf{T}_\beta^-[\psi]}\Big] \Big)<\infty\). Since \(\phi=\mathbf{T}_\beta^-[\psi]\), it follows that \(\mu_0\big(\mathbf{T}_\beta^+[u_T^\phi]\big)<\infty\). Hence \(\mathbf{T}_\beta^+[u_T^\phi](x_*)<\infty\) for some \(x_*\in\mathbb{R}^d\), and by definition of \(\mathbf{T}_\beta^+\), we deduce that \(u_T^\phi(y_0)<\infty\) for some \(y_0\in\mathbb{R}^d\). Equivalently, \(({\cal N}_T*e^\phi)(y_0)<\infty\). We continue the proof in several steps.

We first show that \(h^\phi\) and \(\tilde{u}^\phi\) are finite on \(Q_T\).

Fix \(0<t<T\) and \(x\in\mathbb{R}^d\). Then \(\frac{{\cal N}_{T-t}(x-z)}{{\cal N}_T(y_0-z)} = \big(\frac{T}{T-t}\big)^{\frac{d}{2}} e^{q_x(z)}\) with \(q_x(z):= -\frac{|x-z|^2}{2(T-t)}+\frac{|y_0-z|^2}{2T}\). Notice that \(q_x(z)=-\big(\frac{1}{2(T-t)}-\frac{1}{2T}\big)|z|^2 +\big(\frac{x}{T-t}-\frac{y_0}{T}\big)\!\cdot z -\frac{|x|^2}{2(T-t)}+\frac{|y_0|^2}{2T}\) is a concave quadratic polynomial in \(z\) since \(t<T\). Hence it is bounded above on \(\mathbb{R}^d\), so there exists \(C_{t,x}<\infty\) such that \({\cal N}_{T-t}(x-z)\le C_{t,x}\,{\cal N}_T(y_0-z), \;z\in\mathbb{R}^d\). Multiplying by \(e^\phi(z)\) and integrating, we obtain \[\label{eq:positive-time-finite} 0<{\cal N}_{T-t}*e^\phi(x)\le C_{t,x}({\cal N}_T*e^\phi)(y_0)<\infty ~for all~(t,x)\in Q_T.\tag{22}\]

Fix a compact strip \(K=[t_0,\tau]\times B_R\subset (0,T)\times\mathbb{R}^d, \; 0<t_0<\tau<T\). Choose \(s_*\in(T-t_0,T)\). By 22 , \(({\cal N}_{s_*}*e^\phi)(y_0)<\infty\). For every multi-index \(\alpha\) and every \(m\in\{0,1\}\) with \(|\alpha|+2m\le3\), we have \(\partial_s^m D_x^\alpha {\cal N}_s(x-z)=P_{m,\alpha,s}(x-z)\,{\cal N}_s(x-z)\), where \(P_{m,\alpha,s}\) is a polynomial of degree at most \(3\), and its coefficients are bounded for \(s\in[T-\tau,T-t_0]\).

Since \(T-t\in[T-\tau,T-t_0]\) on \(K\), the same Gaussian comparison as in the previous Step 1 gives \(\sup_{(t,x)\in K}\big|\partial_s^m D_x^\alpha {\cal N}_{T-t}(x-z)\big| \le C_{K,m,\alpha}\,{\cal N}_{s_*}(y_0-z)\). Because \(\int_{\mathbb{R}^d}{\cal N}_{s_*}(y_0-z)e^\phi(z)\,dz<\infty\), differentiation under the integral sign is justified on \(K\). Therefore \(h^\phi\in C^{1,3}(Q_T)\), \(\partial_t h^\phi+\frac{1}{2}\Delta h^\phi=0\). Since \(h^\phi>0\), also \(\tilde{u}^\phi=\log h^\phi\in C^{1,3}(Q_T)\), \(\partial_t\tilde{u}^\phi+\frac{1}{2}\big(\Delta\tilde{u}^\phi+|\nabla\tilde{u}^\phi|^2\big)=0\).

Fix \(t\in(0,T)\) and set \(s:=T-t, \; \kappa(t):=\frac{\beta}{1+\beta s}\). Let \(G(y):=\phi(y)+\frac{\beta}{2}|y|^2\). Since \(\phi\) is \(\beta\)–convex, \(G\) is convex. We compute \(\tilde{u}^\phi_t(x) = \log\int_{\mathbb{R}^d}(2\pi s)^{-\frac{d}{2}}e^{\phi(y)-\frac{|x-y|^2}{2s}}\mathrm{d}y\). A direct completion of the square gives \(-\frac{\beta}{2}|y|^2-\frac{|x-y|^2}{2s}+\frac{\kappa(t)}{2}|x|^2 = -\frac{1+\beta s}{2s}\big|y-\frac{x}{1+\beta s}\big|^2\). Therefore \[\begin{align} \tilde{u}^\phi_t(x)+\frac{\kappa(t)}{2}|x|^2 &= -\frac{d}{2}\log(2\pi s) + \log\int_{\mathbb{R}^d} e^{G(y)-\frac{1+\beta s}{2s}|y-\frac{x}{1+\beta s}|^2}\mathrm{d}y. \\ &= -\frac{d}{2}\log(2\pi s) + \log\int_{\mathbb{R}^d} e^{G(z+\frac{x}{1+\beta s})-\frac{1+\beta s}{2s}|z|^2}\mathrm{d}z, \end{align}\] by the change of variables \(z=y-\frac{x}{1+\beta s}\). For each fixed \(z\), the map \(x\longmapsto G\big(z+\frac{x}{1+\beta s}\big)-\frac{1+\beta s}{2s}|z|^2\) is convex. By Hölder’s inequality, the logarithm of the integral of its exponential is convex. Hence \(x\longmapsto \tilde{u}^\phi(t,x)+\frac{\kappa(t)}{2}|x|^2\) is convex, and therefore \[\label{eq:semiconvex-u} D^2\tilde{u}^\phi_t+\kappa(t)I\succeq0.\tag{23}\]

For \((t,x,y)\in(0,T)\times\mathbb{R}^d\times\mathbb{R}^d\), define \(F(t,x,y):=\tilde{u}^\phi_t(y)+\frac{\beta}{2}|x-y|^2\). By 23 , \(D_y^2F(t,x,y)=D^2\tilde{u}^\phi_t(y)+\beta I \succeq (\beta-\kappa(t))I\). Since \(\beta-\kappa(t)=\frac{\beta^2(T-t)}{1+\beta(T-t)}>0\), the map \(y\mapsto F(t,x,y)\) is strongly convex for each fixed \((t,x)\).

To see it attains its minimum, fix \((t,x)\). By Taylor’s formula at \(y=0\), \[F(t,x,y) \ge F(t,x,0)+\nabla_yF(t,x,0)\cdot y+\frac{\beta-\kappa(t)}{2}|y|^2.\] Hence \(F(t,x,y)\ge \frac{\beta-\kappa(t)}{2}|y|^2-C_{t,x}|y|-C_{t,x}\), which tends to \(+\infty\) as \(|y|\to\infty\). Thus \(F(t,x,\cdot)\) is coercive, so it has a unique minimizer \(\mathscr{Y}_t(x)\), and \(v^\phi_t(x):=F\bigl(t,x,\mathscr{Y}_t(x)\bigr)\).

We next prove that \(\mathscr{Y}\) is continuous on \(Q_T\). Let \((t_n,x_n)\to(t,x)\) and set \(y_n:=\mathscr{Y}(t_n,x_n)\). For all large \(n\), we have \(t_n\in[t/2,(t+T)/2]\). On this compact time interval, \(c_*:=\inf_{r\in[t/2,(t+T)/2]}(\beta-\kappa(r))>0\). Using strong convexity at \(y=0\), \[F(t_n,x_n,y)\ge F(t_n,x_n,0)+\nabla_yF(t_n,x_n,0)\cdot y+\frac{c_*}{2}|y|^2.\] As \((t_n,x_n)\) stays in a compact set, both \(F(t_n,x_n,0)\) and \(\nabla_yF(t_n,x_n,0)\) are bounded uniformly in \(n\), so \(F(t_n,x_n,y)\ge \frac{c_*}{2}|y|^2-C|y|-C\). Because \(y_n\) minimizes \(F(t_n,x_n,\cdot)\), \(F(t_n,x_n,y_n)\le F(t_n,x_n,0)\), and the right-hand side is uniformly bounded. Hence \((y_n)\) is bounded. Passing to a convergent subsequence if needed, \(y_n\to y_*\), and continuity of \(F\) gives \(F(t,x,y_*)=\lim_n F(t_n,x_n,y_n)\le \lim_n F(t_n,x_n,y)=F(t,x,y)\), for all \(y\in\mathbb{R}^d\). Thus \(y_*\) minimizes \(F(t,x,\cdot)\), and by uniqueness \(y_*=\mathscr{Y}_t(x)\). Hence \(\mathscr{Y}\) is continuous.

For fixed \(t\), the minimizer satisfies the first-order condition \[\label{eq:first-order-Y} \nabla_yF(t,x,\mathscr{Y}(t,x))=G_t(x,\mathscr{Y}_t(x))=0, ~with~ G_t(x,y):=\nabla\tilde{u}^\phi_t(y)-\beta(x-y).\tag{24}\] Notice that \(G_t\) is \(C^1\) with \(D_yG_t(x,y)=D^2\tilde{u}^\phi_t(y)+\beta I_d\), which is invertible by 23 . Then \(\mathscr{Y}_t\) is \(C^1\) by the implicit functions theorem. Differentiating 24 in \(x\), we obtain \[\label{eq:DxY} \nabla\mathscr{Y}_t(x)=\beta\bigl(D^2\tilde{u}^\phi_t(\mathscr{Y}_t(x))+\beta I_d\bigr)^{-1}.\tag{25}\] We now derive the envelope identities. Since \(v^\phi_t(x)=\tilde{u}^\phi_t(\mathscr{Y}_t(x))+\frac{\beta}{2}|x-\mathscr{Y}_t(x)|^2\), the chain rule and 24 25 provide \[\begin{align} \nabla v^\phi_t(x)&=\beta(x-\mathscr{Y}(t,x))=\nabla\tilde{u}^\phi_t(\mathscr{Y}_t(x)) \tag{26} \\ D^2v^\phi_t(x) &= \beta I_d-\beta^2\bigl(D^2\tilde{u}^\phi_t(\mathscr{Y}(t,x))+\beta I_d\bigr)^{-1}, \tag{27} \end{align}\] and \(D^2v^\phi\) inherits the continuity of \(\mathscr{Y}\).

It remains to compute \(\partial_t v^\phi\). We do not differentiate \(\mathscr{Y}\) in time. Fix \((t,x)\) and let \(h>0\). Since \(\mathscr{Y}_t(x)\) is admissible in the minimization defining \(v^\phi_{t+h}(x)\), \(v^\phi_{t+h}(x)\le F\bigl(t+h,x,\mathscr{Y}_t(x)\bigr)\), hence \[\frac{v^\phi_{t+h}(x)-v^\phi_t(x)}{h} \le \frac{F(t+h,x,\mathscr{Y}_t(x))-F(t,x,\mathscr{Y}_t(x))}{h} = \frac{\tilde{u}^\phi_{t+h}(\mathscr{Y}_t(x))-\tilde{u}^\phi_t(\mathscr{Y}_t(x))}{h},\] and then \(\limsup_{h\downarrow0}\frac{v^\phi_{t+h}(x)-v^\phi_t(x)}{h} \le \partial_t\tilde{u}^\phi_t(\mathscr{Y}_t(x))\). For the reverse inequality, since \(\mathscr{Y}_{t+h}(x)\) is admissible in the minimization defining \(v^\phi_t(x)\), \(v^\phi_t(x)\le F\bigl(t,x,\mathscr{Y}_{t+h}(x)\bigr)\), so \[\frac{v^\phi_{t+h}(x)\!-\!v^\phi_t(x)}{h} \ge \frac{F(t\!+\!h,x,\mathscr{Y}_{t+h}(x))\!-\!F(t,x,\mathscr{Y}_{t+h}(x))}{h} = \frac{\tilde{u}^\phi_{t+h}(\mathscr{Y}_{t+h}(x))\!-\!\tilde{u}^\phi_t(\mathscr{Y}_{t+h}(x))}{h}.\] Because \(\mathscr{Y}\) is continuous and \(\partial_t\tilde{u}^\phi\) is continuous, \(\mathscr{Y}_{t+h}(x)\to \mathscr{Y}_t(x)\) and \(\partial_t\tilde{u}^\phi_t(\mathscr{Y}_{t+h}(x))\to \partial_t\tilde{u}^\phi_t(\mathscr{Y}_t(x))\). Therefore \(\liminf_{h\downarrow0}\frac{v^\phi_{t+h}(x)-v^\phi_t(x)}{h} \ge \partial_t\tilde{u}^\phi_t(\mathscr{Y}_t(x))\). Hence \[\label{eq:dt-v} \partial_t v^\phi_t(x)=\partial_t\tilde{u}^\phi_t(\mathscr{Y}_t(x)).\tag{28}\] Since \(\mathscr{Y}\) and \(\partial_t\tilde{u}^\phi\) are continuous, \(\partial_t v^\phi\) is continuous. Thus \(v^\phi\in C^{1,2}(Q_T)\).

Let \(\lambda\) be an eigenvalue of \(D^2\tilde{u}^\phi_t(\mathscr{Y}_t(x))\), and let \(\ell\) be the corresponding eigenvalue of \(D^2v^\phi_t(x)\). By 27 , \(\ell = \beta-\frac{\beta^2}{\beta+\lambda} = \frac{\beta\lambda}{\beta+\lambda}\). Since 23 gives \(\lambda\ge -\kappa(t)\), and \(\lambda\mapsto \beta\lambda/(\beta+\lambda)\) is increasing on \((-\beta,\infty)\), \(-\frac{\beta\kappa(t)}{\beta-\kappa(t)}\le \ell<\beta\). Using \(\frac{\beta\kappa(t)}{\beta-\kappa(t)}=\frac{1}{T-t}\), we obtain \(-\frac{1}{T-t}I\preceq D^2v^\phi(t,\cdot)\prec \beta I_d\).

Recall that the Hamiltonian for \(A<\beta I_d\), \(H(p,A) = \frac{1}{2}|p|^2+\frac{\beta}{2}(\text{tr}(I-\frac{1}{\beta} A)^{-1}-d)\). With \(p=\nabla v^\phi_t(x)\) and \(A=D^2v^\phi_t(x)\), and using 27 , \[\Big(I-\frac{1}{\beta} D^2v^\phi_t(x)\Big)^{-1} = I+\frac{1}{\beta} D^2\tilde{u}^\phi_t(\mathscr{Y}_t(x)).\] Therefore \(H(\nabla v^\phi,D^2v^\phi)(t,x) = \frac{1}{2}|\nabla\tilde{u}^\phi_t(\mathscr{Y}_t(x))|^2+\frac{1}{2}\Delta\tilde{u}^\phi_t(\mathscr{Y}_t(x))\). Using 26 , 28 , and the Cole–Hopf equation for \(\tilde{u}^\phi\), we obtain \(\partial_t v^\phi+H(\nabla v^\phi,D^2v^\phi)=0 \;\text{on }Q_T\).

For \(\psi\in{\cal C}_{w,\uparrow}^{\text{conc}}\), we have \(-Cw\le\psi\le C\) for some constant \(C\), implying that \(\phi\le C\) and \(\phi(y) \ge \sup_{x\in\mathbb{R}^d}\left\{-C(1+|x|^2)-\frac{\beta}{2}|x-y|^2\right\} \ge -C'(1+|y|^2)\) for some constant \(C'<\infty\). As \(\phi\) is bounded from above, we deduce that \(h^\phi_t\to e^\phi\) locally uniformly as \(t\uparrow T\). The local uniform convergence \(\tilde{u}^\phi(t,\cdot)\to \phi\) is inherited from that of \(h^\phi_t\) towards \(e^\phi\). By the stability of the Moreau envelop under local uniform convergence, we see that \(v^\phi_t=\mathbf{T}_\beta^+[\tilde{u}^\phi_t]\to \mathbf{T}_\beta^+[\phi]=\psi\) locally uniformly.

We next justify the claimed growth for \(v^\phi\). First, \(v^\phi_t\le\tilde{u}^\phi_t\le \sup\phi\). Moreover, by the \(\beta-\)convexity of \(\phi\), for each \(y\) there exists \(\xi\in\partial(\phi+\frac{\beta}{2}|\cdot|^2)(y)\) such that \[\phi(z)\ge \phi(y)+(\xi-\beta y)\cdot(z-y)-\frac{\beta}{2}|z-y|^2, ~for all~z\in\mathbb{R}^d.\] Using this inequality with \(s=T-t\) and \(z=y+W_s\), \(W_s\sim {\cal N}(0,sI)\), we get \[\begin{align} h^\phi_t(y) =\mathbb{E}\big[e^{\phi(y+W_s)}\big] &\ge e^{\phi(y)}\;\mathbb{E}\big[e^{(\xi-\beta y)\cdot W_s-\frac{\beta}{2}|W_s|^2}\big] \\ &=e^{\phi(y)}\;(1+\beta s)^{-d/2}e^{\frac{s}{2(1+\beta s)}|\xi-\beta y|^2} \;\ge\;e^{\phi(y)}\;(1+\beta T)^{-d/2}. \end{align}\] Then \(\tilde{u}^\phi_t(y)=\log h^\phi_t(y) \ge \phi(y)-\frac{d}{2}\log(1+\beta T)=:\phi(y)-C_T\), and \[v^\phi_t(x)=\inf_y\Big\{\tilde{u}^\phi_t(y)+\frac{\beta}{2}|x-y|^2\Big\} \ge \inf_y\Big\{\phi(y)+\frac{\beta}{2}|x-y|^2\Big\}-C_T =\psi(x)-C_T.\] Since \(\psi\in{\cal C}_{w,\uparrow}^{\text{conc}}\), this proves the claimed quadratic growth of \(v^\phi\).   \({\cal t}\)  \({\cal u}\)

We now have all ingredients to identify the maps \(V^{\psi}\) and \(v^\phi\) by means of a verification argument. We first note that \[\label{eq:Cwuparrow-in-barC} {\cal C}_{w,\uparrow}^{\text{conc}}\subset \bar{\cal C}^{\text{conc}}.\tag{29}\]

Lemma 5. Let \(\psi\in{\cal C}^{\rm conc}_{w,\uparrow}\) and set \(\phi:=\mathbf{T}_\beta^-[\psi]\). Then \(V^{\psi}=v^\phi\) on \(Q_T\).

In the rest of this proof, we denote \(t_m:=T-\frac{1}{m}\), \(m\in\mathbb{N}\).

Let \(\mathbb{P}\in{\cal P}_t(\delta_x)\) be such that \(\mathbb{E}^\mathbb{P}\int_t^T c(\gamma_s^\mathbb{P})\,\mathrm{d}s<\infty\), and recall from Lemma 1 that \(\mathbb{E}^\mathbb{P}[\sup_{r\in[t,T]}|X_r|^2]<\infty\). By the \(C^{1,2}\) regularity of \(v:=v^\phi\) and its quadratic growth established in Lemma 4 (3), and since the stochastic integral in Itô’s formula is a true martingale, we obtain \[\mathbb{E}^\mathbb{P}[v_{t_m}(\!X_{t_m})]-v_t(x) \!=\! \mathbb{E}^\mathbb{P}\!\Big[\int_t^{t_m}\!\!\!\!\big(\partial_sv_s+\alpha_s^\mathbb{P}\cdot\nabla v_s+\frac{1}{2}\,\sigma_s^\mathbb{P}(\sigma_s^\mathbb{P})^\intercal\!\!:\!\!D^2v_s\big)(X_s)\mathrm{d}s\Big] \le \mathbb{E}^\mathbb{P}\!\Big[\int_t^{t_m}\!\!\!\!\!c(\gamma_s^\mathbb{P})\mathrm{d}s\Big],\] where the last inequality follows from the fact that \(v\) is a supersolution of the HJB equation 8 . Since \(X_{t_m}\to X_T\) \(\mathbb{P}\)-a.s., \(v_{t_m}\to\psi\) locally uniformly, as \(m\nearrow\infty\), and \(|v_{t_m}(X_{t_m})|\le C(1+\sup_{r\in[t,T]}|X_r|^2)\in\mathbb{L}^1\), we see that \(v_{t_m}(X_{t_m})\to\psi(X_T)\) in \(\mathbb{L}^1(\mathbb{P})\) by dominated convergence. Then, as the cost function \(c\ge0\), we may now deduce by monotone convergence that \[v_t(x) \ge \mathbb{E}^\mathbb{P}\Big[\psi(X_T)-\int_t^T c(\gamma_s^\mathbb{P})\,\mathrm{d}s\Big],\] and by arbitrariness of \(\mathbb{P}\in{\cal P}_t(\delta_x)\), we get \(v_t(x)\ge V_t^\psi(x)\).

We next prove the reverse inequality \(v_t(x)\le V_t^\psi(x)\). By Lemma 4 (3), the minimizer \(\mathscr{Y}_t(x)\) in \(v_t(x)={\mathbf{T}}_\beta^+[\tilde{u}_t](x)\) is unique. Set \(y:=\mathscr{Y}_t(x)\), so that \(x=\mathscr{X}_t(y)\). Let \(Y\) be a \(\mathbb{Q}\)-Brownian motion on \([t,t_m]\) with \(Y_t=y\), and notice that the upper boundedness of \(\phi\) implies that \(Z_s:=\frac{h_s^\phi(Y_s)}{h_t^\phi(y)}\) is a positive bounded martingale inducing an equivalent probability measure \(\mathbb{Q}^m\) on \({\cal F}_T\) by \(\mathrm{d}\mathbb{Q}^m:=Z_{t_m}\,\mathrm{d}\mathbb{Q}\). By Girsanov’s theorem the process \(W_s:=Y_s-\int_t^{s\wedge t_m}\nabla\log h_r^\phi(Y_r)\,\mathrm{d}r\) defines a \(\mathbb{Q}^m-\)Brownian motion on \([0,T]\), and we may then write the dynamics of \(Y\) as \[\label{eq:Y95sde} \mathrm{d}Y_s=\nabla\log h_s^\phi(Y_s)\mathbf{1}_{s\in [t,t_m]}\,\mathrm{d}s+\mathrm{d}W_s,\qquad Y_t=y.\tag{30}\] Denote \(X_s:=\mathscr{X}_s(Y_s)\mathbf{1}_{s\in [t,t_m]}+(X_{t_m}+\int_{t_m}^s \mathrm{d}W_s)\mathbf{1}_{s\in (t_m,T]}\), where we recall that \(\mathscr{X}_s:=id+\frac{1}{\beta}\nabla\tilde{u}_s\). By Lemma 4 (3-a), it follows that \[\label{eq:grad-v-XY} \nabla v_s(X_s)=\beta(X_s-Y_s)=\nabla\tilde{u}_s(Y_s)=\nabla\log h^\phi(s,Y_s), \qquad s\in[t,t_m].\tag{31}\] Set \(\widehat\alpha_s:=\nabla v_s, \; \widehat\sigma_s:=\Big( I_d-\frac{D^2v_s}{\beta}\Big)^{-1}\). Since \(\mathscr{Y}_s:=id-\frac{1}{\beta}\nabla v_s\) is the inverse of \(\mathscr{X}_s\), differentiating the identity \(\mathscr{Y}_s(\mathscr{X}_s(y))=y\) yields \(D\mathscr{X}_s(y)=\Big( I_d-\frac{D^2v_s(\mathscr{X}_s(y))}{\beta}\Big)^{-1}\), and therefore \(D\mathscr{X}_s(Y_s)=\widehat\sigma_s(X_s), \;s\in[t,t_m]\). By Lemma 4 (1)–(2), we have \(\tilde{u}^\phi\in C^{1,3}(Q_T)\), and therefore \((s,y)\longmapsto \mathscr{X}_s(y):=y+\frac{1}{\beta}\nabla\tilde{u}^\phi_s(y)\) is of class \(C^{1,2}\) on \(Q_T\). Applying Itô to \(X_s=X(s,Y_s)\) on \([t,t_m]\) using 31 , we obtain \[\begin{align} \mathrm{d}X_s &= \Big( \partial_s\mathscr{X}_s+D\mathscr{X}_s\,\nabla\tilde{u}_s+\frac{1}{2}\Delta\mathscr{X}_s \Big)(Y_s)\,\mathrm{d}s + D\mathscr{X}_s(Y_s)\,\mathrm{d}W_s \\ &= \Big( \nabla\tilde{u}_s+\frac{1}{\beta}\Big( \nabla\partial_s\tilde{u}_s+D^2\tilde{u}_s\,\nabla\tilde{u}_s+\frac{1}{2}\nabla\Delta\tilde{u}_s \Big) \Big)(Y_s)\,\mathrm{d}s + \widehat\sigma_s(X_s)\,\mathrm{d}W_s. \end{align}\] Since \(h^\phi\) solves the backward heat equation, differentiating the PDE of \(\tilde{u}=\log h^\phi\) gives \(\nabla\partial_s\tilde{u}_s+D^2\tilde{u}_s\,\nabla\tilde{u}_s+\frac{1}{2}\nabla\Delta\tilde{u}_s=0\). Hence, using also 31 , we get \[\begin{align} X_t=x:=\mathscr{X}_t(y),~ &dX_s=\widehat\alpha_s(X_s)\,ds+\widehat\sigma_s(X_s)\,dW_s,~ s\in[t,t_m], \\ &and~ X_s=X_{t_m}+W_s-W_{t_m}~for~s\in[t_m,T]. \end{align}\] Let \(\mathbb{P}^{t,x,m}\) be the law of \(X\) under \(\mathbb{Q}^m\). Then \(\mathbb{P}_{t,x}^{m}\in{\cal P}_t(\delta_x)\) with characteristics \(\gamma^{\mathbb{P}_{t,x}^{m}}_s=(\widehat\alpha,\widehat\sigma)(s,X_s)\mathbf{1}_{s\in[t,t_m]}+(0,I_d)\mathbf{1}_{s\in[t_m,T]}\). As \(c(\gamma^{\mathbb{P}_{t,x}^{m}})=0\) on \([t_m,T]\), it follows from the HJB equation 8 satisfied by \(v\) that \[\label{eq:exact95trunc} v_t(x) = \mathbb{E}^{\mathbb{P}_{t,x}^{m}}\Big[ v_{t_m}(X_{t_m})-\int_t^{t_m} c(\gamma^{\mathbb{P}_{t,x}^{m}}_s)\,\mathrm{d}s \Big].\tag{32}\] Let \(\delta_m:=\sqrt{T-t_m}\), let \(Z\) be a standard Gaussian random variable independent of \({\cal F}_{t_m}\), and define \(\psi_m(x):=\mathbb{E}[\psi(x+\delta_m Z)]\). Since \(c(\gamma^{\mathbb{P}_{t,x}^{m}})=0\) on \([t_m,T]\), we have \[\mathbb{E}^{\mathbb{P}_{t,x}^{m}}\Big[ \psi(X_T)-\int_t^T c(\gamma^{\mathbb{P}_{t,x}^{m}}_s)\,\mathrm{d}s \Big] = \mathbb{E}^{\mathbb{P}_{t,x}^{m}}\Big[ \psi_m(X_{t_m})-\int_t^{t_m} c(\gamma^{\mathbb{P}_{t,x}^{m}}_s)\,\mathrm{d}s \Big].\] Consequently, \(v_t(x) = \mathbb{E}^{\mathbb{P}_{t,x}^{m}}\Big[ \psi(X_T)-\int_t^T c(\gamma^{\mathbb{P}_{t,x}^{m}}_s)\,\mathrm{d}s \Big] +\varepsilon_m \le V_t^\psi(x)+\varepsilon_m\), where \(\varepsilon_m := \mathbb{E}^{\mathbb{P}_{t,x}^{m}}\big[ v_{t_m}(X_{t_m})-\psi_m(X_{t_m}) \big]\). Since \((v_{t_m},\psi_m)\to(\psi,\psi)\) locally uniformly as \(m\nearrow\infty\) and both \(v_{t_m}\) and \(|\psi_m|\) have quadratic growth uniformly in \(m\), it remains to justify the uniform square-integrability of \((X_{t_m})_{m\ge1}\) under \((\mathbb{P}_{t,x}^{m})_{m\ge1}\).

We first record a uniform linear-growth bound. Since \(\psi\in{\cal C}^{\rm conc}_{w,\uparrow}\), the function \(\phi=\mathbf{T}_\beta^-[\psi]\) is bounded from above and satisfies a quadratic lower bound. Hence, by the same convexity argument as in Lemma 4, there exists a constant \(C_t<\infty\) such that \(|\nabla u_r^\phi(z)|\le C_t(1+|z|), \; 0<r\le T-t,\;z\in\mathbb{R}^d\). Equivalently, \[|\nabla \tilde{u}_s^\phi(z)|\le C_t(1+|z|), \qquad t\le s<T,\;z\in\mathbb{R}^d .\] Under \(\mathbb{Q}^m\), the process \(Y\) satisfies \(\mathrm{d}Y_s=\nabla\tilde{u}_s^\phi(Y_s)\,\mathrm{d}s+\mathrm{d}W_s,\; s\in[t,t_m]\), and therefore the preceding linear-growth estimate, together with the BDG inequality and Gronwall’s lemma, gives \(\sup_{m\ge1}\mathbb{E}^{\mathbb{Q}^m}\Big[\sup_{t\le s\le t_m}|Y_s|^2\Big]<\infty\). Since \(X_{t_m} = Y_{t_m}+\frac{1}{\beta}\nabla\tilde{u}_{t_m}^\phi(Y_{t_m})\), the same linear-growth bound yields \(\sup_{m\ge1}\mathbb{E}^{\mathbb{P}_{t,x}^{m}}\big[|X_{t_m}|^2\big]<\infty\). Thus, for every \(R>0\), \[\sup_{|z|\le R}|v_{t_m}(z)-\psi_m(z)|\longrightarrow0,\] while the tails are controlled uniformly by the quadratic growth and the above uniform square-integrability. Hence \(\varepsilon_m\to0\), inducing the required inequality \(v_t(x)\le V_t^\psi(x)\).   \({\cal t}\)  \({\cal u}\)

We also record that the relaxed dual has the same value. Indeed, by 29 , we have \({\cal C}_{w,\uparrow}^{\text{conc}}\subset \bar{\cal C}^{\text{conc}}\). Hence \[\bar{\cal V}(\mu_0,\mu_T) \ge \sup_{\psi\in{\cal C}_{w,\uparrow}^{\text{conc}}} \{\mu_T(\psi)-\mu_0(V_0^\psi)\} = {\cal V}(\mu_0,\mu_T).\] On the other hand, since \(\bar{\cal C}^{\text{conc}}\subset{\cal C}^{\text{conc}}\), Lemma 3 gives \(\bar{\cal V}(\mu_0,\mu_T)\le {\cal V}(\mu_0,\mu_T)\). Therefore \[\label{eq:barV-equals-V} \bar{\cal V}(\mu_0,\mu_T)={\cal V}(\mu_0,\mu_T).\tag{33}\]

By Moreau biconjugation \(\{\phi=\mathbf{T}_\beta^-[\psi]:\;for \psi\in{\cal C}^{\rm conc}_{w,\uparrow}\} = {\cal C}^{\rm conv}_{w,\uparrow}\). Combining Proposition 4, Lemma 3, Lemma 5, and 33 , we obtain \[\label{dualuparrow} {\rm SBB}(\mu_0,\mu_T) = {\cal V}(\mu_0,\mu_T) = \bar{\cal V}(\mu_0,\mu_T) = \sup_{\phi\in{\cal C}^{\rm conv}_{w,\uparrow}} \left\{ \mu_T(\mathbf{T}_\beta^+[\phi])-\mu_0(v_0^\phi) \right\}.\tag{34}\] It remains to prove that the value of the supremum on the right-hand side is not affected by enlarging the set of dual potential maps from \({\cal C}^{\rm conv}_{w,\uparrow}\) to \({\cal C}^{\rm conv}_w\). To see this, fix \(\phi\in{\cal C}^{\rm conv}_w\) and set \(\psi:=\mathbf{T}_\beta^+[\phi]\). For \(n\ge1\), define \[\chi_n:=\psi\wedge n, \qquad \phi_n:=\mathbf{T}_\beta^-[\chi_n].\] Since \(\chi_n\in{\cal C}_w\) is bounded from above, we have \(\phi_n\in{\cal C}^{\rm conv}_{w,\uparrow}\). Moreover, \(\mathbf{T}_\beta^+[\phi_n]=\mathbf{T}_\beta^+\mathbf{T}_\beta^-[\chi_n]\) is the \(\beta\)–concave envelope of \(\chi_n\). Hence \(\chi_n \le \mathbf{T}_\beta^+[\phi_n] \le \psi\), where the last inequality follows from the fact that \(\psi\) is itself a \(\beta\)–concave majorant of \(\chi_n\). Since \(\chi_n\uparrow\psi\) and \(\psi\in{\cal C}_w\), we obtain \(\mu_T\big(\mathbf{T}_\beta^+[\phi_n]\big) \longrightarrow \mu_T(\psi)\). On the other hand, \(\chi_n\le\psi\) implies \(\phi_n = \mathbf{T}_\beta^-[\chi_n] \le \mathbf{T}_\beta^-[\psi] = \phi\). Therefore \(u_T^{\phi_n}\le u_T^\phi, \;\mathbf{T}_\beta^+[u_T^{\phi_n}] \le \mathbf{T}_\beta^+[u_T^\phi]\), and consequently \(\mathfrak J(\phi_n) = \mu_T\big(\mathbf{T}_\beta^+[\phi_n]\big) - \mu_0\big(\mathbf{T}_\beta^+[u_T^{\phi_n}]\big) \ge \mu_T\big(\mathbf{T}_\beta^+[\phi_n]\big) - \mu_0\big(\mathbf{T}_\beta^+[u_T^\phi]\big)\). Letting \(n\to\infty\) yields \(\liminf_{n\to\infty}\mathfrak J(\phi_n) \ge \mu_T(\psi) - \mu_0\big(\mathbf{T}_\beta^+[u_T^\phi]\big) = \mathfrak J(\phi)\). This proves that the supremum over \({\cal C}^{\rm conv}_{w,\uparrow}\) is equal to the supremum over \({\cal C}^{\rm conv}_w\).   \({\cal t}\)  \({\cal u}\)

6 Dual attainment↩︎

Our objective in this section is to prove Theorem 3 (ii-a), namely the existence of potential \(\hat{\psi}\) attaining the dual value \({\cal V}(\mu_0,\mu_T)\), and a dual potential \(\hat{\phi}\) attaining the reduced dual value \({\cal V}_{\rm red}(\mu_0,\mu_T)\). We shall denote throughout: \[\label{Udeltaphi} U^\phi_\eta := u_{\eta T}^{\phi}(0) = \log\mathbb{E}[e^{\phi(W_{\eta T})}], ~\text{for some fixed}~\eta\in(0,1).\tag{35}\]

Lemma 6. For every \(\phi\in{\cal C}^{\rm conv}_{w,\uparrow}\) with \(\phi(0)=0\), we have:

  1. There is a constant \(\nu_\eta\) (depending on \(\beta,T,d,\eta\)) such that \[-\nu_\eta\le U^\phi_\eta \le -\frac{d}{2}\log(\eta)+\inf_{x\in\mathbb{R}^d} u^\phi_T(x)+\frac{|x|^2}{2(1-\eta)T}.\]

  2. There exists a constant \(\nu'_\eta\) (depending on \(\beta,T,d,\eta\)), such that for all \(x\in\mathbb{R}^d\): \[\begin{align} u_s^\phi(x) & \ge -\frac{\beta|x|^2}{2} +\mathbf{1}_{\{s>0\}}\Big(\frac{\beta^2|x|^2}{4s}-\frac{(\nu'_\eta\!+\!U^\phi_\eta)^2}{\beta^2s}\Big) \!-\!\mathbf{1}_{\{s=0\}}(\nu'_\eta\!+\!U^\phi_\eta)|x|, ~s\!\in\![0,T],\label{eq:u95low95bound} \\u^\phi_s(x) &\le U^\phi_\eta+\nu_\eta -\frac{d}{2}\log\Big(1-\frac{s}{\eta T}\Big) +\frac{|x|^2}{2(\eta T-s)}, ~s\in[0,\eta T). \label{eq:u95up95bound} \end{align}\] {#eq: sublabel=eq:eq:u95low95bound,eq:eq:u95up95bound}

Proof. (i) Using the Young inequality \(|x-y|^2\le\frac{|x|^2}{1-\eta}+\frac{|y|^2}{\eta}\), we see that \[\begin{align} e^{u_T^\phi(x)} = \int \frac{e^{\phi(y)-\frac{1}{2T}|x-y|^2}}{(2\pi T)^{\frac{d}{2}}}\,\mathrm{d}y \ge \eta^{\frac{d}{2}}e^{-\frac{|x|^2}{2(1-\eta)T}} \int \frac{e^{\phi(y)-\frac{|y|^2}{2\eta T}}}{(2\pi \eta T)^{\frac{d}{2}}}\,\mathrm{d}y = \eta^{\frac{d}{2}}e^{-\frac{|x|^2}{2(1-\eta)T}}e^{U^\phi_\eta}, \end{align}\] which provides the right-hand side inequality in (i).

Next, by the \(\beta\)–convexity of \(\phi\), we obtain the Jensen inequality for all \(R>0\): \[\phi(x) \le -\frac{1}{2}\beta|x|^2 +\frac{1}{|B_R|}\int_{B_R(x)}\Big(\phi+\frac{1}{2}\beta|.|^2\Big)(y)\,\mathrm{d}y \le \frac{1}{2}\beta R^2+\frac{1}{|B_R|}\int_{B_R(x)}\phi(y)\,\mathrm{d}y.\] Applying Jensen’s inequality again to the exponential function, we deduce that \[\begin{align} \phi(x) &\le \frac{1}{2}\beta R^2 +\log\Big(\frac{1}{|B_R|}\int_{B_R(x)}e^{\phi(y)}\,\mathrm{d}y\Big) \nonumber\\ &\le \frac{1}{2}\beta R^2 +\log\Big(\frac{1}{|B_R|}e^{\frac{|x|^2+R^2}{2\eta T}} \int_{B_R(x)}e^{\phi(y)-\frac{|y|^2}{2\eta T}}\,\mathrm{d}y\Big) \nonumber\\ &\le \frac{1}{2}\beta R^2 +\log\Big(\frac{(2\pi\eta T)^{\frac{d}{2}}}{|B_R|} e^{\frac{|x|^2+R^2}{2\eta T}}e^{U^\phi_\eta}\Big) \nonumber\\ &= U^\phi_\eta+\frac{|x|^2}{2\eta T}+\bar\nu_\eta(R), ~\text{with}~ \bar\nu_\eta(R) := \frac{1}{2}\Big(\beta+\frac{1}{\eta T}\Big)R^2 +\log\frac{(2\pi\eta T)^{\frac{d}{2}}}{|B_R|}. \label{phileq1} \end{align}\tag{36}\] As \(\phi(0)=0\), this provides the left-hand side inequality in (i) with the finite constant \[\nu_\eta:=\inf_{R>0}\bar\nu_\eta(R).\]

(ii) Notice that 36 provides the required upper bound in ?? at \(s=0\). To obtain the lower bound in ?? at \(s=0\), note that the \(\beta\)–convexity of \(\phi\) implies that \[\label{eq:phi95lb} \phi(x)+\frac{1}{2}\beta|x|^2\ge p\cdot x ~\text{for all }p\in\partial\phi(0),\tag{37}\] as \(\phi(0)=0\). Therefore \(|p|=p\cdot\frac{p}{|p|} \le \frac{\beta}{2}+\max_{B_1}\phi \le \frac{\beta}{2}+\frac{1}{2\eta T}+\nu_\eta+U^\phi_\eta =:\nu'_\eta+U^\phi_\eta,\) by substituting \(R=1\) in 36 . Then \(\phi(x)+\frac{1}{2}\beta|x|^2\ge p\cdot x\ge-(\nu'_\eta+U^\phi_\eta)|x|,\) which is exactly the case \(s=0\) in ?? , since \(u^\phi_0=\phi\).

By 37 , the definition of \(u_s^\phi\), and a direct Gaussian calculation, \[\begin{align} e^{u_s^\phi(x)} = \int_{\mathbb{R}^d}\frac{e^{\phi(y)-\frac{|x-y|^2}{2s}}}{(2\pi s)^{\frac{d}{2}}}\,\mathrm{d}y \ge \int_{\mathbb{R}^d}\frac{e^{p\cdot y-\frac{\beta}{2}|y|^2-\frac{|x-y|^2}{2s}}}{(2\pi s)^{\frac{d}{2}}}\,\mathrm{d}y = \frac{e^{\frac{|\frac{x}{s}+p|^2}{2(\beta+\frac{1}{s})}-\frac{|x|^2}{2s}}}{(1+\beta s)^{\frac{d}{2}}}. \end{align}\] Using the Young inequality \(\frac{p\cdot x}{1+\beta s} \ge -\frac{\lambda}{2}|x|^2-\frac{1}{2\lambda}\frac{|p|^2}{(1+\beta s)^2},\) \(\lambda:=\frac{\beta^2 s}{2(1+\beta s)}>0,\) this implies \[u_s^\phi(x) \ge -\frac{d}{2}\log(1+\beta s) -\frac{\beta}{4}(2-\beta s)|x|^2 -\frac{|p|^2}{(1+\beta s)\beta^2s},\] which gives the case \(s>0\) in ?? , after increasing \(\nu'_\eta\) if necessary so that \(\nu'_\eta+U^\phi_\eta\ge0\), by using \(|p|^2\le(\nu'_\eta+U^\phi_\eta)^2\). We finally derive ?? by using the right-hand side of 36 : \[\begin{align} u^\phi_s(x) = \log \mathbb{E}\big[e^{\phi(x+W_s)}\big] &\le U^\phi_\eta+\nu_\eta +\log \mathbb{E}\Big[e^{\frac{1}{2\eta T}|x+W_s|^2}\Big] \\ &= U^\phi_\eta+\nu_\eta -\frac{d}{2}\log\Big(1-\frac{s}{\eta T}\Big) +\frac{|x|^2}{2(\eta T-s)}, \qquad s\in[0,\eta T). \end{align}\] ◻

In the rest of this paper, \[\label{assum:existence} we assume~ \delta:=1-\frac{1}{\beta T}>0 ~and we fix~ \eta\in(0,\delta).\qquad{(2)}\] By 34 , we may consider a maximizing sequence \((\phi_n)_{n\ge 1}\) for Problem 10 satisfying: \[\label{maxsequence} \phi_n\in{\cal C}^{\rm conv}_{w,\uparrow}, \phi_n(0)=0, ~and~ \mathfrak{J}(\phi_n) := \mu_T\big(\mathbf{T}^+_\beta[\phi_n]\big) - \mu_0\big(\mathbf{T}^+_\beta[u^{\phi_n}_T]\big) \stackrel{n\to\infty}{\longrightarrow} {\cal V}_{\rm red}.\tag{38}\] As \(\mathfrak{J}(\phi+{\rm const})=\mathfrak{J}(\phi)\), we may set \(\phi_n(0)=0\) for all \(n\ge 1\). We proceed in several steps.

We first prove that the sequence \[\label{Udeltabdd} \big(U^{\phi_n}_\eta\big)_{n\ge 1}is bounded.\tag{39}\] First, \(U^{\phi_n}_\eta\ge-\nu_\eta\) by Lemma 6 (i). To derive a uniform upper bound, we observe that \(\mathbf{T}_\beta^+[\phi_n]\le \phi_n(0)+\frac{\beta}{2}|.|^2=\frac{\beta}{2}|.|^2,\) so that Lemma 6 (i) yields \(u_T^{\phi_n}(y)\ge U_\eta^{\phi_n}-\frac{|y|^2}{2(1-\eta)T}+\frac{d}{2}\log(\eta).\) Then \[\mathbf{T}_\beta^+\big[u_T^{\phi_n}\big](x) \ge U_\eta^{\phi_n}+\frac{d}{2}\log(\eta) + \mathbf{T}_\beta^+\Big[-\frac{|.|^2}{2(1-\eta)T}\Big](x).\] Since \(\eta<\delta\), we have \(\frac{1}{(1-\eta)T}<\beta\), and therefore \(\mathbf{T}_\beta^+\Big[-\frac{|.|^2}{2(1-\eta)T}\Big](x) = -\frac{\beta}{2\big(\beta T(1-\eta)-1\big)}\,|x|^2.\) It follows that, for all sufficiently large \(n\), \[\begin{align} {\cal V}_{\rm red}-1 \le J(\phi_n) &= \int \mathbf{T}_\beta^+[\phi_n](x)\,\mu_T(\mathrm{d}x) - \int \mathbf{T}_\beta^+\big[u_T^{\phi_n}\big](x)\,\mu_0(\mathrm{d}x) \\ &\le \frac{\beta}{2}\int |x|^2\,\mu_T(\mathrm{d}x) - U_\eta^{\phi_n} - \frac{d}{2}\log(\eta) + \frac{\beta}{2\big(\beta T(1-\eta)-1\big)} \int |x|^2\,\mu_0(\mathrm{d}x). \end{align}\] Since \(\mu_0,\mu_T\) have finite second moments and \({\cal V}_{\rm red}<\infty\), this proves 39 .

We next show that, after passing to a subsequence \((n_k)_k\), we may find a limit function: \[\label{phinconv} \phi_{n_k} \longrightarrow \hat{\phi} \in {\cal C}^{\rm conv}_w \quadas k\to\infty, locally uniformly on\mathbb{R}^d.\tag{40}\] Indeed, inequalities ?? –?? of Lemma 6 (ii), evaluated at \(s=0\), together with 39 , show that the sequence \((\phi_n)_n\) is locally bounded. As each \(\phi_n\) is \(\beta\)–convex and locally uniformly bounded, it follows that \((\phi_n)_n\) is locally Lipschitz, uniformly in \(n\); see Rockafellar [12], Theorem 10.6, p. 88. The convergence in 40 is then a direct consequence of the Arzelà–Ascoli theorem. Notice that \(\hat{\phi}\in{\cal C}^{\rm conv}_w\) as it inherits the \(\beta-\)convexity of the maps \(\phi_{n_k},\) together with the quadratic bounds in ?? –?? .

In this step, we provide the proof of existence of a dual potential optimiser for the reduced dual problem \({\cal V}_{\rm red}\). We rely on the following claim which will be justified later in Steps 5 to 7 below: \[\label{Tbeta43phinconv} \mathbf{T}_\beta^+[\phi_{n_k}] \longrightarrow \mathbf{T}_\beta^+[\hat{\phi}] \quadas k\to\infty, locally uniformly on\mathbb{R}^d.\tag{41}\]

  • Since \(\phi_n(0)=0\), we have \(\mathbf{T}_\beta^+[\phi_n]\le \frac{\beta}{2}|.|^2\in\mathbb{L}^1(\mu_T)\) by Assumption ?? . It then follows from 41 and Fatou’s lemma that \[\limsup_{k\to\infty}\mu_T\big( \mathbf{T}_\beta^+[\phi_{n_k}]\big) \le \mu_T\big(\mathbf{T}_\beta^+[\hat{\phi}]\big).\]

  • Using Lemma 6 (ii), ?? , with \(s=T\), and the boundedness of \(\big(U_\eta^{\phi_n}\big)_{n\ge 1}\) proved in Step 1 that \(u_T^{\phi_{n_k}}+\frac{\beta}{2}|.|^2\ge -c_0+\frac{\beta^2}{4T}|.|^2\) for some constant \(c_0\). Then, \[\begin{align} \mathbf{T}_\beta^+\big[u^{\phi_{n_k}}_T\big](x) &= \inf_y\Big\{u^{\phi_{n_k}}_T(y)+\frac{\beta}{2}|x-y|^2\Big\} \\ &\ge -c_0+\frac{\beta}{2}|x|^2+\inf_y\Big\{\frac{\beta^2}{4T}|y|^2-\beta x\cdot y\Big\} = -c_0+\Big(\frac{\beta}{2}-T\Big)|x|^2, \end{align}\] for all \(x\in\mathbb{R}^d\), where the last lower bound is in \(\mathbb{L}^1(\mu_0)\) as \(\mu_0\in{\cal P}_2(\mathbb{R}^d)\). We may then apply Fatou’s lemma to get \[\label{lsc1} \liminf_{k\to\infty}\mu_0\big(\mathbf{T}_\beta^+\big[u^{\phi_{n_k}}_T\big]\big) \ge \mu_0\big(L\big), ~~ L:=\liminf_{k\to\infty}\mathbf{T}_\beta^+\big[u^{\phi_{n_k}}_T\big].\tag{42}\] After passing to a subsequence \(n_{k'}=n_{k'}^x\) for fixed \(x\in\mathbb{R}^d\), we have \[\begin{align} L(x) = \lim_{k'\to\infty}\mathbf{T}_\beta^+\big[u^{\phi_{n_{k'}}}_T\big](x) &\ge \limsup_{k'\to\infty} \Big( \frac{\beta}{2}|x|^2 -\beta x\cdot y_{n_{k'}} -c_0 +\frac{\beta^2}{4T}|y_{n_{k'}}|^2 \Big), \end{align}\] with \(y_{n_{k'}}\) a minimizer of \(\mathbf{T}_\beta^+\big[u^{\phi_{n_{k'}}}_T\big](x)\). Then if \(L(x)<\infty\), this shows that the sequence \((y_{n_{k'}})_{k'}\) is bounded; after possibly passing to a further subsequence, we may therefore assume that \(y_{n_{k'}}\to\hat{y}\in\mathbb{R}^d\). It follows that \[\begin{align} L(x) &= \lim_{k'\to\infty} \Big( u^{\phi_{n_{k'}}}_T(y_{n_{k'}})+\frac{\beta}{2}|x-y_{n_{k'}}|^2 \Big) \\ &\ge \liminf_{k'\to\infty}u^{\phi_{n_{k'}}}_T(y_{n_{k'}}) +\frac{\beta}{2}|x-\hat{y}|^2 \\ &= \log\Big( \liminf_{k'\to\infty} \mathbb{E}\big[e^{\phi_{n_{k'}}(y_{n_{k'}}+W_T)}\big] \Big) +\frac{\beta}{2}|x-\hat{y}|^2 \\ &\ge u^{\hat{\phi}}_T(\hat{y})+\frac{\beta}{2}|x-\hat{y}|^2 \ge \inf_{y\in\mathbb{R}^d} \Big\{ u^{\hat{\phi}}_T(y)+\frac{\beta}{2}|x-y|^2 \Big\} = \mathbf{T}_\beta^+\big[u^{\hat{\phi}}_T\big](x), \end{align}\] by Fatou’s lemma and the local uniform convergence of \(\phi_{n_{k'}}\) to \(\hat{\phi}\), where \(u_T^{\hat{\phi}}(y)\in(-\infty,\infty]\) is understood in the extended sense. The same inequality \(L(x)\ge \mathbf{T}_\beta^+\big[u^{\hat{\phi}}_T\big](x)\) is trivially true if \(L(x)=\infty\). Plugging this into 42 , we obtain \[\label{lsc} \liminf_{k\to\infty} \mu_0\big(\mathbf{T}_\beta^+\big[u^{\phi_{n_k}}_T\big]\big) \ge \mu_0\big(\mathbf{T}_\beta^+\big[u^{\hat{\phi}}_T\big]\big).\tag{43}\]

  • Set \(A_k:=\int \mathbf{T}_\beta^+[\phi_{n_k}]\,\mathrm{d}\mu_T\) and \(B_k:=\int \mathbf{T}_\beta^+\big[u_T^{\phi_{n_k}}\big]\,\mathrm{d}\mu_0,\) \(k\ge 1\). By (3a), the sequence \((A_k)_k\) is bounded above, and by 38 , we see that \(A_k-B_k=\mathfrak{J}(\phi_{n_k})\longrightarrow {\cal V}_{\rm red}\). Hence \((B_k)_k\) is bounded above. Together with 43 , this implies \[\label{hatpsiintegrability} \mu_0\big(\mathbf{T}_\beta^+\big[u_T^{\hat{\phi}}\big]\big)<\infty,\tag{44}\] so that \(\mathfrak{J}(\hat{\phi})\) is well defined. Combining the conclusion from (3a) with 43 , we now obtain \[\begin{align} {\cal V}_{\rm red} = \lim_{k\to\infty}\mathfrak{J}(\phi_{n_k}) &\le \limsup_{k\to\infty}\int \mathbf{T}_\beta^+[\phi_{n_k}]\,\mathrm{d}\mu_T - \liminf_{k\to\infty}\int \mathbf{T}_\beta^+\big[u_T^{\phi_{n_k}}\big]\,\mathrm{d}\mu_0 \\ &\le \int \mathbf{T}_\beta^+[\hat{\phi}]\,\mathrm{d}\mu_T - \int \mathbf{T}_\beta^+\big[u_T^{\hat{\phi}}\big]\,\mathrm{d}\mu_0 = \mathfrak{J}(\hat{\phi}), \end{align}\] which shows that \({\cal V}_{\rm red}=\mathfrak{J}(\hat{\phi})\), as required.

We now prove that the induced potential \(\hat{\psi}:={\mathbf{T}}_\beta^+[\hat{\phi}]\) attains the supremum in the relaxed dual problem \(\bar{\cal V}\).

(a). Fix \((t,x)\in[0,T)\times\mathbb{R}^d\) and let \(\mathbb{P}\in {\cal P}_t(\delta_x)\) satisfy \(\mathbb{E}^\mathbb{P}\int_t^T c(\gamma_r^\mathbb{P})\;\mathrm{d}r<\infty\). Since \(\mathbb{P}\circ X_t^{-1}=\delta_x\), we have \(X_t=x\), \(\mathbb{P}\)-a.s. Hence, after subtracting at time \(t\) and defining \[W_s:=W_s^\mathbb{P}-W_t^\mathbb{P},\qquad \alpha_s:=\alpha_s^\mathbb{P},\qquad \sigma_s:=\sigma_s^\mathbb{P},\qquad s\in[t,T],\] we obtain \(X_s=x+\int_t^s\alpha_r\;dr+\int_t^s\sigma_r\;\mathrm{d}W_r, \;t\le s\le T,\; \mathbb{P}\)-a.s., where \(W=(W_s)_{s\in[t,T]}\) is a \(d\)-dimensional \(\mathbb{P}\)-Brownian motion on \([t,T]\).

Fix \(y\in\mathbb{R}^d\) and define the auxiliary process \(Y_s^y:=y+\int_t^s\alpha_r\;\mathrm{d}r+ W_s,\; s\in[t,T]\). Then \[\mathbb{E}^\mathbb{P}|Y_T^y|^2 \le 3|y|^2+3(T-t)\mathbb{E}^\mathbb{P}\int_t^T |\alpha_r|^2dr+3d(T-t)<\infty.\] Since \(\hat{\phi}\) is finite and \(\beta\)–convex and \(\hat{\phi}(0)=0\), we have for all \(p\in\partial\hat{\phi}(0)\): \[\hat{\phi}(z)\ge p\cdot z-\frac{\beta}{2}|z|^2 \ge -|p|\;|z|-\frac{\beta}{2}|z|^2,\; z\in\mathbb{R}^d.\] Since \(\mathbb{E}^\mathbb{P}|Y_T^y|^2<\infty\), this implies \(\hat{\phi}(Y_T^y)^-\in\mathbb{L}_1(\mathbb{P})\). Let \(\Omega_{t,T}:=C([t,T];\mathbb{R}^d)\), let \(R^{t,y}\) be Wiener measure on \(\Omega_{t,T}\) started from \(y\) at time \(t\), and let \(\mathbb{Q}^y:=\text{Law}_\mathbb{P}(Y^y)\in{\cal P}(\Omega_{t,T})\). We claim that \[\label{entropylessthan} H(\mathbb{Q}^y\;|\;R^{t,y})\le \frac{1}{2}\mathbb{E}^\mathbb{P}\int_t^T |\alpha_r|^2 \mathrm{d}r.\tag{45}\] For this, let \(n\ge1\) and define \(\tau_n:=\inf\Big\{s\ge t:\int_t^s |\alpha_r|^2 \mathrm{d}r\ge n\Big\}\wedge T, \;\alpha_r^{(n)}:=\alpha_r\;1_{\{r\le\tau_n\}}\), and \(Y_s^{(n),y}:=y+\int_t^s\alpha_r^{(n)}\mathrm{d}r+W_s,\;s\in[t,T]\). Set \(Z_n:=e^{ -\int_t^T \alpha_r^{(n)}\cdot \mathrm{d}W_r -\frac{1}{2}\int_t^T |\alpha_r^{(n)}|^2 \mathrm{d}r}\). Since \(\int_t^T |\alpha_r^{(n)}|^2 \mathrm{d}r\le n\), Novikov’s criterion holds, so \(Z_n\) is a true martingale and \(\mathrm{d}\widetilde{\mathbb{P}}_n:=Z_n\;\mathrm{d}\mathbb{P}\) defines a probability measure. Under \(\widetilde{\mathbb{P}}_n\), the process \(\widetilde{W}^{(n)}:=W+\int_t^. \alpha_r^{(n)} \mathrm{d}r\) is a Brownian motion on \([t,T]\). Therefore, if \(\mathbb{Q}_n:=\text{Law}_\mathbb{P}(Y^{(n),y})\), then \(\text{Law}_{\widetilde{\mathbb{P}}_n}(Y^{(n),y})=R^{t,y}\). By contraction of relative entropy under the measurable map \(\omega\mapsto Y^{(n),y}(\omega)\), \[H(\mathbb{Q}_n\;|\;R^{t,y}) \le H(\mathbb{P}\;|\;\widetilde{\mathbb{P}}_n) = \mathbb{E}^\mathbb{P}\!\left[\log\frac{d\mathbb{P}}{d\widetilde{\mathbb{P}}_n}\right] = \mathbb{E}^\mathbb{P}[-\log Z_n] = \frac{1}{2}\mathbb{E}^\mathbb{P}\int_t^T |\alpha_r^{(n)}|^2 \mathrm{d}r,\] because \(\mathbb{E}^\mathbb{P}[\int_t^T \alpha_r^{(n)}\cdot \mathrm{d}W_r]=0\). Hence \(H(\mathbb{Q}_n\;|\;R^{t,y})\le \frac{1}{2}\mathbb{E}^\mathbb{P}\int_t^T |\alpha_r^{(n)}|^2 \mathrm{d}r\).

Also, \(\sup_{s\in[t,T]}|Y_s^{(n),y}-Y_s^y| \le \int_{\tau_n}^T |\alpha_r| \mathrm{d}r\), thus \(\mathbb{E}^\mathbb{P}\Big[\sup_{s\in[t,T]}|Y_s^{(n),y}-Y_s^y|^2\Big] \le (T-t)\mathbb{E}^\mathbb{P}\int_{\tau_n}^T |\alpha_r|^2 \mathrm{d}r\longrightarrow 0\). Thus \(\mathbb{Q}_n\Rightarrow \mathbb{Q}^y\) weakly on \(\Omega_{t,T}\). By lower semicontinuity of relative entropy, \(H(\mathbb{Q}^y\;|\;R^{t,y}) \le \liminf_{n\to\infty} H(\mathbb{Q}_n\;|\;R^{t,y}) \le \frac{1}{2}\mathbb{E}^\mathbb{P}\int_t^T |\alpha_r|^2 \mathrm{d}r\).

For \(m\ge1\), define \(f_m(\omega):=\hat{\phi}(\omega_T)\wedge m,\; \omega\in\Omega_{t,T}\). Then \(f_m^+\le m\) and \(f_m^-= \hat{\phi}(\omega_T)^-\), so \(f_m\in \mathbb{L}^1(\mathbb{Q}^y)\). Moreover, \(e^{f_m(\omega)}\le e^{\hat{\phi}(\omega_T)},\; \omega\in\Omega_{t,T}\), hence \[\mathbb{E}^{R^{t,y}}[e^{f_m}] \le \mathbb{E}^{R^{t,y}}[e^{\hat{\phi}(\omega_T)}] = ({\cal N}_{T-t}*e^{\hat{\phi}})(y) = e^{u_{T-t}^{\hat{\phi}}(y)} <\infty.\] Applying the Donsker–Varadhan variational formula to \(f_m\) under \(R^{t,y}\), we see that \[u_{T-t}^{\hat{\phi}}(y) \ge \mathbb{E}^{\mathbb{Q}^y}[f_m]-H(\mathbb{Q}^y| R^{t,y}) \uparrow \mathbb{E}^{\mathbb{Q}^y}[\hat{\phi}(Y^y_T)]-H(\mathbb{Q}^y| R^{t,y}) \ge \mathbb{E}^{\mathbb{Q}^y}[\hat{\phi}(Y^y_T)] -\frac{1}{2}\mathbb{E}^\mathbb{P}\!\!\int_t^T\!\! |\alpha_r|^2\mathrm{d}r,\] by monotone convergence and 45 . Next, we have \(\hat{\psi}(X_T)\le \hat{\phi}(Y_T^y)+\frac{\beta}{2}|X_T-Y_T^y|^2, \;\mathbb{P}\)-a.s. As \(\mathbb{E}^\mathbb{P}|X_T-Y_T^y|^2 = |x-y|^2+\mathbb{E}^\mathbb{P}\int_t^T |\sigma_r-I_d|^2 \mathrm{d}r\), we get \[\begin{align} \mathbb{E}^\mathbb{P}\Big[\hat{\psi}(X_T)-\!\!\int_t^T\!\!\!\!c(\gamma_r) \mathrm{d}r\Big] &\le \mathbb{E}^\mathbb{P}[\hat{\phi}(Y_T^y)]-\frac{1}{2}\mathbb{E}^\mathbb{P}\!\!\int_t^T\!\!\!\! |\alpha_r|^2\mathrm{d}r +\frac{\beta}{2}\Big(\mathbb{E}^\mathbb{P}|X_T\!-\!Y_T^y|^2\!-\!\mathbb{E}^\mathbb{P}\!\!\int_t^T \!\!\!\!|\sigma_r\!-\!I_d|^2\mathrm{d}r\Big) \\ &\le u_{T-t}^{\hat{\phi}}(y)+\frac{\beta}{2}|x-y|^2. \end{align}\] By the arbitrariness of \(y\in\mathbb{R}^d\) and \(\mathbb{P}\in {\cal P}_t(\delta_x)\), this shows that \[\label{eq:V-upper-final-step4} V_t^{\hat{\psi}}(x)\le \hat{v}_t(x),\tag{46}\] which proves the claim.

(b). Since \(\hat{\phi}\) is finite and \(\beta\)–convex, the map \(\hat{\psi}={\mathbf{T}}_\beta^+[\hat{\phi}]\) is finite and \(\beta\)–concave. Moreover, \(\mathbf{T}_\beta^-[\hat{\psi}] = \mathbf{T}_\beta^-\mathbf{T}_\beta^+[\hat{\phi}] = \hat{\phi}\), and by 44 , \(\mu_0\Big(\mathbf{T}_\beta^+ \big[u_T^{\mathbf{T}_\beta^-[\hat{\psi}]}\big]\Big) = \mu_0\big(\mathbf{T}_\beta^+[u_T^{\hat{\phi}}]\big)<\infty\). Thus \(\hat{\psi}\in\bar{\cal C}^{\text{conc}}\subset{\cal C}^{\text{conc}}\). Moreover, 46 implies \(V_0^{\hat{\psi}}\le \hat{v}_0\), and therefore \(\mu_0(V_0^{\hat{\psi}})\le \mu_0(\hat{v}_0)<\infty\). We may apply Lemma 3 together with 46 to obtain \[\label{eq:hatpsi-final-step4} {\cal V}(\mu_0,\mu_T)\ge \mu_T(\hat{\psi})-\mu_0(V_0^{\hat{\psi}}) \ge \mu_T(\hat{\psi})-\mu_0(\hat{v}_0) = \mathfrak J(\hat{\phi}) ={\cal V}_{\rm red}(\mu_0,\mu_T),\tag{47}\] by Step 3. As \({\cal V}_{\rm red}(\mu_0,\mu_T)={\cal V}(\mu_0,\mu_T)\) by Theorem 3 (i), we get \(\mu_T(\hat{\psi})-\mu_0(\hat{v}_0) = \mathfrak J(\hat{\phi}) = {\cal V}_{\rm red}(\mu_0,\mu_T) = {\cal V}(\mu_0,\mu_T)\). Combining this with 47 , we obtain \[\mu_T(\hat{\psi})-\mu_0(V_0^{\hat{\psi}})={\cal V}(\mu_0,\mu_T), \qquad \mu_0(V_0^{\hat{\psi}})=\mu_0(\hat{v}_0),\] and then \(V_0^{\hat{\psi}}=\hat{v}_0\), \(\mu_0\)–a.s. by 46 , proving that \(\hat{\psi}\) attains the supremum in \(\bar{\cal V}(\mu_0,\mu_T)\).

In order to prove the claimed convergence of \(\mathbf{T}^+_\beta[\phi_{n_k}]\) in 41 , fix \(R>0\) and \(s\in(0,\eta T)\), where \(\eta\in(0,\delta)\) is the parameter fixed in Step 1. Then \[\begin{align} \big\|\mathbf{T}^+_\beta[\phi_{n_k}]-\mathbf{T}^+_\beta[\hat{\phi}]\big\|_{\mathbb{L}^\infty(B_R)} &\le \big\|\mathbf{T}^+_\beta[\phi_{n_k}]-\mathbf{T}^+_\beta[u_s^{\phi_{n_k}}]\big\|_{\mathbb{L}^\infty(B_R)} \\ &\quad +\big\|\mathbf{T}^+_\beta[u_s^{\phi_{n_k}}]-\mathbf{T}^+_\beta[u_s^{\hat{\phi}}]\big\|_{\mathbb{L}^\infty(B_R)} \\ &\quad +\big\|\mathbf{T}^+_\beta[u_s^{\hat{\phi}}]-\mathbf{T}^+_\beta[\hat{\phi}]\big\|_{\mathbb{L}^\infty(B_R)}. \end{align}\] We shall prove in Step 5 below that, for every fixed \(s\in(0,\eta T)\), \[\label{usphiconv} u^{\phi_{n_k}}_s \longrightarrow u^{\hat{\phi}}_s \quadand\quad \mathbf{T}^+_\beta[u^{\phi_{n_k}}_s]\longrightarrow \mathbf{T}^+_\beta[u^{\hat{\phi}}_s], ~~locally uniformly in \mathbb{R}^d.\tag{48}\] Hence, for every fixed \(R>0\) and \(s\in(0,\eta T)\), \[\label{task-1} \limsup_{k\to\infty} \big\|\mathbf{T}^+_\beta[\phi_{n_k}]-\mathbf{T}^+_\beta[\hat{\phi}]\big\|_{\mathbb{L}^\infty(B_R)} \le \big\|\mathbf{T}^+_\beta[u_s^{\hat{\phi}}]-\mathbf{T}^+_\beta[\hat{\phi}]\big\|_{\mathbb{L}^\infty(B_R)} + \theta_s^R,\tag{49}\] where \[\label{theta} \theta_s^R := \sup_{n\ge 1} \big\|\mathbf{T}^+_\beta[u_s^{\phi_n}]-\mathbf{T}^+_\beta[\phi_n]\big\|_{\mathbb{L}^\infty(B_R)}.\tag{50}\] In Step 6 we show that \[\theta_s^R\longrightarrow 0 ~~as s\searrow 0,\] and the same argument, applied to \(\hat{\phi}\), yields \[\big\|\mathbf{T}^+_\beta[u_s^{\hat{\phi}}]-\mathbf{T}^+_\beta[\hat{\phi}]\big\|_{\mathbb{L}^\infty(B_R)} \longrightarrow 0 ~~as s\searrow 0.\] Together with 49 , this proves 41 .

In this step we justify 48 . Fix \(s\in(0,\eta T)\) with \(\eta\in(0,\delta)\) as in Step 1, and set \[\Phi:=\{\hat{\phi},\phi_n,\;n\ge1\}.\] By Lemma 6 (ii), ?? , at time \(0\), together with the boundedness of \(\big(U_\eta^{\phi_n}\big)_n\) from Step 1, there exists a constant \(C_\eta<\infty\) such that \[\label{eq:common-upper-bound} \phi_n\le C_\eta+\frac{|.|^2}{2\eta T},~n\ge1, ~and therefore~ \widehat\phi\le C_\eta+\frac{|.|^2}{2\eta T},\tag{51}\] by 40 .

In particular, as \(s<\eta T\), we have \(u_s^\phi(x)=\log\big({\cal N}_s*e^\phi(x)\big)\in\mathbb{R}\) for all \(x\in\mathbb{R}^d,\;\phi\in\Phi\). We split the remaining arguments into three parts.

  • For \(\phi\in\Phi\) and \(R\ge1\), define \[\zeta^\phi_R(s,x) := \frac{{\cal N}_s*(e^\phi\mathbf{1}_{B_R^c})(x)}{{\cal N}_s*(e^\phi\mathbf{1}_{B_R})(x)}.\] Then there exist constants \(C_s,a_s>0\), depending only on \((\beta,T,d,\eta,s)\), such that \[\label{eq:tail-ratio} \sup_{\phi\in\Phi}\zeta^\phi_R(s,x) \le C_s\exp\!\Big( -\frac{1}{4}\Big(\frac{1}{s}-\frac{1}{\eta T}\Big)R^2+a_s|x|^2 \Big), \qquad x\in\mathbb{R}^d,\;R\ge1.\tag{52}\] Indeed, \(\phi(y)\le C_\eta+\frac{|y|^2}{2\eta T}\) for all \(\phi\in\Phi\) by 51 , and using \(\frac{|x-y|^2}{2s} = \frac{|y|^2}{2\eta T} +\frac{|x-y|^2}{2s} -\frac{|y|^2}{2\eta T} \ge \frac{|y|^2}{2\eta T} +\frac{1}{4}\Big(\frac{1}{s}-\frac{1}{\eta T}\Big)|y|^2-a_s|x|^2,\) for a suitable constant \(a_s>0\), we obtain \[\begin{align} {\cal N}_s\!*\!(e^\phi\mathbf{1}_{B_R^c})(x) \le e^{C_\eta} \int_{B_R^c}\!\!\!\frac{e^{\frac{|y|^2}{2\eta T}-\frac{|x-y|^2}{2s}}}{(2\pi s)^{\frac{d}{2}}}\,\mathrm{d}y \le C_s e^{ -\frac{1}{4}\Big(\frac{1}{s}-\frac{1}{\eta T}\Big)R^2+a_s|x|^2}. \label{ineqeta1} \end{align}\tag{53}\] On the other hand, since \(\Phi\) is locally bounded below on \(B_1\), there exists \(m_1<\infty\) such that \(\phi(y)\ge -m_1\) for all \(y\in B_1,\;\phi\in\Phi.\) Hence, for \(R\ge1\) and every \(\phi\in\Phi\), \[{\cal N}_s*(e^\phi\mathbf{1}_{B_R})(x) \ge {\cal N}_s*(e^\phi\mathbf{1}_{B_1})(x) \ge (2\pi s)^{-\frac{d}{2}}e^{-m_1}\int_{B_1}e^{-\frac{|x-y|^2}{2s}}\,\mathrm{d}y \ge C_s'e^{-a_s'|x|^2},\] for some constants \(C_s'>0\) and \(a_s'\ge0\), independent of \(\phi\) and \(R\). Combining this with 53 yields 52 .

  • Fix \(R>0\), and set \[\gamma_{k,R}:=\|\phi_{n_k}-\hat{\phi}\|_{\mathbb{L}^\infty(B_R)}\longrightarrow 0 \qquadas k\to\infty\] by 40 . For every fixed \(x\in\mathbb{R}^d\), the estimate in (5a) implies that \[\sup_{\phi\in\Phi}\zeta^\phi_R(s,x)\longrightarrow 0 \qquadas R\to\infty.\] We now write \[\begin{align} {\cal N}_s*e^{\phi_{n_k}}(x) &\ge e^{-\gamma_{k,R}}{\cal N}_s*(\mathbf{1}_{B_R}e^{\hat{\phi}})(x) = \frac{e^{-\gamma_{k,R}}}{1+\zeta^{\hat{\phi}}_R(s,x)} \,{\cal N}_s*e^{\hat{\phi}}(x), \\ {\cal N}_s*e^{\phi_{n_k}}(x) &\le \big(1+\zeta^{\phi_{n_k}}_R(s,x)\big){\cal N}_s*(\mathbf{1}_{B_R}e^{\phi_{n_k}})(x) \le \big(1+\zeta^{\phi_{n_k}}_R(s,x)\big)e^{\gamma_{k,R}}{\cal N}_s*e^{\hat{\phi}}(x). \end{align}\] Sending \(k\to\infty\) and then \(R\to\infty\), we see that \(u^{\phi_{n_k}}_s(x)\longrightarrow u^{\hat{\phi}}_s(x)\), for all \(x\in\mathbb{R}^d\). Moreover, the positive-time semiconvexity estimate from Section 5 implies that, for this fixed \(s\), the maps \(u_s^{\phi_{n_k}}\!+\!\frac{\kappa(s)}{2}|.|^2\) are convex on \(\mathbb{R}^d\) for all \(k\), with \(\kappa(s):=\frac{\beta}{1+\beta s}\). Since the limit is finite everywhere, Rockafellar [12], Theorem 10.8, p. 90, implies \[u^{\phi_{n_k}}_s\longrightarrow u^{\hat{\phi}}_s \qquadlocally uniformly in \mathbb{R}^d.\]

  • We next prove the convergence of the Moreau envelopes. For every \(\phi\in\Phi\), define \[G^\phi_s(y):=u_s^\phi(y)+\frac{\beta}{2}|x-y|^2.\] As \(\Phi\) is locally bounded from below on \(B_1\), the same argument as in (5a) gives a quadratic lower bound \(u_s^\phi\ge -c_s-a_s|.|^2\), \(\phi\in\Phi\), for some constants \(c_s,a_s>0\) depending only on \((\beta,T,d,\eta,s)\). Hence, after possibly enlarging constants, \[\label{eq:coercive-step5} u_s^\phi(y)+\frac{\beta}{2}|x-y|^2 \ge -C_{s,r}+c_{s,r}|y|^2, \qquad x\in B_r,\;\phi\in\Phi,\tag{54}\] for some constants \(C_{s,r}>0\) and \(c_{s,r}>0\). In particular, for fixed \((s,r)\), the right-hand side tends to \(+\infty\) as \(|y|\to\infty\), uniformly in \(x\in B_r\) and \(\phi\in\Phi\).

    On the other hand, by 51 and \(s<\eta T\), there exists \(C_{s,r}'<\infty\) such that \(u_s^\phi(0)\le C_{s,r}'\) for all \(\phi\in\Phi\). Therefore, \[\mathbf{T}^+_\beta[u_s^\phi](x) \le u_s^\phi(0)+\frac{\beta}{2}|x|^2 \le C_{s,r}' + \frac{\beta}{2}r^2, \qquad x\in B_r,\;\phi\in\Phi.\] Combining this with 54 , we deduce that every minimiser \(y\) of \(u_s^\phi(\cdot)+\frac{\beta}{2}|x-\cdot|^2\) for \(x\in B_r\) and \(\phi\in\Phi\) satisfies \(|y|\le R_{s,r}\) for some \(R_{s,r}<\infty\) independent of \(\phi\) and \(x\). Thus \[\mathbf{T}^+_\beta[u_s^\phi](x) = \inf_{|y|\le R_{s,r}} \Big\{ u_s^\phi(y)+\frac{\beta}{2}|x-y|^2 \Big\}, \qquad x\in B_r,\;\phi\in\Phi.\] The local uniform convergence of \(u^{\phi_{n_k}}_s\) established in (5b) then implies \[\mathbf{T}^+_\beta[u^{\phi_{n_k}}_s] \longrightarrow \mathbf{T}^+_\beta[u^{\hat{\phi}}_s], ~~locally uniformly in \mathbb{R}^d,\] which proves 48 .

It remains to prove that, for the family \(\Phi:=\{\phi_n,\;n\ge 1\}\), we have \[\label{theta-seq} \theta_s^R := \sup_{n\ge1} \big\|\mathbf{T}^+_\beta[u_s^{\phi_n}]-\mathbf{T}^+_\beta[\phi_n]\big\|_{\mathbb{L}^\infty(B_R)} \longrightarrow 0 \qquadas s\searrow 0, \quadfor all R>0.\tag{55}\] Fix \(R>0\) and \(\rho\in(0,1)\). Since \(\Phi\) is locally bounded and \(\beta\)–convex, it is locally equi-Lipschitz, so there exists a common Lipschitz constant \(L=L_{R,\rho}\) on \(B_{R+\rho}\). Denoting \(\delta_s(\rho):=\mathbb{P}(|W_s|\ge \rho)\), we see that \[\label{step4ineq1} e^{u^\phi_s(x)} = {\cal N}_s*e^\phi(x) \ge {\cal N}_s*(e^\phi\mathbf{1}_{B_\rho(x)})(x) \ge e^{\phi(x)-L\rho}\big(1-\delta_s(\rho)\big), ~ \phi\in\Phi,\;x\in B_R.\tag{56}\] On the other hand, \[e^{u^\phi_s(x)} \le {\cal N}_s*(e^\phi\mathbf{1}_{B_\rho(x)})(x) + {\cal N}_s*(e^\phi\mathbf{1}_{B_\rho(x)^c})(x).\] The first term is bounded above by \({\cal N}_s*(e^\phi\mathbf{1}_{B_\rho(x)})(x)\le e^{\phi(x)+L\rho}.\) For the second term, by the same Gaussian-tail argument as in Step 5(a), using the fixed parameter \(\eta\) from Step 1, there exists a deterministic function \(g_{R,\rho}(s)\), defined for \(0<s<\eta T\), such that \[g_{R,\rho}(s)\longrightarrow 0 ~as s\searrow 0, ~and~{\cal N}_s*(e^\phi\mathbf{1}_{B_\rho(x)^c})(x)\le g_{R,\rho}(s), ~\phi\in\Phi,\;x\in B_R.\] Hence \[e^{u^\phi_s(x)} \le e^{\phi(x)+L\rho}+g_{R,\rho}(s), \qquad \phi\in\Phi,\;x\in B_R.\] Together with 56 , this yields \[-L\rho+\log\big(1-\delta_s(\rho)\big) \le u^\phi_s(x)-\phi(x) \le \log\big(e^{\phi(x)+L\rho}+g_{R,\rho}(s)\big)-\phi(x) \le L\rho + C_{R,\rho} g_{R,\rho}(s),\] for some constant \(C_{R,\rho}>0\), by the Lipschitz property of \(\log\) on a bounded interval away from the origin, using the local uniform boundedness of \(\Phi\) on \(B_R\). Since both \(\delta_s(\rho)\to 0\) and \(g_{R,\rho}(s)\to 0\) as \(s\searrow 0\), we obtain \[\label{theta0} \limsup_{s\searrow 0} \sup_{\phi\in\Phi} \|u^\phi_s-\phi\|_{\mathbb{L}^\infty(B_R)} \le L\rho \underset{\rho\to 0}{\longrightarrow} 0 ~for all R>0.\tag{57}\] We next localize the minimizers in the definition of \(\mathbf{T}_\beta^+[u_s^\phi]\). By the same lower-bound argument as in Step 5(c), for every \(R>0\) there exists \(R'>0\) such that, for all sufficiently small \(s\in(0,\eta T)\), all \(\phi\in\Phi\), and all \(x\in B_R\), every minimizer in the definition of \(\mathbf{T}_\beta^+[u_s^\phi](x)\) belongs to \(B_{R'}\). Denoting by \(y_s^\phi(x)\in B_{R'}\) such a minimizer, we have \[\begin{align} \mathbf{T}_\beta^+[u_s^\phi](x) &= u_s^\phi\big(y_s^\phi(x)\big)+\frac{\beta}{2}\big|x-y_s^\phi(x)\big|^2 \\ &\le \phi\big(y_s^\phi(x)\big)+\frac{\beta}{2}\big|x-y_s^\phi(x)\big|^2 +\|u_s^\phi-\phi\|_{\mathbb{L}^\infty(B_{R'})} \\ &\le \mathbf{T}_\beta^+[\phi](x)+\|u_s^\phi-\phi\|_{\mathbb{L}^\infty(B_{R'})}. \end{align}\] On the other hand, by Jensen’s inequality and the \(\beta\)–convexity of \(\phi\), we have \(u_s^\phi\ge \phi-\frac{\beta d}{2}s\), and therefore \(\mathbf{T}_\beta^+[u_s^\phi](x)\ge \mathbf{T}_\beta^+[\phi](x)-\frac{\beta d}{2}s.\) Combining the last two estimates, we obtain \[\big|\mathbf{T}_\beta^+[u_s^\phi](x)-\mathbf{T}_\beta^+[\phi](x)\big| \le \|u_s^\phi-\phi\|_{\mathbb{L}^\infty(B_{R'})}+\frac{\beta d}{2}s, ~~ \phi\in\Phi,\;x\in B_R,\] for all sufficiently small \(s\in(0,\eta T)\). Taking the supremum over \(\phi\in\Phi\) and \(x\in B_R\), and using 57 on \(B_{R'}\), we obtain 55 .   \({\cal t}\)  \({\cal u}\)

7 Primal Attainment↩︎

To prove existence of a primal optimiser \(\widehat\mathbb{P}\) for the problem SBB\((\mu_0,\mu_T)\) we need the following additional regularity properties of the maps \(\hat{u},\hat{v}\). and the corresponding \(\mathscr{Y}\).

Lemma 7. (i) For \(\mu_0\)-a.e. \(x\), the minimum in \(\mathbf{T}_\beta^+[\hat{u}_T](x)\) is attained by a unique minimizer \(\mathscr{Y}_0(x)\). Moreover, \(\mathscr{Y}_0\) is Lipschitz on this full \(\mu_0\)-measure set; we fix a Lipschitz Borel extension, still denoted \(\mathscr{Y}_0\).
(ii) The minimum value \(\hat{\psi}=\mathbf{T}_\beta^+[\hat{\phi}]\) is attained by a non-empty compact set of minimizers \(\mathscr{Y}_T\).
(iii) For all \(t<T\), we have \(|\frac{\hat{v}}{w}|_{\mathbb{L}^\infty[0,t]}+|\frac{\nabla \hat{u}}{\sqrt{w}}|_{\mathbb{L}^\infty[t,T]}<\infty\).

(i) By Step 3(c) in the proof of Theorem 3 (ii-a) in Section 6, we have \(\hat{v}_0=\mathbf{T}_\beta^+[\hat{u}_T]\in \mathbb{L}^1(\mu_0)\) and in particular \(\hat{v}_0(x)<\infty\) for \(\mu_0\)-a.e.\(x\). Denote, for such \(x\), \(F_x(y):=\hat{u}_T(y)+\frac{\beta}{2}|x-y|^2\) . Moreover \(\hat{u}_T(x)+\frac{|x|^2}{2T} = -\frac{d}{2}\log(2\pi T) + \log\!\int_{\mathbb{R}^d} e^{\hat{\phi}(y)-\frac{|y|^2}{2T}+\frac{x\cdot y}{T}}\,\mathrm{d}y\) is convex as a log-Laplace transform. Then \[\label{eq:hess95uT95final} D^2\hat{u}_T\succeq -\frac{1}{T}\;I_d ~and therefore~ D^2F_x(y)\succeq \Big(\beta-\frac{1}{T}\Big)I_d\succ0,\tag{58}\] as \(\beta T>1\). Thus \(F_x\) attains a unique minimizer, and since \((x,y)\mapsto F_x(y)\) is Borel, the measurable maximum theorem yields a Borel map \(\mathscr{Y}_0\).

For \(x_i\in\mathbb{R}^d\) and \(y_i:=\mathscr{Y}_0(x_i)\), \(i=1,2\), the first-order condition for the minimization writes \(x_i=y_i+\beta^{-1}\nabla \hat{u}_T(y_i)\). Denote \(\delta x:=x_1-x_2\), \(\delta y:=y_1-y_2\), then by subtracting and taking the scalar product with \(\delta y\), we obtain \(\delta x\cdot\delta y =\! |\delta y|^2+\frac{1}{\beta}\big(\nabla\hat{u}_T(y_1)\!-\!\nabla\hat{u}_T(y_2)\big)\!\cdot\! \delta y \ge (1-\frac{1}{\beta T})|\delta y|^2,\) by 58 . By the Cauchy-Schwarz inequality, this implies that \(|\delta y|\le (1-\frac{1}{\beta T})^{-1}|\delta x|\), which proves the required Lipschitz continuity of \(\mathscr{Y}_0\).

(ii) As \(\hat{\phi}\) is \(\beta\)-convex, the map \(g_0:=\hat{\phi}+\frac{\beta}{2}|.|^2\) is proper, l.s.c. and convex. We have \(\hat{\psi}(x)= \frac{\beta}{2}|x|^2-g_0^*(\beta x)\), By the first duality result, the maximizing sequence may be chosen so that \(\psi_{n_k}:={\mathbf{T}}_\beta^+[\phi_{n_k}]\in C_w\), hence each \(\psi_{n_k}\) is finite on \(\mathbb{R}^d\). Together with the local uniform convergence \(\psi_{n_k}\to \hat{\psi}={\mathbf{T}}_\beta^+[\hat{\phi}]\), this implies that \(\hat{\psi}\) is finite on all of \(\mathbb{R}^d\). Hence \(g_0^*(p)<\infty\) for all \(p\in\mathbb{R}^d\), and therefore \(g_0\) is coercive, and the map \(y\mapsto g_0(y)-\beta x\!\cdot\!y\) is coercive and l.s.c. for each fixed \(x\) and, as such, attains its minimum with compact argmin set \(\mathscr{Y}_T(x)\).

(iii) Fix now \(\tau<T\) and let us prove for some constant \(C_\tau\) that \[\begin{align} |\nabla \hat{u}_s(x)|\le C_\tau(1+|x|), \qquad s\in[T-\tau,T],\;x\in\mathbb{R}^d, \tag{59} \\ |\hat{v}_t(x)|\le C_\tau(1+|x|^2), \qquad 0\le t\le \tau,\;x\in\mathbb{R}^d,~for all~\tau<T. \tag{60} \end{align}\] Set \(\kappa_\tau:=\frac{\beta}{1+\beta(T-\tau)}\) and define \(G_s(x):=\hat{u}_s(x)+\frac{\kappa_\tau}{2}|x|^2, \;s\in[T-\tau,T]\). By 23 , each \(G_s\) is convex. Moreover, 22 implies \(G_s(x)\le C_\tau(1+|x|^2), \; s\in[T-\tau,T],\;x\in\mathbb{R}^d\). Since \(s\mapsto \hat{u}_s(0)\) is continuous on \([T-\tau,T]\), we also have \(|G_s(0)|\le C_\tau,\; s\in[T-\tau,T]\). Choose any \(p_s\in \partial G_s(0)\). For every unit vector \(e\), \(p_s\cdot e\le G_s(e)-G_s(0), \;-p_s\cdot e\le G_s(-e)-G_s(0)\). Hence \(|p_s|\le C_\tau,\; s\in[T-\tau,T]\). By convexity, \(G_s(x)\ge G_s(0)+p_s\cdot x\ge -C_\tau(1+|x|)\). Now fix \(x\in\mathbb{R}^d\), \(s\in[T-\tau,T]\), and a unit vector \(e\), and set \(r:=1+|x|\). Since \(G_s\) is convex and differentiable, \(\nabla G_s(x)\cdot e\le \frac{G_s(x+re)-G_s(x)}{r}, \;-\nabla G_s(x)\cdot e\le \frac{G_s(x-re)-G_s(x)}{r}\). Using the quadratic upper bound at \(x\pm re\) and the affine lower bound at \(x\), we obtain \(|\nabla G_s(x)\cdot e| \le \frac{C_\tau(1+|x\pm re|^2)+C_\tau(1+|x|)}{r}\). Since \(r=1+|x|\) and \(|x\pm re|\le |x|+r\le 1+2|x|\), this yields \(|\nabla G_s(x)\cdot e|\le C_\tau(1+|x|)\). Taking the supremum over \(|e|=1\) yields \(|\nabla G_s(x)|\le C_\tau(1+|x|)\). Since \(\nabla \hat{u}_s(x)=\nabla G_s(x)-\kappa_\tau x\), we obtain 59 .

It remains to prove 60 . For the upper bound, choose \(y=0\) in the Moreau envelope defining \(\hat{v}_t=\mathbf{T}_\beta^+[\hat{u}_{T-t}]\). Since \(T-t\in[T-\tau,T]\) for \(0\le t\le\tau\), and since \(s\mapsto \hat{u}_s(0)\) is bounded above on \([T-\tau,T]\), we get \(\hat{v}_t(x) \le \hat{u}_{T-t}(0)+\frac{\beta}{2}|x|^2 \le C_\tau+\frac{\beta}{2}|x|^2, \; 0\le t\le \tau\).

For the lower bound, use the Hessian estimate from Lemma 4(3-a): \(D^2\hat{v}_t \succeq -\frac{1}{T-t}I_d \succeq -\frac{1}{T-\tau}I_d, \; 0\le t\le\tau\). Set \(H_t(x):=\hat{v}_t(x)+\frac{|x|^2}{2(T-\tau)},\; 0\le t\le\tau\). Then \(H_t\) is convex. Since \(\hat{v}_t(0)=\mathbf{T}_\beta^+[\hat{u}_{T-t}](0)\) is bounded from below on \([0,\tau]\), and since the subgradients of \(H_t\) at \(0\) are bounded on \([0,\tau]\) by the same local convexity argument used above, there exists \(C_\tau<\infty\) such that \(H_t(x)\ge -C_\tau(1+|x|), \; 0\le t\le\tau,\;x\in\mathbb{R}^d\). Consequently, \(\hat{v}_t(x)\ge -C_\tau(1+|x|^2), \;0\le t\le\tau,\;x\in\mathbb{R}^d\). Together with the upper bound, this proves 60 .   \({\cal t}\)  \({\cal u}\)

1. Set \(m_0:={\mathscr{Y}_0}_\#\mu_0\) and \(\nu_0(dy):=e^{-\hat{u}_T(y)}m_0(dy)\). Since \(\hat{v}_0=\mathbf{T}_\beta^+[\hat{u}_T]\) is finite \(\mu_0\)-a.e. and \(\mathscr{Y}_0\) is the corresponding minimizer, the measure \(\nu_0\) is well-defined, non-zero, and \(\sigma\)-finite. Define \[\label{eq:def95nu095nuT95mT95final} \nu_T:=\nu_0*{\cal N}_T, \qquad m_T(dy):=e^{\hat{\phi}(y)}\nu_T(dy).\tag{61}\] Then \(m_0=e^{\hat{u}_T}\nu_0, \;m_T=e^{\hat{\phi}}\nu_T\). Since \(\mu_0\in{\cal P}_2(\mathbb{R}^d)\) and \(\mathscr{Y}_0\) is Lipschitz by Lemma 7 (i), it follows that \(m_0={\mathscr{Y}_0}_\#\mu_0\in{\cal P}_2(\mathbb{R}^d)\). Moreover, by Tonelli’s theorem and the definition of \(\hat{u}_T\), \[m_T(\mathbb{R}^d) = \nu_T(e^{\hat{\phi}}) = \nu_0({\cal N}_T*e^{\hat{\phi}}) = m_0(e^{\hat{u}_T}e^{-\hat{u}_T}) = 1.\] Hence \(m_T\in{\cal P}(\mathbb{R}^d)\). Since \(\nu_T\) is absolutely continuous with respect to \(\text{Leb}\) by Gaussian convolution, possibly with an extended density, and since \(m_T=e^{\hat{\phi}}\nu_T\) is a finite measure, we also have \(m_T\ll \text{Leb}\).

2. Fix \(g\in C_b(\mathbb{R}^d)\), let \(\hat{\phi}_\varepsilon:=\hat{\phi}+\varepsilon g, \;\hat{\psi}_\varepsilon:={\mathbf{T}}_\beta^+[\hat{\phi}_\varepsilon]\), \(\tilde{\phi}_\varepsilon:={\mathbf{T}}_\beta^-\!\big({\mathbf{T}}_\beta^+[\hat{\phi}_\varepsilon]\big), \;\bar\phi_\varepsilon:=\tilde{\phi}_\varepsilon-\tilde{\phi}_\varepsilon(0)\). By the Moreau envelope property, \(\tilde{\phi}_\varepsilon\) is \(\beta\)-convex, \(\tilde{\phi}_\varepsilon\le \hat{\phi}_\varepsilon, \;{\mathbf{T}}_\beta^+[\tilde{\phi}_\varepsilon] = {\mathbf{T}}_\beta^+[\hat{\phi}_\varepsilon]\). Since \(g\) is bounded and \(\hat{\phi}\in{\cal C}_w^{\rm conv}\), this gives \(\bar\phi_\varepsilon\in {\cal C}^{\rm conv}_{w}\) and \(\bar\phi_\varepsilon(0)=0\). Moreover \(\mathfrak{J}\) is invariant under addition of constants, so \(\mathfrak{J}(\bar\phi_\varepsilon)=\mathfrak{J}(\tilde{\phi}_\varepsilon)\). Finally, as \(\tilde{\phi}_\varepsilon\le \hat{\phi}_\varepsilon\) and \({\mathbf{T}}_\beta^+[\tilde{\phi}_\varepsilon]={\mathbf{T}}_\beta^+[\hat{\phi}_\varepsilon]\), we have \(\mathfrak{J}(\tilde{\phi}_\varepsilon) \ge \mathfrak{J}(\hat{\phi}_\varepsilon)\). Since \(\hat{\phi}\) maximizes \(\mathfrak{J}\), we obtain, for all small enough \(|\varepsilon|\), \[\label{eq:perturb95opt95final} \mathfrak{J}(\hat{\phi}_\varepsilon)\le \mathfrak{J}(\bar\phi_\varepsilon)=\mathfrak{J}(\tilde{\phi}_\varepsilon)\le \mathfrak{J}(\hat{\phi}).\tag{62}\]

Terminal derivatives. By Lemma 7 (ii), for each \(x\in\mathbb{R}^d\) the set \(\mathscr{Y}_T(x)\) is nonempty and compact. Hence, by Danskin’s theorem, \(\partial_+\hat{\psi}_\varepsilon\big|_{\varepsilon=0}(x)=\min_{y\in\mathscr{Y}_T(x)}g(y), \;\partial_-\hat{\psi}_\varepsilon\big|_{\varepsilon=0}(x)=\max_{y\in\mathscr{Y}_T(x)}g(y)\). Moreover, since \(g\) is bounded, \(\hat{\phi}(y)\!-\!|\varepsilon| |g|_\infty \le \hat{\phi}_\varepsilon(y) \le \hat{\phi}(y)\!+\!|\varepsilon||g|_\infty, \;y\in\mathbb{R}^d\), and therefore, by monotonicity of \({\mathbf{T}}_\beta^+\), \(\hat{\psi}(x)-|\varepsilon||g|_\infty \le \hat{\psi}_\varepsilon(x) \le \hat{\psi}(x)+|\varepsilon||g|_\infty, \; x\in\mathbb{R}^d\). Thus \(\left| \frac{\hat{\psi}_\varepsilon(x)-\hat{\psi}(x)}{\varepsilon} \right| \le |g|_\infty, \; \varepsilon\neq0\). Then dominated convergence yields \[\partial_\varepsilon\mu_T(\hat{\psi}_\varepsilon) \big|_{\varepsilon=0^+} = \int_{\mathbb{R}^d}\min_{y\in\mathscr{Y}_T(x)}g(y)\,\mu_T(dx), \qquad \partial_\varepsilon\mu_T(\hat{\psi}_\varepsilon) \big|_{\varepsilon=0^-} = \int_{\mathbb{R}^d}\max_{y\in\mathscr{Y}_T(x)}g(y)\,\mu_T(dx).\]

Initial derivative. Define \(u_{\varepsilon,T}:=\log({\cal N}_T*e^{\hat{\phi}+\varepsilon g}), \;v_{\varepsilon,0}:={\mathbf{T}}_\beta^+[u_{\varepsilon,T}]\). By Lemma 7 (i), the minimizer \(\mathscr{Y}_0(x)\) is unique for \(\mu_0\)-a.e.\(x\), hence Danskin’s theorem gives \(\partial_\varepsilon v_{\varepsilon,0}(x)\big|_{\varepsilon=0} = \partial_\varepsilon u_{\varepsilon,T}(\mathscr{Y}_0(x))\big|_{\varepsilon=0} \;\mu_0\text{-a.e.\;}x\). Since \(g\) is bounded, differentiation under the Gaussian integral is justified and yields \(G(y):=\partial_\varepsilon u_{\varepsilon,T}(y)\big|_{\varepsilon=0} = \frac{{\cal N}_T*(g e^{\hat{\phi}})(y)}{h_T(y)}, \; |G(y)|\le |g|_\infty\). Also, \(e^{-|\varepsilon||g|_\infty}({\cal N}_T*e^{\hat{\phi}})(y) \le ({\cal N}_T*e^{\hat{\phi}+\varepsilon g})(y) \le e^{|\varepsilon||g|_\infty}({\cal N}_T*e^{\hat{\phi}})(y)\), so \(|u_{\varepsilon,T}(y)-u_{0,T}(y)|\le |\varepsilon||g|_\infty, \;y\in\mathbb{R}^d\). By monotonicity of \({\mathbf{T}}_\beta^+\), this implies \(|v_{\varepsilon,0}(x)-v_{0,0}(x)|\le |\varepsilon||g|_\infty, \;x\in\mathbb{R}^d\). Hence dominated convergence applied to the difference quotients gives \[\partial_\varepsilon\mu_0(v_{\varepsilon,0})\Big|_{\varepsilon=0} = \mu_0\big(G(\mathscr{Y}_0)\big) = m_0(G) = \nu_0\big({\cal N}_T*(g e^{\hat{\phi}})\big) = \nu_T\big(ge^{\hat{\phi}}\big) = m_T(g),\] by 61 and Tonelli’s theorem.

Conclusion. Divide 62 by \(\varepsilon\) and let \(\varepsilon\to0^\pm\). Using the derivative formulas above, we obtain, for every \(g\in C_b(\mathbb{R}^d)\), \[\label{eq:sandwich95final} \mu_T\big(\min_{\mathscr{Y}_T}g\big) \le m_T(g) \le \mu_T\big(\max_{\mathscr{Y}_T}g\big).\tag{63}\]

2. Let \(\Gamma:=\{(x,y):y\in\mathscr{Y}_T(x)\}\) be the graph of \(\mathscr{Y}\). Since \(m_T\ll \text{Leb}\) and \(\hat{\phi}\) is \(\beta\)-convex, \(\hat{\phi}\) is differentiable \(m_T\)-a.e. At every differentiability point \(y\) such that \((x,y)\in\Gamma\), then \(y\) minimizes \(z\mapsto \hat{\phi}(z)+\frac{\beta}{2}|x-z|^2\), so the first-order optimality condition yields \(\nabla\hat{\phi}(y)+\beta(y-x)=0\). Thus \(x=\mathscr{X}_T(y):=y+\beta^{-1}\nabla\hat{\phi}(y)\). Therefore, for \(m_T\)-a.e.\(y\), the section \(\Gamma^y:=\{x\in\mathbb{R}^d:(x,y)\in\Gamma\}\) is contained in the singleton \(\{\mathscr{X}_T(y)\}\). Our objective in this step is to show that \[\label{mu95T61Xc95diezmT} \mu_T={\mathscr{X}_T}_\# m_T.\tag{64}\] Since \(\hat{\phi}\) is l.s.c. the graph \(\Gamma:=\{(x,y):y\in\mathscr{Y}_T(x)\}\) is closed and \(\Gamma=\{c_\Gamma=0\}\), where \(c_\Gamma(x,y):=1\wedge \text{dist}((x,y),\Gamma)\) is bounded, continuous, and nonnegative. By Kantorovich duality, using the standard notation \(f\oplus h=f(x)+h(y)\), we have \[0\le \inf_{\pi\in\Pi(\mu_T,m_T)}\pi(c_\Gamma) = \sup_{(f,h)\in\mathscr{D}}\mu_T(f)+m_T(h), ~~\mathscr{D}:=\{(f,h)\in C_b\times C_b:\;f\oplus h\le c_\Gamma\big\}.\] For \((f,h)\in\mathscr{D}\), we have \(\max_{y\in\mathscr{Y}_T(x)} h(y)\le -f(x),\; x\in\mathbb{R}^d\). Applying the right-hand side inequality in 63 with \(g=h\), we get \(m_T(h) \le \int \max_{y\in\mathscr{Y}_T(x)}h(y)\;\mu_T(\mathrm{d}x) \le -\mu_T(f)\), hence \(\le 0\). By arbitrariness of \((f,h)\in\mathscr{D}\), we deduce that \(\inf_{\pi\in\Pi(\mu_T,m_T)}\int c_\Gamma\;d\pi=0\). Since the cost is bounded and continuous, an optimal coupling \(\pi^\star\in\Pi(\mu_T,m_T)\) exists and satisfies \(\pi^\star(\Gamma)=1\), as the minimum value is \(0\).

Disintegrate \(\pi^\star\) with respect to its second marginal \(m_T\): \(\pi^\star(\mathrm{d}x,\mathrm{d}y)=\pi^\star_y(\mathrm{d}x)\;m_T(\mathrm{d}y)\). Since \(\pi^\star(\Gamma)=1\), for \(m_T\)-a.e.\(y\) the measure \(\pi^\star_y\) is supported on \(\Gamma^y\), hence on \(\{\mathscr{X}_T(y)\}\). Thus \(\pi^\star_y=\delta_{\mathscr{X}_T(y)} \;\text{for }m_T\text{-a.e.\;}y\). Taking first marginals, we obtain 64 .

3. Let \(\mathbf{R}\) be the law of a Brownian motion \(Y\) on \([0,T]\) with initial law \(m_0\). Define \(Z_T:=e^{\hat{\phi}(Y_T)-\hat{u}_T(Y_0)}\). Then \(\mathbb{E}^{\mathbf{R}}[Z_T] = \mathbb{E}^{\mathbf{R}}\Big[ e^{-\hat{u}_T(Y_0)} \mathbb{E}^{\mathbf{R}}\big[e^{\hat{\phi}(Y_T)}\mid Y_0\big]\Big] =1\). We may therefore introduce the equivalent probability measure \(d\hat{\mathbb{Q}}:=Z_T\,d\mathbf{R}\). For any bounded Borel \(f\), we have \(\mathbb{E}^{\hat{\mathbb{Q}}}[f(Y_T)] =\mathbb{E}^{\mathbf{R}}\!\big[f(Y_T)e^{\hat{\phi}(Y_T)-\hat{u}_T(Y_0)}\big] =m_T(f)\). Thus \(Y_0\sim m_0 ~and~Y_T\sim m_T~under~\hat{\mathbb{Q}}\). The density process is \(Z_t:=\mathbb{E}^{\mathbf{R}}[Z_T|\mathcal{F}_t] = e^{\hat{u}_{T-t}(Y_t)-\hat{u}_T(Y_0)}, \; 0\le t<T\). As \((t,y)\mapsto \hat{u}_{T-t}(y)\) solves the backward heat equation, Itô’s formula gives \(dZ_t=Z_t\nabla\hat{u}_{T-t}(Y_t)\cdot dW_t^{\mathbf{R}}, \; t<T\). It follows from Girsanov’s theorem that, under \(\hat{\mathbb{Q}}\), \[\label{eq:Y95SDE95global} Y_t =Y_0+\int_0^t \nabla\hat{u}_{T-s}(Y_s)\,ds+W_t^{\hat{\mathbb{Q}}}, \qquad 0\le t<T.\tag{65}\] Define, for \(0\le t<T\), \(X_t:=Y_t+\frac{1}{\beta}\nabla\hat{u}_{T-t}(Y_t)\). Then \(X_0=\mathscr{X}_0(Y_0)\), so \(X_0\sim\mu_0\) under \(\hat{\mathbb{Q}}\), and \(X\) has continuous paths on \([0,T)\) by the regularity of \(\hat{u}\) on positive heat times. Moreover, for every \(\tau<T\), since \(T-s\in[T-\tau,T]\) for \(0\le s\le\tau\) and \(\nabla\hat{u}\) has linear growth on this heat-time strip by Lemma 7(iii), the BDG inequality and Gronwall’s lemma yield \[\label{eq:X95global95second95moment} \mathbb{E}^{\hat{\mathbb{Q}}}\!\Big[\sup_{0\le s\le \tau}|Y_s|^2\Big] + \mathbb{E}^{\hat{\mathbb{Q}}}\!\Big[\sup_{0\le s\le \tau}|X_s|^2\Big] <\infty, \qquad \tau<T.\tag{66}\]

5. Recall that \(\hat{v}_t=\mathbf{T}_\beta^+[\hat{u}_{T-t}], \;\mathscr{X}_t(y)=y+\frac{1}{\beta}\nabla\hat{u}_{T-t}(y), \;X_t=\mathscr{X}_t(Y_t), \quad t<T\). Applying Itô’s formula to \(X_t=\mathscr{X}_t(Y_t)\) on every compact subinterval of \([0,T)\), and using the equation satisfied by \(\hat{u}\), gives \(dX_t=a_t\,dt+\sigma_t\,dW_t, \; 0\le t<T\), where \[\label{eq:feedback95coeffs95global} \gamma_t=(a_t,\sigma_t), \qquad a_t=\nabla\hat{v}_t(X_t), \qquad \sigma_t=\big(I_d-\beta^{-1}D^2\hat{v}_t(X_t)\big)^{-1}.\tag{67}\] These are precisely the maximizers of the Hamiltonian \(H(\nabla\hat{v}_t,D^2\hat{v}_t)\). Fix \(n\ge1\), \(\tau<T\), and define \(\rho_n:=\tau\wedge\inf\{t\le \tau:\;|X_t|\ge n\}\). Fix also \(0<\varepsilon<\tau\). Since \(\hat{v}\in C^{1,2}(Q_T)\), Itô’s formula on \([\varepsilon\wedge\rho_n,\rho_n]\) gives \[\begin{align}\mathbb{E}^{\hat{\mathbb{Q}}}\big[\hat{v}_{\rho_n}(X_{\rho_n}) -\hat{v}_{\varepsilon\wedge\rho_n}(X_{\varepsilon\wedge\rho_n}) \big] &= \mathbb{E}^{\hat{\mathbb{Q}}}\int_{\varepsilon\wedge\rho_n}^{\rho_n}\!\!\! \Big(\partial_t\hat{v}+a_t\!\cdot\!\nabla\hat{v}+\tfrac12\sigma_t\sigma_t^\top\!:\!D^2\hat{v}\Big)(t,X_t)\mathrm{d}t \\ &=\mathbb{E}^{\hat{\mathbb{Q}}}\!\int_{\varepsilon\wedge\rho_n}^{\rho_n} c(\gamma_t)\;\mathrm{d}t. \label{eq:stopped95cost95identity95global} \end{align}\tag{68}\] Since \(\beta T>1\), choose \(\delta>0\) so small that \(\beta-(T-\delta)^{-1}>0\). Then for \(t\in[0,\delta]\) and \(x\) in a fixed compact set \(K\subset\mathbb{R}^d\), the function \(y\mapsto \hat{u}_{T-t}(y)+\frac{\beta}{2}|x-y|^2\) is uniformly strongly convex, hence has a unique minimizer. Moreover, a Taylor expansion at \(y=0\) shows that these minimizers remain in a compact ball \(B_R\), uniformly for \(t\in[0,\delta]\) and \(x\in K\). Since \(\hat{u}_{T-t}\to \hat{u}_T\) uniformly on \(B_R\) as \(t\downarrow0\), it follows that \[\hat{v}_t(x)=\inf_y\Big(\hat{u}_{T-t}(y)+\frac{\beta}{2}|x-y|^2\Big)\to \inf_y\Big(\hat{u}_T(y)+\frac{\beta}{2}|x-y|^2\Big)=\hat{v}_0(x), ~uniformly on~K.\]

Letting \(\varepsilon\downarrow0\), local uniform convergence of \(v_t\) to \(v_0\) near \(t=0\) and continuity of \(X\) give \(\hat{v}_{\varepsilon\wedge\rho_n}(X_{\varepsilon\wedge\rho_n})\to\hat{v}_0(X_0) \;\hat{\mathbb{Q}}\text{-a.s.}\) Moreover the convergence is dominated by an integrable random variable, since on \(\{\rho_n>0\}\) the argument stays in the compact set \([0,T-\frac{1}{n}]\times \overline{B}_n\), while on \(\{\rho_n=0\}\) the term is exactly \(\hat{v}_0(X_0)\in \mathbb{L}^1(\hat{\mathbb{Q}})\). We then deduce from 68 that \(\mathbb{E}^{\hat{\mathbb{Q}}}\!\big[\hat{v}_{\rho_n}(X_{\rho_n})\big] = \mathbb{E}^{\hat{\mathbb{Q}}}[\hat{v}_0(X_0)] + \mathbb{E}^{\hat{\mathbb{Q}}}\!\int_0^{\rho_n} c(\gamma_t)\;dt,\) and by combining the quadratic growth property of \(\hat{v}\) in Lemma 7 with 66 , we get by dominated and monotone convergence: \[\label{step5v0} \mathbb{E}^{\hat{\mathbb{Q}}}[\hat{v}_\tau(X_\tau)] = \mathbb{E}^{\hat{\mathbb{Q}}}[\hat{v}_0(X_0)] + \mathbb{E}^{\hat{\mathbb{Q}}}\!\Big[\int_0^\tau c(\gamma_t)\;\mathrm{d}t\Big] = \mu_0(v_0)+\mathbb{E}^{\hat{\mathbb{Q}}}\!\Big[\int_0^\tau c(\gamma_t)\;\mathrm{d}t\Big], \; \tau<T.\tag{69}\] In particular, \(\mathbb{E}^{\hat{\mathbb{Q}}}\!\int_0^\tau c(\gamma_t)\;\mathrm{d}t<\infty\) for all \(\tau<T\).

6. Define \(F_s(y):=\hat{u}_s(y)+\frac{\beta}{2}|y|^2,\; s>0,\) and \(F(y):=\hat{\phi}(y)+\frac{\beta}{2}|y|^2\). Recall that \(F_s\to F\) locally uniformly on \(\mathbb{R}^d\text{ as }s\downarrow0\) and \[\label{eq:X95as95gradient95global} \beta X_t=\nabla F_{T-t}(Y_t),\qquad t<T.\tag{70}\] Since \(Y\) has continuous paths and \(Y_T\sim m_T\), we have \(Y_t\to Y_T\) a.s.as \(t\uparrow T\). Because \(m_T\ll\text{Leb}\) and \(F\) is finite convex, \(F\) is differentiable at \(Y_T\) for \(\hat{\mathbb{Q}}\)-a.e.sample point.

Fix an arbitrary sequence \(t_n\uparrow T\). Set \(f_n:=F_{T-t_n}\), \(f:=F\), \(x_n:=Y_{t_n}\), and \(x:=Y_T\). Since \(f_n\to f\) locally uniformly, \(x_n\to x\) \(\hat{\mathbb{Q}}\)-a.s., each \(f_n\) is differentiable, and \(f\) is differentiable at \(x\) for \(\hat{\mathbb{Q}}\)-a.e.sample point, the standard stability result for gradients of convex functions yields \(\nabla F_{T-t_n}(Y_{t_n})\to \nabla F(Y_T)\), \(\hat{\mathbb{Q}}\)-a.s. Hence, by 70 , \(X_{t_n} = \beta^{-1}\nabla F_{T-t_n}(Y_{t_n}) \to \beta^{-1}\nabla F(Y_T) \; \hat{\mathbb{Q}}\)-a.s. By Step 4, for \(m_T\)-a.e.\(y\) we have \(\mathscr{X}_T(y)=\beta^{-1}\nabla F(y)\). Since \(Y_T\sim m_T\), it follows that \(X_{t_n}\to \mathscr{X}_T(Y_T) \; \hat{\mathbb{Q}}\)-a.s. Since the sequence \(t_n\uparrow T\) was arbitrary, we conclude that \(X_t\to X_T:=\mathscr{X}_T(Y_T) \; \text{as }t\uparrow T,\;\hat{\mathbb{Q}}\)-a.s. Thus \(X\) admits a continuous extension to \([0,T]\). As \(Y_T\sim m_T\) under \(\hat{\mathbb{Q}}\) and \(\mu_T={\mathscr{X}_T}_\# m_T\) by Step 4, it follows that \(\text{Law}_{\hat{\mathbb{Q}}}(X_T)=\mu_T\). Consequently \[\hat{\mathbb{P}}:=\text{Law}_{\hat{\mathbb{Q}}}(X) \in {\cal P}(\mu_0,\mu_T).\] 7. Fix \(\tau\in(0,T)\) and let \((\hat{\mathbb{P}}^\omega)_\omega\) be a regular conditional probability distribution of \(\hat{\mathbb{P}}\) given \(\mathcal{F}_\tau\). We first observe that for \(\hat{\mathbb{P}}\)-a.e.\(\omega\), the shifted canonical process on \([\tau,T]\) under \(\hat{\mathbb{P}}^\omega\) starts from \(X_\tau(\omega)\) and is a continuous semimartingale with absolutely continuous characteristics \((a_r,\sigma_r)_{r\in[\tau,T)}\). Therefore, for \(\hat{\mathbb{P}}\)-a.e.\(\omega\), the Step 4 estimate on the shifted interval \([\tau,T)\) with horizon \(T-\tau\) and initial point \(X_\tau(\omega)\) give \[\label{eq:conditional95shifted95moment95bound95final} \mathbb{E}^{\hat{\mathbb{P}}^\omega}\!\Big[\sup_{\tau\le s<T}|X_s|^2\Big]<\infty.\tag{71}\] Fix now \(t\in(\tau,T)\), \(n\ge 1\), and define \(\rho_n^{\tau,t}:=t\wedge \inf\{s\in[\tau,t]: |X_s|\ge n\}\). By the same Itô calculation as in Step 5, we get by conditioning with respect to \(\mathcal{F}_\tau\) that for \(\hat{\mathbb{P}}\)-a.e.\(\omega\), \[\begin{align} \hat{v}_\tau(X_\tau(\omega)) &= \mathbb{E}^{\hat{\mathbb{P}}^\omega}\!\big[\hat{v}_{\rho_n^{\tau,t}}(X_{\rho_n^{\tau,t}})\big] - \mathbb{E}^{\hat{\mathbb{P}}^\omega}\!\Big[\int_\tau^{\rho_n^{\tau,t}} c(\gamma_r)\;\mathrm{d}r\Big]. \\ &\longrightarrow \mathbb{E}^{\hat{\mathbb{P}}^\omega}[\hat{v}_t(X_t)] - \mathbb{E}^{\hat{\mathbb{P}}^\omega}\!\left[\int_\tau^t c(\gamma_r)\;\mathrm{d}r\right]~~ as~n\to\infty, \label{eq:conditional95strip95identity95final} \end{align}\tag{72}\] by dominated convergence, due to the quadratic bound \(|\hat{v}_r(x)|\le C_{\tau,t}(1+|x|^2), \; r\in[\tau,t],\;x\in\mathbb{R}^d\) in Step 5, together with \(\mathbb{E}^{\hat{\mathbb{P}}^\omega}\!\Big[\sup_{\tau\le s\le t}|X_s|^2\Big]<\infty \;\text{for }\hat{\mathbb{P}}\text{-a.e. }\omega\).

Next we let \(t\uparrow T\). Since Step 6 gives \(X_t\to X_T\) \(\hat{\mathbb{P}}^\omega\)-a.s., and \(\hat{v}_t\to\hat{\psi}\) locally uniformly as \(t\uparrow T\), we have \(\hat{v}_t(X_t)\to \hat{\psi}(X_T) \; \hat{\mathbb{P}}^\omega\)-a.s. Also, choosing \(y=0\) in the Moreau envelope yields \(\hat{v}_t(x)\le \hat{u}_{T-t}(0)+\frac{\beta}{2}|x|^2\). For fixed \(\tau<T\), the function \(t\mapsto \hat{u}_{T-t}(0)\) is bounded above on \(t\in[\tau,T)\), and 71 implies that \(\sup_{t\in[\tau,T)} \hat{v}_t(X_t)^+ \le C_\tau+\frac{\beta}{2}\sup_{\tau\le s<T}|X_s|^2 \in \mathbb{L}^1(\hat{\mathbb{P}}^\omega) \;\text{for }\hat{\mathbb{P}}\text{-a.e. }\omega\). Hence dominated convergence gives \(\mathbb{E}^{\hat{\mathbb{P}}^\omega}[\hat{v}_t(X_t)^+]\to \mathbb{E}^{\hat{\mathbb{P}}^\omega}[\hat{\psi}(X_T)^+]\). For the negative parts, Fatou’s lemma gives \(\mathbb{E}^{\hat{\mathbb{P}}^\omega}[\hat{\psi}(X_T)^-] \le \liminf_{t\uparrow T} \mathbb{E}^{\hat{\mathbb{P}}^\omega}[\hat{v}_t(X_t)^-]\). Combining with the monotone convergence on the nonnegative cost term, we then obtain from 72 that for \(\hat{\mathbb{P}}\)-a.e.\(\omega\), \(\hat{v}_\tau(X_\tau(\omega)) + \mathbb{E}^{\hat{\mathbb{P}}^\omega}\!\Big[\int_\tau^T c(\gamma_r)\;\mathrm{d}r\Big] \le \mathbb{E}^{\hat{\mathbb{P}}^\omega}[\hat{\psi}(X_T)]\), and then \[\mathbb{E}^{\hat{\mathbb{P}}}[\hat{v}_\tau(X_\tau)] + \mathbb{E}^{\hat{\mathbb{P}}}\!\Big[\int_\tau^T c(\gamma_r)\;\mathrm{d}r\Big] \le \mathbb{E}^{\hat{\mathbb{P}}}[\hat{\psi}(X_T)] = \mu_T(\hat{\psi}),\] by the tower property. By 69 , this implies \(\mathbb{E}^{\hat{\mathbb{P}}}\!\left[\int_0^T c(\gamma_r)\;\mathrm{d}r\right] \le \mu_T(\hat{\psi})-\mu_0(\hat{v}_0)=\mathfrak{J}(\hat{\phi})\). Since \(\hat{\mathbb{P}}\in{\cal P}(\mu_0,\mu_T)\), we have \(\mathbb{E}^{\hat{\mathbb{P}}}\!\left[\int_0^T c(\gamma_r)\,\mathrm{d}r\right] \ge {\rm SBB}(\mu_0,\mu_T) = \mathfrak{J}(\hat{\phi})\), where the last equality follows from the strong duality already proved and the optimality of \(\hat{\phi}\). Hence \(\mathbb{E}^{\hat{\mathbb{P}}}\!\left[\int_0^T c(\gamma_r)\;\mathrm{d}r\right]=\mathfrak{J}(\hat{\phi})=\mathrm{SBB}(\mu_0,\mu_T)\) by the strong duality of Theorem 3 established in Section 4.   \({\cal t}\)  \({\cal u}\)

References↩︎

[1]
C. Léonard, “A survey of the Schrödinger problem and some of its connections with optimal transport,” Dynamical Systems, vol. 34, no. 4, pp. 1533–1574, 2014.
[2]
Y. Chen, T. Georgiou, and M. Pavon, “Stochastic control liaisons: Richard SInkhorn meets Gaspard Monge on a Schrödinger bridge,” SIAM review, vol. 63, no. 2, 2021.
[3]
M. Nutz, “Introduction to entropic optimal transport,” in Lecture notes, columbia university, 2022.
[4]
A. Conze and P. Henry-Labordère, “Bass construction with multi-marginals: Lightspeed computation in a new local volatility model,” 2021.
[5]
J. Backhoff-Veraguas, W. Schachermayer, and B. Tschiderer, The Bass functional of martingale transport,” The Annals of Applied Probability, vol. 35, no. 6, 2025.
[6]
B. Acciaio, A. Marini, and G. Pammer, “Calibration of the bass local volatility model,” SIAM Journal of Financial Mathematics, vol. 16, no. 3, 2025.
[7]
J. Backhoff-Veraguas, M. Beiglböck, M. Huesmann, and S. Källblad, Martingale-Benamou-Brenier: A probabilistic perspective,” Annals of Probability, vol. 48, no. 5, pp. 2258–2289, 2020.
[8]
X. Tan and N. Touzi, “Optimal transportation under controlled stochastic dynamics,” The Annals of Probability, vol. 41, no. 5, pp. 3201–3240, 2013, doi: 10.1214/12-AOP797.
[9]
I. Guo and G. Loeper, “Path dependent optimal transport and model calibration on exotic derivatives,” The Annals of Applied Probability, vol. 31, no. 3, pp. 1232–1263, 2021.
[10]
M. Hasenbichler, G. Pammer, and S. Thonhauser, “A weak transport approach to the Schrödinger—Bass bridge,” 2026.
[11]
I. Gyöngy, “Mimicking the one-dimensional marginal distributions of processes having an itô differential,” Probability Theory and Related Fields, vol. 71, no. 4, pp. 501–516, 1986.
[12]
R. T. Rockafellar, Convex analysis, vol. 28. Princeton, NJ: Princeton University Press, 1970.

  1. QubeRT. Email: phl@hotmail.com↩︎

  2. BNPP and Monash University. Email: gregoire.loeper@bnpparibas.com↩︎

  3. LPSM, Sorbonne Université and Université Paris Cité. This author was supported by the BNP-Paribas Chair “Futures of Quantitative Finance”. Email: othmane.xx90@gmail.com↩︎

  4. Ecole Polytechnique, CMAP. This author is supported by the Chair “Financial Risks”, by FiME (Laboratory of Finance and Energy Markets), and the EDF–CACIB Chair “Finance and Sustainable Development”. Email: huyen.pham@polytechnique.edu↩︎

  5. NYU Tandon School of Engineering. This author is partially supported by NSF grant \(\#\)DMS-2508581. Email: nizar.touzi@nyu.edu↩︎

  6. During the final stage of the preparation of this paper, G. Pammer brought to our attention his work in [10] with M. Hasenbichler and S. Thonhauser, motivated by early presentations of the results reported here, in which they analyze the same problem through the lens of weak optimal transport.↩︎