Optimal drift optimizer for non-convex optimization


Abstract

We study a finite-horizon stochastic control criterion for non-convex optimization in which Brownian exploration is balanced against a quadratic control cost. Rather than emphasizing the classical Hopf–Cole representation, we isolate the exact drift selected by the criterion and reorganize it in a form adapted to optimization. The key object is the conditional terminal law of the optimal process. We show that this law is a Gibbs measure for a proximally penalized energy, yielding three exact representations of the drift: potential, averaged-gradient, and barycentric. We then analyze two asymptotic regimes relevant for optimization. As terminal time is approached, the drift recovers a scaled gradient-descent field. In the low-temperature regime, assuming a unique global minimizer, the conditional terminal law concentrates on it even in the presence of nonglobal local minima, and the drift converges to an affine attraction field toward it. In the nondegenerate case we also derive Laplace asymptotics for the drift, the value function, and the covariance of the conditional terminal law. Finally, we record a simple gradient-free discretization suggested by the barycentric formula.

Introduction We are interested in the global minimization problem \[x^* \in \operatorname*{argmin}_{x \in \mathrm{I\kern-0.21emR}^d} f(x), \qquad f_* = \min_{x \in \mathrm{I\kern-0.21emR}^d} f(x),\] for a possibly non-convex objective function \(f : \mathrm{I\kern-0.21emR}^d \to \mathrm{I\kern-0.21emR}\).

The starting point of the paper is the finite-horizon stochastic control criterion \[\label{criterion95intro} J^\lambda(u) = \mathbb{E}\bigg[ f(X_T) + \frac{\lambda}{2T}\int_0^T |u_t |^2 \,\mathrm{d}t \bigg], \quad X_t = x_0 + \int_0^t u_s \,\mathrm{d}s + \sqrt{2\beta}B_t, \quad u_s = \mathbf{u}(s,X_s).\tag{1}\] We use control, drift, and — when speaking from the optimization viewpoint — optimizer, for the process \(u\). In the deterministic warm-up below, \(u\) is an open-loop control depending only on time. In the stochastic problem, the relevant objects will be time-dependent Markov feedbacks of the form \(u_t = \mathbf{u}(t,X_t)\), which is standard terminology in finite-horizon control theory (see, e.g., [1][3]). Throughout the paper, \(u_s\) denotes the realized admissible control process, \(\mathbf{u}(t,x)\) denotes a Markov feedback vector field, while \(\mathbf{u}_\lambda^*\) will denote the optimal feedback selected by the finite-horizon criterion associated with \(J^\lambda\) (see Theorem 3 below). We keep the expression “optimal drift optimizer” as a convenient shorthand for the feedback selected by 1 , but always relative to this precise criterion. The paper does not claim a universal optimal algorithm for non-convex optimization.

From the viewpoint of stochastic control, the analytic core of 1 is classical. The logarithmic transform reducing the Hamilton-Jacobi-Bellman equation to a linear parabolic equation belongs to the linearly solvable control literature [4][6]. Closely related exponential reweightings of Brownian motion appear in reciprocal diffusion and Schrödinger-Föllmer constructions [7][12]. The same forward-backward structure also appears in mean-field games [13], [14]. On the PDE and optimization side, related kernels have recently been interpreted through Wasserstein proximal operators and Hamilton-Jacobi regularizations [15][18]. The present viewpoint should also be compared with Gaussian homotopies, consensus-based optimization, and gradient-free integration methods [19][22].

The closed form itself is therefore not the main contribution of the paper. The point of the present note is to isolate the optimization content of that classical formula in a short and coherent framework on \(\mathrm{I\kern-0.21emR}^d\). More precisely, the paper emphasizes the following facts.

  1. The criterion 1 fixes its own scales. It selects both the Gaussian variance \(2\beta(T-t)\) and the effective temperature \(\frac{2\lambda\beta}{T}\), instead of taking them as external schedules. This distinguishes the present construction from Gaussian homotopy and consensus-type methods, where the smoothing scale or the consensus temperature is prescribed by hand.

  2. The stochastic problem is best read as a feedback optimal control problem, and the law-level Fokker-Planck formulation explains naturally why a feedback field appears. The deterministic open-loop toy problem already contains the proximal mechanism behind the whole construction.

  3. The central object is the conditional terminal law of the optimal process. It is a Gibbs law for a penalized energy, and it simultaneously yields the three useful forms of the drift: potential, averaged-gradient, and barycentric.

  4. The low-temperature regime \(\lambda \to 0\) is the global-optimization regime. Under the sole assumption that the global minimizer is unique, with no exclusion of nonglobal local minima, the terminal law concentrates on that minimizer and the drift becomes an affine attraction toward it. The regime \(t \to T\) is the local consistency regime: it shows that the same drift reconnects with gradient descent near the end of the horizon. The two limits are generically non-commutative, reflecting the tension between global exploration and local exploitation.

The main text is written on \(\mathrm{I\kern-0.21emR}^d\) with an integrable Gibbs density. We do not mix a free diffusion on \(\mathrm{I\kern-0.21emR}^d\) with a hard state constraint encoded by setting \(f = +\infty\) outside a bounded set. The bounded-domain framework, namely reflected diffusions and no-flux boundary conditions, is mentioned as a perspective in Section [sec:further95comments95sec].

Before turning to the stochastic feedback problem, it is useful to begin with the deterministic open-loop toy model behind the whole construction.

Proposition 1 (Deterministic warm-up). Fix \(x_0 \in \mathrm{I\kern-0.21emR}^d\), \(T > 0\), and \(\lambda > 0\). Consider \[\inf \bigg\{ f(x(T)) + \frac{\lambda}{2T}\int_0^T |u(t) |^2 \,\mathrm{d}t \;\;\big\vert\;\;\dot{x}(t) = u(t),\;x(0) = x_0 \bigg\}.\] Then \[\label{optim95warmup} \inf_{u \in L^2(0,T;\mathrm{I\kern-0.21emR}^d)} \bigg\{ f(x(T)) + \frac{\lambda}{2T}\int_0^T |u(t) |^2 \,\mathrm{d}t \bigg\} = \inf_{y \in \mathrm{I\kern-0.21emR}^d} \bigg\{ f(y) + \frac{\lambda}{2T^2}|y - x_0 |^2 \bigg\}.\qquad{(1)}\] If the right-hand side admits a minimizer \(y_\lambda(x_0)\), then an optimal control is the constant control \[u_\lambda(t) = \frac{y_\lambda(x_0) - x_0}{T}, \qquad x_\lambda(t) = x_0 + \frac{t}{T}\big( y_\lambda(x_0) - x_0 \big).\]

Proof. Let \(y = x(T)\). Since \(y - x_0 = \int_0^T u(t)\,\mathrm{d}t\), the Cauchy-Schwarz inequality gives \[\int_0^T |u(t) |^2 \,\mathrm{d}t \geqslant\frac{1}{T}\bigg|\int_0^T u(t)\,\mathrm{d}t \bigg|^2 = \frac{1}{T}|y - x_0 |^2,\] with equality if and only if \(u\) is constant. Therefore, among all trajectories ending at \(y\), the minimal kinetic cost equals \(\frac{\lambda}{2T^2}|y - x_0 |^2\). Minimizing over \(y\) gives the result. ◻

Proposition 1 already shows what the quadratic running cost does: it turns the endpoint optimization into a proximal penalization. Equivalently, the right-hand side of ?? is the Moreau envelope of \(f\) with parameter \(\mu = T^2/\lambda\), evaluated at \(x_0\). The stochastic problem below is its noisy feedback analogue, with a terminal law instead of a single terminal point.

The same warm-up already contains the basic low-\(\lambda\) intuition.

Corollary 1 (Low-\(\lambda\) limit in the warm-up). Assume that \(f\) is continuous, bounded from below, coercive, and has a unique minimizer \(x^*\). Let \(y_\lambda(x_0)\) be any minimizer of the problem ?? . Then \(y_\lambda(x_0) \to x^*\) as \(\lambda \downarrow 0\). Consequently, \(u_\lambda(t) \to \frac{x^* - x_0}{T}\) for every \(t \in [0,T]\).

Proof. For every \(\lambda > 0\), \[f\big( y_\lambda(x_0) \big) + \frac{\lambda}{2T^2}|y_\lambda(x_0) - x_0 |^2 \leqslant f(x^*) + \frac{\lambda}{2T^2}|x^* - x_0 |^2.\] Hence \[f\big( y_\lambda(x_0) \big) \leqslant f_* + \frac{\lambda}{2T^2}|x^* - x_0 |^2.\] By coercivity, the family \(\{ y_\lambda(x_0) \}_{\lambda > 0}\) is bounded. Let \(\lambda_n \downarrow 0\) and assume \(y_{\lambda_n}(x_0) \to \bar y\) along a subsequence. Passing to the limit in the previous inequality gives \(f(\bar y) \leqslant f_*\). Hence \(\bar y = x^*\) by uniqueness of the minimizer. Every convergent subsequence has the same limit, so the whole family converges to \(x^*\). ◻

The deterministic warm-up is open-loop. The finite-horizon optimizer selected by 1 will instead be a time-dependent Markov feedback.

The rest of the paper is organized as follows. Section [sec:control95sec] solves the finite-horizon optimal control problem and identifies the exact terminal Gibbs law and the three equivalent forms of the drift. Section [sec:asymp95sec] contains the asymptotic analysis: the regime \(t \to T\) explains the local gradient behavior, while the regime \(\lambda \to 0\) gives the global-selection mechanism and the Laplace asymptotics. Section [sec:numerics95sec] records a gradient-free discretization. Section [sec:further95comments95sec] concludes with further comments and perspectives.

Optimal control and exact drift This section solves the finite-horizon optimal control problem associated with 1 , more generally with the translated criterion \(J^\lambda_{t,x}\) (defined by 2 further) starting from an arbitrary time-space point \((t,x)\). We record both the Hamilton-Jacobi-Bellman and the law-level Fokker-Planck formulations, and we identify the conditional terminal law that carries the whole drift. The closed form itself is classical; what matters here is the way it is reorganized for optimization.

0.1 Fokker-Planck formulation and Gibbs objects↩︎

We fix throughout two parameters \(T > 0\) and \(\beta > 0\).

Assumption 1. The objective function \(f : \mathrm{I\kern-0.21emR}^d \to \mathrm{I\kern-0.21emR}\) satisfies the following conditions:

  1. \(f \in C^2(\mathrm{I\kern-0.21emR}^d)\), \(f\) is bounded from below, and \(f(x) \to +\infty\) as \(|x |\to +\infty\);

  2. for every \(c > 0\), the functions \(e^{-cf}\), \(|\nabla f |e^{-cf}\), and1 \(\| D^2 f \| e^{-cf}\) belong to \(L^1(\mathrm{I\kern-0.21emR}^d)\).

Assumption 1 keeps the whole discussion on \(\mathrm{I\kern-0.21emR}^d\) and avoids artificial hard walls. It is satisfied, for instance, by coercive \(C^2\) functions whose first two derivatives have at most polynomial growth. The low-temperature concentration arguments of Section [sec:asymp95sec] use less than this, but these assumptions are convenient for the exact formulas. The bounded-domain case with reflection is mentioned as a perspective in Section [sec:further95comments95sec].

On a filtered probability space carrying a \(d\)-dimensional Brownian motion \(B\), an admissible control on \([t,T]\) from \(x\) is a progressively measurable process \((u_s)_{s \in [t,T]}\) with values in \(\mathrm{I\kern-0.21emR}^d\) such that \(\displaystyle \mathbb{E}\int_t^T |u_s |^2 \,\mathrm{d}s < +\infty\). Given such a control, we set \[X_s = x + \int_t^s u_r \,\mathrm{d}r + \sqrt{2\beta}(B_s - B_t), \qquad s \in [t,T],\] and \[\label{def95Jtx} J^\lambda_{t,x}(u) = \mathbb{E}\bigg[ f(X_T) + \frac{\lambda}{2T}\int_t^T |u_s |^2 \,\mathrm{d}s \bigg].\tag{2}\] A Markov control, also called a time-dependent feedback control in the finite-horizon setting, is a control of the special form \(u_s = \mathbf{u}(s,X_s)\) for some Borel map \(\mathbf{u} : [t,T] \times \mathrm{I\kern-0.21emR}^d \to \mathrm{I\kern-0.21emR}^d\) (see [1], [3]).

If the control is Markov and \(\mathbf{u}\) is regular enough, then the law \(\rho(s,\cdot)\) of \(X_s\) solves the Fokker-Planck equation \[\label{FP} \partial_s \rho + \nabla \cdot \big( \rho \, \mathbf{u} \big) = \beta \Delta \rho\tag{3}\] in the distributional sense; under smoothness assumptions the derivation is classical, see, for example, [23] or [24]. Consequently, the criterion \(J^\lambda_{t,x}\) may be rewritten at the level of densities as \[\label{FP95control} \inf_{\mathbf{u}} \bigg\{ \int_{\mathrm{I\kern-0.21emR}^d} f(y)\rho(T,y)\,\mathrm{d}y + \frac{\lambda}{2T}\int_t^T \int_{\mathrm{I\kern-0.21emR}^d} |\mathbf{u}(s,y) |^2 \rho(s,y)\,\mathrm{d}y\,\mathrm{d}s \bigg\},\tag{4}\] subject to the dynamical constraint 3 . This formulation is the natural feedback analogue of Proposition 1. Thus, starting from an initial density \(\rho(t,\cdot) = \rho_0\), one seeks a feedback vector field \(\mathbf{u}\) and an associated path of probability densities \(\rho(s,\cdot)\) solving 3 . The density is not an additional modeling assumption; it is the law of the controlled diffusion and therefore records the evolution of mass in configuration space. In the global-optimization regime described in Section [sec:asymp95sec], the terminal mass is expected to concentrate near \(x^*\). This is the sense in which the drift selected by 34 may be called an optimal drift optimizer.

0.2 Exact value function and optimal feedback↩︎

We begin with standard notation. For \(r > 0\), let \[G_r(z) = \frac{1}{(4\pi \beta r)^{d/2}} \exp \bigg( - \frac{|z |^2}{4\beta r} \bigg)\] denote the heat kernel with diffusivity \(\beta\). Fix \(\lambda > 0\) and let \[\pi_\lambda(y) = \frac{1}{M_\lambda} e^{-\alpha_\lambda f(y)}, \qquad \alpha_\lambda = \frac{T}{2\lambda\beta}, \quad M_\lambda = \int_{\mathrm{I\kern-0.21emR}^d} e^{-\alpha_\lambda f(y)}\,\mathrm{d}y.\] \(\pi_\lambda\) denotes the normalized Gibbs density associated with \(f\) at inverse temperature \(\alpha_\lambda\), or equivalently at effective temperature \(\frac{2\lambda\beta}{T}\); see, for example, [25]. In the present setting these quantities are not inserted by hand: they are selected by the Hopf-Cole transform of the HJB equation, as in the linearly solvable control literature [4][6], and they also arise from the static free-energy problem \[- \frac{2\lambda\beta}{T}\log M_\lambda = \inf \bigg\{ \int_{\mathrm{I\kern-0.21emR}^d} f(y)\rho(y)\,\mathrm{d}y + \frac{2\lambda\beta}{T}\int_{\mathrm{I\kern-0.21emR}^d} \rho(y)\log \rho(y)\,\mathrm{d}y \;\mid\;\rho \geqslant 0,\;\int_{\mathrm{I\kern-0.21emR}^d} \rho(y)\,\mathrm{d}y = 1 \bigg\},\] whose unique minimizer is \(\pi_\lambda\) (see [25]). In this sense \(\pi_\lambda\) is the Gibbs law that balances objective value and entropy at temperature \(\frac{2\lambda\beta}{T}\).

Finally, for \(0 \leqslant t < T\) and \(x \in \mathrm{I\kern-0.21emR}^d\), we define

\[\label{def95h95phi} h_\lambda(t,x) = (G_{T-t} * \pi_\lambda)(x), \qquad \Phi_\lambda(t,x) = - 2\beta \log h_\lambda(t,x),\tag{5}\] and \[\label{def95V95lambda} V_\lambda(t,x) = \frac{\lambda}{T}\Phi_\lambda(t,x) - \frac{2\lambda\beta}{T}\log M_\lambda.\tag{6}\] Equivalently, \[\label{direct95value95formula} V_\lambda(t,x) = - \frac{2\lambda\beta}{T}\log \int_{\mathrm{I\kern-0.21emR}^d} G_{T-t}(x-y)e^{-\alpha_\lambda f(y)}\,\mathrm{d}y.\tag{7}\]

Remark 2 (Role of \(T\) and \(\beta\)). Formula 7 makes the role of \(T\) completely explicit. The Gaussian factor has covariance \(2\beta(T-t)I_d\), so the typical radius of the terminal averaging at time \(t\) is \(\sqrt{2\beta(T-t)}\). This is the time-dependent search radius. At the same time, the Gibbs factor uses inverse temperature \(\alpha_\lambda = \frac{T}{2\lambda\beta}\). Hence increasing \(T\) simultaneously enlarges the spatial search region and sharpens the Gibbs concentration. With the present normalization, the limit \(T \to +\infty\) is therefore not a neutral passage to a larger horizon but a different optimal control problem. An infinite-horizon theory would require a differently scaled criterion, typically discounted or ergodic. By a time rescaling one may normalize \(\beta = 1\), but keeping \(\beta\) explicit makes both scales visible.

Theorem 3. Under Assumption 1, the function \(V_\lambda\) defined by 6 belongs to \(C^{1,2}([0,T) \times \mathrm{I\kern-0.21emR}^d) \cap C([0,T] \times \mathrm{I\kern-0.21emR}^d)\) and is the classical solution of the backward viscous eikonal equation \[\label{hjb} \partial_t V_\lambda = \frac{T}{2\lambda}|\nabla V_\lambda |^2 - \beta \Delta V_\lambda \quad \text{on } [0,T) \times \mathrm{I\kern-0.21emR}^d,\tag{8}\] with terminal condition \(V_\lambda(T,x) = f(x)\). Moreover, \[V_\lambda(t,x) = \inf_u J^\lambda_{t,x}(u),\] where the infimum is taken over all admissible controls. The optimal control is the Markov feedback \[\label{optimal95feedback95formula} \mathbf{u}_\lambda^*(t,x) = - \frac{T}{\lambda}\nabla V_\lambda(t,x) = - \nabla \Phi_\lambda(t,x) = 2\beta \nabla \log h_\lambda(t,x).\tag{9}\]

The notation of \(\mathbf{u}_\lambda^*\) as the optimal feedback will be used consistently below.

Proof. Since \(\pi_\lambda \in C^2(\mathrm{I\kern-0.21emR}^d) \cap L^1(\mathrm{I\kern-0.21emR}^d)\), the heat convolution \(h_\lambda\) belongs to \(C^\infty([0,T) \times \mathrm{I\kern-0.21emR}^d)\) and satisfies \[\label{h95lambda95heat} \partial_t h_\lambda + \beta \Delta h_\lambda = 0 \qquad \text{on } [0,T) \times \mathrm{I\kern-0.21emR}^d,\tag{10}\] with terminal trace \(h_\lambda(T,\cdot) = \pi_\lambda\). Hence \[\nabla V_\lambda = - \frac{2\lambda\beta}{T}\frac{\nabla h_\lambda}{h_\lambda}, \qquad \Delta V_\lambda = - \frac{2\lambda\beta}{T}\frac{\Delta h_\lambda}{h_\lambda} + \frac{2\lambda\beta}{T}\frac{|\nabla h_\lambda |^2}{h_\lambda^2}.\] Using 10 , we obtain 8 . Since \(h_\lambda(t,\cdot) \to \pi_\lambda\) locally uniformly as \(t \uparrow T\), the definition of \(V_\lambda\) yields \[V_\lambda(T,x) = - \frac{2\lambda\beta}{T}\log \pi_\lambda(x) - \frac{2\lambda\beta}{T}\log M_\lambda = f(x).\] The Hamiltonian of the finite-horizon optimal control problem is (see, e.g., [2], [3], [26]) \[H(x,p,u) = u \cdot p+ \frac{\lambda}{2T}|u |^2 .\] For each \((x,p)\) it is minimized at \(u = - \frac{T}{\lambda}p\), and the minimum value equals \(- \frac{T}{2\lambda}|p |^2\). Therefore the HJB equation is 8 , and the minimizing selector is 9 . Since \(V_\lambda\) is a classical solution and the minimizer is explicit, the verification theorem for finite-horizon controlled diffusions applies (see [1] or [3]), hence \(V_\lambda\) is the value function and 9 is an optimal feedback. ◻

Theorem 3 is classical in several neighboring literatures. The contribution of the present paper is not the existence of the closed form itself. What is specific here is the optimization reading of the same formula: the deterministic warm-up, the feedback interpretation through the Fokker-Planck equation and the Pontryagin maximum principle, the explicit conditional terminal law, the endogenous schedules set by the criterion, and the low-\(\lambda\) asymptotics stated in a form directly relevant for global optimization.

Remark 4 (Regularized potential). The potential \(\Phi_\lambda(t,\cdot) = - 2\beta \log \big( G_{T-t} * \pi_\lambda \big)\) is a heat-regularized Gibbs landscape. The Gaussian scale is \(\sqrt{2\beta(T-t)}\), so earlier times correspond to broader smoothing. Proposition 12 further shows that this regularized landscape recovers \(f\) as \(t \uparrow T\). This is the natural connection with Gaussian homotopies and Hamilton-Jacobi regularizations, with the important difference that the present variance schedule is selected by 1 rather than prescribed externally.

Remark 5 (Monotonicity in \(\lambda\)). The value function \(V_\lambda(t,x)\) is nondecreasing in \(\lambda\) for each fixed \((t,x)\). In particular, the two limiting values match the monotonicity: as \(\lambda \downarrow 0\) the value function decreases to \(f_*\) (Theorem 16), and as \(\lambda \to +\infty\) the control cost forces \(u \approx 0\), so \(V_\lambda(t,x) \uparrow (G_{T-t} * f)(x)\), the free-diffusion average of \(f\).

Remark 6. The same feedback law also emerges from an optimal control problem posed directly on the density. Consider 4 subject to 3 . Applying the Pontryagin maximum principle (see [2], [27]) on the infinite-dimensional state space of densities, with adjoint field \(p(s,y)\), gives the adjoint equation \[- \partial_s p = \mathbf{u} \cdot \nabla p + \beta \Delta p + \frac{\lambda}{2T}|\mathbf{u} |^2, \qquad p(T,\cdot) = f,\] and pointwise minimization of the Hamiltonian yields \(\mathbf{u} = - \frac{T}{\lambda}\nabla p\). Substituting back gives \[\partial_s p = \frac{T}{2\lambda}|\nabla p |^2 - \beta \Delta p,\] which is the backward Hamilton-Jacobi equation 8 . Above, we have used HJB and verification for the derivation in the present finite-dimensional diffusion setting, but the Pontryagin maximum principle could have been applied as well. Conceptually, it explains why the feedback formula is natural once the state variable is the whole density rather than a single trajectory.

0.3 Conditional terminal law and three equivalent formulas↩︎

The next step is to identify the optimal Markov process and its terminal conditional law. This is where the optimization mechanism becomes transparent.

0.3.0.1 The optimal process as a Doob transform.

Lemma 1. For \(0 \leqslant t < s \leqslant T\) define \[p_\lambda^*(t,x;s,y) = \frac{G_{s-t}(y-x) h_\lambda(s,y)}{h_\lambda(t,x)}.\] Then \(p_\lambda^*\) is a Markov transition kernel. The associated Markov process has infinitesimal generator \[\mathcal{L}_t^* \varphi(x) = \beta \Delta \varphi(x) + \mathbf{u}_\lambda^*(t,x) \cdot \nabla \varphi(x),\] where \(\mathbf{u}_\lambda^*\) is the feedback from Theorem 3. In particular, this process realizes the optimal controlled dynamics. At terminal time, we have \[\label{terminal95eta} p_\lambda^*(t,x;T,y) = \eta_{\lambda,t,x}(y), \qquad \eta_{\lambda,t,x}(y) = \frac{G_{T-t}(x-y)\pi_\lambda(y)}{h_\lambda(t,x)}.\tag{11}\]

Proof. The positivity of \(p_\lambda^*\) is obvious. Moreover, \[\begin{gather} \int_{\mathrm{I\kern-0.21emR}^d} p_\lambda^*(t,x;s,y)\,\mathrm{d}y = \frac{1}{h_\lambda(t,x)} \int_{\mathrm{I\kern-0.21emR}^d} G_{s-t}(y-x)h_\lambda(s,y)\,\mathrm{d}y \\ = \frac{1}{h_\lambda(t,x)} (G_{s-t} * h_\lambda(s,\cdot))(x) = \frac{1}{h_\lambda(t,x)} (G_{T-t} * \pi_\lambda)(x) = 1. \end{gather}\] Similarly, for \(t < r < s\), \[\begin{gather} \int_{\mathrm{I\kern-0.21emR}^d} p_\lambda^*(t,x;r,z)p_\lambda^*(r,z;s,y)\,\mathrm{d}z = \frac{h_\lambda(s,y)}{h_\lambda(t,x)} \int_{\mathrm{I\kern-0.21emR}^d} G_{r-t}(z-x)G_{s-r}(y-z)\,\mathrm{d}z \\ = \frac{G_{s-t}(y-x)h_\lambda(s,y)}{h_\lambda(t,x)} = p_\lambda^*(t,x;s,y). \end{gather}\] Hence \(p_\lambda^*\) is a Markov kernel.

Let \(P_{t,s}^*\) denote the corresponding semigroup: \[P_{t,s}^* \varphi(x) = \int_{\mathrm{I\kern-0.21emR}^d} \varphi(y)p_\lambda^*(t,x;s,y)\,\mathrm{d}y = \frac{1}{h_\lambda(t,x)} P_{s-t}\big( h_\lambda(s,\cdot)\varphi \big)(x),\] where \(P_r = G_r * \cdot\) is the heat semigroup. Differentiating at \(s = t\) and using 10 , we obtain \[\begin{gather} \lim_{s \downarrow t} \frac{P_{t,s}^* \varphi(x) - \varphi(x)}{s-t} = \frac{1}{h_\lambda(t,x)} \Big( \beta \Delta \big( h_\lambda(t,\cdot)\varphi \big)(x) + \partial_t h_\lambda(t,x)\varphi(x) \Big) \\ = \beta \Delta \varphi(x) + 2\beta \frac{\nabla h_\lambda(t,x)}{h_\lambda(t,x)} \cdot \nabla \varphi(x) = \beta \Delta \varphi(x) + \mathbf{u}_\lambda^*(t,x) \cdot \nabla \varphi(x). \end{gather}\] This gives the generator. Setting \(s = T\) yields 11 . ◻

0.3.0.2 Penalized Gibbs interpretation of the conditional terminal law.

The next proposition is the bridge with Proposition 1. It rewrites the conditional terminal law 11 as a Gibbs law on a penalized energy.

Proposition 7. For every \(0 \leqslant t < T\) and \(x \in \mathrm{I\kern-0.21emR}^d\), the density \(\eta_{\lambda,t,x}\) from 11 can be written as \[\label{penalized95eta} \eta_{\lambda,t,x}(y) = \frac{1}{Z_{\lambda,t,x}} \exp \big( - \alpha_\lambda \mathcal{E}_{\lambda,t,x}(y) \big),\qquad{(2)}\] where \[\label{penalized95energy} \mathcal{E}_{\lambda,t,x}(y) = f(y) + \frac{\lambda}{2T(T-t)}|y - x |^2,\qquad{(3)}\] and \(Z_{\lambda,t,x}\) is the normalizing constant \[Z_{\lambda,t,x} = (4\pi \beta (T-t))^{-d/2}\int_{\mathrm{I\kern-0.21emR}^d} \exp \big( - \alpha_\lambda \mathcal{E}_{\lambda,t,x}(z) \big)\,\mathrm{d}z.\]

Proof. Using the definitions of \(G_{T-t}\), \(\pi_\lambda\), and \(\alpha_\lambda\), \[\begin{align} G_{T-t}(x-y)\pi_\lambda(y) &= \frac{1}{(4\pi \beta (T-t))^{d/2}M_\lambda} \exp \bigg( - \frac{|x - y |^2}{4\beta(T-t)} - \alpha_\lambda f(y) \bigg) \\ &= \frac{1}{(4\pi \beta (T-t))^{d/2}M_\lambda} \exp \bigg( - \alpha_\lambda \Big( f(y) + \frac{\lambda}{2T(T-t)}|x - y |^2 \Big) \bigg). \end{align}\] Dividing by \(h_\lambda(t,x)\) yields ?? . ◻

Proposition 7 should be read literally. At the current state \((t,x)\), each candidate terminal point \(y\) receives a weight that rewards low objective values through \(f(y)\) and penalizes kinematic distance from the current position through the quadratic term \(\displaystyle \frac{\lambda}{2T(T-t)}|y - x |^2\). At the initial time \(t = 0\), the penalized energy in ?? is \[y \mapsto f(y) + \frac{\lambda}{2T^2}|y - x |^2,\] namely the same proximal energy as in Proposition 1. In that sense the stochastic problem is a Gibbs relaxation of the deterministic warm-up.

0.3.0.3 Three equivalent formulas for the optimal drift.

We next give three representations of the optimal feedback. The potential form is the Hamilton-Jacobi representation; the averaged-gradient and barycentric forms are the two conditional-expectation representations. The differential and one-sided Lipschitz structure of the barycenter is recorded separately afterwards, because it is a consequence of the barycentric representation rather than a fourth formula for the drift.

Proposition 8 (Three representations of the optimal drift). Under Assumption 1, the optimal feedback admits the following equivalent representations on \([0,T)\times\mathrm{I\kern-0.21emR}^d\). First, it is induced by the potential \[\label{potential95form} \mathbf{u}_\lambda^*(t,x) = - \nabla \Phi_\lambda(t,x), \qquad \Phi_\lambda(t,x) = -2\beta \log (G_{T-t} * \pi_\lambda)(x).\qquad{(4)}\] Second, it has the averaged-gradient form \[\label{gradient95and95barycenter} \mathbf{u}_\lambda^*(t,x) = - \frac{T}{\lambda}\int_{\mathrm{I\kern-0.21emR}^d} \nabla f(y)\,\eta_{\lambda,t,x}(y)\,\mathrm{d}y.\qquad{(5)}\] Third, it has the barycentric form \[\label{def95a95lambda} \mathbf{u}_\lambda^*(t,x) = - \frac{x - a_\lambda(t,x)}{T-t}, \qquad a_\lambda(t,x) = \int_{\mathrm{I\kern-0.21emR}^d} y\, \eta_{\lambda,t,x}(y)\,\mathrm{d}y.\qquad{(6)}\]

Proof. The potential form is just the definition of \(\Phi_\lambda\) together with \(\mathbf{u}_\lambda^* = 2\beta \nabla_x \log h_\lambda\) and \(h_\lambda = G_{T-t}*\pi_\lambda\). For the averaged-gradient form, we integrate by parts: \[\begin{align} \nabla_x h_\lambda(t,x) &=\int_{\mathrm{I\kern-0.21emR}^d} \nabla_x G_{T-t}(x-y)\pi_\lambda(y)\,\mathrm{d}y\\ &=- \int_{\mathrm{I\kern-0.21emR}^d} \nabla_y G_{T-t}(x-y)\pi_\lambda(y)\,\mathrm{d}y =\int_{\mathrm{I\kern-0.21emR}^d} G_{T-t}(x-y)\nabla_y \pi_\lambda(y)\,\mathrm{d}y. \end{align}\] Since \(\nabla \pi_\lambda = - \alpha_\lambda \pi_\lambda \nabla f\), we obtain \[\nabla_x h_\lambda(t,x) = - \alpha_\lambda \int_{\mathrm{I\kern-0.21emR}^d} G_{T-t}(x-y)\pi_\lambda(y)\nabla f(y)\,\mathrm{d}y.\] Multiplying by \(\frac{2\beta}{h_\lambda(t,x)}\) and using \(2\beta\alpha_\lambda = \frac{T}{\lambda}\) yields ?? . Finally, \[\nabla_x h_\lambda(t,x) = \int_{\mathrm{I\kern-0.21emR}^d} \nabla_x G_{T-t}(x-y)\pi_\lambda(y)\,\mathrm{d}y = - \frac{1}{2\beta(T-t)} \int_{\mathrm{I\kern-0.21emR}^d} (x-y)G_{T-t}(x-y)\pi_\lambda(y)\,\mathrm{d}y.\] Hence \[\mathbf{u}_\lambda^*(t,x) = 2\beta \frac{\nabla_x h_\lambda(t,x)}{h_\lambda(t,x)} = - \frac{x - a_\lambda(t,x)}{T-t},\] which proves ?? . ◻

0.3.0.4 Differential structure of the barycenter.

We now turn to the monotonicity and one-sided Lipschitz properties of \(a_\lambda(t,\cdot)\). If \(\eta\) is a probability density on \(\mathrm{I\kern-0.21emR}^d\), we denote by \(Y\) the canonical random variable with law \(\eta(y)\,\mathrm{d}y\), by \[\mathbb{E}_\eta[\psi(Y)] = \int_{\mathrm{I\kern-0.21emR}^d}\psi(y)\eta(y)\,\mathrm{d}y, \qquad m_\eta = \mathbb{E}_\eta[Y],\] and by \[\operatorname{Cov}_\eta(Y) = \mathbb{E}_\eta\big[(Y-m_\eta)(Y-m_\eta)^\top\big] = \int_{\mathrm{I\kern-0.21emR}^d}(y-m_\eta)(y-m_\eta)^\top\eta(y)\,\mathrm{d}y\] the covariance matrix. For \(\xi\in\mathrm{I\kern-0.21emR}^d\), \[\operatorname{Var}_\eta(\xi\cdot Y) = \mathbb{E}_\eta\big[(\xi\cdot(Y-m_\eta))^2\big] = \xi^\top\operatorname{Cov}_\eta(Y)\xi.\] When \(\eta=\eta_{\lambda,t,x}\), this convention gives \(m_\eta=a_\lambda(t,x)\).

Proposition 9 (Barycentric monotonicity and covariance-Hessian identity). Under Assumption 1, \(a_\lambda(t,\cdot)\) is \(C^1\) and satisfies the matrix identity \[\label{Ceta95factorization} D_xa_\lambda(t,x) = I_d-(T-t)D_x^2\Phi_\lambda(t,x) = \frac{1}{2\beta(T-t)}\operatorname{Cov}_{\eta_{\lambda,t,x}}(Y).\qquad{(7)}\] Here the covariance is computed under the law \(\eta_{\lambda,t,x}(y)\,\mathrm{d}y\). In particular, \(D_xa_\lambda(t,x)\) is symmetric nonnegative definite and \(a_\lambda(t,\cdot)\) is monotone.

Proof. Fix \((t,x)\) and let \(Y\) have law \(\eta_{\lambda,t,x}(y)\,\mathrm{d}y\). Writing \(\log \eta_{\lambda,t,x}(y) = -\frac{|x-y|^2}{4\beta(T-t)} + \log\pi_\lambda(y) - \log h_\lambda(t,x)\) and using \(\nabla_x \log h_\lambda(t,x) = -\frac{1}{2\beta}\nabla\Phi_\lambda(t,x) = \frac{a_\lambda(t,x) - x}{2\beta(T-t)}\), we obtain \[\nabla_x\log \eta_{\lambda,t,x}(y) = \frac{y-a_\lambda(t,x)}{2\beta(T-t)}.\] Therefore, for every direction \(\xi\in\mathrm{I\kern-0.21emR}^d\), \[\begin{gather} D_xa_\lambda(t,x)\xi = \int_{\mathrm{I\kern-0.21emR}^d} y\,\xi\cdot\nabla_x\log\eta_{\lambda,t,x}(y)\,\eta_{\lambda,t,x}(y)\,\mathrm{d}y \\ = \frac{1}{2\beta(T-t)}\int_{\mathrm{I\kern-0.21emR}^d} y\,\xi\cdot\big(y-a_\lambda(t,x)\big)\eta_{\lambda,t,x}(y)\,\mathrm{d}y = \frac{1}{2\beta(T-t)}\operatorname{Cov}_{\eta_{\lambda,t,x}}(Y)\xi. \end{gather}\] This proves the covariance identity. On the other hand, the barycentric formula and the potential representation give \(a_\lambda(t,x)=x-(T-t)\nabla\Phi_\lambda(t,x)\), hence \(D_xa_\lambda(t,x)=I_d-(T-t)D_x^2\Phi_\lambda(t,x)\). Since covariance matrices are nonnegative definite, \(D_xa_\lambda(t,x)\) is nonnegative definite. The monotonicity follows by integrating \(D_xa_\lambda\) along line segments. ◻

0.3.0.5 Factorized one-sided Lipschitz identities.

Let \(t<T\), \(x_1,x_2\in\mathrm{I\kern-0.21emR}^d\), and \(x(\theta)=x_1+\theta(x_2-x_1)\). We define the trace covariance along this segment by \[\label{eq:Ceta} C_{\eta_\lambda}(t,x_1,x_2) = \int_0^1 \operatorname{tr}\operatorname{Cov}_{\eta_{\lambda,t,x(\theta)}}(Y)\,\mathrm{d}\theta = \int_0^1\int_{\mathrm{I\kern-0.21emR}^d}|y-a_\lambda(t,x(\theta))|^2\eta_{\lambda,t,x(\theta)}(y)\,\mathrm{d}y\,\mathrm{d}\theta.\tag{12}\] Taking the trace in ?? gives, for every \(x\), \[\operatorname{tr}\operatorname{Cov}_{\eta_{\lambda,t,x}}(Y) = 2\beta(T-t)\big(d-(T-t)\Delta\Phi_\lambda(t,x)\big).\] After integration along the segment, this yields the exact factorization \[\label{Ceta95terminal95factorized} C_{\eta_\lambda}(t,x_1,x_2) = 2\beta(T-t)\int_0^1\big(d-(T-t)\Delta\Phi_\lambda(t,x(\theta))\big)\,\mathrm{d}\theta.\tag{13}\]

Corollary 2 (Directional, variance and trace forms of the one-sided Lipschitz bound). Let \(t<T\) and \(x_1,x_2\in \mathrm{I\kern-0.21emR}^d\). With \(x(\theta)=x_1+\theta(x_2-x_1)\), the directional identity \[\label{eq:OSL95directional} \big\langle a_\lambda(t,x_2)-a_\lambda(t,x_1),x_2-x_1\big\rangle = \int_0^1 (x_2-x_1)^\top \big(I_d-(T-t)D_x^2\Phi_\lambda(t,x(\theta))\big)(x_2-x_1)\,\mathrm{d}\theta\tag{14}\] holds. Equivalently, in covariance form, \[\label{eq:OSL95variance} \big\langle a_\lambda(t,x_2)-a_\lambda(t,x_1),x_2-x_1\big\rangle = \frac{1}{2\beta(T-t)}\int_0^1 \operatorname{Var}_{\eta_{\lambda,t,x(\theta)}}\big((x_2-x_1)\cdot Y\big)\,\mathrm{d}\theta,\tag{15}\] where, for each \(\theta\), the random variable \(Y\) has law \(\eta_{\lambda,t,x(\theta)}(y)\,\mathrm{d}y\). Moreover, \[\label{eq:OSLC} 0 \leqslant\big\langle a_\lambda(t,x_2)-a_\lambda(t,x_1),x_2-x_1\big\rangle \leqslant|x_2-x_1|^2\int_0^1\big(d-(T-t)\Delta\Phi_\lambda(t,x(\theta))\big)\,\mathrm{d}\theta.\tag{16}\]

Proof. Integrating ?? along the segment from \(x_1\) to \(x_2\) gives \[a_\lambda(t,x_2)-a_\lambda(t,x_1) = \int_0^1D_xa_\lambda(t,x(\theta))(x_2-x_1)\,\mathrm{d}\theta.\] Taking the scalar product with \(x_2-x_1\) and using the identity \(D_xa_\lambda(t,x)=I_d-(T-t)D_x^2\Phi_\lambda(t,x)\) gives 14 . Using instead \(D_xa_\lambda(t,x)=\frac{1}{2\beta(T-t)}\operatorname{Cov}_{\eta_{\lambda,t,x}}(Y)\) gives 15 , because \[(x_2-x_1)^\top\operatorname{Cov}_{\eta_{\lambda,t,x(\theta)}}(Y)(x_2-x_1) = \operatorname{Var}_{\eta_{\lambda,t,x(\theta)}}\big((x_2-x_1)\cdot Y\big).\] It remains to prove the trace bound. For every \(\theta\), the matrix \[A_\theta=I_d-(T-t)D_x^2\Phi_\lambda(t,x(\theta)) = \frac{1}{2\beta(T-t)}\operatorname{Cov}_{\eta_{\lambda,t,x(\theta)}}(Y)\] is symmetric nonnegative definite. Therefore, for every \(z\in\mathrm{I\kern-0.21emR}^d\), \(z^\top A_\theta z \leqslant|z|^2\operatorname{tr} A_\theta\). Applying this with \(z=x_2-x_1\) and using \(\operatorname{tr}A_\theta=d-(T-t)\Delta\Phi_\lambda(t,x(\theta))\) yields 16 after integration in \(\theta\). ◻

Remark 10 (Near-terminal behavior and Hamilton-Jacobi interpretation). We spell out the near-terminal content of the preceding identities. As \(t\uparrow T\), the conditional law \(\eta_{\lambda,t,x}\) localizes near \(x\) at scale \(\sqrt{T-t}\). Under Assumption 1, \(\pi_\lambda\) is \(C^2\) on \(\mathrm{I\kern-0.21emR}^d\) with locally bounded derivatives. The Gaussian semigroup \(G_\tau\,*\) is a smooth approximation of the identity, so \(G_{T-t}*\pi_\lambda \to \pi_\lambda\) in \(C^2_{\mathrm{loc}}(\mathrm{I\kern-0.21emR}^d)\) as \(t\uparrow T\). Since \(\pi_\lambda\) is positive and the logarithm is smooth on positive functions bounded away from zero on compact sets, this convergence also holds after applying \(-2\beta\log\). Therefore \[\label{terminal95hessian95phi} D_x^2\Phi_\lambda(t,\cdot) \longrightarrow -2\beta D_x^2\log\pi_\lambda = \frac{T}{\lambda}D^2f \qquadlocally uniformly ast\uparrow T.\tag{17}\] Inserting 17 into 14 gives, uniformly for \(x_1,x_2\) in compact sets, \[\begin{gather} \label{terminal95directional95osl} \big\langle a_\lambda(t,x_2)-a_\lambda(t,x_1),x_2-x_1\big\rangle \\ = |x_2-x_1|^2 - \frac{T(T-t)}{\lambda}\int_0^1 (x_2-x_1)^\top D^2f(x(\theta))(x_2-x_1)\,\mathrm{d}\theta + |x_2-x_1|^2\mathrm{o}(T-t). \end{gather}\tag{18}\] Thus the sharp directional coefficient tends to \(1\), consistently with the terminal identity \(a_\lambda(T,x)=x\). Similarly, 13 yields the trace expansion \[\label{Ceta95terminal95expansion} C_{\eta_\lambda}(t,x_1,x_2) = 2d\beta(T-t) - \frac{2\beta T}{\lambda}(T-t)^2\int_0^1\Delta f(x(\theta))\,\mathrm{d}\theta + \mathrm{o}\big((T-t)^2\big).\tag{19}\] Thus the directional coefficient in 14 converges to \(1\), whereas the trace majorant in 16 converges to \(d\). The trace bound is less sharp, but it is useful when only the scalar trace covariance \(C_{\eta_\lambda}\) is tracked.

The Hamilton-Jacobi interpretation is as follows. From ?? , \[I_d-(T-t)D_x^2\Phi_\lambda(t,x)\geqslant 0, \qquad D_x^2\Phi_\lambda(t,x)\leqslant\frac{1}{T-t}I_d.\] This is the usual one-sided semiconcavity barrier for a quadratic Hamilton-Jacobi flow; see, for example, [28][30]. It is an upper barrier, not an assertion that the Hessian actually blows up. Indeed, ?? is equivalently \[D_x^2\Phi_\lambda(t,x) = \frac{1}{T-t}I_d - \frac{1}{2\beta(T-t)^2}\operatorname{Cov}_{\eta_{\lambda,t,x}}(Y).\] Combining this identity with 17 gives, componentwise and locally uniformly in \(x\), \[\operatorname{Cov}_{\eta_{\lambda,t,x}}(Y_i,Y_j) = 2\beta(T-t)\delta_{ij} - \frac{2\beta T}{\lambda}(T-t)^2\partial_{ij}f(x) + \mathrm{o}\big((T-t)^2\big).\] Thus the diagonal covariance contains the leading heat-kernel term \(2\beta(T-t)\), which cancels the diagonal semiconcavity barrier \(\frac{1}{T-t}I_d\) in the Hessian formula; the off-diagonal covariance is only of order \((T-t)^2\), so no off-diagonal singularity is present.

From the geometric viewpoint of optimal control, the inviscid analogue would correspond to a Riccati equation for the Hessian along characteristics, whose finite-time blow-up marks the formation of caustics or conjugate points (see, e.g., [31], [32]). In the present viscous whole-space setting, the Cole-Hopf representation together with the covariance identity ?? shows directly that no such terminal singularity occurs.

Corollary 3 (Conditional terminal expectations). Let \(X^*\) denote the optimal Markov process from Lemma 1. Then, conditionally on \(X_t^* = x\), the terminal law of \(X_T^*\) is \(\eta_{\lambda,t,x}\). Consequently, \[a_\lambda(t,x) = \mathbb{E}\big[ X_T^* \,\vert\, X_t^* = x \big],\] and \[\mathbf{u}_\lambda^*(t,x) = - \frac{T}{\lambda}\mathbb{E}\big[ \nabla f(X_T^*) \,\vert\, X_t^* = x \big] = - \frac{x - \mathbb{E}\big[ X_T^* \,\vert\, X_t^* = x \big]}{T-t}.\]

Proof. The first statement is 11 . The formulas then follow from Proposition 8. ◻

Corollary 3 is the main message of the section. At time \(t\) and current position \(x\), the optimal drift is determined by a weighted family of candidate terminal points. The three formulas say the same thing in three different languages:

  1. Potential language. The optimal feedback is dictated as a heat/Hamilton-Jacobi regularization of the terminal landscape. Namely, \(\mathbf{u}_\lambda^* = - \nabla \Phi_\lambda\), where \(\Phi_\lambda\) is the negative logarithm of a heat-regularized Gibbs density, or equivalently, \(\mathbf{u}_\lambda^*\) is the negative gradient of \(\frac{T}{\lambda}V_\lambda\), up to an irrelevant additive constant in the potential. As \(t \uparrow T\), Proposition 12 gives \(V_\lambda(t,\cdot) \to f\) locally uniformly; equivalently, \(\frac{\lambda}{T}\Phi_\lambda(t,\cdot) - \frac{2\lambda\beta}{T}\log M_\lambda \to f\) locally uniformly.

    a

    b

    Figure 1: Visualization of \(\Phi_\lambda\) and the control \(\mathbf{u}_\lambda^* = -\nabla_x \Phi_\lambda(t,x)\) at different time \(t\). We choose \(f\) to be the 2-dimensional Ackley function. The parameters are set as \(\lambda = 1\). The color-contour represents the landscape of \(\Phi_\lambda(t, x)\) defined in 7 . The red star marks the target minimizer \(x^*\) of \(f\). The white arrows indicate the vector field. The sequence of plots suggests that \(\Phi_\lambda(t,x)\) is a strongly smoothed effective landscape when \(t\) is small, and that this smoothing decreases as \(t\) approaches \(T=1\)..

  2. Averaged-gradient language. The drift is the Gibbs-Gaussian average of the local gradients \(\nabla f(y)\). This weighted form of \(\mathbf{u}_\lambda^*\) involves the normalized weight \[\eta_{\lambda,t,x}(y) \propto \exp \bigg( - \frac{T}{2\lambda\beta}f(y) - \frac{|y - x |^2}{4\beta(T-t)} \bigg).\] The first exponential rewards small values of the objective. The second rewards terminal points compatible with the current position and the remaining time. As \(t \uparrow T\), the optimal feedback recovers a scaled negative gradient field; specifically, Proposition 12 gives \[\lim_{t \uparrow T} \mathbf{u}_\lambda^*(t,x) = -\frac{T}{\lambda}\nabla f(x),\] locally uniformly in \(x\). In the normalization \(T=1\), this becomes \(-\frac{1}{\lambda}\nabla f(x)\). This alignment near \(T=1\) is illustrated in Figure 2.

    a

    b

    Figure 2: Visualization of \((\Phi_\lambda, \mathbf{u}_\lambda^*)\) for \(t = 0.99\) and \((f, -\nabla f)\). They are visually almost indistinguishable. Here \(f\) is the two-dimensional Ackley function..

  3. Barycentric language. The drift points from the current state \(x\) toward the weighted average terminal point \(a_\lambda(t,x)\). The barycentric formula then says: compute the weighted average of all candidate terminal points and move from \(x\) toward that average. It is this gradient-free formulation of \(\mathbf{u}_\lambda^*\) that leads to the computational discretization recorded in Section [sec:numerics95sec] below. The barycentric form also carries the factorized OSL and semiconcavity structure of \(a_\lambda(t,\cdot)\) described in Proposition 9 and Corollary 2.

Two consequences are immediate. First, the drift may be evaluated either from function values alone through the barycentric formula or, when gradients are available, through the averaged-gradient formula. Second, unlike in Gaussian homotopy or consensus-based optimization, the variance \(2\beta(T-t)\) and the effective temperature \(\displaystyle \frac{2\lambda\beta}{T}\) are not chosen manually: they are selected by the control criterion itself.

Corollary 4 (Exact optimal initialization). If \(X_0\) has density \(h_\lambda(0,\cdot)\) and evolves under the optimal feedback \(\mathbf{u}_\lambda^*\), then the law of \(X_s\) is \(h_\lambda(s,\cdot)\) for every \(s \in [0,T]\). In particular, \(\operatorname{Law}(X_T) = \pi_\lambda(y)\,\mathrm{d}y\).

Proof. For \(s \in [0,T]\), the law at time \(s\) is \[\int_{\mathrm{I\kern-0.21emR}^d} p_\lambda^*(0,x;s,\cdot)h_\lambda(0,x)\,\mathrm{d}x = h_\lambda(s,\cdot)\int_{\mathrm{I\kern-0.21emR}^d} G_s(\cdot - x)\,\mathrm{d}x = h_\lambda(s,\cdot),\] by the definition of \(p_\lambda^*\) and the heat semigroup property. ◻

Remark 11 (Path-space reweighting). There is also an equivalent path-space formulation behind the same formulas. Let \(\mathbf{W}_{t,x}\) be Wiener measure on the trajectory space \(C([t,T],\mathrm{I\kern-0.21emR}^d)\) associated with the driftless diffusion \(X_s = x + \sqrt{2\beta}(B_s - B_t)\). Then the value function may be rewritten as \[V_\lambda(t,x) = \inf \left\{ \mathbb{E}_{\mathbf{Q}}[f(X_T)] + \frac{2\lambda\beta}{T}\operatorname{Ent}(\mathbf{Q}\vert \mathbf{W}_{t,x}) \;\;\big\vert\;\;\mathbf{Q}\in \mathcal{P}\big( C([t,T],\mathrm{I\kern-0.21emR}^d) \big) \right\},\] and the minimizer is the exponentially reweighted Brownian law \[\frac{\,\mathrm{d}\mathbf{Q}_{\lambda,t,x}^*}{\,\mathrm{d}\mathbf{W}_{t,x}} = \frac{\exp \big( - \alpha_\lambda f(X_T) \big)}{\mathbb{E}_{\mathbf{W}_{t,x}}\big[ \exp \big( - \alpha_\lambda f(X_T) \big) \big]}.\] This means that Brownian trajectories are reweighted according to their terminal value: paths ending at smaller values of \(f(X_T)\) receive larger weight. This is the bridge with the reciprocal-diffusion and Schrödinger-Föllmer literature [7][12].

Asymptotics and optimization meaning

This is the main optimization section. The low-\(\lambda\) regime is the global regime: the effective temperature \(\frac{2\lambda\beta}{T}\) tends to \(0\), the terminal Gibbs law concentrates on the global minimizer, and the drift becomes an affine attraction toward it. The regime \(t \to T\) is local: it shows that the same finite-horizon drift reconnects with a scaled gradient descent near terminal time. We present the local regime first because its proof is immediate, then the low-temperature regime, and finally the nondegenerate Laplace expansion.

0.4 Near-terminal regime↩︎

The limit \(t \to T\) explains the local meaning of the finite-horizon construction. The Gaussian factor in 11 shrinks to a Dirac mass, so the optimizer stops averaging over distant candidate terminal points and recovers local first-order information.

Proposition 12. Fix \(\lambda > 0\). Under Assumption 1, we have \[V_\lambda(t,x) \to f(x), \qquad \mathbf{u}_\lambda^*(t,x) \to - \frac{T}{\lambda}\nabla f(x),\] as \(t \uparrow T\), locally uniformly in \(x\).

Proof. Since \(h_\lambda(t,\cdot) = G_{T-t} * \pi_\lambda\), the family \(h_\lambda(t,\cdot)\) converges locally uniformly to \(\pi_\lambda\) as \(t \uparrow T\). Because \(\nabla \pi_\lambda \in L^1(\mathrm{I\kern-0.21emR}^d)\) by Assumption 1, we also have \(\nabla h_\lambda(t,\cdot) = G_{T-t} * \nabla \pi_\lambda \to \nabla \pi_\lambda\) locally uniformly. Hence \[V_\lambda(t,x) = - \frac{2\lambda\beta}{T}\log h_\lambda(t,x) - \frac{2\lambda\beta}{T}\log M_\lambda \longrightarrow - \frac{2\lambda\beta}{T}\log \pi_\lambda(x) - \frac{2\lambda\beta}{T}\log M_\lambda = f(x),\] locally uniformly. Likewise, \[\mathbf{u}_\lambda^*(t,x) = 2\beta \frac{\nabla h_\lambda(t,x)}{h_\lambda(t,x)} \longrightarrow 2\beta \frac{\nabla \pi_\lambda(x)}{\pi_\lambda(x)} = - \frac{T}{\lambda}\nabla f(x),\] locally uniformly. ◻

Thus, for fixed \(\lambda\), the finite-horizon drift always recovers a scaled local gradient descent direction at the end of the interval. The nonlocal part of the optimizer is therefore concentrated away from terminal time.

0.5 Low-temperature regime↩︎

We now turn to the main global-optimization regime. The effective temperature of the terminal Gibbs law is \(\displaystyle \frac{2\lambda\beta}{T}\), so the limit \(\lambda \downarrow 0\) is a low-temperature limit. This is the regime in which the terminal law collapses on the global minimizer.

Assumption 2. The function \(f\) has a unique minimizer \(x^* \in \mathrm{I\kern-0.21emR}^d\).

Remark 13 (Local minima are allowed). Assumption 2 requires only that the global minimizer be unique. The function \(f\) may still possess arbitrarily many nonglobal local minima and saddle points. The results below therefore describe a global selection mechanism for the exact finite-horizon optimal control criterion.

The first ingredient is the Laplace principle (see, for example, [25], [33]) for the partition function \(M_\lambda\).

Lemma 2. Under Assumption 1, we have \(- \frac{2\lambda\beta}{T}\log M_\lambda \to f_*\) as \(\lambda \downarrow 0\).

Proof. Recall that \(\alpha_\lambda = \frac{T}{2\lambda\beta}\). Since \(f \geqslant f_*\), for every \(\alpha_\lambda \geqslant 1\) we have \[M_\lambda = \int_{\mathrm{I\kern-0.21emR}^d} e^{-\alpha_\lambda f(y)}\,\mathrm{d}y = \int_{\mathrm{I\kern-0.21emR}^d} e^{-(\alpha_\lambda - 1)f(y)}e^{-f(y)}\,\mathrm{d}y \leqslant e^{-(\alpha_\lambda - 1)f_*}\int_{\mathrm{I\kern-0.21emR}^d} e^{-f(y)}\,\mathrm{d}y.\] Hence \[- \frac{1}{\alpha_\lambda}\log M_\lambda \geqslant f_* - \frac{1}{\alpha_\lambda}\bigg( \log \int_{\mathrm{I\kern-0.21emR}^d} e^{-f(y)}\,\mathrm{d}y + f_* \bigg),\] and the right-hand side tends to \(f_*\).

Conversely, fix \(\varepsilon > 0\). By continuity of \(f\) at any minimizer point \(x^*\), there exists \(\rho > 0\) such that \(f(y) \leqslant f_* + \varepsilon\) for every \(y \in B_\rho(x^*)\). Therefore \(M_\lambda \geqslant\vert B_\rho \vert e^{-\alpha_\lambda(f_* + \varepsilon)}\), which yields \[- \frac{1}{\alpha_\lambda}\log M_\lambda \leqslant f_* + \varepsilon - \frac{1}{\alpha_\lambda}\log \vert B_\rho \vert.\] Letting \(\lambda \downarrow 0\) and then \(\varepsilon \downarrow 0\) gives the conclusion. ◻

The next lemma strengthens the usual weak concentration statement by providing an exponential estimate for \(\pi_\lambda\).

Lemma 3. Under Assumptions 1 and 2, for every \(r > 0\) there exist constants \(C_r > 0\), \(c_r > 0\), and \(\lambda_r > 0\) such that \[\pi_\lambda\big( |y - x^* |\geqslant r \big) \leqslant C_r e^{-c_r/\lambda} \qquad \text{for every } \lambda \in (0,\lambda_r].\] In particular, \(\pi_\lambda(y)\,\mathrm{d}y \Rightarrow \delta_{x^*}\) (narrow convergence) as \(\lambda \downarrow 0\).

Proof. Fix \(r > 0\). Since \(x^*\) is the unique minimizer and \(f\) is continuous, \[m_r = \inf_{|y - x^* |\geqslant r} f(y) > f_*.\] Choose \(\varepsilon_r \in \big( 0, m_r - f_* \big)\) and then \(\rho_r \in (0,r)\) such that \(f(y) \leqslant f_* + \varepsilon_r\) for every \(|y - x^* |\leqslant\rho_r\). As in the proof of Lemma 2, \(M_\lambda \geqslant\vert B_{\rho_r} \vert e^{-\alpha_\lambda(f_* + \varepsilon_r)}\). For \(\lambda\) small enough we have \(\alpha_\lambda \geqslant 1\), and then \[\int_{|y - x^* |\geqslant r} e^{-\alpha_\lambda f(y)}\,\mathrm{d}y = \int_{|y - x^* |\geqslant r} e^{-(\alpha_\lambda - 1)f(y)}e^{-f(y)}\,\mathrm{d}y \leqslant e^{-(\alpha_\lambda - 1)m_r}\int_{\mathrm{I\kern-0.21emR}^d} e^{-f(y)}\,\mathrm{d}y.\] Therefore \[\pi_\lambda\big( |y - x^* |\geqslant r \big) \leqslant\frac{e^{-(\alpha_\lambda - 1)m_r}\int_{\mathrm{I\kern-0.21emR}^d} e^{-f(y)}\,\mathrm{d}y}{\vert B_{\rho_r} \vert e^{-\alpha_\lambda(f_* + \varepsilon_r)}} = \frac{e^{m_r}\int_{\mathrm{I\kern-0.21emR}^d} e^{-f(y)}\,\mathrm{d}y}{\vert B_{\rho_r} \vert} \exp \bigg( - \alpha_\lambda\big( m_r - f_* - \varepsilon_r \big) \bigg).\] Since \(\displaystyle \alpha_\lambda = \frac{T}{2\lambda\beta}\) and \(m_r - f_* - \varepsilon_r > 0\), this gives the desired estimate. ◻

To convert the concentration of \(\pi_\lambda\) into uniform statements in \((t,x)\) on compact subsets of \([0,T) \times \mathrm{I\kern-0.21emR}^d\), we use the following elementary testing lemma.

Lemma 4. Let \(\mu_n\) be a sequence of probability measures on \(\mathrm{I\kern-0.21emR}^d\) converging narrowly to \(\delta_{x^*}\). Let \(\mathscr F\) be an equibounded family of real-valued functions on \(\mathrm{I\kern-0.21emR}^d\) that is equicontinuous at \(x^*\). Then \[\sup_{\varphi \in \mathscr F} \bigg\vert \int_{\mathrm{I\kern-0.21emR}^d} \varphi(y)\mu_n(\,\mathrm{d}y) - \varphi(x^*) \bigg\vert \to 0 \qquad \text{as } n \to +\infty.\]

Proof. Fix \(\varepsilon > 0\). By equicontinuity at \(x^*\), there exists \(r > 0\) such that \[\sup_{\varphi \in \mathscr F} \sup_{|y - x^* |\leqslant r} \vert \varphi(y) - \varphi(x^*) \vert \leqslant\varepsilon.\] Let \(M = \sup_{\varphi \in \mathscr F} |\varphi |_{L^\infty}\). Then \[\sup_{\varphi \in \mathscr F} \bigg\vert \int_{\mathrm{I\kern-0.21emR}^d} \varphi(y)\mu_n(\,\mathrm{d}y) - \varphi(x^*) \bigg\vert \leqslant\sup_{\varphi \in \mathscr F} \int_{\mathrm{I\kern-0.21emR}^d} \vert \varphi(y) - \varphi(x^*) \vert \mu_n(\,\mathrm{d}y) \leqslant\varepsilon + 2M\mu_n\big( |y - x^* |> r \big).\] Since \(\mu_n\) converges narrowly to \(\delta_{x^*}\), the second term tends to \(0\). Letting \(\varepsilon \to 0\) concludes the proof. ◻

The next proposition makes the global-selection mechanism fully explicit at the level of the conditional terminal law.

Proposition 14. Under Assumptions 1 and 2, for every compact subset \(K \Subset [0,T) \times \mathrm{I\kern-0.21emR}^d\) and every \(r > 0\), there exist constants \(C_{K,r} > 0\), \(c_{K,r} > 0\), and \(\lambda_{K,r} > 0\) such that \[\sup_{(t,x) \in K} \eta_{\lambda,t,x}\big( |y - x^* |\geqslant r \big) \leqslant C_{K,r} e^{-c_{K,r}/\lambda} \qquad \text{for every } \lambda \in (0,\lambda_{K,r}].\]

Proof. Fix a compact set \(K \Subset [0,T) \times \mathrm{I\kern-0.21emR}^d\) and \(r > 0\). There exists \(\delta_K > 0\) such that \(T-t \geqslant\delta_K\) for every \((t,x) \in K\), and the \(x\)-projection of \(K\) is contained in some ball \(B_R(0)\). As above, let \[m_r = \inf_{|y - x^* |\geqslant r} f(y) > f_*.\] Choose \(\varepsilon_r \in \big( 0, m_r - f_* \big)\) and then \(\rho_r \in (0,r)\) such that \(f(y) \leqslant f_* + \varepsilon_r\) whenever \(|y - x^* |\leqslant\rho_r\). Write \[D_{\lambda,t,x} = \int_{\mathrm{I\kern-0.21emR}^d} G_{T-t}(x-y)e^{-\alpha_\lambda f(y)}\,\mathrm{d}y.\] Since \((t,x,y)\) ranges over the compact set \(\big\{ (t,x,y) \;\big\vert\;(t,x) \in K,\;|y - x^* |\leqslant\rho_r \big\}\), and \(T-t \geqslant\delta_K\), the heat kernel admits a positive lower bound there: \[m_{K,r} = \inf \big\{ G_{T-t}(x-y) \;\big\vert\;(t,x) \in K,\;|y - x^* |\leqslant\rho_r \big\} > 0.\] Hence \(D_{\lambda,t,x} \geqslant m_{K,r}\vert B_{\rho_r} \vert e^{-\alpha_\lambda(f_* + \varepsilon_r)}\) for all \((t,x) \in K\).

On the other hand, for \(\lambda\) small enough so that \(\alpha_\lambda \geqslant 1\), \[\begin{align} \int_{|y - x^* |\geqslant r} G_{T-t}(x-y)e^{-\alpha_\lambda f(y)}\,\mathrm{d}y &\leqslant (4\pi\beta\delta_K)^{-d/2}\int_{|y - x^* |\geqslant r} e^{-\alpha_\lambda f(y)}\,\mathrm{d}y \\ &\leqslant (4\pi\beta\delta_K)^{-d/2}e^{-(\alpha_\lambda - 1)m_r}\int_{\mathrm{I\kern-0.21emR}^d} e^{-f(y)}\,\mathrm{d}y. \end{align}\] Dividing by the lower bound for \(D_{\lambda,t,x}\) yields \[\eta_{\lambda,t,x}\big( |y - x^* |\geqslant r \big) \leqslant \frac{(4\pi\beta\delta_K)^{-d/2}e^{m_r}\int_{\mathrm{I\kern-0.21emR}^d} e^{-f(y)}\,\mathrm{d}y}{m_{K,r}\vert B_{\rho_r} \vert} \exp \left( - \alpha_\lambda \big( m_r - f_* - \varepsilon_r \big) \right),\] uniformly in \((t,x) \in K\). Since \(\displaystyle \alpha_\lambda = \frac{T}{2\lambda\beta}\), the conclusion follows. ◻

Theorem 15. Under Assumptions 1 and 2, the probability measures \(\eta_{\lambda,t,x}(y)\,\mathrm{d}y\) converge to \(\delta_{x^*}\) as \(\lambda \downarrow 0\) uniformly for \((t,x)\) in compact subsets of \([0,T) \times \mathrm{I\kern-0.21emR}^d\), i.e., \[\sup_{(t,x) \in K} \bigg\vert \int_{\mathrm{I\kern-0.21emR}^d} \varphi(y)\eta_{\lambda,t,x}(y)\,\mathrm{d}y - \varphi(x^*) \bigg\vert \to 0 \qquad \text{as } \lambda \downarrow 0,\] for every compact subset \(K \Subset [0,T) \times \mathrm{I\kern-0.21emR}^d\) and every bounded continuous function \(\varphi : \mathrm{I\kern-0.21emR}^d \to \mathrm{I\kern-0.21emR}\).

Proof. Fix \(K \Subset [0,T) \times \mathrm{I\kern-0.21emR}^d\), a bounded continuous function \(\varphi\), and \(\varepsilon > 0\). By continuity of \(\varphi\) at \(x^*\), there exists \(r > 0\) such that \(\vert \varphi(y) - \varphi(x^*) \vert \leqslant\varepsilon\) whenever \(|y - x^* |\leqslant r\). Then \[\begin{align} \sup_{(t,x) \in K} \bigg\vert \int_{\mathrm{I\kern-0.21emR}^d} \varphi(y)\eta_{\lambda,t,x}(y)\,\mathrm{d}y - \varphi(x^*) \bigg\vert &\leqslant \sup_{(t,x) \in K}\int_{\mathrm{I\kern-0.21emR}^d} \vert \varphi(y) - \varphi(x^*) \vert \eta_{\lambda,t,x}(y)\,\mathrm{d}y \\ &\leqslant \varepsilon + 2|\varphi |_{L^\infty} \sup_{(t,x) \in K}\eta_{\lambda,t,x}\big( |y - x^* |\geqslant r \big). \end{align}\] The second term tends to \(0\) by Proposition 14. Letting \(\varepsilon \downarrow 0\) concludes the proof. ◻

Theorem 15 is the global-optimization content of the finite-horizon criterion. Because Assumption 2 allows any number of nonglobal local minima, the theorem says that these local traps disappear in the low-temperature limit at the level of the exact continuous-time criterion: conditionally on the current state \((t,x)\), the terminal law still collapses onto the global minimizer.

We now pass from the conditional law to the drift field and the value function themselves.

Theorem 16. Under Assumptions 1 and 2, we have \[V_\lambda(t,x) \to f_*, \qquad a_\lambda(t,x) \to x^*, \qquad \mathbf{u}_\lambda^*(t,x) \to - \frac{x - x^*}{T-t},\] as \(\lambda \downarrow 0\), uniformly on any compact subset \(K \Subset [0,T) \times \mathrm{I\kern-0.21emR}^d\).

Proof. Since \(K \Subset [0,T) \times \mathrm{I\kern-0.21emR}^d\), there exists \(\delta_K > 0\) such that \(T-t \geqslant\delta_K\) for all \((t,x) \in K\), and the \(x\)-projection of \(K\) is compact. For \((t,x) \in K\), define \(g_{t,x}(y) = G_{T-t}(x-y)\) and \(b_{t,x}^i(y) = y_i G_{T-t}(x-y)\), for \(i = 1,\dots,d\). Since \(K\) is compact and \(T-t \geqslant\delta_K\), the families \(\mathscr G_K = \big\{ g_{t,x} \;\big\vert\;(t,x) \in K \big\}\) and \(\mathscr B_K^i = \big\{ b_{t,x}^i \;\big\vert\;(t,x) \in K \big\}\) are equibounded and equicontinuous on \(\mathrm{I\kern-0.21emR}^d\). By Lemmas 3 and 4, \[\label{uniform95h} \sup_{(t,x) \in K} \vert h_\lambda(t,x) - G_{T-t}(x-x^*) \vert \to 0,\tag{20}\] and, for each \(i = 1,\dots,d\), \[\label{uniform95num} \sup_{(t,x) \in K} \bigg\vert \int_{\mathrm{I\kern-0.21emR}^d} y_i G_{T-t}(x-y)\pi_\lambda(y)\,\mathrm{d}y - x_i^* G_{T-t}(x-x^*) \bigg\vert \to 0.\tag{21}\] Since \(G_{T-t}(x-x^*)\) is strictly positive on \(K\), 20 and 21 imply \[\label{eq:alambda-limit} \sup_{(t,x) \in K} |a_\lambda(t,x) - x^* |\to 0.\tag{22}\] The barycentric representation ?? then gives \[\sup_{(t,x) \in K} \bigg|\mathbf{u}_\lambda^*(t,x) + \frac{x - x^*}{T-t} \bigg|\to 0.\] Finally, \(\displaystyle V_\lambda(t,x) = - \frac{2\lambda\beta}{T}\log h_\lambda(t,x) - \frac{2\lambda\beta}{T}\log M_\lambda\). By Lemma 2, \(\displaystyle - \frac{2\lambda\beta}{T}\log M_\lambda \to f_*\). By 20 , \(h_\lambda(t,x)\) converges uniformly on \(K\) to the strictly positive function \(G_{T-t}(x-x^*)\), and therefore \(\log h_\lambda\) is uniformly bounded on \(K\) for all sufficiently small \(\lambda\). Since \(\displaystyle \frac{2\lambda\beta}{T} \to 0\), we conclude that \[\sup_{(t,x) \in K} \bigg\vert \frac{2\lambda\beta}{T}\log h_\lambda(t,x) \bigg\vert \to 0.\] Hence \(V_\lambda \to f_*\) uniformly on \(K\). ◻

Theorem 16 should be read as the field-level counterpart of Theorem 15. Away from the terminal singularity, the exact finite-horizon drift converges uniformly on compact sets to the affine field \(\displaystyle x \mapsto - \frac{x - x^*}{T-t}\), which points toward the global minimizer. This is the stochastic feedback analogue of Corollary 1.

Remark 17 (Non-commutativity of the two limits). The two asymptotic regimes of this section do not commute. For fixed \(\lambda > 0\), Proposition 12 gives \(\mathbf{u}_\lambda^*(t,x) \to - \frac{T}{\lambda}\nabla f(x)\) as \(t \uparrow T\). For fixed \(t < T\), Theorem 16 gives \(\displaystyle \mathbf{u}_\lambda^*(t,x) \to - \frac{x - x^*}{T-t}\) as \(\lambda \downarrow 0\). At a point \(x \neq x^*\) where \(\nabla f(x) \neq 0\), these two limits are generically different and point in different directions. At the global minimizer \(x = x^*\), both limits vanish. In other words, the global mechanism (low \(\lambda\)) and the local mechanism (near terminal time) can compete, and the behavior of the drift when \(\lambda \downarrow 0\) and \(t \uparrow T\) simultaneously depends on the relative rates of the two limits.

We now translate these pointwise statements into assertions about the optimally controlled process itself.

Corollary 5. Let \(X^\lambda\) be the optimally controlled process, started from an initial law \(\rho_0\) that does not depend on \(\lambda\) and is supported in a compact set \(K_0 \subset \mathrm{I\kern-0.21emR}^d\). Under Assumptions 1 and 2, we have \(\operatorname{Law}(X_T^\lambda) \Rightarrow \delta_{x^*}\) as \(\lambda \downarrow 0\), i.e., \(\mathbb{E}\big[ \varphi(X_T^\lambda) \big] \to \varphi(x^*)\) for any bounded continuous function \(\varphi\).

Proof. By conditioning on \(X_0^\lambda = x\) and using Lemma 1, \(\operatorname{Law}(X_T^\lambda)(\,\mathrm{d}y) = \int_{K_0} \eta_{\lambda,0,x}(y)\rho_0(\,\mathrm{d}x)\). Hence, for every bounded continuous \(\varphi\), \[\bigg\vert \mathbb{E}\big[ \varphi(X_T^\lambda) \big] - \varphi(x^*) \bigg\vert \leqslant\sup_{x \in K_0} \bigg\vert \int_{\mathrm{I\kern-0.21emR}^d} \varphi(y)\eta_{\lambda,0,x}(y)\,\mathrm{d}y - \varphi(x^*) \bigg\vert.\] The right-hand side tends to \(0\) by Theorem 15, applied to the compact set \(\{0\} \times K_0\). ◻

Corollary 6. Let \(X^\lambda\) be the optimal process started from the exact density \(h_\lambda(0,\cdot)\). Under Assumptions 1 and 2, we have \(\operatorname{Law}(X_T^\lambda) = \pi_\lambda(y)\,\mathrm{d}y\) and \[\mathbb{E}\big[ f(X_T^\lambda) \big] = \int_{\mathrm{I\kern-0.21emR}^d} f(y)\pi_\lambda(y)\,\mathrm{d}y \to f_* \qquad \text{as } \lambda \downarrow 0.\]

Proof. The identity for the terminal law is given by Corollary 4. Since \(f \geqslant f_*\), it remains to prove the upper bound. Fix \(\varepsilon > 0\). By continuity of \(f\) at \(x^*\), there exists \(r > 0\) such that \(f(y) \leqslant f_* + \varepsilon\) whenever \(|y - x^* |\leqslant r\). Then \[\begin{align} \int_{\mathrm{I\kern-0.21emR}^d} f(y)\pi_\lambda(y)\,\mathrm{d}y &\leqslant(f_* + \varepsilon)\pi_\lambda\big( |y - x^* |\leqslant r \big) + \int_{|y - x^* |> r} f(y)\pi_\lambda(y)\,\mathrm{d}y \\ &\leqslant f_* + \varepsilon + \frac{1}{M_\lambda}\int_{|y - x^* |> r} \vert f(y) \vert e^{-\alpha_\lambda f(y)}\,\mathrm{d}y. \end{align}\] By coercivity, \(\vert f \vert e^{-f} \in L^1(\mathrm{I\kern-0.21emR}^d)\). Setting \(m_r = \inf_{|y - x^* |> r} f(y)\), we have \(m_r > f_*\). Arguing as in the proof of Lemma 3, there exists \(\rho_r > 0\) such that \(M_\lambda \geqslant\vert B_{\rho_r} \vert e^{-\alpha_\lambda(f_* + \varepsilon)}\), and therefore \[\frac{1}{M_\lambda}\int_{|y - x^* |> r} \vert f(y) \vert e^{-\alpha_\lambda f(y)}\,\mathrm{d}y \leqslant\frac{e^{-(\alpha_\lambda - 1)m_r}}{\vert B_{\rho_r} \vert e^{-\alpha_\lambda(f_* + \varepsilon)}}\int_{\mathrm{I\kern-0.21emR}^d} \vert f(y) \vert e^{-f(y)}\,\mathrm{d}y \to 0.\] Thus \[\limsup_{\lambda \downarrow 0} \int_{\mathrm{I\kern-0.21emR}^d} f(y)\pi_\lambda(y)\,\mathrm{d}y \leqslant f_* + \varepsilon.\] Since \(\varepsilon > 0\) is arbitrary and the reverse inequality is trivial, the result follows. ◻

Corollaries 5 and 6 give the process-level meaning of the low-temperature limit. The exact continuous-time optimizer selected by 1 drives the terminal law toward the global minimizer as \(\lambda \downarrow 0\); for the exact self-consistent initialization, even the expected terminal objective converges to the global minimum value.

0.6 Nondegenerate case and Laplace asymptotics↩︎

The previous results use only low-temperature concentration. When the minimizer is nondegenerate, one can sharpen them by a leading-order Laplace expansion.

Assumption 3. In addition to Assumptions 1 and 2, \(f \in C^4\) in a neighborhood of \(x^*\) and \(H_* = D^2 f(x^*)\) is positive definite.

Proposition 18. Set \(\displaystyle C_* = \frac{(2\pi)^{d/2}}{\sqrt{\det H_*}}\) and recall \(\displaystyle\alpha_\lambda = \frac{T}{2\lambda\beta}\). Under Assumption 3, we have \[\begin{align} \displaystyle \int_{\mathrm{I\kern-0.21emR}^d} G_{T-t}(x-y)e^{-\alpha_\lambda f(y)}\,\mathrm{d}y &= \displaystyle e^{-\alpha_\lambda f_*}\alpha_\lambda^{-d/2} C_* G_{T-t}(x-x^*)\big( 1 + \mathrm{O}(\lambda) \big), \label{laplace95integral} \\ \displaystyle h_\lambda(t,x) &= \displaystyle G_{T-t}(x-x^*)\big( 1 + \mathrm{O}(\lambda) \big), \label{laplace95h} \\[2mm] \displaystyle a_\lambda(t,x) &= \displaystyle x^* + \mathrm{O}(\lambda), \label{laplace95a} \\ \displaystyle \mathbf{u}_\lambda^*(t,x) &= \displaystyle - \frac{x - x^*}{T-t} + \mathrm{O}(\lambda), \label{laplace95u} \\ \displaystyle V_\lambda(t,x) &= \displaystyle f_* + \frac{2\lambda\beta}{T}\bigg[ \frac{d}{2}\log \alpha_\lambda - \log \Big( C_* G_{T-t}(x-x^*) \Big) \bigg] + \mathrm{O}(\lambda^2), \label{laplace95V} \end{align}\] {#eq: sublabel=eq:laplace95integral,eq:laplace95h,eq:laplace95a,eq:laplace95u,eq:laplace95V} as \(\lambda \downarrow 0\), uniformly on any compact subset \(K \Subset [0,T) \times \mathrm{I\kern-0.21emR}^d\).

Proof. Set \(g_{t,x}(y) = G_{T-t}(x-y)\). Since \(K \Subset [0,T) \times \mathrm{I\kern-0.21emR}^d\), the family \(\{ g_{t,x} \}_{(t,x) \in K}\) is uniformly bounded in \(C^2(U)\) on some fixed neighborhood \(U\) of \(x^*\). The multidimensional Laplace method with parameter-dependent amplitudes (see, for example, [34] and [35]) therefore yields, uniformly on \(K\), \[\int_{\mathrm{I\kern-0.21emR}^d} g_{t,x}(y)e^{-\alpha_\lambda f(y)}\,\mathrm{d}y = e^{-\alpha_\lambda f_*}\alpha_\lambda^{-d/2} C_* \Big( g_{t,x}(x^*) + \mathrm{O}(\alpha_\lambda^{-1}) \Big).\] Since \(g_{t,x}(x^*) = G_{T-t}(x-x^*)\) and \(\alpha_\lambda^{-1} = \frac{2\lambda\beta}{T}\), this proves ?? . Applying the same expansion with amplitude identically equal to \(1\) gives \(M_\lambda = e^{-\alpha_\lambda f_*}\alpha_\lambda^{-d/2} C_* \big( 1 + \mathrm{O}(\lambda) \big)\). Dividing ?? by this expression yields ?? . Applying the same argument once more with amplitudes \(g_{t,x}^{(i)}(y) = y_i\, G_{T-t}(x-y)\) gives \[\int_{\mathrm{I\kern-0.21emR}^d} y_i\, G_{T-t}(x-y)e^{-\alpha_\lambda f(y)}\,\mathrm{d}y = e^{-\alpha_\lambda f_*}\alpha_\lambda^{-d/2} C_* \Big( x_i^* G_{T-t}(x-x^*) + \mathrm{O}(\lambda) \Big),\] uniformly on \(K\). Dividing by the denominator in ?? gives ?? . The barycentric formula from Proposition 8 then yields ?? . Finally, using ?? , \[\begin{align} V_\lambda(t,x) &= - \frac{1}{\alpha_\lambda}\log \bigg( e^{-\alpha_\lambda f_*}\alpha_\lambda^{-d/2} C_* G_{T-t}(x-x^*)\big( 1 + \mathrm{O}(\lambda) \big) \bigg) \\ &= f_* + \frac{1}{\alpha_\lambda}\bigg[ \frac{d}{2}\log \alpha_\lambda - \log \Big( C_* G_{T-t}(x-x^*) \Big) \bigg] - \frac{1}{\alpha_\lambda}\log \big( 1 + \mathrm{O}(\lambda) \big). \end{align}\] Since \(\displaystyle \alpha_\lambda^{-1} = \frac{2\lambda\beta}{T}\) and \(\log \big( 1 + \mathrm{O}(\lambda) \big) = \mathrm{O}(\lambda)\), the last term is \(\mathrm{O}(\lambda^2)\), which proves ?? . ◻

Proposition 18 sharpens the low-temperature theory in two ways. First, the drift correction is of order \(\lambda\) on compact subsets away from terminal time. Second, the value function has an explicit first correction of size \(\lambda \log \big( \frac{1}{\lambda} \big)\) coming from the factor \(\alpha_\lambda^{-d/2}\). This is substantially more informative than a mere convergence statement.

Corollary 7 (Asymptotic covariance of the conditional terminal law). Under Assumption 3, we have \[\operatorname{Cov}_{\eta_\lambda}(t,x) = \frac{2\lambda\beta}{T}H_*^{-1} + \mathrm{O}(\lambda^2),\] as \(\lambda \downarrow 0\), uniformly for \((t,x) \in K\), for any compact subset \(K \Subset [0,T) \times \mathrm{I\kern-0.21emR}^d\), where \[\label{eq:covariance} \operatorname{Cov}_{\eta_\lambda}(t,x)= \int_{\mathrm{I\kern-0.21emR}^d} (y - a_\lambda(t,x))(y - a_\lambda(t,x))^\top \eta_{\lambda,t,x}(y)\,\mathrm{d}y.\tag{23}\]

Proof. The penalized energy ?? satisfies \(\displaystyle D_y^2 \mathcal{E}_{\lambda,t,x}(x^*) = H_* + \frac{\lambda}{T(T-t)}I_d\). Since \(\eta_{\lambda,t,x}\) is a Gibbs law at inverse temperature \(\alpha_\lambda\) for this energy, the standard Laplace argument for second moments (see, for example, [34]) gives \[\operatorname{Cov}_{\eta_\lambda}(t,x) = \frac{1}{\alpha_\lambda}\bigg( H_* + \frac{\lambda}{T(T-t)}I_d \bigg)^{-1} + \mathrm{O}(\alpha_\lambda^{-2}).\] Since \(\displaystyle \alpha_\lambda^{-1} = \frac{2\lambda\beta}{T}\), expanding the inverse matrix to first order gives the result uniformly on \(K\), because \(T-t \geqslant\delta_K > 0\) on \(K\). ◻

Thus, in the nondegenerate regime, the conditional terminal law is approximately Gaussian with mean \(x^* + \mathrm{O}(\lambda)\) and covariance \((2\lambda\beta/T)H_*^{-1}\). The directions in which \(f\) is flatter at \(x^*\) (small eigenvalues of \(H_*\)) produce larger spread in the terminal law, as expected.

0.6.0.1 Low-temperature behavior of the OSL coefficient.

As a direct consequence of Corollary 7, we obtain the low-temperature behavior of the trace coefficient defined in 12 . For \(t\) in a compact subset of \([0,T)\) and \((x_1,x_2)\) in a compact subset of \(\mathrm{I\kern-0.21emR}^d\times\mathrm{I\kern-0.21emR}^d\), the segment \(x(\theta)=x_1+\theta(x_2-x_1)\) stays bounded, and Corollary 7 gives \[\label{Ceta95laplace} C_{\eta_\lambda}(t,x_1,x_2) =\frac{2\lambda\beta}{T}\operatorname{tr}(H_*^{-1})+\mathrm{O}(\lambda^2),\tag{24}\] uniformly in that set. Equivalently, since \(T-t\) is bounded away from zero on such compact subsets, \[\int_0^1\big(d-(T-t)\Delta\Phi_\lambda(t,x(\theta))\big)\,\mathrm{d}\theta = \frac{\lambda}{T(T-t)}\operatorname{tr}(H_*^{-1})+\mathrm{O}(\lambda^2).\] Thus the factorized trace OSL coefficient is of order \(\lambda\) on compact subsets away from the terminal time. Exponential estimates concern the tails away from neighborhoods of \(x^*\), whereas the full covariance is generically of order \(\lambda\) in the nondegenerate Laplace regime.

Gradient-free discretization The continuous finite-horizon theory produces a natural gradient-free drift formula ?? , which explains the underlying exploration-exploitation mechanism. The purpose of this section is not to develop a full complexity theory. It is simply to show how the exact barycentric representation ?? suggests a practical gradient-free discretization and to indicate how the parameters \(T\), \(\lambda\), \(N\), and \(h\) should be interpreted.

Suppose that at time \(t\) the current position is \(x\). To approximate the barycenter \(a_\lambda(t,x)\) from ?? , one may draw \(Y^1,\dots,Y^N\) independently from the Gaussian law \(\mathcal{N}(x,2\beta(T-t)I_d)\) and use the importance-sampling estimator \[\label{mc95barycenter} a_{\lambda,N}(t,x) = \frac{\sum_{j=1}^N Y^j \exp \big( - \alpha_\lambda f(Y^j) \big)}{\sum_{j=1}^N \exp \big( - \alpha_\lambda f(Y^j) \big)}.\tag{25}\] This yields the approximate drift \[\label{approx95drift} \mathbf{u}_{\lambda,N}(t,x) = - \frac{x - a_{\lambda,N}(t,x)}{T-t}.\tag{26}\] A simple Euler-Maruyama discretization (see [36]) on the grid \(t_k = kh\), stopped at \(t_K = T - \delta\) with \(\delta > 0\), reads \[\label{em95scheme} X_{k+1} = X_k + h\mathbf{u}_{\lambda,N}(t_k,X_k) + \sqrt{2\beta h}\xi_k, \qquad \xi_k \sim \mathcal{N}(0,I_d), \qquad k = 0,\dots,K-1.\tag{27}\] The parameters have the following interpretation:

  1. \(T\) fixes the horizon and therefore the search radius \(\sqrt{2\beta(T-t)}\) at time \(t\).

  2. \(\lambda\) fixes the Gibbs concentration through \(\alpha_\lambda = \frac{T}{2\lambda\beta}\): small \(\lambda\) means low temperature and stronger concentration near minimizers.

  3. \(N\) controls the Monte Carlo approximation of the barycenter \(a_\lambda(t,x)\) in 25 .

  4. \(h\) is the time discretization step.

A satisfactory non-asymptotic theorem for 27 should combine the time discretization error, the Monte Carlo variance, the near-terminal stiffness coming from the factor \((T-t)^{-1}\), the low-temperature covariance scaling in 24 and a fair comparison at equal computational budget with other global methods. A detailed description of the computational algorithm along these lines is left to a follow-up work.

Figure 3 visualizes that mechanism on the two-dimensional Ackley landscape. The colormap represents the conditional terminal density \(\eta_{\lambda,t,x}\), the cyan dot is the current position \(x\), the red star is the minimizer, and the white arrow is the drift \(\mathbf{u}_\lambda^*(t,x)\). As \(t\) increases, the conditional terminal law shrinks and the drift passes from a genuinely global barycentric direction to a local direction. We stress that this figure is only qualitative: the Ackley function is used as a familiar visual benchmark, but it does not exactly satisfy Assumption 1.

Figure 3: Visualization of the conditional terminal density \eta_{\lambda,t,x} and of the barycentric drift \mathbf{u}_\lambda^*(t,x) at several times. Early times correspond to broad terminal averaging, while late times recover a local direction.

Conclusion and further comments

This paper studies the finite-horizon criterion 1 as an optimization problem over drift fields. Starting from the deterministic warm-up, we solved the corresponding stochastic control problem on \(\mathrm{I\kern-0.21emR}^d\), identified the exact Markov feedback, and reorganized it around the conditional terminal law of the optimal process. This law is the Gibbs measure of the penalized energy ?? ; it yields the potential, averaged-gradient, and barycentric representations ?? , ?? and ?? .

We then analyzed two asymptotic regimes that are directly relevant for optimization. As \(t \uparrow T\), the drift reconnects with a scaled gradient descent. As \(\lambda \downarrow 0\), the conditional terminal law, the optimal process, and the drift select the unique global minimizer, even when \(f\) has nonglobal local minima; in the nondegenerate case we also obtained leading-order Laplace asymptotics for the drift, the value function, and the covariance of the conditional terminal law. The two regimes are generically non-commutative (Remark 17), which reflects the tension between global exploration and local exploitation.

Finally, we recorded a simple gradient-free discretization suggested by the exact barycentric formula.

The Hopf-Cole closed form and its relation with linearly solvable control, reciprocal diffusions, mean-field games, and recent Hamilton-Jacobi or Wasserstein-proximal viewpoints are classical. The present paper focuses on a narrower aspect which, we believe, enables a sharper message: it isolates the optimization content of that classical formula in a coherent framework on \(\mathrm{I\kern-0.21emR}^d\); it places the conditional terminal law at the center of the discussion; it makes the endogenous search radius \(\sqrt{2\beta(T-t)}\) and temperature \(\displaystyle \frac{2\lambda\beta}{T}\) completely explicit; and it states the low-temperature asymptotics in a form directly relevant for global optimization. Thus, the novelty is not a new Hopf-Cole transform, but the systematic extraction of the optimizer selected by this transform and the precise global-selection interpretation of its conditional terminal law. In particular, the paper highlights that the exact continuous-time criterion selects the global minimizer asymptotically under uniqueness of that minimizer, without any assumption excluding non-global local minima.

We conclude with a few further comments and perspectives.

0.6.0.2 Bounded domains and reflection.

Throughout the paper we have deliberately worked on the whole space \(\mathrm{I\kern-0.21emR}^d\). If one wishes instead to impose a hard state constraint on a smooth bounded domain \(\Omega \subset \mathrm{I\kern-0.21emR}^d\), the coherent way to do so is to replace the free diffusion on \(\mathrm{I\kern-0.21emR}^d\) by a reflected diffusion \[X_t = x_0 + \int_0^t u_s \,\mathrm{d}s + \sqrt{2\beta}B_t - \int_0^t n(X_s)\,\mathrm{d}K_s,\] where \(n\) is the outward unit normal and \(K_t\) is the boundary local time. At the PDE level, the corresponding Fokker-Planck equation carries a no-flux condition and the backward equation for \(h_\lambda\) becomes the backward heat equation with Neumann boundary condition. In that reflected setting, a counterpart of the present analysis should be possible with the Neumann heat semigroup. The local terminal cancellation discussed in Remark 10 should then be understood only away from boundary effects. Near the boundary, reflected heat-kernel contributions and the geometry of the boundary may alter the Riccati/caustic picture, and the covariance factorization used in the whole-space proof must be replaced by its reflected analogue. By contrast, declaring \(f = +\infty\) outside \(\Omega\) while still evolving a free diffusion on \(\mathrm{I\kern-0.21emR}^d\) mixes two different models and is not a satisfactory way to encode state constraints.

0.6.0.3 Several global minimizers.

If \(f\) has several global minimizers, the Gibbs measures \(\pi_\lambda(y)\,\mathrm{d}y\) still concentrate on the set \(\operatorname*{argmin} f\), but the limiting distribution need not be a single Dirac mass. Accordingly, the barycenter \(a_\lambda(t,x)\) may depend in a nontrivial way on the Gaussian factor \(G_{T-t}(x-\cdot)\) and on the geometry of the minimizer set. In that regime one should not expect the single affine limit \(\displaystyle - \frac{x - x^*}{T-t}\) without additional symmetry-breaking assumptions.

0.6.0.4 Unique global minimizer and several local minima.

This is the regime covered by Section [sec:asymp95sec]. The only structural assumption there is the uniqueness of the global minimizer; any number of nonglobal local minima or saddle points is allowed. The exponential concentration estimate of Proposition 14 shows that, at the level of the exact continuous-time criterion, those local minima do not affect the low-temperature selection mechanism: the terminal law still concentrates near \(x^*\). What the paper does not prove is a finite-\(\lambda\) or complexity guarantee for a practical algorithm. For fixed small \(\lambda\) and for the Monte Carlo discretization of Section [sec:numerics95sec], metastability, sampling error, and near-terminal stiffness may still matter. Thus the result is a global-selection theorem for the exact drift in the asymptotic regime \(\lambda \downarrow 0\), not yet a full computational-complexity statement.

0.6.0.5 Other running costs and control constraints.

The quadratic running cost is special because it leads to the Hopf-Cole linearization of the Hamilton-Jacobi equation 8 via the heat equation 10 , and therefore to the explicit formulas of Theorem 3, Proposition 7, and Proposition 8; see also [4][6]. Replacing it by an \(L^1\) running cost, or imposing a hard bound \(|u_t |\leqslant 1\), leads to a completely different optimal control problem. From the Hamiltonian point of view, this is where a Pontryagin maximum principle analysis should become highly relevant. In the \(L^1\) case the minimizing Hamiltonian is no longer quadratic and one expects nonsmooth Hamilton-Jacobi equations and possibly sparse drifts. Under the hard bound \(|u_t |\leqslant 1\), the minimizing controls should saturate the constraint on the region where the adjoint gradient is nonzero, leading to bang-bang type feedbacks. These variants are mathematically natural from the optimization viewpoint, but they fall outside the linearly solvable framework of the present paper.

References↩︎

[1]
W. H. Fleming and H. M. Soner, Controlled Markov Processes and Viscosity Solutions, second edition, Springer, 2006.
[2]
E. Trélat, Control in Finite and Infinite Dimension, SpringerBriefs in PDEs and Data Science, Springer, 2024.
[3]
J. Yong and X. Y. Zhou, Stochastic Controls: Hamiltonian Systems and HJB Equations, Springer, 1999.
[4]
H. J. Kappen, Path integrals and symmetry breaking for optimal control theory, J. Stat. Mech. Theory Exp. 2005, P11011.
[5]
E. Todorov, Linearly-solvable Markov decision problems, in Advances in Neural Information Processing Systems19, 2006, pp. 1369-1376.
[6]
E. Todorov, Efficient computation of optimal actions, Proc. Natl. Acad. Sci. USA106(2009), 11478-11483.
[7]
Y. Dai, Y. Jiao, L. Kang, X. Lu, and J. Z. Yang, Global optimization via Schrödinger-Föllmer diffusion, SIAM J. Control Optim.61(2023), 2953-2980.
[8]
P. Dai Pra, A stochastic control approach to reciprocal diffusion processes, Appl. Math. Optim.23(1991), 313-329.
[9]
H. Föllmer, An entropy approach to the time reversal of diffusion processes, in Stochastic Differential Systems, Filtering and Control, Lecture Notes in Control and Information Sciences 69, Springer, 1985, pp. 156-163.
[10]
J. Huang, Y. Jiao, L. Kang, X. Liao, J. Liu, and Y. Liu, Schrödinger-Föllmer sampler: sampling without ergodicity, arXiv:2106.10880, 2021.
[11]
B. Tzen and M. Raginsky, Theoretical guarantees for sampling and inference in generative models with latent diffusions, in Proceedings of Machine Learning Research99, 2019, pp. 3084-3114.
[12]
Q. Zhang and Y. Chen, Path integral sampler: a stochastic control approach for sampling, in International Conference on Learning Representations, 2022.
[13]
A. Bensoussan, J. Frehse, and P. Yam, Mean Field Games and Mean Field Type Control Theory, Springer, 2013.
[14]
J.-M. Lasry and P.-L. Lions, Mean field games, Japan. J. Math.2(2007), 229-260.
[15]
Y. Huang, D. Kalise, and H. Kouhkouh, Non-convex global optimization as an optimal stabilization problem: dynamical properties, arXiv preprint arXiv:2511.10815, 2025.
[16]
H. Heaton, S. Wu Fung, and S. Osher, Global solutions to nonconvex problems by evolution of Hamilton-Jacobi PDEs, Commun. Appl. Math. Comput.6(2024), 790-810.
[17]
W. Li, S. Liu, and S. Osher, A kernel formula for regularized Wasserstein proximal operators, Res. Math. Sci.10(2023), Article 43.
[18]
S. Osher, H. Heaton, and S. Wu Fung, A Hamilton-Jacobi-based proximal operator, Proc. Natl. Acad. Sci. USA120(2023), e2220469120.
[19]
C. Andrieu, N. Chopin, E. Fincato, and M. Gerber, Gradient-free optimization via integration, arXiv:2408.00888, 2024.
[20]
J. A. Carrillo, Y.-P. Choi, C. Totzeck, and O. Tse, An analytical framework for consensus-based global optimization method, Math. Models Methods Appl. Sci.28(2018), 1037-1066.
[21]
M. Fornasier, T. Klock, and K. Riedl, Consensus-based optimization methods converge globally, SIAM J. Optim.34(2024), 2973-3004.
[22]
H. Iwakiri, Y. Wang, S. Ito, and A. Takeda, Single loop Gaussian homotopy method for non-convex optimization, in Advances in Neural Information Processing Systems35, 2022, pp. 7065-7076.
[23]
G. A. Pavliotis, Stochastic Processes and Applications: Diffusion Processes, the Fokker-Planck and Langevin Equations, Springer, 2014.
[24]
H. Risken, The Fokker-Planck Equation: Methods of Solution and Applications, second edition, Springer, 1996.
[25]
A. Dembo and O. Zeitouni, Large Deviations Techniques and Applications, Springer, 2010.
[26]
L. S. Pontryagin, V. G. Boltyanskii, R. V. Gamkrelidze, and E. F. Mishchenko, The Mathematical Theory of Optimal Processes, Interscience, New York, 1962.
[27]
X. Li and J. Yong, Optimal Control Theory for Infinite Dimensional Systems, Birkhäuser, Boston, 1995.
[28]
M. Bardi and I. Capuzzo-Dolcetta, Optimal Control and Viscosity Solutions of Hamilton-Jacobi-Bellman Equations, Birkhäuser, Boston, 1997.
[29]
P. Cannarsa and C. Sinestrari, Semiconcave Functions, Hamilton-Jacobi Equations, and Optimal Control, Progress in Nonlinear Differential Equations and Their Applications 58, Birkhäuser, Boston, 2004.
[30]
P.-L. Lions, Generalized Solutions of Hamilton-Jacobi Equations, Research Notes in Mathematics 69, Pitman, 1982.
[31]
A. A. Agrachev and Y. L. Sachkov, Control Theory from the Geometric Viewpoint, Encyclopaedia of Mathematical Sciences 87, Springer, 2004.
[32]
B. Bonnard, J.-B. Caillau, and E. Trélat, Second order optimality conditions in the smooth case and applications in optimal control, ESAIM Control Optim. Calc. Var.13(2007), 207-236.
[33]
C.-R. Hwang, Laplace’s method revisited: weak convergence of probability measures, Ann. Probab.8(1980), 1177-1182.
[34]
N. G. de Bruijn, Asymptotic Methods in Analysis, third edition, Dover, 1981.
[35]
R. Wong, Asymptotic Approximations of Integrals, SIAM, 2001.
[36]
P. E. Kloeden and E. Platen, Numerical Solution of Stochastic Differential Equations, Springer, 1992.

  1. Here \(\| \cdot \|\) denotes the usual induced matrix norm.↩︎