Mirror descent for stochastic control problems with measure-valued controls


Abstract

This paper studies the convergence of the mirror descent algorithm for finite horizon stochastic control problems with measure-valued control processes. The control objective involves a convex regularisation function, denoted as \(h\), with regularisation strength determined by the weight \(\tau\ge 0\). The setting covers regularised relaxed control problems. Under suitable conditions, we establish the relative smoothness and convexity of the control objective with respect to the Bregman divergence of \(h\), and prove linear convergence of the algorithm for \(\tau=0\) and exponential convergence for \(\tau>0\). The results apply to common regularisers including relative entropy, \(\chi^2\)-divergence, and entropic Wasserstein costs. This validates recent reinforcement learning heuristics that adding regularisation accelerates the convergence of gradient methods. The proof exploits careful regularity estimates of backward stochastic differential equations in the bounded mean oscillation norm.

Mirror descent , stochastic control , convergence rate analysis, Bregman divergence , Pontryagin’s optimality principle

93E20,49M05 ,68Q25 ,60H30

1 Introduction↩︎

This paper designs a gradient descent algorithm for a finite horizon stochastic control problems where the controlled system is a \(\mathbb{R}^d\)-valued diffusion process, and both the drift and diffusion coefficients of the state process are controlled by measure-valued control processes. These measure-valued controls include as a special case relaxed controls, which are important for several reasons. Employing relaxed controls provides a convenient approach to guarantee the existence of a minimiser for the control problem, even when an optimal classical control may not exist [1]. The relaxed control perspective also offers a principled framework for designing efficient algorithms. It has been observed that the use of relaxed controls, along with additional regularization terms, improves algorithm stability and efficiency [2][5]. The regularised relaxed control formulation further ensures the optimal controls are stable with respect to model perturbations [6]. This stability of optimal controls proves to be essential for optimising the sample efficiency of associated learning algorithms [7][11].

In the sequel, we will give a heuristic overview of the problem, the algorithm and our results. All notation and assumptions will be stated precisely in Section 2.

Problem formulation↩︎

Let \(T\in (0,\infty)\) be a given terminal time, let \((\Omega, \mathcal{F}, \mathbb{F}, \mathbb{P})\) be a filtered probability space satisfying the usual conditions, on which a \(d'\)-dimensional standard Brownian motion \(W=(W_t)_{t\in [0,T]}\) is defined. Let \(\mathcal{P}(A)\) be the space of probability measures on a separable metric space \(A\), and let \(\mathfrak C\) be a convex subset of \(\mathcal{P}(A)\). We consider the admissible control space \(\mathcal{A_{\mathfrak C}}\) containing \(\mathbb{F}\)-progressively measurable processes \(\pi:\Omega\times [0,T]\to \mathfrak C\) satisfying suitable integrability conditions, which will be precisely formulated later in 10 . For each \(\pi\in \mathcal{A_{\mathfrak C}}\), consider the following controlled dynamics: \[\label{sde} dX_s(\pi)=b_s(X_s(\pi),\pi_s)\,ds+\sigma_s(X_s(\pi),\pi_s)\,dW_s\,,\,\, s\in [0,T]\,,\,\,\, X_0 (\pi) = x_0\,,\tag{1}\] where \(x_0\in \mathbb{R}^d\) is a given initial state, \(b: [0,T]\times \mathbb{R}^{d}\times \mathcal{P}(A)\to \mathbb{R}^{d}\) and \(\sigma:[0,T]\times\mathbb{R}^d\times \mathcal{P}(A) \rightarrow\mathbb{R}^{d\times d'}\) are sufficiently regular coefficients such that 1 admits a unique strong solution \(X(\pi)\). For a given regularisation parameter \(\tau\ge 0\), the agent’s objective is to minimise the following cost functional \[\label{problem} J^{\tau}(\pi)\mathrel{\vcenter{:}}= \mathbb{E}\left[\int_{0}^{T}\left( f_s(X_s(\pi),\pi_s)+\tau\,h(\pi_s)\right)ds+g(X_T(\pi))\right]\tag{2}\] over all admissible controls \(\pi\in \mathcal{A_{\mathfrak C}}\), where \(f:[0,T]\times \mathbb{R}^{d}\times \mathcal{P}(A)\to \mathbb{R}\) and \(g: \mathbb{R}^{d} \to \mathbb{R}\) are sufficiently regular cost functions, and \(h:\mathfrak C\to \mathbb{R}\) is a convex regularisation function. Common choices of \(h\) include the relative entropy used in reinforcement learning [3], [5], [12][14], the \(f\)-divergence used in information theory and statistical learning [15], and the (regularised) Wasserstein cost used in optimal transport [16], [17]. Choquet regularization is proposed in [18] for the special case when \(A=\mathbb{R}\). Typically, the regulariser \(h\) does not admit higher-order dertivatives over its domain and in this sense it is less regular than the cost functions \(f\) and \(g\).

Difficulties in gradient-based algorithms.↩︎

This work aims to design a convergent gradient-based algorithm for a (nearly) optimal control of 2 . To this end, for each \(\tau\ge 0\), we equivalently write 2 as \[\label{eq:F95H} J^\tau(\pi)=J^0(\pi) + \tau \mathcal{H}(\pi), \quad \textrm{with\mathcal{H}(\pi)\mathrel{\vcenter{:}}= \mathbb{E}\left[\int_{0}^{T} h(\pi_s) ds \right]}\tag{3}\] and the objective is to minimise this over \(\pi\in \mathcal{A}_{\mathfrak C}\). Such an infinite dimensional optimisation perspective facilitates extending gradient-based algorithms designed for static minimisation problems to the current dynamic setting. The decomposition of \(J^\tau\) into \(J^0\) and \(\mathcal{H}\) further enables treating the non-smooth regulariser \(h\) separately within the algorithms.

A well-established gradient-based algorithm for minimising a functional \(f+ g\) over a Hilbert space \((\mathcal{X},\|\cdot\|_{\mathcal{X}})\) is the proximal gradient algorithm [19]. It is shown in [19] that, if \(f\) is Fréchet differentiable, strongly convex and smooth (i.e., the derivative of \(f\) is Lipschitz continuous) with respect to the norm \(\|\cdot\|_{\mathcal{X}}\), then the non-smooth term \(g\) can be handled by the proximal operator, and a proximal gradient method converges exponentially to the minimiser.

However, incorporating measure-valued controls in 2 poses significant challenges when adapting the proximal gradient algorithm to the present setting. Firstly, the space \(\mathcal{A}_{\mathfrak C}\) of admissible controls is not a Hilbert space, rendering the convergence results in [19] inapplicable. Secondly, computing the gradient of \(J^0\) necessitates introducing a suitable notation of derivatives over \(\mathcal{P}(A)\). Moreover, to establish strong convexity and smoothness of \(J^0\), it is essential to assess its regularity with respect to distinct controls \(\pi\) and \(\pi'\). Since \(\mathcal{P}(A)\) can be equipped with multiple metrics (such as the total variation metric or Wasserstein metrics), it remains unclear which metric is the most natural choice to quantify the regularity of \(J^0\). Lastly, the convergence analysis in [19] overlooks the potential regularisation effect of \(h\). While [19] requires \(J^0:\mathcal{A}_{\mathfrak C}\to \mathbb{R}\) to be strongly convex for exponential convergence despite the existence of the regulariser \(h\), it has been observed in [14], [20], [21] that selecting \(h\) as the relative entropy accelerates the convergence of gradient descent algorithms.

Mirror descent algorithm↩︎

In this work, we adopt an alternative approach and introduce a mirror descent algorithm for 2 . Our algorithm can be viewed as a dynamic extension of [22], which builds on optimisation results in \(\mathbb{R}^d\) [23], [24] and studies mirror descent algorithms for minimising a cost \(f+g\) over spaces of measures. The key observations therein are (i) the Bregman divergence associated with \(g\) provides a natural notion to express regularity of involved functions; and (ii) the relative smoothness and convexity of \(f\) with respect to the Bregman divergence of \(g\) are sufficient to ensure the convergence of mirror descent. We extend these insights to the current dynamic setting.

To derive the mirror descent algorithm for 2 , define the Hamiltonian \(H^{\tau}:[0,T]\times\mathbb{R}^d\times\mathbb{R}^d\times\mathbb{R}^{d\times d'}\times \mathfrak C\rightarrow \mathbb{R}\) by \[\label{eq:32hamiltonian} \begin{align} H_t^{\tau}(x,y,z,m)\mathrel{\vcenter{:}}= b_t(x,m)^\top y +\text{tr}(\sigma_t^\top (x,m)z)+f_t(x,m)+\tau\,h(m)\,. \end{align}\tag{4}\] Under appropriate conditions on the coefficients, we prove in Lemma 6 that for all \(\pi,\pi'\in \mathcal{A}_{\mathfrak C}\), \[\begin{align} \label{eq:first95variation95F} \begin{aligned} d^+ J^0 (\pi)(\pi'-\pi) &\mathrel{\vcenter{:}}= \lim_{\varepsilon\searrow 0}\frac{ J^0(\pi+\varepsilon(\pi'-\pi))-J^0(\pi)}{\varepsilon} =\mathbb{E}\int_0^T\int\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi),\pi_s,a)(\pi'_s-\pi_s)(da)\,ds\,, \end{aligned} \end{align}\tag{5}\] where \(\Theta(\pi)=(X(\pi), Y(\pi), Z(\pi))\) with \(X(\pi)\) being the state process satisfying 1 and \(( Y(\pi), Z(\pi))\) being the adjoint processes satisfying the following backward stochastic differential equation (BSDE): \[\label{eq:32adjoint32equation} \left\{ \begin{align} dY_s(\pi)&=-(D_x H^0_s)(X_s(\pi),Y_s(\pi),Z_s(\pi),\pi_s)\,ds+Z_s(\pi)\,dW_s\,,\,\,s\in[0,T]\,,\\ Y_T(\pi)&=(D_xg)(X_T(\pi))\,, \end{align} \right.\tag{6}\] and for all \(\phi:\mathfrak C\to \mathbb{R}\), we denote by \(\frac{\delta \phi}{\delta m}:\mathfrak C\times A\to \mathbb{R}\) the flat derivative of \(\phi\) (see Definition 1). Equation 5 indicates that \(\frac{\delta H^{0}_\cdot}{\delta m}(\Theta_\cdot(\pi),\pi_\cdot,\cdot)\) characterises the first variation of \(J^0\) at \(\pi\).

Now assume that \(h\) admits a flat derivative \(\frac{\delta h}{\delta m}\) on \(\mathfrak C\), and consider the Bregman divergence associated with \(h\) such that for all \(m, m' \in \mathfrak C\), \[\label{eq:D95h95introd} D_h(m'|m) = h(m')-h(m) - \int \frac{\delta h}{\delta m }(m,a) (m'-m)(da) \,.\tag{7}\] The mirror descent algorithm for 2 is given as follows.

Figure 1: h-Bregman Mirror Descent Algorithm

By 3 and 5 , the minimisation step [eq32update32the32control32bregman32msa] can be formally interpreted is a pointwise representation of the mirror descent update \[\pi^{n+1} \in \mathop{\mathrm{arg\,min}}_{\pi \in \mathcal{A}_{\mathfrak C}} \left( d^+ J^\tau (\pi^n)(\pi-\pi^n) +\lambda D_{\mathcal{H}}(\pi|\pi^n)\right)\,,\] where \(d^+ J^\tau (\pi^n)\) is the first variation of \(J^\tau\) at \(\pi^n\), and \(D_{\mathcal{H}}(\pi|\pi^n) =\mathbb{E}\int_0^T D_h(\pi_s|\pi^n_s) \, ds\) is the Bregman divergence of \(\mathcal{H}\). As a result, [eq32update32the32control32bregman32msa] optimises the first-order approximation of \(J^\tau\) around \(\pi^n\), and the Bregman divergence \(\pi\mapsto \lambda D_{\mathcal{H}}(\pi|\pi^n)\) ensures that the optimisation is over the part of the domain where such an approximation is sufficiently accurate. Similar ideas have been employed in the design of gradient-based algorithms in [25], [26] for problems with finite-dimensional action spaces, and in [9] for linear-quadratic problems with measure-valued controls.

Our contributions↩︎

This paper identifies conditions under which Algorithm 1 converges to an optimal control of 2 and characterises its convergence rate.

  • For a general regularisation function \(h\), we identify regularity conditions on the coefficients such that Algorithm 1 is well-defined and the loss \(J^\tau: \mathcal{A}_{\mathfrak C}\to \mathbb{R}\) is relative smooth with respect to the divergence \((\pi',\pi)\mapsto \mathbb{E}\int_0^T D_h(\pi'_s|\pi_s) \, d s\) (Theorem 4). The coefficient regularity in the measure component is characterised using the Bregman divergence \(D_h\) in 7 , and the diffusion coefficient \(\sigma\) is allowed to be degenerate, state-dependent and controlled.

  • Leveraging the relative smoothness, we prove that \(\{J^\tau(\pi^{n})\}_{n\in \mathbb{N}}\) is decreasing (Theorem 5). If we further assume that the unregularised Hamiltonian \(H^0\) in 4 and the terminal cost \(g\) are convex in \(x\) and \(m\), we prove that Algorithm 1 converges linearly when \(\tau=0\) and exponentially when \(\tau>0\) (Theorem 8). Note that in comparison with [26], the exponential convergence is achieved by leveraging the regularisation effect of \(h\), without imposing strong convexity on the unregularised cost \(J^0\). To the best of our knowledge, this is the first work on the convergence of mirror descent algorithms for continuous-time control problems with measure-valued control processes and general regularisation functions.

  • We demonstrate the applicability of the convergence results with concrete examples of \(h\), including the relative entropy, the \(\chi^2\)-divergence, and the entropic Wasserstein cost.

Here we highlight that analysing mirror descent (i.e., Algorithm 1) in the present dynamic setting introduces additional challenges beyond those encountered in static optimisation problems [22]. Specifically, as the coefficient regularity is measured using the Bregman divergence \(D_h\), which is neither a metric on \(\mathcal{P}(A)\) nor equivalent to commonly used metrics, special care is required for the well-posedness of Algorithm 1 and the differentiability of \(J^0\) (see the discussion above Lemma 7). Moreover, establishing the relative smoothness of \(J^0\) requires fine estimates on the solutions of 6 , which employ the theory of bounded mean oscillation (BMO) martingales (see Remark 18 for details).

We also emphasize that this work analyzes the convergence rate of the discrete-time mirror descent update [eq32update32the32control32bregman32msa], unlike [21], which proves the exponential convergence of a continuous-time Wasserstein gradient flow for the control problem 1 2 with a relative entropy regularizer. Analyzing the discrete-time gradient updates necessitates novel techniques to address time discretization errors, which are not covered in the continuous-time gradient flow analysis in [21]. Moreover, our convergence analysis extends the Pontryagin optimality principle and BSDE analysis from [21] by measuring coefficient regularity with a general Bregman divergence, rather than the Wasserstein metric used in [21]. This approach offers greater flexibility in accommodating general model coefficients and general regularizers for the control problem (see Section 2.4 for more details).

Literature review↩︎

Two primary approaches to solve a (stochastic) control problem are the dynamic programming principle, leading to a nonlinear Hamilton-Jacobi-Bellman (HJB) partial differential equation (PDE) [27], and Pontryagin’s optimality principle [28], [29], leading to a coupled forward-backward stochastic differential equation (FBSDE). Classical numerical methods involve first discretising these nonlinear equations and then solving the resulting discretised equations using iterative methods; see e.g., [30], [31] for the PDE approach and [32][38] for the FBSDE approach.

Recently, there has been a growing interest in designing iterative algorithms to linearise HJB PDEs or to decouple FBSDEs before discretisation. These algorithms have the advantage of solving linear equations at each step, enabling the use of efficient mesh-free algorithms, particularly in high dimensional settings [39][41]. For instance, policy iteration has been utilised to linearise HJB PDEs [42][45], which converges exponentially for drift-controlled problems [42] and super-exponentially for control problems in bounded domains [43]. Additionally, the method of successive approximation and its variants have been proposed to decouple FBSDEs [26], [46][48], which converge linearly for convex losses [47] and exponentially for strongly convex losses [26]. Note that both policy iteration and the method of successive approximation require the exact minimisation of the Hamiltonian over the entire action space at each iteration, which can be computationally expensive for high dimensional action spaces.

Gradient-based methods enhance algorithm efficiency, especially for high dimensional action spaces, by updating controls using the gradient of the Hamiltonian [49], [50]. Most existing works on the convergence of gradient-based algorithms focus on discrete-time control problems (see [14], [20], [51], [52] and references therein). For continuous-time control problems with continuous state and action spaces, theoretical studies on the convergence of gradient-based algorithms are limited. For continuous-time gradient flows, the exponential convergence is proved in [21] for open-loop control problems with sufficiently convex losses, and in [53] for drift-control problems with stochastic policies. For discrete update, the exponential convergence of policy gradient methods is proved in [12] for linear-quadratic problems with possibly nonconvex costs, and in [54] for drift-control problems with finite dimensional action spaces.

Organisation of the paper↩︎

We now conclude Section 1 which forms the introduction. In Section 2, we present the assumptions, main results, and examples of regularisers along with associated Bregman divergences applicable to the paper’s setting. Section 3 collects the proofs of main results. Finally, 4 recalls some known properties of Bregman divergence and 5 proves Lemma 6, the characterisation of the first variation of \(J^0\).

2 Main results↩︎

This section summarises the model assumptions and presents the main results. Throughout this paper, let \(T\in (0,\infty)\), let \((\Omega, \mathcal{F}, \mathbb{F}, \mathbb{P})\) be a filtered probability space satisfying the usual conditions, on which a \(d'\)-dimensional standard Brownian motion \(W=(W_t)_{t\in [0,T]}\) is defined. Let \((A,\rho_A)\) be a separable metric space, and let \(\mathcal{P}(A)\) be the space of probability measures on \(A\) equipped with the topology of the weak convergence of measures and the associated Borel \(\sigma\)-algebra.

2.1 Regulariser \(h\) and the associated Bregman divergence↩︎

Let \(\mathfrak C\) be a convex and measurable subset of \(\mathcal{P}(A)\), and \(h: \mathfrak C\to \mathbb{R}\) be a convex function, i.e., \(h(\varepsilon m + (1-\varepsilon)m')\le \varepsilon h(m)+(1-\varepsilon)h(m')\) for all \(\varepsilon \in [0,1]\) and \(m,m'\in \mathfrak C\). We require the function \(h\) to have a flat derivative \(\frac{\delta h}{\delta m}: \mathfrak C\times A\to \mathbb{R}\), whose precise definition is given as follows.

Definition 1 (Flat derivative on \(\mathfrak C \subseteq \mathcal{P}(A)\)). We say a function \(f:\mathfrak C\to \mathbb{R}^d\) has a flat derivative, if there exists a measurable function \(\frac{\delta f }{\delta m} :\mathfrak C\times A \to \mathbb{R}^d\), called the flat derivative of \(f\), such that for all \(m,m'\in\mathfrak C\), \(\int \big| \frac{\delta f }{\delta m} (m, a) \big| m'(da) <\infty\), \[\label{eq32in32def32of32flat32der} \lim_{\varepsilon \searrow 0 }\frac{f(m^\varepsilon)- f(m) }{\varepsilon} = \int \frac{\delta f }{\delta m} (m, a) ( m'-m) (da) \quad \textrm{with m^\varepsilon =m + \varepsilon (m' - m)\,,}\tag{8}\] and \(\int \frac{\delta f }{\delta m} (m, a) m(da) =0\).

Given a measurable space \(\mathcal{Y}\), we say a function \(f:\mathcal{Y} \times \mathfrak C\to \mathbb{R}^d\) has a flat derivative with respect to \(m\) on \(\mathfrak C\), if there exists a measurable function \(\frac{\delta f }{\delta m}: \mathcal{Y} \times \mathfrak C\times A \to \mathbb{R}^d\) such that \(\frac{\delta f}{\delta m} (y,\cdot)\) is the flat derivative of \(f(y,\cdot)\) for all \(y\in \mathcal{Y}\).

Remark 1. One can show that if \(f:\mathfrak C\to \mathbb{R}^d\) admits a flat derivative \(\frac{\delta f }{\delta m}\), then for all \(m,m'\in \mathfrak C\), the function \([0,1]\ni \varepsilon \mapsto f(m^\varepsilon)\) is continuous on \([0,1]\) and differentiable on \((0,1)\) with derivative \(\frac{d}{d \varepsilon}f(m^\varepsilon) = \int \frac{\delta f }{\delta m} (m^\varepsilon , a) ( m'-m) (da)\) (see [55] and [56]). Hence by the fundamental theorem of calculus, \(f(m')-f(m)=\int_0^1 \int \frac{\delta f }{\delta m} (m^\varepsilon , a) ( m'-m) (da)d \varepsilon\) provided that \(\varepsilon \mapsto \int \frac{\delta f }{\delta m} (m^\varepsilon , a) ( m'-m) (da)\) is integrable.

Given the regulariser \(h\) and its flat derivative \(\frac{\delta h}{\delta m}\), define the Bregman divergence \(D_h(\cdot|\cdot): \mathfrak C\times \mathfrak C\to [0,\infty)\) by \[\label{eq:bregman95def} D_h(m'|m) = h(m')-h(m) - \int \frac{\delta h}{\delta m }(m,a) (m'-m)(da) \,,\tag{9}\] and for a fixed \(\nu\in \mathfrak C\), define the set \(\mathcal{A}_{\mathfrak C}\) of the admissible controls by \[\label{eq:admissible95AC} \begin{align} \mathcal{A_{\mathfrak C}}&\mathrel{\vcenter{:}}= \left\{\pi:\Omega \times [0,T] \rightarrow \mathfrak C \,\middle\vert\, \begin{aligned} &\textrm{\pi is progressively measurable,} \\ &\textrm{\mathbb{E}\int_0^T | h(\pi_t ) |\, d t <\infty and \mathbb{E}\int_0^T D_h(\pi_t|\nu) \, d t <\infty . } \end{aligned} \right\}\,. \end{align}\tag{10}\] Here the condition \(\mathbb{E}\int_0^T D_h(\pi_t|\nu) \, d t <\infty\) ensures for each control \(\pi\in \mathcal{A}_{\mathfrak C}\), the state process is square integrable (see Proposition 2), while the condition \(\mathbb{E}\int_0^T | h(\pi_t ) |\, d t <\infty\) ensures that the regularised cost \(J^\tau(\pi)\) is finite for all \(\tau>0\). Note that \(\mathcal{A_{\mathfrak C}}\) is convex due to the convexity of \(h\) and \(\pi\mapsto D_h(\pi|\nu)\) (see Lemma 9 Item 47 ).

In the sequel, we work with a general regularisation function \(h\), quantify the coefficient regularity using the Bregman divergence \(D_h\), and analyse the convergence of Algorithm 1. Concrete examples of \(h\), including relative entropy, \(\chi^2\)-divergence and entropic optimal transport, and the corresponding Bregman divergence \(D_h\) will be provided in Section 2.4.

2.2 Well-posedness of the iterates↩︎

We start by imposing suitable regularity conditions on the coefficients for the well-posedness of the state dynamics 1 and adjoint dynamics 6 .

Assumption 1 (Regularity of \(b\) and \(\sigma\)). The measurable functions \(b:[0,T]\times\mathbb{R}^d\times \mathcal{P}(A)\to \mathbb{R}^d\) and \(\sigma:[0,T]\times\mathbb{R}^d\times \mathcal{P}(A) \rightarrow\mathbb{R}^{d\times d'}\) are differentiable in \(x\) and have flat derivatives \(\frac{\delta b}{\delta m}\) and \(\frac{\delta \sigma}{\delta m}\), respectively, with respect to \(m\) on \(\mathfrak C\). There exists \(\nu_0\in \mathfrak C\) such that \(\int_0^T |b_t(0,\nu_0)|\, dt<\infty\) and \(\int_0^T |\sigma_t(0,\nu_0)|^2\, dt<\infty\), and there exists \(K\ge 0\) such that for all \(t\in[0,T]\), \(x, x'\in\mathbb{R}^d\) and \(m, m'\in \mathfrak C\), \[\begin{align} |b_t(x,m)-b_t(x',m')|^2 +|\sigma_t(x,m)-\sigma_t(x',m')|^2 & \le K\left( |x-x'|^2+D_h(m|m')\right)\,. \end{align}\]

Assumption 2 (Regularity of \(f\) and \(g\)). The measurable function \(f: [0,T]\times \mathbb{R}^d \times \mathcal{P}(A) \to \mathbb{R}\) is differentiable in \(x\) and has a flat derivative \(\frac{\delta f}{\delta m}\) with respect to \(m\) on \(\mathfrak C\). The function \(g: \mathbb{R}^d\to \mathbb{R}\) is differentiable. There exists \(\nu_0\in \mathfrak C\) such that \(\int_0^T |f_t(0,\nu_0)|\, dt<\infty\), and there exists \(K\ge0\) such that for all \(t\in [0,T]\), \(x\in\mathbb{R}^d\) and \(m\in \mathfrak C\), \(|(D_x g)(x)|+|(D_x f_t)(x,m)|\le K\).

Under Assumptions 1 and 2, for each \(\pi\in \mathcal{A_{\mathfrak C}}\), the equations 1 and 6 admit unique solutions, as shown in the following proposition.

Proposition 2. Suppose Assumptions 1 and 2 hold. For each \(\pi\in\mathcal{A}_{\mathfrak C}\), 1 has a unique strong solution \(X(\pi)\in L^2(\Omega ; C([0,T];\mathbb{R}^d))\) and 6 has a unique solution \((Y(\pi),Z(\pi)) \in L^2(\Omega; C([0,T];\mathbb{R}^d)) \times L^2(\Omega \times [0,T ]; \mathbb{R}^{d\times d'})\).

The proof of Proposition 2 follows directly from standard well-posedness results of SDEs and BSDEs (see e.g., [57]) and hence is omitted. In particular, the square integrability of \(X(\pi)\) follows from the condition \(\mathbb{E}\int_0^T D_h(\pi_t|\nu) \, d t < \infty\) of \(\pi\in\mathcal{A}_{\mathfrak C}\).

We further assume that the mirror descent iterates \(\{\pi^n\}_{n\in \mathbb{N}\cup\{0\}}\) in Algorithm 1 are well-defined.

Assumption 3. Let \(\pi^0\in \mathcal{A}_{\mathfrak C}\) such that for all \(\lambda>0\), the iterates \(\{\pi^n\}_{n\in \mathbb{N}}\) in Algorithm 1 exist, belong to \(\mathcal{A}_{\mathfrak C}\) and satisfy \(\mathbb{E}\int_0^T D_h(\pi^{n+1}_t|\pi^n_t) \, d t < \infty\) for all \(n\in \mathbb{N}\cup\{0\}\).

Assumption 3 assumes that the pointwise update [eq32update32the32control32bregman32msa] maintains the admissible control set \(\mathcal{A}_{\mathfrak C}\). The condition that \(\mathbb{E}\int_0^T D_h(\pi^{n+1}_t|\pi^n_t) d t<\infty\) allows for using the integrated Bregman divergence as a Lyapunov function for the convergence analysis. In general, Assumption 3 needs to be verified depending on the specific choices of \(h\). We refer the reader to Section 2.4 for more details.

2.3 Convergence of the iterates↩︎

We introduce additional regularity assumptions for the coefficients, which will be used to establish the convergence of Algorithm 1.

Assumption 4 (Regularity of spatial derivatives).

There exists \(K\ge 0\) such that for all \(\phi \in \{b,f, \sigma \}\) and for all \(t\in[0,T]\), \(x, x'\in\mathbb{R}^d\) and \(m, m'\in \mathfrak C\), \[\begin{align} |(D_x \phi_t)(x,m)-(D_x \phi_t)(x',m')|^2 & \le K\left( |x-x'|^2+D_h(m|m')\right)\,,\,\,\quad |(D_xg)(x)-(D_xg)(x')| \le K |x - x'|\,. \end{align}\]

Assumption 5 (Regularity of measure derivatives). For each \(\phi \in \{b,\sigma, f \}\), the function \(\frac{\delta \phi}{\delta m}\) is differentiable in \(x\), has a flat derivative \(\frac{\delta^2 \phi}{\delta m^2}\) with respect to \(m\) on \(\mathfrak C\), and satisfies for some \(K\ge 0\) that for all \(t\in[0,T]\), \(x\in\mathbb{R}^d\) and \(m,m',m''\in\mathfrak C\), \[\begin{align} \left|\int \frac{\delta \phi_t}{\delta m}(x,m'',a)(m-m')(da)\right|^2 + \left|\int \left(D_x\frac{\delta \phi_t}{\delta m}\right)(x,m'',a)(m-m')(da)\right|^2 & \le K D_h(m|m')\,. \end{align}\] Moreover, for each \(\phi\in \{b,f\}\), there exists \(K\ge 0\) such that for all \(t\in[0,T]\), \(x\in\mathbb{R}^d\) and \(m,m', m''\in\mathfrak C\), \[\left|\int\int\frac{\delta^2 \phi_t}{\delta m^2}(x,m'',a,a')(m-m')(da')(m-m')(da)\right|\le KD_h(m|m')\,.\] For all \(t\in[0,T]\), \(x\in\mathbb{R}^d\), \(m\in\mathfrak C\) and \(a, a' \in A\), \(\frac{\delta^2 \sigma_t}{\delta m^2}(x,m,a,a') = 0\).

Remark 3. Assumption 5 requires that the derivative \(\frac{\delta \sigma}{\delta m}\) is independent of \(m\). This condition is satisfied when the diffusion coefficient is uncontrolled, or when it depends linearly on the measure-valued control. The latter scenario commonly arises in optimal resource allocation problems, where the control process represents the proportion of resources allocated to various projects, and the state process models the total amount of resources, which evolves according to the returns generated by the invested projects, depending linearly on the allocation strategy.

The condition on \(\frac{\delta \sigma}{\delta m}\) is used to prove Lemma 7, which is crucial for establishing the relative smoothness result in Theorem 4. Specifically, in 30 , we aim to show (assuming for simplicity that \(d=d'=1\)) that for all \(\pi,\pi'\in \mathcal{A}_\mathfrak{C}\), \[\begin{align} \label{eq:Z95estimate95divergence} \int_0^T\mathbb{E}\left[ \int\int \frac{\delta^2 \sigma_s}{\delta m^2}(X_s(\pi),\pi^{\varepsilon',\varepsilon}_s,a,a')Z_s(\pi)(\pi'_s-\pi_s)(da')(\pi'_s-\pi_s)(da) \right] d s\le C \mathbb{E}\int_0^T D_h(\pi'_s|\pi_s)d s\,. \end{align}\tag{11}\] The BMO estimate of \(Z\) in Lemma 1 and the upper bound in Lemma 4 suggest showing \[\begin{align} \mathbb{E}\left[ \left(\int_0^T \left|\int\int \frac{\delta^2 \sigma_s}{\delta m^2}(X_s(\pi),\pi^{\varepsilon',\varepsilon}_s,a,a') (\pi'_s-\pi_s)(da')(\pi'_s-\pi_s)(da) \right|^2 d s\right)^{1/2}\right]\le C \mathbb{E}\int_0^T D_h(\pi'_s|\pi_s)d s\,. \end{align}\] However, this would require bounding the \(L^2\)-norm in time by an \(L^1\)-norm, which is generally infeasible unless the integrand is zero.

The above discussion reflects a fundamental difficulty in analyzing gradient-based algorithms for continuous-time control problems with nonlinearly controlled diffusion coefficients. Existing works focus on cases where the diffusion coefficient is either uncontrolled (see, e.g., [53], [54]) or linearly controlled (see [12], [21], [58], [59]). It remains an interesting open problem to identify sufficient conditions for designing convergent gradient-descent algorithms when diffusion coefficients are nonlinearly controlled.

Using Assumptions 1, 2, 4 and 5, we prove that the performance difference of any two admissible controls can be upper bounded by a first order term and their Bregman divergence. By interpreting \(\frac{\delta H^{\tau}_\cdot}{\delta m}(\Theta_\cdot(\pi),\pi_\cdot,\cdot)\) as the derivative of \(\pi\to J^\tau (\pi)\), the theorem shows that the cost functional is relatively smooth with respect to the Bregman divergence (see e.g., [22]).

Theorem 4 (Relative smoothness). Suppose Assumptions 1, 2, 4 and 5 hold and \(\tau\ge 0\). Then there exists \(L\ge \tau\) such that for any \(\pi,\pi' \in \mathcal{A}_\mathfrak{C}\), \[\label{eq32estimate32for32diff32J32with32explicit32update} \begin{align} J^{\tau}(\pi')-J^{\tau}(\pi) & \le \mathbb{E}\int_0^T \int\frac{\delta H^{\tau}_s}{\delta m}(\Theta_s(\pi),\pi_s,a)(\pi'_s-\pi_s)(d a) ds + L\mathbb{E}\int_0^T\,D_h(\pi_s'|\pi_s)\,ds\,. \end{align}\tag{12}\]

The proof of Theorem 4 is given in Section 3.2. The constant \(L\) depends on the regularisation parameter \(\tau\), the time horizon \(T\), the initial state \(x_0\), and the dimension \(d\) of the state process, the dimension \(d'\) of the Brownian motion and the constants in Assumption 1, 2, 4 and 5.

Based on Theorem 4, we prove that the cost functional decreases along the iterates \(\{\pi^n\}_{n\in \mathbb{N}}\) from Algorithm 1.

Theorem 5 (Energy dissipation). Suppose Assumptions 1, 2, 3, 4 and 5 and \(\tau\ge 0\). Let \(\lambda \ge L\) with \(L\) from Theorem 4, and let \(\{\pi^n\}_{n\in \mathbb{N}}\) be the iterates from Algorithm 1. Then \(J^{\tau}(\pi^{n+1}) \leq J^{\tau}(\pi^{n})\) for all \(n\in \mathbb{N}\cup\{0\}\). Assume further that \(\inf_{\pi \in\mathcal{A}_{\mathfrak C}} J^\tau(\pi)>-\infty\), then \(\lim_{n\rightarrow \infty}\mathbb{E}\int_0^T D_h(\pi^{n+1}_t| \pi^n_t)\,dt = 0\).

The proof of Theorem 5 is given in Section 3.3. Note that Theorem 5 allows the cost functional \(\pi\mapsto J^{\tau}(\pi)\) to be nonconvex and the regularisation parameter \(\tau\) to be zero. However, it does not imply the cost function \(\{J^{\tau} (\pi^{n} )\}_{n\in \mathbb{N}}\) converges to the optimal cost.

To ensure the convergence of Algorithm 1 and determine its convergence rate, we impose additional convexity assumptions on the coefficients under which the Pontryagin optimality principle serves as a sufficient criterion for optimality.

Assumption 6 (Convexity). The function \(g: \mathbb{R}^d\to \mathbb{R}\) is convex, and for all \((t,y,z)\in [0,T]\times \mathbb{R}^d\times \mathbb{R}^{d\times d'}\), the function \((x,m)\mapsto H^0_t(x,y,z,m)\) is convex, i.e., for all \(x,x'\in \mathbb{R}^d\) and \(m,m'\in \mathfrak C\), \[\begin{align} \begin{aligned} &H^0_t(x',y,z,m') - H^0_t(x,y,z,m) \ge D_x H^0_t(x,y,z,m)^\top (x'-x) + \int\frac{\delta H^0_t}{\delta m}(x,y,z,m,a)(m'-m)(da)\,. \end{aligned} \end{align}\]

Using Assumption 6, we prove that the cost functional \(\pi\mapsto J^\tau (\pi)\) is strongly convex with respect to the Bregman divergence \(D_h(\cdot|\cdot)\), and the modulus of convexity is the regularisation parameter \(\tau\ge 0\).

Theorem 6 (Relative convexity). Suppose Assumptions 1, 2 and 6 hold and \(\tau\ge 0\). Then for all \(\pi,\pi'\in \mathcal{A}_{\mathfrak C}\) with \(\mathbb{E}\int_{0}^{T} D_h(\pi'_s|\pi_s)\, ds <\infty\), \[\label{eq:relative-convexity-J} \begin{align} J^{\tau}(\pi')-J^{\tau}(\pi) \ge \mathbb{E}\int_0^T\int\frac{\delta H^{\tau}_s}{\delta m}(\Theta_s(\pi), \pi_s, a)(\pi'_s-\pi_s)(da)\,ds + \tau\mathbb{E}\int_{0}^{T} D_h(\pi'_s|\pi_s) \,ds\,. \end{align}\tag{13}\]

The proof of Theorem 6 is given in Section 3.4.

Remark 7. We contrast Theorems 4 and 6 with the performance difference lemma, a result often used to analyze gradient-based algorithms for Markov feedback controls.

It is known that the control objective is typically non-convex with respect to Markov controls (see [12]). To develop and analyze gradient-based algorithms for Markov controls, the performance difference lemma is a key tool. This lemma provides an exact expression for the difference between the value functions associated with any two Markov controls, in contrast to the approximate bounds presented in Theorems 4 and 6. Such a lemma was first introduced in [60] in the context of controlled Markov chains, and has since become a foundational tool in the analysis of gradient-based methods for Markov policies; see [14], [61] for controlled Markov chains, and [12], [13], [53] for controlled diffusion processes. In particular, the performance difference lemma enables a precise characterization of the cost landscape, which in turn makes it possible to recover optimal convergence rates for gradient flows–as if the objective were (strongly) convex [14], [53].

In contrast, this paper focuses on optimization over open-loop control processes. We show in Theorem 6 that the control objective is convex with respect to the control processes. By leveraging this improved landscape property, the analysis of gradient-based algorithms becomes more direct, and the performance difference lemma is not required.

Moreover, it is unclear whether an analogous performance difference lemma holds in the open-loop setting. The standard proof of the lemma relies on the Feynman–Kac formula, which is applicable under Markovian assumptions. In our setting, however, controls are general adapted processes, potentially non-Markovian, for which such a representation may not be available.

Combining Theorems 4 and 6, we characterise the convergence rate of Algorithm 1 depending on the regularisation parameter \(\tau\). Recall that for all \(\tau \ge 0\), \(\{J^\tau(\pi^{n})\}_{n\in \mathbb{N}}\) is decreasing, as shown in Theorem 5.

Theorem 8 (Linear/Exponential convergence of Algorithm 1). Suppose Assumptions 1, 2, 3, 4, 5 and 6 hold and \(\tau\ge 0\). Let \(\lambda \ge L\) with \(L\) from Theorem 4, and let \(\{\pi^n\}_{n\in \mathbb{N}}\) be the iterates from Algorithm 1.

  • If \(\tau=0\), then for all \(\pi\in \mathcal{A}_{\mathfrak C}\) such that \(\mathbb{E}\int_0^TD_h(\pi_s|\pi^{n}_s)\,ds<\infty\) for all \(n\in \mathbb{N}\cup\{0\}\), \[J^0(\pi^{n})- J^0(\pi)\le\frac{\lambda}{n}\mathbb{E}\int_0^TD_h(\pi_s|\pi^{0}_s)\,ds\,,\quad n\in \mathbb{N}\,,\]

  • If \(\tau>0\) and \(\pi^\ast \in \mathcal{A}_{\mathfrak C}\) satisfies \(J^\tau (\pi^\ast) = \min_{\pi\in A_{\mathfrak C}} J^\tau (\pi)\) and \(\mathbb{E}\int_0^TD_h(\pi^\ast_s|\pi^{n}_s)\,ds<\infty\) for all \(n\in \mathbb{N}\cup\{0\}\), then \[\label{eq:convergence95regularized} \begin{align} 0 \leq J^\tau(\pi^{n})-J^\tau(\pi^\ast)\le \lambda\left(1-\frac{\tau}{\lambda}\right)^n\mathbb{E}\int_0^TD_h(\pi^\ast_s|\pi^{0}_s)\,ds\,,\quad n\in \mathbb{N}\,. \end{align}\tag{14}\]

The proof of Theorem 8 is given in Section 3.5.

Remark 9. The convergence results in Theorem 8 require prior knowledge that the target control is compatible with the iterates \(\{\pi^n\}_{n\in \mathbb{N}}\), as indicated by their finite Bregman divergence.

When \(\tau=0\), the theorem provides an upper bound of the performance difference \(J^0(\pi^n)-J^0(\pi)\) for any compatible control \(\pi\). This indicates that \(\{\pi^n\}_{n\in \mathbb{N}}\) asymptotically outperforms any compatible control \(\pi\) with a rate of \(\mathcal{O}(1/n)\). It implies the convergence of Algorithm 1, if an optimal control of 2 exists and is compatible to the mirror descent iterates. This result extends the convergence result of mirror descent from the static optimization setting [22] to the dynamic framework considered here.

However, as in the static case [22], we note that due to the absence of regularisation (\(\tau=0\)), the compatibility condition may not hold in general and has to be verified on a case-by-case basis, depending on the choice of the Bregman divergence \(D_h\) (and thus on the function \(h\)). See Remarks 12 and 16 for conditions under which the compatibility condition is satisfied, thereby ensuring the \(\mathcal{O}(1/n)\)-convergence rate for the optimal controls.

When \(\tau>0\), the scenario changes as in this case, the existence of an optimal control \(\pi^\ast\) and its compatibility can be verified by exploiting the regularisation effect of \(h\). In particular, under Assumption 6, the optimal control \(\pi^\ast\in \mathcal{A}_{\mathfrak C}\) can be characterised by the following Pontryagin system: for all \(t\in [0,T]\), \[\label{eq:optimal95FBSDE} \left\{ \begin{align} dX_t(\pi^\ast)&=b_t(X_t(\pi^\ast),\pi^\ast_t)\,dt+\sigma_t(X_t(\pi^\ast),\pi^\ast_t)\,dW_t\,, \quad X_0(\pi^\ast) = x\,, \\ dY_t(\pi^\ast)&=-(D_xH^\tau_t)(\Theta_t(\pi^\ast),\pi^\ast_t)\,dt+Z_t(\pi^\ast)\,dW_t\, , \quad Y_T(\pi^\ast)=(D_xg)(X_T(\pi^\ast))\,, \\ \pi^\ast_t &=\mathop{\mathrm{arg\,min}}_{m\in \mathfrak C} H^\tau_t(\Theta_t(\pi^\ast),m)\,. \end{align} \right.\tag{15}\] We anticipate that using the relative convexity of \(\pi\mapsto J^\tau(\pi)\), the well-posedness of 15 can be studied by the Method of Continuation as in [11], [57], which subsequently allows for verifying the finiteness of \(\mathbb{E}\int_0^TD_h(\pi^\ast_s|\pi^{n}_s)\,ds\). We leave a detailed analysis for future work.

Theorem 8 provides the theoretical foundation for solving the unregularized control problem by introducing gradually reducing regularization parameters, known as annealing. For instance, let \(\{\pi^n_\tau\}_{n\in \mathbb{N}}\) denote the iterates generated by Algorithm 1, applied to the regularized objective 3 with some \(\tau>0\). The performance of \(\{\pi^n_\tau\}_{n\in \mathbb{N}}\) in solving the unregularized objective can be decomposed as \[\begin{align} J^0(\pi^n_\tau)-\inf_{\pi \in \mathcal{A}_{\mathfrak C} }J^0(\pi) & =\left(J^0(\pi^n_\tau)-J^\tau(\pi^n_\tau)\right) + \left(J^\tau(\pi^n_\tau)-\inf_{\pi \in \mathcal{A}_{\mathfrak C}} J^\tau (\pi ) \right)+ \left( \inf_{\pi \in \mathcal{A}_{\mathfrak C}} J^\tau(\pi ) - \inf_{\pi \in \mathcal{A}_{\mathfrak C}} J^0(\pi) \right). \end{align}\] Assuming that the regularisation function \(h\) is nonnegative (e.g., \(h\) is chosen as the relative entropy in Proposition 10, or the \(\chi^2\)-divergence in Proposition 13, or the entropic optimal transport in Proposition 15 with \(c \geq 0\)) we have \[\begin{align} 0\le J^0(\pi^n_\tau)-\inf_{\pi \in \mathcal{A}_{\mathfrak C} }J^0(\pi) \le \left(J^\tau(\pi^n_\tau)-\inf_{\pi \in \mathcal{A}_{\mathfrak C}} J^\tau (\pi ) \right)+ \left( \inf_{\pi \in \mathcal{A}_{\mathfrak C}} J^\tau(\pi ) - \inf_{\pi \in \mathcal{A}_{\mathfrak C}} J^0(\pi) \right). \end{align}\] The first term on the right-hand side represents the optimization error for a regularised problem, and can be quantified by 14 in terms of \(n\) and \(\tau\). The second term is the regularization bias resulting from the additional Bregman divergence in 3 . One can choose \(\tau\) to balance these two terms and thereby optimize the convergence (see, e.g., [53] for cases involving Markovian policies and drift-control problems).

2.4 Examples of the regulariser \(h\)↩︎

This section offers various concrete examples of the regularisation function \(h\) commonly utilised in machine learning. We will specify the action set \(\mathfrak C\) and the associated Bregman divergence \(D_h(\cdot|\cdot)\). Additionally, we will verify Assumption 3, i.e., the well-definedness of the mirror descent step [eq32update32the32control32bregman32msa] in Algorithm 1.

The first example of \(h\) is the commonly used relative entropy, as studied in [3], [5], [6], [14], [21]. The following proposition shows that the Bregman divergence \(D_h(m'|m)\) is the Kullback–Leibler divergence between \(m'\) and \(m\).

Proposition 10 (Relative entropy). Let \(\varrho\in \mathcal{P}(A)\), let \(L^\infty(A)\) be the space of bounded measurable functions equipped with the supremum norm, let \(\mathfrak C = \{m \in \mathcal{P}(A) \mid m\ll \varrho, \| \ln \frac{\mathrm d m}{\mathrm d \varrho}\|_{L^\infty(A)}<\infty\}\) and let \(h:\mathfrak C \to \mathbb{R}\) be the relative entropy given by \[h(m)\mathrel{\vcenter{:}}= \int \ln\frac{\mathrm dm}{\mathrm d\varrho}(a) m(da)\,, \quad m\in \mathfrak C\,.\] Then for all \(m,m'\in \mathfrak C\) and \(a\in A\), \(\frac{\delta h}{\delta m}(m,a) = \ln \frac{\mathrm d m}{\mathrm d \varrho}(a)-h(m)\) and \[\label{eq:bregman95KL} D_h(m'|m) = \int \ln \frac{\mathrm d m'}{\mathrm d m }(a)m'(da)\,.\qquad{(1)}\]

Suppose that Assumptions 1 and 2 hold and the flat derivatives \(\frac{\delta b}{\delta m}, \frac{\delta \sigma}{\delta m}\) and \(\frac{\delta f}{\delta m}\) are uniformly bounded. Then Assumption 3 holds for all \(\pi^0\in \mathcal{A}_{\mathfrak C}\) such that \(\mathbb{E}\int_0^T \big\| \ln \frac{\mathrm d \pi^0_t}{\mathrm d \varrho} \big\|_{L^\infty(A)} dt<\infty\). In this case, \(\mathbb{E}\int_0^T \big\| \ln \frac{\mathrm d \pi^n_t}{\mathrm d \varrho} \big\|_{L^\infty(A)} dt<\infty\) for all \(n\in \mathbb{N}\cup\{0\}\).

The proof of Proposition 10 is given in Section 3.6.

Remark 11. By Pinsker’s inequality and ?? , for all \(m,m'\in \mathfrak C\), \(D_h(m'|m)\ge \frac{1}{2}\|m'-m\|^2_{\rm TV}\), where \(\|\cdot\|_{\rm TV}\) is the total variation norm over the space of signed measures. Hence when \(h\) is the relative entropy, the necessary regularity in the measure component in Assumptions 1, 4 and 5 can be stated with the squared norm \(\|\cdot\|^2_{\rm TV}\) instead of the Bregman divergence \(D_h(\cdot|\cdot)\).

Remark 12. Going back to Theorem 8, in the unregularized setting with \(\tau = 0\), a sufficient condition to ensure an optimal control \(\pi^\ast \in \mathcal{A}_{\mathfrak C}\) (if it exists) is compatible to the iterates \(\{\pi^n\}_{n\in \mathbb{N}}\) (i.e. \(\mathbb{E}\int_0^TD_h(\pi_s|\pi^{n}_s)\,ds<\infty\) for all \(n\in \mathbb{N}\cup\{0\}\)), is that the action space \(A\) is of finite cardinality (say \(A=\{a_1,\ldots,a_N\}\)), the Bregman divergence \(D_h\) is chosen as in Proposition 10 with \(\rho\) being the uniform distribution on \(A\), and the initial guess \(\pi^0\) of Algorithm 1 satisfies \(\mathbb{E}\int_0^T \big\| \ln \frac{\mathrm d \pi^0_t}{\mathrm d \varrho} \big\|_{L^\infty(A)} dt<\infty\) (e.g., \(\pi^0 = \rho\)). In this case, the iterates \(\{\pi^n\}_{n\in \mathbb{N}}\) converge to \(\pi^\ast\) with a rate \(\mathcal{O}(1/n)\), as established in Theorem 8.

To see this, for any \(m\in \mathcal{P}(A)\) we have \(0\leq h(m) \leq -\sum_{k=1}^N \ln \rho(a_k) m(a_k) = \ln N < \infty\). Taking \(\nu := \rho\) we similarly have \(0\leq D_h(m|\nu) \leq \ln N\). Thus the set \(\mathcal{A}_{\mathfrak C}\) defined in 10 coincides with the set of all progressively measurable processes taking values in \(\mathcal{P}(A)\). Let \(\pi^*\in \mathcal{A}_{\mathfrak C}\) be an optimal control. To show that \(\pi^*\) is compatible with \(\{\pi^n\}_{n\in \mathbb{N}}\), note that for \(n\in \mathbb{N}\), \[\begin{align} & \mathbb{E}\int_0^T D_h( \pi^*_t|\pi^n_t) dt =\mathbb{E}\int_0^T \left( \sum_{k=1}^N \left(\ln \frac{\mathrm d \pi^*_t}{\mathrm d \varrho }(a_k)- \ln \frac{\mathrm d \pi^n_t}{\mathrm d \varrho }(a_k)\right) \pi^*_t(a_k) \right)dt\\ & \leq - \mathbb{E}\int_0^T \sum_{k=1}^N \ln \frac{\mathrm d \pi^n_t}{\mathrm d \varrho }(a_k) \pi^*_t(a_k) \leq \mathbb{E}\int_0^T \bigg\|\frac{\mathrm d \pi^n_t}{\mathrm d \varrho }\bigg\|_{L^\infty(A)}\,dt < \infty\,, \end{align}\] where the final inequality follows from Proposition 10, as \(\pi_0\) satisfies \(\mathbb{E}\int_0^T \big\| \ln \frac{\mathrm d \pi^0_t}{\mathrm d \varrho} \big\|_{L^\infty(A)} dt<\infty\).

Next we consider \(h\) as Pearson’s \(\chi^2\)-divergence, which plays a key role in information theory and machine learning [62]. The associated Bregman divergence \(D_h(m'|m)\) is the squared \(L^2\)-norm of the density of \(m'-m\).

Proposition 13 (\(\chi^2\)-divergence).

Let \(\varrho \in \mathcal{P}(A)\) and let \(L^2_\varrho(A)\) be the space of square integrable functions with respect to \(\varrho\). Let \(\mathfrak C = \{m \in \mathcal{P}(A) \mid m\ll \varrho, \frac{\mathrm d m}{\mathrm d \varrho} \in L^2_{\varrho}( A)\}\) and let \(h:\mathfrak C \to \mathbb{R}\) be \(\chi^2\)-divergence given by \[h(m) \mathrel{\vcenter{:}}= \int \frac{1}{2} \left(\frac{\mathrm d m }{\mathrm d\varrho }(a) -1\right) ^2 \varrho (da),\,\quad m\in \mathfrak C\, .\] Then for all \(m,m'\in \mathfrak C\) and \(a\in A\), \(\frac{\delta h}{\delta m}(m,a) = \frac{\mathrm d m}{\mathrm d \varrho}(a)-\int \frac{\mathrm d m}{\mathrm d \varrho}(a) m(da)\) and \[\label{eq:bregman95quadratic} D_h(m'|m) = \int \frac{1}{2}\left( \frac{\mathrm d m'}{\mathrm d \varrho }(a) - \frac{\mathrm d m}{\mathrm d \varrho }(a) \right)^2 \varrho(da)\,.\qquad{(2)}\]

Suppose that Assumptions 1 and 2 hold and for each \(\phi\in \{b,\sigma,f\}\), \[a\mapsto \sup_{(t, x,m)\in [0,T]\times \mathbb{R}^d\times \mathfrak C}|\frac{\delta \phi_t}{\delta m}(x,m,a)|\in L^2_\varrho(A)\,.\] Then Assumption 3 holds for all \(\pi^0\in \mathcal{A}_{\mathfrak C}\) such that \(\mathbb{E}\int_0^T\big \| \frac{\mathrm d \pi^0_t}{\mathrm d \varrho} \big\|^2_{L^2_\varrho(A)} dt<\infty\).

The proof of Proposition 13 is given in Section 3.6.

Remark 14. Recall that for all \(m,m'\in \mathfrak C\), \[\|m'-m\|_{\rm TV}=\frac{1}{2}\int \left| \frac{\mathrm d m'}{\mathrm d \varrho }(a) - \frac{\mathrm d m}{\mathrm d \varrho }(a) \right| \varrho (da)\,,\] which along with Jensen’s inequality and ?? implies \(\|m'-m\|^2_{\rm TV}\le \frac{1}{2}D_h(m'|m)\). Hence when \(h\) is the \(\chi^2\)-divergence, the required regularity in the measure component in Assumptions 1, 4 and 5 can be stated with the squared norm \(\|\cdot\|^2_{\rm TV}\).

Finally, we consider \(h\) as the entropic optimal transport studied in [16]. Let \(A\) be a compact metric space, \(c\in C(A\times A)\), and \(\varrho \in \mathcal{P}(A)\). Let \(\mathfrak C = \mathcal{P}(A)\), and let \(h: \mathfrak C\to \mathbb{R}\) be such that \[\label{eq:transport95cost} h(m)\mathrel{\vcenter{:}}= \min_{\gamma\in \Pi(m,\varrho) } \left(\int_{A\times A} c(x,y) d \gamma (x,y)+\kappa \operatorname{KL}(\gamma| m\otimes \varrho)\right),\tag{16}\] where \(\Pi(m,\varrho)\) is the set of measures \(m\in \mathcal{P}( A\times A)\) with \(m\) and \(\varrho\) as respective first and second marginals, \(\kappa>0\) is a fixed parameter, and \(\operatorname{KL}(\gamma| m\otimes \varrho)=\int_{A\times A}\ln \frac{\mathrm d \gamma }{\mathrm d m\otimes \varrho }d\gamma\) if \(\gamma \ll m\otimes \varrho\) and \(\infty\) otherwise. For each \(m\in \mathfrak C\), consider a pair \((\phi,\psi)\in C(A)\times C(A)\) such that for all \((a,a')\in A\times A\), \[\begin{align} \label{eq:schrodinger} \begin{aligned} \phi(a)= -\kappa \ln\left(\int e^{(\psi(a')-c(a,a'))/\kappa }\varrho(d a')\right)\,,\,\,\, \psi(a')= -\kappa \ln\left(\int e^{(\phi(a)-c(a,a'))/\kappa } m(d a)\right)\,. \end{aligned} \end{align}\tag{17}\] The pair \((\phi,\psi)\in C(A)\times C(A)\) satisfying 17 is often referred to as the Schrödinger potentials and is unique up to the transformation \((\phi+c', \psi-c')\) for \(c'\in \mathbb{R}\) (see [16]). In the sequel, we fix \(a_0\in A\) and choose the unique solution to 17 such that \(\phi(a_0)=0\). We denote this specific choice by \((\phi[m],\psi[m])\) to emphasise its dependence on \(m\).

Proposition 15 (Entropic optimal transport). Let \(A\) be a compact separable metric space, \(c: A\times A\to \mathbb{R}\) be Lipschitz continuous and let \(\varrho \in \mathcal{P}(A)\). Define \(\mathfrak C = \mathcal{P}(A)\) and \(h: \mathfrak C\to \mathbb{R}\) as in 16 . Then for all \(m\in \mathfrak C\) and \(a\in A\), \(\frac{\delta h}{\delta m }(m, a)= \phi [m](a)-\int \phi[m](a)m(da)\), where \((\phi[m],\psi[m])\) are the Schrödinger potentials satisfying 17 .

Suppose that Assumptions 1 and 2 hold, and for each \(\phi\in \{b,\sigma,f\}\) and for all \((t, x,m)\in [0,T]\times \mathbb{R}^d\times \mathfrak C\), \(\frac{\delta \phi_t}{\delta m}(x,m,\cdot) \in C(A)\). Then Assumption 3 holds for all \(\pi^0\in \mathcal{A}_{\mathfrak C}\).

The proof of Proposition 15 is given in Section 3.6.

Remark 16.

In the unregularized setting (with \(\tau =0\)), when the function \(h\) in 10 is chosen as the entropic optimal transport map described in Proposition 15, any optimal control \(\pi^\ast \in \mathcal{A}_{\mathfrak C}\) (if it exists) is compatible to the iterates \(\{\pi^n\}_{n\in \mathbb{N}}\) as required by Theorem 8, which implies that \(\{\pi^n\}_{n\in \mathbb{N}}\) converges to \(\pi^\ast\) with a rate \(\mathcal{O}(1/n)\).

To see this, note that by the compactness of \(A\) and continuity of \(c\), \(\mathfrak C = \mathcal{P}(A)\) is compact, and \(\mathfrak C\ni \mu\mapsto \phi[\mu]\in C(A)\) are continuous (see [16]). Then for any \(m\in \mathfrak C\), as \(\operatorname{KL}(\gamma| m\otimes \varrho)\ge 0\), \[\inf_{x,y\in A} c(x,y) \leq h(m) \leq \int_{A\times A} c(x,y)d(m\otimes \rho)(x,y) \leq \sup_{x,y \in A} c(x,y)\,,\] and \(\|\frac{\delta h}{\delta m }(m, \cdot)\|_{L^\infty}= \left|\phi [m]-\int \phi[m](a)m(da)\right\|_{L^\infty}\le 2 \sup_{m\in \mathfrak C }\|\phi[m]\|_{L^\infty}\). Then for all \(n\in \mathbb{N}\cup \{0\}\), \[\begin{align} 0\le D_h( \pi^*_t|\pi^n_t) &=h(\pi^*_t) - h(\pi^n_t) - \int \frac{\delta h}{\delta m}(\pi^n_t,a)(\pi^*_t-\pi^n_t)(da) \\ & \le \sup_{x,y \in A} c(x,y) - \inf_{x,y\in A} c(x,y) +2 \sup_{m\in \mathfrak C }\|\phi[m]\|_{L^\infty}. \end{align}\] This proves that \(\pi^*\) is compatible to the iterates \(\{\pi^n\}_{n\in \mathbb{N}}\).

3 Proofs↩︎

This section proves the main results in Section 2. For notational simplicity, we denote by \(C\in [0,\infty)\) a generic constant, which depends only on the constants appearing in the assumptions and may take a different value at each occurrence.

3.1 Regularity of adjoint processes↩︎

To prove the relative smoothness of \(\pi\mapsto J^\tau (\pi)\) in Theorem 4, we first prove that the adjoint processes \((Y(\pi), Z(\pi))\) depend Lipschitz continuously on \(\pi\). The proof relies crucially on an a-priori uniform bound of \((Z(\pi))_{\pi\in \mathcal{A}_{\mathfrak C}}\) in terms of the bounded mean oscillation (BMO) martingales. Compared with classical \(L^2\) estimates, this improved regularity of \(Z(\pi)\) allows for handling more general dependences in the diffusion coefficient; see Remark 18 for more details.

We first introduce some notation. For each stopping time \(\eta:\Omega\to [0,T]\), we denote by \(\mathbb{E}_\tau [\cdot]\) the conditional expectation with respect to \(\mathcal{F}_\tau\). We define the norm \(\|\cdot\|_{L^\infty}\) such that \(\|Z\|_{L^\infty}\mathrel{\vcenter{:}}= \operatorname{ess}\sup_{\omega}|Z(\omega)|\) for any random variable \(Z:\Omega\to \mathbb{R}\), and the norm \(\|\cdot\|_{\mathbb{H}^{\infty}}\) such that \(\|Z\|_{\mathbb{H}^{\infty}}\mathrel{\vcenter{:}}= \operatorname{ess}\sup_{(t,\omega)}|Z_t(\omega)|\) for any progressively measurable \(Z:\Omega\times [0,T]\to \mathbb{R}\). For a uniformly integrable martingale \(M\) with \(M_0=0\), define \[\|M\|_{\operatorname{BMO}}=\sup_{\eta}\left\|(\mathbb{E}_{\eta}\left[\langle M\rangle_T-\langle M\rangle_\eta\right])^{\frac{1}{2}}\right\|_{L^\infty}\,,\] where \(\langle M\rangle\) is the quadratic variation of \(M\), and the supremum is taken over all stopping times \(\eta:\Omega\to [0,T]\). If \(M'\) is another uniformly integrable martingale with \(M'_0=0\) then we define the cross-variation of \(M\) with \(M'\) as \(\langle M, M'\rangle_t = \frac{1}{4} [\langle M+M' \rangle_t - \langle M - M' \rangle_t]\) for \(t\in[0,T]\). For a square integrable adapted process \(Z\), we denote by \(Z\cdot W\) the stochastic integral of \(Z\) with respect to the Brownian motion \(W\), i.e., \((Z\cdot W)_t\mathrel{\vcenter{:}}= \int_0^t Z_sdW_s\) for \(t\in[0,T]\).

The following lemma proves the boundedness of the adjoint processes.

Lemma 1. Suppose that there exists \(K\ge 0\) such that for all \(t\in[0,T]\), \(x\in\mathbb{R}^d\) and \(m\in \mathfrak C\), \[|(D_x b_t)(x,m)|+|(D_x\sigma_t)(x,m)|+|(D_xf_t)(x,m)|+|(D_x g)(x)|\le K\,.\] Then \(\sup_{\pi\in\mathcal{A}_{\mathfrak C} }\|Y(\pi)\|_{\mathbb{H}^{\infty}}<\infty\) and \(\sup_{\pi\in \mathcal{A}_{\mathfrak C}}\|Z(\pi)\cdot W\|_{\operatorname{BMO}}<\infty\).

Proof. We first establish the bound of \(Y(\pi)\). By the definition of \(H^0\) in 4 , for all \(s\in [0,T]\) and \(i=1,\ldots, d\), \[\begin{align} D_{x_i}H_s^0(\Theta_s(\pi),\pi_s) &= \left(D_{x_i}b^j_s\right)(X_s(\pi),\pi_s) (Y_s(\pi))^j+\left(D_{x_i}\sigma^{jp}_s\right)(X_s(\pi),\pi_s) (Z_s(\pi))^{jp} + \left(D_{x_i}f_s\right)(X_s(\pi),\pi_s)\,, \end{align}\] where we adopt the Einstein summation convention, i.e., repeated equal dummy indices indicate summation over all the values of that index. As \((Y(\pi),Z(\pi))\) satisfies the linear BSDE 6 , by [63], \[Y_t(\pi)=\mathbb{E}_t\left[S_t^{-1}S_T \left(D_xg\right)(X_T(\pi))+\int_t^T S_t^{-1}S_s\left(D_xf_s\right)(X_s(\pi),\pi_s)\,ds\right]\,,\] where the process \(S=(S^{ij})_{i,j=1}^d\) satisfies \(S_0=I_d\) and for all \(i,j=1,\ldots,d\) and \(t\in [0,T]\), \[dS_t^{ij}=S^{il}_t \left(D_{x_l}b^j_t\right)(X_t(\pi),\pi_t)\,dt+S^{il}_t \left(D_{x_l}\sigma^{jp}_t\right)(X_t(\pi),\pi_t)\,dW_t^p\,,\] and \(S^{-1}_t\), the inverse process of \(S_t\) in the sense that \((S^{-1}_t)^{ij} = 1/S_t^{ij}\) exists in the strong sense. By the Cauchy-Schwarz inequality, \[\begin{align} |Y_t(\pi)|&\le\left(\mathbb{E}_t[|S_t^{-1}S_T|^2]\right)^{1/2} \left(\mathbb{E}_t[|\left(D_xg\right)(X_T(\pi))|]^2\right)^{1/2}\\ &\quad+\left(\mathbb{E}_t\left[ \int_t^T |S_t^{-1}S_s|^2\,ds\right]\right)^{1/2}\left(\mathbb{E}_t\left[ \int_t^T |\left(D_xf_s\right)(X_s(\pi),\pi_s)|^2\,ds\right]\right)^{1/2}\,. \end{align}\] Due to the assumptions of the lemma we can apply [63] together with [63] thus obtaining, \({\sup_{\pi \in \mathcal{A}_{\mathfrak C}}}\mathbb{E}_t\sup_{t\le s\le T}|S_t^{-1}S_s|^2<\infty\) for all \(t\in[0,T]\), and hence \[\sup_{\pi\in\mathcal{A}_{\mathfrak C} }\|Y(\pi)\|_{\mathbb{H}^{\infty}}<\infty\,.\]

We proceed to establish the estimate of \(Z(\pi)\) for a fixed \(\pi \in \mathcal{A}_{\mathfrak C}\). Applying Itô’s formula to \(s\mapsto |Y_s(\pi)|^2\), for all \(s\in [0,T]\), \[\begin{align} \int_s^T|Z_t(\pi)|^2\,dt &\le |\left(D_x g\right)(X_T(\pi))|^2-2\int_s^TZ_t(\pi)^\top Y_t(\pi)\,dW_t +\int_s^T2Y_t(\pi)\left(D_xH^{0}_t\right)(\Theta_t(\pi), \pi_t)\,dt\,. \end{align}\] For any stopping time \(\eta:\Omega\to [0,T]\), by taking the conditional expectation, \[\begin{align} \mathbb{E}_\eta \left[ \int_\eta^T|Z_t(\pi)|^2\,dt\right] & \le \mathbb{E}_\eta [| (D_x g )(X_T(\pi))|^2] +\mathbb{E}_\eta \int_\eta ^T\Big[2Y_t(\pi)^\top \Big( (D_xb_t )(X_t(\pi),\pi_t) Y_t(\pi)+ \\ & \qquad \qquad \qquad \qquad \qquad + (D_x\sigma_t )(X_t(\pi),\pi_t) Z_t(\pi)+ (D_xf_t )(X_t(\pi),\pi_t)\Big)\Big]\,dt\,. \end{align}\] Using the assumptions of the lemma and Young’s inequality, \[\begin{align} \mathbb{E}_\eta \left[ \int_\eta ^T|Z_t(\pi)|^2\,dt \right] & \le \| (D_x g )(X_T(\pi))\|^2_{L^\infty} + 2\|Y(\pi)\|_{\mathbb{H}^\infty}^2 \| (D_xb )(X(\pi),\pi)\|_{\mathbb{H}^\infty} + 2\|Y(\pi)\|_{\mathbb{H}^\infty} \| (D_xf )(X(\pi),\pi)\|_{\mathbb{H}^\infty}\\ &\quad +\mathbb{E}_\eta \left[ \int_\eta ^T \|\left(D_x\sigma\right)(X(\pi),\pi)\|_{\mathbb{H}^\infty}(\gamma^{-1}\|Y(\pi)\|_{\mathbb{H}^\infty}^2+\gamma|Z_t(\pi)|^2) \right]\,dt \\ & \le K^2 + 2K \|Y(\pi)\|_{\mathbb{H}^\infty}^2 + 2K \|Y(\pi)\|_{\mathbb{H}^\infty} +K \mathbb{E}_\eta \left[ \int_\eta ^T (\gamma^{-1}\|Y(\pi)\|_{\mathbb{H}^\infty}^2+\gamma|Z_t(\pi)|^2) \right]\,dt\,. \end{align}\] By choosing \(\gamma\) such that \(0<1-K\gamma<1\) and using the fact that \(\sup_{\pi\in\mathcal{A}_{\mathfrak C} }\|Y(\pi)\|_{\mathbb{H}^{\infty}}<\infty\), we conclude that there exists \(C>0\), independent of \(\pi\) and \(\tau\), such that \(\mathbb{E}_\eta \left[\int_\eta ^T|Z_t(\pi)|^2\right]\,dt\le C\). Hence, by the definition of the \(\operatorname{BMO}\) norm, \[\sup_{\pi\in\mathcal{A}_{\mathfrak C} }\|Z(\pi)\cdot W\|_{\operatorname{BMO}}=\sup_{\pi\in\mathcal{A}_{\mathfrak C} }\sup_{\eta}\left\|\left(\mathbb{E}_{\eta}\left[\int_{\eta}^T|Z_t(\pi)|^2\,dt\right]\right)^{1/2}\right\|_{L^\infty}<\infty \,,\] where the supremum is taken over all stopping times \(\eta:\Omega\to [0,T]\). ◻

We proceed to investigate the Lipschitz continuity of \(\pi\mapsto (Y(\pi),Z(\pi))\). The following lemma shows that the state process \(X(\pi)\) is Lipschitz continuous in \(\pi\) and is a standard SDE stability-type estimate. The proof is given in 6.

Lemma 2. Suppose that there exists \(K\ge 0\) such that for all \(t\in[0,T]\), \(x,x'\in\mathbb{R}^d\) and \(m,m'\in \mathfrak C\), \[|b_t(x,m)-b_t(x',m')|^2+|\sigma_t(x,m)-\sigma_t(x',m')|^2\le K(|x-x'|^2+ D_h(m|m'))\,.\] Then there exists \(C\ge 0\) such that for all \(\pi,\pi'\in \mathcal{A}_{\mathfrak C}\), \[\mathbb{E}\left[ \sup_{0\le t\le T}|X_t(\pi)-X_t(\pi')|^2\right]\le C\mathbb{E}\int_0^T\,D_h(\pi_s|\pi'_s)\,ds\,.\]

We further recall the well-known Fefferman’s inequality for BMO martingales (see [64]).

Lemma 3 (Fefferman’s inequality). If \(M\) is a BMO martingale and \(N\) is a martingale such that \(\mathbb{E}\left[ \langle N\rangle_T^{\frac{1}{2}} \right]<\infty\), then \[\mathbb{E}\left[\int_0^T|d\langle M, N\rangle_t|\right]\le\sqrt{2}\|M\|_{\operatorname{BMO}}\,\mathbb{E}\left[\langle N\rangle_T^{\frac{1}{2}}\right]\,.\]

Fefferman’s inequality yields the following lemma, which is important for the regularity analysis of the cost objective.

Lemma 4. Let \(f, g:\Omega\times [0,T]\to \mathbb{R}\) be square-integrable progressively measurable processes such that \(\|f\cdot W\| _{\operatorname{BMO}}<\infty\), where \(W\) is a Brownian motion. Then \(\mathbb{E}[\int_0^T |f_tg_t| d t]\le \sqrt{2} \|f\cdot W\| _{\operatorname{BMO}} \mathbb{E}\left[\left(\int_0^T |g_t|^2 d t \right)^{\frac{1}{2}}\right]\).

Proof. By Itô’s isometry, \(\mathbb{E}[\int_0^T |f_tg_t| d t]=\mathbb{E}[\int_0^Td\langle |f|\cdot W, |g|\cdot W\rangle_t ]\). The desired result follows from Lemma 3 and the fact that \(\||f|\cdot W\| _{\operatorname{BMO}}=\|f\cdot W\| _{\operatorname{BMO}}\). ◻

Remark 17. Lemma 4 leverages the BMO regularity of \(f\cdot W\) and bounds the \(L^1\)-norm of \((\omega, t)\mapsto f_tg_t(\omega)\) using the \(L^2\)-norm of \(g\) in time and \(L^1\)-norm in the probability space. This norm on \(g\) is weaker than the standard \(L^2\)-norm \(\mathbb{E}\left[ \int_0^T |g_t|^2 d t \right]^{{1}/{2}}\) arising from the Cauchy–Schwarz inequality, and is crucial for our regularity estimates, particularly for the bound of \(I_4\) in Lemma 5.

A similar estimate cannot be obtained by replacing Fefferman’s inequality with the Kunita–Watanabe inequality, stated as follows: \[\int_0^T|d\langle M, N\rangle_t| \le \left(\int_0^T d\langle M\rangle_t\right)^{1/2} \left(\int_0^T d\langle N\rangle_t\right)^{1/2}\,,\] since after taking expectations, it is unclear how to isolate the contribution of \(M\) without invoking its BMO norm. Applying the Cauchy–Schwarz inequality in this context leads to the appearance of \(\mathbb{E}[ \langle N\rangle_T]^{1/2}\), which coincides with the undesired \(L^2\)-norm of \(g\), as discussed above.

Based on Lemmas 1, 2 and 3, we prove that the adjoint processes depend Lipschitz continuously on \(\pi\) in terms of the Bregman divergence.

Lemma 5. Suppose Assumption 1, 2 and 4 hold. Then there exists \(C > 0\) such that for all \(\pi,\pi'\in\mathcal{A}_{\mathfrak C}\), \[\mathbb{E}\sup_{t\in[0,T]} |Y_t(\pi)-Y_t(\pi')|^2 + \mathbb{E}\int_0^T |Z_t(\pi)-Z_t(\pi')|^2dt \le C\mathbb{E}\int_0^T D_h(\pi_t|\pi'_t)\,dt\,.\]

Proof. Fix \(\pi,\pi'\in\mathcal{A}_{\mathfrak C}\). Let \(\beta>0\) be a constant whose value will be specified later. By Itô’s formula, \[\label{eq:bsdes95proof95before95BDG} \begin{align} &e^{\beta s}|Y_s(\pi)-Y_s(\pi')|^2 +\int_s^T e^{\beta t} \left(|Z_t(\pi)-Z_t(\pi')|^2+\beta|Y_t(\pi)-Y_t(\pi')|^2\right)dt \\ & = e^{\beta T}| (D_xg )(X_T(\pi))- (D_x g )(X_T(\pi'))|^2 -2\int_s^T e^{\beta t} (Y_t(\pi)-Y_t(\pi'))^\top (Z_t(\pi)-Z_t(\pi')) dW_t\\ &\quad +\int_s^T e^{\beta t}\bigg(2(Y_t(\pi)-Y_t(\pi'))^\top\left(\left(D_x H^{0}_t\right)(\Theta_t(\pi), \pi_t)-\left(D_xH^{0}_t\right)(\Theta_t(\pi'), \pi'_t)\right)\,. \end{align}\tag{18}\] The proof will be divided into two steps. The first step is to choose a sufficiently large \(\beta\) to absorb the terms involving \(|Z_t(\pi)-Z_t(\pi')|^2\) and \(|Y_t(\pi)-Y_t(\pi')|^2\) arising from upper bounding the right-hand side of 18 , and to establish that for any \(\varepsilon > 0\), \[\label{eq:32final32diff32Y32plus32diff32Z32new} \begin{align} &\mathbb{E}\int_0^T \bigg(\frac{1}{2}|Z_t(\pi)-Z_t(\pi')|^2 + |Y_t(\pi)-Y_t(\pi')|^2\bigg) \,dt \\ & \le C(1+\varepsilon^{-1})\mathbb{E}\int_0^T D_h(\pi_t|\pi'_t)\,dt + \varepsilon \hat{C}\mathbb{E}\left[ \sup_{t\in[0,T]} |Y_t(\pi)-Y_t(\pi')|^2\right]\,, \end{align}\tag{19}\] where \(\hat{C}:= \sup_{\pi \in \mathcal{A}_{\mathfrak C}} \|Z(\pi)\cdot W\|_{\text{BMO}} < \infty\), due to Lemma 1, and \(C\) is a constant that depends on \(K\), \(T\), \(d\) and \(d'\) but is independent of \(\varepsilon\). In the second step we will take the supremum in 18 , use Burkholder–David–Gundy inequality and the estimate 19 with a sufficiently small \(\varepsilon>0\) to conclude the proof.

Step 1: From 18 we get that \[\begin{align} & \int_0^T e^{\beta t}\Big[|Z_t(\pi)-Z_t(\pi')|^2 + \beta|Y_t(\pi) - Y_t(\pi')|^2\Big]\,dt \\ & \leq e^{\beta T}| (D_xg )(X_T(\pi))- (D_x g )(X_T(\pi'))|^2 -2\int_0^T e^{\beta t}(Y_t(\pi)-Y_t(\pi'))^\top (Z_t(\pi)-Z_t(\pi')) dW_t\\ &\quad +\int_0^T e^{\beta t}2(Y_t(\pi)-Y_t(\pi'))^\top \Big(\left(D_x H^{0}_t\right)(\Theta_t(\pi), \pi_t)-\left(D_xH^{0}_t\right)(\Theta_t(\pi'), \pi'_t)\Big)\,dt\,. \end{align}\] Observe that \[\mathbb{E}\int_0^T e^{2\beta s}|Y_s(\pi)-Y_s(\pi')|^2 |Z_s(\pi)-Z_s(\pi')|^2\,ds \leq 2e^{2\beta T}\max(\|Y(\pi)\|_{\mathbb{H}^\infty}^2, \|Y(\pi')\|_{\mathbb{H}^\infty}^2)\| Z(\pi)-Z(\pi')\|_{\mathbb{H}^2}^2\] which is finite due to Lemma 1 and Proposition 2. Thus \(\int_0^t e^{\beta s} (Y_s(\pi)-Y_s(\pi'))^\top (Z_s(\pi)-Z_s(\pi')) dW_s\) is a martingale. Hence, taking expectation yields that \[\label{eq32diff32of32Y32and32Z} \begin{align} & \mathbb{E} \int_0^T e^{\beta t}\Big[|Z_t(\pi)-Z_t(\pi')|^2 + \beta|Y_t(\pi) - Y_t(\pi')|^2\Big]\,dt \leq e^{\beta T}\mathbb{E}|(D_x g)(X_T(\pi))- (D_x g )(X_T(\pi'))|^2 \\ &\quad + \mathbb{E}\int_0^T e^{\beta t}2(Y_t(\pi)-Y_t(\pi'))^\top\Big(\left(D_x H^{0}_t\right)(\Theta_t(\pi), \pi_t)-\left(D_xH^{0}_t\right)(\Theta_t(\pi'), \pi'_t)\Big)\,dt\,. \end{align}\tag{20}\] To derive an upper bound of the right-hand side of 20 , observe that \[\begin{align} &\left(D_x H^{0}_t\right)(\Theta_t(\pi), \pi_t)-\left(D_xH^{0}_t\right)(\Theta_t(\pi'), \pi'_t)\\ &=\left(D_x b_t\right)(X_t(\pi), \pi_t) (Y_t(\pi)-Y_t(\pi') ) + \text{tr}\Big[(D_x \sigma_t) (X_t(\pi), \pi_t) (Z_t(\pi)-Z_t(\pi') ) \Big] \\ &\quad + \big( (D_x b_t )(X_t(\pi), \pi_t)- (D_x b_t )(X_t(\pi'), \pi'_t) \big) Y_t(\pi') \\ &\quad + \text{tr}\Big[\Big(\left(D_x \sigma_t\right)(X_t(\pi), \pi_t)-\left(D_x \sigma_t\right)(X_t(\pi'), \pi'_t)\Big) Z_t(\pi')\Big] + \left(D_x f_t\right)(X_t(\pi), \pi_t)-\left(D_x f_t\right)(X_t(\pi'), \pi'_t)\,. \end{align}\] Hence by Assumptions 1, 2 and 4 and by Young’s inequality, for all \(\gamma>0\), \[\label{eq32diff32y32diff32H} \begin{align} &\mathbb{E}\left[\int_0^T 2e^{\beta t}\Big|(Y_t(\pi)-Y_t(\pi'))^\top ( (D_x H^{0}_t )(\Theta_t(\pi), \pi_t)- (D_xH^{0}_t )(\Theta_t(\pi'), \pi'_t))\Big|\,dt\right]\\ &\leq I_1+I_2+I_3+I_4+I_5\,, \end{align}\tag{21}\] where \[\label{eq32bsde32estimate32all32the32I123-terms} \begin{align} I_1&\mathrel{\vcenter{:}}= \mathbb{E}\left[\int_0^T 2e^{\beta t}|(D_x b)(X(\pi),\pi)| \, |Y_t(\pi)-Y_t(\pi')|^2 \, dt\right]\,, \\ I_2 &\mathrel{\vcenter{:}}= \mathbb{E}\left[\int_0^T e^{\beta t} | (D_x\sigma )(X(\pi),\pi)| \left(\gamma|Z_t(\pi)-Z_t(\pi')|^2 +\gamma^{-1}|Y_t(\pi)-Y_t(\pi')|^2\right)\,dt\right]\,, \\ I_3 &\mathrel{\vcenter{:}}= \mathbb{E}\left[\int_0^Te^{\beta t}\Big(\gamma^{-1}|Y_t(\pi)-Y_t(\pi')|^2+\gamma | (D_x f_t )(X_t(\pi), \pi_t)- (D_x f_t )(X_t(\pi'), \pi'_t)|^2\Big)dt\right]\,, \end{align}\tag{22}\] and \[\begin{align} I_4 &\mathrel{\vcenter{:}}= \mathbb{E}\left[ \int_0^T 2e^{\beta t} \Big|(Y_t(\pi)-Y_t(\pi'))^\top \text{tr}\big[\big( (D_x \sigma_t )(X_t(\pi), \pi_t)- (D_x \sigma_t )(X_t(\pi'), \pi'_t)\big) Z_t(\pi') \big]\Big| dt \right]\,, \\ I_5 &\mathrel{\vcenter{:}}= \mathbb{E}\left[\int_0^T2e^{\beta t}\Big|(Y_t(\pi)-Y_t(\pi'))^\top \Big( (D_x b_t )(X_t(\pi), \pi_t)- (D_x b_t )(X_t(\pi'), \pi'_t)\Big)Y_t(\pi')\Big|\,dt\right]\,. \end{align}\] We focus on the most troublesome terms \(I_4\) and \(I_5\). Observe that by Lemma 4, \[\begin{align} & I_4 \leq 2 \sum_{j=1}^d \sum_{i=1}^{d'}\sum_{k=1}^d \mathbb{E}\int_0^T e^{\beta t} \Big|(Y^j_t(\pi)-Y^j_t(\pi'))\big( (D_{x_j} \sigma^{ki}_t )(X_t(\pi), \pi_t)- (D_{x_j} \sigma^{ki}_t )(X_t(\pi'), \pi'_t)\big) \Big|\, \Big| Z^{ki}_t(\pi')\Big| dt \\ & \leq \sum_{j=1}^d \sum_{i=1}^{d'}\sum_{k=1}^d \|Z^{ki}(\pi') \!\cdot\! W^i \|_{\text{BMO}} \mathbb{E}\left[ \bigg(\int_0^T \!\!\! e^{2\beta t} \big|Y_t^j(\pi)-Y_t^j(\pi')\big|^2 \big|(D_{x_j} \sigma_t^{ki} )(X_t(\pi), \pi)- (D_{x_j} \sigma_t^{ki} )(X_t(\pi'), \pi')\big|^2dt\bigg)^{\frac{1}{2}}\!\right] \\ & \leq \|Z(\pi')\!\cdot\! W \|_{\text{BMO}} \sum_{j=1}^d \sum_{i=1}^{d'}\sum_{k=1}^d \mathbb{E}\bigg[ \bigg(\int_0^T \!\! e^{2\beta t} \big|Y_t^j(\pi)-Y_t^j(\pi')\big|^2 \big|(D_{x_j} \sigma_t^{ki} )(X_t(\pi), \pi)- (D_{x_j} \sigma_t^{ki} )(X_t(\pi'), \pi')\big|^2\,dt\bigg)^{\frac{1}{2}}\bigg]\,. \end{align}\] By Young’s inequality, for all \(i\in \{1,\ldots, d'\}\), \(j,k\in \{1,\ldots, d\}\), and \(\varepsilon > 0\), \[\begin{align} & \mathbb{E}\bigg[ \sup_{0\leq t\leq T} \big|Y_t^j(\pi)-Y_t^j(\pi')\big| \bigg(\int_0^T e^{2\beta t} \big|(D_{x_j} \sigma_t^{ki} )(X_t(\pi), \pi)- (D_{x_j} \sigma_t^{ki} )(X_t(\pi'), \pi')\big|^2 dt\bigg)^{1/2}\bigg]\\ & \leq \mathbb{E}\bigg[ \varepsilon \sup_{0\leq t\leq T} \big|Y_t^j(\pi)-Y_t^j(\pi')\big|^2 + \varepsilon^{-1} \int_0^T e^{2\beta t} \big|(D_{x_j} \sigma_t^{ki} )(X_t(\pi), \pi)- (D_{x_j} \sigma_t^{ki} )(X_t(\pi'), \pi')\big|^2\,dt\bigg] \,. \end{align}\] Setting \(\hat{C}:= \sup_{\pi \in \mathcal{A}_{\mathfrak C}} \|Z(\pi)\cdot W\|_{\text{BMO}} < \infty\) (see Lemma 1) and taking the summation over \(i,j,k\) yield \[\begin{align} \label{eq32diff32Y32Z32diff32sigma} \begin{aligned} I_4 & \le \varepsilon \hat{C} d' d \mathbb{E}\bigg[ \sup_{0\leq t\leq T} \big|Y_t(\pi)-Y_t(\pi')\big|^2\bigg] + \hat{C} \mathbb{E}\int_0^T e^{2\beta t} \big|(D_{x} \sigma_t)(X_t(\pi), \pi)- (D_x \sigma_t )(X_t(\pi'), \pi')\big|^2\,dt \\ &\leq \varepsilon \hat{C} d' d \mathbb{E}\left[ \sup_{t\in[0,T]} |Y_t(\pi)-Y_t(\pi')|^2\right] + K \hat{C}\varepsilon ^{-1} \mathbb{E}\left[ \int_0^Te^{2\beta t} \big[|X_t(\pi) - X_t(\pi')|^2 + D_h(\pi_t|\pi'_t)\big] dt \right]\,. \end{aligned} \end{align}\tag{23}\] where the last inequality used the regularity assumption of \(D_x\sigma\). Hereafter, we denote by \(C\ge 0\) a generic constant that depends on \(K\), \(T\), \(d\), \(d'\) but is independent of \(\beta, \gamma\) and \(\varepsilon\), and may vary from line to line. Again, using Lemma 1, Young’s inequality and Assumption 4, \[\label{eq32diffY32Y32diffb} \begin{align} |I_5| & \le \|Y(\pi')\|_{\mathbb{H}^{\infty}}\mathbb{E}\bigg[ \int_0^T e^{\beta t}2|Y_t(\pi)-Y_t(\pi')| | (D_x b_t )(X_t(\pi), \pi_t)- (D_x b_t )(X_t(\pi'), \pi'_t) |\,dt \bigg]\\ & \le C\mathbb{E}\bigg[ \int_0^T e^{\beta t}\left( |Y_t(\pi)-Y_t(\pi')|^2+ | (D_x b_t )(X_t(\pi), \pi_t)- (D_x b_t )(X_t(\pi'), \pi'_t) |^2\right)\,dt \bigg]\\ &\le C\mathbb{E}\bigg[\int_0^T e^{\beta t}\left( |Y_t(\pi)-Y_t(\pi')|^2+ K(|X_t(\pi)-X_t(\pi')|^2+ D_h(\pi_t|\pi'_t) \right)\,dt \bigg]\,. \end{align}\tag{24}\] Hence using 212223 , and 24 and Assumption 2, for all \(\beta\ge 0\) and \(\gamma, \varepsilon>0\), \[\label{eq:32final32diff32Y32plus32diff32Z} \begin{align} &\mathbb{E}\left[\int_0^T 2e^{\beta t}\Big|(Y_t(\pi)-Y_t(\pi'))^\top ( (D_x H^{0}_t )(\Theta_t(\pi), \pi_t)- (D_xH^{0}_t )(\Theta_t(\pi'), \pi'_t))\Big|\,dt\right]\\ & \leq C \mathbb{E}\left[\int_0^T 2e^{\beta t}|Y_t(\pi)-Y_t(\pi')|^2\,dt\right] \\ & \quad + C\mathbb{E}\left[\int_0^T e^{\beta t}\left(\gamma|Z_t(\pi)-Z_t(\pi')|^2 +\gamma^{-1}|Y_t(\pi)-Y_t(\pi')|^2\right)\,dt\right] \\ & \quad + \mathbb{E}\left[\int_0^Te^{\beta t}\Big(\gamma^{-1}|Y_t(\pi)-Y_t(\pi')|^2+ C \gamma \big[|X_t(\pi)-X_t(\pi')|^2 + |D_h(\pi_t|\pi'_t)\big]\Big)dt\right] \\ & \quad + \varepsilon d' d \hat{C}\mathbb{E}\left[ \sup_{t\in[0,T]} |Y_t(\pi)-Y_t(\pi')|^2\right] + C\varepsilon ^{-1} \mathbb{E}\left[ \int_0^Te^{2\beta t} \big[|X_t(\pi) - X_t(\pi')|^2 + D_h(\pi_t|\pi'_t)\big] dt \right] \\ & \quad + C\mathbb{E}\bigg[\int_0^T e^{\beta t}\left( |Y_t(\pi)-Y_t(\pi')|^2+ D_h(\pi_t|\pi'_t)\right)\,dt \bigg]\,. \end{align}\tag{25}\] Therefore, using 25 in 20 and collecting terms along with Lemma 2 shows that \[\label{eq:lipschitz95estimate1} \begin{align} &\mathbb{E}\left[\int_0^T e^{\beta t}\big[|Z_t(\pi)-Z_t(\pi')|^2 + \beta |Y_t(\pi)-Y_t(\pi')|^2 \big]\,dt\right] \\ &\le C \mathbb{E}\left[\int_0^T e^{\beta t}(1 + \gamma^{-1} )|Y_t(\pi)-Y_t(\pi')|^2\,dt\right] + C\mathbb{E}\left[\int_0^T e^{\beta t}\left(\gamma|Z_t(\pi)-Z_t(\pi')|^2 \right)\,dt\right] \\ & \quad + C (1+\gamma+ \varepsilon^{-1}) e^{2\beta T} \mathbb{E}\left[\int_0^T D_h(\pi_t|\pi'_t)dt\right] + \varepsilon d' d \hat{C}\mathbb{E}\left[ \sup_{t\in[0,T]} |Y_t(\pi)-Y_t(\pi')|^2\right]\,, \end{align}\tag{26}\] Taking the constant \(C\) in 26 , and choosing \(\gamma=1/(2C)\) and \(\beta = 1 +C(1 + C)\) yield for all \(\varepsilon>0\), \[\label{eq:32pre32final32diff32Y32plus32diff32Z32new} \begin{align} &\mathbb{E}\left[\int_0^T \bigg(\frac{1}{2}|Z_t(\pi)-Z_t(\pi')|^2 + |Y_t(\pi)-Y_t(\pi')|^2\bigg) \,dt\right] \\ & \le C\left(1+\frac{1}{2C}+\varepsilon^{-1}\right)\mathbb{E}\int_0^T D_h(\pi_t|\pi'_t)\,dt + \varepsilon d' d \hat{C} \mathbb{E}\left[ \sup_{t\in[0,T]} |Y_t(\pi)-Y_t(\pi')|^2\right]\,. \end{align}\tag{27}\] This shows 19 holds thus concluding Step 1.

Step 2: We return to 18 and upper bound \(\mathbb{E}\left[ \sup_{s\in [0,T]} |Y_s(\pi)-Y_s(\pi')|^2 \right]\). For each \(s\in [0,T]\), \[\begin{align} & e^{\beta s}|Y_s(\pi)-Y_s(\pi')|^2 + \int_s^T e^{\beta t}\Big(|Z_t(\pi)-Z_t(\pi')|^2 + \beta|Y_t(\pi)-Y_t(\pi')|^2\Big)\,dt \\ & \leq e^{\beta T}| (D_xg )(X_T(\pi))- (D_x g )(X_T(\pi'))|^2 \\ &\quad +2\sup_{s\in[0,T]}\bigg| \int_s^T e^{\beta t}(Y_t(\pi)-Y_t(\pi'))^\top (Z_t(\pi)-Z_t(\pi')) dW_t\bigg|\\ &\quad + \sup_{s\in[0,T]}\bigg|\int_s^T e^{\beta t} 2(Y_t(\pi)-Y_t(\pi'))^\top \Big( (D_x H^{0}_t )(\Theta_t(\pi), \pi_t)- (D_xH^{0}_t )(\Theta_t(\pi'), \pi'_t) \Big)\,dt\bigg|\,. \end{align}\] Taking the supremum over \(s\in [0,T]\) and then taking the expectation yield \[\begin{align} & \mathbb{E}\left[ \sup_{s\in [0,T]} e^{\beta s}|Y_s(\pi)-Y_s(\pi')|^2 \right] \leq e^{\beta T} \mathbb{E}\Big[|(D_xg )(X_T(\pi))- (D_x g )(X_T(\pi'))|^2\Big] \\ &\qquad \qquad \qquad + 2\mathbb{E} \left[\sup_{s\in[0,T]}\bigg| \int_s^T e^{\beta t} (Y_t(\pi)-Y_t(\pi')) ^\top (Z_t(\pi)-Z_t(\pi')) dW_t\bigg|\right]\\ &\qquad \qquad \qquad + \mathbb{E}\int_0^T e^{\beta t} 2 \big| (Y_t(\pi)-Y_t(\pi'))^\top \Big( (D_x H^{0}_t )(\Theta_t(\pi), \pi_t)- (D_xH^{0}_t )(\Theta_t(\pi'), \pi'_t) \Big) \big|\,dt\,. \end{align}\] Setting \(\beta=0\) and using 25 , Lemma 2 and the conclusion of Step 1, which is 19 , shows that \[\label{eq32proof32of32lem32346432step32232before32BDG} \begin{align} \mathbb{E} \left[\sup_{s\in [0,T]} |Y_s(\pi)-Y_s(\pi')|^2 \right] & \leq 2\mathbb{E} \left[ \sup_{s\in[0,T]}\bigg| \int_s^T (Y_t(\pi)-Y_t(\pi'))^\top (Z_t(\pi)-Z_t(\pi'))dW_t\bigg|\right] \\ &\quad + C(1+\varepsilon^{-1})\mathbb{E} \left[ \int_0^T D_h(\pi_t|\pi'_t)\,dt\right] + \varepsilon \hat{C} \mathbb{E}\left[ \sup_{t\in[0,T]} |Y_t(\pi)-Y_t(\pi')|^2\right]\,. \end{align}\tag{28}\] By the Burkholder–Davis–Gundy inequality and Young’s inequality, for all \(\gamma>0\), \[\begin{align} & \mathbb{E}\left[\sup_{s\in[0,T]}\bigg| \int_s^T (Y_t(\pi)-Y_t(\pi'))^\top (Z_t(\pi)-Z_t(\pi')) dW_t\bigg|\right] \leq C \mathbb{E} \Bigg[ \bigg(\int_0^T |Z_t(\pi)-Z_t(\pi')|^2|Y_t(\pi)-Y_t(\pi')|^2\,dt\bigg)^{1/2} \Bigg]\\ & \leq C \mathbb{E} \Bigg[ \sup_{t\in [0,T]}|Y_t(\pi)-Y_t(\pi')| \bigg(\int_0^T |Z_t(\pi)-Z_t(\pi')|^2\,dt\bigg)^{1/2} \Bigg]\\ & \leq C\left(\gamma \mathbb{E} \left[ \sup_{s\in [0,T]}|Y_s(\pi)-Y_s(\pi')|^2\right] + \gamma^{-1} \mathbb{E} \left[\int_0^T |Z_t(\pi)-Z_t(\pi')|^2\,dt \right] \right)\,. \end{align}\] Choosing \(\gamma := 1/(2C)\), using the above estimate in 28 yield, for all \(\varepsilon > 0\), that \[\begin{align} & \frac{1}{2}\mathbb{E}\left[ \sup_{s\in [0,T]} |Y_s(\pi)-Y_s(\pi')|^2\right]\\ & \le C \mathbb{E}\left[ \int_0^T |Z_t(\pi)-Z_t(\pi')|^2\,dt \right] + C(1+\varepsilon^{-1})\mathbb{E}\left[ \int_0^T D_h(\pi_t|\pi'_t)\,dt \right] + \varepsilon \hat{C}\mathbb{E}\left[ \sup_{t\in[0,T]} |Y_t(\pi)-Y_t(\pi')|^2\right] \end{align}\] Using the conclusion of Step 1, which is 19 , one more time we have, for all \(\varepsilon > 0\), that \[\mathbb{E}\left[ \sup_{s\in [0,T]} |Y_s(\pi)-Y_s(\pi')|^2\right] \leq C(1+\varepsilon^{-1})\mathbb{E}\left[ \int_0^T D_h(\pi_t|\pi'_t)\,dt \right] + \varepsilon 2\hat{C}\mathbb{E}\left[ \sup_{t\in[0,T]} |Y_t(\pi)-Y_t(\pi')|^2\right]\,.\] Finally, taking \(\varepsilon = \frac{1}{2 \hat{C}}\) leads to \[\begin{align} \frac{1}{2}\mathbb{E} \sup_{s\in [0,T]} |Y_s(\pi)-Y_s(\pi')|^2 & \leq C\mathbb{E}\left[ \int_0^T D_h(\pi_t|\pi'_t)\,dt \right]\,. \end{align}\] This, together with 19 concludes the proof. ◻

Remark 18. The above argument leverages the theory of BMO martingales and allows for incorporating more general diffusion coefficients, beyond the scope of the classical \(L^2\)-approach. Indeed, to obtain the desired Lipschitz regularity, standard \(L^2\)-estimate in [57] requires us to bound \(\mathbb{E}\big[\big( \int_0^T |(D_x \sigma_s)(X_s(\pi),\pi_s)-(D_x \sigma_s)(X_s(\pi'),\pi'_s)||Z_s(\pi) | \,ds \big)^2 \big]\). Under Assumption 4, it reduces to upper bound \(\mathbb{E}\big[\sup_{s\in [0,T]}| X_s(\pi)- X_s(\pi')|^2 \big( \int_0^T |Z_s(\pi) | \,ds \big)^2 \big]\). As \(Z(\pi)\) is not uniformly bounded, it is unclear how to bound the term using \(\mathbb{E}\int_0^T D_h(\pi_t|\pi'_t)\,dt\). Here we develop a more refined analysis and handle the \(Z\) dependence using the uniform BMO bound in Lemma 1. Similar remarks apply to the proof of Lemma 7.

3.2 Proof of Theorem 4↩︎

This section proves the relative smoothness of the cost functional \(\pi\mapsto J^\tau(\pi)\), namely Theorem 4.

We start by recalling that the Gâteaux derivative of \(J^0\) can be expressed by the flat derivative of the Hamiltonian. The proof relies on charactersing the pathwise derivative of \(\varepsilon\mapsto X(\pi+\varepsilon(\pi'-\pi))\). Compared with [21], additional complexities arise due to the possible unboundedness of flat derivatives of the coefficients \((b, \sigma, f)\) with respect to the control \(m \in \mathfrak C\) and the fact that the regularity of these coefficients in \(m\) is characterized by the Bregman divergence \(D_h\). The detailed arguments can be found in 5.

Lemma 6.

Suppose Assumptions 1, 2, 4 and 5 hold. Then for all \(\pi,\pi'\in\mathcal{A}_\mathfrak{C}\) such that \(\mathbb{E} \int_0^T D_h(\pi'_t|\pi_t)\,dt < \infty\), \[\frac{d}{d\varepsilon} J^0(\pi+\varepsilon(\pi'-\pi))\big|_{\varepsilon=0}=\mathbb{E}\int_0^T\int\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi),\pi_s,a)(\pi'_s-\pi_s)(da)\,ds\,.\]

The following lemma proves the regularity of \(\frac{\delta H^{0}}{\delta m}\) with respect to \(\pi\), which will be used in the proof of Theorem 4.

Lemma 7.

Suppose Assumptions 1, 2, 4 and 5 hold. Then there exists \(C\ge 0\) such that for all \(\pi,\pi'\in\mathcal{A}_\mathfrak{C}\), \[\begin{align} &\mathbb{E}\int_0^1\int_0^T\left[\int\left(\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi^{\varepsilon}), \pi^{\varepsilon}_s,a)-\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi), \pi_s,a)\right)(\pi'_s-\pi_s)(da) \right]ds\,d\varepsilon \le C\mathbb{E}\int_0^T D_h(\pi'_s|\pi_s)\,ds\,, \end{align}\] where \(\pi^{\varepsilon}=\pi+\varepsilon(\pi'-\pi)\) for \(\varepsilon\in(0,1)\).

Proof. Note that \[\begin{align} &\mathbb{E}\int_0^1\int_0^T \int\left(\frac{\delta H^{0}_s}{\delta m}(\Theta(\pi^{\varepsilon}),\pi_s^{\varepsilon},a) - \frac{\delta H^{0}_s}{\delta m}(\Theta(\pi),\pi_s,a)\right)(\pi'_s-\pi_s)(da)\,ds\,d\varepsilon = I^{(1)}+I^{(2)}\,, \end{align}\] where \[\begin{align} I^{(1)} &\mathrel{\vcenter{:}}= \mathbb{E}\int_0^1\int_0^T \int\left(\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi^{\varepsilon}), \pi^{\varepsilon}_s,a)-\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi), \pi^{\varepsilon}_s,a)\right)(\pi'_s-\pi_s)(da) ds\,d\varepsilon\,,\\ I^{(2)} &\mathrel{\vcenter{:}}= \mathbb{E}\int_0^1\int_0^T \int\left(\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi), \pi^{\varepsilon}_s,a)-\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi), \pi_s,a)\right)(\pi'_s-\pi_s)(da) ds\,d\varepsilon\,. \end{align}\]

To estimate the term \(I^{(1)}\), observe that for each \(\varepsilon\in (0,1)\), \[\begin{align} &\int\left[\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi^{\varepsilon}), \pi^{\varepsilon}_s,a)-\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi), \pi^{\varepsilon}_s,a)\right](\pi'_s-\pi_s)(da) \\ & = \int\Bigg[\left(\frac{\delta b_s}{\delta m}(X_s(\pi^{\varepsilon}),\pi^{\varepsilon}_s,a)-\frac{\delta b_s}{\delta m}(X_s(\pi),\pi^{\varepsilon}_s,a)\right)Y_s(\pi^{\varepsilon}) +\frac{\delta b_s}{\delta m}(X_s(\pi),\pi^{\varepsilon}_s,a)\Big(Y_s(\pi^{\varepsilon})-Y_s(\pi)\Big) \\ & \qquad+\left(\frac{\delta \sigma_s}{\delta m}(X_s(\pi^{\varepsilon}),\pi^{\varepsilon}_s,a)-\frac{\delta \sigma_s}{\delta m}(X_s(\pi),\pi^{\varepsilon}_s,a)\right) Z_s(\pi^{\varepsilon}) +\frac{\delta \sigma_s}{\delta m}(X_s(\pi),\pi^{\varepsilon}_s,a) (Z_s(\pi^{\varepsilon})-Z_s(\pi) )\\ & \qquad+\frac{\delta f_s}{\delta m}(X_s(\pi^{\varepsilon}),\pi^{\varepsilon}_s,a)-\frac{\delta f_s}{\delta m}(X_s(\pi),\pi^{\varepsilon}_s,a)\Bigg](\pi'_s-\pi_s)(a)\,da\,. \end{align}\] Fix \(\varepsilon\in (0,1)\). By the fundamental theorem of calculus we have that \[\frac{\delta b_s}{\delta m}(X_s(\pi^{\varepsilon}),\pi^{\varepsilon}_s,a)-\frac{\delta b_s}{\delta m}(X_s(\pi),\pi^{\varepsilon}_s,a)=\int_0^1 \left( D_x \frac{\delta b_s}{\delta m}(X_s^{\varepsilon, \varepsilon'},\pi^{\varepsilon}_s,a)\right)(X_s(\pi^{\varepsilon})-X_s(\pi))\,d\varepsilon'\,,\] where \(X_s^{\varepsilon, \varepsilon'}\mathrel{\vcenter{:}}= X_s(\pi)+\varepsilon'(X_s(\pi^{\varepsilon})-X_s(\pi))\) for all \(\varepsilon'\in(0,1)\). Thus \[\begin{align} \int&\left[\frac{\delta b_s}{\delta m}(X_s(\pi^{\varepsilon}),\pi^{\varepsilon}_s,a)-\frac{\delta b_s}{\delta m}(X_s(\pi),\pi^{\varepsilon}_s,a)\right] Y_s(\pi^{\varepsilon})(\pi'_s-\pi_s)(d a) \\ &\le |X_s(\pi^{\varepsilon})-X_s(\pi)||Y_s(\pi^{\varepsilon})| \int_0^1\left|\int \left( D_x \frac{\delta b_s}{\delta m}(X_s^{\varepsilon, \varepsilon'},\pi^{\varepsilon}_s,a)\right) (\pi'_s-\pi_s)(da) \right|\,d\varepsilon'\,. \end{align}\] Applying similar arguments to \(\sigma\) and \(f\) and using Young’s inequality, we have that \[\begin{align} &\mathbb{E}\int_0^1\int_0^T\left[\int\left(\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi^{\varepsilon}), \pi^{\varepsilon}_s,a)-\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi), \pi_s^{\varepsilon},a)\right)(\pi'_s-\pi_s)(da) \right]ds\,d\varepsilon \le I_1 + I_2 + I_3 + I_4\,, \end{align}\] where \[\begin{align} I_1 & \mathrel{\vcenter{:}}= \int_0^1\mathbb{E}\int_0^T\int\bigg(\frac{\delta b_s}{\delta m}(X_s(\pi),\pi^{\varepsilon}_s,a)(Y_s(\pi^{\varepsilon})-Y_s(\pi))\\ &\quad +\frac{\delta \sigma_s}{\delta m}(X_s(\pi),\pi^{\varepsilon}_s,a) (Z_s(\pi^{\varepsilon})-Z_s(\pi))\bigg)(\pi'_s-\pi_s)(da) \,ds\,d\varepsilon\,, \\ I_2 & \mathrel{\vcenter{:}}= \int_0^1\mathbb{E}\int_0^T|Y_s(\pi^{\varepsilon})|\bigg[ \frac{1}{2}|X_s(\pi^{\varepsilon})-X_s(\pi)|^2 \\ & \quad + \frac{1}{2}\left(\int_0^1\left|\int \left( D_x \frac{\delta b_s}{\delta m}(X_s^{\varepsilon, \varepsilon'},\pi^{\varepsilon}_s,a)\right) (\pi'_s-\pi_s)(a)\,da\right|\,d\varepsilon'\right)^2\bigg]\,ds\,d\varepsilon\,, \\ I_3 &\mathrel{\vcenter{:}}= \int_0^1\mathbb{E}\int_0^T|Z_s(\pi^{\varepsilon})| |X_s(\pi^{\varepsilon})-X_s(\pi)| \int_0^1\left|\int \left( D_x \frac{\delta \sigma_s}{\delta m}(X_s^{\varepsilon, \varepsilon'},\pi^{\varepsilon}_s,a)\right) (\pi'_s-\pi_s)(da) \right|\,d\varepsilon'\,ds\,d\varepsilon \\ I_4 &\mathrel{\vcenter{:}}= \int_0^1\mathbb{E}\int_0^T\bigg[\frac{1}{2}|X_s(\pi^{\varepsilon})-X_s(\pi)|^2 +\frac{1}{2}\left(\int_0^1\left|\int \left( D_x \frac{\delta f_s}{\delta m}(X_s^{\varepsilon, \varepsilon'},\pi^{\varepsilon}_s,a)\right) (\pi'_s-\pi_s)(da) \right|\,d\varepsilon'\right)^2\bigg]\,ds\,d\varepsilon\,. \end{align}\] Observe that by Fubini’s theorem and the Cauchy–Schwarz inequality, \[\begin{align}\label{eq32prod32of32diff32Y32and32diff32pi} |I_1|&\le \int_0^1\left(\mathbb{E}\int_0^T|Y_s(\pi^{\varepsilon})-Y_s(\pi)|^{2}\,ds\right)^{1/2} \left(\mathbb{E}\int_0^T\left|\int\frac{\delta b_s}{\delta m}(X_s(\pi),\pi^{\varepsilon}_s,a)(\pi'_s-\pi_s)(da) \right|^2\,ds\,\right)^{1/2}\,d\varepsilon\\ &\quad +\int_0^1\left(\mathbb{E}\int_0^T|Z_s(\pi^{\varepsilon})-Z_s(\pi)|^{2}\,ds\right)^{1/2} \left(\mathbb{E}\int_0^T\left|\int\frac{\delta \sigma_s}{\delta m}(X_s(\pi),\pi^{\varepsilon}_s,a)(\pi'_s-\pi_s)(da) \right|^2\,ds\,\right)^{1/2}\,d\varepsilon\,, \end{align}\tag{29}\] which along with Lemma 5, Assumption 5, and the convexity of \(D_h(\cdot|\pi_s)\) implies \[\begin{align} |I_1|& \le 2 \int_0^1\left(C\mathbb{E}\int_0^TD_h(\pi^{\varepsilon}_s|\pi_s) \,ds\right)^{1/2} \left(K\mathbb{E}\int_0^TD_h(\pi'_s|\pi_s)\,ds\,\right)^{1/2}\,d\varepsilon\\ &\le C \int_0^1\varepsilon^{1/2}\left( \mathbb{E}\int_0^T D_h(\pi'_s|\pi_s)\,ds\right) \,d\varepsilon \le C\mathbb{E}\int_0^T D_h(\pi'_s|\pi_s)\,ds\,, \end{align}\] where the last inequality used \(D_h(\pi^{\varepsilon}_s|\pi_s)\le \varepsilon D_h(\pi'_s|\pi_s)+(1-\varepsilon) D_h(\pi_s|\pi_s) = \varepsilon D_h(\pi'_s|\pi_s)\) since \(\pi^{\varepsilon}=\pi+\varepsilon(\pi'-\pi)\) for \(\varepsilon\in(0,1)\).

For the term \(I_2\), recall that by Lemma 1, \(\sup_{\pi\in\mathcal{A}_{\mathfrak C}}\|Y(\pi)\|_{\mathbb{H}^\infty}<\infty\). Then using Jensen’s inequality, Assumption 5, the convexity of the Bregman divergence and Lemma 2, \[\begin{align} I_2&\le \frac{\sup_{\varepsilon\in [0,1]}\|Y(\pi^{\varepsilon})\|_{\mathbb{H}^\infty}}{2} \int_0^1\mathbb{E}\left[\sup_{0\le s\le T}|X_s(\pi^{\varepsilon})-X_s(\pi)|^2+K\int_0^TD_h(\pi'_s|\pi_s)\,ds\right]\,d\varepsilon\\ &\le \frac{C}{2} \int_0^1\mathbb{E}\left[\int_0^TD_h(\pi^{\varepsilon}_s|\pi_s)\,ds+ \int_0^TD_h(\pi'_s|\pi_s)\,ds\right]\,d\varepsilon \le C \mathbb{E}\int_0^TD_h(\pi'_s|\pi_s)\,ds\,. \end{align}\] Similar arguments show that \[\begin{align} I_4\le C\mathbb{E}\int_0^TD_h(\pi'_s|\pi_s)\,ds\,. \end{align}\]

To estimate \(I_3\), we define for all \(s\in [0,T]\) and \(\varepsilon',\varepsilon \in (0,1)\), \[A^{\varepsilon,\varepsilon'}_s \mathrel{\vcenter{:}}= \left|\int \left( D_x \frac{\delta \sigma_s}{\delta m}(X_s^{\varepsilon, \varepsilon'},\pi^{\varepsilon}_s,a)\right) (\pi'_s-\pi_s)(da) \right|\,.\] By Assumption 5, \(|A^{\varepsilon,\varepsilon'}_s|^2\le KD_h(\pi'_s|\pi_s)\). Using Lemmas 1 and 4, Young’s inequality, Lemma 2 and the convexity of the Bregman divergence, \[\begin{align} I_3 &= \int_0^1\mathbb{E}\int_0^T|Z_s(\pi^{\varepsilon})||X_s(\pi^{\varepsilon})-X_s(\pi)| \int_0^1A^{\varepsilon,\varepsilon'}_s\,d\varepsilon'\,ds\,d\varepsilon\\ &\le C\int_0^1\|Z(\pi^{\varepsilon})\cdot W\|_{\operatorname{BMO}}\, \mathbb{E}\left[\left(\int_0^T|X_s(\pi^{\varepsilon})-X_s(\pi)|^2 \left(\int_0^1A^{\varepsilon,\varepsilon'}_s\,d\varepsilon'\right)^2\,ds\right)^{\frac{1}{2}}\right]\,d\varepsilon\\ &\le C\int_0^1\|Z(\pi^{\varepsilon})\cdot W\|_{\operatorname{BMO}}\, \mathbb{E}\left[\sup_{s\in [0,T]}|X_s(\pi^{\varepsilon})-X_s(\pi)|\left(\int_0^T \int_0^1 (A^{\varepsilon,\varepsilon'}_s)^2\,d \varepsilon' ds\right)^{\frac{1}{2}}\right] \,d\varepsilon\\ &\le C\int_0^1 \mathbb{E}\left[\sup_{s\in [0,T]}|X_s(\pi^{\varepsilon})-X_s(\pi)|^2\,+ \int_0^T\int_0^1 (A^{\varepsilon,\varepsilon'}_s)^2\,d \varepsilon' ds\right] d\varepsilon\\ &\le C\int_0^1\mathbb{E}\left[ \int_0^TD_h(\pi^{\varepsilon}_s|\pi_s)\,ds+ \int_0^TD_h(\pi'_s|\pi_s)\,ds\right]\,d\varepsilon \le C \mathbb{E}\int_0^TD_h(\pi'_s|\pi_s)\,ds\,. \end{align}\] Combining estimates for \(I_1,\,I_2,\,I_3\) and \(I_4\) yields that \(I^{(1)}\le C \mathbb{E}\int_0^TD_h(\pi'_s|\pi_s)\,ds\).

To estimate \(I^{(2)}\), by setting \(\pi^{\varepsilon',\varepsilon} = \varepsilon'\pi^\varepsilon+(1-\varepsilon')\pi\) with \(\pi^\varepsilon = \pi + \varepsilon(\pi'-\pi)\), the fundamental theorem of calculus gives \[\begin{align} I^{(2)} &= \mathbb{E}\int_0^1\int_0^T\int\left(\frac{\delta H^{0}_s}{\delta m}(\Theta(\pi),\pi_s^{\varepsilon},a) - \frac{\delta H^{0}_s}{\delta m}(\Theta(\pi),\pi_s,a)\right)(\pi'_s-\pi_s)(da)\,ds\,d\varepsilon\\ & = \int_0^1 \varepsilon\, \mathbb{E}\int_0^T\int \int_0^1 \int \frac{\delta^2 H^{0}_s}{\delta m^2}(\Theta(\pi),\pi_s^{\varepsilon', \varepsilon},a,a')(\pi'_s-\pi_s)(da')\,d\varepsilon'(\pi'_s-\pi_s)(da)\,ds\,d\varepsilon\,, \end{align}\] where we used \(\pi^\varepsilon- \pi = \varepsilon(\pi'-\pi)\). For each \(s\in [0,T]\) and \(\varepsilon, \varepsilon'\in (0,1)\), \[\label{eq:2nd95H} \begin{align} & \int\frac{\delta^2 H^{0}_s}{\delta m^2}(\Theta_s(\pi), \pi^{\varepsilon',\varepsilon}_s,a,a')(\pi'_s-\pi_s)(da')\\ &= \int\Bigg[\frac{\delta^2 b_s}{\delta m^2}(X_s(\pi),\pi^{\varepsilon',\varepsilon}_s,a,a')Y_s(\pi) +\frac{\delta^2 f_s}{\delta m^2}(X_s(\pi),\pi^{\varepsilon',\varepsilon}_s,a,a')\Bigg](\pi'_s-\pi_s)(da')\,, \end{align}\tag{30}\] where we used \(\frac{\delta^2 \sigma_s}{\delta m^2}(X_s(\pi),\pi^{\varepsilon',\varepsilon}_s,a,a') =0\) due to Assumption 5. Applying Assumption 5 and Lemma 1 yields that \(I^{(2)} \le C \mathbb{E}\int_0^TD_h(\pi'_s|\pi_s)\,ds\,.\) This gives the desired estimate and finishes the proof. ◻

Now we are ready to prove Theorem 4.

Proof of Theorem 4. Fix \(\pi,\pi'\in\mathcal{A_{\mathfrak C}}\). We assume without loss of generality that \(\mathbb{E}\int_{0}^{T} D_h(\pi'_s|\pi_s) \, ds<\infty\), since otherwise the inequality holds trivially.

By Lemma 6 and the fundamental theorem of calculus (see Remark 1), \[\begin{align} J^\tau(\pi')-J^\tau(\pi)&=\int_0^1\mathbb{E}\int_0^T\left[\int\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi^{\varepsilon}), \pi^{\varepsilon}_s,a)(\pi'_s-\pi_s)(a)\,da\right]\,ds\,d\varepsilon + \tau\,\mathbb{E}\int_0^T\left[h(\pi'_s)-h(\pi_s)\right]\,ds\,. \end{align}\] where \(\pi^\varepsilon=\pi+\varepsilon(\pi'-\pi)\) for all \(\varepsilon\in(0,1)\). By the definition of \(D_h(\pi'_s|\pi_s)\), \[\begin{align} \mathbb{E}\int_{0}^{T}\left[h(\pi'_s) - h(\pi_s)\right]\,ds =\mathbb{E}\int_{0}^{T}\left(D_h(\pi'_s|\pi_s) + \int \frac{\delta h}{\delta m}(\pi_s,a)(\pi'_s-\pi_s)(da)\right)\,ds\,, \end{align}\] Hence we have \[\label{eq32expansion32of32the32cost32functional} \begin{align} & J^{\tau}(\pi')-J^{\tau}(\pi) - \mathbb{E}\int_0^T\int\frac{\delta H^{\tau}_s}{\delta m}(\Theta_s(\pi),\pi_s,a)(\pi'_s-\pi_s)(da)\,ds\\ & = \int_0^1\mathbb{E}\int_0^T\int\left(\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi^{\varepsilon}),\pi_s^{\varepsilon},a) - \frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi),\pi_s,a)\right)(\pi'_s-\pi_s)(da)\,ds\,d\varepsilon + \tau\,\mathbb{E}\int_{0}^{T}D_h(\pi'_s|\pi_s)\,ds \\ &\le C\,\mathbb{E}\int_{0}^{T}D_h(\pi'_s|\pi_s)\,ds \,, \end{align}\tag{31}\] where the last inequality used Lemma 7 and Fubini’s theorem. This completes the proof. ◻

3.3 Proof of Theorem 5↩︎

Proof of Theorem 5. Let \(\{\pi^{n}\}_{n\ge 0}\) be the iterates from Algorithm 1. Applying Lemma 4 for \(\pi'=\pi^{n+1}\) and \(\pi=\pi^{n}\) gives that \[\label{eq32lq32difference32of32J32in32the32proof} \begin{align} J^\tau(\pi^{n+1})-J^\tau(\pi^{n}) &\le\mathbb{E}\int_0^T\int\frac{\delta H^{\tau}_s}{\delta m}(\Theta_s(\pi^{n}), \pi^{n}_s, a)(\pi^{n+1}_s-\pi^{n}_s)(da)\,ds +L\mathbb{E}\int_0^T D_h(\pi^{n+1}_s|\pi^{n}_s) \,ds\,. \end{align}\tag{32}\] The definition of \(\pi^{n+1}\) in [eq32update32the32control32bregman32msa] shows that for all \(s\in[0,T]\), \[\begin{align} &\int\frac{\delta H^{\tau}_s}{\delta m}(\Theta_s(\pi^{n}), \pi^{n}_s, a)(\pi^{n+1}_s-\pi^{n}_s)(da)+\lambda\,D_h(\pi^{n+1}_s|\pi^{n}_s) \\ &\le \int\frac{\delta H^{\tau}_s}{\delta m}(\Theta_s(\pi^{n}), \pi^{n}_s, a)(\pi^{n}_s-\pi^{n}_s)(da)+\lambda\,D_h(\pi^{n}_s|\pi^{n}_s) = 0\,. \end{align}\] Thus for all \(n\in \mathbb{N}\cup\{0\}\), \[\label{eq32proof32energy32diss321} \begin{align} &J^\tau(\pi^{n+1})-J^\tau(\pi^{n})\le (L-\lambda) \mathbb{E}\int_0^TD_h(\pi^{n+1}_s|\pi^{n}_s)\,ds\,, \end{align}\tag{33}\] which along with \(D_h(\pi^{n+1}_s|\pi^{n}_s)\ge 0\) and \(\lambda\ge L\) shows that \(J^\tau(\pi^{n+1})-J^\tau(\pi^{n})\le 0\).

We further re-write 33 as \[\begin{align} \mathbb{E}\int_0^TD_h(\pi^{n+1}_s|\pi^{n}_s)\,ds \leq (\lambda - L)^{-1}(J^\tau(\pi^{n}) - J^\tau(\pi^{n+1}))\,. \end{align}\] Summing up over \(n=0,1,\ldots,m-1\) and observing the telescoping sum yields \[\begin{align} \sum_{n=0}^{m-1} \mathbb{E}\int_0^TD_h(\pi^{n+1}_s|\pi^{n}_s)\,ds & \leq (\lambda - L)^{-1}(J^\tau(\pi^{0}) - J^\tau(\pi^{m})) \leq (\lambda - L)^{-1}\left(J^\tau(\pi^{0}) - \inf_{\pi \in\mathcal{A}_{\mathfrak C}} J^\tau(\pi)\right)<\infty\,. \end{align}\] Thus by \(\inf_{\pi \in\mathcal{A}_{\mathfrak C}} J^\tau(\pi)>-\infty\), \(\sum_{n=0}^\infty \mathbb{E}\int_0^TD_h(\pi^{n+1}_s|\pi^{n}_s)\,ds < \infty\) which yields the desired conclusion. ◻

3.4 Proof of Theorem 6↩︎

Proof of Theorem 6. Fix \(\pi,\pi'\in\mathcal{A_{\mathfrak C}}\). By the definition of \(J^\tau\), \[\begin{align} J^{\tau}(\pi)-J^{\tau}(\pi') &=\mathbb{E}\int_{0}^{T}\left(f_s(X_s(\pi),\pi_s)-f_s(X_s(\pi'),\pi_s^{'})+\tau\,h(\pi_s)-\tau\,h(\pi_s^{'})\right)\,ds + \mathbb{E}\left[g(X_T(\pi))-g(X_T(\pi'))\right]\,. \end{align}\] Applying the convexity of \(g\) and Itô’s formula to \(t\mapsto Y_t(\pi)^\top X_t(\pi)\) give that \[\begin{align} J^{\tau}(\pi)-J^{\tau}(\pi') &\le\mathbb{E}\int_{0}^{T}\left( f_s(X_s(\pi),\pi_s)-f_s(X_s(\pi'),\pi_s')+\tau\,h(\pi_s)-\tau\,h(\pi_s')\right) \,ds\\ &\qquad +\mathbb{E}\left[(D_xg)(X_T(\pi))^\top (X_T(\pi)-X_T(\pi'))\right]\\ &\le\mathbb{E}\int_{0}^{T}\left( f_s(X_s(\pi),\pi_s)-f_s(X_s(\pi'),\pi_s')+\tau\,h(\pi_s)-\tau\,h(\pi_s')\right) ds\\ &\qquad +\mathbb{E}\left[\int_0^T(Y_s(\pi))^\top d(X_s(\pi)-X_s(\pi'))+\int_0^T(X_s(\pi)-X_s(\pi'))^\top dY_s(\pi)\right]\\ &\qquad +\mathbb{E}\int_0^T\operatorname{tr}((\sigma_s(X_s(\pi),\pi_s)-\sigma_s(X_s(\pi'),\pi'_s))^\top Z_s(\pi))\,ds\,. \end{align}\] This implies that \[\begin{align} & J^{\tau}(\pi)-J^{\tau}(\pi') \le \mathbb{E}\int_{0}^{T}\left( f_s(X_s(\pi),\pi_s)-f_s(X_s(\pi'),\pi_s')+\tau\,h(\pi_s)-\tau\,h(\pi_s')\right) ds\\ & \,\, +\mathbb{E}\int_0^T(Y_s(\pi))^\top (b_s(X_s(\pi),\pi_s)-b_s(X_s(\pi'),\pi_s'))\,ds\\ & \,\, - \mathbb{E}\int_0^T(X_s(\pi)-X_s(\pi'))^\top (D_xH^0_s)(\Theta_s(\pi^{n}),\pi^{n}_s)\,ds +\mathbb{E}\int_0^T\operatorname{tr}((\sigma_s(X_s(\pi),\pi_s)-\sigma_s(X_s(\pi'),\pi'_s))^\top Z_s(\pi))\,ds\,, \end{align}\] where Itô integrals vanish under the expectation due to Lemma 1. Observe that \[\begin{align} & H^{\tau}_s(\Theta_s(\pi), \pi_s) = f_s(X_s(\pi),\pi_s) + \tau\,h(\pi_s) + (Y_s(\pi))^\top b_s(X_s(\pi),\pi_s) + \operatorname{tr}(\sigma_s(X_s(\pi),\pi_s)^\top Z_s(\pi))\,, \end{align}\] and \[\begin{align} & H^{\tau}_s(X_s(\pi'),Y_s(\pi),Z_s(\pi), \pi_s') =f_s(X_s(\pi'),\pi_s') + \tau\,h(\pi_s') + (Y_s(\pi))^\top b_s(X_s(\pi'),\pi_s') + \operatorname{tr}(\sigma_s(X_s(\pi'),\pi'_s)^\top Z_s(\pi))\,. \end{align}\] Thus, \[\label{eq32lq32inequality32with32star} \begin{align} J^{\tau}(\pi)-J^{\tau}(\pi') &\le \mathbb{E}\int_0^T\Big[H^{\tau}_s(\Theta_s(\pi),\pi_s)-H^{\tau}_s(X_s(\pi'),Y_s(\pi),Z_s(\pi), \pi_s')\Big]\,ds \\ &\qquad - \mathbb{E}\left[\int_0^T(X_s(\pi)-X_s(\pi'))^\top (D_xH^0_s)(\Theta_s(\pi),\pi_s)\,ds\right]\,. \end{align}\tag{34}\] By Assumption 6, \[\begin{align} &\mathbb{E}\int_0^T\Big[H^{0}_s(\Theta_s(\pi),\pi_s)-H^{0}_s(X_s(\pi'),Y_s(\pi),Z_s(\pi), \pi_s')\Big]\,ds\\ &\le \mathbb{E}\left[\int_0^T(X_s(\pi)-X_s(\pi'))^\top (D_xH^0_s)(\Theta_s(\pi),\pi_s)\,ds\right] + \mathbb{E}\int_0^T\int\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi^{n}), \pi_s, a)(\pi_s-\pi'_s)(da)\,ds\,. \end{align}\] Substituting the above inequality to 34 yields \[\begin{align} J^{\tau}(\pi)-J^{\tau}(\pi') & \le \mathbb{E}\int_0^T\int\frac{\delta H^{0}_s}{\delta m}(\Theta_s(\pi), \pi_s, a)(\pi_s-\pi'_s)(da)\,ds +\tau \mathbb{E}\int_0^T[h(\pi_s)-h(\pi'_s)]\,ds \\ &= \mathbb{E}\int_0^T \int\frac{\delta H^{\tau}_s}{\delta m}(\Theta_s(\pi), \pi_s, a)(\pi_s-\pi'_s)(da)\,ds - \tau\mathbb{E}\int_{0}^{T}\left[D_h(\pi'_s|\pi_s) \right]\,ds\,. \end{align}\] where the last identity used \[\mathbb{E}\int_{0}^{T}\left[h(\pi_s) - h(\pi'_s)\right]\,ds =\mathbb{E}\int_{0}^{T}\left[-D_h(\pi'_s|\pi_s) + \int \frac{\delta h}{\delta m}(\pi_s,a)(\pi_s-\pi'_s)(da)\right]\,ds\,.\] This completes the proof. ◻

3.5 Proof of Theorem 8↩︎

Proof of Theorem 8. For all \(s\in [0,T]\) and \(\omega \in \Omega\), define \[\mathfrak C \ni m \mapsto G_s(m) = \frac{1}{\lambda}\int \frac{\delta H^\tau_s}{\delta m}(\Theta_s(\pi), \pi_s, a)(m - \pi_s)(da) \in \mathbb{R}\,,\] which is linear and hence convex. The three point lemma applied to \(G_s\) (see Lemma 10 or [22]) implies for all \(s\in [0,T]\), \(\omega \in \Omega\) and \(\pi,\pi'\in \mathfrak C\), \[\begin{align} & \int\frac{\delta H^{\tau}_s}{\delta m}(\Theta_s(\pi), \pi_s, a)( \pi'_s-\pi_s)(da) + \lambda\,D_h( \pi'_s|\pi_s) \\ & \ge \int\frac{\delta H^{\tau}_s}{\delta m}(\Theta_s(\pi), \pi_s, a)(\overline{\pi}_s-\pi_s)(da)+ \lambda \,D_h(\pi'_s|\overline{\pi}_s) + \lambda\,D_h(\overline{\pi}_s|\pi_s) \,, \end{align}\] provided that \(\overline{\pi}\) satisfies \[\overline{\pi} \in \mathop{\mathrm{arg\,min}}_{m\in \mathfrak C}\int\frac{\delta H^{\tau}_s}{\delta m}(\Theta_s(\pi), \pi_s, a)(m -\pi_s)(da) + \lambda\,D_h(m|\pi_s)\,.\] Hence let \(\pi'\in \mathcal{A}_{\mathfrak C}\) such that \(\mathbb{E}\int_0^TD_h(\pi'_s|\pi^{n}_s)\,ds<\infty\) for all \(n\in \mathbb{N}\cup\{0\}\). By taking \(\pi=\pi^n\) and \(\overline{\pi}=\pi^{n+1}\), and using Theorem 4, [eq32update32the32control32bregman32msa] and the fact that \(L\le \lambda\), \[\label{eq32three32point32plus32smoothness} \begin{align} J^\tau(\pi^{n+1}) & \le J^\tau(\pi^{n}) + \mathbb{E}\int_0^T\int\frac{\delta H^{\tau}_s}{\delta m}(\Theta_s(\pi^{n}), \pi^{n}_s, a)(\pi^{n+1}_s-\pi^{n}_s)(da)\,ds + L \mathbb{E}\int_0^TD_h(\pi^{n+1}_s|\pi^{n}_s)\,ds \\ &\le J^\tau(\pi^{n}) + \mathbb{E}\int_0^T\int\frac{\delta H^{\tau}_s}{\delta m}(\Theta_s(\pi^{n}), \pi^{n}_s, a)(\pi'_s-\pi^{n}_s)(da)\,ds\\ &\quad+(L-\lambda)\mathbb{E}\int_0^TD_h(\pi^{n+1}_s|\pi^{n}_s)\,ds + \lambda\mathbb{E}\int_0^TD_h(\pi'_s|\pi^{n}_s)\,ds - \lambda\mathbb{E}\int_0^TD_h(\pi'_s|\pi^{n+1}_s)\,ds\\ &\leq J^\tau(\pi^{n}) + \mathbb{E}\int_0^T\int\frac{\delta H^{\tau}_s}{\delta m}(\Theta_s(\pi^{n}), \pi^{n}_s, a)(\pi'_s-\pi^{n}_s)(da)\,ds\\ &\quad + \lambda\mathbb{E}\int_0^TD_h(\pi'_s|\pi^{n}_s)\,ds - \lambda\mathbb{E}\int_0^TD_h(\pi'_s|\pi^{n+1}_s)\,ds\,. \end{align}\tag{35}\] By Theorem [th32relative32convexity] and the assumption that \(\mathbb{E}\int_{0}^{T}D_h(\pi^{\ast}_s|\pi^n_s)\,ds<\infty\), \[\begin{align} J^{\tau}(\pi') - \tau\mathbb{E}\int_{0}^{T}D_h(\pi'_s|\pi^n_s)\,ds & \geq J^{\tau}(\pi^n) + \mathbb{E}\int_0^T\int\frac{\delta H^{\tau}_s}{\delta m}(\Theta_s(\pi^n), \pi^n_s, a)(\pi'_s-\pi^n_s)(da)\,ds \,. \end{align}\] Plugging this into 35 gives \[\label{eq32key32inequlity32conv} \begin{align} J^\tau(\pi^{n+1}) - J^{\tau}(\pi') &\leq \lambda \left(\mathbb{E}\int_0^TD_h(\pi'_s|\pi^{n}_s)\,ds - \mathbb{E}\int_0^TD_h(\pi'_s|\pi^{n+1}_s)\,ds\, \right) - \tau\mathbb{E}\int_{0}^{T}D_h(\pi'_s|\pi^n_s) \,ds\,. \end{align}\tag{36}\] When \(\tau=0\), 36 implies \[\sum_{n=0}^{m-1} (J^\tau(\pi^{n+1}) - J^{\tau}(\pi^{\ast})) \leq \lambda \left(\mathbb{E}\int_0^TD_h(\pi'_s|\pi^{0}_s)\,ds - \mathbb{E}\int_0^TD_h(\pi'_s|\pi^{m}_s)\,ds\, \right) \,.\] By Theorem 5, \(J^\tau(\pi^{m})\leq J^\tau(\pi^{n+1})\) for \(n=0,\ldots,m-1\), and hence \[m \left( J^\tau(\pi^{m}) - J^{\tau}(\pi') \right) \leq \lambda \mathbb{E}\int_0^TD_h(\pi'_s|\pi^{0}_s)\,ds \,.\]

Let \(\tau>0\) and fix \(\pi^\ast\in \mathop{\mathrm{arg\,min}}_{\pi\in \mathcal{A}_{\mathfrak C}}J^\tau (\pi)\) such that \(\mathbb{E}\int_0^TD_h(\pi^\ast_s|\pi^{n}_s)\,ds<\infty\) for all \(n\in \mathbb{N}\cup\{0\}\). Setting \(\pi'=\pi^\ast\) in 36 and using \(J^\tau(\pi^{n+1}) \ge J^{\tau}(\pi^{\ast})\) implies \[\mathbb{E}\int_0^TD_h(\pi^\ast_s|\pi^{n+1}_s)\,ds \leq \left(1- \frac{\tau}{\lambda}\right)\mathbb{E}\int_{0}^{T} D_h(\pi^{\ast}_s|\pi^n_s) \,ds \,.\] Iterating this inequality for all \(m\) times gives \[\label{} \mathbb{E}\int_0^TD_h(\pi^\ast_s|\pi^{m}_s)\,ds \leq \left(1- \frac{\tau}{\lambda}\right)^m\mathbb{E}\int_{0}^{T} D_h(\pi^{\ast}_s|\pi^0_s) \,ds \,.\tag{37}\] By 36 , \[J^\tau(\pi^{m+1}) - J^{\tau}(\pi^{\ast}) \leq \lambda \mathbb{E}\int_0^TD_h(\pi^\ast_s|\pi^{m}_s)\,ds \,.\] This concludes the proof. ◻

3.6 Proofs of Propositions 10, 13 and 15↩︎

To prove Proposition 10, we first show that the relative entropy has a flat derivative on the set of probability measures whose log-density admits a polynomial growth.

Lemma 8. Let \(p\in \mathbb{N}\cup\{0\}\), let \((A, \rho_A)\) be a metric space, let \(\mathcal{P}_p(A)\) be the set of probability measures on \(A\) with finite \(p\)-moments, and let \(B_p(A)\) be the space of measurable functions \(f :A \rightarrow \mathbb{R}\) such that \(\sup_{a\in A }\frac{|f(a)|}{1+\rho_A(a_0,a)^p}<\infty\) for some \(a_0\in A\). Let \(\lambda \in \mathcal{P}_p(A)\), let \(\mathfrak C = \{m \in \mathcal{P}_p(A) \mid m\ll \lambda, \ln \frac{\mathrm d m}{\mathrm d \lambda}\in B_p( A)\}\) and let \(\operatorname{KL}(\cdot|\cdot):\mathfrak C\times \mathfrak C \to \mathbb{R}\) be such that \[\operatorname{KL}(\mu |\nu) \mathrel{\vcenter{:}}= \int \ln\frac{\mathrm d\mu }{\mathrm d\nu}(a) \mu (da),\,\quad \mu,\nu\in \mathfrak C\, .\] Then for all \(\mu, \nu \in \mathfrak C\) and \(a\in A\), \[\frac{\delta \operatorname{KL}(\cdot|\nu)}{\delta m}(\mu,a) = \ln \frac{\mathrm d\mu}{\mathrm d\nu}(a)- \operatorname{KL}(\mu|\nu)\,.\]

Proof. Note that \(\mathfrak C\) is convex, and for all \(\mu, \mu'\in \mathfrak C\), \(\int \big|\ln \frac{\mathrm d\mu}{\mathrm d\nu}(a)\big|\mu'(da)<\infty\) as well as \(\int \big(\ln \frac{\mathrm d\mu}{\mathrm d\nu}(a)- \operatorname{KL}(\mu|\nu)\big)\mu(da)=0\). Let \(\mu'\in \mathfrak C\), and for each \(\varepsilon\in (0,1)\), let \(\mu^\varepsilon = \mu + \varepsilon(\mu'-\mu)\). It remains to show that \[\label{eq32entropy32der} \lim_{\varepsilon \to 0} \frac{1}{\varepsilon} \left(\operatorname{KL}(\mu^\varepsilon|\nu) - \operatorname{KL}(\mu|\nu)\right) = \int \ln \frac{\mathrm d\mu}{\mathrm d\nu}(a) (\mu' - \mu)(da)\,.\tag{38}\]

We first prove that \[\label{eq32entropy32der32lb} \limsup_{\varepsilon\searrow 0} \frac{1}{\varepsilon} \left(\operatorname{KL}(\mu^\varepsilon|\nu)-\operatorname{KL}(\mu|\nu)\right) \geq \int \ln \frac{\mathrm d\mu}{\mathrm d\nu}(a) (\mu' - \mu)(da)\,.\tag{39}\] Indeed, for all \(\varepsilon \in (0,1)\), \[\begin{align} & \frac{1}{\varepsilon} \left(\operatorname{KL}(\mu^\varepsilon|\nu)-\operatorname{KL}(\mu|\nu)\right) = \frac{1}{\varepsilon}\int\left[ \frac{\mathrm d\mu^\varepsilon}{\mathrm d\nu}(a)\ln \frac{\mathrm d\mu^\varepsilon}{\mathrm d\nu}(a) - \frac{\mathrm d\mu}{\mathrm d\nu}(a)\ln \frac{\mathrm d\mu}{\mathrm d\nu}(a) \right]\nu(da) \\ & = \frac{1}{\varepsilon} \int \bigg[\bigg(\frac{\mathrm d\mu^\varepsilon}{\mathrm d\nu}(a) - \frac{\mathrm d\mu}{\mathrm d\nu}(a)\bigg)\ln \frac{\mathrm d\mu}{\mathrm d\nu}(a) + \frac{\mathrm d\mu^\varepsilon}{\mathrm d\nu}(a)\bigg(\ln \frac{\mathrm d\mu^\varepsilon}{\mathrm d\nu}(a)-\ln \frac{\mathrm d\mu}{\mathrm d\nu}(a)\bigg)\bigg]\nu(da)\\ & = \int \ln \frac{\mathrm d\mu}{\mathrm d\nu}(a)(\mu'-\mu)(da) + \frac{1}{\varepsilon} \int \frac{\mathrm d \mu^\varepsilon}{\mathrm d \mu}(a) \ln \frac{\mathrm d \mu^\varepsilon}{\mathrm d \mu}(a)\mu(da)\,. \end{align}\] Since \(x\ln x \geq x -1\) for \(x> 0\), it holds for all \(\varepsilon\in (0,1)\) that \[\begin{align} \frac{1}{\varepsilon} \left(\operatorname{KL}(\mu^\varepsilon|\nu)-\operatorname{KL}(\mu|\nu)\right) & \geq \int \ln \frac{\mathrm d\mu}{\mathrm d\nu}(a)(\mu'-\mu)(da) + \frac{1}{\varepsilon} \int \bigg(\frac{\mathrm d \mu^\varepsilon}{\mathrm d \mu}(a) -1\bigg)\mu(da) \\ & = \int \ln \frac{\mathrm d\mu}{\mathrm d\nu}(a)(\mu'-\mu)(da) + \frac{1}{\varepsilon} \bigg( \int \mu^\varepsilon(da) - \int \mu(da)\bigg)\,. \end{align}\] Taking the limit inferior for \(\varepsilon\to 0\) yields 39 .

Next we show that \[\label{eq32entropy32der32limsup32ub} \limsup_{\varepsilon \searrow 0}\frac{1}{\varepsilon} \left(\operatorname{KL}(\mu^\varepsilon|\nu)-\operatorname{KL}(\mu|\nu)\right) \leq \int \ln \frac{\mathrm d\mu}{\mathrm d\nu}(a) (\mu' - \mu)(da)\,.\tag{40}\] To prove 40 , note that \[\label{eq32entropy32der32limsup32ub321} \begin{align} & \frac{1}{\varepsilon} \left(\operatorname{KL}(\mu^\varepsilon|\nu) - \operatorname{KL}(\mu|\nu)\right) \\ & = \frac{1}{\varepsilon} \int \left[\big(\ln \frac{\mathrm d \mu^\varepsilon}{\mathrm d \lambda}(a) - \ln\frac{\mathrm d \nu}{\mathrm d\lambda}(a)\big)\frac{\mathrm d\mu^\varepsilon}{\mathrm d\lambda}(a) - \big(\ln \frac{\mathrm d \mu}{\mathrm d \lambda}(a) - \ln \frac{\mathrm d \nu}{\mathrm d \lambda}(a)\big) \frac{\mathrm d \mu}{\mathrm d \lambda}(a)\right] \lambda(da) = \int f_\varepsilon(a) \,\lambda(da)\,, \end{align}\tag{41}\] where \[f_\varepsilon (a) = \frac{1}{\varepsilon}\left[\frac{\mathrm d \mu^\varepsilon}{\mathrm d \lambda} \ln \frac{\mathrm d \mu^\varepsilon}{\mathrm d \lambda}(a) - \frac{\mathrm d \mu}{\mathrm d \lambda}\ln \frac{\mathrm d \mu}{\mathrm d \lambda}(a) - \ln \frac{\mathrm d \nu}{\mathrm d \lambda}(a)\big( \frac{\mathrm d \mu^\varepsilon}{\mathrm d \lambda} - \frac{\mathrm d \mu}{\mathrm d \lambda}(a)\big) \right]\,.\] Observe that \[\label{eq32entropy32der32limsup32ub322} - \frac{1}{\varepsilon} \ln \frac{\mathrm d \nu}{\mathrm d \lambda} (a)\left(\frac{\mathrm d \mu^\varepsilon}{\mathrm d \lambda}(a) - \frac{\mathrm d \mu}{\mathrm d \lambda}(a)\right) = -\ln \frac{\mathrm d \nu}{\mathrm d \lambda}(a)\left(\frac{\mathrm d \mu'}{\mathrm d \lambda}(a) - \frac{\mathrm d \mu}{\mathrm d \lambda}(a)\right)\,.\tag{42}\] Moreover, using the convexity of \(a\mapsto a\ln a\) on \((0,\infty)\) and the definition of \(\mu^\varepsilon\), \[\label{eq32entropy32der32limsup32ub323} \frac{1}{\varepsilon}\left[ \frac{\mathrm d \mu^\varepsilon}{\mathrm d \lambda}(a)\ln\frac{\mathrm d \mu^\varepsilon}{\mathrm d \lambda}(a) - \frac{\mathrm d \mu}{\mathrm d \lambda}(a)\ln \frac{\mathrm d \mu}{\mathrm d \lambda}(a) \right] \leq \frac{\mathrm d \mu'}{\mathrm d \lambda}(a)\ln\frac{\mathrm d \mu'}{\mathrm d \lambda}(a) - \frac{\mathrm d \mu}{\mathrm d \lambda}(a)\ln \frac{\mathrm d \mu}{\mathrm d \lambda}(a)\,.\tag{43}\] This implies that \(f_\varepsilon \le g\) with \(g\mathrel{\vcenter{:}}= \frac{\mathrm d \mu'}{\mathrm d \lambda}\ln\frac{\mathrm d \mu'}{\mathrm d \lambda} - \frac{\mathrm d \mu}{\mathrm d \lambda}\ln \frac{\mathrm d \mu}{\mathrm d \lambda} -\ln \frac{\mathrm d \nu}{\mathrm d \lambda} (\frac{\mathrm d \mu'}{\mathrm d \lambda} - \frac{\mathrm d \mu}{\mathrm d \lambda})\), and \[\begin{align} \begin{aligned} & \int |g(a)|\lambda (da) = \int \left|\frac{\mathrm d \mu'}{\mathrm d \lambda}(a)\ln\frac{\mathrm d \mu'}{\mathrm d \lambda}(a) - \frac{\mathrm d \mu}{\mathrm d \lambda}(a)\ln \frac{\mathrm d \mu}{\mathrm d \lambda}(a) -\ln \frac{\mathrm d \nu}{\mathrm d \lambda}(a) (\frac{\mathrm d \mu'}{\mathrm d \lambda}(a) - \frac{\mathrm d \mu}{\mathrm d \lambda}(a)) \right|\lambda(da) \\ &\le \int \left(\left| \frac{\mathrm d \mu'}{\mathrm d \lambda}(a)\left(\ln\frac{\mathrm d \mu'}{\mathrm d \lambda}(a) - \ln \frac{\mathrm d \nu}{\mathrm d \lambda}(a)\right)\right| + \left| \frac{\mathrm d \mu}{\mathrm d \lambda}(a) \left(\ln \frac{\mathrm d \mu}{\mathrm d \lambda}(a) -\ln \frac{\mathrm d \nu}{\mathrm d \lambda}(a) \right) \right| \right)\lambda(da) \\ &\le \int \left|\ln\frac{\mathrm d \mu'}{\mathrm d \lambda}(a) - \ln \frac{\mathrm d \nu}{\mathrm d \lambda}(a)\right| \mu'(da) + \int \left| \ln \frac{\mathrm d \mu}{\mathrm d \lambda}(a) -\ln \frac{\mathrm d \nu}{\mathrm d \lambda}(a) \right|\mu(da) < \infty\,, \end{aligned} \end{align}\] as \(\nu, \mu,\mu'\in \mathfrak C\). Hence combining 41 and the reverse Fatou’s lemma yields \[\label{eq32entropy32der32limsup32ub324} \begin{align} & \limsup_{\varepsilon\searrow 0} \frac{1}{\varepsilon} \left(\operatorname{KL}(\mu^\varepsilon|\nu) - \operatorname{KL}(\mu|\nu)\right) \\ & \leq \int \limsup_{\varepsilon\searrow 0} \frac{1}{\varepsilon}\left[\frac{\mathrm d \mu^\varepsilon}{\mathrm d \lambda} \ln \frac{\mathrm d \mu^\varepsilon}{\mathrm d \lambda}(a) - \frac{\mathrm d \mu}{\mathrm d \lambda}\ln \frac{\mathrm d \mu}{\mathrm d \lambda}(a) - \ln \frac{\mathrm d \nu}{\mathrm d \lambda}(a)\big( \frac{\mathrm d \mu^\varepsilon}{\mathrm d \lambda} - \frac{\mathrm d \mu}{\mathrm d \lambda}(a)\big) \right]\,\lambda(da)\\ & = \int \left( \limsup_{\varepsilon\searrow 0} \frac{1}{\varepsilon}\left [\frac{\mathrm d \mu^\varepsilon}{\mathrm d \lambda} \ln \frac{\mathrm d \mu^\varepsilon}{\mathrm d \lambda}(a) - \frac{\mathrm d \mu}{\mathrm d \lambda}\ln \frac{\mathrm d \mu}{\mathrm d \lambda}(a)\right] -\ln \frac{\mathrm d \nu}{\mathrm d \lambda}(a) (\frac{\mathrm d \mu'}{\mathrm d \lambda}(a) - \frac{\mathrm d \mu}{\mathrm d \lambda}(a)) \right)\,\lambda(da)\,. \end{align}\tag{44}\] Using the derivative of \(x\mapsto x \log x\) at \(x>0\), \[\lim_{\varepsilon \to 0} \frac{1}{\varepsilon}\left( \frac{\mathrm d \mu^\varepsilon}{\mathrm d \lambda}(a) \ln \frac{\mathrm d \mu^\varepsilon}{\mathrm d \lambda}(a) - \frac{\mathrm d \mu}{\mathrm d\lambda}(a)\ln \frac{\mathrm d \mu}{\mathrm d\lambda}(a) \right) = (1 + \ln \frac{\mathrm d \mu}{\mathrm d\lambda}(a))(\frac{\mathrm d \mu'}{\mathrm d\lambda}(a)-\frac{\mathrm d \mu}{\mathrm d\lambda}(a))\,,\] which together with 44 implies that \[\begin{align} & \limsup_{\varepsilon\searrow 0} \frac{1}{\varepsilon} \left(\operatorname{KL}(\mu^\varepsilon|\nu) - \operatorname{KL}(\mu|\nu)\right) \leq \int \left( \ln \frac{\mathrm d \mu}{\mathrm d\lambda}(a)(\frac{\mathrm d \mu'}{\mathrm d\lambda}(a)-\frac{\mathrm d \mu}{\mathrm d\lambda}(a))-\ln \frac{\mathrm d \nu}{\mathrm d \lambda}(a) (\frac{\mathrm d \mu'}{\mathrm d \lambda}(a) - \frac{\mathrm d \mu}{\mathrm d \lambda}(a)) \right)\,\lambda(da)\\ & = \int \left[\ln \frac{\mathrm d \mu}{\mathrm d\lambda}(a) - \ln \frac{\mathrm d \nu}{\mathrm d \lambda}(a) \right] (\mu'-\mu)(da)\,. \end{align}\] This proves 40 . Combining 39 and 40 yields 38 and finishes the proof. ◻

Proof of Proposition 10. The flat derivative of \(h\) on \(\mathfrak C\) has been shown in Lemma 8. The associated Bregman divergence is given by \[\begin{align} D_h(m'|m) & = h(m')-h(m)-\int \frac{\delta h}{\delta m}(m,a)(m'-m)(da)\\ & = \int \ln \frac{\mathrm d m'}{\mathrm d \varrho }(a)m'(da) - \int \ln \frac{\mathrm d m}{\mathrm d \varrho }(a)m(da) - \int\ln \frac{\mathrm d m}{\mathrm d \varrho }(a)(m'-m)(da) = \int \ln \frac{\mathrm d m'}{\mathrm d m }(a)m'(da)\,. \end{align}\]

We now assume that \(\pi^0\in \mathcal{A}_{\mathfrak C}\) satisfies \(\mathbb{E}\int_0^T \big\| \ln \frac{\mathrm d \pi^0_t}{\mathrm d \varrho} \big\|_{L^\infty(A)} dt<\infty\) and verify Assumption 3. By [65], \(\pi^1\) in [eq32update32the32control32bregman32msa] is given by \[\pi_t^{1}(da)= \frac{\exp{\left(-\frac{1}{\lambda} \frac{\delta H^{\tau}_t}{\delta m}(\Theta_t(\pi^0), \pi^0_t, a) \right)}}{\int \exp{\left(-\frac{1}{\lambda} \frac{\delta H^{\tau}_t}{\delta m}(\Theta_t(\pi^0), \pi^0_t, a') \right)\pi^0_t(da')}}\pi^0_t(da)\,,\] which is progressively measurable. For each \(t\in [0,T]\), \[\begin{align} \ln \frac{\mathrm d \pi^1_t}{\mathrm d \varrho}(a) & =-\frac{1}{\lambda} \frac{\delta H^{\tau}_t}{\delta m}(\Theta_t(\pi^0), \pi^0_t, a) -\ln\left(\int \exp{\left(-\frac{1}{\lambda} \frac{\delta H^{\tau}_t}{\delta m}(\Theta_t(\pi^0), \pi^0_t, a') \right)\pi^0_t(da')}\right) +\ln \frac{\mathrm d \pi^0_t }{\mathrm d \varrho }(a)\,. \end{align}\] Taking the sup-norm over \(a\) on both sides and using the definition of \(H^\tau\) in 4 and the boundedness of \(\frac{\delta b}{\delta m}, \frac{\delta \sigma}{\delta m}\) and \(\frac{\delta f}{\delta m}\) show that there exists \(C\ge 0\) such that \[\begin{align} \left\| \ln \frac{\mathrm d \pi^1_t}{\mathrm d \varrho}\right\|_{L^\infty(A)} & \le C\bigg(1+|Y_t(\pi^0)|+|Z_t(\pi^0)| + \left\| \ln \frac{\mathrm d \pi^0_t }{\mathrm d\varrho }\right\|_{L^\infty(A)} \bigg) \,, \end{align}\] which together with the square integrability of \(Y(\pi^0)\) and \(Z(\pi^0)\) implies that \(\mathbb{E}\int_0^T \big\| \ln \frac{\mathrm d \pi^1_t}{\mathrm d \varrho} \big\|_{L^\infty(A)} dt<\infty\). For fixed \(\nu\in \mathfrak C\), the fact that \(\mathbb{E}\int_0^T D_h(\pi^1_t|\nu) \, d t < \infty\) follows from \[\begin{align} \mathbb{E}\int_0^T h(\pi^1_t)dt &= \mathbb{E}\int_0^T \int\left| \ln \frac{\mathrm d \pi^1_t}{\mathrm d \varrho}(a) \right| \pi^1(da) dt \le \mathbb{E}\int_0^T \left\| \ln \frac{\mathrm d \pi^1_t}{\mathrm d \varrho} \right\|_{L^\infty(A)} dt<\infty\,, \end{align}\] and \[\begin{align} \left|\mathbb{E}\int_0^T \int \frac{\delta h}{\delta m }(\nu,a)\pi^1_t (da) dt \right| &= \left|\mathbb{E}\int_0^T \int \ln \frac{\delta \nu}{\delta \varrho }(a)\pi^1_t (da) dt \right| \le E\int_0^T \left\| \ln \frac{\mathrm d \nu}{\mathrm d \varrho} \right\|_{L^\infty(A)}dt <\infty\, . \end{align}\] Finally, \[\begin{align} \mathbb{E}\int_0^T D_h( \pi^1_t|\pi^0_t) dt & =\mathbb{E}\int_0^T \left( \int \left(\ln \frac{\mathrm d \pi^1_t}{\mathrm d \varrho }(a)-\ln \frac{\mathrm d \pi^0_t}{\mathrm d \varrho }\right) \pi^1_t(da) \right)dt \\ & \le \mathbb{E}\int_0^T \left(\left\| \ln \frac{\mathrm d \pi^1_t}{\mathrm d \varrho }\right\|_{L^\infty(A)}+\left\|\ln \frac{\mathrm d \pi^0_t}{\mathrm d \varrho }\right\|_{L^\infty(A)} \right) dt <\infty\,. \end{align}\] This verifies Assumption 3 for \(n=1\) and also proves that \(\mathbb{E}\int_0^T \big\| \ln \frac{\mathrm d \pi^1_t}{\mathrm d \varrho} \big\|_{L^\infty(A)} dt<\infty\). An inductive argument shows that Assumption 3 is satisfied for all \(n\in \mathbb{N}\cup\{0\}\). ◻

Proof of Proposition 13. Let \(m,m'\in \mathfrak C\) and for each \(\varepsilon\in (0,1)\), let \(m^\varepsilon= m+\varepsilon(m'-m)\). Note that for each \(\varepsilon\in (0,1)\), \[\begin{align} h(m^\varepsilon)&= \int \frac{1}{2} \left(\frac{\mathrm d m }{\mathrm d\varrho }(a) -1+\varepsilon\frac{\mathrm d (m'-m) }{\mathrm d\varrho }(a) \right) ^2 \varrho (da) \\ &= h(m) + \int \left(\frac{\mathrm d m }{\mathrm d\varrho }(a) -1\right) \varepsilon\frac{\mathrm d (m'-m) }{\mathrm d\varrho }(a) \varrho (da) +\varepsilon^2 \int \frac{1}{2} \left( \frac{\mathrm d (m'-m) }{\mathrm d\varrho }(a) \right) ^2 \varrho (da)\, \end{align}\] which along with \(m,m'\in \mathcal{P}(A)\) implies that for all \(c\in \mathbb{R}\), \[\begin{align} \lim_{\varepsilon\to 0}\frac{h(m^\varepsilon)-h(m)}{\varepsilon}&= \int \left(\frac{\mathrm d m }{\mathrm d\varrho }(a) -c \right) (m'-m) (da)\,. \end{align}\] As \(h(m) = \int \frac{1}{2} \left(\frac{\mathrm d m }{\mathrm d\varrho }(a) \right) ^2 \varrho (da)-1\), the divergence \(D_h(m'|m)\) is given by \[\begin{align} D_h(m'|m) &= \int \frac{1}{2} \left[ \left(\frac{\mathrm d m' }{\mathrm d\varrho }(a) \right) ^2 -\left(\frac{\mathrm d m }{\mathrm d\varrho }(a) \right) ^2 -2 \frac{\mathrm d m }{\mathrm d\varrho }(a) \frac{\mathrm d (m'-m) }{\mathrm d\varrho }(a) \right] \varrho (da) \\ & = \int \frac{1}{2} \left(\frac{\mathrm d m' }{\mathrm d\varrho }(a) -\frac{\mathrm d m }{\mathrm d\varrho }(a) \right) ^2 \varrho (da) \,. \end{align}\]

We now assume that \(\pi^0\in \mathcal{A}_{\mathfrak C}\) satisfies \(\mathbb{E}\int_0^T\big \| \frac{\mathrm d \pi^0_t}{\mathrm d \varrho} \big\|^2_{L^2_\varrho(A)} dt<\infty\) and verify Assumption 3. Observe that by the integrability assumptions of \(\frac{\delta b}{\delta m}, \frac{\delta \sigma}{\delta m}\), \(\frac{\delta f}{\delta m}\) and \(\frac{\mathrm d \pi^0}{\mathrm d \varrho}\), for a.e. \((\omega,t)\in \Omega\times [0,T]\), \(\frac{\delta H^{\tau}_t}{\delta m}(\Theta_t(\pi^0), \pi^{0}_t, \cdot)\in L^2_\varrho(A)\), and \(\pi^1\) in [eq32update32the32control32bregman32msa] satisfies \[\pi_t^{1}= \mathop{\mathrm{arg\,min}}_{m\in \mathfrak C}\int \left( \frac{\delta H^{\tau}_t}{\delta m}(\Theta_t(\pi^0), \pi^{0}_t, a)\frac{\mathrm d m}{\mathrm d\varrho}( a) + \frac{\lambda}{2} \left(\frac{\mathrm d m}{\mathrm d\varrho }(a) -\frac{\mathrm d \pi^0_t }{\mathrm d\varrho }(a) \right) ^2 \right) \varrho (da)\,.\] The first order condition (see e.g., [66]) shows that for a.e. \((\omega,t)\in [0,T]\times \Omega\), \[\label{eq:first-order95L2} \left\langle \frac{\delta H^{\tau}_t}{\delta m}(\Theta_t(\pi^0), \pi^{0}_t, \cdot) + \lambda \left(\frac{\mathrm d \pi^1_t}{\mathrm d\varrho } -\frac{\mathrm d \pi^0_t }{\mathrm d\varrho } \right), \phi- \frac{\mathrm d \pi^1_t}{\mathrm d\varrho } \right\rangle_{L^2_\varrho}\ge 0, \quad \forall \phi\in \mathcal{C}\,,\tag{45}\] where \(\left\langle \cdot, \cdot \right\rangle_{L^2_\varrho}\) is the inner product on \(L^2_\varrho(A)\), and \(\mathcal{C}\) is the nonempty closed convex set defined by \[\mathcal{C}=\left\{\phi\in L^2_\varrho(A)\bigg\vert \textrm{\phi\ge 0 \varrho-a.s.~on A and \int \phi(a)\varrho(da)=1}\right\}\,.\] Define the projection map \(\Pi_{\mathcal{C}}: L^2_\varrho(A)\mapsto \mathcal{C}\) such that \(\Pi_{\mathcal{C}}(\varphi)=\mathop{\mathrm{arg\,min}}_{\phi\in \mathcal{C}}\|\phi-\varphi\|_{L^2_\varrho(A)}\) for all \(\varphi \in L^2_\varrho(A)\), which satisfies \[\left\langle \Pi(\varphi)-\varphi, \phi- \Pi(\varphi) \right\rangle_{L^2_\varrho}\ge 0, \quad \forall \phi\in \mathcal{C}\,.\] Then \[\frac{\mathrm d \pi^1}{\mathrm d \varrho} = \Pi_{\mathcal{C}}\left( \frac{\mathrm d \pi^0_t }{\mathrm d\varrho } -\frac{1}{\lambda }\frac{\delta H^{\tau}_t}{\delta m}(\Theta_t(\pi^0), \pi^{0}_t, \cdot) \right)\,.\] As \(\|\Pi_{\mathcal{C}}(\varphi_1)-\Pi_{\mathcal{C}}(\varphi_2)\|_{L^2_\varrho(A)} \le \|\varphi_1-\varphi_2\|_{L^2_\varrho(A)}\) for all \(\varphi_1,\varphi_2\in L^2_\varrho(A)\) (see e.g., [67]), \(\pi^1\) is progressively measurable. Moreover, since \(\frac{\mathrm d \pi^0_t}{\mathrm d \varrho} = \Pi_{\mathcal{C}}\left(\frac{\mathrm d \pi^0_t}{\mathrm d \varrho}\right)\), there exists \(C\ge 0\) such that for a.e. \((\omega,t)\in \Omega\times [0,T]\), \[\begin{align} \left \| \frac{\mathrm d \pi^1_t}{\mathrm d \varrho} \right\|_{L^2_\varrho(A)} & \le \left \| \frac{\mathrm d \pi^0_t }{\mathrm d\varrho }\right\|_{L^2_\varrho(A)} + \frac{1}{\lambda } \left \| \frac{\delta H^{\tau}_t}{\delta m}(\Theta_t(\pi^0), \pi^{0}_t, \cdot)\right\|_{L^2_\varrho(A)} \le C\left(1+ |Y_t(\pi^0)|+|Z_t(\pi^0)|+ \left \| \frac{\mathrm d \pi^0_t }{\mathrm d\varrho }\right\|_{L^2_\varrho(A)} \right)\,, \end{align}\] where the last inequality used the definition of \(H^\tau\) in 4 , the flat derivative of \(h\) at \(\pi^0_t\) and the integrability assumptions of \(\frac{\delta b}{\delta m}, \frac{\delta \sigma}{\delta m}\) and \(\frac{\delta f}{\delta m}\). Using the square integrability of \(Y(\pi^0)\) and \(Z(\pi^0)\) shows that \(\mathbb{E}\int_0^T\big \| \frac{\mathrm d \pi^1_t}{\mathrm d \varrho} \big\|^2_{L^2_\varrho(A)} dt<\infty\). The fact that \(\mathbb{E}\int_0^T D_h(\pi^1_t|\pi^0) \, d t < \infty\) and for fixed \(\nu\in \mathfrak C\), \(\mathbb{E}\int_0^T D_h(\pi^1_t|\nu) \, d t < \infty\) follows directly from the definition of the Bregman divergence. This verifies Assumption 3 for \(n=1\) and proves \(\mathbb{E}\int_0^T\big \| \frac{\mathrm d \pi^1_t}{\mathrm d \varrho} \big\|^2_{L^2_\varrho(A)} dt<\infty\). An inductive argument shows that Assumption 3 is satisfied for all \(n\in \mathbb{N}\cup\{0\}\). ◻

Proof of Proposition 15. Under the assumptions of \(c\) and \(A\), by [16], \(m\to h(m)\) is continuous (with respect to the weak topology) and admits the flat derivative given in the statement.

We now verify Assumption 3. To this end, consider the map \(F: [0,T]\times \mathbb{R}^d\times \mathbb{R}^d\times \mathbb{R}^{d\times d'} \times \mathfrak C \times \mathfrak C \to \mathbb{R}\), \[\begin{align} \begin{aligned} F(t,x,y,z,m, m') &=\int\frac{\delta H^{\tau}_t}{\delta m}(x,y,z, m, a) m'(da) + \lambda \left(h(m')-\int \phi [m](a)m'(da) \right) \\ &=\int \left( \frac{\delta H^{0}_t}{\delta m}(x,y,z, m, a) + (\tau-\lambda) \phi [m](a) \right) m'(da) + \lambda h(m') \,, \end{aligned} \end{align}\] where the last identity used the definition of \(H^\tau\) in 4 . By [16], the map \(\mathfrak C\ni \mu\mapsto \phi[\mu]\in C(A)\) is continuous, where \(\mathfrak C=\mathcal{P}(A)\) is equipped with the topology of weak convergence of measures and \(C(A)\) is equipped with the topology of uniform convergence of functions. By the measurability of \(\frac{\delta H^0}{\delta m}\) and the continuity of \(\mathfrak C\ni \mu\mapsto \phi[\mu]\in C(A)\), for all \(m'\in \mathfrak C\), \((t,x,y,z,m)\mapsto F(t,x,y,z,m, m')\) is measurable. Moreover, by the continuity of \(a\mapsto \frac{\delta H^0_t}{\delta m}(x,y,z,m,a)\) and the continuity of \(m\mapsto h(m)\), for all \((t,x,y,z,m ) \in [0,T]\times \mathbb{R}^d\times \mathbb{R}^d\times \mathbb{R}^{d\times d'} \times \mathfrak C\), \(m'\mapsto F(t,x,y,z,m, m')\) is continuous. This proves that \(F\) is a Carathéodory function. Since \(A\) is compact and separable, \(\mathfrak C = \mathcal{P}(X)\) is a compact and separable metrisable space (see e.g., [68]). Hence, by [68], there exists a measurable function \(\mathfrak f:[0,T]\times \mathbb{R}^d\times \mathbb{R}^d\times \mathbb{R}^{d\times d'} \times \mathfrak C \to \mathfrak C\) such that for all \(t\in [0,T]\), \((x,y,z)\in \mathbb{R}^d\times \mathbb{R}^d\times \mathbb{R}^{d\times d'}\) and \(m\in \mathfrak C\), \[\begin{align} \label{eq:maximum95selector} \mathfrak f(t,x,y,z,m)\in \mathop{\mathrm{arg\,min}}_{m' \in \mathfrak C} F(t,x,y,z,m,m')\,. \end{align}\tag{46}\]

Given \(\pi^0\in \mathcal{A}_{\mathfrak C}\), it is easy to see that \(\pi^1_t \mathrel{\vcenter{:}}= \mathfrak f(t,X_t(\pi^0),Y_t(\pi^0), Z_t(\pi^0),\pi^0_t)\) for a.e. \((\omega,t)\in \Omega\times [0,T]\) satisfies [eq32update32the32control32bregman32msa]. Note that the measurability of \(\mathfrak f\) implies the measurability of \(\pi^1\). Moreover, by the compactness of \(\mathfrak C\) and the continuity of \(h\), for a given \(\nu \in \mathfrak C\), \[\begin{align} \mathbb{E}\int_0^T D_h(\pi_t|\nu) \, d t \le 2 T \left( \sup_{m\in \mathfrak C }h(m)+\|\phi [\nu]\|_{L^\infty} \right)< \infty\,, \end{align}\] and similarly \[\mathbb{E}\int_0^T D_h(\pi^1_t|\pi^0_t) \, d t \le 2 T \sup_{m\in \mathfrak C }\left(h(m)+\|\phi [m]\|_{L^\infty} \right) <\infty,\] where the last inequality used the compactness of \(\mathfrak C\) and the continuity of \(\mathfrak C\ni \mu\mapsto \phi[\mu]\in C(A)\). An inductive argument allows for verifying Assumption 3 for all \(n\in \mathbb{N}\cup\{0\}\). ◻

Acknowledgements↩︎

Supported by the Alan Turing Institute under EPSRC grant no. EP/N510129/1.

The authors would like to thank the anonymous reviewers for their insightful comments. These have helped to greatly improve the revised version of the article.

4 Properties for Bregman divergence over measures↩︎

This section collects basic properties of Bregman divergence used in this paper. Throughout this section, we fix a convex and measurable set \(\mathfrak C\subset \mathcal{P}(A)\) and a convex function \(h:\mathfrak C\to \mathbb{R}\) that has a flat derivative \(\frac{\delta h}{\delta m}: \mathfrak C \times A\to \mathbb{R}\). Define \(D_h(\cdot|\cdot):\mathfrak C\times \mathfrak C\to \mathbb{R}\) as in 9 .

The following properties of \(D_h\) follow from the definition.

Lemma 9. For all \(m,m' \in \mathfrak C\),

  1. \(D_h(m'|m) \geq 0\).

  2. the function \(\mathfrak C \ni m\mapsto D_h(m|m')\in \mathbb{R}\) is convex and has a flat derivative \[\label{item:D95convex} \frac{\delta D_h(\cdot|m')}{\delta m} =\frac{\delta h}{\delta m}(m,a)- \frac{\delta h}{\delta m}(m',a)\,.\tag{47}\]

  3. for all convex functions \(g:\mathfrak C\to \mathbb{R}\) with flat derivatives and all \(\alpha,\beta \in \mathbb{R}\), \[\label{item:32additivity} D_{\alpha h+\beta g}(m'|m) =\alpha D_{h }(m'|m)+\beta D_{g }(m'|m)\,.\tag{48}\]

  4. for all \(\nu \in \mathfrak C\), \[\label{item:breg32of32breg} D_{D_h(\cdot|\nu)}(m'|m ) = D_h(m'| m)\,.\tag{49}\]

Proof. For Item [item:D95positive], note that for all \(0<s<t\le 1\), \(m+s (m'-m) =\frac{s}{t}(m+t (m'-m))+ \left( 1-\frac{s}{t}\right) m\), and hence by the convexity of \(h\), \[h(m+s (m'-m)) \le \frac{s}{t}h(m+t (m'-m) ) + \left( 1-\frac{s}{t}\right) h(m) \,,\] which implies that \([0,1]\ni \varepsilon\mapsto \frac{h(m+\varepsilon (m'-m))-h(m)}{\varepsilon}\in \mathbb{R}\) is increasing. Hence by the definition of \(\frac{\delta h}{\delta m}\), \[{h(m' )-h(m)}\ge \lim_{\varepsilon \searrow 0}\frac{h(m+\varepsilon (m'-m))-h(m)}{\varepsilon}= \int \frac{\delta h}{\delta m}(m'-m)(da)\,,\] which proves \(D_h(m'|m)\ge 0\).

Items 47 and 48 follow directly from the convexity and differentiability of \(h\) and the definition of \(D_h(\cdot|\nu)\).

For Item 49 , by the definitions of \(D_{D_h(\cdot|\nu)}\) and \(\frac{\delta D_h(\cdot|\nu)}{\delta m}\) in Item 48 , \[\begin{align} D_{D_h(\cdot|\nu)}(m'|m ) &= D_{h}(m'|\nu) - D_h( m| \nu ) - \int \frac{\delta D_h(\cdot|\nu)}{\delta m}( m, a )(m'- m)(da) \\ &= D_{h}(m'|\nu) - D_h( m| \nu ) - \int \left(\frac{\delta h}{\delta m}(m,a)- \frac{\delta h}{\delta m}(\nu,a)\right)(m'- m)(da) \\ &= h(m') - h(\nu) - \int \frac{\delta h}{ \delta m}(\nu,a)(m'- \nu)(da) \\ &\quad - \left( h(m) - h(\nu) - \int \frac{\delta h}{ \delta m}(\nu,a)(m - \nu)(da) \right)\\ &\quad - \int \frac{\delta h}{\delta m}(m,a)(m'- m)(da) + \frac{\delta h}{\delta m}(\nu,a)(m'- m)(da) \\ &= h(m') - h(m) - \int \frac{\delta h}{\delta m}(m,a)(m'- m)(da) =D_h(m'|m)\,. \end{align}\] This completes the proof. ◻

We then prove a “three point lemma" for the Bregman divergence \(D_h(\cdot|m)\).

Lemma 10. Let \(\nu\in \mathfrak C\) and let \(G:\mathfrak C \rightarrow \mathbb{R}\) be convex and have a flat derivative \(\frac{\delta G}{\delta m}\). Suppose that there exists \(\overline{m}\in \mathfrak C\) such that \(\overline{m} \in \mathop{\mathrm{arg\,min}}_{m\in \mathfrak C} \left( G(m) + D_h(m|\nu) \right)\). Then for all \(m'\in \mathfrak C\), \[G(m') + D_h(m'|\nu) \geq G(\overline{m} ) + D_h(m'|\overline{m} ) + D_h(\overline{m}| \nu)\,.\]

Proof. Fix \(m'\in \mathfrak C\). Define \(\mathcal{G}: \mathfrak C\to \mathbb{R}\) by \(\mathcal{G}(m) = G(m) + D_h(m|\nu)\) for all \(m\in \mathfrak C\). The convexity and differentability of \(G\) and \(h\) imply that \(\mathcal{G}\) has a flat derivative. As \(\overline{m}\) minimises \(\mathcal{G}\) over \(\mathfrak C\), \[\int \frac{\delta \mathcal{G}}{\delta m}(\overline{m},a)(m'- \overline{m})(da) =\lim_{\varepsilon \searrow 0}\frac{ \mathcal{G} (\overline{m}+\varepsilon (m'-\overline{m} ))- \mathcal{G}(\overline{m})}{\varepsilon}\ge 0\,,\] which implies that \[\begin{align} D_{\mathcal{G}}(m'| \overline{m}) = \mathcal{G}(m') -\mathcal{G} (\overline{m} ) - \int \frac{\delta \mathcal{G}}{ \delta m} (\overline{m} ,a)(m' - \overline{m} )(da) \le \mathcal{G}(m') - \mathcal{G} (\overline{m} )\,. \end{align}\] Adding \(\mathcal{G} (\overline{m} )\) to both sides of the inequality and using the definition of \(\mathcal{G}\) yield that \[\begin{align} G(m')+D_h(m'|\nu) &\ge D_{\mathcal{G}}(m'| \overline{m}) +G(\overline{m}) +D_h( \overline{m} |\nu) \\ &= D_{ G}(m'| \overline{m}) +D_{ D_h(\cdot|\nu)}(m'| \overline{m}) +G(\overline{m}) +D_h( \overline{m} |\nu) \\ &\ge D_{ h}(m'| \overline{m}) +G(\overline{m}) +D_h( \overline{m} |\nu)\,, \end{align}\] where the first identity used Lemma 9 Item 48 and the last inequality used Lemma 9 Items [item:D95positive] and 49 . This completes the proof. ◻

5 Proof of Lemma 6↩︎

Throughout this section, fix \(\pi,\pi' \in \mathcal{A}_{\mathfrak C}\) such that \(\mathbb{E} \int_0^T D_h (\pi'_t|\pi_t)\,dt < \infty\). For each \(\varepsilon\in [0,1]\), let \(\pi^\varepsilon = \pi + \varepsilon(\pi'-\pi)\), which lies in \(\mathcal{A}_{\mathfrak C}\) due to the convexity of \(\mathcal{A}_{\mathfrak C}\). For each \(\varepsilon\in [0,1]\), let \(X(\pi^\varepsilon)\) be a solution to the state dynamics 1 controlled by \(\pi^\varepsilon\). Consider the following SDE: for all \(t\in [0,T]\), \[\label{eq32V32proc32linear} \begin{align} dV_t & = \left[(D_x b_t)(X_t,\pi_t)V_t + \int \frac{\delta b_t}{\delta m}(X_t,\pi_t,a)(\pi'_t - \pi_t)(da) \right] \,dt \\ &\quad + \left[(D_x \sigma_t)(X_t,\pi_t)V_t + \int \frac{\delta \sigma_t}{\delta m}(X_t,\pi_t,a)(\pi'_t - \pi_t)(da) \right]\,dW_t\,;\quad V_0 = 0\,. \end{align}\tag{50}\] Under Assumptions 1 and 5, \(D_x b\) and \(D_x \sigma\) are uniformly bounded, and \[\begin{align} &\left|\int \frac{\delta b_t}{\delta m}(X_t,\pi_t,a)(\pi'_t - \pi_t)(da) \right|^2 +\left|\int \frac{\delta \sigma_t}{\delta m}(X_t,\pi_t,a)(\pi'_t - \pi_t)(da) \right|^2 \\ &\le 2K D_h(\pi'_t|\pi_t)\,. \end{align}\] This, along with \(\mathbb{E} \int_0^T D_h (\pi'_t|\pi_t)\,dt < \infty\) and [57] implies that \(\mathbb{E}[\sup_{t\in [0,T]}|V_t|^2]<\infty\).

The following lemma proves that \(V\) is the derivative of \(\varepsilon\mapsto X^\varepsilon\).

Lemma 11. Suppose that Assumptions 1, 4 and 5 hold. Then \[\lim_{\varepsilon \searrow 0}\mathbb{E} \left[\sup_{t\in[0,T]} \left|\frac{X_t(\pi^\varepsilon) - X_t(\pi)}{\varepsilon} - V_t \right|^2\right] = 0\,.\]

Proof. To simplify the notation, let \(X=X(\pi)\) and for each \(\varepsilon\in (0,1)\), let \(X^\varepsilon=X(\pi^\varepsilon)\) and \(V^\varepsilon=\frac{X^\varepsilon- X}{\varepsilon} - V\). We also assume without loss of generality that all processes are one-dimensional. For each \(\varepsilon\in (0,1)\), using 1 and 50 , we have \(V^\varepsilon=0\) and for all \(t\in [0,T]\), \[\label{eq32X95eps-X-V} \begin{align} dV^\varepsilon_t & = \bigg(\frac{b_t(X^\varepsilon_t,\pi^\varepsilon_t)-b_t(X_t,\pi_t)}{\varepsilon} -(D_x b_t)(X_t,\pi_t)V_t - \int \frac{\delta b_t}{\delta m}(X_t,\pi_t,a)(\pi'_t - \pi_t)(da) \bigg) \,dt \\ & + \bigg( \frac{\sigma_t(X^\varepsilon_t,\pi^\varepsilon_t)-\sigma_t(X_t,\pi_t)}{\varepsilon} - (D_x \sigma_t)(X_t,\pi_t)V_t + \int \frac{\delta \sigma_t}{\delta m}(X_t,\pi_t,a)(\pi'_t - \pi_t)(da) \bigg)\,dW_t\,. \end{align}\tag{51}\] Note that using \(X^\varepsilon=X+\varepsilon(V^\varepsilon+V)\) \[\begin{align} & b_t(X^\varepsilon_t, \pi^\varepsilon_t) - b_t(X_t,\pi_t) = b_t(X^\varepsilon_t, \pi^\varepsilon_t) - b_t(X^\varepsilon_t, \pi_t) + b_t(X^\varepsilon_t,\pi_t) - b_t(X_t,\pi_t) \\ & = \varepsilon\int_0^1 \int \frac{\delta b_t}{\delta m}(X^\varepsilon_t, \pi_t +\lambda (\pi^\varepsilon_t - \pi_t),a)(\pi'_t-\pi_t)(da)\,d\lambda + \varepsilon \int_0^1 (D_x b_t)(X_t + \lambda (X^\varepsilon_t - X_t), \pi_t) (V^\varepsilon_t + V_t)\,d\lambda \,. \end{align}\] Hence, by rearranging the terms, \[\begin{align} & \frac{b_t(X^\varepsilon_t,\pi^\varepsilon_t)-b_t(X_t,\pi_t)}{\varepsilon} -(D_x b_t)(X_t,\pi_t)V_t - \int \frac{\delta b_t}{\delta m}(X_t,\pi_t,a)(\pi'_t - \pi_t)(da) \\ & = \left( \int_0^1 (D_x b_t)(X_t + \lambda (X^\varepsilon_t - X_t), \pi_t)\,d\lambda \right) V^\varepsilon_t + I^{(1)}_t + I^{(2)}_t\,, \end{align}\] where \[\begin{align} I^{(1,\varepsilon)}_t&= \int_0^1 \left[(D_x b_t)(X_t + \lambda (X^\varepsilon_t - X_t), \pi_t) - (D_x b_t)(X_t,\pi_t) \right] V_t\,d\lambda\,, \\ I^{(2,\varepsilon)}_t &= \int_0^1 \int \left[ \frac{\delta b_t}{\delta m}(X^\varepsilon_t, \pi_t +\lambda (\pi^\varepsilon_t - \pi_t),a) - \frac{\delta b_t}{\delta m}(X^\varepsilon_t, \pi_t,a) \right](\pi'_t-\pi_t)(da)\,d\lambda \,. \end{align}\] Similarly, we have \[\begin{align} & \frac{\sigma_t(X^\varepsilon_t,\pi^\varepsilon_t)-\sigma_t(X_t,\pi_t)}{\varepsilon} -(D_x \sigma_t)(X_t,\pi_t)V_t - \int \frac{\delta \sigma_t}{\delta m}(X_t,\pi_t,a)(\pi'_t - \pi_t)(da) \\ & = \left( \int_0^1 (D_x \sigma_t)(X_t + \lambda (X^\varepsilon_t - X_t), \pi_t)\,d\lambda \right) V^\varepsilon_t + J^{(1)}_t + J^{(2)}_t\,, \end{align}\] where \[\begin{align} J^{(1,\varepsilon)}_t&= \int_0^1 \left[(D_x \sigma_t)(X_t + \lambda (X^\varepsilon_t - X_t), \pi_t) - (D_x \sigma_t)(X_t,\pi_t) \right] V_t\,d\lambda\,, \\ J^{(2,\varepsilon)}_t &= \int_0^1 \int \left[ \frac{\delta \sigma_t}{\delta m}(X^\varepsilon_t, \pi_t +\lambda (\pi^\varepsilon_t - \pi_t),a) - \frac{\delta \sigma_t}{\delta m}(X^\varepsilon_t, \pi_t,a) \right](\pi'_t-\pi_t)(da)\,d\lambda \,. \end{align}\] Then by [57] and the uniform boundedness of \(D_x b\) and \(D_x \sigma\), \[\begin{align} \label{eq:V95eps95estimate} \begin{aligned} \mathbb{E} \left[\sup_{t\in[0,T]} \left| V^\varepsilon_t \right|^2\right] & \le C\left[\left(\int_0^T (|I^{(1,\varepsilon)}_t|+|I^{(2,\varepsilon)}_t|)dt\right)^2 + \int_0^T ( |J^{(1,\varepsilon)}_t|^2+|J^{(2,\varepsilon)}_t|^2)dt \right] \\ & \le C\left[\int_0^T \left( |I^{(1,\varepsilon)}_t|^2+|I^{(2,\varepsilon)}_t| ^2 +|J^{(1,\varepsilon)}_t|^2+|J^{(2,\varepsilon)}_t|^2\right)dt \right]\,. \end{aligned} \end{align}\tag{52}\] Above and hereafter, we denote by \(C\) a generic constant independent of \(\varepsilon\).

It remains to prove the terms on the right-hand side of 52 converges to zero as \(\varepsilon\to 0\). Observe that by Lemma 2 and the convexity of \(m\mapsto D_h(m|\pi_s)\), \[\mathbb{E}\left[ \sup_{0\le t\le T}|X^\varepsilon_t -X_t |^2\right]\le C\mathbb{E}\int_0^T\,D_h(\pi^\varepsilon_s|\pi_s)\,ds \le \varepsilon C\mathbb{E}\int_0^T\,D_h(\pi'_s|\pi_s)\,ds \,,\] which along with \(\mathbb{E}\int_0^T\,D_h(\pi'_s|\pi_s)\,ds<\infty\) implies that \(\lim_{\varepsilon\searrow 0}\mathbb{E}\left[ \sup_{t\in [0,T] }|X^\varepsilon_t -X_t |^2\right]=0\). Thus by subtracting a subsequence if necessary, we can assume without loss of generality that \(\lim_{\varepsilon\searrow 0} X^\varepsilon=X\) for a.e. \((\omega ,t)\in \Omega\times [0,T]\). This along with the continuity of \(D_x b\) in \(x\) implies that for a.e. \((\omega ,t)\in \Omega\times [0,T]\), \[\lim_{\varepsilon\searrow 0}|(D_x b_t)(X_t + \lambda (X^\varepsilon_t - X_t), \pi_t) - (D_x b_t)(X_t,\pi_t)|=0\,.\] As \(D_x b\) are bounded by \(K\), \(\sup_{\varepsilon\in [0,1]}|I^{(1,\varepsilon)}_t|^2\le 2K |V_t|^2\), and hence by the square integrability of \(V\), \(\mathbb{E}\int_0^T \sup_{\varepsilon\in [0,1]}|I^{(1,\varepsilon)}_t|^2 dt <\infty\). By Lebesgue’s dominated convergence theorem, \(\lim_{\varepsilon\searrow 0} \int_0^T |I^{(1,\varepsilon)}_t|^2 dt =0\). Similar argument shows that \(\lim_{\varepsilon\searrow 0} \int_0^T |J^{(1,\varepsilon)}_t|^2 dt =0\). For the term \(I^{(2,\varepsilon)}_t\), by Assumption 5 and Fubini’s theorem, \[\begin{align} & |I^{(2,\varepsilon)}_t|^2 = \left|\int_0^1 \int \left( \int_0^1 \int \frac{\delta^2 b_t}{\delta^2 m}(X^\varepsilon_t, \pi_t +\lambda \varepsilon' (\pi^\varepsilon_t - \pi_t),a,a' ) \lambda (\pi^\varepsilon_t-\pi_t)(da') d \varepsilon' \right)(\pi'_t-\pi_t)(da)\,d\lambda \right|^2 \\ &\le \frac{1}{2} \varepsilon K D_h (\pi'_t|\pi_t)\,, \end{align}\] and hence \(\lim_{\varepsilon\searrow 0} \int_0^T |I^{(2,\varepsilon)}_t|^2dt =0\). Finally the condition that \(\frac{\delta^2 \sigma}{\delta ^2 m}=0\) implies that \(J^{(2,\varepsilon)}_t=0\) for all \(\varepsilon\in [0,1]\). This along with 52 yields the desired conclusion. ◻

The next lemma expresses the directional derivative of \(J^0\) using the sensitivity process \(V\).

Lemma 12. Suppose Assumptions 1, 2, 4 and 5. Then \[\begin{align} &\lim_{\varepsilon\searrow 0}\frac{J^0(\pi^\varepsilon)-J^0(\pi)}{\varepsilon} = \mathbb{E} \bigg[\int_0^T \bigg(\int \frac{\delta f_t}{\delta m}(X_t, \pi_t,a)(\pi'_t-\pi_t)(da) + (D_x f)(X_t,\pi_t) V_t\bigg)\,dt + (D_x g)(X_T)V_T \bigg] \,. \end{align}\]

Proof. It suffices to prove \(I_\varepsilon \mathrel{\vcenter{:}}= I^{(1)}_\varepsilon + I^{(2)}_\varepsilon \to 0\) as \(\varepsilon\to 0\), where \[\begin{align} I^{(1)}_\varepsilon &\mathrel{\vcenter{:}}= \mathbb{E} \int_0^T \bigg|\frac{ f_t(X^\varepsilon_t,\pi^\varepsilon_t) - f_t(X_t,\pi_t)}{\varepsilon} - \int \frac{\delta f_t}{\delta m}(X_t,\pi_t,a)\,(\pi'_t-\pi_t)(da) - (D_x f_t)(X_t,\pi_t) V_t \bigg| \,dt \, , \end{align}\] and \[I^{(2)}_\varepsilon \mathrel{\vcenter{:}}= \mathbb{E} \Big| \frac{g(X^\varepsilon_T)-g(X_T)}{\varepsilon} - (D_x g)(X_T)V_T \Big|\,.\] To estimate \(I^{(1)}_\varepsilon\), first note that by Assumption 2, \[\begin{align} & \frac{f_t(X^\varepsilon_t,\pi^\varepsilon_t) - f_t(X_t,\pi_t)}{\varepsilon} = \frac{f_t(X^{\varepsilon}_t, \pi^\varepsilon_t) - f_t(X^{\varepsilon}_t, \pi_t) + f_t(X^{\varepsilon}_t, \pi_t) - f_t(X_t, \pi_t)}{\varepsilon} \\ & = \int_0^1 \int \frac{\delta f_t}{\delta m}(X^\varepsilon_t, \pi_t +\lambda (\pi^\varepsilon_t -\pi_t), a )(\pi'_t - \pi_t)(da) \,d\lambda + \int_0^1(D_x f_t)(X_t + \lambda (X^\varepsilon_t-X_t),\pi_t)(V^{\varepsilon}_t - V_t)\,d\lambda\,, \end{align}\] where \(V^\varepsilon=\frac{X^\varepsilon-X}{\varepsilon}-V\). Hence we can write \(I^{(1)}_\varepsilon \leq I^{(1,1)}_\varepsilon + I^{(1,2)}_\varepsilon + I^{(1,3)}_\varepsilon\), where \[\begin{align} I^{(1,1)}_\varepsilon & \mathrel{\vcenter{:}}= \mathbb{E} \int_0^T \bigg| \int_0^1 \int \bigg[\frac{\delta f_t}{\delta m}(X^\varepsilon_t, \pi_t +\lambda (\pi^\varepsilon_t -\pi_t), a ) -\frac{\delta f_t}{\delta m}(X^\varepsilon_t,\pi_t,a)\bigg] (\pi'_t - \pi_t)(da)d\lambda \bigg|dt\,, \\ I^{(1,2)}_\varepsilon & \mathrel{\vcenter{:}}= \mathbb{E} \int_0^T \bigg| \int_0^1 \int \bigg[\frac{\delta f_t}{\delta m}(X^\varepsilon_t,\pi_t,a)-\frac{\delta f_t}{\delta m}(X_t,\pi_t,a)\bigg] (\pi'_t - \pi_t)(da)d\lambda \bigg|dt\,, \\ I^{(1,3)}_\varepsilon &\mathrel{\vcenter{:}}= \mathbb{E} \int_0^T \bigg|\int_0^1(D_x f_t)(X_t + \lambda (X^\varepsilon_t-X_t),\pi_t)(V^{\varepsilon}_t - V_t) \,d\lambda - (D_x f_t)(X_t,\pi_t) V_t \bigg|\,dt \,. \end{align}\] Recall that Lemma 11 shows that \[\lim_{\varepsilon\searrow 0} \left(\mathbb{E}\left[ \sup_{t\in [0,T] }|X^\varepsilon_t -X_t |^2\right] +\mathbb{E}\left[ \sup_{t\in [0,T] }|V^\varepsilon_t |^2\right]\right)=0\,.\] Hence using Assumptions 2 and 5 and Lebesgue’s dominated convergence theorem, \(\lim_{\varepsilon\searrow 0} I^{(1)}=0\). The convergence of \(I^{(2)}\) can be established by similar arguments. This finishes the proof. ◻

Proof of Lemma 6. The proof follows exactly the same lines as those of  [21], by first applying Itô’s formula to \(Y_t^\top V_t\) and using Lemma 12 and the definition of \(H^0\) in 4 . The details are omitted. ◻

6 Proof of Lemma 2↩︎

Proof of Lemma 2. Fix \(\pi,\pi'\in \mathcal{A}_{\mathfrak C}\). Observe that for all \(t\in [0,T]\), \[\begin{align} &X_t(\pi)-X_t(\pi') =\int_0^t\big(b_s(X_s(\pi),\pi_s)-b(X_s(\pi'),\pi'_s)\big)\,ds+\int_0^t\big(\sigma_s(X_s(\pi),\pi_s)-\sigma_s(X_s(\pi'),\pi'_s)\big)\,dW_s\,. \end{align}\] Using the inequality \((a+b)^2\le 2a^2+2b^2\) and Hölder’s inequality, \[\begin{align} &|X_t(\pi)-X_t(\pi')|^2\\ &=\left|\int_0^t\left(b_s(X_s(\pi),\pi_s)-b(X_s(\pi'),\pi'_s)\right)\,ds+\int_0^t\left(\sigma_s(X_s(\pi),\pi_s)-\sigma_s(X_s(\pi'),\pi'_s)\right)\,dW_s \right|^2\\ &\le 2\left|\int_0^t\left(b_s(X_s(\pi),\pi_s)-b(X_s(\pi'),\pi'_s)\right)\,ds\right|^2+2\left|\int_0^t\left(\sigma_s(X_s(\pi),\pi_s)-\sigma_s(X_s(\pi'),\pi'_s)\right)\,dW_s\right|^2\\ &\le 2t\int_0^t\left|b_s(X_s(\pi),\pi_s)-b(X_s(\pi'),\pi'_s)\right|^2\,ds+2\left|\int_0^t\left(\sigma_s(X_s(\pi),\pi_s)-\sigma_s(X_s(\pi'),\pi'_s)\right)\,dW_s\right|^2\,. \end{align}\] Let \(t'\in [0,T]\). Taking supremum over \(t\in[0,t']\) yields \[\begin{align} \sup_{0\le t\le t'} |X_t(\pi)-X_t(\pi')|^2 &\le 2T\int_0^{t'}\left|b_s(X_s(\pi),\pi_s)-b(X_s(\pi'),\pi'_s)\right|^2\,ds\\ &\quad+2\sup_{0\le t\le t'}\left|\int_0^t\left(\sigma_s(X_s(\pi),\pi_s)-\sigma_s(X_s(\pi'),\pi'_s)\right)\,dW_s\right|^2\,. \end{align}\] By taking the expectation on both sides and using the Burkholder-Davis-Gundy inequality, \[\begin{align} \mathbb{E}\left[ \sup_{0\le t\le {t'}}|X_t(\pi)-X_t(\pi')|^2\right] & \le 2T\mathbb{E}\left[ \int_0^{t'}\left|b_s(X_s(\pi),\pi_s)-b(X_s(\pi'),\pi'_s)\right|^2\,ds\right]\\ & \quad+2C \mathbb{E}\left[ \int_0^{t'}\left|\sigma_s(X_s(\pi),\pi_s)-\sigma_s(X_s(\pi'),\pi'_s)\right|^2\,ds\right] \\ & \le 2TK \mathbb{E}\int_0^{t'}\left(|X_s(\pi)-X_s(\pi')|^2+D_h(\pi_s|\pi'_s)\right)\,ds\\ &\quad +2KC \mathbb{E}\int_0^{t'}\left(|X_s(\pi)-X_s(\pi')|^2+D_h(\pi_s|\pi'_s)\right)\,ds\,, \end{align}\] where the last inequality used the assumption of the lemma. We thus have, with some \(C>0\) depending on \(K\) and \(T\), for any \(t'\in[0,T]\), that \(y(t') \leq C\int_0^{t'} y(s)\,ds + b(t')\) where \(y(t):=\mathbb{E}\left[ \sup_{0\le s\le t}|X_{s}(\pi)-X_{s}(\pi')|^2\right]\) and \(b(t):= \int_0^t D_h(\pi_s|\pi'_s)\,ds\). Hence, by Grönwall’s inequality, we have \(y(T) \leq b(T)e^{CT}\). In other words \[\begin{align} \mathbb{E}\left[ \sup_{0\le t\le T}|X_t(\pi)-X_t(\pi')|^2\right]\le C\mathbb{E}\int_0^T\,D_h(\pi_s|\pi'_s)\,ds\,. \end{align}\] This finishes the proof. ◻

References↩︎

[1]
K. Nicole el, D. Hu̇u̇ Nguyen, and J.-P. Monique, Compactification methods in the control of degenerate diffusions: existence of an optimal control, Stochastics, 20 (1987), pp. 169–219.
[2]
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine, Reinforcement learning with deep energy-based policies, in International conference on machine learning, PMLR, 2017, pp. 1352–1361.
[3]
H. Wang, T. Zariphopoulou, and X. Y. Zhou, Reinforcement learning in continuous time and space: A stochastic control approach, Journal of Machine Learning Research, 21 (2020), pp. 1–34.
[4]
H. Wang, T. Zariphopoulou, and X. Zhou, Exploration versus exploitation in reinforcement learning: A stochastic control approach, arXiv preprint arXiv:1812.01552, (2018).
[5]
H. Wang and X. Y. Zhou, Continuous-time mean–variance portfolio selection: A reinforcement learning framework, Mathematical Finance, 30 (2020), pp. 1273–1308.
[6]
C. Reisinger and Y. Zhang, Regularity and stability of feedback relaxed controls, SIAM Journal on Control and Optimization, 59 (2021), pp. 3118–3151.
[7]
M. Basei, X. Guo, A. Hu, and Y. Zhang, Logarithmic regret for episodic continuous-time linear-quadratic reinforcement learning over a finite-time horizon, Journal of Machine Learning Research, 23 (2022), pp. 1–34.
[8]
L. Szpruch, T. Treetanthiploet, and Y. Zhang, Exploration-exploitation trade-off for continuous-time episodic reinforcement learning with linear-convex models, arXiv preprint arXiv:2112.10264, (2021).
[9]
height 2pt depth -1.6pt width 23pt, Optimal scheduling of entropy regularizer for continuous-time linear-quadratic reinforcement learning, SIAM Journal on Control and Optimization, 62 (2024), pp. 135–166.
[10]
W. Tang, Y. P. Zhang, and X. Y. Zhou, Exploratory HJB equations and their convergence, SIAM Journal on Control and Optimization, 60 (2022), pp. 3191–3216.
[11]
X. Guo, A. Hu, and Y. Zhang, Reinforcement learning for linear-convex models with jumps via stability analysis of feedback controls, SIAM Journal on Control and Optimization, 61 (2023), pp. 755–787.
[12]
M. Giegrich, C. Reisinger, and Y. Zhang, Convergence of policy gradient methods for finite-horizon exploratory linear-quadratic control problems, SIAM Journal on Control and Optimization, 62 (2024), pp. 1060–1092.
[13]
H. Zhao, W. Tang, and D. Yao, Policy optimization for continuous reinforcement learning, Advances in Neural Information Processing Systems, 36 (2023), pp. 13637–13663.
[14]
B. Kerimkulov, J.-M. Leahy, D. Šiška, L. Szpruch, and Y. Zhang, A Fisher–Rao gradient flow for entropy-regularised Markov decision processes in Polish spaces, arXiv preprint arXiv:2310.02951, (2023).
[15]
I. Csiszár, Eine informationstheoretische Ungleichung und ihre Anwendung auf Beweis der Ergodizitaet von Markoffschen Ketten, Magyer Tud. Akad. Mat. Kutato Int. Koezl., 8 (1964), pp. 85–108.
[16]
J. Feydy, T. Séjourné, F.-X. Vialard, S.-i. Amari, A. Trouvé, and G. Peyré, Interpolating between optimal transport and MMD using Sinkhorn divergences, in The 22nd International Conference on Artificial Intelligence and Statistics, PMLR, 2019, pp. 2681–2690.
[17]
L. Chizat, Doubly regularized entropic Wasserstein barycenters, arXiv preprint arXiv:2303.11844, (2023).
[18]
X. Han, R. Wang, and X. Y. Zhou, Choquet regularization for continuous-time reinforcement learning, SIAM Journal on Control and Optimization, 61 (2023), pp. 2777–2801.
[19]
A. Beck and M. Teboulle, A fast iterative shrinkage-thresholding algorithm for linear inverse problems, SIAM journal on imaging sciences, 2 (2009), pp. 183–202.
[20]
J. Mei, C. Xiao, C. Szepesvari, and D. Schuurmans, On the global convergence rates of softmax policy gradient methods, in International Conference on Machine Learning, PMLR, 2020, pp. 6820–6829.
[21]
D. Šiška and Ł. Szpruch, Gradient flows for regularized stochastic control problems, SIAM Journal on Control and Optimization, 62 (2024), pp. 2036–2070.
[22]
P.-C. Aubin-Frankowski, A. Korba, and F. Léger, Mirror descent with relative smoothness in measure spaces, with application to Sinkhorn and EM, Advances in Neural Information Processing Systems, 35 (2022), pp. 17263–17275.
[23]
H. H. Bauschke, J. Bolte, and M. Teboulle, A descent lemma beyond Lipschitz gradient continuity: first-order methods revisited and applications, Mathematics of Operations Research, 42 (2017), pp. 330–348.
[24]
H. Lu, R. M. Freund, and Y. Nesterov, Relatively smooth convex optimization by first-order methods, and applications, SIAM Journal on Optimization, 28 (2018), pp. 333–354.
[25]
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, Proximal policy optimization algorithms, arXiv preprint arXiv:1707.06347, (2017).
[26]
D. Sethi and D. Šiška, The modified MSA, a gradient flow and convergence, The Annals of Applied Probability, 34 (2024), pp. 4455–4492.
[27]
N. V. Krylov, Controlled diffusion processes, Springer, 1980. Translated from Russian by A. B. Aries.
[28]
V. G. Boltyanskii, R. V. Gamkrelidze, and L. S. Pontryagin, The theory of optimal processes. i. the maximum principle, Izv. Akad. Nauk SSSR Ser. Mat, 24 (1960), pp. 3–42.
[29]
L. S. Pontryagin, Mathematical theory of optimal processes, CRC press, 1987.
[30]
H. Dong and N. V. Krylov, The rate of convergence of finite-difference approximations for parabolic Bellman equations with Lipschitz coefficients in cylindrical domains, Applied Mathematics and Optimization, 56 (2007), pp. 37–66.
[31]
I. Gyöngy and D. Šiška, On finite-difference approximations for normalized Bellman equations, Applied Mathematics and Optimization, 60 (2009), pp. 297–339.
[32]
J. Douglas Jr, J. Ma, and P. Protter, Numerical methods for forward-backward stochastic differential equations, The Annals of Applied Probability, 6 (1996), pp. 940–968.
[33]
G. Milstein and M. Tretyakov, Discretization of forward–backward stochastic differential equations and related quasi-linear parabolic equations, IMA journal of numerical analysis, 27 (2007), pp. 24–44.
[34]
F. Delarue and S. Menozzi, A forward–backward stochastic algorithm for quasi-linear PDEs, Ann. Appl. Probab., 16 (2006), pp. 140–184.
[35]
J. Ma, J. Shen, and Y. Zhao, On numerical approximations of forward-backward stochastic differential equations, SIAM Journal on Numerical Analysis, 46 (2008), pp. 2636–2661.
[36]
J. Han and J. Long, Convergence of the deep BSDE method for coupled FBSDEs, Probability, Uncertainty and Quantitative Risk, 5 (2020), pp. 1–33.
[37]
S. Ji, S. Peng, Y. Peng, and X. Zhang, Three algorithms for solving high-dimensional fully coupled FBSDEs through deep learning, IEEE Intelligent Systems, 35 (2020), pp. 71–84.
[38]
height 2pt depth -1.6pt width 23pt, A posteriori error estimates for fully coupled McKean–Vlasov forward-backward SDEs, IMA Journal of Numerical Analysis, 44 (2024), pp. 2323–2369.
[39]
J. Han, A. Jentzen, et al., Deep learning-based numerical methods for high-dimensional parabolic partial differential equations and backward stochastic differential equations, Communications in mathematics and statistics, 5 (2017), pp. 349–380.
[40]
D. Kalise and K. Kunisch, Polynomial approximation of high-dimensional Hamilton–Jacobi–Bellman equations and applications to feedback control of semilinear parabolic PDEs, SIAM Journal on Scientific Computing, 40 (2018), pp. A629–A652.
[41]
J. Sirignano and K. Spiliopoulos, DGM: A deep learning algorithm for solving partial differential equations, Journal of computational physics, 375 (2018), pp. 1339–1364.
[42]
B. Kerimkulov, D. Siska, and L. Szpruch, Exponential convergence and stability of Howard’s policy improvement algorithm for controlled diffusions, SIAM Journal on Control and Optimization, 58 (2020), pp. 1314–1340.
[43]
K. Ito, C. Reisinger, and Y. Zhang, A neural network-based policy iteration algorithm with global \(H^2\)-superlinear convergence for stochastic games on domains, Foundations of Computational Mathematics, 21 (2021), pp. 331–374.
[44]
Y.-J. Huang, Z. Wang, and Z. Zhou, Convergence of policy improvement for entropy-regularized stochastic control problems, arXiv preprint arXiv:2209.07059, (2022).
[45]
H. V. Tran, Z. Wang, and Y. P. Zhang, Policy iteration for exploratory hamilton–jacobi–bellman equations, Applied Mathematics & Optimization, 91 (2025), p. 50.
[46]
Q. Li, L. Chen, C. Tai, and E. Weinan, Maximum principle based algorithms for deep learning, Journal of Machine Learning Research, 18 (2018), pp. 1–29.
[47]
B. Kerimkulov, D. Šiška, and L. Szpruch, A modified MSA for stochastic control problems, Applied Mathematics & Optimization, (2021), pp. 1–20.
[48]
S. Ji and R. Xu, A modified method of successive approximations for stochastic recursive optimal control problems, SIAM Journal on Control and Optimization, 60 (2022), pp. 2759–2786.
[49]
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction, MIT press, 2018.
[50]
Y. Jia and X. Y. Zhou, Policy gradient and actor-critic learning in continuous time and space: Theory and algorithms, The Journal of Machine Learning Research, 23 (2022), pp. 12603–12652.
[51]
J. Mei, Y. Gao, B. Dai, C. Szepesvari, and D. Schuurmans, Leveraging non-uniformity in first-order non-convex optimization, in International Conference on Machine Learning, PMLR, 2021, pp. 7555–7564.
[52]
G. Lan, Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes, Mathematical programming, 198 (2023), pp. 1059–1106.
[53]
D. Sethi, D. Šiška, and Y. Zhang, Entropy annealing for policy mirror descent in continuous time and space, arXiv preprint arXiv:2405.20250, (2024).
[54]
C. Reisinger, W. Stockinger, and Y. Zhang, Linear convergence of a policy gradient method for some finite horizon continuous time control problems, SIAM Journal on Control and Optimization, 61 (2023), pp. 3526–3558.
[55]
B. Jourdain and A. Tse, Central limit theorem over non-linear functionals of empirical measures with applications to the mean-field fluctuation of interacting diffusions, Electronic Journal of Probability, 26 (2021), pp. 1–34.
[56]
X. Guo and Y. Zhang, Towards an analytical framework for dynamic potential games, SIAM Journal on Control and Optimization, 63 (2025), pp. 1213–1242.
[57]
J. Zhang, Backward stochastic differential equations, Springer, 2017.
[58]
A. Davey and H. Zheng, Convergence of proximal policy gradient method for problems with control dependent diffusion coefficients, arXiv preprint arXiv:2505.18379, (2025).
[59]
P. Plank and Y. Zhang, Policy optimization for continuous-time linear-quadratic graphon mean field games, arXiv preprint arXiv:2506.05894, (2025).
[60]
R. A. Howard, Dynamic programming and Markov processes., John Wiley, 1960.
[61]
S. Kakade and J. Langford, Approximately optimal approximate reinforcement learning, in Proceedings of the nineteenth international conference on machine learning, 2002, pp. 267–274.
[62]
K. Pearson, On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling, The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, 50 (1900), pp. 157–175.
[63]
J. Harter and A. Richou, A stability approach for solving multidimensional quadratic BSDEs, Electron. J. Probab., 24 (2019), pp. 1–51.
[64]
N. Kazamaki, Continuous exponential martingales and BMO, Springer, 2006.
[65]
P. Dupuis and R. S. Ellis, A weak convergence approach to the theory of large deviations, John Wiley & Sons, Inc., New York, 1997.
[66]
J. F. Bonnans and A. Shapiro, Perturbation analysis of optimization problems, Springer, 2013.
[67]
P. G. Ciarlet, Linear and nonlinear functional analysis with applications, vol. 130, SIAM, 2013.
[68]
A. H. Guide, Infinite dimensional analysis, Springer, 2006.