Convergence of fictitious play for fully coupled FBSDEs in finite-player stochastic differential games


Abstract

In this article we investigate the theoretical convergence properties of the fictitious-play approximation procedure applied to coupled FBSDE systems for finite-player non-zero-sum stochastic differential games. Under one set of assumptions, the convergence is shown to be geometric. Under an additional structural assumption, the geometric convergence rate further improves to a super-exponential rate in a special class of games. To the best of our knowledge, this provides the first convergence analysis of fictitious play for fully coupled FBSDEs. A numerical experiment with a linear-quadratic interbank borrowing and lending problem confirms the geometric convergence.

1 Introduction↩︎

Stochastic differential games (SDG) provide a mathematical framework for modeling interacting agents or players whose actions or controls influence the evolution of a common stochastic system. In such problems, each player optimizes an individual objective while accounting for the strategies of the other players, leading to the famous Nash equilibrium solution concept. Non-zero-sum stochastic differential games arise naturally in multi-agent systems under uncertainty, where each agent optimizes an individual objective, with applications ranging from financial markets and energy systems to economics and engineering. Representative examples include interbank lending and systemic risk models [1], applications in economics such as dynamic oligopoly and resource management problems in environmental and engineering systems [2], multi-agent optimal execution problems with price impact [3], pursuit-evasion [4][7] and missile guidance and control [8][12].

Similarly to stochastic control, stochastic differential games are either solved with the stochastic maximum principle [6], [13], [14], or by dynamic programming [3], [6], [7], [9], [10], [12], [13], [15][18]. In this work, we focus on the latter approach. Approaching SDG via dynamic programming leads to a coupled system of Hamilton–Jacobi–Bellman (HJB) equations, the Nash PDE system, and to a coupled system of forward-backward stochastic differential equations (FBSDEs), the Nash FBSDE system. In regular settings, these two representations are equivalent, but the Nash FBSDE system is a more general representation for regimes with lower regularity.

Only games with highly specific structure, such as the important and well-studied class of linear–quadratic games, admit analytical or semi-analytical solutions. Therefore, for many applications, numerical approximations are necessary. However, accurate, scalable, and practical numerical approximation of SDG remains an open problem. Finite difference methods have been studied for two-player zero-sum games [19] and for mean-field games [20], [21]. Grid-based finite difference and finite element methods for PDEs, however, suffer from the curse of dimensionality (COD) and are therefore limited to low-dimensional settings. To the best of our knowledge, direct approximation of the Nash PDE system for non-zero-sum stochastic differential games has not been studied in the literature, and would in any case be of limited interest because of the aforementioned scaling problem. A Markov chain approximation framework for two-player zero-sum games was introduced in [17], and extended to finite-player non-zero-sum stochastic differential games in [18]. It was later extended to more general dynamics, including regime switching and jump diffusions [13]. However, this approach is also grid-based and thus suffers from the COD.

Han [22], Han and Hu [23], and Han, Hu, and Long [24] proposed so-called deep fictitious play algorithms as a general numerical scheme for the Nash FBSDE system. It is based on the iterative fictitious-play procedure introduced in [25]. It is a Picard-type iteration in which, at each step, every player solves a single-player stochastic control problem with the other players’ control policies fixed. Approximating the resulting single-player fictitious-play FBSDEs using the deep BSDE method [26] leads to deep fictitious-play. In [24], Han, Hu, and Long proved the geometric convergence of fictitious play when the forward process of the FBSDE is decoupled from the solution components of the backward equation. This decoupled setting is relevant because it is possible, by means of a transformation, to remove the coupling in the forward process, after an appropriate modification of the backward equation. However, when the original deep BSDE method is applied directly to coupled FBSDEs from stochastic control, as shown empirically in [27], [28], it fails to converge except for rather obscure parameter choices. Based on the authors’ experience in the work related to [27], [28], the aforementioned decoupling does not appear to remedy the convergence problem. For this reason, it is desirable to work with fully coupled FBSDEs and instead use one of the robust deep FBSDE methods in [27], [28], which are practically applicable to a larger class of problems.

In the present work, we extend the analysis in [24] by showing geometric convergence of fictitious play in the fully coupled setting. Under an alternative set of assumptions, we show that the convergence is super-exponential for a special class of games. We consider both stopped games on bounded domains under additional smoothness assumptions and unrestricted games under either a gradient bound assumption or a typical smallness condition on the coefficients or the terminal time. The former case is closely related to BSDEs on domains, for which deep learning-based approximation methods with error guarantees have recently been developed [29]. The geometric rate is verified experimentally for a linear–quadratic interbank borrowing and lending game with 2 and 20 players.

The paper is organized as follows. Section 2 introduces the game-theoretic setting and presents the Nash PDE and Nash FBSDE formulations. To make the manuscript more self-contained, accessible, and useful to a wider range of readers, this section is intentionally comprehensive and concludes with examples of games. Section 3 presents the fictitious-play approximation, the detailed setting, and convergence analysis including our main result. In Section 4, we present our numerical experiment.

2 General setting and preliminaries on SDG↩︎

To make the manuscript more self-contained, accessible, and useful to a broader range of readers, we introduce stochastic differential games from both the PDE and FBSDE perspectives. We present the assumptions, give a high-level overview of the synthesis through dynamic programming and end with concrete examples which demonstrate our abstract setting.

2.1 Spaces of stochastic processes and preliminaries↩︎

Throughout this paper, let \(T\in(0,\infty)\), \(n\in\mathbb{N}\), and let \((\boldsymbol{W}_t)_{t\in[0,T]}\) be an \(n\)-dimensional standard Brownian motion on a filtered probability space \((\Omega,\mathcal{F},(\mathcal{F}_t)_{t\in[0,T]},\mathbb{P})\). For \(\ell\geq1\), denote by \(\|\cdot\|_{\mathbb{R}^\ell}\) and \(\langle\cdot,\cdot\rangle_{\mathbb{R}^\ell}\) the norm and scalar product on \(\mathbb{R}^\ell\), and by \(\vert\!\vert\!\vert \cdot \vert\!\vert\!\vert_{\mathbb{R}^{\ell \times n}}\) the Frobenius norm on \(\mathbb{R}^{\ell\times n}\). Furthermore, for \(p\in[1,\infty)\), let \(\mathbb{L}^{p,\ell}\) be the space of random variables \(U\colon\Omega\to\mathbb{R}^\ell\) such that \[\begin{align} \|U\|_{\mathbb{L}^{p,\ell}} \vcentcolon=\Big(\mathbb{E}\Big[\|U\|_{\mathbb{R}^\ell}^p\Big]\Big)^\frac{1}{p}<\infty. \end{align}\] For \(t \in (0,T]\), \(p \in [1, \infty)\), and \(\beta \in [0,\infty)\), let \(\mathbb{S}^{p,\ell}_{\beta,t}\), \(\mathbb{H}^{p,\ell}_{\beta,t}\), and \(\mathbb{H}_{\beta,t}^{p,\ell,n}\) denote the Banach spaces of predictable processes \(\mathcal{Y},\mathcal{Z}\colon\Omega\times[0,t]\to\mathbb{R}^\ell\), and \(\mathcal{W}\colon\Omega\times[0,t]\to\mathbb{R}^{\ell\times n}\), that satisfy \[\begin{align} \|\mathcal{Y}\|_{\mathbb{S}^{p,\ell}_{\beta,t}} &\vcentcolon=\bigg(\mathbb{E}\bigg[\sup_{s\in[0,t]}\text{e}^{\beta s}\|\mathcal{Y}_s\|_{\mathbb{R}^{\ell}}^p\bigg]\bigg)^{\frac{1}{p}}<\infty,\\ \|\mathcal{Z}\|_{\mathbb{H}^{p,\ell}_{\beta,t}} &\vcentcolon= \bigg(\mathbb{E}\bigg[\bigg(\int_0^t\text{e}^{\beta s}\|\mathcal{Z}_s\|_{\mathbb{R}^{\ell}}^2\,\text{d}s\bigg)^{\frac{p}{2}}\bigg]\bigg)^\frac{1}{p} <\infty,\\ \|\mathcal{W}\|_{\mathbb{H}^{p,\ell,n}_{\beta,t}} &\vcentcolon= \bigg(\mathbb{E}\bigg[\bigg(\int_0^t\text{e}^{\beta s}\vert\!\vert\!\vert \mathcal{W}_s \vert\!\vert\!\vert_{\mathbb{R}^{\ell\times n}}^2\,\text{d}s\bigg)^{\frac{p}{2}}\bigg]\bigg)^\frac{1}{p} <\infty. \end{align}\] In the case \(\beta=0\), we suppress \(\beta\) from the notation. It follows immediately from the definitions that, for \(\mathcal{Y} \in \mathbb{S}^{2,\ell}_{t}\), \(\mathcal{Z} \in \mathbb{H}^{2,\ell}_{t}\), we have the norm equivalences \[\label{eq:HHH} \begin{align} \|\mathcal{Y}\|^2_{\mathbb{S}^{2,\ell}_{t}} &\le \|\mathcal{Y}\|^2_{\mathbb{S}^{2,\ell}_{\beta,t}} \le e^{\beta t}\|\mathcal{Y}\|^2_{\mathbb{S}^{2,\ell}_{t}}, \\ \|\mathcal{Z}\|^2_{\mathbb{H}^{2,\ell}_{t}} &\le \|\mathcal{Z}\|^2_{\mathbb{H}^{2,\ell}_{\beta,t}} \le e^{\beta t}\|\mathcal{Z}\|^2_{\mathbb{H}^{2,\ell}_{t}}. \end{align}\tag{1}\] Moreover, if \(\mathcal{Y} \in \mathbb{S}^{2,1}_t\) and \(\mathcal{Z} \in \mathbb{H}^{2,n}_t\), then \(\mathcal{Y} \in \mathbb{H}^{2,1}_t\) and \(\mathcal{Y} \mathcal{Z} \in \mathbb{H}^{1,n}_t\), with \[\begin{align} \tag{2} \|\mathcal{Y}\|_{\mathbb{H}^{2,1}_{\beta,t}}^2 &\leq t \|\mathcal{Y}\|_{\mathbb{S}^{2,1}_{\beta,t}}^2,\\ \tag{3} \|\mathcal{Y} \mathcal{Z}\|_{\mathbb{H}^{1,n}_{\beta,T}} &\leq \|\mathcal{Y}\|_{\mathbb{S}^{2,1}_{\beta,T}} \|\mathcal{Z}\|_{\mathbb{H}^{2,n}_{\beta,T}}. \end{align}\] For \(p=2\), the spaces \(\mathbb{H}^{2,\ell}_{\beta, t}\), \(\beta \geq 0\), are Hilbert spaces with equivalent scalar products \[\begin{align} \big\langle \mathcal{Z}^1,\mathcal{Z}^2\big\rangle_{\mathbb{H}^{2,\ell}_{\beta,t}}\vcentcolon= \mathbb{E}\bigg[\int_0^t\text{e}^{\beta s} \big\langle \mathcal{Z}_s^1,\mathcal{Z}_s^2\big\rangle_{\mathbb{R}^{\ell}}\,\text{d}s\bigg] ,\quad \mathcal{Z}^1, \mathcal{Z}^2\in\mathbb{H}_{\beta, t}^{2,\ell}. \end{align}\] Moreover, for \(\mathcal{Z}^1, \mathcal{Z}^2 \in \mathbb{H}^{2,\ell}_T\), a straightforward application of the Cauchy–Schwarz inequality yields \[\begin{align} \bigg\| \int_{\cdot}^T \big\langle \mathcal{Z}^1_s, \mathcal{Z}^2_s \big\rangle_{\mathbb{R}^\ell} \, \mathrm{d}s\bigg\|_{\mathbb{S}^{1,1}_T} \leq \big\| \mathcal{Z}^1 \big\|_{\mathbb{H}^{2,\ell}_T} \big\| \mathcal{Z}^2 \big\|_{\mathbb{H}^{2,\ell}_T}. \label{eq:integral95in95S1195to95Hnorms} \end{align}\tag{4}\] The Burkholder–Davis–Gundy inequality (see [30]) states that for \(p\in\{1,2\}\), there exist constants \(k_p,K_p\in(0,\infty)\) such that, for all \(\phi\in \mathbb{H}^{p,\ell,n}_{t}\), we have \[\begin{align} \label{eq:BDG1} k_p\|\phi\|_{\mathbb{H}_{t}^{p,\ell,n}} \le \bigg\| \int_0^\cdot \phi_s\,\text{d}\boldsymbol{W}_s\bigg\|_{\mathbb{S}_{t}^{p,\ell}} \le K_p\|\phi\|_{\mathbb{H}_{t}^{p,\ell,n}}. \end{align}\tag{5}\] Moreover, if \(\phi\in \mathbb{H}^{1,\ell,n}_{t}\), then \[\begin{align} \label{eq:BDG2} \mathbb{E}\!\left[\int_0^t \phi_s\,\text{d}\boldsymbol{W}_s\right]=0. \end{align}\tag{6}\]

We next introduce notation and basic properties of stopped (or killed) stochastic processes. Let \(\sigma_1,\sigma_2,\sigma_3,\sigma_4\) be \((\mathcal{F}_t)\)-stopping times. For a predictable process \(f\colon\Omega\times[0,T]\to\mathbb{R}^\ell\), define \[\begin{align} f^{[\sigma_1,\sigma_2]} \vcentcolon=f\,\mathbf{1}_{[\sigma_1,\sigma_2]}, \qquad f^{[\sigma_1]} \vcentcolon=f^{[0,\sigma_1]}. \end{align}\] For \(\Phi\in \mathbb{H}^{2,\ell}_{\beta, T}\), \(\Psi\in \mathbb{S}^{2,\ell}_{\beta, T}\), if \(\sigma_1 < \sigma_2 \leq \sigma_3 < \sigma_4\), it holds that \[\begin{align} \tag{7} \Big\| \Phi^{[\sigma_1,\sigma_2]} + \Phi^{[\sigma_3,\sigma_4]} \Big\|_{\mathbb{H}^{2,\ell}_{\beta,T}}^2 &= \Big\| \Phi^{[\sigma_1,\sigma_2]} \Big\|_{\mathbb{H}^{2,\ell}_{\beta,T}}^2 + \Big\| \Phi^{[\sigma_3,\sigma_4]} \Big\|_{\mathbb{H}^{2,\ell}_{\beta,T}}^2,\\ \Big\|\Psi^{[\sigma_1,\sigma_2]}+\Psi^{[\sigma_3,\sigma_4]}\Big\|_{\mathbb{S}^{2,\ell}_{\beta,T}}^2 & \le \Big\|\Psi^{[\sigma_1,\sigma_2]}\Big\|_{\mathbb{S}^{2,\ell}_{\beta,T}}^2 +\Big\|\Psi^{[\sigma_3,\sigma_4]}\Big\|_{\mathbb{S}^{2,\ell}_{\beta,T}}^2. \tag{8} \end{align}\]

2.2 The stochastic differential game↩︎

Throughout this paper, let \(d, n, N \in \mathbb{N}\). Furthermore, let \(\mathcal{D}\subseteq\mathbb{R}^n\) be either \(\mathcal{D}=\mathbb{R}^n\) or an open, bounded spatial domain with boundary \(\partial \mathcal{D}\), and define the space-time domain \(Q_T^0=[0,T)\times \mathcal{D}\), its closure \(Q_T=[0,T]\times \overline{\mathcal{D}}\), and its boundary \(\partial Q_T=[0,T)\times \partial \mathcal{D}\cup \{T\}\times \overline{\mathcal{D}}\). We consider an \(N\)-player game and denote by \(\mathcal{I}\vcentcolon=\{1,2,\ldots,N\}\) the set of all players. The strategy of each player \(i \in \mathcal{I}\) is given by a Markov control policy \(\alpha^i \colon Q_T \to A^i\subset \mathbb{R}^{d}\), where \(A^i\subset \mathbb{R}^d\) is the compact control set for player \(i\). The policy is not given a priori, but is to be optimized. For \(i\in\mathcal{I}\), let \(\mathbb{A}^i\) denote the set of all \(\alpha^i\) for which there exists \(C=C(\alpha^i) \geq0\) such that, for all \((t_1,\boldsymbol{x}_1),(t_2,\boldsymbol{x}_2)\in Q_T\), we have \[\begin{align} \big\|\alpha^i(t_1,\boldsymbol{x}_1)-\alpha^i(t_2,\boldsymbol{x}_2)\big\|_{\mathbb{R}^{d}}^2 &\leq C \big( |t_1-t_2| + \|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n}^2 \big). \end{align}\] We denote by \(A \vcentcolon=A^1 \times \cdots \times A^N \subset \mathbb{R}^{d\times N}\) the set of all possible joint actions. Similarly, the set \(\mathbb{A}\) consists of all joint Markov policies \(\boldsymbol{\alpha}= (\alpha^1,\dots,\alpha^N) \colon Q_T \to A\), where \(\alpha^i \in \mathbb{A}^i\) for each \(i \in \mathcal{I}\). Let \(\bar{\boldsymbol{b}}\colon Q_T\times A\to\mathbb{R}^n\) and \(\boldsymbol{\Sigma}\colon Q_T\to\mathbb{R}^{n\times n}\) be the drift and diffusion coefficients of the common state process. We assume that there exist \(L_{\bar{\boldsymbol{b}}},L_{\boldsymbol{\Sigma}}\geq0\) such that for all \((t_1,\boldsymbol{x}_1,\boldsymbol{a}_1),(t_2,\boldsymbol{x}_2,\boldsymbol{a}_2)\in Q_T\times A\) with \(\boldsymbol{a}_j=(a_j^1,\dots,a_j^N)\), \(j=1,2\), we have \[\begin{align} \tag{9} \big\| \bar \boldsymbol{b}(t_1,\boldsymbol{x}_1,\boldsymbol{a}_1)-\bar \boldsymbol{b}(t_2,\boldsymbol{x}_2,\boldsymbol{a}_2) \big\|_{\mathbb{R}^{n}}^2 &\leq L_{\bar{\boldsymbol{b}}} \big( |t_1-t_2| + \|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n}^2 + \big|\!\big|\!\big| \boldsymbol{a}_1-\boldsymbol{a}_2 \big|\!\big|\!\big|_{\mathbb{R}^{d\times N}}^2 \big),\\ \big|\!\big|\!\big| \boldsymbol{\Sigma}(t_1,\boldsymbol{x}_1)- \boldsymbol{\Sigma}(t_2,\boldsymbol{x}_2) \big|\!\big|\!\big|_{\mathbb{R}^{n\times n}}^2 &\leq L_{\boldsymbol{\Sigma}} \big( |t_1-t_2| + \|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n}^2 \big). \tag{10} \end{align}\] We assume that \(0\in A^i\) for all \(i\in\mathcal{I}\) so that the trivial control \(0\in\mathbb{A}\) is admissible. This implies that \(\bar \boldsymbol{b}\) satisfies a linear growth condition in \(\boldsymbol{x}\), uniformly over \(\boldsymbol{a}\in A\), which we do not state explicitly. In Section 2.4, after a change of variables, we introduce the related coefficient \(\boldsymbol{b}\). We also assume uniform parabolicity, i.e., there exist constants \(0 < c < C\) such that, for all \((t,\boldsymbol{x})\in Q_T\) and \(\boldsymbol{\xi}\in\mathbb{R}^n\), we have \[\begin{align} \label{eq:UP} c\|\boldsymbol{\xi}\|_{\mathbb{R}^n}^2 \leq \boldsymbol{\xi}^\top\boldsymbol{\Sigma}(t,\boldsymbol{x}) \boldsymbol{\Sigma}^\top(t,\boldsymbol{x})\boldsymbol{\xi} \leq C\|\boldsymbol{\xi}\|_{\mathbb{R}^n}^2. \end{align}\tag{11}\] Uniform parabolicity implies that \(\boldsymbol{\Sigma}(t,\boldsymbol{x})\) is invertible for every \((t,\boldsymbol{x})\in Q_T\), with inverse uniformly bounded on \(Q_T\). Together with 10 , this also implies that \[\begin{align} \label{bPhi} \boldsymbol{\Phi}(t,\boldsymbol{x}) \vcentcolon=\big(\boldsymbol{\Sigma}^\top(t,\boldsymbol{x})\big)^{-1} \end{align}\tag{12}\] is uniformly bounded and Lipschitz continuous on \(Q_T\). Consequently, there exist constants \(C_{\boldsymbol{\Phi}}, L_{\boldsymbol{\Phi}} \geq 0\) such that, for all \((t_1,\boldsymbol{x}_1),(t_2,\boldsymbol{x}_2)\in Q_T\), we have \[\begin{align} \label{eq:Lip95Phi} \vert\!\vert\!\vert \boldsymbol{\Phi}(t_1,\boldsymbol{x}_1) \vert\!\vert\!\vert_{\mathbb{R}^{n\times n}}\leq C_{\boldsymbol{\Phi}}, \quad \textrm{and} \quad \vert\!\vert\!\vert \boldsymbol{\Phi}(t_1,\boldsymbol{x}_1)-\boldsymbol{\Phi}(t_2,\boldsymbol{x}_2) \vert\!\vert\!\vert^2_{\mathbb{R}^{n\times n}} \leq L_{\boldsymbol{\Phi}} \big( |t_1-t_2|+\|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n}^2 \big). \end{align}\tag{13}\] The state dynamics are determined by the predictable, square-integrable and up to modification unique stochastic processes \(\boldsymbol{X}^{t,\boldsymbol{x},\boldsymbol{\alpha}}\colon[t,T]\times\Omega\to\overline{\mathcal{D}}\), which for all \(s\in[t,T]\), \(\mathbb{P}\)-almost surely satisfy \[\begin{align} \label{eq:state95alpha95process} \boldsymbol{X}_s^{t,\boldsymbol{x},\boldsymbol{\alpha}} = \boldsymbol{x} + \int_t^{s\wedge\tau} \bar{\boldsymbol{b}}\big(r,\boldsymbol{X}_r^{t,\boldsymbol{x},\boldsymbol{\alpha}},\boldsymbol{\alpha}\big(r,\boldsymbol{X}_r^{t,\boldsymbol{x},\boldsymbol{\alpha}}\big)\big)\,\text{d}r + \int_t^{s\wedge\tau} \boldsymbol{\Sigma}\big(r,\boldsymbol{X}_r^{t,\boldsymbol{x},\boldsymbol{\alpha}}\big)\,\text{d}\boldsymbol{W}_r. \end{align}\tag{14}\] Here, \(\tau=\tau(t,\boldsymbol{x},\boldsymbol{\alpha})\colon\Omega \to[t,T]\) denotes the exit time of \(Q_T^0\), i.e., the first time at which \(\boldsymbol{X}^{t,\boldsymbol{x},\boldsymbol{\alpha}}\) reaches either \(\partial \mathcal{D}\) or the terminal boundary \(\{T\}\times\overline{\mathcal{D}}\). To simplify notation, we suppress the \((t,\boldsymbol{x},\boldsymbol{\alpha})\)-dependence of \(\tau\). Under our assumptions, these processes are well defined by standard SDE theory.

For \(i \in \mathcal{I}\), let \(\bar{f}^i \colon Q_T \times A \to \mathbb{R}\) and \(g^i \colon \partial Q_T \to \mathbb{R}\) denote player \(i\)’s running and terminal cost functions, respectively. The objective of player \(i\) is to minimize its individual cost functional, which depends both on its own control and on the actions of the remaining \(N - 1\) players. These cost functionals \(J^i\colon Q_T\times \mathbb{A}\to \mathbb{R}\), \(i\in\mathcal{I}\), are given by \[\label{eq:cost95alpha} \begin{align} &J^i(t, \boldsymbol{x}, \boldsymbol{\alpha}) \vcentcolon=\mathbb{E}\bigg[ \int_t^\tau \bar{f}^i\big(s, \boldsymbol{X}_s^{t,\boldsymbol{x},\boldsymbol{\alpha}}, \boldsymbol{\alpha}\big(s, \boldsymbol{X}_s^{t,\boldsymbol{x},\boldsymbol{\alpha}}\big)\big) \, \text{d}s + g^i\big(\tau,\boldsymbol{X}_\tau^{t,\boldsymbol{x},\boldsymbol{\alpha}}\big) \bigg]. \end{align}\tag{15}\] Under our assumptions, the processes \(\boldsymbol{X}^{t,\boldsymbol{x},\boldsymbol{\alpha}}\) have finite moments of all orders. Thus, polynomial bounds on \(\bar f^{1},\dots,\bar f^{N}\) and \(g^{1},\dots,g^{N}\) are sufficient to ensure that \(J^{1},\dots,J^{N}\) are well defined. However, much stronger assumptions are required for the solution theory; see, e.g., Examples 34 below. For the convergence analysis in Section 3.3, we impose the following Lipschitz continuity. For each \(i \in \mathcal{I}\), there exist \(L_{\bar{f}^i}, L_{g^i} \geq 0\) such that for all \((t_1, \boldsymbol{x}_1, \boldsymbol{a}_1), (t_2, \boldsymbol{x}_2, \boldsymbol{a}_2) \in Q_T \times A\), \[\begin{align} \big| \bar{f}^i(t_1, \boldsymbol{x}_1, \boldsymbol{a}_1) - \bar{f}^i(t_2, \boldsymbol{x}_2, \boldsymbol{a}_2) \big|^2 &\leq L_{\bar{f}^i} \big( |t_1 - t_2| + \| \boldsymbol{x}_1 - \boldsymbol{x}_2 \|^2_{\mathbb{R}^n} + \vert\!\vert\!\vert \boldsymbol{a}_1 - \boldsymbol{a}_2 \vert\!\vert\!\vert^2_{\mathbb{R}^{d\times N}} \big), \end{align}\] and, for all \((t_1, \boldsymbol{x}_1), (t_2, \boldsymbol{x}_2) \in \partial Q_T\), \[\begin{align} \big| g^i(t_1, \boldsymbol{x}_1) - g^i(t_2, \boldsymbol{x}_2) \big|^2 &\leq L_{g^i} \big( |t_1 - t_2| + \| \boldsymbol{x}_1 - \boldsymbol{x}_2 \|^2_{\mathbb{R}^n} \big). \end{align}\] For \(\boldsymbol{a}\in A\) and \(i\in\mathcal{I}\), let \(A^{-i}\) denote the set of actions \[\begin{align} \boldsymbol{a}^{-i} \vcentcolon=\big(a^1,\dots,a^{i-1},a^{i+1},\dots,a^N\big), \end{align}\] and use the analogous notation \(\boldsymbol{\alpha}^{-i}\in \mathbb{A}^{-i}\) for joint policies. For \(i \in \mathcal{I}\), define \([\cdot,\cdot]_i\colon A^i\times A^{-i}\to A\) by \[\begin{align} [a, \boldsymbol{a}^{-i}]_i=(a^1,\dots,a^{i-1},a,a^{i+1},a^N), \quad a \in A^i,\;\boldsymbol{a}^{-i} \in A^{-i}. \end{align}\] From an optimization perspective, it is not sufficient for a player \(i \in \mathcal{I}\) to find an optimal strategy \(\beta^{i}\) against an arbitrary joint Markov policy \(\boldsymbol{\alpha}^{-i}\) of the other \(N-1\) players, because those players are optimizing their strategies as well. For a solution to be truly optimal, each player \(i \in \mathcal{I}\) must have a strategy \(\beta^{i}=\alpha^{i,*}\) that is optimal given the strategies \(\boldsymbol{\alpha}^{-i,*}=(\alpha^{1,*},\dots,\alpha^{i-1,*},\alpha^{i+1,*},\alpha^{N,*})\) of the others. Otherwise, some player would have an incentive to change strategy, which contradicts optimality. Moreover, the cost functional of player \(i\) depends on the strategies of the other players directly, and indirectly through the controlled state process. Hence, a change in the strategy of any other player may cause \(\alpha^{i,*}\) to lose its optimality. Thus, optimal strategies are generally not stable under unilateral changes unless they are jointly optimized. This leads to the notion of a Nash equilibrium: no player can benefit from unilaterally deviating from their chosen strategy. Formally, an \(N\)-player closed-loop Nash equilibrium is a tuple of strategies \(\boldsymbol{\alpha}^* = (\alpha^{1,*}, \alpha^{2,*}, \ldots, \alpha^{N,*}) \in \mathbb{A}\) such that for every player \(i \in \mathcal{I}\), every Markov policy \(\alpha^i \in \mathbb{A}^i\), and every \((t,\boldsymbol{x})\in Q_T\), we have \[J^i\big(t,\boldsymbol{x},\boldsymbol{\alpha}^{*}\big) \leq J^i\big(t,\boldsymbol{x},\big[\alpha^i,\boldsymbol{\alpha}^{-i,*}\big]_i\big).\]

2.3 Synthesis through dynamic programming↩︎

The synthesis of the game through dynamic programming is based on the players’ value functions \(V^i\colon Q_T\times \mathbb{A}^{-i}\to \mathbb{R}\), \(i\in\mathcal{I}\), defined by \[V^i\big(t,\boldsymbol{x},\boldsymbol{\alpha}^{-i}\big) \vcentcolon= \inf_{\alpha^i\in\mathbb{A}^i} J^{i}\big(t,\boldsymbol{x},\big[\alpha^i,\boldsymbol{\alpha}^{-i}\big]_i\big).\] Consider player \(i \in \mathcal{I}\) and fix an arbitrary joint policy \(\boldsymbol{\alpha}^{-i}\in\mathbb{A}^{-i}\) for the remaining \(N - 1\) players. From the perspective of player \(i\), these players are treated as fixed and therefore become part of the dynamics. Thus, player \(i\) solves a single-player stochastic control problem and the associated Hamilton–Jacobi–Bellman (HJB) equation. More precisely, we seek \(\mathcal{V}^{i}\colon Q_T\times \mathbb{A}^{-i}\to\mathbb{R}\), which for all \((t,\boldsymbol{x},\boldsymbol{\alpha}^{-i})\in Q_T^0\times \mathbb{A}^{-i}\) satisfies \[\label{eq:HJB} \frac{\partial \mathcal{V}^i}{\partial t} \big(t,\boldsymbol{x},\boldsymbol{\alpha}^{-i}\big) + \big(\mathcal{L}_t \mathcal{V}^i\big)\big(t,\boldsymbol{x},\boldsymbol{\alpha}^{-i}\big) + H^i\big(t,\boldsymbol{x},\mathrm{D}_{\boldsymbol{x}}\mathcal{V}^i\big(t,\boldsymbol{x},\boldsymbol{\alpha}^{-i}\big),\boldsymbol{\alpha}^{-i}(t,\boldsymbol{x})\big) = 0,\tag{16}\] and for \((t,\boldsymbol{x})\in\partial Q_T\) satisfies the boundary condition \[\label{eq:HJB95terminal} \mathcal{V}^i\big(t, \boldsymbol{x},\boldsymbol{\alpha}^{-i}\big) = g^i(t,\boldsymbol{x}).\tag{17}\] The Hamiltonian \(H^i\colon Q_T\times\mathbb{R}^n\times A^{-i}\to \mathbb{R}\) is given by \[\begin{align} H^i\big(t,\boldsymbol{x},p,\boldsymbol{a}^{-i}\big):= \inf_{a^i\in A^{i}}\Big(\big\langle\bar{\boldsymbol{b}}\big(t,\boldsymbol{x},\big[a^i,\boldsymbol{a}^{-i}\big]_i\big),p\big\rangle + \bar{f}^i\big(t,\boldsymbol{x},\big[a^i,\boldsymbol{a}^{-i}\big]_i\big)\Big), \end{align}\] and the second-order operators \(\mathcal{L}_t\), \(t\in[0,T]\), which act on \(\phi\in C^2(Q_T,\mathbb{R})\), by \[\begin{align} (\mathcal{L}_t\phi)(t,\boldsymbol{x}):=\frac{1}{2} \mathrm{Tr}\big(\boldsymbol{\Sigma}(t,\boldsymbol{x})\, \mathrm{D}_{\boldsymbol{x}}^2 \phi(t,\boldsymbol{x})\, \boldsymbol{\Sigma}^\top(t,\boldsymbol{x})\big). \end{align}\] Here, \(\mathrm{D}_{\boldsymbol{x}} \mathcal{V}^i\) and \(\mathrm{D}_{\boldsymbol{x}}^2 \mathcal{V}^i\) denote the gradient and Hessian of \(\mathcal{V}^i\) with respect to \(\boldsymbol{x}\), respectively, and \(\mathrm{Tr}\) denotes the trace operator. Verification theorems show that, in settings where the HJB equation 16 admits a sufficiently regular solution, this solution coincides with the value function, i.e., that \(V^i(t,\boldsymbol{x},\boldsymbol{\alpha}^{-i})=\mathcal{V}^i(t,\boldsymbol{x},\boldsymbol{\alpha}^{-i})\) and that the optimal policy \(\beta^{i}\colon Q_T\times \mathbb{A}^{-i}\to A^i\) is given by \[\begin{align} \beta^{i}\big(t,\boldsymbol{x},\boldsymbol{\alpha}^{-i}\big) = \kappa^i\big(t,\boldsymbol{x},\mathrm{D}_{\boldsymbol{x}}\mathcal{V}^i\big(t,\boldsymbol{x},\boldsymbol{\alpha}^{-i}\big),\boldsymbol{\alpha}^{-i}(t,\boldsymbol{x})\big). \end{align}\] For fixed \(\boldsymbol{\alpha}^{-i}\), the policy \((t,\boldsymbol{x}) \mapsto \beta^{i}(t,\boldsymbol{x},\boldsymbol{\alpha}^{-i})\) in \(\mathbb{A}^i\) is given by \(\kappa^i\colon Q_T\times \mathbb{R}^n\times A^{-i}\to A^i\), where \[\begin{align} \kappa^i\big(t,\boldsymbol{x},p,\boldsymbol{a}^{-i}\big)\in\mathop{\mathrm{arg\,min}}_{a^i\in A^i} \Big(\big\langle\bar{\boldsymbol{b}}\big(t,\boldsymbol{x},\big[a^i,\boldsymbol{a}^{-i}\big]_i\big),p\big\rangle + \bar{f}^i\big(t,\boldsymbol{x},\big[a^i,\boldsymbol{a}^{-i}\big]_i\big)\Big). \end{align}\] We assume that \(\kappa^1, \dots, \kappa^N\) exist and are unique, in the sense that, for every \((t,\boldsymbol{x},p,\boldsymbol{a}^{-i})\in Q_T\times \mathbb{R}^n\times A^{-i}\), each minimization problem defining \(\kappa^1, \dots, \kappa^N\) admits a unique solution; see Example 3 below. For the sake of exposition, throughout the present section we also assume the existence of a sufficiently regular \(\mathcal{V}\) satisfying an appropriate verification theorem, and therefore write \(V\) instead of \(\mathcal{V}\). This somewhat imprecise assumption is replaced in Section 3 by concrete and stronger assumptions.

We next introduce the vector-valued mappings \(\boldsymbol{V}\colon Q_T \times \mathbb{A}\to \mathbb{R}^N\), \(\boldsymbol{\kappa} \colon Q_T \times \mathbb{R}^{n \times N} \times A \to A\) and \(\boldsymbol{\beta}\colon Q_T \times \mathbb{A}\to A\). Writing \(\boldsymbol{p}=(p^1,\dots,p^N)\in\mathbb{R}^{n\times N}\), these are defined by \[\begin{align} \boldsymbol{V}(t,\boldsymbol{x},\boldsymbol{\alpha}) &\vcentcolon= \big( V^1(t,\boldsymbol{x},\boldsymbol{\alpha}^{-1}), \dots, V^N(t,\boldsymbol{x},\boldsymbol{\alpha}^{-N}) \big),\\ \mathbf{\boldsymbol{\kappa}}(t,\boldsymbol{x},\boldsymbol{p},\boldsymbol{a}) &\vcentcolon= \big( \kappa^1\big(t,\boldsymbol{x},p^1,\boldsymbol{a}^{-1}\big), \dots, \kappa^N\big(t,\boldsymbol{x},p^N,\boldsymbol{a}^{-N}\big) \big),\\ \boldsymbol{\beta}(t,\boldsymbol{x},\boldsymbol{\alpha}) &\vcentcolon= \mathbf{\boldsymbol{\kappa}}\big(t,\boldsymbol{x},\mathrm{D}_{\boldsymbol{x}} \boldsymbol{V}(t,\boldsymbol{x},\boldsymbol{\alpha}),\boldsymbol{\alpha}(t,\boldsymbol{x})\big) = \big(\beta^1\big(t,\boldsymbol{x},\boldsymbol{\alpha}^{-1}\big),\dots,\beta^N\big(t,\boldsymbol{x},\boldsymbol{\alpha}^{-N}\big)\big). \end{align}\] For \(\boldsymbol{\kappa}\), we assume the following Lipschitz condition. There exist \(L_{\kappa}\) and \(L_{\kappa}^a \in [0,1)\) such that \[\begin{align} \begin{aligned} &\vert\!\vert\!\vert \boldsymbol{\kappa}(t_1, \boldsymbol{x}_1, \boldsymbol{p}_1, \boldsymbol{a}_1) - \boldsymbol{\kappa}(t_2, \boldsymbol{x}_2, \boldsymbol{p}_2, \boldsymbol{a}_2) \vert\!\vert\!\vert^2_{\mathbb{R}^{d \times N}} \\ &\qquad\leq L_{\kappa} \big( |t_1-t_2| + \|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n}^2 + \vert\!\vert\!\vert \boldsymbol{p}_1-\boldsymbol{p}_2 \vert\!\vert\!\vert_{\mathbb{R}^{n \times N}}^2 \big) + L_{\kappa}^a \vert\!\vert\!\vert \boldsymbol{a}_1-\boldsymbol{a}_2 \vert\!\vert\!\vert_{\mathbb{R}^{d\times N}}^2. \end{aligned} \end{align}\] Note that this implies Lipschitz continuity for \(\kappa^i\) of the form \[\begin{align} \begin{aligned} &\| \kappa^i(t_1, \boldsymbol{x}_1, p_1, \boldsymbol{a}^{-i}_1) - \kappa^i(t_2, \boldsymbol{x}_2, p_2, \boldsymbol{a}^{-i}_2) \|^2_{\mathbb{R}^{d}} \\ &\qquad\leq L_{\kappa} \big( |t_1-t_2| + \|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n}^2 + \|p_1-p_2\|_{\mathbb{R}^{n}}^2 \big) + L_{\kappa}^a \vert\!\vert\!\vert \boldsymbol{a}^{-i}_1-\boldsymbol{a}^{-i}_2 \vert\!\vert\!\vert_{\mathbb{R}^{d\times (N-1)}}^2. \end{aligned}\label{eq:kappai95lipschitz} \end{align}\tag{18}\] We remark that, for fixed \(\boldsymbol{\alpha}\in\mathbb{A}\), the joint policy \(\boldsymbol{\beta}(\cdot,\cdot,\boldsymbol{\alpha}) \in \mathbb{A}\) is in general not optimal for the game, since each player \(i\in\mathcal{I}\) optimizes against the fixed policies \(\boldsymbol{\alpha}^{-i}\). Instead, the game is solved by finding a fixed point \(\boldsymbol{\alpha}^*\) which, for all \((t,\boldsymbol{x}) \in Q_T\), satisfies \[\begin{align} \boldsymbol{\alpha}^*(t,\boldsymbol{x}) = \boldsymbol{\beta}(t,\boldsymbol{x}, \boldsymbol{\alpha}^*), \end{align}\] possibly in some subspace of \(\mathbb{A}\). For such a policy \(\boldsymbol{\alpha}^*\), we write with slight abuse of notation \[\begin{align} V^i(t,\boldsymbol{x}) \vcentcolon=V^i(t,\boldsymbol{x},\boldsymbol{\alpha}^{*,-i}), \qquad \boldsymbol{V}(t,\boldsymbol{x}) \vcentcolon=\boldsymbol{V}(t,\boldsymbol{x},\boldsymbol{\alpha}^*), \end{align}\] both defined on \(Q_T\). By the definition of \(\boldsymbol{\beta}\), it follows that if \(\boldsymbol{p}= \mathrm{D}_{\boldsymbol{x}} \boldsymbol{V}(t,\boldsymbol{x}) \in \mathbb{R}^{n\times N}\), then \(\boldsymbol{\alpha}^*\) satisfies \[\begin{align} \boldsymbol{\alpha}^*(t,\boldsymbol{x}) = \boldsymbol{\kappa} \big(t, \boldsymbol{x}, \boldsymbol{p}, \boldsymbol{\alpha}^*(t,\boldsymbol{x})\big), \label{eq:fp2} \end{align}\tag{19}\] for all \((t,\boldsymbol{x}) \in Q_T\). We next extend this relation to arbitrary \(\boldsymbol{p}\in \mathbb{R}^{n\times N}\). To this end, we assume that the generalized Isaacs conditions hold; see, e.g., [15], [16]. That is, for every \(\boldsymbol{p}\in \mathbb{R}^{n\times N}\), there exists a fixed point \(\boldsymbol{\alpha}^*_{\boldsymbol{p}} \in \mathbb{A}\) satisfying 19 , and the associated map \[\begin{align} \boldsymbol{c}(t, \boldsymbol{x}, \boldsymbol{p}) \vcentcolon=\boldsymbol{\alpha}^*_{\boldsymbol{p}}(t, \boldsymbol{x}) \label{eq:c95def} \end{align}\tag{20}\] is a well-defined function \(Q_T \times \mathbb{R}^{n \times N} \to A\). See Example 3 for one instance in which this holds. In addition, we assume that this fixed point is unique and that there exists \(L_{\boldsymbol{c}}\geq0\) such that for all \((t_1, \boldsymbol{x}_1, \boldsymbol{p}_1), (t_2, \boldsymbol{x}_2, \boldsymbol{p}_2) \in Q_T \times \mathbb{R}^{n \times N}\), we have \[\begin{align} \label{eq:lip7} \vert\!\vert\!\vert \boldsymbol{c}(t_1,\boldsymbol{x}_1,\boldsymbol{p}_1)-\boldsymbol{c}(t_2,\boldsymbol{x}_2,\boldsymbol{p}_2) \vert\!\vert\!\vert_{\mathbb{R}^{d\times N}}^2 \leq L_{\boldsymbol{c}} \big( |t_1-t_2| + \|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n}^2 + \vert\!\vert\!\vert \boldsymbol{p}_1 - \boldsymbol{p}_2 \vert\!\vert\!\vert_{\mathbb{R}^{n \times N}}^2 \big). \end{align}\tag{21}\]

The Hamiltonians associated with the Nash equilibrium are given by \[\begin{align} H^i(t,\boldsymbol{x},\boldsymbol{p}) = \big\langle \bar\boldsymbol{b}\big(t,\boldsymbol{x},\boldsymbol{c}(t,\boldsymbol{x},\boldsymbol{p})\big),p^i \big\rangle + \bar{f}^i\big(t,\boldsymbol{x},\boldsymbol{c}(t,\boldsymbol{x},\boldsymbol{p})\big) , \quad i\in\mathcal{I}. \end{align}\] This allows us to write the resulting system of HJB equations explicitly. We seek \(\boldsymbol{V}\colon Q_T\to \mathbb{R}^N\) such that, for all \((t,\boldsymbol{x})\in Q_T^0\), we have \[\begin{align} \label{eq:Nash95PDE95system1} \frac{\partial \boldsymbol{V}}{\partial t} (t,\boldsymbol{x}) + (\boldsymbol{\mathcal{L}}_t \boldsymbol{V})(t,\boldsymbol{x}) + \boldsymbol{H}\big(t,\boldsymbol{x},\mathrm{D}_{\boldsymbol{x}}\boldsymbol{V}(t,\boldsymbol{x})\big)=0, \end{align}\tag{22}\] and for all \((t,\boldsymbol{x})\in\partial Q_T\) the boundary condition \[\begin{align} \label{eq:Nash95PDE95system2} \boldsymbol{V}(t,\boldsymbol{x})=\boldsymbol{g}(t,\boldsymbol{x}). \end{align}\tag{23}\] Here, for \(\boldsymbol{\phi}=(\phi^1,\dots,\phi^N)\in C^2(Q_T,\mathbb{R}^N)\), \((t,\boldsymbol{x})\in Q_T\) and \(\boldsymbol{p}\in \mathbb{R}^{n\times N}\), we have \[\begin{align} (\boldsymbol{\mathcal{L}}_t \boldsymbol{\phi})(t,\boldsymbol{x}) \vcentcolon= \left[ \begin{array}{c} (\mathcal{L}_t\phi^1)(t,\boldsymbol{x})\\ \vdots\\ (\mathcal{L}_t\phi^N)(t,\boldsymbol{x}) \end{array} \right],\quad \boldsymbol{H} (t,\boldsymbol{x},\boldsymbol{p}) \vcentcolon= \left[ \begin{array}{c} H^1(t,\boldsymbol{x},\boldsymbol{p})\\ \vdots\\ H^N(t,\boldsymbol{x},\boldsymbol{p}) \end{array} \right], \end{align}\] and \(\boldsymbol{g}=(g^1,\dots,g^N)\colon \partial Q_T\to \mathbb{R}^N\). Since \(g^i\) is Lipschitz, it follows that \(\boldsymbol{g}\) is Lipschitz with constant \[\begin{align} L_{\boldsymbol{g}} \vcentcolon=\sum_{i \in \mathcal{I}} L_{g^i}. \end{align}\] The system 2223 is called the Nash PDE system. If a solution exists, then each player solves its corresponding HJB equation and is therefore optimal, provided the relevant verification theorem applies. This implies that the resulting solution leads to a Nash equilibrium \(\boldsymbol{\alpha}^*\). For the exposition in this section, we assume existence and uniqueness of a sufficiently regular solution to the Nash PDE system. In Section 3 we replace this by concrete assumptions ensuring these properties.

2.4 FBSDE formulation of the Nash system↩︎

We next introduce the FBSDE formulation of the Nash system. The optimal game dynamics are given by the family of predictable, square-integrable and up to modification unique stochastic processes \(\boldsymbol{X}^{t,\boldsymbol{x}}\colon[t,T]\times\Omega\to \mathbb{R}^n\), \((t,\boldsymbol{x})\in Q_T\), which satisfy, for all \(s\in[t,T]\), \(\mathbb{P}\)-almost surely \[\begin{align} \boldsymbol{X}_s^{t,\boldsymbol{x}} = \boldsymbol{x} + \int_t^{s\wedge \tau} \bar \boldsymbol{b}\big(r,\boldsymbol{X}_r^{t,\boldsymbol{x}},\boldsymbol{c} \big(r,\boldsymbol{X}_r^{t,\boldsymbol{x}},\mathrm{D}_{\boldsymbol{x}}\boldsymbol{V}\big(r,\boldsymbol{X}_r^{t,\boldsymbol{x}}\big)\big)\big) \,\text{d}r + \int_t^{s\wedge \tau} \boldsymbol{\Sigma}\big(r,\boldsymbol{X}_r^{t,\boldsymbol{x}}\big) \,\text{d}\boldsymbol{W}_r. \end{align}\] For \((t,\boldsymbol{x}) \in Q_T\), let \(\boldsymbol{Y}^{t,\boldsymbol{x}}\colon[t,T]\times\Omega\to \mathbb{R}^N\), \(\boldsymbol{Z}^{t,\boldsymbol{x}}\colon[t,T]\times\Omega\to \mathbb{R}^{n\times N}\) be the family of stochastic processes which, for \(s\in[t,T]\), \(\mathbb{P}\)-almost surely are given by \[\label{eq:YZ1} \begin{align} \boldsymbol{Y}_s^{t,\boldsymbol{x}} &= \boldsymbol{V}\big(s,\boldsymbol{X}_s^{t,\boldsymbol{x}}\big),\\ \boldsymbol{Z}_s^{t,\boldsymbol{x}} &= \boldsymbol{\Sigma}^\top\big(s,\boldsymbol{X}_s^{t,\boldsymbol{x}}\big)\mathrm{D}_{\boldsymbol{x}}\boldsymbol{V}\big(s,\boldsymbol{X}_s^{t,\boldsymbol{x}}\big). \end{align}\tag{24}\] Assuming \(\boldsymbol{V}\) is sufficiently regular, Itô’s formula and the HJB equations give \(\mathbb{P}\)-almost surely \[\begin{align} \label{eq:unif95par} \boldsymbol{Y}_s^{t,\boldsymbol{x}} = \boldsymbol{g}\big(\tau,\boldsymbol{X}_\tau^{t,\boldsymbol{x}}\big) + \int_{s\wedge \tau}^\tau \bar\boldsymbol{f} \big(r,\boldsymbol{X}_r^{t,\boldsymbol{x}},\boldsymbol{c}\big(r,\boldsymbol{X}_r^{t,\boldsymbol{x}},\mathrm{D}_{\boldsymbol{x}}\boldsymbol{V}\big(r,\boldsymbol{X}_r^{t,\boldsymbol{x}}\big)\big)\big) \, \text{d}r - \int_{s\wedge \tau}^\tau \big(\boldsymbol{Z}_r^{t,\boldsymbol{x}}\big)^\top \text{d}\boldsymbol{W}_r, \end{align}\tag{25}\] where \(\bar\boldsymbol{f}\vcentcolon=(\bar f^1,\dots,\bar f^N)\colon[0,T]\times\mathbb{R}^n\times A\to \mathbb{R}^N\). Since \(\bar{f}^i\) is Lipschitz, \(\bar \boldsymbol{f}\) is Lipschitz with constant \[\begin{align} L_{\bar \boldsymbol{f}} \vcentcolon=\sum_{i \in \mathcal{I}} L_{\bar{f}^i}. \end{align}\] By uniform parabolicity, the map \(\boldsymbol{p}\mapsto \boldsymbol{\Sigma}^\top(t,\boldsymbol{x})\boldsymbol{p}\) is injective, with inverse \(\boldsymbol{z}\mapsto \boldsymbol{\Phi}(t,\boldsymbol{x}) \boldsymbol{z}\). This allows us to define \(\boldsymbol{b}\colon Q_T\times \mathbb{R}^{n \times N}\to\mathbb{R}^n\) and \(\boldsymbol{f}\colon Q_T\times \mathbb{R}^{n \times N}\to \mathbb{R}^N\) by \[\begin{align} \boldsymbol{b}(t,\boldsymbol{x},\boldsymbol{z}) &\vcentcolon= \bar \boldsymbol{b}\big(t,\boldsymbol{x},\boldsymbol{c}(t,\boldsymbol{x},\boldsymbol{\Phi}(t,\boldsymbol{x})\boldsymbol{z})\big),\\ \boldsymbol{f}(t,\boldsymbol{x},\boldsymbol{z}) &\vcentcolon= \bar \boldsymbol{f}\big(t,\boldsymbol{x},\boldsymbol{c}(t,\boldsymbol{x},\boldsymbol{\Phi}(t,\boldsymbol{x})\boldsymbol{z})\big), \end{align}\] and write the equations for \(\boldsymbol{X}^{t,\boldsymbol{x}}\), \(\boldsymbol{Y}^{t,\boldsymbol{x}}\) and \(\boldsymbol{Z}^{t,\boldsymbol{x}}\) as a coupled FBSDE, i.e., for all \((t,\boldsymbol{x})\in Q_T\), \(s\in[t,T]\), it holds \(\mathbb{P}\)-almost surely \[\begin{align} \label{eq:FBSDE95full} \begin{aligned} \boldsymbol{X}_s^{t,\boldsymbol{x}}&=\boldsymbol{x}+ \int_t^{s\wedge \tau}\boldsymbol{b}\big(r,\boldsymbol{X}_r^{t,\boldsymbol{x}},\boldsymbol{Z}_r^{t,\boldsymbol{x}}\big)\,\text{d}r + \int_t^{s\wedge \tau}\boldsymbol{\Sigma}\big(r,\boldsymbol{X}_r^{t,\boldsymbol{x}}\big)\,\text{d}\boldsymbol{W}_r,\\ \boldsymbol{Y}_s^{t,\boldsymbol{x}} &=\boldsymbol{g}\big(\tau,\boldsymbol{X}_\tau^{t,\boldsymbol{x}}\big) +\int_{s\wedge \tau}^\tau\boldsymbol{f}\big(r,\boldsymbol{X}_r^{t,\boldsymbol{x}},\boldsymbol{Z}_r^{t,\boldsymbol{x}}\big)\,\text{d}r - \int_{s\wedge \tau}^\tau\big(\boldsymbol{Z}_r^{t,\boldsymbol{x}}\big)^\top\,\text{d}\boldsymbol{W}_r. \end{aligned} \end{align}\tag{26}\] We refer to 26 as the Nash FBSDE system. Our minimum assumptions on \(\boldsymbol{f}\) and \(\boldsymbol{g}\) are that, for all \((t,\boldsymbol{x})\in Q_T\), there exists a unique solution \[\begin{align} \big(\boldsymbol{X}^{t,\boldsymbol{x}},\boldsymbol{Y}^{t,\boldsymbol{x}},\boldsymbol{Z}^{t,\boldsymbol{x}}\big)\in \mathbb{S}^{2,n}_{t,T}\times \mathbb{S}^{2,N}_{t,T} \times \mathbb{H}^{2,n,N}_{t,T}, \end{align}\] where \(\mathbb{S}^{2,n}_{t,T}\), \(\mathbb{S}^{2,N}_{t,T}\) and \(\mathbb{H}^{2,n,N}_{t,T}\) are defined analogously to \(\mathbb{S}^{2,n}_{T}, \mathbb{S}^{2,N}_{T}\) and \(\mathbb{H}^{2,n,N}_{T}\) for processes defined on the interval \([t,T]\). In addition, for every \(p\in[2,\infty)\), we assume that \(\boldsymbol{X}^{t,\boldsymbol{x}}\in \mathbb{S}^{p,n}_{t,T}\). In Section 3, we impose concrete assumptions that guarantee these properties.

2.5 Additional notation for non-equilibrium play↩︎

For later use, we introduce the control problem for player \(i\in\mathcal{I}\) when the policies of the other players are fixed. For this purpose, define \(\tilde{\boldsymbol{b}}^i\colon Q_T\times \mathbb{R}^n\times A^{-i}\to\mathbb{R}^n\) and \(\tilde{\ell}^i\colon Q_T\times \mathbb{R}^n\times A^{-i}\to\mathbb{R}\) by \[\label{xfguvenr} \begin{align} \tilde{\boldsymbol{b}}^i\big(t,\boldsymbol{x},z,\boldsymbol{a}^{-i}\big) &\vcentcolon= \bar \boldsymbol{b}\big(t,\boldsymbol{x},\big[\kappa^i(t,\boldsymbol{x},\boldsymbol{\Phi}(t,\boldsymbol{x})z,\boldsymbol{a}^{-i}),\boldsymbol{a}^{-i}\big]_i\big),\\ \tilde{\ell}^i\big(t,\boldsymbol{x},z,\boldsymbol{a}^{-i}\big) &\vcentcolon= \bar f^i\big(t,\boldsymbol{x},\big[\kappa^i(t,\boldsymbol{x},\boldsymbol{\Phi}(t,\boldsymbol{x})z,\boldsymbol{a}^{-i}),\boldsymbol{a}^{-i}\big]_i\big). \end{align}\tag{27}\] Given any admissible control \(\boldsymbol{\beta}^{-i}\in\mathbb{A}^{-i}\) for the other \(N-1\) players, player \(i\in\mathcal{I}\) solves the FBSDE \[\begin{align} \label{qcuhyzdw} \begin{aligned} \boldsymbol{X}_t^{i,\boldsymbol{\beta}^{-i}} &= \boldsymbol{x}_0 + \int_0^{t\wedge \tau} \tilde{\boldsymbol{b}}^i\big( s,\boldsymbol{X}_s^{i,\boldsymbol{\beta}^{-i}}, Z_s^{i,\boldsymbol{\beta}^{-i}}, \boldsymbol{\beta}^{-i}_s \big) \,\text{d}s + \int_0^{t\wedge \tau} \boldsymbol{\Sigma}\big(s,\boldsymbol{X}_s^{i,\boldsymbol{\beta}^{-i}}\big)\,\text{d}\boldsymbol{W}_s, \\ Y_t^{i,\boldsymbol{\beta}^{-i}} &= g^i\big(\tau,\boldsymbol{X}_\tau^{i,\boldsymbol{\beta}^{-i}}\big) + \int_{t\wedge \tau}^{\tau} \tilde{\ell}^i\big( s,\boldsymbol{X}_s^{i,\boldsymbol{\beta}^{-i}}, Z_s^{i,\boldsymbol{\beta}^{-i}}, \boldsymbol{\beta}^{-i}_s \big) \,\text{d}s - \int_{t\wedge \tau}^{\tau} \big(Z_s^{i,\boldsymbol{\beta}^{-i}}\big)^\top\,\text{d}\boldsymbol{W}_s. \end{aligned} \end{align}\tag{28}\] Since the state variable is shared among all players, we use boldface notation for \(\boldsymbol{X}^{i,\boldsymbol{\beta}^{-i}}\). In contrast, we use non-bold notation for \(Y^{i,\boldsymbol{\beta}^{-i}}\) and \(Z^{i,\boldsymbol{\beta}^{-i}}\), since these are player-specific value and control processes. However, the processes \(\boldsymbol{X}^{1,\boldsymbol{\beta}^{-1}},\dots,\boldsymbol{X}^{N,\boldsymbol{\beta}^{-N}}\) are generally different, since each player optimizes its own control given \(\boldsymbol{\beta}^{-i}\).

2.6 Examples of games↩︎

This section presents a sequence of examples with progressively stronger assumptions. We begin with a setting based only on smoothness, growth, and boundedness conditions, which yields rather restrictive results. We then introduce additional structure on the domain, dynamics, and costs, which leads to a linear–quadratic class for which the Nash system can be reduced to a coupled Riccati system.

Example 1. Let \(\mathcal{D}=\mathbb{R}^n\), so that \(\partial Q_T=\{T\}\times\mathbb{R}^n\) and \(\mathbb{P}(\tau=T)=1\). Assume \(A\) to be compact and the functions \(\bar \boldsymbol{b}\), \(\boldsymbol{\Sigma}\), \(\bar{f}^i\), and \(g^i\), \(i\in\mathcal{I}\), to be bounded. Then, under some additional assumptions, [15] guarantees existence and uniqueness of a solution to the Nash system 2223 . The boundedness is a very strong condition one has to pay for not imposing more structure.

Example 2. We consider a bounded domain \(\mathcal{D}\) with boundary \(\partial \mathcal{D}\) of class \(C^2\), and assume that \(A\) is compact. In addition to our standing assumptions, assume that \(\boldsymbol{\Sigma}\) is bounded and continuously differentiable in time and space with bounded derivatives, and that \(\boldsymbol{g}\) is twice differentiable in the sense required in  [16]. Then [16] guarantees the existence of a (nonclassical) solution to the Nash system 2223 . These conditions are considerably less restrictive than in the unbounded case \(\mathcal{D}=\mathbb{R}^n\). For further refined results, see [31].

Example 3. Here, we consider a subclass of games satisfying the setting of Example 2 and thus have a unique solution to the Nash system. The drift coefficient has the form \[\begin{align} \bar \boldsymbol{b}(t,\boldsymbol{x},\boldsymbol{a}) = \boldsymbol{b}_1(t,\boldsymbol{x}) + \sum_{i\in\mathcal{I}} b_2^i(t,\boldsymbol{x})a^i, \end{align}\] where \(\boldsymbol{b}_1\colon Q_T\to\mathbb{R}^n\) and \(b_2^i\colon Q_T\to \mathbb{R}^{n\times d}\), \(i\in\mathcal{I}\), satisfy for all \((t_1,\boldsymbol{x}_1),(t_2,\boldsymbol{x}_2)\in Q_T\) the bounds \[\begin{align} \big\|\boldsymbol{b}_1(t_1,\boldsymbol{x}_1)-\boldsymbol{b}_1(t_2,\boldsymbol{x}_2)\big\|_{\mathbb{R}^{n}} &\leq C \big( |t_1-t_2|^{\frac{1}{2}} + \|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n} \big),\\ \vert\!\vert\!\vert b_2^i(t_1,\boldsymbol{x}_1)-b_2^i(t_2,\boldsymbol{x}_2) \vert\!\vert\!\vert_{\mathbb{R}^{n\times d}} &\leq C \big( |t_1-t_2|^{\frac{1}{2}} + \|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n} \big). \end{align}\] In addition, \(b_2^i\) is bounded on \(Q_T\) for each \(i \in \mathcal{I}\). Under these assumptions \(\bar\boldsymbol{b}\) satisfies 9 . For each \(i \in \mathcal{I}\), we consider the running cost \[\begin{align} \bar{f}^i(t,\boldsymbol{x},\boldsymbol{a}) = f_1^i(t,\boldsymbol{x}) + \frac{1}{2} \sum_{j\in\mathcal{I}} \big\langle K_a^{i,j} a^j, a^j \big\rangle + \sum_{j\in\mathcal{I}} \big\langle K_{a\boldsymbol{x}}^{i,j} a^j, \boldsymbol{x} \big\rangle. \end{align}\] Here, \(f_1^i\colon Q_T \to \mathbb{R}\) determines the state-based cost, whereas \(K_a^{i,j}\in\mathbb{R}^{d\times d}\), \(K_{a\boldsymbol{x}}^{i,j}\in\mathbb{R}^{n\times d}\), \(i,j\in\mathcal{I}\), penalize control and control-state relation, respectively. We assume \(f_1^i\) to have polynomial growth and \(K_a^{i,i}\) to be positive definite and invertible. With these data, the related Hamiltonian for player \(i\) reads \[\begin{align} H^i(t,\boldsymbol{x},p,\boldsymbol{a}^{-i}) &= \bigg\langle \boldsymbol{b}_1(t,\boldsymbol{x}) + \sum_{j\in\mathcal{I}\setminus\{i\}} b_2^j(t,\boldsymbol{x})a^j ,p \bigg\rangle + f_1^i(t,\boldsymbol{x}) + \frac{1}{2} \sum_{j\in\mathcal{I}\setminus\{i\}} \big\langle K_a^{i,j} a^j, a^j \big\rangle\\ &\quad+ \sum_{j\in\mathcal{I}\setminus\{i\}} \big\langle K_{a\boldsymbol{x}}^{i,j} a^j,\boldsymbol{x} \big\rangle + \inf_{a^i\in A^{i}} \bigg( \big\langle b_2^i(t,\boldsymbol{x})a^i,p \big\rangle + \frac{1}{2} \big\langle K_a^{i,i} a^i, a^i \big\rangle + \big\langle K_{a\boldsymbol{x}}^{i,i} a^i, \boldsymbol{x} \big\rangle \bigg). \end{align}\] The minimization problem is quadratic and when the minimum is attained in the interior of \(A^i\), its solution is explicit. This yields the optimum \[\begin{align} \kappa^i \big( t,\boldsymbol{x},p,\boldsymbol{a}^{-i} \big) = \Phi^i(t,\boldsymbol{x},p) + \mathcal{E}^i_A(t,\boldsymbol{x},p), \end{align}\] where \(\Phi^i\colon Q_T\times \mathbb{R}^n\to\mathbb{R}^d\), \(i\in\mathcal{I}\) is given by \[\begin{align} \Phi^i(t,\boldsymbol{x},p) \vcentcolon= -\big( K_{a}^{i,i} \big)^{-1} \Big( b_2^i(t,\boldsymbol{x})^\top p + \big( K_{a\boldsymbol{x}}^{i,i} \big)^\top \boldsymbol{x} \Big), \end{align}\] and the function \(\mathcal{E}^i_A\) corrects \(\kappa^i\) when the minimum is attained on the boundary of \(A^i\). Since \(\kappa^i\) and \(\boldsymbol{\kappa}\) do not depend on \(\boldsymbol{a}\), the existence of the function \(\boldsymbol{c}\) through the fixed-point problem 19 is trivial. For the same reason, the Nash system is coupled only through \(\bar\boldsymbol{b}\) and \(\bar{f}^i\) and not through \(\boldsymbol{\kappa}\). This leads to \(V^i(t,\boldsymbol{x},\boldsymbol{a}^{-i})=V^i(t,\boldsymbol{x})\) for all \(i\in\mathcal{I}\), \(\boldsymbol{a}^{-i}\in A^{-i}\) and \((t,\boldsymbol{x})\in Q_T^0\). Inserting the gradient, we obtain the resulting Nash system: find \(\boldsymbol{V}=(V^1,\dots,V^N)\colon Q_T\to\mathbb{R}^N\) that, for all \((t,\boldsymbol{x})\in Q_T^0\) and \(i\in\mathcal{I}\), satisfies \[\begin{align} -\frac{\partial V^i}{\partial t} &= \frac{1}{2} \mathrm{Tr}\big(\boldsymbol{\Sigma}\, \mathrm{D}_{\boldsymbol{x}}^2 V^i\, \boldsymbol{\Sigma}^\top\big) + \langle \boldsymbol{b}_1,\mathrm{D}_{\boldsymbol{x}}V^i \rangle + \sum_{j\in\mathcal{I}} \big\langle \Phi^{j}(\mathrm{D}_{\boldsymbol{x}} V^j) , \big(b_2^j\big)^\top \mathrm{D}_{\boldsymbol{x}}V^i \big\rangle + f_1^i\\ &\quad + \frac{1}{2} \sum_{j\in\mathcal{I}} \big\langle K_{a}^{i,j} \Phi^{j}(\mathrm{D}_{\boldsymbol{x}}V^j), \Phi^{j}(\mathrm{D}_{\boldsymbol{x}}V^j) \big\rangle + \sum_{j\in\mathcal{I}} \big\langle K_{a\boldsymbol{x}}^{i,j} \Phi^{j}(\mathrm{D}_{\boldsymbol{x}}V^j),\boldsymbol{x} \big\rangle + \boldsymbol{\Xi}^i(\mathrm{D}_{\boldsymbol{x}}\boldsymbol{V}), \end{align}\] and for all \((t,\boldsymbol{x})\in \partial Q_T\) the boundary condition \(V^i(t,\boldsymbol{x}) = g^i(t,\boldsymbol{x})\) holds. Here, \(\boldsymbol{\Xi}^i\) corrects for replacing the optimal control \(\kappa^i(t,\boldsymbol{x},\mathrm{D}_{\boldsymbol{x}}\boldsymbol{V}(t,\boldsymbol{x}))\), restricted to \(A^i\), in the explicit terms of the equation, with the unrestricted control \(\Phi^i(t,\boldsymbol{x},\mathrm{D}_{\boldsymbol{x}}\boldsymbol{V}(t,\boldsymbol{x}))\). This in particular reveals the structure of the unconstrained problem with \(A=\mathbb{R}^{d\times N}\) or the main term when \(A\) is compact but large.

Example 4 (Linear–quadratic games). We consider the setting of Example 3, with the modifications that \(\mathcal{D}=\mathbb{R}^n\) and \(A=(\mathbb{R}^d)^N\). Here, the non-compact set \(A\) is not covered by our assumptions. We nevertheless include this example, since it is a central case, admits a semi-explicit solution, and is used in our numerical experiment. We further specify \[\begin{align} \boldsymbol{b}_1(t,\boldsymbol{x}) &= F(\boldsymbol{\xi}-\boldsymbol{x}), \quad b_2^i(t,\boldsymbol{x}) = B^i, \quad \boldsymbol{\Sigma}(t,\boldsymbol{x}) = \mathbf{C}, \quad f_1^i(t,\boldsymbol{x}) = \big\| K^{i}_{\boldsymbol{x}}\boldsymbol{x}\big\|_{\mathbb{R}^n}^2, \quad g^i(\boldsymbol{x}) = \big\|G^i\boldsymbol{x}\big\|_{\mathbb{R}^n}^2, \end{align}\] for \(i \in \mathcal{I}\), where \(F\in\mathbb{R}^{n\times n}\), \(\boldsymbol{\xi}\in\mathbb{R}^n\), \(\mathbf{C}\in\mathbb{R}^{n\times n}\), \(B^i\in\mathbb{R}^{n\times d}\), \(K^{i}_{\boldsymbol{x}}\in\mathbb{R}^{n\times n}\), and \(G^i\in\mathbb{R}^{n\times n}\). The Nash system reads: find \(\boldsymbol{V}=(V^1,\dots,V^N)\colon Q_T\to\mathbb{R}^N\) that, for all \((t,\boldsymbol{x})\in Q_T^0\) and \(i\in\mathcal{I}\), satisfies \[\begin{align} \begin{aligned} -\frac{\partial V^i}{\partial t} &= \frac{1}{2} \mathrm{Tr}\big(\mathbf{C}\, \mathrm{D}_{\boldsymbol{x}}^2 V^i\, \mathbf{C}^\top\big) + \big\langle F(\boldsymbol{\xi}-\boldsymbol{x}),\mathrm{D}_{\boldsymbol{x}}V^i \big\rangle + \sum_{j\in \mathcal{I}} \big\langle \tilde{\Phi}^{j}, \big( B^j \big)^\top \mathrm{D}_{\boldsymbol{x}} V^i \big\rangle \\ &\qquad + \big\|K^{i}_{\boldsymbol{x}}\boldsymbol{x}\big\|_{\mathbb{R}^n}^2 + \frac{1}{2} \sum_{j\in\mathcal{I}} \big\langle K_{a}^{i,j} \tilde{\Phi}^{j}, \tilde{\Phi}^{j} \big\rangle + \sum_{j\in \mathcal{I}} \big\langle K_{a\boldsymbol{x}}^{i,j} \tilde{\Phi}^{j}, \boldsymbol{x} \big\rangle, \end{aligned} \label{eq:HJB95equation95LQG95example} \end{align}\tag{29}\] and for all \((t,\boldsymbol{x})\in\partial Q_T\) the boundary condition \(V^i(t,\boldsymbol{x}) = \|G^i\boldsymbol{x}\|_{\mathbb{R}^n}^2\). Here, \(\tilde{\Phi}^{i}\) is the straightforward modification of \(\Phi^i\) given by \[\begin{align} \tilde{\Phi}^{i} = - \big( K_{a}^{i,i} \big)^{-1} \Big( \big( B^i \big)^\top \mathrm{D}_{\boldsymbol{x}} V^i + \big( K_{a\boldsymbol{x}}^{i,i} \big)^\top \boldsymbol{x} \Big). \label{eq:Phii95Ex} \end{align}\tag{30}\] The system 29 admits a semi-analytical solution of the form \[\begin{align} \label{eq:LQ95Vi} V^i(t,\boldsymbol{x}) = \frac{1}{2} \left\langle \boldsymbol{x}, P^i_t \boldsymbol{x} \right\rangle + \left\langle Q^i_t, \boldsymbol{x} \right\rangle + R^i_t, \end{align}\tag{31}\] where \(P^i \colon [0,T]\to \mathbb{R}^{n\times n}\), \(Q^i \colon[0,T]\to\mathbb{R}^n\) and \(R^i\colon[0,T]\to\mathbb{R}\), \(i\in\mathcal{I}\), solve the coupled Riccati system \[\begin{align} \begin{aligned}\label{eq:Riccati95system} \dot{P}^i &= F^\top P^i + P^i F - \big( K^{i}_{\boldsymbol{x}} \big)^\top K^{i}_{\boldsymbol{x}} - 2 \sum_{j\in \mathcal{I}} \big( \tilde{P}^i \big( \Theta_1^{j} \tilde{P}^j + \Theta_2^{j} \big) + \tilde{P}^j \Theta_3^{i,j} \tilde{P}^j + \Theta_{4,5,7}^{i,j} \tilde{P}^j + \Theta_{6,8}^{i,j} \big), \\ \dot{Q}^i &= - \tilde{P}^i F \boldsymbol{\xi} + F^\top Q^i - \sum_{j\in \mathcal{I}} \big( \tilde{P}^i \Theta_1^{j} Q^j + \tilde{P}^j \big( \Theta_1^{j} \big)^\top Q^i + \big( \Theta_2^{j} \big)^\top Q^i + \tilde{P}^j \tilde{\Theta}_3^{i,j} Q^j + \Theta_{4,5,7}^{i,j} Q^j \big), \\ \dot{R}^i &= - \frac{1}{2} \mathrm{Tr} \big( \mathbf{C} \tilde{P}^i \mathbf{C}^\top \big) - \left\langle F \boldsymbol{\xi}, Q^i \right\rangle - \sum_{j\in \mathcal{I}} \big( \big\langle Q^i, \Theta_1^{j} Q^j \big\rangle + \big\langle Q^j, \Theta_3^{i,j} Q^j \big\rangle \big). \end{aligned} \end{align}\tag{32}\] Here, the notation \(\tilde{P}^i = \tfrac{1}{2} \big( P^i + (P^i)^\top\big)\) has been used for brevity, as well as the matrices \[\begin{align} \Theta_1^{j} &= - B^j \left( K^{j,j}_{a} \right)^{-1} \left( B^j \right)^\top , \\ \Theta_2^{j} &= - B^j \left( K^{j,j}_{a} \right)^{-1} \left( K^{j,j}_{a\boldsymbol{x}} \right)^\top, \\ \Theta_3^{i,j} &= \tfrac{1}{2} B^j \left( K^{j,j}_{a} \right)^{-\top} K^{i,j}_{a} \left( K^{j,j}_{a} \right)^{-1} \left( B^j \right)^\top, \\ \tilde{\Theta}_3^{i,j} &= \tfrac{1}{2} B^j \left( K^{j,j}_{a} \right)^{-\top} \Big( \left( K^{i,j}_{a} \right)^\top + K^{i,j}_{a} \Big) \left( K^{j,j}_{a} \right)^{-1} \left( B^j \right)^\top, \\ \Theta_{4,5,7}^{i,j} &= \Big( \tfrac{1}{2} K^{j,j}_{a\boldsymbol{x}} \big( K^{j,j}_{a} \big)^{-\top} \Big( \big( K^{i,j}_{a} \big)^\top + K^{i,j}_{a} \Big) - K^{i,j}_{a\boldsymbol{x}} \Big) \left( K^{j,j}_{a} \right)^{-1} \left( B^j \right)^\top, \\ \Theta_{6,8}^{i,j} &= \Big( \tfrac{1}{2} K^{j,j}_{a\boldsymbol{x}} \left( K^{j,j}_{a} \right)^{-\top} K^{i,j}_{a} - K^{i,j}_{a\boldsymbol{x}} \Big) \left( K^{j,j}_{a} \right)^{-1} \left( K^{j,j}_{a\boldsymbol{x}} \right)^\top. \end{align}\]

Remark 1. For FBSDEs with one-dimensional backward process, i.e., for stochastic control, [32] develops a theory of (probabilistically) weak solutions to the corresponding FBSDE. Under general and for applications very natural assumptions, existence and uniqueness of weak solutions are guaranteed, and through the FBSDE also a solution to the HJB equation. It would be desirable with an analogous theory for the Nash FBSDE system, but, to the best of our knowledge, it has not been established.

3 Fictitious play and its convergence↩︎

System 26 is a coupled FBSDE with an \(N\)-dimensional BSDE component, leading to significant computational complexity. To address this, we adopt the concept of fictitious play, originally introduced for zero-sum games in [33] and further developed in [25], [34]. The approach has since been extended to stochastic differential games (see, e.g., [14]), and has recently gained renewed attention in related computational approaches [22], [23].

Fictitious play is an iterative best-response scheme: at each iteration, every player updates their strategy by solving an optimal control problem assuming the opponents’ strategies are fixed. In this stochastic differential game setting, fixing the strategies of the other players reduces each player’s problem to a single-agent stochastic control problem. The control problem admits an FBSDE representation, whose solution determines the player’s optimal feedback response. Repeating this procedure generates a sequence of strategies that, under suitable conditions, converges to a Nash equilibrium.

Throughout this section, we fix the initial value \(\boldsymbol{x}_0\in \mathcal{D}\). Recalling the family \((\boldsymbol{X}^{t,\boldsymbol{x}},\boldsymbol{Y}^{t,\boldsymbol{x}},\boldsymbol{Z}^{t,\boldsymbol{x}})\) from 26 , our aim is to approximate \((\boldsymbol{X},\boldsymbol{Y},\boldsymbol{Z}):=(\boldsymbol{X}^{0,\boldsymbol{x}_0},\boldsymbol{Y}^{0,\boldsymbol{x}_0},\boldsymbol{Z}^{0,\boldsymbol{x}_0})\). The fictitious-play iteration starts from the zero Markov policy \(\alpha^{i,0} = 0 \in \mathbb{A}^i\) for each player \(i \in \mathcal{I}\). For \(m \geq 1\), each player \(i \in \mathcal{I}\) solves [eq:FBSDE95single95player] with \(\boldsymbol{\beta}^{-i} = \boldsymbol{\alpha}^{-i,m-1}\). This yields a non-equilibrium Markov map \(\zeta^{i,m}\colon [0,T]\times\mathbb{R}^n\to\mathbb{R}^n\), and the policy is updated by \[\begin{align} \label{eq:FP95construction} \alpha^{i,m}(t, \boldsymbol{x}) := \kappa^i \big( t,\boldsymbol{x},\boldsymbol{\Phi}(t,\boldsymbol{x})\zeta^{i,m},\boldsymbol{\alpha}^{-i,m-1}(t,\boldsymbol{x}) \big). \end{align}\tag{33}\] The iterative scheme leads to FBSDEs whose solution triples are denoted by \(\big(\boldsymbol{X}^{i,m},\boldsymbol{Y}^{i,m},\boldsymbol{Z}^{i,m}\big)\).

The remainder of this section is devoted to the convergence analysis of the fictitious-play scheme. To quantify the error at iteration \(m\), we introduce \[\begin{align} \delta \boldsymbol{X}_t^{i,m} \vcentcolon=\boldsymbol{X}_t-\boldsymbol{X}_t^{i,m}, \qquad \delta Y_t^{i,m} \vcentcolon=Y_t^i-Y_t^{i,m}. \label{eq:error95quantities} \end{align}\tag{34}\] In Section 3.1, a reformulation of the fictitious-play FBSDE is presented. Then, under the assumptions stated in Section 3.2, we show in Section 3.3 that the error quantities in 34 , together with the associated feedback controls, converge to zero at a geometric rate as \(m \to \infty\).

3.1 A reformulation of the fictitious play FBSDEs↩︎

The coefficients \(\tilde{\boldsymbol{b}}^i\) and \(\tilde{\ell}^i\) in [eq:FBSDE95single95player] depend on arbitrary policies \(\boldsymbol{\beta}^{-i}\in\mathbb{A}^{-i}\). For the convergence analysis, we reformulate the corresponding FBSDEs in terms of the Markov maps generated by the fictitious-play iteration. For each \(i \in \mathcal{I}\) and \(m \geq 1\), we introduce functions \(\boldsymbol{b}^{i,m}\colon Q_T\times \mathbb{R}^n\times\mathbb{R}^{(N-1)\times n}\to\mathbb{R}^n\) and \(\ell^{i,m}\colon Q_T\times \mathbb{R}^n\times\mathbb{R}^{n\times (N-1)}\to\mathbb{R}\) such that, at iteration \(m\), player \(i\) solves the FBSDE \[\begin{align} \label{eq:FP95FBSDE} \begin{aligned} \boldsymbol{X}_t^{i,m}\! &=\! \boldsymbol{x}_0 + \int_0^{t\wedge\tau^{i,m}} \boldsymbol{b}^{i,m}\big( s,\boldsymbol{X}_s^{i,m}, Z_s^{i,m}, \boldsymbol{\zeta}^{-i,m-1}\big(s,\boldsymbol{X}_s^{i,m}\big) \big)\,\text{d}s + \int_0^{t\wedge\tau^{i,m}} \boldsymbol{\Sigma}\big(s,\boldsymbol{X}_s^{i,m}\big)\,\text{d}\boldsymbol{W}_s, \\ Y_t^{i,m}\! &=\! g^{i}\big(\tau^{i,m},\boldsymbol{X}_{\tau^{i,m}}^{i,m}\big) \!+\! \int_{t\wedge\tau^{i,m}}^{\tau^{i,m}}\! \ell^{i,m}\big( s,\boldsymbol{X}_s^{i,m}, Z_s^{i,m}, \boldsymbol{\zeta}^{-i,m-1}(s,\boldsymbol{X}_s^{i,m}) \big)\,\text{d}s \!-\! \int_{t\wedge\tau^{i,m}}^{\tau^{i,m}}\! \big(Z_s^{i,m}\big)^\top \text{d}\boldsymbol{W}_s, \end{aligned} \end{align}\tag{35}\] where \(\zeta^{i,m} : [0,T] \times \mathbb{R}^n \to \mathbb{R}^n\) satisfies for almost all \(t\in[0,T]\), \(\mathbb{P}\)-almost surely \[\begin{align} Z_t^{i,m}=\zeta^{i,m}\big(t,\boldsymbol{X}_t^{i,m}\big) \label{eq:Z95FP}. \end{align}\tag{36}\] We now identify the coefficients in this reformulation. For the moment, assume that the Markov maps \(\zeta^{i,m}\) exist. Sufficient assumptions for this are stated in Section 3.2. For \(i \in \mathcal{I}\), set \(\kappa^{i,0} \equiv 0\) and define for \(m \geq 1\) the functions \(\kappa^{i,m}\colon Q_T\times\mathbb{R}^n\to A^i\) by \[\begin{align} \kappa^{i,m}\big(t,\boldsymbol{x},z\big) \vcentcolon= \kappa^i \big( t, \boldsymbol{x}, \boldsymbol{\Phi}(t,\boldsymbol{x})z, \boldsymbol{\kappa}^{-i,m-1}\big( t,\boldsymbol{x},\boldsymbol{\zeta}^{-i,m-1}(t,\boldsymbol{x}) \big) \big), \end{align}\] where for \((t,\boldsymbol{x},\boldsymbol{z}^{-i})\in Q_T\times\mathbb{R}^{(N-1)\times n}\), \[\begin{align} &\boldsymbol{\kappa}^{-i,m}\big(t,\boldsymbol{x},\boldsymbol{z}^{-i}\big) \vcentcolon= \big( \kappa^{1,m}\big(t,\boldsymbol{x},z^1\big), \ldots, \kappa^{i-1,m}\big(t,\boldsymbol{x},z^{i-1}\big), \kappa^{i+1,m}\big(t,\boldsymbol{x},z^{i+1}\big), \ldots, \kappa^{N,m}\big(t,\boldsymbol{x},z^N\big) \big). \end{align}\] When necessary, we interpret \(\zeta^{i,0} \equiv 0\). We claim that the coefficients \(\boldsymbol{b}^{i,m}\) and \(\ell^{i,m}\) are given by \[\begin{align} \boldsymbol{b}^{i,m}\big(t,\boldsymbol{x},z,\boldsymbol{z}^{-i}\big) &\vcentcolon= \tilde{\boldsymbol{b}}^i\big( t,\boldsymbol{x}, z, \boldsymbol{\kappa}^{-i,m-1}\big( t,\boldsymbol{x},\boldsymbol{z}^{-i} \big) \big), \\ \ell^{i,m}\big(t,\boldsymbol{x},z,\boldsymbol{z}^{-i}\big) &\vcentcolon= \tilde{\ell}^i\big( t,\boldsymbol{x}, z, \boldsymbol{\kappa}^{-i,m-1}\big( t,\boldsymbol{x},\boldsymbol{z}^{-i} \big) \big). \end{align}\] To prove the claim, it suffices to identify the opponent policy appearing in \(\tilde{\boldsymbol{b}}^i\) and \(\tilde{\ell}^i\) at iteration \(m\). By construction, this policy is \(\boldsymbol{\alpha}^{-i,m-1}\). Thus, the claim follows once we show for every \(r\geq 0\) that \[\begin{align} \label{eq:induction95proof95relation95first} \boldsymbol{\alpha}^{-i,r}(t,\boldsymbol{x}) = \boldsymbol{\kappa}^{-i,r}\big(t,\boldsymbol{x},\boldsymbol{\zeta}^{-i,r}(t,\boldsymbol{x})\big). \end{align}\tag{37}\] Equivalently, it is sufficient to prove for every \(i \in \mathcal{I}\) that \[\begin{align} \label{eq:induction95proof95relation} \alpha^{i,r}(t,\boldsymbol{x}) = \kappa^{i,r}\big(t,\boldsymbol{x},\zeta^{i,r}(t,\boldsymbol{x})\big). \end{align}\tag{38}\] We prove the claim by induction. For \(r=0\), the relation 38 holds trivially, since both sides are initialized to the zero Markov policy. Assume that 38 , equivalently 37 , holds for some fixed \(r \geq 0\). We show that it then holds for \(r+1\). Using the construction 33 , the induction hypothesis, and the definition of \(\kappa^{i,r}\), we obtain \[\begin{align} \alpha^{i,r+1}(t,\boldsymbol{x}) &= \kappa^i\bigl( t,x, \boldsymbol{\Phi}(t,x)\zeta^{i,r+1}(t,\boldsymbol{x}), \boldsymbol{\alpha}^{-i,r}(t,\boldsymbol{x}) \bigr)\\ &= \kappa^i\bigl( t,\boldsymbol{x}, \boldsymbol{\Phi}(t,\boldsymbol{x})\zeta^{i,r+1}(t,\boldsymbol{x}), \boldsymbol{\kappa}^{-i,r} \big(t,\boldsymbol{x},\boldsymbol{\zeta}^{-i,r}(t,\boldsymbol{x})\big) \bigr)\\ &= \kappa^{i,r+1} \bigl(t,\boldsymbol{x},\zeta^{i,r+1}(t,\boldsymbol{x})\bigr). \end{align}\] This completes the proof.

3.2 Assumptions↩︎

Let the setting from Sections 2.22.4 hold. For technical reasons, we use the convention that \(\boldsymbol{b},\boldsymbol{\Sigma},\boldsymbol{f},\boldsymbol{\zeta}\) are killed at \(\tau\), while \(\boldsymbol{b}^{i,m},\ell^{i,m},\boldsymbol{\zeta}^{i,m}\) are killed at \(\tau^{i,m}\), by multiplication with the indicator functions \(1_{\{t\leq\tau\}}\) and \(1_{\{t\leq\tau^{i,m}\}}\), respectively. Thus, these coefficients vanish on the boundary \(\partial Q_T\), so that the processes \(\boldsymbol{X},Y\) and \(\boldsymbol{X}^{i,m},Y^{i,m}\) remain constant after \(\tau\) and \(\tau^{i,m}\), respectively. This convention does not redefine the coefficient functions and therefore does not affect their assumptions.

For all \(i \in \mathcal{I}\) and \(m \ge 1\), we assume that 35 admits a unique solution \[\begin{align} \big(\boldsymbol{X}^{i,m}, Y^{i,m}, Z^{i,m}\big) \in \mathbb{S}^{2,n}_{T} \times \mathbb{S}^{2,1}_{T} \times \mathbb{H}^{2,n}_{T}. \end{align}\] Moreover, for all \(i \in \mathcal{I}\) and \(m \geq 0\), there exists \(L_{\zeta^{i,m}}^{\scalebox{0.55}{\mathrm{FP}}}\) such that for all \((t_1,\boldsymbol{x}_1),(t_2,\boldsymbol{x}_2)\in Q_T\) we have \[\begin{align} \big\| \zeta^{i,m}(t_1,\boldsymbol{x}_1) - \zeta^{i,m}(t_2,\boldsymbol{x}_2) \big\|^2_{\mathbb{R}^{n}} &\leq L_{\zeta^{i,m}}^{\scalebox{0.55}{\mathrm{FP}}} \big(|t_1-t_2| + \|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n}^2\big). \label{eq:lipschitz95zetaim} \end{align}\tag{39}\] In addition, we assume that \[\begin{align} \label{eq:LFP} L_\zeta^{\scalebox{0.55}{\mathrm{FP}}} := \sup_{m\in\mathbb{N}} \max_{i\in\mathcal{I}} L_{\zeta^{i,m}}^{\scalebox{0.55}{\mathrm{FP}}} <\infty, \end{align}\tag{40}\] which is an assumption that we do not verify theoretically. Moreover, we assume that there exists \(L_{\boldsymbol{\zeta}}\geq0\) such that for all \((t_1,\boldsymbol{x}_1),(t_2,\boldsymbol{x}_2)\in Q_T\) we have \[\begin{align} \vert\!\vert\!\vert \boldsymbol{\zeta}(t_1,\boldsymbol{x}_1)-\boldsymbol{\zeta}(t_2,\boldsymbol{x}_2) \vert\!\vert\!\vert^2_{\mathbb{R}^{n\times N}} &\leq L_{\boldsymbol{\zeta}} \big(|t_1-t_2| + \|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n}^2\big). \end{align}\] These are abstract assumptions whose validity must be verified in any concrete application. See Remark 2 for further discussion.

We introduce the following Lipschitz condition (L). There exist constants \(L_{\boldsymbol{b}}, L_{\boldsymbol{f}}, L_\kappa^{\scalebox{0.55}{\mathrm{FP}}}, L_{\boldsymbol{b}}^{\scalebox{0.55}{\mathrm{FP}}}, L_{\ell}^{\scalebox{0.55}{\mathrm{FP}}}\geq 0\) such that, for all \(i \in \mathcal{I}\), \(m \geq 1\), and all admissible arguments, \[\begin{align} \begin{aligned} \|\boldsymbol{b}(t_1,\boldsymbol{x}_1,\boldsymbol{z}_1)-\boldsymbol{b}(t_2,\boldsymbol{x}_2,\boldsymbol{z}_2)\|_{\mathbb{R}^{n}}^2 &\leq L_{\boldsymbol{b}} \big( |t_1-t_2| + \|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n}^2 + \vert\!\vert\!\vert \boldsymbol{z}_1 - \boldsymbol{z}_2 \vert\!\vert\!\vert^2_{\mathbb{R}^{n\times N}} \big),\\ \|\boldsymbol{f}(t_1,\boldsymbol{x}_1,\boldsymbol{z}_1)-\boldsymbol{f}(t_2,\boldsymbol{x}_2,\boldsymbol{z}_2)\|_{\mathbb{R}^{N}}^2 &\leq L_{\boldsymbol{f}} \big( |t_1-t_2| + \|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n}^2 + \vert\!\vert\!\vert \boldsymbol{z}_1 - \boldsymbol{z}_2 \vert\!\vert\!\vert^2_{\mathbb{R}^{n\times N}} \big), \\ \big\| \kappa^{i,m}(t_1,\boldsymbol{x}_1,z_1) - \kappa^{i,m}(t_2,\boldsymbol{x}_2,z_2) \big\|_{\mathbb{R}^d}^2 &\leq L_\kappa^{\scalebox{0.55}{\mathrm{FP}}} \big( |t_1-t_2| + \|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n}^2 + \| z_1 - z_2 \|^2_{\mathbb{R}^{n}} \big),\\ \big\| \boldsymbol{b}^{i,m}(t_1,\boldsymbol{x}_1,z_1^i,\boldsymbol{z}_1^{-i})-\boldsymbol{b}^{i,m}(t_2,\boldsymbol{x}_2,z^i_2,\boldsymbol{z}_2^{-i}) \big\|_{\mathbb{R}^{n}}^2 &\leq L_{b}^{\scalebox{0.55}{\mathrm{FP}}} \big( |t_1-t_2| + \|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n}^2 + \vert\!\vert\!\vert \boldsymbol{z}_1 - \boldsymbol{z}_2 \vert\!\vert\!\vert^2_{\mathbb{R}^{n\times N}} \big),\\ \big| \ell^{i,m}(t_1,\boldsymbol{x}_1,z_1^i,\boldsymbol{z}_1^{-i}) - \ell^{i,m}(t_2,\boldsymbol{x}_2,z_2^i,\boldsymbol{z}_2^{-i}) \big|^2 &\leq L_{\ell}^{\scalebox{0.55}{\mathrm{FP}}} \big( |t_1-t_2| + \|\boldsymbol{x}_1-\boldsymbol{x}_2\|_{\mathbb{R}^n}^2 + \vert\!\vert\!\vert \boldsymbol{z}_1 - \boldsymbol{z}_2 \vert\!\vert\!\vert^2_{\mathbb{R}^{n\times N}} \big). \end{aligned} \end{align}\] Here, admissible arguments means either all \(\boldsymbol{z}_1, \boldsymbol{z}_2\in\mathbb{R}^{n\times N}\), in the global case, or all \(\boldsymbol{z}_1,\boldsymbol{z}_2\) in a bounded subset of \(\mathbb{R}^{n\times N}\), in the local case.

We introduce two alternative sets of assumptions and require that at least one of them holds.

  1. Gradient bound condition: The domain is either \(\mathcal{D}= \mathbb{R}^n\) or \(|\mathcal{D}| < \infty\) with boundary \(\partial \mathcal{D}\) of class \(C^2\). For the corresponding fictitious-play iterations, there exists a unique collection \(\{V^{i,m} \}_{i\in\mathcal{I},\; m\ge1} \subset C^{1,2}(\bar Q_T)\) such that, for all \(i\in\mathcal{I}\), \(m\ge1\), and \((t,\boldsymbol{x})\in Q^0_T\), it holds \[\begin{align} \label{eq:HJB-best-response} \frac{\partial V^{i,m}}{\partial t}(t,\boldsymbol{x}) + \mathcal{L}_tV^{i,m}(t,\boldsymbol{x}) + H^{i,m} \big( t,\boldsymbol{x}, \mathrm D_{\boldsymbol{x}}V^{i,m}(t,\boldsymbol{x}), \boldsymbol{\zeta}^{-i,m-1}(t,\boldsymbol{x}) \big) =0, \end{align}\tag{41}\] and for all \((t,\boldsymbol{x})\in\partial Q_T\), \[V^{i,m}(t,\boldsymbol{x})=g^i(t,\boldsymbol{x}).\] Here, for \((t,\boldsymbol{x},p,\boldsymbol{z}^{-i})\in Q_T\times\mathbb{R}^n\times\mathbb{R}^{n\times (N-1)}\), the Hamiltonian is given by \[H^{i,m}(t,\boldsymbol{x},p,\boldsymbol{z}^{-i}) = \big\langle \boldsymbol{b}^{i,m} \big( t,\boldsymbol{x}, \boldsymbol{\Sigma}^\top(t,\boldsymbol{x})p, \boldsymbol{z}^{-i} \big), p \big\rangle + \ell^{i,m} \big( t,\boldsymbol{x}, \boldsymbol{\Sigma}^\top(t,\boldsymbol{x})p, \boldsymbol{z}^{-i} \big).\] Moreover, for \(m\ge1\), we set \[\zeta^{i,m}(t,\boldsymbol{x}) = \boldsymbol{\Sigma}^\top(t,\boldsymbol{x})\mathrm D_{\boldsymbol{x}}V^{i,m}(t,\boldsymbol{x}).\] Finally, there exists \(M^{\scalebox{0.55}{\mathrm{FP}}}< \infty\) such that \[\begin{align} \label{eq:grad95bound} \sup_{m\in\mathbb{N}}\max_{i\in\mathcal{I}}\sup_{(t,\boldsymbol{x})\in Q_T} \big\| \mathrm D_{\boldsymbol{x}} V^{i,m}(t,\boldsymbol{x}) \big\|_{\mathbb{R}^n}^2 \leq M^{\scalebox{0.55}{\mathrm{FP}}}. \end{align}\tag{42}\]

  2. Smallness condition: The domain is \(\mathcal{D}=\mathbb{R}^n\) and there exist \(\beta > 0\) and \(\eta > 0\) such that \[\begin{align} 2C_{\beta} < 1, \qquad \frac{2C_\beta}{1 - 2C_\beta}(N-1) < 1, \qquad (1+\eta)L_\kappa^a < 1, \end{align}\] and \[\begin{align} \big(1 + \eta^{-1}\big) L_\kappa^a L_{\kappa}^{\scalebox{0.55}{\mathrm{FP}}} \frac{2C_\beta' N}{1-2C_\beta} < \Big( 1 - \frac{2C_\beta}{1 - 2C_\beta}(N-1) \Big) \big(1 - (1+\eta)L_\kappa^a \big), \end{align}\] where \[\begin{align} C_\beta &\vcentcolon=\frac{4L_\ell^{\scalebox{0.55}{\mathrm{FP}}}}{\beta} +Ce^{\beta T} \bigg(L_g+\frac{2L_\ell^{\scalebox{0.55}{\mathrm{FP}}} T(1+2L_\zeta^{\scalebox{0.55}{\mathrm{FP}}} N)}{\beta} + L_\zeta^{\scalebox{0.55}{\mathrm{FP}}} T\bigg), \\ C_\beta' &\vcentcolon= \frac{2 L_{\bar{\boldsymbol{f}}} (1+L_\kappa)}{\beta} +Ce^{\beta T} \bigg(L_g+\frac{2L_\ell^{\scalebox{0.55}{\mathrm{FP}}} T(1+2L_\zeta^{\scalebox{0.55}{\mathrm{FP}}} N)}{\beta} + L_\zeta^{\scalebox{0.55}{\mathrm{FP}}} T\bigg). \end{align}\] The constant \(C\) is derived in Lemma 3 and includes the non-explicit constant of the Burkholder–Davis–Gundy inequality. Furthermore, the Lipschitz condition (L) holds globally.

Under Assumption [assump:gradient], the gradient bound 42 yields a uniform bound on the relevant \(\boldsymbol{z}\)-argument, and hence the local version of the Lipschitz condition (L) holds. Under Assumption [assump:smalltime], no such bound is available, so the global version of (L) is imposed directly. A notable case in which the global version can be verified from the preceding Lipschitz assumptions is additive time-independent noise. Appendix 5 contains the detailed verification of the Lipschitz condition (L).

Remark 2. Assumptions 39 , 40 and 42 are abstract assumptions that need to be verified in specific settings. For \(\mathcal{D}=\mathbb{R}^n\), the gradient estimate 42 holds under strong regularity assumptions, for instance by [35], whereas [36] obtains such estimates under substantially weaker assumptions, namely Lipschitz continuity and linear growth of the coefficients.

For \(|\mathcal{D}|<\infty\), however, we are not aware of a directly applicable result in the literature for the semilinear fictitious-play HJB system considered here. The parabolic Hölder-space theory for linear equations (see, e.g., [37]) provides existence, uniqueness, and a priori estimates, which in particular yield bounds on spatial derivatives under appropriate assumptions. Extending such estimates to the present semilinear setting would require an additional argument, most naturally based on a fixed-point approach combined with Schauder estimates, for instance via Banach’s fixed point theorem, Schauder’s fixed point theorem, or the Leray–Schauder fixed point theorem.

3.3 Convergence analysis↩︎

The goal of this subsection is to establish the main convergence result for the fictitious-play scheme. To quantify the error at iteration \(m\), we define \[\begin{align} \Psi^{m} \vcentcolon=\max_{i\in\mathcal{I}}\big\|\delta\boldsymbol{X}^{i,m}\big\|^2_{\mathbb{S}^{2,n}_{T}} +& \sum_{i\in\mathcal{I}}\big\|\delta Y^{i,m}\big\|^2_{\mathbb{S}^{2,1}_{T}} + \sum_{i\in\mathcal{I}}\big\|\zeta^{i,[\tau]}(\cdot,\boldsymbol{X}_\cdot)-\zeta^{i,m,[\tau^{i,m}]}\big(\cdot,\boldsymbol{X}_\cdot^{i,m}\big)\big\|_{\mathbb{H}^{2,n}_{T}}^2. \label{eq:Psim} \end{align}\tag{43}\] Our main theorem shows that, under the assumptions of Section 3.2, this error decays geometrically as \(m \to \infty\). More precisely, we prove an estimate of the form \[\begin{align} \Psi^{m} \lesssim CK^m, \end{align}\] for \(C < \infty\) and \(K \in (0,1)\). Under Assumption [assump:gradient], in special cases where the policy-mismatch error vanishes, the estimate can be sharpened to a super-exponential rate. This refinement is discussed after the main theorem.

The proof is divided into three steps. We first derive a decomposition of the error \(\Psi^{m}\) into stopping-time and Markov-map contributions. Then, we bound the stopping-time error in terms of the Markov-map error, and finally estimate the Markov-map error itself. These results are combined in the last part, where we state and prove the main theorem. We begin by introducing the notation used throughout the convergence analysis.

3.3.1 Notation↩︎

For \(i \in \mathcal{I}\), \(m \geq 1\) and \(t \in [0,T]\), define the coefficient differences by \[\begin{align} \delta \boldsymbol{b}^{i,m}_t&\vcentcolon=\boldsymbol{b}\big(t,\boldsymbol{X}_t,\zeta^i(t,\boldsymbol{X}_t),\boldsymbol{\zeta}^{-i}(t,\boldsymbol{X}_t)\big)-\boldsymbol{b}^{i,m}\big(t,\boldsymbol{X}_t^{i,m},\zeta^{i,m}\big(t,\boldsymbol{X}_t^{i,m}\big),\boldsymbol{\zeta}^{-i,m-1}\big(t,\boldsymbol{X}_t^{i,m}\big)\big), \tag{44} \\ \delta \boldsymbol{\Sigma}_t^{i,m} &\vcentcolon= \boldsymbol{\Sigma}(t,\boldsymbol{X}_t) - \boldsymbol{\Sigma}\big(t,\boldsymbol{X}_t^{i,m}\big), \tag{45} \\ \delta \ell_t^{i,m} &\vcentcolon= f^i\big(t,\boldsymbol{X}_t,\zeta^i(t,\boldsymbol{X}_t),\boldsymbol{\zeta}^{-i}(t,\boldsymbol{X}_t)\big)- \ell^{i,m}\big(t,\boldsymbol{X}_t^{i,m},\zeta^{i,m}\big(t,\boldsymbol{X}_t^{i,m}\big),\boldsymbol{\zeta}^{-i,m-1}\big(t,\boldsymbol{X}_t^{i,m}\big)\big),\\ \delta g^{i,m} &\vcentcolon= g^i(\tau,\boldsymbol{X}_\tau)- g^i\big(\tau^{i,m},\boldsymbol{X}_{\tau^{i,m}}^{i,m}\big). \end{align}\] The Markov maps require a more delicate analysis, for which we need to distinguish between the errors in the Markov maps themselves and errors caused by evaluating the same map along different state processes. For \(i,j\in\mathcal{I}\), \(m,k\in\mathbb{N}\) and \(t\in[0,T]\) we therefore define \[\begin{align} \delta \zeta_t^{j,i,m,k} &\vcentcolon= \zeta^{j}(t,\boldsymbol{X}_t)-\zeta^{j,m}\big(t,\boldsymbol{X}_t^{i,k}\big), &\qquad \delta\boldsymbol{\zeta}_t^{-i,m,k} &\vcentcolon=\boldsymbol{\zeta}^{-i}(t,\boldsymbol{X}_t)-\boldsymbol{\zeta}^{-i,m}\big(t,\boldsymbol{X}_t^{i,k}\big),\\ \delta \bar{\zeta}_t^{i,m} &\vcentcolon=\zeta^{i}(t,\boldsymbol{X}_t)-\zeta^{i,m}(t,\boldsymbol{X}_t), &\qquad \delta \bar{\boldsymbol{\zeta}}_t^{-i,m} &\vcentcolon=\boldsymbol{\zeta}^{-i}(t,\boldsymbol{X}_t)-\boldsymbol{\zeta}^{-i,m}(t,\boldsymbol{X}_t),\\ \delta \tilde{\zeta}_t^{j,i,m,k} &\vcentcolon=\zeta^{j,m}(t,\boldsymbol{X}_t)-\zeta^{j,m}\big(t,\boldsymbol{X}_t^{i,k}\big), &\qquad \delta \tilde{\boldsymbol{\zeta}}_t^{-i,m,k} &\vcentcolon=\boldsymbol{\zeta}^{-i,m}(t,\boldsymbol{X}_t)-\boldsymbol{\zeta}^{-i,m}\big(t,\boldsymbol{X}_t^{i,k}\big). \end{align}\] Here, the bar notation denotes the difference between the true Markov map and its current approximation evaluated at the same state process, whereas the tilde notation captures the effect of evaluating the same approximation at different state processes. In particular, we have the decomposition \[\begin{align} \label{eq:split} \delta\zeta_t^{j,i,m,k} = \delta\bar{\zeta}_t^{j,m}+\delta\tilde{\zeta}_t^{j,i,m,k}. \end{align}\tag{46}\] We also use the shorthand notation \[\begin{align} \delta\zeta^{i,m,k}_t \vcentcolon=\delta\zeta^{i,i,m,k}_t, \quad \delta\zeta^{j,i,m}_t \vcentcolon=\delta\zeta^{j,i,m,m}_t, \quad \delta\zeta^{i,m}_t \vcentcolon=\delta\zeta^{i,i,m,m}_t, \end{align}\] and analogously for the tilde notation. Furthermore, for the stopping times, we write \[\begin{align} \delta \tau^{i,m} \vcentcolon=\tau - \tau^{i,m}, \quad \tau_{\min}^{i,m} \vcentcolon=\tau\wedge\tau^{i,m}, \quad \tau_{\max}^{i,m} \vcentcolon=\tau\vee\tau^{i,m}. \end{align}\] Finally, we define two error quantities that are central in the analysis. For \(i \in \mathcal{I}\) and \(m \geq 0\), denote the player-specific policy mismatch by \[\begin{align} \Delta^{i,m}(t,\boldsymbol{x}) := \alpha^{*,i}(t,\boldsymbol{x}) - \kappa^{i,m} \big( t, \boldsymbol{x}, \zeta^i(t,\boldsymbol{x}) \big), \end{align}\] and let the policy error be quantified by \[\begin{align} E_\Delta^m := \sum_{i \in \mathcal{I}} \big\| \Delta^{i,m}(\cdot, \boldsymbol{X}_{\cdot}) \big\|^2_{\mathbb{H}^{2,d}_T}. \end{align}\] For the global Markov-map error, we introduce \[\begin{align} E_\zeta^m \vcentcolon=\sum_{i \in \mathcal{I}} \big\| \delta \bar\zeta^{i,m} \big\|^2_{\mathbb{H}^{2,n}_T}. \end{align}\]

3.3.2 State error decomposition↩︎

The purpose of this subsection is to estimate the different terms in the error quantity \(\Psi^{m}\) defined in 43 . The main goal is to express the approximation errors in \(\boldsymbol{X}\), \(Y\) and \(\zeta\) in terms of the stopping-time error \(\delta \tau\), the Markov-map error \(\delta \bar \zeta\) and the policy-mismatch error.

First, Lemma 1 provides a decomposition of \(\delta \zeta\) that will be used repeatedly throughout the analysis. Second, Lemma 2 derives bounds for the coefficient differences \(\delta \boldsymbol{b}^{i,m}_t\) and \(\delta \ell^{i,m}_t\). Lemma 3 then estimates the forward error \(\delta \boldsymbol{X}\), while Lemma 4 controls the backward error \(\delta Y\) together with the associated \(\zeta\)-term. Combining these estimates yields Proposition 1, which gives the desired decomposition of the overall error \(\Psi^{m}\).

Lemma 1. For any \(t\in[0,T]\), \(i,j\in\mathcal{I}\), \(\beta\geq 0\), and \(m,l\in\mathbb{N}\), it holds \[\Big\|\delta\zeta^{j,i,l,m,\,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{\beta,t}}^{2} \;\le\; 2\,\Big\|\delta\bar\zeta^{j,l,\,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{\beta,t}}^{2} \;+\; 2\,L_\zeta^{\scalebox{0.55}{\mathrm{FP}}}\, \Big\|\delta \boldsymbol{X}^{i,m,\,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{\beta,t}}^{2}.\]

Proof. The result follows from the split 46 , the elementary inequality \[\begin{align} \left\| a_1 + a_2 \right\|^2_{\mathcal{V}} \leq 2 \left\| a_1 \right\|^2_{\mathcal{V}} + 2 \left\| a_2 \right\|^2_{\mathcal{V}}, \quad a_1, a_2 \in \mathcal{V}, \label{ineq:sumsq} \end{align}\tag{47}\] for any normed vector space \(\mathcal{V}\), and the Lipschitz continuity of \(\zeta^{j,l}\). ◻

Lemma 2. Fix \(i \in \mathcal{I}\) and \(m \geq 1\). Then, there exists \(C < \infty\) such that \[\begin{align} &\Big\| \delta \boldsymbol{b}^{i,m,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,n}_T} + \Big\| \delta \ell^{i,m,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,1}_T} \\ &\qquad \leq C \bigg( \Big\| \delta \boldsymbol{X}^{i,m,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,n}_T} + \Big\| \delta \bar{\zeta}^{i,m,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,n}_T} + \sum_{j\in \mathcal{I}\backslash\{i\}} \Big\| \delta \bar{\zeta}^{j,m-1,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,n}_T} + E_\Delta^{m-1} \bigg). \end{align}\]

Proof. Using 47 , we can split the error as \[\begin{align} \Big\| \delta \boldsymbol{b}^{i,m,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,n}_T} \leq 2\, \Big\| \delta\bar{\boldsymbol{b}}^{i,m,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,n}_T} + 2\, \Big\| \delta \tilde{\boldsymbol{b}}^{i,m,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,n}_T} \end{align}\] where \[\begin{align} \delta \bar{\boldsymbol{b}}^{i,m}_{t} &:= \boldsymbol{b} \big( t,\boldsymbol{X}_t,\zeta^i(t,\boldsymbol{X}_t),\boldsymbol{\zeta}^{-i}(t,\boldsymbol{X}_t) \big) - \boldsymbol{b}^{i,m} \big( t,\boldsymbol{X}_t,\zeta^{i}\big(t,\boldsymbol{X}_t\big),\boldsymbol{\zeta}^{-i}\big(t,\boldsymbol{X}_t\big) \big), \\ \delta \tilde{\boldsymbol{b}}^{i,m}_{t} &:= \boldsymbol{b}^{i,m} \big( t,\boldsymbol{X}_t,\zeta^{i}\big(t,\boldsymbol{X}_t\big),\boldsymbol{\zeta}^{-i}\big(t,\boldsymbol{X}_t\big) \big) - \boldsymbol{b}^{i,m} \big( t,\boldsymbol{X}_t^{i,m},\zeta^{i,m}\big(t,\boldsymbol{X}_t^{i,m}\big),\boldsymbol{\zeta}^{-i,m-1}\big(t,\boldsymbol{X}_t^{i,m}\big) \big). \end{align}\] The bound for \(\delta \tilde{\boldsymbol{b}}^{i,m}_{t}\) follows immediately from the Lipschitz continuity of \(\boldsymbol{b}^{i,m}\). More precisely, we get \[\begin{align} \big\| \delta \tilde{\boldsymbol{b}}^{i,m}_{t} \big\|_{\mathbb{R}^n}^2 \leq C \bigg( \big\| \delta \boldsymbol{X}^{i,m}_t \big\|_{\mathbb{R}^n}^2 + \big\| \zeta^i(t,\boldsymbol{X}_t) - \zeta^{i,m}(t,\boldsymbol{X}_t^{i,m}) \big\|_{\mathbb{R}^n}^2 + \sum_{j\in \mathcal{I}\backslash\{i\}} \big\| \zeta^j(t,\boldsymbol{X}_t) - \zeta^{j,m-1}(t,\boldsymbol{X}_t^{i,m}) \big\|_{\mathbb{R}^n}^2 \bigg). \end{align}\] After killing at \(\tau^{i,m}_{\min}\), integrating, taking expectation and applying Lemma 1, this gives \[\begin{align} \label{eq:B95b95bound} \Big\| \delta\tilde{\boldsymbol{b}}^{i,m,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,n}_T} \leq C \bigg( \Big\| \delta \boldsymbol{X}^{i,m,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,n}_T} + \Big\| \delta \bar{\zeta}^{i,m,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,n}_T} + \sum_{j\in \mathcal{I}\backslash\{i\}} \Big\| \delta \bar{\zeta}^{j,m-1,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,n}_T} \bigg). \end{align}\tag{48}\] For \(\delta\bar{\boldsymbol{b}}^{i,m}_{t}\), it follows from definitions that \[\begin{align} \delta\bar{\boldsymbol{b}}^{i,m}_{t} = \bar\boldsymbol{b}(t,\boldsymbol{X}_t,\boldsymbol{\alpha}^{*}(t,\boldsymbol{X}_t)) - \bar\boldsymbol{b}(t,\boldsymbol{X}_t,\boldsymbol{\alpha}^{i,m}(t,\boldsymbol{X}_t)), \end{align}\] where \[\begin{align} \boldsymbol{\alpha}^{i,m}(t,\boldsymbol{x}) = \big[ \kappa^i \big( t, \boldsymbol{x}, \boldsymbol{\Phi}(t,\boldsymbol{x})\zeta^i(t,\boldsymbol{x}), \boldsymbol{\kappa}^{-i,m-1} ( t, \boldsymbol{x}, \boldsymbol{\zeta}^{-i}(t,\boldsymbol{x}) ) \big) , \boldsymbol{\kappa}^{-i,m-1} ( t, \boldsymbol{x}, \boldsymbol{\zeta}^{-i}(t,\boldsymbol{x}) ) \big]_i. \end{align}\] By the Lipschitz continuity of \(\bar\boldsymbol{b}\), we have \[\begin{align} \big\| \delta\bar{\boldsymbol{b}}^{i,m}_{t} \big\|_{\mathbb{R}^n}^2 &\leq L_{\bar\boldsymbol{b}} \sum_{j \in \mathcal{I}} \big\| \alpha^{*,j}(t,\boldsymbol{X}_t) - \alpha^{i,m,j}(t,\boldsymbol{X}_t) \big\|_{\mathbb{R}^d}^2. \end{align}\] Note that for \(j \neq i\) the summation term is \(\Delta^{j,m-1}\). For the term \(j = i\), we have \[\begin{align} \big\| \alpha^{*,j}(t,\boldsymbol{X}_t) &- \alpha^{i,m,j}(t,\boldsymbol{X}_t) \big\|_{\mathbb{R}^d}^2 \\ &\qquad = \big\| \alpha^{*,j}(t,\boldsymbol{X}_t) - \kappa^i \big( t, \boldsymbol{X}_t, \boldsymbol{\Phi}(t,\boldsymbol{X}_t)\zeta^i(t,\boldsymbol{X}_t), \boldsymbol{\kappa}^{-i,m-1} ( t, \boldsymbol{X}_t, \boldsymbol{\zeta}^{-i}(t,\boldsymbol{X}_t) ) \big) \big\|_{\mathbb{R}^d}^2 \\ &\qquad \leq L_{\kappa} \sum_{j \in \mathcal{I}\backslash\{i\}} \big\| \alpha^{*,j}(t,\boldsymbol{X}_t) - \kappa^{j,m-1} ( t, \boldsymbol{X}_t, \zeta^{j}(t,\boldsymbol{X}_t) ) \big\|_{\mathbb{R}^d}^2 \\ &\qquad = L_{\kappa} \sum_{j \in \mathcal{I}\backslash\{i\}} \big\| \Delta^{j,m-1}(t,\boldsymbol{X}_t) \big\|_{\mathbb{R}^d}^2. \end{align}\] Collecting terms, we get \[\begin{align} \big\| \delta\bar{\boldsymbol{b}}^{i,m}_{t} \big\|_{\mathbb{R}^n}^2 &\leq L_{\bar\boldsymbol{b}}(1+L_\kappa) \sum_{j\in\mathcal{I}\backslash\{i\}} \big\| \Delta^{j,m-1} \big\|_{\mathbb{R}^d}^2. \end{align}\] After killing at \(\tau^{i,m}_{\min}\), integrating and taking expectation, we have \[\begin{align} \Big\| \delta\bar{\boldsymbol{b}}^{i,m,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,n}_T} \leq L_{\bar\boldsymbol{b}} (1 + L_{\kappa}) \sum_{j \in \mathcal{I}\backslash\{i\}} \Big\| \Delta^{j,m-1}(\cdot,\boldsymbol{X}_{\cdot})^{[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,d}_T} \leq L_{\bar\boldsymbol{b}}(1 + L_{\kappa}) E_{\Delta}^{m-1}. \end{align}\] Next, we estimate the \(\delta \ell^{i,m}\) part of the lemma. The error can similarly be split as \[\begin{align} \Big\| \delta \ell^{i,m,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,1}_T} \leq 2\, \Big\| \delta \bar{\ell}^{i,m,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,1}_T} + 2\, \Big\| \delta \tilde{\ell}^{i,m,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,1}_T} \end{align}\] where \[\begin{align} \delta \bar{\ell}^{i,m}_{t} &:= f^i \big( t,\boldsymbol{X}_t,\zeta^i(t,\boldsymbol{X}_t),\boldsymbol{\zeta}^{-i}(t,\boldsymbol{X}_t) \big) - \ell^{i,m} \big( t,\boldsymbol{X}_t,\zeta^{i}\big(t,\boldsymbol{X}_t\big),\boldsymbol{\zeta}^{-i}\big(t,\boldsymbol{X}_t\big) \big), \\ \delta\tilde{\ell}^{i,m}_{t} &:= \ell^{i,m} \big( t,\boldsymbol{X}_t,\zeta^{i}\big(t,\boldsymbol{X}_t\big),\boldsymbol{\zeta}^{-i}\big(t,\boldsymbol{X}_t\big) \big) - \ell^{i,m} \big( t,\boldsymbol{X}_t^{i,m},\zeta^{i,m}\big(t,\boldsymbol{X}_t^{i,m}\big),\boldsymbol{\zeta}^{-i,m-1}\big(t,\boldsymbol{X}_t^{i,m}\big) \big). \end{align}\] For \(\delta\tilde{\ell}^{i,m}_{t}\), we use the Lipschitz continuity of \(\ell^{i,m}\), and repeat the computations used for \(\delta\tilde{\boldsymbol{b}}^{i,m}_{t}\) to get a bound on the same form as 48 . For \(\delta \bar{\ell}^{i,m}_{t}\), we have from definitions \[\begin{align} \delta \bar{\ell}^{i,m}_{t} = \bar f^i \big( t,\boldsymbol{X}_t,\boldsymbol{\alpha}^{*}(t,\boldsymbol{X}_t) \big) - \bar f^i \big( t,\boldsymbol{X}_t,\boldsymbol{\alpha}^{i,m}(t,\boldsymbol{X}_t) \big). \end{align}\] Thus, repeating the computations used for \(\delta \bar{\boldsymbol{b}}^{i,m}_{t}\), we arrive at the bound \[\begin{align} \Big\| \delta \bar{\ell}^{i,m,[\tau^{i,m}_{\min}]} \Big\|^2_{\mathbb{H}^{2,1}_T} \leq L_{\bar{\boldsymbol{f}}}(1 + L_{\kappa})E_\Delta^{m-1}, \end{align}\] where \(L_{\bar{\boldsymbol{f}}}\) comes from using the Lipschitz continuity of \(\bar{\boldsymbol{f}}\) instead of \(\bar\boldsymbol{b}\). This completes the proof. ◻

Lemma 3. There exists \(C>0\) such that for all \(i\in\mathcal{I}\) and \(m\geq1\) it holds \[\begin{align} \Big\|\delta \boldsymbol{X}^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{S}^{2,n}_{T}}^2 \leq C\bigg( \Big\| \delta \bar\zeta^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2 + \sum_{j\in\mathcal{I}\setminus\{i\}}\Big\| \delta \bar\zeta^{j,m-1,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2 + E_\Delta^{m-1} \bigg). \label{eq:delta95x95bound951} \end{align}\tag{49}\] Moreover, for all \(\varepsilon\in(0,1)\) there exists \(C>0\) such that for all \(i\in\mathcal{I}\) and \(m\geq1\) it holds \[\begin{align} \big\|\delta \boldsymbol{X}^{i,m}\big\|_{\mathbb{S}^{2,n}_{T}}^2 \leq C\bigg(\big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}}^{1-\varepsilon} + \Big\| \delta\bar \zeta^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2 + \sum_{j\in\mathcal{I}\setminus\{i\}}\Big\| \delta \bar\zeta^{j,m-1,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2 + E_\Delta^{m-1} \bigg). \label{eq:delta95x95bound952} \end{align}\tag{50}\]

Proof. Let \(i \in \mathcal{I}\) and \(m \geq 1\) be fixed for the remainder of the proof. For any \(s \in [0,T]\), we have \(\mathbb{P}\)-almost surely \[\begin{align} \delta\boldsymbol{X}^{i,m,[\tau_{\min}^{i,m}]}_s &= \int_0^s \delta \boldsymbol{b}^{i,m,[\tau_{\min}^{i,m}]}_r\,\text{d}r + \int_0^s \delta \boldsymbol{\Sigma}^{i,m,[\tau_{\min}^{i,m}]}_r\,\text{d}\boldsymbol{W}_r. \end{align}\] Taking the squared \(\mathbb{R}^n\)-norm of both sides, applying 47 followed by the Cauchy–Schwarz inequality for the first term, we obtain \[\begin{align} \Big\|\delta \boldsymbol{X}^{i,m,[\tau_{\min}^{i,m}]}_s\Big\|_{\mathbb{R}^n}^2 \leq\;& 2s\int_0^s\Big\|\delta\boldsymbol{b}^{i,m,[\tau_{\min}^{i,m}]}_r \Big\|_{\mathbb{R}^n}^2\,\text{d}r + 2\, \bigg\|\int_0^s\delta \boldsymbol{\Sigma}^{i,m,[\tau_{\min}^{i,m}]}_r\,\text{d}\boldsymbol{W}_r\bigg\|_{\mathbb{R}^n}^2 . \end{align}\] Next, take the supremum over \(s \in [0,t]\), first on the right-hand side and then on the left-hand side, and take expectations on both sides. Recalling the definitions of \(\mathbb{S}^{2,n}_{t}\)- and \(\mathbb{H}^{2,n}_{t}\)-norms, we have \[\begin{align} \label{eq:295terms}\begin{aligned} \Big\|\delta \boldsymbol{X}^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{S}^{2,n}_{t}}^2 \lesssim\;& \Big\|\delta\boldsymbol{b}^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{t}}^2 + \bigg\|\int_0^\cdot \delta \boldsymbol{\Sigma}^{i,m,[\tau_{\min}^{i,m}]}_r\,\text{d}\boldsymbol{W}_r\bigg\|_{\mathbb{S}_{t}^{2,n}}^2 . \end{aligned} \end{align}\tag{51}\] We bound the two terms on the right-hand side of 51 separately. For the first term, we use the bound from Lemma 2. For the second term of 51 , we apply the Burkholder–Davis–Gundy inequality 5 , use the definition of \(\delta \boldsymbol{\Sigma}^{i,m}\) in 45 together with the Lipschitz continuity of \(\boldsymbol{\Sigma}\) from 10 , to get \[\begin{align} \bigg\|\int_0^\cdot\delta \boldsymbol{\Sigma}^{i,m,[\tau_{\min}^{i,m}]}_r\,\text{d}\boldsymbol{W}_r\bigg\|_{\mathbb{S}_{t}^{2,n}}^2 \lesssim \Big\|\delta \boldsymbol{\Sigma}^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n,n}_{t}}^2 \lesssim \Big\| \delta \boldsymbol{X}^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{t}}^2 \leq \int_0^t\Big\|\delta \boldsymbol{X}^{i,m,[\tau_{\min}^{i,m}]}\Big\|^2_{\mathbb{S}^{2,n}_{s}}\, \text{d}s, \label{eq:second95bound} \end{align}\tag{52}\] where the last inequality follows from the definitions of the \(\mathbb{H}^{2,n}_{t}\)- and \(\mathbb{S}^{2,n}_{s}\)-norms. In total, we have \[\begin{align} \Big\|\delta \boldsymbol{X}^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{S}^{2,n}_{t}}^2 &\lesssim \Big\|\delta \bar{\zeta}^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{t}}^2 + \sum_{j\in\mathcal{I}\setminus\{i\}} \Big\|\delta \bar{\zeta}^{j,m-1,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{t}}^2 + \int_0^t\Big\|\delta \boldsymbol{X}^{i,m,[\tau_{\min}^{i,m}]}\Big\|^2_{\mathbb{S}^{2,n}_{s}}\, \text{d}s + E_\Delta^{m-1}. \end{align}\] Since this inequality holds for all \(t \in [0,T]\), Gronwall’s lemma yields \[\begin{align} \Big\|\delta \boldsymbol{X}^{i,m,[\tau_\text{min}^{i,m}]}\Big\|_{\mathbb{S}^{2,n}_{T}}^2 \lesssim \Big\|\delta \bar{\zeta}^{i,m,[\tau_\text{min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2 + \sum_{j\in\mathcal{I}\setminus\{i\}}\Big\|\delta \bar{\zeta}^{j,m-1,[\tau_\text{min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2 + E_\Delta^{m-1}. \end{align}\] This proves the first claim of the lemma.

For the second claim, we first use 47 as \[\begin{align} \label{eq:delta95S95S2} \begin{aligned} \big\|\delta \boldsymbol{X}^{i,m}\big\|_{\mathbb{S}^{2,n}_{T}}^2 &\le 2\Big\|\delta\boldsymbol{X}^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{S}^{2,n}_{T}}^2 +2\Big\|\delta\tilde{\boldsymbol{X}}^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}\Big\|_{\mathbb{S}^{2,n}_{T}}^2, \end{aligned} \end{align}\tag{53}\] where the tilde-notation denotes the increment of the state-error dynamics over \([\tau^{i,m}_{\min}, \tau^{i,m}_{\max}]\) initialized at zero at \(\tau^{i,m}_{\min}\), that is \[\begin{align} \delta\tilde{\boldsymbol{X}}^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}_s \vcentcolon= \int_{0}^{s} \delta \boldsymbol{b}^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}_r \, \text{d}r + \int_{0}^{s} \delta \boldsymbol{\Sigma}^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}_r \, \text{d}\boldsymbol{W}_r. \end{align}\] Note that the first term in 53 is already bounded from the first part of the lemma. For the second term, similarly to 51 , we have \[\begin{align} \Big\|\delta \tilde{\boldsymbol{X}}^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}\Big\|_{\mathbb{S}^{2,n}_{T}}^2 \lesssim\;& \Big\|\delta\boldsymbol{b}^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2 + \bigg\|\int_0^\cdot\delta \boldsymbol{\Sigma}^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}_r\,\text{d}\boldsymbol{W}_r\bigg\|_{\mathbb{S}_{T}^{2,n}}^2 . \label{eq:295terms95again} \end{align}\tag{54}\] Again, we bound the two terms on the right-hand side separately. For the first term, taking \(\varepsilon\in(0,1)\), applying Hölder’s inequality with \(p=(1-\varepsilon)^{-1}\) and \(q=\varepsilon^{-1}\) leading to \(p^{-1}+q^{-1}=1\) yields \[\begin{align} \begin{aligned} \Big\|\delta \boldsymbol{b}^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2 &= \mathbb{E}\!\int_0^T \mathbf{1}_{\{\tau_{\min}^{i,m}< s\le \tau_{\max}^{i,m}\}} \Big\|\delta \boldsymbol{b}^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}_s\Big\|_{\mathbb{R}^n}^2\,\text{d}s \\ &\leq \bigg(\mathbb{E}\!\int_0^T \mathbf{1}_{\{\tau_{\min}^{i,m}< s\le \tau_{\max}^{i,m}\}}\,\text{d}s\bigg)^{\frac{1}{p}} \bigg(\mathbb{E}\int_0^T\Big\|\delta \boldsymbol{b}^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}_s\Big\|_{\mathbb{R}^n}^{2q}\,\text{d}s\bigg)^{\frac{1}{q}}\\ &\leq \|\delta\tau^{i,m}\|_{\mathbb{L}^{1,1}}^{1-\varepsilon} \bigg(\mathbb{E}\int_0^T\Big\|\delta \boldsymbol{b}_s^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}\Big\|_{\mathbb{R}^n}^{\frac{2}{\varepsilon}}\, \text{d}s\bigg)^{\varepsilon}. \label{eq:second95counterpart} \end{aligned} \end{align}\tag{55}\] For the second term, the Burkholder–Davis–Gundy inequality 5 and the Hölder inequality gives \[\begin{align} \begin{aligned} \bigg\|\int_0^{\cdot} \delta \boldsymbol{\Sigma}^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}_r\,\text{d}\boldsymbol{W}_r\bigg\|_{\mathbb{S}_{T}^{2,n}}^2 &\lesssim\; \Big\|\delta \boldsymbol{\Sigma}^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}\Big\|_{\mathbb{H}^{2,n,n}_{T}}^2\\ &\leq\big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}}^{1-\varepsilon}\bigg(\mathbb{E}\int_0^T\Big\|\delta \boldsymbol{\Sigma}_t^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}\Big\|_{\mathbb{R}^n}^{\frac{2}{\varepsilon}}\, \text{d}t\bigg)^{\varepsilon}. \end{aligned} \label{eq:tmptmptmp} \end{align}\tag{56}\] Insert 55 and 56 into 54 , and note that by the Lipschitz continuity of \(\boldsymbol{b}, \boldsymbol{b}^{i,m}\), \(\boldsymbol{\Sigma}\), and the Markov maps, the compactness of \(A\), and the finite moments of \(\boldsymbol{X}\) and \(\boldsymbol{X}^{i,m}\), it follows that \[\begin{align} \sup_{m\geq 1}\, \mathbb{E}\bigg[\int_0^T\Big\|\delta \boldsymbol{b}_t^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}\Big\|_{\mathbb{R}^n}^{\frac{2}{\varepsilon}}\, \text{d}t\bigg]<\infty \quad \textrm{and} \quad \sup_{m\geq 1}\, \mathbb{E}\bigg[\int_0^T\Big\|\delta \boldsymbol{\Sigma}_t^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}\Big\|_{\mathbb{R}^n}^{\frac{2}{\varepsilon}}\, \text{d}t\bigg]<\infty. \end{align}\] This concludes the proof. ◻

Lemma 4. For all \(\varepsilon\in(0,1)\) there exists \(C>0\) such that for all \(i\in\mathcal{I}\) and \(m\geq1\) we have \[\begin{align} \sum_{i\in \mathcal{I}}\big\|\delta Y^{i,m}\big\|_{\mathbb{S}^{2,1}_{T}}^2 &\lesssim E_\zeta^m + E_\zeta^{m-1} + E_\Delta^{m-1} + \sum_{i\in \mathcal{I}} \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}} + \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}}^{1-\varepsilon}, \end{align}\] and \[\begin{align} \sum_{i\in \mathcal{I}} \Big\| \zeta^{i,[\tau]}(\cdot,\boldsymbol{X}_\cdot) - \zeta^{i,m,[\tau^{i,m}]}\big(\cdot,\boldsymbol{X}_\cdot^{i,m}\big) \Big\|_{\mathbb{H}^{2,n}_{T}}^2 \lesssim E_\zeta^m + E_\zeta^{m-1} + E_\Delta^{m-1} + \sum_{i\in \mathcal{I}} \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}}^{1-\varepsilon} . \end{align}\]

Proof. Let \(i \in \mathcal{I}\) and \(m \geq 1\). For all \(t \in [0,T]\), we have \(\mathbb{P}\)-almost surely \[\begin{align} \delta Y_t^{i,m} =\;& \delta g^{i,m} + \int_t^T \delta \ell^{i,m}_s \, \text{d}s - \int_t^T \big( \delta \zeta^{i,m}_s\big)^\top\, \text{d}\boldsymbol{W}_s . \end{align}\] Note that each backward process is frozen after its respective stopping time. When the equations are extended to the fixed horizon \([0,T]\) via killed coefficients, the processes therefore remain constant beyond their stopping times. In particular, \[\begin{align} \delta Y_T^{i,m} = Y_{\tau}^{i} - Y_{\tau^{i,m}}^{i,m} = g^i(\tau,\boldsymbol{X}_\tau) - g^i\big(\tau^{i,m},\boldsymbol{X}_{\tau^{i,m}}^{i,m}\big) = \delta g^{i,m}. \label{eq:Y95T95gi} \end{align}\tag{57}\] Applying Itô’s lemma to the map \(y\mapsto|y|^2\) along \(\delta Y_t^{i,m}\), together with 57 , we have \[\begin{align} \big|\delta Y_t^{i,m}\big|^2 =& \big|\delta g^{i,m}\big|^2 + 2\int_t^T \delta Y_s^{i,m}\delta \ell^{i,m}_s\, \text{d}s - \int_t^T\big\| \delta \zeta^{i,m}_s\big\|_{\mathbb{R}^n}^2 \, \text{d}s - 2\int_t^T \delta Y_s^{i,m}\big( \delta \zeta^{i,m}_s \big)^\top\, \text{d}\boldsymbol{W}_s. \end{align}\] Dropping the third term and taking absolute values, we obtain \[\begin{align} \big|\delta Y_t^{i,m}\big|^2 &\leq \big|\delta g^{i,m}\big|^2 + 2\, \bigg|\int_t^T \delta Y_s^{i,m} \delta \ell^{i,m}_s \, \text{d}s\bigg| +2\, \bigg| \int_{t}^T \delta Y_s^{i,m}\big( \delta \zeta^{i,m}_s \big)^\top\, \text{d}\boldsymbol{W}_s\,\bigg|. \end{align}\] Taking the supremum over \(t\in[0,T]\), first on the right-hand side and then on the left-hand side, and then expectations yields \[\begin{align} \big\|\delta Y^{i,m}\big\|^2_{\mathbb{S}^{2,1}_{T}} &\leq \big\|\delta g^{i,m}\big\|^2_{\mathbb{L}^{2,1}} + 2\, \bigg\|\int_{\cdot}^T \delta Y_s^{i,m} \delta \ell^{i,m}_s \, \text{d}s\bigg\|_{\mathbb{S}^{1,1}_T} + 2\, \bigg\| \int_\cdot^T\delta Y_s^{i,m}\big(\delta \zeta_s^{i,m} \big)^\top \, \text{d}\boldsymbol{W}_s\bigg\|_{\mathbb{S}^{1,1}_{T}}. \label{eq:deltay95first95bound} \end{align}\tag{58}\] We bound the two integral terms separately. For the first one, we use 4 together with 2 , followed by Young’s inequality with parameter \(\gamma > 0\), such that \[\begin{align} 2\, \bigg\|\int_{\cdot}^T \delta Y_s^{i,m} \delta \ell^{i,m}_s \, \text{d}s\bigg\|_{\mathbb{S}^{1,1}_T} &\leq 2T \big\| \delta Y^{i,m} \big\|_{\mathbb{S}^{2,1}_{T}} \big\| \delta \ell^{i,m} \big\|_{\mathbb{H}^{2,1}_{T}} \leq \gamma^{-1} T^2\big\| \delta Y^{i,m} \big\|^2_{\mathbb{S}^{2,1}_{T}} + \gamma \big\| \delta \ell^{i,m} \big\|_{\mathbb{H}^{2,1}_{T}}^2. \label{eq:young95first} \end{align}\tag{59}\] For the second one, we use the Burkholder–Davis–Gundy inequality 5 with \(p=1\), then 3 and finally Young’s inequality, to get \[\begin{align} \begin{aligned} 2\, \bigg\| \int_\cdot^T\delta Y^{i,m}\big(\delta \zeta_s^{i,m} \big)^\top \, \text{d}\boldsymbol{W}_s\bigg\|_{\mathbb{S}^{1,1}_{T}} &\leq 2K_1\big\| \delta Y^{i,m}\delta \zeta^{i,m} \big\|_{\mathbb{H}^{1,n}_{T}} \\ &\leq 2K_1\big\| \delta Y^{i,m} \big\|_{\mathbb{S}^{2,1}_T} \big\| \delta \zeta^{i,m} \big\|_{\mathbb{H}^{2,n}_{T}} \\ &\leq K_1 \lambda^{-1} \big\|\delta Y^{i,m}\big\|_{\mathbb{S}^{2,1}_{T}}^2 + K_1\lambda \big\|\delta \zeta^{i,m}\big\|_{\mathbb{H}^{2,n}_{T}}^2. \end{aligned} \label{eq:young95second} \end{align}\tag{60}\] Substituting 59 and 60 into 58 , choosing \(\gamma = 4T^2\) and \(\lambda = 4K_1\), and moving the terms involving \(\delta Y^{i,m}\) to the left-hand side, we obtain \[\begin{align} \big\|\delta Y^{i,m}\big\|^2_{\mathbb{S}^{2,1}_{T}} &\lesssim \big\|\delta g^{i,m}\big\|^2_{\mathbb{L}^{2,1}} + \big\| \delta \ell^{i,m} \big\|_{\mathbb{H}^{2,1}_{T}}^2 + \big\|\delta \zeta^{i,m}\big\|_{\mathbb{H}^{2,n}_{T}}^2. \label{eq:Ybound95second} \end{align}\tag{61}\] We now estimate each of the three terms on the right-hand side of 61 separately. For the first term, the Lipschitz continuity of \(g^i\) and 47 imply \[\begin{align} \big\|\delta g^{i,m}\big\|_{\mathbb{L}^{2,1}}^2 &\lesssim \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}} + \Big\|\boldsymbol{X}_\tau-\boldsymbol{X}_{\tau^{i,m}}^{i,m}\Big\|_{\mathbb{L}^{2,n}}^2\\ &= \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}} + \Big\|\boldsymbol{X}_\tau-\boldsymbol{X}_{\tau^{i,m}} +\delta\boldsymbol{X}_{\tau^{i,m}}^{i,m}\Big\|_{\mathbb{L}^{2,n}}^2\\ &\leq \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}} + 2\big\|\boldsymbol{X}_\tau-\boldsymbol{X}_{\tau^{i,m}}\big\|_{\mathbb{L}^{2,n}}^2 + 2\Big\|\delta\boldsymbol{X}_{\tau^{i,m}}^{i,m}\Big\|_{\mathbb{L}^{2,n}}^2\\ &\lesssim \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}} +\big\|\delta\boldsymbol{X}^{i,m}\big\|_{\mathbb{S}^{2,n}_{T}}^2 \\ &\lesssim \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}} + \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}}^{1-\varepsilon} + \Big\| \delta\bar \zeta^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2 + \sum_{j\in\mathcal{I}\setminus\{i\}}\Big\| \delta \bar\zeta^{j,m-1,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2 + E_\Delta^{m-1}, \end{align}\] where the fourth step uses the estimates \[\begin{align} \big\|\boldsymbol{X}_\tau-\boldsymbol{X}_{\tau^{i,m}}\big\|_{\mathbb{L}^{2,n}}^2 &\lesssim \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}}, \\ \Big\|\delta\boldsymbol{X}_{\tau^{i,m}}^{i,m}\Big\|_{\mathbb{L}^{2,n}}^2 &\leq \big\|\delta\boldsymbol{X}^{i,m}\big\|_{\mathbb{S}^{2,n}_{T}}^2, \end{align}\] while the final inequality follows from Lemma 3. For the second term in 61 , we use 7 to split the process over two disjoint intervals. The first part is then estimated using Lemma 2, and the second part is treated as in 55 . This gives \[\begin{align} \big\|\delta\ell^{i,m}\big\|_{\mathbb{H}^{2,1}_{T}}^2 &= \Big\|\delta\ell^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,1}_{T}}^2 + \Big\|\delta\ell^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}\Big\|_{\mathbb{H}^{2,1}_{T}}^2 \\ &\lesssim \Big\|\delta \boldsymbol{X}^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{S}^{2,n}_{T}}^2 + \Big\|\delta\bar\zeta^{i,m,[\tau_{\min}^{i,m}]}\Big\|^2_{\mathbb{H}^{2,n}_{T}} + \sum_{j\in\mathcal{I}\setminus\{i\}}\Big\|\delta\bar\zeta^{j,m-1,[\tau_{\min}^{i,m}]}\Big\|^2_{\mathbb{H}^{2,n}_{T}} + E_\Delta^{m-1} + \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}}^{1-\varepsilon} \\ &\lesssim \Big\|\delta\bar\zeta^{i,m,[\tau_{\min}^{i,m}]}\Big\|^2_{\mathbb{H}^{2,n}_{T}} + \sum_{j\in\mathcal{I}\setminus\{i\}}\Big\|\delta\bar\zeta^{j,m-1,[\tau_{\min}^{i,m}]}\Big\|^2_{\mathbb{H}^{2,n}_{T}} + E_\Delta^{m-1} + \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}}^{1-\varepsilon}, \end{align}\] where the final inequality follows from Lemma 3. The third term in 61 is handled analogously, as \[\begin{align} \big\|\delta\zeta^{i,m}\big\|_{\mathbb{H}^{2,n}_{T}}^2 &= \Big\|\delta\zeta^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2 + \Big\|\delta\zeta^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2 \\ &\lesssim \Big\|\delta \boldsymbol{X}^{i,m,[\tau_\text{min}^{i,m}]}\Big\|_{\mathbb{S}^{2,n}_{T}}^2 + \Big\|\delta \bar{\zeta}^{i,m,[\tau_\text{min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2 + \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}}^{1-\varepsilon}\\ &\lesssim \Big\|\delta\bar\zeta^{i,m,[\tau_{\min}^{i,m}]}\Big\|^2_{\mathbb{H}^{2,n}_{T}} + \sum_{j\in\mathcal{I}\setminus\{i\}}\Big\|\delta\bar\zeta^{j,m-1,[\tau_{\min}^{i,m}]}\Big\|^2_{\mathbb{H}^{2,n}_{T}} + \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}}^{1-\varepsilon}. \end{align}\] Combining the bounds for the three terms in 61 and summing over \(i\in\mathcal{I}\), we obtain \[\begin{align} \sum_{i\in\mathcal{I}} \big\| \delta Y^{i,m} \big\|^2_{\mathbb{S}^{2,1}_{T}} \lesssim E_\zeta^{m} + E_\zeta^{m-1} + E_\Delta^{m-1} + \sum_{i\in\mathcal{I}} \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}} + \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}}^{1-\varepsilon}. \end{align}\] This concludes the first claim of the lemma. For the second claim, we have the decomposition \[\Big\|\zeta^{i,[\tau]}(\cdot,\boldsymbol{X}_\cdot)-\zeta^{i,m,[\tau^{i,m}]}\big(\cdot,\boldsymbol{X}_\cdot^{i,m}\big)\Big\|_{\mathbb{H}^{2,n}_{T}}^2 = \Big\|\delta\zeta^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2 + \Big\|\delta\zeta^{i,m,[\tau_{\min}^{i,m},\tau_{\max}^{i,m}]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2.\] Applying Lemma 1, 2 , Lemma 3 and summing over \(i\in\mathcal{I}\) yields the desired estimate. ◻

Proposition 1. Recall the error quantity \(\Psi^{m}\) from 43 . For all \(\varepsilon\in(0,1)\), \(m\geq1\), it holds that \[\begin{align} \Psi^{m} \lesssim E_\zeta^{m} + E_\zeta^{m-1} + E_\Delta^{m-1} + \sum_{i\in\mathcal{I}} \big\|\delta\tau^{i,m}\big\|_{\mathbb{L}^{1,1}} + \big\|\delta\tau^{i,m}\big\|_{\mathbb{L}^{1,1}}^{1-\varepsilon}. \end{align}\]

Proof. The result follows immediately from Lemma 3 and Lemma 4. ◻

3.3.3 Stopping-time error↩︎

We now estimate the stopping-time error \(\delta \tau^{i,m}\). The following proposition shows that this error is controlled by the Markov-map error, and disappears entirely in the case \(\mathcal{D}= \mathbb{R}^n\).

Proposition 2. For \(|\mathcal{D}|<\infty\), there exists \(C<\infty\) such that for all \(i\in\mathcal{I}\) and \(m\geq 1\), it holds \[\begin{align} \big\|\delta\tau^{i,m}\big\|_{\mathbb{L}^{1,1}} \leq C \bigg( \Big\| \delta\bar \zeta^{i,m,[\tau_{\min}^{i,m}]} \Big\|_{\mathbb{H}^{2,n}_{T}}^2 + \sum_{j\in\mathcal{I}\setminus\{i\}} \Big\| \delta \bar{\zeta}^{j,m-1,[\tau_{\min}^{i,m}]} \Big\|_{\mathbb{H}^{2,n}_{T}}^2 + E_\Delta^{m-1} \bigg)^{1/2}. \end{align}\] For \(\mathcal{D}=\mathbb{R}^n\), it holds \(\mathbb{P}\)-almost surely that \(\delta\tau^{i,m}=0\).

Proof. In the case \(|\mathcal{D}|<\infty\), by [38], the Cauchy–Schwarz inequality and taking the supremum over \([0,T]\) we have for \(i\in\mathcal{I}\) and \(m\geq1\) that \[\begin{align} \big\|\delta \tau^{i,m}\big\|_{\mathbb{L}^{1,1}} \lesssim \Big\|\delta\boldsymbol{X}_{\tau_{\min}^{i,m}}^{i,m}\Big\|_{\mathbb{L}^{1,n}} \leq \Big\|\delta\boldsymbol{X}_{\tau_{\min}^{i,m}}^{i,m}\Big\|_{\mathbb{L}^{2,n}} \leq \Big\|\delta\boldsymbol{X}^{i,m,[\tau_{\min}^{i,m}]}\Big\|_{\mathbb{S}^{2,n}_{T}}. \end{align}\] The claim now follows by applying Lemma 3. For \(\mathcal{D}= \mathbb{R}^n\), we have \(\partial\mathcal{D}=\emptyset\) and thus \(\tau=\tau^{i,m}\) \(\mathbb{P}\)-almost surely. In particular, \(\|\delta \tau^{i,m}\|_{\mathbb{L}^{1,1}}=0\). ◻

3.3.4 Markov-map error↩︎

We next turn to the Markov-map error. The analysis also involves the policy-mismatch error \(E_\Delta^m\), such that \(E_\zeta^m\) and \(E_\Delta^m\) must be estimated jointly. Lemmas 5 and 6 provide the Markov-map estimates under Assumptions [assump:gradient] and [assump:smalltime], respectively. Lemma 7 gives the corresponding bound for the policy error. Lemma 8 then shows that the resulting coupled estimate is contractive, and Proposition 3 summarizes the resulting convergence estimate.

Lemma 5. Consider the setting in Assumption [assump:gradient]. There exists \(\beta_0 \in (0,\infty)\) such that for every \(\beta > \beta_0\), there exist constants \(C^{\scalebox{0.55}{\mathrm{A}}}_\Delta(\beta) < \infty\) and \(\tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta) \in (0,1)\) for which, for all \(m \geq 1\), \[\begin{align} E^m_{\zeta, \beta} \leq \tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta) E_{\zeta,\beta}^{m-1} + C^{\scalebox{0.55}{\mathrm{A}}}_\Delta(\beta) E_{\Delta,\beta}^{m-1}. \end{align}\]

Proof. Let \(i \in \mathcal{I}\) and \(m \geq 1\) be fixed. To facilitate the analysis, we introduce an auxiliary FBSDE whose stochastic integral has integrand \(\zeta^{i,m}_t = \boldsymbol{\Sigma}^\top \mathrm{D}_{\boldsymbol{x}} V^{i,m}(t,\boldsymbol{X}_t)\). For notational convenience, we define \(V^{i,0} \equiv 0,\) and \(\zeta^{i,0} \equiv 0.\) We stress that \(V^{i,0}\) is not obtained from a fictitious-play PDE. Since the initialization of fictitious play is given by the zero policy, the coefficients \(b^{i,1}\) and \(\ell^{i,1}\) are independent of \(V^{i,0}\). The above definition is introduced solely to simplify the notation in the convergence analysis.

We introduce the notation \[\begin{align} \delta\hat{\boldsymbol{b}}^{i,m}_t \vcentcolon= \boldsymbol{b}\big(t,\boldsymbol{X}_t,\zeta^{i}_t,\boldsymbol{\zeta}^{-i}_t\big) - \boldsymbol{b}^{i,m}\big(t,\boldsymbol{X}_t,\zeta^{i,m}_t,\boldsymbol{\zeta}^{-i,m-1}_t\big), \quad t\in[0,T], \end{align}\] and let \(\Gamma_t^{i,m}\vcentcolon=\mathrm{D}_x V^{i,m}(t,\boldsymbol{X}_t)\). Itô’s lemma applied to \(\hat{Y}^{i,m}_t \vcentcolon=V^{i,m}(t,\boldsymbol{X}_t)\) yields \(\mathbb{P}\)-almost surely \[\begin{align} \hat{Y}_t^{i,m} &= g^i(\tau,\boldsymbol{X}_\tau) + \int_{t\wedge \tau}^\tau \ell^{i,m}\big(s,\boldsymbol{X}_s,\zeta^{i,m}_s,\boldsymbol{\zeta}^{-i,m-1}_s\big) \, \mathrm{d}s - \int_{t\wedge \tau}^\tau\big\langle \delta\hat{\boldsymbol{b}}_s^{i,m},\Gamma_s^{i,m}\big\rangle \, \mathrm{d}s - \int_{t \wedge \tau}^\tau(\zeta^{i,m}_s)^\top \text{d}\boldsymbol{W}_s. \end{align}\] For the corresponding difference process \(\delta\hat{Y}^{i,m}_t:=Y^i_t-\hat{Y}^{i,m}_t\), we introduce \[\begin{align} \delta\hat{\ell}^{i,m}_t \vcentcolon= f^{i}\big(t,\boldsymbol{X}_t,\zeta^{i}_t,\boldsymbol{\zeta}^{-i}_t\big) - \ell^{i,m}\big(t,\boldsymbol{X}_t,\zeta^{i,m}_t,\boldsymbol{\zeta}^{-i,m-1}_t\big), \quad t\in[0,T]. \end{align}\] Then, for all \(t \in [0,T]\), we have \(\mathbb{P}\)-almost surely \[\begin{align} \delta \hat{Y}_t^{i,m} &= \int_{t\wedge \tau}^\tau \delta\hat{\ell}^{i,m}_s \, \mathrm{d}s + \int_{t\wedge \tau}^\tau\big\langle \delta\hat{\boldsymbol{b}}_s^{i,m},\Gamma_s^{i,m}\big\rangle \, \mathrm{d}s - \int_{t \wedge \tau}^\tau(\delta\bar\zeta^{i,m}_s)^\top\text{d}\boldsymbol{W}_s. \end{align}\] For \(\beta>0\) we apply Itô’s lemma to \((t,y) \mapsto e^{\beta t}|y|^2\) along \(\delta\bar Y^{i,m}_t\) to obtain \(\mathbb{P}\)-almost surely \[\begin{align} \int_{0}^\tau\text{e}^{\beta s}\big\|\delta\bar\zeta_s^{i,m}\big\|^2_{\mathbb{R}^n} \, \mathrm{d}s &= 2\int_{0}^\tau \text{e}^{\beta s} \delta \hat{Y}_s^{i,m}\delta \hat{\ell}^{i,m}_s \, \mathrm{d}s + 2\int_{0}^\tau \text{e}^{\beta s} \delta \hat{Y}_s^{i,m}\big\langle \delta\hat{\boldsymbol{b}}_s^{i,m},\Gamma_s^{i,m}\big\rangle \, \mathrm{d}s\\ &\quad - \beta\int_{0}^\tau\text{e}^{\beta s}\big|\delta \hat{Y}_s^{i,m}\big|^2 \, \mathrm{d}s - 2\int_{0}^\tau \text{e}^{\beta s}\delta \hat{Y}_s^{i,m}\big(\delta\bar \zeta^{i,m}_s \big)^\top\text{d}\boldsymbol{W}_s - |\delta\hat{Y}_{0}^{i,m}|^2 . \end{align}\] Taking expectations, inserting absolute values, and dropping the non-positive \(\delta\hat{Y}_{0}^{i,m}\)-term, we get \[\begin{align} \Big\|\delta\bar\zeta^{i,m,[\tau]}\Big\|^2_{\mathbb{H}^{2,n}_{\beta,T}}\! &\leq 2\,\Big| \big\langle \delta \hat{Y}^{i,m},\delta \hat{\ell}^{i,m,[\tau]} \big\rangle_{\mathbb{H}^{2,1}_{\beta,T}} \Big| + 2\,\Big| \big\langle \delta \hat{Y}^{i,m}, \big\langle \delta\hat{\boldsymbol{b}}^{i,m,[\tau]},\Gamma^{i,m}\big\rangle \big\rangle_{\mathbb{H}^{2,1}_{\beta,T}} \Big| - \beta\big\|\delta \hat{Y}^{i,m}\big\|^2_{\mathbb{H}^{2,1}_{\beta,T}}. \end{align}\] For the first term, by Cauchy–Schwarz and Young’s inequality with parameter \(\beta/2\), we have \[\begin{align} 2 \Big| \big\langle \delta \hat{Y}^{i,m},\delta \hat{\ell}^{i,m,[\tau]} \big\rangle_{\mathbb{H}^{2,1}_{\beta,T}} \Big| \leq 2 \big\| \delta \hat{Y}^{i,m} \big\|_{\mathbb{H}^{2,1}_{\beta,T}} \big\|\delta \hat{\ell}^{i,m,[\tau]} \big\|_{\mathbb{H}^{2,1}_{\beta,T}} \leq \frac{\beta}{2} \big\| \delta \hat{Y}^{i,m} \big\|_{\mathbb{H}^{2,1}_{\beta,T}}^2 + \frac{2}{\beta} \big\|\delta \hat{\ell}^{i,m,[\tau]} \big\|_{\mathbb{H}^{2,1}_{\beta,T}}^2. \end{align}\] The second term is treated analogously. Combining the estimates and canceling the \(\delta \hat{Y}^{i,m}\)-terms yields \[\begin{align} \label{eq:thre95terms95functional} \big\|\delta\bar\zeta^{i,m,[\tau]}\big\|^2_{\mathbb{H}^{2,n}_{\beta,T}} \leq \frac{2}{\beta}\Big(\big\|\delta \hat{\ell}^{i,m,[\tau]}\big\|^2_{\mathbb{H}^{2,1}_{\beta,T}} + \big\|\big\langle \delta\hat{\boldsymbol{b}}^{i,m,[\tau]},\Gamma^{i,m}\big\rangle\big\|^2_{\mathbb{H}^{2,1}_{\beta,T}}\Big). \end{align}\tag{62}\] For the first term, note that we can split the coefficient difference as \[\begin{align} \big\|\delta \hat{\ell}^{i,m,[\tau]}\big\|^2_{\mathbb{H}^{2,1}_{\beta,T}} \leq 2\big\| \delta \bar{\ell}^{i,m} \big\|^2_{\mathbb{H}^{2,1}_{\beta,T}} + 2\big\| \ell^{i,m}\big(\cdot,\boldsymbol{X}_{\cdot},\zeta^{i},\boldsymbol{\zeta}^{-i}\big) - \ell^{i,m}\big(\cdot,\boldsymbol{X}_{\cdot},\zeta^{i,m},\boldsymbol{\zeta}^{-i,m-1}\big) \big\|^2_{\mathbb{H}^{2,1}_{\beta,T}}. \end{align}\] The first part has already been bounded in the proof of Lemma 2, such that \[\begin{align} \big\| \delta \bar{\ell}^{i,m} \big\|^2_{\mathbb{H}^{2,1}_{\beta,T}} \leq L_{\bar{\boldsymbol{f}}} (1 + L_\kappa) E_{\Delta,\beta}^{m-1}. \end{align}\] For the second part, we use the Lipschitz continuity of \(\ell^{i,m}\), which gives \[\begin{align} \big\| \ell^{i,m}\big(\cdot,\boldsymbol{X}_{\cdot},\zeta^{i},\boldsymbol{\zeta}^{-i}\big) - \ell^{i,m}\big(\cdot,\boldsymbol{X}_{\cdot},\zeta^{i,m},\boldsymbol{\zeta}^{-i,m-1}\big) \big\|^2_{\mathbb{H}^{2,1}_{\beta,T}} &\leq L_\ell^{\scalebox{0.55}{\mathrm{FP}}}\bigg( \big\|\delta \bar\zeta^{i,m,[\tau]}\big\|^2_{\mathbb{H}^{2,n}_{\beta,T}} + \sum_{j\in\mathcal{I}\setminus\{i\}}\big\|\delta \bar\zeta^{j,m-1,[\tau]}\big\|^2_{\mathbb{H}^{2,n}_{\beta,T}}\bigg). \end{align}\] In total, this gives the estimate \[\begin{align} \label{eq:lemma3695first95term} \big\|\delta \hat{\ell}^{i,m,[\tau]}\big\|^2_{\mathbb{H}^{2,1}_{\beta,T}} \leq 2e^{\beta T}L_{\bar{\boldsymbol{f}}} (1 + L_\kappa) E_\Delta^{m-1} + 2 L_\ell^{\scalebox{0.55}{\mathrm{FP}}}\bigg( \big\|\delta \bar\zeta^{i,m,[\tau]}\big\|^2_{\mathbb{H}^{2,n}_{\beta,T}} + \sum_{j\in\mathcal{I}\setminus\{i\}}\big\|\delta \bar\zeta^{j,m-1,[\tau]}\big\|^2_{\mathbb{H}^{2,n}_{\beta,T}}\bigg) \end{align}\tag{63}\] For the second term in 62 , we use Cauchy–Schwarz and the uniform boundedness 42 , which yields \[\begin{align} \begin{aligned} \big\|\big\langle \delta\hat{\boldsymbol{b}}^{i,m,[\tau]}, \Gamma^{i,m}\big\rangle\big\|^2_{\mathbb{H}^{2,1}_{\beta,T}} &= \mathbb{E}\bigg[\int_0^T e^{\beta t}\big |\big\langle \delta\hat{\boldsymbol{b}}^{i,m,[\tau]}_t, \Gamma_t^{i,m}\big\rangle\big|^2\, \mathrm{d}t\bigg]\\ &\leq \mathbb{E}\bigg[\int_0^T e^{\beta t}\big\| \delta\hat{\boldsymbol{b}}^{i,m,[\tau]}_t\big\|^2_{\mathbb{R}^n} \big\|\Gamma_t^{i,m} \big\|^2_{\mathbb{R}^n}\, \mathrm{d}t\bigg]\\ &\leq M^{\scalebox{0.55}{\mathrm{FP}}} \big\|\delta\hat{\boldsymbol{b}}^{i,m,[\tau]}\big\|^2_{\mathbb{H}^{2,n}_{\beta,T}}. \end{aligned} \label{eq:lemma3695second95term} \end{align}\tag{64}\] Using a similar split as for \(\delta\hat{\ell}^{i,m}\) and using the same arguments, we have the bound \[\begin{align} \big\|\delta\hat{\boldsymbol{b}}^{i,m,[\tau]}\big\|^2_{\mathbb{H}^{2,n}_{\beta,T}} \leq 2 e^{\beta T} L_{\bar{\boldsymbol{b}}} (1+L_\kappa) E_\Delta^{m-1} + 2L_b^{\scalebox{0.55}{\mathrm{FP}}}\bigg( \big\|\delta \bar\zeta^{i,m,[\tau]}\big\|^2_{\mathbb{H}^{2,n}_{\beta,T}} + \sum_{j\in\mathcal{I}\setminus\{i\}}\big\|\delta \bar\zeta^{j,m-1,[\tau]}\big\|^2_{\mathbb{H}^{2,n}_{\beta,T}}\bigg). \end{align}\] Inserting 63 and 64 into 62 , summing over \(i\in\mathcal{I}\) and rearranging terms, we have \[\begin{align} E^m_{\zeta, \beta} \leq C^{\scalebox{0.55}{\mathrm{A}}}_\Delta(\beta) E_{\Delta,\beta}^{m-1} + \tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta) E_{\zeta,\beta}^{m-1}, \end{align}\] where \[\begin{align} \begin{aligned} C^{\scalebox{0.55}{\mathrm{A}}}_\Delta(\beta) &\vcentcolon= \frac{4N(1+L_\kappa)(L_{\bar{\boldsymbol{f}}} + M^{\scalebox{0.55}{\mathrm{FP}}}L_{\bar{\boldsymbol{b}}})}{\beta}, \\ \tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta) &\vcentcolon= \frac{2(N-1)(L_\ell^{\scalebox{0.55}{\mathrm{FP}}} + M^{\scalebox{0.55}{\mathrm{FP}}} L_b^{\scalebox{0.55}{\mathrm{FP}}})}{\beta - 2(L_\ell^{\scalebox{0.55}{\mathrm{FP}}} +M^{\scalebox{0.55}{\mathrm{FP}}}L_b^{\scalebox{0.55}{\mathrm{FP}}})}. \end{aligned} \end{align}\] This completes the proof. ◻

Lemma 6. Consider the setting in Assumption [assump:smalltime]. For all \(m \geq 1\), there exist \(\tilde{K}_{\scalebox{0.55}{\mathrm{B}}}(\beta) \in (0,1)\) and \(C^{\scalebox{0.55}{\mathrm{B}}}_{\Delta}(\beta) < \infty\) such that \[\begin{align} E_{\zeta,\beta}^{m} \leq \tilde{K}_{\scalebox{0.55}{\mathrm{B}}}(\beta) E_{\zeta,\beta}^{m-1} + C^{\scalebox{0.55}{\mathrm{B}}}_{\Delta}(\beta) E_{\Delta,\beta}^{m-1}. \end{align}\]

Proof. We apply a classical FBSDE analysis based on the smallness condition in Assumption [assump:smalltime]. Let \(i\in \mathcal{I}\) and \(m \geq 1\) be fixed. By the decomposition 46 and the inequality 47 , we have \[\begin{align} \big\|\delta\bar\zeta^{i,m}\big\|^2_{\mathbb{H}^{2,n}_{\beta,T}} &\leq 2\,\big\|\delta\zeta^{i,m}\big\|^2_{\mathbb{H}^{2,n}_{\beta,T}} + 2\, \big\|\delta\tilde{\zeta}^{i,m}\big\|^2_{\mathbb{H}^{2,n}_{\beta,T}}. \label{eq:unbounded95lemma95two95terms} \end{align}\tag{65}\] We estimate the two terms on the right-hand side of 65 separately. For the first term, recall that \(\delta Y^{i,m} = Y^i-Y^{i,m}\) satisfies for all \(t\in[0,T]\), \(\mathbb{P}\)-almost surely \[\begin{align} \delta Y_t^{i,m} = \delta g^{i,m} + \int_t^T \delta \ell_s^{i,m} \, \mathrm{d}s - \int_t^T \big(\delta \zeta_s^{i,m}\big)^\top \,\text{d}\boldsymbol{W}_s . \end{align}\] For \(\beta > 0\), we apply Itô’s lemma to \((t,y)\mapsto e^{\beta t}\lvert y\rvert^2\) along \(\delta Y_t^{i,m}\), which gives \(\mathbb{P}\)-almost surely \[\begin{align} \int_0^T\text{e}^{\beta s}\big\| \delta \zeta_s^{i,m}\big\|_{\mathbb{R}^n}^2\, \mathrm{d}s &= \text{e}^{\beta T}\big| \delta g^{i,m}\big|^2 + \int_0^T \text{e}^{\beta s}\big( 2\,\delta Y_s^{i,m}\,\delta \ell_s^{i,m} - \beta \big| \delta Y_s^{i,m}\big|^2 \big)\,\text{d}s \\ &\quad - 2\int_0^T \text{e}^{\beta s}\,\delta Y_s^{i,m}\,\big(\delta \zeta_s^{i,m}\big)^\top \,\text{d}\boldsymbol{W}_s - \big| \delta Y_0^{i,m}\big|^2. \end{align}\] Applying Young’s inequality to \(2\,\delta Y_s^{i,m}\,\delta \ell_s^{i,m}\) with parameter \(\beta\), and discarding the final (non-positive) term, we get \[\begin{align} \int_0^T\text{e}^{\beta s}\big\| \delta \zeta_s^{i,m}\big\|_{\mathbb{R}^n}^2 \, \mathrm{d}s \leq \text{e}^{\beta T}\big| \delta g^{i,m}\big|^2 + \frac{1}{\beta}\int_0^T \text{e}^{\beta s}\big| \delta \ell_s^{i,m}\big|^2 \, \mathrm{d}s - 2\int_0^T \text{e}^{\beta s}\,\delta Y_s^{i,m}\, \big(\delta \zeta_s^{i,m}\big)^\top \,\text{d}\boldsymbol{W}_s . \end{align}\] Since \(\delta Y^{i,m} \in \mathbb{S}^{2,1}_{\beta,T}\) and \(\delta \zeta^{i,m} \in \mathbb{H}^{2,n}_{\beta,T}\), we have \(\delta Y^{i,m}\delta \zeta^{i,m} \in \mathbb{H}^{1,n}_{\beta,T}\) and by 6 the Itô integral has zero expectation. Taking expectations on both sides, we obtain \[\begin{align} \label{eq:delta95zeta95195new} \big\| \delta \zeta^{i,m}\big\|_{\mathbb{H}^{2,n}_{\beta,T}}^2 &\leq \mathbb{E}\Big[\text{e}^{\beta T}\big| \delta g^{i,m}\big|^2\Big] + \frac{1}{\beta}\big\| \delta \ell^{i,m}\big\|_{\mathbb{H}^{2,1}_{\beta,T}}^2 . \end{align}\tag{66}\] By the Lipschitz continuity of \(g\), we have \[\begin{align} \mathbb{E}\Big[\text{e}^{\beta T}\big| \delta g^{i,m}\big|^2\Big] &\leq L_g \, \mathbb{E}\Big[e^{\beta T}\big\|\delta\boldsymbol{X}^{i,m}_T\big\|_{\mathbb{R}^n}^2\Big] \leq L_g \mathbb{E}\bigg[\sup_{t\in[0,T]} e^{\beta t}\big\|\delta\boldsymbol{X}^{i,m}_t\big\|_{\mathbb{R}^n}^2\bigg] = L_g \, \big\|\delta\boldsymbol{X}^{i,m}\big\|_{\mathbb{S}^{2,n}_{\beta,T}}^2. \end{align}\] For the second term of 66 , we again split the error \(\delta \ell^{i,m}_t = \delta\bar{\ell}^{i,m}_{t} + \delta\tilde{\ell}^{i,m}_{t}\) and get \[\begin{align} \big\| \delta \bar{\ell}^{i,m} \big\|^2_{\mathbb{H}^{2,1}_{\beta,T}} \leq L_{\bar{\boldsymbol{f}}} (1 + L_\kappa) E_{\Delta,\beta}^{m-1}, \end{align}\] by repeating calculations from Lemma 2. Moreover, by the Lipschitz continuity of \(\ell^{i,m}\), the norm inequality 2 , and Lemma 1, we have \[\begin{align} \big\| \delta\tilde{\ell}^{i,m} \big\|^2_{\mathbb{H}^{2,1}_{\beta,T}} &\leq L_\ell^{\scalebox{0.55}{\mathrm{FP}}}\bigg( T\big\|\delta\boldsymbol{X}^{i,m}\big\|_{\mathbb{S}^{2,n}_{\beta,T}}^2 + \big\|\delta \zeta^{i,m}\big\|_{\mathbb{H}^{2,n}_{\beta,T}}^2 + \sum_{j\in\mathcal{I}\setminus\{i\}} \big\|\delta \zeta^{j,i,m-1,m}\big\|_{\mathbb{H}^{2,n}_{\beta,T}}^2 \bigg)\\ &\leq L_\ell^{\scalebox{0.55}{\mathrm{FP}}} T(1 + 2NL_\zeta^{\scalebox{0.55}{\mathrm{FP}}} ) \big\| \delta \boldsymbol{X}^{i,m} \big\|_{\mathbb{S}^{2,n}_{\beta,T}}^2 + 2L_\ell^{\scalebox{0.55}{\mathrm{FP}}} \big\| \delta \bar\zeta^{i,m} \big\|_{\mathbb{H}^{2,n}_{\beta,T}}^2 + 2L_\ell^{\scalebox{0.55}{\mathrm{FP}}} \sum_{j\in\mathcal{I}\setminus\{i\}} \big\| \delta \bar \zeta^{j,m-1} \big\|_{\mathbb{H}^{2,n}_{\beta,T}}^2. \end{align}\] Hence, using 47 , we get \[\begin{align} \big\| \delta \ell^{i,m}\big\|_{\mathbb{H}^{2,1}_{\beta,T}}^2 &\leq 2L_\ell^{\scalebox{0.55}{\mathrm{FP}}} T(1 + 2NL_\zeta^{\scalebox{0.55}{\mathrm{FP}}} ) \big\| \delta \boldsymbol{X}^{i,m} \big\|_{\mathbb{S}^{2,n}_{\beta,T}}^2 + 4L_\ell^{\scalebox{0.55}{\mathrm{FP}}} \big\| \delta \bar\zeta^{i,m} \big\|_{\mathbb{H}^{2,n}_{\beta,T}}^2 \\ &\qquad+ 4L_\ell^{\scalebox{0.55}{\mathrm{FP}}} \sum_{j\in\mathcal{I}\setminus\{i\}} \big\| \delta \bar \zeta^{j,m-1} \big\|_{\mathbb{H}^{2,n}_{\beta,T}}^2 + 2 L_{\bar{\boldsymbol{f}}} (1+L_\kappa) E_{\Delta,\beta}^{m-1}. \end{align}\] For the second term in 65 , the Lipschitz continuity of \(\zeta^{i,m}\) and 2 yield \[\begin{align} \label{eq:delta95zeta95295new} \big\|\delta\tilde{\zeta}^{i,m}\big\|^2_{\mathbb{H}^{2,n}_{\beta,T}} \leq L_\zeta^{\scalebox{0.55}{\mathrm{FP}}} T \big\|\delta\boldsymbol{X}^{i,m}\big\|^2_{\mathbb{S}^{2,n}_{\beta,T}}. \end{align}\tag{67}\] Inserting 66 and 67 into 65 gives \[\label{eq:smallness95ineq} \begin{align} \bigg(\frac{1}{2}-\frac{4L_\ell^{\scalebox{0.55}{\mathrm{FP}}}}{\beta} \bigg) \big\| \delta \bar \zeta^{i,m}\big\|_{\mathbb{H}^{2,n}_{\beta,T}}^2 &\leq \bigg(L_g+\frac{2L_\ell^{\scalebox{0.55}{\mathrm{FP}}} T(1+2L_\zeta^{\scalebox{0.55}{\mathrm{FP}}} N)}{\beta} +L_\zeta^{\scalebox{0.55}{\mathrm{FP}}} T\bigg) \big\|\delta\boldsymbol{X}^{i,m}\big\|_{\mathbb{S}^{2,n}_{\beta,T}}^2 \\ &\qquad + \frac{4L_\ell^{\scalebox{0.55}{\mathrm{FP}}} }{\beta}\sum_{j\in\mathcal{I}\setminus\{i\}} \big\|\delta \bar \zeta^{j,m-1}\big\|_{\mathbb{H}^{2,n}_{\beta,T}}^2 + \frac{2L_{\bar{\boldsymbol{f}}} (1+L_\kappa)}{\beta} E^{m-1}_{\Delta,\beta}. \end{align}\tag{68}\] Since \(\delta\tau^{i,m}=0\) \(\mathbb{P}\)-almost surely in the unbounded domain setting, the norm equivalence 1 followed by Lemma 3 gives \[\label{eq:dzetacalc4} \begin{align} &\bigg( \frac{1}{2} - C_\beta \bigg) \big\| \delta \bar \zeta^{i,m}\big\|_{\mathbb{H}^{2,n}_{\beta,T}}^2 \leq C_\beta \sum_{j\in\mathcal{I}\setminus\{i\}} \big\|\delta \bar \zeta^{j,m-1}\big\|_{\mathbb{H}^{2,n}_{\beta,T}}^2 + C_{\beta}' E_{\Delta,\beta}^{m-1}, \end{align}\tag{69}\] where \[\begin{align} C_\beta &\vcentcolon=\frac{4L_\ell^{\scalebox{0.55}{\mathrm{FP}}}}{\beta} +Ce^{\beta T} \bigg(L_g+\frac{2L_\ell^{\scalebox{0.55}{\mathrm{FP}}} T(1+2L_\zeta^{\scalebox{0.55}{\mathrm{FP}}} N)}{\beta} + L_\zeta^{\scalebox{0.55}{\mathrm{FP}}} T\bigg), \\ C_\beta' &\vcentcolon= \frac{2 L_{\bar{\boldsymbol{f}}} (1+L_\kappa)}{\beta} +Ce^{\beta T} \bigg(L_g+\frac{2L_\ell^{\scalebox{0.55}{\mathrm{FP}}} T(1+2L_\zeta^{\scalebox{0.55}{\mathrm{FP}}} N)}{\beta} + L_\zeta^{\scalebox{0.55}{\mathrm{FP}}} T\bigg). \end{align}\] Now define \(\tilde{K}_{\scalebox{0.55}{\mathrm{B}}}(\beta) \vcentcolon=\frac{2C_\beta}{1-2C_\beta}(N-1)\) and \(C^{\scalebox{0.55}{\mathrm{B}}}_{\Delta}(\beta) \vcentcolon=\frac{2C_\beta' N}{1-2C_\beta}\). By Assumption [assump:smalltime], it holds that \(\tilde{K}_{\scalebox{0.55}{\mathrm{B}}}(\beta) < 1\). Following rearrangement of terms and summation over \(i \in \mathcal{I}\), we have \[\begin{align} \begin{aligned} E_{\zeta,\beta}^{m} \leq \tilde{K}_{\scalebox{0.55}{\mathrm{B}}}(\beta) E_{\zeta,\beta}^{m-1} + C^{\scalebox{0.55}{\mathrm{B}}}_{\Delta}(\beta) E_{\Delta,\beta}^{m-1}. \end{aligned} \qedhere \end{align}\] ◻

Lemma 7. For every \(m \geq 1\), there exist \(C_\zeta(\eta) < \infty\) and \(\rho(\eta) < 1\) such that the policy error satisfies \[\begin{align} E_{\Delta,\beta}^{m} \leq \rho(\eta) E_{\Delta,\beta}^{m-1} + C_\zeta(\eta) E_{\zeta,\beta}^{m-1}. \end{align}\]

Proof. Fix \(m \geq 1\). For \((t,\boldsymbol{x}) \in Q_T\), we have by the Nash fixed-point relation, \[\begin{align} \alpha^{*,i}(t,\boldsymbol{x}) = \kappa^i \big( t, \boldsymbol{x}, \boldsymbol{\Phi}(t,\boldsymbol{x}) \zeta^i(t,\boldsymbol{x}), \boldsymbol{\alpha}^{*,-i}(t,\boldsymbol{x}) \big). \end{align}\] By the definition of \(\kappa^{i,m}\) and the contraction property 18 , we have \[\begin{align} \sum_{i \in \mathcal{I}} \big\| \Delta^{i,m}(t,\boldsymbol{x}) \big\|_{\mathbb{R}^d}^2 \leq L_{\kappa}^a \sum_{i \in \mathcal{I}} \big\| \alpha^{*,i}(t,\boldsymbol{x}) - \kappa^{i,m-1}(t,\boldsymbol{x},\zeta^{i,m-1}(t,\boldsymbol{x})) \big\|_{\mathbb{R}^d}^2. \end{align}\] For each \(i \in \mathcal{I}\), add and subtract \(\kappa^{i,m-1}(t,\boldsymbol{x},\zeta^{i}(t,\boldsymbol{x}))\). By Young’s inequality, for any \(\eta > 0\), we have \[\begin{align} \sum_{i \in \mathcal{I}} \big\| \Delta^{i,m}(t,\boldsymbol{x}) \big\|_{\mathbb{R}^d}^2 &\leq (1+\eta)L_{\kappa}^a \sum_{i \in \mathcal{I}} \big\| \Delta^{i,m-1}(t,\boldsymbol{x}) \big\|_{\mathbb{R}^d}^2 \\ &\qquad + \big( 1 + \eta^{-1} \big) L_{\kappa}^a \sum_{i \in \mathcal{I}} \big\| \kappa^{i,m-1}(t,\boldsymbol{x},\zeta^i(t,\boldsymbol{x})) - \kappa^{i,m-1}(t,\boldsymbol{x},\zeta^{i,m-1}(t,\boldsymbol{x})) \big\|_{\mathbb{R}^d}^2. \end{align}\] For the final sum, we apply the uniform Lipschitz continuity of \(\kappa^{i,m-1}\), such that \[\begin{align} \sum_{i \in \mathcal{I}} \big\| \Delta^{i,m}(t,\boldsymbol{x}) \big\|_{\mathbb{R}^d}^2 &\leq (1+\eta)L_{\kappa}^a \sum_{i \in \mathcal{I}} \big\| \Delta^{i,m-1}(t,\boldsymbol{x}) \big\|_{\mathbb{R}^d}^2 + \big( 1 + \eta^{-1} \big) L_{\kappa}^aL_{\kappa}^{\scalebox{0.55}{\mathrm{FP}}} \sum_{i \in \mathcal{I}} \big\| \zeta^i(t,\boldsymbol{x}) - \zeta^{i,m-1}(t,\boldsymbol{x}) \big\|_{\mathbb{R}^n}^2. \end{align}\] Evaluating at \(\boldsymbol{x}= \boldsymbol{X}_t\), multiplying by \(e^{\beta t}\), integrating over \([0,T]\), and taking expectation gives \[\begin{align} E^m_{\Delta,\beta} \leq (1+\eta)L_{\kappa}^a E_{\Delta,\beta}^{m-1} + \big( 1 +\eta^{-1} \big) L^a_{\kappa}L_{\kappa}^{\scalebox{0.55}{\mathrm{FP}}} E_{\zeta,\beta}^{m-1}. \end{align}\] Since \(L^a_\kappa < 1\), we can choose \(\eta > 0\) such that \[\begin{align} \rho(\eta) := (1 + \eta) L_\kappa^a < 1. \end{align}\] With \(C_\zeta(\eta) := (1+\eta^{-1})L_\kappa^a L_{\kappa}^{\scalebox{0.55}{\mathrm{FP}}}\), we obtain \[\begin{align} \begin{aligned} E_{\Delta,\beta}^{m} \leq \rho(\eta) E_{\Delta,\beta}^{m-1} + C_\zeta(\eta) E_{\zeta,\beta}^{m-1}. \end{aligned} \qedhere \end{align}\] ◻

Lemma 8. The matrices \[\begin{align} M_{\scalebox{0.55}{\mathrm{A}}} \vcentcolon= \begin{bmatrix} \tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta) & C_\Delta^{\scalebox{0.55}{\mathrm{A}}}(\beta) \\ C_\zeta(\eta) & \rho(\eta) \end{bmatrix}, \quad \text{ and } \quad M_{\scalebox{0.55}{\mathrm{B}}} \vcentcolon= \begin{bmatrix} \tilde{K}_{\scalebox{0.55}{\mathrm{B}}}(\beta) & C_\Delta^{\scalebox{0.55}{\mathrm{B}}}(\beta) \\ C_\zeta(\eta) & \rho(\eta) \end{bmatrix}, \end{align}\] under Assumption [assump:gradient] and Assumption [assump:smalltime], respectively, have spectral radius \(r(M_{\scalebox{0.55}{\mathrm{A}}}), r(M_{\scalebox{0.55}{\mathrm{B}}}) < 1\).

Proof. We begin by proving the result for \(M_{\scalebox{0.55}{\mathrm{A}}}\). Since all entries of \(M_{\scalebox{0.55}{\mathrm{A}}}\) are non-negative, its spectral radius is given by its Perron root. For a \(2 \times 2\) matrix with non-negative entries, this root is given by \[\begin{align} r(M_{\scalebox{0.55}{\mathrm{A}}}) = \frac{\tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta) + \rho(\eta) + \sqrt{(\tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta) - \rho(\eta))^2 + 4 C_\Delta^{\scalebox{0.55}{\mathrm{A}}}(\beta) C_\zeta(\eta)}}{2}. \end{align}\] The statement \(r(M_{\scalebox{0.55}{\mathrm{A}}}) < 1\) thus holds if we can show that \[\begin{align} \sqrt{(\tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta) - \rho(\eta))^2 + 4 C_\Delta^{\scalebox{0.55}{\mathrm{A}}}(\beta) C_\zeta(\eta)} < 2 - \tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta) - \rho(\eta). \end{align}\] Note that, since both \(\tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta) < 1\) and \(\rho(\eta) < 1\), the right-hand side is strictly positive. Squaring both sides and rearranging terms give the condition \[\begin{align} C_\Delta^{\scalebox{0.55}{\mathrm{A}}}(\beta) C_\zeta(\eta) < (1-\tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta))(1 - \rho(\eta)). \end{align}\] We note that, given an \(\eta\) such that \(\rho(\eta) < 1\), we can consider \(C_\zeta(\eta)\) and \(\rho(\eta)\) fixed. Since \(C^{\scalebox{0.55}{\mathrm{A}}}_\Delta(\beta)\) and \(\tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta)\) both vanish as \(\beta \rightarrow \infty\), it is always possible to choose \(\beta\) sufficiently large such that this condition holds. Thus, it holds that \(r(M_{\scalebox{0.55}{\mathrm{A}}}) < 1\). By Assumption [assump:smalltime], the same conditions are satisfied for \(M_{\scalebox{0.55}{\mathrm{B}}}\), and it follows that \(r(M_{\scalebox{0.55}{\mathrm{B}}}) < 1\). ◻

Proposition 3. Under either Assumption [assump:gradient] or [assump:smalltime], there exist \(\tilde{K} \in (0,1)\) and \(C < \infty\) such that for all \(m \geq 1\), it holds \[\begin{align} E_{\zeta}^{m} + E_{\Delta}^{m} \leq C \tilde{K}^m.\label{eq:prop95functional95error} \end{align}\qquad{(1)}\]

Proof. Under Assumption [assump:gradient], Lemma 5 combined with Lemma 7 gives the coupled bound \[\begin{align} E_{\zeta,\beta}^{m} &\leq \tilde{K}_{\scalebox{0.55}{\mathrm{A}}} E^{m-1}_{\zeta,\beta} + C_{\Delta}^{\scalebox{0.55}{\mathrm{A}}} E^{m-1}_{\Delta,\beta}, \\ E_{\Delta,\beta}^{m} &\leq C_\zeta(\eta) E^{m-1}_{\zeta,\beta} + \rho(\eta) E^{m-1}_{\Delta,\beta}. \end{align}\] On vector form, this can be written as \[\begin{align} U_m \leq M_{\scalebox{0.55}{\mathrm{A}}} U_{m-1} \leq M_{\scalebox{0.55}{\mathrm{A}}}^m U_0, \end{align}\] where \(U_m \vcentcolon=[E_{\zeta,\beta}^{m},E_{\Delta,\beta}^{m}]^\top\), and \(M_{\scalebox{0.55}{\mathrm{A}}}\) is the matrix from Lemma 8. Taking the induced \(\ell^1\)-norm on both sides, we have \[\begin{align} \| U_m \|_1 \leq \| M_{\scalebox{0.55}{\mathrm{A}}}^m \|_1 \| U_0 \|_1.\label{eq:Um95estimate} \end{align}\tag{70}\] By Lemma 8, it holds under Assumption [assump:gradient] that \(r(M_{\scalebox{0.55}{\mathrm{A}}}) < 1\). Thus, from the Gelfand formula, for every \(\tilde{K} \in (r(M_{\scalebox{0.55}{\mathrm{A}}}),1)\) there exists \(C_K \geq 1\) such that \[\begin{align} \|M_{\scalebox{0.55}{\mathrm{A}}}^m\|_1 \leq C_K \tilde{K}^m. \end{align}\] Inserting this into 70 gives \[\begin{align} E_{\zeta,\beta}^{m} + E_{\Delta,\beta}^{m} \leq C_K \tilde{K}^m \big( E_{\zeta,\beta}^{0} + E_{\Delta,\beta}^{0} \big). \end{align}\] The norm equivalence 1 finishes the proof. The proof under Assumption [assump:smalltime] follows analogously. ◻

3.3.5 Main theorem↩︎

We now conclude the convergence analysis by combining the results from Propositions 13. The following theorem states the main contribution of our paper.

Theorem 1. Let either Assumption [assump:gradient] or [assump:smalltime] hold. Then there exist \(K \in (0,1)\) and \(C < \infty\) such that for all \(m \geq 1\), it holds \[\begin{align} \label{eq:XYZ} \max_{i\in\mathcal{I}}\big\|\delta\boldsymbol{X}^{i,m}\big\|^2_{\mathbb{S}^{2,n}_{T}} +& \sum_{i\in\mathcal{I}}\big\|\delta Y^{i,m}\big\|^2_{\mathbb{S}^{2,1}_{T}} + \sum_{i\in\mathcal{I}}\Big\|\zeta^{i,[\tau]}(\cdot,\boldsymbol{X}_\cdot)-\zeta^{i,m,[\tau^{i,m}]}\big(\cdot,\boldsymbol{X}_\cdot^{i,m}\big)\Big\|_{\mathbb{H}^{2,n}_{T}}^2 \leq CK^{m}. \end{align}\tag{71}\]

Proof. Recall that \(\Psi^{m}\) from 43 is the left-hand side of 71 . From Proposition 1, it holds that \[\begin{align} \Psi^{m} \lesssim E_\zeta^{m} + E_\zeta^{m-1} + E_\Delta^{m-1} + \sum_{i\in\mathcal{I}} \big\|\delta\tau^{i,m}\big\|_{\mathbb{L}^{1,1}} + \big\|\delta\tau^{i,m}\big\|_{\mathbb{L}^{1,1}}^{1-\varepsilon}.\label{eq:mainthm95bound} \end{align}\tag{72}\] The last two terms can be bounded using Proposition 2. More precisely, for \(\alpha \in \{1, 1 - \varepsilon\}\), we have \[\begin{align} \big\| \delta\tau^{i,m} \big\|_{\mathbb{L}^{1,1}}^{\alpha} &\lesssim \bigg( \Big\| \delta\bar \zeta^{i,m,[\tau_{\min}^{i,m}]} \Big\|_{\mathbb{H}^{2,n}_{T}}^2 + \sum_{j\in\mathcal{I}\setminus\{i\}} \Big\| \delta \bar{\zeta}^{j,m-1,[\tau_\text{min}^{i,m}]} \Big\|_{\mathbb{H}^{2,n}_{T}}^2 + E_{\Delta}^{m-1} \bigg)^{\frac{\alpha}{2}}\\ &\leq \bigg( \Big\| \delta\bar \zeta^{i,m,[\tau_{\min}^{i,m}]} \Big\|_{\mathbb{H}^{2,n}_{T}}^2 \bigg)^{\frac{\alpha}{2}} + \sum_{j\in\mathcal{I}\setminus\{i\}} \bigg( \Big\| \delta \bar{\zeta}^{j,m-1,[\tau_\text{min}^{i,m}]} \Big\|_{\mathbb{H}^{2,n}_{T}}^2 \bigg)^{\frac{\alpha}{2}} + \big( E_{\Delta}^{m-1} \big)^{\frac{\alpha}{2}} \\ &\leq \bigg( \Big\| \delta\bar \zeta^{i,m,[\tau]} \Big\|_{\mathbb{H}^{2,n}_{T}}^2 \bigg)^{\frac{\alpha}{2}} + \sum_{j\in\mathcal{I}\setminus\{i\}} \bigg( \Big\| \delta \bar{\zeta}^{j,m-1,[\tau]} \Big\|_{\mathbb{H}^{2,n}_{T}}^2 \bigg)^{\frac{\alpha}{2}} + \big( E_{\Delta}^{m-1} \big)^{\frac{\alpha}{2}}, \end{align}\] where the second inequality follows from subadditivity of the function \(\varphi(y) = y^{\alpha/2}\), and the last follows from the fact that \(\tau^{i,m}_{\min} \leq \tau\). By summing over all players \(i \in \mathcal{I}\), the above inequality becomes \[\begin{align} \sum_{i\in\mathcal{I}} \big\|\delta\tau^{i,m}\big\|_{\mathbb{L}^{1,1}}^{\alpha} &\lesssim \sum_{i\in\mathcal{I}} \bigg( \Big\|\delta\bar \zeta^{i,m,[\tau]} \Big\|_{\mathbb{H}^{2,n}_{T}}^2\bigg)^{\frac{\alpha}{2}} + \sum_{i\in \mathcal{I}} \sum_{j\in\mathcal{I}\setminus\{i\}}\bigg( \Big\|\delta \bar{\zeta}^{j,m-1,[\tau]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2\bigg)^{\frac{\alpha}{2}} + N\big( E_\Delta^{m-1} \big)^{\frac{\alpha}{2}}. \label{eq:mainthm95before95jensen} \end{align}\tag{73}\] Next, recall that from Jensen’s inequality for concave functions it holds that \[\begin{align} \sum_{i=1}^N \varphi(y_i) \leq N \varphi\bigg( \frac{1}{N} \sum_{i=1}^N y_i \bigg). \end{align}\] Applying this to 73 yields \[\begin{align} \sum_{i\in\mathcal{I}} \big\|\delta\tau^{i,m}\big\|_{\mathbb{L}^{1,1}}^{\alpha} &\lesssim N^{1-\frac{\alpha}{2}} \bigg(\sum_{i\in\mathcal{I}} \Big\|\delta\bar \zeta^{i,m,[\tau]} \Big\|_{\mathbb{H}^{2,n}_{T}}^2\bigg)^{\frac{\alpha}{2}} + N(N-1)^{1-\frac{\alpha}{2}} \bigg(\sum_{j\in\mathcal{I}\setminus\{i\}} \Big\|\delta \bar{\zeta}^{j,m-1,[\tau]}\Big\|_{\mathbb{H}^{2,n}_{T}}^2\bigg)^{\frac{\alpha}{2}} + N\big( E_\Delta^{m-1} \big)^{\frac{\alpha}{2}}\\ &\lesssim \big( E_\zeta^m \big)^{\frac{\alpha}{2}} + \big( E_{\zeta}^{m-1} \big)^{\frac{\alpha}{2}} + \big( E_\Delta^{m-1}\big)^{\frac{\alpha}{2}}. \end{align}\] Inserting this bound into 72 and applying Proposition 3 gives for the case \(|\mathcal{D}|<\infty\) the bound \[\begin{align} \Psi^{m}& \lesssim \sum_{\alpha\in\{1-\varepsilon,1,2\}} \big( E_\zeta^m + E_\zeta^{m-1} + E_\Delta^{m-1} \big)^\frac{\alpha}{2} \lesssim \sum_{\alpha\in\{1-\varepsilon,1,2\}} \big( \tilde{K}^{m} + \tilde{K}^{m-1} \big)^{\frac{\alpha}{2}} \lesssim \tilde{K}^{\frac{(1-\varepsilon)m}{2}} = K^m, \end{align}\] where \(K=\tilde{K}^{\frac{1-\varepsilon}{2}}\). For the case \(\mathcal{D}=\mathbb{R}^n\), where \(\|\delta\tau^{i,m}\|_{\mathbb{L}^{1,1}}=0\), we instead have the bound \[\begin{align} \begin{aligned} \Psi^{m} \lesssim E_\zeta^m + E_\zeta^{m-1} + E_\Delta^{m-1} \lesssim \tilde{K}^m + \tilde{K}^{m-1} \lesssim K^m. \end{aligned} \qedhere \end{align}\] ◻

Remark 3. Theorem 1 establishes a geometric convergence rate of the fictitious-play scheme under both Assumption [assump:gradient] and Assumption [assump:smalltime]. Under Assumption [assump:gradient], this rate can be sharpened to a super-exponential rate in the special case where the best-response map does not depend on the opponents’ actions. This is the case, for instance, in Example 3.

Suppose that \(\kappa^i(t,\boldsymbol{x},p,\boldsymbol{a}^{-i})\) is independent of \(\boldsymbol{a}^{-i}\) for each player \(i \in \mathcal{I}\). Then, by definition, \(\Delta^i(t,\boldsymbol{x}) \equiv 0\), and in particular \(E_\Delta^m \equiv 0\). Consequently, Lemma 5, which measures the Markov-map error under Assumption [assump:gradient], reduces to \[\begin{align} E_\zeta^m \leq E_{\zeta,\beta}^m \leq \tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta) E_{\zeta,\beta}^{m-1} \leq \big(\tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta)\big)^m E_{\zeta,\beta}^{0} \leq e^{\beta T} \big(\tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta)\big)^m E_{\zeta}^0. \end{align}\] As in Lemma 5, \(\tilde{K}_{\scalebox{0.55}{\mathrm{A}}}(\beta)\) is of order \(\mathcal{O}(1/\beta)\). Choosing \(\beta\) of order \(m/T\) therefore gives \[\begin{align} E_\zeta^m \lesssim \exp(-cm\log(m)) \end{align}\] for some constant \(c > 0\). Inserting this estimate into the final error bound used in the proof of Theorem 1, and using \(E_\Delta^m \equiv 0\), gives the same super-exponential decay for the fictitious-play error.

4 Numerical example↩︎

We illustrate the convergence properties of fictitious play using the linear–quadratic interbank borrowing and lending model introduced in [1]. This example provides a convenient benchmark for the numerical method, since both the Nash equilibrium and the fictitious-play iterates admit explicit representations. In particular, this allows for a direct comparison between the numerical approximations and the corresponding analytical solutions.

We consider an \(N\)-player game, where the setting is a special case of the linear–quadratic game from Example 4. The state space is \(\mathcal{D}= \mathbb{R}^N\), and each player has action space \(A^1 = \cdots = A^N = \mathbb{R}\). Denoting by \(\mathbf{1} \in \mathbb{R}^N\) the vector of ones, we define the empirical mean \(\mu \colon \mathbb{R}^N \to \mathbb{R}\) by \[\begin{align} \mu(\boldsymbol{x}) \vcentcolon=\frac{1}{N} \mathbf{1}^\top \boldsymbol{x}, \end{align}\] and the corresponding vectorized function \(\boldsymbol{\mu}(\boldsymbol{x}) \vcentcolon=\mu(\boldsymbol{x}) \mathbf{1}\). Let \(\beta, \rho, \sigma, \delta, \varepsilon, \gamma \in \mathbb{R}\) be scalar parameters, whose values are specified below. For \((t,\boldsymbol{x}) \in Q_T\) and \(\boldsymbol{a}\in \mathbb{R}^N\), the drift coefficient is given by \[\begin{align} \bar{\boldsymbol{b}}(t,\boldsymbol{x},\boldsymbol{a}) = \beta ( \boldsymbol{\mu}(\boldsymbol{x}) - \boldsymbol{x} ) + \boldsymbol{a}, \end{align}\] and the diffusion coefficient by \[\begin{align} \boldsymbol{\Sigma}(t,\boldsymbol{x}) = \sigma \begin{pmatrix} \rho & \sqrt{1-\rho^2} & 0 & \cdots & 0\\ \rho & 0 & \sqrt{1-\rho^2} & \cdots & 0\\ \vdots & \vdots & \vdots & \ddots & \vdots\\ \rho & 0 & 0 & \cdots & \sqrt{1-\rho^2} \end{pmatrix}. \end{align}\] For \((t,\boldsymbol{x}) \in Q_T\) and \(\boldsymbol{a}\in \mathbb{R}^N\), the running and terminal costs of player \(i \in \mathcal{I}\) are given by \[\begin{align} \bar{f}^i(t,\boldsymbol{x},\boldsymbol{a}) &= \frac{1}{2} \big( a^i \big)^2 -\delta\,a^i \big( \mu(\boldsymbol{x})-x^i \big) +\frac{\varepsilon}{2} \big( \mu(\boldsymbol{x})-x^i \big)^2,\\ g^i(\boldsymbol{x}) &= \frac{\gamma}{2} \big(\mu(\boldsymbol{x})-x^i\big)^2. \end{align}\] Recall that the HJB equation admits a semi-analytical solution of the form 31 . Using the ansatz \[\begin{align} P^i_t = \eta_t \Bigl( \boldsymbol{e}_i - \frac{1}{N}\mathbf{1} \Bigr) \Bigl( \boldsymbol{e}_i - \frac{1}{N}\mathbf{1} \Bigr)^\top, \end{align}\] as a solution to the Riccati system, one can show that the value function of player \(i\) takes the form \[\begin{align} V^i(t,\boldsymbol{x}) = \frac{1}{2} \eta_t \left( \mu(\boldsymbol{x})-x^i \right)^2 + \nu_t, \end{align}\] where the scalar functions \(\eta_t\) and \(\nu_t\) satisfy \[\begin{align} \dot{\eta}_t &= \Bigl(1-\frac{1}{N^2}\Bigr)\eta_t^2 + 2(\beta+\delta)\eta_t - \varepsilon+\delta^2, \\ \nu_t &= \frac{1}{2}\,\sigma^2\big(1-\rho^2\big)\Bigl(1-\frac{1}{N}\Bigr) \int_t^T \eta_s\,\text{d}s, \end{align}\] with the terminal condition \(\eta_T = \gamma\). From this we obtain the semi-analytical reference solution for the processes \(\boldsymbol{Y}\) and \(\boldsymbol{Z}\). The forward process \(\boldsymbol{X}\) can then be solved for explicitly.

We initialize the fictitious-play scheme with \(\boldsymbol{\zeta}^{0} \equiv 0\). For \(m \geq 1\), we consider the fictitious-play HJB equations 41 with an ansatz analogous to the equilibrium case. This leads to a value function on the similar form \[\begin{align} V^{i,m}(t,\boldsymbol{x}) = \frac{1}{2} \eta^{m}_t \left( \mu(\boldsymbol{x})-x^i \right)^2 + \nu^{m}_t, \end{align}\] where the scalar function \(\eta^{m}_t\) satisfies \[\begin{align} \dot{\eta}^{m}_t &= \Bigl(1-\frac{1}{N}\Bigr)^2\big(\eta^{m}_t\big)^2 + 2\bigg(\beta+\frac{K^{m-1}(t)}{N} + \Bigl(1-\frac{1}{N}\Bigr)\delta\bigg)\eta^{m}_t - \varepsilon+\delta^2, \end{align}\] with \(K^{m}(t) \vcentcolon=\delta+(1-\frac{1}{N})\eta^{m}(t)\), and where \[\begin{align} \nu^{m}_t = \frac{1}{2}\,\sigma^2\big(1-\rho^2\big)\Bigl(1-\frac{1}{N}\Bigr) \int_t^T \eta^{m}_s\,\text{d}s. \end{align}\] This again yields a semi-analytical solution for the fictitious-play iterates. Since the derivation is standard in the linear–quadratic setting, we omit the details.

In the numerical example, we fix \[\begin{align} \beta = 0.1, \quad \delta = 0.2, \quad \varepsilon = 0.5, \quad \gamma = 0.5, \quad \rho = 0.2, \quad \sigma = 1, \quad T = 10, \end{align}\] and consider the cases \(N = 2\) and \(N = 20\). We define the error metrics by \[\begin{align} \mathrm{Err}_{\scalebox{0.65}{X}} \vcentcolon=\max_{i\in\mathcal{I}}\big\|\delta \boldsymbol{X}^{i,m}\big\|_{\mathbb{S}^{2,N}_{T}}^2, \quad \mathrm{Err}_{\scalebox{0.65}{Y}} \vcentcolon=\sum_{i\in\mathcal{I}}\big\|\delta Y^{i,m}\big\|_{\mathbb{S}^{2,1}_{T}}^2 \quad\textrm{and}\quad \mathrm{Err}_{\scalebox{0.65}{Z}} \vcentcolon=\sum_{i\in\mathcal{I}}\big\|\delta Z^{i,m}\big\|_{\mathbb{H}^{2,N+1}_{T}}^2. \end{align}\] Since these norms involve expectations, the reported values are their Monte Carlo estimators \(\widehat{\mathrm{Err}}_{\scalebox{0.65}{X}}\), \(\widehat{\mathrm{Err}}_{\scalebox{0.65}{Y}}\), and \(\widehat{\mathrm{Err}}_{\scalebox{0.65}{Z}}\).

The Riccati equation for \(\eta\) admits a closed-form solution, which we evaluate explicitly. For \(\eta^{m}\), we solve the Riccati equation on each time interval by using its explicit solution with the previous iterate frozen on that interval. The quantities \(\nu\) and \(\nu^{m}\) are approximated on the same uniform grid using the trapezoidal rule. For the Monte Carlo approximation of the Brownian motion, we use \(2^{6}\) samples, and we keep the same set of trajectories at every fictitious-play iteration so that the reported errors reflect only the dependence on the iteration index \(m\). We use \(5 \cdot 10^6\) time steps. The initial conditions are \(\boldsymbol{x}_0 = (-4.95, 1.84)^\top\) and \[\begin{align} \boldsymbol{x}_0 = \big( -&4.95,\, -1.84,\, 6.44,\, 0.97,\, 4.60,\, 2.89,\, -3.18,\, 2.71,\, -1.58,\, -1.61, \\ &0.49,\, -7.63,\, 5.96,\, -3.36,\, 5.00,\, 0.68,\, 7.66,\, -3.30,\, -1.56,\, 1.69 \big)^\top \end{align}\] for \(N = 2\) and \(N = 20\), respectively.

Figure 1: Convergence of the fictitious-play scheme for the interbank model in the cases N=2 (left) and N=20 (right).

Figure 1 shows the convergence results for the numerical example. For both population sizes and all error metrics, the fictitious-play iterates exhibit clear exponential convergence until the errors reach machine precision, or values close to it. For \(N = 2\) and \(N = 20\), the errors decrease by factors of approximately \(23\) and \(1000\) per iteration, respectively. The faster convergence for \(N=20\) is consistent with the weaker coupling between players as the population size increases. Indeed, in the limit \(N \to \infty\), the game converges to a mean-field control problem. These results are consistent with our theoretical findings, even though this example does not satisfy the assumptions under which the convergence theory is established. This suggests that fictitious play may converge beyond the range of our theoretical assumptions.

5 Notes on the Lipschitz condition↩︎

In Section 3.2, we introduced the Lipschitz condition (L) for the functions \(\boldsymbol{b}\), \(\boldsymbol{f}\), \(\boldsymbol{b}^{i,m}\), \(\ell^{i,m}\) and \(\kappa^{i,m}\). Under Assumption [assump:smalltime], this condition is imposed directly. The purpose of this appendix is to verify that the local version of (L) follows under Assumption [assump:gradient], and to identify a simple setting in which the global version follows from the standard Lipschitz assumptions.

The obstruction to global Lipschitz continuity is the variable \(\boldsymbol{\Phi}(t,\boldsymbol{x})\boldsymbol{z}\). The estimates below show that condition (L) holds in either of the following two cases:

  1. The relevant \(\boldsymbol{z}\) arguments are uniformly bounded. That is, \(\boldsymbol{z}\) belongs to the set of matrices with norm bounded by some \(R_{\boldsymbol{z}} < \infty\). Under Assumption [assump:gradient], this follows from the gradient bound.

  2. The noise is additive and time-independent, so that \(\boldsymbol{\Phi}\) is independent of \((t,\boldsymbol{x})\).

Throughout this section, we use the compact notation \[\begin{align} D_{t\boldsymbol{x}} = |t_1 - t_2| + \|\boldsymbol{x}_1 - \boldsymbol{x}_2\|^2_{\mathbb{R}^n}, \qquad D_{\boldsymbol{z}} = \vert\!\vert\!\vert \boldsymbol{z}_1 - \boldsymbol{z}_2 \vert\!\vert\!\vert^2_{\mathbb{R}^{n \times N}}, \qquad D_{\boldsymbol{a}^{-i}} = \vert\!\vert\!\vert \boldsymbol{a}^{-i}_1 - \boldsymbol{a}^{-i}_2 \vert\!\vert\!\vert^2_{\mathbb{R}^{d \times (N-1)}}, \end{align}\] and \(D_{z}\) the adaptation of \(D_{\boldsymbol{z}}\) onto \(\mathbb{R}^n\). We show that the functions satisfy the bounds \[\begin{align} \begin{aligned} \|\boldsymbol{b}(t_1,\boldsymbol{x}_1,\boldsymbol{z}_1)-\boldsymbol{b}(t_2,\boldsymbol{x}_2,\boldsymbol{z}_2)\|_{\mathbb{R}^{n}}^2 &\leq L_{\boldsymbol{b}} \big( D_{t\boldsymbol{x}} + D_{\boldsymbol{z}} \big),\\ \|\boldsymbol{f}(t_1,\boldsymbol{x}_1,\boldsymbol{z}_1)-\boldsymbol{f}(t_2,\boldsymbol{x}_2,\boldsymbol{z}_2)\|_{\mathbb{R}^{N}}^2 &\leq L_{\boldsymbol{f}} \big( D_{t\boldsymbol{x}} + D_{\boldsymbol{z}} \big),\\ \big\| \kappa^{i,m}(t_1,\boldsymbol{x}_1,z_1) - \kappa^{i,m}(t_2,\boldsymbol{x}_2,z_2) \big\|_{\mathbb{R}^d}^2 &\leq L_\kappa^{\scalebox{0.55}{\mathrm{FP}}} \big( D_{t\boldsymbol{x}} + D_{z} \big),\\ \big\| \boldsymbol{b}^{i,m}(t_1,\boldsymbol{x}_1,z_1^i,\boldsymbol{z}_1^{-i})-\boldsymbol{b}^{i,m}(t_2,\boldsymbol{x}_2,z^i_2,\boldsymbol{z}_2^{-i}) \big\|_{\mathbb{R}^{n}}^2 &\leq L_{b}^{\text{FP}} \big( D_{t\boldsymbol{x}} + D_{\boldsymbol{z}} \big),\\ \big| \ell^{i,m}(t_1,\boldsymbol{x}_1,z_1^i,\boldsymbol{z}_1^{-i}) - \ell^{i,m}(t_2,\boldsymbol{x}_2,z_2^i,\boldsymbol{z}_2^{-i}) \big|^2 &\leq L_{\ell}^{\text{FP}} \big( D_{t\boldsymbol{x}} + D_{\boldsymbol{z}} \big). \end{aligned}\label{eq:lipschtiz95bounds} \end{align}\tag{74}\] The following lemma shows the Lipschitz continuity for the functions \(\boldsymbol{b}\) and \(\boldsymbol{f}\).

Lemma 9. If case (I) holds, the functions \(\boldsymbol{b}\) and \(\boldsymbol{f}\) satisfy the bounds in 74 with constants \[\begin{align} L_{\boldsymbol{b}} &= L_{\bar{\boldsymbol{b}}} \max \big\{ 1 + L_{\boldsymbol{c}} + 2L_{\boldsymbol{c}}L_{\boldsymbol{\Phi}}R_{\boldsymbol{z}}^2, 2L_{\boldsymbol{c}}C_{\boldsymbol{\Phi}}^2 \big\}, \\ L_{\boldsymbol{f}} &= L_{\bar{\boldsymbol{f}}} \max \big\{ 1 + L_{\boldsymbol{c}} + 2L_{\boldsymbol{c}}L_{\boldsymbol{\Phi}}R_{\boldsymbol{z}}^2, 2L_{\boldsymbol{c}}C_{\boldsymbol{\Phi}}^2 \big\}. \end{align}\] Alternatively, if case (II) holds, the functions \(\boldsymbol{b}\) and \(\boldsymbol{f}\) satisfy the bounds with constants \[\begin{align} L_{\boldsymbol{b}} &= L_{\bar{\boldsymbol{b}}} \max \big\{ 1 + L_{\boldsymbol{c}}, 2L_{\boldsymbol{c}}C_{\boldsymbol{\Phi}}^2 \big\}, \\ L_{\boldsymbol{f}} &= L_{\bar{\boldsymbol{f}}} \max \big\{ 1 + L_{\boldsymbol{c}}, 2L_{\boldsymbol{c}}C_{\boldsymbol{\Phi}}^2 \big\}. \end{align}\]

Proof. Using the definition of \(\boldsymbol{b}\), and the Lipschitz continuity of \(\bar{\boldsymbol{b}}\) and \(\boldsymbol{c}\), we obtain \[\begin{align} &\big\| \boldsymbol{b}(t_1, \boldsymbol{x}_1, \boldsymbol{z}_1) - \boldsymbol{b}(t_2, \boldsymbol{x}_2, \boldsymbol{z}_2) \big\|_{\mathbb{R}^n}^2 \\ &\qquad= \big\| \bar{\boldsymbol{b}}(t_1, \boldsymbol{x}_1, \boldsymbol{c}(t_1,\boldsymbol{x}_1,\boldsymbol{\Phi}(t_1,\boldsymbol{x}_1)\boldsymbol{z}_1)) - \bar{\boldsymbol{b}}(t_2, \boldsymbol{x}_2, \boldsymbol{c}(t_2,\boldsymbol{x}_2,\boldsymbol{\Phi}(t_2,\boldsymbol{x}_2)\boldsymbol{z}_2)) \big\|^2_{\mathbb{R}^n} \\ &\qquad\leq L_{\bar{\boldsymbol{b}}} \big( D_{t\boldsymbol{x}} + \vert\!\vert\!\vert \boldsymbol{c}(t_1,\boldsymbol{x}_1,\boldsymbol{\Phi}(t_1,\boldsymbol{x}_1)\boldsymbol{z}_1)) - \boldsymbol{c}(t_2,\boldsymbol{x}_2,\boldsymbol{\Phi}(t_2,\boldsymbol{x}_2)\boldsymbol{z}_2)) \vert\!\vert\!\vert_{\mathbb{R}^{d\times N}}^2 \big) \\ &\qquad\leq L_{\bar{\boldsymbol{b}}} \Big( D_{t\boldsymbol{x}} + L_{\boldsymbol{c}} \big( D_{t\boldsymbol{x}} + \vert\!\vert\!\vert \boldsymbol{\Phi}(t_1,\boldsymbol{x}_1)\boldsymbol{z}_1 - \boldsymbol{\Phi}(t_2,\boldsymbol{x}_2)\boldsymbol{z}_2 \vert\!\vert\!\vert^2_{\mathbb{R}^{n \times N}} \big) \Big). \end{align}\] For the final term, adding and subtracting \(\boldsymbol{\Phi}(t_1,\boldsymbol{x}_1)\boldsymbol{z}_2\) gives \[\begin{align} \vert\!\vert\!\vert \boldsymbol{\Phi}(t_1,\boldsymbol{x}_1)\boldsymbol{z}_1 - \boldsymbol{\Phi}(t_2,\boldsymbol{x}_2)\boldsymbol{z}_2 \vert\!\vert\!\vert^2_{\mathbb{R}^{n \times N}} &\leq 2 \vert\!\vert\!\vert \boldsymbol{\Phi}(t_1,\boldsymbol{x}_1)(\boldsymbol{z}_1 - \boldsymbol{z}_2) \vert\!\vert\!\vert^2_{\mathbb{R}^{n \times N}} + 2 \vert\!\vert\!\vert \big( \boldsymbol{\Phi}(t_1, \boldsymbol{x}_1) - \boldsymbol{\Phi}(t_2, \boldsymbol{x}_2) \big) \boldsymbol{z}_2 \vert\!\vert\!\vert^2_{\mathbb{R}^{n \times N}} \\ &\leq 2C_{\boldsymbol{\Phi}}^2 D_{\boldsymbol{z}} + 2L_{\boldsymbol{\Phi}}D_{t\boldsymbol{x}} \vert\!\vert\!\vert \boldsymbol{z}_2 \vert\!\vert\!\vert_{\mathbb{R}^{n\times N}}^2. \end{align}\] Consequently, we have \[\begin{align} \big\| \boldsymbol{b}(t_1, \boldsymbol{x}_1, \boldsymbol{z}_1) - \boldsymbol{b}(t_2, \boldsymbol{x}_2, \boldsymbol{z}_2) \big\|_{\mathbb{R}^n}^2 &\leq L_{\bar{\boldsymbol{b}}} \big( D_{t\boldsymbol{x}} + L_{\boldsymbol{c}} \big( D_{t\boldsymbol{x}} + 2C_{\boldsymbol{\Phi}}^2 D_{\boldsymbol{z}} + 2L_{\boldsymbol{\Phi}}D_{t\boldsymbol{x}}\vert\!\vert\!\vert \boldsymbol{z}_2 \vert\!\vert\!\vert_{\mathbb{R}^{n\times N}}^2 \big) \big) \\ &\leq L_{\bar{\boldsymbol{b}}} \big( (1 + L_{\boldsymbol{c}} + 2L_{\boldsymbol{c}}L_{\boldsymbol{\Phi}}\vert\!\vert\!\vert \boldsymbol{z}_2 \vert\!\vert\!\vert_{\mathbb{R}^{n\times N}}^2) D_{t\boldsymbol{x}} + 2L_{\boldsymbol{c}}C_{\boldsymbol{\Phi}}^2 D_{\boldsymbol{z}} \big). \end{align}\] Thus, if case (I) holds, then \(\vert\!\vert\!\vert \boldsymbol{z}_2 \vert\!\vert\!\vert_{\mathbb{R}^{n \times N}} \leq R_{\boldsymbol{z}}\), and \(\boldsymbol{b}\) is Lipschitz with constant \[\begin{align} L_{\boldsymbol{b}} = L_{\bar{\boldsymbol{b}}} \max \big\{ 1 + L_{\boldsymbol{c}} + 2L_{\boldsymbol{c}}L_{\boldsymbol{\Phi}}R_{\boldsymbol{z}}^2, 2L_{\boldsymbol{c}}C_{\boldsymbol{\Phi}}^2 \big\}. \end{align}\] Alternatively, if case (II) holds, then \(\boldsymbol{\Phi}\) is independent of \((t,\boldsymbol{x})\) and \(L_{\boldsymbol{\Phi}}=0\). The bound then becomes global with constant \[\begin{align} L_{\boldsymbol{b}} = L_{\bar{\boldsymbol{b}}} \max \big\{ 1 + L_{\boldsymbol{c}}, 2L_{\boldsymbol{c}}C_{\boldsymbol{\Phi}}^2 \big\}. \end{align}\] The same arguments apply to \(\boldsymbol{f}\), with \(L_{\bar{\boldsymbol{b}}}\) replaced by \(L_{\bar{\boldsymbol{f}}}\). ◻

Next, we focus on the Lipschitz continuity of \(\kappa^{i,m}\). The following lemma first shows Lipschitz continuity of \(\boldsymbol{\alpha}^m\), from which the result for \(\kappa^{i,m}\) follows.

Lemma 10. If case (I) or (II) holds, then there exists \(L_\alpha > 0\), independent of \(m\), such that \[\begin{align} \vert\!\vert\!\vert \boldsymbol{\alpha}^m(t_1, \boldsymbol{x}_1) - \boldsymbol{\alpha}^m(t_2, \boldsymbol{x}_2) \vert\!\vert\!\vert^2_{\mathbb{R}^{d \times N}} \leq L_\alpha D_{t\boldsymbol{x}}, \end{align}\] and, for all \(i\in \mathcal{I}\), \(m\geq 1\), the functions \(\kappa^{i,m}\) satisfy 74 .

Proof. We introduce the compact notation \[\begin{align} p^i_k = \boldsymbol{\Phi}(t_k, \boldsymbol{x}_k) \zeta^{i,m}(t_k, \boldsymbol{x}_k), \qquad \alpha^{i,m}_k = \alpha^{i,m}(t_k, \boldsymbol{x}_k), \end{align}\] for \(k = 1,2\). We want to show that \[\begin{align} A_m \vcentcolon=\sum_{i \in \mathcal{I}} \| \alpha^{i,m}_1 - \alpha^{i,m}_2 \|^2_{\mathbb{R}^d} \leq L_\alpha D_{t\boldsymbol{x}}, \end{align}\] for all \(m\geq 1\). Using the definition of \(\alpha^{i,m}\), we have \[\begin{align} A_m &= \sum_{i \in \mathcal{I}} \big\| \kappa^i \big( t_1, \boldsymbol{x}_1, p^i_1, \boldsymbol{\alpha}^{-i,m-1}_1 \big) - \kappa^i \big( t_2, \boldsymbol{x}_2, p^i_2, \boldsymbol{\alpha}^{-i,m-1}_2 \big) \big\|^2_{\mathbb{R}^d} \\ &\leq (1 + \eta) \sum_{i \in \mathcal{I}} \big\| \kappa^i \big( t_1, \boldsymbol{x}_1, p^i_1, \boldsymbol{\alpha}^{-i,m-1}_1 \big) - \kappa^i \big( t_1, \boldsymbol{x}_1, p^i_1, \boldsymbol{\alpha}^{-i,m-1}_2 \big) \big\|^2_{\mathbb{R}^d} \\ &\qquad + (1 + \eta^{-1}) \sum_{i \in \mathcal{I}} \big\| \kappa^i \big( t_1, \boldsymbol{x}_1, p^i_1, \boldsymbol{\alpha}^{-i,m-1}_2 \big) - \kappa^i \big( t_2, \boldsymbol{x}_2, p^i_2, \boldsymbol{\alpha}^{-i,m-1}_2 \big) \big\|^2_{\mathbb{R}^d}, \end{align}\] where the second step is Young’s inequality. The first term is controlled by the joint contraction \[\begin{align} \sum_{i \in \mathcal{I}} \big\| \kappa^i \big( t_1, \boldsymbol{x}_1, p^i_1, \boldsymbol{\alpha}^{-i,m-1}_1 \big) - \kappa^i \big( t_1, \boldsymbol{x}_1, p^i_1, \boldsymbol{\alpha}^{-i,m-1}_2 \big) \big\|^2_{\mathbb{R}^d} \leq L_\kappa^a \sum_{i \in \mathcal{I}} \big\| \alpha^{i,m-1}_1 - \alpha^{i,m-1}_2 \big\|^2_{\mathbb{R}^d} = L_\kappa^a A_{m-1}. \end{align}\] For the second term, the Lipschitz continuity of \(\kappa^i\) gives \[\begin{align} \sum_{i \in \mathcal{I}} \big\| \kappa^i \big( t_1, \boldsymbol{x}_1, p^i_1, \boldsymbol{\alpha}^{-i,m-1}_2 \big) - \kappa^i \big( t_2, \boldsymbol{x}_2, p^i_2, \boldsymbol{\alpha}^{-i,m-1}_2 \big) \big\|^2_{\mathbb{R}^d} \leq \sum_{i \in \mathcal{I}} L_{\kappa}\big(D_{t\boldsymbol{x}} + \big\| p^i_1 - p^i_2 \big\|_{\mathbb{R}^n}^2 \big). \end{align}\] As in the proof of Lemma 9, we get \[\begin{align} \big\| p^i_1 - p^i_2 \big\|_{\mathbb{R}^n}^2 &= \big\| \boldsymbol{\Phi}(t_1, \boldsymbol{x}_1) \zeta^{i,m}(t_1, \boldsymbol{x}_1) - \boldsymbol{\Phi}(t_2, \boldsymbol{x}_2) \zeta^{i,m}(t_2, \boldsymbol{x}_2) \big\|_{\mathbb{R}^n}^2 \\ &\leq 2C^2_{\boldsymbol{\Phi}} \big\| \zeta^{i,m}(t_1, \boldsymbol{x}_1) - \zeta^{i,m}(t_2, \boldsymbol{x}_2) \big\|^2_{\mathbb{R}^n} + 2 L_{\boldsymbol{\Phi}} D_{t\boldsymbol{x}} \big\| \zeta^{i,m}(t_2, \boldsymbol{x}_2) \big\|^2_{\mathbb{R}^n}. \end{align}\] For the first term we use the uniform Lipschitz bound for \(\zeta^{i,m}\). Under case (I), we have a uniform bound on \(\zeta^{i,m}\) for the second term. Alternatively, under case (II), we have \(L_{\boldsymbol{\Phi}} = 0\) and the second term vanishes completely. This gives the bound \[\begin{align} \| p_1^i - p_2^i \|^2_{\mathbb{R}^n} \leq L_p D_{t\boldsymbol{x}}. \label{eq:phiz95bound} \end{align}\tag{75}\] Thus, the bound for \(A_m\) becomes \[\begin{align} A_m \leq (1+\eta)L_\kappa^a A_{m-1} + (1+\eta^{-1})L_\kappa \sum_{i\in\mathcal{I}} \big( D_{t\boldsymbol{x}} + L_p D_{t\boldsymbol{x}} \big) = q A_{m-1} + BD_{t\boldsymbol{x}}, \end{align}\] where we have defined \[\begin{align} q \vcentcolon=(1+\eta)L_\kappa^a, \qquad B \vcentcolon=(1+\eta^{-1})L_\kappa N(1+L_p), \end{align}\] and choose \(\eta\) such that \(q < 1\). Iterating the bound, we get \[\begin{align} A_m &\leq q^m A_0 + \Bigg( \sum_{r=0}^{m-1} q^r \Bigg) B D_{t\boldsymbol{x}} \leq \frac{B}{1-q} D_{t\boldsymbol{x}}, \end{align}\] where in the last inequality we use \(A_0=0\) by zero-policy initialization, and the geometric sum formula. We thus have Lipschitz continuity of \(\boldsymbol{\alpha}^{m}\) for all \(m\), with \[\begin{align} L_\alpha \vcentcolon=\frac{B}{1-q}. \end{align}\] This concludes the \(\boldsymbol{\alpha}^m\)-part of the lemma. The Lipschitz continuity of \(\kappa^{i,m}\) now follows by using the Lipschitz continuity of \(\kappa^i\), the same type of estimate for the \(\boldsymbol{\Phi}(t_k,\boldsymbol{x}_k)z_k\) terms as above, and the Lipschitz continuity of \(\boldsymbol{\alpha}^m\), \[\begin{align} \big\| \kappa^{i,m}(t_1, \boldsymbol{x}_1, z_1) - \kappa^{i,m}(t_2, \boldsymbol{x}_2, z_2) \big\|^2_{\mathbb{R}^d} &\leq \big\| \kappa^{i}(t_1, \boldsymbol{x}_1, \boldsymbol{\Phi}_1z_1, \boldsymbol{\alpha}^{-i,m-1}_1) - \kappa^{i}(t_2, \boldsymbol{x}_2, \boldsymbol{\Phi}_2z_2, \boldsymbol{\alpha}^{-i,m-1}_2) \big\|^2_{\mathbb{R}^d} \\ &\leq L_{\kappa^i} \big( D_{t\boldsymbol{x}} + \big\| \boldsymbol{\Phi}_1 z_1 - \boldsymbol{\Phi}_2 z_2 \big\|_{\mathbb{R}^n}^2 \big) \!+\! L_\kappa^{a} \big|\!\big|\!\big| \boldsymbol{\alpha}^{-i,m-1}_1 - \boldsymbol{\alpha}^{-i,m-1}_2 \big|\!\big|\!\big|_{\mathbb{R}^{d\times (N-1)}}^2 \\ &\leq L_{\kappa}^{\scalebox{0.55}{\mathrm{FP}}} \big( D_{t\boldsymbol{x}} + D_{z} \big). \end{align}\] ◻

It remains to show the Lipschitz continuity of \(\boldsymbol{b}^{i,m}\) and \(\ell^{i,m}\). This requires Lipschitz continuity of the coefficient functions \(\tilde{\boldsymbol{b}}^i\) and \(\tilde{\ell}^i\), which follows under similar calculations and assumptions as for \(\boldsymbol{b}^i\) and \(\ell^i\) in Lemma 9. The following lemma concludes the appendix.

Lemma 11. If case (I) or (II) holds, then the functions \(\boldsymbol{b}^{i,m}\) and \(\ell^{i,m}\) satisfy 74 .

Proof. Using the definition of \(\boldsymbol{b}^{i,m}\) and the Lipschitz continuity of \(\kappa^{i,m}\), we obtain \[\begin{align} &\big\| \boldsymbol{b}^{i,m}(t_1, \boldsymbol{x}_1, z^i_1, \boldsymbol{z}_1^{-i}) - \boldsymbol{b}^{i,m}(t_2, \boldsymbol{x}_2, z^i_2, \boldsymbol{z}_2^{-i}) \big\|^2 \\ &\qquad = \big\| \tilde{\boldsymbol{b}}^{i}(t_1, \boldsymbol{x}_1, z^i_1, \boldsymbol{\kappa}^{-i,m-1}(t_1, \boldsymbol{x}_1, \boldsymbol{z}_1^{-i})) - \tilde{\boldsymbol{b}}^{i}(t_2, \boldsymbol{x}_2, z^i_2,\boldsymbol{\kappa}^{-i,m-1}(t_2, \boldsymbol{x}_2, \boldsymbol{z}_2^{-i})) \big\|^2 \\ &\qquad \leq L_{\tilde{\boldsymbol{b}}^i} \bigg( D_{t\boldsymbol{x}} + \| z^i_1 - z_2^i \|^2_{\mathbb{R}^n} + L_\kappa^{\scalebox{0.55}{\mathrm{FP}}} \bigg( (N-1) D_{t\boldsymbol{x}} + \sum_{i\in\mathcal{I}\backslash\{i\}} \| z^j_1 - z^j_2 \|^2_{\mathbb{R}^n} \bigg) \bigg) \\ &\qquad \leq L^{\scalebox{0.55}{\mathrm{FP}}}_{\boldsymbol{b}} \left( D_{t\boldsymbol{x}} + D_{\boldsymbol{z}} \right). \end{align}\] The same argument gives Lipschitz continuity for \(\ell^{i,m}\). ◻

6 Synthesis of the LQ-problem↩︎

This section contains a detailed derivation of the analytical solution to the LQG problem presented in Example 4. To lighten the notation, the explicit dependence on \((t,\boldsymbol{x})\) is suppressed. The value function \(V^i\) satisfies the HJB equation \[\begin{align} \begin{aligned} \partial_t V^i &+ \frac{1}{2} \mathrm{Tr} \left( \mathbf{C} \mathrm{D}_{\boldsymbol{x}}^2 V^i \mathbf{C}^\top \right) + \left\langle F \left( \boldsymbol{\xi} - \boldsymbol{x} \right), \mathrm{D}_{\boldsymbol{x}} V^i \right\rangle + \big\| K^{i}_{\boldsymbol{x}} \boldsymbol{x} \big\|_{\mathbb{R}^n}^2 \\ &\quad + \sum_{j\in \mathcal{I}} \left\langle \tilde{\Phi}^{j}, \left( B^j \right)^\top \mathrm{D}_{\boldsymbol{x}} V^i \right\rangle + \frac{1}{2} \sum_{j\in \mathcal{I}} \left\langle K^{i,j}_{a} \tilde{\Phi}^{j}, \tilde{\Phi}^{j} \right\rangle + \sum_{j\in \mathcal{I}} \left\langle K^{i,j}_{a\boldsymbol{x}} \tilde{\Phi}^{j}, \boldsymbol{x} \right\rangle = 0, \end{aligned}\label{eq:HJB95LQG95Example95new} \end{align}\tag{76}\] where we have used the convenient notation \[\begin{align} \tilde{\Phi}^{i} = - \left( K^{i,i}_{a} \right)^{-1} \left( \left( B^i \right)^\top \mathrm{D}_{\boldsymbol{x}} V^i + \left( K^{i,i}_{a\boldsymbol{x}} \right)^\top \boldsymbol{x} \right). \label{eq:Phii95new} \end{align}\tag{77}\] We make the quadratic ansatz \[\begin{align} V^i(t,\boldsymbol{x}) = \frac{1}{2} \left\langle \boldsymbol{x}, P^i \boldsymbol{x} \right\rangle + \left\langle Q^i, \boldsymbol{x} \right\rangle + R^i. \label{eq:ex95ansatz95new} \end{align}\tag{78}\] Differentiating \(V^i\) in time and space, respectively, we get \[\begin{align} \partial_t V^i &= \frac{1}{2} \big\langle \boldsymbol{x}, \dot{P}^i \boldsymbol{x} \big\rangle + \big\langle \dot{Q}^i, \boldsymbol{x} \big\rangle + \dot{R}^i. \\ \mathrm{D}_{\boldsymbol{x}}V^i &= \tilde{P}^i \boldsymbol{x} + Q^i, \label{eq:gradient95V} \\ \mathrm{D}_{\boldsymbol{x}}^2V^i &= \tilde{P}^i, \end{align}\tag{79}\] where the notation \(\tilde{P}^i = \tfrac{1}{2} \big( P^i + (P^i)^\top\big)\) is used for brevity. Next, we evaluate each term in 76 by explicitly inserting 77 followed by the ansatz 78 . The first term immediately becomes \[\begin{align} \frac{1}{2} \mathrm{Tr} \left( \mathbf{C} \mathrm{D}_{\boldsymbol{x}}^2 V^i \mathbf{C}^\top \right) = \frac{1}{2} \mathrm{Tr} \big( \mathbf{C} \tilde{P}^i \mathbf{C}^\top \big). \end{align}\] Inserting 79 into the second term and rearranging terms, we get \[\begin{align} \left\langle F \left( \boldsymbol{\xi} - \boldsymbol{x} \right), \mathrm{D}_{\boldsymbol{x}} V^i \right\rangle &= - \left\langle \boldsymbol{x}, F^\top \tilde{P}^i \boldsymbol{x} \right\rangle + \left\langle \boldsymbol{x}, \tilde{P}^i F \boldsymbol{\xi} - F^\top Q^i \right\rangle + \left\langle F \boldsymbol{\xi}, Q^i \right\rangle. \end{align}\] For the first sum in 76 , note that each term can be written as \[\begin{align} \left\langle \tilde{\Phi}^{j}, \left( B^j \right)^\top\! \mathrm{D}_{\boldsymbol{x}} V^i \right\rangle &= - \left\langle \mathrm{D}_{\boldsymbol{x}} V^i, B^j \left( K^{j,j}_{a} \right)^{-1} \left( \left( B^j \right)^\top \mathrm{D}_{\boldsymbol{x}} V^j + \left( K^{j,j}_{a\boldsymbol{x}} \right)^\top \boldsymbol{x} \right) \right\rangle \\ &= \left\langle \mathrm{D}_{\boldsymbol{x}} V^i, \Theta_1^{j} \mathrm{D}_{\boldsymbol{x}} V^j + \Theta_2^{j} \boldsymbol{x} \right\rangle \\ &= \left\langle\! \boldsymbol{x},\! \big( \tilde{P}^i \Theta_1^{j} \tilde{P}^j \! + \! \tilde{P}^i \Theta_2^{j} \big) \boldsymbol{x}\! \right\rangle \! + \! \left\langle\! \boldsymbol{x},\! \tilde{P}^i \Theta_1^{j} Q^j \! + \! \tilde{P}^j \big( \Theta_1^{j} \big)^\top\! Q^i \! + \! \big( \Theta_2^{j} \big)^\top\! Q^i\! \right\rangle \! + \! \left\langle\! Q^i, \Theta_1^{j} Q^j\! \right\rangle, \end{align}\] where we introduced the notation \[\begin{align} \Theta_1^{j} &\vcentcolon= - B^j \left( K^{j,j}_{a} \right)^{-1} \left( B^j \right)^\top, \quad \Theta_2^{j} \vcentcolon= - B^j \left( K^{j,j}_{a} \right)^{-1} \left( K^{j,j}_{a\boldsymbol{x}} \right)^\top. \end{align}\] Next, we define \[\begin{align} \rho^{i,j} \vcentcolon= \tfrac{1}{2} \big( K^{j,j}_{a} \big)^{-\top} K^{i,j}_{a} \big( K^{j,j}_{a} \big)^{-1} \end{align}\] For each term in the second sum of 76 , we insert the expressions for \(\tilde{\Phi}^{j}\) and \(\mathrm{D}_{\boldsymbol{x}}V^j\), and rearrange terms to get \[\begin{align} \frac{1}{2} \left\langle K^{i,j}_{a} \tilde{\Phi}^{j}, \tilde{\Phi}^{j} \right\rangle &= \left\langle \mathrm{D}_{\boldsymbol{x}} V^j, B^j \rho^{i,j} \left( B^j \right)^\top \mathrm{D}_{\boldsymbol{x}} V^j \right\rangle + \left\langle \boldsymbol{x}, K^{j,j}_{a\boldsymbol{x}} \rho^{i,j} \left( B^j \right)^\top \mathrm{D}_{\boldsymbol{x}} V^j \right\rangle \\ & + \left\langle \mathrm{D}_{\boldsymbol{x}} V^j, B^j \rho^{i,j} \left( K^{j,j}_{a\boldsymbol{x}} \right)^\top \boldsymbol{x} \right\rangle + \left\langle \boldsymbol{x}, K^{j,j}_{a\boldsymbol{x}} \rho^{i,j} \left( K^{j,j}_{a\boldsymbol{x}} \right)^\top \boldsymbol{x} \right\rangle \\ &= \left\langle \mathrm{D}_{\boldsymbol{x}} V^j, \Theta_3^{i,j} \mathrm{D}_{\boldsymbol{x}} V^j \right\rangle + \left\langle \boldsymbol{x}, \Theta_4^{i,j} \mathrm{D}_{\boldsymbol{x}} V^j \right\rangle + \left\langle \mathrm{D}_{\boldsymbol{x}} V^j, \Theta_5^{i,j} \boldsymbol{x} \right\rangle + \left\langle \boldsymbol{x}, \Theta_6^{i,j} \boldsymbol{x} \right\rangle \\ &= \left\langle \boldsymbol{x}, \big( \tilde{P}^j \Theta_3^{i,j} \tilde{P}^j + \Theta_4^{i,j} \tilde{P}^j + \tilde{P}^j \Theta_5^{i,j} + \Theta_6^{i,j} \big) \boldsymbol{x} \right\rangle \\ & + \left\langle \boldsymbol{x}, \tilde{P}^j \big( \Theta_3^{i,j} \big)^\top Q^j + \tilde{P}^j \Theta_3^{i,j} Q^j + \Theta_4^{i,j} Q^j + \big( \Theta_5^{i,j} \big)^\top Q^j \right\rangle + \left\langle Q^j, \Theta_3^{i,j} Q^j \right\rangle. \end{align}\] The matrices introduced in the second equality are defined as \[\begin{align} \Theta_3^{i,j} \vcentcolon= B^j \rho^{i,j} \left( B^j \right)^\top, \quad \Theta_4^{i,j} \vcentcolon= K^{j,j}_{a\boldsymbol{x}} \rho^{i,j} \left( B^j \right)^\top, \quad \Theta_5^{i,j} \vcentcolon= B^j \rho^{i,j} \left( K^{j,j}_{a\boldsymbol{x}} \right)^\top, \quad \Theta_6^{i,j} \vcentcolon= K^{j,j}_{a\boldsymbol{x}} \rho^{i,j} \left( K^{j,j}_{a\boldsymbol{x}} \right)^\top. \end{align}\] Each term in the final sum of 76 can be written as \[\begin{align} \left\langle K^{i,j}_{a\boldsymbol{x}} \tilde{\Phi}^{j}, \boldsymbol{x} \right\rangle &= - \left\langle \boldsymbol{x}, K^{i,j}_{a\boldsymbol{x}} \left( K^{j,j}_{a} \right)^{-1} \left( \left( B^j \right)^\top \mathrm{D}_{\boldsymbol{x}} V^j + \left( K^{j,j}_{a\boldsymbol{x}} \right)^\top \boldsymbol{x} \right) \right\rangle = \left\langle \boldsymbol{x}, \Theta_7^{i,j} \mathrm{D}_{\boldsymbol{x}} V^j + \Theta_8^{i,j} \boldsymbol{x} \right\rangle \\ &= \left\langle \boldsymbol{x}, \big( \Theta_7^{i,j} \tilde{P}^j + \Theta_8^{i,j} \big) \boldsymbol{x} \right\rangle + \left\langle \boldsymbol{x}, \Theta_7^{i,j} Q^j \right\rangle, \end{align}\] where we have introduced the matrices \[\begin{align} \Theta_7^{i,j} &\vcentcolon= - K^{i,j}_{a\boldsymbol{x}} \left( K^{j,j}_{a} \right)^{-1} \left( B^j \right)^\top, \quad \Theta_8^{i,j} \vcentcolon= - K^{i,j}_{a\boldsymbol{x}} \left( K^{j,j}_{a} \right)^{-1} \left( K^{j,j}_{a\boldsymbol{x}} \right)^\top. \end{align}\] Now, we combine the earlier defined matrices according to \[\begin{align} \tilde{\Theta}_3^{i,j}\vcentcolon=\Theta_3^{i,j}+ \big( \Theta_3^{i,j}\big)^\top, \quad \Theta_{4,5,7}^{i,j}\vcentcolon=\Theta_4^{i,j}+ \big( \Theta_5^{i,j}\big)^\top + \Theta_7^{i,j}, \quad \Theta_{6,8}^{i,j}\vcentcolon=\Theta_6^{i,j}+ \Theta_8^{i,j}. \end{align}\] Inserting the derived expressions into 76 , we then have \[\begin{align} 0 &= \Big\langle \boldsymbol{x}, \Big( \tfrac{1}{2} \dot{P}^i - F^\top \tilde{P}^i + (K^{i}_{\boldsymbol{x}})^\top K^{i}_{\boldsymbol{x}} + \sum_{j\in \mathcal{I}} \big( \tilde{P}^i ( \Theta_1^{j} \tilde{P}^j + \Theta_2^{j} ) + \tilde{P}^j \Theta_3^{i,j} \tilde{P}^j + \Theta_{4,5,7}^{i,j} \tilde{P}^j + \Theta_{6,8}^{i,j} \big) \Big) \boldsymbol{x} \Big\rangle \\ &+ \Big\langle \boldsymbol{x}, \dot{Q}^i + \tilde{P}^i F \boldsymbol{\xi} - F^\top Q^i + \sum_{j\in \mathcal{I}} \Big( \tilde{P}^i \Theta_1^{j} Q^j + \tilde{P}^j \big( \Theta_1^{j} \big)^\top Q^i + \big( \Theta_2^{j} \big)^\top Q^i + \tilde{P}^j \tilde{\Theta}_3^{i,j} Q^j + \Theta_{4,5,7}^{i,j} Q^j \Big) \Big\rangle \\ &+ \dot{R}^i + \frac{1}{2} \mathrm{Tr} \left( \mathbf{C} \tilde{P}^i \mathbf{C}^\top \right) + \left\langle F \boldsymbol{\xi}, Q^i \right\rangle + \sum_{j\in \mathcal{I}} \big( \big\langle Q^i, \Theta_1^{j} Q^j \big\rangle + \big\langle Q^j, \Theta_3^{i,j} Q^j \big\rangle \big). \end{align}\] Setting each inner product to zero, and using the fact that \[\begin{align} \big\langle \boldsymbol{x}, 2 F^\top \tilde{P}^i \boldsymbol{x} \big\rangle &= \big\langle \boldsymbol{x}, \big( F^\top P^i + P^i F \big) \boldsymbol{x} \big\rangle, \end{align}\] yields the Riccati system 32 , that is \[\begin{align} \begin{aligned}\label{eq:Riccati95system95derived} \dot{P}^i &= F^\top P^i + P^i F - (K^{i}_{\boldsymbol{x}})^\top K^{i}_{\boldsymbol{x}} - 2 \sum_{j\in \mathcal{I}} \Big( \tilde{P}^i \big( \Theta_1^{j} \tilde{P}^j + \Theta_2^{j} \big) + \tilde{P}^j \Theta_3^{i,j} \tilde{P}^j + \Theta_{4,5,7}^{i,j} \tilde{P}^j + \Theta_{6,8}^{i,j} \Big) \\ \dot{Q}^i &= - \tilde{P}^i F \boldsymbol{\xi} + F^\top Q^i - \sum_{j\in \mathcal{I}} \Big( \tilde{P}^i \Theta_1^{j} Q^j + \tilde{P}^j \big( \Theta_1^{j} \big)^\top Q^i + \big( \Theta_2^{j} \big)^\top Q^i + \tilde{P}^j \tilde{\Theta}_3^{i,j} Q^j + \Theta_{4,5,7}^{i,j} Q^j \Big) \\ \dot{R}^i &= - \frac{1}{2} \mathrm{Tr} \left( \mathbf{C} \tilde{P}^i \mathbf{C}^\top \right) - \left\langle F \boldsymbol{\xi}, Q^i \right\rangle - \sum_{j\in \mathcal{I}} \Big( \big\langle Q^i, \Theta_1^{j} Q^j \big\rangle + \big\langle Q^j, \Theta_3^{i,j} Q^j \big\rangle \Big) \end{aligned} \end{align}\tag{80}\] where the \(\Theta\)-matrices are \[\begin{align} \Theta_1^{j} &= - B^j \left( K^{j,j}_{a} \right)^{-1} \left( B^j \right)^\top , \\ \Theta_2^{j} &= - B^j \left( K^{j,j}_{a} \right)^{-1} \left( K^{j,j}_{a\boldsymbol{x}} \right)^\top, \\ \Theta_3^{i,j} &= \tfrac{1}{2} B^j \left( K^{j,j}_{a} \right)^{-\top} K^{i,j}_{a} \left( K^{j,j}_{a} \right)^{-1} \left( B^j \right)^\top, \\ \tilde{\Theta}_3^{i,j} &= \tfrac{1}{2} B^j \left( K^{j,j}_{a} \right)^{-\top} \Big( \left( K^{i,j}_{a} \right)^\top + K^{i,j}_{a} \Big) \left( K^{j,j}_{a} \right)^{-1} \left( B^j \right)^\top, \\ \Theta_{4,5,7}^{i,j} &= \Big( \tfrac{1}{2} K^{j,j}_{a\boldsymbol{x}} \big( K^{j,j}_{a} \big)^{-\top} \Big( \big( K^{i,j}_{a} \big)^\top + K^{i,j}_{a} \Big) - K^{i,j}_{a\boldsymbol{x}} \Big) \left( K^{j,j}_{a} \right)^{-1} \left( B^j \right)^\top, \\ \Theta_{6,8}^{i,j} &= \Big( \tfrac{1}{2} K^{j,j}_{a\boldsymbol{x}} \left( K^{j,j}_{a} \right)^{-\top} K^{i,j}_{a} - K^{i,j}_{a\boldsymbol{x}} \Big) \left( K^{j,j}_{a} \right)^{-1} \left( K^{j,j}_{a\boldsymbol{x}} \right)^\top. \end{align}\]

Remark 4. We consider the single-player case, i.e., \(\mathcal{I}= \{1\}\), and show that the derived system corresponds to the generalized Riccati differential equation (see, e.g., [39]). Since there is only one player, we can remove the superscript indices of all variables. Assuming \(P\) and \(K_a\) symmetric, we have \[\begin{align} \dot{P} &= F^\top P + P F - K_{\boldsymbol{x}}^\top K_{\boldsymbol{x}} - 2 \Big( P \big( \Theta_1 + \Theta_3 \big) P + P \Theta_2 + \Theta_{4,5,7} P + \Theta_{6,8} \Big). \label{eq:P95equation95single95agent} \end{align}\qquad{(2)}\] In this setting, the matrices can be simplified according to \[\begin{align} \Theta_1 + \Theta_3 &= - B K_{a}^{-1} B^\top + \tfrac{1}{2} B K_{a}^{-\top} K_{a} K_{a}^{-1} B^\top= - \tfrac{1}{2} B K_{a}^{-1} B^\top, \\ \Theta_2 &= - B K_a^{-1} K_{a\boldsymbol{x}}^\top, \\ \Theta_{4,5,7} &= \Big( \tfrac{1}{2} K_{a\boldsymbol{x}} K_{a}^{-\top} \Big( K_{a}^\top + K_{a} \Big) - K_{a\boldsymbol{x}} \Big) K_{a}^{-1} B^\top = 0, \\ \Theta_{6,8} &= \Big( \frac{1}{2} K_{a\boldsymbol{x}} K_{a}^{-\top} K_{a} - K_{a\boldsymbol{x}} \Big) K_{a}^{-1} K_{a\boldsymbol{x}}^\top = - \frac{1}{2} K_{a\boldsymbol{x}} K_{a}^{-1} K_{a\boldsymbol{x}}^\top \end{align}\] Inserting these expressions into ?? , we obtain the single agent Riccati equation \[\begin{align} \dot{P} &= F^\top P + P F - K_{\boldsymbol{x}}^\top K_{\boldsymbol{x}} + P B K_{a}^{-1} B^\top P + 2 P B K_{a}^{-1} K_{a\boldsymbol{x}}^\top + K_{a\boldsymbol{x}} K_{a}^{-1} K_{a\boldsymbol{x}}^\top. \end{align}\] Compare this to the generalized Riccati differential equation (see, e.g., [39]), which can be written on the form \[\begin{align} \dot{P} = - P A - A^\top P + P B R^{-1} B^\top P + S R^{-1} S^\top + 2 S R^{-1} B^\top P - Q. \end{align}\] Note that, with \(Q = K_{\boldsymbol{x}}^\top K_{\boldsymbol{x}}\), \(S = K_{a\boldsymbol{x}}\) and \(R=K_{a}\), the generalized Riccati equation coincides with our single-player Riccati equation up to the sign convention for the drift term.

References↩︎

[1]
J.-P. Fouque, R. A. Carmona, and L. H. Sun, “Mean field games and systemic risk,” Communications in Mathematical Sciences, vol. 13, no. 4, pp. 911–933, 2015.
[2]
E. J. Dockner, S. Jorgensen, N. V. Long, and G. Sorger, Differential games in economics and management science. Cambridge University Press, 2000.
[3]
A. Schied and T. Zhang, “A state-constrained differential game arising in optimal portfolio liquidation,” Mathematical Finance, vol. 27, no. 3, pp. 779–802, 2017.
[4]
L. D. Berkovitz, “Differential games of generalized pursuit and evasion,” SIAM Journal on Control and Optimization, vol. 24, no. 3, pp. 361–373, 1986.
[5]
R. Isaacs, Differential games: A mathematical theory with applications to warfare and pursuit, control and optimization. New York: John Wiley & Sons, 1965.
[6]
K. M. Ramachandran and C. P. Tsokos, Stochastic differential games. Theory and applications, vol. 2. Atlantis Press, 2012.
[7]
Y. Yavin, “Stochastic pursuit-evasion differential games in the plane,” Journal of Optimization Theory and Applications, vol. 50, no. 3, pp. 495–523, 1986.
[8]
G. M. Anderson, “Comparison of optimal control and differential game intercept missile guidance laws,” Journal of Guidance and Control, vol. 4, no. 2, pp. 109–115, 1981.
[9]
F. A. Faruqi, P. Belobaba, J. Cooper, and A. Seabridge, Differential game theory with applications to missiles and autonomous systems guidance. John Wiley & Sons, 2017.
[10]
M. Pontani and B. A. Conway, “Optimal interception of evasive missile warheads: Numerical solution of the differential game,” Journal of Guidance, Control, and Dynamics, vol. 31, no. 4, pp. 1111–1122, 2008.
[11]
I. E. Weintraub, M. Pachter, and E. Garcia, “An introduction to pursuit-evasion differential games,” in 2020 american control conference (ACC), 2020, pp. 1049–1066.
[12]
Y. Yavin and R. de Villiers, “Application of stochastic differential games to medium-range air-to-air missiles,” Journal of Optimization Theory and Applications, vol. 67, no. 2, pp. 355–367, 1990.
[13]
T. Bui, X. Cheng, Z. Jin, and G. Yin, “Approximation of a class of non-zero-sum investment and reinsurance games for regime-switching jump–diffusion models,” Nonlinear Analysis: Hybrid Systems, vol. 32, pp. 276–293, 2019.
[14]
R. J. Elliott, “Stochastic differential games and alternate play,” in Control theory, numerical methods and computer systems modelling: International symposium, rocquencourt, june 17–21, 1974, 1975, pp. 97–106.
[15]
R. Carmona and F. Delarue, Probabilistic theory of mean field games with applications i and II. Springer, 2018.
[16]
A. Friedman, “Stochastic differential games,” Journal of Differential Equations, vol. 11, no. 1, pp. 79–108, 1972.
[17]
H. J. Kushner, “Numerical approximations for stochastic differential games,” SIAM Journal on Control and Optimization, vol. 41, no. 2, pp. 457–486, 2002.
[18]
H. J. Kushner, “Numerical approximations for nonzero-sum stochastic differential games,” SIAM Journal on Control and Optimization, vol. 46, no. 6, pp. 1942–1971, 2007.
[19]
M. Falcone, “Numerical methods for differential games based on partial differential equations,” International Game Theory Review, vol. 8, no. 2, pp. 231–272, 2006.
[20]
Y. Achdou, F. Camilli, and I. Capuzzo-Dolcetta, “Mean field games: Convergence of a finite difference method,” SIAM Journal on Numerical Analysis, vol. 51, no. 5, pp. 2585–2612, 2013.
[21]
Y. Achdou and I. Capuzzo-Dolcetta, “Mean field games: Numerical methods,” SIAM Journal on Numerical Analysis, vol. 48, no. 3, pp. 1136–1162, 2010.
[22]
R. Hu, “Deep fictitious play for stochastic differential games,” Communications in Mathematical Sciences, vol. 19, no. 2, pp. 325–353, 2021, doi: 10.4310/CMS.2021.v19.n2.a2.
[23]
J. Han and R. Hu, “Deep fictitious play for finding Markovian Nash equilibrium in multi-agent games,” in Mathematical and scientific machine learning, 2020, pp. 221–245.
[24]
J. Han, R. Hu, and J. Long, “Convergence of deep fictitious play for stochastic differential games,” Frontiers of Mathematical Finance, vol. 1, no. 2, pp. 287–319, 2022, doi: 10.3934/fmf.202101.
[25]
G. W. Brown, “Iterative solution of games by fictitious play,” in Activity analysis of production and allocation, T. C. Koopmans, Ed. New York: Wiley, 1951, pp. 374–376.
[26]
J. Han, A. Jentzen, and W. E, “Solving high-dimensional partial differential equations using deep learning,” Proceedings of the National Academy of Sciences, vol. 115, no. 34, pp. 8505–8510, 2018.
[27]
K. Andersson, A. Andersson, and C. W. Oosterlee, “Convergence of a robust deep FBSDE method for stochastic control,” SIAM Journal on Scientific Computing, vol. 45, no. 1, pp. A226–A255, 2023.
[28]
K. Andersson, A. Andersson, and C. W. Oosterlee, “The deep multi-FBSDE method: A robust deep learning method for coupled FBSDEs,” Journal of Scientific Computing, vol. 106, no. 3, p. 77, 2026.
[29]
K. Andersson and A. Gnoatto, “Multi-layer deep xVA: Structural credit models, measure changes and convergence analysis,” arXiv preprint arXiv:2502.14766, 2025.
[30]
I. Karatzas and S. E. Shreve, Brownian motion and stochastic calculus, Second. New York: Springer, 1991.
[31]
A. Bensoussan and J. Frehse, “Smooth solutions of systems of quasilinear parabolic equations,” ESAIM: Control, Optimisation and Calculus of Variations, vol. 8, pp. 169–193, 2002.
[32]
F. Delarue and G. Guatteri, “Weak existence and uniqueness for forward–backward SDEs,” Stochastic Processes and their Applications, vol. 116, no. 12, pp. 1712–1742, 2006.
[33]
G. W. Brown, “Some notes on computation of games solutions,” RAND Corporation, RM-125-PR, 1949.
[34]
J. Robinson, “An iterative method of solving a game,” Annals of Mathematics, vol. 54, no. 2, pp. 296–301, 1951.
[35]
W. H. Fleming and H. M. Soner, Controlled markov processes and viscosity solutions, Second., vol. 25. New York: Springer, 2006.
[36]
L.-P. Chaintron, “Existence and global Lipschitz estimates for unbounded classical solutions of a HamiltonJacobi equation,” Annales de la Faculté des sciences de Toulouse : Mathématiques, vol. 34, no. 4, pp. 851–876, 2025, doi: 10.5802/afst.1824.
[37]
N. V. Krylov, Lectures on elliptic and parabolic equations in hölder spaces, vol. 12. American Mathematical Society, 1996.
[38]
N. Dokuchaev, “Estimates for distances between first exit times via parabolic equations in unbounded cylinders,” Probability Theory and Related Fields, vol. 129, no. 2, pp. 290–314, 2004.
[39]
A. Ferrante and L. Ntogramatzidis, “The generalised continuous algebraic Riccati equation and impulse-free continuous-time LQ optimal control,” Automatica, vol. 50, no. 4, pp. 1176–1180, 2014, doi: 10.1016/j.automatica.2014.02.014.

  1. Saab AB, Gothenburg, Sweden↩︎

  2. Department of Mathematical Sciences, Chalmers University of Technology & University of Gothenburg↩︎

  3. Department of Economics, University of Verona. Email: kristofferherbert.andersson@univr.it↩︎

  4. Department of Electrical Engineering, Chalmers University of Technology↩︎