Risk-averse mean field games: exploitability and non-asymptotic analysis 3


Abstract

In this paper, we use mean field games (MFGs) to investigate approximations of \(N\)-player games (\(N\)pGs) with uniformly symmetrically continuous heterogeneous closed-loop actions. To incorporate agents’ risk aversion (beyond the classical expected utility of total costs), we use an abstract evaluation functional for their performance criteria. Centered around the notion of exploitability, we conduct non-asymptotic analysis on the approximation capability of MFGs from the perspective of state-action distributions without requiring the uniqueness of equilibria. Under suitable assumptions, we first show that scenarios in the \(N\)pGs with large \(N\) and small average exploitabilities can be well approximated by approximate solutions of MFGs with relatively small exploitabilities. We then show that \(\delta\)-mean field equilibria can be used to construct \(\varepsilon\)-equilibria in \(N\)pGs. Furthermore, in this general setting, we prove the existence of mean field equilibria. This proof reveals a possible avenue for incorporating penalization for randomized action into MFGs.

Mean field games, risk averse, exploitability, non-asymptotics, closed loop controls, non-uniqueness

91A16, 91A06, 91A25, 93E20

1 Introduction↩︎

Since the pioneering work of [1], [2] and [3], the literature on mean field games (MFGs) has gone through, and continues to experience, a massive expansion. MFGs can be viewed as symmetric games with infinitely many players and are primarily used to approximate \(N\)-player games (\(N\)pGs) with large population where the impact of each individual is “small”. The majority of the literature focuses on analysis in continuous time. We refer to [4][8] and the reference therein for studies of MFGs under different settings in continuous time, such as linear-quadratic games, major-minor agent games, robust games, partially observed games, and exploratory games. We also refer to [9] and [10] for overviews.

While MFGs are used to approximate large games, it is crucial to understand their approximation capabilities. Following [1], there has been enormous work developing the idea that uses mean field equilibriia (MFEs) to construct \(\varepsilon\)-Nash equilibria in \(N\)pGs. Conversely, the results of [3] suggests that certain equilibria in \(N\)pGs, for large \(N\), can be captured by MFEs. Research in this direction is indispensable because it helps to answer whether using MFGs to approximate \(N\)pGs misses any equilibria. Below we highlight some work that progress the understanding on this matter. Assuming the uniqueness of MFEs, [11] uses the master equation to study closed-loop4 \(N\)pGs, and shows that the equilibria converge to the MFE as \(N\to\infty\). [13] and [14] study the convergence of open-loop equlibria under various settings without assuming uniqueness. [15] and [16] further extend the convergence results from [13] to closed-loop equilibria with or without common noise. We note that the convergence results in [15] and [16] are established for weak MFEs. Weak MFEs consists of a random flow of measures. Convergence to weak MFEs is possible even with considerably weak assumptions on the actions of players. On the contrary, the more common-seen MFEs in [1], [3], [14], among many others, consists of a deterministic flow of measures, and is called strong MFEs, or simply, MFEs. It often requires suitable assumptions on players’ actions to facilitate the convergences to MFEs. Apart from the above, [17] pioneers the study of the value function approximation of MFGs within the discrete time framework without requiring the uniqueness of MFEs. Unlike previous studies that have focused on the capturing capabilities of MFGs, the analysis conducted in [17] is predominantly non-asymptotic. Our work draws inspiration from [15] and [17], among others, and we provide a detailed comparison in 2.2. For completeness, we also refer to [18][21] and the references therein for other studies that address the non-uniqueness of equilibria, although capturing capability is not their primary focus. Finally, we would like to point to [22][24] for recent work on approximation analysis on extended MFGs with unique equilibrium. Extended MFGs consider interactions through players’ actions.

This paper’s goal is to further explore the approximation capability of MFGs from the perspective of state-action distributions – with a focus on (strong) MFEs and non-asymp-
totics description, while allowing for multiple equilibria, certain heterogeneous closed-loop actions, as well as \(\varepsilon\)-relaxation in individual decision making. Readers familiar with MFGs may realize that some features in our exploration has appeared in the extant literature, however, not all together in a single setting. To avoid excessive technicalities, we choose to work in discrete time and finite horizon with state and action spaces being Polish.

The idea of MFGs in discrete time dates back to [25], although the connection to \(N\)pGs was left unstudied for quite a while. Notably, [26] considers an infinite horizon \(N\)pG with discounted unbounded costs on Polish state and action spaces, proves the existence of MFEs, and constructs \(\varepsilon\)-equilibra in \(N\)pGs using MFEs. Analogous results are established for risk sensitive games [27], partially observed games [28], and risk averse games [29]. There has been recent progress on machine learning algorithms for MFEs in discrete time, see, e.g., [30] and [31]. We would also like to point out [32] for a causal transport-based method potentially useful for handling weak MEFs.

Another aspect we aim to address in our study is the incorporation of risk aversion into MFG theory. In practical scenarios, it is often observed—and sometimes preferred—that decision-makers prioritize avoiding unfavorable outcomes over pursuing favorable ones. A performance criterion that aligns with this concept is introduced by [33], with further extensions allowing for randomized actions discussed in [34]. This criterion offers the advantage of inherent time consistency, and in the meantime, naturally generalizes the standard risk-neutral framework, which uses expected total cost as the performance metric. This setup facilitates subsequent analyses from the standpoint of dynamic programming. For related applications in areas such as automation and finance, we refer to [35][37] and the references therein. Although risk-aversion theory has been extensively explored in the context of single-agent decision-making, risk aversion in a MFG framework remains relatively uncharted. Whether one chooses to follow the principles proposed by [33] or not, the discourse on this topic in MFG literature is notably sparse. For an study that aligns with [33], we direct readers to [29]. Regarding other forms of risk aversion within MFGs, we refer to [38] and the reference therein.

In our setup, we assume all players are subject to the same Markovian controlled dynamics and consider weak interaction, that is, the impact of the other players on an individual is through the empirical distribution of the population states. Regarding the performance criteria, we abstract the procedure of backward induction, and define an evaluation functional as a composition of operators that maps value functions backward. This abstraction not only encompasses the Bellman operators commonly studied in risk-neutral mean field Markov Decision Process (MDP) frameworks (e.g., [26] and [39]), but also facilitates the examination of risk-averse variants tied to the aforementioned risk-averse performance criteria. Additionally, it contributes to a more concise presentation and simplifies notation. To approximate \(N\)pGs with MFGs, we impose further assumptions. In particular, we consider \(N\)pGs where all players use \(\vartheta\)-symmetrically continuous policy (see 7). An example illustrating the necessity5 of this assumption is provided in an example in 8.1. The example is constructed in a degenerate environment with absorption, which is in spirit similar to that in [40].

Our main results are centered around the notion called exploitability. In \(N\)pGs, exploitability quantifies the best possible improvement a player can achieve by revising her original policy, while assuming the policies of other players remain unchanged. It effectively describes the scale of disequilibrium in a game scenario. Note that \(\varepsilon\)-equilibria in an \(N\)pG can be equivalently defined via exploitability. The introduction of exploitability paves ways for investigating the global properties of games, in that it provides insights beyond the existence and uniqueness of equilibria. In MFGs, exploitability can be defined analogously. More precisely, exploitability of a policy is defined as the difference between the outcome under the considered policy and the outcome under the optimal policy, where both outcomes are computed under the mean field flow induced by the considered policy. \(\varepsilon\)-MFE can be defined using exploitability.

The notion of exploitability is frequently used for analyzing the convergence of algorithms that approximate MFEs, e.g., [41][46]. Notably, [45] develops an algorithm that approximate MFEs by minimizing exploitability. This algorithm is capable of handling the presence of multiple MFEs.

Despite its widespread use in equilibrium-finding algorithms, exploitability has rarely been the focal point in the approximation analysis of MFGs, as evidenced by the literature review above. This oversight may stem from the current focus in approximation analysis, which often overlooks non-asymptotic considerations surrounding approximate \(N\)-player equilibria (\(N\)pEs). To the best of our knowledge, the only relevant study that concerns capturing approximate \(N\)pEs non-asymptotically thus far is [17]. They primarily concentrates on value functions and does not involve exploitability.

To advocate for the use of exploitability, we contend that in more realistic considerations of \(N\)pGs, equilibria are often attained in an approximate sense. Without sufficiently strong assumptions, there is no assurance that an approximate equilibrium closely aligns with any exact equilibrium (in terms of, for example, the empirical state distribution), even when there is a unique equilibrium. To substantiate that MFGs approximate \(N\)pGs effectively, one approach is to relate each approximate \(N\)pE to an approximate MFE, if feasible, or to identify the reasons for failure if not. The use of exploitability is naturally justified, as it is anticipated that related approximate equilibria, if they exist, should exhibit comparable scales of exploitability.

The main contributions of this paper can be summarized as follows:

  1. Various types of exploitabilities are proposed and their interrelationships are scrutinized. In particular, to incorporate the notion of exploitability in \(N\)pGs and MFGs into our setting that involves abstract evaluation operations, we introduce the total (stepwise) exploitability (see 53 and 60 ). The relationship between total exploitability and the standard notion of exploitability, here called end exploitability (see 52 and 59 ), is examined in [prop:EstEndExploitabilityCont] and [prop:EstMFEndExploitability]. We also consider end exploitability based on the best symmetrically continuous policy in \(N\)pG (see 4.1), as this aligns naturally with our setup. In [prop:EstEndExploitabilityCont], we demonstrate that the difference between the constrained end exploitability and the unconstrained one is negligible under suitable conditions.

  2. We construct a state-action mean field flow from an \(N\)-player scenario, and show that the expected bounded-Lipschitz distance between the empirical flow in the \(N\)-player scenario and the mean field flow is small when \(N\) is large. As long as certain continuity assumptions hold true, we prove that the total exploitability in a mean field scenario is dominated by the average of the total exploitabilities of all \(N\) players plus a small error term. Under stronger conditions, the difference between two quantities can be bounded by the same error term. Detailed statements are presented in 3. Conversely, in 4, we use \(\varepsilon\)-MFEs to construct an open-loop \(N\)pG where all players have exploitabilities bounded by \(\varepsilon\) plus a small error term. These results collectively lead to a non-asymptotic description on the approximation capability of MFGs.

  3. We prove in an abstract formulation the existence of MFEs. This formulation encompasses a range of performance criteria beyond the classical expected total cost, such as risk aversions via nested compositions (see 8.3 for example), certain extended forms of cost (e.g., 48 ). The proof uses the Kakutani fixed point theorem, and is motivated by earlier work on the existence of MFEs in discrete time (cf. [25] and [26]). Due to our abstract formulation, however, we must develop significant modifications to the original analysis. Our approach provides substantial extensions of the previous methods and our techniques could be of interest on their own. As an unexpected byproduct, our proof uncovers a potential avenue for incorporating penalization for randomized actions within the framework of MFGs. We refer to 5 and thereafter for related discussions.

The organization of the remainder of this paper is as follows. To facilitate understanding, we begin with heuristic discussions of our approximation results in 2. In 3, we describe our model setup and introduce the necessary notations. Specifically, the technical assumptions are detailed in 3.7. We then define and examine various notions of exploitability for \(N\)pGs and MFGs in 4. Our findings on the approximation capabilities of MFGs are presented in 5. The subsequent section, 6, addresses the existence of MFE. The proofs for the results in 4, 5, and 6 are provided in 7. Additionally, 8 contains several examples, and 9 presents technical results. The proofs of preliminary results mentioned in 3 are included in 10. Finally, a glossary of notations is provided for reference in 1.

2 Heuristics on the approximation capability of mean field games↩︎

In 2.1, we heuristically derive our key findings in a discrete setting, focusing on how \(N\)pEs can be captured by MFEs, in the absence of uniqueness. We place more emphasis on this aspect rather than on constructing \(N\)pEs from MFEs, as the latter is generally better understood, while understanding of the former is still evolving. In 2.2, we comment on the approximation capabilities of MFGs. Additionally, we discuss how our approach differs from existing ones. Finally in 2.3, we highlight some of the technical challenges we face when rigorously establishing these heuristics in a more general setting.

The discussion in this section primarily pertains to the results presented in 4 and 5, which regards various notions of exploitability and the approximation capabilities of MFGs. For insights into the existence of MFE, we refer the reader to 6.

2.1 Derivation↩︎

For simplicity, we consider finite state space \(\mathbb{X}\), finite action space \(\mathbb{A}\), and finite horizon \(T\in\mathbb{N}\). Let \(N\) represent the number of players. As we consider non-asymptotics, we suppress \(N\) from the notations introduced below unless necessary. Suppose all players start from the same initial position \(x_1\in\mathbb{X}\) at \(t=1\). We consider a Markov environment with transition kernels \(P_t:\mathbb{X}^N\times\mathbb{A}\to\mathcal{P}(\mathbb{X})\) for \(t=1,\dots,T-1\). Here, the state transition of an individual player depends only on the positions of all players and the action of the player herself. The case where the transition is also affected by other players’ actions will eventually lead to extended mean field games, which is out of the scope of this paper. Similarly, we consider cost functions \(C_t:\mathbb{X}^N\times\mathbb{A}\to\mathbb{R}\) for \(t=1,\dots,T\). In what follows, we let \(\mathfrak{p}^n_t:\mathbb{X}^N\to\mathcal{P}(\mathbb{A})\) represent the randomized actions of player-\(n\) at time \(t\), and let \(\mathfrak{p}^n=(\mathfrak{p}^n_1,\dots,\mathfrak{p}^n_T)\) denote the player-\(n\)’s policy. By convention, we use subscripts for the inputs of measure-valued functions. Moreover, we set \(\boldsymbol{\mathfrak{P}}=(\mathfrak{p}^1,\dots,\mathfrak{p}^N)\) to be the collection of all player’s policies, which we call a scenario in the \(N\)pG.

Let us consider a probability space \((\Omega,\mathscr{F},\mathbb{P}^{\boldsymbol{\mathfrak{P}}})\). For notational convenience, we suppress \(\boldsymbol{\mathfrak{P}}\) from \(\mathbb{P}^{\boldsymbol{\mathfrak{P}}}\) and \(\mathbb{E}^{\boldsymbol{\mathfrak{P}}}\). Under \(\mathbb{P}\), we have the following dynamics \[\begin{align} \boldsymbol{X}_1=(x_1,\dots,x_1),\quad \boldsymbol{A}_t\sim\bigotimes_{n=1}^N\mathfrak{p}^n_{t,\boldsymbol{X}_t}\,,\quad \boldsymbol{X}_{t+1}\sim\bigotimes_{n=1}^N P_{t,\boldsymbol{X}_t, A^n_t}\,, \end{align}\] where \(\boldsymbol{X}_t = (X^1_t,\dots,X^N_t)\) and \(\boldsymbol{A}_t = \left(A^1_t,\dots,A^N_t\right)\) represent the state and actions of the players at time \(t\), respectively. Furthermore, player-\(n\) aims to find \[\begin{align} \min_{\mathfrak{p}^n} \mathbb{E}\left[ \sum_{t=1}^T C_t(\boldsymbol{X}_t, A^n_t) \right]. \end{align}\]

To ensure that MFGs effectively approximate the aforementioned \(N\)pG, additional structure is needed. We impose the required structure shortly. For the moment, however, let us introduce a key concept in our analysis: the exploitability of player-\(n\) in an \(N\)pG. To do so, consider a hypothetical case where all other players adhere to their original policies, while player-\(n\) has complete observation of \(\boldsymbol{X}_t\) at each decision-making point in time \(t\) and is allowed to consider a different policy than her original one. From player-\(n\)’s viewpoint, this problem may be posed as a MDP. We next define player-\(n\)’s exploitability as \[\begin{align} \label{eq:HeurDefEndExploitability} R^n(\boldsymbol{\mathfrak{P}}) := \mathbb{E}\left[ \sum_{t=1}^T C_t(\boldsymbol{X}_t, A^n_t) \right] - \min_{\mathfrak{p}^n} \mathbb{E}\left[ \sum_{t=1}^T C_t(\boldsymbol{X}_t, A^n_t) \right]. \end{align}\tag{1}\] It is straightforward to see that \(\boldsymbol{\mathfrak{P}}\) is a \(\varepsilon\)-Nash equilibrium if and only if \(R^n(\boldsymbol{\mathfrak{P}})\le\varepsilon\) for all \(n\). Later, we also use \(\frac{1}{N}\sum_{n=1}^N R^n(\boldsymbol{\mathfrak{P}})\) to characterize approximate equilibrium in the average sense. We introduce below a few terms for the Dynamic Programing Principle (DPP) related to the aforementioned MDP. For \(\boldsymbol{x} = (x^1,\dots,x^N)\), we define \[\begin{align} V^{n*}_t(\boldsymbol{x}) := \min_{\mathfrak{p}^n} \mathbb{E}\left[ \sum_{r=t}^T C(\boldsymbol{X}_r,A^n_r) \,\Bigg|\,\boldsymbol{X}_r = \boldsymbol{x} \right], \quad t=1,\dots, T, \end{align}\] i.e., \(V^{n*}_t\) is the optimal value function of player-\(n\) at time \(t\). For convenience, we stipulate that \(V^{n*}_t\equiv 0\) for \(t>T\). The following DPP in terms of the Bellman equation is well-known \[\begin{align} \label{eq:NpDPP} V^{n*}_t(\boldsymbol{x}) = \min_{\lambda\in\mathcal{P}(\mathbb{A})} \mathbb{E}\big[ C_t(\boldsymbol{x}, A^n) + V^{n*}_{t+1}(\boldsymbol{X}) \big],\quad t=1,\dots,T-1, \end{align}\tag{2}\] where we have employed dummy random variables \(\boldsymbol{A}=(A^1,\dots,A^N)\) and \(\boldsymbol{X}=(X^1,\dots,X^N)\) with dynamics:

\[\begin{align} \boldsymbol{A}\sim\mathfrak{p}^1_{t,\boldsymbol{x}}\otimes\dots\otimes \mathfrak{p}^{n-1}_{t,\boldsymbol{x}}\otimes\;\lambda\;\otimes\mathfrak{p}^{n+1}_{t,\boldsymbol{x}}\otimes\dots\otimes \mathfrak{p}^{n+1}_{t,\boldsymbol{x}}\;, \qquad \boldsymbol{X}\sim\bigotimes_{i=1}^N P_{t,\boldsymbol{x},A^i}\;. \end{align}\]

Next, in this risk neutral setting, we develop an equivalent form of \(R^n\). We first expand \(R^n\) by subtracting and adding auxiliary terms, \[\begin{align} \label{eq:ExploitabilityTelescopingSum} R^n(\boldsymbol{\mathfrak{P}}) &= \mathbb{E}\left[ \sum_{r=1}^{T-1} C_r(\boldsymbol{X}_r, A^n_r) + C_T(\boldsymbol{X}_T, A^n_T) + V^{n*}_{T+1}(\boldsymbol{X}_{T+1}) \right] - \mathbb{E}\left[ \sum_{r=1}^{T-1}C_r(\boldsymbol{X}_r, A^n_r) + V^{n*}_T(\boldsymbol{X}_T) \right] \nonumber\\ &\quad + \mathbb{E}\left[ \sum_{r=1}^{T-1}C(\boldsymbol{X}_r, A^n_r) + V^{n*}_T(\boldsymbol{X}_T) \right] - V^{n*}_1(x_1,\dots,x_1) \nonumber\\ &= \mathbb{E}\left[ C_T(\boldsymbol{X}_T, A^n_T) + V^{n*}_{T+1}(\boldsymbol{X}_{T+1}) - V^{n*}_T(\boldsymbol{X}_T) \right]\nonumber\\ &\quad +\mathbb{E}\left[ \sum_{r=1}^{T-2} C_r(\boldsymbol{X}_r, A^n_r) + C_{T-1}(\boldsymbol{X}_{T-1}, A^n_{T-1}) + V^{n*}_{T}(\boldsymbol{X}_{T}) \right] \!-\! \mathbb{E}\left[ \sum_{r=1}^{T-2}C_r(\boldsymbol{X}_r, A^n_r) + V^{n*}_{T-1}(\boldsymbol{X}_{T-1}) \right] \nonumber\\ &\quad + \mathbb{E}\left[ \sum_{r=1}^{T-2}C_r(\boldsymbol{X}_r, A^n_r) + V^{n*}_{T-1}(\boldsymbol{X}_{T-1}) \right] - V^{n*}_1(x_1,\dots,x_1) \nonumber\\ &=\dots\dots \end{align}\tag{3}\] Eventually, we obtain the following expression, \[\begin{align} \label{eq:HeurDefStepwiseExploitability} R^n(\boldsymbol{\mathfrak{P}}) &= \sum_{t=1}^T \mathbb{E}\left[ C_t(\boldsymbol{X}_t,A^n_t) + V^{n*}_{t+1}(\boldsymbol{X}_{t+1}) - V^{n*}_t(\boldsymbol{X}_t) \right] \nonumber\\ &= \sum_{t=1}^T \mathbb{E}\Big[ \mathbb{E}\big[ C_t(\boldsymbol{X}_t,A^n_t) + V^{n*}_{t+1}(\boldsymbol{X}_{t+1}) \big| \boldsymbol{X}_t \big] - V^{n*}_t(\boldsymbol{X}_t) \Big] . \end{align}\tag{4}\] The reader might notice that 4 may be derived more concisely using telescoping sums under the expectation, rather than using 3 . The approach we here take, however, generalizes to the risk-averse case while the simpler approach does not. We refer to 68 for the related application in proofs. Moreover, we stress that the risk-averse analogue to the right hand side of 4 , see 4, provides a crucial foundation for deriving our main results which concerns risk-aversion beyond expectation. Therefore, we maintain working through 4 although simpler approach is possible in the risk-neural setting.

To distinguish the two representations, we call the right hand side of 4 total (stepwise) exploitability, while the original definition of \(R^n\) in 1 we call end exploitability. For clarification, we note that end exploitability compares the (aggregated) outcome of the original policy against the optimal (aggregated) outcome. Contrastingly, total exploitability compares the expected one-step cost plus the one-step ahead optimal value, conditional on the state at time \(t\), against the optimal value at time \(t\) (see also 2 ), and then averages over state at time \(t\). The total stepwise exploitability is obtained by then summing over these individual stepwise exploitabilities.

Next, we, heuristically, provide a few technical assumptions and lay the foundation for the upcoming discussion on how, in the absence of uniqueness, \(N\)pEs can be captured by MFEs. Let \(\delta_x\) be the Dirac delta measure concentrated at \(x\). For \(\boldsymbol{x}=(x_1,\dots,x_n)\), we define \(\overline{\delta}_{\boldsymbol{x}}:=\frac{1}{N}\sum_{n=1}^N\delta_{x^n}\) – the empirical measure. The conditions below are used to ensure the impact of an individual player is small, given a sufficiently large \(N\).

Assumption 1. The following is true for all players at all \(t\),

  • \(P_{t,\boldsymbol{x}, a} = P_{t, x^n, \overline{\delta}_{\boldsymbol{x}}, a}\) ,

  • \(C_t(\boldsymbol{x}, a) = C_t(x^n, \overline{\delta}_{\boldsymbol{x}}, a)\) ,

  • \(\mathfrak{p}^n_{t,\boldsymbol{x}} = \mathfrak{p}^n_{t,x^n,\overline{\delta}_{\boldsymbol{x}}}\),

  • \(\xi\mapsto P_{x,\xi,a}\), \(\xi\mapsto C_t(x,\xi,a)\), and \(\xi\mapsto\mathfrak{p}^n_{t,x,\xi}\) are continuous in \(\xi\in\mathcal{P}(\mathbb{X})\).6

Under this assumption, the dynamics of the \(N\)pG becomes \[\begin{align} \boldsymbol{X}_1=(x_1,\dots,x_1),\quad \boldsymbol{A}_t\sim\bigotimes_{n=1}^N\mathfrak{p}^n_{t,X^n_t,\overline{\delta}_{\boldsymbol{X}_t}}\,,\quad \boldsymbol{X}_{t+1}\sim\bigotimes_{n=1}^N P_{t,X^n_t,\overline{\delta}_{\boldsymbol{X}_t},A^n_t}\,. \end{align}\] We note that, eventually, the accuracy of the mean field approximation hinges on the modulus of continuity in 1 (iv). If the continuity is excessively rough, the approximation bound derived may become vacuous. For an illustrative example, we refer to 8.1.

Next, we introduce an analogous scenario but now in the MFG setting. We propose to characterize MFG from the perspective of state-action distribution, rather than through the policy of the representative player. In this approach, the MFG is described by a time-indexed vector of state-action distributions from \(\mathcal{P}(\mathbb{X}\times\mathbb{A})\). This perspective is frequently adopted in the discrete-time MFG literature for various reasons. For instance, in [26], it facilitates the application of the Kakutani fixed-point theorem, and in [45], it aids in algorithm development.

Based on \(\boldsymbol{\mathfrak{P}}\), we introduce an auxiliary probability \(\overline{\mathbb{P}}\) and auxiliary \(\overline{\xi}_1,\dots,\overline{\xi}_T\in\mathcal{P}(\mathbb{X})\). As previously, we omit the explicit dependence of these terms on \(\boldsymbol{\mathfrak{P}}\) in our notation for convenience. We construct \(\overline{\mathbb{P}}\) and \(\overline{\xi}_1,\dots,\overline{\xi}_T\in\mathcal{P}(\mathbb{X})\) such that, under \(\overline{\mathbb{P}}\), \[\begin{gather} \label{eq:MFPDynamics} \boldsymbol{X}_1=(x_1,\dots,x_1)\,,\quad\boldsymbol{A}_t\sim\bigotimes_{n=1}^N\mathfrak{p}^n_{t, X^n_t, \overline{\xi}_t}\,,\quad \boldsymbol{X}_{t+1}\sim\bigotimes_{n=1}^N P_{t, X^n_t, \overline{\xi}_t, A^n_t}\,. \end{gather}\tag{5}\] and \[\begin{gather} \overline{\xi}_1:=\delta_{x_1}\,,\quad \overline{\xi}_{t+1}(\cdot) := \frac{1}{N} \sum_{n=1}^N \overline{\mathbb{P}}\left[X^n_{t+1}\in\cdot\right]\,. \end{gather}\] Note that \(\overline{\xi}_t\)’s are (deterministic) elements from \(\mathcal{P}(\mathbb{X})\). It follows that the players are independent under \(\overline{\mathbb{P}}\). In short, \(\overline{\xi}_t\)’s represent the average across players of their individual state distributions, where players are mutually independent and use \(\boldsymbol{\mathfrak{P}}:=(\mathfrak{p}^1,\dots,\mathfrak{p}^N)\). For future reference, under \(\overline{\mathbb{P}}\), we have the following dynamics, independently across \(n\), \[\begin{align} \label{eq:DynPbar} X^n_1=x_1,\quad A^n_t\sim \mathfrak{p}^n_{t, X^n_t, \overline{\xi}_t},\quad X^{n}_{t+1}\sim P_{t, X^n_t, \overline{\xi}_t, A^n_t}\,. \end{align}\tag{6}\]

We next introduce a heuristic proposition, which essentially follow from Assumption 1, Hoeffding’s inequality, and some meticulous application of induction. The “\(\approx\)” symbol below is a notation for approximation. For rigorous counterparts of [prop:HeurMFApprox], including a precise meaning of “\(\approx\)", please see [prop:ProcConc] and [prop:EstDiffNPMean]. We also refer to [17] for a similar result, where they consider deterministic Lipschitz path-dependent policies.

Under 1, the following is true:

  • \(\mathbb{P}\left[\overline{\delta}_{\mathbf{X}_t} \approx \overline{\xi}_t\right]\approx 1\) for all \(t\). Moreover, this approximation remains valid uniformly if we vary one \(\mathfrak{p}^n\) in \(\mathbb{P}\),7 while still defining \(\overline{\xi}_t\) with the original \(\boldsymbol{\mathfrak{P}}\).

  • \(\mathbb{P}\left[X^n_{t}\in\cdot\right] \approx \overline{\mathbb{P}}\left[X^n_t\in\cdot\right]\) for each \(n\) and \(t\).

Note in [prop:HeurMFApprox] (a), we consider approximation under a \(\mathbb{P}\) where we may vary one \(\mathfrak{p}^n\). This aligns with the aforementioned hypothetical case in defining the end exploitability, where all other players maintain their original policies, while player-\(n\) may alter hers. The approximation accuracy relies on the modulous of continuity in 1 (iv) as well as other model parameters, such as \(N\) and the cardinality of \(\mathbb{X}\) and \(\mathbb{A}\). We refer to [prop:ProcConc] and [prop:EstDiffNPMean] for the related details.

The construction above only specifies the state marginal distributions of a vector on \(\mathcal{P}(\mathbb{X}\times\mathbb{A})\). To provide a complete description of the MFG scenario from the perspective of state-action distributions, we procedd to construct the full state-action distributions. For notational convenience, we write \[\begin{align} \mathbb{P}^{X^n_{t}}:=\mathbb{P}[X^n_{t}\in\cdot] \quad\text{and}\quad \overline{\mathbb{P}}^{X^n_t}:=\overline{\mathbb{P}}[X^n_t\in\cdot]. \end{align}\] For \(t=1,\dots,T\), we define \(\overline{\psi}_t\in\mathcal{P}(\mathbb{X}\times\mathbb{A})\) by setting \[\begin{align} \label{eq:HearDefpsi} \overline{\psi}_t(B\times A) = \frac{1}{N} \sum_{n=1}^N \int_B \mathfrak{p}^n_{t,x,\overline{\xi}_t}(A)\; \overline{\mathbb{P}}^{X^n_{t}}(\mathop{\mathrm{d \!}}x). \end{align}\tag{7}\] Clearly, the marginal distribution of \(\overline{\psi}_t\) on \(\mathbb{X}\), denoted by \(\xi^{\overline{\psi}_t}\), equals \(\overline{\xi}_t\). Additionally, by 6 , \[\begin{align} \label{eq:MFScene} \xi^{\overline{\psi}_{t+1}}(B) &= \frac{1}{N} \sum_{n=1}^N \overline{\mathbb{P}}^{X^n_{t+1}}(B) = \frac{1}{N} \sum_{n=1}^N \int_\mathbb{X}\int_\mathbb{A}P_{x,\overline{\xi}_t,a}(B) \mathfrak{p}^n_{t,x,\overline{\xi}_t}(\mathop{\mathrm{d \!}}a)\; \overline{\mathbb{P}}^{X^n_{t}}(\mathop{\mathrm{d \!}}x) \nonumber\\ &\quad = \int_{\mathbb{X}\times\mathbb{A}} P_{x,\overline{\xi}_t,a}(B) \; \overline{\psi}_t(\mathop{\mathrm{d \!}}x \mathop{\mathrm{d \!}}a). \end{align}\tag{8}\] We then construct the mean field scenario as \[\begin{align} \label{eq:HeurDefMFF} \overline{\Psi}:=(\overline{\psi}_1,\dots,\overline{\psi}_{T}). \end{align}\tag{9}\] Note that \(\overline{\Psi}\) depends solely on \(\boldsymbol{\mathfrak{P}}\). Moreover, \(\overline{\Psi}\) indeed represents a scenario in MFGs: in light of the law of large numbers, when all the infinitely many players adopt the same (randomized) policy, it leads to a deterministic flow of state-action distributions, which in turn statistically summarizes the movement and behaviour of the players. Furthermore, by letting \(\overline{\pi}^{\overline{\psi}_t}\) be the conditional distribution of action given the state induced by \(\overline{\psi}_t\), we have \[\begin{align} \label{eq:InducedMFPolicy} \overline{\pi}^{\overline{\psi}_t}_{x}(\{a\}) = \sum_{n=1}^N \frac{ \overline{\mathbb{P}}^{X^n_{t}}(\{x\})}{\sum_{m=1}^N \overline{\mathbb{P}}^{X^m_{t}}(\{x\})} \mathfrak{p}^n_{t,x,\overline{\xi}_t}(\{a\}) ,\quad x\in\mathbb{X},\;a\in\mathbb{A}. \end{align}\tag{10}\] This depicts the policy of the representative player in the MFG scenario. It can be shown that 9 and 10 provide equivalence characterization of the mean field scenario (see also 3).

In order to proceed, we introduce another auxiliary MDP in the MFG setting. Consider a \(\overline{\mathfrak{p}}= (\overline{\mathfrak{p}}_1,\dots,\overline{\mathfrak{p}}_T)\), where \(\overline{\mathfrak{p}}_t:\mathbb{X}\to\mathcal{P}(\mathbb{A})\) for \(t=1,\dots,T\). Additionally, for the representative player in MFG scenario \(\overline{\Psi}\), we introduce state-action processes \(\overline{X}_t\) and \(\overline{A}_t\). We extend the dependence of \(\overline{\mathbb{P}}\) from \(\boldsymbol{\mathfrak{P}}\) to \((\boldsymbol{\mathfrak{P}}, \overline{\mathfrak{p}})\) so that, under \(\overline{\mathbb{P}}\), the following dynamics hold: \[\begin{align} \label{eq:ReprDynamics} \overline{X}_1 = x_1, \quad \overline{A}_t \sim \overline{\mathfrak{p}}_{t, \overline{X}_t}, \quad \overline{X}_{t+1} \sim P_{t, \overline{X}_t, \overline{\xi}_t, \overline{A}_t}. \end{align}\tag{11}\] Following the discussion leading to 10 , the representative player naturally employs the policy induced by \(\overline{\Psi}\). More precisely, \[\begin{align} \label{eq:ReprPolicy} \overline{\mathfrak{p}}= \left(\overline{\pi}^{\overline{\psi}_1},\dots,\overline{\pi}^{\overline{\psi}_T}\right). \end{align}\tag{12}\] However, this policy need not to be optimal in the hypothetical case where the representative player may revise her policy while the others maintain their original policies. This leads to the following auxiliary optimization problem \[\begin{align} \min_{\overline{\mathfrak{p}}} \overline{\mathbb{E}}\left[ \sum_{t=1}^T C_t(\overline{X}_t, \overline{\xi}_t, \overline{A}_t) \right], \end{align}\] where we emphasize that \(\overline{\xi}_1,\dots,\overline{\xi}_T\) depend solely on \(\boldsymbol{\mathfrak{P}}\) but not \(\overline{\mathfrak{p}}\). We subsequently introduce the corresponding optimal value function \[\begin{align} \overline{V}^*_t(x) := \min_{\overline{\mathfrak{p}}} \overline{\mathbb{E}}\left[ \sum_{r=t}^T C_t(\overline{X}_r, \overline{\xi}_r, \overline{A}_r) \,\Bigg| \,\overline{X}_r=x \right], \quad t=1,\dots,T. \end{align}\] For convenience, we stipulate that \(\overline{V}^*_t\equiv 0\) for \(t>T\). We then have the following DPP \[\begin{align} \label{eq:MFDPP} \overline{V}^*_t(x) = \min_{\lambda} \overline{\mathbb{E}}\left[ \, C_t(x,\overline{\xi}_t,\overline{A}) + \overline{V}^*_{t+1}(\overline{X}) \,\right], \quad t=1,\dots,T, \end{align}\tag{13}\] where we have employed dummy random variables with dynamics: \[\begin{align} \overline{A}\sim\lambda,\quad \overline{X}\sim P_{t,x,\overline{\xi}_t,\overline{A}}. \end{align}\] The related mean field exploitabilities can be defined in a manner similar to earlier discussions. However, their introduction is postponed to slightly later for smoother narrative flow. Instead, we introduce below a crucial proposition. We refer to [prop:EstfTDiffcSstar] for the corresponding formal statement.

Under 1, we have \(\mathbb{P}\big[V^{n*}_t(\boldsymbol{X}_t) \approx \overline{V}^*_t(X^n_t)\big] \approx 1\).

Heuristically, establishing [prop:HeurApproxOptV] requires examining an altered \(\boldsymbol{\mathfrak{P}}\), where the original \(\mathfrak{p}^n\) is replaced by a concatenation of the original \(\mathfrak{p}^n\) up to time \(t\) with the optimal policy obtained from 2 after time \(t\). Under this altered \(\boldsymbol{\mathfrak{P}}\), [prop:HeurMFApprox] (a) together with Assumption 1 and a careful backward induction allows, with a high probability, an approximate reduction of 2 to 13 when \(\boldsymbol{x}\) in 2 is substituted by \(\boldsymbol{X}_t\), which eventually leads to [prop:HeurApproxOptV].

We are ready to heuristically derive one of our main results. By 4 , 1, [prop:HeurMFApprox] (a) and [prop:HeurApproxOptV], we have \[\begin{align} R^n(\boldsymbol{\mathfrak{P}}) \approx \sum_{t=1}^T \mathbb{E}\Big[ \overline{\mathbb{E}}\big[ C_t(X^n_t,\overline{\xi}_t,A^n_t) + \overline{V}^{*}_{t+1}(X^n_{t+1}) \big| X^n_t \big] - \overline{V}^{*}_t(X^n_t) \Big]. \end{align}\] Above, note that the only randomness inside \(\mathbb{E}[\cdot]\) is \(X^n_t\). Expanding the right hand side into integral form and invoking [prop:HeurMFApprox] (b), we yield \[\begin{align} R^n(\boldsymbol{\mathfrak{P}}) &\approx \sum_{t=1}^T \int_{\mathbb{X}} \left( \int_{\mathbb{A}} \left( C_t(x,\overline{\xi}_t,a) + \int_{\mathbb{X}}\overline{V}^*_{t+1}(y)P_{x,\overline{\xi}_t,a}(\mathop{\mathrm{d \!}}y) \right) \mathfrak{p}^n_{t,x,\overline{\xi}_t}(\mathop{\mathrm{d \!}}a) - \overline{V}^*_t(x) \right) \; \overline{\mathbb{P}}^{X^n_{t}} (\mathop{\mathrm{d \!}}x), \label{eqn:average-stepwise-approx} \end{align}\tag{14}\] Then, 14 together with 7 and 10 implies \[\begin{align} \label{eq:HeurPreCapture} \frac{1}{N} \sum_{n=1}^N R^n(\boldsymbol{\mathfrak{P}}) &\approx \sum_{t=1}^T \frac{1}{N} \sum_{n=1}^N \int_{\mathbb{X}} \left( \int_{\mathbb{A}} {!}{ \left( C_t(x,\overline{\xi}_t,a) + \int_{\mathbb{X}}\overline{V}^*_{t+1}(y)P_{x,\overline{\xi}_t,a}(\mathop{\mathrm{d \!}}y) \right) } \mathfrak{p}^n_{t,x,\overline{\xi}_t}(\mathop{\mathrm{d \!}}a) - \overline{V}^*_t(x) \right) \; \overline{\mathbb{P}}^{X^n_{t}} (\mathop{\mathrm{d \!}}x) \nonumber\\ &= \sum_{t=1}^T \int_{\mathbb{X}\times\mathbb{A}} \left( C_t(x,\overline{\xi}_t, a) + \int_{\mathbb{X}}\overline{V}^*_{t+1}(y)P_{x,\overline{\xi}_t,a}(\mathop{\mathrm{d \!}}y) - \overline{V}^*_t(x) \right) \overline{\psi}_t(\mathop{\mathrm{d \!}}x\mathop{\mathrm{d \!}}a) \nonumber\\ &= \sum_{t=1}^T \int_{\mathbb{X}} \left( \int_\mathbb{A}\left(C_t(x,\overline{\xi}_t, a) + \int_{\mathbb{X}}\overline{V}^*_{t+1}(y)P_{x,\overline{\xi}_t,a}(\mathop{\mathrm{d \!}}y) \right) \overline{\pi}^{\overline{\psi}_t}_x(\mathop{\mathrm{d \!}}a) - \overline{V}^*_t(x) \right) \xi^{\overline{\psi}_t}(\mathop{\mathrm{d \!}}x). \end{align}\tag{15}\] In view of 11 and 12 , we subsequently rewrite 14 into \[\begin{align} \label{eq:HeurCaptureExpn} \frac{1}{N} \sum_{n=1}^N R^n(\boldsymbol{\mathfrak{P}}) &\approx \sum_{t=1}^T \overline{\mathbb{E}}\left[ \overline{\mathbb{E}}\left[C_t(\overline{X}_t, \overline{\xi}_t, \overline{A}_t) + \overline{V}^*_{t+1}(\overline{X}_{t+1})\Big| \overline{X}_t \right] - \overline{V}^*_t(\overline{X}_{t}) \right], \end{align}\tag{16}\] In view of 13 , using the same reasoning that leads to 4 , we obtain the following approximation as one of our keys results \[\begin{align} \label{eq:ApproxAveNplayerExploitability} \frac{1}{N} \sum_{n=1}^N R^n(\boldsymbol{\mathfrak{P}}) &\approx \overline{\mathbb{E}}\left[ \sum_{t=1}^T C_t(\overline{X}_t, \overline{\xi}_t, \overline{A}_t) \right] - \min_{\overline{\mathfrak{p}}} \overline{\mathbb{E}}\left[ \sum_{t=1}^T C_t(\overline{X}_t, \overline{\xi}_t, \overline{A}_t) \right] =: \overline{R}(\overline{\Psi}). \end{align}\tag{17}\] Above, we call \(\overline{R}\) the mean field end exploitability. Here, we emphasize that that, in analogy with 1 , \(\overline{R}\) is also computed under the hypothetical case in MFG setting where all (infinitely many) other players maintain their original policies while the representative player may vary hers. Analogously to \(N\)pG, we term the right hand side of 16 as mean field total (stepwise) exploitability. Clearly, \(\overline{\Psi}\) is a \(\varepsilon\)-MFE if and only if \(\overline{R}(\overline{\Psi})\le\varepsilon\).

To proceed, we briefly highlight two important results. These results complement 17 in offering a comprehensive description of the approximation capabilities of MFGs. Firstly, as a (partial) enhancement of [prop:HeurMFApprox] (a), we have that \[\begin{align} \label{eq:ApproxStateActionEmp} \mathbb{P}\left[\overline{\delta}_{(\boldsymbol{X}, \boldsymbol{A})_t} \approx \overline{\psi}_t\right] \approx 1, \end{align}\tag{18}\] where the notation \(\overline{\delta}_{(\boldsymbol{X}, \boldsymbol{A})_t}\) is the empirical measure of \((\boldsymbol{X}, \boldsymbol{A})_t\). Secondly, we consider a \(\Psi=(\psi_1,\dots,\psi_T)\in\mathcal{P}(\mathbb{X}\times\mathbb{A})^{T}\) that satisfies \[\begin{align} \label{eq:HeurDefMFF2} \xi^{\psi_1}=\delta_{x_1} \quad\text{and}\quad \xi^{\psi_{t+1}}(B)=\int_{\mathbb{X}\times\mathbb{A}} P_{x,\xi^{\psi_t},a}(B)\psi_t(\mathop{\mathrm{d \!}}x \mathop{\mathrm{d \!}}a),\;B\in\mathcal{B}(\mathbb{X}),\,t=1,\dots,T-1, \end{align}\tag{19}\] i.e., \(\Psi\) can be related to some scenario in the MFG setting. For \(\psi \in \mathcal{P}(\mathbb{X}\times \mathbb{A})\), we define \(\overline{\pi}^{\psi}\) as the conditional distribution of actions given the state. This definition naturally leads to an action kernel that is solely dependent on the state of the representative player, i.e., the law of the random action remains invariant under changes in the empirical state distribution. We call \(\mathfrak{p}^{\Psi}=(\overline{\pi}^{\psi_1},\dots,\overline{\pi}^{\psi_T})\) the the policy induced by \(\Psi\). It can be shown that the homogeneous \(N\)-player scenario \(\overline{\boldsymbol{\mathfrak{P}}}^{\Psi} = (\mathfrak{p}^{\Psi},\dots,\mathfrak{p}^{\Psi})\) satisfy \[\begin{align} \label{eq:ApproxMFExploitability} \overline{R}(\Psi) \approx R^n(\overline{\boldsymbol{\mathfrak{P}}}^{\Psi}),\quad n=1,\dots,N. \end{align}\tag{20}\] Results such as 18 and 20 are generally better understood in extant literature compared to 17 . Nonetheless, proving 18 in the context of \(N\)pG with features like heterogeneous agents and non-asymptotics demands some additional treatment. Regarding 20 , while it is well-established in many risk-neutral settings (at least for when \(\Psi\) is a MFE), it still requires case-by-case analysis in various risk-averse or risk-sensitive settings. Establishing 20 in a general setting is another objective of this paper.

We summarize the discussion above into following heuristic theorems. For formal statements, we refer to 3 and 4, respectively.

Theorem 1. Suppose 1 and let \(\overline{\Psi}\) be constructed in 9 . Then, we have \(\mathbb{P}\left[\overline{\delta}_{(\boldsymbol{X}, \boldsymbol{A})_t} \approx \overline{\psi}_t\right] \approx 1\) and \(\frac{1}{N} \sum_{n=1}^N R^n(\boldsymbol{\mathfrak{P}})\approx\overline{R}(\overline{\Psi})\).

Theorem 2. For any \(\Psi=(\psi_1,\dots,\psi_T)\) that satisfies 19 , consider the induced \(N\)-player scenario \(\overline{\boldsymbol{\mathfrak{P}}}^{\Psi} = (\mathfrak{p}^{\Psi},\dots,\mathfrak{p}^{\Psi})\). Then, \(\overline{R}(\Psi) \approx R^n(\overline{\boldsymbol{\mathfrak{P}}}^{\Psi})\) for \(n=1,\dots,N\).

2.2 Comments on 1 and 2↩︎

In general, for an \(N\)pG with a large \(N\) that allows heterogeneous policies, it is exceedingly challenging to characterize the entire spectrum of (approximate) equilibria. It is, however, plausible that different combinations of individual policies could still result in similar population statistics in terms of state-action distribution. This leads to the natural emergence of an approximate refinement of equilibria in \(N\)pGs via state-action distributions. Under 1, in view of 1, we can categorize different \({\boldsymbol{\mathfrak{P}}}\)’s based on the corresponding \(\overline{\Psi}\)’s defined in 9 . Such approximate refinement eradicates redundancy stemming from permutations of player labels. It also conserves effort by accounting for the situation where distinct \(\mathfrak{P}\)’s could lead to similar state-action distributions. Beyond the refinement perspective, 1 and 2 together advocate for the use of \(\delta\text{-}\mathop{\mathrm{arg\,min}}_{\Psi}\overline{R}(\Psi)\), and ultimately the examination of \(\Psi\mapsto\overline{R}(\Psi)\). Indeed, 1 asserts that whenever an \(\varepsilon\)-equilibrium in the average sense is achieved in an \(N\)pG, there must be some \[\overline{\Psi}\in\delta\text{-}\mathop{\mathrm{arg\,min}}_{\Psi}\overline{R}(\Psi)\] that approximates the (empirical) state-action distributions of said \(\varepsilon\)-equilibrium. In other words, MFGs do not overlook equilibria in \(N\)pGs. Conversely, 2 demonstrates that any element from \(\varepsilon\text{-}\mathop{\mathrm{arg\,min}}_{\Psi}\overline{R}(\Psi)\) can be associated with a scenario in an \(N\)pG that is \(\delta\)-equilibrium, implying that \(\varepsilon\text{-}\mathop{\mathrm{arg\,min}}_{\Psi}\overline{R}(\Psi)\) does not produce nonsensical equilibria.

In concluding this section, we juxtapose 1 with results established in [15] and [17]. These results represent important progress in understanding the capturing capabilities of MFGs without requiring the uniqueness of equilibria, offering valuable insights that have informed our current research. [15] explores continuous-time games driven by diffusion with controlled drift. It is shown in [15] that, under suitable conditions, any subsequential limits of \(N\)-player approximate equilibria are weak MFE, a relaxed notation of MFE. A key strength of this result lies in its minimal assumptions on player policies in \(N\)pGs, in contrast to our 1 (iii) (iv). However, the asymptotic approach of [15] leaves some ambiguity in the relationship between an \(N\)pE for a finite \(N\) and the corresponding MFE. In contrast, 1 offer a non-asymptotic perspective to this issue under additional assumptions. [17] works in mostly in discrete-time setups, and utilizes a set of conditions that arguably parallels those in 1. They focus on the similarity between value functions from \(N\)pGs and MFGs at approximate equilibria. Our interest, however, lies more in linking (empirical) state-action distributions from both games with exploitabilities, although we do share some common heuristics as the concepts of propagation of chaos is usually indispensable in such analysis. We would also like to highlight our technical endeavor. For instance, we extend the discussion from discrete spaces to Polish spaces (though still in discrete time). We also remove some conditions that we do not deem necessary for our interest, such as the restriction to deterministic policies in \(N\)pGs.

2.3 Technical concerns↩︎

In this section, we will outline some of the challenges associated with formally establishing the heuristic results presented in 2.1 in a more general setting.

2.3.0.1 Exploitability under continuity constrain

In light of 1, it is natural to consider in \(N\)pGs an exploitability with candidate policies that also satisfy 1 (iii) (iv). Here we simply call this the constrained exploitability. Unfortunately, due to the lack of DPP for the related constrained problem in MDPs, the running optimal value function is not well-defined, hindering derivation of an analogue to 17 for constrained exploitability. However, this issue can be side-stepped by arguing that the constrained exploitability is approximately no different from the unconstrained one. We refer to 4 for more discussion. The proof of this side-stepping argument in fact largely overlaps with that of eq:ApproxAveNplayerExploitability.

2.3.0.2 Risk aversion

Following [33], we would like to incorporate into our models risk averse agents. Since[33] does not consider randomized actions, we use the extended version in [34] for illustration. Heuristically, a typical form of policy evaluation in risk averse MDPs reads (likewise for the Bellman equation) \[\begin{align} v_t(x) = \int_{\mathbb{A}} \big(c( x,a) + \rho(v_{t+1}(X^{x,a})) \big) \mathfrak{p}_{t,x}(\mathop{\mathrm{d \!}}a), \end{align}\] where \(X^{x,a}\) is generated from some distribution that may depends on \((x,a)\), and \(\rho\) could be a convex risk measures, say average value-at-risk (see [47]). The usual absence of linearity in \(\rho\) hinders the derivation of total stepwise exploitability, which is crucial for establishing 17 . This issue will be discussed with more details in 4.

In the following analysis, we use complete separable metric spaces (e.g., \(\mathbb{R}^d\)) for state and action spaces. This introduces additional technical complexity. In particular, we need to choose between weak or strong continuity in replacement of 1 (i) (iv), an issue that is trivial in finite environment as the two concepts coincides.8 We note that, weak continuity are in general a weaker condition, allowing, e.g., dynamics with atoms, a feature that is often desirable but not compatible with strong continuity. However, equipping \(({x,\xi, a})\mapsto P_{x,\xi, a}\) with weak continuity typically requires such continuity to present in not only all inputs jointly but also other components of the models, such as \(C\) and \(\mathfrak{p}^n_t\). In contrast, we later find out that imposing strong continuity on \((\xi,a)\mapsto P_{x,\xi,a}\) allows discontinuity in the \(x\) argument of \(P, C, \mathfrak{p}^n_t\), while still enabling certain desirable approximation property. This ultimately enables us to encompass in our approximation analysis a broader spectrum of \(N\)-player scenarios (at the expense of more stringent assumptions on \(\xi,a\) inputs of \(P\)). In the meantime, we would like to point out that jointly weak continuity, though not a necessity in approximation analysis, will be eventually needed in showing the existence of MFEs via Kakutani-Fan-Glicksberg fixed point theorem. There is not a universal rule on which continuity is more suitable. In our approximation analysis, we have opted to work with a mixture of continuities, while assuming joint weak continuity for establishing the existence of MFEs. Of course, the ideal case would be that all continuity assumptions are satisfied simultaneously.

2.3.0.4 Complex notations and formulas

The heuristic derivation in 2.1 already involves complex notations. A major reason for this complexity is the need to work with difference processes when comparing a scenario in \(N\)pGs against the counterpart in MFGs. These processes are potentially defined on different probability bases, which in turn demands further notations for handling related calculations. For the sake of simplicity, we have decided to focus solely on the operator perspective for the remainder of the paper. Another challenge arises from lengthy expressions like 15 , which, combined with numerous upcoming applications of the triangle inequality, as well as the consideration of risk aversion, add to the complexity. To reduce the notational complexity, we propose to adopt an abstract formulation. This approach proves particularly beneficial in the context of triangle inequalities, as it narrows down the search area for identifying key differences.

3 Setup and Preliminaries↩︎

We initiate our discussion by introducing some notations in 3.1 that are used throughout this paper. The dynamics of players are summarized in 3.2, followed by the definitions of transition operators in 3.3. In 3.4, we present a series of score operators that formulate the performance criteria for the players. The concept of state-action mean field flow is formally introduced in 3.5. 3.6 presents a crucial construction of the mean field flow associated with a specific \(N\)-player scenario. We compile the technical assumptions in 3.7. An important technical lemma that underscores our main results on mean field approximation is presented in 3.8. Lastly, in 3.9, we introduce several additional error terms.

3.1 Preliminary notations↩︎

We first summarize the rules of notations. Bold-faced notations are associated with \(N\)pGs, while overlined notations pertain to MFGs. Subscripts typically refer to arguments or parameters whose effects are more directly defined—such as time or (empirical) population distributions. Conversely, superscripts indicate the involvement of more complex operations during definition, like computing marginal or conditional distributions. Additionally, superscripts are used to index individual players in \(N\)pGs.

In what follows, we let \(t\in\{1,2,\dots,T\}\) be the time index. We fix the total number of players in the \(N\)pG for the remainder of the paper.

We say \(\beta\) is a subadditive modulus of continuity if \(\beta:[0,\infty]\to\overline{[}0,\infty]\) is non-decreasing, subaditive and satisfies \(\lim_{\ell\to0+}\beta(\ell) = 0\). It can be shown that \(\beta\) is concave. For convenience, we sometimes view \(\beta\) as an operator so that \(\beta f=\beta(f(\cdot))\) for any \([0,\infty]\)-valued function \(f\).

The state space \(\mathbb{X}\) is a complete separable metric space and \(\mathcal{B}(\mathbb{X})\) is the corresponding Borel \(\sigma\)-algebra. \(\mathcal{P}(\mathbb{X})\) is the set of probability measures on \(\mathcal{B}(\mathbb{X})\), we endow \(\mathcal{P}(\mathbb{X})\) with the weak topology, \(\mathcal{B}(\mathcal{P}(\mathbb{X}))\) is the corresponding Borel \(\sigma\)-algebra. We let \(\mathcal{E}(\mathcal{P}(\mathbb{X}))\) be the evaluation \(\sigma\)-algebra, i.e., \(\mathcal{E}(\mathcal{P}(\mathbb{X}))\) is the \(\sigma\)-algebra generated by sets \(\{\xi\in\mathcal{P}(\mathbb{X}):\int_{\mathbb{Y}}f(y)\xi(\mathop{\mathrm{d \!}}y) \in B\}\,\) for any real-valued bounded \(\mathcal{B}(\mathbb{X})\)-\(\mathcal{B}(\mathbb{R})\) measurable \(f\) and \(B\in\mathcal{B}(\mathbb{R})\). Due to 21, \(\mathcal{B}(\mathcal{P}(\mathbb{X}))=\mathcal{E}(\mathcal{P}(\mathbb{X}))\).

The action domain \(\mathbb{A}\) is another complete separable metric space and \(\mathcal{B}(\mathbb{A})\) is the corresponding Borel \(\sigma\)-algebra. \(\mathcal{P}(\mathbb{A})\) is the set of probability measures on \(\mathcal{B}(\mathbb{A})\), we endow \(\mathcal{P}(\mathbb{A})\) with the weak topology, \(\mathcal{B}(\mathcal{P}(\mathbb{A}))\) is the corresponding Borel \(\sigma\)-algebra, \(\mathcal{E}(\mathcal{P}(\mathbb{A}))\) is the evaluation \(\sigma\)-algebra. Again, by 21, we have \(\mathcal{B}(\mathcal{P}(\mathbb{A}))=\mathcal{E}(\mathcal{P}(\mathbb{A}))\).

The Cartesian product of metric spaces in this paper are always equipped with the metric equal to the sum of metrics of the component spaces and endowed with the corresponding Borel \(\sigma\)-algebra.

We let \(\mathcal{P}(\mathbb{X}\times\mathbb{A})\) be the set of probability measures on \(\mathcal{B}(\mathbb{X}\times\mathbb{A})\) endowed with weak topology. For \(\psi\in\mathcal{P}(\mathbb{X}\times\mathbb{A})\), we let \(\xi^{\psi}\) be the marginal measure of \(\psi\) on \(\mathbb{X}\). Additionally, we find it convenient to denote \(\Psi = (\psi_1, \dots, \pi_{T-1}) \in \mathcal{P}(\mathbb{X}\times \mathbb{A})^{T-1}\) and \(\Xi^\Psi = (\xi^{\psi_1}, \dots, \xi^{\psi_{T-1}})\).

We let \(\delta_y(B):=\mathbb{1}_{B}(y)\) be the Dirac measure at \(y\). Additionally, we write \(\overline{\delta}_{\boldsymbol{y}}:=\frac{1}{N}\sum_{n=1}^N\delta_{y^n}\), where \(\boldsymbol{y}=(y^1,\dots,y^N)\).

Let \(\mathbb{Y}\) be a complete separable metric space and \(\mathcal{B}(\mathbb{Y})\) be the corresponding Borel \(\sigma\)-algebra. Let \(\mathcal{P}(\mathbb{Y})\) be the set of probability measures on \(\mathcal{B}(\mathbb{Y})\), endowed with weak topology and Borel \(\sigma\)-algebra. \(B_b(\mathbb{Y})\) is the set of Borel measurable real-valued bounded functions on \(\mathbb{Y}\). Subsequently, \(C_b(\mathbb{Y})\subseteq B_b(\mathbb{Y})\) is the set of continuous functions, equipped with sup-norm. Let \(L>0\) and let \(C_{L-BL}(\mathbb{Y})\) be the set of \(L\)-Lipschitz continuous functions bounded by \(L\). For finite measures \(m,m'\) on \(\mathcal{B}(\mathbb{Y})\), we define the following norm and metric: \[\begin{gather} \|m\|_{L-BL} := \sup_{h\in C_{L-BL}}\frac{1}{2} \int_\mathbb{Y}h(y) m(\mathop{\mathrm{d \!}}y),\\ d_{L-BL}(m,m') := \|m-m'\|_{L-BL} = \sup_{h\in C_{L-BL}(\mathbb{Y})}\frac{1}{2}\left|\int_{\mathbb{Y}}h(y)m(\mathop{\mathrm{d \!}}y) - \int_{\mathbb{Y}}h(y)m'(\mathop{\mathrm{d \!}}y)\right|. \end{gather}\] We write \(C_{BL}\) and \(\|\cdot\|_{BL}\) for abbreviations of \(C_{1-BL}\) and \(\|\cdot\|_{1-BL}\). Clearly, \(\|m-m'\|_{L-BL} = L \|m-m'\|_{BL}\) and \(\|m-m'\|_{BL}\le 1\). Below we recall a well-known result regarding the relationship between \(\|\,\cdot\,\|_{BL}\) and weak convergence of probabilities; see, for example, [48].

Lemma 1. Let \((\mu_n)_{n\in\mathbb{N}}\subseteq\mathcal{P}(\mathbb{Y})\) and \(\mu\in\mathcal{P}(\mathbb{Y})\). Then, \(\lim_{n\to\infty}\|\mu_n-\mu\|_{BL}=0\) if and only if \(\mu_n\) converges weakly to \(\mu\).

Throughout the rest of the paper, we will consistently associate spaces with their respective Borel \(\sigma\)-algebras and consider exclusively measurable mappings. Consequently, notations like \(\mathcal{B}(\mathbb{X})\) and \(\mathcal{B}(\mathcal{P}(\mathbb{X}))\) will be omitted, except where specifically noted.

3.2 Controlled dynamics of the \(N\)pG↩︎

Throughout the rest of the paper, we will fix \(N\), the total number of players in the finite-player game. The scripts indicating this number is omitted for brevity. The only exception is the peripheral result, 6, which connects our non-asymptotic analysis to an asymptotic setting.

For \(t=1,\dots,T\), let \(P_t:\mathbb{X}\times\mathcal{P}(\mathbb{X})\times\mathbb{A}\to\mathcal{P}(\mathbb{X})\) be the transition kernel.

We define \(\Pi\) as the set of \(\pi:\mathbb{X}^N\to\mathcal{P}(\mathbb{A})\), i.e., \(\Pi\) is the set of Markovian action kernels. Let \(\widetilde{\Pi}\) be the set of \(\tilde{\pi}:\mathbb{X}\times\mathcal{P}(\mathbb{X})\to\mathcal{P}(\mathbb{A})\), i.e., \(\widetilde{\Pi}\) is the set of symmetric Markovian action kernels. Additionally, \(\overline{\Pi}\) is the set of \(\overline{\pi}:\mathbb{X}\to\mathcal{P}(\mathbb{A})\), i.e., \(\overline{\Pi}\) is the set of oblivious Markovian action kernels [49]. “Oblivious” refers to a player’s decision-making process that considers only their own state, ignoring the states of other players. Sometimes, it is convenient to extend elements in \(\overline{\Pi}\) (resp. \(\widetilde{\Pi}\)) to elements in \(\widetilde{\Pi}\) (resp. \(\Pi\)) by considering \(\overline{\pi}\) as constant in \(\mathcal{P}(\mathbb{A})\) (resp. redacting \(\boldsymbol{x}\) to \((x^n,\overline{\delta}_{\boldsymbol{x}})\) from some \(n\), depending on the context). In this sense, we have \(\overline{\Pi}\subseteq\widetilde{\Pi}\subseteq\Pi\).

Let \(N\in\mathbb{N}\) be the number of players. For the remainder of the paper, we will fix \(N\) and, unless necessary, omit it from the notation concerning the entire \(N\)pG. We use \(\mathfrak{p}^n=(\mathfrak{p}^n_1,\dots,\mathfrak{p}^n_{T-1})\in\Pi^{T-1}\) to denote the policy of player-\(n\), and \(\boldsymbol{\mathfrak{P}}=(\mathfrak{p}^1,\dots,\mathfrak{p}^N)\in\Pi^{N\times(T-1)}\) to denote the collection of policies in an \(N\)pG. Occasionally, we will use \(\boldsymbol{\mathfrak{P}}=(\boldsymbol{\mathfrak{P}}_1,\dots,\boldsymbol{\mathfrak{P}}_{T-1})\), where \(\boldsymbol{\mathfrak{P}}_t=(\mathfrak{p}^1_t,\dots,\mathfrak{p}^N_t)\in\Pi^N\) – this notation shows the time index in the policy for all players, while the former one shows the player index for each player, however, the set of action kernels considered is the same in both cases; this notation is also more convenient for compositions of operators, see e.g., 25 and 33 later.

A \(\boldsymbol{\mathfrak{P}}\in\Pi^{N\times(T-1)}\) corresponds to an \(N\)pG scenario with the following dynamics. The position of each player at \(t=1\) is drawn independently from \(\mathring{\xi}\in\mathcal{P}(\mathbb{X})\). At time \(t=2,\dots,T-1\), given the realization of state \(\boldsymbol{x}=(x^1,\dots,x^N)\), the generation of the actions of all player is governed by \(\bigotimes_{n=1}^N \mathfrak{p}^n_{t,\boldsymbol{x}}\). In particular, for \(\widetilde{\boldsymbol{\mathfrak{P}}}\in\widetilde{\Pi}^{N\times(T-1)}\), the generation is then governed by \(\bigotimes_{n=1}^N\tilde{\mathfrak{p}}^n_{t,x^n,\overline{\delta}_{\boldsymbol{x}}}\). Moreover, given the realization of actions \(\boldsymbol{a} = (a^1,\dots,a^N)\), the transition of the states of all \(N\) players is governed by \(\bigotimes_{n=1}^N P_{t,\boldsymbol{x},\overline{\delta}_{\boldsymbol{x}},a^n}\). In summary, we have the following transition dynamics: \[\begin{align} \label{eq:NpDyn} \boldsymbol{X}_1=(x_1,\dots,x_1),\quad \boldsymbol{A}_t\sim\bigotimes_{n=1}^N\tilde{\mathfrak{p}}^n_{t,X^n_t,\overline{\delta}_{\boldsymbol{X}_t}}\,,\quad \boldsymbol{X}_{t+1}\sim\bigotimes_{n=1}^N P_{t,X^n_t,\overline{\delta}_{\boldsymbol{X}_t}, A^n_t}. \end{align}\tag{21}\] For notional convenience, in what follows, we use \(\boldsymbol{\lambda}=(\lambda^1,\dots,\lambda^N)\in\mathcal{P}(\mathbb{A})^N\) for an \(N\)-tuple of elements from \(\mathcal{P}(\mathbb{A})\).

3.3 Transition operators↩︎

In this section, we introduce several operators related to state transitions. All measurability issues can be addressed by referring to 20, and as such, the related details are not reiterated here.

In order to introduce the transition and expectation operators for \(N\)pGs, we first define an auxiliary notation. With \(t=1,\dots,T-1\) and \((x,\xi,\lambda)\in\mathbb{X}\times\mathcal{P}(\mathbb{X})\times\mathcal{P}(\mathbb{A})\), we let \[\begin{align} \label{eq:DefQ} Q^\lambda_{t,x,\xi}(B) := \int_{\mathbb{A}} P_{t,x,\xi,a}(B)\lambda(\mathop{\mathrm{d \!}}a), \quad B\in\mathcal{B}(\mathbb{X}). \end{align}\tag{22}\] Consider additionally \(\boldsymbol{x}\in\mathbb{X}^N\) and \(\boldsymbol{\lambda}\in\mathcal{P}(\mathbb{A})^N\). By utilizing standard tools in probability (e.g., [50]), we obtain \[\begin{align} \label{eq:IndIntProdQ} &\int_{\mathbb{X}^N} f\left(\boldsymbol{y}\right) \left[\bigotimes_{n=1}^N Q^{\lambda^n}_{t,x^n,\xi}\right]\left(\mathop{\mathrm{d \!}}\boldsymbol{y}\right)\nonumber\\ &\quad=\int_{\mathbb{A}^N} \int_{\mathbb{X}^N}f\left(\boldsymbol{y}\right)\left[\bigotimes_{n=1}^N P_{t,x^n,\overline{\delta}_{\boldsymbol{x}},a^n}\right]\left(\mathop{\mathrm{d \!}}\boldsymbol{y}\right) \left[\bigotimes_{n=1}^N\lambda^n\right]\left(\mathop{\mathrm{d \!}}\boldsymbol{a}\right),\quad f\in B_b(\mathbb{X}^N). \end{align}\tag{23}\]

We are ready to introduce the transition and expectation operators for \(N\)pGs. Given \(\boldsymbol{\pi}=(\pi^1,\dots,\pi^N)\in\Pi^N\), we define \[\begin{align} \label{eq:DefT} \boldsymbol{T}^{\boldsymbol{\pi}}_t f\left(\boldsymbol{x}\right) := \int_{\mathbb{X}^N}f(\boldsymbol{y}) \left[\bigotimes_{n=1}^N Q^{\pi^n_{\boldsymbol{x}}}_{t,x^n,\overline{\delta}_{\boldsymbol{x}}}\right]\left(\mathop{\mathrm{d \!}}\boldsymbol{y}\right),\; t=1,\dots,T-1, \quad f\in B_b(\mathbb{X}^N). \end{align}\tag{24}\] With \(\boldsymbol{\mathfrak{P}}\in\Pi^{N\times(T-1)}\), we continue to define \[\begin{align} \label{eq:DefcT} {\boldsymbol{\mathcal{T}}}^{\boldsymbol{\mathfrak{P}}}_{s,s}f := f \quad\text{and}\quad {\boldsymbol{\mathcal{T}}}^{\boldsymbol{\mathfrak{P}}}_{s,t}f\left(\boldsymbol{x}\right) := \boldsymbol{T}^{\boldsymbol{\mathfrak{P}}_s}_s\circ\cdots\circ\boldsymbol{T}^{\boldsymbol{\mathfrak{P}}_{t-1}}_{t-1} f\left(\boldsymbol{x}\right), \;1\le s < t \le T, \quad f\in B_b(\mathbb{X}^N). \end{align}\tag{25}\] We also define the following expectation operator \[\begin{align} \label{eq:DeffT} \mathring{\boldsymbol{\mathbb{T}}} f := \int_{\mathbb{X}^N}f(\boldsymbol{y})\mathring{\xi}^{\otimes N}(\mathop{\mathrm{d \!}}\boldsymbol{y}) \quad\text{and}\quad {\boldsymbol{\mathbb{T}}}^{\boldsymbol{\mathfrak{P}}}_{t}f := \mathring{\boldsymbol{\mathbb{T}}} \circ {\boldsymbol{\mathcal{T}}}^{\boldsymbol{\mathfrak{P}}}_{1,t}f,\; t=1,\dots,T, \quad f\in B_b(\mathbb{X}^N). \end{align}\tag{26}\] It is clear that \({\boldsymbol{\mathbb{T}}}^{\boldsymbol{\mathfrak{P}}}_{t}\) is a non-negative linear functional on \(B_b(\mathbb{X}^N)\) with operator norm of \(1\). It follows that \({\boldsymbol{\mathbb{T}}}^{\boldsymbol{\mathfrak{P}}}_{t}\) is also non-decreasing. Additionally, the definitions provided above are applicable to symmetric/oblivious Markovian policies.

Below we introduce the analogous transition and expectation operators for MFGs. With \(\xi\in\mathcal{P}(\mathbb{X})\) and \(\tilde{\pi}\in\widetilde{\Pi}\), we define \[\begin{align} \label{eq:DefMeanT} \overline{T}^{\tilde{\pi}}_{t,\xi} h(x) := \int_{\mathbb{X}} h(y) Q^{\tilde{\pi}_{x,\xi}}_{t,x,\xi}(\mathop{\mathrm{d \!}}y),\; t=1,\dots,T-1, \quad h\in B_b(\mathbb{X}). \end{align}\tag{27}\] Furthermore, with \(\Xi=(\xi_1,\cdots,\xi_{T-1})\in\mathcal{P}(\mathbb{X})^T\) and \(\tilde{\mathfrak{p}}=(\tilde{\mathfrak{p}}_1,\dots,\tilde{\mathfrak{p}}_{T-1})\in\widetilde{\Pi}^{T-1}\), we define \[\begin{align} \label{eq:DefMeancT} \overline{\mathcal{T}}^{\tilde{\mathfrak{p}}}_{s,s,\Xi}h:=h \quad\text{and}\quad \overline{\mathcal{T}}^{\tilde{\mathfrak{p}}}_{s,t,\Xi}h := \overline{T}^{\tilde{\mathfrak{p}}_s}_{s,\xi_s}\circ\cdots\circ\overline{T}^{\tilde{\mathfrak{p}}_{t-1}}_{t-1,\xi_{t-1}} h,\; 1\le s<t\le T,\quad h\in B_b(\mathbb{X}). \end{align}\tag{28}\] In a similar manner as before, we define the expectation operator in MFG as \[\begin{align} \label{eq:DefMeanfT} \mathring{\overline{\mathbb{T}}}h:=\int_\mathbb{X}h(x)\mathring\xi(\mathop{\mathrm{d \!}}x) \quad\text{and}\quad \overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{t,\Xi}h := \mathring{\overline{\mathbb{T}}} \circ \overline{\mathcal{T}}^{\tilde{\mathfrak{p}}}_{1,t,\Xi}h,\, t=1,\dots,T,\quad h\in B_b(\mathbb{X}). \end{align}\tag{29}\] Clearly, \(\overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\) is a non-negative linear functional on \(B_b(\mathbb{X})\), and thus \(\overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\) is non-decreasing. Moreover, it can be verified via induction that \(D\mapsto\overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\mathbb{1}_{D}\) is a probability measure on \(\mathcal{B}(\mathbb{X})\), and we denote it by \(\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\). Clearly, \(\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}}_{1,\Xi}=\mathring{\xi}\).

3.4 Score operators↩︎

The operators discussed in this section, combined with the assumptions detailed in 3.7 later, abstractly depict the performance evaluation process that adheres to a finite horizon dynamic programming principle.

For \(N\)pGs, with \(\boldsymbol{\lambda}\in\mathcal{P}(\mathbb{A})^N\), we define an operators \(\boldsymbol{G}^{\boldsymbol{\lambda}}_{t}: B_b(\mathbb{X}^N) \to B_b(\mathbb{X}^N)\). For MFGs, with \(\lambda\in\mathcal{P}(\mathbb{A})\) and \(\xi\in\mathcal{P}(\mathbb{X})\), we consider \(\overline{G}^{\lambda}_{t,\xi}: B_b(\mathbb{X})\to B_b(\mathbb{X})\). Both \(\boldsymbol{G}^{\boldsymbol{\lambda}}_t\) and \(\overline{G}^\lambda_{t,\xi}\) are employed to construct mappings that facilitate the backward induction of value functions, transitioning them from time step \(t+1\) to time step \(t\). Additionally, we consider real-valued functionals \(\mathring{\boldsymbol{\mathbb{G}}}\) and \(\mathring{\overline{\mathbb{G}}}\) on \(B_b(\mathbb{X}^N)\) and \(B_b(\mathbb{X})\), respectively. These functionals facilitate the final step of backward induction, yielding a real number as the outcome score. Our abstract formulation with \(\boldsymbol{G}^{\boldsymbol{\lambda}}_t\) and \(\overline{G}^\lambda_{t,\xi}\) aims to encompasses not only Bellman operators commonly studied in risk-neutral mean field MDPs (e.g., [26] and [39]), but also risk-averse variant such as the example provided in 8.3. As a by-product, certain extended form of cost is also allowed, for example, 48 . However, we note that, while Bellman operators incorporating Kullback–Leibler divergence (e.g., [51]) can also be cast into our abstraction, the application is currently limited to discrete settings because Kullback–Leibler divergence is not compatible with latter assumptions (6 and 10) in continuous spaces such as \(\mathbb{R}^d\).

Throughout the paper, we always assume the measurabilities below:

  • For any \(\boldsymbol{u}\in B_b(\mathbb{X}^N)\), the mapping \((\boldsymbol{x},\boldsymbol{\lambda})\mapsto\boldsymbol{G}^{\boldsymbol{\lambda}}_{t}\boldsymbol{u}(\boldsymbol{x})\) is \(\mathcal{B}(\mathbb{X}^N\times\mathcal{P}(\mathbb{A})^N)\)-\(\mathcal{B}(\mathbb{R})\) measurable.

  • For any \(v\in B_b(\mathbb{X})\), the mapping \((x,\xi,\lambda)\mapsto \overline{G}^{\lambda}_{t,\xi} v(x)\) is \(\mathcal{B}(\mathbb{X}\times\mathcal{P}(\mathbb{X})\times\mathcal{P}(\mathbb{A}))\)-\(\mathcal{B}(\mathbb{R})\) measurable.

In most cases, these measurabilities can be verified using standard tools such as listed in [50].

\(\boldsymbol{G}^{\boldsymbol{\lambda}}_t\) is designed to evaluate the performance of player-1 in an \(N\)pG. Under certain symmetry assumptions (see 4 (iv) later), \(\boldsymbol{G}^{\boldsymbol{\lambda}}_t\) can also be applied to assess other players through permutations. It is crucial that \(\boldsymbol{G}^{\boldsymbol{\lambda}}_t\) and \(\overline{G}^\lambda_{t,\xi}\) maintain a specific relationship to facilitate approximating an \(N\)pG with a MFG. These assumptions will be introduced in 3.7. For examples of \(\boldsymbol{G}^{\boldsymbol{\lambda}}_t\) and \(\overline{G}^\lambda_{t,\xi}\), we refer to 8.3.

In the \(N\)pG, we assume all players aim to minimize their scores. More precisely, we let \(\boldsymbol{U}\in B_b(\mathbb{X})\) be the terminal cost function. With the score operator \(\boldsymbol{\mathbb{S}}^{\boldsymbol{\mathfrak{P}}}_{T}\) introduced later in 34 , player-\(1\) aims to find \[\begin{align} \label{eq:NplayerObj} \inf_{\mathfrak{p}^1\in\widetilde{\Pi}^{T-1}_\vartheta}\boldsymbol{\mathbb{S}}^{\boldsymbol{\mathfrak{P}}}_{T}\boldsymbol{U}, \end{align}\tag{30}\] where \(\vartheta\) is a subadditive modular of continuity, and \(\widetilde{\Pi}_\vartheta\) is an subset of \(\widetilde{\Pi}\) such that for any \(\tilde{\pi}\in\widetilde{\Pi}_\vartheta\) we have \(\xi\mapsto\tilde{\pi}(x,\xi)\) is \(\vartheta\)-continuous in \(\mathcal{P}(\mathbb{X})\) for any \(x\in\mathbb{X}\), where \(\mathcal{P}(\mathbb{X})\) and \(\mathcal{P}(\mathbb{A})\) are equipped with \(d_{BL}\). We call \(\tilde{\mathfrak{p}}\in\widetilde{\Pi}^{T-1}_\vartheta\) the \(\vartheta\)-symmetrically continuous policy. This objective aligns with the scenario we aim to explore in this paper, where all players employ a \(\vartheta\)-symmetrically continuous policy. Heuristically, this scenario might arise from practical situations where players have limited accuracy or certainty in perceiving the empirical population distribution, making them reluctant to drastically alter their behavior in response to changes in the population. Besides, to highlight the technical importance of symmetrical continuity, we note that a mild \(\vartheta\) is typically required for the mean field approximation to be effective. For an illustrative example, please refer to 8.1.

To facilitate the comparison between \(N\)pGs and MFGs, we consider the following hypothetical objective in MFG from the perspective of the representative player. Let \(V\in B_b(\mathbb{X}\times\mathcal{P}(\mathbb{X}))\) be the terminal cost and write \(V_\xi(x):=V(x,\xi)\). Given a \(\Xi\in\mathcal{P}(\mathbb{X})^{T-1}\) and \(\xi_T\in\mathcal{P}(\mathbb{X})\), with the score operator \(\overline{\mathbb{S}}^{\tilde{\mathfrak{p}}}_{T,\Xi}\) introduced later in 40 , the representative player aims to find \[\begin{align} \label{eq:MFObj} \inf_{\tilde{\mathfrak{p}}\in\widetilde{\Pi}^{T-1}}\overline{\mathbb{S}}^{\tilde{\mathfrak{p}}}_{T,\Xi} V_{\xi_T}, \end{align}\tag{31}\] Above, \(\tilde{\mathfrak{p}}\in\widetilde{\Pi}^{T-1}\) can be effectively replaced by \(\bar\mathfrak{p}\in\overline{\Pi}^{T-1}\) as \(\Xi\) and \(\xi_{T}\) are fixed. Note that 31 corresponds to the hypothetical case where the representative player is allowed to revise her policy while other players maintain their original behaviour so that \(\Xi\) and \(\xi_{T}\) remain unchanged.

In the remainder of this section, we will provide the definitions of the aforementioned score operators.

For the \(N\)pG, we let \(\boldsymbol{\pi}=(\pi^1,\dots,\pi^N)\in\Pi^N\) and \(\boldsymbol{\mathfrak{P}}=(\boldsymbol{\mathfrak{P}}_1,\dots,\boldsymbol{\mathfrak{P}}_{T-1})\in\Pi^{N\times(T-1)}\). For \(\boldsymbol{u}\in B_b(\mathbb{X}^N)\) we define \[\begin{gather} \boldsymbol{S}^{\boldsymbol{\pi}}_{t}\boldsymbol{u}({\boldsymbol{x}}) := \boldsymbol{G}^{(\pi^n_{\boldsymbol{x}})_{n=1}^N}_{t} \boldsymbol{u} (\boldsymbol{x}), \quad t=1,\dots,T-1,\tag{32}\\ \boldsymbol{\mathcal{S}}^{\boldsymbol{\mathfrak{P}}}_{s,s}\boldsymbol{u} = \boldsymbol{u} \quad\text{and}\quad \boldsymbol{\mathcal{S}}^{\boldsymbol{\mathfrak{P}}}_{s,t}\boldsymbol{u} := \boldsymbol{S}^{\boldsymbol{\mathfrak{P}}_s}_{s}\circ\cdots\circ \boldsymbol{S}^{\boldsymbol{\mathfrak{P}}_{t-1}}_{t-1}\boldsymbol{u},\quad 1\le s<t\le T,\tag{33}\\ \boldsymbol{\mathbb{S}}^{\boldsymbol{\mathfrak{P}}}_{t} \boldsymbol{u} := \mathring{\boldsymbol{\mathbb{G}}} \circ \boldsymbol{\mathcal{S}}^{\boldsymbol{\mathfrak{P}}}_{1,t}\boldsymbol{u},\quad t=1,\dots,T.\tag{34} \end{gather}\] We present the associated optimization operators below. Note that these optimization operators are not directly related to 30 , as the constrain in 30 hinders the derivation of dynamic programming principle. Nevertheless, the following operators serve as important auxiliary tools in our analysis. We first define the (one-step) Bellman operator \[\begin{align} \label{eq:DefSstar} \boldsymbol{S}^{*\boldsymbol{\pi}}_t \boldsymbol{u}(\boldsymbol{x}) := \inf_{\lambda\in\mathcal{P}(\mathbb{A})} G^{(\lambda,\pi^2_{\boldsymbol{x}},\dots,\pi^N_{\boldsymbol{x}})}_t \boldsymbol{u}(\boldsymbol{x}), \quad t=1,\dots,T-1. \end{align}\tag{35}\] We note that \(\boldsymbol{S}^{*\boldsymbol{\pi}}_t \boldsymbol{u}\) does not depend on \(\pi^1\) but we insist on this notation for the sake of neatness. Under suitable conditions (see 14), we further define: \[\begin{gather} \boldsymbol{\mathcal{S}}^{*\boldsymbol{\mathfrak{P}}}_{s,s} \boldsymbol{u} := \boldsymbol{u} \quad\text{and}\quad \boldsymbol{\mathcal{S}}^{*\boldsymbol{\mathfrak{P}}}_{s,t} \boldsymbol{u} := \boldsymbol{S}^{*\boldsymbol{\mathfrak{P}}_s}_s \circ \cdots \circ \boldsymbol{S}^{*\boldsymbol{\mathfrak{P}}_{t-1}}_{t-1} \boldsymbol{u},\quad 1\le s<t\le T,\tag{36}\\ \boldsymbol{\mathbb{S}}^{*\boldsymbol{\mathfrak{P}}}_{t} \boldsymbol{u} := \mathring{\boldsymbol{\mathbb{G}}}\circ\boldsymbol{\mathcal{S}}^{*\boldsymbol{\mathfrak{P}}}_{1,t} \boldsymbol{u},\quad t=1,\dots,T.\tag{37} \end{gather}\] Again, we note that \(\boldsymbol{\mathcal{S}}^{*\boldsymbol{\mathfrak{P}}}_{s,t}\) and \(\boldsymbol{\mathbb{S}}^{*\boldsymbol{\mathfrak{P}}}_{t}\) are constant in \(\mathfrak{p}^1\).

Regarding MFGs, let \(\xi\in\mathcal{P}(\mathbb{X})\), \(\tilde{\pi}\in\widetilde{\Pi}\), \(\Xi=(\xi_1,\dots,\xi_{T-1})\in\mathcal{P}(\mathbb{X})^{T-1}\), and \(\tilde{\mathfrak{p}}=(\tilde{\mathfrak{p}}_1,\dots,\tilde{\mathfrak{p}}_{T-1})\in{\widetilde{\Pi}}^{T-1}\). To align with previous setting, unless specified otherwise, we set \(\xi_1=\mathring\xi\). For \(v\in B_b(\mathbb{X})\) we define \[\begin{gather} \overline{S}^{\tilde{\pi}}_{t,\xi} v(x) := \overline{G}^{\tilde{\pi}_{x,\xi}}_{t,\xi} v(x),\quad t=1,\dots,T-1,\tag{38}\\ \overline{\mathcal{S}}^{\tilde{\mathfrak{p}}}_{s,s,\Xi}v := v\quad\text{and}\quad \overline{\mathcal{S}}^{\tilde{\mathfrak{p}}}_{s,t,\Xi} v := \overline{S}^{\tilde{\mathfrak{p}}_s}_{s,\xi_s} \circ \cdots \circ \overline{S}^{\tilde{\mathfrak{p}}_{t-1}}_{t-1,\xi_{t-1}} v , \; 1\le s<t\le T,\tag{39}\\ \overline{\mathbb{S}}^{\tilde{\mathfrak{p}}}_{t,\Xi} := \mathring{\overline{\mathbb{G}}} \circ \overline{\mathcal{S}}^{\tilde{\mathfrak{p}}}_{1,t,\Xi} v, \quad t=1,\dots,T. \tag{40} \end{gather}\] Note that \(\overline{\mathcal{S}}^{\tilde{\mathfrak{p}}}_{s,t,\Xi}\) depend on \(\Xi\) only through \(\xi_s,\dots,\xi_{t-1}\). In addition, we define the corresponding Bellman operator as \[\begin{align} \label{eq:DefMeanSstar} \overline{S}^{*}_{t,\xi} v(x) := \inf_{\lambda\in\mathcal{P}(\mathbb{A})}\overline{G}^{\lambda}_{t,\xi} v(x),\quad t=1,\dots,T-1. \end{align}\tag{41}\] Under suitable conditions (see 14), we further define: \[\begin{gather} \overline{\mathcal{S}}^*_{s,s,\Xi} v := v\quad\text{and}\quad \overline{\mathcal{S}}^*_{s,t,\Xi} v := \overline{S}^*_{s,\xi_s} \circ\cdots\circ \overline{S}^*_{t-1,\xi_{t-1}} v,\; 1\le s<t\le T,\tag{42}\\ \overline{\mathbb{S}}^*_{t,\Xi} v := \mathring{\overline{\mathbb{G}}} \circ \overline{\mathcal{S}}^*_{t,\Xi} v,\quad t=1,\dots,T.\tag{43} \end{gather}\]

Below we make a few obvious observations. Under the setting of 14 (with technical assumptions to be introduced in 3.7), we have \[\begin{align} \boldsymbol{\mathcal{S}}^{*\boldsymbol{\mathfrak{P}}}_{s,t} \boldsymbol{u} \le \boldsymbol{\mathcal{S}}^{(\mathfrak{p},\mathfrak{p}^2,\dots,\mathfrak{p}^N)}_{s,t} \boldsymbol{u},\;\mathfrak{p}\in\boldsymbol{\Pi} \quad\text{and}\quad \overline{\mathcal{S}}^*_{s,t,\Xi}v \le \overline{\mathcal{S}}^{\tilde{\mathfrak{p}}}_{s,t,\Xi}v,\;\tilde{\mathfrak{p}}\in\widetilde{\boldsymbol{\Pi}}. \end{align}\] Similar results holds for \(\boldsymbol{\mathbb{S}}^{*\boldsymbol{\mathfrak{P}}}_t\) and \(\overline{\mathbb{S}}^*_{t,\Xi}\) as well. Moreover, there exist \(\mathfrak{p}^*\in\Pi^{T-1}\) and \(\overline{\mathfrak{p}}^*\in\overline{\Pi}^{T-1}\) that attain the optimality above (for fixed \(\boldsymbol{\mathfrak{P}}, \boldsymbol{u}\) and \(\Xi, v\)), that is, \[\begin{gather} \boldsymbol{\mathcal{S}}^{*\boldsymbol{\mathfrak{P}}}_{s,t} \boldsymbol{u}=\boldsymbol{\mathcal{S}}^{(\mathfrak{p}^*,\mathfrak{p}^2,\dots,\mathfrak{p}^N)}_{s,t} \boldsymbol{u},\quad \boldsymbol{\mathbb{S}}^{*\boldsymbol{\mathfrak{P}}}_{t} \boldsymbol{u} = \boldsymbol{\mathbb{S}}^{(\mathfrak{p}^*,\mathfrak{p}^2,\dots,\mathfrak{p}^N)}_{t} \boldsymbol{u},\\ \overline{\mathcal{S}}^*_{s,t,\Xi} v = \overline{\mathcal{S}}^{\overline{\mathfrak{p}}^*}_{s,t,\Xi} v,\quad \overline{\mathbb{S}}^*_{t,\Xi} v = \overline{\mathbb{S}}^{\overline{\mathfrak{p}}^*}_{t,\Xi} v. \end{gather}\]

For the rest of this paper, functions from \(B_b(\mathbb{X})\), when acted upon by \(\boldsymbol{G}^{\boldsymbol{\lambda}}_t\) and related operators, are treated as functions from \(B_b(\mathbb{X}^N)\) that are constant in \(x^2,\dots,x^N\).

3.5 State-action mean field flow↩︎

It can be beneficial to describe a scenario in MFG using a flow of state-action distributions. This section is dedicated to explaining this concept. Recall that \(\xi^{\psi}\) denotes the marginal distribution of \(\psi\) on \(\mathbb{X}\), where \(\psi\in\mathcal{P}(\mathbb{X}\times\mathbb{A})\).

Let \(\Psi=(\psi_{1},\dots,\psi_{T-1})\in\mathcal{P}(\mathbb{X}\times\mathbb{A})^{T-1}\). We say \(\Psi\) is a mean field flow (on \(\mathbb{X}\times\mathbb{A}\) with initial marginal state distribution \(\mathring{\xi}\)), if \(\Psi\) satisfies \[\begin{align} \label{eq:DefMFF} \xi^{\psi_1}=\mathring{\xi}\quad\text{and}\quad \xi^{\psi_{t+1}}(B) = \int_{\mathbb{X}\times\mathbb{A}} P_{t,y,\xi^{\psi_t},a}(B) \psi_t(\mathop{\mathrm{d \!}}y\mathop{\mathrm{d \!}}a),\; B\in\mathcal{B}(\mathbb{X}),\, t=1,\dots,T-2. \end{align}\tag{44}\] For clarification, we point out that, based on the heuristics of the law of large numbers, all of the infinitely many independent players employing the same (randomized) policy will result in a deterministic flow of state-action distributions. This flow should comply with the evolution equation 44 imposed on the state marginal. Moreover, these distributions summarize the movements and behaviors of the players and can be viewed as a scenario in MFGs, possibly not at equilibrium.

The policy of the representative player can be characterized by a mean field flow via conditioning. To elucidate this procedure, we first revisit a pertinent notion in current context, that of the regular conditional distribution. Let \(\psi\in\mathcal{P}(\mathbb{X}\times\mathbb{A})\) and consider \(Y(x,a):=a\). By [50], we have \(\mathcal{B}(\mathbb{X}\times\mathbb{A})=\mathcal{B}(\mathbb{X})\otimes\mathcal{B}(\mathbb{A})\). Additionally, \(\mathcal{B}(\mathbb{X})\otimes\{\emptyset,\mathbb{A}\}\) is a sub-\(\sigma\)-algebra of \(\mathcal{B}(\mathbb{X})\otimes\mathcal{B}(\mathbb{A})\). By [48] Theorem 10.4.8, Example 10.4.9 and Example 6.5.2, under \(\psi\), there exists a \(\varpi^\psi:(\mathbb{X}\times\mathbb{A},\mathcal{B}(\mathbb{X})\otimes\{\emptyset,\mathbb{A}\})\to(\mathcal{P}(\mathbb{A}),\mathcal{E}(\mathcal{P}(\mathbb{A})))\) such that \[\begin{align} \psi({B}\cap\{Y \in A\}) = \int_{\mathbb{X}\times\mathbb{A}} \mathbb{1}_{{B}}(y,a) \varpi^\psi_{y,a}(A) \psi(\mathop{\mathrm{d \!}}y \mathop{\mathrm{d \!}}a),\quad B\in\mathcal{B}(\mathbb{X})\otimes\{\emptyset,\mathbb{A}\},\; A\in\mathcal{B}(\mathbb{A}), \end{align}\] i.e., \(\varpi^\psi\) is a regular conditional distribution of \(Y\) given \(\mathcal{B}(\mathbb{X})\otimes\{\emptyset,\mathbb{A}\}\). By definition, for any \(A\in\mathcal{B}(\mathbb{A})\), \(\varpi^\psi_{y,a}(A)\) is constant in \(a\in\mathbb{A}\). Therefore, we let \(\overline{\pi}^\psi\) be \(\varpi^\psi\) restrained on \(\mathbb{X}\) and call \(\overline{\pi}^\psi\) the action kernel induced by \(\psi\). It follows that \(\overline{\pi}^\psi\) is \(\mathcal{B}(\mathbb{X})\)-\(\mathcal{E}(\mathcal{P}(\mathbb{A}))\) measurable, i.e. \(\overline{\pi}^{\psi}\in\overline{\Pi}\). Furthermore, we have \[\begin{align} \label{eq:IndInthpsi} \psi(B\times A) = \int_{\mathbb{X}} \mathbb{1}_{B}(y) \int_{\mathbb{A}} \mathbb{1}_{A}(a)\overline{\pi}^\psi_{y}(\mathop{\mathrm{d \!}}a) \,\xi^\psi(\mathop{\mathrm{d \!}}y), \quad B\in\mathcal{B}(\mathbb{X}),\;A\in\mathcal{B}(\mathbb{A}). \end{align}\tag{45}\] Below, we present a result asserting that 45 uniquely characterizes the induced action kernel. The proof of this statement can be found in 10.1.

Lemma 2. If 45 holds with \(\overline{\pi}^\psi\) replaced by \(\hat{\pi}\in\overline{\Pi}\), then \(\hat{\pi}(x)=\overline{\pi}^\psi(x)\) for \(\xi^{\psi}\)-almost every \(x\).

Subsequently, we denote \(\overline{\mathfrak{p}}^\Psi=(\overline{\pi}^{\psi_1},\dots,\overline{\pi}^{\psi_{T-1}})\) and refer to this as the policy induced by \(\Psi\). Additionally, recall that \(\Xi^\Psi = (\xi^{\psi_1}, \dots, \xi^{\psi_{T-1}})\). Below we establish an equivalence between the mean field flow, \(\Psi\), and the induced policy, \(\overline{\mathfrak{p}}^\Psi\), using the transition operators defined in 3.3. Furthermore, this result confirms that, regardless of the version of the induced policy, it consistently reproduces the mean field flow from which it is derived. The proof is deferred to 10.2.

Lemma 3. Let \(\Psi\) be a mean field flow. Then, for any \(t=1,\dots,T\), \(h\in B_b(\mathbb{X})\), and any version of \(\overline{\mathfrak{p}}^\Psi\), we have \(\overline{\mathbb{T}}^{\overline{\mathfrak{p}}^\Psi}_{t,\Xi^\Psi}h = \int_{\mathbb{X}} h(y)\xi^{\psi_t}(\mathop{\mathrm{d \!}}y)\), or equivalently, \(\overline{\mathbb{Q}}^{\overline{\mathfrak{p}}^\Psi}_{t,\Xi^\Psi}=\xi^{\psi_t}\), where \(\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\) is defined below 29 .

3.6 An \(N\)-player scenario and its mean field counterpart↩︎

For remainder of the paper, we will fix the collection of policies of all players, \(\widetilde{\boldsymbol{\mathfrak{P}}}\in\widetilde{\Pi}^{N\times(T-1)}\), in the \(N\)pG, dubbed the \(N\)-player scenario.

To construct the corresponding scenario in MFG, we first introduce the auxiliary terms \(\overline{\Xi}=(\overline{\xi}_1,\dots,\overline{\xi}_{T-1})\in\mathcal{P}(\mathbb{X})^{T-1}\) and \(\overline{\xi}_T\in\mathcal{P}(\mathbb{X})\) by defining \[\begin{align} \label{eq:Defxibar} \overline{\xi}_1 := \mathring{\xi},\quad\text{and}\quad \overline{\xi}_{t} := \frac{1}{N}\sum_{n=1}^N\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}^n}_{t,\overline{\Xi}},\; t=2,\dots,T, \end{align}\tag{46}\] where we recall that \(\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\) is introduced below 29 , and we note that \(\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}^n}_{t,\overline{\Xi}}\) depends on \(\overline{\Xi}\) only through \((\overline{\xi}_1,\cdots,\overline{\xi}_{t-1})\). We then specify the mean field scenario corresponding to \(\widetilde{\boldsymbol{\mathfrak{P}}}\) as a \(\overline{\Psi}=(\overline{\psi}_1,\dots,\overline{\psi}_{T-1})\in\mathcal{P}(\mathbb{X}\times\mathbb{A})^{T-1}\) satisfying9 \[\begin{align} \label{eq:DefPsiBar} \overline{\psi}_{t}(B\times A) = \frac{1}{N}\sum_{n=1}^N\int_{\mathbb{X}} \int_{\mathbb{A}} \mathbb{1}_B(y)\mathbb{1}_A(a) \tilde{\mathfrak{p}}^n_{t,y,\overline{\xi}_{t}}(\mathop{\mathrm{d \!}}a) \;\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}^n}_{t,\overline{\Xi}} (\mathop{\mathrm{d \!}}y),\quad B\in\mathcal{B}(\mathbb{X}),\,A\in\mathcal{B}(\mathbb{A}). \end{align}\tag{47}\]

The lemma below shows that \(\overline{\Psi}\) is a bona fide mean field flow. The proof is deferred to 10.3.

Lemma 4. With the notations above, we have \(\xi^{\overline{\psi}_t}=\overline{\xi}_t\) for \(t=1,\dots,T-1\), and \[\begin{align} \overline{\xi}_T(B)=\int_{\mathbb{X}\times\mathbb{A}} P_{T-1,x,\xi^{\overline{\psi}_{T-1}},a}(B)\overline{\psi}_{T-1}(\mathop{\mathrm{d \!}}x \mathop{\mathrm{d \!}}a), \quad B\in\mathcal{B}(\mathbb{X}). \end{align}\] Moreover, \(\overline{\Psi}\) is a mean field flow.

3.7 Technical assumptions↩︎

In this section, we consolidate the technical assumptions that are utilized in the forthcoming study. We note that not all of these assumptions are applied concurrently. Instead, they are invoked individually as and when required.

The first assumption regards the compactness of the action domain.

Assumption 2. \(\mathbb{A}\) is compact.

Under 2, we also have that \(\mathcal{P}(\mathbb{A})\) is also compact under weak topology due to Prokhorov’s theorem (see [50]).

The second assumption regards the uniform tightness of \(\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\), where we recall \(\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\) is defined below 29 .

Assumption 3. There is a sequence of increasing compact subsets of \(\mathbb{X}\), denoted by \(\mathfrak{K}=(K_i)_{i\in\mathbb{N}}\), such that \((\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}}_{t,\Xi})_{(t,\Xi,\tilde{\mathfrak{p}})\in\{1,\dots,T\}\times\mathcal{P}(\mathbb{X})^{T-1}\times\widetilde{\Pi}^{T-1}}\) is uniformly tight with respect to \(\mathfrak{K}\) in the following sense \[\begin{align} \overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}}_{t,\Xi}(K_i^c) \le i^{-1},\quad (t,\Xi,\tilde{\mathfrak{p}})\in\{1,\dots,T\}\times\mathcal{P}(\mathbb{X})^{T-1}\times\widetilde{\Pi}^{T-1},\quad i\in\mathbb{N}. \end{align}\]

Clearly, 3 holds if \(\mathbb{X}\) is compact. When \(\mathbb{X}\) is non-compact, the moment condition may serve as a sufficient condition; see, e.g., [52]. We refer to 8 and 15 for additional discussion related to moment functions.

Below, we outline several technical conditions for \(\mathring{\boldsymbol{\mathbb{G}}}\), \(\mathring{\overline{\mathbb{G}}}\), \(\boldsymbol{G}^{\boldsymbol{\lambda}}_t\), and \(\overline{G}^{\lambda}_{t,\xi}\). In particular, 4 supports the abstract formulations introduced in 3.4. We refer to 8.3 for related examples.

Assumption 4. The following is true for any \(t\in\{0,1,\dots,T-1\}\), \(\boldsymbol{\lambda}\in\mathcal{P}(\mathbb{A})^N\), \(\lambda,\lambda'\in\mathcal{P}(\mathbb{A})\), \(\xi\in\mathcal{P}(\mathbb{X})\), \(\boldsymbol{u},{\boldsymbol{u}}'\in B_b(\mathbb{X}^N)\), and \(v,v'\in B_b(\mathbb{X})\):

  • \(\mathring{\boldsymbol{\mathbb{G}}}\), \(\mathring{\overline{\mathbb{G}}}\), \(\boldsymbol{G}^{\boldsymbol{\lambda}}_t\), and \(\overline{G}^{\lambda}_{t,\xi}\) are non-decreasing (i.e., \(\boldsymbol{G}^{\boldsymbol{\lambda}}_{t} \boldsymbol{u} \ge \boldsymbol{G}^{\boldsymbol{\lambda}}_{t} \boldsymbol{u}'\) if \(\boldsymbol{u}\ge \boldsymbol{u}'\)).

  • There are constants \(c_0,c_1>0\) such that \[\begin{align} \left\|\boldsymbol{G}^{\boldsymbol{\lambda}}_{t} \boldsymbol{u}\right\|_\infty \le c_0 + c_1\|\boldsymbol{u}\|_\infty, \quad \left\|\overline{G}^\lambda_{t,\xi} v\right\|_\infty\le c_0 + c_1\|v\|_\infty. \end{align}\]

  • There is a constant \(\bar{c}>0\) such that, for any \(\boldsymbol{x}\in\mathbb{X}^N\), \[\begin{gather} \left|\mathring{\boldsymbol{G}} \boldsymbol{u} - \mathring{\boldsymbol{G}} \boldsymbol{u}'\right| \le \bar{c}\int_{\mathbb{X}^N} \left|\boldsymbol{u}(\boldsymbol{y})-\boldsymbol{u}'(\boldsymbol{y})\right| \mathring{\xi}^{\otimes N}(\mathop{\mathrm{d \!}}\boldsymbol{y}),\\ \left|\boldsymbol{G}^{\boldsymbol{\lambda}}_t \boldsymbol{u}(\boldsymbol{x}) - \boldsymbol{G}^{\boldsymbol{\lambda}}_t \boldsymbol{u}'(\boldsymbol{x})\right| \le \bar{c}\int_{\mathbb{X}^N} \left|\boldsymbol{u}(\boldsymbol{y})-\boldsymbol{u}'(\boldsymbol{y})\right| \left[\bigotimes_{n=1}^N Q^{\lambda^n}_{t,x^n,\overline{\delta}_{\boldsymbol{x}}}\right](\mathop{\mathrm{d \!}}\boldsymbol{y}), \end{gather}\] and, for any \(x\in\mathbb{X}\), \[\begin{gather} \left|\mathring{\overline{G}} v - \mathring{\overline{G}} v'\right| \le \bar{c}\int_{\mathbb{X}} \left|v(y)-v'(y)\right| \mathring{\xi}(\mathop{\mathrm{d \!}}y),\\ \left|\overline{G}^{\lambda}_{t,\xi} v(x) - \overline{G}^{\lambda}_{t,\xi} v'(x)\right| \le \bar{c}\int_{\mathbb{X}} \left|v(y)-v'(y)\right| Q^{\lambda}_{t,x,\xi}(\mathop{\mathrm{d \!}}y). \end{gather}\]

  • If \(\boldsymbol{u}\) satisfies \(\boldsymbol{u}(\boldsymbol{x})=\boldsymbol{u}(\phi \boldsymbol{x})\) for any \(\phi\) that permutes \(z^2,\dots,z^N\) in \((z^1,z^2,\dots,z^N)\), then for any such permutation \(\phi\) we have \[\begin{align} \mathring{\boldsymbol{G}} u = \mathring{\boldsymbol{G}} u(\phi\,\cdot),\quad \boldsymbol{G}^{\boldsymbol{\lambda}}_{t}\boldsymbol{u}(\boldsymbol{x}) = \boldsymbol{G}^{\phi\boldsymbol{\lambda}}_{t} \boldsymbol{u}(\phi\boldsymbol{x}). \end{align}\]

  • If there is \(v\) such that \(u(\boldsymbol{x}) = v(x^1)\) for any \(\boldsymbol{x}\in\mathbb{X}^N\), then \[\begin{align} \mathring{\boldsymbol{G}} \boldsymbol{u} = \mathring{\overline{G}} v,\quad \boldsymbol{G}^{\boldsymbol{\lambda}}_{t}\boldsymbol{u}(\boldsymbol{x}) = \overline{G}^{\lambda^1}_{t,\overline{\delta}_{\boldsymbol{x}}} v(x^1). \end{align}\]

  • For any \(\gamma\in(0,1)\), we have \[\begin{align} \overline{G}^{\gamma \lambda + (1-\gamma) \lambda'}_{t,\xi}v \le \gamma \overline{G}^{ \lambda}_{t,\xi}v + (1-\gamma) \overline{G}^{ \lambda'}_{t,\xi}v. \end{align}\]

4 (i) and (ii) are common for operators involved in performance evaluation. In particular, 4 (ii) imposes boundedness for simplicity. 4 (iii) - (vi) are pivotal for our approximation of \(N\)pGs using MFGs. 4 (iii) allows us to control the propagation of errors in the value functions when approximating \(N\)pGs; we refer to 17 for further discussion. The symmetries in 4 (iv) and (v) facilitate mean field approximations, although one might consider using approximate equality to relax these conditions. 4 (vi) is designed to encourage randomized action. Reversing this inequality could cause our constructed mean field approximation, as done in 2.1 or 3.6, to miss the \(N\)pE. For more discussion, please refer to [rmk:MFPathExploitability] and the example in 8.2.

In addition, we would like to highlight the roles of 4 (iii) and (vi) in the existence of MFEs. 4 (iii) is crucial for proving the existence via the Kakutani–Fan–Glicksberg fixed point theorem (e.g., [50]). Together with the continuity conditions to introduced later, it facilitates the construction of a set-valued map with a closed graph, a requirement of the theorem; see 12 and thereafter. 4 (vi) is generally necessary for the existence. For an example supporting this claim, we again refer to 8.2. Lastly, we observe that 4 (vi) permits costs that extend beyond the usual functional form \(C(x,\xi,a)\), provided they also satisfy other conditions listed in this section. For example, with predefined \(g_1,g_2:\mathbb{X}\times\mathcal{P}(\mathbb{X})\times\mathbb{A}\to\mathbb{R}\), we can incorporate a cost component of the form \[\begin{align} \label{eq:ExmpExtendedCost} \lambda \mapsto \max_{g\in\{g_1,g_2\}} \int_{\mathbb{A}} g(x,\xi,a)\lambda(\mathop{\mathrm{d \!}}a), \end{align}\tag{48}\] as opposed to \(\lambda\mapsto\int_\mathbb{A}C(x,\xi,a)\lambda(\mathop{\mathrm{d \!}}a)\) in the usual setting.

Following are two sets of assumptions concerning the continuities of the transition kernels and the score operators. These two sets of assumptions serve different results and can be considered independently. However, it is ideal when both sets of conditions are met.

The first set of assumptions pertains to the mean field approximation, as presented in 5. Recall that the total-variation norm of a finite measure \(m\) on \(\mathcal{B}(\mathbb{X})\) is defined as \[\begin{align} \|m\|_{TV} := \sup_{h\in B_b(\mathbb{X}),\|h\|_\infty\le1}\frac{1}{2}\int_{\mathbb{X}}h(y)m(\mathop{\mathrm{d \!}}y). \end{align}\] Accordingly, the total-variation distance between two finite measures is defined as \[\begin{align} d_{TV}(m,m') := \|m-m'\|_{TV} = \sup_{h\in B_b(\mathbb{X}),\|h\|_\infty\le 1}\frac{1}{2}\left|\int_{\mathbb{X}}h(y)m(\mathop{\mathrm{d \!}}y)-\int_{\mathbb{X}}h(y)m'(\mathop{\mathrm{d \!}}y)\right|. \end{align}\]

Assumption 5. There is a subadditive modulus of continuity \(\eta\) such that \[\begin{align} \big\|P_t(x,\xi,a) - P_t(x,\xi',a')\big\|_{TV} \le \eta_t\big(d_{\mathbb{A}}(a,a')\big) + \eta_t\big(\big\|\xi-\xi'\big\|_{BL}\big) \end{align}\] for any \(t=1,\dots,T-1\), \(x\in\mathbb{X}\), \(\xi,\xi'\in\mathcal{P}(\mathbb{X})\) and \(a,a'\in\mathbb{A}\).

Assumption 6. The following is true for any \(t=1,\dots,T-1\), \(N\in\mathbb{N}\), \(\boldsymbol{\lambda},\boldsymbol{\lambda}'\in\mathcal{P}(\mathbb{A})^N\), \(\boldsymbol{u}\in B_b(\mathbb{X}^N)\), \(\lambda,\lambda'\in\mathcal{P}(\mathbb{A})\), \(\xi,\xi'\in\mathcal{P}(\mathbb{X})\), and \(v\in B_b(\mathbb{X})\):10

  • There are constants \(c_0,c_1>0\) and a subadditive modulus of continuity \(\zeta\) such that \[\begin{align} \left\|\boldsymbol{G}^{\boldsymbol{\lambda}}_t \boldsymbol{u} - \boldsymbol{G}^{\boldsymbol{\lambda}'}_t \boldsymbol{u}\right\|_\infty \le \left(c_0 + c_1\|\boldsymbol{u}\|_\infty\right)\;\zeta(\|\lambda^1-{\lambda'}^1\|_{BL}). \end{align}\]

  • There are constants \(c_0,c_1>0\) and a subadditive modulus of continuity \(\zeta\) such that \[\begin{align} \left\|\overline{G}^{\lambda}_{t,\xi}v - \overline{G}^{\lambda'}_{t,\xi'}v\right\|_\infty \le \left(c_0 + c_1\|v\|_\infty\right)\;\big(\zeta(\|\xi-\xi'\|_{BL}) + \zeta(\|\lambda-\lambda'\|_{BL}) \big). \end{align}\]

Recall the definitions of \(\widetilde{\Pi}_\vartheta\) and a symmetrically continuous policy from below 30 . We say \(\widetilde{\boldsymbol{\mathfrak{P}}}\) is \(\vartheta\)-symmetrically continuous if it belongs to \(\widetilde{\Pi}^{N\times(T-1)}_\vartheta\). The analogous definition also applies to \(V\in B_b(\mathbb{X}\times\mathcal{P}(\mathbb{X}))\).

Assumption 7. Let \(\vartheta,\iota\) be subaddtive modulus of continuity. The \(N\)-player scenario \(\widetilde{\boldsymbol{\mathfrak{P}}}\) is \(\vartheta\)-symmetrically continuous. All players are subject to the same terminal cost \(\boldsymbol{U}\in B_b(\mathbb{X}^N)\). Moreover, there is an \(\iota\)-symmetrically continuous \(V\in B_b(\mathbb{X}\times\mathcal{P}(\mathbb{X}))\) such that \(\boldsymbol{U}(\boldsymbol{x})=V(x^1,\overline{\delta}_{\boldsymbol{x}})\) for all \(\boldsymbol{x}\in\mathbb{X}^N\).

It’s important to note that 5 does not enforce joint weak continuity. Instead, it imposes strong continuity with respect to \((\xi,a)\). Strong continuities are frequently assumed in games where approximations are concerned (cf. [53]).11 We acknowledge that alternative sets of continuity assumptions may be applicable. For instance, in [26], which parallels our 2 or 4, 5 is imposed, but solely with respect to \(\xi\). However, they also assume joint weak continuity of \(P_t\) and the continuity of the representative player’s policy with respect to the state. Determining which continuity assumptions are most appropriate, or whether it is feasible to relax these conditions, may necessitate additional modeling details that are beyond the scope of this study.

6 plays a role similar to 5. In certain settings, 6 can be derived as a result of 5. For a concrete example, we refer to 8.3. In order to emphasize the importance of \(\vartheta\) in 7, we refer to 8.1 for an example illustrating how an overly rough \(\vartheta\) can lead to a vacuous approximation.

Lastly, we repeat that the symmetrically continuous policies in 7 is reasonable in some realistic considerations. For example, when players have limited accuracy or certainty in perceiving the empirical population distribution, they might become reluctant to drastically alter their behavior in response to changes in the population. This concern could be represented by the aforementioned symmetric continuity.

The following constitutes the second set of assumptions. These assumptions are pivotal when establishing the existence of a MFE through Kakutani–Fan–Glicksberg fixed point theorem (e.g., [50]). 8, as adopted from [26], implies 3; see 15. Clearly, this assumption holds if \(\mathbb{X}\) is compact.12 In addition, we refer to 8.4 for another example with \(\mathbb{X}=\mathbb{R}\). 8 is crucial in the construction of a set-valued mapping that maps into its own domain, and this mapping will be ultimately used to construct a MFE. We note that 8 is also used in [27] for the existence of a MFE. 9 and 10 impose joint continuity on the model’s components and are crucial in ensuring that the aforementioned mapping has a closed graph, a condition required by the fixed point theorem. In 8.3, we provide an example where 9 under a suitable setting implies 10.

Assumption 8. There are an increasing sequence of compact subsets of \(\mathbb{X}\), denoted by \(\check\mathfrak{K}=(\check K_i)_{i\in\mathbb{N}}\), a constant \(\check c \ge 1\), and a continuous \(\sigma:\mathbb{X}\to\mathbb{R}\) such that \(\inf_{x\in \check K_i^c}\sigma(x) \ge i\) for \(i\in\mathbb{N}\), and \[\begin{gather} \int_{\mathbb{X}}\sigma(y)\mathring{\xi}(\mathop{\mathrm{d \!}}y) \le \check c,\label{eq:MomentCondInit}\\ \int_{\mathbb{X}}\sigma(y)P_{t,x,\xi,a}(\mathop{\mathrm{d \!}}y) \le \check c\sigma(x),\quad (t,x,\xi,a)\in\{1,\dots,T-1\}\times\mathbb{X}\times\mathcal{P}(\mathbb{X})\times\mathbb{A}.\label{eq:MomentCond} \end{gather}\] {#eq: sublabel=eq:eq:MomentCondInit,eq:eq:MomentCond}

Assumption 9. \((x,\xi,a)\mapsto P_{t,x,\xi,a}\) is weakly continuous in \(\mathbb{X}\times\mathcal{P}(\mathbb{X})\times\mathbb{A}\).

Assumption 10. \((x,\xi,\lambda)\mapsto \overline{G}^{\lambda}_{t,\xi}v(x)\) is continuous in \(\mathbb{X}\times\mathcal{P}(\mathbb{X})\times\mathcal{P}(\mathbb{A})\) for any \(t=1,\dots,T-1\) and \(v\in C_b(\mathbb{X})\).

3.8 Empirical measures under independent sampling↩︎

In this section, we present a technical lemma, 5, concerning the convergence of the empirical measure derived from independent sampling. It’s important to note that 5 does not assume identical distribution. 5 is inherently required by one of our main results, 3 (see also 1), which pertains to capturing approximate \(N\)pEs with approximate MFEs. This necessity arises from our setup where, in the \(N\)pG, players are allowed to use different policies.

Let \(\mathbb{Y}\) be a complete separable metric space. For \(j\in\mathbb{N}\) and \(A\subseteq\mathbb{Y}\) we define \[\begin{align} \mathfrak{N}_j(A) := \min\left\{n\in\mathbb{N}: C_{BL}(A)\subseteq\bigcup_{i=1}^n B_{j^{-1}}(h_i) \text{ for some } h_1,\dots,h_n\in C_{BL}(A)\right\}, \end{align}\] i.e., \(\mathfrak{N}_j(A)\) is the smallest number of \(j^{-1}\)-open balls needed to cover \(C_{BL}(A)\). If \(A\) is compact, by Ascoli-Arzelá theorem (cf. [54]) and the fact that compact metric space is totally bounded (cf. [54]), \(\mathfrak{N}_j(A)\) is finite for \(j\in\mathbb{N}\).

Let \(N\in\mathbb{N}\) be fixed. Suppose \((Y^n)_{n\in\mathbb{N}}\) are independent \(\mathbb{Y}\)-valued random variables and let \(\upsilon^n\) be the law of \(Y^n\). We denote \(\boldsymbol{Y}:=(Y^n)_{n=1}^N\). Consider the empirical measure \(\overline{\delta}_{\boldsymbol{Y}}\) and the corresponding intensity measure \(\overline{\upsilon} := \frac{1}{N}\sum_{n=1}^N\upsilon^n\). Note that \(\left\|\overline{\delta}_{\boldsymbol{Y}}-\overline{\upsilon}\right\|_{BL}\) is \(\mathscr{A}\)-\(\mathcal{B}(\mathbb{R})\) measurable because \(\|\cdot\|_{BL}\) is continuous. 5 below provides an upper bound for the expectation of the bounded-Lipschitz distance between the empirical measure under independent sampling and the corresponding intensity measure. The proof of 5, which is deferred to 10.5, primarily involves applying Hoeffding’s inequality following a discretization enabled by the assumed uniform tightness. 5 embodies the concept of a law of large numbers without identical distribution (cf. [55]). The lack of a citation in 5 should not be interpreted as a claim of originality, but rather as an indication that this is a commonly expected result (if not sharper).

Lemma 5. Let \((\upsilon^n)_{n=1}^N\), \(\overline{\upsilon}\), and \(\boldsymbol{Y}\) be as introduced earlier. Suppose there is an increasing sequence of compact subsets of \(\mathbb{Y}\), denoted by \(\mathfrak{A}:=(A_i)_{i\in\mathbb{N}}\), such that \(\sup_{n}\upsilon_n(A_i^c)\le i^{-1}\) for all \(i\in\mathbb{N}\), i.e., \((\upsilon^n)_{n\in\mathbb{N}}\) is uniformly tight with respect to \(\mathfrak{A}\). Then, \[\begin{align} \label{eq:EmpMeasConc} \mathbb{E}\left(\left\|\overline{\delta}_{\boldsymbol{Y}}-\overline{\upsilon}\right\|_{BL}\right) \le \inf_{i,j\in\mathbb{N}} \left\{ \frac{1}{2}j^{-1} + i^{-1} + \frac{\sqrt{\pi}\;\mathfrak{N}_{j}(A_i)}{\sqrt{2 N}} \right\} =: \mathfrak{r}_{\mathfrak{A}}(N). \end{align}\qquad{(1)}\] Moreover, we have \(\lim_{N\to\infty}\mathfrak{r}_{\mathfrak{A}}(N)=0\) as \(N\to\infty\).

The convergence rate established above leaves significant room for improvement. It is important to note that the forthcoming main results, while involving \(\mathfrak{r}_\mathfrak{A}\), are independent of the specific procedure used to derive \(\mathfrak{r}_\mathfrak{A}\). Therefore, any improvement in the convergence rate can directly enhance these results simply by substituting with the improved rate. A better rate will be pursuit elsewhere.

Results akin to 5 can be found in [56][59] and the reference therein. These results provides sharper estimations but requires identical distribution and more specific spatial structure.

3.9 Error terms↩︎

In this section, we will introduce several error terms, primarily for the sake of notational convenience.

Recall that \(\overline{\delta}_{\boldsymbol{x}}=\frac{1}{N}\sum_{n=1}^N \delta_{x^n}\) for \(\boldsymbol{x}=(x^1,\dots,x^N)\in\mathbb{X}^N\). Let \(\overline{\xi}_t\) and \(\overline{\psi}_t\) be as introduced in 3.6. we define \[\begin{align} \label{eq:DefEmpErr} {\boldsymbol{e}}_t({\boldsymbol{x}}):=\left\|\overline{\delta}_{\boldsymbol{x}} - \overline{\xi}_t\right\|_{BL},\quad t=1,\dots,T. \end{align}\tag{49}\] We additionally define \[\begin{align} \label{eq:DefStateActionEmpErr} \breve {\boldsymbol{e}}_t({\boldsymbol{x}}):=\int_{\mathbb{A}^N}\left\|\overline{\delta}_{((x^1,a^1),\dots,(x^N,a^N))}-{\overline{\psi}}_t\right\|_{BL} \left[\bigotimes_{n=1}^N{\tilde{\mathfrak{p}}^n_{t,x^n,\overline{\delta}_{\boldsymbol{x}}}}\right](\mathop{\mathrm{d \!}}{\boldsymbol{a}}), \quad t=1,\dots,T-1. \end{align}\tag{50}\] It follows from definition that \(\boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_t{\boldsymbol{e}}_t\) calculates, in the \(N\)-player scenario \(\widetilde{\boldsymbol{\mathfrak{P}}}\) at time \(t\), the expectation of bounded-Lipschitz distance between the empirical state distribution and \(\overline{\xi}_t\). Similarly, \(\boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_t\breve{\boldsymbol{e}}_t\) calculates the expectation of bounded-Lipschitz distance between the empirical state-action distribution and \(\overline{\psi}_t\).

Below we introduce a few more constants \[\begin{gather} \mathfrak{e}_t\,, \quad \breve\mathfrak{e}_t\,, \quad \mathfrak{e}^0_t\,, \quad t=1,\dots, T,\\ \underline \mathfrak{E}\,, \quad \mathfrak{E}\,, \quad \mathfrak{E}^0\,,\quad \mathfrak{E}^\diamond\,. \end{gather}\] We defer the detailed definitions of these constants to 10.4. These constants depend on the total number of players \(N\), \(\mathfrak{K}\) in 3, \(c_0,c_1,\bar{c}\) in 4, \(\eta\) in 5, \(\zeta\) in 6, \(\vartheta,\iota,\|V\|_\infty\) in 7, and the convergence rate established in 5. It is important to note that values of these constants do not hinge on the particular choice of \(\widetilde{\boldsymbol{\mathfrak{P}}}\).

The forthcoming lemma offers a qualitative understanding of the aforementioned constants. Unlike the other parts of this paper, in this lemma, we explicitly incorporate the dependence on the total number of players, \(N\), into our notations and allow \(N\) to approach infinity. We refer to 10.4 for the proof.

Lemma 6. Suppose 2 and consider the aforementioned notations. Then, for any \(t=1,\dots, T\), we have \(\lim_{N\to\infty}{\mathfrak{e}_t}(N)=0\), \(\lim_{N\to\infty}{\breve\mathfrak{e}_t}(N)=0\), and \(\lim_{N\to\infty}\mathfrak{e}^0_t(N)=0\). Moreover, we have \(\lim_{N\to\infty}\underline{\mathfrak{E}}(N)=0\), \(\lim_{N\to\infty}{\mathfrak{E}}(N)=0\), \(\lim_{N\to\infty}{\mathfrak{E}^0}(N)=0\), and \(\lim_{N\to\infty}{\mathfrak{E}^\diamond}(N)=0\).

6 implicitly requires that the regularities of the transition and score operators, as imposed in 3 - 7, are preserved as the size of the \(N\)pG increases.

4 Exploitabilities in \(N\)pGs and MFGs↩︎

This section is devoted to the introduction of various notions of exploitabilities and the examination of their interrelationships. Exploitabilities for \(N\)pG are discussed in 4.1, while exploitabilities for MFGs are the focus of 4.2.

4.1 Exploitabilities in \(N\)pGs↩︎

In light of the setting in 3.6 and 7, from the player-\(1\)’s perspective, a natural definition of exploitability emerges as follows. For any \(N\)-player scenario \({\boldsymbol{\mathfrak{P}}}=(\mathfrak{p}^1,\dots,\mathfrak{p}^N)\in\Pi^{N\times(T-1)}\) and terminal cost function \({\boldsymbol{U}}\in B_b(\mathbb{X}^N)\). we define \[\begin{align} \label{eq:DefSymContEndExploitability} \boldsymbol{\mathcal{R}}_{\vartheta}({\boldsymbol{\mathfrak{P}}};{\boldsymbol{U}}):= \boldsymbol{\mathbb{S}}^{{\boldsymbol{\mathfrak{P}}}}_T {\boldsymbol{U}} - \inf_{\tilde{\mathfrak{p}}\in\widetilde{\Pi}_\vartheta}\boldsymbol{\mathbb{S}}^{(\tilde{\mathfrak{p}},\mathfrak{p}^2,\dots,\mathfrak{p}^N)}_{t} {\boldsymbol{U}}. \end{align}\tag{51}\] We note that, while \(\boldsymbol{\mathcal{R}}_{\vartheta}\) is defined for generic \({\boldsymbol{\mathfrak{P}}}\) and \({\boldsymbol{U}}\), our approximation result developed later is limited to \({\boldsymbol{\mathfrak{P}}}\) and \({\boldsymbol{U}}\) that satisfies 7. Let \(\varphi^{1}\) be the identity permeation. For \(n=2,\dots,N\), let \(\varphi^{n}\) be a permutation such that \(\varphi^{n}(z^1,\dots,z^N) = (z^n,z^1,\dots,z^{n-1},z^{n+1},\dots,z^N)\). Under 4 (iv), definition 51 also applies to player-\(n\) via \[\begin{align} \boldsymbol{\mathcal{R}}^n_{\vartheta}({\boldsymbol{\mathfrak{P}}};{\boldsymbol{U}}) := \boldsymbol{\mathcal{R}}_{\vartheta}(\varphi^{n}{\boldsymbol{\mathfrak{P}}};{\boldsymbol{U}}), \quad n=1,\dots,N. \end{align}\]

Although \({\boldsymbol{\mathcal{R}}_{\vartheta}}\) is inherently suited to the scenario described in 3.6, the associated constrained optimization problem presents a significant challenge. To circumvent this, we employ an approximation approach. To aid in this process, we introduce the following auxiliary definitions of exploitability. Suppose 2, 4 (i) (ii), and 6 for the validity of the upcoming definitions (see also 14). We define \[\begin{gather} \boldsymbol{\mathcal{R}}({\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}) := \boldsymbol{\mathbb{S}}^{{\boldsymbol{\mathfrak{P}}}}_T {\boldsymbol{U}} - \boldsymbol{\mathbb{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{T} {\boldsymbol{U}},\label{eq:DefEndExploitability} \end{gather}\tag{52}\] and \[\begin{gather} \boldsymbol{\mathfrak{R}}({\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}) := \sum_{t=1}^{T-1} \bar{c}^t \boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_t \left(\boldsymbol{S}^{{\boldsymbol{\mathfrak{P}}}_t}_t \circ \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t+1,T} {\boldsymbol{U}} - \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}}\right),\label{eq:DefStepExploitability} \end{gather}\tag{53}\] where we recall \(\bar{c}\) from 4 (iii). Similarly as before, with 4 (iv), we further define \[\begin{align} \label{eq:DefPermExploitabilities} \boldsymbol{\mathcal{R}}^n({\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}) := \boldsymbol{\mathcal{R}}(\varphi^{n}{\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}\circ\varphi^n),\quad \boldsymbol{\mathfrak{R}}^n({\boldsymbol{\mathfrak{P}}};{\boldsymbol{U}}) := \boldsymbol{\mathfrak{R}}(\varphi^{n}{\boldsymbol{\mathfrak{P}}};{\boldsymbol{U}}\circ\varphi^n), \quad n=1,\dots,N. \end{align}\tag{54}\]

Following the discussion in 2.1, we refer to \(\boldsymbol{\mathcal{R}}\) and \(\boldsymbol{\mathfrak{R}}\) as end exploitability and total (stepwise) exploitability, respectively. Accordingly, we refer to \({\boldsymbol{\mathcal{R}}_{\vartheta}}\) as the end exploitability constrained to \(\vartheta\)-symmetrically continuous policy, or simply, the constrained end exploitability. By definition (recall 3.4), both end exploitability and total exploitability are non-negative. Technically, the constrained end exploitability could be negative if the individual player employs a policy that does not meet the continuity constrain.

In this remark, we discuss a modeling issue of the end exploitability within a risk-averse framework, which ultimately prompts the exploration of total exploitability. For certain risk-averse performance criterion, such as 85 with \(J=1\), \(w_1=1\), and \(\kappa_1=\kappa\in(0,1)\), when combined with specific environments, achieving \(0\) end exploitability may not necessarily optimize sample paths that yield superior overall outcomes. This is in fact an immediate consequence of the robust representation of average value at risk (see, e.g., [47]). This phenomenon of under-optimization appears somewhat problematic, as it is typically anticipated that a player would continue to optimize, regardless of a strong initial start. To mitigate this issue, one approach is to modify the criterion by employing a blend with expectation to prevent the negligence of sample paths. Alternatively, considering total exploitability could be a solution, as it demands optimal actions across almost every sample paths to achieve a value of 0.

Below we summarize the relations between different notions of \(N\)-player exploitabilities introduced above, the proof of which is deferred to 7.4 .

Suppose 2, 4 (i) (ii), and 6 for \(\boldsymbol{\mathcal{R}},\boldsymbol{\mathfrak{R}},\boldsymbol{\mathcal{R}}_\vartheta\) to be well-defined. The following is true:

  • If 4 (iii) holds, then \[\begin{align} \label{eq:EndvsStep} \boldsymbol{\mathcal{R}}({\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}})\le \mathfrak{R}({\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}), \quad \boldsymbol{\mathfrak{P}}\in\Pi^{N\times(T-1)},\, \boldsymbol{U}\in B_b(\mathbb{X}^N). \end{align}\tag{55}\]

  • Suppose 4 (i). Additionally, assume the existence of a \(\underline c>0\) such that, for any \({\boldsymbol{\lambda}}\in\mathcal{P}(\mathbb{A})^N\) and \({\boldsymbol{u}}\ge{\boldsymbol{u}}'\), \[\begin{gather} \mathring{\boldsymbol{G}}{\boldsymbol{u}} - \mathring{\boldsymbol{G}}{\boldsymbol{u}}' \ge \underline c\int_{\mathbb{X}^N} \left({\boldsymbol{u}}({\boldsymbol{y}})-{\boldsymbol{u}}'({\boldsymbol{y}})\right) \mathring{\xi}^{\otimes N}(\mathop{\mathrm{d \!}}{\boldsymbol{y}}),\\ \boldsymbol{G}^{{\boldsymbol{\lambda}}}_t{\boldsymbol{u}}({\boldsymbol{x}}) - \boldsymbol{G}^{{\boldsymbol{\lambda}}}_t{\boldsymbol{u}}'({\boldsymbol{x}}) \ge \underline c\int_{\mathbb{X}^N} \left({\boldsymbol{u}}({\boldsymbol{y}})-{\boldsymbol{u}}'({\boldsymbol{y}})\right) \left[\bigotimes_{n=1}^N Q^{\lambda^n}_{t,x^n,\overline{\delta}_{\boldsymbol{x}}}\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}}), \end{gather}\] then \[\begin{align} \label{eq:EndvsStep2} \boldsymbol{\mathcal{R}}({\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}) \ge \left({\underline c}\middle/{\bar{c}}\right)^{T-1}\boldsymbol{\mathfrak{R}}({\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}). \end{align}\tag{56}\]

  • Suppose 3, 4 (iii) (v), 5, 6 and 7. With \(\underline\mathfrak{E}\) introduced in 3.9, we have \[\begin{align} \left|{\boldsymbol{\mathcal{R}}_{\vartheta}}(\widetilde{\boldsymbol{\mathfrak{P}}};{\boldsymbol{U}}) - \boldsymbol{\mathcal{R}}(\widetilde{\boldsymbol{\mathfrak{P}}};{\boldsymbol{U}})\right| \le \underline\mathfrak{E}. \end{align}\]

For an example where \(\mathring{\boldsymbol{G}}\) and \(\boldsymbol{G}^{\boldsymbol{\lambda}}_t\) satisfy the assumptions in [prop:EstEndExploitabilityCont] (b), we refer to the expressions in 85 and 86 with \(w_1\in(0,1]\) and \(\kappa_1=1\), i.e., we blend average value at risk’s with the expectation. This is related to the spectral risk measure [60]. In this case, we can set \(\underline c=w_1\). For another example, we refer to the standard risk-neutral setting as used in 2.1, where the expectation of total cost is used as performance criteria (i.e., \(w_1=1\) and \(\kappa_1=1\) in 85 and 86 ). Consequently, we have \(\underline c=\bar{c}=1\), and thus end exploitability and total exploitability coincide, \[\begin{align} \label{eq:IndEndStepExploitability} \boldsymbol{\mathcal{R}}({\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}) = \boldsymbol{\mathfrak{R}}({\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}). \end{align}\tag{57}\]

4.2 Exploitabilities in mean field games↩︎

We first recall the notations introduced in 3.5. In particular, \(\Psi\in\mathcal{P}(\mathbb{X}\times\mathbb{A})^{T-1}\) and \(\Xi^{\Psi}\in\mathcal{P}(\mathbb{X})^{T-1}\) is the corresponding vector of state marginals. For \(V\in B_b(\mathbb{X}\times\mathcal{P}(\mathbb{X}))\), we write13 \[\begin{align} \label{eq:DefVM} V_\Psi(x):=V\left(x,\int_{\mathbb{X}\times\mathbb{A}} P_{t,y,\xi_t,a}(\cdot) \psi_{T-1}(\mathop{\mathrm{d \!}}y\mathop{\mathrm{d \!}}a)\right). \end{align}\tag{58}\] We define the mean field end exploitability as \[\begin{align} \label{eq:DefMeanEndExploitability} \overline{\mathcal{R}}(\Psi;V) := \overline{\mathbb{S}}^{{\overline{\mathfrak{p}}^\Psi}}_{T,\Xi^{\Psi}} V_{\Psi} - \overline{\mathbb{S}}^{*}_{T,{\Xi^{\Psi}}} V_{\Psi}. \end{align}\tag{59}\] The mean field total (stepwise) exploitability is \[\begin{align} \label{eq:DefMeanStepExploitability} \overline{\mathfrak{R}}(\Psi;V) &:= \sum_{t=1}^{T-1} \bar{c}^t \overline{\mathbb{T}}^{{\overline{\mathfrak{p}}^\Psi}}_{t,{\Xi^{\Psi}}} \left( \overline{S}^{{\overline{\pi}^{\psi_t}}}_{t,{\xi^{\psi_t}}} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\Xi^{\Psi}}} V_{\Psi} - \overline{\mathcal{S}}^{*}_{t,T,{\overline{\Xi}^{\Psi}}} V_{\Psi} \right)\nonumber\\ &= \sum_{t=1}^{T-1} \bar{c}^t \int_{\mathbb{X}} \left(\overline{S}^{{\overline{\pi}^{\psi_t}}}_{t,{\xi^{\psi_t}}} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\Xi^{\Psi}}} V_{\Psi}(y) - \overline{\mathcal{S}}^{*}_{t,T,{\overline{\Xi}^{\Psi}}} V_{\Psi}(y)\right) {\xi^{\psi_t}}(\mathop{\mathrm{d \!}}y), \end{align}\tag{60}\] where we have used 3 in the last equality. In view of 2, 4 (i) (ii), and 6 (see also 14), both \(\overline{\mathfrak{R}}(\Psi;V)\) and \(\overline{\mathcal{R}}(\Psi;V)\) are well-defined and non-negative.

The process of defining \(\overline{\mathfrak{R}}\) entails deriving regular conditional distributions from specific joint distributions, which may be cumbersome in certain instances. In 8.5, we present an example illustrating how a more explicit form of \(\overline{G}^{\lambda}_{t,\xi}\) can lead to a simplification of \(\overline{\mathfrak{R}}\).

The lemma below shows \(\overline{\mathcal{R}}(\Psi;V)\) and \(\overline{\mathfrak{R}}(\Psi;V)\) are version-independent in \({\overline{\mathfrak{p}}^\Psi}\), the proof of which is deferred to 7.3.

Lemma 7. Suppose 4 (iii). Consider a mean field system \(\Psi\) and \(v\in B_b(\mathbb{X})\). Let \({\overline{\mathfrak{p}}^\Psi}\) and \({\overline{\mathfrak{p}}^\Psi}'\) be two version of policy induced by \(\Psi\). Then, \[\begin{align} \overline{\mathbb{S}}^{{\overline{\mathfrak{p}}^\Psi}}_{T,{\Xi^{\Psi}}} v = \overline{\mathbb{S}}^{{\overline{\mathfrak{p}}^\Psi}'}_{T,{\Xi^{\Psi}}} v,\quad \int_\mathbb{X}\overline{S}^{{\overline{\pi}^{\psi_t}}}_{t,{\xi^{\psi_t}}} v(y) {\xi^{\psi_t}}(\mathop{\mathrm{d \!}}y) = \int_\mathbb{X}\overline{S}^{{\overline{\pi}^{\psi_t}}'}_{t,{\xi^{\psi_t}}} v(y) {\xi^{\psi_t}}(\mathop{\mathrm{d \!}}y),\,t=1,\dots,T-1. \end{align}\]

Analogous to previous results for \(N\)pGs, we observe a similar relationship between mean field end exploitability and mean field total exploitability. The proof mirrors the argument used to prove [prop:EstEndExploitabilityCont], and is therefore omitted here.

Suppose 2, 4 (i) (ii), and 6 for \(\overline{\mathcal{R}}\) and \(\overline{\mathfrak{R}}\) to be well-defined. The following is true for any mean field flow \(\Psi\in\mathcal{P}(\mathbb{X}\times\mathbb{A})^{T-1}\) and \(V\in B_b(\mathbb{X}\times\mathcal{P}(\mathbb{A}))\):

  • If 4 (iii) holds, then \[\begin{align} \label{eq:MeanEndvsStep} \overline{\mathcal{R}}(\Psi;V) \le \overline{\mathfrak{R}}(\Psi;V). \end{align}\tag{61}\]

  • Suppose 4 (i). Additionally, assume the existence of a \(\underline c>0\) such that, for any \(\lambda\in\mathcal{P}(\mathbb{A})\) and \(v\ge{v'}\), \[\begin{gather} \mathring{\overline{G}} v - \mathring{\overline{G}} {v'} \ge \underline c\int_{\mathbb{X}^N} \left(v(y)-{v'}(y)\right) \mathring{\xi}(\mathop{\mathrm{d \!}}y),\\ \overline{G}^{\lambda}_{t,\xi} v(x) - \overline{G}^{\lambda}_{t,\xi} {v'}(x) \ge \underline c\int_{\mathbb{X}} \left(v(y)-{v'}(y)\right) Q^{\lambda}_{t,x,\xi}(u). \end{gather}\] Then, \[\begin{align} \overline{\mathcal{R}}(\Psi;V) \ge \left({\underline c}\middle/{\bar{c}}\right)^{T-1}\overline{\mathfrak{R}}(\Psi;V). \end{align}\]

Similarly as before, for an example where \(\mathring{\overline{G}}\) and \(\overline{G}^\lambda_{t,\xi}\) satisfy the assumptions in [prop:EstMFEndExploitability] (b), we refer to 87 and 88 with \(w_1\in(0,1]\) and \(\kappa_1=1\). In this case, we may set \(\underline c=w_1\). Additionally, in the standard risk-neutral setting as in 2.1, we have \(\underline c=\bar{c}=1\), and thus \[\begin{align} \label{eq:IndMFEndStepExploitability} \overline{\mathcal{R}}({\Psi}; V) = \overline{\mathfrak{R}}({\Psi}; V). \end{align}\tag{62}\]

5 Approximation capability of mean field games↩︎

In the following theorem, we examine an \(N\)pG where all players employ symmetrically continuous policies. The theorem asserts that, given a sufficiently large \(N\), the empirical state-action flow can be closely approximated by the corresponding mean field flow. Furthermore, the average of stepwise exploitabilities in the \(N\)pG dominates the stepwise exploitability of the corresponding MFG up to a minor error term. The proof of this theorem is deferred to 7.5 and 7.6.

Theorem 3. Let \(\widetilde{\boldsymbol{\mathfrak{P}}}\) and \(\overline{\Psi}\) be as introduced in 3.6. Suppose 2, 3, 4, 5, 6 and 7. Then, with the error terms defined in 3.9, we have\[\begin{align} \label{eq:GoodDescription} \boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{t}\breve {\boldsymbol{e}}_t \le {\breve\mathfrak{e}_t} \end{align}\qquad{(2)}\] and \[\begin{align} \label{eq:NoMiss} \overline{\mathfrak{R}}({\overline{\Psi}};V) \le \frac{1}{N}\sum_{n=1}^N{\boldsymbol{\mathfrak{R}}^n}(\widetilde{\boldsymbol{\mathfrak{P}}};{\boldsymbol{U}}) + {\mathfrak{E}}. \end{align}\qquad{(3)}\]

As seen in the proof of 3, if we have equality in 4 (vi),14 we achieve equality instead of inequality in 76 . As a result, we obtain a stronger version of 3 that \[\begin{align} \label{eq:RNNpRMFR} \left|\overline{\mathfrak{R}}({\overline{\Psi}};V) - \frac{1}{N}\sum_{n=1}^N{\boldsymbol{\mathfrak{R}}^n}(\widetilde{\boldsymbol{\mathfrak{P}}};{\boldsymbol{U}})\right| \le {\mathfrak{E}}. \end{align}\tag{63}\] On the other hand, if we assume concavity in 4 (vi), by using in 76 Jensen’s inequality for concave functions, we eventually yield \[\frac{1}{N}\sum_{n=1}^N{\boldsymbol{\mathfrak{R}}^n}(\widetilde{\boldsymbol{\mathfrak{P}}};{\boldsymbol{U}}) \le \overline{\mathfrak{R}}({\overline{\Psi}};V) + {\mathfrak{E}}.\] This suggests that, in the study of an \(N\)pG that encourages deterministic actions over randomized actions, approximation via \(\overline{\Psi}\), as constructed in 3.6, may result in overlooking approximate equilibria in the \(N\)pG.

We conjecture that 3 (b) remain trues if \(\boldsymbol{\mathfrak{R}}^N\) and \(\overline{\mathfrak{R}}\) are replaced by \(\boldsymbol{\mathcal{R}}^N\) and \(\overline{\mathcal{R}}\); whether additional assumption is needed remains unknown. Alternatively, we can utilize [prop:EstEndExploitabilityCont] and [prop:EstMFEndExploitability] to derive less sharp estimates.

The following theorem asserts that any mean field flow with a small mean field scenario can be used to construct a homogeneous open-loop \(N\)-player scenario, where the exploitability of each player closely approximates the mean field exploitability. We note that 4 (vi) is not required in this context. We refer to 7.7 for the proof.

Theorem 4. Let \(\Psi\) be a mean field flow as defined in 3.5. For \(n=1,\dots,N\), let \({\overline{\mathfrak{p}}^\Psi}^n\) be a version of policy induced by \(\Psi\). Consider the \(N\)-player scenario given by \(\overline{\boldsymbol{\mathfrak{P}}}:=({\overline{\mathfrak{p}}^\Psi}^1,\dots,{\overline{\mathfrak{p}}^\Psi}^N)\). If 2, 3, 4 (i) - (v), 5, 6, and 7 hold, then \[\left|{\boldsymbol{\mathfrak{R}}^n}(\overline{\boldsymbol{\mathfrak{P}}};{\boldsymbol{U}}) - \overline{\mathfrak{R}}(\Psi;V)\right| \le{\mathfrak{E}^0}, \quad n=1,\dots,N,\] and \[\left|\boldsymbol{\mathcal{R}^n}(\overline{\boldsymbol{\mathfrak{P}}};{\boldsymbol{U}}) - \overline{\mathcal{R}}(\Psi;V)\right| \le {\mathfrak{E}^\diamond}, \quad n=1,\dots,N,\] where \(\mathfrak{E}^0\) and \(\mathfrak{E}^\diamond\) are introduced in 3.9.

6 Existence of mean field equilibrium↩︎

Recall that, in our context, a mean field equilibrium is a mean field flow \(\Psi\) with \(\overline{\mathcal{R}}(\Psi;V)=0\). In light of 61 , 5 below reveals the existence of a mean field equilibrium. We refer to 7.8 for the detailed proof.

Theorem 5. Let \(V\in C_b(\mathbb{X}\times\mathcal{P}(\mathbb{X}))\). Suppose 2, 4 (i) (ii) (iii) (vi), 8, 9 and 10. Then, there is a mean field flow \(\Psi^*\) such that \(\overline{\mathfrak{R}}(\Psi^*;V)=0\).

Below we would like to share some insights we gleaned from the proof of 5. These insights pertain to incorporating penalization for randomized action into mean field game theory. The proof of 5 hinges significantly on 4 (vi). We note that 4 (vi) encourages the use of randomized actions and is crucial for both the approximation and the existence of equilibrium in the current mean field game theory framework; see the discussion following 4 (vi). On the other hand, the reluctance to embrace randomized actions is commonly observed in real world, prompting a pertinent discussion. In our proof of the existence of an MFE, 4 (vi) is not utilized until the final step. This observation could illuminate potential pathways for incorporating penalization for randomized actions into mean field game theory. To delve deeper, we subsequently highlight an important intermediate result, [prop:GammaFixedpt], which does not require 4 (vi).

Let \(\mathcal{P}(\mathbb{X}\times\mathcal{P}(\mathbb{A}))\) be the set of probability measures on \(\mathcal{B}(\mathbb{X})\otimes\mathcal{B}(\mathcal{P}(\mathbb{A}))\). To connect \(\mu\in\mathcal{P}(\mathbb{X}\times\mathcal{P}(\mathbb{A}))\) with state-action distributions, which are fundamental in constructing mean field flows, we introduce a transformation that collapses the \(\mathcal{P}(\mathbb{A})\) component of \(\mu\). Specifically, we define \(\psi^\mu\in\mathcal{P}(\mathbb{X}\times\mathbb{A})\) by setting \[\begin{align} \psi^\mu(B\times A) = \int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})} \mathbb{1}_{B}(y) \lambda(A) \mu(\mathop{\mathrm{d \!}}y\mathop{\mathrm{d \!}}\lambda),\quad B\in\mathcal{B}(\mathbb{X}),\, A\in\mathcal{B}(\mathbb{A}). \end{align}\] Heuristically, \(\mu \in \mathcal{P}(\mathbb{X}\times \mathcal{P}(\mathbb{A}))\) represents the state-action-kernel distribution, in contrast to \(\psi \in \mathcal{P}(\mathbb{X}\times \mathbb{A})\), which represents the state-action distribution. Note the \(\mu\) can be viewed as a statistics that summaries the the state distribution of players, and the (conditional) distribution of action kernels (instead of actions) players at a given state employ. The process of defining \(\psi^\mu\) entails averaging the action kernels of players who share the same state, in an expected sense.

In what follows, we let \(\xi^\mu\) be the marginal distribution of \(\mu\) on \(\mathbb{X}\). Additionally, for \(\mathfrak{m}=(\mu_1,\dots,\mu_{T-1})\in\mathcal{P}(\mathbb{X}\times\mathcal{P}(\mathbb{A}))^{T-1}\), we denote \(\Xi^\mathfrak{m}=(\xi^{\mu_1},\dots,\xi^{\mu_{T-1}})\) and \(\Psi^\mathfrak{m}=(\psi^{\mu_1},\dots,\psi^{\mu_{T-1}})\). [prop:GammaFixedpt] reveals the existence of an \(\mathfrak{m}\in\mathcal{P}(\mathbb{X}\times\mathcal{P}(\mathbb{A}))^{T-1}\) that satisfies \(\xi^{\mu_1}=\mathring\xi\), \[\begin{align} \label{eq:HeurIndGammaMarginal} \mu_{t+1}(B) = \int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})}\int_{\mathbb{A}}P_{t,x,\xi^{\mu_t},a}(B)\lambda(\mathop{\mathrm{d \!}}a)\mu_t(\mathop{\mathrm{d \!}}x\mathop{\mathrm{d \!}}\lambda),\quad B\in\mathcal{B}(\mathbb{X}),\, t=1,\dots,{T-2}, \end{align}\tag{64}\] and \[\begin{align} \label{eq:HeurIndGammaOpt} \int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})} \left(\overline{G}^{\lambda}_{t,\xi^{\mu_t}} \circ \overline{\mathcal{S}}^{*}_{t+1,T,\Xi^\mathfrak{m}} V_{\mathfrak{m}}(y) - \overline{\mathcal{S}}^{*}_{t,T,\Xi^\mathfrak{m}} V_{\mathfrak{m}}(y)\right) \mu_t(\mathop{\mathrm{d \!}}y \mathop{\mathrm{d \!}}\lambda) = 0, \end{align}\tag{65}\] where, in light of the definition given in 58 , and with a slight abuse of notation, we denote \(V_\mathfrak{m}:=V_{\Psi^\mathfrak{m}}\). This suggests that an ‘equilibrium’, characterized by a flow of state-action-kernel distributions, can exist even when 4 (vi) is not met. Note that, by 64 , \(\Psi^\mathfrak{m}\) is a genuine mean field flow; see 83 for detailed deduction. Additionally, 65 guarantees that \(\mu_t\) only assign masses to optimal action kernels. However, 65 does not necessarily ensure that \(\Psi^\mathfrak{m}\) exhibits zero exploitability. Indeed, without 4 (vi), averaging (optimal) action kernels could potentially degrade performance. We refer to 84 for the implied optimality when 4 (vi) is assumed.

In conclusion, we posit that integrating penalization for randomized actions into mean field game theory is viable. However, we advise against adopting the state-action mean field flow notation, or its equivalent, the representative agent framework. A more suitable approach involves considering the broader concept, flows of state-action-kernel distributions. For a related example in discrete spaces, we refer to 8.2.

7 Proofs↩︎

In 7.1 and 7.2, we initially present some auxiliary results that lay the groundwork for the proofs of [prop:EstEndExploitabilityCont], 3, and 4. Specifically, we first establish a technical result, [prop:ProcConc], based on 5, which characterizes the concentration of empirical state-action distributions at time \(t\) in the \(N\)pG. Leveraging [prop:ProcConc], we then demonstrate in [prop:EstDiffRMF] that total exploitabilities in the \(N\)pG can be approximated by quantities computed solely by mean field-type operators. [prop:EstEndExploitabilityCont], 3, and 4 are subsequently proved in 7.4 through 7.7. Lastly, 7.8 is devoted to the proof of 5, which is done independently of the preceding results.

7.1 Empirical measures in the \(N\)pG↩︎

Let \(\widetilde{\boldsymbol{\mathfrak{P}}}\) and \(\overline{\Xi}\) be as introduced in 3.6. Consider a \({\boldsymbol{\mathfrak{P}}}\in\Pi^{N\times(T-1)}\) that satisfies \(\mathfrak{p}^n=\tilde{\mathfrak{p}}^n\) for \(n=2,\dots,N\). If 3, 5, and 7 hold, then for any \(t=1,\dots,T\) and subadditive modulus of continuity \(\beta\), we have \({\boldsymbol{\mathbb{T}}}^{{\boldsymbol{\mathfrak{P}}}}_{t} \circ \beta {\boldsymbol{e}}_t \le \beta({\mathfrak{e}_t})\).

To facilitate proving [prop:ProcConc], we first introduce some auxiliary notations. For any \(\tilde{\boldsymbol{\pi}}=(\tilde{\pi}^1,\dots,\tilde{\pi}^N)\in\widetilde{\Pi}^N\) and \(\xi\in\mathcal{P}(\mathbb{X})\), we define \[\begin{gather} \check{\boldsymbol{T}}^{\tilde{\boldsymbol{\pi}}}_{t,\xi} \boldsymbol{u}({\boldsymbol{x}}):= \int_{\mathbb{X}^{N}} \boldsymbol{u}({\boldsymbol{y}}) \left[\bigotimes_{n=1}^N Q^{\tilde{\pi}^n_{x^n, \xi}}_{t,x^n,\xi}\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}}), \quad \boldsymbol{u} \in B_b(\mathbb{X}). \end{gather}\] In contrary to 21 , \(\check{\boldsymbol{T}}^{\tilde{\boldsymbol{\pi}}}_{t,\xi}\) suggests the following auxiliary dynamics where the empirical population distribution in the \(N\)pG is replaced by a pre-specified \(\xi\), \[\begin{align} \boldsymbol{A}_t\sim\bigotimes_{n=1}^N\tilde{\pi}^n_{t,X^n_t,\xi}\,,\quad \boldsymbol{X}_{t+1}\sim\bigotimes_{n=1}^N P_{t,X^n_t,\xi, A^n_t}. \end{align}\] With the notations introduced in 3.6, we further define \[\begin{align} \label{eq:DefMeancTN} {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s,s,\overline{\Xi}}\boldsymbol{u}=\boldsymbol{u} \quad\text{and}\quad {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s,t,{\overline{\Xi}}} \boldsymbol{u} := {\check{\boldsymbol{T}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}_s}_{s,\overline{\xi}_s}\circ\cdots\circ{\check{\boldsymbol{T}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}_{t-1}}_{t-1,\overline{\xi}_{t-1}} \boldsymbol{u},\; 1\le s<t \le T, \quad \boldsymbol{u}\in B_b(\mathbb{X}^N), \end{align}\tag{66}\] where we note the definition is also valid for generic \(\widetilde{\boldsymbol{\mathfrak{P}}}\in\widetilde{\Pi}^{N\times(T-1)}\) and \(\Xi\in\mathcal{P}(\mathbb{X})^{T-1}\), although the generic case will not be considered. Observe that, if \(\boldsymbol{u}({\boldsymbol{x}})=\prod_{n=1}^N h_n(x^n)\) for some \(h_1,\dots,h_N\in B_b(\mathbb{X})\), we have \[\begin{align} \label{eq:DecompMeanT} {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s,t,{\overline{\Xi}}}\boldsymbol{u}({\boldsymbol{x}}) = \prod_{n=1}^N\overline{\mathcal{T}}^{\mathfrak{p}^n}_{s,t,{\overline{\Xi}}}h_n(x^n). \end{align}\tag{67}\] This in particular implies that, for fixed \({\boldsymbol{x}}\in\mathbb{X}^N\), the probability measure on \(\mathcal{B}(\mathbb{X}^{N})\) characterized by \(D\mapsto{\check{\boldsymbol{\mathcal{T}}}}^{{\boldsymbol{\mathfrak{P}}}}_{s,t,{\overline{\Xi}}}\mathbb{1}_D({\boldsymbol{x}})\) is the independent product of the probability measures on \(\mathcal{B}(\mathbb{X})\) characterized by \(B\mapsto \overline{\mathcal{T}}^{\mathfrak{p}^n}_{s,t,{\overline{\Xi}}}\mathbb{1}_B(x^n)\), \(n=1,\dots,N\).

Lemma 8. Under the setting of [prop:ProcConc], we have \[\begin{align} \left|{\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s,t,{\overline{\Xi}}} {\boldsymbol{e}}_t({\boldsymbol{x}}) - \inf_{x^n\in\mathbb{X}}{\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s,t,{\overline{\Xi}}} {\boldsymbol{e}}_t({\boldsymbol{x}})\right| \le \frac{1}{N},\quad {\boldsymbol{x}}\in\mathbb{X}^N,\;n=1,\dots,N. \end{align}\]

Proof. The statement is an immediate consequence of the independence described in 67 , 24 and the definition of \({\boldsymbol{e}}_t\) in 3.9. ◻

Lemma 9. Under the setting of [prop:ProcConc], for \(s< t\le T\), we have \[\begin{align} \left| \boldsymbol{T}^{(\mathfrak{p}^1_t,\tilde{\mathfrak{p}}^2_t,\dots,\tilde{\mathfrak{p}}^N_t)}_{s} \circ {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s+1,t,{\overline{\Xi}}} {\boldsymbol{e}}_{t}({\boldsymbol{x}}) - {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s,t,{\overline{\Xi}}} {\boldsymbol{e}}_{t}({\boldsymbol{x}}) \right| \le \frac{1}{N} + 2\left((L\vartheta+\eta){\boldsymbol{e}}_s({\boldsymbol{x}}) + \eta(2L^{-1})\right),\quad \boldsymbol{x}\in\mathbb{X}^N. \end{align}\]

Proof. Recall the definition of \(\boldsymbol{T}^{\boldsymbol{\pi}}_t\) from 24 as well as the definitions in 66 . Notice that \[\begin{align} \boldsymbol{T}^{(\mathfrak{p}^1_t,\tilde{\mathfrak{p}}^2_t,\dots,\tilde{\mathfrak{p}}^N_t)}_{s} \circ {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s+1,t,{\overline{\Xi}}} {\boldsymbol{e}}_{t}({\boldsymbol{x}}) = \int_{\mathbb{X}} {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s+1,t,{\overline{\Xi}}} {\boldsymbol{e}}_t(\boldsymbol{y})\left[Q^{\mathfrak{p}^1_{s,{\boldsymbol{x}}}}_{t,x^1,\overline{\delta}_{\boldsymbol{x}}}\otimes\bigotimes_{n=2}^N Q^{\tilde{\mathfrak{p}}^n_{s,x^n,\overline{\delta}_{\boldsymbol{x}}}}_{s,x^n,\overline{\delta}_{\boldsymbol{x}}}\right](\mathop{\mathrm{d \!}}\boldsymbol{y}). \end{align}\] Additionally, \[\begin{align} &{\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s,t,{\overline{\Xi}}} {\boldsymbol{e}}_{t}({\boldsymbol{x}}) = {\check{\boldsymbol{T}}}^{\tilde{\boldsymbol{\mathfrak{P}}}_s}_{s,\overline{\xi}_{s}}\circ{\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s+1,t,{\overline{\Xi}}} {\boldsymbol{e}}_{t}({\boldsymbol{x}}) = \int_{\mathbb{X}} {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s+1,t,{\overline{\Xi}}} {\boldsymbol{e}}_t(\boldsymbol{y}) \left[\bigotimes_{n=1}^N Q^{\tilde{\mathfrak{p}}^n_{s,x^n,\overline{\xi}_{s}}}_{s,x^n,\overline{\xi}_{s}}\right](\mathop{\mathrm{d \!}}\boldsymbol{y}). \end{align}\] The above together with 24 implies that \[\begin{align} &\left| \boldsymbol{T}^{(\mathfrak{p}^1_t,\tilde{\mathfrak{p}}^2_t,\dots,\tilde{\mathfrak{p}}^N_t)}_{s}\circ{\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s+1,t,{\overline{\Xi}}} {\boldsymbol{e}}_{t}({\boldsymbol{x}}) - {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s,t,{\overline{\Xi}}} {\boldsymbol{e}}_{t}({\boldsymbol{x}}) \right|\\ &\quad\le \sup_{y^2,\dots,y^N} \left|\int_{\mathbb{X}} {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s+1,t,{\overline{\Xi}}} {\boldsymbol{e}}_t({\boldsymbol{y}}) \left( Q^{\mathfrak{p}^1_{s,{\boldsymbol{x}}}}_{s,x^1,\overline{\delta}_{\boldsymbol{x}}}(\mathop{\mathrm{d \!}}y^1) - Q^{\tilde{\mathfrak{p}}^1_{s,x^n,\overline{\xi}_{s}}}_{s,x^1,\overline{\xi}_{s}}(\mathop{\mathrm{d \!}}y^1) \right)\right|\\ &\qquad+ \sum_{n=2}^N \sup_{\{y^1,\dots,y^N\}\setminus\{y^n\}} \left|\int_{\mathbb{X}} {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s+1,t,{\overline{\Xi}}} {\boldsymbol{e}}_t({\boldsymbol{y}}) \left( Q^{\tilde{\mathfrak{p}}^n_{s,x^n,\overline{\delta}_{\boldsymbol{x}}}}_{s,x^n,\overline{\delta}_{\boldsymbol{x}}}(\mathop{\mathrm{d \!}}y^n) - Q^{\tilde{\mathfrak{p}}^n_{s,x^n,\overline{\xi}_{s}}}_{s,x^n,\overline{\xi}_{s}}(\mathop{\mathrm{d \!}}y^n) \right)\right|\\ &\quad= \sup_{y^2,\dots,y^N} \left|\int_{\mathbb{X}} \left( {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s+1,t,{\overline{\Xi}}} {\boldsymbol{e}}_t({\boldsymbol{y}}) - \inf_{y^1} {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s+1,t,{\overline{\Xi}}} {\boldsymbol{e}}_t({\boldsymbol{y}}) \right) \left( Q^{\mathfrak{p}^1_{s,{\boldsymbol{x}}}}_{s,x^1,\overline{\delta}_{\boldsymbol{x}}}(\mathop{\mathrm{d \!}}y^1) - Q^{\tilde{\mathfrak{p}}^1_{s,x^1,\overline{\xi}_{s}}}_{s,x^1,\overline{\xi}_{s}}(\mathop{\mathrm{d \!}}y^1) \right)\right|\\ &\qquad+ \sum_{n=2}^N \sup_{\{y^1,\dots,y^N\}\setminus\{y^n\}} \left|\int_{\mathbb{X}} \left( {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s+1,t,{\overline{\Xi}}} {\boldsymbol{e}}_t({\boldsymbol{y}}) - \inf_{y^n} {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s+1,t,{\overline{\Xi}}} {\boldsymbol{e}}_t({\boldsymbol{y}}) \right) \left( Q^{\tilde{\mathfrak{p}}^n_{s,x^n,\overline{\delta}_{\boldsymbol{x}}}}_{s,x^n,\overline{\delta}_{\boldsymbol{x}}}(\mathop{\mathrm{d \!}}y^n) - Q^{\tilde{\mathfrak{p}}^n_{s,x^n,\overline{\xi}_{s}}}_{s,x^n,\overline{\xi}_{s}}(\mathop{\mathrm{d \!}}y^n) \right)\right|. \end{align}\] It follows from 8 and 18 that \[\begin{align} &\left| \boldsymbol{T}^{(\mathfrak{p}^1_t,\tilde{\mathfrak{p}}^2_t,\dots,\tilde{\mathfrak{p}}^N_t)}_{s} \circ {\check{\boldsymbol{\mathcal{T}}}}^{{\boldsymbol{\mathfrak{P}}}}_{s+1,t,{\overline{\Xi}}} {\boldsymbol{e}}_{t}({\boldsymbol{x}}) - {\check{\boldsymbol{\mathcal{T}}}}^{{\boldsymbol{\mathfrak{P}}}}_{s,t,{\overline{\Xi}}} {\boldsymbol{e}}_{t}({\boldsymbol{x}}) \right|\\ &\quad\le \frac{1}{N} + \frac{2}{N}\sum_{n=2}^N \left( L \|\tilde{\mathfrak{p}}^n(x^n,\overline{\delta}_{\boldsymbol{x}})-\tilde{\mathfrak{p}}^n(x^n,\overline{\xi}_s)\|_{BL} + \eta(2L^{-1}) + \eta(\|\overline{\delta}_{\boldsymbol{x}}-\overline{\xi}_s\|_{BL}) \right)\\ &\quad\le \frac{1}{N} + 2\left(L\vartheta({\boldsymbol{e}}_s({\boldsymbol{x}})) + \eta(2L^{-1}) + \eta({\boldsymbol{e}}_s({\boldsymbol{x}}))\right) = \frac{1}{N} + 2\left((L\vartheta + \eta){\boldsymbol{e}}_s({\boldsymbol{x}}) + \eta(2L^{-1})\right). \end{align}\] ◻

Lemma 10. Under the setting of [prop:ProcConc], for \(t\ge 2\) and \(L>1+\eta(2L^{-1})\) we have \[\begin{align} {\boldsymbol{\mathbb{T}}}^{{\boldsymbol{\mathfrak{P}}}}_{t} {\boldsymbol{e}}_t \le 2\sum_{r=1}^{t-1} (L\vartheta + \eta) \circ {\boldsymbol{\mathbb{T}}}^{{\boldsymbol{\mathfrak{P}}}}_{r}{\boldsymbol{e}}_{r} + (t-1)\left(\frac{1}{N} + 2\eta(2L^{-1})\right) + \mathfrak{r}_{\mathfrak{K}}(N), \end{align}\] where \(\mathfrak{r}_{\mathfrak{A}}\) is defined in 5.

Proof. Recall the definitions of \({\boldsymbol{\mathbb{T}}}^{\boldsymbol{\mathfrak{P}}}_t\) and \({\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_t\) in 26 and 66 , respectively. Expanding \({\boldsymbol{\mathbb{T}}}^{{\boldsymbol{\mathfrak{P}}}}_{t} {\boldsymbol{e}}_t\) via telescope sum, we yield \[\begin{align} {\boldsymbol{\mathbb{T}}}^{{\boldsymbol{\mathfrak{P}}}}_{t} {\boldsymbol{e}}_t &\le \sum_{r=1}^{t-1} {\boldsymbol{\mathbb{T}}}^{{\boldsymbol{\mathfrak{P}}}}_{r} \left| \boldsymbol{T}^{(\mathfrak{p}^1_r, \tilde{\mathfrak{p}}^2_r,\dots,\tilde{\mathfrak{p}}^N_r)}_{r} \circ {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r+1,t,{\overline{\Xi}}} {\boldsymbol{e}}_{t} - {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r,t,{\overline{\Xi}}} {\boldsymbol{e}}_{t} \right| + \mathring{\boldsymbol{T}}\circ{\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{1,t,{\overline{\Xi}}} {\boldsymbol{e}}_t. \end{align}\] This together with 9 implies \[\begin{align} {\boldsymbol{\mathbb{T}}}^{{\boldsymbol{\mathfrak{P}}}}_{t} {\boldsymbol{e}}_t &\le 2\sum_{r=1}^{t-1} {\boldsymbol{\mathbb{T}}}^{{\boldsymbol{\mathfrak{P}}}}_{r} \circ (L\vartheta+\eta){\boldsymbol{e}}_r ({\boldsymbol{x}}) + (t-1)\left(2\eta(2L^{-1})+\frac{1}{N}\right) + \mathring{\boldsymbol{T}} \circ {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{1,t,{\overline{\Xi}}} {\boldsymbol{e}}_t\\ &\le 2\sum_{r=1}^{t-1} (L\vartheta+\eta) \circ {\boldsymbol{\mathbb{T}}}^{{\boldsymbol{\mathfrak{P}}}}_{r} {\boldsymbol{e}}_r ({\boldsymbol{x}}) + (t-1)\left(2\eta(2L^{-1})+\frac{1}{N}\right) + \mathring{\boldsymbol{T}} \circ {\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{1,t,{\overline{\Xi}}} {\boldsymbol{e}}_t \end{align}\] where we have use Jensen’s inequality (cf. [50]) in the last line. Finally, recall the definition of \(\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\) below 29 . In view of the independence in 67 , by 5, we obtain \[\begin{align} \mathring{\boldsymbol{T}}\circ{\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{t,{\overline{\Xi}}} {\boldsymbol{e}}_t = \int_{\mathbb{X}^N} {\boldsymbol{e}}_t({\boldsymbol{y}}) \left[\bigotimes_{n=1}^N\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}^n}_{t,\overline{\Xi}}\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}}) \le \mathfrak{r}_{\mathfrak{K}}(N), \end{align}\] which completes the proof. ◻

We are ready to prove Proposition [prop:ProcConc].

Proof of [prop:ProcConc]. Let \(\mathfrak{e}_t\) be defined in 3.9. For \(t=1\), the statement is obvious from 26 and 5. We proceed by induction. Suppose that there is a \(t=1,\dots,T-1\) such that \({\boldsymbol{\mathbb{T}}}^{{\boldsymbol{\mathfrak{P}}}}_{r}{\boldsymbol{e}}_r\le \mathfrak{e}_{r}\) for all \(r=1,\dots,t\). By 10, for \(L>1+\eta(2L^{-1})\) we have \[\begin{align} {\boldsymbol{\mathbb{T}}}^{{\boldsymbol{\mathfrak{P}}}}_{t+1} {\boldsymbol{e}}_{t+1} \le 2\sum_{r=1}^{t} \big(L\vartheta(\mathfrak{e}_{r}) + \eta(\mathfrak{e}_{r})\big) + t\left(\frac{1}{N}+2\eta(2L^{-1})\right) + \mathfrak{r}_{\mathfrak{K}}(N), \end{align}\] thus \({\boldsymbol{\mathbb{T}}}^{{\boldsymbol{\mathfrak{P}}}}_{t+1} {\boldsymbol{e}}_{t+1}\le\mathfrak{e}_{t+1}\). Finally, invoking Jensen’s inequality (cf. [50]), we conclude that \({\boldsymbol{\mathbb{T}}}^{{\boldsymbol{\mathfrak{P}}}}_{t} \circ \beta {\boldsymbol{e}}_{t} \le \beta\circ {\boldsymbol{\mathbb{T}}}^{{\boldsymbol{\mathfrak{P}}}}_{t} {\boldsymbol{e}}_{t} \le \beta(\mathfrak{e}_{t}(N))\). ◻

Apart from Proposition [prop:ProcConc], the following result is also useful.

Let \(\widetilde{\boldsymbol{\mathfrak{P}}}\) and \(\overline{\Xi}\) be as introduced in 3.6. If 3, 5 and 7 holds, then for any \(h\in B_b(\mathbb{X})\) and \(t=1,\dots,T\), we have \[\left|\boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{t}h-\overline{\mathbb{T}}^{\tilde{\mathfrak{p}}^1}_{t,{\overline{\Xi}}}h\right| \le \|h\|_\infty \mathfrak{e}_{t},\] where \(\mathfrak{e}_t\) is defined in 3.9, and when acted by \(\boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{t}\), \(h\) is treated as a function from \(B_b(\mathbb{X}^N)\) that is constant in \((x^2,\dots,x^N)\) .

Proof. Recall the definition of \(\boldsymbol{\mathbb{T}}^{\boldsymbol{\mathfrak{P}}}_t\) and \(\overline{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_t\) from 26 and 29 , respectively. Note that \[\begin{align} \left|\boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{t} h - \overline{\mathbb{T}}^{\tilde{\mathfrak{p}}^1}_{t,{\overline{\Xi}}} h\right| &\le \sum_{r=1}^{t-1} \left|\boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r+1} \circ \overline{\mathcal{T}}^{\tilde{\mathfrak{p}}^1}_{r+1,t,{\overline{\Xi}}} h - \boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r} \circ \overline{\mathcal{T}}^{\tilde{\mathfrak{p}}^1}_{r,t,{\overline{\Xi}}} h\right| \le \sum_{r=1}^{t-1} \boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r} \left| \left( \boldsymbol{T}^{(\tilde{\mathfrak{p}}^1_r,\dots,\tilde{\mathfrak{p}}^N_r)}_r - \overline{T}^{\tilde{\mathfrak{p}}^1_r}_{r,\overline{\xi}_r} \right) \circ \overline{\mathcal{T}}^{\tilde{\mathfrak{p}}^1}_{r+1,t,{\overline{\Xi}}} h\right|. \end{align}\] Let \(g=\overline{\mathcal{T}}^{\tilde{\mathfrak{p}}^1}_{r+1,t,{\overline{\Xi}}} h\). Clearly, \(\|g\|_\infty\le\|h\|_\infty\). In view of 24 and 27 , for \(L>1+\eta(2L^{-1})\), \[\begin{align} &\left|\boldsymbol{T}^{(\tilde{\mathfrak{p}}^1_r,\dots,\tilde{\mathfrak{p}}^N_r)}_r g ({\boldsymbol{x}}) - \overline{T}^{\tilde{\mathfrak{p}}^1_r}_{r,\overline{\xi}_r} g ({\boldsymbol{x}})\right| = \left|\int_\mathbb{X}g(y) Q^{\tilde{\mathfrak{p}}^1_{t,x^1,\overline{\delta}_{\boldsymbol{x}}}}_{r,x^1,\overline{\delta}_{\boldsymbol{x}}}(\mathop{\mathrm{d \!}}y) - \int_\mathbb{X}g(y) Q^{\tilde{\mathfrak{p}}^1_{r,x^1,\overline{\xi}_t}}_{r,x^1,\overline{\xi}_r}(\mathop{\mathrm{d \!}}y)\right|\\ &\quad\le \|h\|_\infty \left( ( L\vartheta + \eta) {\boldsymbol{e}}_t({\boldsymbol{x}}) + 2\eta(2L^{-1}) \right), \end{align}\] where we have used 18 in the last inequality. It follows from [prop:ProcConc] that for \(L>1+\eta(2L^{-1})\), \[\begin{align} \left|\boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{t} h - \overline{\mathbb{T}}^{\tilde{\mathfrak{p}}^1}_{t,{\overline{\Xi}}} h\right| \le 2\|h\|_\infty \left( \sum_{r=1}^{t-1}( L\vartheta(\mathfrak{e}_t) + \eta(\mathfrak{e}_t)) + (t-1)\eta(2L^{-1}) \right). \end{align}\] Taking infimum over \(L>1+\eta(2L^{-1})\) on both hand sides above, we conclude the proof. ◻

7.2 Approximating \(N\)-player stepwise exploitability↩︎

In this section, we write \(V_{\overline{\xi}_{T}} := V(\cdot,\overline{\xi}_{T})\).

Lemma 11. Let \(t=2,\dots,T-1\). Let \(\widetilde{\boldsymbol{\mathfrak{P}}}\), \(\overline{\Xi}\) and \(\overline{\xi}_T\) be as introduced in 3.6. Consider a \({\boldsymbol{\mathfrak{P}}}\in\Pi^{N\times(T-1)}\) such that \(\mathfrak{p}^1_r\in\widetilde{\Pi}\) for \(r=t,\dots,T-1\) is \(\vartheta^1\)-symmetrically continuous, and \(\mathfrak{p}^n=\tilde{\mathfrak{p}}^n\) for \(n=2,\dots,N\). Suppose 2, 3, 4 (i)-(iii) (v), 5, 6 and 7. Then,15 \[\begin{align} &\boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{t}\left|\boldsymbol{\mathcal{S}}^{{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}} - \overline{\mathcal{S}}^{\mathfrak{p}^1}_{t,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}}\right| \le \sum_{r=t}^{T-1} \bar{c}^{r-t}\mathcal{C}_{T-r}(\|V\|_\infty) \big(\zeta(\vartheta^1(\mathfrak{e}_{r})) + \zeta(\mathfrak{e}_{r}) \big) + \bar{c}^{T-t}\iota({\mathfrak{e}_T}). \end{align}\]

Proof. Observe that \[\begin{align} &\boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{t}\left|\boldsymbol{\mathcal{S}}^{{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}} - \overline{\mathcal{S}}^{\mathfrak{p}^1}_{t,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} \right|\\ &\quad\le \sum_{r=t}^{T-1} \boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{t} \left| \boldsymbol{\mathcal{S}}^{{\boldsymbol{\mathfrak{P}}}}_{t,r+1} \circ \overline{\mathcal{S}}^{\mathfrak{p}^1}_{r+1,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}} - \boldsymbol{\mathcal{S}}^{{\boldsymbol{\mathfrak{P}}}}_{t,r} \circ \overline{\mathcal{S}}^{\mathfrak{p}^1}_{r,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}} \right| + \boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{t} \left|\boldsymbol{\mathcal{S}}^{{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}} - \boldsymbol{\mathcal{S}}^{{\boldsymbol{\mathfrak{P}}}}_{t,T} V_{\overline{\xi}_{T}}\right|\\ &\quad\le \sum_{r=t}^{T-1} \bar{c}^{r-t} \boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{r} \left| \boldsymbol{S}^{(\mathfrak{p}^1_t,\dots,\mathfrak{p}^N_t)}_{r} \circ \overline{\mathcal{S}}^{\mathfrak{p}^1}_{r+1,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}} - \overline{\mathcal{S}}^{\mathfrak{p}^1}_{r,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}} \right| + \bar{c}^{T-t}\boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{T}\!\! \circ \iota {\boldsymbol{e}}_T, \end{align}\] where we have used 33 , 17 and 26 in last inequality. In addition, due to 32 , 4 (v), 38 and 39 , we have \[\begin{align} &\left| \left[\boldsymbol{S}^{(\mathfrak{p}^1_t,\dots,\mathfrak{p}^N_t)}_{r} \circ \overline{\mathcal{S}}^{\mathfrak{p}^1}_{r+1,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}}\right]({\boldsymbol{x}}) - \left[\overline{\mathcal{S}}^{\mathfrak{p}^1}_{r,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}}\right](x^1) \right|\\ &\quad= \left| \left[\overline{G}^{\mathfrak{p}^1_{r,x^1,\overline{\delta}_{\boldsymbol{x}}}}_{r,\overline{\delta}_{\boldsymbol{x}}} \circ \overline{\mathcal{S}}^{\mathfrak{p}^1}_{r+1,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}}\right](x^1) - \left[\overline{G}^{\mathfrak{p}^1_{r,x^1,\overline{\xi}_r}}_{r,\overline{\xi}_r}\circ\overline{\mathcal{S}}^{\mathfrak{p}^1}_{r+1,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}}\right](x^1) \right|\\ &\quad\le \mathcal{C}_{T-r}(\|V\|_\infty)\left(\zeta(\vartheta^1(\|\overline{\delta}_{\boldsymbol{x}}-\overline{\xi}_r\|_{BL})) + \zeta(\|\overline{\delta}_{\boldsymbol{x}}-\overline{\xi}_r\|_{BL})\right), \end{align}\] where we have used 6 (ii) and 16 in the last inequality. Note \(\ell\mapsto\zeta(\vartheta^1(\ell))+\zeta(\ell)\) is also a subadditive modulus of continuity. The above together with [prop:ProcConc] completes the proof. ◻

Let \(\widetilde{\boldsymbol{\mathfrak{P}}}\), \(\overline{\Xi}\) and \(\overline{\xi}_T\) be as introduced in 3.6. Consider a \({\boldsymbol{\mathfrak{P}}}\in\Pi^{N\times(T-1)}\) that satisfies \(\mathfrak{p}^n=\tilde{\mathfrak{p}}^n\) for \(n=2,\dots,N\). Suppose 2, 3, 4 (i)-(iii) (v), 5, 6 and 7. Then, we have \(\boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{T}\left|{\boldsymbol{U}} - V_{\overline{\xi}_T} \right| \le \iota({\mathfrak{e}_T})\), and \[\begin{align} &\boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{t} \left| \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}} - \overline{\mathcal{S}}^{*}_{t,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}} \right| \le (T+1-t)\left(\sum_{r=t}^{T-1} \bar{c}^{r-t} \mathcal{C}_{T-r}(\|V\|_\infty) \zeta(\mathfrak{e}_{r}) + \bar{c}^{T-t}\iota({\mathfrak{e}_T})\right),\quad t=1,\dots,T-1. \end{align}\]

Proof. For \(t=T\), owing Proposition [prop:ProcConc] and the hypothesis that \(V\) is \(\iota\)-symmetrically continuous, we have \(\boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{T}\left|{\boldsymbol{U}} - V_{\overline{\xi}_T} \right| \le \iota({\mathfrak{e}_T})\), which proves the statement for \(t=T\). We proceed by backward induction. Suppose the statement is true at \(t+1\) for some \(t=1,\dots,T-1\). In view of 14, let \(\mathfrak{p}^*\in\Pi^{T-1}\) and \(\overline{\mathfrak{p}}^*\in\overline{\Pi}^{T-1}\) be the optimal policies attaining \(\boldsymbol{\mathbb{S}}^{*\mathfrak{P}}_TV\) and \(\overline{\mathbb{S}}^{*}_{T,(\overline{\xi}_{1},\dots,\overline{\xi}_{T-1})}V_{\overline{\xi}_{T}}\), respectively. It follows that \[\begin{align} \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}}({\boldsymbol{x}}) - \overline{\mathcal{S}}^{*}_{t,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}}(x^1) &\le \boldsymbol{\mathcal{S}}^{(\overline{\mathfrak{p}}^*,\tilde{\mathfrak{p}}^2,\dots,\tilde{\mathfrak{p}}^N)}_{t,T} {\boldsymbol{U}}({\boldsymbol{x}}) - \overline{\mathcal{S}}^{\mathfrak{p}^*}_{t,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}}(x^1)\\ &\le \left|\boldsymbol{\mathcal{S}}^{(\overline{\mathfrak{p}}^*,\tilde{\mathfrak{p}}^2,\dots,\tilde{\mathfrak{p}}^N)}_{t,T} {\boldsymbol{U}}({\boldsymbol{x}}) - \overline{\mathcal{S}}^{\mathfrak{p}^*}_{t,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}}(x^1)\right|. \end{align}\] Meanwhile, \[\begin{align} \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}}({\boldsymbol{x}}) - \overline{\mathcal{S}}^{*}_{t,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}}(x^1) &\ge \boldsymbol{S}^{(\mathfrak{p}^*_t,\tilde{\mathfrak{p}}^2_t,\dots,\tilde{\mathfrak{p}}^N_t)}_{t} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}}({\boldsymbol{x}}) - \overline{\mathcal{S}}^{*}_{t,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}}(x^1)\\ &\qquad- \left|\boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}}({\boldsymbol{x}}) - \boldsymbol{S}^{(\mathfrak{p}^*_t,\tilde{\mathfrak{p}}^2_t,\dots,\tilde{\mathfrak{p}}^N_t)}_{t} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}}({\boldsymbol{x}})\right|, \end{align}\] For the first term in the right hand side above, in view of 4 (v), 6 (ii) and 16, we further yield \[\begin{align} &\left[\boldsymbol{S}^{(\mathfrak{p}^*_t,\tilde{\mathfrak{p}}^2_t,\dots,\tilde{\mathfrak{p}}^N_t)}_{t} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}}\right] ({\boldsymbol{x}}) = \left[\overline{G}^{\mathfrak{p}^*_{t,{\boldsymbol{x}}}}_{t,\overline{\delta}_{\boldsymbol{x}}} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}}\right] (x^1)\\ &\quad\ge \left[\overline{G}^{\mathfrak{p}^*_{t,{\boldsymbol{x}}}}_{t,\overline{\xi}_t} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}}\right] (x^1) - \left(c_0 + c_1\mathcal{C}_{T-(t+1)}(\|V\|_\infty)\right)\zeta(\|\overline{\delta}_{\boldsymbol{x}}-\overline{\xi}_{t}\|_{BL})\\ &\quad\ge \left[\overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}}\right](x^1) - \mathcal{C}_{T-t}(\|V\|_\infty)\zeta({\boldsymbol{e}}_t({\boldsymbol{x}})). \end{align}\] Therefore, \[\begin{align} \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}}({\boldsymbol{x}}) - \overline{\mathcal{S}}^{*}_{t,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}}(x^1) \ge - \mathcal{C}_{T-t}(\|V\|_\infty)\zeta {\boldsymbol{e}}_t({\boldsymbol{x}}) - \left|\boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}}({\boldsymbol{x}}) - \boldsymbol{S}^{(\mathfrak{p}^*_t,\tilde{\mathfrak{p}}^2_t,\dots,\tilde{\mathfrak{p}}^N_t)}_{t} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}}({\boldsymbol{x}})\right|. \end{align}\] Note additionally that for any \(c,\underline c,\bar{c}\in\mathbb{R}\) with \(\underline c\le 0\le \bar{c}\) and \(\underline c \le c \le \bar{c}\), we have \(|c|\le|\underline c|+|\bar{c}|\). By combining this with the estimation above, we yield \[\begin{align} \boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{t} \left| \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}} - \overline{\mathcal{S}}^{*}_{t,T,{\overline{\Xi}}}V_{\overline{\xi}_T} \right| \le I_1 + I_2, \end{align}\] where \[\begin{gather} I_1 := \boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{t}\left|\boldsymbol{\mathcal{S}}^{(\overline{\mathfrak{p}}^*,\tilde{\mathfrak{p}}^2,\dots,\tilde{\mathfrak{p}}^N)}_{t,T} {\boldsymbol{U}} - \overline{\mathcal{S}}^{\overline{\mathfrak{p}}^*}_{t,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} \right|\\ I_2 := \mathcal{C}_{T-t}(\|V\|_\infty)\boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{t} \circ \zeta {\boldsymbol{e}}_t + \boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{t} \left|\boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}} - \boldsymbol{S}^{(\mathfrak{p}^*_t,\tilde{\mathfrak{p}}^2_t,\dots,\tilde{\mathfrak{p}}^N_t)}_{t} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}}\right|. \end{gather}\] As for \(I_1\), by defining \(\mathfrak{p}^1\oplus_{t}\overline{\mathfrak{p}}^*:=(\mathfrak{p}^1,\dots,\mathfrak{p}^1_{t-1},\overline{\mathfrak{p}}^*_t,\dots,\overline{\mathfrak{p}}^*_{T-1})\) if \(t\ge 2\) and \(\mathfrak{p}^1\oplus_{t}\overline{\mathfrak{p}}^*:=\overline{\mathfrak{p}}^*\) if \(t=1\), we have \(\boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{t} = \boldsymbol{\mathbb{T}}^{(\mathfrak{p}^1\oplus_{t}\overline{\mathfrak{p}}^*,\tilde{\mathfrak{p}}^2,\dots,\tilde{\mathfrak{p}}^N)}_{t}\) and \(\boldsymbol{\mathcal{S}}^{\overline{\mathfrak{p}}^*,{\boldsymbol{\mathfrak{P}}}}_{t,T} = \boldsymbol{\mathcal{S}}^{(\mathfrak{p}^1\oplus_{t}\overline{\mathfrak{p}}^*,\tilde{\mathfrak{p}}^2,\dots,\tilde{\mathfrak{p}}^N)}_{t,T}\). This together with 11 provides an estimation on \(I_1\) (note \(\vartheta^1\equiv 0\) in this case). As for \(I_2\), by [prop:ProcConc], 26 , 4 (iii) and the induction hypothesis, we yield16 \[\begin{align} I_2 &\le \mathcal{C}_{T-t}(\|V\|_\infty)\zeta({\mathfrak{e}_t}) - \bar{c}\boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{t} \circ \boldsymbol{T}^{(\mathfrak{p}^*_t,\tilde{\mathfrak{p}}^2_t,\dots,\tilde{\mathfrak{p}}^N_t)}_{t} \left| \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t+1,T} {\boldsymbol{U}} - \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}} \right|\\ &\le (T-t)\left(\mathcal{C}_{T-t}(\|V\|_\infty)\zeta({\mathfrak{e}_t}) - \sum_{r=t+1}^{T-1} \bar{c}^{r-t} \mathcal{C}_{T-r}(\|V\|_\infty) \zeta(\mathfrak{e}_{r}) + c_1^{T-t}\iota({\mathfrak{e}_T})\right)\\ &= (T-t)\left(\sum_{r=t}^{T-1} \bar{c}^{r-t} \mathcal{C}_{T-r}(\|V\|_\infty) \zeta(\mathfrak{e}_{r}) + \overline{C}^{T-t}\iota({\mathfrak{e}_T})\right). \end{align}\] By combining the estimates of \(I_1\) and \(I_2\), the proof is complete. ◻

Let \(\widetilde{\boldsymbol{\mathfrak{P}}}\), \(\overline{\Xi}\) and \(\overline{\xi}_T\) be as introduced in 3.6. Suppose 2, 3, 4 (i)-(iii) (v), 5, 6 and 7. Then, with \(\mathfrak{E}\) defined in 3.9, we have \[\begin{align} \left|\boldsymbol{\mathfrak{R}}(\widetilde{\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}) - \sum_{t=1}^{T-1} \bar{c}^t \overline{\mathbb{T}}^{\tilde{\mathfrak{p}}^1}_{t,{\overline{\Xi}}} \left( \overline{S}^{\tilde{\mathfrak{p}}^1}_{t,\overline{\xi}_t} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} - \overline{\mathcal{S}}^{*}_{t,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} \right)\right| \le {\mathfrak{E}}. \end{align}\]

Proof. Let \(t=1,\dots,T-1\). Observe that, by 26 and 17, \[\begin{align} \left|\boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_t \circ \boldsymbol{S}^{\widetilde{\boldsymbol{\mathfrak{P}}}_t}_t \circ \boldsymbol{\mathcal{S}}^{*\widetilde{\boldsymbol{\mathfrak{P}}}}_{t+1,T} {\boldsymbol{U}} - \boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_t \circ \boldsymbol{S}^{\widetilde{\boldsymbol{\mathfrak{P}}}_t}_t \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}}\right| \le \bar{c} \boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{t+1}\left| \boldsymbol{\mathcal{S}}^{*\widetilde{\boldsymbol{\mathfrak{P}}}}_{t+1,T} {\boldsymbol{U}} - \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}}\right|. \end{align}\] In addition, by 32 and 18, we have \[\begin{align} &\left| \boldsymbol{S}^{\widetilde{\boldsymbol{\mathfrak{P}}}_t}_t \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}} ({\boldsymbol{x}}) - \overline{S}^{\tilde{\mathfrak{p}}^1_t}_{t,\overline{\xi}_t} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}}V_{\overline{\xi}_{T}} ({\boldsymbol{x}}) \right|\\ &\quad= \left| \overline{G}^{\tilde{\mathfrak{p}}^1_{t,x^1,\overline{\delta}_{\boldsymbol{x}}}}_{t,\overline{\delta}_{\boldsymbol{x}}} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} ({\boldsymbol{x}}) - \overline{G}^{\tilde{\mathfrak{p}}^1_{t,x^1,\overline{\xi}_t}}_{t,\overline{\xi}_t} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} ({\boldsymbol{x}}) \right| \le \mathcal{C}_{T-t}(\|V\|_\infty)(\zeta\vartheta+\zeta){\boldsymbol{e}}_t({\boldsymbol{x}}). \end{align}\] By combining the above, we yield \[\begin{gather} \left|\boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_t \circ \boldsymbol{S}^{\widetilde{\boldsymbol{\mathfrak{P}}}_t}_t \circ \boldsymbol{\mathcal{S}}^{*\widetilde{\boldsymbol{\mathfrak{P}}}}_{t+1,T} {\boldsymbol{U}} - \boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_t \circ \overline{S}^{\tilde{\mathfrak{p}}^1_t}_{t,{\overline{\xi}_t}} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} ({\boldsymbol{x}})\right|\\ \le \bar{c} \boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{t+1}\left| \boldsymbol{\mathcal{S}}^{*\widetilde{\boldsymbol{\mathfrak{P}}}}_{t+1,T} {\boldsymbol{U}} - \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}}V_{{\overline{\Psi}}}\right| + \mathcal{C}_{T-t}(\|V\|_\infty) \boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_t \circ (\zeta\vartheta+\zeta){\boldsymbol{e}}_t. \end{gather}\] This together with [prop:EstfTDiffcSstar] and [prop:ProcConc] implies that \[\begin{align} & \left| \boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_t \left(\boldsymbol{S}^{\widetilde{\boldsymbol{\mathfrak{P}}}_t}_t \circ \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t+1,T} {\boldsymbol{U}} - \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}}\right) - \boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_t \left( \overline{S}^{\tilde{\mathfrak{p}}^1}_{t,{\overline{\xi}_t}} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} - \overline{\mathcal{S}}^{*}_{t,T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} \right) \right| \\ &\quad\le \bar{c} (T-t)\left(\sum_{r=t+1}^{T-1} \bar{c}^{r-(t+1)} \mathcal{C}_{T-r}(\|V\|_\infty) \zeta(\mathfrak{e}_{r}) + \bar{c}^{T-(t+1)}\iota({\mathfrak{e}_T})\right)\\ &\qquad{!}{ + \mathcal{C}_{T-t}(\|V\|_\infty) \big(\zeta\vartheta({\mathfrak{e}_t})+\zeta({\mathfrak{e}_t})\big) + (T+1-t) \left(\sum_{r=t}^{T-1} \bar{c}^{r-t} \mathcal{C}_{T-r}(\|V\|_\infty) \zeta(\mathfrak{e}_{r}) + \bar{c}^{T-t} \iota({\mathfrak{e}_T})\right) }. \end{align}\] Finally, recall the definition of \(\boldsymbol{\mathfrak{R}}\) and \(\mathfrak{E}\) from 53 and 3.9, respectively, the proof is complete. ◻

7.3 Proof of 7↩︎

Proof of 7. Due to 2, \(\int_\mathbb{X}\overline{S}^{{\overline{\pi}^{\psi_t}}}_{t,{\xi^{\psi_t}}} v(y) {\xi^{\psi_t}}(\mathop{\mathrm{d \!}}y) = \int_\mathbb{X}\overline{S}^{{\overline{\pi}^{\psi_t}}'}_{t,{\xi^{\psi_t}}} v(y) {\xi^{\psi_t}}(\mathop{\mathrm{d \!}}y)\). Next, in view of 40 , 17 and 3, expanding the left hand side below by telescope sum, we yield \[\begin{align} \left|\overline{\mathbb{S}}^{{\overline{\mathfrak{p}}^\Psi}'}_{T,{\Xi^{\Psi}}} v - \overline{\mathbb{S}}^{{\overline{\mathfrak{p}}^\Psi}}_{T,{\Xi^{\Psi}}} v\right| &\le \sum_{t=1}^{T-1} \left|\overline{\mathbb{S}}^{{\overline{\mathfrak{p}}^\Psi}}_{t,{\Xi^{\Psi}}}\circ\overline{\mathcal{S}}^{{\overline{\mathfrak{p}}^\Psi}'}_{t,T,{\Xi^{\overline{\Psi}}}} v - \overline{\mathbb{S}}^{{\overline{\mathfrak{p}}^\Psi}}_{t+1,{\Xi^{\overline{\Psi}}}}\circ\overline{\mathcal{S}}^{{\overline{\mathfrak{p}}^\Psi}'}_{t+1,T,{\Xi^{\Psi}}} v\right| \\ &\le \sum_{t=1}^{T-1} \bar{c}^{t} \overline{\mathbb{T}}^{{\overline{\mathfrak{p}}^\Psi}}_{t,{\Xi^{\Psi}}} \left|\overline{\mathcal{S}}^{{\overline{\mathfrak{p}}^\Psi}'}_{t,T,{\Xi^{\overline{\Psi}}}} v - \overline{S}^{{\overline{\pi}^{\psi_t}}}_{t,{\xi^{\psi_t}}} \circ \overline{\mathcal{S}}^{{\overline{\mathfrak{p}}^\Psi}'}_{t+1,T,{\Xi^{\Psi}}} v\right|\\ &= \sum_{t=1}^{T-1} \bar{c}^{t}\int_{\mathbb{X}} \left|\overline{\mathcal{S}}^{{\overline{\mathfrak{p}}^\Psi}'}_{t,T,{\Xi^{\overline{\Psi}}}} v(y) - \overline{S}^{{\overline{\pi}^{\psi_t}}}_{t,{\xi^{\psi_t}}} \circ \overline{\mathcal{S}}^{{\overline{\mathfrak{p}}^\Psi}'}_{t+1,T,{\Xi^{\Psi}}} v(y)\right| {\xi^{\psi_t}}(\mathop{\mathrm{d \!}}y) = 0. \end{align}\] ◻

7.4 Proof of [prop:EstEndExploitabilityCont]↩︎

Proof of [prop:EstEndExploitabilityCont](a). Note that \[\begin{align} \label{eq:EndExploitabilityTelescope} \boldsymbol{\mathcal{R}}({\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}) &= \sum_{t=1}^{T-1} \left(\boldsymbol{\mathbb{S}}^{{\boldsymbol{\mathfrak{P}}}}_{t+1} \circ \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t+1,T} {\boldsymbol{U}} - \boldsymbol{\mathbb{S}}^{{\boldsymbol{\mathfrak{P}}}}_{t} \circ \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}}\right) = \sum_{t=1}^{T-1} \left(\boldsymbol{\mathbb{S}}^{{\boldsymbol{\mathfrak{P}}}}_{t} \circ \boldsymbol{S}^{\boldsymbol{\mathfrak{P}}_t}_t \circ \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t+1,T} {\boldsymbol{U}} - \boldsymbol{\mathbb{S}}^{{\boldsymbol{\mathfrak{P}}}}_{t} \circ \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}}\right). \end{align}\tag{68}\] Then, by 17, we yields. \[\begin{align} \boldsymbol{\mathcal{R}}({\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}) \le \sum_{t=1}^{T-1} \bar{c}^t \boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_t \left(\boldsymbol{S}^{{\boldsymbol{\mathfrak{P}}}_t}_t \circ \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t+1,T} {\boldsymbol{U}} - \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}}\right) = \boldsymbol{\mathfrak{R}}({\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}). \end{align}\] ◻

Proof of [prop:EstEndExploitabilityCont](b). With a similar induction as in the proof of 17, for any \({\boldsymbol{u}}\ge{\boldsymbol{u}}'\) we have \[\begin{align} \boldsymbol{\mathbb{S}}^{{\boldsymbol{\mathfrak{P}}}}_{t}{\boldsymbol{u}} - \boldsymbol{\mathbb{S}}^{{\boldsymbol{\mathfrak{P}}}}_{t}{\boldsymbol{u}}' \ge \underline c^{t} \boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{t} \left(u -{\boldsymbol{u}}'\right), \quad t=1,\dots,T. \end{align}\] This together with 68 implies that \[\begin{align} \boldsymbol{\mathcal{R}}({\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}) \ge \sum_{t=1}^{T-1} \underline c^t \boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_t \left(\boldsymbol{S}^{{\boldsymbol{\mathfrak{P}}}_t}_t \circ \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t+1,T} {\boldsymbol{U}} - \boldsymbol{\mathcal{S}}^{*{\boldsymbol{\mathfrak{P}}}}_{t,T} {\boldsymbol{U}}\right) \ge \left({\underline c}\middle/{\bar{c}}\right)^{T-1}\boldsymbol{\mathfrak{R}}({\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}). \end{align}\] ◻

Proof of [prop:EstEndExploitabilityCont](c). To start with, observe that \[\begin{align} \label{eq:EstEndExploitabilityCont} &\left|{\boldsymbol{\mathcal{R}}_{\vartheta}}(\widetilde{\boldsymbol{\mathfrak{P}}};\boldsymbol{U}) - \left( \overline{\mathbb{S}}^{\tilde{\mathfrak{p}}^1}_{T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} - \overline{\mathbb{S}}^{*}_{T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} \right)\right|\nonumber\\ &\quad= \left| \left( \boldsymbol{\mathbb{S}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_T {\boldsymbol{U}} - \inf_{\tilde{\mathfrak{p}}\in\widetilde{\Pi}_\vartheta}\left\{\boldsymbol{\mathbb{S}}^{(\tilde{\mathfrak{p}},\tilde{\mathfrak{p}}^2,\dots,\tilde{\mathfrak{p}}^N)}_{t} {\boldsymbol{U}} \right\} \right) - \left( \overline{\mathbb{S}}^{\tilde{\mathfrak{p}}^1}_{T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} - \inf_{\tilde{\mathfrak{p}}\in\widetilde{\Pi}_\vartheta}\left\{\overline{\mathbb{S}}^{\tilde{\mathfrak{p}}}_{T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} \right\} \right) \right|\nonumber \\ &\quad\le \left| \boldsymbol{\mathbb{S}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_T \boldsymbol{U} - \overline{\mathbb{S}}^{\tilde{\mathfrak{p}}^1}_{T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} \right| + \sup_{\tilde{\mathfrak{p}}\in\widetilde{\Pi}_\vartheta}\left| \boldsymbol{\mathbb{S}}^{(\tilde{\mathfrak{p}},\tilde{\mathfrak{p}}^2,\dots,\tilde{\mathfrak{p}}^N)}_T \boldsymbol{U} - \overline{\mathbb{S}}^{\tilde{\mathfrak{p}}}_{T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} \right| \nonumber\\ &\quad\le 2\left(\sum_{r=1}^{T-1} \bar{c}^{r}\mathcal{C}_{T-r}(\|V\|_\infty) \big(\zeta(\vartheta(\mathfrak{e}_{r})) + \zeta(\mathfrak{e}_{r}) \big) + \bar{c}^{T}\iota({\mathfrak{e}_T})\right), \end{align}\tag{69}\] where we have used 4 (v), ?? in 17 with \(t=1\), and 11 in the last inequality.

In order to proceed, in view of 14, we let \(\mathfrak{p}^*\) and \(\bar\mathfrak{p}^*\) be the optimal policies that attains \(\boldsymbol{\mathbb{S}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}{\boldsymbol{U}}\) and \(\overline{\mathbb{S}}^{*}_{T,{\overline{\Xi}}} V_{\overline{\xi}_{T}}\), respectively. Then, \[\begin{align} &\left|\boldsymbol{\mathcal{R}}(\widetilde{\boldsymbol{\mathfrak{P}}};\boldsymbol{U}) - \left( \overline{\mathbb{S}}^{\tilde{\mathfrak{p}}^1}_{T,\overline{\Xi}} V_{\overline{\xi}_{T}} - \overline{\mathbb{S}}^{*}_{T,\overline{\Xi}} V_{\overline{\xi}_{T}} \right) \right| \le \left|\boldsymbol{\mathbb{S}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_T {\boldsymbol{U}} - \overline{\mathbb{S}}^{\tilde{\mathfrak{p}}^1}_{T,\overline{\Xi}} V_{\overline{\xi}_{T}} \right| + \left| \boldsymbol{\mathbb{S}}^{*\widetilde{\boldsymbol{\mathfrak{P}}}}_{T} {\boldsymbol{U}} - \overline{\mathbb{S}}^{*}_{T,\overline{\Xi}} V_{\overline{\xi}_{T}} \right|. \end{align}\] By 4 (v), ?? in 17 with \(t=1\), 11, and [prop:EstfTDiffcSstar], we yield \[\begin{align} &\left|\boldsymbol{\mathcal{R}}(\widetilde{\boldsymbol{\mathfrak{P}}};V) - \left( \overline{\mathbb{S}}^{\tilde{\mathfrak{p}}^1}_{T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} - \overline{\mathbb{S}}^{*}_{T,{\overline{\Xi}}} V_{\overline{\xi}_{T}} \right) \right|\\ &\quad \le \sum_{r=1}^{T-1} \bar{c}^{r}\mathcal{C}_{T-r}(\|V\|_\infty) \big(\zeta(\vartheta(\mathfrak{e}_{r})) + \zeta(\mathfrak{e}_{r}) \big) + \bar{c}^{T}\iota({\mathfrak{e}_T}) + T\left(\sum_{r=1}^{T-1} \bar{c}^{r} \mathcal{C}_{T-r}(\|V\|_\infty) \zeta(\mathfrak{e}_{r}) + \bar{c}^{T}\iota({\mathfrak{e}_T})\right) \end{align}\] This together with 69 and triangle inequality completes the proof. ◻

7.5 Proof of ?? in 3↩︎

Proof of ?? . Recall the definition of \(\breve{\boldsymbol{e}}_t\) from 50 . The statement for \(t=1\) follows immediately from 5. We now suppose \(t\ge 2\). Note for any \(\{(x^k,a^k)\}_{k=1}^N\), we have \[\begin{align} \label{eq:BLBound} \left\|\overline{\delta}_{(x^k,a^k)_{k=1}^N}-{\overline{\psi}}_t\right\|_{BL} - \inf_{(x^n,a^n)}\left\|\overline{\delta}_{(x^k,a^k)_{k=1}^N}-{\overline{\psi}}_t\right\|_{BL} \le \frac{1}{N},\quad n=1,\dots,N, \end{align}\tag{70}\] and \((x^n,a^n) \mapsto \left\|\overline{\delta}_{(x^k,a^k)_{k=1}^N}-{\overline{\psi}}_t\right\|_{BL}\) is \(N^{-1}\)-Lipschitz continuous. We define \[\begin{align} \bar {\boldsymbol{e}}_t({\boldsymbol{x}}) := \int_{\mathbb{A}^N}\left\|\overline{\delta}_{(x^k,a^k)_{k=1}^N}-{\overline{\psi}}_t\right\|_{BL} \left[\bigotimes_{n=1}^N\tilde{\mathfrak{p}}_t^n(x^n,\overline{\xi}_t)\right](\mathop{\mathrm{d \!}}{\boldsymbol{a}}). \end{align}\] The above together with 24 and the 7 that \(\tilde{\mathfrak{p}}^n\)’s are \(\vartheta\)-symmetrically continuous, implies \[\begin{align} &\left|\breve {\boldsymbol{e}}_t({\boldsymbol{x}}) - \bar {\boldsymbol{e}}_t({\boldsymbol{x}})\right|\\ &\quad\le \sum_{n=1}^N \sup_{\{a_1,\dots,a_N\}\setminus\{a_n\}}\left|\int_\mathbb{Y}\left\|\overline{\delta}_{(x^k,a^k)_{k=1}^N}-{\overline{\psi}}_t\right\|_{BL}\left({\tilde{\mathfrak{p}}^n_{t,x^n,\overline{\delta}_{\boldsymbol{x}}}}(\mathop{\mathrm{d \!}}a^n) - \tilde{\mathfrak{p}}_t^n(x^n,\overline{\xi}_t)(\mathop{\mathrm{d \!}}a^n)\right) \right|\\ &\quad\le \frac{2}{N}\sum_{n=1}^N\left\|{\tilde{\mathfrak{p}}^n_{t,x^n,\overline{\delta}_{\boldsymbol{x}}}} - \tilde{\mathfrak{p}}_t^n(x^n,\overline{\xi}_t)\right\|_{BL} \le 2\vartheta({\boldsymbol{e}}_t({\boldsymbol{x}})),\quad{\boldsymbol{x}}\in\mathbb{X}^N, \end{align}\] where we have replaced the integrand by the left hand side of 70 to get the second inequality. In view of the estimation above and [prop:ProcConc], what is left to bound is \(\boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_t\bar e^t\); the procedure is similar to the proof [prop:ProcConc]. We continue to finish the proof for the sake of completeness. To this end we fix arbitrarily \(i\in\{1,\dots,N\}\) and consider \(\hat{\boldsymbol{x}}=(\hat{x}^1,\dots,\hat{x}^N)\in\mathbb{X}^N\) such that \(\hat{x}^k=x^k\) except the \(i\)-th entry. Note that \[\begin{align} \left|\bar {\boldsymbol{e}}_t(\hat{\boldsymbol{x}}) - \int_{\mathbb{A}^N}\left\|\overline{\delta}_{(x^k,a^k)_{k=1}^N}-{\overline{\psi}}_t\right\|_{BL} \left[\bigotimes_{k=1}^N\tilde{\mathfrak{p}}_t^k(\hat{x}^k,\overline{\xi}_t)\right](\mathop{\mathrm{d \!}}{\boldsymbol{a}})\right| \le \frac{1}{N}. \end{align}\] Then, by triangle inequality and a similar reasoning as above, we yield \[\begin{align} &\left|\bar {\boldsymbol{e}}_t({\boldsymbol{x}}) - \bar {\boldsymbol{e}}_t(\hat{\boldsymbol{x}})\right|\\ &\quad\le \frac{1}{N} + \sup_{(x^k,a^k),k\neq i}\left|\int_{\mathbb{A}^N}\left\|{\overline{\delta}_{(x^n,a^n)_{n=1}^N}}-{\overline{\psi}}_t\right\|_{BL} \left({\tilde{\mathfrak{p}}^i_{t,x^i,\overline{\xi}_t}}(\mathop{\mathrm{d \!}}a^i)-{\tilde{\mathfrak{p}}^i_{t,\hat{x}^i,\overline{\xi}_t}}(\mathop{\mathrm{d \!}}a^i)\right)\right| \le \frac{3}{2N}. \end{align}\] Recall the definitions of \(\overline{\boldsymbol{T}}\) and \(\check{\boldsymbol{\mathcal{T}}}\) from and below 66 . Since the above is true for any \(i\in\{1,\dots,N\}\), by 67 and 24, we have \[\begin{align} \label{eq:EstMeancTinfMeancT} \left|{\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s,t,{\overline{\Xi}}}\bar {\boldsymbol{e}}_t({\boldsymbol{x}}) - \inf_{x^i\in\mathbb{X}}{\check{\boldsymbol{\mathcal{T}}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{s,t,{\overline{\Xi}}}\bar {\boldsymbol{e}}_t({\boldsymbol{x}})\right| \le \frac{3}{2N},\quad {\boldsymbol{x}}\in\mathbb{X}^N,\;i=1,\dots,N. \end{align}\tag{71}\] Next, recall the definition of \(\boldsymbol{\mathbb{T}}\) from 26 and note that \[\begin{align} \boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{t}\bar {\boldsymbol{e}}_t \le \mathring{\boldsymbol{T}} \circ \overline{\boldsymbol{\mathcal{T}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{1,t,{\overline{\Xi}}}\bar {\boldsymbol{e}}_t + \sum_{r=1}^{t-1} \boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r} \left| \boldsymbol{T}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r+1} \circ \check{\boldsymbol{\mathcal{T}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r+1,t,{\overline{\Xi}}}\bar {\boldsymbol{e}}_t - \check{\boldsymbol{\mathcal{T}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r,t,{\overline{\Xi}}}\bar {\boldsymbol{e}}_t \right|. \end{align}\] Due to 5 and 15, for the first term in the right hand side above, we have \(\mathring{\boldsymbol{T}} \circ \overline{\boldsymbol{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{1,t,{\overline{\Xi}}}\bar {\boldsymbol{e}}_t \le \mathfrak{r}_{\mathfrak{K}\times\mathbb{A}}(N)\). We proceed to estimate the quantity with absolute sign in the second term, \[\begin{align} &\left| \boldsymbol{T}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r} \circ \check{\boldsymbol{\mathcal{T}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r+1,t,{\overline{\Xi}}}\bar {\boldsymbol{e}}_t ({\boldsymbol{x}}) - \check{\boldsymbol{\mathcal{T}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r,t,{\overline{\Xi}}}\bar {\boldsymbol{e}}_t ({\boldsymbol{x}}) \right|\\ &\quad = \left| \int_{\mathbb{X}^N} \check{\boldsymbol{\mathcal{T}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r+1,t,{\overline{\Xi}}}\bar {\boldsymbol{e}}_t({\boldsymbol{y}}) \left( \left[\bigotimes Q^{{\tilde{\mathfrak{p}}^n_{t,x^n,\overline{\delta}_{\boldsymbol{x}}}}}_{t,x^n,\overline{\delta}_{\boldsymbol{x}}}\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}}) - \left[\bigotimes Q^{{\tilde{\mathfrak{p}}^n_{t,x^n,\overline{\xi}_t}}}_{t,x^n,\overline{\xi}_t}\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}}) \right) \right|\\ &\quad\le \sum_{n=1}^N \sup_{\{y^1,\dots,y^N\}\setminus\{y^N\}} \left| \int_{\mathbb{X}^N} \check{\boldsymbol{\mathcal{T}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r+1,t,{\overline{\Xi}}}\bar {\boldsymbol{e}}_t({\boldsymbol{y}}) \left( Q^{{\tilde{\mathfrak{p}}^n_{t,x^n,\overline{\delta}_{\boldsymbol{x}}}}}_{t,x^n,\overline{\delta}_{\boldsymbol{x}}}(\mathop{\mathrm{d \!}}y^n) - Q^{{\tilde{\mathfrak{p}}^n_{t,x^n,\overline{\xi}_t}}}_{t,x^n,\overline{\xi}_t}(\mathop{\mathrm{d \!}}y^n) \right) \right| \end{align}\] where we have used 24 in the last inequality. It follows from 18 and 71 that, for \(L\ge 1+\eta(2L^{-1})\), \[\begin{align} \left| \boldsymbol{T}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r} \circ \check{\boldsymbol{\mathcal{T}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r+1,t,{\overline{\Xi}}}\bar {\boldsymbol{e}}_t ({\boldsymbol{x}}) - \check{\boldsymbol{\mathcal{T}}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r,t,{\overline{\Xi}}}\bar {\boldsymbol{e}}_t ({\boldsymbol{x}}) \right| \le 3 \left((L\vartheta+\eta)e_t({\boldsymbol{x}})+\eta(2L^{-1})\right). \end{align}\] Finally, by combining the above, we yield \[\begin{align} \boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{t}\bar {\boldsymbol{e}}_t \le \mathfrak{r}_{\mathfrak{K}\times\mathbb{A}}(N) + 3\sum_{r=1}^{t-1} \boldsymbol{\mathbb{T}}^{\widetilde{\boldsymbol{\mathfrak{P}}}}_{r} \circ (L\vartheta+\eta)e_{t-1} + 3(t-1)\eta(2L^{-1}). \end{align}\] Invoking [prop:ProcConc], the proof is complete. ◻

7.6 Proof of ?? in 3↩︎

Proof of ?? . Recall the definition of \({\boldsymbol{\mathfrak{R}}^n}\) and \(\overline{\mathfrak{R}}\) from 53 and 60 , respectively. In view of 3, below we always use \(\overline{\xi}_t\) and \(\overline{\Xi}\) instead of \(\xi^{\overline{\psi}_t}\) and \(\Xi^{\overline{\psi}_t}\). By [prop:EstDiffRMF], we have \[\begin{align} \label{eq:EstDiffRMF} \left|{\boldsymbol{\mathfrak{R}}^n}({\boldsymbol{\mathfrak{P}}}; {\boldsymbol{U}}) - \sum_{t=1}^{T-1} \bar{c}^t \overline{\mathbb{T}}^{\tilde{\mathfrak{p}}^n}_{t,{\overline{\Xi}}} \left( \overline{S}^{\tilde{\mathfrak{p}}^n}_{t,{\overline{\xi}_t}} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{{\overline{\Psi}}} - \overline{\mathcal{S}}^{*}_{t,T,{\overline{\Xi}}} V_{{\overline{\Psi}}} \right)\right| \le {\mathfrak{E}}. \end{align}\tag{72}\] Therefore, it is sufficient to show \[\begin{align} \label{eq:IneqSuff} \frac{1}{N}\sum_{n=1}^N\sum_{t=1}^{T-1} \bar{c}^t \overline{\mathbb{T}}^{\tilde{\mathfrak{p}}^n}_{t,{\overline{\Xi}}} \left( \overline{S}^{\tilde{\mathfrak{p}}^n}_{t,{\overline{\xi}_t}} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\overline{\Xi}}} V_{{\overline{\Psi}}} - \overline{\mathcal{S}}^{*}_{t,T,{\overline{\Xi}}} V_{{\overline{\Psi}}} \right) \ge \overline{\mathfrak{R}}(\overline{\Psi};V). \end{align}\tag{73}\] To this end notice that, by 4, \[\begin{align} \label{eq:IndAveMeanfTS} \frac{1}{N}\sum_{n=1}\overline{\mathbb{T}}^{\tilde{\mathfrak{p}}^n}_{t,{\overline{\Xi}}} \circ \overline{\mathcal{S}}^{*}_{t,T,{\overline{\Xi}}}V_{{\overline{\Psi}}} = \int_{\mathbb{X}} \overline{\mathcal{S}}^{*}_{t,T,{\overline{\Xi}}}V_{{\overline{\Psi}}}(y)\,\xi^{{\overline{\psi}}_t}(\mathop{\mathrm{d \!}}y) = \overline{\mathbb{T}}^{\overline{\mathfrak{p}}^{\overline{\Psi}}}_{t,\overline{\Xi}} \circ \overline{\mathcal{S}}^{*}_{t,T,\overline{\Xi}} V_{{\overline{\Psi}}} \end{align}\tag{74}\] Next, let \([N]:=\{1,\dots,N\}\) and consider the probability space \(([N]\times\mathbb{X},2^{[N]}\otimes\mathcal{B}(\mathbb{X}),\rho)\), where \(\rho\) satisfies \(\rho(\{n\}\times B) := \frac{1}{N} \overline{\mathbb{T}}^{\tilde{\mathfrak{p}}^n}_{t,\overline{\Xi}}\mathbb{1}_B\) for any \(B\in\mathcal{B}(\mathbb{X})\). We write \(\rho^n(B):=\rho(\{n\}\times B)\). Let \(q\) be the regular conditional probability of \(Z(n,x):=n\) given \(\{[N],\emptyset\}\otimes\mathcal{B}(\mathbb{X})\). Note \(q\) is constant in \(n\) and \(q_x\) can be treated as a \(N\)-dimensional probability vector. We subsequently write \(q^n(x):=q_x(\{n\})\). By partial averaging property17 and 4, for \(n\in [N]\), \(q^n\) satisfies \[\begin{align} \label{eq:IndPAP} \frac{1}{N} \overline{\mathbb{T}}^{\tilde{\mathfrak{p}}^n}_{t,\overline{\Xi}}h = \sum_{k=1}^N \int_{\mathbb{X}}h(y) q^n(y) \rho^k(\mathop{\mathrm{d \!}}y) = \int_{\mathbb{X}}h(y) q^n(y) \overline{\xi}_t(\mathop{\mathrm{d \!}}y),\; h\in B_b(\mathbb{X}). \end{align}\tag{75}\] For \(\tilde{\pi}^1,\dots,\tilde{\pi}^N\in\widetilde{\Pi}\), let \(\left[\sum_{n=1}^N q^n\tilde{\pi}^n\right]_{y,\xi}:=\sum_{n=1}^N q^n(y)\tilde{\pi}^n_{y,\xi}\) and note that it is \(\mathcal{B}(\mathbb{X}\times\mathcal{P}(\mathbb{X}))\)-\(\mathcal{E}(\mathcal{P}(\mathbb{A}))\) measurable. Recall the definition of \(\overline{S}^{\tilde{\boldsymbol{\pi}}}_{t,\xi}\) from 38 . It follows from 75 , 4 (vi) and generalized Jensen’s inequality [61] that \[\begin{align} \label{eq:EstAveMeanfTS} \frac{1}{N}\sum_{n=1}^N \overline{\mathbb{T}}^{\tilde{\mathfrak{p}}^n}_{t,\overline{\Xi}} \circ \overline{S}^{\tilde{\mathfrak{p}}^n_t}_{t,\overline{\xi}_t} \circ \overline{\mathcal{S}}^{*}_{t+1,T,\overline{\Xi}}V_{{\overline{\Psi}}} &= \sum_{n=1}^N\int_{\mathbb{X}} \overline{S}^{\tilde{\mathfrak{p}}^n_t}_{t,\overline{\xi}_t} \circ \overline{\mathcal{S}}^{*}_{t+1,T,\overline{\Xi}}V_{{\overline{\Psi}}}(y)\; q^n(y) \overline{\xi}_t(\mathop{\mathrm{d \!}}y)\nonumber\\ &\quad\ge \int_{\mathbb{X}} \overline{S}^{\sum_{n=1}^N q^n \tilde{\mathfrak{p}}^n_t}_{t,\xi^{{\overline{\psi}}_t}} \circ \overline{\mathcal{S}}^{*}_{t+1,T,\overline{\Xi}}V_{{\overline{\Psi}}}(y) \overline{\xi}_t(\mathop{\mathrm{d \!}}y). \end{align}\tag{76}\] On the other hand, note that for any \(B\in\mathcal{B}(\mathbb{X})\), \[\begin{align} &\int_{\mathbb{X}} \mathbb{1}_{B}(y) \int_{\mathbb{A}} \mathbb{1}_{A}(a) \left[\sum_{n=1}^N q^n(y)\tilde{\mathfrak{p}}^n_{t,y,\overline{\xi}_t}\right](\mathop{\mathrm{d \!}}a)\, \overline{\xi}_t(\mathop{\mathrm{d \!}}y)\\ &\quad= \sum_{n=1}^N \int_{\mathbb{X}} \mathbb{1}_{B}(y)\tilde{\mathfrak{p}}^n_{t,y,\overline{\xi}_t}(A)\; q^n(y) \overline{\xi}_t(\mathop{\mathrm{d \!}}y) = \frac{1}{N}\sum_{n=1}^N \int_{\mathbb{X}} \mathbb{1}_{B}(y)\tilde{\mathfrak{p}}^n_{t,y,\overline{\xi}_t}(A)\; \overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}^n}_{t,\overline{\Xi}}(\mathop{\mathrm{d \!}}y) = {\overline{\psi}}_t(B\times A), \end{align}\] where in the second equality we have used 75 and definition of \(\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}^n}_{t,\Xi}\) from below 29 , and the last equality follows from the definition of \(\overline{\Psi}\) in 47 . Therefore, by 2, \(y\mapsto\left[\sum_{n=1}^N q^n(y)\tilde{\mathfrak{p}}^n_{t,y,\overline{\xi}_t}\right]\) is a version of \(\overline{\pi}^{{\overline{\psi}}_t}\). This together with 76 , 7 and 3 implies that \[\begin{align} \label{eq:EstAveMeanfTSRe} \frac{1}{N}\sum_{n=1} \overline{\mathbb{T}}^{\tilde{\mathfrak{p}}^n}_{t,\overline{\Xi}} \circ \overline{S}^{\tilde{\mathfrak{p}}^n}_{t,\overline{\xi}_t} \circ \overline{\mathcal{S}}^{*}_{t+1,T,\overline{\Xi}}V_{{\overline{\Psi}}} \ge \overline{\mathbb{T}}^{\overline{\mathfrak{p}}^{\overline{\Psi}}}_{t,\overline{\Xi}} \circ \overline{S}^{\overline{\pi}^{\overline{\psi}_t}}_{t,\overline{\xi}_t} \circ \overline{\mathcal{S}}^{*}_{t+1,T,\overline{\Xi}} V_{{\overline{\Psi}}}. \end{align}\tag{77}\] Finally, by combining 74 and 77 , we have verified 73 . ◻

7.7 Proof of 4↩︎

Proof of 4. To start with note that \(\overline{\boldsymbol{\mathfrak{P}}}^\Psi\) is \(0\)-symmetrically continuous. Let \(\overline{\Xi}\) and \({\overline{\Psi}}\) be induced by \(\overline{\boldsymbol{\mathfrak{P}}}^\Psi\) as in 3.6 but with \(\widetilde{\boldsymbol{\mathfrak{P}}}=\overline{\boldsymbol{\mathfrak{P}}}^{\Psi}\).

We claim that \({\overline{\Psi}}=\Psi\). Recall the definition of \(\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}}_{t,\overline{\Xi}}\) from below 29 . By 3, we have \(\overline{\mathbb{Q}}^{{\overline{\mathfrak{p}}^\Psi}^n}_{t,{\Xi^{\Psi}}}=\xi^{\psi_{t}}\) for \(n=1,\dots,N\) and \(t=1,\dots,T-1\). It follows from 46 and induction that \(\overline{\xi}_{t}={\xi^{\psi_t}}\) for all \(n\) and \(t\). In addition, by 4, \(\xi^{\overline{\psi}_t}=\overline{\xi}_t=\xi^{\psi_t}\) and \({\overline{\Psi}}\) is a mean field flow. This together with the homogeneous policies and 47 implies \(\overline{\psi}_t(B\times A)=\psi_t(B\times A)\) for \(B\in\mathcal{B}(\mathbb{X})\) and \(A\in\mathcal{B}(\mathbb{A})\) for each \(t\). An application of monotone class lemma ([50]) shows \(\overline{\psi}_t=\psi_t\) for each \(t\), and thus \({\overline{\Psi}}=\Psi\).

Recall the definitions of \(\boldsymbol{\mathfrak{R}}^n\) and \(\overline{\mathfrak{R}}\) from 54 and 60 , respectively, invoking [prop:EstDiffRMF], we yield \[\begin{align} \left|{\boldsymbol{\mathfrak{R}}^n}(\overline{\boldsymbol{\mathfrak{P}}}^\Psi;{\boldsymbol{U}}) - \overline{\mathfrak{R}}(\Psi;V)\right| = \left|{\boldsymbol{\mathfrak{R}}^n}(\overline{\boldsymbol{\mathfrak{P}}}_\Psi;{\boldsymbol{U}}) - \overline{\mathfrak{R}}({\overline{\Psi}};V)\right| \le {\mathfrak{E}^0}, \end{align}\] where \(\mathfrak{E}^0\) is introduced in 3.9.

In order to finish the proof, by 34 , 37 , 40 and 43 , we have \[\begin{align} \left|\boldsymbol{\mathcal{R}^n}(\overline{\boldsymbol{\mathfrak{P}}}^\Psi;{\boldsymbol{U}}) - \overline{\mathcal{R}}(\Psi;V)\right| &\le \left|\mathring{\boldsymbol{G}} \circ\boldsymbol{\mathcal{S}}^{\overline{\boldsymbol{\mathfrak{P}}}^\Psi}_{1,T}{\boldsymbol{U}} - \mathring{\boldsymbol{G}} \circ \overline{\mathcal{S}}^{{\overline{\mathfrak{p}}^\Psi}}_{T,{\Xi^{\Psi}}} \right| + \left| \mathring{\boldsymbol{G}} \circ \boldsymbol{\mathcal{S}}^{*,\overline{\boldsymbol{\mathfrak{P}}}^\Psi}_{1,T}{\boldsymbol{U}} - \mathring{\boldsymbol{G}} \circ \overline{\mathcal{S}}^{*}_{T,{\Xi^{\Psi}}} \right|. \end{align}\] Due to what is proved above, we can replace \(\Xi^\Psi\) with \(\overline{\Xi}\). Invoking 4 (iii), 11 and [prop:EstfTDiffcSstar] (with \(\vartheta^1\equiv0\) and \(\vartheta\equiv 0\) thus \(\mathfrak{e}_t=\mathfrak{e}^0_t\)), the proof is complete. ◻

7.8 Proofs of 5↩︎

To prepare for the proof of 5, we introduce a few more notations. Note that \(\mathcal{P}(\mathbb{A})\) with weak topology is a Polish space because \(\mathbb{A}\) is (cf. [50]). It follows that \(\mathbb{X}\times\mathcal{P}(\mathbb{A})\) is also Polish. We let \(\mathcal{P}(\mathbb{X}\times\mathcal{P}(\mathbb{A}))\) be the set of probability measure on \(\mathcal{B}(\mathbb{X})\otimes\mathcal{B}(\mathcal{P}(\mathbb{A}))\) endowed with weak topology. Consider \(\mathfrak{m}=(\mu_1,\dots,\mu_{T-1})\in\mathcal{P}(\mathbb{X}\times\mathcal{P}(\mathbb{A}))^{T-1}\). We equip \(\mathcal{P}(\mathbb{X}\times\mathcal{P}(\mathbb{A}))^{T-1}\) with product topology.

For \(\mu\in\mathcal{P}(\mathbb{X}\times\mathcal{P}(\mathbb{A}))\), we let \({\psi^\mu}\in\mathcal{P}(\mathbb{X}\times\mathbb{A})\) satisfy \[\begin{align} \label{eq:Defpsifu} {\psi^\mu}(B\times A) = \int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})} \int_{\mathbb{A}} \mathbb{1}_{B}(y) \mathbb{1}_{A}(a) \lambda(\mathop{\mathrm{d \!}}a) \mu(\mathop{\mathrm{d \!}}y \mathop{\mathrm{d \!}}\lambda), \quad B\in\mathcal{B}(\mathbb{X}),\, A\in\mathcal{B}(\mathbb{A}), \end{align}\tag{78}\] and denote \({\Psi^\mathfrak{m}} = (\psi^{\mu_1},\dots,\psi^{\mu_{T-1}})\). Thanks to Carathéodory extension theorem (cf. [50]), it is valid to define a measure on \(\mathcal{B}(\mathbb{X}\times\mathbb{A})\) by imposing 78 . Using monotone class lemma [50], simple function approximation [50] and monotone convergence [50], for any \(f\in B(\mathbb{X}\times\mathbb{A})\) we have \[\begin{align} \label{eq:PsiUpsilon} \int_{\mathbb{X}\times\mathbb{A}}f(x,a){\psi^\mu}(\mathop{\mathrm{d \!}}x\mathop{\mathrm{d \!}}a) = \int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})} \int_{\mathbb{A}} f(x,a) \lambda(\mathop{\mathrm{d \!}}a) \mu(\mathop{\mathrm{d \!}}a \mathop{\mathrm{d \!}}\lambda). \end{align}\tag{79}\] It is well-known that \(\mathcal{P}(\mathbb{X}\times\mathcal{P}(\mathbb{A}))^{T-1}\) endowed with product of weak topologies is a locally convex topological vector space [50].

Recall the notations introduced in 8. In what follows, we let \(\check K_{t,i} := \check K_{\lceil\check c^t i\rceil}\) and \(\check\mathfrak{K}_t:=(\check K_{t,i})_{i\in\mathbb{N}}\). We let \({\mathbb{M}}_0\) be a subset of \(\mathcal{P}(\mathbb{X}\times\mathcal{P}(\mathbb{A}))^{T-1}\) such that for any \(\mathfrak{m}=(\mu_1,\dots,\mu_{T-1})\in{\mathbb{M}}_0\). Note that \(\xi^{\mu_1}=\mathring\xi\) and \(\xi^{\mu_t}\) is \(\check\mathfrak{K}_t\)-tight for \(t\ge 2\) by 15.18 In addition, under 2, due to the compactness of \(\mathcal{P}(\mathbb{A})\), we yield that \(\mu_t\) is \(\check\mathfrak{K}\times\{\mathcal{P}(\mathbb{A})\}\)-tight, for \(t=1,\dots,T-1\). It follows that \(\mathbb{M}_0\) is \((\check\mathfrak{K}_T\times\{\mathcal{P}(\mathbb{A})\})^{T-1}\)-uniformly tight, and thus pre-compact due to Prokhorov’s theorem [50]. Therefore, the completion of \(\mathbb{M}_0\), denoted by \({\mathbb{M}}\), is compact. We also note that \({\mathbb{M}}_0\) is convex and so is \({\mathbb{M}}\).

In what follows, we consider \(\mathfrak{m}=(\mu_1,\dots,\mu_{T-1})\) and \(\mathfrak{n}=(\nu_1,\dots,\nu_{T-1})\). For \(\mathfrak{m}\in{\mathbb{M}}\), we define \[\begin{align} \Gamma\mathfrak{m}&:= \bigg\{\mathfrak{n}\in{\mathbb{M}}:\; \xi^{\nu_{t+1}}(B) = \int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})}\int_{\mathbb{A}}P_{t,x,\xi^{\mu_t},a}(B)\lambda(\mathop{\mathrm{d \!}}a)\mu_t(\mathop{\mathrm{d \!}}x\mathop{\mathrm{d \!}}\lambda), B\in\mathcal{B}(\mathbb{X}),\tag{80}\\ &\qquad \int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})} \left(\overline{G}^{\lambda}_{t,\xi^{\mu_t}} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\Xi^\mathfrak{m}}} V_{\mathfrak{m}}(y) - \overline{\mathcal{S}}^{*}_{t,T,{\Xi^\mathfrak{m}}} V_{\mathfrak{m}}(y)\right) \nu_t(\mathop{\mathrm{d \!}}y \mathop{\mathrm{d \!}}\lambda) = 0 \bigg\}, \tag{81} \end{align}\] where we recall 58 and 79 , and define with a sight abuse of notation \[\begin{align} \label{eq:DefVfu} V_\mathfrak{m}(x) := V_{{\Psi^\mathfrak{m}}}(x) = V\left(x,\int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})} \int_{\mathbb{A}} P_{T-1,y,\xi_{\mu_{T-1}},a}(\cdot)\lambda(\mathop{\mathrm{d \!}}a)\; \mu_{T-1}(\mathop{\mathrm{d \!}}y\mathop{\mathrm{d \!}}\lambda)\right). \end{align}\tag{82}\] We are aware that \(\Psi^{\mathfrak{n}}\) need not to be a mean field flow even if \(\mathfrak{n}\in\Gamma\mathfrak{m}\).

The lemmas below will prove useful.

Lemma 12. Suppose 2, 4 (i) (ii) (iii), 9 and 10. Let \(V\in C_b(\mathbb{X}\times\mathcal{P}(\mathbb{X}))\). Then, \((x,\mu_{T-1})\mapsto V_\mathfrak{m}(x)\) is continuous. Moreover, for \(t=1,\dots,T-1\), both \[\begin{gather} (\xi,\lambda,\mu_{t+1},\dots,\mu_{T-1},x)\mapsto\overline{G}^{\lambda}_{t,\xi} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\Xi^\mathfrak{m}}} V_{\mathfrak{m}}(x)\\ (\mu_{t},\dots,\mu_{T-1},x)\mapsto \overline{\mathcal{S}}^{*}_{t,T,{\Xi^\mathfrak{m}}} V_{\mathfrak{m}}(x) \end{gather}\] are continuous.

Proof. In what follows, we consider \(\mathfrak{m}^k\) and \(\mathfrak{m}^0\) such that \(\mu^k_{t}\) converge weakly to \(\mu^0_{t}\) as \(k\to\infty\).19 Note that weak convergence of a sequence of joint probability measures implies weak convergence of the marginal probability measures. Therefore, \(\lim_{k\to\infty}\xi^{\mu^k_t}=\xi^{\mu^0_t}\).

We first show the continuity of \((x,\mu_{T-1})\mapsto V_\mathfrak{m}(x)\). Note that under 9, for any \(h\in C_b(\mathbb{X})\), \((y,\xi,a)\mapsto\int_{\mathbb{X}} h(z) P_{T-1,y,\xi,a}(\mathop{\mathrm{d \!}}z)\) is continuous. By 22, we yield the continuity of \((y,\xi,\lambda)\mapsto\int_{\mathbb{A}} \int_{\mathbb{X}} h(z) P_{T-1,y,\xi,a}(\mathop{\mathrm{d \!}}z) \lambda(\mathop{\mathrm{d \!}}a)\). By 22 again, we have \[\begin{align} &\lim_{k\to\infty}\int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})}\int_{\mathbb{A}} \int_{\mathbb{X}} h(z)P_{T-1,y,\xi^{\mu_{T-1}^k},a}(\mathop{\mathrm{d \!}}z)\; \lambda(\mathop{\mathrm{d \!}}a)\; \mu_{T-1}^k(\mathop{\mathrm{d \!}}y\mathop{\mathrm{d \!}}\lambda)\\ &\quad= \int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})}\int_{\mathbb{A}} \int_{\mathbb{X}} h(z)P_{T-1,y,\xi^{\mu_{T-1}^0},a}(\mathop{\mathrm{d \!}}z)\; \lambda(\mathop{\mathrm{d \!}}a)\; \mu_{T-1}^0(\mathop{\mathrm{d \!}}y\mathop{\mathrm{d \!}}\lambda). \end{align}\] This proves the continuity of \((x,\mu_{T-1})\mapsto V_\mathfrak{m}(x)\).

Next, we proceed by backward induction. Suppose for some \(t=1,\dots,T-1\) we have \((\mu_{t+1},\dots,\mu_{T-1},x) \mapsto \overline{\mathcal{S}}^{*}_{t+1,T,{\Xi^\mathfrak{m}}} V_{\mathfrak{m}}(x)\) is continuous, where we recall \(\overline{\mathcal{S}}^{*}_{T,T,{\Xi^\mathfrak{m}}}\) is the identity operator. Below we also consider \(x^k\) and \(\xi^k\) such that \(\lim_{k\to\infty} x^k=x^0\) and \(\lim_{k\to\infty} \xi^k=\xi^0\). Then, by 4 (iii), \[\begin{align} &\left| \overline{G}^{\lambda^0}_{t,\xi^0} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\Xi^{\mathfrak{m}^0}}} V_{\mathfrak{m}^0}(x^0) - \overline{G}^{\lambda^k}_{t,\xi^k} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\Xi^{\mathfrak{m}^k}}} V_{\mathfrak{m}^k}(x^k) \right|\\ &\quad\le \left| \overline{G}^{\lambda^0}_{t,\xi^0} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\Xi^{\mathfrak{m}^0}}} V_{\mathfrak{m}^0}(x^k) - \overline{G}^{\lambda^k}_{t,\xi^k} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\Xi^{\mathfrak{m}^0}}} V_{\mathfrak{m}^0}(x)(x^k) \right| \\ &\qquad+ \bar{c} \int_{\mathbb{X}} \left| \overline{\mathcal{S}}^{*}_{t+1,T,{\Xi^{\mathfrak{m}^0}}} V_{\mathfrak{m}^0}(y) - \overline{\mathcal{S}}^{*}_{t+1,T,{\Xi^{\mathfrak{m}^k}}} V_{\mathfrak{m}^k}(y) \right| Q^{\lambda^k}_{t,x^k,\xi^k}(\mathop{\mathrm{d \!}}y). \end{align}\] Due to the induction hypothesis, Assumption 10, 19 and 22, the right hand side above vanishes as \(k\to\infty\). We have shown the continuity of \((\xi,\lambda,\mu_{t+1},\dots,\mu_{T-1},x)\mapsto\overline{G}^{\lambda}_{t,\xi} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\Xi^\mathfrak{m}}} V_{\mathfrak{m}}(x)\).

Finally, by combining the above with 2 and [50], we obtain the continuity of \((\mu_{t},\dots,\mu_{T-1},x)\mapsto \overline{\mathcal{S}}^{*}_{t,T,{\Xi^\mathfrak{m}}} V_{\mathfrak{m}}(x)\). ◻

Suppose 2, 4 (i) (ii) (iii), 8, 9 and 10. Then, the set of fixed points of \(\Gamma\) is non-empty and compact.

Proof. We will apply Kakutani–Fan–Glicksberg fixed point theorem [50]. Below we fix \(\mathfrak{m}\in{\mathbb{M}}\) arbitrarily and verify the conditions of the theorem.

Note that \({\mathbb{M}}\) is compact by definition.

We then argue that \(\Gamma\mathfrak{m}\) is non-empty. Thanks to 8 and 15, the right hand side of 80 , as a probability on \(\mathcal{B}(\mathbb{X})\), is \(\check\mathfrak{k}_{t+1}\)-tight. In view of Lemma 12 and measurable maximum theorem (cf. [50]), let \(\overline{\pi}^*_t:\mathbb{X}\to\mathcal{P}(\mathbb{A})\) be an optimal selector that attains \(\inf_{\lambda\in\mathcal{P}(\mathbb{A})}\overline{G}^{\lambda}_{t,\xi^{\mu_t}} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\Xi^\mathfrak{m}}} V_{\mathfrak{m}}(y)\) for \(y\in\mathbb{X}\). In view of Carathéodory extension theorem [50], let \(\nu_t\in\mathcal{P}(\mathbb{X}\times\mathcal{P}(\mathbb{A}))\) be characterized by \(B\times\Lambda \mapsto \int_{\mathbb{X}} \mathbb{1}_{B}(x)\mathbb{1}_{\Lambda}(\overline{\pi}^*_t(x)) \xi_t(\mathop{\mathrm{d \!}}x)\) for \(B\in\mathcal{B}(\mathbb{X})\) and \(\Lambda\in\mathcal{B}(\mathcal{P}(\mathbb{A}))\). Consequently, this \(\nu_t\) belongs to \(\mathbb{M}\) and satisfies both 80 and 81 .

Next, we verify that \(\Gamma\mathfrak{m}\) is convex. To see this, let \(\gamma\in(0,1)\), and \(\mathfrak{n}^1,\mathfrak{n}^2 \in \Gamma\mathfrak{m}\). We consider a \(\mathfrak{n}\) satisfying \(\nu_t=\gamma\nu_t^1+(1-\gamma)\nu_t^2\) for \(t=1,\dots,T-1\). Clearly, \(\nu_t\) satisfies 80 as 80 specifies the state marginal uniquely. It follows from the linearity with respect to the integrating measure that \(\nu_t\) also satisfies 81 .

Finally, we show that \(\Gamma\) has a closed graph. To this end let \((\mathfrak{m}^k)_{k\in\mathbb{N}}\subseteq{\mathbb{M}}\), \(\mathfrak{n}^k\in\Gamma\mathfrak{m}^k\) for \(k\in\mathbb{N}\), and suppose \(\mathfrak{m}^k\) and \(\mathfrak{n}^k\) converges to \(\mathfrak{m}^0\) and \(\mathfrak{n}^0\), respectively, as \(k\to\infty\). For any \(h\in C_b(\mathbb{X})\), invoking 9 and 80 , then applying 22 twice, we yield \[\begin{align} \int_{\mathbb{X}}h(y)\xi^{\nu^{0}_{t+1}}\!(\mathop{\mathrm{d \!}}y) &= \lim_{k\to\infty}\int_{\mathbb{X}}h(y)\xi_{\nu^{k}_{t+1}}\!(\mathop{\mathrm{d \!}}y) = \lim_{n\to\infty} \int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})}\int_{\mathbb{A}} \int_{\mathbb{X}}h(y)P_{t,x,\xi^{\mu^k_t},a}(\mathop{\mathrm{d \!}}y)\lambda(\mathop{\mathrm{d \!}}a)\mu^k_t(\mathop{\mathrm{d \!}}x\mathop{\mathrm{d \!}}\lambda)\\ &= \int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})}\int_{\mathbb{A}} \int_{\mathbb{X}}h(y)P_{t,x,\xi^{\mu^0_t},a}(\mathop{\mathrm{d \!}}y)\lambda(\mathop{\mathrm{d \!}}a)\mu^0_t(\mathop{\mathrm{d \!}}x\mathop{\mathrm{d \!}}\lambda), \end{align}\] i.e., 80 is also true for \(\mathfrak{m}^0\) and \(\mathfrak{n}^0\). By combining 12 and 22, we have verified 81 for \(\mathfrak{m}^0\) and \(\mathfrak{n}^0\). The proof is complete. ◻

Let \(Y(x,\lambda):=\lambda\) for \((x,\lambda)\in\mathbb{X}\times\mathcal{P}(\mathbb{A})\) and \(Z^\mu\) be the Bochner conditional expectation of \(Y\) given \(\mathcal{B}(\mathbb{X})\otimes\{\emptyset,\mathcal{P}(\mathbb{A})\}\) under \(\mu\in\mathcal{P}(\mathbb{X}\times\mathcal{P}(\mathbb{A}))\). We refer to [62] for the validity and definition of Bochner conditional expectation. Clearly, \(Z^\mu\) is constant in \(\lambda\in\mathcal{P}(\mathbb{A})\). Recall the definition of induced action kernel from 3.5. The next lemma will be useful.

Lemma 13. We have \(\overline{\pi}^{{\psi^\mu}}=Z^\mu\) for \(\xi_{\mu}\) almost every \(x\in\mathbb{X}\).

Proof. In view of [62], it is sufficient to prove the partial averaging principle in Bochner’s sense, \[\begin{align} \int_{B\times\mathcal{P}(\mathbb{A})} Y(y,\lambda) \mu(\mathop{\mathrm{d \!}}y \mathop{\mathrm{d \!}}\lambda) = \int_{B\times\mathcal{P}(\mathbb{A})} \overline{\pi}^{{\psi^\mu}}_y \mu(\mathop{\mathrm{d \!}}y \mathop{\mathrm{d \!}}\lambda) = \int_{B} \overline{\pi}^{{\psi^\mu}}_y \xi^{{\psi^\mu}}(\mathop{\mathrm{d \!}}y),\quad B\in\mathcal{B}(\mathbb{X}), \end{align}\] where the second equality is due to the fact that \(y\mapsto\overline{\pi}^{{\psi^\mu}}_y\) is constant in \(\lambda\). The integrals above are indeed well-defined in Bochner’s sense due to [50]. To proceed, notice that, for any \(h\in B_b(\mathbb{A})\), \[\begin{align} &\int_{\mathbb{A}} h(a)\left[\int_{B} \overline{\pi}^{{\psi^\mu}}_y \xi^{{\psi^\mu}}(\mathop{\mathrm{d \!}}y)\right](\mathop{\mathrm{d \!}}a) = \int_{B} \int_{\mathbb{A}}h(a) \overline{\pi}^{{\psi^\mu}}_y (\mathop{\mathrm{d \!}}a) \xi_{{\psi^\mu}}(\mathop{\mathrm{d \!}}y) = \int_{B\times\mathbb{A}} h(a) {\psi^\mu}(\mathop{\mathrm{d \!}}y \mathop{\mathrm{d \!}}a) \\ &\quad = \int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})} \int_{\mathbb{A}} h(a) \lambda(\mathop{\mathrm{d \!}}a) \mu(\mathop{\mathrm{d \!}}y \mathop{\mathrm{d \!}}\lambda) = \int_{\mathbb{A}} h(a)\left[\int_{B\times\mathcal{P}(\mathbb{A})} Y(y,\lambda) \mu(\mathop{\mathrm{d \!}}y \mathop{\mathrm{d \!}}\lambda)\right](\mathop{\mathrm{d \!}}a), \end{align}\] where we have used [50] in the first and last equality, 79 in the third. Since \(h\in B_b(\mathbb{A})\) is arbitrary, the proof is complete. ◻

We are in the position to prove 5.

Proof of 5. In view of [prop:GammaFixedpt], we consider \(\mathfrak{m}^*\in{\mathbb{M}}\) such that \(\mathfrak{m}^*=\Gamma\mathfrak{m}^*\) and let \(\Psi^*=\Psi^{\mathfrak{m}^*}\) in the sense of 78 . We will show that \(\Psi^*\) is a MFF with zero exploitability.

To start with, note that, by definition, \(\xi^{\psi^*_t}=\xi^{\psi^{\mu^*_t}}=\xi^{\mu^*_t}\) for \(t=1,\dots,T-1\). Then, by 80 and 79 , we have \[\begin{align} \label{eq:ImpliesMFF} \xi^{\psi^*_{t+1}}(B) &= \xi^{\mu^*_{t+1}}(B) = \int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})}\int_{\mathbb{A}}P_{t,y,\xi^{\mu^*_t},a}(B)\lambda(\mathop{\mathrm{d \!}}a)\mu^*_t(\mathop{\mathrm{d \!}}y\mathop{\mathrm{d \!}}\lambda) = \int_{\mathbb{X}\times\mathbb{A}} P_{t,y,\xi^{\psi^*_t},a}(B) \psi^*_t(\mathop{\mathrm{d \!}}y\mathop{\mathrm{d \!}}a) \end{align}\tag{83}\] for any \(B\in\mathcal{B}(\mathbb{X})\), i.e., \(\Psi^*\) satisfies 44 and is a mean field flow.

Note that \(\xi^{\mu^*_t}=\xi^{\psi^*_t}\), \(V_{\mathfrak{m}^*}=V_{\Psi^*}\), and \(\Xi^{\mathfrak{m}^*}=\Xi^{\Psi^*}\) by definition. In view of 60 and 81 , it remains to show \[\begin{align} \label{eq:ImpliesOpt} \int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})} \overline{G}^{\lambda}_{t,\xi^{\psi^*_t}} \circ \overline{\mathcal{S}}^{*}_{t+1,T,\Xi^{\Psi^*}} V_{\Psi^*}(y) \mu^*_t(\mathop{\mathrm{d \!}}y \mathop{\mathrm{d \!}}\lambda) \ge \int_{\mathbb{X}} \overline{S}^{\overline{\pi}^{\psi^*_t}}_{t,\xi^{\psi^*_t}} \circ \overline{\mathcal{S}}^{*}_{t+1,T,\Xi^{\Psi^*}} V_{\Psi^*}(y) \xi^{\psi^*_t}(\mathop{\mathrm{d \!}}y). \end{align}\tag{84}\] for arbitrarily fixed \(t=1,\dots,T-1\). Let \(Y(x,\lambda):=\lambda\) for \((x,\lambda)\in\mathbb{X}\times\mathcal{P}(\mathbb{A})\) and \(Z^{\mu^*_t}\) be the Bochner conditional expectation of \(Y\) given \(\mathcal{B}(\mathbb{X})\otimes\{\emptyset,\mathcal{P}(\mathbb{A})\}\) under \(\mu^*_t\). In view of 4 (vi), by generalized Jensen’s inequality [61], \[\begin{align} \int_{\mathbb{X}\times\mathcal{P}(\mathbb{A})} \! \overline{G}^{\lambda}_{t,\xi^{\psi^*_t}} \!\! \circ\! \overline{\mathcal{S}}^{*}_{t+1,T,\Xi^{\Psi^*}} V_{\Psi^*}(y) \mu^*_t(\mathop{\mathrm{d \!}}y \mathop{\mathrm{d \!}}\lambda) \ge \int_{\mathbb{X}} \overline{G}^{Z^{\mu^*_t}(y)}_{t,\xi^{\psi^*_t}} \!\! \circ\! \overline{\mathcal{S}}^{*}_{t+1,T,\Xi^{\Psi^*}} V_{\Psi^*}(y) \xi^{\psi^*_t}(\mathop{\mathrm{d \!}}y). \end{align}\] Finally, we conclude the proof by observing that, by 13, \(Z^{\mu^*_t}=\overline{\pi}^{\psi^*_t}\) for \(\xi^{\mu^*_t}=\xi^{\psi^*_t}\) almost every \(x\in\mathbb{X}\). ◻

8 Examples↩︎

8.1 No-one-get-it game↩︎

Consider a deterministic \(N\)pG with \(T=3\) and \(\mathbb{X}=\mathbb{A}=\{0,1\}\). All players start from \(0\) at \(t=1\). At \(t=1,2\), by choosing \(a\in\mathbb{A}\), each player will transit to the corresponding state for the next epoch. If a player arrives at state \(1\), the player must stay in state \(1\) till the end of the game, i.e., state \(1\) is an absorption state. At \(t=2\), if a player reaches state \(1\), a cost of \(-1\) (i.e., a reward of \(1\)) incurs; \(0\) otherwise. At \(t=3\), each player is penalized by a cost of the size \(10\,\,\times\) the portion of players on the same level. All costs are constant in actions. Players are allowed to use randomized policy. The performance criteria is the expectation of sum of the costs.

Let us consider a scenario where all players use the following policy: \[\begin{align} \tilde{\mathfrak{p}}_1(0,\delta_0) = 0, \quad \tilde{\mathfrak{p}}_2(0,(1-p)\delta_0+p\delta_1) = \begin{cases}0.5\delta_0+0.5\delta_1,&p=0\\ \delta_1, &p\in(0,1]\end{cases}. \end{align}\] The above policy implies that every player chooses to stay on state \(0\) for \(t=2\). At \(t=2\), if no player is at state \(1\), all players randomly go to state \(0\) or \(1\) with even chances for \(t=3\). But if there is at least one player at state \(1\) at \(t=2\), the other players go to state \(1\) at \(t=3\), inducing massive penalization to all players.

In this \(N\)-player scenario, all players have \(0\) exploitabilities, i.e., it is an equilibrium. Note that the policy can be made continuous in \(p\) by interpolating over \(p=0\) and \(p=\frac{1}{n}\), but the modular of continuity has a steep slope of size \(\frac{N}{2}\) near \(0\). This scenario can not be well approximated in the sense of 2.1. Indeed, when lifted to MFG via 9 , we have \[\begin{align} \overline{\psi}_1=\delta_0\otimes\delta_0,\quad \overline{\psi}_2 = \delta_0 \otimes (0.5\delta_0 + 0.5\delta_1),\quad \overline{\xi}_3 = 0.5\delta_0 + 0.5\delta_1. \end{align}\] The population distribution concentrates on state \(0\) at \(t=1,2\), and distributes evenly on states \(0\) and \(1\) at \(t=3\) while the induced policy is to stay at state \(0\) at \(t=1,2\) and going randomly to either \(0\) or \(1\) with even chances. Consequently, the representative player has incentive to reach \(1\) at \(t=2\), as the population state distribution is unaffected in this mean field setting. The resulting mean field exploitability is \(1\), indicating a vacuous mean field approximation.

8.2 An illustrating example for 4 (vi)↩︎

In this example, we consider a single-period (\(T=2\)) MFG with \(\mathbb{X}=\mathbb{A}=\{0,1\}\). Suppose all players start from the same position \(x_1=0\). Upon realizing an action \(a\), a player will be transmitted to the corresponding state at \(t=2\), regardless of other players, i.e., \(P_{1,\boldsymbol{x}, a}=\delta_a\). Note that, under such setting, all (state-action) mean field flows take the form of \(\delta_0\otimes\lambda\), where \(\lambda=(p,1-p)\in\mathcal{P}(\mathbb{A})\). For convenience, we let \(p\in[0,1]\) represent the representative player’s policy and the associated mean field flow. The corresponding state distribution at \(t=2\) is then given by \((p,1-p)\).

Consider a MFG that imposes a congestion penalty at time \(t=2\). For \(i\in\mathbb{X}\) and \(\xi=(p,1-p)\in\mathcal{P}(\mathbb{X})\), we specify the terminal cost as \(V(i,\xi)=p\mathbb{1}_{\{0\}}(i) + (1-p)\mathbb{1}_{\{1\}}(i)\). Furthermore, we consider an entropic cost that violates 4 (vi): \(c_1(p)= - p\ln(p) - (1-p)\ln(1-p)\). Notably, this cost penalizes randomized actions at \(t=1\). Suppose all players aim to find \[\begin{align} \mathop{\mathrm{arg\,min}}_{p\in[0,1]}\mathbb{E}\left[C(p)+V(X_2,(p,1-p))\right], \end{align}\] where \(X_2\sim\textsf{Binomial}(1-p)\).

We claim that, in this MFG setting, there is no mean field equilibrium in terms of mean field flow. Indeed, for any \(p\in[0,1]\) \[\begin{align} \overline{\mathcal{R}}(p) = - p\ln(p) - (1-p)\ln(1-p) + p^2 + (1-p)^2 - \min\{p,1-p\} > 0, \end{align}\] because both the entropic cost \(C(p)\) and \(p^2 + (1-p)^2 - \min\{p,1-p\}\) are non-negative for \(p\in[0,1]\), and \(p^2 + (1-p)^2=\min\{p,1-p\}\) only at \(p=\frac{1}{2}\) while \(C(\frac{1}{2})=-\ln(\frac{1}{2})>0\). However, it is possible to construct \(N\)pE under a similar setting. For example, let \(N\) be even. Then, the \(N\)-player scenario where player-\(n\) employ action \((n\bmod2)\) deterministically is an equilibrium.

Lastly, to connect with the discussion following 5, we can heuristically construct an equilibrium situation in the mean field game setting, and relate this to the idea of state-action-kernel distribution. Consider a situation where half of the (infinitely many) players deterministically move to \(x_2=0\), while the remaining half deterministically move to \(x_2=1\). This situation is an equilibrium. Moreover, it admits a representation in \(\mathcal{P}(\mathbb{X}\otimes\mathcal{P}(\mathbb{A}))\). That is, \[\begin{align} \delta_0 \otimes \left( \frac{1}{2} \delta_{\delta_0} + \frac{1}{2} \delta_{\delta_1} \right). \end{align}\] Above, the \(\delta_0\) to the left of \(\otimes\) indicates the initial state marginal. The parenthesis to the right of \(\otimes\) provides a statistical summary of the action-kernel employed by players (located at \(x_1=0\)): half of the players use a deterministic policy symbolized by \(\delta_0\) (which entails a deterministic move to \(x_2 = 0\)), whereas the remaining half use a deterministic policy symbolized by \(\delta_1\).

8.3 Examples of scores operators↩︎

Consider a cost function \(C:\mathbb{X}\times\mathcal{P}(\mathbb{X})\times\mathbb{A}\to\mathbb{R}\) with \(\|C\|_\infty\le c_0\). Let \(J>0\) be an integer. Consider a probability vector \((w_1,\dots,w_J)\), and \(\kappa_j\in(0,1]\) for \(j=1,\dots,J\). We define \[\begin{align} \mathring{\boldsymbol{G}}{\boldsymbol{u}} &:= \sum_{j=1}^J w_j\inf_{q\in\mathbb{R}}\left\{ q+\kappa_j^{-1}\int_{\mathbb{X}^N} \left({\boldsymbol{u}}({\boldsymbol{y}})-q\right)_{+} \left[\bigotimes_{n=1}^N\mathring{\xi}\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}}) \right\} \tag{85}\\ \boldsymbol{G}^{{\boldsymbol{\lambda}}}_t{\boldsymbol{u}}({\boldsymbol{x}}) &:= \int_{\mathbb{A}} C(x,\overline{\delta}_{\boldsymbol{x}},a^1)\lambda^1(\mathop{\mathrm{d \!}}a^1),\nonumber\\ &\quad + \int_{\mathbb{A}^N}\!\!\left( \sum_{j=1}^J\!w_j\!\inf_{q\in\mathbb{R}}\left\{ {!}{ q+\kappa_j^{-1}\!\!\!\int_{\mathbb{X}^N} \left({\boldsymbol{u}}({\boldsymbol{y}})-q\right)_{+}\! \left[\bigotimes_{n=1}^N\! P_{t,x^n,\overline{\delta}_{\boldsymbol{x}},a^n}\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}}) } \right\} \right) \left[\bigotimes_{n=1}^N\lambda^n\right](\mathop{\mathrm{d \!}}{\boldsymbol{a}}\!)\tag{86}, \end{align}\] and \[\begin{align} \mathring{\overline{G}} v &:= \sum_{j=1}^J w_j \inf_{q\in\mathbb{R}}\left\{ q+\kappa_j^{-1}\int_{\mathbb{X}} \left(v(y)-q\right)_{+} \mathring{\xi}(\mathop{\mathrm{d \!}}y) \right\},\tag{87}\\ \overline{G}^{\lambda}_{t,\xi} v(x) &:= \int_{\mathbb{A}} c(x,\xi,a)\lambda(\mathop{\mathrm{d \!}}a) + \int_{\mathbb{A}}\left( \sum_{j=1}^J w_j \inf_{q\in\mathbb{R}}\left\{ q+\kappa_j^{-1}\int_{\mathbb{X}} \left(v(y)-q\right)_{+} P_{t,x,\xi,a}(\mathop{\mathrm{d \!}}y) \right\}\right) \lambda (\mathop{\mathrm{d \!}}a).\tag{88} \end{align}\] Note that \[\begin{align} {\boldsymbol{u}} \mapsto \inf_{q\in\mathbb{R}}\left\{ {!}{ q+\kappa^{-1}\int_{\mathbb{X}^N} \left({\boldsymbol{u}}({\boldsymbol{y}})-q\right)_{+} \left[\bigotimes_{n=1}^N P_{t,x^n,\overline{\delta}_{\boldsymbol{x}},a^n}\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}}) } \right\} \end{align}\] and \[\begin{align} v \mapsto \inf_{q\in\mathbb{R}}\left\{ q+\kappa^{-1}\int_{\mathbb{X}} \left(v(y)-q\right)_{+} P_{t,x,\xi,a}(\mathop{\mathrm{d \!}}y) \right\} \end{align}\] are the average value at risks (cf. [47]) of the random variables \({\boldsymbol{y}}\mapsto{\boldsymbol{u}}({\boldsymbol{y}})\) and \(y\mapsto v(y)\) under the probabilities \(\bigotimes_{n=1}^N P_{t,x^n,\overline{\delta}_{\boldsymbol{x}},a^n}\) and \(P_{t,x,\xi,a}\), respectively. Additionally, \(\inf_{q\in\mathbb{R}}\) in \(\boldsymbol{G}^{{\boldsymbol{\lambda}}}_t{\boldsymbol{u}}\) can be replaced by \(\inf_{q\in[-\|{\boldsymbol{u}}\|_\infty,\|{\boldsymbol{u}}\|_\infty]}\).The analogue holds true for \(\overline{G}^{\lambda}_{t,\xi} v\). Additionally, \(w_j\) and \(\kappa_j\) may depend on \((t,x,\xi)\), although this possibility is not considered here for simplicity.

For illustration, by setting \(J=1\) and \(\kappa_1=1\), we recover the case of risk neutral decision making. In particular, with definition 34 , we have \[\begin{align} \boldsymbol{\mathbb{S}}^{\boldsymbol{\boldsymbol{\mathfrak{P}}}}_{T}{\boldsymbol{u}} = \mathbb{E}\left[\sum_{t=1}^{T-1} C\left(X^1_t,\overline{\delta}_{(X^1_t,\dots,X^N_t)}, A^1_t\right) +{\boldsymbol{u}}\left(X^1_t,\dots,X^N_t\right)\right], \end{align}\] where under \(\mathbb{P}\), \[\begin{gather} \left(X^1_1,\dots,X^N_1\right)\sim\mathring{\xi}^{\otimes N}\,,\quad \left(A^1_t,\dots,A^N_t\right) \sim \bigotimes_{n=1}^N\mathfrak{p}^n_{t,(X^1_t,\dots,X^N_t)}\,,\quad \left(X^1_{t+1},\dots,X^N_{t+1}\right) \sim \bigotimes_{n=1}^N P_{t,X^n_t,\overline{\delta}_{(X^1_t,\dots,X^N_t)},A^n_t}\,. \end{gather}\] Similarly, with definition 40 , given \(\Xi=(\xi_1,\dots,\xi_{T-1})\in\mathcal{P}(\mathbb{X})^{T-1}\) and \(\xi_T\in\mathcal{P}(\mathbb{X})\), we have \[\begin{align} \overline{\mathbb{S}}^{\tilde{\mathfrak{p}}}_{T,\Xi} v = \overline{\mathbb{E}}\left[\sum_{t=1}^{T-1}C\left(X_t,\xi_{t}, A_t\right) + v(\xi_{T})\right], \end{align}\] where under \(\overline{\mathbb{P}}\), \[\begin{gather} X_1\sim{\xi}_1\,,\quad A_t\sim\tilde{\mathfrak{p}}_{t,X_t,\xi_t}\,,\quad X_{t+1}\sim P_{t,X_t,\xi_t,A_t}\,. \end{gather}\]

For the remainder of this example, we will provide some discussion on when the aforementioned score operators satisfies the assumptions in 3.7.

In view of 85 - 88 , it is straightforward to verify 4 (i) (ii) (iv) (v) (vi). Below we verify 4 (iii). It is sufficient to consider the simplified case where \(J=1\) and \(\kappa_1=\kappa\in(0,1]\). Observe that \[\begin{align} &\left|\boldsymbol{G}^{{\boldsymbol{\lambda}}}_t{\boldsymbol{u}}({\boldsymbol{x}}) - \boldsymbol{G}^{{\boldsymbol{\lambda}}}_t{\boldsymbol{u}}'({\boldsymbol{x}})\right|\\ &\quad= \left| \int_{\mathbb{A}^N}\inf_{q\in\mathbb{R}}\left\{ {!}{ q+\kappa^{-1}\int_{\mathbb{X}^N} \left({\boldsymbol{u}}({\boldsymbol{y}})-q\right)_{+} \left[\bigotimes_{n=1}^N P_{t,x^n,\overline{\delta}_{\boldsymbol{x}},a^n}\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}}) } \right\} \left[\bigotimes_{n=1}^N\lambda^n\right](\mathop{\mathrm{d \!}}{\boldsymbol{a}}) \right.\\ &\qquad - \left. \int_{\mathbb{A}^N} \inf_{q\in\mathbb{R}}\left\{ {!}{ q+\kappa^{-1}\int_{\mathbb{X}^N} \left({\boldsymbol{u}}'({\boldsymbol{y}})-q\right)_{+} \left[\bigotimes_{n=1}^N P_{t,x^n,\overline{\delta}_{\boldsymbol{x}},a^n}\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}}) } \right\} \left[\bigotimes_{n=1}^N\lambda^n\right](\mathop{\mathrm{d \!}}{\boldsymbol{a}}) \right|\\ &\quad\le \int_{\mathbb{A}^N} \inf_{q\in\mathbb{R}}\left\{ {!}{ q+\kappa^{-1}\int_{\mathbb{X}^N} \left( \left|{\boldsymbol{u}}({\boldsymbol{y}})-{\boldsymbol{u}}'({\boldsymbol{y}})\right|-q\right)_{+} \left[\bigotimes_{n=1}^N P_{t,x^n,\overline{\delta}_{\boldsymbol{x}},a^n}\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}}) } \right\} \left[\bigotimes_{n=1}^N\lambda^n\right](\mathop{\mathrm{d \!}}{\boldsymbol{a}}) \end{align}\] due to the convexity of average value at risk (cf. [47]). By taking \(q=0\) and invoking 23 , we continue to obtain \[\begin{align} &\left|\boldsymbol{G}^{{\boldsymbol{\lambda}}}_t{\boldsymbol{u}}({\boldsymbol{x}}) - \boldsymbol{G}^{{\boldsymbol{\lambda}}}_t{\boldsymbol{u}}'({\boldsymbol{x}})\right|\\ &\quad\le \kappa^{-1}\int_{\mathbb{A}^N} \int_{\mathbb{X}^N} \left|{\boldsymbol{u}}({\boldsymbol{y}})-{\boldsymbol{u}}'({\boldsymbol{y}})\right| \left[\bigotimes_{n=1}^N P_{t,x^n,\overline{\delta}_{\boldsymbol{x}},a^n}\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}}) \left[\bigotimes_{n=1}^N\lambda^n\right](\mathop{\mathrm{d \!}}{\boldsymbol{a}})\\ &\quad= \kappa^{-1} \int_{\mathbb{X}^N} \left|{\boldsymbol{u}}({\boldsymbol{y}})-{\boldsymbol{u}}'({\boldsymbol{y}})\right| \left[\bigotimes_{n=1}^N Q^{\lambda^n}_{t,x^n,\overline{\delta}_{\boldsymbol{x}}}\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}}). \end{align}\] Letting \(\bar{c}=\kappa^{-1}\), we have verified 4 (iii) for \(\boldsymbol{G}^{\boldsymbol{\lambda}}_t\). The assumption for \(\overline{G}\) can be verified with similar argument.

For 6 to hold, we additionally assume 5 and that, there is a subadditive modulus of continuity \(\zeta_0\) such that \[|C(x,\xi, a) - C(x,\xi, a')| \le \zeta_0(d_\mathbb{A}(a,a')),\quad (a,a')\in\mathbb{A}^2,\,(x,\xi)\in\mathbb{X}\times\mathcal{P}(\mathbb{X}).\] With similar reasoning leading to 18, for \(L>1+\eta(2L^{-1})\) we have \[\begin{align} \left|\int_\mathbb{A}c(x,\xi,a) \lambda^1(\mathop{\mathrm{d \!}}a) - \int_\mathbb{A}c(x,\xi,a) {\lambda'}^1(\mathop{\mathrm{d \!}}a)\right| \le 2c_0(L\|\lambda^1-{\lambda'}^1\|_{BL}+\zeta_0(2L^{-1})). \end{align}\] Next, by 5 and 24, \[\begin{align} a^1\mapsto q+\kappa^{-1}\int_{\mathbb{X}^N} \left({\boldsymbol{u}}({\boldsymbol{y}})-q\right)_{+} \left[\bigotimes_{n=1}^N P_{t,x^n,\overline{\delta}_{\boldsymbol{x}},a^n}\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}}) \end{align}\] has the modular of continuity of \(2\kappa^{-1}\|{\boldsymbol{u}}\|_\infty \eta(d_\mathbb{A}(a,a'))\). This modular holds regardless of \(q\in[-\|{\boldsymbol{u}}\|_\infty,\|{\boldsymbol{u}}\|_\infty]\), \({\boldsymbol{x}}\in\mathbb{X}^N\) and \(a^2,\dots,a^N\in\mathbb{A}\). Morevoer, this modulus preserves after acted by \(\inf_{q\in[-\|{\boldsymbol{u}}\|_\infty,\|{\boldsymbol{u}}\|_\infty]}\). Let \(\hat{\zeta}\) be a subadditive modular of continuity with \(\hat{\zeta}(\ell)\ge\max\{\zeta_0(\ell),\eta(\ell)\}\) for \(\ell\in\mathbb{R}_+\) (e.g., \(\hat{\zeta}=\zeta_0 + \eta\)). It follows from 23 and 18 that \[\begin{align} \left\|\boldsymbol{G}^{{\boldsymbol{\lambda}}}_t{\boldsymbol{u}} - \boldsymbol{G}^{{\boldsymbol{\lambda}}'}_t{\boldsymbol{u}}\right\|_\infty \le 2(c_0+2\kappa^{-1}\|{\boldsymbol{u}}\|_\infty)\inf_{L>1+\hat{\zeta}(2L^{-1})}\left(L\|\lambda^1-{\lambda'}^1\|_{BL} + \hat{\zeta}(2L^{-1})\right). \end{align}\] We define \(\zeta(\ell):=\inf_{L>1+\hat{\zeta}(2L^{-1})}\left\{L\ell + \hat{\zeta}(2L^{-1})\right\}\). Clearly, \(\zeta\) is non-decreasing. Moreover, for \(\ell\) small enough, letting \(L(\ell)=\ell^{-\frac{1}{2}}\) we yield \(\zeta(\ell)\le \ell^{\frac{1}{2}} + \zeta_0(2\ell^{\frac{1}{2}}) \xrightarrow[\ell\to0+]{}0\). As for the subadditivity, note that \(\hat{\zeta}\) is continuous (due to the concavity) and \(L\ell + \hat{\zeta}(2L^{-1})\) tends to infinity as \(L\to\infty\), we have \(\inf_{L>1+\hat{\zeta}(2L^{-1})}\left\{L\ell + \hat{\zeta}(2L^{-1})\right\}\) is attainable for \(\ell\in\mathbb{R}_+\). Let \(L_1^*\) and \(L_2^*\) be the points of minimum for \(\inf_{L>1+\hat{\zeta}(2L^{-1})}\left\{L\ell_1 + \hat{\zeta}(2L^{-1})\right\}\) and \(\inf_{L>1+\hat{\zeta}(2L^{-1})}\left\{L\ell_2 + \hat{\zeta}(2L^{-1})\right\}\), respectively. Without loss of generality we assume \(L_1^*\le L_2^*\). Thus, \[\begin{align} \zeta(\ell_1)+\zeta(\ell_2) \ge L_1^*(\ell_1+\ell_2) + \zeta_0(2L_1^{*-1}) \ge \zeta(\ell_1+\ell_2). \end{align}\] The above verifies 6 (i). 6 (ii) can be verified by similar reasoning with an additional application of triangle inequality.

Finally, to validate 10, we assume 9 and that \(C\) is jointly continuous. It follows immediately from 10 and 22 that \[\begin{align} (x,\xi,\lambda) \mapsto \int_{\mathbb{A}} C(x,\xi,a)\lambda(\mathop{\mathrm{d \!}}a) \end{align}\] is continuous. Next, notice that \(\inf_{q\in\mathbb{R}}\) in 88 can be replaced by \(\inf_{q\in[-\|v\|_\infty,\|v\|_\infty}\). This together with 9 and [50] implies that \[\begin{align} (x,\xi,a)\mapsto\inf_{q\in\mathbb{R}}\left\{ q+\kappa^{-1}\int_{\mathbb{X}} \left(v(y)-q\right)_{+} P_{t,x,\xi,a}(\mathop{\mathrm{d \!}}y) \right\} \end{align}\] is continuous. The rest follows from 22 again.

8.4 Ancillary example for 8↩︎

Let \(\mathbb{X}=\mathbb{R}\) and consider \(f:\mathbb{X}\times\mathcal{P}(\mathbb{X})\times\mathbb{A}\to[-1,1]\). Consider the transition kernels represented by a dummy random process \(X=(X_t)_{t=1,\dots,T}\) \[\begin{align} X_{t+1} = X_t + f(X_t,\xi,a) + Z_t,\quad t=1,\dots,T-1 \end{align}\] with \(X_1=0\), where \(Z_1,\dots,Z_{T-1}\stackrel{i.i.d.}{\sim}\mathcal{N}(0,1)\). Additionally, let \(\sigma(x) = x^2+1\) for \(x\in\mathbb{R}\) and \(\check c=4\). Clearly, ?? holds true. For ?? , we have \[\begin{align} \int_{\mathbb{X}}\sigma(y)P_{t,x,\xi,a}(\mathop{\mathrm{d \!}}y) &= \int_{\mathbb{R}} (y^2+1) \frac{1}{\sqrt{2\pi}} \exp\left(-\frac{(y+x+f(x,\xi,a))^2}{2}\right)\mathop{\mathrm{d \!}}y\\ &= \int_{\mathbb{R}} (y-(x+f(x,\xi,a)))^2\frac{1}{\sqrt{2\pi}} \exp\left(-\frac{(y)^2}{2}\right)\mathop{\mathrm{d \!}}y + 1 \\ &= (x+f(x,\xi,a))^2 + 2 \le x^2 + 2|x| + 3 \le \check c \sigma(x). \end{align}\]

8.5 Simplifying \(\overline{\mathfrak{R}}\)↩︎

This example compliments [rmk:MFPathExploitability]. Recall from 32 and 60 that \[\begin{align} \overline{\mathfrak{R}}(\Psi;V) = \sum_{t=1}^{T-1} \overline{C}^t \int_{\mathbb{X}} \left(\overline{G}^{\overline{\mathfrak{p}}_{\psi_t}(y)}_{t,{\xi^{\psi_t}}} \circ \overline{\mathcal{S}}^{*}_{t+1,T,{\Xi^{\Psi}}} V_{\Psi}(y) - \overline{\mathcal{S}}^{*}_{t,T,{\overline{\Xi}^{\Psi}}} V_{\Psi}(y)\right) {\xi^{\psi_t}}(\mathop{\mathrm{d \!}}y). \end{align}\] Let \(\overline{G}\) be defined in 88 . For convenience, we write \(V^*_{t+1}:=\overline{\mathcal{S}}^{*}_{t+1,T,{\Xi^{\Psi}}} V_{\Psi}\). Then, \[\begin{align} &\int_{\mathbb{X}} \overline{G}^{{\overline{\pi}^{\psi_t}}(y)}_{t,{\xi^{\psi_t}}} V^*_{t+1}(y) {\xi^{\psi_t}}(\mathop{\mathrm{d \!}}y)\\ &\quad= \int_{\mathbb{X}}\int_{\mathbb{A}} c(x,\xi,a)\left[{\overline{\pi}^{\psi_t}}(y)\right](\mathop{\mathrm{d \!}}a){\xi^{\psi_t}}(\mathop{\mathrm{d \!}}y)\\ &\qquad + \int_{\mathbb{X}}\int_{\mathbb{A}}\inf_{q\in\mathbb{R}}\left\{ q+\kappa^{-1}\int_{\mathbb{X}} \left(V^*_{t+1}(y)-q\right)_{+} \left[P_t(x,\xi,a)\right](\mathop{\mathrm{d \!}}y) \right\} \left[{\overline{\pi}^{\psi_t}}(y)\right](\mathop{\mathrm{d \!}}a) {\xi^{\psi_t}}(\mathop{\mathrm{d \!}}y)\\ &\quad= \int_{\mathbb{X}\times\mathbb{A}} c(x,\xi,a)\psi_t(\mathop{\mathrm{d \!}}y \mathop{\mathrm{d \!}}a)\\ &\qquad + \int_{\mathbb{X}\times\mathbb{A}}\inf_{q\in\mathbb{R}}\left\{ q+\kappa^{-1}\int_{\mathbb{X}} \left(V^*_{t+1}(y)-q\right)_{+} \left[P_t(x,\xi,a)\right](\mathop{\mathrm{d \!}}y) \right\} \psi_t(\mathop{\mathrm{d \!}}y \mathop{\mathrm{d \!}}a). \end{align}\]

9 Technical results↩︎

9.1 Consequences of various assumptions↩︎

The following dynamic programming principle is an immediate consequence of measurable maximal theorem (e.g., [50]) and the fact that \(\mathcal{B}(\mathcal{P}(\mathbb{A}))=\mathcal{E}(\mathcal{P}(\mathbb{A}))\) due to 21.

Lemma 14. Under 2, 4 (i) (ii) and 6, the following is true for any \(t=1,\dots,T-1\), \(\xi\in\mathcal{P}(\mathbb{X})\), \(N\in\mathbb{N}\), \(\boldsymbol{\pi}=(\pi^1,\dots,\pi^N)\in\Pi^N\), \({\boldsymbol{u}}\in B_b(\mathbb{X}^N)\) and \(v\in B_b(\mathbb{X})\):

  • \(\boldsymbol{S}^{*\boldsymbol{\pi}}_t{\boldsymbol{u}}\) is \(\mathcal{B}(\mathbb{X}^N)\)-\(\mathcal{B}(\mathbb{R})\) measurable and there is a \(\pi^{*}:(\mathbb{X}^N,\mathcal{B}(\mathbb{X}^N))\to(\mathcal{P}(\mathbb{A}),\mathcal{E}(\mathcal{P}(\mathbb{A})))\) such that \(\boldsymbol{S}^{(\pi^{*},\pi^2,\dots,\pi^N)}_t{\boldsymbol{u}} = \boldsymbol{S}^{*\boldsymbol{\pi}}_t{\boldsymbol{u}}\);

  • \(\overline{S}^*_{t,\xi}v\) is \(\mathcal{B}(\mathbb{X})\)-\(\mathcal{B}(\mathbb{R})\) measurable, and there is a \(\overline{\pi}^*:(\mathbb{X},\mathcal{B}(\mathbb{X}))\to(\mathcal{P}(\mathbb{A}),\mathcal{E}(\mathcal{P}(\mathbb{A})))\) such that \(\overline{S}^{\overline{\pi}^*}_{t,\xi} v = \overline{S}^*_{t,\xi} v\).

Alternatively, under 2, 4 (i) (ii) and 10, statement (b) holds true for \(v\in C_b(\mathbb{X})\).

Recall the definition of \(\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\) from below 29 . The next lemma pertains to the relationship between 3 and 8.

Lemma 15. 8 implies 3 with \(K_i=\check K_{\lceil\check c^T i\rceil},\, i\in\mathbb{N}\).

Proof. We first show that \(\overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\sigma \le \check c^t\) for \(t=1,\cdots,T\). For \(t=1\), this follows immediately from 29 and ?? . We proceed by induction. Suppose the statement is true for some \(t=1,\cdots,T-1\). Then, by 27 and ?? , \[\begin{align} \overline{T}^{\tilde{\mathfrak{p}}}_{t,\xi_t}\sigma(x) = \int_{\mathbb{A}}\int_{\mathbb{X}}\sigma(y)P_{t,x,\xi,a}(\mathop{\mathrm{d \!}}y)\tilde{\mathfrak{p}}_{t,x,\xi_t}(\mathop{\mathrm{d \!}}a) \le \check c\int_{\mathbb{A}}\sigma(x)\tilde{\mathfrak{p}}_{t,x,\xi_t}(\mathop{\mathrm{d \!}}a) = \check c\sigma(x). \end{align}\] Consequently, by 29 again, we obtain \[\begin{align} \overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{t+1,\Xi}\sigma = \overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\circ \overline{T}^{\tilde{\mathfrak{p}}}_{t,\xi_t}\sigma \le \check c\overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\sigma\le \check c^{t+1}. \end{align}\] The above proves \(\overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\sigma \le \check c^{t}\) for \(t=1,\cdots,T\). To finish the proof, we consider \(h(x)=i\mathbb{1}_{\check K_{\lceil\check c^T i\rceil}^c}(x)\). Note that \(h\le\check c^{-T}\sigma\) by the setting in 8. Due to the linearity and monotonicity of \(\overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\), we have \[\begin{align} i\; \overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\mathbb{1}_{\check K_{\lceil\check c^T i\rceil}^c} = \overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{t,\Xi} h \le \check c^{-T}\overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\sigma \le 1. \end{align}\] It follows that \(\overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{t,\Xi}\mathbb{1}_{\check K_{\lceil\check c^T i\rceil}^c} \le i^{-1}\). The proof is complete. ◻

Below are some important bounds.

Lemma 16. Suppose 4 (ii). For any \(\xi_1,\dots,\xi_{T-1}\in\mathcal{P}(\mathbb{X})\), \(\tilde{\mathfrak{p}}\in\widetilde{\Pi}\) and \(v\in B_b(\mathbb{X})\), we have \[\begin{gather} \left\|\overline{\mathcal{S}}^{\tilde{\mathfrak{p}}}_{s,t,(\xi_s,\dots,\xi_{t-1})} v\right\|_\infty \le \mathcal{C}_{t-s}(\|v\|_\infty),\quad 1\le s<t\le T,\\ \left|\overline{\mathbb{S}}^{\tilde{\mathfrak{p}}}_{t,(\xi_1,\dots,\xi_{t-1})} v\right| \le \mathcal{C}_{t}(\|v\|_\infty),\quad t=1,\dots,T, \end{gather}\] where \(\mathcal{C}_{r}(z) := c_0\sum_{k=1}^{r} c_1^{k-1} + c_1^{r} z\) for \(r\in\mathbb{N}_+\) and note \(\mathcal{C}_{r+1}(z) = c_0 + c_1 \mathcal{C}_{r}(z)\).

Proof. The statement can be established through a straightforward backward inductive argument. The details are therefore omitted. ◻

Lemma 17. Suppose 4 (ii) (iii). For any \({\boldsymbol{\mathfrak{P}}}\in\Pi^{N\times(T-1)}\), \({\boldsymbol{u}},{\boldsymbol{u}}'\in B_b(\mathbb{X}^N)\), and \({\boldsymbol{x}}\in\mathbb{X}^N\), we have \[\begin{gather} \left|\boldsymbol{\mathcal{S}}^{{\boldsymbol{\mathfrak{P}}}}_{s,t}{\boldsymbol{u}} - \boldsymbol{\mathcal{S}}^{{\boldsymbol{\mathfrak{P}}}}_{s,t}{\boldsymbol{u}}'\right| \le \bar{c}^{t-s} \boldsymbol{\mathcal{T}}^{{\boldsymbol{\mathfrak{P}}}}_{s,t} \left|u -{\boldsymbol{u}}'\right|, \quad 1\le s\le t\le T,\label{eq:EstDiffcS}\\ \left|\boldsymbol{\mathbb{S}}^{{\boldsymbol{\mathfrak{P}}}}_{t}{\boldsymbol{u}} - \boldsymbol{\mathbb{S}}^{{\boldsymbol{\mathfrak{P}}}}_{t}{\boldsymbol{u}}'\right| \le \bar{c}^{t} \boldsymbol{\mathbb{T}}^{{\boldsymbol{\mathfrak{P}}}}_{t} \left|u -{\boldsymbol{u}}'\right|, \quad t=1,\dots,T.\label{eq:EstDifffS} \end{gather}\] {#eq: sublabel=eq:eq:EstDiffcS,eq:eq:EstDifffS} For any \(\Xi=(\xi_1,\dots,\xi_{T-1})\in\mathcal{P}(\mathbb{X})^{T-1}\), \(\tilde{\mathfrak{p}}\in\boldsymbol{\widetilde{\Pi}}\), and \(v,v'\in B_b(\mathbb{X})\), we have \[\begin{gather} \left|\overline{\mathcal{S}}^{\tilde{\mathfrak{p}}}_{s,t,\Xi} v - \overline{\mathcal{S}}^{\tilde{\mathfrak{p}}}_{s,t,\Xi} v'\right| \le \bar{c}^{t-s}\overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{s,t,\Xi}|v-v'|,\quad 1\le s<t\le T,\\ \left|\overline{\mathbb{S}}^{\tilde{\mathfrak{p}}}_{t,\Xi} v - \overline{\mathbb{S}}^{\tilde{\mathfrak{p}}}_{t,\Xi} v'\right| \le \bar{c}^t\overline{\mathbb{T}}^{\tilde{\mathfrak{p}}}_{t,\Xi}|v-v'|,\quad t=1,\dots,T. \end{gather}\]

Proof. The statement is primarily a consequence of 4 (iii), while 4 (ii) is needed for the operators to be well-defined. For \(t=s\), \(\boldsymbol{\mathcal{S}}^{{\boldsymbol{\mathfrak{P}}}}_{s,s}\) and \(\boldsymbol{\mathcal{T}}^{{\boldsymbol{\mathfrak{P}}}}_{s,s}\) are the identity operators and thus \[\begin{align} \left|\boldsymbol{\mathcal{S}}^{{\boldsymbol{\mathfrak{P}}}}_{s,s}{\boldsymbol{u}}({\boldsymbol{x}}) - \boldsymbol{\mathcal{S}}^{{\boldsymbol{\mathfrak{P}}}}_{s,s}({\boldsymbol{x}}){\boldsymbol{u}}'({\boldsymbol{x}})\right| = \boldsymbol{\mathcal{T}}^{{\boldsymbol{\mathfrak{P}}}}_{s,s} \left|u -{\boldsymbol{u}}'\right| ({\boldsymbol{x}}). \end{align}\] We proceed by induction. Suppose ?? is true for some \(1\le s\le t<T\). Then, by 33 \[\begin{align} \left|\boldsymbol{\mathcal{S}}^{{\boldsymbol{\mathfrak{P}}}}_{s,t+1}{\boldsymbol{u}} - \boldsymbol{\mathcal{S}}^{{\boldsymbol{\mathfrak{P}}}}_{s,t+1}{\boldsymbol{u}}'\right| &= \left|\boldsymbol{\mathcal{S}}^{{\boldsymbol{\mathfrak{P}}}}_{s,t}\circ\boldsymbol{S}^{(\mathfrak{p}^n_t)_{n=1}^N}_{t}{\boldsymbol{u}} - \boldsymbol{\mathcal{S}}^{{\boldsymbol{\mathfrak{P}}}}_{s,t}\circ\boldsymbol{S}^{(\mathfrak{p}^n_t)_{n=1}^N}_{t}{\boldsymbol{u}}'\right| \\ &\le \bar{c}^{t-s} \boldsymbol{\mathcal{T}}^{{\boldsymbol{\mathfrak{P}}}}_{s,t} \left|\boldsymbol{S}^{(\mathfrak{p}^n_t)_{n=1}^N}_t{\boldsymbol{u}} - \boldsymbol{S}^{(\mathfrak{p}^n_t)_{n=1}^N}_t{\boldsymbol{u}}'\right| \le \bar{c}^{t+1-s} \boldsymbol{\mathcal{T}}^{{\boldsymbol{\mathfrak{P}}}}_{s,t+1} \left|{\boldsymbol{u}} -{\boldsymbol{u}}'\right|, \end{align}\] where we have used the induction hypothesis in the first inequality, 4 (iii) and 26 in the second inequality. This proves ?? . By combining ?? and 4 (iii), we yield ?? . A similar reasoning finishes the rest of the proof for score operators in mean field settings. ◻

The next two lemmas are consequence of the continuity assumed in 5.

Lemma 18. Under 5, for any \(h\in B_b(\mathbb{X})\), \(\xi,\xi'\in\mathcal{P}(\mathbb{X})\), \(\lambda,\lambda'\in\mathcal{P}(\mathbb{A})\) and \(L\ge 1+\eta(2L^{-1})\), we have \[\begin{align} \left|\int_{\mathbb{X}}h(y)Q_{t,x,\xi}^{\lambda}(\mathop{\mathrm{d \!}}y) - \int_{\mathbb{X}}h(y)Q_{t,x,\xi'}^{\lambda'}(\mathop{\mathrm{d \!}}y)\right| \le 2\|h\|_\infty\left( L\|\lambda-\lambda'\|_{BL} + \eta(2L^{-1}) + \eta(\|\xi-\xi'\|_{BL}) \right). \end{align}\]

Proof. Without loss of generality, we assume \(\|h\|_\infty=1\). Notice that \[\begin{align} &\left|\int_{\mathbb{X}}h(y)Q_{t,x,\xi}^{\lambda}(\mathop{\mathrm{d \!}}y) - \int_{\mathbb{X}}h(y)Q_{t,x,\xi'}^{\lambda'}(\mathop{\mathrm{d \!}}y)\right|\\ &\quad=\left|\int_{\mathbb{A}}\int_{\mathbb{X}}h(y)P_{t,x,\xi,a}(\mathop{\mathrm{d \!}}y)\lambda(\mathop{\mathrm{d \!}}a) - \int_{\mathbb{A}}\int_{\mathbb{X}}h(y)P_{t,x,\xi',a}(\mathop{\mathrm{d \!}}y)\lambda'(\mathop{\mathrm{d \!}}a)\right|\\ &\quad\le \left|\int_{\mathbb{A}}\int_{\mathbb{X}}h(y)P_{t,x,\xi,a}(\mathop{\mathrm{d \!}}y)\lambda(\mathop{\mathrm{d \!}}a) - \int_{\mathbb{A}}\int_{\mathbb{X}} h(y)P_{t,x,\xi,a}(\mathop{\mathrm{d \!}}y)\lambda'(\mathop{\mathrm{d \!}}a)\right|\\ &\qquad+ \left|\int_{\mathbb{A}}\int_{\mathbb{X}}h(y)P_{t,x,\xi,a}(\mathop{\mathrm{d \!}}y)\lambda'(\mathop{\mathrm{d \!}}a) - \int_{\mathbb{A}}\int_{\mathbb{X}} h(y)P_{t,x,\xi',a}(\mathop{\mathrm{d \!}}y)\lambda'(\mathop{\mathrm{d \!}}a)\right|\\ &\quad=: I_1 + I_2. \end{align}\] Regarding \(I_1\), we define \[H_{x,\xi}(a):=\int_{\mathbb{X}}h(y)P_{t,x,\xi,a}(\mathop{\mathrm{d \!}}y).\] By 5, we have \(\left|H_{x,\xi}(a)-H_{x,\xi}(a')\right| \le \eta(d_{\mathbb{A}}(a,a'))\), i.e., \(H_{x,\xi}\) has a modular of continuity being \(\eta\). For \(L\ge 1+\eta(2L^{-1})\), we let \(H^L_{x,\xi}\) be defined as in 23 with \(g=H_{x,\xi}\). Note that \(\|H^L_{x,\xi}\|_\infty\le 1+\eta(2L^{-1}) \le L\). Then, \[\begin{align} I_1 &\le \left|\int_{\mathbb{A}} H^L_{x,\xi}(a)\lambda(\mathop{\mathrm{d \!}}a) - \int_{\mathbb{A}} H^L_{x,\xi}(a)\lambda'(\mathop{\mathrm{d \!}}a)\right| + 2\eta\left(2L^{-1}\right) \\ &\le 2\|\lambda-\lambda'\|_{L-BL} + 2\eta\left(2L^{-1}\right) = 2L\|\lambda-\lambda'\|_{BL} + 2\eta\left(2L^{-1}\right). \end{align}\] Finally, regarding \(I_2\), by 5 again we yield \[\begin{align} I_2 &\le \sup_{x\in\mathbb{X}, a\in\mathbb{A}}\left|\int_{\mathbb{X}} h(y)P_{t,x,\xi,a}(\mathop{\mathrm{d \!}}y) - \int_{\mathbb{X}} h(y)P_{t,x,\xi',a}(\mathop{\mathrm{d \!}}y)\right| \le 2\eta(\|\xi-\xi'\|_{BL}). \end{align}\] The proof is complete. ◻

We recall the definition of \(Q^\lambda_{t,x,\xi}\) from 22 .

Lemma 19. Under 9, \((x,\xi,\lambda)\mapsto Q^{\lambda}_{t,x,\xi}\) is weakly continuous.

Proof. Due to simple function approximation (cf. [50]) and monotone convergence, we have \[\begin{align} \int_{\mathbb{X}}h(y)Q^{\lambda}_{t,x,\xi}(\mathop{\mathrm{d \!}}y) = \int_{\mathbb{A}}\int_{\mathbb{X}} h(y) P_{t,x,\xi,a}(\mathop{\mathrm{d \!}}y) \lambda(\mathop{\mathrm{d \!}}a), \quad h\in B_b(\mathbb{X}). \end{align}\] The rest follows from 22. ◻

9.2 Auxiliary technical results↩︎

Lemma 20. Let \((\mathbb{X},\mathscr{X})\) and \((\mathbb{Y},\mathscr{Y})\) be measurable spaces. Consider nonnegative \(f:(\mathbb{X}\times\mathbb{Y},\mathscr{X}\otimes\mathscr{Y})\to(\mathbb{R},\mathcal{B}(\mathbb{R}))\) and \(M:(\mathbb{X},\mathscr{X})\to(\mathcal{P},\mathcal{E}(\mathcal{P}))\), where \(\mathcal{P}\) is the set of probability measure on \(\mathscr{Y}\) and \(\mathcal{E}(\mathcal{P})\) is the corresponding evaluation \(\sigma\)-algebra, i.e., \(\mathcal{E}(\mathcal{P})\) is the \(\sigma\)-algebra generated by sets \(\{\zeta\in\mathcal{P}:\int_{\mathbb{Y}}f(y)\zeta(\mathop{\mathrm{d \!}}y) \in B\}\,\) for any real-valued bounded \(\mathscr{Y}\)-\(\mathcal{B}(\mathbb{R})\) measurable \(f\) and \(B\in\mathcal{B}(\mathbb{R})\). Then, \(x\mapsto\int_\mathbb{Y}f(x,y)\,{M_x}(\mathop{\mathrm{d \!}}y)\) is \(\mathscr{X}\)-\(\mathcal{B}(\mathbb{R})\) measurable.

Proof. We first consider \(f(x,y)=\mathbb{1}_D(x,y)\), where \(D\in\mathscr{X}\otimes\mathscr{Y}\). Let \(\mathscr{D}\) consist of \(D\in\mathscr{X}\otimes\mathscr{Y}\) such that \(x\mapsto\int_\mathbb{Y}f(x,y) {M_x}(\mathop{\mathrm{d \!}}y)\) is \(\mathscr{X}\)-\(\mathcal{B}(\mathbb{R})\) measurable. Note \(\{A\times B:A\in\mathscr{X},B\in\mathscr{Y}\}\subseteq\mathscr{D}\) because \(\int_\mathbb{Y}\mathbb{1}_{A\times B}(x,y) {M_x}(\mathop{\mathrm{d \!}}y) = \mathbb{1}_{A}(x) \, {M_x}(B)\) and \[\begin{align} \{x\in\mathbb{X}:{M_x}(B)\in C\} = \{x\in\mathbb{X}:{M_x}\in\{\zeta\in\mathcal{P}:\zeta(B)\in C\}\}\in\mathscr{X},\quad C\in\mathcal{B}([0,1]). \end{align}\] If \(D^1,D^2\in\mathscr{D}\) and \(D^1\subseteq D^2\), then \[\begin{align} \int_{\mathbb{Y}}\mathbb{1}_{D^2\setminus D^1}(x,y)\,{M_x}(\mathop{\mathrm{d \!}}y) = \int_{\mathbb{Y}}\mathbb{1}_{D^2}(x,y)\,{M_x}(\mathop{\mathrm{d \!}}y) - \int_{\mathbb{Y}}\mathbb{1}_{D^1}(x,y)\,{M_x}(\mathop{\mathrm{d \!}}y) \end{align}\] is also \(\mathscr{X}\)-\(\mathcal{B}(\mathbb{R})\) measurable. Similarly, if \(D^1,D^2\in\mathscr{D}\) are disjoint, then \(D^1\cup D^2\in\mathscr{D}\). The above shows that the algebra generated by \(\{A\times B:A\in\mathscr{X},B\in\mathscr{Y}\}\) is included by \(\mathscr{D}\). Now we consider an increasing sequence \((D^n)_{n\in\mathbb{N}}\subseteq\mathscr{D}\) and set \(D^0=\bigcup_{n\in\mathbb{N}}D^0\), then by monotone convergence [50], \[\begin{align} \lim_{n\to\infty}\int_{\mathbb{Y}}\mathbb{1}_{D^n}(x,y)\,{M_x}(\mathop{\mathrm{d \!}}y) = \int_{\mathbb{Y}}\mathbb{1}_{D^0}(x,y)\,{M_x}(\mathop{\mathrm{d \!}}y),\quad x\in\mathbb{X}. \end{align}\] It follows from [50] that \(D^0\in\mathscr{D}\). Invoking monotone class lemma [50], we yield \(\mathscr{D}=\mathscr{X}\otimes\mathscr{Y}\).

Now let \(f\) be any non-negative measurable function. Note that \(f\) can be approximated by a sequence of simple function \((f^n)_{n\in\mathbb{N}}\) such that \(f^n\uparrow f\) [50]. Since \(\mathscr{D}=\mathscr{X}\otimes\mathscr{Y}\), for \(n\in\mathbb{N}\), \(x\mapsto\int_\mathbb{Y}f^n(x,y) {M_x}(\mathop{\mathrm{d \!}}y)\) is \(\mathscr{X}\)-\(\mathcal{B}(\mathbb{R})\) measurable. Finally, by monotone convergence [50] and the fact that pointwise convergence preserves measurability [50], we conclude the proof. ◻

Let \(\mathbb{Y}\) be a topological space. Let \(\mathcal{B}(\mathbb{Y})\) be the Borel \(\sigma\)-algebra of \(\mathbb{Y}\), and \(\mathcal{P}\) be the set of probability measures on \(\mathcal{B}(\mathbb{Y})\). We endow \(\mathcal{P}\) with weak topology and let \(\mathcal{B}(\mathcal{P})\) be the corresponding Borel \(\sigma\)-algebra. Let \(\mathcal{E}(\mathcal{P})\) be the \(\sigma\)-algebra on \(\mathcal{P}\) generated by sets \(\{\zeta\in\mathcal{P}:\zeta(A)\in B\},\,A\in\mathcal{B}(\mathbb{Y}),\,B\in\mathcal{B}([0,1])\). Equivalently, \(\mathcal{E}(\mathcal{P})\) is the \(\sigma\)-algebra generated by sets \(\{\zeta\in\mathcal{P}:\int_{\mathbb{Y}}f(y)\zeta(\mathop{\mathrm{d \!}}y) \in B\}\,\) for any real-valued bounded \(\mathcal{B}(\mathbb{Y})\)-\(\mathcal{B}(\mathbb{R})\) measurable \(f\) and \(B\in\mathcal{B}(\mathbb{R})\). The lemma below regards the equivalence of \(\mathcal{B}(\mathcal{P})\) and \(\mathcal{E}(\mathcal{P})\).

Lemma 21. If \(\mathbb{Y}\) is a separable metric space, then \(\mathcal{B}(\mathcal{P})=\mathcal{E}(\mathcal{P})\).

Proof. Notice that the weak topology of \(\mathcal{P}\) is generated by sets \(\{\zeta\in\mathcal{P}:\int_{\mathbb{Y}}f(y)\zeta(\mathop{\mathrm{d \!}}y)\in U\}\) for any \(f\in C_b(\mathbb{Y})\) and open \(U\subseteq\mathbb{R}\). Therefore, \(\mathcal{B}(\mathcal{P})\subseteq\mathcal{E}(\mathcal{P})\). On the other hand, by [50], for any bounded real-valued \(\mathcal{B}(\mathbb{Y})\)-\(\mathcal{B}(\mathbb{R})\) measurable \(f\) and any \(B\in\mathcal{B}(\mathbb{R})\), we have \(\{\zeta\in\mathcal{P}:\int_{\mathbb{Y}}f(y)\zeta(\mathop{\mathrm{d \!}}y) \in B\}\in\mathcal{B}(\mathcal{P})\). The proof is complete. ◻

Lemma 22. Let \(\mathbb{Y}\) and \(\mathbb{Z}\) separable metric spaces. Let \(\mathbb{Y}\times\mathbb{Z}\) be endowed with product Borel \(\sigma\)-algebra \(\mathcal{B}(\mathbb{Y})\otimes\mathcal{B}(\mathbb{Z})\). Let \(f\in\ell^\infty(\mathbb{Y}\times\mathbb{Z},\mathcal{B}(\mathbb{Y})\otimes\mathcal{B}(\mathbb{Z}))\) be continuous. Let \((y^n)_{n\in\mathbb{N}}\subset\mathbb{Y}\) converges to \(y^0\in\mathbb{Y}\) and let \((\upsilon^n)_{n\in\mathbb{N}}\) be a sequence of probability measure on \(\mathcal{B}(\mathbb{Z})\) converging weakly to \(\upsilon^0\). Then, \[\begin{align} \lim_{n\to\infty}\int_\mathbb{Z}f(y^n,z)\, \upsilon^n(\mathop{\mathrm{d \!}}z) = \int_\mathbb{Z}f(y^0,z)\, \upsilon^0(\mathop{\mathrm{d \!}}z). \end{align}\]

Proof. To start with, note that \((\delta_{y^n})_{n\in\mathbb{N}}\) converges weakly to \(\delta_{y^0}\). Therefore, by [63], \((\delta_{y^n}\otimes\upsilon^n)_{n\to\mathbb{N}}\) converges to \(\delta_{y^0}\otimes\upsilon^0\). It follows that \[\begin{align} \lim_{n\to\infty}\int_\mathbb{Z}f(y^n,z)\, \upsilon^n(\mathop{\mathrm{d \!}}z) &= \lim_{n\to\infty}\int_{\mathbb{Y}\times\mathbb{Z}} f(y,z)\, \delta_{y^n}\otimes\upsilon^n(\mathop{\mathrm{d \!}}y\,\mathop{\mathrm{d \!}}z)\\ &= \int_{\mathbb{Y}\times\mathbb{Z}} f(y,z)\, \delta_{y^0}\otimes\upsilon^0(\mathop{\mathrm{d \!}}y\,\mathop{\mathrm{d \!}}z) = \int_\mathbb{Z}f(y^0,z)\, \upsilon^0(\mathop{\mathrm{d \!}}z). \end{align}\] The proof is complete. ◻

Lemma 23. Let \(\mathbb{Y}\) be a metric space. Suppose \(g:\mathbb{Y}\to\mathbb{R}\) is a continuous function satisfying \(|g(y)-g(y')| \le \beta(d(y,y'))\) for any \(y,y'\in\mathbb{Y}\) for some non-decreasing \(\beta:\overline{\mathbb{R}}_+\to\overline{\mathbb{R}}_+\). Then, for any \(L>0\), \(g_L(y):=\sup_{q\in\mathbb{A}}\{g(q) - L d(y,q)\}\) is a \(L\)-Lipschitz continuous function \(g_L\) satisfying \(\|g-g_L\|_\infty \le \beta(2L^{-1}\|g\|_\infty)\).

Proof. Because \[\begin{align} g_L(y) \ge \sup_{q\in\mathbb{A}}\{g(q) - L(d(y',q) + d_{\mathbb{A}}(y,y') )\} = g_L(y') - L d_{\mathbb{A}}(y,y'), \end{align}\] \(g_L\) is \(L\)-Lipschitz continuous \(L d_{\mathbb{A}}(y,y')\). Note \(g\le g_L\) dy definition. Finally, \[\begin{align} g(y) & = \sup_{q\in B_{2L^{-1}\|g\|_\infty}(y)}\left\{g(q) - (g(q)-g(y)) \right\} \ge \sup_{q\in B_{2L^{-1}\|g\|_\infty}(y)}\left\{g(q) - \beta\left(2L^{-1}\|g\|_\infty\right)\right\} \\ &\ge \sup_{q\in B_{2L^{-1}\|g\|_\infty}(y)}\left\{g(q) - Ld_\mathbb{A}(y,y')\right\} - \beta\left(2L^{-1}\|g\|_\infty\right) = g_L(y) - \beta\left(2L^{-1}\|g\|_\infty\right), \end{align}\] where we have used in the last equality the fact that \(q\) must stay within the \(2L^{-1}\|g\|_\infty\)-ball centered at \(y\) to be within \([-\|g\|_\infty,\|g\|_\infty]\). The proof is complete. ◻

Lemma 24. Let \((\mathbb{Y},\mathscr{Y})\) be a measurable space. Let \(m^1,\dots,m^N, {m^1}',\dots,{m^N}'\) be probability measures on \(\mathscr{Y}\). Then, for any \(\boldsymbol{u}\in\ell^\infty(\mathbb{Y}^N,\mathscr{Y})\), we have \[\begin{align} &\left|\int_{\mathbb{Y}^N} \boldsymbol{u}({\boldsymbol{y}}) \left[\bigotimes_{n=1}^N m^n\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}}) - \int_{\mathbb{Y}^N} f({\boldsymbol{y}}) \left[\bigotimes_{n=1}^N {m^n}'\right](\mathop{\mathrm{d \!}}{\boldsymbol{y}})\right|\\ &\quad\le \sum_{n=1}^N \sup_{\{y_1,\dots,y_N\}\setminus\{y_n\}}\left|\int_\mathbb{Y}\boldsymbol{u}(y_1,\dots,y_N)m^n(\mathop{\mathrm{d \!}}y_n) - \int_\mathbb{Y}f(y_1,\cdots,y_N){m^n}'(\mathop{\mathrm{d \!}}y_n)\right|. \end{align}\]

Proof. For \(N=2\), we have \[\begin{align} &\left|\int_{\mathbb{Y}^2} \boldsymbol{u}(y_1,y_2)m^1(\mathop{\mathrm{d \!}}y_1)m^2(\mathop{\mathrm{d \!}}y_2) - \int_{\mathbb{Y}^2} \boldsymbol{u}(y_1,y_2){m^1}'(\mathop{\mathrm{d \!}}y_1){m^2}'(\mathop{\mathrm{d \!}}y_2)\right|\\ &\quad \begin{multlined}[b] \le \left|\int_{\mathbb{Y}^2} \boldsymbol{u}(y_1,y_2)m^1(\mathop{\mathrm{d \!}}y_1)m^2(\mathop{\mathrm{d \!}}y_2) - \int_{\mathbb{Y}^2} \boldsymbol{u}(y_1,y_2)m^1(\mathop{\mathrm{d \!}}y_1){m^2}'(\mathop{\mathrm{d \!}}y_2)\right|\\ + \left|\int_{\mathbb{Y}^2} \boldsymbol{u}(y_1,y_2)m^1(\mathop{\mathrm{d \!}}y_1){m^2}'(\mathop{\mathrm{d \!}}y_2) - \int_{\mathbb{Y}^2} \boldsymbol{u}(y_1,y_2){m^1}'(\mathop{\mathrm{d \!}}y_1){m^2}'(\mathop{\mathrm{d \!}}y_2)\right| \end{multlined}\\ &\quad \begin{multlined}[b] = \left|\int_{\mathbb{Y}} \left(\int_{\mathbb{Y}} \boldsymbol{u}(y_1,y_2){m^2}(\mathop{\mathrm{d \!}}y_2) - \int_{\mathbb{Y}} \boldsymbol{u}(y_1,y_2){m^2}'(\mathop{\mathrm{d \!}}y_2)\right){m^1}(\mathop{\mathrm{d \!}}y_1) \right|\\ + \left|\int_{\mathbb{Y}} \left(\int_{\mathbb{Y}} \boldsymbol{u}(y_1,y_2)m^1(\mathop{\mathrm{d \!}}y_1) - \int_{\mathbb{Y}} \boldsymbol{u}(y_1,y_2){m^1}'(\mathop{\mathrm{d \!}}y_1)\right){m^2}'(\mathop{\mathrm{d \!}}y_2) \right| \end{multlined}\\ &\quad\le \sup_{y_1\in\mathbb{Y}}\left|\int_{\mathbb{Y}} \boldsymbol{u}(y_1,y_2)m^2(\mathop{\mathrm{d \!}}y_2) - \int_{\mathbb{Y}}\boldsymbol{u}(y_1,y_2){m^2}'(\mathop{\mathrm{d \!}}y_2)\right| + \sup_{y_2\in\mathbb{Y}}\left|\int_{\mathbb{Y}} \boldsymbol{u}(y_1,y_2)m^1(\mathop{\mathrm{d \!}}y_1) - \int_{\mathbb{Y}}\boldsymbol{u}(y_1,y_2){m^1}'(\mathop{\mathrm{d \!}}y_1)\right|. \end{align}\] The proof for \(N>2\) follows analogously by considering the corresponding telescoping sum. ◻

10 Supplementary to 3↩︎

10.1 Proof of 2↩︎

Proof of 2. Note that \(\mathcal{B}(\mathbb{A})=\sigma((A_n)_{n\in\mathbb{N}})\) for some \(A_n\subseteq\mathbb{A},\,n\in\mathbb{N}\), i.e., \(\mathcal{B}(\mathbb{A})\) is countably generated (cf. [48]). This together with 45 implies that \[\begin{align} \xi^\psi\left(\left\{x\in\mathbb{X}:\overline{\pi}^\psi_x(A_n)=\hat{\pi}_x(A_n),\; n\in\mathbb{N}\right\}\right) = 1. \end{align}\] Then, with probability \(1\), \(\overline{\pi}^\psi\) and \(\hat{\pi}\) coincide in the algrebra generated by \((A_n)_{n\in\mathbb{N}}\). By monotone class lemma ([50]), we have \(\overline{\pi}^\psi_x(A_n)=\hat{\pi}_x(A_n)\) for all \(n\in A_n\) implies \(\overline{\pi}^\psi_x(B)=\hat{\pi}_x(B)\) for all \(B\in\mathcal{B}(\mathbb{A})\). It follows that \[\begin{align} \left\{x\in\mathbb{X}:\overline{\pi}^\psi_x(A_n)=\hat{\pi}_x(A_n),\; n\in\mathbb{N}\right\} = \left\{x\in\mathbb{X}:\overline{\pi}^\psi_x=\hat{\pi}_x \right\}, \end{align}\] and thus \(\xi^\psi(\{x:\overline{\pi}^\psi_x=\hat{\pi}_x\})=1\). ◻

10.2 Proof of 3↩︎

Proof of 3. It is sufficient to consider \(h(x)=\mathbb{1}_B(x)\), where \(B\in\mathcal{B}(\mathbb{X})\). Note for \(t=1\), by 29 and 44 , we have \(\overline{\mathbb{T}}^{\overline{\mathfrak{p}}^\Psi}_{1,\emptyset}\mathbb{1}_B=\mathring\xi(B)=\xi_{1}(B)\). We proceed by induction. Suppose \(\overline{\mathbb{T}}^{\overline{\mathfrak{p}}^\Psi}_{t,\Xi^\Psi}\mathbb{1}_B=\xi^{\psi_t}(B)\) for \(B\in\mathcal{B}(\mathbb{X})\) for some \(t=1,\dots,T-1\). By 29 and the induction hypothesis implies that \[\begin{align} &\overline{\mathbb{T}}^{\overline{\mathfrak{p}}^\Psi}_{t+1,\Xi^\Psi}\mathbb{1}_B = \overline{\mathbb{T}}^{\overline{\mathfrak{p}}^\Psi}_{t,\Xi^\Psi} \circ \overline{T}^{\overline{\pi}^{\psi_t}}_{t,\xi^{\psi_t}} \mathbb{1}_B = \int_{\mathbb{X}} T^{\overline{\pi}^{\psi_t}}_{t,\xi^{\psi_t}} \mathbb{1}_B (y) \, \xi^{\psi_{t}}(\mathop{\mathrm{d \!}}y)\\ &\quad= \int_{\mathbb{X}} \int_{\mathbb{A}} P_{t,y,\xi^{\psi_t},a}(B) \overline{\pi}^{\psi_t}_y(\mathop{\mathrm{d \!}}a) \, \xi^{\psi_{t}}(\mathop{\mathrm{d \!}}y) = \int_{\mathbb{X}\times\mathbb{A}} P_{t,y,\xi^{\psi_t},a}(B) \, \psi_t(\mathop{\mathrm{d \!}}y\mathop{\mathrm{d \!}}a) = \xi^{\psi_{t+1}}(B), \end{align}\] where in the last equality we have used the fact that, for any \(h'\in\ell^\infty(\mathbb{X}\times\mathbb{A},\mathcal{B}(\mathbb{X}\times\mathbb{A}))\) and \(\psi\in\mathcal{P}(\mathbb{X})\), \[\begin{align} \int_{\mathbb{X}\times\mathbb{A}} h'(y,a) \psi(\mathop{\mathrm{d \!}}y\mathop{\mathrm{d \!}}a) = \int_{\mathbb{X}} \int_{\mathbb{A}} h'(y,a)\overline{\pi}^\psi_y(\mathop{\mathrm{d \!}}a)\, \xi^\psi(\mathop{\mathrm{d \!}}y) \end{align}\] due to a combination of 45 , monotone class lemma ([50]), simple function approximation (cf. [50]), and monotone convergence (cf. [50]). The proof is complete. ◻

10.3 Proof of 4↩︎

Proof of 4. The first statement that \(\xi^{\overline{\psi}_t}=\overline{\xi}_t\) is an immediate consequence of the definition of \(\overline{\mathbb{Q}}^{t,\tilde{\mathfrak{p}}}_\Xi\) below 29 . What is left to prove regards the mean field flow defined in 44 : \[\begin{align} \overline{\xi}_{t+1}(B) = \int_{\mathbb{X}\times\mathbb{A}} P_{t,y,\xi^{\overline{\psi}_t},a}(B)\; {\overline{\psi}}_t(\mathop{\mathrm{d \!}}y\mathop{\mathrm{d \!}}a),\quad B\in\mathcal{B}(\mathbb{X}). \end{align}\] Notice that by definition, \[\begin{align} \overline{\psi}_{t}(B\times A) = \frac{1}{N}\sum_{n=1}^N\int_{\mathbb{X}} \int_{\mathbb{A}} \mathbb{1}_B(y)\mathbb{1}_A(a) \tilde{\mathfrak{p}}^n_{t,y,\overline{\xi}_{t}}(\mathop{\mathrm{d \!}}a) \;\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}^n}_{t,\overline{\Xi}} (\mathop{\mathrm{d \!}}y). \end{align}\] Then, by monotone class lemma ([50]), simple function function approximation (cf. [50]) and monotone convergence (cf. [50]), for any \(B\in\mathcal{B}(\mathbb{X})\) we have \[\begin{align} &\int_{\mathbb{X}\times\mathbb{A}} \left[P_t(y,{\xi^{\psi_t}},a)\right](B) \;{\overline{\psi}}_t(\mathop{\mathrm{d \!}}y\mathop{\mathrm{d \!}}a)\\ &\quad = \frac{1}{N}\sum_{n=1}^N \int_{\mathbb{X}}\int_{\mathbb{A}} \left[P_t(y,\overline{\xi}_{t},a)\right](B) \;\left[\tilde{\mathfrak{p}}^n_t(y,\overline{\xi}_{t})\right](\mathop{\mathrm{d \!}}a)\;\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}^n}_{t,(\overline{\xi}_1,\cdots,\overline{\xi}_{t-1})} (\mathop{\mathrm{d \!}}y)\\ &\quad= \frac{1}{N}\sum_{n=1}^N \int_{\mathbb{X}}\overline{T}^{\tilde{\mathfrak{p}}^n_t}_{t,\overline{\xi}_t}\mathbb{1}_B(y)\;\overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}^n}_{t,(\overline{\xi}_1,\cdots,\overline{\xi}_{t-1})} (\mathop{\mathrm{d \!}}y) = \frac{1}{N}\sum_{n=1}^N \overline{\mathbb{T}}^{\tilde{\mathfrak{p}}^n}_{t,(\overline{\xi}_1,\cdots,\overline{\xi}_{t-1})} \circ \overline{T}^{\tilde{\mathfrak{p}}^n_t}_{t,\overline{\xi}_t}\mathbb{1}_B \\ &\quad= \frac{1}{N}\sum_{n=1}^N \overline{\mathbb{Q}}^{\tilde{\mathfrak{p}}^n}_{t+1,(\overline{\xi}_1,\cdots,\overline{\xi}_{t})}(B) = \overline{\xi}_{t+1}(B), \end{align}\] where we have also used 23 in the first equality, and 27 , 29 in the second last equality. This completes the proof. ◻

10.4 Details on error terms introduced in 3.9↩︎

We recall that \(\mathfrak{K}\) are introduced in 3, \(c_0,c_1,\bar{c}\) in 5 4, \(\eta\) in 5, \(\zeta\) in 6, \(\vartheta,\iota,V\) in 7, and \(\mathfrak{r}_\mathfrak{A}\) in 5. Additionally, we let \(\mathcal{C}_{r}(z) := c_0\sum_{k=1}^{r} c_1^{k-1} + c_1^{r} z\).

We let \(\mathfrak{e}_1(N):=\mathfrak{r}_{\mathfrak{K}}(N)\), and \[\begin{align} \mathfrak{e}_{t}(N)&:=2\inf_{L>1+\eta(2L^{-1})}\left\{\sum_{r=1}^{t-1}(L\vartheta(\mathfrak{e}_{r})+\eta(\mathfrak{e}_{r})) + \eta(2L^{-1}) \right\} + \frac{t}{N} + \mathfrak{r}_{\mathfrak{K}}(N),\quad t=2,\dots,T. \end{align}\] Additionally, we define \[\begin{align} {\breve\mathfrak{e}_t} := 2\vartheta({\mathfrak{e}_t}) + \frac{3}{2}{\mathfrak{e}_t} + \mathfrak{r}_{\mathfrak{k}\times\mathbb{A}}(N)-\mathfrak{r}_{\mathfrak{k}}(N), \end{align}\] where \(\mathfrak{k}\times\mathbb{A}=(K_i\times\mathbb{A})_{i\in\mathbb{N}}\). We also define \[\begin{align} \mathfrak{e}^0_1(N):=\mathfrak{e}_1(N),\quad\text{and}\quad \mathfrak{e}^0_{t}(N) &:=2\sum_{r=1}^{t-1}\eta(\mathfrak{e}^0_r(N)) + \frac{t}{N} + \mathfrak{r}_{\mathfrak{K}}(N),\; t=2,\dots,T. \end{align}\] Note that \(\mathfrak{e}^0_t=\mathfrak{e}_{t}\) if \(\vartheta\equiv 0\).

Based on \(\mathfrak{e}_t\) and \(\mathfrak{e}^0_t\), we further define \[\begin{align} \underline{\mathfrak{E}}(N) := \sum_{r=1}^{T-1} \bar{c}^{r}\mathcal{C}_{T-r}(\|V\|_\infty) \big(3\zeta(\vartheta(\mathfrak{e}_{r})) + (3+T)\zeta(\mathfrak{e}_{r}) \big) + (3+T)\bar{c}^{T}\iota({\mathfrak{e}_T}),\nonumber\\{\mathfrak{E}}(N) &:= (1+\bar{c}) \sum_{t=1}^{T-1} (T+1-t)\left( \sum_{r=t}^{T-1} \bar{c}^{r-t} \mathcal{C}_{r,T}(\|V\|_\infty) \zeta({\mathfrak{e}_t}) + \bar{c}^{T-t}\iota({\mathfrak{e}_T}) \right)\nonumber\\ &\quad+ \bar{c} \iota({\mathfrak{e}_T}) + \sum_{t=1}^{T-1} \mathcal{C}_{T-t}(\|V\|_\infty)\big(\zeta(\vartheta({\mathfrak{e}_t}))+\zeta({\mathfrak{e}_t}) + \mathfrak{e}_{t}(N)\big),\nonumber \\ \begin{align} {\mathfrak{E}^0}(N) &:= (1+\bar{c}) \sum_{t=1}^{T-1} (T+1-t)\left( \sum_{r=t}^{T-1} \bar{c}^{r-t} \mathcal{C}_{r,T}(\|V\|_\infty) \zeta(\mathfrak{e}^0_t(N)) + \bar{c}^{T-t}\iota(\mathfrak{e}^0_T(N)) \right)\nonumber\\ &\quad+ \bar{c} \iota(\mathfrak{e}^0_T(N)) + \sum_{t=1}^{T-1} \mathcal{C}_{T-t}(\|V\|_\infty)\big(\zeta(\mathfrak{e}^0_t(N)) + \mathfrak{e}^0_{t}(N)\big),\nonumber \end{align}\\ {\mathfrak{E}^\diamond}(N) := (T+1)\left(\sum_{r=1}^{T-1} \bar{c}^{r} \mathcal{C}_{T-r}(\|V\|_\infty) \zeta(\mathfrak{e}^0_r(N)) + \bar{c}^{T-t}\iota(\mathfrak{e}^0_T(N))\right).\nonumber \end{align}\] Above, we note that \(\mathfrak{E}^0=\mathfrak{E}\) if \(\vartheta\equiv 0\).

Proof of 6. We will only prove \(\lim_{N\to\infty}\mathfrak{e}_{t}(N)=0\) for \(t=1,\dots,T\), as the rest of the proof follows automatically. In view of 5, the claim is obviously true for \(t=1\). We will proceed by induction. Suppose for some \(t\ge 1\) we have \(\lim_{N\to\infty}\mathfrak{e}_{t}(N)=0\). For sufficiently large \(N\), we let \[L(N):=\left(\max_{r=1,\dots,t}\{\vartheta(\mathfrak{e}_{r})\}\right)^{-\frac{1}{2}}.\] Then, by 5 and the induction hypothesis, we have \[\begin{align} \mathfrak{e}_{t+1}(N) \le \frac{2t}{L(N)} + 2\sum_{r=1}^t \eta(\mathfrak{e}_{r}) + t\left(\frac{1}{N}+2\eta\left(\frac{2}{L(N)}\right)\right) + \mathfrak{r}_{\mathfrak{K}}(N) \xrightarrow[]{N\to\infty}0. \end{align}\] ◻

10.5 Proof of 5↩︎

Proof of 5. To start with, we point out a useful observation \[\begin{align} \label{eq:EstEMeandelta} \mathbb{E}\left[\overline{\delta}_{\boldsymbol{Y}}(A_i^c)\right]=\frac{1}{N}\sum_{n=1}^N\mathbb{E}\left[\mathbb{1}_{A_i^c}(Y^n)\right]= \frac{1}{N}\sum_{n=1}^N\upsilon^n(A_i^c)\le i^{-1}. \end{align}\tag{89}\] Next, by the uniform tightness of \((\upsilon^n)_{n\in\mathbb{N}}\), we have \[\begin{align} \label{eq:deltaintensityBLUpperBound} \left\|\overline{\delta}_{\boldsymbol{Y}}-\overline{\upsilon}\right\|_{BL} &\le \sup_{h\in C_{BL}(\mathbb{Y})} \frac{1}{2}\left|\int_{A_i}h(y)\overline{\delta}_{\boldsymbol{Y}}(\mathop{\mathrm{d \!}}y) - \int_{A_i}h(y)\overline{\upsilon}(\mathop{\mathrm{d \!}}y) \right| + \frac{1}{2}\left(\overline{\delta}_{\boldsymbol{Y}}(A_i^c) + i^{-1}\right). \end{align}\tag{90}\] Regarding the first term in the right hand side of 90 , we observe \[\begin{align} \label{eq:supCYsupCK} \sup_{h\in C_{BL}(\mathbb{Y})}\left|\int_{A_i}h(y) \overline{\delta}_{\boldsymbol{Y}}(\mathop{\mathrm{d \!}}y) - \int_{A_i}h(y)\overline{\upsilon}(\mathop{\mathrm{d \!}}y) \right| &\quad\le \sup_{h\in C_{BL}(A_i)}\left|\int_{A_i}h(y)\overline{\delta}_{\boldsymbol{Y}}(\mathop{\mathrm{d \!}}y) - \int_{A_i}h(y)\overline{\upsilon}(\mathop{\mathrm{d \!}}y) \right|\nonumber\\ &\quad\le \max_{h\in\mathfrak{c}_{j}(A_i)}\left|\int_{A_i}h(y)\overline{\delta}_{\boldsymbol{Y}}(\mathop{\mathrm{d \!}}y) - \int_{A_i}h(y)\overline{\upsilon}(\mathop{\mathrm{d \!}}y) \right| + j^{-1}, \end{align}\tag{91}\] where \(\mathfrak{c}_{j}(A_i)\) is a finite set with cardinality of \(\mathfrak{N}_{j}(A_i)\) such that \(C_{BL}(A_i)\subseteq\bigcup_{h\in\mathfrak{c}_{j}(A_i)}B_{j^{-1}}(h)\). By 90 , 91 and 89 , we have \[\begin{align} \label{eq:EstExpectedBL} &\mathbb{E}\left[\left\|\overline{\delta}(\boldsymbol{Y})-\overline{\upsilon}\right\|_{BL}\right]\nonumber\\ &\quad\le \mathbb{E}\left[\sup_{h\in C_{BL}(A_i)}\frac{1}{2}\left|\int_{A_i}h(y)\delta_{\boldsymbol{Y}}(\mathop{\mathrm{d \!}}y) - \int_{A_i}h(y)\overline{\upsilon}(\mathop{\mathrm{d \!}}y) \right|\right] + \frac{1}{2}\left(\mathbb{E}\left( \delta_{\boldsymbol{Y}}(A_i^c) \right) + i^{-1}\right)\nonumber\\ &\quad\le \mathbb{E}\left(\max_{h\in\mathfrak{c}_{j}(A_i)}\frac{1}{2}\left|\int_{A_i}h(y)\delta_{\boldsymbol{Y}}(\mathop{\mathrm{d \!}}y) - \int_{A_i}h(y)\overline{\upsilon}(\mathop{\mathrm{d \!}}y)\right|\right) + \frac{1}{2}j^{-1} + i^{-1}, \end{align}\tag{92}\] In order to proceed, for \(\varepsilon>0\) and \(h\in\mathfrak{c}_{j}(A_i)\), we invoke Heoffding’s inequality to yield \[\begin{align} &\mathbb{P}\left[\frac{1}{2}\left|\int_{A_i}h(y)\delta_{\boldsymbol{Y}}(\mathop{\mathrm{d \!}}y) - \int_{A_i}h(y)\overline{\upsilon}(\mathop{\mathrm{d \!}}y) \right| \ge \varepsilon\right] \\ &\quad= \mathbb{P}\left[\frac{1}{2}\left|\sum_{n=1}^N\mathbb{1}_{A_i}(Y_n)h(Y_n) - \sum_{n=1}^N\mathbb{E}\left[\mathbb{1}_{A_i}(Y_n)h(Y_n)\right] \right| \ge N\varepsilon \right] \le 2\exp\left(-2N\varepsilon^2\right). \end{align}\] Thus, \[\begin{align} \mathbb{P}\left[\max_{h\in\mathfrak{c}_{j}(A_i)}\frac{1}{2}\left|\int_{K_i}h(y)\delta_{\boldsymbol{Y}}(\mathop{\mathrm{d \!}}y) - \int_{A_i}h(y)\overline{\upsilon}(\mathop{\mathrm{d \!}}y)\right| \ge \varepsilon\right] \le 2 \mathfrak{N}_{j}(A_i) \exp\left(-2N\varepsilon^2\right). \end{align}\] It follows from the fact that \(\mathbb{E}[Z]=\int_{\mathbb{R}_+}(1-F_{Z}(r))\mathop{\mathrm{d \!}}r\) for any non-negative random variable \(Z\) that \[\begin{align} \mathbb{E}\left[\max_{h\in\mathfrak{c}_{j}(A_i)}\frac{1}{2}\left|\int_{K_i}h(y)\delta_{\boldsymbol{Y}}(\mathop{\mathrm{d \!}}y) - \int_{A_i}h(y)\overline{\upsilon}(\mathop{\mathrm{d \!}}y)\right|\right] \le 2\mathfrak{N}_{j}(A_i) \int_{\mathbb{R}_+} \exp(-2Nr^2) \mathop{\mathrm{d \!}}r = \frac{\sqrt{\pi}\;\mathfrak{N}_{j}(A_i)}{\sqrt{2 N}}. \end{align}\] This together with 92 proves ?? .

Finally, in order to show \(\lim_{N\to\infty}\mathfrak{r}_\mathfrak{A}(N)=0\), it is sufficient to construct non-decreasing \(i(N),j(N)\) such that \[\begin{align} \hat{\mathfrak{r}}_\mathfrak{A}(N) := \frac{1}{2}j(N)^{-1} + i(N)^{-1} + \frac{\sqrt{\pi}\;\mathfrak{N}_{j(N)}(A_{i(N)})}{\sqrt{2 N}} \xrightarrow[N\to\infty]{} 0. \end{align}\] We set \(i(1)=j(1)=1\) and construct \(i(N),j(N)\) in a way such that they increase simultaneously with increments of size \(1\). We define \(N_0:=1\) and \(N_\ell:=\min\{N>N_{\ell-1}:i(N)\neq i(N-1)\}\) for \(\ell\in\mathbb{N}\), i.e., \(N_\ell\) is the \(N\) where \(i(N),j(N)\) increase for the \(\ell\)-th time. Notice that \(\frac{\mathfrak{N}_{j}(A_{i})}{2\sqrt{2\pi N}}\) can be arbitrarily small as long as \(i,j\) are fixed and \(N\) is large enough, we construct \(i(N),j(N)\) such that \[\begin{align} N_{\ell} = \min\left\{N>N_{\ell-1}: \frac{\sqrt{\pi}\;\mathfrak{N}_{j(N)}(A_{i(N)})}{\sqrt{2N}} \le \frac{1}{\ell} \right\}. \end{align}\] Therefore, \(\hat{\mathfrak{r}}_\mathfrak{A}(N)\) is decreasing in \(N=N_{\ell},\dots,N_{\ell+1}-1\) and \(\hat{\mathfrak{r}}_\mathfrak{A}(N_\ell)\le \frac{3}{2}\ell^{-1} + \ell^{-1} = \frac{5}{2}\ell^{-1}\). The proof is complete. ◻

Figure 1: image.

References↩︎

[1]
M. Huang, R. P. Malhemé, and P. E. Caines, “Large population stochastic dymaic games: Closed-loop McKean-vlasov systems and the nash certainty equivalience principle,” Communications in Information and Systems, vol. 6, no. 3, pp. 221–252, 2006.
[2]
M. Huang, P. E. Caines, and R. P. Malhamé, “Large-population cost-coupled LQG problems with nonuniform agents: Individual-mass behavior and decentralized \(\epsilon\)-Nash equilibria,” IEEE transactions on automatic control, vol. 52, no. 9, pp. 1560–1571, 2007.
[3]
J.-M. Larsry and P.-L. Lions, “Mean field games,” Japanese Journal of Mathematics, vol. 2, no. 1, pp. 229–260, 2007.
[4]
A. Bensoussan, K. C. J. Sung, S. C. P. Yam, and S. P. Yung, “Linear-quadratic mean field games,” Journal of Optimization Theory and Applications, vol. 169, no. 496–529, 2016.
[5]
Y. Ma and M. Huang, “Linear quadratic mean field games with a major player: The multi-scale approach,” Automatica, vol. 133, 2020.
[6]
J. Moon and T. Başar, “Linear quadratic risk-sensitive and robust mean field games,” IEEE Transactions on Automatic Control, vol. 62, no. 3, 2017.
[7]
D. Firoozi and P. Caines, \(\varepsilon\)-nash equilibria for major–minor LQG mean field games with partial observations of all agents,” IEEE Transactions on Automatic Control, vol. 66, no. 6, 2021.
[8]
D. Firoozi and S. Jaimungal, “Exploratory LQG mean field games with entropy regulations,” Automatica, vol. 139, 2022.
[9]
P. Caines, M. Huang, and R. P. Malhamé, “Mean field games,” Handbook of Dynamic Game Theory, 2018.
[10]
R. Carmona and F. Delarue, Probabilistic theory of mean field games with application i: Mean field FBSDEs, control, and games. Springer, 2018.
[11]
L. Cardaliaguet, F. Delarue, J. M. Lasry, and P. L. Lions, The master equation and the convergence problem in mean field games. Arxiv:1509.02505, 2015.
[12]
D. Fudenberg and D. K. Lavine, “Open-loop and closed-loop equilibria in dynamic games with many players,” Journal of Economic Theory, vol. 44, no. 1, 1988.
[13]
D. Lacker, “A general characterization of the mean field limit for stochastic differential games,” Probability Theory Related Field, vol. 165, pp. 581–648, 2016.
[14]
M. Fischer, “On the connection between symmetric \(N\)-player games and mean field games,” Annals of Applied Probability, vol. 27, no. 2, pp. 757–810, 2017.
[15]
D. Lacker, “On the convergence of closed-loop nash equilibria to the mean field game limit,” Annals of Applied Probability, vol. 30, no. 4, pp. 1693–1761, 2020.
[16]
D. Lacker and L. L. Flem, “Closed-loop convergence for mean field games with common noise,” arXiv:2107.03273, 2022.
[17]
M. Iseri and J. Zhang, “Set values for mean field games,” Transactions of the American Mathematical Society, no. 377, pp. 7117–7174, 2024.
[18]
E. Bayraktar and X. Zhang, “On non-uniqueness in mean field games,” Proceedings of the American Mathematical Society, vol. 148, 2020.
[19]
A. Cecchin and F. Delarue, “Selection by vanishing common noise for potential finite state mean field games,” Communications in Partial Differential Equations, vol. 47, no. 1, pp. 89–168, 2022.
[20]
J. Dianetti, G. Ferrari, M. Fischer, and M. Nendel, “A unifying framework for submodular mean field games,” Mathematics of Operations Research, vol. 48, no. 3, pp. 1679–1710, 2023.
[21]
S. Sanjari, N. Saldi, and S. Yuksel, “Optimality of decentralized symmetric policies for stochastic teams with mean-field information sharing,” arXiv:2404.04957, 2024.
[22]
M. F. Djete, “Large population games with interactions through controls and common noise: Convergence results and equivalence between open–loop and closed–loop controls,” arXiv:2108.02992, 2021.
[23]
M. F. Djete, “Mean field games of controls: On the convergence of nash equilibria,” arXiv:2006.12993, 2021.
[24]
M. Laurière and L. Tangpi, “Convergence of large population games to mean field games with interaction through controls,” arXiv:2108.02992, 2022.
[25]
B. Jovanovic and R. W. Rosenthal, “Anonymous sequential games,” Journal of Mathematical Economics, vol. 17, pp. 77–87, 1988.
[26]
N. Saldi, T. Başar, and M. Raginsky, “Markov–nash equilibria in mean-field games with discounted cost,” SIAM Journal on Control and Optimization, vol. 56, no. 6, pp. 4256--4287, 2018.
[27]
N. Saldi, T. Başar, and M. Raginsky, “Approximate markov-nash equilibria for discrete-time risk-sensitive mean-field games,” Mathematics on Operation Research, vol. 45, no. 4, 2020.
[28]
N. Saldi, T. Başar, and M. Raginsky, “Partially observed discrete-time risk-sensitive mean field games,” Dynamic Games and Applications, 2022.
[29]
N. Bonnans, P. Lavigne, and L. Pfeiffer, “Discrete time mean field games with risk averse agents,” ESAIM: Control, Optimisation and Calculus of Variations, vol. 27, no. 44, 2021.
[30]
X. Guo, A. Hu, R. Xu, and J. Zhang, “A general framework for learning mean-field games,” Mathematics of Operations Research, 2022.
[31]
R. Hu and M. Laurière, “Recent developments in machine learning methods for stochastic control and games,” hal-03656245, 2022.
[32]
B. Acciaio, J. B. Veraguas, and J. Jia, “Cournot–nash equilibrium and optimal transport in a dynamic setting,” SIAM Journal on Control and Optimization, vol. 59, no. 3, 2021.
[33]
A. Ruszczyński, “Risk-averse dynamic programming for markov decision processes,” Mathematical Programming, Series B, vol. 125, pp. 235–261, 2010.
[34]
S. Chu and Y. Zhang, “Markov decision processes with iterated coherent risk measures,” International Journal of Control, vol. 88, no. 11, pp. 2286–2293, 2014.
[35]
A. Majumdar and M. Pavone, “How should a robot assess risk? Towards an axiomatic theory of risk in robotics,” Robotics Research, pp. 75–84, 2019.
[36]
Y. Wang and M. P. Chapman, “Risk-averse autonomous systems: A brief history and recent developments from the perspective of optimal control,” Artificial Intelligence, vol. 311, 2022.
[37]
A. Coache, S. Jaimungal, and Á. Cartea, “Conditionally elicitable dynamic risk measures for deep reinforcement learning,” SIAM Journal on Financial Mathematics, vol. 14, no. 4, pp. 1249–1289, 2023.
[38]
N. Saldi, T. Başar, and M. Raginsky, “Approximate Markov-Nash equilibria for discrete-time risk-sensitive mean-field games,” Mathematics of Operations Research, vol. 45, no. 4, pp. 1596–1620, 2020.
[39]
R. Carmona, M. Laurière, and Z. Tan, “Model-free mean-field reinforcement learning: Mean-field MDP and mean-field q-learning,” The Annals of Applied Probability, vol. 33, no. 6B, pp. 5334–5381, 2023.
[40]
L. Campi and M. Fischer, \(N\)-player games and mean field games with absorption,” Annuals of Applied Probability, vol. 28, no. 4, pp. 2188–2242, 2018.
[41]
S. Perrin, J. Perolat, M. Laurière, M. Geist, R. Elie, and Q. Pietquin, “Fictitious play for mean field games: Continuous time analysis and applications,” 34th Conference on Neural Information Processing Systems, 2020.
[42]
K. Cui and K. Koeppl, “Approximately solving mean field games via entropy-regularized deep reinforcement learning,” Proceedings of the 24th International Conference on Artificial Intelligence and Statistics, vol. 130, 2021.
[43]
J. F. Bonnans, P. Lavigne, and L. Pfeiffer, “Generalized conditional gradient and learning in potential mean field games,” arXiv:2109.05785, 2021.
[44]
R. Dumitrescu, M. Leutscher, and P. Tankov, “Linear programming fictitious play algorithm for mean field games with optimal stopping and absorption,” arXiv:2202.11428, 2022.
[45]
X. Guo, A. Hu, and J. Zhang, MF-OMO: An optimization formulation of mean-field games,” arXiv:2206.09608, 2022.
[46]
A. Hu and J. Zhang, “MF-OML: Online mean-field reinforcement learning with occupation measures for large population games,” arXiv:2405.00282, 2024.
[47]
A. Shapiro, D. Dentcheva, and A. Ruszczynski, Lectures on stochastic programming: Modeling and theory, third edition. Springer, 2021.
[48]
V. I. Bogachev, Measure theory volume II. Springer-Verlag Berlin Heidelberg, 2007.
[49]
G. Y. Weintraub, L. Benkard, and B. V. Roy, “Oblivious equilibrium: A mean field approximation for large-scale dynamic games,” Advances in Neural Information Processing Systems 18, 2005.
[50]
C. D. Aliprantis and K. C. Border, Infinite dimensional analysis: A hitchhiker’s guide. Springer-Verlag Berlin Heidelberg, 2006.
[51]
J. M. Leahy, B. Kerimkulov, D. Siska, and L. Szpruch, “Convergence of policy gradient for entropy regularized MDPs with neural network approximation in the mean-field regime,” International Conference on Machine Learning, pp. 5334–5381, 2022.
[52]
O. Hernández-Lerma and J. B. Lasserre, Discrete-time markov control processes: Basic optimality criteria. Springer, 1996.
[53]
A. Jaśkiewicz and A. S. Nowak, “Non-zero-sum stochastic games,” Handbook of Dynamic Game Theory, 2020.
[54]
H. L. Royden, Real analysis. Macmillian Publishing Company, New York, 1988.
[55]
A. N. Shiryaev, Probability-2. Springer, 2019.
[56]
R. M. Dudley, “The speed of mean glivenko-cantelli convergence,” The Annals of Mathematical Statistics, vol. 40, no. 1, pp. 40–50, 1969.
[57]
N. Fournier and A. Guillin, “On the rate of convergence in wasserstein distance of the empirical measure,” Probability Theory and Related Fields, vol. 162, pp. 707–738, 2015.
[58]
J. Lei, “Convergence and concentration of empirical measures under wasserstein distance in unbounded functional spaces,” Bernoulli, vol. 26, no. 1, pp. 767–798, 2020.
[59]
B. R. Kloeckner, “Empirical measures: Regularity is a counter-curse to dimensionality,” ESAIM: Probability and Statistics, vol. 24, pp. 408–434, 2020.
[60]
C. Acerbi, “Spectral measures of risk: A coherent representation of subjective risk aversion,” Journal of Banking and Finance, vol. 26, pp. 1505–1518, 2002.
[61]
O. T. Ting and K. W. Yip, “A generalized jensen’s inequality,” Pacific Journal of Mathematics, vol. 58, no. 1, 1975.
[62]
F. S. Scalora, “Abstract martinagle convergence theorem,” ProQuest Dissertations Publishing, 1958.
[63]
P. Billingsley, Convergence of probability measures. John Wiley & Sons, Inc, 1999.

  1. Financial Technology Thrust, The Hong Kong University of Science and Technology (Guangzhou) ().↩︎

  2. Department of Statistical Sciences, University of Toronto (, http://sebastian.statistics.utoronto.ca)↩︎

  3. SJ would like to acknowledge support from the Natural Sciences and Engineering Research Council of Canada (grants RGPIN-2018-05705 and RGPAS-2018-522715). ZC would like to acknowledge support from the Guangzhou-HKUST(GZ) Joint Funding Program (No. 2024A03J0630). ZC’s work on this project was mostly done during his postdoctoral appointment at UofT.↩︎

  4. Roughly speaking, closed-loop means each player has access to information of other players, while open-loop assumes no access to such information. The precise definitions vary across the literature. We refer to [12] and [10] for more discussions.↩︎

  5. It is necessary in the sense that, getting rid of this assumption requires adding other assumptions on other components of the games.↩︎

  6. Since in this section, \(\mathbb{X}\) and \(\mathbb{A}\) are finite, the continuity on \(\mathbb{X}\) and \(\mathbb{A}\) are given by default.↩︎

  7. Recall that \(\mathbb{P}\) is in fact \(\mathbb{P}^{\boldsymbol{\mathfrak{P}}}\).↩︎

  8. Roughly speaking, for measure-valued function, weak continuity associates the output (probability) space with weak convergence, while strong continuity associates the output space with set-wise convergence.↩︎

  9. By a version of Carathéodory extension theorem (cf. [50]), \(\overline{\psi}_t\) exists and is specified uniquely.↩︎

  10. Without loss of generality, we employ the same \(c_0, c_1\) from 4, and maintain \(\zeta_t\) for both \(N\)-play and mean field settings.↩︎

  11. These approximations, in a broad sense, can include cases such as using finite horizon games to approximate infinite horizon games.↩︎

  12. By convention, \(\inf\emptyset=\infty\)↩︎

  13. For convenience, we slightly abuse the notation here. This should not be confused with \(V_\xi\) introduced above 31 .↩︎

  14. This is true in, for example, the risk neutral setup. See also 2.1 and 4.1.↩︎

  15. For \(t\ge 2\), \(\overline{\mathcal{S}}^{\mathfrak{p}^1}_{t,T,{\overline{\Xi}}}v\) is understood as \(\overline{S}^{\mathfrak{p}^1_t}_{t,\overline{\xi}_t}\circ\cdots\circ\overline{S}^{\mathfrak{p}^1_{T-1}}_{T-1,\overline{\xi}_{T_1}} v\), and does not depend on \(\mathfrak{p}^1_1,\dots,\mathfrak{p}^1_{t-1}\).↩︎

  16. For the sake of neatness, we set \(\sum_{r=T}^{T-1}=0\).↩︎

  17. In terms of expectation under \(\rho\), it writes \(\mathbb{E}\left(H \mathbb{1}_{\{n\}}(Z) \right) = \mathbb{E}\left(H \mathbb{E}\left(\mathbb{1}_{\{n\}}(Z)\big|\{[N],\emptyset\}\otimes\mathcal{B}(\mathbb{X})\right) \right)\), where \(H(n,x)=h(x)\).↩︎

  18. We continue using \(\xi\) for the marginal measure on \(\mathcal{B}(\mathbb{X})\).↩︎

  19. In the current setting where we aim to prove the existence of MFE, superscript is no longer related to players in the \(N\)pG.↩︎