Delta-Epsilon-Common Knowledge and Quantitative Agreement Theorems


Abstract

Aumann [1] defined common knowledge mathematically and established his now famous Agreement Theorem. We present a novel approach to quantifying how close individuals are to commonly knowing events, \((\delta,\varepsilon)\)-common knowledge, which is defined for any (and not just countable) probability spaces, and provide quantitative versions of the key results in this field. Specifically, we do this for Aumann’s [1] Agreement Theorem and Nielsen’s [2] extension thereof to random variables, as well as for the setting in which posteriors are communicated back and forth between individuals. Our results apply in particular to noisy communication settings.

1 Introduction↩︎

Fifty years after its publication, Aumann’s [1] Agreement Theorem and the notion of common knowledge introduced therein have not ceased to intrigue and inspire (see, for instance, recent contributions by Billot and Vergopoulos [3], Di Tillio et al. [4], Gizatulina and Hellman [5], Hellman and Pinter [6], Geanakoplos and Polemarchakis [7], Gonczarowski and Moses [8]). The critique that in real-life scenarios, we are often closer to something like “almost” common knowledge (Halpern [9], Halpern and Moses [10], Fagin et al. [11], Rubinstein [12]) has led theorists to investigate notions of common belief. Monderer and Samet [13] show that under common \(p\)-belief, an approximate agreement result holds. Geanakoplos [14] and Morris [15] extend this result to relaxed versions of common \(p\)-belief. These accounts, however, do not cover a different line of generalization of Aumann’s theorem that consists in moving from countable to general probability spaces and from knowledge of an event to knowledge of a random variable (Nielsen [2]). The present article brings these two strands together by introducing two novel notions of approximate common knowledge, formulated in the language of \(\sigma\)-algebras and applicable to general probability spaces.

The first proposed concept, \((\delta, \varepsilon)\)-common knowledge of an event, relies on two relaxations of elementary set-theoretic notions: (1) an event \(B\) being \(\delta\)-nearly contained in a \(\sigma\)-algebra, and (2) an event \(B\) being \(\varepsilon\)-nearly contained in another event \(A\). We establish the following results:

  • Equivalence of the \(\sigma\)-algebra-based definition of \((\delta, \varepsilon)\)-common knowledge of an event to both a hierarchical and an alternating hierarchical definition thereof, generalizing Aumann’s [1] argument showing the equivalence of the partition-based definition and the informal alternating hierarchical definition of common knowledge (Proposition 8).

  • A generalization of Aumann’s [1] Agreement Theorem, showing that when two individuals have \((\delta, \varepsilon)\)-common knowledge of their posteriors of an event—or more generally, of a random variable taking values in the unit interval—then the distance between these posteriors is bounded as a function of \(\delta\) and \(\varepsilon\) (Theorem 9 and Theorem 11).

These results are related to those of Geanakoplos [14] and Morris [15] for weak common \(p\)-belief, which have been established for countable probability spaces. In contrast, all results of the present paper hold for general probability spaces. Section 3.3 provides a more detailed discussion of the relationship between \((\delta, \varepsilon)\)-common and weak common \(p\)-belief, and shows in particular how one notion can be converted into the other (Proposition 12 and Proposition 13).

For the second proposed concept, which is central to this work, we extend Nielsen’s [2] notion of common knowledge of a random variable to a notion of \((\delta, \varepsilon)\)-common knowledge of a random variable, formulated in terms of conditional variances (Definition 17). In this setting, we establish the following results:

  • Reformulations of \((\delta, \varepsilon)\)-common knowledge of a random variable with both a hierarchical and an alternating hierarchical definition of \((\delta, \varepsilon)\)-common knowledge of a random variable (Proposition 19).

  • An agreement theorem for \((\delta, \varepsilon)\)-common knowledge of a random variable which states that if the posteriors of a random variable \(X\) are \((\delta,\varepsilon)\)-common knowledge on an event \(B\), then the \(L_2\)-distance between the posteriors on \(B\) is bounded as a function of \(\delta\) and \(\varepsilon\) (Theorem 20).

  • Finally, in extension of Geanakoplos and Polemarchakis [16] and Nielsen [2], a dynamic approximate agreement result (Theorem 22), which shows that, when individuals keep learning information of a random variable \(X\), it suffices to merely approach \(\varepsilon\)-common information (Definition 15) to guarantee that \(\varepsilon\)-common information of \(X\) is reached at infinity. We illustrate this by a scenario of communication with noise (Example 1).

The versatility of \((\delta, \varepsilon\))-common knowledge rests on a conceptual generalization: instead of defining \((\delta, \varepsilon)\)-knowledge and common knowledge of an event at a particular state \(\omega\), we define it on an event. We begin in Section 2 by motivating this approach and reformulating Aumann’s [1] results within this framework.

2 Preliminaries: knowledge “on” an event↩︎

Aumann uses partitions to model individuals’ knowledge and common knowledge and states his—now iconic—theorem in that language. As shown by Nielsen [2], the result can be generalized to a framework where individuals’ knowledge is modeled by \(\sigma\)-algebras. In this more general framework, Nielsen extends the notion of common knowledge of an event to common knowledge of a random variable and shows an agreement result for knowing a random variable. We follow Nielsen in this approach, but generalize it one step further to accommodate nearby knowledge.

Nielsen utilizes the notion (and existence) of the largest set on which an individual knows an event. For our relaxation of knowledge and common knowledge, this approach is not viable, as there are typically many different events of the same size that provide common knowledge of an event up to a given margin. To address this, we define what it means for an individual to know an event \(A\) on an event \(B\).

2.1 The basic setup↩︎

We fix an ambient probability space \((\Omega, \mathcal{F}, \mathbb{P})\), where \(\Omega\) represents all possible states of the world. Individuals who reason about the world are each identified with a \(\sigma\)-algebra representing their knowledge of the world. An event is in an individual’s \(\sigma\)-algebra precisely if the individual can discern whether that event occurs.

We assume all \(\sigma\)-algebras to be complete, that is, every subset of every null set is an element of the \(\sigma\)-algebra. Furthermore, we treat sets that differ only on null sets as equivalent.

To avoid notational complexity, all definitions and results are stated for two individuals. The main results can all be reformulated to allow for more than two individuals and can be shown with virtually the same proofs as those presented here.

Definition 1. Let \(\mathcal{G}\) be a \(\sigma\)-algebra and \(A, B\) events, \(\mathbb{P}(B)>0\). Then, we say that \(\mathcal{G}\) knows \(A\) on \(B\) if \(B \in \mathcal{G}\) and \(B \subseteq A\). This property is denoted by \(K(\mathcal{G},A,B)\).

With this, we can also define common knowledge on an event.

Definition 2.

Let \(\mathcal{G}_1, \mathcal{G}_2\) be \(\sigma\)-algebras and \(A, B\) events, \(\mathbb{P}(B)>0\). Then, \(A\) is common knowledge on \(B\) if \(B \in \mathcal{G}_1\cap \mathcal{G}_2\) and \(B \subseteq A\). This property is denoted by \(K(A,B).\)

Remark 1. In the case that \(\Omega\) satisfies \(\mathbb{P}(\{ \omega \})>0\) for all \(\omega \in \Omega\), this notion of common knowledge is essentially equivalent to Aumann’s notion. If \(A\) is common knowledge on an event \(B\), then \(A\) is common knowledge at every \(\omega \in B\). Conversely, if \(A\) is common knowledge on some \(\omega\), then \(A\) is common knowledge on \(B\), where \(B\) is the smallest event in \(\mathcal{G}_1 \cap \mathcal{G}_2\) that contains \(\omega\).

As is well known, it is possible to give an alternative characterization of \(K(A,B)\) based on the informal notion originally suggested by Lewis [17], according to which an event \(A\) is common knowledge, in a group of people, if everyone knows that it occurs, everyone knows that everyone knows that it occurs, and so on.4 The following definition, which we refer to as the hierarchical definition of common knowledge, formalizes this verbal description.

Definition 3. Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras and \(A,B\) events, \(\mathbb{P}(B)>0\). Then, \(A\) is hierarchically common knowledge on \(B\) if there exist sequences \((C_n)_{n\in \mathbb{N}}\) and \((D_n)_{n\in \mathbb{N}}\) such that the following conditions hold for all \(n \in \mathbb{N}\):

  • \(K(\mathcal{G}_1, A, C_n)\) and \(K(\mathcal{G}_2,A,D_n)\),

  • \(C_{n+1} \subseteq C_n \cap D_n\) and \(D_{n+1} \subseteq C_n \cap D_n\),

  • \(B=\bigcap_{n\in \mathbb{N}} \left( C_n \cap D_n \right)\).

Remark 2. In the above, \(C_1\) is an event on which individual 1 knows \(A\); \(D_1\) an event on which individual 2 knows \(A\); \(C_2\) an event on which individual 1 knows that both individuals know \(A\); \(D_2\) an event on which individual 2 knows that both individuals know \(A\); and so on.

Remark 3. Note that \(\bigcap_{n\in \mathbb{N}} \left( C_n \cap D_n \right) = \bigcap_{n\in \mathbb{N}} C_n = \bigcap_{n\in \mathbb{N}} D_n\) and thus the condition \(B=\bigcap_{n\in \mathbb{N}} \left( C_n \cap D_n \right)\) could be equivalently posed as \(B=\bigcap_{n\in \mathbb{N}} C_n\) or \(B=\bigcap_{n\in \mathbb{N}} D_n\).

For some generalizations of common knowledge it is useful to consider a variant of the hierarchical definition encapsulating the idea that an event is common knowledge between two individuals if individual \(1\) knows that it occurs, individual \(2\) knows that individual \(1\) knows that it occurs, individual \(1\) knows that individual \(2\) knows that individual \(1\) knows it occurs, and so on ad infinitum. This is the verbal description given by Aumann.

To distinguish it from the notion given in Definition 3, we refer to it as the alternating hierarchical definition of common knowledge.

Definition 4. Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras and \(A,B\) events, \(\mathbb{P}(B)>0\). Then, \(A\) is alternatingly hierarchically common knowledge on \(B\) if there exists a sequence of events \((B_n)_{n\in \mathbb{N}}\) such that \(B=\bigcap_{n\in \mathbb{N}} B_n\), \(B_i \subseteq B_j\) for \(j < i\), and \(K(\mathcal{G}_1, A, B_n)\) if \(n\) is odd and \(K(\mathcal{G}_2,A,B_n)\) if \(n\) is even.

Proposition 4. (Aumann, Nielsen)Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras, and \(A,B\) events, \(\mathbb{P}(B)>0\). Then, the following conditions are equivalent:

  • \(A\) is common knowledge on \(B\).

  • \(A\) is hierarchically common knowledge on \(B\).

  • \(A\) is alternatingly hierarchically common knowledge on \(B\).

The proof of this equivalence is provided in a more general setting in Proposition 8.

In the language “on \(B\),” Aumann’s agreement theorem can be restated as follows.

Theorem 5. (Aumann) Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras, \(A\) an event, and \(q_1,q_2 \in [0,1]\). If the events \(\{\mathbb{P}(A|\mathcal{G}_1)=q_1\}\), \(\{ \mathbb{P}(A|\mathcal{G}_2)=q_2\}\) are common knowledge on some event \(B\), \(\mathbb{P}(B)>0\), then \(q_1=q_2\).

These definitions and results will serve as templates for the approximate versions of knowledge and common knowledge introduced in Sections 3 and 4.

3 \({(\delta,\varepsilon)}\)-common knowledge of an event and agreement↩︎

This section introduces \((\delta,\varepsilon)\)-common knowledge of an event defined on the basis of \(\sigma\)-algebras, shows that this notion is equivalent to both a hierarchical and an alternating hierarchical definition, and extends Aumann’s [1] result to that framework.

3.1 \((\delta,\varepsilon)\)-common knowledge on an event↩︎

We begin by introducing a notion that captures that an event is almost an element of a \(\sigma\)-algebra.

Definition 5. Let \(\mathcal{G}\) be a \(\sigma\)-algebra and \(B\) an event. Then, \(B\) is \(\delta\)-nearly in \(\mathcal{G}\) if there exists an event \(G \in \mathcal{G}\) such that \[\mathbb{P}(B \triangle G)\leq \mathbb{P}(B) \delta,\] where the symmetric difference between \(B\) and \(G\) is defined as \[B\triangle G := (B\setminus G) \cup (G \setminus B) .\] For this property, we write \(B \in \mathcal{G}^\delta\).

Note that \(A \in \mathcal{G}^0\) is equivalent to \(A \in \mathcal{G}\) for any complete \(\sigma\)-algebra \(\mathcal{G}\). Next, we define what it means for an event to be nearly contained in another event.

Definition 6. Let \(A,B\) be events. Then, \(B\) is \(\varepsilon\)-nearly a subset of \(A\) if \[\mathbb{P}(B\setminus A) \leq \mathbb{P}(B)\varepsilon.\] For this property, we write \(B \subseteq_\varepsilon A\).

Note that \(\subseteq_\varepsilon\) is not a transitive relation.

Given an event \(A\) and a \(\sigma\)-algebra \(\mathcal{G}\), for what follows, it is helpful to know if there is a smallest \(\delta\) such that \(A \in \mathcal{G}^\delta\). The following lemma establishes that such a \(\delta\) always exists by a constructive argument explicitly determining it.

Lemma 1. Let \(\mathcal{G}\) be a \(\sigma\)-algebra and \(A\) an event. Define \(G\in \mathcal{G}\) as \(G:=\{\mathbb{P}(A|\mathcal{G}) \geq \frac{1}{2}\}\). Then, \[\mathbb{P}(A \triangle G) = \inf_{\tilde{G} \in \mathcal{G}} \mathbb{P}(A \triangle \tilde{G} ).\]

In particular, \(\mathbb{P}(A \triangle G)/\mathbb{P}(A)\) is the smallest element of the set \(\{\delta : A \in \mathcal{G}^\delta\}\).

Proof. We need to show that \(\mathbb{P}(A \Delta G) \le \mathbb{P}(A \Delta \tilde{G})\) for every \(\tilde{G} \in \mathcal{G}\). To that end, fix \(\tilde{G} \in \mathcal{G}\). For any events \(B,C\), the inequality \(\mathbb{P}(B) \leq \mathbb{P}(C)\) is equivalent to \(\mathbb{P}(B\setminus C) \leq \mathbb{P}(C\setminus B)\). Hence, \(\mathbb{P}(A\triangle G) \leq \mathbb{P}(A \triangle \tilde{G})\) holds if and only if the inequality \[\begin{align} \mathbb{P}(A \cap (\tilde{G}\setminus G)) + \mathbb{P}(A^\complement \cap (G\setminus \tilde{G})) = \mathbb{P}((A \triangle G) \setminus(A\triangle \tilde{G}) ) \\ \leq \mathbb{P}((A \triangle \tilde{G}) \setminus(A\triangle G) ) = \mathbb{P}(A \cap (G\setminus \tilde{G})) + \mathbb{P}(A^\complement \cap (\tilde{G}\setminus G)) \end{align}\] holds. To prove this inequality, we show that \[\begin{align} \label{align:minimal95delta951} \mathbb{P}(A \cap (\tilde{G}\setminus G)) \leq \mathbb{P}(A^\complement \cap (\tilde{G}\setminus G)) \end{align}\tag{1}\] as well as \[\begin{align} \label{align:minimal95delta952} \mathbb{P}(A^\complement \cap (G\setminus \tilde{G})) \leq \mathbb{P}(A \cap (G\setminus \tilde{G})). \end{align}\tag{2}\] To prove inequality 1 , we calculate \[\begin{align} \label{align:minimal95delta953} \mathbb{P}(A \cap (\tilde{G} \setminus G)| \mathcal{G})= \mathbf{1}_{\tilde{G} \setminus G} \mathbb{P}(A|\mathcal{G}) \leq \frac{1}{2} \mathbf{1}_{\tilde{G} \setminus G} \leq \mathbf{1}_{\tilde{G} \setminus G} \mathbb{P}(A^\complement|\mathcal{G}) = \mathbb{P}(A^\complement \cap (\tilde{G} \setminus G)|\mathcal{G}), \end{align}\tag{3}\] where the inequalities follow directly from the definition of \(G\). Now, taking expectations over inequality 3 proves inequality 1 . Inequality 2 can be shown by a similar argument. ◻

After this preparation, we can formulate relaxed notions of knowledge and common knowledge.

Definition 7. Let \(\mathcal{G}\) be a \(\sigma\)-algebra, \(A,B\) events, \(\mathbb{P}(B)>0\), and \(\delta,\varepsilon\geq 0\). We say that \(A\) is \((\delta,\varepsilon)\)-known on \(B\) by \(\mathcal{G}\) if \(B\in \mathcal{G}^\delta\) and \(B \subseteq_\varepsilon A\). We denote this property by \(K_{\delta,\varepsilon}(\mathcal{G}, A,B).\)

Definition 8. Let \(\mathcal{G}_1, \mathcal{G}_2\) be \(\sigma\)-algebras, \(A, B\) events, \(\mathbb{P}(B)>0\), and \(\delta,\varepsilon\geq 0\). Then, \(A\) is \((\delta, \varepsilon)\)-common knowledge on \(B\) if \(B\in \mathcal{G}_1^\delta \cap \mathcal{G}_2^\delta\) and \(B \subseteq_\varepsilon A.\) We denote this property by \(K_{\delta,\varepsilon}( A,B)\).

Remark 6. If \(A\) is (common) knowledge on \(B\), then \(A\) is also \((0,0)\)-(common) knowledge. For the converse, if \(A\) is \((0,0)\)-(common) knowledge on \(B\), then it is also (common) knowledge on \(B\).

Remark 7. An event \(A\) is \((\delta, \varepsilon)\)-common knowledge on \(B\) if and only if \(K_{\delta,\varepsilon}(\mathcal{G}_1, A,B)\) and \(K_{\delta,\varepsilon}(\mathcal{G}_2, A,B)\), which splits the notion of \((\delta, \varepsilon)\)-common knowledge into two conditions, each referring to only one individual at a time. This reflects a well-known property of classical common knowledge, when looking at it through the lens of self-evident events: An event \(A\) is common knowledge at \(\omega\) if there exists an event \(B\), so that \(\omega \in B\) and \(B\) is self-evident to both individuals, that is, \(B \in \mathcal{G}_i\) for both \(i=1\) and \(i=2\) (see, for instance, Geanakoplos [14]).

Next, we confirm that \((\delta,\varepsilon)\)-common knowledge can be equivalently formulated by both a hierarchical and an alternating hierarchical definition.

Definition 9. Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras, \(A,B\) events, \(\mathbb{P}(B)>0\), and \(\delta,\varepsilon\geq 0\). Then, \(A\) is hierarchically \((\delta, \varepsilon)\)-common knowledge on \(B\) if there exist sequences \((C_n)_{n\in \mathbb{N}}\) and \((D_n)_{n\in \mathbb{N}}\) such that the following conditions hold for all \(n \in \mathbb{N}\):

  • \(K_{\delta,\varepsilon}(\mathcal{G}_1, A, C_n)\) and \(K_{\delta,\varepsilon}(\mathcal{G}_2,A,D_n)\),

  • \(C_{n+1} \subseteq C_n \cap D_n\) and \(D_{n+1} \subseteq C_n \cap D_n\),

  • \(B=\bigcap_{n\in \mathbb{N}} \left( C_n \cap D_n \right)\).

Definition 10. Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras, \(A,B\) events, \(\mathbb{P}(B)>0\), and \(\delta,\varepsilon\geq 0\). Then, \(A\) is alternatingly hierarchically \((\delta,\varepsilon)\)-common knowledge on \(B\) if there exists a sequence of events \((B_n)_{n\in \mathbb{N}}\) such that \(B=\bigcap_{n\in \mathbb{N}} B_n\), \(B_i \subseteq B_j\) for \(j < i\), and \(K_{\delta,\varepsilon}(\mathcal{G}_1, A,B_n)\) if \(n\) is odd and \(K_{\delta,\varepsilon}(\mathcal{G}_2, A,B_n)\) if \(n\) is even.

To show that these notions of \((\delta,\varepsilon)\)-common knowledge agree, we start with a preparatory lemma.

Lemma 2. Let \(\mathcal{G}\) be a \(\sigma\)-algebra, \(A\) an event, \((B_i)_{i\in \mathbb{N}}\) a sequence of events such that \(B_i \subseteq B_j\) for \(j<i\), and \(\delta, \varepsilon\geq 0.\) If for all \(i\in \mathbb{N}\), \[K_{\delta,\varepsilon}(\mathcal{G}, A,B_i),\] then \[K_{\delta,\varepsilon}\left(\mathcal{G}, A, \bigcap_{i\in \mathbb{N}} B_i\right).\]

Proof. Let \(B:= \bigcap_{i\in \mathbb{N}}B_i.\) We have to show that \(B \in \mathcal{G}^\delta\) and \(B \subseteq_\varepsilon A.\)

We start with \(B \in \mathcal{G}^\delta\). Let \(\tilde{\varepsilon} >0\) be given. As probability measures are continuous from above, we can find \(N \in \mathbb{N}\), such that \[\mathbb{P}(B_n \triangle B) = \mathbb{P}(B_n)- \mathbb{P}(B) \leq \tilde{\varepsilon},\] for all \(n > N\). In particular, \(\mathbb{P}(B_n) \leq \tilde{\varepsilon} + \mathbb{P}(B)\) for such \(n\).

As \(B_i \in \mathcal{G}^\delta\) for \(i \in \mathbb{N}\), there exists \(G_i \in \mathcal{G}\) such that \(\mathbb{P}(B_i \triangle G_i) \leq \delta \mathbb{P}(B_i).\) Using the triangle inequality, we find that for all \(n > N\), \[\begin{align} \mathbb{P}(B \triangle G_n) &\leq \mathbb{P}(B \triangle B_n) + \mathbb{P}(B_n \triangle G_n) \leq \tilde{\varepsilon} + \delta \mathbb{P}(B_n) \\ &\leq \tilde{\varepsilon} + \delta (\tilde{\varepsilon} + \mathbb{P}(B))= \delta \mathbb{P}(B) + \tilde{\varepsilon}(1+\delta). \end{align}\] As \(\tilde{\varepsilon}\) was arbitrarily small, we conclude \(B \in \mathcal{G}^{\tilde{\delta}}\) for all \(\tilde{\delta}> \delta\). Lemma 1 then implies that also \(B \in \mathcal{G}^\delta\).

To show that \(B \subseteq_\varepsilon A\), we simply observe that for all \(i \in \mathbb{N}\), \[\mathbb{P}(B_n \setminus A) \leq \varepsilon\mathbb{P}(B_n),\] and hence, by taking limits and by the continuity of probability measures from above, \[\mathbb{P}(B\setminus A) \leq \varepsilon\mathbb{P}(B). \qedhere\] ◻

Proposition 8. Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras, \(A,B\) events, and \(\delta, \varepsilon\geq 0\). Then, the following conditions are equivalent:

  • \(A\) is \((\delta, \varepsilon)\)-common knowledge on \(B\).

  • \(A\) is hierarchically \((\delta, \varepsilon)\)-common knowledge on \(B\).

  • \(A\) is alternatingly hierarchically \((\delta, \varepsilon)\)-common knowledge on \(B\).

Proof. To show that \(A\) being \((\delta,\varepsilon)\)-common knowledge on \(B\) implies \(A\) being hierarchically \((\delta,\varepsilon)\)-common knowledge on \(B\), simply set \(C_n:=D_n:=B\) for all \(n \in \mathbb{N}\).

To show that \(A\) being hierarchically \((\delta,\varepsilon)\)-common knowledge on \(B\) implies \(A\) being alternatingly hierarchically \((\delta,\varepsilon)\)-common knowledge on \(B\), let \((C_n)_{n\in \mathbb{N}}\) and \((D_n)_{n\in \mathbb{N}}\) be given as in the definition of hierarchical \((\delta,\varepsilon)\)-common knowledge. Now, define the sequence of events \((B_n)_{n \in \mathbb{N}}\) as \(B_n:=C_n\) for odd \(n\) and \(B_n:=D_n\) for even \(n\). It is now straightforward to check that the sequence \((B_n)_{n \in \mathbb{N}}\) shows that \(A\) is alternatingly hierarchically \((\delta,\varepsilon)\)-common knowledge on \(B\).

Lastly, we have to check that alternating hierarchical \((\delta,\varepsilon)\)-common knowledge implies \((\delta,\varepsilon)\)-common knowledge. For this, let \((B_n)_{n \in \mathbb{N}}\) be a sequence of events as in the definition of alternating hierarchical \((\delta,\varepsilon)\)-common knowledge. We introduce the monotonically decreasing sequence of events \(C_n:=B_{2n+1}\) for, \(n \in \mathbb{N}\), and note that \(B=\bigcap_{i \in \mathbb{N}}C_i.\) By assumption, \(K_{\delta,\varepsilon}(\mathcal{G}_1,A,C_n)\) for all \(n \in \mathbb{N}\), and hence, by Lemma 2, \(K_{\delta,\varepsilon}(\mathcal{G}_1,A,B)\). The property \(K_{\delta,\varepsilon}(\mathcal{G}_2,A,B)\) is proven similarly, using the sequence \((D_n)_{n\in \mathbb{N}}\) defined by \(D_n:=B_{2n}\) instead of the sequence \((C_n)_{n \in \mathbb{N}}.\) ◻

Note that we obtain Proposition 4 by setting \(\delta=\varepsilon=0\) in Proposition 8.

3.2 Aumann’s agreement theorem under \((\delta,\varepsilon)\)-common knowledge↩︎

The aim of this section is to establish results that show the robustness of Aumann’s Agreement Theorem. Our results show that even when common knowledge is weakened, substantial disagreement about the likelihood of an event is impossible.

Specifically, Theorem 9 establishes the following: If it is \((\delta,\varepsilon)\)-common knowledge that the two individuals’ posterior probabilities of a given event lie in some intervals \([a_1,b_1]\) and \([a_2,b_2]\), respectively, then the difference between their posteriors is uniformly bounded in terms of \(\delta\), \(\varepsilon\) and the widths of the intervals.

Theorem 9. Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras and \(C\) an event. Define \[A:=\{\mathbb{P}(C|\mathcal{G}_1)\in [a_1,b_1] \} \cap \{\mathbb{P}(C|\mathcal{G}_2)\in [a_2,b_2] \}\] for some \(a_1,b_1,a_2,b_2 \in [0,1]\). If there exists an event \(B\), \(\mathbb{P}(B)>0\), and \(\delta, \varepsilon\in [0,1]\), \(3\delta + 2\varepsilon\leq 1\), such that \(A\) is \((\delta,\varepsilon)\)-common knowledge on \(B\), then there exist \(q_1 \in [a_1,b_1]\) and \(q_2 \in [a_2,b_2]\) such that \[\begin{align} \label{align:Set95Aumann951} |q_1-q_2| \leq \frac{2\delta+ \varepsilon}{1+\delta} \leq 2(\delta+\varepsilon). \end{align}\tag{4}\] In particular, on \(A\) we obtain \[|\mathbb{P}(C|\mathcal{G}_1)-\mathbb{P}(C|\mathcal{G}_2)| \leq b_1-a_1+b_2-a_2+2(\delta+\varepsilon).\]

Proof. Theorem 9 is an easy consequence of the more general Theorem 11 below. ◻

As the choice of \(a_1=b_1\) and \(a_2=b_2\) is admissible in Theorem 9, considering events of the type \(\{\mathbb{P}(C|\mathcal{G}) \in [a,b]\}\) is a generalization of the usual Aumann result which concerns events of the form \(\{\mathbb{P}(C|\mathcal{G})=q\}\). The motivation for this is twofold. First, in continuous models, events of the form \(\{\mathbb{P}(C|\mathcal{G})=q\}\) are typically null sets. Hence, such events are usually not \((\delta,\varepsilon)\)-common knowledge on any set \(B\) with \(\mathbb{P}(B)>0\), and Aumann-type theorems based on such conditions are therefore typically not meaningful in continuous settings.

Second, for \(\delta=\varepsilon=0\), Theorem 9 is in its own right an extension of Aumann’s agreement theorem. It covers the situation where the precise values of the posteriors of individuals 1 and 2 are not common knowledge, but it is however common knowledge that they are within certain ranges \([a_1,b_1]\) and \([a_2,b_2]\). In this case these ranges necessarily overlap, i.e. the individuals cannot disagree entirely. The following Corollary states this explicitly.

Corollary 1. Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras and \(C\) be an event. Define \[A:= \{\mathbb{P}(C | \mathcal{G}_1) \in [a_1,b_1] \} \cap \{ \mathbb{P}( C| \mathcal{G}_2) \in [a_2,b_2] \},\] for some \(a_1,b_1,a_2,b_2 \in [0,1]\). If there exists an event \(B\), \(\mathbb{P}(B)>0\), such that \(A\) is common knowledge on \(B\), then \([a_1,b_1] \cap [a_2,b_2] \neq \emptyset.\)

In particular, on \(A\) we obtain \[|\mathbb{P}(C | \mathcal{G}_1) - \mathbb{P}(C|\mathcal{G}_2)| \leq b_1 -a_1 + b_2 +a_2.\]

Proof. This is an immediate consequence of Theorem 9 with \(\delta=\varepsilon=0\). ◻

Remark 10. For countable probability spaces, results similar to Theorem 9 have been shown by Monderer and Samet [13] in terms of common \(p\)-belief (bounds for which were improved by Neeman [18]), and by Geanakoplos [14] and Morris [15] in terms of weak common \(p\)-belief. Our results apply to general probability spaces. A comparison of weak common \(p\)-belief to \((\delta,\varepsilon)\)-common knowledge of an event is provided in Section 3.3.

Theorem 9, which, as Aumann’s theorem, is about the individuals’ posteriors attributed to an event, is a special case of the following more general agreement theorem about the individuals’ posterior expectation of a random variable.

Theorem 11. Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras and \(X\) a random variable taking values in \([0,1]\). Define \[A:=\{\mathbb{E}[X|\mathcal{G}_1]\in [a_1,b_1] \} \cap \{\mathbb{E}[X|\mathcal{G}_2]\in [a_2,b_2] \}\] for some real numbers \(a_1,b_1,a_2,b_2 \in [0,1]\). If there exists an event \(B\), \(\mathbb{P}(B)>0\), and \(\delta, \varepsilon\in [0,1]\), \(3\delta + 2\varepsilon\leq 1\), such that \(A\) is \((\delta,\varepsilon)\)-common knowledge on \(B\), then there exist \(q_1 \in [a_1,b_1]\) and \(q_2 \in [a_2,b_2]\) such that \[\begin{align} \label{vkbifydx} |q_1-q_2| \leq \frac{2\delta+ \varepsilon}{1+\delta} \leq 2 (\delta+\varepsilon). \end{align}\tag{5}\]

We give the proof for the more general Theorem 11. For this, it is useful to isolate the following lemma, which is also an important step in the proof of Neeman’s [18] Aumann-type result.

Lemma 3. Let \(Z\) be a random variable taking values in \([0,1]\) and \(E,F\) events such that \(\mathbb{P}(E\cap F) > 0\). Then, \[\frac{\mathbb{E}[Z \mathbf{1}_{E}]}{\mathbb{P}(E)} \geq \mathbb{P}(F|E) \frac{\mathbb{E}[Z \mathbf{1}_{E\cap F}]}{\mathbb{P}(E\cap F)}.\]

Proof. We calculate \[\frac{\mathbb{E}[Z \mathbf{1}_{E}]}{\mathbb{P}(E)} \geq \frac{\mathbb{E}[Z \mathbf{1}_{E\cap F}]}{\mathbb{P}(E)} = \frac{\mathbb{P}(E\cap F)}{\mathbb{P}(E)} \frac{\mathbb{E}[Z \mathbf{1}_{E\cap F}]}{\mathbb{P}(E\cap F)}=\mathbb{P}(F|E) \frac{\mathbb{E}[Z \mathbf{1}_{E \cap F}]}{\mathbb{P}(E\cap F)}. \qedhere\] ◻

Proof of Theorem 11. For \(i\in \{1,2\}\), let \(A_i:= \{\mathbb{P}(A|\mathcal{G}_i) > 0\} \in \mathcal{G}_i\), \(B_i \in \mathcal{G}_i\) such that \(\mathbb{P}(B \triangle B_i) \leq \delta \mathbb{P}(B)\), and \(C_i:=A_i\cap B_i \in \mathcal{G}_i\).

We want to apply Lemma 3 to \(X\) and the sets \(E=C_i\) and \(F=C_j\) for \(i,j\in \{1,2\}\). We first give a lower bound for \(\mathbb{P}(C_i|C_j)\): \[\begin{align} \mathbb{P}(C_i|C_j) = & \frac{\mathbb{P}(C_i \cap C_j)}{\mathbb{P}(C_j)} = \frac{\mathbb{P}(A_i \cap B_i \cap A_j \cap B_j)}{\mathbb{P}(A_j \cap B_j)} \geq \frac{\mathbb{P}(B_i \cap B_j \cap A)}{\mathbb{P}(A_j \cap B_j)}\\ &\geq \frac{\mathbb{P}(B_i\cap B_j \cap A)}{\mathbb{P}(B_j)} \geq \frac{\mathbb{P}(B_i\cap B_j \cap A \cap B)}{\mathbb{P}(B_j)} \\ &= \frac{\mathbb{P}(B_i\cap B_j \cap B) - \mathbb{P}(B_i \cap B_j \cap B \cap A^{\complement} )}{\mathbb{P}(B_j)} \\ & \geq \frac{\mathbb{P}(B) - \mathbb{P}(B\setminus B_i) - \mathbb{P}(B\setminus B_j)- \mathbb{P}(B_i \cap B_j \cap B \cap A^{\complement} )}{\mathbb{P}(B_j)} \\ & \geq \frac{\mathbb{P}(B) - \mathbb{P}(B\setminus B_i) - \mathbb{P}(B\setminus B_j)- \mathbb{P}( B \cap A^{\complement})}{\mathbb{P}(B_j)} \\ &\geq \frac{\mathbb{P}(B) (1-\delta-\varepsilon) - \mathbb{P}(B\setminus B_j)}{\mathbb{P}(B_j)}. \end{align}\] As \(\mathbb{P}(B \triangle B_j) \leq \mathbb{P}(B) \delta\), we can write \(\mathbb{P}(B\setminus B_j) = a \mathbb{P}(B)\) and \(\mathbb{P}(B_j\setminus B)=b \mathbb{P}(B)\) with \(a+b \leq \delta\), and so \(\mathbb{P}(B_j)=\mathbb{P}(B) - \mathbb{P}(B\setminus B_j) + \mathbb{P}(B_j \setminus B) = \mathbb{P}(B) (1-a+b)\). An easy computation shows that, under the condition that \(3\delta + 2 \varepsilon\leq 1\), \[\frac{\mathbb{P}(B) (1-\delta-\varepsilon) - \mathbb{P}(B\setminus B_j)}{\mathbb{P}(B_j)}=\frac{\mathbb{P}(B) (1-\delta-\varepsilon-a)}{\mathbb{P}(B)(1-a+b)} =\frac{ (1-\delta-\varepsilon-a)}{(1-a+b)}\] is minimal for \(a=0, b=\delta\). Hence,

\[\begin{align} \label{align:Set95Aumann952} \mathbb{P}(C_i|C_j) \geq \frac{ 1-\delta-\varepsilon}{1+\delta}. \end{align}\tag{6}\] Note that this also implies that \(\mathbb{P}(C_1 \cap C_2)>0\). Now, we confirm that \(C_i \subseteq \{\mathbb{E}[X|\mathcal{G}_i] \in [a_i,b_i] \}\). As \(A \subseteq \{\mathbb{E}[X|\mathcal{G}_i] \in [a_i,b_i] \}\), \[\begin{align} \label{align:Set95Aumann952465} \mathbb{P}(A|\mathcal{G}_i) \leq \mathbb{P}( \{\mathbb{E}[X|\mathcal{G}_i] \in [a_i,b_i] \}|\mathcal{G}_i) = \mathbf{1}_{ \{\mathbb{E}[X|\mathcal{G}_i] \in [a_i,b_i] \}}, \end{align}\tag{7}\] using that \(\{\mathbb{E}[X|\mathcal{G}_i] \in [a_i,b_i] \} \in \mathcal{G}_i\). On \(C_i\), we know that \(\mathbb{P}(A|\mathcal{G}_i) > 0\), inequality 7 hence implies that \(\mathbf{1}_{\{\mathbb{E}[X|\mathcal{G}_i] \in [a_i,b_i] \}}=1\) on \(C_i\), that is, \(\mathbb{E}[X|\mathcal{G}_i] \in [a_i,b_i]\) on \(C_i\).

Next, we calculate \[\begin{align} \label{align:Set95Aumann953} q_i:= \frac{\mathbb{E}[X\mathbf{1}_{C_i}]}{\mathbb{P}(C_i)} = \frac{\mathbb{E}[\mathbb{E}[X\mathbf{1}_{C_i}|\mathcal{G}_i]]}{\mathbb{P}(C_i)} = \frac{\mathbb{E}[\mathbf{1}_{C_i}\mathbb{E}[X|\mathcal{G}_i]]}{\mathbb{P}(C_i)} \in [a_i,b_i], \end{align}\tag{8}\] where we used that \(C_i \in \mathcal{G}_i\), the tower property,5 and that \(\mathbb{E}[X|\mathcal{G}_i] \in [a_i,b_i]\) on \(C_i\).

Lemma 3, together with 6 and 8 , yields \[q_i \geq \frac{1-\delta-\varepsilon}{1+\delta} \frac{\mathbb{E}[X \mathbf{1}_{ C_1\cap C_2}]}{\mathbb{P}(C_1\cap C_2)}.\] Applying Lemma 3 now to \(1-X\) and keeping \(E=C_i\) and \(F=C_j\), we get \[\begin{align} 1-q_i &= \frac{\mathbb{E}[(1-X) \mathbf{1}_{C_i}]}{\mathbb{P}(C_i)} \geq \mathbb{P}(C_i|C_j) \frac{\mathbb{E}[(1-X) \mathbf{1}_{ C_1\cap C_2}]}{\mathbb{P}(C_1\cap C_2)} \\ &\geq \frac{1-\delta-\varepsilon}{1+\delta}\frac{\mathbb{E}[(1-X) \mathbf{1}_{ C_1\cap C_2}]}{\mathbb{P}(C_1\cap C_2)} = \frac{1-\delta-\varepsilon}{1+\delta} \left(1- \frac{\mathbb{E}[X \mathbf{1}_{ C_1\cap C_2}]}{\mathbb{P}(C_1\cap C_2)}\right) \end{align}\] and hence an upper bound for \(q_i\), \[\begin{align} &q_i \leq 1 - \frac{1-\delta-\varepsilon}{1+\delta} + \frac{1-\delta-\varepsilon}{1+\delta} \frac{\mathbb{E}[X \mathbf{1}_{ C_1\cap C_2}]}{\mathbb{P}(C_1\cap C_2)} \\ &= \frac{1+\delta}{1+\delta} - \frac{1-\delta-\varepsilon}{1+\delta} + \frac{1-\delta-\varepsilon}{1+\delta} \frac{\mathbb{E}[X \mathbf{1}_{ C_1\cap C_2}]}{\mathbb{P}(C_1\cap C_2)} = \frac{2\delta+\varepsilon}{1+\delta} + \frac{1-\delta-\varepsilon}{1+\delta} \frac{\mathbb{E}[X \mathbf{1}_{ C_1\cap C_2}]}{\mathbb{P}(C_1\cap C_2)}. \end{align}\] We see that both the upper and the lower bounds are independent of \(i\), and thus hold for both \(q_1\) and \(q_2\), which means that \[|q_1-q_2| \leq \frac{2\delta+\varepsilon}{1+\delta}.\qedhere\] ◻

3.3 Comparison of weak common \(p\)-belief and \((\delta,\varepsilon)\)-common knowledge of an event↩︎

Monderer and Samet [13] introduce the concept of common \(p\)-belief, which Geanakoplos [14] generalizes to weak \(p\)-common knowledge, also referred to as weak common \(p\)-belief. We state the definition of weak common \(p\)-belief here in a slightly adapted version from Morris [15], in our language “on \(B\).”

Definition 11. Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras, \(A,B\) events, \(\mathbb{P}(B)>0\), and \(p \in [0,1]\). The event \(A\) is weak common \(p\)-belief on \(B\) if \(B\) can be written as \(B=B_1\cap B_2\) with \(B_i \in \mathcal{G}_i\) and

  • \(\mathbb{P}(B_i|B_j) \geq p\)

  • \(\mathbb{P}(A|B_i) \geq p\),

for \(i,j \in \{1,2\}\).

Weak common \(p\)-belief is closely related to the notion of \((\delta,\varepsilon)\)-common knowledge of an event. In particular, whenever an event is \((\delta,\varepsilon)\)-common knowledge with small \(\delta\) and small \(\varepsilon\), then it is weak common \(p\)-belief with large \(p\), and vice versa, as the following two propositions show.

Proposition 12. Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras, \(A,B\) events, \(\mathbb{P}(B)>0\), and \(\delta,\varepsilon\geq 0\) such that \(\delta \leq 1/3\) and \(2\varepsilon+ \delta \leq 1\). If \(A\) is \((\delta, \varepsilon)\)-common knowledge on \(B\) with \(B_1 \in \mathcal{G}_1, B_2 \in \mathcal{G}_2\) such that \(\mathbb{P}(B \triangle B_i) \leq \mathbb{P}(B) \delta\), for \(i \in \{1,2\}\), then \(A\) is weak common \(\frac{1-\max(\delta,\varepsilon)}{1+\delta}\)-belief on \(B_1 \cap B_2\).

Proof. As \(\mathbb{P}(B \triangle B_i) \leq \delta \mathbb{P}(B)\), for \(i \in \{1,2\}\), we can define non-negative numbers \(a,b,c,d\) by \[\begin{align} &\mathbb{P}(B_1 \setminus B) = a \mathbb{P}(B), \\ &\mathbb{P}(B \setminus B_1) = b \mathbb{P}(B), \\ &\mathbb{P}(B_2 \setminus B) = c \mathbb{P}(B), \\ &\mathbb{P}(B \setminus B_2) = d \mathbb{P}(B), \end{align}\] with \(a+b \leq \delta\) and \(c+d \leq \delta\). We first prove \(\mathbb{P}(B_2|B_1) \geq (1-\delta)/(1+\delta)\). The inequality with \(B_1\) and \(B_2\) swapped then follows in the same way. We calculate \[\begin{align} \mathbb{P}(B_2|B_1) &= \frac{\mathbb{P}(B_1\cap B_2)}{\mathbb{P}(B_1)} \geq \frac{\mathbb{P}(B_1\cap B_2 \cap B)}{\mathbb{P}(B)+\mathbb{P}(B_1 \setminus B)- \mathbb{P}(B \setminus B_1)} \\ &\geq \frac{\mathbb{P}(B) - \mathbb{P}(B\setminus B_1) -\mathbb{P}(B \setminus B_2) }{\mathbb{P}(B)(1+a-b)} = \frac{\mathbb{P}(B) (1-b-d) }{\mathbb{P}(B)(1+a-b)} \\ &= \frac{1-b-d}{1+a-b}. \end{align}\] An easy calculation shows that \(\frac{1-b-d}{1+a-b}\) is minimal for \(a=d=\delta, b=c=0\), if \(\delta \leq 1/3\). Hence, \[\mathbb{P}(B_2|B_1) \geq \frac{1-\delta}{1+\delta}.\] Next, we show that \(P(A|B_1) \geq (1-\varepsilon)/(1+\delta)\). The inequality \(P(A|B_2) \geq (1-\varepsilon)/(1+\delta)\) then follows in the same way. We thus calculate \[\begin{align} \mathbb{P}(A|B_1) &= \frac{P(A\cap B_1)}{\mathbb{P}(B_1)} \geq \frac{\mathbb{P}(B \cap B_1 \cap A)}{\mathbb{P}(B) + \mathbb{P}(B_1 \setminus B) - \mathbb{P}(B \setminus B_1)} \\ &\geq \frac{\mathbb{P}(B)- \mathbb{P}(B \setminus B_1) - \mathbb{P}(B \setminus A)}{\mathbb{P}(B) + \mathbb{P}(B_1 \setminus B) - \mathbb{P}(B \setminus B_1)} \geq \frac{\mathbb{P}(B)(1-b-\varepsilon)}{\mathbb{P}(B)(1+a-b)}=\frac{1-b-\varepsilon}{1+a-b}. \end{align}\] Under the assumption that \(2\varepsilon+ \delta \leq 1\), the expression \((1-b-\varepsilon)/(1+a-b)\) is minimal for \(a=\delta,b=0\), and hence \[\mathbb{P}(A|B_1) \geq \frac{1-\varepsilon}{1+\delta}. \qedhere\] ◻

Proposition 13. Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras, \(A\) an event, \(B_1 \in \mathcal{G}_1, B_2 \in \mathcal{G}_2\) events such that \(\mathbb{P}(B_1 \cap B_2)>0\), and \(p \in [0,1]\). If \(A\) is weak common \(p\)-belief on \(B:=B_1\cap B_2\), then \(A\) is \((\frac{1-p}{p},\frac{1-p}{p})\)-common knowledge on \(B\).

Proof. We first show that \[\mathbb{P}(B_1 \triangle B) \leq \mathbb{P}(B) \frac{1-p}{p}.\] The analogous inequality for \(B_2\) can be proved in the same way. First, observe that the assumption \(\mathbb{P}(B_2|B_1)\geq p\) can be rewritten as \[\begin{align} \label{align:p95to95epsilon} \mathbb{P}(B_1) \leq \mathbb{P}(B_1 \cap B_2)/p. \end{align}\tag{9}\] Hence, \[\begin{align} &\mathbb{P}(B \triangle B_1) = \mathbb{P}((B_1 \cap B_2) \triangle B_1) = \mathbb{P}(B_1) - \mathbb{P}(B_1 \cap B_2) \\ &\leq \frac{\mathbb{P}(B_1 \cap B_2)}{p} - \mathbb{P}(B_1 \cap B_2) = \mathbb{P}(B_1 \cap B_2) \frac{1-p}{p} =\mathbb{P}(B) \frac{1-p}{p}. \end{align}\] What is left to prove is that \(B \subseteq_\varepsilon A.\) Note that \(\mathbb{P}(A|B_1) \geq p\) is equivalent to \(\mathbb{P}(A \cap B_1) \geq p \mathbb{P}(B_1)\), and thus \[\begin{align} \mathbb{P}(B \setminus A) &= \mathbb{P}((B_1 \cap B_2) \setminus A) \leq \mathbb{P}(B_1 \setminus A) = \mathbb{P}(B_1)- \mathbb{P}(A\cap B_1)\\ &\leq \mathbb{P}(B_1) - p\mathbb{P}(B_1) = \mathbb{P}(B_1) (1-p) \leq \mathbb{P}(B) \frac{1-p}{p}, \end{align}\] where we used inequality 9 in the last step. ◻

Remark 14. Note that, while each of the two previous lemmas is sharp, alternating applications of these conversions between weak common \(p\)-belief and \((\delta,\varepsilon)\)-common knowledge result in worse constants. For instance, converting weak common \(p\)-belief back and forth leads to weak common \((2p-1)\)-belief.

This implies that the two concepts can be applied in the same settings. An appealing feature of \((\delta,\varepsilon)\)-common knowledge is the equivalence of the \(\sigma\)-algebra-based definition to the two hierarchical definitions (Proposition 8), generalizing a characteristic of common knowledge (Proposition 4), which does not seem to be available for (weak) common \(p\)-belief (see, Morris [15]). A further advantage of \((\delta,\varepsilon)\)-common knowledge is that it is not restricted to the framework of countable probability spaces, but is defined for arbitrary probability spaces. This can be useful for extending results about the role of common \(p\)-belief for equilibrium selection in Bayesian games to Bayesian games with infinite type spaces.

4 Approximate common knowledge of a random variable and agreement↩︎

Nielsen [2] extends the concept of common knowledge to random variables. Following this approach, we introduce a notion of \((\delta,\varepsilon)\)-common knowledge for random variables. We begin by reviewing Nielsen’s account.

4.1 Nielsen’s extension of common knowledge to random variables↩︎

Nielsen introduces the following notions of knowledge and common knowledge of a random variable.

Definition 12. (Nielsen)Let \(X\) be a random variable and \(\mathcal{G}_1, \mathcal{G}_2\) \(\sigma\)-algebras. The event \(K(\mathcal{G}_i,X)\) that \(\mathcal{G}_i\) knows \(X\), for \(i \in \{1,2\}\), is defined as the largest set \(A \in \mathcal{G}_i\) such that \(X\mathbf{1}_{A}\) is \(\mathcal{G}_i\)-measurable. The event \(K(X)\) that \(X\) is common knowledge is defined as \(K(\mathcal{G}_1 \cap \mathcal{G}_2,X)\).

With these notions, Nielsen shows the following generalization of Aumann’s Agreement Theorem (5).

Theorem 15. (Nielsen)Let \(X\) be a random variable and \(\mathcal{G}_1,\mathcal{G}_2\) \(\sigma\)-algebras. For \(i\in \{1,2\}\) we have \(\mathbb{E}[X|\mathcal{G}_i]=\mathbb{E}[X| \mathcal{G}_1 \cap \mathcal{G}_2]\) almost surely on \(K(\mathbb{E}[X|\mathcal{G}_i])\). In particular, \(\mathbb{E}[X|\mathcal{G}_1]=\mathbb{E}[X|\mathcal{G}_2]\) almost surely on \(K(\mathbb{E}[X|\mathcal{G}_1]) \cap K(\mathbb{E}[X|\mathcal{G}_2])\).

As knowing a random variable \(X\) is defined in terms of the measurability of \(X\), we need to relax measurability to reach a relaxed notion of knowing \(X\). To quantify how well \(X\) is known by an individual with information \(\mathcal{G}\)—that is, how close it is to the \(\mathcal{G}\)-measurable posterior \(\mathbb{E}[X|\mathcal{G}]\)—we use the conditional variance \(\mathrm{Var}(X|\mathcal{G})= \mathbb{E}[(X-\mathbb{E}[X|\mathcal{G}])^2|\mathcal{G}]\). We begin by applying this approach to the special case of common knowledge of a random variable on the entire state space and establish an agreement result for this notion. Building on this, we introduce a relaxation of common knowledge of a random variable, \((\delta,\varepsilon)\)-common knowledge of a random variable, on an event \(B\), and show that an approximate agreement result holds, which constitutes the main theorem of the section.

4.2 Common information and \(\varepsilon\)-common information↩︎

Nielsen introduces notions for the particular case that (common) knowledge of a random variable occurs on the entire state space.

Definition 13. (Nielsen)Let \(X\) be a random variable, \(\mathcal{G}_1, \mathcal{G}_2\) \(\sigma\)-algebras and \(i \in \{1,2\}\). If \(K(\mathcal{G}_i,X)=\Omega\), then \(\mathcal{G}_i\) is said to be informed about \(X\). If \(K(X)=\Omega\), then \(X\) is said to be common information.

Remark 16. The property that a random variable \(X\) is common information can equivalently be expressed through notions that only refer to a single individual, namely, that each individual is informed about \(X\) (see also Remark 7).

We begin by defining what it means for an individual to be \(\varepsilon\)-informed about a random variable \(X\) and what it means for \(X\) to be \(\varepsilon\)-common information.

Definition 14. Let \(X \in L_2(\Omega)\), \(\mathcal{G}\) a \(\sigma\)-algebra, and \(\varepsilon\geq 0\). We say that \(\mathcal{G}\) is \(\varepsilon\)-informed about \(X\) if \[\mathbb{E}[\mathrm{Var}(X|\mathcal{G})] \leq \varepsilon.\]

Definition 15.

Let \(X\in L_2(\Omega)\) and \(\mathcal{G}_1, \mathcal{G}_2\) \(\sigma\)-algebras. We say that \(X\) is \(\varepsilon\)-common information if \[\mathbb{E}[\mathrm{Var}(X|\mathcal{G}_i)] \leq \varepsilon,\] for \(i \in \{1,2\}\).

Note that \(0\)-common information coincides with common information, as defined by Nielsen (Definition 13).

First, we show a preparatory lemma about the difference of posteriors.

Lemma 4. Let \(X \in L_2(\Omega)\) be a random variable and \(\mathcal{G}_1,\mathcal{G}_2\) \(\sigma\)-algebras. Then, \[\begin{align} \mathbb{E}[(\mathbb{E}[X|\mathcal{G}_1]-\mathbb{E}[X|\mathcal{G}_2])^2] \leq ||\mathbb{E}[X|\mathcal{G}_1]+ \mathbb{E}[X|\mathcal{G}_2]-\mathbb{E}[\mathbb{E}[X|\mathcal{G}_1]|\mathcal{G}_2]-\mathbb{E}[\mathbb{E}[X|\mathcal{G}_2]|\mathcal{G}_1]||_2 \cdot ||X||_2. \end{align}\]

Proof. Using several times that conditional expectations are self-adjoint, that is, \[\mathbb{E}[Y \cdot \mathbb{E}[Z|\mathcal{G}]] = \mathbb{E}[\mathbb{E}[Y|\mathcal{G}] \cdot Z],\] for any random variables \(Y,Z \in L_2(\Omega)\) and \(\sigma\)-algebra \(\mathcal{G}\), we calculate \[\begin{align} &\mathbb{E}[(\mathbb{E}[X|\mathcal{G}_1]-\mathbb{E}[X|\mathcal{G}_2])^2] = \mathbb{E}[\mathbb{E}[X|\mathcal{G}_1]^2] + \mathbb{E}[\mathbb{E}[X|\mathcal{G}_2]^2] - 2 \mathbb{E}[\mathbb{E}[X|\mathcal{G}_1]\mathbb{E}[X|\mathcal{G}_2]]\\ &=\mathbb{E}[ \mathbb{E}[X|\mathcal{G}_1] \cdot X] + \mathbb{E}[\mathbb{E}[X|\mathcal{G}_2] \cdot X] - \mathbb{E}[ \mathbb{E}[\mathbb{E}[X|\mathcal{G}_1]|\mathcal{G}_2] \cdot X] - \mathbb{E}[ \mathbb{E}[\mathbb{E}[X|\mathcal{G}_2]|\mathcal{G}_1] \cdot X] \\ &=\mathbb{E}[(\mathbb{E}[X|\mathcal{G}_1] + \mathbb{E}[X|\mathcal{G}_2] - \mathbb{E}[\mathbb{E}[X|\mathcal{G}_1]|\mathcal{G}_2] - \mathbb{E}[\mathbb{E}[X|\mathcal{G}_2]|\mathcal{G}_1]) \cdot X] \\ &\leq ||\mathbb{E}[X|\mathcal{G}_1]+ \mathbb{E}[X|\mathcal{G}_2]-\mathbb{E}[\mathbb{E}[X|\mathcal{G}_1]|\mathcal{G}_2]-\mathbb{E}[\mathbb{E}[X|\mathcal{G}_2]|\mathcal{G}_1]||_2 \cdot ||X||_2, \end{align}\] where the last inequality is due to the Cauchy–Schwarz inequality. ◻

We next establish an Aumann-type result for the particular case of \(\varepsilon\)-common information, showing that if the posteriors of a random variable \(X\) are \(\varepsilon\)-common information, then the expected squared difference of posteriors is at most \(2\sqrt{\varepsilon\mathrm{Var}(X)}\).

Theorem 17. Let \(\mathcal{G}_1, \mathcal{G}_2\) be \(\sigma\)-algebras, \(X\in L_2(\Omega)\), \(\varepsilon\geq 0\). If the posteriors \(\mathbb{E}[X|\mathcal{G}_1]\) and \(\mathbb{E}[X|\mathcal{G}_2]\) are both \(\varepsilon\)-common information, then \[\mathbb{E}[ (\mathbb{E}[X|\mathcal{G}_1]-\mathbb{E}[X|\mathcal{G}_2])^2 ] \leq 2\sqrt{\varepsilon\mathrm{Var}(X)}.\]

Proof. Without loss of generality, we may assume that \(\mathbb{E}[X]=0\). Using Lemma 4 and the triangle inequality in \(L_2(\Omega)\), we find \[\begin{align} || \mathbb{E}[X|\mathcal{G}_1]- \mathbb{E}[X|\mathcal{G}_2]||_2^2 & \leq ||\mathbb{E}[X|\mathcal{G}_1] - \mathbb{E}[ \mathbb{E}[X|\mathcal{G}_1] |\mathcal{G}_2] + \mathbb{E}[X|\mathcal{G}_2] - \mathbb{E}[ \mathbb{E}[X|\mathcal{G}_2] |\mathcal{G}_1]||_2 \cdot ||X||_2 \\ &\leq (\sqrt{\mathbb{E}[\mathrm{Var}( \mathbb{E}[X|\mathcal{G}_1]|\mathcal{G}_2)]} + \sqrt{\mathbb{E}[\mathrm{Var}( \mathbb{E}[X|\mathcal{G}_2]|\mathcal{G}_1)]}\;) \cdot \sqrt{\mathrm{Var}(X)} \\ &\leq 2 \sqrt{\varepsilon\mathrm{Var}(X)}.\qedhere \end{align}\] ◻

Note that for random variables in \(L_2(\Omega)\) that are common information, setting \(\varepsilon=0\) in Theorem 17 recovers Nielsen’s theorem (Theorem 15).

4.3 \((\delta,\varepsilon)\)-common knowledge of a random variable↩︎

In this subsection, we give a notion of approximate common knowledge of a random variable on an event B, show that this notion allows different equivalent formulations, and establish an Aumann-type result for it. To that end, we recall the definition of the restriction of a probability space.

Let \((\Omega, \mathcal{F}, \mathbb{P})\) be a probability space and \(B\) an event, \(\mathbb{P}(B)>0\). The probability space \((B, \mathcal{F}_B, \mathbb{P}_B)\) is called the restriction of \((\Omega, \mathcal{F}, \mathbb{P})\) to \(B\), where \(\mathcal{F}_B := \{ A \cap B : A \in \mathcal{F}\}\) is the trace \(\sigma\)-algebra of \(\mathcal{F}\) in \(B\) and \({\mathbb{P}_B(A):=\mathbb{P}(A\cap B)/\mathbb{P}(B)}\), for any event \(A\). We further denote the restriction of a random variable \(X\) to \(B\) by \(X_{|B}\). When referring to an individual’s \(\sigma\)-algebra \(\mathcal{G}\) on \((B, \mathcal{F}_B, \mathbb{P}_B)\), we implicitly refer to the trace \(\sigma\)-algebra \(\mathcal{G}_B\).

With this, we can define \((\delta,\varepsilon)\)-knowledge and \((\delta,\varepsilon)\)-common knowledge of a random variable \(X\) on an event \(B\).

Definition 16. Let \(X\in L_\infty(\Omega)\), \(\mathcal{G}\) a \(\sigma\)-algebra, \(B\) an event, \(\mathbb{P}(B)>0\), and \(\delta, \varepsilon\geq 0\). We say that \(X\) is \((\delta,\varepsilon)\)-known on \(B\) by \(\mathcal{G}\) if \(B \in \mathcal{G}^\delta\) and \(\mathcal{G}_{B}\) is \(\varepsilon\)-informed about \(X_{|B}\) on the space \((B, \mathcal{F}_B, \mathbb{P}_B)\), that is, \(\mathbb{E}_{\mathbb{P}_B}[\mathrm{Var}_{\mathbb{P}_B}(X_{|B}|\mathcal{G}_{B})]\leq \varepsilon\). For this property, we write \(K_{ \delta, \varepsilon}(\mathcal{G}, X, B).\)

Definition 17. Let \(X\in L_\infty(\Omega)\), \(\mathcal{G}_1, \mathcal{G}_2\) \(\sigma\)-algebras, \(B\) an event, \(\mathbb{P}(B)>0\), and \(\delta, \varepsilon\geq 0\). We say that \(X\) is \((\delta,\varepsilon)\)-common knowledge on \(B\) if \(K_{\delta,\varepsilon}(\mathcal{G}_1, X, B)\) and \(K_{\delta,\varepsilon}(\mathcal{G}_2, X, B).\) For this notion, we write \(K_{\delta,\varepsilon}( X, B).\)

Remark 18. On \(B\), we have \[\mathbb{E}_{\mathbb{P}_B}[X_{|B}|\mathcal{G}_{B}]=\mathbb{E}[X|\sigma(\mathcal{G}\cup \{B\}) ].\]

To establish alternative hierarchical characterizations and an agreement theorem for \((\delta,\varepsilon)\)-common knowledge of a random variable, we need to define a distance between complete \(\sigma\)-algebras. We adopt a notion introduced by Rogge [20]. A slightly different, equivalent metric was previously studied by Boylan [21] and by Neveu [22].

Definition 18. Let \((\Omega, \mathcal{F}, \mathbb{P})\) be a probability space. The set of all complete sub-\(\sigma\)-algebras of \(\mathcal{F}\) is endowed with the metric \(d\) given by \[d(\mathcal{G}_1,\mathcal{G}_2):= \max \left\{ \sup_{G_1 \in \mathcal{G}_1} \inf_{G_2\in\mathcal{G}_2} \mathbb{P}(G_1 \triangle G_2) , \sup_{G_2\in \mathcal{G}_2} \inf_{G_1\in\mathcal{G}_1} \mathbb{P}(G_1\triangle G_2)\right\},\] for complete \(\sigma\)-algebras \(\mathcal{G}_1,\mathcal{G}_2 \subseteq \mathcal{F}\).

An important property of this distance is the following result of Rogge [20].

Lemma 5. (Rogge)Let \(X\) be a random variable taking values in \([0,1]\) and \(\mathcal{G}_1, \mathcal{G}_2\) sub-\(\sigma\)-algebras of \(\mathcal{F}\). Then, \[|| \mathbb{E}[X|\mathcal{G}_1]-\mathbb{E}[X|\mathcal{G}_2]||_2 \leq \sqrt{2 d(\mathcal{G}_1,\mathcal{G}_2)(1-d(\mathcal{G}_1,\mathcal{G}_2))}.\]

This shows that both conditional expectations and conditional variances of given bounded random variables are continuous with respect to the \(\sigma\)-algebra. In particular, \[\begin{align} \label{align:bedingte95varianz95ist95stetig} ||\mathrm{Var}(X|\mathcal{G}_1)-\mathrm{Var}(X|\mathcal{G}_2)||_2 \leq 4\sqrt{2d(\mathcal{G}_1,\mathcal{G}_2) (1-d(\mathcal{G}_1,\mathcal{G}_2))}, \end{align}\tag{10}\] for every random variable \(X\) taking values in \([0,1]\) and \(\sigma\)-algebras \(\mathcal{G}_1,\mathcal{G}_2\).

We proceed by collecting results on \(\sigma\)-algebras augmented by an additional event, which we will rely on in the proof of the main theorem of this section.

Definition 19. Let \(\mathcal{A}\) be a collection of events. We denote the smallest \(\sigma\)-algebra that contains \(\mathcal{A}\) by \(\sigma(\mathcal{A})\).

Lemma 6. Let \(\mathcal{G}\) be a \(\sigma\)-algebra and \(A\) an event. Then,

\[\sigma(\mathcal{G}\cup \{A\})= \{(G \cap A) \cup (H \cap A^\complement) | G,H\in \mathcal{G}\}.\]

Proof. The result is obtained by a direct computation. ◻

Lemma 7. Let \(\mathcal{G}\) be a \(\sigma\)-algebra, \(A\) an event, and \(\varepsilon\geq 0\) such that \[\inf_{G\in \mathcal{G}} \mathbb{P}(A \triangle G) \leq \varepsilon.\] Then, \[d(\mathcal{G}, \sigma(\mathcal{G}\cup \{A\})) \leq \varepsilon.\]

Proof. We begin by showing that \[\sup_{\tilde{G} \in \sigma(\mathcal{G}\cup \{A\})} \inf_{G'\in \mathcal{G}} \mathbb{P}(\tilde{G} \triangle G') \leq \varepsilon.\] By Lemma 6, any event \(\tilde{G} \in \sigma(\mathcal{G}\cup \{A\})\) can be expressed as \[\tilde{G}=(G \cap A) \cup (H \cap A^\complement),\] for some \(G,H \in \mathcal{G}.\) Furthermore, we know by Lemma 1 that the event \(B:=\{\mathbb{P}(A|\mathcal{G})\geq \frac{1}{2}\}\) fulfills \[\mathbb{P}(A\triangle B) \leq \varepsilon.\] We proceed by approximating \(\tilde{G}\) by \((G \cap B) \cup (H \cap B^\complement) \in \mathcal{G}\) and find \[\begin{align} &\inf_{G'\in \mathcal{G}} \mathbb{P}(((G \cap A) \cup (H \cap A^\complement) ) \triangle G') \leq \mathbb{P}( ((G \cap A) \cup (H \cap A^\complement)) \triangle( (G \cap B) \cup (H \cap B^\complement) ) ) \\ &\leq \mathbb{P}((( G \cap A) \triangle (G \cap B)) \cup ((H \cap A^\complement) \triangle (H \cap B^\complement))) \\ &\leq \mathbb{P}( (A\triangle B) \cup (A^\complement \triangle B^\complement)) = \mathbb{P}(A\triangle B) \leq \varepsilon. \end{align}\] As \(\tilde{G}\) was chosen arbitrarily, we conclude \[\sup_{\tilde{G} \in \sigma(\mathcal{G}\cup \{A\})} \inf_{G'\in \mathcal{G}} \mathbb{P}(\tilde{G} \triangle G') \leq \varepsilon.\] Finally, to conclude the proof, we note that \[\sup_{G'\in \mathcal{G}}\inf_{\tilde{G} \in \sigma(\mathcal{G}\cup \{A\})} \mathbb{P}(G' \triangle \tilde{G}) =0,\] as \(\mathcal{G}\subseteq \sigma(\mathcal{G}\cup \{A\})\), and thus \[d(\mathcal{G}, \sigma(\mathcal{G}\cup \{A\})) \leq \varepsilon.\qedhere\] ◻

Lemma 8. Let \(\mathcal{G}\) be a \(\sigma\)-algebra, \(A,B\) events. Then, \[d(\sigma(\mathcal{G}\cup \{A\}), \sigma(\mathcal{G}\cup \{B\})) \leq \mathbb{P}(A \triangle B ).\]

Proof. We show that any event \(A' \in \sigma(\mathcal{G}\cup \{A\})\) can be approximated by an event \(B' \in \sigma(\mathcal{G}\cup \{B\})\) so that \(\mathbb{P}(A' \triangle B') \leq \mathbb{P}(A \triangle B )\). Let \(A' \in \sigma(\mathcal{G}\cup \{A\})\). By Lemma 6 we know that there exist events \(G, H \in \mathcal{G}\) so that \(A'= (G \cap A) \cup (H \cap A^\complement)\). By the same lemma, we also know that the event \(B'= (G \cap B) \cup (H \cap B^\complement)\) is an element of \(\sigma(\mathcal{G}\cup \{B\})\). We estimate the symmetric difference between \(A'\) and \(B'\) by \[\begin{align} \mathbb{P}(A' \triangle B')& = \mathbb{P}(( (G \cap A) \cup (H \cap A^\complement)) \triangle ((G \cap B) \cup (H \cap B^\complement))) \\ &\leq \mathbb{P}(( (G \cap A) \triangle (G \cap B) ) \cup ( (H \cap A^\complement) \triangle (H \cap B^\complement) ) )\\ & \leq \mathbb{P}((A \triangle B) \cup (A^\complement \triangle B^\complement)) = \mathbb{P}(A\triangle B). \end{align}\]

As this is true for every \(A' \in \sigma(\mathcal{G}\cup \{A\})\), we find \[\sup_{A' \in \sigma(\mathcal{G}\cup \{A\})} \inf_{B'\in \sigma(\mathcal{G}\cup \{A\})} \mathbb{P}(A' \triangle B') \leq \mathbb{P}(A \triangle B ).\]

The same inequality holds true with the roles of \(A\) and \(B\) reversed. Hence, \(d(\sigma(\mathcal{G}\cup \{A\}), \sigma(\mathcal{G}\cup \{B\})) \leq \mathbb{P}(A \triangle B ).\) ◻

We proceed by establishing that our new notion of approximate common knowledge of a random variable \(X\) admits alternative hierarchical characterizations.

Definition 20. Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras, \(X \in L_\infty(\Omega)\) a random variable, \(B\) an event, \(\mathbb{P}(B)>0\), and \(\delta,\varepsilon\geq 0\). Then, \(X\) is hierarchically \((\delta, \varepsilon)\)-common knowledge on \(B\) if there exist sequences \((C_n)_{n\in \mathbb{N}}\) and \((D_n)_{n\in \mathbb{N}}\) such that the following conditions hold for all \(n \in \mathbb{N}\):

  • \(K_{\delta,\varepsilon}(\mathcal{G}_1, X, C_n)\) and \(K_{\delta,\varepsilon}(\mathcal{G}_2,X,D_n)\),

  • \(C_{n+1} \subseteq C_n \cap D_n\) and \(D_{n+1} \subseteq C_n \cap D_n\),

  • \(B=\bigcap_{n\in \mathbb{N}} \left( C_n \cap D_n \right)\).

Definition 21. Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras, \(X \in L_\infty(\Omega)\) a random variable, \(B\) an event, \(\mathbb{P}(B)>0\), and \(\delta,\varepsilon\geq 0\). Then, \(X\) is alternatingly hierarchically \((\delta,\varepsilon)\)-common knowledge on \(B\) if there exists a sequence of events \((B_n)_{n\in \mathbb{N}}\) such that \(B=\bigcap_{n\in \mathbb{N}} B_n\), \(B_i \subseteq B_j\) for \(j < i\), and \(K_{\delta,\varepsilon}(\mathcal{G}_1, X,B_n)\) if \(n\) is odd and \(K_{\delta,\varepsilon}(\mathcal{G}_2, X,B_n)\) if \(n\) is even.

Proposition 19. Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras, \(X \in L_\infty(\Omega)\) a random variable, \(B\) an event, and \(\delta, \varepsilon\geq 0\). Then, the following conditions are equivalent:

  • \(X\) is \((\delta, \varepsilon)\)-common knowledge on \(B\).

  • \(X\) is hierarchically \((\delta, \varepsilon)\)-common knowledge on \(B\).

  • \(X\) is alternatingly hierarchically \((\delta, \varepsilon)\)-common knowledge on \(B\).

Proof. The implications \((1) \Rightarrow (2)\) and \((2) \Rightarrow (3)\) can be shown with the same proof as in Proposition 8.

What is left to show is that \(X\) being alternatingly hierarchically \((\delta, \varepsilon)\)-common knowledge on \(B\) implies \(X\) being \((\delta, \varepsilon)\)-common knowledge on \(B\). Let \((B_n)_{n\in \mathbb{N}}\) be a sequence of events as in the definition of alternating hierarchical \((\delta,\varepsilon)\)-common knowledge. Applying Lemma 2 to the sequence \((B_{2n+1})_{n\in \mathbb{N}}\) and \(A:=\Omega\), we see that \(B \in \mathcal{G}_1^\delta\). The same argument applied to the sequence \((B_{2n})_{n\in \mathbb{N}}\) gives \(B \in \mathcal{G}_2^\delta\) and thus \(B \in \mathcal{G}_1^\delta \cap \mathcal{G}_2^\delta.\)

It remains to be shown that \(\mathbb{E}_{\mathbb{P}_B}[\mathrm{Var}_{\mathbb{P}_B}(X_{|B}|\mathcal{G}_{i_{B}})] \leq \varepsilon\) for \(i \in \{1,2\}\). As \(\lim_{n\to \infty} \mathbb{P}(B_n \triangle B) =0\), Lemma 5 together with Lemma 8 implies that \[\lim_{n\to \infty} \mathbb{E}[X | \sigma(\mathcal{G}_i \cup \{B_n\})]= \mathbb{E}[X|\sigma(\mathcal{G}_i \cup \{B\})]\] in \(L_2(\Omega)\). From this, it is easy to see that also \[\begin{align} &\lim_{n \to \infty} \mathbb{E}_{\mathbb{P}_{B_{n}}}[\mathrm{Var}_{\mathbb{P}_{B_{n}}}(X_{|B_{n}}|\mathcal{G}_{i_{B_n}})]=\lim_{n \to \infty} \frac{\mathbb{E}[(X-\mathbb{E}[X|\sigma(\mathcal{G}_i \cup \{B_n\})])^2\mathbf{1}_{B_n}]}{\mathbb{P}(B_n)} \\ &=\frac{\mathbb{E}[(X-\mathbb{E}[X|\sigma(\mathcal{G}_i \cup \{B\})])^2\mathbf{1}_{B}]}{\mathbb{P}(B)}= \mathbb{E}_{\mathbb{P}_{B}}[\mathrm{Var}_{\mathbb{P}_{B}}(X_{|B}|\mathcal{G}_{i_{B}})] \end{align}\] in \(L_2(\Omega)\). This concludes the proof as \[\mathbb{E}_{\mathbb{P}_B}[\mathrm{Var}_{\mathbb{P}_B}(X_{|B}|\mathcal{G}_{1_{B}})] = \lim_{n\to \infty} \mathbb{E}_{\mathbb{P}_{B_{2n+1}}}[\mathrm{Var}_{\mathbb{P}_{B_{2n+1}}}(X_{|B_{2n+1}}|\mathcal{G}_{1_{B_{2n+1}}})] \leq \varepsilon\] as well as \[\mathbb{E}_{\mathbb{P}_B}[\mathrm{Var}_{\mathbb{P}_B}(X_{|B}|\mathcal{G}_{2_{B}})] = \lim_{n\to \infty} \mathbb{E}_{\mathbb{P}_{B_{2n}}}[\mathrm{Var}_{\mathbb{P}_{B_{2n}}}(X_{|B_{2n}}|\mathcal{G}_{2_{B_{2n}}})] \leq \varepsilon. \qedhere\] ◻

We are ready to prove the main theorem of this section.

Theorem 20. Let \(\mathcal{G}_1,\mathcal{G}_2\) be \(\sigma\)-algebras, \(X\in L_\infty(\Omega)\), \(B\) an event, \(\mathbb{P}(B)>0\), and \(\delta,\varepsilon\geq 0\) such that \(K_{\delta,\varepsilon}(\mathbb{E}_{\mathbb{P}_B}[X|\mathcal{G}_1], B)\) and \(K_{\delta,\varepsilon}(\mathbb{E}_{\mathbb{P}_B}[X|\mathcal{G}_2], B).\) Then, the following inequality holds \[\mathbb{E}_{\mathbb{P}_B}[(\mathbb{E}[X|\mathcal{G}_1]-\mathbb{E}[X|\mathcal{G}_2])^2] \leq 64 \delta ||X||_{L_\infty(\mathbb{P})}^2 +4 \sqrt{\varepsilon} ||X||_{L_\infty(\mathbb{P})}.\]

Proof. Using the triangle inequality in \(L_2(\mathbb{P}_B)\), we find

\[\begin{align} & ||\mathbb{E}[X|\mathcal{G}_1]-\mathbb{E}[X|\mathcal{G}_2]||_{L_2(\mathbb{P}_B)} \leq ||\mathbb{E}[X|\mathcal{G}_1]-\mathbb{E}_{\mathbb{P}_B}[X|\mathcal{G}_1]||_{L_2(\mathbb{P}_B)} \\ &+ ||\mathbb{E}_{P_B}[X|\mathcal{G}_1]-\mathbb{E}_{P_B}[X|\mathcal{G}_2]||_{L_2(\mathbb{P}_B)} +||\mathbb{E}[X|\mathcal{G}_2]-\mathbb{E}_{\mathbb{P}_B}[X|\mathcal{G}_2]||_{L_2(\mathbb{P}_B)}. \end{align}\] We first bound the first and the third term. For \(i \in \{1,2\}\), \[\begin{align} &||\mathbb{E}[X|\mathcal{G}_i]-\mathbb{E}_{\mathbb{P}_B}[X|\mathcal{G}_i]||_{L_2(\mathbb{P}_B)} = \frac{1}{\sqrt{\mathbb{P}(B)}} ||(\mathbb{E}[X|\mathcal{G}_i]-\mathbb{E}_{\mathbb{P}_B}[X|\mathcal{G}_i]) \mathbf{1}_{B}||_{L_2(\mathbb{P})} \\ &=\frac{1}{\sqrt{\mathbb{P}(B)}} ||(\mathbb{E}[X|\mathcal{G}_i]-\mathbb{E}[X|\sigma(\mathcal{G}_i \cup \{B\})] ) \mathbf{1}_{B}||_{L_2(\mathbb{P})} \\ &\leq \frac{1}{\sqrt{\mathbb{P}(B)}} ||\mathbb{E}[X|\mathcal{G}_i]-\mathbb{E}[X|\sigma(\mathcal{G}_i \cup \{B\})] ||_{L_2(\mathbb{P})} \\ &\leq \frac{2||X||_{L_\infty(\mathbb{P})}}{\sqrt{\mathbb{P}(B)}} \sqrt{2\delta \mathbb{P}(B)}= 2\sqrt{2\delta}||X||_{L_\infty(\mathbb{P})} , \end{align}\] where the last inequality is due to Lemma 5 and Lemma 7.

For the second term, we use that, by assumption, \(K_{\delta,\varepsilon}(\mathbb{E}_{\mathbb{P}_B}[X|\mathcal{G}_1], B)\) and \(K_{\delta,\varepsilon}(\mathbb{E}_{\mathbb{P}_B}[X|\mathcal{G}_2], B)\), and so we can apply Theorem 17, and get \[||\mathbb{E}_{P_B}[X|\mathcal{G}_1]-\mathbb{E}_{P_B}[X|\mathcal{G}_2]||_{L_2(\mathbb{P}_B)} \leq \sqrt{2\sqrt{\varepsilon} ||X||_{L_2(\mathbb{P}_B)}} \leq \sqrt{2\sqrt{\varepsilon} ||X||_{L_\infty(\mathbb{P})}}.\]

What is left to do is to bring the estimates above together. For this, we use the fact that \((a+b)^2 \leq 2(a^2+b^2)\) for all real numbers \(a,b\), and find \[\begin{align} &\mathbb{E}_{\mathbb{P}_B}[(\mathbb{E}[X|\mathcal{G}_1]-\mathbb{E}[X|\mathcal{G}_2])^2] \leq \left(4\sqrt{2\delta} ||X||_{L_\infty(\mathbb{P})} +\sqrt{2\sqrt{\varepsilon} ||X||_{L_\infty(\mathbb{P})}}\right)^2 \\ &\leq 64\delta ||X||_{L_\infty(\mathbb{P})}^2 + 4 \sqrt{\varepsilon}||X||_{L_\infty(\mathbb{P})}.\qedhere \end{align}\] ◻

Theorem 20 is a local version of Theorem 17 in the following sense: Let both individuals pick an event \(B \in \mathcal{G}_1^\delta \cap \mathcal{G}_2^\delta\) of which they assume that it occurs, and let them calculate their posteriors of a random variable \(X\) given \(B\). If these posteriors, while assuming \(B\), are \(\varepsilon\)-common information, then the \(L_2\)-distance on \(B\) of the true posteriors (without assuming \(B\) as given) is bounded as a function of \(\delta\) and \(\varepsilon\).

5 Dialogues and “almost” convergence of posteriors↩︎

An Aumann-type result for dynamic settings, in which individuals continue to exchange information of their posteriors of a given random variable \(X\), is given by Geanakoplos and Polemarchakis [16] for discrete probability spaces. Nielsen [2] generalizes this result to arbitrary probability spaces.

5.1 Bayesian Dialogues over a random variable↩︎

To model the passing of time and the individuals’ ability to learn new information, following Nielsen, we associate each individual \(i\) with a complete filtration \((\mathcal{G}^i_n)_{n\in \mathbb{N}}\). Furthermore, we set \(\mathcal{G}^i_\infty:= \sigma(\bigcup_{i=1}^\infty \mathcal{G}^i_n)\).

Theorem 21. (Nielsen)Let \(X\in L_1(\Omega)\) and filtrations \((\mathcal{G}_n^1)_{n \in \mathbb{N}}\) and \((\mathcal{G}_n^2)_{n\in \mathbb{N}}\) be given, such that for each \(n\in \mathbb{N}\), there exists an \(m>n\) such that \(\mathbb{E}[X|\mathcal{G}^1_n]\) and \(\mathbb{E}[X|\mathcal{G}^2_n]\) are common information at time \(m\). Then, \[\lim_n \mathbb{E}[X|\mathcal{G}_n^1] = \lim_n \mathbb{E}[X|\mathcal{G}_n^2] \, a.s.\]

In other words, if at each point in time \(n\), for each individual \(i\), the other individual will, at some later point in time, \(m>n\), know all information pertaining to \(X\) that \(i\) knew at time \(n\), then their posteriors converge to the same value, almost surely.

The condition that each individual eventually learns the posterior of the other can be expressed in terms of conditional variances, namely, that for \(i,j \in \{1,2\}\) and \(n \in \mathbb{N}\), there exists \(N_n \in \mathbb{N}\) such that for all \(m >N_n\), we have \(\mathbb{E}[\mathrm{Var}(\mathbb{E}[X|\mathcal{G}^i_n]|\mathcal{G}^j_m)]=0.\)

The following theorem uses this condition to establish a relaxed version of Nielsen’s result: If each individual approximately knows the posterior of the other individual in the limit as time progresses, then their posteriors at infinity are approximately the same.

Theorem 22. Let \(X \in L_2(\Omega)\) and filtrations \((\mathcal{G}^1_n)_{n \in \mathbb{N}}\) and \((\mathcal{G}^2_n)_{n \in \mathbb{N}}\) be given. Assume that for all \(i,j \in \{1,2\}\) \[\begin{align} \label{align:dialoge95epsilon} \lim_{n\to\infty} \lim_{m\to\infty} \mathbb{E}[ \mathrm{Var}(\mathbb{E}[X |\mathcal{G}^i_n]|\mathcal{G}^j_m)] \le \varepsilon. \end{align}\tag{11}\] Then, \[\mathbb{E}[( \mathbb{E}[X|\mathcal{G}^1_\infty] - \mathbb{E}[X| \mathcal{G}^2_\infty] )^2] \le 2\sqrt{\varepsilon\mathrm{Var}(X)}.\]

Proof. Let \(i,j \in \{1,2\}\). By Doob’s martingale convergence theorem, \[\lim_{m\to\infty} \mathbb{E}[ \mathbb{E}[X|\mathcal{G}^i_n]|\mathcal{G}^j_m] = \mathbb{E}[ \mathbb{E}[X|\mathcal{G}^i_n]|\mathcal{G}^j_\infty]\] in \(L_2(\Omega)\), for all \(n \in \mathbb{N}\), and hence, \[\begin{align} \label{equ:dialoge95eps1} \begin{aligned} \lim_{m\to\infty} \mathbb{E}[ \mathrm{Var}(\mathbb{E}[X |\mathcal{G}^i_n]|\mathcal{G}^j_m)] &= \lim_{m\to\infty} || \mathbb{E}[X|\mathcal{G}^i_n] - \mathbb{E}[ \mathbb{E}[X|\mathcal{G}^i_n]|\mathcal{G}^j_m] ||^2_2 \\ &= || \mathbb{E}[X|\mathcal{G}^i_n] - \mathbb{E}[ \mathbb{E}[X|\mathcal{G}^i_n]|\mathcal{G}^j_\infty] ||^2_2. \end{aligned} \end{align}\tag{12}\]

Applying Doob’s martingale inequality once more to \(\mathbb{E}[X|\mathcal{G}^i_n]\), we see \[\begin{align} \label{equ:dialoge95eps2} \lim_{n\to\infty} \mathbb{E}[X|\mathcal{G}^i_n] = \mathbb{E}[X|\mathcal{G}^i_\infty] \end{align}\tag{13}\] in \(L_2(\Omega)\). This, together with Jensen’s inequality (see, for instance, Williams [19]), implies that \[\begin{align} \begin{aligned} &\lim_{n\to\infty} || \mathbb{E}[ \mathbb{E}[X|\mathcal{G}^i_n]|\mathcal{G}^j_\infty] - \mathbb{E}[ \mathbb{E}[X|\mathcal{G}^i_\infty]|\mathcal{G}^j_\infty] ||^2_2 = \lim_{n\to\infty} \mathbb{E}[ (\mathbb{E}[\mathbb{E}[X|\mathcal{G}^i_n] - \mathbb{E}[X|\mathcal{G}^i_\infty] |\mathcal{G}^j_\infty])^2] \\ \leq & \lim_{n\to\infty} \mathbb{E}[ \mathbb{E}[(\mathbb{E}[X|\mathcal{G}^i_n] - \mathbb{E}[X|\mathcal{G}^i_\infty])^2 |\mathcal{G}^j_\infty]] = \lim_{n\to\infty} ||\mathbb{E}[X|\mathcal{G}^i_n] - \mathbb{E}[X|\mathcal{G}^i_\infty] ||^2_2 =0, \end{aligned} \end{align}\] and hence \[\begin{align} \label{equ:dialoge95eps3} \lim_{n\to\infty} \mathbb{E}[ \mathbb{E}[X|\mathcal{G}^i_n]|\mathcal{G}^j_\infty] = \mathbb{E}[ \mathbb{E}[X|\mathcal{G}^i_\infty]|\mathcal{G}^j_\infty] \end{align}\tag{14}\] in \(L_2(\Omega)\). Next, we use inequality 11 , together with equations 12 , 13 and 14 to calculate \[\begin{align} \varepsilon\geq &\lim_{n\to\infty} \lim_{m\to\infty} \mathbb{E}[ \mathrm{Var}(\mathbb{E}[X |\mathcal{G}^i_n]|\mathcal{G}^j_m)] = \lim_{n\to\infty} || \mathbb{E}[X|\mathcal{G}^i_n] - \mathbb{E}[ \mathbb{E}[X|\mathcal{G}^i_n]|\mathcal{G}^j_\infty] ||^2_2 \\ =& || \mathbb{E}[X|\mathcal{G}^i_\infty] - \mathbb{E}[ \mathbb{E}[X|\mathcal{G}^i_\infty]|\mathcal{G}^j_\infty] ||^2_2 = \mathbb{E}[\mathrm{Var}(\mathbb{E}[X|\mathcal{G}^i_\infty]|\mathcal{G}^j_\infty)] , \end{align}\] which means that \(\mathbb{E}[X|\mathcal{G}^i_\infty]\) is \(\varepsilon\)-common information for all \(i \in \{1,2\}\). Applying Theorem 17 to \(\mathbb{E}[X|\mathcal{G}^1_\infty]\) and \(\mathbb{E}[X|\mathcal{G}^2_\infty]\) proves that \[\mathbb{E}[( \mathbb{E}[X|\mathcal{G}^1_\infty] - \mathbb{E}[X|\mathcal{G}^2_\infty] )^2] \le 2\sqrt{\varepsilon\mathrm{Var}{(X)}}. \qedhere\] ◻

If \(\varepsilon=0\), we also get that the posteriors converge a.s.to each other.

Corollary 2. Let \(X \in L_2(\Omega)\) and \((\mathcal{G}^1_n)_{n\in \mathbb{N}}, (\mathcal{G}^2_n)_{n\in \mathbb{N}}\) be filtrations. Furthermore, assume that \[\lim_n \lim_m \mathbb{E}[\mathrm{Var}(\mathbb{E}[X |\mathcal{G}^1_n]|\mathcal{G}^2_m)] =0 \;\text{and } \ \lim_n \lim_m \mathbb{E}[\mathrm{Var}(\mathbb{E}[X |\mathcal{G}^2_n]|\mathcal{G}^1_m)] =0.\] Then, \[\lim_n\mathbb{E}[X|\mathcal{G}^1_n]= \lim_n\mathbb{E}[X|\mathcal{G}^2_n] \, a.s.\]

Proof. By Theorem 22, \[\mathbb{E}[( \mathbb{E}[X|\mathcal{G}^1_\infty]- \mathbb{E}[X|\mathcal{G}^2_\infty] )^2]=0,\] hence \(\mathbb{E}[X|\mathcal{G}^1_\infty] = \mathbb{E}[X|\mathcal{G}^2_\infty]\) a.s. By Doob’s martingale convergence theorem, we know that \[\lim_{n\to\infty} \mathbb{E}[X|\mathcal{G}^1_n] = \mathbb{E}[X|\mathcal{G}^1_\infty] \, a.s.\] Similarly for \((\mathcal{G}^2_n)_{n\in \mathbb{N}}\), and thus \[\lim_{n\to\infty} \mathbb{E}[X|\mathcal{G}^1_n] = \mathbb{E}[X|\mathcal{G}^1_\infty] = \mathbb{E}[X|\mathcal{G}^2_\infty] = \lim_{n\to\infty} \mathbb{E}[X|\mathcal{G}^2_n] \, a.s. \qedhere\] ◻

5.2 Communication with noise↩︎

Theorem 22 has the following important application.

Example 1. Consider communication between two individuals, individual 1 and individual 2, through a noisy channel, where the individuals tell each other their current posteriors of a given random variable \(X\). The noise in the channel is given by the random variables \((\eta^1_n)_{n\in\mathbb{N}}\) and \((\eta^2_n)_{n\in\mathbb{N}}\) with a uniform bound on their variances, \(\mathrm{Var}(\eta^i_n) \leq \varepsilon\) for \(i \in \{1,2\}\), \(n \in \mathbb{N}\) and some \(\varepsilon> 0\).

The individuals start with the \(\sigma\)-algebras \(\mathcal{G}^1_0,\mathcal{G}^2_0\), encoding their knowledge before communication starts, and update their knowledge as follows: \[\mathcal{G}^1_{n+1} = \mathcal{G}_n^1 \vee \sigma ( \mathbb{E}[X|\mathcal{G}^2_n] + \eta_n^2 ),\] and analogously for individual 2. For natural numbers \(n \leq m\), we calculate \[\begin{align} \mathbb{E}[ \mathrm{Var}( \mathbb{E}[X | \mathcal{G}^i_n] | \mathcal{G}^j_{m} )] &\leq \mathbb{E}[ \mathrm{Var}( \mathbb{E}[X | \mathcal{G}^i_n] | \mathcal{G}^j_{n+1} )] \le \mathbb{E}[ \mathrm{Var}( \mathbb{E}[X | \mathcal{G}^i_n] | \mathbb{E}[X|\mathcal{G}^i_n] + \eta_n^i )] \\ &=\mathbb{E}[ \mathrm{Var}(\eta_n^i | \mathbb{E}[X|\mathcal{G}^i_n] + \eta_n^i )] \le \mathrm{Var}(\eta_n^i) \le \varepsilon. \end{align}\] Applying Theorem 22, we see that the difference in the individuals’ posteriors at infinity is controlled by \(\varepsilon\).

Remark 23. In this example, we did not need to assume anything about the shape of the noise, only that its variance is uniformly bounded.

Acknowledgements. We are grateful to Mathias Beiglböck for his comments throughout this work and for initiating the collaboration that resulted in this paper. We also benefited from discussions with Michael Greinecker. C.P. gratefully acknowledges the hospitality of the Faculty of Mathematics at the University of Vienna.

This research was funded in whole or in part by the Austrian Science Fund (FWF) 10.55776/ P34743 and 10.55776/J4981 as well as the Austrian National Bank [Jubiläumsfond, project 18983]. For open access purposes, the author has applied a CC BY public copyright license to any author accepted manuscript version arising from this submission.

References↩︎

[1]
R. J. Aumann. Agreeing to disagree. The Annals of Statistics, 4:1236–1239, 1976.
[2]
L. T. Nielsen. Common knowledge, communication, and convergence of beliefs. Mathematical Social Sciences, 8(1):1–14, 1984.
[3]
A. Billot and V. Vergopoulos. Weak agreement and the properties of beliefs under ambiguity. International Journal of Game Theory, 55, 2026.
[4]
A. Di Tillio, E. Lehrer, and D. Samet. Monologues, dialogues, and common priors. Theoretical Economics, 17(2):587–615, 2022.
[5]
A. Gizatulina and Z. Hellman. No trade and yes trade theorems for heterogeneous priors. Journal of Economic Theory, 182:161–184, 2019.
[6]
Z. Hellman and M. Pintér. Charges and bets: a general characterisation of common priors. International Journal of Game Theory, 51(3-4):567–587, 2022.
[7]
J. Geanakoplos and H. Polemarchakis. Rational dialogues. Revue économique, 74(4):559–568, 2023.
[8]
Y. A. Gonczarowski and Y. Moses. Common knowledge, regained. In Proceedings of the 25th ACM Conference on Economics and Computation, EC ’24, page 208, New York, NY, USA, 2024. Association for Computing Machinery.
[9]
J. Y. Halpern. Reasoning about knowledge: An overview. In Theoretical aspects of reasoning about knowledge, pages 1–17. Elsevier, 1986.
[10]
J. Y. Halpern and Y. Moses. Knowledge and common knowledge in a distributed environment. Journal of the ACM (JACM), 37(3):549–587, 1990.
[11]
R. Fagin, J. Y. Halpern, Y. Moses, and M. Vardi. Reasoning about knowledge. MIT press, 2004.
[12]
A. Rubinstein. The electronic mail game: Strategic behavior under "almost common knowledge". The American Economic Review, 79(3):385–391, 1989.
[13]
D. Monderer and D. Samet. Approximating common knowledge with common beliefs. Games and Economic Behavior, 1(2):170–190, 1989.
[14]
J. Geanakoplos. Common knowledge. Handbook of Game Theory with Economic Applications (Aumann R, Hart S (eds.)), 2:1437–1496, 1994.
[15]
S. Morris. Approximate common knowledge revisited. International Journal of Game Theory, 28:385–408, 1999.
[16]
J. D. Geanakoplos and H. M. Polemarchakis. We can’t disagree forever. Journal of Economic Theory, 28(1):192–200, 1982.
[17]
D. K. Lewis. Convention: A Philosophical Study. Wiley-Blackwell, Cambridge, MA, USA, 1969.
[18]
Z. Neeman. Approximating agreeing to disagree results with common p-beliefs. Games and Economic Behavior, 12(1):162–164, 1996.
[19]
D. Williams. Probability with Martingales. Cambridge Mathematical Textbooks. Cambridge University Press, Cambridge, 1991.
[20]
L. Rogge. Uniform inequalities for conditional expectations. The Annals of Probability, 2(3):486–489, 1974.
[21]
E. S. Boylan. Equiconvergence of martingales. The Annals of Mathematical Statistics, 42(2):552–559, 1971.
[22]
J. Neveu. Note on the tightness of the metric on the set of complete sub-\(\sigma\)-algebras of a probability space. The annals of mathematical statistics, 43(4):1369–1371, 1972.

  1. University Paris–Panthéon–Assas, Laboratoire de Mathématique Économique, and Institut Léon Walras: christina.pawlowitsch@assas-universite.fr↩︎

  2. University of Münster, Institute for Mathematical Stochastics: stefan.schrott@uni-muenster.de↩︎

  3. University of Vienna, Faculty of Mathematics: daniel.toneian@univie.ac.at↩︎

  4. See Lewis [17], p. 52 and following.↩︎

  5. If \(\mathcal{H}\) is a sub \(\sigma\)-algebra of \(\mathcal{G}\), then \(\mathbb{E}[\mathbb{E}[X \mid \mathcal{G}]\mid \mathcal{H} ] = \mathbb{E}[X \mid \mathcal{H}]\), see, for instance, Williams [19].↩︎