January 01, 1970
This paper presents general strong duality results when testing hypotheses by betting against them. A bet is an e-variable for a composite null hypothesis \(\mathcal{P}\): a nonnegative random variable \(X\) whose expected value is at most one under every \({\mathsf P}\in \mathcal{P}\). Following Kelly, Breiman, Cover, Shafer, Grünwald and others, we study a natural minimax log-optimality criterion: given a composite alternative \(\mathcal{Q}\), we characterize the “GROW value” \(\sup_{X} \inf_{{\mathsf Q}} \mathbb{E}_{{\mathsf Q}}[\log X]\). This paper generalizes the results of [1] from (arbitrary \(\mathcal{P}\) and) simple \(\mathcal{Q}\) to arbitrary \(\mathcal{Q}\). We identify a weak-\(*\) joint information projection pair between arbitrary \(\mathcal{P}\) and \(\mathcal{Q}\) that always exists and show that the GROW value for bounded e-variables always equals the relative entropy of this pair, without any restrictions on \(\mathcal{P}\) or \(\mathcal{Q}\). We also prove a similarly general strong duality for the REGROW criterion with bounded e-variables and arbitrary bounded offsets. Under various assumptions our results extend to unbounded e-variables, and examples show that without any assumptions such extensions fail. Our results are analogous to those in [2], swapping tests for bounded e-variables, minimax risk for the GROW criterion, and total variation for relative entropy.
Given a measurable space \((\Omega,\mathcal{F})\), we write \(\mathcal{M}\) for the set of finite signed (countably additive) measures on \(\mathcal{F}\), and \(\mathcal{M}_+\) and \(\mathcal{M}_1\) for the subsets of nonnegative measures and probability measures. Let \(\mathcal{P}\subseteq\mathcal{M}_1\) be an arbitrary composite null hypothesis, and \(\mathcal{Q}\subseteq\mathcal{M}_1\) an arbitrary composite alternative. An e-variable for \(\mathcal{P}\) is a \([0,\infty]\)-valued random variable \(X\) such that \(\mathbb{E}_{\mathsf P}[X] \leq 1\) for all \({\mathsf P}\in \mathcal{P}\). The set of all e-variables for \(\mathcal{P}\) is denoted \(\mathcal{E}\), and the subset of e-variables bounded from above is denoted \(\mathcal{E}_b\).
The influential work of [3] proposed studying the GROW (Growth Rate that is Optimal in the Worst case) value, \[\label{eq:G} \mathsf{G}=\sup_{X\in\mathcal{E}}\inf_{{\mathsf Q}\in\mathcal{Q}} \mathbb{E}_{{\mathsf Q}}[\log X],\tag{1}\] and in particular, establishing duality results that equate \(\mathsf{G}\) to the infimum relative entropy between (appropriate extensions of) \(\mathcal{Q}\) and \(\mathcal{P}\). For example, [1] point out that for simple \(\mathcal{Q}=\{{\mathsf Q}\}\), \(\mathsf{G}\) equals a particular infimum relative entropy between \({\mathsf Q}\) and the effective null hypothesis corresponding to \(\mathcal{P}\). Similarly, for simple \(\mathcal{P}= \{{\mathsf P}\}\), \(\mathsf{G}\) equals the infimum relative entropy between \(\mathcal{Q}\) and \({\mathsf P}\); we present this in Theorem 6.
When \(\mathcal{P}\) and \(\mathcal{Q}\) are both composite, such a “universal” (i.e., without any additional restrictions) strong duality result is not known to hold. For example, the results of [4] imply that if \(\mathcal{P}\) and \(\mathcal{Q}\) have a least favorable distribution pair (a strong assumption!), then \[\mathsf{G}= \inf_{{\mathsf P}\in \mathcal{P}} \inf_{{\mathsf Q}\in \mathcal{Q}} H({\mathsf Q}\mid {\mathsf P}),\] where \(H\) is the Kullback–Leibler divergence, or relative entropy; see [5] for a self-contained proof. [3] provide a different (also strong) sufficient condition under which the same statement holds. Our paper introduces several other conditions under which such characterizations hold; for example, under suitable compactness or domination assumptions on \(\mathcal{P}\) and \(\mathcal{Q}\), many of which are weaker or incomparable to existing ones.
However, our investigations of the above quantity led us to a related notion that remarkably satisfies a strong duality without any restrictions on \(\mathcal{P}\) or \(\mathcal{Q}\). In particular, we obtain a complete characterization of the bounded GROW value, where the supremum is taken over the set \(\mathcal{E}_b\) of all bounded e-variables, \[\label{eq:Gb} \mathsf{G}_b=\sup_{X\in\mathcal{E}_b}\inf_{{\mathsf Q}\in\mathcal{Q}} \mathbb{E}_{{\mathsf Q}}[\log X].\tag{2}\] There are at least three reasons for considering bounded e-variables. First, [2] provide a complete characterization of testable hypotheses, that is, of when there exists a nontrivial test between \(\mathcal{P}\) and \(\mathcal{Q}\); they point out this is equivalent to the existence of bounded e-variables. Second, as we formalize below, when \(\mathcal{Q}\) is a singleton, we have \(\mathsf{G}=\mathsf{G}_b\), so both 1 and 2 may equally be considered generalizations of the singleton case. Third, \(\mathsf{G}_b\) permits a strong duality result without any restrictions, while this appears not to be the case for \(\mathsf{G}\).
The term GROW was originally motivated by the following betting context. We imagine a forecaster claiming that the data are well described by some \({\mathsf P}\in \mathcal{P}\). A skeptic, believing instead that some \({\mathsf Q}\in \mathcal{Q}\) is a better description, would like to bet against the forecaster. The e-variable corresponds to an available bet or wealth multiplier, meaning that the skeptic puts forward (say) one dollar along with a particular e-variable before seeing the data, and the forecaster must pay back the realized value of the e-variable (in dollars) after the data is observed. The forecaster is willing to accept any e-variable as a bet, because under their own belief \(\mathcal{P}\), the Skeptic will lose money. So which bet should the Skeptic choose? A long tradition dating back to [6]–[9] suggests that the Skeptic should maximize the expected logarithm of the wealth under the alternative \({\mathsf Q}\). Picking the worst \({\mathsf Q}\in \mathcal{Q}\) results in the GROW criterion/value. Since \(\mathbb{E}_{{\mathsf Q}}[\log X]\) is also sometimes called the e-power of \(X\) under \({\mathsf Q}\) (see, for example, [10]), one can also refer to \(\mathsf{G}\) as the minimax e-power, but we stick to the former terminology for simplicity.
As mentioned above, a universal strong duality result was recently established for singleton \(\mathcal{Q}=\{{\mathsf Q}\}\) and arbitrary \(\mathcal{P}\) by [1], extending earlier work by [11] and [12]. Defining the effective null generated by \(\mathcal{P}\) as \[\begin{align} \label{eq:260523} \mathcal{P}_\mathrm{eff}=\bigl\{{\mathsf P}\in\mathcal{M}_+:\mathbb{E}_{\mathsf P}[X]\le 1\text{ for all } X\in\mathcal{E}\bigr\}, \end{align}\tag{3}\] that paper showed that under no restrictions whatsoever on \({\mathsf Q}\) or \(\mathcal{P}\), \[\label{eq:numeraire-duality} \sup_{X \in \mathcal{E}} \mathbb{E}_{\mathsf Q}[\log X] = \inf_{{\mathsf P}\in \mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P}),\tag{4}\] and identified a (sub-probability) measure \({\mathsf P}^* \in \mathcal{P}_\mathrm{eff}\) called the reverse information projection such that the right-hand side above equals \(H({\mathsf Q}\mid {\mathsf P}^*)\). One can view our work as completely generalizing these results to the setting of composite \(\mathcal{P}\) and \(\mathcal{Q}\).
To set the stage, we argue in Lemma 3 that for a singleton \({\mathsf Q}\), \[\label{eq:e-eb-ebb-pointQ} \sup_{X \in \mathcal{E}} \mathbb{E}_{\mathsf Q}[\log X] = \sup_{X \in \mathcal{E}_b} \mathbb{E}_{\mathsf Q}[\log X].\tag{5}\] So the left-hand side in 4 could also have used bounded e-variables. Further, by monotone convergence the right-hand side of 4 , which depends on \(\mathcal{E}\) through the definition of the effective null, could also have used bounded e-variables (see 8 ). Unsurprisingly, the story is more complicated for general composite \(\mathcal{Q}\).
Obviously, \(\mathsf{G}\geq \mathsf{G}_b,\) but the inequality can be strict — remarkably, even for a point null \(\mathcal{P}= \{{\mathsf P}\}\) and composite \(\mathcal{Q}\), both supported on the naturals \(\mathbb{N}\) and thus having a common dominating measure; see Example 9. It turns out that \(\mathsf{G}_b\) has a duality theorem without any restrictions.
Our main result, Theorem 3, shows that under no assumptions on \(\mathcal{P}\) or \(\mathcal{Q}\), there always exists a pair \(({\mathsf P}^*,{\mathsf Q}^*) \in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\times \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\) such that \[\label{eq:intro-strong-duality-main} \mathsf{G}_b = \inf_{{\mathsf P}\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})} \inf_{{\mathsf Q}\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})} H({\mathsf Q}\mid {\mathsf P}) = H({\mathsf Q}^* \mid {\mathsf P}^*),\tag{6}\] where \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(S)\) is the closed convex hull of \(S \subseteq \mathcal{M}_1\), where we take the weak-\(*\) closure in \(\operatorname{ba}\), the space of bounded finitely additive measures. (These are standard concepts from topology that will be recapped later.) One can also prove that for any \({\mathsf Q}\in \mathcal{M}_1\), \[\label{eq:kl-equality-costar-peff} \min_{{\mathsf P}\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})} H({\mathsf Q}\mid {\mathsf P}) = \min_{{\mathsf P}\in \mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P}),\tag{7}\] so that 6 , in conjunction with 7 and 5 , implies 4 , justifying our claim that our main result generalizes the work of [1] to arbitrary \(\mathcal{Q}\).
Following the suggestion of [3], we also study a generalization of GROW, called REGROW (relative GROW): \[\sup_{X\in\mathcal{E}_b} \inf_{{\mathsf Q}\in \mathcal{Q}} \left( \mathbb{E}_{\mathsf Q}[\log X] - \xi({\mathsf Q}) \right),\] for some bounded functional \(\xi: \mathcal{M}_1 \to \mathbb{R}\), for example \(\xi({\mathsf Q}) = \sup_X \mathbb{E}_{\mathsf Q}[\log X]\) (the GROW criterion for a singleton \({\mathsf Q}\), called GRO). Theorem 4 establishes a strong duality result for arbitrary \(\mathcal{P},\mathcal{Q},\xi\) that involves the concave biconjugate of \(\xi\).
Next, we examine several interesting special cases where one can use unbounded e-variables.
First, if \(\mathcal{P}= \{{\mathsf P}\}\) is a singleton, and \(\mathcal{Q}\) is convex with \(H({\mathsf Q}\mid {\mathsf P}) < \infty\) for all \({\mathsf Q}\in \mathcal{Q}\), we prove a strong duality result using a generalized information projection of \({\mathsf P}\) onto \(\mathcal{Q}\): \[\mathsf{G}= \inf_{{\mathsf Q}\in \mathcal{Q}} H({\mathsf Q}\mid {\mathsf P}).\] Example 9 shows that despite strong duality holding for both \(\mathsf{G}\) and \(\mathsf{G}_b\), the two quantities are not equal in general. Example 10 shows that if \(H({\mathsf Q}\mid {\mathsf P}) < \infty\) does not hold for all \({\mathsf Q}\in \mathcal{Q}\), then strong duality can fail, even though \({\mathsf Q}\ll {\mathsf P}\) for all \({\mathsf Q}\in \mathcal{Q}\).
We generalize the above result to the case of arbitrary \(\mathcal{P}\), assuming that a joint information projection (JIPr) exists. A JIPr is defined as a pair \(({\mathsf P}^*,{\mathsf Q}^*)\) such that the subprobability \({\mathsf P}^*\) is the reverse information projection of the probability \({\mathsf Q}^*\) onto \(\mathcal{P}_\mathrm{eff}\) and \({\mathsf Q}^*\) is the generalized information projection (in the sense of [13]) of \({\mathsf P}^*\) onto \(\mathcal{Q}\). When a JIPr exists, and (say) \(H({\mathsf Q}\mid {\mathsf P}^*) < \infty\) for all \({\mathsf Q}\in \mathcal{Q}\), we show that \[\mathsf{G}= \inf_{{\mathsf Q}\in \mathcal{Q}}\inf_{{\mathsf P}\in \mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P}),\] and the left-hand side is achieved by the e-variable \(d{\mathsf Q}^*/d{\mathsf P}^*\).
Additionally, if \(\mathcal{Q}\) is convex and compact in the setwise topology, then we show that \[\mathsf{G}= \mathsf{G}_b = \min_{{\mathsf Q}\in \mathcal{Q}} \min_{{\mathsf P}\in \mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P}).\] In the above two cases, if we further assume that (on Polish sample spaces) \(\mathcal{P}\) is convex and weakly compact, we show that \(\mathcal{P}_\mathrm{eff}\) can be replaced by \(\mathcal{P}\).
When working on Polish sample spaces, we show that if \(\mathcal{P}\) and \(\mathcal{Q}\) are both convex and weakly compact, then again \[\mathsf{G}= \mathsf{G}_b = \min_{{\mathsf Q}\in \mathcal{Q}} \min_{{\mathsf P}\in \mathcal{P}} H({\mathsf Q}\mid {\mathsf P}),\] but this time one can also further restrict to bounded continuous e-variables.
Finally, we show that if \(\Omega\) is finite, and \(\mathcal{P},\mathcal{Q}\) are compact, convex sets in the finite-dimensional probability simplex, then \[\mathsf{G}= \mathsf{G}_b = \min_{{\mathsf Q}\in \mathcal{Q}} \min_{{\mathsf P}\in \mathcal{P}} H({\mathsf Q}\mid {\mathsf P}).\] Further the GROW value is finite, attained by a bounded e-variable, which can be represented as the likelihood ratio of JIPr pairs that minimize the above right-hand side.
The first and sixth result above have an intriguing strong parallel to the main results of [2], after swapping e-variables for tests, the GROW criterion for minimax risk, and the relative entropy for total variation distance.
We show a number of other results of independent interest, along with several examples and counterexamples that demonstrate various subtleties in the aforementioned conditions.
The paper is organized as follows. Section 2 recaps some basic definitions and properties, especially of the relative entropy when operating on finitely additive measures. Section 3 presents a strong duality for bounded GROW (and REGROW) without restriction on the sets of measures. Section 4 provides strong duality results for the unbounded GROW criterion, under some relatively mild restrictions. Section 5 discusses some interesting and nontrivial examples that demonstrate certain subtleties or challenges in these strong duality statements. The appendix contains all proofs not presented in the paper.
Recall that the set of e-variables for \(\mathcal{P}\) is \[\mathcal{E}= \left\{\text{all measurable } f \colon \Omega \to [0,\infty] \text{ such that } \int f d\mu \le 1 \text{ for all } \mu \in \mathcal{P}\right\},\] and the set of all bounded e-variables is given by \[\mathcal{E}_b = \left\{f \in \mathcal{E}\colon \sup_{\omega \in \Omega} f(\omega) < \infty\right\} = \mathcal{E}\cap \mathcal{B}_b,\] where \(\mathcal{B}_b\) is the Banach space of all bounded measurable functions \(f:\Omega\to \mathbb{R}\) with the supremum norm \(\|f\|_\infty\).
Since \(f = 1\) is an e-variable, all elements of \(\mathcal{P}_\text{eff}\), defined in 3 , are sub-probabilities. Any sub-probability which is setwise dominated by some element of \(\mathcal{P}_\text{eff}\) also belongs to \(\mathcal{P}_\text{eff}\). Moreover, \(\mathcal{P}\) and \(\mathcal{P}_\textrm{eff}\) have the same nullsets \(A\) since if \(\mu(A) = 0\) for all \(\mu \in \mathcal{P}\), then \(\infty \boldsymbol{1}_A \in \mathcal{E}\), and hence \(\mu(A) = 0\) for all \(\mu \in \mathcal{P}_\textrm{eff}\). Finally, in the definition of \(\mathcal{P}_\textrm{eff}\) it is enough, thanks to the monotone convergence theorem, to let \(X\) range over the set \(\mathcal{E}_b\). Hence we have \[\begin{align} \label{eq:260525} \mathcal{P}_\textrm{eff} = \left\{\mu \in \mathcal{M}_+ \colon \int f d\mu \le 1 \textrm{ for all } f \in \mathcal{E}\right\} = \left\{\mu \in \mathcal{M}_+ \colon \int f d\mu \le 1 \textrm{ for all } f \in \mathcal{E}_b\right\}. \end{align}\tag{8}\]
We will also write \(\operatorname{ba}=\operatorname{ba}(\Omega,\mathcal{F})\) for the Banach space of all bounded finitely additive signed measures on \((\Omega,\mathcal{F})\), endowed with the total variation norm. Its positive cone is denoted by \(\operatorname{ba}_+\). Moreover, \(\operatorname{ba}_1=\{\mu\in\operatorname{ba}_+:\mu(\Omega)=1\}\) denotes the set of all finitely additive probability measures on \((\Omega,\mathcal{F})\). Each \(\mu\in\operatorname{ba}_1\) acts on \(\mathcal{B}_b\) through finitely additive integration; for \(f\in\mathcal{B}_b\), we write \(\mathbb{E}_\mu[f]=\int f\,d\mu\). Thus, \(\operatorname{ba}_1\) may equivalently be viewed as the positive normalized part of the dual of \((\mathcal{B}_b,\|\cdot\|_\infty)\). For \({\mathsf P}\in\operatorname{ba}_1\), we will always let \({\mathsf P}= {\mathsf P}_c + {\mathsf P}_p\) denote its Yosida-Hewitt decomposition, where \({\mathsf P}_c\in\mathcal{M}_+\) is countably additive and \({\mathsf P}_p\) is purely finitely additive. Moreover, let \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*\) denote the weak-\(*\) closed convex hull, that is, for a set \(A\subset \mathcal{M}_1\), \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(A)\) denotes the closure of the convex hull of \(A\), taken in the space \(\operatorname{ba}\) with the topology \(\sigma(\operatorname{ba}, \mathcal{B}_b)\).
Remark 1. Unless specified otherwise, no topology is imposed on \(\Omega\); we work on a general measurable space \((\Omega,\mathcal{F})\). If \(\Omega\) is a topological space and \(\mathcal{F}\) is its Borel \(\sigma\)-algebra, then the usual weak topology on \(\mathcal{M}\) is \(\sigma(\mathcal{M},C_b)\), where \(C_b\) denotes the bounded continuous functions. This topology is coarser than the setwise topology \(\sigma(\mathcal{M},\mathcal{B}_b)\), since \(C_b\subseteq\mathcal{B}_b\). Thus setwise convergence implies weak convergence, but not conversely; for example, on \(\mathbb{R}\), \(\delta_{1/n}\) converges weakly to \(\delta_0\), but not setwise. On the other hand, the total variation topology is finer than \(\sigma(\mathcal{M},\mathcal{B}_b)\). Indeed, \[\|\mu\|_{\mathrm{TV}} = \sup_{\|f\|_\infty\le1} \left|\int f\,d\mu\right|,\] so total variation convergence is uniform convergence over the unit ball of \((\mathcal{B}_b,\|\cdot\|_\infty)\), whereas convergence in \(\sigma(\mathcal{M},\mathcal{B}_b)\) is only pointwise convergence against each fixed \(f\in\mathcal{B}_b\).
For a nonnegative measurable function \(f: \Omega\to [0,\infty]\), and \(\mu\in\operatorname{ba}_1\), we define the extended finitely additive integral as \(\mathbb{E}_\mu[f]=\sup_{n\ge 1}\mathbb{E}_{\mu}[f\wedge n]\in[0,\infty]\); equivalently, \(\mathbb{E}_{\mu}[f]=\sup\{\mathbb{E}_{\mu}[g]:g\in\mathcal{B}_b, 0\le g\le f\}\). Clearly, this definition coincides with the familiar Lebesgue integral whenever \(\mu\in\mathcal{M}_1\subset\operatorname{ba}_1\). Let us adopt the conventions \(\log 0=-\infty\) and \(\log \infty = \infty\). Whenever \(\mathbb{E}_{\mu}[(\log f)^{-}]<\infty\), we set \(\mathbb{E}_{\mu}[(\log f)] = \mathbb{E}_{\mu}[(\log f)^+] - \mathbb{E}_{\mu}[(\log f)^-]\in(-\infty,\infty]\). Whenever \(\mathbb{E}_{\mu}[(\log f)^-]=\infty\), we define \(\mathbb{E}_{\mu}[\log f]=-\infty\). Thus, by convention, this quantity equals \(-\infty\) whenever the negative logarithmic part has infinite finitely additive integral.
The following convention is used when we evaluate the expression \(q \log(q/p)\) for \(p,q \in [0,1]\): \[\label{eq95q95log95q95by95p95convention} q \log\frac{q}{p} = \begin{cases} 0, & q = 0, \\ \infty, & p = 0 \text{ and } q>0. \end{cases}\tag{9}\] This ensures that \(q \log(q/p)\) is jointly convex and lower semicontinuous in \((p,q)\).
Definition 1 (Relative entropy over \(\operatorname{ba}_1\)). We write \(\Pi\) for the set of all finite measurable partitions of \(\Omega\). For any \({\mathsf Q}\in \operatorname{ba}_1\), \({\mathsf P}\in \operatorname{ba}_+\), and finite measurable partition \(\pi = \{A_1,\ldots,A_n\}\) we define, using the convention 9 , \[H_\pi({\mathsf Q}\mid {\mathsf P}) = \sum_{i=1}^n {\mathsf Q}(A_i) \log \frac{{\mathsf Q}(A_i)}{{\mathsf P}(A_i)}.\] Their Kullback–Leibler divergence or relative entropy is then given by \[H({\mathsf Q}\mid {\mathsf P}) = \sup_{\pi \in \Pi} H_\pi({\mathsf Q}\mid {\mathsf P}).\]
Even in \(\operatorname{ba}_+\), zero relative entropy implies equality. All proofs for the statements below in this section are deferred to Appendix 7.
Lemma 1. For \(\mu,\nu\in\operatorname{ba}_1\), \(H(\nu\mid \mu)=0\) if and only if \(\mu=\nu\).
For \({\mathsf Q}\in\mathcal{M}_1\) and \({\mathsf P}\in\mathcal{M}_+\), it can be shown that Definition 1 reduces to the usual definition \[H({\mathsf Q}\mid {\mathsf P})= \begin{cases} \displaystyle \int\log\left(\frac{d{\mathsf Q}}{d{\mathsf P}}\right)\,d{\mathsf Q}, &{\mathsf Q}\ll {\mathsf P},\\ \infty, &{\mathsf Q}\not\ll {\mathsf P}. \end{cases}\]
There is also a Donsker–Varadhan dual representation of relative entropy.
Lemma 2. For \({\mathsf Q}\in \operatorname{ba}_1\) and \({\mathsf P}\in \operatorname{ba}_+\), we have \[H({\mathsf Q}\mid{\mathsf P})=\sup_{g\in\mathcal{B}_b}\left(\int gd{\mathsf Q}-\log\int e^g d{\mathsf P}\right).\]
Proposition 2. For \({\mathsf Q}\in \mathcal{M}_1\) and \({\mathsf P}\in \operatorname{ba}_+\), we have \(H({\mathsf Q}\mid {\mathsf P}) = H({\mathsf Q}\mid {\mathsf P}_c).\) For \({\mathsf Q}\in\operatorname{ba}_{1}\setminus\mathcal{M}_1\) and \({\mathsf P}\in\mathcal{M}_+\), we have \(H({\mathsf Q}\mid {\mathsf P})=\infty\). Thus, for \({\mathsf Q}\in \operatorname{ba}_1\) and \({\mathsf P}\in\mathcal{M}_+\), \(H({\mathsf Q}\mid {\mathsf P}) < \infty\) implies that \({\mathsf Q}\in \mathcal{M}_1\) and \({\mathsf Q}\ll {\mathsf P}\).
Last, we note that for simple alternatives, it is enough to optimize over bounded e-variables.
Lemma 3. For any \({\mathsf Q}\in \mathcal{M}_1\), \(\sup_{h \in \mathcal{E}}\mathbb{E}_{\mathsf Q}[\log h] = \sup_{h \in \mathcal{E}_b} \mathbb{E}_{\mathsf Q}[\log h]\).
Note however that the supremum on the left-hand side is achieved (by the numeraire), but on the right-hand side in general it is not.
This section will present and prove the main strong duality result. The reader may find the following weak duality useful for intuition for the results that follow: \[\begin{align} \sup_{f \in \mathcal{E}} \inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log f] &= \sup_{f \in \mathcal{E}} \inf_{{\mathsf Q}\in \mathop{\mathrm{co}}(\mathcal{Q})} \mathbb{E}_{\mathsf Q}[\log f] \leq \inf_{{\mathsf Q}\in \mathop{\mathrm{co}}(\mathcal{Q})} \sup_{f \in \mathcal{E}} \mathbb{E}_{\mathsf Q}[\log f] \nonumber \\ & = \inf_{\substack{{\mathsf P}\in \mathcal{P}_\mathrm{eff}\\[.1ex] {\mathsf Q}\in \mathop{\mathrm{co}}(\mathcal{Q})}} H({\mathsf Q}\mid {\mathsf P}) \leq \inf_{\substack{{\mathsf P}\in \mathop{\mathrm{co}}(\mathcal{P}) \\[.1ex] {\mathsf Q}\in \mathop{\mathrm{co}}(\mathcal{Q})}} H({\mathsf Q}\mid {\mathsf P}), \label{weak95duality95general} \end{align}\tag{10}\] and note that \(\mathcal{E}\) could have been replaced by \(\mathcal{E}_b\) above. Our aim in this paper is to establish when strong duality holds, i.e., when we have equality in suitably adjusted versions of the above display.
Theorem 3. Let \(\mathcal{P}\) and \(\mathcal{Q}\) be arbitrary nonempty subsets of \(\mathcal{M}_1\), and let \(\mathcal{E}_b\) be the set of bounded e-variables for \(\mathcal{P}\). Then the following strong duality holds: \[\label{eq95strong95duality95general} \sup_{f \in \mathcal{E}_b} \inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log f] = \min_{\substack{{\mathsf P}\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P}) \\[.1ex] {\mathsf Q}\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})}} H({\mathsf Q}\mid {\mathsf P}),\tag{11}\] even if one side, and then also the other side, equals infinity. Further, there exists a pair \(({\mathsf P}^*,{\mathsf Q}^*) \in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\times\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\) that achieves the minimum on the right-hand side.
This theorem follows from Theorem 4 by setting \(\xi=0\) therein (in which case \(\xi^{**}_\mathcal{Q}=0\)).
One may ask whether considering the effective null or closures in \(\operatorname{ba}\) is really necessary. Here we present an example where \(\mathcal{Q}\) is a singleton, \(\mathcal{P}\) is convex, \(\inf_{{\mathsf P}\in \mathcal{P}} H({\mathsf Q}\mid {\mathsf P})\) is achieved in \(\mathcal{P}\) and yet this value is vastly different from \(\inf_{{\mathsf P}\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})} H({\mathsf Q}\mid {\mathsf P})\) and from \(\inf_{{\mathsf P}\in \mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P})\).
Example 1. Let \(\Omega = [0,1]\). Let \(\mathcal{P}\) be the convex hull of \({\mathsf P}_0\) and all \(\{\delta_z\}_{z \in [0,1]}\), where \(\delta_z\) denotes a Dirac delta mass at \(z\) and \({\mathsf P}_0 = \varepsilon \mathsf U + (1-\varepsilon)\delta_0\) for some small constant \(\varepsilon > 0\). Let \({\mathsf Q}= \mathsf U\). Clearly, \(\widetilde{\mathsf P}= {\mathsf P}_0\) achieves the infimum \(\inf_{{\mathsf P}\in \mathcal{P}} H({\mathsf Q}\mid {\mathsf P})\), which can be made arbitrarily large by letting \(\varepsilon\) tend to zero. However, \(\mathcal{P}_\mathrm{eff}\) contains all distributions on \([0,1]\), and thus contains \({\mathsf Q}\), causing \(\inf_{{\mathsf P}\in \mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P}) = 0\). [2] shows that \(\mathcal{P}_\mathrm{eff}\cap \mathcal{M}_1=\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P}) \cap \mathcal{M}_1\) and so \({\mathsf Q}\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\), yielding \(\inf_{{\mathsf P}\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})} H({\mathsf Q}\mid {\mathsf P}) = 0\).
The following example shows that every minimizing pair \(({\mathsf P}^*,{\mathsf Q}^*)\) on the right-hand side of Theorem 3 can in fact be purely finitely additive.
Example 2. Let \(\Omega=\mathbb{N}\) be equipped with its power set \(\sigma\)-field. For each \(n\in\mathbb{N}\), define \[{\mathsf P}_n=\frac{1}{n} \sum_{k=1}^n \delta_k,\quad {\mathsf Q}_n=\frac{1}{n} \sum_{k=2}^{n+1} \delta_k\] and set \(\mathcal{P}=\{{\mathsf P}_n:n\in\mathbb{N}\}\) and \(\mathcal{Q}=\{{\mathsf Q}_n:n\in\mathbb{N}\}\). Then the right-hand side of 11 is equal to zero. Moreover, every minimizing pair \(({\mathsf Q}^*,{\mathsf P}^*) \in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\times\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\) satisfies \({\mathsf Q}^*={\mathsf P}^* \in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\cap\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\). Every such common minimizer is purely finitely additive. In particular, there is no minimizing pair with either \({\mathsf P}^*\in\mathcal{M}_1\) or \({\mathsf Q}^*\in\mathcal{M}_1\). These claims are argued in Appendix 8.
[3] point out that \(\mathsf{G}\) can be too pessimistic as a criterion, as it is a worst case over all alternatives, and thus any (approximately) GROW e-variable may achieve poor e-power against an “easy” \({\mathsf Q}\) in order to preserve optimal e-power against a worst-case \({\mathsf Q}\). To address this, they suggest the \(\operatorname{REGROW}\) criterion, which normalizes the objective by the best achievable growth rate for that alternative. In this sense, a REGROW e-variable attempts to achieve (as much as possible) nearly optimal e-power for every \({\mathsf Q}\in \mathcal{Q}\), and thus is a possibly more appropriate criterion to work with, motivating us to present the following theorem (for general bounded offsets \(\xi\)).
Theorem 4. Follow the same setup as Theorem 3 and let \(\xi:\mathcal{Q}\to\mathbb{R}\) be a bounded function. Define the concave biconjugate \(\xi^{**}_\mathcal{Q}\) of \(\xi\) for each \(\nu \in \operatorname{ba}_+\) by \[\label{eq:260531461} \xi_{\mathcal{Q}}^{\star\star}(\nu)=\inf_{g\in\mathcal{B}_b}\left(\int g\,d\nu - \inf_{{\mathsf Q}\in\mathcal{Q}}\left(\mathbb{E}_{\mathsf Q}[g]-\xi({\mathsf Q})\right)\right).\tag{12}\] Then the following strong duality holds: \[\begin{align} \sup_{f\in\mathcal{E}_{b}}\inf_{{\mathsf Q}\in\mathcal{Q}}(\mathbb{E}_{\mathsf Q}[\log f]-\xi({\mathsf Q})) = \min_{\substack{\mu \in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P}) \\[.1ex] \nu \in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})}} \left(H(\nu\mid\mu)-\xi^{\star\star}_{\mathcal{Q}}(\nu)\right),\label{eq:weak-star-offset-form} \end{align}\tag{13}\] even if one side, and then also the other side, equals infinity. Further, there exists a pair \((\mu^*,\nu^*) \in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\times\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\) that achieves the minimum on the right-hand side.
Before proving this theorem we provide a remark and some corollaries and auxiliary lemmata.
Remark 5. Since the offset function \(\xi\) will typically be convex rather than concave, it cannot be used directly in the minimax argument. We therefore work with its concave biconjugate, given by 12 . Let us write \(\mathsf{G}({\mathsf Q})\) for the quantity in 2 with \(\mathcal{Q}=\{{\mathsf Q}\}\). This is the maximal e-power against the simple alternative \({\mathsf Q}\) and coincides with the \(\operatorname{GRO}\) value in the terminology of [3]. Let us now assume \[\sup_{{\mathsf Q}\in\mathcal{Q}}\mathsf{G}({\mathsf Q})<\infty\] and set \(\xi =\mathsf{G}_b({\mathsf Q})\). Then, the value in 13 presents REGROW (for bounded-variables): \[\sup_{f\in\mathcal{E}_{b}}\inf_{{\mathsf Q}\in\mathcal{Q}}(\mathbb{E}_{\mathsf Q}[\log f]-\mathsf{G}({\mathsf Q})) = \min_{\substack{\mu \in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P}) \\[.1ex] \nu \in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})}} \left(H(\nu\mid\mu)-\mathsf{G}^{\star\star}(\nu)\right) = \min_{\substack{\nu \in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})} } (\mathsf{G}(\nu) - \mathsf{G}^{\star\star}(\nu))\] where the final equality follows by Lemma 6, presented later.
Corollary 1. For any \({\mathsf Q}\in \mathcal{M}_1\), \[\min_{{\mathsf P}\in \mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P}) = \min_{{\mathsf P}\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})} H({\mathsf Q}\mid {\mathsf P}).\]
Proof. We get the chain of equalities \[\min_{{\mathsf P}\in \mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P}) = \sup_{f \in \mathcal{E}} \mathbb{E}_{\mathsf Q}[\log f] = \sup_{f \in \mathcal{E}_b} \mathbb{E}_{\mathsf Q}[\log f] = \min_{{\mathsf P}\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})} H({\mathsf Q}\mid {\mathsf P}),\] where the first equality follows from [1], the second from Lemma 3, and the last by taking \(\mathcal{Q}=\{{\mathsf Q}\}\) in 11 . ◻
There is also an interesting corollary regarding the existence of a nontrivial test, which we define as any test whose worst case power exceeds its worst case level. To set the stage, a test is a measurable function \(\phi\) whose range is \([0,1]\), its worst case type-I error is \(\sup_{{\mathsf P}\in \mathcal{P}} \mathbb{E}_{\mathsf P}[\phi]\) and its worst case power is \(\inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{\mathsf Q}[\phi]\). A classical theorem by [14] (credited also to Le Cam) asserts that if \(\mathcal{P}\cup \mathcal{Q}\) are dominated by a common reference measure, then a nontrivial test exists if and only if \(\mathop{\mathrm{co}}(\mathcal{P})\) and \(\mathop{\mathrm{co}}(\mathcal{Q})\) are separated in the total variation distance. Recently, [1] proved that the aforementioned reference measure assumption can be dropped if \(\mathop{\mathrm{co}}(\mathcal{P})\) and \(\mathop{\mathrm{co}}(\mathcal{Q})\) are replaced by \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\) and \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\), respectively. Our theorem above delivers the same corollary.
Corollary 2. A nontrivial test for \(\mathcal{P}\) against \(\mathcal{Q}\) exists if and only if \(d_{\mathrm{TV}}(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P}),\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q}))>0\).
Proof. The sets \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\) and \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\) are weak-\(\star\) compact, and the map \((\lambda,\rho)\mapsto \|\lambda-\rho\|_{\mathrm{TV}}\) is weak-\(\star\) lower semicontinuous. Hence the total variation distance between the two sets is attained. Therefore this distance is positive if and only if \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\) and \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\) are disjoint.
By [10], a nontrivial test exists if and only if there exists \(f\in\mathcal{E}_b\) such that \(\inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log f] > 0\). Equivalently, the left-hand side of 11 is strictly positive, and thus so is the right-hand side. Since relative entropy between two measures in \(\operatorname{ba}_1\) equals zero if and only they are equal by Lemma 1, this is equivalent to \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\) and \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\) being disjoint, and this concludes the proof. ◻
Before proving Theorem 4, we establish a few facts concerning the concave biconjugate.
Lemma 4. Consider a nonempty set \(\mathcal{Q}\subset \operatorname{ba}_+\) and suppose that \(\xi:\mathcal{Q}\to\mathbb{R}\) is bounded. Define \(\xi^{**}_\mathcal{Q}\) as in 12 . Then \(\xi^{\star\star}_{\mathcal{Q}}\) is weak-\(*\)-upper semicontinuous and concave on \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\), and for every \(g\in\mathcal{B}_b\), \[\label{vcrztkqy} \inf_{\nu\in\mathcal{Q}}\left(\int g\,d\nu-\xi(\nu)\right) = \inf_{\nu\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})} \left( \int g\,d\nu - \xi^{\star\star}_{\mathcal{Q}}(\nu) \right).\tag{14}\]
Proof. For \(g\in\mathcal{B}_b\), define the concave conjugate of \(\xi\in\mathcal{B}_b\) by \[a_{\xi}(g)=\inf_{\nu\in\mathcal{Q}}\left(\int g\,d\nu-\xi(\nu)\right)\] and set \[\psi_g(\nu)=\int g\,d\nu-a_{\xi}(g), \qquad \nu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q}).\] Then \(\psi_g\) is affine and weak-\(*\) continuous on \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\) and we have \(\xi^{\star\star}_{\mathcal{Q}}=\inf_{g\in\mathcal{B}_b}\psi_g\). The pointwise infimum of continuous affine functions is upper semicontinuous and concave. Hence \(\xi^{\star\star}_{\mathcal{Q}}\) is weak-\(*\) upper semicontinuous and concave on \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\).
It remains to prove [eq:offset-transform-inf-preserving]. Fix \(g\in\mathcal{B}_b\). From the definition of \(\xi^{\star\star}_{\mathcal{Q}}\), for every \(\nu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\), \[\label{eq:260530} a_{\xi}(g) \leq \inf_{\nu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})} \left( \int g\,d\nu-\xi^{\star\star}_{\mathcal{Q}}(\nu) \right).\tag{15}\] For the reverse inequality, first note that for every \(\nu_0\in\mathcal{Q}\) and every \(h\in\mathcal{B}_b\), \[a_\xi(h) = \inf_{\nu\in\mathcal{Q}}\left(\int h\,d\nu-\xi(\nu)\right) \le \int h\,d\nu_0-\xi(\nu_0).\] Thus \[\int h\,d\nu_0-a_\xi(h)\ge \xi(\nu_0).\] Taking the infimum over \(h\in\mathcal{B}_b\) gives \(\xi^{\star\star}_{\mathcal{Q}}(\nu_0)\ge \xi(\nu_0)\). Consequently, \[\begin{align} \inf_{\nu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})} \left( \int g\,d\nu-\xi^{\star\star}_{\mathcal{Q}}(\nu) \right) \le \inf_{\nu\in\mathcal{Q}} \left( \int g\,d\nu-\xi^{\star\star}_{\mathcal{Q}}(\nu) \right) \le \inf_{\nu\in\mathcal{Q}} \left( \int g\,d\nu-\xi(\nu) \right) = a_\xi(g). \end{align}\] Combining this with 15 proves [eq:offset-transform-inf-preserving]. ◻
The following short lemma will be useful below.
Lemma 5. Consider the setup of Theorem 4 and let \[\begin{align} \label{eq:Gamma} \Gamma = \left\{ g\in\mathcal{B}_b: \sup_{{\mathsf P}\in\mathcal{P}}\mathbb{E}_{\mathsf P}[e^g]\le 1 \right\}. \end{align}\tag{16}\] Then \(\Gamma\) is convex and \[\sup_{f\in\mathcal{E}_{b}}\inf_{{\mathsf Q}\in\mathcal{Q}}(\mathbb{E}_{\mathsf Q}[\log f]-\xi({\mathsf Q})) = \sup_{g\in\Gamma}\inf_{{\mathsf Q}\in\mathcal{Q}}(\mathbb{E}_{\mathsf Q}[g]-\xi({\mathsf Q})).\]
Proof. We first observe that \(\Gamma\) is convex due to Hölder’s inequality. The inequality “\(\ge\)” is clear. For the reverse inequality, fix \(f\in\mathcal{E}_b\) and \(\varepsilon\in(0,1)\), and set \(f_\varepsilon=\varepsilon+(1-\varepsilon)f \in \mathcal{E}_b\). Then \(g_\varepsilon = \log f_\varepsilon\in\Gamma\). Since \(f_\varepsilon \geq (1-\epsilon)f\), we have \(g_\varepsilon \geq \log (1-\epsilon) + \log f\). Therefore \[\inf_{{\mathsf Q}\in \mathcal{Q}}(\mathbb{E}_{\mathsf Q}[g_\varepsilon] - \xi({\mathsf Q})) \geq \inf_{{\mathsf Q}\in \mathcal{Q}}(\mathbb{E}_{\mathsf Q}[\log f] - \xi({\mathsf Q})) + \log(1-\epsilon).\] Sending \(\varepsilon\) to zero and taking the supremum over \(f \in \mathcal{E}_b\) proves the reverse inequality. ◻
In preparation for the proof of Theorem 4, we also provide a simple-alternative duality. The main differences between this result and 4 are that \(\nu\) is allowed to be in \(\operatorname{ba}_1\) (instead of just \(\mathcal{M}_1\)), and the right-hand side minimizes over \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\) instead of \(\mathcal{P}_\mathrm{eff}\).
Lemma 6. Let \(\Gamma\) be as in 16 . For \(\nu\in\operatorname{ba}_1\), \[\sup_{f\in \mathcal{E}_b}\int \log f\,d\nu = \sup_{g\in \Gamma}\int g\,d\nu =\min_{\mu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})} H(\nu\mid\mu).\]
Proof. By Lemma 5, applied with \(\xi\equiv 0\), it suffices to argue \[\label{eq:single-measure-duality2} \sup_{g\in\Gamma}\int g\,d\nu=\min_{\mu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})} H(\nu\mid\mu).\tag{17}\]
To make headway, for \(\mu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\) and \(g\in\mathbb{B}_b\), define \[\Phi(g,\mu) = \int g\,d\nu-\log\int e^g\,d\mu .\] For each fixed \(\mu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\), the map \(g\mapsto \Phi(g,\mu)\) is concave on \(\mathbb{B}_b\). Indeed, \(g\mapsto \int g\,d\nu\) is affine, while \(g\mapsto \log\int e^g\,d\mu\) is convex by Hölder’s inequality. Moreover, this map is continuous with respect to the supremum norm. For each fixed \(g\in\mathbb{B}_b\), the map \(\mu\mapsto \Phi(g,\mu)\) is convex and continuous on \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\), since \(\mu\mapsto\int e^g\,d\mu\) is affine and continuous, takes values in \((0,\infty)\), and \(x\mapsto -\log x\) is continuous and convex on \((0,\infty)\).
Since \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\) is compact by the Banach–Alaoglu theorem and convex, and \(\mathbb{B}_b\) is convex, Sion’s minimax theorem applies. Together with the variational formula for relative entropy (Lemma 2), this yields \[\begin{align} \min_{\mu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})} H(\nu\mid\mu) &= \min_{\mu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})} \sup_{g\in\mathbb{B}_b}\Phi(g,\mu) = \sup_{g\in\mathbb{B}_b} \min_{\mu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})}\Phi(g,\mu) \notag\\ &= \sup_{g\in\mathbb{B}_b} \left( \int g\,d\nu - \sup_{\mu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})} \log\int e^g\,d\mu \right). \label{eq:single-measure-after-sion} \end{align}\tag{18}\]
For \(g\in\mathcal{B}_b\) set \(c(g)=\sup_{\mu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})}\log\int e^g\,d\mu\). Then we have \(g-c(g)\in\Gamma\), yielding \[\int g\,d\nu-c(g)=\int (g-c(g))\,d\nu\le \sup_{h\in\Gamma}\int h\, d\nu.\] From 18 we get hence get \[\label{eq:single-measure-upper} \min_{\mu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})} H(\nu\mid\mu)\le \sup_{h\in\Gamma}\int h\,d\nu.\tag{19}\] Conversely, suppose that \(h\in\Gamma\). We know that \(c(h)\le 0\). Hence 18 yields \[\int h\,d\nu\le \int h\,d\nu-c(h)\le \sup_{g\in\mathcal{B}_b}\left(\int g\, d\nu-c(g)\right)=\min_{ \mu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})}H(\nu\mid\mu).\] Together with 19 , this yields 17 , concluding the proof. ◻
Proof of Theorem 4. Let \(\Gamma\) be as in 16 . We first note \[\sup_{f\in\mathcal{E}_{b}}\inf_{{\mathsf Q}\in\mathcal{Q}}(\mathbb{E}_{\mathsf Q}[\log f]-\xi({\mathsf Q})) = \sup_{g\in\Gamma}\inf_{{\mathsf Q}\in\mathcal{Q}}(\mathbb{E}_{\mathsf Q}[g]-\xi({\mathsf Q})) = \sup_{g\in\Gamma} \inf_{\nu\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})} \left( \int g\,d\nu - \xi^{\star\star}_{\mathcal{Q}}(\nu) \right)\] by Lemmata 5 and [lem:offset-transform]. We now intend to apply Sion’s minimax theorem. To this end, recall that \(\Gamma\) is convex and \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\) is compact by Lemma 5 and the Banach–Alaoglu theorem, respectively. We next define \[L(g,\nu)=\int gd\nu-\xi^{\star\star}_{\mathcal{Q}}(\nu).\] By Lemma [lem:offset-transform], \(\nu \mapsto L(g,\nu)\) is lower semicontinuous and convex on \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\). Moreover, \(g \mapsto L(g,\nu)\) is linear and continuous. Therefore, Sion’s minimax theorem applies and we obtain \[\begin{align} \sup_{g\in\Gamma} \inf_{\nu\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})} \left( \int g\,d\nu - \xi^{\star\star}_{\mathcal{Q}}(\nu) \right) = \inf_{\nu\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})} \sup_{g\in\Gamma} \left( \int g\,d\nu - \xi^{\star\star}_{\mathcal{Q}}(\nu) \right). \end{align}\] By Lemma 6, we get 13 . Finally, attainment follows from lower semi-continuity and compactness of \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P}) \times \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\). ◻
To provide further intuition for the concave biconjugate, we discuss the simpler case of a finite full simplex.
Example 3. Suppose that \(\Omega=\{1,\ldots,n\}\), and let \({\mathsf P}\in\Delta_n\) have full support, that is, \(p_i={\mathsf P}(\{i\})>0\) for every \(i=1,\ldots,n\). Let \(\mathcal{P}=\{{\mathsf P}\}\) and \(\mathcal{Q}=\Delta_n\). For \({\mathsf Q}\in\mathcal{Q}\), we set again \[\mathsf{G}({\mathsf Q})=H({\mathsf Q}\mid{\mathsf P}).\] We first compute the concave biconjugate of \(\mathsf{G}\) on \(\Delta_n\). Write \(q_i={\mathsf Q}(\{i\})\). Then \[\mathsf{G}({\mathsf Q}) = \sum_{i=1}^n q_i\log\frac{q_i}{p_i}.\] For \(g=(g_1,\ldots,g_n)\in\mathbb{R}^n\), the concave conjugate of \(\mathsf{G}\) is \[a_{\mathsf{G}}(g) = \inf_{q\in\Delta_n} \left( \sum_{i=1}^n q_i(g_i+\log p_i) - \sum_{i=1}^n q_i\log q_i \right).\] Since \(-\sum_i q_i\log q_i\ge0\), this expression is bounded below by \(\min_i(g_i+\log p_i)\), and this lower bound is attained by choosing \(q\) to be a Dirac mass at a minimizer of \(g_i+\log p_i\). Hence \[a_{\mathsf{G}}(g)=\min_{1\le i\le n}(g_i+\log p_i).\] Therefore \[\begin{align} \mathsf{G}^{\star\star}({\mathsf Q}) &= \inf_{g\in\mathbb{R}^n} \left( \sum_{i=1}^n q_i g_i - \min_{1\le i\le n}(g_i+\log p_i) \right) = -\sum_{i=1}^n q_i\log p_i + \inf_{h\in\mathbb{R}^n} \left( \sum_{i=1}^n q_i h_i-\min_i h_i \right), \end{align}\] where \(h_i=g_i+\log p_i\). The last infimum is zero, since \(\sum_i q_i h_i\ge\min_i h_i\) and equality is attained by taking \(h_1=\cdots=h_n\). Thus \[\mathsf{G}^{\star\star}({\mathsf Q}) = -\sum_{i=1}^n {\mathsf Q}(\{i\})\log p_i, \qquad {\mathsf Q}\in\Delta_n.\]
Consequently, \[\mathsf{G}({\mathsf Q})-\mathsf{G}^{\star\star}({\mathsf Q}) = \sum_{i=1}^n q_i\log q_i.\] This quantity is minimized over \(\Delta_n\) at the uniform distribution \(\mathsf U\), and the minimum value is \(-\log n\). Thus the offset relative entropy side of 13 has value \(-\log n\).
On the e-variable side, let \(E^*={d\mathsf U}/{d{\mathsf P}}\). Then \(E^*\) is an e-variable for \({\mathsf P}\), and \(E^*(i)=1/(np_i)\). For \({\mathsf Q}\in\Delta_n\), \[\begin{align} \mathbb{E}_{\mathsf Q}[\log E^*]-\mathsf{G}({\mathsf Q}) &= \sum_{i=1}^n q_i\log\frac{1}{np_i} - \sum_{i=1}^n q_i\log\frac{q_i}{p_i} = -\log n-\sum_{i=1}^n q_i\log q_i. \end{align}\] Taking the infimum over \({\mathsf Q}\in\Delta_n\) gives \(-\log n\), attained at any vertex of \(\Delta_n\). Hence the REGROW value is \(-\log n\), and the REGROW e-variable is \(d\mathsf U/d{\mathsf P}\). Notice that the minimizing \({\mathsf Q}\) on the e-variable side may be any Dirac mass, whereas the minimizing \({\mathsf Q}\) on the offset relative entropy side is the uniform distribution \(\mathsf U\).
While we do not have an “assumption-free” theorem for unbounded GROW, i.e. when using unbounded e-variables, we do identify four important special cases where strong duality does hold. We present these below in the following order, one case per subsection. First, we handle the case of a singleton \(\mathcal{P}= \{{\mathsf P}\}\) together with convex \(\mathcal{Q}\) such that \(H({\mathsf Q}\mid {\mathsf P}) < \infty\) for all \({\mathsf Q}\in\mathcal{Q}\). We then extend this to the case where a joint information projection (defined later) exists. Third, we handle the case of setwise compact and convex \(\mathcal{Q}\). Fourth, we handle the case of weakly compact \(\mathcal{P}\) and \(\mathcal{Q}\) on a Polish space \(\Omega\). Finally, we handle the case of finite \(\Omega\). In all cases, the final conclusions avoid finitely additive measures.
Fix a convex family \(\mathcal{Q}\) of probability measures and a simple null \(\mathcal{P}= \{{\mathsf P}\}\) such that \[\inf_{{\mathsf Q}\in \mathcal{Q}} H({\mathsf Q}\mid {\mathsf P}) < \infty.\] Recall that then Csiszár’s I-projection (IPr), provided it exists, is the probability measure \({\mathsf Q}^{\mathrm{IPr}}\in \mathcal{Q}\) that achieves \(\inf_{{\mathsf Q}\in\mathcal{Q}} H({\mathsf Q}\mid {\mathsf P})\) [15], so that \[H({\mathsf Q}^{\mathrm{IPr}} \mid {\mathsf P}) = \inf_{{\mathsf Q}\in\mathcal{Q}} H({\mathsf Q}\mid {\mathsf P}).\] Csiszár proved that the I-projection of \({\mathsf P}\) onto \(\mathcal{Q}\) exists if \(\mathcal{Q}\) is convex and TV-closed and \(\inf_{{\mathsf Q}\in \mathcal{Q}} H({\mathsf Q}\mid {\mathsf P}) < \infty\), in which case clearly \(H({\mathsf Q}^{\mathrm{IPr}}\mid {\mathsf P}) < \infty\). Since the IPr does not always exist, the following generalization was conceived.
Following [16] and [13], the generalized I-projection (GIPr) \({\mathsf Q}^{\mathrm{GIPr}}\) of \({\mathsf P}\in\mathcal{M}_1\) onto \(\mathcal{Q}\subset\mathcal{M}_1\) is the unique probability measure that satisfies the following Pythagorean inequality for any \({\mathsf Q}\in \mathcal{Q}\): \[\label{eq:pythagorean} H({\mathsf Q}\mid {\mathsf Q}^{\mathrm{GIPr}}) + \inf_{{\mathsf R}\in\mathcal{Q}}H({\mathsf R}\mid {\mathsf P}) \leq H({\mathsf Q}\mid {\mathsf P}).\tag{20}\] The GIPr \({\mathsf Q}^{\mathrm{GIPr}}\) was shown by the above authors to always exist for any convex \(\mathcal{Q}\) with \(\inf_{{\mathsf Q}\in \mathcal{Q}} H({\mathsf Q}\mid {\mathsf P}) < \infty\). However, \({\mathsf Q}^{\mathrm{GIPr}}\notin \mathcal{Q}\) in general, but 20 implies that it does lie in the I-closure of \(\mathcal{Q}\), given by \[\overline{\mathcal{Q}}^I = \left\{{\mathsf R}: \inf_{{\mathsf Q}\in \mathcal{Q}} H({\mathsf Q}\mid {\mathsf R})=0\right\}.\] Further, one has \[\label{eq:existence-LR} H({\mathsf Q}^{\mathrm{GIPr}}\mid {\mathsf P})\le \inf_{{\mathsf Q}\in\mathcal{Q}}H({\mathsf Q}\mid {\mathsf P}),\tag{21}\] where strict inequality can occur as shown by [13]. He also shows that one cannot obtain the generalized I-projection as a minimizer over some closure of \(\mathcal{Q}\). Thanks to 21 , and our standing assumption that \(\inf_{{\mathsf Q}\in \mathcal{Q}} H({\mathsf Q}\mid {\mathsf P}) < \infty\), we have \(H({\mathsf Q}^{\mathrm{GIPr}}\mid {\mathsf P}) < \infty\), and thus \({\mathsf Q}^\mathrm{GIPr} \ll {\mathsf P}\), meaning that \(d{\mathsf Q}^\mathrm{GIPr}/d{\mathsf P}\) is well defined, and it is clearly an e-variable for \({\mathsf P}\). We now show that under some conditions, \(d{\mathsf Q}^\mathrm{GIPr}/d{\mathsf P}\) is the log-optimal (unbounded) e-variable.
Assumption 1. \(\inf_{{\mathsf R}\in \mathcal{Q}} H({\mathsf R}\mid {\mathsf P}) < \infty\) and for every \({\mathsf Q}\in \mathcal{Q}\), one of these conditions holds:
There exists a \({\mathsf P}\)-version \(X^*\) of \(d{\mathsf Q}^\mathrm{GIPR}/d{\mathsf P}\) such that \(\inf_{{\mathsf R}\in\mathcal{Q}}H({\mathsf R}\mid {\mathsf P}) \leq \mathbb{E}_{\mathsf Q}[\log X^*]\).
\(H({\mathsf Q}\mid {\mathsf Q}^{\mathrm{GIPr}}) < \infty\).
\(H({\mathsf Q}\mid {\mathsf P}) < \infty\).
The Pythagorean inequality 20 yields the implication \((2) \implies (1)\). Condition (1) implies \({\mathsf Q}\ll {\mathsf Q}^{\mathrm{GIPr}} \ll {\mathsf P}\). Thus, \((1) \implies (0)\) by rearranging 20 , which one can check is allowed even when \(H({\mathsf Q}\mid {\mathsf P})=\infty\).
Theorem 6 (Singleton null strong duality). For singleton \(\mathcal{P}= \{{\mathsf P}\}\) and convex \(\mathcal{Q}\) such that Assumption 1 holds, the e-variable \(X^* = d{\mathsf Q}^\mathrm{GIPr}/d{\mathsf P}\) achieves the supremum on the left-hand side of the following strong duality: \[\mathsf{G} = \inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X^*] = \inf_{{\mathsf Q}\in\mathcal{Q}}H({\mathsf Q}\mid {\mathsf P}).\] Further, if the GIPr is an IPr (i.e.it lies in \(\mathcal{Q}\)), then \(X^*\) is the \({\mathsf P}\)-a.s.unique GROW e-variable.
Proof. Since condition (0) in Assumption 1 is the weakest one, we prove the result under that condition for each \({\mathsf Q}\in \mathcal{Q}\). Note that then \[\inf_{{\mathsf Q}\in\mathcal{Q}}H({\mathsf Q}\mid {\mathsf P}) \leq \inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X^*] \leq \sup_{X \in \mathcal{E}} \inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X] \leq \inf_{{\mathsf Q}\in\mathcal{Q}}H({\mathsf Q}\mid {\mathsf P})< \infty ,\] where the first inequality follows from Assumption 1, the second one holds by Noting \(X^* \in \mathcal{E}\), the third by weak duality 10 , and the last by assumption. This yields the strong duality. The uniqueness claim follows by Proposition 9. ◻
The above result also corrects some inaccuracies in the statement and proof of Proposition 1 in [17]. In particular, the inequality in their equation (1.11) would only follow if the GIPr belonged to \(\mathcal{Q}\), which need not hold in general. This leads them to conclude in their equation (1.10) that \(\mathsf{G}\) equals \(H({\mathsf Q}^{\mathrm{GIPr}} \mid {\mathsf P})\), which is not true in general. It is worth noting that their assumption is weaker than the usual Pythagorean inequality, because their equation (1.6) is equivalent to assuming our 20 but with the term \(\inf_{{\mathsf R}\in \mathcal{Q}} H({\mathsf R}\mid {\mathsf P})\) being replaced by the quantity \(H({\mathsf Q}^{\mathrm{GIPr}} \mid {\mathsf P})\), which 21 implies can be smaller. Thus the error leads to a weaker assumption but a stronger conclusion. Example 11 provides an explicit counterexample: it satisfies their assumptions but not their conclusion.
It is natural to ask whether we can relax the assumption to only needing \(\inf_{{\mathsf Q}\in\mathcal{Q}}H({\mathsf Q}\mid {\mathsf P}) < \infty\) (which was anyway required for the GIPr to exist). The following example shows that this condition does not suffice, and we can still have \[\label{eq:strict-inequality} \sup_{X \in \mathcal{E}} \inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X] < \inf_{{\mathsf Q}\in\mathcal{Q}}H({\mathsf Q}\mid {\mathsf P}).\tag{22}\]
Example 4. Let \(\Omega=[0,1]\), equipped with the Borel sigma algebra, and consider \({\mathsf P}=\mathsf U\), the uniform. Define \({\mathsf Q}_0\) by \({d{\mathsf Q}_0}/{d{\mathsf P}}(\omega)=2\omega\) so that \({\mathsf Q}_0\) is a probability measure, and set \[\mathcal{Q}=\mathop{\mathrm{co}}(\{{\mathsf Q}_0\}\cup\{\delta_\omega:\omega\in[0,1]\}).\] Let \(X\) be any e-variable with respect to \({\mathsf P}\). Then necessarily \(\inf_{\omega\in[0,1]}X(\omega)\le1\); otherwise \(X>1\) everywhere and hence \(\mathbb{E}_{\mathsf P}[X]>1\). Therefore, the left-hand side of 22 equals zero, achieved by the constant e-variable \(1\). On the other hand, among the elements of \(\mathcal{Q}\) the only one absolutely continuous with respect to \({\mathsf P}\) is \({\mathsf Q}_0\), while \(H({\mathsf Q}_0\mid{\mathsf P}) = \log 2-1/2 >0\). Hence the right-hand side of 22 is finite and strictly positive, being attained by \({\mathsf Q}_0\).
The above example does not have \({\mathsf Q}\ll {\mathsf P}\) for every \({\mathsf Q}\in \mathcal{Q}\). So it leaves open the possibility that assuming \(\inf_{{\mathsf Q}\in\mathcal{Q}}H({\mathsf Q}\mid {\mathsf P}) < \infty\) along with \({\mathsf Q}\ll {\mathsf P}\) for every \({\mathsf Q}\in \mathcal{Q}\) may possibly suffice. However, such hopes are quickly dashed by an extension of the above example, which due to its length is provided later as Example 10.
We end by noting that Assumption 1 is not necessary for strong duality to hold. Rather, it additionally ensures that the supremum over e-variables is attained. The following example emphasizes this point. (See also Example 12 for a more statistical relevant instance.)
Example 5. Let \(\Omega = \mathbb{N}\cup \{0\}\), equipped with its power set \(\sigma\)-algebra, and let \(\mathcal{P}= \{{\mathsf P}\}\) with \({\mathsf P}(n) = 2^{-(n+1)}\) for all \(n \in \mathbb{N}_0\). Define two probability measures \({\mathsf Q}_0 = \delta_0\) and \({\mathsf Q}_\infty(n) = 1/(n(n+1))\) for \(n \in \mathbb{N}\), and set \[\mathcal{Q}= \mathop{\mathrm{co}}({\mathsf Q}_0,{\mathsf Q}_\infty) = \{t{\mathsf Q}_0+(1-t){\mathsf Q}_\infty:t\in[0,1]\}.\] We have \(H({\mathsf Q}_0 \mid {\mathsf P}) = \log 2\), but \(H({\mathsf Q}_\infty \mid {\mathsf P}) = \infty\), so for any \(t \in [0,1)\), \(H(t {\mathsf Q}_0 + (1-t){\mathsf Q}_\infty \mid {\mathsf P}) = \infty\). Consequently, \({\mathsf Q}_0\) is the IPr of \({\mathsf P}\), with relative entropy \(\log 2\). The likelihood ratio \(X^* = d{\mathsf Q}_0/d{\mathsf P}\) is uniquely determined since \({\mathsf P}\) has full support. It satisfies \(X^*(0) = 2\) and \(X^*(n) = 0\) for all \(n \in \mathbb{N}\). Thus, \(\mathbb{E}_{{\mathsf Q}_\infty}[\log X^*] = -\infty\), in particular \(\inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X^*] = -\infty\). To see that strong duality holds, for a small \(\epsilon > 0\), define the e-variable \(E^\epsilon\) by \(E^{\epsilon}(0) = 2(1-\epsilon)\) and \(E^{\epsilon}(n) = \epsilon {\mathsf Q}_\infty(n)/{\mathsf P}(n)\). Further, \(\mathbb{E}_{{\mathsf Q}_\infty}[\log E^\epsilon]=\infty\), but \(\mathbb{E}_{{\mathsf Q}_0}[\log E^\epsilon]=\log(2(1-\epsilon))\). Thus, \(\inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log E^\epsilon] = \log(2(1-\epsilon))\) which approaches \(\log 2\) as \(\epsilon \downarrow 0\), meaning that strong duality holds. (In fact, Proposition 9, proved later, implies that no e-variable achieves strong duality.)
Remark 7. We explain the key idea behind the above example. Let \({\mathsf Q}^*\) denote the generalized I-projection of \({\mathsf P}\) onto some convex \(\mathcal{Q}\) with \(I = \inf_{{\mathsf Q}\in \mathcal{Q}} H({\mathsf Q}\mid {\mathsf P}) < \infty\). Let \(X^*\) be any fixed version of \(d{\mathsf Q}^*/d{\mathsf P}\); clearly \(X^* \in \mathcal{E}\). Call a \({\mathsf Q}\in \mathcal{Q}\) “bad” if \(\mathbb{E}_{\mathsf Q}[\log X^*] < I\), and “good” otherwise. Suppose there exists an e-variable \(Y\) which satisfies \(\mathbb{E}_{\mathsf Q}[\log Y] = \infty\) for every bad \({\mathsf Q}\in \mathcal{Q}\). Then for \(X_\delta = (1-\delta)X^* + \delta Y \in \mathcal{E}\), the mixture with \(Y\) provides protection against the bad \({\mathsf Q}\). Since \(X_\delta \geq \delta Y\), we have \(\mathbb{E}_{\mathsf Q}[\log X_\delta] \geq \mathbb{E}_{\mathsf Q}[\log Y] + \log \delta = \infty\) for bad \({\mathsf Q}\). And for the “good” \({\mathsf Q}\), \(X_\delta \geq (1-\delta)X^*\) gives us \(\mathbb{E}_{\mathsf Q}[\log X_\delta] \geq I + \log(1-\delta)\). Put together, we get that \(\inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{{\mathsf Q}} [\log X_\delta] \geq I + \log(1-\delta)\), and letting \(\delta \downarrow 0\) yields strong duality. We formalize these ideas further in Theorem 15.
[1] showed that for an arbitrary composite null \(\mathcal{P}\) and singleton alternative \({\mathsf Q}\), a log-optimal e-variable \(X^\mathrm{num}\), called the numeraire, always exists and is \({\mathsf Q}\)-almost-surely unique and positive. Then, they define the reverse information projection (RIPr) of \({\mathsf Q}\) onto \(\mathcal{P}_\mathrm{eff}\) as the sub-probability measure defined by \(d{\mathsf P}^*/d{\mathsf Q}= 1/X^\mathrm{num}\), and show that \({\mathsf P}^* \in \mathcal{P}_\mathrm{eff}\).
We define the GIPr (or IPr, if it exists) of a nonzero sub-probability \({\mathsf P}\in \mathcal{M}_+\) onto a convex set \(\mathcal{Q}\), as the GIPr (or IPr) of \(\tilde{\mathsf P}= {\mathsf P}/{\mathsf P}(\Omega)\) onto \(\mathcal{Q}\) (assuming, as before, that \(\inf_{{\mathsf Q}\in \mathcal{Q}} H({\mathsf Q}\mid \tilde{\mathsf P}) < \infty\)). Since the Pythagorean inequality 20 contains \(\tilde{\mathsf P}\) on both sides, one finds that the normalization constant cancels out. Hence the inequality holds for \({\mathsf P}\), despite \({\mathsf P}\) being a sub-probability, a fact that we record as a lemma for easier reference.
Lemma 7. Let \(\mathcal{Q}\) be convex. For any nonzero sub-probability measure \({\mathsf P}\in \mathcal{M}_+\) with \(\inf_{{\mathsf Q}\in \mathcal{Q}} H({\mathsf Q}\mid {\mathsf P}) < \infty\), its GIPr satisfies 20 verbatim, i.e., \[H({\mathsf Q}\mid {\mathsf Q}^{\mathrm{GIPr}}) + \inf_{{\mathsf R}\in\mathcal{Q}}H({\mathsf R}\mid {\mathsf P}) \leq H({\mathsf Q}\mid {\mathsf P}).\]
Now, we can define the joint information projection (JIPr).
Definition 2. Given arbitrary \(\mathcal{P}\) and convex \(\mathcal{Q}\), we will say that a pair \(({\mathsf P}^*,{\mathsf Q}^*) \in \mathcal{M}_+\times\mathcal{M}_1\) is a joint information projection (JIPr) if
\({\mathsf P}^* \in \mathcal{P}_\mathrm{eff}\) is the RIPr of \({\mathsf Q}^*\) onto \(\mathcal{P}_\mathrm{eff}\);
\(\inf_{{\mathsf Q}\in \mathcal{Q}} H({\mathsf Q}\mid {\mathsf P}^*) < \infty\) and \({\mathsf Q}^* \in \overline{\mathcal{Q}}^I\) is the GIPr of \({\mathsf P}^*\) onto \(\mathcal{Q}\).
Also, assuming \((i,ii)\) is weaker than assuming
One has the implication \((iii) \implies (i,ii)\), but in general the reverse implication may fail. We provide an example to illustrate that \((i,ii)\) does not imply \((iii)\), even when the GIPr is an IPr.
Example 6. Let \(\Omega=\{0,1,2\}\) and set \({\mathsf P}^*=\frac{1}{2}\delta_1+\frac{1}{2}\delta_2\) and \({\mathsf Q}^*=\delta_1\). Define the convex sets \[\mathcal{P}= \{(1-s){\mathsf P}^*+s\delta_0:s\in[0,1]\}; \qquad \mathcal{Q} = \{(1-s){\mathsf Q}^*+s\delta_0:s\in[0,1]\}.\] Then the absolutely continuous part of \({\mathsf P}^*\) with respect to \({\mathsf Q}^*\) is \({\mathsf P}^*_a=\frac{1}{2}\delta_1.\) We will show that the pair \(({\mathsf P}^{\star}_a, {\mathsf Q}^{\star})\) satisfies both \((i)\) and \((ii)\). Since the e-variable \(X^* = 2 \mathbf{1}_{\{1\}}\) is the reciprocal of \(d{\mathsf P}_a^* / d{\mathsf Q}^*\), \({\mathsf P}^*_a\) is the RIPr of \({\mathsf Q}^*\) onto \(\mathcal{P}_\mathrm{eff}\) [1], so condition \((i)\) holds. Next, \({\mathsf Q}^*\) is the IPr of \({\mathsf P}^*_a\) onto \(\mathcal{Q}\). Indeed, \(H({\mathsf Q}^*\mid{\mathsf P}^*_a)=\log2<\infty,\) whereas every element \({\mathsf Q}_s = (1-s){\mathsf Q}^*+s\delta_0\) with \(s>0\) assigns positive mass to \(\{0\}\), so \(H({\mathsf Q}_s | {\mathsf P}^*_a) = \infty\) is infinite. Hence condition \((ii)\) holds. However, condition \((iii)\) fails. Indeed, \(\delta_0\in\mathcal{P}\cap \mathcal{Q}\), and therefore \(\inf_{{\mathsf P}\in\mathcal{P}_\mathrm{eff},\;{\mathsf Q}\in\mathcal{Q}}H({\mathsf Q}\mid{\mathsf P})=0.\) On the other hand, \(H({\mathsf Q}^*\mid{\mathsf P}^*_a)=\log2>0.\) Thus \(({\mathsf P}^*_a,{\mathsf Q}^*)\) does not attain \((iii)\), even though conditions \((i)\) and \((ii)\) hold.
Another illustration for (i) and (ii) holding but (iii) failing is given in Example 11. This example is a slightly longer but in addition also satisfies (0) in Assumption 2 below, hence Theorem 8 applies for that example.
Condition \((i)\) implies that \(d{\mathsf Q}^*/d{\mathsf P}^*\) is an e-variable for \(\mathcal{P}\) (it is the numeraire for \(\mathcal{P}\) versus \({\mathsf Q}^*\)). However, it may not achieve the desired strong duality; recall Example 4, where \(({\mathsf P},{\mathsf Q}_0)\) was a JIPr pair satisfying condition \((iii)\). Thus, for strong duality to hold, we need additional conditions, analogous to Assumption 1.
Assumption 2. If a JIPr pair \(({\mathsf P}^*, {\mathsf Q}^*)\) exists, then every \({\mathsf Q}\in \mathcal{Q}\) satisfies one of the following:
there exists a \({\mathsf P}^*\)-version \(X^*\) of \(d{\mathsf Q}^*/d{\mathsf P}^*\) such that \(\inf_{{\mathsf R}\in\mathcal{Q}}H({\mathsf R}\mid {\mathsf P}^*) \leq \mathbb{E}_{\mathsf Q}[\log X^*]\).
\(H({\mathsf Q}\mid {\mathsf Q}^*) < \infty\).
\(H({\mathsf Q}\mid {\mathsf P}^*) < \infty\).
As before, the Pythagorean inequality 20 yields the implication \((2) \implies (1)\). Condition (1) implies \({\mathsf Q}\ll {\mathsf Q}^* \ll {\mathsf P}^*\). Thus, \((1) \implies (0)\) by rearranging 20 , even when \(H({\mathsf Q}\mid {\mathsf P}^*)=\infty\).
Theorem 8 (Strong duality for composite null and alternative). For arbitrary \(\mathcal{P}\) and convex \(\mathcal{Q}\), suppose there exists a joint information projection \(({\mathsf P}^*,{\mathsf Q}^*)\) such that Assumption 2 holds. Then \(X^* = d{\mathsf Q}^*/d{\mathsf P}^*\) is a GROW e-variable, which satisfies the strong duality \[\begin{align} \label{eq:260619} \mathsf{G}= \sup_{X \in \mathcal{E}}\inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{{\mathsf Q}}[\log X] = \inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X^*] = \inf_{{\mathsf Q}\in \mathcal{Q}} H({\mathsf Q}\mid {\mathsf P}^*) = \inf_{{\mathsf Q}\in \mathcal{Q}} \min_{{\mathsf P}\in \mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P}). \end{align}\tag{23}\] Further, if \({\mathsf Q}^*\) is an IPr (i.e.it lies in \(\mathcal{Q}\)), then \(X^*\) is the \({\mathsf P}^*\)-a.s.unique GROW e-variable.
The above theorem eliminates certain conditions in [3], such as assuming that \(H({\mathsf Q}\mid {\mathsf Q}') < \infty\) for all \({\mathsf Q},{\mathsf Q}' \in \mathcal{Q}\), and also that these have full support.
We note above that if the GIPr actually was an IPr (meaning that the infimum relative entropy was achieved by \({\mathsf Q}^* \in \mathcal{Q}\)), then we can add to the strong duality a further equality to \(H({\mathsf Q}^* \mid {\mathsf P}^*)\).
Proof of Theorem 8. Since condition (0) in Assumption 2 is the weakest, we prove the result under that condition. We begin with the weak duality \[\sup_{X\in\mathcal{E}} \inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[\log X] \leq \inf_{{\mathsf Q}\in\mathcal{Q}} \sup_{X\in\mathcal{E}} \mathbb{E}_{\mathsf Q}[\log X] = \inf_{{\mathsf Q}\in\mathcal{Q}} \min_{{\mathsf P}\in \mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid{\mathsf P}) \leq \inf_{{\mathsf Q}\in\mathcal{Q}} H({\mathsf Q}\mid {\mathsf P}^*),\] where the sole equality follows from 4 , and the two inequalities are immediate.
In the opposite direction, we have \[\sup_{X \in \mathcal{E}}\inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[\log X] \geq \inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[\log X^*] \geq \inf_{{\mathsf Q}\in \mathcal{Q}} H({\mathsf Q}\mid {\mathsf P}^*),\] where the first inequality is immediate, and the second follows by taking infimum over \(\mathcal{Q}\) in condition (0). Combining the two displays completes the proof of strong duality. The uniqueness claim follows by Proposition 9. ◻
Example 7. Let \(\mathcal{P}\) be \(\{N(\mu,1): \mu \leq 0\}\) and \(\mathcal{Q}\) be (the convex hull of) \(\{N(\mu,1): \mu \geq 1\}\). Then the JIPr pair is given by \((N(0,1),N(1,1))\) and the GIPr is actually an IPr. In this case, the JIPr pair is also a least favorable distribution pair in the sense of [4]; indeed [5] already applies to yield the GROW e-variable, which is the likelihood ratio of the JIPr pair.
Now take \(\mathcal{P}\) to be set of all \(1\)-sub-Gaussian distributions with nonpositive mean, and \(\mathcal{Q}\) to be (the convex hull of) all \(1\)-sub-Gaussian distributions with mean at least \(1\) and finite relative entropy to the standard Gaussian. For this nonparametric class, a least favorable distribution pair in the sense of Huber–Strassen does not exist, and it is not true that \(H({\mathsf Q}\mid {\mathsf Q}') < \infty\) for all \({\mathsf Q},{\mathsf Q}'\in \mathcal{Q}\) (take, for example, \({\mathsf Q}=\mathsf{U}[1,2],{\mathsf Q}'=\mathsf{U}[3,4]\)), so the results in [3] do not apply. Nevertheless, the JIPr pair is still \((N(0,1),N(1,1))\), and the conditions of Theorem 8 are fulfilled, so the GROW e-variable is still their likelihood ratio.
Proposition 9. Assume that the following strong duality holds: \[\begin{align} \mathsf{G}= \sup_{X \in \mathcal{E}}\inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{{\mathsf Q}}[\log X] = \inf_{{\mathsf Q}\in \mathcal{Q}, {\mathsf P}\in \mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P}) < \infty, \end{align}\] and that there exists \(({\mathsf P}^*,{\mathsf Q}^*) \in \mathcal{P}_\mathrm{eff}\times \mathcal{Q}\), for example a JIPR pair, such that \(H({\mathsf Q}^* \mid {\mathsf P}^*) = \inf_{{\mathsf Q}\in \mathcal{Q}, {\mathsf P}\in \mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P})\). If any e-variable \(X^*\) achieves the supremum above, then we must have \(X^*={d{\mathsf Q}^*}/{d{\mathsf P}^*}\), \({\mathsf P}^*\text{-a.s.}\)
Proof. Since \(\mathbb{E}_{{\mathsf Q}^*}[\log X^*]\ge \mathsf{G}\) because \(X^* \in \mathcal{E}\) witnesses strong duality, applying the Donsker–Varadhan inequality (Lemma 2) yields \[\mathsf{G}\leq \mathbb{E}_{{\mathsf Q}^*}[\log X^*] \leq H({\mathsf Q}^* \mid {\mathsf P}^*) + \log \mathbb{E}_{{\mathsf P}^*}[X^*] \leq H({\mathsf Q}^* \mid {\mathsf P}^*) = \mathsf{G}.\] It follows that equality holds throughout. In particular, \(\log\mathbb{E}_{{\mathsf P}^*}[X^*]=0\), hence \(\mathbb{E}_{{\mathsf P}^*}[X^*]=1\), and \(\mathbb{E}_{{\mathsf Q}^*}[\log X^*] = H({\mathsf Q}^*\mid{\mathsf P}^*)\).
Define now a probability measure \({\mathsf Q}'\) by \(\frac{d{\mathsf Q}'}{d{\mathsf P}^*}=X^*\). Since \(H({\mathsf Q}^*\mid{\mathsf P}^*)<\infty\), we have \({\mathsf Q}^*\ll{\mathsf P}^*\). Therefore, using the chain rule for relative entropy, \[H({\mathsf Q}^*\mid{\mathsf Q}') = H({\mathsf Q}^*\mid{\mathsf P}^*)-\mathbb{E}_{{\mathsf Q}^*}[\log X^*] = 0.\] Thus \({\mathsf Q}^*={\mathsf Q}'\), proving the claim. ◻
Remark 10. If \({\mathsf Q}^*\) in Proposition 9 is not in \(\mathcal{Q}\) but in the I-closure \(\overline{\mathcal{Q}}^I\) only, then the conclusion of the statement still holds provided one assumes in addition \(\mathbb{E}_{{\mathsf Q}^*}[\log X^*]\ge \mathsf{G}\).
We recall Example 5, which points out that Assumption 2 is not necessary. Indeed, this example shows that the JIPr can exist (satisfying the stronger condition (iii)) — it uniquely equals \(({\mathsf Q}_0,\delta_0/2)\) — and strong duality can hold, but the likelihood ratio of the JIPr pair does not achieve the strong duality.
Definition 3. The setwise topology on \(\mathcal{M}\) is the initial topology \(\sigma(\mathcal{M},\mathcal{B}_b)\) induced by the maps \[\mu\mapsto \langle \mu,f\rangle=\int f\,d\mu, \qquad \mu\in\mathcal{M}, f\in\mathcal{B}_b.\] A set \(K\subseteq\mathcal{M}\) is called setwise closed/compact if it is closed/compact in \((\mathcal{M},\sigma(\mathcal{M},\mathcal{B}_b))\).
The setwise topology is the subspace topology induced by the weak-\(*\) topology in \(\operatorname{ba}\): \[\sigma(\mathcal{M},\mathcal{B}_b) = \sigma(\operatorname{ba},\mathcal{B}_b)\big|_{\mathcal{M}}.\] Indeed, \(\sigma(\operatorname{ba},\mathcal{B}_b)\) is the coarsest topology on \(\operatorname{ba}\) making all maps \(\nu\mapsto \int f\,d\nu, f\in\mathcal{B}_b,\) continuous. Restricting these maps to \(\mathcal{M}\subseteq\operatorname{ba}\) gives exactly the maps defining \(\sigma(\mathcal{M},\mathcal{B}_b)\).
If we equip \(\mathcal{M}\) with the total variation (TV) norm \(\|\cdot\|_{\mathrm{TV}}\) and its norm topology, \(\mathrm{TV}\) convergence implies setwise convergence, because \(\|\mu_n - \mu\|_{\mathrm{TV}}\to 0\) implies \(\mu_n\to\mu\) setwise. In addition, if we further suppose that \(\Omega\) is Polish and \(\mathcal{F}\) is the Borel sigma-algebra, then setwise convergence implies weak convergence: clearly, \(\mu_n \to \mu\) in \(\sigma(\mathcal{M},\mathcal{B}_b)\) implies that \(\mu_n\to \mu\) in \(\sigma(\mathcal{M}, C_b)\). All three topologies are Hausdorff.
Recall that \(\mathcal{Q}\subseteq\mathcal{M}_1\) is said to be uniformly absolutely continuous with respect to a finite \(\mu \in \mathcal{M}_+\) if \({\mathsf Q}\ll \mu\) for all \({\mathsf Q}\in\mathcal{Q}\), and for every \(\varepsilon>0\) there exists \(\delta>0\) such that for all \(A\in\mathcal{F}\), \(\mu(A)<\delta \;\Longrightarrow\; \sup_{{\mathsf Q}\in\mathcal{Q}} {\mathsf Q}(A)\le \varepsilon.\) It follows directly from [18] that \(\operatorname{cl}^{\sigma(\mathcal{M},\mathcal{B}_b)}(\operatorname{co}(\mathcal{Q}))\) is setwise compact if and only if there exists a finite measure \(\mu\in\mathcal{M}_+\) such that \(\mathcal{Q}\) is uniformly absolutely continuous with respect to \(\mu\). From this we see that assuming \(\mathcal{Q}\) to be convex and setwise compact is a rather strong assumption.
Two examples of convex and setwise compact \(\mathcal{Q}\) are (i) the convex hull of Gaussians with means in \(\{-1,1\}\) and unit variance and (ii) all Bernoulli distributions with parameters in \([0.2,0.8]\).
We shall use the following two lemmata.
Lemma 8. If \(\mathcal{Q}\) is convex and setwise compact, then \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q}) = \mathcal{Q}\).
Proof. Since the setwise topology on \(\mathcal{M}_1\) is the subspace topology induced by \(\sigma(\operatorname{ba},\mathcal{B}_b)\), the set \(\mathcal{Q}\) is compact also as a subset of \((\operatorname{ba},\sigma(\operatorname{ba},\mathcal{B}_b))\). The latter space is Hausdorff, so compact subsets are closed. Hence \(\mathcal{Q}\) is \(\sigma(\operatorname{ba},\mathcal{B}_b)\)-closed, yielding the statement. ◻
Lemma 9 (Distribution-uniform monotone convergence). If \(\mathcal{Q}\) is setwise compact, then \[\text{f_n \ge 0 and f_n \uparrow f} \quad \Rightarrow \quad \inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{\mathsf Q}[f_n] \uparrow \inf_{{\mathsf Q}\in \mathcal{Q}} \mathbb{E}_{\mathsf Q}[f].\]
Proof of Lemma 9. Set \[a_n=\inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[f_n], \qquad a=\inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[f].\] Then \((a_n)_{n\ge1}\) is increasing and \(a_n\le a\). Let \(L=\lim_n a_n\). Suppose, to the contrary, that \(L<a\). For each \(n\), choose \({\mathsf Q}_n\in\mathcal{Q}\) such that \[\mathbb{E}_{{\mathsf Q}_n}[f_n]\le a_n+\frac{1}{n}.\] By setwise compactness, there is a subnet \(({\mathsf Q}_{n_\alpha})_\alpha\) converging setwise to some \({\mathsf Q}^*\in\mathcal{Q}\). Since \((n_\alpha)_\alpha\) is a subnet of the sequence of indices, we have \(n_\alpha\to\infty\).
Fix now \(k,m\in\mathbb{N}\). For all sufficiently large \(\alpha\), \(n_\alpha\ge k\), and hence \[\mathbb{E}_{{\mathsf Q}_{n_\alpha}}[f_k\wedge m] \le \mathbb{E}_{{\mathsf Q}_{n_\alpha}}[f_{n_\alpha}] \le a_{n_\alpha}+\frac{1}{n_\alpha}.\] Since \(f_k\wedge m\in\mathcal{B}_b\), setwise convergence yields \(\mathbb{E}_{{\mathsf Q}^*}[f_k\wedge m]\le L\). By monotone convergence we get \(\mathbb{E}_{{\mathsf Q}^*}[f_k]\le L\). Letting now \(k\to\infty\) and using monotone convergence once more gives \(\mathbb{E}_{{\mathsf Q}^*}[f]\le L\). Therefore \[a=\inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[f]\le \mathbb{E}_{{\mathsf Q}^*}[f]\le L,\] contradicting \(L<a\). Hence \(L=a\), proving the claim. ◻
With these two lemmata in place we can now prove the following.
Theorem 11. If \(\mathcal{P}\) is arbitrary and \(\mathcal{Q}\) is convex and setwise compact, then \[\mathsf{G}= \mathsf{G}_b =\min_{{\mathsf Q}\in \mathcal{Q}}\min_{{\mathsf P}\in\mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P}).\]
Proof. We first argue that \(\mathsf{G}_b = \min_{{\mathsf Q}\in \mathcal{Q}}\min_{{\mathsf P}\in\mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P}).\) Lemma 8 allows us to replace \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\) in Theorem 3 with just \(\mathcal{Q}\). Corollary 1 shows that for \({\mathsf Q}\in \mathcal{M}_1\), \(\min_{{\mathsf P}\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})} H({\mathsf Q}\mid{\mathsf P})= \min_{{\mathsf P}\in\mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P})\). This proves \(\mathsf{G}_b=\min_{{\mathsf Q}\in \mathcal{Q}}\min_{{\mathsf P}\in\mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P})\).
Clearly \(\mathsf{G}_b\le \mathsf{G}\), so we focus on the reverse inequality. Fix an arbitrary e-variable \(X\). For \(\varepsilon\in(0,1)\) and \(n\in\mathbb{N}\), define \[X_{\varepsilon,n} = \varepsilon+(1-\varepsilon)(X\wedge n), \qquad X_\varepsilon = \varepsilon + (1-\varepsilon) X.\] Then \(X_{\varepsilon,n}\) is strictly positive and \(X_{\varepsilon,n} \in \mathcal{E}_b\). For fixed \(\varepsilon\), we have \(\log X_{\varepsilon,n}\uparrow \log X_\varepsilon\) and \(\log X_{\varepsilon,n}\ge \log \varepsilon\). Hence Lemma 9 (applied to the nonnegative functions \(\log X_{\varepsilon,n} - \log \varepsilon\)) yields \[\mathsf{G}_b \ge \sup_n \inf_{{\mathsf Q}\in\mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X_{\varepsilon,n}] = \inf_{{\mathsf Q}\in\mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X_\varepsilon] \ge \inf_{{\mathsf Q}\in\mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X] + \log(1-\varepsilon).\] Letting \(\varepsilon\downarrow 0\), we get \(\mathsf{G}_b \ge \inf_{{\mathsf Q}\in\mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X].\) Since \(X \in \mathcal{E}\) was arbitrary, we conclude that \(\mathsf{G}_b\ge \mathsf{G}\). ◻
We end this section with a slightly strengthened form of [19]. In preparation, define the solid hull of \(\mathcal{P}\) as \[\operatorname{sol}(\mathcal{P})=\{{\mathsf R}\in \mathcal{M}_+: {\mathsf R}\le {\mathsf P}\text{ for some } {\mathsf P}\in\mathcal{P}\} \subseteq \mathcal{P}_\mathrm{eff}.\]
Proposition 12. Assume \((\Omega, \mathcal{F})\) is a Polish space with its Borel \(\sigma\)-algebra. If \(\mathcal{P}\) is convex and weakly compact, then \(\mathcal{P}_\mathrm{eff}\) equals the solid hull of \(\mathcal{P}\), and thus for any \({\mathsf Q}\in \mathcal{M}_1\), \[\min_{{\mathsf P}\in \mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P}) = \min_{{\mathsf P}\in \mathcal{P}} H({\mathsf Q}\mid {\mathsf P}).\] Thus, in Theorems 8 and 11, if \(\mathcal{P}\) is convex and weakly compact, we can replace \(\mathcal{P}_\mathrm{eff}\) by \(\mathcal{P}\) on the right-hand side.
Proof. To begin with, on a Polish space, the weak topology on probability measures is metrizable and thus the topology on uniformly bounded subprobabilities is metrizable also. Thus, we may work with sequences instead of nets. Weak compactness of \(\mathcal{P}\) implies that \(\operatorname{sol}(\mathcal{P})\) is weakly closed. Indeed, suppose \({\mathsf P}_n \in \operatorname{sol}(\mathcal{P})\) with \({\mathsf P}_n \le {\mathsf P}'_n\) for some \({\mathsf P}'_n \in \mathcal{P}\) and suppose \({\mathsf P}_n \to {\mathsf P}\). By weak compactness of \(\mathcal{P}\), we may pass to a subsequence such that \({\mathsf P}'_n \to {\mathsf P}' \in \mathcal{P}\). Then, \({\mathsf P}_n' - {\mathsf P}_n \to {\mathsf P}' - {\mathsf P}\) and since each \({\mathsf P}_n'-{\mathsf P}_n\) is positive, so is \({\mathsf P}'-{\mathsf P}\). This implies that \({\mathsf P}\le {\mathsf P}'\) and thus \({\mathsf P}\in \operatorname{sol}(\mathcal{P})\). This shows that \(\operatorname{sol}(\mathcal{P})\) is weakly closed.
We next prove \(\operatorname{sol}(\mathcal{P})\supseteq \mathcal{P}_\mathrm{eff}\). To this end, suppose there exists \({\mathsf R}\in \mathcal{P}_\mathrm{eff}\) with \({\mathsf R}\notin \operatorname{sol}(\mathcal{P})\). Now, \(\operatorname{sol}(\mathcal{P})\) being weakly closed and convex, there exists an \(f\in C_b\) such that \[\int f_+ d{\mathsf R}\geq \int f d{\mathsf R}> \sup_{{\mathsf P}\in \operatorname{sol}(\mathcal{P})} \int f d{\mathsf P} = \sup_{{\mathsf P}\in \operatorname{sol}(\mathcal{P})} \int f d{\mathsf P}.\] Solidity yields \[\sup_{{\mathsf P}\in \operatorname{sol}(\mathcal{P})} \int f d{\mathsf P}= \sup_{{\mathsf P}\in\mathcal{P}}\int f_+ d{\mathsf P}.\] Indeed, this is true since for each \({\mathsf P}\in\mathcal{P}\) the supremum of \(\int f d \mu\) over \(\mu \le {\mathsf P}\) is achieved by \(\mu = \mathbf{1}_{f>0}{\mathsf P}\). Hence, we have the existence of \(f \in C_b\) such that \[\int f_+ d{\mathsf R}> \sup_{{\mathsf P}\in \mathcal{P}} \int f_+ d{\mathsf P}.\] Now set \(M=\sup_{{\mathsf P}\in\mathcal{P}} \int f_+ d{\mathsf P}\). If \(M>0\) then \(X = f_+ / M \in \mathcal{E}_b\) with \(\mathbb{E}_{{\mathsf R}}[X]>1\). If \(M=0\), then \(c = \int f_+ d{\mathsf R}>0\) but \(X= 2 f_+ / c \in \mathcal{E}_b\) with \(\mathbb{E}_{{\mathsf R}}[X]=2>1\). This proves that \({\mathsf R}\notin \mathcal{P}_\mathrm{eff}\). We conclude that \(\mathcal{P}_\mathrm{eff}= \operatorname{sol}(\mathcal{P})\).
Now, if \({\mathsf R}\le {\mathsf P}\), then \(H({\mathsf Q}\mid {\mathsf R})\ge H({\mathsf Q}\mid {\mathsf P})\). Moreover, since \(\mathcal{P}\subseteq \mathcal{P}_\mathrm{eff}\), it follows that \(\inf_{{\mathsf R}\in \mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf R}) = \inf_{{\mathsf P}\in\mathcal{P}} H({\mathsf Q}\mid {\mathsf P})\). Finally, since the map \({\mathsf P}\mapsto H({\mathsf Q}\mid {\mathsf P})\) is weakly lower semicontinuous and \(\mathcal{P}\) is weakly compact, we have the last infimum is attained. ◻
In this section, we will prove a countably additive version of our main result by imposing a Polish sample space \(\Omega\) and a weak compactness assumption on \(\mathcal{P}\) and \(\mathcal{Q}\), which allows one to work with continuous e-variables and avoid assuming a common dominating reference measure.
To this end, in this section alone we will assume that \(\Omega\) is a Polish space, equipped with the Borel \(\sigma\)-field. We will equip \(\mathcal{M}_1\) with the weak topology \(\sigma(\mathcal{M}, C_b)\) meaning that a sequence \({\mathsf P}_n\) converges weakly to \({\mathsf P}\) if and only if for every \(h\in C_b\), \(\int h, d{\mathsf P}_n \to \int h\, d{\mathsf P}\). Also, we will write \[\mathcal{E}_{bc} =\mathcal{E}_{bb}\cap C_b,\] where \(\mathcal{E}_{bb}\) is the subset of e-variables that are bounded from above and below. We denote the corresponding GROW values as \(\mathsf{G}_{bb}\) and \(\mathsf{G}_{bc}\).
Theorem 13. Suppose that \(\Omega\) is Polish, equipped with the Borel \(\sigma\)-field, and assume \(\mathcal{P}\) and \(\mathcal{Q}\) are both nonempty, convex, and weakly compact. Then, \[\begin{align} \mathsf{G}= \mathsf{G}_b = \mathsf{G}_{bb} = \mathsf{G}_{bc} =\min_{{\mathsf Q}\in\mathcal{Q}}\min_{{\mathsf P}\in\mathcal{P}} H({\mathsf Q}\mid {\mathsf P}). \end{align}\]
By Theorem 3, \(\mathsf{G}_b\) (and hence \(\mathsf{G},\mathsf{G}_{bc},\mathsf{G}_{bb}\)) also equals \(\min_{{\mathsf Q}\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})}\min_{{\mathsf P}\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})} H({\mathsf Q}\mid {\mathsf P})\). The proof uses the Donsker–Varadhan variational representation on Polish spaces [20], \[\label{eq:Donsker-Varadhan-Cb} H({\mathsf Q}\mid{\mathsf P}) = \sup_{g\in C_b} \left( \mathbb{E}_{\mathsf Q}[g]-\log \mathbb{E}_{\mathsf P}[e^g] \right).\tag{24}\] We first record a robust version over compact convex null classes, obtained by an application of Sion’s minimax theorem.
Lemma 10 (Robust Donsker–Varadhan formula under compact convex null). Suppose that \(\Omega\) is Polish, equipped with its Borel \(\sigma\)-field, and let \(\mathcal{P}\subseteq\mathcal{M}_1\) be nonempty, convex, and weakly compact. Then, for every \({\mathsf Q}\in\mathcal{M}_1\), \[\sup_{g\in C_b} \inf_{{\mathsf P}\in\mathcal{P}} \left( \mathbb{E}_{\mathsf Q}[g]-\log \mathbb{E}_{\mathsf P}[e^g] \right) = \min_{{\mathsf P}\in\mathcal{P}} H({\mathsf Q}\mid{\mathsf P}).\]
Proof. Fix \({\mathsf Q}\in\mathcal{M}_1\) and define, for \(g\in C_b\) and \({\mathsf P}\in\mathcal{P}\), \(\Phi(g,{\mathsf P}) = \mathbb{E}_{\mathsf Q}[g]-\log \mathbb{E}_{\mathsf P}[e^g]\). For fixed \(g\in C_b\), the map \({\mathsf P}\mapsto \Phi(g,{\mathsf P})\) is weakly continuous, because \(e^g\in C_b\). It is also convex in \({\mathsf P}\), since \({\mathsf P}\mapsto \mathbb{E}_{\mathsf P}[e^g]\) is affine and \(x\mapsto-\log x\) is convex on \((0,\infty)\). For fixed \({\mathsf P}\in\mathcal{P}\), the map \(g\mapsto \Phi(g,{\mathsf P})\) is continuous (with respect to the supremum norm) and concave on \(C_b\). Hence Sion’s minimax theorem and 24 yield \[\sup_{g\in C_b}\inf_{{\mathsf P}\in\mathcal{P}}\Phi(g,{\mathsf P}) = \inf_{{\mathsf P}\in\mathcal{P}}\sup_{g\in C_b}\Phi(g,{\mathsf P}) = \inf_{{\mathsf P}\in\mathcal{P}}H({\mathsf Q}\mid{\mathsf P}).\] Finally, the map \({\mathsf P}\mapsto H({\mathsf Q}\mid{\mathsf P})\) is weakly lower semicontinuous, again by 24 , since it is the supremum over \(g\in C_b\) of weakly continuous functions of \({\mathsf P}\). As \(\mathcal{P}\) is weakly compact, the infimum is attained. ◻
Proof of Theorem 13. The inclusions \(\mathcal{E}_{bc}\subseteq\mathcal{E}_{bb}\subseteq\mathcal{E}_b \subseteq \mathcal{E}\), together with 10 , imply \[\mathsf{G}_{bc}\le \mathsf{G}_{bb}\le \mathsf{G}_b\le \mathsf{G}\leq \inf_{{\mathsf Q}\in\mathcal{Q}} \inf_{{\mathsf P}\in\mathcal{P}} H({\mathsf Q}\mid{\mathsf P}).\] It is therefore enough to prove the reverse inequality \[\mathsf{G}_{bc} \ge \inf_{{\mathsf Q}\in\mathcal{Q}} \inf_{{\mathsf P}\in\mathcal{P}} H({\mathsf Q}\mid{\mathsf P}) .\] Indeed, this will imply equality throughout. The attainment of the two infima will follow from weak compactness and weak lower semicontinuity of relative entropy.
To make headway, for \(g\in C_b\), define \[c(g) = \log \sup_{{\mathsf P}\in\mathcal{P}}\mathbb{E}_{\mathsf P}[e^g], \qquad \Phi({\mathsf Q},g) = \mathbb{E}_{\mathsf Q}[g]-c(g), \qquad {\mathsf Q}\in\mathcal{Q}.\] The function \(c\) is finite-valued and continuous. It is also convex since for each \({\mathsf P}\in\mathcal{P}\) the map \(g\mapsto \log \mathbb{E}_{\mathsf P}[e^g]\) is convex and \(c\) is the pointwise supremum of these convex functions. Consequently, for each \({\mathsf Q}\in\mathcal{Q}\), the map \(g\mapsto\Phi({\mathsf Q},g)\) is concave and continuous on \(C_b\). For each \(g\in C_b\), the map \({\mathsf Q}\mapsto\Phi({\mathsf Q},g)\) is affine and weakly continuous on \(\mathcal{Q}\).
We now claim that \[\label{eq:continuous-payoff-reduction} \mathsf{G}_{bc} = \sup_{g\in C_b} \inf_{{\mathsf Q}\in\mathcal{Q}} \Phi({\mathsf Q},g).\tag{25}\] For \(g\in C_b\), define \(X_g = \exp(g-c(g))\). Then \(X_g\) is continuous and bounded from above and below. Moreover, \[\sup_{{\mathsf P}\in\mathcal{P}}\mathbb{E}_{\mathsf P}[X_g] = \exp(-c(g)) \sup_{{\mathsf P}\in\mathcal{P}}\mathbb{E}_{\mathsf P}[e^g] = 1,\] so \(X_g\in\mathcal{E}_{bc}\). Since \(\mathbb{E}_{\mathsf Q}[\log X_g]= \Phi({\mathsf Q},g)\), we obtain \(\mathsf{G}_{bc} \ge \sup_{g\in C_b} \inf_{{\mathsf Q}\in\mathcal{Q}} \Phi({\mathsf Q},g)\). Conversely, let \(X\in\mathcal{E}_{bc}\) and set \(g=\log X\). Then \(g\in C_b\) and \(c(g) = \log\sup_{{\mathsf P}\in\mathcal{P}}\mathbb{E}_{\mathsf P}[X] \le 0\). Hence, for every \({\mathsf Q}\in\mathcal{Q}\), \[\mathbb{E}_{\mathsf Q}[\log X] = \mathbb{E}_{\mathsf Q}[g] \le \mathbb{E}_{\mathsf Q}[g]-c(g) = \Phi({\mathsf Q},g).\] It follows that \(\inf_{{\mathsf Q}\in\mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X] \le \sup_{h\in C_b} \inf_{{\mathsf Q}\in\mathcal{Q}} \Phi({\mathsf Q},h)\). Taking the supremum over \(X\in\mathcal{E}_{bc}\) gives the reverse inequality in 25 .
For fixed \(g\in C_b\), the map \({\mathsf Q}\mapsto\Phi({\mathsf Q},g)\) is affine and weakly continuous. For fixed \({\mathsf Q}\in\mathcal{Q}\), the map \(g\mapsto\Phi({\mathsf Q},g)\) is concave and continuous. Since \(\mathcal{Q}\) is convex and weakly compact, Sion’s theorem gives \[\sup_{g\in C_b}\inf_{{\mathsf Q}\in\mathcal{Q}}\Phi({\mathsf Q},g) = \inf_{{\mathsf Q}\in\mathcal{Q}}\sup_{g\in C_b}\Phi({\mathsf Q},g).\] Combining this with 25 , we obtain \[\begin{align} \mathsf{G}_{bc} &= \inf_{{\mathsf Q}\in\mathcal{Q}} \sup_{g\in C_b} \left( \mathbb{E}_{\mathsf Q}[g] - \log\sup_{{\mathsf P}\in\mathcal{P}}\mathbb{E}_{\mathsf P}[e^g] \right) = \inf_{{\mathsf Q}\in\mathcal{Q}}\inf_{{\mathsf P}\in\mathcal{P}}H({\mathsf Q}\mid{\mathsf P}), \end{align}\] where the last equality follows from Lemma 10. Noting that relative entropy is jointly weakly lower semicontinuous and \(\mathcal{Q}\times \mathcal{P}\) is weakly compact, we conclude that the infimum is attained. ◻
The following corollary gives a simple criterion for when an e-variable witnesses strong duality. Its proof is immediate and hence omitted.
Corollary 3. Assume either the setup of Theorem 11 or of Theorem 13, and let \(({\mathsf P}^*,{\mathsf Q}^*)\) denote any pair that achieves the right-hand side of the strong duality. If \(\mathsf{G}< \infty\), then \(({\mathsf P}^*_a,{\mathsf Q}^*)\) is a joint information projection. (Here, \({\mathsf P}^*_a\) is the absolutely continuous part of \({\mathsf P}^*\) with respect to \({\mathsf Q}^*\).) If we further assume that Assumption 2 holds, then Theorem 8 implies that the e-variable \(X^* = d{\mathsf Q}^*/d{\mathsf P}^*_a\) achieves the supremum in \(\mathsf{G}\).
We also recall Example 5 to remind us that even with a simple null and setwise and weakly compact and convex \(\mathcal{Q}\), strong duality may hold, but the likelihood ratio of the JIPr pair can have an arbitrarily bad worst-case e-power, and no e-variable witnesses the strong duality; see also Example 12. We conclude with an example frequently encountered in the recent e-variable literature.
Example 8 (Bounded means). Let \(\Omega=[0,1]\), let \(\mathcal{P}=\{{\mathsf P}: \mathbb{E}_{{\mathsf P}}[X] \leq a \}\) and \(\mathcal{Q}=\{{\mathsf P}: \mathbb{E}_{{\mathsf P}}[X] \geq b \}\), for some \(0 < a<b < 1\). [21] show that every admissible e-variable is of the form \(1+\lambda(X-a)\) for \(\lambda \in [0,1/a]\), and is bounded by \(1/a\). Therefore, \(\mathsf{G}=\mathsf{G}_b < \infty\) is not surprising. Since \(\mathcal{P}\) and \(\mathcal{Q}\) are both convex and weakly compact, Corollary 3 applies to conclude that a JIPr pair \(({\mathsf P}^*,{\mathsf Q}^*)\) exists (in this case, uniquely): it is in fact just the pair of Bernoulli distributions with parameters \(a,b\). The continuous e-variable obtained by picking \(\lambda^* = (b-a)/(a(1-a))\) is easily checked to be a \({\mathsf P}^*\)-version of their likelihood ratio.
Condition (0) of Assumption 2 can be verified as follows. Since \({\mathsf P}^* = (1-a)\delta_0+a\delta_1\), \(\inf_{{\mathsf R}\in \mathcal{Q}} H({\mathsf R}\mid {\mathsf P}^*)\) is achieved by \({\mathsf Q}^*\). (Indeed, any \({\mathsf R}\) not supported on \(\{0,1\}\) has infinite relative entropy, so we only check amongst Bernoullis.) So we need to check that for every \({\mathsf Q}\in \mathcal{Q}\), \(\mathbb{E}_{\mathsf Q}[\log X^*] \geq H({\mathsf Q}^* \mid {\mathsf P}^*) = b \log (b/a) + (1-b)\log[(1-b)/(1-a)]\). Denote \(\phi(x) = \log X^*(x) = \log(1+ \lambda^*(x-a))\). Since \(\phi\) is concave on \([0,1]\), it lies above its chord between 0 and 1: \(\phi(x) \geq (1-x)\phi(0) + x\phi(1)\). So for any \({\mathsf Q}\in \mathcal{Q}\), \(\mathbb{E}_{\mathsf Q}[\phi(X)] \geq \mathbb{E}_{\mathsf Q}[1-X]\phi(0) + \mathbb{E}_{\mathsf Q}[X]\phi(1)\). Since \(\mathbb{E}_{\mathsf Q}[X] \geq b\) and \(\phi(1) > \phi(0)\), we get \(\mathbb{E}_{\mathsf Q}[\log X^*] \geq (1-b)\phi(0) + b\phi(1) = H({\mathsf Q}^* \mid {\mathsf P}^*)\), as required.
The GROW optimality of \(1+\lambda^*(X-a)\) thus follows from Corollary 3. [22] proved the GROW optimality of this e-variable as well, but it is reassuring that we can derive the same result as a special case of our general theory. To emphasize this point, we present similar calculations for a slightly more complicated problem in Appendix 9, where unbounded data have bounded second moment, and the null and alternative classes differ in their means.
Other natural statistical examples that are convex and weakly compact include relative entropy balls \(\mathcal{P}=\{{\mathsf P}: H({\mathsf P}\mid {\mathsf P}_0) \leq c_0\}\) and \(\mathcal{Q}=\{{\mathsf Q}: H({\mathsf Q}\mid {\mathsf Q}_0) \leq c_1\}\), for some \({\mathsf P}_0,{\mathsf Q}_0\in\mathcal{M}_1\) and \(c_0,c_1 > 0\), or Hellinger balls on compact Polish \(\Omega\).
Below, we consider the case when \(\Omega=\{1,\ldots,d\}\), and identify probability measures on \(\Omega\) with vectors in the simplex \[\Delta_d=\left\{p\in[0,1]^d:\sum_{i=1}^d p_i=1\right\}.\]
Theorem 14 (Finite sample spaces with no null coordinates). Let \(\mathcal{P},\mathcal{Q}\subset\Delta_d\) be convex and compact sets. Assume that \(\mathcal{P}\) is supported on all coordinates, meaning that \(\sup_{{\mathsf P}\in \mathcal{P}} {\mathsf P}_i > 0\) for all \(i \in \{1, \ldots, d\}\). Then \(\mathcal{E}=\mathcal{E}_b\) and \[\mathsf{G}=\mathsf{G}_b = \max_{X\in\mathcal{E}_b} \inf_{{\mathsf Q}\in\mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X] = \min_{{\mathsf Q}\in\mathcal{Q}}\min_{{\mathsf P}\in\mathcal{P}}H({\mathsf Q}\mid {\mathsf P}) < \infty,\] meaning that the GROW value is finite and is attained by some bounded e-variable \(X^*\in\mathcal{E}_b\). Moreover, if \(({\mathsf P}^*,{\mathsf Q}^*)\in\mathcal{P}\times\mathcal{Q}\) minimizes the right-hand side, then \(X^*={d{\mathsf Q}^*}/{d{\mathsf P}^*}\), \({\mathsf P}^*\text{-a.s.}\)
Proof. Any \(X \in \mathcal{E}\) takes only finitely many values, all finite due to the support assumption on \(\mathcal{P}\). This shows \(\mathcal{E}=\mathcal{E}_b\), so \(\mathsf{G}= \mathsf{G}_b\). Next, since \(\Omega\) is finite, weak convergence on \(\Delta_d\) is the same as coordinatewise, hence Euclidean, convergence. Thus the assumed compactness of \(\mathcal{P}\) and \(\mathcal{Q}\) implies weak compactness. Moreover, weak-\(*\) closure agrees with Euclidean closure in finite dimension, and therefore \[\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})=\mathcal{P}, \qquad \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})=\mathcal{Q}.\] Theorem 3 then gives \[\begin{align} \label{eq:260623} \mathsf{G}= \mathsf{G}_b = \min_{{\mathsf Q}\in\mathcal{Q}}\min_{{\mathsf P}\in\mathcal{P}}H({\mathsf Q}\mid {\mathsf P}). \end{align}\tag{26}\] To show that the right-hand side above finite, for each \(i\) choose \({\mathsf P}^{(i)}\in\mathcal{P}\) with \({\mathsf P}_i^{(i)}>0\), and define \(\overline{\mathsf P}=\frac{1}{d}\sum_{i=1}^d {\mathsf P}^{(i)} \in \mathcal{P}.\) by convexity. Then every \({\mathsf Q}\in\mathcal{Q}\) satisfies \({\mathsf Q}\ll \overline{\mathsf P}\), so \(H({\mathsf Q}\mid \overline{\mathsf P})<\infty\).
Finally, we prove that the supremum over e-variables is attained. Note that \[\mathcal{E}\subseteq \prod_{i=1}^d \left[0,\frac{1}{{\mathsf P}_i^{(i)}}\right].\] Moreover, \(\mathcal{E}\) is closed, since it is defined by the closed linear inequalities \[\sum_{i=1}^d {\mathsf P}_iX_i\le1, \qquad {\mathsf P}\in\mathcal{P}.\] Thus \(\mathcal{E}\) is compact. For fixed \({\mathsf Q}\in\mathcal{Q}\), the map \(X\mapsto \mathbb{E}_{\mathsf Q}[\log X]\) is upper semicontinuous, hence so is \(X\mapsto \inf_{{\mathsf Q}\in\mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X]\). Therefore this objective attains its maximum on the compact set \(\mathcal{E}\).
Let \(({\mathsf P}^*,{\mathsf Q}^*)\in\mathcal{P}\times\mathcal{Q}\) now minimize the right-hand side of 26 , and let \(X^*\in\mathcal{E}\) be any GROW optimizer. By Proposition 12 we have \[\min_{{\mathsf Q}\in \mathcal{Q}} \min_{{\mathsf P}\in \mathcal{P}_\mathrm{eff}} H({\mathsf Q}\mid {\mathsf P}) = H({\mathsf Q}^* \mid {\mathsf P}^*).\] Applying Proposition 9 then yields the final claim. ◻
While our main Theorem 3 establishes strong duality for \(\mathcal{E}_b\) without imposing any assumptions on \(\mathcal{P}\) or \(\mathcal{Q}\), the strong duality results for \(\mathcal{E}\) in Section 4 require additional restrictions on these families. A natural question is whether such restrictions are necessary. We answer this in part by presenting a small modification of Example 3.15 in [10], which shows that \(\mathsf{G}>\mathsf{G}_b\) can occur. Notably, the example only involves a singleton null \(\mathcal{P}=\{{\mathsf P}\}\) (which is hence compact in any reasonable topology), and a countable alternative \(\mathcal{Q}=\{{\mathsf Q}_n\}_{n\ge1}\) on \(\mathbb{N}\). In particular, these measures admit a common dominating measure. Thus, even in a very regular setting, the equality \(\mathsf{G}=\mathsf{G}_b\) may fail.
Example 9. Let \(\Omega=\mathbb{N}_0=\mathbb{N}\cup\{0\}\), let \({\mathsf P}=(1/2)\delta_0+(1/2)\delta_1\), where \(\delta_x\) denotes the point mass at \(x\), and let \(\mathcal{Q}=\{{\mathsf Q}_n:n\in\mathbb{N}\}\), where \[{\mathsf Q}_n=\frac{1}{n}\delta_n+\left(1-\frac{1}{n}\right)\delta_0.\] Clearly, the family admits a common dominating reference measure. Moreover, \[\inf_{{\mathsf Q}\in\mathcal{Q}}H({\mathsf Q}\mid{\mathsf P})=H({\mathsf Q}_1\mid{\mathsf P})=\log2<\infty,\] since \({\mathsf Q}_1=\delta_1\), whereas \(H({\mathsf Q}_n\mid{\mathsf P})=\infty\) for all \(n\ge2\). The same remains true after passing to \(\mathop{\mathrm{co}}(\mathcal{Q})\), and hence the (generalized) information projection of \({\mathsf P}\) onto \(\mathcal{Q}\) or \(\mathop{\mathrm{co}}(\mathcal{Q})\) equals \({\mathsf Q}_1\). We further have the following observations.
For \(0<\varepsilon<1\), define \[E_\varepsilon(n) = \exp\left(n\log(2-\varepsilon)-(n-1)\log\varepsilon\right), \qquad n\in\mathbb{N}_0.\] Then \(E_\varepsilon(0)=\varepsilon\) and \(E_\varepsilon(1)=2-\varepsilon\), hence \(\mathbb{E}_{\mathsf P}[E_\varepsilon] = 1\) and \(E_\varepsilon\) is an e-variable for \({\mathsf P}\). Moreover, for every \(n\in \mathbb{N}\), \[\mathbb{E}_{{\mathsf Q}_n}[\log E_\varepsilon] = \frac{1}{n}\left(n\log(2-\varepsilon)-(n-1)\log\varepsilon\right) + \left(1-\frac{1}{n}\right)\log\varepsilon = \log(2-\varepsilon).\] Therefore \(\inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[\log E_\varepsilon] = \log(2-\varepsilon)\), which tends to \(\log 2\) as \(\varepsilon\downarrow0\). By weak duality (recall 10 ), this is the largest possible value. Hence \(\mathsf{G}=\log2\), and the strong duality of Theorem 6 holds.
There does not exist any bounded e-variable \(E\) for which \(\mathbb{E}_{\mathsf Q}[E] > 1\) for all \({\mathsf Q}\in \mathcal{Q}\). Indeed, meeting this requirement for \({\mathsf Q}_1\) alone would imply \(E(1)>1\). To meet this requirement for \({\mathsf Q}_n\), we would need \((1/n) E(n) + (1-1/n) E(0)>1\) for all \(n\in \mathbb{N}\). Since \(E\) is bounded, this then implies that \(E(0)\ge 1\). Together with \(E(1)>1\), we have \(\mathbb{E}_{\mathsf P}[E]>1\), conflicting the fact that \(E\) is an e-variable for \({\mathsf P}\).
Point (ii) above then implies that there is no bounded e-variable \(E\) for which \(\inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[\log E] > 0\). This means that \(\mathsf{G}_b=0\).
Note that here \({\mathsf P}\) does not lie in \(\mathop{\mathrm{co}}(\mathcal{Q})\), but it does lie within its TV-closed convex hull (hence also within its weak-\(*\) closure). To see this, note that by sending \(n\to\infty\), \({\mathsf Q}_n\) converges to \(\delta_0\) in total variation. Averaging \(\delta_0\) with \({\mathsf Q}_1\) yields \({\mathsf P}\). Thus the information projection of \({\mathsf P}\) onto \(\operatorname{co}(\mathcal{Q})\), which equals \({\mathsf Q}_1\), differs from its information projection onto the TV-closure of \(\operatorname{co}(\mathcal{Q})\), which equals \({\mathsf P}\) itself. This reinforces a point made by [13] that the generalized I-projection of \({\mathsf P}\) onto \(\operatorname{co}(\mathcal{Q})\) cannot in general be seen as the I-projection onto some closure of \(\operatorname{co}(\mathcal{Q})\).
The above example shows that strong duality alone does not imply \(\mathsf{G}=\mathsf{G}_b\). Indeed, strong duality holds for \(\mathsf{G}\) by (i), while Theorem 3 gives the corresponding strong-duality representation for \(\mathsf{G}_b\); nevertheless, in this example \(\mathsf{G}>\mathsf{G}_b\). This failure occurs despite substantial structure: the null class is a singleton, the families admit a common dominating reference measure, and the relevant information projection is explicit.
We now provide the promised extension of Example 4, which additionally maintains \({\mathsf Q}\ll {\mathsf P}\) for all \({\mathsf Q}\in \mathcal{Q}\). Here, an IPr exists, the supremum in \(\mathsf{G}\) is achieved, but strong duality fails.
Example 10. As in Example 4, we consider \(\Omega=[0,1]\), equipped with the Borel sigma algebra, \({\mathsf P}=\mathsf U\), and the probability measure \({\mathsf Q}_0\) given by \({d{\mathsf Q}_0}/{d{\mathsf P}}(\omega)=2\omega.\) Let \[\mathcal{I}_\infty = \{{\mathsf Q}\ll {\mathsf P}: H({\mathsf Q}\mid {\mathsf P})=\infty\},\] and define \[\mathcal{Q} = \{{\mathsf Q}_0\} \cup \bigl\{ (1-t){\mathsf Q}_0+t{\mathsf Q} : 0<t\le 1,\; {\mathsf Q}\in\mathcal{I}_\infty \bigr\}.\] Note that \(\mathcal{I}_\infty\) is convex, since positive mixtures of elements of \(\mathcal{I}_\infty\) remain in \(\mathcal{I}_\infty\): if \(H({\mathsf Q}\mid{\mathsf P})=\infty\), then \(H((1-t){\mathsf R}+t{\mathsf Q}\mid{\mathsf P})=\infty\) for every \({\mathsf R}\ll{\mathsf P}\) and \(t\in(0,1]\). Thus \(\mathcal{Q}\) is also convex. By the same token, the only finite-entropy element of \(\mathcal{Q}\) is \({\mathsf Q}_0\), so \[\begin{align} \label{eq:260603461} \inf_{{\mathsf Q}\in\mathcal{Q}} H({\mathsf Q}\mid {\mathsf P}) = H({\mathsf Q}_0\mid {\mathsf P}) = \log 2-\frac{1}{2}. \end{align}\tag{27}\] Moreover, \({\mathsf Q}^{\mathrm{gen}}={\mathsf Q}_0\). Indeed, the Pythagorean inequality 20 holds with equality at \({\mathsf Q}_0\), while for every other element of \(\mathcal{Q}\) the left-hand side is infinite.
Now, let \(X\) be any e-variable for \({\mathsf P}\). Define \[A=\{\omega:X(\omega)\le 1\}.\] As in Example 4 we have \({\mathsf P}(A)>0\). Since \({\mathsf P}\) is nonatomic, there exists a probability measure \({\mathsf Q}_A\ll {\mathsf P}\) supported on \(A\) such that \(H({\mathsf Q}_A\mid {\mathsf P})=\infty\). Consequently \({\mathsf Q}_A\in\mathcal{Q}\). Since \(X\le 1\) \({\mathsf Q}_A\)-almost surely, \(\mathbb{E}_{{\mathsf Q}_A}[\log X] \le 0,\) and therefore \[\inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[\log X] \le \mathbb{E}_{{\mathsf Q}_A}[\log X] \le 0.\] Since this holds for every e-variable \(X\), \(\sup_{X\in\mathcal{E}} \inf_{{\mathsf Q}\in\mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X] \le 0.\) On the other hand, choosing \(X\equiv 1\) yields \(\inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[\log X] = 0.\) Hence \[\sup_{X\in\mathcal{E}} \inf_{{\mathsf Q}\in\mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X] = 0.\] Combining this with 27 , \[\sup_{X\in\mathcal{E}} \inf_{{\mathsf Q}\in\mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X] = 0 < \log 2-\frac{1}{2} = \inf_{{\mathsf Q}\in\mathcal{Q}} H({\mathsf Q}\mid {\mathsf P}).\]
The above example violates strong duality in Theorem 6, but (as expected) the strong duality in Theorem 3 still holds. To see this, first note that \(\sup_{X\in\mathcal{E}} \inf_{{\mathsf Q}\in\mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X] = 0\) implies immediately that \(\sup_{X\in\mathcal{E}_b} \inf_{{\mathsf Q}\in\mathcal{Q}} \mathbb{E}_{\mathsf Q}[\log X] = 0\), both achieved by \(X\equiv1\). But the weak-\(*\) closure \(\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\) contains \({\mathsf P}\), so the right-hand side of Theorem 3 also equals zero.
We now revisit Example 3.2 from [13], which makes the point that \(\inf_{{\mathsf Q}\in\mathcal{Q}}H({\mathsf Q}\mid {\mathsf P}) > H({\mathsf Q}^* \mid {\mathsf P})\) for the GIPr \({\mathsf Q}^*\). We now complete the duality picture for this example. We show that strong duality for \(\mathsf{G}\) holds, but despite having a common dominating (Lebesgue) measure and a singleton \(\mathcal{P}=\{{\mathsf P}\}\), \(\mathsf{G}\) may not equal \(H({\mathsf Q}\mid {\mathsf P})\) for any \({\mathsf Q}\in \mathcal{Q}\), nor equal \(H({\mathsf Q}^* \mid {\mathsf P})\).
Example 6 already shows that it is possible to satisfy conditions (i) and (ii) in the JIPr Definition 2 without satisfying condition (iii). The following example goes one step further. It shows that even if Assumption 1 is satisfied, it is possible to satisfy conditions (i) and (ii) in the JIPr definition 2 without satisfying condition (iii).
Example 11. Let \(\mathcal{Q}_a = \{{\mathsf Q}: \int s d{\mathsf Q}(s) \geq a\}\) for some \(a > 2/3\), which is convex. Let \({\mathsf P}\) be defined by the density \(p(s) = \exp(-s)(s+4)/(s+1)^4\) for \(s\geq 0\), which has mean \(\approx 0.298 < 2/3\), and thus lies outside \(\mathcal{Q}_a\). Csiszár proves that the GIPr of \({\mathsf P}\) onto \(\mathcal{Q}_a\) is a distribution \({\mathsf Q}^*\) with density \(q^*(s) = \frac{2}{3} (s+4)/(s+1)^4\) for \(s\geq 0\), which remarkably does not depend on \(a\), and its mean is \(2/3\), so \({\mathsf Q}^* \notin \mathcal{Q}_a\). We have the following facts.
For any \({\mathsf Q}\in\mathcal{\mathsf Q}_a\) with \({\mathsf Q}\ll {\mathsf Q}^*\), the chain rule gives \[H({\mathsf Q}\mid {\mathsf P}) = H({\mathsf Q}\mid {\mathsf Q}^*)+ \mathbb{E}_{\mathsf Q}\left[\log\frac{d{\mathsf Q}^*}{d{\mathsf P}}\right] = H({\mathsf Q}\mid {\mathsf Q}^*)+ \int s d {\mathsf Q}(s) +\log\frac{2}{3} \ge a+\log\frac{2}{3}.\] If \({\mathsf Q}\not\ll {\mathsf Q}^*\), then \(H({\mathsf Q}\mid {\mathsf Q}^*)=\infty\), and the same lower bound is trivial. Put together, \(\inf_{{\mathsf Q}\in\mathcal{Q}_a}H({\mathsf Q}\mid {\mathsf P}) \ge a+\log\frac{2}{3}.\) In fact, equality holds, as justified in Appendix 10.
Since \[\frac{q^*(s)}{p(s)} = \frac{\frac{2}{3} (s+4)/(s+1)^4}{e^{-s}(s+4)/(s+1)^4} = \frac{2}{3} e^s, \qquad s\geq 0,\] we choose the Borel version of the Radon–Nikodym derivative \(d{\mathsf Q}^*/d{\mathsf P}\) given by \(X^*(s)=\frac{2}{3} e^s\). Then \(X^*=\frac{d{\mathsf Q}^*}{d{\mathsf P}}\), \({\mathsf P}\text{-a.s.}\), hence \(X^*\in\mathcal{E}\). For every \({\mathsf Q}\in \mathcal{Q}_a\), \[\mathbb{E}_{\mathsf Q}[\log X^*] = \int s d {\mathsf Q}(s)+\log\frac{2}{3} \ge a+\log\frac{2}{3} = \inf_{{\mathsf R}\in\mathcal{Q}_a}H({\mathsf R}|{\mathsf P}) .\] Thus, condition (0) in Assumption 1 is satisfied, and, by Theorem 6, strong duality holds and \(X^*\) is GROW optimal.
\(H({\mathsf Q}^* \mid {\mathsf P}) =2/3 - \log(3/2) < a + \log(2/3)\).
To summarize, we have \[a + \log\frac{2}{3}= \sup_{X\in\mathcal{E}}\inf_{{\mathsf Q}\in\mathcal{Q}_a} \mathbb{E}_{{\mathsf Q}}[\log X] = \inf_{{\mathsf Q}\in \mathcal{Q}_a} H({\mathsf Q}\mid {\mathsf P}) > H({\mathsf Q}^* \mid {\mathsf P}),\] meaning that strong duality holds, and the left-hand side supremum is achieved by the e-variable \(X^* = d{\mathsf Q}^*/d{\mathsf P}\). Furthermore, \(\inf_{{\mathsf Q}\in\mathcal{Q}_a} \mathbb{E}_{{\mathsf Q}}[\log X^*]\) is not achieved by \({\mathsf Q}^*\), but it is achieved by any \({\mathsf Q}\in \mathcal{Q}_a\) with mean equal to \(a\). Finally, \(({\mathsf P},{\mathsf Q}^*)\) also forms a JIPr pair satisfying conditions \((i),(ii)\) but not condition \((iii)\) in Definition 2.
We now formalize and generalize Remark 7, which allows us to extend Example 5 to a more statistically meaningful context.
Theorem 15 (Shielding theorem). Suppose that \(\mathcal{Q}\) is partitioned as \(\mathcal{Q}=\mathcal{Q}_I\cup \mathcal{Q}_\infty,\) where \(\mathcal{Q}_I\) is nonempty. Assume that strong duality holds over \(\mathcal{Q}_I\) with a finite value, i.e., \[I = \sup_{X\in\mathcal{E}} \inf_{{\mathsf Q}\in\mathcal{Q}_I}\mathbb{E}_{\mathsf Q}[\log X] = \inf_{{\mathsf Q}\in\mathcal{Q}_I} \inf_{P\in\mathcal{P}_\mathrm{eff}}H({\mathsf Q}\mid {\mathsf P}) <\infty.\] Assume further that there exists \(Y\in\mathcal{E}\) such that \[\label{eq:shield-assumption} \mathbb{E}_{\mathsf Q}[\log Y]=\infty \qquad \text{for every }{\mathsf Q}\in\mathcal{Q}_\infty.\tag{28}\] Then strong duality holds over the full alternative \(\mathcal{Q}\), i.e., \[\label{eq:shielded-strong-duality} I = \sup_{X\in\mathcal{E}} \inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[\log X] = \inf_{{\mathsf Q}\in\mathcal{Q}} \inf_{{\mathsf P}\in\mathcal{P}_\mathrm{eff}}H({\mathsf Q}\mid {\mathsf P}).\tag{29}\]
Appendix 10 provides the proof of Theorem 15.
Example 12 (Gaussian mean shift with heavy-tailed outliers). Let \(\Omega=\mathbb{R}\) and let \(\mathcal{P}=\{{\mathsf P}\}\) with \({\mathsf P}=N(0,1)\). Fix \(\theta>0\) and let \({\mathsf Q}_0=N(\theta,1).\) The likelihood ratio of \({\mathsf Q}_0\) with respect to \({\mathsf P}\) is \(L(x)= \exp\left(\theta x-\frac{\theta^2}{2}\right),\) which is an e-variable with \(\mathbb{E}_{{\mathsf Q}_0}[\log L] = H({\mathsf Q}_0\mid {\mathsf P}) = \frac{\theta^2}{2}.\) Thus strong duality holds for the singleton alternative \(\mathcal{Q}_I=\{{\mathsf Q}_0\},\) with value \(I=\frac{\theta^2}{2}\). Now let \({\mathsf R}\) be a centered Student-\(t_\nu\) distribution with \(1<\nu<2\). For \(t\in[0,1]\), define \({\mathsf Q}_t=(1-t){\mathsf Q}_0+t{\mathsf R},\) and set \[\mathcal{Q} = \{{\mathsf Q}_t:0\le t\le1\}, \qquad \mathcal{Q}_\infty = \{{\mathsf Q}_t:0<t\le1\}.\] This gives the partition \(\mathcal{Q}=\mathcal{Q}_I\cup\mathcal{Q}_\infty,\) and we note that \(\mathcal{Q}\) is both setwise and weakly compact, so Theorem 11 already gives strong duality, but we show below that it is not attained by any e-variable, and shielding helps verify this strong duality.
Choose now \(a\in(0,1/2)\) and define the random variable \[Y_a(\omega) = \sqrt{1-2a}\exp(a\omega^2).\] Since \(\mathbb{E}_{\mathsf P}[Y_a]=1\), we get \(Y_a \in \mathcal{E}\). Moreover, \(\log Y_a(\omega) = \frac{1}{2}\log(1-2a)+a\omega^2\). Therefore, for every \(t>0\), \[\mathbb{E}_{{\mathsf Q}_t}[\log Y_a] = (1-t)\mathbb{E}_{{\mathsf Q}_0}[\log Y_a]+t \mathbb{E}_{\mathsf R}[\log Y_a] = \infty,\] because \({\mathsf R}\) has infinite second moment. Thus \(Y_a\) is a shielding e-variable for \(\mathcal{Q}_\infty\). By Theorem 15, strong duality holds: \[\frac{\theta^2}{2} = \sup_{X\in\mathcal{E}} \inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[\log X] = \inf_{{\mathsf Q}\in\mathcal{Q}}H({\mathsf Q}\mid {\mathsf P}).\] Proposition 9 implies that this strong duality is not attained by any e-variable. Indeed, if it was obtained, such e-variable would have to equal \(d{\mathsf Q}_0/d{\mathsf P}\), \({\mathsf P}\)-a.s., and hence also \({\mathsf R}\)-a.s., but \(\mathbb{E}_\mathbb{R}[\log d{\mathsf Q}_0/d{\mathsf P}] = \mathbb{E}_{\mathsf R}[\theta X - \theta^2/2] = -\theta^2/2 < I\).
In the context of Subsection 4.4, \(\mathsf{G}_b \geq \mathsf{G}_{bc}\) holds always. We provide an example where \(\mathsf{G}_{b} > \mathsf{G}_{bc}\) in order to show that some assumptions on \(\mathcal{Q}\) are needed for \(\mathsf{G}_{b} = \mathsf{G}_{bc}\) to hold.
Example 13. Let \(\Omega=[0,1]\) with its Borel \(\sigma\)-field, let \(\mathcal{P}=\{\mathsf U\}\) and let \(\mathcal{Q}=\{\delta_q:q\in\mathbb{Q}\cap[0,1]\}\). Set \(A=\mathbb{Q}\cap[0,1]\). For \(N\ge1\), define \(X_N=N\mathbf{1}_A+\mathbf{1}_{A^c} \in \mathcal{E}_b\). Moreover, for every \(q\in\mathbb{Q}\cap[0,1]\), \(\mathbb{E}_{\delta_q}[\log X_N]=\log N\). Hence \[\sup_{X\in\mathcal{E}_b}\inf_{Q\in\mathcal{Q}}\mathbb{\mathbb{E}}_{\mathsf Q}[\log X] =\infty.\]
On the other hand, let \(X\in\mathcal{E}_{bc}\). Then \(X\) is continuous, nonnegative, and satisfies \(\mathbb{E}_{\mathsf U}[X]\le1\). Therefore \(\inf_{\omega\in[0,1]}X(\omega)\le1\). Since \(\mathbb{Q}\cap[0,1]\) is dense and \(X\) is continuous, \[\inf_{q\in\mathbb{Q}\cap[0,1]}X(q) = \inf_{x\in[0,1]}X(\omega) \le1.\] Consequently, \[\inf_{Q\in\mathcal{Q}}\mathbb{\mathbb{E}}_{\mathsf Q}[\log X] = \inf_{q\in\mathbb{Q}\cap[0,1]}\log X(q) \le0.\] Since the constant function \(X\equiv1\) belongs to \(\mathcal{E}_{bc}\), we obtain \[\sup_{X\in\mathcal{E}_{bc}}\inf_{Q\in\mathcal{Q}}\mathbb{\mathbb{E}}_{\mathsf Q}[\log X] =0.\] Thus \(\mathsf{G}_b>\mathsf{G}_{bc}\) in general.
Theorem 3, establishes strong duality for the minimax e-power of bounded e-variables, equivalently for the bounded GROW value \(\mathsf{G}_b\). It shows that \(\mathsf{G}_b\) equals the minimum relative entropy between the weak-\(*\) closed convex hulls of \(\mathcal{P}\) and \(\mathcal{Q}\), without imposing any restrictions on either family. We also extend this result to the REGROW criterion for arbitrary offsets. We then prove several strong duality results for the unbounded GROW value \(\mathsf{G}\). Theorem 6 treats the case of a singleton null \(\mathcal{P}\), improving on previous work. Theorem 8 extends this result to the case that a joint information projection exists. We then prove strong duality results for unbounded GROW under assumptions such as setwise or weak compactness, that do not require taking weak-\(*\) closures. It remains open to what extent such restrictions are necessary. While our examples show that \(\mathsf{G}\) can be strictly larger than \(\mathsf{G}_b\), we do not know whether a universal strong duality theorem for \(\mathsf{G}\) holds for arbitrary \(\mathcal{P}\) and \(\mathcal{Q}\). Indeed, Examples 4 and 10 demonstrate that a naive entropy duality does not hold without assumptions; it remains open whether some alternative universal dual representation for unbounded GROW exists.
AR acknowledges support from the National Science Foundation under grant DMS-2310718. ML acknowledges support from the National Science Foundation under grant NSF DMS-2510965. ChatGPT was used to simplify some proofs and construct some examples, but the authors take responsibility for the correctness of all claims in the paper.
Proof of Lemma 1. The implication \(\mu=\nu \Rightarrow H(\nu\mid\mu)=0\) is immediate from the definition. Conversely, suppose that \(H(\nu\mid\mu)=0\) and consider some \(A\in\mathcal{F}\), along with the partition \(\pi=\{A,A^c\}\). Then the finite-dimensional entropy \[H_\pi(\nu\mid\mu) = \sum_{A\in\pi}\nu(A)\log\frac{\nu(A)}{\mu(A)}\] is nonnegative, since it is the relative entropy between the two probability vectors \((\nu(A), \nu(A^c))\) and \((\mu(A), \mu(A^c))\). Hence, we must have \(H_\pi(\nu\mid\mu)=0\). The finite-dimensional relative entropy of two probability vectors is zero if and only if the two vectors are equal. Therefore \[\nu(A)=\mu(A) \qquad\text{and}\qquad \nu(A^c)=\mu(A^c).\] Since \(A\in\mathcal{F}\) was arbitrary, it follows that \(\nu=\mu\). ◻
Proof of Lemma 2. Write \[F(g)=\int g\,d{\mathsf Q}-\log\int e^g\,d{\mathsf P}, \qquad g\in\mathcal{B}_b.\]
We first show that \(F(g)\le H({\mathsf Q}\mid{\mathsf P})\) for every \(g\in\mathcal{B}_b\). Suppose first that \(g\) is simple, say \(g=\sum_{i=1}^n c_i 1_{A_i}\) for some finite measurable partition \(\pi=\{A_1,\dots,A_n\}\). Set \[q_i={\mathsf Q}(A_i),\qquad p_i={\mathsf P}(A_i),\qquad Z=\sum_{i=1}^n p_i e^{c_i}.\] If \(q_i>0\) and \(p_i=0\) for some \(i\), then \(H_\pi({\mathsf Q}\mid{\mathsf P})=\infty\), so there is nothing to prove. Otherwise \(Z>0\) and we define \(r_i={p_i e^{c_i}}/{Z}\). Then \((r_i)\) is a probability vector, and recalling convention 9 , \[F(g) = \sum_{i=1}^n q_i (c_i-\log Z) = \sum_{i=1}^n q_i \log\frac{r_i}{p_i} = \sum_{i=1}^n q_i \log\frac{q_i}{p_i} - \sum_{i=1}^n q_i \log\frac{q_i}{r_i}.\] Since \(\sum_{i=1}^n q_i \log({q_i}/{r_i})\ge 0\), it follows that \[F(g)\le \sum_{i=1}^n q_i \log\frac{q_i}{p_i}=H_\pi({\mathsf Q}\mid{\mathsf P})\le H({\mathsf Q}\mid{\mathsf P}).\]
Now let \(g\in\mathcal{B}_b\) be arbitrary. Choose simple functions \(g_m\in\mathcal{B}_b\) such that \(\|g_m-g\|_\infty\to 0\). Then \[\left|\int g_m\,d{\mathsf Q}-\int g\,d{\mathsf Q}\right|\le \|g_m-g\|_\infty\to 0.\] Also, if \(\varepsilon_m=\|g_m-g\|_\infty\), then \[e^{-\varepsilon_m}\int e^{g_m}\,d{\mathsf P} \le \int e^g\,d{\mathsf P} \le e^{\varepsilon_m}\int e^{g_m}\,d{\mathsf P},\] so \[\left|\log\int e^g\,d{\mathsf P}-\log\int e^{g_m}\,d{\mathsf P}\right| \le \varepsilon_m\to 0.\] Therefore \(F(g_m)\to F(g)\). Since \(F(g_m)\le H({\mathsf Q}\mid{\mathsf P})\) for every \(m\), we conclude that \[\sup_{g\in\mathcal{B}_b}F(g)\le H({\mathsf Q}\mid{\mathsf P}).\]
For the reverse inequality, fix a finite measurable partition \(\pi=\{A_1,\dots,A_n\}\), and again write \(q_i={\mathsf Q}(A_i)\) and \(p_i={\mathsf P}(A_i)\). If \(q_i>0\) and \(p_i=0\) for some \(i\), then \(H_\pi({\mathsf Q}\mid{\mathsf P})=\infty\). In that case, letting \(g_m=m\, \mathbf{1}_{A_i}\), we obtain \[F(g_m)\ge m q_i-\log\int e^{g_m}\,d{\mathsf P}= m q_i - \log {\mathsf P}(A_i^c) \to \infty.\] Hence \(\sup_{g\in\mathcal{B}_b}F(g)=\infty=H_\pi({\mathsf Q}\mid{\mathsf P})\). Now assume instead that \(q_i=0\) whenever \(p_i=0\). For \(m\in\mathbb{N}\), define \[c_i^{(m)}= \begin{cases} \log(q_i/p_i), & q_i>0,\\ -m, & q_i=0, \end{cases} \qquad g_m=\sum_{i=1}^n c_i^{(m)} \mathbf{1}_{A_i}\in\mathcal{B}_b.\] Then \[\int g_m\,d{\mathsf Q}=\sum_{q_i>0} q_i\log\frac{q_i}{p_i},\] while \[\int e^{g_m}\,d{\mathsf P} = \sum_{q_i>0} p_i\frac{q_i}{p_i} + \sum_{q_i=0} p_i e^{-m} = 1+e^{-m}\sum_{q_i=0}p_i.\] Therefore \[F(g_m) = \sum_{q_i>0} q_i\log\frac{q_i}{p_i} - \log\left(1+e^{-m}\sum_{q_i=0}p_i\right) \rightarrow \sum_{i=1}^n q_i\log\frac{q_i}{p_i}=H_\pi({\mathsf Q}\mid{\mathsf P}).\] Hence \(H_\pi({\mathsf Q}\mid{\mathsf P})\le \sup_{g\in\mathcal{B}_b}F(g)\). Since \(\pi\) was arbitrary, taking the supremum over all finite measurable partitions gives \(H({\mathsf Q}\mid{\mathsf P})\le \sup_{g\in\mathcal{B}_b}F(g)\), concluding the proof. ◻
Proof of Proposition 2. Fix \({\mathsf P}\in\operatorname{ba}_+\) and let \({\mathsf P}={\mathsf P}_c+{\mathsf P}_p\) denote its Yosida–Hewitt decomposition. Since \({\mathsf P}_c\le {\mathsf P}\), we immediately have \(H({\mathsf Q}\mid{\mathsf P})\le H({\mathsf Q}\mid{\mathsf P}_c)\).
We next prove the reverse inequality. Let \(\lambda={\mathsf Q}+{\mathsf P}_c\). Since \(\lambda\) is countably additive and \({\mathsf P}_p\) is purely finitely additive, there is no nonzero positive finitely additive measure dominated by both \({\mathsf P}_p\) and \(\lambda\). Equivalently, we may choose a sequence \((B_n)\subseteq\mathcal{F}\) such that \[{\mathsf P}_p(B_n)+\lambda(B_n^c)\longrightarrow0.\] Fix now a finite measurable partition \(\pi=\{A_1,\ldots,A_m\}\). For each \(n\), consider the finite measurable partition \(\pi_n = \{A_1\cap B_n,\ldots,A_m\cap B_n,B_n^c\}\). Then \(H({\mathsf Q}\mid{\mathsf P})\ge H_{\pi_n}({\mathsf Q}\mid{\mathsf P})\). For each \(i=1,\ldots,m\), we have \({\mathsf Q}(A_i\cap B_n)\to {\mathsf Q}(A_i)\), and \[{\mathsf P}(A_i\cap B_n) = {\mathsf P}_c(A_i\cap B_n)+{\mathsf P}_p(A_i\cap B_n) \to {\mathsf P}_c(A_i).\] Hence, using 9 , which implies the lower semicontinuity of \(q\log(q/p)\), yields \[\liminf_{n\to\infty} \sum_{i=1}^m {\mathsf Q}(A_i\cap B_n) \log\frac{{\mathsf Q}(A_i\cap B_n)}{{\mathsf P}(A_i\cap B_n)} \ge \sum_{i=1}^m {\mathsf Q}(A_i)\log\frac{{\mathsf Q}(A_i)}{{\mathsf P}_c(A_i)} = H_\pi({\mathsf Q}\mid{\mathsf P}_c).\] The remaining atom \(B_n^c\) does not affect the lower bound, since \({\mathsf Q}(B_n^c)\to0\) and \({\mathsf Q}(B_n^c)\log({{\mathsf Q}(B_n^c)}/{{\mathsf P}(B_n^c)})\) has nonnegative limit inferior. Therefore, \[H({\mathsf Q}\mid{\mathsf P}) \ge \liminf_{n\to\infty} H_{\pi_n}({\mathsf Q}\mid{\mathsf P}) \ge H_\pi({\mathsf Q}\mid{\mathsf P}_c).\] Taking the supremum over all finite measurable partitions \(\pi\) yields \(H({\mathsf Q}\mid{\mathsf P})\ge H({\mathsf Q}\mid{\mathsf P}_c)\), hence the first claim is established.
We now consider \({\mathsf Q}\in\operatorname{ba}_1 \setminus \mathcal{M}_1\). Then continuity from above fails for \({\mathsf Q}\). Hence there exists a nonincreasing sequence \((A_n)\subseteq\mathcal{F}\) such that \(A_n\downarrow\varnothing\) but there exists some \(\delta>0\) such that \({\mathsf Q}(A_n) \geq \delta\) for all \(n\). On the other hand, since \({\mathsf P}\in\mathcal{M}_+\) is countably additive and finite, we have \({\mathsf P}(A_n)\downarrow0\). Consider the two-point partitions \(\pi_n=\{A_n,A_n^c\}\). Then \[H({\mathsf Q}\mid{\mathsf P}) \ge H_{\pi_n}({\mathsf Q}\mid{\mathsf P}) \geq {\mathsf Q}(A_n)\log\frac{{\mathsf Q}(A_n)}{{\mathsf P}(A_n)} + {\mathsf Q}(A_n^c)\log\frac{{\mathsf Q}(A_n^c)}{{\mathsf P}(\Omega)}.\] The first term now tends to \(\infty\), whereas the second term is bounded from below. This yields the claim. ◻
Proof of Lemma 3. Since \(\mathcal{E}_b\subseteq\mathcal{E}\), we only have to argue \(\sup_{g\in\mathcal{E}_b}\mathbb{E}_{\mathsf Q}[\log g] \geq \mathbb{E}_{\mathsf Q}[\log f]\) for all \(f \in \mathcal{E}\) such that \(\mathbb{E}_{\mathsf Q}[(\log f)^-] < \infty\). We fix such an \(f\) and define \(f_m=f\wedge m \in \mathcal{E}_b\) for \(m\ge1\). Then by monotone convergence, \(\mathbb{E}_{\mathsf Q}[\log f_m] \uparrow \mathbb{E}_{\mathsf Q}[\log f]\). This yields the statement. ◻
We first show that the weak-\(*\) closed convex hulls of \(\mathcal{P}\) and \(\mathcal{Q}\) have a common finitely additive element. Since \(\operatorname{ba}_1\) is weak-\(*\) compact, the sequence \(({\mathsf P}_n)_{n\ge1}\) has a weak-\(*\) convergent subnet \(({\mathsf P}_{n(\alpha)})_\alpha\), say with limit \(\mu\in\operatorname{ba}_1\). Since \((n(\alpha))_\alpha\) is a subnet of the sequence of integers, we have \(n(\alpha)\to\infty\). For any \(f\in\mathcal{B}_b\), \[\int f\,d{\mathsf Q}_{n(\alpha)}-\int f\,d{\mathsf P}_{n(\alpha)} = \frac{1}{n(\alpha)} \sum_{k=2}^{n(\alpha)+1} f(k) - \frac{1}{n(\alpha)} \sum_{k=1}^{n(\alpha)} f(k) = \frac{f(n(\alpha)+1)-f(1)}{n(\alpha)}.\] Thus \[\left| \int f\,d{\mathsf Q}_{n(\alpha)}-\int f\,d{\mathsf P}_{n(\alpha)} \right| \le \frac{2\|f\|_\infty}{n(\alpha)} \longrightarrow 0.\] Since \({\mathsf P}_{n(\alpha)}\to\mu\) in \(\sigma(\operatorname{ba},\mathcal{B}_b)\), it follows that also \({\mathsf Q}_{n(\alpha)}\to\mu\) in \(\sigma(\operatorname{ba},\mathcal{B}_b)\). Hence \(\mu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\cap\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\).
We next show that \(\mu\) is not countably additive. Fix \(m\in\mathbb{N}\). Since \(n(\alpha)\to\infty\), we have \(n(\alpha)\ge m\) eventually, and hence \({\mathsf P}_{n(\alpha)}(\{m\})=\frac{1}{n(\alpha)}\) eventually. Therefore \(\mu(\{m\}) = \lim_\alpha {\mathsf P}_{n(\alpha)}(\{m\}) = 0\). If \(\mu\) were countably additive, then \[1=\mu(\mathbb{N})=\sum_{m=1}^\infty \mu(\{m\})=0,\] a contradiction. Thus \(\mu\notin\mathcal{M}_1\).
Lemma 1 yields \(H(\mu\mid\mu) = 0\). Since \(\mu\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\cap\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\) and since relative entropy is nonnegative, we conclude that \[\label{eq:example-weak-star-value-zero} \min_{\rho\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})}\min_{\lambda\in \mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})} H(\rho\mid\lambda) = 0.\tag{30}\] Let \((\rho^*,\lambda^*)\) be now any minimizing pair in 30 . By Lemma 1, this implies \(\rho^*=\lambda^*\). Hence every minimizing pair is of the form \((\lambda^*,\lambda^*)\), with \(\lambda^*\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\cap\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\). It remains to show that this intersection contains only purely finitely additive probability measures.
Suppose, to the contrary, that \(\lambda\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{P})\cap\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\cap\mathcal{M}_1\). For every \(n,k\in\mathbb{N}\), we have \({\mathsf P}_n(\{k\})\ge {\mathsf P}_n(\{k+1\})\). This inequality is preserved under finite convex combinations and under weak-\(*\) limits, since the maps \(\rho\mapsto\rho(\{k\})\) and \(\rho\mapsto\rho(\{k+1\})\) are \(\sigma(\operatorname{ba},\mathcal{B}_b)\)-continuous. Hence, for every \(k\in\mathbb{N}\), \(\lambda(\{k\})\ge \lambda(\{k+1\})\). On the other hand, every \({\mathsf Q}_n\) satisfies \({\mathsf Q}_n(\{1\})=0\), and this property is also preserved under finite convex combinations and weak-\(*\) limits. Since \(\lambda\in\mathop{\mathrm{\overline{\mathop{\mathrm{co}}}}}^*(\mathcal{Q})\), it follows that \(\lambda(\{1\})=0\). Therefore, \[0\le \lambda(\{k\})\le \lambda(\{1\})=0, \qquad k\in\mathbb{N},\] and hence \(\lambda(\{k\})=0\) for every \(k\). Let now \(\lambda=\lambda_c+\lambda_p\) be the Yosida–Hewitt decomposition of \(\lambda\) into its countably additive and purely finitely additive parts. Since \(\lambda_c\le\lambda\), we have \(\lambda_c(\{k\})=0\) for every \(k\). Countable additivity of \(\lambda_c\) then yields \[\lambda_c(\mathbb{N})=\sum_{k=1}^\infty \lambda_c(\{k\})=0.\] Hence \(\lambda_c=0\), and therefore \(\lambda\) is purely finitely additive.
Example 8 presented a simple example of \(\mathcal{P},\mathcal{Q}\) that are convex and weakly compact, and showed how to apply this paper’s techniques to derive strong duality, along with the GROW e-variable. We now do the same for a slightly more sophisticated example, which involves mean shifts of unbounded random variables with bounded second moment.
Example 14 (Mean shift with second-moment constraints and no domination). Let \(\Omega=\mathbb{R}\) and define \[\mathcal{P} = \left\{ {\mathsf P}\in \mathcal{M}_1: \mathbb{E}_{\mathsf P}[X]\le0,\;\mathbb{E}_{\mathsf P}[X^2]\le1 \right\},\] and \[\mathcal{Q} = \left\{ {\mathsf Q}\in \mathcal{M}_1: \mathbb{E}_{\mathsf Q}[X]\ge1,\;\mathbb{E}_{\mathsf Q}[X^2]\le2 \right\}.\] Both classes are convex and weakly compact. Indeed, the second-moment bounds imply tightness, while the moment constraints are closed under weak convergence because the second moments are uniformly bounded. The classes are not dominated by any common finite reference measure, since \(\mathcal{P}\) contains \(\delta_x\) for every \(x\in[-1,0]\) and \(\mathcal{Q}\) contains \(\delta_x\) for every \(x\in[1,\sqrt2]\).
Let \[\varphi=\frac{1+\sqrt5}{2}, \qquad x_-=-\frac{1}{\varphi}, \qquad x_+=\varphi.\] Define \[{\mathsf P}^* = \frac{\varphi^2}{\varphi^2+1}\delta_{x_-} + \frac{1}{\varphi^2+1}\delta_{x_+},\] and \[{\mathsf Q}^* = \frac{1}{\varphi^2+1}\delta_{x_-} + \frac{\varphi^2}{\varphi^2+1}\delta_{x_+}.\] Then \(\mathbb{E}_{{\mathsf P}^*}[X]=0, \mathbb{E}_{{\mathsf P}^*}[X^2]=1,\) so \({\mathsf P}^*\in\mathcal{P}\), while \(\mathbb{E}_{{\mathsf Q}^*}[X]=1, \mathbb{E}_{{\mathsf Q}^*}[X^2]=2,\) so \({\mathsf Q}^*\in\mathcal{Q}\).
On the support \(\{x_-,x_+\}\), the likelihood ratio is \[\frac{d{\mathsf Q}^*}{d{\mathsf P}^*}(x_-)=\frac{1}{\varphi^2}, \qquad \frac{d{\mathsf Q}^*}{d{\mathsf P}^*}(x_+)=\varphi^2,\] yielding \(I= H({\mathsf Q}^*\mid {\mathsf P}^*) \frac{2\log\varphi}{\sqrt5}\). We next choose a useful pointwise version of the likelihood ratio. [21] indicates that all admissible e-variables have a quadratic form, so below we choose the quadratic to match the value of \(\frac{d{\mathsf Q}^*}{d{\mathsf P}^*}\) at \(x_-\) and \(x_+\). Let \[A = \frac{2}{5}\left(1+\frac{4\log\varphi}{\sqrt5}\right),\] and define \[X^*(x)=A(1+x)+(1-A)x^2, \qquad x\in\mathbb{R}.\] Then \(2/3<A<3/4\), and \(X^*(x)\ge0\) for all \(x\in\mathbb{R}\). Moreover, since \(1+\varphi=\varphi^2\) and \(1-1/\varphi=1/\varphi^2\), we have \[X^*(x_-)=\frac{1}{\varphi^2}, \qquad X^*(x_+)=\varphi^2.\] Thus \(X^*\) is a \({\mathsf P}^*\)-version of \(d{\mathsf Q}^*/d{\mathsf P}^*\), and one can check that \(X^* \in \mathcal{E}\).
We now show that condition (0) of Assumption 2 holds. First note that \[\inf_{{\mathsf R}\in\mathcal{Q}}H({\mathsf R}\mid {\mathsf P}^*)=H({\mathsf Q}^*\mid {\mathsf P}^*).\] Indeed, since \({\mathsf P}^*\) is supported on \(\{x_-,x_+\}\), any \({\mathsf R}\) with \(H({\mathsf R}\mid {\mathsf P}^*)<\infty\) must also be supported on \(\{x_-,x_+\}\). Writing \[{\mathsf R}=(1-q)\delta_{x_-}+q\delta_{x_+},\] the constraints \(\mathbb{E}_{\mathsf R}[X]\ge1\) and \(\mathbb{E}_{\mathsf R}[X^2]\le2\) force \(q=\frac{\varphi^2}{\varphi^2+1},\) and hence \({\mathsf R}={\mathsf Q}^*\).
It remains to check that \(\mathbb{E}_{\mathsf Q}[\log X^*]\ge H({\mathsf Q}^*\mid {\mathsf P}^*) \text{ for every }{\mathsf Q}\in\mathcal{Q}\). For the above choice of \(A\), a direct one-dimensional calculus check shows that there exist constants \(c\in\mathbb{R}\) and \[\lambda=4A-2>0, \qquad \mu=\frac{3}{2} A-1>0,\] such that \[h(x)=c+\lambda x-\mu x^2 \le \log X^*(x) \qquad \text{for all }x\in\mathbb{R},\] with equality at \(x_-\) and \(x_+\). Therefore, for any \({\mathsf Q}\in\mathcal{Q}\), \[\begin{align} \mathbb{E}_{\mathsf Q}[\log X^*] \ge \mathbb{E}_{\mathsf Q}[h(X)] = c+\lambda \mathbb{E}_{\mathsf Q}[X]-\mu \mathbb{E}_{\mathsf Q}[X^2] \ge c+\lambda-2\mu. \end{align}\] Since \({\mathsf Q}^*\) has \(\mathbb{E}_{{\mathsf Q}^*}[X]=1\) and \(\mathbb{E}_{{\mathsf Q}^*}[X^2]=2\), and since \(h(X)=\log(X^*)\) on the support of \({\mathsf Q}^*\), \[c+\lambda-2\mu = \mathbb{E}_{{\mathsf Q}^*}[h(X)] = \mathbb{E}_{{\mathsf Q}^*}[\log X^*] = H({\mathsf Q}^*\mid {\mathsf P}^*).\] Thus \[\mathbb{E}_{\mathsf Q}[\log X^*]\ge \inf_{{\mathsf R}\in\mathcal{Q}}H({\mathsf R}\mid P^*) \qquad \text{for every }{\mathsf Q}\in\mathcal{Q}.\] This is condition (0). Consequently, the likelihood-ratio version \(X^*\) witnesses strong duality: \[\mathsf{G} = \sup_{X\in\mathcal{E}} \inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[\log X] = H({\mathsf Q}^*\mid {\mathsf P}^*) = \frac{2\log\varphi}{\sqrt5}.\] Clearly \(X^*\) is also continuous, though not bounded, but Theorem 13 yields that \(\mathsf{G}= \mathsf{G}_b = \mathsf{G}_{bc}\) as well.
Derivation of \(\inf_{{\mathsf Q}\in\mathcal{Q}_a}H({\mathsf Q}\mid {\mathsf P})\) in Example 11. Let \(A_m=[m,\infty), \alpha_m={\mathsf Q}^*(A_m),\) and let \({\mathsf R}_m\) be the conditional distribution of \({\mathsf Q}^*\) on \(A_m\), i.e., \({\mathsf R}_m(\cdot)={\mathsf Q}^*(\cdot\mid A_m).\) A direct calculation gives \(\alpha_m = \frac{1}{3(m+1)^2}+\frac{2}{3(m+1)^3}.\) Moreover, \(M_m=\int s d {\mathsf R}_m(s) \to\infty.\) Indeed, \[\frac{1}{\alpha_m} \int_m^\infty s\,q^*(s)\,ds = \frac{2}{3\alpha_m} \left( \frac{1}{m+1} + \frac{1}{(m+1)^2} - \frac{1}{(m+1)^3} \right) \geq 2m.\] For all large \(m\), define \(\eta_m = \frac{a-\frac{2}{3}}{M_m-\frac{2}{3}} \in(0,1)\) and set \({\mathsf Q}_m=(1-\eta_m){\mathsf Q}^*+\eta_m {\mathsf R}_m.\) Then \[\int s d {\mathsf Q}_m(s) = (1-\eta_m)\frac{2}{3}+\eta_m M_m = a,\] so \({\mathsf Q}_m\in\mathcal{Q}_a\). By convexity of relative entropy we get, \[H({\mathsf Q}_m\mid {\mathsf Q}^*) \le (1-\eta_m)H({\mathsf Q}^*\mid {\mathsf Q}^*)+\eta_m H({\mathsf R}_m\mid {\mathsf Q}^*) = \eta_m\log\frac{1}{\alpha_m} \rightarrow 0.\] Therefore, \[H({\mathsf Q}_m\mid {\mathsf P}) = H({\mathsf Q}_m\mid {\mathsf Q}^*)+ \mathbb{E}_{{\mathsf Q}_m}\left[\log\frac{d{\mathsf Q}^*}{d{\mathsf P}}\right] = H({\mathsf Q}_m\mid {\mathsf Q}^*)+ a+\log\frac{2}{3} \longrightarrow a+\log\frac{2}{3}.\] Hence \(\inf_{{\mathsf Q}\in\mathcal{Q}_a}H({\mathsf Q}\mid {\mathsf P}) = a+\log\frac{2}{3}.\) ◻
Proof of Theorem 15. For every \({\mathsf Q}\in\mathcal{Q}_\infty\), \(\sup_{X \in \mathcal{E}} \mathbb{E}_{\mathsf Q}[\log X] \geq \mathbb{E}_{\mathsf Q}[\log Y] = \infty\), and thus from 4 , \(\inf_{{\mathsf P}\in\mathcal{P}_\mathrm{eff}}H({\mathsf Q}\mid {\mathsf P})=\infty.\) Therefore, \[\inf_{{\mathsf Q}\in\mathcal{Q}} \inf_{{\mathsf P}\in\mathcal{P}_\mathrm{eff}}H({\mathsf Q}\mid {\mathsf P}) = \min\left( \inf_{{\mathsf Q}\in\mathcal{Q}_I} \inf_{P\in\mathcal{P}_\mathrm{eff}}H({\mathsf Q}\mid {\mathsf P}), \, \inf_{{\mathsf Q}\in\mathcal{Q}_\infty} \inf_{{\mathsf P}\in\mathcal{P}_\mathrm{eff}}H({\mathsf Q}\mid {\mathsf P}) \right) = I.\] Thus, we only have to prove the first equality in 29 , and indeed only \(\sup_{X\in\mathcal{E}} \inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[\log X] \ge I,\) needs to be argued. To this end \(\varepsilon>0\). By the definition of \(I\), there exists \(X_\varepsilon\in\mathcal{E}\) such that \[I - \varepsilon \leq \inf_{{\mathsf Q}\in\mathcal{Q}_I}\mathbb{E}_{\mathsf Q}[\log X_\varepsilon].\] For \(\delta\in(0,1)\) define now \[Z_{\varepsilon,\delta} = (1-\delta)X_\varepsilon+\delta Y \in \mathcal{E}.\] If \({\mathsf Q}\in\mathcal{Q}_I\), then \(Z_{\varepsilon,\delta}\ge (1-\delta)X_\varepsilon,\) and therefore \[\mathbb{E}_{\mathsf Q}[\log Z_{\varepsilon,\delta}] \ge \mathbb{E}_{\mathsf Q}[\log X_\varepsilon]+\log(1-\delta) \ge I-\varepsilon+\log(1-\delta).\] If \({\mathsf Q}\in\mathcal{Q}_\infty\), then \(Z_{\varepsilon,\delta}\ge \delta Y,\) and hence, by 28 , \[\mathbb{E}_{\mathsf Q}[\log Z_{\varepsilon,\delta}] \ge \log\delta+\mathbb{E}_{\mathsf Q}[\log Y] = \infty.\] The last two displays together yield \[\inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[\log Z_{\varepsilon,\delta}] \ge I-\varepsilon+\log(1-\delta).\] Letting \(\delta\downarrow0\) and \(\varepsilon\downarrow0\) gives \(\sup_{X\in\mathcal{E}} \inf_{{\mathsf Q}\in\mathcal{Q}}\mathbb{E}_{\mathsf Q}[\log X] \ge I,\) concluding the proof. ◻