Consistency of variational approximations
under bounded Kullback–Leibler divergence

Hien Duy Nguyen
Department of Mathematics and Physical Science,
La Trobe University, Melbourne, Australia
Institute of Mathematics for Industry, Kyushu University, Fukuoka, Japan
hien@imi.kyushu-u.ac.jp

,

Jacob Westerhout
School of Mathematics and Physics, University of Queensland, Brisbane, Australia
j.westerhout@uq.edu.au

,

Thomas Guilmeau and Julyan Arbel
Univ. Grenoble Alpes, Inria, CNRS, Grenoble INP, LJK, 38000 Grenoble, France
thomas.guilmeau@inria.fr, julyan.arbel@inria.fr


Abstract

Variational methods are widely used to approximate posterior distributions in Bayesian inference when exact computation is infeasible. We study when such approximations inherit posterior consistency. Our first result shows that, on a general metric space, a uniform bound on the Kullback–Leibler divergence from the approximating measures to a tight sequence of target measures forces the approximating sequence to be tight. It follows that if the target posteriors converge weakly to a Dirac mass at the true parameter, then any variational sequence with bounded Kullback–Leibler divergence to the targets is also consistent. We also give simple logarithmic-moment conditions that verify this boundedness condition, and illustrate them for smooth generalised posterior distributions.

Generalised Bayesian inference, Kullback–Leibler divergence, Posterior consistency, Variational inference, Weak convergence.

1 Introduction↩︎

Variational methods are widely used to approximate posterior probability measures arising in Bayesian and generalised Bayesian inference, particularly when exact computation is infeasible [1], [2]. A posterior sequence is consistent when the posterior distributions \(\mu_n\) converge to a Dirac mass at the true parameter \(\theta_0\) as the number of observations \(n\) grows to infinity. A fundamental theoretical question is as follows: under what conditions do variational approximations \(\nu_n\) inherit this convergence?

This question is the focus of a growing literature. However, most existing works impose restrictions on one or more of the following aspects: (i) the approximating family, (ii) the statistical model or posterior, and (iii) the variational optimisation procedure. [3] study mean-field variational Bayes for mixture estimation and model selection, while [4] study mean-field spike-and-slab approximations for high-dimensional sparse linear regression. [5] treats sparse deep learning and model selection over neural-network architectures. [6] and [7] work with fractional or \(\alpha\)-variational posteriors; the latter extends the framework to latent-variable models. [8] obtain asymptotic variational Bernstein–von Mises results for parametric latent-variable models, and [9] give convergence rates under prior-mass and testing conditions, including conditions tailored to mean-field approximations. These contributions provide statistical guarantees in important settings, but their assumptions typically encode some combination of model structure, variational-family structure, exact variational optimisation, or fractional posterior form. As a result, the basic mechanism by which posterior consistency is preserved under variational approximation remains unclear.

A typical variational inference construction produces a sequence of approximations by (approximately) minimising the Kullback–Leibler (KL) divergence with respect to the posterior \(\mu_n\) over a variational class. In this paper, we show that uniform control of this KL divergence induces tightness of variational sequences, leading to a general consistency principle in metric spaces. The resulting argument, building on a technical device introduced by [4], is concise and relies only on fundamental properties of the KL divergence and weak convergence of probability measures. In particular, we impose no structural assumption on the approximating family, the possibly tempered posteriors only need to converge to a Dirac mass, and variational approximations that only approximately minimise the KL divergence are allowed.

To illustrate the usefulness of this principle, we derive simple sufficient conditions ensuring boundedness of the KL divergence based on moment bounds, and apply these results to smooth generalised posterior distributions. Thus, consistency of variational approximations can be reduced to verifying an interpretable moment condition, rather than a detailed analysis of the variational optimisation problem itself. These results hold for a large class of variational families, including in particular location-scale families, such as Gaussian variational approximations.

In Section 2, we first establish a general tightness property for KL level sets in metric spaces. We show in Theorem [thm:metric:asympt-tight] that uniform KL control relative to a tight sequence of target measures induces tightness of the corresponding approximating sequence. Our main contribution, Theorem [thm:metric:delta], establishes consistency for variational sequences and follows directly from Theorem [thm:metric:asympt-tight], under a sufficient boundedness condition on the KL divergence between \(\nu_n\) and \(\mu_n\). We then use three examples to show that natural weakenings of this boundedness condition do not lead to consistency and that this condition is not necessary. Section 3 addresses the complementary question of verifying the KL boundedness condition in practice. We show in Proposition [prop:log-mom-oneQ] that boundedness of the optimal variational KL divergence can be reduced to a simple moment condition under a single reference distribution. We then provide a sufficient condition in Proposition [prop:LAN-like] based on a quadratic lower envelope, while Proposition [prop:quadratic-envelope] shows how to handle this envelope condition. We illustrate the approach on generalised posterior distributions in Theorem [thm:gp-envelope]. Finally, Section 4 applies these results in the context of generalised linear models.
Notations. Let \((\mathcal{X},d)\) be a metric space and let \(\mathcal{P}(\mathcal{X})\) denote the set of Borel probability measures on \(\mathcal{X}\). We write weak convergence as \(\nu_n\rightsquigarrow\nu\). For \(\mu,\nu\in\mathcal{P}(\mathcal{X})\), the KL divergence is \(\mathrm{KL}(\nu\Vert\mu)=\int_{\mathcal{X}}\log({\mathrm{d}\nu}/{\mathrm{d}\mu})\,\mathrm{d}\nu\), with \(\mathrm{KL}(\nu\Vert\mu)=\infty\) if \(\nu\not\ll\mu\). We say that \(\mu\) and \(\nu\) are equivalent if \(\mu\ll\nu\) and \(\nu\ll\mu\). For \(\delta>0\) and \(A\subseteq \mathcal{X}\), we denote the \(\delta\)-expansion of \(A\) by \(A^\delta = \{x\in\mathcal{X}: d(x,A) < \delta\}\), and denote the \(\delta\)-shrinkage of \(A\) by \(A_\delta = \{x\in A: B(x,\delta)\subseteq A\}\), where \(B(x,\delta)=\{y\in\mathcal{X}:d(y,x)<\delta\}\) is the open ball of radius \(\delta\) centred at \(x\). We write \(\overline{B}(x,\delta)\) for the corresponding closed ball. Note that \((A^c)_\delta = (A^\delta)^c\). The indicator function of a set \(A\) is denoted by \(\mathbb{1}_A\), and \(x_+=\max(x,0)\) for \(x\in\mathbb{R}\).

2 Consistency of variational approximations in metric spaces through tightness↩︎

This section collects the main results on consistency of variational sequences in metric spaces. The idea of the following technical device is due to [4].

For probability measures \(\mu,\nu\) on a common measurable space \((\mathcal{X},\mathfrak X)\), for every \(A\in\mathfrak X\) and every \(\delta>0\), \[\nu(A) \le \frac{1}{\delta}\left\{\mathrm{KL}(\nu\Vert\mu) + \mu(A) e^{\delta}\right\}.\]

Proof. If \(\nu\not\ll \mu\) then \(\mathrm{KL}(\nu\Vert\mu)=\infty\) and the bound is trivial. Otherwise, the Donsker–Varadhan variational bound [10] implies that for any measurable \(f\), \[\int f\,\mathrm{d}\nu \le \mathrm{KL}(\nu\Vert\mu) + \ln\left(\int e^{f}\,\mathrm{d}\mu\right).\] Taking \(f=\delta\,\mathbb{1}_A\) and using \(\ln(1+x)\le x\) yields the result. ◻

2.1 A KL level-set tightness statement↩︎

Definition 1. Let \((\mathcal{X},d)\) be a metric space and for every \(n\in \mathbb{N}\) let \(S_n\) be a set of Borel probability measures. We say that \((S_n)\) is asymptotically tight if for every \(\eta>0\) there exists a compact \(K\subseteq\mathcal{X}\) such that for every sequence \((\nu_n)\) with \(\nu_n\in S_n\) and every \(\delta>0\), \[\liminf_{n\to\infty}\nu_n(K^\delta)\ge 1-\eta.\] When every \(S_n\) is a singleton, this reduces to the standard definition of an asymptotically tight sequence of probability measures [11].

Let \((\mathcal{X},d)\) be a metric space and let \((\mu_n)\) be a sequence of asymptotically tight Borel probability measures on \(\mathcal{X}\). For any sequence \((\epsilon_n)\subseteq [0,\infty]\) with \(\limsup_n \epsilon_n<\infty\), let the sets \[S_n=\{\nu\in\mathcal{P}(\mathcal{X}):\;\mathrm{KL}(\nu\Vert\mu_n)\le \epsilon_n\}.\] Then the sequence \((S_n)\) is asymptotically tight.

Proof. Let \(M=\limsup_n\epsilon_n<\infty\) and fix \(\eta>0\). Choose \(\zeta>0\) such that \((M+1)/\ln(1/\zeta)\le\eta\). By asymptotic tightness of \((\mu_n)\) there exists a compact \(K\subseteq\mathcal{X}\) such that, for every \(a>0\), \[\limsup_{n\to\infty} \mu_n\left((K^c)_a\right)\le\zeta.\] Let \((\nu_n)\) be any sequence with \(\nu_n\in S_n\). For fixed \(a>0\), apply Lemma [lem:j:1] with \(A=(K^c)_a\) and with the scalar parameter \(u=\ln(1/\zeta)\). Since \(\mathrm{KL}(\nu_n\Vert\mu_n)\le\epsilon_n\), \[\limsup_{n\to\infty}\nu_n((K^c)_a)\le \frac{M+\zeta e^u}{u}=\frac{M+1}{\ln(1/\zeta)}\le\eta.\] Using \((K^c)_a=(K^a)^c\), this is equivalent to \(\liminf_n\nu_n(K^a)\ge1-\eta\) for every \(a>0\). ◻

2.2 A consistency consequence↩︎

Let \((\mathcal{X},d)\) be a metric space and let \(\mu_n\rightsquigarrow\delta_{\theta_0}\). For any \(\nu_n\in\mathcal{P}(\mathcal{X})\) such that \(\limsup_n \mathrm{KL}(\nu_n\Vert\mu_n)< \infty\), we have \(\nu_n\rightsquigarrow\delta_{\theta_0}\).

Proof. Since \(\delta_{\theta_0}\) is tight, \((\mu_n)\) is asymptotically tight [11]. Applying Theorem [thm:metric:asympt-tight] with \(\epsilon_n=\mathrm{KL}(\nu_n\Vert\mu_n)\) shows that \((\nu_n)\) is asymptotically tight. Hence, by Prohorov’s theorem in metric spaces [11], every subsequence of \((\nu_n)\) has a further weakly convergent subsequence. Let \(\nu_{n_k}\rightsquigarrow\nu\) be such a subsequential limit. Lower semicontinuity of KL under weak convergence [10] gives \[\mathrm{KL}(\nu\Vert\delta_{\theta_0}) \le \liminf_k \mathrm{KL}(\nu_{n_k}\Vert\mu_{n_k}) < \infty.\] Thus \(\nu\ll\delta_{\theta_0}\), and hence \(\nu=\delta_{\theta_0}\). Every subsequence of \((\nu_n)\) therefore has a further subsequence converging weakly to \(\delta_{\theta_0}\), which implies \(\nu_n\rightsquigarrow\delta_{\theta_0}\). ◻

Note that many natural weakenings of the condition \(\limsup_{n}\mathrm{KL}(\nu_{n}\Vert\mu_{n})<\infty\) do not imply \(\nu_{n}\rightsquigarrow\delta_{\theta_{0}}\). Example 1 shows that finiteness of \(\mathrm{KL}(\nu_{n}\Vert\mu_{n})\) for each \(n\) is not even sufficient for \(\nu_n\) to converge. The same example also shows that these failures can occur with \(\mathrm{KL}(\nu_{n}\Vert\mu_{n})\) diverging arbitrarily slowly. Note also that boundedness of the KL divergence is not necessary. Example 2 satisfies \(\mu_{n},\nu_{n}\rightsquigarrow\delta_{0}\) with \(\mu_n\) equivalent to \(\nu_n\), but \(\mathrm{KL}(\nu_{n}\Vert\mu_{n})=\infty\) for every \(n\). Finally, Example 3 shows that even if \(\mu_{n},\nu_{n}\rightsquigarrow\delta_{0}\) and \(\mathrm{KL}(\nu_{n}\Vert\mu_{n})<\infty\) for every \(n\), one may still have \(\limsup_{n\to\infty}\mathrm{KL}(\nu_{n}\Vert\mu_{n})=\infty\).

Example 1. Take any positive sequence \(\epsilon_{n}\to\infty\), any sequence \((a_n)\subseteq \mathbb{R}\) and let \(\mu_{n}\) and \(\nu_{n}\) be the measures of laws \({\rm N}(0,(2\epsilon_{n})^{-1})\) and \({\rm N}(a_n,(2\epsilon_{n})^{-1})\), respectively. Then \(\mu_{n}\) and \(\nu_{n}\) are equivalent, \(\mu_{n}\rightsquigarrow\delta_{0}\), and \(\mathrm{KL}(\nu_{n}\Vert\mu_{n})=\epsilon_{n}a_n^2<\infty\) for every \(n\). However, \(\nu_{n}\) converges if and only if \(a_n\) converges to, say, \(a\), in which case \(\nu_n \rightsquigarrow \delta_a\).

Example 2. Let \(\phi_{\sigma}\) and \(\psi\) denote, respectively, the densities of \({\rm N}(0,\sigma^{2})\) and the standard Cauchy distribution. Define \[q_{n}(x)=\left(1-\frac{1}{n}\right)\phi_{1/n}(x)+\frac{1}{n}\psi(x),\; p_{n}(x)=Z_{n}^{-1}e^{-|x|}q_{n}(x),\; Z_{n}=\int_{\mathbb{R}}e^{-|x|}q_{n}(x)\,dx\in(0,1),\] and let \(\nu_{n}\) and \(\mu_{n}\) have densities \(q_{n}\) and \(p_{n}\), respectively. Then \(\mu_{n}\) and \(\nu_{n}\) are equivalent, \(\nu_{n}\rightsquigarrow\delta_{0}\), and, since \(p_{n}\propto e^{-|x|}q_{n}\) with \(e^{-|x|}\) bounded and continuous, also \(\mu_{n}\rightsquigarrow\delta_{0}\). Moreover, \[\mathrm{KL}(\nu_{n}\Vert\mu_{n})=\int\ln\!\left({q_{n}}/{p_{n}}\right)\,d\nu_{n}=\ln Z_{n}+\int|x|\,d\nu_{n}(x)=\infty,\] because \(\nu_{n}\) has a Cauchy component of weight \(1/n\).

Example 3. Let \(\mu_{n}\) and \(\nu_{n}\) be the laws of \({\rm N}(0,n^{-2})\) and \({\rm N}(0,n^{-1})\), respectively. Then \(\mu_{n},\nu_{n}\rightsquigarrow\delta_{0}\), whereas \[\mathrm{KL}(\nu_{n}\Vert\mu_{n})=\frac{1}{2}\left(n-1-\ln n\right)\to\infty.\]

2.3 Relation to the consistency of variational sequences↩︎

The following theorem, which is a direct corollary of Theorem [thm:metric:delta], states the sufficient Assumption (A1) which links the approximating families \(\mathcal{F}_n\) and the target distributions \(\mu_n\).

Let \((\mathcal{X},d)\) be a metric space, \(\mu_n\) a sequence of Borel probability measures on \(\mathcal{X}\) converging weakly to \(\delta_{\theta_0}\) with \(\theta_0 \in \mathcal{X}\). Let \(\mathcal{F}_n\) denote a sequence of sets of Borel probability measures on \(\mathcal{X}\) and assume

  • \(m_{n}=\inf_{\nu\in\mathcal{F}_{n}}\mathrm{KL}(\nu\Vert\mu_{n})\) satisfies \(\limsup_n m_n < \infty\).

Then for any sequence \(\nu_n \in \mathcal{F}_n\) with \(\mathrm{KL}(\nu_n\Vert \mu_n) \leq m_n + \epsilon_n\) where \(\limsup_n \epsilon_n<\infty\), \(\nu_n\) converges weakly to \(\delta_{\theta_0}\): \(\nu_n\rightsquigarrow \delta_{\theta_0}\).

Proof. This is immediate by Theorem [thm:metric:delta]. ◻

Conditions under which Assumption (A1) is satisfied are discussed in Section 3. The assumption that \(\limsup_n \epsilon_n < \infty\) depends on how \(\nu_n\) is generated. This assumption does not require \(\nu_n\) to exactly minimise the KL divergence and can be verified by leveraging convergence guarantees for variational inference algorithms. See for instance [12], [13] for gradient descent algorithms over location-scale families, [14] for Wasserstein gradient descent algorithms over Gaussian densities, or [15] for natural gradient algorithms over exponential families.

3 A logarithmic-moment approach for verifying Assumption (A1)↩︎

This section presents sufficient conditions for Assumption (A1) to hold, that is, for \(m_{n}=\inf_{\nu\in\mathcal{F}_{n}}\mathrm{KL}(\nu\Vert\mu_{n})\) to remain bounded over a family of approximating probability measures \(\mathcal{F}_{n}\). We do this by constructing a sequence \(\nu_{n}^{*}\in\mathcal{F}_{n}\) such that \(\limsup_{n\to\infty}\mathrm{KL}(\nu_{n}^{*}\Vert\mu_{n})<\infty\). Proofs of the results in this section and Section 4 are collected in Appendix 5.

In the Euclidean setting, let \(\lambda\) denote the Lebesgue measure on \(\mathbb{R}^{p}\) and let \(\nu_{n}\ll\mu_{n}\), with \(\mu_n\) equivalent to \(\lambda\). Then \[\mathrm{KL}(\nu_{n}\Vert\mu_{n})=\int\log\left(\frac{{\rm d}\nu_{n}}{{\rm d}\lambda}\right){\rm d}\nu_{n}-\int\log\left(\frac{{\rm d}\mu_{n}}{{\rm d}\lambda}\right){\rm d}\nu_{n}. \label{eq:KL95expansion}\tag{1}\] To bound \(\mathrm{KL}(\nu_{n}\Vert\mu_{n})\), we must then control log-moments of the densities of \(\nu_{n}\) and \(\mu_{n}\) under \(\nu_{n}\). Direct application of 1 requires control of expectations under varying measures, which is often technically delicate. As a simplifying device we use measurable bijections to convert this to the problem of bounding the log-moments of the densities of \(\nu_{n}\) and \(\mu_{n}\) under a single measure \(\tilde{\nu}\).

3.1 Setup↩︎

Let each \(\mu_{n}\) be a Borel probability measure on \(\mathbb{R}^{p}\), and fix Borel measurable bijections \(\gamma_{n}:\mathbb{R}^{p}\to\mathbb{R}^{p}\). Define the push-forward measures \(\tilde{\mu}_n=\mu_{n}\circ\gamma_{n}^{-1}\). We write \(\tilde{p}_{n}\) for the Lebesgue density of \(\tilde{\mu}_n\).

Let \(\tilde{\nu}\) be a Borel probability measure on \(\mathbb{R}^{p}\) with Lebesgue density \(\tilde{q}\). Assume that, for some \(n_{0}\) and for all \(n\ge n_{0}\), the push-forward \(\nu_{n}^{*}=\tilde{\nu}\circ\gamma_{n}\) belongs to \(\mathcal{F}_{n}\) and satisfies \(\nu_{n}^{*}\ll\mu_{n}\). Then, for all \(n\ge n_{0}\), \[\mathrm{KL}(\nu_{n}^{*}\Vert\mu_{n})=\mathrm{KL}(\tilde{\nu}\Vert\tilde{\mu}_{n})=\mathrm{E}_{\tilde{\nu}}\left[\log\left(\frac{\tilde{q}}{\tilde{p}_{n}}\right)\right].\] In particular, if \(\mathrm{E}_{\tilde{\nu}}|\log \tilde{q}|<\infty\) and \(\limsup_{n\to\infty}\mathrm{E}_{\tilde{\nu}}(-\log \tilde{p}_{n})_+<\infty\), then (A1) holds.

The feasibility condition \(\nu_{n}^{*}\in\mathcal{F}_{n}\) for all large \(n\) is usually easy to verify for common variational classes. A common choice of bijection is the affine recentering and rescaling \(\gamma_{n}(\theta)=t_{n}(\theta-\theta_{n})\), with centring points \(\theta_{n}\in\mathbb{R}^{p}\) and positive scales \(t_{n}\to\infty\). Then, for any probability measure \(\tilde{\nu}\) on \(\mathbb{R}^{p}\) and random variable \(Z\sim \tilde{\nu}\), \(\tilde{\nu}\circ\gamma_{n}\) is the measure of the law \(\mathcal{L}(\theta_{n}+t_{n}^{-1}Z)\). Suppose there is a class \(\mathcal{V}\) of measures on \(\mathbb{R}^{p}\) (e.g.a location-scale or elliptical family) such that, for each \(n\), \[\mathcal{F}_{n}\supseteq\left\{ \mathcal{L}\!\left(\theta_{n}+t_{n}^{-1}Z\right):Z\sim \tilde{\nu},\;\tilde{\nu}\in\mathcal{V}\right\} .\] Then, for any fixed \(\tilde{\nu}\in\mathcal{V}\), the reference measure \(\nu_{n}^{*}=\tilde{\nu}\circ\gamma_{n}\) belongs to \(\mathcal{F}_{n}\) for all \(n\). Moreover, if \(\mu_{n}\) has Lebesgue density \(p_{n}\), then the corresponding push-forward \(\tilde{\mu}_n=\mu_{n}\circ\gamma_{n}^{-1}\) has density \[\tilde{p}_{n}(z)=t_{n}^{-p}p_{n}\!\left(\theta_{n}+t_{n}^{-1}z\right).\]

3.2 A simple sufficient envelope condition for positive-part log-moments↩︎

A convenient way to verify the positive-part log-moment bound \(\limsup_{n}\mathrm{E}_{\tilde{\nu}}(-\log \tilde{p}_{n})_+<\infty\) is to exhibit a pointwise lower envelope for \(\tilde{p}_{n}\). Since \(x\mapsto-\log x\) is decreasing, any bound \(\tilde{p}_{n} \geq \exp (-h)\) implies \[(-\log \tilde{p}_{n})_+\le h_+.\] In particular, it is enough that \(\tilde{p}_{n}(z)\ge\exp\left(-h(z)\right)\) for all \(n\ge n_{0}\) and some measurable \(h\) with \(\mathrm{E}_{\tilde{\nu}}h_+<\infty\). This shows that the positive-part log-moment condition only prevents \(\tilde{p}_{n}\) from becoming too small on sets carrying non-negligible \(\tilde{\nu}\)-mass. The following concrete application of this principle is used in the sequel.

If there exist \(c,C>0\) and \(n_{0}\in\mathbb{N}\) such that \(\tilde{p}_{n}(z)\ge c\exp\left(-\frac{C}{2}\|z\|^{2}\right)\) for all \(z\in\mathbb{R}^{p}\) and all \(n\ge n_0\), and \(\tilde{\nu}\) has finite second moment, then \(\limsup_{n\to\infty}\mathrm{E}_{\tilde{\nu}}(-\log \tilde{p}_{n})_+<\infty\).

The next result isolates a mechanism that yields the required lower envelope.

Let \(\tilde{\mu}_{n}\) be Borel probability measures on \(\mathbb{R}^{p}\) with strictly positive Lebesgue densities \(\tilde{p}_{n}\). Write \(g_{n}=\log \tilde{p}_{n}\). Assume that there exist constants \(L,M,\alpha,r>0\) and \(n_{0}\in\mathbb{N}\) such that, for every \(n\ge n_{0}\), \(\mathrm{(B1)}\) \(\nabla g_{n}\) is \(L\)-Lipschitz on \(\mathbb{R}^{p}\), \(\mathrm{(B2)}\) \(\|\nabla g_{n}(0)\|\le M\), \(\mathrm{(B3)}\) \(\tilde{\mu}_{n}(B(0,r))\ge\alpha\). Then there exist constants \(c,C>0\), depending only on \(L,M,\alpha,r\) and \(p\), such that \[\tilde{p}_{n}(z)\ge c\exp\left(-\frac{C}{2}\|z\|^{2}\right)\qquad\text{for all }z\in\mathbb{R}^{p}\text{ and all }n\ge n_{0}.\]

A sufficient condition for \(\mathrm{(B3)}\) is the existence of a weak limit \(\mu_{\infty}\) of \((\tilde{\mu}_n)\) with \(\mu_{\infty}(B(0,r))>0\) for some \(r>0\). Indeed, the Portmanteau theorem then yields \[\liminf_{n\to\infty}\tilde{\mu}_{n}(B(0,r))\ge\mu_{\infty}(B(0,r))>0,\] so any \(0<\alpha<\mu_{\infty}(B(0,r))\) satisfies assumption \(\mathrm{(B3)}\) for all sufficiently large \(n\).

3.3 Satisfying the envelope condition in the case of generalised posteriors↩︎

Let \((\Omega,\mathfrak{A},\operatorname{pr})\) be the data-generating probability space, with expectation \(\mathrm{E}\), and let \((X_{i})_{i\ge1}\) be i.i.d.\(\mathbb{X}\)-valued random elements with common law \(\mathrm{P}_{0}\). This probability notation is used for the random data, while \(\mu\) and \(\nu\) continue to denote Borel probability measures on parameter spaces. Given a loss \(\ell:\mathbb{R}^{p}\times\mathbb{X}\to\mathbb{R}\), a prior density \(\pi\) on \(\mathbb{R}^{p}\), and temperatures \(\eta_{n}>0\), consider the generalised posterior [16] \[\begin{align} \mu_{n}(\mathrm{d}\theta)\propto\exp\left\{ -\eta_{n}\sum_{i=1}^{n}\ell(\theta,X_{i})\right\} \pi(\theta)\,\mathrm{d}\theta,\label{eq:Gibbs-posterior} \end{align}\tag{2}\] with random Lebesgue density \(p_{n}\). Let \(\hat{\theta}_{n}:\Omega\to\mathbb{R}^{p}\) be a measurable centring sequence. For each \(\omega\in\Omega\), define \[\gamma_{n}^{\omega}(\theta)=\sqrt{n}\,(\theta-\hat{\theta}_{n}(\omega)),\qquad\tilde{\mu}_n(\omega)=\mu_{n}(\omega)\circ(\gamma_{n}^{\omega})^{-1},\] and let \(\tilde{p}_{n}(\omega,\cdot)\) denote the Lebesgue density of \(\tilde{\mu}_n(\omega)\), so that \[\tilde{p}_{n}(\omega,z)=n^{-p/2}p_{n}\!\left(\omega,\hat{\theta}_{n}(\omega)+n^{-1/2}z\right).\] Below, \(B(x,r)\) and \(\overline{B}(x,r)\) are understood with respect to the Euclidean metric on \(\mathbb{R}^{p}\). The next result verifies the hypotheses of Proposition [prop:quadratic-envelope] pathwise.

Assume:

  1. For every \(x\in\mathbb{X}\), the map \(\theta\mapsto\ell(\theta,x)\) is twice continuously differentiable, and there exists a measurable \(H:\mathbb{X}\to[0,\infty)\) such that \[\sup_{\theta\in\mathbb{R}^{p}}\|\nabla_{\theta}^{2}\ell(\theta,x)\|\le H(x),\qquad\mathrm{E}[H(X_{1})]<\infty.\]

  2. \(\pi\) is strictly positive and twice continuously differentiable on \(\mathbb{R}^{p}\), and \[M_{\pi}=\sup_{\theta\in\mathbb{R}^{p}}\|\nabla_{\theta}^{2}\log\pi(\theta)\|<\infty.\]

  3. \(\bar{\eta}=\sup_{n\ge1}\eta_{n}<\infty\).

  4. There exist almost sure events \(\Omega_{\mathrm{loss}},\Omega_{\nabla}\in\mathfrak{A}\) such that, for every \(\omega\in\Omega_{\mathrm{loss}}\) and every \(n\ge1\), \(\hat{\theta}_{n}(\omega)\) is a global minimiser of \(\theta\mapsto\sum_{i=1}^{n}\ell(\theta,X_{i}(\omega))\), and, for every \(\omega\in\Omega_{\nabla}\), \[n^{-1/2}\left\|\nabla_{\theta}\log\pi\left(\hat{\theta}_{n}(\omega)\right)\right\|\to0.\]

  5. There exist a Borel probability measure \(\mu_{\infty}\) on \(\mathbb{R}^{p}\), a radius \(r>0\), and an event \(\Omega_{\mathrm{w}}\in\mathfrak{A}\) with \(\operatorname{pr}(\Omega_{\mathrm{w}})=1\) such that, for every \(\omega\in\Omega_{\mathrm{w}}\), \[\tilde{\mu}_n(\omega)\rightsquigarrow\mu_{\infty}\qquad\text{and}\qquad\mu_{\infty}(B(0,r))>0.\]

Then there exist constants \(c,C>0\) and an event \(\Omega_{0}\in\mathfrak{A}\) with \(\operatorname{pr}(\Omega_{0})=1\) such that, for every \(\omega\in\Omega_{0}\), there exists \(n_{0}(\omega)\in\mathbb{N}\) for which \[\tilde{p}_{n}(\omega,z)\ge c\exp\left(-\frac{C}{2}\|z\|^{2}\right)\qquad\text{for all }z\in\mathbb{R}^{p}\text{ and all }n\ge n_{0}(\omega).\label{eq:Gen95Bayes95Lower95bound}\tag{3}\]

4 Application to generalised linear models in generalised Bayes↩︎

We now specialise the generalised posterior (2 ) to canonical generalised linear models (GLMs) for which the loss is defined on all of \(\mathbb{R}^{p}\). Let \(X_{i}=(Y_{i},W_{i})\) be i.i.d.with law \(\mathrm{P}_{0}\), where \(Y_{i}:\Omega\to\mathbb{R}\) and \(W_{i}:\Omega\to\mathbb{R}^{p}\). Assume that, conditional on \(W_{i}=w\), the response \(Y_{i}\) has a canonical one-parameter exponential-family conditional density with natural parameter \(\theta^{\top}w\), namely \[p_{\theta}(y\mid w)=\exp\{y\theta^{\top}w-b(\theta^{\top}w)\},\qquad\theta\in\mathbb{R}^{p},\] where \(b\) is three times continuously differentiable and \(b''(u)>0\) for all \(u\in\mathbb{R}\). Up to an additive term depending only on \(y\), the associated loss is \[\ell(\theta,(y,w))=b(\theta^{\top}w)-y\theta^{\top}w.\] Taking \(\tilde{\nu}=\mathrm{N}(0,\mathbf{I}_{p})\), the next theorem verifies the hypotheses needed to apply Theorem [thm:gp-envelope], Proposition [prop:LAN-like], and the positive-part logarithmic-moment step in Proposition [prop:log-mom-oneQ] pathwise in this GLM setting.

Assume that the prior density \(\pi\) and the measurable centring sequence \((\hat{\theta}_{n})\) satisfy (C2) and (C4) of Theorem [thm:gp-envelope], that \(\eta_{n}\to\eta_{*}\in(0,\infty)\), and that \[H(X_{1})=\sup_{\theta\in\mathbb{R}^{p}}\|\nabla_{\theta}^{2}\ell(\theta,(Y_{1},W_{1}))\|\le\sup_{\theta\in\mathbb{R}^{p}}b''(\theta^{\top}W_{1})\|W_{1}\|^{2}\qquad\text{satisfies}\qquad\mathrm{E}[H(X_{1})]<\infty.\] Assume moreover that there exist \(\theta_{0}\in\mathbb{R}^{p}\) and \(r_{0}>0\) such that \[\mathrm{E}\!\left[W_{1}\{b'(\theta_{0}^{\top}W_{1})-Y_{1}\}\right]=0,\qquad\mathrm{E}[\|W_{1}\|\,|Y_{1}|]<\infty,\qquad\mathrm{E}|b(\theta^{\top}W_{1})|<\infty\quad\text{for all }\theta\in\mathbb{R}^{p},\] \[\mathrm{E}[W_{1}W_{1}^{\top}]\text{ exists and is positive definite},\qquad\mathrm{E}\left[\sup_{\theta\in\overline{B}(\theta_{0},r_{0})}|b'''(\theta^{\top}W_{1})|\,\|W_{1}\|^{3}\right]<\infty.\]

Under the preceding assumptions, the hypotheses of Theorem [thm:gp-envelope] hold. Consequently, there exist constants \(c,C>0\) and an event \(\Omega_{0}\in\mathfrak{A}\) with \(\operatorname{pr}(\Omega_{0})=1\) such that, for every \(\omega\in\Omega_{0}\), there exists \(n_{0}(\omega)\in\mathbb{N}\) for which 3 holds.

Gaussian location-scale variational classes satisfy the feasibility condition \(\nu_{n}^{*}=\tilde{\nu}\circ\gamma_{n}\in\mathcal{F}_{n}\) for all large \(n\). In particular, Theorem [thm:GLM] applies to the Gaussian variational GLM approximations considered by [17], [18].

5 Proofs↩︎

of Proposition [prop:log-mom-oneQ]. By [10], \(\mathrm{KL}(\nu_{n}^{*}\Vert\mu_{n})=\mathrm{KL}(\tilde{\nu}\Vert\tilde{\mu}_n)\), and bijectivity of \(\gamma_{n}\) implies \(\tilde{\nu}\ll\tilde{\mu}_n\). Writing \(\tilde{p}_{n}={\rm d}\tilde{\mu}_n/{\rm d}\lambda\) and \(\tilde{q}={\rm d}\tilde{\nu}/{\rm d}\lambda\), the chain rule for Radon–Nikodym derivatives gives that on \(\{\tilde{p}_{n}\neq 0\}\), \({\rm d}\tilde{\nu}/{\rm d}\tilde{\mu}_n=\tilde{q}/\tilde{p}_{n}\). As \(\{\tilde{p}_{n}=0\}\) is a \(\tilde{\nu}\) null set, \[\mathrm{KL}(\tilde{\nu}\Vert\tilde{\mu}_n)=\int\log\left(\frac{\tilde{q}}{\tilde{p}_{n}}\right){\rm d}\tilde{\nu}=\mathrm{E}_{\tilde{\nu}}\left[\log\left(\frac{\tilde{q}}{\tilde{p}_{n}}\right)\right].\] If the two sufficient log-moment bounds hold, then \[\log\left(\frac{\tilde{q}}{\tilde{p}_n}\right)\le |\log \tilde{q}|+(-\log \tilde{p}_n)_+,\] and therefore \[\limsup_{n\to\infty} m_n \leq \limsup_{n\to\infty} \mathrm{KL}(\nu_{n}^{*}\Vert\mu_{n}) < \infty.\] This establishes (A1). ◻

of Proposition [prop:LAN-like]. The bound on \(\tilde{p}_{n}\) implies \(-\log \tilde{p}_{n}(\cdot)\le-\log c+\frac{C}{2}\|\cdot\|^{2}\), hence uniformly over \(n\), \[\mathrm{E}_{\tilde{\nu}}(-\log \tilde{p}_{n})_+\le |\log c|+\frac{C}{2}\int\| z\|^{2}~\tilde{\nu}(\mathrm{d}z)<\infty.\] ◻

of Proposition [prop:quadratic-envelope]. Fix \(n\ge n_{0}\). Since \(\nabla g_{n}\) is \(L\)-Lipschitz, for every \(z\in\mathbb{R}^{p}\), \[g_{n}(z)=g_{n}(0)+\int_{0}^{1}\langle\nabla g_{n}(tz),z\rangle\,\mathrm{d}t.\] Hence \[\begin{align}g_{n}(z) & =g_{n}(0)+\langle\nabla g_{n}(0),z\rangle+\int_{0}^{1}\langle\nabla g_{n}(tz)-\nabla g_{n}(0),z\rangle\,\mathrm{d}t\\ & \ge g_{n}(0) - \|\nabla g_n(0)\|\, \|z\| -\int_{0}^{1}\|\nabla g_{n}(tz)-\nabla g_{n}(0)\|\,\|z\|\,\mathrm{d}t\\ & \ge g_{n}(0)-M\|z\|-\int_{0}^{1}Lt\,\|z\|^{2}\,\mathrm{d}t\\ & \ge g_{n}(0)-M\|z\|-\frac{L}{2}\|z\|^{2}. \end{align}\] Therefore \[\tilde{p}_{n}(z)\ge \tilde{p}_{n}(0)\exp\left(-M\|z\|-\frac{L}{2}\|z\|^{2}\right).\label{eq:quadratic-envelope-lower-step}\tag{4}\]

Applying the same argument but instead using that \(\langle\nabla g_{n}(x),z\rangle \leq \|\nabla g_n(x)\|\,\|z\|\), for every \(x\in\mathbb{R}^{p}\), \[g_{n}(x)\le g_{n}(0)+M\|x\|+\frac{L}{2}\|x\|^{2}.\] Hence, for every \(x\in B(0,r)\), \[\tilde{p}_{n}(x)\le \tilde{p}_{n}(0)\exp\left(Mr+\frac{L}{2}r^{2}\right).\] Integrating over \(B(0,r)\) and using assumption (B3) gives \[\alpha\le\tilde{\mu}_{n}(B(0,r))=\int_{B(0,r)}\tilde{p}_{n}(x)\,\mathrm{d}x\le\lambda\left(B(0,r)\right)\exp\left(Mr+\frac{L}{2}r^{2}\right)\tilde{p}_{n}(0),\] so that \[\tilde{p}_{n}(0)\ge\frac{\alpha}{\lambda\left(B(0,r)\right)}\exp\left(-Mr-\frac{L}{2}r^{2}\right)=:m_{0}>0.\] Substituting this bound into 4 yields \[\tilde{p}_{n}(z)\ge m_{0}\exp\left(-M\|z\|-\frac{L}{2}\|z\|^{2}\right).\] Finally, by Young’s inequality, \[M\|z\|\le\frac{M^{2}}{2}+\frac{\|z\|^{2}}{2},\] so \[\tilde{p}_{n}(z)\ge m_{0}\,e^{-M^{2}/2}\exp\left(-\frac{L+1}{2}\|z\|^{2}\right).\] Thus the conclusion holds with \[c=\frac{\alpha}{\lambda\left(B(0,r)\right)}\exp\left(-Mr-\frac{L}{2}r^{2}-\frac{M^{2}}{2}\right),\qquad C=L+1.\] ◻

of Theorem [thm:gp-envelope]. By the strong law of large numbers, \[\Omega_{H}=\left\{\omega\in\Omega: \frac{1}{n}\sum_{i=1}^{n}H(X_{i}(\omega))\to\mathrm{E}[H(X_{1})]\right\}\] has probability one. Let \[\Omega_{0}=\Omega_{H}\cap\Omega_{\mathrm{loss}}\cap\Omega_{\nabla}\cap\Omega_{\mathrm{w}},\] so that \(\operatorname{pr}(\Omega_{0})=1\). Fix \(\omega\in\Omega_{0}\) for the remainder of the proof and suppress the \(\omega\)-dependence when no confusion arises.

For each \(n\), define \[g_{n}(z)=\log \tilde{p}_{n}(z),\qquad z\in\mathbb{R}^{p}.\] Since \[g_{n}(z)=-\eta_{n}\sum_{i=1}^{n}\ell\!\left(\hat{\theta}_{n}+n^{-1/2}z,X_{i}\right)+\log\pi\!\left(\hat{\theta}_{n}+n^{-1/2}z\right)-\log\mathcal{Z}_{n}-\frac{p}{2}\log n,\] for a normalising constant \(\mathcal{Z}_{n}\) independent of \(z\), assumptions (C1) and (C2) imply that \(g_{n}\) is twice continuously differentiable and \[\nabla_{z}^{2}g_{n}(z)=\frac{1}{n}\left\{ -\eta_{n}\sum_{i=1}^{n}\nabla_{\theta}^{2}\ell\!\left(\hat{\theta}_{n}+n^{-1/2}z,X_{i}\right)+\nabla_{\theta}^{2}\log\pi\!\left(\hat{\theta}_{n}+n^{-1/2}z\right)\right\} .\] Hence \[\sup_{z\in\mathbb{R}^{p}}\|\nabla_{z}^{2}g_{n}(z)\|\le\eta_{n}\frac{1}{n}\sum_{i=1}^{n}H(X_{i})+\frac{M_{\pi}}{n}.\] Since \(\omega\in\Omega_{H}\), there exists \(n_{1}(\omega)\in\mathbb{N}\) such that, with \(L=\bar{\eta}\left(\mathrm{E}[H(X_{1})]+1\right)+M_{\pi}\), \[\sup_{z\in\mathbb{R}^{p}}\|\nabla_{z}^{2}g_{n}(z)\|\le L\qquad\text{for all }n\ge n_{1}(\omega).\] In particular, \(\nabla g_{n}\) is \(L\)-Lipschitz on \(\mathbb{R}^{p}\) for all \(n\ge n_{1}(\omega)\).

Next, \[\nabla_{z}g_{n}(0)=n^{-1/2}\left\{ -\eta_{n}\sum_{i=1}^{n}\nabla_{\theta}\ell(\hat{\theta}_{n},X_{i})+\nabla_{\theta}\log\pi(\hat{\theta}_{n})\right\} .\] Because \(\hat{\theta}_{n}\) is a global minimiser of \(\theta\mapsto\sum_{i=1}^{n}\ell(\theta,X_{i})\) and this criterion is continuously differentiable on \(\mathbb{R}^{p}\), we have \[\sum_{i=1}^{n}\nabla_{\theta}\ell(\hat{\theta}_{n},X_{i})=0.\] Therefore \[\nabla_{z}g_{n}(0)=n^{-1/2}\nabla_{\theta}\log\pi(\hat{\theta}_{n}).\] Since \(\omega\in\Omega_{\nabla}\), there exists \(n_{2}(\omega)\in\mathbb{N}\) such that \[\|\nabla_{z}g_{n}(0)\|\le1\qquad\text{for all }n\ge n_{2}(\omega).\]

Since \(\omega\in\Omega_{\mathrm{w}}\), we have \[\tilde{\mu}_n\rightsquigarrow\mu_{\infty}\qquad\text{and}\qquad\mu_{\infty}(B(0,r))>0.\] The Portmanteau theorem gives \[\liminf_{n\to\infty}\tilde{\mu}_n(B(0,r))\ge\mu_{\infty}(B(0,r))>0.\] Hence, with \[\alpha=\frac{1}{2}\mu_{\infty}(B(0,r))>0,\] there exists \(n_{3}(\omega)\in\mathbb{N}\) such that \[\tilde{\mu}_n(B(0,r))\ge\alpha\qquad\text{for all }n\ge n_{3}(\omega).\]

Applying Proposition [prop:quadratic-envelope] to the deterministic sequence \((\tilde{\mu}_n(\omega))_{n\ge1}\) with the constants \(L\), \(M=1\), \(\alpha\), and \(r\), we obtain constants \(c,C>0\), depending only on \(\bar{\eta},~\mathrm{E}[H(X_{1})],~M_{\pi},~r,~\mu_{\infty}(B(0,r)),~p\), and an index \[n_{0}(\omega)=\max\{n_{1}(\omega),n_{2}(\omega),n_{3}(\omega)\}\] such that \[\tilde{p}_{n}(\omega,z)\ge c\exp\left(-\frac{C}{2}\|z\|^{2}\right)\qquad\text{for all }z\in\mathbb{R}^{p}\text{ and all }n\ge n_{0}(\omega).\] ◻

of Theorem [thm:GLM]. We verify (C1), (C3), and (C5) of Theorem [thm:gp-envelope]; assumptions (C2) and (C4) are part of the statement. Write \[B_0=B(\theta_{0},r_{0}).\]

For \((y,w)\in\mathbb{R}\times\mathbb{R}^{p}\), \[\nabla_{\theta}\ell(\theta,(y,w))=\{b'(\theta^{\top}w)-y\}w,\qquad\nabla_{\theta}^{2}\ell(\theta,(y,w))=b''(\theta^{\top}w)\,ww^{\top}.\] Hence \[\|\nabla_{\theta}^{2}\ell(\theta,(Y_{1},W_{1}))\|=\|b''(\theta^{\top}W_{1})W_{1}W_{1}^{\top}\|\le H(X_{1}),\] since \(\|W_{1}W_{1}^{\top}\|=\|W_{1}\|^{2}\). Therefore (C1) holds. Since \(\eta_{n}\to\eta_{*}\in(0,\infty)\), we also have \(\bar{\eta}=\sup_{n\ge1}\eta_{n}<\infty\), so (C3) holds.

To verify (C5), set \[g_{n}(\theta)=\frac{1}{n}\sum_{i=1}^{n}\ell(\theta,X_{i})=\frac{1}{n}\sum_{i=1}^{n}\{b(\theta^{\top}W_{i})-Y_{i}\theta^{\top}W_{i}\},\qquad g(\theta)=\mathrm{E}[\ell(\theta,X_{1})],\] and define \[\tilde{g}_{n}(\theta)=\eta_{n}g_{n}(\theta),\qquad\tilde{g}(\theta)=\eta_{*}g(\theta).\] Since \[-\log p_{\theta}(Y_{i}\mid W_{i})=b(\theta^{\top}W_{i})-Y_{i}\theta^{\top}W_{i},\] up to an additive term independent of \(\theta\), the criterion \(g_{n}\) is exactly the empirical criterion appearing in [19] for canonical GLMs. The assumptions in the statement verify the hypotheses of that theorem. The only points requiring verification are identifiability and the third-derivative bound.

First, if \(a\in\mathbb{R}^{p}\) satisfies \(a^{\top}W_{1}=0\) almost surely, then \[\mathrm{E}[(a^{\top}W_{1})^{2}]=a^{\top}\mathrm{E}[W_{1}W_{1}^{\top}]a=0.\] Since \(\mathrm{E}[W_{1}W_{1}^{\top}]\) is positive definite, it follows that \(a=0\). Thus the identifiability condition in [19] holds. Second, for every \(j,k,\ell\in\{1,\ldots,p\}\), \[|W_{1j}W_{1k}W_{1\ell}|\le\|W_{1}\|^{3},\] and hence \[\mathrm{E}\!\left[\sup_{\theta\in\overline{B}(\theta_{0},r_{0})}|b'''(\theta^{\top}W_{1})|\,|W_{1j}W_{1k}W_{1\ell}|\right]\le\mathrm{E}\!\left[\sup_{\theta\in\overline{B}(\theta_{0},r_{0})}|b'''(\theta^{\top}W_{1})|\,\|W_{1}\|^{3}\right]<\infty.\]

Therefore [19] implies that, on an event \(\Omega_{\mathrm{M}}\in\mathfrak{A}\) with \(\operatorname{pr}(\Omega_{\mathrm{M}})=1\), the sequence \((g_{n}(\omega,\cdot))_{n\ge1}\) satisfies the hypotheses of case (2) of [19] for every \(\omega\in\Omega_{\mathrm{M}}\). Since \(\eta_{n}\to\eta_{*}\in(0,\infty)\), it follows that, for every \(\omega\in\Omega_{\mathrm{M}}\), the sequence \((\tilde{g}_{n}(\omega,\cdot))_{n\ge1}\) also satisfies the hypotheses of case (2) of [19], with limit \(\tilde{g}\). In particular, by [19], for every \(\omega\in\Omega_{\mathrm{M}}\), \[\tilde{g}_{n}(\omega,\cdot)\to\tilde{g}\qquad\text{and}\qquad\nabla_{\theta}^{2}\tilde{g}_{n}(\omega,\cdot)\to\nabla_{\theta}^{2}\tilde{g}\] uniformly on \(B_0\).

Let \[H_{3}(X_{1})=\sup_{\theta\in\overline{B}(\theta_{0},r_{0})}|b'''(\theta^{\top}W_{1})|\,\|W_{1}\|^{3}.\] By assumption, \(\mathrm{E}[H_{3}(X_{1})]<\infty\). Hence the strong law of large numbers yields an event \(\Omega_{3}\in\mathfrak{A}\) with \(\operatorname{pr}(\Omega_{3})=1\) such that \[\frac{1}{n}\sum_{i=1}^{n}H_{3}(X_{i})\to\mathrm{E}[H_{3}(X_{1})]\qquad\text{on }\Omega_{3}.\]

Fix \(\omega\in\Omega_{\mathrm{M}}\cap\Omega_{3}\cap\Omega_{\mathrm{loss}}\), where \(\Omega_{\mathrm{loss}}\) is the event in the first part of (C4), and suppress the \(\omega\)-dependence for the remainder of the proof. Since \(\eta_{n}>0\), the minimisers of \[\theta\mapsto\sum_{i=1}^{n}\ell(\theta,X_{i}),\qquad g_{n}(\cdot),\qquad\text{and}\qquad\tilde{g}_{n}(\cdot)\] coincide. Hence \(\hat{\theta}_{n}\) is a global minimiser of \(\tilde{g}_{n}\) for every \(n\). Because \(\mathbb{R}^{p}\) is open and \(\tilde{g}_{n}\) is continuously differentiable, we have \[\nabla_{\theta}\tilde{g}_{n}(\hat{\theta}_{n})=0\qquad\text{for every }n.\]

We next show that \(\hat{\theta}_{n}\to\theta_{0}\). Since case (2) of [19] implies case (1), there exists a compact set \(\mathbb{K}\subset B_0\) with \(\theta_{0}\in\mathrm{int}(\mathbb{K})\) such that \[\tilde{g}(\theta)>\tilde{g}(\theta_{0})\qquad\text{for all }\theta\in\mathbb{K}\setminus\{\theta_{0}\},\] and \[\liminf_{n\to\infty}\inf_{\theta\in\mathbb{R}^{p}\setminus\mathbb{K}}\tilde{g}_{n}(\theta)>\tilde{g}(\theta_{0}).\] Choose \(\rho>0\) such that \[\overline{B}(\theta_{0},\rho)\subset\mathrm{int}(\mathbb{K}).\] Since \(\tilde{g}\) is continuous on \(B_0\) and strictly larger than \(\tilde{g}(\theta_{0})\) on \(\mathbb{K}\setminus\{\theta_{0}\}\), there exists \(\delta>0\) such that \[\inf_{\theta\in\mathbb{K}\setminus\overline{B}(\theta_{0},\rho)}\tilde{g}(\theta)\ge\tilde{g}(\theta_{0})+4\delta.\] By uniform convergence of \(\tilde{g}_{n}\) to \(\tilde{g}\) on \(B_0\), by convergence of \(\tilde{g}_{n}(\theta_{0})\) to \(\tilde{g}(\theta_{0})\), and by the displayed liminf bound outside \(\mathbb{K}\), there exists \(n_{1}\in\mathbb{N}\) such that, for all \(n\ge n_{1}\), \[\tilde{g}_{n}(\theta_{0})\le\tilde{g}(\theta_{0})+\delta,\] \[\inf_{\theta\in\mathbb{K}\setminus\overline{B}(\theta_{0},\rho)}\tilde{g}_{n}(\theta)\ge\tilde{g}(\theta_{0})+3\delta,\qquad\inf_{\theta\in\mathbb{R}^{p}\setminus\mathbb{K}}\tilde{g}_{n}(\theta)\ge\tilde{g}(\theta_{0})+3\delta.\] Hence \[\inf_{\theta\in\mathbb{R}^{p}\setminus\overline{B}(\theta_{0},\rho)}\tilde{g}_{n}(\theta)\ge\tilde{g}(\theta_{0})+3\delta>\tilde{g}_{n}(\theta_{0}).\] Since \(\hat{\theta}_{n}\) is a global minimiser of \(\tilde{g}_{n}\), it follows that \[\hat{\theta}_{n}\in\overline{B}(\theta_{0},\rho)\qquad\text{for all }n\ge n_{1}.\] Because \(\rho>0\) may be chosen arbitrarily small, we conclude that \[\hat{\theta}_{n}\to\theta_{0}.\]

We now verify the hypotheses of [19]. Condition (2) holds because \[\nabla_{\theta}^{2}\tilde{g}_{n}(\theta_{0})\to\nabla_{\theta}^{2}\tilde{g}(\theta_{0}),\] and, for every \(a\in\mathbb{R}^{p}\setminus\{0\}\), \[\begin{align}a^{\top}\nabla_{\theta}^{2}\tilde{g}(\theta_{0})a & =\eta_{*}\,a^{\top}\mathrm{E}\!\left[b''(\theta_{0}^{\top}W_{1})W_{1}W_{1}^{\top}\right]a\\ & =\eta_{*}\,\mathrm{E}\!\left[b''(\theta_{0}^{\top}W_{1})(a^{\top}W_{1})^{2}\right]>0. \end{align}\] Here we used that \(\eta_{*}>0\), that \(b''>0\) by assumption, and that \(a^{\top}W_{1}\) is not almost surely zero by the identifiability argument above. Hence \(\nabla_{\theta}^{2}\tilde{g}(\theta_{0})\) is positive definite.

To verify Condition (3), note that for every \(j,k,\ell\in\{1,\ldots,p\}\), \[\sup_{\theta\in B_0}\left|\partial_{jk\ell}^{3}\tilde{g}_{n}(\theta)\right|\le\eta_{n}\frac{1}{n}\sum_{i=1}^{n}\sup_{u\in\overline{B}(\theta_{0},r_{0})}|b'''(u^{\top}W_{i})|\,|W_{ij}W_{ik}W_{i\ell}|\le\eta_{n}\frac{1}{n}\sum_{i=1}^{n}H_{3}(X_{i}).\] Since \(\omega\in\Omega_{3}\) and \(\eta_{n}\to\eta_{*}\), the right-hand side is eventually bounded in \(n\). Therefore the third derivatives of \(\tilde{g}_{n}\) are uniformly bounded on \(B_0\), so Condition (3) of [19] holds. Consequently, Assumption (1) of [19] is satisfied for \(\tilde{g}_{n}\) with centring sequence \(\hat{\theta}_{n}\).

To verify Assumption (2) of [19], fix \(\varepsilon>0\). Since case (2) of [19] implies case (1), and \(\theta_{0}\in\mathrm{int}(\mathbb{K})\), there exists \(\delta>0\) such that \[\inf_{\theta\in\mathbb{K}\cap B(\theta_{0},\varepsilon/2)^{c}}\tilde{g}(\theta)\ge\tilde{g}(\theta_{0})+4\delta.\] By uniform convergence of \(\tilde{g}_{n}\) to \(\tilde{g}\) on \(B_0\), by convergence of \(\tilde{g}_{n}(\theta_{0})\) to \(\tilde{g}(\theta_{0})\), and by the displayed liminf bound outside \(\mathbb{K}\), there exists \(n_{2}\in\mathbb{N}\) such that, for all \(n\ge n_{2}\), \[\tilde{g}_{n}(\theta_{0})\le\tilde{g}(\theta_{0})+\delta,\] \[\inf_{\theta\in\mathbb{K}\cap B(\theta_{0},\varepsilon/2)^{c}}\tilde{g}_{n}(\theta)\ge\tilde{g}(\theta_{0})+3\delta,\qquad\inf_{\theta\in\mathbb{R}^{p}\setminus\mathbb{K}}\tilde{g}_{n}(\theta)\ge\tilde{g}(\theta_{0})+3\delta.\] Hence \[\inf_{\theta\in B(\theta_{0},\varepsilon/2)^{c}}\tilde{g}_{n}(\theta)\ge\tilde{g}(\theta_{0})+3\delta\qquad\text{for all }n\ge n_{2}.\] Since \(\hat{\theta}_{n}\to\theta_{0}\), we have \[B(\theta_{0},\varepsilon/2)\subset B(\hat{\theta}_{n},\varepsilon)\] for all sufficiently large \(n\). Therefore, for all sufficiently large \(n\), \[\begin{align} \inf_{\theta\in B(\hat{\theta}_{n},\varepsilon)^{c}}\{\tilde{g}_{n}(\theta)-\tilde{g}_{n}(\hat{\theta}_{n})\} & \ge\inf_{\theta\in B(\theta_{0},\varepsilon/2)^{c}}\{\tilde{g}_{n}(\theta)-\tilde{g}_{n}(\hat{\theta}_{n})\}\\ & \ge\inf_{\theta\in B(\theta_{0},\varepsilon/2)^{c}}\{\tilde{g}_{n}(\theta)-\tilde{g}_{n}(\theta_{0})\}, \end{align}\] since \(\hat{\theta}_{n}\) minimises \(\tilde{g}_{n}\). It follows from the two preceding displays that \[\liminf_{n\to\infty}\inf_{\theta\in B(\hat{\theta}_{n},\varepsilon)^{c}}\{\tilde{g}_{n}(\theta)-\tilde{g}_{n}(\hat{\theta}_{n})\}>0.\] Thus Assumption (2) of [19] holds.

Since \(\pi\) is strictly positive and twice continuously differentiable by (C2), the prior assumptions in [19] are also satisfied. Finally, \[\mu_{n}(\mathrm{d}\theta)\propto\exp\{-n\tilde{g}_{n}(\theta)\}\pi(\theta)\,\mathrm{d}\theta.\] Hence [19] applies pathwise and yields a nondegenerate centred Gaussian probability measure \(\Phi\) on \(\mathbb{R}^{p}\) such that \[\sup_{A\in\mathfrak{B}(\mathbb{R}^{p})}\left|\tilde{\mu}_n(A)-\Phi(A)\right|\to0.\] In particular, \(\tilde{\mu}_n\rightsquigarrow\Phi\), and \[\Phi\left(B(0,r)\right)>0\qquad\text{for every }r>0.\] Thus (C5) of Theorem [thm:gp-envelope] holds on the almost sure event \[\Omega_{\mathrm{w}}=\Omega_{\mathrm{M}}\cap\Omega_{3}\cap\Omega_{\mathrm{loss}}.\] Theorem [thm:gp-envelope] therefore yields the conclusion. ◻

References↩︎

[1]
Blei, D. M., Kucukelbir, A. and McAuliffe, J. D. (2017). Variational inference: A review for statisticians. Journal of the American Statistical Association112, 859–877.
[2]
Zhang, C., Bütepage, J., Kjellström, H. and Mandt, S. (2019). Advances in variational inference. IEEE Transactions on Pattern Analysis and Machine Intelligence41, 2008–2026.
[3]
Chérief-Abdellatif, B.-E. and Alquier, P. (2018). Consistency of variational Bayes inference for estimation and model selection in mixtures. Electronic Journal of Statistics12, 2995–3035.
[4]
Ray, K. and Szabó, B. (2022). Variational Bayes for high-dimensional linear regression with sparse priors. Journal of the American Statistical Association117, 1270–1281.
[5]
Chérief-Abdellatif, B.-E. (2020). Convergence rates of variational inference in sparse deep learning. In International Conference on Machine Learning.
[6]
Alquier, P. and Ridgway, J. (2020). Concentration of tempered posteriors and of their variational approximations. The Annals of Statistics48, 1475–1497.
[7]
Yang, Y., Pati, D. and Bhattacharya, A. (2020). \(\alpha\)-variational inference with statistical guarantees. The Annals of Statistics48, 886–905.
[8]
Wang, Y. and Blei, D. M. (2019). Frequentist consistency of variational Bayes. Journal of the American Statistical Association114, 1147–1161.
[9]
Zhang, F. and Gao, C. (2020). Convergence rates of variational posterior distributions. The Annals of Statistics48, 2180–2207.
[10]
Polyanskiy, Y. and Wu, Y. (2025). Information Theory: From Coding to Learning. Cambridge University Press.
[11]
van der Vaart, A. W. and Wellner, J. A. (2023). Weak Convergence and Empirical Processes: With Applications to Statistics. Springer.
[12]
Domke, J., Garrigos, G. and Gower, R. (2023). Provable convergence guarantees for black-box variational inference. In Advances in Neural Information Processing Systems.
[13]
Kim, K., Oh, J., Wu, K., Ma, Y.-A. and Gardner, J. R. (2023). On the convergence of black-box variational inference. In Advances in Neural Information Processing Systems.
[14]
Lambert, M., Chewi, S., Bach, F., Bonnabel, S. and Rigollet, P. (2022). Variational inference via Wasserstein gradient flows. In Advances in Neural Information Processing Systems.
[15]
Wu, K. and Gardner, J. (2024). Understanding stochastic natural gradient variational inference. In International Conference on Machine Learning.
[16]
Bissiri, P. G., Holmes, C. C. and Walker, S. G. (2016). A general framework for updating belief distributions. Journal of the Royal Statistical Society: Series B (Statistical Methodology)78, 1103–1130.
[17]
Jaakkola, T. S. and Jordan, M. I. (2000). Bayesian parameter estimation via variational methods. Statistics and Computing10, 25–37.
[18]
Knowles, D. A. and Minka, T. (2011). Non-conjugate variational message passing for multinomial and binary regression. In Advances in Neural Information Processing Systems.
[19]
Miller, J. W. (2021). Asymptotic normality, concentration, and coverage of generalized posteriors. Journal of Machine Learning Research22, 1–53.