Beyond Global Divergences: A Local-Mass Perspective on Bayesian Inference


Abstract

Global objectives, such as KL divergence and ELBO, are widely used in Bayesian inference for measuring distributional discrepancy. This paper studies their “local-mass behaviour” that is not directly captured by such objectives. We introduce and use two mathematical tools: (1) Mass Index for recording the polynomial and logarithmic decay scales of local mass, and (2) regularised extended KL (RE-KL), a set-localised divergence that can be formulated in the presence of singular components. Mass Indices help characterise how Bayesian updating changes local mass: (1) power-log likelihood factors shift it explicitly, and (2) parameter-dependent supports, or their smooth softenings, may change the local scale through the amount of mass that remains near the parameter value. Using local RE-KL, we prove absolute, relative, and directional inequalities for comparing local small-ball masses under the two KL directions. Together, these results provide a local theoretical account of local mass behaviour. Experiments provide controlled illustrations of the local behaviour. Code is available at https://github.com/Forsythia0604/Local-Mass-Framework.

1 Introduction↩︎

Global objectives, such as Kullback-Leibler (KL) and the evidence lower bound (ELBO), are standard in Bayesian inference [1][4]. However, global closeness does not necessarily determine how probability mass behaves near a specified parameter value. Bayesian asymptotic theory has long made this local issue explicit through prior mass and small-ball conditions for posterior contraction, while sparse, singular, and constrained Bayesian models provide common settings in which the behaviour of mass near special parameter values is substantively important [5][7]. Such questions are not fully captured by a single global divergence value. A sequence of distributions can converge in KL divergence while still exhibiting different small-ball scaling around a fixed point. This paper develops a local-mass framework for Bayesian inference. The basic object of study is the small-ball probability \(p(B_r(\boldsymbol{\theta}))\), \(r \downarrow 0\), where \(B_r(\boldsymbol{\theta})\) denotes the ball of radius \(r\) centred at a parameter value \(\boldsymbol{\theta}\). The asymptotic behaviour of this quantity records how probability mass accumulates around \(\boldsymbol{\theta}\).

We use the term Mass Index for a concentration-oriented parametrisation of local small-ball behaviour. The object of interest is a probability measure near a fixed parameter value, and the notion is closely related to the local dimension theory of a measure [8][10]. Specifically, suppose that, as \(r\downarrow 0\), \(\mu(B_r(\boldsymbol{\theta})) \asymp r^a\). Here \(d\) is the ambient dimension, and \(a\) is the usual local scaling exponent. The power component of the Mass Index records the normalised reciprocal scale \(d/a\). Thus it reverses the usual local-dimension scale and expresses it as a concentration scale. The logarithmic component records the first log-order correction, when the power scale is well defined. This plays a role analogous to slowly varying corrections in regular variation [11][13]. The notation therefore distinguishes measures with the same leading polynomial order but different logarithmic local mass behaviour, such as smooth priors and horseshoe-type priors near sparse parameter values.

To compare local masses, we use regularised extended KL (RE-KL), a set-localised regularised \(f\)-divergence [14][16]. Its finite recession slope permits an extension to possibly singular measures by separating absolutely continuous and singular contributions. We therefore use RE-KL as a local \(f\)-divergence-based tool for comparing small-ball masses, not as a new global divergence. In the absolutely continuous case, it relates to Tsallis-type divergences, includes squared Hellinger distance at \(\alpha=1/2\), and converges to KL as \(\alpha\uparrow1\).

We prove that Bayesian updating preserves both power and logarithmic local mass scales, if the likelihood is locally regular around the parameter value. Thus Bayesian updating can be decomposed into a prior local scale and a likelihood-induced local correction. When the support is parameter-dependent, the likelihood may contain a hard support constraint that removes part of the prior mass near the parameter value. We prove that now local mass is governed by the fraction of prior mass that survives the constraint in shrinking neighbourhoods. If this surviving fraction decays at a polynomial or logarithmic rate, the posterior Mass Index shifts by the corresponding amount. We also show that smoothing hard constraints may not preserve this local structure: the zero-temperature limit of a softened likelihood and the small-neighbourhood limit need not commute.

We then study local discrepancy between two probability measures. RE-KL gives both absolute and relative bounds on local probability masses. These bounds imply directional sufficient conditions for comparing Mass Indices. The approximation-to-target direction is rigid: sufficiently small normalised local RE-KL forces equality of the local power scale, and also of the logarithmic scale when defined. The target-to-approximation direction is weaker: bounded local divergence prevents the approximation from being locally thinner than the target, but still allows extra local concentration. This expresses the usual asymmetry between KL directions directly in terms of local small-ball mass.

Numerical experiments illustrate the theory. Synthetic examples display regular mass, depletion, cusps, logarithmic corrections, and atoms. A UCI Bayesian logistic-regression experiment shows that regular Bayesian updating can greatly increase posterior mass near the Laplace mean while preserving the leading local power order. A final toy example demonstrates the directionality of local RE-KL, with one direction remaining bounded and the other diverging on shrinking neighbourhoods.

2 Related work↩︎

Variational inference and mode coverage. Divergence choice is known to affect the behaviour of variational approximations [2], [3]. Inclusive-KL and related \(f\)-divergence methods often encourage broader mass coverage than reverse-KL objectives [3], [14], [16][18]. These approaches are mainly global. Our use of local RE-KL is closer in spirit to the local analysis of Csiszár \(\phi\)-divergences by Avlogiaris et al. [15]. We study a narrower question: how Bayesian updating and variational approximation affect small-ball probabilities \(p(B_r(\boldsymbol{\theta}))\) as \(r\downarrow0\). This is also related to mode-coverage questions, where global criteria may fail to capture set-level or pointwise coverage [19], [20].

2.0.0.1 Regular variation and local asymptotic scales.

Our notation draws on two classical sources. The first is regular variation, which provides the language for separating a leading power law from slowly varying, including logarithmic, corrections [11][13]. The second is the local, or pointwise, dimension theory of measures, where the small-ball behaviour of \(p(B_r(\boldsymbol{\theta}))\) is summarised through limits of \(\log p(B_r(\boldsymbol{\theta}))/\log r\) [8][10]. We use these ideas only as a local small-ball bookkeeping device: for a fixed parameter value \(\boldsymbol{\theta}\), \(p(B_r(\boldsymbol{\theta}))\) records the leading power scale and, when present, the first logarithmic correction of the probability mass.

Bayesian pruning. Bayesian sparsification methods use priors to control how mass is allocated across parameter regions while avoiding excessive shrinkage of large signals [21][23]. These works motivate our local perspective, but they typically evaluate architecture size, contraction, objective value, or prediction [6], [7], [24]. We instead isolate a measure-level object: the asymptotic scale of local small-ball mass, independent of a particular architecture or optimisation procedure.

3 Preliminaries↩︎

3.0.0.1 Basics.

For non-negative functions \(f\) and \(g\), with respect to the limit under consideration, \(f=O(g)\), \(f=\Omega(g)\), and \(f\asymp g\) denote upper, lower, and two-sided comparison up to positive constants, respectively. For probability measures \(p\) and \(q\), \(p\ll q\) means that \(q(E)=0\Rightarrow p(E)=0\) for every measurable set \(E\). We write \(p\perp q\) if there exists measurable \(E\) such that \(p(\mathbb{R}^d\setminus E)=0\) and \(q(E)=0\). These notions are referred to as absolute continuity and mutual singularity, respectively.

Lemma 1 (Lebesgue decomposition; Theorem 3.8 [25]). For probability measures \(p\) and \(q\), there are unique measures \(q_{\ll p}\ll p\) and \(q_{\perp p}\perp p\) such that \(q=q_{\ll p}+q_{\perp p}\).

Definition 1 (Total variation [26]). Let \(p\) and \(q\) be probability measures on a measurable space \((\Theta,\mathcal{F})\). The total variation distance between \(p\) and \(q\) is \[\|p-q\|_{\mathrm{TV}} := \sup_{E\in\mathcal{F}}|p(E)-q(E)|.\]

Definition 2 (KL divergence [27]). Let \(p\) and \(q\) be probability measures on a measurable space \((\Theta,\mathcal{F})\). If \(q\ll p\), the KL divergence from \(q\) to \(p\) is \[D_{\mathrm{KL}}(q\|p) := \int_{\Theta} \log\!\left(\frac{\mathrm d q}{\mathrm d p}\right)\,\mathrm d q .\] Otherwise, we set \(D_{\mathrm{KL}}(q\|p)=\infty\).

3.0.0.2 Problem formulation.

Let \(\Theta={\mathbb{R}^d}\) be the parameter space, and let \(\mathcal{D}=(z_1,\ldots,z_n)\in\mathcal{Z}^n\) denote the observed data, where \(z_i\sim P_{\boldsymbol{\theta}},\,i=1,\ldots,n.\) The likelihood of the observed data under parameter \(\boldsymbol{\theta}\) is denoted by \(L(\boldsymbol{\theta}\mid\mathcal{D}).\) Let \(\pi_0\) be a prior probability measure on \(\Theta\). Whenever the marginal \(Z_{\mathcal{D}} := \int_{\Theta} L(\boldsymbol{\theta}\mid\mathcal{D})\,\pi_0(\mathrm d\boldsymbol{\theta})\in(0,\infty)\), the posterior is defined by \(\pi_{\mathcal{D}}(E) =Z_{\mathcal{D}}^{-1} {\int_E L(\boldsymbol{\theta}\mid\mathcal{D})\,\pi_0(\mathrm d\boldsymbol{\theta})} .\) This map from the prior \(\pi_0\) to the posterior \(\pi_{\mathcal{D}}\) is called the Bayesian update.

In this paper, Bayesian inference refers to the construction and representation of the posterior distribution induced by the prior and the observed data. The Bayesian update gives the target posterior \(\pi_{\mathcal{D}}\). Since this posterior is often analytically or computationally intractable, one may further replace it by an approximating measure \(q\) chosen from a tractable variational family \(\mathcal{Q}\subseteq\mathcal{P}(\Theta)\).

The variational approximation problem is to find \(q^*\) such that \(q^{*}\in \mathrm{argmin}_{q\in\mathcal{Q}} D_{\mathrm{KL}}(q\|\pi_{\mathcal{D}}).\) Or equivalently, maximising the evidence lower bound \(\mathcal{L}_{\mathcal{D}}(q) := \int_{\Theta}\log L(\boldsymbol{\theta}\mid\mathcal{D})\,q(\mathrm d\boldsymbol{\theta}) - D_{\mathrm{KL}}(q\|\pi_0).\)

4 The local mass framework↩︎

We now introduce the local mass framework. Throughout this paper, \(p\) and \(q\) denote two arbitrary probability measures on \(\Theta\).

4.1 Mass Index: On local small-ball scale↩︎

We now introduce the Mass Index as an order-scale summary of the local small-ball mass \(p(B_r(\boldsymbol{\theta})),r\downarrow0.\) The power component records the reciprocal of the sharp local power scale, normalised by the ambient dimension \(d\). The logarithmic component is defined only after the power scale is fixed, and records the first log-power correction. This construction is deliberately weaker than regular variation: it uses upper and lower power gauges, and does not require an exact asymptotic representation or a slowly varying factor.

Definition 3 (Power and Logarithmic Mass Indices). Let \(B_r(\boldsymbol{\theta})\) be the standard open ball of radius \(r\) centred at \(\boldsymbol{\theta}\). We adopt the convention that \(\sup \varnothing = 0\) and \(\inf \varnothing = \infty\).

Power Mass Index. As \(r\downarrow 0\), define \[\begin{align} \mathcal{U}_{\mathrm{pow}} := \Big\{\eta>0: p(B_r(\boldsymbol{\theta}))=O(r^{1/\eta}) \Big\},\\ \mathcal{L}_{\mathrm{pow}} := \Big\{\eta>0: p(B_r(\boldsymbol{\theta}))=\Omega(r^{1/\eta}) \Big\}. \end{align}\] The upper and lower Power Mass Indices are \[\overline{\mathrm{MI}}_{\mathrm{pow}}(p,\boldsymbol{\theta}) :=d\inf\mathcal{U}_{\mathrm{pow}},\, \underline{\mathrm{MI}}_{\mathrm{pow}}(p,\boldsymbol{\theta}) :=d\sup\mathcal{L}_{\mathrm{pow}}.\] If \(\overline{\mathrm{MI}}_{\mathrm{pow}}(p,\boldsymbol{\theta}) = \underline{\mathrm{MI}}_{\mathrm{pow}}(p,\boldsymbol{\theta}),\) then the common value is denoted by \(\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta}).\)

Logarithmic Mass Index. Assume that \(\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})\) exists and lies in \((0,\infty)\). Let \(a_{p,\boldsymbol{\theta}}:=\frac{d}{\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})}.\) As \(r\downarrow 0\), define \[\begin{align} &\mathcal{U}_{\mathrm{log}} := \big\{ \eta\in\mathbb{R}: p(B_r(\boldsymbol{\theta}))=O\bigl(r^{a_{p,\boldsymbol{\theta}}}(-\log r)^\eta\bigr) \big\},\\ &\mathcal{L}_{\mathrm{log}} := \big\{ \eta\in \mathbb{R}: p(B_r(\boldsymbol{\theta}))=\Omega\bigl(r^{a_{p,\boldsymbol{\theta}}}(-\log r)^\eta\bigr) \big\}. \end{align}\] The upper and lower Logarithmic Mass Indices are \[\overline{\mathrm{MI}}_{\mathrm{log}}(p,\boldsymbol{\theta}) := \inf\mathcal{U}_{\mathrm{log}},\, \underline{\mathrm{MI}}_{\mathrm{log}}(p,\boldsymbol{\theta}) := \sup\mathcal{L}_{\mathrm{log}}.\] If \(\overline{\mathrm{MI}}_{\mathrm{log}}(p,\boldsymbol{\theta}) = \underline{\mathrm{MI}}_{\mathrm{log}}(p,\boldsymbol{\theta}),\) then the common value is denoted by \(\mathrm{MI}_{\mathrm{log}}(p,\boldsymbol{\theta}).\)

We compare \((\mathrm{MI}_\mathrm{pow}(p,\boldsymbol{\theta}),\mathrm{MI}_\mathrm{log}(p,\boldsymbol{\theta}))\) using the lexicographic order. Throughout the paper, we denote \(\mathrm{MI}_{\mathrm{pow}}^{-1}(p,\boldsymbol{\theta}) := \frac{1}{\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})}\) whenever \(\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})\in(0,\infty)\).

4.1.1 Basic calibrations↩︎

The following proposition records the basic local calibrations of the Mass Index.

Proposition 1 (Basic local small-ball calibrations). Let \(p\) be a probability measure on \(\mathbb{R}^d\). For some \(r_0>0\):

(1) Local hole. If \(p(B_{r_0}(\boldsymbol{\theta}))=0\), then \(\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})=0\).

(2) Regular continuous mass. If \(p\) admits a continuous density \(\rho_p\) on \(B_{r_0}(\boldsymbol{\theta})\), and \(\rho_p(\boldsymbol{\theta})>0\), then \(\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})=1\).

(3) Power-log asymptotics. If \(p(B_r(\boldsymbol{\theta}))\asymp r^a(-\log r)^b, \, r\downarrow 0\) with \(a>0\), then \(\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})=da^{-1},\mathrm{MI}_{\mathrm{log}}(p,\boldsymbol{\theta})=b.\)

(4) Atomic mass. If \(p(\{\boldsymbol{\theta}\})>0\), then \(\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})=\infty\).

Interpretation. Proposition 1 gives basic local small-ball calibrations of the Mass Index. A local hole, regular continuous mass, and an atom correspond to \(\mathrm{MI}_{\mathrm{pow}}=0,1,\infty\), respectively. The factor \(d\) normalises the reciprocal exponent against the ordinary full-dimensional scaling \(r^d\).

Table 1 gives representative one-dimensional examples at \(\boldsymbol{\theta}=0\).

Table 1: Power and Logarithmic Mass Indices at \(0\).
Distribution \(\MI_{\mathrm{pow}}\) \(\MI_{\mathrm{log}}\)
Gaussian \(1\) \(0\)
Student-\(t_\nu\) \(1\) \(0\)
Laplace \(1\) \(0\)
Horseshoe \(1\) \(1\)
Spike-and-slab \(\infty\) -

4pt

4.1.2 Scope of the definition↩︎

Although the Horseshoe prior and the Gaussian prior both have \(\mathrm{MI}_{\mathrm{pow}}=1\), the Horseshoe prior is widely used in Bayesian compression [22], and its behaviour near zero is closely connected with an additional logarithmic-order factor. Thus, a logarithmic refinement is needed if one wants to distinguish these two priors beyond their leading power order.

The cases \(\mathrm{MI}_{\mathrm{pow}}=0\) and \(\mathrm{MI}_{\mathrm{pow}}=\infty\) correspond to behaviours outside the polynomial scale, such as super-polynomial or logarithmic-type decay. Thus, the Power Mass Index should be viewed as a modest power-scale calibration. It captures the dominant polynomial order when meaningful, but it does not distinguish all finer sub-polynomial or super-polynomial behaviours.

4.2 RE-KL: A set-localised divergence↩︎

Avlogiaris et al. [15] introduced local divergences based on Csiszár \(\phi\)-divergences, while Póczos et al. [28] studied global Tsallis-\(\alpha\) divergences and their nonparametric estimation. Our construction is related to both: we use a Tsallis-type generator within the extended \(f\)-divergence framework [14], [16] and localise it to a measurable set \(E\). The resulting quantity separates absolutely continuous and singular contributions and relates them to small-ball mass behaviour.

Definition 4 (Regularised extended Kullback-Leibler divergence). For \(\alpha\in(0,1)\), define \[f_\alpha(x) :=\frac{x^\alpha-\alpha x+\alpha-1}{{\alpha-1}}, \, x\in[0,\infty).\] The Lebesgue decomposition (Lemma 1) of \(q\) with respect to \(p\) is \(q=q_{\ll p}+q_{\perp p}.\) Then the RE-KL divergence on any measurable set \(E\) is defined as \[D_\alpha^E(q\|p):=\underbrace{\int_E\,f_\alpha\left(\frac{\mathrm dq_{\ll p}}{\mathrm dp}\right)\,\mathrm dp}_\text{absolutely continuous part}+\underbrace{\frac{\alpha}{1-\alpha}q_{\perp p}(E)}_\text{singular penalty}.\]

4.2.1 Relation to standard divergences↩︎

On the whole space \(\Theta\) with \(q\ll p\):

(1) KL divergence. When \(\alpha\uparrow 1\), \(D_\alpha^\Theta(q\|p)\uparrow D_{\mathrm{KL}}(q\|p)\).

(2) Hellinger distance \(\mathcal{H}(q,p)\) [29]. When \(\alpha=1/2\), \(D_{1/2}^\Theta(q\|p)=2\mathcal{H}^2(q,p)=2\bigl(1-\int_\Theta \sqrt{\mathrm dq/\mathrm dp}\,\mathrm dp\bigr)\).

(3) Tsallis-\(\alpha\) divergence \(T_\alpha\) [28]. When \(\alpha\in(0,1)\), \(D_{\alpha}^\Theta(q\|p)=T_\alpha(q\|p) = (\alpha-1)^{-1} \left( \int_{\Theta} q^\alpha(x)p^{1-\alpha}(x)\,\mathrm dx -1 \right).\)

Thus, the present definition can be viewed as a mild localisation of the standard extended \(f\)-divergence, isolating the contribution of a measurable set \(E\) while keeping track of the corresponding singular part.

4.3 Why global KL is not enough↩︎

Global KL divergence need not determine local probability mass. Related but different concerns appear in the GAN literature under mode collapse and incomplete mode coverage [19], [20]. We use this connection only as motivation for separating global closeness from local small-ball behaviour. The result below formulates this distinction through the Power Mass Index and shows that global KL convergence alone need not determine the local power order at \(\boldsymbol{\theta}\).

Theorem 2 (Sequential discontinuity of the Power Mass Index under KL convergence). Fix \(\boldsymbol{\theta}\in\Theta\). Define \(\mathcal{D}_{\boldsymbol{\theta}}:= \{q\mkern-2mu\in\mkern-2mu\mathcal{P}(\Theta): \mathrm{MI}_{\mathrm{pow}}(q,\boldsymbol{\theta})\in(0,\infty)\}.\)

Let \(p\in\mathcal{D}_{\boldsymbol{\theta}}\). Then, for every \(k>0\) with \(k\neq \mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})\), there exists a sequence \(\{q_n\}_{n\ge1}\subseteq \mathcal{D}_{\boldsymbol{\theta}}\) such that \(D_{\mathrm{KL}}(q_n\|p)\to0,\) but \(\mathrm{MI}_{\mathrm{pow}}(q_n,\boldsymbol{\theta})=k \, \text{for every } n\ge1.\) In particular, the functional \(\Phi_{\boldsymbol{\theta}}:\mathcal{D}_{\boldsymbol{\theta}}\to(0,\infty), \, \Phi_{\boldsymbol{\theta}}(q):=\mathrm{MI}_{\mathrm{pow}}(q,\boldsymbol{\theta})\) is not sequentially continuous at \(p\) under global KL convergence.

Interpretation. Theorem 2 gives a cautionary example for variational approximation. Even when \(D_{\mathrm{KL}}(q_n\|p)\to0\), global KL convergence alone need not preserve the local small-ball power order near \(\boldsymbol{\theta}\). KL-small perturbations may still change the Power Mass Index at \(\boldsymbol{\theta}\). This does not mean that larger variational families necessarily cause local distortion. It shows only that, without additional structural restrictions, such stability cannot be derived from the KL objective alone.

Proof sketch. Modify \(p\) only on \(B_{\rho_n}(\boldsymbol{\theta})\), where \(\rho_n\downarrow0\), and keep \(q_n=p\) outside. Since \(p\) has no atom at \(\boldsymbol{\theta}\), the modified mass tends to zero. Redistribute this mass so that the local power index at \(\boldsymbol{\theta}\) equals \(k\). Then \(\mathrm{MI}_{\mathrm{pow}}(q_n,\boldsymbol{\theta})=k\) for all \(n\), while the KL cost is confined to a set of vanishing \(p\)-mass, so \(D_{\mathrm{KL}}(q_n\|p)\to0\).

5 Local mass preservation theory↩︎

Using the two local tools developed in Section 4, we analyse the mathematical structure of local probability mass and local discrepancy under Bayesian update and variational approximation.

5.1 Bayesian updating and local mass scaling↩︎

We first study how Bayesian updating changes local mass.

5.1.1 Stability under Bayesian updating↩︎

We prove a result that characterises how the likelihood affects the preservation of local mass under Bayesian update. The analysis is based on the following two assumptions.

Assumption 1 (Finite marginal). \(0<Z_\mathcal{D}<\infty\).

Assumption 1 ensures that the posterior exists and is well-defined as a probability measure.

Assumption 2 (Well-defined Power Mass Indices). At \(\boldsymbol{\theta}\), \(\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})\) exists.

Assumption 2 rules out scale-wise oscillations of small-ball probabilities by requiring the upper and lower Power Mass Indices to coincide.

Example 1 shows that the MI may fail to exist. It is deliberately artificial: it builds infinite oscillations between two local power laws. Such behaviour is absent from standard priors with atoms, regular densities, finite mixtures, or regularly varying small-ball masses. Hence Assumption 2 rules out pathological oscillations, not ordinary Bayesian models.

Theorem 3 (Local likelihood scaling). Under Assumption 1 and 2 with \(p=\pi_0\), suppose that for some \(\gamma,\beta\in\mathbb{R}\), the likelihood satisfies \[L(\boldsymbol{x}\mid \mathcal{D}) \asymp\! \|\boldsymbol{x}-\boldsymbol{\theta}\|^\gamma \left(-\log {\|\boldsymbol{x}-\boldsymbol{\theta}\|}\right)^\beta,\,\boldsymbol{x}\to\boldsymbol{\theta}.\] When \(\mathrm{MI}_{\mathrm{pow}}(\pi_0,\boldsymbol{\theta}) \in(0,\infty)\), if \(\gamma >-{d}\,{\mathrm{MI}^{-1}_{\mathrm{pow}}(\pi_0,\boldsymbol{\theta})} ,\) then \(\mathrm{MI}_{\mathrm{pow}}(\pi_{\mathcal{D}},\boldsymbol{\theta})\) exists and \[\,{\mathrm{MI}^{-1}_{\mathrm{pow}}(\pi_{\mathcal{D}},\boldsymbol{\theta})} = \,{\mathrm{MI}^{-1}_{\mathrm{pow}}(\pi_0,\boldsymbol{\theta})} + d^{-1}\gamma.\]

Furthermore, if \(\mathrm{MI}_{\mathrm{log}}(\pi_0,\boldsymbol{\theta})\) exists, then \(\mathrm{MI}_{\mathrm{log}}(\pi_{\mathcal{D}},\boldsymbol{\theta})\) exists and \[\mathrm{MI}_{\mathrm{log}}(\pi_{\mathcal{D}},\boldsymbol{\theta}) = \mathrm{MI}_{\mathrm{log}}(\pi_0,\boldsymbol{\theta}) + \beta.\] When \(\mathrm{MI}_{\mathrm{pow}}(\pi_0,\boldsymbol{\theta})=0\), then \(\mathrm{MI}_{\mathrm{pow}}(\pi_{\mathcal{D}},\boldsymbol{\theta})\) exists and \[\mathrm{MI}_{\mathrm{pow}}(\pi_{\mathcal{D}},\boldsymbol{\theta}) = 0.\] The condition \(\gamma>-d\,\mathrm{MI}^{-1}_{\mathrm{pow}}(\pi_0,\boldsymbol{\theta})\) is a natural sufficient condition required by Assumption 1. We therefore restrict the theorem to this non-degenerate case and do not discuss the remaining degenerate cases in this paper.

Interpretation. Theorem 3 describes Bayesian updating as a local reweighting of the prior by the likelihood. If the likelihood is locally of power-log order, then the polynomial factor shifts the posterior Power Mass Index, while the logarithmic factor shifts only the posterior Logarithmic Mass Index.

In the regular case, where the likelihood is bounded above and below by positive constants near \(\boldsymbol{\theta}\), we have \(\gamma=\beta=0\). So Bayesian updating preserves both local mass scales. Thus local mass distortion under exact Bayesian updating arises only from non-regular local likelihood behaviour, such as zeros, singularities, logarithmic corrections, or support constraints.

Proof sketch. By Bayes’ formula, the problem reduces to estimating the likelihood-weighted prior mass of \(B_r(\boldsymbol{\theta})\). The upper estimates are obtained by decomposing \(B_r(\boldsymbol{\theta})\) into a sequence of shrinking annuli: on each annulus, the likelihood is controlled according to its distance scale, and the upper Mass Index bounds are applied to the prior mass. For the lower estimates, we restrict the integral to an outer annulus. On these annuli, the likelihood has a uniform lower bound, while the discarded inner part is negligible, yielding the matching lower order.

5.1.2 Parameter-dependent support and surviving local mass↩︎

In economics, parameter-dependent support arises in certain parametric auction models, search models, among others [30]. In these models, the boundary of the support of the observed data depends on some parameters of interest and on regressor variables. The likelihood can then be written as \[L(\boldsymbol{x}\mid \mathcal{D}) = g(\boldsymbol{x})\mathbf{1}_E(\boldsymbol{x}),\,\text{with}\,g(\boldsymbol{x})\asymp 1,\boldsymbol{x}\to \boldsymbol{\theta}.\label{eq:g1}\tag{1}\] The indicator \(\mathbf{1}_E\) may remove part, or all, of the local prior mass near \(\boldsymbol{\theta}\). The following theorem characterises the effect of such removal on local-mass preservation.

Theorem 4 (Acceptance-ratio scaling of local Mass Indices). Under Assumption 1, assume also that \[0<\mathrm{MI}_{\mathrm{pow}}(\pi_0,\boldsymbol{\theta})<\infty .\] The likelihood satisfies 1 . Define the local acceptance ratio \[A_{E,\boldsymbol{\theta}}(r) :=\frac{\pi_0(B_r(\boldsymbol{\theta})\cap E)}{\pi_0(B_r(\boldsymbol{\theta}))}\in[0,1],\] whenever \(\pi_0(B_r(\boldsymbol{\theta}))>0\). If \[A_{E,\boldsymbol{\theta}}(r)\asymp r^\delta\left(-\log{r}\right)^\lambda,\,r\to 0,\] where \(\delta>0\) or \(\delta=0,\lambda\le 0\), then \(\mathrm{MI}_{\mathrm{pow}}(\pi_{\mathcal{D}},\boldsymbol{\theta})\) exists and \[{\mathrm{MI}^{-1}_{\mathrm{pow}}(\pi_{\mathcal{D}},\boldsymbol{\theta})} = {\mathrm{MI}^{-1}_{\mathrm{pow}}(\pi_0,\boldsymbol{\theta})} +d^{-1} \delta.\]

Furthermore, if \(\mathrm{MI}_{\mathrm{log}}(\pi_0,\boldsymbol{\theta})\) exists, then \(\mathrm{MI}_{\mathrm{log}}(\pi_{\mathcal{D}},\boldsymbol{\theta})\) exists and \[\mathrm{MI}_{\mathrm{log}}(\pi_{\mathcal{D}},\boldsymbol{\theta}) = \mathrm{MI}_{\mathrm{log}}(\pi_0,\boldsymbol{\theta}) + \lambda.\]

Interpretation. The quantity \(A_{E,\boldsymbol{\theta}}(r)\) measures the fraction of the local prior mass around \(\boldsymbol{\theta}\) that is not removed by the hard constraint \(\mathbf{1}_E\). Since \(g(\boldsymbol{x})\asymp 1\) near \(\boldsymbol{\theta}\), the smooth part of the likelihood does not change the local mass scale. Hence the only local effect of the hard constraint is the multiplicative factor \(A_{E,\boldsymbol{\theta}}(r) \asymp r^\delta(-\log r)^\lambda.\) Thus the hard constraint adds \(\delta\) to the local power exponent and \(\lambda\) to the local logarithmic exponent.

Proof sketch. The proof is a direct analogue of Theorem 3. Since \(g(\boldsymbol{x})\asymp 1\) near \(\boldsymbol{\theta}\), Bayes’ formula gives \[\pi_{\mathcal{D}}(B_r(\boldsymbol{\theta})) \asymp \pi_0(B_r(\boldsymbol{\theta})\cap E) = A_{E,\boldsymbol{\theta}}(r)\pi_0(B_r(\boldsymbol{\theta})).\] Thus \(A_{E,\boldsymbol{\theta}}(r)\asymp r^\delta(-\log r)^\lambda\) shifts the local power and logarithmic exponents by \(\delta\) and \(\lambda\), respectively. The condition on \(\delta\) guarantees that the posterior power exponent remains positive. The identities follow.

In practice, hard constraints, indicator factors, and parameter-dependent supports are often smoothed for optimisation or sampling. Such softening is not locally innocuous: it can change small-ball mass scales and hence the local mass concentration induced by the hard formulation.

Corollary 1 (Softening and non-commuting local limits). Let \(\{s_\tau\}_{\tau>0}\) be positive continuous functions. Assume that there exists \(C>0\) such that \(0<s_\tau\le C\) for every \(\tau>0\), and that \[\lim_{\tau\downarrow0}s_\tau(\boldsymbol{x})=\mathbf{1}_E(\boldsymbol{x}) \quad\text{for }\pi_0\text{-a.e. }\boldsymbol{x} .\] Assume that Assumption 1 holds for the hard likelihood \(L=\mathbf{1}_E\). Assume also that \(0<\mathrm{MI}_{\mathrm{pow}}(\pi_0,\boldsymbol{\theta})<\infty .\) For each \(\tau>0\), denote by \(\pi_\tau\) the posterior obtained from the softened likelihood \(s_\tau\), and denote by \(\pi_{\mathcal{D}}\) the hard posterior obtained from \(\mathbf{1}_E\). Then \[\|\pi_\tau-\pi_{\mathcal{D}}\|_{\mathrm{TV}}\to0, \qquad \tau\downarrow0 .\] If \[A_{E,\boldsymbol{\theta}}(r)\asymp r^\delta(-\log r)^\lambda, \qquad r\downarrow0,\] with \(\delta>0\), then \[\mathrm{MI}_{\mathrm{pow}}(\pi_{\mathcal{D}},\boldsymbol{\theta}) \neq \lim_{\tau\downarrow0} \mathrm{MI}_{\mathrm{pow}}(\pi_\tau,\boldsymbol{\theta}).\] Furthermore, assume that \(\mathrm{MI}_{\mathrm{log}}(\pi_0,\boldsymbol{\theta})\) exists. If \[A_{E,\boldsymbol{\theta}}(r)\asymp (-\log r)^\lambda, \qquad r\downarrow0,\] with \(\lambda<0\), then \[\mathrm{MI}_{\mathrm{log}}(\pi_{\mathcal{D}},\boldsymbol{\theta}) \neq \lim_{\tau\downarrow0} \mathrm{MI}_{\mathrm{log}}(\pi_\tau,\boldsymbol{\theta}).\]

Interpretation. This corollary shows that the local limit \(r\to0\) and the zero-temperature limit \(\tau\downarrow0\) need not commute. For fixed \(\tau>0\), the soft factor \(s_\tau\) is locally positive near \(\boldsymbol{\theta}\), so the local MI is preserved. In the hard limit, the constraint may retain only a fraction \(A_{E,\boldsymbol{\theta}}(r)\asymp r^\delta(-\log r)^\lambda\) of the local prior mass, and therefore changes the local mass scale unless \(\delta=\lambda=0\). Thus softening is not locally innocuous: it may fail to reproduce the local mass-concentration structure of the hard-constrained model.

5.2 Local RE-KL analysis for Mass Indices↩︎

MI provides a way to characterise the local mass of a probability distribution. The local RE-KL bounds below give local bounds and sufficient conditions for comparing the mass behaviour of two measures.

5.2.0.1 Local RE-KL bounds.

We first show how the absolutely continuous component and the singular component jointly contribute to the absolute error.

Theorem 5 (Absolute error bound). For any measurable set \(E\), \[\bigl|q(E)-p(E)\bigr| \le 2\sqrt{C_\alpha D_\alpha^E(q\|p)},\] where \(C_\alpha=\max\left\{1,{\alpha}^{-1}({1-\alpha})\right\}.\)

Interpretation. Theorem 5 gives a local absolute error control for the mass assigned to a measurable region \(E\). Small local RE–KL discrepancy on \(E\) forces \(q(E)\) and \(p(E)\) to be close in additive mass. Importantly, the bound remains meaningful even when \(q\) has a \(p\)-singular component: singular mass does not make the right-hand side automatically infinite, but is charged through the finite recession penalty in \(D_\alpha^E(q\|p)\).

Proof sketch. Decompose \(q\) into its \(p\)-absolutely continuous and \(p\)-singular parts. For the former, rewrite the local mass error in Hellinger form and apply Cauchy–Schwarz. The RE–KL term dominates the resulting quantity up to \(C_\alpha\). The singular part is already charged by the recession term in \(D_\alpha^E(q\|p)\). Combining these two estimates yields the bound.

Theorem 5 leads to Corollary 2 directly.

Corollary 2. Assume that for some \(\delta>0\), \(D_{\alpha}^{B_r(\boldsymbol{\theta})}(q\|p)=O(r^{\delta}),\,r\to 0.\) We have \[\left|q(B_{r}(\boldsymbol{\theta}))-p(B_{r}(\boldsymbol{\theta}))\right| =O(r^{\delta/2}).\]

We next derive a relative bound for the local mass ratio \(q(E)/p(E)\).

Theorem 6 (Relative error bound). Denote the normalised RE-KL divergence as \(\overline{D}_\alpha^E(q\|p):={D_{\alpha}^E(q\|p)}/{p(E)}\) whenever \(p(E)>0\). Then \[f_{\alpha}\!\left(\frac{q(E)}{p(E)}\right) \le \overline{D}_\alpha^E(q\|p) .\] Furthermore, we can obtain a relative error bound: \[\left|1-\frac{q(E)}{p(E)}\right| \le\,\frac{2}{\alpha}\sqrt{{ \overline{D}_\alpha^E(q\|p)}}\sqrt{{ \overline{D}_\alpha^E(q\|p)}+2}.\]

Interpretation. Theorem 6 upgrades absolute error control to relative error control. This is needed for MI: when \(p(E)\) is itself small, absolute error is not enough to prevent \(q(E)\) from deviating too much in relative terms, whereas controlling \(q(E)/p(E)\) prevents asymptotic local under- or over-coverage.

Proof sketch. Jensen’s inequality controls the absolutely continuous part, while the recession slope of \(f_\alpha\) absorbs the singular part. Then a scalar lower bound for \(f_\alpha\) near its minimum at \(1\) converts this control into a quadratic inequality for \(|1-q(E)/p(E)|\), whose solution gives the displayed relative error bound.

Theorem 6 leads to Corollary 3 directly.

Corollary 3. Suppose \(p(B_r(\boldsymbol{\theta})) > 0\) for all sufficiently small \(r\). Assume that for some \(\delta>0\), \(\overline{D}_{\alpha}^{B_r(\boldsymbol{\theta})}(q\|p)=O(r^{\delta}),\,r\to 0.\) We have \[\left|1-\frac{q(B_{r}(\boldsymbol{\theta}))}{p(B_{r}(\boldsymbol{\theta}))}\right| =O(r^{\delta/2}).\]

5.2.0.2 Directional local RE–KL implications for Mass Indices.

We next use Theorem 6 to compare the local small-ball scales of two measures. In variational approximation, \(p=\pi_{\mathcal{D}}\) denotes the posterior and \(q\) denotes the approximation. The comparison of interest is the one-sided relation \(\mathrm{MI}_{\mathrm{pow}}(q,\boldsymbol{\theta})\ge\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})\). The following result gives directional sufficient conditions for this relation, and for equality in the stronger case.

Theorem 7 (Directional sufficient conditions for Power Mass Index comparison). Assume \(p(B_r(\boldsymbol{\theta}))>0\) and \(q(B_r(\boldsymbol{\theta}))>0\) for all sufficiently small \(r>0\). Suppose Assumption 2 holds with \(p\) and \(q\).

(i) If \(\limsup_{r\to0}{\overline{D}_\alpha^{B_r(\boldsymbol{\theta})}(q\|p)}<1,\)

then \(\mathrm{MI}_{\mathrm{pow}}(q,\boldsymbol{\theta})=\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})\). If \(\mathrm{MI}_{\mathrm{log}}(p,\boldsymbol{\theta})\) exists, then \(\mathrm{MI}_{\mathrm{log}}(q,\boldsymbol{\theta})=\mathrm{MI}_{\mathrm{log}}(p,\boldsymbol{\theta})\).

(ii) If \(\limsup_{r\to0}{\overline{D}_\alpha^{B_r(\boldsymbol{\theta})}(p\|q)}<\infty,\) then \(\mathrm{MI}_{\mathrm{pow}}(q,\boldsymbol{\theta})\ge\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})\). Furthermore, if \(\mathrm{MI}_{\mathrm{pow}}(q,\boldsymbol{\theta})=\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})\), and both \(\mathrm{MI}_{\mathrm{log}}(p,\boldsymbol{\theta})\) and \(\mathrm{MI}_{\mathrm{log}}(q,\boldsymbol{\theta})\) exist, then \(\mathrm{MI}_{\mathrm{log}}(q,\boldsymbol{\theta})\ge\mathrm{MI}_{\mathrm{log}}(p,\boldsymbol{\theta})\).

Interpretation. The two parts describe different strengths of local information. The \(q\|p\) direction is rigid: a merely finite value of \(\overline{D}_\alpha^{B_r(\boldsymbol{\theta})}(q\|p)\) is not enough at the power scale, and one needs the subcritical condition \(\limsup_{r\to0}\overline{D}_\alpha^{B_r(\boldsymbol{\theta})}(q\|p)<1\). Once this condition is imposed, there is no remaining freedom at the leading small-ball scale: \(q\) and \(p\) have the same power index, and also the same logarithmic index whenever the latter is defined for \(p\).

The \(p\|q\) direction is weaker but more permissive. Finiteness of \(\limsup_{r\to0}\overline{D}_\alpha^{B_r(\boldsymbol{\theta})}(p\|q)\) is enough only for a one-sided comparison: it precludes \(q\) from being locally thinner than \(p\) at the polynomial scale, but it still allows \(q\) to assign more mass near \(\boldsymbol{\theta}\). Thus this direction leaves room for local over-concentration of the approximation. This asymmetry is consistent with the usual distinction between KL directions in variational approximation [3], [17], [18]. These are conditional local implications, not a characterisation of minimisers of the global variational objective.

Proof sketch. Apply the local RE–KL mass bound to \(B_r(\boldsymbol{\theta})\). In the \(q\|p\) direction, the subcritical bound forces equality of the power scale, and then of the logarithmic scale. In the \(p\|q\) direction, finite normalised divergence only rules out polynomially smaller \(q(B_r(\boldsymbol{\theta}))\), giving the one-sided Power Mass Index bound. When the power scales coincide, the logarithmic comparison gives the final claim.

Example 2 illustrates that the two directions encode genuinely different local comparisons. The condition in the \(p\|q\) direction permits the conclusion that \(q\) has at least as large a local power mass index as \(p\), while the reverse condition is stronger here and is not satisfied.

6 Experiments↩︎

We give three small-scale checks of the proposed local quantities. These experiments are not intended as scalable diagnostics for large neural networks, but as controlled sanity checks. For the real-data experiment, we use three UCI binary tasks: Breast Cancer, Iris \(0\) vs \(1\), and Wine \(0\) vs \(1\). Each model is Bayesian logistic regression with four PCA covariates and an intercept, so \(d=5\). We use a Gaussian prior with scale \(2\), approximate the posterior by a Laplace Gaussian, set \(\boldsymbol{\theta}_0\) to the Laplace posterior mean, and estimate Euclidean small-ball masses by Sobol quadrature over \(r\in[0.03,0.45]\). UCI results are averaged over five seeds. Full experimental details are given in Appendix 10.

6.0.0.1 Experiment 1: Synthetic calibration.

Figure 1 shows the basic regimes used in the theory. A regular Gaussian follows the reference law \(p(B_r(\boldsymbol{\theta}_0))\asymp r^d\). The power-law examples \(p(B_r(\boldsymbol{\theta}_0))\asymp r^{d+\beta}\) show how the parameter \(\beta\) changes the local small-ball mass order. The log-singular and spike-slab examples illustrate logarithmic correction and atomic mass. The exact curves used in this calibration are listed in Appendix 10.1.

Figure 1: Synthetic small-ball calibration.

6.0.0.2 Experiment 2: UCI Bayesian sanity check.

Figure 2 shows that posterior small-ball mass around the Laplace mean is much larger than prior mass. As shown in Table 2, at the smallest radii, both prior and posterior slopes remain close to \(d=5\). This matches the regular case: Bayesian updating changes the local density level, but not the leading local power order. The posterior slope decrease at larger radii is a finite-radius effect. Dataset preprocessing, posterior fitting, quadrature, and slope computation are described in Appendix 10.2.

Figure 2: UCI prior/posterior small-ball masses and finite-radius slopes.
Table 2: Small-radius slopes fitted over \(r\leq0.06\).
Dataset \(n\) prior posterior
Breast Cancer 398 \(5.000\pm0.000\) \(4.969\pm0.005\)
Iris \(0\) vs \(1\) 70 \(5.000\pm0.000\) \(4.999\pm0.000\)
Wine \(0\) vs \(1\) 91 \(5.000\pm0.000\) \(4.993\pm0.001\)

3pt

6.0.0.3 Experiment 3: Directionality of local RE-KL.

Figure 3 illustrates that local RE-KL control is directional. On the same shrinking neighbourhoods, the \(p\|q\) direction stays bounded while the \(q\|p\) direction diverges. Thus the order of the two arguments determines which local mismatch is detected. The explicit construction is given in Appendix 10.3.

Figure 3: Directionality of local RE-KL.

7 Conclusion↩︎

This paper studied local small-ball mass \(p(B_r(\boldsymbol{\theta}))\) as a measure-level quantity in Bayesian inference. The aim was not to propose a new global divergence or a scalable diagnostic method, but to make precise how local mass scales can change under Bayesian updating and approximation.

We made three concrete contributions. First, we used the Mass Index to record the leading power order and the first logarithmic correction of local small-ball probabilities. Second, we derived local scale identities for Bayesian updating, showing how regular likelihoods preserve the local scale, how power-log likelihood factors shift it, and how parameter-dependent supports act through the surviving fraction of prior mass. Third, we used local RE-KL to prove absolute, relative, and directional bounds for comparing small-ball masses, yielding sufficient conditions for equality or one-sided comparison of Mass Indices under the two KL directions.

These results should be read as local asymptotic scale comparisons rather than a characterisation of global variational minimisers. The experiments provide finite-radius illustrations of the quantities and regimes appearing in the theory.

8 Examples and Counterexamples↩︎

Example 1 (Oscillating small-ball order). This example shows why the upper and lower Power Mass Indices need not coincide. We construct a probability measure whose small-ball mass behaves like \(r^a\) along one sequence of radii, but like \(r^b\) along another sequence of radii, where \(0<a<b\).

For simplicity, take \(\Theta=\mathbb{R}^d\). The same construction works whenever \(\Theta\) contains a sufficiently small line segment starting from \(\boldsymbol{\theta}\). Fix \(0<a<b\), and set \[t_n:=\left(\frac{2b}{a}\right)^n, \qquad n=0,1,2,\ldots .\] We define an increasing continuous function \(\phi\) on large positive values of \(t\) by prescribing its values on the sequence \((t_n)\), and then interpolating linearly between consecutive points: \[\phi(t_{2m})=a t_{2m}, \qquad \phi(t_{2m+1})=b t_{2m+1}.\] The prescribed values are strictly increasing. Indeed, \[b t_{2m+1}>a t_{2m}, \qquad a t_{2m+2}=2b t_{2m+1}.\] Hence the piecewise-linear interpolation is increasing.

Now define, for small \(r>0\), \[m(r):=\exp\{-\phi(\log(1/r))\}.\] As \(r\downarrow0\), we have \(\log(1/r)\to\infty\). Since \(\phi\) is increasing and tends to infinity, \(m(r)\downarrow0\). Also, as a function of \(r\), \(m(r)\) is increasing. Thus \(m\) can be used as the distribution function of radial mass near \(\boldsymbol{\theta}\).

Choose \(\varepsilon>0\) such that \(m(\varepsilon)<1\). By the Lebesgue–Stieltjes construction, there exists a finite measure \(\nu\) on \((0,\varepsilon]\) such that \[\nu((0,r])=m(r), \qquad 0<r\le \varepsilon .\] Push this measure forward by the map \[s\mapsto \boldsymbol{\theta}+s e_1,\] where \(e_1\) is the first coordinate vector. Finally, place the remaining mass \(1-m(\varepsilon)\) outside \(B_\varepsilon(\boldsymbol{\theta})\). This defines a probability measure \(p\).

Since \(m\) is continuous, the measure \(\nu\) has no atoms. Therefore no mass is placed on spheres centred at \(\boldsymbol{\theta}\). Consequently, for every \(0<r\le \varepsilon\), \[p(B_r(\boldsymbol{\theta}))=m(r).\]

Let \[r_n:=e^{-t_n}.\] Then \(\log(1/r_n)=t_n\). Along the even subsequence, we obtain \[p(B_{r_{2m}}(\boldsymbol{\theta})) = \exp\{-\phi(t_{2m})\} = \exp\{-a t_{2m}\} = r_{2m}^a.\] Along the odd subsequence, we obtain \[p(B_{r_{2m+1}}(\boldsymbol{\theta})) = \exp\{-\phi(t_{2m+1})\} = \exp\{-b t_{2m+1}\} = r_{2m+1}^b.\]

Moreover, by construction, the ratio \(\phi(t)/t\) stays between \(a\) and \(b\) on each interpolation interval. Hence, for all sufficiently small \(r\), \[r^b \le p(B_r(\boldsymbol{\theta})) \le r^a .\] The even subsequence shows that the upper power order cannot be better than \(a\), while the odd subsequence shows that the lower power order cannot be better than \(b\). Therefore \[\overline{\mathrm{MI}}_{\mathrm{pow}}(p,\boldsymbol{\theta})=\frac{d}{a}, \qquad \underline{\mathrm{MI}}_{\mathrm{pow}}(p,\boldsymbol{\theta})=\frac{d}{b}.\] Since \(a<b\), these two values are different. Hence the Power Mass Index at \(\boldsymbol{\theta}\) does not exist.

Example 2 (Directionality of local RE–KL conditions). This example illustrates that the local RE–KL conditions are directional. We work at the point \(0\). Let \(p\) and \(q\) be probability measures on \([-1,1]\) with densities \[\frac{\mathrm d p}{\mathrm d x}(x)=\frac{1}{2}, \qquad \frac{\mathrm d q}{\mathrm d x}(x)=\frac{1}{4} |x|^{-1/2}.\] Both are probability measures, since \[\int_{-1}^{1} \frac{1}{2} \,\mathrm dx =1, \qquad \int_{-1}^{1} \frac{1}{4} |x|^{-1/2}\,\mathrm dx =1.\] For \(0<r<1\), writing \(B_r=B_r(0)=(-r,r)\), we have \[p(B_r) = \int_{-r}^{r}\frac{1}{2}\,\mathrm dx = r,\] whereas \[q(B_r) = \int_{-r}^{r}\frac{1}{4} |x|^{-1/2}\,\mathrm dx = r^{1/2}.\] Therefore \[\mathrm{MI}_{\mathrm{pow}}(p,0)=1, \qquad \mathrm{MI}_{\mathrm{pow}}(q,0)=2.\] Thus \(q\) has more local mass near \(0\) than \(p\), in the power-order sense.

We first consider the direction \(p\|q\). Since \[\frac{\mathrm d p}{\mathrm d q}(x) = 2|x|^{1/2},\] the likelihood ratio tends to \(0\) as \(x\to0\). Moreover, on \(B_r\), \[0\le \frac{\mathrm d p}{\mathrm d q}(x)\le 2r^{1/2}.\] Hence the ratio \(\mathrm d p/\mathrm d q\) converges uniformly to \(0\) on \(B_r\) as \(r\downarrow0\). By continuity of \(f_\alpha\) at \(0\), \[f_\alpha\!\left(\frac{\mathrm d p}{\mathrm d q}(x)\right) \longrightarrow f_\alpha(0) \qquad \text{uniformly on }B_r.\] Consequently, \[\overline{D}_\alpha^{B_r}(p\|q) := \frac{1}{q(B_r)} \int_{B_r} f_\alpha\!\left(\frac{\mathrm d p}{\mathrm d q}\right) \,\mathrm d q \longrightarrow f_\alpha(0) = 1.\] Thus the local RE–KL condition in the direction \(p\|q\) is satisfied. The resulting one-sided conclusion is exactly the expected one: \[\mathrm{MI}_{\mathrm{pow}}(q,0) \ge \mathrm{MI}_{\mathrm{pow}}(p,0).\]

The reverse direction behaves differently. In the direction \(q\|p\), we have \[\frac{\mathrm d q}{\mathrm d p}(x) = \frac{1}{2} |x|^{-1/2},\] which diverges as \(x\to0\). For large \(t\), \[f_\alpha(t) \sim \frac{\alpha}{1-\alpha}t.\] Therefore, near \(0\), \[f_\alpha\!\left(\frac{\mathrm d q}{\mathrm d p}(x)\right) \asymp \frac{\mathrm d q}{\mathrm d p}(x).\] It follows that \[\int_{B_r} f_\alpha\!\left(\frac{\mathrm d q}{\mathrm d p}\right) \,\mathrm d p \asymp \int_{B_r} \frac{\mathrm d q}{\mathrm d p} \,\mathrm d p = q(B_r).\] After normalisation by \(p(B_r)\), we obtain \[\overline{D}_\alpha^{B_r}(q\|p) := \frac{1}{p(B_r)} \int_{B_r} f_\alpha\!\left(\frac{\mathrm d q}{\mathrm d p}\right) \,\mathrm d p \asymp \frac{q(B_r)}{p(B_r)} = \frac{r^{1/2}}{r} = r^{-1/2} \longrightarrow \infty.\] Hence the corresponding \(q\|p\) local RE–KL condition fails. This shows that the two directions are not interchangeable: \(p\|q\) controls a different local comparison from \(q\|p\).

9 Proofs↩︎

Lemma 2 (Gauge identification). Let \(m(r)\ge0\) be a small-ball mass function and let \(\ell(r):=-\log r\).

(i) Assume \(a>0\). If, for every \(\varepsilon>0\), \[m(r)=O(r^{a-\varepsilon}), \qquad m(r)=\Omega(r^{a+\varepsilon}), \qquad r\downarrow0,\] then the Power Mass Index exists and equals \(d/a\).

(ii) Assume that the power exponent is \(a>0\). If, for some \(b\in\mathbb{R}\) and every \(\varepsilon>0\), \[m(r)=O\!\left(r^a\ell(r)^{b+\varepsilon}\right), \qquad m(r)=\Omega\!\left(r^a\ell(r)^{b-\varepsilon}\right), \qquad r\downarrow0,\] then the Logarithmic Mass Index exists and equals \(b\).

Proof. For (i), the upper bounds imply that every exponent strictly below \(a\) is an upper power gauge, while the lower bounds imply that every exponent strictly above \(a\) is a lower power gauge. Therefore the sharp threshold in exponent form is \(a\). Since Definition 3 parametrises power bounds as \(r^{1/\eta}\), this gives \[\inf\mathcal{U}_{\mathrm{pow}}=\sup\mathcal{L}_{\mathrm{pow}}=\frac{1}{a},\] and hence \(\mathrm{MI}_{\mathrm{pow}}=d/a\).

For (ii), the definition of the logarithmic gauge is already taken at the identified power scale \(a\). The upper bounds imply that every logarithmic exponent strictly above \(b\) is an upper log gauge, and the lower bounds imply that every logarithmic exponent strictly below \(b\) is a lower log gauge. Thus \[\inf\mathcal{U}_{\mathrm{log}}=\sup\mathcal{L}_{\mathrm{log}}=b,\] so \(\mathrm{MI}_{\mathrm{log}}=b\). ◻

Lemma 3 (Radial power transform has bounded KL). Let \(\mu\) be a probability measure on \([0,\rho)\), and write \(F(r):=\mu([0,r))\). Fix \(t>0\), and let \(\nu\) be the probability measure on \([0,\rho)\) determined by \[\nu([0,r))=F(r)^t, \qquad 0<r<\rho .\] Then \(\nu\ll\mu\), and \[D_{\mathrm{KL}}(\nu\|\mu) \le \log t+\frac{1}{t}-1 .\]

Proof. Let \(\lambda\) be Lebesgue measure on \([0,1]\), and let \(\lambda_t\) have density \[\frac{\mathrm d\lambda_t}{\mathrm d\lambda}(u)=t u^{t-1}, \qquad 0<u<1 .\] Let \(Q\) be a quantile map that pushes \(\lambda\) forward to \(\mu\). Then the measure defined above is the pushforward \(Q_\#\lambda_t\). Since \(\lambda_t\ll\lambda\), we have \(\nu\ll\mu\). By data processing for KL divergence, \[D_{\mathrm{KL}}(\nu\|\mu) \le D_{\mathrm{KL}}(\lambda_t\|\lambda).\] The right-hand side is explicit: \[D_{\mathrm{KL}}(\lambda_t\|\lambda) =\int_0^1t u^{t-1}\log(tu^{t-1})\,\mathrm d u =\log t+\frac{1}{t}-1 .\] This proves the claim. ◻

Proof of Proposition 1↩︎

Proof. Let \[m(r):=p(B_r(\boldsymbol{\theta})), \qquad \ell(r):=-\log r .\]

(1) Local hole. If \(p(B_{r_0}(\boldsymbol{\theta}))=0\), then \(m(r)=0\) for all \(0<r<r_0\). Hence \(m(r)=O(r^{1/\eta})\) for every \(\eta>0\), so \(\inf\mathcal{U}_{\mathrm{pow}}=0\). On the other hand, \(m(r)=\Omega(r^{1/\eta})\) holds for no \(\eta>0\). By the convention \(\sup\emptyset=0\), we also have \(\sup\mathcal{L}_{\mathrm{pow}}=0\). Thus \[\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})=0 .\]

(2) Regular continuous mass. Suppose that \(p\) admits a continuous density \(\rho_p\) on \(B_{r_0}(\boldsymbol{\theta})\) and \(\rho_p(\boldsymbol{\theta})>0\). Then, for some \(r_1\in(0,r_0)\) and constants \(0<c<C<\infty\), \[c\le \rho_p(x)\le C, \qquad x\in B_{r_1}(\boldsymbol{\theta}).\] Therefore, for \(0<r<r_1\), \[c\,\mathrm{Vol}(B_r(\boldsymbol{\theta})) \le p(B_r(\boldsymbol{\theta})) \le C\,\mathrm{Vol}(B_r(\boldsymbol{\theta})).\] Since \(\mathrm{Vol}(B_r(\boldsymbol{\theta}))\asymp r^d\), we get \[p(B_r(\boldsymbol{\theta}))\asymp r^d .\] The definition of the Power Mass Index then gives \[\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})=1 .\]

(3) Power-log asymptotics. Assume \[p(B_r(\boldsymbol{\theta}))\asymp r^a \ell(r)^b, \qquad r\downarrow0 .\] For every \(\varepsilon>0\), logarithmic factors are slower than powers, hence \[r^a\ell(r)^b=O(r^{a-\varepsilon}), \qquad r^a\ell(r)^b=\Omega(r^{a+\varepsilon}).\] Lemma 2(i) gives \[\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})=\frac{d}{a} .\] At the corresponding power scale \(a_{p,\boldsymbol{\theta}}=a\), the two-sided asymptotic also gives, for every \(\varepsilon>0\), \[p(B_r(\boldsymbol{\theta}))=O\!\left(r^a\ell(r)^{b+\varepsilon}\right), \qquad p(B_r(\boldsymbol{\theta}))=\Omega\!\left(r^a\ell(r)^{b-\varepsilon}\right).\] Lemma 2(ii) gives \[\mathrm{MI}_{\mathrm{log}}(p,\boldsymbol{\theta})=b .\]

(4) Atomic mass. If \(p(\{\boldsymbol{\theta}\})>0\), then \(m(r)\ge p(\{\boldsymbol{\theta}\})>0\) for every \(r>0\). Hence \(m(r)=O(r^{1/\eta})\) holds for no \(\eta>0\), so \(\inf\mathcal{U}_{\mathrm{pow}}=\infty\). Also \(m(r)=\Omega(r^{1/\eta})\) holds for every \(\eta>0\), so \(\sup\mathcal{L}_{\mathrm{pow}}=\infty\). Therefore \[\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})=\infty .\] The proof is complete. ◻

Proof of the RE-KL relations↩︎

Proof. Assume first that \(q\ll p\), and write \[R:=\frac{\mathrm d q}{\mathrm d p}.\] Then \[D_\alpha^\Theta(q\|p) =\int_\Theta f_\alpha(R)\,\mathrm d p =\frac{\int_\Theta R^\alpha\,\mathrm d p-1}{\alpha-1} =\frac{1-\int_\Theta R^\alpha\,\mathrm d p}{1-\alpha}.\] If both measures are represented by densities with respect to a common dominating measure, this is exactly \[(\alpha-1)^{-1} \left(\int q(x)^\alpha p(x)^{1-\alpha}\,\mathrm d x-1\right),\] which is the Tsallis-\(\alpha\) form used in the main text.

For \(\alpha=1/2\), \[f_{1/2}(x)=(\sqrt x-1)^2 .\] Therefore \[D_{1/2}^\Theta(q\|p) =\int_\Theta (\sqrt R-1)^2\,\mathrm d p =2\left(1-\int_\Theta \sqrt R\,\mathrm d p\right) =2\mathcal{H}^2(q,p),\] under the Hellinger convention used in the paper.

It remains to justify the monotone KL limit. For \(x>0\), differentiating \[f_\alpha(x)=\frac{\alpha x-x^\alpha+1-\alpha}{1-\alpha}\] with respect to \(\alpha\) gives \[\partial_\alpha f_\alpha(x) =\frac{x-x^\alpha\{1+(1-\alpha)\log x\}}{(1-\alpha)^2} =\frac{x^\alpha}{(1-\alpha)^2} \left(x^{1-\alpha}-1-\log x^{1-\alpha}\right).\] The elementary inequality \(u-1-\log u\ge0\), \(u>0\), therefore implies \(\partial_\alpha f_\alpha(x)\ge0\). At \(x=0\), \(f_\alpha(0)=1\) for all \(\alpha\in(0,1)\), so monotonicity is trivial. Moreover, for each fixed \(x\ge0\), a Taylor expansion of \(x^\alpha\) at \(\alpha=1\), with the endpoint values interpreted by continuity, gives \[f_\alpha(x)\uparrow x\log x-x+1, \qquad \alpha\uparrow1 .\] By monotone convergence, \[\lim_{\alpha\uparrow1}D_\alpha^\Theta(q\|p) =\int_\Theta (R\log R-R+1)\,\mathrm d p =\int_\Theta R\log R\,\mathrm d p =D_{\mathrm{KL}}(q\|p),\] because \(\int R\,\mathrm d p=1\). If \(q\not\ll p\), then the singular part is charged by \[\frac{\alpha}{1-\alpha}q_{\perp p}(\Theta),\] which diverges as \(\alpha\uparrow1\), matching \(D_{\mathrm{KL}}(q\|p)=\infty\). This proves the stated relations. ◻

Proof of Theorem 2↩︎

Proof. Let \[m:=\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})\in(0,\infty).\] Since the Power Mass Index of \(p\) at \(\boldsymbol{\theta}\) exists, Definition 3 gives \[\inf\mathcal{U}_{\mathrm{pow}}^p = \sup\mathcal{L}_{\mathrm{pow}}^p = \frac{m}{d},\] where the superscript indicates that the corresponding sets are defined with respect to \(p\).

Fix \(k>0\) with \(k\neq m\), and set \[t:=\frac{m}{k}>0 .\] Choose a sequence \(\rho_n\downarrow0\), and define \[E_n:=B_{\rho_n}(\boldsymbol{\theta}), \qquad a_n:=p(E_n).\] Because \(m<\infty\), the measure \(p\) has no atom at \(\boldsymbol{\theta}\), and \(a_n\to0\). Since \(m>0\), there is no local hole around \(\boldsymbol{\theta}\), so \(a_n>0\) for all sufficiently large \(n\). Passing to this tail of the sequence, define the conditional probability measure \[p_n(A):=\frac{p(A\cap E_n)}{a_n}.\]

Let \[R(x):=\|x-\boldsymbol{\theta}\|, \qquad F_n(r):=p_n(B_r(\boldsymbol{\theta})), \qquad 0<r<\rho_n .\] Let \(\mu_n\) be the radial distribution of \(R\) under \(p_n\). To keep the open-ball convention used in the Mass Index definition, we use half-open radial intervals and write \[\mu_n([0,r))=F_n(r)=p_n(B_r(\boldsymbol{\theta})).\] This convention avoids any ambiguity from possible mass on spheres. Define a new radial probability measure \(\nu_n\) by \[\nu_n([0,r))=F_n(r)^t, \qquad 0<r<\rho_n .\] Disintegrate \(p_n\) with respect to the radial variable: \[p_n(\mathrm d x)=\mu_n(\mathrm d s)K_n(s,\mathrm d\omega), \qquad s=\|x-\boldsymbol{\theta}\|,\] where \(K_n(s,\cdot)\) is a regular conditional angular distribution. Define \[\eta_n(\mathrm d x):=\nu_n(\mathrm d s)K_n(s,\mathrm d\omega).\] Then \(\eta_n\ll p_n\), and, for \(0<r<\rho_n\), \[\eta_n(B_r(\boldsymbol{\theta})) =\nu_n([0,r)) =F_n(r)^t =\left(\frac{p(B_r(\boldsymbol{\theta}))}{a_n}\right)^t .\]

By Lemma 3, \[D_{\mathrm{KL}}(\nu_n\|\mu_n) \le C_t:=\log t+\frac{1}{t}-1<\infty .\] Since \(\eta_n\) and \(p_n\) have the same conditional angular distributions, the Radon–Nikodym derivative is \[\frac{\mathrm d\eta_n}{\mathrm d p_n}(x) = \frac{\mathrm d\nu_n}{\mathrm d\mu_n}(\|x-\boldsymbol{\theta}\|) \quad p_n\text{-a.e.}\] Therefore \[\begin{align} D_{\mathrm{KL}}(\eta_n\|p_n) &=\int \log\!\left(\frac{\mathrm d\eta_n}{\mathrm d p_n}\right) \,\mathrm d\eta_n \\ &=\int \log\!\left(\frac{\mathrm d\nu_n}{\mathrm d\mu_n}\right) \,\mathrm d\nu_n =D_{\mathrm{KL}}(\nu_n\|\mu_n) \le C_t . \end{align}\] This bound is uniform in \(n\).

Now define \[q_n(A):=p(A\cap E_n^c)+a_n\eta_n(A\cap E_n).\] Then \(q_n\) is a probability measure, \(q_n\ll p\), and \(q_n=p\) on \(E_n^c\). For \(0<r<\rho_n\), \[q_n(B_r(\boldsymbol{\theta})) =a_n\eta_n(B_r(\boldsymbol{\theta})) =a_n^{1-t}p(B_r(\boldsymbol{\theta}))^t .\] The positive constant \(a_n^{1-t}\) does not affect the local \(O\)- or \(\Omega\)-classes as \(r\downarrow0\). Therefore, for every \(\eta>0\), \[q_n(B_r(\boldsymbol{\theta}))=O(r^{1/\eta}) \quad\Longleftrightarrow\quad p(B_r(\boldsymbol{\theta}))=O(r^{1/(t\eta)}),\] and similarly \[q_n(B_r(\boldsymbol{\theta}))=\Omega(r^{1/\eta}) \quad\Longleftrightarrow\quad p(B_r(\boldsymbol{\theta}))=\Omega(r^{1/(t\eta)}).\] Thus \[\inf\mathcal{U}_{\mathrm{pow}}^{q_n} =\frac{1}{t}\inf\mathcal{U}_{\mathrm{pow}}^p =\frac{1}{t}\frac{m}{d} =\frac{k}{d},\] and \[\sup\mathcal{L}_{\mathrm{pow}}^{q_n} =\frac{1}{t}\sup\mathcal{L}_{\mathrm{pow}}^p =\frac{1}{t}\frac{m}{d} =\frac{k}{d}.\] Consequently, \[\mathrm{MI}_{\mathrm{pow}}(q_n,\boldsymbol{\theta})=k\] for every \(n\).

Finally, since \(q_n=p\) outside \(E_n\) and the conditional law inside \(E_n\) is changed from \(p_n\) to \(\eta_n\), \[D_{\mathrm{KL}}(q_n\|p) =a_nD_{\mathrm{KL}}(\eta_n\|p_n) \le a_nC_t.\] As \(a_n=p(B_{\rho_n}(\boldsymbol{\theta}))\to0\), we obtain \[D_{\mathrm{KL}}(q_n\|p)\to0 .\] Thus \(q_n\in\mathcal{D}_{\boldsymbol{\theta}}\), \(D_{\mathrm{KL}}(q_n\|p)\to0\), and \(\mathrm{MI}_{\mathrm{pow}}(q_n,\boldsymbol{\theta})=k\) for every \(n\). This proves the theorem. ◻

Proof of Theorem 3↩︎

Proof. Write \[m_0(r):=\pi_0(B_r(\boldsymbol{\theta})), \qquad \ell(r):=\log\frac{1}{r} .\] First assume that \[0<\mathrm{MI}_{\mathrm{pow}}(\pi_0,\boldsymbol{\theta})<\infty, \qquad a:=d\,\mathrm{MI}_{\mathrm{pow}}^{-1}(\pi_0,\boldsymbol{\theta}).\] The condition in the theorem is exactly \[a+\gamma>0 .\] Let \[I(r):= \int_{B_r(\boldsymbol{\theta})} \|x-\boldsymbol{\theta}\|^\gamma \left(\log\frac{1}{\|x-\boldsymbol{\theta}\|}\right)^\beta \pi_0(\mathrm d x).\] Since \(0<\mathrm{MI}_{\mathrm{pow}}(\pi_0,\boldsymbol{\theta})<\infty\), the prior has no atom at \(\boldsymbol{\theta}\), so the value of the integrand at \(\boldsymbol{\theta}\) is irrelevant. The local likelihood assumption gives constants \(0<c<C<\infty\) and \(r_0>0\) such that, for \(x\in B_{r_0}(\boldsymbol{\theta})\setminus\{\boldsymbol{\theta}\}\), \[c\|x-\boldsymbol{\theta}\|^\gamma \left(\log\frac{1}{\|x-\boldsymbol{\theta}\|}\right)^\beta \le L(x\mid\mathcal{D}) \le C\|x-\boldsymbol{\theta}\|^\gamma \left(\log\frac{1}{\|x-\boldsymbol{\theta}\|}\right)^\beta .\] By Bayes’ formula, \[\pi_{\mathcal{D}}(B_r(\boldsymbol{\theta})) =Z_{\mathcal{D}}^{-1} \int_{B_r(\boldsymbol{\theta})}L(x\mid\mathcal{D})\,\pi_0(\mathrm d x).\] Since \(Z_{\mathcal{D}}\in(0,\infty)\), the normalising constant is fixed, and therefore \[\pi_{\mathcal{D}}(B_r(\boldsymbol{\theta}))\asymp I(r), \qquad r\downarrow0 .\] It remains to identify the power and logarithmic orders of \(I(r)\).

For the power order, the definition of the Power Mass Index gives, for every sufficiently small \(u>0\), \[m_0(r)=O(r^{a-u}), \qquad m_0(r)=\Omega(r^{a+u}).\] We claim that, for every \(\varepsilon>0\), \[I(r)=O(r^{a+\gamma-\varepsilon}), \qquad I(r)=\Omega(r^{a+\gamma+\varepsilon}).\] For the upper bound, choose \(u,v>0\) such that \[u+v<\varepsilon, \qquad a+\gamma-u-v>0 .\] Since logarithmic factors are slower than powers, \[\left(\log\frac{1}{t}\right)^\beta\le C_vt^{-v}\] for all sufficiently small \(t>0\), with \(C_v\) independent of the annulus index used below. Let \(r_j=e^{-j}r\), and set \[A_j:=B_{r_j}(\boldsymbol{\theta})\setminus B_{r_{j+1}}(\boldsymbol{\theta}), \qquad j\ge0 .\] Then \(B_r(\boldsymbol{\theta})\setminus\{\boldsymbol{\theta}\}=\bigcup_{j\ge0}A_j\). On \(A_j\), \(r_{j+1}\le \|x-\boldsymbol{\theta}\|<r_j\). Hence, after absorbing the fixed factor \(e^{|\gamma-v|}\) when \(\gamma-v<0\), \[\|x-\boldsymbol{\theta}\|^\gamma \left(\log\frac{1}{\|x-\boldsymbol{\theta}\|}\right)^\beta \le C r_j^{\gamma-v},\] where the constant is uniform in \(j\) and in all sufficiently small \(r\). Therefore \[\begin{align} I(r) &\le C\sum_{j=0}^{\infty} r_j^{\gamma-v}\pi_0(A_j) \\ &\le C\sum_{j=0}^{\infty} r_j^{\gamma-v}m_0(r_j) \\ &\le C\sum_{j=0}^{\infty} r_j^{a+\gamma-u-v} \le C r^{a+\gamma-u-v}. \end{align}\] This proves \(I(r)=O(r^{a+\gamma-\varepsilon})\).

For the lower bound, choose \(u,v>0\) and \(s>1\), with \(s\) sufficiently close to one, such that \[s(a-u)>a+u, \qquad u+(s-1)\max\{\gamma,0\}+v<\varepsilon .\] Let \[A(r):=B_r(\boldsymbol{\theta})\setminus B_{r^s}(\boldsymbol{\theta}).\] The prior power bounds imply \[\pi_0(A(r)) =m_0(r)-m_0(r^s) \ge c r^{a+u}-C r^{s(a-u)} \ge c r^{a+u}\] for all sufficiently small \(r\). On \(A(r)\), \[\|x-\boldsymbol{\theta}\|^\gamma \ge C r^{\gamma+(s-1)\max\{\gamma,0\}}, \qquad \left(\log\frac{1}{\|x-\boldsymbol{\theta}\|}\right)^\beta \ge C r^v .\] Consequently, \[I(r) \ge C r^{a+\gamma+u+(s-1)\max\{\gamma,0\}+v} .\] Because \[u+(s-1)\max\{\gamma,0\}+v<\varepsilon,\] we have \[r^{a+\gamma+u+(s-1)\max\{\gamma,0\}+v} =\Omega(r^{a+\gamma+\varepsilon}),\] and hence \(I(r)=\Omega(r^{a+\gamma+\varepsilon})\). The preceding two estimates satisfy Lemma 2(i) with exponent \(a+\gamma\). Therefore \[d\,\mathrm{MI}_{\mathrm{pow}}^{-1}(\pi_{\mathcal{D}},\boldsymbol{\theta}) =a+\gamma =d\,\mathrm{MI}_{\mathrm{pow}}^{-1}(\pi_0,\boldsymbol{\theta})+\gamma .\] Equivalently, \[\mathrm{MI}_{\mathrm{pow}}^{-1}(\pi_{\mathcal{D}},\boldsymbol{\theta}) =\mathrm{MI}_{\mathrm{pow}}^{-1}(\pi_0,\boldsymbol{\theta})+d^{-1}\gamma .\]

Now assume that \(\mathrm{MI}_{\mathrm{log}}(\pi_0,\boldsymbol{\theta})\) exists, and set \[b:=\mathrm{MI}_{\mathrm{log}}(\pi_0,\boldsymbol{\theta}), \qquad \lambda:=a+\gamma>0 .\] Then, for every sufficiently small \(u>0\), \[m_0(r)=O\!\left(r^a\ell(r)^{b+u}\right), \qquad m_0(r)=\Omega\!\left(r^a\ell(r)^{b-u}\right).\] We show that, for every \(\varepsilon>0\), \[I(r)=O\!\left(r^\lambda\ell(r)^{b+\beta+\varepsilon}\right), \qquad I(r)=\Omega\!\left(r^\lambda\ell(r)^{b+\beta-\varepsilon}\right).\] For the upper bound, use the same annuli \(A_j\), and choose \(u\in(0,\varepsilon)\). On \(A_j\), the relations \(r_{j+1}\le \|x-\boldsymbol{\theta}\|<r_j\) and \(\log(1/\|x-\boldsymbol{\theta}\|)\asymp\ell(r_j)\) give \[\|x-\boldsymbol{\theta}\|^\gamma \left(\log\frac{1}{\|x-\boldsymbol{\theta}\|}\right)^\beta \le C r_j^\gamma\ell(r_j)^\beta ,\] with a constant uniform in \(j\) and in all sufficiently small \(r\). Hence \[\begin{align} I(r) &\le C\sum_{j=0}^{\infty}r_j^\gamma\ell(r_j)^\beta m_0(r_j)\\ &\le C\sum_{j=0}^{\infty}r_j^{a+\gamma}\ell(r_j)^{b+\beta+u}\\ &=C r^\lambda\sum_{j=0}^{\infty}e^{-j\lambda}(\ell(r)+j)^{b+\beta+u}. \end{align}\] For \(c=b+\beta+u\), the elementary geometric estimate \[\sum_{j=0}^{\infty}e^{-j\lambda}(\ell(r)+j)^c \le C\ell(r)^c\] holds uniformly for all sufficiently small \(r\). Indeed, if \(c<0\), then \((\ell(r)+j)^c\le \ell(r)^c\); if \(c\ge0\), then \((\ell(r)+j)^c\le \ell(r)^c(1+j)^c\), and \(\sum_{j\ge0}e^{-j\lambda}(1+j)^c<\infty\). Therefore \[I(r)\le C r^\lambda\ell(r)^{b+\beta+u}.\] This gives the desired upper logarithmic bound.

For the lower bound, write \(\gamma_+:=\max\{\gamma,0\}\). Choose \(u>0\) and \(K>0\) such that \[Ka>2u, \qquad u+K\gamma_+<\varepsilon .\] Set \[\rho(r):=r\ell(r)^{-K}, \qquad A(r):=B_r(\boldsymbol{\theta})\setminus B_{\rho(r)}(\boldsymbol{\theta}).\] Using the lower and upper logarithmic bounds for \(m_0\), \[\begin{align} \pi_0(A(r)) &=m_0(r)-m_0(\rho(r))\\ &\ge c r^a\ell(r)^{b-u} -C r^a\ell(r)^{b+u-Ka} \ge c r^a\ell(r)^{b-u} \end{align}\] for all sufficiently small \(r\). On \(A(r)\), \[\|x-\boldsymbol{\theta}\|^\gamma\ge C r^\gamma\ell(r)^{-K\gamma_+}, \qquad \log\frac{1}{\|x-\boldsymbol{\theta}\|}\asymp\ell(r).\] Therefore \[I(r) \ge C r^{a+\gamma}\ell(r)^{b+\beta-u-K\gamma_+} =\Omega\!\left(r^\lambda\ell(r)^{b+\beta-\varepsilon}\right).\] The upper and lower logarithmic bounds satisfy Lemma 2(ii) at the already identified power scale \(\lambda\). Hence \[\mathrm{MI}_{\mathrm{log}}(\pi_{\mathcal{D}},\boldsymbol{\theta}) =b+\beta =\mathrm{MI}_{\mathrm{log}}(\pi_0,\boldsymbol{\theta})+\beta .\]

It remains to treat the case \(\mathrm{MI}_{\mathrm{pow}}(\pi_0,\boldsymbol{\theta})=0\). In this case, for every \(M>0\), \[m_0(r)=O(r^M), \qquad r\downarrow0 .\] Fix any \(N>0\). Choose \(v>0\), and then choose \(M\) so large that \[M+\gamma-v>N, \qquad M+\gamma-v>0 .\] Using the same annular decomposition as above and the bound \(m_0(r_j)=O(r_j^M)\), we obtain \[\begin{align} I(r) &\le C\sum_{j=0}^{\infty} r_j^{\gamma-v}m_0(r_j) \\ &\le C\sum_{j=0}^{\infty} r_j^{M+\gamma-v} \le C r^{M+\gamma-v} =O(r^N). \end{align}\] Thus \(\pi_{\mathcal{D}}(B_r(\boldsymbol{\theta}))=O(r^N)\) for every \(N>0\). Consequently the posterior small-ball mass is faster than every polynomial power. Hence every positive power is an upper gauge and no positive power is a lower gauge, which gives \[\mathrm{MI}_{\mathrm{pow}}(\pi_{\mathcal{D}},\boldsymbol{\theta})=0 .\] The proof is complete. ◻

Proof of Theorem 4↩︎

Proof. Set \[m_0(r):=\pi_0(B_r(\boldsymbol{\theta})), \qquad m_{\mathcal{D}}(r):=\pi_{\mathcal{D}}(B_r(\boldsymbol{\theta})), \qquad \ell(r):=-\log r .\] By the theorem hypothesis, \(0<\mathrm{MI}_{\mathrm{pow}}(\pi_0,\boldsymbol{\theta})<\infty\). Set \[a:=d\,\mathrm{MI}_{\mathrm{pow}}^{-1}(\pi_0,\boldsymbol{\theta}).\] Since \(g(x)\asymp1\) as \(x\to\boldsymbol{\theta}\), there exist constants \(0<c_g<C_g<\infty\) and \(r_0>0\) such that \[c_g\le g(x)\le C_g, \qquad x\in B_{r_0}(\boldsymbol{\theta}).\] By Bayes’ formula and the finite marginal assumption, the normalising constant only changes small-ball masses by a fixed multiplicative constant. Hence, for \(0<r<r_0\), \[m_{\mathcal{D}}(r) \asymp \pi_0(B_r(\boldsymbol{\theta})\cap E) =A_{E,\boldsymbol{\theta}}(r)m_0(r).\] Using the assumed acceptance-ratio scaling, \[\label{eq:acceptance-proof-comparison-new} m_{\mathcal{D}}(r) \asymp r^\delta\ell(r)^\lambda m_0(r), \qquad r\downarrow0 .\tag{2}\] Because \(A_{E,\boldsymbol{\theta}}(r)\in[0,1]\), the regime stated in the theorem is the non-degenerate regime compatible with this two-sided comparison.

For every \(\varepsilon>0\), the prior Power Mass Index gives \[m_0(r)=O(r^{a-\varepsilon}), \qquad m_0(r)=\Omega(r^{a+\varepsilon}).\] Since logarithmic factors are sub-polynomial, for every \(\eta>0\), \[\ell(r)^\lambda=O(r^{-\eta}), \qquad \ell(r)^\lambda=\Omega(r^{\eta}) .\] Given \(\varepsilon>0\), apply the prior bounds with \(\varepsilon/2\) and the last display with \(\eta=\varepsilon/2\). Then 2 gives \[m_{\mathcal{D}}(r)=O(r^{a+\delta-\varepsilon}), \qquad m_{\mathcal{D}}(r)=\Omega(r^{a+\delta+\varepsilon}).\] By Lemma 2(i), the posterior power exponent is \(a+\delta\), and \[d\,\mathrm{MI}_{\mathrm{pow}}^{-1}(\pi_{\mathcal{D}},\boldsymbol{\theta}) =a+\delta =d\,\mathrm{MI}_{\mathrm{pow}}^{-1}(\pi_0,\boldsymbol{\theta})+\delta .\] Equivalently, \[\mathrm{MI}_{\mathrm{pow}}^{-1}(\pi_{\mathcal{D}},\boldsymbol{\theta}) =\mathrm{MI}_{\mathrm{pow}}^{-1}(\pi_0,\boldsymbol{\theta})+d^{-1}\delta .\]

Now assume that \(\mathrm{MI}_{\mathrm{log}}(\pi_0,\boldsymbol{\theta})\) exists and set \[b:=\mathrm{MI}_{\mathrm{log}}(\pi_0,\boldsymbol{\theta}).\] Then, for every \(\varepsilon>0\), \[m_0(r)=O\!\left(r^a\ell(r)^{b+\varepsilon}\right), \qquad m_0(r)=\Omega\!\left(r^a\ell(r)^{b-\varepsilon}\right).\] Combining these bounds with 2 gives \[m_{\mathcal{D}}(r) =O\!\left(r^{a+\delta}\ell(r)^{b+\lambda+\varepsilon}\right), \qquad m_{\mathcal{D}}(r) =\Omega\!\left(r^{a+\delta}\ell(r)^{b+\lambda-\varepsilon}\right).\] The posterior power exponent has already been identified as \(a+\delta\). Hence Lemma 2(ii) gives \[\mathrm{MI}_{\mathrm{log}}(\pi_{\mathcal{D}},\boldsymbol{\theta}) =b+\lambda =\mathrm{MI}_{\mathrm{log}}(\pi_0,\boldsymbol{\theta})+\lambda .\] The proof is complete. ◻

Proof of Corollary 1↩︎

Proof. For each \(\tau>0\), define \[Z_\tau:=\int s_\tau(x)\,\pi_0(\mathrm d x), \qquad Z_E:=\int \mathbf{1}_E(x)\,\pi_0(\mathrm d x).\] For each \(\tau>0\), the softened normalising constant satisfies \(0<Z_\tau\le C\). By Assumption 1 applied to the hard likelihood \(L=\mathbf{1}_E\), we have \(Z_E\in(0,\infty)\). By uniform boundedness, pointwise convergence \(s_\tau\to\mathbf{1}_E\) \(\pi_0\)-a.e., and dominated convergence, \[Z_\tau\to Z_E .\] For any measurable set \(A\), \[\pi_\tau(A)=Z_\tau^{-1}\int_A s_\tau(x)\,\pi_0(\mathrm d x), \qquad \pi_{\mathcal{D}}(A)=Z_E^{-1}\int_A \mathbf{1}_E(x)\,\pi_0(\mathrm d x).\] Therefore \[\|\pi_\tau-\pi_{\mathcal{D}}\|_{\mathrm{TV}} \le |Z_\tau^{-1}-Z_E^{-1}|Z_\tau +Z_E^{-1}\int |s_\tau-\mathbf{1}_E|\,\mathrm d\pi_0,\] and both terms converge to zero. Hence \[\|\pi_\tau-\pi_{\mathcal{D}}\|_{\mathrm{TV}}\to0 .\]

Now fix \(\tau>0\). Since \(s_\tau\) is positive and continuous, there are constants \(0<c_\tau<C_\tau<\infty\) and \(r_\tau>0\) such that \[c_\tau\le s_\tau(x)\le C_\tau, \qquad x\in B_{r_\tau}(\boldsymbol{\theta}).\] Thus \[\pi_\tau(B_r(\boldsymbol{\theta}))\asymp\pi_0(B_r(\boldsymbol{\theta})), \qquad r\downarrow0,\] for each fixed \(\tau>0\). Consequently \[\mathrm{MI}_{\mathrm{pow}}(\pi_\tau,\boldsymbol{\theta}) =\mathrm{MI}_{\mathrm{pow}}(\pi_0,\boldsymbol{\theta}), \qquad \tau>0 .\] If \(A_{E,\boldsymbol{\theta}}(r)\asymp r^\delta\ell(r)^\lambda\) with \(\delta>0\), then Theorem 4 applied to the hard posterior gives \[d\,\mathrm{MI}_{\mathrm{pow}}^{-1}(\pi_{\mathcal{D}},\boldsymbol{\theta}) =d\,\mathrm{MI}_{\mathrm{pow}}^{-1}(\pi_0,\boldsymbol{\theta})+\delta .\] This differs from the softened exponent \(d\,\mathrm{MI}_{\mathrm{pow}}^{-1}(\pi_0,\boldsymbol{\theta})\). Hence \[\mathrm{MI}_{\mathrm{pow}}(\pi_{\mathcal{D}},\boldsymbol{\theta}) \neq \lim_{\tau\downarrow0}\mathrm{MI}_{\mathrm{pow}}(\pi_\tau,\boldsymbol{\theta}).\]

For the logarithmic statement, assume \(\delta=0\), \(\lambda<0\), and that \(\mathrm{MI}_{\mathrm{log}}(\pi_0,\boldsymbol{\theta})\) exists. For each fixed \(\tau>0\), the local positivity of \(s_\tau\) gives \[\mathrm{MI}_{\mathrm{log}}(\pi_\tau,\boldsymbol{\theta}) =\mathrm{MI}_{\mathrm{log}}(\pi_0,\boldsymbol{\theta}).\] For the hard posterior, Theorem 4 gives \[\mathrm{MI}_{\mathrm{log}}(\pi_{\mathcal{D}},\boldsymbol{\theta}) =\mathrm{MI}_{\mathrm{log}}(\pi_0,\boldsymbol{\theta})+\lambda .\] Since \(\lambda<0\), the two logarithmic indices differ, and therefore \[\mathrm{MI}_{\mathrm{log}}(\pi_{\mathcal{D}},\boldsymbol{\theta}) \neq \lim_{\tau\downarrow0}\mathrm{MI}_{\mathrm{log}}(\pi_\tau,\boldsymbol{\theta}).\] The proof is complete. ◻

Proof of Theorem 5↩︎

Proof. Let \(q=q_{\ll p}+q_{\perp p}\), and write \[R:=\frac{\mathrm d q_{\ll p}}{\mathrm d p}.\] For the restrictions of \(q\) and \(p\) to \(E\), define the local Hellinger-type quantity \[H_E^2(q,p) :=p(E)+q(E)-2\int_E\sqrt R\,\mathrm d p.\] Using the Lebesgue decomposition, this can be written as \[H_E^2(q,p) =\int_E(\sqrt R-1)^2\,\mathrm d p+q_{\perp p}(E).\] We first spell out the finite-measure Cauchy–Schwarz step. Since \[\int_E\sqrt R\,\mathrm d p \le \sqrt{q_{\ll p}(E)}\sqrt{p(E)} \le \sqrt{q(E)p(E)},\] we have \[\begin{align} H_E^2(q,p) &=p(E)+q(E)-2\int_E\sqrt R\,\mathrm d p \\ &\ge p(E)+q(E)-2\sqrt{p(E)q(E)} =\left(\sqrt{q(E)}-\sqrt{p(E)}\right)^2 . \end{align}\] Hence \[\left|\sqrt{q(E)}-\sqrt{p(E)}\right| \le H_E(q,p),\] and therefore \[|q(E)-p(E)| \le \bigl(\sqrt{q(E)}+\sqrt{p(E)}\bigr)H_E(q,p) \le 2H_E(q,p),\] because \(p(E),q(E)\le1\).

It remains to compare \(H_E^2(q,p)\) with \(D_\alpha^E(q\|p)\). We prove the scalar inequality \[(\sqrt x-1)^2\le C_\alpha f_\alpha(x), \qquad x\ge0, \qquad C_\alpha:=\max\left\{1,\frac{1-\alpha}{\alpha}\right\}.\] Set \(t=\sqrt x\). If \(0<\alpha\le1/2\), then \(C_\alpha=(1-\alpha)/\alpha\), and the desired inequality is equivalent to \[2\alpha t-t^{2\alpha}-2\alpha+1\ge0 .\] The derivative of the left-hand side is \(2\alpha(1-t^{2\alpha-1})\), which is negative on \((0,1)\) and positive on \((1,\infty)\). Since the value at \(t=1\) is zero, the inequality follows. If \(1/2\le\alpha<1\), then \(C_\alpha=1\), and the desired inequality is equivalent to \[(2\alpha-1)t^2+2(1-\alpha)t\ge t^{2\alpha}.\] This follows from weighted AM–GM, because the weights \(2\alpha-1\) and \(2(1-\alpha)\) are non-negative and sum to one. The endpoint \(x=0\) is covered by continuity.

Consequently, \[\int_E(\sqrt R-1)^2\,\mathrm d p \le C_\alpha\int_E f_\alpha(R)\,\mathrm d p .\] Moreover, since \(C_\alpha\ge (1-\alpha)/\alpha\), \[q_{\perp p}(E) \le C_\alpha\frac{\alpha}{1-\alpha}q_{\perp p}(E).\] Therefore \[H_E^2(q,p) \le C_\alpha\left\{ \int_E f_\alpha(R)\,\mathrm d p +\frac{\alpha}{1-\alpha}q_{\perp p}(E) \right\} =C_\alpha D_\alpha^E(q\|p).\] Combining the two displays yields \[|q(E)-p(E)| \le 2\sqrt{C_\alpha D_\alpha^E(q\|p)}.\] This proves the theorem. ◻

Proof of Corollary 2↩︎

Proof. Apply Theorem 5 with \(E=B_r(\boldsymbol{\theta})\). If \[D_\alpha^{B_r(\boldsymbol{\theta})}(q\|p)=O(r^\delta),\] then \[|q(B_r(\boldsymbol{\theta}))-p(B_r(\boldsymbol{\theta}))| \le 2\sqrt{C_\alpha D_\alpha^{B_r(\boldsymbol{\theta})}(q\|p)} =O(r^{\delta/2}).\] ◻

Proof of Theorem 6↩︎

Proof. Assume \(p(E)>0\). Let \(q=q_{\ll p}+q_{\perp p}\), and write \[R:=\frac{\mathrm d q_{\ll p}}{\mathrm d p}, \qquad x:=\frac{1}{p(E)}\int_E R\,\mathrm d p, \qquad h:=\frac{q_{\perp p}(E)}{p(E)}.\] Then \[\frac{q(E)}{p(E)}=x+h.\] Since \[f_\alpha'(u)=\frac{\alpha}{1-\alpha}(1-u^{\alpha-1}) \le \frac{\alpha}{1-\alpha},\] we have \[f_\alpha(x+h) \le f_\alpha(x)+\frac{\alpha}{1-\alpha}h .\] By Jensen’s inequality, \[f_\alpha(x) \le \frac{1}{p(E)}\int_E f_\alpha(R)\,\mathrm d p.\] Combining these bounds gives \[f_\alpha\!\left(\frac{q(E)}{p(E)}\right) \le \frac{D_\alpha^E(q\|p)}{p(E)} =\overline{D}_\alpha^E(q\|p).\] This proves the first claim.

For the relative error bound, we derive the scalar inequality \[f_\alpha(y) \ge \frac{\alpha |y-1|^2}{2(1+|y-1|)}, \qquad y\ge0 .\] Since \(f_\alpha(1)=f_\alpha'(1)=0\) and \(f_\alpha''(u)=\alpha u^{\alpha-2}\), Taylor’s formula with integral remainder gives, for \(0\le y\le1\), \[f_\alpha(y) =\alpha\int_y^1 (u-y)u^{\alpha-2}\,\mathrm d u \ge \frac{\alpha}{2}(1-y)^2 \ge \frac{\alpha(1-y)^2}{2(1+1-y)}.\] For \(y=1+s\ge1\), the same formula gives \[f_\alpha(1+s) =\alpha\int_0^s (s-u)(1+u)^{\alpha-2}\,\mathrm d u .\] Because \(\alpha>0\), \((1+u)^{\alpha-2}\ge(1+u)^{-2}\). Hence \[f_\alpha(1+s) \ge \alpha\int_0^s \frac{s-u}{(1+u)^2}\,\mathrm d u =\alpha\{s-\log(1+s)\}.\] The elementary bound \[s-\log(1+s)\ge \frac{s^2}{2(1+s)}, \qquad s\ge0,\] follows by differentiating the difference between the two sides. Therefore the same scalar inequality holds for \(y\ge1\).

Set \[y:=\frac{q(E)}{p(E)}, \qquad s:=|1-y|, \qquad z:=\overline{D}_\alpha^E(q\|p).\] The first part gives \(f_\alpha(y)\le z\), so \[\frac{\alpha s^2}{2(1+s)}\le z .\] Solving this quadratic inequality for \(s\ge0\) yields \[s \le \frac{z}{\alpha} + \sqrt{\left(\frac{z}{\alpha}\right)^2+\frac{2z}{\alpha}} .\] The last display is bounded by \[\frac{2}{\alpha}\sqrt{z}\sqrt{z+2}.\] Therefore \[\left|1-\frac{q(E)}{p(E)}\right| \le \frac{2}{\alpha} \sqrt{\overline{D}_\alpha^E(q\|p)} \sqrt{\overline{D}_\alpha^E(q\|p)+2}.\] This proves the theorem. ◻

Proof of Corollary 3↩︎

Proof. Set \[t_r:=\overline{D}_\alpha^{B_r(\boldsymbol{\theta})}(q\|p).\] By assumption, \[t_r=O(r^\delta), \qquad r\downarrow0,\] and hence \(t_r\to0\). Applying Theorem 6 with \(E=B_r(\boldsymbol{\theta})\) gives \[\left|1-\frac{q(B_r(\boldsymbol{\theta}))}{p(B_r(\boldsymbol{\theta}))}\right| \le \frac{2}{\alpha}\sqrt{t_r}\sqrt{t_r+2} =O(t_r^{1/2}) =O(r^{\delta/2}).\] ◻

Proof of Theorem 7↩︎

Proof. Let \[E_r:=B_r(\boldsymbol{\theta}), \qquad m_p(r):=p(E_r), \qquad m_q(r):=q(E_r).\] By assumption, \(m_p(r)>0\) and \(m_q(r)>0\) for all sufficiently small \(r\). We use two elementary facts about \(f_\alpha\). First, \(f_\alpha\) is continuous on \([0,\infty)\), satisfies \(f_\alpha(1)=0\), \(f_\alpha(0)=1\), and diverges as \(x\to\infty\). Hence, for every \(\lambda<1\), the sublevel set \[\{x\ge0:f_\alpha(x)\le\lambda\}\] is contained in an interval \([c,C]\) with \(0<c<C<\infty\). Second, for every finite \(M\), the sublevel set \[\{x\ge0:f_\alpha(x)\le M\}\] is bounded above.

We first prove (i). Suppose \[\limsup_{r\downarrow0}\overline{D}_\alpha^{E_r}(q\|p)<1 .\] Choose \(\lambda<1\) such that \[\overline{D}_\alpha^{E_r}(q\|p)\le\lambda\] for all sufficiently small \(r\). By Theorem 6, \[f_\alpha\!\left(\frac{m_q(r)}{m_p(r)}\right) \le \overline{D}_\alpha^{E_r}(q\|p) \le\lambda .\] The first scalar fact gives constants \(0<c<C<\infty\) such that \[c\le\frac{m_q(r)}{m_p(r)}\le C\] for all sufficiently small \(r\). Equivalently, \[m_q(r)\asymp m_p(r), \qquad r\downarrow0 .\] Thus the upper and lower power gauge classes of \(p\) and \(q\) coincide. For any small-ball function, \(\sup\mathcal{L}_{\mathrm{pow}}\le \inf\mathcal{U}_{\mathrm{pow}}\). Since Assumption 2 gives equality of these two quantities for the relevant measures, the common gauge classes identify \[\mathrm{MI}_{\mathrm{pow}}(q,\boldsymbol{\theta})=\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta}).\]

If \(\mathrm{MI}_{\mathrm{log}}(p,\boldsymbol{\theta})\) exists, then the common power exponent is \[a:=d\,\mathrm{MI}_{\mathrm{pow}}^{-1}(p,\boldsymbol{\theta}) =d\,\mathrm{MI}_{\mathrm{pow}}^{-1}(q,\boldsymbol{\theta}).\] Since \(m_q(r)\asymp m_p(r)\), for every \(\eta\in\mathbb{R}\), \[m_p(r)=O\!\left(r^a(-\log r)^\eta\right) \quad\Longleftrightarrow\quad m_q(r)=O\!\left(r^a(-\log r)^\eta\right),\] and similarly \[m_p(r)=\Omega\!\left(r^a(-\log r)^\eta\right) \quad\Longleftrightarrow\quad m_q(r)=\Omega\!\left(r^a(-\log r)^\eta\right).\] Hence the logarithmic upper and lower gauge classes also coincide. Since the Logarithmic Mass Index of \(p\) exists, and the common gauge classes force the same upper and lower logarithmic thresholds for \(q\), \[\mathrm{MI}_{\mathrm{log}}(q,\boldsymbol{\theta})=\mathrm{MI}_{\mathrm{log}}(p,\boldsymbol{\theta}).\]

We now prove (ii). Suppose \[\limsup_{r\downarrow0}\overline{D}_\alpha^{E_r}(p\|q)<\infty .\] Then there exists \(M<\infty\) such that \[\overline{D}_\alpha^{E_r}(p\|q)\le M\] for all sufficiently small \(r\). Applying Theorem 6 with \(p\) and \(q\) interchanged gives \[f_\alpha\!\left(\frac{m_p(r)}{m_q(r)}\right) \le \overline{D}_\alpha^{E_r}(p\|q) \le M .\] By the second scalar fact, \(m_p(r)/m_q(r)\) is bounded above. Therefore there exists \(C<\infty\) such that \[m_q(r)\ge C^{-1}m_p(r)\] for all sufficiently small \(r\).

Consequently, every lower power bound for \(m_p\) is also a lower power bound for \(m_q\): \[\mathcal{L}_{\mathrm{pow}}(p,\boldsymbol{\theta}) \subseteq \mathcal{L}_{\mathrm{pow}}(q,\boldsymbol{\theta}).\] Taking suprema and using the dimensional normalisation in Definition 3 gives \[\sup\mathcal{L}_{\mathrm{pow}}(q,\boldsymbol{\theta}) \ge \sup\mathcal{L}_{\mathrm{pow}}(p,\boldsymbol{\theta}).\] Assumption 2 identifies these lower thresholds with the corresponding Power Mass Indices, because \(\sup\mathcal{L}_{\mathrm{pow}}=\inf\mathcal{U}_{\mathrm{pow}}\) for both measures. Therefore \[\mathrm{MI}_{\mathrm{pow}}(q,\boldsymbol{\theta}) \ge \mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta}).\]

Finally, assume \[\mathrm{MI}_{\mathrm{pow}}(q,\boldsymbol{\theta})=\mathrm{MI}_{\mathrm{pow}}(p,\boldsymbol{\theta})\] and that both Logarithmic Mass Indices exist. Let \[a:=d\,\mathrm{MI}_{\mathrm{pow}}^{-1}(p,\boldsymbol{\theta}) =d\,\mathrm{MI}_{\mathrm{pow}}^{-1}(q,\boldsymbol{\theta}).\] The inequality \(m_q(r)\ge C^{-1}m_p(r)\) implies that every lower logarithmic bound for \(m_p\) at the common power scale is also a lower logarithmic bound for \(m_q\). Hence \[\mathcal{L}_{\mathrm{log}}(p,\boldsymbol{\theta}) \subseteq \mathcal{L}_{\mathrm{log}}(q,\boldsymbol{\theta}).\] Taking suprema and using the existence of both Logarithmic Mass Indices gives \[\mathrm{MI}_{\mathrm{log}}(q,\boldsymbol{\theta}) \ge \mathrm{MI}_{\mathrm{log}}(p,\boldsymbol{\theta}).\] The proof is complete. ◻

10 Additional experimental details↩︎

This appendix gives the concrete settings used for the experiments in Section 6. All experiments are finite-radius illustrations. We do not estimate asymptotic Mass Indices directly.

10.1 Experiment 1: synthetic calibration↩︎

Experiment 1 plots closed-form small-ball curves over \[r\in[10^{-3},0.6].\] The purpose is to visualize representative local regimes rather than to fit a statistical model.

For the regular reference case, we use the one-dimensional Gaussian small-ball mass \[p_{\mathrm{Gauss}}(B_r(0)) = \operatorname{erf}\left(\frac{r}{\sqrt{2}}\right).\] This behaves as \(r\) for small \(r\).

For the local-depletion example, we plot \[p_{\mathrm{dep}}(B_r(0))=r^2.\] This has faster small-radius decay than the regular reference.

For the cusp example, we plot \[p_{\mathrm{cusp}}(B_r(0))=\sqrt r.\] This has slower small-radius decay than the regular reference.

For the logarithmically corrected example, we plot \[p_{\log}(B_r(0))=r(1-\log r).\] This has the same leading power order as the regular reference, but differs by a logarithmic factor.

For the spike-slab example, we use spike weight \(0.25\) and slab scale \(0.35\), giving \[p_{\mathrm{spike}}(B_r(0)) = 0.25 + 0.75\, \operatorname{erf}\left( \frac{r}{0.35\sqrt{2}} \right).\] The nonzero spike produces a small-ball mass that does not vanish as \(r\to0\).

10.2 Experiment 2: UCI Bayesian logistic-regression sanity check↩︎

Experiment 2 uses three binary classification datasets: Breast Cancer, Iris \(0\) vs \(1\), and Wine \(0\) vs \(1\). For Iris and Wine, only classes \(0\) and \(1\) are retained. Each dataset is repeated over five random seeds, \[\{0,1,2,3,4\}.\] The seed controls the stratified train-test split. Conditional on the training set, preprocessing, Laplace optimisation, and quadrature settings are fixed. For each dataset and seed, we use a stratified train-test split with test fraction \(0.30\). Only the training set is used for the Bayesian logistic-regression fit.

The covariates are standardized on the training set and then projected to four principal components using PCA. An intercept is added, so the parameter dimension is \[d=5.\] The model is Bayesian logistic regression. Given design matrix \(X\), labels \(y_i\in\{0,1\}\), and parameter \(\boldsymbol{\theta}\in\mathbb{R}^5\), the negative log posterior objective is \[\sum_i \left[ \log(1+\exp(x_i^\top\boldsymbol{\theta})) - y_i x_i^\top\boldsymbol{\theta} \right] + \frac{1}{2\sigma_0^2}\|\boldsymbol{\theta}\|^2, \qquad \sigma_0=2.\] Equivalently, the prior is \[\boldsymbol{\theta}\sim N(0,4I_5).\]

The posterior is approximated by a Laplace Gaussian approximation. The posterior mean \(\hat{\boldsymbol{\theta}}\) is obtained by minimizing the negative log posterior. We use a trust-region Newton optimizer with analytic gradient and Hessian. If this optimizer fails, we use an L-BFGS-B fallback. The Laplace covariance is the inverse Hessian at \(\hat{\boldsymbol{\theta}}\), with eigenvalues clipped below \(10^{-10}\) for numerical stability. We set \[\boldsymbol{\theta}_0=\hat{\boldsymbol{\theta}}.\]

For each run, prior and posterior Euclidean small-ball masses are estimated for \[B_r(\boldsymbol{\theta}_0) = \{\boldsymbol{\theta}:\|\boldsymbol{\theta}-\boldsymbol{\theta}_0\|_2\leq r\}\] over a geometric grid of \(24\) radii in \[[0.03,0.45].\] The prior covariance is \(4I_5\), while the posterior covariance is the Laplace covariance. The prior mean is \(0\), and the posterior mean is \(\boldsymbol{\theta}_0\).

The small-ball masses are estimated by Sobol quadrature over the Euclidean unit ball. The default number of Sobol points is \[2^{15}=32768\] per run. The unit-ball Sobol points are generated in dimension \(d+1\). The first \(d\) coordinates are transformed into Gaussian directions and normalized to obtain directions on the sphere. The final coordinate is transformed as \(U^{1/d}\) to obtain the radial component. For a radius \(r\), the quadrature points are \[\boldsymbol{\theta}_0+r z_j, \qquad z_j\in B_1(0).\]

For a Gaussian distribution \(N(m,\Sigma)\), the estimated log small-ball mass is \[\log |B_1(0)| + d\log r + \log\left[ \frac{1}{M} \sum_{j=1}^{M} \varphi_{\Sigma}(\boldsymbol{\theta}_0+r z_j-m) \right],\] where \(M=32768\), and \(\varphi_{\Sigma}\) denotes the Gaussian density with covariance \(\Sigma\).

Finite-radius slopes are computed by first differences on the log-log scale: \[s_k = \frac{ \log p(B_{r_{k+1}}(\boldsymbol{\theta}_0)) - \log p(B_{r_k}(\boldsymbol{\theta}_0)) }{ \log r_{k+1}-\log r_k }.\] The slope location is reported at the geometric midpoint \[(r_k r_{k+1})^{1/2}.\] The main-text table reports small-radius slopes fitted over \(r\leq0.06\).

Across datasets and seeds, the plotted aggregate curves use means and standard errors over all runs. In total, the aggregate UCI figure uses \[3\times5=15\] runs.

10.3 Experiment 3: local RE-KL directionality↩︎

Experiment 3 is a one-dimensional toy example designed to show that local RE-KL control is directional. On \((-1,1)\), define \[p(x)=\frac{1}{2}, \qquad q(x)=\frac{1}{4} |x|^{-1/2}.\] Both are probability densities on \((-1,1)\). Around \(\boldsymbol{\theta}_0=0\), their small-ball masses are \[p(B_r(0))=r, \qquad q(B_r(0))=\sqrt r.\]

The plotted local RE-KL quantities use mass-ratio weighted conditional terms. The conditional constants are \[C_{p\|q}=\log 2-\frac{1}{2}, \qquad C_{q\|p}=1-\log 2.\] The plotted finite-radius quantities are \[D_r(p\|q) = \sqrt r \left(\log 2-\frac{1}{2}\right),\] and \[D_r(q\|p) = \frac{1-\log 2}{\sqrt r}.\] Therefore, \[D_r(p\|q)\to0, \qquad D_r(q\|p)\to\infty \qquad (r\to0).\] This example shows that the two directions can behave differently on the same shrinking neighbourhoods. Hence the direction in a local RE-KL assumption is substantive.

References↩︎

[1]
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114, 2013.
[2]
Y. Li and R. Turner, “Rényi Divergence Variational Inference,” in Advances in Neural Information Processing Systems, 2016, vol. 29, Accessed: May 21, 2026. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2016/hash/7750ca3559e5b8e1f44210283368fc16-Abstract.html.
[3]
T. P. Minka, “Divergence measures and message passing,” Microsoft Research, Technical Report MSR-TR-2005-173, 2005.
[4]
D. M. Blei, A. Kucukelbir, and J. D. McAuliffe, “Variational inference: A review for statisticians,” Journal of the American statistical Association, vol. 112, no. 518, pp. 859–877, 2017.
[5]
S. Ghosal, J. K. Ghosh, and A. W. Van Der Vaart, “Convergence rates of posterior distributions,” Annals of Statistics, pp. 500–531, 2000.
[6]
I. Castillo, J. Schmidt-Hieber, and A. van der Vaart, “Bayesian linear regression with sparse priors,” The Annals of Statistics, vol. 43, no. 5, pp. 1986–2018, 2015, doi: 10.1214/15-AOS1334.
[7]
A. W. van der Vaart and J. H. van Zanten, “Rates of contraction of posterior distributions based on gaussian process priors,” The Annals of Statistics, vol. 36, no. 3, pp. 1435–1463, 2008, doi: 10.1214/009053607000000613.
[8]
K. J. Falconer and K. Falconer, Techniques in fractal geometry, vol. 3. Wiley Chichester, 1997.
[9]
Y. Heurteaux, “Dimension of measures: The probabilistic approach,” Publicacions Matemàtiques, vol. 51, no. 2, pp. 243–290, 2007.
[10]
P. Mattila, Geometry of sets and measures in euclidean spaces: Fractals and rectifiability, vol. 44. Cambridge: Cambridge University Press, 1999.
[11]
J. Karamata, “Sur un mode de croissance régulière des fonctions,” Mathematica (Cluj), vol. 4, pp. 38–53, 1930.
[12]
J. Karamata, “Sur un mode de croissance régulière. Théorèmes fondamentaux,” Bulletin de la Société Mathématique de France, vol. 61, pp. 55–62, 1933.
[13]
N. H. Bingham, C. M. Goldie, and J. L. Teugels, Regular variation, vol. 27. Cambridge university press, 1989.
[14]
N. Wan, D. Li, and N. Hovakimyan, “F-divergence variational inference,” Advances in neural information processing systems, vol. 33, pp. 17370–17379, 2020.
[15]
G. Avlogiaris, A. Micheas, and K. Zografos, “On local divergences between two probability measures,” Metrika, vol. 79, no. 3, pp. 303–333, 2016.
[16]
M. Zhang, T. Bird, R. Habib, T. Xu, and D. Barber, arXiv:1907.11891 [stat.ML]“Variational f-divergence Minimization.” arXiv, Dec. 2024, doi: 10.48550/arXiv.1907.11891.
[17]
C. Naesseth, F. Lindsten, and D. Blei, “Markovian score climbing: Variational inference with KL (p|| q),” Advances in Neural Information Processing Systems, vol. 33, pp. 15499–15510, 2020.
[18]
D. McNamara, J. Loper, and J. Regier, “Sequential Monte Carlo for Inclusive KL Minimization in Amortized Variational Inference,” in Proceedings of The 27th International Conference on Artificial Intelligence and Statistics, Apr. 2024, pp. 4312–4320, Accessed: May 21, 2026. [Online]. Available: https://proceedings.mlr.press/v238/mcnamara24a.html.
[19]
Z. Lin, A. Khetan, G. Fanti, and S. Oh, “Pacgan: The power of two samples in generative adversarial networks,” Advances in neural information processing systems, vol. 31, 2018.
[20]
P. Zhong, Y. Mo, C. Xiao, P. Chen, and C. Zheng, “Rethinking generative mode coverage: A pointwise guaranteed approach,” Advances in Neural Information Processing Systems, vol. 32, 2019.
[21]
C. Louizos, K. Ullrich, and M. Welling, “Bayesian compression for deep learning,” in Advances in neural information processing systems, 2017, vol. 30, [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2017/file/69d1fc78dbda242c43ad6590368912d4-Paper.pdf.
[22]
S. Ghosh, J. Yao, and F. Doshi-Velez, “Model selection in bayesian neural networks via horseshoe priors,” Journal of Machine Learning Research, vol. 20, no. 182, pp. 1–46, 2019.
[23]
V. Fortuin et al., “Bayesian neural network priors revisited,” in International conference on learning representations, 2022, [Online]. Available: https://openreview.net/forum?id=xkjqJYqRJy.
[24]
N. G. Polson and V. Ročková, “Posterior Concentration for Sparse Deep Learning,” in Advances in Neural Information Processing Systems, 2018, vol. 31, Accessed: May 23, 2026. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2018/hash/59b90e1005a220e2ebc542eb9d950b1e-Abstract.html.
[25]
G. B. Folland, Real analysis: Modern techniques and their applications. John Wiley & Sons, 1999.
[26]
P. Billingsley, Convergence of probability measures. John Wiley & Sons, 2013.
[27]
S. Kullback and R. A. Leibler, “On information and sufficiency,” The annals of mathematical statistics, vol. 22, no. 1, pp. 79–86, 1951.
[28]
B. Póczos and J. Schneider, “On the estimation of \(\alpha\)-divergences,” in Proceedings of the fourteenth international conference on artificial intelligence and statistics, 2011, pp. 609–617.
[29]
R. Beran, “Minimum hellinger distance estimates for parametric models,” The annals of Statistics, pp. 445–463, 1977.
[30]
K. Hirano and J. R. Porter, “Asymptotic efficiency in parametric structural models with parameter-dependent support,” Econometrica, vol. 71, no. 5, pp. 1307–1338, 2003.