A Simple Bivariate Example of Fast Convergence Rates for Maximum Likelihood Estimates


Abstract

We present a one-parameter family of bivariate absolutely continuous distributions based on location-scale family of variance Gaussian mixtures, with continuous densities with the same support (effective domain). The maximum likelihood estimation of the location parameter converges to the true value faster than the classic square root rate. In fact, we can obtain any convergence rate given by a regularly varying function with index greater than 0.5, and some convergence rates given by regularly varying functions with index 0.5 but faster than the classic square root rate.

1 Introduction↩︎

1.1 Classic results↩︎

Take a (univariate or multivariate) family of absolutely continuous distributions, parameterized by \(\theta \in \mathbb{R}\). Their Lebesgue density is \(f(x, \theta)\). Take a sample of independent identically distributed (IID) random variables \(X_1, \ldots, X_n\) from this distribution. Our goal is to study the maximum likelihood estimate (MLE) \(\hat{\theta}_n\). Under classic results by Sir Ronald Fisher and others, this estimate is asymptotically normal: \[\sqrt{n}(\hat{\theta}_n - \theta) \stackrel{d}{\to} \mathcal{N}_d(\mathbf{0}, \Sigma^{-1}),\] where \(\Sigma\) is the Fisher information, defined as the variance of the score function \(\nabla_{\theta}\log f(x, \theta)\) with \(x := X\) having density \(f(x, \theta)\). See for example [1]. Therefore, the classic convergence rate is \(\sqrt{n}\): The distance between the MLE and the true value is inversely proportional to the square root of the number of data points. However, there exist cases when the convergence rate is faster: The classic example is the uniform distribution on \([0, \theta]\), where \(\theta > 0\) is an unknown parameter. The MLE is then \(\hat{\theta}_n = \max(X_1, \ldots, X_n)\), and converges to the true value with fast rate \(n\) instead of \(\sqrt{n}\), see [1]: \[n(\theta - \hat{\theta}_n) \stackrel{d}{\to} \mathrm{Exp}(1),\quad n \to \infty.\] The reason is dependence of the effective domain \([0, \theta]\) (and not just the likelihood function) upon the unknown parameter \(\theta\). One can construct other such examples, with rates intermediate between \(\sqrt{n}\) and \(n\). See, for example, leave-one-out likelihood in [2], with super-effective rates \(n^{\varepsilon + 1/2}\). A case when the Fisher information is infinite and the score function converges with slower-than-standard rate is considered in [3].

1.2 Our contribution↩︎

A natural question is to find a family of continuous densities with the same effective domain, for which the MLE exists and is unique, but converges to the true value faster than \(\sqrt{n}\). In this short note, we present a simple example of a bivariate family of densities with this property: Theorem 1. We can get any rate faster than \(n^{0.5+\varepsilon}\) for example, \(n\) or \(n^2\). We can also get rates of the type \(n\ln^{1 + \varepsilon}n\) or \(n\ln\ln n\), or others. Take a random variable \(W > 0\) with continuous Lebesgue density \(g_W : (0, \infty) \to (0, \infty)\). Denote its distribution by \(Q\). The Gaussian variance mixture is given by \[\label{eq:repr} X = \mu + \sqrt{W}Z,\quad Z \sim \mathcal{N}(0, 1),\quad W \sim Q,\tag{1}\] for independent \(Z\) and \(W\). These mixtures (and their generalization, the Gaussian mean-variance mixture) are a classic study object. These include, for example, a large class of generalized hyperbolic distributions, [4]. For a multivariate generalization, see [5]. For semiparametric estimation, see [6]. Then \(X\) has a continuous Lebesgue density \(g_X : (0, \infty) \to (0, \infty)\). The conditional distribution of \(X\) given \(W\) is \(X\mid W \sim \mathcal{N}(\mu, W)\). The bivariate random variable \((X, W)\) has joint density \[\label{eq:joint} g(x, w; \mu) = g_W(w)\cdot\frac{1}{\sqrt{2\pi w}}\exp\left[-\frac{(x - \mu)^2}{2w}\right].\tag{2}\] Assume we observe \((X_1, W_1), \ldots, (X_n, W_n)\). We then have: \[X_i\mid W_i \sim \mathcal{N}(\mu, W_i), \quad i = 1, \ldots, n.\] And \(W_1, \ldots, W_n \sim Q\) are independent identically distributed. Applications include:

  • Measuring a quantity with observable variance (observation quality) \(W_i\) dependent on the observation \(X_i\);

  • Observing stock index returns \(X_i\) and time-dependent volatility \(W_i\) (measured by VIX, for example), and estimating long-term mean returns;

  • Connections with empirical Bayes estimation of multivariate normal.

1.3 Motivation↩︎

This research was inspired by our recent manuscript on \((X, W)\), where \(W\) is Gamma and \(X\) is Gaussian mean-variance mixture: \(X\mid W \sim \mathcal{N}(\mu + \delta W, \sigma^2W)\). To our surprise, we established faster-than-usual convergence rate \(\hat{\mu}_n \to \mu\) for some cases. In particular, if \(W\) is exponential (and then \(X\) is asymmetric Laplace), then the rate is \(\sqrt{n\ln n}\). This works for the case of only variance mixture, without mixing using the mean, with \(\delta = 0\) and symmetric Laplace \(X\). There, the joint density of \((X, W)\) is continuous on \([0, \infty)\times \mathbb{R}\).

In this short note, we generalize this idea for various densities of \(W\) other than exponential or Gamma. We have only one parameter instead of five: De facto, we assume \(\delta = 0\) and \(\sigma = 1\); and the marginal distribution of \(W\) is assumed to be completely known and does not contain any parameters to be estimated.

1.4 Further research↩︎

We stress that we could have generalized this for Gaussian mean-variance mixtures. This would involve three parameters \(\mu, \delta, \sigma\) instead of only \(\mu\). Moreover, we could have included other unknown parameters in the distribution of \(W\), such as the shape and the scale for the Gamma distribution. However, we deliberately chose to present the simplest possible example in this short note. Generalizations are left for further research.

One could connect this to empirical Bayes method. Also, we could consider the case when the mean \(\mu\) differs from observation to observation: \(X_i\mid W_i \sim \mathcal{N}(\mu_i, W_i)\), for \(i = 1, \ldots, n\), but the vector \((\mu_1, \ldots, \mu_n)\) of such means belongs to a certain linear subspace.

2 Main Results↩︎

2.1 Preliminary remarks↩︎

This density 2 is a continuous function \(g : \mathbb{R}\times (0, \infty) \to (0, \infty)\). The effective region is always the same: \(\mathbb{R}\times(0, \infty)\), irrespective of the values of the parameter \(\mu\). Inside this effective region, there are no singularities.

For \(x \ne \mu\), as \(w \to 0\), then \(g(x, w) \to 0\). But for \(x = \mu\), as \(w \to 0\), we get \(g(\mu, w) \sim (2\pi w)^{-1/2}g_W(w)\). Only at \((\mu, 0)\), we might have some singularity, depending on the behavior of the density \(g_W\) at zero. The marginal density of \(X\) does not have any singularities.

We can rewrite the density 2 as a curved exponential family \[\frac{g_W(w)}{\sqrt{2\pi w}}\exp\left[-\frac{x^2}{2w} + \frac{x}{w}\mu - \frac{1}{2w}\mu^2\right]\] with sufficient statistic \(\mathbf{T} = (\mathbf{T}_1, \mathbf{T}_2) := (x/w, -1/(2w))\). The family is curved since we have a two-dimensional sufficient statistic but only one-dimensional parameter.

2.2 Fisher information↩︎

Let us compute the score function and the Fisher information for the joint density 2 . The score function is \[s(x, w, \mu) = \frac{\partial \ln g(x, w, \mu)}{\partial \mu} = -\frac{x-\mu}{w}.\] The Fisher information is \[\label{eq:F} \Sigma = \mathbb{E}[s^2(X, W, \mu)] = \mathbb{E}\left[\frac{(X - \mu)^2}{W^2}\right].\tag{3}\] Plugging 1 in 3 , we get: \(\Sigma = \mathbb{E}[1/W]\). If this is finite, we have classic convergence rate \(\sqrt{N}\). But we are interested in the case when \(\Sigma = \infty\).

2.3 Computation of the estimate↩︎

We have one unknown parameter \(\mu\). Now, pick a sample from this distribution: a sequence of IID random variables \((X_1, W_1), \ldots, (X_n, W_n)\). Write the log-likelihood function \[\label{eq:LL} \ell(\mu) := \sum\limits_{i=1}^n\ln g(X_i, W_i, \mu).\tag{4}\] The maximum likelihood estimate (MLE) is the value \(\hat{\mu}_N\) which maximizes the log-likelihood function \(\ell\) from 4 . If we had only one random variable \(X\) observed, then it would not be so easy to solve for MLE. But knowledge of \(W\) gives us critical additional information, so the solution is straightforward and explicit, below in 5 . Such MLE exists and is unique, see [7] or [8]: \[\label{eq:MLE} \hat{\mu}_n = \frac{X_1/W_1 + \ldots + X_n/W_n}{1/W_1 + \ldots + 1/W_n}.\tag{5}\]

2.4 The main result↩︎

Define the cumulative distribution function (CDF) of \(W\): \[G(u) = \int_0^ug_W(w)\,\mathrm{d}w = \mathbb{P}(W \le u),\quad u \in [0, \infty).\] It is strictly increasing on \([0, \infty)\), and its range is \([0, 1)\). Therefore, we can define its inverse \(H : [0, 1) \to [0, \infty)\) as a strictly increasing function such that \(H(G(u)) \equiv u\). Also, define \(B(t) = 1/H(1/t)\), and let \[\label{eq:A-def} A(t) = t\int_0^{B(t)}G(1/u)\,\mathrm{d}u,\quad t > 0,\tag{6}\] Below is the main result of this short note. Let \(Z \sim \mathcal{N}(0, 1)\) be a standard normal random variable. For \(\beta \in (0, 1)\), define \(U_{\beta} > 0\) to be the stable subordinator with index \(\beta\), independent of \(Z\) (see [9] or any other standard reference): \[\mathbb{E}[e^{-\lambda U_{\beta}}] = \exp(-\Gamma(1 - \beta)\lambda^{\beta}),\quad \lambda > 0.\] A function \(f : (0, \infty) \to (0, \infty)\) is called regularly varying at \(a\) with index \(b\), where \(a = 0\) or \(a = \infty\), if for any \(s > 0\) we have: \[\lim\limits_{t \to a}\frac{f(st)}{f(t)} = s^b\quad for all\quad s > 0.\] For the case \(b = 0\), we say \(f\) is slowly varying at \(a\). We say that \(f(t) \sim g(t)\) as \(t \to a\) for \(a = 0, \infty\), if \(f(t)/g(t) \to 1\) as \(t \to a\).

Theorem 1. Assume \(G\) is regularly varying at zero with index \(\beta \in (0, 1]\).

(a) The function \(B\) is regularly varying with index \(1/\beta\) at infinity. If \(\beta = 1\), then the function \(A\) is regularly varying with index 1 at infinity.

(b) If \(\beta \in (0, 1)\), then \[\label{eq:beta601} \sqrt{B(n)}\cdot(\hat{\mu}_n - \mu) \stackrel{d}{\to} U_{\beta}Z,\quad n \to \infty,\qquad{(1)}\]

(c) If \(\beta = 1\), then we get: \[\label{eq:beta611} \sqrt{A(n)}\cdot(\hat{\mu}_n - \mu) \stackrel{d}{\to} Z,\quad n \to \infty,\qquad{(2)}\]

under the additional assumption: \[\label{eq:cond} \mathbb{E}\left[W^{-1}\right] = \int_0^{\infty}G(1/u)\,\mathrm{d}u = \infty.\qquad{(3)}\]

In particular, if the density \(g\) is regularly varying at zero with index \(\beta - 1\), then the CDF \(G\) is regularly varying at zero with index \(\beta\), [10].

2.5 Choice of convergence rates↩︎

Let us discuss special cases of Theorem 1. Various choices of \(g_W\) and therefore \(G\) lead to various convergence rates.

First, consider the case \(\beta < 1\). Take any regularly varying function \(C(t) > 0\) at infinity with index \(\alpha > 0.5\), and with \(C' > 0\) everywhere. Then there exists a distribution \(Q\) on \((0, \infty)\) such that \(B(t) = C(t)\) in Theorem 1. Indeed, just take an inverse function \(F\) to \(C^2\): \(F(C^2(t)) = t\). The function \(C^2\) is strictly positive and has strictly positive derivative. Therefore, the inverse function also has strictly positive derivative. By the results of [10], this function is \(1/(2\alpha)\) at infinity. Then define \(G(u) = 1/F(1/u)\) to be the cumulative distribution function of \(W\). Following [10], this is a regularly varying function at zero with index \(\beta = 1/(2\alpha) \in (1/2, 1)\).

It is more complicated to find all possible convergence rates for \(\beta = 1\). We leave it for the future resesarch. Some examples are below.

2.6 Examples↩︎

Fix \(\gamma \in \mathbb{R}\), and assume \[\label{eq:G-asymp} G(u) \sim \mathrm{const}\cdot u^{\beta}|\ln u|^{\gamma},\quad u \to 0.\tag{7}\] Take the logarithm: \(\ln G(u) \sim \beta\ln u\). Therefore, \(\ln v \sim \beta\ln H(v)\). A standard exercise shows that the inverse function \(H\) of \(G\) satisfies \(H(u) \sim \mathrm{const} \cdot u^{1/\beta}|\ln u|^{-\gamma/\beta}\) as \(u \to 0\). And the function \(B(t) = 1/H(1/t)\) satisfies \[B(t) \sim \mathrm{const}\cdot t^{1/\beta}(\ln t)^{\gamma/\beta},\quad t \to \infty.\] Therefore, the convergence rate for \(\beta \in (1/2, 1)\) is \(\sqrt{B(t)} \sim \mathrm{const}\cdot t^{1/(2\beta)}(\ln t)^{\gamma/(2\beta)}\). By choosing \(a = 1/(2\beta)\) and \(b = \gamma/(2\beta)\), we can make any convergence rate \(t^a(\ln t)^b\) for any \(a \in (0.5, 1)\) and \(b \in \mathbb{R}\).

For \(\beta = 1\), it is harder to compute the convergence rate. From 7 , we get: \(G(1/u) \sim \mathrm{const}\cdot u^{-1}(\ln u)^{\gamma}\). The second condition in ?? holds if and only if \(\gamma \le -1\). Here, in turn, we consider two cases: \(\gamma = -1\) and \(\gamma < -1\). For \(\gamma = -1\), we compute \[\int_0^{t}G(1/u)\,\mathrm{d}u \sim \mathrm{const}\cdot \ln\ln t,\quad t \to \infty.\] Since \(B(t) \sim \mathrm{const}\cdot t(\ln t)^{-1}\), we get: \(A(t) \sim \mathrm{const}\cdot t\ln\ln t\) as \(t \to \infty\). The convergence rate is \(\sqrt{A(t)} \sim \mathrm{const}\cdot \sqrt{t\ln\ln t}\). Now, for \(\gamma < -1\), we get: \[\int_0^{t}G(1/u)\,\mathrm{d}u \sim \mathrm{const}\cdot (\ln t)^{1+\gamma},\quad t \to \infty.\] Also, \(B(t) \sim \mathrm{const}\cdot t(\ln t)^{\gamma}\). Therefore, \[A(t) = t\int_0^{B(t)}G(1/u)\,\mathrm{d}u = \mathrm{const}\cdot t(\ln t)^{1 + \gamma},\quad t \to \infty.\] Therefore, the convergence rate is \(\sqrt{A(t)} \sim \mathrm{const}\cdot t^{1/2}(\ln t)^{(1 + \gamma)/2}\). By choosing \(b = (1+\gamma)/2\), we can make any convergence rate of the type \(t^{1/2}\cdot (\ln t)^b\) for any \(b > 0\).

Similarly, we can have convergence rates of the type \(t^a\cdot (\ln t)^b\cdot (\ln\ln t)^c\), for \(a \ge 1/2\), \(b, c \in \mathbb{R}\), so that if \(a = 1/2\), then either \(b > 0\), or \(b = 0\) and \(c > 0\). We can have even more iterations of logarithms, as long as the rate is asymptotically faster than the standard convergence rate \(t^{1/2}\).

2.7 Conclusions↩︎

We have created a bivariate family of distributions with one unknown parameter and MLE convergence rates faster than \(\sqrt{N}\). We could generalize it for \(Z \sim \mathcal{N}(0, \sigma^2)\) with unknown variance \(\sigma^2\), or normal mean-variance mixture \[X \mid W \sim \mathcal{N}(\mu + \theta W, \sigma^2W)\] with unknown parameters \(\mu, \theta, \sigma\). We anticipate that the same non-standard convergence rates will hold for the MLE of \(\mu\). But parameters \(\theta\) or \(\sigma\) have MLEs which are asymptotically normal with standard convergence rates \(\sqrt{n}\). Moreover, we can add some unknown parameters in the distribution \(Q\) of \(W\), just like we did for the Gamma distribution in [11]. Again, we anticipate that the convergence rates for the MLE will stay the same.

3 Proof of Theorem 1↩︎

3.1 Overview of the proof↩︎

The proof is based on classic results about domains of attraction and stable limit theorems, taken from [10], also stated and proved in [10]. The proof of (A) is based on results from another Appendix 6 of the same monograph [10]. First, we show we can rewrite with a \(Z \sim \mathcal{N}(0, 1)\) independent of \(W_1, \ldots, W_n \sim Q\): \[\label{eq:final-repr} \hat{\mu}_n - \mu = S_n^{-1/2}Z,\quad S_n := \sum\limits_{i=1}^nW_i^{-1}.\tag{8}\] Next, we show: For \(\beta \in (0, 1)\), \[\label{eq:main-A} \frac{S_n}{B(n)} \stackrel{d}{\to} U_{\beta},\quad n \to \infty.\tag{9}\] And for \(\beta = 1\), it sufffices to show that \[\label{eq:main-B} \frac{S_n}{A(n)} \stackrel{d}{\to} 1,\quad n \to \infty.\tag{10}\] Combining 9 with 8 , we get ?? . Combining 10 with 8 , we get ?? . Thus, we complete the proof of the main statement of Theorem 1.

3.2 Proof of (A)↩︎

We apply [10] to get that the inverse function \(H(1/t)\) is regularly varying with index \(-1/\beta\). Therefore, \(B(t) = 1/H(1/t)\) is regularly varying with index \(1/\beta\). Finally, let us show that \(A(t)\) is regularly varying with index 1 at infinity. This follows from successive applications of classic results on regularly varying functions from the same monograph: [10]. The function \(G(1/u)\) is regularly varying at infinity with index \(-1\). Therefore, the function \(K: t \mapsto \int_0^tG(1/u)\,\mathrm{d}u\) is slowly varying at infinity. But the function \(B\) is regularly varying at infinity with index \(1\). Thus, the function \(t \mapsto K(B(t))\) is also slowly varying at infinity. Finally, the function \(A(t) = tK(B(t))\) is regularly varying at infinity with index 1.

3.3 A technical lemma↩︎

Let us show that, if \(\beta = 1\), then \[\label{eq:A-B} \frac{A(t)}{B(t)} \to \infty,\quad t \to \infty.\tag{11}\] Apply [10] to the function \(V(s) = G(1/s)\) (regularly varying at infinity with index \(-1\)), we get: \[\label{eq:L-def} L(t) := \frac{1}{tV(t)}\int_0^tV(s)\,\mathrm{d}s \to \infty,\quad t \to \infty,\tag{12}\] Rewrite 12 as follows: \[\label{eq:L-int} \int_0^tG(1/s)\,\mathrm{d}s = tV(t)L(V(t)) = tG(1/t)L(t).\tag{13}\] Plugging 13 into 6 , and using \(G(1/B(t)) = G(H(1/t)) = 1/t\), we get: \[\frac{A(t)}{B(t)} = \frac{t}{B(t)}\int_0^{B(t)}G(1/s)\,\mathrm{d}s = \frac{t}{B(t)}B(t)G(1/B(t))L(B(t)) = L(B(t)) \to \infty.\]

3.4 Conditional distribution↩︎

The conditional distribution of \(\hat{\mu}_n\) given \(W_1, \ldots, W_n\) is \[\label{eq:conditional} \hat{\mu}_n \mid W_1, \ldots, W_n \sim \mathcal{N}(\mu, S_n^{-1}),\tag{14}\] with \(S_n\) given 8 . Let us show 14 . Apply 1 to each \(X_i\): For two independent sequences of independent identically distributed random variables \[Z_1, \ldots, Z_n \sim \mathcal{N}(0, 1),\quad W_1, \ldots, W_n \sim Q,\] we can represent each observation \(X_i\) as \[\label{eq:X-W-i} X_i = \mu + \sqrt{W_i}Z_i,\quad i = 1, \ldots, n.\tag{15}\] Plugging 15 into the first equation in 5 , we get: \[\begin{align} \label{eq:big} \hat{\mu}_n = \frac{\mu/W_1 + \ldots + \mu/W_n}{1/W_1 + \ldots + 1/W_n} + \frac{Z_1/\sqrt{W_1} + \ldots + Z_n/\sqrt{W_n}}{1/W_1 + \ldots + 1/W_n}. \end{align}\tag{16}\] The first term in the right-hand side of 16 is equal to \(\mu\), after everything else in the numerator and the denominator cancels out. The second term, conditional on the values of \(W_1, \ldots, W_n\), is normally distributed with mean zero and variance \[\frac{(1/\sqrt{W_1})^2 + \ldots + (1/\sqrt{W_n})^2}{(1/W_1 + \ldots + 1/W_n)^2} = \frac{1}{1/W_1 + \ldots + 1/W_n}.\] This completes the proof of 14 . Using this, we get 8 .

3.5 Proof of 9↩︎

To show 9 , we apply the aforementioned results from [10]. We need to match the notation: \(\xi_k := W_k^{-1}\), \[F_+(t) = \mathbb{P}(1/W \ge t) = \mathbb{P}(W \le 1/t) = G(1/t);\] and \(F_-(t) = 0\) because the random variable \(1/W\) is positive; \(F_0(t) = F_+(t) + F_-(t) = G(1/t)\); and this function is regularly varying at infinity with index \(-1/\beta\). Moreover, \(b(n) = F_0^{(-1)}(1/n)\), where \(F_0^{(-1)}(t) = 1/H(t)\) is the inverse of \(F_0(t) = G(1/t)\). Thus \(b(n) := B(n)\). Next, \(V = F_+\) and \(W = 0\). Therefore, \(\rho_+ = 1\) and \(\rho = 1\). Also, we use discussion on one-sided stable distributions after [10]. Finally, we need to show ?? , otherwise, per discussion before [10], we would need to subtract this mean to get such weak convergence. But ?? follows from the assumption that \(G\) is regularly varying at zero with index \(\beta < 1\); therefore, \(G(1/u)\) (which is the CDF of \(1/W\)) is regularly varying at infinity with index \(-\beta > -1\), and \[\mathbb{E}[W^{-1}] = \int_0^{\infty}G(1/u)\,\mathrm{d}u = \infty.\] The rest of ?? simply follows from the aforementioned theorem [10].

3.6 Proof of 10↩︎

Similarly to the case (B), we have the same notation, and \[V_I(t) = \int_0^tV(y)\,\mathrm{d}y = \int_0^tG(1/y)\,\mathrm{d}y,\quad W_I = 0.\] Define \(C \approx 0.5772\) to be the Euler’s constant. We get \(\rho = 1\), since there is no left tail. From ?? , we get \(\mathbb{E}[1/W] = \infty\); this discussion was already done in the proof of (B). In the notation of [10], \[\label{eq:appendix} A_n = \frac{n}{b(n)}\left[V_I(b(n)) - W_I(b(n))\right] - \rho C = \frac{n}{B(n)}\int_0^{B(n)}G(1/y)\,\mathrm{d}y - C.\tag{17}\] Applying 11 to 17 , we get: \[A_n = \frac{A(n)}{B(n)} - C \sim \frac{A(n)}{B(n)},\quad n \to \infty.\] The main result from [10]: A weak convergence to a limit \(\zeta\) (whose exact distribution is not important for us): \[\label{eq:fundamental} \frac{1}{B(n)}S_n - A_n \stackrel{d}{\to} \zeta,\quad n \to \infty.\tag{18}\] It suffices to note that \(B(n)A_n \sim A(n)\) as \(n \to \infty\). Divide 18 by \(A_n\), and apply 11 again to finish the proof of ?? : \[\frac{1}{B(n)A_n}S_n \stackrel{d}{\to} 1,\quad n \to \infty.\]

References↩︎

[1]
George Casella, Roger L. Berger(2002). Statistical Inference. Second edition. Duxbury.
[2]
Krzysztof Podgorski, Jonas Wallin(2015). Maximizing Leave-one-out Likelihood for the Location Parameter of Unbounded Densities. Annals of the Institute of Statistical Mathematics67(1), 19–38.
[3]
Svetlana Litvinova, Mervyn J. Silvapulle(2012). A Test When The Fisher Information May Be Infinite, Exemplified by a Test for Marginal Independence in Extreme Value Distributions. Journal of Statistical Planning and Inference142(6), 1506–1515.
[4]
Ole Eiler Barndorff-Nielsen, John T. Kent, Michael Sorensen(1982). Normal Variance-Mean Mixtures and \(z\) Distributions. International Statistical Review50, 145–159.
[5]
Erik Hintz, Marius Hofert, Christiane Lemieux(2021). Normal Variance Mixtures: Distribution, Density and Parameter Estimation. Computational Statistics and Data Analysis157, 107175.
[6]
Denis Belomestny, Vladimir Panov(2018). Semiparametric Estimation in the Normal Variance-Mean Mixture Model. Statistics: A Journal of Theoretical and Applied Statistics52(3), 571–589.
[7]
Yanjun Han, Abhishek Shetty, Jacob Shkrob(2026). An Empirical Bayes Perspective on Heteroskedastic Mean Estimation. arXiv:2603.13499.
[8]
Asaf Weinstein, Ma Zhuang, Lawrence D. Brown, Cun-Hui Zhang(2018). Group-Linear Empirical Bayes Estimates for a Heteroscedastic Normal Mean. Journal of the American Statistical Association113(522), 698–710.
[9]
Fred W. Steutel, Klaas van Harn(2004). Infinite Divisibility of Probability Distributions on the Real Line. Marcel Dekker. Monographs and Textbooks in Pure and Applied Probability259.
[10]
Alexander Borovkov(2013). Probability Theory. Springer.
[11]
Tomasz Kozubowski, Andrey Sarantsev, James Spiker(2026). Modeling Stock Returns and Volatility Using Bivariate Gamma Generalized Laplace Law.