A scale-free density bound for Gaussian maxima


Abstract

We derive a scale-free bound on the density of the maximum of a centered Gaussian vector. The basic bound is non-uniform, depends logarithmically on the dimension, and allows any covariance matrix. When the largest marginal variance is separated from zero, it implies that the density of the maximum is uniformly controlled at all quantiles above \(\frac{2}{3}\), which is sufficient for many hypothesis testing applications; it yields validity of Gaussian and bootstrap approximations for maxima of high-dimensional sums at test levels \(\alpha < \frac{1}{3}\) without further restricting the covariance. The result also implies uniform anti-concentration bounds and control of the variance of the maximum with optimal dimension dependence, in terms of the expectation of the maximum and the largest marginal variance. We discuss implications for high-dimensional correlation testing, time-uniform sequential testing, and non-parametric inference under latent, low-dimensional structure.

1 Introduction↩︎

Gaussian and bootstrap approximations for maxima of sums have become an essential component of modern statistics, providing practical methods for inference in problems of high-dimensional selection, multiple testing, and non-parametric estimation. These tools are especially useful when the distribution being approximated has a complex form and its covariance is difficult to characterize, yet simulation-based inference remains effective [1].

A certain limitation of the theory, however, concerns degeneracy: settings where coordinates have vanishing variance, a high degree of correlation, and where standardization is not viable. In these cases, it becomes difficult to verify that the approximating law has a bounded density (or is anti-concentrated), which is needed to justify conventional statistical inference. This setting arises in a wide range of modern statistical problems, from multiple testing with genomic or gene-expression data, to anytime-valid sequential testing, to the construction of uniform confidence bands when data have latent, low-dimensional structure.

In each of these cases, the standard remedy of dividing by coordinate-wise standard deviations can fail. This happens for two primary reasons: as \(\sigma(x)\) approaches zero, uniform control of \(|\hat{\sigma}(x)/\sigma(x)-1|\) becomes challenging, and, moreover, the standardized process may not be tight. In the context of high-dimensional correlation testing, the relevant “coordinates” are sparse linear combinations \(\left\langle \beta,\,X\right\rangle\), and estimating the variance of each such combination is non-trivial; indeed, certifying that the smallest restricted eigenvalue exceeds any fixed threshold is NP-hard [2], [3]. In anytime-valid sequential testing, standardizing the time-dependent variance results in an approximating Gaussian law that is not tight. In non-parametric inference on data with latent low-dimensional structure, pointwise variances depend strongly on the unknown intrinsic dimension and local geometry, causing the two problems to compound.

1.1 Primary contributions↩︎

This paper studies anti-concentration and Gaussian approximation under degeneracy. We state density bounds for the maximum which depend logarithmically on the dimension and, importantly, place minimal restrictions on the covariance. This leads to Gaussian and bootstrap approximations under similarly permissive conditions.

The most general bound covers the upper quantiles of the distribution, including the most relevant range for hypothesis testing. As a consequence, we demonstrate validity of standard Gaussian and bootstrap critical values for maxima of high-dimensional sums at levels \(\alpha < \frac{1}{3}\) under the simple requirement of a single non-degenerate coordinate.

These arguments further imply variance bounds and global anti-concentration bounds (i.e., ones that hold at all quantiles of the distribution), which have optimal dependence on the ambient dimension and which remain sharp under degeneracy. The following summarizes the main results.

Theorem 1 (Main results, informal). Let \((Z_1, \ldots, Z_p) \sim N(0,\Sigma)\) be a centered Gaussian vector in \(\mathbb{R}^p\) for \(p \ge 3\), with \(\Sigma \ne 0\). Let \(q_{\frac{2}{3}}^M\) denote the \(\frac{2}{3}\) quantile of the distribution of \(M = \max_{1 \le i \le p}Z_i\). Put \(\sigma_{\text{max}}^2 = \max_{1 \le i \le p} \mathbb{E}[Z_i^2]\). The law of \(M\) has a density \(f\) on \((0,\infty)\)1, and for \(t > 0\), \[f(t) \le \frac{4\log p}{t}.\] This further implies:

  1. in general, \(f(t) \lesssim \sigma_\text{max}^{-1}\log p\) holds for all \(t \ge q_{\frac{2}{3}}^M\);

  2. if \(\mu = \mathbb{E}[M]\) is positive, then \(\sqrt{\mathrm{Var}(M)} \gtrsim \mu / (\log p);\)

  3. if \(\mu > 0\), then for every \(t \in \mathbb{R}\) and every \(\varepsilon \ge \sigma_{\mathrm{max}}\exp\{-\mu^2/(8\sigma_{\mathrm{max}}^2)\}\), \[\mathbb{P}\{t \le M \le t+\varepsilon\} \;\lesssim\; \frac{\varepsilon \log p}{\mu}.\]

The first stated bound imposes no condition on the covariance \(\Sigma\) whatsoever; it is scale-free in this sense. Implication [item:variance-weak] shows that the density is bounded at testing-relevant quantiles whenever \(\max_i\sigma_i\) is bounded away from zero, regardless of how many entries of \(Z\) have near-zero variance or how correlated they may be (Proposition 2). Implication [item:variance-lb] is a lower bound on the variance of the maximum in terms of its expectation alone (Proposition 3).

Implication [item:anticon] is a uniform anti-concentration bound which replaces the usual requirement \(\min_i \sigma_i \gtrsim 1\) with a condition that \(Z\) is sufficiently complex: the distribution of the maximum is not dominated by the largest single coordinate. Thus, it extends standard anti-concentration bounds to degenerate vectors in the relevant regime \(\mu \gg \sigma_{\mathrm{max}}\), which arises broadly in high-dimensional and non-parametric statistics. In the case of \(M^* = \max_{1 \le i \le p}|Z_i|\), this condition can be dropped and only \(\mathbb{E}[M^*] > 0\) is needed. See Corollary 2 for the precise statement.

The conclusions of Theorem 1 sharpen when further information on \(\Sigma\) is available, such as the weak variance decay considered by [4]; more detailed versions of the result and further discussion, including a characterization of anti-concentration under weak variance decay and an investigation of the density at extremal quantiles, are contained in Section 2.

By combining Theorem 1 with established couplings (see e.g. [5]), we recover validity of critical values based on Gaussian and bootstrap approximations of maxima of high-dimensional sums under degeneracy. Formal results of this type, and analogous statements for the multiplier bootstrap, are given in Section 3.

Corollary 1 (Approximation under degeneracy, informal). Let \(S_n = \frac{1}{\sqrt n}\sum_{i=1}^n X_i\) be an i.n.i.d. sum of centered random vectors in \(\mathbb{R}^p\) with unrestricted variance \(\Sigma\), such that \(\|X_{ij}\|_{\psi_1} \lesssim 1\). Let \(Z\) be a Gaussian vector with variance \(\Sigma\), and let \(t_q\) be the \(q^{\mathrm{th}}\) quantile of \(\max_{1 \le j \le p} Z_j\). Let \(\sigma_{\max}^2\) be the largest diagonal entry of \(\Sigma\). Then, with \(\Delta_{n,p} = (\log^7 p )/ n\), \[\sup_{q \ge \frac{2}{3}}\left|\,\mathbb{P}\Big\{\max_{1 \le j \le p}(S_n)_j \le t_q\Big\} - \mathbb{P}\Big\{\max_{1 \le j \le p} Z_j \le t_q\Big\}\right| \;\lesssim\; \sigma_{\max}^{-2/3}\Delta_{n,p}^{1/6}\] Similarly, put \(\mu^* = \mathbb{E}\{\max_{1 \le j \le p}| Z_j|\}\) and note that \(\mu^* \gtrsim \sigma_{\max}\). Then, \[\sup_{t \in \mathbb{R}}\left|\,\mathbb{P}\Big\{\max_{1 \le j \le p}|(S_n)_j| \le t\Big\} - \mathbb{P}\Big\{\max_{1 \le j \le p}| Z_j| \le t \Big\}\right| \;\lesssim\; (\mu^*)^{-2/5}\Delta_{n,p}^{1/10}\]

Corollary 1 and its bootstrap analogs (formally stated as Corollaries 3-8) show that a single non-degenerate coordinate is sufficient for high-dimensional approximation of the upper quantiles of \(\max_j (S_n)_j\), and for all quantiles of \(\max_j |(S_n)_j|\), in the regime \((\log^7 p)/ n \to 0\).

To motivate our results, we describe three examples of modern statistical problems involving approximation by degenerate Gaussian vectors. A detailed investigation is left to forthcoming work.

Example 1: robustness of sparse correlation detection↩︎

Detecting sparse correlationswhether some subset \(S \subset [p]\) of at most \(s\) entries of the covariate vector \(X \in \mathbb{R}^p\) carries a non-trivial linear association with an outcome \(Y\)is a fundamental task in high-dimensional inference. [6] consider the null hypothesis of no association, which amounts to a moment restriction that for all \(\beta \in \mathbb{R}^p\) with \(\|\beta\|_2 = 1\) and \(\|\beta\|_0 \le s\), \(\mathbb{E}[\left\langle \beta,\,X_i\right\rangle Y_i] = 0.\)2 They show that the null distribution of \[\label{eq:max-sparse-corr} M_{s,n} \;\mathrel{\vcenter{:}}= \; \sup_{\|\beta\|_2 = 1 \atop \|\beta\|_0 \le s} \frac{1}{n}\sum_{i=1}^n \left\langle \beta,\,X_i\right\rangle\, Y_i,\tag{1}\] is approximated by \(\sup_\beta \left\langle \beta,\,Z\right\rangle\) for \(Z \sim N(0, \Sigma_n)\) with \(\Sigma_n\) the sample covariance of \((Y_i X_i)_{i=1}^n\).

The resulting test was shown to be valid provided that the restricted eigenvalue \[\label{eq:gwas-restricted-eig} \nu_s \;=\; \inf_{\|\beta\|_2 = 1 \atop \|\beta\|_0 \le s}\, \mathbb{E}\langle \beta, X_i\rangle^2\tag{2}\] is bounded away from zero. In their framework, the condition provides a natural route to the anti-concentration needed for bootstrap validity, as it also underlies methods used to discover sparse correlations. However, it is challenging to verify: lower-bounding \(\nu_s\) from data is itself NP-hard [2], [3]. It also matters in practice: in genetic data, meiosis produces blocks of highly correlated coordinates whose effective dimension is well below the block length [7], yielding \(s\)-sparse contrasts \(\beta\) with \(\mathbb{E}\langle \beta, X_i\rangle^2 \approx 0\). A similar phenomenon arises in gene-expression data [8].

Corollary 1, combined with a standard discrete approximation, suggests that bootstrap critical values for \(M_{s,n}\) are asymptotically valid at every level \(\alpha < \frac{1}{3}\) with no condition on \(\nu_s\) whatsoever. In this case, a sparse correlation flagged at level \(\alpha\) by the bootstrap test retains its usual interpretation, even when the restricted eigenvalue condition fails.

Example 2: asymptotic confidence sequences with arbitrary spending↩︎

Group sequential designs, in which interim analyses are performed at successive stopping times with the option to terminate early, are standard practice in clinical trials and are employed by agencies including the U.S. Food and Drug Administration [9][11]. Such designs are remarkably flexible: the trialist may specify an alpha-spending policy \(\alpha(t)\) governing how the type-I error budget is allocated across time, and may use a multivariate test statistica joint test across outcomes, dose levels, or pre-specified subgroupswithin a common inferential framework.

Recent results on asymptotic confidence sequences study the high-frequency and infinite-horizon limit of this construction. In the formulation of [12], the monitored statistic converges to a Brownian path tested against a time-dependent boundary \(\psi(t)\); specifying \(\alpha(t)\) is equivalent to specifying the desired null law of the first crossing time for \(\psi(t)\). These constructions currently emphasize analytically tractable, one-dimensional boundaries.3

A natural alternative is simulation-based calibration: draw the limiting Gaussian random walk, compute the boundary-crossing statistic, and use its empirical \((1-\alpha)\)-quantile [14]. Doing so requires approximating the distribution of a degenerate Gaussian supremum of the form \[M_{A,T,\psi} \mathrel{\vcenter{:}}= \sup_{t \in T \atop a \in A} \psi(t)\left\langle a,\,B_t\right\rangle\] where \(B_t\) is a \(d\)-dimensional Brownian path, \(T\) is a fine grid of monitoring times, and \(A \subset \mathbb{R}^d\) is e.g. the unit sphere. For standard boundaries, the variance is proportional to \(t\psi(t)^2\) and it is degenerate at initial times or in the infinite horizon. Our results, via Corollary 1, support calibration of critical values as long as the largest pointwise variance is nontrivial. In principle, this can allow arbitrary implicit spending rules and multivariate outcomes, suggesting a practical bridge between FDA-style group sequential monitoring and asymptotic anytime-valid inference.

Example 3: confidence bands under latent structure↩︎

Many modern datasets concentrate on low-complexity subsets of a much higher-dimensional ambient space, and a growing literature studies estimators that can adapt to the latent geometry of the data [15][18]. Variance degeneracy is a natural consequence. As a canonical example, consider the ambient kernel density estimator on \(\mathbb{R}^D\) with kernel \(k\) and bandwidth \(h\), given by \(\hat{p}_h(x) = (n h^D)^{-1}\sum_i k\!\left((x - X_i)/h\right)\), whose stochastic fluctuations were analyzed by [15] under the assumption that the underlying distribution \(P\) has latent, low-dimensional structure. Their analysis shows that the pointwise variance \(\sigma^2(x)\) of \(\hat{p}_h(x)\) is of order at most \[n^{-1} h^{-2D} P(B(x,h)),\] where \(B(x,h)\) denotes the Euclidean ball of radius \(h\) centered at \(x\). When the distribution satisfies the natural local scaling relation \[P(B(x,r)) \asymp r^{d(x)},\] as occurs when \(P\) has local “volume dimension” \(d(x) \ll D\), the variance scales as \(n^{-1} h^{-(2D-d(x))}.\) Consequently, fluctuations are much smaller in regions of low volume dimension and vanish entirely away from the effective support. More broadly, modern statistical problems increasingly involve structured or non-standard data for which degeneracy of the approximating process is unavoidable [19]. This creates a mismatch with confidence band constructions motivated by classical non-parametric estimation, where it is natural for pointwise variances to be standardized [20]. Standardization produces conditions such as \(\sup_x|\hat{\sigma}(x)/\sigma(x) - 1| = o_p(1)\), which are prohibitive when \(D\) is large and \(\sigma(x)\) can be tiny. In contrast, Corollary 1 can be applied to the discretized function class without rescaling, avoiding the need to learn \(\sigma(x)\) uniformly at multiplicative scale.

1.2 Related work↩︎

Densities of maxima of Gaussian vectors are classical objects of study, and [20] demonstrate their central role in high-dimensional and non-parametric statistics. They showed that controlling the density of the Gaussian maximumtogether with an appropriate high-dimensional couplingsuffices to establish distributional approximation in several problems of practical and theoretical interest. A general anti-concentration bound was obtained for centered vectors by [21], yielding a density bound of order \((\min_i \sigma_i)^{-1}\sqrt{\log p}\), which is proved to be essentially optimal when all marginal variances are of constant order. Further work extended this result to the non-centered case [22], [23] via Nazarov’s inequality on the Gaussian surface area of polytopes [24], [25].

Subsequent work has sought to relax the dependence on \(\min_i \sigma_i\). [4], in establishing improved bootstrap approximations under weak variance decay, also obtain improved density bounds by separating the contributions of large and small variances. [26], who provide novel non-Gaussian bootstrap couplings, further improve anti-concentration estimates in certain regimes including for non-centered Gaussian vectors (see also [22]). More recently, [27] studies general order statistics and relates the problem to control of the variance. With the exception of [4], who exploit structured variance decay and control of correlations, dependence on the minimum variance persists in each of these works.

This paper eliminates the dependence on \(\min_i\sigma_i\) for quantiles above \(\frac{2}{3}\), and provides uniform anti-concentration bounds of optimal order which remain valid under degeneracy. Our approach exploits the geometry of centered maximaspecifically, that facets of the max polytope corresponding to small variance coordinates are moved far from the origin by standardization.

Organization↩︎

The remainder of the paper is organized as follows. Section 2 contains the main technical density bounds for Gaussian maxima: a non-uniform bound that requires essentially no condition on the covariance, location-uniform bounds and variance control under mild conditions on \(\mathbb{E}[\max_i Z_i]\), a characterization of the density under weak variance decay, and a bound at extremal quantiles. Section 3 works out the implications for Gaussian and multiplier-bootstrap approximation of maxima of high-dimensional sums under weakened variance restrictions. Section 5 gives the proofs of the main density bounds; Appendix 6 collects auxiliary results, and Appendix 7 gives the proofs of the implications.

2 Anti-concentration of degenerate maxima↩︎

We begin with a density bound for arbitrary finite-dimensional Gaussian vectors. The novelty of the bound is that it is non-uniform in the location \(t\), but uniform in the covariance \(\Sigma \succeq 0\). We subsequently show how to “lift” the inequality to bounds which are uniform over \(t \in \mathbb{R}\), or over relevant quantiles.

2.1 A non-uniform, scale-free density bound↩︎

Proposition 2. Let \(Z \sim N(0,\Sigma)\) be a centered Gaussian vector in \(\mathbb{R}^p\) for \(p \ge 3\), with \(\sigma_i^2 = \mathbb{E}[Z_i^2]\), and write \(M = \max_{1 \le i \le p} Z_i\). Define the effective dimension \(p^*(t) \le p\) according to \[p^*(t) = \#\Big\{1 \le i \le p : \sigma_i \ge t(2 \log p + 2\log\log p)^{-\frac{1}{2}}\Big\} \vee2.\] Then, for any \(t > 0\), \[\lim_{\delta \downarrow 0} \frac{1}{\delta} \mathbb{P}\left\{ t \le M < t+ \delta \right\} \le 4 \min\left\{ \frac{\log p^*(t) + \frac{1}{4}}{t}, \frac{ \log p}{t}\right\}.\label{eq:general-finite-bound}\tag{3}\] In particular, with \(\sigma_{\max} = \max_{1 \le i \le p} \sigma_i\), \[\lim_{\delta \downarrow 0} \frac{1}{\delta} \mathbb{P}\left\{ t \le \sigma_{\max}^{-1}\max_{1 \le i\le p}Z_i < t+ \delta \right\} \le \frac{4 \log p}{t},\label{eq:max-normalized-finite-bound}\tag{4}\] so the probability density of \(M/\sigma_{\max}\) is bounded by \(4 \log p\) at all points \(t \ge 1\).

Proposition 2 follows from modifying Nazarov’s bound on the Gaussian surface area of a polytope [24], [25]. Unlike other applications of Nazarov’s argument to anti-concentration of maxima, however, our modification does not easily extend to non-centered vectors, \(Z\). We emphasize that the bound \((4 \log p) / t\) in 3 provides a fixed envelope that does not depend on the covariance \(\Sigma\) at all. Indeed, 4 follows trivially from applying 3 to the vector \(Z/\sigma_{\max}\).

The lack of dependence upon the covariance matrix \(\Sigma\) may appear surprising. For example, one could choose a large number \(N \gg 1\) and apply the bound 3 to \(Z/N\), thus obtaining the same upper bound for the densities of both \(Z\) and \(Z/N\) at the point \(t\). However, this “free lunch” is counteracted by the non-uniform nature of the bound: interesting quantiles of \(\max_i Z_i/N\) will be very close to \(0\), and the bound scales as \(t^{-1}\), so the bound increases by the necessary factor \(N\) at the corresponding quantiles \(t/N\) of the law of \(Z/N\).

Still, for \(t \ge 1\), we obtain an \(O(\log p)\) anti-concentration bound with no restriction on \(\Sigma\), and the bound improves for large values of \(t\). Moreover, the normalization in 4 is enough to ensure that many interesting quantiles of \(\max_i Z_i / \sigma_{\max}\) are at least \(1\); this follows by comparison to the standard normal law, since quantiles of the maximum are larger than that of each coordinate. Thus, our result implies the density of \(\max_i Z_i\) is well behaved at all quantiles \(q > 2/3\), provided \(Z\) satisfies the simple condition \(\max_i \sigma_i \gtrsim 1\).

2.2 Uniform anti-concentration under degeneracy↩︎

In fact, the same argument underlying Proposition 2 also implies uniform anti-concentration boundsones which cover all \(t \in \mathbb{R}\). The following Corollary 2 differs from standard anti-concentration bounds in that it depends upon the global size of \(Z\) via \(\mu\) and \(\sigma_{\mathrm{max}}\), as opposed to local size via \(\min_i \sigma_i\). This makes it applicable to degenerate vectors and processes.

Corollary 2. Let \(Z \sim N(0,\Sigma)\) be a non-trivial, centered Gaussian vector in \(\mathbb{R}^p\) with \(p \ge 3\), let \(M = \max_{1 \le i \le p} Z_i\), and write \(\mu = \mathbb{E}[M]\) and \(\sigma_{\mathrm{max}}^2 = \max_{1 \le i \le p}\mathbb{E}[Z_i^2]\). Suppose that \(\mu > 0\). Then, for every \(t \in \mathbb{R}\) and every \(\varepsilon \ge \sigma_{\mathrm{max}}\exp\{-\mu^2/(8\sigma_{\mathrm{max}}^2)\}\), \[\label{eq:finite-window-eps} \mathbb{P}\{t \le M \le t+\varepsilon\} \;\lesssim\; \frac{\varepsilon\log p}{\mu}.\tag{5}\] Moreover, writing \(M^* = \max_{1 \le i \le p }|Z_i|\) and \(\mu^* = \mathbb{E}[M^*]\), we have for all \(0< \varepsilon < \mu^*/\log(2p)\), \[\label{eq:finite-window-unsigned} \mathbb{P}\{t \le M^* \le t+\varepsilon\} \;\lesssim\; \left(\frac{\varepsilon\log (2p)}{\mu^*}\right)^{\frac{1}{2}}.\tag{6}\]

Proof. Fix \(t \in \mathbb{R}\) and \(\varepsilon > 0\). For any \(a \in (0,\mu)\), splitting the window \([t,t+\varepsilon]\) at \(a\) gives \[\label{eq:finite-window-split} \mathbb{P}\{t \le M \le t+\varepsilon\} \;\le\; \mathbb{P}\{M \le a\} + \int_{t \mathrel{\vee} a}^{t + \varepsilon} f(s)\,ds \le \mathbb{P}\{M \le a\} + \frac{4\varepsilon\log\,p}{a},\tag{7}\] by our density bound 3 . For 5 , Gaussian concentration bounds the first term as \[\label{eq:borell-tis} \mathbb{P}\{M \le a\} \;\le\; \exp\left\{-\frac{(\mu-a)^2}{2\sigma_{\mathrm{max}}^2}\right\}.\tag{8}\] Take \(a = \mu/2\); given our lower restriction on \(\varepsilon\), the concentration gives \(\mathbb{P}\{M \le a\} \le \varepsilon/\sigma_{\mathrm{max}}\). Summing the two terms in 7 and using \(\mu \le \sigma_{\mathrm{max}}\sqrt{2\log p}\) yields 5 .

For 6 , note that \(M^*\) is the signed maximum of a Gaussian vector in dimension \(2p\), so that 7 holds with \(M\) and \(\log p\) replaced by \(M^*\) and \(\log(2p)\). In place of Gaussian concentration we use the following small-ball estimate, a consequence of the S-inequality of [28]; it relies on symmetry of the \(\ell^\infty\) ball to characterize the events \(\{M^* \le t\}\).

Lemma 1. There exists a universal constant \(C\) such that for all \(r \in (0,1)\), \[\mathbb{P}\{M^* \le r\mu^*\} \le C r.\]

We give details of Lemma 1 in Appendix 6.1. Using it in 7 with \(a = r\mu^*\) and optimizing over \(r \in (0,1)\) gives 6 . ◻

Remark 1 (Tightness). The dependence on \(p\) in 5 is optimal. For \(Z\sim N(0,I_p)\) one has \(\mu = \mathbb{E}[M] \asymp \sqrt{\log p}\) and \(\sup_t f(t) \asymp \sqrt{\log p}\); since the bound is of order \(\log p/\mu \asymp \sqrt{\log p}\), and for such non-degenerate vectors the finite-window anti-concentration in 5 is of the same order as \(\sup_t f\), the bound is attained. Rescaling \(Z\) by the factor \((\log p)^{-\frac{1}{2}}\) gives \(\mathbb{E}[M] \asymp 1\) and \(\sup_t f(t) \asymp \log p\), matching the resulting order \(\log p\).

Remark 2 (Comparisons). [21] give a bound of order \((\min_i \sigma_i)^{-1}\mu\) and prove its optimality in the case where \(\sigma_i \asymp 1\). The above shows that \(\min_i\sigma_i \gtrsim 1\) can be dropped whenever \(\mu/\sigma_{\mathrm{max}} \gtrsim \sqrt{\log p}\) while preserving the same (optimal) order of the bound. The improved anti-concentration bound of [26] is of order \((\min_i \sigma_i)^{-1}\) in its most favorable cases (roughly, when \(\sigma_i \gtrsim \sigma_p\sqrt{\log\{p-i\}}\) for each \(1 \le i < p\)). In this case, the bound 5 may be worse by a factor \(\sqrt{\log p}\) if coordinates of \(Z\) are highly correlated. On the other hand, as \(\min_i \sigma_i \downarrow 0\), our Corollary 2 can be sharper by an arbitrarily large factor. A major difference is that [26] also consider non-centered vectors (see also [22]). Another difference is the implicit condition \(\mu \gg \sigma_{\max}\), which is unavoidable for degenerate vectors due to possible mass at \(0\), and which is eliminated by 6 for unsigned maxima.

2.3 Bounds on the variance↩︎

Proposition 2 also implies bounds on the variance of the maximum which depend only upon its expectation and the dimension, eliminating further dependence on \(\Sigma\).

Proposition 3. Let \(Z \sim N(0,\Sigma)\) be a centered Gaussian vector in \(\mathbb{R}^p\) with \(p \ge 3\), and let \(M = \max_{1 \le i \le p} Z_i\), and let \(\mu= \mathbb{E}[M]\) denote its expectation. Then, if \(\mu > 0\), then \[\sqrt{\mathrm{Var}(M)} \ge \frac{\mu}{15 \log p}.\label{eq:variance-bound}\tag{9}\]

Proposition 3 complements [29], who show the bound \(\sqrt{\mathrm{Var}(M)} \gtrsim 1/\mu\) when all coordinate variances are of equal order (see [21] for a related result). Interestingly, here the role of \(\mu\) is reversed: whereas in the homogeneous case, \(\mu\) must be large in order for \(M\) to be very concentrated, here we find that large \(\mu\) limits concentration of \(M\): due to the density envelope of Proposition 2, large spikes can only lie close to \(0\), forcing a variance contribution of order \(\mu^2\). The two bounds agree when \(\mu \asymp \sqrt{\log p}\).

Proof of Proposition 3. With our scale-free bound 3 in hand, lower-bounding the variance becomes a straightforward variational problem that can be solved exactly.

Lemma 2. Let \(X\) be a random variable with a density \(g\) on \((0, \infty)\), such that (i) for all \(t > 0\), \(g(t) \le B/t\), and (ii) \(\mathbb{E}[X] > 0\). Recall the definition \(\coth(x) = (e^{2x} + 1)/(e^{2x}-1)\). Then \[\mathrm{Var}(X) \ge \mathbb{E}[X]^2\left\{\frac{1}{2B}\coth\left(\frac{1}{2B}\right) - 1 \right\} \ge \frac{\mathbb{E}[X]^2}{13 B^2},\] where the first inequality is tight, and the last inequality holds whenever \(B \ge 1\).

Lemma 2 follows from a short tilting argument (Lemma 6), after ruling out mass on \((-\infty,0]\). One can show that the minimum variance is attained by a density of the form \(g_*(t) = (B/t)\mathbb{1}\{a \le t \le b\}\), and then optimize over \(b > a > 0\) subject to \(\int g_*(t)\,dt = 1\). The full proof of the lemma is postponed to Section 5.2. Proposition 3 is then proved by applying Lemma 2 with \(X = M\) and \(B = 4\log(p)\). ◻

Variance-density bounds↩︎

One may consider lifting Proposition 3 into a uniform density bound using a statement of the following form (cf. [27]).

Assumption 1. It holds that \(\sup_{t \in \mathbb{R}} f(t) \le \sqrt{\kappa / \mathrm{Var}(M)}\).

A variance-density principle of this form was recently proposed [27]. However, as stated, it requires additional qualification for any (even dimension-dependent) \(\kappa\): take \(Z \sim N(0, \Sigma)\) with \(\Sigma = \operatorname{diag}(\delta,1, \ldots, 1)\) in any dimension \(p \ge 2\). As \(\delta \downarrow 0\), both mean and variance of \(\max_i Z_i\) converge to positive constants, while the maximum density diverges as \(\delta^{-\frac{1}{2}}\).4

Nevertheless, we record some interesting implications of Assumption 1 below. In particular, under Assumption 1, we recover a general principle for lifting non-uniform bounds into uniform bounds by evaluating them at \(\mu\).

Lemma 3. Suppose that Assumption 1 holds. Then, in the setting of Proposition 3, \[\sup_{t \in \mathbb{R}} f(t) \le \frac{15\sqrt{\kappa}\log p}{\mu}.\] Moreover, under Assumption 1, the density of the maximum of any non-trivial, not necessarily centered Gaussian vector, satisfies the relation \[\sup_{t\in \mathbb{R}} f(t) \lesssim \kappa^2 f(\mu).\label{eq:uniform-lift}\tag{10}\]

Proof. The first claim is an immediate consequence of Proposition 3 and Assumption 1. For the second, we combine Assumption 1 with Ehrhard-concavity of the CDF of \(M\) (Lemma 5 below) and a variational argument; a full proof is given in Appendix 6. ◻

Remark 3 (Unsigned maxima). For unsigned maxima, \(M^* = \max_i|Z_i|\), one may take \(Z \sim N(0,\Sigma)\) for \(\Sigma = \operatorname{diag}\{4, 1/(\log p), \ldots, 1/(\log p)\}\) to obtain \(\sup_t f(t) \asymp \log p\) with \(\mathrm{Var}(M^*)\) of constant order. This leaves open the question of whether a uniform density bound of order \(\log p\), and a variance-density principle with \(\kappa\) depending on the dimension, may hold for unsigned maxima.

2.4 Improved bounds under pointwise decay↩︎

We now consider the behavior of our bounds under further structural assumptions, particularly the weak variance decay considered by [4]. Without loss of generality, suppose that the standard deviations \(\{\sigma_q\}_{1 \le q \le p}\) are given in decreasing order. The bounds of Section 2.1 admit the following extension, implying a dimension-independent bound when \(\sigma_q = o(1/\sqrt{\log q})\).

Proposition 4. For \(u > 0\), let \(S_u = \{q: \sigma_q \ge u\}\) be the set of coordinate indices whose standard deviations exceed \(u\) and \(S_u^c\) be the complementary set. We then have \[\label{eq:weak} f(t) \lesssim \inf_{u > 0}\left\{\frac{\sqrt{\log (\#S_u \vee 2)}}{u} + \sum_{q \in S_u^c} \frac{1}{\sigma_q} \exp\left(\frac{-t^2}{2\sigma_q^2}\right)\right\}.\tag{11}\]

Proof. See Section ¿sec:sec:proof-prop1?. ◻

Proposition 4 recovers the following sharp boundary in the behavior of the density under weak variance decay. The result is based upon convergence of the sum appearing in 11 .

Example 1 (Boundedness under weak variance decay). Suppose that for some \(\alpha > 0\), the vector \(Z \in \mathbb{R}^p\) satisfies (after a possible rearrangement of the coordinates): \[\sigma_k \le \frac{t}{\sqrt{2\alpha\log k}} \label{eq:very-weak-variance-decay}.\tag{12}\] If \(\alpha > 1\), then \(f(t)\) is bounded by an envelope which depends only upon \(\alpha\) (and not \(p\)).

On the other hand, for any \(p\) and any \(t > 0\), we may take \(Z\) to be a Gaussian with variance \(t^2(2\log p)^{-1}I_p\). This construction satisfies 12 with the precise constant \(\alpha = 1\), and it has \(f(t) \asymp \sqrt{\log p}/t\) for each \(p\), so that the density is unbounded as \(p\) grows. In particular, this construction has \(\mu \sim 1\) as \(p \uparrow \infty\), and \(M - \mu = O_p(\log^{-1}p)\) [30].

Thus, 11 captures the threshold for uniform boundedness of the density under this variance profile, up to the sharp constant \(\alpha = 1\) in the regime \(\sigma_k \asymp (2\alpha\log k)^{-\frac{1}{2}}\).

2.5 Density at extremal quantiles↩︎

Arguments based on concavity of the Gaussian law can provide a more complete picture of the density at extremal points. As usual, we let \(f\) denote the density of \(M = \max_i Z_i\) on \((0,\infty)\) and we let \(F(t) = \int_{-\infty}^t f(s)\,ds\) be its cumulative distribution function.

Lemma 4. Let \(Z \sim N(0,\Sigma)\) be a centered Gaussian vector in \(\mathbb{R}^p\), let \(M = \max_{1 \le i \le p} Z_i\) with density \(f\) on \((0,\infty)\), and let \(m\) be a median of \(M\). Then, whenever \(m > 0\), \(f\) is non-increasing on \([m,\infty)\), and for all \(h \ge 0\), \[\label{eq:extremal-density-bound} f(m + h\sigma_{\mathrm{max}}) \le \sqrt{2\pi}\, f(m)\, \varphi(h) = f(m)\, e^{-h^2/2},\tag{13}\] where \(\varphi\) is the standard Gaussian density. Thus, the non-uniform bound 3 gives the envelope \[\frac{4\log p}{m}\, e^{-h^2/2} \quad (h \ge 0).\]

Lemma 4 provides a differential analog of the Gaussian concentration inequality, which states that \(1 - F(m + h\sigma_{\mathrm{max}}) \le 1 - \Phi(h)\). It follows from the celebrated inequality of [31], which implies the following.

Lemma 5 (Ehrhard concavity). The map \(G(t) \!=\! \Phi^{-1}\{F(t)\}\) is concave on the support of \(F\).

[32] showed, moreover, that \(G(t)\) is concave if and only if \(F\) is the cumulative distribution function of the supremum of a separable Gaussian process.

Proof of Lemma 4. Let \(G(s) = \Phi^{-1}\{F(s)\}\), so that \(F(s) = \Phi(G(s))\). By Lemma 5, \(G\) is concave, hence differentiable almost everywhere with \(G'\) non-increasing; at these points, \(f(s) = G'(s)\,\varphi(G(s)).\) We restrict to points of differentiability and extend to all \(h \ge 0\) by continuity. This immediately gives \[\label{eq:G-prime-at-median} G'(m) = f(m)/\varphi(0) = \sqrt{2\pi}\,f(m).\tag{14}\] Gaussian concentration gives the slope-one lower bound \[G(m + h\sigma_{\mathrm{max}}) \;\ge\; h, \qquad h \ge 0.\] Combining this with \(G' \downarrow\) and \(G(m+h\sigma_{\mathrm{max}}) \ge h \ge 0\), which implies \(\varphi(G(m + h\sigma_{\mathrm{max}})) \le \varphi(h)\), we obtain \[f(m + h\sigma_{\mathrm{max}}) \;=\; G'(m + h\sigma_{\mathrm{max}})\,\varphi(G(m + h\sigma_{\mathrm{max}})) \;\le\; G'(m)\,\varphi(h) \;=\; \sqrt{2\pi}\, f(m)\,\varphi(h),\] which is 13 . For the envelope, note that if \(m > 0\) the non-uniform bound 3 applies at \(t = m\) to give \(f(m) \le 4\log p/m\); combined with the monotonicity of \(f\) on \([m,\infty)\) established above, this yields \(\sup_{t \ge m} f(t) = f(m) \le 4\log p/m\) and the displayed Gaussian envelope. ◻

3 Some implications↩︎

In this section, we work out the implications of the bounds of Section 2 for Gaussian and bootstrap approximation of maxima of high-dimensional sums. The novelty is that these approximations place weaker restrictions on the pointwise variances, and thus they remain applicable under degeneracy.

We rely on the high-dimensional couplings of [5], which build on the seminal Gaussian and bootstrap approximation bounds of [23] using a randomized Lindeberg construction. The coupling is the substantive ingredient; our density bounds enter only as a replacement for Koike’s application of Nazarov’s surface area inequality, and the only work is subsequent optimization of the bound. In other words, the corollaries below follow by combining the two results with little additional effort. We focus on this particular coupling because its proof needs only anti-concentration of centered Gaussian maximathe regime in which our results apply. Subsequent improvements to the high-dimensional CLT [33] achieve tighter dimension dependence but rely internally on anti-concentration of non-centered Gaussian vectors, which we do not address.

Throughout this section, \(X_1, \ldots, X_n\) are independent, centered random vectors in \(\mathbb{R}^p\) whose entries satisfy the sub-exponential bound \(\|X_{ij}\|_{\psi_1} \le K\) for some constant \(K \ge 1\). We write \(S_n = \frac{1}{\sqrt n}\sum_{i=1}^n X_i\) for the normalized sum and \(\Sigma = \mathbb{E}[S_n S_n^\top]\) for its covariance, with largest marginal variance \[\label{eq:nontrivial-var} \sigma_{\mathrm{max}}^2 = \max_{1 \le j \le p} \Sigma_{jj} = \max_{1 \le j \le p} \frac{1}{n}\sum_{i=1}^n \mathbb{E}[X_{ij}^2].\tag{15}\] We let \(Z \sim N(0,\Sigma)\) denote the Gaussian analog of \(S_n\), and \(t_q^\Sigma\) the \(q^{\mathrm{th}}\) quantile of \(\max_{1 \le j \le p} Z_j\). For the bootstrap results we additionally fix multipliers \((w_i)_{i=1}^n\), i.i.d.and independent of \(X = (X_1, \ldots, X_n)\), with \(\mathbb{E}[w_i] = 0\), \(\mathbb{E}[w_i^2] = 1\), and \(|w_i| \le b\) almost surely for some constant \(b \ge 1\), and form the multiplier (wild) bootstrap statistic \[S_n^{\mathrm{WB}} = \frac{1}{\sqrt n}\sum_{i=1}^n w_i (X_i - \bar X_n), \qquad \bar X_n = \frac{1}{n}\sum_{i=1}^n X_i.\]

We additionally impose the following mild large-sample condition, which ensures that the lower-order remainder term in the coupling of [5] is dominated.

Assumption 2. The quantities above satisfy \[b\,K^2(\log p)^{1/2}(\log n)^2 \;\le\; \sqrt n,\] with \(b = 1\) understood for the Gaussian results of Section 3.1 (which involve no multipliers).

Assumption 2 can be dropped whenever \(K, b = O(1)\), as a non-trivial rate already requires \(\log^7 p \lesssim n\), and then \((\log p)^{1/2}(\log n)^2 = o(\sqrt n)\). It plays no role beyond ensuring the second remainder term of the coupling is of lower order; see the proof of Corollaries 36 in Appendix 7.

3.1 High-dimensional CLT for degenerate vectors↩︎

The utility of the approximations below is their immediate ability to handle vectors \(X_i\), which may have many coordinates with very small variance. In fact, the main requirement is that \(\sigma_{\mathrm{max}}^2 = \max_{j} \frac{1}{n}\sum_{i=1}^n \mathbb{E}X_{ij}^2\) is not too small, which is enforced by a simple, global rescaling; note that 16 is scale-invariant.

Corollary 3 (Non-uniform CLT). Under Assumption 2, we have \[\label{eq:nonuniform-clt} \sup_{q \ge \frac{2}{3}} \left| \mathbb{P}\left\{\max_{1 \le j \le p} Z_j \le t_{q}^\Sigma\right\} - \mathbb{P}\left\{\max_{1 \le j \le p} \frac{1}{\sqrt n} \sum_{i=1}^{n}X_{ij}\le t_{q}^\Sigma \right\}\right| \lesssim {\sigma_{\mathrm{max}}^{-2/3}}\left(\frac{K^4\log^7p}{n}\right)^{\frac{1}{6}} .\tag{16}\]

Our results also imply a high-dimensional CLT in Kolmogorov’s distance analogous to [23], though restricted to centered maxima, in the regime \(\mu \gg \sigma_{\mathrm{max}}\).

Corollary 4 (Uniform CLT). Put \(\mu = \mathbb{E}[\max_{1 \le j \le p} Z_j] > 0\). Under Assumption 2, we have \[\label{eq:uniform-clt} \sup_{t \in \mathbb{R}} \left| \mathbb{P}\left\{\max_{1 \le j \le p} Z_j \le t\right\} - \mathbb{P}\left\{\max_{1 \le j \le p} \frac{1}{\sqrt n} \sum_{i=1}^{n}X_{ij}\le t \right\}\right| \lesssim e^{\frac{-\mu^2}{8\sigma_{\mathrm{max}}^2}} + \mu^{-2/3}\left(\frac{K^4\log^7p}{n\,}\right)^{\frac{1}{6}}.\tag{17}\]

3.2 Bootstrap approximation↩︎

A parallel result holds for the multiplier (or wild) bootstrap statistic \(S_n^{\mathrm{WB}}\) introduced above, analyzed by [5], whose conditional law given the data \(X = (X_1, \ldots, X_n)\) approximates that of \(\max_{1 \le j \le p} Z_j\).

Corollary 5 (Non-uniform bootstrap). Under Assumption 2, \[\label{eq:nonuniform-bootstrap} \mathbb{E}\!\left[\sup_{q \ge \frac{2}{3}}\, \left|\mathbb{P}\!\left\{\max_{1 \le j \le p}\, (S_n^{\mathrm{WB}})_j \le t_q^\Sigma\,\big|\, X\right\} - \mathbb{P}\!\left\{\max_{1 \le j \le p} Z_j \le t_q^\Sigma\right\}\right|\right] \lesssim \sigma_{\mathrm{max}}^{-2/3}\left(\frac{b^2 K^4\log^7 p}{n}\right)^{\frac{1}{6}}.\tag{18}\]

Corollary 6 (Uniform bootstrap). With \(\mu\) as in Corollary 4, under Assumption 2, \[\label{eq:uniform-bootstrap} \mathbb{E} \sup_{t \in \mathbb{R}}\, \left|\mathbb{P}\!\left\{\max_{1 \le j \le p}\, (S_n^{\mathrm{WB}})_j \le t \,\big|\, X\right\} - \mathbb{P}\!\left\{\max_{1 \le j \le p} Z_j \le t\right\}\right| \lesssim e^{\frac{-\mu^2}{8\sigma_{\mathrm{max}}^2}}+ \mu^{-2/3}\left(\frac{b^2 K^4\log^7 p}{n}\right)^{\frac{1}{6}}.\tag{19}\]

3.3 Approximation of unsigned maxima↩︎

The term \(e^{-\mu^2/8\sigma_{\mathrm{max}}^2}\) in Corollaries 4 and 6 is informative only when \(\mu \gg \sigma_{\mathrm{max}}\). For the unsigned maximum it disappears: the small-ball estimate of Lemma 1 replaces Gaussian concentration and yields the unrestricted bound 6 . In what follows, we write \(\mu^* = \mathbb{E}\{\max_{1 \le j \le p}|Z_j|\}\).

Corollary 7 (Uniform CLT, unsigned maxima). Under Assumption 2, \[\label{eq:uniform-clt-unsigned} \sup_{t \in \mathbb{R}} \left| \mathbb{P}\Big\{\max_{1 \le j \le p}|(S_n)_j| \le t\Big\} - \mathbb{P}\Big\{\max_{1 \le j \le p}|Z_j| \le t\Big\}\right| \;\lesssim\; (\mu^*)^{-2/5}\left(\frac{K^4 \log^7 p}{n}\right)^{\frac{1}{10}}.\tag{20}\]

Corollary 8 (Uniform bootstrap, unsigned maxima). Under Assumption 2, \[\label{eq:uniform-bootstrap-unsigned} \mathbb{E}\sup_{t \in \mathbb{R}} \left| \mathbb{P}\Big\{\max_{1 \le j \le p}|(S_n^{\mathrm{WB}})_j| \le t \,\big|\, X\Big\} - \mathbb{P}\Big\{\max_{1 \le j \le p}|Z_j| \le t\Big\}\right| \;\lesssim\; (\mu^*)^{-2/5}\left(\frac{b^2 K^4 \log^7 p}{n}\right)^{\frac{1}{10}}.\tag{21}\]

Remark 4. Unlike Corollaries 4 and 6, the bounds 2021 apply even when \(\mu^* \asymp \sigma_{\mathrm{max}}\); a single non-degenerate coordinate (implying \(\mu^* > 0\)) suffices. The price is the exponent \(\frac{1}{10}\) in place of \(\frac{1}{6}\), driven entirely by \(t\) near zerofor \(t\) bounded away from \(0\) the density envelope 3 alone applies and the rate exponent \(\frac{1}{6}\) is recovered.

Corollaries 3-8 require no lower bound on the minimum variance \(\min_j \sigma_j\), in contrast to similar results which depend explicitly on \((\min_j \sigma_j)^{-1}\) [5], [20], [23].

4 Conclusion↩︎

This paper established a scale-free density bound for the maximum of a centered Gaussian vector: on \((0,\infty)\) the density satisfies \(f(t) \le 4\log p / t\), with no restriction on the covariance. From this we derived uniform anti-concentration and variance bounds for the maximum, control of its density at extremal quantiles, and Gaussian and bootstrap approximation guarantees for high-dimensional maxima that remain valid under degeneracy.

Two questions are left open. First, the corresponding theory for unsigned maxima \(\max_i |Z_i|\) is unresolved: as noted in Remark 3, the signed counterexamples do not fully transfer, and it is possible that a uniform density bound of order \(\log p\) and a variance–density principle with a dimension-dependent constant can be verified. Second, our density bounds enter the approximation results of Section 3 only as a drop-in replacement for Nazarov’s inequality within the coupling of [5]; it would be of interest to incorporate our anti-concentration results more intrinsically into Gaussian comparison and coupling arguments, as was done by [33] in the non-degenerate case.

5 Proofs of main results↩︎

5.1 The master bound↩︎

We begin by stating a more general result, from which the bounds of Sections 2.1 and 2.4 follow quite easily. Its proof will be given in Section 5.3.

Proposition 5. Let \(Z\) be a centered Gaussian random vector in \(\mathbb{R}^p\) with covariance \(\Sigma\). Let \(\varphi\) be the standard normal density. Then, for any \(u,t > 0\), it holds that \[\label{eq:master-bound} \lim_{\delta \downarrow 0} \frac{1}{\delta}\left( \mathbb{P}\left\{\max_{1 \le i\le p}Z_i \le t + \delta\right\} - \mathbb{P}\left\{\max_{1 \le i\le p}Z_i \le t \right\} \right) \le \left(\frac{u + t}{u^2}\right) + \sum_{0<\sigma_i < u} \frac{1}{\sigma_i} \varphi\left(\frac{t}{\sigma_i}\right).\tag{22}\] In particular, \(M = \max_{1 \le i \le p} Z_i\) admits a density on \((0,\infty)\).

Proof of Proposition 2 (scale-free bound)↩︎

We begin by explaining how Proposition 5 may be used to recover the main results. Note that the function \(x e^{-x^2}\) is decreasing for \(x \ge 1/\sqrt{2}\) as its derivative is \((1-2x^2)e^{-x^2}\). Put \(x = x(\sigma) = t/(\sigma\sqrt{2})\), which is a decreasing function of \(\sigma\). It follows that if \(\sigma \le t\), so that \(x (\sigma) \ge 1/\sqrt{2}\), then \[\begin{align} \frac{1}{\sigma} \exp\left(\frac{-t^2}{2\sigma^2}\right) = \sqrt{2} t^{-1}xe^{-x^2} \end{align}\] is decreasing in \(x\), hence the left-hand side is an increasing function of \(\sigma\). Thus, whenever \(u \le t\) the bound 22 simplifies to \[\begin{align} \lim_{\delta \downarrow 0} \frac{1}{\delta}\left( \mathbb{P}\left\{\max_{1 \le i\le p}Z_i \le t + \delta\right\} - \mathbb{P}\left\{\max_{1 \le i\le p}Z_i \le t \right\} \right) &\le \left(\frac{u + t}{u^2}\right) + \frac{p}{\sqrt{2\pi}}\frac{1}{u} \exp\left(\frac{-t^2}{2u^2}\right) \\ & = \frac{2\log p}{t} + \left(\sqrt{2} + \frac{1}{\sqrt \pi}\right)\frac{\sqrt{\log p}}{t} \le \frac{4\log p}{t}, \end{align}\] where the last line follows from taking \(u = (2\log p)^{-\frac{1}{2}}t\) and simplifying. This proves one of the two claimed bounds.

Next, put \(u' = t(2 \log p + 2\log\log p)^{-\frac{1}{2}}\) and note that for \(\sigma \le u',\) we have \(\sigma^{-1}\varphi(t/\sigma) \le (\sqrt{2\pi}\,p\,t)^{-1}.\) Thus, our upper bound 22 may be analyzed as \[\begin{align} \left(\frac{u + t}{u^2}\right) + \sum_{\sigma_i < u} \frac{1}{\sigma_i} \varphi\left(\frac{t}{\sigma_i}\right) &= \left(\frac{u + t}{u^2}\right) + \sum_{u > \sigma_i > u'} \frac{1}{\sigma_i} \varphi\left(\frac{t}{\sigma_i}\right) + \sum_{u' \ge \sigma_i } \frac{1}{\sigma_i} \varphi\left(\frac{t}{\sigma_i}\right) \\ &\le \left(\frac{u + t}{u^2}\right) + \sum_{u > \sigma_i > u'} \frac{1}{\sigma_i} \varphi\left(\frac{t}{\sigma_i}\right) + \frac{1}{\sqrt{2\pi}t}. \nonumber \end{align}\] Note that optimizing the bound given by the first two summands is analogous to the computation performed in the preceding paragraph, where we now take \(u = \{2\log p^*(t)\}^{-\frac{1}{2}}t\). Here, \(p^*(t)\) is the number of \(\sigma_i\) that exceed \(u'\). This proves the second claim of Proposition 2.

Proof of Proposition 4 (weak variance decay)↩︎

We adapt the argument used to prove Proposition 2. Fix the outer threshold \(u > 0\) appearing in 11 and set \(k := \#S_u \vee 2= \#\{i : \sigma_i \ge u\} \vee2,\) so that every coordinate of \(S_u\) has \(\sigma_i \ge u\). The master bound (Proposition 5) holds at an arbitrary inner threshold \(v > 0\), \[\label{eq:prop3-master} f(t) \le \frac{v + t}{v^2} + \sum_{\sigma_i < v} \frac{1}{\sigma_i}\,\varphi\!\left(\frac{t}{\sigma_i}\right);\tag{23}\] we bound the contribution of the first term and the summands \(\sigma_i \ge u\) in two ways.

First, take \(v = u_* := t/\sqrt{2\log k}\). The first term of 23 satisfies \((u_* + t)/u_*^2 \le 2t/u_*^2 = 4 \log k / t\), since \(u_* \le t\) for \(k \ge 2\). We split the remaining sum at \(u\), the second sum below being empty when \(u \ge u_*\): \[\label{eq:prop3-split} \sum_{\sigma_i < u_*} \frac{1}{\sigma_i}\,\varphi\!\left(\frac{t}{\sigma_i}\right) \;\le\; \sum_{u \le \sigma_i < u_*} \frac{1}{\sigma_i}\,\varphi\!\left(\frac{t}{\sigma_i}\right) \;+\; \sum_{\sigma_i < u} \frac{1}{\sigma_i}\,\varphi\!\left(\frac{t}{\sigma_i}\right).\tag{24}\] The function \(\sigma \mapsto \sigma^{-1}\varphi(t/\sigma)\) is increasing on \((0,t]\), so each of the at most \(k\) summands in the first sum is bounded by \(u^{-1}\varphi(t/u_*) = u^{-1}/(\sqrt{2\pi}\,k)\) by the definition of \(u_*\); hence that sum is at most \(1/(\sqrt{2\pi}\,u)\). Thus the contribution is at most \(4\log k / t + 1/(\sqrt{2\pi}\,u)\).

Second, we may optimize the inner threshold in \(t\). If \(t \le u\sqrt{2\log k}\), take \(v = u\): no coordinate of \(S_u\) enters the sum in 23 , and the first term obeys \((u + t)/u^2 \le (1 + \sqrt{2\log k})/u \lesssim \sqrt{\log k}/u\), using \(k \ge 2\). If instead \(t > u\sqrt{2\log k}\), take \(v = u_* > u\); then \((u_* + t)/u_*^2 \le 4\log k / t < 2\sqrt{2}\,\sqrt{\log k}/u\), while the at most \(k\) leftover coordinates in \([u, u_*)\) contribute, exactly as above, at most \(1/(\sqrt{2\pi}\,u) \le \sqrt{\log k}/u\). Either way the contribution is \(\lesssim \sqrt{\log k}/u\).

Taking the smaller of the two, and recalling \(k = \#S_u \vee 2\), we obtain 11 up to absolute constants. Taking the infimum over \(u > 0\) completes the proof. 0◻

5.2 Proofs from Section 2.3 (variance bounds).↩︎

In this subsection we prove Lemma 2, the variational bound underlying Proposition 3. Lower-bounding \(\mathrm{Var}(M)\) amounts to a variational problem over densities under the pointwise restriction \(f(t) \le B/t\) supplied by Proposition 2 (with \(B = 4\log p\)), which is solved exactly by the tilting argument stated next.

Lemma 6 (Bathtub principle). Let \(I \subseteq \mathbb{R}\) be measurable, \(a: I \to [0,\infty]\), and \(c, w_1, \ldots, w_k : I \to \mathbb{R}\) measurable. Suppose \(f_*\) satisfies \(0 \le f_* \le a\) and, for some multipliers \(\lambda_1, \ldots, \lambda_k \in \mathbb{R}\), the tilted integrand \(P := c - \sum_{m=1}^k \lambda_m w_m\) obeys \[f_*(t) = a(t) \;\text{ where } P(t) < 0, \qquad f_*(t) = 0 \;\text{ where } P(t) > 0 .\] Then \(f_*\) minimizes \(\int_I c\,f\) among all measurable \(f\) with \(0 \le f \le a\) and \(\int_I w_m f = \int_I w_m f_*\) for every \(m\). For \(k = 1\), \(w_1 \equiv 1\), this is the classical bathtub principle [34].

Proof. Let \(f\) be any competitor. Pointwise \(P\,(f - f_*) \ge 0\): where \(P < 0\) we have \(f_* = a \ge f\), and where \(P > 0\) we have \(f_* = 0 \le f\). Integrating and using \(\int_I w_m(f - f_*) = 0\) for each \(m\), \[0 \le \int_I P\,(f - f_*) = \int_I c\,(f - f_*) - \sum_{m=1}^k \lambda_m \int_I w_m (f - f_*) = \int_I c\,(f - f_*),\] so \(\int_I c\,f \ge \int_I c\,f_*\). ◻

Proof of Lemma 2↩︎

Since any feasible distribution can be transported to a feasible distribution on the nonnegative half-line while reducing the cost \(\mathbb{E} X^2\) and preserving the mean, via \(X \mapsto X_+\,\mathbb{E}[X]/\mathbb{E}[X_+]\) with \(X_+ = X\mathbf{1}\{X\ge 0\}\), it suffices to consider the problem over distributions on the nonnegative half-line carrying mass \(q\) on \((0,\infty)\) and an atom of mass \(1-q\) at the origin, which affects neither the mean nor the second moment.

The extremal distribution must place its mass as close as possible to the mean while respecting the upper bound \(f(t)\le B/t\). We apply Lemma 6 on \(I=(0,\infty)\) with cost \(c(t)=t^2\), box \(a(t)=B/t\), and the two constraints \(w_1\equiv 1\) (mass \(q\)) and \(w_2(t)=t\) (mean \(A\)). The tilt \(P(t)=t^2-\lambda_2 t-\lambda_1\) is an upward parabola, so \(\{P<0\}\) is an interval \((a,b)\) and the minimizer saturates the box there: \[f_*(t)=\frac{B}{t} \mathbf{1}\{a\le t\le b\}.\] The constants \(a,b\) are fixed by the two constraints, \[\int_a^b \frac{B}{t}\,dt=B\log(b/a)=q, \qquad \int_a^b t\,\frac{B}{t}\,dt=B(b-a)=A,\] that is, \[a=\frac{A}{B(e^{q/B}-1)}, \qquad b=a\,e^{q/B}=\frac{Ae^{q/B}}{B(e^{q/B}-1)}.\] The second moment of the extremal distribution is \[\mathbb{E} X_*^2 = \int_a^b t^2\frac{B}{t}\,dt = \frac{B}{2}(b^2-a^2).\] After plugging in the values of \(a\) and \(b\), a further manipulation yields \[\mathbb{E} X_*^2 = \frac{A^2}{2B} \coth\!\left(\frac{q}{2B}\right),\] and hence any such distribution satisfies \[\operatorname{Var}(X) \ge \operatorname{Var}(X_*) = \mathbb{E} X_*^2-A^2 = A^2\left[ \frac{1}{2B}\coth\!\left(\frac{q}{2B}\right)-1 \right] \ge A^2\left[ \frac{1}{2B}\coth\!\left(\frac{1}{2B}\right)-1 \right],\] the final inequality because \(q\le 1\) and \(\coth\) is decreasing. If \(B\ge 1\), the lower bound \(1/(13B^2)\) then follows by monotonicity of \(u^{-2}(u\coth u - 1)\). 0◻

5.3 Proof of Proposition 5 (master bound)↩︎

Let \(Z \in \mathbb{R}^p\) be a Gaussian random vector with covariance \(\Sigma\) of rank \(s \le p\). Then, we may write \(Z = Ag\) for \(g\) a standard Gaussian vector in \(\mathbb{R}^{s}\), where \(A\) is a \(p \times s\) real matrix such that \(AA^*=\Sigma\). Note that the \(i^\text{th}\) row of \(A\), which we denote \(a_i\), satisfies \(\|a_i\|_2^2 = \sigma_i^2\), and we write its normalization as \(\nu_i = a_i/\|a_i\|_2\). We have for any \(\delta \ge 0\), \[\begin{align} \mathbb{P}\left\{\max_{1 \le i\le p}Z_i \le t + \delta\right\} &= \mathbb{P}\left\{\max_{1 \le i\le p} \left\langle a_i,\,g\right\rangle \le t + \delta\right\} \nonumber \\ &=\mathbb{P}\left\{\max_{1 \le i\le p} \left\langle \frac{a_i}{\|a_i\|_2},\,g\right\rangle \le \frac{t}{\|a_i\|_2} + \frac{\delta}{\|a_i\|_2}\right\} \nonumber \\ &= \mathbb{P}\left\{ \max_{1 \le i\le p} \left\langle \nu_i,\,g\right\rangle \le \sigma_i^{-1}t +\sigma_i^{-1}\delta\right\}. \label{eq:standardize-polytope} \end{align}\tag{25}\]

Without loss of generality, we may assume that all \(\sigma_i\) are strictly positive, as the remaining coordinates can be dropped from the event in 25 ; since the final bound increases in the number of inequalities, \(p\), this is harmless. We use the following Lemma 7, mirroring the standard reduction to Nazarov’s inequality in [22]. For completeness, its proof is given in the appendix.

Lemma 7. Let \(K \subset \mathbb{R}^{s}\) denote the polytope defined by the inequalities \(\left\langle \nu_i,\,x\right\rangle \le t/\sigma_i\) for \(1 \le i \le p\), and let \(F_i \subset \partial K\) be the facet contained in the hyperplane \(H_i\) defined by \(\left\langle \nu_i,\,x\right\rangle = t/\sigma_i\). Then \[\label{eq:inhomogeneous-decomp} \lim_{\delta \downarrow 0} \frac{1}{\delta}\left( \mathbb{P}\left\{\max_{1 \le i\le p}Z_i \le t + \delta\right\} - \mathbb{P}\left\{\max_{1 \le i\le p}Z_i \le t \right\} \right) \le \sum_{i=1}^p \frac{1}{\sigma_i} \int_{F_i} \varphi_s(x) \mathscr{H}^{s-1}(dx) ,\tag{26}\] where \(\mathscr{H}^{s-1}\) is the \((s-1)\)-dimensional Hausdorff measure on \(\partial K \subset \mathbb{R}^{s}\).

Proof. See Appendix 6. ◻

The crucial observation in the proof is that the facets \(F_i\) in 26 are far away from the origin precisely when the corresponding \(\sigma_i\) are small. Accounting for this correspondence drastically improves the final bound.

Following [35], for a set \(S \subset \mathbb{R}^s\), we may define \(d(x,S) = \inf_{y \in S} \|x-y\|\) and, for any facet \(F_i\) as defined in Lemma 7, we define its outward normal column as \[N_i = \left\{x + y\nu_i \, \middle|\, x \in \mathrm{relint}(F_i), y > 0 \right\}.\] We also have the following estimates, which are taken from the proof of Nazarov’s inequality given by [35]. Proofs are provided in Appendix 6.

Lemma 8. The sets \(N_i=\left\{x + y\nu_i \, \middle|\, x \in \mathrm{relint}(F_i), y > 0 \right\}\) are disjoint, so \(\sum_{i}\gamma_s(N_i) \le 1\).

Lemma 9. For any facet \(F_i \subset \partial K\) as defined in Lemma 7, we have \[\begin{align} \int_{F_i} \varphi_s(x) \mathscr{H}^{s-1}(dx) \le [1+d(0,H_i)]\gamma_s(N_i)\, \text{ and } \tag{27}\\ \int_{F_i} \varphi_s(x) \mathscr{H}^{s-1}(dx) \le \frac{1}{\sqrt{2\pi}}\exp\left\{-\frac{d(0,H_i)^2}{2}\right\}.\tag{28} \end{align}\]

Finally, we apply the main novel idea in the proof: we show that facets \(F_i\) for which \(\sigma_i\) is small (corresponding to \(i \in S_1\) below) cannot contribute much to the sum on the right hand side of inequality 26 . To do so, we partition \(\sigma_i\) into two complementary subsets \(S_1(u)\) and \(S_2(u)\) defined by: \[\begin{align} S_1(u) &\mathrel{\vcenter{:}}= \left\{i \in [p]: \sigma_i \le u \right\}; \\ S_2(u) &\mathrel{\vcenter{:}}= \left\{i \in [p]: \sigma_i > u \right\}. \end{align}\] Crucially, note that the distance between the hyperplane \(H_i\) and the origin satisfies \[d(0,H_i) =t/{\sigma_i}\] by construction. For \(i \in S_1(u)\), inequality 28 implies that \[\begin{align} \frac{1}{\sigma_i}\int_{F_i} \varphi_s(x) \mathscr{H}^{s-1}(dx) \le \frac{1}{\sigma_i\sqrt{2\pi}}\exp\left\{-\frac{t^2}{2\sigma_i^2}\right\} = \frac{1}{\sigma_i}\varphi\left(\frac{t}{\sigma_i}\right), \end{align}\] so that \[\sum_{i\in S_1(u)}\frac{1}{\sigma_i}\int_{F_i} \varphi_s(x) \mathscr{H}^{s-1}(dx) \le \sum_{i\in S_1(u)}\frac{1}{\sigma_i}\varphi\left(\frac{t}{\sigma_i}\right).\] Meanwhile, for \(i \in S_2(u)\), it follows by 27 that \[\begin{align} \sum_{i\in S_2(u)}\frac{1}{\sigma_i}\int_{F_i} \varphi_s(x) \mathscr{H}^{s-1}(dx) &\le \sum_{i\in S_2(u)} \frac{1}{\sigma_i}[1 + d(0,H_i)]\gamma_s(N_i) \\ &\le \frac{u + t}{u^2} \sum_{i\in S_2(u)} \gamma_s(N_i) \le \frac{u + t}{u^2}. \end{align}\] This completes the proof.

6 Additional proofs of density bounds↩︎

6.1 Proof of Lemma 1↩︎

Corollary 2 uses the following linear small-ball bound, a standard consequence of the Gaussian S-inequality of [28]: among symmetric convex sets of a given Gaussian measure, the symmetric strip is extremal under dilation. The argument applies to any seminorm of a Gaussian vector; we state it for \(M^* = \max_{1 \le i \le p}|Z_i| = \|Z\|_\infty\).

Lemma 10. Let \(Z \sim N(0,\Sigma)\) be non-trivial and put \(M^* = \max_{1 \le i \le p}|Z_i|\) and \(\mu^* = \mathbb{E}[M^*] > 0\). There is a universal constant \(C\) such that \[\mathbb{P}\{M^* \le r\mu^*\} \le C r, \qquad 0 < r \le 1.\]

Proof. Let \(g\) be standard Gaussian with \(Z = \Sigma^{1/2}g\) in distribution (restricting to the support of \(Z\) if \(\Sigma\) is singular), and write \(\gamma\) for the law of \(g\). Let \(m\) be a median of \(M^*\). Then \(K = \{x : \|\Sigma^{1/2}x\|_\infty \le m\}\) is a symmetric convex set with \(\gamma(K) = \mathbb{P}\{M^* \le m\} = \frac{1}{2}\) and \(\mathbb{P}\{M^* \le sm\} = \gamma(sK)\) for \(s \ge 0\). The S-inequality [28] compares the dilates of \(K\) with those of a symmetric strip \(P = \{x : |x_1| \le w\}\) of equal measure \(\gamma(P) = \frac{1}{2}\), that is \(w = \Phi^{-1}(\frac{3}{4})\): \[\gamma(tK) \le \gamma(tP) \;\;(0 \le t \le 1), \qquad \gamma(tK) \ge \gamma(tP) \;\;(t \ge 1).\] For \(0 < t \le 1\) the small-dilation side gives \[\mathbb{P}\{M^* \le tm\} = \gamma(tK) \le \gamma(tP) = \mathbb{P}\{|g_1| \le tw\} \le \sqrt{\tfrac{2}{\pi}}\,w\,t =: C_1 t,\] using that the density of \(|g_1|\) is at most \(\sqrt{2/\pi}\). For \(t \ge 1\) the large-dilation side gives \(\mathbb{P}\{M^* > tm\} \le \mathbb{P}\{|g_1| > tw\}\), so that \[\mu^* = \int_0^\infty \mathbb{P}\{M^* > s\}\,ds = m\int_0^\infty \mathbb{P}\{M^* > tm\}\,dt \le m\Big(1 + \int_1^\infty \mathbb{P}\{|g_1| > tw\}\,dt\Big) =: C_0\,m.\] Hence for \(0 < r \le 1/C_0\), \(\mathbb{P}\{M^* \le r\mu^*\} \le \mathbb{P}\{M^* \le C_0 r\,m\} \le C_1 C_0\,r\), while for \(1/C_0 < r \le 1\) the claim is trivial since \(1 \le C_0 r\). Enlarging the constant completes the proof. ◻

6.2 Proof of Lemma 3↩︎

It remains to establish the lifting relation 10 , the first claim being immediate from Proposition 3. The relation follows from the result below, applied with \(X = M\): hypothesis (i) is Ehrhard’s inequality (Lemma 5), and hypothesis (ii) holds with \(C = \sqrt{\kappa}\) under Assumption 1, since then \(\|f\|_\infty\sqrt{\mathrm{Var}(M)} \le \sqrt{\kappa}\). Neither ingredient requires the Gaussian vector to be centered, so the conclusion also applies to non-centered \(Z\). With \(C = \sqrt{\kappa}\) the result gives \(f(\mu) \gtrsim \kappa^{-2}\|f\|_\infty\), i.e.\(\sup_t f(t) \lesssim \kappa^2 f(\mu)\).

Lemma 11. Let \(X\) have density \(f\), CDF \(F\), mean \(\mu\) and variance \(\sigma^2.\) Assume (i) that \(F\) is Ehrhard-concave, i.e.  \(G(t) = \Phi^{-1}\{F(t)\}\) is concave on the support of \(F\), and (ii) that \(\|f\|_\infty \sigma \le C\) for some \(C > 0\). Then \(f(\mu) \gtrsim C^{-4}\|f\|_\infty.\)

Proof. By translating and rescaling, we may assume that \(\mu = 0\) and \(\|f\|_\infty = 1.\) Then (ii) gives \(\sigma^2 \le C^2\), and it suffices to prove that \(f(0) \gtrsim 1\).

6.2.0.1 Step 1: \(f(t) \le 2f(0)\) for \(t \ge 0\)

We first note that Ehrhard-concavity gives a one-sided median bound. In particular, since \(G\) is increasing and concave, \(h=G^{-1}\) is increasing and convex. Being Ehrhard-concave, \(F\) is log-concave with connected support, hence strictly increasing there; thus \(G\) is a strictly increasing continuous bijection onto its range and \(h=G^{-1}\) is well defined. If \(Z \sim N(0,1)\), then \(h(Z)\) has distribution function \(F\). Therefore, by Jensen’s inequality, \[0=\mathbb{E} X=\mathbb{E} h(Z)\ge h(\mathbb{E} Z)=h(0).\] Since \(F(h(0))=\Phi(0)=1/2\), it follows that \(F(0) \ge \frac{1}{2}.\) Next, since \(F\) is Ehrhard-concave, it is a log-concave function. Thus \({f(t)}/{F(t)}\) is nonincreasing as a function of \(t\). Consequently, for every \(t\ge 0\), \[\frac{f(t)}{F(t)} \le \frac{f(0)}{F(0)} \le 2f(0).\] Since \(F(t)\le 1\), we obtain the one-sided bound \(f(t)\le 2f(0)\) for all \(t \ge 0\).

6.2.0.2 Step 2: \(\mathbb{E}[X\mathbb{1}\{X \ge 0\}] \ge 1/8\)

Since \(\mathbb{E} X=0\), \[D = \int_0^\infty t f(t)\,dt = \int_{-\infty}^0 (-t)f(t)\,dt = \frac{1}{2}\mathbb{E} |X|.\] Since \(\|f\|_\infty\le 1\) and \(\int f(t)\,dt=1\), the bathtub principle (Lemma 6) implies \(\mathbb{E}|X|\) is bounded below by \(1/4\), and \(D \ge 1/8\).

6.2.0.3 Step 3: Minimize variance subject to \(f(t) \le 2 f(0)\) and \(\mathbb{E}[X\mathbb{1}\{X \ge 0\}] \ge 1/8\)

Among all nonnegative functions \(q(x)\) on the nonnegative real line satisfying \[q(x) \le 2f(0), \qquad \int_0^\infty x q(x)\,dx = D,\] another application of Lemma 6 (with cost \(t^2\), box \(2f(0)\), and the single first-moment constraint \(\int_0^\infty t\,q=D\), so that the tilt \(t^2-\lambda t\) is negative on an initial interval) shows that the second moment \(\int_0^\infty t^2 q(t)\,dt\) is minimized, up to universal constants, by filling an initial interval at maximal height. Since \(f\) restricted to \([0,\infty)\) is itself feasible for this programit satisfies \(f(t) \le 2f(0)\) by Step 1 and \(\int_0^\infty t f(t)\,dt = D\)its second moment is at least this minimal value. Therefore \[C^2 \ge \mathrm{Var}(X) \ge \int_0^\infty t^2 f(t)\,dt \gtrsim \frac{D^{3/2}}{\sqrt{f(0)}} \gtrsim \frac{1}{\sqrt{f(0)}}.\] Thus \(f(0) \gtrsim C^{-4},\) as needed. ◻

6.3 Proofs from Section 5.3 (Proposition 5)↩︎

We provide proofs of the various lemmas used in proving Proposition 5. These results are fairly standard, and the proofs follow closely after [35].

6.3.1 Proof of Lemma 7↩︎

We begin by restating Lemma 7 to keep track of its dependence on the parameter \(t\in \mathbb{R}\).

Lemma 12. For \(t > 0\), let \(K^t \subset \mathbb{R}^s\) denote the polytope defined by the inequalities \(\left\langle \nu_i,\,x\right\rangle \le t/\sigma_i\) for \(1 \le i \le p\) and let \(F_i^t \subset \partial K^t\) be the (possibly redundant) facet contained in the hyperplane \(H_i^t\) defined by \(\left\langle \nu_i,\,x\right\rangle = t/\sigma_i\). Then \[\label{eq:inhomogeneous-decomp-apx} \lim_{\delta \downarrow 0} \frac{1}{\delta}\left( \mathbb{P}\left\{\max_{1 \le i\le p}Z_i \le t + \delta\right\} - \mathbb{P}\left\{\max_{1 \le i\le p}Z_i \le t \right\} \right) \le \sum_{i=1}^p \frac{1}{\sigma_i} \int_{F_i^t} \varphi_s(x)\, \mathscr{H}^{s-1}(dx) ,\tag{29}\] where \(\mathscr{H}^{s-1}\) denotes the \((s-1)\)-dimensional Hausdorff measure on \(\partial K^t \subset \mathbb{R}^s\).

6.3.2 Proofs of Lemmas 8 and 9↩︎

7 Proofs of approximation results↩︎

The Gaussian and bootstrap approximation results stated in Section 3 follow by combining the density bounds of Section 2 with the coupling results of [5], in the notation fixed at the start of Section 3.

7.1 One-sided couplings↩︎

The proofs rely on the following one-sided Prokhorov-type couplings of [5], which we restate for convenience. The first compares \(S_n\) to its Gaussian analog \(Z \sim N(0,\Sigma)\).

Lemma 13 (Lemma 5.7 of [5]). Assume that there exists a constant \(K \ge 1\) such that for all \(1 \le i \le n\) and \(1 \le j \le p\), \(\|X_{ij}\|_{\psi_1} \le K.\) Then there exists a universal constant \(C>0\) such that for every Borel set \(A \subset \mathbb{R}\), every \(y \in \mathbb{R}^p\), and every \(\varepsilon \ge CK^2 n^{-\frac{1}{2}}(\log n)(\log p),\) we have \[\label{eq:gaussian-onesided} \begin{align} \mathbb{P}\!\left\{\max_{1 \le j \le p}(S_{n})_j - y_j \in A\right\} &\le \mathbb{P}\!\left\{\max_{1 \le j \le p} Z_j - y_j \in A^{6\varepsilon}\right\}\\ &\qquad + C\varepsilon^{-2}\!\left(\frac{K^{2}(\log p)^{3/2}}{\sqrt{n}} + \frac{K^{4}(\log p)^2(\log n)^2}{n}\right), \end{align}\tag{30}\] where \(A^\delta = \cup_{a\in A} [a-\delta,a+\delta]\) denotes the \(\delta\)-enlargement of \(A\).

Remark 5. Koike’s bound is stated in terms of a constant \(B_n \ge 1\) bounding both \(\max_{i,j}\|X_{ij}\|_{\psi_1}\) and \((\max_j n^{-1}\sum_i \mathbb{E}[X_{ij}^4])^{1/2}\). We bound the fourth moment using the \(\psi_1\)-norm, so \(B_n \asymp K^2\); this is the source of the powers \(K^2\) and \(K^4\) in 30 .

The corresponding one-sided coupling for the multiplier bootstrap statistic \(S_n^{\mathrm{WB}}\) was also derived by [5].

Lemma 14 (Bootstrap version of [5], Thm. 3.1). Under the assumptions of Lemma 13, with multipliers \((w_i)_{i=1}^n\) as in Section 3.2 and \(b \ge 1\) such that \(|w_i| \le b\) almost surely, there exist a universal constant \(C' > 0\) and an event \(\Omega_0\) with \(\mathbb{P}(\Omega_0) = 1\) such that the following holds for every \(\varepsilon \ge C'bK^2 n^{-\frac{1}{2}}(\log n)(\log p)\). There is a nonnegative random variable \(\rho(X)\), depending on \(X\) but not on the choice of \(A\) or \(y\), with \[\label{eq:bootstrap-remainder} \mathbb{E}[\rho(X)] \le C'\varepsilon^{-2}\!\left(\frac{b\, K^{2}(\log p)^{3/2}}{\sqrt{n}} + \frac{b^2 K^{4}(\log p)^2(\log n)^2}{n}\right),\tag{31}\] such that, on \(\Omega_0\), for every Borel set \(A \subset \mathbb{R}\) and every \(y \in \mathbb{R}^p\), \[\label{eq:bootstrap-onesided} \mathbb{P}\!\left\{\max_{1 \le j \le p}(S_n^{\mathrm{WB}})_j - y_j \in A\,\middle|\, X\right\} \le \mathbb{P}\!\left\{\max_{1 \le j \le p} Z_j - y_j \in A^{6\varepsilon}\right\} + \rho(X).\tag{32}\]

Proof. The statement is the one-sided, pre-optimization form of [5]; since we use the intermediate coupling rather than its optimized conclusion, we recall the steps from the proof in [5]. We abbreviate the remainders \[\delta_{n,1} = \frac{B_n^2(\log p)^3}{n}, \qquad \delta_{n,2} = \frac{B_n^2(\log p)^2(\log n)^2}{n},\] and use \(B_n \asymp K^2\) (Remark 5); thus \(\sqrt{\delta_{n,1}} \asymp K^2(\log p)^{3/2}/\sqrt n\) and \(\delta_{n,2} \asymp K^4(\log p)^2(\log n)^2/n\) are the two terms in 31 . Set \(\kappa_n = 2K^2\log n\) and, exactly as in the truncation underlying Lemma 13 [5], put \(\tilde{X}_{ij} = X_{ij}\mathbb{1}\{|X_{ij}|\le\kappa_n\} - \mathbb{E}[X_{ij}\mathbb{1}\{|X_{ij}|\le\kappa_n\}]\) and \(\tilde{Y}_i = \tilde{X}_i - \bar{\tilde{X}}_n\), so that \(\max_{i,j}|w_i\tilde{Y}_{ij}|\le 4b\kappa_n\) almost surely.

7.1.0.1 Truncation split.

Writing \(Y_i = X_i - \bar X_n\) (so that \(S_n^{wY} := n^{-1/2}\sum_i w_i Y_i = S_n^{\mathrm{WB}}\)), for every Borel \(A\) and \(y \in \mathbb{R}^p\), \[\label{eq:wb-split} \mathbb{P}\!\left\{\max_j (S_n^{\mathrm{WB}})_j - y_j \in A \,\middle|\, X\right\} \le \mathbb{P}\!\left\{\max_j (S_n^{w\tilde{Y}})_j - y_j \in A^{\varepsilon} \,\middle|\, X\right\} + \mathbb{P}\!\left\{\|S_n^{w(Y-\tilde{Y})}\|_\infty > \varepsilon \,\middle|\, X\right\}.\tag{33}\]

7.1.0.2 Conditional coupling.

Conditionally on \(X\), the vectors \((w_i\tilde{Y}_i)_{i=1}^n\) are independent and centered, with conditional covariance \(n^{-1}\sum_i \mathbb{E}[w_i^2]\,\tilde{Y}_i\tilde{Y}_i^\top = n^{-1}\sum_i \tilde{Y}_i\tilde{Y}_i^\top\) since \(\mathbb{E}[w_i^2]=1\). Applying the coupling of [5]proved through the third-moment-matching randomized Lindeberg interpolation of [5] together with a Stein-kernel estimateto \((w_i\tilde{Y}_i)\) given \(X\), with Gaussian target \(Z \sim N(0,\Sigma)\), \(\Sigma = \mathbb{E}[S_n S_n^\top]\), yields a probability-one event \(\Omega_0\) on which, for every \(\varepsilon\) above the stated threshold, every Borel \(A\) and every \(y\), \[\label{eq:wb-conditional} \mathbb{P}\!\left\{\max_j (S_n^{w\tilde{Y}})_j - y_j \in A^{\varepsilon} \,\middle|\, X\right\} \le \mathbb{P}\!\left\{\max_j Z_j - y_j \in A^{6\varepsilon}\right\} + C\varepsilon^{-2}\!\left(\Delta^*_{n,0}\log p + \Delta^*_{n,1}\sqrt{\tfrac{(\log p)^3}{n}}\right),\tag{34}\] where the \(5\varepsilon\) enlargement of [5] turns \(A^\varepsilon\) into \(A^{6\varepsilon}\), and the conditional remainder quantities are \[\Delta^*_{n,0} = \mathbb{E}\!\left[\max_{1\le j,k\le p}\Big|\tfrac1n\textstyle\sum_i w_i^2\tilde{Y}_{ij}\tilde{Y}_{ik} - \Sigma_{jk}\Big| \,\middle|\, X\right], \qquad \Delta^*_{n,1} = \left(\tfrac1n\,\mathbb{E}\!\left[\max_j \textstyle\sum_i w_i^4\tilde{Y}_{ij}^4 \,\middle|\, X\right]\right)^{1/2}.\] Crucially, \(\Delta^*_{n,0}\) and \(\Delta^*_{n,1}\) depend on \(X\) but not on \((A,y)\), so 34 holds for all \((A,y)\) simultaneously on \(\Omega_0\) with an \(A\)-independent error; this is what later permits taking a supremum over \(t\) inside the conditional probability.

7.1.0.3 Truncation error.

Since \(|w_i|\le b\) and \(\sqrt n(\bar X_n - \bar{\tilde{X}}_n) = S_n^{X-\tilde{X}}\), the second term of 33 is bounded exactly as the truncation estimate underlying Lemma 13 [5]: \[\mathbb{P}\!\left\{\|S_n^{w(Y-\tilde{Y})}\|_\infty > \varepsilon \,\middle|\, X\right\} \le C\varepsilon^{-2} b^2\delta_{n,2}.\] Combining the three displays gives 32 on \(\Omega_0\) with the nonnegative, \(A\)-independent remainder \[\rho(X) = \mathbb{P}\!\left\{\|S_n^{w(Y-\tilde{Y})}\|_\infty > \varepsilon \,\middle|\, X\right\} + C\varepsilon^{-2}\!\left(\Delta^*_{n,0}\log p + \Delta^*_{n,1}\sqrt{\tfrac{(\log p)^3}{n}}\right).\] The threshold \(\varepsilon \ge C'bK^2 n^{-1/2}(\log n)(\log p)\) corresponds to the requirement \(\varepsilon \ge 12\,b\kappa_n(\log p)/\sqrt n\), ensuring the truncation level is compatible with the coupling.

7.1.0.4 Remainder in expectation.

Taking expectations over \(X\) bounds each conditional term. As \(\mathbb{E}[w_i^2]=1\), the conditional mean of \(n^{-1}\sum_i w_i^2\tilde{Y}_{ij}\tilde{Y}_{ik}\) is the truncated sample covariance, whose deviation from \(\Sigma\) is controlled by the fourth-moment concentration of [5] and a maximal inequality for \(\|\bar X_n\|_\infty\); the computation of [5] gives \(\mathbb{E}[\Delta^*_{n,0}]\log p \lesssim b\sqrt{\delta_{n,1}} + b^2\delta_{n,2}\). For \(\Delta^*_{n,1}\), the bound \(\mathbb{E}[w_i^4]\le b^2\mathbb{E}[w_i^2] = b^2\) together with the Jensen and Lyapunov inequalities reduces it to the Gaussian fourth-moment term, giving \(\mathbb{E}[\Delta^*_{n,1}]\sqrt{(\log p)^3/n}\lesssim b\sqrt{\delta_{n,1}} + b^2\delta_{n,2}\). Collecting the three contributions yields \(\mathbb{E}[\rho(X)] \le C'\varepsilon^{-2}(b\sqrt{\delta_{n,1}} + b^2\delta_{n,2})\), which is 31 . ◻

7.2 Proof of Corollaries 38↩︎

The signed corollaries are proved by combining the one-sided couplings of Lemmas 1314 with the density bounds of Section 2; the unsigned corollaries are treated at the end. We write the argument for the Gaussian case; the bootstrap argument is identical after replacing Lemma 13 with Lemma 14 and tracking the additional factor of \(b\).

Step 1: a two-sided Prokhorov bound. Lemma 13 provides a one-sided bound; to derive a two-sided bound on a CDF difference, we apply it twice with \(y = 0\). First, taking \(A = (-\infty, t]\) in 30 , we obtain \[\label{eq:clt-onesided-upper} \mathbb{P}\!\left\{\max_j (S_n)_j \le t\right\} - \mathbb{P}\!\left\{\max_j Z_j \le t\right\} \le \mathbb{P}\!\left\{t < \max_j Z_j \le t + 6\varepsilon\right\} + R_n(\varepsilon),\tag{35}\] where \(R_n(\varepsilon) := C\varepsilon^{-2}\!\left(K^{2}(\log p)^{3/2}/\sqrt{n} + K^{4}(\log p)^2(\log n)^2/n\right)\). Second, taking \(A = (t,\infty)\), we get \[\label{eq:clt-onesided-lower} \begin{align} \mathbb{P}\!\left\{\max_j Z_j \le t - 6\varepsilon\right\} - \mathbb{P}\!\left\{\max_j (S_n)_j \le t\right\} &\le R_n(\varepsilon). \end{align}\tag{36}\] Combining 35 and 36 yields \[\label{eq:two-sided-bound} \left|\mathbb{P}\!\left\{\max_j (S_n)_j \le t\right\} - \mathbb{P}\!\left\{\max_j Z_j \le t\right\}\right| \le \mathbb{P}\!\left\{t - 6\varepsilon < \max_j Z_j \le t + 6\varepsilon\right\} + R_n(\varepsilon).\tag{37}\]

Step 2: anti-concentration of \(\max_j Z_j\). The first term on the right-hand side of 37 is controlled by the density of \(\max_j Z_j\) on \([t-6\varepsilon, t+6\varepsilon]\). Throughout, \(C\) denotes a universal constant whose value may change between occurrences. We distinguish two cases.

7.2.0.1 Case I (Corollaries 3 and 5, \(t = t_q^\Sigma\) for \(q \ge \frac{2}{3}\)).

By stochastic dominance, \(t_q^\Sigma \ge q^M_{\frac{2}{3}} \ge \sigma_{\mathrm{max}}\Phi^{-1}(\frac{2}{3}) \ge \frac{2 \sigma_{\mathrm{max}}}{5}\) for every \(q \ge \frac{2}{3}\). Hence, provided \(\varepsilon \le \frac{\sigma_{\mathrm{max}}}{40}\), the interval \([t-6\varepsilon, t+6\varepsilon]\) lies above \(\frac{\sigma_{\mathrm{max}}}{4}\), and Proposition 2 gives \(f(s) \le C\log p/\sigma_{\mathrm{max}}\) there, so \[\mathbb{P}\!\left\{t-6\varepsilon < \max_j Z_j \le t+6\varepsilon\right\} \le \frac{C\,\varepsilon \log p}{\sigma_{\mathrm{max}}}.\]

7.2.0.2 Case II (Corollaries 4 and 6, \(t \in \mathbb{R}\) arbitrary).

Splitting \([t-6\varepsilon, t+6\varepsilon]\) at the fixed cut \(a = \mu/2\), as in the proof of Corollary 2 but without tying the cut to the window width, the mass below \(a\) is at most \(e^{-\mu^2/8\sigma_{\mathrm{max}}^2}\) by 8 , while above \(a\) the density is at most \(C\log p/\mu\) by 3 . Hence \[\mathbb{P}\!\left\{t-6\varepsilon < \max_j Z_j \le t+6\varepsilon\right\} \le e^{-\mu^2/8\sigma_{\mathrm{max}}^2} + \frac{C\,\varepsilon\log p}{\mu}.\] The additive Gaussian-tail term does not shrink with the coupling scale \(\varepsilon\); it is independent of \(n\) and depends only on the ratio \(\mu/\sigma_{\mathrm{max}}\).

7.2.0.3 Step 3: optimization in \(\varepsilon\).

Write \(\nu = \sigma_{\mathrm{max}}\), \(P = 0\) in Case I, and \(\nu = \mu\), \(P = e^{-\mu^2/8\sigma_{\mathrm{max}}^2}\) in Case II. Combining 37 with Step 2, for every admissible \(\varepsilon\), \[\label{eq:pre-optimization} \left|\mathbb{P}\!\left\{\max_j (S_n)_j \le t\right\} - \mathbb{P}\!\left\{\max_j Z_j \le t\right\}\right| \le P + C\Big(\frac{\varepsilon \log p}{\nu} + \varepsilon^{-2}\Delta_1 + \varepsilon^{-2}\Delta_2\Big),\tag{38}\] where \(\Delta_1 = K^2(\log p)^{3/2}/\sqrt n\) and \(\Delta_2 = K^4(\log p)^2(\log n)^2/n\) are the two terms of \(R_n(\varepsilon)\). Balancing the first two non-penalty terms gives the optimal scale \[\varepsilon^* \asymp \nu^{1/3} K^{2/3} (\log p)^{1/6} n^{-1/6}, \qquad \frac{\varepsilon^*\log p}{\nu} + \varepsilon^{*-2}\Delta_1 \asymp R := \nu^{-2/3}\Big(\frac{K^4 \log^7 p}{n}\Big)^{1/6}.\] Two observations complete the proof. First, the remaining term obeys \[\varepsilon^{*-2}\Delta_2 = R \cdot K^2(\log p)^{1/2}(\log n)^2 n^{-1/2} \le R\] by Assumption 2 (with \(b = 1\)). Second, the optimal \(\varepsilon^*\) may be inadmissibleeither \(\varepsilon^* > \sigma_{\mathrm{max}}/40\) (violating the requirement of Case I), or \(\varepsilon^*\) below the lower threshold \(\asymp\sqrt{\Delta_2}\) of Lemma 13but in either case \(R\) exceeds a universal positive constant (in the first, \(R \ge (\log p)/40\); in the second, Assumption 2 forces \(R \gtrsim 1\)), so the asserted bound holds trivially because the left-hand side is at most \(1\). In all cases, \[\left|\mathbb{P}\!\left\{\max_j (S_n)_j \le t\right\} - \mathbb{P}\!\left\{\max_j Z_j \le t\right\}\right| \lesssim P + R.\] Taking the supremum over \(t\)over \(t \ge q^M_{\frac{2}{3}}\) in Case I (\(P = 0\)), and over \(t \in \mathbb{R}\) in Case IIgives Corollaries 3 and 4; in Case II the displayed rate dominates \(P\) precisely when \(\mu \gtrsim \sigma_{\mathrm{max}}\sqrt{\log n}\).

For the bootstrap corollaries we argue from the coupling of Lemma 14. On the probability-one event \(\Omega_0\), the two-sided bound 37 holds for every \(t\) simultaneously with \(R_n(\varepsilon)\) replaced by the \(A\)-independent remainder \(\rho(X)\). Since \(\rho(X)\) does not depend on \(t\), we may take the supremum over \(t\) on \(\Omega_0\) inside the expectation, after which \(\mathbb{E}[\rho(X)]\) obeys 31 . As the remainder carries an extra factor \(b\) in \(\Delta_1\) and \(b^2\) in \(\Delta_2\), the same optimization yields \(\sigma_{\mathrm{max}}^{-2/3}(b^2 K^4\log^7 p/n)^{1/6}\), and the analogue with \(\mu\), giving Corollaries 5 and 6.

Unsigned maxima (Corollaries 7 and 8). Apply the one-sided couplings to the reflected vectors \((X_i, -X_i)\) (and \(w_i(X_i - \bar X_n)\) likewise) in dimension \(2p\), whose coordinate maxima are \(\max_j|(S_n)_j|\) and \(\max_j|(S_n^{\mathrm{WB}})_j|\) with Gaussian analog \(\max_j|Z_j|\); this only replaces \(\log p\) by \(\log 2p\) in the remainder. In Step 2, in place of the cut at \(\mu/2\) we invoke the unsigned window bound 6 : for \(\varepsilon \lesssim \mu^*/\log p\) and every \(t \in \mathbb{R}\), \[\mathbb{P}\Big\{t - 6\varepsilon < \max_j|Z_j| \le t + 6\varepsilon\Big\} \;\lesssim\; \Big(\frac{\varepsilon \log p}{\mu^*}\Big)^{1/2},\] with no additive penalty. Substituting into 37 and balancing \((\varepsilon\log p/\mu^*)^{1/2} \asymp \varepsilon^{-2}\Delta_1\) gives \[\varepsilon^* \asymp \Big(\frac{\mu^* K^4 \log^2 p}{n}\Big)^{1/5}, \qquad \Big(\frac{\varepsilon^*\log p}{\mu^*}\Big)^{1/2} + \varepsilon^{*-2}\Delta_1 \asymp (\mu^*)^{-2/5}\Big(\frac{K^4 \log^7 p}{n}\Big)^{1/10} =: R^*.\] As before \(\varepsilon^{*-2}\Delta_2 \le R^*\) under Assumption 2; and if \(\varepsilon^*\) is inadmissibleeither \(\varepsilon^* \gtrsim \mu^*/\log p\) or below the lower threshold of Lemma 13then \(R^* \gtrsim 1\) (indeed \(\varepsilon^* \lesssim \mu^*/\log p \iff R^* \lesssim 1\)) and the bound holds trivially. The bootstrap version carries the extra factors \(b\) and \(b^2\) exactly as for Corollary 6. 0◻

References↩︎

[1]
V. Chernozhukov, D. Chetverikov, K. Kato, and Y. Koike, “High-dimensional data bootstrap,” Annual Review of Statistics and Its Application, vol. 10, no. 1, pp. 427–449, 2023.
[2]
A. S. Bandeira, E. Dobriban, D. G. Mixon, and W. F. Sawin, “Certifying the restricted isometry property is hard,” IEEE transactions on information theory, vol. 59, no. 6, pp. 3448–3450, 2013.
[3]
E. Dobriban and J. Fan, “Regularity properties for sparse regression: A tribute to professor Xiru Chen,” Communications in mathematics and statistics, vol. 4, no. 1, pp. 1–19, 2016.
[4]
M. E. Lopes, Z. Lin, and H.-G. Müller, “BOOTSTRAPPING MAX STATISTICS IN HIGH DIMENSIONS,” The Annals of Statistics, vol. 48, no. 2, pp. 1214–1229, 2020.
[5]
Y. Koike, “Notes on the dimension dependence in high-dimensional central limit theorems for hyperrectangles,” Japanese Journal of Statistics and Data Science, vol. 4, pp. 257–297, 2021.
[6]
J. Fan, Q.-M. Shao, and W.-X. Zhou, “Are discoveries spurious? Distributions of maximum spurious correlations and their applications,” The Annals of Statistics, vol. 46, no. 3, p. 989, 2018.
[7]
D. Altshuler, P. Donnelly, et al., “A haplotype map of the human genome,” Nature, vol. 437, no. 7063, pp. 1299–1320, 2005.
[8]
H. Wang, B. J. Lengerich, B. Aragam, and E. P. Xing, “Precision Lasso: Accounting for correlations and linear dependencies in high-dimensional genomic data,” Bioinformatics, vol. 35, no. 7, pp. 1181–1187, 2019.
[9]
S. J. Pocock, “Group sequential methods in the design and analysis of clinical trials,” Biometrika, vol. 64, no. 2, pp. 191–199, 1977, doi: 10.1093/biomet/64.2.191.
[10]
P. C. O’Brien and T. R. Fleming, “A multiple testing procedure for clinical trials,” Biometrics, vol. 35, no. 3, pp. 549–556, 1979, doi: 10.2307/2530245.
[11]
K. K. G. Lan and D. L. DeMets, “Discrete sequential boundaries for clinical trials,” Biometrika, vol. 70, no. 3, pp. 659–663, 1983, doi: 10.1093/biomet/70.3.659.
[12]
I. Waudby-Smith, D. Arbour, R. Sinha, E. H. Kennedy, and A. Ramdas, “Time-uniform central limit theory and asymptotic confidence sequences,” The Annals of Statistics, vol. 52, no. 6, pp. 2613–2640, 2024.
[13]
S. R. Howard, A. Ramdas, J. McAuliffe, and J. Sekhon, “Time-uniform, nonparametric, nonasymptotic confidence sequences,” The Annals of Statistics, vol. 49, no. 2, pp. 1055–1080, 2021.
[14]
F. Gnettner and C. Kirch, “A new and flexible class of sharp asymptotic time-uniform confidence sequences,” Statistics & Probability Letters, vol. 226, p. 110462, 2025.
[15]
J. Kim, J. Shin, A. Rinaldo, and L. Wasserman, “Uniform convergence rate of the kernel density estimator adaptive to intrinsic volume dimension,” in International conference on machine learning, 2019, pp. 3398–3407.
[16]
G. Cleanthous, A. G. Georgiadis, G. Kerkyacharian, P. Petrushev, and D. Picard, “Kernel and wavelet density estimators on manifolds and more general metric spaces,” Bernoulli, vol. 26, no. 3, pp. 1832–1862, 2020.
[17]
C. Berenfeld, P. Rosa, and J. Rousseau, “Estimating a density near an unknown manifold: A Bayesian nonparametric approach,” The Annals of Statistics, vol. 52, no. 5, pp. 2081–2108, 2024.
[18]
A. Green, S. Balakrishnan, and R. J. Tibshirani, “Minimax optimal regression over Sobolev spaces via Laplacian regularization on neighborhood graphs,” in International conference on artificial intelligence and statistics, 2021, pp. 2602–2610.
[19]
R. Singh and S. Vijaykumar, “Kernel ridge regression inference,” arXiv preprint arXiv:2302.06578, 2023.
[20]
V. Chernozhukov, D. Chetverikov, and K. Kato, “ANTI-CONCENTRATION AND HONEST, ADAPTIVE CONFIDENCE BANDS,” The Annals of Statistics, pp. 1787–1818, 2014.
[21]
V. Chernozhukov, D. Chetverikov, and K. Kato, “Comparison and anti-concentration bounds for maxima of Gaussian random vectors,” Probability Theory and Related Fields, vol. 162, no. 1, pp. 47–70, 2015.
[22]
V. Chernozhukov, D. Chetverikov, and K. Kato, “Empirical and multiplier bootstraps for suprema of empirical processes of increasing complexity, and related Gaussian couplings,” Stochastic Processes and their Applications, vol. 126, no. 12, pp. 3632–3651, 2016.
[23]
V. Chernozhukov, D. Chetverikov, and K. Kato, “CENTRAL LIMIT THEOREMS AND BOOTSTRAP IN HIGH DIMENSIONS,” The Annals of Probability, pp. 2309–2352, 2017.
[24]
F. Nazarov, “On the maximal perimeter of a convex set in with respect to a Gaussian measure,” in Geometric aspects of functional analysis: Israel seminar 2001-2002, 2004, pp. 169–187.
[25]
A. R. Klivans, R. O’Donnell, and R. A. Servedio, “Learning geometric concepts via Gaussian surface area,” in 2008 49th annual IEEE symposium on foundations of computer science, 2008, pp. 541–550.
[26]
H. Deng and C.-H. Zhang, “BEYOND GAUSSIAN APPROXIMATION,” The Annals of Statistics, vol. 48, no. 6, pp. 3643–3671, 2020.
[27]
A. Giessing, “Anti-concentration of suprema of Gaussian processes and Gaussian order statistics,” arXiv preprint arXiv:2310.12119, 2023.
[28]
R. Latala and K. Oleszkiewicz, “Gaussian measures of dilatations of convex symmetric sets,” The Annals of Probability, vol. 27, no. 4, pp. 1922–1938, 1999.
[29]
J. Ding, R. Eldan, and A. Zhai, On multiple peaks and moderate deviations for the supremum of a Gaussian field,” The Annals of Probability, vol. 43, no. 6, pp. 3468–3493, 2015.
[30]
S. Chatterjee, Superconcentration and related topics, vol. 15. Springer, 2014.
[31]
A. Ehrhard, “Symétrisation dans l’espace de Gauss,” Mathematica Scandinavica, vol. 53, no. 2, pp. 281–301, 1983.
[32]
S. Bobkov, “A note on the distributions of the maximum of linear Bernoulli processes,” Electronic Communications in Probability, vol. 13, pp. 266–271, 2008, doi: 10.1214/ECP.v13-1375.
[33]
V. Chernozhukov, D. Chetverikov, K. Kato, and Y. Koike, “Improved central limit theorem and bootstrap approximations in high dimensions,” The Annals of Statistics, vol. 50, no. 5, pp. 2562–2586, 2022.
[34]
E. H. Lieb and M. Loss, Analysis, 2nd ed., vol. 14. Providence, RI: American Mathematical Society, 2001.
[35]
V. Chernozhukov, D. Chetverikov, and K. Kato, “Detailed proof of Nazarov’s inequality,” arXiv preprint arXiv:1711.10696, 2017.

  1. By our convention, this does not restrict the law of \(M\) on the negative line and allows a mass point at zero.↩︎

  2. Here \(\|\beta\|_0\) denotes the number of non-zero coordinates of the vector \(\beta\).↩︎

  3. The situation is markedly different for non-asymptotic confidence sequences, though these do not share as close a connection to group sequential designs used in practice [13].↩︎

  4. By our Proposition 2, such a density spike must occur in a neighborhood of \(0\).↩︎