Empirical tail dependence functions in high dimensions:
uniform linearizations and inference
April 01, 2026
The analysis of extremal dependence in high dimensions is a key challenge in modern extreme-value statistics. Existing methodology primarily focuses on modeling and estimation of extremal dependence structures, often supported by concentration bounds for empirical tail quantities. However, comparatively little is known about general inferential procedures in high-dimensional extremes. In this paper, we develop foundational results that enable inference for rank-based empirical tail dependence coefficients, stable tail dependence functions, and functionals derived from them. We start by establishing finite-sample probability bounds that quantify the linearization error for such estimators uniformly over collections of coordinates. Moreover, we derive high-dimensional central limit theorems and establish the validity of multiplier bootstrap procedures for collections of empirical tail dependence statistics. Within an asymptotic framework, our results allow the dimension to grow exponentially with the effective sample size. We illustrate the usefulness of the results through two applications: uniform expansions and normal approximations for M-estimators of tail dependence parameters and inference for spatial isotropy based on collections of tail dependence functions.
Extreme value theory studies the probabilistic behavior and statistical analysis of rare events, that is, realizations of a random sample occurring at unusually high (or low) levels [1], [2]. A central object of interest is tail dependence, which describes the strength and structure of dependence between components of a random vector when some coordinates take extreme values. Understanding tail dependence is crucial for analyzing events driven or amplified by simultaneous extreme values accross multiple variables, with examples ranging from floods [3], [4] over climate extremes [5] to financial crises [6], [7]. Mathematically, tail dependence can be characterized using various equivalent objects, including stable tail dependence functions (STDF) and tail copulas, exponent and spectral measures, and Pickands dependence functions; see Chapters 8 and 9 in [1] and Chapters 6 and 7 in [2].
Motivated by applications involving large spatial fields or high-dimensional financial data, there has been rapidly growing interest in modeling and analyzing high-dimensional extremes. In such settings, fully nonparametric approaches are often difficult to interpret and may be computationally infeasible. Moreover, extreme value methods are particularly susceptible to the curse of dimensionality, as estimation relies solely on tail observations. These challenges have led to a variety of approaches that provide parsimonious and structured descriptions of tail dependence in high dimensions [8]. Popular approaches include clustering methods [9]–[12], principal component analysis [13], [14], factor models [15], graphical modeling and structure learning based on directed and undirected graphs [16]–[22] and vine copula constructions tailored to extremes [23].
When it comes to a formal mathematical analysis of the methods, some of the above works explicitly allow the dimension to grow with the sample size, a setting that is arguably most relevant for many modern applications. However, the available theoretical guarantees in this regime remain limited: either the proposed methods lack a rigorous theoretical analysis altogether, or they rely predominantly on concentration inequalities. The latter have been established for empirical (rank-based) tail dependence quantities by [24], with subsequent refinements in [25], [26] and [22]. While such results provide non-asymptotic bounds that quantify stochastic fluctuations and thus yield useful performance guarantees, they do not deliver distributional approximations and are therefore inherently insufficient for non-conservative inference in the form of confidence intervals or hypothesis tests.
To the best of our knowledge, the few existing contributions that address inference for extremes in growing dimensions do not cover the problem of tail dependence. [27] develop tests for marginal tail parameters of high-dimensional random vectors, relying on techniques specific to univariate extremes. [28] study a regression framework with high-dimensional predictors, focusing on the tail behavior of a univariate response conditional on covariates. Neither approach provides tools for inference on the extremal dependence structure.
The present paper develops tools for inference on tail dependence measures that comes with formal theoretical guarantees. Our focus is on STDFs and tail copulas, which are key building blocks in many modern methodologies for both low- and high-dimensional extremes. In fixed dimensions, the statistical properties of their empirical counterparts are well understood, typically through large-sample asymptotics in the form of (functional) central limit theorems. Foundational contributions were made by [29]–[31]; their results have been extended in various directions by [32]–[35]. Complementary bootstrap methods were developed in [36], and the resulting theory has been applied to parametric estimation in spatial models by [37]. A key challenge in this line of work is that the estimators are rank-based, which complicates the analysis as one must account for the stochastic fluctuations of empirical ranks in addition to those arising from the unknown tail dependence.1 However, the established theoretical tools and results do not readily extend to growing dimensions. In particular, (functional) weak convergence is no longer meaningful when the dimension of the ambient space increases. Moreover, existing results provide no quantitative insight into how the dimension affects the accuracy of distributional approximations.
We overcome these challenges through a two-step approach. In the first step, we derive linear representations of the empirical estimators, where the leading term is expressed as a sum of independent random variables. We establish convergence rates and provide explicit finite-sample probability bounds for the remainder terms. In particular, we identify regimes in which the remainder is asymptotically negligible relative to the leading term, even as the dimension grows. Our approach is inspired by related developments for empirical copulas in [39], with a key application consisting of linearizations that hold uniformly over large collections of lower-dimensional margins, such as all bivariate margins. This type of result is particularly relevant for high-dimensional models characterized by pairwise dependence structures, including the Hüsler–Reiss model. In the second step, we leverage recent advances in high-dimensional Gaussian approximation [40]–[42], combined with multiplier bootstrap techniques [43], to enable inference for the leading term. In this way, we extend bootstrap-based inferential methods for STDFs from the fixed-dimensional setting [36] to the high-dimensional regime.
We illustrate the scope of the results in two applications. First, we study M-estimators for tail dependence parameters in the spirit of [32], [44] and derive uniform asymptotic expansions and normal approximation in high dimensions. Second, we consider testing isotropy in spatial extremal dependence structures, where the proposed multiplier bootstrap enables inference for large collections of tail dependence functions. Simulation experiments illustrate the finite-sample performance of the procedures.
The remaining parts of this paper are organized as follows. Section 2 introduces tail dependence functions and their empirical counterparts. Section 3 establishes the uniform linearization results that form the basis of our analysis. Section 4 derives high-dimensional central limit theorems and establishes the validity of multiplier bootstrap procedures. Section 5 discusses two applications, namely M-estimation for tail dependence parameters and testing spatial isotropy. Proofs of the main results are collected in Appendix 6, while auxiliary technical results are deferred to Appendix 7.
For \(d\in\mathbb{N}\), we write \([d]=\{1, \dots, d\}\). For a real-valued function \(f\) defined on a set \(B \subseteq\mathbb{R}^d\) and \(\varepsilon>0\), let \[\begin{align} \label{eq:definition-modulus} \omega_{f}(\varepsilon;B) = \sup \big\{ |f(\boldsymbol{u}) - f(\boldsymbol{v})|: \boldsymbol{u}, \boldsymbol{v} \in B, \| \boldsymbol{u} - \boldsymbol{v}\|_\infty \le \varepsilon\big\} \end{align}\tag{1}\] denote the modulus of continuity with respect to the maximum norm on \(\mathbb{R}^d\). For \(\emptyset \ne I\subseteq[d]\) and \(\boldsymbol{x} \in [-\infty, \infty]^d\) write \(\boldsymbol{x}_I=(x_i)_{i \in I} \in [-\infty, \infty]^{I}\) for the vector made up by the coordinates of \(\boldsymbol{x}\) that belong to \(I\); note that we consider the vector to be indexed by \(I\) and not by \(\{1, \dots, |I|\}\). The same convention is applied for functions \(f_I\) defined on a subset \(B_I\) of \(\mathbb{R}^I\). If existent, we denote the partial derivative of \(f_I\) at \(\boldsymbol{x}_I \in B_I\) with respect to the \(j\)th coordinate (\(j\in I\)) by \(\partial_jf_I(\boldsymbol{x}_I)=\lim_{h \to 0} h^{-1} \{ f_I(\boldsymbol{x}_I + h \boldsymbol{e}_{I,j}) - f_I(\boldsymbol{x}_I)\}\), where \(\boldsymbol{e}_{I,j} \in \mathbb{R}^I\) has coordinates \(\boldsymbol{1}(i=j)\) for \(i \in I\). For a set \(A \subseteq[0,\infty)^d\) and \(\varepsilon>0\), let \(A^{\oplus \varepsilon}= \{ \boldsymbol{x} \in [0, \infty)^d : \mathrm{dist}(\boldsymbol{x}, A) \le \varepsilon\}\) denote the \(\varepsilon\)-enlargement of \(A\) in \([0,\infty)^d\), where \(\mathrm{dist}(\boldsymbol{x}, A) := \inf\{\|\boldsymbol{x} - \boldsymbol{y}\|_\infty: \boldsymbol{y} \in A\}\) is based on maximum-norm \(\| \cdot \|_\infty\) on \(\mathbb{R}^d\). Finally, \(\|\cdot\|_p\) denotes the \(p\)-norm, for \(p \ge 1\), and \(\mathcal{N}_d(\boldsymbol{\mu}, \boldsymbol{\Sigma})\) denotes the \(d\)-variate normal distribution with mean \(\boldsymbol{\mu}\) and variance matrix \(\boldsymbol{\Sigma}\).
Let \(\boldsymbol{X}=(X_1, \dots, X_d)^\top \in \mathbb{R}^d\) denote a \(d\)-variate random vector with common cumulative distribution function (cdf) \(F\) and continuous marginal cdfs \(F_1,\ldots,F_d\). As is standard in multivariate extremes, we assume that the dependence structure of \(\boldsymbol{X}\) stabilizes in the tail. Formally, this can be characterized through the existence of the stable tail dependence function \(L:[0,\infty)^d \to [0,\infty)\) or the tail copula \(R:[0,\infty]^d \setminus \{ \boldsymbol{\infty} \} \to [0,\infty)\) of \(\boldsymbol{X}\), which are defined by \[\begin{align} L(\boldsymbol{x}) \tag{2} &= \lim_{t \to 0} t^{-1} \mathbb{P}\big( \exists j \in [d]: F_j(X_j)>1-tx_j\big), \\ \tag{3} R(\boldsymbol{x}) &= \lim_{t \to 0} t^{-1} \mathbb{P}\big( \forall j \in [d]: F_j(X_j)>1-tx_j\big), \end{align}\] respectively. Both functions characterize the extremal dependence of \(\boldsymbol{X}\), and by inclusion-exclusion, we have \[L(\boldsymbol{x}) = \sum_{\emptyset \ne I \subseteq[d]} (-1)^{|I|+1} R_I(\boldsymbol{x}_I), \qquad R(\boldsymbol{x}) = \sum_{\emptyset \ne I \subseteq[d]} (-1)^{|I|+1} L_I(\boldsymbol{x}_I),\] where \(L_I(\boldsymbol{x}_I) = L(\boldsymbol{x}_I^0)\) and \(R_I(\boldsymbol{x}_I)=R( \boldsymbol{x}_I^\infty)\) with \(\boldsymbol{x}_I^a\) the vector having coordinates \(x_j\) for \(j \in I\) and \(x_j = a\) for \(j \in [d] \setminus I\), for \(a \in \{0,\infty\}\). Note that \[\begin{align} L_I(\boldsymbol{x}_I) &= \lim_{t \to 0} t^{-1} \mathbb{P}\big( \exists j \in I: F_j(X_j)>1-tx_j\big) \\ R_I(\boldsymbol{x}_I) &= \lim_{t \to 0} t^{-1} \mathbb{P}\big( \forall j \in I: F_j(X_j)>1-tx_j\big) \end{align}\] are nothing else than the stable tail dependence function and the tail copula of the sub-vector \(\boldsymbol{X}_I=(X_i)_{i\in I}\), which are formally functions \(L_I:[0,\infty)^I \to [0,\infty)\) and \(R_I:[0,\infty]^I \setminus \{ \boldsymbol{\infty} \} \to [0,\infty)\).
Evaluating \(L_I\) and \(R_I\) at the \(\boldsymbol{1}\)-vector, we obtain the extremal coefficient \(\theta_I\) [45] and the joint tail coefficient \(\chi_I\), that is, \[\begin{align} \label{eq:tail-coefficients} \theta_I = L_I(\boldsymbol{1}_I), \qquad \chi_I = R_I(\boldsymbol{1}_I). \end{align}\tag{4}\] Note that \(\chi_I=2-\theta_I = \lim_{t \to 0} \mathbb{P}(F_j(X_j) > 1-t \mid F_{j'}(X_{j'})>1-t)\) for \(I=\{j,j'\}\) of cardinality \(|I|=2\), which is also known as the upper tail dependence coefficient [46] or the tail correlation. The matrix of pairwise tail correlations \((\chi_I)_{I \subseteq[d]: |I|=2}\) plays a fundamental role in multivariate extreme value analysis [22].
Example 1 (Hüsler-Reiss distributions). The Hüsler-Reiss distribution has played a central role in recent developments on graphical modeling for extremes [16]. Its STDF is parametrized in terms of a \(d\)-dimensional symmetric, conditionally negative definite matrix \(\Gamma=(\gamma_{j\ell})\) with non-negative entries satisfying \(\gamma_{jj}=0\) for each \(j \in [d]\), and is given by \[\begin{align} L(\boldsymbol{x}; \Gamma) = \sum_{j=1}^d x_j\, \Phi_{d-1}\!\Big( \Big( \log\frac{x_j}{x_\ell} + \frac{\gamma_{\ell j}}{2} \Big)_{\ell\neq j}; \, \Sigma^{(j)} \Big), \end{align}\] where \(\Phi_d(\cdot; \Sigma)\) is the cdf of the \((d-1)\)-variate normal distribution with covariance matrix \(\Sigma\) and where \(\Sigma^{(j)} = (\Sigma^{(j)}_{\ell m})_{\ell,m \in [d] \setminus \{j\}}\) has entries \(\Sigma^{(j)}_{\ell m} = (\gamma_{\ell j} + \gamma_{mj} - \gamma_{\ell m})/2\) [47]. The bivariate marginal STDFs are given by \[\begin{align} L_{I}(x_j,x_\ell; \gamma_{j\ell}) = x_j\, \Phi\!\Big( \frac{\log(x_j/x_\ell)}{\sqrt{\gamma_{j\ell}}} + \frac{\sqrt{\gamma_{j\ell}}}{2} \Big) + x_\ell\, \Phi\!\Big( \frac{\log(x_\ell/x_j)}{\sqrt{\gamma_{j\ell}}} + \frac{\sqrt{\gamma_{j\ell}}}{2} \Big), \qquad I = \{j,\ell\}, \end{align}\] which shows that the parameter matrix \(\Gamma\) can be fully recovered from the bivariate margins only. Note that \(\lim_{\gamma \to +\infty} L_{I}(x_j,x_\ell; \gamma)= x_j+x_\ell\) and \(\lim_{\gamma\to 0} L_{I}(x_j,x_\ell;\gamma)= x_j \vee x_\ell\).
Example 2 (Factor models and max-linear models). As argued in [32], factor models with heavy-tailed factors and light-tailed noise lead to a STDF of the form \[L(\boldsymbol{x};B) = \sum_{j=1}^r \max_{\ell=1}^d (b_{j\ell} x_\ell), \qquad \boldsymbol{x} \in [0,\infty)^d,\] where \(B=(b_{j\ell})_{j \in [r], \ell \in [d]} \in [0,1]^{r \times d}\) has column sums 1. Such STDFs also arise in max-linear models on directed acyclic graphs which have recently gained popularity in modeling causal structural relationships in the tail [48].
We next introduce empirical tail dependence functions. Let \(\boldsymbol{X}_1, \dots, \boldsymbol{X}_n\) denote an i.i.d.sample of \(\boldsymbol{X}\), with \(\boldsymbol{X}_i=(X_{i1}, \dots, X_{id})^\top\). For \(j\in \{1, \dots, d\}\), let \(R_{ij}\) denote the rank of \(X_{ij}\) among \(X_{1j}, \dots, X_{nj}\). The empirical stable tail dependence function and the empirical tail copula are defined as \[\begin{align} \tag{5} \widehat L_{n}(\boldsymbol{x}) &:= \frac{1}{k} \sum_{i=1}^n \boldsymbol{1}\big(\exists j \in [d]: R_{ij} > n+1-kx_j\big), \\ \tag{6} \widehat R_{n}(\boldsymbol{x}) &:= \frac{1}{k} \sum_{i=1}^n \boldsymbol{1}\big(\forall j \in [d]: R_{ij} > n+1-kx_j\big), \end{align}\] where \(k \in [n]\) denotes a parameter to be chosen by the statistician that controls the size of the presumed tail area. Note that those estimators can be interpreted as ‘plug-in’ versions of the limiting relations in 2 and 3 . Indeed, replacing \(t\) by \(k/n\), \(F_j\) by the marginal empirical CDF and probabilities by their empirical counterparts leads to expressions that are almost identical to 5 and 6 . In order to obtain consistent estimators for \(L\) and \(R\), one typically needs to select an intermediate sequence \(k = k_n\) which satisfies \(k_n \to \infty, k_n/n\to 0\). The challenges in analyzing the estimators \(\widehat L_n, \widehat R_n\) are thus two-fold. First, taking ranks introduces dependence across all terms in the sum. Second, the sum is normalized by \(1/k\) rather than \(1/n\), and the distribution of the summands depends on \(n\) and \(k\).
In the finite-dimensional case where \(d\) is a fixed integer, the asymptotic behavior of \(\widehat L_n\) and \(\widehat R_n\) is well-studied [29], [32], [33]. We present one possible result in a way that is instructive for the developments in later sections. Let \[\begin{align} \label{eq:l-process-r-process-def} \mathbb{L}_n = \sqrt k(\widehat L_n - L ), \qquad \mathbb{R}_n = \sqrt k(\widehat R_n - R ) \end{align}\tag{7}\] denote the processes of rescaled estimation errors.
Let \(\Lambda\) denote the measure on the Borel subsets of \(\mathbb{E}_\infty := [0, \infty]^d \setminus \{ \boldsymbol{\infty}\}\) determined by \(\Lambda(A(\boldsymbol{x})) = L(\boldsymbol{x})\) where \[A(\boldsymbol{x}) := \big\{ \boldsymbol{y} \in \mathbb{E}_\infty \mid \exists j \in [d]: y_j < x_j\big\}.\] Let \(\mathbb{W}_\Lambda\) denote a zero-mean Gaussian process indexed by the Borel sets of \(\mathbb{E}_\infty\) with covariance function \(\mathbb{E}[\mathbb{W}_\Lambda(A) \mathbb{W}_\Lambda(B)] = \Lambda(A \cap B)\). The process shall be chosen in such a way that \([0,\infty)^d \to \mathbb{R}, \boldsymbol{x} \mapsto \mathbb{W}_L(\boldsymbol{x}) := \mathbb{W}_\Lambda(A(\boldsymbol{x}))\) is continuous almost surely. Finally, define \(\boldsymbol{V}_i=(V_{i1}, \dots, V_{id})^\top\) with \(V_{ij}=1-F_j(X_{ij})\) for \(j\in[d]\) and \(i\in[n]\), and let \[\begin{align} \tag{8} \widetilde{L}_n(\boldsymbol{x}) &= \frac{1}{k} \sum_{i=1}^n \boldsymbol{1}\Big( \exists j \in [d]: V_{ij} < \frac{k}{n}x_j \Big) \\ \tag{9} \widetilde{\mu}_n(\boldsymbol{x}) &= \frac{n}{k} \mathbb{P}\Big(\exists j \in [d]: V_{ij} < \frac{k}{n}x_j \Big) \end{align}\] and \(\widetilde{\mathbb{L}}_n(\boldsymbol{x}) = \sqrt k\big\{ \widetilde{L}_n(\boldsymbol{x}) - \widetilde{\mu}_n(x) \big\}\). Note that \(\widetilde{\mathbb{L}}_n(\boldsymbol{x})\) has expectation zero. We then have the following result.
Theorem 1 (Linearization and weak convergence for fixed \(d\), [32]). Suppose that the following conditions are met:
There exists \(\alpha>0\) such that \(\sup_{\boldsymbol{x} \in \Delta_{d-1}} \big| t^{-1} \mathbb{P}(F_1(X_1) > 1-tx_1 \text{ or } \dots \text{ or }F_d( X_d) > 1-tx_d) - L(\boldsymbol{x} ) \big| = O(t^\alpha)\) as \(t\to0\), where \(\Delta_{d-1} = \{ \boldsymbol{x} \in [0,1]^d: x_1 + \dots + x_d=1\}\).
\(k\to\infty\) and \(k=o(n^{2\alpha/(1+2\alpha)})\), with \(\alpha\) from [cond:second-order].
For all \(j\in[d]\), the first order partial derivative of \(L\) with respect to \(x_j\), say \(\partial_{j} L\), exists and is continuous on the set of points \(\boldsymbol{x}\) such that \(x_j>0\).
Then, for any fixed \(T \in \mathbb{N}\), we have \[\begin{align} \label{eq:linearization} \sup_{\boldsymbol{x} [0,T]^d} \big| \mathbb{L}_n(\boldsymbol{x}) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x}) \big| = o_\mathbb{P}(1), \end{align}\tag{10}\] where \[\begin{align} \label{eq:definition-widebar-Ln} \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x}) = \widetilde{\mathbb{L}}_n(\boldsymbol{x}) - \sum_{j =1 }^d \partial_{j} L(\boldsymbol{x}) \widetilde{\mathbb{L}}_{nj}(x_j). \end{align}\tag{11}\] Here, \(\widetilde{\mathbb{L}}_{nj}(x_j)=\widetilde{\mathbb{L}}_n(0, \dots, 0, x_j, 0, \dots, 0)\), and \(\partial_{j} L(\boldsymbol{x})\) is defined as the right-hand derivative at points \(\boldsymbol{x}\) with \(x_j=0\). Moreover, we have \(\widetilde{\mathbb{L}}_n = \sqrt k (\widetilde{L}_n - \widetilde{\mu}_n) \rightsquigarrow\mathbb{W}_L\) in \(\ell^\infty([0,T]^d),\) and hence \[\begin{align} \label{eq:lnw} \mathbb{L}_n = \sqrt k(\widehat L_n - L ) \rightsquigarrow\mathbb{B}_L \qquad \text{ in } \ell^\infty([0,T]^d), \end{align}\tag{12}\] where the limit process \(\mathbb{B}_L\) has the representation \[\mathbb{B}_L(\boldsymbol{x}) = \mathbb{W}_L(\boldsymbol{x}) - \sum_{j=1}^d \partial_{j} L(\boldsymbol{x}) \mathbb{W}_{L,j}(x_j)\] with \(\mathbb{W}_{L,j}(x_j) = \mathbb{W}_L(0, \dots, 0, x_j, 0 \dots, 0)\) for \(x_j \ge 0\).
While this result is not stated in any paper in this exact form, it can essentially be extracted from the proofs in [32]. Note that the weak convergence in 12 does not make sense if \(d\) changes with \(n\), whereas the representation in 10 can be reasonable. The proofs in [32] and related works, however, rely on the fact that the dimension \(d\) is fixed. In the following section, we derive a quantitative version of 10 that gives an explicit rate and tail bound for the difference in there and allows for increasing dimensions \(d=d_n\to\infty\). Finally, we note that a simple calculation shows that Assumption [cond:smoothness-1] holds if \(L\) is the STDF of a Hüsler-Reiss distribution from Example 1 but fails for the STDF corresponding to factor models in Example 2.
The main results in this section are two theorems that derive linearizations of the empirical tail dependence process \(\mathbb{L}_n\) under two different regularity assumptions on the partial derivatives of \(L\). For the first theorem, we fix an interesting set \(A\), for instance \(A=\{\boldsymbol{1}\}\) to handle the extremal coefficient \(\theta=\theta_{[d]}\) from 4 , and then demand sufficient regularity of \(L\) in a small extension of \(A\). For the second one, we start with \(L\), and derive uniform linearizations on sets that are adapted to the regularity of \(L\) and that are as large as possible. Either approach can be useful, depending on the application. For given \(T \in \mathbb{N}, \delta \in (0,e^{-1})\) and \(k\in\mathbb{N}\), let \[\begin{align} \label{eq:definition-r} r = r(\delta, T, k) = \sqrt{\frac{T}{k} \log\Big(\frac{1}{\delta}\Big)}. \end{align}\tag{13}\] Further, let \[\begin{align} \label{eq:bias} B_n(\boldsymbol{x}) = \sqrt k\big\{\widetilde{\mu}_n(\boldsymbol{x}) - L(\boldsymbol{x}) \big\}, \qquad \boldsymbol{x} \in [0,\infty)^d. \end{align}\tag{14}\] denote the rescaled difference between the preasymptotic STDF and the STDF itself, and write \[\begin{align} \label{eq:bias-A} B_{n,k}(L;S) := \sup_{\boldsymbol{x} \in S} |B_n(\boldsymbol{x})| \end{align}\tag{15}\] for \(S \subseteq[0,\infty)^d\). Our first result will be stated under the following regularity assumption on the pair \((A,L)\).
Theorem 2. Let \(L\) be a \(d\)-variate STDF and let \(A \subseteq[0,T]^d\) (with \(T \in \mathbb{N}\)) be a fixed set such that the pair \((A,L)\) satisfies Assumption [cond:smoothness-hoelder]. Then, there exist constants \(D_1=D_1(d), D_2=D_2(d)\) and \(D_3=D_3(d,K_L,\alpha_L)\) such that, for any \(n\in\mathbb{N}, k \in [n], \delta \in (0, e^{-1})\) satisfying \(\log(d/\delta)\le 2kT/7, n/k \ge T\) and \(r \le \kappa_L/C_{s}\) with \(C_{s}\) the universal constant from Lemma 11, we have \[\begin{align} \sup_{\boldsymbol{x} \in A } \big| \mathbb{L}_n(\boldsymbol{x}) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x}) \big| &\le B_{n,k}(L ; A^{\oplus \kappa_L}) + \frac{d}{\sqrt{k}} + D_{1} \sqrt{r \log\Big(\frac{TD_{2}}{\delta r}\Big)} + D_3 r^{\alpha_L} \sqrt{T \log\Big(\frac{1}{\delta}\Big)}. \end{align}\] with probability at least \(1- (6d+5)\delta\), with \(r\) from 13 . More specifically, the constant \(D_1\) depends on \(d\) via \(d^{3/2}\), while \(D_2\) and \(D_3\) depend linearly on \(d\) (precisely, \(D_3 = C_{s}^{1+\alpha_L} K_L d\)).
We provide an explicit discussion of the bias term, the smoothness condition [cond:smoothness-hoelder], and the domain parameter \(T\) in Remarks 4, 5, and 6, respectively.
In contrast to Theorem 1, Theorem 2 provides non-asymptotic control of the error in approximating \(\mathbb{L}_n\) by \(\begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n\) and also explicitly characterizes the effect of the dimension \(d\) on the approximation error. Another salient feature is that \(\delta\) only enters the bound logarithmically. This is crucial for considering many estimators simultaneously since the maximum error is still controllable by using union bound type arguments.
The upper bound \(d/\sqrt k\) prevents \(d\) from being of the order \(\sqrt k\) or larger. Much of the recent methodology for high-dimensional extremes does not attempt to estimate the entire joint tail of a large number of variables non-parametrically. For instance, the structure learning approaches in [17], [19], [22] are based on a large number of estimators of bivariate tail dependence. To perform statistical inference in such settings, one needs results that hold uniformly in a growing number of low-dimensional estimators rather than one high-dimensional estimator. Theorem 2 readily yields such results as we demonstrate next.
For \(I\subseteq[d]\) with \(|I|\ge 2\) and \(\boldsymbol{x}_I =(x_i)_{i \in I} \in [0,\infty)^I\), let \[\begin{align} \widehat L_{n,I}(\boldsymbol{x}_I) &= \frac{1}{k} \sum_{i=1}^n \boldsymbol{1}\big(\exists j \in I: R_{ij} > n+1-kx_j\big) = \widehat L_n(\boldsymbol{x}_I^0) \\ \widetilde{L}_{n,I}(\boldsymbol{x}_I) &= \frac{1}{k} \sum_{i=1}^n \boldsymbol{1}\Big( \exists j \in I: V_{ij} < \frac{k}{n}x_j \Big) = \widetilde{L}_{n}(\boldsymbol{x}_I^0) \\ \widetilde{\mu}_{n,I}(\boldsymbol{x}_I) &= \frac{n}{k} \mathbb{P}\Big(\exists j \in I: V_{ij} < \frac{k}{n}x_j \Big) = \widetilde{\mu}_{n}(\boldsymbol{x}_I^0) \end{align}\] denote the \(I\)-variate margin of \(\widehat L_n\), \(\widetilde{L}_{n}\) and \(\widetilde{\mu}_{n}\), respectively. Recall that \(\boldsymbol{x}_I^0\) has \(x_j\) for \(j \in I\) and \(x_j = 0\) for \(j \in [d] \setminus I\). Further, let \(\mathbb{L}_{n,I}=\sqrt k(\widehat L_{n,I} - L_I )\), \(\widetilde{ \mathbb{L}}_{n,I} = \sqrt k(\widetilde{L}_{n,I} - \widetilde{\mu}_I )\) and \[\begin{align} \label{eq:definition-widebar-LnI} \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}(\boldsymbol{x}_I) = \widetilde{\mathbb{L}}_{n,I}(\boldsymbol{x}_I) - \sum_{j\in I} \partial_{j} L_I(\boldsymbol{x}_I) \widetilde{\mathbb{L}}_{nj}(x_j). \end{align}\tag{16}\] The following result shows that we obtain linearizations that are uniform over collections of margins. It follows from the union bound and Theorem 2 applied to each \((A_I,L_I)\).
Corollary 1. Let \(\mathcal{I}\) be a collection of index sets \(I \subseteq[d]\) with \(|I| \ge 2\), and write \(m=\max_{I \in \mathcal{I}} |I|\). Fix \(T\in\mathbb{N}\), let \((A_I)_{I \in \mathcal{I}}\) be a collection of sets with \(A_I \subseteq[0,T]^I\), and suppose that, for each \(I\in \mathcal{I}\), \(\boldsymbol{X}_I\) has STDF \(L_I\) such that [cond:smoothness-hoelder] is met for \((A_I, L_I)\), with constants \(\kappa_I, K_{I}\) and exponent \(\alpha_I\). Then, with \(\kappa_L = \min_{I \in \mathcal{I}} \kappa_I, K_L=\max_{I \in \mathcal{I}} K_{I}\) and \(\alpha_L=\min_{I \in \mathcal{I}}a_I\), there exist constants \(D_1=D_1(m)\) and \(D_2=D_2(m)\) and \(D_3=D_3(m,K_L,\alpha_L)\) such that, for any \(n\in\mathbb{N}, k \in [n], \delta \in (0, e^{-1})\) satisfying \(\log(m/\delta)\le 2kT/7\), \(n/k \ge T\) and \(r \le \kappa_L/C_{s}\) with \(C_{s}\) from Lemma 11, we have \[\begin{align} \max_{I \in \mathcal{I}} \sup_{\boldsymbol{x} \in A_I} \big| \mathbb{L}_{n,I}(\boldsymbol{x}) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}(\boldsymbol{x}) \big| &\le \Big(\max_{I \in \mathcal{I}} B_{n,k}(L_I; A_I^{\oplus \kappa_L})\Big) + \frac{m}{\sqrt{k}} \\ &\qquad + D_{1} \sqrt{r\log\Big(\frac{TD_{2}}{\delta r}\Big)} + D_3 r^{\alpha_L}\sqrt{T \log\Big(\frac{1}{\delta}\Big)} \end{align}\] with probability at least \(1-|\mathcal{I}|(6m + 5)\delta\), with \(r\) from 13 and \(B_{n,k}\) from 15 .
To see the power of this result in applications with large \(|\mathcal{I}|\), let \(T=1, \alpha_L=1/2\) and write \(p\) for \(m|\mathcal{I}|\) to lighten the notation. Picking \(\delta = (9pk)^{-1}\) (recall that \(m \ge 2\), such that \(|\mathcal{I}|(6m+5) \le 9p\)) shows that, with probability at least \(1-k^{-1}\) \[\begin{align} \max_{I \in \mathcal{I}} \sup_{\boldsymbol{x} \in A_I} \big| \mathbb{L}_{n,I}(\boldsymbol{x}) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}(\boldsymbol{x}) \big| &\lesssim \Big(\max_{I \in \mathcal{I}} B_{n,k}(L_I; A_I^{\oplus \kappa_L})\Big) + \Big( \frac{\log^3(pk)}{k}\Big)^{1/4}, \end{align}\] where the implicit constant in \(\lesssim\) only depends on \(m\) and \(K_L\) and where we have used that \(r = \sqrt{k^{-1} \log(1/\delta)} \lesssim \sqrt{k^{-1} \log(pk)}\) and \(\log(D_2 / \delta r) \lesssim \log (D_2\sqrt k/\delta) \lesssim \log(pk)\). In an asymptotic framework with \(p=p_n, k= k_n, n \to \infty\) the upper bound vanishes provided that \(\log p = o(k^{1/3})\), i.e. even when the number of estimators we consider grows faster than any polynomial of \(k\). An important special case is \(\mathcal{I} = \{I \subseteq[d]: |I|=2\}\) and \(A_I=\{ \boldsymbol{1}_I\}\), which corresponds to uniform linearizations for all bivariate empirical extremal coefficients \((\theta_I)_{|I|=2}\).
For the next result, let \(E_j = \{ \boldsymbol{x} \in [0,\infty)^d: x_j>0\}\), and for a \(d\)-variate STDF \(L\), write \[\begin{align} G_j^{(1)} &= \big\{ \boldsymbol{x} \in E_j \mid \partial_{j} L (\boldsymbol{x}) \text{ exists and is continuous} \big\}, \\ G_{j\ell}^{(2)} &= \big\{ \boldsymbol{x} \in E_j \cap E_\ell \mid \partial_{j\ell} L (\boldsymbol{x}) \text{ exists and is continuous} \big\}, \end{align}\] where \(j,\ell \in [d]\). Moreover, write \(\mathfrak B_j^{(1)} = E_j \setminus G_j^{(1)}, \mathfrak B_{j\ell}^{(2)} = (E_j \cap E_\ell) \setminus G_{j\ell}^{(2)}\), and let \[\label{eq:defbadB} \mathfrak B = \Big(\bigcup_{j \in [d]} \mathfrak B_j^{(1)}\Big) \cup \Big(\bigcup_{j,\ell \in [d]} \mathfrak B_{j\ell}^{(2)}\Big)\tag{17}\] denote a set of ‘bad points’, where \(L\) is not sufficiently regular. The next theorem provides uniform linearizations of \(\mathbb{L}_n(\boldsymbol{x})\) over collections of points \(\boldsymbol{x}\) that are not too close to such ‘bad’ points. Consider the following smoothness condition on \(L\).
A detailed comparison of this condition with condition [cond:smoothness-hoelder] is given in Remark 5 below.
Theorem 3. Let \(L\) be a \(d\)-variate stable tail dependence function satisfying [cond:smoothness-good]. Fix \(T\in \mathbb{N}\). Then, there exist constants \(D_1=D_1(d,K_L)\) and \(D_2=D_2(d,K_L)\) such that, for any \(n\in\mathbb{N}, k \in [n], \delta \in (0, e^{-1})\) satisfying \(\log(d/\delta)\le 2kT/7\) and \(n/k \ge 2T\), we have \[\begin{align} \sup_{\boldsymbol{x} \in [0,T]^d \setminus (\mathfrak B^{\oplus C_{s}r})} \big| \mathbb{L}_n(\boldsymbol{x}) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x}) \big| &\le B_{n,k}(L; [0,T+C_{s}r]^d) + \frac{d}{\sqrt{k}} + D_{1} \sqrt{r\log\Big(\frac{TD_{2}}{\delta r}\Big)} \end{align}\] with probability at least \(1- (6d + 5)\delta\), where \(C_{s}\) is the universal constant from Lemma 11 and where \(r\) is from 13 . Here, the constant \(D_1\) depends quadratically on \(d\), while \(D_2\) depends linearly on \(d\).
For many models, the set \(\mathfrak B\) of bad points from 17 is actually empty. The derived linearization then holds uniformly on \([0,T]^d=[0,T]^d \setminus (\emptyset^{\oplus C_{s}r})\). Similar as for Theorem 2, the upper bound \(d /\sqrt k\) prevents \(d\) from being exponentially large, which can be avoided by treating \(m\)-dimensional margins only. The following result follows by combining the tail bounds in Theorem 3 with the union bound.
Corollary 2. Let \(\mathcal{I}\) be a collection of index sets \(I \subseteq[d]\) with \(|I| \ge 2\), and write \(m=\max_{I \in \mathcal{I}} |I|\). Suppose that, for each \(I\in \mathcal{I}\), \(\boldsymbol{X}_I\) has STDF \(L_I\) satisfying [cond:smoothness-good]; denote the respective set of bad points from 17 by \(\mathfrak B_{I}\). Fix \(T\in\mathbb{N}\). Then, with \(K_L=\max_{I \in \mathcal{I}} K_{L_I}\), there exist constants \(D_1=D_1(m,K_L)\) and \(D_2=D_2(m,K_L)\) such that, for any \(n\in\mathbb{N}, k \in [n], \delta \in (0, e^{-1})\) satisfying \(\log(m/\delta)\le 2kT/7\) and \(n/k\ge 2T\), we have \[\begin{align} &\max_{I \in \mathcal{I}} \sup_{\boldsymbol{x} \in [0,T]^I \setminus (\mathfrak B_I^{\oplus C_{s}r})} \big| \mathbb{L}_{n,I}(\boldsymbol{x}) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}(\boldsymbol{x}) \big| \\ & \le \Big(\max_{I \in \mathcal{I}} B_{n,k}(L_I; [0,T+C_{s}r]^I)\Big) + \frac{m}{\sqrt{k}} + D_{1} \sqrt{r\log\Big(\frac{TD_{2}}{\delta r}\Big)} \end{align}\] with probability at least \(1-|\mathcal{I}|(6m + 5)\delta\), where \(C_{s}\) is from Lemma 11, where \(r\) is from 13 and where \(B_{n,k}\) is from 15 .
Remark 4 (On the bias term). Most of the literature that deals with inference for multivariate extremes is based on second order conditions which control the speed of convergence in 2 or 3 , see for instance [17], [22], [32], [49] among many others. For many typical models, the speed of convergence in 2 or 3 is a power of \(t\). Consequently the bias \(k^{-1/2} B_n(\boldsymbol{x}) = \widetilde{\mu}_n(\boldsymbol{x}) - L(\boldsymbol{x})\) from 14 is a power of \(k/n\). In some settings, it is possible to establish the exact scaling and an exact asymptotic expansion for the bias, see Section 4 in [49] for details and further references.
Remark 5 (Comparison of [cond:smoothness-hoelder] and [cond:smoothness-good]). Conditions [cond:smoothness-hoelder] and [cond:smoothness-good] are different in nature, and neither condition is weaker than the other. Condition [cond:smoothness-hoelder] fails on sets of points that are not bounded away from zero, unless \(L\) is the STDF corresponding to tail independence.
Indeed, by homogeneity of \(L\), i.e. \(L(\lambda \boldsymbol{x}) = \lambda L(\boldsymbol{x})\) for all \(\boldsymbol{x} \in (0,\infty)^d\) and \(\lambda>0\), we have \(\partial_{j} L(\lambda \boldsymbol{x}) = \partial_{j} L(\boldsymbol{x})\) for every \(\boldsymbol{x}\) for which \(\partial_{j} L(\boldsymbol{x})\) exists. Suppose now that \(A\) from [cond:smoothness-hoelder] is not bounded away from zero. In that case, \(A\) contains a null-sequence \(\boldsymbol{x}_n\). If \(\boldsymbol{y}_1, \boldsymbol{y}_2 \in [0,1]^d\) are arbitrary, then \(\max_{i \in [2]}\| \boldsymbol{x}_n - \boldsymbol{y}_i/n\|_\infty \le \|\boldsymbol{x}_n\|_\infty + 1/n \le \kappa_L\) for sufficiently large \(n\), and [cond:smoothness-hoelder] then implies that \[\begin{align} \forall j \in [d]: \qquad |\partial_{j} L(\boldsymbol{y}_1) - \partial_{j} L(\boldsymbol{y}_2)| &= |\partial_{j} L(\boldsymbol{y}_1/n) - \partial_{j} L(\boldsymbol{y}_2/n)| \\&\le |\partial_{j} L(\boldsymbol{x}_n) - \partial_{j} L(\boldsymbol{y}_2/n)| + |\partial_{j} L(\boldsymbol{x}_n) - \partial_{j} L(\boldsymbol{y}_1/n)| \\&\le 2K_L \big(\|\boldsymbol{x}_n\|_\infty +1/n\big)^{\alpha_L} = o(1) \qquad (n \to \infty). \end{align}\] Hence, \(L\) must be linear on \([0,1]^d\), and the only linear STDF is the one corresponding to tail independence, \(L(\boldsymbol{x}) = \sum_{j\in[d]} x_j\).
In contrast, condition [cond:smoothness-good] can often be verified with \(\mathfrak B = \emptyset\), see Lemma 1 for an example in the bivariate case. When \((0, \infty)^d \subseteq G_{jl}^{(2)}\), Condition [cond:smoothness-good] implies Lipschitz continuity of the partial derivatives when all coordinates are away from zero, which is more restrictive than the Hölder assumption in [cond:smoothness-hoelder]. Condition [cond:smoothness-hoelder] is thus most useful for establishing expansions at individual points \(\boldsymbol{x}\) with entries bounded away form zero under minimal assumptions, or on sets of such points. Important applications include the extremal coefficient or tail correlation.
We next discuss Condition [cond:smoothness-good], which is related to Assumption 2 in [22]. By homogeneity of \(L\), that is, \(L(\lambda \boldsymbol{x}) = \lambda L(\boldsymbol{x})\) for all \(\boldsymbol{x} \in [0,\infty)^d\) and \(\lambda>0\), we have \(\partial_{j} L(\lambda \boldsymbol{x}) = \partial_{j} L(\boldsymbol{x})\) and \(\partial_{j\ell} L(\lambda \boldsymbol{x}) = \lambda^{-1} \partial_{j\ell} L(\boldsymbol{x})\) for all \(j, \ell\in[d]\). It is hence sufficient to check the required bound for \(\boldsymbol{x} \in G_{j\ell}^{(2)} \cap [0,1]^d\), as it then automatically holds for all \(\boldsymbol{x} \in G_{j\ell}^{(2)}\) with the same constant \(K_L\). The following lemma provides a simple sufficient condition for the bivariate case.
Lemma 1. Suppose \(L\) is a bivariate stable tail dependence function, and let \(A(t) = L(1-t,t)\), \(t \in [0,1]\), denote the associated Pickands dependence function. If \(A\) is twice continuously differentiable on \((0,1)\) and if \(A_{\infty} := \sup_{t\in(0,1)} t (1-t) A''(t) < \infty\), then Condition [cond:smoothness-good] is met for \(L\), with \(\mathfrak B=\emptyset\) and with \(K_L=A_{\infty}\).
If, for instance, \(L\) is the stable tail dependence function of the \(d\)-variate Hüsler-Reiss-copula with parameter matrix \(\Gamma=(\gamma_{j\ell})_{j,\ell\in[d]}\) satisfying \(\lambda_0:= \min_{j \ne \ell} \gamma_{j\ell} >0\) (i.e., the bivariate margins are bounded away from perfect dependence; see Example 1), then each bivariate marginal Pickands dependence function \(A_I\) satisfies \(A_{I,\infty} \le C_A\) for some constant \(C_A=C_A(\lambda_0)\) [39]. As a consequence, Corollary 2 is applicable with \(\mathcal{I}=\{ I \subseteq[d]: |I|=2\}\), with \(\mathfrak B_I = \emptyset\), and with \(K_L= \max_{|I|=2}A_{I,\infty}\le C_A\).
Remark 6 (On the domain parameter \(T\)). It is possible to derive Theorem 2 with general \(T \in \mathbb{N}\) as stated from the version with \(T=1\) only by utilizing certain homogeneity properties. To make explicit the dependence of the estimator \(\widehat L_n\) on \(k\), we will write \(\widehat L_{n,k}\) throughout this remark. For example, for any \(\eta > 0\) such that \(\eta k\) is an integer, a straightforward calculation yields \(\widehat L_{n,k}(\eta\boldsymbol{x}) = \eta \widehat L_{n,k\eta}(\boldsymbol{x})\). Together with homogeneity of \(L\), this implies \[\mathbb{L}_{n,k}(\eta \boldsymbol{x}) = \sqrt{k}\big(\widehat L_{n,k}(\eta\boldsymbol{x}) - L(\eta\boldsymbol{x}) \big) = \sqrt{\eta} \sqrt{k\eta} \big(\widehat L_{n,k\eta}(\boldsymbol{x}) - L(\boldsymbol{x}) \big) = \sqrt{\eta} \mathbb{L}_{n,k\eta}( \boldsymbol{x}).\] Similar computations show \(\widetilde{\mathbb{L}}_{n,k}(\eta \boldsymbol{x}) = \sqrt{\eta} \widetilde{\mathbb{L}}_{n,k\eta}(\boldsymbol{x})\), \(\begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,k}(\eta \boldsymbol{x}) = \sqrt{\eta} \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,k\eta}(\boldsymbol{x})\). We still choose to state the version for general \(T\) directly since the full conversion requires some tedious work. A similar comment applies to some of the other results in this section.
Let \(\mathcal{I}\) be a finite collection of index sets \(I \subseteq[d]\) with \(|I| \ge 2\), let \(m = \max_{I \in \mathcal{I}} |I|\). For each \(I \in \mathcal{I}\), assume that \(L_I\) exists, let \(A_I = \{\boldsymbol{x}_{I,1}, \dots, \boldsymbol{x}_{I,p_I}\}\) be a finite set of vectors in \((0,1]^{I}\), and let \(p = \sum_{I \in \mathcal{I}} p_I \ge |\mathcal{I}|\). Note that we restrict ourselves to \(T=1\), which is not restrictive by homogeneity of STDFs. Our goal is to derive Gaussian approximations for the \(p\)-dimensional random vector \[\begin{align} \label{eq:Sn-clt} \boldsymbol{S}_n = (\mathbb{L}_{n,I}(\boldsymbol{x}_{I, \ell}))_{I \in \mathcal{I}, \ell \in [p_I]}. \end{align}\tag{18}\] Writing \(\boldsymbol{y}_{I, \ell} = (\boldsymbol{x}_{I, \ell}, \boldsymbol{0}_{I^c}) \in [0,1]^d\) and \(A = \bigcup_{I \in \mathcal{I}} \{ \boldsymbol{y}_{I, \ell}: j \in [p_I]\}\), we can write \[\boldsymbol{S}_n = (\mathbb{L}_{n}(\boldsymbol{y}))_{y \in A} \in \mathbb{R}^p.\] Such high-dimensional vectors arise naturally, for instance, when considering the extremal coefficient matrix with elements \(\theta_I=L_I(\boldsymbol{1}_I)\) for \(I \subseteq[d]\) with \(|I|=2\). The rescaled estimation error of the empirical counterpart is \(\sqrt k(\hat{\theta}_I - \theta_I) = \mathbb{L}_{n,I}(\boldsymbol{1}_I)\). Collecting these errors in a vector corresponds to considering \(\mathcal{I} = \{ I \subseteq[d]: |I|=2\}\) and \(A_I =\{\boldsymbol{1}_I\}\), with \(m=2\) and \(p = d(d-1)/2\).
Let \[\begin{align} \boldsymbol{G}_n \sim \mathcal{N}_{p}(\boldsymbol{0}, \Sigma_n), \quad \text{ where } \Sigma_n = \operatorname{Var}(\boldsymbol{T}_n) \text{ with } \boldsymbol{T}_n = ( \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}(\boldsymbol{x}_{I, \ell}))_{I \in \mathcal{I}, \ell \in [p_I]} \in \mathbb{R}^p. \end{align}\] and with \(\begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}\) from 16 . Specific formulas for the entries of \(\Sigma_n\) are given in 54 . Write \(\sigma_{n,q}^2\) for \(q\)th diagonal element of \(\Sigma_n\). For random vectors \(\boldsymbol{S}\) and \(\boldsymbol{T}\) of the same dimension \(p\in\mathbb{N}\), let \[d_K(\boldsymbol{S}, \boldsymbol{T}) = \sup_{\boldsymbol{x} \in \mathbb{R}^{p}} \big| \mathbb{P}(\boldsymbol{S} \le \boldsymbol{x}) - \mathbb{P}(\boldsymbol{T} \le \boldsymbol{x}) \big|\] denote the Kolmogorov distance between \(\boldsymbol{S}\) and \(\boldsymbol{T}\). The following result provides a bound on \(d_K(\boldsymbol{S}_n, \boldsymbol{G}_n)\) under a condition as in Corollary 1; adaptations to the conditions of Corollary 2 follow along similar lines and are omitted for the sake of brevity. The obtained upper bound has similar features as the bounds in classical high-dimensional Gaussian approximation results in [43]. However, there is an additional bias term which is due to the fact that we do not directly observe data from \(L\) but rather work with domain of attraction conditions. Note also that \(n\) in the upper bound in [43] is replaced by \(k\) in our setting. Intuitively, this is because we effectively only use \(k\) observations to compute \(\widehat L\).
Theorem 7. Let \(\mathcal{I}\) and \((A_I)_{I \in \mathcal{I}}\) be as described in the beginning of Section 4 and suppose that the STDF \(L_I\) of \(\boldsymbol{X}_I\) exists for every \(I \in \mathcal{I}\). Assume that there exist \(\kappa_L, K_L\in(0,\infty)\) and \(\alpha_L \in (1/2,1]\) such that \[\begin{align} \forall I \in \mathcal{I}, & \forall j \in I, \forall \boldsymbol{x}_I \in A_I, \forall \boldsymbol{y}_I \in [0,\infty)^I \text{ with } \|\boldsymbol{x}_I - \boldsymbol{y}_I\|_\infty \le \kappa_L: \\ &\partial_{j} L_I(\boldsymbol{x}_I), \partial_{j} L_I(\boldsymbol{y}_I) \text{ exist and satisfy } |\partial_{j} L_I(\boldsymbol{x}_I)-\partial_{j} L_I(\boldsymbol{y}_I)| \le K_L\|\boldsymbol{x}_I - \boldsymbol{y}_I\|_\infty^{\alpha_L}. \end{align}\] Moreover, assume that \(m|\mathcal{I}| \ge 3, n \ge 2,p \ge 2\) and
\(\sigma_{\min}^2 := \min_{q \in [p]} \sigma_{n,q}^2>0\).
\(\log(m^2 |\mathcal{I}|k^{1/4} ) \le 2k/7\).
\(\log(m|\mathcal{I}|k^{1/4}) \le \kappa_L^2 k / C_{s}^2\) with \(C_{s}\) from Lemma 11.
Then there exists a constant \(c = c(\sigma_{\min}^2, m, K_L, \alpha_L) \ge 1\) such that \[d_K(\boldsymbol{S}_n, \boldsymbol{G}_n) \le c \Big[ \sqrt{\log p} \Big(\max_{I \in \mathcal{I}} B_{n,k}(L_I; A_I^{\oplus \kappa_L}) \Big) + \Big( \frac{\log^5(pn)}{k} \Big)^{1/4} \Big].\]
We briefly discuss the assumptions and the result. First, the smoothness condition on the collection \((L_I)_I\) essentially requires [cond:smoothness-hoelder] to hold for each pair \((A_I, L_I)\), see also Corollary 1. The assumptions \(m|\mathcal{I}| \ge 3, n \ge 2,p \ge 2\) are very mild; they can be omitted at the cost of more technical arguments within the proof. The variance condition in (i) is required for high-dimensional CLTs as in [42]; as shown in Remark 10 below, it is a very mild and natural requirement if \(m=2\). Finally, the conditions in (ii) and (iii) can best be interpreted in an asymptotic (triangular array) framework where \(\mathcal{I}=\mathcal{I}_n\) and \(k=k_n\) is allowed to depend on \(n\): both conditions are satisfied for sufficiently large \(n\) if \(\log(|\mathcal{I}_n|) = o(k_n)\). In such an asymptotic framework, the upper bound on the Kolmogorov distance converges to zero if \(\log^5(p_n) = o(k_n)\) and if the (uniform) bias term is of smaller order that \(\sqrt{\log(p_n)}\). Finally, note that the factor \(\sqrt{\log p}\) in front of the bias term is natural in view of Lemma 1 in [43].
Remark 8 (Other possible versions). We note that the proof of Theorem 7 utilizes a particular version of a high-dimensional Gaussian approximation result from [42]. Specifically, the proof proceeds by applying Theorem 19 to the collection \(\boldsymbol{T}_n = ( \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}(\boldsymbol{x}_{I, \ell}))_{I \in \mathcal{I}, \ell \in [p_I]}\) and controlling the error in approximating \(\boldsymbol{S}_n\) by \(\boldsymbol{T}_n\). Depending on the assumptions, other versions of high-dimensional Gaussian approximation results can be applied. Here, we briefly mention two possible versions without going into details. First, [50], [51] have shown that even better rates for the error are possible if the covariance matrix of the vector \(\boldsymbol{T}_n\) has smallest eigenvalue bounded away form zero. In that case a rate of \(n^{-1/2}\) up to poly-logarithmic factors can be achieved. Second, Theorem 7 requires a lower bound on the variance. This is because the smallest variance appears in both, Theorem 19 and in the proof via Theorem 18. This restricts the set of admissible values for \(\boldsymbol{x}\) away from the origin and also rules out asymptotic independence. At the cost of a slower rate and for a restricted Kolmogorov distance, it is possible to drop the minimum variance assumption by utilizing Lemma 7 and Theorem 3 in [52].
Remark 9 (On supremum statistics). The result in Theorem 7 is sufficiently strong to cover distributional approximations for supremum-statistics. It is instructive to study the bivariate case first, and more specifically, we are then interested in approximations for the cdf of \(\sup_{\boldsymbol{x} \in B} \mathbb{L}_n(\boldsymbol{x})\) with \(B \subseteq[0,1]^2\). In view of the fact that \(\widehat L_n\) is a piecewise constant function that is constant on intervals of the form \([\ell/k, (\ell+1)/k) \times [\ell'/k, (\ell'+1)/k)\), we have \(\sup_{\boldsymbol{x} \in B} \mathbb{L}_n(\boldsymbol{x}) = \max_{ \boldsymbol{x} \in B \cap G} \mathbb{L}_n(\boldsymbol{x}) ,\) where \(G\) contains all vectors in \([0,1]^2\) of the form \((\ell/k, \ell'/k)\) with \(\ell, \ell' \in \mathbb{N}_0\). Note that \(|G| \le (k+1)^2\). As a consequence, \[\mathbb{P}\Big( \sup_{\boldsymbol{x} \in B} \mathbb{L}_n(\boldsymbol{x}) \le t \Big) = \mathbb{P}\Big(\max_{ \boldsymbol{x} \in B \cap G} \mathbb{L}_n(\boldsymbol{x}) \le t \Big) = \mathbb{P}\Big( (\mathbb{L}_n(\boldsymbol{x}))_{\boldsymbol{x} \in B \cap G} \le \boldsymbol{t} \Big),\] where \(\boldsymbol{t}=(t,\dots, t)\in \mathbb{R}^{ B \cap G}\). We can hence apply Theorem 7 with \(p=| B \cap G| \le (k+1)^2\), and the approach could easily be extended to the multivariate case, which each margin under consideration contribution at most \((k+1)^m\) to \(p\).
Remark 10 (On the variance condition). A generic diagonal element \(\sigma_{n,q}^2\) of \(\Sigma_n\) can be written as \(\sigma_{n,I}^2(\boldsymbol{x}_I) = \mathbb{E}[ \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}^2(\boldsymbol{x}_I)]\) for certain \(I \in \mathcal{I}\) and \(\boldsymbol{x}_I \in A_I\). A straightforward calculation, carried out in Section 6.2, shows that, if \(I\) and \(\boldsymbol{x}_I\) are fixed and if \(k=k_n\) satisfies \(k_n =o(n)\) as \(n \to \infty\), \[\begin{align} \sigma_I^2 (\boldsymbol{x}_I) &= \lim_{n \to \infty}\sigma_{n,I}^2(\boldsymbol{x}_I) = - L_I(\boldsymbol{x}_I) + (\nabla L_I(\boldsymbol{x}_I))^\top \mathcal{R}_I(\boldsymbol{x}_I) (\nabla L_I(\boldsymbol{x}_I)) , \end{align}\] where \(\nabla L_I(\boldsymbol{x}_I)= (\partial_{j} L_I(\boldsymbol{x}_I))_{j \in I} \in \mathbb{R}^I\) and where \(\mathcal{R}_I(\boldsymbol{x}_I) = (R_{\{j,\ell\}}(x_{I,j}, x_{I,\ell}))_{j,\ell \in I}\) is a \(|I| \times |I|\) matrix, with diagonal entries \(R_{\{j,j\}}(x_{I,j}, x_{I,j}) = x_{I,j}\) and with \(R_{\{j, \ell\}}\) the tail copula of the bivariate sub-vector \(X_{\{j,\ell\}}\) of \(\boldsymbol{X}_I\). The variance condition in (i) of Theorem 7 would be satisfied for sufficiently large \(n\) (more precisely, for sufficiently small \(k/n\)) if \(\sigma_I^2 (\boldsymbol{x}_I)\) is bounded away from zero, uniformly in \(I\) and \(\boldsymbol{x}_I\). We show in Section 6.2 that, in the case \(|I|=2\), \(\sigma_I^2 (\boldsymbol{x}_I)\) is non-zero if and only if \(R_I \notin\{ R_{{\text{ind}}}, R_{\text{pd}}\}\), where \(R_{{\text{ind}}} \equiv 0\) and \(R_{\text{pd}}(x,y) = x \wedge y\) correspond to tail independence and perfect tail dependence, respectively. As a consequence, (i) would be satisfied for sufficiently large \(n\) if all \(R_I\) are bounded away from these two extreme cases.
Next, we derive bootstrap approximations, following the multiplier approach from [36], whose validity will be transferred to the high-dimensional setting by combining arguments from [43] with a careful analysis of the impact of estimating the partial derivatives \(\partial_j L\) in the bootstrap procedure. The presence of the latter means that the high-dimensional bootstrap result in Theorem 3 of [43] is not directly applicable and additional arguments are needed. The approach requires suitable estimates of the partial derivatives of \(L_I\), for which one may follow a simple finite-differencing technique: for \(\boldsymbol{x}_I \in (0,\infty)^I\), \(j \in I\), and a bandwidth parameter \(h>0\) such that \(0< h < x_j\), define \[\widehat{\partial_{j} L}_I(\boldsymbol{x}_I) = \widehat{\partial_{j} L}{}_{n,h,I}(\boldsymbol{x}) = \min\Big\{ \frac{\widehat L_{n,I}(\boldsymbol{x} + h \boldsymbol{e}_j) - \widehat L_{n,I}(\boldsymbol{x} - h \boldsymbol{e}_j)}{2h}, 1\Big\}.\] Next, note that \[\begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}(\boldsymbol{x}_I) = \sum_{i=1}^n Y_{i,I}(\boldsymbol{x}_I),\] where \[\begin{align} \label{eq:YiI} Y_{i,I}(\boldsymbol{x}_I) \nonumber &= \frac{1}{\sqrt k} \Big[ \boldsymbol{1}(\exists j \in I: V_{ij} < kx_j/n) - \mathbb{P}(\exists j \in I: V_{ij} < kx_j/n) \\ & - \sum_{j\in I} \partial_{j} L_I(\boldsymbol{x}_I) \big\{\boldsymbol{1}(V_{ij} < kx_j/n) - kx_j/n \big\} \Big]. \end{align}\tag{19}\] Define observable counterparts of \(Y_{i,I}(\boldsymbol{x}_I)\) by \[\begin{align} \label{eq:hatYI} \widehat Y_{i,I} (\boldsymbol{x}_I) &= \nonumber \frac{1}{\sqrt k} \Big[ \boldsymbol{1}(\exists j \in I: \hat{V}_{ij} < kx_j/n) - (k/n) \widehat L_{n,I}(\boldsymbol{x}_I) \\ & - \sum_{j\in I} \widehat{\partial_{j} L}_{I}(\boldsymbol{x}_I) \big\{\boldsymbol{1}(\hat{V}_{ij} < kx_j/n) - kx_j/n \big\} \Big], \end{align}\tag{20}\] where \(\hat{\boldsymbol{V}}_i=(\hat{V}_{i1}, \dots, \hat{V}_{id})^\top\) has coordinates \(\hat{V}_{ij} = 1 +n^{-1}-n^{-1}R_{ij}\). For \(e_1, e_2, \dots\) iid standard normal and independent of the observations \(\boldsymbol{X}_i\), we propose to use \[\begin{align} \label{eq:Sn42-boot} \boldsymbol{S}_n^* = ( \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup ^*_{n,I}(\boldsymbol{x}_{I,\ell}))_{I \in \mathcal{I}, \ell \in [p_I]}, \qquad \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup ^*_{n,I}(\boldsymbol{x}_I) = \sum_{i=1}^n e_i \widehat Y_{i,I} (\boldsymbol{x}_I) \end{align}\tag{21}\] as a bootstrap approximation for \(\boldsymbol{S}_n\) from 18 . The following result provides high-probability bounds for \[d_K( \mathcal{L}(\boldsymbol{S}_n^* \mid \mathrm{data}), \boldsymbol{G}_n)\] under a suitable Hölder smoothness assumption on each \(L_I\). Unlike for the CLT from Theorem 7, we restrict attention to the case where the Hölder exponent is 1; extensions to other exponents or smoothness assumptions as in Corollary 2 are possible but are omitted for the sake of a clear exposition.
Theorem 11. Let \(\mathcal{I}\) and \((A_I)_{I \in \mathcal{I}}\) be as described in the beginning of Section 4 and suppose that the STDF \(L_I\) of \(\boldsymbol{X}_I\) exists for every \(I \in \mathcal{I}\). Assume that there exist \(\kappa_L, K_L\in(0,\infty)\) such that \[\begin{align} \forall I \in \mathcal{I},&\forall j \in I, \forall \boldsymbol{x}_I \in A_I^{\oplus \min(1,\kappa_L/2)}, \forall \boldsymbol{y}_I \in [0,\infty)^I \text{ with } \|\boldsymbol{x}_I - \boldsymbol{y}_I\|_\infty \le \kappa_L: \\ &\partial_{j} L_I(\boldsymbol{x}_I), \partial_{j} L_I(\boldsymbol{y}_I) \text{ exist and satisfy } |\partial_{j} L_I(\boldsymbol{x}_I)-\partial_{j} L_I(\boldsymbol{y}_I)| \le K_L\|\boldsymbol{x}_I - \boldsymbol{y}_I\|_\infty. \end{align}\] Assume the conditions (i)–(iii) of Theorem 7 are met with the condition \(\log(m|\mathcal{I}| k^{1/4}) \le \kappa_L^2k/ C_{s}^2\) replaced by \(\log(m|\mathcal{I}| k^{1/4}) \le \kappa_L^2k/ (8C_{s}^2)\), and with \(n/k\ge 2\). Let \(0<c_h< c_h'< \infty\) be constants, and assume that the bandwidth \(h<(\min_{I \in \mathcal{I}} \min_{\boldsymbol{x}_I \in A_I} \min_{j \in I}x_{I,j}) \wedge (\kappa_L/2)\) satisfies \[c_h\Big( \frac{\log(p+k)}{k}\Big)^{1/2} \le h \le c_h' \Big( \frac{\log(p+k)}{k}\Big)^{1/4}.\] Then, there exist constants \(c_i = c_i(m,K_L,\sigma_{\mathrm{min}},c_h,c_h'), i = 1,2\) such that, with probability at least \(1-c_1\delta_n\) \[d_K( \mathcal{L}(\boldsymbol{S}_n^* \mid \mathrm{data}), \boldsymbol{G}_n) \le c_2 \Big\{\delta_n + \sqrt{\log(p+k)} B_{n,k}(L_I ; A_I^{\oplus\kappa_L}) \Big\},\] where \(\delta_n = [k^{-1}\log^5(pn)]^{1/4}\).
We briefly comment on the conditions and the result. The smoothness condition is a slightly stronger version of the one imposed for Theorem 7: first, we restrict attention to \(\alpha_L=1\) for simplicity, and second, the third \(\forall\)-quantor requires \(\boldsymbol{x}_I\) to be from a small extension of \(A_I\) rather than from \(A_I\) only. This extension is needed in the proofs when passing from estimated partial derivatives to true unknown partial derivatives. The strengthening of condition (iii) from Theorem 7 is mild. Finally, the condition on the bandwidth is mild in the sense that the same approximation bound is obtained for a large range of bandwidth choices. The obtained rate is almost the same as in Theorem 7, with a factor \(\sqrt{\log(p+k)}\) instead of \(\sqrt{\log(p)}\) in front of the bias term; in particular, the same ‘rate’ is obtained in the (high-dimensional) case where \(k \lesssim p\).
As an application of the uniform linearizations established in Section 3, we derive corresponding linearizations for moment estimators based on integrals of \(\widehat L_n\). We first consider estimators constructed from the full \(d\)-variate function \(\widehat L_n\) and subsequently turn to a collection of M-estimators based on lower-dimensional margins of \(\widehat L_n\). In the latter setting, we additionally establish high-dimensional central limit theorems.
In defining the estimators, we follow the setup in [32]. Let \(\{L(\cdot; \theta) \colon \theta \in \Theta\}\) be a parametric family of STDFs, with a parameter space \(\Theta \subseteq \mathbb{R}^s\). Next, let \[Q_n(\theta) := \Big\|\int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x}) \big(L(\boldsymbol{x};\theta)-\widehat L_n(\boldsymbol{x})\big) \mathrm d\mu(\boldsymbol{x})\Big\|_2\] for a (known) measure \(\mu\) on \([0,1]^d\) and a (known) function \(\boldsymbol{g}: [0,1]^d \to \mathbb{R}^q\) with \(q \in \mathbb{N}_{\ge s}\) such that \[\label{eq:defCg} C_g \mathrel{\vcenter{:}}= \;\int_{[0,1]^d} \|\boldsymbol{g}(\boldsymbol{x})\|_2 \,\mathrm d\mu(\boldsymbol{x})<\infty.\tag{22}\] For the subsequent analysis, we also define the population version of \(Q_n\) which is given by \[Q_L(\theta) := \Big\|\int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x})\big( L(\boldsymbol{x};\theta)-L(\boldsymbol{x}) \big) \mathrm d\mu(\boldsymbol{x}) \Big\|_2.\] [32] assume that \(\theta \mapsto \int \boldsymbol{g}L(\cdot;\theta)\mathrm d\mu\) is a homeomorphism between \(\Theta\) and its codomain and show that, under certain conditions, \(Q_{n}\) has a unique minimizer in \(\Theta\) with probability going to one when the sample sizes grows to infinity. We will take a different route and instead prove results for any sufficiently good approximate minimizer of \(Q_n\), i.e. any \(\hat{\theta}_n\) that satisfies \[Q_n(\hat{\theta}_n)-\inf_{\theta \in \Theta} Q_n(\theta) < \eta\] for \(\eta\) ‘small’ in a sense made precise below. This allows us to give statistical guarantees for estimators that are computed by numerical optimization, which is a common scenario in practice. We will work under the following assumptions.
Assumption 1. There exist constants \(\kappa>0, \gamma_{h}\in (0,1]\) and \(C_{h}>0\) such that the tuple \((L, \{L(\cdot; \theta): \theta \in \Theta\}, \boldsymbol{g}, \mu)\) satisfies the following:
The function \(\theta \mapsto Q_L(\theta)\) has a unique minimum in \(\theta_0\).
The closed ball \(\overline{B}_\kappa(\theta_0) := \{\theta \colon \| \theta - \theta_0 \|_2\leq \kappa\}\) is contained in \(\Theta\).
The function \(\boldsymbol{\varphi} \colon \Theta \subseteq\mathbb{R}^s \to \mathbb{R}^q\) defined by \(\boldsymbol{\varphi}(\theta) = \int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x})L(\boldsymbol{x};\theta)\mathrm d\mu(\boldsymbol{x})\) is twice differentiable on \(B_\kappa(\theta_0) := \{\theta \colon \| \theta - \theta_0 \|_2< \kappa\}\).
All mixed second order partial derivatives of \(\boldsymbol{\varphi}\) are uniformly Hölder continuous at \(\theta_0\) in the following sense: \[\forall \theta \in B_{\kappa}(\theta_0): \quad \max_{j,\ell \in [s], p \in [q]}\left|\partial_{j\ell} \varphi_p(\theta)-\partial_{j\ell}\varphi_p(\theta_0)\right| \leq C_{h}\left\lVert\theta -\theta_0\right\rVert_2^{\gamma_{h}},\] where, for \(\theta' \in B_\kappa(\theta_0)\), \(j,\ell \in [s]\) and \(p \in [q]\), \[\partial_j \varphi_p(\theta') = \frac{\partial}{\partial \theta_j} \varphi_p(\theta)\bigg|_{\theta = \theta'} \quad \text{ and }\quad \partial_{j\ell} \varphi_p(\theta') = \frac{\partial^2}{\partial \theta_j \partial \theta_\ell} \varphi_p(\theta)\bigg|_{\theta = \theta'}.\]
The function \(\theta \mapsto d_Q(\theta) = Q_L^2(\theta) - Q_L^2(\theta_0)\) (which has a unique minimum in \(\theta_0\) by (i) and which is twice differentiable on \(B_\kappa(\theta_0)\) by (iii)) is bounded away from zero on \(\Theta \setminus B _{\kappa}(\theta_0)\), has invertible Hessian \(V_{\theta_0} \in \mathbb{R}^{s \times s}\) at \(\theta_0\) and satisfies \[\forall \theta \in B_{\kappa}(\theta_0): \quad d_Q(\theta) \ge \frac{\lambda_{\min}(V_{\theta_0})}{4} \| \theta- \theta_0\|^2_2.\]
Parts (ii)–(iv) are standard smoothness assumptions. Under sufficient smoothness, Parts (i) and (v) are essentially equivalent to requiring that \(\theta_0\) be the unique well-separated minimizer of \(Q_L\) (and thus of \(d_Q\)); this follows from a standard Taylor expansion of \(d_Q\) around \(\theta_0\). In the present assumption, this property is formulated in a more quantitative manner to facilitate later arguments. Finally, note that we do not assume that \(Q_L(\theta_0)=0\). Consequently, the subsequent result also applies to misspecified models, that is, to situations where \(L \notin \{L(\cdot;\theta): \theta \in \Theta\}\).
Before providing a linear representation for \(\hat{\theta}_n- \theta_0\), we need to introduce some additional notation. Denote by \(J_\theta \in \mathbb{R}^{q \times s}\) the Jacobian matrix of \(\boldsymbol{\varphi}\) evaluated at \(\theta\). Let \(V_{n, \theta}\) denote the Hessian matrix of the map \(\theta \mapsto Q_n^2(\theta)\) evaluated at \(\theta\). Let \(\partial_j \widetilde{L}(\boldsymbol{x})\) denote the partial derivative of \(L\) where it exists and the right-side directional partial derivative with respect to \(x_j\) otherwise; note that the right-hand partial derivative always exists by convexity of \(L\). For \(i \in [n]\), define \[\label{eq:defZjn} Z_{i,n} := 2V_{\theta_0}^{-1} J_{\theta_0}^\top \int_{[0,1]^d} \Big\{\boldsymbol{1}\Big( \exists j \in [d]: V_{ij} < \frac{k}{n}x_j \Big) - \sum_{j=1}^d\partial_j \widetilde{L}(\boldsymbol{x}) \boldsymbol{1}\Big( V_{ij} < \frac{k}{n}x_j \Big) \Big\} \boldsymbol{g}(\boldsymbol{x}) \, \mathrm d\mu(\boldsymbol{x})\tag{23}\] and note that \(Z_{1,n}, \dots, Z_{n,n}\) are iid. Finally, note that Assumption 1 implies that the constants \[\begin{align} \label{eq:C95partial} C_\partial &:= \max_{j \in [s], p \in [q]}\sup_{\theta \in B_\kappa(\theta_0)} |\partial_j \varphi_p(\theta)| , \qquad C_{\partial^2} := \max_{j,\ell \in [s], p \in [q]} \sup_{\theta \in B_\kappa(\theta_0)} |\partial_{j\ell} \varphi_p(\theta)|, \end{align}\tag{24}\] are finite, while the following two constants are positive: \[\begin{align} \label{eq:lambda-min} C_{V} := \lambda_{\min}(V_{\theta_0}), \qquad C_Q := \inf_{\theta \in \Theta \setminus B_{\kappa}(\theta_0)} d_Q(\theta). \end{align}\tag{25}\]
Theorem 12. Let \(L\) be a \(d\)-variate STDF satisfying [cond:smoothness-good], and assume that the tuple \((L, \{L(\cdot; \theta): \theta \in \Theta\}, \boldsymbol{g}, \mu)\) satisfies Assumption 1. Then, there exist constants \(D_1, D_2>0\) only depending on \(d\) and \(K_L\) (from Theorem 3) and \(\tilde{C}_\beta, \tilde{C}_\eta\in (0,1], \tilde{C}_{r1}, \tilde{C}_{r2}>0\) only depending on \(d,s,q\), the constant \(C_g\) from 22 , the three parameters \(\kappa,C_{h},\gamma_{h}\) from Assumption 1 and the four constants defined in 24 and 25 such that, for any \(n \in \mathbb{N}, k \in [n], \delta \in (0,e^{-1})\) satisfying \(\log(d/\delta)\le 2k/7\), \(C_{s}r \le 1\) and \(n/k \ge 2\) with \(C_{s}\) the universal constant from Lemma 11, the following holds with probability at least \(1- 7(d+1) \delta\):
If \(\eta\in (0, \tilde{C}_\eta)\) and if \[\begin{align} \zeta_{n,1} := k^{-1/2} \sup_{\boldsymbol{x} \in [0,2]^d} |B_n(\boldsymbol{x})| +(C_{s}+188\sqrt 2/3)\cdot d r \le \tilde{C}_\beta \end{align}\] with \(B_n\) from 14 and with \(r = \sqrt{ k^{-1}\log(1/\delta)}\) as in 13 , then \[\sqrt{k} \big( \hat{\theta}_n- \theta_0 \big) = \frac{1}{\sqrt{k}} \sum_{i=1}^n \big( Z_{i,n} - \mathbb{E}[Z_{i,n}] \big) + \boldsymbol{r}_{n,1} + \boldsymbol{r}_{n,2}\] where \(\|\boldsymbol{r}_{n,1} \|_2^2 \le \tilde{C}_{r1} k (\zeta_{n,1}^{2+\gamma_h} + \eta)\) and \[\begin{align} \|\boldsymbol{r}_{n,2}\|_2 &\le \tilde{C}_{r2}\Big(\sup_{\boldsymbol{x} \in [0,2]^d} |B_n(\boldsymbol{x})| + \frac{d}{\sqrt{k}} + D_{1} \sqrt{r\log\Big(\frac{D_{2}}{\delta r}\Big)} \\&+ \sqrt k \zeta_{n,1}\int_{\mathfrak B^{\oplus C_{s}r}} \|\boldsymbol{g}(\boldsymbol{x})\|_2\, \mathrm d\mu(\boldsymbol{x})\Big). \end{align}\]
Theorem 12 can be combined with the central limit theorem to yield an alternative proof of Theorem 4.2 in [32] for approximate, rather than exact, M-estimators, albeit under stronger smoothness assumptions on \(L\). In contrast to [32], which establishes only weak convergence, our result additionally provides non-asymptotic remainder bounds with explicit rates. We further note that the part of the remainder term involving the integral over \(\mathfrak B^{\oplus C_{s}r}\) can be small even in irregular models. For instance, for the factor models from Example 2, a straightforward computation shows that the Lebesgue measure of \(\mathfrak B^{\oplus C_{s}r} \cap [0,1]^d\) is bounded by a constant multiple of \(r\). For functions \(g\) with uniformly bounded norm and for \(\mu\) corresponding to Lebesgue measure, we thus have \[\int_{\mathfrak B^{\oplus C_{s}r}} \|\boldsymbol{g}(\boldsymbol{x})\|_2\, \mathrm d\mu(\boldsymbol{x}) \lesssim r.\]
Similarly to Corollaries 1 and 2, Theorem 12 can be combined with the union bound to obtain uniform linearizations for collections of M-estimators based on lower-dimensional margins, where the number of estimators may grow at a rate of the form \(\exp(k^a)\) for sufficiently small \(a\). Such settings naturally arise when a multivariate tail dependence model is characterized through parametric bivariate dependencies only, as is the case for the Hüsler–Reiss model from Example 1. Moreover, in the same framework, the result also provides the basis for a high-dimensional Gaussian approximation analogous to the results derived in Section 4. We conclude this section by establishing such a result.
Specifically, let \(\mathcal{I}\) be a collection of index sets \(I \subseteq[d]\) with \(|I| \ge 2\), and write \(m=\max_{I \in \mathcal{I}} |I|\). For each \(I \in \mathcal{I}\), let \(\{L_I(\cdot; \theta^I) \colon \theta^I \in \Theta^I\}\) be a parametric family of STDFs, with a parameter space \(\Theta^I \subseteq \mathbb{R}^{s^I}\); note that \(\Theta^I\) has a different meaning than the notation \(A^I\) for \(A \subseteq[-\infty, \infty]\) introduced in Section 1.1. Let \[\begin{align} Q_n^I(\theta^I) \coloneq \Big\|\int_{[0,1]^{I}} \boldsymbol{g}^I(\boldsymbol{x}_I)\big(L_I(\boldsymbol{x}_I; \theta^I)-\hat{L}_{n,I}(\boldsymbol{x}_I)\big) \,\mathrm d\mu^I(\boldsymbol{x}_I)\Big\|_2 \end{align}\] for known measures \(\mu^I\) on \([0,1]^I\) and known functions \(\boldsymbol{g}^I \colon [0,1]^I \to \mathbb{R}^{q^I}\) with \(q^I \in \mathbb{N}_{\geq s^I}\). Suppose that \(C_g^I := \int_{[0,1]^{I}} \|\boldsymbol{g}^I\|_2 \,\mathrm d\mu^I < \infty\) for any \(I \in \mathcal{I}\). Likewise, define \[\begin{align} Q_{L_I}^I(\theta^I) \mathrel{\vcenter{:}}= \Big\|\int_{[0,1]^{I}} \boldsymbol{g}^I(\boldsymbol{x}_I)\big( L_I(\boldsymbol{x}_I;\theta^I)-L_I(\boldsymbol{x}_I) \big) \, \mathrm d\mu^I(\boldsymbol{x}_I)\Big\|_2. \end{align}\] Let \(\hat{\theta}_n^I\) be an approximate minimizer of \(Q_n^I\) in the sense that for some \(\eta>0\), \[\begin{align} Q_n^I(\hat{\theta}_n^I)-\inf_{\theta^I \in \Theta^I} Q_n^I(\theta^I)< \eta. \end{align}\]
Assumption 2. There exist constants \(\kappa>0\) and \(C_{h}>0\) such that, for each \(I \in \mathcal{I}\) the tuple \((L_I, \{L_I(\cdot; \theta^I): \theta^I \in \Theta^I\}, \boldsymbol{g}^I, \mu^I)\) satisfies the conditions (i) - (v) from Assumption 1 with \(\gamma_{h}= 1\) and with \(\theta_0 = \theta_0^I\), \(\boldsymbol{\varphi} = \boldsymbol{\varphi}^I\), \(d_Q=d_Q^I\) and \(V_{\theta_0} = V_{\theta_0^I}\).
Write \(C_\partial^{I}, C_\partial^{I}, C_{V}^I\) and \(C_Q^I\) for the constants in 24 and 25 when applied for \((L^I, \{L^I(\cdot; \theta^I): \theta^I \in \Theta^I\}, \boldsymbol{g}^I, \mu^I)\), and let \[\begin{align} \label{eq:constants-Ic} C_\partial^{\mathcal{I}} = \max_{I \in \mathcal{I}} C_\partial^{I}, \quad C_{\partial^2}^{\mathcal{I}} = \max_{I \in \mathcal{I}} C_{\partial^2}^{I}, \quad C_g^{\mathcal{I}} = \max_{I \in \mathcal{I}} C_g^{I}, \quad C_{V}^{\mathcal{I}} = \min_{I \in \mathcal{I}} C_V^I, \quad C_{Q}^{\mathcal{I}} = \min_{I \in \mathcal{I}} C_Q^I. \end{align}\tag{26}\] Let \(J_{\theta^I} \in \mathbb{R}^{q^I \times s^I}\) denote the Jacobian matrix of \(\boldsymbol{\varphi}^I\) evaluated at \(\theta^I\). Let \(\partial_j \widetilde{L}_I(\boldsymbol{x}_I)\) denote the partial derivative of \(L_I\) where it exists and the right-side directional partial derivative with respect to \(x_j\) otherwise; note that the right-hand partial derivative always exists by convexity of \(L_I\). For each \(i \in [n]\) and \(I \in \mathcal{I}\), let \[\begin{align} \boldsymbol{A}_{i,n}^I &= \label{eq:def-Ani} \int_{[0,1]^{I}} \Big\{\mathbf{1}\Big(\exists j\in I:\; V_{ij}<\frac{k}{n}x_j\Big) - \sum_{j\in I} \partial_j \widetilde{L}_I(\boldsymbol{x}_I)\, \mathbf{1}\Big(V_{ij}<\frac{k}{n}x_j\Big)\Big\} \boldsymbol{g}^I(\boldsymbol{x}_I)\,\mathrm d\mu^I(\boldsymbol{x}_I) \\ \boldsymbol{Z}_{i,n}^I &= 2 \,\big(V_{\theta_0^I}\big)^{-1} J_{\theta_0^I}^{\!\top} \boldsymbol{A}_{i,n}^I \end{align}\tag{27}\] For \(t \in [s^I]\), let \(Z_{i,n}^{I,t}\) be the \(t\)-th component of the random vector \(\boldsymbol{Z}_{i,n}^I\). Define \[B_n^I(\boldsymbol{x}) := \sqrt{k}\{\widetilde{\mu}_{n,I}(\boldsymbol{x}) - L_I(\boldsymbol{x})\}\] and, with \(s= \sum_{I \in \mathcal{I}} s^I\), let \[\begin{align} \boldsymbol{S}_n = \big(\boldsymbol{S}_n^I\big)_{I \in \mathcal{I}} = \big(\sqrt{k}(\hat{\theta}^I_n-\theta_0^I) \big)_{I \in \mathcal{I}} \in \mathbb{R}^{s} \end{align}\] and \[\begin{align} \boldsymbol{T}_n = \big(\boldsymbol{T}_n^I\big)_{I \in \mathcal{I}} = \bigg(\frac{1}{\sqrt{k}} \sum_{i=1}^n \big(Z_{i,n}^I - \mathbb{E}[Z_{i,n}^I]\big) \bigg)_{I\in \mathcal{I}} \in \mathbb{R}^{s}. \end{align}\] Let \(\Sigma_n \coloneq \text{Var}(\boldsymbol{T}_n)\) and \(\boldsymbol{G}_n \sim \mathcal{N}_{s}(\boldsymbol{0}, \Sigma_n)\).
Theorem 13. Suppose that, for each \(I\in \mathcal{I}\), \(\boldsymbol{X}_I\) has STDF \(L_I\) satisfying [cond:smoothness-good]; denote the respective set of bad points by \(\mathfrak B_{I}\) and the constants by \(K_{L_I}\). Suppose Assumption 2 holds. Moreover, assume that \(m|\mathcal{I}| \geq 3, k \geq 2, s \geq 3, n/k \ge 2\) and that
\(\sigma_{\min}^2 \coloneq \min_{i\in [s]} (\Sigma_n)_{ii}>0\),
\(\log(m^2 |\mathcal{I}|k^{1/4} ) \le 2k/{7}\),
\(\log(m |\mathcal{I}|k^{1/4} ) \le k/{C_{s}^2}\),
with \(C_{s}\) the universal constant from Lemma 11. Then there exist constants \(\tilde{C}_{k}^{\mathcal{I}},\tilde{C}_\eta^{\mathcal{I}}>0\) and \(\tilde{C}_N>0\) only depending on \(m = \max_{I \in \mathcal{I}} m^I, s^{\mathcal{I}} = \max_{I \in \mathcal{I}} s^I, q^{\mathcal{I}} = \max_{I \in \mathcal{I}} q^I, K_L^{\mathcal{I}} = \max_{I \in \mathcal{I}} K_{L_I}\) and \(\kappa, C_{h}\) and the constants in 26 such that, for any \(\eta \le \tilde{C}_\eta^{\mathcal{I}}\) and any \(k\) that additionally satisfies \(k^{-1}\log(m |\mathcal{I}|k^{1/4} ) \le \tilde{C}_{k}^{\mathcal{I}}\), we have \[\begin{align} d_K(\boldsymbol{S}_n, \boldsymbol{G}_n) & \le \tilde{C}_N \Big\{ \sqrt{k\eta\log s} + \sqrt{\log s} B_{n,k}^{\mathcal{I}} + \Big(\frac{\log^5(sn)}{k}\Big)^{1/4} \\&+ \log(sk) \max_{I \in \mathcal{I}} \int_{\mathfrak B_I^{\oplus \zeta_n}} \|{\boldsymbol{g}}^I(\boldsymbol{x}_I)\|_2 \, \mathrm d\mu^I(\boldsymbol{x}_I) \Big\}, \end{align}\] where \[B_{n,k}^{\mathcal{I}} = \max_{I \in \mathcal{I}} \sup_{\boldsymbol{x}_I \in [0,2]^{I}} \left|B_n^I(\boldsymbol{x}_I)\right|, \qquad \zeta_n := C_{s}\sqrt{1+\log m}\sqrt{\frac{\log (sk)}{k}}\] If \(\mathfrak B_I = \emptyset\) and if \(\eta\) is additionally chosen smaller than \(k^{-3/2}\), the bound simplifies: \[\begin{align} d_K(\boldsymbol{S}_n, \boldsymbol{G}_n) & \le \tilde{C}_N \Big\{ \sqrt{\log s} B_{n,k}^{\mathcal{I}} + \Big(\frac{\log^5(sn)}{k}\Big)^{1/4} \Big \}, \end{align}\] which is analogous to the bound in Theorem 7.
As noted earlier, Theorem 13 is useful for studying high-dimensional parametric models whose parameter is identifiable from lower-dimensional margins. A prominent example is the Hüsler–Reiss model from Example 1, for which the full parameter is identifiable from the collection of bivariate margins. As discussed after Lemma 1, each bivariate Hüsler–Reiss margin with \(\gamma_{j\ell}>0\) satisfies Condition [cond:smoothness-good] with \(\mathfrak B_{j\ell}=\emptyset\).
Suppose \(\mathbb{X}=\{ X(\boldsymbol{s}): \boldsymbol{x} \in \mathcal{S}\}\) is a random field indexed by a spatial domain \(\mathcal{S} \subseteq\mathbb{R}^2\); for instance, \(X(\boldsymbol{s})\) could correspond to daily maximal wind speed at location \(\boldsymbol{s}\) during a winter day. We assume that, for each pair \((\boldsymbol{s}, \boldsymbol{s}')\), the stable tail dependence function \(L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}\) of \((X(\boldsymbol{s}_1), X(\boldsymbol{s}_2))\) exists. (Bivariate) extremal isotropy refers to the assumption that \(L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}\) depends on \(\boldsymbol{s}_1, \boldsymbol{s}_2\) only through the spatial domain distance \(\| \boldsymbol{s}_1 - \boldsymbol{s}_2\|_2\); an assumption that is met for many max-stable models like the Smith model [53] or Schlather’s model [54]. In this section, we illustrate how the assumption can be tested (non-parametrically) based on repeated observations of \(\mathbb{X}\) at a finite set of locations \(\mathcal{S}_d=\{\boldsymbol{s}_1, \dots, \boldsymbol{s}_d\}\). In the non-extreme world, tests for isotropy are used routinely for model building and diagnostics [55].
More formally, let \(\mathcal{P}_d = \{ (\boldsymbol{s}_1, \boldsymbol{s}_2) \in \mathcal{S}_d \times \mathcal{S}_d: \boldsymbol{s}_1 \ne \boldsymbol{s}_2\}\) denote the set of (ordered) pairs of unequal locations, with \(|\mathcal{P}_d | =d^2(d^2-1)\). For a given spatial distance \(\rho >0\), let \[\mathcal{P}_d(\rho) = \{ (\boldsymbol{s}_1, \boldsymbol{s}_2) \in \mathcal{P}_d: \| \boldsymbol{s}_1 - \boldsymbol{s}_2 \|_2 = \rho\}\] denote the set of (ordered) pairs of locations whose Euclidean distance is \(\rho\); note that \(\mathcal{P}_d(\rho)\) is non-empty for a finite set of distances only. For such a distance, consider the null hypothesis of extremal isotropy at spatial distance \(\rho\) defined as \[H(\rho): \quad L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)} = L_{(\boldsymbol{s}_1', \boldsymbol{s}_2')} \quad \text{ for all } \quad (\boldsymbol{s}_1,\boldsymbol{s}_2), (\boldsymbol{s}_1',\boldsymbol{s}_2') \in \mathcal{P}_d(\rho);\] note each equality in the hypothesis essentially corresponds to the hypothesis considered in Section 4.2 in [36]. The intersection hypothesis \(H = \bigcap_{\rho >0} H(\rho)\) then corresponds to (bivariate) extremal isotropy.
In the following, and for simplicity, we restrict ourselves to the case of gridded observations on a rectangular domain; without loss of generality, \(\mathcal{S}_d = \{1, \dots, d\}^2\). In that case, \(|\mathcal{P}_d(1)| = 4d(d-1)\), \(|\mathcal{P}_d(\sqrt 2)| = 4(d-1)^2\), and so on. We will concentrate on testing for \(H(\rho)\) for \(\rho \in \{1, \sqrt 2\}\) only, and illustrate how these tests can be combined to test for the intersection hypothesis \(H(1, \sqrt 2) := H(1) \cap H(\sqrt 2)\). The resulting combination test can be interpreted as a test for extremal isotropy that is able to detect non-isotropic behavior for ‘small’ distances (\(\rho \le \sqrt 2)\).
A natural test statistic for \(H(\rho)(L)\) is given by \[\widetilde{T}_n^{(\rho)} = \max_{(\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1',\boldsymbol{s}_2') \in \mathcal{P}_d(\rho)} \sup_{t \in [0,1] } \sqrt k\big\{ \hat{L}_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(1-t,t) - \hat{L}_{(\boldsymbol{s}_1', \boldsymbol{s}_2')}(1-t,t) \big\},\] where \(\hat{L}_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}\) denotes the empirical STDF corresponding to the bivariate sample \((X_i(\boldsymbol{s}_1), X_i( \boldsymbol{s}_2))_{i \in [n]}\) and where we restrict attention to evaluation points \((1-t,t)\) since the population counterparts \(L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}\) are uniquely determined by their restriction to the unit simplex. To reduce the computational complexity, we further approximate the supremum by a finite maximum, and consider \[\begin{align} T_n^{(\rho)} &= \sqrt k \max_{(\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1',\boldsymbol{s}_2') \in \mathcal{P}_d(\rho)} \max_{t \in A} \big| \hat{L}_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(1-t,t) - \hat{L}_{(\boldsymbol{s}_1', \boldsymbol{s}_2')}(1-t,t)\big| \\ &= \sqrt k \max_{t \in A} \Big\{ \max_{(\boldsymbol{s}_1, \boldsymbol{s}_2)\in \mathcal{P}_d(\rho)} \hat{L}_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(1-t,t) - \min_{(\boldsymbol{s}_1', \boldsymbol{s}_2')\in \mathcal{P}_d(\rho)} \hat{L}_{(\boldsymbol{s}_1', \boldsymbol{s}_2')}(1-t,t) \Big\} \end{align}\] instead, where \(A=\{1/12, 2/12, \dots, 11/12\}\). Bootstrap versions of this statistic can be obtained as in Section 4. Specifically, as in 20 , for some small positive bandwidth parameter \(h\), let \[\begin{align} \label{eq:hatY-isotropy} \hat{Y}_{i,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2) &= \nonumber \frac{1}{\sqrt k} \Big[ \boldsymbol{1}(\exists j \in [2]: \hat{V}_{i}(\boldsymbol{s}_j) < kx_j/n) - (k/n) \widehat L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(\boldsymbol{x}) \\ & -\sum_{j \in [2]}\widehat{\partial_{j} L}_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2) \big\{\boldsymbol{1}( \hat{V}_{i}(\boldsymbol{s}_j) < kx_j/n) - kx_j/n \big\} \Big], \end{align}\tag{28}\] and with iid standard normal multipliers \(e_1, \dots, e_n\) define \[T_{n,b}^{(\rho),*} = \max_{(\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1',\boldsymbol{s}_2') \in \mathcal{P}_d(\rho)} \max_{t \in A} \sum_{i=1}^n e_{i} \big\{ \hat{Y}_{i,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(1-t,t) - \hat{Y}_{i,(\boldsymbol{s}_1', \boldsymbol{s}_2')}(1-t,t) \big\}.\] Theorem 11 suggests that, under the null hypothesis \(H(\rho)\), the distribution of \(T_n^{(\rho)}\) can be approximated by the conditional distribution of \(T_{n}^{(\rho),*}\) given the data. Under fixed alternatives, however, \(T_n^{(\rho)}\) explodes while the bootstrap can be expected to stay stochastically bounded. Overall, these considerations suggest to reject \(H(\rho)\) if the \(p\)-value \[\hat{p}_{n}^{(\rho)} = 1-F_{n}^{(\rho),*}(T_n^{(\rho)} )\] is smaller than the nominal level \(\alpha\); here, \(F_{n}^{(\rho),*}\) denotes the conditional cdf of \(T_n^{(\rho),*}\) given the data (in practice, the latter can be approximated using repeated simulation of \(T_n^{(\rho),*}\)).
The intersection hypothesis \(H(1, \sqrt 2) := H(1) \cap H(\sqrt 2)\) can be tested using the approach described in [56]. More specifically, let \[C_n = \hat{p}_{n}^{(1)} \wedge \hat{p}_{n}^{(\sqrt 2)},\] and note that small values of \(C_n\) provide evidence against the intersection hypothesis. Critical values will be obtained using the bootstrap analogue \[C_n^* = \hat{p}_{n}^{(1),*} \wedge \hat{p}_{n}^{(\sqrt 2),*}, \qquad \hat{p}_{n}^{(\rho), *} =1-F_{n}^{(\rho),*}(T_n^{(\rho),*} ),\] where it is crucial that the bootstrap expressions with \(\rho \in \{1, \sqrt 2\}\) are based on the same multipliers. Specifically, if \(\hat{q}_{n,\alpha}^{*}\) denotes the conditional \(\alpha\)-quantile of \(C_n^*\), given the data (again, the latter can be approximated by repeated simulation) we propose to reject \(H(1,\sqrt 2)\) if \(C_n \le \hat{q}_{n,\alpha}^{*}\).
The following result yields (approximate) finite-sample level control. Let \(\mathcal{P}_d(1, \sqrt 2) = \mathcal{P}_d(1) \cup \mathcal{P}_d(\sqrt 2)\) and \(D(1,\sqrt2 ) = D(1) \cup D(\sqrt 2)\) with \(p = |D(1,\sqrt2 )|\), where \[\begin{align} \label{eq:Drho} D(\rho) = \Big\{ (t, \boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2') : t \in A, (\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1', \boldsymbol{s}_2') \in \mathcal{P}_d(\rho) , (\boldsymbol{s}_1, \boldsymbol{s}_2) \ne (\boldsymbol{s}_1', \boldsymbol{s}_2')\Big\}. \end{align}\tag{29}\] Further, writing \(V(\boldsymbol{s}) = 1-F_{\boldsymbol{s}}(X(\boldsymbol{s}))\) with \(F_{\boldsymbol{s}}\) the cdf of \(X(\boldsymbol{s})\), introduce the bias term \[\begin{align} \label{eq:bias-isotropy} B_{n, (\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2) = \sqrt k \Big\{ \frac{n}{k} \mathbb{P}\Big( \exists j \in \{1,2\}: V(\boldsymbol{s}_j) \le \frac{k}{n} x_j\Big) - L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2) \Big\}. \end{align}\tag{30}\]
Theorem 14. Suppose \(H(1,\sqrt 2)\) holds and that there exist \(\kappa_L, K_L\in(0,\infty)\) such that \[\begin{align} &\forall (\boldsymbol{s}_1, \boldsymbol{s}_2) \in \mathcal{P}_d(1, \sqrt 2), \forall \boldsymbol{x} \in A(1 \wedge (\kappa_L/2)), \boldsymbol{y} \in [0,\infty)^2 \text{ with } \| \boldsymbol{x}- \boldsymbol{y}\|_\infty \le \kappa_L, \forall j \in [2]: \\ & \partial_{j} L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(\boldsymbol{x}), \partial_{j} L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(\boldsymbol{y}) \text{ exist and satisfy } \\ & |\partial_{j} L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(\boldsymbol{x})-\partial_{j} L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(\boldsymbol{y})| \le K_L\|\boldsymbol{x} - \boldsymbol{y}\|_\infty, \end{align}\] where \(A(\kappa) = \{(1-t,t): t \in A\}^{\oplus \kappa}\). Moreover, assume that \(n/k \ge 2, |\mathcal{P}_d(1, \sqrt 2)|\ge 3\) and
\(\sigma_{\min}^2 := \min_{(t,\boldsymbol{s}_1, \boldsymbol{s}_2,\boldsymbol{s}_1',\boldsymbol{s}_2') \in D(1,\sqrt2 )} \operatorname{Var}( \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{D}} \endgroup _{n, (\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1', \boldsymbol{s}_2')}(t) )>0\) with \(\begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{D}} \endgroup _{n}\) from 74 .
\(\log(2 |\mathcal{P}_d(1, \sqrt 2)|k^{1/4} ) \le 2k/7\).
\(\log(|\mathcal{P}_d(1, \sqrt 2)| k^{1/4}) \le \kappa_L^2 k / (8C_{s}^2)\) with \(C_{s}\) from Lemma 11,
Let \(0<c_h< c_h'< \infty\) be constants, and assume that the bandwidth \(h<(\min_{t\in A} (1-t) \wedge t) \wedge (\kappa_L/2)\) satisfies \[c_h\Big( \frac{\log(p+k)}{k}\Big)^{1/2} \le h \le c_h' \Big( \frac{\log(p+k)}{k}\Big)^{1/4}.\] There exists a constant \(c_0\) depending on \(K_L, \sigma_{\min}^2c_h, c_h'\) only such that \[\Big|\mathbb{P}( C_n \le \hat{q}_{n,\alpha}^* ) - \alpha \Big| \le c_0 \Big[ \sqrt{\log(p+k)}\, B_{n,k} + \Big( \frac{\log^5(pn)}{k} \Big)^{1/4} \Big],\] where \(B_{n,k}= \max_{\rho \in \{1, \sqrt 2\},(\boldsymbol{s}_1, \boldsymbol{s}_2) \in \mathcal{P}_d(\rho),(x_1, x_2) \in A(\kappa_L)} |B_{n, (\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2)|\) with \(B_{n, (\boldsymbol{s}_1, \boldsymbol{s}_2)}\) from 30 .
We end this section by illustrating the performance of the above tests in a small simulation study. For that purpose, we consider data generated from the max-stable Brown-Resnick random field [57], whose bivariate STDF at location pair \((\boldsymbol{s}_1, \boldsymbol{s}_2)\) is given by that of the bivariate Hüsler-Reiss distribution from Example 1, i.e., \[L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1,x_2) = x_1 \Phi\left( \frac{a}{2} + \frac{1}{a}\log\left(\frac{x_1}{x_2}\right)\right) + x_2\Phi\left( \frac{a}{2} + \frac{1}{a}\log\left(\frac{x_2}{x_1}\right)\right),\] where \(\Phi\) denotes the c.d.f. of standard normal distribution and where \[a^2 = \gamma_{\xi, \beta}(\boldsymbol{s}_1, \boldsymbol{s}_2) = \beta \Big[ (\boldsymbol{s}_1 - \boldsymbol{s}_2)^\top \Sigma^{-1} (\boldsymbol{s}_1 - \boldsymbol{s}_2) \Big]^{\xi/2}\] for some \(\Sigma \in \mathbb{R}^{2\times 2}\) positive definite and parameters \(\beta>0\) and \(\xi \in (0,2]\). Note that the respective extremal coefficients are given by \[\chi(\boldsymbol{s}_1, \boldsymbol{s}_2) = 2-L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(1,1) = 2 - 2 \Phi\Big( \frac{\gamma_{\xi, \beta}(\boldsymbol{s}_1, \boldsymbol{s}_2)^{1/2}}{2}\Big).\] For the simulation study, we consider the choices \(\xi \in \{0.9,1.8\}, \beta = 0.5\) and covariance matrices \[\Sigma_1 = \begin{pmatrix} 1 & 0 \\ 0 & 1 \end{pmatrix} \text{ (isotropic)}, \qquad \Sigma_2 = \begin{pmatrix} 0.5 & 0.25 \\ 0.25 & 1 \end{pmatrix} \text{ (anisotropic)},\] The resulting extremal coefficients only depend on the (linear span of the) spatial lag \(\boldsymbol{\rho}=\boldsymbol{s}_1 - \boldsymbol{s}_2\); they are explicitly provided in Table 1 for the case where \(\| \boldsymbol{\rho} \|_2 \in \{1, \sqrt 2\}\).
| \(\Sigma\) | \(\xi\) | hor | vert | dia1 | dia2 |
|---|---|---|---|---|---|
| \(\Sigma_1\) | 0.9 | 0.72 | 0.68 | ||
| 1.8 | 0.72 | 0.63 | |||
| \(\Sigma_2\) | 0.9 | 0.67 | 0.72 | 0.67 | 0.62 |
| 1.8 | 0.61 | 0.71 | 0.61 | 0.48 | |
For the simulation study, we consider a sample size of \(n = 10^4\) and a spatial grid \(\mathcal{S}_{10} = [10]^2\). The number of equations to be tested for the hypothesis \(H(\rho)\) is \(\binom{|\mathcal{P}_d(1)|}{2}=360\cdot 359/2= 64\, 620\) for \(\rho = 1\) and \(52\,326\) for \(\rho = \sqrt{2}\), yielding a total of \(116\,946\) equations for the combined intersection hypothesis. For each parameter configuration, we generate 200 datasets and evaluate the three tests corresponding to \(H(1)\), \(H(\sqrt 2)\), and \(H(1, \sqrt 2)\). In each case, we employ \(B = 500\) bootstrap replications and consider threshold parameters \(k \in \{200, 350, 500\}\). The results are summarized in Table 2, which reports rejection frequencies at significance level \(0.05\). The findings are consistent with theoretical expectations: all tests maintain the nominal level. Moreover, the power increases from \(H(1)\) to \(H(\sqrt 2)\) to \(H(1, \sqrt 2)\) and is also increasing in \(\xi\).
| \(\Sigma_1\) | \(\Sigma_2\) | ||||||
|---|---|---|---|---|---|---|---|
| 3-5 (lr)6-8 \(\xi\) | \(\Xi\) | \(k=200\) | \(k=350\) | \(k=500\) | \(k=200\) | \(k=350\) | \(k=500\) |
| 0.9 | \(1\) | 2.5 | 4.0 | 4.5 | 22.5 | 50.0 | 75.5 |
| \(\sqrt{2}\) | 1.5 | 4.5 | 5.5 | 29.5 | 62.0 | 85.0 | |
| \(1, \sqrt{2}\) | 1.0 | 2.0 | 4.5 | 31.0 | 73.0 | 88.0 | |
| 1.8 | \(1\) | 3.0 | 3.0 | 4.0 | 90.5 | 100.0 | 100.0 |
| \(\sqrt{2}\) | 3.5 | 5.0 | 4.0 | 99.0 | 100.0 | 100.0 | |
| \(1, \sqrt{2}\) | 4.0 | 3.0 | 4.0 | 99.0 | 100.0 | 100.0 | |
AB and KE were supported by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation; Project-ID 520388526; TRR 391: Spatio-temporal Statistics for the Transition of Energy and Transport) which is gratefully acknowledged. YC and SV gratefully acknowledge support by a Discovery Grant from the Natural Sciences and Engineering Research Council of Canada (NSERC, grant RGPIN-2024-05528). Calculations for this publication were performed on the HPC cluster Elysium of the Ruhr University Bochum, subsidised by the DFG (INST 213/1055-1).
SUPPLEMENT TO THE PAPER:
“EMPIRICAL TAIL DEPENDENCE FUNCTIONS IN HIGH DIMENSIONS: UNIFORM LINEARIZATIONS AND INFERENCE”
By Axel Bücher, Yeonjoon Choi, Katharina Effertz and Stanislav Volgushev
Ruhr-Universität Bochum and University of Toronto
Figure 1:
.
Proof of Lemma 1. Since \(L(x_1, x_2) = (x_1+x_2) A(x_2/(x_1+x_2))\) for all \(\boldsymbol{x} =(x_1, x_2) \in [0,\infty)^d\) such that \(x_1 + x_2>0\), we have, for \(\boldsymbol{x} \in (0,\infty)^2\), \[\begin{align} \partial_{1} L(x_1, x_2) &= A\Big(\frac{x_2}{x_1+x_2}\Big) -\frac{x_2}{x_1+x_2}A'\Big(\frac{x_2}{x_1+x_2}\Big), \\ \partial_{2} L(x_1, x_2) &= A\Big(\frac{x_2}{x_1+x_2}\Big) + \frac{x_1}{x_1+x_2}A'\Big(\frac{x_2}{x_1+x_2}\Big). \end{align}\] Moreover, \(\partial_{1} L(x_1, 0) = \partial_{2} L(0,x_2) = 1\) for \(x_1,x_2>0\). Continuity of \(\partial_{1} L\) on \((0,\infty)^2\) is immediate. Further, for a sequence \(\boldsymbol{x}_n\) in \((0,\infty)^2\) converging to \(\boldsymbol{x}=(x_1, 0)\) with \(x_1>0\), we have \(\lim_{n\to\infty}x_{n2}/(x_{n1}+x_{n2})=0\), which implies \(\lim_{n\to\infty} \partial_{1} L(\boldsymbol{x}_n) = A(1)=1=\partial_{1} L(\boldsymbol{x})\) by continuity of \(A\) on \([0,1]\) and boundedness of \(A'\) on \((0,1)\). Hence, \(\partial_{1} L\) is continuous on \(E_1\), and the same arguments show continuity of \(\partial_{2} L\) on \(E_2\). Regarding the second-order partial derivatives, note that, for \(\boldsymbol{x} \in (0,\infty)^2\), \[\begin{align} \partial_{11} L (x_1,x_2) &= \frac{x_2^2}{(x_1+x_2)^3}A''\Big(\frac{x_2}{x_1+x_2}\Big) = \frac{t^2 A''(t)}{x_1+x_2} \\ \partial_{22} L (x_1,x_2) &= \frac{x_1^2}{(x_1+x_2)^3}A''\Big(\frac{x_2}{x_1+x_2}\Big) = \frac{(1-t)^2 A''(t)}{x_1+x_2} \\ \partial_{12} L(x_1,x_2) &= -\frac{x_1x_2}{(x_1+x_2)^3}A''\Big(\frac{x_2}{x_1+x_2}\Big) = - \frac{t(1-t) A''(t)}{x_1+x_2}, \end{align}\] where we write \(t=x_2/(x_1+x_2)\). Continuity on \((0,\infty)^2\) is immediate. Moreover, \[\begin{align} \frac{t^2 A''(t)}{x_1+x_2} &= t(1-t)A''(t) \frac{x_2}{x_1+x_2} \frac{1}{x}_1 \le A_\infty \frac{1}{x_1} \\ \frac{(1-t)^2 A''(t)}{x_1+x_2} &= t(1-t)A''(t) \frac{x_1}{x_1+x_2} \frac{1}{x}_2 \le A_\infty \frac{1}{x_2} \end{align}\] and \[\begin{align} \Big|-\frac{t(1-t) A''(t)}{x_1+x_2} \Big| \le \frac{A_\infty}{x_1+x_2} \le \frac{A_\infty}{x_1 \vee x_2}, \end{align}\] which finalizes the proof. ◻
Proof of Theorem 2 and Theorem 3. We start by noting that our assumption \(n/k \ge T\) implies that, for any \(\boldsymbol{x} \in [0,T]^d\), we have \(kx_j/n \le 1\) for all \(j \in [d]\). In the subsequent proof, we will only consider such \(\boldsymbol{x}\).
Recall the definition \(\boldsymbol{V}_i=(V_{i1}, \dots, V_{id})^\top\) with \(V_{ij}=1-F_j(X_{ij})\) for \(j\in[d]\) and \(i\in[n]\). Let \(V_{1:n,j} \le V_{2:n, j} \le \dots \le V_{n:n,j}\) denote the order statistics of \(V_{1j}, \dots V_{nj}\), and define \(Q_{nj}(v_j) = V_{\lceil nv_j \rceil:n, j}\) for \(v_j \in (0,1]\), where \(\lceil a \rceil\) denote the smallest integer not smaller than \(a\). For completeness, we define \(Q_{nj}(0)=0\). Note that \(Q_{nj}(v_j) = G_{nj}^\leftarrow(v_j)\) with \(G_{nj}(u_j) = \frac{1}{n} \sum_{i=1}^n \boldsymbol{1}(V_{ij} \le u_j)\) the empirical cdf of \(V_{1j}, \dots, V_{nj}\) and \[\begin{align} \label{eq:generalized-inverse} H^\leftarrow(v) = \inf\{u \in[0,\infty): H(u) \ge v\} \end{align}\tag{31}\] the left-continuous generalized inverse of a non-decreasing function \(H:[0,\infty) \to [0,\infty)\).
Observing that the rank of \(V_{ij}\) among \(V_{1j}, \dots, V_{nj}\) is equal to \(n+1-R_{ij}\), we have \(V_{ij} < V_{\lceil kx_j \rceil:n, j}\) if and only if \(n+1-R_{ij} < \lceil kx_j \rceil\), which in turn is equivalent to \(R_{ij}>n+1-kx_j\) 2. We may therefore write \(\widehat L_n(\boldsymbol{x})=\widetilde{L}_n(S_n(\boldsymbol{x}))\) for \(\boldsymbol{x} \in [0, T]^d\), where \(\widetilde{L}_n\) is from 8 and where \(S_n(\boldsymbol{x}) = (S_{n1}(x_1), \dots, S_{nd}(x_d))^\top\) with \[\begin{align} \label{eq:snj} S_{nj}(x_j) &= \frac{n}{k} Q_{nj}\Big(\frac{k}{n} x_j \Big) = \frac{n}{k} V_{\lceil kx_j \rceil:n, j} \boldsymbol{1}(x_j>0), \qquad j \in [d]. \end{align}\tag{32}\] Further, let \[\begin{align} \label{eq:Lnj-tilde} \widetilde{L}_{nj}(x_j) &:= \widetilde{L}_n(0, \dots, 0, x_j, 0, \dots 0) = \frac{1}{k} \sum_{i=1}^n \boldsymbol{1}\Big(V_{ij} < \frac{k}{n}x_j \Big) \end{align}\tag{33}\] and note that \(\widetilde{L}_{nj}^\leftarrow(x_j) = S_{nj}(x_j)\). Finally, recalling the definition of \(\widetilde{\mu}_n\) from 9 , note that \(\mathbb{E}[\widetilde{L}_n(\boldsymbol{x})]=\widetilde{\mu}_n(\boldsymbol{x})\) and that \(\widetilde{\mu}_{nj}(x_j):= \widetilde{\mu}_n(0, \dots, 0, x_j, 0, \dots 0)\) satisfies \(\widetilde{\mu}_{nj}(x_j)=\widetilde{\mu}^\leftarrow_{nj}(x_j)=x_j\).
The above definitions and identities imply the decomposition \[\begin{align} \label{eq:decomposition1-Lbn} \mathbb{L}_n = \sqrt k (\widehat L_n - L ) &= \nonumber \sqrt k (\widetilde{L}_n \circ S_n - \widetilde{\mu}_n \circ S_n) + \sqrt k (L \circ S_n - L) + \sqrt k (\widetilde{\mu}_n \circ S_n - L \circ S_n) \\&= \widetilde{\mathbb{L}}_n \circ S_n + \sqrt k (L \circ S_n - L) + \sqrt k (\widetilde{\mu}_n - L) \circ S_n. \end{align}\tag{34}\] By Lemma 115 , we have, on an event \(\Omega_0\) with probability at least \(1-(d+1)\delta\), \[\begin{align} \label{eq:Snj-uniform-bound} \max_{j\in[d]} \sup_{x_j \in [0,T]} |S_{nj}(x_j) - x_j| \le C_{s}r(\delta,T,k), \end{align}\tag{35}\] where \(C_{s}\approx 89.18\) is from Lemma 11 and where \(r\) is defined in 13 . Subsequently, we work on this event.
We now distinguish between the two theorems: under the conditions of Theorem 2, we have \(C_{s}r \le \kappa_L\) by our assumption \(r \le \kappa_L/C_{s}\). Hence, for any \(\boldsymbol{x} \in A\), we have \(S_n(\boldsymbol{x}) \in A^{\oplus \kappa_L}\), whence we can apply [cond:smoothness-hoelder] and the mean value theorem to conclude that there exists a (random) \(t^* := t^*_n(\boldsymbol{x}) \in [0,1]\) such that \[\begin{align} \sqrt k \{ L(S_n(\boldsymbol{x})) - L(\boldsymbol{x}) \} &= \sum_{j \in [d]} \partial_{j} L(\boldsymbol{x} + t^*(S_n(\boldsymbol{x}) - \boldsymbol{x}))\sqrt k\{ S_{nj}(x_j) - x_j \}. \end{align}\] Likewise, under the conditions of Theorem 3, for any \(\boldsymbol{x} \in [0,T]^d \setminus (\mathfrak B^{\oplus C_{s}r})\), we have \(S_n(\boldsymbol{x}) \in [0,T+C_{s}r]^d \setminus \mathfrak B\) by 35 , and [cond:smoothness-good] and the mean value theorem allows to conclude that the previous display holds for any \(\boldsymbol{x} \in [0,T]^d \setminus (\mathfrak B^{\oplus C_{s}r})\).
In the following, we either consider \(\boldsymbol{x} \in A\) (Theorem 2), or \(\boldsymbol{x} \in [0,T]^d \setminus (\mathfrak B^{\oplus C_{s}r})\) (Theorem 3). In both cases, the previous display and 34 , together with the definitions \(\begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x}) = \widetilde{\mathbb{L}}_n(\boldsymbol{x}) - \sum_{j=1}^d \partial_{j} L(\boldsymbol{x}) \widetilde{\mathbb{L}}_{nj}(x_j)\) and \(B_n(\boldsymbol{x}) = \sqrt k \{ \widetilde{\mu}_n(\boldsymbol{x}) - L(\boldsymbol{x})\}\), imply the fundamental decomposition \[\begin{align} \label{eq:decomposition2-Lbn} \mathbb{L}_n(\boldsymbol{x}) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n (\boldsymbol{x}) - B_n(S_n(\boldsymbol{x})) = D_{n1}(\boldsymbol{x}) + D_{n2}(\boldsymbol{x}) + D_{n3}(\boldsymbol{x}), \end{align}\tag{36}\] where \[\begin{align} D_{n1}(\boldsymbol{x}) \tag{37} & = \widetilde{\mathbb{L}}_n \circ S_n(\boldsymbol{x}) - \widetilde{\mathbb{L}}_n(\boldsymbol{x}), \\ D_{n2}(\boldsymbol{x}) \tag{38} &= \sum_{j \in [d]} \partial_{j} L\big(\boldsymbol{x} + t^*(S_n(\boldsymbol{x}) - \boldsymbol{x})\big) \big[ \sqrt k\{ S_{nj}(x_j) - x_j \} + \widetilde{\mathbb{L}}_{nj}(x_j) \big] \\ D_{n3}(\boldsymbol{x}) \tag{39} &=\sum_{j \in [d]} \big[ \partial_{j} L(\boldsymbol{x})- \partial_{j} L\big(\boldsymbol{x} + t^*(S_n(\boldsymbol{x}) - \boldsymbol{x})\big) \big] \widetilde{\mathbb{L}}_{nj}(x_j). \end{align}\] Moreover, since the partial derivatives of \(L\) are bounded by \(1\) (whenever they exist), we have \[\begin{align} \label{eq:definition-dn2-prime} |D_{n2}(\boldsymbol{x})| \le \sum_{j\in [d]} \Big|\sqrt k\{ S_{nj}(x_j) - x_j \} + \widetilde{\mathbb{L}}_{nj}(x_j) \Big| =: D_{n2}'(\boldsymbol{x}); \end{align}\tag{40}\] note that \(D_{n2}'\) is well-defined on \([0,\infty)^d\).
Regarding Theorem 2, its first result is now an immediate consequence of Lemma 2, 3 and 4. Moreover, \[\sup_{\boldsymbol{x} \in A} |B_n(S_n(\boldsymbol{x}))| \le \sup_{\boldsymbol{x} \in A^{\oplus C_{s}r}} |B_n(\boldsymbol{x})|\] is an immediate consequence of 35 .
Regarding Theorem 3, its first result is an immediate consequence of Lemma 2, 3 and 5. ◻
Lemma 2. Fix \(d \in \mathbb{N}_{\ge 2}\). There exist constants \(D_{1,1} = D_{1,1}(d) \ge 1\) and \(D_{1,2} = D_{1,2}(d) \ge 1\) only depending on \(d\) such that, for any \(n \in \mathbb{N},k \in [n], T \in \mathbb{N}\) and \(\delta \in (0,e^{-1})\) satisfying \(\log(d/ \delta) \le 2kT/7\), we have \[\begin{align} \label{eq:dn1-l} \sup_{\boldsymbol{x} \in [0,T]^d} |D_{n1}(\boldsymbol{x}) | \le D_{1,1} \sqrt{r\log\Big(\frac{T D_{1,2}}{\delta r}\Big)} =: \lambda_{n,k,d,T}^{(1)}(\delta) \end{align}\tag{41}\] with probability at least \(1-(d+2) \delta\), where \(D_{n1}(\boldsymbol{x})\) is from 37 and where \(r=r(\delta, T, k)\) is from 13 .
Lemma 3. There exist universal constants \(D_{2,1} \ge 1\) and \(D_{2,2} \ge 1\) such that, for any \(n \in \mathbb{N},k \in [n], d \in \mathbb{N}, T \in \mathbb{N}\) and \(\delta \in (0,e^{-1})\) satisfying \(\log(d/ \delta) \le 2kT/7\) and \(n/k \ge T\), \[\begin{align} \label{eq:dn2-l} \sup_{\boldsymbol{x} \in [0,T]^d} D_{n2}'(\boldsymbol{x}) \le \frac{d}{\sqrt{k}} + D_{2,1}d \sqrt{r\log\Big(\frac{T D_{2,2}}{\delta r}\Big)} =: \lambda_{n,k,d,T}^{(2)}(\delta) \end{align}\tag{42}\] with probability at least \(1-(2d + 1)\delta\), where \(D_{n2}'(\boldsymbol{x})\) is from 40 and where \(r=r(\delta, T, k)\) is from 13 .
Lemma 4. Fix \(d,T\in \mathbb{N}\) and let \((A,L)\) with \(A \subseteq[0,T]^d\) satisfy [cond:smoothness-hoelder] from Theorem 2. Then there exist some constant \(D_{3} = D_{3}(d,K_L,\alpha_L) \ge 1\) only depending on \(d,K_L\) and \(\alpha\) such that, for any \(n \in \mathbb{N},k \in [n]\) and \(\delta \in (0,e^{-1})\) satisfying \(\log(d/\delta)\le 2kT/7\) and \(r \le \kappa_L/C_{s}\) with \(C_{s}\approx 89.18\) from Lemma 11, \[\sup_{\boldsymbol{x} \in A} |D_{n3}(\boldsymbol{x})| \le D_{3} r^{\alpha_L}\sqrt{T\log(1/\delta)} =: \lambda_{n,k,d,T,K_L,\alpha_L}^{(3)}\] with probability at least \(1-(2d + 1)\delta\), where \(D_{n3}(\boldsymbol{x})\) is from 39 and where \(r=r(\delta, T, k)\) is from 13 .
Lemma 5. Fix \(d,T\in \mathbb{N}\) and assume that [cond:smoothness-good] from Theorem 3 if met. Then, there exists a constant \(D_{4} = D_{4}(d,K_L) \ge 1\) such that, for any \(n \in \mathbb{N},k \in [n]\) and \(\delta \in (0,e^{-1})\) satisfying \(\log(d/\delta)\le 2kT/7\) and \(n/k\ge 2T\), we have \[\sup_{\boldsymbol{x} \in [0,T]^d \setminus (\mathfrak B^{\oplus C_{s}r})} |D_{n3}(\boldsymbol{x})| \le D_{4} \sqrt{r \log\Big(\frac{T}{\delta r}\Big)} =: \lambda_{n,k,d,T,K_L}^{(4)}\] with probability at least \(1-(3d + 1)\delta\), where \(D_{n3}(\boldsymbol{x})\) is from 39 and where \(r=r(\delta, T, k)\) is from 13 .
Proof of Lemma 2. Subsequently, let \(\Omega_0\) denote the event of probability at least \(1-(d+1)\delta\) on which 114 and 115 are met, and let \(C_{s}\approx 89.18\) denote the universal constant in 115 .
Let \(\boldsymbol{x} \in [0,T]^d\). Then, on \(\Omega_0\), we have \[\begin{align} \sup_{\boldsymbol{x}\in [0,T]^d} | D_{n1}(\boldsymbol{x}) | = \sup_{\boldsymbol{x}\in [0,T]^d} | \widetilde{\mathbb{L}}_n(S_n(\boldsymbol{x})) - \widetilde{\mathbb{L}}_n(\boldsymbol{x}) | \le \omega_{\widetilde{\mathbb{L}}_{n}}\Big(\max_{j \in [d]} \sup_{x_j \in [0,T]} |S_{nj}(x_j) - x_j |; [0,2T]^d\Big) \end{align}\] where \(\omega_{f}(\varepsilon;B)\) denotes the modulus of continuity of \(f\) with respect to the maximum norm as defined in 1 , and where we used 114 .
We next distinguish two cases. First, suppose that \(C_{s}r \le 2T\), where \(r=\sqrt{(T/k) \log(1/\delta)}\) is from 13 . Then, on the event \(\Omega_0\), by 114 and 115 , \[\begin{align} \sup_{\boldsymbol{x}\in [0,T]^d} | D_{n1}(\boldsymbol{x}) | \le \omega_{\widetilde{\mathbb{L}}_{n}}(C_{s}r; [0,2T]^d) = \sqrt{\frac{n}{k} } \omega_{ \beta_{n}}\Big(\frac{k}{n}C_{s}r; [0,2Tk/n]^d\Big), \end{align}\] with \(\beta_n\) from 118 . Next, by 119 from Lemma 12 (which is applicable since \(C_{s}r \le 2T\)), there exists a set \(\Omega_1\) with probability at least \(1-\delta\) such that, on \(\Omega_1\), \[\begin{align} \label{eq:lemma1-bound1} \sqrt{\frac{n}{k} }\omega_{ \beta_{n}}\Big(\frac{k}{n}C_{s}r; [0,2Tk/n]^d\Big) \le \kappa \sqrt{C_{s}r \log\Big(\frac{4dT}{C_{s}r\delta}\Big)}, \end{align}\tag{43}\] where \[\kappa = 2d \Big[ \sqrt{\frac{4}{9C_{s}kr} \log\Big( \frac{4dT}{C_{s}r\delta}\Big)} + 2 +60 \sqrt{2d} \Big].\] Since \(\log(x) \le x\) and \(1 \le \log(1/\delta) \le 2kT/7\), we have \[\begin{align} \label{eq:bound-kappa} \frac{4}{9C_{s}kr} \log\Big( \frac{4dT}{C_{s}r\delta}\Big) &= \nonumber \frac{4}{9C_{s}kr} \Big\{ \log \Big(\frac{4dT}{C_{s}r}\Big) + \log(1/\delta) \Big\} \\&\le \nonumber \frac{4}{9C_{s}kr} \Big\{ \frac{4dT}{C_{s}r} + \sqrt{\log(1/\delta)\cdot2kT/7}\Big\} \\&= \frac{4}{9C_{s}kr} \Big\{ \frac{4d rk}{C_{s}\log(1/\delta)} + \sqrt{2/7}rk\Big\} \le \frac{4}{9C_{s}} \Big\{ \frac{4d}{C_{s}} + \sqrt{2/7}\Big\}; \end{align}\tag{44}\] note that the upper bound only depends on \(d\). As a consequence, by 43 , there exist constants \(D_{1,1}=D_{1,1}(d)\) and \(D_{1,2}=D_{1,2}(d)\) only depending on \(d\) such that, on \(\Omega_1\), \[\begin{align} \sqrt{\frac{n}{k} }\omega_{ \beta_{n}}\Big(\frac{k}{n}C_{s}r; [0,2Tk/n]^d\Big) \le D_{1,1} \sqrt{r \log\Big(\frac{TD_{1,2}}{r\delta}\Big)} = \lambda_{n,k,d,T}^{(1)}(\delta), \end{align}\] which in turn implies 41 on the event \(\Omega_0 \cap \Omega_1\) and in the case \(C_{s}r \le 2T\). The assertion follows from the fact that this event has probability at least \(1-(d+2)\delta\).
It remains to treat the case \(C_{s}r > 2T\). In that case, on \(\Omega_0\), by the triangle inequality, \[\begin{align} \sup_{\boldsymbol{x}\in [0,T]^d} | D_{n1}(\boldsymbol{x}) | \le 2 \sup_{\boldsymbol{x} \in [0,2T]^d} |{\widetilde{\mathbb{L}}_{n}} (\boldsymbol{x}) |. \end{align}\] By Lemma 10, there exists an event \(\Omega_1'\) that has probability at least \(1-\delta\) such that, on \(\Omega_1'\) and with \(C_{s}\) from 115 , \(\sup_{\boldsymbol{x} \in [0,2T]^d} |{\widetilde{\mathbb{L}}_{n}} (\boldsymbol{x}) | \le (188/3) \cdot \sqrt 2 \cdot d \sqrt{T \log(1/\delta)} \le C_{s}d \sqrt{T\log(1/ \delta)}\). Hence, on \(\Omega_0 \cap \Omega_1'\), we have \[\begin{align} \sup_{\boldsymbol{x}\in [0,T]^d} | D_{n1}(\boldsymbol{x}) | \le 2 C_{s}d \sqrt{T \log(1/\delta)} &\le \sqrt{2} C_{s}^{3/2} d \sqrt{r \log(1/\delta)} \\&\le \sqrt{2} C_{s}^{3/2} d \sqrt{r\log\Big(\frac{\sqrt{2/7} \cdot T}{r\delta} \Big)}, \end{align}\] where we used that \(T \le C_{s}r/2\) and \(r \le \sqrt{2/7} \cdot T\) at the last two inequalities. By possibly increasing \(D_{1,1}\) and \(D_{1,2}\), the upper bound is bounded by \(\lambda_{n,k,d,T}^{(1)}(\delta)\). Overall, we have shown that 41 holds on the event \(\Omega_0 \cap \Omega_1'\) and in the case \(C_{s}r > 2T\). The assertion follows from the fact that this event has probability at least \(1-(d+2)\delta\). ◻
Proof of Lemma 3. We start by writing \[\begin{align} \sqrt k\{ S_{nj}(x_j) - x_j \} &=- \sqrt k \{ \widetilde{L}_{nj}(S_{nj}(x_j)) - S_{nj}(x_j) \} + \sqrt k \{ \widetilde{L}_{nj}(S_{nj}(x_j)) - x_j \} \\ &= - \widetilde{\mathbb{L}}_{nj} \circ S_{nj}(x_j) + \sqrt k \{ \widetilde{L}_{nj}(\widetilde{L}_{nj}^\leftarrow(x_j)) - x_j \} \end{align}\] A picture reveals that \(|\widetilde{L}_{nj}(\widetilde{L}_{nj}^\leftarrow(x_j)) - x_j| \le k^{-1}\) for all \(x_j \le n/k\). Hence, since \(n/k\ge T\) by assumption, we obtain the bound \[D_{n2}'(\boldsymbol{x}) \le \sum_{j\in [d]} \Big|\sqrt k\{ S_{nj}(x_j) - x_j \} + \widetilde{\mathbb{L}}_{nj}(x_j) \Big| \le \frac{d}{\sqrt k} + \sum_{j\in [d]} \Big| \widetilde{\mathbb{L}}_{nj}(x_j) - \widetilde{\mathbb{L}}_{nj} \circ S_{nj}(x_j) \Big|.\]
We now argue as in the proof of Lemma 2: let \(\Omega_0\) denote the event of probability at least \(1-(d+1)\delta\) on which 114 and 115 are met, and let \(C_{s}\ge 1\) denote the universal constant in 115 . In the case where \(C_{s}r\le 2T\), we then have, on \(\Omega_0\), \[\begin{align} \sup_{\boldsymbol{x} \in [0,T]^d} D_{n2}'(\boldsymbol{x}) &\le \frac{d}{\sqrt{k}} + d \max_{j \in [d]} \omega_{\widetilde{\mathbb{L}}_{n,j}}(C_{s}r; [0,2T]) = \frac{d}{\sqrt{k}} + d \max_{j \in [d] } \sqrt{\frac{n}{k} } \omega_{ \beta_{n,j}}\Big(\frac{k}{n}C_{s}r; [0,2Tk/n]\Big), \end{align}\] where \(r=\sqrt{(T/k)\log(1/\delta)}\) is as in 13 and where \(\beta_{n,j}\) is the \(j\)th margin of \(\beta_n\) from 118 . As a consequence, by Lemma 12 and the union bound, \[\begin{align} \sup_{\boldsymbol{x} \in [0,T]^d} D_{n2}'(\boldsymbol{x}) &\le \frac{d}{\sqrt{k}} + d \kappa \sqrt{C_{s}r\log\Big(\frac{4T}{C_{s}\delta r}\Big)} \end{align}\] with probability at least \(1-(2d+1)\delta\), where \[\kappa=2\Big[ \sqrt{\frac{4}{9C_{s}kr} \log\Big( \frac{4T}{C_{s}r\delta}\Big)} + 2 +60 \sqrt{2} \Big] \le 2\Big[ \frac{2\sqrt{4+C_{s}(2/7)^{1/2}}}{3C_{s}} + 2 + 60\sqrt 2\Big]\] and where we used 44 with \(d=1\) for the last inequality. We hence find universal constants \(D_{2,1}\) and \(D_{2,2}\) such that that 42 holds with with probability at least \(1-(2d+1)\delta\), for the case \(C_{s}r\le 2T\).
For the case \(C_{s}r > 2T\), note that \(\widetilde{\mathbb{L}}_{nj}(x_j) = \widetilde{\mathbb{L}}_{nj}(0,\dots,0,x_j,0,\dots,0)\) and thus \[\max_{j \in [d]} \omega_{\widetilde{\mathbb{L}}_{n,j}}(C_{s}r; [0,2T]) \le \omega_{\widetilde{\mathbb{L}}_{n}}(C_{s}r; [0,2T]^d).\] Using the bound \[\max_{j \in [d]} \omega_{\widetilde{\mathbb{L}}_{n,j}}(C_{s}r; [0,2T]) \le 2 \max_{j \in [d]} \sup_{x_j \in [0,2T]}\big|\widetilde{\mathbb{L}}_{n,j}(x_j)\big|\] and then arguing similarly to the case \(C_{s}r > 2T\) in the proof of Lemma 2 completes the proof after possibly enlarging \(D_{2,1}\) and \(D_{2,2}\). ◻
Proof of Lemma 4. Recall that, for \(\boldsymbol{x} \in A\), \[\begin{align} D_{n3}(\boldsymbol{x}) =\sum_{j \in [d]} \big[ \partial_{j} L(\boldsymbol{x})- \partial_{j} L\big(\boldsymbol{x} + t^*(S_n(\boldsymbol{x}) - \boldsymbol{x})\big) \big] \widetilde{\mathbb{L}}_{nj}(x_j), \end{align}\] with \(t^*=t^*(n,x)\in[0,1]\). By Lemma 11, it holds that \(\max_{j\in[d]} \sup_{x_j \in [0,T]} |S_{nj}(x_j) - x_j| \le C_{s}r\) on a set \(\Omega_0\) of probability at least \(1-(d+1)\delta.\) Hence, on this set, the assumption \(C_{s}r \le \kappa_L\) and [cond:smoothness-hoelder] imply that \[\big| \partial_{j} L(\boldsymbol{x})- \partial_{j} L\big(\boldsymbol{x} + t^*(S_n(\boldsymbol{x}) - \boldsymbol{x}) \big| \le K_L\|\boldsymbol{x} - S_{n}(\boldsymbol{x})\|^{\alpha_L}_{\infty} \le K_L(C_{s}r)^{\alpha_L}\] for all \(\boldsymbol{x}\in A\). As a consequence, \[|D_{n3}(\boldsymbol{x})| \le K_L(C_{s}r)^{\alpha_L} \sum_{j \in[d]} |\widetilde{\mathbb{L}}_{nj}(x_j)| \le d K_L(C_{s}r)^{\alpha_L} \max_{j\in [d]} \sup_{x_j \in [0,T]} |\widetilde{\mathbb{L}}_{nj}(x_j) |.\] By Lemma 10, with probability at least \(1-d\delta\), \[\max_{j \in [d]} \sup_{x_j \in [0,T]} |\tilde{\mathbb{L}}_{nj}(x_j) | \le (188/3) \sqrt{T\log(1/\delta)} \le C_{s}\sqrt{T\log(1/\delta)}.\] Combining the previous displays, we find that \[\sup_{\boldsymbol{x} \in A} |D_{n3}(\boldsymbol{x})| \le C_{s}^{1+\alpha_L} K_Ld r^{\alpha_L} \sqrt{T \log(1/\delta)}\] with probability at least \(1-(2d+1)\delta\). Choosing \(D_{3} = C_{s}^{1+\alpha_L} K_Ld\) yields the desired bound. ◻
Proof of Lemma 5. Subsequently, let \(\Omega_0\) denote the event of probability at least \(1-(d+1)\delta\) on which 114 and 115 are met, and let \(C_{s}\approx 89.18\) denote the universal constant in 115 .
Recall that, for \(\boldsymbol{x} \in [0,T]^d \setminus(\mathfrak B^{\oplus C_{s}r})\), \[\begin{align} D_{n3}(\boldsymbol{x}) =\sum_{j \in [d]} \big[ \partial_{j} L(\boldsymbol{x})- \partial_{j} L\big(\boldsymbol{x} + t^*(S_n(\boldsymbol{x}) - \boldsymbol{x})\big) \big] \widetilde{\mathbb{L}}_{nj}(x_j), \end{align}\] with \(t^*=t^*(n,x)\in[0,1]\). We now distinguish two cases, according to whether \(4C_{s}r \le T\) or \(4C_{s}r>T\). In the latter case, using that \(0 \le \partial_{j} L(\cdot ) \le 1\) and Lemma 10 (which is applicable since \(\log(1/\delta) \le \log(d/\delta) \le 2kT/7 \le Tk\)), we have \[\sup_{\boldsymbol{x} \in [0,T]^d \setminus(\mathfrak B^{\oplus C_{s}r})} |D_{n3}(\boldsymbol{x})| \le d \max_{j\in [d]} \sup_{x_j < T} \big| \widetilde{\mathbb{L}}_{nj}(x_j) \big| \le d(188/3) \sqrt{T \log(1/\delta)} = d C_{s}\sqrt{T\log(1/\delta)/2}\] with probability at least \(1-d\delta\). Since \(T<4C_{s}r\) and \(r \le \sqrt{2/7} \cdot T \le T\), the upper bound satisfies \[d C_{s}\sqrt{T \log(1/\delta) /2 } \le d C_{s}^{3/2} \sqrt{2r \log(1/\delta)} \le d C_{s}^{3/2} \sqrt{2r \log\Big(\frac{T}{r\delta} \Big)} \le \lambda_{n,k,m,T,K_L}^{(4)},\] provided we choose \(D_{4}\ge d C_{s}^{3/2}\sqrt{2}\). Note that we do not need any smoothness assumptions on \(L\) here.
It remains to treat the case \(4C_{s}r \le T\). For each \(\boldsymbol{x} \in [0,T]^d \setminus(\mathfrak B^{\oplus C_{s}r})\), we may decompose \[D_{n3}(\boldsymbol{x}) = D_{n3}^0(\boldsymbol{x}) + D_{n3}^+(\boldsymbol{x}) := \sum_{j\in [d]} A_{nj}(\boldsymbol{x}) \boldsymbol{1}(x_j < 2C_{s}r) + \sum_{j\in [d]} A_{nj}(\boldsymbol{x}) \boldsymbol{1}(x_j \in [2C_{s}r,T])\] where \[A_{nj}(\boldsymbol{x}) := \big[ \partial_{j} L(\boldsymbol{x})- \partial_{j} L\big(\boldsymbol{x} + t^*(S_n(\boldsymbol{x}) - \boldsymbol{x})\big) \big] \widetilde{\mathbb{L}}_{nj}(x_j).\] We start by bounding \(D_{n3}^0(\boldsymbol{x})\). Again using that \(0 \le \partial_{j} L(\cdot) \le 1\), we have, for any \(j \in [d]\), \[|A_{nj}(\boldsymbol{x})| \boldsymbol{1} (x_j < 2C_{s}r) \le \sup_{0 < x_j < 2C_{s}r} \big|\widetilde{\mathbb{L}}_{nj}(x_j) \big|.\] As a consequence, again by Lemma 10 applied with \(T = 2C_{s}r\) and \(d=1\), the union bound and the fact that \(r \le T\), we have \[\label{eq:bound-dn3-0} \sup_{\boldsymbol{x} \in [0,T]^d \setminus(\mathfrak B^{\oplus C_{s}r})} |D_{n3}^0(\boldsymbol{x})| \le d C_{s}^{3/2} \sqrt{r \log(1/\delta)} \le d C_{s}^{3/2} \sqrt{r\log\Big(\frac{T}{r\delta} \Big)}\tag{45}\] with probability at least \(1-d\delta\); note that Lemma 10 can be applied with \(T = 2C_{s}r\) here because \(\log(1/\delta) = r\sqrt{k\log(1/\delta) /T} \le \sqrt{2/7}\cdot rk = [\sqrt{2/7} / (2C_{s})] \cdot 2C_{s}r k \le 2C_{s}r k\) by assumption.
We continue by bounding \(\sup_{\boldsymbol{x} \in [0,T]^d \setminus(\mathfrak B^{\oplus C_{s}r})} |D_{n3}^+(\boldsymbol{x})|\). Again working on the set \(\Omega_0\), note that \(\boldsymbol{x} \in [0,T]^d \setminus(\mathfrak B^{\oplus C_{s}r})\) implies that \([\boldsymbol{x}, S_n(\boldsymbol{x})] \subseteq G :=[0,T]^d \setminus \mathfrak B\). Further, the condition \(x_j \ge 2C_{s}r\) implies that \(S_{nj}(x_j) \ge C_{s}r >0\). As a consequence, we may apply Lemma 13 to obtain the bound \[\begin{align} \big| A_{nj}(\boldsymbol{x}) \big| \boldsymbol{1}(x_j \in [2C_{s}r,T]) &\le K_L\max\Big\{\frac{1}{x_j},\frac{1}{S_{nj}(x_j)}\Big\} \| S_n(\boldsymbol{x}) - \boldsymbol{x}\|_1 \big| \widetilde{\mathbb{L}}_{nj}(x_j)\big| \boldsymbol{1} (x_j \in [2C_{s}r , T ]) \\ &\le K_Ld \times C_{n1} \times C_{n2}, \end{align}\] where \[\begin{align} C_{n1} &= \max_{\ell \in [d]} \sup_{x_\ell \in [0,T]} \big| S_{n\ell}(x_\ell) - x_\ell \big|, \\ C_{n2} &= \max_{j \in [d]} \sup_{x_j \in [2C_{s}r,T]} \max \Big\{\frac{1}{x_j},\frac{1}{S_{nj}(x_j)} \Big\} \big| \widetilde{\mathbb{L}}_{nj}(x_j)\big|, \end{align}\] which in turn yields \[\sup_{\boldsymbol{x} \in [0,T]^d \setminus(\mathfrak B^{\oplus C_{s}r})} |D_{n3}^+(\boldsymbol{x})| \le K_Ld^2 \times C_{n1} \times C_{n2}.\] Since we are working on \(\Omega_0\), we have \(C_{n1} \le C_{s}r\). Concerning \(C_{n2}\), note that for \(x_j\ge 2C_{s}r\), \[\begin{align} S_{nj}(x_j) = x_j \Big(1+\frac{S_{nj}(x_j) - x_j}{x_j} \Big) &\ge x_j \Big(1- \frac{\max_{\ell \in [d]} \sup_{x_\ell \in [0,T]}|S_{n\ell}(x_j) - x_j| }{2C_{s}r} \Big) \\&= x_j \Big(1-\frac{C_{n1}}{2C_{s}r} \Big) \ge \frac{x_j}{2}, \end{align}\] where we have used that \(C_{n1} \le C_{s}r\) on the event \(\Omega_0\). As a consequence, with \(\beta_{nj}(u_j)\) the \(j\)th coordinate of \(\beta_n\) from 118 , \[\begin{align} C_{n2} \le 2\max_{j \in [d]} \sup_{x_j \in [2C_{s}r,T]} \frac{1}{x_j} \big| \widetilde{\mathbb{L}}_{nj}(x_j)\big| &\le 2 (2C_{s}r)^{-1/2} \max_{j \in [d]} \sup_{x_j \in [2C_{s}r,T]}\frac{1}{x_j^{1/2}} \big| \widetilde{\mathbb{L}}_{nj}(x_j)\big| \\&= 2^{1/2} (C_{s}r)^{-1/2} \max_{j \in [d]} \sup_{x_j \in [2C_{s}r,T]}\frac{1}{x_j^{1/2}} \Big|\sqrt{\frac{n}{k}}\beta_{nj}\Big(\frac{k}{n} x_j\Big)\Big| \\&= 2^{1/2} (C_{s}r)^{-1/2} \max_{j \in [d]} \sup_{x_j \in [ 2C_{s}r\frac{k}{n}, T\frac{k}{n}]} \frac{|\beta_{nj}(x_j)|}{{x_j}^{1/2}}. \end{align}\] Thus, on \(\Omega_0\), we obtain the upper bound \[\sup_{\boldsymbol{x} \in [0,T]^d \setminus(\mathfrak B^{\oplus C_{s}r})} |D_{n3}^+(\boldsymbol{x})| \le K_Ld^2 (2C_{s}r)^{1/2} \max_{j \in [d]} \sup_{x_j \in [ 2C_{s}r\frac{k}{n}, T\frac{k}{n}]} \frac{|\beta_{nj}(x_j)|}{{x_j}^{1/2}}.\] By Corollary 11.2.1 on page 446 in [58] (with \(\delta=1/2\) in the notation of that reference; it should also be noted that some considerations show that the result also applies with our definition of \(\beta_{nj}\) that is based on ‘\(<\)’ instead of ‘\(\le\)’ inside the indicators), which is applicable since \(n/k\ge 2T\) by assumption and since \(2C_{s}r\tfrac{k}n / (T \tfrac{k}n) =2C_{s}r/T \le 1/2\) in our current case \(4C_{s}r \le T\), we have, for any \(\varepsilon>0\), \[\label{eq:showel95alpha47x-new} \mathbb{P}\Big( \sup_{x_1 \in [ 2C_{s}r\frac{k}{n}, T\frac{k}{n}]} \frac{\beta_{n1}(x_1)^{\pm}}{{x_1}^{1/2}} \ge \varepsilon\Big) \le 6 \log\Big( \frac{T}{2C_{s}r}\Big)\exp\Big(- \gamma_{\pm} \frac{\varepsilon^2}{8} \Big),\tag{46}\] where \(a^+ = \max(a,0)\) and \(a^- = \max(-a,0)\) for \(a \in \mathbb{R}\) and \(\gamma_- =1\) and \[\gamma_+ = \begin{cases} \frac{1}{2} & \text{if } \varepsilon\le \frac{3}{2}(2C_{s}kr)^{1/2},\\ \frac{3}{4} \frac{(2C_{s}kr)^{1/2}}{\varepsilon} & \text{if } \varepsilon> \frac{3}{2}(2C_{s}kr)^{1/2}. \end{cases}\]
We will later show that for \(\varepsilon=\lambda/(K_Ld^2(2C_{s}r)^{1/2})\) and our choice of \(\lambda\) below it holds that \(\varepsilon\le \frac{3}{2} \sqrt{2C_{s}rk}\). Then, since \(\gamma_- = 1 \ge 1/2 = \gamma_+\) and \(|a| = a^+ \vee a^-\) for any \(a\in \mathbb{R}\), Equation 46 implies that \[\mathbb{P}\Big( \sup_{x_1 \in [2r\frac{k}{n} ,T \frac{k}{n}]} \frac{|\beta_{n1}(x_1)|}{{x_1}^{1/2}} > \varepsilon\Big) \le 12 \log\Big( \frac{T}{2C_{s}r}\Big) \exp\Big(- \frac{\varepsilon^2}{16} \Big) .\] As a result, \[\begin{align} \mathbb{P}\Big( \Big\{\sup_{\boldsymbol{x} \in [0,T]^d \setminus(\mathfrak B^{\oplus C_{s}r})} |D_{n3}^+(\boldsymbol{x})| > \lambda \Big\} \cap \Omega_0\Big) \le 12 d \log\Big( \frac{T}{2C_{s}r}\Big) \exp\Big(-\frac{\lambda^2}{32C_{s}K_L^2d^4 r} \Big) \end{align}\] which is equal to \(d\delta\) if we set \[\begin{align} \label{eq:lambda-dn543} \lambda = 4\sqrt{2C_{s}}K_Ld^2 \sqrt{r \log\Big(\frac{12 \log(T/(2C_{s}r))}{\delta}\Big)}. \end{align}\tag{47}\] Overall, \[\begin{align} \mathbb{P}\Big( \sup_{\boldsymbol{x} \in [0,T]^d \setminus(\mathfrak B^{\oplus C_{s}r})} |D_{n3}^+(\boldsymbol{x})| > \lambda \Big) &\le \mathbb{P}\Big( \Big\{\sup_{\boldsymbol{x} \in [0,T]^d \setminus(\mathfrak B^{\oplus C_{s}r})} |D_{n3}^+(\boldsymbol{x})| > \lambda \Big\} \cap \Omega_0 \Big) + \mathbb{P}(\Omega_0^c) \\&\le (2d+1)\delta, \end{align}\] and together with 45 , we get \[\sup_{\boldsymbol{x} \in [0,T]^d \setminus(\mathfrak B^{\oplus C_{s}r})}|D_{n3}(\boldsymbol{x})| \le d C_{s}^{3/2} \sqrt{r\log\Big(\frac{T}{\delta r}\Big)} + 4K_Ld^2 \sqrt{2C_{s}r \log\Big(\frac{12 \log(T/(2C_{s}r))}{\delta}\Big)}\] with probability at least \(1-(3d+1)\delta\). Since \(\log(x) \le x/e\) for \(x\ge 1\) \[\label{eq:lambda5bnd-new} \frac{12\log(T/(2C_{s}r))}{\delta} \le \frac{6e^{-1} T}{ \delta C_{s}r} \le \frac{T}{ \delta r},\tag{48}\] we obtain that, with probability at least \(1-(3d+1)\delta\), \[\sup_{\boldsymbol{x} \in W_m^0(T)} |D_{n3}(\boldsymbol{x})| \le \Big(d C_{s}^{3/2} + 4 K_Ld^2 \sqrt{2C_{s}} \Big) \sqrt{r \log\Big(\frac{T}{\delta r} \Big)},\] which is bounded by \(\lambda_{n,k,m,T,K_L}^{(4)}\) if we choose \(D_{4}\) at least as large as the term in round brackets. This yields the claim for the case \(4C_{s}r \le T\). The two cases \(4C_{s}r \le T\) and \(4C_{s}r> T\) can then easily be merged by choosing \(D_4\) appropriately.
Finally, we need to show that \(\varepsilon= \lambda/(K_Ld^2(2C_{s}r)^{1/2}) \le \frac{3}{2} \sqrt{2C_{s}kr}\) holds for \(\lambda\) in 47 , provided that \(4C_{s}r \le T\). Using 48 , we have \[\varepsilon= \frac{\lambda}{K_Ld^2\sqrt{2C_{s}r}} = 4\sqrt{ \log\Big(\frac{12 \log(T/(2C_{s}r))}{\delta}\Big)} \le 4\sqrt{\log\Big(\frac{6e^{-1}T}{C_{s}\delta r}\Big)}.\] Next, using \(r=\sqrt{T\log(1/\delta)/k} \ge\sqrt{T/k}\), \(C_{s}\ge 1\) and \(6 /e \ge 1\), and again using that \(\log(x)\le x/e\) for \(x \ge 1\), it follows that \[\log\Big(\frac{6e^{-1}T}{\delta C_{s}r}\Big)\le \log\Big(\frac{6e^{-1}\sqrt{Tk}}{\delta}\Big)\le 6e^{-2}\sqrt{Tk}+ \log(1/\delta).\] By assumption, we also have \(1\le \log(1/\delta) \le \sqrt{Tk\log(1/\delta)}\sqrt{2/7}\), which yields the upper bound \[6e^{-2}\sqrt{Tk}+ \log(1/\delta) \le \sqrt{Tk \log(1/\delta)}\big(6e^{-2} + \sqrt{2/7}\big).\] With \(16(6e^{-2} + \sqrt{2/7}) = 21.54\ldots < 22\) we obtain that \(\varepsilon^2 \le 22\sqrt{Tk\log(1/\delta)} = 22 rk\), which is bounded by \((9/2) C_{s}kr\) by definition of \(C_{s}\approx 89.18\) in Lemma 11. ◻
Proof of Theorem 7. Without loss of generality, we can assume that \(\log^5(pn) / k \le 1\); otherwise, the result is trivial.
The triangle inequality yields \[d_K(\boldsymbol{S}_n, \boldsymbol{G}_n) \le d_K(\boldsymbol{S}_n, \boldsymbol{T}_n) + d_K(\boldsymbol{T}_n, \boldsymbol{G}_n).\] We start by bounding \(d_K(\boldsymbol{S}_n , \boldsymbol{T}_n)\). An application of Lemma 14 yields, for any \(\lambda>0\), \[\begin{align} \label{eq:d-s-t} d_K(\boldsymbol{S}_n , \boldsymbol{T}_n) \le \mathbb{P}\big(\|\boldsymbol{S}_n-\boldsymbol{T}_n\|_\infty \ge \lambda\big) + \sup_{\boldsymbol{x} \in \mathbb{R}^p} \mathbb{P}( \boldsymbol{T}_n \le \boldsymbol{x}+\lambda \boldsymbol{1} ) - \mathbb{P}( \boldsymbol{T}_n \le \boldsymbol{x}-\lambda \boldsymbol{1}). \end{align}\tag{49}\] The first term can be dealt with using Corollary 1. Denote by \(\lambda=\lambda_{n,k}(\delta)\) the upper bound in Corollary 1 for suitable \(\delta\) chosen below and for \(T=1\); we justify below that the corollary can be applied. With this, we obtain that \[\begin{align} \mathbb{P}\Big(\| \boldsymbol{S}_n - \boldsymbol{T}_n \|_\infty > \lambda \Big) = \mathbb{P}\Big( \max_{\boldsymbol{y} \in A} \big|{\mathbb{L}}_{n}(\boldsymbol{y}) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n}(\boldsymbol{y}) \big| > \lambda \Big) \le |\mathcal{I}|(6m+5) \delta \le 11|\mathcal{I}|m\delta. \end{align}\] Regarding the supremum on the right of 49 , we have, by Theorem 18, \[\begin{align} &\phantom{{}={}} \nonumber \mathbb{P}( \boldsymbol{T}_n \le \boldsymbol{x}+\lambda\boldsymbol{1} ) - \mathbb{P}( \boldsymbol{T}_n \le \boldsymbol{x}-\lambda\boldsymbol{1} ) \\&= \nonumber \mathbb{P}( \boldsymbol{G}_n \le \boldsymbol{x}+\lambda\boldsymbol{1} ) - \mathbb{P}( \boldsymbol{G}_n \le \boldsymbol{x}-\lambda\boldsymbol{1} ) + \big\{ \mathbb{P}( \boldsymbol{T}_n \le \boldsymbol{x}+\lambda\boldsymbol{1} ) - \mathbb{P}( \boldsymbol{G}_n \le \boldsymbol{x}+\lambda\boldsymbol{1} ) \big\} \\&\nonumber + \big\{ \mathbb{P}( \boldsymbol{G}_n \le \boldsymbol{x}-\lambda\boldsymbol{1} ) - \mathbb{P}( \boldsymbol{T}_n \le \boldsymbol{x}-\lambda\boldsymbol{1} ) \big\} \\&\le\nonumber \frac{2\lambda}{\sigma_{\min}^2} \big\{ 2+\sqrt{2\log p} \big\} + 2 d_K(\boldsymbol{T}_n, \boldsymbol{G}_n) \\&\le \label{eq:bound-dK} \frac{8\lambda}{\sigma_{\min}^2} \sqrt{\log p} + 2 d_K(\boldsymbol{T}_n, \boldsymbol{G}_n) \end{align}\tag{50}\] where we have used that \(p \ge 2\) and that \(2/\sqrt{\log(2)}+\sqrt 2 \approx 3.81 \le 4\) at the last inequality. Overall, \[\label{eq:cltXZ} d_K(\boldsymbol{S}_n, \boldsymbol{G}_n) \le 11|\mathcal{I}|m\delta + \frac{8\lambda_{n,k}(\delta)}{\sigma_{\min}^2} \sqrt{\log(p)} + 3 d_K(\boldsymbol{T}_n, \boldsymbol{G}_n).\tag{51}\] We proceed by bounding \(d_K(\boldsymbol{T}_n, \boldsymbol{G}_n)\). Note that the coordinates of \(\boldsymbol{T}_n\) are of the form \(\sum_{i=1}^n Y_{i,n,I}(\boldsymbol{x}_I)\), where \[\begin{align} Y_{i,n,I}(\boldsymbol{x}_I) &= \frac{1}{\sqrt k} \Big[ \boldsymbol{1}(\exists j \in I: V_{ij} < kx_j/n) - \mathbb{P}(\exists j \in I: V_{ij} < kx_j/n) \\ & - \sum_{j\in I} \partial_{j} L_I(\boldsymbol{x}_I) \big\{\boldsymbol{1}(V_{ij} < kx_j/n) - kx_j/n \big\} \Big], \end{align}\] with \(\mathbb{E}[Y_{i,n,I}(\boldsymbol{x}_I) ]=0\) and \(\sum_{i=1}^n \mathbb{E}[|Y_{i,n,I}(\boldsymbol{x}_I)|^2 ]\) equal to one of the diagonal entries of \(\Sigma_n\). We are going to apply the CCK-result from Theorem 19, and need to check its conditions. The first conditions holds with \(b_1=\sigma_{\min}^2\). The second and third condition hold with \(B_n = (m+1) (\log 2)^{-1} \sqrt{n/k}\) and \(b_2 = 4(1+m)m(\log2)^2\); indeed, \[\begin{align} \sum_{i=1}^n \mathbb{E}[|Y_{i,n,I}(\boldsymbol{x}_I) |^{4}] &\le \nonumber (1+m)^3\frac{n}{k^{3/2}} \mathbb{E}[|Y_{i,n,I}(\boldsymbol{x}_I)|] \\&\le \nonumber 2(1+m)^3\frac{1}{k} \Big[ \widetilde{\mu}_{n,I}(\boldsymbol{x}_I) + \sum_{j \in I} x_j \Big] \\&\le \label{eq:bound-yni4} 4(1+m)^3m \frac{1}{k}= b_2 B_n^2 \frac{1}{n}, \end{align}\tag{52}\] where we used the triangle inequality, the fact that that for a Bernoulli\((p)\) random variable \(X\) we have \(\mathbb{E}|X-p| = 2p(1-p) \le 2p\), and \(|\widetilde{\mu}_{n,I}(\boldsymbol{x}_I)| \le \sum_{j \in I} x_j \le m\) by the union bound. Moreover \[\sqrt n|Y_{i,n,I}(\boldsymbol{x}_I) | / B_n \le \sqrt{n/k} (m+1)/B_n = \log(2).\] An application of Theorem 19 then yields \[3d_K(\boldsymbol{T}_n , \boldsymbol{G}_n) \le c_1 \Big( \frac{\log^5(pn)}{k} \Big)^{1/4},\] for some constant \(c_1\) depending on \(\sigma_{\min}^2\) and \(m\) only.
It remains to bound the first and second term in 51 , for which we use \[\delta=\frac{1}{m|\mathcal{I}|}\Big( \frac{\log^5(pn)}{k} \Big)^{1/4}\] to balance the first and the last term. Indeed, the first term in 51 then satisfies \[11|\mathcal{I}|m\delta \le 11 \Big( \frac{\log^5(pn)}{k} \Big)^{1/4}.\]
Finally, regarding the second summand in 51 , we start by justifying the application of Corollary 1 with the above choice of \(\delta\) and with \(T=1\). First, our assumption \(\log^5(pn) / k \le 1\) from the beginning of the proof implies that \(\delta \le 1/(m|\mathcal{I}|) < 1/e\), while the assumption \(\log(m^2 |\mathcal{I}|k^{1/4} ) \le 2k/7\) yields, \[\begin{align} \log(m/\delta) &= \log\Big( \frac{m^2|\mathcal{I}|k^{1/4}}{\log^{5/4}(pn)}\Big) \le \log(m^2 |\mathcal{I}| k^{1/4}) \le 2k/7. \end{align}\] Finally, the assumption \(\log(m |\mathcal{I}|k^{1/4} ) \le \kappa_L^2 k/ C_{s}^2\) yields \[r = \sqrt{\frac{1}{k}\log\Big( \frac{1}{\delta}\Big)} = \sqrt{\frac{1}{k}\log\Big( \frac{m|\mathcal{I}|k^{1/4}}{\log^{5/4}(pn)} \Big)} \le \sqrt{\frac{1}{k}\log( m|\mathcal{I}|k^{1/4})} \le \kappa_L/C_{s}.\] Overall, all Conditions of Corollary 1 are met.
It remains to bound the second summand in 51 , which is \[\begin{align} \label{eq:bound-on-logp-term} \frac{8\lambda_{n,k}(\delta)}{\sigma_{\min}^2} \sqrt{\log(p)} &= \nonumber \frac{8}{\sigma_{\min}^2} \sqrt{\log(p)} \Big\{ \max_{I \in \mathcal{I}} B_{n,k,T}(L_I; A_I^{\oplus \kappa_L})+ \frac{m}{\sqrt{k}} \\&+ D_{1} \sqrt{r\log\Big(\frac{D_{2}}{\delta r}\Big)} + D_3 r^{\alpha_L}\sqrt{ \log\Big(\frac{1}{\delta}\Big)} \Big\}. \end{align}\tag{53}\] First, since \(\log(p)/k \le 1\) by our assumption at the beginning of the proof, we have \[\sqrt{\frac{\log p}{k}} \le \Big( \frac{\log p}{k} \Big)^{1/4} \le \Big( \frac{\log^5(pn)}{k} \Big)^{1/4}.\] Next, with our above choice of \(\delta\), we have, using \(|\mathcal{I}| \le p\) and the fact that \(pk \ge 2\) implies \(\log(mpk) \le C_{1,m}^2 \log(pk)\) with \(C_{1,m} = \{1 + \log(m)/\log(2)\}^{1/2}\), \[r = \sqrt{\frac{1}{k} \log\Big( \frac{m|\mathcal{I}|k^{1/4}}{\log^{5/4}(pn)}\Big)} \le \sqrt{\frac{1}{k} \log\big(m p k^{1/4}\big)} \le C_{1,m} \sqrt{\frac{\log( p k)}{k}}\] Also, \[\delta = \frac{1}{m|\mathcal{I}|} \Big( \frac{\log^5(pn)}{k} \Big)^{1/4} \ge \frac{1}{m|\mathcal{I}|k^{1/4}} \ge \frac{1}{mpk^{1/4}}\] and \(r \ge k^{-1/2}\) (since \(\delta < 1/e\)). Hence, the last two terms in 53 can be bounded as follows: first, \[\begin{align} \sqrt{r\log\Big(\frac{D_{2}}{\delta r}\Big)} \sqrt{\log p} &\le \Big(\frac{C_{1,m}^2\log(pk)}{k} \Big)^{1/4} \sqrt{\log(D_2mpk^{3/4}) \log p} \\&\le \Big(\frac{C_{1,m}^2\log(pk)}{k} \Big)^{1/4} \sqrt{D_2'\log(pk) \log p} \\&\le (C_{1,m}D_2')^{1/2} \Big(\frac{\log^5(pk)}{k} \Big)^{1/4}\le (C_{1,m}D_2')^{1/2} \Big(\frac{\log^5(pn)}{k} \Big)^{1/4}, \end{align}\] where \(D_2'=1+ \log(D_2m)/\log (2)\) only depends on \(m\). Second, \[\begin{align} r^{\alpha_L}\sqrt{\log\Big(\frac{1}{\delta }\Big)} \sqrt{\log p} &\le C_{1,m}^{\alpha_L}\Big(\frac{\log(pk)}{k} \Big)^{\alpha_L/2} \sqrt{\log(mpk^{1/4}) \log p} \\&\le C_{1,m}^{\alpha_L}\Big(\frac{\log(pk)}{k} \Big)^{\alpha_L/2} \sqrt{C_{1,m}^2\log(pk) \log p} \\&\le C_{1,m}^{1+\alpha_L}\Big(\frac{\log(pk)}{k} \Big)^{1/4} \sqrt{\log(pk) \log p} \\&\le C_{1,m}^2 \Big(\frac{\log^5(pn)}{k} \Big)^{1/4}, \end{align}\] where we used that \(\alpha_L \in [1/2,1]\) and that \(\log(pk)/k \le 1\) (which is a consequence of our assumption at the beginning of the proof). Assembling terms starting from 51 , we have shown that \[\begin{align} d_K(\boldsymbol{S}_n, \boldsymbol{G}_n) &\le \frac{8}{\sigma_{\min}^2} \sqrt{\log p} \Big(\max_{I \in \mathcal{I}} B_{n,k,T}(L_I; A_I^{\oplus \kappa_L}) \Big) \\ &+ \Big(c_1 + 11 + 8\frac{m + D_1(C_{1,m}D_2')^{1/2} + D_3 C_{1,m}^2}{\sigma_{\min}^2} \Big) \Big(\frac{\log^5(pn)}{k} \Big)^{1/4}, \end{align}\] which implies the assertion. ◻
Proof of Remark 10. A generic element of \(\Sigma_n\), say the entry at position \((q,q') \in [p]^2\), can be written as \[\begin{align} \sigma_{n,I,J}(\boldsymbol{x}_I, \boldsymbol{x}_J) = \mathbb{E}[ \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}(\boldsymbol{x}_I) \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,J}(\boldsymbol{x}_J)] \end{align}\] for certain \(I, J \in \mathcal{I}\) and \(\boldsymbol{x}_I \in A_I, \boldsymbol{x}_J \in A_J\). Write \[\begin{align} Y_{I}(\boldsymbol{x}_I) = \frac{1}{\sqrt k}\Big[ \boldsymbol{1}(J_{I}(\boldsymbol{x}_I)) - \mathbb{P}(J_{I}(\boldsymbol{x}_I)) - \sum_{j \in I} \partial_{j} L_I(\boldsymbol{x}_I) \big\{\boldsymbol{1}(J_{j}(x_{I,j})) - kx_{I,j}/n \big\} \Big], \end{align}\] where \(\boldsymbol{x}_I=(x_{I,j})_{j \in I} \in (0,1]^I\), \(J_{I}(\boldsymbol{x}_I) = \{ \exists j \in I: V_{j} < kx_{I,j}/n\}\) and \(J_{j}(x_{I,j})=J_{\{j\}}(x_{I,j})=\{V_{j} < kx_{I,j}/n\}\). We then have \[\begin{align} \sigma_{n,I,J}(\boldsymbol{x}_I, \boldsymbol{x}_J) &=\nonumber n \mathbb{E}[Y_{I}(\boldsymbol{x}_I) Y_{J}(\boldsymbol{x}_J)] \\&= \nonumber \frac{n}{k} \bigg[ \mathbb{P}[J_{I}(\boldsymbol{x}_I) \cap J_{J}(\boldsymbol{x}_J)] - \mathbb{P}[J_{I}(\boldsymbol{x}_I)] \mathbb{P}[J_{J}(\boldsymbol{x}_J)] \\ &- \nonumber \sum_{\ell \in I}\partial_{\ell} L_I(\boldsymbol{x}_I) \Big\{ \mathbb{P}[J_{\ell}( x_{I,\ell}) \cap J_{J}(\boldsymbol{x}_J)] - \frac{kx_{I,\ell}}{n} \mathbb{P}[J_{J}(\boldsymbol{x}_J)] \Big\} \\ &- \nonumber \sum_{j \in J}\partial_{j} L_J(\boldsymbol{x}_J) \Big\{ \mathbb{P}[J_{j}(x_{J,j}) \cap J_{I}(\boldsymbol{x}_I)] - \frac{kx_{J,j}}{n} \mathbb{P}[J_{I}(\boldsymbol{x}_I)] \Big\} \\ &+ \label{eq:covariance-formula} \sum_{\ell \in I, j \in J}\partial_{\ell} L_I(\boldsymbol{x}_I)\partial_{j} L_J(\boldsymbol{x}_J) \Big\{ \mathbb{P}[J_{\ell}(x_{I,\ell}) \cap J_{j}(x_{J,j})] - \frac{k^2x_{I,\ell}x_{J,j}}{n^2} \Big\} \bigg]. \end{align}\tag{54}\] The variance is obtained for \(I=J\) and \(\boldsymbol{x}_I = \boldsymbol{x}_J\), which yields \[\begin{align} \sigma_{n,I}^2(\boldsymbol{x}_I) &= \frac{n}{k} \bigg[ \mathbb{P}[J_{I}(\boldsymbol{x}_I)] - \mathbb{P}[J_{I}(\boldsymbol{x}_I)]^2 \\ &- 2 \sum_{\ell \in I}\partial_{\ell} L_I(\boldsymbol{x}_I) \Big\{ \frac{kx_{I,\ell}}{n} - \frac{kx_{I,\ell}}{n} \mathbb{P}[J_{I}(\boldsymbol{x}_I)] \Big\} \\ &+ \sum_{\ell \in I}\{ \partial_{\ell} L_I(\boldsymbol{x}_I) \}^2 \Big\{ \frac{kx_{I,\ell}}{n} - \frac{k^2x_{I,\ell}^2}{n^2} \Big\} \\ &+ \sum_{j,\ell \in I, j \ne \ell}\partial_{\ell} L_I(\boldsymbol{x}_I)\partial_{j} L_I(\boldsymbol{x}_I) \Big\{ \mathbb{P}[J_{\ell}(x_{I,\ell}) \cap J_{j}(x_{I,j})] - \frac{k^2x_{I,\ell}x_{I,j}}{n^2} \Big\} \bigg], \end{align}\] where we have used that \(\mathbb{P}[J_{\ell}(x_{I,\ell}) \cap J_{I}(\boldsymbol{x}_I)] =\mathbb{P}[J_{\ell}(x_{I,\ell})] = kx_{I,\ell}/n\). As a consequence, \[\begin{align} \sigma_I^2 (\boldsymbol{x}_I) = \lim_{n \to \infty}\sigma_{n,I}^2(\boldsymbol{x}_I) &= L_I(\boldsymbol{x}_I) - \sum_{\ell \in I}x_{I,\ell}\partial_{\ell} L_I(\boldsymbol{x}_I) \{ 2-\partial_{\ell} L_I(\boldsymbol{x}_I)\} \\ &+ 2\sum_{j,\ell \in I, j < \ell}\partial_{\ell} L_I(\boldsymbol{x}_I)\partial_{j} L_I(\boldsymbol{x}_I) R_{\{j, \ell\}}(x_{I,j}, x_{I,\ell}). \end{align}\] Homogeneity of \(L_I\) implies that the directional derivative of \(L_I\) in \(\boldsymbol{x}_I\) in direction \(\boldsymbol{v}=\boldsymbol{x}_I/\| \boldsymbol{x}_I\|_2\) is given by \[\partial_{\boldsymbol{v}} L_I(\boldsymbol{x}_I) = \lim_{h \to 0}h^{-1}\{ L_I(\boldsymbol{x}_I+h\boldsymbol{x}_I/\| \boldsymbol{x}_I\|_2) - L_I(\boldsymbol{x}_I)\} = L_I(\boldsymbol{x}_I) / \| \boldsymbol{x}_I\|_2.\] If \(L_I\) is differentiable at \(\boldsymbol{x}_I\) (a consequence of convexity and existing continuous partial derivatives in neighbourhood of \(\boldsymbol{x}_I\); see Lemma 15), we obtain that \[L_I(\boldsymbol{x}_I) = \| \boldsymbol{x}_I\|_2 \cdot \partial_{\boldsymbol{v}} L_I(\boldsymbol{x}_I) = \| \boldsymbol{x}_I\|_2 \cdot \langle \boldsymbol{v}, \nabla L_I(\boldsymbol{x}_I)) = \sum_{\ell \in I} x_{I,\ell} \partial_{\ell} L_I(\boldsymbol{x}_I).\] As a consequence, we may write \[\begin{align} \sigma_I^2 (\boldsymbol{x}_I) &= - \boldsymbol{x}_I^\top \nabla L_I(\boldsymbol{x}_I) + (\nabla L_I(\boldsymbol{x}_I))^\top \mathcal{R}_I (\nabla L_I(\boldsymbol{x}_I)) \\&= - L_I(\boldsymbol{x}_I) + (\nabla L_I(\boldsymbol{x}_I))^\top \mathcal{R}_I (\nabla L_I(\boldsymbol{x}_I)) , \end{align}\] where \(\mathcal{R}_I = (R_{j,\ell}(x_{I,j}, x_{I,\ell}))_{j,\ell \in I}\) is a \(|I| \times |I|\) matrix, with diagonal entries \(R_{j,j}(x_{I,j}, x_{I,j}) = x_{I,j}\). Suppose that \(\mathcal{R}_I\) is positive definite. Then, by the Cauchy-Schwarz-inequality, \[(\nabla L_I(\boldsymbol{x}_I))^\top \mathcal{R}_I (\nabla L_I(\boldsymbol{x}_I)) \ge \frac{(\boldsymbol{x}_I^T \nabla L_I(\boldsymbol{x}_I))^2}{\boldsymbol{x}_I^\top \mathcal{R}_I^{-1} \boldsymbol{x}_I} = \frac{L_I^2(\boldsymbol{x}_I)}{\boldsymbol{x}_I^\top \mathcal{R}_I^{-1} \boldsymbol{x}_I},\] which yields \[\sigma_I^2 (\boldsymbol{x}_I) \ge - L_I(\boldsymbol{x}_I) + \frac{L_I^2(\boldsymbol{x}_I)}{\boldsymbol{x}_I^\top \mathcal{R}_I^{-1} \boldsymbol{x}_I}.\] In the bivariate case \(I=\{j,\ell\}\) and \(\boldsymbol{x}_I=(x_j,x_\ell)\), some tedious but straightforward calculation shows that the right-hand side is equal to \[\frac{r(x_j+x_\ell - r)(x_j-r)(x_\ell-r)}{(x_j+x_\ell-2r)x_j x_\ell}\] where \(r=R_I(x_j,x_\ell)\) denotes the off-diagonal element of \(\mathcal{R}_I\). Since \(0 \le r \le x_j \wedge x_\ell\), the expression is strictly positive if an only if \(R_I \notin\{ R_{{\text{ind}}}, R_{\text{pd}}\}\), where \(R_{{\text{ind}}} \equiv 0\) and \(R_{\text{pd}}(x,y) = x \wedge y\) correspond to tail independence and perfect tail dependence, respectively. ◻
The bootstrap consistency result in Theorem 11 will be an immediate consequence of the following proposition, which in turn will follow from a couple of intermediate results stated below.
Proposition 15. Let \(L\) be a \(d\)-variate stable tail dependence function and let \(\mathcal{I}\) and \((A_I)_{I \in \mathcal{I}}\) be as described in the beginning of Section 4. Assume that there exist \(\kappa_L, K_L\in(0,\infty)\) such that \[\begin{align} \forall I \in \mathcal{I}, \forall j \in I,& \forall \boldsymbol{x}_I \in A_I^{\oplus \min(1,\kappa_L/2)}, \forall \boldsymbol{y}_I \in [0,\infty)^I \text{ with } \|\boldsymbol{x}_I - \boldsymbol{y}_I\|_\infty \le \kappa_L: \\ &\partial_{j} L_I(\boldsymbol{x}_I), \partial_{j} L_I(\boldsymbol{y}_I) \text{ exist and satisfy } |\partial_{j} L_I(\boldsymbol{x}_I)-\partial_{j} L_I(\boldsymbol{y}_I)| \le K_L\|\boldsymbol{x}_I - \boldsymbol{y}_I\|_\infty. \end{align}\] Assume the conditions (i)–(iii) of Theorem 7 are met with the condition \(\log(m|\mathcal{I}| k^{1/4}) \le \kappa_L^2k/ C_{s}^2\) replaced by \(\log(m|\mathcal{I}| k^{1/4}) \le \kappa_L^2k/ (8C_{s}^2)\), and with \(n/k\ge 2\). Let \[h<(\min_{I \in \mathcal{I}} \min_{\boldsymbol{x}_I \in A_I} \min_{j \in I}\boldsymbol{x}_{I,j}) \wedge (\kappa_L/2).\] Then, there exist constants \(c_i = c_i(m,K_L,\sigma_{\mathrm{min}}) \ge 1, i = 1,2\), such that, with probability at least \(1-c_1\delta_n\) \[\begin{gather} \label{eq:prop-boot} d_K( \mathcal{L}(\boldsymbol{S}_n^* \mid \mathrm{data}), \boldsymbol{G}_n) \leq c_2\delta_n + c_2\log(p+k) \\ \times \Big( h + \sqrt{r_{2,n}} + \frac{r_{2,n}}{\sqrt{h}} + \frac{r_{2,n}^2}{h} + \frac{1}{h\sqrt{k}} \Big\{ B_{n,k}(L_I ; A_I^{\oplus\kappa_L}) + \Big[\frac{\log^3(pk)}{k} \Big]^{1/4} \Big\}\Big) \end{gather}\qquad{(1)}\] where \(\delta_n := [k^{-1}\log^5(pn)]^{1/4}\) and \(r_{2,n} := \sqrt{k^{-1}\log(pk)}\).
Proof of Theorem 11. The conditions of Proposition 15 are a subset of the conditions of Theorem 11, whence it suffices to show that the upper bound in ?? can be bounded as claimed in the theorem. Since \(n,p \ge 2\) we may assume without loss of generality that \(k \ge 2\), which yields \(\log(p+k) \le \log(pk) \le \log(pn)\). Hence, \[\begin{align} h\log(p+k) & \le c_h' [k^{-1}\log (p+k)]^{1/4} \log(p+k) \le c_h' \delta_n, \\ \frac{r_{n,2}\log(p+k)}{\sqrt{h}} &\le c_h^{-1/2}\frac{k^{-1/2}\log^{1/2}(pk) \log(p+k)}{k^{-1/4}\log^{1/4}(p+k)} \le c_h^{-1/2}\frac{\log^{5/4}(pk)}{k^{1/4}} \le c_h^{-1/2}\delta_n, \\ \frac{r_{n,2}^2\log(p+k)}{h} &\le c_h^{-1}\frac{k^{-1}\log(pk)\log(p+k)}{k^{-1/2}\log^{1/2}(p+k)} \le c_h^{-1}\frac{\log^{3/2}(pk)}{k^{1/2}} \le c_h^{-1} \delta_n^2, \\ \frac{\log(p+k)}{h\sqrt{k}} \Big[\frac{\log^3(pk)}{k} \Big]^{1/4} &\le c_h^{-1}\frac{\log(p+k)\log^{3/4}(pk)}{k^{3/4}k^{-1/2} \log^{1/2}(p+k)} \le c_h^{-1} \frac{\log^{5/4}(pk)}{k^{1/4}} \le c_h^{-1} \delta_n. \end{align}\] Finally \[\frac{\log(p+k)}{h\sqrt{k}} \le c_h^{-1}\frac{\log(p+k)}{k^{1/2}k^{-1/2} \log^{1/2}(p+k)} = c_h^{-1} \sqrt{\log(p+k)},\] so \[\frac{1}{h\sqrt{k}} B_{n,k}(L_I ; A_I^{\oplus\kappa_L})\log(p+k) \le c_h^{-1} \sqrt{\log(p+k)} B_{n,k}(L_I ; A_I^{\oplus\kappa_L}).\] Combining the above and noting that we can assume \(\delta_n \le 1\) since otherwise the bound is trivial by setting \(c_2=1\) completes the proof. ◻
The proof of Proposition 15 and the subsequent lemmas require additional notation. Recall \(\boldsymbol{S}_n\) and \(\boldsymbol{S}_n^*\) from 18 and 21 , respectively, and let \[\begin{align} \label{eq:Scirc-boot} \boldsymbol{S}_n^\circ = ( \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup ^\circ_{n,I}(\boldsymbol{x}_{I,\ell}))_{I \in \mathcal{I}, \ell \in [p_I]}, \qquad \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup ^\circ_{n,I}(\boldsymbol{x}_I) =\sum_{i=1}^n e_i \Big\{Y_{i,I}(\boldsymbol{x}_I) - \frac{1}{n} \sum_{i'=1}^n Y_{i',I}(\boldsymbol{x}_I) \Big\} \end{align}\tag{55}\] which is unobservable.
Proof of Proposition 15. Throughout the proof we assume \(k^{-1}\log(pk) \le 1\) as the statement is trivial otherwise. By Lemma 6 we have with probability one \[\label{eq:boot-bound-KS-first} d_K( \mathcal{L}(\boldsymbol{S}_n^* \mid \mathrm{data}), \boldsymbol{G}_n) \lesssim \frac{1}{k} + \frac{\Delta \cdot \log(p+k)}{\sigma_{\mathrm{min}}^2} + d_K( \mathcal{L}(\boldsymbol{S}_n^\circ \mid \mathrm{data}), \boldsymbol{G}_n).\tag{56}\] Set \[\delta := \frac{1}{m|\mathcal{I}|} \Big( \frac{\log^5(pn)}{k} \Big)^{1/4}.\] In the proof of Theorem 7 we verify that the conditions of Corollary 1 hold with this choice of \(\delta\). Moreover, \(n/k\ge 2\) by assumption, and using that \(|\mathcal{I}| \le p\) and \(\log(pn) \ge 1\), the assumption \(\log(mpk^{1/4}) \le \kappa_L^2/(8 C_{s}^2)\) implies \(r=\sqrt{k^{-1}\log(1/\delta)} \le \kappa_L/(2^{3/2} C_{s})\). Hence all conditions of Lemma 8 hold with this choice of \(\delta\). The latter lemma shows that, with probability at least \(1 - |\mathcal{I}|(6m+7)\delta\) \[\Delta \lesssim h + \sqrt{r} + \frac{r^2}{h} + \frac{r}{\sqrt{h}} + \frac{1}{h\sqrt{k}} \Big\{ B_{n,k}(L_I ; A_I^{\oplus\kappa_L}) + \sqrt{r \log\Big(\frac{1}{\delta r}\Big)} \Big\}\] where the implicit constant depends on \(m\) and \(K_L\) only.
The assumption \(p \ge 2\) implies \(\log(mpk) \le C_{1,m}^2 \log(pk)\), where \(C_{1,m} = \{1 + \log(m)/\log(2)\}^{1/2}\) only depends on \(m\). Recalling that \(p,n \ge 2\) and \(k^{-1}\log(pn) \le 1\), and noting \(p \ge |\mathcal{I}|\) by definition of \(\mathcal{I}\), we find \[r = \sqrt{\frac{1}{k} \log\Big( \frac{m|\mathcal{I}|k^{1/4}}{\log^{5/4}(pn)}\Big)} \le \sqrt{\frac{1}{k} \log\big(m p k^{1/4}\big)} \le C_{1,m} \sqrt{ \frac{\log(pk)}{k}}\] and \[\delta = \frac{1}{m|\mathcal{I}|} \Big( \frac{\log^5(pn)}{k} \Big)^{1/4} \ge \frac{1}{m |\mathcal{I}| k^{1/4} } \ge \frac{1}{mpk^{1/4}}.\] Thus, noting that \(r \ge k^{-1/2}\) (this follows from \(\delta < e^{-1}\)) \[r\log\Big(\frac{1}{r\delta}\Big) \le r \log(mpk^{3/4}) \le r \log(mpk) \le C_{1,m}^3 \Big[\frac{\log^3(pk)}{k} \Big]^{1/2}.\] In summary, there exists a universal constant \(c_1\) and constant \(c_{2,m}\) depending only on \(m\) and \(K_L\) such that, with probability at least \(1-c_1\delta_n\), \[\label{eq:boot-bound-Delta-final} \Delta \leq c_{2,m} \Big[ h + \sqrt{r_{2,n}} + \frac{r_{2,n}^2}{h} + \frac{r_{2,n}}{\sqrt{h}} + \frac{1}{h\sqrt{k}} \Big\{ B_{n,k}(L_I ; A_I^{\oplus\kappa_L}) + \Big[\frac{\log^3(pk)}{k} \Big]^{1/4} \Big\} \Big]\tag{57}\] where \(r_{2,n} = \sqrt{k^{-1}\log(pk)}\) as defined in the theorem..
To bound \(d_K( \mathcal{L}(\boldsymbol{S}_n^\circ \mid \mathrm{data}), \boldsymbol{G}_n)\) we apply Theorem 3 from [43]. In the proof of Theorem 7, we verified that the conditions of that theorem are satisfied by \(X_i\) in their notation replaced with \(\sqrt{n}\boldsymbol{Y}_{i,n}\) in our notation with \(\underline{\sigma}^2 = \sigma^2_{\mathrm{min}}\), \(B_n = (m+1)(\log 2)^{-1}\sqrt{n/k}\) and \(\overline{\sigma}^2 = 4 (\log 2)^2 m (m+1)\). From this we obtain, for constants \(c_{3,m}, c_{4,m}\) that depend on \(m, \sigma_{\mathrm{min}}\) only, \[\label{eq:boot-bound-KS} d_K( \mathcal{L}(\boldsymbol{S}_n^\circ \mid \mathrm{data}), \boldsymbol{G}_n) \leq c_{3,m}\delta_n\tag{58}\] with probability at least \(1-c_{4,m}\delta_n\). Combining the bounds in 56 –58 completes the proof. ◻
Lemma 6. Recall the definitions of \(\boldsymbol{S}_n^*\) and \(\boldsymbol{S}_n^\circ\) from 21 and 55 , respectively.
If \(p\ge 2\), we have with probability one \[d_K( \mathcal{L}(\boldsymbol{S}_n^* \mid \mathrm{data}), \boldsymbol{G}_n) \lesssim \frac{1}{k} + \frac{\Delta \cdot \log(p+k)}{\sigma_{\mathrm{min}}^2} + d_K( \mathcal{L}(\boldsymbol{S}_n^\circ \mid \mathrm{data}), \boldsymbol{G}_n),\] where the constant in \(\lesssim\) is universal and where \[\begin{align} \label{eq:def-delta-and-si} \Delta^2 := \max_{I \in \mathcal{I}} \max_{\boldsymbol{x}_I \in A_I} \sum_{i=1}^n S_{i,I}^2(\boldsymbol{x}_I), \quad S_{i,I}(\boldsymbol{x}_I) := \widehat Y_{i,I}(\boldsymbol{x}_I) - Y_{i,I}(\boldsymbol{x}_I) + \frac{1}{n} \sum_{i'=1}^n Y_{i',I}(\boldsymbol{x}_I). \end{align}\tag{59}\]
Proof of Lemma 6. By the triangle inequality, we have \[d_K( \mathcal{L}(\boldsymbol{S}_n^* \mid \mathrm{data}), \boldsymbol{G}_n) \le d_K( \mathcal{L}(\boldsymbol{S}_n^* \mid \mathrm{data}), \mathcal{L}(\boldsymbol{S}_n^\circ \mid \mathrm{data})) + d_K( \mathcal{L}(\boldsymbol{S}_n^\circ \mid \mathrm{data}), \boldsymbol{G}_n).\] To bound \(d_K( \mathcal{L}(\boldsymbol{S}_n^* \mid \mathrm{data}), \mathcal{L}(\boldsymbol{S}_n^\circ \mid \mathrm{data}))\) we will apply Lemma 14 conditionally on the data. Write \(\mathbb{P}_e\) and \(\mathbb{E}_e\) for the conditional probability/expectation given the data \((\boldsymbol{X}_1, \dots, \boldsymbol{X}_n)\). Then, for any \(\lambda>0\), \[\begin{align} d_K( \mathcal{L}(\boldsymbol{S}_n^* \mid \mathrm{data}), \mathcal{L}(\boldsymbol{S}_n^\circ \mid \mathrm{data})) &\le \mathbb{P}_e( \| \boldsymbol{S}_n^* - \boldsymbol{S}_n^\circ\|_\infty \ge \lambda) \\ &+ \sup_{\boldsymbol{x} \in \mathbb{R}^p} \mathbb{P}_e( \boldsymbol{S}_n^\circ \le \boldsymbol{x}+\lambda \boldsymbol{1} ) - \mathbb{P}_e( \boldsymbol{S}_n^\circ \le \boldsymbol{x}-\lambda \boldsymbol{1}), \end{align}\] By the same calculation as in 50 in the proof of Theorem 7, we have \[\begin{align} &\phantom{{}={}} \mathbb{P}_e( \boldsymbol{S}_n^\circ \le \boldsymbol{x}+\lambda\boldsymbol{1} ) - \mathbb{P}_e( \boldsymbol{S}_n^\circ \le \boldsymbol{x}-\lambda\boldsymbol{1} ) \\&= \mathbb{P}( \boldsymbol{G}_n \le \boldsymbol{x}+\lambda\boldsymbol{1} ) - \mathbb{P}( \boldsymbol{G}_n \le \boldsymbol{x}-\lambda\boldsymbol{1} ) + \big\{ \mathbb{P}_e( \boldsymbol{S}_n^\circ \le \boldsymbol{x}+\lambda\boldsymbol{1} ) - \mathbb{P}( \boldsymbol{G}_n \le \boldsymbol{x}+\lambda\boldsymbol{1} ) \big\} \\& + \big\{ \mathbb{P}( \boldsymbol{G}_n \le \boldsymbol{x}-\lambda\boldsymbol{1} ) - \mathbb{P}_e( \boldsymbol{S}_n^\circ \le \boldsymbol{x}-\lambda\boldsymbol{1} ) \big\} \\&\le \frac{8\lambda}{\sigma_{\min}^2} \sqrt{\log p} + 2 d_K(\mathcal{L}(\boldsymbol{S}_n^\circ \mid \mathrm{data}), \boldsymbol{G}_n) \end{align}\] where we have used Theorem 18. Overall, \[\begin{align} \label{eq:bound-boot} d_K( \mathcal{L}(\boldsymbol{S}_n^* \mid \mathrm{data}), \boldsymbol{G}_n) \le \mathbb{P}_e( \| \boldsymbol{S}_n^* - \boldsymbol{S}_n^\circ\|_\infty \ge \lambda) + \frac{8\lambda}{\sigma_{\min}^2} \sqrt{\log p} + 3 d_K( \mathcal{L}(\boldsymbol{S}_n^\circ \mid \mathrm{data}), \boldsymbol{G}_n), \end{align}\tag{60}\] and it remains to choose \(\lambda\) appropriately and to bound the first summand on the right. For that purpose, write \[\|\boldsymbol{S}_n^* - \boldsymbol{S}_n^\circ\|_\infty = \max_{I \in \mathcal{I}} \max_{\boldsymbol{x}_I \in A_I} |D_{I}(\boldsymbol{x}_I)| ,\] where \[\begin{align} D_{I}(\boldsymbol{x}_I) := \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup ^*_{n,I}(\boldsymbol{x}_I) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup ^\circ_{n,I}(\boldsymbol{x}_I) = \sum_{i=1}^n e_i S_{i,I}(\boldsymbol{x}_I) \end{align}\] with \(S_{i,I}(\boldsymbol{x}_I)\) defined in the statement of the lemma. We also let \[\Delta_I^2(\boldsymbol{x}_I) := \sum_{i=1}^n S_{i,I}^2(\boldsymbol{x}_I)\] and note that \(\Delta^2 = \max_{I \in \mathcal{I}} \max_{\boldsymbol{x}_I \in A_I} \Delta_I^2(\boldsymbol{x}_I)\).
Since the multipliers \(e_1,\dots,e_n\) are standard Gaussian, we have \[\mathbb{P}_e( D_{I}(\boldsymbol{x}_I) \in \cdot) = \mathcal{N}(0, \Delta_I^2(\boldsymbol{x}_I))(\cdot).\] For \(\eta>0\), let \[\lambda = \mathbb{E}_e[\max_{I \in \mathcal{I}} \max_{\boldsymbol{x}_I \in A_I} |D_{I}(\boldsymbol{x}_I)| ] + \eta.\] The Borell-TIS inequality [59] then yields \[\begin{align} \mathbb{P}_e\Big( \max_{I \in \mathcal{I}} \max_{\boldsymbol{x}_I \in A_I} |D_{I}(\boldsymbol{x}_I)| > \lambda\Big) &= \mathbb{P}_e\Big( \max_{I \in \mathcal{I}} \max_{\boldsymbol{x}_I \in A_I} |D_{I}(\boldsymbol{x}_I)| > \mathbb{E}_e[\max_{I \in \mathcal{I}} \max_{\boldsymbol{x}_I \in A_I} |D_{I}(\boldsymbol{x}_I)| ] + \eta\Big) \\&\le \exp\Big(- \frac{\eta^2}{2 \max_{I \in \mathcal{I}} \max_{\boldsymbol{x}_I \in A_I} \mathbb{E}_e[|D_{I}(\boldsymbol{x}_I)|^2 ]}\Big) \\&= \exp\Big(- \frac{\eta^2}{2 \Delta^2}\Big). \end{align}\] Moreover, by the inequality at the beginning of Section 2.5 in [60], we have \[\mathbb{E}_e\Big[\max_{I \in \mathcal{I}} \max_{\boldsymbol{x}_I \in A_I} |D_{I}(\boldsymbol{x}_I)| \Big] \le \Delta \sqrt{2\log(2p)} \le 2 \Delta \sqrt{\log p},\] where the last inequality follows from \(p \ge 2\). Using these bounds and definitions, 60 yields \[\begin{align} d_K( \mathcal{L}(\boldsymbol{S}_n^* \mid \mathrm{data}), \boldsymbol{G}_n) & \le \nonumber \exp\Big(- \frac{\eta^2}{2 \Delta^2}\Big) + \frac{8}{\sigma_{\min}^2} \eta\sqrt{\log p} + \frac{16}{\sigma_{\min}^2} \Delta \log p \\&+ 3 d_K(\mathcal{L}(\boldsymbol{S}_n^\circ \mid \mathrm{data}), \boldsymbol{G}_n). \end{align}\] Setting \(\eta = \Delta \sqrt{2\log k}\) and noting that \(\log k, \log p \leq \log(p+k)\) completes the proof. ◻
The following two lemmas provide bounds on \(\sum_{i=1}^n S_{i,I}^2(\boldsymbol{x}_I)\) with \(S_{i,I}\) from 59 . Note that the first one is non-stochastic.
Lemma 7. Let \(I\subseteq[d]\), \(\boldsymbol{x}_I \in (0,1]^I\), and \(n/k \ge 2\). Assume there exists an \(\varepsilon\in(0,1)\) such that on the set \(\bar B_\varepsilon(\boldsymbol{x}_I)=\{ \boldsymbol{y}_I \in (0,\infty)^I: \| \boldsymbol{x}_I - \boldsymbol{y}_I \|_\infty \le \varepsilon\}\), all partial derivatives \(\partial_{j} L_I\) with \(j \in I\) exist and are Lipschitz-continuous with constant \(K_L\). Then, for any \(0< h < (\min_{j \in I} x_j) \wedge \varepsilon\), we have \[\begin{align} \label{eq:bound-on-si} \Delta_I^2(\boldsymbol{x}_I) = \sum_{i=1}^n S_{i,I}^2(\boldsymbol{x}_i) \lesssim |I|^2h^2 + \frac{|I|^2}{k} &{} \nonumber + \frac{|I|^2}{\sqrt k} \max_{j \in I}\sup_{y_j\in [x_j-h,x_j+h]} \big|\widetilde{\mathbb{L}}_{nj}(y_j)\big| \\&{} \nonumber + \frac{|I|^4}{k} \max_{j \in I}\sup_{y_j\in [x_j-h,x_j+h]} \big|\widetilde{\mathbb{L}}_{nj}(y_j)\big|^2 \\&{} \nonumber + \frac{1}{\sqrt{k}} | \widetilde{\mathbb{L}}_{n,I}(\boldsymbol{x}_I) | \\&{} \nonumber + \frac{|I|^2}{h^2k}\sup_{\boldsymbol{y}_I \in \bar B_{h}(\boldsymbol{x}_I)} \big|\mathbb{L}_{n,I}(\boldsymbol{y}_I) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}(\boldsymbol{y}_I)\big|^2 \\&{} + \frac{|I|^2}{h^2k} \omega_{\widetilde{\mathbb{L}}_{n,I}}(2h;\bar B_h(\boldsymbol{x}_I))^2. \end{align}\tag{61}\] where the implicit constant in \(\lesssim\) depends on \(K_L\) only.
Proof of Lemma 7. We start by introducing the notation \[\begin{align} \label{eq:J95iI} J_{i,I} = \{ \exists j \in I: V_{ij} < kx_j/n \}, \qquad \hat{J}_{i,I} = \{ \exists j \in I: \hat{V}_{ij} < kx_j/n \}, \end{align}\tag{62}\] and note that \(\mathbb{P}(J_{i,I}) = (k/n) \widetilde{\mu}_{n,I}(\boldsymbol{x}_I)\). Hence, \[\begin{align} S_{i,I}(\boldsymbol{x}_I) &\equiv \widehat Y_{i,I}(\boldsymbol{x}_I) - Y_{i,I}(\boldsymbol{x}_I) + \frac{1}{n} \sum_{i'=1}^n Y_{i',I}(\boldsymbol{x}_I) =\frac{1}{\sqrt k } \big( A_{i,I} - B_{i,I} - C_{i,I} + D_{i,I} \big) \end{align}\] where \[\begin{align} A_{i,I} &= \boldsymbol{1}(\hat{J}_{i,I}) - \boldsymbol{1}(J_{i,I}) \\ B_{i,I} &= \frac{k}{n} \Big\{ \widehat L_{n,I}(\boldsymbol{x}_I) - \widetilde{\mu}_{n,I}(\boldsymbol{x}_I) \Big\} \\ C_{i,I} &= \sum_{j \in I} \widehat{\partial_{j} L}_{I}(\boldsymbol{x}_I) \Big\{ \boldsymbol{1}(\hat{V}_{ij} < kx_j/n) - kx_j/n \Big\} - \partial_{j} L_{I}(\boldsymbol{x}_I) \Big\{ \boldsymbol{1}(V_{ij} < kx_j/n) - kx_j/n \Big\} \\ D_{i,I} &= \frac{1}{n} \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}(\boldsymbol{x}_I); \end{align}\] note that \(B_{i,I}\) and \(D_{i,I}\) do not depend on \(i\). As a consequence, since \((a+b+c+d)^2 \le 4(a^2+b^2+c^2+d^2)\), we obtain that \(\sum_{i=1}^n S_{i,I}^2(\boldsymbol{x}_i) \le 4(A^2+B^2+C^2+D^2)\), where \[A^2 = \frac{1}{k}\sum_{i=1}^n A_{i,I}^2, \qquad B^2 =\frac{n}{k} B_{1,I}^2, \qquad C^2 = \frac{1}{k}\sum_{i=1}^n C_{i,I}^2, \qquad D^2 = \frac{n}{k} D_{1,I}^2.\] A direct computation yields \[D^2 \le \frac{1}{kn} | \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}(\boldsymbol{x}_I)|^2 \le \frac{2}{kn} | \widetilde{\mathbb{L}}_{n,I}(\boldsymbol{x}_I) |^2 + \frac{2|I|^2}{kn} \max_{j\in I} | \widetilde{\mathbb{L}}_{nj}(x_j)|^2.\] We will further show below that \[\begin{align} \tag{63} A^2 &\le \frac{|I|}{\sqrt k} \max_{j \in I} | \widetilde{\mathbb{L}}_{nj}(x_j)| + \frac{|I|}{k}, \\ \tag{64} B^2 &\le \frac{3|I|^2}{n} \max_{j\in I} | \widetilde{\mathbb{L}}_{nj}(x_j)|^2 + \frac{3}{n} | \widetilde{\mathbb{L}}_{n,I}(\boldsymbol{x}_I) |^2 + \frac{3|I|^2}{kn}, \\ \tag{65} C^2 &\le 2|I|^2\max_{j\in I} \big| \widehat{\partial_{j} L}_{I}(\boldsymbol{x}_I) - \partial_{j} L_{I}(\boldsymbol{x}_I) \big|^2 + \frac{2|I|^2}{\sqrt k}\max_{j\in I} |\widetilde{\mathbb{L}}_{nj}(x_j)| + \frac{2|I|^2}{k}, \end{align}\] which in turn implies \[\begin{align} \sum_{i=1}^n S_{i,I}^2(\boldsymbol{x}_i) \le \frac{4|I|+(8+12/n)|I|^2}{k} &{} + \frac{4|I|+8|I|^2}{\sqrt k} \max_{j\in I} | \widetilde{\mathbb{L}}_{nj}(x_j)| \\&{} + \frac{(12+8/k)|I|^2}{n} \max_{j\in I} | \widetilde{\mathbb{L}}_{nj}(x_j)|^2 \\&{} + \frac{12+8/k}{n} | \widetilde{\mathbb{L}}_{n,I}(\boldsymbol{x}_I) |^2 \\&{} + 8|I|^2\max_{j\in I} \big| \widehat{\partial_{j} L}_{I}(\boldsymbol{x}_I) - \partial_{j} L_{I}(\boldsymbol{x}_I) \big|^2. \end{align}\] The squared terms involving \(| \widetilde{\mathbb{L}}_{nj}(x_j)|^2\) and \(| \widetilde{\mathbb{L}}_{n,I}(\boldsymbol{x}_I) |^2\) can be absorbed into the non-squared ones by using the trivial bounds \(| \widetilde{\mathbb{L}}_{nj}(x_j)| \le n/\sqrt k\) and \(| \widetilde{\mathbb{L}}_{n,I}(\boldsymbol{x}_I) | \le n/\sqrt k\). Further, it follows from Lemma 9 that \[\begin{align} \big| \widehat{\partial_{j} L}_{I}(\boldsymbol{x}_I)- \partial_{j} L_{I}(\boldsymbol{x}_I) \big|^2 & \le 4K_L^2 h^2 + \frac{4}{h^2k}\sup_{\boldsymbol{y}_I \in \bar B_{h}(\boldsymbol{x}_I)} \big|\mathbb{L}_{n,I}(\boldsymbol{y}_I) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}(\boldsymbol{y}_I)\big|^2 \\ &+ 4K_L^2 \frac{|I|^2}{k} \max_{j \in I}\sup_{y_j\in [x_j-h,x_j+h]} \big|\widetilde{\mathbb{L}}_{nj}(y_j)\big|^2 \\ &+ \frac{4}{h^2k} \omega_{\widetilde{\mathbb{L}}_{n,I}}(2h;\bar B_h(\boldsymbol{x}_I))^2. \end{align}\] Assembling terms we find the claimed bound in the formulation of the lemma.
It remains to show 63 65 . We start by showing 63 . For that purpose, note that \[\begin{align} \phantom{{}={}} \big| \boldsymbol{1}(\hat{J}_{i,I}) - \boldsymbol{1}(J_{i,I}) \big| &\le \sum_{j\in I} \big|\boldsymbol{1}(\hat{V}_{ij} < kx_j/n) - \boldsymbol{1}(V_{ij} < kx_j/n)\big|. \end{align}\] Subsequently, we fix \(j \in I\). By definition of \(\hat{V}_{ij}\), we have \(\hat{V}_{ij} < kx_j/n\) if and only if \(R_{ij}>n+1-kx_j\), which in turn is equivalent to \(V_{ij}<V_{\lceil kx_j \rceil:n,j}\), as shown at the beginning of the proof of Theorem 3. Hence, depending on whether \(V_{\lceil kx_j \rceil:n,j} < kx_j/n\) or not, we either have ‘\(\{\hat{V}_{ij} < kx_j/n\} \subseteq\{V_{ij} < kx_j/n\}\) for all \(i\in[n]\)’ or ‘\(\{V_{ij} < kx_j/n\} \subseteq\{\hat{V}_{ij} < kx_j/n\}\) for all \(i\in [n]\)’. It follows that all differences \(\boldsymbol{1}(\hat{V}_{ij} < kx_j/n) -\boldsymbol{1}(V_{ij}< k/n)\) with \(i \in [n]\) have the same sign, and we can rewrite \[\begin{align} \label{eq:sum-Ai-2} \sum_{i=1}^n \big|\boldsymbol{1}(\hat{V}_{ij} < kx_j/n) -\boldsymbol{1}(V_{ij} < kx_j/n)\big| &= \nonumber \Big|\sum_{i=1}^n \boldsymbol{1}(\hat{V}_{ij} < kx_j/n) -\boldsymbol{1}(V_{ij} < kx_j/n)\Big| \\&=\nonumber \Big|\sum_{i=1}^n \boldsymbol{1}(R_{ij} > n+1-\lceil kx_j \rceil ) -\boldsymbol{1}(V_{ij} < kx_j/n)\Big| \\&=\nonumber \Big|(\lceil kx_j \rceil -1) - k \widetilde{L}_{nj}(x_j)\Big| \\&\le \nonumber k |\widetilde{L}_{nj}(x_j) - x_j| + \big|(\lceil kx_j \rceil -1) - kx_j\big| \\&\le \sqrt k |\widetilde{\mathbb{L}}_{nj}(x_j) | + 1. \end{align}\tag{66}\] The previous two displays yield 63 .
We next show 64 . Note that \[\begin{align} B_{1,I} \nonumber = \frac{k}{n} \Big\{ \widehat L_{n,I}(\boldsymbol{x}_I) - \widetilde{\mu}_{n,I}(\boldsymbol{x}_I) \Big\} &= \frac{k}{n} \Big\{ \widehat L_{n,I}(\boldsymbol{x}_I) - \widetilde{L}_{n,I}(\boldsymbol{x}_I) + \widetilde{L}_{n,I}(\boldsymbol{x}_I) - \widetilde{\mu}_{n,I}(\boldsymbol{x}_I) \Big\} \\ &= \frac{k}{n} \Big\{ \widehat L_{n,I}(\boldsymbol{x}_I) - \widetilde{L}_{n,I}(\boldsymbol{x}_I) \Big\} + \frac{\sqrt{k}}{n}\widetilde{\mathbb{L}}_{n,I}(\boldsymbol{x}_I). \end{align}\] By the triangle inequality, we have \[\begin{align} \big|\widehat L_{n,I}(\boldsymbol{x}_I) - \widetilde{L}_{n,I}(\boldsymbol{x}_I) \big| \le \frac{1}{k} \sum_{i=1}^n \big| \boldsymbol{1}(\hat{J}_{i,I}) - \boldsymbol{1}(J_{i,I}) \big| &\le \label{eq:bound-for-b} \frac{|I|}{\sqrt k} \max_{j \in I} | \widetilde{\mathbb{L}}_{nj}(x_j)| + \frac{|I|}{k} \end{align}\tag{67}\] where we used 63 at the last inequality. The claimed identity in 64 then follows from combining the previous two displays and the inequality \((a+b+c)^2 \le 3(a^2+b^2+c^2)\).
We next show 65 , and for that purpose, note that \(C_{i,I}=\sum_{j \in I} C_{i,I,j}\), where \[\begin{align} C_{i,I,j} &\equiv \widehat{\partial_{j} L}_{I}(\boldsymbol{x}_I) \Big\{ \boldsymbol{1}(\hat{V}_{ij} < kx_j/n) - kx_j/n \Big\} - \partial_{j} L_{I}(\boldsymbol{x}_I) \Big\{ \boldsymbol{1}(V_{ij} < kx_j/n) - kx_j/n \Big\} \\&= \Big\{ \widehat{\partial_{j} L}_{I}(\boldsymbol{x}_I)- \partial_{j} L_{I}(\boldsymbol{x}_I) \Big\} \Big\{ \boldsymbol{1}(\hat{V}_{ij} < kx_j/n) - kx_j/n \Big\} \\ & + \partial_{j} L_{I}(\boldsymbol{x}_I) \Big\{\boldsymbol{1}(\hat{V}_{ij} < kx_j/n) - \boldsymbol{1}(V_{ij} < kx_j/n) \Big\} . \end{align}\] Next, \[\begin{align} \frac{1}{k} \sum_{i=1}^n \Big| \boldsymbol{1}(\hat{V}_{ij} < kx_j/n) - kx_j/n \Big|^2 &= \frac{1}{k} \Big\{ (1-2kx_j/n) \Big( \sum_{i=1}^n \boldsymbol{1}(\hat{V}_{ij} < kx_j/n) \Big) + k^2x_j^2/n \Big\} \\&= \frac{1}{k} \Big\{(1-2kx_j/n)(\lceil kx_j \rceil -1) + k^2x_j^2/n \Big\} \\& \le x_j (1-kx_j/n) \le x_j \le 1. \end{align}\] where we used the assumption that \(x_j \le 1 \le n/(2k)\) and the fact that \((\lceil kx_j \rceil -1) \le kx_j\). As a consequence, since \(0 \le \partial_{j} L(\boldsymbol{x}_I) \le 1\) and \((a+b)^2 \le 2(a^2+b^2)\), we obtain the bound \[\begin{align} \frac{1}{k} \sum_{i=1}^n C_{i,I,j}^2 &\le 2 \big| \widehat{\partial_{j} L}_{I}(\boldsymbol{x}_I)- \partial_{j} L_{I}(\boldsymbol{x}_I) \big|^2 + \frac{2}{k} \sum_{i=1}^n \big| \boldsymbol{1}(\hat{V}_{ij} \le kx_j/n) - \boldsymbol{1}(V_{ij} \le kx_j/n) \big|^2 \\&\le 2 \big| \widehat{\partial_{j} L}_{I}(\boldsymbol{x}_I)- \partial_{j} L_{I}(\boldsymbol{x}_I) \big|^2 + \frac{2}{\sqrt k} | \widetilde{\mathbb{L}}_{nj}(x_j)| + \frac{2}{k}, \end{align}\] where the last bound follows from 66 . This inequality, combined with \[\frac{1}{k} \sum_{i=1}^n C_{i,I}^2 \le \frac{1}{k} \sum_{i=1}^n |I| \sum_{j \in I} C_{i,I,j}^2 \le |I|^2 \max_{j \in I} \frac{1}{k} \sum_{i=1}^n C_{i,I,j}^2\] yields 65 . ◻
Lemma 8. Let \(L\) be a \(d\)-variate stable tail dependence function. Let \(\mathcal{I}\) be a collection of index sets \(I \subseteq[d]\) with \(|I| \ge 2\), and write \(m=\max_{I \in \mathcal{I}} |I|\). Let \((A_I)_{I \in \mathcal{I}}\) be a collection of sets with \(A_I \subseteq(0,1]^I\), and suppose that there exist \(\kappa_L, K_L\in(0,\infty)\) such that \[\begin{align} \forall I \in \mathcal{I}, \forall j \in I,& \forall \boldsymbol{x}_I \in {A_I^{\oplus \min(1,\kappa_L/2)}}, \forall \boldsymbol{y}_I \in [0,\infty)^I \text{ with } \|\boldsymbol{x}_I - \boldsymbol{y}_I\|_\infty \le \kappa_L: \\ &\partial_{j} L_I(\boldsymbol{x}_I), \partial_{j} L_I(\boldsymbol{y}_I) \text{ exist and satisfy } |\partial_{j} L_I(\boldsymbol{x}_I)-\partial_{j} L_I(\boldsymbol{y}_I)| \le K_L\|\boldsymbol{x}_I - \boldsymbol{y}_I\|_\infty. \end{align}\] Suppose further that \(n\in\mathbb{N}_{\ge 2}, k \in \mathbb{N}, \delta \in (0, e^{-1})\) satisfy \(\log(m/\delta)\le 2k/7\), \(n/k \ge 2\) and \(r = \sqrt{k^{-1} \log(1/\delta)} \le \kappa_L/(2^{3/2} C_{s})\) with \(C_{s}\) from Lemma 11. Then, for any \(h\) satisfying \[h < (\min_{I \in \mathcal{I}} \min_{\boldsymbol{x}_I \in A_I} \min_{j \in I} x_{I,j}) \wedge (\kappa_L/2),\] we have \[\Delta = \max_{I \in \mathcal{I}} \max_{\boldsymbol{x}_I \in A_I} \Delta_I(\boldsymbol{x}_I) \lesssim h + \sqrt{r} + \frac{r}{\sqrt{h}} + \frac{r^2}{h} + \frac{1}{h\sqrt{k}} \Big\{ B_{n,k}(L_I ; A_I^{\oplus\kappa_L}) + \sqrt{r \log\Big(\frac{1}{\delta r}\Big)} \Big\}\] with probability at least \(1-|\mathcal{I}|(6m+7)\delta\), where the implicit constant in \(\lesssim\) only depends on \(m\) and \(K_L\).
Proof of Lemma 8. Throughout the proof, \(\lesssim\) denotes inequality up to a constant only depending on \(m\) and \(K_L\). Fix some \(I \in \mathcal{I}\), and recall that \(|I| \le m\). We apply Lemma 7 with \(\varepsilon= (\kappa_L/2) \wedge 1\) and \(\boldsymbol{x}_I \in A_I\) to obtain that \[\begin{align} \label{eq:bound-on-si-uniform} \sup_{\boldsymbol{x}_I \in A_I} \Delta_I^2(\boldsymbol{x}_I) \lesssim h^2 + \frac{1}{k} + \frac{1}{\sqrt k}\sup_{\boldsymbol{y}_I \in [0,2]^I}\big|\widetilde{\mathbb{L}}_{n,I}(\boldsymbol{y}_I)\big| &{} \nonumber + \frac{1}{k}\sup_{\boldsymbol{y}_I \in [0,2]^I}\big|\widetilde{\mathbb{L}}_{n,I}(\boldsymbol{y}_I)\big|^2 \\&{} \nonumber + \frac{1}{h^2k}\sup_{\boldsymbol{y}_I \in A_I^{\oplus h}} \big|\mathbb{L}_{n,I}(\boldsymbol{y}_I) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}(\boldsymbol{y}_I)\big|^2 \\&{} + \frac{1}{h^2k} \sup_{\boldsymbol{x}_I \in A_I} \omega_{\widetilde{\mathbb{L}}_{n,I}}(2h;\bar B_h(\boldsymbol{x}_I))^2. \end{align}\tag{68}\] where we have used that, for each \(\boldsymbol{x}_I \in A_I \subseteq(0,1]^I\), \[\max \Big(| \widetilde{\mathbb{L}}_{n,I}(\boldsymbol{x}_I) |, \max_{j \in I}\sup_{y_j\in [x_j-h,x_j+h]} \big|\widetilde{\mathbb{L}}_{nj}(y_j)\big| \Big) \le \sup_{\boldsymbol{y}_I \in [0,2]^I}\big|\widetilde{\mathbb{L}}_{n,I}(\boldsymbol{y}_I)\big|,\] (recall that \(h < \varepsilon\le 1\)). We need to bound each term on the right-hand side of 68 . First, by Lemma 10, we have \[\begin{align} \label{eq:sni-proof-1} \frac{1}{\sqrt k}\sup_{\boldsymbol{y}_I \in [0,2]^I}\big|\widetilde{\mathbb{L}}_{n,I}(\boldsymbol{y}_I)\big| \lesssim \sqrt{\frac{2}{k} \log\Big(\frac{1}{\delta}\Big)} \lesssim r \end{align}\tag{69}\] on an event \(\Omega_{I,1}\) with probability at least \(1-\delta\). Moreover, since \(r = \sqrt{k^{-1} \log(1/\delta)} \le \sqrt{2/7} < 1\) by our assumption \(\log(m/\delta) \le 2k/7\), the same upper bound holds true for the squared term \(k^{-1}\sup_{\boldsymbol{y}_I \in [0,2]^I}|\widetilde{\mathbb{L}}_{n,I}(\boldsymbol{y}_I)|^2\).
Next, we apply Theorem 2 with \(T=2\) (note that \(n/k \ge 2\) by assumption), \(L=L_I\) and \(A= A_I^{\oplus h}\); note that \(A_I^{\oplus h} \subseteq{A_I^{\oplus \min(1,\kappa_L/2)}}\) such that \((A_I^{\oplus h}, L_I)\) satisfies [cond:smoothness-hoelder] with \(\alpha_L =1\) by our assumption on \(L\). Further note that \(r(\delta,2,k)\) in Theorem 2 is equal to \(\sqrt 2 r = \sqrt{2} r(\delta,1,k)\) in our current notation. We obtain that \[\sup_{\boldsymbol{y}_I \in A_I^{\oplus h}} \big|\mathbb{L}_{n,I}(\boldsymbol{x}) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}(\boldsymbol{x})\big| \lesssim B_{n,k}(L_I ; A_I^{\oplus h+C_{s}\sqrt 2r}) + \frac{1}{\sqrt{k}} + \sqrt{r \log\Big(\frac{1}{\delta r}\Big)} + r\sqrt{\log\Big(\frac{1}{\delta}\Big)}\] on an event \(\Omega_{I,2}\) with probability at least \(1-(6m+5)\delta\). Since \(r \le \sqrt{2/7}<1\) as noted earlier, and \(\delta < 1/e\), we have \[\frac{1}{\sqrt{k}} + r\sqrt{\log\Big(\frac{1}{\delta}\Big)} \lesssim \sqrt{r \log\Big(\frac{1}{\delta r}\Big)}.\] Next, since \(h + C_{s}\sqrt 2 r \le \kappa_L/2+\kappa_L/2=\kappa_L\) by assumption, we have \[B_{n,k}(L_I ; A_I^{\oplus h+C_{s}\sqrt{2} r}) \le B_{n,k}(L_I ; A_I^{\oplus \kappa_L}).\] Overall, \[\begin{align} \label{eq:sni-proof-2} \frac{1}{h^2k}\sup_{\boldsymbol{y}_I \in A_I^{\oplus h}} \big|\mathbb{L}_{n,I}(\boldsymbol{x}) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,I}(\boldsymbol{x})\big|^2 \lesssim \frac{1}{h^2k} \Big\{ B_{n,k}^2(L_I ; A_I^{\oplus\kappa_L}) + r \log\Big(\frac{1}{\delta r}\Big) \Big\}. \end{align}\tag{70}\] Next, from Lemma 12 we get \[\omega_{\widetilde{\mathbb{L}}_{n,I}}(2h;\bar B_h(\boldsymbol{x}_I)) = \sqrt{\frac{n}{k}} \omega_{\beta_{n,I}}\Big(\frac{k}{n} 2h ; \frac{k}{n}[\boldsymbol{x}_I - h \boldsymbol{1}_I,\boldsymbol{x}_I + h \boldsymbol{1}_I]\Big) \le \kappa \sqrt{2h \log(2 |I|/\delta)}\] on an event \(\Omega_{I,3}\) with probability at least \(1-\delta\), where \[\kappa = 2|I|\bigg[\sqrt{ \frac{2}{9kh}\log({2|I|}/{\delta})} + 2 + 60 \sqrt{2|I|}\bigg] \lesssim \Big(\frac{\log(1/\delta)}{kh}\Big)^{1/2} + 1.\] As a consequence, on \(\Omega_{I,3}\), \[\begin{align} \label{eq:sni-proof-3} \frac{1}{h^2k} \omega_{\widetilde{\mathbb{L}}_{n,I}}(2h;\bar B_h(\boldsymbol{x}_I))^2 \lesssim \frac{1}{kh} \kappa^2 \log(1/\delta) \lesssim \Big(\frac{\log(1/\delta)}{kh}\Big)^2 + \Big(\frac{\log(1/\delta)}{kh}\Big) = \frac{r^4}{h^2} + \frac{r^2}{h}. \end{align}\tag{71}\] Overall, combining 68 with 69 , 70 and 71 and the fact that \(k^{-1/2} \le r\), we find that, on the event \(\Omega_{I,1} \cap \Omega_{I,2} \cap \Omega_{I,3}\), \[\sup_{\boldsymbol{x}_I \in A_I} \Delta_I^2(\boldsymbol{x}_I) \lesssim h^2 + r + \frac{r^2}{h} + \frac{r^4}{h^2} + \frac{1}{h^2k} \Big\{ B_{n,k}^2(L_I ; A_I^{\oplus \kappa_L}) + r \log\Big(\frac{1}{\delta r}\Big) \Big\}.\] Moreover, \(\mathbb{P}(\Omega_{I,1} \cap \Omega_{I,2} \cap \Omega_{I,3}) \ge 1-(6m+7)\delta\). The assertion regarding the maximum over \(I \in \mathcal{I}\) then follows from the union bound. ◻
Lemma 9. Let \(L\) be a \(d\)-variate stable tail dependence function and let \(\boldsymbol{x} \in (0,\infty)^d\). Assume there exists an \(\varepsilon>0\) such that on the set \(\bar B_\varepsilon(\boldsymbol{x})=\{ \boldsymbol{y} \in (0,\infty)^d: \| \boldsymbol{x} - \boldsymbol{y} \|_\infty \le \varepsilon\}\), the partial derivatives \(\partial_{j} L\) exist and are Lipschitz-continuous with constant \(K_L\). Then, for any \(0<h< \varepsilon\wedge (\min_{j \in [d]}x_j)\), we have \[\begin{align} \max_{j \in [d]} \big|\widehat{\partial_{j} L}(\boldsymbol{x}) - \partial_{j} L(\boldsymbol{x}) \big| \le K_Lh &+ \frac{1}{h\sqrt{k}}\sup_{\boldsymbol{y} \in \bar B_{h}(\boldsymbol{x})} \big|\mathbb{L}_{n}(\boldsymbol{y}) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n}(\boldsymbol{y})\big| \\ &+ K_L\frac{d}{\sqrt k} \max_{j \in [d]}\sup_{y_j\in [x_j-h,x_j+h]} \big|\widetilde{\mathbb{L}}_{nj}(y_j)\big|\\ &+ \frac{1}{h\sqrt k} \omega_{\widetilde{\mathbb{L}}_{n}}(2h;\bar B_h(\boldsymbol{x})). \end{align}\]
Proof. Note that \(|\min(a,1) - b| \le |a-b|\) for \(a \in \mathbb{R}, b \in [0,1]\). Together with the triangle inequality this yields \[\begin{align} \label{eq:pd-bound} |\widehat{\partial_{j} L}(\boldsymbol{x}) - \partial_{j} L(\boldsymbol{x}) | &\le \Big|\frac{\widehat L_n(\boldsymbol{x} + h \boldsymbol{e}_j) - L(\boldsymbol{x} + h \boldsymbol{e}_j)}{2h} - \frac{\widehat L_n(\boldsymbol{x} - h \boldsymbol{e}_j) - L(\boldsymbol{x} - h \boldsymbol{e}_j)}{2h}\Big| \nonumber\\ & + \Big|\frac{L(\boldsymbol{x} + h \boldsymbol{e}_j) - L(\boldsymbol{x} - h \boldsymbol{e}_j)}{2h} - \partial_{j} L(\boldsymbol{x})\Big| \nonumber\\ &= \Big|\frac{\mathbb{L}_n(\boldsymbol{x} + h \boldsymbol{e}_j) - \mathbb{L}_n(\boldsymbol{x} - h \boldsymbol{e}_j)}{2h\sqrt{k}} \Big| + \Big|\frac{L(\boldsymbol{x} + h \boldsymbol{e}_j) - L(\boldsymbol{x} - h \boldsymbol{e}_j)}{2h} - \partial_{j} L(\boldsymbol{x})\Big|. \end{align}\tag{72}\] We start with the second term on the right hand side. By the mean value theorem, there exists some \(t\in (-1,1)\) such that \[\frac{L(\boldsymbol{x} + h \boldsymbol{e}_j) - L(\boldsymbol{x} - h \boldsymbol{e}_j)}{2h} = \partial_{j} L(\boldsymbol{x} + t h \boldsymbol{e}_j).\] Using the Lipschitz continuity of \(\partial_{j} L\), we obtain \[\Big| \frac{L(\boldsymbol{x} + h \boldsymbol{e}_j) - L(\boldsymbol{x} - h \boldsymbol{e}_j)}{2h} - \partial_{j} L(\boldsymbol{x}) \Big| \le K_L|t| h \le K_Lh.\] For the first term on the right hand side of 72 , again using the triangle inequality, we have \[\begin{align} &\phantom{{}={}} \big| \mathbb{L}_n(\boldsymbol{x} + h \boldsymbol{e}_j) - \mathbb{L}_n(\boldsymbol{x} - h \boldsymbol{e}_j) \big| \\ &\le \big| \mathbb{L}_n(\boldsymbol{x} + h \boldsymbol{e}_j) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x} + h \boldsymbol{e}_j) \big| + \big| \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x} + h \boldsymbol{e}_j) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x} - h \boldsymbol{e}_j) \big| \\& + \big| \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x} - h \boldsymbol{e}_j) - \mathbb{L}_n(\boldsymbol{x} - h \boldsymbol{e}_j) \big| \\& \le 2\sup_{\boldsymbol{y} \in \bar B_{h}(\boldsymbol{x})} \big| \mathbb{L}_n(\boldsymbol{y}) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{y}) \big| + \big| \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x} + h \boldsymbol{e}_j) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x} - h \boldsymbol{e}_j) \big|. \end{align}\] It remains to show that \[\big| \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x} + h \boldsymbol{e}_j) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x} - h \boldsymbol{e}_j) \big| \le {2K_Ldh} \max_{j \in [d]} \sup_{y_j\in [x_j-h,x_j]} \big|\widetilde{\mathbb{L}}_{nj}(y_j)\big| + 2 \omega_{\widetilde{\mathbb{L}}_{n}}(2h;\bar B_{h}(\boldsymbol{x}))\] By definition of \(\begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n}\), for any \(\boldsymbol{y}, \boldsymbol{y}' \in \bar B_\varepsilon(\boldsymbol{x})\), we have \[\begin{align} &\phantom{{}={}} \big| \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n}(\boldsymbol{y}) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n}(\boldsymbol{y}')\big| \\ &\le \big|\widetilde{\mathbb{L}}_{n}(\boldsymbol{y}) - \widetilde{\mathbb{L}}_{n}(\boldsymbol{y}')\big| + \sum_{\ell \in [d]} \big|\partial_{\ell} L(\boldsymbol{y})\widetilde{\mathbb{L}}_{n\ell}(y_\ell) - \partial_{\ell} L(\boldsymbol{y}')\widetilde{\mathbb{L}}_{n\ell}(y_\ell')\big| \\ &\le \big|\widetilde{\mathbb{L}}_{n}(\boldsymbol{y}) - \widetilde{\mathbb{L}}_{n}(\boldsymbol{y}')\big| + \sum_{\ell \in [d]} \Big\{ \big|\partial_{\ell} L(\boldsymbol{y}) \big|\times\big|\widetilde{\mathbb{L}}_{n\ell}(y_\ell) - \widetilde{\mathbb{L}}_{n\ell}(y_\ell')\big| \\ & + \big|\widetilde{\mathbb{L}}_{n\ell}(y_\ell')\big|\times\big|\partial_{\ell} L(\boldsymbol{y}) - \partial_{\ell} L(\boldsymbol{y}')\big|\Big\} \\ &\le \big|\widetilde{\mathbb{L}}_{n}(\boldsymbol{y}) - \widetilde{\mathbb{L}}_{n}(\boldsymbol{y}')\big| + \sum_{\ell \in [d]} \Big\{\big|\widetilde{\mathbb{L}}_{n\ell}(y_\ell) - \widetilde{\mathbb{L}}_{n\ell}(y_\ell')\big| + \big|\widetilde{\mathbb{L}}_{n\ell}(y_\ell')\big|\times K_L\big\|\boldsymbol{y} -\boldsymbol{y}'\big\|_\infty\Big\}, \end{align}\] where we used \(|\partial_{\ell} L|\le 1\) and Lipschitz-continuity of the partial derivatives. For \(\boldsymbol{y} = \boldsymbol{x} + h \boldsymbol{e}_j\) and \(\boldsymbol{y}' = \boldsymbol{x} - h \boldsymbol{e}_j\), we obtain \[\big|{\widetilde{\mathbb{L}}}_{n}(\boldsymbol{y}) - {\widetilde{\mathbb{L}}}_{n}(\boldsymbol{y}')\big| \le \omega_{{\widetilde{\mathbb{L}}}_{n}}(2h;\bar B_{h}(\boldsymbol{x})).\] The term \(|\widetilde{\mathbb{L}}_{n\ell}(y_\ell) - \widetilde{\mathbb{L}}_{n\ell}(y_\ell')|\) equals zero for \(\ell \ne j\) and is bounded by \(\omega_{{\widetilde{\mathbb{L}}}_{n}}(2h;\bar B_{h}(\boldsymbol{x}))\) for \(\ell=j\). Finally, it holds that \[|\widetilde{\mathbb{L}}_{n\ell}(y_\ell')| \le \sup_{y_\ell \in [x_\ell-h, x_\ell]} |\widetilde{\mathbb{L}}_{n\ell}(y_\ell)|\] and \(\big\|\boldsymbol{y} -\boldsymbol{y}'\big\|_\infty = 2h.\) Combining the previous results yields the assertion. ◻
Proof of Theorem 14. Without loss of generality, we can assume that \(\log^5(pn) / k \le 1\); otherwise, the result is trivial. Under \(H(\rho)\), we can rewrite \[\begin{align} T_n^{(\rho)} &= \max_{(t,\boldsymbol{s}_1, \boldsymbol{s}_2,\boldsymbol{s}_1',\boldsymbol{s}_2') \in D(\rho)} \mathbb{D}_{n, (\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1', \boldsymbol{s}_2')}(t) \end{align}\] where \(D(\rho)\) is from 29 and where \[\begin{align} \mathbb{D}_{n, (\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1', \boldsymbol{s}_2')}(t) &= \mathbb{L}_{n,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(1-t,t) - \mathbb{L}_{n,(\boldsymbol{s}_1', \boldsymbol{s}_2')}(1-t,t) \end{align}\] with \[\mathbb{L}_{n,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2) = \sqrt k \{ \hat{L}_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2) - L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2) \}.\] Likewise, we can rewrite \[T_n^{(\rho),*} = \max_{(t,\boldsymbol{s}_1, \boldsymbol{s}_2,\boldsymbol{s}_1',\boldsymbol{s}_2') \in D(\rho)} \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{D}} \endgroup ^*_{n, (\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1', \boldsymbol{s}_2')}(t)\] where \[\begin{align} \label{eq:dbstar} \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{D}} \endgroup ^*_{n, (\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1', \boldsymbol{s}_2')}(t) = \sum_{i=1}^n e_{i} \big\{ \hat{Y}_{i,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(1-t,t) - \hat{Y}_{i,(\boldsymbol{s}_1', \boldsymbol{s}_2')}(1-t,t) \big\} \end{align}\tag{73}\] with \(\hat{Y}_{i, (\boldsymbol{s}_1, \boldsymbol{s}_2)}\) from 28 .
Next, recall \(D= D(1, \sqrt 2) = D(1) \cup D(\sqrt 2)\) with \(p = |D|\), and consider the stacked vectors \[\begin{align} \widetilde{\boldsymbol{S}}_n &= \big(\mathbb{D}_{n, (\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1', \boldsymbol{s}_2')}(t)\big)_{(t, \boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2') \in D} \in \mathbb{R}^p, \\ \widetilde{\boldsymbol{T}}_n &= \big( \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{D}} \endgroup _{n, (\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1', \boldsymbol{s}_2')}(t)\big)_{(t, \boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2') \in D} \in \mathbb{R}^p, \\ \widetilde{\boldsymbol{S}}_n^* &= \big( \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{D}} \endgroup ^*_{n, (\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1', \boldsymbol{s}_2')}(t)\big)_{(t, \boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2') \in D} \in \mathbb{R}^p, \end{align}\] where \(\begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{D}} \endgroup ^*_{n, (\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1', \boldsymbol{s}_2')}(t)\) is from 73 and where \[\begin{align} \label{eq:bardbn} \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{D}} \endgroup _{n, (\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1', \boldsymbol{s}_2')}(t) &= \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(1-t,t) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,(\boldsymbol{s}_1', \boldsymbol{s}_2')}(1-t,t) \end{align}\tag{74}\] with \[\begin{align} \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2) = \widetilde{\mathbb{L}}_{n,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2) & - \sum_{j \in [2]}\partial_j L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2) \widetilde{\mathbb{L}}_{n,\boldsymbol{s}_j}(x_j), \end{align}\] and, with \(V_i(\boldsymbol{s}) = 1- F_{\boldsymbol{s}}(X_i(\boldsymbol{s}))\), \[\begin{align} \widetilde{\mathbb{L}}_{n,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2) &= \frac{1}{\sqrt k} \sum_{i=1}^n \Big\{ \boldsymbol{1}\Big(\exists j \in \{1,2\}: V_i(\boldsymbol{s}_j) \le \frac{k}{n} x_j\Big) \\ & - \mathbb{P}\Big( \exists j \in \{1,2\}: V_i(\boldsymbol{s}_j) \le \frac{k}{n} x_j\Big)\Big\}, \\ \widetilde{\mathbb{L}}_{n,\boldsymbol{s}}(x) &= \frac{1}{\sqrt k} \sum_{i=1}^n \Big\{ \boldsymbol{1}\Big( V_{i}(\boldsymbol{s}) \le \frac{k}{n} x\Big) - \mathbb{P}\Big( V_{i}(\boldsymbol{s}) \le \frac{k}{n} x\Big)\Big\}. \end{align}\] Further, let \[\widetilde{\boldsymbol{G}}_n \sim \mathcal{N}_{p}(\boldsymbol{0}, \operatorname{Var}(\widetilde{\boldsymbol{T}}_n)).\] Note that \[T_n^{(\rho)} = \max_{q \in P(\rho)} \widetilde{S}_{nq}, \qquad T_n^{(\rho),*} = \max_{q \in P(\rho)} \widetilde{S}_{nq}^*\] and define \[\begin{align} T_n^{(\rho),g} = \max_{q \in P(\rho)} \widetilde{G}_{nq}^{(\rho)}, \end{align}\] where \(P(\rho)\) corresponds to all coordinate indices of \(\widetilde{\boldsymbol{S}}_n\) for which \({(t, \boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2') \in D(\rho)}\). Define bivariate random vectors \[\boldsymbol{Y}_n = (T_n^{(1)}, T_n^{(\sqrt 2)}), \quad \boldsymbol{Y}_n^{g} = (T_n^{(1),g}, T_n^{(\sqrt 2),g}), \quad \boldsymbol{Y}_n^* = (T_n^{(1),*}, T_n^{(\sqrt 2),*}),\] and note that \[\begin{align} \mathbb{P}\big(Y_{n1} \le x, Y_{n2} \le y\big) &= \mathbb{P}\big( \widetilde{\boldsymbol{S}}_n \le \boldsymbol{s}_{x,y} \big), \\ \mathbb{P}\big( Y_{n1}^{g} \le x, Y_{n2}^{g} \le y) &= \mathbb{P}\big(\widetilde{\boldsymbol{G}}_n \le \boldsymbol{s}_{x,y} \big), \\ \mathbb{P}\big( Y_{n1}^*\le x, Y_{n2}^* \le y \mid \mathrm{data}\big) &= \mathbb{P}\big( \widetilde{\boldsymbol{S}}_n^* \le \boldsymbol{s}_{x,y} \mid \mathrm{data} \big), \end{align}\] where \(\boldsymbol{s}_{x,y} \in \mathbb{R}^p\) is the vector with coordinates \(s_j = x\) for \(j \in P(1)\) and \(s_j = y\) for \(j \in P(\sqrt 2)\). Hence, \[\begin{align} \label{eq:bound-K-dist} d_K\Big( \boldsymbol{Y}_n, \boldsymbol{Y}_n^{g} \Big) &\le d_K(\widetilde{\boldsymbol{S}}_n, \widetilde{\boldsymbol{G}}_n), \qquad d_K\Big( \mathcal{L}\big( \boldsymbol{Y}_n^* \mid \mathrm{data} \big), \boldsymbol{Y}_n^{g} \Big) \le d_K( \mathcal{L}(\widetilde{\boldsymbol{S}}_n^* \mid \mathrm{data} ), \widetilde{\boldsymbol{G}}_n), \end{align}\tag{75}\] which will eventually allow to apply Proposition 20.
In the following, let \[\begin{align} \label{eq:deltan} \delta_n = \Big( \frac{\log^5(pn)}{k} \Big)^{1/4} . \end{align}\tag{76}\] We will show below that there exist constants \(c_1, c_2, c_3\) only depending on \(K_L, \sigma_{\min}^2c_h, c_h'\) such that \[\begin{align} \tag{77} d_K(\widetilde{\boldsymbol{S}}_n, \widetilde{\boldsymbol{G}}_n) &\le c_1 \big[ \delta_n + \sqrt{\log p} B_{n,k} \big], \\ \tag{78} d_K( \mathcal{L}(\widetilde{\boldsymbol{S}}_n^* \mid \mathrm{data} ), \widetilde{\boldsymbol{G}}_n) &\le c_2 \big[\delta_n + \sqrt{\log(p+k)}\, B_{n,k} \big] \end{align}\] the latter holding with probability at least \(1-c_3\delta_n\). In view of 75 , an application of Proposition 20 implies that, for some constant \(c_0\) depending on \(c_1, c_2, c_3\), \[\Big|\mathbb{P}( C_n \le \hat{q}_{n,\alpha}^* ) - \alpha \Big| \le c_0 \big[\delta_n + \sqrt{\log(p+k)}\, B_{n,k} \big]\] as asserted.
It remains to show 77 and 78 . We start with the former and begin by observing that \[\begin{align} \label{eq:bound-dn1} \max_{(t,\boldsymbol{s}_1, \boldsymbol{s}_2,\boldsymbol{s}_1',\boldsymbol{s}_2') \in D(1,\sqrt 2)} & \big| \mathbb{D}_{n, (\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1', \boldsymbol{s}_2')}(t) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{D}} \endgroup _{n, (\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1', \boldsymbol{s}_2')}(t) \big| \nonumber\\&\le 2 \max_{(\boldsymbol{s}_1, \boldsymbol{s}_2) \in \mathcal{P}_d(1, \sqrt 2)} \max_{t \in A} \big| \mathbb{L}_{n,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(1-t,t) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(1-t,t) \big|. \end{align}\tag{79}\] An application of Theorem 2 and the union bound implies that there exist constants \(D_1=D_1(K_L)\) and \(D_2\) such that, for \(\delta>0\) specified below, \[\begin{align} \label{eq:bound-dn2} &\max_{(\boldsymbol{s}_1, \boldsymbol{s}_2) \in \mathcal{P}_d(\rho)} \max_{t \in A} \big| \mathbb{L}_{n,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(1-t,t) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _{n,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(1-t,t) \big| \nonumber \\&\le B_{n,k} + \frac{2}{\sqrt k} + D_1 \sqrt{r \log \Big( \frac{D_2}{\delta r} \Big)} =: \lambda_{n,k}(\delta) \end{align}\tag{80}\] with probability at least \(1-17|\mathcal{P}_d(1, \sqrt 2)| \delta\). Here, we note that the last term in the final bound of 2 can be absorbed into the second-to last term at the cost of possibly increasing the constant since \(\alpha_L = 1\).
We now proceed as in the proof of Theorem 7, and obtain that, for any \(\lambda>0\), \[d_K(\widetilde{\boldsymbol{S}}_n,\widetilde{\boldsymbol{G}}_n) \le \mathbb{P}(\| \widetilde{\boldsymbol{S}}_n -\widetilde{\boldsymbol{T}}_n \|_\infty \ge \lambda) + \frac{8\lambda}{{\sigma}^2_{\min}}\sqrt{\log p} + 3d_K(\widetilde{\boldsymbol{T}}_n , \widetilde{\boldsymbol{G}}_n);\] see the derivations in 49 and 50 . With \(\lambda = \lambda_{n,k}(\delta)\) from 80 , we obtain that, for \(\delta\) chosen below, \[d_K(\widetilde{\boldsymbol{S}}_n,\widetilde{\boldsymbol{G}}_n) \le 17|\mathcal{P}_d(1, \sqrt 2)| \delta + \frac{8\lambda_{n,k}(\delta)}{{\sigma}^2_{\min}}\sqrt{\log p} + 3d_K(\widetilde{\boldsymbol{T}}_n , \widetilde{\boldsymbol{G}}_n).\] We proceed by bounding \(d_K(\widetilde{\boldsymbol{T}}_n , \widetilde{\boldsymbol{G}}_n)\). The coordinates of \(\widetilde{\boldsymbol{T}}_n\) are of the form \[\begin{align} \label{eq:yitssss} \sum_{i=1}^n Y_{i,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(1-t,t) - Y_{i,(\boldsymbol{s}_1', \boldsymbol{s}_2')}(1-t,t) =: \sum_{i=1}^n Y_{i,(t,\boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2')} \end{align}\tag{81}\] where \[\begin{align} Y_{i,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2) &= \frac{1}{\sqrt k} \Big[ \boldsymbol{1}(\exists j \in [2]: V_{i}(\boldsymbol{s}_j) < kx_j/n) - (k/n) L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(\boldsymbol{x}) \\ & -\sum_{j \in [2]}\partial_{j} L_{(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2) \big\{\boldsymbol{1}( V_{i}(\boldsymbol{s}_j) < kx_j/n) - kx_j/n \big\} \Big]. \end{align}\] Note the resemblance between \(Y_{i,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2)\) and \(Y_{I}(\boldsymbol{x}_I)\) from 19 , which allows us to make use of results in the proof of Theorem 7. For instance, using that \((a-b)^4 \le 8 (a^4 + b^4)\), we have \[\sum_{i=1}^n \mathbb{E}[|Y_{i,(t,\boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2')}|^4] \le 16 \max_{(\boldsymbol{s}_1, \boldsymbol{s}_2) \in \mathcal{P}_d(1, \sqrt 2)} \max_{t \in A} \sum_{i=1}^n \mathbb{E}[|Y_{i,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(1-t, t)|^4 ],\] and the right-hand side has been bounded in the proof of Theorem 7 by some constant times \(n^{-1}\); see 52 . Similar elementary computations bound \(|Y_{i,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2)|\), and the result follows by applying Theorem 19 to obtain \(d_K(\widetilde{\boldsymbol{T}}_n , \widetilde{\boldsymbol{G}}_n) \le \tilde{c}_1\delta_n,\) where \(\tilde{c}_1\) depends only on \(\sigma_{\min}^2\) and where \(\delta_n\) is from 76 .
Finally, letting \(\delta = |\mathcal{P}_d(1, \sqrt 2)|^{-1} \delta_n \le 3^{-1} < e^{-1}\), we will verify that the conditions of Theorem 2 hold. Specifically, using (ii) and (iii) we see \[\begin{align} \log(2/\delta) &\le \log (2 |\mathcal{P}_d(1, \sqrt 2)| k^{1/4}) \le 2k/7, \\ r := \sqrt{k^{-1} \log(\delta^{-1})}&\le \sqrt{k^{-1}\log (|\mathcal{P}_d(1, \sqrt 2)|k^{1/4})} \le \kappa_L/C_{s}. \end{align}\] Finally, combining the bounds \(r \ge k^{-1/2}, |\mathcal{P}_d(1, \sqrt 2)| \ge 3 \gtrsim D_2\) and \(|\mathcal{P}_d(1, \sqrt 2)| \le p\), elementary computations show that \[\sqrt{\log p} \sqrt{r \log\Big(\frac{D_2}{\delta r}\Big)} \le D_2' \delta_n\] for some constant \(D_2'\) depending on \(D_2\). Assembling terms, this implies 77 .
It remains to show 78 , for which we proceed as in the proof of Theorem 11, which itself is a consequence of Proposition 15. The proof of Proposition 15 is based on Lemmas 6-9. We will now discuss how these lemmas and their proofs can be adapted to the present setting.
Let \[\begin{align} \widetilde{\boldsymbol{S}}_n^\circ = \big( \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{D}} \endgroup ^\circ_{n, (\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1', \boldsymbol{s}_2')}(t)\big)_{(t, \boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2') \in D} \in \mathbb{R}^p, \end{align}\] where \[\begin{align} \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{D}} \endgroup ^\circ_{n, (\boldsymbol{s}_1, \boldsymbol{s}_2), (\boldsymbol{s}_1', \boldsymbol{s}_2')}(t) = \sum_{i=1}^n e_i \Big\{Y_{i,(t,\boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2')} - \frac{1}{n} \sum_{i'=1}^n Y_{i',(t,\boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2')} \Big\}, \end{align}\] with \(Y_{i,(t,\boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2')}\) as defined in 81 . A careful inspection of the proof of Lemma 6 shows that it continues to holds with \(\boldsymbol{S}^*_n, \boldsymbol{S}_n^\circ\) and \(\boldsymbol{G}_n\) replaced by the respective tilde-versions. More specifically, we have \[d_K( \mathcal{L}(\widetilde{\boldsymbol{S}}_n^* \mid \mathrm{data}), \widetilde{\boldsymbol{G}}_n) \lesssim \frac{1}{k} + \frac{\widetilde{\Delta} \cdot \log(p+k)}{\sigma_{\mathrm{min}}^2} + d_K( \mathcal{L}(\widetilde{\boldsymbol{S}}_n^\circ \mid \mathrm{data}), \widetilde{\boldsymbol{G}}_n),\] where the constant in \(\lesssim\) is universal and where \[\begin{align} \widetilde{\Delta}^2 := \max_{(t, \boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2') \in D} \widetilde{\Delta}^2_{t, \boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2'}, \end{align}\] with \[\begin{align} \widetilde{\Delta}^2_{t, \boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2'} = \sum_{i=1}^n \Big\{\hat{Y}_{i,(t,\boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2')} - Y_{i,(t,\boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2')} - \frac{1}{n} \sum_{i'=1}^n Y_{i',(t,\boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2')} \Big\}^2 \end{align}\] and \[\hat{Y}_{i,(t,\boldsymbol{s}_1, \boldsymbol{s}_2, \boldsymbol{s}_1', \boldsymbol{s}_2')} = \hat{Y}_{i,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(1-t,t) - \hat{Y}_{i,(\boldsymbol{s}_1', \boldsymbol{s}_2')}(1-t,t).\]
Next, note that Lemma 9 can be applied as is, after proper identification of the notation. Moreover, in view of the inequality \((a+b)^2 \le 2(a^2+b^2)\) and the fact that \(Y_{i,(\boldsymbol{s}_1, \boldsymbol{s}_2)}(x_1, x_2)\) essentially corresponds to \(Y_{I}(\boldsymbol{x}_I)\) from 19 , the bounds in Lemma 7 and Lemma 8 continue to hold, with the hidden universal constant multiplied by \(4\) and 2, respectively. More specifically, the tilde-version of Lemma 8 is as follows: \[\widetilde{\Delta} \lesssim h + \sqrt{r} + \frac{r}{\sqrt{h}} + \frac{r^2}{h} + \frac{1}{h\sqrt{k}} \Big\{ B_{n,k} + \sqrt{r \log\Big(\frac{1}{\delta r}\Big)} \Big\}\] with probability at least \(1-19 |\mathcal{P}_d(1, \sqrt 2)| \delta\). The rest of the proof follows the arguments in the proofs of Proposition 15 and Theorem 11. ◻
The main purpose of this section is to prove Theorem 12. Along the way, we also establish two intermediate results; the following one is useful for proving consistency.
Proposition 16. Suppose that the tuple \((L, \{L(\cdot; \theta): \theta \in \Theta\}, \boldsymbol{g}, \mu)\) satisfies the following: there exists some \(\theta_0 \in \Theta\) such that for every \(\varepsilon>0\), we have that \[f_{Q,L}(\varepsilon) \mathrel{\vcenter{:}}= \inf_{\theta \in \Theta : \| \theta-\theta_0\|_2 \geq \varepsilon} \Big\{ {Q_L(\theta)} - {Q_L(\theta_0)}\Big\}>0,\] where the infimum over an empty set is defined to be infinity. Let \(\eta>0\). Then, for any estimator \(\hat{\theta}_n\) that is a near minimizer of \(\theta \mapsto Q_n(\theta)\) in the sense that \(Q_n(\hat{\theta}_n)-\inf_{\theta \in \Theta} Q_n(\theta) < \eta\), we have \[\big\| \hat{\theta}_n-\theta_0 \big\|_2 \leq f_{Q,L}^{\leftarrow}\Big(\eta+ 2 C_g \sup_{\boldsymbol{x}\in [0,1]^d} \big|\widehat L_n(\boldsymbol{x})-L(\boldsymbol{x})\big|\Big),\] where \(f_{Q,L}^{\leftarrow}\) denotes the generalized inverse of \(f_{Q,L}\) defined in 31 and where \(C_g\) is from 22 .
Note that Proposition 16 is formulated in a general, non-stochastic framework that does not put any assumptions on the observations. Such assumptions will be needed to control the order of \(\sup_{\boldsymbol{x}\in [0,1]^d} \big|\widehat L(\boldsymbol{x})-L(\boldsymbol{x})\big|\) which appears in the upper bound. The proposition also provides a key step in the proof of the following result.
Theorem 17. Suppose that Assumption 1 is met. For \(\eta>0\), let \(\hat{\theta}_n\) be an estimator that satisfies \(Q_n(\hat{\theta}_n)-\inf_{\theta \in \Theta} Q_n(\theta) < \eta.\) For \(\beta >0\), consider the event \[\label{eq:Omega1} \Omega_1(n, \beta) \mathrel{\vcenter{:}}= \Big\{\sup_{\boldsymbol{x}\in [0,1]^d} k^{-\frac{1}{2}} \left|\mathbb{L}_{n}(\boldsymbol{x})\right| \leq \beta \Big\}.\tag{82}\] There exist constants \(\tilde{C}_{r1}, \tilde{C}_{r2}>0\) and \(\tilde{C}_\beta, \tilde{C}_\eta\in (0,1]\) only depending on \(d,s,q\), the constant \(C_g\) from 22 , the three parameters \(\kappa,C_{h},\gamma_{h}\) from Assumption 1 and the four constants defined in 24 and 25 such that, for any \(\beta \in (0, \tilde{C}_\beta)\) and \(\eta\in (0, \tilde{C}_\eta)\), we have, on the event \(\Omega_1(n, \beta)\), \[\label{eq:linMest1} \sqrt{k} \big( \hat{\theta}_n- \theta_0 \big) = 2V_{\theta_0}^{-1} J_{\theta_0}^\top \int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x}) \mathbb{L}_n(\boldsymbol{x}) \, \mathrm d\mu(\boldsymbol{x}) + \sqrt{k} \boldsymbol{r}_{n,1}(\beta, \eta)\tag{83}\] where \(\left\lVert\boldsymbol{r}_{n,1}(\beta, \eta)\right\rVert_2^2 \leq \tilde{C}_{r1}\big( \beta^{2+\gamma_{h}}+\eta\big)\). Moreover, for any measurable set \(A \subseteq[0,1]^d\) such that \(\begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n\) is defined on \([0,1]^d \setminus A\) \[\sqrt{k} \big( \hat{\theta}_n- \theta_0 \big) = 2V_{\theta_0}^{-1} J_{\theta_0}^\top \int_{[0,1]^d\setminus A} \boldsymbol{g}(\boldsymbol{x}) \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x}) \,\mathrm d\mu(\boldsymbol{x}) + \sqrt{k} \boldsymbol{r}_{n,1}(\beta, \eta) + \boldsymbol{r}_{n,2}(A)\] where \[\|\boldsymbol{r}_{n,2}(A)\|_2 \le \tilde{C}_{r2} \Big\{ \sup_{\boldsymbol{x}\in [0,1]^d\setminus A} \big| \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x}) - {\mathbb{L}}_n(\boldsymbol{x}) \big| + \sqrt{k}\beta\int_{A} \|\boldsymbol{g}(\boldsymbol{x})\|_2 \, \mathrm d\mu(\boldsymbol{x}) \Big\}.\]
In the following, we successively prove Proposition 16, Theorem 17 and then Theorem 12.
Proof of Proposition 16. Throughout, we write \(Q=Q_L\). By definition of the generalized inverse, it suffices to prove that \[\label{eq:bound-f} f_{Q,L}\big(\big\| \hat{\theta}_n-\theta_0 \big\|_2\big) < \eta+2C_g \sup_{\boldsymbol{x}\in [0,1]^d} \big| \widehat L_{n}(\boldsymbol{x})-L(\boldsymbol{x}) \big|.\tag{84}\] Note that, by the definition of \(\hat{\theta}_n\) and \(\eta\), \[\begin{align} \eta& > Q_n(\hat{\theta}_n) - Q_n(\theta_0) \\ &= \big(Q(\hat{\theta}_n) - Q(\theta_0)\big) - \big(Q(\hat{\theta}_n) - Q_n(\hat{\theta}_n)\big) - \big(Q_n(\theta_0) - Q(\theta_0)\big) \\ &\geq \big(Q(\hat{\theta}_n) - Q(\theta_0)\big) - \big|Q(\hat{\theta}_n) - Q_n(\hat{\theta}_n)\big| - \big|Q_n(\theta_0) - Q(\theta_0)\big|. \end{align}\] Thus \[f_{Q,L}\big(\big\| \hat{\theta}_n-\theta_0 \big\|_2\big) \le Q(\hat{\theta}_n) - Q(\theta_0) < 2 \sup_{\theta \in \Theta} \big| Q_n(\theta) - Q(\theta)\big|+ \eta.\] For each \(\theta \in \Theta\), the reverse triangle inequality implies that \[\begin{align} &\big| Q_n(\theta)-Q(\theta) \big|_2 \\ & = \bigg| \Big\| \int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x})\big( L(\boldsymbol{x};\theta)-\widehat L_n(\boldsymbol{x}) \big) \mathrm d\mu(\boldsymbol{x}) \Big\|_2 - \Big\| \int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x})\big( L(\boldsymbol{x};\theta)- L(\boldsymbol{x}) \big) \mathrm d\mu(\boldsymbol{x}) \Big\|_2 \bigg| \\& \leq \Big\| \int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x})\big( L(\boldsymbol{x};\theta)-\widehat L_n(\boldsymbol{x}) \big) \mathrm d\mu(\boldsymbol{x}) - \int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x})\big( L(\boldsymbol{x};\theta)- L(\boldsymbol{x}) \big) \mathrm d\mu(\boldsymbol{x}) \Big\|_2 \\ &= \Big\| \int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x})\big( L(\boldsymbol{x}) - \widehat L_n(\boldsymbol{x})\big) \mathrm d\mu(\boldsymbol{x}) \Big\|_2. \end{align}\] By the triangle inequality for integrals, \[\begin{align} \label{eq:qn-q-hoelder} \Big\| \int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x})\big( L(\boldsymbol{x}) - \widehat L_n(\boldsymbol{x})\big) \mathrm d\mu(\boldsymbol{x})\Big\|_2 &\leq \sup_{\boldsymbol{x}\in [0,1]^d} \big| \widehat L_n(\boldsymbol{x})-L(\boldsymbol{x}) \big| \times \int_{[0,1]^d} \| \boldsymbol{g}(\boldsymbol{x}) \|_2\, \mathrm d\mu(\boldsymbol{x}). \end{align}\tag{85}\] Combining the last three displayed formulas establishes 84 and completes the proof. ◻
Proof of Theorem 17. Throughout, we write \(Q=Q_L\) and utilize the following additional notation \[\boldsymbol{\psi}:= \int_{[0,1]^d } \boldsymbol{g}(\boldsymbol{x}) L(\boldsymbol{x}) \mathrm d\mu(\boldsymbol{x}), \quad \widehat \boldsymbol{\psi}:= \int_{[0,1]^d } \boldsymbol{g}(\boldsymbol{x})\widehat L_{n}(\boldsymbol{x}) \mathrm d\mu(\boldsymbol{x}).\] For a matrix \(A\), let \({\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert A \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2\) denote the spectral norm of \(A\), that is, \({\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert A \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2\) is largest singular value of \(A\). Further, \({\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert A \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_1\) is the maximum of the absolute column sums of \(A\), while \({\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert A \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\infty\) is the maximum of the absolute row sums of \(A\); note that \({\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert A \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2^2 \leq {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert A \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_1 \cdot {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert A \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_\infty\). For either a vector or a matrix, \(\left\lVert\cdot\right\rVert_\infty\) refers to the absolute maximum entry; note that the previous inequality then yields \({\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert A \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2 \leq \sqrt{sq} \|A\|_\infty\) for \(A \in \mathbb{R}^{s \times q}\). Further, \(\|A\boldsymbol{b}\|_2 \le {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert A \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2 \|\boldsymbol{b}\|_2\) for \(A\in\mathbb{R}^{s \times q}\) and \(\boldsymbol{b} \in\mathbb{R}^{q}\). Finally, if \(A\) is a square matrix and \(\boldsymbol{b}\) a vector, we have \(|\boldsymbol{b}^\top A \boldsymbol{b}| \le {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert A \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2 \| \boldsymbol{b}\|_2^2\).
In what follows, we will without loss of generality assume that \(\kappa\leq 1\). Moreover, we will choose \(\tilde{C}_\beta\) and \(\tilde{C}_\eta\) not larger than \(1\), which implies \(\beta, \eta\le 1\).
Let \(\theta \in B_{\kappa}(\theta_0)\) and define \(\Delta_{\theta} \mathrel{\vcenter{:}}= \theta-\theta_0\). Under Assumption 1, we have the Taylor expansion \[\begin{align} \label{eq:expansion-qn2} Q_n^2(\theta)-Q_n^2(\theta_0) \nonumber &= \left[\nabla Q_n^2(\theta_0)\right]^{\top}\Delta_{\theta}+\frac{1}{2}\Delta_{\theta}^{\top} V_{n, \tilde{\theta}} \Delta_{\theta} \\ &= \frac{1}{2}\Delta_{\theta}^{\top} V_{\theta_0} \Delta_{\theta} + r_{n,1}(\theta) +r_{n,2}(\theta) +r_{n,3}(\theta), \end{align}\tag{86}\] where \(\tilde{\theta}\) is a convex combination of \(\theta\) and \(\theta_0\) and where \[\begin{align} r_{n,1}(\theta) &:= \left[\nabla Q_n^2(\theta_0)\right]^{\top}\Delta_{\theta}, \\ r_{n,2}(\theta) &:= \frac{1}{2}\Delta_{\theta}^\top(V_{n,\theta_0} - V_{\theta_0})\Delta_{\theta}, \\ r_{n,3}(\theta) &:= \frac{1}{2}\Delta_{\theta}^{\top} (V_{n, \tilde{\theta}}-V_{n,\theta_0}) \Delta_{\theta} . \end{align}\] We will show below that, on the event \(\Omega_1(n,\beta)\), \[\begin{align} \label{eq:bound-rn} r_{n,1}(\theta) &\le C_1 \beta \left\lVert\Delta_{\theta}\right\rVert_2, \qquad r_{n,2}(\theta) \le C_2 \beta \left\lVert\Delta_{\theta}\right\rVert_2^2, \qquad r_{n,3}(\theta) \le C_3(\beta) \left\lVert\Delta_{\theta}\right\rVert_2^{2+\gamma_{h}}, \end{align}\tag{87}\] where \(C_1=2 \sqrt{sq}C_\partial C_g\), \(C_2=s\sqrt q C_g C_{\partial^2}\), and \(C_3(\beta) := 3sq^{3/2}C_\partial C_{\partial^2} + sqC_{h}(C_g \beta + d_{\theta_0})\) with \(d_{\theta_0} := \max_{p \in [q]} |\varphi_p(\theta_0) - \psi_p|\); note that \(C_3(\beta)\) is increasing in \(\beta\) and hence bounded by \(C_3(1)\). Note that by using \(0\le L(\cdot),L(\cdot;\theta)\le d\), we have \(d_{\theta_0} \le dC_g\) which is an upper bound that does not depend on \(L\).
Regarding \(r_{n,1}(\theta)\), recall that \(\theta_0\) is the global minimizer of \(\theta \mapsto Q^2(\theta) = \|\boldsymbol{\varphi}(\theta)-\boldsymbol{\psi} \|_2^2\) and so \[0 = \nabla Q^2(\theta_0) = 2 \sum_{p \in [q]} \big(\varphi_p(\theta_0)-\psi_p\big)\nabla \varphi_p(\theta_0).\] Thus \[\begin{align} \label{eq:nableqn2} \nabla Q_n^2(\theta_0) &= \nonumber 2\sum_{p \in [q]} (\varphi_p(\theta_0) - \widehat \psi_p)\nabla \varphi_p(\theta_0) \\ &= 2\sum_{p \in [q]} \big( \psi_p - \widehat \psi_p \big)\nabla \varphi_p(\theta_0) = -\frac{2}{\sqrt k} J_{\theta_0}^\top \int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x}) \mathbb{L}_n(\boldsymbol{x}) \mathrm d\mu(\boldsymbol{x}). \end{align}\tag{88}\] As a consequence, on the event \(\Omega_1(n,\beta)\), recalling the definition of \(C_g\) and \(C_\partial\) in 22 and 24 , respectively, we have the bound \[\begin{align} \big| r_{n,1}(\theta) \big| &= \frac{2}{\sqrt k} \Big| \Big( J_{\theta_0}^\top \int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x}) \mathbb{L}_n(\boldsymbol{x}) \mathrm d\mu(\boldsymbol{x})\Big)^\top \Delta_{\theta} \Big| \\&\leq \frac{2}{\sqrt k} \Big\| J_{\theta_0}^\top \int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x}) \mathbb{L}_n(\boldsymbol{x}) \mathrm d\mu(\boldsymbol{x})\Big\|_2 \left\lVert\Delta_{\theta}\right\rVert_2 \\ &\leq \frac{2}{\sqrt k}{\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert J_{\theta_0}^\top \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2 \times \Big\|\int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x}) \mathbb{L}_n(\boldsymbol{x}) \mathrm d\mu(\boldsymbol{x}) \Big\|_2 \left\lVert\Delta_\theta\right\rVert_2 \\ &\leq \frac{2}{\sqrt k} \sqrt {sq}C_\partial C_g \Big(\sup_{\boldsymbol{x}\in [0,1]^d} \left|\mathbb{L}_n(\boldsymbol{x})\right|\Big) \left\lVert\Delta_\theta\right\rVert_2 \\ & \le 2 \sqrt{sq}C_\partial C_g \beta \left\lVert\Delta_{\theta}\right\rVert_2, \end{align}\] as claimed in 87 .
Next, regarding \(r_{n,2}(\theta)\), note that the \((j, \ell)\)-entry of \(V_{n, \theta_0}-V_{\theta_0}\in \mathbb{R}^{s \times s}\) is given by \[\begin{align} [V_{n,\theta_0}-V_{\theta_0}]_{j\ell} &= 2 \sum_{p \in [q]} \Big((\varphi_p(\theta_0) - \widehat \psi_p)\partial_{j\ell} \varphi_p(\theta_0) + \partial_j \varphi_p(\theta_0) \partial_\ell \varphi_p(\theta_0)\Big) \\ &\quad \quad - 2 \sum_{p \in [q]} \Big((\varphi_p(\theta_0) -\psi_p)\partial_{j\ell} \varphi_p(\theta_0) + \partial_j \varphi_p(\theta_0) \partial_\ell \varphi_p(\theta_0)\Big) \\ &= -2 (\widehat \boldsymbol{\psi}- \boldsymbol{\psi})^\top \partial_{j\ell} \boldsymbol{\varphi}(\theta_0) = - \frac{2}{\sqrt k} \Big[\int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x}) \mathbb{L}_n(\boldsymbol{x}) \mathrm d\mu(\boldsymbol{x})\Big]^\top \partial_{j\ell} \boldsymbol{\varphi}(\theta_0). \end{align}\] Hence, on the event \(\Omega_{1}(n, \beta)\) \[\left\lVert V_{n, \theta_0}-V_{\theta_0}\right\rVert_\infty \leq 2 \sqrt q C_g C_{\partial^2} \beta,\] which in turn implies \[\begin{align} \label{eq:32V95n32-32V95theta32bound} \big|r_{n,2}(\theta) \big| = \left|\frac{1}{2} \Delta_{\theta}^\top (V_{n,\theta_0}-V_{\theta_0}) \Delta_{\theta}\right| &\leq \frac{1}{2} {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{n, \theta_0}-V_{\theta_0} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2\left\lVert\Delta_{\theta}\right\rVert^2_2 \leq s\sqrt q C_g C_{\partial^2} \beta \left\lVert\Delta_{\theta}\right\rVert_2^2 \end{align}\tag{89}\] as claimed in 87 .
Finally, regarding \(r_{n,3}(\theta)\), a similar calculation shows that the \((j, \ell)\)-entry of \(V_{n, \tilde{\theta}}- V_{n, \theta_0}\) can be written as \[\begin{align} [V_{n, \tilde{\theta}}- V_{n, \theta_0}]_{j\ell} &= 2 \sum_{p \in [q]} \Big(\partial_j \varphi_p(\tilde{\theta}) \partial_\ell \varphi_p(\tilde{\theta}) - \partial_j \varphi_p(\theta_0) \partial_\ell \varphi_p(\theta_0)\Big) \\ &\quad \quad + 2 \sum_{p \in [q]} \Big((\varphi_p(\tilde{\theta}) - \widehat \psi_p)\partial_{j\ell} \varphi_p(\tilde{\theta}) - (\varphi_p(\theta_0) -\widehat \psi_p)\partial_{j\ell} \varphi_p(\theta_0) \Big). \end{align}\] First, since \(|ab-cd| \le |a||b-d|+|d||a-c|\) and \(\tilde{\theta} \in B_{\kappa}(\theta_0)\), \[\begin{align} & \left|\partial_j \varphi_p(\tilde{\theta}) \partial_\ell \varphi_p(\tilde{\theta}) - \partial_j \varphi_p(\theta_0) \partial_\ell \varphi_p(\theta_0)\right| \\ &\leq \left|\partial_j \varphi_p(\tilde{\theta})\right| \left|\partial_\ell \varphi_p(\tilde{\theta})-\partial_\ell \varphi_p(\theta_0)\right| + \left|\partial_\ell \varphi_p(\theta_0)\right| \left|\partial_j \varphi_p(\tilde{\theta})-\partial_j \varphi_p(\theta_0)\right| \\ & \leq C_{\partial} \Big(\left|\partial_\ell \varphi_p(\tilde{\theta})-\partial_\ell \varphi_p(\theta_0)\right| + \left|\partial_j \varphi_p(\tilde{\theta})-\partial_j \varphi_p(\theta_0)\right|\Big) \\ & \leq 2 \sqrt{q}C_{\partial} C_{\partial^2}\|\tilde{\theta}-\theta_0\|_2, \end{align}\] where we have used that, by the mean value inequality and the fact that the partial derivatives of \(\theta \mapsto \partial_j \varphi_m(\theta)\) are bounded by \(C_{\partial^2}\) on \(B_{\kappa}(\theta_0)\), \[\begin{align} \label{eq:bound-pds-phi} \left|\partial_j \varphi_p(\tilde{\theta})-\partial_{j} \varphi_p(\theta_0)\right| &\le \nonumber \sup_{t \in (0,1)} \Big| \frac{d}{dt} \partial_j \varphi_p(\theta_0 + t(\tilde{\theta}-\theta_0)) \Big| \\&\le \sup_{t \in (0,1)} \big\|\nabla[\partial_j \varphi_p](\theta_0 + t(\tilde{\theta}-\theta_0))\big\|_2 \big\|\tilde{\theta} - \theta_0\big\|_2 \leq \sqrt{q}C_{\partial^2} \big\|\tilde{\theta} - \theta_0\big\|_2. \end{align}\tag{90}\] Second, recalling \(d_{\theta_0} := \max_{p \in [q]} |\varphi_p(\theta_0) - \psi_p|\), \[\begin{align} &\Big| (\varphi_p(\tilde{\theta}) - \widehat \psi_p)\partial_{j\ell} \varphi_p(\tilde{\theta}) - (\varphi_p(\theta_0) -\widehat \psi_p)\partial_{j\ell} \varphi_p(\theta_0) \Big| \\ &\leq \big| \partial_{j\ell}\varphi_p(\tilde{\theta}) \big| \left|\varphi_p(\tilde{\theta})-\varphi_p(\theta_0)\right| + \big| \varphi_p(\theta_0)-\widehat \psi_p \big| \left|\partial_{j\ell} \varphi_p(\tilde{\theta}) - \partial_{j\ell} \varphi_p(\theta_0)\right| \\&\leq \sqrt{q} C_\partial C_{\partial^2} \big\| \tilde{\theta}-\theta_0 \big\|_2 + C_{h}\big\| \tilde{\theta}-\theta_0 \big\|_2^{\gamma_{h}}\Big(\big| \varphi_p(\theta_0)- \psi_p \big| + \big| \widehat \psi_p - \psi_p \big| \Big) \\ &\leq [ \sqrt{q} C_\partial C_{\partial^2} + C_{h}(C_g \beta + d_{\theta_0})] \times \left\lVert\Delta_{\theta}\right\rVert_2^{\gamma_{h}}, \end{align}\] where we used that \(\|\tilde{\theta}- \theta_0\|_2 \le \|\theta- \theta_0\|_2= \left\lVert\Delta_{\theta}\right\rVert_2 \le \kappa\le 1\), and that \[\hat{\psi}_p - \psi_p = \frac{1}{\sqrt k} \int_{[0,1]^d}g_p(\boldsymbol{x}) \mathbb{L}_n(\boldsymbol{x}) \mathrm d\mu(\boldsymbol{x})\] is bounded by \(C_g \beta\) on the event \(\Omega_1(n, \beta)\), and that \(|\varphi_p(\tilde{\theta}) - \varphi_p(\theta_0)| \le \sqrt{q} C_{\partial} \big\|\tilde{\theta} - \theta_0\big\|_2\), which follows from the same arguments that were used in 90 . Combining the bounds so far we obtain \[\label{eq:32bound32for32V95n32-32V95n32theta} {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{n, \tilde{\theta}}-V_{n,\theta_0} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2\leq s \left\lVert V_{n, \tilde{\theta}}-V_{n,\theta_0}\right\rVert_\infty \leq 2C_3(\beta) \left\lVert\Delta_{\theta}\right\rVert_2^{\gamma_{h}},\tag{91}\] where \(C_3(\beta) = 3sq^{3/2}C_\partial C_{\partial^2} + sqC_{h}(C_g \beta + d_{\theta_0})\), which in turn implies \[\begin{align} \big|r_{n,3}(\theta) \big| = \left|\frac{1}{2} \Delta_{\theta}^\top (V_{n,\tilde{\theta}}-V_{\theta_0}) \Delta_{\theta}\right| &\leq \frac{1}{2} {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{n, \tilde{\theta}}-V_{\theta_0} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2\left\lVert\Delta_{\theta}\right\rVert^2_2 \leq C_3(\beta) \left\lVert\Delta_{\theta}\right\rVert_2^{2+\gamma_{h}} \end{align}\] as claimed in 87 .
Next, we will show that \[\label{eq:bounddiffQn942} \forall \theta \in \Theta: \qquad Q_n^2(\hat{\theta}_n) - Q_n^2(\theta) < 2dC_g \eta=:C_4 \eta.\tag{92}\] For that purpose, note that our assumption on \(\hat{\theta}_n\) yields \(Q_n(\hat{\theta}_n)-Q_n(\theta) < \eta\) for any \(\theta\in \Theta\). Moreover, by a similar calculation as in 85 , we have for any \(\theta \in \Theta\) (in particular, for \(\theta=\hat{\theta}_n\)) \[0 \le Q_n(\theta) \le C_g \sup_{\boldsymbol{x} \in [0,1]^d}\Big| \widehat L_n(\boldsymbol{x}) - L(\boldsymbol{x};\theta) \Big| \le dC_g,\] where we used that \(\widehat L(\boldsymbol{x}),L(\boldsymbol{x};\theta) \le \| \boldsymbol{x}\|_1\). As a consequence \[Q_n^2(\hat{\theta}_n) - Q_n^2(\theta) = \big(Q_n(\hat{\theta}_n)-Q_n(\theta)\big)\big(Q_n(\hat{\theta}_n)+Q_n(\theta)\big) < 2dC_g \eta\] as asserted in 92 .
We will next apply Proposition 16, and for that purpose, we need to check that \(f_{Q,L}(\varepsilon)>0\) for all \(\varepsilon>0\). In fact, for later purposes, we will need a precise lower bound on that function, and more specifically, we will now show that there exists a constant \(C_f>0\) depending on \(d,C_g,C_V\) and \(C_Q\) only such that \[\begin{align} \label{eq:fql-quadratic-bound} f_{Q,L}(\varepsilon) \equiv \inf_{\theta: \|\theta-\theta_0\|_2 \ge \varepsilon} \big\{ Q(\theta) - Q(\theta_0)\big\} \ge C_f (\varepsilon^2 \wedge 1). \end{align}\tag{93}\] Indeed, note that \(Q(\theta) \le d \int \| \boldsymbol{g}\|_2 \, \mathrm d\mu \le dC_g\) for all \(\theta \in \Theta\), whence \[Q^2(\theta) - Q^2(\theta_0) = (Q(\theta) - Q(\theta_0))(Q(\theta) + Q(\theta_0)) \le 2d C_g (Q(\theta) - Q(\theta_0)).\] As a consequence, with \(\tilde{C}_f = (2d C_g)^{-1}\) and with \(d_Q\) as defined in Assumption 1(v). \[Q(\theta) - Q(\theta_0) \ge \tilde{C}_f \big\{ Q^2(\theta) - Q^2(\theta_0) \big\} = \tilde{C}_f d_Q(\theta).\] Hence, \[\begin{align} f_{Q,L}(\varepsilon) &\ge \tilde{C}_f \inf_{\theta: \|\theta-\theta_0\|_2 \ge \varepsilon} d_Q(\theta) \\&= \tilde{C}_f \min\Big\{ \inf_{\theta: \kappa\ge \|\theta-\theta_0\|_2 \ge \varepsilon} d_Q(\theta), \inf_{\theta: \|\theta-\theta_0\|_2 > \kappa} d_Q(\theta) \Big\} \ge \tilde{C}_f \min\Big( \frac{C_V}{4} \varepsilon^2, C_Q\Big), \end{align}\] where we have used Assumption 1(v). This implies 93 with \(C_f = \tilde{C}_f\min(C_V/4, C_Q)\).
Next, note that \(f_{Q,L}^{\leftarrow}(u) \le \sqrt{u/C_f}\) for \(0 < u \le C_f\) by 93 . Choosing \(\tilde{C}_\eta\le (\kappa^2\wedge 1)C_f/2\) and \(\tilde{C}_\beta \le (\kappa^2\wedge 1)C_f/(4 C_g)\) ensures \(\eta + 2C_g \beta \le (\kappa^2 \wedge 1) C_f \le C_f\) for \(\beta \in (0,\tilde{C}_\beta)\) and \(\eta\in (0,\tilde{C}_\eta)\), and so \[\big\| \hat{\theta}_n-\theta_0\big\|_2 \leq f_{Q,L}^{\leftarrow}\Big(\eta+2 C_g \beta \Big) \leq \Big(\frac{\eta+2 C_g \beta}{C_f} \Big)^{1/2} \le (\kappa^2 \wedge 1)^{1/2} \le \kappa.\] by Proposition 16.
As a consequence, we can apply 86 and 87 with \(\hat{\Delta}_n=\Delta_{\hat{\theta}_n}=\hat{\theta}_n-\theta_0\) to obtain that \[\label{eq:qn2-mest} Q_n^2(\hat{\theta}_n) - Q_n^2(\theta_0) = \frac{1}{2} \hat{\Delta}_n^{\top} V_{\theta_0} \hat{\Delta}_n + r_{n,1}(\hat{\theta}_n) + r_{n,2}(\hat{\theta}_n) + r_{n,3}(\hat{\theta}_n),\tag{94}\] with the three error terms satisfying \[\begin{align} \label{eq:error-mest} |r_{n,1}(\hat{\theta}_n)| \le C_1 \beta \big\| \hat{\Delta}_n \big\|_2, \qquad |r_{n,2}(\hat{\theta}_n)| + |r_{n,3}(\hat{\theta}_n) | \leq C_5(\beta, \eta) \big\| \hat{\Delta}_n \big\|_2^2, \end{align}\tag{95}\] with \(C_5(\beta, \eta) := C_2\beta + C_3(\beta) \{ (\eta+2 C_g \beta)/{C_f}\}^{\gamma_{h}/2}\). Combining 92 (with \(\theta=\theta_0\)) with 94 and 95 , we obtain that \[\begin{align} C_{4} \eta & \ge \frac{1}{2} \hat{\Delta}_n^{\top} V_{\theta_0} \hat{\Delta}_n + r_{n,1}(\hat{\theta}_n) + r_{n,2}(\hat{\theta}_n) + r_{n,3}(\hat{\theta}_n) \\ &> \frac{1}{2} \lambda_{\text{min}}(V_{\theta_0}) \big\| \hat{\Delta}_n \big\|_2^2 - C_1 \beta \big\| \hat{\Delta}_n \big\|_2 - C_5(\beta, \eta) \big\| \hat{\Delta}_n \big\|_2^2. \end{align}\] Decreasing \(\tilde{C}_\beta\) and \(\tilde{C}_\eta\) if necessary, we can guarantee that \(C_{5}(\beta, \eta) \le \lambda_{\min}(V_{\theta_0})/4\) for any \(\beta \in (0,\tilde{C}_\beta)\) and \(\eta\in (0,\tilde{C}_\eta)\). Hence, \[\big\| \hat{\Delta}_n \big\|_2^2 < \frac{4}{\lambda_{\min}(V_{\theta_0})} \big( C_4 \eta+ C_1 \beta \big\| \hat{\Delta}_n \big\|_2 \big).\] For \(a,b>0\) and \(x \ge 0\), we have that \(x^2 \le ax+b\) implies \(x \le a+\sqrt b\); indeed, if \(x>a+\sqrt b\), we have \(x^2 > x(a+\sqrt b) >ax+(a +\sqrt b)\sqrt b > ax+b\). Thus, \[\label{eq:boundhatdelta1} \big\| \hat{\Delta}_n \big\|_2 \le \frac{2\sqrt{C_4 \eta}}{\sqrt{\lambda_{\min}(V_{\theta_0})}} + \frac{4C_1 \beta}{\lambda_{\min}(V_{\theta_0})}.\tag{96}\] As a consequence, \(\| \hat{\Delta}_n \|_2^2 \le C_6 \big( \eta+ \beta^2\big)\) with \(C_6 = \{8C_4/\lambda_{\min}(V_{\theta_0}) \} \vee \{32C_1^2/\lambda_{\min}^2(V_{\theta_0})\}\), which, using 87 with \(\theta=\hat{\theta}_n\), yields \[\begin{align} \label{eq:bound-errors-hat} |r_{n,2}(\hat{\theta}_n)| + |r_{n,3}(\hat{\theta}_n) | &\le \nonumber C_2\beta \big\| \hat{\Delta}_{n} \big\|_2^{2} + C_3(\beta)\big\| \hat{\Delta}_{n} \big\|_2^{2+\gamma_{h}} \\&\le\nonumber \Big( C_2C_6 \frac{\beta }{(\eta+ \beta^2)^{\gamma_{h}/2}}+C_3(\beta)C_6^{1+\gamma_{h}/2} \Big) (\eta+ \beta^2)^{1+\gamma_{h}/2} \\&\le C_{7}(\beta) (\eta+ \beta^2)^{1+\gamma_{h}/2}, \end{align}\tag{97}\] where \(C_7(\beta)= C_2C_6\beta^{1-\gamma_{h}} + C_3(\beta) C_6^{1+\gamma_{h}/2}\).
Next, let \[\widetilde{\Delta}_n = 2k^{-\frac{1}{2}}V_{\theta_0}^{-1} J_{\theta_0}^{\top} \int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x}) \mathbb{L}_n(\boldsymbol{x}) \mathrm d\mu(\boldsymbol{x}) = - V_{\theta_0}^{-1} \nabla Q_n^2(\theta_0)\] where the second equality follows from 88 . Note that we need to find \(\tilde{C}_r>0\) such that \(\| \hat{\Delta}_n - \widetilde{\Delta}_n \|_2^2 \le \tilde{C}_r(\eta+ \beta^{2+\gamma_{h}})\). On \(\Omega_1(n, \beta)\), we have \[\begin{align} \label{eq:boundhatdelta2} \big\| \widetilde{\Delta}_n \big\|_2 \le 2C_g {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{\theta_0}^{-1} J_{\theta_0}^{\top} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2 \beta \le 2C_g {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{\theta_0}^{-1} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2 {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert J_{\theta_0}^{\top} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2 \beta \le 2C_g C_V^{-1}\sqrt{sq} C_{\partial} \beta =: C_{8} \beta, \end{align}\tag{98}\] where we have used 85 . Further decreasing \(\tilde{C}_\beta\) if necessary, the right hand-side is bounded by \(\kappa\) for all \(\beta \in (0, \tilde{C}_\beta)\), which implies that \(\widetilde{\theta}_n := \theta_0 + \widetilde{\Delta}_n \in B_\kappa(\theta_0)\). We can hence apply the expansions and bounds derived at the beginning of this proof, specifically 86 , with \(\theta= \widetilde{\theta}_n\) and \(\Delta_{\widetilde{\theta}_n}= \widetilde{\Delta}_n\) to deduce that \[\label{eq:qn2-mest-tilde} Q_n^2(\widetilde{\theta}_n)-Q_n^2(\theta_0) = \frac{1}{2} \widetilde{\Delta}_{n}^\top V_{\theta_0} \widetilde{\Delta}_{n}+ r_{n,1}(\widetilde{\theta}_n) + r_{n,2}(\widetilde{\theta}_n) + r_{n,3}(\widetilde{\theta}_n),\tag{99}\] where, using 87 and 98 , \[\begin{align} \label{eq:bound-errors-tilde} |r_{n,2}(\widetilde{\theta}_n)| + |r_{n,3}(\widetilde{\theta}_n) | \le C_2\beta \big\| \widetilde{\Delta}_{n} \big\|_2^{2} + C_3(\beta)\big\| \widetilde{\Delta}_{n} \big\|_2^{2+\gamma_{h}} \le C_{9}(\beta) \beta^{2+\gamma_{h}}, \end{align}\tag{100}\] where \(C_{9}(\beta) = C_{8}^2\{ C_2\beta^{1-\gamma_{h}} + C_8^{\gamma_{h}}C_3(\beta)\}\). Overall, from 92 applied with \(\theta= \widetilde{\theta}_n\) and 94 and 99 , we find that \[\begin{align} C_{4} \eta &> Q_n^2(\hat{\theta}_n)-Q_n^2(\widetilde{\theta}_n) = \big( Q_n^2(\hat{\theta}_n)-Q_n^2(\theta_0)\big)-\big(Q_n^2(\widetilde{\theta}_n)-Q_n^2(\theta_0)\big) = M_n + \tilde{r}_n \end{align}\] where \[\begin{align} M_n &= \frac{1}{2} \hat{\Delta}_n^{\top} V_{\theta_0} \hat{\Delta}_n - \frac{1}{2} \widetilde{\Delta}_n^{\top} V_{\theta_0} \widetilde{\Delta}_n + \big[ \nabla Q_n^2(\theta_0)\big]^\top \big(\hat{\Delta}_n- \widetilde{\Delta}_n\big), \\ \tilde{r}_n &= r_{n,2}(\hat{\theta}_n)-r_{n,2}(\widetilde{\theta}_n) + r_{n,3}(\hat{\theta}_n)-r_{n,3}(\widetilde{\theta}_n). \end{align}\] In view of 97 and 100 , the remainder term satisfies \[|\tilde{r}_n| \le C_{7}(\beta) (\eta+ \beta^2)^{1+\gamma_{h}/2} + C_{9}(\beta) \beta^{2+\gamma_{h}} \le C_{10} (\eta+\beta^2)^{1+\gamma_{h}/2}\] with \(C_{10}=C_7(\tilde{C}_\beta) + C_{9}(\tilde{C}_\beta)\). Moreover, since \(\nabla Q_n^2(\theta_0) = - V_{\theta_0}\widetilde{\Delta}_n\), we find that \[\begin{align} M_n = \frac{1}{2} \hat{\Delta}_n^{\top} V_{\theta_0} \hat{\Delta}_n + \frac{1}{2} \widetilde{\Delta}_n^{\top} V_{\theta_0} \widetilde{\Delta}_n - \widetilde{\Delta}_n^{\top} V_{\theta_0} \hat{\Delta}_n &= \frac{1}{2} \Big\| V_{\theta_0}^{1/2}(\hat{\Delta}_n-\widetilde{\Delta}_n) \Big\|_2^2 \\&\ge \frac{1}{2} \lambda_{\min}(V_{\theta_0}) \big\|\hat{\Delta}_n-\widetilde{\Delta}_n\big\|_2^2. \end{align}\] Overall, \[C_4 \eta> \frac{1}{2} \lambda_{\min}(V_{\theta_0}) \big\|\hat{\Delta}_n-\widetilde{\Delta}_n\big\|_2^2 - C_{10}(\eta+\beta^2)^{1+\gamma_{h}/2}.\] Convexity of \(x \mapsto x^{1+\gamma_{h}/2}\) and the fact that \(\eta\le 1\) yields \[\big\|\hat{\Delta}_n-\widetilde{\Delta}_n\big\|_2^2 \le \frac{2}{\lambda_{\min}(V_{\theta_0})} \Big[ (C_4 + 2^{\gamma_{h}/2}C_{10}) \eta+ 2^{\gamma_{h}/2}C_{10} \beta^{2+\gamma_{h}} \Big].\] This proves 83 with \(\tilde{C}_{r1} = 2 (C_4 + 2^{\gamma_{h}/2}C_{10})/ {\lambda_{\min}(V_{\theta_0})}\).
To prove the second half of the theorem, note that \[\begin{align} & \Big\|\int_{[0,1]^d} V_{\theta_0}^{-1} J_{\theta_0}^\top\boldsymbol{g}(\boldsymbol{x}) \mathbb{L}_n(\boldsymbol{x}) \, \mathrm d\mu(\boldsymbol{x}) - \int_{[0,1]^d\setminus A} V_{\theta_0}^{-1} J_{\theta_0}^\top\boldsymbol{g}(\boldsymbol{x}) \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x}) \, \mathrm d\mu(\boldsymbol{x})\Big\|_2 \\ \le~& \int_{[0,1]^d\setminus A} \|V_{\theta_0}^{-1} J_{\theta_0}^\top\boldsymbol{g}(\boldsymbol{x})\|_2 \cdot \big| \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x}) - {\mathbb{L}}_n(\boldsymbol{x}) \big| \, \mathrm d\mu(\boldsymbol{x}) + \int_{A} \|V_{\theta_0}^{-1} J_{\theta_0}^\top\boldsymbol{g}(\boldsymbol{x})\|_2 \cdot |{\mathbb{L}}_n(\boldsymbol{x})| \, \mathrm d\mu(\boldsymbol{x}) \\ \le~& \Big(\sup_{\boldsymbol{x}\in [0,1]^d\setminus A} \big| \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x}) - {\mathbb{L}}_n(\boldsymbol{x}) \big|\Big) \times \int_{[0,1]^d} \|V_{\theta_0}^{-1} J_{\theta_0}^\top\boldsymbol{g}(\boldsymbol{x})\|_2 \, \mathrm d\mu(\boldsymbol{x}) \\ &+ \sqrt{k} \beta \int_{A} \|V_{\theta_0}^{-1} J_{\theta_0}^\top\boldsymbol{g}(\boldsymbol{x})\|_2 \, \mathrm d\mu(\boldsymbol{x}). \end{align}\] The bound \[\begin{align} \label{eq:bound-vjg} \|V_{\theta_0}^{-1} J_{\theta_0}^\top\boldsymbol{g}(\boldsymbol{x})\|_2 \le {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{\theta_0}^{-1}J_{\theta_0}^\top \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2 \| \boldsymbol{g}(\boldsymbol{x}) \|_2 &\le {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{\theta_0}^{-1} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2 {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert J_{\theta_0}^\top \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2\| \boldsymbol{g}(\boldsymbol{x}) \|_2 \le C_V^{-1} \cdot \sqrt{sq} C_{\partial} \cdot \| \boldsymbol{g} (x)\|_2 \end{align}\tag{101}\] completes the proof of Theorem 17, with \(\tilde{C}_{r2} = C_V^{-1} \sqrt{sq} C_{\partial}(C_g \vee 1)\). ◻
Proof of Theorem 12. First, all assumptions of Theorem 3 are satisfied, and an application of that theorem implies that there exist constants \(D_1 = D_1(d,K_L)\) and \(D_2 = D_2(d,K_L)\) and an event \(\Omega_2\) that has probability at least \(1-(6d+5)\delta\) on which \[\sup_{\boldsymbol{x} \in [0,1]^d \setminus (\mathfrak B^{\oplus C_{s}r})} \big| \mathbb{L}_n(\boldsymbol{x}) - \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x}) \big| \le \zeta_{n,2} := B_{n,k}(L; [0,1+C_{s}r]^d) + \frac{d}{\sqrt{k}} + D_{1} \sqrt{r\log\Big(\frac{D_{2}}{\delta r}\Big)}.\]
On the same event, by 35 , \[\begin{align} \max_{j\in[d]} \sup_{x_j \in [0,1]} |S_{nj}(x_j) - x_j| \le C_{s}r, \end{align}\] and in view of the decomposition \[\begin{align} \mathbb{L}_n = \widetilde{\mathbb{L}}_n \circ S_n + \sqrt k (L \circ S_n - L) + B_n \circ S_n \end{align}\] from 34 , we obtain that \[\begin{align} \sup_{\boldsymbol{x}\in [0,1]^d}|\mathbb{L}_n(\boldsymbol{x})| &\le \sup_{\boldsymbol{x}\in [0,1+C_{s}r]^d}|\widetilde{\mathbb{L}}_n(\boldsymbol{x})| + C_{s}d r \sqrt{k} + B_{n,k}(L;[0,1+C_{s}r]^d) \\&\le \sup_{\boldsymbol{x}\in [0,2]^d}|\widetilde{\mathbb{L}}_n(\boldsymbol{x})| + C_{s}d r \sqrt{k} + B_{n,k}(L;[0,2]^d) \end{align}\] by Lipschitz continuity of \(L\) and using that \(C_{s}r \le 1\) by assumption.
The current choice of \(\delta\) also satisfies the conditions of Lemma 10 with \(T=2\). Hence there exists an event \(\Omega_3\) with probability at least \(1-\delta\) on which \[\sup_{\boldsymbol{x}\in [0,2]^d} |\widetilde{\mathbb{L}}_n(\boldsymbol{x})| \le (188/3) \cdot d \cdot \sqrt{2 \log(1/\delta)} = (188\sqrt 2/3) \cdot d r \sqrt{k}.\] Combining the above, we find that on \(\Omega_2\cap\Omega_3\) \[\sup_{\boldsymbol{x}\in [0,1]^d} \big|\mathbb{L}_n(\boldsymbol{x}) \big| \le (C_{s}+188\sqrt 2/3)dr \sqrt{k} + B_{n,k}(L;[0,2]^d) =\sqrt k \zeta_{n,1}.\] As a consequence, \(\Omega_2\cap\Omega_3 \subseteq\Omega_1(n,\zeta_{n,1})\) with \(\Omega_1(\cdot,\cdot)\) from 82 . By an application of the second part of Theorem 17 with \(A=\mathfrak B^{\oplus C_{s}r}\) we obtain \[\sqrt{k} \big( \hat{\theta}_n- \theta_0 \big) = 2V_{\theta_0}^{-1} J_{\theta_0}^\top \int_{[0,1]^d\setminus \mathfrak B^{\oplus C_{s}r}} \boldsymbol{g}(\boldsymbol{x}) \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x}) \, \mathrm d\mu(\boldsymbol{x}) + \sqrt{k} \boldsymbol{r}_{n,1}(\beta,\eta) + \boldsymbol{r}_{n,2}(\mathfrak B^{\oplus C_{s}r})\] where \[\|r_{n,2}(\mathfrak B^{\oplus C_{s}r})\|_2 \le \tilde{C}_{r2} \Big( \zeta_{n,2} + \sqrt{k} \zeta_{n,1}\int_{\mathfrak B^{\oplus C_{s}r}} \|\boldsymbol{g}(\boldsymbol{x})\|_2 \, \mathrm d\mu(\boldsymbol{x}) \Big).\] In the following, with a slight abuse of notation, we extend the definition of \(\begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n\) to \([0,1]^d\) by replacing the partial derivatives of \(L\) by the their right-hand side counterparts as described in the paragraph before Theorem 12. Then, by an application of Lemma 10 we have, on an event \(\Omega_4\) that has probability at least \(1-(d+1)\delta\), \[\begin{align} \sup_{\boldsymbol{x}\in [0,1]^d}| \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x})| \le \sup_{\boldsymbol{x}\in [0,1]^d}|\widetilde{\mathbb{L}}_n(\boldsymbol{x})| + \sum_{j \in [d]}\sup_{x \in [0,1]} |\widetilde{\mathbb{L}}_{nj}(\boldsymbol{x})| \le 2\cdot(188/3)\cdot dr \cdot \sqrt{k}. \end{align}\] Thus, on the event \(\Omega_2\cap\Omega_3\cap\Omega_4\) and using 101 , we have \[\begin{align} &\Big\|2V_{\theta_0}^{-1} J_{\theta_0}^\top \int_{[0,1]^d\setminus \mathfrak B^{\oplus C_{s}r}} \boldsymbol{g}(\boldsymbol{x}) \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x}) \, \mathrm d\mu(\boldsymbol{x}) - 2V_{\theta_0}^{-1} J_{\theta_0}^\top \int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x}) \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x}) \, \mathrm d\mu(\boldsymbol{x})\Big\|_2 \\ & \le 4 \cdot(188/3)\cdot dr \cdot \sqrt{k} \cdot \int_{\mathfrak B^{\oplus C_{s}r}} \|V_{\theta_0}^{-1} J_{\theta_0}^\top \boldsymbol{g}(\boldsymbol{x})\|_2 \, \mathrm d\mu(\boldsymbol{x}) \\ & \le 4 \sqrt{k} \zeta_{n,1} \int_{\mathfrak B^{\oplus C_{s}r}} \|V_{\theta_0}^{-1} J_{\theta_0}^\top \boldsymbol{g}(\boldsymbol{x})\|_2 \, \mathrm d\mu(\boldsymbol{x}) \\&\le 4C_V^{-1} \sqrt{sq} C_{\partial} \cdot \sqrt{k} \zeta_{n,1}\int_{\mathfrak B^{\oplus C_{s}r}} \| \boldsymbol{g}(\boldsymbol{x})\|_2 \, \mathrm d\mu(\boldsymbol{x}) . \end{align}\] Noting that \(\Omega_2\cap\Omega_3\cap\Omega_4\) has probability at least \(1 - 7(d+1)\delta\) and that \[2V_{\theta_0}^{-1} J_{\theta_0}^\top \int_{[0,1]^d} \boldsymbol{g}(\boldsymbol{x}) \begingroup \def\mathaccent##1##2{ \kern 0.8\dimexpr\macc@kerna \overline{\kern-0.8\dimexpr\macc@kerna\macc@nucleus\kern 0.2\dimexpr\macc@kerna} \kern-0.2\dimexpr\macc@kerna } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \macc@nested@a\relax 111{\mathbb{L}} \endgroup _n(\boldsymbol{x}) \, \mathrm d\mu(\boldsymbol{x}) = \frac{1}{\sqrt{k}} \sum_{i=1}^n \big( Z_{i,n} - \mathbb{E}[Z_{i,n}] \big)\] by definition of \(Z_{i,n}\) in 23 completes the proof, after increasing \(\tilde{C}_{r2}\). ◻
Proof of Theorem 13. Without loss of generality, \[\begin{align} \label{eq:helpMestBmax} \frac{\log^5(sn)}{k} \leq 1, \quad B_{n,k}^{\mathcal{I}} \le 1 \end{align}\tag{102}\] as otherwise the right-hand side in the theorem is greater than 1. We start by bounding \(d_K(\boldsymbol{S}_n, \boldsymbol{T}_n)\). By Lemma 14, for any \(\lambda>0\), \[\begin{align} d_K(\boldsymbol{S}_n , \boldsymbol{T}_n) &\le \mathbb{P}\big(\|\boldsymbol{S}_n-\boldsymbol{T}_n\|_\infty \ge \lambda\big) + \sup_{\boldsymbol{x} \in \mathbb{R}^{s}} \mathbb{P}( \boldsymbol{T}_n \le \boldsymbol{x}+\lambda \boldsymbol{1} ) - \mathbb{P}( \boldsymbol{T}_n \le \boldsymbol{x}-\lambda \boldsymbol{1}) \\ &\leq \sum_{I \in \mathcal{I}} \mathbb{P}\big( \|\boldsymbol{S}_n^I - \boldsymbol{T}_n^I\|_2 \geq \lambda\big)+\sup_{\boldsymbol{x} \in \mathbb{R}^{s}} \mathbb{P}( \boldsymbol{T}_n \le \boldsymbol{x}+\lambda \boldsymbol{1} ) - \mathbb{P}( \boldsymbol{T}_n \le \boldsymbol{x}-\lambda \boldsymbol{1}), \end{align}\] where we have used the union bound and the fact that \(\| \cdot \|_\infty \le \| \cdot \|_2\). By the assumption that \(\sigma^2_{\min}>0\), with the same reasoning as in the proof of Theorem 7, specifically 50 , \[\begin{align} \sup_{\boldsymbol{x} \in \mathbb{R}^{s}} \mathbb{P}( \boldsymbol{T}_n \le \boldsymbol{x}+\lambda \boldsymbol{1} ) - \mathbb{P}( \boldsymbol{T}_n \le \boldsymbol{x}-\lambda \boldsymbol{1}) \leq \frac{8\lambda}{\sigma_{\min}^2} \sqrt{\log s} + 2 d_K(\boldsymbol{T}_n, \boldsymbol{G}_n). \end{align}\] Thus \[\begin{align} \label{eq:dk-sn-gn-m-estimators} d_K(\boldsymbol{S}_n, \boldsymbol{G}_n) \leq \sum_{I \in \mathcal{I}} \mathbb{P}(\left\lVert\boldsymbol{S}_n^I - \boldsymbol{T}_n^I\right\rVert_2 \geq \lambda) + \frac{8\lambda}{\sigma_{\min}^2} \sqrt{\log(s)} + 3 d_K(\boldsymbol{T}_n, \boldsymbol{G}_n). \end{align}\tag{103}\] In the remaining proof, we will choose a suitable \(\lambda\) to balance the second and third term and bound \(d_K(\boldsymbol{T}_n, \boldsymbol{G}_n)\).
Bounding \(d_K(\boldsymbol{T}_n, \boldsymbol{G}_n)\). We will apply Theorem 19 to the random vector \(\boldsymbol{Y}_{i,n} \in \mathbb{R}^{s}\) with entries given by \(Y_{i,n,(I,t)} = k^{-1/2} (Z_{i,n}^{I,t} - \mathbb{E}[Z_{i,n}^{I,t}])\), enumerating over \(I \in \mathcal{I}\) and \(t \in [s^I]\). By assumption, \(\mathbb{E}[Y_{i,n,(I,t)}^2] \geq \sigma^2_{\min}\).
Next, note that \[|Z_{i,n}^{I,t}| \le \| \boldsymbol{Z}_{i,n}^I\|_2\le 2 {\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert V_{\theta_0^I}^{-1}J_{\theta_0^I}^\top \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2 \| \boldsymbol{A}^I_{i,n}\|_2 \le 2 (C_V^{\mathcal{I}})^{-1}\sqrt{s^{\mathcal{I}} q^{\mathcal{I}}} C_{\partial}^{\mathcal{I}} \cdot \| \boldsymbol{A}^I_{i,n}\|_2\] where \(s^{\mathcal{I}} =\max_{I \in \mathcal{I}} s^I\) and \(q^{\mathcal{I}} = \max_{I \in \mathcal{I}} q^I\). Further, by the same argumentation that lead to 85 , we have \[\begin{align} \left\lVert\boldsymbol{A}^I_{i,n}\right\rVert_2 &\leq C_g^{\mathcal{I}} \sup_{\boldsymbol{x}\in [0,1]^I} \Big| \mathbf{1}\!\Big(\exists j\in I:\; V_{ij}<\frac{k}{n}x_j\Big) - \sum_{j\in I} \partial_j \widetilde{L}_I(\boldsymbol{x}_I)\, \mathbf{1}\!\Big(V_{ij}<\frac{k}{n}x_j\Big) \Big| \\ &\leq C_g^{\mathcal{I}} \left|I\right| \leq C_g^{\mathcal{I}} m. \end{align}\] Combining the previous two inequalities, we obtain that \[\begin{align} \label{eq:deftildeC} \sup_{I \in \mathcal{I}, t \in [q^I]}\big| Z_{i,n}^{I,t}-\mathbb{E}[Z_{i,n}^{I,t}]\big| \leq 2 (C_V^{\mathcal{I}})^{-1}\sqrt{s^{\mathcal{I}} q^{\mathcal{I}}} C_{\partial}^{\mathcal{I}}C_g^{\mathcal{I}} m =: \tilde{C}. \end{align}\tag{104}\] Also, \[\begin{align} \left\lVert\boldsymbol{A}^I_{i,n}\right\rVert_2 &\leq C_g^{\mathcal{I}} \Big\{ \sup_{\boldsymbol{x}_I \in [0,1]^I} \mathbf{1}\Big(\exists j\in I:\; V_{ij}<\frac{k}{n}x_j\Big) + \sum_{j \in I} \sup_{\boldsymbol{x}_I \in [0,1]^I} \mathbf{1}\Big(V_{ij}<\frac{k}{n}x_j\Big) \Big\} \\ &\leq 2C_g^{\mathcal{I}} \sum_{j \in I} \sup_{\boldsymbol{x}_I \in [0,1]^I} \mathbf{1}\Big(V_{ij}<\frac{k}{n}x_j\Big) \\ &\leq 2C_g^{\mathcal{I}} \sum_{j \in I} \mathbf{1}\Big(V_{ij}<\frac{k}{n}\Big), \end{align}\] which yields \[\begin{align} \mathbb{E}\big[\big| Z_{i,n}^{I,t} -\mathbb{E}[Z_{i,n}^{I,t}] \big| \big] \le 2 \mathbb{E}\big| Z_{i,n}^{I,t}\big| &\le 4(C_V^{\mathcal{I}})^{-1}\sqrt{s^{\mathcal{I}} q^{\mathcal{I}}} C_{\partial}^{\mathcal{I}} \cdot \mathbb{E}\| \boldsymbol{A}_{i,n}^{I} \|_2 \\&\le 8(C_V^{\mathcal{I}})^{-1}\sqrt{s^{\mathcal{I}} q^{\mathcal{I}}} C_{\partial}^{\mathcal{I}} C_g^{\mathcal{I}} m\cdot \frac{k}{n} =4 \tilde{C} \frac{k}{n}. \end{align}\] Hence, \[\begin{align} \sum_{i=1}^n \mathbb{E}\Big| k^{-1/2} \big(Z_{i,n}^{I,t}-\mathbb{E}[ Z^{I,t}_{i,n}] \big)\Big|^4 &\leq \tilde{C}^3 \frac{n}{k^2} \mathbb{E}\Big[\big| Z_{1,n}^{I,t} -\mathbb{E}[Z_{1,n}^{I,t}] \big| \Big] \leq 4\tilde{C}^4 \frac{1}{k} = \frac{b_2 B_n^2}{n}, \end{align}\] where \(B_n = (\log 2)^{-1}\tilde{C}\sqrt{n/k}\) and \(b_2 = 4(\log 2)^2 \tilde{C}^{2}\). With these choices, we also have \[\frac{\sqrt{n}\big| k^{-1/2}(Z_{i,n}^{I,t}-\mathbb{E}[Z_{i,n}^{I,t}])\big| }{B_n} \le \frac{\sqrt{n/k} \tilde{C}}{B_n} = \log2,\] and hence the conditions of Theorem 19 hold with \(b_1 = \sigma^2_{\min}\). An application of the theorem yields \[\begin{align} \label{eq:dk-tn-gn-m-estimators} d_K(\boldsymbol{T}_n, \boldsymbol{G}_n) \leq C_1 \Big( \frac{\log^5(s n)}{k}\Big)^{1/4}, \end{align}\tag{105}\] with \(C_1\) depending only on \(\sigma^2_{\min}\) and \(\tilde{C}\).
Choosing \(\lambda\). Let \[\begin{align} \label{eq:helpMestdefdelta} \delta = \frac{1}{m |\mathcal{I}|}\Big(\frac{\log^5(sn)}{k}\Big)^{1/4} \end{align}\tag{106}\] and recall \(r = r(\delta, 1,k) = \sqrt{k^{-1}\log(1/\delta)}\) from 13 . By our assumption in 102 from the beginning of the proof, we have \(\delta \le 1/(m|\mathcal{I}|) < e^{-1}\), and we will later verify that this \(\delta\) also satisfies the conditions \(\log(m/\delta)\le 2k/7\) and \(C_{s}r \le 1\). We may therefore apply Theorem 12 for each tuple \((L_I, \{L_I(\cdot; \theta^I): \theta^I \in \Theta^I\}, \boldsymbol{g}^I, \mu^I)\), which yields the existence of certain constants \(D_1^I, D_2^I>0, \tilde{C}_\beta^I, \tilde{C}_\eta^I \in (0,1]\) and \(\tilde{C}_{r1}^I, \tilde{C}_{r2}^I>0\) such that certain claims hold with probability at least \(1-7(|I|+1)\delta \ge 1-7(m+1)\delta\). In the following, we will make use of these claims.
Let \(\tilde{C}_\beta^{\mathcal{I}} = \min_{I \in \mathcal{I}} \tilde{C}_\beta^I\) and \[\begin{align} \label{eq:definition-zeta95n1-mest} \zeta_{n,1}^{\mathcal{I}} := \max_{I \in \mathcal{I}} \zeta_{n,1}^I, \qquad \zeta_{n,1}^I := \Big( k^{-1/2} \sup_{\boldsymbol{x}_I \in [0,2]^{I}} \left|B_n^I(\boldsymbol{x}_I)\right| +(C_{s}+188\sqrt 2/3)\cdot |I| r \Big). \end{align}\tag{107}\] We will later verify that \(\zeta_{n,1}^{\mathcal{I}} \le \tilde{C}_\beta^{\mathcal{I}},\) which in turn implies that \(\zeta_{n,1}^I \le C_\beta^{I}\) for each \(I \in \mathcal{I}\). Theorem 12 therefore guarantees that, for any \(\eta< \tilde{C}_{\eta}^{\mathcal{I}} := \min_{I \in \mathcal{I}} \tilde{C}_\eta^I\), \[\begin{align} \label{eq:dk-sni-tni-m-estimators} \sum_{I \in \mathcal{I}} \mathbb{P}\big( \|\boldsymbol{S}_n^I - \boldsymbol{T}_n^I\|_2 \geq \lambda_{n,k}(\delta) \big) \leq 7|\mathcal{I}|(m+1)\delta \leq 14|\mathcal{I}|m\delta = 14\cdot \Big(\frac{\log^5(sn)}{k}\Big)^{1/4} \end{align}\tag{108}\] where \(\lambda_{n,k}(\delta) := \zeta_{n,2}^{\mathcal{I}} + \zeta_{n,3}^{\mathcal{I}}\) with \[\begin{align} \zeta_{n,2}^{\mathcal{I}} & := \max_{I \in \mathcal{I}} \zeta_{n,2}^I, \qquad \zeta_{n,3}^{\mathcal{I}} := \max_{I \in \mathcal{I}} \zeta_{n,3}^I, \end{align}\] where, recalling that \(\gamma_{h}=1\) by assumption, \[\begin{align} \zeta_{n,2}^I &= \sqrt{\tilde{C}_{r1}^I k \{ (\zeta_{n,1}^{I})^{3} + \eta\} }, \\ \zeta_{n,3}^I &= \tilde{C}_{r2}^I \Big(\sup_{\boldsymbol{x}_I \in [0,2T]^{I}} |B_n^I(\boldsymbol{x}_I)| + \frac{\left|I\right|}{\sqrt{k}} + D_{1}^I \sqrt{r\log\Big(\frac{D_{2}^I}{\delta r}\Big)} \Big) \\ & + \sqrt k \zeta_{n,1}^I \int_{\mathfrak B_I^{\oplus C_{s}r}} \|{\boldsymbol{g}}^I(\boldsymbol{x}_I)\|_2 \, \mathrm d\mu^I(\boldsymbol{x}_I) \Big). \end{align}\] In summary, from 103 , 105 and 108 , \[\begin{align} \label{eq:bounddKprelim1} d_K(\boldsymbol{S}_n, \boldsymbol{G}_n) \lesssim (\zeta_{n,2}^{\mathcal{I}} + \zeta_{n,3}^{\mathcal{I}}) \sqrt{\log(s)} + \Big(\frac{\log^5(sn)}{k}\Big)^{1/4}, \end{align}\tag{109}\] where the constant in \(\lesssim\) depends on \(\sigma_{\min}\) and \(\tilde{C}\) from 104 .
We now bound \(\zeta_{n,2}^{\mathcal{I}}\). First, by our choice of \(\delta\) in 106 , \[\begin{align} \label{eq:boundhelpr951} r = \sqrt{\frac{1}{k} \log\Big( \frac{m|\mathcal{I}|k^{1/4}}{\log^{5/4}(sn)}\Big)} \leq \sqrt{\frac{1}{k} \log(m s k^{1/4})} \leq C_2 \sqrt{\frac{\log(sk)}{k}} \end{align}\tag{110}\] with \(C_2 = \sqrt{1+\log m}\) as \[\log(msk^{1/4}) \leq \log(m)+\log(sk) \leq \log(sk) \autobracket*{1 + \log m} = C_2^2 \log(sk),\] using that \(\log(sk) \geq \log(6) \ge 1\). Recall our assumption \(B_{n,k}^{\mathcal{I}} \le 1\) from the beginning of the proof. Moreover, since \(\delta \le e^{-1}\), we have \(k^{-1/2} \le r \le mr\) and thus, by 107 and 110 , \[\begin{align} \label{eq:helpboundzetan1} \zeta_{n,1}^{\mathcal{I}} \le (1+C_{s}+188\sqrt 2/3) \cdot mr \le C_3 \sqrt{\frac{\log(sk)}{k}} \end{align}\tag{111}\] where \(C_3 := C_2 (C_{s}+1+188\sqrt 2/3)m\). Hence, by subadditivity of \(x \mapsto \sqrt x\) on \([0,\infty)\), \[\begin{align} \zeta_{n,2}^{\mathcal{I}} &\le \sqrt{\tilde{C}_{r1}^{\mathcal{I}} k} \times \big\{ (\zeta_{n,1}^{\mathcal{I}})^{3/2} + \sqrt{\eta} \big\} \leq C_4 \Big\{ \Big(\frac{\log^3(sk)}{k}\Big)^{1/4} + \sqrt{k\eta}\Big\}, \label{eq:boundzetan2} \end{align}\tag{112}\] where \(\tilde{C}_{r1}^{\mathcal{I}} = \max_{I \in \mathcal{I}}\tilde{C}_{r1}^{I}\) and where \(C_4 = (\tilde{C}_{r1}^{\mathcal{I}})^{1/2}C_3^{3/2}\).
Next, we bound \(\zeta_{n,3}^{\mathcal{I}}\). First, using that \[\delta = \frac{1}{m|\mathcal{I}|} \Big( \frac{\log^5(sn)}{k} \Big)^{1/4} \ge \frac{1}{m|\mathcal{I}|k^{1/4}} \ge \frac{1}{msk^{1/4}}\] and \(r \geq k^{-1/2}\), we obtain, recalling 110 , \[\begin{align} \max_{I \in \mathcal{I}} D_1^I\sqrt{r\log\Big(\frac{D_{2}^I}{\delta r}\Big)} &\leq D_1^{\mathcal{I}}C_2^{1/2} \Big(\frac{\log(sk)}{k}\Big)^{1/4} \sqrt{\log(D_2^{\mathcal{I}} ms k^{3/4})} \leq C_5 \Big( \frac{\log^3(sk)}{k}\Big)^{1/4} \end{align}\] where \(D_j^{\mathcal{I}} = \max_{I \in \mathcal{I}} D_j^I\) and \(C_5 = D_1^{\mathcal{I}}C_2^{1/2} \{1 + \log(D_2^{\mathcal{I}} m)\}^{1/2}\) and we used \(sk\ge 3\) so that \[\log(D_2^{\mathcal{I}} ms k^{3/4}) = \log(D_2^{\mathcal{I}}m) + \log(s k^{3/4}) \leq \log(sk)\{ 1 + \log(D_2^{\mathcal{I}} m) \}.\] Together with 111 and \(C_{s}r \le \zeta_n\) by 110 with \(\zeta_n\) from the formulation of the theorem, we obtain that \[\begin{align} \label{eq:boundzetan3} \zeta_{n,3}^{\mathcal{I}} \le\tilde{C}_{r2}^{\mathcal{I}} \Big\{ B_{n,k}^{\mathcal{I}} + (C_5 +m)\Big( \frac{\log^3(sk)}{k}\Big)^{1/4} + C_3 \sqrt{\log(sk)} \int_{\mathfrak B_I^{\oplus \zeta_n}} \|{\boldsymbol{g}}^I(\boldsymbol{x}_I)\|_2 \, \mathrm d\mu^I(\boldsymbol{x}_I) \Big\} \end{align}\tag{113}\]
Combining the bounds in 109 , 112 and 113 we obtain \[\begin{align} d_K(\boldsymbol{S}_n, \boldsymbol{G}_n) &\lesssim \Big(\frac{\log^5(sn)}{k}\Big)^{1/4} + \sqrt{\log s}\Big\{ B_{n,k}^{\mathcal{I}} + \Big(\frac{\log^3(sk)}{k}\Big)^{1/4} \\ &+ \sqrt{k\eta} + \sqrt{\log(sk)} \int_{\mathfrak B_I^{\oplus \zeta_n}} \|{\boldsymbol{g}}^I(\boldsymbol{x}_I)\|_2 \, \mathrm d\mu^I(\boldsymbol{x}_I)\Big\}, \end{align}\] which implies the assertion.
It remains to verify that \(\delta\) as defined in 106 satisfies \(\log(m/\delta)\le 2k/7\) and \(C_{s}r \le 1\) and \(\zeta_{n,1}^{\mathcal{I}} \le \tilde{C}_\beta^{\mathcal{I}}\). First, \[\begin{align} \log\autobracket*{\frac{m}{\delta}} = \log\Big( \frac{m^2 \left|\mathcal{I}\right| k^{1/4}}{\log^{5/4}(sn)} \Big) \leq \log(m^2 \left|\mathcal{I}\right| k^{1/4}) \leq 2k/7 \end{align}\] by assumption (ii). Next, by assumption (iii), \[\begin{align} r = \sqrt{\frac{1}{k}\log\Big( \frac{1}{\delta}\Big)} = \sqrt{\frac{1}{k}\log\Big( \frac{m|\mathcal{I}|k^{1/4}}{\log^{5/4}(sn)} \Big)} \le \sqrt{\frac{1}{k}\log( m|\mathcal{I}|k^{1/4})} \le \frac{1}{C_{s}} \end{align}\] Finally, from 110 and 111 , \[\zeta_{n,1}^{\mathcal{I}} \le m(1+C_{s}+ 188\sqrt{2}/3) \sqrt{ \frac{1}{k}\log(m|\mathcal{I}|k^{1/4})}.\] The right-hand side is upper bounded by \(\tilde{C}_\beta^{\mathcal{I}}\) if we choose \(\tilde{C}_{k}^{\mathcal{I}} = [\tilde{C}_\beta^{\mathcal{I}} / \{m(1+C_{s}+ 188\sqrt{2}/3)\}]^2\). This completes the proof. ◻
The following lemma is a version of the argument on page 7 in [24], with the precise constant \(188/3\) deduced from [26].
Lemma 10. Let \(n \in \mathbb{N},k \in [n],d \in \mathbb{N}, T >0,\delta \in (0,e^{-1})\) and \(\emptyset \ne I \subseteq[d]\) satisfy \(\log(1/\delta)\le |I|^2 T k\). Then \[\sup_{\boldsymbol{x} \in [0,T]^I} | \widetilde{\mathbb{L}}_{n,I}(\boldsymbol{x})| \le (188/3) \cdot |I| \cdot \sqrt{T \log(1/\delta)}\] with probability at least \(1-\delta.\)
Proof. Fix \(I \subseteq[d]\), write \(m=|I|\) and define \(\mu_{n,I} = \frac{1}{n} \sum_{i=1}^n \delta_{V_{i,I}}\) and let \(\mu_I\) denote the distribution of \(V_{i,I}\). Then we can write \[\sup_{\boldsymbol{x} \in [0,T]^I} | \widetilde{\mathbb{L}}_{n,I}(\boldsymbol{x})| = \frac{n}{\sqrt k} \sup_{A \in \mathcal{A}} | \mu_{n,I}(A) - \mu_I(A)|\] where \(\mathcal{A}\) contains all sets of the form \(A_{\boldsymbol{x}} = \{ \boldsymbol{z} \in [0,\infty)^I \mid \exists j \in I: z_j < (k/n) x_j \}\) with \(\boldsymbol{x} \in [0, T]^I\). Let \(\mathbb{A} := \bigcup_{A\in \mathcal{A}} A\), with \(p = \mu(\mathbb{A}) = \mathbb{P}(\exists j \in I : V_{ij} \le \frac{k}{n} T) \le mTk/n\). By Theorem A.1 in [26] we have, with probability at least \(1-\delta\), \[\sup_{A \in \mathcal{A}} | \mu_n(A) - \mu(A)| \le \frac{2}{3n} \log(1/\delta) + \sqrt{\frac{mTk}{n^2}} \Big\{ 2 \sqrt{\log(1/\delta)}+ 60 \sqrt m\Big\},\] where we have used that the VC-dimension of \(\mathcal{A}\) is \(m\). Since \(1 \le \log(1/\delta) \le m^2 T k\), we get the upper bound \[\begin{align} \sup_{\boldsymbol{x} \in [0,T]^I} | \tilde{\mathbb{L}}_{n,I}(\boldsymbol{x})| &\le \frac{2}{3 \sqrt k} \log(1/\delta) + \sqrt{mT} \Big\{ 2 \sqrt{\log(1/\delta)}+ 60 \sqrt m\Big\} \\&\le \frac{2}{3\sqrt k} m\sqrt{Tk \log(1/\delta)} + \sqrt{mT} \Big\{2\sqrt{\log(1/\delta)} + 60\sqrt{m}\Big\} \\& \le m\sqrt{T\log(1/\delta)}\Big\{ \frac{2}{3} + \frac{2}{\sqrt m} + 60 \Big\} \le (188/3) m \sqrt{T \log(1/\delta)} \end{align}\] with probability at least \(1-\delta\). ◻
Recall \(S_{nj}(x_j) = (n/k) \cdot V_{\lceil kx_j \rceil, j} \cdot \boldsymbol{1}(x_j>0)\) from 32 . The following lemma is akin to Lemma 9 in [24].
Lemma 11 (Bound on order statistics). Let \(C_{s}=188\sqrt{2}/3 + \sqrt{1-\log2} \approx 89.18\). For any \(n,d,k,T \in \mathbb{N}\) and \(\delta \in (0, e^{-1})\) with \(k\in[n]\) and \(\log(d/ \delta) \le (1-\log 2) kT \approx 0.31\cdot kT\) we have \[\begin{align} \label{eq:bound-on-order-statistics-1} \max_{j\in[d]} \sup_{x_j \in [0,T]} S_{nj}(x_j) \le 2T \end{align}\tag{114}\] with probability larger than \(1-\delta\). Moreover, we have \[\begin{align} \label{eq:bound-on-order-statistics-2} \max_{j\in[d]} \sup_{x_j \in [0,T]} |S_{nj}(x_j) - x_j| \le C_{s}\sqrt{\frac{T}{k} \log\Big(\frac{1}{\delta}\Big)} \end{align}\tag{115}\] with probability larger than \(1-(d+1)\delta\), and on the latter event where 115 is met we also have 114 .
Proof of Lemma 11. First, note that \(\sup_{x_j \in [0,T]} S_{nj}(x_j) = (n/k) \cdot V_{kT:n, j}\) by monotonicity. Moreover, writing \(G_{nj}(v_j)=n^{-1}\sum_{i=1}^n \boldsymbol{1}(V_{ij} \le v_j)\), we have \(V_{\ell:n} \le x\) iff \(G_{nj}(x) \ge \ell/n\) for all \(\ell \in[n]\) and \(x \in \mathbb{R}\), which implies \[\frac{n}{k} V_{kT:n, j} \le 2T \quad \Longleftrightarrow \quad G_{nj}\Big( 2\frac{kT}{n} \Big) \ge \frac{kT}{n}.\] As a consequence, by the union bound, \[\begin{align} \mathbb{P}\Big(\max_{j\in[d]} \sup_{x_j \in [0,T]} S_{nj}(x_j) > 2 T \Big) \le d \cdot \mathbb{P}\Big( G_{nj}\Big( 2\frac{kT}{n} \Big) < \frac{kT }{n} \Big) &\le d \cdot \big(\sqrt 2 e^{-1/2}\big)^{2kT} \\&= d \cdot \exp\big( - (1-\log2) kT \big\}, \end{align}\] where the second inequality follows from the multiplicative Chernoff bound; see, for instance, Exercise 2.11 in [60]. By our assumption \(\log(d/\delta) \le (1-\log2) kT\), the upper bound in the previous display is smaller than \(\delta\). This proves 114 .
We may now proceed analogously to the proof of Lemma 9 in [24] to show that \[\begin{align} \label{eq:bound-gauss} \max_{j\in[d]} \sup_{x_j \in [0,T]} \Big|S_{nj}(x_j) - \frac{\lceil kx_j\rceil}{k}\Big| \le (188\sqrt{2}/3) \sqrt{\frac{T}{k} \log\Big(\frac{1}{\delta}\Big)} \end{align}\tag{116}\] with probability at least \(1-(d+1)\delta\). Indeed, by the definition of \(S_{nj}\) in 32 , we have, on the event in 114 , \[\begin{align} \sup_{x_j \in [0,T]} \Big|S_{nj}(x_j) - \frac{\lceil kx_j\rceil}{k}\Big| &= \sup_{x_j \in (0,T]} \Big| S_{nj}(x_j) - \frac{n}{k} G_{nj}\big(V_{\lceil kx_j\rceil:n, j}\big) \Big| \\&= \frac{n}{k} \sup_{x_j \in (0,T]} \Big| \frac{k}{n} S_{nj}(x_j) - G_{nj}\Big( \frac{k}{n} S_{nj}(x_j) \Big) \Big| \\&\le \frac{n}{k} \sup_{x_j \in [0, 2T]} \Big| \frac{k}{n} x_j - G_{nj}\Big( \frac{k}{n} x_j \Big) \Big| \\&= \sup_{x_j \in [0, 2T]} \Big| x_j - \widetilde{L}_{nj}( x_j ) \Big| =\frac{1}{\sqrt k}\sup_{x_j \in [0,2T]} | \tilde{\mathbb{L}}_{nj}(x_j)| \end{align}\] where we used that \(\frac{n}{k} G_{nj}(\frac{k}{n} x_j - ) = \widetilde{L}_{nj}(x_j)\). As a result, since \(\log(1/\delta) \le \log(d/\delta) \le ( 1-\log2) kT \le 2Tk\), the assertion in 116 follows from Lemma 10, applied with \(T\) replaced by \(2T\), and the union bound. Finally, the result in 115 follows from the triangular inequality, observing that \[\sup_{x_j \in [0,T]} \Big| \frac{\lceil kx_j\rceil}{k} - x_j\Big| \le \frac{1}{k} \le \sqrt{1-\log2} \sqrt{\frac{T}{k} \log \Big(\frac{1}{\delta}\Big)},\] again using that \(\log(1/\delta) \le ( 1-\log2) kT\). ◻
Recall that \(\boldsymbol{V}_1, \boldsymbol{V}_2, \dots\) are iid random vectors in \([0,1]^d\) with standard uniform margins. For \(\boldsymbol{u} \in \mathbb{R}^d\), the interesting points being \(\boldsymbol{u} \in [0,1]^d\), let \[\begin{align} \alpha_{n}(\boldsymbol{u}) &= \tag{117} \frac{1}{\sqrt n} \sum_{i=1}^n \big[ \boldsymbol{1}(\forall j \in [d]: V_{ij} < u_j) - \mathbb{P}(\forall j \in [d]: V_{ij} < u_j) \big], \\ \beta_{n}(\boldsymbol{u}) &= \tag{118} \frac{1}{\sqrt n} \sum_{i=1}^n \big[ \boldsymbol{1}(\exists j \in [d]: V_{ij} < u_j) - \mathbb{P}(\exists j \in [d]: V_{ij} < u_j) \big]. \end{align}\]
Lemma 12. Fix \(d \in \mathbb{N}\), \(0 \le a_j < b_j \le 1\) for \(j \in [d]\), \(\varepsilon\in (0, \min_{j \in [d]}(b_j-a_j)]\), and \(\delta \in (0,e^{-1})\). Then, for any \(n\in\mathbb{N}\), there exists an event \(\Omega\) of probability at least \(1-\delta\) such that, on \(\Omega\), \[\begin{align} \label{eq:modulus95beta1} \omega_{\alpha_{n}}(\varepsilon; [\boldsymbol{a}, \boldsymbol{b}]) &\le \nonumber 2d \Big[ \frac{2}{3\sqrt n} \log\Big( \frac{2 \| \boldsymbol{b} - \boldsymbol{a} \|_1}{\varepsilon\delta}\Big) +\Big\{2 \sqrt{\varepsilon\log\Big( \frac{2\| \boldsymbol{b} - \boldsymbol{a} \|_1}{\varepsilon\delta}\Big)}+ 60\sqrt{2d\varepsilon} \Big\} \Big] \\&\le \kappa \sqrt{\varepsilon\log\Big( \frac{2\| \boldsymbol{b} - \boldsymbol{a} \|_1}{\varepsilon\delta}\Big)}, \end{align}\tag{119}\] where \(\omega_{\alpha_n}\) is the modulus of continuity defined in 1 and where \[\kappa = 2d \Big[ \sqrt{\frac{4}{9n\varepsilon} \log\Big( \frac{2\| \boldsymbol{b} - \boldsymbol{a} \|_1}{\varepsilon\delta}\Big)} + 2 + 60 \sqrt{2d} \Big].\] The same inequality holds with \(\alpha_{n}\) replaced by \(\beta_{n}\), also with probability at least \(1-\delta\).
Proof. The proof is largely inspired by [61]. For \(j \in [d]\) and \(k \in K_j := \{ 1, \dots, \lceil (b_j-a_j)/\varepsilon\rceil \}\) define \[\mathcal{A}_{j,k} = \Big\{ [\boldsymbol{x},\boldsymbol{y}) \subseteq[0,1)^d : \;a_j+\varepsilon(k - 1) \le x_j < y_j \le a_j+\varepsilon k \Big\},\] which has VC-dimension \(2d\). Next, let \(\mathbb{A}_{j,k} = \bigcup_{A \in \mathcal{A}_{j,k}} A\), and note that for all \(j \in [d], k \in K_j\) we have \(\mathbb{P}(\boldsymbol{V} \in \mathbb{A}_{j,k}) \le \mathbb{P}(V_j \in [a_j+\varepsilon(k-1), a_j+\varepsilon k]) \le \varepsilon\).
Let \(\tilde{\delta} >0\). Then, by Theorem A.1 in [26], applied with \(B = \mathbb{A}_{j,k}\), there exists an event \(\Omega_{j,k}\) with probability at least \(1-\tilde{\delta}\) such that, on \(\Omega_{j,k}\), \[\sup_{A\in \mathcal{A}_{j,k}} |\mu_{n}(A) - \mu(A) | \le \frac{2}{3n} \log(1/{\tilde{\delta}}) + \sqrt{\frac{\varepsilon}{n}}\Big\{2 \sqrt{\log(1/{\tilde{\delta}})}+ 60\sqrt{2d} \Big\},\] where \(\mu_{n} = n^{-1} \sum_{i=1}^n \delta_{\boldsymbol{V}_{i}}\) and where \(\mu\) is the distribution of \(\boldsymbol{V}_{i}\). Note that \(|K_j| = \lceil (b_j-a_j)/\varepsilon\rceil \le (b_j-a_j)/\varepsilon+ 1 \le 2(b_j-a_j)/\varepsilon\). On the intersection set \(\Omega_1 = \bigcap_{j \in [d]} \bigcap_{k \in K_j} \Omega_{j,k}\), which has probability at least \(1- \sum_{j \in [d]}|K_j| \tilde{\delta} \ge 1-2 \| \boldsymbol{b} - \boldsymbol{a} \|_1 \tilde{\delta} /\varepsilon\), we obtain that \[\max_{j \in [d]} \max_{k \in K_j }\sup_{A\in \mathcal{A}_{j,k}} |\mu_{n}(A) - \mu(A) | \le \Big[ \frac{2}{3n} \log(1/{\tilde{\delta}}) + \sqrt{\frac{\varepsilon}{n}}\Big\{2 \sqrt{\log(1/{\tilde{\delta}})}+ 60\sqrt{2d} \Big\} \Big].\] Let \[\mathcal{A}:= \Big\{\;R_{k_1, \dots, k_d} := \bigtimes_{j = 1}^d \big[ a_j+(k_j - 1)\varepsilon, (a_j+k_j\varepsilon)\wedge b_j \big] \;: \;k_j \in K_j \Big\}\] denote a cover of \([\boldsymbol{a}, \boldsymbol{b}]\) consisting of axis aligned hyper-rectangles \(R_{k_1, \dots, k_d}\) with edge length at most \(\varepsilon\), and note that \[\begin{align} \omega_{\alpha_{n}}(\varepsilon, [\boldsymbol{a}, \boldsymbol{b}]) &= \sup_{\|\boldsymbol{x} - \boldsymbol{y} \|_\infty \le \varepsilon, \boldsymbol{x},\boldsymbol{y} \in [\boldsymbol{a}, \boldsymbol{b}]} |\alpha_{n}(\boldsymbol{x}) - \alpha_{n}(\boldsymbol{y})| \le 2 \max_{R \in \mathcal{A}} \sup_{\boldsymbol{x}, \boldsymbol{y} \in R} |\alpha_{n}(\boldsymbol{x}) - \alpha_{n}(\boldsymbol{y})| \end{align}\] by the triangle inequality for the \(\| \cdot \|_\infty\)-norm.3
Next, for fixed \(R=R_{k_1 \dots k_d}\) and \(\boldsymbol{x}, \boldsymbol{y} \in R = R_{k_1,\dots,k_d} \subseteq[0,1]^d\) we have \[\begin{align} \alpha_n(\boldsymbol{x}) -\alpha_n(\boldsymbol{y}) &= \alpha_n(x_1, \dots, x_d) \pm \alpha(y_1, x_2, \dots, x_d) \pm \alpha_n(y_1, y_2, x_3, \dots, x_d) \\&\pm \dots \pm \alpha_n(y_1, \dots, y_{d-1}, x_d) - \alpha_n(y_1, \dots, y_d) \\&= \sum_{j \in [d]} \alpha_n(y_{1:j-1},x_{j:d}) - \alpha_n(y_{1:j},x_{j+1:d}) \end{align}\] where \(x_{i:j}=(x_i, \dots, x_j)\) for \(i \le j\), and where \(x_{i:j}\) should be interpreted as ‘not being there’ for \(i>j\). In what follows, with a slight abuse of notation, write \(\alpha_n(A) = \sqrt n\{\mu_n(A)-\mu(A)\}\) for Borel sets \(A\). This defines a finite signed measure. Fix \(j \in [d]\). First consider the case \(x_j > y_j\). Then \[\begin{align} T_{nj}(\boldsymbol{x}, \boldsymbol{y}) :=&\;\alpha_n(y_{1:j-1},x_{j:d}) - \alpha_n(y_{1:j},x_{j+1:d}) \\=&\; \alpha_n(y_{1:j-1}, x_j, x_{j+1:d}) - \alpha_n(y_{1:j-1},y_j, x_{j+1:d}) \\=&\; \alpha_n(A_{j>,\boldsymbol{x}, \boldsymbol{y}}), \end{align}\] with \[A_{j>,\boldsymbol{x}, \boldsymbol{y}} := [0,y_1) \times \dots \times [0,y_{j-1}) \times [y_j, x_j) \times [0,x_{j+1}) \times \dots \times [0,x_{d}) \in \mathcal{A}_{j,k_j}.\] Likewise, if \(x_j < y_j\), we have \[T_{nj}(\boldsymbol{x}, \boldsymbol{y}) = - \alpha_n( A_{j<,\boldsymbol{x}, \boldsymbol{y}}),\] where \[A_{j<,\boldsymbol{x}, \boldsymbol{y}} := [0,y_1) \times \dots \times [0,y_{j-1}) \times [x_j, y_j) \times [0,x_{j+1}) \times \dots \times [0,x_{d}) \in \mathcal{A}_{j,k_j},\] and if \(x_j=y_j\), we have \(T_{nj}(\boldsymbol{x}, \boldsymbol{y}) =0\). Overall, \(|T_{nj}(\boldsymbol{x}, \boldsymbol{y}) | \le \sup_{A \in \mathcal{A}_{j,k_j}} | \alpha_n(A)|\), which implies \[\begin{align} \sup_{\boldsymbol{x}, \boldsymbol{y} \in R} |\alpha_{n}(\boldsymbol{x}) - \alpha_{n}(\boldsymbol{y})| \le \sum_{j \in [d]} \sup_{A \in \mathcal{A}_{j,k_j}} | \alpha_{n}(A) | \le d \max_{j \in [d]}\max_{k_j \in K_j} \sup_{A \in \mathcal{A}_{j,k_j}} | \alpha_{n}(A) |. \end{align}\] Hence, \[\begin{align} \omega_{\alpha_{n}}(\varepsilon, [\boldsymbol{a}, \boldsymbol{b}]) \le 2d \max_{j \in [d]} \max_{k_j \in K_j}\sup_{A\in \mathcal{A}_{j,k_j}} | \alpha_{n}(A) |, \end{align}\] and thus, with probability at least \(1-2 \| \boldsymbol{b} - \boldsymbol{a} \|_1 \tilde{\delta} /\varepsilon\), \[\omega_{\alpha_{n}}(\varepsilon, [\boldsymbol{a}, \boldsymbol{b}]) \le 2d \sqrt n\Big[ \frac{2}{3n} \log(1/{\tilde{\delta}}) + \sqrt{\frac{\varepsilon}{n}}\Big\{2 \sqrt{\log(1/{\tilde{\delta}})}+ 60\sqrt{2d} \Big\} \Big].\] With \(\tilde{\delta} = \varepsilon\delta / ( 2 \| \boldsymbol{b} - \boldsymbol{a} \|_1 )\), the upper bound can be rewritten as \[2d \Big[ \frac{2}{3\sqrt n} \log\Big( \frac{2\| \boldsymbol{b} - \boldsymbol{a} \|_1}{\varepsilon\delta}\Big) +\Big\{2 \sqrt{\varepsilon\log\Big( \frac{2\| \boldsymbol{b} - \boldsymbol{a} \|_1}{\varepsilon\delta}\Big)}+ 60\sqrt{2d\varepsilon} \Big\} \Big],\] which is the first statement of the lemma.
Regarding the second statement concerning \(\beta_{n}\), note that the events of interest in its definition satisfy \[\big\{ \exists j \in [d]: V_{ij}<u_j \big\} = \big\{ \forall j \in [d]: V_{ij} \ge u_j \big\}^c = \big\{ \forall j \in [d]: U_{ij} \le 1-u_j \big\}^c\] where \(U_{ij}=1-V_{ij}\). As a consequence, \[\beta_{n}(\boldsymbol{u}) = - \tilde{\alpha}_{n}^\circ(\boldsymbol{1}-\boldsymbol{u})\] where \[\tilde{\alpha}_{n}^\circ(\boldsymbol{u}) = \frac{1}{\sqrt n} \sum_{i=1}^n \big[ \boldsymbol{1}(\forall j \in [d]: U_{ij} \le u_j) - \mathbb{P}(\forall j \in [d]: U_{ij} \le u_j) \big].\] Hence, \(\omega_{\beta_{n}}(\varepsilon; [\boldsymbol{a}, \boldsymbol{b}]) = \omega_{\tilde{\alpha}_{n}^\circ}(\varepsilon; [\boldsymbol{1}- \boldsymbol{b}, \boldsymbol{1} - \boldsymbol{a}])\). Define \(\alpha_{n}^\circ\) as in 117 , but with \(\boldsymbol{V}_i\) replaced by \(\boldsymbol{U}_i\), and note that the derived probability bound holds for \(\alpha_n^\circ\). Further note that \(\tilde{\alpha}_n^\circ(\boldsymbol{u}) = \lim_{\eta\downarrow 0} \alpha_n^\circ(\boldsymbol{u} + \eta \boldsymbol{1})\) for any \(\boldsymbol{u} \in [0,1)^d\), so that \(\omega_{\tilde{\alpha}_{n}^\circ}(\varepsilon; (\boldsymbol{1}- \boldsymbol{b}, \boldsymbol{1} - \boldsymbol{a})) = \omega_{ \alpha_{n}^\circ}(\varepsilon; (\boldsymbol{1}- \boldsymbol{b}, \boldsymbol{1} - \boldsymbol{a}))\). Moreover, for fixed \(\boldsymbol{a}, \boldsymbol{b}\) we have with probability one \(\tilde{\alpha}_{n}^\circ(\boldsymbol{u}) = \alpha_{n}^\circ(\boldsymbol{u})\) for all \(\boldsymbol{u}\) on the boundary of the set \([\boldsymbol{1}- \boldsymbol{b}, \boldsymbol{1} - \boldsymbol{a}]\), so that in fact \(\omega_{\tilde{\alpha}_{n}^\circ}(\varepsilon; (\boldsymbol{1}- \boldsymbol{b}, \boldsymbol{1} - \boldsymbol{a})) = \omega_{ \alpha_{n}^\circ}(\varepsilon; [\boldsymbol{1}- \boldsymbol{b}, \boldsymbol{1} - \boldsymbol{a}])\) with probability one. The assertion for \(\omega_{\beta_{n}}\) now follows from the probability bound on \(\omega_{\alpha_{n}^\circ}\). ◻
Lemma 13. Let \(L\) be an \(d\)-variate stable tail dependence function satisfying [cond:smoothness-good], and let \(j \in [d]\). Then, for any \(\boldsymbol{y}, \boldsymbol{z} \in E_j\) such that the rectangle \([\boldsymbol{y}, \boldsymbol{z}] = \{ \boldsymbol{x} \in [0,\infty)^d: y_\ell \le x_\ell \le z_\ell \text{ for all } \ell \in [d] \}\) is contained in \(G_j :=G_j^{(1)} \cap \bigcap_{\ell \in [d]} G^{(2)}_{j\ell}\), we have \[|\partial_{j} L(\boldsymbol{y}) - \partial_{j} L(\boldsymbol{z})| \le K_L\max\Big\{\frac{1}{y_j},\frac{1}{z_j}\Big\} \|\boldsymbol{y} - \boldsymbol{z}\|_1.\]
Proof of Lemma 13. For \(t\in[0,1]\), let \(\boldsymbol{x}(t) = \boldsymbol{y} + t (\boldsymbol{z} - \boldsymbol{y})\) denote the line segment connecting \(\boldsymbol{y}\) and \(\boldsymbol{z}\). Note that \(x_j(t)>0\). Since \(\boldsymbol{x}(t) \in [\boldsymbol{y}, \boldsymbol{z}] \subseteq G_j\) by assumption, the function \(f(t) = \partial_{j} L(\boldsymbol{x}(t))\) is well-defined, continuous on \([0,1]\) and continuously differentiable on \((0,1)\) with derivative \[f'(t) = \sum_{\ell \in [d] : y_\ell>0 \text{ or } z_\ell>0 } (z_\ell-y_\ell) \partial_{j\ell} L( \boldsymbol{x}(t)).\] By the mean-value theorem, there exists some \(t^* \in (0,1)\) such that \[\partial_{j} L(\boldsymbol{z}) - \partial_{j} L(\boldsymbol{y}) = f(1) - f(0) = f'(t^*) = \sum_{\ell \in [d] : y_\ell>0 \text{ or } z_\ell>0 } (z_\ell-y_\ell) \partial_{j\ell} L( \boldsymbol{x}(t^*)).\] Hence, by Condition [cond:smoothness-good], \[\begin{align} |\partial_{j} L(\boldsymbol{y}) - \partial_{j} L(\boldsymbol{z})| &\le \max_{\ell \in [d] : y_\ell>0 \text{ or } z_\ell>0 } \sup_{t \in (0,1)} |\partial_{j\ell} L( \boldsymbol{x}(t))| \times \sum_{\ell \in [d] : y_\ell>0 \text{ or } z_\ell>0 } |y_\ell-z_\ell| \\&\le K_L\Big( \sup_{t \in (0,1)} \frac{1}{x_j(t)}\Big) \times \sum_{\ell \in [d]} |y_\ell - z_\ell| \end{align}\] Since the denominator in the supremum on the right-hand side is an affine linear function, the supremum must be attained at one of the boundary points 0 or 1, with \(1/x_j(0)=1/y_j\) and \(1/x_j(1) = 1/z_j\). As a consequence, \(\sup_{t \in (0,1)} 1/x_j(t) =\max(1/y_j, 1/z_j)\), which yields the assertion. ◻
Lemma 14. Suppose \(\boldsymbol{X}, \boldsymbol{Y}\) are \(d\)-variate random vectors defined on the same probability space. Then, for all \(\delta>0\), \[\sup_{\boldsymbol{x} \in \mathbb{R}^d} \big| \mathbb{P}(\boldsymbol{X} \le \boldsymbol{x}) - \mathbb{P}(\boldsymbol{Y}\le \boldsymbol{x}) \big| \le \mathbb{P}\big(\|\boldsymbol{X}-\boldsymbol{Y}\|_\infty \ge \delta\big) + \sup_{\boldsymbol{x} \in \mathbb{R}^d} \mathbb{P}( \boldsymbol{Y} \le \boldsymbol{x}+\delta \boldsymbol{1} ) - \mathbb{P}( \boldsymbol{Y} \le \boldsymbol{x}-\delta \boldsymbol{1}),\] where \(\| \cdot \|_\infty\) is the maximum norm on \(\mathbb{R}^d\).
Proof of Lemma 14. Let \(\Delta = \big\{\|\boldsymbol{X}- \boldsymbol{Y}\|_\infty \geq \delta\big\}\). Then, for any \(\boldsymbol{x} \in \mathbb{R}^d\), \[\begin{align} \mathbb{P}\big(\boldsymbol{X} \le \boldsymbol{x} \big) \ge \mathbb{P}\big(\boldsymbol{X} \le \boldsymbol{x}, \Delta^c\big) &\ge \mathbb{P}\big( \boldsymbol{Y} \le \boldsymbol{x} - \delta \boldsymbol{1}, \Delta^c\big) \\ &= \mathbb{P}\big( \boldsymbol{Y}\le \boldsymbol{x}-\delta\boldsymbol{1} \big) - \mathbb{P}\big(\boldsymbol{Y} \le \boldsymbol{x}-\delta \boldsymbol{1}, \Delta\big) \\& \ge \mathbb{P}\big(\boldsymbol{Y}\le \boldsymbol{x}-\delta\boldsymbol{1} \big) - \mathbb{P}(\Delta). \end{align}\] As a consequence, \[\begin{align} \mathbb{P}\big(\boldsymbol{Y} \leq \boldsymbol{x}\big) -\mathbb{P}\big(\boldsymbol{X} \leq \boldsymbol{x} \big) & \le \mathbb{P}\big(\boldsymbol{Y} \leq \boldsymbol{x}\big) - \mathbb{P}\big(\boldsymbol{Y} \le \boldsymbol{x} - \delta \boldsymbol{1} \big) + \mathbb{P}(\Delta) \\ &\le \mathbb{P}\big(\boldsymbol{Y} \leq \boldsymbol{x}+\delta \boldsymbol{1}\big) - \mathbb{P}\big(\boldsymbol{Y} \le \boldsymbol{x} - \delta \boldsymbol{1} \big) + \mathbb{P}(\Delta). \end{align}\] Likewise, \[\begin{align} \mathbb{P}\big(\boldsymbol{X} \le \boldsymbol{x}, \Delta^c\big) \le \mathbb{P}\big(\boldsymbol{Y} \le \boldsymbol{x} + \delta\boldsymbol{1}, \Delta^c \big) \le \mathbb{P}\big(\boldsymbol{Y} \le \boldsymbol{x} + \delta \boldsymbol{1}), \end{align}\] which implies \[\begin{align} \mathbb{P}\big(\boldsymbol{X} \leq \boldsymbol{x}\big) -\mathbb{P}\big(\boldsymbol{Y} \leq \boldsymbol{x} \big) &= \mathbb{P}\big(\boldsymbol{X} \leq \boldsymbol{x}, \Delta\big) + \mathbb{P}\big(\boldsymbol{X} \leq \boldsymbol{x}, \Delta^c\big) -\mathbb{P}\big(\boldsymbol{Y} \leq \boldsymbol{x} \big) \\&\le \mathbb{P}\big(\boldsymbol{X} \leq \boldsymbol{x}, \Delta\big) + \mathbb{P}\big(\boldsymbol{Y} \le \boldsymbol{x} + \delta\boldsymbol{1}) - \mathbb{P}(\boldsymbol{Y} \le \boldsymbol{x}) \\&\le\mathbb{P}(\Delta) + \mathbb{P}\big(\boldsymbol{Y} \leq \boldsymbol{x}+\delta \boldsymbol{1}\big) - \mathbb{P}\big(\boldsymbol{Y} \le \boldsymbol{x} - \delta \boldsymbol{1} \big). \end{align}\] This concludes the proof. ◻
Theorem 18 (Nazarov). Suppose \(\boldsymbol{Z} \sim \mathcal{N}_d(\boldsymbol{0}, \Sigma)\) such that \(\min_{j=1}^d \operatorname{Var}(Z_j) \ge \sigma_{\min}^2>0\). Then, for every \(\delta>0\), \[\sup_{\boldsymbol{x} \in \mathbb{R}^d} \mathbb{P}( \boldsymbol{Z} \le \boldsymbol{x}+\delta \boldsymbol{1} ) - \mathbb{P}( \boldsymbol{Z} \le \boldsymbol{x}-\delta \boldsymbol{1} ) \le \frac{2\delta}{\sigma_{\min}}\big( 2+ \sqrt{2\log d} \big).\]
Proof. This is Nazarov’s inequality, see [62]. ◻
Theorem 19 ([43]). Let \(\boldsymbol{S}_n = \sum_{i=1}^n \boldsymbol{Y}_{i,n}\) with \(\boldsymbol{Y}_{1,n}, \dots, \boldsymbol{Y}_{n,n}\) independent and with \(\mathbb{E}[\boldsymbol{Y}_{i,n}]=0, \mathbb{E}[Y_{i,n,j}^2]<\infty\), where \(\boldsymbol{Y}_{i,n}=(Y_{i,n,1}, \dots, Y_{i,n,p})^\top\). Further suppose that \(b_1,b_2>0\) and \(B_n\ge 1\) are constants such that
\(\sum_{i=1}^n \mathbb{E}[Y_{i,n,j}^2] \ge b_1\) for all \(j \in [p]\).
\(\sum_{i=1}^n \mathbb{E}[|Y_{i,n,j}|^{4}] \le b_2 B_n^2/n\) for all \(j \in [p]\).
\(\mathbb{E}[\exp(\sqrt n|Y_{i,n,j}|/B_n)] \le 2\) for all \(i \in [n], j \in [p]\).
Let \(\Sigma_n=\operatorname{Var}(\boldsymbol{S}_n)\) and \(\boldsymbol{Z}_n \sim \mathcal{N}_p(\boldsymbol{0}, \Sigma_n)\). Then there exists a constant \(C_{g}\) only depending on \(b_1\) and \(b_2\) such that \[\sup_{\boldsymbol{x} \in \mathbb{R}^p} \big| \mathbb{P}(\boldsymbol{S}_n \le \boldsymbol{x}) - \mathbb{P}(\boldsymbol{Z}_n \le \boldsymbol{x}) \big| \le C_{g}\Big( \frac{B_n^2 \log^5(pn)}{n} \Big)^{1/4}.\]
Proof. This is Theorem 1 in [43], with their \(X_{i}\) equal to our \(\sqrt n \boldsymbol{Y}_{i,n}\). ◻
Lemma 15. Let \(U\subseteq\mathbb{R}^d\) be an open convex set and \(f:U \to \mathbb{R}\) a convex function. If for some \(\boldsymbol{x}\in U\) all partial derivatives \(\partial_i f(\boldsymbol{x})\) exist, then \(f\) is (totally) differentiable at \(\boldsymbol{x}\).
Proof. Since \(U\) is an open set, there exists an \(\varepsilon>0\) such that \(\mathcal{B}_\varepsilon(\boldsymbol{x}) \subseteq U\). For \(\boldsymbol{h}\in \mathbb{R}^d\) with \(\|\boldsymbol{h}\|\le \varepsilon\), define \(\varphi(\boldsymbol{h}) = f(\boldsymbol{x} + \boldsymbol{h}) - f(\boldsymbol{x}) - \langle \nabla f(\boldsymbol{x}) , \boldsymbol{h} \rangle.\) Convexity of \(f\) implies that \(\varphi\) is convex as well. Denote by \(\boldsymbol{e}_i\) the standard basis vectors of \(\mathbb{R}^d\) so that \(\boldsymbol{h}\in \mathbb{R}^d\) can be written as \(\boldsymbol{h} = h_1 \boldsymbol{e}_1 + \dots + h_d \boldsymbol{e}_d\). Then, \[{\varphi(\boldsymbol{h})} = {\varphi\Big(\frac{1}{d} \sum_{i=1}^d dh_i \boldsymbol{e}_i\Big)} \le \frac{1}{d} \sum_{i=1}^d \varphi(d h_i \boldsymbol{e}_i) \le \frac{1}{d} \sum_{i=1}^d |\varphi(d h_i \boldsymbol{e}_i)|\] and as a result, using \(\|\boldsymbol{h}\| \ge |h_i|\), \[\frac{\varphi(\boldsymbol{h})}{\|\boldsymbol{h}\|} \le \frac{1}{d} \sum_{i=1}^d \frac{|\varphi(d h_i \boldsymbol{e}_i)|}{\|\boldsymbol{h}\|} \le \frac{1}{d} \sum_{i=1}^d \frac{|\varphi(d h_i \boldsymbol{e}_i)|}{|h_i|}.\] Next, \(\varphi(\boldsymbol{0}) = 0\) together with the convexity of \(\varphi\) implies \(0=\varphi(\boldsymbol{h}/2 - \boldsymbol{h}/2)\le (\varphi(\boldsymbol{h}) + \varphi(-\boldsymbol{h}))/2\) and thus \(-\varphi(\boldsymbol{h}) \le \varphi(-\boldsymbol{h})\). It follows that \[-\frac{\varphi(\boldsymbol{h})}{\|\boldsymbol{h}\|} \le \frac{\varphi(-\boldsymbol{h})}{\|-\boldsymbol{h}\|} \le \frac{1}{d} \sum_{i=1}^d \frac{|\varphi(-d h_i \boldsymbol{e}_i)|}{|-h_i|}.\] All that remains to show is that \(|\varphi(d h_i e_i)|/{| dh_i|}\) converges to \(0\) for \(h_i \to 0\), for each \(i\in [d]\). We have \[\frac{|\varphi(d h_i e_i)|}{d|h_i|} =\Big| \frac{f(x + de_i h_i)-f(x) - \partial_i f(x) dh_i}{dh_i}\Big| = \Big| \frac{f(x + de_i h_i)-f(x)}{dh_i} - \partial_i f (x) \Big| \to 0\] for \(h_i \to 0\) by definition of the partial derivatives. ◻
The following result provides finite sample guarantees on the level of a test that is obtained by combining dependent p-values. The set-up is as follows: suppose \(\boldsymbol{Y}_n \in \mathbb{R}^p\) is an observable vector of test statistics (it is instructive to consider each coordinate \(Y_{nq}, q \in [p]\), as a test statistic for which large value provide evidence against some hypothesis \(H_q\)), \(\boldsymbol{Y}_n^{g} \in \mathbb{R}^p\) is some unobservable random vector, and \(\boldsymbol{Y}_n^* \in \mathbb{R}^p\) is some observable bootstrap vector, to be thought of as approximating \(\boldsymbol{Y}_n\). For \(q \in [p]\), let \[\hat{p}_{nq} = 1-F_{nq}^*(Y_{nq}), \quad \hat{p}_{nq}^* = 1-F_{nq}^*(Y_{nq}^*),\] with \(F_{nq}^*\) the conditional cdf of \(Y_{nq}^*\) given the data. Let \[C_n = \min_{q \in [p]} \hat{p}_{nq}, \quad C_n^* = \min_{q \in [p]} \hat{p}_{nq}^*.\] It is instructive to think of small values of \(C_n\) providing evidence against some intersection hypothesis \(H_1 \cap \dots \cap H_p\).
Proposition 20. Let \(\lambda, \lambda^*, \delta >0\), and suppose that \(\boldsymbol{Y}_n^{g}\) has a continuous cdf and that \[\begin{align} \label{eq:gauss-approx-p-values} d_K(\boldsymbol{Y}_n, \boldsymbol{Y}_n^{g}) \le \lambda, \qquad d_K(\mathcal{L}(\boldsymbol{Y}_n^* \mid \mathrm{data}), \boldsymbol{Y}_n^{g}) \le \lambda^* , \end{align}\qquad{(2)}\] the latter holding on a set \(\Omega_n^*\) with \(\mathbb{P}(\Omega_n^*) \ge 1 -\delta\). Then, for every \(\alpha \in (0,1)\), \[\Big| \mathbb{P}(C_n \le \hat{q}_{n,\alpha}^*) - \alpha \Big| \le 2\delta + \lambda + (2p+1)\lambda^* ,\] where \(q_{n,\alpha}^*\) is the \(\alpha\)-quantile of \(\mathbb{P}(C_n^* \in \cdot \mid \mathrm{data})\).
Proof of Proposition 20. Throughout, we denote the cdf of \(\boldsymbol{Y}_n\) by \(F_n\), the cdf of \(\boldsymbol{Y}_n^{g}\) by \(F_n^{g}\) and the conditional cdf of \(\boldsymbol{Y}_n^*\) given the data by \(F_n^*\). Moreover, we define \[C_n^{g} = \min_{q \in [p]} p_{nq}^{g}, \quad \text{ with } \quad p_{nq}^{g} = 1 - F_{nq}^{g}(Y_{nq}^{g}).\]
We start by noting that, if \(F\) and \(G\) are cdfs on the real line satisfying \(\sup_{x \in \mathbb{R}} |F(x) - G(x) | \le \lambda\), then \[\label{eq:marginal-quantiles2} \forall \alpha \in (0,1) : \qquad F^-(\alpha - \lambda) \le G^-(\alpha) \le F^-(\alpha + \lambda);\tag{120}\] here, \(F^{-}(v) = \inf\{ u \in \mathbb{R}: F(u) \ge v\}\) for \(v \in (0,1]\) and \(F^-(v)=-\infty\) for \(v \le 0\) and \(F^-(v)=+\infty\) for \(v>1\); note the slight difference to the generalized inverse used in 31 .
We will show below that \[\begin{align} \tag{121} \sup_{t \in \mathbb{R}} \Big| \mathbb{P}(C_n \le t) - \mathbb{P}(C_n^{g} \le t) \Big| \le \delta + \lambda + p \lambda^*, \\ \tag{122} \text{On } \Omega_n^*: \qquad \sup_{t \in \mathbb{R}} \Big| \mathbb{P}(C_n^* \le t \mid \mathrm{data}) - \mathbb{P}( C_n^{g} \le t ) \Big| \le (p+1) \lambda^*. \end{align}\] Combined with 120 , the bound in 122 then implies that, on \(\Omega_n^*\), \[F_{C_n^{g}}^{-} \big(\alpha - (p+1)\lambda^*\big) \le q_{n,\alpha}^* \le F_{C_n^{g}}^{-} \big(\alpha + (p+1)\lambda^*\big),\] where \(F_{C_n^{g}}\) is the cdf of \(C_n^{g}\). As a consequence, by 121 \[\begin{align} \mathbb{P}(C_n \le q^*_{n,\alpha}) &\le \mathbb{P}\big(C_n \le F_{C_n^{g}}^{-} \big(\alpha + (p+1)\lambda^*\big) \big) + \mathbb{P}\big((\Omega_n^*)^c\big) \\&\le \mathbb{P}\big(C_n^{g} \le F_{C_n^{g}}^{-} \big(\alpha + (p+1)\lambda^*\big) \big) + 2\delta + \lambda + p \lambda^* \le \alpha + 2\delta + \lambda + (2p+1)\lambda^* . \end{align}\] A lower bound can be obtain by similar arguments, and this yields the claim.
It remains to prove 121 and 122 . We start with the former, fix \(t\in \mathbb{R}\) and note that \[\begin{align} \mathbb{P}(C_n \le t) = 1 - \mathbb{P}\Big( \forall q \in [p]: 1-F_{nq}^*(Y_{nq}) > t \Big) &= 1- \mathbb{P}\Big(\forall q \in [p]: Y_{nq} < [F_{nq}^*]^-(1-t) \Big). \end{align}\] Our assumptions imply that, on the event \(\Omega_n^*\), we have \(\max_{p \in [q]}\sup_{x \in \mathbb{R}} |F_{nq}^*(x) - F_{nq}^{g}(x)| \le \lambda^*,\) where \(F_{nq}^{g}\) is the cdf of \(Y_{nq}^{g}\). As a consequence, on the same event and by 120 , \[\begin{align} \label{eq:fn42-fng-quantiles} \forall q \in [p]: \qquad [F_{nq}^*]^-(1-t) \ge [F_{nq}^{g}]^-(1-t-\lambda^*) \end{align}\tag{123}\] Combining the previous inequalities and using that \(\mathbb{P}((\Omega_n^*)^c) \le \delta\), we obtain that \[\begin{align} \mathbb{P}(C_n \le t) &\le 1 - \mathbb{P}\Big(\forall q \in [p]: Y_{nq} < [F_{nq}^{g}]^-(1-t-\lambda^*)\Big) + \mathbb{P}((\Omega_n^*)^c) \\& \le 1- \mathbb{P}\Big(\forall q \in [p]: Y_{nq}^{g} < [F_{nq}^{g}]^-(1-t-\lambda^*)\Big) + \delta + \lambda \\&= 1-F_n^{g}\big( [F_{n1}^{g}]^-(1-t-\lambda^*), \dots, [F_{np}^{g}]^-(1-t-\lambda^*) \big) + \delta + \lambda \\&= 1 - \mathfrak C_n^{g}(1-t-\lambda^*, \dots, 1-t-\lambda^*) + \delta + \lambda \end{align}\] where \(\mathfrak C_n^{g}\) is the copula of \(\boldsymbol{Y}_n^{g}\) (considered as a function on \(\mathbb{R}^p\)) and where we have used ?? at the second inequality. Since copulas are Lipschitz-continuous, we obtain that \[\begin{align} \mathbb{P}(C_n \le t) &\le 1- \mathfrak C_n^{g}(1-t, \dots, 1-t) + \delta + \lambda + p \lambda^* \\&= \mathbb{P}( C_n^{g} \le t ) + \delta + \lambda + p \lambda^*, \end{align}\] where the last equality follows from a straightforward calculation similar to the ones done above. With the same arguments, \[\mathbb{P}(C_n \le t) \ge \mathbb{P}( C_n^{g} \le t ) - \delta - \lambda - p \lambda^*.\] The previous two equations imply 121 .
It remains to prove 122 . By the same calculations as before, and on the event \(\Omega_n^*\) and for each fixed \(t \in \mathbb{R}\), \[\begin{align} \mathbb{P}(C_n^* \le t \mid \mathrm{data}) &= 1 - \mathbb{P}\Big( \forall q \in [p]: Y_{nq}^* < [F_{nq}^*]^-(1-t) \, \Big| \, \mathrm{data}\Big) \\&\le 1 - \mathbb{P}\Big( \forall q \in [p]: Y_{nq}^{g} < [F_{nq}^*]^-(1-t) \, \Big| \, \mathrm{data}\Big) + \lambda^* \\ &= 1 - \mathbb{P}\Big( \forall q \in [p]: Y_{nq}^{g} \le [F_{nq}^*]^-(1-t) \, \Big| \, \mathrm{data}\Big) + \lambda^* \\&\le 1 - \mathbb{P}\Big( \forall q \in [p]: Y_{nq}^{g} \le [F_{nq}^{g}]^-(1-t-\lambda^*) \, \Big| \, \mathrm{data}\Big) + \lambda^* \\&= 1 - \mathfrak C_n^{g}(1-t-\lambda^*, \dots, 1-t-\lambda^*) + \lambda^* \\&\le 1- \mathfrak C_n^{g}(1-t, \dots, 1-t) + (p+1) \lambda^* \\&= \mathbb{P}( C_n^{g} \le t ) + (p+1) \lambda^*, \end{align}\] where we have used 123 at the second inequality. The respective lower bound can be deduced similarly, and 122 thus follows from \(\mathbb{P}(\Omega_n^*) \ge 1- \delta\). ◻
At the same time, rank-based methods are attractive because they avoid modeling marginal tails and can be more efficient than corresponding oracle procedures based on the true marginal distributions [38].↩︎
\(n+1-kx_j \in [ n+1-\lceil kx_j \rceil, n+2-\lceil kx_j \rceil)\) and \(R_{ij} \in \mathbb{N}\), thus \(R_{ij}>n+1-\lceil kx_j \rceil\) implies \(R_{ij} \ge n+2-\lceil kx_j \rceil > n+1-kx_j\), also conversely \(R_{ij} > n+1-kx_j \ge n+1-\lceil kx_j \rceil\)↩︎
In [61], the constant in front of the max-sup is \(2^d\), but it can be replaced by \(2\). Indeed, note that if \(\boldsymbol{x}, \boldsymbol{y} \in [\boldsymbol{a}, \boldsymbol{b}]\) with \(\|\boldsymbol{x} - \boldsymbol{y}\|_\infty \leq \varepsilon\) then there must exists rectangles \(R, \tilde{R} \in \mathcal{A}\) with a non-empty intersection such that \(\boldsymbol{x} \in R, \boldsymbol{y} \in \tilde{R}\). Since each rectangle has diameter \(\varepsilon\) with respect to the sup norm, the claim follows from the triangle inequality.↩︎