Linear Functional Testing with General Loadings in Sparse Regression: Separation Rates and Computational Barriers


Abstract

We study the problem of testing \(H_0: \xi^\top\beta=t_0\) in high-dimensional sparse linear regression with Gaussian random design and unknown design covariance. The loading vector \(\xi\) is arbitrary, and the exact sparsity level \(k\) is unknown but bounded by a known value \(k_u\). Tests are required to control Type I error uniformly over the \(k_u\)-sparse null, while power is evaluated against \(k\)-sparse alternatives. We construct a computationally efficient mixed test that gives an upper bound on the adaptive separation distance and establish an information-theoretic lower bound calibrated to the magnitude profile of \(\xi\). In the ultra-sparse regime \(k_u\lesssim \sqrt n/\log p\), these bounds characterize the adaptive separation rate up to logarithmic factors for arbitrary \(\xi\). In the moderately sparse regime \(\sqrt n/\log p\ll k_u\lesssim n/\log p\), these bounds match for several classes of loading vectors but may differ in general. In this regime, we further prove a low-degree lower bound that matches the upper bound up to logarithmic factors. This provides evidence that improving on the rate of the mixed test, if statistically possible, may be computationally hard. For flat sparse loadings, we complement this evidence with a polynomial-time reduction from sparse CCA. Finally, we examine how information about the design covariance affects the adaptive separation rate in two settings. Under a sparse signed-spiked covariance model, the information-theoretic lower bound is attainable up to logarithmic factors by a computationally inefficient procedure, while the low-degree lower bound and sparse-CCA reduction continue to apply, providing evidence for a statistical-computational gap. When the design covariance is known and diagonal, the adaptive separation rate takes the same form as in the ultra-sparse regime.

1 Introduction↩︎

High-dimensional regression is commonly studied in the regime where the number of covariates \(p\) exceeds the sample size \(n\). Consider the linear regression model \[\label{eq:32linear32regression} Y = X\beta + \varepsilon, \qquad \varepsilon \sim {\mathcal{N}}(0, \sigma^2 \mathbf{I}_n),\tag{1}\] where \(Y \in \mathbb{R}^n\), \(X \in \mathbb{R}^{n \times p}\), and \(\beta \in \mathbb{R}^p\). Under suitable conditions on \(X\), regularized estimators such as the Lasso [1] can attain the minimax-optimal rate \(k\log p/n\) for the squared estimation error over \(k\)-sparse vectors with \(k \le c n/\log p\) for some constant \(c>0\); see, for example, [2], [3].

Beyond estimation, uncertainty quantification and hypothesis testing are also important tasks in high-dimensional regression. In this paper, we study testing problems of the form \[\label{eq:32hypothesis32test} H_0: \xi^\top \beta = t_0,\tag{2}\] where \(t_0 \in \mathbb{R}\) is fixed and \(\xi \in \mathbb{R}^p\) is a loading vector. Through test inversion, this testing problem is related to confidence interval construction for the linear functional \(L(\beta)=\xi^\top \beta\). This formulation includes coordinate-wise inference as the special case where \(\xi\) is a standard basis vector, and also covers prediction-type targets, where \(\xi\) represents a test-point covariate vector and \(\xi^\top\beta\) is the corresponding conditional mean.

For coordinate-wise inference, Zhang and Zhang [4], van de Geer et al.[5], and Javanmard and Montanari [6] develop debiasing Lasso methods that yield asymptotically valid confidence intervals. However, these procedures are primarily designed for the ultra-sparse regime \(k \lesssim \sqrt{n}/\log p\), which is substantially more restrictive than the condition \(k \lesssim n/\log p\) required for consistent estimation.

Inference for general linear functionals is more delicate, and existing optimality theory is largely tied to structured loading vectors. Cai and Guo [7] study confidence intervals for \(\xi^\top\beta\) under the regular loading condition \[\label{eq:32tony32loading32vector} \frac{\max_{i \in {\rm supp}(\xi)} |\xi_i|}{\min_{i \in {\rm supp}(\xi)} |\xi_i|} \le \bar{c}, \qquad \text{for some constant } \bar{c} \ge 1,\tag{3}\] and distinguish two settings: the sparse-loading setting where \(\|\xi\|_0\lesssim k\) and the dense-loading setting where \(k \asymp p^{\gamma}\), \(\|\xi\|_0 \asymp p^{\gamma_\xi}\), and \(\gamma_\xi>2\gamma\). For both settings, they derive the minimax expected length of confidence intervals for \(\xi^\top \beta\). Cai, Cai and Guo [8] develop an inference procedure that is valid for arbitrary \(\xi\), but their lower-bound theory is established for loadings satisfying 3 , and for a specific class of polynomially decaying loadings.

We study the problem 2 in an adaptive testing framework, where the true sparsity level \(k\) is unknown and only an upper bound \(k_u\) is available. This question lies at the core of adaptive inference, a central theme in high-dimensional statistics [7][10]. Following the adaptive testing formulation of [8], we require the test to control Type I error uniformly over the enlarged parameter space with sparsity \(k_u\), while its power is evaluated over alternatives with the true sparsity level \(k\le k_u\). An adaptive separation distance quantifies the smallest signal size needed to distinguish the null from the local alternative when the test is only allowed to use \(k_u\).

Within this adaptive framework, our goal is to characterize the separation distance for general loading vectors. As discussed above, existing optimality theory is largely tied to structured loading classes. In particular, prior results mainly cover loadings that satisfy the regular loading condition 3 and whose support size falls into either the sparse-loading regime or the dense-loading regime. The intermediate regime between these two cases remains open even under regular loading. More generally, existing results do not describe the adaptive separation distance for heterogeneous loading vectors. Such heterogeneity is natural in prediction problems, where \(\xi\) may be a test-point covariate vector. For instance, if the coordinates of \(\xi\) are i.i.d.Gaussian, then the nonzero coordinates need not be of the same order, so \(\xi\) violates the regular loading condition 3 with high probability and does not belong to the polynomial-decay classes. This motivates the central question studied in this paper:

1.1 Major Contributions↩︎

This paper develops a loading-dependent theory to characterize the adaptive separation distance for testing a linear functional in high-dimensional sparse regression with Gaussian design. Notably, we identify a potential computational barrier that arises in the moderately sparse regime \(\sqrt{n}/\log p \ll k_u \lesssim n/\log p\) when the design covariance matrix \(\Sigma\) is unknown.

Our theory has three parts: adaptive separation distances for general loadings in the ultra-sparse regime, computational evidence for the gaps in the moderately sparse regime, and refined analysis to identify the source of the gaps. We next describe the main results in more detail.

Adaptive separation distances for general loading vectors in the ultra-sparse regime. In the ultra-sparse regime \(k_u \lesssim \sqrt{n}/\log p\), we characterize the adaptive separation rate, up to logarithmic factors, for testing 2 with an arbitrary loading vector \(\xi \in \mathbb{R}^p\). This extends prior work [7], [8], which mainly treats loading vectors satisfying the regular loading condition 3 or a polynomial-decay condition. Our upper bound is obtained by decomposing \(\xi\) into large and small coordinates and mixing debiased and plug-in tests; the lower bound uses least favorable priors calibrated to the magnitude profile of \(\xi\). Extending into the moderately sparse regime, the upper and lower bounds continue to match for some classes of loading vectors, but may differ in general.

Evidence for computational barriers in the moderately sparse regime. In the moderately sparse regime, we prove a low-degree lower bound showing that no degree-\(D\) polynomial weakly separates the null and alternative at and below a certain separation level. For \(D=O(\log p)\), the low-degree lower bound matches the computationally feasible upper bound up to logarithmic factors. Under the standard low-degree heuristic [11], [12], this result provides evidence for a computational barrier. In addition, for flat sparse loadings, we derive an explicit polynomial-time reduction to demonstrate that the testing problem 2 is at least as hard as sparse canonical correlation analysis (CCA). This connection is non-obvious: testing a linear functional in sparse linear regression has no apparent structural resemblance to sparse CCA or sparse PCA, and the existing literature on linear functional inference gives little indication that such a connection should arise. Since sparse CCA is widely believed to exhibit computational barriers [13], [14], our reduction provides complementary evidence for the low-degree barrier.

To the best of our knowledge, this is the first work to provide low-degree lower-bound evidence and sparse-CCA reduction evidence for computational barriers in testing linear functionals in high-dimensional sparse linear regression. Existing computational lower bounds for sparse regression mainly concern recovery, estimation, or detection of the sparse signal itself [15], [16]; by contrast, our problem concerns inference for a linear functional, where the difficulty arises from the interaction between the unknown covariance structure and sparsity level.

Refined analysis with knowledge of covariance. We calibrate the scope of the lower-bound analysis through two benchmarks with additional information about the design covariance. When the design covariance is known and diagonal, the upper bound is refined to match the lower bound up to logarithmic factors. When \(\Sigma\) remains unknown but satisfies a sparse signed-spiked assumption, the lower bound is again sharp up to logarithmic factors, but now the attaining optimal test is computationally inefficient. In this setting, the low-degree lower bound and sparse-CCA reduction continue to apply, thereby providing evidence for a statistical–computational gap.

1.2 Related literature↩︎

We review several lines of research related to our setting.

Minimax linear functional estimation. The minimax theory for estimating linear functionals has been extensively developed under various statistical models; see, for example, [17][21]. The work closest to ours is [21], which establishes minimax rates for estimating a general linear functional \(\xi^\top\mu\), with an arbitrary loading vector \(\xi\), in the Gaussian sequence model with a sparse mean vector \(\mu\). Our analysis differs from theirs in three respects: the observations come from a random-design regression model, the quantity of interest is an adaptive testing boundary rather than an estimation risk, and the unknown design covariance plays an important role in the moderately sparse regime.

Linear hypothesis testing in high-dimensional linear models. Beyond the debiased Lasso approach discussed earlier, linear hypothesis testing in high-dimensional regression has also been studied in [10], [22][24]. Javanmard and Lee [22] reduce testing a general linear functional to testing its projections onto a selected orthogonal basis, and then construct debiased estimators for the corresponding basis coefficients. Their approach, however, relies critically on sparsity of the loading vector \(\xi\), and therefore does not extend to general loadings. Zhao et al. [24] propose an estimator that adapts to the sparsity of \(\beta\) and to the strength of correlations among predictors, but their objectives and assumptions differ substantially from those considered here.

Zhu and Bradic [23] construct confidence intervals for linear functionals under the assumption that \[\mathbb{E}\!\left[\xi^\top X_i \,\middle|\, \omega_1^\top X_i,\ldots,\omega_{p-1}^\top X_i\right]\] is a sparse linear combination of \(\omega_1^\top X_i,\ldots,\omega_{p-1}^\top X_i\), where \(\{\omega_j\}_{1\le j\le p-1}\) is an orthonormal basis of the subspace orthogonal to \(\xi\). When this conditional sparsity level is sufficiently small, their procedure attains the parametric rate \(n^{-1/2}\) for estimating the linear functional. Bradic et al. [10] develop related ideas for inference on a single coefficient \(\beta_1\), assuming that \(X_1\mid X_{-1}\) admits a sparse linear representation in terms of \(X_{-1}\), and establish corresponding minimax rates. Although these structural conditions lead to elegant inferential procedures, they need not hold in our setting, where \(\xi\) is arbitrary and \(\Sigma\) is only assumed to have bounded spectrum.

Computational barriers in high-dimensional problems. Several lines of work address computational barriers in high-dimensional statistical problems. One common strategy is to establish polynomial-time reductions from conjecturally hard problems to the statistical task of interest. Such reductions have been used to provide evidence of computational barriers in a variety of high-dimensional statistical problems, including sparse PCA [25], submatrix detection [26], and sparse CCA [13], typically under the assumption that certain instances of the planted clique problem cannot be solved by randomized polynomial-time algorithms [27], [28]. However, constructing these reductions is often delicate: the resulting transformations can be fragile, may require additional structural assumptions, and can limit the scope of the resulting hardness conclusions.

A complementary line of work establishes lower bounds for broad classes of algorithms, including sum-of-squares algorithms [29][31], statistical query algorithms [28], and low-degree polynomial algorithms [11], [12]. Although these classes do not capture all polynomial-time algorithms, they are sufficiently expressive to model many practical procedures and often provide strong evidence of computational hardness. Among them, the low-degree polynomial framework has proven particularly powerful, as the behavior of such algorithms can be analyzed directly and sharply. This framework has been successfully applied to a wide range of high-dimensional testing problems, including sparse PCA [11], [12], tensor PCA [11], sparse CCA [14], graphon estimation [32], and independent component analysis [33]. In this work, our computational results rely on these two complementary forms of evidence: a direct low-degree lower bound for general loading vectors and a polynomial-time reduction from sparse CCA for flat sparse loading vectors.

1.3 Organization of the paper↩︎

The remainder of the paper is organized as follows. 2 introduces the framework for analyzing the minimax properties of linear hypothesis testing and provides a preview of the main results. 3 develops the general upper and lower bounds for the adaptive separation distance under arbitrary loading vectors and identifies conditions under which these bounds match. 4 provides evidence for computational barriers in the moderately sparse regime \(\sqrt{n}/\log p \ll k_u \lesssim n/\log p\). 5 examines two cases with prior knowledge about design covariance, including sparse signed-spiked covariance and known design covariance. 6 concludes with further discussions and open directions.

1.4 Notation↩︎

For a positive integer \(p\), let \([p]=\{1,\ldots,p\}\). For a set \(S\), denote its cardinality by \(|S|\). For \(x,y\in\mathbb{R}\), let \(x\vee y=\max\{x,y\}\), \(x\wedge y=\min\{x,y\}\), and \(x_+=x\vee 0\). For \(x\in\mathbb{R}\), let \(\lfloor x\rfloor\) denote the greatest integer less than or equal to \(x\). For a vector \(x\in\mathbb{R}^p\) and a subset \(J\subset[p]\), \(x_J\) denotes the subvector of \(x\) indexed by \(J\), and \(x_{-J}\) denotes the subvector indexed by \(J^c\). We write \({\rm supp}(x)\) for the support of \(x\). For \(q>0\), define \(\|x\|_q=\left(\sum_{i=1}^p |x_i|^q\right)^{1/q}\). We also use the conventions \(\|x\|_0=|{\rm supp}(x)|\) and \(\|x\|_\infty=\max_{1\le j\le p}|x_j|\). Let \(e_i\) denote the \(i\)th standard basis vector in \(\mathbb{R}^p\). For a matrix \(X\in\mathbb{R}^{n\times p}\), \(X_{i\cdot}\), \(X_{\cdot j}\), and \(X_{ij}\) denote, respectively, its \(i\)th row, \(j\)th column, and \((i,j)\) entry. The notation \(X_{i,-j}\) denotes the \(i\)th row of \(X\) with the \(j\)th coordinate removed, and \(X_{-j}\) denotes the submatrix of \(X\) obtained by removing its \(j\)th column. For \(J\subset[p]\), \(X_{\cdot J}\) denotes the submatrix of \(X\) formed by the columns indexed by \(J\). For a symmetric matrix \(A\), \(\lambda_{\min}(A)\) and \(\lambda_{\max}(A)\) denote its smallest and largest eigenvalues, respectively. We use \(c\) and \(C\) to denote generic positive constants whose values may vary from line to line. For two positive sequences \(a_n\) and \(b_n\), we write \(a_n\lesssim b_n\) if there exists a constant \(C>0\) such that \(a_n\le Cb_n\) for all \(n\). We write \(a_n\gtrsim b_n\) if \(b_n\lesssim a_n\), and \(a_n\asymp b_n\) if both \(a_n\lesssim b_n\) and \(b_n\lesssim a_n\). We write \(a_n\ll b_n\) if \(a_n/b_n\to 0\), and \(a_n\gg b_n\) if \(b_n\ll a_n\). The notation \(a_n\asymp_{\log} b_n\) means that \(a_n\lesssim b_n\) and \(b_n\lesssim a_n\) hold up to logarithmic factors in \(p\). For a measure \(\pi\), we denote by \(\pi^{\otimes 2}\) the product measure \(\pi\otimes\pi\).

2 Preliminaries and Preview↩︎

In this section, we introduce the framework for linear hypothesis testing in high-dimensional linear regression and state several technical preliminaries used in the subsequent analysis. We also provide an informal summary of the main results in 2.4, which serves as a theorem map for the rest of the paper.

2.1 Problem setup↩︎

Throughout the article, we focus on the high-dimensional linear model 1 with random design, where the rows of \(X\) satisfy \(X_{i\cdot}\stackrel{\text{i.i.d.}}{\sim} \mathcal{N}(0,\Sigma),i=1,\cdots,n\), and are independent of \(\varepsilon\). Both \(\Sigma\) and the noise level \(\sigma\) are treated as unknown. The observed data are \({\cal Z}=\{Z_1,\cdots,Z_n\}\), where \(Z_i=(Y_i,X_{i \cdot})\in \mathbb{R}^{p+1}\) for \(i=1,\cdots,n\). The distribution of the data is now indexed by the parameter \[\theta=(\beta,\Sigma,\sigma),\] which consists of the signal \(\beta\), the covariance matrix \(\Sigma=\mathbb{E}[X_{i \cdot}X_{i \cdot}^\top]\) for the random design, and the noise level \(\sigma\). We consider the following collection of parameter spaces: \[\label{eq:32parameter32space} \begin{align} \Theta(k)=&\left\{\theta=(\beta,\Sigma,\sigma):\left\|\beta\right\|_0\le k,\frac{1}{M_1}\le \lambda_{\min}(\Sigma)\le \lambda_{\max}(\Sigma)\le M_1, 0<\sigma\le M_2\right\}, \end{align}\tag{4}\] where \(M_1>1\) and \(M_2>0\) are some positive constants. The eigenvalue bound and noise-variance bound are two mild regularity conditions [7], [9], [34]. We use \(\mathbb{P}_\theta^n\) and \(\mathbb{E}_\theta\) to denote the probability and expectation with respect to the distribution of the data \({\cal Z}\) indexed by \(\theta\).

2.2 Optimality framework for hypothesis testing: minimaxity and adaptivity↩︎

For \(0<\alpha<1\) and a given parameter space \(\Theta\), the class of tests of nominal level \(\alpha\) for testing the null hypothesis \(\theta\in\Theta\) is defined as \[\label{eq:32null32test} \Psi_{\alpha}(\Theta) = \left\{ \psi : \mathcal{Z} \mapsto [0,1] :\; \sup_{\theta \in \Theta} \mathbb{E}_{\theta} \psi \le \alpha \right\},\tag{5}\] see, for example, Lehmann and Romano [35]. Here, we allow for both randomized and non-randomized tests. We consider the linear hypothesis testing problem 2 . The corresponding null parameter space is given by \[\label{eq:32null32space} \Theta(k;\xi,t_0) = \left\{ \theta=(\beta,\Sigma,\sigma)\in \Theta(k) :\; \xi^\top \beta = t_0 \right\}.\tag{6}\] To characterize local alternatives, we define \[\label{eq:32alternative32space} \Theta_{\pm\tau}(k;\xi,t_0) = \left\{ \theta=(\beta,\Sigma,\sigma)\in \Theta(k) :\; \lvert \xi^\top \beta - t_0 \rvert \ge \tau \right\},\tag{7}\] where \(\tau>0\) is a given separation level. Larger values of \(\tau\) correspond to alternatives that are farther from the null and are therefore easier to distinguish.

For fixed \(0<\alpha,\eta<1\), we define the minimax separation distance by \[\label{eq:32minimax32separation} \tau_{\text{mini}}(k\,;\,\xi) = \inf\left\{ \tau :\; \sup_{\psi\in\Psi_\alpha(\Theta(k;\xi,t_0))} \inf_{\theta\in\Theta_{\pm\tau}(k;\xi,t_0)} \mathbb{E}_\theta \psi \ge 1-\eta \right\}.\tag{8}\] Thus, \(\tau_{\mathrm{mini}}(k\,;\,\xi)\) is the smallest separation level at which there exists a test with Type I error at most \(\alpha\) uniformly over \(\Theta(k;\xi,t_0)\) and power at least \(1-\eta\) uniformly over \(\Theta_{\pm\tau}(k;\xi,t_0)\). The separation distance depends on \(\xi,k,t_0,\alpha\) and \(\eta\). Throughout the paper, we suppress the dependence on \(t_0\), \(\alpha\), and \(\eta\) in the notation \(\tau_{\text{mini}}(k\,;\,\xi)\), as these quantities do not affect the rate of the separation distance. The constants \(\alpha\) and \(\eta\) can be chosen arbitrarily small.

A test is said to be valid if its Type I error over the null parameter space does not exceed \(\alpha\), and powerful if its power over the local alternative parameter space is at least \(1-\eta\). We say that a test \(\psi\) is minimax optimal if \[\label{eq:32optimal} \psi \in \Psi_\alpha(\Theta(k;\xi,t_0)) \quad\text{and}\quad \inf_{\theta\in\Theta_{\pm\tau}(k;\xi,t_0)} \mathbb{E}_\theta \psi \ge 1-\eta \quad\text{for}\quad \tau \asymp \tau_{\text{mini}}(k\,;\,\xi).\tag{9}\]

The minimax separation distance in 8 is defined for a fixed sparsity level \(k\). To formalize adaptivity, we follow [8] and consider two sparsity levels \(k\le k_u\), where \(k\) is the unknown true sparsity level and \(k_u\) is a known upper bound. When prior information about sparsity is limited, \(k_u\) may be substantially larger than \(k\). The test must therefore control its Type I error uniformly over the larger null parameter space \(\Theta(k_u;\xi,t_0)\), while its power is evaluated over alternatives with sparsity level \(k\). Accordingly, we define the adaptive separation distance by \[\label{eq:32adaptive32separation} \tau_{\text{adap}}(k_u, k; \xi) = \inf \left\{ \tau : \sup_{\psi\in \Psi_\alpha(\Theta(k_u;\xi,t_0))} \inf_{\theta\in \Theta_{\pm \tau}(k;\xi, t_0)} \mathbb{E}_\theta \psi\ge 1 - \eta \right\}.\tag{10}\] This work focuses on the asymptotic regime in which \(n\), \(p\), and \(k_u\) diverge. For simplicity, we use the notation \(\lim\), \(\asymp\), \(\to\), \(o(\cdot)\), and \(O(\cdot)\) to denote limits and asymptotic relations.

Comparing 10 with 8 , we observe that the lack of precise knowledge about the sparsity level affects only the size control over a larger parameter space, while the power functions in 8 and 10 are evaluated over the same parameter space. It is evident that \(\tau_{\text{mini}}(k\,;\,\xi) = \tau_{\text{adap}}(k,k; \xi)\). Analogous to 9 , a test \(\psi\) is defined as adaptively optimal if it satisfies: \[\label{eq:32adaptively32optimal} \psi\in \Psi_\alpha(\Theta(k_u;\xi,t_0))\quad \text{and}\quad \inf_{\theta\in \Theta_{\pm \tau}(k;\xi, t_0)}\mathbb{E}_\theta \psi \ge 1 - \eta \quad \text{for}\quad \tau \asymp \tau_{\text{adap}}(k_u, k; \xi).\tag{11}\]

The quantities \(\tau_{\text{mini}}(k\,;\,\xi)\) and \(\tau_{\text{adap}}(k_u,k\,;\,\xi)\) do not depend on a specific testing procedure but instead reflect the intrinsic difficulty of the testing problem 2 , which is determined by the parameter space and the loading vector \(\xi\).

We can tell whether one pays a statistical price for not knowing \(k\) by comparing \(\tau_{\text{mini}}(k\,;\,\xi)\) and \(\tau_{\text{adap}}(k_u,k\,;\,\xi)\). We focus on the nontrivial case \(k\ll k_u\). When \(k\asymp k_u\), the upper bound \(k_u\) already localizes the sparsity level sufficiently well, and one typically has \(\tau_{\text{mini}}(k;\xi)\asymp \tau_{\text{adap}}(k_u,k;\xi)\). If \[\tau_{\text{mini}}(k\,;\,\xi) \asymp \tau_{\text{adap}}(k_u,k\,;\,\xi),\] then we say the hypothesis testing problem 2 is adaptive to sparsity; that is, even without knowing the exact sparsity level, it is possible to construct a test that achieves the same separation rate as if the sparsity level were known. In contrast, if \[\tau_{\text{mini}}(k\,;\,\xi) \ll \tau_{\text{adap}}(k_u,k\,;\,\xi),\] then the hypothesis testing problem 2 is non-adaptive, indicating that information about the sparsity level is essential. In this case, the adaptive separation distance \(\tau_{\text{adap}}(k_u,k\,;\,\xi)\) is of interest, as it quantifies the best achievable performance in the absence of precise sparsity information.

As pointed out in [8], the minimax detection boundary quantifies the intrinsic difficulty of the testing problem when the sparsity level is known, whereas the adaptive separation distance addresses a more challenging setting where the sparsity level is unknown. In practice, adaptively optimal tests that satisfy condition 11 are more useful than minimax optimal tests, as the exact sparsity level \(k\) is typically unknown in real applications. Furthermore, the definitions immediately yield the identity \(\tau_{\text{mini}}(k\,;\,\xi)=\tau_{\text{adap}}(k, k; \xi)\), so it suffices to focus on characterizing the adaptive separation distance.

In addition, we consider the setting in which the covariance matrix is known and fixed as \(\Sigma = \Sigma_0\) for some \(\Sigma_0\in \mathbb{R}^{p \times p}\) satisfying the eigenvalue condition in 4 . In this case, we specify the parameter space as \[\Theta(k,\Sigma_0) = \left\{ \theta = (\beta,\Sigma_0,\sigma) : \theta \in \Theta(k) \right\}.\] Analogous to 6 , 7 , 8 , and 10 , we define \[\begin{align} \Theta(k,\Sigma_0;\xi,t_0) &= \left\{ \theta = (\beta,\Sigma_0,\sigma) \in \Theta(k;\xi,t_0) \right\},\nonumber \\ \Theta_{\pm \tau}(k,\Sigma_0;\xi, t_0) &= \left\{ \theta = (\beta,\Sigma_0,\sigma) \in \Theta_{\pm \tau}(k;\xi, t_0) \right\}, \nonumber \\ \tau_{\text{mini}}(k,\Sigma_0\,;\,\xi) &= \inf \left\{ \tau :\; \sup_{\psi \in \Psi_\alpha(\Theta(k,\Sigma_0;\xi,t_0))} \inf_{\theta \in \Theta_{\pm \tau}(k,\Sigma_0;\xi, t_0)} \mathbb{E}_\theta \psi \ge 1 - \eta \right\}, \tag{12} \\ \tau_{\text{adap}}(k_u,k,\Sigma_0\,;\,\xi) &= \inf \left\{ \tau :\; \sup_{\psi \in \Psi_\alpha(\Theta(k_u,\Sigma_0;\xi,t_0))} \inf_{\theta \in \Theta_{\pm \tau}(k,\Sigma_0;\xi, t_0)} \mathbb{E}_\theta \psi \ge 1 - \eta \right\}. \tag{13} \end{align}\]

2.3 Estimation and sparsity conditions↩︎

Before turning to the inference problem, we first state several conditions used throughout this article.

To develop the inference procedures, we assume the existence of estimators \(\hat{\beta}\) and \(\hat{\sigma}^2\) satisfying the following conditions.

Condition 1. With probability approaching 1, the estimator \(\hat{\beta}\) satisfies \[\label{eq:32linear32estimator} \|\hat{\beta}-\beta\|_1 \le c_\beta \sigma k_u \sqrt{\frac{\log p}{n}}, \quad \text{and} \quad \|\hat{\beta}-\beta\|_2 \le C_\beta \sigma\sqrt{\frac{k_u \log p}{n}},\tag{14}\] for some constants \(c_\beta, C_\beta > 0\).

Condition 2. \(\hat{\sigma}^2\) is a consistent estimator of \(\sigma^2\), i.e., \(|\hat{\sigma}^2/\sigma^2-1|\stackrel{p}{\to}0\).

[cdt: linear estimator,cdt: linear variance] are used in constructing upper bounds on adaptive separation distances. Various polynomial-time regularized estimators satisfy Conditions 1 and 2; one example is the scaled Lasso, defined by \[\{\hat{\beta},\hat{\sigma}\} = \mathop{\mathrm{arg\,min}}_{\beta \in \mathbb{R}^p,\, \sigma \in \mathbb{R}^+} \left\{ \frac{\|Y - X\beta\|_2^2}{2n\sigma} + \frac{\sigma}{2} + \sqrt{\frac{2.01 \log p}{n}} \sum_{j=1}^p \frac{\|X_{\cdot j}\|_2}{\sqrt{n}} |\beta_j| \right\},\] which Sun and Zhang [36] showed satisfies these conditions. A key assumption required for general regularized estimators to meet these conditions is the restricted eigenvalue condition on the design matrix \(X\), originally introduced by [2]: \[\label{eq:32restricted32eigenvalue} \kappa(X, k_u, \alpha_0) = \min_{\substack{J_0 \subseteq \{1, \ldots, p\} \\ |J_0| \le k_u}} \;\min_{\substack{\boldsymbol{\delta}\neq 0 \\ \|\boldsymbol{\delta}_{J_0^c}\|_1 \le \alpha_0 \|\boldsymbol{\delta}_{J_0}\|_1}} \frac{\|X \boldsymbol{\delta}\|_2}{\sqrt{n} \|\boldsymbol{\delta}_{J_0}\|_2}.\tag{15}\]

In the classical analysis of sparsity-inducing estimators [2], the constants \(c_\beta\) and \(C_\beta\) in 14 typically depend on the restricted eigenvalue condition; in particular, the error bound deteriorates as \(\kappa(X,k_u,\alpha_0)\) decreases. Furthermore, directly verifying the restricted eigenvalue condition is computationally difficult [37], [38].

In our analysis, the conditions can be handled probabilistically. Specifically, for Gaussian random designs whose covariance matrix \(\Sigma\) satisfies the eigenvalue condition in 4 , we can show that, for each fixed \(\alpha_0>0\), there exists a constant \(c_{\rm RE}>0\), depending only on \(M_1\) and \(\alpha_0\), such that if \(k_u\log p/n \le c_{\rm RE}\), then \(\kappa(X,k_u,\alpha_0)\) is bounded away from zero with high probability. This high-probability lower bound can then be used to choose admissible values of \(c_\beta\) and \(C_\beta\). A more detailed discussion is deferred to 11.2.

Remark 1. Although the estimation error in 14 is adaptive to the unknown exact sparsity level \(k\), Cai and Guo [9] showed that the accuracy assessment of the \(\ell_q\)-loss, for \(1 \le q \le 2\), is hard and generally non-adaptive. Consequently, we bound the estimation loss using the upper bound in 14 , stated in terms of the sparsity upper bound \(k_u\) rather than the unknown exact sparsity level \(k\). While such a bound may seem non-tight, we will later show that it still leads to an optimal test, thereby justifying the use of the \(k_u\)-based error bound.

Condition 3. There exists a constant \(\gamma\in[0,1/2)\) such that \(k_u \lesssim p^\gamma.\)

Condition 3 is standard and it places \(k_u\) in the sparsity regime commonly assumed when deriving minimax rates in high-dimensional settings [7], [9], [10].

2.4 Informal theorem map↩︎

We now give an informal summary of the main results. Throughout the paper, we assume without loss of generality that the loading vector \(\xi=(\xi_1,\ldots,\xi_p)^\top\in\mathbb{R}^p\) is ordered by decreasing magnitude, \(|\xi_1|\ge |\xi_2|\ge \cdots \ge |\xi_p|.\) We also write \(k_\xi=\|\xi\|_0\) for the support size of \(\xi\). For \(t>0\), define the top-\(\lceil t\rceil\) norm of \(\xi\) by \[\label{eq:32H32definition} H(t;\xi) := \sqrt{\sum_{j\le \lceil t\rceil \wedge p}\xi_j^2} = \max_{\substack{A\subseteq[p]\\ |A|\le \lceil t\rceil}} \|\xi_A\|_2,\tag{16}\] and \(H(0;\xi)=0\). The main rate statements are summarized in 1.

Table 1: Informal summary of the main separation rates. All comparisons areunderstood up to logarithmic factors.
Setting Informal rate statement
\(k_u \lesssim \sqrt n/\log p\) \(\displaystyle \tau_{\text{adap}}(k_u,k;\xi) \asymp_{\log} \frac{H(k_u^2\log p;\xi)}{\sqrt n}\)
\(\sqrt n/\log p \ll k_u \lesssim n/\log p\) \(\displaystyle \tau_{\text{com}}(k_u,k;\xi) \asymp_{\log} H({n}/{\log p};\xi)\frac{k_u\log p}{n}\)
All sparsity levels: \(k_u\lesssim n/\log p\) \(\displaystyle \tau^{\text{spike}}_{\text{adap}}(k_u,k;\xi) \asymp_{\log} \frac{H(k_u^2\log p;\xi)}{\sqrt n} + H(k_u;\xi)\frac{k_u\log p}{n}\)
All sparsity levels: \(k_u\lesssim n/\log p\) \(\displaystyle \tau_{\text{adap}}(k_u,k,\Sigma_0^{\textrm{diag}};\xi) \asymp_{\log} \frac{H(k_u^2\log p;\xi)}{\sqrt n}\)

The entries in Table 1 should be read as follows. The first, third, and fourth rows give adaptive separation rates matched by statistical lower bounds up to logarithmic factors. The second row concerns the moderately sparse unknown-covariance regime and should instead be interpreted as a computational separation distance: it is achieved by computationally feasible tests and is supported by computational lower-bound evidence.

The sparse signed-spiked row shows that, under additional covariance structure, the adaptive separation distance can be characterized up to logarithmic factors. This rate is no larger than the computational separation distance in the general unknown-covariance setting. Indeed, since \(k_u\lesssim n/\log p\), we have \(H(k_u;\xi)\lesssim H(n/\log p;\xi)\). Moreover, in the moderately sparse regime \(\sqrt n/\log p\ll k_u\), we have \(n/\log p\lesssim k_u^2\log p\). By the definition of \(H(t;\xi)\) and the decreasing ordering of \(|\xi_j|\), this implies \[\frac{H(k_u^2\log p;\xi)}{\sqrt{k_u^2\log p}} \lesssim \frac{H(n/\log p;\xi)}{\sqrt{n/\log p}} .\] Consequently, the sparse signed-spiked rate is no larger, up to logarithmic factors, than the computational separation rate in the moderately sparse regime \(\sqrt{n}/\log p\ll k_u\lesssim n/\log p\). However, the attaining procedure relies on computationally inefficient covariance estimation, as discussed in 5.1. Finally, when the covariance matrix \(\Sigma_0\) is known and diagonal, the separation distance can be further reduced. This is consistent with the intuition that, even under sparse signed-spiked structure, the covariance matrix must still be estimated, whereas the known-diagonal benchmark removes covariance uncertainty entirely.

We next relate these results to the known-sparsity minimax rate \(\tau_{\text{mini}}(k;\xi)\) and discuss adaptivity for testing 2 . Consider the ultra-sparse unknown-covariance setting, corresponding to the first row of Table 1. In this setting, \[\tau_{\text{mini}}(k;\xi) \asymp_{\log} \frac{H(k^2\log p;\xi)}{\sqrt n}, \qquad \tau_{\text{adap}}(k_u,k;\xi) \asymp_{\log} \frac{H(k_u^2\log p;\xi)}{\sqrt n}.\] Thus, testing 2 is adaptive, in the sense that unknown \(k\) incurs no additional cost up to logarithmic factors, whenever, for \(k\ll k_u\), \[H(k_u^2\log p;\xi) \asymp_{\log} H(k^2\log p;\xi).\] If \(\xi\) satisfies the regular loading condition 3 , this condition reduces, up to logarithmic factors, to \(k_\xi=\|\xi\|_0 \lesssim k^2\). In this case, the minimax separation rate simplifies to \(\|\xi\|_2/\sqrt n\), consistent with the adaptivity results in [7], [8].

3 Adaptive separation distance↩︎

In this section, we investigate the adaptive separation distance for the linear hypothesis testing problem.

Since \(\xi\) is nonzero, we have \(k_\xi=\|\xi\|_0\ge 1\). Let \(\zeta \in \mathbb{R}\) be the solution to the equation \[\label{eq:32equation32sol} \frac{\sum_{j=1}^{k_\xi} |\xi_j| \exp(-\zeta / \xi_j^2)}{\sqrt{\sum_{j=1}^{k_\xi} \xi_j^2 \exp(-\zeta / \xi_j^2)}} = \frac{k_u}{2}, \qquad \text{and set } \lambda = \sqrt{\zeta_+}.\tag{17}\] 22 guarantees that 17 admits a unique solution, since the left-hand side is continuous, strictly decreasing, and tends to \(+\infty\) as \(\zeta \to -\infty\) and to \(0\) as \(\zeta \to +\infty\).

The first key quantity in our analysis is \[\label{eq:32nu132definition} \nu_1=\nu_1(k_u;\xi) = \lambda k_u + \sqrt{\sum_{j=1}^{k_\xi} \xi_j^2\exp\{-\lambda^2/\xi_j^2\}} .\tag{18}\] The second key quantity is \[\label{eq:32nu232definition} \nu_2=\nu_2(k_u;\xi) = H(k_u;\xi),\tag{19}\] where \(H(t;\xi)\) is defined in 16 .

3.1 General upper and lower bounds↩︎

For arbitrary loading vectors, 1 below gives general upper and lower bounds for the adaptive separation distance of testing 2 .

Theorem 1. Under 3, there exists some constant \(c>0\) such that if \(n\ge c\,k_u\log p\), then for any \(1\le k\le k_u\), the following statements hold:

  1. Upper Bound: Let \(\xi_{p+1}=0\), then \[\label{eq:32upper32bound} \tau_{\text{adap}}(k_u, k; \xi) \lesssim \min_{0\le m\le p}\left( H(m\,;\,\xi) \left(\frac{1}{\sqrt{n}}+\frac{k_u\log p}{n}\right) +|\xi_{m+1}|k_u\sqrt{\frac{\log p}{n}}\right).\tag{20}\]

  2. Lower Bound: \[\label{eq:32lower32bound} \tau_{\text{adap}}(k_u, k; \xi) \gtrsim \nu_1\frac{1}{\sqrt{n}}\vee \nu_2\frac{k_u\log p}{n}.\tag{21}\]

We discuss the implications of these bounds and identify conditions where they match. The following proposition rewrites the upper bound in an equivalent and more interpretable form.

Proposition 1. We write \(\xi_{p+1}=0\). For any \(t\ge 1\), it holds that \[\min_{0\le m\le p} \left\{ H(m\,;\,\xi) + |\xi_{m+1}|\,t \right\} \asymp H(t^2\,;\,\xi).\] Consequently, the upper bound in 20 is equivalent to \[\label{eq:32upper32bound32equivalent} \tau_{\text{adap}}(k_u,k\,;\,\xi) \;\lesssim\; \begin{cases} \displaystyle \frac{1}{\sqrt n} H(k_u^2\log p\,;\,\xi), & \displaystyle k_u \lesssim \frac{\sqrt n}{\log p}, \\[1.2em] \displaystyle \frac{k_u\log p}{n} H(n/\log p\,;\,\xi), & \displaystyle k_u \gg \frac{\sqrt n}{\log p}. \end{cases}\qquad{(1)}\]

1 separates the upper bound into two sparsity regimes: the ultra-sparse regime \(k_u \lesssim \sqrt n/\log p\) and the moderately sparse regime \(\sqrt n/\log p \ll k_u \lesssim n/\log p\). We next discuss the corresponding lower bounds in these two regimes.

Ultra-sparse regime. In this regime, the quantity \(\nu_1\) in the lower bound in 21 plays an important role. It has a more explicit characterization that relates to the upper bound.

Proposition 2. Let \(j_1 = \max\left\{j\in[p]: |\xi_j|\ge \lambda\right\}\) with the convention that \(j_1=0\) if the set is empty. Then, for \(\nu_1\) defined in 18 , \[\nu_1 \asymp H(j_1\,;\,\xi) + \lambda k_u .\]

In the ultra-sparse regime \(k_u \lesssim \sqrt n/\log p\), the upper bound in 20 can be bounded by the choice of \(m=j_1\) so that \[\begin{align} \tau_{\text{adap}}(k_u,k\,;\,\xi) & \;\lesssim\; \frac{1}{\sqrt n} H(j_1\,;\,\xi) + \lambda k_u\sqrt{\frac{\log p}{n}} \;\lesssim\; \nu_1\sqrt{\frac{\log p}{n}} . \end{align}\] Therefore, the term \(\nu_1/\sqrt{n}\) in the lower bound in 21 for the adaptive separation distance is sharp up to logarithmic factors, as summarized in the following corollary.

Corollary 1. Assume the conditions in 1 hold. In the ultra-sparse regime \(k_u \lesssim \sqrt n/\log p\), it holds for all \(1\le k\le k_u\) that \[\tau_{\text{adap}}(k_u,k\,;\,\xi)\asymp_{\log} \frac{\nu_1}{\sqrt n} \asymp_{\log} \frac{H(k_u^2\log p\,;\,\xi)}{\sqrt n} .\] If further \(k_u^2 \log p \lesssim j_1\), it holds that \[\tau_{\text{adap}}(k_u,k\,;\,\xi)\asymp \frac{\nu_1}{\sqrt n},\] that is, the logarithmic gap disappears.

The closest existing result is [8], which gives a lower bound for the adaptive separation distance in the testing problem 2 . However, even in the ultra-sparse regime \(k_u \lesssim \sqrt n/\log p\), they only establish the sharpness of their lower bound for a few specific loading classes considered there. By contrast, the lower bound based on \(\nu_1\) applies to arbitrary loading vectors and is tight up to logarithmic factors throughout the ultra-sparse regime.

Moderately sparse regime. In this regime, \(\nu_1/\sqrt{n}\) can still be sharp in a subclass of loading vectors while the quantity \(\nu_2\) also becomes relevant and could be the dominant term when \(\xi\) is light-tailed.

Corollary 2. Assume the conditions in 1 hold. Suppose that \(k_u \gg \sqrt {n}/{\log p}\) and that the loading vector \(\xi\) satisfies that \[\label{eq:sharp-cond-nu1-moderatesparse} \nu_1 \gtrsim \frac{k_u\log p}{\sqrt{n}} H( n/\log p \,;\,\xi).\tag{22}\] Then for all \(1\le k\le k_u\), \[\tau_{\text{adap}}(k_u,k\,;\,\xi)\asymp \frac{\nu_1}{\sqrt n}.\]

For regular loading vectors, Appendix F shows that the dense side \(k_\xi\gtrsim k_u^2\) gives \[\tau_{\text{adap}}(k_u,k\,;\,\xi) \asymp_{\log} \|\xi\|_\infty k_u\sqrt{\frac{\log p}{n}} .\] If, more strongly, \(k_\xi/k_u^2\ge p^c\) for some constant \(c>0\), then 22 holds and the logarithmic equivalence can be strengthened to \[\tau_{\text{adap}}(k_u,k\,;\,\xi) \asymp \|\xi\|_\infty k_u\sqrt{\frac{\log p}{n}} .\]

The quantity \(\nu_2\) in the lower bound of 21 is relevant in the moderately sparse regime. As revealed by the proof, this term arises from the unknown covariance among covariates. Furthermore, 5 shows that when the design covariance is known and diagonal, the adaptive separation distance reduces to \(\nu_1/\sqrt{n}\) and \(\nu_2\) no longer plays a role.

3 characterizes an important subclass of loading vectors for which the adaptive separation distance is determined by the \(\nu_2\) quantity.

Corollary 3. Assume the conditions in 1 hold. Suppose that \(k_u \gg \sqrt {n}/{\log p}\) and that the loading vector \(\xi\) satisfies that \[\label{eq:sharp-cond-nu2-moderatesparse} H( k_u \,;\,\xi) \asymp H( n/\log p \,;\,\xi).\tag{23}\] Then, for any \(1\le k\le k_u\), \[\tau_{\text{adap}}(k_u,k\,;\,\xi) \asymp \frac{k_u\log p}{n}\nu_2.\]

The condition in 23 means that the top-\(k_u\) subvector of \(\xi\) captures nearly as much energy as the enlarged top-\((n/\log p)\) counterpart. This condition holds, for example, when the loading vector \(\xi\) is sufficiently sparse or when its coordinates decay sufficiently fast. This includes the sparse-loading case \(k_\xi\lesssim k_u\) considered by [7]. For such loading vectors, the lower bound involving \(\nu_2\) is sharp in the moderately sparse regime.

To clarify the scope of our bounds, Appendix F in the supplementary material provides details on several loading-profile examples. Except for the regular loading example in Appendix F.1, existing theory in [7], [8] does not apply the other examples. Dense nonregular profiles in Appendix F.2 show that the upper and lower bounds in 1 can match even when \(\xi\) is not regular, whereas the multiscale example in Appendix F.3 illustrates that a gap can exist in the moderately sparse regime. Appendix F.4 considers loadings with random coordinates that have sub-Weibull tails, including Gaussian and exponential distributions.

For general loading vectors, there may be a gap between the information-theoretic lower bound 21 and the computationally feasible upper bound 20 in the moderately sparse regime \(k_u \gg {\sqrt n}/{\log p}\). 4 examines this issue from a computational perspective.

3.2 Phase diagram for regular loading vectors↩︎

To provide a concrete interpretation of our results, we consider the case in which the loading vector \(\xi \in \mathbb{R}^p\) satisfies 3 for some constant \(\bar{c}>0\) with sparsity level \(k_\xi = \|\xi\|_0\). We compare the resulting adaptive separation distances with existing results in [7], [8].

Figure 1: Phase diagram for the loading vector in 3 with sparsity level k_\xi. We parameterize k_u = p^{\gamma_u}, k_\xi = p^{\gamma_\xi}, n = p^{\gamma_n}, and the rescaled alternative shift \sqrt{n}\tau=\|\xi\|_\infty p^{\gamma_\tau}. Left panel (a): Ultra-sparse regime, k_u\lesssim \sqrt{n}/\log p. Right panel (b): Moderately sparse regime, \sqrt{n}/\log p \ll k_u \lesssim n/\log p.

Figure 1 presents the upper and lower bounds for adaptive separation distance for testing \(H_0: \xi^\top \beta = t_0\) versus \(H_1: |\xi^\top \beta - t_0| = \tau\) under various \(k_\xi\). We parametrize the problem as \[k_\xi = p^{\gamma_\xi}, \qquad k_u = p^{\gamma_u}, \qquad n = p^{\gamma_n},\text{ and }\tau \asymp\|\xi\|_\infty \, p^{\gamma_\tau - \gamma_n/2}.\] The parametrization of \(\tau\) can be interpreted as a rescaled version of the alternative shift, where the exponent \(\gamma_\tau\) characterizes the signal strength after normalizing for the common loading magnitude \(\|\xi\|_\infty\) and the classical parametric rate \(n^{-1/2}\).

We first explain the left panel, corresponding to the ultra-sparse regime \(k_u \lesssim \sqrt{n}/\log p\). In this regime, the statistical boundary is sharp and completely characterizes the phase transition between distinguishability and indistinguishability. Specifically, the separation distance is given by \[\label{eq:32separation32under32regular32and32ultrasparse} \tau_{\text{adap}}(k_u,k\,;\,\xi) \asymp \begin{cases} \|\xi\|_2/\sqrt{n}, & \gamma_\xi \le 2\gamma_u,\\[0.4em] \|\xi\|_\infty k_u \sqrt{\log p/n}, & \gamma_\xi > 2\gamma_u; \end{cases}\tag{24}\] at the boundary \(\gamma_{\xi}=2 \gamma_u\), the expression should be read up to logarithmic factors.

Notably, when \(\gamma_\xi < 2\gamma_u\), \(\tau_{\text{adap}}(k_u,k\,;\,\xi)\) does not depend on \(k_u\), indicating that full adaptivity to the unknown exact sparsity is achievable. This is in agreement with the existing results by [8].

The right panel depicts the moderately sparse regime \(\sqrt{n}/\log p \ll k_u \lesssim n/\log p\), where the geometry becomes more intricate. In this regime, the phase diagram should be interpreted through two curves: an information-theoretic lower-bound curve, below which detection is statistically impossible, and a computationally feasible upper-bound curve, above which there exist computationally efficient procedures that succeed. When the two curves coincide, their common value gives the adaptive separation distance. When they differ, the region between them is not resolved by the statistical bounds alone. Accordingly, there are three different regimes:

  • When the loading sparsity satisfies \(\gamma_\xi \le \gamma_u\) or \(\gamma_\xi \ge 2\gamma_u\), the two curves coincide, and the adaptive separation distance can be explicitly characterized as \[\label{eq:32separation32under32regular32and32moderatesparse} \tau_{\text{adap}}(k_u,k\,;\,\xi) \asymp \begin{cases} \|\xi\|_2 \, k_u \log p / n, & \gamma_\xi \le \gamma_u,\\[0.4em] \|\xi\|_\infty k_u \sqrt{\log p/n}, & \gamma_\xi \ge 2\gamma_u. \end{cases}\tag{25}\] In these loading settings, the boundary is achievable by computationally efficient tests.

    Comparing with 24 , the only difference lies in the sparse-loading regime \(\gamma_\xi\le \gamma_u\), where the rate is increased by a factor of \(k_u\log p /\sqrt{n} \gg 1\). This is driven by the \(\nu_2\) quantity resulting from the unknown design covariance.

    Accordingly, the statistical boundary now has a nonzero intercept at \((0,\gamma_u-\gamma_n/2)\). This intercept corresponds to a separation distance of order \(k_u\log p/n\) when the loading vector \(\xi\) has only a constant number of nonzero entries. Thus, even for very sparse \(\xi\), the parametric rate \(n^{-1/2}\) is no longer attainable, and adaptivity with respect to \(k_u\) necessarily fails.

  • When the loading sparsity satisfies \(\gamma_u < \gamma_\xi < 2\gamma_u\), the available information-theoretic lower bound and the computationally feasible upper bound do not match. More precisely, \[\label{eq:32gap32region} \|\xi\|_\infty \left[\frac{k_u^{3/2}\log p}{n}+\sqrt{\frac{k_\xi}{n}}\right] \;\lesssim\; \tau_{\text{adap}}(k_u,k\,;\,\xi) \;\lesssim\; \|\xi\|_\infty \sqrt{\frac{n}{\log p}\wedge k_\xi}\, \frac{k_u\log p}{n}.\tag{26}\] The middle shaded region in the right panel corresponds to this non-negligible gap between these two curves.

The characterization summarized in 1 is consistent with the results of [7], while the intermediate loading-sparsity setting on the right panel remains unresolved in the existing literature. In 4, we show that the upper curve is matched by a low-degree lower bound up to logarithmic factors. Accordingly, under the low-degree heuristic, the shaded region is interpreted as a computationally hard region.

3.3 Computationally feasible upper bound via loading decomposition↩︎

The upper bound in 1 is proven by inverting confidence intervals with desired lengths. Suppose that \[{\rm CI}_{\alpha_1}(Z) = [\hat{L}(Z\,;\,\xi)-r(Z\,;\,\xi),\,\hat{L}(Z\,;\,\xi)+r(Z\,;\,\xi)]\] is a \((1-\alpha_1)\)-level confidence interval for \(L(\beta\,;\,\xi)=\xi^\top\beta\), and that \(r(Z\,;\,\xi)\le \tilde{r}(\xi)\) with probability at least \(1-\alpha_2\). Then the test \[\label{eq:32inverted-ci-test} \psi(Z) = \mathbf{1}\{|t_0-\hat{L}(Z\,;\,\xi)|>\tilde{r}(\xi)\}\tag{27}\] belongs to \(\Psi_{\alpha_1+\alpha_2}(\Theta(k_u;\xi,t_0))\) and has power at least \(1-(\alpha_1+\alpha_2)\) over \(\Theta_{\pm 2\tilde{r}(\xi)}(k;\xi,t_0)\).

Taking \(\alpha'=\min\{\alpha,\eta\}\), it is sufficient to construct a \((1-\alpha'/2)\)-level confidence interval for \(L(\beta\,;\,\xi)=\xi^\top\beta\) whose radius is bounded by \(\tilde{r}(\xi)\) with probability at least \(1-\alpha'/2\). The resulting inverted test has Type I error at most \(\alpha\) and power at least \(1-\eta\) over alternatives separated from the null by \(2\tilde{r}(\xi)\).

To obtain a confidence interval that is effectively tailored to the particular loading profile of any \(\xi\), we propose a mixed construction that generalizes two existing approaches:

  • Plug-in confidence interval. By 3 and the discussion following [cdt: linear estimator,cdt: linear variance], we can construct an interval centered at \(\hat{L}_{\mathrm{pi}}(Z\,;\,\xi)=\xi^\top\hat{\beta}\) with radius bounded with probability tending to 1 by \[\label{eq:32plug32in32length} \tilde{r}_{\mathrm{pi}}(\xi) \asymp \sigma \|\xi\|_\infty k_u\sqrt{\frac{\log p}{n}} .\tag{28}\]

  • Debiased confidence interval. The debiased interval is centered at \[\label{eq:32debiased32center} \hat{L}_{\mathrm{db}}(Z\,;\,\xi) = \xi^\top\hat{\beta} + \hat{u}^\top \frac{1}{n}X^\top(Y-X\hat{\beta}),\tag{29}\] where \[\label{eq:32projection32vector} \hat{u} = \mathop{\mathrm{arg\,min}}_{u\in\mathbb{R}^p} \left\{ u^\top\hat{\Sigma} u: \|\hat{\Sigma} u-\xi\|_\infty \le C_\xi\|\xi\|_2\sqrt{\frac{\log p}{n}} \right\},\tag{30}\] \(\hat{\Sigma}=n^{-1}X^\top X\), and \(C_\xi>0\) is sufficiently large. If the feasible set in 30 is empty, we set \(\hat{u}=0\); by 7, this event has vanishing probability for some large enough \(C_\xi\). The radius of the debiased interval admits the high-probability deterministic bound \[\label{eq:32debiased32length} \tilde{r}_{\mathrm{db}}(\xi) \asymp \sigma \|\xi\|_2 \left( \frac{1}{\sqrt n} + k_u\frac{\log p}{n} \right).\tag{31}\]

Comparing 28 and 31 , the plug-in confidence interval is effective when \(\|\xi\|_\infty\) is small, whereas the debiased confidence interval is effective when \(\|\xi\|_2\) is small. To combine these complementary advantages, we decompose the loading vector according to coordinate magnitude. For any \(0\le m\le p\), write \(\xi=\xi^{(1)}+\xi^{(2)}\), where \[\xi^{(1)} = (\xi_1,\ldots,\xi_m,0,\ldots,0)^\top, \qquad \xi^{(2)} = (0,\ldots,0,\xi_{m+1},\ldots,\xi_p)^\top .\] At the level of \(1-\alpha'/4\), we construct the debiased confidence interval for \((\xi^{(1)})^\top\beta\), which corresponds to the \(m\) largest coordinates of \(\xi\) in absolute value, and the plug-in confidence interval for \((\xi^{(2)})^\top\beta\), which corresponds to the remaining coordinates.

We call the Minkowski sum of these two intervals the mixed confidence interval, which has level \(1-\alpha'/2\), and the corresponding inverted test the mixed test.

By [eq: plug in length,eq: debiased length], the length of the mixed confidence interval is at the scale of \[\label{eq:upper-length-mixed} \sigma\sqrt{\sum_{j\le m}\xi_j^2} \left( \frac{1}{\sqrt n} + k_u\frac{\log p}{n} \right) + \sigma |\xi_{m+1}|k_u\sqrt{\frac{\log p}{n}}.\tag{32}\]

The choices \(m=0\) and \(m=p\) recover the plug-in confidence interval and the debiased confidence interval, respectively. Optimizing 32 over \(m\) yields the upper bound in 1. Moreover, 1 shows that the rate-optimal cutoff can be taken as \[\label{eq:32optimal32cutoff} m_* = \begin{cases} \lceil k_u^2\log p\rceil ~ \wedge ~p , & k_u\lesssim \sqrt n/\log p ,\\ \lceil n/\log p\rceil ~ \wedge ~p , & k_u\gg \sqrt n/\log p. \end{cases}\tag{33}\]

3.4 Proof idea of lower bound↩︎

The lower bound in 1 is obtained from two distinct least-favorable prior constructions. Both constructions are governed by the magnitude profile of \(\xi\), but they exploit different components of the model to turn this profile into a testing difficulty.

  • The first construction yields the term \(\nu_1/\sqrt{n}\). This construction uses the full magnitude profile of \(\xi\) to calibrate the local separation that remains hard even when the design covariance is fixed at \(\Sigma=\mathbf{I}_p\).

  • The second construction yields the term \(\nu_2 k_u\log p/n\). Unlike the first term, this term exploits the unknown-covariance component of the model, because the construction couples the leading coordinates of \(\xi\) with a sparse perturbation of \(\Sigma\).

Both constructions are analyzed through Le Cam’s method [39]. Since Type I error must be controlled over the enlarged null space \(\Theta(k_u;\xi,t_0)\), it is natural to use a mixture-over-null versus point-alternative comparison; see, for example, [8]. Concretely, we construct a prior \(\pi_1\) supported on \(\Theta(k_u;\xi,t_0)\) and compare the induced mixture distribution \(\mathbb{P}^n_{\pi_1} = \int \mathbb{P}^n_\theta\,\pi_1({\rm d}\theta)\) with the distribution \(\mathbb{P}^n_{\theta_\star}\) generated by a fixed alternative point \(\theta_\star \in \Theta_{\pm\tau}(k;\xi,t_0)\). If their total variation \({\rm TV}\bigl(\mathbb{P}^n_{\pi_1},\mathbb{P}^n_{\theta_\star}\bigr)\) is small, then every test with uniform Type I error control over \(\Theta(k_u;\xi,t_0)\) has limited power at \(\theta_\star\).

The main technical step is the construction of \(\pi_1\) for arbitrary loading vectors. The prior must simultaneously satisfy the sparsity constraint, the exact null constraint, and the eigenvalue restrictions on \(\Sigma\), while keeping \({\rm TV}\bigl(\mathbb{P}^n_{\pi_1},\mathbb{P}^n_{\theta_\star}\bigr)\) small.

Below we outline the key ideas behind our two constructions. For exposition, we describe the constructions after a deterministic translation that cancels out \(t_0\). Accordingly, the fixed alternative has \(\beta_\star=0\) and the null prior is supported on the space with \(L(\beta\,;\,\xi)=\tau\). This translation preserves the relevant total variation and \(\chi^2\)-divergence, so the argument applies to a general \(t_0\).

Lower bound via the quantity \(\nu_1\). To capture the term involving \(\nu_1\), we use a random-sparsity prior, following the ideas of [21], [40]. Under such a prior, the coordinates are generated independently: the \(j\)th coordinate is assigned the value \(\gamma_j\) with probability \(q_j\), and is set to zero with probability \(1-q_j\). The sequences \(\{q_j\}\) and \(\{\gamma_j\}\) are calibrated according to the magnitude profile of the loading vector \(\xi\).

This calibration is essential for general loading vectors. The classical least favorable prior construction used in [7], [8] selects a support uniformly and assigns equal magnitudes on the selected coordinates. It is sharp for sufficiently homogeneous loading vectors, such as those satisfying 3 , but can be suboptimal for heterogeneous \(\xi\). The random-sparsity prior instead adapts \(\{q_j\}\) and \(\{\gamma_j\}\) to the magnitude profile of \(\xi\), which yields the lower-bound term involving \(\nu_1\).

A further adjustment is needed because the vanilla random-sparsity construction induces a random value of the target functional \(L(\beta\,;\,\xi)\). To obtain a valid null prior that satisfies \(L(\beta\,;\,\xi)=\tau\), we modify the construction through a scalar parameter to ensure that every draw satisfies the linear constraint exactly, while preserving the statistical closeness between the induced null and alternative distributions. This adjustment step is specific to the current testing problem and does not follow directly from the constructions in [21], [40].

Lower bound associated with \(\nu_2\). To obtain the term involving \(\nu_2\), we exploit the uncertainty in the design covariance to construct random \(\Sigma\) in the null prior. The idea is to balance two goals: (1) align the directions of \(\beta\) and \(\xi\) to produce a large shift in \(L(\beta\,;\,\xi)\); (2) adjust \(\Sigma\) accordingly so that the induced mixture distribution of \((Y,X)\) stays close to the alternative distribution with \(\Sigma=\mathbf{I}_p\) and \(\beta=\mathbf{0}\).

Let \(p_1=\lfloor k_u/4\rfloor\). The null prior constructs a block-structured covariance matrix \[\label{eq:32covariance32construction} \Sigma = \begin{pmatrix} \mathbf{I}_{p_1\times p_1} & \boldsymbol{\delta}_1\boldsymbol{\delta}_2^\top \\ \boldsymbol{\delta}_2\boldsymbol{\delta}_1^\top & \mathbf{I}_{(p-p_1)\times(p-p_1)} \end{pmatrix},\tag{34}\] where the vector \(\boldsymbol{\delta}_1\in\mathbb{R}^{p_1}\) is deterministic with entries \[(\boldsymbol{\delta}_1)_j = -\frac{\xi_j}{\sqrt{\sum_{i=1}^{p_1}\xi_i^2}}, \qquad j\in[p_1],\] and \(\boldsymbol{\delta}_2\in\mathbb{R}^{p-p_1}\) is sampled with a uniform random support \(S\) of size \(p_1\) and nonzero entries \[(\boldsymbol{\delta}_2)_j \asymp \operatorname{sign}(\xi_{p_1+j})\sqrt{\frac{\log p}{n}}, \qquad j\in S.\] The prior also chooses \(\beta=\Sigma^{-1}(0, \kappa \boldsymbol{\delta}_2^\top)^\top\) where \(\kappa=\kappa(b)\in[0,1]\) is chosen to enforce the translated null constraint \(L(\beta \,;\,\xi)=\tau\). The leading term of \(L(\beta \,;\,\xi)\) is of order \[- \kappa \xi_{[p_1]}^\top \boldsymbol{\delta}_1\,\|\boldsymbol{\delta}_2\|_2^2 = \kappa H(p_1;\xi)\|\boldsymbol{\delta}_2\|_2^2,\] whereas \(\xi_{p_1+S}^\top \beta_{S}\) is nonnegative. Since \(p_1\asymp k_u\), we have \(H(p_1;\xi)\asymp H(k_u;\xi)\) by the monotone ordering of the coordinates. Furthermore, since \(\|\boldsymbol{\delta}_2\|_2^2\asymp k_u\log p/n\), the null constraint can be satisfied if \(\tau\) is at the scale of \[H(k_u;\xi)\frac{k_u\log p}{n} = \nu_2(k_u;\xi)\frac{k_u\log p}{n}.\] Meanwhile, the random support of \(\boldsymbol{\delta}_2\) keeps the mixture statistically close to the fixed alternative, as the resulting \(\chi^2\)-divergence is controlled by a hypergeometric overlap bound.

The particular forms of \(\Sigma\) and \(\beta\) are crucial for this tractable analysis. Furthermore, the control on the \(\chi^2\)-divergence reveals an interesting connection under this prior construction: the resulting functional testing comparison is as hard as detecting the sparse rank-two covariance perturbation, because the \(\chi^2\)-divergence in the former problem can be bounded by the one associated with the latter problem. This connection also foreshadows our computational-hardness argument in 4, where a closely related sparse covariance perturbation is used to connect the testing problem with computationally hard sparse covariance detection problems.

4 Computational Barriers in the Moderately Sparse Regime↩︎

In this section, we study computational barriers for testing 2 in the moderately sparse regime \({\sqrt n}/{\log p} \ll k_u \lesssim {n}/{\log p}\). The discussion following 3 shows that, in this regime, the upper and lower bounds in Theorem 1 need not coincide for general loading vectors. Our results provide two forms of computational lower-bound evidence that match the upper bound: a direct low-degree lower bound for general loadings and a reduction from sparse CCA for flat sparse loadings.

4.1 Frameworks for computational-barrier evidence↩︎

We introduce two frameworks for establishing evidence for computational barriers. The first is the low-degree polynomial method, which yields lower bounds against low-degree polynomial algorithms. The second is polynomial-time reduction, which transfers hardness from a conjecturally hard source problem to the testing problem 2 .

Low-degree polynomial framework. The low-degree framework has been successfully applied to a wide range of high-dimensional testing problems, including sparse PCA, tensor PCA, sparse CCA, graphon estimation, and independent component analysis; see [11], [12], [14], [32], [33] and references therein.

Let \(\mathbb{R}[\mathcal{Z}]_{\le D}\) denote the space of multivariate polynomials in the observed data \(\mathcal{Z}\) of degree at most \(D\). Following [41], we use the following notion of weak separation.

Definition 1. We say that \(f\in\mathbb{R}[\mathcal{Z}]_{\le D}\) weakly separates two distributions \(\mathbb{Q}_1\) and \(\mathbb{Q}_2\) if, as \(n\to\infty\), \[\sqrt{ \max\left\{ {\rm Var}_{\mathbb{Q}_1}(f(\mathcal{Z})), {\rm Var}_{\mathbb{Q}_2}(f(\mathcal{Z})) \right\} } = O\left( \left| \mathbb{E}_{\mathbb{Q}_1}f(\mathcal{Z}) - \mathbb{E}_{\mathbb{Q}_2}f(\mathcal{Z}) \right| \right).\]

Accordingly, no degree-\(D\) polynomial weakly separates two parameter spaces \(\Theta_1\) and \(\Theta_2\) if there exist priors \(\pi_1\) and \(\pi_2\), supported on \(\Theta_1\) and \(\Theta_2\), respectively, such that no \(f\in\mathbb{R}[\mathcal{Z}]_{\le D}\) weakly separates the induced mixture distributions \[\mathbb{P}_{\pi_1}^n = \int \mathbb{P}_\theta^n\,\pi_1({\rm d}\theta), \qquad \mathbb{P}_{\pi_2}^n = \int \mathbb{P}_\theta^n\,\pi_2({\rm d}\theta).\]

The following standard criterion reduces low-degree hardness to bounding the low-degree likelihood-ratio norm.

Proposition 3 ([41]). Let \(\mathbb{Q}_1\) and \(\mathbb{Q}_2\) be probability measures with \(\mathbb{Q}_1\ll \mathbb{Q}_2\). Define \[L=\frac{{\rm d}\mathbb{Q}_1}{{\rm d}\mathbb{Q}_2}, \qquad \mathrm{LD}(D)=\|L^{\le D}\|_{L^2(\mathbb{Q}_2)}^2,\] where \(L^{\le D}\) is the orthogonal projection of \(L\) onto \(\mathbb{R}[\mathcal{Z}]_{\le D}\) in \(L^2(\mathbb{Q}_2)\). If \[\mathrm{LD}(D)=1+o(1),\] then no degree-\(D\) polynomial weakly separates \(\mathbb{Q}_1\) and \(\mathbb{Q}_2\).

When \(D=\infty\), we have \(\mathrm{LD}(\infty) = \mathbb{E}_{\mathbb{Q}_2}[L^2] = 1+\chi^2(\mathbb{Q}_1\|\mathbb{Q}_2)\). Thus, \(\mathrm{LD}(D)=1+o(1)\) is a computational analogue of the usual \(\chi^2\)-based information-theoretic indistinguishability condition. In particular, it may hold even when the full \(\chi^2\)-divergence diverges. It has been conjectured that, for many average-case high-dimensional testing problems, low-degree polynomials of degree \(D=O(\log p)\) capture the limits of polynomial-time computation [11], [12]. Although this conjecture is not universal [42], [43], low-degree lower bounds provide strong evidence of computational hardness.

Reduction from sparse CCA. The second approach is based on polynomial-time reductions. In this section, we use sparse canonical correlation analysis as the source problem, since sparse CCA is widely conjectured to exhibit intrinsic computational barriers [13], [14], [44]. The role of the reduction is to show that, for a special class of loading vectors, an efficient algorithm for 2 would imply an efficient algorithm for an appropriately parameterized sparse CCA detection problem. The precise sparse CCA model and the reduction are given in 4.3.

4.2 Low-degree lower bound↩︎

We first establish a direct low-degree lower bound for general loading vectors. The construction is motivated by the information-theoretic lower bound associated with \(\nu_2\), where the difficulty comes from uncertainty in the design covariance matrix \(\Sigma\). Here, the same covariance-perturbation mechanism in 34 is calibrated to the polynomial degree \(D\), leading to the following low-degree lower bound.

Theorem 2. Suppose 3 holds, \({\sqrt n}/{\log p} \ll k_u \lesssim {n}/{\log p}\), and \(D\lesssim p\). Let \[k_{\mathrm{eff}} = \left\lfloor \frac{n}{\log p} \wedge \frac{k_u^2}{D\log p} \right\rfloor, \qquad \nu_3 = H(k_{\mathrm{eff}};\xi).\] Then for any \(1\le k\leq k_u\), no degree-\(D\) polynomial weakly separates \(\Theta(k_u;\xi,t_0)\) and \(\Theta_{\pm\tau}(k;\xi,t_0)\) when \[\tau = c\,\nu_3\,\frac{k_u\log p}{n},\] where \(c>0\) is a sufficiently small constant.

Theorem 2 can be viewed as the computational analogue of the lower bound associated with \(\nu_2\). Taking \(D=O(\log p)\), we have \[k_{\mathrm{eff}} \asymp_{\log} \frac{n}{\log p} \wedge k_u^2.\] In the moderately sparse regime \(k_u\gg \sqrt n/\log p\), the lower bound in Theorem 2 matches the computationally feasible upper bound in ?? up to logarithmic factors. Moreover, when \(k_u\gtrsim \sqrt{n\log p}\), the lower bound in Theorem 2 matches the upper bound in Theorem 1 based on 1.

To prove 2, we construct a hidden covariance perturbation analogously to the one in 34 for proving the information-theoretic lower bound based on \(\nu_2\). We briefly outline the argument for bounding \(\mathrm{LD}(D)\), with full details deferred to 8.3. Let \(\pi_2\) be a point mass at a carefully chosen alternative \(\theta^*\), and let \(\pi_1\) be a prior supported on the null space with the covariance perturbation. For \(L_\theta = {{\rm d}\mathbb{P}_\theta^n}/{{\rm d}\mathbb{P}_{\theta^*}^n}\), linearity of the low-degree projection gives \[\mathrm{LD}(D) = \left\| \left(\mathbb{E}_{\theta\sim\pi_1}L_\theta\right)^{\le D} \right\|_{L^2(\mathbb{P}_{\theta^*}^n)}^2 = \mathbb{E}_{(\theta_1,\theta_2)\sim\pi_1^{\otimes2}} \mathbb{E}_{\mathbb{P}_{\theta^*}^n} \left[ L_{\theta_1}^{\le D} L_{\theta_2}^{\le D} \right].\]

The key step is to decompose the last expectation according to a high-probability event \(A\) under \(\pi_1^{\otimes2}\). On \(A\), the two hidden covariance perturbations have a tractable alignment structure, which allows us to extend the proof of Proposition 3.6 in [41] and to show that \[\mathbb{E} \left[ L_{\theta_1}^{\le D} L_{\theta_2}^{\le D} \mathbf{1}_A \right] \le \mathbb{E} \left[ L_{\theta_1} L_{\theta_2} \mathbf{1}_A \right];\] see 16. On \(A^c\), we combine a uniform bound on \(\|L_\theta^{\le D}\|_{L^2(\mathbb{P}_{\theta^*}^n)}\) with the small probability of \(A^c\). These two bounds imply \[\mathrm{LD}(D)=1+o(1),\] which proves Theorem 2 by Proposition 3.

This argument is not a standard low-degree calculation. The null prior must satisfy both the sparsity constraint and the scalar constraint \(\xi^\top\beta=t_0\). Moreover, the proof establishes \(\mathrm{LD}(D)=1+o(1)\) even though the full \(\chi^2\)-divergence may diverge. The event decomposition isolates the covariance-perturbation pairs that contribute to the low-degree likelihood-ratio norm and is the key step behind the weak-separation lower bound.

4.3 Reduction from sparse CCA↩︎

We next give a complementary hardness result through a polynomial-time reduction. This result applies to a special class of loading vectors and gives additional evidence that the low-degree lower bound in 4.2 reflects a genuine computational limitation.

We restrict attention to loading vectors \(\xi\) with sparsity \(k_\xi=\|\xi\|_0\) and unit nonzero entries. After relabeling coordinates, we may take \(\xi_j=1\) for \(1\le j\le k_\xi\) and \(\xi_j=0\) for \(j>k_\xi\). Thus, 2 reduces to \[\label{eq:32falt32testing32problem} H_0:\;\sum_{j=1}^{k_\xi}\beta_j=t_0.\tag{35}\] This restriction is natural for the reduction because sparse CCA hardness is typically formulated for equal-magnitude sparse vectors [13], [14]. It is also the setting where the remaining gap between the statistical lower and upper bounds is most transparent.

As shown in 3.2, the adaptive separation distance is already characterized outside the intermediate regime \(\sqrt{n}/\log p\ll k_u \ll k_\xi \ll k_u^2.\) For the equal-magnitude loading vectors considered here, the computational rate suggested by the feasible upper bound is \[\label{eq:32computational32upper32tony} H(\frac{n}{\log p};\xi)\frac{k_u\log p}{n}=\sqrt{\frac{n}{\log p}\wedge k_\xi} \frac{k_u\log p}{n},\tag{36}\] which is increasing in \(k_\xi\) until it saturates at \(k_\xi\asymp n/\log p\). Here we focus on the range of \(k_\xi\) before the saturation: \[\sqrt n /\log p\ll k_u \ll k_\xi \lesssim {n}/{\log p}.\]

In the following, we define the source and target problems used in the reduction. Let \({\rm LT}(n,k_u,k,k_\xi,p,\tau)\) denote the linear testing problem \[H_0:\theta\in\Theta(k_u;\xi,t_0) \qquad\text{vs.}\qquad H_1:\theta\in\Theta_{\pm\tau}(k;\xi,t_0),\] where \(\xi\) has \(k_\xi\) nonzero entries, all equal to one. The total error probability of a test is defined as the sum of its size and its worst-case Type II error probability. Let \({\rm SCCA}(n,s,p_1,p_2,\lambda)\) denote the sparse CCA detection problem \[\label{eq:32sparse32cca32detection} H_0: \begin{pmatrix} U_1\\ U_2 \end{pmatrix} \sim \mathcal{N}(0,\mathbf{I}_{p_1+p_2})^{\otimes n} \quad \text{vs.} \quad H_1: \begin{pmatrix} U_1\\ U_2 \end{pmatrix} \sim \mathcal{N}\left( 0, \begin{pmatrix} \mathbf{I}_{p_1} & \lambda\boldsymbol{\delta}_1\boldsymbol{\delta}_2^\top \\ \lambda\boldsymbol{\delta}_2\boldsymbol{\delta}_1^\top & \mathbf{I}_{p_2} \end{pmatrix} \right)^{\otimes n},\tag{37}\] where \(\boldsymbol{\delta}_1\) and \(\boldsymbol{\delta}_2\) are drawn uniformly from \[\mathcal{V}_{s,p} = \left\{ v\in\mathbb{R}^p: \|v\|_0=s,\; v_i=s^{-1/2}\;\text{for all } i\in{\rm supp}(v) \right\}.\]

Theorem 3. Suppose 3 holds, and \(\sqrt{n}/\log p\ll k_u \ll k_\xi \lesssim {n}/{\log p}\). Assume there exists a polynomial-time algorithm that solves \({\rm LT}(n,k_u,k,k_\xi,p,\tau)\) for some \(1\le k\le k_u\) with total error probability at most \(\alpha+\eta\), where \[\tau = c\,\rho_n^2\,\frac{k_u}{\sqrt{k_\xi}},\] for a sufficiently small constant \(c>0\) and some \(\rho_n<1/2\). Then there exists a polynomial-time algorithm that solves \[{\rm SCCA}\left( 2n,\, \lfloor k_u/4\rfloor,\, k_\xi,\, p-k_\xi,\, \rho_n \right)\] with total error probability at most \(\alpha+\eta\).

If we take \(\rho_n \asymp \sqrt{{k_\xi\log p}/{n}},\) Theorem 3 implies that under the conjectured polynomial-time hardness of the following asymmetric sparse CCA instance \[\label{eq:32relevant32SCCA} {\rm SCCA}\left( 2n,\, \lfloor k_u/4\rfloor,\, k_\xi,\, p-k_\xi,\, \sqrt{\frac{k_\xi\log p}{n}} \right),\tag{38}\] \({\rm LT}(n,k_u,k,k_\xi,p,\tau)\) cannot be solved efficiently at the separation scale \[\tau \asymp \frac{k_u\sqrt{k_\xi}\log p}{n}.\] This rate matches the computationally feasible upper bound in 36 when \(k_\xi\lesssim n/\log p\). Therefore, Theorem 3 provides conditional evidence for the computational barrier in the flat sparse loading subclass.

Connection with known computational barriers for sparse CCA. The sparse CCA instance in 38 is asymmetric, with dimensions \(k_\xi\) and \(p-k_\xi\). This asymmetry is important because many existing sparse CCA reductions are formulated for balanced dimensions and hence do not directly apply to the instance induced by 3. The closest related result is [14], which establishes low-degree barriers for asymmetric sparse CCA. However, a direct combination of [14] with 3 would only rule out strong separation for the linear testing problem, at the feasible upper bound in 36 up to logarithmic factors. By contrast, Theorem 2 rules out weak separation by low-degree polynomials. Thus, the direct low-degree analysis in 4.2 is not a straightforward consequence of known sparse CCA results.

Algorithmic evidence for computational barriers in sparse CCA. We also connect the above reduction to standard algorithmic evidence for computational barriers in sparse CCA detection. In the sparse CCA formulation, the empirical cross-covariance matrix contains the planted cross-correlation signal, so it is natural to compare procedures that threshold different functionals of this matrix. Exhaustive scan statistics can detect signals at the information-theoretic boundary \(\rho_n\asymp \sqrt{k_u\log p/n}\) given in [45], but they require a combinatorial search over \(k_u\times k_u\) submatrices. By contrast, simple polynomial-time procedures, such as max-column statistics, require stronger signals, of order \(\rho_n\asymp \sqrt{k_\xi\log p/n}\) up to logarithmic factors, in the parameter regime used in 3. Thus, the algorithmic behavior of natural sparse CCA tests is consistent with the computational barrier suggested by the reduction. A detailed comparison of these test statistics is deferred to 11.3.

5 Benchmarks with structured or known covariance↩︎

The lower bound in 3 is sharp up to logarithmic factors in the ultra-sparse regime for arbitrary loading vectors and in the moderately sparse regime for certain loading subclasses. For arbitrary loading vectors in the moderately sparse regime, it is of interest to assess the sharpness of the lower bound. This section calibrates the scope of the lower-bound analysis through two covariance benchmarks. In both benchmarks, we consider any sparsity level \(k_u\lesssim n/\log p\) and any loading vector \(\xi\).

First, when \(\Sigma\) is unknown but has a sparse signed-spiked structure, the adaptive testing boundary is shown to be \(\nu_1/\sqrt n+\nu_2 k_u\log p/n\) up to logarithmic factors. Thus the information-theoretic lower bound in 3 is nearly attainable in this structured unknown-covariance model. Meanwhile, as the computational results in 4 continue to apply, we obtain evidence for a statistical–computational gap.

Second, when \(\Sigma\) is known, covariance uncertainty is removed from the testing problem and we obtain a sharper adaptive upper bound than in the unknown-covariance setting. When \(\Sigma_0\) is diagonal, this upper bound matches the lower bound \(\nu_1/\sqrt{n}\) up to logarithmic factors.

5.1 \(\Sigma\) with sparse signed-spiked structures↩︎

We consider a structured covariance class under which the lower bound in Theorem 1 is statistically attainable, up to logarithmic factors, even in the moderately sparse regime \(\sqrt{n}/\log p \ll k_u \lesssim n/\log p\). This structural assumption is imposed to isolate an information-theoretic phenomenon: for sparse signed-spiked covariance matrices, \(\Sigma\) can be estimated at a rate sufficiently fast to attain the lower bound. The restriction is also natural from the perspective of high-dimensional covariance modeling. Sparse signed-spiked covariance models are standard in sparse PCA, principal subspace estimation, and principal component regression; see, for example, [46][49].

To formulate the sparse signed-spiked structure, we first introduce some notation. For a matrix \(V \in \mathbb{R}^{p \times r}\), let \(V_{j*}\) denote its \(j\)th row. The row support of \(V\) is defined by \[\label{eq:suppV} {\rm supp}(V) = \{ j \in [p] : V_{j*} \neq 0 \},\tag{39}\] and its cardinality is denoted by \(|{\rm supp}(V)|\). Let \[\mathbb{O}(p,r) = \{ V \in \mathbb{R}^{p \times r} : V^\top V = \mathbf{I}_r \}\] denote the set of \(p \times r\) matrices with orthonormal columns. For the covariance matrix \(\Sigma\), we define the following sparse signed-spiked parameter space: \[\label{eq:32sparse32spiked} \begin{align} \Pi_0(k,p) = \Bigl\{ \Sigma = &V \operatorname{diag}(\lambda_1,\ldots,\lambda_r) V^\top + \mathbf{I}_p : \text{ for some } r \le k, \\ &\frac{1}{M_1}-1 \le \lambda_r \le \cdots \le \lambda_1 \le M_1-1,\; V \in \mathbb{O}(p,r),\; |\operatorname{supp}(V)| \le k \Bigr\}. \end{align}\tag{40}\] It is immediate that any \(\Sigma \in \Pi_0(k,p)\) satisfies the eigenvalue condition required in 4 that \(1/M_1 \le \lambda_{\min}(\Sigma) \le \lambda_{\max}(\Sigma) \le M_1\).

We now define the null and alternative spaces, as well as the corresponding adaptive separation distance, under the additional structural assumption \(\Sigma \in \Pi_0(k,p)\). Analogously to 6 , 7 , 8 , and 10 , we define \[\left\{ \begin{align} \Theta^{\text{spike}}(k;\xi,t_0) &= \Bigl\{ \theta=(\beta,\Sigma,\sigma)\in\Theta(k;\xi,t_0): \Sigma\in \Pi_0(k,p) \Bigr\},\\ \Theta^{\text{spike}}_{\pm \tau}(k;\xi,t_0) &= \Bigl\{ \theta=(\beta,\Sigma,\sigma)\in\Theta_{\pm \tau}(k;\xi,t_0): \Sigma\in \Pi_0(k,p) \Bigr\},\\ \tau^{\text{spike}}_{\text{adap}}(k_u,k\,;\,\xi) &= \inf\Bigl\{ \tau : \sup_{\psi \in \Psi_\alpha(\Theta^{\text{spike}}(k_u;\xi,t_0))} \inf_{\theta \in \Theta^{\text{spike}}_{\pm \tau}(k;\xi,t_0)} \mathbb{E}_\theta \psi \ge 1 - \eta \Bigr\}. \end{align} \right.\] As before, the rank \(r\) and the true sparsity level \(k\) are not assumed to be known to the test.

Theorem 4. Under 3, there exists some constant \(c>0\) such that if \(n\ge c\,k_u\log p\), then for any \(1\le k\le k_u\), we have \[\tau^{\text{spike}}_{\text{adap}}(k_u,k\,;\,\xi) \;\asymp_{\log}\; \frac{\nu_1}{\sqrt{n}} + \nu_2\,\frac{k_u \log p}{n}.\]

4 gives the optimal adaptive separation rate under the sparse signed-spiked covariance assumption \(\Sigma \in \Pi_0(k,p)\), up to logarithmic factors. The lower-bound part follows directly from the constructions used for \(\nu_1\) and \(\nu_2\): in those constructions, \(\Sigma\) is either the identity matrix or has the form in 34 , and therefore belongs to the sparse signed-spiked class in 40 .

The upper bound is more delicate and differs substantially from the upper-bound argument in 1. It is attained by a statistically optimal but computationally inefficient procedure. The key ingredient of this procedure is covariance estimation over the parameter space \(\Pi_0(k_u,p)\). In the proof, we construct an estimator of \(\Sigma\) based on exhaustive search over sparse supports. This estimator is specific to the sparse signed-spiked class considered here, and its operator-norm error is of order \(\sqrt{(k_u\log p)/n}\), which is minimax rate-optimal because the standard sparse-spiked submodel yields a matching lower bound [47]. Consequently, under the sparse signed-spiked structure, the information-theoretic limits governed by \(\nu_1\) and \(\nu_2\) are statistically attainable up to logarithmic factors, although the resulting procedure is not computationally efficient.

Computational barrier. The computational evidence in 4 is compatible with the sparse signed-spiked covariance class in 40 . Since all covariance constructions used in the proofs of [thm: computational lower bound,thm: reduction] are of the form 34 , the proofs of [thm: computational lower bound,thm: reduction] remain valid when we replace the used parameter spaces by their sparse signed-spiked counterparts. Therefore, the low-degree lower bounds for general loadings and the sparse CCA reduction for flat sparse loadings continue to apply without change under the sparse signed-spiked covariance restriction.

Accordingly, 4 should be interpreted together with the evidence for computational barriers in the moderately sparse regime \(\sqrt{n}/\log p\ll k_u \lesssim n/\log p\): although the sparse signed-spiked adaptive separation distance is statistically attainable up to logarithmic factors, achieving this rate appears to require computationally intractable procedures. In particular, when the low-degree lower bound \(H(n/\log p; \xi) k_u\log p/n\) is much larger than \(\nu_1/\sqrt{n}+\nu_2 k_u \log p /n\), the testing problem is predicted to exhibit a statistical–computational gap under the standard low-degree heuristic.

5.2 \(\Sigma\) known↩︎

Thus far, our analysis has focused on the practically relevant setting in which the covariance matrix \(\Sigma\) is unknown. We now consider the idealized setting in which \(\Sigma=\Sigma_0\) is known a priori. Although such knowledge is rarely available in practice, as noted by [34], [50], this setting can be viewed as an extreme case of the semi-supervised framework, corresponding to an infinite amount of unlabeled data for estimating \(\Sigma\). Comparing this benchmark with the unknown-design case helps isolate how uncertainty in the covariance matrix \(\Sigma\) affects the testing problem 2 .

The following theorem gives a general upper bound for arbitrary known \(\Sigma_0\) satisfying the eigenvalue condition. When \(\Sigma_0\) is diagonal, including the identity covariance as a special case, the theorem further characterizes the minimax adaptive separation distance up to logarithmic factors.

Theorem 5. Under 3, there exists some constant \(c>0\) such that if \(n\ge c\,k_u\log p\), then for any \(1\le k\le k_u\) and any \(\Sigma_0\in\mathbb{R}^{p\times p}\) satisfying \(1/M_1 \le \lambda_{\min}(\Sigma_0) \le \lambda_{\max}(\Sigma_0) \le M_1\), we have \[\tau_{\text{adap}}(k_u,k,\Sigma_0\,;\,\xi) \;\lesssim\; \min_{0\le m\le p} \left( H(m;\xi)\,\frac{1}{\sqrt{n}} + |\xi_{m+1}|\,k_u\sqrt{\frac{\log p}{n}} \right).\] Furthermore, when \(\Sigma_0\) is diagonal, we have \[\tau_{\text{adap}}(k_u,k,\Sigma_0\,;\,\xi) \;\asymp_{\log}\; \frac{\nu_1}{\sqrt{n}}.\]

The proof of 5 follows the same plug-in \(+\) debiasing decomposition as in 3.3, but with an oracle debiasing direction. For the coordinates of \(\xi\) with small magnitudes, the plug-in construction is unchanged and contributes the term \(|\xi_{m+1}|\,k_u\sqrt{{\log p}/{n}} .\) For the leading coordinates, knowing \(\Sigma_0\) allows us to use the oracle population direction \(\Omega_0\xi^{(1)}=\Sigma_0^{-1}\xi^{(1)}\) directly in the debiased estimator. This removes the error introduced by estimating the bias-correction direction from the sample covariance. Combining the oracle debiased interval for the leading coordinates with the plug-in interval for the remaining coordinates yields the displayed upper bound after optimizing over \(m\). For brevity, we defer the proof to 7.3.

This explains the improvement over the unknown-design upper bound in 1. In the moderately sparse regime \(\sqrt{n}/\log p \ll k_u \lesssim n/\log p\), uncertainty in \(\Sigma\) creates an additional cost through the construction of the bias-correction direction. When \(\Sigma_0\) is known, this cost disappears, and the plug-in \(+\) debiasing construction attains a sharper rate. In particular, when \(\Sigma_0\) is diagonal, the minimax adaptive separation distance is \(\nu_1/\sqrt n\), up to logarithmic factors.

The known-\(\Sigma\) benchmark also has a broader semi-supervised interpretation: exact knowledge of \(\Sigma\) is stronger than necessary. It is enough to construct an estimator \(\widehat{\Omega}\) of \(\Omega_0\) satisfying \[\|\widehat{\Omega}-\Omega_0\|_{2\to\infty} \;\lesssim\; \frac{1}{k_u\sqrt{\log p}} .\] Here, for a matrix \(A\in\mathbb{R}^{p\times p}\), \(\|A\|_{2\to\infty}:=\max_{1\le j\le p}\|A_{j\cdot}\|_2\) denotes the maximum row \(\ell_2\)-norm. Such an estimator may be available, for example, when \(\Sigma\) has additional structure or when sufficiently many unlabeled samples of \(X\) are available. Under this condition, the same debiasing argument applies with \(\widehat{\Omega}\) in place of \(\Omega\), and the induced error is negligible. Consequently, the adaptive rate involving \(\nu_1\) remains attainable up to logarithmic factors.

Related phenomena have been observed in [7], [51]. In particular, [51] construct asymptotically normal debiased estimators for individual regression coefficients, while [7] use data splitting to obtain confidence intervals for \(\xi^\top\beta\) with optimal length for certain structured loading vectors. However, their results do not directly imply the adaptive upper bound above for general loading vectors \(\xi\), and our plug-in \(+\) debiasing construction is not a direct consequence of these existing methods.

6 Discussion↩︎

We investigated the minimax and adaptive separation distances for testing a general linear functional \(\xi^\intercal\beta\) in high-dimensional linear regression with Gaussian random design. Our results reveal a transition in both statistical and computational behavior as the sparsity upper bound \(k_u\) moves from the ultra-sparse regime \(k_u \lesssim \sqrt{n}/\log p\) to the moderately sparse regime \(\sqrt{n}/\log p \ll k_u \lesssim n/\log p\).

In the ultra-sparse regime, the testing problem admits a statistical characterization, \(\nu_1(k_u;\xi)/\sqrt{n}\), that is precise up to logarithmic factors. By contrast, the moderately sparse regime exhibits qualitatively different behavior. In this regime, another lower-bound component \(\nu_2(k_u;\xi)k_u\log p/n\) becomes relevant. This lower-bound construction is based on coupling the leading coordinates of \(\xi\) with a hidden sparse perturbation in the design covariance. A closely related perturbation mechanism is used in our low-degree lower bound for general loadings and our sparse CCA reduction for flat sparse loadings.

Several directions remain open. First, closing the remaining gaps in both the ultra-sparse and moderately sparse regimes remains an important direction for future work. Our results partially address this question by identifying structured classes of loading vectors for which the upper and lower bounds match sharply. Second, it would be of interest to extend the theory beyond \(\ell_0\)-sparsity to more general structured sparsity classes, such as \(\ell_q\)-sparsity [3] and sparse-group structures [52]. Third, one may replace the sparse signed-spiked covariance structure with a sparse precision matrix structure, under which each row of \(\Omega=\Sigma^{-1}\) has only a small number of nonzero entries [5]. Since this class is generally larger than the sparse signed-spiked covariance class considered in 5.1, an important question is whether a comparable minimax upper bound can still be attained by statistically optimal (but possibly computationally inefficient) procedures. Finally, developing computational lower bounds for sparse CCA with more general latent vectors would strengthen the theoretical basis for the computational barrier identified here. It also remains open whether analogous statistical–computational phase transitions arise for quadratic or more general nonlinear functionals.

Supplement to “Linear Functional Testing with General Loadings in Sparse Regression: Separation Rates and Computational Barriers”

This supplement collects proofs and auxiliary results deferred from the main text. We structure the material into three parts: (i) upper bounds, (ii) information-theoretic lower bounds, and (iii) additional technical lemmas.

7 establishes the upper bounds in [thm: hypothesis,thm: hypothesis design] (as well as the related upper bound in 4). We first compile concentration tools in 7.1. We then derive confidence-interval constructions under unknown design in 7.2, and under known design via data splitting in 7.3. Finally, 7.4 adapts these arguments to the setting of 4.

8 proves the lower bounds in [thm: hypothesis,thm: computational lower bound,thm: reduction,thm: hypothesis design] via the standard “least favorable prior + \(\chi^2\)-divergence” method. Specifically, we construct priors supported on (or predominantly supported on) the null parameter sets and bound the resulting \(\chi^2\)-divergence using explicit Gaussian calculations; these bounds imply power limitations through 9.

The technical ingredients required for the upper-bound arguments (in particular, feasibility of the projection step) are proved in 9. Additional tools for the lower bounds and checks of prior validity (e.g., eigenvalue/sparsity control and \(\chi^2\) integral identities) are collected in 10. Finally, 11 contains complementary discussions, and 12 contains selected loading-profile examples, including one random-predictor example.

Table 2 gives a cross-reference from each main result to its proof and selected lemmas. For 4, the table lists the separate upper-bound proof; its lower-bound proof follows from the same lower-bound argument used for 1.

Table 2: Proof map: where each main result is proved and the key auxiliary lemmas it uses. Parentheses after a lemma indicate where it is proved.
Result Location Key ingredients (selected)
[thm:32hypothesis] (upper) [sec:sec:32proof32upper32unknown32design] [lem:32feasibility32of32optimization] ([sec:sec:32proof32feasibility32of32optimization]).
[thm:32hypothesis32design] (upper) [sec:sec:32proof32upper32design] Debiased estimator under known design [eq:32bias32estimator32known32design].
[thm:32sparse32spiked] (upper) [sec:sec:32proof32upper32np32hard] Precision matrix estimation [lem:32sparse32spiked] ([sec:sec:32proof32sparse32spiked]) and the corresponding debiased estimator [eq:32debiased32center32np32hard].
[thm:32hypothesis] (lower) [sec:sec:32proof32lower32ultra32sparse] Prior validity [lem:32valid32covariance] ([sec:sec:32proof32valid32covariance]); \(\chi^2\) integral bound [lem:32chi32square32integral321] ([sec:sec:32proof32chi32square32integral]).
[thm:32hypothesis32design] (lower) [sec:sec:32proof32design32lower] Prior validity [lem:32valid32covariance322] ([sec:sec:32proof32valid32covariance322]); \(\chi^2\) integral bound [lem:32chi32square32integral322]([sec:sec:32proof32chi32square32integral322]).
[thm:32computational32lower32bound] [sec:sec:32proof32computational32lower32bound] Prior validity [lem:32valid32covariance323] ([sec:sec:32proof32valid32covariance323]); Low-degree quantity control [lem:32low32degree32integral] ([sec:sec:32proof32low32degree32integral]).
[thm:32reduction] [sec:sec:32proof32reduction] Polynomial-time reduction [lem:32reduction].

7 Proof of the upper bounds↩︎

In this section, we prove the upper bounds in 1 and 5. Define \(\mathcal{A}_0\) to be the event such that the following holds: \[\label{eq:32event32A0} 1.1\hat{\sigma}>\sigma>0.9\hat{\sigma},\|\hat{\beta}-\beta\|_1\le c_\beta \sigma k_u\sqrt{\frac{\log p}{n}},\|\hat{\beta}-\beta\|_2\le C_\beta\sigma \sqrt{\frac{k_u\log p}{n}}.\tag{41}\] For any prescribed error probability \(\bar\alpha\in(0,1)\), Conditions 1 and 2 allow \(n\) to be taken large enough so that \[\inf_{\theta\in\Theta(k_u)} \mathbb{P}_\theta\{\mathcal{A}_0\} \ge 1-\bar\alpha.\]

Throughout this supplement, we write \(\Omega=\Sigma^{-1}\) for the precision matrix and \(z_\alpha\) for the \(\alpha\)-quantile of the standard normal distribution. We use the convention \(\operatorname{sign}(x)=1\) for \(x\ge0\) and \(\operatorname{sign}(x)=-1\) for \(x<0\).

7.1 Auxiliary lemmas for the proof of upper bounds↩︎

Lemma 1 ([53]). Let \(Z\sim \chi^2(n)\). For any \(x>0\), we have \[\mathbb{P}\left(Z-n\ge 2\sqrt{nx}+2x\right)\le e^{-x},\] and \[\mathbb{P}\left(Z-n\le -2\sqrt{nx}\right)\le e^{-x}.\]

We introduce the following definitions. The sub-Gaussian norm of a random variable \(U\) is defined as \[\|U\|_{\psi_2} := \sup_{q \ge 1} \frac{1}{\sqrt{q}} \bigl( \mathbb{E}|U|^{q} \bigr)^{1/q},\] and the sub-Gaussian norm of a random vector \(U \in \mathbb{R}^p\) is defined as \[\|U\|_{\psi_2} := \sup_{v \in S^{p-1}} \|\langle v, U \rangle\|_{\psi_2},\] where \(S^{p-1}\) denotes the unit sphere in \(\mathbb{R}^p\).

Similarly, the sub-exponential norm of a random variable \(U\) is defined as \[\|U\|_{\psi_1} := \sup_{q \ge 1} \frac{1}{q} \bigl( \mathbb{E}|U|^{q} \bigr)^{1/q},\] and the sub-exponential norm of a random vector \(U \in \mathbb{R}^p\) is defined as \[\|U\|_{\psi_1} := \sup_{v \in S^{p-1}} \|\langle v, U \rangle\|_{\psi_1}.\]

Lemma 2 ([54]). For a sub-exponential random variable \(U\), we have \[\|U-\mathbb{E}U\|_{\psi_1} \le 2 \|U\|_{\psi_1}.\]

Lemma 3 ([54]). Let \(X_1,\ldots,X_N\) be independent centered sub-exponential random variables, and \[K = \max_{i} \|X_i\|_{\psi_1}.\] Then for every \(a = (a_1,\ldots,a_N) \in \mathbb{R}^N\) and every \(t \ge 0\), we have \[\mathbb{P}\!\left( \left| \sum_{i=1}^N a_i X_i \right| \ge t \right) \le 2 \exp\!\left[ -c_0 \min\!\left( \frac{t^2}{K^2 \|a\|_2^2},\; \frac{t}{K \|a\|_\infty} \right) \right],\] where \(c_0 > 0\) is an absolute constant.

Lemma 4. Let \(X\) and \(Y\) be sub-Gaussian random variables. Then \(XY\) is sub-exponential. Moreover, \[\|X Y\|_{\psi_1} \le 2\|X\|_{\psi_2}\|Y\|_{\psi_2}.\]

Proof. Fix any \(q\ge 1\). By Hölder’s inequality, \[\begin{align} \left(\mathbb{E}|XY|^q\right)^{1/q} &\le \left(\mathbb{E}|X|^{2q}\right)^{1/(2q)} \left(\mathbb{E}|Y|^{2q}\right)^{1/(2q)}\\ &\le \bigl(\|X\|_{\psi_2}\sqrt{2q}\bigr)\bigl(\|Y\|_{\psi_2}\sqrt{2q}\bigr) = 2q\,\|X\|_{\psi_2}\,\|Y\|_{\psi_2}. \end{align}\] Dividing by \(q\) and taking the supremum over \(q\ge 1\) yields the desired inequality. ◻

Lemma 5 (Mixed confidence interval). Let \(\mathcal{P}\subseteq \Theta(k_u)\) and let \(\xi\in\mathbb{R}^p\) satisfy \(|\xi_1|\ge \cdots \ge |\xi_p|\). Let \(\alpha_1,\alpha_2\in(0,1)\) satisfy \(\alpha_1+\alpha_2<1\). Assume that, for every fixed \(v\in\mathbb{R}^p\), we have the following two intervals with nonnegative random radii \[\begin{align} {\rm CI}_{\rm db}(v) &= [\hat{h}_{\rm db}(v)-r_{\rm db}(v),\, \hat{h}_{\rm db}(v)+r_{\rm db}(v)],\\ {\rm CI}_{\rm pi}(v) &= [v^\top\hat{\beta}-r_{\rm pi}(v),\, v^\top\hat{\beta}+r_{\rm pi}(v)], \end{align}\] and that they satisfy the following high-probability coverage and radius bounds uniformly over \(\mathcal{P}\): \[\inf_{\theta\in\mathcal{P}} \mathbb{P}_\theta \left( v^\top\beta\in{\rm CI}_{\rm db}(v),\; r_{\rm db}(v)\le R_{\rm db}(v) \right) \ge 1-\alpha_1\] and \[\inf_{\theta\in\mathcal{P}} \mathbb{P}_\theta \left( v^\top\beta\in{\rm CI}_{\rm pi}(v),\; r_{\rm pi}(v)\le R_{\rm pi}(v) \right) \ge 1-\alpha_2,\] where \(R_{\rm db}(v)\) and \(R_{\rm pi}(v)\) are nonnegative non-random radius envelopes. The debiased envelope \(R_{\rm db}(v)\) may depend on \(v\), \(\sigma\), and fixed model constants, but not on the data or on the parameter \(\theta\in\mathcal{P}\). The plug-in envelope is given by \[\qquad R_{\rm pi}(v) = C_{\rm pi}^+\sigma\|v\|_\infty k_u\sqrt{\frac{\log p}{n}},\] where \(C_{\rm pi}^+>0\) is a fixed constant that may depend on fixed estimator constants such as \(c_{\beta}\) but not on \(n,p,k_u,v\), or on the parameter \(\theta\in\mathcal{P}\).

For a fixed deterministic \(0\le m\le p\), define \[\xi^{(1)}_m=(\xi_1,\ldots,\xi_m,0,\ldots,0)^\top, \qquad \xi^{(2)}_m=\xi-\xi^{(1)}_m ,\] with the convention \(\xi_{p+1}=0\), and set \[{\rm CI}_m(\xi) = {\rm CI}_{\rm db}(\xi^{(1)}_m) + {\rm CI}_{\rm pi}(\xi^{(2)}_m),\] where the sum denotes the Minkowski sum of two intervals. Then \({\rm CI}_m(\xi)\) is a \((1-\alpha_1-\alpha_2)\)-level confidence interval for \(\xi^\top\beta\) over \(\mathcal{P}\). Moreover, uniformly over \(\mathcal{P}\), with probability at least \(1-\alpha_1-\alpha_2\) its radius is bounded by \[R_m(\xi) = R_{\rm db}(\xi^{(1)}_m) + C_{\rm pi}^+\sigma |\xi_{m+1}| k_u\sqrt{\frac{\log p}{n}} .\] Consequently, the test \[\psi_m = \mathbf{1}\{t_0\notin {\rm CI}_m(\xi)\}\] has Type-I error at most \(\alpha_1+\alpha_2\) over \(\mathcal{P}\cap\Theta(k_u;\xi,t_0)\) and has power at least \(1-\alpha_1-\alpha_2\) at every \(\theta\in\mathcal{P}\) satisfying \(|\xi^\top\beta-t_0|>2R_m(\xi)\).

Proof. Let \(E_m\) be the intersection of the two events in the two assumed bounds, with \(v=\xi_m^{(1)}\) in the debiased bound and \(v=\xi_m^{(2)}\) in the plug-in bound. In other words, \(E_m\) is the event that \[(\xi_m^{(1)})^\top\beta\in{\rm CI}_{\rm db}(\xi_m^{(1)}),\; r_{\rm db}(\xi_m^{(1)})\le R_{\rm db}(\xi_m^{(1)}), (\xi_m^{(2)})^\top\beta\in{\rm CI}_{\rm pi}(\xi_m^{(2)}),\; r_{\rm pi}(\xi_m^{(2)})\le R_{\rm pi}(\xi_m^{(2)}).\] The union bound gives \(\mathbb{P}_\theta(E_m)\ge 1-\alpha_1-\alpha_2\) uniformly over \(\mathcal{P}\). On \(E_m\), \[\xi^\top\beta = (\xi_m^{(1)})^\top\beta+(\xi_m^{(2)})^\top\beta \in {\rm CI}_{\rm db}(\xi_m^{(1)}) + {\rm CI}_{\rm pi}(\xi_m^{(2)})\] and both radius bounds also hold. Therefore, the radius of the interval sum is at most \[R_{\rm db}(\xi^{(1)}_m) + R_{\rm pi}(\xi^{(2)}_m) = R_{\rm db}(\xi^{(1)}_m) + C_{\rm pi}^+\sigma\|\xi^{(2)}_m\|_\infty k_u\sqrt{\frac{\log p}{n}},\] which equals \(R_m(\xi)\) because \(\|\xi^{(2)}_m\|_\infty=|\xi_{m+1}|\). If \(\xi^\top\beta=t_0\), coverage implies \(t_0\in{\rm CI}_m(\xi)\); hence the Type-I error is at most \(\alpha_1+\alpha_2\). If \(|\xi^\top\beta-t_0|>2R_m(\xi)\), then under \(E_m\), the coverage and radius bound both hold and they imply that \(t_0\notin {\rm CI}_m(\xi)\). Therefore, the power is at least \(\mathbb{P}_\theta(E_m)\ge 1-\alpha_1-\alpha_2\). ◻

Lemma 6 (Plug-in confidence interval). Let \(\mathcal{P}\subseteq\Theta(k_u)\) and fix \(v\in\mathbb{R}^p\). Suppose that \((\hat{\beta},\hat{\sigma})\) is computed from \(N\) observations. Assume that, for a prescribed \(\bar\alpha\in(0,1)\), there exists an event \(\mathcal{E}_{\rm pi}\) such that \[\inf_{\theta\in\mathcal{P}} \mathbb{P}_\theta(\mathcal{E}_{\rm pi}) \ge 1-\bar\alpha,\] and on \(\mathcal{E}_{\rm pi}\), the following bounds hold: \[1.1\hat{\sigma}>\sigma>0.9\hat{\sigma}, \qquad \|\hat{\beta}-\beta\|_1 \le c_\beta\sigma k_u\sqrt{\frac{\log p}{N}} .\] Choose constants \(C_{\rm pi}\ge 1.1c_\beta\) and \(C_{\rm pi}^+\ge C_{\rm pi}/0.9\), and define \[{\rm CI}_{\rm pi,N}(v) = [v^\top\hat{\beta}-r_{\rm pi,N}(v),\, v^\top\hat{\beta}+r_{\rm pi,N}(v)], \qquad r_{\rm pi,N}(v) = C_{\rm pi}\hat{\sigma}\|v\|_\infty k_u\sqrt{\frac{\log p}{N}} .\] Then \[\inf_{\theta\in\mathcal{P}} \mathbb{P}_\theta \left( v^\top\beta\in{\rm CI}_{\rm pi,N}(v),\; r_{\rm pi,N}(v) \le C_{\rm pi}^+\sigma\|v\|_\infty k_u\sqrt{\frac{\log p}{N}} \right) \ge 1-\bar\alpha .\] In particular, if \(N\asymp n\), the same statement gives the plug-in assumption in 5 after enlarging \(C_{\rm pi}^+\) by a fixed factor.

Proof. On \(\mathcal{E}_{\rm pi}\), \[|v^\top\hat{\beta}-v^\top\beta| \le \|v\|_\infty\|\hat{\beta}-\beta\|_1 \le c_\beta\sigma\|v\|_\infty k_u\sqrt{\frac{\log p}{N}} \le C_{\rm pi}\hat{\sigma}\|v\|_\infty k_u\sqrt{\frac{\log p}{N}},\] where the last inequality uses \(\sigma<1.1\hat{\sigma}\) and \(C_{\rm pi}\ge1.1c_\beta\). Thus \(v^\top\beta\in{\rm CI}_{\rm pi,N}(v)\) on \(\mathcal{E}_{\rm pi}\). The same event also gives \(\hat{\sigma}<\sigma/0.9\), and therefore \[r_{\rm pi,N}(v) \le \frac{C_{\rm pi}}{0.9}\sigma\|v\|_\infty k_u\sqrt{\frac{\log p}{N}} \le C_{\rm pi}^+\sigma\|v\|_\infty k_u\sqrt{\frac{\log p}{N}}.\] Taking probabilities and using the assumed bound for \(\mathbb{P}_\theta(\mathcal{E}_{\rm pi})\) proves the claim. ◻

7.2 Proof of the upper bound in Theorem 1↩︎

Proof. Let \(\alpha_\star=\min\{\alpha,\eta\}\). As discussed in 3.3, to establish the upper bound in 1 it suffices to show that for any loading vector \(\xi \in \mathbb{R}^p\) one can construct:

  • a \((1-\alpha_\star/2)\)-level plug-in confidence interval with length \[O\!\left( \sigma \|\xi\|_\infty\, k_u \sqrt{\frac{\log p}{n}} \right),\]

  • a \((1-\alpha_\star/2)\)-level debiased confidence interval with length \[O\!\left( \sigma \|\xi\|_2 \left( \frac{1}{\sqrt{n}} + \frac{k_u \log p}{n} \right) \right).\]

Indeed, once these two ingredients are available, we decompose \(\xi\) at a cutoff \(m \in [p]\) as \(\xi=\xi^{(1)}+\xi^{(2)}\), where \[\xi^{(1)}=(\xi_1,\ldots,\xi_m,0,\ldots,0), \qquad \xi^{(2)}=(0,\ldots,0,\xi_{m+1},\ldots,\xi_p).\] We then construct a \((1-\alpha_\star/2)\)-level debiased confidence interval for \((\xi^{(1)})^\top\beta\) and a \((1-\alpha_\star/2)\)-level plug-in confidence interval for \((\xi^{(2)})^\top\beta\), and combine them to obtain a valid \((1-\alpha_\star)\)-level confidence interval for \(\xi^\top\beta\). The resulting interval length is of order \[\sigma\left( \|\xi^{(1)}\|_2\Bigl(\frac{1}{\sqrt{n}}+\frac{k_u\log p}{n}\Bigr) + \|\xi^{(2)}\|_\infty\, k_u\sqrt{\frac{\log p}{n}} \right).\] Using the condition \(\sigma\le M_2\) in 4 , we obtain an upper bound on the adaptive separation distance for any given \(m\). Since such an upper bound holds for all \(m\in [p]\), minimizing over \(m\) yields the desired upper bound in 1.

The construction of the plug-in confidence interval is an application of 6 with \(N=n\), \(\mathcal{P}=\Theta(k_u)\), \(v=\xi\), and \(\mathcal{E}_{\rm pi}=\mathcal{A}_0\). Therefore, we focus on the construction of the debiased confidence interval.

Let \(\widetilde{\mathcal{A}}\) be the event that the vector \(u=\Omega\xi\) is a feasible point for the optimization problem 30 . The following lemma guarantees that \(\widetilde{\mathcal{A}}\) happens with probability close to 1.

Lemma 7. For any \(\alpha\in(0,1)\), there exists a constant \(C_\xi>0\) such that \(\widetilde{\mathcal{A}}\) holds with probability at least \(1-\alpha/24\) for all sufficiently large \(n\).

Recall the debiased estimator \(\hat{L}_{\mathrm{db}}(Z;\xi)\) defined in 29 . Its estimation error admits the decomposition \[\hat{L}_{\mathrm{db}}(Z;\xi)-\xi^\top\beta = \underbrace{\frac{1}{n}\hat{u}^\top X^\top\varepsilon}_{\mathrm{I}} + \underbrace{(\xi-\hat{\Sigma}\hat{u})^\top(\hat{\beta}-\beta)}_{\mathrm{II}},\] where \(\hat{\Sigma}=n^{-1}X^\top X\) and \(\varepsilon=Y-X\beta\sim\mathcal{N}(0,\sigma^2{\mathbf{I}}_n)\) is independent of \(X\) and \(\hat{u}\).

Bounding \(\mathrm{I}\). Define the event \[\mathcal{A}_1 = \left\{ \left|\frac{1}{n}\hat{u}^\top X^\top\varepsilon\right| \le \frac{\sigma}{\sqrt{n}} \sqrt{\hat{u}^\top\hat{\Sigma}\hat{u}}\, z_{1-\alpha/8} \right\}.\] Conditional on \(X\) and \(\hat{u}\), \[\frac{1}{n}\hat{u}^\top X^\top\varepsilon \;\Big|\; X \sim \mathcal{N}\!\left(0, \frac{\sigma^2}{n}\hat{u}^\top\hat{\Sigma}\hat{u} \right),\] so \(\mathcal{A}_1\) holds with probability at least \(1-\alpha/4\).

We next control \(\hat{u}^\top\hat{\Sigma}\hat{u}\). Let \[\mathcal{A}_2 = \left\{ u^\top\hat{\Sigma}u \le 1.1\,M_1^2\|\xi\|_2^2 \right\} \cap \widetilde{\mathcal{A}}.\] For \(u=\Omega\xi\), we have \[\frac{n}{\xi^\top\Omega\xi}\,u^\top\hat{\Sigma}u \stackrel{d}{=}\chi^2(n).\] By Lemma 1, there exists \(n_1\) such that for all \(n\ge n_1\), the following holds with probability at least \(1-\alpha/24\): \[\label{eq:32bound32oracle32quadform32revised} u^\top\hat{\Sigma}u \le 1.1\,\xi^\top\Omega\xi \le 1.1\,M_1^2\|\xi\|_2^2,\tag{42}\] where the second inequality follows from the eigenvalue bounds in 4 . By 7, there exists \(n_2\) such that for \(n\ge n_2\), \(\widetilde{\mathcal{A}}\) holds with probability at least \(1-\alpha/24\). Therefore, \(\mathcal{A}_2\) holds with probability at least \(1-\alpha/12\).

Since \(u\) is feasible on \(\widetilde{\mathcal{A}}\), the definition of \(\hat{u}\) implies that on the event \(\mathcal{A}_2\), we have \[\hat{u}^\top\hat{\Sigma}\hat{u} \le u^\top\hat{\Sigma}u.\] Therefore, \[\mathcal{A}_2 \subseteq \left\{ \hat{u}^\top\hat{\Sigma}\hat{u} \le 1.1\,M_1^2\|\xi\|_2^2 \right\} \cap \widetilde{\mathcal{A}}.\]

Bounding \(\mathrm{II}\). Recall the event \(\mathcal{A}_0\) in 41 , which holds with probability at least \(1-\alpha/6\). On \(\mathcal{A}_0\cap\widetilde{\mathcal{A}}\), by the definition of \(\hat{u}\), \[|(\xi-\hat{\Sigma}\hat{u})^\top(\hat{\beta}-\beta)| \le \|\hat{\Sigma}\hat{u}-\xi\|_\infty\|\hat{\beta}-\beta\|_1 \le \sigma c_\beta C_\xi\|\xi\|_2\frac{k_u\log p}{n}.\]

Synthesis. Combining the above bounds, for sufficiently large \(n\), the event \(\mathcal{A}_0\cap\mathcal{A}_1\cap\mathcal{A}_2\) holds with probability at least \(1-\alpha/2\), and on this event, \[\begin{align} \left| \hat{L}_{\mathrm{db}}(Z;\xi)-\xi^\top\beta \right| &\le \frac{\sigma}{\sqrt{n}} \sqrt{\hat{u}^\top\hat{\Sigma}\hat{u}}\, z_{1-\alpha/8} + \sigma c_\beta C_\xi\|\xi\|_2\frac{k_u\log p}{n} \\ &\lesssim \sigma\|\xi\|_2 \left( \frac{1}{\sqrt{n}}+\frac{k_u\log p}{n} \right). \end{align}\] Replacing \(\sigma\) by \(1.1\hat{\sigma}\) yields a valid confidence interval with the stated length, completing the proof. More explicitly, for any fixed loading vector \(v\), let \(\hat{u}_v\) and \(\hat{L}_{\rm db}(Z;v)\) denote the quantities in 30 and 29 with \(\xi\) replaced by \(v\). For any prescribed component error probability \(\bar\alpha\in(0,1)\), define \[{\rm CI}_{\rm db}(v) = [\hat{L}_{\rm db}(Z;v)-r_{\rm db}(v),\, \hat{L}_{\rm db}(Z;v)+r_{\rm db}(v)]\] with \[r_{\rm db}(v) = 1.1\hat{\sigma} \left\{ \frac{\sqrt{\hat{u}_v^\top\hat{\Sigma}\hat{u}_v}}{\sqrt n} z_{1-\bar\alpha/8} + c_\beta C_\xi\|v\|_2\frac{k_u\log p}{n} \right\}.\] The preceding argument, applied with \(\alpha=\bar\alpha\), gives \[\inf_{\theta\in\Theta(k_u)} \mathbb{P}_\theta \left\{ v^\top\beta\in{\rm CI}_{\rm db}(v),\; r_{\rm db}(v)\le R_{\rm db}(v) \right\} \ge 1-\bar\alpha ,\] where one may take \[R_{\rm db}(v) = C_{\rm db}\sigma\|v\|_2 \left( \frac{1}{\sqrt n} + \frac{k_u\log p}{n} \right)\] for a fixed constant \(C_{\rm db}>0\) depending only on the fixed model constants and on \(\bar\alpha\). This verifies the debiased assumption required by 5. ◻

7.3 Proof of the upper bound in Theorem 5↩︎

Proof. The argument follows the same general strategy as the proof of the upper bound under unknown design in 7.2. The key difference is that, when the design covariance matrix \(\Sigma=\Sigma_0\) is known, one can construct a debiased confidence interval for an arbitrary loading vector \(\xi\in\mathbb{R}^p\) with confidence level \(1-\alpha/2\) and length of order \(\sigma\|\xi\|_2/\sqrt{n}\).

Following [7], we employ a data-splitting strategy. Randomly split the sample into two independent halves \(Z^{(1)}=(X^{(1)},Y^{(1)})\) and \(Z^{(2)}=(X^{(2)},Y^{(2)})\) of sizes \(n_1\) and \(n_2\), respectively. Without loss of generality, assume \(n\) is even and \(n_1=n_2=n/2\).

Using the first half of the data \(Z^{(1)}\), we compute the lasso estimator \(\hat{\beta}\) and the noise level estimator \(\hat{\sigma}\) in [cdt: linear estimator,cdt: linear variance]. Consequently, the event \[\widetilde{\mathcal{A}}_0 = \left\{ 1.1\hat{\sigma}>\sigma>0.9\hat{\sigma},\; \|\hat{\beta}-\beta\|_1 \le c_\beta\sigma k_u\sqrt{\frac{\log p}{n_1}},\; \|\hat{\beta}-\beta\|_2 \le C_\beta\sigma\sqrt{\frac{k_u\log p}{n_1}} \right\}\] holds with probability approaching 1, because \(n_1=n/2\) and Conditions 1 and 2 give the same high-probability guarantee, after changing only fixed constants. The result for plug-in intervals in 6 is then applied with \(\mathcal{E}_{\rm pi}=\widetilde{\mathcal{A}}_0\) and \(N=n_1\). Since the estimators are based on the first half of the data \(Z^{(1)}\), the event \(\widetilde{\mathcal{A}}_0\) is independent of the second half of data \(Z^{(2)}\).

In the following, we condition on \(Z^{(1)}\) and assume \(\widetilde{\mathcal{A}}_0\) happens.

We construct the debiased estimator as follows: \[\label{eq:32bias32estimator32known32design} \hat{L}_{0}(Z;\xi)=\xi^\top \hat{\beta}+\frac{1}{n_2}\xi^\top \Sigma_0^{-1}\left(X^{(2)}\right)^\top \left(Y^{(2)} - X^{(2)} \hat{\beta}\right)\tag{43}\] Its estimation error admits the decomposition \[\label{eq:32error32decomp32known32design} \hat{L}_{0}(Z;\xi)-\xi^\top \beta= \underbrace{\left(\xi-\frac{1}{n_2}\left(X^{(2)}\right)^\top X^{(2)} \Sigma_0^{-1}\xi\right)^\top (\hat{\beta}-\beta)}_{\mathrm{I}} + \underbrace{\frac{1}{n_2}\xi^\top \Sigma_0^{-1}\left(X^{(2)}\right)^{\top} \varepsilon^{(2)}}_{\mathrm{II}},\tag{44}\] where \(\varepsilon^{(2)}=Y^{(2)}-X^{(2)}\beta\sim \mathcal{N}(0, \sigma^2 {\mathbf{I}}_{n_2})\).

Bounding \(\mathrm{I}\) in 44 : Let \[q_{i}^\prime=\xi^\top \Sigma_0^{-1}X_{i\cdot}^{(2)}\left(X_{i\cdot}^{(2)}\right)^\top (\hat{\beta}-\beta).\] By 4, \[\begin{align} \|q_{i}^\prime\|_{\psi_1} & \le 2 \left\|\xi^{\top} \Sigma_0^{-1} X_{i \cdot}^{(2)}\right\|_{\psi_2} \left\|\left(X_{i \cdot}^{(2)}\right)^{\top}(\hat{\beta}-\beta)\right\|_{\psi_2} \\ & \le 2\| \Sigma_0^{-1}\xi\| \|(\hat{\beta}-\beta)\| \left\| X_{i \cdot}^{(2)}\right\|_{\psi_2}^2 \le 2 M_1^2 \|\xi\|_2\|\hat{\beta}-\beta\|_2, \end{align}\] which suggests \(q_{i}^\prime\) is sub-exponential. Consequently, there is a constant \(C_1>0\) such that conditional on \(\widetilde{\mathcal{A}}_0\), it holds that \[\|q_{i}^\prime -\mathbb{E}q_{i}^\prime\|_{\psi_1} \le C_1\sigma \|\xi\|_2\sqrt{\frac{k_u\log p}{n_1}}.\] Furthermore, it holds that \(\mathbb{E}q_{i}^\prime=\xi^\top (\hat{\beta}-\beta)\).

For any \(c>0\), applying Lemma 3 to \(n_2^{-1}\sum_{i}(q_{i}^\prime-\mathbb{E}q_{i}^\prime)\) with \(t=c\sigma\|\xi\|_2\sqrt{\frac{k_u\log p}{ n_1 n_2 }}>0\), we have \[\begin{align} \mathbb{P}\left(\bigg|\left(\xi-\frac{1}{n_2}\left(X^{(2)}\right)^\top X^{(2)} \Sigma_0^{-1} \xi\right)^\top (\hat{\beta}-\beta)\bigg|\ge t\right) \le 2\exp\left(-c_0 \min\left(\frac{c^2}{C_1^2}, \frac{c \sqrt{n_2}}{C_1}\right)\right). \end{align}\] We fix a value of \(c\) large enough such that \(2\exp(-c_0 c^2/C_1^2)\le \alpha/12\).

Since \(n_1 = n_2 = n/2\) under the data splitting and \(n \gtrsim k_u\log p\), we have \(n_1 \gtrsim k_u\log p\). Hence, \[t=c\sigma\|\xi\|_2\sqrt{\frac{k_u\log p}{n_1n_2}} =\frac{c\sigma\|\xi\|_2}{\sqrt{n_2}}\sqrt{\frac{k_u\log p}{n_1}} \le \frac{C_2\sigma\|\xi\|_2}{\sqrt{n_2}},\] where \(C_2:=c\sup_n\sqrt{\frac{k_u\log p}{n_1}}\) is a finite constant. Therefore, the event \[\begin{align} \mathcal{A}_3 =\Bigg\{ \bigg|\Big(\xi-\frac{1}{n_2}(X^{(2)})^\top X^{(2)} \Sigma_0^{-1}\xi\Big)^\top (\hat{\beta}-\beta)\bigg| \le C_2\|\xi\|_2\frac{\sigma}{\sqrt{n_2}} \Bigg\} \end{align}\] holds with probability at least \(1-\alpha/12\) when \(n_2\ge c^2/C_1^2\).

Bounding \(\mathrm{II}\) in 44 : Since \(\xi^\top \Sigma_0^{-1}X_{i\cdot}^{(2)}\mid \varepsilon^{(2)}\stackrel{{\rm i.i.d.}}{\sim }{\mathcal{N}}(0,\xi^\top \Sigma_0^{-1}\xi),i=1,\cdots,n_2\), we have \[\frac{1}{n_2}\xi^\top \Sigma_0^{-1}\left(X^{(2)}\right)^{\top} \varepsilon^{(2)}\mid \varepsilon^{(2)} \stackrel{{\rm i.i.d.}}{\sim} \mathcal{N}\left(0, \frac{\xi^\top \Sigma_0^{-1}\xi}{n_2^2}\|\varepsilon^{(2)}\|_2^2 \right).\] Note that \(\xi^\top \Sigma_0^{-1}\xi\le M_1\|\xi\|_2^2\) by the assumption on the eigenvalues of \(\Sigma_0\) and \(\|\varepsilon^{(2)}\|_2^2/\sigma^2\sim \chi^2(n_2)\). By Lemma 1, we have, with probability at least \(1-\alpha/12\), that \[\mathcal{A}_4=\left\{\bigg|\frac{1}{n_2}\xi^\top \Sigma_0^{-1}\left(X^{(2)}\right)^{\top} \varepsilon^{(2)}\bigg|\le C_3\|\xi\|_2\frac{\sigma}{\sqrt{n_2}}\right\}\] holds for some constant \(C_3>0\) and \(n\) sufficiently large.

Synthesis. The event \(\widetilde{\mathcal{A}}_0\cap\mathcal{A}_3\cap\mathcal{A}_4\) holds with probability at least \(1-\alpha/2\), and on this event, \[|\hat{L}_{0}(Z;\xi)-\xi^\top\beta| \le (C_2+C_3)\|\xi\|_2\frac{\sigma}{\sqrt{n_2}} \le 1.1(C_2+C_3)\|\xi\|_2\frac{\hat{\sigma}}{\sqrt{n_2}}.\] This yields a valid \((1-\alpha/2)\)-level debiased confidence interval of length \(O(\sigma\|\xi\|_2/\sqrt{n})\), completing the proof for the debiased confidence interval. The final step makes use of 5 in the same way as in 7.2, so we omit the details. ◻

7.4 Proof of the upper bound in Theorem 4↩︎

Proof. We follow the argument at the beginning of 7.3 using the same balanced data split with \(n_1=n_2=n/2\) and the same plug-in interval result. To apply 5, the only difference is a new debiased confidence interval.

In particular, we need to prove that for a general loading vector \(\xi\in\mathbb{R}^p\), we can construct a \((1-\alpha/2)\)-level debiased confidence interval with length of order \[\label{eq:upper-spike-debias-rate} \sigma\left( \frac{\|\xi\|_2}{\sqrt{n}} + \sqrt{\sum_{j\le k_u}\xi_j^2}\, \frac{k_u\log p}{n} \right).\tag{45}\]

Using the first half of the sample \(Z^{(1)}\), we compute the lasso estimator \(\hat{\beta}\), the noise level estimator \(\hat{\sigma}\), and an estimator \(\hat{\Omega}\) of the precision matrix \(\Omega=\Sigma^{-1}\) such that the event \[\begin{align} \widetilde{\mathcal{A}}_0 = \Big\{& 1.1\hat{\sigma}>\sigma>0.9\hat{\sigma},\; \|\hat{\beta}-\beta\|_1 \le c_\beta\sigma k_u\sqrt{\tfrac{\log p}{n_1}},\; \|\hat{\beta}-\beta\|_2 \le C_\beta\sigma\sqrt{\tfrac{k_u\log p}{n_1}},\\ & \|\hat{\Omega}-\Omega\|_2 \le C_\Omega\sqrt{\tfrac{k_u\log p}{n_1}},\; \hat{\Omega}\in\Pi_0(k_u,p) \Big\} \end{align}\] holds with probability tending to one.

Lemma 8. If \(\Sigma\in\Pi_0(k_u,p)\) and \(n\ge c\,k_u\log p\) for a sufficiently large constant \(c>0\), then there exists an estimator \(\hat{\Omega}\) such that \[\|\hat{\Omega}-\Omega\|_2 \le C_\Omega\sqrt{\frac{k_u\log p}{n}},\] with probability tending to one, where \(C_\Omega\) is a constant. Moreover, \(\hat{\Omega}\) coincides with \({\mathbf{I}}_p\) outside an index set of cardinality at most \(k_u\).

8 is proven in 9.2.

Using \(\hat{\Omega}\) and the second half of the sample \(Z^{(2)}\), define \[\label{eq:32debiased32center32np32hard} \hat{L}_0(Z;\xi) = \xi^\top\hat{\beta} + \frac{1}{n_2}\xi^\top\hat{\Omega} (X^{(2)})^\top \bigl(Y^{(2)}-X^{(2)}\hat{\beta}\bigr).\tag{46}\] Its error admits the decomposition \[\label{eq:32error32decomp32np32hard} \begin{align} \hat{L}_0(Z;\xi)-\xi^\top\beta =&\; \underbrace{ \Big(\xi-\tfrac{1}{n_2}(X^{(2)})^\top X^{(2)}\Omega\xi\Big)^\top (\hat{\beta}-\beta) }_{\mathrm{I}} + \underbrace{ \tfrac{1}{n_2}\xi^\top\hat{\Omega}(X^{(2)})^\top\varepsilon^{(2)} }_{\mathrm{II}}\\ &- \underbrace{ \tfrac{1}{n_2}\xi^\top(\hat{\Omega}-\Omega) (X^{(2)})^\top X^{(2)}(\hat{\beta}-\beta) }_{\mathrm{III}}, \end{align}\tag{47}\] where \(\varepsilon^{(2)}=Y^{(2)}-X^{(2)}\beta\sim\mathcal{N}(0,\sigma^2 I_{n_2})\).

Terms \(\mathrm{I}\) and \(\mathrm{II}\) are controlled as in 7.3, with high probability. The only difference for term \(\mathrm{II}\) is that \(\Omega=\Sigma_0^{-1}\) is replaced by \(\widehat\Omega\). This replacement does not affect the argument, since the analysis only requires spectral-norm control of the precision matrix. Indeed, on the event \(\widetilde{\mathcal{A}}_0\), we have \[\|\widehat\Omega-\Omega\|_2 \lesssim \sqrt{\frac{k_u\log p}{n}},\] which implies that \(\|\widehat\Omega\|_2\) is uniformly controlled. Hence the same proof as in 7.3 applies, yielding a bound of order \({\sigma\|\xi\|_2}/{\sqrt n}.\)

For term \(\mathrm{III}\), write \[q_i^{\prime\prime} = \xi^\top(\hat{\Omega}-\Omega) X_{i\cdot}^{(2)}(X_{i\cdot}^{(2)})^\top(\hat{\beta}-\beta), \qquad i=1,\ldots,n_2.\] Then \[\mathbb{E} q_i^{\prime\prime} = \xi^\top(\hat{\Omega}-\Omega)\Sigma(\hat{\beta}-\beta) \le \|\xi^\top(\hat{\Omega}-\Omega)\|_2\, \|\hat{\beta}-\beta\|_2.\] Moreover, \[\|q_i^{\prime\prime}\|_{\psi_1} \le 2\|\xi^\top(\hat{\Omega}-\Omega)X_{i\cdot}^{(2)}\|_{\psi_2} \|(X_{i\cdot}^{(2)})^\top(\hat{\beta}-\beta)\|_{\psi_2} \lesssim \|\xi^\top(\hat{\Omega}-\Omega)\|_2 \|\hat{\beta}-\beta\|_2.\]

On \(\widetilde{\mathcal{A}}_0\), \(\|\hat{\beta}-\beta\|_2\lesssim \sigma\sqrt{\tfrac{k_u\log p}{n}}\). Since \(\hat{\Omega}-\Omega\) is supported on at most \(2k_u\) rows and columns, \[\|\xi^\top(\hat{\Omega}-\Omega)\|_2 \le \sqrt{\sum_{j\le 2k_u}\xi_j^2}\, \|\hat{\Omega}-\Omega\|_2 \lesssim \sqrt{\sum_{j\le k_u}\xi_j^2} \sqrt{\tfrac{k_u\log p}{n}}.\] Hence \[\|q_i^{\prime\prime}\|_{\psi_1} \lesssim \sigma\sqrt{\sum_{j\le k_u}\xi_j^2}\, \frac{k_u\log p}{n}.\]

Applying Lemma 3 to \(n_2^{-1}\sum_i(q_i^{\prime\prime}-\mathbb{E} q_i^{\prime\prime})\) with \(t=c\sigma\sqrt{\sum_{j\le k_u}\xi_j^2}\frac{k_u\log p}{n}\) for sufficiently large \(c\), we obtain \[\left| \frac{1}{n_2}\xi^\top(\hat{\Omega}-\Omega) (X^{(2)})^\top X^{(2)}(\hat{\beta}-\beta) \right| \lesssim \sigma\sqrt{\sum_{j\le k_u}\xi_j^2}\, \frac{k_u\log p}{n}\] with high probability.

Combining the bounds for \(\mathrm{I}\)-\(\mathrm{III}\) yields the claimed rate in 45 .

We now complete the proof of the upper bound claimed in 4 by applying 5 with \[\mathcal{P}_u := \Bigl\{ \theta=(\beta,\Sigma,\sigma)\in\Theta(k_u): \Sigma\in\Pi_0(k_u,p) \Bigr\}.\] The condition on the plug-in interval follows from 6 with \(N=n_1\). The preceding debiased construction, with \(v\) in place of \(\xi\), verifies the condition on the debiased interval over \(\mathcal{P}_u\). For \(v=\xi_m^{(1)}\), the debiased radius is bounded by \[R_{\rm db}^{\rm spike}(\xi_m^{(1)}) \lesssim \sigma\left\{ \frac{(\sum_{j\le m}\xi_j^2)^{1/2}}{\sqrt n} + \nu_2\frac{k_u\log p}{n} \right\},\] since \(\sum_{j\le k_u}(\xi_m^{(1)})_j^2\le \sum_{j\le k_u}\xi_j^2=\nu_2^2\). Moreover, \[\mathcal{P}_u\cap\Theta(k_u;\xi,t_0) = \Theta^{\text{spike}}(k_u;\xi,t_0), \qquad \Theta_{\pm\tau}^{\text{spike}}(k;\xi,t_0)\subseteq \mathcal{P}_u\] for \(1\le k\le k_u\). Thus 5 gives, for every fixed deterministic \(m\),

\[\tau_{\rm adap}^{\rm spike}(k_u,k;\xi) \lesssim \frac{\left(\sum_{j\le m}\xi_j^2\right)^{1/2}}{\sqrt n} + |\xi_{m+1}|k_u\sqrt{\frac{\log p}{n}} + \nu_2\frac{k_u\log p}{n},\] where we have used \(\sigma\le M_2\) for \(\theta\in \Theta(k)\). Optimizing over \(m\) and using 2 gives \[\inf_{0\le m\le p} \left\{ \frac{\left(\sum_{j\le m}\xi_j^2\right)^{1/2}}{\sqrt n} + |\xi_{m+1}|k_u\sqrt{\frac{\log p}{n}} \right\} \lesssim_{\log} \frac{\nu_1}{\sqrt n}.\] Thus, we have \[\tau_{\rm adap}^{\rm spike}(k_u,k;\xi) \lesssim_{\log} \frac{\nu_1}{\sqrt n} + \nu_2\frac{k_u\log p}{n}.\] This completes the proof. ◻

8 Proof of the lower bounds↩︎

In this section, we present the proofs of the lower bounds stated in [thm: hypothesis,thm: computational lower bound,thm: reduction,thm: hypothesis design]. Our proof strategy largely follows the framework developed in [8], [10]. Specifically, for a given candidate parameter under the alternative hypothesis, we construct a prior over the null space and show that the corresponding \(\chi^2\)-divergences between their induced mixture distribution and that of the alternative parameter are sufficiently small. This implies that the power of any valid test must necessarily be limited, thereby establishing the desired lower bound.

We summarize below several key technical tools used in the proof. For a probability measure \(\pi\) over the parameter space \(\Theta(k_u)\), we denote \[\mathbb{P}^n_\pi=\int \mathbb{P}^n_\theta\,{\rm d}\,\pi(\theta).\]

For two probability measures \(P, Q\) defined on the same measurable space \((\mathcal{X}, \mathcal{U})\), we define: \[\mathrm{TV}(P,Q) \;=\; \sup_{B \in \mathcal{U}} \big| P(B) - Q(B) \big| ,\] the total variation distance and \[\chi^2(P \,\|\, Q) \;=\; \int \!\left( \frac{dP}{dQ} - 1 \right)^{\!2} dQ ,\] the chi-square divergence.

Lemma 9. For any test \(\psi_*\) and a probability measure \(\pi\) over the parameter space \(\Theta(k_u)\), we have \[\left|\mathbb{E}_{\pi}\psi_* -\mathbb{E}_{\theta_*}\psi_*\right|\le {\rm TV}(\mathbb{P}^n_\pi, \mathbb{P}^n_{\theta_*})\le \frac{1}{2}\sqrt{\chi^2(\mathbb{P}^n_\pi \,\|\, \mathbb{P}^n_{\theta_*})}.\]

8.1 Proof of the lower bound in Theorem 1↩︎

There are two quantities in the lower bound of Theorem 1. The quantity involving \(\nu_1\) appears in the lower bound of 5 and will be proved in 8.2. In this section, we focus on proving the lower bound involving \(\nu_2\).

Proof. Since \(k_u\to\infty\), we assume \(k_u\ge 4\).

Following the definition of \(\tau_{\mathrm{adap}}(k_u, k; \xi, t_0)\) in 10 , we aim to show that for any constant \(c > 0\), there exists a constant \(c' > 0\) such that for \(\tau = c' \nu_2 k_u \tfrac{\log p}{n}\) and for any test \(\psi\), it holds that \[\label{eq:32lower32dense32lower32bound} \mathbb{E}_{\theta_*}\psi \le \sup_{\theta \in \Theta(k_u; \xi, t_0)} \mathbb{E}_{\theta}\psi + c,\tag{48}\] where \(\theta_* = (\beta_*, \Sigma_*, \sigma_*) \in \Theta_{\pm\tau}(k; \xi, t_0)\) is a fixed alternative point to be specified later.

According to 9, it suffices to construct a point \(\theta_* = (\beta_*, \Sigma_*, \sigma_*) \in \Theta_{\pm\tau}(k; \xi, t_0)\) and a prior distribution \(\pi\) over \(\Theta(k_u; \xi, t_0)\) such that \[\label{eq:32simpe32target32for32lower32bound} \chi^2(\mathbb{P}^n_\pi \,\|\, \mathbb{P}^n_{\theta_*}) \le 4c^2 .\tag{49}\]

To facilitate the construction, we introduce a transformation of the parameter space. For any \(k \ge 1\) and any \(t_0,\tau\), choose an index \(j_0\in \operatorname{supp}(\xi)\) and set \[\beta_0 = \frac{t_0-\tau}{\xi_{j_0}} e_{j_0} \in \mathbb{R}^p,\] with the convention \(\beta_0=0\) when \(t_0-\tau=0\) and \(e_{j_0}\) is the \(j_0\)-th standard basis. Then \(\|\beta_0\|_0\le 1\le k\) and \(\xi^\top\beta_0=t_0-\tau\). Define the transition mapping \(\varphi\) on the parameter space as \[\label{eq:32lower32bound32transform32on32theta} \varphi(\beta, \Sigma, \sigma) = (\beta - \beta_0, \Sigma, \sigma).\tag{50}\] Its inverse is given by \(\varphi^{-1}(\tilde{\beta}, \Sigma, \sigma) = (\tilde{\beta} + \beta_0, \Sigma, \sigma).\) We have the following properties:

  • When changing the parameter from \(\theta\) to \(\varphi(\theta)\), the linear functional changes from \(L(\beta;\xi)\) to \(L(\beta;\xi)-t_0+\tau\).

  • When changing the parameter from \(\theta\) to \(\varphi(\theta)\), the induced transformation on the data and noise is \[(Y,X,\varepsilon)\mapsto (Y-X\beta_0,X,\varepsilon),\] which is bijective. Therefore \(\varphi\) preserves the total variation and \(\chi^2\)-divergence between the induced distributions.

  • The pre-image of the point \((0,\Sigma,\sigma)\) is \((\beta_0,\Sigma,\sigma)\) and belongs to \(\Theta_{\pm\tau}(k;\xi,t_0)\).

  • Since \(k_u\ge 4\), the pre-image of \(\Theta(k_u/2;\xi,\tau)\) is contained in \(\Theta(k_u;\xi,t_0)\).

Therefore, it suffices to establish 49 in the \(\varphi\)-transformed space with \(\beta_* = 0\) and \(\pi\) supported on \(\Theta(k_u/2; \xi, \tau)\). We can then apply the inverse transformation \(\varphi^{-1}\) to obtain the desired construction in the pre-transformed parameter space.

Let \(p_1 = \lfloor k_u/4\rfloor\) and \(p_2=p-\lfloor k_u/4\rfloor\). Let \(S_1 = [p_1]\) and \(S_2 = [p] \setminus S_1\); \(p_1\) and \(p_2\) are the sizes of \(S_1\) and \(S_2\), respectively. Under the Gaussian design model, \(Z_i=(Y_i,X_{i\cdot})\in \mathbb{R}^{p+1}\) follows a joint Gaussian distribution with mean \(0\). Let \(\Sigma^z\) denote the covariance of \(Z_i\). Decompose \(\Sigma^{z}\) into blocks \(\begin{pmatrix} \Sigma_{yy}^{z}& \left(\Sigma_{xy}^{z}\right)^{\top}\\ \Sigma_{xy}^{z}& \Sigma_{xx}^{z} \end{pmatrix},\) where \(\Sigma_{yy}^{z}\), \(\Sigma_{xx}^{z}\) and \(\Sigma_{xy}^{z}\) denote the variance of \(y_i\), the variance of \(X_{i\cdot}\) and the covariance of \(y_i\) and \(X_{i\cdot}\), respectively. We define the function \(h : \Sigma^z \rightarrow \theta=\left(\beta,\Sigma,\sigma\right)\) as \[h(\Sigma^z)=\left( \left(\Sigma_{xx}^{z}\right)^{-1}\Sigma_{xy}^{z}, \Sigma_{xx}^{z}, \sqrt{\Sigma_{yy}^{z}-\left(\Sigma_{xy}^{z}\right)^{\top}\left(\Sigma_{xx}^{z}\right)^{-1}\Sigma_{xy}^{z}} \right).\] The function \(h\) is bijective and its inverse mapping \(h^{-1}: \theta=\left(\beta,\Sigma,\sigma\right) \rightarrow \Sigma^z\) is \[h^{-1}\left(\left(\beta,\Sigma,\sigma\right)\right)=\begin{pmatrix} \beta^{\top}\Sigma\beta+\sigma^2& \beta^{\top}\Sigma\\ \Sigma\beta& \Sigma\end{pmatrix}.\]

Step 1: Construct the least favorable prior. We set the alternative point as \[\theta_*=(\beta_*={\mathbf{0}}_{p}, \Sigma_*={\mathbf{I}}_{p\times p}, \sigma_*=M_2/2)\] and we have \[\label{eq:32alter32matrix} \Sigma_*^z=h^{-1}(\theta_*)=\left(\begin{array}{c|c|c} \sigma_*^2&{\mathbf{0}}_{1\times p_1}& {{\mathbf{0}}}_{1\times p_2}\\ \hline {\mathbf{0}}_{p_1\times 1}&{\mathbf{I}}_{p_1\times p_1}&{\mathbf{0}}_{p_1\times p_2}\\ \hline {\mathbf{0}}_{p_2\times 1}&{\mathbf{0}}_{p_2\times p_1}&{\mathbf{I}}_{p_2\times p_2} \end{array}\right)\tag{51}\] Let \(I\) be chosen randomly and uniformly from all subsets of \([p_2]\) with size \(p_1\). We consider the random vectors \(\boldsymbol{\delta}_2\in \mathbb{R}^{p_2}\) defined coordinate-wise as \[\label{eq:32moderate32sparse32delta} (\boldsymbol{\delta}_2)_j=c_1\mathrm{sign}(\xi_{p_1+j})\sqrt{\frac{\log p}{n}}\mathbf{1}\left\{j\in I\right\},\quad \forall j\in [p_2],\tag{52}\] where \(c_1>0\) is a small positive constant to be specified later. We further define a fixed vector \(\boldsymbol{\delta}_1\in \mathbb{R}^{p_1}\) by \[(\boldsymbol{\delta}_1)_j=-\frac{\xi_j}{\sqrt{\sum_{i \in S_1} \xi_i^2}},\quad\forall j\in S_1.\]

Given \(\boldsymbol{\delta}_2\), we can construct the following corresponding covariance matrix: \[\label{eq:32null32point32mat} \Sigma^z= \left( \begin{array}{c|c|c} \sigma_*^2 & {\mathbf{0}}_{1\times p_1}&\kappa\boldsymbol{\delta}_2^\top \\ \hline {\mathbf{0}}_{p_1\times 1}& {\mathbf{I}}_{p_1\times p_1} & \boldsymbol{\delta}_1\boldsymbol{\delta}_2^\top \\ \hline \kappa\boldsymbol{\delta}_2& \boldsymbol{\delta}_2\boldsymbol{\delta}_1^\top & {\mathbf{I}}_{p_2\times p_2} \end{array} \right)\stackrel{\triangle}{=}g_1(\boldsymbol{\delta}_2),\tag{53}\] where \(\kappa = \kappa_1(\boldsymbol{\delta}_2) \in \mathbb{R}\) is a function of \(\boldsymbol{\delta}_2\) defined such that \(\Sigma^z\) corresponds to a point in the transformed null space, i.e., \(h(\Sigma^z) \in \Theta(k_u/2; \xi, \tau)\). We will construct the prior in a way such that this \(\kappa\) exists a.s. Furthermore, since the null space imposes a linear constraint on \(\beta\), whenever such a \(\kappa\) exists, it is uniquely determined since \(\tau>0\).

Denote by \(\pi_1\) the distribution of \(\boldsymbol{\delta}_2\) defined in 52 and \(\pi_{}\) the induced prior of \(\pi_1\) under the mapping \(h\circ g_1\). The following lemma shows that \(\pi_{}\) is supported on \(\Theta(k_u/2; \xi, \tau)\), which is the least favorable prior we aim to construct.

Lemma 10. If we choose \(c_1\) sufficiently small, then there exists a constant \(c_2>0\) such that the induced prior \(\pi_{}\) is supported on \(\Theta(k_u/2;\xi,\tau)\), where \(\tau = c_2 \nu_2 \frac{k_u \log p}{n}\), and moreover the associated coefficient \(\kappa=\kappa_1(\boldsymbol{\delta}_2)\) satisfies \(0<\kappa\le 1\) \(\pi_1\)-almost surely.

The proof is given in 10.2.

Step 2: Control the \(\chi^2\)-divergence. For the \(\chi^2\)-divergence, by Fubini’s theorem, we have \[\label{eq:32chi32square32decomp} \begin{align} \mathbb{E}_{\theta_*}\left(\frac{\,{\rm d}\,\mathbb{P}^n_{\pi_{}}}{\,{\rm d}\,\mathbb{P}^n_{\theta_*}}-1\right)^2&=\mathbb{E}_{\theta_*}\left[\left(\frac{\,{\rm d}\,\mathbb{P}^n_{\pi_{}}}{\,{\rm d}\,\mathbb{P}^n_{\theta_*}}\right)^2\right]-1 =\mathbb{E}_{(\theta,\tilde{\theta})\sim \pi_{}\otimes \pi_{}}\int_{\mathbb{R}^n}\frac{\,{\rm d}\,\mathbb{P}^n_{\theta}\,{\rm d}\,\mathbb{P}^n_{\tilde{\theta}}}{\,{\rm d}\,\mathbb{P}^n_{\theta_*}}-1. \end{align}\tag{54}\]

The following lemma provides an upper bound for the integral in 54 , and its proof is given in 10.3.

Lemma 11. For any pair of parameters \((\boldsymbol{\delta}_2,\tilde{\boldsymbol{\delta}}_2)\) over the support of \({\pi}_1\) constructed in 10, let \(\theta = h(g_1( \boldsymbol{\delta}_2))\) and \(\tilde{\theta} = h(g_1( \tilde{\boldsymbol{\delta}}_2))\). If \(c_1\) is chosen sufficiently small, there exists some constants \(c_3>0\) only depending on \(\sigma_*\) such that \[\mathbb{E}_{\theta_*}\left(\frac{\,{\rm d}\,\mathbb{P}^n_{\pi_{}}}{\,{\rm d}\,\mathbb{P}^n_{\theta_*}}-1\right)^2 \le \mathbb{E}_{(\boldsymbol{\delta}_2,\tilde{\boldsymbol{\delta}}_2)\sim \pi_1\otimes \pi_1} \exp\left(c_3 n\boldsymbol{\delta}_2^\top \tilde{\boldsymbol{\delta}}_2\right)-1.\]

The following lemma is useful in controlling the right-hand side of the above equation in 11, and its proof is given in 10.4.

Lemma 12. Let \(J\) be a \(\mathrm{Hypergeometric}(p, k, k)\) variable with \[\mathbb{P}(J = j) = \frac{\binom{k}{j} \binom{p - k}{k - j}}{\binom{p}{k}}.\] If \(k\le p^\gamma\) for some constant \(\gamma\in [0,1/2)\), then for constant \(c\in (0,1-2\gamma)\), we have \[\label{eq:32hypergeo32bound} \lim_{p\to\infty}\mathbb{E}\big[\exp(c\log p \cdot J)\big]=1.\tag{55}\]

Let \(I\) and \(\tilde{I}\) denote the supports of \(\boldsymbol{\delta}_2\) and \(\tilde{\boldsymbol{\delta}}_2\), respectively. By the construction of \(\boldsymbol{\delta}_2\) in 52 , both \(I\) and \(\tilde{I}\) are independently and uniformly sampled from all subsets of \([p_2]\) of size \(p_1\). Consequently, the intersection size \(|I \cap \tilde{I}|\) follows a \(\mathrm{Hypergeometric}(p_2, p_1, p_1)\) distribution. Since \[\boldsymbol{\delta}_2^\top \tilde{\boldsymbol{\delta}}_2 = c_1^{2} \frac{\log p}{n} \, |I \cap \tilde{I}|,\] we may apply 12. Recall the sparsity condition \(p_1 \lesssim k_u \le p_2^\gamma\) for some \(\gamma \in (0, 1/2)\) from 3. [lem: chi square integral 1,lem: hypergeometric] together imply that \[\begin{align} \mathbb{E}_{\theta_*}\!\left(\frac{\,{\rm d}\,\mathbb{P}_{\pi_{}}^n}{\,{\rm d}\,\mathbb{P}_{\theta_*}^n} - 1\right)^{2} &\le \mathbb{E}\exp\!\left( c_3 c_1^{2} \log p \, |I \cap \tilde{I}| \right) - 1 = o(1), \end{align}\] where \(c_3\) is the constant from 11 and \(c_1\) is chosen sufficiently small such that \(c_3c_1^2<1-2\gamma\). ◻

8.2 Proof of the lower bound in Theorem 5↩︎

Proof. It suffices to prove the lower bound for the identity covariance case. Indeed, if \(\Sigma_0\) is diagonal and satisfies the eigenvalue condition in 4 , then a coordinatewise rescaling transforms the design covariance to \(\mathbf{I}_p\). Specifically, with \(\widetilde{X}=X\Sigma_0^{-1/2}\), \(\widetilde{\beta}=\Sigma_0^{1/2}\beta\), and \(\widetilde{\xi}=\Sigma_0^{-1/2}\xi\), we have \[\xi^\top\beta=\widetilde{\xi}^\top\widetilde{\beta}, \qquad \|\widetilde{\beta}\|_0=\|\beta\|_0,\] and the bounded eigenvalues of \(\Sigma_0\) imply that the loading-profile quantities for \(\widetilde{\xi}\) and \(\xi\) are equivalent up to constants. Therefore, the diagonal known-covariance case reduces to the identity covariance case up to constant factors. We focus below on \(\Sigma_0=\mathbf{I}_p\).

We follow the same argument in 8.1 and consider the transition mapping \(\varphi\) defined in 50 . In the transformed alternative space, we fix \(\theta_* = (\mathbf{0}_{p\times 1}, \mathbf{I}_{p\times p}, \sigma_*)\) with the corresponding matrix \(\Sigma_*^{z}\) defined in 51 . Suppose that we can construct a prior \(\pi\) satisfying \[\pi\!\left( \Theta(k_u/2; \xi, \tau) \right) \ge 1 - \frac{c}{2}, \qquad\text{and}\qquad \chi^{2}\!\left( \mathbb{P}_{\pi}^{n} \,\|\, \mathbb{P}_{\theta_*}^{n} \right) \le c^{2},\] where \(\tau=c\nu_1/\sqrt{n}\) for some constant \(c > 0\) to be specified later. Define \(\tilde{\pi}\) as the restriction of \(\pi\) to a subset of \(\Theta(k_u/2;\xi,\tau)\). Then, we have \[\label{eq:32restriction32TV} \begin{align} \mathrm{TV}\!\left(\mathbb{P}_{\tilde{\pi}}^{n},\, \mathbb{P}_{\theta_*}^{n}\right) &\le \mathrm{TV}\!\left(\mathbb{P}_{\pi}^{n},\, \mathbb{P}_{\theta_*}^{n}\right) + \mathrm{TV}\!\left(\mathbb{P}_{\pi}^{n},\, \mathbb{P}_{\tilde{\pi}}^{n}\right) \\ &\le \mathrm{TV}(\tilde{\pi}, \pi) + \frac{1}{2}\sqrt{ \chi^{2}\!\left( \mathbb{P}_{\pi}^{n} \,\|\, \mathbb{P}_{\theta_*}^{n} \right) } \\ &\le c, \end{align}\tag{56}\] and the lower bound can be established by 9 and the properties of \(\varphi\).

In the following, we focus on the construction of \(\pi\). We first assume that \(\nu_1 \ge C_4 |\xi_1|\) for some sufficiently large constant \(C_4>0\). The complementary case \(\nu_1 < C_4 |\xi_1|\) is technically simpler; we therefore defer its analysis to the end of this section.

Step 1: Construct the least favorable prior. We define \(\pi_3\) to be the probability measure of \(\boldsymbol{\delta}\in \mathbb{R}^p\) constructed as follows. Let \(k_\xi=\|\xi\|_0\) and let \(\boldsymbol{\delta}_j = 0\) for all \(j \notin [k_\xi]\). For each \(j \in [k_\xi]\), we specify the coordinates independently as \[\begin{align} \boldsymbol{\delta}_j &= \frac{c_5}{\sqrt{n}}\, b^{(1)}_j \gamma^{(1)}_j, \end{align}\] where \(b^{(1)}_j={\rm Ber}(q^{(1)}_j)\), \[q^{(1)}_j=c_4\cdot\frac{ |\xi_j|\exp(-\lambda^2/\xi_j^2)}{\sqrt{\sum_{i=1}^p \xi_i^2\exp(-\lambda^2/\xi_i^2)}},\quad\text{and}\quad \gamma^{(1)}_j=\left\{\begin{array}{lc} {\rm sign}(\xi_j)& j\le j_1,\\ {\lambda}/{\xi_j}& j> j_1, \end{array}\right.\] where \(c_4,c_5 > 0\) are some constants to be specified later, \(\lambda\) is defined in 17 , and \(j_1=\max\left\{j\in [p]:|\xi_j|\ge \lambda\right\}\).

Given \(\boldsymbol{\delta}\), we construct the covariance matrix as \[\label{eq:32null32point32mat322} \Sigma^z= \left( \begin{array}{c|c} \sigma_*^2 & \kappa\boldsymbol{\delta}^\top \\ \hline \kappa\boldsymbol{\delta}& {\mathbf{I}}_{p\times p} \end{array} \right)\stackrel{\triangle}{=}g_2(\boldsymbol{\delta}),\tag{57}\] where \(\kappa\) is a value to be chosen so that either \(\Sigma^z\) is the identity or the above matrix corresponds to a parameter in the transformed null space. Concretely, for any \(\tau>0\), define \[\mathcal{G}_\tau=\{ \boldsymbol{\delta}\in \mathbb{R}^p : \xi^\top \boldsymbol{\delta}\ge \tau, ~\|\boldsymbol{\delta}\|_0\le k_u/2, ~\|\boldsymbol{\delta}\|_2^2\le (\xi^\top \boldsymbol{\delta})^2\sigma_*^2/(2\tau^2)\}.\] Given the value of \(\tau>0\), we set \[\label{eq:32identity32cov32kappa32choice} \kappa= \begin{cases} \frac{\tau}{\xi^\top \boldsymbol{\delta}}& \quad \boldsymbol{\delta}\in \mathcal{G}_\tau, \\ 0 & \quad \text{otherwise}. \end{cases}\tag{58}\] Then \(\kappa\in [0,1]\). Furthermore, when \(\pi_3(\mathcal{G}_\tau)>0\), denote the restriction of \(\pi_3\) on \(\mathcal{G}_\tau\) by \(\pi_3(\cdot\mid G_\tau)\).

Define \(\pi_4\) and \(\widetilde{\pi}_4\) as the pushforward of \(\pi_3\) and \(\pi_3(\cdot\mid G_\tau)\), respectively, under the map \(h\circ g_2\).

The next lemma states that for the case \(\nu_1 \ge C_4 |\xi_1|\), there exists a choice of the constants \(c_4,c_5>0\) and a corresponding \(c_6\) such that \(\mathcal{G}_\tau\) holds with a probability close to 1 and \(\widetilde{\pi}_4\) is supported on the transformed null space \(\Theta(k_u/2;\xi,c_6\nu_1/\sqrt{n})\).

Lemma 13. Consider \(\boldsymbol{\delta}\sim \pi_3\). For any constants \(c>0\) and \(c_5\in (0,1)\), one can choose the constants \(c_4\) sufficiently small, \(C_4\) sufficiently large, and \(c_6\) such that for \(\tau=c_6\nu_1/\sqrt n\), if \(\nu_1\ge C_4|\xi_1|\) and \(k_u\ge C_4\), then \[\pi_3(G_\tau)\ge 1-c/2,\] and \(\widetilde{\pi}_4\) is supported on \(\Theta(k_u/2;\xi,\tau)\).

Step 2: Control the \(\chi^2\)-divergence. The null prior is given by \(\widetilde{\pi}_4\). Since total variation distance cannot increase under a measurable pushforward, we have \[\mathrm{TV}(\widetilde{\pi}_4, \pi_4) \le \mathrm{TV}\left(\pi_3, \pi_3(\cdot\mid \mathbf{G}_\tau) \right) = 1 - \pi_3(\mathbf{G}_\tau) \le c/2,\] where the last inequality follows from 13. By 56 , it remains to show that the \(\chi^2\)-divergence between the mixture distribution induced by \(\pi_4\) and that of \(\theta_*\) is small. Similar to 11, we have the following result.

Lemma 14. For any pair of \(\boldsymbol{\delta}\) and \(\tilde{\boldsymbol{\delta}}\) sampled independently from \({\pi}_3\), let \(\theta = h(g_2( \boldsymbol{\delta}))\) and \(\tilde{\theta} = h(g_2( \tilde{\boldsymbol{\delta}}))\). If \(c_4\) is chosen sufficiently small, there exists some constant \(c_7>0\) only depending on \(\sigma_*\) such that \[\mathbb{E}_{\theta_*}\left(\frac{\,{\rm d}\,\mathbb{P}^n_{\pi_4}}{\,{\rm d}\,\mathbb{P}^n_{\theta_*}}-1\right)^2 \le \mathbb{E}_{(\boldsymbol{\delta},\tilde{\boldsymbol{\delta}})\sim \pi_3\otimes \pi_3} \exp\left(c_7 n\boldsymbol{\delta}^\top \tilde{\boldsymbol{\delta}}\right)-1.\]

For \((\boldsymbol{\delta},\tilde{\boldsymbol{\delta}})\sim \pi_3^{\otimes 2}\), we have \[\forall j\in [k_\xi], \qquad \boldsymbol{\delta}_j\tilde{\boldsymbol{\delta}}_j=\left\{ \begin{array}{cc} c_5^2(\gamma^{(1)}_j)^2/n & \text{w.p. } (q^{(1)}_j)^2,\\ 0 & \text{w.p. } 1-(q^{(1)}_j)^2,\\ \end{array} \right.\] and \(\boldsymbol{\delta}_j\tilde{\boldsymbol{\delta}}_j=0\) for all \(j>k_\xi\). Therefore, we have \[\begin{align} \mathbb{E}_{(\boldsymbol{\delta},\tilde{\boldsymbol{\delta}})\sim \pi_3\otimes \pi_3} \exp\left(c_7 n\boldsymbol{\delta}^\top \tilde{\boldsymbol{\delta}}\right)&=\prod_{j=1}^p \mathbb{E}_{(\boldsymbol{\delta}_j,\tilde{\boldsymbol{\delta}}_j)\sim \pi_3\otimes \pi_3} \exp\left(c_7 n\boldsymbol{\delta}_j \tilde{\boldsymbol{\delta}}_j\right)\\ &=\prod_{j=1}^{k_\xi} \left[1+(q^{(1)}_j)^2\left(\exp\left(c_7 c_5^2(\gamma^{(1)}_j)^2\right)-1\right)\right]\\ &\le \exp\left(\sum_{j=1}^{k_\xi} (q^{(1)}_j)^2\left(\exp\left(c_7 c_5^2(\gamma^{(1)}_j)^2\right)-1\right)\right). \end{align}\] By the definition of \(q^{(1)}_j\) and \(\gamma^{(1)}_j\), we have \[\begin{align} &\sum_{j=1}^{k_\xi} (q^{(1)}_j)^2\left(\exp\left(c_7c_5^2 (\gamma^{(1)}_j)^2\right)-1\right)\\ \le &c_4^2\frac{ \sum_{j\le j_1} \xi_j^2\exp(-2\lambda^2/\xi_j^2)\left(\exp\left(c_7 c_5^2\right)-1\right)+\sum_{j>j_1}\xi_j^2\exp(-2\lambda^2/\xi_j^2)\left(\exp(c_7 c_5^2\lambda^2/\xi_j^2)-1\right)}{\sum_{i=1}^{k_\xi} \xi_i^2\exp(-\lambda^2/\xi_i^2)}. \end{align}\] If we choose \(c_5\) sufficiently small such that \(c_7 c_5^2<\ln 2\), then we have \[\begin{align} \exp(-2\lambda^2/\xi_j^2)\left(\exp\left(c_7 c_5^2\right)-1\right)&\le \exp(-\lambda^2/\xi_j^2),\quad \forall j\le j_1,\\ \exp(-2\lambda^2/\xi_j^2)\left(\exp\left(c_7 c_5^2\lambda^2/\xi_j^2\right)-1\right)&\le \exp(-\lambda^2/\xi_j^2),\quad \forall j>j_1. \end{align}\] Therefore, we have \[\mathbb{E}_{(\boldsymbol{\delta},\tilde{\boldsymbol{\delta}})\sim \pi_3\otimes \pi_3} \exp\left(c_7 n\boldsymbol{\delta}^\top \tilde{\boldsymbol{\delta}}\right)\le \exp\left(c_4^2\right).\] By choosing \(c_4\) sufficiently small such that \(c_4^2 \in (0, \ln(1+c^2))\), we have \(\chi^{2}\!\left( \mathbb{P}_{\pi}^{n} \,\|\, \mathbb{P}_{\theta_*}^{n} \right) \le c^{2}\).

In the remaining case where \(\nu_1 < C_4 |\xi_1|\), we only need to prove that \(\tau_{\mathrm{adap}}(k_u; \xi, t_0) \gtrsim |\xi_1|/\sqrt{n}\). In this case, we choose the point mass prior at \(\theta_o\) with \[\Sigma^z_o=\left(\begin{array}{c|c} \sigma_*^2& \frac{\kappa_0}{\sqrt{n}}\mathrm{sign}(\xi_1){\boldsymbol{e}}_1^\top \\ \hline \frac{\kappa_0}{\sqrt{n}} \mathrm{sign}(\xi_1){\boldsymbol{e}}_1& {\mathbf{I}}_{p\times p} \end{array}\right),\] where \({\boldsymbol{e}}_1\) is the first standard basis vector, i.e., \[\left({\boldsymbol{e}}_1\right)_j=\left\{ \begin{array}{cc} 1,&j=1\\ 0,&j\neq 1 \end{array} \right.,\] and \(\kappa_0>0\) is some constant specified later. It is easy to verify that the above matrix corresponds to a point in the transformed null space with \(\tau=\kappa_0|\xi_1|/\sqrt{n}\). By an argument similar to that in 14, we have \[\chi^2(\mathbb{P}_{\theta_o}^n\|\mathbb{P}_{\theta_*}^n) = \exp(\kappa_0^2)-1.\] Therefore we finish the proof by choosing the constant \(\kappa_0\) sufficiently small. ◻

8.3 Proof of Theorem 2↩︎

Proof. Without loss of generality we assume \(8\le 2k_u<k_{\rm eff}\) and \[\sum_{k_u<j\le k_{\rm eff}}\xi_j^2\ge \sum_{j\le k_u}\xi_j^2.\] In the complementary case, the low-degree lower bound in 2 is an immediate consequence of the statistical lower bound in 1, so no additional proof is needed. Since \(D\lesssim p\) and the standing dimensional assumptions \(\sqrt n/\log p\lesssim k_u \lesssim p^\gamma\) imply \(\log n\lesssim \log p\), we have \(\log(6npD)\lesssim \log p .\) Consequently, \(D\log(6npD)\lesssim D\log p .\) On the other hand, by the definition of \(k_{\mathrm{eff}}\), \[k_{\mathrm{eff}} \le \frac{k_u^2}{D\log p}.\] Therefore, \[\frac{k_u^2}{k_{\mathrm{eff}}} \gtrsim D\log p \gtrsim D\log(6npD).\] Thus the exponent appearing in the Chernoff bound for the event \(\mathcal{C}^c\) is at least of order \(D\log(6npD)\), which is sufficient for the subsequent union bound over degree-\(D\) polynomial terms.

The proof is closely analogous to the argument in 8.1. We again take the alternative point \(\theta_* = (\mathbf{0}_{p\times 1}, \mathbf{I}_{p\times p}, \sigma_*)\) together with the associated matrix \(\Sigma_*^{z}\) defined in 51 . Our strategy is to construct a prior \(\pi\) that places overwhelming probability mass on the set \(\Theta(\lfloor k_u/2\rfloor; \xi, \tau)\) with \(\tau = c\, \nu_3\, k_u \frac{\log p}{n}\) for a constant \(c>0\) to be chosen later. We then show that the low-degree likelihood ratio satisfies \(\mathrm{LD}(D) = 1 + o(1)\) for \(\mathbb{Q}_1 = \mathbb{P}_{\pi}^{n}\) and \(\mathbb{Q}_0 = \mathbb{P}_{\theta_*}^{n}.\) The conclusion of the theorem then follows immediately from 3.

Step 1: Construct the least favorable prior. Let \(S_3 = [ k_{\rm eff}]\), \(S_4 = [p] \setminus [k_{\rm eff}]\), and \(S_5=[k_{\rm eff}]\setminus [k_u]\). Denote by \(p_3\), \(p_4\), and \(p_5\) the sizes of \(S_3\), \(S_4\), and \(S_5\), respectively; that is, \(p_3 = k_{\rm eff}\), \(p_4=p-k_{\rm eff}\), and \(p_5=k_{\rm eff}-k_u\). Let \(I\) be chosen randomly and uniformly from all subsets of \([p_4]\) with size \(\lfloor k_u/4\rfloor\). We consider the random vectors \(\boldsymbol{\delta}_1\in \mathbb{R}^{p_4},\boldsymbol{\delta}_2\in \mathbb{R}^{p_3}\) defined as \[\label{eq:32moderate32sparse32delta322} \left\{ \begin{align} (\boldsymbol{\delta}_1)_j&=c_8\mathrm{sign}(\xi_{p_3+j})\sqrt{\frac{\log p}{n}}\mathbf{1}\left\{j\in I\right\},\quad \forall j\in [p_4], \\ (\boldsymbol{\delta}_2)_j&=-\frac{\sqrt{p_5}}{k_u}\mathrm{sign}(\xi_j) b^{(2)}_j,\quad \forall j\in S_3, \end{align}\right.\tag{59}\] for a small enough constant \(c_8>0\) to be specified later, where \(b^{(2)}_j\) are independent Bernoulli variables with parameters \[q^{(2)}_j=\frac{|\xi_j|}{8\sqrt{\sum_{i\in S_5}\xi_i^2}}\cdot \frac{k_u}{\sqrt{p_5}}\mathbf{1}\left\{j\in S_5\right\},\quad \forall j\in S_3.\] Specifically, we have \(\sum_{i\in S_5}\xi_i^2\ge \sum_{i\le k_u}\xi_i^2\ge k_u\xi_{k_u}^2\) and \(p_5>k_u\) by our assumption. Therefore, we have \[q^{(2)}_j\le \frac{1}{8}\sqrt{\frac{k_u}{p_5}}< \frac{1}{8}, \qquad \forall j\in S_5\] and \(q^{(2)}_j=0\) for all \(j\in S_3\setminus S_5\). This verifies that the above construction of \(q_j^{(2)}\) is well-defined.

Given a pair \((\boldsymbol{\delta}_1, \boldsymbol{\delta}_2)\), we can construct the following corresponding covariance matrix: \[\label{eq:32null32point32mat323} \Sigma^z= \left( \begin{array}{c|c|c} \sigma_*^2 & {\mathbf{0}}_{1\times p_3}&\kappa\boldsymbol{\delta}_1^\top \\ \hline {\mathbf{0}}_{p_3\times 1}& {\mathbf{I}}_{p_3\times p_3} & \boldsymbol{\delta}_2\boldsymbol{\delta}_1^\top \\ \hline \kappa\boldsymbol{\delta}_1& \boldsymbol{\delta}_1\boldsymbol{\delta}_2^\top & {\mathbf{I}}_{p_4\times p_4} \end{array} \right)\stackrel{\triangle}{=}g_3(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2),\tag{60}\] where \(\kappa = \kappa_3(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2) \in \mathbb{R}\) is defined such that the above matrix corresponds to a point in the transformed null space, i.e., \(h(\Sigma^z) \in \Theta(k_u/2; \xi, \tau)\) for some nonzero \(\tau\) specified later. We additionally require \(\kappa\) to be bounded by \(1\). If no such value exists, we set \(\kappa = 0\). This definition is well-posed since the null space imposes a linear constraint on \(\beta\), and whenever such a \(\kappa\) exists, it is uniquely determined.

Denote by \(\pi_5\) the joint distribution of \((\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)\) defined in 59 and \(\pi_6\) the induced prior of \(\pi_5\) under the mapping \(h\circ g_3\). The following lemma shows that \(\pi_6\) is mostly supported on the null space.

Lemma 15. If we choose \(c_8\) sufficiently small, then there exists a constant \(c_9>0\) such that \(\pi_6\) assigns probability \(1-o(1)\) to the transformed null space \(\Theta(\lfloor k_u/2\rfloor;\xi,c_ 9\nu_3\frac{k_u\log p}{n})\) and \(\kappa_3\in (0,1]\).

The proof is given in 10.7.

The low-degree comparison must be made with a prior supported on the null space. For the mapping \[\theta(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2) := h\circ g_3(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2),\] define the validity event \[\mathcal{A}_n = \left\{ (\boldsymbol{\delta}_1,\boldsymbol{\delta}_2): \theta(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2) \in \Theta\left(k_u;\xi,c_9\nu_3\frac{k_u\log p}{n}\right) \;\text{and}\;0<\kappa_3(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)\le 1 \right\}.\] By 15, the exception probability rate \(\varepsilon_n:=\pi_5(\mathcal{A}_n^c)\) satisfies \(\varepsilon_n=o(1)\). Let \(\pi_7\) be the restricted probability measure of \(\pi_5\) on \(\mathcal{A}_n\), i.e., the conditional distribution of \((\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)\) given \(\mathcal{A}_n\). Define \[\pi_8 := \theta_{\#}\pi_7 = (h\circ g_3)_{\#}\pi_7 .\] Then \(\pi_8\) is supported on \[\Theta\left(k_u;\xi,c_9\nu_3\frac{k_u\log p}{n}\right)\] and is the least favorable prior we aim to construct.

We also verify here that the same construction is supported on the sparse signed-spiked covariance class used in 5.1. On the validity event \(\mathcal{A}_n\), the \(X\)-covariance block of \(g_3(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)\) is \[\Sigma(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2) = \begin{pmatrix} \mathbf{I}_{p_3} & \boldsymbol{\delta}_2\boldsymbol{\delta}_1^\top\\ \boldsymbol{\delta}_1\boldsymbol{\delta}_2^\top & \mathbf{I}_{p_4} \end{pmatrix} = \mathbf{I}_p+ \begin{pmatrix} 0 & \boldsymbol{\delta}_2\boldsymbol{\delta}_1^\top\\ \boldsymbol{\delta}_1\boldsymbol{\delta}_2^\top & 0 \end{pmatrix}.\] The perturbation has rank at most two and is supported on \({\rm supp}(\boldsymbol{\delta}_2)\cup(k_{\rm eff}+{\rm supp}(\boldsymbol{\delta}_1))\). The construction gives \(\|\boldsymbol{\delta}_1\|_0=\lfloor k_u/4\rfloor\), and the proof of 15 gives \(\|\boldsymbol{\delta}_2\|_0\le k_u/4\) with probability \(1-o(1)\). Conditioning additionally on this latter event changes the restricted prior only by a \(1+o(1)\) factor, by the same comparison as in 61 . Hence we may take the null prior \(\pi_8\) to be supported on \[\Theta^{\mathrm{spike}}\!\left(k_u;\xi,c_9\nu_3\frac{k_u\log p}{n}\right).\] The alternative point \(\theta_*=(\mathbf{0},\mathbf{I}_p,\sigma_*)\) also has covariance in \(\Pi_0(k,p)\). Therefore the low-degree likelihood-ratio bound proved below applies without change to the sparse signed-spiked null and alternative spaces.

Moreover, for any nonnegative measurable function \(F\), \[\label{eq:restriction-measure-comparison} \begin{align} \mathbb{E}_{\pi_7^{\otimes2}}F &= \frac{ \mathbb{E}_{\pi_5^{\otimes2}} \left[ F\, \mathbf{1}({\mathcal{A}_n})(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2) \mathbf{1}({\mathcal{A}_n})(\tilde{\boldsymbol{\delta}}_1,\tilde{\boldsymbol{\delta}}_2) \right] }{(1-\varepsilon_n)^2} \\ &\le \frac{1}{(1-\varepsilon_n)^2} \mathbb{E}_{\pi_5^{\otimes2}}F = (1+o(1))\mathbb{E}_{\pi_5^{\otimes2}}F . \end{align}\tag{61}\] This comparison will be used below to evaluate bounds under the simpler unrestricted latent prior \(\pi_5\).

Step 2: Control \(\mathrm{LD}(D)\). Set \[\mathbb{Q}_0=\mathbb{P}^n_{\theta_*}, \qquad \mathbb{Q}_1=\mathbb{P}^n_{\pi_8} = \int \mathbb{P}^n_{\theta(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)} \,\pi_7(d\boldsymbol{\delta}_1,d\boldsymbol{\delta}_2).\] For \[L_\theta=\frac{d\mathbb{P}^n_\theta}{d\mathbb{P}^n_{\theta_*}},\] the likelihood ratio of \(\mathbb{Q}_1\) with respect to \(\mathbb{Q}_0\) is \[L = \frac{d\mathbb{Q}_1}{d\mathbb{Q}_0} = \mathbb{E}_{(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)\sim\pi_7} L_{\theta(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)} .\] By linearity of the orthogonal projection onto polynomials of degree at most \(D\), \[L^{\le D} = \mathbb{E}_{(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)\sim\pi_7} L_{\theta(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)}^{\le D}.\] Therefore, \[\begin{align} \mathrm{LD}(D) &= \|L^{\le D}\|_{L^2(\mathbb{Q}_0)}^2 \\ &= \mathbb{E}_{(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2),(\tilde{\boldsymbol{\delta}}_1,\tilde{\boldsymbol{\delta}}_2) \sim\pi_7^{\otimes 2}} \mathbb{E}_{\theta_*} \left[ L_{\theta(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)}^{\le D} L_{\theta(\tilde{\boldsymbol{\delta}}_1,\tilde{\boldsymbol{\delta}}_2)}^{\le D} \right]. \end{align}\]

The expectation above is with respect to the restricted latent prior \(\pi_7^{\otimes2}\). Hence both parameter pairs lie in the validity event \(\mathcal{A}_n\), so the Gaussian covariance matrices are well defined and 16 applies. The following lemma provides some properties for the \(\mathbb{E}_{\theta_*}(L_\theta^{\le D}L_{\tilde{\theta}}^{\le D})\).

Lemma 16. Let \((\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)\) and \((\tilde{\boldsymbol{\delta}}_1,\tilde{\boldsymbol{\delta}}_2)\) belong to the validity event \(\mathcal{A}_n\), and define \[\theta=h(g_3(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)), \qquad \tilde{\theta}=h(g_3(\tilde{\boldsymbol{\delta}}_1,\tilde{\boldsymbol{\delta}}_2)).\] Then \[\mathbb{E}_{\theta_*}\!\left( L_\theta^{\le D}L_{\tilde{\theta}}^{\le D} \right) \le \mathbb{E}_{\theta_*}\!\left( L_\theta L_{\tilde{\theta}} \right),\] and \[\mathbb{E}_{\theta_*}\! \left[L_\theta^{\le D}\right]^2 \le 9(6npD)^{4D}.\]

The proof of 16 is provided in 10.8.

We next split the low-degree second moment according to whether the two latent covariance perturbations have unusually large overlap. Write \[\mathcal{C} = \left\{ (\boldsymbol{\delta}_1, \boldsymbol{\delta}_2, \tilde{\boldsymbol{\delta}}_1, \tilde{\boldsymbol{\delta}}_2) : \boldsymbol{\delta}_2^{\top} \tilde{\boldsymbol{\delta}}_2 \le C_5 \right\},\] where \(C_5\) is an absolute constant to be specified later. Its corresponding event on the parameter space is the pushforward event \[\widetilde{\mathcal{C}} = \left\{ \bigl( \theta(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2), \theta(\tilde{\boldsymbol{\delta}}_1,\tilde{\boldsymbol{\delta}}_2) \bigr): (\boldsymbol{\delta}_1,\boldsymbol{\delta}_2,\tilde{\boldsymbol{\delta}}_1,\tilde{\boldsymbol{\delta}}_2)\in\mathcal{C} \right\},\] where we recall that \(\theta(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2):=h(g_3(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)).\) Since \(\pi_8=\theta_\#\pi_7\), expectations over \(\pi_8^{\otimes2}\) can be evaluated through the latent prior \(\pi_7^{\otimes2}\). In particular, for any nonnegative measurable function \(F\), \[\begin{align} &\mathbb{E}_{\pi_8^{\otimes2}} \left[ F(\theta,\tilde{\theta}) \mathbf{1}(\widetilde{\mathcal{C}})(\theta,\tilde{\theta}) \right] \\ &\qquad = \mathbb{E}_{\pi_7^{\otimes2}} \left[ F\!\left( \theta(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2), \theta(\tilde{\boldsymbol{\delta}}_1,\tilde{\boldsymbol{\delta}}_2) \right) \mathbf{1}(\mathcal{C}) (\boldsymbol{\delta}_1,\boldsymbol{\delta}_2,\tilde{\boldsymbol{\delta}}_1,\tilde{\boldsymbol{\delta}}_2) \right]. \end{align}\] Thus the likelihood-ratio terms are naturally written on the parameter space, whereas the probability of \(\mathcal{C}\) and \(\mathcal{C}^c\) is controlled on the latent space.

Define \[T_{\widetilde{\mathcal{C}}} := \mathbb{E}_{\pi_8^{\otimes2}} \left[ \mathbf{1}(\widetilde{\mathcal{C}}) \mathbb{E}_{\theta_*} \left( L_\theta^{\le D}L_{\tilde{\theta}}^{\le D} \right) \right],\] and \[T_{\widetilde{\mathcal{C}}^c} := \mathbb{E}_{\pi_8^{\otimes2}} \left[ \mathbf{1}(\widetilde{\mathcal{C}}^c) \mathbb{E}_{\theta_*} \left( L_\theta^{\le D}L_{\tilde{\theta}}^{\le D} \right) \right].\] Then \[\mathrm{LD}(D) = T_{\widetilde{\mathcal{C}}} + T_{\widetilde{\mathcal{C}}^c}.\]

For the term \(T_{\widetilde{\mathcal{C}}^c}\): By Cauchy’s inequality and 16, \[\begin{align} T_{\widetilde{\mathcal{C}}^c} &\le \mathbb{E}_{\pi_8^{\otimes2}} \left[ \mathbf{1}(\widetilde{\mathcal{C}}^c) \left\{ \mathbb{E}_{\theta_*}\left(L_\theta^{\le D}\right)^2 \mathbb{E}_{\theta_*}\left(L_{\tilde{\theta}}^{\le D}\right)^2 \right\}^{1/2} \right] \\ &\le 9(6npD)^{4D}\, \mathbb{P}_{\pi_7^{\otimes2}}(\mathcal{C}^c) \\ &\le 9(6npD)^{4D}(1+o(1))\, \mathbb{P}_{\pi_5^{\otimes2}}(\mathcal{C}^c), \end{align}\] where the last inequality follows from 61 . Let \[J := \sum_{j\in S_5} b_j^{(2)}\tilde{b}_j^{(2)} .\] Since \(q_j^{(2)}=0\) for \(j\in S_3\setminus S_5\), this is also the sum over \(S_3\). By the definition of \(\boldsymbol{\delta}_2\), \[\boldsymbol{\delta}_2^\top\tilde{\boldsymbol{\delta}}_2 = \frac{p_5}{k_u^2}J .\] The variables \(b_j^{(2)}\tilde{b}_j^{(2)}\) are independent Bernoulli random variables with success probabilities \((q_j^{(2)})^2\), and hence \[\mu_1 := \mathbb{E}_{\pi_5^{\otimes 2}}J = \sum_{j\in S_5}(q_j^{(2)})^2 = \frac{k_u^2}{64p_5} \ge \frac{k_u^2}{64k_{\rm eff}}.\] Therefore, \[\mathcal{C}^c = \left\{ \boldsymbol{\delta}_2^\top\tilde{\boldsymbol{\delta}}_2>C_5 \right\} = \left\{ J>C_5\frac{k_u^2}{p_5} \right\} = \left\{ J>64C_5\mu_1 \right\}.\] Thus \(\mathcal{C}^c\) is an upper-tail event. For \(a>1\), let \[\psi(a)=a\log a-a+1 .\] The Chernoff bound for sums of independent Bernoulli random variables gives \[\mathbb{P}_{\pi_5^{\otimes 2}}(J\ge a\mu_1) \le \exp\{-\mu_1\psi(a)\}.\] Taking \(a=64C_5\) and using the comparison between \(\pi_7\) and \(\pi_5\) in 61 , we have \[\begin{align} \mathbb{P}_{\pi_7^{\otimes 2}}(\mathcal{C}^c) &\le (1-\varepsilon_n)^{-2} \mathbb{P}_{\pi_5^{\otimes 2}}(\mathcal{C}^c) \\ &\le (1+o(1)) \exp\{-\mu_1\psi(64C_5)\}. \end{align}\] Consequently, \[T_{\widetilde{\mathcal{C}}^c} \le 9(1+o(1)) \exp\left\{ 4D\log(6npD) - \mu_1\psi(64C_5) \right\}.\] The growth condition above gives \[\mu_1\ge c_\mu D\log(6npD)\] for some absolute constant \(c_\mu>0\). Choose \(C_5\) large enough so that \(c_\mu\psi(64C_5)>8\). Then \[T_{\widetilde{\mathcal{C}}^c} \le 9(1+o(1)) \exp\{-4D\log(6npD)\} = o(1).\] We now fix such a value of \(C_5\). The constant \(c_8\) in the prior construction will be chosen sufficiently small below, after \(C_5\) is fixed.

For the term \(T_{\widetilde{\mathcal{C}}}\): On this event, the overlap between the two \(\boldsymbol{\delta}_2\)-perturbations is bounded by \(C_5\), so the projected likelihood-ratio inner product can be compared with the full likelihood-ratio inner product. By the first inequality in 16, for every parameter pair in the support of \(\pi_8^{\otimes2}\), \[\mathbb{E}_{\theta_*} \left( L_\theta^{\le D}L_{\tilde{\theta}}^{\le D} \right) \le \mathbb{E}_{\theta_*} \left( L_\theta L_{\tilde{\theta}} \right).\] Hence \[T_{\widetilde{\mathcal{C}}}\le \mathbb{E}_{\pi_8^{\otimes2}}\left[ \mathbf{1}(\widetilde{\mathcal{C}}) \mathbb{E}_{\theta_*} \left( L_\theta L_{\tilde{\theta}} \right) \right].\] Using the pushforward representation, the right-hand side equals \[\mathbb{E}_{\pi_7^{\otimes2}} \left[ \mathbf{1}({\mathcal{C}}) \mathbb{E}_{\theta_*} \left( L_{\theta(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)} L_{\theta(\tilde{\boldsymbol{\delta}}_1,\tilde{\boldsymbol{\delta}}_2)} \right) \right].\] As in 11, for \(\theta=\theta(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)\) and \(\tilde{\theta}=\theta(\tilde{\boldsymbol{\delta}}_1,\tilde{\boldsymbol{\delta}}_2)\), \[\mathbb{E}_{\theta_*} \left( L_\theta L_{\tilde{\theta}} \right) = \left[ 1- \left( \frac{\kappa\tilde{\kappa}}{\sigma_*^2} + \boldsymbol{\delta}_2^\top\tilde{\boldsymbol{\delta}}_2 \right) \boldsymbol{\delta}_1^\top\tilde{\boldsymbol{\delta}}_1 \right]^{-n}.\] On \(\mathcal{C}\) and on \(\mathcal{A}_n\), we have \[0<\kappa,\tilde{\kappa}\le 1, \qquad \boldsymbol{\delta}_2^\top\tilde{\boldsymbol{\delta}}_2\le C_5 .\] Therefore, with \[A_5:=C_5+\sigma_*^{-2},\] we obtain \[T_{\mathcal{C}} \le \mathbb{E}_{\pi_7^{\otimes 2}} \left[ \left( 1-A_5\boldsymbol{\delta}_1^\top\tilde{\boldsymbol{\delta}}_1 \right)^{-n} \right],\] where the indicator is dropped because the integrand is nonnegative. Using 61 , \[T_{\mathcal{C}} \le (1-\varepsilon_n)^{-2} \mathbb{E}_{\pi_5^{\otimes 2}} \left[ \left( 1-A_5\boldsymbol{\delta}_1^\top\tilde{\boldsymbol{\delta}}_1 \right)^{-n} \right].\] This is where \(\varepsilon_n=o(1)\) is needed.

Let \[s_1=\lfloor k_u/4\rfloor, \qquad H=|I\cap \tilde{I}|.\] Under \(\pi_5^{\otimes 2}\), \[H\sim {\rm Hypergeometric}(p_4,s_1,s_1),\] and by the construction of \(\boldsymbol{\delta}_1\), \[\boldsymbol{\delta}_1^\top\tilde{\boldsymbol{\delta}}_1 = c_8^2\frac{\log p}{n}H.\] Choose \(c_8>0\) sufficiently small, after \(C_5\) has been fixed, so that \[A_5c_8^2\frac{\log p}{n}s_1\le \frac{1}{2}\] and \[2A_5c_8^2<1-2\gamma',\] where \(\gamma'<1/2\) satisfies \(s_1\le p_4^{\gamma'}\) for all sufficiently large \(n\). Such a \(\gamma'\) exists because \(k_u\le p^\gamma\) with \(\gamma<1/2\) and \(p_4=p-k_{\rm eff}\sim p\). Then \[0\le A_5\boldsymbol{\delta}_1^\top\tilde{\boldsymbol{\delta}}_1\le \frac{1}{2},\] and \((1-x)^{-1}\le \exp(2x)\) for \(0\le x\le 1/2\) gives \[\begin{align} \left( 1-A_5\boldsymbol{\delta}_1^\top\tilde{\boldsymbol{\delta}}_1 \right)^{-n} &\le \exp\left( 2nA_5c_8^2\frac{\log p}{n}H \right) \\ &= \exp\left( 2A_5c_8^2\log p\cdot H \right). \end{align}\] By 12, since \(2A_5c_8^2<1-2\gamma'\), \[\mathbb{E}_{\pi_5^{\otimes 2}} \exp\left( 2A_5c_8^2\log p\cdot H \right) = 1+o(1).\] Therefore, \[T_{\widetilde{\mathcal{C}}} \le (1-\varepsilon_n)^{-2}(1+o(1)) = 1+o(1).\] Combining this bound with \(T_{\mathcal{C}^c}=o(1)\) gives \[\mathrm{LD}(D) = T_{\mathcal{C}}+T_{\mathcal{C}^c} \le 1+o(1).\] Since the degree-\(D\) polynomial space contains the constant function, \[\mathrm{LD}(D)\ge \|1\|_{L^2(\mathbb{Q}_0)}^2=1.\] Consequently, \[\mathrm{LD}(D)=1+o(1).\] By 3, no degree-\(D\) polynomial weakly separates \(\mathbb{Q}_1=\mathbb{P}^n_{\pi_8}\) from \(\mathbb{Q}_0=\mathbb{P}^n_{\theta_*}\). Since \(\pi_8\) is supported on the null space and \(\theta_*\) lies in the alternative space at separation \[c_9\nu_3\frac{k_u\log p}{n},\] the desired low-degree lower bound follows. ◻

8.4 Proof of Theorem 3↩︎

Proof. The proof is based on constructing a polynomial-time reduction from a sparse canonical correlation analysis (SCCA) detection problem to the linear hypothesis testing problem considered in this paper. For a generic detection problem \(\mathcal{P}\), we use \(\mathcal{L}_{H_0}(X)\) and \(\mathcal{L}_{H_1}(X)\) to denote the distributions of the instance \(X\) under the null \(H_0\) and the alternative \(H_1\), respectively. The reduction principle we rely on is summarized in the following lemma, which is adapted from [55].

Lemma 17. Let \(\mathcal{P}\) and \(\mathcal{P}'\) be detection problems with hypotheses \((H_0, H_1)\) and \((H_0', H_1')\), and let \(X\) and \(Y\) be instances of \(\mathcal{P}\) and \(\mathcal{P}'\), respectively. Suppose there exists a polynomial-time computable map \(\phi\) and a prior \(\pi\) on \(H_1^\prime\) such that \[d_{\mathrm{TV}}\!\left(\mathcal{L}_{H_0}(\phi(X)), \mathcal{L}_{H_0'}(Y)\right) \;+\;\inf d_{\mathrm{TV}}\!\left( \mathcal{L}_{H_1}(\phi(X)), \int_{H_1^\prime}\mathcal{L}_{\mathbb{P}^\prime}(Y)\,{\rm d}\,\pi(\mathbb{P}^\prime) \right) \;\le\; \delta .\] If there exists a polynomial-time algorithm solving \(\mathcal{P}'\) with Type I+II error at most \(\epsilon\), then there exists a polynomial-time algorithm solving \(\mathcal{P}\) with Type I+II error at most \(\epsilon+\delta\).

The proof of 17 follows directly from the definition of total variation distance and is therefore omitted.

For notational convenience, we set \(p_6 = k_\xi,\; p_7 = p-k_\xi,\) and \(k_* = \lfloor k_u/4 \rfloor.\) Fix \(\sigma_*=M_2/2\), so \((\mathbf{0}_{p\times1},\mathbf{I}_{p\times p},\sigma_*)\in \Theta(\lfloor k_u/2\rfloor)\) by 4 . Let \(\tau_{\rm red}\) be the exact value of \(\xi^\top\beta\) produced by the \(h\)-transformation under the alternative covariance in 63 ; it is computed in 64 . Since \(\xi_1=1\), set \[\beta_0=(t_0-\tau_{\rm red})e_1, \qquad \varphi(\beta,\Sigma,\sigma)=(\beta-\beta_0,\Sigma,\sigma).\] This is the transition argument from 8.1, with \(j_0=1\) and \(\tau\) replaced by \(\tau_{\rm red}\). The associated data transformation is \((Y,X)\mapsto(Y-X\beta_0,X)\), and its inverse sends a translated sample \(Z'=(Y',X)\) to \((Y'+X\beta_0,X)\). For the present reduction we need two consequences. First, the inverse image of \((\mathbf{0},\mathbf{I},\sigma_*)\) is \((\beta_0,\mathbf{I},\sigma_*)\), whose linear functional is \(t_0-\tau_{\rm red}\). Once \(c\) is chosen so that \(\tau\le\tau_{\rm red}\), this inverse image belongs to the LT alternative class \(\Theta_{\pm\tau}(k;\xi,t_0)\). Second, if \(\tilde{\theta}\in\Theta(\lfloor k_u/2\rfloor;\xi,\tau_{\rm red})\), then \[\xi^\top(\tilde{\beta}+\beta_0) = \tau_{\rm red}+t_0-\tau_{\rm red} =t_0,\] and the sparsity increases by at most one, so \(\varphi^{-1}(\tilde{\theta})\in\Theta(k_u;\xi,t_0)\). Thus the SCCA null is sent to an LT alternative point, whereas the SCCA alternative is sent to the LT null class. If \(\psi\) solves \({\rm LT}(n,k_u,k,k_\xi,p,\tau)\), the SCCA decision uses \[1-\psi(Y'+X\beta_0,X).\] It remains to construct an exact reduction from SCCA to \[\label{eq:32reduction32target} H_0^{\rm tr}:\theta=(\mathbf{0},\mathbf{I},\sigma_*), \qquad H_1^{\rm tr}:\theta\in \Theta(\lfloor k_u/2\rfloor;\xi,\tau_{\rm red}).\tag{62}\] The prior on \(H_1^{\rm tr}\) is induced by the random supports of \(\widetilde{\boldsymbol{\delta}}_1\) and \(\widetilde{\boldsymbol{\delta}}_2\) in 63 ; the construction below gives exact law matching under both hypotheses, hence the total-variation slack is \(\delta=0\).

We now specialize to the regime \(\tau \asymp \rho_n^2 k_u / \sqrt{k_\xi}\) for some \(\rho_n \le 1\). Consider the Gaussian testing problem with \(V_1\in\mathbb{R}, V_2\in\mathbb{R}^{p_6}\), and \(V_3\in\mathbb{R}^{p_7}\): \[\label{eq:32reduction32alternative} \begin{align} & H_0:\; \begin{pmatrix} V_1\\ V_2\\ V_3 \end{pmatrix} &\sim \mathcal{N}\!\left( {{\mathbf{0}}}, \left(\begin{array}{c|c|c} \sigma_*^2 & {{\boldsymbol{0}}}_{1\times p_6} & {\boldsymbol{0}}_{1\times p_7} \\ \hline {{\boldsymbol{0}}}_{p_6\times 1} & {{\mathbf{I}}}_{p_6\times p_6} & {\boldsymbol{0}}_{p_6\times p_7} \\ \hline {\boldsymbol{0}}_{p_7\times 1} & {\boldsymbol{0}}_{p_7\times p_6} & {{\mathbf{I}}}_{p_7\times p_7} \end{array}\right)\right)\\ & \text{vs.} \\[0.5em] & H_1:\; \begin{pmatrix} V_1\\ V_2\\ V_3 \end{pmatrix} &\sim \renewcommand{\arraystretch}{1.4} \mathcal{N}\!\left( {{\mathbf{0}}}, \left(\begin{array}{c|c|c} \sigma_*^2 & \mathbf{0}_{1\times p_6} & \widetilde{\boldsymbol{\delta}}_2^\top \\ \hline \mathbf{0}_{p_6\times 1} & \mathbf{I}_{p_6} & \widetilde{\boldsymbol{\delta}}_1\widetilde{\boldsymbol{\delta}}_2^\top \\ \hline \widetilde{\boldsymbol{\delta}}_2 & \widetilde{\boldsymbol{\delta}}_2\widetilde{\boldsymbol{\delta}}_1^\top & \mathbf{I}_{p_7} \end{array}\right) \right), \end{align}\tag{63}\] where \(\widetilde{\boldsymbol{\delta}}_1\in\mathbb{R}^{p_6}\) and \(\widetilde{\boldsymbol{\delta}}_2\in\mathbb{R}^{p_7}\) are \(k_*\)-sparse vectors with supports drawn uniformly at random from all subsets of \([p_6]\) and \([p_7]\) of size \(k_*\), respectively. Their nonzero entries are defined by \[(\widetilde{\boldsymbol{\delta}}_1)_j = -\frac{c_{10}}{\sigma_*}\frac{\sqrt{p_6}}{k_*}, \quad j\in{\rm supp}(\widetilde{\boldsymbol{\delta}}_1), \qquad (\widetilde{\boldsymbol{\delta}}_2)_j = \frac{\sigma_*\rho_n}{\sqrt{2p_6}}, \quad j\in{\rm supp}(\widetilde{\boldsymbol{\delta}}_2),\] for a sufficiently small \(c_{10}>0\).

Under the mapping \(h\), the induced parameter \(\theta=(\beta,\Sigma,\sigma)\) corresponding to the covariance matrix in the alternative \(H_1\) of 63 admits the explicit expressions \[\begin{align} \beta &= \begin{pmatrix} \mathbf{I}_{p_6} & \widetilde{\boldsymbol{\delta}}_1\widetilde{\boldsymbol{\delta}}_2^\top \\ \widetilde{\boldsymbol{\delta}}_2\widetilde{\boldsymbol{\delta}}_1^\top & \mathbf{I}_{p_7} \end{pmatrix}^{-1} \begin{pmatrix} \mathbf{0}_{p_6} \\ \widetilde{\boldsymbol{\delta}}_2 \end{pmatrix} = \begin{pmatrix} -\dfrac{\|\widetilde{\boldsymbol{\delta}}_2\|_2^2} {1-\|\widetilde{\boldsymbol{\delta}}_1\|_2^2\|\widetilde{\boldsymbol{\delta}}_2\|_2^2} \,\widetilde{\boldsymbol{\delta}}_1 \\[0.6em] \dfrac{1} {1-\|\widetilde{\boldsymbol{\delta}}_1\|_2^2\|\widetilde{\boldsymbol{\delta}}_2\|_2^2} \,\widetilde{\boldsymbol{\delta}}_2 \end{pmatrix},\\ \Sigma &= \begin{pmatrix} \mathbf{I}_{p_6} & \widetilde{\boldsymbol{\delta}}_1\widetilde{\boldsymbol{\delta}}_2^\top \\ \widetilde{\boldsymbol{\delta}}_2\widetilde{\boldsymbol{\delta}}_1^\top & \mathbf{I}_{p_7} \end{pmatrix}, \qquad \sigma = \sqrt{ \sigma_*^2 - \frac{\|\widetilde{\boldsymbol{\delta}}_2\|_2^2}{1-\|\widetilde{\boldsymbol{\delta}}_1\|_2^2\|\widetilde{\boldsymbol{\delta}}_2\|_2^2} }. \end{align}\]

It is immediate that \(\beta\) is \(2k_*\)-sparse, and hence at most \(\lfloor k_u/2\rfloor\)-sparse. By choosing \(c_{10}\) sufficiently small, \(\Sigma\) and \(\sigma\) are well defined and satisfy the constraints in the parameter space 4 . Moreover, since \(\|\widetilde{\boldsymbol{\delta}}_1\|_2^2 = c_{10}^2p_6/(k_*\sigma_*^2)\) and \(\|\widetilde{\boldsymbol{\delta}}_2\|_2^2 = \sigma_*^2\rho_n^2 k_*/(2p_6),\) we obtain

\[\label{eq:32reduction32tau32red} \tau_{\rm red} := \xi^\top\beta = \frac{ \frac{c_{10}}{\sigma_*} \sqrt{p_6}\,\|\widetilde{\boldsymbol{\delta}}_2\|_2^2}{1-\|\widetilde{\boldsymbol{\delta}}_1\|_2^2\|\widetilde{\boldsymbol{\delta}}_2\|_2^2} = \frac{c_{10} \sigma_*}{2-c_{10}^2\rho_n^2}\, \rho_n^2\frac{k_*}{\sqrt{p_6}}.\tag{64}\] Since \(p_6=k_\xi\), \(k_*=\lfloor k_u/4\rfloor\), and \(\rho_n<1/2\), the denominator in 64 is at most \(2\), and \(k_*\ge k_u/8\) for all sufficiently large \(n\). Hence \[\tau_{\rm red} \ge \frac{c_{10}\sigma_*}{16}\, \rho_n^2\frac{k_u}{\sqrt{k_\xi}}.\] After fixing \(c_{10}\in(0,1)\) sufficiently small for the covariance and noise constraints above, choose the theorem constant \(c\le c_{10}\sigma_*/16\). Then the theorem’s separation level \(\tau=c\rho_n^2 k_u/\sqrt{k_\xi}\) satisfies \(\tau\le \tau_{\rm red}\), as required in the translated testing problem 62 .

It remains to construct a polynomial-time mapping from \({\rm SCCA}(2n,k_*,p_6,p_7,\rho_n)\) to the testing problem in 63 . Since the observations are i.i.d.under both models, it suffices to describe the transformation at the level of a single pair of observations. Specifically, we describe how two independent samples \((U_1,U_2)\) and \((U_1',U_2')\) from the SCCA model are mapped to a single observation \((V_1,V_2,V_3)\).

The mapping is defined as follows.

  • Compute \(W_1 = \sum_{j=1}^{p_6}(U_1')_j/\sqrt{p_6}\in\mathbb{R}\).

  • Set \[V_1 = \sigma_* W_1,\qquad V_2 = -c_{10} U_1 + \sqrt{1-c_{10}^2}\,V_*,\qquad V_3 = (U_2+U_2')/\sqrt{2},\] where \(V_* \sim \mathcal{N}(0,\mathbf{I}_{p_6})\) is independent of all other random variables.

Then under the null, since \(U_1,U_1^\prime\sim \mathcal{N}(0,\mathbf{I}_{p_6})\) and \(U_2,U_2^\prime\sim \mathcal{N}(0,\mathbf{I}_{p_7})\) are independent, we have \[\begin{pmatrix} V_1\\ V_2\\ V_3 \end{pmatrix}\sim \mathcal{N}\!\left( 0, \left(\begin{array}{c|c|c} \sigma_*^2 & {{\boldsymbol{0}}}_{1\times p_6} & {\boldsymbol{0}}_{1\times p_7} \\ \hline {{\boldsymbol{0}}}_{p_6\times 1} & {{\mathbf{I}}}_{p_6\times p_6} & {\boldsymbol{0}}_{p_6\times p_7} \\ \hline {\boldsymbol{0}}_{p_7\times 1} & {\boldsymbol{0}}_{p_7\times p_6} & {{\mathbf{I}}}_{p_7\times p_7} \end{array}\right)\right).\] Under the alternative, we first have: \[\begin{pmatrix} U_1\\U_2 \end{pmatrix}\sim \mathcal{N}\!\left( 0, \begin{pmatrix} \mathbf{I}_{p_6} & \rho_n \boldsymbol{\delta}_1 \boldsymbol{\delta}_2^\top \\[0.2em] \rho_n \boldsymbol{\delta}_2 \boldsymbol{\delta}_1^\top & \mathbf{I}_{p_7} \end{pmatrix} \right),\qquad \begin{pmatrix} W_1\\U_2^\prime \end{pmatrix} \sim \mathcal{N}\!\left( 0, \begin{pmatrix} 1 & \rho_n\sqrt{\frac{ k_*}{p_6}} \boldsymbol{\delta}_2^\top \\[0.2em] \rho_n\sqrt{\frac{ k_*}{p_6}} \boldsymbol{\delta}_2 & \mathbf{I}_{p_7} \end{pmatrix} \right)\] are independent, where we have used the fact that \(\boldsymbol{\delta}_1\) has exactly \(k_*\) nonzero entries each equal to \(1/\sqrt{k_*}\) under the alternative of SCCA. Therefore, we have \[\renewcommand{\arraystretch}{1.2} \begin{pmatrix} V_1\\ V_2\\ V_3 \end{pmatrix}\sim \mathcal{N}\!\left( {{\mathbf{0}}}, \left(\begin{array}{c|c|c} \sigma_*^2 & \mathbf{0}_{1\times p_6} & \rho_n\sigma_*\sqrt{\frac{k_*}{2p_6}}{\boldsymbol{\delta}}_2^\top \\ \hline \mathbf{0}_{p_6\times 1} & \mathbf{I}_{p_6} & -\frac{c_{10}}{\sqrt{2}}\rho_n{\boldsymbol{\delta}}_1{\boldsymbol{\delta}}_2^\top \\ \hline \rho_n\sigma_*\sqrt{\frac{k_*}{2p_6}}{\boldsymbol{\delta}}_2 & -\frac{c_{10}}{\sqrt{2}}\rho_n{\boldsymbol{\delta}}_2{\boldsymbol{\delta}}_1^\top & \mathbf{I}_{p_7} \end{array}\right) \right).\] Consequently, \((V_1,V_2,V_3)\) follows the alternative distribution in 63 , completing the reduction. ◻

9 Proof of Additional Technical Results for the Upper Bounds↩︎

9.1 Proof of Lemma 7↩︎

Proof of 7. For \(u=\Omega\xi\), define \[d=\hat{\Sigma}u-\xi=\frac{1}{n}\sum_{i=1}^n X_{i\cdot}X_{i\cdot}^{\top}\Omega\xi-\xi\in\mathbb{R}^p,\] with coordinates \[d_j=\frac{1}{n}\sum_{i=1}^n X_{ij}X_{i\cdot}^{\top}\Omega\xi-\xi_j,\qquad j\in[p].\] For any \(j\in[p]\) and \(i\in [n]\), define \(f_{ij}=X_{ij}X_{i\cdot}^{\top}\Omega\xi\). Then \(\mathbb{E}f_{ij}=\xi_j\), and \(f_{ij}\) is sub-exponential. Indeed, by 4, \[\begin{align} \|f_{ij}\|_{\psi_1} =\|X_{ij}\,X_{i\cdot}^{\top}\Omega\xi\|_{\psi_1} &\le 2\,\|X_{ij}\|_{\psi_2}\,\|X_{i\cdot}^{\top}\Omega\xi\|_{\psi_2}. \end{align}\] Since \(X_{ij}\sim \mathcal{N}(0, \Sigma_{jj})\) and \(X_{i\cdot}^{\top}\Omega\xi\sim \mathcal{N}(0, \xi^\top\Omega\xi)\), we have \(\|X_{ij}\|_{\psi_2}\le \sqrt{M_1}\) and \(\|X_{i\cdot}^{\top}\Omega\xi\|_{\psi_2}\le \sqrt{M_1}\|\xi\|\) where \(M_1\) is the constant in the assumption on the eigenvalues of \(\Sigma\) (see 4 ).

Hence, there exists a constant \(C'>0\) such that \(\|f_{ij}\|_{\psi_1}\le C'\|\xi\|_2\). Furthermore, Lemma 2 implies that \[\| f_{ij} - \mathbb{E} f_{ij} \|_{\psi_1} \le 2 \| f_{ij} \|_{\psi_1} \le 2C' \|\xi\|_2 .\] We pick a constant \(c\) sufficiently large such that \(2^{2-c^2c_0}<\alpha/24\), where \(c_0\) is the constant in Lemma 3. For any \(j\in [p]\), we apply Lemma 3 with \(d_j= n^{-1}\sum_{i=1}^n f_{ij}\) and \(t = 2c C' \|\xi\|_2 \sqrt{\tfrac{\log p}{n}}\) to conclude that for sufficiently large \(n\), it holds that \[\mathbb{P}\!\left( |d_j| \ge 2c C' \|\xi\|_2 \sqrt{\frac{\log p}{n}} \right) \le 2 \exp\!\left( - c_0 \min\!\left( c^2 \log p,\; c \sqrt{n \log p} \right) \right) \le 2 p^{-c^2 c_0},\] where the last inequality holds as long as \(n>c^2 \log p\), which is guaranteed because \(n / \log p \gtrsim k_u \to \infty\). Taking a union bound over all \(j\), we obtain \[\mathbb{P}\!\left( \|d\|_\infty \ge 2c C' \|\xi\|_2 \sqrt{\frac{\log p}{n}} \right) \le 2 p^{\,1 - c^2 c_0}\le 2^{2-c^2c_0}\le \alpha/24.\] Choosing \(C_\xi=2c C'\) completes the proof. ◻

9.2 Proof of Lemma 8↩︎

Proof. Before proceeding, we record two concentration tools used repeatedly in the proof.

Lemma 18 ([56]). Let \(Y\) be an \(n\times k\) matrix with i.i.d.\(\mathcal{N}(0,1)\) entries. For any \(t>0\), \[\mathbb{P}\!\left( \left\| \frac{1}{n}Y^\top Y - {\mathbf{I}}_k \right\|_2 \le 2\left(\sqrt{\frac{k}{n}}+t\right) + \left(\sqrt{\frac{k}{n}}+t\right)^2 \right) \;\ge\; 1-2e^{-nt^2/2}.\]

Lemma 19 ([57]). Let A be an \(N \times n\) matrix whose entries are independent standard normal random variables. Then for every \(t \ge 0\), with probability at least \(1 - 2 \exp(-t^2/2)\), one has \[\|A\|_2\le \sqrt{N}+\sqrt{n}+t\]

We now prove 8. Let \[\Sigma = V\Lambda V^\top + {\mathbf{I}}_p, \qquad \Lambda=\operatorname{diag}(\lambda_1,\ldots,\lambda_r)\in\mathbb{R}^{r\times r},\] and let \(B_*={\rm supp}(V)\subseteq[p]\) denote the row support of \(V\). Note that \(k:=|B_*|\le k_u\) for a known upper bound \(k_u\). Define the sample covariance matrix based on the first half of the data by \[\label{eq:32sample32cov321} \hat{\Sigma}^{(1)}=\frac{1}{n_1}\bigl(X^{(1)}\bigr)^\top X^{(1)}.\tag{65}\]

For any symmetric matrix \({\mathbf{A}}=(a_{ij})\in\mathbb{R}^{p\times p}\) and any index set \(B\subseteq[p]\), define \(\Gamma_B({\mathbf{A}})\) to be the \(p\times p\) matrix whose \(B\times B\) principal submatrix equals \({\mathbf{A}}_{BB}\), whose diagonal entries indexed by \(B^c\) are equal to one, and whose remaining off-diagonal entries are zero: \[\label{eq:AB} (\Gamma_B({\mathbf{A}}))_{ij} = a_{ij}\,\mathbf{1}\{i\in B,\, j\in B\} + \mathbf{1}\{i=j\in B^c\}.\tag{66}\] Equivalently, after permuting indices so that \(B\) appears first, \[\Gamma_B({\mathbf{A}}) = \begin{bmatrix} {\mathbf{A}}_{BB} & {\mathbf{0}}\\ {\mathbf{0}}& {\mathbf{I}}_{B^cB^c} \end{bmatrix}.\] For two index sets \(I,J\subseteq[p]\), we write \({\mathbf{A}}_{IJ}\) for the \(|I|\times|J|\) submatrix of \({\mathbf{A}}\) with rows indexed by \(I\) and columns indexed by \(J\).

Following [47], define the candidate support class \[\label{eq:supp-set} \begin{align} \mathbb{B}_{k_u} = \Biggl\{ &B\subseteq [p]:~ |B|\le k_u, ~\text{and for all }D\subseteq B^c\text{ with }|D|\le k_u,\\ & \|\Gamma_D(\hat{\Sigma}^{(1)})-{\mathbf{I}}_p\|_2 \le 2\left(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{\gamma_*|D|\log p}{n_1}}\right) +\left(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{\gamma_*|D|\log p}{n_1}}\right)^2,\\ &\|\hat{\Sigma}^{(1)}_{DB}\|_2 \le \sqrt{\|\Gamma_B(\hat{\Sigma}^{(1)})\|_2}\biggl(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{|B|}{n_1}}+\sqrt{\frac{\gamma_* |D|\log p}{n_1}}\biggr) \Biggr\}. \end{align}\tag{67}\] (Here \(\Gamma_D(\hat{\Sigma}^{(1)})\) and \(\Gamma_B(\hat{\Sigma}^{(1)})\) are interpreted in the sense of 66 .)

Define \({\cal E}_1=\{B_*\in\mathbb{B}_{k_u}\}\).

Lemma 20. Suppose \(\Sigma\in \Pi_0(k,p)\) and \(\hat{\Sigma}^{(1)}\) is constructed from \(n_1\) i.i.d.samples drawn from \(\mathcal{N}(0,\Sigma)\) as in 65 . If \(n_1\ge c k_u \log p\) for some sufficiently large constant \(c>0\), then for \(\gamma_*\ge 3\) and \(p\ge 2\), we have \[\mathbb{P}\bigl({\cal E}_1\bigr)\ge 1-8 p^{ 1-\gamma_*/2 }.\] In particular, on \({\cal E}_1\), the class \(\mathbb{B}_{k_u}\) is nonempty.

By 20, \(\mathbb{P}({\cal E}_1)\ge 1-8p^{1-\gamma_*/2}\). On the event \({\cal E}_1\), the class \(\mathbb{B}_{k_u}\) is nonempty and we choose any \(\widehat B\in\mathbb{B}_{k_u}\). We then construct \[\label{eq:32spike32estimators} \hat{\Sigma}_{\text{spike}} := \Gamma_{\widehat{B}}(\hat{\Sigma}^{(1)})\mathbf{1}\left\{{\cal E}_1\right\}+{{\mathbf{I}}}_p\mathbf{1}\left\{{\cal E}_1^c\right\}, \qquad \widehat{\Omega} := \hat{\Sigma}_{\text{spike}}^{-1}.\tag{68}\]

By triangle inequality, \[\begin{align} \label{eq:32triangle32inequality32matrix} \bigl\|\hat{\Sigma}_{{\text{spike}}}-\Sigma\bigr\|_2 \mathbf{1}\left\{{\cal E}_1\right\} &\le \underbrace{\,\bigl\|\Gamma_{\widehat{B}}(\hat{\Sigma}^{(1)})-\Gamma_{B_*}(\hat{\Sigma}^{(1)})\bigr\|_2\mathbf{1}\left\{{\cal E}_1\right\}}_{\mathrm{I}} + \underbrace{\,\bigl\|\Gamma_{B_*}(\hat{\Sigma}^{(1)})-\Sigma\bigr\|_2\mathbf{1}\left\{{\cal E}_1\right\}}_{\mathrm{II}} \end{align}\tag{69}\]

It remains to control \(\|\widehat{\Omega}-\Omega\|_2\) on \({\cal E}_1\).

Bounding \(\mathrm{I}\) in 69 . Let \[G=B_*\cap \widehat{B},\qquad M=B_*\cap \widehat{B}^{\,c},\qquad O=B_*^{\,c}\cap \widehat{B}. \label{eq:decomp-GMO}\tag{70}\] These correspond to correctly identified, missing, and overly identified coordinates, respectively. Writing the matrix in the block order \((G,M,O)\), the difference \(\Gamma_{\widehat{B}}(\hat{\Sigma}^{(1)})-\Gamma_{B_*}(\hat{\Sigma}^{(1)})\) can be decomposed as a sum of four block-sparse matrices, \[\Gamma_{\widehat{B}}(\hat{\Sigma}^{(1)})-\Gamma_{B_*}(\hat{\Sigma}^{(1)}) = A_M + A_O + A_{GM} + A_{GO},\] where \[A_M= \begin{pmatrix} 0&0&0\\ 0&{\mathbf{I}}_{MM}-\hat{\Sigma}^{(1)}_{MM}&0\\ 0&0&0 \end{pmatrix},\quad A_O= \begin{pmatrix} 0&0&0\\ 0&0&0\\ 0&0&\hat{\Sigma}^{(1)}_{OO}-{\mathbf{I}}_{OO} \end{pmatrix},\] and \[A_{GM}= \begin{pmatrix} 0&-\hat{\Sigma}^{(1)}_{GM}&0\\ -\hat{\Sigma}^{(1)}_{MG}&0&0\\ 0&0&0 \end{pmatrix},\quad A_{GO}= \begin{pmatrix} 0&0&\hat{\Sigma}^{(1)}_{GO}\\ 0&0&0\\ \hat{\Sigma}^{(1)}_{OG}&0&0 \end{pmatrix}.\] By the triangle inequality for the operator norm, \[\bigl\|\Gamma_{\widehat{B}}(\hat{\Sigma}^{(1)})-\Gamma_{B_*}(\hat{\Sigma}^{(1)})\bigr\|_2 \le \|A_M\|_2+\|A_O\|_2+\|A_{GM}\|_2+\|A_{GO}\|_2.\] Moreover, since \(A_M\) and \(A_O\) are block-diagonal, we have \[\|A_M\|_2=\|{\mathbf{I}}_{MM}-\hat{\Sigma}^{(1)}_{MM}\|_2, \qquad \|A_O\|_2=\|\hat{\Sigma}^{(1)}_{OO}-{\mathbf{I}}_{OO}\|_2.\] For the off-diagonal blocks, applying the triangle inequality again yields \[\left\| \begin{pmatrix} 0 & X\\ X^\top & 0 \end{pmatrix} \right\|_2 \le \left\| \begin{pmatrix} 0 & X\\ 0 & 0 \end{pmatrix} \right\|_2 + \left\| \begin{pmatrix} 0 & 0\\ X^\top & 0 \end{pmatrix} \right\|_2 = \|X\|_2+\|X^\top\|_2 = 2\|X\|_2,\] and thus \(\|A_{GM}\|_2\le 2\|\hat{\Sigma}^{(1)}_{GM}\|_2\) and \(\|A_{GO}\|_2\le 2\|\hat{\Sigma}^{(1)}_{GO}\|_2\).

Combining these bounds gives \[\begin{align} \label{eq:32decomp32variance32error} \|\Gamma_{\widehat{B}}(\hat{\Sigma}^{(1)})-\Gamma_{B_*}(\hat{\Sigma}^{(1)})\|_2 &\le \|{\mathbf{I}}_{MM}-\hat{\Sigma}^{(1)}_{MM}\|_2 +\|\hat{\Sigma}^{(1)}_{OO}-{\mathbf{I}}_{OO}\|_2 +2\|\hat{\Sigma}^{(1)}_{GM}\|_2 +2\|\hat{\Sigma}^{(1)}_{GO}\|_2. \end{align}\tag{71}\]

We now define the event \[{\cal E}_2=\left\{ \left\|\Gamma_{\widehat{B}}(\hat{\Sigma}^{(1)}) -\Gamma_{B_*}(\hat{\Sigma}^{(1)})\right\|_2 \lesssim C_{E,1}\sqrt{\frac{k_u\log p}{n_1}} \right\}.\] The following analysis of the right-hand side of 71 establishes that, for a sufficiently large constant \(C_{E,1}\), the event \({\cal E}_2\) occurs with probability \(1-o(1)\).

  1. Since \(M\subseteq \widehat{B}^{\,c}\) and \(|M|\le k_u\), the defining property of \(\widehat{B}\in\mathbb{B}_{k_u}\) yields, for every such subset \(D\subseteq \widehat{B}^{\,c}\) with \(|D|\le k_u\), \[\|\hat{\Sigma}^{(1)}_{D}-{\mathbf{I}}_p\|_2\le 2\left(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{\gamma_*|D|\log p}{n_1}}\right) +\left(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{\gamma_*|D|\log p}{n_1}}\right)^2,\] where, since we have \(n_1\ge c k_u \log p\) for some sufficiently large constant \(c\), the leading rate in the right-hand side is \(\sqrt{(\gamma_*|D|\log p)/{n_1}}\).

    Applying this with \(D=M\) gives \[\begin{align} \label{eq:32leading32I32in32covariance} \|\hat{\Sigma}^{(1)}_{MM}-{\mathbf{I}}_{MM}\|_2 = \|\hat{\Sigma}^{(1)}_{M}-{\mathbf{I}}_p\|_2 \lesssim \sqrt{\frac{k_u\log p}{n_1}}. \end{align}\tag{72}\]

  2. For any fixed \(D\subseteq B_*^{\,c}\) with \(|D|\le k_u\), the submatrix \(X^{(1)}_{\cdot D}\) has i.i.d.\(\mathcal{N}(0,{\mathbf{I}}_{|D|})\) columns, hence \[\hat{\Sigma}^{(1)}_{DD}=\frac{1}{n_1}(X^{(1)}_{\cdot D})^\top X^{(1)}_{\cdot D} = \frac{1}{n_1}Z^\top Z,\] for \(Z\in\mathbb{R}^{n_1\times |D|}\) with i.i.d.\(\mathcal{N}(0,1)\) entries. Therefore, by similar calculation in 9.3, with probability at least \(1-4(ep)^{1-\gamma_*/2}\), we have \[\|\hat{\Sigma}^{(1)}_{DD}-{\mathbf{I}}_{DD}\|_2\le 2\left(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{\gamma_*|D|\log p}{n_1}}\right) +\left(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{\gamma_*|D|\log p}{n_1}}\right)^2,\] for all \(D\subseteq B_*^c\) and \(|D|\le k_u\). We take \(D=O\) and conclude that the leading term on the right-hand side of the above display is of order \(\sqrt{k_u\log p/n_1}\). Therefore, with probability at least \(1-4(ep)^{1-\gamma_*/2}\), it holds that \[\begin{align} \label{eq:32leading32II32in32covariance} \|\hat{\Sigma}^{(1)}_{OO}-{\mathbf{I}}_{OO}\|_2 \lesssim \sqrt{\frac{k_u\log p}{n_1}}. \end{align}\tag{73}\]

  3. Since \(M\subseteq \widehat{B}^{\,c}\) and \(G\subseteq \widehat{B}\), we have \(\hat{\Sigma}^{(1)}_{GM}\) as a submatrix of \(\hat{\Sigma}^{(1)}_{\widehat{B}M}\). By monotonicity of the operator norm under taking submatrices, \(\|\hat{\Sigma}^{(1)}_{GM}\|_2 \le \|\hat{\Sigma}^{(1)}_{\widehat{B}M}\|_2\). The defining property of \(\widehat{B}\in\mathbb{B}_{k_u}\) states that for every \(D\subseteq \widehat{B}^{\,c}\) with \(|D|\le k_u\), \[\|\hat{\Sigma}^{(1)}_{D\widehat{B}}\|_2 \le \sqrt{\|\Gamma_{\widehat{B}}(\hat{\Sigma}^{(1)})\|_2}\biggl(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{|\widehat{B}|}{n_1}}+\sqrt{\frac{\gamma_* |D|\log p}{n_1}}\biggr).\] Applying this with \(D=M\) gives \[\begin{align} \|\hat{\Sigma}^{(1)}_{GM}\|_2 \le \|\hat{\Sigma}^{(1)}_{\widehat{B}M}\|_2 \le \sqrt{\|\hat{\Sigma}^{(1)}_{\hat{B}\hat{B}}\|_2}\biggl(\sqrt{\frac{|M|}{n_1}}+\sqrt{\frac{|\widehat{B}|}{n_1}}+\sqrt{\frac{\gamma_* |M|\log p}{n_1}}\biggr). \end{align}\]

    Moreover, we have \[\max_{|B|\le k_u}\|\hat{\Sigma}_{BB}^{(1)}\|_2\le \max_{|B|\le k_u}\|\Sigma_{BB}\|\cdot \|\Sigma_{BB}^{-1/2}\hat{\Sigma}_{BB}^{(1)}\Sigma_{BB}^{-1/2}\|_2\le M_1\max_{|B|\le k_u }\|\Sigma_{BB}^{-1/2}\hat{\Sigma}_{BB}^{(1)}\Sigma_{BB}^{-1/2}\|_2,\] where we have used the eigenvalue condition \(\lambda_{\max}(\Sigma)\le M_1\). Therefore, Lemma 18 with \(k=k_u\), together with a union bound, gives (similarly to 9.3) that \[\max_{|B|\le k_u}\|\hat{\Sigma}_{BB}^{(1)}\|_2\le M_1\left(1+\sqrt{\frac{k_u}{n_1}}+\sqrt{\frac{\gamma_* k_u\log p}{n_1}}\right)=O(1)\] with probability \(1-4p^{1-\gamma_*/2}\). Therefore, we have \[\begin{align} \label{eq:32leading32III32in32covariance} \|\hat{\Sigma}^{(1)}_{GM}\|_2\lesssim \sqrt{\frac{k_u\log p}{n_1}}. \end{align}\tag{74}\]

  4. We next control \(\|\hat{\Sigma}^{(1)}_{GO}\|_2\). On \(\mathcal{E}_1\), since \[G\subseteq B_*, \qquad O\subseteq B_*^c, \qquad |O|\le |\widehat B|\le k_u,\] the defining property of \(B_*\in\mathbb{B}_{k_u}\) in 67 , applied with \(D=O\), gives \[\|\hat{\Sigma}^{(1)}_{O B_*}\|_2 \le \sqrt{\|\Gamma_{B_*}(\hat{\Sigma}^{(1)})\|_2} \biggl( \sqrt{\frac{|O|}{n_1}} +\sqrt{\frac{|B_*|}{n_1}} +\sqrt{\frac{\gamma_* |O|\log p}{n_1}} \biggr).\] Moreover, \[\|\hat{\Sigma}^{(1)}_{GO}\|_2 = \|\hat{\Sigma}^{(1)}_{OG}\|_2 \le \|\hat{\Sigma}^{(1)}_{O B_*}\|_2.\] Using the same argument with Lemma 18 in the bound on \(\|\hat{\Sigma}_{GM}^{(1)}\|_2\), it holds with probability \(1-4p^{1-\gamma_*/2}\) that \[\|\Gamma_{B_*}(\hat{\Sigma}^{(1)})\|_2=O(1).\] On the intersection of both events, we have \[\label{eq:32leading32IV32in32covariance} \|\hat{\Sigma}^{(1)}_{GO}\|_2 \lesssim \sqrt{\frac{k_u\log p}{n_1}}.\tag{75}\]

Combining 71 with 72 , 73 , 74 , and 75 yields that with probability \(1-20 p^{1-\gamma_*/2}\), we have \[\begin{align} \label{eq:32bound32I32in32covariance} \|\Gamma_{\widehat{B}}(\hat{\Sigma}^{(1)})-\Gamma_{B_*}(\hat{\Sigma}^{(1)})\|_2 \lesssim \sqrt{\frac{k_u\log p}{n_1}}. \end{align}\tag{76}\]

Bounding \(\mathrm{II}\) in 69 . Note that \[\|\Gamma_{B_*}(\hat{\Sigma}^{(1)})-\Sigma\|_2=\|\hat{\Sigma}^{(1)}_{B_*B_*}-\Sigma_{B_*B_*}\|_2\le \|\Sigma_{B_*B_*}\|_2\cdot \|\Sigma_{B_*B_*}^{-1/2}\hat{\Sigma}_{B_*B_*}^{(1)}\Sigma_{B_*B_*}^{-1/2}-{\mathbf{I}}_{|B_*|}\|_2,\] where we have \(\|\Sigma_{B_*B_*}\|_2\le M_1\) by the eigenvalue condition in 4 and Lemma 18 with \(k=|B_*|\) implies that with probability at least \(1-4p^{1-\gamma_*/2}\), we have \[\|\|\Sigma_{B_*B_*}^{-1/2}\hat{\Sigma}_{B_*B_*}^{(1)}\Sigma_{B_*B_*}^{-1/2}-{\mathbf{I}}_{|B_*|}\|_2\lesssim \sqrt{\frac{k_u\log p}{n_1}}.\] We denote this event by \({\cal E}_3\).

Combining the above bounds for \(\mathrm{I}\) and \(\mathrm{II}\) in 69 , we have arrived at the spectral norm of the estimation error of \(\hat{\Sigma}_{\text{spike}}\): \[\begin{align} \label{eq:32final32covariance32error} \bigl\|\hat{\Sigma}_{\text{spike}}-\Sigma\bigr\|_2 \lesssim \sqrt{\frac{k_u\log p}{n_1}} \end{align}\tag{77}\] conditioning on \({\cal E}_1\cap {\cal E}_2\cap {\cal E}_3\).

If \(n_1\ge c k_u\log p\) for a sufficiently large constant \(c>0\), then 77 implies \(\|\hat{\Sigma}_{\text{spike}}-\Sigma\|_2\le (2M_1)^{-1}\) with high probability. On \({\cal E}_1\cap {\cal E}_2\cap {\cal E}_3\), using \[\widehat{\Omega}-\Omega = \hat{\Sigma}_{\text{spike}}^{-1}-\Sigma^{-1} = \hat{\Sigma}_{\text{spike}}^{-1}\,(\Sigma-\hat{\Sigma}_{\text{spike}})\,\Sigma^{-1},\] we obtain \[\|\widehat{\Omega}-\Omega\|_2 \le \|\widehat{\Omega}\|_2\,\|\Omega\|_2\,\|\hat{\Sigma}_{\text{spike}}-\Sigma\|_2.\] By Weyl’s inequality and \(\lambda_{\min}(\Sigma)\ge 1/M_1\), \(\lambda_{\min}(\hat{\Sigma}_{\text{spike}}) \ge \lambda_{\min}(\Sigma)-\|\hat{\Sigma}_{\text{spike}}-\Sigma\|_2 \ge 1/M_1-\|\hat{\Sigma}_{\text{spike}}-\Sigma\|_2,\) hence \[\|\widehat{\Omega}\|_2 = \frac{1}{\lambda_{\min}(\hat{\Sigma}_{\text{spike}})} \le \frac{1}{1/M_1-\|\hat{\Sigma}_{\text{spike}}-\Sigma\|_2}.\] Combining the displays yields \[\|\widehat{\Omega}-\Omega\|_2 \le \frac{M_1\|\hat{\Sigma}_{\text{spike}}-\Sigma\|_2}{1/M_1-\|\hat{\Sigma}_{\text{spike}}-\Sigma\|_2} \lesssim \sqrt{\frac{k_u\log p}{n_1}}.\] Since \(n_1\asymp n\), this gives the stated rate \(\|\widehat{\Omega}-\Omega\|_2\lesssim \sqrt{(k_u\log p)/n}\). Moreover, by construction \(\hat{\Sigma}_{\text{spike}}\) equals \({\mathbf{I}}_p\) outside \(\widehat{B}\) and \(|\widehat{B}|\le k_u\). The proof is completed by choosing \(\gamma_*>2\). ◻

9.3 Proof of Lemma 20↩︎

Proof. Recall that \(\Sigma\in\Pi_0(k,p)\) with \(\Sigma={\mathbf{I}}_p+V\Lambda V^\top\) and \(B_*={\rm supp}(V)\) satisfying \(|B_*|=k\le k_u\). Let \(\hat{\Sigma}^{(1)}\) be defined in 65 . To prove that \(\mathbb{B}_{k_u}\neq\varnothing\) with high probability, it suffices to show that \(B_*\in\mathbb{B}_{k_u}\) with high probability.

From the definition of \(\mathbb{B}_{k_u}\) in 67 , \(\mathbb{P}(B_*\notin \mathbb{B}_{k_u})\) can be bounded by the sum of \[\label{eq:eq:32prob32B32notin32B32revised-ondiag} \mathbb{P}\Biggl(\exists D\subseteq B_*^c, |D|\le k_u:\; \|\hat{\Sigma}^{(1)}_D-{\mathbf{I}}_p\|_2> 2\left(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{\gamma_*|D|\log p}{n_1}}\right) +\left(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{\gamma_*|D|\log p}{n_1}}\right)^2 \Biggr),\tag{78}\] and \[\label{eq:eq:32prob32B32notin32B32revised-offdiag} \mathbb{P}\Biggl(\exists D\subseteq B_*^c, |D|\le k_u:\; \|\hat{\Sigma}^{(1)}_{DB_*}\|_2> \sqrt{\|\Gamma_{B_*}(\hat{\Sigma}^{(1)})\|_2}\biggl(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{|B_*|}{n_1}}+\sqrt{\frac{\gamma_* |D|\log p}{n_1}}\biggr) \Biggr).\tag{79}\]

We bound these two terms separately.

For the term in 78 , since \(D\subseteq B_*^c\) and \(\Sigma_{DD}={\mathbf{I}}_{DD}\), we have \(\hat{\Sigma}^{(1)}_{DD}\stackrel{d}{=}n_1^{-1}Y^\top Y\) for a matrix \(Y\) with i.i.d.\(\mathcal{N}(0,1)\) entries. Therefore, by 18 and a union bound, \[\begin{align} &\mathbb{P}\Biggl(\exists D\subseteq B_*^c,\;|D|\le k_u:\; \|\hat{\Sigma}^{(1)}_D-{\mathbf{I}}_p\|_2> 2\left(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{\gamma_*|D|\log p}{n_1}}\right) +\left(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{\gamma_*|D|\log p}{n_1}}\right)^2 \Biggr)\\ \le\;& \sum_{\substack{D\subseteq B_*^c\\ |D|\le k_u}} \mathbb{P}\Biggl(\|\hat{\Sigma}^{(1)}_D-{\mathbf{I}}_p\|_2> 2\left(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{\gamma_*|D|\log p}{n_1}}\right) +\left(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{\gamma_*|D|\log p}{n_1}}\right)^2 \Biggr)\\ \le\;& \sum_{\ell=1}^{k_u} { \binom{p-k_u}{\ell} } \,2\exp\!\left(-\frac{\gamma_*}{2}\,\ell\log p\right) \le 2\sum_{\ell=1}^{k_u} p^{\ell (1-\gamma_*/2)} \le 4 p^{1-\gamma_*/2}, \end{align}\] where the last inequality holds for all \(\gamma_*\ge 3\) and \(p\ge 2\).

For the term in 79 , fix \(D\subseteq B_*^c\) with \(|D|=\ell\le k_u\). Let \(W\) be the left singular vector matrix of \(X_{\cdot B_*}^{(1)}\). Then \[\bigl\|\hat{\Sigma}^{(1)}_{DB_*}\bigr\|_2 \le \frac{1}{n_1}\bigl\|\bigl(X_{\cdot D}^{(1)}\bigr)^\top W\bigr\|_2\, \bigl\|X_{\cdot B_*}^{(1)}\bigr\|_2 \stackrel{d}{=} \frac{1}{\sqrt{n_1}}\|Y\|_2\,\sqrt{\bigl\|\Gamma_{B_*}(\hat{\Sigma}^{(1)})\bigr\|_2},\] where \(Y\) is a \(\ell\times |B_*|\) matrix with i.i.d.\(\mathcal{N}(0,1)\) entries. The last equality in distribution is understood conditionally on \(X_{\cdot B_*}^{(1)}\), or equivalently conditionally on \(\Gamma_{B_*}(\hat{\Sigma}^{(1)})\). Indeed, by the block-diagonal structure of \(\Sigma_*\) over the \(D\)- and \(B_*\)-blocks, \(X_{\cdot D}^{(1)}\) is independent of \(X_{\cdot B_*}^{(1)}\), and hence is independent of \(W\), since \(W\) is constructed from \(X_{\cdot B_*}^{(1)}\). Therefore, conditional on \(W\), the Gaussian matrix \[\bigl(X_{\cdot D}^{(1)}\bigr)^\top W\] has the same distribution as \[\sqrt{n_1}\,Y\,\Gamma_{B_*}(\hat{\Sigma}^{(1)})^{1/2},\] with \(Y\) independent of \(\Gamma_{B_*}(\hat{\Sigma}^{(1)})\). Taking spectral norms and using \[\bigl\|X_{\cdot B_*}^{(1)}\bigr\|_2 = \sqrt{n_1}\sqrt{\bigl\|\Gamma_{B_*}(\hat{\Sigma}^{(1)})\bigr\|_2}\] gives the displayed bound.

Applying the Davidson–Szarek bound (19) and then taking a union bound over \(D\) gives \[\begin{align} &\mathbb{P}\left(\exists D\subseteq B_*^c,\;|D|\le k_u:\;\bigl\|\hat{\Sigma}^{(1)}_{DB_*}\bigr\|_2> \sqrt{\bigl\|\Gamma_{B_*}(\hat{\Sigma}^{(1)})\bigr\|}\biggl(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{|B_*|}{n_1}}+\sqrt{\frac{\gamma_* |D|\log p}{n_1}}\biggr)\right)\\ \le& \sum_{\substack{D\subseteq B_*^c\\ |D|\le k_u}} \mathbb{P}\left(\bigl\|\hat{\Sigma}^{(1)}_{DB_*}\bigr\|> \sqrt{\bigl\|\Gamma_{B_*}(\hat{\Sigma}^{(1)})\bigr\|}\biggl(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{|B_*|}{n_1}}+\sqrt{\frac{\gamma_* |D|\log p}{n_1}}\biggr)\right)\\ \le& \sum_{\substack{D\subseteq B_*^c\\ |D|\le k_u}} \mathbb{P}\left(\|Y\|_2>\sqrt{n_1}\biggl(\sqrt{\frac{|D|}{n_1}}+\sqrt{\frac{|B_*|}{n_1}}+\sqrt{\frac{\gamma_* |D|\log p}{n_1}}\biggr)\right)\\ \le& \sum_{\ell=1}^{k_u}\binom{p-k_u}{\ell}\,2\exp\!\left(-\frac{\gamma_*}{2}\,\ell\log p\right) \le 2\sum_{\ell=1}^{k_u} p^{\ell (1- {\gamma_*}/{2})} \le 4 p^{1-\gamma_*/2}, \end{align}\] where the last inequality holds for all \(\gamma_*\ge 3\) and \(p\ge 2\).

Combining the two bounds yields \(\mathbb{P}\bigl({\cal E}_1^c\bigr)=\mathbb{P}(B_*\notin\mathbb{B}_{k_u})\le 8p^{1-\gamma_*/2}\). Since \(B_*\in\mathbb{B}_{k_u}\) implies \(\mathbb{B}_{k_u}\neq\varnothing\), this completes the proof. ◻

10 Proof of Additional Technical Results for the Lower Bounds↩︎

10.1 Technical lemmas↩︎

Lemma 21 ([58]). Let \(g_i\) be the density function of \({\mathcal{N}}(0, \Sigma_i)\) for \(i = 0, 1, 2\), respectively. Then \[\int \frac{g_1 g_2}{g_0} = \left( \det \left( {\mathbf{I}}- \Sigma_0^{-1} (\Sigma_1 - \Sigma_0) \Sigma_0^{-1} (\Sigma_2 - \Sigma_0) \right) \right)^{-1/2}.\]

Lemma 22 ([21]). The following function \(\phi_\beta:\mathbb{R}\to\mathbb{R}\) is continuous and strictly decreasing: \[\phi_\beta(\beta)=\frac{\sum_{j=1}^p|\xi_j| \exp(-\beta/\xi_j^2)}{\sqrt{\sum_{j=1}^p \xi_j^2\exp(-\beta/\xi_j^2)}}.\]

The following lemma records a basic comparison between expectations under a probability measure and under the corresponding conditional measure on an event.

Lemma 23. Let \((\Omega,\mathcal{F},P)\) be a probability space. Let \(Q=P(\,\cdot\,\mid \mathcal{A})\) for some \(\mathcal{A}\in\mathcal{F}\) with \(P(\mathcal{A})>0\). Then for any nonnegative measurable \(X\), \[\mathbb{E}_{Q}[X] =\frac{1}{P(\mathcal{A})}\mathbb{E}_{P}\!\left[X\mathbf{1}\{\mathcal{A}\}\right] \le \frac{1}{P(\mathcal{A})}\mathbb{E}_{P}[X].\] If additionally \(P(\mathcal{A})>1/2\), then \[\mathbb{E}_{Q}[X]\le \bigl(1+2P(\mathcal{A}^{c})\bigr)\mathbb{E}_{P}[X].\]

Proof. The identity follows from the definition of \(Q\), and the first inequality uses \(X\mathbf{1}\{\mathcal{A}\}\le X\). For the second inequality, write \(P(\mathcal{A})=1-\varepsilon\) with \(\varepsilon=P(\mathcal{A}^c)<1/2\) and note that \(1/(1-\varepsilon)\le 1+2\varepsilon\). ◻

10.2 Proof of Lemma 10↩︎

Proof. We verify the conditions in the definition of the null space in 6 hold with probability 1 under the prior distribution \(\pi_1\).

Eigenvalues Control of \(\Sigma\): The covariance matrix of \(X\) is given by the lower-right \((p\times p)\) block of the covariance matrix of \(Z\), i.e. \(\Sigma^z\) defined in 34 : \[\Sigma = \left(\begin{array}{c|c} {\mathbf{I}}_{p_1\times p_1} & \boldsymbol{\delta}_1\boldsymbol{\delta}_2^\top \\ \hline \boldsymbol{\delta}_2\boldsymbol{\delta}_1^\top & {\mathbf{I}}_{p_2\times p_2} \end{array}\right).\] The largest and smallest eigenvalues of \(\Sigma\) are given by \(1\pm \|\boldsymbol{\delta}_1\|_2\|\boldsymbol{\delta}_2\|_2\), and the other \(p-2\) eigenvalues are all equal to 1. Specifically, we have \(\|\boldsymbol{\delta}_1\|_2=1\) and \[\label{eq:32valid32cov32alphab2} \|\boldsymbol{\delta}_2\|_2=c_1\sqrt{p_1\log p/n} \in [3^{-1}c_1\sqrt{k_u \log p/n}, 2^{-1}c_1\sqrt{k_u \log p/n}],\tag{80}\] where we have used \(p_1=\lfloor k_u/4\rfloor \in [k_u/9, k_u/4]\) for \(k_u\ge 4\). Since \(k_u\lesssim n/\log p\) from 3, we can choose \(c_1\) sufficiently small such that \(1/M_1\le \lambda_{\min}(\Sigma)\le \lambda_{\max}(\Sigma)\le M_1\) with probability 1.

Sparsity control of \(\beta\): For the covariance matrix \(\Sigma^z\) defined in 53 , we have \[\beta = \left( \begin{array}{cc} {\mathbf{I}}_{p_1\times p_1}& \boldsymbol{\delta}_1\boldsymbol{\delta}_2^\top\\ \boldsymbol{\delta}_2\boldsymbol{\delta}_1^\top & {\mathbf{I}}_{p_2\times p_2} \end{array} \right)^{-1} \left(\begin{array}{c} {{\mathbf{0}}}_{p_1\times 1}\\ \kappa\boldsymbol{\delta}_2 \end{array}\right),\] which leads to \[\label{eq:32beta32expression} \left\{ \begin{align} \beta_{S_1}&=-\frac{\kappa \|\boldsymbol{\delta}_2\|_2^2}{1-\|\boldsymbol{\delta}_2\|_2^2\|\boldsymbol{\delta}_1\|_2^2}\boldsymbol{\delta}_1,\\ \beta_{S_2}&=\frac{\kappa}{1-\|\boldsymbol{\delta}_2\|_2^2\|\boldsymbol{\delta}_1\|_2^2}\boldsymbol{\delta}_2. \end{align} \right.\tag{81}\] Therefore, we have \(\|\beta\|_0=\|\beta_{S_1}\|_0+\|\beta_{S_2}\|_0\le \|\boldsymbol{\delta}_1\|_0+\|\boldsymbol{\delta}_2\|_0\le k_u/4+k_u/4=k_u/2\).

Control of \(\sigma\): We verify that \(\sigma\) is bounded between \(0\) and \(M_2\), which implies the constructed \(\Sigma^z\) is positive definite. A direct calculation yields \[\begin{align} \sigma^2 & = \sigma_*^2-\beta^\top \Sigma\beta\\ & =\sigma_*^2-\frac{\kappa^2}{1-\|\boldsymbol{\delta}_2\|_2^2\|\boldsymbol{\delta}_1\|_2^2}\|\boldsymbol{\delta}_2\|_2^2\\ &\ge M_2^2/4 - \frac{2^{-2}\kappa^2 c_1^2 k_u\log p/n }{1-2^{-2}c_1^2 k_u\log p/n} \end{align}\] Since \(k_u\lesssim n/\log p\), we can choose \(c_1\) sufficiently small such that \[\label{eq:32temp32bound32of32delta1} 2^{-2} c_1^2 k_u\log p/n < \min(1/2, M_2^2/16).\tag{82}\]

Therefore, we have \[\label{eq:32bound32of32sigma32in32lem32of32valid32covariance} \sigma^2 \ge M_2^2/4 - \kappa^2 M_2^2/8.\tag{83}\]

Control of \(\kappa\): Recall that \(\kappa\) is defined as the unique solution to the linear equation \[\label{eq:32kappa32equation} \xi^\top \beta= c_2\sqrt{\sum_{j\le k_u}\xi_j^2}\frac{k_u\log p}{n}.\tag{84}\] Note that \[(\xi_{S_1})^\top \beta_{S_1}=\kappa \frac{-\|\boldsymbol{\delta}_2\|_2^2(\xi_{S_1})^\top \boldsymbol{\delta}_1}{1-\|\boldsymbol{\delta}_2\|_2^2\|\boldsymbol{\delta}_1\|_2^2}=\kappa \cdot \frac{\|\boldsymbol{\delta}_2\|_2^2 }{1-\|\boldsymbol{\delta}_2\|_2^2 } \sqrt{\sum_{j\le p_1}\xi_j^2}\] and \(\xi_{S_2}^\top \beta_{S_2}\) has the same sign as \(\kappa\), which is positive, we see that the coefficient of \(\kappa\) in 84 is positive. Consequently, the solution to the linear equation 84 exists and is unique. From 80 and 82 , we have \[3^{-2}c_1^2 \frac{k_u\log p}{n}\le \|\boldsymbol{\delta}_2\|_2^2 \le 2^{-2}c_1^2 \frac{k_u\log p}{n}\le 1/2.\] Therefore, the solution satisfies that \[c_2\sqrt{\sum_{j\le k_u}\xi_j^2}\frac{k_u\log p}{n} \ge (\xi_{S_1})^\top \beta_{S_1} \ge 2\kappa \|\boldsymbol{\delta}_2\|_2^2 \sqrt{\sum_{j\le p_1}\xi_j^2} \ge \frac{2}{9} \kappa c_1^2\sqrt{\sum_{j\le p_1}\xi_j^2}\frac{k_u\log p}{n}.\] Note that \(\sum_{p_1<j\le k_u}\xi_j^2\) is a sum of at most \(4p_1\) terms and each term is upper bounded by \(\xi_{p_1}^2\). Therefore, we have \[\sqrt{\sum_{j\le k_u}\xi_j^2}= \sqrt{\sum_{j\le p_1}\xi_j^2+\sum_{p_1<j\le k_u}\xi_j^2}\le \sqrt{5\sum_{j\le p_1}\xi_j^2}.\] Therefore, the solution \(\kappa\) satisfies \[\kappa\le 9\sqrt{5}c_2/(2c_1^2).\] Consequently, for any given value of \(c_1\), we choose \(c_2\) sufficiently small such that \(9\sqrt{5}c_2/(2c_1^2)\le 1\). It follows that \(\kappa\le 1\). Furthermore, the inequality in 83 guarantees that \(\sigma\in (M_2/\sqrt{8}, M_2)\). ◻

10.3 Proof of Lemma 11↩︎

Proof. For \(\Sigma_*^z\) defined in 51 and \(\Sigma^z=g_1(\boldsymbol{\delta}_2),\tilde{\Sigma}^z=g_1(\tilde{\boldsymbol{\delta}}_2)\) defined in 53 , we have \[\left(\Sigma_*^z\right)^{-1}=\left(\begin{array}{c|c|c} 1/{\sigma_*^2}&{\mathbf{0}}_{1\times p_1}&{\mathbf{0}}_{1\times p_2}\\ \hline {\mathbf{0}}_{p_1\times 1}&{\mathbf{I}}_{p_1\times p_1}&{\mathbf{0}}_{p_1\times p_2}\\ \hline {\mathbf{0}}_{p_2\times 1}&{\mathbf{0}}_{p_2\times p_1}&{\mathbf{I}}_{p_2\times p_2} \end{array}\right),\] and \[\left(\Sigma_*^z\right)^{-1}\left(\Sigma^z-\Sigma_*^z\right)=\left( \begin{array}{c|c|c} 0&{\mathbf{0}}_{1\times p_1}&{\kappa\boldsymbol{\delta}_2^\top}/{\sigma_*^2}\\ \hline {\mathbf{0}}_{p_1\times 1}&{\mathbf{0}}_{p_1\times p_1}&\boldsymbol{\delta}_1\boldsymbol{\delta}_2^\top\\ \hline \kappa\boldsymbol{\delta}_2&\boldsymbol{\delta}_2\boldsymbol{\delta}_1^\top&{\mathbf{0}}_{p_2\times p_2} \end{array}, \right)\] and \[\left(\Sigma_*^z\right)^{-1}\left(\Sigma^z-\Sigma_*^z\right)\left(\Sigma_*^z\right)^{-1}\left(\tilde{\Sigma}^z-\Sigma_*^z\right)=\left( \begin{array}{c|c} \Delta&{\mathbf{0}}_{(1+p_1)\times p_2}\\ \hline {\mathbf{0}}_{p_2\times (1+p_1)}&\displaystyle\left[\frac{\kappa\tilde{\kappa}}{\sigma_*^2}+\|\boldsymbol{\delta}_1\|_2^2\right]\boldsymbol{\delta}_2\tilde{\boldsymbol{\delta}}_2^\top \end{array} \right),\] where the upper left \((1+p_1)\times (1+p_1)\) block matrix is \[\Delta=\left(\begin{array}{c} \kappa \boldsymbol{\delta}_2^\top/\sigma_*^2\\ \boldsymbol{\delta}_1\boldsymbol{\delta}_2^\top \end{array}\right)\left(\begin{array}{cc} \tilde{\kappa} \tilde{\boldsymbol{\delta}}_2 \,&\, \tilde{\boldsymbol{\delta}}_2\boldsymbol{\delta}_1^\top \end{array}\right),\] and the right lower \(p_2\times p_2\) block matrix has rank no larger than 1 with the only nonzero eigenvalue (if \(\tilde{\boldsymbol{\delta}}_2^\top\boldsymbol{\delta}_2\neq 0\)) \[\left[\frac{\kappa\tilde{\kappa}}{\sigma_*^2}+\|\boldsymbol{\delta}_1\|_2^2\right]\tilde{\boldsymbol{\delta}}_2^\top\boldsymbol{\delta}_2.\] Similarly, \(\Delta\) has rank no larger than 1 and its only nonzero eigenvalue is given by \[\left(\begin{array}{cc} \tilde{\kappa} \tilde{\boldsymbol{\delta}}_2&\tilde{\boldsymbol{\delta}}_2\boldsymbol{\delta}_1^\top \end{array}\right)\left(\begin{array}{c} \kappa \boldsymbol{\delta}_2^\top/\sigma_*^2\\ \boldsymbol{\delta}_1\boldsymbol{\delta}_2^\top \end{array}\right)=\left[\frac{\kappa\tilde{\kappa}}{\sigma_*^2}+\|\boldsymbol{\delta}_1\|_2^2\right]\tilde{\boldsymbol{\delta}}_2^\top\boldsymbol{\delta}_2.\] Consequently, the matrix \[\left({\mathbf{I}}_{p+1}-\left(\Sigma_*^z\right)^{-1}\left(\Sigma^z-\Sigma_*^z\right)\left(\Sigma_*^z\right)^{-1}\left(\tilde{\Sigma}^z-\Sigma_*^z\right)\right)\] has two eigenvalues given by \(1-\left[\frac{\kappa\tilde{\kappa}}{\sigma_*^2}+\|\boldsymbol{\delta}_1\|_2^2\right]\tilde{\boldsymbol{\delta}}_2^\top\boldsymbol{\delta}_2\) and the rest \(p-1\) eigenvalues are all equal to 1; if \(\tilde{\boldsymbol{\delta}}_2^\top\boldsymbol{\delta}_2=0\), then all its eigenvalues are 1. Therefore, with 21, we have \[\label{eq:32chi32square32integral32bound} \begin{align} &\mathbb{E}_{(\theta,\tilde{\theta})\sim \pi_{}\times \pi_{}}\int_{\mathbb{R}^n}\frac{\,{\rm d}\,\mathbb{P}_{\theta}\,{\rm d}\,\mathbb{P}_{\tilde{\theta}}}{\,{\rm d}\,\mathbb{P}_{\theta_*}}\\ =\, &\mathbb{E}_{(\boldsymbol{\delta}_2,\tilde{\boldsymbol{\delta}}_2)\sim \pi_1\times \pi_1}\left[1-\left(\frac{\kappa\tilde{\kappa}}{\sigma_*^2}+\|\boldsymbol{\delta}_1\|_2^2\right)\tilde{\boldsymbol{\delta}}_2^\top\boldsymbol{\delta}_2\right]^{-n}\\ \le \,& \mathbb{E}_{(\boldsymbol{\delta}_2,\tilde{\boldsymbol{\delta}}_2)\sim {\pi}_1\times {\pi}_1}\exp\left(2n\left(\frac{\kappa\tilde{\kappa}}{\sigma_*^2}+\|\boldsymbol{\delta}_1\|_2^2\right)\tilde{\boldsymbol{\delta}}_2^\top\boldsymbol{\delta}_2\right), \end{align}\tag{85}\] where the last inequality follows from the fact that \((1-x)^{-1} \le \exp(2x)\) for \(x \in [0,1/2]\) and based on 52 we have \[0\le \left(\frac{\kappa\tilde{\kappa}}{\sigma_*^2}+\|\boldsymbol{\delta}_1\|_2^2\right)\tilde{\boldsymbol{\delta}}_2^\top\boldsymbol{\delta}_2\le \left(\frac{\kappa\tilde{\kappa}}{\sigma_*^2}+\|\boldsymbol{\delta}_1\|_2^2\right)c_1^2\frac{k_u\log p}{n}\le \frac{1}{2}\] if \(c_1\) is chosen sufficiently small (recall that \(k_u\log p/n\) is bounded). Since we have \(\kappa, \tilde{\kappa}\le 1\), \(\sigma_*=M_2/2\), and \(\|\boldsymbol{\delta}_1\|_2^2=1\), we can choose the constant \(c_3\) (for example \(c_3 = 2 (4/M_2^2 +1)\)) to complete the proof. ◻

10.4 Proof of Lemma 12↩︎

Proof. Since \(J\) is non-negative, we only need to show that \(\mathbb{E}[\exp (c(\log p) J)]\le 1+o(1)\) as \(p\) goes to infinity. Given any constant \(c\in (0, 1-2\gamma)\), let \[A_m=p^{cm}\mathbb{P}(J=m)=\exp(c\log p\cdot m) \frac{\binom{k}{m}\binom{p-k }{k - m}}{\binom{p }{k}}.\] Then we have \(\mathbb{E}[\exp (c(\log p) J)]=\sum_{m=0}^k A_m\). Notice that for \(0 \le m \le k-1\), \[\begin{align} \frac{A_{m+1}}{A_m} &= p^{c}\left[\frac{\binom{k}{m+1}}{\binom{k}{m}} \cdot \frac{\binom{p-k}{k-m-1}}{\binom{p-k}{k-m}}\right]\\ &= p^c \left[ \frac{(k-m)^2}{(m + 1)(p+m+1-2k)} \right] \\ &\le p^c \frac{(k-m)^2}{p-2k} \\ &\le p^c \frac{p^{2\gamma}}{p - 2p^\gamma} \\ &= \frac{p^{c + 2\gamma - 1}}{1 - 2p^{\gamma - 1}} =: r_p. \end{align}\] Since the constant \(c+2\gamma-1<0\) and \(\gamma-1<-1/2\), \(r_p\to 0\). In particular, for sufficiently large \(p\), \(r_p\le 1/2\) and \(A_m\le A_1 r_p^{m-1}\), which implies \[\label{eq:32Ak32sum32bound} \sum_{m=0}^{k} A_k = A_0 + \sum_{m=1}^{k} A_m \le A_0 + A_1 \sum_{m=0}^{k-1} 2^{-m} \le A_0 + 2A_1.\tag{86}\]

Notice that \[A_0 = \frac{\binom{p - k}{k}}{\binom{p }{k}} = \prod_{j=1}^{k} \frac{p - 2k + j}{p - k+j} = \prod_{j=1}^{k} \left(1 - \frac{k}{p - k + j}\right).\] Hence, \[\left(1 - \frac{k}{p - k}\right)^k \le A_0 \le \left(1 - \frac{k}{p}\right)^k.\] Since \(k^2/p \le p^{2\gamma - 1} \to 0\), both bounds converge to \(1\), and thus \(A_0 \to 1\). Furthermore, the inequality \(A_1\le r_p A_0\) implies that \(A_1=o(1)\). By 86 , the proof is completed. ◻

10.5 Proof of Lemma 13↩︎

Proof. The proof is similar to that of 10. It remains to verify that the following conditions hold with probability \(1-c/2\) under the prior distribution \(\pi_3\).

Eigenvalues control of \(\Sigma\): For the joint covariance matrix \(\Sigma^z\) defined in 57 , the covariance matrix for \(X\) is \(\Sigma={\mathbf{I}}_{p\times p}\), so the eigenvalues of \(\Sigma\) are always controlled.

Sparsity control of \(\beta\): Define \[\mu\stackrel{\triangle}{=}\sum_{j=1}^{k_\xi} q^{(1)}_j=c_4\frac{\sum_{j=1}^{k_\xi} |\xi_j|\exp(-\lambda^2/\xi_j^2)}{\sqrt{\sum_{j=1}^{k_\xi}\xi_j^2\exp(-\lambda^2/\xi_j^2)}}.\] Based on 17 , we have \[c_4\le \mu\le 2^{-1} c_4 k_u,\] where the second inequality follows from 22 and the fact that \(\zeta\le \lambda^2\).

Choosing \(c_4\le 1/2\), the bound \(\mu\le c_4 k_u/2\) gives \(\mu\le k_u/4\). Let \(S=\|\boldsymbol{\delta}\|_0\). \(S\) is a sum of independent Bernoulli variables with \[\mathbb{E}_{\pi_3}S=\mathbb{E}_{\pi_3} \sum_{j=1}^{k_\xi} b^{(1)}_j=\mu \le k_u/4.\] Since \(\beta=\kappa\boldsymbol{\delta}\), we have \(\|\beta\|_0\le S\). By Chernoff’s inequality for sums of independent Bernoulli variables (see [59]), we have \[\begin{align} \mathbb{P}_{\pi_4}\!\left(\|\beta\|_0\ge \frac{k_u}{2}\right) \le & \; \mathbb{P}_{\pi_3}(S\ge k_u/2) \\ \le & \exp\left\{-\mu+\frac{k_u}{2}\log\left(\frac{e\mu}{k_u/2}\right)\right\}\\ \le & \exp\left\{- \frac{k_u}{4}+\frac{k_u}{2}\log\left(\frac{e}{2}\right)\right\} \\ \le & \exp\left\{- \frac{\log(4/e)k_u}{8}\right\}, \end{align}\] where the third inequality follows from the fact that the function \(x\mapsto -x+(k_u/2)\log(ex/(k_u/2))\) is increasing on \((0, (k_u/2)]\) and \(\mu\le k_u/4\). Finally, choose \(C_4\) sufficiently large so that \[\exp\left(- \frac{\log(4/e)C_4}{8}\right)\le \frac{c}{6}.\] Under the condition that \(k_u\ge C_4\), we have \[\mathbb{P}_{\pi_4}\!\left(\|\beta\|_0\ge \frac{k_u}{2}\right) \le \frac{c}{6}.\]

Control of \(\kappa\): Note that \(\beta=\kappa \boldsymbol{\delta}\). To show that the unique solution to \(\xi^\top \beta =\tau\) satisfies \(\kappa\le 1\) for \(\tau=c_6\nu_1/\sqrt{n}\), we only need to show \(\xi^\top \boldsymbol{\delta}\ge c_6\nu_1/\sqrt{n}\) holds. Here \(c_6\) is determined by the values of \(c_4\) and \(c_5\).

For \(\boldsymbol{\delta}\sim \pi_3\), we have \[\xi^\top \boldsymbol{\delta}\stackrel{{\rm d}}{=}\sum_{j=1}^{k_\xi}\frac{c_5\gamma^{(1)}_j\xi_j}{\sqrt{n}}b^{(1)}_j,\] where \(b^{(1)}_j\sim {\rm Bernoulli}(q^{(1)}_j)\). Since \(\xi_j\gamma^{(1)}_j= \max(|\xi_j|, \lambda)\) is non-increasing in \(j\), we have \[\begin{align} {\rm Var}_{\pi_3} [\xi^\top \boldsymbol{\delta}] & = \frac{c_5^2}{n}\sum_{j=1}^{k_\xi}\xi_j^2(\gamma^{(1)}_j)^2q^{(1)}_j[1-q^{(1)}_j] \\ & \le \frac{c_5^2\xi_1\gamma_1^{(1)}}{n}\sum_{j=1}^{k_\xi}\xi_j\gamma^{(1)}_jq^{(1)}_j\\ & = \frac{c_5}{\sqrt{n}}\max(|\xi_1|, \lambda)\mathbb{E}_{\pi_3}\xi^\top \boldsymbol{\delta}. \end{align}\] Using \(\xi_j\gamma^{(1)}_j\ge |\xi_j|\), we have \[\mathbb{E}_{\pi_3} \xi^\top \boldsymbol{\delta}=\frac{c_5}{\sqrt{n}}\sum_{j=1}^{k_\xi}\xi_j\gamma^{(1)}_jq^{(1)}_j\ge \frac{c_5}{\sqrt{n}} \sum_{j=1}^{k_\xi}|\xi_j|q^{(1)}_j= \frac{c_4c_5}{\sqrt{n}}\sqrt{\sum_{j=1}^{k_\xi}\xi_j^2\exp(-\lambda^2/\xi_j^2)}.\] Using \(\xi_j\gamma^{(1)}_j\ge \lambda\), we have \[\mathbb{E}_{\pi_3} \xi^\top \boldsymbol{\delta}=\frac{c_5}{\sqrt{n}} \sum_{j=1}^{k_\xi}\xi_j\gamma^{(1)}_jq^{(1)}_j\ge \frac{c_5}{\sqrt{n}} \lambda \sum_{j=1}^{k_\xi}q^{(1)}_j.\] If \(\lambda>0\), 17 implies that \(\sum_{j=1}^{k_\xi}q^{(1)}_j= c_4 k_u\) and thus \(\mathbb{E}_{\pi_3} \xi^\top \boldsymbol{\delta}\ge \frac{c_4 c_5 }{\sqrt{n}} \lambda k_u\). This inequality also holds when \(\lambda=0\). Consequently, we have \[\label{eq:32control32mean32of32LF32in32lem32of32valid32covariance322} \mathbb{E}_{\pi_3}\xi^\top \boldsymbol{\delta}\ge \frac{c_4 c_5}{2\sqrt{n}} \nu_1>0.\tag{87}\]

By Chebyshev’s inequality, we have \[\label{eq:32bounding32LF32in32lem32of32valid32covariance322} \begin{align} \mathbb{P}_{\pi_3}\left(\xi^\top \boldsymbol{\delta}< \frac{1}{2}\mathbb{E}_{\pi_3}[\xi^\top \boldsymbol{\delta}]\right)& \le \frac{4{\rm Var}_{\pi_3} (\xi^\top \boldsymbol{\delta})}{\left[\mathbb{E}_{\pi_3}(\xi^\top \boldsymbol{\delta})\right]^2}\le \frac{8}{c_4}\frac{\max(|\xi_1|,\lambda)}{\nu_1}. \end{align}\tag{88}\] We pick \(C_4 = 48/(c c_4)\). Recall the conditions that \(\nu_1\ge C_4 |\xi_1|\), \(k_u\ge C_4\), and \(\nu_1\ge \lambda k_u\), we have \(\max(|\xi_1|,\lambda)/\nu_1\le \max(1/C_4,1/k_u)\le 1/C_4\le c c_4 /48\). 87 and 88 together imply that \[\mathbb{P}_{\pi_4}\left(\kappa\in (0,1]\right)\ge \mathbb{P}_{\pi_3}\left(\xi^\top \delta\ge c_6\nu_1/\sqrt{n}\right)\ge 1- c/6\] with \(c_6=c_4 c_5 /4\).

Control of \(\sigma\): By the construction of \(\Sigma^z\), we have \(\sigma^2=\sigma_*^2-\beta^{\top} \Sigma \beta=\sigma_*^2-\kappa^2 \boldsymbol{\delta}^\top \boldsymbol{\delta}\); our goal is to show that the right-hand side is positive, so that \(\Sigma^z\) is positive definite. In the last part, we have shown that with probability \(1-c/6\), the event \(\{\xi^\top \delta\ge c_6\nu_1/\sqrt{n}\}\) (and thus \(\kappa\le 1\)) holds. It remains to show that \(\|\boldsymbol{\delta}\|_2\) is smaller than \(\sigma_*\) with probability close to 1. Note that \[\|\boldsymbol{\delta}\|_2^2\stackrel{{\rm d}}{=}\frac{c_5^2}{n}\sum_{j=1}^{k_\xi}b^{(1)}_j(\gamma^{(1)}_j)^2\] with \(b^{(1)}_j\sim {\rm Bernoulli}(q^{(1)}_j)\). By the definitions of \(q^{(1)}_j\) and \(\gamma^{(1)}_j\), we have \[\begin{align} &\sum_{j=1}^{k_\xi}(q^{(1)}_j)^2\exp((\gamma^{(1)}_j)^2)\\ \le &c_4^2\frac{\sum_{j\le j_1} \xi_j^2\exp(-2\lambda^2/\xi_j^2) e +\sum_{j>j_1}\xi_j^2\exp(-2\lambda^2/\xi_j^2) \exp(\lambda^2/\xi_j^2)}{\sum_{i=1}^p\xi_i^2\exp(-\lambda^2/\xi_i^2)}\\ \le &c_4^2 e . \end{align}\] Therefore, we have \((\gamma^{(1)}_j)^2\le \log(c_4^2 e / (q^{(1)}_j)^2)\) and \[\begin{align} \mathbb{E}\|\boldsymbol{\delta}\|_2^2&=\frac{c_5^2}{n}\sum_{j=1}^{k_\xi}q^{(1)}_j(\gamma^{(1)}_j)^2\le \frac{c_5^2}{n}\sum_{j=1}^{k_\xi}q^{(1)}_j\log \left(\frac{c_4^2 e}{(q^{(1)}_j)^2}\right). \end{align}\] Since the function \(x\log(c_4^2 e /x^2)\) is concave, by Jensen’s inequality, we have \[\begin{align} \mathbb{E}\|\boldsymbol{\delta}\|_2^2&\le \frac{c_5^2}{n}\left(\sum_{j=1}^{k_\xi}q_j^{(1)}\right)\log \left(\frac{ e c_4^2 k_\xi^2}{\left(\sum_{j=1}^{k_\xi} q_j^{(1)}\right)^2}\right) \end{align}\] In the previous analysis, we have \(\mu=\sum_{j=1}^{k_\xi}q^{(1)}_j\in [c_4, c_4 k_u]\), so we further have \[\begin{align} \mathbb{E}\|\boldsymbol{\delta}\|_2^2\le 2 c_4c_5^2\frac{k_u\log ( e p)}{n}. \end{align}\] By Markov’s inequality, \[\mathbb{P}_{\pi_3}\left( \|\boldsymbol{\delta}\|_2 > \sigma_*/\sqrt{2} \right) \le 2\sigma_*^{-2} \mathbb{E}\|\boldsymbol{\delta}\|_2^2 \le 16 M_2^{-2} c_4c_5^2\frac{k_u\log ( e p)}{n}\] Note that \(k_u\log p\lesssim n\) based on 3. Given \(c_5\), we can choose \(c_4\) sufficiently small so that the right-hand side on the last inequality is bounded by \(c/6\).

Combining the above analyses, we complete the proof. ◻

10.6 Proof of Lemma 14↩︎

Proof. For \(\Sigma_*^z\) defined in 51 and \(\Sigma^z=g_2(\boldsymbol{\delta}),\tilde{\Sigma}^z=g_2(\tilde{\boldsymbol{\delta}})\) defined in 57 , we have \(\kappa,\tilde{\kappa}\in [0,1]\) and \[\left(\Sigma_*^z\right)^{-1}\left(\Sigma^z-\Sigma_*^z\right)=\left( \begin{array}{c|c} 0& {\kappa\boldsymbol{\delta}^\top}/{\sigma_*^2}\\ \hline \kappa\boldsymbol{\delta}& {\mathbf{0}}_{p\times p} \end{array} \right).\] Therefore, we have \[\left(\Sigma_*^z\right)^{-1}\left(\Sigma^z-\Sigma_*^z\right)\left(\Sigma_*^z\right)^{-1}\left(\tilde{\Sigma}^z-\Sigma_*^z\right)=\left( \begin{array}{c|c} \frac{\kappa\tilde{\kappa}}{\sigma_*^2}\boldsymbol{\delta}^\top \tilde{\boldsymbol{\delta}}& {\mathbf{0}}_{1\times p}\\ \hline {\mathbf{0}}_{p\times 1}& \frac{\kappa\tilde{\kappa}}{\sigma_*^2}\boldsymbol{\delta}\tilde{\boldsymbol{\delta}}^\top \end{array} \right).\] Consequently, the matrix \[\left({\mathbf{I}}_{p+1}-\left(\Sigma_*^z\right)^{-1}\left(\Sigma^z-\Sigma_*^z\right)\left(\Sigma_*^z\right)^{-1}\left(\tilde{\Sigma}^z-\Sigma_*^z\right)\right)\] has two non-unit eigenvalues given by \(1-\frac{\kappa\tilde{\kappa}}{\sigma_*^2}\boldsymbol{\delta}^\top \tilde{\boldsymbol{\delta}}\) and the rest \(p-1\) eigenvalues are all equal to 1.

When \(\tilde{\kappa}\kappa=0\), the desired result holds obviously.

When \(\tilde{\kappa} \kappa\in (0,1]\), by definition, we have \(\tilde{\boldsymbol{\delta}}, \boldsymbol{\delta}\in \mathcal{G}_\tau\). By definition of \(\mathcal{G}_\tau\), we have \[\frac{\tilde{\kappa} \kappa}{\sigma_*^2}\boldsymbol{\delta}^\top \tilde{\boldsymbol{\delta}}\le \frac{1}{2}.\] Therefore, with 21, we have \[\label{eq:32chi32square32integral32bound322} \begin{align} &\mathbb{E}_{(\theta,\tilde{\theta})\sim \pi_4\times \pi_4}\int_{\mathbb{R}^n}\frac{\,{\rm d}\,\mathbb{P}_{\theta}\,{\rm d}\,\mathbb{P}_{\tilde{\theta}}}{\,{\rm d}\,\mathbb{P}_{\theta_*}}\\ =&\mathbb{E}_{(\boldsymbol{\delta},\tilde{\boldsymbol{\delta}})\sim \pi_3\times \pi_3}\left[1-\frac{\kappa\tilde{\kappa}}{\sigma_*^2}\boldsymbol{\delta}^\top \tilde{\boldsymbol{\delta}}\right]^{-n}\\ \le &\mathbb{E}_{(\boldsymbol{\delta},\tilde{\boldsymbol{\delta}})\sim {\pi}_3\times {\pi}_3}\exp\left(\frac{2n\kappa\tilde{\kappa}}{\sigma_*^2}\boldsymbol{\delta}^\top \tilde{\boldsymbol{\delta}}\right), \end{align}\tag{89}\] where the last inequality follows from the fact that \((1-x)^{-1} \le \exp(2x)\) for \(x \in [0,1/2]\). Since \(\kappa\tilde{\kappa}\le1\), we complete the proof with \(c_7=2/\sigma_*^2\). ◻

10.7 Proof of Lemma 15↩︎

Proof. The proof is similar to that of 10. It remains to verify that the following conditions hold with probability arbitrarily close to 1 under the prior distribution \(\pi_5\):

Sparsity control of \(\boldsymbol{\delta}_2\): For the random vector \(\boldsymbol{\delta}_2\) defined in 59 , the sparsity level of \(\boldsymbol{\delta}_2\) is a sum of independent Bernoulli random variables with parameters \(q^{(2)}_j\). Let \[\mu_2\stackrel{\triangle}{=}\sum_{j\in S_3}q_j^{(2)}=\frac{k_u}{8}\frac{\sum_{j\in S_5}|\xi_j|}{\sqrt{p_5\sum_{j\in S_5}\xi_j^2}}\le \frac{k_u}{8},\] where the inequality follows from Cauchy-Schwarz inequality. Consequently, based on Chernoff’s inequality [59], we have \[\begin{align} \mathbb{P}(\|\boldsymbol{\delta}_2\|_0\ge \frac{k_u}{4})&\le \exp\left[-\mu_2+\frac{k_u}{4}\ln\left(\frac{4e\mu_2}{k_u}\right)\right]\\ &\stackrel{(*)}\le \exp\left[-\frac{k_u}{8}+\frac{k_u}{4}\ln(\frac{e}{2})\right]\le \exp\left[- \frac{\ln(4/e) }{8}k_u\right], \end{align}\] where (*) follows from the fact that the function \(x\mapsto -x+(k_u/4)\ln x\) is increasing when \(x\le k_u/4\). Therefore, we have \(\|\boldsymbol{\delta}_2\|_0\le k_u/4\) with high probability.

Eigenvalues Control of \(\Sigma\): Similar to 10.2, \(\Sigma\) has largest and smallest eigenvalues given by \(1\pm \|\boldsymbol{\delta}_2\|_2\|\boldsymbol{\delta}_1\|_2\) and the remaining \(p-2\) eigenvalues are all equal to 1. Specifically, we have \[\|\boldsymbol{\delta}_1\|_2=c_8\sqrt{\frac{k_u\log p}{n}},\quad \text{and}\quad \|\boldsymbol{\delta}_2\|_2=\sqrt{\|\boldsymbol{\delta}_2\|_0}\frac{\sqrt{p_5}}{k_u}.\] Since \(\|\boldsymbol{\delta}_2\|_0\le k_u/4\) with high probability from the analysis above, we have \[\|\boldsymbol{\delta}_1\|_2\|\boldsymbol{\delta}_2\|_2\le \frac{c_8}{2} \sqrt{\frac{p_5\log p}{n}}.\] Since \(p_5\le k_{\rm eff}\lesssim n/\log p\), by choosing \(c_8\) sufficiently small, we have \(1/M_1\le \lambda_{\min}(\Sigma)\le \lambda_{\max}(\Sigma)\le M_1\) with high probability.

Sparsity control of \(\beta\): Following the analysis in 10.2, we have \[\begin{align} \beta_{S_3}&=-\frac{\kappa \|\boldsymbol{\delta}_1\|_2^2}{1-\|\boldsymbol{\delta}_1\|_2^2\|\boldsymbol{\delta}_2\|_2^2}\boldsymbol{\delta}_2,\\ \beta_{S_4}&=\frac{\kappa}{1-\|\boldsymbol{\delta}_1\|_2^2\|\boldsymbol{\delta}_2\|_2^2}\boldsymbol{\delta}_1. \end{align}\] Note that \(\|\boldsymbol{\delta}_1\|_0\le k_u/4\) and \(\|\boldsymbol{\delta}_2\|_0\le k_u/4\) with high probability. Therefore, we have \(\|\beta\|_0= \|\boldsymbol{\delta}_2\|_0+\|\boldsymbol{\delta}_1\|_0\le k_u/2\) with high probability.

Control of \(\kappa\): We just need to show that there exists some constant \(c_9>0\) such that \[\label{eq:32kappa32equation322} \frac{-\|\boldsymbol{\delta}_1\|_2^2\xi_{S_3}^\top \boldsymbol{\delta}_2}{1-\|\boldsymbol{\delta}_1\|_2^2\|\boldsymbol{\delta}_2\|_2^2}+\frac{\xi_{S_4}^\top \boldsymbol{\delta}_1}{1-\|\boldsymbol{\delta}_1\|_2^2\|\boldsymbol{\delta}_2\|_2^2}\ge c_9\nu_3\frac{k_u\log p}{n}\tag{90}\] with high probability. Specifically, we have \(\xi_{S_4}^\top \boldsymbol{\delta}_1\ge 0\) based on 59 and \[\frac{\|\boldsymbol{\delta}_1\|_2^2}{1-\|\boldsymbol{\delta}_1\|_2^2\|\boldsymbol{\delta}_2\|_2^2}\asymp \|\boldsymbol{\delta}_1\|_2^2=c_8^2\frac{k_u\log p}{n}.\] Moreover, we have \[\begin{align} -\xi_{S_3}^\top \boldsymbol{\delta}_2&=\sum_{j\in S_3}|\xi_j| \frac{\sqrt{p_5}}{k_u}b^{(2)}_j, \end{align}\] where \(b^{(2)}_j\sim {\rm Bernoulli}(q^{(2)}_j)\). Therefore, we have \[\begin{align} \mathbb{E}(-\xi_{S_3}^\top \boldsymbol{\delta}_2)&=\frac{\sqrt{p_5}}{k_u}\sum_{j\in S_5}|\xi_j| q^{(2)}_j=\frac{1}{8}\sqrt{\sum_{j\in S_5}\xi_j^2}.\\ {\rm Var}(-\xi_{S_5}^\top \boldsymbol{\delta}_2)&=\frac{p_5}{k_u^2}\sum_{j\in S_3}\xi_j^2q^{(2)}_j(1-q^{(2)}_j)\le \frac{\sqrt{p_5}}{8k_u}\frac{\sum_{j \in S_5}|\xi_j|^3}{\sqrt{\sum_{j\in S_5}\xi_j^2}}. \end{align}\] Note that \(p_5\le k_{\rm eff}\lesssim k_u^2/\log p\) and \[\sum_{j\in S_5}|\xi_j|^3\le |\xi_1|\sum_{j\in S_5}\xi_j^2\le \left(\sum_{j\in S_5}\xi_j^2\right)^{3/2}.\] Therefore, we have \[{\rm Var}(-\xi_{S_3}^\top \boldsymbol{\delta}_2)\lesssim \frac{1}{\log p}\left(\mathbb{E}(-\xi_{S_3}^\top \boldsymbol{\delta}_2)\right)^2.\] By Chebyshev’s inequality, we have for any constant \(c>0\), with probability asymptotically at least \(1-c/4\), we have \[-\xi_{S_3}^\top \boldsymbol{\delta}_2\ge \frac{1}{16}\sqrt{\sum_{j\in S_5}\xi_j^2}.\] Based on the assumption at the beginning of 8.3, we have \[\sum_{j\in S_5}\xi_j^2\ge \frac{1}{2}\sum_{j\in S_3}\xi_j^2=\frac{1}{2}\nu_3^2\] Therefore, we can choose \(c_9>0\) in 90 sufficiently small such that we have \(0<\kappa\le 1\) with large probability.

Control of \(\sigma\): Finally, we verify that \(\sigma\) is bounded between \(0\) and \(M_2\). Note that \[\sigma^2=\sigma_*^2-\beta^\top \Sigma\beta=\sigma_*^2-\frac{\kappa^2}{1-\|\boldsymbol{\delta}_1\|_2^2\|\boldsymbol{\delta}_2\|_2^2}\|\boldsymbol{\delta}_1\|_2^2.\] Combining the previous analysis, we can choose \(c_8\) sufficiently small such that \(\sigma\in (0,M_2)\). ◻

10.8 Proof of Lemma 16↩︎

Proof. There are two statements to prove: \[\label{eq:32statement321} \mathbb{E}_{\theta_*}(L_\theta^{\le D}, L_{\tilde{\theta}}^{\le D})\le \mathbb{E}_{\theta_*}(L_\theta, L_{\tilde{\theta}})\tag{91}\] and \[\label{eq:32statement322} \mathbb{E}_{\theta_*}\left((L_{\theta}^{\le D})\right)^2\le 9(6npD)^{4D}.\tag{92}\] Without loss of generality, we scale the problem such that \(\sigma_*=1\). Then \(\mathbb{P}_{\theta_*}^n\) is the standard Gaussian distribution on \(N={n\times(p+1)}\) dimensions and we denote it by \(\mathbb{Q}\). Then the space \(L^2(\mathbb{Q})\) admits the orthogonal basis of Hermite polynomials, see [60] for a standard reference. We denote these by \((H_{\alpha})_{\alpha\in\mathbb{N}^{N}}\) (where \(0\in\mathbb{N}\) by convention), where \[H_{\alpha}(z) = \prod_{i=1}^{N} h_{\alpha_{i}}(z_{i})\] for univariate Hermite polynomials \((h_{j})_{j\in\mathbb{N}}\), and \(z \in \mathbb{R}^{N}\). We adopt the normalization where \(\|H_{\alpha}\|_{\mathbb{Q}} = 1\), which is not usually the standard convention in the literature. This basis is graded in the sense that for any \(D \in \mathbb{N}\), \((H_{\alpha})_{\alpha\in\mathbb{N}^{N},\,|\alpha|\le D}\) is an orthonormal basis for the polynomials of degree at most \(D\), where \(|\alpha| := \sum_{i=1}^{N} \alpha_{i}\).

Regarding \(\mathbb{P}_\theta^n\), we will make use of the following fact: Since \(\theta=h(g_3(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2))\), after the scaling \(\sigma_*=1\), the diagonal entries of the covariance matrix in 60 are all equal to one. Hence each coordinate \(Z_i\) has standard normal marginal distribution under \(\mathbb{P}_\theta^n\).

Proof of 91 : Throughout this argument, \(\theta\) and \(\tilde{\theta}\) are generated by the restricted prior used in 8.3, and we work on the validity event from 15; in particular, their associated coefficients satisfy \[0<\kappa\le 1, \qquad 0<\tilde{\kappa}\le 1.\] Expanding in the orthonormal Hermite basis \(\{H_{\alpha}\}\), we have \[\label{eq:32low32degree32expansion} \begin{align} \left\langle L^{\le D}_\theta,\, L^{\le D}_{\tilde{\theta}} \right\rangle_{\mathbb{Q}} &= \sum_{\alpha\in\mathbb{N}^{N},\, |\alpha|\le D} \left\langle L_\theta, H_{\alpha} \right\rangle_{\mathbb{Q}} \left\langle L_{\tilde{\theta}}, H_{\alpha} \right\rangle_{\mathbb{Q}} \\ &= \sum_{\alpha\in\mathbb{N}^{N},\, |\alpha|\le D} \mathbb{E}_{Z \sim \mathbb{P}_\theta^n}\!\left[H_{\alpha}(Z)\right]\, \mathbb{E}_{Z \sim \mathbb{P}_{\tilde{\theta}}^n}\!\left[H_{\alpha}(Z)\right], \end{align}\tag{93}\] where we have used the change-of-measure identity \(\mathbb{E}_{\mathbb{Q}}[ L_\theta f ] = \mathbb{E}_{\mathbb{P}_\theta^n}[ f ]\).

We claim that for every Hermite multi-index \(\alpha\), \[\label{eq:32claim32low32degree} \mathbb{E}_{Z \sim \mathbb{P}_\theta^n}\!\left[H_{\alpha}(Z)\right]\, \mathbb{E}_{Z \sim \mathbb{P}_{\tilde{\theta}}^n}\!\left[H_{\alpha}(Z)\right] \ge 0 .\tag{94}\]

Lemma 24 (Rank-one Hermite covariance identity). Let \((U,V)\) be a centered Gaussian vector satisfying \[\operatorname{Cov}(U)=I_a,\qquad \operatorname{Cov}(V)=I_b,\qquad \operatorname{Cov}(U,V)=r c^\top .\] For multi-indices \(\mu\in\mathbb{N}^{a}\) and \(\nu\in\mathbb{N}^{b}\), write \[h_\mu(U)=\prod_{\ell=1}^a h_{\mu_\ell}(U_\ell), \qquad h_\nu(V)=\prod_{j=1}^b h_{\nu_j}(V_j),\] where the univariate Hermite polynomials have the normalization fixed above. Then \[\mathbb{E}\{h_\mu(U)h_\nu(V)\}=0 \qquad \text{if } |\mu|\ne|\nu|,\] and, if \(|\mu|=|\nu|=m\), then \[\mathbb{E}\{h_\mu(U)h_\nu(V)\} = \frac{m!}{\sqrt{\mu!\nu!}}\, r^\mu c^\nu ,\] where \[\mu!=\prod_{\ell=1}^a \mu_\ell!, \qquad \nu!=\prod_{j=1}^b \nu_j!, \qquad r^\mu=\prod_{\ell=1}^a r_\ell^{\mu_\ell}, \qquad c^\nu=\prod_{j=1}^b c_j^{\nu_j}.\]

Proof of 24. Let \[G(s,t) = \mathbb{E}\left[ \prod_{\ell=1}^a \exp\{s_\ell U_\ell-s_\ell^2/2\} \prod_{j=1}^b \exp\{t_j V_j-t_j^2/2\} \right] .\] Since \((U,V)\) is centered Gaussian, \[\begin{align} G(s,t) = &\mathbb{E}\left[ \prod_{\ell=1}^a \exp\{s_\ell U_\ell-s_\ell^2/2\} \prod_{j=1}^b \exp\{t_j V_j-t_j^2/2\} \right] \\ \qquad = & \mathbb{E}\left[\exp(s^{\top} U+t^{\top} V)\right] \exp(-\|s\|^2/2-\|t\|^2/2)\\ \qquad = & \exp(s^\top r c^\top t), \end{align}\] where the last equation follows from \(\operatorname{Var}\left(s^{\top} U+t^{\top} V\right)=\|s\|^2+\|t\|^2+2 s^{\top} r c^{\top} t\).

Hence \[\begin{align} G(s,t) &= \sum_{k=0}^\infty \frac{1}{k!}(s^\top r)^k(c^\top t)^k \\ &= \sum_{k=0}^\infty \sum_{|\mu|=k} \sum_{|\nu|=k} \frac{k!}{\mu!\nu!} r^\mu c^\nu s^\mu t^\nu . \end{align}\] Thus the coefficient of \(s^\mu t^\nu\) in \(G(s,t)\) is zero unless \(|\mu|=|\nu|\). If \(|\mu|=|\nu|=m\), that coefficient is \[\frac{m!}{\mu!\nu!}r^\mu c^\nu .\]

On the other hand, by the generating function of the normalized Hermite polynomials, \[\exp\{sx-s^2/2\} = \sum_{k=0}^\infty h_k(x)\frac{s^k}{\sqrt{k!}},\] we have \[\begin{align} G(s,t) &= \sum_{\mu,\nu} \mathbb{E}\{h_\mu(U)h_\nu(V)\} \frac{s^\mu t^\nu}{\sqrt{\mu!\nu!}} . \end{align}\] Comparing the coefficient of \(s^\mu t^\nu\) gives \[\frac{\mathbb{E}\{h_\mu(U)h_\nu(V)\}}{\sqrt{\mu!\nu!}} = \frac{m!}{\mu!\nu!}r^\mu c^\nu\] when \(|\mu|=|\nu|=m\), and gives zero otherwise. Therefore \[\mathbb{E}\{h_\mu(U)h_\nu(V)\} = \frac{m!}{\sqrt{\mu!\nu!}}r^\mu c^\nu\] when \(|\mu|=|\nu|=m\), while the expectation is zero when \(|\mu|\ne |\nu|\). This proves the lemma. ◻

We apply 24 row by row. After the rescaling \(\sigma_*=1\), write each observation under the covariance model \(g_3\) as \[Z_i=(Y_i,X_{i,S_3},X_{i,S_4}),\qquad i=1,\ldots,n.\] Following the coordinate convention used in the construction of \(g_3\) in 8.3, \(\boldsymbol{\delta}_1\in\mathbb{R}^{p_4}\) is the coefficient vector paired with \(X_{i,S_4}\), and \(\boldsymbol{\delta}_2\in\mathbb{R}^{p_3}\) is the coefficient vector paired with \(X_{i,S_3}\). For \(\theta=h(g_3(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2))\), the covariance matrix of a single observation, denoted by \[\Sigma_\theta^z= \begin{pmatrix} 1 & 0 & \kappa\boldsymbol{\delta}_1^\top\\ 0 & I_{p_3} & \boldsymbol{\delta}_2\boldsymbol{\delta}_1^\top\\ \kappa\boldsymbol{\delta}_1 & \boldsymbol{\delta}_1\boldsymbol{\delta}_2^\top & I_{p_4} \end{pmatrix}.\] Define \[U_i=(Y_i,X_{i,S_3})\in\mathbb{R}^{1+p_3}, \qquad V_i=X_{i,S_4}\in\mathbb{R}^{p_4}.\] Then \[\operatorname{Cov}_\theta(U_i)=I_{1+p_3}, \qquad \operatorname{Cov}_\theta(V_i)=I_{p_4}, \qquad \operatorname{Cov}_\theta(U_i,V_i)=r_\theta c_\theta^\top,\] where \[r_\theta=(\kappa,\boldsymbol{\delta}_2^\top)^\top\in\mathbb{R}^{1+p_3}, \qquad c_\theta=\boldsymbol{\delta}_1\in\mathbb{R}^{p_4}.\] The same representation holds for \(\tilde{\theta}\) with \[r_{\tilde{\theta}}=(\tilde{\kappa},\tilde{\boldsymbol{\delta}}_2^\top)^\top, \qquad c_{\tilde{\theta}}=\tilde{\boldsymbol{\delta}}_1.\]

The signs of the corresponding coordinates of these rank-one factors agree. Let \(\tilde{I}\) and \(\tilde{b}_j^{(2)}\) denote the support and Bernoulli variables used to construct \(\tilde{\boldsymbol{\delta}}_1\) and \(\tilde{\boldsymbol{\delta}}_2\). By the construction in 59 , \[(\boldsymbol{\delta}_1)_j = c_8\operatorname{sign}(\xi_{p_3+j}) \sqrt{\frac{\log p}{n}}\,\mathbf{1}\{j\in I\},\] and the same formula holds for \(\tilde{\boldsymbol{\delta}}_1\) with \(\tilde{I}\) in place of \(I\). Hence \[(\boldsymbol{\delta}_1)_j(\tilde{\boldsymbol{\delta}}_1)_j\ge 0 \qquad \text{for every } j.\] Similarly, \[(\boldsymbol{\delta}_2)_j = -\frac{\sqrt{p_5}}{k_u}\operatorname{sign}(\xi_j)b_j^{(2)}\] and the same formula holds for \(\tilde{\boldsymbol{\delta}}_2\), so \[(\boldsymbol{\delta}_2)_j(\tilde{\boldsymbol{\delta}}_2)_j\ge 0 \qquad \text{for every } j.\] Together with \(0<\kappa,\tilde{\kappa}\le 1\), this gives \[\label{eq:32rank-one32sign32relations} (r_\theta)_\ell(r_{\tilde{\theta}})_\ell\ge 0 \quad\text{for every } \ell, \qquad (c_\theta)_j(c_{\tilde{\theta}})_j\ge 0 \quad\text{for every } j.\tag{95}\]

Let \(\alpha^{(i)}\) be the part of \(\alpha\) corresponding to the \(i\)th observation, and write \[\alpha^{(i)}=(\mu^{(i)},\nu^{(i)})\] according to the coordinates \(U_i=(Y_i,X_{i,S_3})\) and \(V_i=X_{i,S_4}\), where \(\mu^{(i)}\in\mathbb{N}^{1+p_3}\) and \(\nu^{(i)}\in\mathbb{N}^{p_4}\). With this notation, \[H_{\alpha^{(i)}}(Z_i) = h_{\mu^{(i)}}(U_i)h_{\nu^{(i)}}(V_i).\] Since the observations are independent across \(i=1,\ldots,n\), \[\label{eq:hermite-product-form} H_\alpha(Z) = \prod_{i=1}^n H_{\alpha^{(i)}}(Z_i), \qquad \mathbb{E}_{\theta} H_\alpha(Z) = \prod_{i=1}^n \mathbb{E}_{\theta} H_{\alpha^{(i)}}(Z_i),\tag{96}\]

and the same factorization holds under \(\tilde{\theta}\). Fix a row \(i\). If \(|\mu^{(i)}|\ne|\nu^{(i)}|\), then 24 gives \[\mathbb{E}_\theta H_{\alpha^{(i)}}(Z_i) = \mathbb{E}_{\tilde{\theta}} H_{\alpha^{(i)}}(Z_i) = 0.\] If \(|\mu^{(i)}|=|\nu^{(i)}|=m_i\), set \[K_i=\frac{m_i!}{\sqrt{\mu^{(i)}!\nu^{(i)}!}}>0.\] Then 24 gives \[\mathbb{E}_\theta H_{\alpha^{(i)}}(Z_i) = K_i r_\theta^{\mu^{(i)}}c_\theta^{\nu^{(i)}}, \qquad \mathbb{E}_{\tilde{\theta}} H_{\alpha^{(i)}}(Z_i) = K_i r_{\tilde{\theta}}^{\mu^{(i)}}c_{\tilde{\theta}}^{\nu^{(i)}}.\] Therefore, in the nonzero case, \[\begin{align} &\mathbb{E}_\theta H_{\alpha^{(i)}}(Z_i)\, \mathbb{E}_{\tilde{\theta}} H_{\alpha^{(i)}}(Z_i)\\ &\qquad = K_i^2 \prod_{\ell=1}^{1+p_3} \bigl\{(r_\theta)_\ell(r_{\tilde{\theta}})_\ell\bigr\}^{\mu^{(i)}_\ell} \prod_{j=1}^{p_4} \bigl\{(c_\theta)_j(c_{\tilde{\theta}})_j\bigr\}^{\nu^{(i)}_j}. \end{align}\] The inequality 95 and the nonnegative integer exponents \(\mu^{(i)}_\ell,\nu^{(i)}_j\) imply that the last display is nonnegative. Thus, in both cases, \[\mathbb{E}_\theta H_{\alpha^{(i)}}(Z_i)\, \mathbb{E}_{\tilde{\theta}} H_{\alpha^{(i)}}(Z_i) \ge 0 \qquad \text{for every } i.\] Taking the product over \(i=1,\ldots,n\) proves 94 .

We now complete the proof of 91 . By 94 , every summand in 93 is nonnegative. Hence \[\begin{align} \left\langle L^{\le D}_\theta, L^{\le D}_{\tilde{\theta}} \right\rangle_{\mathbb{Q}} &= \sum_{\alpha\in\mathbb{N}^{N},\,|\alpha|\le D} \mathbb{E}_{\theta}H_\alpha(Z)\, \mathbb{E}_{\tilde{\theta}}H_\alpha(Z)\\ &\le \sum_{\alpha\in\mathbb{N}^{N}} \mathbb{E}_{\theta}H_\alpha(Z)\, \mathbb{E}_{\tilde{\theta}}H_\alpha(Z). \end{align}\] For the valid covariance matrices considered here, the likelihood ratios \(L_\theta\) and \(L_{\tilde{\theta}}\) belong to \(L^2(\mathbb{Q})\). Since the Hermite polynomials form a complete orthonormal basis of \(L^2(\mathbb{Q})\), the right-hand side equals \[\left\langle L_\theta,L_{\tilde{\theta}}\right\rangle_{\mathbb{Q}} = \mathbb{E}_{\theta_*}\!\left(L_\theta L_{\tilde{\theta}}\right).\] Therefore, \[\mathbb{E}_{\theta_*}\!\left( L_\theta^{\le D}L_{\tilde{\theta}}^{\le D} \right) \le \mathbb{E}_{\theta_*}\!\left( L_\theta L_{\tilde{\theta}} \right),\] which proves 91 .

Proof of 92 : Expanding in the orthonormal basis \(\{H_{\alpha}\}\), we have for any \(\theta\), \[\|L^{\le D}_\theta\|^{2}_{\mathbb{Q}} = \sum_{|\alpha|\le D} \langle L_{\theta}, H_{\alpha}\rangle_{\mathbb{Q}}^{2} = \sum_{|\alpha|\le D} \left( \mathbb{E}_{Z\sim \mathbb{P}_\theta^n}\, H_{\alpha}(Z) \right)^{2} \le (N+1)^{D}\, \max_{\alpha:\,|\alpha|\le D} \left( \mathbb{E}_{Z\sim \mathbb{P}_\theta^n}\, H_{\alpha}(Z) \right)^{2}.\]

For each \(a \in \mathbb{N}\), the function \(h_a\) admits the expansion \[h_{a}(z) = \frac{1}{\sqrt{a!}} \sum_{j=0}^{a} c_{a,j} z^{j},\] where the coefficients satisfy \(\sum_{j=0}^{a} |c_{a,j}| = T(a)\). Here, \(T(a)\) is known as the telephone number, which counts the number of involutions on \(a\) elements [61]. In particular, we have the trivial upper bound \(T(a) \le a!\). This means for any \(a \ge 1\) and \(q \in [1,\infty)\), \[\begin{align} \mathbb{E}|h_{a}(z)|^{q} &= \mathbb{E}\left| \frac{1}{\sqrt{a!}} \sum_{j=0}^{a} c_{a,j} z^{j} \right|^{q} \le \mathbb{E}\left( \sqrt{a!} \max_{0 \le j \le a} |z|^{j} \right)^{q} = (a!)^{q/2} \mathbb{E}\left( \max\{1, |z|^{a}\} \right)^{q}\\ &= (a!)^{q/2}\, \mathbb{E}\max\{1, |z|^{aq}\} \le (a!)^{q/2}\,\left(1 + \mathbb{E}|z|^{aq}\right). \end{align}\] Using the formula for Gaussian moments, and that for all \(x \ge 1\), \(\Gamma(x) \le x^{x}\) (see e.g.[62]), we have \[\label{eq:hermite-h-q-bound} \begin{align} \mathbb{E}|h_{a}(z)|^{q}&\le (a!)^{q/2} \left( 1 + \pi^{-1/2} 2^{aq/2} \Gamma\!\left( \frac{aq+1}{2} \right) \right)\\ &\le a^{aq/2} \left( 1 + 2^{aq/2} \left( \frac{aq+1}{2} \right)^{(aq+1)/2} \right)\\ &\le a^{aq/2}\bigl(1 + 2^{aq/2} (aq)^{aq}\bigr) \qquad \text{since } \frac{aq+1}{2} \le aq\\ &\le 2 a^{aq/2}\, 2^{aq/2} (aq)^{aq} \le 2 (2aq)^{2aq}. \end{align}\tag{97}\]

Now fix \(\alpha\) with \(|\alpha|\le D\), and set \(d=|\alpha|=\sum_i\alpha_i\). If \(d=0\), then \(H_\alpha\equiv1\), so the desired bound is immediate. Hence assume \(d\ge1\).

Using the product form in 96 , \(h_0\equiv 1\), Hölder’s inequality with exponents \(d/\alpha_i\) for indices \(i\) such that \(\alpha_i>0\), we obtain \[\begin{align} \left| \mathbb{E}_{Z\sim\mathbb{P}_\theta^n}H_\alpha(Z) \right| &\le \mathbb{E}_{Z\sim\mathbb{P}_\theta^n} \prod_{i:\alpha_i>0}|h_{\alpha_i}(Z_i)| \\ &\le \prod_{i:\alpha_i>0} \left( \mathbb{E}_{Z\sim\mathbb{P}_\theta^n} |h_{\alpha_i}(Z_i)|^{d/\alpha_i} \right)^{\alpha_i/d}, \end{align}\] because \(\sum_{i:\alpha_i>0}\alpha_i/d=1\).

For each \(i\) with \(\alpha_i>0\), set \(a=\alpha_i\) and \(q=d/\alpha_i\). Then \(a\ge1\), \(q\ge1\), and \(aq=d\). Applying 97 to the standard normal variable \(Z_i\) yields \[\mathbb{E}_{Z\sim\mathbb{P}_\theta^n}|h_{\alpha_i}(Z_i)|^{d/\alpha_i} \le 2(2d)^{2d}.\] Consequently, \[\left| \mathbb{E}_{Z\sim\mathbb{P}_\theta^n}H_\alpha(Z) \right| \le \prod_{i:\alpha_i>0} \left(2(2d)^{2d}\right)^{\alpha_i/d} = \left(2(2d)^{2d}\right)^{\sum_{i:\alpha_i>0}\alpha_i/d} = 2(2d)^{2d}.\] Since \(d=|\alpha|\le D\), this gives \[\left| \mathbb{E}_{Z\sim\mathbb{P}_\theta^n}H_\alpha(Z) \right| \le 2(2D)^{2D}.\]

Finally, using the bound \(N+1 = n(p+1)+1 \le 3np\), we have \[\|L^{\le D}_{\theta}\|_{Q}^{2} \le (N+1)^{D} \bigl[2(2D)^{2D}\bigr]^2 \le 4 (3np)^{D} (2D)^{4D} \le 4 (6npD)^{4D},\] which completes the proof. ◻

11 Additional Results↩︎

11.1 Low-degree polynomial framework↩︎

In this section, we briefly recall the low-degree polynomial method, which provides a widely used formal framework for studying computational limits in high-dimensional inference problems. The key idea is to analyze algorithms that can be represented (or well-approximated) by evaluating polynomials of bounded degree in the observed data; we refer readers to the survey [43] for broader background and further details.

Setup. Let \(\{\mathbb{P}^n_{\pi_1}\}_{n\ge1}\) and \(\{\mathbb{P}^n_{\pi_2}\}_{n\ge1}\) be two sequences of probability measures on an observation space \((\Omega_n,\mathcal{F}_n)\), corresponding to the null and alternative (or to two priors supported on two parameter spaces). Let \({\cal Z}\in\Omega_n\) denote the observed data. For \(D\in\mathbb{N}\), write \(\mathbb{R}[{\cal Z}]_{\le D}\) for the space of multivariate polynomials in the coordinates of \({\cal Z}\) of total degree at most \(D\). As in the main text, the degree \(D=D_n\) is allowed to grow with \(n\), and with a slight abuse of notation we view a polynomial as a sequence \(f=(f_n)_{n\ge1}\) with \(f_n\in\mathbb{R}[{\cal Z}]_{\le D_n}\).

Weak and strong separation. The low-degree method quantifies the ability of degree-bounded polynomials to distinguish \(\mathbb{P}^n_{\pi_1}\) and \(\mathbb{P}^n_{\pi_2}\) through signal-to-noise separation notions. We recall weak separation in Definition 1 and state the corresponding strong notion.

Definition 2 (Strong separation). We say that \(f\in \mathbb{R}[{\cal Z}]_{\le D}\) strongly separates \(\mathbb{P}^n_{\pi_1}\) and \(\mathbb{P}^n_{\pi_2}\) if, as \(n \to \infty\), \[\sqrt{ \max\!\left\{ {\rm Var}_{\mathbb{P}^n_{\pi_1}}(f({\cal Z})), {\rm Var}_{\mathbb{P}^n_{\pi_2}}(f({\cal Z})) \right\} } = o\!\left( \left| \mathbb{E}_{\mathbb{P}^n_{\pi_1}}[f({\cal Z})] - \mathbb{E}_{\mathbb{P}^n_{\pi_2}}[f({\cal Z})] \right| \right).\]

If \(f\) weakly separates, then thresholding \(f({\cal Z})\) at an appropriate level yields a test with nontrivial advantage (weak detection). If \(f\) strongly separates, thresholding yields a test whose sum of type-I and type-II errors tends to zero (strong detection). We do not reproduce these standard reductions here.

Low-degree likelihood ratio norm. Set \(\mathbb{Q}_1:=\mathbb{P}^n_{\pi_1}\) and \(\mathbb{Q}_2:=\mathbb{P}^n_{\pi_2}\), and assume \(\mathbb{Q}_1\) is absolutely continuous with respect to \(\mathbb{Q}_2\). Define the likelihood ratio \[L \;=\; \frac{\,{\rm d}\,\mathbb{Q}_1}{\,{\rm d}\,\mathbb{Q}_2}.\] Endow \(L^2(\mathbb{Q}_2)\) with inner product \(\langle f,g\rangle := \mathbb{E}_{\mathbb{Q}_2}[f({\cal Z})g({\cal Z})]\). Let \(L^{\le D}\) denote the \(L^2(\mathbb{Q}_2)\)-orthogonal projection of \(L\) onto the polynomial subspace \(\mathbb{R}[{\cal Z}]_{\le D}\). The associated low-degree quantity is \[\label{eq:LD-def-app} \mathrm{LD}(D) \;:=\; \|L^{\le D}\|_{L^2(\mathbb{Q}_2)}^2 \;=\; \mathbb{E}_{\mathbb{Q}_2}\!\left[(L^{\le D}({\cal Z}))^2\right].\tag{98}\] By construction, \(\mathrm{LD}(D)\ge 1\) since \(1\in\mathbb{R}[{\cal Z}]_{\le D}\) and \(\mathbb{E}_{\mathbb{Q}_2}[L]=1\).

Interpreting \(\mathrm{LD}(D)\): weak vs strong low-degree indistinguishability. The following implication formalizes the meaning of the two regimes \(\mathrm{LD}(D)\to 1\) versus \(\mathrm{LD}(D)=O(1)\); we refer to [41], [63] and the aforementioned survey for proofs and refinements.

Proposition 4 (Low-degree obstruction to separation). Let \(\mathbb{Q}_1=\mathbb{P}^n_{\pi_1}\) and \(\mathbb{Q}_2=\mathbb{P}^n_{\pi_2}\) with \(\mathbb{Q}_1\ll \mathbb{Q}_2\). Fix \(D=D_n\).

  1. If \(\mathrm{LD}(D)=1+o(1)\) as \(n\to\infty\), then no degree-\(D\) polynomial weakly separates \(\mathbb{Q}_1\) and \(\mathbb{Q}_2\).

  2. If \(\mathrm{LD}(D)=O(1)\) as \(n\to\infty\), then no degree-\(D\) polynomial strongly separates \(\mathbb{Q}_1\) and \(\mathbb{Q}_2\).

Remark. Statement (i) asserts that, when \(\mathrm{LD}(D)=1+o(1)\), every degree-\(D\) polynomial test is asymptotically no better than the trivial constant statistic in the sense of weak separation. Statement (ii) is stronger in that it rules out vanishing total error: boundedness of \(\mathrm{LD}(D)\) permits at most a constant-factor signal-to-noise ratio and therefore precludes strong separation. For additional perspectives and related connections (e.g.to sum-of-squares lower bounds and pseudo-calibration), we refer readers to [11] and the survey references above.

11.2 Restricted eigenvalue condition for Gaussian designs↩︎

We briefly justify the restricted eigenvalue condition used after 15 under Gaussian random designs, and explain how it implies the estimator bounds in Condition 1. Recall that \[\kappa(X,k,\alpha_0) = \min_{\substack{J_0\subseteq\{1,\ldots,p\}\\ |J_0|\le k}} \; \min_{\substack{\delta\neq 0\\ \|\delta_{J_0^c}\|_1\le \alpha_0\|\delta_{J_0}\|_1}} \frac{\|X\delta\|_2}{\sqrt n\,\|\delta_{J_0}\|_2}.\] Direct verification of this condition is computationally difficult [37], [38]. For Gaussian random designs, however, it follows from standard uniform concentration results. In particular, [64] show that if \(X\in\mathbb{R}^{n\times p}\) has i.i.d.\(N(0,\Sigma)\) rows and \(\rho(\Sigma)=\sqrt{\max_{1\le j\le p}\Sigma_{jj}}\), then with probability at least \(1-c'\exp(-cn)\), \[\frac{\|Xv\|_2}{\sqrt n} \ge \frac{1}{4}\|\Sigma^{1/2}v\|_2 - 9\rho(\Sigma)\sqrt{\frac{\log p}{n}}\|v\|_1, \qquad \text{for all } v\in\mathbb{R}^p .\]

Suppose further that \[M_1^{-1}\le \lambda_{\min}(\Sigma) \le \lambda_{\max}(\Sigma)\le M_1 .\] For any vector \(\delta\) in the cone \(\|\delta_{J_0^c}\|_1\le \alpha_0\|\delta_{J_0}\|_1\) with \(|J_0|\le k_u\), we have \[\|\delta\|_1 \le (1+\alpha_0)\|\delta_{J_0}\|_1 \le (1+\alpha_0)\sqrt{k_u}\|\delta_{J_0}\|_2 .\] Therefore, on the event above, \[\frac{\|X\delta\|_2}{\sqrt n\,\|\delta_{J_0}\|_2} \ge \frac{1}{4\sqrt{M_1}} - 9\sqrt{M_1}(1+\alpha_0) \sqrt{\frac{k_u\log p}{n}} .\] Consequently, for each fixed \(\alpha_0>0\), there exists a constant \(c_{\rm RE}>0\), depending only on \(M_1\) and \(\alpha_0\), such that if \[\frac{k_u\log p}{n}\le c_{\rm RE},\] then \[\kappa(X,k_u,\alpha_0) \ge \frac{1}{8\sqrt{M_1}}\] with probability at least \(1-c'\exp(-cn)\).

We next explain how this lower bound yields the constants \(c_\beta\) and \(C_\beta\) in Condition 1. Consider the Lasso estimator \[\hat{\beta} \in \mathop{\mathrm{arg\,min}}_{b\in\mathbb{R}^p} \left\{ \frac{1}{2n}\|Y-Xb\|_2^2+\lambda\|b\|_1 \right\}, \qquad \lambda=A\sigma\sqrt{\frac{\log p}{n}},\] where \(A>0\) is sufficiently large. On the event \[\frac{1}{n}\left\|X^\top\varepsilon\right\|_\infty \le \frac{\lambda}{2},\] the usual basic inequality implies that \(\Delta=\hat{\beta}-\beta\) satisfies the cone condition \[\|\Delta_{S^c}\|_1\le 3\|\Delta_S\|_1, \qquad S={\rm supp}(\beta),\quad |S|\le k_u .\] Thus, if \(\kappa(X,k_u,3)>0\), the restricted eigenvalue condition in 15 gives \[\frac{\|X\Delta\|_2}{\sqrt n} \ge \kappa(X,k_u,3)\|\Delta_S\|_2 .\] Combining this with the standard Lasso oracle inequality yields \[\|\hat{\beta}-\beta\|_1 \le \frac{C A}{\kappa^2(X,k_u,3)} \sigma k_u\sqrt{\frac{\log p}{n}},\] and \[\|\hat{\beta}-\beta\|_2 \le \frac{C A}{\kappa^2(X,k_u,3)} \sigma\sqrt{\frac{k_u\log p}{n}},\] where \(C>0\) is a universal constant. Hence, if \[\kappa(X,k_u,3)\ge \kappa_0>0\] with probability approaching one, then Condition 1 holds with admissible constants \[c_\beta=\frac{C A}{\kappa_0^2}, \qquad C_\beta=\frac{C A}{\kappa_0^2}.\] In particular, under Gaussian designs with bounded population eigenvalues, the previous concentration argument gives such a lower bound with \(\kappa_0=1/(8\sqrt{M_1})\), provided \(k_u\log p/n\le c_{\rm RE}\). Thus, the restricted eigenvalue argument gives fixed constants \(c_\beta\) and \(C_\beta\) depending only on \(A\), \(M_1\), and the cone constant.

[cdt: linear estimator,cdt: linear variance] are satisfied by the scaled Lasso under Gaussian designs provided \[k_u \log p / n \le c_{\mathrm{RE}}\] where \(c_{\mathrm{RE}}>0\) depends only on \(M_1\) and the cone constant. This is a small-constant version of the sparsity scaling \(k_u \lesssim n / \log p\).

The scaled Lasso satisfies analogous bounds with \(\lambda\asymp \hat{\sigma}\sqrt{\log p/n}\), and additionally yields a consistent estimator of \(\sigma^2\). Hence, under the same restricted eigenvalue scaling, the scaled Lasso verifies both Conditions 1 and 2.

11.3 Performance of some test statistics for \({\rm SCCA}(n,s,p_1,p_2,\lambda)\)↩︎

In this section we justify the performance of several test statistics for \({\rm SCCA}(n,s,p_1,p_2,\lambda)\) defined in 37 . We focus on the regime \[s \lesssim p_1 \lesssim n \ll \sqrt{p_2}, \qquad \text{with } s,p_1,p_2,n\to\infty,\] which differs from the classical sparse submatrix detection scaling where typically \(p_1\asymp p_2\) (e.g.[26], [45]). As discussed in 4.3, it suffices to analyze statistics based on the sample cross-covariance matrix \[\label{eq:Rhatscca} \widehat R \;:=\; \frac{1}{2n}\sum_{i=1}^{2n} U_{1}^{(i)}\bigl(U_{2}^{(i)}\bigr)^\top \;\in\; \mathbb{R}^{p_1\times p_2},\tag{99}\] where under \(H_0\) the pairs \(\{(U_1^{(i)},U_2^{(i)})\}_{i=1}^{2n}\) are i.i.d.from \(N(0,I_{p_1})\otimes N(0,I_{p_2})\).

Under \(H_1\), conditional on \((\boldsymbol{\delta}_1,\boldsymbol{\delta}_2)\), we have \[\label{eq:meanRhat} \mathbb{E}[\widehat R \mid \boldsymbol{\delta}_1,\boldsymbol{\delta}_2] \;=\; \lambda \boldsymbol{\delta}_1\boldsymbol{\delta}_2^\top.\tag{100}\] Since \(\boldsymbol{\delta}_1,\boldsymbol{\delta}_2\) are \(s\)-sparse and flat-on-support, each nonzero entry of \(\lambda \boldsymbol{\delta}_1\boldsymbol{\delta}_2^\top\) equals \(\lambda/s\), and the signal is supported on an unknown \(s\times s\) submatrix.

Scan test. The scan statistic is given by \[\label{eq:Tscan-scca} T_{\text{scan}} \;=\; \max_{\substack{S_1 \subseteq [p_1],\, |S_1| = s\\ S_2 \subseteq [p_2],\, |S_2| = s}} \frac{1}{s^2}\sum_{i \in S_1} \sum_{j \in S_2} \widehat R_{ij}.\tag{101}\] The test rejects \(H_0\) when \(T_{\text{scan}}\) exceeds a threshold \(\tau_{\text{scan}}\).

Fix \((S_1,S_2)\) with \(|S_1|=|S_2|=s\) and define the averaged submatrix sum \[Z(S_1,S_2) \;:=\; \frac{1}{s^2}\sum_{i \in S_1}\sum_{j \in S_2}\widehat R_{ij} \;=\; \frac{1}{2n}\sum_{\ell=1}^{2n}\underbrace{\left(\frac{1}{s}\sum_{i\in S_1}U_{1,i}^{(\ell)}\right) \left(\frac{1}{s}\sum_{j\in S_2}U_{2,j}^{(\ell)}\right)}_{=: \;A_\ell(S_1)\,B_\ell(S_2)}.\] Under \(H_0\), conditional on \((S_1,S_2)\), the random variables \[A_\ell(S_1)=\frac{1}{s}\sum_{i\in S_1}U_{1,i}^{(\ell)},\qquad B_\ell(S_2)=\frac{1}{s}\sum_{j\in S_2}U_{2,j}^{(\ell)}\] are independent Gaussians with \[A_\ell(S_1)\sim N\!\left(0,\frac{1}{s}\right), \qquad B_\ell(S_2)\sim N\!\left(0,\frac{1}{s}\right),\] hence \(A_\ell(S_1)B_\ell(S_2)\) is sub-exponential with scale \(\asymp 1/s\). Applying Bernstein’s inequality yields: for all \(t\ge 0\), \[\label{eq:Z-tail} \mathbb{P}_{H_0}\!\left(|Z(S_1,S_2)|\ge t\right) \le 2\exp\!\left(-c\,n\,\min\{s^2 t^2,\, s t\}\right).\tag{102}\] In particular, for \(t\le 1/s\), \[\label{eq:Z-subg} \mathbb{P}_{H_0}\!\left(|Z(S_1,S_2)|\ge t\right) \le 2\exp(-c n s^2 t^2).\tag{103}\]

Now take a union bound over all \((S_1,S_2)\): the number of candidates is \[N_{\text{scan}} = \binom{p_1}{s}\binom{p_2}{s}, \qquad \log N_{\text{scan}} \le s\log\!\left(\frac{e p_1}{s}\right)+s\log\!\left(\frac{e p_2}{s}\right).\] Setting \[\label{eq:tauscan} \tau_{\text{scan}} = C\sqrt{\frac{\log N_{\text{scan}}}{n s^2}} \;\asymp\; \sqrt{\frac{\log N_{\text{scan}}}{n s^2}},\tag{104}\] and assuming \(\tau_{\text{scan}}\le 1/s\) (which holds in the regime of interest), 103 implies \[\label{eq:type1-scan} \mathbb{P}_{H_0}\!\left(T_{\text{scan}}>\tau_{\text{scan}}\right) \le 2N_{\text{scan}}\exp(-c n s^2 \tau_{\text{scan}}^2) \le 2\exp\!\left(\log N_{\text{scan}}-c C^2\log N_{\text{scan}}\right) \to 0\tag{105}\] for \(C\) large enough. Therefore, we can choose the rejection threshold \(\tau_{\text{scan}}\) as in 104 to control the type-I error.

Under \(H_1\), let \(S_1^\star=\mathrm{supp}(\boldsymbol{\delta}_1)\) and \(S_2^\star=\mathrm{supp}(\boldsymbol{\delta}_2)\). By 100 , the corresponding submatrix average satisfies \[\mathbb{E}_{H_1}\!\left[ Z(S_1^\star,S_2^\star)\mid \boldsymbol{\delta}_1,\boldsymbol{\delta}_2\right] = \frac{1}{s^2}\sum_{i\in S_1^\star}\sum_{j\in S_2^\star}\frac{\lambda}{s} = \frac{\lambda}{s}.\] Moreover, the concentration bound 103 continues to hold under \(H_1\) up to absolute-constant changes in \(c\) (since the model remains Gaussian with bounded covariance operator norm). Thus, if \[\label{eq:scan-power-cond} \frac{\lambda}{s} \;\ge\; 2\tau_{\text{scan}} \;\asymp\; \sqrt{\frac{\log N_{\text{scan}}}{n s^2}},\tag{106}\] then with probability \(1-o(1)\) we have \(Z(S_1^\star,S_2^\star)\ge \tau_{\text{scan}}\), hence \(T_{\text{scan}}\ge \tau_{\text{scan}}\) and the scan test rejects. Equivalently, the scan test is powerful whenever \[\label{eq:lambda-scan-rate} \lambda \;\gtrsim\; \sqrt{\frac{\log N_{\text{scan}}}{n}} \;\asymp\; \sqrt{\frac{s\log(p_1/s)+s\log(p_2/s)}{n}} \;\asymp\; \sqrt{\frac{s\log p_2}{n}},\tag{107}\] where the last simplification uses \(p_2\gg p_1\gtrsim s\) so that \(\log(p_2/s)\) dominates. This matches the information-theoretically optimal benchmark, but computing \(T_{\text{scan}}\) is combinatorial.

Entrywise maximum statistic. Consider \[\label{eq:Tinf-scca} T_{\infty}:=\max_{i\in[p_1],\,j\in[p_2]}\widehat R_{ij}.\tag{108}\] Under \(H_0\), for each \((i,j)\), we have \[\widehat R_{ij}=\frac{1}{2n}\sum_{\ell=1}^{2n} W_\ell^{(ij)}, \qquad W_\ell^{(ij)}:=U_{1,i}^{(\ell)}U_{2,j}^{(\ell)}.\] where \(U_{1,i}^{(\ell)}\sim N(0,1)\) and \(U_{2,j}^{(\ell)}\sim N(0,1)\) are independent, so \(W_\ell^{(ij)}\) is a product of independent standard normals and is sub-exponential. Consequently, by Bernstein’s inequality for averages of i.i.d.sub-exponential variables, there exist absolute constants \(c,C>0\) such that for all \(t\ge 0\), \[\label{eq:Rij-tail-scca} \mathbb{P}_{0}\!\left(\left|\widehat R_{ij}\right|\ge t\right) \le 2\exp\!\left(-c\,n\,\min\{t^2,t\}\right).\tag{109}\] In particular, for \(0\le t\le 1\), \[\label{eq:Rij-subg-small-scca} \mathbb{P}_{0}\!\left(\left|\widehat R_{ij}\right|\ge t\right) \le 2\exp(-c n t^2),\tag{110}\] which matches sub-Gaussian behavior at the relevant scale \(t\asymp \sqrt{(\log p_2)/n}\).

Now take a union bound over all \(p_1p_2\) entries, \[\label{eq:Tinf-null} T_{\infty} = O_{\mathbb{P}_{H_0}}\!\left(\sqrt{\frac{\log(p_1p_2)}{n}}\right) = O_{\mathbb{P}_{H_0}}\!\left(\sqrt{\frac{\log p_2}{n}}\right),\tag{111}\] using \(p_2\gg p_1\). Under \(H_1\), on the signal support \((S_1^\star\times S_2^\star)\), \[\mathbb{E}_{H_1}[\widehat R_{ij}\mid \boldsymbol{\delta}_1,\boldsymbol{\delta}_2]=\lambda/s, \qquad (i,j)\in S_1^\star\times S_2^\star.\] Therefore, if \[\label{eq:lambda-entrywise} \frac{\lambda}{s} \;\gg\; \sqrt{\frac{\log(p_1p_2)}{n}},\tag{112}\] then \(T_\infty\) exceeds any null-calibrated threshold with probability \(1-o(1)\), and the test is powerful. Equivalently, the entrywise maximum test requires \[\label{eq:lambda-rate-entrywise} \lambda \;\gtrsim\; s\sqrt{\frac{\log p_2}{n}},\tag{113}\] which is strictly weaker (i.e., needs larger \(\lambda\)) than 107 for \(s\to\infty\).

Max-column statistic. Define the max-column statistic \[\label{eq:Tmaxcol-scca} T_{\text{max-col}} := \max_{j\in[p_2]}\frac{1}{s}\sum_{i=1}^{p_1}\widehat R_{ij}.\tag{114}\]

Under \(H_0\), for fixed \(j\), write \[C_j:=\frac{1}{s}\sum_{i=1}^{p_1}\widehat R_{ij} = \frac{1}{2n}\sum_{\ell=1}^{2n} \left(\frac{1}{s}\sum_{i=1}^{p_1}U_{1,i}^{(\ell)}\right)U_{2,j}^{(\ell)}.\] Under \(H_0\), \(\sum_{i=1}^{p_1}U_{1,i}^{(\ell)}\sim N(0,p_1)\) and is independent of \(U_{2,j}^{(\ell)}\sim N(0,1)\), so the summand is a product of independent Gaussians with standard deviations \(\sqrt{p_1}/s\) and \(1\). Hence \(C_j\) is an average of i.i.d.sub-exponential variables with scale \(\asymp \sqrt{p_1}/s\). Bernstein’s inequality yields, for all \(t\ge 0\), \[\label{eq:Cj-tail} \mathbb{P}_{H_0}\!\left(|C_j|\ge t\right) \le 2\exp\!\left(-c\,n\,\min\left\{\frac{s^2 t^2}{p_1},\,\frac{s t}{\sqrt{p_1}}\right\}\right).\tag{115}\] In particular, for \(t\le \sqrt{p_1}/s\), \[\label{eq:Cj-subg} \mathbb{P}_{H_0}\!\left(|C_j|\ge t\right) \le 2\exp\!\left(-c n \frac{s^2 t^2}{p_1}\right).\tag{116}\] Taking a union bound over \(j\in[p_2]\) and choosing \[\label{eq:taumaxcol} \tau_{\text{max-col}} = C\sqrt{\frac{p_1\log p_2}{n s^2}},\tag{117}\] we obtain \(\mathbb{P}_{H_0}(T_{\text{max-col}}>\tau_{\text{max-col}})\to 0\) for \(C\) large enough.

Under \(H_1\), let \(S_1^\star=\mathrm{supp}(\boldsymbol{\delta}_1)\) and \(S_2^\star=\mathrm{supp}(\boldsymbol{\delta}_2)\). For any \(j\in S_2^\star\), using 100 , \[\mathbb{E}_{H_1}[C_j\mid \boldsymbol{\delta}_1,\boldsymbol{\delta}_2] = \frac{1}{s}\sum_{i\in S_1^\star}\frac{\lambda}{s} = \frac{\lambda}{s}.\] Thus, if \(\lambda/s \ge 2\tau_{\text{max-col}}\), then with probability \(1-o(1)\) we have \(T_{\text{max-col}}\ge \tau_{\text{max-col}}\) and the max-column test is powerful. Equivalently, the max-column statistic requires \[\label{eq:lambda-rate-maxcol} \lambda \;\gtrsim\; \sqrt{\frac{p_1\log p_2}{n}}.\tag{118}\]

Max-row statistic↩︎

Define the max-row statistic \[\label{eq:Tmaxrow-scca} T_{\text{max-row}} := \max_{i\in[p_1]}\frac{1}{s}\sum_{j=1}^{p_2}\widehat R_{ij}.\tag{119}\] This statistic is polynomial-time computable.

Under \(H_0\), for a fixed \(i\), write \[R_i := \frac{1}{s}\sum_{j=1}^{p_2}\widehat R_{ij} = \frac{1}{2n}\sum_{\ell=1}^{2n} U_{1,i}^{(\ell)}\left(\frac{1}{s}\sum_{j=1}^{p_2}U_{2,j}^{(\ell)}\right).\] Under \(H_0\), \(U_{1,i}^{(\ell)}\sim N(0,1)\) and \(\sum_{j=1}^{p_2}U_{2,j}^{(\ell)}\sim N(0,p_2)\) are independent, hence the summand is a product of independent Gaussians with standard deviations \(1\) and \(\sqrt{p_2}/s\). Therefore \(R_i\) is an average of i.i.d.sub-exponential variables with scale \(\asymp \sqrt{p_2}/s\). By Bernstein’s inequality, there exist absolute constants \(c,C>0\) such that for all \(t\ge 0\), \[\label{eq:Ri-tail} \mathbb{P}_0\!\left(|R_i|\ge t\right) \le 2\exp\!\left(-c\,n\,\min\left\{\frac{s^2 t^2}{p_2},\,\frac{s t}{\sqrt{p_2}}\right\}\right).\tag{120}\] In particular, for \(t\le \sqrt{p_2}/s\), \[\label{eq:Ri-subg} \mathbb{P}_0\!\left(|R_i|\ge t\right) \le 2\exp\!\left(-c n \frac{s^2 t^2}{p_2}\right).\tag{121}\] Taking a union bound over \(i\in[p_1]\) and choosing the threshold \[\label{eq:taumaxrow} \tau_{\text{max-row}} = C\sqrt{\frac{p_2\log p_1}{n s^2}},\tag{122}\] we obtain \(\mathbb{P}_0(T_{\text{max-row}}>\tau_{\text{max-row}})\to 0\) for \(C\) large enough.

Under \(H_1\), let \(S_1^\star=\mathrm{supp}(\boldsymbol{\delta}_1)\) and \(S_2^\star=\mathrm{supp}(\boldsymbol{\delta}_2)\). For any \(i\in S_1^\star\), using 100 , \[\mathbb{E}_1[R_i\mid \boldsymbol{\delta}_1,\boldsymbol{\delta}_2] = \frac{1}{s}\sum_{j\in S_2^\star}\frac{\lambda}{s} = \frac{\lambda}{s}.\] Thus, if \(\lambda/s \ge 2\tau_{\text{max-row}}\), then with probability \(1-o(1)\) \(T_{\text{max-row}}\ge \tau_{\text{max-row}}\) and the max-row test is powerful. Equivalently, the max-row statistic requires \[\label{eq:lambda-rate-maxrow} \lambda \;\gtrsim\; \sqrt{\frac{p_2\log p_1}{n}}.\tag{123}\]

Global sum statistic. Finally, consider the global sum \[\label{eq:Tsum-scca} T_{\text{sum}}:= \frac{1}{p_1p_2}\sum_{i=1}^{p_1}\sum_{j=1}^{p_2}\widehat R_{ij}.\tag{124}\] Under \(H_0\), \[T_{\text{sum}} = \frac{1}{2n}\sum_{\ell=1}^{2n} \left(\frac{1}{p_1}\sum_{i=1}^{p_1}U_{1,i}^{(\ell)}\right) \left(\frac{1}{p_2}\sum_{j=1}^{p_2}U_{2,j}^{(\ell)}\right),\] where the two parentheses are independent Gaussians with variances \(1/p_1\) and \(1/p_2\). Hence \(T_{\text{sum}}\) concentrates at scale \(\sqrt{1/(n p_1 p_2)}\) (sub-exponential tails analogously).

Under \(H_1\), the mean shift equals \[\mathbb{E}_{H_1}[T_{\text{sum}}\mid \boldsymbol{\delta}_1,\boldsymbol{\delta}_2] = \frac{1}{p_1p_2}\sum_{i\in S_1^\star}\sum_{j\in S_2^\star}\frac{\lambda}{s} = \frac{\lambda s}{p_1p_2},\] which is heavily diluted for sparse alternatives. Balancing the signal mean against the null standard deviation suggests that power requires \[\label{eq:lambda-rate-sum} \frac{\lambda s}{p_1p_2} \;\gtrsim\; \sqrt{\frac{1}{n p_1 p_2}} \qquad\Longleftrightarrow\qquad \lambda \;\gtrsim\; \sqrt{\frac{p_1p_2}{n}}\cdot \frac{1}{s},\tag{125}\] which is far worse than 107 in the sparse regime \(s\ll \sqrt{p_1p_2}\).

Summary of thresholds. In the regime \(s\lesssim p_1\lesssim n\ll p_2\), the above calculations yield the detection boundary for all test statistics: \[\begin{align} \lambda_{\text{scan}} \asymp \sqrt{\frac{s\log p_2}{n}}, \quad \lambda_{\infty} \asymp &s\sqrt{\frac{\log p_2}{n}}, \quad \lambda_{\text{max-col}} \asymp \sqrt{\frac{p_1\log p_2}{n}}\\ \lambda_{\text{max-row}}\asymp\sqrt{\frac{p_2\log p_1}{n}} ,&\quad \lambda_{\text{sum}}\asymp \sqrt{\frac{p_1 p_2}{ns^2}}. \end{align}\]

12 Loading-profile examples and consequences↩︎

This appendix works out several concrete loading profiles to illustrate the scope of the profile-based theory developed in the main text. The calculations below serve two purposes. First, they translate the abstract quantities \(H(\cdot~;~\xi)\), \(\nu_1\), and \(\nu_2\) into explicit separation rates. Second, they identify loading profiles that are not covered by existing regular-loading or exact polynomial-decay theories.

The regular-loading example in 12.1 recovers the phase diagram in 1 and highlights the intermediate moderately sparse range \(k_u\ll K\ll k_u^2\). 12.2 gives dense nonregular profiles for which the adaptive separation distance can still be determined. 12.3 constructs a multiscale profile for which the low-degree rate exceeds the statistical rate by a polynomial factor, which illustrates a statistical–computational gap in sparse signed-spiked models. 12.4 is related to the case where the loading vector is a random test point; it treats random loadings by conditioning on the realized vector and evaluating the resulting profile quantities.

Throughout this appendix, we assume [cdt: linear estimator,cdt: linear variance,cdt: sparsity assumption] hold.

12.1 Regular loading vectors↩︎

We derive the rates displayed in 1 for loading vectors satisfying the regular loading condition 3 . Throughout this subsection, let \[K=\|\xi\|_0, \qquad a=\|\xi\|_\infty .\] Since the coordinates of \(\xi\) are ordered in decreasing absolute value, the regular loading condition implies that, for some constant \(\bar c>0\), \[a/\bar c\le |\xi_j|\le a,\qquad 1\le j\le K, \qquad \xi_j=0,\qquad j>K.\] All constants below may depend on \(\bar c\), but not on \(n,p,k_u,K\), or \(a\).

We first state several elementary consequences of the regular loading condition.

Lemma 25. Suppose that \(\xi\) satisfies the regular loading condition 3 . Then, for all \(t\ge 1\), \[H(t~;~\xi) = \left(\sum_{j\le \lceil t\rceil}\xi_j^2\right)^{1/2} \asymp a\sqrt{t\wedge K},\] and thus \(\nu_2 = H(k_u~;~\xi) \asymp a\sqrt{k_u\wedge K}\). Moreover, the quantity \(\nu_1\) satisfies \[\nu_1 \asymp \begin{cases} a\sqrt K, & K\lesssim k_u^2,\\[4pt] a k_u\left\{1+\sqrt{\log\left(eK/k_u^2\right)}\right\}, & K\gtrsim k_u^2 . \end{cases}\]

Proof. The first two displays follow immediately from the regular loading condition. We prove the assertion for \(\nu_1\).

Let \(x_j=|\xi_j|\), and define \[F(z) = \frac{\sum_{j=1}^{K} x_j \exp(-z/x_j^2)}{\left(\sum_{j=1}^{K} x_j^2 \exp(-z/x_j^2)\right)^{1/2}}, \qquad z\in\mathbb{R}.\] By the definition of \(\lambda\) in 17 , \(\lambda=\sqrt{\zeta_+}\), where \(\zeta\) is the unique solution to \(F(\zeta)=k_u/2\). By 3 , we have \(F(0)\asymp \sqrt K\).

For \(z\ge0\), by Cauchy–Schwarz inequality, we have \[\label{eq:32regular32example32CS} F(z) \le \left(\sum_{j=1}^{K}\exp(-z/x_j^2)\right)^{1/2} \le \sqrt K \exp(-z/(2a^2)).\tag{126}\]

Setting 1: \(K\lesssim k_u^2\). If \(F(0)\leq k_u/2\), then \(\zeta\le 0\) and \(\lambda=0\).

If \(F(0)>k_u/2\), we can prove that \(\zeta\) is bounded above by a constant multiple of \(a^2\). Indeed, by 126 and \(x_j\leq a\), we have \(F(ta^2)\leq \sqrt{K}\exp(-t/2)\) for any \(t>0\). Since \(k_u/(2\sqrt{K})\asymp \frac{k_u}{2F(0)}\lesssim 1\) we can choose a constant \(t>0\) such that \(\exp(-t/2)\leq k_u/(2\sqrt{K})\).

In both cases, we have \(\lambda\lesssim a\), and thus \(\exp(-\lambda^2/\xi_j^2)\) is bounded by some constant for all \(j\leq K\). Therefore, \[\left(\sum_{j=1}^K \xi_j^2 \exp(-\lambda^2/\xi_j^2)\right)^{1/2} \asymp a\sqrt K .\] Moreover, we have \(\lambda k_u=0\) in the first case and \(\lambda k_u\lesssim a k_u\lesssim a\sqrt K\) in the second case. Therefore \(\nu_1\asymp a\sqrt K\).

Setting 2: \(K\gtrsim k_u^2\). If \(K\asymp k_u^2\), the preceding argument gives \(\lambda\lesssim a\), and the second term in \(\nu_1\) is of order \(ak_u\); hence \(\nu_1\asymp ak_u\), which agrees with the displayed bound because \(\log(eK/k_u^2)\asymp 1\). It remains to consider the case in which \(K/k_u^2\) is larger than a sufficiently large constant. In this case the solution is positive and we show that \[\lambda^2 \asymp a^2\log\left(eK/k_u^2\right).\] Since \(x_j\ge a/\bar c\) for \(1\le j\le K\), \[\label{eq:32regular32example32F32lower} F(z) \ge \frac{(a/\bar c)K\exp(-\bar c^2 z/a^2)}{a\sqrt K} = \bar c^{-1}\sqrt K \exp(-\bar c^2 z/a^2).\tag{127}\]

Combining this with 126 , we see that the solution of \(F(z)=k_u/2\) satisfies \[\zeta \asymp a^2\log\left(eK/k_u^2\right),\] and therefore \[\lambda \asymp a\sqrt{\log\left(eK/k_u^2\right)} .\]

It remains to evaluate the second term in \(\nu_1\). Let \[W_\lambda = \sum_{j=1}^K \exp(-\lambda^2/x_j^2).\] Since \(a/\bar{c}\leq x_j\leq a\), we can follow the proofs for [eq: regular example CS,eq: regular example F lower] to see that \[F(\lambda^2) \asymp W_\lambda^{1/2}.\] Since \(F(\lambda^2)=k_u/2\), we have \(W_\lambda\asymp k_u^2\). Consequently, \[\left(\sum_{j=1}^K \xi_j^2\exp(-\lambda^2/\xi_j^2)\right)^{1/2} \asymp a W_\lambda^{1/2} \asymp a k_u .\] Combining this with the bound for \(\lambda\) gives \[\nu_1 = \lambda k_u+ \left(\sum_{j=1}^K \xi_j^2\exp(-\lambda^2/\xi_j^2)\right)^{1/2} \asymp a k_u\left\{1+\sqrt{\log\left(eK/k_u^2\right)}\right\}.\] This proves the lemma. ◻

We now derive the rates in the ultra-sparse and moderately sparse regimes.

Proposition 5. Suppose that \(\xi\) satisfies the regular loading condition 3 . Let \(K=\|\xi\|_0\) and \(a=\|\xi\|_\infty\).

  1. Suppose \(k_u\lesssim \sqrt n/\log p\). If \(K\lesssim k_u^2\), then \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp \frac{a\sqrt K}{\sqrt n}.\] If \(K\gtrsim k_u^2\), then \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp_{\log} a k_u\sqrt{\frac{\log p}{n}}.\] Moreover, if \(K/k_u^2\ge p^c\) for some constant \(c>0\), then the logarithmic equivalence can be strengthened to \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp a k_u\sqrt{\frac{\log p}{n}} .\]

  2. Suppose \(k_u\gg \sqrt n/\log p\). If \(K\lesssim k_u\), then \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp a\sqrt K\,\frac{k_u\log p}{n}.\] If \(K\gtrsim k_u^2\), then \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp_{\log} a k_u\sqrt{\frac{\log p}{n}}.\] Moreover, if \(K/k_u^2\ge p^c\) for some constant \(c>0\), then \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp a k_u\sqrt{\frac{\log p}{n}} .\] In the intermediate regime \(k_u\ll K\ll k_u^2\), the general bounds in [thm: hypothesis,prop: upper bound equivalent] give \[a\left\{ \frac{\sqrt K}{\sqrt n} + \frac{k_u^{3/2}\log p}{n} \right\} \lesssim \tau_{\mathrm{adap}}(k_u,k~;~\xi) \lesssim a\sqrt{K\wedge \frac{n}{\log p}}\, \frac{k_u\log p}{n}.\] If, in addition, \(K\le n/\log p\), and both \(k_u\log p/\sqrt n\) and \(K/k_u\) diverge by polynomial factors, then these available upper and lower bounds do not match up to logarithmic factors.

All the statements above hold uniformly over \(k\le k_u\).

Proof. We repeatedly use [thm: hypothesis,prop: upper bound equivalent], together with 25.

1. Ultra-sparse regime \(k_u\lesssim \sqrt n/\log p\).

By 1, \[\tau_{\mathrm{adap}}(k_u,k~;~\xi)\lesssim \frac{1}{\sqrt n}H(k_u^2\log p~;~\xi).\] If \(K\lesssim k_u^2\), then \(H(k_u^2\log p~;~\xi)\asymp a\sqrt K\), and the lower bound \(\nu_1/\sqrt n\) satisfies \(\nu_1/\sqrt n\asymp a\sqrt K/\sqrt n\) by 25. Hence \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp \frac{a\sqrt K}{\sqrt n}.\]

If \(K\gtrsim k_u^2\), 25 gives \[\nu_1 \gtrsim a k_u .\] Therefore, 1 gives \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp_{\log} a k_u\sqrt{\frac{\log p}{n}}.\] If \(K/k_u^2\ge p^c\) for \(c>0\), then \(H(k_u^2\log p~;~\xi)\asymp a\sqrt{k_u^2\log p}\). Furthermore, 25 gives \[\nu_1 \asymp a k_u\sqrt{\log p},\] which implies that the lower bound \(\nu_1/\sqrt{n}\) matches the upper bound \(a k_u\sqrt{\frac{\log p}{n}}\) up to constants.

2. Moderately sparse regime \(k_u\gg \sqrt n/\log p\).

In this case, 1 yields the upper bound \[\label{eq:regular-moderate-sparse-upper} \tau_{\mathrm{adap}}(k_u,k~;~\xi) \lesssim \frac{k_u\log p}{n}H(n/\log p~;~\xi) \asymp a\sqrt{K\wedge \frac{n}{\log p}}\, \frac{k_u\log p}{n}.\tag{128}\] The lower bound in 1 is \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \gtrsim \frac{\nu_1}{\sqrt n} \vee \nu_2\frac{k_u\log p}{n}.\]

If \(K\lesssim k_u\), then \[H(n/\log p~;~\xi)\asymp H(k_u~;~\xi)\asymp a\sqrt K.\] Recall that \(\nu_2=H(k_u;\xi)\). We see that both the upper bound and the lower bound \(\nu_2 k_u\log p/n\) are at the scale of \[a\sqrt K\,\frac{k_u\log p}{n}.\] Thus \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp a\sqrt K\,\frac{k_u\log p}{n}.\]

If \(K\gtrsim k_u^2\), then the upper bound in 128 satisfies \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \lesssim a\sqrt{\frac{n}{\log p}}\frac{k_u\log p}{n} = a k_u\sqrt{\frac{\log p}{n}}.\] By 25, the lower bound involving \(\nu_1\) gives at least \[\frac{\nu_1}{\sqrt n} \gtrsim \frac{a k_u}{\sqrt n}.\] Hence the upper and lower bounds match up to logarithmic factors.

Furthermore, if \(K/k_u^2\ge p^c\) for some \(c>0\), then 25 gives \[\frac{\nu_1}{\sqrt n} \asymp a k_u\sqrt{\frac{\log p}{n}},\] which matches the upper bound up to constants and thus \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp a k_u\sqrt{\frac{\log p}{n}}.\]

Finally, suppose \(k_u\ll K\ll k_u^2\). Then \(\nu_2=H(k_u;\xi)\asymp a\sqrt{k_u}\), while 25 gives \[\nu_1\asymp a\sqrt K.\] The lower bound in 1 therefore gives \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \gtrsim \frac{a\sqrt K}{\sqrt n} \vee a\sqrt{k_u}\frac{k_u\log p}{n} \asymp a\left\{ \frac{\sqrt K}{\sqrt n} + \frac{k_u^{3/2}\log p}{n} \right\},\] where the last equivalence uses that the maximum and the sum are the same up to a universal constant. If \(K\le n/\log p\), the upper bound in 128 reduces to \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \lesssim a\sqrt{K\wedge \frac{n}{\log p}}\, \frac{k_u\log p}{n} \asymp a\sqrt K\,\frac{k_u\log p}{n}.\]

The ratios of the upper bound to the two lower-bound terms are \[\frac{a\sqrt K\, k_u\log p/n}{a\sqrt K/\sqrt n} = \frac{k_u\log p}{\sqrt n}, \qquad \frac{a\sqrt K\, k_u\log p/n}{a k_u^{3/2}\log p/n} = \sqrt{\frac{K}{k_u}}.\] Thus the additional polynomial-divergence assumptions in the proposition make the upper bound polynomially larger than each available lower-bound term.

This proves the intermediate-regime display and completes the proof. ◻

12.2 Dense nonregular loading profiles↩︎

This subsection contains two dense profiles that violate the regular loading condition. Since 1 has already determined the adaptive separation distance \(\nu_1/\sqrt{n}\) up to logarithmic factors, we will focus on the moderately sparse regime in this subsection. Let \[N=\left\lfloor \frac{n}{\log p}\right\rfloor.\]

The common purpose of the two examples is to demonstrate that the regular loading condition 3 is sufficient but not necessary for determining the adaptive separation rate through \(\nu_1/\sqrt n\) in the moderately sparse regime.

In both examples, the leading coordinates at the scale \(N\) carry dense, nearly flat energy even though the full support of \(\xi\) is nonregular. The examples differ in how this nonregularity appears: the first has a nearly flat leading block and allows a small nonzero tail, while the second lets the coordinate magnitudes vary logarithmically across the full support.

We again assume \(\{|\xi_j|\}\) are sorted in the decreasing order.

Example 1: nearly flat leading block with a small tail.

For the first example, suppose that there exist constants \(c_0,C_0,C_1,c_1>0\), a scale \(a>0\), and an integer \(K\) such that \(N\le K\ll p\), \(K\ge k_u^2 p^{c_1}\), \(k_u\leq p^{c_2}\) with \(c_2<c_1\), and the loading vector satisfies \[\label{eq:nearly-flat-leading-block} \begin{cases} c_0a\le |\xi_j|\le C_0a, \quad \text{ for } 1\le j\le K, \\ \xi_j\neq 0, \quad \text{ for } j > K, \\ \sum_{j>K}\xi_j^2\le C_1Ka^2. \end{cases}\tag{129}\] In other words, the leading \(K\) elements are of the same order while the remaining coordinates do not dominate the total energy.

We note that the loading vector fails the regular loading condition 3 . Since \(0<\xi_p^2\leq (p-K)^{-1}\sum_{j>K}\xi_j^2\lesssim a^2 K/(p-K)\), we have \(|\xi_p|/a\to0\) while \(|\xi_p|>0\). Then \[\frac{\max_{j\in\operatorname{supp}(\xi)}|\xi_j|}{\min_{j\in\operatorname{supp}(\xi)}|\xi_j|} \to\infty .\]

Furthermore, 129 excludes the exact polynomially decaying profile considered in [8], since they assume \(\xi_j\asymp j^{-\alpha}\) with fixed \(\alpha>0\). This is because the ratio between its first and \(K\)th coordinates is \(K^\alpha\) and does not meet 129 .

Thus, in this nonregular dense example, the regular-loading optimality theory of [7] and the polynomial-decay analysis of [8] do not apply.

Proposition 6. Suppose \(k_u\gg \sqrt n/\log p\) and the loading vector satisfies 129 . Then, uniformly over \(k\le k_u\), \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp a k_u\sqrt{\frac{\log p}{n}}.\]

Proof. Since \(N\le K\) and the first \(K\) coordinates are all of order \(a\), \[H(N~;~\xi) = \left(\sum_{j\le N}\xi_j^2\right)^{1/2} \asymp a\sqrt N.\] Because \(k_u\gg \sqrt n/\log p\), 1 gives \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \lesssim \frac{k_u\log p}{n}H(N~;~\xi) \asymp \frac{k_u\log p}{n}a\sqrt{\frac{n}{\log p}} = a k_u\sqrt{\frac{\log p}{n}}.\]

It remains to show that the \(\nu_1\)-term gives the matching lower bound. Let \(x_j=|\xi_j|\), and define \[F(z) = \frac{\sum_{j=1}^{k_\xi}x_j\exp(-z/x_j^2)}{\left(\sum_{j=1}^{k_\xi}x_j^2\exp(-z/x_j^2)\right)^{1/2}}, \qquad z\in\mathbb{R}.\] By 17 , \(\lambda=\sqrt{\zeta_+}\), where \(F(\zeta)=k_u/2\).

We claim that \(\lambda\gtrsim a\sqrt{\log p}\). Let \(z=c_\star a^2\log p\), where \(c_\star>0\) is a sufficiently small constant. For \(j\leq K\), we have \(x_j\exp(-z/x_j^2)\geq c_0 a \exp(-c_0^2 z/a^2)\). Combining this with \(\sum_{j>K}\xi_j^2\le C_1Ka^2\), we have \[F(c_\star a^2\log p) \gtrsim \frac{Ka\exp(-c_0^2 z/a^2)}{a\sqrt K} = \sqrt K \exp(-c_0^2 c_\star\log p).\] Since \(K/k_u^2\ge p^{c_1}\), we can choose \(c_\star>0\) sufficiently small such that \(c_1-c_0^2c_\star>c_2\). Then for sufficiently large \(p\), we have \[F(c_\star a^2\log p) \geq \exp( c_2 \log p) \ge k_u.\] As \(F\) is decreasing and \(F(\zeta)=k_u/2\), this implies \[\lambda^2=\zeta_+\gtrsim a^2\log p.\] Consequently, \[\nu_1 = \lambda k_u+ \left(\sum_{j=1}^{k_\xi}\xi_j^2\exp(-\lambda^2/\xi_j^2)\right)^{1/2} \ge \lambda k_u \gtrsim a k_u\sqrt{\log p}.\] The lower bound in 1 therefore yields \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \gtrsim \frac{\nu_1}{\sqrt n} \gtrsim a k_u\sqrt{\frac{\log p}{n}}.\] This matches the upper bound and proves the proposition. ◻

This example reveals that, even if a small nonzero tail makes the loading vector nonregular, the adaptive separation distance is the same as for a dense regular loading vector because the leading \(N\) coordinates already contain dense, nearly flat energy.

Example 2: logarithmically varying dense loading vectors.

The second example places the nonregularity across the full support rather than in a small tail. It gives a dense loading profile that violates the regular loading condition only through a logarithmic variation across coordinates.

Let \(b>0\) be fixed and suppose \[|\xi_j|=a\{\log(e+j)\}^{-b}, \qquad 1\le j\le p,\] where \(a>0\). Then \[\frac{|\xi_1|}{|\xi_p|} \asymp (\log p)^b\to\infty,\] so the regular loading condition fails. This logarithmically varying profile is also outside the exact polynomial-decay class in [8]. For instance, we can check the ratio of the \(j\)th and \(2j\)th coordinates for \(j=\lfloor \sqrt{p}\rfloor\): \[\frac{|\xi_j|}{|\xi_{2j}|} = \left\{\frac{\log(e+2j)}{\log(e+j)}\right\}^{b} \to 1, \text{ as } p\to \infty,\] whereas a polynomial-decay profile \(j^{-\alpha}\) with \(\alpha>0\) has a constant dyadic ratio \(2^\alpha>1\). Hence neither the regular-loading condition used in [7] nor the exact polynomial-decay condition treated in [8] covers this example.

Proposition 7. Assume that \(N\to\infty\), \(\log N\asymp \log p\), and that \(k_u\le Cp^\gamma\) for some \(\gamma<1/2\). Suppose \(k_u\gg \sqrt n/\log p\). Then, uniformly over \(k\le k_u\), \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp a k_u\sqrt{\frac{\log p}{n}}\frac{1}{(\log p)^b}.\]

Proof. We first compute the top-\(r\) norm for \(r\) sufficiently large.

A simple comparison gives \[\sum_{j\le r}\{\log(e+j)\}^{-2b} \asymp \frac{r}{(\log r)^{2b}}.\] Indeed, the lower bound follows by summing over \(r/2<j\le r\). For the upper bound, split the sum at \(\lfloor\sqrt r\rfloor\): \[\sum_{j\le r}\{\log(e+j)\}^{-2b} \le \sqrt r + \sum_{\sqrt r<j\le r} \left(\frac{1}{2}\log r\right)^{-2b} \lesssim \frac{r}{(\log r)^{2b}},\] where the last step uses \((\log r)^{2b}=o(\sqrt r)\). Thus \[H(r~;~\xi)\asymp a\frac{\sqrt r}{(\log r)^b}.\] In the moderately sparse regime, 1 gives \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \lesssim \frac{k_u\log p}{n}H(N~;~\xi).\] Using the condition that \(\log N\asymp\log p\) and the definition of \(N=n/\log p\) (up to integer rounding), we have \[\frac{k_u\log p}{n}H(N~;~\xi) \asymp \frac{k_u\log p}{n} \cdot a\frac{\sqrt{n/\log p}}{(\log N)^b} \asymp a k_u\sqrt{\frac{\log p}{n}}\frac{1}{(\log p)^b}.\]

We now prove the matching lower bound through \(\nu_1\). Let \(x_j=|\xi_j|\). We have \[\label{eq:log-varying-example-numerator} x_j\asymp a_p:=a(\log p)^{-b}, ~~ p/2\le j\le p,\tag{130}\] and \[\label{eq:log-varying-example-denominator} \sum_{j=1}^p \xi_j^2 = H^2(p;\xi) \asymp a_p^2 \cdot p .\tag{131}\] As in the proof of 6, define \[F(z) = \frac{\sum_{j=1}^{p}x_j\exp(-z/x_j^2)}{\left(\sum_{j=1}^{p}x_j^2\exp(-z/x_j^2)\right)^{1/2}}.\] Let \(z=c_\star a_p^2\log p\), where \(c_\star>0\) is sufficiently small. Using 130 for the block \(p/2\le j\le p\) in the numerator and 131 for the total energy bound in the denominator, we have \[F(z) \gtrsim \sqrt p\,\exp(-Cc_\star\log p),\] where \(C>0\) is a constant. Since \(k_u\le Cp^\gamma\) with \(\gamma<1/2\), choosing \(c_\star\) sufficiently small gives \[F(z)\gtrsim k_u.\] Therefore \(\lambda^2\gtrsim a_p^2\log p\), and hence \[\lambda\gtrsim a(\log p)^{-b}\sqrt{\log p}.\] It follows that \[\nu_1\ge \lambda k_u \gtrsim a k_u\frac{\sqrt{\log p}}{(\log p)^b}.\] 1 then gives \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \gtrsim \frac{\nu_1}{\sqrt n} \gtrsim a k_u\sqrt{\frac{\log p}{n}}\frac{1}{(\log p)^b}.\] This matches the upper bound. ◻

There is an interesting equivalence between this example and a dense flat loading profile. Although the coordinate ratio diverges as \((\log p)^b\) in the current example, the loading vector behaves the same as a dense loading vector if we replace the common scale \(a\) by the effective scale \(a(\log p)^{-b}\) on the critical dense block.

12.3 A multiscale example for the statistical–computational gap↩︎

This example shows that the statistical–computational gap in the moderately sparse regime can occur for loading vectors outside the regular-loading phase diagram in 3.2. The construction in 134 puts equal \(\ell_2\)-energy on \(L\) blocks whose sizes increase with the block index. The example is interpreted together with [thm: computational lower bound,thm: sparse spiked]: 2 gives a low-degree lower bound for the general unknown-covariance problem, whereas 4 gives a smaller adaptive separation distance under the sparse signed-spiked covariance restriction.

Fix a degree \(D\) satisfying the assumptions of 2, and define \[k_{\mathrm{eff}} = \left\lfloor \frac{n}{\log p}\wedge \frac{k_u^2}{D\log p} \right\rfloor .\] Assume that \[k_u\gg \frac{\sqrt n}{\log p}.\] Let \(L=L_n\to\infty\) be an integer such that \[\label{eq:multiscale-L-conditions} L^3\le c_0 k_u, \qquad k_uL^3\le c_0 k_{\mathrm{eff}},\tag{132}\] where \(c_0>0\) is a sufficiently small absolute constant. Define block sizes \[\label{eq:multiscale-block-sizes} m_\ell=\lceil k_u\ell^2\rceil, \qquad \ell=1,\ldots,L,\tag{133}\] and cumulative indices \[M_0=0, \qquad M_\ell=\sum_{r=1}^{\ell}m_r, \qquad \ell=1,\ldots,L.\] For a scale \(a>0\), define the loading vector by \[\label{eq:multiscale-loading} |\xi_j| = \frac{a}{\sqrt{m_\ell}}, \qquad M_{\ell-1}<j\le M_\ell, \quad \ell=1,\ldots,L, \qquad \xi_j=0 \quad \text{for } j>M_L .\tag{134}\] Since \(m_\ell\) is increasing in \(\ell\), the coordinates in 134 are ordered in decreasing absolute value. Moreover, 133 and 132 give \[\label{eq:multiscale-support-size} M_L = \sum_{\ell=1}^L m_\ell \asymp k_u\sum_{\ell=1}^L\ell^2 \asymp k_uL^3 \le k_{\mathrm{eff}} \le \frac{n}{\log p}.\tag{135}\] The loading vector in 134 is not regular when \(L\to\infty\), because the ratio between the largest and smallest nonzero coordinates is \[\frac{a/\sqrt{m_1}}{a/\sqrt{m_L}} = \sqrt{\frac{m_L}{m_1}} \asymp L.\] It is also not an exact polynomially decaying loading vector in the sense of [8]: the profile in 134 has flat blocks whose lengths diverge. In particular, the first block contains \(m_1=\lceil k_u\rceil\) coordinates with magnitude \(a/\sqrt{m_1}\). In contrast, an exact polynomial-decay profile \(a_0j^{-\alpha}\), with \(\alpha>0\), would satisfy \(|\xi_1|/|\xi_{m_1}|=m_1^\alpha\to\infty\). Thus the construction in 134 is not an exact deterministic polynomial-decay profile. Consequently, the example is not covered directly by either the regular-loading theory of [7] or the exact polynomial-decay analysis of [8].

Proposition 8. For the multiscale loading vector in 134 , the following relations hold: \[H(k_{\mathrm{eff}}~;~\xi) \asymp H(n/\log p~;~\xi) \asymp a\sqrt L, \qquad \nu_2\asymp a, \qquad \nu_1\asymp a\sqrt L.\] Consequently, the upper bound in 1 is of order \[\label{eq:multiscale-computational-scale} a\sqrt L\,\frac{k_u\log p}{n},\qquad{(2)}\] and 2 gives a low-degree lower bound of the order in ?? . By contrast, under the sparse signed-spiked covariance model of 4, we have \[\tau^{\text{spike}}_{\text{adap}}(k_u,k~;~\xi) \asymp_{\log} \frac{a\sqrt L}{\sqrt n} + a\frac{k_u\log p}{n}.\]

If \(\sqrt L\wedge \frac{k_u\log p}{\sqrt n}\) diverges polynomially in \(p\), the low-degree scale exceeds the sparse signed-spiked rate by a polynomial factor.

Proof. For each block \(\ell\), 134 gives constant \(\ell_2\)-energy: \[\label{eq:multiscale-block-energy} \sum_{j=M_{\ell-1}+1}^{M_\ell}\xi_j^2 = m_\ell\frac{a^2}{m_\ell} = a^2.\tag{136}\] By 135 , the support of \(\xi\) is a subset of the first \(k_{\mathrm{eff}}\) coordinates and also a subset of the first \(n/\log p\) coordinates. Therefore, the definition of \(H\) and the block-energy identity 136 give \[H(k_{\mathrm{eff}}~;~\xi)^2 = H(n/\log p~;~\xi)^2 = \sum_{\ell=1}^L a^2 = La^2,\] which implies \[\label{eq:multiscale-H-result} H(k_{\mathrm{eff}}~;~\xi) \asymp H(n/\log p~;~\xi) \asymp a\sqrt L.\tag{137}\]

The block with \(\ell=1\) has size \(m_1=\lceil k_u\rceil\) and coordinate magnitude \(a/\sqrt{m_1}\). Therefore, we have \[\label{eq:multiscale-nu2} \nu_2 = H(k_u~;~\xi) = \left(m_1\frac{a^2}{m_1}\right)^{1/2} = a.\tag{138}\]

It remains to compute \(\nu_1\). Let \(x_j=|\xi_j|\), and define \[F(z) = \frac{\sum_j x_j\exp(-z/x_j^2)}{\left(\sum_j x_j^2\exp(-z/x_j^2)\right)^{1/2}}.\] This \(F\) is the left-hand side of 17 with \(z\) in place of \(\zeta\). At \(z=0\), \[F(0) = \frac{\sum_j x_j}{\left(\sum_jx_j^2\right)^{1/2}}.\] For block \(\ell\), 133 and 134 imply \[\sum_{j=M_{\ell-1}+1}^{M_\ell}x_j = m_\ell\frac{a}{\sqrt{m_\ell}} = a\sqrt{m_\ell} \asymp a\sqrt{k_u}\,\ell .\] Summing this blockwise identity over \(\ell=1,\ldots,L\) gives \[\sum_jx_j \asymp a\sqrt{k_u}\sum_{\ell=1}^L \ell \asymp a\sqrt{k_u}L^2.\] The same block-energy identity 136 gives \[\label{eq:multiscale-total-l2} \left(\sum_jx_j^2\right)^{1/2} = a\sqrt L.\tag{139}\] Combining the formula for \(F(0)\), the \(\ell_1\) computation, and 139 , we obtain \[F(0)\asymp \sqrt{k_u}L^{3/2}.\] By choosing \(c_0>0\) sufficiently small in 132 , this gives \[F(0)\le k_u/2.\] Since \(F\) is decreasing by 22 and the defining equation 17 is \(F(\zeta)=k_u/2\), the inequality \(F(0)\le k_u/2\) implies \(\zeta\le0\). Therefore \(\lambda=\sqrt{\zeta_+}=0\). Substituting \(\lambda=0\) into 18 and using 139 gives \[\label{eq:multiscale-nu1} \nu_1 = \left(\sum_j\xi_j^2\right)^{1/2} = a\sqrt L.\tag{140}\] Equations 137 , 138 , and 140 prove the three estimates stated at the beginning of the proposition.

By the moderately sparse assumption \(k_u\gg \sqrt n/\log p\), the moderately sparse branch of 1 applies. Combining that branch with 137 gives the computationally feasible upper bound \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \lesssim H(n/\log p~;~\xi)\frac{k_u\log p}{n} \asymp a\sqrt L\frac{k_u\log p}{n}.\] Similarly, 2 and 137 give a low-degree lower bound at separation \[H(k_{\mathrm{eff}}~;~\xi)\frac{k_u\log p}{n} \asymp a\sqrt L\frac{k_u\log p}{n}.\] Thus the computationally feasible upper bound and the low-degree scale both have order \(a\sqrt L\,k_u\log p/n\).

Finally, 4, 140 , and 138 give, under the sparse signed-spiked covariance model, \[\tau^{\text{spike}}_{\text{adap}}(k_u,k~;~\xi) \asymp_{\log} \frac{\nu_1}{\sqrt n} + \nu_2\frac{k_u\log p}{n} \asymp_{\log} \frac{a\sqrt L}{\sqrt n} + a\frac{k_u\log p}{n}.\] Dividing this sparse signed-spiked rate by the low-degree scale \(H(k_{\mathrm{eff}}~;~\xi)k_u\log p/n\), and using 137 , gives \[\frac{ \tau^{\text{spike}}_{\text{adap}}(k_u,k~;~\xi) }{ H(k_{\mathrm{eff}}~;~\xi)k_u\log p/n } \asymp_{\log} \frac{ a\sqrt L/\sqrt n+a k_u\log p/n }{ a\sqrt L\,k_u\log p/n } = \frac{1}{\sqrt L}+\frac{\sqrt n}{k_u\log p}.\] If \(\min(\sqrt L,~ k_u\log p/\sqrt n)\) diverges polynomially in \(p\), then the displayed ratio is polynomially small, which proves the claimed separation. ◻

12.4 I.i.d. sub-Weibull random-predictor loadings↩︎

This subsection treats random-predictor loadings, where \(\xi\) is the covariate vector of a new test point. For this example, the additional randomness is the draw of \(\xi\). After the test point is observed, we condition on its realized value, relabel coordinates so that \[|\xi_1|\ge |\xi_2|\ge\cdots\ge |\xi_p|,\] and evaluate the deterministic quantities \(H(\cdot~;~\xi)\), \(\nu_1\), and \(\nu_2\) from 16 , 18 , and 19 . Thus the probability statements below concern the draw of \(\xi\). Throughout 12.4, let \[N=\left\lfloor \frac{n}{\log p}\right\rfloor .\]

Let \(W_1,\ldots,W_p\) be i.i.d. nonnegative random variables, and let \(\xi\) be the decreasing rearrangement of \((W_1,\ldots,W_p)\). Assume that \(W\) has a two-sided sub-Weibull tail: for some constants \(q>0\), \(0<c_1,c_2,C_1,C_2<\infty\), and \(t_0>0\), \[\label{eq:subweibull-tail} c_1\exp(-C_1 t^q) \le \mathbb{P}(W\ge t) \le C_2\exp(-c_2 t^q), \qquad t\ge t_0.\tag{141}\] This tail condition includes absolute Gaussian coordinates \((q=2)\), absolute sub-exponential coordinates \((q=1)\), and other light-tailed predictors, up to constants in the exponent. Assume \[\label{eq:subweibull-scaling} 1\ll k_u\le N, \qquad N=o(p), \qquad k_u\le p^\gamma \quad\text{for some fixed }\gamma<1/2.\tag{142}\]

With probability tending to one, the realized loading vector lies outside the two deterministic classes most closely related to this example.

  • It fails the regular-loading condition of [7]. The lower bound in 141 implies there exists some constant \(c_1>0\) such that \(\max_{j\le p}W_j\ge c_1 (\log p)^{1/q}\) with probability tending to 1. Furthermore, it also implies that there exists a fixed \(M\) such that the interval \([t_0,M]\) receives positive probability mass. Therefore, with probability tending to 1, there is at least one nonzero coordinate with magnitude at most \(M\). When both events happen, we have \[\frac{\max_{j\in{\rm supp}(\xi)}|\xi_j|}{\min_{j\in{\rm supp}(\xi)}|\xi_j|} \gtrsim (\log p)^{1/q},\] which is unbounded.

  • It is also not the exact deterministic polynomial-decay profile treated in [8]. The order-statistic comparison in 143 in the proof of 9 shows that for \(j\to\infty\) with \(2j\le N\), the ratio between \(|\xi_j|\) and \(|\xi_{2j}|\) is of order \[\frac{\{\log(ep/j)\}^{1/q}}{\{\log(ep/(2j))\}^{1/q}} \to 1,\] whereas \(a_0j^{-\alpha}\) with \(\alpha>0\) has dyadic ratio \(2^\alpha\).

Therefore, the regular-loading theory of [7] and the exact polynomial-decay analysis of [8] do not apply directly to this random-predictor loading.

Proposition 9. Under 141 and 142 , with probability tending to one, we have \[\label{eq:subweibull-H-rate} H(r~;~\xi) = \left(\sum_{j\le r}\xi_j^2\right)^{1/2} \asymp \sqrt r\,\left\{\log\left(\frac{ep}{r}\right)\right\}^{1/q},\qquad{(3)}\] uniformly for \(1\le r\le N\). Moreover, we have \[\label{eq:subweibull-nu1-rate} \nu_1 \asymp k_u\left\{\log\left(\frac{ep}{k_u^2}\right)\right\}^{1/2+1/q}.\qquad{(4)}\] Consequently, in the moderately sparse regime \(k_u\gg \sqrt n/\log p\), we have \[\label{eq:subweibull-upper-rate} \tau_{\mathrm{adap}}(k_u,k~;~\xi) \lesssim k_u\sqrt{\frac{\log p}{n}} \left\{\log\left(\frac{ep}{N}\right)\right\}^{1/q},\qquad{(5)}\] while 1 gives \[\label{eq:subweibull-lower-rate} \tau_{\mathrm{adap}}(k_u,k~;~\xi) \gtrsim \frac{k_u}{\sqrt n} \left\{\log\left(\frac{ep}{k_u^2}\right)\right\}^{1/2+1/q}.\qquad{(6)}\] In the regime where it holds that \[\label{eq:subweibull-matched-cond} \log(ep/N)\asymp \log(ep/k_u^2)\asymp \log p,\qquad{(7)}\] we have \[\label{eq:subweibull-matched-rate} \tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp k_u\frac{(\log p)^{1/2+1/q}}{\sqrt n}.\qquad{(8)}\]

For the moderately sparse regime \(\sqrt{n}/\log p \ll k_u\lesssim n/\log p\), ?? holds if \(n=O(p^{1-\varepsilon})\) for some \(\varepsilon>0\) and 3 holds.

For Gaussian coordinates \(q=2\), ?? becomes \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp k_u\frac{\log p}{\sqrt n}\] under polynomial aspect-ratio scaling. For sub-exponential coordinates \(q=1\), ?? becomes \(\tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp k_u(\log p)^{3/2}/\sqrt n\). Formally, bounded predictors correspond to the limiting case \(q=\infty\), and the displayed rate then reduces to \(\tau_{\mathrm{adap}}(k_u,k~;~\xi) \asymp k_u\sqrt{\log p/n}\), which is the same as the rate for dense regular loadings.

12.4.1 Proof of Proposition 9↩︎

The proof of 9 makes use of two auxiliary facts: a high-probability order-statistic comparison and a deterministic logarithmic-sum comparison.

Lemma 26 (Order-statistic envelope). Under 141 and 142 , there exist constants \(0<c<C<\infty\), depending only on the constants in 141 , such that, with probability tending to one, \[\label{eq:subweibull-order-stat-event} c\left\{\log\left(\frac{ep}{j}\right)\right\}^{1/q} \le |\xi_j| \le C\left\{\log\left(\frac{ep}{j}\right)\right\}^{1/q}, \qquad 1\le j\le N.\tag{143}\]

Lemma 27 (Logarithmic sum comparison). Suppose \(N=o(p)\). Then, for every fixed \(q>0\), the following holds uniformly for \(1\le r\le N\): \[\sum_{j\le r} \left\{\log\left(\frac{ep}{j}\right)\right\}^{2/q} \asymp r\left\{\log\left(\frac{ep}{r}\right)\right\}^{2/q}.\]

By 26, with probability tending to one, 143 holds uniformly for \(1\le j\le N\). We work on this event in the following.

Using 143 , uniformly for \(1\le r\le N\), \[H(r~;~\xi)^2 \asymp \sum_{j\le r} \left\{\log\left(\frac{ep}{j}\right)\right\}^{2/q}.\] By 27, the sum on the right satisfies \[\sum_{j\le r} \left\{\log\left(\frac{ep}{j}\right)\right\}^{2/q} \asymp r\left\{\log\left(\frac{ep}{r}\right)\right\}^{2/q}, \qquad 1\le r\le N.\] Taking square roots proves ?? .

We next compute \(\nu_1\). As in the preceding examples, let \(F_\xi(z)\) denote the left-hand side of 17 : \[F_\xi(z) = \frac{\sum_{j=1}^{k_\xi}|\xi_j| \exp(-z/\xi_j^2)}{\left(\sum_{j=1}^{k_\xi}\xi_j^2 \exp(-z/\xi_j^2)\right)^{1/2}} .\] For the tail estimates below, it is more convenient to work with the squared version on the \(t=\sqrt z\) scale. For \(t>0\), define \[\label{eq:subweibull-D-definition} D_\xi(t) = F_\xi(t^2)^2 = \frac{ \left(\sum_{j=1}^{k_\xi}|\xi_j|e^{-t^2/\xi_j^2}\right)^2 }{ \sum_{j=1}^{k_\xi}\xi_j^2 e^{-t^2/\xi_j^2} }.\tag{144}\] The equation 17 is equivalent to \[D_\xi(\lambda)=k_u^2/4 \text{ if } \lambda>0.\]

We state the following result but defer its proof to the end of the section.

Lemma 28 (Localization of the solution defining \(\nu_1\)). Assume 141 and 142 . Let \[L=\log\left(\frac{ep}{k_u^2}\right), \qquad s=L^{(q+2)/(2q)}.\] The following properties hold with probability tending to one:

  1. There exist constants \(0<c_-<C_+<\infty\) such that \[\label{eq:subweibull-upper-crossing} D_\xi(C_+s)\le \frac{k_u^2}{8},\tag{145}\] and \[\label{eq:subweibull-lower-crossing} D_\xi(c_-s)\ge \frac{k_u^2}{2}.\tag{146}\] Consequently, the solution \(\lambda\) in 17 is positive and satisfies \[\label{eq:subweibull-lambda-localization} c_-s\le \lambda\le C_+s.\tag{147}\]

  2. It holds that \[\label{eq:subweibull-S2-at-lambda} \left(\sum_{j=1}^{k_\xi}\xi_j^2\exp(-\lambda^2/\xi_j^2)\right)^{1/2} \lesssim k_u L^{1/q}.\tag{148}\]

Let \[S_2(t) = \sum_{j=1}^{k_\xi}\xi_j^2\exp(-t^2/\xi_j^2).\]

By 28, we have \[\lambda \asymp \left\{\log\left(\frac{ep}{k_u^2}\right)\right\}^{1/2+1/q},\] and \[S_2(\lambda)^{1/2} \lesssim k_u \left\{\log\left(\frac{ep}{k_u^2}\right)\right\}^{1/q}.\] Therefore, by the definition of \(\nu_1\) in 18 , \[\begin{align} \nu_1 &= k_u\lambda+S_2(\lambda)^{1/2}\\ &\asymp k_u \left\{\log\left(\frac{ep}{k_u^2}\right)\right\}^{1/2+1/q}. \end{align}\] This proves ?? .

In the moderately sparse regime, 1 and \(N=\lfloor n/\log p\rfloor\) give \[\tau_{\mathrm{adap}}(k_u,k~;~\xi) \lesssim \frac{k_u\log p}{n}H(N~;~\xi).\] Using ?? with \(r=N\), the above upper bound becomes \[\begin{align} \tau_{\mathrm{adap}}(k_u,k~;~\xi) &\lesssim \frac{k_u\log p}{n} \sqrt{N} \left\{\log\left(\frac{ep}{N}\right)\right\}^{1/q}\\ &\asymp k_u\sqrt{\frac{\log p}{n}} \left\{\log\left(\frac{ep}{N}\right)\right\}^{1/q}, \end{align}\] which proves ?? .

The lower bound in ?? follows from the \(\nu_1/\sqrt n\) term in 1 and ?? .

If \[\log(ep/N)\asymp \log(ep/k_u^2)\asymp\log p,\] then the upper bound ?? and the lower bound ?? have the common order \[k_u\frac{(\log p)^{1/2+1/q}}{\sqrt n}\] up to logarithmic factors. This proves ?? and completes the proof of 9.

12.4.2 Proofs for Lemma 26 and Lemma 27↩︎

Proof of 26. For \(1\le j\le N\), put \[u_j=\left\{\log\left(\frac{ep}{j}\right)\right\}^{1/q}.\] Since \(N=o(p)\), we have \(u_N\to\infty\). Hence all thresholds used below exceed \(t_0\) uniformly over \(1\le j\le N\), for all sufficiently large \(p\).

We first prove the upper bound. Choose \(A_+>0\) so large that \[a_+ := c_2A_+^q>1.\] For each \(j\le N\), define the exceedance count \[X_j^+ = \sum_{i=1}^p \mathbf{1}\{W_i>A_+u_j\}.\] If \(|\xi_j|>A_+u_j\), then \(X_j^+\ge j\). Since \(X_j^+\) is binomial with mean \(\mu_j^+=p\mathbb{P}(W>A_+u_j)\), the upper tail bound in 141 gives \[\mu_j^+ \le C_2p\exp(-c_2A_+^q u_j^q) = C_2p\left(\frac{ep}{j}\right)^{-a_+}.\] Therefore, uniformly for \(1\le j\le N\), we have \[\frac{e\mu_j^+}{j} \le K_+\left(\frac{N}{p}\right)^{a_+-1} =:\rho_p,\] where \(K_+<\infty\) is a constant and \(\rho_p\to0\). Recall the standard binomial tail bound \[\mathbb{P}(X\ge j)\le(e\mu/j)^j\] where \(X\) is a binomial random variable with mean \(\mu\). Together with the fact that \(\rho_p<1\) for all sufficiently large \(p\), the standard binomial tail bound implies that \[\mathbb{P}\!\left(\exists\,1\le j\le N:\;|\xi_j|>A_+u_j\right) \le \sum_{j=1}^N \mathbb{P}(X_j^+\ge j) \le \sum_{j=1}^N \rho_p^j = o(1).\] Thus, with probability tending to one, \(|\xi_j|\le A_+u_j\) for all \(1\le j\le N\).

We next prove the lower bound. Choose \(A_->0\) so small that \[a_- :=C_1A_-^q<1.\] For each \(j\le N\), define \[X_j^- = \sum_{i=1}^p \mathbf{1}\{W_i\ge A_-u_j\}.\] If \(X_j^-\ge j\), then \(|\xi_j|\ge A_-u_j\). Denote the mean as \(\mu_j^-=p\mathbb{P}(W\ge A_-u_j)\). The lower tail bound in 141 gives the bound \[\mu_j^- \ge c_1p\exp(-C_1A_-^q u_j^q) = c_1p\left(\frac{ep}{j}\right)^{-a_-}.\] Consequently, denoting \(K_- = c_1e^{-a_-}>0\), we have \(\mu_j^-\ge K_-p^{1-a_-}\). Moreover, \[\frac{\mu_j^-}{j} \ge \left(\frac{p}{j}\right)^{1-a_-} \ge K_- \left(\frac{p}{N}\right)^{1-a_-} \to\infty,\] uniformly for \(1\le j\le N\). Consequently, for all sufficiently large \(p\), we have \(j\le \mu_j^-/2\) uniformly over \(1\le j\le N\). Chernoff’s lower-tail bound therefore gives \[\mathbb{P}(X_j^-<j) \le \mathbb{P}(X_j^-<\mu_j^-/2) \le \exp(-\mu_j^-/8).\] Using \(N=o(p)\) and \(\mu_j^-\ge K_-p^{1-a_-}\), we obtain \[\mathbb{P}\!\left(\exists\,1\le j\le N:\;|\xi_j|<A_-u_j\right) \le \sum_{j=1}^N \mathbb{P}(X_j^-<j) \le N\exp(-K_-p^{1-a_-}/8) = o(1).\] Combining the upper and lower events proves 143 with \(c=A_-\) and \(C=A_+\). ◻

Proof of 27. The lower bound is immediate when \(r=1\). When \(r\ge2\), the indices \(r/2<j\le r\) contribute at least a constant multiple of \(r\) terms, and each such term is at least \(\{\log(ep/r)\}^{2/q}\). Hence the sum is bounded below by a constant multiple of \(r\{\log(ep/r)\}^{2/q}\). This gives the lower bound.

For the upper bound, write \(L_r=\log(ep/r)\). Since \(r\le N=o(p)\), we have \(L_r\to\infty\) uniformly over \(1\le r\le N\). The function \(x\mapsto \{\log(ep/x)\}^{2/q}\) is decreasing on \((0,r]\), and hence \[\sum_{j\le r}\left\{\log\left(\frac{ep}{j}\right)\right\}^{2/q} \le \int_0^r \left\{\log\left(\frac{ep}{x}\right)\right\}^{2/q}\,dx .\] The integral is finite. With the change of variables \(y=\log(ep/x)\), \[\int_0^r \left\{\log\left(\frac{ep}{x}\right)\right\}^{2/q}\,dx = ep\int_{L_r}^{\infty} y^{2/q}e^{-y}\,dy .\] Putting \(y=L_r+u\), we obtain \[\int_{L_r}^{\infty} y^{2/q}e^{-y}\,dy = e^{-L_r}L_r^{2/q} \int_0^\infty \left(1+\frac{u}{L_r}\right)^{2/q}e^{-u}\,du \lesssim e^{-L_r}L_r^{2/q},\] where the last inequality holds uniformly for large \(p\), because \(L_r\to\infty\) and \(\int_0^\infty(1+u)^{2/q}e^{-u}\,du<\infty\). Since \(ep\,e^{-L_r}=r\), the upper bound is \(O(rL_r^{2/q})\). This proves the lemma. ◻

12.4.3 Preliminary results for proving Lemma 28↩︎

To prove 28, we need two more results.

Lemma 29 (A deterministic geometric saddle comparison). Fix \(q>0\), \(B>1\), and constants \(a,b>0\). For \(\ell=0,1,2\), there exist constants \(0<c<C<\infty\), depending only on \(q,B,a,b,\ell\), such that for all \(t\ge1\), \[\label{eq:subweibull-geometric-saddle} \sum_{m\in\mathbb{Z}} B^{m\ell} \exp\left\{-a\frac{t^2}{B^{2m}}-bB^{mq}\right\} \le C\,x(t)^\ell \exp\{-c\Phi(t)\},\tag{149}\] where \[\label{eq:subweibull-x-phi-def} x(t)=t^{2/(q+2)},\qquad \Phi(t)=x(t)^q=t^{2q/(q+2)}.\tag{150}\] Moreover, if \(m(t)\) is any integer satisfying \(B^{m(t)}\asymp x(t)\), then \[\label{eq:subweibull-geometric-saddle-lower} B^{m(t)\ell} \exp\left\{-a\frac{t^2}{B^{2m(t)}}-bB^{m(t)q}\right\} \ge c\,x(t)^\ell \exp\{-C\Phi(t)\}.\tag{151}\]

Lemma 30 (Grid-count event for sub-Weibull coordinates). Under 141 , there exist constants \(B>1\), \(A>0\), \(C_M<\infty\), and \(0<c<C<\infty\) such that the following event has probability tending to one:

  1. \[\label{eq:subweibull-max-bound} \max_{j\le p}|\xi_j|\le C_M(\log p)^{1/q}.\tag{152}\]

  2. For every grid interval \(I_m=[B^m,B^{m+1})\) intersecting \([t_0,C_M(\log p)^{1/q}]\), \[\label{eq:subweibull-grid-upper-count} \#\{j:|\xi_j|\in I_m\} \le (\log p)^A p\exp(-cB^{mq}).\tag{153}\]

  3. Whenever \(p\exp(-CB^{mq})\ge(\log p)^A\), \[\label{eq:subweibull-grid-lower-count} \#\{j:|\xi_j|\in I_m\} \ge (\log p)^{-A}p\exp(-CB^{mq}).\tag{154}\]

Proof of 29. Set \(x=x(t)\). We have \(t^2=x^{q+2}\).

We first prove the upper bound 149 .

Choose \(m_0\in\mathbb{Z}\) such that \[B^{m_0}\le x < B^{m_0+1}.\] For \(m=m_0+k\), we have \[B^m = B^{m_0}B^k = x\theta B^k, \qquad \theta:=\frac{B^{m_0}}{x}\in[B^{-1},1].\] It follows that \[\begin{align} \frac{t^2}{B^{2m}} &= \frac{x^{q+2}}{x^2\theta^2B^{2k}} = x^q\theta^{-2}B^{-2k},\\ B^{mq} &= x^q\theta^qB^{qk}. \end{align}\] Since \(\theta\in[B^{-1},1]\), there are constants \(c_1,C_1>0\), depending only on \(B,a,b,q\), such that \[\label{eq:geom-saddle-exponent-lower} a\frac{t^2}{B^{2m}}+bB^{mq} \ge c_1x^q\left(B^{-2k}+B^{qk}\right).\tag{155}\] Also, because \(\theta\le1\), and \(\ell\in \{0,1,2\}\), \[B^{m\ell} = x^\ell \theta^\ell B^{k\ell} \le x^\ell B^{k\ell}.\] Therefore \[\begin{align} &\sum_{m\in\mathbb{Z}} B^{m\ell} \exp\left\{-a\frac{t^2}{B^{2m}}-bB^{mq}\right\} \\ &\qquad\le x^\ell \sum_{k\in\mathbb{Z}} B^{k\ell} \exp\left\{-c_1x^q\left(B^{-2k}+B^{qk}\right)\right\}. \end{align}\] It remains to show that the last sum is at most a constant multiple of \(\exp(-c x^q)\).

We split the sum into \(k\ge0\) and \(k<0\).

For \(k\ge0\), the term with \(B^{qk}\) in the exponent dominates. Since \(t\ge1\), we have \(x^q\ge1\). Because \(\ell\) is fixed, the polynomial prefactor \(B^{k\ell}\) can be absorbed into the exponential: there exists \(c_2\in(0,c_1)\) such that, for all \(k\ge0\), \[B^{k\ell}\exp\{-c_1x^qB^{qk}\} \le \exp\{-c_2x^qB^{qk}\}.\] Hence \[\sum_{k\ge0} B^{k\ell} \exp\left\{-c_1x^q\left(B^{-2k}+B^{qk}\right)\right\} \le \sum_{k\ge0}\exp\{-c_2x^qB^{qk}\}.\] Since \(B^{qk}\) grows geometrically, the last sum is bounded by a constant multiple of its first term. Thus, for some constants \(C_2,c_3>0\), \[\sum_{k\ge0}\exp\{-c_2x^qB^{qk}\} \le C_2\exp(-c_3x^q).\]

For \(k<0\), write \(h=-k\ge1\). Then \(B^{-2k}=B^{2h}\), and \(B^{k\ell}\leq 1\). We have \[\begin{align} &\sum_{k<0} B^{k\ell} \exp\left\{-c_1x^q\left(B^{-2k}+B^{qk}\right)\right\} \\ &\qquad\le \sum_{h\ge1} \exp\{-c_1x^qB^{2h}\} \le C_3\exp(-c_4x^q), \end{align}\] again because \(B^{2h}\) grows geometrically.

Combining the bounds for \(k\ge0\) and \(k<0\), we obtain \[\sum_{m\in\mathbb{Z}} B^{m\ell} \exp\left\{-a\frac{t^2}{B^{2m}}-bB^{mq}\right\} \le C x^\ell \exp(-c x^q) = Cx(t)^\ell\exp\{-c\Phi(t)\}.\] This proves 149 .

We now prove the lower bound 151 . Let \(m(t)=\lceil \frac{\log(x(t))}{\log B}\rceil\). Then \(m(t)\) satisfies that \(B^{m(t)}\asymp x(t)\); that is, there are constants \(0<c_5<C_5<\infty\) such that \[c_5x\le B^{m(t)}\le C_5x.\] Hence \[B^{m(t)\ell}\asymp x^\ell.\] Moreover, \[\frac{t^2}{B^{2m(t)}}+B^{m(t)q} \asymp \frac{x^{q+2}}{x^2}+x^q \asymp x^q = \Phi(t).\] Therefore the single summand at \(m=m(t)\) satisfies \[B^{m(t)\ell} \exp\left\{-a\frac{t^2}{B^{2m(t)}}-bB^{m(t)q}\right\} \ge c x^\ell\exp(-C x^q),\] which is exactly 151 . ◻

Proof of 30. Choose \(C_M\) sufficiently large. By the upper tail bound in 141 , \[\mathbb{P}\!\left(\max_{j\le p} W_j>C_M(\log p)^{1/q}\right) \le pC_2\exp(-c_2C_M^q\log p) = C_2p^{1-c_2C_M^q} = o(1),\] which proves 152 .

Next choose \(B>1\) large enough that \[c_2B^q>C_1\] and \[\sup_{u\ge t_0} \exp\{C_1u^q-c_2B^qu^q\} < \frac{c_1}{2C_2}.\] Then, for all \(u\ge t_0\), \[\label{eq:subweibull-annulus-mass} \begin{align} \mathbb{P}(u\le W\le Bu) &= \mathbb{P}(W\ge u)-\mathbb{P}(W>Bu)\\ &\ge c_1\exp(-C_1u^q)-C_2\exp(-c_2B^q u^q)\\ &\ge \frac{c_1}{2}\exp(-C_1u^q). \end{align}\tag{156}\] Thus each grid interval \(I_m=[B^m,B^{m+1})\) with \(B^m\ge t_0\) has probability at least a constant multiple of \(\exp(-CB^{mq})\), while the upper tail bound gives \[\label{eq:subweibull-grid-mean-upper} \mathbb{P}(W\in I_m)\le \mathbb{P}(W\ge B^m)\le C_2\exp(-c_2B^{mq}).\tag{157}\]

We now justify the simultaneous count bounds. Let \[\mathcal{M}_p = \{m\in\mathbb{Z}: I_m=[B^m,B^{m+1}) \text{ intersects } [t_0,C_M(\log p)^{1/q}]\}.\] Then \(|\mathcal{M}_p|=O(\log\log p)\). For \(m\in\mathcal{M}_p\), define \[N_m = \#\{j:|\xi_j|\in I_m\} = \sum_{j=1}^p\mathbf{1}\{W_j\in I_m\}.\] Thus \(N_m\) is binomial with mean \(\mu_m:= p\,\mathbb{P}(W\in I_m)\). For the finitely many boundary intervals with \(B^m<t_0\), the bounds below can be absorbed by changing constants. Hence we focus on intervals with \(B^m\ge t_0\).

Choose a constant \(c>0\) smaller than \(c_2\), and also small enough that for all \(m\in\mathcal{M}_p\) and all sufficiently large \(p\), it holds that \[p\exp(-cB^{mq})\ge p^{1/2}.\] This is possible because \(B^{mq}\lesssim \log p\) on \(\mathcal{M}_p\).

For any \(A>0\) to be determined, set \[T_m = (\log p)^A p\exp(-cB^{mq}).\] By 157 , \[\frac{T_m}{e\mu_m} \ge C(\log p)^A \exp\{(c_2-c)B^{mq}\} \ge C(\log p)^A.\] By taking \(A\) large enough, we can ensure that \(T_m\ge e\mu_m\) for all \(m\in\mathcal{M}_p\). The binomial Chernoff bound yields \[\mathbb{P}(N_m\ge T_m) \le \left(\frac{e\mu_m}{T_m}\right)^{T_m}.\] Therefore, for some constant \(c'\) and sufficiently large \(p\), we have \[\mathbb{P}(N_m\ge T_m) \le \exp\{-c'T_m\log\log p\} \le \exp\{-c'p^{1/2}\log\log p\}.\] Taking a union bound over \(\mathcal{M}_p\), we obtain \[\mathbb{P}\left( \exists m\in\mathcal{M}_p \text{ such that } N_m> T_m \right) \le o(1),\] since \(|\mathcal{M}_p|=O(\log\log p)\). This proves the simultaneous upper count bound 153 .

We next prove the simultaneous lower count bound. By 156 , there exist constants \(c_\ell,C_\ell>0\) such that \[\label{eq:subweibull-grid-mean-lower} \mu_m = p\,\mathbb{P}(W\in I_m) \ge c_\ell p\exp(-C_\ell B^{mq})\tag{158}\] for all \(m\in \mathcal{M}_p\). Choose \(C\ge C_\ell\). For any \(A>0\) to be determined, consider only those \(m\in\mathcal{M}_p\) for which \[p\exp(-CB^{mq})\ge(\log p)^A.\] Set \[L_m = (\log p)^{-A}p\exp(-CB^{mq}).\] By 158 , \[\frac{L_m}{\mu_m} \le c_\ell^{-1}(\log p)^{-A} \exp\{-(C-C_\ell)B^{mq}\} \le c_\ell^{-1}(\log p)^{-A}.\] Taking \(A\) sufficiently large gives \(L_m\le \mu_m/2\) uniformly over these intervals. Hence Chernoff’s lower-tail bound implies \[\mathbb{P}(N_m<L_m) \le \mathbb{P}(N_m<\mu_m/2) \le \exp(-\mu_m/8).\] Moreover, since \(C\ge C_\ell\) and \(p\exp(-CB^{mq})\ge(\log p)^A\), we have \[\mu_m \ge c_\ell p\exp(-C_\ell B^{mq}) \ge c_\ell p\exp(-CB^{mq}) \ge c_\ell(\log p)^A.\] Therefore, for some constant \(c>0\), we have \[\mathbb{P}(N_m<L_m) \le \exp\{-c(\log p)^A\}.\] A union bound over \(O(\log\log p)\) intervals gives \[\mathbb{P}\left( \exists m\in\mathcal{M}_p: p\exp(-CB^{mq})\ge(\log p)^A \text{ and } N_m< L_m \right) =o(1).\] This proves the simultaneous lower count bound 154 . ◻

12.4.4 Proof of Lemma 28↩︎

To prove 28, we work on the event in 30. For \(t>0\), define \[S_\ell(t) = \sum_{j=1}^{k_\xi} |\xi_j|^\ell \exp(-t^2/\xi_j^2), \qquad \ell=0,1,2.\] Then \[D_\xi(t)=\frac{S_1(t)^2}{S_2(t)}.\]

First we prove 145 . By Cauchy’s inequality, \[S_1(t)^2\le S_0(t)S_2(t),\] and hence \[\label{eq:subweibull-D-by-S0} D_\xi(t)\le S_0(t).\tag{159}\] Set \(t_+=C_+s\). Coordinates with \(|\xi_j|\le t_0\) contribute at most \(p\exp(-t_+^2/t_0^2)\) to \(S_0(t_+)\). For the grid intervals above \(t_0\) as defined in 30, 153 implies that \[\label{eq:subweibull-S1-upper-tminus} \begin{align} S_0(t_+) &\le p\exp(-t_+^2/t_0^2) + \sum_m \#\{j:|\xi_j|\in I_m\} \exp(-t_+^2/B^{2m+2})\\ &\le p\exp(-t_+^2/t_0^2) + (\log p)^A p \sum_m \exp\{-t_+^2/B^{2m+2}-cB^{mq}\}. \end{align}\tag{160}\]

By 29, \[\label{eq:subweibull-S0-upper} S_0(t_+) \lesssim_{\log} p\exp\{-c_0 t_+^{2q/(q+2)}\} = p\exp\{-c_0 C_+^{2q/(q+2)}L\},\tag{161}\] for a constant \(c_0>0\). Choose \(C_+\) so large that \[a_+:=c_0 C_+^{2q/(q+2)}>1.\] Then \[\frac{p\exp(-a_+L)}{k_u^2} = e^{-a_+} \left(\frac{k_u^2}{p}\right)^{a_+-1} \to0,\] because \(k_u\le p^\gamma\) with \(\gamma<1/2\). The logarithmic factor in 161 is also negligible compared with the resulting polynomial decay. Combining this with 159 proves 145 .

We next prove 146 . Set \(t_-=c_-s\) and \[x_-=x(t_-)=t_-^{2/(q+2)}.\] Choose a grid interval \(I_m=[B^m,B^{m+1})\) such that \(B^m\asymp x_-\). Since \(x_-\asymp L^{1/q}\to\infty\), this interval lies above \(t_0\) for large \(p\). Also, \[B^{mq}\asymp x_-^q = t_-^{2q/(q+2)} = c_-^{2q/(q+2)}L.\] Taking \(c_->0\) sufficiently small, we have \[p\exp(-CB^{mq}) \ge (\log p)^A,\] so the lower count bound 154 applies to this interval. Write \(N_m=\#\{j:|\xi_j|\in I_m\}\). We have \[N_m \ge (\log p)^{-A}p\exp(-CB^{mq}).\] For every coordinate with \(|\xi_j|\in I_m=[B^m,B^{m+1})\), \[|\xi_j|\ge B^m, \qquad \exp(-t_-^2/\xi_j^2)\ge \exp(-t_-^2/B^{2m}).\] Therefore \[\begin{align} S_1(t_-) &= \sum_{j=1}^{k_\xi}|\xi_j|\exp(-t_-^2/\xi_j^2) \notag\\ &\ge \sum_{j:\,|\xi_j|\in I_m} |\xi_j|\exp(-t_-^2/\xi_j^2) \notag\\ &\ge N_m B^m\exp(-t_-^2/B^{2m}) \notag\\ &\ge (\log p)^{-A}pB^m \exp\{-t_-^2/B^{2m}-CB^{mq}\}. \label{eq:subweibull-S1-selected-interval} \end{align}\tag{162}\] The choice \(B^m\asymp x_-\) implies, after changing constants only by factors depending on \(B\), \[B^m\gtrsim x_-, \qquad \frac{t_-^2}{B^{2m}}\lesssim \frac{t_-^2}{x_-^2}, \qquad B^{mq}\lesssim x_-^q .\] Substituting these three comparisons into 162 gives \[S_1(t_-) \gtrsim_{\log} p\,x_-\, \exp\{-C_0(t_-^2/x_-^2+x_-^q)\}.\] Finally, since \(x_-=t_-^{2/(q+2)}\), we have \(\frac{t_-^2}{x_-^2} = t_-^{2q/(q+2)}\) and \(x_-^q = t_-^{2q/(q+2)}\). After increasing \(C_0\) if necessary, we have \[\label{eq:subweibull-S1-lower} S_1(t_-) \gtrsim_{\log} p\,x_-\, \exp\{-C_0t_-^{2q/(q+2)}\}.\tag{163}\]

For \(S_2(t_-)\), we now have \(\ell=2\) and we can still use the same argument in [eq:subweibull-S1-upper-tminus,eq:subweibull-S0-upper] together with 29 to get \[\label{eq:subweibull-S2-upper-tminus} \begin{align} S_2(t_-) &\le p t_0^2\exp(-t_-^2/t_0^2) + (\log p)^A p \sum_m B^{2m+2}\exp\{-t_-^2/B^{2m+2}-cB^{mq}\}\\ &\lesssim_{\log} p\,x_-^2\,\exp\{-c_0't_-^{2q/(q+2)}\}, \end{align}\tag{164}\] where \(c_0'\) is a positive constant. Combining 163 and 164 , we get \[\label{eq:subweibull-D-lower-general} D_\xi(t_-) = \frac{S_1(t_-)^2}{S_2(t_-)} \gtrsim_{\log} p\exp\{-C_1't_-^{2q/(q+2)}\}.\tag{165}\] Since \[t_-^{2q/(q+2)} = c_-^{2q/(q+2)}L,\] we can choose \(c_->0\) sufficiently small so that \[a_-:=C_1'c_-^{2q/(q+2)}<1.\] Then \[\frac{p\exp(-a_-L)}{k_u^2} = e^{-a_-} \left(\frac{p}{k_u^2}\right)^{1-a_-} \to\infty,\] again because \(k_u\le p^\gamma\) with \(\gamma<1/2\). This dominates the logarithmic loss in 165 , proving 146 .

Since 22 shows that \(F_\xi(z)\) is nonincreasing in \(z\), the function \(D_\xi(t)=F_\xi(t^2)^2\) is nonincreasing in \(t>0\). The two crossing inequalities 145 and 146 therefore imply 147 .

It remains to prove 148 . By 152 , \[\|\xi\|_\infty \le C_M(\log p)^{1/q} \asymp L^{1/q},\] where the last comparison follows from \(k_u\le p^\gamma\) with \(\gamma<1/2\), which implies \(L=\log(ep/k_u^2)\asymp\log p\). Since \[S_2(\lambda) = \sum_j \xi_j^2e^{-\lambda^2/\xi_j^2} \le \|\xi\|_\infty \sum_j |\xi_j|e^{-\lambda^2/\xi_j^2} = \|\xi\|_\infty S_1(\lambda),\] and since \(D_\xi(\lambda)=S_1(\lambda)^2/S_2(\lambda)=k_u^2/4\), we have \[S_2(\lambda) \le \|\xi\|_\infty \{D_\xi(\lambda)S_2(\lambda)\}^{1/2}.\] Therefore \[S_2(\lambda)^{1/2} \le \|\xi\|_\infty D_\xi(\lambda)^{1/2} \lesssim k_u L^{1/q},\] which proves 148 .

References↩︎

[1]
Robert Tibshirani, Regression shrinkage and selection via the lasso, Journal of the Royal Statistical Society: Series B (Statistical Methodology) 58(1996), no. 1, 267–288.
[2]
Peter J. Bickel, Ya’acov Ritov, and Alexandre B. Tsybakov, Simultaneous analysis of lasso and dantzig selector, The Annals of Statistics 37(2009), no. 4, 1705–1732.
[3]
, Minimax rates of estimation for high-dimensional linear regression over lq-balls, IEEE transactions on information theory 57(2011), no. 10, 6976–6994.
[4]
Cun-Hui Zhang and Stephanie S. Zhang, Confidence intervals for low dimensional parameters in high dimensional linear models, Journal of the Royal Statistical Society: Series B (Statistical Methodology) 76(2014), no. 1, 217–242.
[5]
Sara van de Geer, Peter Bühlmann, Ya’acov Ritov, and Ruben Dezeure, On asymptotically optimal confidence regions and tests for high-dimensional models, The Annals of Statistics 42(2014), no. 3, 1166–1202.
[6]
Adel Javanmard and Andrea Montanari, Confidence intervals and hypothesis testing for high-dimensional regression, Journal of Machine Learning Research 15(2014), no. 1, 2869–2909.
[7]
T. Tony Cai and Zijian Guo, Confidence intervals for high-dimensional linear regression: Minimax rates and adaptivity, The Annals of Statistics 45(2017), no. 2, 615–646.
[8]
Tianxi Cai, T. Tony Cai, and Zijian Guo, Optimal statistical inference for individualized treatment effects in high-dimensional models, Journal of the Royal Statistical Society: Series B (Statistical Methodology) 83(2021), no. 4, 669–719.
[9]
, Accuracy assessment for high-dimensional linear regression, The Annals of Statistics 46(2018), no. 4, 1807–1836.
[10]
Jelena Bradic, Jianqing Fan, and Yinchu Zhu, Testability of high-dimensional linear models with nonsparse structures, The Annals of Statistics 50(2022), no. 2, 615–639.
[11]
Samuel Hopkins, Statistical inference and the sum of squares method, Ph.D. thesis, 2018, p. 430.
[12]
Dmitriy Kunisky, Alexander S Wein, and Afonso S Bandeira, Notes on computational hardness of hypothesis testing: Predictions using the low-degree likelihood ratio, ISAAC Congress (International Society for Analysis, its Applications and Computation), Springer, 2019, pp. 1–50.
[13]
Chao Gao, Zongming Ma, and Harrison H. Zhou, Sparse CCA: Adaptive estimation and computational barriers, The Annals of Statistics 45(2017), no. 5, 2074 – 2101.
[14]
Nilanjana Laha and Rajarshi Mukherjee, On support recovery with sparse CCA: Information theoretic and computational limits, IEEE transactions on information theory 69(2023), no. 3, 1695–1738.
[15]
Galen Reeves, Jiaming Xu, and Ilias Zadik, The all-or-nothing phenomenon in sparse linear regression, Conference on Learning Theory, PMLR, 2019, pp. 2652–2663.
[16]
Guy Bresler, Sung Min Park, and Madalina Persu, Sparse pca from sparse linear regression, Advances in Neural Information Processing Systems, 2018.
[17]
Ildar Abdullovich Ibragimov and Rafail Zalmanovich Khas’ minskii, On nonparametric estimation of the value of a linear functional in gaussian white noise, Theory of Probability & Its Applications 29(1985), no. 1, 18–32.
[18]
T. Tony Cai and Mark G. Low, Minimax estimation of linear functionals over nonconvex parameter spaces, The Annals of Statistics 32(2004), no. 2, 552 – 576.
[19]
, On adaptive estimation of linear functionals, The Annals of Statistics 33(2005), no. 5, 2311 – 2343.
[20]
Olivier Collier, Laëtitia Comminges, and Alexandre B. Tsybakov, Minimax estimation of linear and quadratic functionals on sparsity classes, The Annals of Statistics 45(2017), no. 3, 923 – 958.
[21]
Jie Xie and Dongming Huang, Minimax and adaptive estimation of general linear functionals under sparsity, arXiv preprint arXiv:2509.25595 (2025).
[22]
Adel Javanmard and Jason D Lee, A flexible framework for hypothesis testing in high dimensions, Journal of the Royal Statistical Society Series B: Statistical Methodology 82(2020), no. 3, 685–718.
[23]
Yinchu Zhu and Jelena Bradic, Linear hypothesis testing in dense high-dimensional linear models, Journal of the American Statistical Association 113(2018), no. 524, 1583–1600.
[24]
Junlong Zhao, Yang Zhou, and Yufeng Liu, Estimation of linear functionals in high-dimensional linear models: From sparsity to nonsparsity, Journal of the American Statistical Association 119(2024), no. 546, 1579–1591.
[25]
Quentin Berthet and Philippe Rigollet, Complexity theoretic lower bounds for sparse principal component detection, Conference on learning theory, PMLR, 2013, pp. 1046–1066.
[26]
Zongming Ma and Yihong Wu, Computational barriers in minimax submatrix detection, The Annals of Statistics 43(2015), no. 61, 1089–1116.
[27]
Benjamin Rossman, Average-case complexity of detecting cliques, Ph.D. thesis, Massachusetts Institute of Technology, 2010.
[28]
Vitaly Feldman, Elena Grigorescu, Lev Reyzin, Santosh S Vempala, and Ying Xiao, Statistical algorithms and a lower bound for detecting planted cliques, Journal of the ACM (JACM) 64(2017), no. 2, 1–37.
[29]
Yash Deshpande and Andrea Montanari, Improved sum-of-squares lower bounds for hidden clique and hidden submatrix problems, Conference on Learning Theory, PMLR, 2015, pp. 523–562.
[30]
Aaron Potechin and Goutham Rajendran, Sub-exponential time sum-of-squares lower bounds for principal components analysis, Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 35724–35740.
[31]
Tengyu Ma and Avi Wigderson, Sum-of-squares lower bounds for sparse pca, Advances in Neural Information Processing Systems, vol. 28, 2015.
[32]
Yuetian Luo and Chao Gao, Computational lower bounds for graphon estimation via low-degree polynomials, The Annals of Statistics 52(2024), no. 5, 2318–2348.
[33]
Arnab Auddy and Ming Yuan, Large-dimensional independent component analysis: Statistical optimality and computational tractability, The Annals of Statistics 53(2025), no. 2, 477–505.
[34]
, Semi-supervised inference for explained variance in high-dimensional linear regression, Journal of the Royal Statistical Society: Series B (Statistical Methodology) 82(2020), no. 2, 391–419.
[35]
Erich L. Lehmann and Joseph P. Romano, Testing statistical hypotheses, 3rd ed., Springer Texts in Statistics, Springer, New York, 2005.
[36]
T. Sun and C.-H. Zhang, Scaled sparse linear regression, Biometrika 99(2012), no. 4, 879–898.
[37]
Afonso S Bandeira, Edgar Dobriban, Dustin G Mixon, and William F Sawin, Certifying the restricted isometry property is hard, IEEE transactions on information theory 59(2013), no. 6, 3448–3450.
[38]
Andreas M Tillmann and Marc E Pfetsch, The computational complexity of the restricted isometry property, the nullspace property, and related concepts in compressed sensing, IEEE Transactions on Information Theory 60(2013), no. 2, 1248–1259.
[39]
Alexandre B. Tsybakov, Introduction to nonparametric estimation, Springer Series in Statistics, Springer, 2009.
[40]
Julien Chhor, Rajarshi Mukherjee, and Subhabrata Sen, Sparse signal detection in heteroscedastic gaussian sequence models: sharp minimax rates, Bernoulli 30(2024), no. 3, 2127–2153.
[41]
Afonso S Bandeira, Ahmed El Alaoui, Samuel Hopkins, Tselil Schramm, Alexander S Wein, and Ilias Zadik, The franz-parisi criterion and computational trade-offs in high dimensional statistics, Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 33831–33844.
[42]
Rares-Darius Buhai, Jun-Ting Hsieh, Aayush Jain, and Pravesh K. Kothari, The quasi-polynomial low-degree conjecture is false, 2025 IEEE 66th Annual Symposium on Foundations of Computer Science (FOCS), 2025, pp. 2577–2590.
[43]
Alexander S Wein, Computational complexity of statistics: New insights from low-degree polynomials, arXiv preprint arXiv:2506.10748 (2025).
[44]
Chao Gao, Zongming Ma, Zhao Ren, and Harrison H. Zhou, Minimax estimation in sparse canonical correlation analysis, The Annals of Statistics 43(2015), no. 5, 2168 – 2197.
[45]
Cristina Butucea and Yuri I. Ingster, Detection of a sparse submatrix of a high-dimensional noisy matrix, Bernoulli 19(2013), no. 5B, 2652 – 2688.
[46]
Iain M. Johnstone and Arthur Yu Lu, On consistency and sparsity for principal components analysis in high dimensions, Journal of the American Statistical Association 104(2009), no. 486, 682–693.
[47]
T. Tony Cai, Zongming Ma, and Yihong Wu, Optimal estimation and rank detection for sparse spiked covariance matrices, Probability theory and related fields 161(2015), no. 3, 781–815.
[48]
Alden Green and Elad Romanov, The high-dimensional asymptotics of principal component regression, The Annals of Statistics 53(2025), no. 4, 1697–1727.
[49]
Yixuan Wu, Yilun Zhu, Lei Cao, and Naichen Shi, Calibrated principal component regression, The 29th International Conference on Artificial Intelligence and Statistics, 2026.
[50]
T. Tony Cai, Zijian Guo, and Rong Ma, Statistical inference for high-dimensional generalized linear models with binary outcomes, Journal of the American Statistical Association 118(2023), no. 542, 1319–1332.
[51]
Adel Javanmard and Andrea Montanari, Debiasing the lasso: Optimal sample size for Gaussian designs, The Annals of Statistics 46(2018), no. 6A, 2593 – 2622.
[52]
T. Tony Cai, Anru R. Zhang, and Yuchen Zhou, Sparse group lasso: Optimal sample complexity, convergence rate, and statistical inference, IEEE Transactions on Information Theory 68(2022), no. 9, 5975–6002.
[53]
Beatrice Laurent and Pascal Massart, Adaptive estimation of a quadratic functional by model selection, The Annals of statistics (2000), 1302–1338.
[54]
Roman Vershynin, Introduction to the non-asymptotic analysis of random matrices, Compressed sensing, Cambridge University Press, 2012, pp. 210–268.
[55]
Matthew Brennan, Guy Bresler, and Wasim Huleihel, Reducibility and computational lower bounds for problems with planted sparse structure, Conference On Learning Theory, PMLR, 2018, pp. 48–166.
[56]
Zongming Ma, Sparse principal component analysis and iterative thresholding, The Annals of Statistics 41(2013), no. 2, 772 – 801.
[57]
Kenneth R Davidson and Stanislaw J Szarek, Local operator theory, random matrices and banach spaces, Handbook of the geometry of Banach spaces, vol. 1, Elsevier, 2001, pp. 317–366.
[58]
T. Tony Cai and Harrison H. Zhou, Optimal rates of convergence for sparse covariance matrix estimation, The Annals of Statistics 40(2012), no. 5, 2389 – 2420.
[59]
, High-dimensional probability: An introduction with applications in data science, vol. 47, Cambridge university press, 2018.
[60]
Gabor Szeg, Orthogonal polynomials, vol. 23, American Mathematical Soc., 1939.
[61]
Cyril Banderier, Mireille Bousquet-Mélou, Alain Denise, Philippe Flajolet, Daniele Gardy, and Dominique Gouyou-Beauchamps, Generating functions for generating trees, Discrete mathematics 246(2002), no. 1-3, 29–55.
[62]
Xin Li and Chao-Ping Chen, Inequalities for the gamma function., JIPAM. Journal of Inequalities in Pure & Applied Mathematics 8(2007), no. 1, 554–563.
[63]
Tselil Schramm and Alexander S Wein, Computational barriers to estimation from low-degree polynomials, The Annals of Statistics 50(2022), no. 3, 1833–1858.
[64]
Garvesh Raskutti, Martin J Wainwright, and Bin Yu, Restricted eigenvalue properties for correlated gaussian designs, The Journal of Machine Learning Research 11(2010), 2241–2259.