June 04, 2026
In many settings one is often interested in determining whether two networks share some joint structural connectivity patterns such as communities. However, while communities may be shared across networks, edge probabilities may differ significantly. Therefore, in this paper we consider testing a general null hypothesis that two networks have the same underlying subspace, which in particular includes the setting that communities are the same for either stochastic blockmodels or mixed-membership stochastic blockmodels (even if edge probabilities are different). We propose a test statistic based on the Frobenius norm of the difference of the leading subspace projection matrices, and we prove that our test statistic, after appropriate centering and scaling, converges in distribution to a Gaussian random variable as long as the average expected degree grows at least logarithmically in the number of vertices. We then provide estimators for the asymptotic mean and variance and show consistency under a stronger signal condition, and we give the local power of our test when the networks are sufficiently dense. Our theoretical results are based on a limit theorem for the projection difference of empirical and true eigenvectors which can also be viewed as the one-sample version of our test statistic, and this result may be of independent interest. We demonstrate our results through numerical simulations and an application to US Flight data.
Keywords: Low-rank matrices, Subspace perturbation, Network data, Signal-plus-noise, Random matrix theory.
In the modern era, network-valued data has become central to scientific inquiry across diverse fields, such as social science[1], [2] or neuroscience[3], [4]. In these settings, the primary analytical goal is often to understand the underlying latent structure (such as communities) that dictates network behavior. A critical challenge in these settings is determining whether two networks share this structural geometry, even when their overall edge probabilities differ. For example, communities may remain the same between social networking platforms, but edge formations can differ due to the particular platform.
To address this, we propose a spectral two-sample test for the equality of underlying population-level subspaces, assuming only that the edge probability matrices are low-rank. In the case of networks with community structure such as stochastic blockmodels (2) or mixed-membership stochastic blockmodels (3), this hypothesis is equivalent to testing whether the communities are equal across networks. Our theoretical results rely on a novel limit theorem for the projection distance of empirical eigenvectors, a result that may be of independent interest.
The primary contributions of this work are threefold:
A Two-Sample Test for Common Subspaces: We propose a test statistic for the hypothesis that two networks arise from probability matrices with the same column space. We prove that, after appropriate centering and scaling, our test statistic converges in distribution to a standard Gaussian random variable. These limit theorems are valid over a wide parameter space regarding the network dimension \(n\) and the sparsity parameter \(\rho_n\). We also provide consistent plug-in estimators under a slightly stronger sparsity assumption.
Application to Synthetic and Real-World Data: We apply our test procedure to synthetic data, showing that the test exhibits strong power in distinguishing stochastic blockmodels with different community structures and is capable of detecting subtle shifts in mixed-membership models. Furthermore, we apply our methodology to United States domestic flight data. Our analysis reveals statistically significant shifts in the structural connectivity of the airport network corresponding to the onset of the COVID-19 pandemic.
A Limit Theorem for Projection Distances: Our main technical results for our test are based on a corresponding limit theorem for the projection distance between subspaces, which can be viewed as the “one-sample” analogue of our test. This limit theorem, with associated plug-in estimators, can be used for confidence intervals for the \(\sin \Theta\) distance between the observed and population subspaces. This result generalizes several previous results in the literature to Bernoulli noise, and may be of independent interest.
The problem of hypothesis testing for network data has received considerable attention in recent years. Our work resides at the intersection of two-sample inference for random graphs and high-dimensional spectral perturbation theory.
A foundational nonparametric framework for determining if two networks are generated from the same random dot product graph (RDPG) latent positions was established by [5], later generalized by [6]. Other nonparametric approaches include assumption-lean inference [7] and tests for arbitrary vertices [8], [9]. While these tests are robust, they generally test for the equality of the precise generating parameters rather than the underlying subspace up to rotation or scaling.
Another significant strand of literature has focused on moment-based, subsampling, and bootstrap approaches. [10], [11], [12], and [13] leverage subsampling techniques to construct test statistics, while others utilize spectral moments, graph cumulants, or U-statistics [14]–[19], or employ bootstrap methods [20], [21]. Alternatively, [22] and [23] propose tests based on matrix norms. In contrast to these strategies, our approach is based on direct spectral analysis. Instead of relying on empirical subsamples or higher-order moment matching, we establish the asymptotic normality of our test statistic directly via different technical matrix analysis tools. Our test statistic admits a closed-form limiting distribution, obviating the need for computationally intensive resampling or moment-computation procedures.
Related work has addressed hypothesis testing problems similar to ours, though primarily under the assumption of a Stochastic Blockmodel (SBM). For instance, [24] and [25] address change-point detection in dynamic networks, which can be viewed as a sequential extension of two-sample testing. However, these approaches typically monitor global deviations in model parameters via least-squares objectives or cumulative sums, rather than directly testing the stability of the invariant subspace. Consequently, our statistic is distinct in its specific focus on the distance between spectral projectors, ensuring robustness to nuisance fluctuations in edge density that might otherwise be identified as change points. Methodological differences are also pronounced in comparison to [26], who propose an extreme-value test based on the maximum standardized deviation of edges from estimated block probabilities. While their strategy is contingent on consistent community recovery, our statistic utilizes the Frobenius norm of the difference between spectral projectors. This \(L_2\)-type functional captures global discrepancies in the underlying subspace geometry, avoiding reliance on local fluctuations or the explicit alignment of community labels.
Beyond these specific formulations, recent literature has expanded network inference to address varied constraints. [27] and [28], [29] develop tests for scenarios involving unknown vertex correspondence, unequal graph sizes, or degree-corrected mixed-membership assumptions. Other complementary contributions focus on rank estimation [30], mesoscale comparisons [31], and the sensitivity of testing power to label misalignment [32]. Distinct from these specialized focuses, our work establishes a comprehensive asymptotic theory for the global projection distance itself. Our work is closely related to the semiparametric framework of [33], which motivates a similar spectral statistic but relies on permutation-based validity arguments. We, conversely, derive the exact Gaussian limiting distribution of the test statistic, enabling a direct asymptotic pivot for rigorous Type I error control.
The work most relevant to our work is [34], which develops a two-sample test specifically for community memberships in weighted SBMs. However, their approach is restricted to a specific class of blockmodels. In contrast, our subspace-based formulation is significantly more general: by testing the null hypothesis of invariant column spaces, our method remains valid for both SBMs and Mixed Membership Stochastic Blockmodels (MMSBMs) and is robust to nuisance variation in edge probabilities.
Our theoretical analysis relies on the behavior of empirical eigenvectors; see [35], [36], and [37] for overviews. Recent advances have established entrywise perturbation bounds under asymmetric noise [38], small eigengaps [39], [40], and incoherence [41]. [42] provides perturbation bounds results for a slightly different model that support our technical analysis.
While many results provide perturbative bounds, our inference tools require distributional theory. The works [43], [44], and [45] provide limit theorems for individual eigenvectors or eigenvalues. [46] extends this to Laplacian matrices, and [47] and [48] apply these results to testing membership equality of two individual vertices. [49] and [50] address small eigengaps and asymmetry.
In contrast to row-wise or “local” inference, our statistic depends on the global projection matrix. The primary theoretical antecedents are [51] and the closely related [52], which derive normal approximations for singular subspaces. We also note complementary results for \(\ell_{2,\infty}\) norms [53] and tensor methods [54]–[56]. [51] provides explicit representation formulas for empirical spectral projectors. However, that work assumes i.i.d. Gaussian noise. In contrast, our technical analysis is tailored to heteroskedastic Bernoulli noise, and we leverage our results to develop our two-sample test.
We write \(J_n\) for the \(n\times n\) all-ones matrix, \(I_n\) for the \(n\times n\) identity matrix, and \(\mathbf{1}_n\) for the all-ones vector of length \(n\). For a vector \(\boldsymbol{v} \in \mathbb{R}^n\), the operator \(\mathop{\mathrm{Diag}}(\boldsymbol{v})\) denotes the diagonal matrix with diagonal entries given by the components of \(\boldsymbol{v}\). For \(M \in \mathbb{R}^{n\times n}\), the operator \(\mathrm{diag}(M)\) returns the vector of diagonal entries of \(M\) and the operator \(\mathop{\mathrm{Diag}}(M)\) returns a diagonal matrix with only the diagonal entries of \(M\).
For a matrix \(M\), \(M_{i \cdot}\) represents the \(i\)-th row and \(M_{\cdot i}\) the \(i\)-th column of \(M\). The \(\ell_{2,\infty}\)-norm is defined as \[\|M\|_{2,\infty} = \max_{i} \|M_{i\cdot}\|.\] The Frobenius norm is denoted by \(\|M\|_F\), the spectral norm by \(\|M\|\), the infinity norm by \(\|M\|_{\infty}=\max_{i,j} |M_{ij}|\), and the \(\ell_0\)-norm by \(\|M\|_0\) (which denotes the number of non-zero elements in the matrix, and similarly for a vector \(v\), \(\|v\|_0\) denotes the number of non-zero elements in that vector). The column space of a matrix \(M\) is denoted by \(\mathop{\mathrm{col}}(M)\), the trace by \(\mathop{\mathrm{tr}}(M)\), and the least absolute eigenvalue by \(\lambda_{\min}(M)\). For two matrices \(M_1\) and \(M_2\), we denote their Frobenius inner product by \(\langle M_1, M_2 \rangle\), defined as \(\langle M_1, M_2 \rangle = \mathop{\mathrm{tr}}(M_1 M_2^{\top})\). Their Hadamard (entrywise) product is denoted by \(M_1 \circ M_2\). We write \(A \sim \mathop{\mathrm{Ber}}(P)\) to indicate that \[A_{ij} = \begin{cases} \mathrm{Bernoulli}(P_{ij}), & i \le j,\\ A_{ji}, & i > j, \end{cases}\] with the collection \(\{A_{ij} : i \le j\}\) being mutually independent.
We write \(f(n) \lesssim g(n)\) to indicate that there exists a sufficiently large constant \(C>0\) such that \(f(n) \le C\, g(n)\) for all sufficiently large \(n\). We write \(f(n) \ll g(n)\) if \(f(n)/g(n)\to 0\) as \(n\to\infty\). Similarly, \(f(n) = O(g(n))\) denotes \(f(n)\lesssim g(n)\), and \(f(n) = o(g(n))\) denotes \(f(n)\ll g(n)\). We use \(f(n) \asymp g(n)\) to mean both \(f(n) = O(g(n))\) and \(g(n) = O(f(n))\).
If \(X\) is a random variable, then \(X = O_p(f(n))\) means \(\mathbb{P}\left( X \lesssim f(n) \right) > 1 - O(n^{-c}) \quad \text{for some constant } c>1,\) and \(X=o_p(f(n))\) means \(\mathbb{P}\left( X \ll f(n) \right) > 1 - O(n^{-c}) \quad \text{for some constant } c>0.\) For a sequence of random variables \(\{X_n\}_{n=1}^{\infty}\), we say \(X_n \xrightarrow{D} X\) if \(X_n\) converges to \(X\) in distribution. Finally, we say that a random variable \(X\) satisfies \(X \asymp f(n)\) if for every \(\epsilon >0\) there exist constants \(a,b>0\), independent of \(n\), such that \(\mathbb{P}\left( a\,f(n) \le X \le b\,f(n) \right) \ge 1-\epsilon .\)
The remainder of this article is organized as follows. 2 formulates the two-sample hypothesis testing problem, defines the test statistic, and presents the main asymptotic normality result and associated estimators. 3 validates the method through numerical simulations, while [sec:airport] applies it to US flight network data. 5 provides the theoretical foundation by establishing a limit theorem for the projection distance between empirical and population eigenspaces. Finally, 6 concludes with a discussion, and proofs are deferred to the Appendices.
Suppose one observes two independent adjacency matrices \({A^{(1)}}\) and \({A^{(2)}}\), with \({A^{(i)}} \in \{0,1\}^{n\times n}\) assumed to be symmetric. We consider the case where \({A^{(i)}} \sim \mathop{\mathrm{Ber}}(P^{(i)})\) for \(i = 1,2\), and each \(P^{(i)}\) is a symmetric, low-rank probability matrix of dimension \(n \times n\) and rank \(k\). Let \(P^{(i)}\) have spectral decomposition \(P^{(i)} = {V^{(i)}} \Lambda^{(i)} {V^{(i)}}^{\top}\), where \({V^{(i)}} \in \mathbb{R}^{n \times k}\) contains the \(k\) leading orthonormal eigenvectors associated with the \(k\) eigenvalues of \(P^{(i)}\), collected in the diagonal matrix \(\Lambda^{(i)}\). For simplicity of analysis, we allow self-loops (i.e., \({A^{(i)}}\) need not be hollow), but the results do not materially change is self-loops are disallowed.
Formally, we consider the hypotheses \[\begin{align} \label{hypo2} H_0 : {V^{(1)}} {V^{(1)}}^{\top}= {V^{(2)}} {V^{(2)}}^{\top} \quad \text{versus} \quad H_1 : {V^{(1)}} {V^{(1)}}^{\top}\neq {V^{(2)}} {V^{(2)}}^{\top}. \end{align}\tag{1}\] In words, this hypothesis test seeks to determine whether the subspaces spanned by the columns of \({V^{(1)}}\) and \({V^{(2)}}\) coincide, which is equivalent to 1 . To contextualize our hypothesis test within the framework of block models and latent position models, we provide formal definitions for the generalized random dot product graph, the stochastic blockmodel, and the mixed-membership stochastic blockmodel.
Definition 1 (Generalized Random Dot Product Graph (GRDPG)). Let \(n\in\mathbb{N}\) be the number of vertices and \(k\ge 1\) be the latent dimension. Let \(I_{p,q} \in \mathbb{R}^{k \times k}\) be a diagonal matrix with \(p\) ones and \(q\) minus ones on the diagonal, where \(p+q=k\). Let \(\Gamma\in\mathbb{R}^{n\times k}\) denote the matrix of latent positions. Suppose the rows of \(\Gamma\) satisfy the constraint that for all \(i,j \in \{1,\dots,n\}\), \(0 \le ( \Gamma I_{p,q} \Gamma^{\top})_{ij} \le 1\). We say \(A\) is an instantiation of a generalized random dot product graph, denoted \(A \sim \text{GRDPG}_{p,q}(\Gamma)\), if \(A \sim \mathop{\mathrm{Ber}}(P)\) with \(P = \Gamma I_{p,q} \Gamma^{\top}\).
Definition 2 (Stochastic Blockmodel (SBM)). Let \(n\in\mathbb{N}\) be the number of vertices and let \(k\ge 1\) be the number of communities. Let \(B\in[0,1]^{k\times k}\) denote the symmetric connectivity matrix. Suppose each vertex \(i\) belongs to some community \(C(i) \in \{1, \dots, k\}\). Define the matrix \(Z \in \{0,1\}^{n \times k}\) via \(Z_{ik} = 1\) if vertex \(i\) belongs to community \(k\), and \(Z_{ik} = 0\) otherwise. We say \(A\) is an instantiation of a stochastic blockmodel if \(A \sim \mathop{\mathrm{Ber}}( P)\) with \(P=Z B Z^{\top}.\)
Definition 3 (Mixed membership stochastic blockmodel (MMSBM)). Let \(n\in\mathbb{N}\) be the number of vertices and \(k\ge 1\) be the number of communities. Let \(B\in[0,1]^{k\times k}\) be the symmetric connectivity matrix. For each vertex \(i\in\{1,\dots,n\}\), let \[\begin{align} \label{simplex} Z_{i\cdot}=(Z_{i1},\dots,Z_{ik})\in\Delta^{k-1}:=\bigg\{x\in[0,1]^k:\sum_{j=1}^k x_j=1\bigg\} \end{align}\tag{2}\] be a membership vector representing the fractional affiliation of vertex \(i\) to each of the \(k\) communities. We say \(A\) is an instantiation of a mixed membership stochastic blockmodel if \(A \sim \mathop{\mathrm{Ber}}( P)\) with \(P=Z B Z^{\top}.\)
All the models described above satisfy \(\mathbb{E} \left(A\right) = P\), where \(P\) has rank at most \(k\). The following propositions examine how the general null hypothesis 1 applies to specific model cases. We begin with the GRDPG, where the test admits a geometric interpretation regarding the column space of the latent positions.
Proposition 1 (Equivalence of Latent Subspaces in GRDPG). Let \(A^{(1)}\) and \(A^{(2)}\) be independent instantiations of generalized random dot product graphs with latent position matrices \(\Gamma^{(1)}, \Gamma^{(2)} \in \mathbb{R}^{n \times k}\). Assume \(\Gamma^{(1)}\) and \(\Gamma^{(2)}\) have full column rank \(k\). Let the associated diagonal matrices be \(I_{p_1,q_1}\) and \(I_{p_2,q_2}\), where \(p_i+q_i = k\), such that \(P^{(i)} = \Gamma^{(i)} I_{p_i,q_i} {\Gamma^{(i)}}^{\top}\) for \(i=1,2\). Let \(V^{(i)} \in \mathbb{R}^{n \times k}\) be the matrix of \(k\) orthonormal eigenvectors corresponding to the non-zero eigenvalues of \(P^{(i)}\). Then \[V^{(1)} {V^{(1)}}^{\top}= V^{(2)} {V^{(2)}}^{\top}\] if and only if there exists a non-singular matrix \(W \in \mathbb{R}^{k \times k}\) such that \(\Gamma^{(1)} = \Gamma^{(2)} W.\)
Proof. See 15.1. ◻
In essence, 1 states that for GRDPGs, the hypothesis test determines whether the latent positions of the two graphs span the same feature space. The test is invariant to the linear transformation \(W\), meaning it detects fundamental changes in the geometric configuration of the latent positions rather than simple rotations or scaling.
When we restrict our attention to community based models, this geometric equivalence forces a stricter structural correspondence. For two matrices \(M_1,M_2\), we write \(M_1 \mathrel{ \vcenter{ \offinterlineskip \halign{##\cr to 1.2ex{\hfil\rule{0.5pt}{1.1ex}\hfil}\cr \noalign{\vskip 0.2ex} \rule{1.2ex}{0.4pt}\cr \noalign{\vskip 0.3ex} \rule{1.2ex}{0.4pt}\cr } } }M_2\) if there exists a permutation matrix \(\Pi\) such that \(M_1 = M_2\Pi\).
Proposition 2. Let \(P^{(i)}\) be one of the blockmodels with a full rank symmetric connectivity matrix as defined in 3 2. Let \({V^{(i)}}\in\mathbb{R}^{n\times k}\) be an orthonormal matrix whose columns span the column space of \(P^{(i)}\) (the nonzero population eigenvectors). We further assume that there is a pure node for each community for both \(Z^{(1)}\) and \(Z^{(2)}\). Then \[{V^{(1)}}{V^{(1)}}^{\top}= {V^{(2)}}{V^{(2)}}^{\top}\] if and only if \(Z^{(1)} \mathrel{ \vcenter{ \offinterlineskip \halign{##\cr to 1.2ex{\hfil\rule{0.5pt}{1.1ex}\hfil}\cr \noalign{\vskip 0.2ex} \rule{1.2ex}{0.4pt}\cr \noalign{\vskip 0.3ex} \rule{1.2ex}{0.4pt}\cr } } }Z^{(2)}\).
Proof. See 15.2. ◻
Thus, it follows from 2 that 1 is equivalent to \[\begin{align} \label{hypo1} H_0 : Z^{(1)} \mathrel{ \vcenter{ \offinterlineskip \halign{##\crto 1.2ex{\hfil\rule{0.5pt}{1.1ex}\hfil}\cr \noalign{\vskip 0.2ex} \rule{1.2ex}{0.4pt}\cr \noalign{\vskip 0.3ex} \rule{1.2ex}{0.4pt}\cr } } }Z^{(2)} \quad \text{versus} \quad H_1 : Z^{(1)} \mathrel{ \ooalign{\mathrel{ \vcenter{ \offinterlineskip \halign{##\crto 1.2ex{\hfil\rule{0.5pt}{1.1ex}\hfil}\cr \noalign{\vskip 0.2ex} \rule{1.2ex}{0.4pt}\cr \noalign{\vskip 0.3ex} \rule{1.2ex}{0.4pt}\cr } } }\cr \hidewidth\mkern 2mu/\mkern 2mu\hidewidth\cr } }Z^{(2)} \end{align}\tag{3}\] under both the SBM and MMSBM. Testing \({V^{(1)}}{V^{(1)}}^{\top}= {V^{(2)}}{V^{(2)}}^{\top}\) is, in effect, a test of whether the two networks encode the same underlying community structure even if the probability matrices \(P_1\) and \(P_2\) are different. The airport network analysis in [sec:airport] offers a concrete illustration: our procedure successfully distinguishes between periods of stable community structure and periods of severe structural disruption, such as the COVID-19 pandemic. Full empirical results and their interpretation are presented in Section [sec:airport]. Hence, a test of 1 furnishes a principled and interpretable mechanism for detecting structural changes in community organization while remaining robust to transient edge-level fluctuations.
Before proposing our test statistics and studying their asymptotics, we impose several mild regularity conditions on the parameter space. For notational simplicity, we state the assumptions using a single adjacency matrix \(A\) and its population counterpart \(P=\mathbb{E}(A)\). We assume that \(P\) admits the spectral decomposition \(P = V \Lambda V^\top\), where \(V \in \mathbb{R}^{n \times k}\) is the matrix of eigenvectors and \(\Lambda\) is the diagonal matrix of eigenvalues with the \(i\)-th eigenvalue given by \(\lambda_i\).
Our first assumption concerns the asymptotic regime and sparsity of each network.
Assumption 1 (Asymptotic regime). The edge probabilities are uniformly of order \(\rho_n\); that is, there exist constants \(0<a<b<\infty\) such that \(a\,\rho_n \le P_{ij} \le b\,\rho_n\) for all \(1\le i,j\le n.\) The number of communities \(k\) remains fixed as \(n\to\infty\), and the sparsity parameter satisfies \(\rho_n\to 0\).
Assumption 2 (Sparse regime). The sparsity parameter satisfies \(n\rho_n \gg \log n\) for sufficiently large \(n\).
It is well-known that if \(n\rho_n \ll \log n\), then the network is disconnected with high probability, and the Frobenius distance between the empirical and population subspace projection matrices fails to converge to zero [36], [57]. Thus, our assumption that \(n\rho_n \gg \log n\) is relatively mild, as we focus on inference, not estimation. For simplicity we assume that \(\rho_n\) is of the same order for both networks.
Our next assumption places conditions on the leading \(k\) eigenvalues of each matrix.
Assumption 3 (Eigenvalue scaling). There exist positive constants \(C_i,D_i\) (independent of \(n\)) such that, for each \(i=1,\dots,k\), \(C_i < \frac{|\lambda_i|}{n\rho_n} < D_i\).
This assumption imposes only a mild regularity condition on the eigenvalues of the probability matrix \(P\), and is satisfied in virtually all of the canonical instances of the blockmodel family. Its validity is formalized in the proposition below. We first include the following definition.
Definition 4 (Balanced Community Sizes). We say that the community sizes are balanced if the following conditions hold:
For SBM: Let \(n_l\) denote the size of community \(l\), such that \(n_l = n\pi_l\), where proportions \(\pi_l\) satisfy \[\begin{align} C_{1} \le \min_{l}\pi_{l} \le \max_{l}\pi_{l} \le C_{2}, \end{align}\] for some fixed constants \(C_{1}, C_{2} > 0\).
For MMSB: The membership matrix \(Z\) satisfies \[\lambda_{\min}(Z^\top Z) \ge c\,\frac{n}{k},\] for some constant \(c>0\).
Proposition 3. Let \(P\) be one of the blockmodels as defined in 3 2. Define \(\widetilde{B} = \frac{1}{\rho_n} B\) where we assume that the eigenvalues of \(\widetilde{B}\) are bounded; specifically, there exist constants \(0 < a_1 < a_2 < \infty\) such that \(a_1 < |\lambda_{\min}(\widetilde{B})| \le |\lambda_{\max}(\widetilde{B})| < a_2.\) Furthermore, assume that the community sizes are balanced in the sense of 4 and that Assumptions 1 and 2 are satisfied. Then the \(k\) nonzero eigenvalues \(\lambda_{1},\dots,\lambda_{k}\) of \(P\) satisfy 3.
Proof. See 15.3. ◻
Finally, the following assumption imposes the condition that the rows of \(V\) (denoted by \(V_{1 \cdot},\dots,V_{n \cdot}\)) are uniformly spread, a concept known as incoherence in the literature. In the particular case of the SBM or MMSBM, this assumption follows if we assume that the community sizes are balanced as we will demonstrate in 20 21.
Assumption 4 (Incoherence). The matrix of population eigenvectors \(V\) satisfies the following incoherence condition: there exist constants \(\psi_1>0\) and \(\psi_2>0\) (independent of \(n\)) such that for all \(j=1, \dots,n,\) \[\sqrt{\frac{\psi_1 k}{n}} \le \|V_{j\cdot}\|_{2} \le \sqrt{\frac{\psi_2 k}{n}}.\] In addition, \(|(VV^{\top})_{ij}|\asymp \frac{k}{n}\) for all \(i,j=1, \cdots ,n\)
Let \(\widehat V^{(i)}\in\mathbb{R}^{n\times k}\) denote the matrix of leading \(k\) eigenvectors of \({A^{(i)}}\) with \(\widehat{\Lambda}^{(i)}\) the corresponding eigenvalue matrix. We consider the test statistic \[\begin{align} \label{2sampleTS} T_n := \bigl\|(\widehat V^{(1)})(\widehat V^{(1)})^{\top}- (\widehat V^{(2)})(\widehat V^{(2)})^{\top}\bigr\|_F^2. \end{align}\tag{4}\] The following theorem characterizes its null distribution after appropriate centering and scaling. The full proof can be found in 8. Here, \(\widetilde{\mu}_2\) and \(\widetilde{\sigma}_2\) denote the sample mean and sample standard deviation of 4 , respectively, with the subscript indicating the two-sample setting. Subsequently, we will consider a one-sample analogue of the following result; see 5. In that case, \(\mu_1\) and \(\sigma_1\) will denote the corresponding one-sample mean and standard deviation.
Theorem 1. Suppose 1 2 3 4 hold. Then, under the null as defined in 1 , \[\label{eq:clt2} \frac{T_n - \widetilde{\mu}_{2}}{\widetilde{\sigma}_2} \xrightarrow{D} \mathcal{N}(0,1),\tag{5}\] where the asymptotic centering and scaling terms satisfy \[\begin{align} \widetilde{\mu}_{2} &= \underbrace{2\sum_{i=1}^{2}\,\operatorname{tr}\biggl[\beta_i^{-2}\Bigl({\Sigma^{(i)}}\circ\beta_i^\perp + \operatorname{Diag}\bigl({\Sigma^{(i)}}\cdot d_i - \operatorname{diag}({\Sigma^{(i)}}) \circ d_i\bigr)\Bigr)\biggr]}_{\mu_2} + O_p\left(\frac{k}{n^2 \rho_n^2}\right), \tag{6} \\[1em] \widetilde{\sigma}_2^2 &= \sigma_2^2 + o\left(\frac{k^2}{n^3 \rho_n^2}\right) \nonumber \\[0.5em] &= \begin{aligned}[t] &\sum_{i=1}^{2} \biggl( 8 \Bigl\langle (J_n - I_n)\circ(\beta_i^{-2}\circ\beta_i^{-2}),({\Sigma^{(i)}})^2\Bigr\rangle + 4 \,\alpha_i^\top K^{(i)}\,\alpha_i \biggr) \\ &+ 4 \Bigl\langle G \circ G,{\Sigma^{(1)}}^\top{\Sigma^{(2)}} + \left(J_n-I_n\right) \circ \left({\Sigma^{(1)}} \circ {\Sigma^{(2)}}\right)\Bigr\rangle + o\left(\frac{k^2}{n^3 \rho_n^2}\right). \end{aligned} \tag{7} \end{align}\] Here we define \[\begin{align} \label{last00} &\beta_i^k = {V^{(i)}}({\Lambda^{(i)}})^k {V^{(i)}}^{\top}, \quad \beta_i^\perp = I_n - {V^{(i)}} {V^{(i)}}^{\top}, \quad d_i = \mathrm{diag}(\beta_i^\perp),\notag \\ &G=\beta_1^{-1}\beta_2^{-1} + \beta_2^{-1}\beta_1^{-1}, \quad {\Sigma^{(i)}} = P^{(i)}\circ(J_n-P^{(i)}), \quad K^{(i)} = 2\,{\Sigma^{(i)}}\circ(J_n - 2P^{(i)}), \quad \alpha_i = \mathrm{diag}(\beta_i^{-2}). \end{align}\tag{8}\]
Proof. See 7. ◻
The residual contributions \(O_p\bigl(\tfrac{k}{n^{2}\rho_n^{2}}\bigr)\) and \(o\bigl(\tfrac{k^{2}}{n^{3}\rho_n^{2}}\bigr)\) appearing in 6 and 7 arise from the accumulation of all higher–order terms in the expansion underlying our statistic. These terms remain controlled under the asymptotic regime of 1 2, but they do not vanish automatically when \(n\rho_n \gg \log n\). In particular, when the graph is moderately sparse, their magnitude is asymptotically negligible relative to the leading terms, but can still influence the centering and scaling terms. If, however, we strengthen the sparsity condition to a denser regime, then all such residual higher–order terms become negligible. This phenomenon is rigorously verified in 2, and motivates the additional assumption stated below.
Assumption 5 (Dense Regime). The sparsity parameter satisfies \[n\rho_n \gg \sqrt n .\]
Theorem 2. Suppose 1 5 3 4 hold. Then, under the null as defined in 1 , \[\label{eq:clt295dense} \frac{T_n - \mu_{2}}{ \sigma_2} \xrightarrow{D} \mathcal{N}(0,1),\tag{9}\] where the asymptotic centering and scaling terms are given by \[\begin{align} \mu_{2} &= 2\sum_{i=1}^{2}\, \mathop{\mathrm{tr}}\Bigl[\beta_i^{-2}\bigl({\Sigma^{(i)}}\circ\beta_i^\perp + \mathop{\mathrm{Diag}}\bigl({\Sigma^{(i)}}\cdot d_i - \mathrm{diag}({\Sigma^{(i)}})\circ d_i\bigr)\bigr)\Bigr],\\[4pt] \sigma_2^2 &= \sum_{i=1}^{2} \Bigl( 8\left\langle (J_n-I_n)\circ(\beta_i^{-2}\circ\beta_i^{-2}),{\Sigma^{(i)}}^{2}\right\rangle + 4\,\alpha_i^{\top}K^{(i)} \alpha_i \Bigr) \\ &\qquad\qquad + 4\Bigl\langle G\circ G, {\Sigma^{(1)}}^{\top}{\Sigma^{(2)}} + (J_n-I_n)\circ({\Sigma^{(1)}}\circ {\Sigma^{(2)}}) \Bigr\rangle , \end{align}\] where all notation coincides with that of 1.
Proof. See 15.4. ◻
We emphasize that the validity of our two–sample limit theorem does not require 5; 1 already establishes the asymptotic normality under the weaker sparsity conditions in 2. The role of 5 is instead methodological: in the dense regime, the dominant components in 6 and 7 can be estimated with sufficient accuracy so that the resulting data–driven test remains consistent. This observation motivates the construction of the estimators in the next section, which are designed specifically to recover the leading terms appearing in the asymptotic centering and scaling expressions.
To carry out valid inference using our proposed two-sample test, we require consistent, data-driven estimators of the dominant mean and variance terms \(\mu_{2}\) and \(\sigma_{2}\) appearing in the Gaussian limits established in 2. For the plug-in procedure to remain valid, the estimators \(\widehat{\mu}_2\) and \(\widehat{\sigma}_2\) must satisfy: \[\begin{align} \label{rates} \frac{\widehat{\mu}_2 - \mu_2}{\widehat{\sigma}_2} \xrightarrow{p} 0, \quad \text{and} \quad \frac{\widehat{\sigma}_2}{\sigma_2} \xrightarrow{p} 1. \end{align}\tag{10}\] These requirements ensure that the plug-in centering and scaling do not distort the limiting distribution, thereby yielding a fully data-driven inference procedure whose Type-I error and power remain asymptotically valid.
Below we describe a principled plug-in procedure that satisfies 10 . The parameters \(\mu_2\) and \(\sigma_2\) are functions of the population eigen-structure \(({V^{(i)}},\Lambda^{(i)})\) for \(i=1,2\). Since \(({V^{(i)}},\Lambda^{(i)})\) are not observed, we estimate them from the data; for the eigenvectors we use the empirical eigenvectors \(\widehat V^{(i)}\), while for the eigenvalues we use the empirical eigenvalues.
Motivated by the explicit expressions for \(\mu_{2}\) and \(\sigma_{2}\) in 2, we next define the following data-driven estimators by plug-in: \[\label{eq:estimators} \begin{align} \widehat{\mu}_2 &= 2\sum_{i=1}^{2}\,\mathop{\mathrm{tr}}\Bigl[\widehat{\beta}_i^{-2}\Bigl((\widehat{\Sigma}^{(i)})\circ \widehat\beta_i^\perp + \mathop{\mathrm{Diag}}\bigl((\widehat{\Sigma}^{(i)})\cdot \widehat d_i \bigr)\Bigr)\Bigr], \\[4pt] \widehat{\sigma}^2_2 &= \sum_{i=1}^{2} \left(8 \bigl\langle (J_n - I_n)\circ(\widehat{\beta}_i^{-2}\circ\widehat{\beta}_i^{-2}),(\widehat{\Sigma}^{(i)})^2\bigr\rangle + 4 \, \widehat{\alpha}_i^{\top}{\widehat{K}^{(i)}}\,\widehat{\alpha}_i \right)\\[3pt] &\qquad\qquad+ 4 \bigl\langle \widehat{G}\circ\widehat{G}, (\widehat{\Sigma}^{(1)})^{\top}(\widehat{\Sigma}^{(2)}) + (J_n-I_n)\circ((\widehat{\Sigma}^{(1)})\circ(\widehat{\Sigma}^{(2)}))\bigr\rangle, \end{align}\tag{11}\] where, for \(i=1,2\), \[\begin{align}\label{thm3460} {\widehat{P}^{(i)}} = (\widehat V^{(i)}) (\widehat{\Lambda}^{(i)}) (\widehat V^{(i)})^{\top}, \qquad \widehat{\beta}_i^k = (\widehat V^{(i)}) (\widehat{\Lambda}^{(i)})^k (\widehat V^{(i)})^{\top}, \qquad \widehat\beta_i^\perp = I_n - (\widehat V^{(i)})(\widehat V^{(i)})^{\top}, \qquad \widehat d_i = \mathrm{diag}(\widehat\beta_i^\perp),\\ \widehat{\Sigma}^{(i)} = {\widehat{P}^{(i)}}\circ(J_n-{\widehat{P}^{(i)}}), \qquad {\widehat{K}^{(i)}} = 2\,(\widehat{\Sigma}^{(i)})\circ(J_n - 2{\widehat{P}^{(i)}}), \qquad \widehat{\alpha}_i = \mathrm{diag}(\widehat{\beta}_i^{-2}), \quad \widehat{G} = \widehat{\beta}_1^{-1}\widehat{\beta}_2^{-1} + \widehat{\beta}_2^{-1}\widehat{\beta}_1^{-1}. \end{align}\tag{12}\] 3 establishes the consistency properties of the estimators \(\widehat{\mu}_{2}\) and \(\widehat{\sigma}_{2}\).
Theorem 3 (Mean and Variance consistency). Let \(\mu_2\) and \(\sigma_2^2\) denote the theoretical mean and variance of the two-sample test statistic as defined in 1. Define the plug-in mean estimator \(\widehat\mu_2\) and variance estimator \(\widehat\sigma_2^2\) as in 11 . Then under 1 5 3 4, \[|\widehat{\mu}_2 - \mu_2| =O_p \left(\max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right)\right), \qquad |\widehat{\sigma}_2^2 - \sigma_2^2| = o_p \left(\frac{k^2}{n^3 \rho_n^2}\right).\] In particular, \(\frac{\widehat{\mu}_2 - \mu_2}{\widehat{\sigma}_2} \xrightarrow{p} 0,\) and \(\frac{\widehat{\sigma}_2}{\sigma_2} \xrightarrow{p} 1\) as required.
Proof. See 9. ◻
Remark 1 (Proof Strategy). The proof of 3 proceeds via a systematic decomposition: (i) expanding the plug-in estimation error into sums involving deviations of empirical eigenvectors and eigenvalues from their population counterparts; (ii) controlling the empirical eigenvector and eigenvalue fluctuations; and (iii) verifying that all remainder terms are asymptotically negligible in the dense regime. While this high-level roadmap is standard, the novelty of our analysis lies in the intricate handling of the higher-order terms required by the Bernoulli noise model. Unlike homoskedastic Gaussian settings where perturbation terms are often amenable to simpler bounds, the adjacency matrix entails heteroskedasticity and discreteness. Our analysis necessitates a delicate higher-order expansion to ensure that the complex variance structure of the Bernoulli entries does not invalidate the consistency rates. Crucially, despite these theoretical challenges, the resulting estimators remain algorithmically attractive: they are closed-form functions of the observed empirical eigen-structure, permitting valid inference without requiring complex auxiliary optimization or regularization.
Our full testing procedure is summarized in 1.
Under the null hypothesis \(H_0: {V^{(1)}} {V^{(1)}}^{\top}= {V^{(2)}} {V^{(2)}}^{\top}\), 2 guarantees that the test statistic is asymptotically standard normal, and the estimator consistency results in the previous section guarantee asymptotic Type I error control at any desired level \(\alpha\). Therefore, in this section we consider the consistency of our full testing procedure under general alternatives and the important case of blockmodel families.
We study the local alternative hypothesis \(H_1:{V^{(1)}} {V^{(1)}}^{\top}\neq {V^{(2)}} {V^{(2)}}^{\top}\). The following result establishes the power of our test under local alternatives.
Theorem 4. Suppose 1 3 4 5 hold. Let the two–sample test statistic be as in 4 and consider the testing problem stated in 1 . If the alternative satisfies the condition \[\label{eq:lowerbound95Frobenius} \|{V^{(1)}}{V^{(1)}}^{\top}- {V^{(2)}}{V^{(2)}}^{\top}\|_F^2 \gg \frac{k}{n^{3/2}\,\rho_n},\tag{13}\] then the test is consistent in the sense that the power converges to one as \(n\to\infty\).
Proof. See 10. ◻
We now proceed to investigate what this separation condition entails in the context of specific blockmodel families. Our first result establishes the local power for SBM as in 2.
Corollary 1. Let \({A^{(i)}}\), \(i = 1, 2\), denote the adjacency matrices of two SBMs with \(k\) communities and balanced community sizes satisfying 4. Suppose that \(P^{(i)} = {Z^{(i)}} B^{(i)} {Z^{(i)}}^{\top}\) for \(i = 1, 2\), where \({Z^{(i)}} \in \{0,1\}^{n \times k}\) is the community membership matrix, and \(B^{(i)}\) is the corresponding block matrix under model \(i\). Let \({Z_{\cdot l}^{(i)}} \in \{0,1\}^n\) be as in 2. Then, under 1 5, the proposed two-sample test statistic in 4 has power converging to one for testing \[H_0 : Z^{(1)} \mathrel{ \vcenter{ \offinterlineskip \halign{##\crto 1.2ex{\hfil\rule{0.5pt}{1.1ex}\hfil}\cr \noalign{\vskip 0.2ex} \rule{1.2ex}{0.4pt}\cr \noalign{\vskip 0.3ex} \rule{1.2ex}{0.4pt}\cr } } }Z^{(2)} \quad \text{versus} \quad H_1 : Z^{(1)} \mathrel{ \ooalign{\mathrel{ \vcenter{ \offinterlineskip \halign{##\crto 1.2ex{\hfil\rule{0.5pt}{1.1ex}\hfil}\cr \noalign{\vskip 0.2ex} \rule{1.2ex}{0.4pt}\cr \noalign{\vskip 0.3ex} \rule{1.2ex}{0.4pt}\cr } } }\cr \hidewidth\mkern 2mu/\mkern 2mu\hidewidth\cr } }Z^{(2)},\] provided that the following condition holds: \(\quad \max_{1 \leq l \leq k} \sum_{j=1}^{n} \left| Z_{jl}^{(1)} - Z_{jl}^{(2)} \right| \gg \frac{k}{n^{1/2}\rho_n}.\)
Proof. See 10. ◻
The above corollary reveals that even if we swap the community memberships of as few as \(n_0 = k\) vertices, the null hypothesis \(H_0: Z^{(1)} \mathrel{ \vcenter{ \offinterlineskip \halign{##\cr to 1.2ex{\hfil\rule{0.5pt}{1.1ex}\hfil}\cr \noalign{\vskip 0.2ex} \rule{1.2ex}{0.4pt}\cr \noalign{\vskip 0.3ex} \rule{1.2ex}{0.4pt}\cr } } }Z^{(2)}\) will be rejected with high probability, indicating strong power of the proposed test. This highlights the sensitivity of the test in detecting small but structured perturbations in community assignments.
We will now consider MMSBMs.
Corollary 2. Let \({A^{(i)}}\) denote the adjacency matrices of two MMSBMs for \(i=1,2.\) Suppose that the membership matrices satisfy \[\begin{align} \lambda_{\min}({Z^{(i)}}^{\top}{Z^{(i)}}) \ge c\,\frac{n}{k}, \end{align}\] where \(c\) is some positive constant. Suppose that \(P^{(i)} = {Z^{(i)}} B^{(i)} {Z^{(i)}}^{\top}\) for \(i = 1, 2\), where \({Z^{(i)}} \in [0,1]^{n \times k}\) is the mixed-membership matrix as in 3, and \(B^{(i)}\) is the corresponding block matrix under model \(i\). Then, under 1 5, the proposed two-sample test statistic in 4 has power converging to one for testing \[H_0 : Z^{(1)} \mathrel{ \vcenter{ \offinterlineskip \halign{##\crto 1.2ex{\hfil\rule{0.5pt}{1.1ex}\hfil}\cr \noalign{\vskip 0.2ex} \rule{1.2ex}{0.4pt}\cr \noalign{\vskip 0.3ex} \rule{1.2ex}{0.4pt}\cr } } }Z^{(2)} \quad \text{versus} \quad H_1 : Z^{(1)} \mathrel{ \ooalign{\mathrel{ \vcenter{ \offinterlineskip \halign{##\crto 1.2ex{\hfil\rule{0.5pt}{1.1ex}\hfil}\cr \noalign{\vskip 0.2ex} \rule{1.2ex}{0.4pt}\cr \noalign{\vskip 0.3ex} \rule{1.2ex}{0.4pt}\cr } } }\cr \hidewidth\mkern 2mu/\mkern 2mu\hidewidth\cr } }Z^{(2)},\] provided that the following condition holds: \(\quad \|Z^{(1)} - Z^{(2)}\|_F^2 \gg \frac{k}{n^{1/2}\rho_n}.\)
Proof. See 10. ◻
The corollary states that, in the MMSBM setting, detectability of a difference between two network models is governed by the squared Frobenius distance between their membership matrices, with the scaling determined by the sample size \(n\), the rank \(k\), and the sparsity parameter \(\rho_n\). This condition naturally extends the SBM result to the soft–assignment regime: larger deviations of the entire membership profile (measured in Frobenius norm) are required for reliable discrimination when the networks are sparser.
In this section we consider simulations under both the SBM and the MMSBM.
First, we apply 1 to SBMs. For each specification of \(k\) (the number of communities), we construct two block matrices \(B^{(1)}\) and \(B^{(2)}\). The diagonal entries of these matrices are independently sampled from \([0.8,1]\), while the off-diagonal entries of \(B^{(1)}\) are sampled from \([0.3,0.5]\) and those of \(B^{(2)}\) are sampled from \([0.1,0.3]\).
For \(k=3\), the two block matrices (rounded to two decimal places) are given below:
\[B^{(1)} = \begin{bmatrix} 0.98 & 0.36 & 0.46 \\ 0.36 & 0.99 & 0.38 \\ 0.46 & 0.38 & 0.81 \end{bmatrix}, \qquad B^{(2)} = \begin{bmatrix} 0.97 & 0.12 & 0.14 \\ 0.12 & 0.96 & 0.25 \\ 0.14 & 0.25 & 0.87 \end{bmatrix}.\] For \(k=5\), the two block matrices (rounded to two decimal places) are given below: \[B^{(1)} = \begin{bmatrix} 0.99 & 0.36 & 0.46 & 0.48 & 0.41 \\ 0.36 & 0.89 & 0.38 & 0.49 & 0.48 \\ 0.46 & 0.38 & 0.94 & 0.31 & 0.41 \\ 0.48 & 0.49 & 0.31 & 0.91 & 0.39 \\ 0.41 & 0.48 & 0.41 & 0.39 & 0.82 \end{bmatrix}, \qquad B^{(2)} = \begin{bmatrix} 0.87 & 0.12 & 0.14 & 0.27 & 0.12 \\ 0.12 & 0.84 & 0.25 & 0.26 & 0.16 \\ 0.14 & 0.25 & 0.95 & 0.17 & 0.15 \\ 0.27 & 0.26 & 0.17 & 0.96 & 0.18 \\ 0.12 & 0.16 & 0.15 & 0.18 & 0.92 \end{bmatrix}.\]
To construct the community assignment for \(Z^{(1)}\), we allocate the \(n\) nodes uniformly across the \(k\) communities, such that each community contains exactly \(n/k\) nodes. The subsequent assignment \(Z^{(2)}\) is generated by initializing \(Z^{(2)} = Z^{(1)}\), selecting \(n_0\) nodes uniformly at random, and reassigning their community labels to a different class, while leaving the memberships of the remaining \(n - n_0\) nodes unchanged. We then form the two probability matrices as \(P^{(i)} = {Z^{(i)}} B^{(i)} {Z^{(i)}}^{\top}\), and generate the observed adjacency matrices according to \({A^{(i)}} \sim \text{Ber}(P^{(i)})\). For each fixed combination of \(n,k,\rho_n\), and \(n_0\), we simulate the data set 100 times and apply the test at \(\alpha=0.05\) level of significance, recording the number of rejections. The empirical power of the test at those parameters is then estimated by the fraction of rejections. 2 4 report these empirical powers, with rows corresponding to different values of \(\rho_n\) and columns corresponding to different values of \(n_0\). The first row of each table (\(n_0 = 0\)) reflects the performance of the proposed test under the null hypothesis in 1 . While finite-sample empirical sizes may exhibit slight deviations from the exact nominal level \(\alpha = 0.05\), they remain closely aligned. Furthermore, as anticipated by our theoretical analysis, this size calibration steadily improves, converging toward 0.05, as the density parameter \(\rho_n\) increases toward 1 for a fixed \(k\). Conversely, under the alternative hypothesis (\(n_0 > 0\)), the empirical power exhibits a sharp transition, converging rapidly to \(1\) as the perturbation size \(n_0\) increases down the rows, demonstrating the high sensitivity of the procedure. While we do not keep track of the dependence on \(k\) in our theoretical results, we see that larger \(k\) has weaker detection ability.
0.49
3.5pt
| \(n_0\) | \(\rho_n=0.4\) | \(\rho_n=0.5\) | \(\rho_n=0.6\) | \(\rho_n=0.7\) | \(\rho_n=0.8\) | \(\rho=0.9\) |
|---|---|---|---|---|---|---|
| 0 | 0.05 | 0.04 | 0.06 | 0.05 | 0.05 | 0.06 |
| 2 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 4 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 6 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 8 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 10 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
0.49
3.5pt
| \(n_0\) | \(\rho_n=0.4\) | \(\rho_n=0.5\) | \(\rho_n=0.6\) | \(\rho_n=0.7\) | \(\rho_n=0.8\) | \(\rho_n=0.9\) |
|---|---|---|---|---|---|---|
| 0 | 0.03 | 0.03 | 0.05 | 0.08 | 0.03 | 0.04 |
| 2 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 4 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 6 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 8 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 10 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
0.49
3.5pt
| \(n_0\) | \(\rho_n=0.4\) | \(\rho_n=0.5\) | \(\rho_n=0.6\) | \(\rho_n=0.7\) | \(\rho_n=0.8\) | \(\rho_n=0.9\) |
|---|---|---|---|---|---|---|
| 0 | 0.08 | 0.13 | 0.13 | 0.12 | 0.12 | 0.06 |
| 2 | 0.36 | 0.53 | 0.77 | 0.93 | 1.00 | 1.00 |
| 4 | 0.90 | 0.98 | 1.00 | 1.00 | 1.00 | 1.00 |
| 6 | 0.99 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 8 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 10 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
0.49
3.5pt
| \(n_0\) | \(\rho_n=0.4\) | \(\rho_n=0.5\) | \(\rho_n=0.6\) | \(\rho_n=0.7\) | \(\rho_n=0.8\) | \(\rho_n=0.9\) |
|---|---|---|---|---|---|---|
| 0 | 0.08 | 0.06 | 0.05 | 0.08 | 0.12 | 0.07 |
| 2 | 0.35 | 0.58 | 0.84 | 0.98 | 1.00 | 1.00 |
| 4 | 0.97 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 6 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 8 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 10 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
Next, we apply 1 to MMSBMs. For each specification of the number of communities \(k\), we construct two block matrices \(B^{(1)}\) and \(B^{(2)}\). The diagonal entries of these matrices are independently sampled from \([0.8,1]\), while the off-diagonal entries of \(B^{(1)}\) are sampled from \([0.3,0.5]\) and those of \(B^{(2)}\) are sampled from \([0.1,0.3]\)
For \(k=3\), the block matrices are given below: \[B^{(1)} = \begin{bmatrix} 0.98 & 0.53 & 0.58 \\ 0.53 & 0.99 & 0.54 \\ 0.58 & 0.54 & 0.81 \end{bmatrix}, \qquad B^{(2)} = \begin{bmatrix} 0.97 & 0.11 & 0.12 \\ 0.11 & 0.96 & 0.17 \\ 0.12 & 0.17 & 0.87 \end{bmatrix}.\] For \(k=5\), the two block matrices are given below: \[B^{(1)} = \begin{bmatrix} 0.99 & 0.53 & 0.58 & 0.59 & 0.55 \\ 0.53 & 0.89 & 0.54 & 0.59 & 0.59 \\ 0.58 & 0.54 & 0.94 & 0.50 & 0.56 \\ 0.59 & 0.59 & 0.50 & 0.91 & 0.55 \\ 0.55 & 0.59 & 0.56 & 0.55 & 0.82 \end{bmatrix}, \qquad B^{(2)} = \begin{bmatrix} 0.87 & 0.11 & 0.12 & 0.19 & 0.11 \\ 0.11 & 0.84 & 0.17 & 0.18 & 0.13 \\ 0.12 & 0.17 & 0.95 & 0.13 & 0.12 \\ 0.19 & 0.18 & 0.13 & 0.96 & 0.14 \\ 0.11 & 0.13 & 0.12 & 0.14 & 0.92 \end{bmatrix}.\]
To construct the membership matrices, we proceed as follows. For \(Z^{(1)}\), we randomly select \(n_0\) nodes and assign their community membership vectors to be \(\left( \tfrac{1}{k-1}, \tfrac{1}{k-1}, \ldots, \tfrac{1}{k-1}, 0 \right),\) while the membership vectors of all remaining nodes are independently sampled from a Dirichlet distribution with parameter vector \(\left( \tfrac{1}{k}, \tfrac{1}{k}, \ldots, \tfrac{1}{k} \right).\) Similarly, for \(Z^{(2)}\), we choose the same \(n_0\) nodes as before and assign their community membership vectors to be \(\left( \tfrac{1}{k}, \tfrac{1}{k}, \ldots, \tfrac{1}{k} \right),\) while the remaining nodes are kept identical to those in \(Z^{(1)}\). As before, we form the two probability matrices \(P^{(i)} = {Z^{(i)}} B^{(i)} {Z^{(i)}}^{\top}\) for \(i=1,2,\) and generate the observed adjacency matrices according to \({A^{(i)}} \sim \text{Ber}(P^{(i)}).\)
For each fixed combination of \(n,k,\rho_n\), and \(n_0\), we simulate the data set 100 times and apply the test at \(\alpha=0.05\) level of significance, recording the number of rejections. The empirical power at those parameters is then estimated by the fraction of rejections. 6 8 report these empirical powers, with rows corresponding to different values of \(\rho_n\) and columns corresponding to different values of \(n_0\). Consistent with the results for the SBM, we observe analogous trends under the MMSBM. The first row (\(n_0 = 0\)) reflects the performance of the test under the null hypothesis in 1 . The finite-sample empirical sizes remain closely aligned with the nominal level \(\alpha = 0.05\), and this calibration steadily improves as the density parameter \(\rho_n\) increases for a fixed \(k\). Furthermore, under the alternative hypothesis (\(n_0 > 0\)), the empirical power again exhibits a stable convergence to \(1\) as the perturbation size \(n_0\) increases, aligning with our theoretical expectations. The power with \(k = 5\) tends to one more slowly than in the previous subsection, which suggests that detecting changes in this setup is more difficult than the case where networks are discrete stochastic blockmodels.
0.49
3.5pt
| \(n_0\) | \(\rho_n=0.75\) | \(\rho_n=0.8\) | \(\rho_n=0.85\) | \(\rho_n=0.9\) | \(\rho_n=0.95\) | \(\rho_n=1\) |
|---|---|---|---|---|---|---|
| 0 | 0.03 | 0.06 | 0.13 | 0.07 | 0.09 | 0.06 |
| 20 | 0.81 | 0.86 | 0.88 | 0.98 | 1.00 | 1.00 |
| 40 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 60 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 80 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 100 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 150 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
0.49
3.5pt
| \(n_0\) | \(\rho_n=0.75\) | \(\rho_n=0.8\) | \(\rho_n=0.85\) | \(\rho_n=0.9\) | \(\rho_n=0.95\) | \(\rho_n=1\) |
|---|---|---|---|---|---|---|
| 0 | 0.12 | 0.09 | 0.10 | 0.06 | 0.06 | 0.04 |
| 20 | 0.80 | 0.87 | 0.93 | 1.00 | 1.00 | 1.00 |
| 40 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 60 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 80 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 100 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| 150 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
0.49
3.5pt
| \(n_0\) | \(\rho_n=0.75\) | \(\rho_n=0.8\) | \(\rho_n=0.85\) | \(\rho_n=0.9\) | \(\rho_n=0.95\) | \(\rho_n=1\) |
|---|---|---|---|---|---|---|
| 0 | 0.87 | 0.71 | 0.37 | 0.09 | 0.08 | 0.02 |
| 20 | 0.92 | 0.86 | 0.51 | 0.26 | 0.13 | 0.11 |
| 40 | 0.95 | 0.93 | 0.66 | 0.41 | 0.35 | 0.30 |
| 60 | 0.99 | 0.95 | 0.73 | 0.64 | 0.62 | 0.55 |
| 80 | 1.00 | 0.98 | 0.88 | 0.83 | 0.78 | 0.71 |
| 100 | 1.00 | 0.97 | 0.95 | 0.93 | 0.90 | 0.90 |
| 150 | 1.00 | 1.00 | 1.00 | 0.99 | 1.00 | 1.00 |
0.49
3.5pt
| \(n_0\) | \(\rho_n=0.75\) | \(\rho_n=0.8\) | \(\rho_n=0.85\) | \(\rho_n=0.9\) | \(\rho_n=0.95\) | \(\rho_n=1\) |
|---|---|---|---|---|---|---|
| 0 | 0.70 | 0.43 | 0.21 | 0.06 | 0.09 | 0.07 |
| 20 | 0.77 | 0.62 | 0.37 | 0.18 | 0.19 | 0.09 |
| 40 | 0.87 | 0.80 | 0.56 | 0.38 | 0.36 | 0.27 |
| 60 | 0.93 | 0.89 | 0.71 | 0.66 | 0.56 | 0.51 |
| 80 | 0.96 | 0.94 | 0.81 | 0.80 | 0.78 | 0.74 |
| 100 | 0.98 | 0.97 | 0.92 | 0.92 | 0.90 | 0.89 |
| 150 | 0.99 | 0.99 | 1.00 | 0.99 | 1.00 | 1.00 |
We also apply 1 to U.S.domestic flight data publicly available from the Bureau of Transportation Statistics [58] and previously analyzed in [42] and [59].. After restricting to the largest connected component, we obtain \(n=343\) airports and data for \(T=69\) months (January 2016 through September 2021). For each month \(t=1,\dots,T\), we have a count matrix \(F_t=(F_t(i,j))_{i,j=1}^n\in\mathbb{N}^{n\times n},\) where \(F_t(i,j)\) denotes the number of flights from airport \(i\) to airport \(j\) during month \(t\). To obtain a binary network representation for each month, we symmetrize and threshold the counts: define the adjacency matrix \(A_t=(A_t(i,j))_{i,j=1}^n\in\{0,1\}^{n\times n}\) by \[A_t(i,j) = \mathbb{1}\{\,F_t(i,j)+F_t(j,i)\ge 1\,\},\qquad i\neq j,\] and set \(A_t(i,i)=0\) for all \(i\). In words, \(A_t(i,j)=1\) if there exists at least one flight between airports \(i\) and \(j\) (in either direction) during month \(t\), and \(A_t(i,j)=0\) otherwise. For detecting the number of communities \(k\), we select the embedding dimension using the “elbow” method of [60]; the resulting choice is \(k=5\).
The months of January, June, and November provide a particularly informative contrast for assessing structural stability in the U.S. flight network. June 2020 corresponds to the period of most pronounced disruption in national air-traffic patterns due to the COVID shock, and thus serves as a natural stress point at which one might expect substantive alterations in the community structure. By comparison, January and November exhibit far more regular seasonal behavior across years, and consequently their underlying connectivity patterns are expected to display a high degree of structural persistence. This separation of regimes allows for a clear evaluation of the sensitivity of our methodology to both stable and significantly perturbed network environments.
In 9 10 11, the number of rows \(m\) in each table corresponds to the frequency of that specific month within the study period. Consequently, there are \(m=6\) occurrences for January and June, and \(m=5\) occurrences for November. The \((i,j)\)-th entry, \(i\neq j\), reports the \(p\)-value from a hypothesis test comparing the adjacency matrices of that month in year \(i\) and year \(j\). Because multiple pairwise comparisons are performed (there are \(m(m-1)/2\) tests per month), we apply a Bonferroni correction: at nominal level \(\alpha=0.05\), the significance threshold is \({2\alpha}/({m(m-1))}.\) Cells with \(p\)-values at or below this threshold are displayed in red, and cells with \(p\)-values above the threshold are displayed in green.
| 2016-01 | 2017-01 | 2018-01 | 2019-01 | 2020-01 | 2021-01 | |
|---|---|---|---|---|---|---|
| 2016-01 | 0.0726 | 0.0875 | 0.2437 | 0.8976 | 0.2951 | |
| 2017-01 | 0.0726 | 0.0615 | 0.1810 | 0.4708 | 0.3407 | |
| 2018-01 | 0.0875 | 0.0615 | 0.0634 | 0.2615 | 0.8945 | |
| 2019-01 | 0.2437 | 0.1810 | 0.0634 | 0.0787 | 0.8612 | |
| 2020-01 | 0.8976 | 0.4708 | 0.2615 | 0.0787 | 0.8358 | |
| 2021-01 | 0.2951 | 0.3407 | 0.8945 | 0.8612 | 0.8358 |
| 2016-06 | 2017-06 | 2018-06 | 2019-06 | 2020-06 | 2021-06 | |
|---|---|---|---|---|---|---|
| 2016-06 | 0.1815 | 0.5459 | 0.9343 | 0.0002 | 0.0760 | |
| 2017-06 | 0.1815 | 0.2123 | 0.4802 | 0.0005 | 0.1777 | |
| 2018-06 | 0.5459 | 0.2123 | 0.0674 | 0.0001 | 0.0919 | |
| 2019-06 | 0.9343 | 0.4802 | 0.0674 | 0.0000 | 0.3596 | |
| 2020-06 | 0.0002 | 0.0005 | 0.0001 | 0.0000 | 0.2915 | |
| 2021-06 | 0.0760 | 0.1777 | 0.0919 | 0.3596 | 0.2915 |
| 2016-11 | 2017-11 | 2018-11 | 2019-11 | 2020-11 | |
|---|---|---|---|---|---|
| 2016-11 | 0.0101 | 0.2853 | 0.5370 | 0.4457 | |
| 2017-11 | 0.0101 | 0.0891 | 0.1971 | 0.8067 | |
| 2018-11 | 0.2853 | 0.0891 | 0.0876 | 0.6468 | |
| 2019-11 | 0.5370 | 0.1971 | 0.0876 | 0.9973 | |
| 2020-11 | 0.4457 | 0.8067 | 0.6468 | 0.9973 |
For January and November, all cells are shaded green, indicating that the corresponding flight networks are not significantly different across years at the Bonferroni-corrected level. In contrast, for June the entries involving year 2020 are shaded red, while all other entries remain green. This pattern is attributable to the COVID–19 pandemic, which peaked in June 2020 and caused a severe disruption of U.S.domestic air traffic. During this period many airports experienced no incoming or outgoing flights, leading to a flight network structure that is markedly different from the corresponding networks in other years. These findings illustrate that our proposed testing methodology is capable of distinguishing between genuinely similar network structures (as in January and November) and substantive structural changes (as in June 2020).
In order to understand the proof of 1 for the two‐sample setting, we find it beneficial to study its one‐sample analogue. In this simpler setting, we observe a single adjacency matrix \(A \sim \mathop{\mathrm{Ber}}(P),\) with population spectral decomposition \(P = V \Lambda V^{\top}\). We wish to test whether the principal subspace of \(P\) coincides with a given reference subspace spanned by an orthonormal matrix \(V^*\in\mathbb{R}^{n\times k}\). Formally, we consider \[\begin{align} \label{hypo3} H_0: V V^{\top}= V^* V^{*^{\top}} \quad\text{versus}\quad H_1: V V^{\top}\neq V^* V^{*^{\top}} \end{align}\tag{14}\] Let \(\widehat V\in\mathbb{R}^{n\times k}\) denote the matrix of the leading \(k\) empirical eigenvectors of \(A\). We define the one‐sample test statistic \[T_{n}^{({\sf os})} := \bigl\|\widehat V\,\widehat V^{\top}- V^* V^{*^{\top}}\bigr\|_F^2.\] Note that we can equivalently reject the null hypothesis at level \(\alpha\) if we can form an appropriate \(1 - \alpha\) confidence interval. The following theorem establishes the Gaussian limit under the null hypothesis in 14 .
Theorem 5. Suppose 1 2 3 4 hold. Then under \(H_0\) of 1 , \[\label{eq:clt1} \frac{T_{n}^{({\sf os})} - \mu_1}{\sigma_1} \xrightarrow{D} \mathcal{N}(0,1),\tag{15}\] where \[\begin{align} \mu_1 &= 2\,\mathrm{tr}\Bigl(\beta^{-2}\left[\Sigma\circ\beta^\perp + \mathop{\mathrm{Diag}}(\Sigma \cdot d-\mathrm{diag}(\Sigma) \circ d)\right]\Bigr) + O_p\left(\frac{k}{n^2 \rho_n^2}\right), \tag{16}\\ \sigma^2_1 &= 8 \Bigl\langle (J_n - I_n)\circ(\beta^{-2}\circ\beta^{-2}),\Sigma^2\Bigr\rangle + 4 \,\alpha^{\top}K\,\alpha+o \left(\frac{k^2}{n^3 \rho_n^2} \right), \tag{17} \end{align}\] with \(V V^{\top}= V^* V^{*^{\top}}\). Here we define \[P = V \Lambda V^{\top}, \quad \beta^k = V\Lambda^kV^{\top}, \quad \beta^\perp = I_n - VV^{\top}, \quad d = \mathrm{diag}(\beta^\perp),\] \[\Sigma = P\circ(J_n-P), \quad K = 2\,\Sigma\circ(J_n - 2P), \quad \alpha = \mathrm{diag}(\beta^{-2}).\]
Proof. See 7. ◻
The above one–sample result serves as the cornerstone for deriving the more delicate two–sample asymptotic limit stated in Theorem 1. In particular, the expansion arguments and concentration inequalities developed here form the analytical basis for the proof of the two–sample result.
Remark 2 (Proof Overview). We provide a brief outline of the proof of 5. The argument begins with a series expansion of the projection distance \(\|\widehat V \widehat V^{\top}- V^* {V^*}^{\top}\|_F^2\) first developed in [51]. We show the statistic decomposes as \[T_{n}^{({\sf os})} = T_{1}^{(S)} + T_{1}^{(3)} + T^{(H)}_{1},\] where we denote the second–order term by \(T_{1}^{(S)}\), the third–order term by \(T_{1}^{(3)}\) and the remaining higher–order terms collectively by \(T^{(H)}_{1}\). The first order term vanishes. The dominant component is \(T_{1}^{(S)}\), and the central limit behavior follows from a martingale central limit theorem (CLT), building on ideas from [43]. In particular, we can show that \[\frac{T_{1}^{(S)} - \mathbb{E}\big(T_{1}^{(S)}\big)}{\sqrt{\mathop{\mathrm{Var}}\big(T_{1}^{(S)}\big)}} \xrightarrow{D} \mathcal{N}(0,1).\] The variance comparison is then crucial: we establish that \(\mathop{\mathrm{Var}}( T_{1}^{(S)}) \gg \mathop{\mathrm{Var}}(T^{(3)}_{1})\). For the expectations, we verify that \(\mathbb{E}(T^{(3)}_{1})\) is of order \(o(\frac{k}{n^2 \rho_n^2})\). Furthermore, by a straightforward application of Chebyshev’s inequality, \(\big(T^{(3)}_{1}-\mathbb{E}\big(T^{(3)}_{1}\big)\big)\big/\sqrt{\mathop{\mathrm{Var}}\big(T_{1}^{(S)}\big)}\) vanish in probability. Finally from Theorem 1 of [51], we show that \(\|T^{(H)}_1\|_F \lesssim \frac{k}{n^2 \rho_n^2}\) with high probability.
Putting these ingredients together, we arrive at the desired Gaussian limit under 2: \[\frac{T_{n}^{({\sf os})} - \mu_1}{\sigma_1} \xrightarrow{D} \mathcal{N}(0,1),\] where \(\mu_1=\mathbb{E}\left(T_{1}^{(S)}\right)+O_p \left(\frac{k}{n^2 \rho_n^2} \right)\) and \(\sigma_1^2= \mathop{\mathrm{Var}}\left(T_{1}^{(S)} \right)\).
For the test implementation, it remains to consistently estimate the mean and variance appearing in 15 –17 . The estimation procedure mirrors that in the two–sample setting: population eigenvectors \(V\) are replaced by their empirical estimates \(\widehat{V}\), and the population eigenvalues by their empirical counterparts \(\widehat{\Lambda}\).
Define the plug-in estimators \[\label{eq:estimatorsone} \begin{align} \widehat{\mu}_1 &=2\mathop{\mathrm{tr}}\Bigl[\widehat{\beta}^{-2}\Bigl(\widehat\Sigma\circ \widehat\beta^\perp + \mathop{\mathrm{Diag}}\bigl(\widehat{\Sigma}\cdot \widehat d \bigr)\Bigr)\Bigr], \\[6pt] \widehat{\sigma}_1^2 &= 8 \Bigl\langle (J_n - I_n)\circ(\widehat{\beta}^{-2}\circ\widehat{\beta}^{-2}),\widehat{\Sigma}^2\Bigr\rangle + 4 \,\widehat{\alpha}^{\top}\widehat{K}\,\widehat{\alpha}, \end{align}\tag{18}\] where \[\begin{align} \label{thm6461} \widehat{P}=\widehat V\,\widehat{\Lambda}\,\widehat V^{\top}, \qquad \widehat{\beta}^k=\widehat V\,\widehat{\Lambda}^{k}\,\widehat V^{\top}, \qquad \widehat\beta^\perp=I_n-\widehat V\widehat V^{\top}, \qquad \widehat d=\mathrm{diag}(\widehat\beta^\perp), \end{align}\tag{19}\] and \[\begin{align} \label{thm6462} \widehat{\Sigma}=\widehat{P}\circ(J_n-\widehat{P}), \qquad \widehat{K}=2\,\widehat{\Sigma}\circ(J_n-2\widehat{P}), \qquad \widehat{\alpha}=\mathrm{diag}(\widehat{\beta}^{-2}). \end{align}\tag{20}\]
Theorem 6 (Mean and variance estimators for the one–sample test). The estimators in 18 satisfy the following rate under 1 5 3 4: \[|\widehat{\mu}_1 - \mu_1| =O_p \left(\max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right)\right), \qquad |\widehat{\sigma}_1^2 - \sigma_1^2| = o_p \left(\frac{k^2}{n^3 \rho_n^2}\right).\]
Proof. See 9.3.1. ◻
The following proposition establishes that the upper bound on \(|\widehat{\mu}_1 - \mu_1|\) as stated in 6 is sharp and therefore cannot be improved for plug-in estimators.
Proposition 4. Suppose \(P\) is generated from a SBM. Let \(\widehat{\mu}_1\) be the estimator defined as in 18 . Then, under 1 3 4 5, \[\bigl|\widehat{\mu}_1 - \mu_1\bigr| \asymp \max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right).\]
Proof. See 15.5. ◻
An immediate implication is that the proposed inference procedure cannot remain valid outside the dense regime (5). In particular, when \(n\rho_{n} \lesssim \sqrt{n}\), the estimation error \(\frac{k}{n^{2}\rho_{n}^{2}}\) causes the test statistic to lose consistency. This precisely motivates the necessity of 5.
In this work, we have developed a rigorous and computationally efficient two-sample testing framework for assessing the equality of the underlying low-rank subspaces of two networks. At the core of our methodology are novel Gaussian limit theorems for the projection distance between estimated spectral subspaces. Notably, these limit theorems are the first of their kind established explicitly under a Bernoulli noise model. Supported by easily computable, data-driven plug-in estimators, our proposed test demonstrates strong empirical power in detecting subtle structural shifts. Furthermore, the generality of our approach extends beyond detecting community structure changes in SBMs. For instance, within the framework of GRDPGs, our test inherently determines whether the latent positions of two graphs span the exact same feature space, rendering the procedure robust to arbitrary linear transformations of the latent geometry.
While our findings establish a theoretically grounded foundation for network hypothesis testing, they simultaneously open several broader avenues for future research. For instance, one could extend the test statistic to incorporate information beyond the principal angles between the estimated subspaces. Our current angle-based procedure provides robustness against global density fluctuations; however, incorporating the spectral distance between the underlying connectivity matrices could further enhance statistical power.
Furthermore, while our theoretical framework establishes rigorous guarantees under bounded, independent Bernoulli noise, extending these limit theorems to accommodate more complex noise structures presents an exciting frontier. Two immediate avenues emerge:
Heavy-Tailed Distributions: Developing limit theorems for spectral projectors under heavy-tailed noise would broaden the applicability of spectral testing methods to weighted networks.
Dependent Edges: Investigating the robustness of the test statistic under weak, local edge dependencies (e.g., transitivity or reciprocity) poses a challenging and valuable open problem.
Finally, our current analysis assumes the rank \(k\) of the probability matrices is known a priori. While several consistent estimators for \(k\) are readily available in the literature, formally incorporating the theoretical uncertainty of rank selection into the asymptotic distribution of the test statistic is a mathematically non-trivial task. Developing a fully adaptive testing procedure that remains robust to an estimated rank, or constructing a test intrinsically agnostic to the specific choice of \(k\) would be another interesting avenue for future research.
In this section we give the full proof of 5. Our results are based on three steps. In 7.1 we apply the matrix series expansion of [51] to identify the leading-order and higher-order terms. In 7.2 we study the asymptotic behavior of the leading-order term, and in 7.3 we show that the higher-order terms are asymptotically negligible relative to the leading-order term. Combining all these ingredients gives the final proof of 5.
In order to write the test statistic \(\|\widehat{V}\widehat{V}^{\top}- VV^{\top}\|_F^2\) using Theorem 1 of [51], we first define the following good event and show it holds with high probability.
Lemma 1. Suppose \(A \sim \mathop{\mathrm{Ber}}(P)\), and let \(X = A - P\). Define the event \[\mathcal{E}_{{\sf good}} = \{\|X\|\lesssim \sqrt{n\rho_n}\,\}.\] Then under 1 3 2, \(\mathbb{P}(\mathcal{E}_{{\sf good}})\ge1-O\left(n^{-19}\right).\) Moreover, on the event \(\mathcal{E}_{{\sf good}}\) the following bound holds: \[\big\|\widehat V\widehat V^{\top}- V V^{\top}\big\|_F^2 \lesssim \frac{k}{n\rho_n},\] where \(\widehat V\) and \(V\) are the matrices of estimated and population top-\(k\) eigenvectors, respectively. Thus, we have \[\big\|\widehat V\widehat V^{\top}- V V^{\top}\big\|_F^2 = O_p \left(\frac{k}{n\rho_n}\right) \quad \text{and} \quad \|X\|=O_p( \sqrt{n \rho_n}).\]
Proof. See 11.1. ◻
Define \[\begin{align} \beta^0 = \beta^\perp &= I - VV^{\top}; \\ \beta^{-l} &= V \Lambda^{-l} V^{\top}, \end{align}\] and for each integer \(l \geq 1\), \[\begin{align} \label{Sk951sample} S_{l}(X) &=\sum_{\substack{s=(s_1,\dots,s_{l+1}):\\ s_1+\cdots+s_{l+1}=l}} (-1)^{1+\tau(s)}\; \beta^{-s_1}\, X\, \beta^{-s_2}\, X\cdots \beta^{-s_{l}} X\, \beta^{-s_{l+1}}, \\ \tau(s)&=\sum_{j=1}^{l+1} 1\{s_j>0\}, \end{align}\tag{21}\] where we recall \(X:=A-P\) denotes the perturbation matrix. On the event \(\mathcal{E}_{{\sf good}}\), 3 impies that \(\| X \| \ll \lambda_k\), and hence the conditions of Theorem 1 of [51] are satisfied. Applying this result, we have that \[\label{eq1} \begin{align} \|\widehat{V}\widehat{V}^{\top}- VV^{\top}\|_F^2 &=\,-2\,\sum_{l\ge2}\,\langle VV^{\top},\,S_l(X)\rangle_F \\ &=2\,\|\beta^{\perp}X\beta^{-1}\|_F^2 \;-\;2\,\sum_{l\ge3}\,\langle VV^{\top},\,S_l(X)\rangle_F. \end{align}\tag{22}\] The leading nonvanishing contribution to the eigenspace projection error thus arises from the second–order term \(2\,\|\beta^{\perp}X\beta^{-1}\|_F^2\). We derive its asymptotic distribution in 7.2, and study the higher-order terms in 7.3.
We begin this subsection by deriving the mean and variance of the leading contribution to the test statistic. Define \[\begin{align} T_{1}^{(S)} = 2\big\|\beta^{\perp} X \beta^{-1}\big\|_F^2 = 2\mathop{\mathrm{tr}}\big(\beta^{-1}X^{\top}\beta^{\perp} X\beta^{-1}\big). \end{align}\] Our first result demonstrates that the variance of this term can be approximated by the variance of a similar term with \(\beta^{\perp}\) taken to be the identity.
Lemma 2. Suppose 1 2 3 4 hold. Then \[\frac{\mathop{\mathrm{Var}}\big(\|\beta^{\perp} X \beta^{-1}\|_F^2\big)}{\mathop{\mathrm{Var}}\big(\|X \beta^{-1}\|_F^2\big)} \longrightarrow 1,\]
Proof. See 11.2. ◻
The following result computes this mean and variance.
Lemma 3 (Second–order term: mean and variance of one-sample test). Suppose 1 2 3 4 hold. Define \[\Sigma := P\circ(J_n-P),\qquad d:=\mathrm{diag}(\beta^\perp),\qquad \alpha:=\mathrm{diag}(\beta^{-2}),\qquad K:=2\,\Sigma\circ(J_n-2P).\] The mean and variance of \(T_{1}^{(S)}\) admit the following expressions: \[\mathbb{E}\bigl[T_{1}^{(S)}\bigr] = 2\,\mathop{\mathrm{tr}}\Bigl[\beta^{-2}\bigl(\Sigma\circ\beta^\perp + \mathrm{Diag}(\Sigma\cdot d-\mathrm{diag}(\Sigma) \circ d)\bigr)\Bigr],\] and \[\mathop{\mathrm{Var}}\bigl(T_{1}^{(S)}\bigr) = 8\big\langle (J_n-I_n)\circ(\beta^{-2}\circ\beta^{-2}),\Sigma^2 \big\rangle + 4\,\alpha^{\top}K \alpha + o \left( \frac{k^2}{n^3 \rho_n^2} \right).\] The quantities coincide with the leading mean and variance contributions stated in 16 17 . In particular, we also have \(\mathop{\mathrm{Var}}\bigl(T_{1}^{(S)}\bigr) \asymp \frac{k^{2}}{n^{3}\rho_n^{2}}.\)
Proof. See 11.3. ◻
Having calculated these moments, we now study the asymptotic distribution of the first-order approximation \(T_1^{(S)}\). The following result uses the martingale central limit theorem to establish the Gaussian limit for the centered and scaled second-order term.
Lemma 4 (Second–order asymptotic distribution of one-sample test). Suppose 1 2 3 4 hold. The quantity \(T_{1}^{(S)}\) satisfies \[\frac{ T_{1}^{(S)} - \,\mathbb{E}[T_{1}^{(S)}] }{ \sqrt{\mathop{\mathrm{Var}}(T_{1}^{(S)})} } \xrightarrow{D} \mathcal{N}(0,1)\] as \(n \to \infty\).
Proof. See 11.4. ◻
This completes the derivation of the Gaussian limit for the second–order term and provides the principal probabilistic ingredient in the proof of 5.
In this section we study the higher-order terms from the decomposition 22 . For \(l \ge 3\), each term in the expansion 22 takes the form \(\langle V V^{\top}, S_{l}(X) \rangle_F\). Using the definition of \(S_l(X)\), this can be written as \[\begin{align}\label{bigvar} \langle V V^{\top}, S_{l}(X) \rangle_F &= \sum_{\substack{s=(s_1,\dots,s_{l+1}):\\ s_1+\cdots+s_{l+1}=l}} (-1)^{1+\tau(s)}\; \langle V V^{\top}, \beta^{-s_1}\, X\, \beta^{-s_2}\, X\cdots \beta^{-s_{l}} X\, \beta^{-s_{l+1}}\rangle. \end{align}\tag{23}\]
We separate the expansion into the third-order term and higher order terms. We first analyze the contribution of \(T_1^{(3)}= \langle V V^{\top}, S_{3}(X) \rangle_F\).
Lemma 5. Assume that 1 2 3 4 hold. Let \(S_{3}(X)\) be as specified in 21 . Then, we have \[\mathbb{E}\langle V V^{\top},S_3(X)\rangle = o \left(\frac{k}{n^2 \rho_n^2}\right).\]
Proof. See 11.5. ◻
We next consider the variance of \(\langle VV^{\top}, S_3(X)\rangle\).
Lemma 6. Assume that 1 2 3 4 hold. Then \(\mathop{\mathrm{Var}}(\langle V V^{\top},S_3(X)\rangle) \asymp \frac{k^2}{n^4 \rho_n^3}\).
Proof. See 11.6. ◻
A direct application of Chebyshev’s inequality gives \[\begin{align} \mathbb{P}\left[\, \frac{\big|T_{1}^{(3)} - \mathbb{E}\left(T_{1}^{(3)}\right)\big|}{\sqrt{\mathop{\mathrm{Var}}\left(T_{1}^{(S)}\right)}} \;\ge\; \delta \right] &\le \frac{\mathop{\mathrm{Var}}\left(T_{1}^{(3)}\right)}{\delta^2 \mathop{\mathrm{Var}}(T_{1}^{(S)})}\\ &\lesssim \frac{1}{\delta^2 n \rho_n}, \end{align}\] for any \(\delta > 0\). Therefore by 2, \[\begin{align} \label{eq:cheby1} \frac{\left|T_{1}^{(3)} - \mathbb{E}\left(T_{1}^{(3)}\right)\right|}{\sqrt{\mathop{\mathrm{Var}}\left(T_{1}^{(S)}\right)}} \;\to\; 0 \qquad \text{in probability}. \end{align}\tag{24}\] The higher-order terms are significantly simpler to analyze. We note that whenever \(A\in\mathcal{E}_{{ \sf good}}\), we have \(\|S_l(X)\| \le \Big(\tfrac{c}{n\rho_n}\Big)^{l/2}\) and \(\mathop{\mathrm{rank}}\big(S_l(X)\big)\le{2l\choose l}k\;\le\;4^{l}k,\) for some constant \(c>0\). Consequently, \[\|S_l(X)\|_{F} \;\le\; \sqrt{\mathop{\mathrm{rank}}(S_l(X))}\,\|S_l(X)\| \;\le\; \big(4^{l}k\big)^{1/2}\Big(\tfrac{c}{n\rho_n}\Big)^{l/2} = \sqrt{k} \Big(\tfrac{4c}{n\rho_n}\Big)^{l/2}.\]
Using this bound termwise in the higher order terms, \(T_1^{(H)} = \big\langle VV^{\top}, \sum_{l\ge4} S_l(X) \big\rangle,\) we obtain \[|T_1^{(H)}| \;\le\; k\,\Big\|\sum_{l\ge4} S_l(X)\Big\|_{F} \;\le\; \sum_{l\ge4} k\,\|S_l(X)\|_{F} \;\le\; k\sum_{l\ge4}\Big(\tfrac{4c}{n\rho_n}\Big)^{l/2} \frac{16 k c^{2}}{n^{2}\rho_n^{2}}.\] Therefore \(T_1^{(H)} = O_{p}\left(\frac{k}{n^{2}\rho_n^{2}}\right).\) Combining this result with 4 and 24 yields the claim of 5. This completes the proof.
We now turn to the proof of 1. Following the same methodology as in the proof of 5 in the previous section, we begin by developing the matrix expansion of the two-sample statistic. The analysis proceeds in several steps. First, we derive the explicit form of the expansion and isolate the second-order contribution, which captures the leading fluctuation behavior of the statistic. We then compute the mean and variance of this second-order term and establish its asymptotic Gaussianity under the stated assumptions. Finally, we demonstrate that the higher-order terms possess variances of smaller order compared to that of the second-order component. Consequently, these higher-order terms affect only the mean of the test statistic while leaving its asymptotic variance unchanged.
The two–sample projection–distance statistic considered in 1 is \(\|(\widehat V^{(1)}) (\widehat V^{(1)})^{\top}- (\widehat V^{(2)}) (\widehat V^{(2)})^{\top}\|_F^2\). Define \[\begin{align} \beta_i^k = {V^{(i)}} (\Lambda^{(i)})^k {V^{(i)}}^{\top}, \quad \beta^{\perp} = I_n - {V^{(i)}} {V^{(i)}}^{\top}. \end{align}\] Applying the matrix expansion of [51] to each empirical projector and collecting terms gives the decomposition \[\begin{align}\label{eq2} \|(\widehat V^{(1)}) (\widehat V^{(1)})^{\top}- (\widehat V^{(2)}) (\widehat V^{(2)})^{\top}\|_F^2 &= \big\|\widehat{V}_{1} \widehat{V}_{1}^{\top}- {V^{(1)}} {V^{(1)}}^{\top}\big\|_F^2 + \big\|\widehat{V}_{2} \widehat{V}_{2}^{\top}- {V^{(2)}} {V^{(2)}}^{\top}\big\|_F^2 \notag \\ &\qquad - 2\, \mathop{\mathrm{tr}}\Big(\big(\widehat{V}_{1} \widehat{V}_{1}^{\top}- {V^{(1)}} {V^{(1)}}^{\top}\big)\big(\widehat{V}_{2} \widehat{V}_{2}^{\top}- {V^{(2)}} {V^{(2)}}^{\top}\big)\Big) \\ &= 2\big\|\beta_1^{\perp} {X^{(1)}} \beta_1^{-1}\big\|_F^2 - 2\,\sum_{j\ge3}\,\langle {V^{(1)}} {V^{(1)}}^{\top},\,S_{1,j}({X^{(1)}})\rangle_F + 2\big\|\beta_2^{\perp} {X^{(2)}} \beta_2^{-1}\big\|_F^2 \\ &\quad - 2\,\sum_{j\ge3}\,\langle {V^{(1)}} {V^{(1)}}^{\top},\,S_{2,j}({X^{(2)}})\rangle_F - 2\, \mathop{\mathrm{tr}}\bigl(S_{1,1}({X^{(1)}})\, S_{2,1}({X^{(2)}})\bigr)\\ &\quad - 2\sum_{l \geq 4} \sum_{l_1+l_2=l} \, \mathop{\mathrm{tr}}\bigl(S_{1,1}({X^{(1)}})\, S_{2,1}({X^{(2)}})\bigr), \end{align}\tag{25}\] where \({X^{(1)}}={A^{(1)}}-{P^{(1)}}\), \({X^{(2)}}={A^{(2)}}-{P^{(2)}}\) and \[\begin{align} \label{Sk952sample} S_{i,k}(X^{(i)}) &=\sum_{\substack{s=(s_1,\dots,s_{k+1}):\\ s_1+\cdots+s_{k+1}=k}} (-1)^{1+\tau(s)}\; \beta_i^{-s_1}\, {X^{(i)}}\, \beta_i^{-s_2}\, {X^{(i)}}\cdots \beta_i^{-s_{k}}{X^{(i)}}\, \beta_i^{-s_{k+1}}, \\ \tau(s)&=\sum_{j=1}^{k+1} 1\{s_j>0\} \notag. \end{align}\tag{26}\] Thus, the second-order term takes the form \[\begin{align} \label{T952} T_2^{(S)}&= 2\big\|\beta_1^{\perp} {X^{(1)}} \beta_1^{-1}\big\|_F^2 + 2\big\|\beta_2^{\perp} {X^{(2)}} \beta_2^{-1}\big\|_F^2 - 2\, \mathop{\mathrm{tr}}\bigl(S_{1,1}({X^{(1)}})\, S_{2,1}({X^{(2)}})\bigr) \\[4pt] &= 2\big\|\beta^{\perp} {X^{(1)}} \beta_1^{-1}\big\|_F^2 + 2\big\|\beta^{\perp} {X^{(2)}} \beta_2^{-1}\big\|_F^2 - 2\, \mathop{\mathrm{tr}}\big(\beta^{\perp} {X^{(1)}} \beta_1^{-1} \beta_2^{-1} {X^{(2)}}\big) - 2\, \mathop{\mathrm{tr}}\big(\beta_1^{-1} {X^{(1)}} \beta^{\perp} {X^{(2)}} \beta_2^{-1} \big) \\[4pt] &= 2\bigl\|\beta^{\perp}\bigl( {X^{(1)}} \beta_1^{-1} - {X^{(2)}} \beta_2^{-1} \bigr)\bigr\|_F^2 \\[4pt] &= 2\,\mathop{\mathrm{tr}}\left( U^{\top} \begin{pmatrix} (Y^{(1)})^{\top}(Y^{(1)}) & -(Y^{(1)})^{\top}(Y^{(2)})\\[4pt] -(Y^{(2)})^{\top}(Y^{(1)}) & (Y^{(2)})^{\top}(Y^{(2)}) \end{pmatrix} U \right), \end{align}\tag{27}\] where \(Y^{(i)} = \beta_i^{\perp}{X^{(i)}}\) for \(i=1,2\) and \(U = \begin{pmatrix} \beta_1^{-1} & \beta_2^{-1} \end{pmatrix}\) is the corresponding block matrix (whose columns we denote by \(U_{1 \cdot},\dots,U_{2n \cdot}\in\mathbb{R}^n\)). The representation in 27 exhibits the two–sample second–order contribution as a quadratic form in the block matrix of vectors \(U_{j \cdot}\), and provides the starting point for the subsequent mean/variance calculations and the martingale central limit argument.
As in the one-sample case, we first derive the mean and variance of the leading contribution to the two–sample projection–distance statistic. Recall the definition: \[T_{2}^{(S)} = 2\bigl\|\beta_1^{\perp} {X^{(1)}} \beta_1^{-1}\bigr\|_F^2 + 2\bigl\|\beta_2^{\perp} {X^{(2)}} \beta_2^{-1}\bigr\|_F^2 - 2\, \mathop{\mathrm{tr}}\bigl(S_{1,1}({X^{(1)}})\, S_{2,1}({X^{(2)}})\bigr),\] equivalently written in the block form appearing in 27 . Just like in the one sample setting, we have a similar version of 2 for the two-sample setting under the same set of assumptions which is stated as follows.
Lemma 7. In the context of 3, it holds that \[\frac{\mathop{\mathrm{Var}}\Big(2\,\mathop{\mathrm{tr}}\Big( U^{\top} \begin{pmatrix} (Y^{(1)})^{\top}(Y^{(1)}) & -(Y^{(1)})^{\top}(Y^{(2)})\\[4pt] -(Y^{(2)})^{\top}(Y^{(1)}) & (Y^{(2)})^{\top}(Y^{(2)}) \end{pmatrix} U \Big)\Big)}{\mathop{\mathrm{Var}}\Big(2\,\mathop{\mathrm{tr}}\Big( U^{\top} \begin{pmatrix} {X^{(1)}}^{\top}{X^{(1)}} & -{X^{(1)}}^{\top}{X^{(2)}}\\[4pt] -{X^{(2)}}^{\top}{X^{(1)}} & {X^{(2)}}^{\top}{X^{(2)}} \end{pmatrix} U \Big)\Big)} \;\longrightarrow\; 1,\] where \(U = \begin{pmatrix} \beta_1^{-1} & \beta_2^{-1} \end{pmatrix}\) and \(Y^{(i)} = \beta_i^{\perp} {X^{(i)}}\) for \(i = 1,2\).
Proof. See 12.1. ◻
We similarly calculate the mean and variance of the second-order contribution.
Lemma 8 (Second–order term: mean and variance of the two-sample test statistic). Define the second–order contribution \[\begin{align} T_{2}^{(S)} &= 2\bigl\|\beta^{\perp} {X^{(1)}} \beta_1^{-1}\bigr\|_F^2 + 2\bigl\|\beta^{\perp} {X^{(2)}} \beta_2^{-1}\bigr\|_F^2 - 2\, \mathop{\mathrm{tr}}\bigl(S_{1,1}({X^{(1)}})\, S_{2,1}({X^{(2)}})\bigr) \\ &= 2\,\mathop{\mathrm{tr}}\left( U^{\top} \begin{pmatrix} (Y^{(1)})^{\top}(Y^{(1)}) & -(Y^{(1)})^{\top}(Y^{(2)})\\[4pt] -(Y^{(2)})^{\top}(Y^{(1)}) & (Y^{(2)})^{\top}(Y^{(2)}) \end{pmatrix} U \right), \end{align}\] where \(Y^{(i)} = \beta^{\perp} {X^{(i)}}\) and \(U = \begin{pmatrix} \beta_1^{-1} & \beta_2^{-1} \end{pmatrix}\). Assume that 1 2 3 4 hold. Recall the definitions from 8 \[\begin{align} &\quad d_i = \mathrm{diag}(\beta_i^{\perp}), \quad G = \beta_1^{-1}\beta_2^{-1} + \beta_2^{-1}\beta_1^{-1}, \quad {\Sigma^{(i)}} = P^{(i)} \circ (J_n - P^{(i)}), \\ &\quad K^{(i)} = 2\,{\Sigma^{(i)}} \circ (J_n - 2P^{(i)}), \quad \alpha_i = \mathrm{diag}(\beta_i^{-2}). \end{align}\] Then the mean and variance of \(T_{2}^{(S)}\) satisfy \[\mathbb{E}\bigl[T_{2}^{(S)}\bigr] = 2\sum_{i=1}^{2} \mathop{\mathrm{tr}}\Bigl[ \beta_i^{-2} \Bigl( {\Sigma^{(i)}} \circ \beta^{\perp} + \mathop{\mathrm{Diag}}({\Sigma^{(i)}} \cdot d_i - \mathrm{diag}({\Sigma^{(i)}}) \circ d_i) \Bigr) \Bigr],\] and \[\mathop{\mathrm{Var}}\bigl(T_{2}^{(S)}\bigr) = \sum_{i=1}^{2} \Bigl( 8\,\bigl\langle (J_n - I_n)\circ(\beta_i^{-2}\circ\beta_i^{-2}),\;({\Sigma^{(i)}})^2 \bigr\rangle + 4\,\alpha_i^{\top}K^{(i)} \alpha_i \Bigr) + 4\,\bigl\langle G \circ G,\; {\Sigma^{(1)}}^{\top}{\Sigma^{(2)}} + (J_n - I_n) \circ ({\Sigma^{(1)}} \circ {\Sigma^{(2)}}) \bigr\rangle + o \left( \frac{k^2}{n^3 \rho_n^2}\right).\] In particular, we also have \(\mathop{\mathrm{Var}}\bigl(T_{2}^{(S)}\bigr) \;\asymp\; \frac{k^{2}}{n^{3}\rho_n^{2}}.\)
Proof. See 12.2. ◻
Combining these moment formulae with the martingale central limit argument yields the Gaussian limit for the centered and scaled second–order term.
Lemma 9 (Second–order asymptotic distribution of the two-sample test). Suppose 1 2 3 4 hold. The quantity \(T_{2}^{(S)}\) satisfies \[\frac{ T_{2}^{(S)} - \mathbb{E}[T_{2}^{(S)}] }{ \sqrt{\mathop{\mathrm{Var}}(T_{2}^{(S)})} } \xrightarrow{D} \mathcal{N}(0,1).\]
Proof. See 12.3. ◻
This completes the derivation of the Gaussian limit for the two–sample second–order contribution and provides the principal stochastic ingredient in the proof of 1.
Like in the proof of 5, let \(T_{2}^{(H)}\) denote the higher–order contributions in the asymptotic expansion of the test statistic: \[T_{2}^{(H)} = \sum_{l \geq 3} T_{2}^{(l)} = \sum_{l \geq 3} \left( \langle {V^{(1)}} {V^{(1)}}^{\top}, S_{1,l}({X^{(1)}}) \rangle + \langle {V^{(2)}} {V^{(2)}}^{\top}, S_{2,l}({X^{(2)}}) \rangle + \sum_{l_1 + l_2 = l} \langle S_{1,l_1}({X^{(1)}}), S_{2,l_2}({X^{(2)}}) \rangle \right).\] As in the one-sample setting, we first separate out the third-order contribution. The following result bounds its mean.
Lemma 10. Assume 1 2 3 4. Let \(S_{i,3}\) denote the third–order remainder terms in 26 . Then we have \[\mathbb{E}\langle V V^{\top},S_{i,3}(X)\rangle = o \left(\frac{k}{n^2 \rho_n^2}\right).\]
Proof. See 12.4. ◻
Next we bound the variance of the third–order contribution. By 6 we already have tight control of the pure projection term \(\mathop{\mathrm{Var}}\big(\langle VV^{\top},S_{i,3}(X)\rangle\big)\), so it suffices to control the mixed or cross–term variances of the form \(\mathop{\mathrm{Var}}\big(\langle S_{1,l_1}(X),S_{2,l_2}(X)\rangle\big)\) for \(l_1,l_2\ge 1,\) because these are the only remaining contributions appearing in the expansion of the third–order part. The next lemma provides the required uniform bound on these cross–term variances, which together with 6, yields the desired bound on the overall third–order variance.
Lemma 11. Assume that 1 2 3 4 hold. Then \(\mathop{\mathrm{Var}}(\langle S_{1,l_1}(X),S_{2,l_2}(X)\rangle) \asymp \frac{k^2}{n^4 \rho_n^3}\) where \(l_1+l_2=3\) and \(l_1,l_2 \in \mathbb{Z}^{+}.\)
Proof. See 12.5. ◻
A simple application of Chebyshev’s inequality, as in 24 , yields the following result in probability: \[\frac{\left|T_{2}^{(3)} - \mathbb{E}\left(T_{2}^{(3)}\right)\right|}{\sqrt{\mathop{\mathrm{Var}}\left(T_{2}^{(S)}\right)}} \to 0.\] For the higher order terms, we follow the same proof as in one-sample case, which together with the asymptotic normality established in 9, yields the conclusion of 1.
To prove these result we start with the following lemma that shows the empirical eigenvectors are sufficiently incoherent.
Lemma 12. Let \(\mathcal{E}_{{\sf good}}\) be the event from 1 and define the event \[\mathcal{E}_{{\sf very \;good}} = \mathcal{E}_{{\sf good}} \cap \bigg\{\\ \|\widehat{V}\|_{2,\infty} \lesssim \sqrt{\frac{k}{n}}\,\bigg\}.\] Then \(\mathbb{P}(\mathcal{E}_{{\sf very \;good}})\ge1 - O(n^{-19}).\)
Proof. See 13.1. ◻
The event \(\mathcal{E}_{{\sf very \;good}}\) defines a regime where both the spectral norm bound \(\|X\| \lesssim \sqrt{n \rho_n}\) and the incoherence condition on \(\widehat V\) hold simultaneously. These are crucial for controlling higher-order error terms and ensuring the consistency of the proposed estimator.
Our estimator consistency results are built upon the following series expansion for the difference \({\widehat P} - P\), where \({\widehat P}\) is the rank \(k\) approximation of \(A\).
Lemma 13 (Series expansion for the low-rank estimator). Let \({\widehat P}\) denote the best (with respect to Frobenius norm) rank-\(k\) approximation of \(A\). Then under 1 2 3 4, \({\widehat P}-P\) admits the following series expansion: \[{\widehat P} - P = \sum_{l=1}^{\infty} T_{l}(X),\] where the first-order term is \[T_{1}(X) = V V^{\top}X (I - V V^{\top})+(I-V V^{\top}) X V V^{\top}+V V^{\top}X V V^{\top},\] and, for each integer \(l\ge 2\), \[\begin{align} T_{l}(X) &= S_{l}(X)\,P + P\,S_{l}(X) + \sum_{l_1+l_2=l} S_{l_1}(X)\,P\,S_{l_2}(X)\\ &\qquad + S_{l-1}(X)\,X\,V V^{\top}+ V V^{\top}X S_{l-1}(X) + \sum_{l_1+l_2=l-1} S_{l_1}(X)\,X\,S_{l_2}(X). \end{align}\] Here \(S_{l}(X)\) are the polynomial operators defined as in 21 . Moreover, the higher-order terms satisfy the operator-norm bound \[\|T_{l}(X)\| \lesssim c_l \left(\frac{\|X\|}{n\rho_n}\right)^{\!l}\,(n\rho_n),\] for constants \(c_l>0\) such that \(\log c_l = O(1)\).
Proof. See 13.2. ◻
We now study the plug-in estimates we use in the definition of our mean estimator.
Lemma 14. Define \(\Sigma, \widehat{\Sigma}\) as in 5 20 . Define the matrices \[M := \mathop{\mathrm{Diag}}(\Sigma \mathbf{1}_n), \quad \widehat{M} := \mathop{\mathrm{Diag}}(\widehat{\Sigma} \mathbf{1}_n), \quad \Delta M := M - \widehat{M}, \quad \Delta\mathop{\mathrm{Diag}}(\Sigma) := \mathop{\mathrm{Diag}}(\Sigma) - \mathop{\mathrm{Diag}}(\widehat{\Sigma}).\] Then, under 1 5 3 4, the following bounds hold: \[\begin{align} \big| \mathop{\mathrm{tr}}\bigl(VV^{\top}\,\Delta M\bigr) \big| &= O_p\bigl(\max (k,k \sqrt{\rho_n \log n})\bigr); \tag{28} \\ \big| \mathop{\mathrm{tr}}\bigl(VV^{\top}\,\Delta\mathop{\mathrm{Diag}}(\Sigma)\bigr)\big| &= O_p\left(\max \left(\frac{k}{n}, \frac{k \sqrt{\rho_n \log n}}{n}\right)\right). \tag{29} \end{align}\]
Proof. See 13.3. ◻
Lemma 15. Defining the quantities as in 14, the following bounds hold: \[\begin{align} \big| \mathop{\mathrm{tr}}\bigl((V \Lambda^{-2} V^{\top}-\widehat V \widehat \Lambda^{-2} \widehat V^{\top})\,\mathop{\mathrm{Diag}}(\widehat{\Sigma}\bigr)\big|&=O_p \left(\max \left(\frac{k}{n^3 \rho_n^2},\frac{k \sqrt{\log n}}{n^3 \rho_n^{3/2}}\right)\right);\tag{30}\\ \big| \mathop{\mathrm{tr}}\bigl((V \Lambda^{-2} V^{\top}-\widehat V \widehat \Lambda^{-2} \widehat V^{\top})\,\mathop{\mathrm{Diag}}(\widehat{\Sigma}\cdot \mathbf{1}_n)\bigr)\big|&=O_p \left(\max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right)\right).\tag{31} \end{align}\]
Proof. See 13.4. ◻
Next, we will study the plug-in estimates we use in the definition of our variance estimator.
Lemma 16. Define the vectors \(\alpha := \mathrm{diag}(\beta^{-2})\) and \(\widehat{\alpha} := \mathrm{diag}(\widehat{\beta}^{-2})\). Then, under 1 5 3 4, the following bounds hold: \[\begin{align} \|\widehat{\alpha}-\alpha\| &=O_p\left( \frac{k}{n^{3}\rho_n^{2.5}}\right);\tag{32}\\ \big\|\widehat{G}\circ\widehat{G}-G\circ G\big\|_F &=O_p\left( \frac{k^{3/2}}{n^{5.5}\rho_n^{4.5}}\right);\tag{33}\\ \left\| \widehat\beta^{-2} - \beta^{-2} \right\|_F &=O_p\left(\frac{k^{1/2}}{n^{2.5}\rho_n^{2.5}}\right).\tag{34} \end{align}\] where \(G\) and \(\widehat{G}\) are as defined in 1 12 .
Proof. See 13.5. ◻
Lemma 17. Defining the quantities as in 13, the following bounds hold: \[\begin{align} \left\| \widehat{P}^2 - P^2 \right\|_F &=O_p\left( k^{1/2} n^{3/2} \rho_n^{3/2}\right);\tag{35}\\ \left\| \widehat{P} ({\widehat P} \circ {\widehat P}) - P ( P \circ P) \right\|_F &=O_p\left( k^{1/2} n^{3/2} \rho_n^{5/2}\right);\tag{36}\\ \left\| ({\widehat P} \circ \widehat{P})^2 - ( P \circ P)^2 \right\|_F &=O_p\left( k^{1/2} n^{1.5} \rho_n^{3.5}\right).\tag{37} \end{align}\]
Proof. See 13.6. ◻
We first establish the consistency of the one–sample estimators, which will subsequently be used to derive the consistency of the two–sample estimators. Throughout the proof, we work on the event \(\mathcal{E}_{{\sf very \;good}}\) which has been defined in 12.
Proof. We now establish 6, which concerns the consistency of the mean and variance estimators in the one–sample setting under 5.
Analyzing the estimated mean. From 16 , the population mean admits the expansion \[\mu_1
=\underbrace{2\mathop{\mathrm{tr}}\bigl[\beta^{-2}\bigl(\Sigma \circ\beta^\perp\bigr)\bigr]}_{\mu_1^{(1)}}
+\underbrace{2\mathop{\mathrm{tr}}\bigl[\beta^{-2}\bigl(\mathop{\mathrm{Diag}}(\Sigma\cdot d)\bigr)\bigr]}_{\mu_1^{(2)}}
+ O_p\left(\frac{k}{n^2\rho_n^2}\right),\] where the remainder term vanishes in 5. Note that the original mean expression in 16 contained an
additional term \(2\,\mathrm{tr}\Bigl(\beta^{-2}\left[\mathop{\mathrm{Diag}}(\mathrm{diag}(\Sigma)\circ d)\right]\Bigr)\), which is absent in the above expression. This is because \[\begin{align}
\mathrm{tr}\Bigl(\beta^{-2}\left[\mathop{\mathrm{Diag}}(\mathrm{diag}(\Sigma)\circ d)\right]\Bigr)
&\lesssim \mathrm{tr}\Bigl(\beta^{-2}\left[\mathop{\mathrm{Diag}}(\mathrm{diag}(\Sigma)\right]\Bigr) + \mathrm{tr}\Bigl(\beta^{-2}\left[\mathop{\mathrm{Diag}}\left(\mathrm{diag}(\Sigma)\circ \mathrm{diag}(VV^{\top})\right)\right]\Bigr)\\
& \lesssim \rho_n\,\mathrm{tr}(\beta^{-2}) \\
& \lesssim \frac{k}{n^2\rho_n},
\end{align}\] where the second line follows from 4 1. Hence, this
term is absorbed within the remainder term \(O_p\left(\tfrac{k}{n^2\rho_n^2}\right)\). The plug–in estimator of \(\mu_1\), defined in 18 , is given by
\[\widehat{\mu}_1
=\underbrace{2\mathop{\mathrm{tr}}\Bigl[\widehat{\beta}^{-2}\bigl(\widehat{\Sigma}\circ\widehat{\beta}^\perp\bigr)\Bigr]}_{\widehat{\mu}_1^{(1)}}
+\underbrace{2\mathop{\mathrm{tr}}\Bigl[\widehat{\beta}^{-2}\mathop{\mathrm{Diag}}\bigl(\widehat{\Sigma}\cdot\widehat{d}\bigr)\Bigr]}_{\widehat{\mu}_1^{(2)}},\] where all notation follows that of 6. To control the estimation error, we first consider the leading component \(\mu_1^{(1)}\) and decompose the difference as \[\begin{align}
\mu_1^{(1)} - \widehat{\mu}_1^{(1)}
&= \mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\cdot \Sigma\circ\beta^\perp - V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\Sigma) \right) \\
&\quad + \mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\Sigma) - V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma}) \right) \\
&\quad + \mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma}) - \widehat{V} \widehat{\Lambda}^{-2} \widehat{V}^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma}) \right) \\
&\quad + \mathop{\mathrm{tr}}\left( \widehat{V} \widehat{\Lambda}^{-2} \widehat{V}^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma}) - \widehat{V} \widehat{\Lambda}^{-2} \widehat{V}^{\top}\cdot \widehat{\Sigma}\circ\widehat{\beta}^\perp \right).
\end{align}\]
An analogous decomposition applies to \(\mu_1^{(2)}\): \[\begin{align} \mu_1^{(2)} - \widehat{\mu}_1^{(2)} &= \mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\Sigma \cdot d) - V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\Sigma \cdot \mathbf{1}_n) \right) \\ &\quad + \mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\Sigma \cdot \mathbf{1}_n) - V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma} \cdot \mathbf{1}_n) \right) \\ &\quad + \mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma} \cdot \mathbf{1}_n) - \widehat{V} \widehat{\Lambda}^{-2} \widehat{V}^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma} \cdot \mathbf{1}_n) \right)\\ &\quad + \mathop{\mathrm{tr}}\left( \widehat{V} \widehat{\Lambda}^{-2} \widehat{V}^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma} \cdot \mathbf{1}_n) - \widehat{V} \widehat{\Lambda}^{-2} \widehat{V}^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma} \cdot \widehat{d}) \right), \end{align}\] where \(d=\mathrm{diag}\left(I-V V^{\top}\right)\) and \(\widehat{d}=\mathrm{diag}\left(I-\widehat{V} \widehat{V}^{\top}\right)\).
Bounding \(|\mu_1^{(1)} - \widehat\mu_1^{(1)} |\). We first establish that \[\begin{align} \label{mu1bound} |\mu_1^{(1)} - \widehat{\mu}_1^{(1)} |=o_p\left(\max \left(\frac{k}{n^2 \rho_n^2}, \frac{k \sqrt{\rho_n \log n}}{n^2 \rho_n^2}\right)\right). \end{align}\tag{38}\] Consider \[\begin{align} \label{eq:ab1} \big|\mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\Sigma \circ \beta^\perp - V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\Sigma) \right)\big| &= |\mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\Sigma \circ (V V^{\top}) \right)| \\ &\lesssim \frac{1}{n^2 \rho_n^2}\, |\mathop{\mathrm{tr}}\left( V V^{\top}\cdot \Sigma \circ (V V^{\top}) \right)|\\ &\lesssim \frac{k}{n^2 \rho_n^2}\, \|V V^{\top}\|_F\, \|\Sigma \circ (V V^{\top})\|_F\\ &\lesssim \frac{k}{n^3 \rho_n^2}\, \|V V^{\top}\|_F\, \|\Sigma\|_F\\ &\lesssim \frac{k^{3/2}}{n^2 \rho_n} \\ &=o_p\left(\max \left(\frac{k}{n^2 \rho_n^2}, \frac{k \sqrt{\rho_n \log n}}{n^2 \rho_n^2}\right)\right), \end{align}\tag{39}\] where the fourth line follows from 4 and the final line follows from 1.
Next, we bound the term involving the deviation between \(\Sigma\) and \(\widehat{\Sigma}\): \[\begin{align} \big|\mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\Sigma - \widehat{\Sigma}) \right)\big| &\lesssim \frac{1}{(n\rho_n)^2}\, \big|\mathop{\mathrm{tr}}\left( V V^{\top}\mathop{\mathrm{Diag}}(\Sigma - \widehat{\Sigma}) \right)\big|\\ &=O_p\left(\max \left(\frac{k}{n^3 \rho_n^2}, \frac{k \sqrt{\rho_n \log n}}{n^3 \rho_n^2}\right)\right)\\ &=o_p\left(\max \left(\frac{k}{n^2 \rho_n^2}, \frac{k \sqrt{\rho_n \log n}}{n^2 \rho_n^2}\right)\right), \end{align}\] where the second line uses 29 of 14.
We then control the eigenvalue and eigenvector perturbation component directly using 30 of 15 via \[\begin{align} \big|\mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma}) - \widehat V \widehat\Lambda^{-2} \widehat V^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma}) \right)\big| &=O_p \left(\max \left(\frac{k}{n^3 \rho_n^2},\frac{k \log n}{n^3 \rho_n^{3/2}}\right)\right)=o_p\left(\max \left(\frac{k}{n^2 \rho_n^2}, \frac{k \sqrt{\rho_n \log n}}{n^2 \rho_n^2}\right)\right). \end{align}\] Finally, for the term \[\mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma}) - V \Lambda^{-2} V^{\top}\cdot \widehat{\Sigma}\circ\beta^\perp \right)=o_p\left(\max \left(\frac{k}{n^2 \rho_n^2}, \frac{k \sqrt{\rho_n \log n}}{n^2 \rho_n^2}\right)\right),\] the same reasoning as in 39 follows, with the incoherence of \(\widehat{V}\) (guaranteed by 12) replacing that of \(V\).
Combining all four bounds above yields the desired rate in 38 .
Bounding \(| \mu_1^{(2)} - \widehat\mu_1^{(2)} |\). We next establish that \[\begin{align} \label{mu12bound} |\mu_1^{(2)} - \widehat{\mu}_1^{(2)}| =O_p\left(\max \left(\frac{k}{n^2 \rho_n^2}, \frac{k \sqrt{\rho_n\log n} }{n^2 \rho_n^2}\right)\right). \end{align}\tag{40}\] Consider first \[\begin{align}\label{eq:ab2} \big | \mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\Sigma \cdot d) - V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\Sigma \cdot \mathbf{1}_n) \right) \big | &= \mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\Sigma \cdot \mathrm{diag}(V V^{\top})) \right) \\ &\lesssim \frac{1}{n^2 \rho_n^2} \mathop{\mathrm{tr}}\left( V V^{\top}\mathop{\mathrm{Diag}}(\Sigma \cdot \mathrm{diag}(V V^{\top})) \right) \\ &\lesssim \frac{1}{n^2 \rho_n^2} \left( \frac{k}{n} \right) \mathop{\mathrm{tr}}\Big( \mathop{\mathrm{Diag}}(\Sigma \cdot \mathrm{diag}(V V^{\top})) \Big) \\ &\lesssim \frac{k}{n^3 \rho_n^2}\, \big( \mathbf{1}_n^{\top}\cdot \Sigma \cdot \mathrm{diag}(V V^{\top}) \big) \\ &\lesssim \frac{k}{n^3 \rho_n^2} \left( \frac{k}{n} \right) \big( \mathbf{1}_n^{\top}\cdot \Sigma \cdot \mathbf{1}_n \big) \\ &\lesssim \frac{k^2}{n^4 \rho_n^2} (n^2 \rho_n) \\ &= \frac{k^2}{n^2 \rho_n}\\ &= o_p\left(\max \left(\frac{k}{n^2 \rho_n^2}, \frac{k \sqrt{\rho_n \log n}}{n^2 \rho_n^2}\right)\right), \end{align}\tag{41}\] where we have used 4 and the final line follows from 1.
Next, for the perturbation in \(\Sigma\): \[\begin{align} \big| \mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}(\mathop{\mathrm{Diag}}(\Sigma \cdot \mathbf{1}_n) - \mathop{\mathrm{Diag}}(\widehat{\Sigma} \cdot \mathbf{1}_n)) \right) \big| &\lesssim \frac{1}{(n\rho_n)^2} \big| \mathop{\mathrm{tr}}\left( V V^{\top}(\mathop{\mathrm{Diag}}(\Sigma \cdot \mathbf{1}_n) - \mathop{\mathrm{Diag}}(\widehat{\Sigma} \cdot \mathbf{1}_n)) \right) \big| \\ &=O_p\left(\max \left(\frac{k}{n^2 \rho_n^2}, \frac{k \sqrt{\rho_n\log n} }{n^2 \rho_n^2}\right)\right), \end{align}\] where the last line follows from 28 of 14.
We then control the eigenvalue and eigenvector perturbation component directly using 31 of 15: \[\begin{align} &\Big|\mathop{\mathrm{tr}}\big( V\Lambda^{-2}V^{\top}\mathop{\mathrm{Diag}}(\widehat\Sigma \cdot \mathbf{1}_n)-\widehat V\widehat\Lambda^{-2}\widehat V^{\top}\mathop{\mathrm{Diag}}(\widehat\Sigma \cdot \mathbf{1}_n)\big)\Big| =O_p\left(\max \left(\frac{k}{n^2 \rho_n^2}, \frac{k \sqrt{\rho_n\log n} }{n^2 \rho_n^2}\right)\right). \end{align}\] Finally, for the term \(\mathop{\mathrm{tr}}\left( \widehat{V} \widehat{\Lambda}^{-2} \widehat{V}^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma} \cdot \mathbf{1}_n) - \widehat{V} \widehat{\Lambda}^{-2} \widehat{V}^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma} \cdot \widehat{d}) \right),\) we have the same argument as in 41 , using incoherence of \(\widehat V\) which is guaranteed on the event \(\mathcal{E}_{{\sf very \;good}}\) as per 12. Combining all four bounds yields the desired rate in 40 .
Thus, from 38 and 40 we have \[\begin{align}
\label{mubd} \big|\widehat{\mu}_1 - \mu_1 \big| =O_p\left(\max \left(\frac{k}{n^2 \rho_n^2}, \frac{k \sqrt{\rho_n\log n} }{n^2 \rho_n^2}\right)\right),
\end{align}\tag{42}\] under 1 5 3 4.
Analyzing the Estimated Variance. From 17 , the population variance admits the expansion \[\begin{align}
\sigma_1^2 = 8\,\underbrace{\Bigl\langle (J_n - I_n)\circ(\beta^{-2}\circ\beta^{-2}),\Sigma^2\Bigr\rangle}_{(\sigma^{11})^{2}} + 4\,\underbrace{\alpha^{\top}K\,\alpha}_{(\sigma^{12})^{2}} + o\left(\frac{k^2}{n^3 \rho_n^2}\right).
\end{align}\] The corresponding plug–in estimator, as defined in 18 , takes the form \[\widehat{\sigma}_1^2
= 8\,\underbrace{\Bigl\langle (J_n - I_n)\circ(\widehat{\beta}^{-2}\circ\widehat{\beta}^{-2}),\widehat{\Sigma}^2\Bigr\rangle}_{(\widehat\sigma^{11})^{2}} +
4\,\underbrace{\widehat{\alpha}^{\top}\widehat{K}\,\widehat{\alpha}}_{(\widehat\sigma^{12})^{2}},\] where all notation follows that of 6. To analyze the
estimation error, we first focus on the first component \(\left(\sigma^{11}\right)^{2}\) and decompose the difference as \[\begin{align} (\widehat\sigma^{11})^{2} - ( \sigma^{11})^{2} &=
\underbrace{\left\langle (J_n - I_n)\circ(\widehat{\beta}^{-2}\circ\widehat{\beta}^{-2}),\,\widehat{\Sigma}^2 - \Sigma^2 \right\rangle}_{I_1} + \underbrace{\left\langle (J_n - I_n)\circ(\widehat{\beta}^{-2}\circ\widehat{\beta}^{-2} -
\beta^{-2}\circ\beta^{-2}),\,\Sigma^2 \right\rangle}_{I_2}.
\end{align}\] An analogous decomposition applies to \(( \sigma^{12})^{2}\): \[\begin{align}
(\widehat\sigma^{12})^{2} - ( \sigma^{12})^{2}
&= \alpha^{\top}\widehat{K} \alpha - \widehat{\alpha}^{\top}K \widehat{\alpha} \nonumber \\
&= \alpha^{\top}(\widehat{K} - K)\alpha + 2\,\alpha^{\top}\widehat{K} (\widehat{\alpha} - \alpha) + (\widehat{\alpha} - \alpha)^{\top}\widehat{K} (\widehat{\alpha} - \alpha). \label{sigma12decomposition}
\end{align}\tag{43}\]
Bounding \(|(\widehat\sigma^{11})^{2} - ( \sigma^{11})^{2}|\). We will establish that \[\begin{align}\label{sigma11bound} \big|(\widehat\sigma^{11})^{2} - ( \sigma^{11})^{2}\big| =o_p \left( \frac{k^2}{n^3 \rho_n^2}\right). \end{align}\tag{44}\] Applying the Cauchy–Schwarz inequality and noting that \(\|(J_n-I_n)\circ M\|_F \le \|M\|_F\), we obtain \[\begin{align} \label{eq:i1} \big| I_1 \big| &= \left\langle (J_n - I_n)\circ(\widehat{\beta}^{-2}\circ\widehat{\beta}^{-2}),\,\widehat{\Sigma}^2 - \Sigma^2 \right\rangle \\ &\le \left\| (J_n - I_n)\circ(\widehat{\beta}^{-2}\circ\widehat{\beta}^{-2}) \right\|_F \left\| \widehat{\Sigma}^2 - \Sigma^2 \right\|_F \\ &\lesssim \left\| \widehat{\beta}^{-2}\circ\widehat{\beta}^{-2} \right\|_F \left\| \widehat{\Sigma}^2 - \Sigma^2 \right\|_F \\ &\lesssim \left\| \widehat{\beta}^{-2}\circ\widehat{\beta}^{-2} \right\|_F \left( \left\| \widehat{P}^2 - P^2 \right\|_F + 2 \left\| \widehat{P} ({\widehat P} \circ {\widehat P}) - P ( P \circ P) \right\|_F + \left\| ({\widehat P} \circ \widehat{P})^2 - ( P \circ P)^2 \right\|_F\right)\\ &\lesssim \frac{1}{n^4 \rho_n^4} \left\| ((\widehat{V}^{(i)}) (\widehat{V}^{(i)})^{\top}) \circ ((\widehat{V}^{(i)}) (\widehat{V}^{(i)})^{\top}) \right\|_F \left( \left\| \widehat{P}^2 - P^2 \right\|_F\right) \\ &=O_p \left( \frac{(k/n)}{n^4 \rho_n^4}\, k^{1/2} (n\rho_n)^{1.5} \right)\\ &= o_p \left( \frac{k^{3/2}}{n^3 \rho_n^2} \right). \end{align}\tag{45}\] Here, the third to last line and the second to last line follow from 35 36 37 in 17. The bound \(\|((\widehat{V}^{(i)}) (\widehat{V}^{(i)})^{\top}) \circ ((\widehat{V}^{(i)}) (\widehat{V}^{(i)})^{\top})\|_F \asymp k/n\) in the second to last line is ensured by 4.
Note the following bounds: \[\begin{align} \label{note01} \|\widehat{\beta}^{-2}+\beta^{-2}\|_{\infty} \le \|\widehat{\beta}^{-2}\|_{\infty} + \|\beta^{-2}\|_{\infty}=O_p \left(\frac{1}{n^3 \rho_n^2} \right), \end{align}\tag{46}\] \[\begin{align}\label{note02} \|\Sigma^2\|_F &\le \|\Sigma^2-P^2\|_F+\|P^2\|_F \\ &\le \|(\Sigma-P)\Sigma+P(\Sigma-P) \|_F+\|P^2\|_F\\ &\lesssim n\rho_n\|\Sigma-P\|_F+k n^2 \rho_n^2\\ &\lesssim n\rho_n\|P\circ P\|_F+k n^2 \rho_n^2\\ &\lesssim n\rho_n^2\|P\|_F+k n^2 \rho_n^2\\ &\lesssim k n^2 \rho_n^2. \end{align}\tag{47}\] Now, consider the term \(I_2\). We have \[\begin{align}\label{eq:i2} \big| I_2 \big| &= \Big| \big\langle (J_n-I_n)\circ(\widehat{\beta}^{-2}\circ\widehat{\beta}^{-2}-\beta^{-2}\circ\beta^{-2}), \,\Sigma^2 \big\rangle \Big| \\[4pt] &\lesssim \Big| \big\langle (\widehat{\beta}^{-2}-\beta^{-2})\circ(\widehat{\beta}^{-2}+\beta^{-2}), \,\Sigma^2 \big\rangle \Big| \\[4pt] &\lesssim \| (\widehat{\beta}^{-2}-\beta^{-2})\circ(\widehat{\beta}^{-2}+\beta^{-2}) \|_F\, \|\Sigma^2\|_F \\[4pt] &\lesssim \|\Sigma^2\|_F\, \|\widehat{\beta}^{-2}+\beta^{-2}\|_{\infty}\, \|\widehat{\beta}^{-2}-\beta^{-2}\|_F \\[4pt] &\lesssim k n^2 \rho_n^2 \cdot \|\widehat{\beta}^{-2}+\beta^{-2}\|_{\infty}\, \|\widehat{\beta}^{-2}-\beta^{-2}\|_F \\[4pt] &=O_p \left( \frac{k^{3/2}}{n^{7/2}\rho_n^{5/2}} \right)\\ &= o_p \left(\frac{k^{3/2}}{n^3\rho_n^2}\right). \end{align}\tag{48}\] The third to last line and the second to last line follow from 34 in 16 and 46 47 . Combining 45 48 yields 44 .
Bounding \(|(\widehat\sigma^{12})^{2} - ( \sigma^{12})^{2}|\). We will now establish that \[\begin{align} |\widehat{\sigma}_{12}^2 - \sigma_{12}^2 |=o_p \left( \frac{k^2}{n^3 \rho_n^2}\right). \label{sigma12bound} \end{align}\tag{49}\] Recalling that \(\alpha = \mathrm{diag}(\beta^{-2})\), we have that \(\|\alpha\|_{\infty} \asymp \frac{k}{n^3 \rho_n^{2}},\) and \(\|\alpha\|_{2} \asymp \frac{k}{n^{2.5} \rho_n^{2}}.\) For the first term on the right hand side of 43 , we have \[\begin{align}\label{kbd01} |{\alpha}^{\top}(\widehat{K} - K) {\alpha}| &\lesssim \|\alpha\|^2 \cdot \|\widehat{K} - K\| \\ &\lesssim \left( \frac{k^2}{n^5 \rho_n^4} \right) \cdot \|P - {\widehat P}\| \\ &=O_p \left( \frac{k^2}{n^{4.5} \rho_n^{3.5}} \right) = o_p \left( \frac{k^2}{n^3 \rho_n^2} \right), \end{align}\tag{50}\] which holds on the event \(\mathcal{E}_{{\sf good}}\) as explained in 1. To establish the bound \(\|\widehat K-K\| \lesssim \|P-{\widehat P}\|\), we first observe that \[\|\widehat K-K\| \lesssim \|\widehat\Sigma-\Sigma\| + \|\widehat\Sigma\circ\widehat P-\Sigma\circ P\|.\] By the triangle inequality, the second term on the right-hand side satisfies \[\begin{align} \label{note03} \big\|\widehat\Sigma\circ\widehat P-\Sigma\circ P \big\| \le \big\|\widehat\Sigma \circ (\widehat P- P)\big\| + \big\| P \circ (\widehat \Sigma- \Sigma)\big\|. \end{align}\tag{51}\] For the second term on the right hand side of 51 , 4 3 1 imply \[\label{schur01} \begin{align} \big\| P \circ (\widehat \Sigma- \Sigma)\big\| &= \bigg\|\sum_{l=1}^{k} \lambda_l \mathop{\mathrm{Diag}}(V_{l.}) (\widehat \Sigma- \Sigma) \mathop{\mathrm{Diag}}(V_{l.})\bigg\| \\ &\le \left(\sum_{l=1}^{k} \lambda_l \|V_{l.}\|_{\infty}^2 \right) \big\|\widehat \Sigma- \Sigma \big\| \\ &\le k^2 \rho_n \big\|\widehat \Sigma- \Sigma \big\| \ll \big\|\widehat \Sigma- \Sigma \big\|. \end{align}\tag{52}\] Similarly, the first term on the right hand side of 51 can be bounded as \[\label{schur02} \begin{align} \big\|\widehat\Sigma \circ (\widehat P- P)\big\| &\le \big\|\widehat P \circ (\widehat P- P)\big\| + \big\|(\widehat P \circ \widehat P) \circ (\widehat P- P)\big\| \\ &\ll \big\|\widehat P- P \big\| \end{align}\tag{53}\] with probability \(1-O(n^{-19})\), where the final step follows from 12, 3 1, and the fact that \(\mathop{\mathrm{rank}}(\widehat P \circ \widehat P) \le k^2\). Under the same assumptions, an analogous argument yields \(\|\widehat\Sigma-\Sigma\| \lesssim \|P-{\widehat P}\|\). Combining this with 52 and 53 establishes the desired bound \(\|\widehat K-K\| \lesssim \|P-{\widehat P}\|\).
For the cross term in 43 , it follows that \[\begin{align}\label{kbd02} |{\alpha}^{\top}\widehat{K} (\widehat{\alpha} - {\alpha})| &\leq \|\alpha\| \, \|\widehat{K}\| \, \|\widehat{\alpha} - {\alpha}\| \\ &=O_p \left( \frac{k^{2}}{n^{4.5} \rho_n^{3.5}} \right) = o_p \left( \frac{k^{2}}{n^3 \rho_n^2}\right), \end{align}\tag{54}\] where the second line uses 32 of 16.
For the quadratic term in 43 , again by 32 , \[\begin{align}\label{kbd03} |(\widehat{\alpha} - {\alpha})^{\top}\widehat{K} (\widehat{\alpha} - {\alpha})| &\leq \|\widehat{K}\| \cdot \|\widehat{\alpha} - {\alpha}\|^2 \\ &=O_p \left( n \rho_n \cdot \left( \frac{k}{n^{3} \rho_n^{2.5}} \right)^2 \right) =o_p \left( \frac{k^2}{n^3 \rho_n^2}\right). \end{align}\tag{55}\] Combining 50 54 55 43 yields \[\begin{align} \label{wivbmtyg} \big|(\widehat\sigma^{12})^{2} - ( \sigma^{12})^{2} \big| = o_p\left( \frac{k^{2}}{n^3 {\rho_n}^2} \right). \end{align}\tag{56}\]
Consequently, 49 and 44 imply that \[\begin{align} \label{sigbd} |\widehat{\sigma}_1^2 - \sigma_1^2| = o_p \left( \frac{k^2}{n^3 \rho_n^2} \right). \end{align}\tag{57}\]
3 implies that \(\widehat{\sigma}_1^2 \asymp \frac{k^{2}}{n^{3}\rho_n^{2}}\). Combining this with 57 , we obtain \(\frac{\widehat{\sigma}_1}{\sigma_1} \xrightarrow{p} 1.\) Furthermore, 5, 42 and 3 imply that \(\frac{\widehat{\mu}_1 - \mu_1}{\widehat{\sigma}_1} \xrightarrow{p} 0,\) which verifies the conditions required for the consistency of the estimator. ◻
We will next establish 3.
Proof. Computing the mean. The population mean for the two–sample statistic admits the decomposition \[\mu_2=2\sum_{i=1}^{2}\mu_2^{(i)} =2\sum_{i=1}^{2}\mathop{\mathrm{tr}}\Bigl[\beta_i^{-2}\Bigl({\Sigma^{(i)}}\circ\beta_i^\perp + \mathop{\mathrm{Diag}}({\Sigma^{(i)}}\cdot d_i)\Bigr)\Bigr],\] where \(\mu_2^{(i)}:=\mathop{\mathrm{tr}}\Bigl[\beta_i^{-2}\Bigl({\Sigma^{(i)}}\circ\beta_i^\perp + \mathop{\mathrm{Diag}}({\Sigma^{(i)}}\cdot d_i)\Bigr)\Bigr].\) The corresponding plug–in estimator introduced in 11 is \[\widehat{\mu}_2 = \sum_{i=1}^{2}\widehat{\mu}_2^{(i)} = 2\sum_{i=1}^{2}\mathop{\mathrm{tr}}\Bigl[\widehat{\beta}_i^{-2}\Bigl({\widehat{P}^{(i)}}\circ(J_n-{\widehat{P}^{(i)}})\circ\widehat\beta_i^\perp + \mathop{\mathrm{Diag}}\bigl(({\widehat{P}^{(i)}}\circ(J_n-{\widehat{P}^{(i)}}))\cdot\widehat d_i\bigr)\Bigr)\Bigr],\] with notation as in 3.
Since the two–sample mean decomposes as the sum of the two one–sample contributions, the consistency and rate established in 6 apply componentwise. In
particular, for each \(i\in\{1,2\}\) we have \(\big|\widehat{\mu}_2^{(i)}-\mu_2^{(i)}\big| = O_p\Bigl(\frac{k}{n^2\rho_n^2}\Bigr),\) and hence \[\begin{align}
\label{mu2bd} \big|\widehat{\mu}_2-\mu_2\big| = O_p \left(\max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right)\right),
\end{align}\tag{58}\] which establishes the claimed mean consistency for the two–sample statistic and thus proves the first part of 3.
Computing the Variance. From 2, the population variance for the two–sample statistic admits the decomposition \[\begin{align}
\sigma^2_2
&= \sum_{i=1}^{2} \bigl(\sigma_2^{(i)}\bigr)^2 + \sigma_2^{(1,2)} \\
&= \underbrace{\sum_{i=1}^{2} \!\left(
8 \,\bigl\langle (J_n - I_n)\!\circ\!(\beta_i^{-2}\!\circ\!\beta_i^{-2}),\,({\Sigma^{(i)}})^2 \bigr\rangle
+ 4\,\alpha_i^{\top}K^{(i)}\,\alpha_i \right)}_{\sum_{i=1}^{2} (\sigma_2^{(i)})^2}
+ \underbrace{4 \,\bigl\langle G\!\circ\!G,
{\Sigma^{(1)}}^{\top}{\Sigma^{(2)}} + (J_n-I_n)\!\circ\!({\Sigma^{(1)}}\!\circ\!{\Sigma^{(2)}}) \bigr\rangle}_{\sigma_2^{(1,2)}}.
\end{align}\] The corresponding plug–in estimator introduced in 11 takes the form \[\begin{align} \widehat{\sigma}^2_2 &= \sum_{i=1}^{2} \bigl(\widehat{\sigma}_2^{(i)}\bigr)^2 +
\widehat{\sigma}_2^{(1,2)} \\ &= \underbrace{ \begin{aligned} \sum_{i=1}^{2} \biggl[ \, & 8\,\bigl\langle (J_n - I_n)\circ(\widehat{\beta}_i^{-2}\circ\widehat{\beta}_i^{-2}), \,(\widehat{\Sigma}^{(i)})^2 \bigr\rangle + 4\,\widehat{\alpha}_i^\top
{\widehat{K}^{(i)}}\,\widehat{\alpha}_i \biggr] \end{aligned} }_{\sum_{i=1}^{2} (\widehat{\sigma}_2^{(i)})^2} \\ &\quad + \underbrace{ \begin{align} 4\,\bigl\langle \widehat{G}\circ\widehat{G}, (\widehat{\Sigma}^{(1)})^\top(\widehat{\Sigma}^{(2)}) +
(J_n-I_n)\circ((\widehat{\Sigma}^{(1)})\circ(\widehat{\Sigma}^{(2)}))\bigr\rangle \end{align} }_{\widehat{\sigma}_2^{(1,2)}}
\end{align}\] with notation as in 3.
From the one–sample variance consistency result as in 57 , it follows that \(| (\widehat{\sigma}_2^{(i)})^2 - (\sigma_2^{(i)})^2| \ll \frac{k^2}{n^3\rho_n^2},\) for \(i = 1,2\). To establish the full variance consistency of the two–sample statistic, it remains to verify that \(|\widehat{\sigma}_2^{(1,2)} - \sigma_2^{(1,2)}| =\; o_p \left(\frac{k^2}{n^3\rho_n^2}\right).\) We decompose the difference as \[\begin{align} \widehat{\sigma}_2^{(1,2)} - \sigma_2^{(1,2)} &= \Big\langle \widehat{G}\!\circ\!\widehat{G},\; (\widehat{\Sigma}^{(1)})^{\top}(\widehat{\Sigma}^{(2)}) + (J_n-I_n)\!\circ\!((\widehat{\Sigma}^{(1)})\!\circ\!(\widehat{\Sigma}^{(2)}))\Big\rangle - \Big\langle G\!\circ\!G,\; {\Sigma^{(1)}}^{\top}{\Sigma^{(2)}} + (J_n-I_n)\!\circ\!({\Sigma^{(1)}}\!\circ\!{\Sigma^{(2)}}) \Big\rangle \\ &= \underbrace{\Big\langle \widehat{G}\!\circ\!\widehat{G} - G \circ G,\; (\widehat{\Sigma}^{(1)})^{\top}(\widehat{\Sigma}^{(2)}) + (J_n-I_n)\!\circ\!((\widehat{\Sigma}^{(1)})\!\circ\!(\widehat{\Sigma}^{(2)}))\Big\rangle}_{Q_2} \\ &- \underbrace{\Big\langle G\!\circ\!G,\; ({\Sigma^{(1)}}^{\top}{\Sigma^{(2)}} - (\widehat{\Sigma}^{(1)})^{\top}(\widehat{\Sigma}^{(2)})) + (J_n-I_n)\!\circ\!({\Sigma^{(1)}}\!\circ\!{\Sigma^{(2)}} - (\widehat{\Sigma}^{(1)})\!\circ\!(\widehat{\Sigma}^{(2)})) \Big\rangle}_{Q_1}. \end{align}\] We first bound \(Q_2\). Observe that \[\begin{align}\label{newG1} \big| \langle \widehat{G}\!\circ\!\widehat{G} - G \circ G,\; (\widehat{\Sigma}^{(1)})^{\top}(\widehat{\Sigma}^{(2)}) + (J_n-I_n)\!\circ\!((\widehat{\Sigma}^{(1)})\!\circ\!(\widehat{\Sigma}^{(2)}))\rangle \big| &\lesssim \big| \langle \widehat{G}\!\circ\!\widehat{G} - G \circ G,\; (\widehat{\Sigma}^{(1)})^{\top}(\widehat{\Sigma}^{(2)}) \rangle \big| \\ &\lesssim \| \widehat{G}\!\circ\!\widehat{G} - G \circ G\|_F \|(\widehat{\Sigma}^{(1)})^{\top}(\widehat{\Sigma}^{(2)})\|_F \\ &\lesssim \sqrt{k} \| \widehat{G}\!\circ\!\widehat{G} - G \circ G\|_F \|(\widehat{\Sigma}^{(1)})\|_2 \|^{\top}(\widehat{\Sigma}^{(2)})\|_2 \\ &=O_p \left(\sqrt{k} \cdot \frac{k^{3/2}}{n^{5.5}\rho_n^{4.5}} \cdot n^2 \rho_n^2\right) \\ &=O_p \left( \frac{k^{2}}{n^{3.5}\rho_n^{2.5}} \right)\\ &= o_p \left(\frac{k^2}{n^3 \rho_n^2}\right), \end{align}\tag{59}\] where the fourth line follows from 33 of 16.
Next, we bound \(Q_1\). Observe that \[\begin{align}\label{newG} & \big| \langle G\!\circ\!G,\; ({\Sigma^{(1)}}{\Sigma^{(2)}} - (\widehat{\Sigma}^{(1)})(\widehat{\Sigma}^{(2)})) + (J_n-I_n)\!\circ\!({\Sigma^{(1)}}\!\circ\!{\Sigma^{(2)}} - (\widehat{\Sigma}^{(1)})\!\circ\!(\widehat{\Sigma}^{(2)})) \rangle \big|\\ \lesssim & \big| \langle G\!\circ\!G,\; {\Sigma^{(1)}}{\Sigma^{(2)}} - (\widehat{\Sigma}^{(1)})(\widehat{\Sigma}^{(2)}) \rangle \big| + \big| \langle G\!\circ\!G, ({\Sigma^{(1)}}\!\circ\!{\Sigma^{(2)}} - (\widehat{\Sigma}^{(1)})\!\circ\!(\widehat{\Sigma}^{(2)})) \rangle \big|. \end{align}\tag{60}\] By Hölder’s inequality for Frobenius inner products, \[\big| \langle G\!\circ\!G,\; {\Sigma^{(1)}}{\Sigma^{(2)}} - (\widehat{\Sigma}^{(1)})(\widehat{\Sigma}^{(2)}) \rangle \big| \le \|G\!\circ\!G\|_F\, \|{\widehat{P}^{(1)}}{\widehat{P}^{(2)}} - {P^{(1)}}{P^{(2)}}\|_F .\] Using the identity \({\widehat{P}^{(1)}}{\widehat{P}^{(2)}} - {P^{(1)}}{P^{(2)}} = {\widehat{P}^{(1)}}({\widehat{P}^{(2)}}-{P^{(2)}}) - ({\widehat{P}^{(1)}}-{P^{(1)}}){P^{(2)}} ,\) together with \(\|G\!\circ\!G\|_F \le \|G\|_\infty \|G\|_F\), we obtain \[\begin{align} \label{right1} \|G\!\circ\!G\|_F\, \|{\widehat{P}^{(1)}}{\widehat{P}^{(2)}} - {P^{(1)}}{P^{(2)}}\|_F &\lesssim \|G\|_\infty \|G\|_F \Big( \|{\widehat{P}^{(1)}}\|_F \|{X^{(2)}}\| + \|{X^{(1)}}\| \|{P^{(2)}}\|_F \Big). \end{align}\tag{61}\] Invoking 1 3 4 and 12, we have \[\|G\|_\infty \le \|\beta_1^{-1}\beta_2^{-1}\|_{\infty} + \|\widehat \beta_1^{-1} \widehat \beta_2^{-1}\|_{\infty} \lesssim \frac{k}{n^3\rho_n^2}, \qquad \|G\|_F \le \|\beta_1^{-1}\beta_2^{-1}\|_F + \|\widehat \beta_1^{-1} \widehat \beta_2^{-1}\|_F \le \frac{\sqrt{k}}{n^2\rho_n^2}.\] Thus, 61 is bounded by \[\frac{k}{n^3\rho_n^2}\cdot\frac{\sqrt{k}}{n^2\rho_n^2} \Big( \|{\widehat{P}^{(1)}}\|_F \|{X^{(2)}}\| + \|{X^{(1)}}\| \|{P^{(2)}}\|_F \Big) = O_p \left(\frac{k^2}{n^{3.5}\rho_n^{2.5}}\right) = o_p\!\left(\frac{k^2}{n^3\rho_n^2}\right),\] which completes the bound for the first term on the right hand side of 60 . Before bounding the second term on the right hand side of 60 , we observe that \(({\Sigma^{(1)}}\!\circ\!{\Sigma^{(2)}} - (\widehat{\Sigma}^{(1)})\!\circ\!(\widehat{\Sigma}^{(2)}))\) is an \(n \times n\) matrix whose elements are \(O(\rho_n^2).\) Therefore, \[\begin{align}\label{right2} \big| \langle G\!\circ\!G,\; ({\Sigma^{(1)}}\!\circ\!{\Sigma^{(2)}} - (\widehat{\Sigma}^{(1)})\!\circ\!(\widehat{\Sigma}^{(2)})) \rangle \big| & \lesssim \| G\!\circ\!G \|_F \|{\Sigma^{(1)}}\!\circ\!{\Sigma^{(2)}} - (\widehat{\Sigma}^{(1)})\!\circ\!(\widehat{\Sigma}^{(2)}) \|_F \\ &= O_p\left( \frac{k^{3/2}}{n^5 \rho_n^4} \cdot \sqrt{k} n \rho_n \right) = o_p\!\left(\frac{k^2}{n^3\rho_n^2}\right). \end{align}\tag{62}\] Thus, 62 60 give the following bound on \(Q_1\) \[\big| \langle G\!\circ\!G,\; ({\Sigma^{(1)}}{\Sigma^{(2)}} - (\widehat{\Sigma}^{(1)})(\widehat{\Sigma}^{(2)})) + (J_n-I_n)\!\circ\!({\Sigma^{(1)}}\!\circ\!{\Sigma^{(2)}} - (\widehat{\Sigma}^{(1)})\!\circ\!(\widehat{\Sigma}^{(2)})) \rangle \big|=o_p\!\left(\frac{k^2}{n^3\rho_n^2}\right).\] Combining it with 59 where we bounded \(Q_2\), we conclude that \(\big| \widehat{\sigma}_2^{(1,2)} - \sigma_2^{(1,2)} \big| \ll \frac{k^2}{n^3\rho_n^2},\) implying \[\begin{align} \label{sig2bd} \big| \widehat{\sigma}_2^2 - \sigma_2^2 \big| \ll \frac{k^2}{n^3\rho_n^2}. \end{align}\tag{63}\]
8 implies that \(\widehat{\sigma}_2^2 \asymp \frac{k^{2}}{n^{3}\rho_n^{2}}\). Combining this with 63 , we obtain \(\frac{\widehat{\sigma}_2}{\sigma_2} \xrightarrow{p} 1.\) Furthermore, 5, 58 and 8 imply that \(\frac{\widehat{\mu}_2 - \mu_2}{\widehat{\sigma}_2} \xrightarrow{p} 0,\) which completes the proof of 3. ◻
We first state some lemmas which will be used to prove the corollaries for the consistency of our test.
Lemma 18 (Cross–term moments). Suppose 1 3 4 5 hold. Then \[\frac{\langle {V^{(1)}}{V^{(1)}}^{\top}- {V^{(2)}}{V^{(2)}}^{\top},\; \widehat{V^{(2)}}\widehat{V^{(2)}}^{\top}- {V^{(2)}}{V^{(2)}}^{\top}\rangle}{\widehat\sigma_2} \asymp 1,\] where \(\widehat\sigma_2\) is the variance estimator defined in 11 .
Proof. See 14.1. ◻
Lemma 19. Suppose 1 3 4 5 hold. Let \(\Delta= {V^{(1)}} {V^{(1)}}^{\top}- {V^{(2)}} {V^{(2)}}^{\top}.\) Then the quantity \[L(\Delta):=\big\langle (\widehat V^{(1)}) (\widehat V^{(1)})^{\top}- {V^{(1)}} {V^{(1)}}^{\top},\;\Delta\big\rangle -\big\langle (\widehat V^{(2)}) (\widehat V^{(2)})^{\top}- {V^{(2)}} {V^{(2)}}^{\top},\Delta\big\rangle,\] satisfies \(L(\Delta) = O_p\left(\max \left(\frac{\|\Delta\|_F}{n\rho_n},\frac{\|\Delta\|_F \sqrt{\rho_n\log n}}{n \rho_n}\right)\right).\)
Proof. See 14.2. ◻
Lemma 20 (Balanced communities imply incoherence). Let \(P=ZBZ^{\top}\) denote the population probability matrix of a SBM, where \(Z\in\{0,1\}^{n\times k}\) is the membership matrix with exactly one unit entry per row and \(B\in\mathbb{R}^{k\times k}\) is a symmetric block matrix. Let \(n_a\) denote the size of community \(a\) and set \(N=\mathrm{diag}(n_1,\dots,n_k)=Z^{\top}Z\). Assume \(\mathop{\mathrm{rank}}(P)=k\). If the community sizes are balanced in the sense that \(n_a \asymp \frac{n}{k},\) then 4 holds.
Proof. See 14.3. ◻
Lemma 21 (Balanced mixed memberships imply incoherence). Let \(P=ZBZ^{\top}\) denote the population probability matrix of a MMSBM, where \(Z\in[0,1]^{n\times k}\) has rows \({Z^{(i)}}^{\top}\) lying in the probability simplex (\(\sum_{a=1}^k z_{i,a}=1\) for each \(i\)), \(B\in\mathbb{R}^{k\times k}\) is symmetric, and \(\mathop{\mathrm{rank}}(P)=k\). Define the Gram matrix \(N:=Z^{\top}Z\in\mathbb{R}^{k\times k}\) and let \(V\in\mathbb{R}^{n\times k}\) be an orthonormal basis of \(\mathop{\mathrm{col}}(P)\) (the population eigenvectors associated to the nonzero eigenvalues). Suppose there exist constants \(c_1,c_2>0\) (independent of \(n\)) such that \(\lambda_{\min}(N)\ge c_1\,\frac{n}{k}\) and \(\lambda_{\max}(N)\le c_2\,\frac{n}{k}.\) Then 4 holds.
Proof. See 14.4. ◻
Proof. We begin by decomposing the test statistic in terms of its population and estimation components. Using the expression of \(\widehat{\mu}_2\) from 11 , we obtain \[\begin{align} \label{eq:Power1} \frac{\|\widehat{V}_{1}\widehat{V}_{1}^{\top}- \widehat{V}_{2}\widehat{V}_{2}^{\top}\|_F^2 - \widehat{\mu}_2}{\widehat{\sigma}_2} &= \frac{\|\widehat{V}_{1}\widehat{V}_{1}^{\top}- {V^{(1)}}{V^{(1)}}^{\top}\|_F^2 - \widehat{\mu}_2^{(1)}}{\widehat{\sigma}_2} + \frac{\|\widehat{V}_{2}\widehat{V}_{2}^{\top}- {V^{(2)}}{V^{(2)}}^{\top}\|_F^2 - \widehat{\mu}_2^{(2)}}{\widehat{\sigma}_2} \\ &\quad + \frac{\|{V^{(1)}}{V^{(1)}}^{\top}- {V^{(2)}}{V^{(2)}}^{\top}\|_F^2}{\widehat{\sigma}_2} - 2\,\frac{\langle \widehat{V}_{1}\widehat{V}_{1}^{\top}- {V^{(1)}}{V^{(1)}}^{\top},\, \widehat{V}_{2}\widehat{V}_{2}^{\top}- {V^{(2)}}{V^{(2)}}^{\top}\rangle}{\widehat{\sigma}_2} + 2\,\frac{L(\Delta)}{\widehat{\sigma}_2}, \end{align}\tag{64}\] where \(L(\Delta)\) is defined in 19. From 5, the first two terms on the right hand side of 64 \(\frac{\|\widehat{V}_{i}\widehat{V}_{i}^{\top}- {V^{(i)}}{V^{(i)}}^{\top}\|_F^2 - \widehat{\mu}_2^{(i)}}{\widehat{\sigma}_2}\asymp 1\) for \(i = 1, 2\). Furthermore, by 18, the cross–sample interaction term satisfies \[\frac{\langle \widehat{V}_{1}\widehat{V}_{1}^{\top}- {V^{(1)}}{V^{(1)}}^{\top},\, \widehat{V}_{2}\widehat{V}_{2}^{\top}- {V^{(2)}}{V^{(2)}}^{\top}\rangle}{\widehat{\sigma}_2} \asymp 1.\] Turning to the component involving \(L(\Delta)\), 19 yields \(L(\Delta) = O_p\left(\max \left(\frac{\|\Delta\|_F}{n\rho_n},\frac{\|\Delta\|_F \sqrt{\rho_n \log n}}{n \rho_n}\right)\right).\) Finally, recalling from 8 that \(\widehat{\sigma}_2 \asymp \dfrac{k}{n^{3/2}\rho_n}\), the decomposition in 64 reduces to the asymptotic expansion \[\label{eq:Power2} \frac{\|\widehat{V}_{1}\widehat{V}_{1}^{\top}- \widehat{V}_{2}\widehat{V}_{2}^{\top}\|_F^2 - \widehat{\mu}_2}{\widehat{\sigma}_2} \asymp \zeta_1 + \frac{n^{3/2}\rho_n}{k}\,\|\Delta\|_F^2 + \zeta_2,\tag{65}\] where \(\zeta_1\asymp 1\) and \(\zeta_2=O_p\left(\max \left(\frac{\sqrt n\|\Delta\|_F}{k},\frac{\|\Delta\|_F \sqrt{n\rho_n \log n}}{k}\right)\right)\).
Therefore, for the left–hand side of 65 to diverge it suffices that \(\dfrac{n^{3/2}\rho_n}{k}\|\Delta\|_F^2\) diverges; equivalently, \(\|\Delta\|_F^2 \gg \frac{k}{n^{3/2}\rho_n}.\) We note that forcing divergence via \(\zeta_2\) would require \(\|\Delta\|_F^2 \gg \min\left(\dfrac{k^2}{n},\dfrac{k^2}{n \rho_n \log n}\right)\) and under 5, we have \(\min \left(\dfrac{k^2}{n},\dfrac{k^2}{n \rho_n \log n} \right)\gg \dfrac{k}{n^{3/2}\rho_n}\).
Thus, in the dense regime (5) the test attains power tending to one whenever \[\|{V^{(1)}}{V^{(1)}}^{\top}- {V^{(2)}}{V^{(2)}}^{\top}\|_F^2 \gg \frac{k}{n^{3/2}\rho_n},\] which verifies the signal–strength condition stated in 13 and completes the argument. ◻
Proof. Using Lemma 1 from [36], in the SBM setting with balanced communities, the population eigenspace projector admits the explicit representation \(Q_i = {V^{(i)}} {V^{(i)}}^{\top}= {Z^{(i)}} ({Z^{(i)}}^{\top}{Z^{(i)}})^{-1} {Z^{(i)}}^{\top} = \sum_{l=1}^k \frac{1}{{n^{(i)}_{l}}}\,({Z^{(i)}})_{\cdot l}({Z^{(i)}})_{\cdot l}^{\top}\) for \(i=1,2\), where \({n^{(i)}_{l}}=n{\pi^{(i)}_{l}}\) denotes the size of community \(l\) under model \(i\). Since each \(Q_i\) is an orthogonal projector of rank \(k\), we may write \[\begin{align} \label{eq:Q1} \|Q_1-Q_2\|_F^2 = \mathop{\mathrm{tr}}(Q_1) + \mathop{\mathrm{tr}}(Q_2) - 2\mathop{\mathrm{tr}}(Q_1Q_2) = 2k - 2\mathop{\mathrm{tr}}(Q_1Q_2). \end{align}\tag{66}\] Hence it suffices to analyse \(\mathop{\mathrm{tr}}(Q_1Q_2)\) using the block expansions above. Under the balanced–community assumption (so that \({\pi^{(i)}_{l}}\) are bounded away from \(0\) and \(\infty\)) we have \[\begin{align}\label{trq1q2} \mathop{\mathrm{tr}}(Q_1Q_2) &= \sum_{l_1,l_2=1}^k \frac{((Z^{(1)})_{\cdot l_1}^{\top}(Z^{(2)})_{\cdot l_2})^2}{{n^{(1)}_{l}}{n^{(2)}_{l}}} \\ &= \frac{1}{n^2}\sum_{l_1,l_2=1}^k \frac{((Z^{(1)})_{\cdot l_1}^{\top}(Z^{(2)})_{\cdot l_2})^2}{{\pi_{l_1}^{(1)}}{\pi_{l_2}^{(2)}}}\\ &=\frac{1}{n^2}\sum_{l=1}^k \frac{((Z^{(1)})_{\cdot l}^{\top}(Z^{(2)})_{\cdot l})^2}{{\pi^{(1)}_{l}}{\pi^{(2)}_{l}}} + O \left(\frac{\|Z^{(1)}-Z^{(2)}\|_0^2}{n^2}\right). \end{align}\tag{67}\] The last big-O term arises from the fact that \(((Z^{(1)})_{\cdot l_1}^{\top}(Z^{(2)})_{\cdot l_2})^2 \lesssim \|(Z^{(1)})_{\cdot l_1}-(Z^{(2)})_{\cdot l_2}\|_0^2\) if \(l_1 \neq l_2\) and corresponds to the cross term contributions. These cross-terms count nodes that moved from community \(l_1\) to community \(l_2\).
Let \(w_l := (Z^{(1)})_{\cdot l}-(Z^{(2)})_{\cdot l}\). Then \[\|w_l\|^2 = (Z^{(1)})_{\cdot l}^{\top}(Z^{(1)})_{\cdot l}+(Z^{(2)})_{\cdot l}^{\top}(Z^{(2)})_{\cdot l}-2\,(Z^{(1)})_{\cdot l}^{\top}(Z^{(2)})_{\cdot l} = n({\pi^{(1)}_{l}}+{\pi^{(2)}_{l}}) - 2\,(Z^{(1)})_{\cdot l}^{\top}(Z^{(2)})_{\cdot l},\] so that \((Z^{(1)})_{\cdot l}^{\top}(Z^{(2)})_{\cdot l} = \frac{{\pi^{(1)}_{l}}+{\pi^{(2)}_{l}}}{2}\,n - \tfrac12\|w_l\|^2.\) Squaring gives the expansion \[((Z^{(1)})_{\cdot l}^{\top}(Z^{(2)})_{\cdot l})^2 = \Big(\frac{{\pi^{(1)}_{l}}+{\pi^{(2)}_{l}}}{2}\Big)^2 n^2 - \Big(\frac{{\pi^{(1)}_{l}}+{\pi^{(2)}_{l}}}{2}\Big)n\,\|w_l\|^2 + \tfrac14\|w_l\|^4.\] Substituting this into the expression for \(\mathop{\mathrm{tr}}(Q_1Q_2)\) in 67 and using \({n^{(i)}_{l}}=n{\pi^{(i)}_{l}}\) yields \[\begin{align} \mathop{\mathrm{tr}}(Q_1Q_2) &=\sum_{l=1}^k \frac{\Big(\frac{{\pi^{(1)}_{l}}+{\pi^{(2)}_{l}}}{2}\Big)^2}{{\pi^{(1)}_{l}}{\pi^{(2)}_{l}}} + \frac{1}{n^2}\sum_{l=1}^k \frac{ - \Big(\frac{{\pi^{(1)}_{l}}+{\pi^{(2)}_{l}}}{2}\Big)n\,\|w_l\|^2 + \tfrac14\|w_l\|^4}{{\pi^{(1)}_{l}}{\pi^{(2)}_{l}}} + O \left(\frac{\|Z^{(1)}-Z^{(2)}\|_0^2}{n^2}\right).\\ &= k + \sum_{l=1}^{k} \frac{({\pi^{(1)}_{l}}-{\pi^{(2)}_{l}})^2}{4 {\pi^{(1)}_{l}} {\pi^{(2)}_{l}}} + \sum_{l=1}^k \Bigg[ - \frac{{\pi^{(1)}_{l}}+{\pi^{(2)}_{l}}}{2n{\pi^{(1)}_{l}}{\pi^{(2)}_{l}}}\,\|w_l\|^2 + \frac{1}{4n^2{\pi^{(1)}_{l}}{\pi^{(2)}_{l}}}\,\|w_l\|^4\Bigg] + O \left(\frac{\|Z^{(1)}-Z^{(2)}\|_0^2}{n^2} \right) \\ &= k + \sum_{l=1}^k \Bigg[ - \frac{{\pi^{(1)}_{l}}+{\pi^{(2)}_{l}}}{2n{\pi^{(1)}_{l}}{\pi^{(2)}_{l}}}\,\|w_l\|^2 + \frac{1}{4n^2{\pi^{(1)}_{l}}{\pi^{(2)}_{l}}}\,\|w_l\|^4\Bigg] + O \left(\frac{\|Z^{(1)}-Z^{(2)}\|_0^2}{n^2} \right). \end{align}\] The leading \(k\)–term of the above expression cancels with the \(2k\) appearing in 66 . Hence \[\begin{align} \label{q1mq2} \|Q_1-Q_2\|_F^2 &= 2k - 2\mathop{\mathrm{tr}}(Q_1Q_2)\\ \notag &= 2\sum_{l=1}^k \Bigg[ \frac{{\pi^{(1)}_{l}}+{\pi^{(2)}_{l}}}{2n{\pi^{(1)}_{l}}{\pi^{(2)}_{l}}}\,\|w_l\|^2 - \frac{1}{4n^2{\pi^{(1)}_{l}}{\pi^{(2)}_{l}}}\,\|w_l\|^4 \Bigg] + O \left(\frac{\|Z^{(1)}-Z^{(2)}\|_0^2}{n^2} \right). \end{align}\tag{68}\] Now, if \(\frac{\|Z^{(1)}-Z^{(2)}\|_0^2}{n^2} \asymp 1\), then the condition \(\max_{1 \leq l \leq k} \sum_{j=1}^{n} | Z_{jl}^{(1)} - Z_{jl}^{(2)} | \gtrsim \frac{k}{n^{1/2}\rho_n}\) is already satisfied trivially. So, we will assume \(\frac{\|Z^{(1)}-Z^{(2)}\|_0}{n} \ll 1\). From the definition of \(w_l\), we have \[\begin{align} &\|w_l\|^2 \leq n_l^{(1)} + n_l^{(2)} = n{\pi^{(1)}_{l}} + n{\pi^{(2)}_{l}} = n({\pi^{(1)}_{l}} + {\pi^{(2)}_{l}}) \end{align}\] which implies that \[\begin{align} & \frac{1}{4n^2{\pi^{(1)}_{l}}{\pi^{(2)}_{l}}}\,\|w_l\|^4 \lesssim \frac{{\pi^{(1)}_{l}}+{\pi^{(2)}_{l}}}{2n{\pi^{(1)}_{l}}{\pi^{(2)}_{l}}}\,\|w_l\|^2 \end{align}\] under the balanced–community assumption that \({\pi^{(i)}_{l}}\) are bounded above and below by constants. Thus, discarding the smaller quartic correction and the big-O term gives \[\|Q_1-Q_2\|_F^2 \asymp \frac{1}{n}\sum_{l=1}^k \|(Z^{(1)})_{\cdot l}-(Z^{(2)})_{\cdot l}\|^2.\] Combining this representation with the condition from 4, \(\|Q_1-Q_2\|_F^2 \gg \frac{k}{n^{3/2}\rho_n},\) we obtain the necessary lower bound \(\sum_{l=1}^k \|(Z^{(1)})_{\cdot l}-(Z^{(2)})_{\cdot l}\|^2 \gg \frac{k}{n^{1/2}\rho_n},\) and in particular since community sizes are balanced, \[\max_{1 \leq l \leq k} \sum_{j=1}^{n} \left| Z_{jl}^{(1)} - Z_{jl}^{(2)} \right| \gg \frac{k}{n^{1/2}\rho_n},\] and each \(\|(Z^{(1)})_{\cdot l}-(Z^{(2)})_{\cdot l}\|^2\) equals twice the Hamming distance for binary indicator vectors. This completes the proof. ◻
Proof. Adopting the notation from the statement, for \(i=1,2\) write \(Q_i = {V^{(i)}} {V^{(i)}}^{\top}= {Z^{(i)}} ({Z^{(i)}}^{\top}{Z^{(i)}})^{-1} {Z^{(i)}}^{\top},\) where the second equality follows from the mixed–membership representation \({V^{(i)}} = {Z^{(i)}} ({Z^{(i)}}^{\top}{Z^{(i)}})^{-1/2}T\) with \(T\) orthogonal.
We begin with the decomposition \[\begin{align} \label{MMSBM95a95final} \|Q_1-Q_2\|_F &=\|Z^{(1)} ({Z^{(1)}}^{\top}Z^{(1)})^{-1}{Z^{(1)}}^{\top}- Z^{(2)} ({Z^{(2)}}^{\top}Z^{(2)})^{-1}{Z^{(2)}}^{\top}\|_F \\ &\le \|(Z^{(1)}-Z^{(2)})({Z^{(1)}}^{\top}Z^{(1)})^{-1}{Z^{(1)}}^{\top}\|_F + \|Z^{(2)}\bigl(({Z^{(1)}}^{\top}Z^{(1)})^{-1}-({Z^{(2)}}^{\top}Z^{(2)})^{-1}\bigr){Z^{(1)}}^{\top}\|_F \\ &\qquad\qquad + \|Z^{(2)}({Z^{(2)}}^{\top}Z^{(2)})^{-1}(Z^{(1)}-Z^{(2)})^{\top}\|_F. \end{align}\tag{69}\] Under the spectral gap hypothesis \(\lambda_{\min}({Z^{(i)}}^{\top}{Z^{(i)}})\ge c\,\frac{n}{k}\), we have \[\begin{align} \label{bdneed1} \|({Z^{(i)}}^{\top}{Z^{(i)}})^{-1}\| \lesssim \frac{k}{n} \quad \text{and} \quad \|{Z^{(i)}}\|_F \lesssim \sqrt{n} \end{align}\tag{70}\] for \(i=1,2.\) Hence the first and third terms on the right hand side of 69 are immediately bounded by \[\begin{align} \|(Z^{(1)}-Z^{(2)})({Z^{(1)}}^{\top}Z^{(1)})^{-1}{Z^{(1)}}^{\top}\|_F &\lesssim \frac{k\,\|Z^{(1)}-Z^{(2)}\|_F}{\sqrt{n}}, \\ \|Z^{(2)}({Z^{(2)}}^{\top}Z^{(2)})^{-1}(Z^{(1)}-Z^{(2)})^{\top}\|_F &\lesssim \frac{k\,\|Z^{(1)}-Z^{(2)}\|_F}{\sqrt{n}}. \end{align}\] To handle the middle term in 69 set \(A={Z^{(1)}}^{\top}Z^{(1)}\) and \(B={Z^{(2)}}^{\top}Z^{(2)}\). Using the identity \(A^{-1}-B^{-1}=A^{-1}(B-A)B^{-1}\) we obtain \[\begin{align} \label{cross95MMSBM95final} \|Z^{(2)}(A^{-1}-B^{-1}){Z^{(1)}}^{\top}\|_F &\le \|Z^{(2)}\|_F\,\|A^{-1}(B-A)B^{-1}\|_F\,\|Z^{(1)}\|_F \\ &\le \|Z^{(2)}\|_F \|A^{-1}\| \|B-A\|_F \|B^{-1}\| \|Z^{(1)}\|_F. \end{align}\tag{71}\] A direct expansion gives \[\begin{align} \label{cross95MMSBM95diff} \|B-A\|_F &= \|{Z^{(2)}}^{\top}Z^{(2)} - {Z^{(1)}}^{\top}Z^{(1)}\|_F = \|{Z^{(2)}}^{\top}(Z^{(2)}-Z^{(1)})\|_F + \|(Z^{(2)}-Z^{(1)})^{\top}Z^{(1)}\|_F \\ &\lesssim \|Z^{(2)}\|_F\|Z^{(2)}-Z^{(1)}\|_F + \|Z^{(1)}\|_F\|Z^{(2)}-Z^{(1)}\|_F \lesssim \sqrt{n}\,\|Z^{(1)}-Z^{(2)}\|_F. \end{align}\tag{72}\] Combining 72 with 71 70 yields \(\|Z^{(2)}(A^{-1}-B^{-1}){Z^{(1)}}^{\top}\|_F \lesssim \frac{k^2\,\|Z^{(1)}-Z^{(2)}\|_F}{\sqrt{n}}.\) Collecting the bounds for the three terms on the right hand side of 69 we obtain \(\|Q_1-Q_2\|_F \lesssim \frac{k^2}{\sqrt{n}}\,\|Z^{(1)}-Z^{(2)}\|_F.\) Recalling the signal requirement from 4, \(\|Q_1-Q_2\|_F^2 \gg \frac{k}{n^{3/2}\rho_n},\) and substituting the preceding inequality gives the condition on the membership matrices \(\|Z^{(1)}-Z^{(2)}\|_F^2 \gg \frac{1}{n^{1/2}\rho_n},\) as claimed. ◻
In this section we prove all of the technical results stated in 7.
Proof. The probability bound for \(\mathcal{E}_{{\sf good}}\) follows from a direct application of Corollary 3.12 of [61]. Specifically, for a random matrix \(X=A-P\) with independent, mean-zero entries, we have \(\mathbb{P}\left(\|X\|\ge (1+\epsilon)\,2\widetilde{\sigma}+t\right) \le \exp\left(\log n - \frac{t^2}{c_{\epsilon}^2\,\widetilde{\sigma}_*^{\,2}}\right),\) where \(\widetilde{\sigma}=\max_i\sqrt{\sum_j \mathbb{E}({X_{ij}}^2)}\) and \(\widetilde{\sigma}_*=\max_{i,j}\|X_{ij}\|_{\infty}\). In our setting, since the entries of \(A\) are Bernoulli with mean \(P_{ij} \asymp \rho_n\), we have \(\widetilde{\sigma}\asymp \sqrt{n\rho_n}\) and \(\widetilde{\sigma}_*\le 1\). Choosing \(t=\sqrt{n\rho_n}\) gives \(\mathbb{P}\left(\|X\|\gtrsim \sqrt{n\rho_n}\right) \le \exp\left(\log n - \frac{n\rho_n}{C^2}\right),\) for some universal constant \(C>0\). Under 1 2, the right-hand side can be shown as \(O(n^{-19})\) .
The Davis–Kahan perturbation bound on the event \(\mathcal{E}_{{\sf good}}\) follows from Corollary 2.8 of [37]. Since \(\|X\|\lesssim \sqrt{n\rho_n}\) on \(\mathcal{E}_{{\sf good}}\), and by 3 we have \(\|X\|<(1-\frac{1}{\sqrt{2}})\lambda_k\), the eigengap condition required by Davis–Kahan holds. Consequently, \(\|\widehat V\widehat V^{\top}- VV^{\top}\|_F^2 \lesssim \frac{k\,\|X\|^2}{(n\rho_n)^2}.\) Substituting \(\|X\|\lesssim \sqrt{n\rho_n}\) yields the desired bound \(\big\|\widehat V\widehat V^{\top}- VV^{\top}\big\|_F^2 \lesssim \frac{k}{n\rho_n},\) which holds on the event \(\mathcal{E}_{{\sf good}}\). ◻
Proof. Define \[M(X) = X^\top X, \qquad M(Y) = Y^\top Y,\] where \(Y = \beta^{\perp} X\). It therefore suffices to analyze \[\frac{\mathop{\mathrm{Var}}\big(\|\beta^{\perp} X \beta^{-1}\|_F^2-\|X \beta^{-1}\|_F^2\big)}{\mathop{\mathrm{Var}}\big(\|X \beta^{-1}\|_F^2\big)}=\frac{\mathop{\mathrm{Var}}\big( \mathop{\mathrm{tr}}\big( \beta^{-1} (M(Y) - M(X)) \beta^{-1} \big) \big)}{\mathop{\mathrm{Var}}\big( \mathop{\mathrm{tr}}\big( \beta^{-1} M(X) \beta^{-1} \big) \big)}.\] We can write \[M(X) - M(Y) = X^\top X - X^\top \beta^{\perp} \beta^{\perp} X = X^\top (I - \beta^{\perp}) X = X^\top V V^\top X.\] Denote the inner matrix by \(W = V V^\top\). Then \[\mathop{\mathrm{Var}}\left( \mathop{\mathrm{tr}}\big( \beta^{-1} (M(Y) - M(X)) \beta^{-1} \big) \right) = \mathop{\mathrm{Var}}\left( \mathop{\mathrm{tr}}\left( V V^\top X \beta^{-2} X^\top \right) \right).\] Hence, to demonstrate that \(\mathop{\mathrm{Var}}\big(\mathop{\mathrm{tr}}(\beta^{-1} (M(Y) - M(X)) \beta^{-1})\big)\) is asymptotically negligible compared to \(\mathop{\mathrm{Var}}\big(\mathop{\mathrm{tr}}(\beta^{-1} M(X) \beta^{-1})\big)\), it suffices to establish that \[\frac{ \mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(V V^\top X \beta^{-2} X^\top\right)\right) }{ \mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(X \beta^{-2} X^\top\right)\right) } \to 0.\] The trace appearing in the numerator can be expanded as \[\mathop{\mathrm{tr}}\big(V V^\top X \beta^{-2} X\big) = \sum_{a,b,c} (V V^\top)_{a,b}\, (\beta^{-2})_{c,c}\, X_{a,c} X_{b,c}.\] We partition the above sum into diagonal (\(a=b\)) and off-diagonal (\(a \neq b\)) terms. For the diagonal terms (\(a=b\)), \[\begin{align} \label{need02} \mathop{\mathrm{Var}}\bigg(\sum_{a,c} (V V^\top)_{a,a} (\beta^{-2})_{c,c} X_{a,c}^2 \bigg) & = \sum_{a,c} (V V^\top)_{a,a}^2 (\beta^{-2})_{c,c}^2 \mathop{\mathrm{Var}}(X_{a,c}^2)\\ & \asymp \rho_n \sum_{a,c} (V V^\top)_{a,a}^2 (\beta^{-2})_{c,c}^2. \end{align}\tag{73}\] For the off-diagonal terms (\(a \neq b\)), \[\begin{align} \label{need03} \mathop{\mathrm{Var}}\bigg(\sum_{a \neq b, c} (V V^\top)_{a,b} (\beta^{-2})_{c,c} X_{a,c} X_{b,c} \bigg) &= \sum_{a \neq b, c} (V V^\top)_{a,b}^2 (\beta^{-2})_{c,c}^2 \mathop{\mathrm{Var}}(X_{a,c} X_{b,c})\\ & \asymp \rho_n \sum_{a,c} (V V^\top)_{a,a}^2 (\beta^{-2})_{c,c}^2 + \rho_n^2 \|V V^\top\|_F^2 \|\beta^{-2}\|_F^2. \end{align}\tag{74}\] We compare the above to the variance of the unweighted trace in the denominator. By a perfectly analogous variance expansion for \(\mathop{\mathrm{tr}}(X \beta^{-2} X) = \sum_{a,c} (\beta^{-2})_{c,c} X_{a,c}^2\), one obtains \[\begin{align} \label{need04} \mathop{\mathrm{Var}}\big(\mathop{\mathrm{tr}}(X \beta^{-2} X^\top)\big) = \sum_{a,c} (\beta^{-2})_{c,c}^2 \mathop{\mathrm{Var}}(X_{a,c}^2) \asymp \rho_n\, n\, \|\beta^{-2}\|_F^2. \end{align}\tag{75}\] Combining the results from 73 74 75 yields \[\begin{align} \frac{\mathop{\mathrm{Var}}\big(\mathop{\mathrm{tr}}(V V^\top X \beta^{-2} X^\top)\big)}{\mathop{\mathrm{Var}}\big(\mathop{\mathrm{tr}}(X \beta^{-2} X^\top)\big)} &\lesssim \frac{\rho_n \sum_{a} (V V^\top)_{a,a}^2 \|\beta^{-2}\|_F^2 + \rho_n^2 \|V V^\top\|_F^2 \|\beta^{-2}\|_F^2}{\rho_n\, n\, \|\beta^{-2}\|_F^2}\\ &= \frac{\sum_{a} (V V^\top)_{a,a}^2 + \rho_n \|V V^\top\|_F^2}{n}\\ &\lesssim \frac{k(1+\rho_n)}{n} \to 0, \end{align}\] where the last line follows from the fact that \(VV^{\top}\) is a projection matrix and 1. This completes the proof of the lemma. ◻
Proof. Computing the mean: The expectation of the second–order term can be written in a compact operator form. Using linearity of trace and expectation, we have \[\label{eq:mean1}
\begin{align}
\mathbb{E}\bigl[\,2\|\beta^{\perp} X \beta^{-1}\|_F^2\bigr]
&= 2\,\mathbb{E}\bigl[\mathop{\mathrm{tr}}(\beta^{-2} X \beta^{\perp} X)\bigr]
= 2\,\mathop{\mathrm{tr}}\bigl[\beta^{-2}\,\mathbb{E}(X \beta^{\perp} X)\bigr].
\end{align}\tag{76}\] To evaluate \(\mathbb{E}(X \beta^{\perp} X)\) we use the independence and centering of the entries of \(X\). Write \(\Sigma:=P\circ(J_n-P)\) for the entrywise variances of the Bernoulli observations. A straightforward bookkeeping of the contributing index configurations yields the decomposition of \(\mathbb{E}(X
\beta^{\perp} X)\) into an off–diagonal Hadamard term and a diagonal correction: \(\mathbb{E}(X \beta^{\perp} X) = \Sigma\circ \beta^{\perp} + \mathrm{Diag}(\Sigma \cdot d - \mathrm{diag}(\Sigma) \circ d),\) where
\(d:=\mathrm{diag}(\beta^{\perp}).\) Substituting this identity into 76 produces the final mean expression \[\begin{align}
\mathbb{E}\bigl[\,2\|\beta^{\perp} X \beta^{-1}\|_F^2\bigr]
&= 2\,\mathop{\mathrm{tr}}\Bigl[\beta^{-2}\bigl(\Sigma\circ \beta^{\perp} + \mathrm{Diag}(\Sigma \cdot d - \mathrm{diag}(\Sigma) \circ d))\bigr)\Bigr] \\
&=2\,\mathop{\mathrm{tr}}\Bigl[\beta^{-2}\bigl(\Sigma\circ\beta^\perp + \mathrm{Diag}(\Sigma\cdot d - \mathrm{diag}(\Sigma) \circ d))\bigr)\Bigr].
\end{align}\] This identity is the deterministic mean contribution that appears in 16 .
Computing the variance: We now compute the variance of the leading second–order contribution. Observe that since \(\beta^{\perp} = I - VV^{\top}\), by the cyclic property of trace we have that \[\begin{align}
\|\beta^{\perp}X\beta^{-1}\|_F^2
= \mathop{\mathrm{tr}}(\beta^{-1}X\beta^{\perp}X\beta^{-1})
= \mathop{\mathrm{tr}}(X\beta^{-2}X) - \mathop{\mathrm{tr}}(VV^{\top}X\beta^{-2}X).
\end{align}\] From 2, it suffices to analyze the variance of \(2\|X \beta^{-1}\|_F^2 = 2\,\mathop{\mathrm{tr}}(X \beta^{-2}
X^{\top})
= \sum_{i=1}^n 2\,X_{i \cdot}^{\top}\beta^{-2} X_{i \cdot},\) For a fixed \(i\), define the quadratic form \[\begin{align}
Q_i := 2\,X_{i \cdot}^{\top}\beta^{-2} X_{i \cdot}
= 4\sum_{s < t} (\beta^{-2})_{st}\,{X_{is}}\,{X_{it}}
+ 2\sum_{s=1}^n (\beta^{-2})_{ss}\,{X_{is}}^2,
\end{align}\] so that \(2\|X \beta^{-1}\|_F^2 = \sum_{i=1}^n Q_i\). The variance then decomposes as \[\begin{align}
\label{eq:varq1}
\mathop{\mathrm{Var}}\left(2\|X \beta^{-1}\|_F^2\right)
= \sum_{i=1}^n \mathop{\mathrm{Var}}(Q_i) + 2\sum_{i<k} \mathop{\mathrm{Cov}}(Q_i,Q_k).
\end{align}\tag{77}\] For the diagonal terms, using independence and the standard fourth moment identity for centered Bernoulli variables, we obtain \[\begin{align}
\mathop{\mathrm{Var}}(Q_i)
= 16 \sum_{s < t} (\beta^{-2})_{st}^2\,\Sigma_{is}\,\Sigma_{it}
+ 4 \sum_{s=1}^n (\beta^{-2})_{ss}^2\,K_{is},
\end{align}\] where \(K_{is} := \mathbb{E}({X_{is}}^4)-\bigl(\mathbb{E}({X_{is}}^2)\bigr)^2
= 2\,\Sigma_{is}(1-2P_{is})\) is the fourth–moment correction term.
For the off–diagonal contributions, note that \(Q_i\) and \(Q_k\) share only the symmetric variable \(x_{ik}=x_{ki}\). A direct calculation shows \(\mathop{\mathrm{Cov}}(Q_i,Q_k) = 4\,K_{ik}\,(\beta^{-2})_{ii}\,(\beta^{-2})_{kk}.\) Combining the two components and rewriting the double sums from 77 yields \[\begin{align} \mathop{\mathrm{Var}}\left(2\|X \beta^{-1}\|_F^2\right) &=\sum_{i=1}^{n} \left( 16 \sum_{s < t} (\beta^{-2})_{st}^2\,\Sigma_{is}\,\Sigma_{it} + 4 \sum_{s=1}^n (\beta^{-2})_{ss}^2\,K_{is} \right) + 8 \sum_{i<k} \,K_{ik}\,(\beta^{-2})_{ii}\,(\beta^{-2})_{kk}\\ &=8\sum_{i=1}^{n} \left( \sum_{s \ne t} (\beta^{-2})_{st}^2\,\Sigma_{is}\,\Sigma_{it} \right) + \left( 4 \sum_{s=1}^n (\beta^{-2})_{ss}^2\,K_{is} + 4 \sum_{i \ne k} \,K_{ik}\,(\beta^{-2})_{ii}\,(\beta^{-2})_{kk} \right)\\ &= 8\big\langle (J_n-I_n)\circ(\beta^{-2}\circ\beta^{-2}),\,\Sigma^2 \big\rangle + 4\,\alpha^{\top}K \alpha. \end{align}\] This expression coincides with the variance term in 17 .
Finally, we compute the asymptotic order of the variance obtained above. From 1 3 4, we have \(\bigl|(\beta^{-2})_{ij}\bigr| \asymp \dfrac{k}{n^{3}\rho_n^{2}},\) \(K_{ij} \asymp \rho_n,\) and \(\Sigma_{ij} \asymp \rho_n.\) It follows that \(\mathop{\mathrm{Var}}\left(2\|X \beta^{-1}\|_F^{2}\right) \asymp \frac{k^{2}}{n^{3}\rho_n^{2}}.\) Hence, by 2, we obtain \[\mathop{\mathrm{Var}}\left(2\|\beta^{\perp} X \beta^{-1}\|_F^{2}\right) = 8\big\langle (J_n - I_n)\circ(\beta^{-2}\circ \beta^{-2}),\,\Sigma^{2}\big\rangle + 4\,\alpha^{\top}K \alpha + o\left(\frac{k}{n^{3}\rho_n^{2}}\right).\] ◻
Proof. Suppose that \(\|\beta^{\perp} X \beta^{-1}\|_{F}^{2} = \|X\beta^{-1}\|_{F}^{2} + R.\) By 2, we have \(\mathop{\mathrm{Var}}(R) / \mathop{\mathrm{Var}}(T_{1}^{(S)}) \to 0,\) and consequently \((R - \mathbb{E}R)/\sqrt{\mathop{\mathrm{Var}}(T_{1}^{(S)})} \to 0\) in probability. Hence the contribution of \(R\) is asymptotically negligible. Therefore, in establishing the limiting distribution of \(\|\beta^{\perp} X \beta^{-1}\|_{F}^{2},\) it suffices to derive the asymptotic distribution of \(\|X\beta^{-1}\|_{F}^{2}\) after appropriate centering and scaling. The desired result then follows by a direct application of Slutsky’s theorem. For notational convenience set \(U=\beta^{-1}\). Then \(\|X\beta^{-1}\|_F^2=\mathop{\mathrm{tr}}(\beta^{-1}X^2\beta^{-1})=\mathop{\mathrm{tr}}(UX^2U),\) and hence \(n\mathop{\mathrm{tr}}(UX^2U) = n \sum_{j=1}^{n} \sum_{1 \leq i,r,q \leq n} {X_{ri}} {X_{qi}} U_{jr} U_{jq}.\) Following the decomposition strategy employed in the proof of Lemma 2 in [43], the right–hand side can be reorganized into the form used for a martingale central limit theorem (CLT). Concretely, \[\label{eq:martingale-decomp} \begin{align} \sum_{j=1}^{n} \sum_{1 \leq i,r,q \leq n} n\,{X_{ri}} {X_{qi}} U_{jr} U_{jq} &= \sum_{j=1}^{n} \sum_{1 \leq r < i \leq n} 2 {X_{ri}} \Bigl( n\,U_{jr} \sum_{1 \leq q < r} {X_{iq}} U_{jq} + n\,U_{ji} \sum_{1 \leq q < i} {X_{rq}} U_{jq} \Bigr) \\ &\quad + \sum_{j=1}^{n} \sum_{1 \leq q < r \leq n} 2n \,{X_{rr}} {X_{rq}} \,U_{jq} U_{jr} \\ &\quad + \sum_{j=1}^{n} \sum_{1 \leq r < i \leq n} {X_{ri}}^2 \bigl(n\,U_{jr}^2 + n\,U_{ji}^2\bigr) + \sum_{j=1}^{n} \sum_{r = 1}^{n} {X_{rr}}^2 \,n\,U_{jr}^2 \\ &= \sum_{1 \leq r < i \leq n} 2 {X_{ri}} \Bigl( \sum_{1 \leq q < r} {X_{iq}} \gamma_{rq} + \sum_{1 \leq q < i} {X_{rq}} \gamma_{iq} \Bigr) \\ &\quad + \sum_{1 \leq q < r \leq n} 2 \gamma_{rq} {X_{rr}} {X_{rq}} + \sum_{1 \leq r < i \leq n} {X_{ri}}^2 (\gamma_{rr} + \gamma_{ii}) + \sum_{r = 1}^{n} {X_{rr}}^2 \gamma_{rr}, \end{align}\tag{78}\] where we have set \(\gamma_{r,q} = \sum_{j=1}^{n} n\, U_{jr} U_{jq}.\) After centering we obtain \[\label{eq:centered} \begin{align} \sum_{j=1}^{k} U_{j \cdot}^{\top}(X^2 - \mathbb{E}X^2) U_{j \cdot} &= \sum_{1 \leq r < i \leq n} 2 {X_{ri}} \Bigl( \sum_{1 \leq q < r} {X_{iq}} \gamma_{rq} + \sum_{1 \leq q < i} {X_{rq}} \gamma_{iq} \Bigr) + \sum_{1 \leq q < r \leq n} 2 \gamma_{rq} {X_{rr}} {X_{rq}} \\ &+ \sum_{1 \leq r < i \leq n} ({X_{ri}}^2-\nu_{ri}^2) (\gamma_{rr} + \gamma_{ii}) + \sum_{r = 1}^{n} ({X_{rr}}^2 - \sigma_{rr}^2)\,\gamma_{rr}. \end{align}\tag{79}\] To set up a martingale difference array, we introduce the triangular filtration \[\mathscr{F}_t := \sigma(w_1, \ldots, w_t), \qquad t = k + \tfrac{l(l-1)}{2},\quad 1\le k\le l\le n,\] so that \(\mathscr{F}_t\) records the entries \({X_{ij}}\) revealed up to position \(t\) in the standard lexicographic ordering of the upper–triangular matrix. Under this filtration 79 can be written as a sum of martingale differences. Each summand (the increment associated with the revealing of the entry \({X_{ri}}\)) admits the decomposition \[{X_{ri}} b_{ri} + ({X_{ri}}^2 - \nu_{ri}^2)\, c_{ri},\] with coefficients given by \[\label{eq:coefficients} \begin{align} b_{ri} &= \begin{cases} 2 \displaystyle\Bigl( \sum_{1 \leq q < r} {X_{iq}} \gamma_{rq} + \sum_{1 \leq q < i} {X_{rq}} \gamma_{iq} \Bigr), & r<i, \\[8pt] \displaystyle 2\sum_{1 \leq q < r} {X_{rq}} \gamma_{rq} , & i=r , \end{cases} \\[10pt] c_{ri} &= \begin{cases} \gamma_{rr} + \gamma_{ii}, & r<i, \\[6pt] \gamma_{rr}, & i=r . \end{cases} \end{align}\tag{80}\] The conditional variance (given the \(\sigma\)–field just prior to revealing \({X_{ri}}\)) takes the explicit form \[\label{eq:cond-var} \sum_{1 \leq r \leq i \leq n} \nu_{ri}^2 b_{ri}^2 + 2 \sum_{1 \leq r \leq i \leq n} \theta_{ri} b_{ri} c_{ri} + \sum_{1 \leq r \leq i \leq n} \kappa_{ri} c_{ri}^2,\tag{81}\] where we adopt the notation \(\nu_{ri}^2 = \mathbb{E}[{X_{ri}}^2], \theta_{ri} = \mathbb{E}[{X_{ri}}^3], \kappa_{ri} = \mathbb{E}\bigl[({X_{ri}}^2-\nu_{ri}^2)^2\bigr].\)
The verification of the martingale CLT requires control of the first four moments of the coefficients \(b_{ri}\) and the second and fourth moments of the conditional variance of 81 . We proceed to record the necessary moment estimates. Using the definitions above, one checks that the leading order behavior of the coefficients is captured by the following asymptotic relations (the algebraic manipulations are elementary and follow from counting the contributing index configurations).
We calculate the necessary moment estimates. First, note that \(\mathbb{E}[b_{ri}]=0\) for all \(r,i\). The second moment of \(b_{ri}\) satisfies \[\mathbb{E}[b_{ri}^2] \asymp \sum_{1\le q<r}\gamma_{rq}^2\,\nu_{iq}^2 + \sum_{1\le q<i}\gamma_{iq}^2\,\sigma_{rq}^2 \asymp n^2\,\gamma^2\,\rho_n,\] where \(\gamma_{rq}^2=(\sum_{r=1}^{n} n U_{jr} U_{jq})^2 \asymp \frac{k^2}{n^4 \rho_n^4}\) which follows from 1 4 3. For notational simplicity, we refer to it as \(\gamma^2\) in the proof. Apart from that, we also have \(\nu_{ri}^2, \, \theta_{ri}, \kappa_{ri} \asymp \rho_n\).
The fourth moment of \(b_{ri}\) admits the expansion \[\begin{align} \mathbb{E}[b_{ri}^4] &\asymp \sum_{1\le q<r}\gamma_{rq}^4\,\mathbb{E}[{X_{iq}}^4] + \sum_{1\le q<i}\gamma_{iq}^4\,\mathbb{E}[{X_{rq}}^4] \\ &\quad {}+ \sum_{\substack{q_1\neq q_2\\1\le q_1,q_2<r}} \gamma_{rq_1}^2\gamma_{rq_2}^2\, \mathbb{E}[{X_{iq_1}}^2]\,\mathbb{E}[{X_{iq_2}}^2] + \sum_{\substack{q_1\neq q_2\\1\le q_1,q_2<i}} \gamma_{iq_1}^2\gamma_{iq_2}^2\, \mathbb{E}[{X_{rq_1}}^2]\,\mathbb{E}[{X_{rq_2}}^2] \\ &\asymp n^2\,\gamma^4\,\rho_n^2 + n\,\gamma^4\,\rho_n^2 \\ &\asymp n^2\,\gamma^4\,\rho_n^2. \end{align}\] Finally, for the \(c_{ri}\) terms we have the simpler bounds: \[\begin{align} c_{ri}^2 \asymp \gamma^2, \quad c_{ri}^4 \asymp \gamma^4. \end{align}\] With these moment estimates in hand we evaluate the first two moments of the conditional variance of 81 . Denote by \(s_V^2\) the expectation of the conditional variance and by \(\kappa_V^2\) its variance. A straightforward aggregation over indices yields the asymptotic relations \[\begin{align} s_V^2 &= \mathbb{E}\Big(\sum_{1 \leq r \leq i \leq n}\nu_{ri}^2 b_{ri}^2 + 2 \sum_{1 \leq r \leq i \leq n} \theta_{ri} b_{ri} c_{ri} + \sum_{1 \leq r \leq i \leq n} \kappa_{ri} c_{ri}^2\Big)\\ &\asymp \sum_{1 \leq r \leq i \leq n} p_{ri}\,\mathbb{E}[b_{ri}^2] + \sum_{1 \leq r \leq i \leq n} p_{ri}\, c_{ri}^2 \asymp n^4 \gamma^2 \rho_n^2 + n^2 \rho_n \gamma^2, \end{align}\] and \[\begin{align} \kappa_V &= \mathop{\mathrm{Var}}\Big(\sum_{r\le i}\nu_{ri}^2 b_{ri}^2 + 2\sum_{r\le i}\theta_{ri}b_{ri}c_{ri} + \sum_{r\le i}\kappa_{ri}c_{ri}^2\Big) \\[4pt] &\asymp \sum_{r\le i}\mathop{\mathrm{Var}}(\nu_{ri}^2 b_{ri}^2) + \sum_{r\le i}\mathop{\mathrm{Var}}(2\theta_{ri}b_{ri}c_{ri}) \asymp n^4\gamma^4\rho_n^4 + n^4\gamma^4\rho_n^3 \asymp n^4\gamma^4\rho_n^4. \end{align}\] Combining the above estimates we obtain the Lyapunov–type ratio used to verify the martingale CLT: \[\frac{s_V}{\kappa_V^{1/4}} = \frac{n^2 \gamma \rho_n + n \gamma \sqrt{\rho_n}}{n \gamma \rho_n} \gg 1,\] and simultaneously \(s_V^2 \asymp n^4 \gamma^2 \rho_n + n^2 \rho_n\gamma^2 \asymp \frac{k^4}{\rho_n^2} + \frac{k^2}{\rho_n} \to \infty\) as \(n\to\infty\). Thus the conditional variance diverges and the Lyapunov type condition as used in Lemma 9.12 of [62] is satisfied.
Therefore the martingale central limit theorem applies and we conclude that the appropriately centered and scaled second–order term converges in distribution to a standard normal: \[\frac{ 2\bigl\|\beta^{\perp} X \beta^{-1}\bigr\|_F^2 - \,\mathbb{E}\left[\,2\bigl\|\beta^{\perp} X \beta^{-1}\bigr\|_F^2\right] }{ \sqrt{\mathop{\mathrm{Var}}\left( 2\bigl\|\beta^{\perp} X \beta^{-1}\bigr\|_F^2 \right)} } \xrightarrow{D} \mathcal{N}(0,1).\] ◻
Proof. We first decompose \(\langle VV^{\top},\,S_{3}(X)\rangle\): \[\begin{align} \mathbb{E}\left(\langle VV^{\top},S_3(X)\rangle\right) &= \mathbb{E}\left(\mathop{\mathrm{tr}}\big(\beta^{-1} X \beta^{\perp} X \beta^{-1} X \beta^{-1}\big)\right) +\mathbb{E}\left(\mathop{\mathrm{tr}}\big(\beta^{-1} X \beta^{-1} X \beta^{\perp} X \beta^{-1}\big)\right) \notag\\ &\quad -\mathbb{E}\left(\mathop{\mathrm{tr}}\big(\beta^{-2} X \beta^{\perp} X \beta^{\perp} X \beta^{-1}\big)\right) -\mathbb{E}\left(\mathop{\mathrm{tr}}\big(\beta^{-1} X \beta^{\perp} X \beta^{\perp} X \beta^{-2}\big)\right). \end{align}\] We compute the orders of \(\mathbb{E}\left(\mathop{\mathrm{tr}}(\beta^{-1} X \beta^{\perp} X \beta^{-1} X \beta^{-1})\right)\) and \(\mathbb{E}\left(\mathop{\mathrm{tr}}(\beta^{-2} X \beta^{\perp} X \beta^{\perp} X \beta^{-1})\right)\); the remaining two terms follow identically. Note that
\[\begin{align} \label{eq:mean31} \bigg|\mathbb{E}\left(\mathop{\mathrm{tr}}\big(\beta^{-1} X \beta^{\perp} X \beta^{-1} X \beta^{-1}\big)\right)\bigg| &= \mathbb{E}\bigg|\sum_{i,j} (\beta^{-2})_{ji}\,(\beta^{\perp})_{ji}\,(\beta^{-1})_{ji}\,(X_{ij})^{3}\bigg| \\[2pt] &=O_p \left( \frac{k^{2}}{n^{4}\rho_{n}^{2}}\right), \end{align}\tag{82}\] where the final line follows from 1 3 4. Also, we have
\[\begin{align} \label{eq:mean32} \bigg|\mathbb{E}\left(\mathop{\mathrm{tr}}\big(\beta^{-2} X \beta^{\perp} X \beta^{\perp} X \beta^{-1}\big)\right)\bigg| &= \mathbb{E}\bigg|\sum_{i,j} (\beta^{-3})_{ji}\,(\beta^{\perp})_{ji}^{2}\,(X_{ij})^{3}\bigg| \\[2pt] &= O_p \left( \frac{k}{n^{3}\rho_{n}^{2}}\right). \end{align}\tag{83}\]
Combining 82 83 establishes the claim and completes the proof of 5. ◻
Proof. To evaluate the order of \(\mathop{\mathrm{Var}}\!\big(\langle VV^{\top}, S_3(X)\rangle\big)\), it suffices to show that \[\frac{\mathop{\mathrm{Var}}\!\big(\langle VV^{\top}, S_3(X)\rangle\big)}{\mathop{\mathrm{Var}}\!\left(\mathop{\mathrm{tr}}\!\big(\beta^{-3} X^{3}\big)\right)} \asymp 1\] and calculate the order of \(\mathop{\mathrm{Var}}\!\left(\mathop{\mathrm{tr}}\!\big(\beta^{-3} X^{3}\big)\right).\) We proceed by comparing cyclically reduced representatives. Define \[R_1 := \mathop{\mathrm{tr}}\!\big(\beta^{-3} X^{3}\big), \qquad R_2 := \mathop{\mathrm{tr}}\!\big(\beta^{-3} X(\beta^{\perp}X)^{2}\big).\] We first establish that \(\frac{\mathop{\mathrm{Var}}(R_1)}{\mathop{\mathrm{Var}}(R_2)} \to 1 .\) To this end, insert the decomposition \(I = VV^{\top}+ \beta^{\perp}\) in each interior gap of \(R_1\) to obtain \[\begin{align} \label{R1} R_1 = \sum_{\sigma \in \{VV^{\top},\,\beta^{\perp}\}^{\,2}} \mathop{\mathrm{tr}}\big(\beta^{-3} X M_{\sigma,1} X M_{\sigma,2} X\big), \end{align}\tag{84}\] where each \(M_{\sigma,r}\in\{VV^{\top},\beta^{\perp}\}\). The unique summand for which every \(M_{\sigma,r}=\beta^{\perp}\) is exactly \(R_2\); every other summand (henceforth a remainder) contains at least one factor \(VV^{\top}\). Since the total number of summands is finite, it suffices to prove that each remainder has variance \(o(\mathop{\mathrm{Var}}(R_2))\); the claim then follows by finite summation.
Fix a remainder monomial \(\mathcal{M} := \mathop{\mathrm{tr}}\big(\beta^{-3} X M_{1} X M_{2} X\big).\) Expanding the trace yields \[\begin{align} \label{eq:monomial1} \mathcal{M} &= \sum_{i_{1},i_{2},i_{3} \atop j_{1},j_{2},j_{3}} (\beta^{-3})_{j_{1},i_{1}}\, X_{i_{1},j_{2}}\, (M_{1})_{j_{2},i_{2}}\, X_{i_{2},j_{3}}\, (M_{2})_{j_{3},i_{3}}\, X_{i_{3},j_{1}} \\ &= \sum_{i_{1},i_{2},i_{3} \atop j_{1},j_{2},j_{3}} w(\mathbf{i},\mathbf{j})\, X_{i_{1},j_{2}}\, X_{i_{2},j_{3}}\, X_{i_{3},j_{1}}, \end{align}\tag{85}\] where \(w(\mathbf{i},\mathbf{j}) := (\beta^{-3})_{j_{1},i_{1}} (M_{1})_{j_{2},i_{2}} (M_{2})_{j_{3},i_{3}}.\) In order for \[\mathop{\mathrm{Cov}}\big( w(\mathbf{i},\mathbf{j})\, X_{i_{1},j_{2}}\, X_{i_{2},j_{3}}\, X_{i_{3},j_{1}}, w(\mathbf{k},\mathbf{l})\, X_{k_{1},l_{2}}\, X_{k_{2},l_{3}}\, X_{k_{3},l_{1}} \big) \neq 0\] to hold, the underlying edge configurations must coincide. Without loss of generality, the edge \((k_{2},l_{3})\) must coincide with \((i_{2},j_{3})\), and the subgraph \(\{(k_{1},l_{2}), (k_{3},l_{1})\}\) must coincide with \(\{(i_{1},j_{2}), (i_{3},j_{1})\}\). Alternatively, the edge \((k_{1},l_{2})\) is identified with \((i_{1},j_{2})\), or \((k_{3},l_{1})\) with \((i_{3},j_{1})\), with the corresponding subgraph matching accordingly. Consequently, these constraints imply that the index tuples \((\mathbf{k},\mathbf{l})\) agree with \((\mathbf{i},\mathbf{j})\) up to the natural symmetries of the configuration, a relationship we denote by \((\mathbf{k},\mathbf{l}) \sim (\mathbf{i},\mathbf{j})\). Thus, we obtain \[\mathop{\mathrm{Var}}(\mathcal{M}) \asymp \sum_{(\mathbf{k},\mathbf{l}) \sim (\mathbf{i},\mathbf{j})} w(\mathbf{i},\mathbf{j}) w(\mathbf{k},\mathbf{l})\, \rho_{n}^{3}.\] Define the quantity \[\Sigma(\mathcal{M}) := \sum_{(\mathbf{k},\mathbf{l}) \sim (\mathbf{i},\mathbf{j})} w(\mathbf{i},\mathbf{j}) w(\mathbf{k},\mathbf{l}) = \sum_{(\mathbf{k},\mathbf{l}) \sim (\mathbf{i},\mathbf{j})} \Big( (\beta^{-3})_{j_{1},i_{1}}\, (M_{1})_{j_{1},i_{2}}\, (M_{2})_{j_{2},i_{3}} \Big) \Big( (\beta^{-3})_{l_{1},k_{1}}\, (M_{1})_{l_{1},k_{2}}\, (M_{2})_{l_{2},k_{3}} \Big).\] Invoking 3 1, we find that \[\Sigma(\mathcal{M}) \asymp \frac{k^2}{n^8 \rho_n^6}\sum_{(\mathbf{k},\mathbf{l}) \sim (\mathbf{i},\mathbf{j})} \Big( (M_{1})_{j_{1},i_{2}}\, (M_{2})_{j_{2},i_{3}} \Big) \Big( (M_{1})_{l_{1},k_{2}}\, (M_{2})_{l_{2},k_{3}} \Big).\] Consider the configuration where \(M_1 = V V^\top\) and \(M_2 = \beta^\perp\). Evaluating \(\Sigma(\mathcal{M})\) requires summing over six structural indices: \(i_1, j_1, i_2, j_2, i_3, \text{ and } j_3\). However, because \(M_2 = \beta^\perp\), it enforces an index collapse (specifically, \(j_2 = i_3\)). This structural constraint reduces the effective degrees of freedom in the summation from six to five, yielding a factor of \(O(n^5)\). Coupling this with the entrywise bound of \((k/n)^2\) for \(M_1\), guaranteed by the incoherence condition in 4, we directly obtain the upper bound: \[\Sigma(\mathcal{M}) \asymp \frac{k^2}{n^8 \rho_n^6}\cdot n^5 \cdot \left( \frac{k}{n} \right)^2 \asymp \frac{k^4}{n^5 \rho_n^6}.\] Similarly, if \(M_1=M_2=V V^{\top}\), we have \[\Sigma(\mathcal{M}) \asymp \frac{k^2}{n^8 \rho_n^6}\cdot n^6 \cdot \left( \frac{k}{n} \right)^4 \asymp \frac{k^6}{n^6 \rho_n^6}.\] Finally, if \(M_1=M_2=\beta^{\perp}\), which corresponds to the case \(\mathcal{M}=R_2\), we obtain \(\Sigma(\mathcal{M}) \asymp \frac{k^2}{n^4 \rho_n^6}\). It follows that whenever at least one of \(M_{1}\) or \(M_{2}\) equals \(VV^{\top}\), the variance of the monomial satisfies \[\begin{align} \label{mono1} \mathop{\mathrm{Var}}(\mathcal{M}) = o(\mathop{\mathrm{Var}}(R_2)). \end{align}\tag{86}\]
Analogously, \(\mathop{\mathrm{Var}}(R_1) \asymp \Sigma(R_1) \rho_n^3\), and \[\begin{align}\label{varu1} \Sigma(R_1) &\asymp \frac{k^2}{n^8 \rho_n^6}\sum_{(\mathbf{k},\mathbf{l}) \sim (\mathbf{i},\mathbf{j})} \Big( (\delta)_{j_{1},i_{2}}\, (\delta)_{j_{2},i_{3}} \Big) \Big( (\delta)_{l_{1},k_{2}}\, (\delta)_{l_{2},k_{3}} \Big) \\ & \asymp \frac{k^2}{n^4 \rho_n^6}, \end{align}\tag{87}\] where \(\delta_{i,j}=\mathbb{1}[i=j]\). Thus 87 implies that \[\begin{align} \label{ord1} \mathop{\mathrm{Var}}\!\left(\mathop{\mathrm{tr}}\!\big(\beta^{-3} X^{3}\big)\right) = \mathop{\mathrm{Var}}\!\left(R_1\right) \asymp \frac{k^2}{n^4 \rho_n^3}. \end{align}\tag{88}\] Combining 84 88 with 86 establishes the claim that \[\begin{align} \label{later01} \frac{\mathop{\mathrm{Var}}(R_1)}{\mathop{\mathrm{Var}}(R_2)} = \frac{\mathop{\mathrm{Var}}(\sum_{\sigma \in \{VV^{\top},\,\beta^{\perp}\}^{\,2}} \mathop{\mathrm{tr}}\big(\beta^{-3} X M_{\sigma,1} X M_{\sigma,2} X\big))}{\mathop{\mathrm{Var}}(R_2)} = \frac{\mathop{\mathrm{Var}}(R_2) + o(\mathop{\mathrm{Var}}(R_2))}{\mathop{\mathrm{Var}}(R_2)}\to 1. \end{align}\tag{89}\] In exactly the same manner, we can show that \[\frac{\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta^{-t} X^t \beta^{-(3-t)} X^{3-t}\right) \right)}{\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left( \beta^{-t} X (\beta^{\perp}X)^{t-1}\beta^{-(3-t)}X(\beta^{\perp}X)^{3-t-1}\right) \right)} \to 1\] for \(t=1,2\) and \[\frac{\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta^{-1} X \beta^{-1} X \beta^{-1} X\right) \right)}{\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta^{-1} X \beta^{-1} (\beta^\perp X) \beta^{-1} (\beta^\perp X)\right) \right)} \to 1.\] To establish that the dominant variance contributions in \(\langle VV^{\top},\, S_{3}(X)\rangle\) arise from terms of the form \(\mathop{\mathrm{tr}}\big(\beta^{-s_{1}} X\,\beta^{\perp} X \cdots \beta^{\perp}X\,\beta^{-s_{4}}\big),\) it therefore suffices to show that for \(t =1,2\), \[\frac{\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta^{-t} X^{t}\, \beta^{-(3-t)} X^{3-t}\right)\right)}{\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta^{-3} X^{3}\right)\right)} \to 0, \qquad \frac{\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta^{-1} X \beta^{-1} X \beta^{-1} X\right)\right)}{\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta^{-3} X^{3}\right)\right)} \to 0.\] The proof technique for both ratios is identical; hence we concentrate on the first.
Write the numerator trace as \(R_{1,\mathrm{num}} := \mathop{\mathrm{tr}}\big(\beta^{-t} X^{t}\, \beta^{-(3-t)} X^{3-t}\big).\) Proceeding as in the preceding argument, we obtain \[\begin{align} \mathop{\mathrm{Var}}(R_{1,\mathrm{num}}) &\asymp \rho_{n}^{3} \sum_{i_{1}, i_{2} \atop j_{1}, j_{2}} \left[ (\beta^{-t})_{i_{1},j_{1}}\, (\beta^{-(3-t)})_{i_{2},j_{2}} \right]^{2} \\ &\asymp \rho_{n}^{3} \sum_{i_{1}, i_{2} \atop j_{1}, j_{2}} \frac{k^{4}}{n^{10}\rho_{n}^{6}} \asymp \frac{k^{4}}{n^{6}\rho_{n}^{3}}. \end{align}\] From 87 , we therefore have \[\frac{ \mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta^{-t} X^{t}\, \beta^{-(3-t)} X^{3-t}\right)\right) }{ \mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta^{-3} X^{3}\right)\right) } = \frac{\frac{k^{4}}{n^{6}\rho_{n}^{3}}}{\frac{k^{2}}{n^{4}\rho_{n}^{3}}} \to 0,\] as required. This establishes that the dominant contributions to \(\mathop{\mathrm{Var}}\!\big(\langle VV^{\top}, S_3(X)\rangle\big)\) arise precisely from summands of the form \(\mathop{\mathrm{tr}}\big(\beta^{-s_{1}} X\,\beta^{\perp} X \cdots \beta^{\perp}X\,\beta^{-s_{4}}\big)\). In particular, we have shown that \[\frac{\mathop{\mathrm{Var}}\!\big(\langle VV^{\top}, S_3(X)\rangle\big)}{\mathop{\mathrm{Var}}\!\left(\mathop{\mathrm{tr}}\!\big(\beta^{-3} X^{3}\big)\right)} \asymp 1 .\] Combining this with the variance order in 88 completes the proof of the lemma. ◻
In this section we prove the auxiliary lemmas required for the proof of 1 in 8.
Proof. Define \[M(X) = \begin{pmatrix} {X^{(1)}}^{\top}{X^{(1)}} & -{X^{(1)}}^{\top}{X^{(2)}} \\ - {X^{(2)}}^{\top}{X^{(1)}} & {X^{(2)}}^{\top}{X^{(2)}} \end{pmatrix}, \qquad M(Y) = \begin{pmatrix} (Y^{(1)})^{\top}(Y^{(1)}) & -(Y^{(1)})^{\top}(Y^{(2)}) \\ - (Y^{(2)})^{\top}(Y^{(1)}) & (Y^{(2)})^{\top}(Y^{(2)}) \end{pmatrix}.\] It therefore suffices to analyze \[\begin{align} \frac{\mathop{\mathrm{Var}}\left( 2\, \mathop{\mathrm{tr}}\big( U (M(Y) - M(X)) U^{\top}\big) \right)}{\mathop{\mathrm{Var}}\left( 2\, \mathop{\mathrm{tr}}\big( U M(X) U^{\top}\big) \right)}. \end{align}\] Observe that \[\begin{align} \label{Wijdef} M(X) - M(Y) = \begin{pmatrix} {X^{(1)}}^{\top}W_{11} {X^{(1)}} & {X^{(1)}}^{\top}W_{12} {X^{(2)}} \\ {X^{(2)}}^{\top}W_{21} {X^{(1)}} & {X^{(2)}}^{\top}W_{22} {X^{(2)}} \end{pmatrix}. \end{align}\tag{90}\] Denote the \((i,j)\)-th block of \(M(X) - M(Y)\) by \({X^{(i)}}^{\top}W_{ij} {X^{(j)}}\), where \[\begin{align}\label{note04} W_{11} &= {V^{(1)}} {V^{(1)}}^{\top}, \\ W_{12} = W_{21} &= -({V^{(1)}} {V^{(1)}}^{\top}+ {V^{(2)}} {V^{(2)}}^{\top}- {V^{(1)}} {V^{(1)}}^{\top}{V^{(2)}} {V^{(2)}}^{\top}), \\ W_{22} &= {V^{(2)}} {V^{(2)}}^{\top}. \end{align}\tag{91}\] Then, \[\begin{align} \mathop{\mathrm{Var}}\left( \mathop{\mathrm{tr}}\big( U (M(Y) - M(X)) U^{\top}\big) \right) = \mathop{\mathrm{Var}}\left( \sum_{i,j=1}^{2} \mathop{\mathrm{tr}}\left( W_{ij} {X^{(j)}} \beta_j^{-1} \beta_i^{-1} {X^{(i)}}^{\top}\right) \right). \end{align}\] Hence, to demonstrate that \(\mathop{\mathrm{Var}}\big(\mathop{\mathrm{tr}}(U (M(Y) - M(X)) U^{\top})\big)\) is asymptotically negligible compared to \(\mathop{\mathrm{Var}}\big(\mathop{\mathrm{tr}}(U M(X) U^{\top})\big)\), it suffices to consider any fixed pair \((i,j)\) and establish that \[\begin{align} \label{need01} \frac{ \mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(W_{ij} {X^{(j)}} \beta_j^{-1} \beta_i^{-1} {X^{(i)}}^{\top}\right)\right) }{ \mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left({X^{(j)}} \beta_j^{-1} \beta_i^{-1} {X^{(i)}}^{\top}\right)\right) } \to 0. \end{align}\tag{92}\] By conditioning on \({X^{(i)}}\) (recall \({X^{(j)}}\) is independent of \({X^{(i)}}\) and has mean zero) we have \[\begin{align} \mathop{\mathrm{Var}}\big(\mathop{\mathrm{tr}}\big(W_{ij} {X^{(j)}} \beta_j^{-1}\beta_i^{-1} {X^{(i)}}^{\top}\big)\big) =\mathbb{E}\Big[\mathop{\mathrm{Var}}\big(\mathop{\mathrm{tr}}\big(W_{ij} {X^{(j)}} \beta_j^{-1}\beta_i^{-1} {X^{(i)}}^{\top}\big)\bigm| {X^{(i)}}\big)\Big]. \end{align}\] Writing the trace in entrywise form yields \(\mathop{\mathrm{tr}}\big(W_{ij} {X^{(j)}} \beta_j^{-1}\beta_i^{-1} {X^{(i)}}^{\top}\big) =\sum_{b,c} ({X^{(j)}})_{b,c} s_{b c}\big(({X^{(i)}})\big),\) where, for each pair \((b,c)\), \(s_{b c}\big(({X^{(i)}})\big) :=\sum_{a,d} W_{a b}\,(\beta_j^{-1}\beta_i^{-1})_{c d}\,({X^{(i)}})_{a,d}.\) By independence (up to symmetry), there exist positive constants \(C_1,C_2\) (independent of \(n\)) such that for every fixed realization of \({X^{(i)}}\), \[\label{eq:cond-var-bounds} C_1\sum_{b,c}\mathop{\mathrm{Var}}\big(({X^{(j)}})_{b,c}\big)\,s_{b c}\big(({X^{(i)}})\big)^2 \le \mathop{\mathrm{Var}}\big(\mathop{\mathrm{tr}}(W_{ij} {X^{(j)}} \beta_j^{-1}\beta_i^{-1} {X^{(i)}}^{\top})\bigm|{X^{(i)}}\big) \le C_2\sum_{b,c}\mathop{\mathrm{Var}}\big(({X^{(j)}})_{b,c}\big)\,s_{b c}\big(({X^{(i)}})\big)^2.\tag{93}\] Taking expectation with respect to \({X^{(i)}}\), and using \(\mathop{\mathrm{Var}}(({X^{(j)}})_{b,c})\asymp \rho_n\) uniformly in \((b,c)\), we obtain \[\label{eq:var-W-sum} \mathop{\mathrm{Var}}\big(\mathop{\mathrm{tr}}(W_{ij} {X^{(j)}} \beta_j^{-1}\beta_i^{-1} {X^{(i)}}^{\top})\big) \asymp \rho_n \sum_{b,c}\mathbb{E}\big[s_{b c}\big(({X^{(i)}})\big)^2\big].\tag{94}\] We now expand \(\mathbb{E}\big[s_{b c}(({X^{(i)}}))^2\big]\). Using independence (up to symmetry) of the entries of \({X^{(i)}}\) and the uniform variance bound \(\mathop{\mathrm{Var}}(({X^{(i)}})_{a,d})\asymp\rho_n\), it holds that \[\begin{align} \mathbb{E}\big[s_{b c}\big(({X^{(i)}})\big)^2\big] = \sum_{a,d} \bigg(W_{a b}^2\,(\beta_j^{-1}\beta_i^{-1})_{c d}^2 + W_{d b}^2\,(\beta_j^{-1}\beta_i^{-1})_{c a}^2\bigg)\,\mathbb{E}\big[({X^{(i)}})_{a,d}^2\big], \end{align}\] and hence \[\label{eq:s-square-expectation} \mathbb{E}\big[s_{b c}\big(({X^{(i)}})\big)^2\big]\asymp \rho_n \sum_{a,d} \bigg(W_{a b}^2\,(\beta_j^{-1}\beta_i^{-1})_{c d}^2 + W_{d b}^2\,(\beta_j^{-1}\beta_i^{-1})_{c a}^2\bigg).\tag{95}\] Summing 95 over \(b,c\) and reordering the sums yields \[\begin{align} \sum_{b,c}\mathbb{E}\big[s_{b c}\big(({X^{(i)}})\big)^2\big] \asymp \rho_n\sum_{a,b}\sum_{c,d} W_{a b}^2\,(\beta_j^{-1}\beta_i^{-1})_{c d}^2 \asymp \rho_n\,\|W_{ij}\|_F^2\,\|\beta_j^{-1}\beta_i^{-1}\|_F^2. \end{align}\] Combining this with 94 shows that \[\label{eq:var-weighted-final} \mathop{\mathrm{Var}}\big(\mathop{\mathrm{tr}}(W_{ij} {X^{(j)}} \beta_j^{-1}\beta_i^{-1} {X^{(i)}}^{\top})\big) \asymp \rho_n^2\,\|W_{ij}\|_F^2\,\|\beta_j^{-1}\beta_i^{-1}\|_F^2.\tag{96}\] We compare the above to the variance of the unweighted trace. By an identical conditioning argument, now with the coefficient \(r_{m p}\big(({X^{(i)}})\big):=\sum_{q}(\beta_j^{-1}\beta_i^{-1})_{p q}\,({X^{(i)}})_{m,q},\) appearing in the expansion \(\mathop{\mathrm{tr}}({X^{(j)}} \beta_j^{-1}\beta_i^{-1} {X^{(i)}}^{\top})=\sum_{m,p} ({X^{(j)}})_{m,p}\,r_{m p}(({X^{(i)}}))\), one obtains \[\label{eq:var-unweighted-final} \mathop{\mathrm{Var}}\big(\mathop{\mathrm{tr}}({X^{(j)}} \beta_j^{-1}\beta_i^{-1} {X^{(i)}}^{\top})\big) \asymp \rho_n^2\,n\,\|\beta_j^{-1}\beta_i^{-1}\|_F^2.\tag{97}\] Dividing 96 by 97 yields \[\frac{\mathop{\mathrm{Var}}\big(\mathop{\mathrm{tr}}(W_{ij} {X^{(j)}} \beta_j^{-1}\beta_i^{-1} {X^{(i)}}^{\top})\big)}{\mathop{\mathrm{Var}}\big(\mathop{\mathrm{tr}}({X^{(j)}} \beta_j^{-1}\beta_i^{-1} {X^{(i)}}^{\top})\big)} \asymp \frac{\|W_{ij}\|_F^2}{n}.\] Recalling the definition of \(W_{ij}\) in 91 , we note that \(\|{V^{(1)}}{V^{(1)}}^{\top}\|_F^2 = k\) and \(\|{V^{(2)}}{V^{(2)}}^{\top}\|_F^2 = k\), which follow immediately from the orthonormality of the columns of \({V^{(1)}}\) and \({V^{(2)}}\). Moreover, \[\|{V^{(1)}}{V^{(1)}}^{\top}+{V^{(2)}}{V^{(2)}}^{\top}-{V^{(1)}}{V^{(1)}}^{\top}{V^{(2)}}{V^{(2)}}^{\top}\|_F^2 = \|{V^{(1)}}{V^{(1)}}^{\top}\|_F^2 + \|{V^{(2)}}{V^{(2)}}^{\top}-{V^{(1)}}{V^{(1)}}^{\top}{V^{(2)}}{V^{(2)}}^{\top}\|_F^2,\] which is bounded above by \(2k\) and below by \(k\). Consequently, \(\|W_{ij}\|_F^2 \asymp k\). It follows that \[\frac{\mathop{\mathrm{Var}}\!\big(\mathop{\mathrm{tr}}(W_{ij} {X^{(j)}} \beta_j^{-1}\beta_i^{-1} {X^{(i)}}^{\top})\big)}{\mathop{\mathrm{Var}}\!\big(\mathop{\mathrm{tr}}({X^{(j)}} \beta_j^{-1}\beta_i^{-1} {X^{(i)}}^{\top})\big)} \asymp \frac{k}{n} \longrightarrow 0,\] which completes the argument for 92 , thereby proving the lemma. ◻
Proof of 8. Step 1: computing the mean. For independent, mean-zero matrices \({X^{(1)}}\) and \({X^{(2)}}\), \[\mathbb{E}\bigl\|\beta^{\perp}({X^{(1)}} \beta_1^{-1} - {X^{(2)}} \beta_2^{-1})\bigr\|_F^2
= \mathbb{E}\,\mathop{\mathrm{tr}}\bigl(\beta^{\perp} {X^{(1)}} \beta_1^{-2} {X^{(1)}}\bigr)
+ \mathbb{E}\,\mathop{\mathrm{tr}}\bigl(\beta^{\perp} {X^{(2)}} \beta_2^{-2} {X^{(2)}}\bigr),\] since the cross terms vanish by independence. Using the one-sample expectation result from 3, for each \(i=1,2\), \[\begin{align}
\mathbb{E}\,({X^{(i)}} \beta_i^{\perp} {X^{(i)}})
= {\Sigma^{(i)}} \circ \beta_i^{\perp}
+ \mathop{\mathrm{Diag}}\bigl({\Sigma^{(i)}} \cdot d_i - \mathrm{diag}\left({\Sigma^{(i)}}\right) \circ d_i \bigr),
\end{align}\] where \(d_i = \mathrm{diag}(\beta_i^{\perp})\). Consequently, the leading mean term satisfies \[\begin{align}
\mu_2
= 2\sum_{i=1}^{2} \mathop{\mathrm{tr}}\left[ \beta_i^{-2} \Bigl( {\Sigma^{(i)}} \circ \beta_i^{\perp}
+ \mathop{\mathrm{Diag}}\bigl({\Sigma^{(i)}} \cdot d_i - \mathrm{diag}\left({\Sigma^{(i)}}\right) \circ d_i \bigr) \Bigr) \right].
\end{align}\]
Step 2: Computing the variance. According to 7, it suffices to compute \[\mathop{\mathrm{Var}}\bigg(
2\,\mathop{\mathrm{tr}}\left(
U^{\top}
\begin{pmatrix}
{X^{(1)}}^{\top}{X^{(1)}} & -{X^{(1)}}^{\top}{X^{(2)}}\\[4pt]
-{X^{(2)}}^{\top}{X^{(1)}} & {X^{(2)}}^{\top}{X^{(2)}}
\end{pmatrix}
U
\right) \bigg).\] Let \[\begin{align}
Q_i = {(X^{(1)})_{i.}}^{\top}\beta_1^{-2} {(X^{(1)})_{i.}} + {(X^{(2)})_{i.}}^{\top}\beta_2^{-2} {(X^{(2)})_{i.}} - {(X^{(1)})_{i.}}^{\top}\beta_1^{-1}\beta_2^{-1} {(X^{(2)})_{i.}} - {(X^{(2)})_{i.}}^{\top}\beta_2^{-1}\beta_1^{-1} {(X^{(1)})_{i.}},
\end{align}\] and define \(\widetilde{T}_{2}^{(S)} = 2\sum_{i=1}^{n} Q_i.\) Then \[\begin{align} \label{eq:ind1}
\mathop{\mathrm{Var}}(\widetilde{T}_{2}^{(S)})
= 4\sum_{i=1}^{n} \mathop{\mathrm{Var}}(Q_i) + 8\sum_{1\le i<k\le n} \mathop{\mathrm{Cov}}(Q_i,Q_k).
\end{align}\tag{98}\] First, we will calculate \(\mathop{\mathrm{Var}}(Q_i).\) For each sample \(r\in\{1,2\}\), the variance of the individual quadratic form \(Q_i^{(r)} = {(X^{(r)})_{i.}}^{\top}\beta_r^{-2} {(X^{(r)})_{i.}}\) follows directly from the one-sample calculation (see 11.3) \[\begin{align}
\label{eq:ind2}
\mathop{\mathrm{Var}}(Q_i^{(r)})
= 4 \sum_{s<t} (\beta_r^{-2})_{st}^2\, \Sigma_{is}^{(r)} \Sigma_{it}^{(r)} + \sum_{s=1}^n (\beta_r^{-2})_{ss}^2\, K_{is}^{(r)},
\end{align}\tag{99}\] where \(\Sigma_{ij}^{(r)}=\mathop{\mathrm{Var}}({(X^{(r)})_{ij}})\) and \(K_{ij}^{(r)} = \mathbb{E}\left({(X^{(r)})_{ij}}^4\right) -
\mathbb{E}\left({(X^{(r)})_{ij}}^2\right)^2\). Since the vectors \({(X^{(r)})_{i.}}\) are centered, it follows immediately that \({(X^{(r_1)})_{i.}}^{\top}\,\beta_{1}^{-1}\beta_{2}^{-1}\,{(X^{(r_2)})_{i.}}\) and \({(X^{(r_1)})_{i.}}^{\top}\,\beta_{1}^{-1}\beta_{2}^{-1}\,{(X^{(r_1)})_{i.}}\) are independent whenever \((r_{1},r_{2})=(1,2)\). Hence, for the purpose of evaluating \(\mathop{\mathrm{Var}}(Q_{i})\), it suffices to compute \(\mathop{\mathrm{Var}}\left({(X^{(1)})_{i.}}^{\top}\,\beta_{1}^{-1}\beta_{2}^{-1}\,{(X^{(2)})_{i.}}
+ {(X^{(2)})_{i.}}^{\top}\,\beta_{2}^{-1}\beta_{1}^{-1}\,{(X^{(1)})_{i.}}\right).\)
A direct calculation shows that \(\; \mathop{\mathrm{Var}}\left({(X^{(1)})_{i.}}^{\top}\,\beta_{1}^{-1}\beta_{2}^{-1}\,{(X^{(2)})_{i.}}\right) = \sum_{s,t} M_{s,t}^2 \Sigma_{is}^{(1)} \Sigma_{is}^{(2)},\) where \(M := \beta_{1}^{-1}\beta_{2}^{-1}.\) Therefore, the total variance contribution from the bilinear cross–sample term equals \[\begin{align} \label{eq:ind3} \mathop{\mathrm{Var}}\left({(X^{(1)})_{i.}}^{\top}\,\beta_{1}^{-1}\beta_{2}^{-1}\,{(X^{(2)})_{i.}} + {(X^{(2)})_{i.}}^{\top}\,\beta_{2}^{-1}\beta_{1}^{-1}\,{(X^{(1)})_{i.}}\right) = \sum_{s,t} G_{st}^{2}\,\Sigma_{is}^{(1)}\Sigma_{it}^{(2)}, \end{align}\tag{100}\] where \(G := M + M^{\top},\) which completes the calculation.
We now compute the covariance terms. Decompose \(Q_i=Q_i^{(1)}+Q_i^{(2)}-Q_i^{(1,2)}\) such that \(Q_i^{(1)} = {(X^{(1)})_{i.}}^{\top}\beta_1^{-2} {(X^{(1)})_{i.}}, Q_i^{(2)} = {(X^{(2)})_{i.}}^{\top}\beta_2^{-2} {(X^{(2)})_{i.}}, Q_i^{(1,2)} = {(X^{(1)})_{i.}}^{\top}M {(X^{(2)})_{i.}} + {(X^{(2)})_{i.}}^{\top}M^{\top}{(X^{(1)})_{i.}}.\) By bilinearity of covariance, \[\begin{align} \mathop{\mathrm{Cov}}(Q_i, Q_k) &= \mathop{\mathrm{Cov}}(Q_i^{(1)}, Q_k^{(1)}) + \mathop{\mathrm{Cov}}(Q_i^{(2)}, Q_k^{(2)}) - \mathop{\mathrm{Cov}}(Q_i^{(1)}, Q_k^{(1,2)}) - \mathop{\mathrm{Cov}}(Q_i^{(2)}, Q_k^{(1,2)}) \\[3pt] &\quad - \mathop{\mathrm{Cov}}(Q_i^{(1,2)}, Q_k^{(1)}) - \mathop{\mathrm{Cov}}(Q_i^{(1,2)}, Q_k^{(2)}) + \mathop{\mathrm{Cov}}(Q_i^{(1,2)}, Q_k^{(1,2)}). \end{align}\] Observe that \(\mathop{\mathrm{Cov}}(Q_i^{(1)},Q_k^{(1,2)})=0\) because \(Q_k^{(1,2)}\) is a linear combination of products \({X^{(1)}_{ks}}{X^{(2)}_{kt}}\), and by independence of \({X^{(1)}}\) and \({X^{(2)}}\) together with \(\mathbb{E}[{X^{(2)}_{kt}}]=0\), every term in \(\mathbb{E}[Q_i^{(1)} Q_k^{(1,2)}]\) vanishes. The same reasoning yields \(\mathop{\mathrm{Cov}}(Q_i^{(2)},Q_k^{(1,2)})=\mathop{\mathrm{Cov}}(Q_i^{(1,2)},Q_i^{(1)})=\mathop{\mathrm{Cov}}(Q_i^{(1,2)},Q_k^{(2)})=0\), and independence across samples gives \(\mathop{\mathrm{Cov}}(Q_i^{(1)},Q_k^{(2)})=\mathop{\mathrm{Cov}}(Q_i^{(2)},Q_k^{(1)})=0\). Hence, \[\begin{align} \mathop{\mathrm{Cov}}(Q_i,Q_k) = \mathop{\mathrm{Cov}}(Q_i^{(1)},Q_k^{(1)}) + \mathop{\mathrm{Cov}}(Q_i^{(2)},Q_k^{(2)}) + \mathop{\mathrm{Cov}}(Q_i^{(12)},Q_k^{(12)}). \end{align}\] Expanding \(Q_i^{(1)} = \sum_{s,t} (\beta_1^{-2})_{st}\, {X^{(1)}_{is}} {X^{(1)}_{it}},\) and \(Q_k^{(1)} = \sum_{u,v} (\beta_1^{-2})_{uv}\, {X^{(1)}_{ku}} {X^{(1)}_{kv}},\) the only shared random variable for \(i\neq k\) is \({X^{(1)}_{ik}}={X^{(1)}_{ki}}\). Thus all cross-terms vanish except the product involving \(\left({X^{(1)}_{ik}}\right)^2\), leading to \[\begin{align} \mathop{\mathrm{Cov}}(Q_i^{(1)},Q_k^{(1)}) = \left(\mathbb{E}\left[\left({X^{(1)}_{ik}}\right)^4\right] - \mathbb{E}\left[\left({X^{(1)}_{ik}}\right)^2\right]^2\right)\, \left(\beta_1^{-2}\right)_{ii} \left(\beta_1^{-2}\right)_{kk} = K_{ik}^{(1)} (\beta_1^{-2})_{ii} (\beta_1^{-2})_{kk}. \end{align}\] Analogously, \(\mathop{\mathrm{Cov}}(Q_i^{(2)},Q_k^{(2)}) = K_{ik}^{(2)} (\beta_2^{-2})_{ii} (\beta_2^{-2})_{kk}.\) Writing \(Q_i^{(12)} = \sum_{s,t} G_{st}\, {X^{(1)}_{is}} {X^{(2)}_{it}},\) and exploiting the independence of \({X^{(1)}}\) and \({X^{(2)}}\), \[\begin{align} \mathbb{E}[Q_i^{(12)} Q_k^{(12)}] = \sum_{s,t,u,v} G_{st} G_{uv}\, \mathbb{E}[{X^{(1)}_{is}}{X^{(1)}_{ku}}]\,\mathbb{E}[{X^{(2)}_{it}}{X^{(2)}_{kv}}]. \end{align}\] For \(i\neq k\), the only index combination contributing a nonzero expectation is \((s,u)=(k,i)\) and \((t,v)=(i,k)\), yielding \(\mathop{\mathrm{Cov}}(Q_i^{(12)}, Q_k^{(12)}) = G_{ik}^2\, \Sigma_{ik}^{(1)} \Sigma_{ik}^{(2)}.\) Combining all nonzero contributions, for \(i\neq k\), \[\begin{align} \label{eq:ind4} \mathop{\mathrm{Cov}}(Q_i,Q_k) = K_{ik}^{(1)} (\beta_1^{-2})_{ii} (\beta_1^{-2})_{kk} + K_{ik}^{(2)} (\beta_2^{-2})_{ii} (\beta_2^{-2})_{kk} + G_{ik}^2\, \Sigma_{ik}^{(1)} \Sigma_{ik}^{(2)}. \end{align}\tag{101}\] Let \(\alpha_r=\mathrm{diag}(\beta_r^{-2})\). Summing the within and between index contributions from 98 99 100 101 , and including the cross-term variance from the bilinear component, yields \[\begin{align} \mathop{\mathrm{Var}}(\widetilde{T}_{2}^{(S)}) &= 8 \sum_{r=1}^{2} \Big\langle (J_n-I_n) \circ (\beta_r^{-2}\circ\beta_r^{-2}),\, (\Sigma^{(r)})^2 \Big\rangle + 4 \sum_{r=1}^{2} \alpha_r^{\top}K^{(r)} \alpha_r \\ &\quad + 4 \Big\langle G\circ G, {\Sigma^{(1)}}^{\top}{\Sigma^{(2)}} + (J_n-I_n)\circ({\Sigma^{(1)}}\circ{\Sigma^{(2)}}) \Big\rangle, \end{align}\] where \(G = \beta_1^{-1}\beta_2^{-1} + \beta_2^{-1}\beta_1^{-1}\). From 1 3 4, we have \(\bigl|(\beta^{-2})_{ij}\bigr| \asymp \dfrac{k}{n^{3}\rho_n^{2}},\) \(K^{(r)}_{ij} \asymp \rho_n,\) and \(\Sigma^{(r)}_{ij} \asymp \rho_n\) for \(r=1,2.\) Plugging these values in our expression just like in the last paragraph of 11.3, it follows that \(\mathop{\mathrm{Var}}(\widetilde{T}_{2}^{(S)}) \asymp \frac{k^2}{n^3 \rho_n^2}\). Hence by 7, we have \[\mathop{\mathrm{Var}}\bigl(T_{2}^{(S)}\bigr) = \sum_{i=1}^{2} \Bigl( 8\,\bigl\langle (J_n - I_n)\circ(\beta_i^{-2}\circ\beta_i^{-2}),({\Sigma^{(i)}})^2 \bigr\rangle + 4\,\alpha_i^{\top}K^{(i)} \alpha_i \Bigr) + 4\,\bigl\langle G \circ G, {\Sigma^{(1)}}^{\top}{\Sigma^{(2)}} + (J_n - I_n) \circ ({\Sigma^{(1)}} \circ {\Sigma^{(2)}}) \bigr\rangle + o \left( \frac{k^2}{n^3 \rho_n^2}\right).\] ◻
Proof of 9. We will prove this using a very similar approach to the proof of 4. Suppose that \[\underbrace{2\,\mathop{\mathrm{tr}}\left( U^{\top} \begin{pmatrix} (Y^{(1)})^{\top}(Y^{(1)}) & -(Y^{(1)})^{\top}(Y^{(2)})\\[4pt] -(Y^{(2)})^{\top}(Y^{(1)}) & (Y^{(2)})^{\top}(Y^{(2)}) \end{pmatrix} U \right)}_{T_2^{(S)}} = \underbrace{2\,\mathop{\mathrm{tr}}\left( U^{\top} \begin{pmatrix} {X^{(1)}}^{\top}{X^{(1)}} & -{X^{(1)}}^{\top}{X^{(2)}}\\[4pt] -{X^{(2)}}^{\top}{X^{(1)}} & {X^{(2)}}^{\top}{X^{(2)}} \end{pmatrix} U \right)}_{\widetilde{T}_2^{(S)}} + R,\] with the same notation as in 8. By 2, we have \(\mathop{\mathrm{Var}}(R) / \mathop{\mathrm{Var}}(T_{2}^{(S)}) \to 0,\) and consequently \((R - \mathbb{E}(R))/\sqrt{\mathop{\mathrm{Var}}(T_{2}^{(S)})} \to 0\) in probability. Hence the contribution of \(R\) is asymptotically negligible. Therefore, in establishing the limiting distribution of \(T_2^{(S)},\) it suffices to derive the asymptotic distribution of \(\widetilde{T}_2^{(S)}\) after appropriate centering and scaling. The desired result then follows by a direct application of Slutsky’s theorem. We expand the leading-order term via \[\label{break95step1} \begin{align} n \mathrm{tr}\bigg( U^{\top}& \begin{pmatrix} {X^{(1)}}^{\top}{X^{(1)}} & -{X^{(1)}}^{\top}{X^{(2)}} \\ -{X^{(2)}}^{\top}{X^{(1)}} & {X^{(2)}}^{\top}{X^{(2)}} \end{pmatrix} U \bigg) \\ &= \sum_{i=1}^{n} \sum_{j=1}^{n} \sum_{k=1}^{n} \sum_{l=1}^{n} n\Big[ U_{k j}U_{l j}(x_1)_{i k}(x_1)_{i l} + U_{n+k, j}U_{n+l,j}(x_2)_{i k}(x_2)_{i l} - 2\,U_{k j}U_{n+l,j}(x_1)_{i k}(x_2)_{i l} \Big]. \end{align}\tag{102}\] Regrouping terms using \(\gamma_{a,b} := n\sum_{j=1}^n U_{a j}U_{b j}\) and rearranging, we obtain \[\begin{align} \begin{aligned} n \mathrm{tr}\bigg( U^{\top}& \begin{pmatrix} {X^{(1)}}^{\top}{X^{(1)}} & -{X^{(1)}}^{\top}{X^{(2)}} \\ -{X^{(2)}}^{\top}{X^{(1)}} & {X^{(2)}}^{\top}{X^{(2)}} \end{pmatrix} U \bigg) \\ &= \sum_{i=1}^{n} \sum_{k=1}^{n} \sum_{l=1}^{n} \Big[ \gamma_{k,l}(x_1)_{i k}(x_1)_{i l} + \gamma_{n+k,n+l}(x_2)_{i k}(x_2)_{i l} - 2\,\gamma_{k,n+l}(x_1)_{i k}(x_2)_{i l} \Big] \end{aligned} \notag\\ \begin{align} &= \sum_{i=1}^{n} \sum_{\substack{k,l=1 \\ k>l}}^{n} \Big[ 2\gamma_{k,l}(x_1)_{i k}(x_1)_{i l} + 2\gamma_{n+k,n+l}(x_2)_{i k}(x_2)_{i l} - 2\,\gamma_{k,n+l}(x_1)_{i l}(x_2)_{i k} - 2\,\gamma_{l,n+k}(x_1)_{i k}(x_2)_{i l} \Big] \\ &\quad + \sum_{i=1}^{n} \sum_{k=1}^{n} \Big[ \gamma_{k,k}(x_1)_{i k}^2 + \gamma_{n+k,n+k}(x_2)_{i k}^2 - 2\,\gamma_{k,n+k}(x_1)_{i k}(x_2)_{i k} \Big]\end{align} \notag\\ \begin{align} &= \sum_{\substack{i,k,l=1 \\ k>l,\, k<i}}^{n} \Big[ 2\gamma_{k,l}(x_1)_{i k}(x_1)_{i l} + 2\gamma_{n+k,n+l}(x_2)_{i k}(x_2)_{i l} - 2\,\gamma_{k,n+l}(x_1)_{i l}(x_2)_{i k} - 2\,\gamma_{l,n+k}(x_1)_{i k}(x_2)_{i l} \Big] \\ &\quad + \sum_{\substack{i,k,l=1 \\ l<i,\, k<i}}^{n} \Big[ 2\gamma_{i,l}(x_1)_{i k}(x_1)_{k l} + 2\gamma_{n+i,n+l}(x_2)_{i k}(x_2)_{k l} - 2\,\gamma_{i,n+l}(x_1)_{k l}(x_2)_{i k} - 2\,\gamma_{l,n+i}(x_1)_{i k}(x_2)_{l k} \Big] \\ &\quad + \sum_{\substack{k,l=1 \\ l<k}}^{n} \Big[ 2\gamma_{k,l}(x_1)_{k k}(x_1)_{k l} + 2\gamma_{n+k,n+l}(x_2)_{k k}(x_2)_{k l} - 2\,\gamma_{k,n+l}(x_1)_{k l}(x_2)_{k k} - 2\,\gamma_{l,n+k}(x_1)_{k k}(x_2)_{l k} \Big] \\ &\quad + \sum_{i,k=1}^{n} \Big[ \gamma_{k,k}(x_1)_{i k}^2 + \gamma_{n+k,n+k}(x_2)_{i k}^2 - 2\gamma_{k,n+k}(x_1)_{i k}(x_2)_{i k} \Big] \\ &= \sum_{1 \leq k < i \leq n} \Bigg[ \sum_{1 \leq l < k \leq n} \big( 2\gamma_{k,l}(x_1)_{i k}(x_1)_{i l} + 2\gamma_{n+k,n+l}(x_2)_{i k}(x_2)_{i l} - 2\,\gamma_{k,n+l}(x_1)_{i l}(x_2)_{i k} - 2\,\gamma_{l,n+k}(x_1)_{i k}(x_2)_{i l} \big) \\ &\quad + \sum_{1 \leq l < i \leq n} \big( 2\gamma_{i,l}(x_1)_{i k}(x_1)_{k l} + 2\gamma_{n+i,n+l}(x_2)_{i k}(x_2)_{k l} - 2\,\gamma_{i,n+l}(x_1)_{k l}(x_2)_{i k} - 2\,\gamma_{l,n+i}(x_1)_{i k}(x_2)_{l k} \big) \Bigg] \\ &\quad + \sum_{1 \leq l < k \leq n} \big( 2\gamma_{k,l}(x_1)_{k k}(x_1)_{k l} + 2\gamma_{n+k,n+l}(x_2)_{k k}(x_2)_{k l} - 2\,\gamma_{k,n+l}(x_1)_{k l}(x_2)_{k k} - 2\,\gamma_{l,n+k}(x_1)_{k k}(x_2)_{l k} \big) \\ &\quad + \sum_{1 \leq k < i \leq n} \big[ (\gamma_{k,k} + \gamma_{i,i})(x_1)_{i k}^2 + (\gamma_{n+k,n+k} + \gamma_{n+i,n+i})(x_2)_{i k}^2 - 2(\gamma_{k,n+k} + \gamma_{i,n+i})(x_1)_{i k}(x_2)_{i k} \big] \\ &\quad + \sum_{1 \leq k \leq n} \big[ \gamma_{k,k} (x_1)_{k k}^2 + \gamma_{n+k,n+k} (x_2)_{k k}^2 - 2 \gamma_{n+k,k} (x_1)_{k k} (x_2)_{k k} \big]. \end{align} \notag \end{align}\] Now consider the centered version of the quadratic form above. The centering ensures that each term has mean zero, which is essential for the martingale difference structure that follows. Explicitly, we write \[\begin{align} \label{centered2} &\sum_{1 \leq k < i \leq n} \Bigg[ \sum_{1 \leq l < k \leq n} \big( 2\gamma_{k,l}(x_1)_{i k}(x_1)_{i l} + 2\gamma_{n+k,n+l}(x_2)_{i k}(x_2)_{i l} - 2\,\gamma_{k,n+l}(x_1)_{i l}(x_2)_{i k} - 2\,\gamma_{l,n+k}(x_1)_{i k}(x_2)_{i l} \big) \\ &\quad + \sum_{1 \leq l < i \leq n} \big( 2\gamma_{i,l}(x_1)_{i k}(x_1)_{k l} + 2\gamma_{n+i,n+l}(x_2)_{i k}(x_2)_{k l} - 2\,\gamma_{i,n+l}(x_1)_{k l}(x_2)_{i k} - 2\,\gamma_{l,n+i}(x_1)_{i k}(x_2)_{l k} \big) \Bigg] \\ &\quad + \sum_{1 \leq l < k \leq n} \big( 2\gamma_{k,l}(x_1)_{k k}(x_1)_{k l} + 2\gamma_{n+k,n+l}(x_2)_{k k}(x_2)_{k l} - 2\,\gamma_{k,n+l}(x_1)_{k l}(x_2)_{k k} - 2\,\gamma_{l,n+k}(x_1)_{k k}(x_2)_{l k} \big) \\ &\quad + \sum_{1 \leq k < i \leq n} \big[ (\gamma_{k,k} + \gamma_{i,i})\big((x_1)_{i k}^2 - (\sigma_1)_{i,k}^2\big) + (\gamma_{n+k,n+k} + \gamma_{n+i,n+i})\big((x_2)_{i k}^2 - (\sigma_2)_{i,k}^2\big) \\ &\quad\quad - 2(\gamma_{k,n+k} + \gamma_{i,n+i})(x_1)_{i k}(x_2)_{i k} \big] \\ &\quad + \sum_{1 \leq k \leq n} \big[ \gamma_{k,k} \big((x_1)_{k k}^2 - (\sigma_1)_{k,k}^2\big) + \gamma_{n+k,n+k} \big((x_2)_{k k}^2 - (\sigma_2)_{k,k}^2\big) - 2 \gamma_{n+k,k} (x_1)_{k k} (x_2)_{k k} \big], \end{align}\tag{103}\] where \(\mathbb{E}\,(x_i)_{k_1,k_2}^2=(\sigma_i)_{k_1,k_2}^2\). To employ the martingale central limit theorem, we introduce a filtration that orders the terms appropriately. Define \(\mathscr{F}_t := \sigma({w}_1, \ldots, {w}_t), \; \text{where } t = k + \frac{l(l - 1)}{2}, \; 1 \leq k \leq l \leq n,\) so that \[\mathscr{F}_t = \sigma\big( (x_m)_{ij} : 1 \leq i \leq j < l \;\text{or} \;1 \leq k \leq j < l, \;m=1,2 \big).\] Under this ordering, each summand in 103 can be expressed as a martingale difference with respect to \(\{\mathscr{F}_t\}\). Indeed, for \(1 \leq k \leq i \leq n\) and \(t = k + \frac{i(i-1)}{2}\), \[\begin{align} \label{single2} &\mathbb{E}\Big[(x_1)_{i,k}(b_1)_{i,k} + (x_2)_{i,k}(b_2)_{i,k} + \big((x_1)^2_{i,k}-(\sigma_1)^2_{i,k}\big)(c_1)_{i,k} \\ &\quad + \big((x_2)^2_{i,k}-(\sigma_2)^2_{i,k}\big)(c_2)_{i,k} + (x_1)_{i,k}(x_2)_{i,k}(c_3)_{i,k} \;\big| \;\mathscr{F}_{t-1} \Big] = 0, \end{align}\tag{104}\] where \[\begin{align} (b_1)_{i,k} &= \sum_{1 \leq l < k \leq n} \big( 2 \gamma_{k,l} (x_1)_{i,l} - 2 \gamma_{l,n+k} (x_2)_{i,l} \big) + \sum_{1 \leq l < i \leq n} \big( 2 \gamma_{i,l} (x_1)_{k,l} - 2 \gamma_{l,n+i}(x_2)_{i,k} \big), \\ (b_2)_{i,k} &= \sum_{1 \leq l < k \leq n} \big( 2 \gamma_{n+k,n+l}(x_2)_{i,l} - 2\gamma_{k,n+l}(x_1)_{i,l} \big) + \sum_{1 \leq l < i \leq n} \big( 2 \gamma_{n+i,n+l}(x_2)_{k,l} - 2 \gamma_{i,n+l} (x_1)_{k,l} \big), \\ (c_1)_{i,k} &= \sum_{1 \leq k < i \leq n} \big( \gamma_{k,k} + \gamma_{i,i} \big), \\ (c_2)_{i,k} &= \sum_{1 \leq k < i \leq n} \big( \gamma_{n+k,n+k} + \gamma_{n+i,n+i} \big), \\ (c_3)_{i,k} &= \sum_{1 \leq k < i \leq n} \big( 2\gamma_{k,n+k} + 2\gamma_{i,n+i} \big). \end{align}\] From 104 , the conditional variance of the sum in 103 with respect to \(\{\mathscr{F}_t\}\) denoted by \(\eta\) is given by: \[\begin{align}\label{cond95var2} \eta = \sum_{1\leq k<i \leq n} &\Big[ (b_1)^2_{i,k} (\sigma_1)^2_{i,k} + (b_2)^2_{i,k} (\sigma_2)^2_{i,k} + (c_1)^2_{i,k} (\kappa_1)_{i,k} + (c_2)^2_{i,k} (\kappa_2)_{i,k} \\ &\quad + (c_3)^2_{i,k} (\sigma_1)^2_{i,k} (\sigma_2)^2_{i,k} + (c_1)_{i,k}(b_1)_{i,k} (\theta_1)_{i,k} + (c_2)_{i,k}(b_2)_{i,k} (\theta_2)_{i,k} \Big], \end{align}\tag{105}\] where \((\kappa_m)_{i,k} = \mathbb{E}\big[ (x_m)^2_{i,k} - (\sigma_m)^2_{i,k} \big]^2\) and \((\theta_m)_{i,k} = \mathbb{E}\big[ (x_m)_{i,k} \big( (x_m)^2_{i,k} - (\sigma_m)^2_{i,k} \big) \big]\).
This explicit martingale difference decomposition, together with the variance representation in 105 , provides the necessary framework to invoke the martingale central limit theorem. From the definition \(U = \begin{pmatrix} \beta_1^{-1} & \beta_2^{-1} \end{pmatrix}\) and 1 4 3, it follows that \(U_{aj} \asymp \frac{1}{n^2 \rho_n},\) and hence \(\gamma_{a,b} = n \sum_{j=1}^{n} U_{aj} U_{bj} \asymp\frac{1}{n^2 \rho_n^2} .\) For notational convenience in our subsequent order calculations, we will simply denote \(\gamma_{a,b} \asymp\gamma\).
We first compute the order of \(\mathbb{E}(\eta)\). From the definitions of \(b_i\) and \(c_i\), it is straightforward to verify that \[\begin{align} \mathbb{E}(b_m)^2_{i,k} &\asymp n^2 \gamma^2 \rho_n; \\ \mathbb{E}(c_m)^2_{i,k} &\asymp n^2 \gamma^2, \quad \text{for} \quad m = 1,2,3. \end{align}\] In addition, we have the bounds \((\kappa_m)_{i,k} \asymp \rho_n, \; (\theta_m)_{i,k} \asymp \rho_n, \; (\sigma_m)^2_{i,k} \asymp \rho_n.\) Combining these estimates, we find \(\mathbb{E}(\eta) \asymp n^4 \gamma^2 \rho_n.\)
Next, we determine the order of \(\mathop{\mathrm{Var}}(\eta)\). Since the terms \((c_m)_{i,k}\) are deterministic, they do not contribute directly to the variance; only the stochastic terms involving \((b_m)_{i,k}\) enter into \(\mathop{\mathrm{Var}}(\eta)\). We obtain \[\begin{align} \mathop{\mathrm{Var}}\big( (b_m)^2_{i,k} (\sigma_m)^2_{i,k} \big) &\asymp n^4 \gamma^4 \rho_n^3; \\ \mathop{\mathrm{Var}}\big( (c_1)_{i,k}(b_1)_{i,k} \big) &\asymp n^6 \gamma^4 \rho_n^3, \end{align}\] and \[\begin{align} \mathop{\mathrm{Cov}}\big( (c_1)_{i,k}(b_1)_{i,k}, (b_m)^2_{i,k} (\sigma_m)^2_{i,k} \big) \asymp n^5 \gamma^4 \rho_n^3. \end{align}\] Aggregating these contributions, we conclude that \(\mathop{\mathrm{Var}}(\eta) \asymp n^8 \gamma^4 \rho_n^3.\)
Consequently, we have \[\begin{align} \frac{\mathbb{E}(\eta)}{\sqrt{\mathop{\mathrm{Var}}(\eta)}} = \frac{n^4 \gamma^2 \rho_n}{n^4 \gamma^2 \rho_n^{3/2}} \gg 1. \end{align}\] Furthermore, \(\mathbb{E}(\eta) \asymp n^4 \gamma^2 \rho_n = \frac{1}{\rho_n^3} \to \infty\) while, on the other hand, \(\gamma_{a,b} \asymp \frac{1}{n^2 \rho_n^2} \to 0\) as \(n \to \infty\).
These results together verify the conditions required for the martingale CLT as stated in Lemma 9.12 of [62]. Therefore, the appropriately scaled second-order approximation of our test statistic satisfies \[\begin{align} \frac{ T_{2}^{(S)} - \mathbb{E}[T_{2}^{(S)}] }{\sqrt{\mathop{\mathrm{Var}}(T_{2}^{(S)})}} \xrightarrow{D} \mathcal{N}(0,1) \end{align}\] which establishes the result. ◻
Proof. The third order term is \[\left( \langle {V^{(1)}} {V^{(1)}}^{\top}, S_{1,3}({X^{(1)}}) \rangle + \langle {V^{(2)}} {V^{(2)}}^{\top}, S_{2,3}({X^{(2)}}) \rangle + \sum_{l_1 + l_2 = 3} \langle S_{1,l_1}({X^{(1)}}), S_{2,l_2}({X^{(2)}}) \rangle \right).\] The first two terms are directly \(o \left(\frac{k}{n^2 \rho_n^2} \right)\) using 5. For the mixed terms, we have that \[\begin{align} \label{eq:van} \mathbb{E}\!\left\langle S_{1,1}({X^{(1)}}), S_{2,2}({X^{(2)}}) \right\rangle &= \mathbb{E}_{{X^{(2)}}}\!\left[ \mathbb{E}_{{X^{(1)}}}\!\left( \left\langle S_{1,1}({X^{(1)}}), S_{2,2}({X^{(2)}}) \right\rangle \,\big|\, {X^{(2)}} \right) \right] \\ &= \mathbb{E}_{{X^{(2)}}}\!\left[ \left\langle \mathbb{E}_{{X^{(1)}}}\!\left( S_{1,1}({X^{(1)}}) \right), S_{2,2}({X^{(2)}}) \right\rangle \right] \\ &= 0 . \end{align}\tag{106}\] By the same argument as in 106 , the remaining mixed term \(\mathbb{E}\langle S_{1,2}({X^{(1)}}), S_{2,1}({X^{(2)}})\rangle\) also vanishes identically. Indeed, since \({X^{(1)}}\) and \({X^{(2)}}\) are independent and centered, the inner conditional expectation with respect to either random matrix is zero. This completes the proof of 10. ◻
Proof. This proof follows by identical index–counting applied component-wise as in 6. Without loss of generality, take \(l_1 = 1\) and \(l_2 = 2\). First we expand \(\langle S_{1,1}({X^{(1)}}),\, S_{2,2}({X^{(2)}})\rangle\): \[\begin{align} \label{eq:exp} \langle S_{1,1}({X^{(1)}}),\, S_{2,2}({X^{(2)}})\rangle & = \mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} \beta^{\perp} {X^{(2)}} \beta^{\perp} {X^{(2)}} \beta_2^{-2} + \beta^{\perp} {X^{(1)}} \beta_1^{-1} \beta_2^{-2} {X^{(2)}} \beta^{\perp} {X^{(2)}} \right) \\ & - \mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} \beta^{\perp} {X^{(2)}} \beta_2^{-1} {X^{(2)}} \beta_2^{-1} + \beta^{\perp} {X^{(1)}} \beta_1^{-1} \beta_2^{-1} {X^{(2)}} \beta_2^{-1} {X^{(2)}} \right). \end{align}\tag{107}\] From the expansion in 107 , we will establish that the term whose variance dominates the other terms and hence is the dominant variance contributor to \(\langle S_{1,1}({X^{(1)}}),\, S_{2,2}({X^{(2)}})\rangle,\) arises from the fully separated monomials \(\mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} \beta^{\perp} {X^{(2)}} \beta^{\perp} {X^{(2)}} \beta_2^{-2}\right)\) and \(\mathop{\mathrm{tr}}\left(\beta^{\perp} {X^{(1)}} \beta_1^{-1} \beta_2^{-2} {X^{(2)}} \beta^{\perp} {X^{(2)}} \right)\), and then calculate its order. We first show that \[\frac{ \mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} (\beta^{\perp} {X^{(2)}})(\beta^{\perp} {X^{(2)}})\beta_2^{-2}\right)\right) }{ \mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} {X^{(2)}}^{2}\beta_2^{-2}\right)\right) } \to 1,\] exactly paralleling the one–sample setting as in 11.6.
Consider the mixed trace \(R_{3}
:= \mathop{\mathrm{tr}}\big(\beta_1^{-1} {X^{(1)}} ({X^{(2)}})^{2}\beta_2^{-2}\big).\) Decompose \({X^{(2)}} = \left({V^{(2)}} {V^{(2)}}^{\top}+ \beta^{\perp}\right){X^{(2)}}\) in the definition of \(R_{3}\). The unique term in which all insertions equal \(\beta^{\perp}\) is the fully separated monomial \[R_{4}
:= \mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} (\beta^{\perp} {X^{(2)}})(\beta^{\perp} {X^{(2)}})\beta_2^{-2}\right),\] the two–sample analogue of \(R_{2}\) in 11.6. Every
other summand (a remainder term) contains at least one factor \({V^{(2)}} {V^{(2)}}^{\top}\). Fix a remainder term \(\mathcal{N}
:= \mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} (N_{1} {X^{(2)}})(N_{2} {X^{(2)}})\beta_2^{-2}\right),\) where each of \(N_{1}, N_{2}\) is either \({V^{(2)}}{V^{(2)}}^{\top}\) or
\(\beta^{\perp}\). Expanding the trace into its index form, exactly as in 11.6, yields \[\mathcal{N}
=
\sum_{i_1,i_2,i_3 \atop j_1,j_2,j_3}
c(\mathbf{i},\mathbf{j})\,
({X^{(1)}})_{i_{1},j_{2}}\,
({X^{(2)}})_{i_{2},j_{3}}\,
({X^{(2)}})_{i_{3},j_{1}},\] where \(c(\mathbf{i},\mathbf{j})
:=
(\beta_2^{-2}\beta_1^{-1})_{j_1,i_1}\,
(N_{1})_{j_2,i_2}\,
(N_{2})_{j_3,i_3}.\) Using the same arguments as in 11.6, we define \(\Sigma(\mathcal{N})
:=
\sum_{(\mathbf{k},\mathbf{l}) \sim (\mathbf{i},\mathbf{j})}
c(\mathbf{i},\mathbf{j}) c(\mathbf{k},\mathbf{l})\) which enables us to write \(\mathop{\mathrm{Var}}(\mathcal{N})\asymp \Sigma(\mathcal{N}) \rho_n^3.\) The exact same argument of 11.6 holds here as well. Consider the case where \(N_1=V V^{\top}\) and \(N_2= \beta^{\perp}\). The entrywise bounds for \(N_1\) and \(N_2\) implied by 4 yield \(\Sigma(\mathcal{N}) \asymp
\frac{k^4}{n^5 \rho_n^6}\). Similarly, if \(N_1=N_2=V V^{\top}\), we have \(\Sigma(\mathcal{N}) \asymp \frac{k^6}{n^6 \rho_n^6}\). Finally, if \(N_1=N_2=\beta^{\perp},\) which corresponds to the case \(\mathcal{N}=R_4\) we obtain \(\Sigma(R_4) \asymp \frac{k^2}{n^4 \rho_n^6}\) which implies \[\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} (\beta^{\perp} {X^{(2)}})(\beta^{\perp} {X^{(2)}})\beta_2^{-2}\right)\right) \asymp \Sigma(R_4) \rho_n^3 \asymp \frac{k^2}{n^4 \rho_n^3}.\] It follows
that whenever at least one of \(N_{1}\) or \(N_{2}\) equals \(VV^{\top}\), the variance of the monomial satisfies \(\mathop{\mathrm{Var}}(\mathcal{N}) = o(\mathop{\mathrm{Var}}(R_4)),\) as in 86 .
In contrast, just like in 87 , \[\mathop{\mathrm{Var}}(\mathop{\mathrm{tr}}\big(\beta_1^{-1} {X^{(1)}} {X^{(2)}}^{2}\beta_2^{-2}\big)) \asymp \mathop{\mathrm{Var}}(R_{3})
\asymp
\Sigma(R_{3})\,\rho_n^{3}
\asymp
\frac{k^2}{n^{4}\rho_n^{3}}.\] Using the same logic as in 89 , this immediately proves \[\frac{
\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} (\beta^{\perp} {X^{(2)}})(\beta^{\perp} {X^{(2)}})\beta_2^{-2}\right)\right)
}{
\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} {X^{(2)}}^{2}\beta_2^{-2}\right)\right)
}
\to 1.\] The same argument establishes \[\frac{
\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} \beta^\perp {X^{(2)}} \beta_2^{-1} {X^{(2)}} \beta_2^{-1}\right)\right)
}{
\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} {X^{(2)}} \beta_2^{-1} {X^{(2)}} \beta_2^{-1}\right)\right)
}
\to 1.\] Thus, it remains to show \[\frac{
\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} {X^{(2)}} \beta_2^{-1} {X^{(2)}} \beta_2^{-1}\right)\right)
}{
\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\big(\beta_1^{-1} {X^{(1)}} {X^{(2)}}^{2}\beta_2^{-2}\big)\right)
}
\to 0.\] Write the trace in the numerator as \[R_{2,\mathrm{num}}
:=
\mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} {X^{(2)}} \beta_2^{-1} {X^{(2)}} \beta_2^{-1}\right).\] Using the same argument employed in 11.6, \[\begin{align}
\mathop{\mathrm{Var}}(R_{2,\mathrm{num}})
&\asymp
\rho_n^{3}
\sum_{i_1,i_2 \atop j_1,j_2}
\left[
(\beta_2^{-1}\beta_1^{-1})_{i_1,j_1}\,
(\beta_2^{-1})_{i_2,j_2}
\right]^{2}
\\
&\asymp
\rho_n^{3}
\sum_{i_1,i_2 \atop j_1,j_2}
\frac{k^{6}}{n^{10}\rho_n^{6}}
\asymp
\frac{k^{6}}{n^{6}\rho_n^{3}}.
\end{align}\] Thus, \[\frac{
\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} {X^{(2)}} \beta_2^{-1} {X^{(2)}} \beta_2^{-1}\right)\right)
}{
\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\big(\beta_1^{-1} {X^{(1)}} {X^{(2)}}^{2}\beta_2^{-2}\big)\right)
} = \frac{\mathop{\mathrm{Var}}(R_{2,\mathrm{num}})}{\mathop{\mathrm{Var}}(R_{3})} \asymp \frac{\frac{k^6}{n^6 \rho_n^3}}{\frac{k^2}{n^4 \rho_n^3}}
\to 0.\] Similarly we can show \[\frac{
\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left( \beta^{\perp} {X^{(1)}} \beta_1^{-1} \beta_2^{-1} {X^{(2)}} \beta_2^{-1} {X^{(2)}} \right) \right)
}{
\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} (\beta^{\perp} {X^{(2)}})(\beta^{\perp} {X^{(2)}})\beta_2^{-2}\right)\right)
} \to 0.\] So, in 107 , the variance of the terms \(\mathop{\mathrm{tr}}\left(\beta_1^{-1} {X^{(1)}} \beta^{\perp} {X^{(2)}} \beta^{\perp} {X^{(2)}} \beta_2^{-2}\right)\) and \(\mathop{\mathrm{tr}}\left(\beta^{\perp} {X^{(1)}} \beta_1^{-1} \beta_2^{-2} {X^{(2)}} \beta^{\perp} {X^{(2)}} \right)\) is of order \(k^2/(n^4 \rho_n^3)\). The remaining terms of \(\mathop{\mathrm{Var}}\!\big(\langle S_{1,1}(X), S_{2,2}(X)\rangle\big)\) have variance of order \(o\!\left(k^2/(n^4 \rho_n^3)\right)\). Since the number of summands in \(\langle S_{1,1}(X), S_{2,2}(X)\rangle\) is finite, the claim of the lemma follows. ◻
In this section we prove the additional results from 9.
Proof of 12. Write \(\mathcal{E}_{22}=\{\, \|\widehat{V}\|_{2,\infty} \lesssim \sqrt{k/n}\,\}\). By union bound, \(\mathbb{P}(\mathcal{E}_{{\sf very \;good}}^c)\le\mathbb{P}(\mathcal{E}_{{\sf good}}^c)+\mathbb{P}(\mathcal{E}_{22}^c).\) The first term was controlled in 1, namely \(\mathbb{P}(\mathcal{E}_{{\sf good}}^c)=O(n^{-19}).\) To bound \(\mathbb{P}(\mathcal{E}_{22}^c)\) we invoke Lemma C.6 of [42] (where we have that \(\theta_i \asymp \sqrt{\rho_n}\) for all \(n\) as well as the fact that their result holds under the same assumptions we impose herein). Under 1 and 3 that lemma guarantees that, with probability at least \(1-O(n^{-19})\), there exists an orthogonal alignment matrix \(W_*\) such that \(\|\widehat{V}-V W_*\|_{2,\infty} \lesssim \sqrt{\frac{\log n}{n\rho_n}}\|V\|_{2,\infty}.\) Combining this with the triangle inequality gives, with the same high probability, \(\|\widehat{V}\|_{2,\infty} \le \|\widehat{V}-V W_*\|_{2,\infty} + \|V\|_{2,\infty} \lesssim \Big(1+\sqrt{\frac{\log n}{n\rho_n}}\Big)\|V\|_{2,\infty}.\) 4 implies \(\|V\|_{2,\infty}\lesssim\sqrt{k/n}\) and, since \(\sqrt{\frac{\log n}{n\rho_n}}=o(1)\) under 2, the right-hand side is \(\lesssim\sqrt{k/n}\). Hence \(\mathbb{P}(\mathcal{E}_{22}^c)=O(n^{-19}).\) Combining the two bounds yields \(\mathbb{P}(\mathcal{E}_{{\sf very \;good}}) \ge 1- O(n^{-19}),\) as claimed. ◻
Proof of 13. We begin with the following identity for the difference between the true probability matrix and its rank-\(k\) approximation: \[\begin{align} P - \widehat{A} = P - \widehat{V}\widehat{V}^{\top}P \widehat{V}\widehat{V}^{\top}- \widehat{V}\widehat{V}^{\top}(A-P)\widehat{V}\widehat{V}^{\top}. \end{align}\] Invoking Theorem 1 of [51], we can expand the projection operator \(\widehat{V}\widehat{V}^{\top}\) as \[\begin{align} \widehat{V}\widehat{V}^{\top}= V V^{\top}+ \sum_{l \geq 1} S_{l}(X), \end{align}\] where the operators \(S_{l}(X)\) are defined in 7. Substituting this expansion, we obtain \[\begin{align} \widehat{A} - P &= \widehat{V}\widehat{V}^{\top}P \widehat{V}\widehat{V}^{\top}+ \widehat{V}\widehat{V}^{\top}X \widehat{V}\widehat{V}^{\top}- P \\ &= \big(V V^{\top}+ \sum_{l \geq 1} S_{l}(X)\big)\,P\,\big(V V^{\top}+ \sum_{l \geq 1} S_{l}(X)\big) + \big(V V^{\top}+ \sum_{l \geq 1} S_{l}(X)\big)\,X\,\big(V V^{\top}+ \sum_{l \geq 1} S_{l}(X)\big) - P. \end{align} \label{eq:series-expansion}\tag{108}\] Expanding and collecting terms, we obtain \[\begin{align} \widehat{A} - P &= \sum_{l \geq 1} S_{l}(X)\,P + \sum_{l \geq 1} P\,S_{l}(X) + \sum_{l_1,l_2 \geq 1} S_{l_1}(X)\,P\,S_{l_2}(X) \\ &\quad + V V^{\top}X V V^{\top}+ \sum_{l \geq 1} S_{l}(X)\,X\,V V^{\top}+ \sum_{l \geq 1} V V^{\top}X S_{l}(X) \\ &\quad + \sum_{l_1,l_2 \geq 1} S_{l_1}(X)\,X\,S_{l_2}(X). \end{align}\] From this decomposition, it follows that the first-order term (i.e., the term linear in \(X\)) is \[\begin{align} T_{1}(X) = V V^{\top}X (I - V V^{\top}) + (I-V V^{\top}) X V V^{\top}+ V V^{\top}X V V^{\top}, \end{align}\] while for general \(l \geq 2\), \[\begin{align} T_{l}(X) &= S_{l}(X)\,P + P\,S_{l}(X) + \sum_{l_1+l_2=l} S_{l_1}(X)\,P\,S_{l_2}(X) \\ &\quad + S_{P,l-1}(X)\,X\,V V^{\top}+ X V V^{\top}S_{P,l-1}(X) \\ &\quad + \sum_{l_1+l_2=l-1} S_{l_1}(X)\,X\,S_{l_2}(X). \end{align}\] By Theorem 1 of [51], we have the operator norm bound \(\|S_{l}(X)\| \lesssim \left(\frac{\|X\|}{n\rho_n}\right)^{\!l}.\) Combining this with the decomposition above, we obtain \[\begin{align} \|T_{l}(X)\| &\lesssim (l+1)\,2^{\,l-1}\, (n\rho_n)\, \|S_{l}(X)\| + l\,2^{\,l-2}\,\|X\|\,\|S_{l}(X)\| \\ &\lesssim c_l \left(\frac{\|X\|}{n\rho_n}\right)^{\!l}\,(n\rho_n), \end{align}\] for constants \(c_l>0\) such that \(\log c_l = O(1)\). This completes the proof. ◻
Proof of 14. In this proof, we will assume that \(A\) (in the one-sample case) and \({A^{(i)}}\) (in the
two-sample case) both belong to \(\mathcal{E}_{{\sf very \;good}}\), as defined in 12, which happens with high probability as
already proved.
\[\begin{align} \label{intereq01}
\mathop{\mathrm{tr}}\bigl(VV^{\top}\,\Delta M\bigr) &= \sum_{i,j} (V V^{\top})_{ii}(P_{ij}(1-P_{ij}) - {\widehat P}_{ij}(1-{\widehat P}_{ij}) )\\
&= \sum_{i,j} (V V^{\top})_{ii} (P_{ij}-{\widehat P}_{ij}) + \sum_{i,j} (V V^{\top})_{ii} ({\widehat P}_{ij}^2 - P_{ij}^2 ).
\end{align}\tag{109}\] We will start the proof by first proving that \[\begin{align} \label{claim01} \bigg|\sum_{i,j} (V V^{\top})_{ii} (P_{ij}-{\widehat P}_{ij})\bigg| =
\bigg|\mathop{\mathrm{tr}}\left(VV ^{\top}\mathop{\mathrm{Diag}}(P - \widehat P \right)\bigg| =O_p(\max (k,k \sqrt{\rho_n \log n})).
\end{align}\tag{110}\] By 13, we can expand \(\mathop{\mathrm{Diag}}((\widehat{P}-P)\cdot \mathbf{1}_n)\). We
first use Bernstein’s inequality to bound the first–order term, and then control the combined contribution of higher–order terms.
Define \(u:=V V^{\top}\cdot \mathbf{1}_n, y:=(I - V V^{\top})X u\). By 13, the term of \(\mathop{\mathrm{Diag}}((\widehat{P}-P)\cdot \mathbf{1}_n)\) that is linear in \(X\), is \(\mathop{\mathrm{Diag}}(((I-VV^{\top})XVV^{\top}+ X VV^{\top})\cdot \mathbf{1}_n)\). We will control the first term which is \(\mathop{\mathrm{Diag}}((I-VV^{\top})XVV^{\top}\cdot \mathbf{1}_n)) = \mathop{\mathrm{Diag}}(y)\). The bound for the second term \(\mathop{\mathrm{Diag}}(XVV^{\top}\cdot \mathbf{1}_n))\) follows in exactly the same way. Let \(w:=\mathrm{diag}(V V^{\top})\) and set \(\alpha:=w^{\top}(I - V V^{\top})\). Then \[\label{T} \begin{align} \tau_1 &:=\mathop{\mathrm{tr}}\bigl(V V^{\top}\mathop{\mathrm{Diag}}(y)\bigr) =\sum_{i=1}^n (V V^{\top})_{ii} y_i = w^{\top}y = w^{\top}(I - V V^{\top})X u = \alpha^{\top}X u. \end{align}\tag{111}\] Since \(X\) is symmetric and centered with independent off–diagonal entries \({X_{ij}}\), \(\tau_1\) can be rewritten as \[\label{T1} \begin{align} \tau_1 &=\sum_{i,j=1}^n \alpha_{i} {X_{ij}} u_j =\sum_{i< j}\epsilon_{ij}{X_{ij}} + \sum_{i}\epsilon_{ii}x_{ii}, \end{align}\tag{112}\] where \(\epsilon_{ii}:=\alpha_{i}u_i\) and \(\epsilon_{ij}:=\alpha_{i}u_j+\alpha_{j}u_i\). To bound \(\epsilon_{ij}\) and \(\mathop{\mathrm{Var}}(\tau_1)\), note that \[\label{bds} \|w\| \lesssim kn^{-1/2}, \quad \|\alpha\|=\|w(I-V V^{\top})\| \lesssim kn^{-1/2}, \quad \|u\|_{2}\lesssim \sqrt {kn}, \quad \|u\|_{\infty}\lesssim \sqrt {k}.\tag{113}\]
Next, we have \(\|\alpha\|_\infty \le \|w\|_{\infty} + \|w^{\top}V V^{\top}\|_{\infty} = O\left( \frac{k}{n} \right)\) from 4. Hence \(|\epsilon_{ij}| \le 2\|\alpha\|_\infty\|u\|_\infty =O(k^2 n^{-1}).\) Therefore, \[\begin{align} \mathop{\mathrm{Var}}(\tau_1) =\sum_{i\le j}\epsilon_{ij}^2\,\mathop{\mathrm{Var}}({X_{ij}}) \lesssim\sum_{i\le j}(\alpha_{i}u_j+\alpha_{j}u_i)^2\rho_n \lesssim k^2\rho_n, \end{align}\] and moreover \(\max_{i,j}|\epsilon_{ij}{X_{ij}}|\lesssim \frac{k^2}{n}\). By Bernstein’s inequality, \[\begin{align} \mathbb{P}(|\tau_1|>t) \lesssim \exp\Biggl(-\frac{t^2/2}{k^2\rho_n+k^2t/(3n)}\Biggr). \end{align}\] So by choosing \(t \gtrsim k \sqrt{\rho_n \log n}\), we have \[\begin{align} \label{claim02} \tau_1 = O_p \left( k\sqrt{\rho_n \log n}\right). \end{align}\tag{114}\] This establishes the bound on the first–order term for \(\mathop{\mathrm{tr}}(V V^{\top}\Delta M)\).
The higher–order contributions are given by \(\sum_{l \ge 2} \mathop{\mathrm{tr}}\bigl(V V^{\top}\,\mathop{\mathrm{Diag}}(T_{l}(X)\cdot \mathbf{1}_n)\bigr).\) For each \(l \ge 2\), we have \(\mathop{\mathrm{tr}}\bigl(V V^{\top}\,\mathop{\mathrm{Diag}}(T_{l}(X)\cdot \mathbf{1}_n)\bigr) = \sum_{i=1}^{n} (V V^{\top})_{ii}\,\bigl(T_{l}(X)\cdot \mathbf{1}_n\bigr)_{i} \lesssim \sum_{i=1}^{n} \frac{k}{n}\,\bigl(T_{l}(X)\cdot \mathbf{1}_n\bigr)_{i},\) where the inequality follows from 4. Now \(\sum_{i=1}^{n} \bigl(T_{l}(X)\cdot \mathbf{1}_n\bigr)_{i}\) is the sum of all entries in the vector \(T_{l}(X)\cdot \mathbf{1}_n\). Thus \(\sum_{i=1}^{n} \bigl(T_{l}(X)\cdot \mathbf{1}_n\bigr)_{i} = \langle \mathbf{1}_n,(T_l(X) \mathbf{1}_n) \rangle \le n \|T_l(X)\|\). Consequently, \[\begin{align} \label{claim03} \sum_{l \ge 2} \mathop{\mathrm{tr}}\bigl(V V^{\top}\,\mathop{\mathrm{Diag}}(T_{l}(X)\cdot \mathbf{1}_n)\bigr) \lesssim \frac{k}{n}\, \sum_{i=1}^{n} \bigl(T_{l}(X)\cdot \mathbf{1}_n\bigr)_{i} \le \sum_{l \ge 2} \|T_{l}(X)\| = O_p(k), \end{align}\tag{115}\] where the final bound follows directly from 1 13, which control the operator norms of the higher–order terms in the expansion. Thus 114 115 together prove 110 .
Next, we will bound the second term on the right hand side of 109 by proving that \[\begin{align} \label{claim04} \bigg|\sum_{i,j} (V V^{\top})_{ii} (P_{ij}^2-{\widehat P}_{ij}^2)\bigg| = o_p(\max (k,k \sqrt{\rho_n \log n})). \end{align}\tag{116}\] We decompose the term further: \[\begin{align} \label{intereq02} \bigg|\sum_{i,j} (V V^{\top})_{ii} (P_{ij}^2-{\widehat P}_{ij}^2)\bigg| &= \bigg|\sum_{i,j} (V V^{\top})_{ii} P_{ij}(P_{ij} - \widehat P_{ij}) - \sum_{i,j} (V V^{\top})_{ii}(P_{ij} - \widehat P_{ij})^2\bigg| \\ & \le \bigg|\sum_{i,j} (V V^{\top})_{ii} P_{ij}(P_{ij} - \widehat P_{ij}) \bigg| + \bigg| \sum_{i,j} (V V^{\top})_{ii}(P_{ij} - \widehat P_{ij})^2\bigg|. \end{align}\tag{117}\] The second term on the right hand side of 117 can be bounded using 13: \[\begin{align}\label{intereq03} \bigg| \sum_{i,j} (V V^{\top})_{ii}(P_{ij} - \widehat P_{ij})^2\bigg| &\le \max_i |(V V^{\top})_{ii}| \|P - \widehat P\|_F^2 \\ & \lesssim \frac{k^{3/2}}{n} \|P - \widehat P\|_2^2 = O_p (k^{3/2} \rho_n) = o_p (k ). \end{align}\tag{118}\] The first term on the right hand side of 117 can be expressed as \[\bigg| \sum_{i,j} (V V^{\top})_{ii} P_{ij}(P_{ij} - \widehat P_{ij}) \bigg| = \bigg| \mathop{\mathrm{tr}}(P \mathop{\mathrm{Diag}}(VV ^{\top}) (P - \widehat P)) \bigg|.\] We calculate its order in a manner very similar to what we did for the first term of 109 earlier in the proof. By 13, we can expand \(\mathop{\mathrm{Diag}}(\widehat{P}-P)\). We first use Bernstein’s inequality to bound the first-order term, and then control the combined contribution of higher-order terms: \[\mathop{\mathrm{tr}}(P \mathop{\mathrm{Diag}}(VV ^{\top}) (P - \widehat P)) = \mathop{\mathrm{tr}}(P \mathop{\mathrm{Diag}}(VV ^{\top}) T_1(X)) + \sum_{l \ge 2} \mathop{\mathrm{tr}}(P \mathop{\mathrm{Diag}}(VV ^{\top}) T_l(X)).\] Substituting \(T_1(X)\) from 13, we see that \[\begin{align} \label{intereq04} \tau_2 = \mathop{\mathrm{tr}}(P \mathop{\mathrm{Diag}}(VV ^{\top}) T_1(X)) = \mathop{\mathrm{tr}}(P \mathop{\mathrm{Diag}}(VV ^{\top}) X) = \sum_{i,j}(VV ^{\top})_{ii} P_{ij} X_{ij}. \end{align}\tag{119}\] We can easily check that \(\mathop{\mathrm{Var}}(\tau_2) \asymp k^2 \rho_n^3\) from 1 4 and \(|(VV ^{\top})_{ii} P_{ij} X_{ij}| \le \frac{k \rho_n}{n}\). By Bernstein’s inequality, \[\begin{align} \mathbb{P}(|\tau_2|>t) \lesssim \exp\Biggl(-\frac{t^2/2}{k^2\rho_n^3+k^2 t \rho_n/(3n)}\Biggr). \end{align}\] Choosing \(t \gtrsim k \sqrt{\rho_n^3 \log n}\), we obtain \[\begin{align} \label{claim05} \tau_2 = O_p \left( k\sqrt{\rho_n^3 \log n}\right) = o_p (k ). \end{align}\tag{120}\] The higher-order terms can be bounded using 13 1: \[\begin{align}\label{rcxkhyoz} \sum_{l \ge 2} \mathop{\mathrm{tr}}(P \mathop{\mathrm{Diag}}(VV ^{\top}) T_l(X)) &\le \sum_{l \ge 2} k^{3/2} \|P\| \|\mathop{\mathrm{Diag}}(VV ^{\top})\| \|T_l(X)\| \\ & =O_p \left( \sum_{l \ge 2} k^{5/2} c_l \left(\frac{\|X\|}{n\rho_n}\right)^{\!l}\,n\rho_n^2 \right) = O_p \left( k^{5/2} \rho_n \right) = o_p(k ). \end{align}\tag{121}\] Thus, 120 119 118 together prove 116 . Since we have proved both 116 110 , 28 also follows.
An identical argument applies to 29 , where again we expand \(\Delta\Sigma\) and proceed analogously to the steps used for 28 . ◻
Proof of 15. In this proof, we will assume that \(A\) (in the one-sample case) and \({A^{(i)}}\) (in the
two-sample case) both belong to \(\mathcal{E}_{{\sf very \;good}},\) as defined in 12, which happens with high probability as
already proved.
Proof of 30 . Just as in the proof of 28 , it suffices to show \[\begin{align} \label{claim06}
\bigg|\mathop{\mathrm{tr}}\Bigl((V\Lambda^{-2}V^{\top}-\widehat V\widehat\Lambda^{-2}\widehat V^{\top})\,\mathop{\mathrm{Diag}}({\widehat P})\Bigr)\bigg| = O_p \left(\max \left(\frac{k}{n^3 \rho_n^2},\frac{k \sqrt{\log n}}{n^3 \rho_n^{3/2}}\right)\right),
\end{align}\tag{122}\] because then we can similarly show that \[\begin{align} \bigg|\mathop{\mathrm{tr}}\Bigl((V\Lambda^{-2}V^{\top}-\widehat V\widehat\Lambda^{-2}\widehat
V^{\top})\,\mathop{\mathrm{Diag}}({\widehat P} \circ {\widehat P})\Bigr)\bigg| = o_p \left(\max \left(\frac{k}{n^3 \rho_n^2},\frac{k \sqrt{\log n}}{n^3 \rho_n^{3/2}}\right)\right).
\end{align}\] To that end we first work with \(D:=\mathop{\mathrm{Diag}}(P)\) and prove \[\bigg|\mathop{\mathrm{tr}}\Bigl((V\Lambda^{-2}V^{\top}-\widehat V\widehat\Lambda^{-2}\widehat
V^{\top})\,\mathop{\mathrm{Diag}}(P)\Bigr)\bigg|
=O_p \left(\max \left(\frac{k}{n^3 \rho_n^2},\frac{k \sqrt{\log n}}{n^3 \rho_n^{3/2}}\right)\right).\] Using the algebraic identity \[\begin{align}
\label{dseries}
\widehat V\widehat\Lambda^{-2}\widehat V^\top - V\Lambda^{-2}V^\top
= \big(\widehat V\widehat V^\top - VV^\top\big)\widehat V\widehat\Lambda^{-2}\widehat V^\top
+ V\big(V^\top\widehat V\widehat\Lambda^{-2}-\Lambda^{-2}V^\top\widehat V\big)\widehat V^\top
+ V\Lambda^{-2}V^\top\big(\widehat V\widehat V^\top - VV^\top\big),
\end{align}\tag{123}\] and observing that \[\begin{align}
V\big(V^\top\widehat V\widehat\Lambda^{-2}-\Lambda^{-2}V^\top\widehat V\big)\widehat V^\top
&=V\Lambda^{-2}\big(\Lambda^2V^\top\widehat V - V^\top\widehat V\widehat\Lambda^2\big)\widehat\Lambda^{-2}\widehat V^\top\\
&=V\Lambda^{-2}\big(V^\top P^2\widehat V - V^\top A^2\widehat V\big)\widehat\Lambda^{-2}\widehat V^\top\\
&=V\Lambda^{-2}V^\top\big(P^2-A^2\big)\widehat V\widehat\Lambda^{-2}\widehat V^\top\\
&=-V\Lambda^{-2}V^\top\big(PX+XP+X^2\big)\widehat V\widehat\Lambda^{-2}\widehat V^\top,
\end{align}\] we expand the leading (first–order) contributions using the series representation \(\widehat V\widehat V^{\top}- VV^{\top}= S_1(X) + \sum_{j\ge 2} S_j(X),\) with \(S_1(X)=\beta^{\perp}X\beta^{-1}+\beta^{-1}X\beta^{\perp}.\) Collecting the principal terms and using 123 yields the following bound: \[\label{dseries1}
\begin{align}
\bigg|\mathop{\mathrm{tr}}&\Bigl((\widehat V\widehat\Lambda^{-2}\widehat V^{\top}-V\Lambda^{-2}V^{\top})\,\mathop{\mathrm{Diag}}(P)\Bigr)\bigg| \\
&\le \Big|\mathop{\mathrm{tr}}\Big(\big[V\Lambda^{-2}V^{\top}S_1(X)+S_1(X)V\Lambda^{-2}V^{\top}-\,V\Lambda^{-2}V^{\top}(XP+PX)V\Lambda^{-2}V^{\top}\big]\mathop{\mathrm{Diag}}(P)\Big)\Big|\\[4pt]
&\quad+\Big|\mathop{\mathrm{tr}}\Big(S_1(X)\big(\widehat V\widehat\Lambda^{-2}\widehat V^{\top}-V\Lambda^{-2}V^{\top}\big)\mathop{\mathrm{Diag}}(P)\Big)\Big|\\[4pt]
&\quad+\Big|\mathop{\mathrm{tr}}\Big(V\Lambda^{-2}V^{\top}(XP+PX)\big(\widehat V\widehat\Lambda^{-2}\widehat V^{\top}-V\Lambda^{-2}V^{\top}\big)\mathop{\mathrm{Diag}}(P)\Big)\Big|\\[4pt]
&\quad+\Big|\mathop{\mathrm{tr}}\Big(\big(\sum_{j\ge2}V\Lambda^{-2}V^{\top}S_j(X)+\sum_{j\ge2}S_j(X)\widehat V\widehat\Lambda^{-2}\widehat V^{\top}-\,V\Lambda^{-2}V^{\top}X^2\widehat V\widehat\Lambda^{-2}\widehat
V^{\top}\big)\mathop{\mathrm{Diag}}(P)\Big)\Big|.
\end{align}\tag{124}\] The first term on the right hand side of 124 corresponds precisely to those components in the expansion 123 that are linear in the noise matrix \(X\). We bound these contributions using a standard Bernstein–type concentration argument.
Consider first the term \(\mathop{\mathrm{tr}}\left(V\Lambda^{-2}V^{\top}S_{1}(X)\,\mathop{\mathrm{Diag}}(P)\right).\) By direct expansion, this term can be written as \[\mathop{\mathrm{tr}}\left(V\Lambda^{-2}V^{\top}S_{1}(X)\,\mathop{\mathrm{Diag}}(P)\right) =\sum_{t,q=1}^{n} w_{tq}\,x_{tq} =\sum_{t>q}(w_{tq}+w_{qt})x_{tq}+\sum_{t} w_{tt}x_{tt},\] where the deterministic weights \(w_{tq}\) are given by \(w_{tq} =\sum_{s=1}^{k}\sum_{r=1}^{n} \frac{1}{\lambda_{ss}^{3}}\, v_{rs}v_{ts}\,(I-VV^{\top})_{qr}\,p_{rr}.\) Under 1 3 4, these coefficients satisfy \[\frac{k}{n^{4}\rho_{n}^{2}} \lesssim \bigg| \sum_{s=1}^{k} \frac{1}{\lambda_{ss}^{3}}\, v_{rs}v_{ts}\,p_{qq} \bigg| \lesssim |w_{tq}| \lesssim \bigg| \sum_{s=1}^{k} \frac{1}{\lambda_{ss}^{3}}\, v_{qs}v_{ts}\,p_{qq} \bigg| \lesssim \frac{k}{n^{4}\rho_{n}^{2}}.\] Moreover, since the entries \(\{x_{tq}\}\) are independent (up to symmetry) with variance of order \(\rho_n\), we have \(\mathop{\mathrm{Var}}\left(\sum_{t,q=1}^{n} w_{tq}x_{tq}\right) \asymp\frac{k^{2}}{n^{6}\rho_{n}^{3}}.\) Applying Bernstein’s inequality yields, for any \(z>0\), \[\mathbb{P}\left( \left|\sum_{t,q=1}^{n} w_{tq}x_{tq}\right|\ge z \right) \le 2\exp\left( -\frac{z^{2}/2}{ \frac{k^{2}}{n^{6}\rho_{n}^{3}}+\frac{kz}{3n^{4}\rho_{n}^{2}} } \right).\] Taking \(z\asymp \frac{k \sqrt {\log n}}{n^{3}\rho_{n}^{3/2}}\) gives \[\begin{align} \label{claim001} \mathop{\mathrm{tr}}\left(V\Lambda^{-2}V^{\top}S_{1}(X)\,\mathop{\mathrm{Diag}}(P)\right) =O_{p}\left(\frac{k\sqrt {\log n}}{n^{3}\rho_{n}^{3/2}}\right). \end{align}\tag{125}\] A completely analogous argument applies to \(\mathop{\mathrm{tr}}\left(V\Lambda^{-2}V^{\top}X V\Lambda^{-1}V^{\top}\,\mathop{\mathrm{Diag}}(P)\right),\) which can again be written in the form \(\sum_{t,q=1}^{n} w_{tq}x_{tq} =\sum_{t>q}(w_{tq}+w_{qt})x_{tq}+\sum_{t} w_{tt}x_{tt},\) with coefficients \[w_{tq} =\sum_{s,j=1}^{k}\sum_{r=1}^{n} \frac{1}{\lambda_{ss}^{2}\lambda_{jj}}\, v_{rs}v_{ts}v_{qj}v_{rj}\,p_{rr}.\] Using the same approach one obtains \[\begin{align} \label{claim002} \mathop{\mathrm{tr}}\left(V\Lambda^{-2}V^{\top}X V\Lambda^{-1}V^{\top}\,\mathop{\mathrm{Diag}}(P)\right) =O_{p}\left(\frac{k\sqrt {\log n}}{n^{3}\rho_{n}^{3/2}}\right). \end{align}\tag{126}\] The remaining linear terms, \(\mathop{\mathrm{tr}}\left(S_{1}(X)V\Lambda^{-2}V^{\top}\,\mathop{\mathrm{Diag}}(P)\right)\) and \(\mathop{\mathrm{tr}}\left(V\Lambda^{-1}V^{\top}X V\Lambda^{-2}V^{\top}\,\mathop{\mathrm{Diag}}(P)\right),\) are handled in exactly the same manner. Collecting all such contributions from 125 126 , we conclude that the total first–order (linear in \(X\)) component in 124 is \(O_{p}\left(\frac{k\sqrt {\log n}}{n^{3}\rho_{n}^{3/2}}\right).\)
Bounding the final term in 124 is straightforward. We first establish a high–probability bound for \[\Bigg|\mathop{\mathrm{tr}}\Bigg(\sum_{j\ge 2} V\Lambda^{-2}V^{\top}S_j(X)\,\mathop{\mathrm{Diag}}(P)\Bigg)\Bigg|.\] By 1, on the corresponding high–probability event, we have \(\|S_j(X)\| \lesssim \frac{1}{(n\rho_n)^{j/2}},\) for \(j\ge 2.\) Consequently, \[\begin{align}\label{dseries4} \Bigg|\mathop{\mathrm{tr}}\Bigg(\sum_{j\ge 2} V\Lambda^{-2}V^{\top}S_j(X)\,\mathop{\mathrm{Diag}}(P)\Bigg)\Bigg| &\le \sum_{j\ge 2} \Big|\mathop{\mathrm{tr}}\big(V\Lambda^{-2}V^{\top}S_j(X)\,\mathop{\mathrm{Diag}}(P)\big)\Big| \\ &\le k\sum_{j\ge 2} \|\mathop{\mathrm{Diag}}(P)V\Lambda^{-2}V^{\top}\| \, \|S_j(X)\| \\ &=O_p \left( \sum_{j\ge 2} k \frac{\rho_n}{n^2\rho_n^2}\, \frac{4^{j/2}}{(n\rho_n)^{j/2}} \right) \\ &=O_p \left( \frac{k}{n^3\rho_n^2}\right). \end{align}\tag{127}\] Applying an identical argument, we have \[\Bigg|\mathop{\mathrm{tr}}\Bigg(\sum_{j\ge 2} S_j(X) \widehat V \widehat\Lambda^{-2} \widehat V^{\top}\,\mathop{\mathrm{Diag}}(P)\Bigg)\Bigg| =O_p \left( \frac{k}{n^3 \rho_n^2} \right).\] Finally, consider the term \(\mathop{\mathrm{tr}}\Big(V\Lambda^{-2}V^{\top}X^2\widehat V\widehat\Lambda^{-2}\widehat V^{\top}\,\mathop{\mathrm{Diag}}(P)\Big).\) By the Cauchy–Schwarz inequality for the Frobenius inner product, \[\begin{align}\label{dseries5} \Big| \mathop{\mathrm{tr}}\!\Big( V\Lambda^{-2}V^{\top} X^{2}\, \widehat V \widehat\Lambda^{-2}\widehat V^{\top}\, \mathop{\mathrm{Diag}}(P) \Big) \Big| &\le k\, \big\|V\Lambda^{-2}V^{\top} X^{2}\big\|\, \big\|\widehat V \widehat\Lambda^{-2}\widehat V^{\top}\mathop{\mathrm{Diag}}(P)\big\| \\ &\le k\, \|X\|^{2}\, \|V\Lambda^{-2}V^{\top}\|\, \big\|\widehat V \widehat\Lambda^{-2}\widehat V^{\top}\mathop{\mathrm{Diag}}(P)\big\| \\ &=O_p \left( k\, (n\rho_n)\, \frac{1}{n^{2}\rho_n^{2}}\, \frac{\rho_n}{n^{2}\rho_n^{2}} \right) = O_p \left( \frac{k^{2}}{n^{3}\rho_n^{2}} \right), \end{align}\tag{128}\] where the final bound follows from 12 3 1. Combining the above estimates, we conclude that the last term on the right hand side of 124 is of order \(O_p\left(\frac{k}{n^3\rho_n^2}\right).\)
We now turn to bounding the second term on the right hand side of 124 . To this end, we once again decompose \(\widehat V\widehat\Lambda^{-2}\widehat V^\top - V\Lambda^{-2}V^\top\) using the expansion in 123 . This yields \[\begin{align}\label{dseries2} \Big| \mathop{\mathrm{tr}}\Big( S_1(X)\big(\widehat V\widehat\Lambda^{-2}\widehat V^\top - V\Lambda^{-2}V^\top\big)\mathop{\mathrm{Diag}}(P) \Big) \Big| &\le \Big| \mathop{\mathrm{tr}}\Big( S_1(X)\big(\widehat V\widehat V^\top - VV^\top\big) \widehat V\widehat\Lambda^{-2}\widehat V^\top \mathop{\mathrm{Diag}}(P) \Big) \Big| \\ &\quad + \Big| \mathop{\mathrm{tr}}\Big( S_1(X) V\big(V^\top\widehat V\widehat\Lambda^{-2}-\Lambda^{-2}V^\top\widehat V\big) \widehat V^\top \mathop{\mathrm{Diag}}(P) \Big) \Big| \\ &\quad + \Big| \mathop{\mathrm{tr}}\Big( S_1(X) V\Lambda^{-2}V^\top \big(\widehat V\widehat V^\top - VV^\top\big)\mathop{\mathrm{Diag}}(P) \Big) \Big| \\ &\le \sum_{j\ge 1} \Big| \mathop{\mathrm{tr}}\Big( S_1(X) S_j(X)\widehat V\widehat\Lambda^{-2}\widehat V^\top \mathop{\mathrm{Diag}}(P) \Big) \Big| \\ &\quad + \Big| \mathop{\mathrm{tr}}\Big( S_1(X) V\Lambda^{-2}V^\top \big(PX+XP+X^2\big) \widehat V\widehat\Lambda^{-2}\widehat V^\top \mathop{\mathrm{Diag}}(P) \Big) \Big| \\ &\quad + \sum_{j\ge 1} \Big| \mathop{\mathrm{tr}}\Big( S_1(X) V\Lambda^{-2}V^\top S_j(X)\mathop{\mathrm{Diag}}(P) \Big) \Big|. \end{align}\tag{129}\] All terms appearing on the right hand side of 129 can be controlled using 1 5 3 4, together with the high–probability bounds established in 12. Indeed, \[\begin{align} \sum_{j\ge 1} \Big| \mathop{\mathrm{tr}}\Big( S_1(X) S_j(X)\widehat V\widehat\Lambda^{-2}\widehat V^\top \mathop{\mathrm{Diag}}(P) \Big) \Big| &\le \sum_{j\ge 1} k\|S_1(X)\| \, \|S_j(X)\| \, \|\widehat V\widehat\Lambda^{-2}\widehat V^\top \mathop{\mathrm{Diag}}(P)\| \\ &= O_p \left( \sum_{j\ge 1} k \frac{\rho_n}{n^{2.5}\rho_n^{2.5}}\, \frac{4^{j/2}}{(n\rho_n)^{j/2}}\right)\\ &= O_p \left(\frac{k}{n^3\rho_n^2} \right), \end{align}\] where the second line follows similarly as in 127 . An identical bound holds for \[\sum_{j\ge 1} \Big| \mathop{\mathrm{tr}}\Big( S_1(X) V\Lambda^{-2}V^\top S_j(X)\mathop{\mathrm{Diag}}(P) \Big) \Big|\] by the same argument. Moreover, 128 yields the same order of magnitude for \[\Big| \mathop{\mathrm{tr}}\Big( S_1(X) V\Lambda^{-2}V^\top \big(PX+XP+X^2\big) \widehat V\widehat\Lambda^{-2}\widehat V^\top \mathop{\mathrm{Diag}}(P) \Big) \Big|.\] Combining these bounds, we conclude that the second term on the right hand side of 124 is of order \(O_p\left(\frac{k}{n^3\rho_n^2}\right).\)
By an entirely analogous argument, one can show that the third term on the right hand side of 124 is also \(O_p\left(\frac{k}{n^3\rho_n^2}\right).\) Specifically, one again decomposes \(\widehat V\widehat\Lambda^{-2}\widehat V^\top - V\Lambda^{-2}V^\top\) using 123 and applies 1 5 3 4 together with 12. The resulting calculations mirror those used for the second term and therefore yield the stated bound.
We now turn to the deviation term involving \(\mathop{\mathrm{Diag}}(\widehat P)\). Observe that \[\begin{align}\label{claim07}
\mathop{\mathrm{tr}}\bigl((\widehat V \widehat \Lambda^{-2} \widehat V^\top - V \Lambda^{-2} V^\top)\mathop{\mathrm{Diag}}(\widehat P)\bigr)
&=
\mathop{\mathrm{tr}}\bigl((\widehat V \widehat \Lambda^{-2} \widehat V^\top - V \Lambda^{-2} V^\top) \mathop{\mathrm{Diag}}(P)\bigr)
+
\mathop{\mathrm{tr}}\bigl((\widehat V \widehat \Lambda^{-2} \widehat V^\top - V \Lambda^{-2} V^\top)( \mathop{\mathrm{Diag}}(\widehat P - P)\bigr) \\
&=
O_p \left(\max \left(\frac{k}{n^3 \rho_n^2},\frac{k \sqrt{\log n}}{n^3 \rho_n^{3/2}}\right)\right)
+
\mathop{\mathrm{tr}}\bigl((\widehat V \widehat \Lambda ^{-2}\widehat V^\top - V \Lambda^{-2} V^\top)\mathop{\mathrm{Diag}}({\widehat P} - P)\bigr),
\end{align}\tag{130}\] where the bound for the first term follows from 122 . It therefore suffices to control the second term. To this end, we invoke both the decomposition in 123 and the
series expansion in 13, and obtain \[\begin{align}\label{dseries6}
\Big|
\mathop{\mathrm{tr}}\bigl((\widehat V \widehat \Lambda^{-2} \widehat V^\top - V \Lambda^{-2} V^\top)\mathop{\mathrm{Diag}}({\widehat P} - P)\bigr)
\Big|
&\le
\Big|
\mathop{\mathrm{tr}}\bigl((\widehat V\widehat V^\top - VV^\top)
\widehat V\widehat\Lambda^{-2}\widehat V^\top
\mathop{\mathrm{Diag}}({\widehat P} - P)\bigr)
\Big| \\
&\quad +
\Big|
\mathop{\mathrm{tr}}\bigl(
V\big(V^\top\widehat V\widehat\Lambda^{-2}-\Lambda^{-2}V^\top\widehat V\big)
\widehat V^\top
\mathop{\mathrm{Diag}}({\widehat P} - P)\bigr)
\Big| \\
&\quad +
\Big|
\mathop{\mathrm{tr}}\bigl(
V\Lambda^{-2}V^\top(\widehat V\widehat V^\top - VV^\top)
\mathop{\mathrm{Diag}}({\widehat P} - P)\bigr)
\Big| \\
&\le
\sum_{l\ge 1}\sum_{j\ge 1}
\Big|
\mathop{\mathrm{tr}}\bigl(
S_j(X)\widehat V\widehat\Lambda^{-2}\widehat V^\top
\mathop{\mathrm{Diag}}(T_l(X))\bigr)
\Big| \\
&\quad +
\sum_{l\ge 1}
\Big|
\mathop{\mathrm{tr}}\bigl(
V\Lambda^{-2}V^\top(PX+XP+X^2)
\widehat V\widehat\Lambda^{-2}\widehat V^\top
\mathop{\mathrm{Diag}}(T_l(X))\bigr)
\Big| \\
&\quad +
\sum_{l\ge 1}\sum_{j\ge 1}
\Big|
\mathop{\mathrm{tr}}\bigl(
V\Lambda^{-2}V^\top S_j(X)\mathop{\mathrm{Diag}}(T_l(X))\bigr)
\Big|.
\end{align}\tag{131}\] We now derive high–probability bounds for the terms on the right hand side of 131 . Using 1 5 3 4 together with the event in 12 and the series representation in 13, we obtain \[\begin{align}
\sum_{l\ge 1}\sum_{j\ge 1}
\Big|
\mathop{\mathrm{tr}}\bigl(
S_j(X)\widehat V\widehat\Lambda^{-2}\widehat V^\top
\mathop{\mathrm{Diag}}(T_l(X))\bigr)
\Big|
&\le k
\sum_{l,j\ge 1}
\|S_j(X)\| \,
\|\widehat V\widehat\Lambda^{-2}\widehat V^\top\mathop{\mathrm{Diag}}(T_l(X))\| \\
&= O_p \left(
\sum_{l,j\ge 1}
\frac{k}{(n\rho_n)^2}
\frac{4^{j/2}}{(n\rho_n)^{j/2}}
\frac{4^{l/2}}{(n\rho_n)^{l/2}}
\rho_n \right)
= O_p \left(
\frac{k}{n^3\rho_n^2} \right).
\end{align}\] The final inequality follows from the bound \((T_l(X))_{ij}\lesssim
\rho_n\Big(\frac{\|X\|}{n\rho_n}\Big)^{l},\) which holds by the definition of \(T_l(X)\) in 13. The remaining two
terms in 131 can be bounded in an identical manner by applying the same norm inequalities and high–probability controls. Combining all these bounds for the four terms in 124 and 130 completes the proof of 122 .
Proof of 31 This bound follows directly from 30 . Indeed, \[\begin{align}
\big|
\mathop{\mathrm{tr}}\bigl((V \Lambda^{-2} V^{\top}- \widehat V \widehat \Lambda^{-2} \widehat V^{\top})\,
\mathop{\mathrm{Diag}}(\widehat{\Sigma}\cdot \mathbf{1}_n)\bigr)
\big|
&\lesssim
n\,\big|
\mathop{\mathrm{tr}}\bigl((V \Lambda^{-2} V^{\top}- \widehat V \widehat \Lambda^{-2} \widehat V^{\top})\,
\mathop{\mathrm{Diag}}(\widehat{\Sigma}\bigr)
\big| \\
&=O_p \left(\max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right)\right),
\end{align}\] which completes the proof. ◻
Proof of 16. Just like in 13.3, we will assume that \(A\) (in the one-sample case) and \({A^{(i)}}\) (in the two-sample case) both belong to \(\mathcal{E}_{{\sf very \;good}},\) as defined in 12, which happens with high probability as already proved.
Proof of 32 We have \(\alpha=\mathrm{diag}(\beta^{-2})\) and \(\widehat{\alpha}=\mathrm{diag}(\widehat{\beta}^{-2}).\) Once again, we invoke 123 to obtain the bound \[\label{dseriesnew1}
\begin{align}
\bigg\|&
\mathrm{diag}\Bigl(
\widehat V \widehat\Lambda^{-2}\widehat V^{\top}
-
V\Lambda^{-2}V^{\top}
\Bigr)
\bigg\| \\
&\le
\Big\|
\mathrm{diag}\Big(
V\Lambda^{-2}V^{\top} S_1(X)
+
S_1(X)V\Lambda^{-2}V^{\top}
-
V\Lambda^{-2}V^{\top}(XP+PX)V\Lambda^{-2}V^{\top}
\Big)
\Big\| \\[4pt]
&\quad+
\Big\|
\mathrm{diag}\Big(
S_1(X)
\big(
\widehat V \widehat\Lambda^{-2}\widehat V^{\top}
-
V\Lambda^{-2}V^{\top}
\big)
\Big)
\Big\| \\[4pt]
&\quad+
\Big\|
\mathrm{diag}\Big(
V\Lambda^{-2}V^{\top}(XP+PX)
\big(
\widehat V \widehat\Lambda^{-2}\widehat V^{\top}
-
V\Lambda^{-2}V^{\top}
\big)
\Big)
\Big\| \\[4pt]
&\quad+
\Big\|
\mathrm{diag}\Big(
\sum_{j\ge2} V\Lambda^{-2}V^{\top} S_j(X)
+
\sum_{j\ge2} S_j(X)\widehat V \widehat\Lambda^{-2}\widehat V^{\top} -
V\Lambda^{-2}V^{\top} X^2 \widehat V \widehat\Lambda^{-2}\widehat V^{\top}
\Big)
\Big\|.
\end{align}\tag{132}\] Consider first the vector \(r
=
\mathrm{diag}\left(
V\Lambda^{-2}V^{\top} S_{1}(X)
\right)
=
\mathrm{diag}\left(
\beta^{-3} X \beta^{\perp}
\right).\) For each \(i\), \(r_i
=
u_i^{\top} \Lambda^{-3} w_i,\) where \(u_i = V^{\top} e_i\) and \(w_i = V^{\top} X \beta^{\perp} e_i\). By the Cauchy–Schwarz inequality, \(|r_i|^2
\le
\|\Lambda^{-3}\|_2^2
\|u_i\|_2^2
\|w_i\|_2^2.\) Consequently, \[\|r\|_2^2
=
\sum_{i=1}^n r_i^2
\le
\|\Lambda^{-3}\|_2^2
\sum_{i=1}^n
\|u_i\|_2^2
\|w_i\|_2^2
\lesssim
\frac{k}{n}
\|\Lambda^{-3}\|_2^2
\,
\|V^{\top} X \beta^{\perp}\|_F^2,\] where the final inequality follows from the incoherence assumption in 4.
Using standard norm inequalities, \(\|V^{\top} X \beta^{\perp}\|_F \le \|V\|_F \|X\|_2 \|\beta^{\perp}\|_2 \le \sqrt{k}\,\|X\|_2.\) Therefore, \(\|r\|_2 \lesssim \frac{k}{\sqrt{n}} \|X\|_2 \|\Lambda^{-3}\|_2.\) On the event \(\mathcal{E}_{{\sf good}}\), together with the eigen-scaling assumption in 3, we conclude that \(\|r\|_2 = O_p \left( \frac{k}{n^{3}\rho_n^{2.5}}\right).\) Similarly we can also show that \(\|\mathrm{diag}\left(S_{1}(X) V\Lambda^{-2}V^{\top} \right)\| =O_p \left( \frac{k}{n^{3}\rho_n^{2.5}}\right)\).
Next, consider the vector \(r=\mathrm{diag}\left(\beta^{-2} X \beta^{-1}\right),\) for which, for each \(i\), \(r_i = (V^{\top} e_i)^{\top}\Lambda^{-2}(V^{\top}X\beta^{-1}e_i).\) Proceeding exactly as in the previous argument, using Cauchy-Schwarz inequality, the incoherence condition in 4, standard norm inequalities, the bound on event \(\mathcal{E}_{{\sf good}}\) as in 1, together with the eigenvalue scaling in 3, we obtain \(\|\mathrm{diag}(V\Lambda^{-2}V^{\top} X P V\Lambda^{-2}V^{\top})\| =O_p \left( \frac{k}{n^{3}\rho_n^{2.5}} \right).\) The same bound holds for \(\|\mathrm{diag}(V\Lambda^{-2}V^{\top} P X V\Lambda^{-2}V^{\top})\|.\) Consequently, the entire first term on the right-hand side of 132 is bounded by \(O_p \left(\frac{k}{n^{3}\rho_n^{2.5}}\right) .\)
We now turn to bounding the second term in 132 . To this end, we once again decompose \(\widehat V\widehat\Lambda^{-2}\widehat V^\top - V\Lambda^{-2}V^\top\) using the expansion in 123 . This yields \[\begin{align}\label{dseriesnew2} \Big| \mathrm{diag}\Big( S_1(X)\big(\widehat V\widehat\Lambda^{-2}\widehat V^\top - V\Lambda^{-2}V^\top\big) \Big) \Big| &\le \Big| \mathrm{diag}\Big( S_1(X)\big(\widehat V\widehat V^\top - VV^\top\big) \widehat V\widehat\Lambda^{-2}\widehat V^\top \Big) \Big| \\ &\quad + \Big| \mathrm{diag}\Big( S_1(X) V\big(V^\top\widehat V\widehat\Lambda^{-2}-\Lambda^{-2}V^\top\widehat V\big) \widehat V^\top \Big) \Big| \\ &\quad + \Big| \mathrm{diag}\Big( S_1(X) V\Lambda^{-2}V^\top \big(\widehat V\widehat V^\top - VV^\top\big) \Big) \Big| \\ &\le \sum_{j\ge 1} \Big| \mathrm{diag}\Big( S_1(X) S_j(X)\widehat V\widehat\Lambda^{-2}\widehat V^\top \Big) \Big| \\ &\quad + \Big| \mathrm{diag}\Big( S_1(X) V\Lambda^{-2}V^\top \big(PX+XP+X^2\big) \widehat V\widehat\Lambda^{-2}\widehat V^\top \Big) \Big| \\ &\quad + \sum_{j\ge 1} \Big| \mathrm{diag}\Big( S_1(X) V\Lambda^{-2}V^\top S_j(X) \Big) \Big|. \end{align}\tag{133}\] We first focus on the first term \(r_j =\mathrm{diag}\Big( S_1(X)\, S_j(X)\,\widehat V\widehat\Lambda^{-2}\widehat V^{\top} \Big).\) For each \(i\), we may write \[(r_{1j})_i = u_i^{\top}\widehat\Lambda^{-2} w_i, \qquad u_i := e_i^{\top} S_1(X) S_j(X) \widehat V, \quad w_i := \widehat V^{\top} e_i .\] By Cauchy-Schwarz inequality, \((r_{1j})_i^2 \le \|\widehat\Lambda^{-2}\|_2^{\,2}\, \|u_i\|_2^{\,2}\, \|w_i\|_2^{\,2}.\) Summing over \(i\) and using the incoherence of \(\widehat V\) from 12, we obtain \[\label{side1} \begin{align} \|r_{1j}\| &\lesssim \frac{\sqrt{k}}{\sqrt{n}}\, \|\widehat\Lambda^{-2}\|\, \|S_1(X) S_j(X)\widehat V\|_F \\[4pt] &\lesssim \frac{k}{\sqrt{n}}\, \|\widehat\Lambda^{-2}\|\, \|S_1(X) S_j(X)\| \\[4pt] &\lesssim \frac{k}{\sqrt{n}}\, \|\widehat\Lambda^{-2}\|\, \|S_1(X)\|\, \|S_j(X)\| \\[4pt] &=O_p \left( \frac{k}{\sqrt{n}}\, \frac{1}{n^{2}\rho_n^{2}} \left(\frac{4}{n\rho_n}\right)^{(j+1)/2}\right). \end{align}\tag{134}\] Summing the bound in 134 over \(j\ge 1\) yields \[\sum_{j\ge 1} \Big\| \mathrm{diag}\Big( S_1(X)\, S_j(X)\,\widehat V\widehat\Lambda^{-2}\widehat V^{\top} \Big) \Big\| =O_p \left( \frac{k}{n^{3.5}\rho_n^{3}}\right) .\] Similarly, consider \(r_{2j} =\mathrm{diag}\Big( S_1(X)\, V \Lambda^{-2} V^{\top} S_j(X) \Big).\) For each \(i\), we may write \[(r_{2j})_i = \big(e_i^{\top}S_1(X) \big)\, \big(V \Lambda^{-2} V^{\top}\big)\, \big(S_j(X) e_i\big).\] Proceeding analogously to the previous case and invoking Cauchy–Schwarz inequality together with the incoherence of \(V\), we obtain
\[\label{side2} \begin{align} \|r_{2j}\| &\le \|\Lambda^{-2}\|\, \left( \max_{1 \le i \le n} \|e_i^\top S_1(X) V\|_2 \right)\, \|V^\top S_j(X)\|_F \\ &\le \frac{1}{\sqrt{n}}\, \|\Lambda^{-2}\|\, \|V^{\top} S_j(X)\|_F\, \|S_1(X) V\|_F \\[4pt] &\le \frac{k}{n^{2.5}\rho_n^{2}}\, \|S_1(X)\|\, \|S_j(X)\| \\[4pt] &=O_p \left( \frac{k}{n^{2.5}\rho_n}\, \left(\frac{4}{n\rho_n}\right)^{(j+1)/2}\right). \end{align}\tag{135}\] Summing the bound in 135 over \(j\ge 1\) yields \[\sum_{j\ge 1} \Big\| \mathrm{diag}\Big( S_1(X)\, V \Lambda^{-2} V^{\top} S_j(X) \Big) \Big\| =O_p \left( \frac{k}{n^{3.5}\rho_n^{3}}\right) .\] A completely analogous argument applies to the middle term in 133 , \[\mathrm{diag}\Big( S_1(X)\, V\Lambda^{-2}V^{\top} \big(PX+XP+X^2\big) \widehat V\widehat\Lambda^{-2}\widehat V^{\top} \Big),\] and yields the same rate. Consequently, the entire second term in 133 is bounded by \(O_p\left(\frac{k}{n^{3.5}\rho_n^{3}}\right).\)
By an entirely analogous argument, one can show that the third term in 132 is also \(O_p\left(\frac{k^{3/2}}{n^{3.5}\rho_n^{3}}\right).\) Specifically, one again decomposes \(\widehat V\widehat\Lambda^{-2}\widehat V^\top - V\Lambda^{-2}V^\top\) using 123 and applies 1 5 3 4 together with 12. The resulting calculations mirror those used for the second term and therefore yield the stated bound.
For the final term in 132 , we first focus on \(r_{3j} = \mathrm{diag}\Big( V\Lambda^{-2}V^{\top} S_j(X) \Big).\) For each \(i\), we may write \[(r_{3j})_i = u_i^{\top}\Lambda^{-2} w_i, \qquad u_i = V^{\top} e_i, \quad w_i = V^{\top} S_j(X) e_i .\] Proceeding along the same lines as in 134 , and using Cauchy–Schwarz inequality together with the incoherence of \(V\), we obtain a bound on each \(\|r_{3j}\|\). Summing over \(j\ge 2\) then yields \(\sum_{j\ge 2} \|r_{3j}\| = O_p\!\left(\frac{k}{n^{3.5}\rho_n^{3}}\right).\)
The same argument applies to the remaining two terms in the final line of 132 , namely \[\sum_{j\ge 2}
\mathrm{diag}\Big(
S_j(X)\widehat V \widehat\Lambda^{-2}\widehat V^{\top}
\Big)
\quad\text{and}\quad
\mathrm{diag}\Big(
V\Lambda^{-2}V^{\top} X^2 \widehat V \widehat\Lambda^{-2}\widehat V^{\top}
\Big),\] and both admit the bound \(O_p\!\left(\frac{k}{n^{3.5}\rho_n^{3}}\right).\) Adding up the bounds obtained for each term in 132 , we conclude that \[\|\widehat{\alpha}-\alpha\| = O_p \left(\frac{k}{n^{3}\rho_n^{2.5}} \right),\] which completes the proof of 32 .
Proof of 33 . From 11 , \(G = \beta_1^{-1}\beta_2^{-1} + \beta_2^{-1}\beta_1^{-1}\) and \(\widehat{G} =
\widehat{\beta}_1^{-1}\widehat{\beta}_2^{-1} + \widehat{\beta}_2^{-1}\widehat{\beta}_1^{-1}.\) Hence, \[\begin{align}
\|\widehat{G} \circ \widehat{G} - G \circ G\|_F
&= \biggl\|\sum_{i_1 \neq i_2}\sum_{j_1 \neq j_2}
\Bigl[\bigl(\widehat{\beta}_{i_1}^{-1}\widehat{\beta}_{i_2}^{-1}\bigr)\circ\bigl(\widehat{\beta}_{j_1}^{-1}\widehat{\beta}_{j_2}^{-1}\bigr)
-\bigl(\beta_{i_1}^{-1}\beta_{i_2}^{-1}\bigr)\circ\bigl(\beta_{j_1}^{-1}\beta_{j_2}^{-1}\bigr)\Bigr]\biggr\|_F \\
&\le \sum_{i_1 \neq i_2}\sum_{j_1 \neq j_2}
\bigl\|\bigl(\widehat{\beta}_{i_1}^{-1}\widehat{\beta}_{i_2}^{-1}\bigr)\circ\bigl(\widehat{\beta}_{j_1}^{-1}\widehat{\beta}_{j_2}^{-1}\bigr)
-\bigl(\beta_{i_1}^{-1}\beta_{i_2}^{-1}\bigr)\circ\bigl(\beta_{j_1}^{-1}\beta_{j_2}^{-1}\bigr)\bigr\|_F.
\end{align}\] We present the argument for one representative configuration \((i_1,i_2,j_1,j_2)=(1,2,1,2)\); all other cases follow identically. Then \[\begin{align} \label{intereq}
\bigl\|\bigl(\widehat{\beta}_1^{-1}\widehat{\beta}_2^{-1}\bigr)\circ\bigl(\widehat{\beta}_1^{-1}\widehat{\beta}_2^{-1}\bigr)
-\bigl(\beta_1^{-1}\beta_2^{-1}\bigr)\circ\bigl(\beta_1^{-1}\beta_2^{-1}\bigr)\bigr\|_F
&\le
\bigl\|\bigl(\widehat{\beta}_1^{-1}\widehat{\beta}_2^{-1}\bigr)\circ
\bigl(\widehat{\beta}_1^{-1}\widehat{\beta}_2^{-1}-\beta_1^{-1}\beta_2^{-1}\bigr)\bigr\|_F \\
&\quad + \bigl\|\bigl(\widehat{\beta}_1^{-1}\widehat{\beta}_2^{-1}-\beta_1^{-1}\beta_2^{-1}\bigr)\circ
\bigl({\beta}_1^{-1}{\beta}_2^{-1}\bigr)\bigr\|_F.
\end{align}\tag{136}\] We begin by bounding the term \(\|\widehat{\beta}_1^{-1}\widehat{\beta}_2^{-1}-\beta_1^{-1}\beta_2^{-1}\|_F .\) By the triangle inequality, \[\begin{align}
\label{ds}
\|\widehat\beta_1^{-1} \widehat\beta_2^{-1} - \beta_1^{-1} \beta_2^{-1}\|_F
&\le
\|\widehat\beta_1^{-1}(\widehat\beta_2^{-1} - \beta_2^{-1})\|_F
+
\| (\widehat\beta_1^{-1} - \beta_1^{-1})\beta_2^{-1}\|_F \notag\\
&\le \|\widehat\beta_1^{-1} \|_F\|\widehat\beta_2^{-1} - \beta_2^{-1}\|_F
+
\| \widehat\beta_1^{-1} - \beta_1^{-1}\|_F \|\beta_2^{-1}\|_F .
\end{align}\tag{137}\] Since we know \(\|\beta^{-1}_2\|_2 \asymp \frac{1}{n\rho_n}\) and \(\|\widehat \beta^{-1}_1\|_2 \asymp \frac{1}{n\rho_n}\) from 3, it suffices to derive a bound for the generic quantity \(\|\widehat\beta^{-1} - \beta^{-1}\|_F .\)
To this end, we employ the identity \[\begin{align} \label{dseriesmod} \widehat V\widehat\Lambda^{-1}\widehat V^\top - V\Lambda^{-1}V^\top &= \big(\widehat V\widehat V^\top - VV^\top\big)\widehat V\widehat\Lambda^{-1}\widehat V^\top \\ &\quad + V\big(V^\top\widehat V\widehat\Lambda^{-1}-\Lambda^{-1}V^\top\widehat V\big)\widehat V^\top \nonumber\\ &\quad + V\Lambda^{-1}V^\top\big(\widehat V\widehat V^\top - VV^\top\big), \nonumber \end{align}\tag{138}\] which is a first–order analogue of 123 . Observe that \[\begin{align} V\big(V^\top\widehat V\widehat\Lambda^{-1}-\Lambda^{-1}V^\top\widehat V\big)\widehat V^\top &= V\Lambda^{-1}\big(\Lambda V^\top\widehat V - V^\top\widehat V\widehat\Lambda\big) \widehat\Lambda^{-1}\widehat V^\top \\ &= -\,V\Lambda^{-1}V^\top X \widehat V\widehat\Lambda^{-1}\widehat V^\top . \end{align}\] Using 138 , we obtain \[\label{ds1} \begin{align} \|\widehat\beta^{-1} - \beta^{-1}\|_F &\le \|\widehat V\widehat V^\top - VV^\top\|_F \, \|\widehat\beta^{-1}\|_2 + \|\widehat\beta^{-1}\|_2 \, \|X\|_2 \, \|\beta^{-1}\|_F \\ &\qquad + \|\widehat V\widehat V^\top - VV^\top\|_F \, \|\beta^{-1}\|_2 \\ &=O_p \left( \frac{k^{1/2}}{n^{3/2}\rho_n^{3/2}} \right), \end{align}\tag{139}\] where the final line follows from the Davis–Kahan bound in 1 together with 3. 3 also implies \(\|\beta^{-1}\|_2 \asymp \frac{1}{n\rho_n}\) and \(\|\widehat \beta^{-1}\|_2 \asymp \frac{1}{n\rho_n}\).
Thus, substituting 139 into 137 yields \[\|\widehat\beta_1^{-1} \widehat\beta_2^{-1} - \beta_1^{-1} \beta_2^{-1}\|_F \lesssim \frac{k^{1/2}}{n^{5/2}\rho_n^{5/2}} .\] Inserting this bound into 136 and invoking 3 4 along with the incoherence of \(\widehat V\), guaranteed by 12, we conclude that \[\begin{align} \bigl\| (\widehat{\beta}_1^{-1}\widehat{\beta}_2^{-1})\circ (\widehat{\beta}_1^{-1}\widehat{\beta}_2^{-1}) - (\beta_1^{-1}\beta_2^{-1})\circ (\beta_1^{-1}\beta_2^{-1}) \bigr\|_F &\lesssim \max\!\left( \|\widehat{\beta}_1^{-1}\widehat{\beta}_2^{-1}\|_{\infty}, \|\beta_1^{-1}\beta_2^{-1}\|_{\infty} \right) \|\widehat\beta_1^{-1} \widehat\beta_2^{-1} - \beta_1^{-1} \beta_2^{-1}\|_F \\ &\lesssim \frac{k^{3/2}}{n^{11/2}\rho_n^{9/2}}. \end{align}\] This is because \[\begin{align} \big|(\beta_1^{-1}\beta_2^{-1})_{ij}\big| &\le \|e_i^{\top}V_1\|_2 \|\Lambda_1^{-1}\|_2 \|V_1^{\top}V_2\|_2 \|\Lambda_2^{-1}\|_2 \|V_2^{\top}e_j\|_2 \\ &\lesssim \frac{k}{n^3 \rho_n^2}. \end{align}\] Here, we used 3 4 to obtain the bound. Similarly, in order to provide the same bound for \(\|\widehat{\beta}_1^{-1}\widehat{\beta}_2^{-1}\|_{\infty}\), we will use the incoherence of \(\widehat V\), guaranteed by 12.
Summing over all index combinations yields the desired bound \[\|\widehat{G} \circ \widehat{G} - G \circ G\|_F
=O_p \left(
\frac{k^{3/2}}{n^{11/2}\rho_n^{9/2}}\right) .\]
Proof of 34 . Once again, invoking 123 yields \[\label{dseriesnew10}
\begin{align}
\bigg\|
\widehat V &\widehat\Lambda^{-2}\widehat V^{\top}
-
V\Lambda^{-2}V^{\top}
\bigg\|_F \\
&\le
\Big\|
V\Lambda^{-2}V^{\top} S_1(X)
+
S_1(X)V\Lambda^{-2}V^{\top}
-
V\Lambda^{-2}V^{\top}(XP+PX)V\Lambda^{-2}V^{\top}
\Big\|_F \\[4pt]
&\quad+
\Big\|
S_1(X)
\big(
\widehat V \widehat\Lambda^{-2}\widehat V^{\top}
-
V\Lambda^{-2}V^{\top}
\big)
\Big\|_F \\[4pt]
&\quad+
\Big\|
V\Lambda^{-2}V^{\top}(XP+PX)
\big(
\widehat V \widehat\Lambda^{-2}\widehat V^{\top}
-
V\Lambda^{-2}V^{\top}
\big)
\Big\|_F \\[4pt]
&\quad+
\Big\|
\sum_{j\ge2} V\Lambda^{-2}V^{\top} S_j(X)
+
\sum_{j\ge2} S_j(X)\widehat V \widehat\Lambda^{-2}\widehat V^{\top}
-
V\Lambda^{-2}V^{\top} X^2 \widehat V \widehat\Lambda^{-2}\widehat V^{\top}
\Big\|_F .
\end{align}\tag{140}\] By 1 3, together with bound on
\(\|X\|_2\) from 1 and the fact that \(\|S_1(X)\|_2 \lesssim \frac{1}{\sqrt{n \rho_n}}\), the first term
on the right-hand side of 140 satisfies \[\begin{align}
\label{smllbd1} \Big\|
V\Lambda^{-2}V^{\top} S_1(X)
+
S_1(X)V\Lambda^{-2}V^{\top}
-
V\Lambda^{-2}V^{\top}(XP+PX)V\Lambda^{-2}V^{\top}
\Big\|_F
= O_p \left(
\frac{\sqrt k}{n^{5/2}\rho_n^{5/2}}\right).
\end{align}\tag{141}\] The second and third terms can be bounded using submultiplicativity of the Frobenius norm and the operator-norm control of \(S_1(X)\) and \(XP+PX\),
yielding \[\begin{align}
\label{smllbd2} \Big\|
S_1(X)
\big(
\widehat V \widehat\Lambda^{-2}\widehat V^{\top}
-
V\Lambda^{-2}V^{\top}
\big)
\Big\|_F
+
\Big\|
V\Lambda^{-2}V^{\top}(XP+PX)
\big(
\widehat V \widehat\Lambda^{-2}\widehat V^{\top}
-
V\Lambda^{-2}V^{\top}
\big)
\Big\|_F\\ \notag
\lesssim
\frac{1}{\sqrt{n\rho_n}}
\bigg\|
\widehat V \widehat\Lambda^{-2}\widehat V^{\top}
-
V\Lambda^{-2}V^{\top}
\bigg\|_F .
\end{align}\tag{142}\] Finally, for the last term in 140 , 1 implies \[\begin{align}
\label{smllbd3} \Big\|
\sum_{j\ge2} V\Lambda^{-2}V^{\top} S_j(X)
\Big\|_F
= O_p \left(
\frac{k}{n^{3}\rho_n^{3}}\right),
\end{align}\tag{143}\] and the same bound holds for \(\big\|
\sum_{j\ge2} S_j(X)\widehat V \widehat\Lambda^{-2}\widehat V^{\top}
\big\|_F\) and \(\big\|
V\Lambda^{-2}V^{\top} X^2 \widehat V \widehat\Lambda^{-2}\widehat V^{\top}
\big\|_F\). Collecting 141 142 143 completes the proof. ◻
Proof of 17. Just like in 13.4, we will assume that \(A\) (in the one-sample case) and \({A^{(i)}}\) (in the two-sample case) both belong to \(\mathcal{E}_{{\sf very \;good}}\), as defined in 12, which happens with high probability as already proved.
Proof of 35 . We prove the claim directly. Using the triangle inequality we obtain \[\begin{align}
\|{\widehat P}^2 - P^2\|_F
&\le \|{\widehat P} \|_F \|{\widehat P}-P\|_2 + \|{\widehat P}-P\|_2 \| P \|_F\\
&\lesssim k^{1/2}\,n\rho_n\,\|X\|_2\\
&= O_p \left(
k^{1/2}\,n^{3/2}\rho_n^{3/2}\right),
\end{align}\] where the last inequality follows from 13 1.
Proof of 36 . We begin by decomposing the difference as \[{\widehat P}({\widehat P}\circ{\widehat P})-P(P\circ P)
=
({\widehat P}-P)({\widehat P}\circ{\widehat P})
+
P\bigl[({\widehat P}\circ{\widehat P})-(P\circ P)\bigr].\] Note that \(\|{\widehat P}\circ{\widehat P}\|_F^2 = \sum_{i,j} \widehat P_{ij}^4 = O_p \left(n^2 \rho_n^4 \right)\). Thus, for the first term on the right
hand side, \[\begin{align} \|({\widehat P}-P)({\widehat P}\circ{\widehat P})\|_F &\le k^{1/2}\|{\widehat P}-P\|_2\,\|{\widehat P}\circ{\widehat P}\|_F \\ &=O_p \left( k^{1/2}\,n^{3/2}\rho_n^{5/2}\right),
\end{align}\] where the bound follows from 1 13 1.
For the second term, observe that \[\begin{align} \label{smllbd4} \|({\widehat P}\circ{\widehat P})-(P\circ P)\|_F &= \|({\widehat P}-P)\circ({\widehat P}+P)\|_F \\ &\lesssim k^{1/2}
\|{\widehat P}-P\|_2 \|{\widehat P}+P\|_{\infty} \\ &=O_p \left( k^{1/2}\,n^{1/2}\rho_n^{3/2}\right).
\end{align}\tag{144}\] Consequently, \(\|P\bigl[({\widehat P}\circ{\widehat P})-(P\circ P)\bigr]\|_F
=O_p\left(
k^{1/2}\,n^{3/2}\rho_n^{5/2}\right).\) Combining the above bounds yields the claim of 36 .
Proof of 37 . We again use the inequality \[\begin{align} \label{smllbd5} \|({\widehat P}\circ{\widehat P})^2-(P\circ P)^2\|_F
\le
\|({\widehat P}\circ{\widehat P})-(P\circ P)\|_F\,
\|({\widehat P}\circ{\widehat P})+(P\circ P)\|_F.
\end{align}\tag{145}\] The first factor is bounded by \(k^{1/2}\,n^{1/2}\rho_n^{3/2}\) as shown in the last part, while 1 3 imply \(\|({\widehat P}\circ{\widehat P})+(P\circ P)\|_F = O_p \left( n\rho_n^2
\right)\) because of \(\|{\widehat P}\circ{\widehat P}\|_F^2 = \sum_{i,j} \widehat P_{ij}^4 = O_p \left(n^2 \rho_n^4 \right)\) as already shown in the last part. Combining this bound with 144 145 , we obtain \[\|({\widehat P}\circ{\widehat P})^2-(P\circ P)^2\|_F
=O_p \left(
k^{1/2}\,n^{3/2}\rho_n^{7/2}\right).\] ◻
Proof. Throughout this proof we assume that each \({A^{(i)}}\) lies in the set \(\mathcal{E}_{{\sf very \;good}}\) of 12, which holds with high probability. From Theorem 1 of [51], we have \((\widehat V^{(i)}) (\widehat V^{(i)})^{\top}- {V^{(i)}} {V^{(i)}}^{\top} = \sum_{l \ge 1} S_{i,l}({X^{(i)}}),\) where \(S_{i,l}({X^{(i)}})\) is defined exactly as in 26 . Consequently, the second and third order terms of \({X^{(i)}}\) in \(\mathbb{E}\big\langle \widehat{V^{(1)}}\widehat{V^{(1)}}^{\top}- {V^{(1)}}{V^{(1)}}^{\top}, \widehat{V^{(2)}}\widehat{V^{(2)}}^{\top}- {V^{(2)}}{V^{(2)}}^{\top}\big\rangle\) are \[\mathbb{E}\left( \mathop{\mathrm{tr}}\big(S_{1,1}({X^{(1)}}) S_{2,1}({X^{(2)}})\big) \right) \quad\text{and}\quad \mathbb{E}\left( \mathop{\mathrm{tr}}\big(S_{1,2}({X^{(1)}}) S_{2,1}({X^{(2)}}) + S_{1,1}({X^{(1)}}) S_{2,2}({X^{(2)}})\big) \right),\] both of which vanish immediately, as shown in the proof of 10. Moreover, again from that proof, we have the bound \[\begin{align} \label{smllbd6} \big\langle S_{1,l_1}({X^{(1)}}),\, S_{2,l_2}({X^{(2)}}) \big\rangle \le \|S_{1,l_1}({X^{(1)}})\|_F \, \|S_{2,l_2}({X^{(2)}})\|_F =O_p \left( \frac{k^2}{(n\rho_n)^{\,l/2}}\right), \end{align}\tag{146}\] where \(l = l_1 + l_2\). Next, by 8 and 3, it holds that \(\widehat\sigma_2^2 \asymp k^2/(n^3 \rho_n^2).\) We therefore obtain \[\begin{align}\label{note05} \frac{1}{\widehat\sigma_2} \big\langle {V^{(1)}}{V^{(1)}}^{\top}- {V^{(2)}}{V^{(2)}}^{\top}, \widehat{V^{(2)}}\widehat{V^{(2)}}^{\top}- {V^{(2)}}{V^{(2)}}^{\top} \big\rangle &= \frac{1}{\widehat\sigma_2} \displaystyle\sum_{l=2,3}\;\sum_{l_1 + l_2 = l} \big\langle S_{1,l_1}({X^{(1)}}),\, S_{2,l_2}({X^{(2)}}) \big\rangle \\ &\qquad + \frac{1}{\widehat\sigma_2} \displaystyle\sum_{l \ge 4}\;\sum_{l_1 + l_2 = l} \big\langle S_{1,l_1}({X^{(1)}}),\, S_{2,l_2}({X^{(2)}}) \big\rangle. \end{align}\tag{147}\] We will show that the second term in 147 is \(o_p(1)\) under 5. This is because from 146 8, we have the following bound with probability \(1-O(n^{-c})\) for a constant \(c>1\), \[\begin{align} \frac{1}{\widehat\sigma_2} \displaystyle\sum_{l \ge 4}\;\sum_{l_1 + l_2 = l} \big\langle S_{1,l_1}({X^{(1)}}),\, S_{2,l_2}({X^{(2)}}) \big\rangle &\lesssim \frac{n^{3/2} \rho_n}{k}\sum_{l \ge 4}\;2^l \frac{k^2}{(n\rho_n)^{\,l/2}}\\ &\lesssim \frac{k}{n^{1/2} \rho_n}=o_p(1). \end{align}\] For the first term on the right hand side of 147 , we invoke Chebyshev’s inequality. As argued earlier, \[\mathbb{E}\left( \sum_{l=2,3}\;\sum_{l_1+l_2=l} \big\langle S_{1,l_1}({X^{(1)}}),\, S_{2,l_2}({X^{(2)}}) \big\rangle \right)=0.\] We therefore compute the order of \(\mathop{\mathrm{Var}}\!\left( \sum_{l=2,3}\;\sum_{l_1+l_2=l} \big\langle S_{1,l_1}({X^{(1)}}),\, S_{2,l_2}({X^{(2)}}) \big\rangle \right).\) The first term in the above sum is \[\big\langle S_{1,1}({X^{(1)}}),\, S_{2,1}({X^{(2)}}) \big\rangle= \mathop{\mathrm{tr}}\!\left( \beta_1^{-1} {X^{(1)}} \beta^{\perp} {X^{(2)}} \beta_2^{-1} + \beta^{\perp} {X^{(1)}} \beta_1^{-1}\beta_2^{-1} {X^{(2)}} \right) = 2\,\mathop{\mathrm{tr}}\!\left(\beta_1^{-1}\beta_2^{-1} {X^{(1)}} \beta^{\perp} {X^{(2)}}\right).\] Hence it suffices to determine the order of \(4\,\mathbb{E}\!\left[ \bigl(\mathop{\mathrm{tr}}(\beta_1^{-1}\beta_2^{-1} {X^{(1)}} \beta^{\perp} {X^{(2)}})\bigr)^2 \right]\) since it is mean-zero.
Conditioning on \({X^{(1)}}\) and writing \(M_1 := \beta_1^{-1}\beta_2^{-1} {X^{(1)}} \beta^{\perp}\), we obtain \[\begin{align}\label{eq:e1a}
\mathbb{E}_{{X^{(2)}}}\!\left[
\bigl(\mathop{\mathrm{tr}}(M_1 {X^{(2)}})\bigr)^2
\right] &= \sum_{i < j} \mathbb{E}\left(X^{(2)}_{ij}\right)^2 \bigl( (M_1)_{ij} + (M_1)_{ji} \bigr)^2 + \sum_{i=1}^n \mathbb{E}\left(X^{(2)}_{ii}\right)^2 (M_1)^2_{ii}\\
&= \sum_{i,j} \sigma_{ij}^2 (M_1)_{ij}^2 + \sum_{i \neq j} \sigma_{ij}^2 (M_1)_{ij}(M_1)_{ji}.
\end{align}\tag{148}\] Taking expectation with respect to \({X^{(1)}}\), we will bound the first term on the right hand side of 148 : \[\begin{align}\label{eq:e1b}
\sum_{i,j} \sigma_{ij}^2 \mathbb{E}_{X^{(1)}}[(M_1)_{ij}^2] &\asymp \rho_n \sum_{i,j} \mathbb{E}_{X^{(1)}}[(M_1)_{ij}^2]\\
&= \rho_n \mathbb{E}_{X^{(1)}}\bigl[\|M_1\|_F^2\bigr] = \rho_n \mathbb{E}_{X^{(1)}}\bigl[\text{tr}(M_1 M_1^\top)\bigr].
\end{align}\tag{149}\] For the second term of 148 , we apply the Cauchy-Schwarz inequality to the sum: \[\begin{align}\label{eq:e1c}
\left| \sum_{i \neq j} \sigma_{ij}^2 \mathbb{E}_{X^{(1)}}[(M_1)_{ij}(M_1)_{ji}] \right| &\le \sum_{i \neq j} \sigma_{ij}^2 \mathbb{E}_{X^{(1)}}\bigl[|(M_1)_{ij}| |(M_1)_{ji}|\bigr]\\
&\le \frac{1}{2} \sum_{i \neq j} \sigma_{ij}^2 \mathbb{E}_{X^{(1)}}\bigl[ (M_1)_{ij}^2 + (M_1)_{ji}^2 \bigr] \lesssim \rho_n \mathbb{E}_{X^{(1)}}\bigl[\text{tr}(M_1 M_1^\top)\bigr].
\end{align}\tag{150}\] Thus, 149 150 imply that to find the order of \(\mathbb{E}_{{X^{(2)}}}\!\left[
\bigl(\mathop{\mathrm{tr}}(M_1 {X^{(2)}})\bigr)^2
\right]\) as in 148 , it suffices to evaluate that of \(\rho_n \mathbb{E}_{X^{(1)}}\bigl[\text{tr}(M_1 M_1^\top)\bigr]\). To evaluate the expectation rigorously, note that both \(\beta_2^{-1}\beta_1^{-2}\beta_2^{-1}\) and \(\beta^\perp\) are deterministic, symmetric, and positive semi-definite. Thus, \[\begin{align}
\mathbb{E}_{X^{(1)}}\bigl[\text{tr}(M_1 M_1^\top)\bigr] &= \mathbb{E}_{X^{(1)}} \bigl[ \mathop{\mathrm{tr}}\bigl( \beta_2^{-1}\beta_1^{-2}\beta_2^{-1} X^{(1)} \beta^\perp X^{(1)} \bigr)\bigr] \\ &= \sum_{i,j=1}^n \sigma_{ij}^2 \bigl(
(\beta_2^{-1}\beta_1^{-2}\beta_2^{-1})_{ii} (\beta^\perp)_{jj} + (\beta_2^{-1}\beta_1^{-2}\beta_2^{-1})_{ij} (\beta^\perp)_{ij} \bigr) \\ &\asymp \rho_n \sum_{i,j=1}^n \bigl( (\beta_2^{-1}\beta_1^{-2}\beta_2^{-1})_{ii} (\beta^\perp)_{jj} +
(\beta_2^{-1}\beta_1^{-2}\beta_2^{-1})_{ij} (\beta^\perp)_{ij} \bigr).
\end{align}\] Using 1 3 4 we obain: \[\mathbb{E}_{X^{(1)}}\bigl[\operatorname{tr}(M_1 M_1^\top)\bigr] \asymp \rho_n \cdot \frac{k}{(n\rho_n)^4} \cdot n = \frac{k}{n^3 \rho_n^3}.\]
Consequently, \[\begin{align} \label{need1} \mathop{\mathrm{Var}}\!\left(
\big\langle S_{1,1}({X^{(1)}}),\, S_{2,1}({X^{(2)}}) \big\rangle
\right) &= 4\,\mathbb{E}\!\left[
\bigl(\mathop{\mathrm{tr}}(\beta_1^{-1}\beta_2^{-1} {X^{(1)}} \beta^{\perp} {X^{(2)}})\bigr)^2
\right] \\ \notag
& \asymp \rho_n \mathbb{E}_{X^{(1)}}\bigl[\text{tr}(M_1 M_1^\top)\bigr] \\ \notag
&\asymp
\frac{k}{n^{3}\rho_n^{2}}.
\end{align}\tag{151}\] We now proceed to bound \(\operatorname{Var}\big( \langle S_{1,2}({X^{(1)}}),\, S_{2,1}({X^{(2)}}) \rangle \big)\) using analogous arguments. We expand the inner product as
\[\label{eq:e2a}
\begin{align}
\langle S_{1,2}({X^{(1)}}),\, S_{2,1}({X^{(2)}}) \rangle &= \operatorname{tr}\left( \beta^\perp {X^{(1)}} \beta^\perp {X^{(1)}} \beta_1^{-2} \beta_2^{-1} {X^{(2)}} \right) \\ &\quad - \operatorname{tr}\left( \beta^\perp {X^{(1)}} \beta_1^{-1}
{X^{(1)}} \beta_1^{-1} \beta_2^{-1} {X^{(2)}} \right) \\ &\quad + \operatorname{tr}\left( \beta_1^{-2} {X^{(1)}} \beta^\perp {X^{(1)}} \beta^\perp {X^{(2)}} \beta_2^{-1} \right) \\ &\quad - \operatorname{tr}\left( \beta_1^{-1} {X^{(1)}}
\beta_1^{-1} {X^{(1)}} \beta^\perp {X^{(2)}} \beta_2^{-1} \right).
\end{align}\tag{152}\] To analyze the variance, it suffices to determine the order of the second moment of the first term on the right hand side of 152 because all other terms can be tackled in the same manner: \(\mathbb{E}\left[\operatorname{tr}\left( \beta^\perp {X^{(1)}} \beta^\perp {X^{(1)}} \beta_1^{-2} \beta_2^{-1} {X^{(2)}} \right)^2 \right].\) Conditioning on \({X^{(1)}}\) as in 148 and defining \(N_1 := \beta^\perp {X^{(1)}} \beta^\perp {X^{(1)}} \beta_1^{-2} \beta_2^{-1}\), we obtain \[\label{eq:e2b}
\begin{align}
\mathbb{E}_{{X^{(2)}}}\left[ \bigl(\operatorname{tr}(N_1 {X^{(2)}})\bigr)^2 \right] \asymp \sum_{i,j} \sigma_{ij}^2 (N_1)_{ij}^2 + \sum_{i \neq j} \sigma_{ij}^2 (N_1)_{ij}(N_1)_{ji}.
\end{align}\tag{153}\] Taking the expectation with respect to \({X^{(1)}}\) and following the logic of 151 , 153 implies
\[\label{eq:e2c}
\begin{align}
\mathbb{E}\left[\operatorname{tr}\left( \beta^\perp {X^{(1)}} \beta^\perp {X^{(1)}} \beta_1^{-2} \beta_2^{-1} {X^{(2)}} \right)^2 \right] &\asymp \rho_n\,\mathbb{E}_{{X^{(1)}}}\left[\operatorname{tr}(N_1 N_1^{\top})\right] \\ &\asymp
\rho_n\,\mathbb{E}_{{X^{(1)}}}\left[ \operatorname{tr}\bigl( \beta_1^{-2}\beta_2^{-2}\beta_1^{-2} {X^{(1)}} (\beta^{\perp} {X^{(1)}})^3 \bigr) \right] \\ &\asymp \frac{k}{n^4\rho_n^3}.
\end{align}\tag{154}\] To derive the asymptotic order of \(\mathbb{E}_{{X^{(1)}}}\left[ \operatorname{tr}\bigl( \beta_1^{-2}\beta_2^{-2}\beta_1^{-2} {X^{(1)}} (\beta^{\perp} {X^{(1)}})^3 \bigr) \right]\) in 154 , we expand the trace term: \[\sum_{i_1,\dots,i_8} \mathbb{E} \left( (\beta_1^{-2}\beta_2^{-2}\beta_1^{-2})_{i_1,i_2} ({X^{(1)}})_{i_2,i_3} (\beta^\perp)_{i_3,i_4} ({X^{(1)}})_{i_4,i_5}
(\beta^\perp)_{i_5,i_6} ({X^{(1)}})_{i_6,i_7} (\beta^\perp)_{i_7,i_8} ({X^{(1)}})_{i_8,i_1} \right).\] We identify the index pairings for which this expectation is non-zero. One case is \(\{i_2, i_3\} = \{i_4,
i_5\}\) and \(\{i_6, i_7\} = \{i_8, i_1\}\). The contribution from this case is given by \[\label{spl}
\begin{align} \sum_{\substack{i_1,\dots,i_8 \\ \{i_2, i_3\} \sim \{i_4, i_5\} \\ \{i_6, i_7\} \sim \{i_8, i_1\}}} \mathbb{E} &\Bigg[ (\beta_1^{-2}\beta_2^{-2}\beta_1^{-2})_{i_1,i_2} ({X^{(1)}})_{i_2,i_3} (\beta^\perp)_{i_3,i_4}
({X^{(1)}})_{i_4,i_5}(\beta^\perp)_{i_5,i_6} ({X^{(1)}})_{i_6,i_7} (\beta^\perp)_{i_7,i_8} ({X^{(1)}})_{i_8,i_1} \Bigg] \\ &= \sum_{\substack{i_1,\dots,i_8 \\ \{i_2, i_3\} \sim \{i_4, i_5\} \\ \{i_6, i_7\} \sim \{i_8, i_1\}}} \Bigg[
(\beta_1^{-2}\beta_2^{-2}\beta_1^{-2})_{i_1,i_2} \underbrace{\mathbb{E} \left( ({X^{(1)}})_{i_2,i_3}^2 (\beta^\perp)_{i_3,i_2} + ({X^{(1)}})_{i_2,i_3}^2 (\beta^\perp)_{i_3,i_3} \right)}_{\text{Block 1}} \\ &\qquad\qquad\qquad\qquad \times
(\beta^\perp)_{i_5,i_6} \underbrace{\mathbb{E} \left( ({X^{(1)}})_{i_6,i_7}^2 (\beta^\perp)_{i_7,i_6} + ({X^{(1)}})_{i_6,i_7}^2 (\beta^\perp)_{i_7,i_7} \right)}_{\text{Block 2}} \Bigg] ,
\end{align}\tag{155}\] where the final estimate follows from Assumptions 1, 3, and 4.
To evaluate the order of 155 , we first compute the expectations within Block 1 and Block 2. By the independence of the entries of \(X^{(1)}\) and noting that \(\mathbb{E}[({X^{(1)}})_{ij}^2] = P_{ij}(1-P_{ij}) \asymp \rho_n\), summing over the inner index \(i_3\) in Block 1 gives: \[\begin{align}
\label{spl1}
\sum_{i_3}\mathbb{E} \left( ({X^{(1)}})_{i_2,i_3}^2 (\beta^\perp)_{i_3,i_2} + ({X^{(1)}})_{i_2,i_3}^2 (\beta^\perp)_{i_3,i_3} \right) &=
\sum_{i_3} P_{i_2 i_3}(1-P_{i_2 i_3}) \bigl( (\beta^\perp)_{i_3,i_2} + (\beta^\perp)_{i_3,i_3} \bigr) \\ \notag
& \asymp n\rho_n,
\end{align}\tag{156}\] by 4 1. By symmetry, we have
\[\begin{align}
\label{spl2}
\sum_{i_7}\mathbb{E} \left( ({X^{(1)}})_{i_6,i_7}^2 (\beta^\perp)_{i_7,i_6} + ({X^{(1)}})_{i_6,i_7}^2 (\beta^\perp)_{i_7,i_7} \right)
& \asymp n\rho_n.
\end{align}\tag{157}\] Using 1 3 4, we also have \[\begin{align}
\label{spl3} (\beta_1^{-2}\beta_2^{-2}\beta_1^{-2})_{i_1,i_2} \asymp \frac{k}{n(n \rho_n)^6}.
\end{align}\tag{158}\] Combining 155 156 157 158 , we obtain: \[\begin{align} &\sum_{\substack{i_1,\dots,i_8 \\ \{i_2,
i_3\} \sim \{i_4, i_5\} \\ \{i_6, i_7\} \sim \{i_8, i_1\}}} \mathbb{E} \Bigg[ (\beta_1^{-2}\beta_2^{-2}\beta_1^{-2})_{i_1,i_2} ({X^{(1)}})_{i_2,i_3} (\beta^\perp)_{i_3,i_4} ({X^{(1)}})_{i_4,i_5} \\ &\qquad\qquad\qquad\quad \times
(\beta^\perp)_{i_5,i_6} ({X^{(1)}})_{i_6,i_7} (\beta^\perp)_{i_7,i_8} ({X^{(1)}})_{i_8,i_1} \Bigg] \\ &\quad \asymp n^2 \rho_n^2 \cdot \frac{k}{(n\rho_n)^6} = \frac{k}{n^4 \rho_n^4}.
\end{align}\]
The other non-zero pairings, \(\{i_2, i_3\} = \{i_6, i_7\}\), \(\{i_4, i_5\} = \{i_8, i_1\}\) and \(\{i_2, i_3\} = \{i_8, i_1\}\), \(\{i_4, i_5\} = \{i_6, i_7\}\), are handled identically to 155 . Thus, multiplying the result of 155 by the factor \(\rho_n\) from 153 , the variance of the first term in 152 is of order \(\frac{k}{n^{4}\rho_n^{3}}\).
The remaining terms in 152 are bounded analogously. Consequently, \[\label{need2} \operatorname{Var}\big( \langle S_{1,2}({X^{(1)}}),\, S_{2,1}({X^{(2)}}) \rangle \big) \asymp \frac{k}{n^{4}\rho_n^{3}}.\tag{159}\] Similarly, we have \[\label{need3} \operatorname{Var}\big( \langle S_{1,1}({X^{(1)}}),\, S_{2,2}({X^{(2)}}) \rangle \big) \asymp \frac{k}{n^{4}\rho_n^{3}}.\tag{160}\] Using 151 159 160 and applying Cauchy–Schwarz inequality to the covariances, we bound the variance of the total sum: \[\begin{align} \operatorname{Var}\left( \sum_{l=2,3}\;\sum_{l_1+l_2=l} \langle S_{1,l_1}({X^{(1)}}),\, S_{2,l_2}({X^{(2)}}) \rangle \right) &\le \sum_{l=2,3}\;\sum_{l_1+l_2=l} \operatorname{Var}\left( \langle S_{1,l_1}({X^{(1)}}),\, S_{2,l_2}({X^{(2)}}) \rangle \right) \\ & \quad + \mathop{\mathrm{Cov}}\left(\langle S_{1,1}({X^{(1)}}),\, S_{2,2}({X^{(2)}}) \rangle, \langle S_{1,2}({X^{(1)}}),\, S_{2,1}({X^{(2)}}) \rangle \right) \\ & \quad + \mathop{\mathrm{Cov}}\left(\langle S_{1,1}({X^{(1)}}),\, S_{2,2}({X^{(2)}}) \rangle, \langle S_{1,1}({X^{(1)}}),\, S_{2,1}({X^{(2)}}) \rangle \right) \\ & \quad + \mathop{\mathrm{Cov}}\left(\langle S_{1,1}({X^{(1)}}),\, S_{2,1}({X^{(2)}}) \rangle, \langle S_{1,1}({X^{(1)}}),\, S_{2,2}({X^{(2)}}) \rangle \right) \\ &\le \sum_{l=2,3}\;\sum_{l_1+l_2=l} \operatorname{Var}\left( \langle S_{1,l_1}({X^{(1)}}),\, S_{2,l_2}({X^{(2)}}) \rangle \right) \\ & \quad + \sqrt {\mathop{\mathrm{Var}}\left(\langle S_{1,1}({X^{(1)}}),\, S_{2,2}({X^{(2)}}) \rangle\right) \mathop{\mathrm{Var}}\left(\langle S_{1,2}({X^{(1)}}),\, S_{2,1}({X^{(2)}}) \rangle \right)} \\ & \quad + \sqrt {\mathop{\mathrm{Var}}\left(\langle S_{1,1}({X^{(1)}}),\, S_{2,2}({X^{(2)}}) \rangle\right) \mathop{\mathrm{Var}}\left(\langle S_{1,1}({X^{(1)}}),\, S_{2,1}({X^{(2)}}) \rangle \right)} \\ & \quad + \sqrt {\mathop{\mathrm{Var}}\left(\langle S_{1,1}({X^{(1)}}),\, S_{2,1}({X^{(2)}}) \rangle\right) \mathop{\mathrm{Var}}\left(\langle S_{1,1}({X^{(1)}}),\, S_{2,2}({X^{(2)}}) \rangle \right)}\\ &\le \frac{k}{n^3\rho_n^2}. \end{align}\] Therefore, by Chebyshev’s inequality, we conclude \[\frac{\sum_{l=2,3}\;\sum_{l_1 + l_2 = l} \langle S_{1,l_1}({X^{(1)}}),\, S_{2,l_2}({X^{(2)}}) \rangle}{\widehat\sigma_2} \asymp 1,\] which completes the proof. ◻
Proof. Assume that \({A^{(i)}}\) belongs to \(\mathcal{E}_{{\sf very \;good}}\), as defined in 12, which happens with high probability as already proved. Using the expansion in 26 , we have \[\begin{align} L(\Delta) &= \langle (\widehat V^{(1)}) (\widehat V^{(1)})^{\top}- {V^{(1)}} {V^{(1)}}^{\top},\, \Delta \rangle - \langle (\widehat V^{(2)}) (\widehat V^{(2)})^{\top}- {V^{(2)}} {V^{(2)}}^{\top},\, \Delta \rangle \\ &= \sum_{k \geq 1} \langle S_{1,k}({X^{(1)}}) - S_{2,k}({X^{(2)}}),\, \Delta \rangle. \end{align}\] We first determine the order of \(\langle S_{1,1}({X^{(1)}}),\,\Delta\rangle\). Consider the scalar quantity \(T := \mathop{\mathrm{tr}}\big(\Delta\,\beta_1^{-1}{X^{(1)}}\beta_1^\perp\big).\) By cyclicity of the trace, \(T = \mathop{\mathrm{tr}}\big(\beta_1^\perp\Delta\beta_1^{-1}{X^{(1)}}\big) = \langle W,\,{X^{(1)}}\rangle,\) where \(W := \beta_1^\perp\Delta\beta_1^{-1}.\) Expanding over unordered indices gives \(T = \sum_{p} W_{pp}X_{pp} + \sum_{p<q} (W_{pq}+W_{qp})X_{pq}\) with \(a_{pp}=W_{pp}, a_{pq}=W_{pq}+W_{qp}(p<q).\) Applying Bernstein’s inequality to the independent, mean–zero summands \(a_{pq}X_{pq}\) yields \[\mathbb{P}(|T|\ge u) \le 2\exp\Big( -\frac{u^2}{2\sigma^2 + \tfrac{2}{3}Mu} \Big),\] where \(\sigma^2 = \sum_{p\le q}\mathop{\mathrm{Var}}(X_{pq})a_{pq}^2,\) and \(M = \max_{p\le q}|a_{pq}|\). Since \(\mathop{\mathrm{Var}}(X_{pq})\lesssim \rho_n\) and \(\sum_{p\le q}a_{pq}^2\asymp \|W\|_F^2\), we have \(\sigma^2 \lesssim \rho_n \|W\|_F^2.\) By submultiplicativity of norms, \(\|W\|_F = \|\beta_1^\perp \Delta \beta_1^{-1}\|_F \lesssim \frac{\|\Delta\|_F}{n\rho_n},\) which gives \(\sigma^2 \lesssim \rho_n \Big(\frac{\|\Delta\|_F}{n\rho_n}\Big)^2 = \frac{\|\Delta\|_F^2}{n^2\rho_n},\) and \(M \le \|W\|_F \lesssim \frac{\|\Delta\|_F}{n\rho_n}.\) Choosing \(u = C\sigma \sqrt {\log n}\) with a sufficiently large constant \(C\), we obtain \(T = O_p(\sigma\sqrt{\log n} )\). Hence, \[\begin{align} \label{start01} &\langle S_{1,1}({X^{(1)}}),\,\Delta\rangle = O_p\left( \frac{\|\Delta\|_F \sqrt{\log n}}{n \sqrt \rho_n} \right), \end{align}\tag{161}\] which implies that \[\begin{align} &\langle S_{1,1}({X^{(1)}}) - S_{2,1}({X^{(2)}}),\, \Delta \rangle = O_p\left(\frac{\|\Delta\|_F \sqrt{\rho_n\log n}}{n \rho_n}\right). \end{align}\] For higher–order terms (\(l \ge 2\)), Theorem 1 of [51] implies \[\bigg|\langle S_{i,l}({X^{(i)}}),\, \Delta \rangle \bigg| \le \| S_{i,l}({X^{(i)}})\|_F \|\Delta\|_F \lesssim \left(\frac{\|X^{(i)}\|}{n \rho_n} \right)^l\|\Delta\|_F =O_p \left(\frac{\|\Delta\|_F}{(n\rho_n)^{l/2}}\right).\] Combining the sharp Bernstein bound for the first–order term from 161 with the rate for higher–order terms yields \(L(\Delta) = O_p\left(\max \left(\frac{\|\Delta\|_F}{n\rho_n},\frac{\|\Delta\|_F \sqrt{\rho_n\log n}}{n \rho_n}\right)\right).\) ◻
Proof. Because \(\mathop{\mathrm{col}}(P)=\mathop{\mathrm{col}}(Z)\) and \(\mathop{\mathrm{rank}}(P)=k\), the \(k\) nonzero population eigenvectors of \(P\) span the column space of \(Z\). Hence there exists a full-rank matrix \(W\in\mathbb{R}^{k\times k}\) such that \(V = ZW.\) By construction each row of \(Z\) is a standard basis vector: if node \(i\) belongs to community \(a\) then the \(i\)-th row of \(Z\) equals \(e_a^{\top}\). Therefore the rows of \(V\) are constant on communities, namely for any node \(i\) in community \(a\) we have \(V_{i,\cdot} = W_{a,\cdot}.\) Orthonormality of the columns of \(V\) gives \(I_k= V^{\top}V = W^{\top}Z^{\top}Z\, W =: W^{\top}N W.\) Define \(U := N^{1/2} W\). Then \(U^{\top}U = W^{\top}N W = I_k,\) so \(U\) is a \(k\times k\) orthonormal matrix. Consequently \(W = N^{-1/2} U.\) Examining a row \(a\) of \(W\) yields \(\|W_{a,\cdot}\| = \frac{1}{\sqrt{n_a}} \,\|U_{a,\cdot}\|.\) Since \(U\) is square and orthonormal, each row of \(U\) has Euclidean norm equal to \(1\). Hence for every community \(a\), \(\|W_{a,\cdot}\| = \frac{1}{\sqrt{n_a}}.\) Therefore for any node \(i\) in community \(a\), \(\|V_{i,\cdot}\| = \|W_{a,\cdot}\| = \frac{1}{\sqrt{n_a}}.\) Under the balanced-size assumption \(n_a \asymp \,n/k\) we obtain \(\|V\|_{2,\infty} \asymp \sqrt{\frac{k}{c\,n}}.\) Next, we check that \[V V^{\top}= Z W W^{\top}Z^{\top}= Z N^{-1/2} U U^{\top}{N^{-1/2}} Z^{\top}= Z N^{-1}Z^{\top},\] implying that \((VV^\top)_{ij} = e_{C(i)}^\top N^{-1} e_{C(j)}=\frac{1}{n_a} \asymp \frac{k}{n}\), which establishes the claimed incoherence result. ◻
Proof. The proof follows the same high-level steps as 20 for the SBM. Because \(\mathop{\mathrm{col}}(P)=\mathop{\mathrm{col}}(Z)\) and \(\mathop{\mathrm{rank}}(P)=k\), there exists a full–rank matrix \(W\in\mathbb{R}^{k\times k}\) such that \(V = ZW.\) Orthonormality of the columns of \(V\) yields the same relation \(I_k= W^{\top}N W.\) Setting \(U:=N^{1/2}W\) we have \(W = N^{-1/2} U\) implying \(V = Z N^{-1/2} U.\) Right multiplication by the orthonormal matrix \(U\) preserves row Euclidean norms. Therefore for each node \(i\), \(\|V_{i,\cdot}\| = \|\,Z_{i,\cdot} N^{-1/2} U\,\| = \|\,Z_{i,\cdot} N^{-1/2}\,\|.\) By the operator norm inequality, \(\|V_{i,\cdot}\| \le \|Z_{i,\cdot}\| \,\|N^{-1/2}\| = \frac{\|Z_{i,\cdot}\|}{\sqrt{\lambda_{\min}(N)}}.\) Since each row \(Z_{i,\cdot}\) is a probability vector we have \(\|Z_{i,\cdot}\|\le\|Z_{i,\cdot}\|_1=1\), and thus \(\|V_{i,\cdot}\| \le \frac{1}{\sqrt{\lambda_{\min}(N)}}\) for all \(i.\) Similarly, \(\|V_{i,\cdot}\| = \|Z_{i,.} N^{-1/2}\| \ge \frac{\|Z_{i,.}\|}{\sqrt{\lambda_{\max}(N)}} \ge \frac{1}{\sqrt{k\lambda_{\max}(N)}}\) for all \(i\) and \(V V^{\top}=Z N^{-1}Z^{\top}.\) Finally applying the assumed spectral bounds \(\lambda_{\min}(N)\ge c_1\,n/k\) and \(\lambda_{\max}(N)\le c_2\,n/k\) completes the proof. ◻
Proof. Since \(V^{(i)}\) contains the eigenvectors corresponding to the non-zero eigenvalues of \(P^{(i)}\), we have \(\mathop{\mathrm{col}}(V^{(i)}) = \mathop{\mathrm{col}}(P^{(i)})\). Substituting the structure of the GRDPG probability matrix, we have \[\mathop{\mathrm{col}}(P^{(i)}) = \mathop{\mathrm{col}}(\Gamma^{(i)} I_{p_i,q_i} {\Gamma^{(i)}}^{\top}).\] \(I_{p_i,q_i}\) is a diagonal matrix with entries \(\pm 1\) and hence, is non-singular. Furthermore, because \(\Gamma^{(i)}\) has full column rank \(k\), the product \(I_{p_i,q_i} {\Gamma^{(i)}}^{\top}\) has full row rank \(k\). Right-multiplication by a full row-rank matrix preserves the column space of the left matrix. Therefore, \(\mathop{\mathrm{col}}(P^{(i)}) = \mathop{\mathrm{col}}(\Gamma^{(i)}).\) Consequently, the condition \(V^{(1)} {V^{(1)}}^{\top}= V^{(2)} {V^{(2)}}^{\top}\) holds if and only if \(\mathop{\mathrm{col}}(\Gamma^{(1)}) = \mathop{\mathrm{col}}(\Gamma^{(2)})\). For two full-rank matrices \(\Gamma^{(1)}, \Gamma^{(2)} \in \mathbb{R}^{n \times k}\), their column spaces are identical if and only if there exists an invertible matrix \(W \in \mathbb{R}^{k \times k}\) such that \(\Gamma^{(1)} = \Gamma^{(2)} W\). ◻
Proof. First note that, for each \(i\), \(\mathop{\mathrm{col}}(P^{(i)})\subseteq\mathop{\mathrm{col}}({Z^{(i)}})\) because \(P^{(i)}={Z^{(i)}}B^{(i)}{Z^{(i)}}^{\top}\). Since \(B^{(i)}\) is full-rank and \(\mathop{\mathrm{rank}}(P^{(i)})=k\) by assumption, we must have \(\mathop{\mathrm{rank}}({Z^{(i)}})=k\) and therefore \(\mathop{\mathrm{col}}(P^{(i)})=\mathop{\mathrm{col}}({Z^{(i)}})\) for \(i=1,2.\) If \({V^{(1)}}{V^{(1)}}^{\top}={V^{(2)}}{V^{(2)}}^{\top}\), then the two projection (equivalently column) spaces coincide: \[\mathop{\mathrm{col}}({P^{(1)}})=\mathop{\mathrm{col}}({V^{(1)}})=\mathop{\mathrm{col}}({V^{(2)}})=\mathop{\mathrm{col}}({P^{(2)}}).\] Using the above equation, this is equivalent to \(\mathop{\mathrm{col}}(Z^{(1)})=\mathop{\mathrm{col}}(Z^{(2)})\). Conversely, if \(\mathop{\mathrm{col}}(Z^{(1)})=\mathop{\mathrm{col}}(Z^{(2)}),\) then \(\mathop{\mathrm{col}}({P^{(1)}})=\mathop{\mathrm{col}}({P^{(2)}})\) and hence \({V^{(1)}}{V^{(1)}}^{\top}={V^{(2)}}{V^{(2)}}^{\top}\). Because \(\mathop{\mathrm{col}}(Z^{(1)})=\mathop{\mathrm{col}}(Z^{(2)})\) and both \(Z^{(1)}\) and \(Z^{(2)}\) have full column rank, there exists an invertible \(k\times k\) matrix \(R\) such that \(Z^{(2)} = Z^{(1)} R.\)
For the SBM, each canonical basis vector \(e_j^{\top}\) appears as a row of \(Z^{(1)}\); applying \(R\) yields that the rows of \(Z^{(2)}\) are the vectors \(e_j^{\top}R\). But each \(e_j^{\top}R\) must itself be a canonical basis vector (since rows of \(Z^{(2)}\) are one-hot); hence \(R\) sends the standard basis to the standard basis and therefore \(R\) is a permutation matrix. Thus \(Z^{(2)} = Z^{(1)}\Pi\) for some permutation \(\Pi\), i.e., the membership matrices coincide up to relabeling of community indices.
For the MMSBM case, consider the pure rows of \(Z^{(1)}\). If \(i_a\) is an index such that \(Z^{(1)}_{i_a \cdot} = e_a^\top\), then evaluating \(Z^{(2)}=Z^{(1)}R\) at row \(i_a\) yields \[Z^{(2)}_{i_a \cdot} = e_a^\top R.\] This implies that the rows of \(R\) exactly coincide with the rows of \(Z^{(2)}\) corresponding to the pure nodes of \(Z^{(1)}\). Because the rows of \(Z^{(2)}\) lie in the \((k-1)\)-simplex, every row of \(R\) must also reside in the simplex; in particular, \(R\) is element-wise non-negative and its rows sum to one.
\(Z^{(2)}\) also contains pure nodes for each community. Since \(R\) is invertible, we can express the latent position relationship as \(Z^{(1)} = Z^{(2)} R^{-1}\). Evaluating this equation at the pure nodes of \(Z^{(2)}\) demonstrates that the rows of \(R^{-1}\) similarly correspond to rows of \(Z^{(1)}\), which implies that \(R^{-1}\) is also an element-wise non-negative matrix.
Theorem 5.1 of [63] establishes that a matrix and its inverse are both non-negative if and only if the matrix can be expressed as the product of a diagonal matrix with strictly positive diagonal entries and a permutation matrix. This result implies that \(R\) must be a generalized permutation matrix. Because we have already established that the rows of \(R\) must sum to one, its non-zero entries are forced to be exactly one. Hence, \(R\) is a standard permutation matrix, and consequently, \(Z^{(2)}=Z^{(1)}\Pi\) for some permutation \(\Pi\). ◻
Proof. Write \(B=\rho_n\widetilde{B}\) and \(N=Z^{\top}Z\). Then \(P = Z B Z^{\top}= \rho_n\, Z\widetilde{B} Z^{\top}.\) The nonzero eigenvalues of \(Z\widetilde{B} Z^{\top}\) coincide with the eigenvalues of the \(k\times k\) matrix \(\widetilde{B}^{1/2} N \widetilde{B}^{1/2},\) so that for each \(1\le j\le k\), \(\lambda_j(P) = \rho_n\cdot \lambda_j\big(\widetilde{B}^{1/2} N \widetilde{B}^{1/2}\big).\) Suppose \(v_j\) is the eigenvector of \(\widetilde{B}^{1/2} N \widetilde{B}^{1/2}\) corresponding to the eigenvalue \(\lambda_j\big(\widetilde{B}^{1/2} N \widetilde{B}^{1/2}\big).\) Since \(\widetilde{B}\) and \(N\) are symmetric positive definite, we have \[\begin{align} \lambda_j\big(\widetilde{B}^{1/2} N \widetilde{B}^{1/2}\big) &= v_j^{\top}\big(\widetilde{B}^{1/2} N \widetilde{B}^{1/2}\big)v_j. \\ \intertext{which yields} \lambda_{\min}(N) \|\widetilde{B}^{1/2} v_j\|^2 &\le \lambda_j\big(\widetilde{B}^{1/2} N \widetilde{B}^{1/2}\big) \le \lambda_{\max}(N) \|\widetilde{B}^{1/2} v_j\|^2, \\ \intertext{and further bounding \|\widetilde{B}^{1/2} v_j\|^2 gives} \lambda_{\min}(N) \lambda_{\min}(\widetilde{B}) &\le\lambda_j\big(\widetilde{B}^{1/2} N \widetilde{B}^{1/2}\big) \le \lambda_{\max}(N) \lambda_{\max}(\widetilde{B}). \end{align}\] Combining this result with 1, the eigenvalue bounds on \(B\), and the assumption of balanced community sizes (which imply there exist constants \(b_1,b_2>0\) such that \(b_1 n \lesssim \lambda_{\min}(N) \le \lambda_{\max}(N) \lesssim b_2 n\)), we obtain \[\begin{align} a_1 b_1 \, n &\le \lambda_j\big(\widetilde{B}^{1/2} N \widetilde{B}^{1/2}\big) \le a_2 b_2 \, n. \\ \intertext{This implies that} a_1 b_1 \, n\rho_n &\le \lambda_j(P) \le a_2 b_2 \, n\rho_n, \\ \intertext{which upon dividing by n \rho_n yields} a_1 b_1 &\le \frac{\lambda_j(P)}{n \rho_n} \le a_2 b_2, \end{align}\] completing the proof. ◻
Proof. The conclusion follows directly from 1 together with 8. The latter establishes that \(\widetilde{\sigma}_2^{2} \asymp k^{2}/(n^{3}\rho_{n}^{2})\). Moreover, under 5 we have \(\frac{k}{n^{3/2}\rho_{n}} \gg \frac{k}{n^{2}\rho_{n}^{2}},\) which, in combination with the preceding variance characterization, yields the desired result. ◻
We will break down the proof into several steps.
Assume that \(A\) belongs to \(\mathcal{E}_{{\sf very \;good}}\) as defined in 12. We slightly refine the proof of 6 where it is already shown that \(|\widehat{\mu}_1 - \mu_1| =O_p \left(\max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right)\right).\) We will now show that \(|\widehat{\mu}_1 - \mu_1| \asymp \max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right).\) We will decompose \(\mu_1^{(2)}\) and \(\widehat\mu_1^{(2)}\) (defined in 9.3.1) in a slightly different way for this. We have that \[\begin{align}\label{inite95dec01} \mu_1^{(2)} - \widehat{\mu}_1^{(2)} &= \mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\Sigma \cdot d) - V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\Sigma \cdot \mathbf{1}_n) \right) \\ &\quad + \mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\Sigma \cdot \mathbf{1}_n) - \widehat V \widehat\Lambda^{-2} \widehat V^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma} \cdot \mathbf{1}_n) \right) \\ &\quad + \mathop{\mathrm{tr}}\left( \widehat{V} \widehat{\Lambda}^{-2} \widehat{V}^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma} \cdot \mathbf{1}_n) - \widehat{V} \widehat{\Lambda}^{-2} \widehat{V}^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma} \cdot \widehat{d}) \right). \end{align}\tag{162}\] All the terms of 162 except the second one \(\mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(\Sigma \cdot \mathbf{1}_n) - \widehat V \widehat\Lambda^{-2} \widehat V^{\top}\mathop{\mathrm{Diag}}(\widehat{\Sigma} \cdot \mathbf{1}_n) \right)\) have already been shown to be \(o_p \left(\max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right) \right)\) in the proof of 6 in 9.3.1. Our target will be to prove that \[\mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\,\mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) - \widehat V \widehat\Lambda^{-2} \widehat V^{\top}\,\mathop{\mathrm{Diag}}({\widehat P} \cdot \mathbf{1}_n) \right) \asymp \max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right),\] because we can show that \[\mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\,\mathop{\mathrm{Diag}}((P \circ P) \cdot \mathbf{1}_n) - \widehat V \widehat\Lambda^{-2} \widehat V^{\top}\,\mathop{\mathrm{Diag}}(({\widehat P} \circ {\widehat P}) \cdot \mathbf{1}_n) \right) = o_p \left(\max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right) \right)\] by an argument entirely analogous to that used earlier in the proof of 14 in 13.3.
We have that \[\begin{align} \label{nseries1} \mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\,\mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) - \widehat V \widehat\Lambda^{-2} \widehat V^{\top}\,\mathop{\mathrm{Diag}}({\widehat P} \cdot \mathbf{1}_n) \right) &= \mathop{\mathrm{tr}}\left( \left( V \Lambda^{-2} V^{\top} - \widehat V \widehat\Lambda^{-2} \widehat V^{\top}\right)\,\mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) \right) \\ & \quad - \mathop{\mathrm{tr}}\left( \left(\widehat V \widehat\Lambda^{-2} \widehat V^{\top}- V \Lambda^{-2} V^{\top}\right)\,\mathop{\mathrm{Diag}}(( {\widehat P}-P) \cdot \mathbf{1}_n) \right)\\ & \quad - \mathop{\mathrm{tr}}\! \left( V \Lambda^{-2} V^{\top}\,\mathop{\mathrm{Diag}}(( {\widehat P}-P) \cdot \mathbf{1}_n) \right). \end{align}\tag{163}\]
We now expand the terms in 163 one by one using 13, 123 , and Theorem 1 of [51]. This yields the following decompositions: \[\begin{align} \label{nseries2} \mathop{\mathrm{tr}}\bigg(& \bigl(\widehat V \widehat \Lambda^{-2} \widehat V^{\top} - V \Lambda^{-2} V^{\top}\bigr)\,\mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) \bigg) \\ &= \mathop{\mathrm{tr}}\Big( \bigl[ V\Lambda^{-2}V^{\top}S_1(X) + S_1(X)V\Lambda^{-2}V^{\top} - V\Lambda^{-2}V^{\top}(XP+PX)V\Lambda^{-2}V^{\top} \bigr]\mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) \Big) \\[4pt] &\quad+ \mathop{\mathrm{tr}}\Big( S_1(X)\bigl(\widehat V\widehat\Lambda^{-2}\widehat V^{\top} - V\Lambda^{-2}V^{\top}\bigr) \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) \Big) \\[4pt] &\quad- \mathop{\mathrm{tr}}\Big( V\Lambda^{-2}V^{\top}(XP+PX) \bigl(\widehat V\widehat\Lambda^{-2}\widehat V^{\top} - V\Lambda^{-2}V^{\top}\bigr) \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) \Big) \\[4pt] &\quad+ \mathop{\mathrm{tr}}\Big( \bigl( \sum_{j\ge2}V\Lambda^{-2}V^{\top}S_j(X) + \sum_{j\ge2}S_j(X)\widehat V\widehat\Lambda^{-2}\widehat V^{\top}- V\Lambda^{-2}V^{\top}X^2\widehat V\widehat\Lambda^{-2}\widehat V^{\top} \bigr)\mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) \Big). \end{align}\tag{164}\] Similarly, \[\begin{align} \label{nseries3} \mathop{\mathrm{tr}}\bigg( & \bigl(\widehat V \widehat \Lambda^{-2} \widehat V^{\top} - V \Lambda^{-2} V^{\top}\bigr)\, \mathop{\mathrm{Diag}}\bigl(({\widehat P}-P)\cdot \mathbf{1}_n\bigr) \bigg) \\ &= \sum_{l \ge 1} \mathop{\mathrm{tr}}\Big( \bigl[ V\Lambda^{-2}V^{\top}S_1(X) + S_1(X)V\Lambda^{-2}V^{\top} - V\Lambda^{-2}V^{\top}(XP+PX)V\Lambda^{-2}V^{\top} \bigr]\mathop{\mathrm{Diag}}(T_l(X)\cdot \mathbf{1}_n) \Big) \\[4pt] &\quad+ \sum_{l \ge 1} \mathop{\mathrm{tr}}\Big( S_1(X)\bigl(\widehat V\widehat\Lambda^{-2}\widehat V^{\top} - V\Lambda^{-2}V^{\top}\bigr) \mathop{\mathrm{Diag}}(T_l(X)\cdot \mathbf{1}_n) \Big) \\[4pt] &\quad- \sum_{l \ge 1} \mathop{\mathrm{tr}}\Big( V\Lambda^{-2}V^{\top}(XP+PX) \bigl(\widehat V\widehat\Lambda^{-2}\widehat V^{\top} - V\Lambda^{-2}V^{\top}\bigr) \mathop{\mathrm{Diag}}(T_l(X)\cdot \mathbf{1}_n) \Big) \\[4pt] &\quad+ \sum_{l \ge 1} \mathop{\mathrm{tr}}\Big( \bigl( \sum_{j\ge2}V\Lambda^{-2}V^{\top}S_j(X) + \sum_{j\ge2}S_j(X)\widehat V\widehat\Lambda^{-2}\widehat V^{\top}- V\Lambda^{-2}V^{\top}X^2\widehat V\widehat\Lambda^{-2}\widehat V^{\top} \bigr)\mathop{\mathrm{Diag}}(T_l(X)\cdot \mathbf{1}_n) \Big). \end{align}\tag{165}\] Finally, \[\begin{align} \label{nseries4} \mathop{\mathrm{tr}}\left( V\Lambda^{-2}V^{\top}\, \mathop{\mathrm{Diag}}\bigl(({\widehat P}-P)\cdot \mathbf{1}_n\bigr) \right) = \sum_{l \ge 1} \mathop{\mathrm{tr}}\left( V\Lambda^{-2}V^{\top}\, \mathop{\mathrm{Diag}}(T_l(X)\cdot \mathbf{1}_n) \right). \end{align}\tag{166}\] Thus, we have shown that \[\begin{align} \label{note06} \mu_1^{(2)} - \widehat\mu_1^{(2)} &= \mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\,\mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) - \widehat V \widehat\Lambda^{-2} \widehat V^{\top}\,\mathop{\mathrm{Diag}}({\widehat P} \cdot \mathbf{1}_n) \right) + o_p \left(\max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right) \right). \end{align}\tag{167}\] As in the proofs of 30 28 , one can verify that all linear in \(X\) contributions of the first term in 167 (expanded term-wise in 163 164 165 166 ) are of order \(O_p\left(\frac{k \log n}{n^{2}\rho_n^{3/2}}\right).\) Next, we will find the quadratic contributions of the same trace term in 167 .
We now demonstrate that the combined contribution of the second-order terms in 163 is of order \(\max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right)\). Recall the term-by-term expansions of 163 provided in 164 165 166 . Focusing first on 164 and isolating all terms that depend quadratically on \(X\), we obtain: \[\begin{align} \label{nseries5} \text{Term 1} &= \mathop{\mathrm{tr}}\Big( S_1(X)^2\, V\Lambda^{-2}V^{\top} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) \Big) - \mathop{\mathrm{tr}}\Big( S_1(X)V\Lambda^{-2}V^\top (PX+XP)V\Lambda^{-2}V^\top \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) \Big)\\ &\quad+ \mathop{\mathrm{tr}}\Big( S_1(X) V\Lambda^{-2}V^{\top}S_1(X) \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) \Big) - \mathop{\mathrm{tr}}\Big( V\Lambda^{-2}V^{\top}(PX+XP) S_1(X) V\Lambda^{-2}V^{\top} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) \Big)\\ &\quad+ \mathop{\mathrm{tr}}\Big( V\Lambda^{-2}V^{\top}(PX+XP) V\Lambda^{-2}V^\top (PX+XP) V\Lambda^{-2}V^\top \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) \Big)\\ &\quad- \mathop{\mathrm{tr}}\Big( V\Lambda^{-2}V^{\top}(PX+XP) V\Lambda^{-2}V^{\top}S_1(X) \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) \Big)\\ &\quad+ \mathop{\mathrm{tr}}\Big( \bigl( V\Lambda^{-2}V^{\top}S_2(X) + S_2(X)V\Lambda^{-2}V^{\top} - V\Lambda^{-2}V^{\top}X^2V\Lambda^{-2}V^{\top} \bigr)\mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) \Big)\\ &= \mathop{\mathrm{tr}}\Bigg( \Big( \beta^{\perp} X \beta^{-2} X \beta^{-2} - \beta^{\perp} X \beta^{-3} X \beta^{-1} + \beta^{\perp} X \beta^{-4} X \beta^{\perp} - \beta^{-1} X \beta^{\perp} X \beta^{-3} + \beta^{-2} X \beta^{-1} X \beta^{-1} \\ &\qquad\quad - \beta^{-2} X \beta^{-2} X \beta^{\perp} - \beta^{-2} X \beta^{\perp} X \beta^{-2} + \beta^{-1} X \beta^{-1} X \beta^{-2} + \beta^{-1} X \beta^{-2} X \beta^{-1} - \beta^{-1} X \beta^{-3} X \beta^{\perp} \\ &\qquad\quad + \beta^{-4} X \beta^{\perp} X \beta^{\perp} - \beta^{-3} X \beta^{\perp} X \beta^{-1} - \beta^{-3} X \beta^{-1} X \beta^{\perp} + \beta^{\perp} X \beta^{\perp} X \beta^{-4} - \beta^{\perp} X \beta^{-1} X \beta^{-3} \Big) \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) \Bigg)\\ &= \text{Term 0} + \mathop{\mathrm{tr}}\left(\beta^{-4} X \beta^{\perp} X \beta^{\perp} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\right) + \mathop{\mathrm{tr}}\left(\beta^{\perp} X \beta^{\perp} X \beta^{-4} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\right) \\ & \qquad \qquad + \mathop{\mathrm{tr}}\left(\beta^{\perp} X \beta^{-4} X \beta^{\perp} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\right), \end{align}\tag{168}\] where we group the remaining components into \(\text{Term 0}\), defined as \[\begin{align} \text{Term 0} \mathrel{\vcenter{:}}= \mathop{\mathrm{tr}}\Bigg( \Big( &\beta^{\perp} X \beta^{-2} X \beta^{-2} - \beta^{\perp} X \beta^{-3} X \beta^{-1} - \beta^{-1} X \beta^{\perp} X \beta^{-3} + \beta^{-2} X \beta^{-1} X \beta^{-1} - \beta^{-2} X \beta^{-2} X \beta^{\perp} \\ &+ \beta^{-1} X \beta^{-1} X \beta^{-2} + \beta^{-1} X \beta^{-2} X \beta^{-1} - \beta^{-1} X \beta^{-3} X \beta^{\perp} - \beta^{-3} X \beta^{\perp} X \beta^{-1} \\ &- \beta^{-3} X \beta^{-1} X \beta^{\perp} - \beta^{\perp} X \beta^{-1} X \beta^{-3} - \beta^{-2} X \beta^{\perp} X \beta^{-2} \Big) \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) \Bigg). \end{align}\] We now proceed to bound all terms in 168 starting first with \[\mathop{\mathrm{tr}}\left(\beta^{-4} X \beta^{\perp} X \beta^{\perp} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\right) = \sum_{j,k,l,m} w^{(1)}_{jklm}\, X_{jk} X_{lm},\] where the coefficients are defined as \[\begin{align} w^{(1)}_{jklm} = \sum_{i,q} (\beta^{-4})_{ij} (\beta^\perp)_{kl} (\beta^\perp)_{mi} P_{iq}. \end{align}\] Under 1 3 4, these coefficients satisfy the bounds \[\bigl| w^{(1)}_{jklm} \bigr| \asymp \begin{cases} \dfrac{k^2}{n^4 \rho_n^3}, & k=l, \\[10pt] \dfrac{k^2}{n^5 \rho_n^3}, & k\neq l. \end{cases}\] Define \[\begin{align} \label{e461} E_1 = \mathbb{E}\left(\mathop{\mathrm{tr}}\left(\beta^{-4} X \beta^{\perp} X \beta^{\perp} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\right) \right). \end{align}\tag{169}\] Consequently, it is straightforward to verify that the expectation satisfies \[\begin{align} \big|E_1\big| \lesssim \frac{k^2}{n^2 \rho_n^2}. \end{align}\] Next, we determine the asymptotic order of the corresponding variance: \[\begin{align}\label{varref01} \mathop{\mathrm{Var}}\left(\sum_{j,k,l,m} w^{(2)}_{jklm} X_{jk} X_{lm}\right) &= \sum_{j,k,l,m} \left(w^{(2)}_{jklm} \right)^2 \mathop{\mathrm{Var}}\left(X_{jk} X_{lm} \right) + \sum_{(j,k,l,m) \neq \atop (j',k',l',m')} w^{(2)}_{jklm} w^{(2)}_{j' k' l' m'} \mathop{\mathrm{Cov}}\left(X_{jk} X_{lm},X_{j'k'} X_{l'm'} \right) \\ &= \sum_{j,k,l,m \atop k \neq l} \left(w^{(2)}_{jklm} \right)^2 \mathop{\mathrm{Var}}\left(X_{jk} X_{lm} \right) + \sum_{j,k,l,m \atop k = l} \left(w^{(2)}_{jklm} \right)^2 \mathop{\mathrm{Var}}\left(X_{jk} X_{lm} \right) \\ &\quad + \sum_{(j,k,l,m) \neq \atop (j',k',l',m')} w^{(2)}_{jklm} w^{(2)}_{j' k' l' m'} \mathop{\mathrm{Cov}}\left(X_{jk} X_{lm},X_{j'k'} X_{l'm'} \right) \\ &\lesssim \frac{k^4}{n^5 \rho_n^4}. \end{align}\tag{170}\] Therefore, an application of Chebyshev’s inequality yields that, for some constant \(C_1>0\), \[\begin{align} \label{nseries6464} \mathbb{P} \left(\left|\mathop{\mathrm{tr}}\left(\beta^{-4} X \beta^{\perp} X \beta^{\perp} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\right)- E_1 \right| \le \frac{k^2}{n^{5/2} \rho_n^2} \right) \ge C_1. \end{align}\tag{171}\] By analogous reasoning as in 171 , there exists a constant \(C_2>0\) such that \[\begin{align} \label{nseries6465} \mathbb{P} \left(\left|\mathop{\mathrm{tr}}\left(\beta^{\perp} X \beta^{\perp} X \beta^{-4} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\right)- E_2 \right| \le \frac{k^2}{n^{5/2} \rho_n^2} \right) \ge C_2, \end{align}\tag{172}\] where \[\begin{align} \label{e462} E_2 = \mathbb{E}\left(\mathop{\mathrm{tr}}\left(\beta^{\perp} X \beta^{\perp} X \beta^{-4} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\right)\right). \end{align}\tag{173}\] Next, we will bound the term \(\mathop{\mathrm{tr}}\left(\beta^{\perp} X \beta^{-4} X \beta^{\perp} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\right)\), where we we again do our routine decomposition: \[\mathop{\mathrm{tr}}\left(\beta^{\perp} X \beta^{-4} X \beta^{\perp} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\right)=\sum_{j,k,l,m} w^{(4)}_{jklm}\, X_{jk} X_{lm}.\] The coefficients are defined as \[\begin{align} w^{(3)}_{jklm} = \sum_{i,q} (\beta^\perp)_{ij} (\beta^{-4})_{kl} (\beta^\perp)_{mi} P_{iq}. \end{align}\] Under 1 3 4, \[\bigl| w^{(3)}_{jklm} \bigr| \asymp \begin{cases} \dfrac{k^2}{n^4 \rho_n^3}, & j=m, \\[10pt] \dfrac{k^2}{n^5 \rho_n^3}, & j\neq m. \end{cases}\] By analogous reasoning as in 171 , there exists a constant \(C_3>0\) such that \[\begin{align} \label{nseries6466} \mathbb{P} \left(\left|\mathop{\mathrm{tr}}\left(\beta^{\perp} X \beta^{-4} X \beta^{\perp} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\right)- E_3 \right| \lesssim \frac{k^2}{n^{5/2} \rho_n^2} \right) \ge C_3, \end{align}\tag{174}\] where \[\begin{align} \label{e463} E_3 = \mathbb{E}\left(\mathop{\mathrm{tr}}\left(\beta^{\perp} X \beta^{-4} X \beta^{\perp} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\right)\right). \end{align}\tag{175}\] Identical asymptotic bounds hold for Term 0 in 168 . Suppose we take \(\mathop{\mathrm{tr}}\left(\beta^\perp X \beta^{-3}X \beta^{-1} \mathop{\mathrm{Diag}}(P \mathbf{1}_n) \right)\) as a representative term. Expanding the term, we get \[\mathop{\mathrm{tr}}\left(\beta^\perp X \beta^{-3}X \beta^{-1} \mathop{\mathrm{Diag}}(P \mathbf{1}_n) \right) = \sum_{j,u,v,l} w^{(0)}_{juvl}X_{ju}X_{vl},\] where \(w^{(0)}_{juvl}=\sum_{i,s} (\beta^\perp)_{ij}(\beta^{-3})_{uv}(\beta^{-1})_{li}P_{is}\). We can check that under 1 3 4, \[\bigl| w^{(0)}_{juvl} \bigr| \asymp \frac{k^2}{n^5 \rho_n^3}.\] Thus, \[\left|\mathbb{E}\left[ \mathop{\mathrm{tr}}\left(\beta^\perp X \beta^{-3}X \beta^{-1} \mathop{\mathrm{Diag}}(P \mathbf{1}_n) \right) \right] \right| = O \left(\frac{k^2}{n^3 \rho_n^2}\right),\] and, proceeding similarly to 170 , we have \[\mathop{\mathrm{Var}}\left(\mathop{\mathrm{tr}}\left(\beta^\perp X \beta^{-3}X \beta^{-1} \mathop{\mathrm{Diag}}(P \mathbf{1}_n) \right) \right) = O \left(\max \left(\frac{k^2}{n^3 \rho_n^2}, \frac{k^2}{n^4 \rho_n^{5/2}}\right)\right).\] Applying Chebyshev’s inequality, we obtain \[\label{nseries6462} \mathbb{P} \left(\big| \text{Term 0} \big| \lesssim \frac{k^2}{n^3 \rho_n^2}\right) \ge C_4\tag{176}\] for some constant \(C_4>0\). Therefore, combining the decomposition of 168 with the concentration bounds established in 171 172 174 176 , there exists a constant \(C_5 > 0\) such that \[\label{nseries6469} \mathbb{P} \left( \big| \text{Term 1} - (E_1 + E_2 + E_3) \big| \ll \frac{k^2}{n^2 \rho_n^2} \right) \ge C_5.\tag{177}\]
Next, we move on to the second-order contributions from 165 : \[\begin{align} \label{nseries6}
\text{Term 2}
&=
\mathop{\mathrm{tr}}\Big(
\bigl[
V\Lambda^{-2}V^{\top}S_1(X)
+
S_1(X)V\Lambda^{-2}V^{\top}
-
V\Lambda^{-2}V^{\top}(XP+PX)V\Lambda^{-2}V^{\top}
\bigr]\mathop{\mathrm{Diag}}(T_1(X)\cdot \mathbf{1}_n)
\Big) \\
&=
\mathop{\mathrm{tr}}\Big(
\bigl[
\beta^{-3} X \beta^{\perp}
+
\beta^{\perp} X \beta^{-3}
-
\beta^{-2} X \beta^{-1}
-
\beta^{-1} X \beta^{-2}
\bigr]
\mathop{\mathrm{Diag}}(T_1(X)\cdot \mathbf{1}_n)
\Big).
\end{align}\tag{178}\] Focusing on the first term of 178 , we expand the trace as \[\begin{align} \label{nseries6461} \mathop{\mathrm{tr}}\left(
\beta^{-3} X \beta^{\perp} \mathop{\mathrm{Diag}}(T_1(X)\cdot \mathbf{1}_n) \right) & = \sum_{i,j,k,l} (\beta^{-3})_{ij} (\beta^\perp)_{ki} X_{jk} (T_1(X))_{il}\\ & \lesssim \sum_{i,j,k,m,l} (\beta^{-3})_{ij} (\beta^\perp)_{ki} (VV^{\top})_{im}
X_{jk} X_{ml}\\ & \lesssim \sum_{j,k,m,l} \underbrace{\left(\sum_i(\beta^{-3})_{ij} (\beta^\perp)_{ki} (VV^{\top})_{im} \right)}_{w^{(4)}_{jklm}} X_{jk} X_{ml}.
\end{align}\tag{179}\] Under 1 3 4, it follows that \(w^{(4)}_{jklm} \asymp \frac{k^2}{n^2(n \rho_n)^3}\). Consequently, it is straightforward to verify the expectation
\[\left|\mathbb{E}\left(\sum_{j,k,m,l} w^{(4)}_{jklm} X_{jk} X_{ml}\right) \right| \lesssim \frac{k^2}{n^3 \rho_n^2}.\] Next, we determine the order of the corresponding variance: \[\begin{align}
\mathop{\mathrm{Var}}\left(\sum_{j,k,m,l} w^{(4)}_{jklm} X_{jk} X_{ml}\right) & = \sum_{j,k,m,l} \left(w^{(4)}_{jklm} \right)^2 \mathop{\mathrm{Var}}\left(X_{jk} X_{ml} \right) + \sum_{(j,k,m,l) \neq \atop (j',k',m',l')} w^{(4)}_{jklm}
w^{(4)}_{j'k'l'm'} \mathop{\mathrm{Cov}}\left(X_{jk} X_{ml},X_{j'k'} X_{m'l'} \right) \\ & \lesssim \frac{k^4}{n^6 \rho_n^4}.
\end{align}\] Therefore, an application of Chebyshev’s inequality to 179 implies that with constant probability \[\begin{align} \label{nseries64620}
\big | \mathop{\mathrm{tr}}\left( \beta^{-3} X \beta^{\perp} \mathop{\mathrm{Diag}}(T_1(X)\cdot \mathbf{1}_n) \right) \big| \lesssim \frac{k^2}{n^3 \rho_n^2}.
\end{align}\tag{180}\] By analogous reasoning, identical asymptotic bounds hold for the remaining components in 178 , which ultimately yields \[\begin{align}
\label{nseries6463} \mathbb{P} \left(\big| \text{Term 2} \big| \lesssim \frac{k^2}{n^3 \rho_n^2}\right) \ge C_6.
\end{align}\tag{181}\] for some constant \(C_6>0\).
Finally, the sole second-order contribution from 166 is given by \[\text{Term 3} = \mathop{\mathrm{tr}}\left(\beta^{-2}\mathop{\mathrm{Diag}}(T_2(X)\cdot \mathbf{1}_n)\right).\] Invoking 13, we obtain \[T_2(X) = \bigl(\beta^{\perp} X \beta^{\perp} X \beta^{-1} + \beta^{-1} X \beta^{\perp} X \beta^{\perp} +
\beta^{\perp} X \beta^{-1} X \beta^{\perp}\bigr),\] which implies that \[\text{Term 3} = \sum_{a,b,c,d} w^{(5)}_{abcd}\, X_{ab} X_{cd},\] where the coefficients are defined as \[\begin{align} w^{(5)}_{abcd} &= \sum_{i,j} (\beta^{-2})_{ii} \Big[(\beta^\perp)_{ia} (\beta^\perp)_{bc} (\beta^{-1})_{dj} + (\beta^\perp)_{ia} (\beta^{-1})_{bc} (\beta^\perp)_{dj} + (\beta^{-1})_{ia} (\beta^\perp)_{bc}
(\beta^\perp)_{dj}\Big].
\end{align}\] Under 1 3 4, these coefficients satisfy the bounds \[\bigl| w^{(5)}_{abcd} \bigr| \asymp \begin{cases} \dfrac{k^2}{n^4 \rho_n^3}, & b=c, \\[10pt] \dfrac{k^2}{n^5 \rho_n^3},
& b\neq c. \end{cases}\] Define \[\begin{align}
\label{e464} E_4 &= \mathbb{E}\left(\mathop{\mathrm{tr}}\left(\beta^{-2}\mathop{\mathrm{Diag}}(T_2(X)\cdot \mathbf{1}_n)\right) \right) \\ \notag &= \mathbb{E}\left(\mathop{\mathrm{tr}}\left(\beta^{-2}\mathop{\mathrm{Diag}}(\left( \beta^{\perp} X
\beta^{\perp} X \beta^{-1} + \beta^{-1} X \beta^{\perp} X \beta^{\perp} + \beta^{\perp} X \beta^{-1} X \beta^{\perp}\right)\cdot \mathbf{1}_n)\right) \right).
\end{align}\tag{182}\] Consequently, it is straightforward to verify that the expectation satisfies \[\left |E_4 \right| = \left |\mathbb{E}\left(\sum_{a,b,c,d} w^{(5)}_{abcd} X_{ab} X_{cd}\right) \right| \lesssim
\frac{k^2}{n^2 \rho_n^2}.\] Next, we determine the asymptotic order of the corresponding variance: \[\begin{align} \mathop{\mathrm{Var}}\left(\sum_{a,b,c,d} w^{(5)}_{abcd} X_{ab} X_{cd}\right) &= \sum_{a,b,c,d}
\left(w^{(5)}_{abcd} \right)^2 \mathop{\mathrm{Var}}\left(X_{ab} X_{cd} \right) + \sum_{(a,b,c,d) \neq \atop (a',b',c',d')} w^{(5)}_{abcd} w^{(5)}_{a'b'c'd'} \mathop{\mathrm{Cov}}\left(X_{ab} X_{cd},X_{a'b'}
X_{c'd'} \right) \\ &= \sum_{a,b,c,d \atop b \neq c} \left(w^{(5)}_{abcd} \right)^2 \mathop{\mathrm{Var}}\left(X_{ab} X_{cd} \right) + \sum_{a,b,c,d \atop b = c} \left(w^{(5)}_{abcd} \right)^2 \mathop{\mathrm{Var}}\left(X_{ab} X_{cd} \right) \\
&\quad + \sum_{(a,b,c,d) \neq \atop (a',b',c',d')} w^{(5)}_{abcd} w^{(5)}_{a'b'c'd'} \mathop{\mathrm{Cov}}\left(X_{ab} X_{cd},X_{a'b'} X_{c'd'} \right) \\ &\lesssim \frac{k^4}{n^5 \rho_n^4}.
\end{align}\] Therefore, an application of Chebyshev’s inequality yields that, for some constant \(C_7>0\), \[\begin{align} \label{nseries6467}
\mathbb{P} \left(\left|\text{Term 3}- E_4 \right| \le \frac{k^2}{n^{5/2} \rho_n^2} \right) \ge C_7.
\end{align}\tag{183}\] Combining all the second-order terms from 177 181 183 , we conclude that there exists a constant \(C_8>0\) such that \[\begin{align} \label{nseries6468} \mathbb{P} \left(\left|\text{Term 4}- E \right| \ll \frac{k^2}{n^{2} \rho_n^2} \right) \ge C_8,
\end{align}\tag{184}\] where \(\text{Term 4}=\text{Term 1}-\text{Term 2}-\text{Term 3}\) contains all the quadratic terms of \(X\) in 163 , and \(E=E_1+E_2+E_3-E_4\) as defined earlier in 169 173 175 182 .
Combining these arguments, we have therefore shown that with at least constant probability, \[\begin{align} \label{note07} \mu_1^{(2)} - \widehat{\mu}_1^{(2)} &= E_1 + E_2 + E_3 - E_4 + (\text{terms of degree \ge 3 in X}) \\ &\quad + o_p \left(\max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right) \right).\notag \end{align}\tag{185}\]
Using the identity \[\mathbb{E}\left(X \beta^\perp X \right)= \Sigma \circ \beta^\perp + \mathop{\mathrm{Diag}}\left( \left(\Sigma - \mathop{\mathrm{Diag}}\left(\Sigma \right) \right) \mathrm{diag}\left(\beta^\perp \right) \right),\] we can evaluate \(E_1, E_2, E_3\), and \(E_4\): \[\begin{align} \label{e123} E_1 &= \mathop{\mathrm{tr}}\left(\beta^{-4} \left[ \Sigma \circ \beta^\perp + \mathop{\mathrm{Diag}}\left( \left(\Sigma - \mathop{\mathrm{Diag}}\left(\Sigma \right) \right) \mathrm{diag}\left(\beta^\perp \right) \right)\right] \beta^{\perp} \mathop{\mathrm{Diag}}\left(P \cdot \mathbf{1}_n \right) \right), \\ E_2 &= \mathop{\mathrm{tr}}\left(\beta^{\perp} \left[ \Sigma \circ \beta^{\perp} + \mathop{\mathrm{Diag}}\left( \left(\Sigma - \mathop{\mathrm{Diag}}\left(\Sigma \right) \right) \mathrm{diag}\left(\beta^{\perp} \right) \right)\right] \beta^{-4} \mathop{\mathrm{Diag}}\left(P \cdot \mathbf{1}_n \right) \right),\\ E_3 &= \mathop{\mathrm{tr}}\left(\beta^{\perp} \left[ \Sigma \circ \beta^{-4} + \mathop{\mathrm{Diag}}\left( \left(\Sigma - \mathop{\mathrm{Diag}}\left(\Sigma \right) \right) \mathrm{diag}\left(\beta^{-4} \right) \right)\right] \beta^{\perp} \mathop{\mathrm{Diag}}\left(P \cdot \mathbf{1}_n \right) \right), \end{align}\tag{186}\] and \[\begin{align} \label{e4} E_4 &= \mathop{\mathrm{tr}}\left(\beta^{-2} \mathop{\mathrm{Diag}}\left( \beta^\perp \left[ \Sigma \circ \beta^\perp + \mathop{\mathrm{Diag}}\left( \left( \Sigma - \mathop{\mathrm{Diag}}\left( \Sigma \right) \right) \mathrm{diag}(\beta^\perp) \right) \right] \beta^{-1} \cdot \mathbf{1}_n \right) \right) \\ &\quad + \mathop{\mathrm{tr}}\left(\beta^{-2} \mathop{\mathrm{Diag}}\left( \beta^\perp \left[ \Sigma \circ \beta^{-1} + \mathop{\mathrm{Diag}}\left( \left( \Sigma - \mathop{\mathrm{Diag}}\left( \Sigma \right) \right) \mathrm{diag}(\beta^{-1}) \right) \right] \beta^{\perp} \cdot \mathbf{1}_n \right) \right) \\ &\quad + \mathop{\mathrm{tr}}\left(\beta^{-2} \mathop{\mathrm{Diag}}\left( \beta^{-1} \left[ \Sigma \circ \beta^\perp + \mathop{\mathrm{Diag}}\left( \left( \Sigma - \mathop{\mathrm{Diag}}\left( \Sigma \right) \right) \mathrm{diag}(\beta^\perp) \right) \right] \beta^{\perp} \cdot \mathbf{1}_n \right) \right). \end{align}\tag{187}\] For the specific case of the SBM family, we show that \(E_1=E_3=E_4=0\) and \(E_2 \asymp \frac{k}{n^2\rho_n^2}\).
Here we prove that \(\beta^{\perp} \mathop{\mathrm{Diag}}\left(P \cdot \mathbf{1}_n \right) = \mathop{\mathrm{Diag}}\left(P \cdot \mathbf{1}_n \right) \beta^{\perp}\). We first analyze the structure of \(\beta^\perp\) under the SBM: \[\beta^\perp = I - VV^{\top}= I - Z(Z^{\top}Z)^{-1}Z^{\top}.\] Thus, \(\beta^\perp\) is a block-diagonal matrix with the \(m\)th block given by \((\beta^\perp)^{(m)}=I_{n_m}-\frac{1}{n_m}J_{n_m}\), where \(n_m\) is the size of the \(m\)th community. Similarly, we have \[\begin{align} \label{last1} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n) = \mathop{\mathrm{Diag}}( Z B Z^{\top}\mathbf{1}_n) = \mathop{\mathrm{Diag}}( Z \widetilde{d} ), \end{align}\tag{188}\] where \(\widetilde{d}=B Z^{\top}\mathbf{1}_n\) is a \(k \times 1\) vector with elements \(\widetilde{d}_m\) for \(m=1,\dots,k\). Thus, 188 implies that \(\mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\) also has a block structure, with the \(m\)th block being \(\mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)^{(m)}=\widetilde{d}_m I_{n_m}\). Since both \(\beta^\perp\) and \(\mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\) are block-diagonal matrices with matching block sizes, their product is also block-diagonal. The \(m\)th block of \(\beta^{\perp} \mathop{\mathrm{Diag}}\left(P \cdot \mathbf{1}_n \right)\) is therefore \(\widetilde{d}_m \left(I_{n_m}-\frac{1}{n_m}J_{n_m}\right)\), which is identical to the \(m\)th block of \(\mathop{\mathrm{Diag}}\left(P \cdot \mathbf{1}_n \right)\beta^{\perp}\). This proves that \(\beta^{\perp} \mathop{\mathrm{Diag}}\left(P \cdot \mathbf{1}_n \right) = \mathop{\mathrm{Diag}}\left(P \cdot \mathbf{1}_n \right) \beta^{\perp}\).
Consequently, for \(E_1\), we obtain \[\begin{align} \label{last02} E_1 &= \mathop{\mathrm{tr}}\left(\beta^{-4} \left[ \Sigma \circ \beta^\perp + \mathop{\mathrm{Diag}}\left( \left(\Sigma - \mathop{\mathrm{Diag}}\left(\Sigma \right) \right) \mathrm{diag}\left(\beta^\perp \right) \right)\right] \beta^{\perp} \mathop{\mathrm{Diag}}\left(P \cdot \mathbf{1}_n \right) \right)\\ &= \mathop{\mathrm{tr}}\left(\left[ \Sigma \circ \beta^\perp + \mathop{\mathrm{Diag}}\left( \left(\Sigma - \mathop{\mathrm{Diag}}\left(\Sigma \right) \right) \mathrm{diag}\left(\beta^\perp \right) \right)\right] \beta^{\perp} \mathop{\mathrm{Diag}}\left(P \cdot \mathbf{1}_n \right)\beta^{-4} \right)\\ &= \mathop{\mathrm{tr}}\left(\left[ \Sigma \circ \beta^\perp + \mathop{\mathrm{Diag}}\left( \left(\Sigma - \mathop{\mathrm{Diag}}\left(\Sigma \right) \right) \mathrm{diag}\left(\beta^\perp \right) \right)\right] \mathop{\mathrm{Diag}}\left(P \cdot \mathbf{1}_n \right) \beta^{\perp} \beta^{-4} \right) \\ &= 0. \end{align}\tag{189}\] The exact same argument applies to \(E_2\) as well.
For the SBM, since \(Z \cdot \mathbf{1}_k=\mathbf{1}_n\), we have \(\mathbf{1}_n \in \mathop{\mathrm{col}}(Z)=\mathop{\mathrm{col}}(V)\) because the block matrix \(B\) has full rank. This directly forces the second and third terms of 187 to vanish since \(\beta^\perp \cdot \mathbf{1}_n = 0\).
Next, we evaluate the first term: \[\mathop{\mathrm{tr}}\left(\beta^{-2} \mathop{\mathrm{Diag}}\left( \beta^\perp \left[ \Sigma \circ \beta^\perp + \mathop{\mathrm{Diag}}\left( \left( \Sigma - \mathop{\mathrm{Diag}}\left( \Sigma \right) \right) \mathrm{diag}(\beta^\perp) \right) \right] \beta^{-1} \cdot \mathbf{1}_n \right) \right).\] Our target is to show that the components \(\mathop{\mathrm{Diag}}\left( \beta^\perp \left[ \Sigma \circ \beta^\perp \right] \beta^{-1} \cdot \mathbf{1}_n \right)\), \(\mathop{\mathrm{Diag}}\left( \beta^\perp \left[ \mathop{\mathrm{Diag}}\left( \Sigma \, \mathrm{diag}(\beta^\perp) \right) \right] \beta^{-1} \cdot \mathbf{1}_n \right)\), and \(\mathop{\mathrm{Diag}}\left( \beta^\perp \left[ \mathop{\mathrm{Diag}}\left( \mathop{\mathrm{Diag}}\left( \Sigma \right) \mathrm{diag}(\beta^\perp) \right) \right] \beta^{-1} \cdot \mathbf{1}_n \right)\) vanish separately.
We established in the calculation for \(E_1\) that \((\beta^\perp)^{(m)}=I_{n_m}-\frac{1}{n_m}J_{n_m}\). The matrix \(\Sigma=P \circ (J_n-P)\) also possesses a block structure. Thus, \(\Sigma \circ \beta^\perp\) is a block-diagonal matrix with the \(m\)th block given by \((\Sigma \circ \beta^\perp)^{(m)}=P_{mm}(1-P_{mm})\left(I_{n_m}-\frac{1}{n_m}J_{n_m}\right)\). Furthermore, \(\beta^{-1}\cdot \mathbf{1}_n=V \Lambda^{-1} V^{\top}\cdot \mathbf{1}_n \in \mathop{\mathrm{col}}(Z)\) is a block-constant vector, with \(q_m\) denoting the value of the \(m\)th block. Evaluating the \(m\)th block of the vector \(\left[ \Sigma \circ \beta^\perp \right] \beta^{-1} \cdot \mathbf{1}_n\), we find \[\begin{align} \label{last03} \left(\left[ \Sigma \circ \beta^\perp \right] \beta^{-1} \cdot \mathbf{1}_n \right)^{(m)} &= \left(\left[ \Sigma \circ \beta^\perp \right]\right)^{(m)} q_m \mathbf{1}_{n_m}\\ &= P_{mm}(1-P_{mm})q_m\left(I_{n_m}-\frac{1}{n_m}J_{n_m}\right)\mathbf{1}_{n_m} = 0. \end{align}\tag{190}\] Thus, 190 proves that \(\left[ \Sigma \circ \beta^\perp \right] \beta^{-1} \cdot \mathbf{1}_n = 0\).
Next, \(\mathrm{diag}(\beta^\perp)\) is a block-constant vector with its \(m\)th block equal to \((1-\frac{1}{n_m})\). Consequently, \(\Sigma \, \mathrm{diag}(\beta^\perp)\) is also a block-constant vector with its \(m\)th block equal to \(P_{mm}(1-P_{mm})(1-\frac{1}{n_m})\). This yields \[\left(\mathop{\mathrm{Diag}}\left( \Sigma \, \mathrm{diag}(\beta^\perp) \right) \right)^{(m)} = P_{mm}(1-P_{mm})(1-\frac{1}{n_m})I_{n_m}.\] We then have \[\begin{align} \label{last04} \left(\beta^\perp \left[ \mathop{\mathrm{Diag}}\left( \Sigma \, \mathrm{diag}(\beta^\perp) \right) \right] \beta^{-1} \cdot \mathbf{1}_n\right)^{(m)} &= (\beta^\perp)^{(m)} \left(\mathop{\mathrm{Diag}}\left( \Sigma \, \mathrm{diag}(\beta^\perp) \right) \right) ^{(m)} \left(\beta^{-1} \cdot \mathbf{1}_n\right)^{(m)}\\ &= P_{mm}(1-P_{mm}) q_{m} \left(1-\frac{1}{n_m}\right) (\beta^\perp)^{(m)}\mathbf{1}_{n_m} = 0. \end{align}\tag{191}\] This proves that \(\mathop{\mathrm{Diag}}\left( \beta^\perp \left[ \mathop{\mathrm{Diag}}\left( \Sigma \, \mathrm{diag}(\beta^\perp) \right) \right] \beta^{-1} \cdot \mathbf{1}_n \right) = 0\). By a nearly identical argument, one can show that \(\mathop{\mathrm{Diag}}\left( \beta^\perp \left[ \mathop{\mathrm{Diag}}\left( \mathop{\mathrm{Diag}}\left( \Sigma \right) \mathrm{diag}(\beta^\perp) \right) \right] \beta^{-1} \cdot \mathbf{1}_n \right) = 0\). Combining this result with 190 191 confirms that \(E_4 = 0\).
From 175 , we know that \[E_3 = \mathbb{E}\left(\mathop{\mathrm{tr}}\left(\beta^{\perp} X \beta^{-4} X \beta^{\perp} \mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\right)\right) = \mathbb{E}\left(\|\beta^{-2}X \beta^{\perp} \left(\mathop{\mathrm{Diag}}(P \cdot \mathbf{1}_n)\right)^{1/2}\|_F^2 \right)>0.\] Now, we evaluate the exact asymptotic order of \(E_3\). We already know that \(\Sigma\) is a block-constant matrix. Additionally, \(\beta^{-4}= V \Lambda^{-4} V^{\top}\in \mathop{\mathrm{col}}(Z)\) is also a block-constant matrix. Thus, \(\Sigma \circ \beta^{-4}\) is a block-constant matrix, meaning that \(\mathop{\mathrm{col}}(\Sigma \circ \beta^{-4}) \subseteq \mathop{\mathrm{col}}(Z)\). This implies that \(\beta^\perp (\Sigma \circ \beta^{-4})=0\). Thus, from 186 , we have \[\begin{align} \label{last05} E_3 &= \mathop{\mathrm{tr}}\left(\beta^{\perp} \left[ \mathop{\mathrm{Diag}}\left( \left(\Sigma - \mathop{\mathrm{Diag}}\left(\Sigma \right) \right) \mathrm{diag}\left(\beta^{-4} \right) \right)\right] \beta^{\perp} \mathop{\mathrm{Diag}}\left(P \cdot \mathbf{1}_n \right) \right)\\ &= \sum_{i=1}^n (\mathop{\mathrm{Diag}}(P\cdot \mathbf{1}_n))_{ii} \cdot \left(\mathop{\mathrm{Diag}}\left( \left(\Sigma - \mathop{\mathrm{Diag}}\left(\Sigma \right) \right) \mathrm{diag}\left(\beta^{-4} \right) \right)\right)_{ii} \cdot (\beta^\perp)_{ii}. \end{align}\tag{192}\] Now, \[\left(\mathop{\mathrm{Diag}}\left( \left(\Sigma - \mathop{\mathrm{Diag}}\left(\Sigma \right) \right) \mathrm{diag}\left(\beta^{-4} \right) \right)\right)_{ii}= \sum_{j \neq i} \Sigma_{ij} \beta_{jj}^{-4} \asymp \frac{k}{n^4 \rho_n^3}, \quad (\mathop{\mathrm{Diag}}(P\cdot \mathbf{1}_n))_{ii}=\sum_{j} P_{ij} \asymp n \rho_n.\] Since we have \((\beta^{\perp})_{ii} \asymp 1\) from 4, 192 implies that \(E_3 \asymp \frac{k}{n^2 \rho_n^2}\).
Consequently, we have established that under the SBM, the combined term \(E = E_1 + E_2 + E_3 - E_4 \asymp \frac{k}{n^2 \rho_n^2}\). Combining this result with 185 yields \[\begin{align} \label{note08} \mu_1^{(2)} - \widehat{\mu}_1^{(2)} &\asymp \frac{k}{n^2 \rho_n^2} + (\text{terms of degree \ge 3 in X}) + o_p \left(\max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right) \right). \end{align}\tag{193}\] In the final step, we bound the remaining higher-order terms to complete the proof.
To bound the higher-order terms (those with power of \(X\) greater than or equal to \(3\)), we begin with the final term in 163 which has been expanded in 166 . Using 13, we have \[\begin{align} \mathop{\mathrm{tr}} \left( V \Lambda^{-2} V^{\top}\,\mathop{\mathrm{Diag}}(( {\widehat P}-P) \cdot \mathbf{1}_n) \right) &= \sum_{l \ge 3} \mathop{\mathrm{tr}}\left( V \Lambda^{-2} V^{\top}\mathop{\mathrm{Diag}}(T_l(X)\cdot \mathbf{1}_n) \right) \\ &\lesssim \sum_{l \ge 3} \frac{k}{(n\rho_n)^2} \left(\frac{4}{n\rho_n}\right)^{l/2} \, n\rho_n \\ &\lesssim \frac{k}{(n\rho_n)^{2.5}}. \end{align}\]
To bound the remaining higher-order contributions from 163 , we revisit the decomposition in 164 and follow an argument entirely analogous to the proof of 30 in 13.4. The sole distinction is that in 13.4, after applying Bernstein’s inequality to the linear terms, we established a high-probability bound for terms of degree two or higher in \(X\); here, however, we restrict our attention to terms of degree three or higher. The remainder of the proof remains unchanged, yielding a bound of \(O_p\left(\frac{k}{(n\rho_n)^{2.5}}\right)\). The higher-order terms in 165 can be handled analogously using the same reasoning as in the proof of 30 . Again, the only modification is that we now discard all terms of order \(3\) or higher in \(X\). This gives the same bound, \(O_p\left(\frac{k}{(n\rho_n)^{2.5}}\right).\)
Combining the above bounds for terms of order \(3\) or higher in \(X\) expanded in 164 165 166 , we conclude that the total contribution of all terms with power of \(X\) greater than or equal to \(3\) in 163 is \(O_p\left(\frac{k}{(n\rho_n)^{2.5}}\right) = o_p\left(\frac{k}{n^2\rho_n^2}\right).\) Thus, 193 now can be updated to conclude that \[\begin{align} \big|\mu_1^{(2)} - \widehat{\mu}_1^{(2)}\big| &\asymp \frac{k}{n^2 \rho_n^2}+ o_p\left(\frac{k}{n^2\rho_n^2}\right) + o_p \left(\max \left(\frac{k}{n^2 \rho_n^2},\frac{k \sqrt{\log n}}{n^2 \rho_n^{3/2}}\right) \right) \\ &\asymp \frac{k}{n^2 \rho_n^2}, \end{align}\] which completes the proof.