Interactive fixed effects are routinely controlled for in linear panel models. While an analogous fixed effects (FE) estimator for nonlinear models has been available in the literature [1], it sees much more limited use in applied research because its implementation involves solving a high-dimensional non-convex problem. In this paper, we complement the theoretical analysis of
[1] by providing a new computationally efficient estimator that is asymptotically equivalent to their estimator. Unlike the previously
proposed FE estimator, our estimator avoids solving a high-dimensional non-convex optimization problem and can be feasibly computed in large nonlinear panels. Our proposed method involves two steps. In the first step, we convexify the optimization problem
using nuclear norm regularization (NNR) and obtain preliminary NNR estimators of the parameters, including the fixed effects. Then, we find the global solution of the original optimization problem using a standard gradient descent method initialized at
these preliminary estimates. To make our method readily applicable in practice, we also propose specific numerical algorithms for solving the involved optimization problems, establish their convergence, and provide their efficient implementation in our R
package NNRPanel.
The importance of accounting for interactive unobserved heterogeneity in panel and network models is well recognized. For example, in linear panel models, interactive fixed effects are routinely controlled for using the seminal approaches of [2] or [3]. While analogous methods for nonlinear models have been developed in the literature (e.g., [1]), they see much more limited use in empirical research due to their rapidly growing computational complexity or the lack of inferential theory.4
The main goal of this paper is to bridge the gap between the recent theoretical developments by [1] and empirical work by providing a new computationally efficient estimator that can be feasibly implemented in a wide range of nonlinear (semiparametric) settings with unobserved effects following a linear factor structure. We demonstrate that our estimator has two important properties. First, unlike the approach of [1], our method does not require solving a high-dimensional non-convex optimization problem, so our estimator can be efficiently computed when the cross-sectional and time dimensions, \(N\) and \(T\), are large. Second, we argue that our estimator is asymptotically equivalent to the fixed effects (FE) estimator of [1]. This means that, in practice, one can combine our computationally efficient estimator with the inferential theory provided in [1] to construct confidence intervals for various objects of interest, including structural parameters and average partial effects.
Our proposed estimation procedure involves the following two steps. In the first step, we construct preliminary estimators of the parameters of interest, including the loadings and the factors, by solving a convex relaxation of the original (non-convex) optimization problem in [1]. Following the literature, we convexify the original problem by replacing the low-rank constraint imposed on the unobserved effects by the factor model with a nuclear norm penalty. Then, we compute our final estimator by solving the original optimization problem using a standard gradient descent method initialized at the preliminary nuclear norm regularized (NNR) estimator obtained in the first step. To demonstrate that our final estimator is asymptotically equivalent to the FE estimator of [1] defined as the global solution of the original high-dimensional and non-convex optimization problem, we show that the original problem is locally convex in a shrinking neighborhood around the true values of the parameters. Importantly, in the general nonlinear setting studied in this paper (with a growing number of factors and loadings as \(N,T \rightarrow \infty\)), the size of this neighborhood shrinks at a certain rate. To establish the desired result, we characterize the rate of convergence of our preliminary NNR estimator, and demonstrate that this rate is sufficiently fast to ensure that our NNR estimator, as well as the FE estimator, falls into that shrinking neighborhood with probability approaching one. The idea of using a preliminary NNR estimator to initialize local optimization in (globally) non-convex problems has been previously explored in the econometrics literature. For example, [5] originally proposed an analogous two-step approach for estimating linear panel models with interactive fixed effects. In particular, [5] also demonstrate that their two-step estimator is asymptotically equivalent to the LS estimator of [2]. However, extending these ideas and formally establishing an analogous equivalence result in the general nonlinear setting of [1] is a non-trivial task involving additional technical challenges.
As highlighted above, the main conceptual and technical difference is that, in the general nonlinear case, the objective function is locally convex only in a shrinking neighborhood of the true parameters value. In particular, unlike in the linear case, one cannot simply profile out the fixed effects using the singular value decomposition, and demonstrate that the profiled objective function (only depending on the common parameters \(\beta\)) is locally convex. Since, in the nonlinear case, we cannot work with the profiled objective function directly, we establish local convexity of the original objective function by inspecting its Hessian taken with respect to all of the parameters including the loadings and the factors. The analysis is further complicated by the fact that the dimension of the parameter space and hence the dimension of the Hessian grows with \(N,T \rightarrow \infty\). As a result, local convexity of the objective function can only be established in a shrinking neighborhood of the true parameters. Establishing local convexity in that neighborhood and characterizing at which rate it shrinks is a technical innovation of the paper that has important practical implications. Specifically, it imposes an additional requirement on the preliminary estimator’s rate of convergence: unless the preliminary estimator falls into that shrinking convexity region with probability approaching one, one cannot guarantee that the second step local optimization finds the global solution. In particular, it turns out that the rate obtained by [5] for the NNR estimator in single-index models is not sufficiently fast to satisfy this requirement.
To take advantage of the local convexity result described above, we provide a new improved error bound for the NNR estimator in nonlinear models with interactive fixed effects. Following the literature, we derive this result under a version of the restricted strong convexity (RSC) condition. While various variations of the RSC condition are routinely employed for deriving analogous results in low-rank models (e.g., [5]–[7]), these conditions are often difficult to verify. Unlike most previous studies, we provide a set of primitive conditions which can be used to verify the RSC condition in a wide range of panel models allowing, in particular, for predetermined covariates.
To further facilitate the practicality of our approach, we supplement it with concrete implementation details. In particular, we provide specific optimization algorithms, which can be used to efficiently compute the preliminary NNR and the final
estimators, and establish their convergence. We also propose data-driven ways of choosing the regularization parameter involved in the first step and determining the unknown number of factors. To make the proposed method readily applicable, we provide its
efficient implementation in an optimized R package, NNRPanel.5
Finally, we study the finite sample properties of our estimator in a number of numerical experiments, documenting its excellent performance and computational efficiency even in large panels with \({(N,T) = (1000,200)}\), and revisit the empirical application of [1].
This paper contributes to the literature on estimation of panel (and network) models with interactive fixed effects in two important ways.
First, we complement the theoretical analysis of nonlinear panel models provided by [1] by proposing a new estimator that is asymptotically equivalent to their FE estimator. Importantly, unlike their FE estimator, our estimator does not involve solving a high-dimensional non-convex optimization problem, making it an attractive, if not the only available, computationally efficient alternative, which can be feasibly implemented even when both \(N\) and \(T\) are large. Our two-step approach to solving a non-convex optimization problem essentially extends the proposal of [5] to nonlinear settings. However, as explained above, establishing the asymptotic equivalence between the two-step and FE estimators in nonlinear settings is more nuanced: it involves careful establishing of local convexity of the criterion function in a shrinking neighborhood of the true parameters value, resulting in additional requirements imposed on the preliminary NNR estimator’s convergence rate absent in the linear case studied by [5].
Second, we also contribute to the literature on nuclear norm regularized estimation of low-rank models by extending the previously available results established by [5] and [6] for linear panel models to nonlinear settings. By verifying the RSC condition in a wide class of nonlinear panel models, we improve on the result of [5], which extends their original analysis to single-index models. Importantly, our analysis allows for predetermined covariates such as the outcome’s lags, which are routinely used in panels, whereas the existing studies providing error bounds for the NNR estimator either only consider strictly exogenous covariates or do not verify the RSC condition at all.
The idea of using nuclear norm regularization to turn the estimation of low-rank models, such as factor models, into a convex problem has been extensively applied in various settings in statistics and econometrics. In econometrics, its numerous recent applications include estimation of pure factor models [8], [9], estimation of linear [5], [6], [10], [11] and quantile panel regressions [12]–[14], and treatment effect estimation [15], [16]. Nuclear norm relaxations have also been proved instrumental in constructing estimation and inference methods robust to weak factors [17] and missing data [18]. Other recent applications of nuclear norm regularization also include, among others, network recovery and community detection ([19] and [7]), and estimation of panel threshold models and high-dimensional VARs ([20] and [21]).
Recent papers employing similar two-step estimation procedures in nonlinear panels also include [22] and [23]. [22] study estimation of logistic panel models without covariates, whereas we (mainly) focus on the common parameters \(\beta\). In a closely related and independently developed paper, [23] also studies estimation of nonlinear panel models with interactive fixed effects. In particular, [23] allows for nonconvex model specifications such as the important class of random-coefficients models. This generality, however, comes at a cost: even after the low-rank constraint is convexified using the nuclear norm penalty, the first step of [23]’s procedure still requires solving a nonconvex problem. Since the goal of our paper is to provide a theoretically justified and computationally feasible estimator that avoids nonconvex optimization, we focus on single-index models with convex link functions. As a result, we are able to verify a global version of the RSC condition, allowing us to formally study the global properties of the NNR estimator. We also provide specific numerical algorithms for solving both first and second step optimization problems and establish their convergence.
For any vector \(u\in \mathbb{R}^n\), its Euclidean norm is denoted as \(\|u\| = \left(u'u\right)^{\frac{1}{2}}\). For any matrix \(A\in \mathbb{R}^{m\times n}\), we use \(A'\) to denote the transpose of \(A\), and use \(\|A\|_{\mathrm{F}} = \left(\mathrm{trace}( A'A )\right)^{\frac{1}{2}}\) to denote the Frobenius norm. Furthermore, the singular values of \(A\) are arranged in non-increasing order: \(\psi_1\left(A\right)\geq \psi_2\left(A\right) \geq \ldots \geq \psi_{\min\{m, n\}}\left(A\right) \geq 0\). The \(\ell^2\) operator norm, \(\|A\|_{\mathrm{op}} = \psi_1\left(A\right)\), is the maximum singular value of the matrix, and the nuclear norm is the sum of all singular values: \(\|A\|_{\mathrm{nuc}} = \sum_{i=1}^{\min\{m, n\}}\psi_i\left(A\right)\). We also use \(\|A\|_{\max} = \max_{i,j} |A_{ij}|\) to denote the element-wise max norm. When \(A\) is a square matrix, we use \(\sigma_i(A)\) to denote \(A\)’s \(i\)-th largest eigenvalue. We also use \(\psi_{\max}\), \(\psi_{\min}\), \(\sigma_{\max}\), \(\sigma_{\min}\) to denote the max/min singular values and max/min eigenvalues respectively. For any matrix \(A\), define the coprojection matrix as \(M_{A}: = \mathbb{I} - A(A'A)^{\dagger }A'\), where \(\mathbb{I}\) denotes the identity matrix of appropriate size and the super-script \(\dagger\) denotes the Moore-Penrose generalized inverse. Finally, for any two square matrices \(A\) and \(B\) of the same dimension, we use \(A \geq B\) to denote that \(A - B\) is positive semi-definite, and \(A > B\) to denote that \(A - B\) is positive definite. We use the abbreviation wpa1 instead of with probability approaching one.
The remainder of the paper is organized as follows. Section 2 introduces the model, highlights the computational challenges of the FE estimator, and describes the proposed two-step estimator. Section 3 presents the asymptotic equivalence between our two-step estimator and the FE estimator. Section 4 provides practical implementation details, including optimization algorithms, data-driven selection of the regularization parameter, and determination of the number of factors. Section 5 provides numerical and empirical illustrations. Additional technical and numerical results, along with all proofs, are contained in the appendix.
We observe data \(\{(Y_{it}, X_{it})\}_{ 1 \leq i \leq N, 1 \leq t\leq T}\), where \(Y_{it}\) is a scalar outcome variable and \(X_{it} \in \mathbb{R}^{d_X}\) is a vector of covariates. For concreteness, we adopt the standard panel notation with \(i\) indexing units and \(t\) indexing time periods, but it should be understood that the considered framework applies to general two-way settings. For example, in a directed network \(i\) and \(t\) could index senders and receivers (e.g., exporters and importers in an international trade network). The covariates \(X_{it}\) could be strictly exogenous or predetermined, e.g., our framework also accommodates lagged outcomes as covariates in panels.
Following [1], we assume that the (conditional) distribution of \(Y_{it}\) belongs to a known family of distributions and is determined by the latent index \(Y_{it}^*\), i.e., we assume that the (conditional) log-likelihood takes the form \[\begin{align} \label{eq:true95model} \log f (Y_{it}|X_{it},\lambda_{0,i},\gamma_{0,t}) = \ell (Y_{it}|Y^*_{it}), \quad Y^*_{it} = X'_{it}\beta_0 + \lambda_{0, i}' \gamma_{0, t}, \end{align}\tag{1}\] where \(\ell(\cdot|Y_{it}^*)\) is a known log-likelihood function, and \(\beta_0 \in \mathbb{R}^{d_X}\) is the parameter of interest. Here, \({\lambda_{0,i} \in \mathbb{R}^R}\) and \(\gamma_{0, t} \in \mathbb{R}^R\) are unobserved interactive unit and time effects, commonly referred to as loadings and factors. This formulation is substantially more flexible than the routinely employed two-way fixed effects (TWFE) model, \(\lambda_{0, i} + \gamma_{0, t}\), because it allows incorporating multidimensional heterogeneous individual responses \(\lambda_{0, i}\) to time-varying aggregate shocks \(\gamma_{0, t}\).6 In particular, the TWFE model corresponds to the special case of the interactive fixed effects model with \(R = 2\), where \(\lambda_{0, i} = (\lambda_{i1}, 1)'\) and \(\gamma_{0, t} = (1, \gamma_{ t1})'\).
While the single-index formulation 1 is restrictive, it covers a number of important nonlinear models, including binary response models such as Probit and Logit, and Poisson regression.
Example 1 (Binary response panel model). Consider the following binary response panel model \[\begin{align} \label{eq:32binary32choice32example} Y_{it} = \boldsymbol{1}(X_{it}'\beta + \lambda_i'\gamma_t- u_{it} \geq 0), \end{align}\qquad{(1)}\] where the idiosyncratic errors \(\{u_{it}\}\) are independent draws from a distribution with cumulative distribution function \(F(\cdot)\).7
Special variations of ?? with unobserved heterogeneity fully captured by additive unit and/or time effects are routinely used in empirical studies of panels with binary outcomes such as labor force participation. As explained above, the general specification ?? allows not only for TWFE but also for heterogeneous responses of units to aggregate shocks. E.g., in the labor force participation example, \(\lambda_i\) and \(\gamma_t\) can represent multidimensional human capital and its market prices that vary over time, respectively. Importantly, our framework accommodates predetermined covariates \(X_{it}\) such as lagged outcomes \(Y_{i,t-1}\), \(Y_{i,t-2}\), etc., which are often included in binary response panel models either as primary variables of interest or as controls.
In this example, the conditional distribution of \(Y_{it}\) given \(Y_{it}^*\) takes the form \[\begin{align} \mathbb{P}(Y_{it} = y\mid Y^*_{it}) = F(Y^*_{it})^{y} (1-F(Y^*_{it}))^{(1-y)},\quad y\in \{0, 1\}, \end{align}\] and the conditional log-likelihood is given by \[\begin{align} \ell (Y_{it}|Y_{it}^*) = Y_{it} \log F(Y_{it}^*) + (1 - Y_{it}) \log (1 - F(Y_{it}^*)). \end{align}\]
Example 2 (Poisson network regression). Our analysis applies to network settings. For example, the studied framework generalizes the TWFE Poisson regression model (e.g., [24]) commonly employed to analyze international trade data. Specifically, consider \[\begin{align} Y_{ij} \sim \mathrm{ Poisson }\left(\exp\left( X_{ij}'\beta + \lambda_i'\gamma_j\right)\right), \end{align}\] where \(Y_{ij}\) represents the export volume from country \(i\) to \(j\), and \(X_{ij}\) is a collection of trade determinants used in the gravity analysis, e.g., the geographic distance between countries \(i\) and \(j\). This specification allows for rich patterns of unobserved heterogeneity, including degree heterogeneity and homophily based on latent characteristics. For example, homophily based on the quadratic distance \((\xi_i - \xi_j)^2\) can be captured by \(\lambda_{i} \propto (\xi_i^2, 1, -2\xi_i)'\) and \(\gamma_{j} \propto (1, \xi_j^2, \xi_j)'\).
In this example, the conditional distribution of \(Y_{ij}\) given \(Y_{ij}^*\) takes the form \[\begin{align} \mathbb{P}(Y_{ij} = y\mid Y^*_{ij}) = \frac{\exp (-\exp(Y^*_{ij}))(\exp(Y^*_{ij}))^{y}}{y!}, \quad y = 0,1,2,\ldots, \end{align}\] and the conditional log-likelihood is given by \[\begin{align} \ell (Y_{ij}|Y_{ij}^*) = Y_{ij} Y_{ij}^* - \exp (Y_{ij}^*) - \log (Y_{ij}!). \end{align}\] We will revisit this specification in our empirical application in Section 5.2.
As in [1], we consider the so-called large \(N,T\) asymptotics with \(N,T \rightarrow \infty\), whereas we treat both \(d_X\) and \(R\) as fixed. For now, we will also assume that the number of factors \(R\) is known; we will discuss the estimation of \(R\) in Section 4. Finally, we do not put additional restrictions on the relationship between the covariates and the unobserved effects, i.e., we adopt the fixed effects approach.
[1] proposed estimating the studied model by the fixed effects (FE) MLE estimator maximizing the conditional log-likelihood jointly over the common parameters \(\beta\), loadings \(\{\lambda_{i}\}_{ 1\leq i\leq N}\) and factors \(\{\gamma_{ t}\}_{1\leq t\leq T}\). Specifically, the FE estimator \((\hat{\beta}_{\mathrm{FE}}, \hat{\Lambda}_{\mathrm{FE}}, \hat{\Gamma}_{\mathrm{FE}})\) solves \[\label{eq:FE95estimator} (\hat{\beta}_{\mathrm{FE}}, \hat{\Lambda}_{\mathrm{FE}}, \hat{\Gamma}_{\mathrm{FE}}) \in \mathop{\mathrm{argmin}}_{ \beta, \Lambda, \Gamma} \underbrace{-\frac{1}{NT} \sum_{i=1}^{N}\sum_{t=1}^{T} \ell(Y_{it}\mid X_{it}'\beta + \lambda_{ i}'\gamma_{ t})}_{\mathcal{L}_{NT}(\beta, \Lambda, \Gamma)},\tag{2}\] where, for notational simplicity, we collect the unobserved effects \(\{\lambda_{i}\}_{ 1\leq i\leq N}\) and \(\{\gamma_{ t}\}_{1\leq t\leq T}\) into matrices \(\Lambda = \left(\lambda_{1}, \lambda_{2},\ldots, \lambda_{N} \right)'\in \mathbb{R}^{N\times R}\) and \(\Gamma = \left(\gamma_{1}, \gamma_{2},\ldots, \gamma_{T} \right)'\in \mathbb{R}^{T\times R}\). Note that problem 2 does not have a unique solution for \(\hat{\Lambda}_{\mathrm{FE}}\) and \(\hat{\Gamma}_{\mathrm{FE}}\) and thus requires a normalization. We will abstract from this issue for now and discuss it in more detail in Section 3.
[1] showed that the FE estimator of \(\beta_0\) is \(\sqrt{NT}\)-consistent and asymptotically normal, with an asymptotic incidental parameter bias that can be corrected using various bias reduction methods. However, despite these well-established theoretical properties, implementing the FE estimator remains a significant computational challenge.
The key computational difficulty is the non-convexity of the objective function \(\mathcal{L}_{NT}(\beta, \Lambda, \Gamma)\). To better understand this issue, we reformulate the original optimization problem into an alternative but equivalent form. Let \(\theta_{it} = \lambda_i'\gamma_t\) and collect \(\theta_{it}\) into a matrix \(\Theta\in \mathbb{R}^{N\times T}\). Note that since matrices \(\Lambda\) and \(\Gamma\) have at most rank \(R\), the rank of \(\Theta = \Lambda \Gamma'\) is also at most \(R\). Likewise, any matrix \(\Theta \in \mathbb{R}^{N \times T}\) such that \(\mathrm{rank}(\Theta)\leq R\) can be represented as \(\Lambda \Gamma'\) for some \(\Lambda \in \mathbb{R}^{N \times R}\) and \(\Gamma \in \mathbb{R}^{T \times R}\).8 Thus, problem 2 can be equivalently reformulated as \[\begin{align} \label{eq:FE95rank} (\hat{\beta}_{\mathrm{FE}}, \hat{\Theta}_{\mathrm{FE}}) \in \mathop{\mathrm{argmin}}_{\beta\in \mathbb{R}^{d_X}, \Theta \in \mathbb{R}^{N\times T} } \underbrace{ -\frac{1}{NT} \sum_{i=1}^{N}\sum_{t = 1}^{T} \ell(Y_{it}\mid X_{it}'\beta + \theta_{it})}_{\mathcal{L}_{NT}(\beta, \Theta )}, \quad \text{s.t. } \mathrm{rank}(\Theta)\leq R. \end{align}\tag{3}\] The non-convexity arises from the rank constraint \(\mathrm{rank}(\Theta)\leq R\): the set of matrices satisfying it is not convex since the sum of two rank-\(R\) matrices could have a rank up to \(2R\).
The high-dimensional parameter space further exacerbates the computational challenges. When dealing with non-convex optimization problems, it is common practice to start the optimization process with multiple initial values and select the solution that minimizes the objective function. This approach is generally considered effective for finding the global minimum with sufficient trials. However, for problems 2 and 3 involving \(d_X + R (N+T)\) parameters, this approach becomes intractable even for moderate values of \(N\) and \(T\).
Remark 1. [1] proposed solving optimization problem 2 using the EM-algorithm of [25] initialized at multiple starting values. Unfortunately, this method does not overcome the computational challenge discussed above because the EM-algorithm of [25] as well as EM-algorithms in general does not have global convergence guarantees in non-convex problems.
To overcome the computational challenges faced by the FE estimator, we propose an alternative two-step estimation procedure. Our procedure does not involve solving a non-convex problem and can be efficiently computed even for large values of \(N\) and \(T\). Importantly, in Section 3, we demonstrate that, under standard regularity conditions, our two-step estimator is asymptotically equivalent to the FE estimator, whose asymptotic properties have been established in [1]. This means that, instead of trying to solve the non-convex and high-dimensional optimization problem 2 directly, one could compute our two-step estimator and then combine it with the asymptotic theory developed by [1] to construct confidence intervals for various parameters of interest, including \(\beta\) and policy relevant counterfactuals such as average partial effects (APEs).
Our estimation procedure involves the following two steps.
The goal of the first step is to construct an easily computable preliminary estimator of \((\beta_0, \Lambda_0, \Gamma_0)\) that is sufficiently close to the global minimizer in 2 . To this end, we consider a convex relaxation of problem 3 of the form \[\begin{align} \label{eq:nnr95definition} \left(\hat{\beta}_{\mathrm{nuc}}, \hat{\Theta}_{\mathrm{nuc}}\right) = \mathop{\mathrm{argmin}}_{\beta \in \mathbb{R}^{d_X}, \Theta\in \mathbb{R}^{N\times T}} \left\{ \mathcal{L}_{NT} \left( \beta, \Theta\right) + \frac{\varphi_{NT}}{\sqrt{NT}} \|\Theta\|_{\mathrm{nuc}}\right\}, \end{align}\tag{4}\] where \(\|\Theta\|_{\mathrm{nuc}}\) denotes the nuclear norm of matrix \(\Theta\), and \(\varphi_{NT} > 0\) is a regularization parameter. We will refer to the solution of this problem \((\hat{\beta}_{\mathrm{nuc}}, \hat{\Theta}_{\mathrm{nuc}})\) as the nuclear norm regularized (NNR) estimator.
Since \(\|\Theta\|_{\mathrm{nuc}}\) is a convex function of \(\Theta\), problem 4 is convex when \(\mathcal{L}_{NT} \left( \beta, \Theta\right)\) is a convex function of \(\beta\) and \(\Theta\). This condition is satisfied in important nonlinear models such as Logit, Probit, and Poisson models. Thanks to the convexity of problem 4 , the NNR estimator can be efficiently computed using, for example, a proximal gradient descent method (e.g., [26]) even when the parameter space is high-dimensional. We provide a specific optimization algorithm and a data-dependent recommendation for choosing the regularization parameter \(\varphi_{NT}\) in Section 4.
Notice that problem 4 can be equivalently rewritten as \[\begin{align} \left(\hat{\beta}_{\mathrm{nuc}}, \hat{\Theta}_{\mathrm{nuc}}\right) = \mathop{\mathrm{argmin}}_{\beta \in \mathbb{R}^{d_X}, \Theta\in \mathbb{R}^{N\times T}} \mathcal{L}_{NT} \left( \beta, \Theta\right), \quad \text{s.t. } \|\Theta\|_{\mathrm{nuc}} \leqslant C_{\varphi_{NT}} \end{align}\] for an appropriately chosen \(C_{\varphi_{NT}} > 0\) determined by \(\varphi_{NT}\). Thus, problem 4 can be seen as a convexification of problem 3 , where the non-convex rank constraint is replaced by the slightly looser yet convex constraint \(\|\Theta\|_{\mathrm{nuc}} \leqslant C_{\varphi_{NT}}\). Analogously to LASSO using the \(\ell_1\)-regularization to induce sparsity of the solution in a high-dimensional regression, the nuclear norm regularization (i.e., the \(\ell_1\)-regularization of the singular values of \(\Theta\)) induces \(\hat{\Theta}_{\mathrm{nuc}}\) to have low rank (i.e., sparsity of its singular values).
Finally, the nuclear norm regularized estimators \((\hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) are obtained through the singular value decomposition of \(\hat{\Theta}_{\mathrm{nuc}}\). Specifically, let \(\hat{\Theta}_{\mathrm{nuc}}/\sqrt{NT} = \hat{U}\hat{D}\hat{V}'\), where \(\hat{U}\in \mathbb{R}^{N\times \min\{N, T\}}\) and \(\hat{V}\in \mathbb{R}^{T\times \min\{N, T\}}\) are matrices of left and right (orthonormal) singular vectors of \(\hat{\Theta}_{\mathrm{nuc}}\), and \(\hat{D}\) is a diagonal matrix with the singular values of \(\hat{\Theta}_{\mathrm{nuc}}/\sqrt{NT}\) (arranged in non-increasing order) on its diagonal. Let \(\hat{U}_{[:, 1:R]}\) and \(\hat{V}_{[:, 1:R]}\) denote the matrices containing the first \(R\) columns of \(\hat{U}\) and \(\hat{V}\), respectively, and \(\hat{D}_{[1:R, 1:R]}\) denote the upper-left \(R\times R\) diagonal block of \(\hat{D}\). We compute \((\hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) as follows: \[\begin{gather} \label{eq:space95definition} \hat{\Lambda}_{\mathrm{nuc}} = \sqrt{N} \hat{U}_{[:, 1:R]}\hat{D}^{1/2}_{[1:R, 1:R]}, \quad \hat{\Gamma}_{\mathrm{nuc}} = \sqrt{T} \hat{V}_{[:, 1:R]}\hat{D}^{1/2}_{[1:R, 1:R]}. \end{gather}\tag{5}\]
While, with appropriately chosen \(\varphi_{NT}\), the NNR estimator is consistent for the true values \((\beta_0, \Lambda_0, \Gamma_0)\), it suffers from the regularization bias. To improve on the NNR estimator, in the second step, we solve the original optimization problem 2 using a standard gradient descent method with \((\hat{\beta}_{\mathrm{nuc}}, \hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) as the initial values. While the original problem 2 is non-convex, availability of the NNR estimator allows us to guarantee that standard local optimization methods initialized at \((\hat{\beta}_{\mathrm{nuc}}, \hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) converge to the global solution \((\hat{\beta}_{\mathrm{FE}}, \hat{\Lambda}_{\mathrm{FE}}, \hat{\Gamma}_{\mathrm{FE}})\). In particular, in Section 4, we provide a specific gradient descent algorithm and establish its convergence guarantees.9
Specifically, to demonstrate that our two-step estimator is (asymptotically) equivalent to the FE estimator, in Section 3, we show that, with probability approaching one, (i) the objective function \(\mathcal{L}_{NT} (\beta, \Lambda, \Gamma)\) is strictly convex in a shrinking neighborhood around the true values \((\beta_0, \Lambda_0, \Gamma_0)\), after normalization of the loadings and factors, and (ii) both the NNR and the FE estimators fall into this neighborhood. The technical difficulty here is that, since the dimension of the parameter space grows with \(N,T \rightarrow \infty\), the size of the local neighborhood, in which \(\mathcal{L}(\beta,\Lambda,\Gamma)\) remains convex, shrinks at a certain rate. To establish the desired result, we characterize this rate, and demonstrate that the NNR and the FE estimator have sufficiently fast rates of convergence to fall into that neighborhood with probability approaching one.
Since our two-step estimator is asymptotically equivalent to the FE estimator, it also follows the same asymptotic distribution previously derived by [1]. In particular, similar to the FE estimator, the two-step estimator of \(\beta_0\) also suffers from the incidental parameter bias caused by the necessity to estimate a large number of nuisance parameters. Implementation details for the analytical and split-panel jackknife bias correction methods are provided in Appendix 7.2. For a more detailed discussion of various bias correction and inference methods available for the FE estimator, we refer the reader to [1] and [27].
We provide additional practical implementation details in Section 4, including specific optimization algorithms used for both steps, as well as data=driven procedures for choosing the regularization
parameter \(\varphi_{NT}\) and the number of factors \(R\). An optimized implementation of our two-step estimator in R, including bias correction and calculation of standard errors, is also
readily available in our package NNRPanel.
In this section, we establish the consistency of the NNR estimator and derive its convergence rate. We also demonstrate that the original optimization problem 2 is locally convex in a (shrinking) neighborhood of the true values of the parameters. Combining these results, we argue that our two-step estimator is asymptotically equivalent to the FE estimator.
In this section, we establish the consistency of \((\hat{\beta}_{\mathrm{nuc}}, \hat{\Theta}_{\mathrm{nuc}})\) as in 4 , as well as of the associated nuisance estimators \((\hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) defined in 5 . To fix ideas, we will first focus on strictly exogenous covariates \(X_{it}\). We formally extend our analysis to more general settings allowing for predetermined covariates (e.g., the lagged outcomes) in Appendix 6.2.
For notational simplicity, we collect \(X_{it, d}\) into covariate matrices \(X_{d} \in \mathbb{R}^{N\times T}\) for each \(d= 1,2,\ldots, d_X\), and let \(X\) be the collection of all covariate matrices \(X= \{X_{1}, \ldots, X_{d_X}\}\). Whenever it does not cause confusion, for any \(Y^*\), we abbreviate \(\ell_{it}(Y^*) = \ell (Y_{it}\mid Y^* )\) to denote the log-likelihood evaluated at index \(Y^*\). We further denote the derivatives of the log-likelihood with respect to the index by \(\dot{\ell}_{it}(\cdot), \ddot{\ell}_{it}(\cdot), \ldots\). In addition, we use \(\mathbb{P}_{X, \Lambda_0, \Gamma_0}(\cdot) = \mathbb{P}(\cdot\mid X, \Lambda_0, \Gamma_0)\) and \(\mathbb{E}_{X, \Lambda_0, \Gamma_0}(\cdot) = \mathbb{E}(\cdot\mid X, \Lambda_0, \Gamma_0)\) to denote the conditional probability and expectation given \(X, \Lambda_0, \Gamma_0\), respectively.
We will maintain the following regularity conditions throughout this section.
Assumption 1 (Regularity conditions).
(Sampling) For all \(i = 1,2,\ldots, N\) and \(t=1,2,\ldots, T\), conditional on \((X, \Lambda_{0}, \Gamma_{0})\), \(\{Y_{it}\}_{1\leq i \leq N, 1\leq t \leq T}\) are independent across \(i\) and \(t\) and distributed as in 1 .
(Boundedness) The parameter spaces for \(\beta\), \(\lambda_i\), and \(\gamma_t\) are uniformly bounded for all \(i, t, N, T\). In addition, there exists a constant \(\rho_X>0\) such that \(\max_{d=1,\ldots, d_X}\|X_d\|_{\max} \leq \rho_X\) for all \(i, t, N, T\).
(Smoothness and convexity) \(-\ell_{it}(\cdot)\) is four times continuously differentiable. Furthermore, \(-\ell_{it}(\cdot)\) is strictly convex with \(0 < b_{\min} \leq -\ddot{\ell}_{it}( X_{it}' \beta + \lambda_i' \gamma_t) \leq b_{\max} <\infty\) almost surely for all \(\beta, \lambda_i, \gamma_t\) in the parameter space uniformly over \(i, t, N, T\).
(Strong factors) \(\frac{1}{N} \sum_{i=1}^{N}\lambda_{0, i}\lambda_{0, i}' \stackrel{p}{\longrightarrow}\Sigma_{\lambda}\) and \(\frac{1}{T} \sum_{t=1}^{T}\gamma_{0, t}\gamma_{0, t}'\stackrel{p}{\longrightarrow}\Sigma_{\gamma}\), where \(\Sigma_{\lambda} >0\) and \(\Sigma_{\gamma} >0\). In addition, the eigenvalues of \(\Sigma_{\lambda}\Sigma_{\gamma}\) are distinct.
(Generalized non-collinearity) For any \(\Gamma\in \mathbb{R}^{T\times R}\), let \(M_{\Lambda_0}\) and \(M_{\Gamma}\) be coprojection matrices of \(\Lambda_0\) and \(\Gamma\), respectively. The \(d_{X}\times d_{X}\) matrix \(D(\Gamma)\) with elements \[\begin{align} D(\Gamma)_{d_1, d_2} = \frac{1}{NT} \mathrm{Tr} \left(M_{\Lambda_0}X_{d_1} M_{\Gamma}X_{d_2}'\right), \quad d_{1}, d_{2} = 1, \ldots, d_X, \end{align}\] satisfies \(\inf_{\Gamma \in \mathbb{R}^{T\times R}} \sigma_{\min} (D(\Gamma) )>0\) wpa1.
Assumption 1[item:sampling] imposes the conditional independence of \(Y_{it}\) across \(i\) and \(t\). This aligns with the sampling assumption in [1] and is primarily applicable in contexts where \(X_{it}\) is strictly exogenous. It is also natural in network settings when the nature of interactions is primarily bilateral and the interdependencies of the outcomes are captured by the fixed effects. We consider a more general version of the sampling process allowing for predetermined covariates in Appendix 6.2.
Assumption 1[item:boundedness] imposes boundedness on the parameter spaces for \(\beta\), \(\Lambda\), and \(\Gamma\), as well as the boundedness on covariates. This assumption is widely adopted in the literature to derive concentration bounds (see, for example, [6], [28], and [7]). It is worth noting that [29] and [1] do not impose the boundedness of \(\lambda_i\), \(\gamma_t\), or \(X_{it}\) because their analyses focus on the local properties of the objective function. In contrast, we impose stronger assumptions on the parameter space to study global properties and establish global convergence guarantees for the NNR estimator.10
Assumption 1[item:smoothness] is commonly adopted in the nonlinear panel regression literature (see [29] and [1]) and is satisfied by Logit, Probit, and Poisson models.
Assumption 1[item:strong95factors] requires the factors to be strong. This condition is standard in the factor model literature.11 Additionally, the boundedness of the nuisance parameter space ensures that the maximum eigenvalues of \(\Sigma_{\lambda}\) and \(\Sigma_{\gamma}\) are bounded. The additional assumption that \(\Sigma_{\lambda}\Sigma_{\gamma}\) has distinct eigenvalues is not crucial and is introduced purely to simplify the discussion of technical aspects in the main text. In Appendix 6.5, we demonstrate that this assumption can be relaxed without affecting our main results.
Assumption 1[item:X95generalized95nonlinearity] is a generalized non-collinearity condition that rules out covariates that do not display variation in the individual and time dimensions, such as time-invariant or individual-invariant regressors. This condition is identical to Assumption 1(vii) in [1], and we the reader to that paper for further discussion. In fact, we do not use this assumption in our proofs but we impose it here in order to be able to invoke the results of [1] characterizing the asymptotic properties of the FE estimator.
We now turn to a key condition for establishing the consistency of our NNR estimator, the restricted strong convexity (RSC) condition.12 As in models with a fixed number of parameters, establishing error bounds for NNR estimators requires analyzing the Hessian of the objective function \(\mathcal{L}(\beta, \Theta)\) and demonstrating that it is positive definite, implying that the objective function is strongly convex. However, in the context of problem 4 , it is not possible for the Hessian matrix to be positive definite, as the number of parameters grows with \(N\) and \(T\) and exceeds the number of observations. Nonetheless, thanks to the nuclear norm penalty in 4 , it is sufficient to require \(\mathcal{L}(\beta, \Theta)\) to be strictly convex only over a restricted part of the parameter space in which \(\Theta\) is approximately low-rank.
To elaborate, for any \((\beta, \Theta)\), the second-order Taylor remainder of the objective function \(\mathcal{L}(\beta, \Theta)\) around the true parameter \((\beta_0, \Theta_0)\) satisfies \[\label{eqref:second95order95convexity} \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} \frac{-\ddot{\ell}_{it}(X_{it}'\tilde{\beta} + \tilde{\theta}_{it})}{2} (X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 \geq \frac{b_{\min}}{2} \underbrace{\frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} (X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2}_{\mathcal{E}_{NT} (\Delta_{\beta}, \Delta_{\Theta}) },\tag{6}\] where \((\tilde{\beta}, \tilde{\Theta})\) lies between \((\beta, \Theta)\) and \((\beta_0, \Theta_0)\), \(\Delta_{\beta} = \beta - \beta_0\), \(\Delta_{\theta_{it}} = \theta_{it} - \theta_{0,it}\), and the inequality follows from the fact that the second-order derivative \(-\ddot{\ell}_{it}(\cdot)\) is uniformly bounded below by \(b_{\min}\) (Assumption 1[item:smoothness]). Thus, it suffices to study the properties of \(\mathcal{E}_{NT}(\Delta_{\beta}, \Delta_{\Theta})\) to ensure that the objective function is strongly convex over a restricted parameter space. Specifically, we introduce the following assumption.
Assumption 2 (Restricted strong convexity (RSC)). For any \(c_0 > 0\), there exist constants \({\kappa, \eta >0}\), independent of \(N, T\), such that for any \((\Delta_{\beta}, \Delta_{\Theta})\in \mathcal{C}_1 \cap \mathcal{C}_2\), \[\begin{align} \label{eq:RSC} \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 \geq \kappa \left(\|\Delta_{\beta}\|^2 + \frac{1}{NT}\|\Delta_{\Theta}\|_{\mathrm{F}}^2\right) - \eta \frac{N+T}{NT} (\log(NT))^2, \quad \text{wpa1}, \end{align}\qquad{(2)}\] where \[\begin{gather} \mathcal{C}_1 := \left\{(\Delta_{\beta}, \Delta_{\Theta})\in (\mathbb{R}^{d_X}\times \mathbb{R}^{N\times T})\mid \|M_{\Lambda_0}\Delta_{\Theta}M_{\Gamma_0}\|_{\mathrm{nuc}} \leq c_0 (\sqrt{NT}\|\Delta_{\beta}\| + \|\Delta_{\Theta} - M_{\Lambda_0}\Delta_{\Theta}M_{\Gamma_0}\|_{\mathrm{nuc}})\right\}, \\ \mathcal{C}_2 := \left\{ (\Delta_{\beta}, \Delta_{\Theta})\in (\mathbb{R}^{d_X}\times \mathbb{R}^{N\times T})\mid \|\Delta_{\beta}\|^2 + \frac{1}{NT} \|\Delta_{\Theta}\|_{\mathrm{F}}^2 \geq \sqrt{\frac{\log (NT)}{NT}}\right\}. \end{gather}\]
Assumption 2 controls the behavior of \(\mathcal{E}_{NT} (\Delta_{\beta}, \Delta_{\Theta})\) defined in 6 over the restricted parameter space of interest \(\mathcal{C}_1\cap \mathcal{C}_2\).13 \(\mathcal{C}_1\) can be viewed as an approximately low-rank space. The term \({\Delta_{\Theta} - M_{\Lambda_0}\Delta_{\Theta}M_{\Gamma_0}}\) on the right-hand side represents the component that can be explained by \(\Lambda_0\) and \(\Gamma_0\), serving as a low-rank approximation. In contrast, the left-hand side, \(M_{\Lambda_0}\Delta_{\Theta}M_{\Gamma_0}\), corresponds to the residual of \(\Delta_{\Theta}\) that cannot be explained by \(\Lambda_0\) and \(\Gamma_0\), interpreted as the low-rank approximation residual. Therefore, \(\mathcal{C}_1\) consists of matrices whose low-rank approximation residuals (in terms of nuclear norm) are small compared to their low-rank approximation (along with the estimation error of \(\beta\)). Importantly, it is sufficient to control the behavior of \(\mathcal{E}_{NT} (\Delta_{\beta}, \Delta_{\Theta})\) over \(\mathcal{C}_1\) because the appropriately chosen nuclear norm penalty in 4 enforces \((\Delta \hat{\beta}_{\mathrm{nuc}}, \Delta \hat{\Theta}_{\mathrm{nuc}}) \in \mathcal{C}_1\) for some \(c_0\) wpa1.14
The set \(\mathcal{C}_2\) is introduced to restrict our attention to scenarios of primary interest. Since we can directly obtain the bounds \(\|\Delta_{\beta}\|_{2}^2 \leq \sqrt{\frac{\log (NT)}{NT}}\) and \(\frac{1}{NT} \|\Delta_{\Theta}\|_F^2 \leq \sqrt{\frac{\log (NT)}{NT}}\) for matrices that do not belong to this space, focusing on \(\mathcal{C}_2\) simplifies the analysis without affecting the results of this section. Notice that inequality ?? also introduces an additional tolerance term \(\eta \frac{N+T}{NT} (\log(NT))^2\) to account for the randomness in \(X_{it}\). Combined with 6 , ?? ensures that \(\mathcal{L}(\beta, \Theta)\) is strongly convex (up to the tolerance term) over the restricted space \(\mathcal{C}_1 \cap \mathcal{C}_2\), even in the high-dimensional setting.
While various versions of the RSC condition are routinely employed in the low-rank estimation literature (e.g., [5]–[7]), they are typically introduced as high-level assumptions and can be challenging to verify. To address this issue, [6] provided sufficient conditions for verifying RSC in panels with strictly exogenous covariates. In this paper, we extend the related results of [6] and provide a set of easily verifiable primitive conditions sufficient for the RSC to hold in settings with predetermined \(X_{it}\), including the outcome’s lags. Since predetermined regressors are ubiquitous in applied work, we believe that this result, formalized by the lemma below, is an important technical contribution of the paper, broadening the applicability of low-rank estimators and enabling empirical researchers to adopt them with greater confidence.
Lemma 1 (Sufficient conditions for RSC). Under Assumptions 4 and 6 provided in Appendix 6, Assumption 2 is satisfied.
To streamline the exposition, we will still state our main results of this section using Assumption 2 and defer the discussion of Lemma 1 and the related primitive conditions to Appendix 6.3.
For notational simplicity, we impose the sign constraint (e.g., fixing the sign of each factor) and the normalization constraint on the nuisance parameters \((\Lambda_0, \Gamma_0)\) such that \(\Lambda_0'\Lambda_0/N\) and \(\Gamma_0'\Gamma_0/T\) are diagonal, with \(\Lambda_0'\Lambda_0/N = \Gamma_0'\Gamma_0/T\). These constraints are consistent with the construction of \((\hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) in 5 . The feasibility of these constraints is provided in Appendix 6.1.
The following theorem establishes the convergence rates of the NNR estimators.
Theorem 1. For any \(\alpha>0\) such that \(\varphi_{NT} \geq (1+\alpha) \max\{ \|\nabla_{\beta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)\|_{2}, \sqrt{NT}\|\nabla_{\Theta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)\|_{\mathrm{op}} \}\), under Assumptions 1 and 2, there exist constants \(c_1, c_2>0\) that do not depend on \(N, T\) such that, as \(N, T\rightarrow\infty\), wpa1, \[\begin{align} \| \hat{\beta}_{\mathrm{nuc}} - \beta_0\| & \leq c_1 \left(\varphi_{NT} + \log (NT)/\sqrt{\min\{N, T\}}\right), \\ \frac{1}{\sqrt{NT}}\|\hat{\Theta}_{\mathrm{nuc}} - \Theta_0\|_{\mathrm{F}} & \leq c_1 \left(\varphi_{NT} + \log (NT)/\sqrt{\min\{N, T\}}\right). \end{align}\] In addition, wpa1, \[\begin{align} \frac{1}{\sqrt{N}}\|\hat{\Lambda}_{\mathrm{nuc}} - \Lambda_0 \|_{\mathrm{F}} & \leq c_2 \left(\varphi_{NT} + \log (NT)/\sqrt{\min\{N, T\}} \right), \\ \frac{1}{\sqrt{T}}\|\hat{\Gamma}_{\mathrm{nuc}} - \Gamma_0 \|_{\mathrm{F}} & \leq c_2 \left(\varphi_{NT} + \log (NT)/\sqrt{\min\{N, T\}} \right). \end{align}\]
Theorem 1 establishes that, for a sufficiently large regularization parameter \(\varphi_{NT}\), the NNR estimator \(\hat{\beta}_{\mathrm{nuc}}\) converges to \(\beta_0\) at a rate of order \(\varphi_{NT} + \log(NT)/\sqrt{\min\{N,T\}}\). This convergence rate coincides with that obtained by [5] for the linear panel model. The estimator \(\hat{\Theta}_{\mathrm{nuc}}\) achieves the same convergence rate under the normalized Frobenius norm, and the associated estimators of the latent factors \((\hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) inherit this rate as well.
The convergence rates in Theorem 1 depend on the choice of the regularization parameter \(\varphi_{NT}\). While \(\varphi_{NT}\) must be sufficiently large to guarantee consistency, choosing \(\varphi_{NT}\) larger than \(O\left(\log (NT)/ \sqrt{ \min\{N, T\}}\right)\) would lead to slower rates of convergence. To further investigate this issue, we demonstrate that choosing \(\varphi_{NT} = O\left(\log (NT)/ \sqrt{ \min\{N, T\}}\right)\) is compatible with the restriction imposed by Theorem 1, and thus achieves the fastest rate implied by the theorem. This result is formalized by the following corollary.
Corollary 1. Suppose that the hypotheses of Theorem 1 hold, and let \(\varphi_{NT} = O\left(\log (NT)/ \sqrt{ \min\{N, T\}}\right)\). Then there exist constants \(c_3, c_4>0\) that do not depend on \(N, T\) such that, as \(N, T\rightarrow\infty\), wpa1, \[\begin{align} \| \hat{\beta}_{\mathrm{nuc}} - \beta_0\| & \leq c_3 \log (NT)/\sqrt{\min\{N, T\}}, \\ \frac{1}{\sqrt{NT}}\|\hat{\Theta}_{\mathrm{nuc}} - \Theta_0\|_{\mathrm{F}} & \leq c_3 \log (NT)/\sqrt{\min\{N, T\}}. \end{align}\] In addition, wpa1, \[\begin{align} \frac{1}{\sqrt{N}}\|\hat{\Lambda}_{\mathrm{nuc}} - \Lambda_0 \|_{\mathrm{F}} & \leq c_4 \log (NT)/\sqrt{\min\{N, T\}}, \\ \frac{1}{\sqrt{T}}\|\hat{\Gamma}_{\mathrm{nuc}} - \Gamma_0 \|_{\mathrm{F}} & \leq c_4 \log (NT)/\sqrt{\min\{N, T\}}. \end{align}\]
Corollary 1 establishes the convergence rate for the studied NNR estimator under the optimal choice of \(\varphi_{NT}\). It is instructive to compare this result with the existing rates for NNR estimators previously obtained in the literature for different settings. In particular, up to an additional logarithmic factor, our rate coincides with the rates established by [5], [6], and [18] for linear panel models. In nonlinear models, [7] established the same rate in a network formation model. Importantly, Corollary 1 improves on [5], which extends their original analysis to single-index models. As we explain in the following section, this improvement in the convergence rate is crucial for demonstrating that the first-step estimator \((\hat{\beta}_{\mathrm{nuc}}, \hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) is sufficiently close to the true values of the parameters, so the second-step optimization problem initialized at the NNR estimator is (locally) convex.
In this section, we establish the asymptotic equivalence between our two-step estimator and the FE estimator. This result builds on the consistency of the NNR estimator and the local convexity of the objective function \(\mathcal{L}_{NT}(\beta, \Lambda, \Gamma)\) around the true values of the parameters.
Since the NNR estimator \((\hat{\beta}_{\mathrm{nuc}}, \hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) is used to initialize the second-step optimization problem 2 , its consistency ensures that the starting values lie within a shrinking neighborhood of the true parameters as \(N, T \rightarrow \infty\). Thus, to establish the desired result, it suffices to focus on studying the local properties of \(\mathcal{L}_{NT}(\beta, \Lambda, \Gamma)\) within these shrinking neighborhoods. To formalize the argument, let \(\{\delta_{NT}\}\) be a sequence of shrinking radii such that \(\delta_{NT}\rightarrow 0\) as \(N, T\rightarrow \infty\), and define the shrinking neighborhood around the true parameters \((\beta_0, \Lambda_0, \Gamma_0)\) as follows: \[\begin{align} \label{eq:neighborhood} \mathcal{B}_{\delta_{NT}} := \bigg\{(\beta, \Lambda, \Gamma) \mid & \|\beta-\beta_0\|, \frac{1}{\sqrt{N}}\|\Lambda - {\Lambda}_0 \|_{\mathrm{F}} , \frac{1}{\sqrt{T}}\|\Gamma - {\Gamma}_0 \|_{\mathrm{F}} \leq \delta_{NT} \bigg\}. \end{align}\tag{7}\] The neighborhood \(\mathcal{B}_{\delta_{NT}}\) consists of parameters whose distances from the true values \((\beta_0, \Lambda_0, \Gamma_0)\) are less than \(\delta_{NT}\).
We establish the desired asymptotic equivalence result by demonstrating that with a properly chosen \(\delta_{NT}\), the following conditions hold: (1) the NNR estimator falls within the shrinking neighborhood, i.e., \((\hat{\beta}_{\mathrm{nuc}}, \hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\in \mathcal{B}_{\delta_{NT}}\) wpa1, (2) the FE estimator, as the global minimizer of problem 2 , also lies within the shrinking neighborhood (up to rotation) wpa1, and (3) the objective function \(\mathcal{L}_{NT}(\beta, \Lambda, \Gamma)\) is strictly convex within the neighborhood \(\mathcal{B}_{\delta_{NT}}\). Together, these conditions ensure that, wpa1, (1) any locally convergent solver initialized at the NNR estimator is guaranteed to find the global minimum of problem 2 , and, consequently, (2) our two-step estimator is asymptotically equivalent to the FE estimator.
Let \(\delta_{NT} = \log(NT) \min\{N^{-3/8}, T^{-3/8}\}\). The NNR estimator falls within the corresponding shrinking neighborhood \(\mathcal{B}_{\delta_{NT}}\) wpa1 by Corollary 1 (provided that \(\varphi_{NT}\) is chosen appropriately). Similarly, the FE estimator also lies within \(\mathcal{B}_{\delta_{NT}}\) wpa1, as shown in Lemma 1 in [1]. Hence, to establish the desired equivalence result, it would be sufficient to establish local convexity of the original objective function \(\mathcal{L}_{NT}(\beta, \Lambda, \Gamma)\) within \(\mathcal{B}_{\delta_{NT}}\), which will be the main focus of this section.
Specifically, consider the following local optimization problem \[\label{eq:definition95local} \begin{align} (\hat{\beta}_{\mathrm{local}}, \hat{\Lambda}_{\mathrm{local}}, \hat{\Gamma}_{\mathrm{local}}) \in \mathop{\mathrm{argmin}}_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}} } \mathcal{L}_{NT}(\beta, \Lambda, \Gamma), \end{align}\tag{8}\] where the parameter space is restricted to the shrinking neighborhood \(\mathcal{B}_{\delta_{NT}}\). Notice that, as in the linear case, this problem does not have a unique solution for \(\hat{\Lambda}_{\mathrm{local}}\) and \(\hat{\Gamma}_{\mathrm{local}}\) and requires imposing \(R^2\) additional constraints to identify these parameters. Building on the normalization proposed in [1], we introduce the following restricted parameter set \[\begin{align} \Phi_{NT} = \left\{(\Lambda, \Gamma)\mid \hat{\Lambda}_{\mathrm{nuc}}' \Lambda/N = \Gamma' \hat{\Gamma}_{\mathrm{nuc}} /T \right\} \end{align}\] consistent with the construction in 5 . While in practice researchers could employ alternative normalizations, focusing on \(\Phi_{NT}\) greatly facilitates the following theoretical analysis thanks to the linearity of the imposed restriction. Specifically, instead of imposing \(\Phi_{NT}\) directly, we convert it into a quadratic penalization term, \(\| \hat{\Lambda}_{\mathrm{nuc}}' \Lambda/N - \Gamma' \hat{\Gamma}_{\mathrm{nuc}} /T \|_{\mathrm{F}}^2\), whose Hessian matrix only depends on \((\hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) due to the aforementioned linearity. In particular, consider the associated (penalized) optimization problem: \[\label{eq:local95estimators} \begin{align} (\hat{\beta}_{\mathrm{local}}, \hat{\Lambda}_{\mathrm{local}}, \hat{\Gamma}_{\mathrm{local}}) = \mathop{\mathrm{argmin}}_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}} } \left\{\mathcal{L}_{NT}(\beta, \Lambda, \Gamma) + \frac{1}{2} \| \hat{\Lambda}_{\mathrm{nuc}}' \Lambda/N - \Gamma' \hat{\Gamma}_{\mathrm{nuc}} /T \|_{\mathrm{F}}^2 \right\}, \end{align}\tag{9}\] which is equivalent to solving 8 with the normalization constraint \((\Lambda, \Gamma)\in\Phi_{NT}\). The differentiability of the penalty simplifies the analysis of the associated (penalized) Hessian, providing a more tractable alternative to directly studying 8 under the hard constraint.
To demonstrate that problem 9 is convex, we need to inspect the sample Hessian of its objective function with respect to \((\beta, \Lambda, \Gamma)\) given by \[\begin{align} \mathcal{H}_{NT}(\beta, \Lambda, \Gamma) := \nabla^2 \mathcal{L}_{NT}(\beta, \Lambda, \Gamma) + \frac{1}{2} \nabla^2 \| \hat{\Lambda}_{\mathrm{nuc}}' \Lambda/N - \Gamma' \hat{\Gamma}_{\mathrm{nuc}} /T \|_{\mathrm{F}}^2. \end{align}\] It is a \((d_X + R(N+T))\)-dimensional square matrix.15 Establishing local convexity is equivalent to showing that \(\mathcal{H}_{NT}(\beta, \Lambda, \Gamma)\) is positive definite uniformly over \((\beta, \Lambda, \Gamma)\in \mathcal{B}_{\delta_{NT}}\). Consider the following decomposition: \[\begin{align} \label{eq:hessian32decomposition32main} \mathcal{H}_{NT}(\beta, \Lambda, \Gamma) = \underbrace{\mathbb{E}_{X, \Lambda_0, \Gamma_0}\mathcal{H}_{NT}(\beta_0, \Lambda_0, \Gamma_0)}_{\text{population Hessian at true parameters}} + \underbrace{\mathcal{H}_{NT}(\beta, \Lambda, \Gamma)- \mathbb{E}_{X, \Lambda_0, \Gamma_0}\mathcal{H}_{NT}(\beta_0, \Lambda_0, \Gamma_0)}_{\text{deviation}}. \end{align}\tag{10}\] The first term represents the population Hessian evaluated at the true parameters \((\beta_0, \Lambda_0, \Gamma_0)\), whose smallest eigenvalue is strictly positive under general conditions. The second term captures the deviation arising from the sampling error and the perturbations of \((\beta, \Lambda, \Gamma)\) from the true parameters. Thus, by Weyl’s theorem, we can demonstrate that problem 9 is locally convex provided that the deviation term is negligible (in terms of the operator norm) compared to \(\mathbb{E}_{X, \Lambda_0, \Gamma_0}\mathcal{H}_{NT}(\beta_0, \Lambda_0, \Gamma_0)\).
Following the literature, we control the behavior of the population Hessian using the condition below.
Assumption 3 (Diagonal structure). There exists a constant \(C>0\) (that does not depend on \(N, T\)) such that \[\begin{align} \mathbb{E}_{X, \Lambda_0, \Gamma_0}\mathcal{H}_{NT}(\beta_0, \Lambda_0, \Gamma_0) \geq C \mathrm{diag}\left\{\mathbb{I}_{d_X}, \frac{1}{N}\mathbb{I}_{N R}, \frac{1}{T}\mathbb{I}_{TR}\right\}. \end{align}\]
Assumption 3 is closely related to the asymptotic diagonal structure condition in [1], [33], and [18]. Notice that Assumption 3 is mild since it imposes conditions only on the population Hessian evaluated at the true parameters and does not restrict the behavior of the sample Hessian \(\mathcal{H}_{NT}(\beta, \Lambda, \Gamma)\) directly. Easily verifiable sufficient conditions for Assumption 3 are provided in Lemma 3 in Appendix 6.4.
Equipped with Assumption 3, we are now ready to inspect the behavior of the sample Hessian \(\mathcal{H}_{NT}(\beta, \Lambda, \Gamma)\) using the decomposition in 10 . Specifically, to establish the desired result, we demonstrate that the contribution of the deviation term in 10 is asymptotically negligible uniformly over \((\beta, \Lambda, \Gamma)\in \mathcal{B}_{\delta_{NT}}\). This result is formalized by the following theorem.
Theorem 2. Suppose that the hypotheses of Corollary 1 are satisfied, and that \(N\) and \(T\) have the same order. Let \(\delta_{NT}= \log (NT) \min\{N^{-3/8}, T^{-3/8}\}\), and \(\mathcal{B}_{\delta_{NT}}\) be the neighborhood defined in 7 . Then, under Assumption 3, the following results hold:
the local optimization problem 9 is strictly convex wpa1;
our two-step estimator is asymptotically equivalent to the FE estimator;
the local optimization problem 9 is strongly convex wpa1 uniformly over \(\mathcal{B}_{\delta_{NT}}\), i.e., there exists a constant \(c_5>0\) independent of \(N, T\) such that for any \((\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}\), \[\begin{align} \mathcal{H}_{NT}(\beta, \Lambda, \Gamma) > c_5 \mathrm{diag}\left\{\mathbb{I}_{d_X}, \frac{1}{N}\mathbb{I}_{NR}, \frac{1}{T}\mathbb{I}_{TR}\right\}, \quad \text{wpa1}. \end{align}\]
Since both the NNR and the FE estimators fall into \(\mathcal{B}_{\delta_{NT}}\) wpa1, Theorem 2[item:1] guarantees that any locally convergent solver initialized at \((\hat{\beta}_{\mathrm{nuc}}, \hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) finds the global minimum, and thus it formally demonstrates that our two-step estimator is asymptotically equivalent to the FE estimator. Moreover, Theorem 2[item:3] further establishes strong local convexity of 9 , implying that simple gradient descent can be effectively applied in the second step to find the global minimum, even in high-dimensional settings.
Remark 2. [5] and [18] obtained analogous asymptotic results for their post-NNR estimators in linear panel models by inspecting the second-step optimization problem and establishing its local convexity. Thus, Theorem 2 generalizes their analysis to nonlinear settings. However, this extension is technically nontrivial and requires a more elaborate treatment. In particular, in linear panel models, the original objective function is locally convex in \(\mathcal{B}_{\delta_{NT}}\) whenever \(\delta_{NT}=o(1)\), and thus consistency of the NNR estimator alone is sufficient to guarantee convergence of the second-step optimization to the global minimum. In contrast, establishing local convexity is more delicate in the studied nonlinear setting, requiring the neighborhood to shrink at a faster rate than in the linear cases. Therefore, unless the preliminary estimator has a sufficiently fast convergence rate to fall into the shrinking convexity region wpa1, the convergence of the second-step local optimization to the global solution cannot be guaranteed. In particular, it turns out that the convergence rate obtained by [5] for the NNR estimator for single-index models is not sufficiently fast to satisfy this requirement.
In this section, we provide practical implementation details for the proposed method, including specific optimization algorithms for both estimation steps (along with their convergence guarantees), as well as data-dependent procedures for selecting the
regularization parameter \(\varphi_{NT}\) and determining the number of factors. The specific algorithms and procedures described in this section are also implemented in our R package NNRPanel.
We compute the NNR estimator defined in 4 using the proximal gradient descent method following [26]. Specifically, given the k-step estimates \((\beta^{(k)}, \Theta^{(k)})\), the \(k+1\)-step estimates are updated by solving \[\begin{align} \beta^{(k+1)}, \Theta^{(k+1)} \in \arg\min_{\beta, \Theta} \Big\{& \mathcal{L}_{NT}(\beta^{(k)}, \Theta^{(k)}) + \langle\nabla_{\beta}\mathcal{L}_{NT}(\beta^{(k)}, \Theta^{(k)}), \beta - \beta^{(k)}\rangle+ \langle\nabla_{\Theta}\mathcal{\mathcal{L}}_{NT}(\beta^{(k)}, \Theta^{(k)}), \Theta - \Theta^{(k)} \rangle\\ & + \frac{1}{2s_{\beta}}\|\beta - \beta^{(k)}\|^2 + \frac{1}{2s_{\theta}} \|\Theta - \Theta^{(k)}\|_{\mathrm{F}}^2 + \frac{\varphi_{NT}}{\sqrt{NT}}\|\Theta\|_{\mathrm{nuc}}\Big\}, \end{align}\] where \(\nabla_{\beta}\mathcal{L}_{NT}(\cdot, \cdot)\) is a \(d_X\)-dimensional vector of gradients with respect to \(\beta\), \(\nabla_{\Theta}\mathcal{L}_{NT}(\cdot, \cdot)\in \mathbb{R}^{N\times T}\) is a matrix of gradients with respect to \(\theta_{it}\), \(\langle\cdot, \cdot\rangle\) denotes the inner product between two vectors or two matrices, and \(s_{\beta}, s_{\theta}>0\) are step sizes.
Importantly, the solution to the optimization problem is available in closed form. In particular, let \({\mathcal{S}^*_{s_{\theta}\frac{\varphi_{NT}}{\sqrt{NT}}}: \mathbb{R}^{N\times T}\mapsto \mathbb{R}^{N\times T}}\) denote the soft-thresholding operator applied to the singular values of an \(N \times T\) matrix, where \(s_{\theta}\frac{\varphi_{NT}}{\sqrt{NT}}\) is the threshold value. Specifically, for any matrix \(A\in \mathbb{R}^{N\times T}\) with singular value decomposition \(A = U\Sigma V'\), let \[\mathcal{S}^*_{s_{\theta}\frac{\varphi_{NT}}{\sqrt{NT}}}(A) = U\mathrm{diag}\left\{\max\left\{\Sigma_{rr} - s_{\theta}\frac{\varphi_{NT}}{\sqrt{NT}}, 0\right\}_{r=1,\ldots, \min\{N, T\}}\right\}V'.\] Then, the NNR estimator in 4 can be obtained using the following algorithm.
Figure 1:
.
We establish the convergence of Algorithm 1 using a proof strategy similar to the one provided in [34].
Theorem 3. Suppose that the hypotheses of Theorem 1 are satisfied. Then, Algorithm 1 is guaranteed to converge to the global minimizer if \(0 < s_{\beta} < \frac{1}{L_{\beta}}\) and \(0 < \frac{s_{\theta}}{NT} < \frac{1}{L_{\theta} }\), where \(L_{\beta} = 2 d_X b_{\max} \rho_X^2\), \(L_{\theta} = 2b_{\max}\), and \({\rho_X = \max_{1\leq d\leq d_X} \|X_d\|_{\max}}\).
The theorem establishes the algorithm’s convergence to the global minimizer provided that \((s_{\beta}, s_{\theta}/(NT))\) is sufficiently small. Notice that the step sizes scale differently: \(s_{\beta}\) and \(s_{\theta}/(NT)\) should be of the same order, reflecting their distinct influence on the objective function. This difference arises because a change in \(\beta\) affects \(\ell_{it}\) for all \((i,t)\), whereas a change in \(\theta_{it}\) affects only the corresponding \(\ell_{it}\).
To apply Algorithm 1 in practice, one needs to specify the initial values \((\beta^{(0)},\Theta^{(0)})\) and step sizes \((s_\beta, s_\theta)\). Since the optimization problem is convex, the algorithm converges to a global minimizer regardless of the choice of its initial values. Thus, researchers may simply initialize with \(\beta^{(0)} = 0\) and \(\Theta^{(0)} = 0\). While Theorem 3 characterizes how small \(s_\beta\) and \(s_\theta\) should be to guarantee the convergence, it offers limited practical guidance, since \(b_{\max}\) is typically unknown. In practice, we recommend starting each iteration with step sizes \(s_{\beta} = 1\) and \(s_{\theta} = NT\). If the objective function increases, the step sizes are iteratively halved until a decrease in the objective function is achieved.
Finally, we note that solving the nuclear norm-regularized problem is usually more computationally demanding than the second step. The computational bottleneck is computing the singular value decomposition of an \(N \times T\) matrix at each iteration. Our R package implementation remains computationally efficient even for matrices of size \(N, T = 1000\). In principle, to speed up the computation for much larger values of \(N\) or \(T\), it is also possible to employ the accelerated proximal gradient method (see, e.g., [34]), but we do not explore this direction in this paper.
In the second step, we employ gradient descent to search for a minimizer of 9 . The following algorithm practically minimizes \(\mathcal{L}_{NT}(\beta, \Lambda, \Gamma)\) using the NNR estimator as the initial value. Step 3 is introduced to address the rotational invariance of the loadings and factors. It is optional since the penalization introduced in 9 is only used to facilitate the theoretical analysis.
Figure 2:
.
The following theorem provides conditions under which Algorithm 2 is guaranteed to converge.
Theorem 4. Suppose that the hypotheses of Theorem 2 are satisfied. Then, Algorithm 2 is guaranteed to converge to a global minimizer when \(0 < s_{\beta} < \frac{1}{L_{\beta}}\), \(0 < \frac{s_{\lambda}}{N} < \frac{1}{L_{\lambda} }\), and \(0 < \frac{s_{\gamma}}{T} < \frac{1}{L_{\gamma} }\), where \(L_{\beta}, L_{\lambda}, L_{\gamma}\) are sufficiently large constants independent of \(N, T\).
Similarly to the analogous requirement of Theorem 3, to guarantee the convergence, the step sizes need to be scaled differently. Specifically, we require \(s_{\beta}\sim s_{\lambda}/N \sim s_{\gamma}/T\), reflecting their respective influence on the objective function. In practice, in each step, we recommend starting with \(s_{\beta} = 1\), \(s_{\lambda} = N\), and \(s_{\gamma} = T\). If the objective function increases, we iteratively halve the step sizes \((s_{\beta}, s_{\lambda}, s_{\gamma})\) until the objective function decreases.
Remark 3. [1] propose using an EM algorithm, a Newton-Raphson-type method that theoretically achieves faster convergence through second-order accuracy. Despite this potential advantage, implementation of their method requires calculating and inverting a high-dimensional Hessian matrix, which can computationally expensive but also numerically unstable in the studied nonlinear setting. Thus, we adopt a more robust, albeit slower, gradient-based algorithm.
In this section, we propose a data-driven procedure for selecting the regularization parameter \(\varphi_{NT}\) and determining the number of factors \(R\). Recall that Theorem 1 requires \(\varphi_{NT} > \sqrt{NT}\|\nabla_{\Theta}\mathcal{L}_{NT}(\beta_0, \Theta_0)\|_{\mathrm{op}}\) in order to achieve the desired consistency result.16 At the same time, Theorem 1 also indicates that selecting an excessively large \(\varphi_{NT}\) should be avoided, as it may induce substantial estimation error in the NNR estimator. Hence, a preferable choice for \(\varphi_{NT}\) is one that slightly exceeds \(\sqrt{NT}\|\nabla_{\Theta}\mathcal{L}_{NT}(\beta_0, \Theta_0)\|_{\mathrm{op}}\). This quantity, however, is not immediately available since \(\beta_0\) and \(\Theta_0\) are unknown. To address this, we propose the following three-step procedure for jointly selecting \(\varphi_{NT}\) and \(R\), building on the idea of [6].
Figure 3:
.
Algorithm 3 is computationally simple to implement. Step 1 and Step 2 require solving convex optimization problems, for which efficient algorithms are readily available. Importantly, our procedure avoids using cross-validation for selecting \(\varphi_{NT}\), which substantially reduces the computational cost.
We investigate the performance of Algorithm 3 in numerical experiments and document its consistently good performance for different specifications with sample sizes ranging from \((N,T) =
(50,40)\) to \((N,T) = (1000,200)\). This routine is already implemented in our R package NNRPanel and only requires the user to specify an ex-ante upper bound on the number of factors \(R_{\max}\) as an input, which is a standard requirement even in linear factor models.
In this section, we investigate the finite sample properties of our method in numerical experiments and revisit the empirical application studied in [1].
In this section, we evaluate the performance of our estimator in canonical binary response panel models. For brevity, here we will present and discuss results for specifications with strictly exogenous covariates (conditional on the fixed effects). Additional numerical results for dynamic specifications with predetermined covariates are provided and discussed in Appendix 7.3.
In the considered numerical experiments, the data are generated from the following binary response model with \(R = 2\): \[\label{eq:logit95static} \begin{align} Y_{it} & = \boldsymbol{1}\left(\beta_{1}X_{it} + \lambda_{i}' \gamma_{t} + \epsilon_{Y, it}>0\right), \\ X_{it} & = \lambda_{i}' \gamma_{t} + \lambda_{i}' \iota + \gamma_{ t}' \iota + \lambda_{X, i} \gamma_{X, t} + \epsilon_{X, it}. \end{align}\tag{11}\] The error terms \(\{\epsilon_{Y, it}\}_{1\leq i\leq N, 1\leq t\leq T}\) are i.i.d. over both dimensions. We consider two canonical designs: (i) Probit, where \(\epsilon_{Y, it}\) follows the standard normal distribution \(N(0, 1)\); and (ii) Logit, where \(\epsilon_{Y, it}\) follows the standard logistic distribution.
In both designs, \(\lambda_{i} = (\lambda_{i1}, \lambda_{i2})'\), \(\gamma_{t} = (\gamma_{t1}, \gamma_{t2})'\). \(\{ \lambda_{ir}\}_{ 1\leq i\leq N, r=1,2}\), \(\{ \gamma_{tr}\}_{1\leq t\leq T, r=1,2}\), \(\{ \lambda_{X, i}\}_{ 1\leq i\leq N}\), and \(\{ \gamma_{X, t}\}_{ 1\leq t\leq T}\) consist of independent random variables drawn from the standard normal distribution \(N(0, 1)\). In addition, these collections are mutually independent. For both designs, \(\{\epsilon_{X, it}\}_{1\leq i\leq N, 1\leq t\leq T}\) are i.i.d. (over both dimensions) draws from \(N(0, 4)\), and \(\beta_1 = 0.2\). We vary the sample size from \((N,T) = (50,40)\) to \((N,T) = (1000,200)\) and perform 1000 replications for each design.
We report results for two versions of our two-step estimator. The first is implemented using the correct number of factors \(R = 2\) (\(\mathrm{TS}^*\)), whereas the second involves estimating the number of factors (\(\mathrm{TS}\)). For both of these estimators, we use Algorithm 3 to choose \(\varphi_{NT}\) (setting \(\alpha = 0.05\)), and we specify \(R_{\max} = 5\) when we also need to estimate the number of factors.17 For both of these estimators, we also report results for their analytical and jackknife bias-corrected versions (\(\mathrm{ABC}^*\) and \(\mathrm{JBC}^*\) for \(\mathrm{TS}^*\), and \(\mathrm{ABC}\) and \(\mathrm{JBC}\) for \(\mathrm{TS}\)).18 We also report results for the naive pooled estimator that ignores individual and time latent factors (POOL) and for our first-step nuclear norm-regularized estimator (NNR).
Table 1 reports the results for the Probit specification. As expected, the pooled estimator (POOL) has a substantial bias that does not diminish as the sample size increases, whereas the bias of the NNR estimator decreases slowly towards zero as \(N, T \to \infty\), supporting the findings of Theorem 1. The proposed two-step estimator using the true number of factors (\(\mathrm{TS}^*\)) substantially improves upon the NNR estimator: its bias is smaller and converges rapidly to zero as \(N\) and \(T\) increase. Nevertheless, due to the incidental parameter problem, the bias and standard deviation of the \(\mathrm{TS}^*\) estimator remain of the same order even for large sample sizes, such as \((N, T) = (1000, 200)\). The analytical bias correction method effectively addresses this problem even in smaller samples, e.g., for \((N, T) = (50, 40)\). While both the analytical and jackknife correction methods effectively reduce the bias in larger sample sizes, the former might still be preferred to the latter because it does not inflate the (asymptotic) variance of the estimator. Finally, when the number of factors is estimated rather than known, the corresponding estimators (\(\mathrm{TS}\), \(\mathrm{ABC}\), and \(\mathrm{JBC}\)) continue to perform well.19
Table 2 reports the results for the Logit specification. These findings are very similar to those for the Probit specification, except for the deterioration in the performance of the analytical bias correction method in smaller sample sizes. The analogous results for dynamic specifications with predetermined covariates are also qualitatively similar and discussed in Appendix 7.3. Overall, the considered numerical experiments demonstrate that our two-step estimator has good finite sample properties and that it can be efficiently computed even in large panels with \((N,T) = (1000,200)\).
| POOL | NNR | \(\mathrm{TS}^*\) | \(\mathrm{ABC}^*\) | \(\mathrm{JBC}^*\) | \(\mathrm{TS}\) | ABC | JBC | \(\bar{R}\) | |
| \(( \times 10^{-2})\) | \(( \times 10^{-2})\) | \((10^{-2})\) | \((\times 10^{-2})\) | \((\times 10^{-2})\) | \((\times 10^{-2})\) | \((\times 10^{-2})\) | \((\times 10^{-2})\) | ||
| N = 50, T = 40 | |||||||||
| BIAS | 2.14 | 3.69 | 3.78 | 0.78 | -3.05 | 3.78 | 0.85 | -2.87 | 1.962 |
| STD | (1.82) | (1.85) | (2.53) | (2.15) | (3.61) | (2.51) | (2.20) | (3.74) | |
| N = 100, T = 40 | |||||||||
| BIAS | 2.04 | 3.25 | 2.29 | 0.18 | -1.56 | 2.30 | 0.18 | -1.55 | 1.999 |
| STD | (1.46) | (1.32) | (1.59) | (1.42) | (2.13) | (1.59) | (1.42) | (2.13) | |
| N = 200, T = 40 | |||||||||
| BIAS | 2.06 | 3.08 | 1.79 | 0.05 | -0.55 | 1.79 | 0.05 | -0.55 | 2.000 |
| STD | (1.35) | (1.11) | (1.06) | (0.96) | (1.36) | (1.06) | (0.96) | (1.36) | |
| N = 100, T = 100 | |||||||||
| BIAS | 2.07 | 2.70 | 1.17 | 0.07 | -0.34 | 1.17 | 0.07 | -0.34 | 2.000 |
| STD | (1.10) | (0.91) | (0.86) | (0.81) | (0.97) | (0.86) | (0.81) | (0.97) | |
| N = 200, T = 100 | |||||||||
| BIAS | 2.01 | 2.41 | 0.83 | 0.02 | -0.13 | 0.83 | 0.02 | -0.13 | 2.000 |
| STD | (0.94) | (0.68) | (0.60) | (0.58) | (0.65) | (0.60) | (0.58) | (0.65) | |
| N = 200, T = 200 | |||||||||
| BIAS | 1.98 | 2.14 | 0.53 | 0.00 | -0.07 | 0.53 | 0.00 | -0.07 | 2.000 |
| STD | (0.75) | (0.49) | (0.41) | (0.40) | (0.44) | (0.41) | (0.40) | (0.44) | |
| N = 1000, T = 200 | |||||||||
| BIAS | 1.98 | 1.93 | 0.31 | 0.00 | -0.01 | 0.31 | 0.00 | -0.01 | 2.000 |
| STD | (0.57) | (0.28) | (0.18) | (0.18) | (0.25) | (0.18) | (0.18) | (0.25) |
| POOL | NNR | \(\mathrm{TS}^*\) | \(\mathrm{ABC}^*\) | \(\mathrm{JBC}^*\) | \(\mathrm{TS}\) | ABC | JBC | \(\bar{R}\) | |
| \(( \times 10^{-2})\) | \(( \times 10^{-2})\) | \((10^{-2})\) | \((\times 10^{-2})\) | \((\times 10^{-2})\) | \((\times 10^{-2})\) | \((\times 10^{-2})\) | \((\times 10^{-2})\) | ||
| N = 50, T = 40 | |||||||||
| BIAS | 7.51 | 6.89 | 5.40 | 3.64 | -3.32 | 5.45 | 3.70 | -3.21 | 1.802 |
| STD | (2.48) | (2.40) | (3.64) | (3.48) | (6.56) | (3.68) | (3.52) | (6.75) | |
| N = 100, T = 40 | |||||||||
| BIAS | 7.57 | 6.64 | 3.09 | 1.70 | -2.25 | 3.10 | 1.71 | -2.23 | 1.959 |
| STD | (1.99) | (1.88) | (2.39) | (2.33) | (4.01) | (2.40) | (2.34) | (4.02) | |
| N = 200, T = 40 | |||||||||
| BIAS | 7.46 | 6.26 | 1.89 | 0.69 | -1.24 | 1.89 | 0.69 | -1.24 | 1.999 |
| STD | (1.67) | (1.50) | (1.56) | (1.50) | (2.33) | (1.56) | (1.50) | (2.33) | |
| N = 100, T = 100 | |||||||||
| BIAS | 7.35 | 5.82 | 1.23 | 0.32 | -0.90 | 1.23 | 0.32 | -0.90 | 2.000 |
| STD | (1.40) | (1.22) | (1.17) | (1.12) | (1.51) | (1.17) | (1.12) | (1.51) | |
| N = 200, T = 100 | |||||||||
| BIAS | 7.36 | 5.39 | 0.76 | 0.11 | -0.29 | 0.76 | 0.11 | -0.29 | 2.000 |
| STD | (1.23) | (1.01) | (0.88) | (0.85) | (1.07) | (0.88) | (0.85) | (1.07) | |
| N = 200, T = 200 | |||||||||
| BIAS | 7.34 | 4.87 | 0.51 | 0.08 | -0.09 | 0.51 | 0.08 | -0.09 | 2.000 |
| STD | (0.95) | (0.74) | (0.59) | (0.57) | (0.68) | (0.59) | (0.57) | (0.68) | |
| N = 1000, T = 200 | |||||||||
| BIAS | 7.36 | 4.07 | 0.27 | 0.02 | -0.01 | 0.27 | 0.02 | -0.01 | 2.000 |
| STD | (0.70) | (0.45) | (0.27) | (0.26) | (0.34) | (0.27) | (0.26) | (0.34) |
In this section, we revisit the empirical application studied by [1], who proposed estimating the gravity equation with interactive fixed effects to examine the determinants of international trade flows. Specifically, our goal is to investigate if our two-step approach can replicate the estimates reported in [1].
The data, originally from [35], include bilateral trade flows and other relevant variables for \(N = 157\) countries. Following [1], we focus on the year \(1986\) with the trade network including 157 countries, with the effective sample size \(157 \times 156 = 24{,}492\) (the number of distinct country pairs). The outcome variable \(Y_{ij}\) represents the volume of trade from country \(i\) to country \(j\) (measured in thousands of constant \(2000\) US dollars). The covariates \(X_{ij}\) include key determinants of \(Y_{ij}\), such as the logarithm of the distance between the capitals of the two countries (Log distance) and binary indicators for shared border (Border), legal system (Legal), common language (Language), colonial ties (Colony), currency union (Currency), regional free-trade agreement (FTA), and religion (Religion). The descriptive statistics are presented in Table 3 below.
| Mean | Standard deviation | |
|---|---|---|
| Trade Volume | 84,542 | 1,082,219 |
| Log distance | 4.18 | 0.78 |
| Border | 0.02 | 0.13 |
| Legal | 0.37 | 0.48 |
| Language | 0.29 | 0.45 |
| Colony | 0.01 | 0.10 |
| Currency | 0.60 | 1.37 |
| FTA | 0.01 | 0.08 |
| Religion | 0.17 | 0.25 |
As in [1], we estimate the following Poisson regression model \[\begin{align} Y_{ij}\mid X_{ij}, \lambda_{1, i} , \gamma_{1, j}, \lambda_{2, i}, \gamma_{2, j} \sim \mathrm{Poisson}(\exp\{\beta' X_{ij} + \lambda_{1, i} + \gamma_{1, j} + \lambda'_{2, i}\gamma_{2, j}\}), \end{align}\] where we explicitly include additive two-way fixed effects \(\lambda_{1,i}+\gamma_{1,j}\) together with an interactive fixed effects component \(\lambda_{2,i}'\gamma_{2,j}\), which allows for richer patterns of unobserved heterogeneity such as homophily based on latent characteristics of countries \(i\) and \(j\). We compute our two-step estimator and select the regularization parameter and the number of factors using the algorithms provided in Section 4.20
In Table 4 below, we provide the estimates produced by our two-step estimator (column \(\mathrm{TS}\)) and the estimates reported by [1] (column \(\mathrm{CFW}\)). Notice that since, using our algorithm, the estimated number of factors is \(\hat{R}=2\), we include the estimates of [1] for \(R = 2\) only. As a baseline benchmark, we also report the estimates produced by the estimator incorporating only two-way fixed effects \(\lambda_{1,i}+\gamma_{1,j}\) (column \(\mathrm{TWFE}\)), which match the analogous results in [1].
We find that the estimates produced by our two-step estimator and the ones obtained by [1] are overall very similar. The main difference is in the estimated effects of countries having the same currency, but this difference is still small relative to the associated standard errors. We also compare the (average) log-likelihood achieved by the methods (multiplied by 100), and find that the value achieved by our two-step estimator, \(0.6711\), is only slightly lower than the one reported by [1], \(0.6714\). Thus, we conclude that our two-step estimator can effectively recover the FE estimates in fairly large-dimensional empirical settings, making it a computationally attractive approach to estimation of nonlinear interactive fixed effects models in practice.
| TWFE | CFW | \(\mathrm{TS}\) | |
| R = 0 | R = 2 | R = 2 | |
| Log distance | -0.64 | -0.71 | -0.71 |
| (0.07) | (0.06) | (0.05) | |
| Border | 0.71 | 0.32 | 0.32 |
| (0.16) | (0.05) | (0.06) | |
| Legal | 0.30 | 0.26 | 0.26 |
| (0.06) | (0.04) | (0.04) | |
| Language | -0.17 | -0.02 | -0.02 |
| (0.10) | (0.06) | (0.06) | |
| Colony | 0.36 | 0.39 | 0.39 |
| (0.12) | (0.09) | (0.10) | |
| Currency | 0.60 | 1.37 | 1.25 |
| (0.09) | (0.41) | (0.34) | |
| FTA | 0.25 | 0.17 | 0.17 |
| (0.13) | (0.07) | (0.06) | |
| Religion | -0.25 | 0.24 | 0.24 |
| (0.12) | (0.13) | (0.08) | |
| Log-Likelihood | -0.44 | 0.67 | 0.67 |
During the preparation of this work, the authors used ChatGPT for assistance with language editing and LaTeX formatting, and Refine.ink to identify typos in the proofs. The authors reviewed and edited the output as needed and take full responsibility for the content of the published article.
This section presents extensions and technical discussions related to the assumptions in the main text: (1) We discuss the normalization of fixed effects. (2) We show how to extend our analysis to the case where \(X_{it}\) contains predetermined variables. (3) We give easily verifiable sufficient conditions for restricted strong convexity (Assumption 2) and diagonal structure (Assumption 3) when predetermined covariates are present. (4) We demonstrate that our method remains applicable even when \(\Sigma_{\lambda}\Sigma_{\gamma}\) contains repeated eigenvalues (relaxing Assumption 1[item:strong95factors]).
Similar to the linear case, we need additional \(R^2\) restrictions to identify fixed effects \((\Lambda_0, \Gamma_0)\)21. In Section 3, we directly impose normalization constraints on fixed effects \((\Lambda_0, \Gamma_0)\) such that \(\Lambda_0'\Lambda/N\) and \(\Gamma_0'\Gamma_0/T\) are diagonal and satisfy \(\Lambda_0'\Lambda_0 /N= \Gamma_0'\Gamma_0/T\). The feasibility of such normalization can be illustrated as follows: for any \((\Lambda_0, \Gamma_0)\) satisfying Assumption 1[item:strong95factors], let \(D\) be the diagonal matrix containing the square roots of the eigenvalues of the matrix \((NT)^{-1} (\Lambda_0'\Lambda_0)^{1/2}\Gamma_0'\Gamma_0 (\Lambda_0'\Lambda_0)^{1/2}\), and let \(\Upsilon\) be the matrix collecting the corresponding eigenvectors. Then there exists a unique transformation matrix \(G= D^{1/2} \Upsilon'(\Lambda_0\Lambda_0/N)^{-1/2}\), such that the normalized nuisance parameters, \[\begin{align} \Lambda^G_0 := \Lambda_0 G', \quad \Gamma^G_0 := \Gamma_0 G^{-1}, \end{align}\] lie in the normalized parameter space \(\Phi_{NT}\)22. Unlike in the main text, in the following analysis and proofs, we distinguish between \((\Lambda_0, \Gamma_0)\) and \((\Lambda_0^G, \Gamma_0^G)\) to ensure a rigorous argument.
Our estimation method is applicable in scenarios where \(X_{it}\) includes predetermined variables (e.g., lags of \(Y_{it}\) in dynamic panels). While the main text focuses on the simpler case where \(X_{it}\) is strictly exogenous, this section extends the analysis to the more complex case involving predetermined variables. Considering that \(X_{it}\) includes predetermined variables is crucial in panel data because the time order is very important—–unlike in network data, where node order is irrelevant. Including these variables allows us to account for dynamic effects common in empirical research.
We partition \(X_{it}\) into \(X_{it} := (W'_{it}, Z'_{it})'\), where \(W_{it}\) represents \(d_{W}\)-dimensional predetermined covariates and \(Z_{it}\) represents \(d_{Z}\)-dimensional exogenous covariates. We collect \(Z_{it, d}\) into the covariate matrix \(Z_{d} \in \mathbb{R}^{N\times T}\) for each \(d= 1,2,\ldots, d_Z\), and let \(Z\) be the collection of all strictly exogenous covariate matrices, \(Z= \{Z_{1}, \ldots, Z_{d_Z}\}\). For each \(i = 1,2,\ldots, N\), denote \(Y_i := (Y_{i1}, Y_{i2}, \ldots, Y_{iT})'\) and \(W_i := \left(W_{i1}, W_{i2}, \ldots, W_{iT}\right)'\). In addition, we use \(\mathbb{P}_{Z, \Lambda_0, \Gamma_0}(\cdot) = \mathbb{P}(\cdot\mid Z, \Lambda_0, \Gamma_0)\) to denote the conditional probability and use \(\mathbb{E}_{Z, \Lambda_0, \Gamma_0}(\cdot) = \mathbb{E}(\cdot\mid Z, \Lambda_0, \Gamma_0)\) to denote the conditional expectation.
Let us now consider the regularity conditions appropriate for predetermined covariates.
Assumption 4 (Regularity conditions - predetermined covariates). Suppose that
(Sampling) Conditional on \((Z, \Lambda_{0}, \Gamma_{0})\), \(\{(Y_{i},W_i)\}_{i=1,\ldots, N}\) are independent across \(i\), and for each \(i\), \(((Y_{i1}, W_{i1}), (Y_{i2}, W_{i2}), \ldots, (Y_{iT}, W_{iT}))\) is \(\phi\)-mixing with mixing coefficient \(\phi_i(\tau)\rightarrow 0\) as \(\tau \rightarrow \infty\), where \[\begin{align} \phi_i(\tau) = \sup_{t} \sup_{A \in \mathcal{A}^{i}_{t}, B\in \mathcal{B}^{i}_{t + \tau}}|\mathbb{P}_{Z, \Lambda_0, \Gamma_0}(B\mid A) - \mathbb{P}_{ Z, \Lambda_0, \Gamma_0}(B)|. \end{align}\] Here, \(\mathcal{A}^{i}_{t}\) is the sigma-field generated by \(\{ \ldots, (Y_{i,t-1}, W_{i,t-1}), (Y_{i,t},W_{i,t})\}\), and \(\mathcal{B}^{i}_{t+\tau}\) is the sigma-field generated by \(\{(Y_{i,t+\tau},W_{i, t+\tau}), (Y_{i,t+\tau+1},W_{i, t+\tau+1}), \ldots\}\). For mixing coefficients \(\phi_i(\tau)\), \(i=1,2,\ldots, N\), we further assume that they exhibit a uniformly exponential decay rate: there exists \(\zeta_0 >0\) such that \(\sup_{1\leq i\leq N}\phi_i(\tau) \leq e^{-\zeta_0 \tau }\).
(Boundedness) The parameter spaces of \(\beta\), \(\lambda_i\), and \(\gamma_t\) are bounded uniformly for all \(i, t, N, T\). In addition, there exists a constant \(\rho_X>0\) such that \(\max_{d=1,\ldots, d_X}\|X_d\|_\infty <\rho_X\) for all \(i, t\) and \(N, T\).
(Smoothness and Convexity) \(-\ell_{it}(\cdot)\) is four times continuously differentiable and strictly convex almost surely. Furthermore, we assume that \(0 < b_{\min} \leq -\ddot{\ell}_{it}( X_{it}' \beta + \lambda_i' \gamma_t)\leq b_{\max}<\infty\) almost surely for all \(\beta, \lambda_i, \gamma_t\) in the parameter space uniformly over \(i, t, N, T\).
(Strong factors) \(\frac{1}{N} \sum_{i=1}^{N}\lambda_{0, i}\lambda_{0, i}' \stackrel{p}{\longrightarrow}\Sigma_{\lambda}\) and \(\frac{1}{T} \sum_{t=1}^{T}\gamma_{0, t}\gamma_{0, t}'\stackrel{p}{\longrightarrow}\Sigma_{\gamma}\), where \(\Sigma_{\lambda} >0\) and \(\Sigma_{\gamma} >0\). In addition, the eigenvalues of \(\Sigma_{\lambda}\Sigma_{\gamma}\) are distinct.
(Generalized non-collinearity) For any \(\Gamma\in \mathbb{R}^{T\times R}\), let \(M_{\Lambda_0}\) and \(M_{\Gamma}\) be coprojection matrices of \(\Lambda_0\) and \(\Gamma\) respectively. The \(d_{X}\times d_{X}\) matrix \(D(\Gamma)\) with elements \[\begin{align} D(\Gamma)_{d_1, d_2} = \frac{1}{NT} \mathrm{Tr} \left(M_{\Lambda_0}X_{d_1} M_{\Gamma}X_{d_2}'\right), \quad d_{1}, d_{2} = 1, \ldots, d_X, \end{align}\] satisfies \(\inf_{\Gamma \in \mathbb{R}^{T\times R}} \sigma_{\min} (D(\Gamma) )>0\) wpa1.
The new regularity assumption differs from the previous one (Assumption 1) only in the sampling assumption. In Assumption 4[item:sampling95pre], we impose weak dependence on the sequence \((Y_{it}, W_{it})\): conditional on \((Z, \Lambda_0, \Gamma_0)\), each sequence is \(\phi\)-mixing with an exponential decay rate. Although this is stricter than necessary, we adopt it to align with the sufficient conditions for restricted strong convexity (RSC). \(\phi\)-mixing can be replaced with \(\alpha\)-mixing, and the exponential decay rate can be relaxed to a sufficiently fast polynomial decay, as in [29]. In addition, we do not need identical distribution or stationarity assumptions on the sequence.
We now give the new consistency results of NNR estimators, extending Theorem 1 to incorporate predetermined covariates:
Theorem 5. For any \(\alpha>0\) such that \(\varphi_{NT} \geq (1+\alpha) \max\{ \|\nabla_{\beta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)\|_{2}, \sqrt{NT}\|\nabla_{\Theta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)\|_{\mathrm{op}} \}\), under Assumption 4 and 2, as \(N, T\rightarrow\infty\), there exist constants \(c_1, c_2>0\) that do not depend on \(N, T\) such that wpa1, \[\begin{align} \| \hat{\beta}_{\mathrm{nuc}} - \beta_0\|_2 & \leq c_1 \left(\varphi_{NT} + \log (NT)/\sqrt{\min\{N, T\}}\right), \\ \frac{1}{\sqrt{NT}}{\vert\kern-0.25ex\vert \hat{\Theta}_{\mathrm{nuc}} - \Theta_0 \vert\kern-0.25ex\vert}_F & \leq c_1 \left(\varphi_{NT} + \log (NT)/\sqrt{\min\{N, T\}}\right). \end{align}\] In addition, wpa1, \[\begin{align} \frac{1}{\sqrt{N}}\|\hat{\Lambda}_{\mathrm{nuc}} - \Lambda^G_0 \|_{\mathrm{F}} & \leq c_2 \left(\varphi_{NT} + \log (NT)/\sqrt{\min\{N, T\}} \right), \\ \frac{1}{\sqrt{T}}{\vert\kern-0.25ex\vert \hat{\Gamma}_{\mathrm{nuc}} - \Gamma^G_0 \vert\kern-0.25ex\vert}_{\mathrm{F}} & \leq c_2 \left(\varphi_{NT} + \log (NT)/\sqrt{\min\{N, T\}} \right). \end{align}\]
The following corollary extends Corollary 1 to include predetermined covariates.
Corollary 2. Under the conditions of Theorem 5, let \(\varphi_{NT} = O\left(\log (NT)/ \sqrt{ \min\{N, T\}}\right)\). Then there exist constants \(c_3, c_4>0\) that do not depend on \(N, T\) such that wpa1, \[\begin{align} \| \hat{\beta}_{\mathrm{nuc}} - \beta_0\|_2 & \leq c_3 \log (NT)/\sqrt{\min\{N, T\}}, \\ \frac{1}{\sqrt{NT}}{\vert\kern-0.25ex\vert \hat{\Theta}_{\mathrm{nuc}} - \Theta_0 \vert\kern-0.25ex\vert}_{\mathrm{F}} & \leq c_3 \log (NT)/\sqrt{\min\{N, T\}}. \end{align}\] In addition, wpa1, \[\begin{align} \frac{1}{\sqrt{N}}\|\hat{\Lambda}_{\mathrm{nuc}} - \Lambda^G_0\|_{\mathrm{F}} & \leq c_4 \log (NT)/\sqrt{\min\{N, T\}}, \\ \frac{1}{\sqrt{T}}\| \hat{\Gamma}_{\mathrm{nuc}} - \Gamma^G_0 \|_{\mathrm{F}} & \leq c_4 \log (NT)/\sqrt{\min\{N, T\}}. \end{align}\]
We state the new diagonal structure assumption which generalizes Assumption 3.
Assumption 5 (Diagonal structure). The population Hessian in 9 evaluated at the true parameters, \(\mathbb{E}_{Z, \Lambda_0, \Gamma_0}\mathcal{H}_{NT}(\beta_0, \Lambda_0, \Gamma_0)\), admits a diagonal structure, i.e., there exists a constant \(C>0\) (does not depend on \(N, T\)) such that \[\begin{align} \mathbb{E}_{Z, \Lambda_0, \Gamma_0}\mathcal{H}_{NT}(\beta_0, \Lambda_0^G, \Gamma_0^G) \geq C \mathrm{diag}\left\{\mathbb{I}_{d_X}, \frac{1}{N}\mathbb{I}_{NR}, \frac{1}{T}\mathbb{I}_{TR}\right\}. \end{align}\]
The local convexity and the asymptotic equivalence can be extended to incorporate predetermined covariates.
Theorem 6. Under Assumption 5 and the conditions in Corollary 2, suppose \(N, T\) have the same order. Let \(\delta_{NT}= \log (NT) \min\{N^{-3/8}, T^{-3/8}\}\), and let \(\mathcal{B}_{\delta_{NT}}\) be the neighborhood defined in 7 . The following results hold:
the local optimization problem 9 is strictly convex wpa1;
our two-step estimator is asymptotically equivalent to the FE estimator;
the local optimization problem 9 is strongly convex wpa1 uniformly on \(\mathcal{B}_{\delta_{NT}}\), i.e., there exists a constant \(c_5>0\) independent of \(N, T\) such that for any \((\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}\), \[\begin{align} \mathcal{H}_{NT}(\beta, \Lambda, \Gamma) > c_5 \mathrm{diag}\left\{\mathbb{I}_{d_X}, \frac{1}{N}\mathbb{I}_{NR}, \frac{1}{T}\mathbb{I}_{TR}\right\}, \quad \text{wpa1}. \end{align}\]
We aim to provide a set of easily verifiable sufficient conditions for the restricted strong convexity (RSC) condition in the panel data setting. Given the critical role of the RSC condition in low-rank estimation, establishing verifiable low-level conditions in the panel data setting is important and may be of independent interest. While [6] offers sufficient conditions for verifying the RSC condition in the panel data context, their approach is limited to strictly exogenous covariates. We extend their proof strategy to accommodate predetermined covariates, thereby broadening the applicability of the RSC condition in panel data. It is worth noting that although we are focusing on low-rank estimation with homogeneous slopes, our proof strategy is quite flexible and is applicable to more general settings. For example, it can be readily adapted to establish the RSC condition under heterogeneous slopes, as in the models considered by [6] and [7].
Assumption 6. There exists a sequence of random vectors \(\{v_{it}\}_{ 1\leqslant i \leqslant N, 1\leqslant t \leqslant T}\), such that
(Conditional weak dependence) Conditional on \(\mathcal{V}: =\left\{v_{it} \mid 1\leq i\leq N, 1\leq t\leq T\right\}\), \(\{X_{i}\}_{1\leq i\leq N}\) are independent across \(i\), and for each \(i\), \((X_{i1}, X_{i2}, \ldots, X_{iT})\) is \(\phi\)-mixing with mixing coefficient \(\phi_i(\tau)\rightarrow 0\) as \(\tau \rightarrow \infty\). We assume that the mixing coefficients exhibit a uniformly exponential decay rate, i.e., there exists \(\zeta_1 >0\) such that \(\sup_{1\leq i\leq N}\phi_i(\tau) \leq e^{-\zeta_1 \tau }\).
(Conditional variability) Conditional on \(\mathcal{V}\), there exists \(\kappa_0 >0\) such that the following inequality holds for any \(N, T\), \[\begin{align} \inf_{1\leq i\leq N, 1\leq t \leq T}\sigma_{\min}\left( \begin{pmatrix} \mathbb{E}(X_{it}X_{it}' \mid \mathcal{V}) & \mathbb{E} (X_{it}\mid \mathcal{V}) \\ \mathbb{E}(X_{it}'\mid \mathcal{V}) & 1 \end{pmatrix} \right) \geq \kappa_0, \end{align}\] where \(\sigma_{\min}(\cdot)\) denotes the minimum eigenvalue of a matrix.
The first part of Assumption 6 states that while it is unrealistic to assume weak dependence of \(\{X_{it}\}_{1\leq t\leq T}\) across \(t\) due to the presence of common factors, we assume that the serial correlation can be significantly reduced by conditioning on a set of latent factors, \(\mathcal{V}\). The second part of Assumption 6 makes sure that \(X_{it}\) cannot be fully explained by \(\mathcal{V}\).
More specifically, Assumption 6[item:conditional95weak95dependence95RSC] imposes weak dependence on the sequence \(\{X_{it}\}_{1\leq t\leq T}\): conditional on the common factors \(\mathcal{V}\), the sequence is \(\phi\)-mixing with an exponential decay rate. The conditional weak dependence assumption is much milder than it appears, and here are some examples:
Example 3 (Factor model). Suppose \(X_{it} = \lambda_x'\gamma_t + \epsilon_{it}\) where \(\lambda_x'\gamma_t\) is the common factor and \(\epsilon_{it}\) is the error term. Let \(v_{it} = \lambda_x'\gamma_t\). The conditional weak dependence condition holds if \(\{\epsilon_{it}\}_{1\leq i\leq N, 1\leq t\leq T}\) is independent across both \(i\) and \(t\) conditional on \(\mathcal{V}\).
Example 4 (Autoregression). Allowing for serial correlation becomes necessary when the model includes predetermined covariates. Suppose \(X_{it} = A X_{i,t-1} + B v_{it} + \epsilon_{it}\) where \(A\) and \(B\) are the coefficient matrices and \(\epsilon_{it}\) is an innovation term with uniform bounded support. We further assume that \(\psi_{\max}(A) < 1\), where \(\psi_{\max}(A)\) denotes the largest singular value of \(A\). Then conditional on \(\mathcal{V}\), the sequence \(\{X_{it}\}_{1\leq t\leq T}\) is \(\phi\)-mixing with exponential decay rate.
Example 5 (Non-separable model). A more interesting case is the non-separable model with serial correlation, i.e., \(X_{it} = h(X_{i,t-1}, v_{it})\). For example, consider the dynamic Logit panel, where \(X_{it} = Y_{i, t-1}\), and \[\begin{align} Y_{it} = 1(\beta Y_{i, t-1} + \lambda_{0, i}'\gamma_{0, t} + \epsilon_{it} > 0). \end{align}\] Here, \(\epsilon_{it}\) follows a standard logistic distribution. Let \(v_{it} = \lambda_{0, i}'\gamma_{0, t}\), and assume that the innovation terms \(\{\epsilon_{it}\}_{1\leq i\leq N, 1\leq t\leq T}\) are independent across both \(i\) and \(t\) and are strictly exogenous. By Assumption 4[item:smoothing95pre], we conclude that, conditional on \(\mathcal{V}\), \(\{X_{it}\}_{1\leq t\leq T}\) is a time-inhomogeneous Markov process with transition matrices whose entries are uniformly bounded away from zero. Furthermore, by verifying Doeblin’s condition, it follows that \(\{X_{it}\}_{1\leq t\leq T}\) is \(\phi\)-mixing with exponential decay rate. We recommend the interested reader in nonlinear time series models to refer to [36] and [37] for more details.
Assumption 6[item:conditional95variability95RSC] imposes that, even after controlling for the set of latent factors \(\mathcal{V}\), \(X_{it}\) preserves sufficient variation. Notably, by applying the Schur decomposition, we directly obtain that \[\begin{align} \begin{pmatrix} \mathbb{E}(X_{it}X_{it}' \mid \mathcal{V}) & \mathbb{E} (X_{it}\mid \mathcal{V}) \\ \mathbb{E}(X_{it}'\mid \mathcal{V}) & 1 \end{pmatrix} > 0 \quad \Longleftrightarrow \quad \mathrm{Var}(X_{it}\mid \mathcal{V})>0. \end{align}\] If \(X_{it}\) given \(\mathcal{V}\) is identically distributed, the conditional variability condition simplifies to \(\mathrm{Var}(X_{it}\mid \mathcal{V})>0\). If we further assume that \(X_{it}\) admits an additive structure, this condition is satisfied if \(\inf_{1\leq i\leq N, 1\leq t\leq T}\mathrm{Var}(\epsilon_{it}\mid \mathcal{V})>0\). Assumption 6[item:conditional95variability95RSC] can be viewed as a generalization of these simpler cases.
We will provide sufficient conditions for verifying Assumption 3 and Assumption 5, which are critical in establishing local convexity in Theorem 2 and Theorem 6. Assumption 3 and Assumption 5 require that the population Hessian, evaluated at the normalized true parameters \((\beta_0, \Lambda_0^G, \Gamma^G_0)\), is positive definite, with a minimum eigenvalue on the order of \((\max\{N, T\})^{-1}\). We will show that these two assumptions are weak and that, under the strong factor condition, it suffices for \(X_{it}\) to exhibit weak dependence after accounting for the influence of common factors and to retain components that cannot be fully explained by these factors.
Assumption 7. There exists a sequence of random vectors \(\{u_{it}\}_{ 1\leqslant i\leqslant N, 1\leqslant t\leqslant T}\), such that
(Conditional weak dependence) Conditional on \(\mathcal{U}: =\left\{u_{it} \mid 1\leq i\leq N, 1\leq t\leq T\right\}\), \(\{X_{i}\}_{1\leq i\leq N}\) is independent across \(i\), and for each \(i\), \((X_{i1}, X_{i2}, \ldots, X_{iT})\) is \(\phi\)-mixing with mixing coefficient \(\phi_i(\tau)\rightarrow 0\) as \(\tau \rightarrow \infty\). We assume that the mixing coefficients exhibit a uniformly exponential decay rate, i.e., there exists \(\zeta_2 >0\) such that \(\sup_{1\leq i\leq N}\phi_i(\tau) \leq e^{-\zeta_2 \tau }\).
(Conditional variability) Conditional on \(\mathcal{U}\), there exists \(0 <\nu <1\) such that the following inequality holds uniformly for all \(1\leq i\leq N\) and \(1\leq t\leq T\), \[\begin{align} \mathbb{E}(\ddot{\ell}_{it}^0X_{it}\mid \mathcal{U})\mathbb{E}(\ddot{\ell}_{it}^0X_{it}'\mid \mathcal{U}) \leq \nu \mathbb{E} (\ddot{\ell}_{it}^0 \mid \mathcal{U} )\mathbb{E} (\ddot{\ell}_{it}^0X_{it}X_{it}'\mid \mathcal{U}), \end{align}\] where \(\ddot{\ell}_{it}^0 =\ddot{\ell}_{it}(X_{it}'\beta_0 + \lambda_{0, i}'\gamma_{0, t})\).
The first part of this assumption is identical to the conditional weak dependence stated in Assumption 7[item:conditional95weak95dependence95hessian]. It is worth noting that the condition requiring \(\{X_{it}\}_{1 \leq t \leq T}\) to be \(\phi\)-mixing with an exponential decay rate uniformly across \(i\) is stronger than necessary and is imposed here for consistency with Assumption 6[item:conditional95weak95dependence95RSC]. In fact, this condition could be relaxed to \(\{X_{it}\}_{1 \leq t \leq T}\) being \(\alpha\)-mixing with a polynomial decay rate, still uniformly across \(i\). Assumption 7[item:conditional95variability95hessian] imposes a weak additional restriction on the model. To elaborate, let us use \(\mathbb{E}_{\mathcal{U}}(\cdot)\) to denote the conditional expectation \(\mathbb{E}(\cdot \mid \mathcal{U})\), and first consider the simplest case where \(\ddot{\ell}_{it}^0\) is a constant (corresponding to a linear panel model) with \(d_X = 1\). In this case, it is straightforward to verify that Assumption 7[item:conditional95variability95hessian] is equivalent to \((\mathbb{E}_{\mathcal{U}}{X}_{it})^2 \leq \nu \mathbb{E}_{\mathcal{U}}({X}^2_{it})\) uniformly for all \(i, t, N, T\). When extending the analysis to the nonlinear model with \(d_X = 1\), it follows (using Hölder’s inequality) that \[\begin{align} \mathbb{E}_{\mathcal{U}}(\ddot{\ell}_{it}^0{X}_{it}) \leq \mathbb{E}_{\mathcal{U}}(|\ddot{\ell}_{it}^0{X}_{it}|) = \mathbb{E}_{\mathcal{U}}(|\ddot{\ell}_{it}^0|^{1/2} |\ddot{\ell}_{it}^0{X}^2_{it}|^{1/2}) \leq \left(\mathbb{E}_{\mathcal{U}}(-\ddot{\ell}_{it}^0)\right)^{1/2} \left(\mathbb{E}_{\mathcal{U}}(-\ddot{\ell}_{it}^0{X}^2_{it})\right)^{1/2}. \end{align}\] Equality holds if and only if \(X_{it}\) is constant given \(\mathcal{U}\), i.e., \(X_{it}\) can be fully explained by \(\mathcal{U}\). Thus, in this case, Assumption 7[item:conditional95variability95hessian] can be viewed as a uniform version of Hölder’s inequality. Extending the analysis from \(d_X = 1\) to \(d_X > 1\) is straightforward and yields the same conclusions.
Lemma 3 (Sufficient conditions for diagonal structure). Under Assumption 4 and Assumption 7, Assumption 3 and Assumption 5 hold.
We now demonstrate that our method remains applicable even when \(\Sigma_{\lambda}\Sigma_{\gamma}\) contains repeated eigenvalues (relaxing Assumption 1[item:strong95factors] and Assumption 4[item:strong95factors95pre]). When \(\Sigma_{\lambda}\Sigma_{\gamma}\) has distinct eigenvalues, there exists a unique \(R\)-dimensional invertible matrix, \(G= D^{1/2} \Upsilon'(\Lambda_0'\Lambda_0/N)^{-1/2}\), which does not depend on the sample \(\{(Y_{it}, X_{it})\}_{1\leq i\leq N, 1\leq t\leq T}\), such that \(\Lambda^{G\prime}_0\Lambda^G_0/N, \Gamma^{G\prime}_0\Gamma^G_0/T\) are diagonal and satisfy \(\Lambda^{G\prime}_0\Lambda^G_0/N = \Gamma^{G\prime}_0\Gamma^G_0/T\). However, when \(\Sigma_{\lambda}\Sigma_{\gamma}\) has repeated eigenvalues, the transformation \(G\) may depend on the sample. In this case, we denote it by \(\hat{G}\): \(\hat{G} = D^{1/2}\hat{O}\Upsilon'(\Lambda_0'\Lambda_0 /N)^{-1/2}\) (see the proof of Theorem 5), where \(\hat{O}\) is an orthogonal matrix that may also depend on the sample.
The dependence of \(\hat{G}\) on the sample adds extra complexity to the proof. Therefore, we aim to construct an objective function that is independent of \(G\), along with its corresponding Hessian, which allows us to directly apply the proof of Theorem 6 to establish the asymptotic equivalence between our estimators and the FE estimators in [1].
Define \[\begin{align} G_{NT} := \mathrm{diag}\left\{\mathbb{I}_{d_X}, \underbrace{\hat{G}, \ldots, \hat{G}}_{N}, \underbrace{\hat{G}^{-1}, \ldots, \hat{G}^{-1}}_{T}\right\}. \end{align}\] It is straightforward to verify that (i) \(\hat{G}\) is invertible wpa1, (ii) the minimum eigenvalue of \(\hat{G}\) is strictly positive wpa1 and does not depend on the sample, and (iii) the maximum eigenvalue of \(\hat{G}\) is uniformly bounded wpa1 (see the proof of Theorem 5) and does not depend on the sample. Therefore, it follows that \(G_{NT}\) is invertible wpa1, the minimum eigenvalue of \(G_{NT}\) is strictly positive wpa1 and independent of the sample, and the maximum eigenvalue of \(G_{NT}\) is uniformly bounded wpa1. In addition, one can verify that neither \[\begin{align} \nabla^2 \mathcal{L}_{NT}(\beta_0, \Lambda_0, \Gamma_0) = G_{NT}' \nabla^2 \mathcal{L}_{NT}(\beta_0, \Lambda_0^G, \Gamma_0^G) G_{NT} \end{align}\] or \(\nabla^2 \|\hat{\Lambda}_{\mathrm{nuc}}'\Lambda - \Gamma'\hat{\Gamma}_{\mathrm{nuc}} \|_{\mathrm{F}}^2\) do not depend on \(G_{NT}\). Let \(a := \min\left\{1, \sigma_{\min}(G_{NT}), 1/\sigma_{\max}(G_{NT})\right\}\). Then, \[\begin{align} \mathcal{H}_{NT}(\beta_0, \Lambda_0^G, \Gamma_0^G) & = G_{NT}^{-1\prime} \nabla^2 \mathcal{L}_{NT}(\beta_0, \Lambda_0, \Gamma_0)G_{NT}^{-1} + \frac{1}{2}G_{NT}^{\prime -1} G_{NT}^{\prime} \nabla^2\|\hat{\Lambda}_{\mathrm{nuc}}'\Lambda - \Gamma'\hat{\Gamma}_{\mathrm{nuc}} \|_{\mathrm{F}}^2 G_{NT} G_{NT}^{-1} \\ & = G_{NT}^{-1\prime} \left(\nabla^2 \mathcal{L}_{NT}(\beta_0, \Lambda_0, \Gamma_0) + \frac{1}{2} G_{NT}^{\prime} \nabla^2 \|\hat{\Lambda}_{\mathrm{nuc}}'\Lambda - \Gamma'\hat{\Gamma}_{\mathrm{nuc}} \|_{\mathrm{F}}^2 G_{NT}\right) G_{NT}^{-1} \\ & \geq G_{NT}^{-1\prime} \underbrace{\left(\nabla^2 \mathcal{L}_{NT}(\beta_0, \Lambda_0, \Gamma_0) + \frac{a^2}{2} \nabla^2 \|\hat{\Lambda}_{\mathrm{nuc}}'\Lambda - \Gamma'\hat{\Gamma}_{\mathrm{nuc}} \|_{\mathrm{F}}^2 \right)}_{\text{does not depend on the sample}} G_{NT}^{-1}. \end{align}\] Therefore, \(\mathbb{E}_{Z, \Lambda_0, \Gamma_0}\mathcal{H}_{NT}(\beta_0, \Lambda_0^G, \Gamma_0^G)\) satisfies Assumption 5 (or Assumption 3) if and only if \[\begin{align} \mathbb{E}_{Z, \Lambda_0, \Gamma_0}\left(\nabla^2 \mathcal{L}_{NT}(\beta_0, \Lambda_0, \Gamma_0) + \frac{a^2}{2} \nabla^2 \|\hat{\Lambda}_{\mathrm{nuc}}'\Lambda - \Gamma'\hat{\Gamma}_{\mathrm{nuc}} \|_{\mathrm{F}}^2 \right) \end{align}\] also satisfies Assumption 5 (or Assumption 3). Using the same technique as in the proof of Lemma 3, Assumption 4 together with Assumption 7 is sufficient to establish that \[\mathbb{E}_{Z, \Lambda_0, \Gamma_0}\left(\nabla^2 \mathcal{L}_{NT}(\beta_0, \Lambda_0, \Gamma_0) + \frac{a^2}{2} \nabla^2 \|\hat{\Lambda}_{\mathrm{nuc}}'\Lambda - \Gamma'\hat{\Gamma}_{\mathrm{nuc}} \|_{\mathrm{F}}^2\right)\] has a diagonal structure. Therefore, by replacing \((\Lambda^G_0, \Gamma_0^G)\) with \((\Lambda_0, \Gamma_0)\) in the proof of Theorem 6, we establish that Theorem 6 still holds even in the presence of repeated eigenvalues.
In this section, (i) we characterize the two-step estimator and present the corresponding algorithm used in the empirical application in Section 5; (ii) we provide procedures for the analytical bias correction and the split-panel Jackknife bias correction; (iii) we present Monte Carlo simulation results for nonlinear panel models with predetermined covariates.
We provide the optimization problem and the associated algorithm used in the empirical application in Section 5, where additive two-way fixed effects are explicitly incorporated into the index. Recall that we aim to estimate the following Poisson model: \[\begin{align} Y_{ij}\mid X_{ij}, \lambda_{1, i} , \gamma_{1, j}, \lambda_{2, i}, \gamma_{2, j} \sim \mathrm{Poisson}(\exp\{\beta' X_{ij} + \lambda_{1, i} + \gamma_{1, j} + \lambda'_{2, i}\gamma_{2, j}\}). \end{align}\] For notational simplicity, define \[\begin{align} \mathcal{L}_{N}\left(\beta, \lambda_1, \gamma_1, \Theta\right)& : = \frac{-1}{N(N-1)}\sum_{\substack{ i, j = 1, \ldots, N \\ i\neq j} } \log f(Y_{ij} \mid \beta'X_{ij} + \lambda_{1, i} + \gamma_{1, j} + \theta_{ij}), \\ \mathcal{L}_{N}\left(\beta, \lambda_1, \gamma_1, \Lambda_2, \Gamma_2 \right)& : = \frac{-1}{N(N-1)}\sum_{\substack{ i, j = 1, \ldots, N \\ i\neq j} } \log f(Y_{ij} \mid \beta'X_{ij} + \lambda_{1, i} + \gamma_{1, j} + \lambda_{2, i}' \gamma_{2, j}). \end{align}\]
In the first step, the NNR estimator \((\hat{\beta}_{\mathrm{nuc}},\hat{\lambda}_{\mathrm{1, nuc}}, \hat{\gamma}_{\mathrm{1, nuc}}, \hat{\Theta}_{\mathrm{nuc}})\) solves \[\begin{align} \left(\hat{\beta}_{\mathrm{nuc}},\hat{\lambda}_{\mathrm{1, nuc}}, \hat{\gamma}_{\mathrm{1, nuc}}, \hat{\Theta}_{\mathrm{nuc}}\right) = \mathop{\mathrm{argmin}}_{\substack{ \beta \in \mathbb{R}^{d_X}, \Theta\in \mathbb{R}^{N\times N} \\ \lambda_1 \in \mathbb{R}^{N}, \gamma_1 \in \mathbb{R}^{N} }} \left\{ \mathcal{L}_{N}\left(\beta, \lambda_1, \gamma_1, \Theta\right) + \frac{\varphi_{N}}{\sqrt{N(N-1)}} \|\Theta\|_{\mathrm{nuc}}\right\}, \end{align}\] where \(\|\Theta\|_{\mathrm{nuc}}\) denotes the nuclear norm of matrix \(\Theta\), and \(\varphi_{N} > 0\) is a regularization parameter. The nuclear norm regularized estimators \((\hat{\Lambda}_{\mathrm{2, nuc}}, \hat{\Gamma}_{\mathrm{2, nuc}})\) are obtained through the singular value decomposition of \(\hat{\Theta}_{\mathrm{nuc}}\) as in 5 .
The local estimator \((\hat{\beta}_{\mathrm{local}},\hat{\lambda}_{\mathrm{1, local}}, \hat{\gamma}_{\mathrm{1, local}}, \hat{\Lambda}_{\mathrm{2, local}}, \hat{\Gamma}_{\mathrm{2, local}})\) solves \[\begin{align} (\hat{\beta}_{\mathrm{local}},\hat{\lambda}_{\mathrm{1, local}}, \hat{\gamma}_{\mathrm{1, local}}, \hat{\Lambda}_{\mathrm{2, local}}, \hat{\Gamma}_{\mathrm{2, local}} ) = \mathop{\mathrm{argmin}}_{\substack{ \beta \in \mathbb{R}^{d_X}, \lambda_1 \in \mathbb{R}^{N}, \gamma_1 \in \mathbb{R}^{N} , \\ \Lambda_2 \in \mathbb{R}^{N\times R}, \Gamma_2 \in \mathbb{R}^{N\times R} }} \mathcal{L}_{N}\left(\beta, \lambda_1, \gamma_1, \Lambda_2, \Gamma_2 \right) \end{align}\] using the NNR estimator as the initial value.
We now provide the algorithm for obtaining the NNR estimators with additive fixed effects as follows.
Figure 4:
.
Regarding the choice of step sizes, we recommend starting each iteration with step sizes \(s_{\beta} = 1\), \(s_{\lambda, \gamma} = N\), and \(s_{\theta} = N^2\). If the objective function increases, the step sizes are iteratively halved until a decrease in the objective function is achieved.
Figure 5:
.
In practice, in each step, we recommend starting with \(s_{\beta} = 1\), \(s_{\lambda, \gamma} = s_{\lambda_2} = s_{\gamma_2} = N\). If the objective function increases, the step sizes are iteratively halved until a decrease in the objective function is achieved.
In this section, we provide implementation details for the analytical bias correction and the split-panel jackknife bias correction.
We follow [1] to perform analytical bias correction. For any \(d = 1,2,\ldots, d_X\), let \[\begin{align} (\Lambda^*_{d}, \Gamma^*_{d}) \in \mathop{\mathrm{argmin}}_{\Lambda_{d}\in \mathbb{R}^{N\times R}, \Gamma_{d}\in \mathbb{R}^{T\times R}}\sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}(-\ddot{\ell}^0_{it})\left(\frac{\mathbb{E}(\ddot{\ell}^0_{it}X_{it, d})}{\mathbb{E}(\ddot{\ell}^0_{it})} - \lambda_{d, i}' \gamma_{0, t} - \lambda_{0, i}' \gamma_{d, t}\right)^2, \end{align}\] and define \[\begin{align} \Xi_{d, it} := \lambda_{d, i}^{*\prime} \gamma_{0, t} + \lambda_{0, i}' \gamma^*_{d, t}, \quad \tilde{X}_{d, it} := X_{d, it} - \Xi_{d, it}. \end{align}\] Here, \(\ddot{\ell}_{it}^0: = \partial^2 \log f(\beta_0'X_{it} + \lambda_0'\gamma_0 )/\partial Y^{*2}\) denotes the second derivative of the log-likelihood \(\ell_{it}(\cdot)\) with respect to the index, evaluated at the true parameter values. In addition, let \(\hat{\Xi}_{it}\) denote the sample analogue of \(\Xi_{it}\), and define \[\begin{align} \widehat{B} & := -\frac{1}{N} \sum_{i=1}^{N} \sum_{t=1}^{T}\widehat{\gamma}'_{\mathrm{local}, t} \left(\sum_{\tau=1}^{T}\widehat{\gamma}_{\mathrm{local}, \tau}\widehat{\gamma}_{\mathrm{local}, \tau}'\widehat{\ddot{\ell}}_{i\tau}\right)^{-1}\widehat{\gamma}_{\mathrm{local}, t} \left( \widehat{\dot{\ell}}_{it}\widehat{\ddot{\ell}}_{it}(X_{it} - \widehat{\Xi}_{it}) + \frac{1}{2}\widehat{\dddot{\ell}}_{it}(X_{it} - \widehat{\Xi}_{it})\right), \\ \widehat{D} & := -\frac{1}{T} \sum_{t=1}^{T} \sum_{i=1}^{N} \widehat{\lambda}'_{\mathrm{local}, i} \left(\sum_{j=1}^{N}\widehat{\lambda}_{\mathrm{local}, j}\widehat{\lambda}_{\mathrm{local}, j}'\widehat{\ddot{\ell}}_{jt}\right)^{-1}\widehat{\lambda}_{\mathrm{local}, i}\left(\widehat{\dot{\ell}}_{it}\widehat{\ddot{\ell}}_{it}(X_{it} - \widehat{\Xi}_{it}) + \frac{1}{2}\widehat{\dddot{\ell}}_{it}(X_{it} - \widehat{\Xi}_{it})\right), \\ \widehat{W} & := -\frac{1}{NT} \sum_{i=1}^{N}\sum_{t=1}^{T} \widehat{\ddot{\ell}}_{it}(X_{it} - \widehat{\Xi}_{it})(X_{it} - \widehat{\Xi}_{it})'. \end{align}\] Here, \(\widehat{\dot{\ell}}_{it}\), \(\widehat{\ddot{\ell}}_{it}\), \(\widehat{\dddot{\ell}}_{it}\) denote the first, second, and third derivatives of the log-likelihood \(\ell_{it}(\cdot)\) with respect to the index, evaluated at the local estimator \((\hat{\beta}_{\mathrm{local}}, \hat{\Lambda}_{\mathrm{local}}, \hat{\Gamma}_{\mathrm{local}})\). The analytical bias correction estimator, \(\hat{\beta}_{\mathrm{ABC}} := \hat{\beta}_{\mathrm{local}} - \frac{1}{T} \widehat{W}^{-1} \widehat{B} - \frac{1}{N} \widehat{W}^{-1} \widehat{D}\), follows \[\begin{align} \sqrt{NT}\left(\hat{\beta}_{\mathrm{ABC}} - \beta_0\right)\stackrel{d}{\longrightarrow}N(0, \widehat{W}^{-1}). \end{align}\]
The split-panel jackknife estimator in [1] is given by \[\begin{align} \hat{\beta}_{JBC} := 3 \hat{\beta}_{\mathrm{local}} - \overline{\beta}_{N, T/2} - \overline{\beta}_{N/2, T}, \end{align}\] where \(\overline{\beta}_{N, T/2}\) is the average of the estimators in the half-panels \(\{(i, t)\mid i=1, \ldots, N, t = 1,\ldots, \lceil T/2 \rceil \}\) and \(\{(i, t)\mid i=1, \ldots, N, t = \lceil T/2 \rceil \;+ 1,\ldots, T\}\), \(\overline{\beta}_{N/2, T}\) is the average of the estimators in the half-panels \(\{(i, t)\mid i=1, \ldots, \lceil N/2 \rceil , t = 1,\ldots, T\}\) and \(\{(i, t)\mid i=\lceil N/2 \rceil + 1\ldots, N , t = 1,\ldots, T\}\).
We evaluate the performance of our estimator in a binary response model with predetermined covariates, where the data are generated from the following single-index model with \(R = 2\), \[\label{eq:logit95dynamic} \begin{align} Y_{it} & = \boldsymbol{1}\left(\beta_{1} Y_{i,t-1} + \beta_{ 2}Z_{it} + \lambda_{i}' \gamma_{t} + \epsilon_{Y, it}>0\right), \\ Z_{it} & = \lambda_{i}' \gamma_{t} + \lambda_{i}' \iota + \gamma_{ t}' \iota + \lambda_{z, i} \gamma_{z, t} + \epsilon_{Z, it}. \end{align}\tag{12}\] Here, \(\lambda_{i} = (\lambda_{i1}, \lambda_{i2})'\), \(\gamma_{t} = (\gamma_{t1}, \gamma_{t2})'\). \(\{ \lambda_{ir}\}_{ 1\leq i\leq N, r=1,2}\), \(\{ \gamma_{tr}\}_{1\leq t\leq T, r=1,2}\), \(\{ \lambda_{z, i}\}_{ 1\leq i\leq N}\), and \(\{ \gamma_{z, t}\}_{ 1\leq t\leq T}\) consist of independent random variables drawn from the standard normal distribution \(N(0, 1)\). The error terms \(\{\epsilon_{Y, it}\}_{1\leq i\leq N, 1\leq t\leq T}\) are i.i.d. random variables over both \(i\) and \(t\) following the standard logistic distribution, and \(\{\epsilon_{Z, it}\}_{1\leq i\leq N, 1\leq t\leq T}\) are i.i.d. random variables over both \(i\) and \(t\) following the normal distribution \(N(0, 4)\). We set \(R_{\max} = 5\) and \((\beta_{1}, \beta_2) = (0.5, 0.2)\). Conditional on \(\Lambda_0\) and \(\Gamma_0\), \(Z_{it}\) is an exogenous covariate while the lagged outcome \(Y_{it-1}\) is a predetermined variable.
We conduct \(1000\) Monte Carlo replications to evaluate the finite-sample performance of our estimator across different sample sizes, ranging from \((N,T) = (50,40)\) to \((N,T) = (1000,200)\). We also report alternative estimators for comparison, as well as bias-corrected estimators to examine whether bias corrections based on our two-step estimator mitigate the incidental parameter problem.
The estimators considered in the Monte Carlo experiments include a pooled estimator that ignores individual and time latent factors (POOL); our first-step nuclear norm-regularized estimator (NNR), defined in 4 ; our two-step estimator using the true number of factors \(R\) (\(\mathrm{TS}^*\)); and its analytical bias-corrected estimator (\(\mathrm{ABC}^*\)). We also report the corresponding two-step estimator using an estimated number of factors \(\hat{R}\) (\(\mathrm{TS}\)), along with its analytical bias-corrected counterpart (\(\mathrm{ABC}\)). We report both the bias and the standard deviation of each estimator. In addition, we provide the average estimated number of factors \(\hat{R}\).
It is worth noting that we do not report results for the split-panel jackknife bias correction method, as it requires an additional homogeneity condition ([29], Assumption 4.3), which we consider too restrictive in dynamic settings. For the analytical correction, we directly apply the analytical bias correction method proposed in [1] to the dynamic setting, despite their method not being specifically designed for dynamic models. Developing a valid analytical bias correction formula for dynamic nonlinear panel models with interactive fixed effects is challenging and beyond the scope of this paper.
Table 5 reports results for dynamic Logit models. The first column (POOL) presents the pooled regression results, which exhibit substantial bias that does not diminish as the sample size increases. The second column (NNR) reports the performance of the NNR estimator, whose bias decreases slowly toward zero as \(N, T \to \infty\), consistent with Theorem 1. Table 5 shows that our proposed two-step estimator using the true number of factors (\(R = 2\)), reported in column \(\mathrm{TS}^*\), substantially improves upon the NNR estimator: its bias is smaller and converges rapidly to zero as \(N\) and \(T\) increase. Nevertheless, due to the incidental parameter problem, the bias and standard deviation of the \(\mathrm{TS}^*\) estimator remain of the same order even for large sample sizes such as \((N, T) = (1000, 200)\). The table also shows that applying analytical bias correction to our two-step estimator effectively reduces bias. The analytical bias-corrected estimator (\(\mathrm{ABC}^*\)) significantly reduces bias even in small samples, such as \((N, T) = (50, 40)\). In addition, we find that the analytical bias correction applied to the \(\mathrm{TS}^*\) estimator is more effective in reducing the bias of the estimator of the dynamic effect \(\beta_1\) than that of \(\beta_2\).
When the number of factors is estimated rather than known, the corresponding estimators (\(\mathrm{TS}\), and \(\mathrm{ABC}\)) continue to perform well. Since the number of factors can be estimated with high accuracy (see the last column \(\bar{R}\)), the performance of the estimators based on \(\hat{R}\) is very similar to that obtained using the true \(R\). In some cases, the bias-corrected estimator using \(\hat{R}\) even performs better than the bias-corrected estimator that uses the true \(R\). For example, when \(N = 200, T = 40\), the bias of \(\hat{\beta}_1\) for \(\mathrm{ABC}^*\) is \(-0.47 \times 10^{-2}\), whereas the bias for \(\mathrm{ABC}\) is \(0.09 \times 10^{-2}\). We attribute this observation to small-sample effects. Overall, the two-step estimator delivers strong finite-sample performance and, when combined with bias correction, achieves substantial bias reduction and supports valid inference.
| POOL | NNR | \(\mathrm{TS}^*\) | \(\mathrm{ABC}^*\) | \(\mathrm{TS}\) | ABC | \(\overline{R}\) | ||
| \((\times 10^{-2})\) | \((\times 10^{-2})\) | \((\times 10^{-2})\) | \((\times 10^{-2})\) | \((\times 10^{-2})\) | \((\times 10^{-2})\) | |||
| N = 50, T = 40 | ||||||||
| BIAS | \(\beta_1\) | -11.99 | -10.65 | 2.25 | -0.90 | 2.68 | -0.37 | 1.94 |
| STD | (12.75) | (8.32) | (10.75) | (10.37) | (10.93) | (10.26) | ||
| BIAS | \(\beta_2\) | 7.53 | 6.97 | 5.57 | 4.04 | 5.61 | 3.75 | |
| STD | (2.38) | (2.34) | (3.69) | (3.48) | (3.81) | (3.52) | ||
| N = 100, T = 40 | ||||||||
| BIAS | \(\beta_1\) | -12.32 | -10.36 | 2.03 | -0.53 | 2.44 | 0.00 | 1.997 |
| STD | (11.27) | (6.10) | (6.76) | (6.52) | (6.74) | (6.42) | ||
| BIAS | \(\beta_2\) | 7.53 | 6.66 | 2.89 | 1.71 | 2.93 | 1.49 | |
| STD | (1.98) | (1.86) | (2.30) | (2.23) | (2.27) | (2.18) | ||
| N = 200, T = 40 | ||||||||
| BIAS | \(\beta_1\) | -12.12 | -10.28 | 1.55 | -0.47 | 1.89 | -0.09 | 1.988 |
| STD | (10.37) | (5.08) | (4.79) | (4.71) | (4.82) | (4.56) | ||
| BIAS | \(\beta_2\) | 7.48 | 6.37 | 1.84 | 0.85 | 1.94 | 0.73 | |
| STD | (1.66) | (1.53) | (1.70) | (1.69) | (1.68) | (1.66) | ||
| N = 100, T = 100 | ||||||||
| BIAS | \(\beta_1\) | -12.12 | -9.81 | 1.53 | 0.09 | 1.53 | 0.09 | 2.000 |
| STD | (7.03) | (3.89) | (3.87) | (3.77) | (3.87) | (3.77) | ||
| BIAS | \(\beta_2\) | 7.47 | 5.95 | 1.17 | 0.27 | 1.17 | 0.27 | |
| STD | (1.45) | (1.29) | (1.31) | (1.26) | (1.31) | (1.26) | ||
| N = 200, T = 100 | ||||||||
| BIAS | \(\beta_1\) | -12.25 | -9.25 | 1.42 | 0.18 | 1.42 | 0.18 | 2.000 |
| STD | (6.64) | (3.02) | (2.72) | (2.63) | (2.72) | (2.63) | ||
| BIAS | \(\beta_2\) | 7.37 | 5.46 | 0.80 | 0.14 | 0.80 | 0.14 | |
| STD | (1.19) | (1.01) | (0.86) | (0.83) | (0.86) | (0.83) | ||
| N = 200, T = 200 | ||||||||
| BIAS | \(\beta_1\) | -11.95 | -8.38 | 1.02 | 0.15 | 1.02 | 0.15 | 2.000 |
| STD | (4.48) | (1.97) | (1.83) | (1.79) | (1.83) | (1.79) | ||
| BIAS | \(\beta_2\) | 7.43 | 4.94 | 0.48 | 0.04 | 0.48 | 0.04 | |
| STD | (0.93) | (0.74) | (0.60) | (0.59) | (0.60) | (0.59) | ||
| N = 1000, T = 200 | ||||||||
| BIAS | \(\beta_1\) | -12.11 | -7.11 | 0.52 | -0.04 | 0.52 | -0.04 | 2.000 |
| STD | (4.09) | (1.16) | (0.80) | (0.79) | (0.80) | (0.79) | ||
| BIAS | \(\beta_2\) | 7.40 | 4.13 | 0.28 | 0.03 | 0.28 | 0.03 | |
| STD | (0.71) | (0.47) | (0.26) | (0.26) | (0.26) | (0.26) |
For readability, the main text presents simplified definitions of the NNR estimators, FE estimators, and shrinking neighborhoods. We begin by providing their fully rigorous definitions here. The NNR estimators \((\hat{\beta}_{\mathrm{nuc}}, \hat{\Theta}_{\mathrm{nuc}})\) solve the following optimization problem with a nuclear-norm penalty: \[\label{eq:nnr95definition95formal} \begin{gather} \left(\hat{\beta}_{\mathrm{nuc}}, \hat{\Theta}_{\mathrm{nuc}}\right) = \mathop{\mathrm{argmin}}_{\beta \in \mathbb{R}^{d_X}, \Theta\in \mathbb{R}^{N\times T}} \left\{ \mathcal{L}_{NT} \left( \beta, \Theta\right) + \frac{\varphi_{NT}}{\sqrt{NT}} \|\Theta\|_{\mathrm{nuc}}\right\}, \\ \text{s.t. } \quad \|\beta\|_{\max} \leq \rho_{\beta}, \quad \|\Theta\|_{\max} \leq \rho_{\theta}. \end{gather}\tag{13}\] Here, we impose additional constraints \(\|\beta\|_{\max} \leq \rho_{\beta}\) and \(\|\Theta\|_{\max} \leq \rho_{\theta}\) compared to optimization 4 . These constraints, which are standard in extremum estimation, ensure that optimization is conducted within a compact parameter space. In addition, the FE estimators \((\hat{\beta}_{\mathrm{FE}}, \hat{\Lambda}_{\mathrm{FE}}, \hat{\Gamma}_{\mathrm{FE}})\) solve \[\label{eq:FE95estimator95formal} \begin{gather} (\hat{\beta}_{\mathrm{FE}}, \hat{\Lambda}_{\mathrm{FE}}, \hat{\Gamma}_{\mathrm{FE}}) \in \mathop{\mathrm{argmin}}_{ \beta, \Lambda, \Gamma} \mathcal{L}_{NT}(\beta, \Lambda, \Gamma), \\ \text{s.t. }\quad \|\beta\|_{\max} \leq \rho_{\beta}, \quad \|\Lambda\|_{\max} \leq \rho_{\lambda}, \quad \|\Gamma\|_{\max} \leq \rho_{\gamma}. \end{gather}\tag{14}\] in which we impose additional constraints to 2 , \(\|\Lambda\|_{\max} \leq \rho_{\lambda}\) and \(\|\Gamma\|_{\max} \leq \rho_{\gamma}\), which are standard in extremum estimation and ensure that the optimization is conducted over a compact parameter space. We also impose such compactness conditions on the definition of neighborhoods \(\mathcal{B}_{\delta_{NT}}\) introduced in 7 : \[\begin{align} \label{eq:neighborhood95formal} \mathcal{B}_{\delta_{NT}} = \bigg\{(\beta, \Lambda, \Gamma)\mid & \|\beta-\beta_0\|, \frac{1}{\sqrt{N}}\|\Lambda - {\Lambda}_0 \|_{\mathrm{F}} , \frac{1}{\sqrt{T}}\|\Gamma - {\Gamma}_0 \|_{\mathrm{F}} \leq \delta_{NT} \text{, } \|\Lambda\|_{\max} \leq \rho_{\lambda}\text{, } \|\Gamma\|_{\max} \leq \rho_{\gamma}\bigg\}, \end{align}\tag{15}\] and we need to make sure the NNR estimator falls within the shrinking neighborhood \(\mathcal{B}_{\delta_{NT}}\) wpa1. One potential concern of adding compactness constraints in 15 is that the entries of \((\hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) are not necessarily uniformly bounded. This is not a substantive issue, as we can truncate and normalize \((\hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) to obtain nuisance estimators that satisfy the uniform boundedness condition. The details of this procedure, along with a proof demonstrating that it does not affect the theoretical results, are provided later.
We introduce some new notations for simplicity. For any two positive real sequences \(\left\{a_n\right\}_{n \geq 1}\) and \(\left\{b_n\right\}_{n\geq 1}\), we use \(a_n\lesssim b_n\) (\(a_n\gtrsim b_n\)) to denote that there exists a positive constant \(c\) such that \(a_n\leq cb_n\) (\(a_n \geq cb_n\)) for all \(n\). We write \(a_n\asymp b_n\) if both \(a_n\lesssim b_n\) and \(a_n\gtrsim b_n\). Recall that there exist constants \((\rho_{\beta}, \rho_{\lambda}, \rho_{\gamma}, \rho_{\theta})\) such that \[\begin{align} \|\beta\|_{\max}\leq \rho_{\beta}, \quad \|\Lambda_0\|_{\max} \leq \rho_{\lambda}, \quad \|\Gamma_0\|_{\max} \leq \rho_{\gamma}, \quad \|\Theta\|_{\max} \leq \rho_{\theta}. \end{align}\] We define the estimation errors as \[\begin{align} \hat{\Delta}_{\beta } := \hat{\beta}_{\mathrm{nuc}} - \beta_0, \quad \hat{\Delta}_{\Lambda } := \hat{\Lambda}_{\mathrm{nuc}} - \Lambda_0, \quad \hat{\Delta}_{\Gamma} := \hat{\Gamma}_{\mathrm{nuc}} - \Gamma_0, \quad \hat{\Delta}_{\Theta } := \hat{\Theta}_{\mathrm{nuc}} - \Theta_0. \end{align}\] It follows from 13 that \[\begin{align} \|\hat{\Delta}_{\beta}\|_{\max} \leq 2\rho_{\beta}, \quad \|\hat{\Delta}_{\Theta}\|_{\max} \leq 2\rho_{\theta}. \end{align}\]
The following lemma collects basic properties of low-rank projections that are used frequently in the subsequent proof.
Lemma 4. Let \(\Delta\) be any \(N\times T\) matrix. The following properties hold:
\(\|\Delta\|_{\mathrm{nuc}} = \|M_{\Lambda_0}\Delta M_{\Gamma_0} \|_{\mathrm{nuc}} + \|\Delta - M_{\Lambda_0}\Delta M_{\Gamma_0}\|_{\mathrm{nuc}}\).
\(\|\Delta\|_{\mathrm{F}}^2 = \|M_{\Lambda_0}\Delta M_{\Gamma_0} \|_{\mathrm{F}}^2 + \|\Delta - M_{\Lambda_0}\Delta M_{\Gamma_0}\|_{\mathrm{F}}^2\).
\(\mathrm{rank}(\Delta - M_{\Lambda_0}\Delta M_{\Gamma_0})\leq 2R\).
The proof is omitted. Readers may refer to [6] and [38] for details.
We provide only the proof of Theorem 5, as it extends Theorem 1 by incorporating predetermined covariates.
Proof of Theorem 5. The proof is based on the following lemma:
Lemma 5. Under the conditions of Theorem 5, \[\begin{align} \| M_{\Lambda_0} \hat{\Delta}_{\Theta} M_{\Gamma_0}\|_{\mathrm{nuc}} \leq \frac{2+\alpha}{\alpha}\left(\sqrt{NT}\|\hat{\Delta}_{\beta}\| + \| \hat{\Delta}_{\Theta} - M_{\Lambda_0} \hat{\Delta}_{\Theta} M_{\Gamma_0} \|_{\mathrm{nuc}} \right). \end{align}\]
The proof of Lemma 5 is standard and will be presented at the end of this subsection. Lemma 5 states that, when the penalization parameter \(\varphi_{NT}\) is sufficiently large, the component of the estimation error of \(\Theta\) that cannot be explained by either \(\Lambda_0\) or \(\Gamma_0\) is relatively small compared to the part of the estimation error of \(\Theta\) that can be explained by \(\Lambda_0\) and \(\Gamma_0\), along with a term that accounts for the estimation error of \(\beta\).
When \(\|\hat{\Delta}_{\beta}\|^2 + \frac{1}{NT}\|\hat{\Delta}_{\Theta}\|_\mathrm{F}^2 \leq \sqrt{\frac{\log (NT)}{NT}}\), we directly obtain \[\begin{align} \|\hat{\Delta}_{\beta}\| & \leq \log (NT) /\sqrt{ \min\{N, T\}}, \quad \frac{1}{\sqrt{NT}}\|\hat{\Delta}_{\Theta}\|_{\mathrm{F}} \leq \log (NT) /\sqrt{ \min\{N, T\}}. \end{align}\] When \(\|\hat{\Delta}_{\beta}\|^2 + \frac{1}{NT}\|\hat{\Delta}_{\Theta}\|_\mathrm{F}^2 > \sqrt{\frac{\log (NT)}{NT}}\), the proof becomes more intricate and requires additional effort.
Since \((\hat{\beta}_{\mathrm{nuc}}, \hat{\Theta}_{\mathrm{nuc}})\) solves optimization problem 13 , we have \[\begin{align} \label{eq:thm95consistency951} \mathcal{L}_{NT}( \beta_0 + \hat{\Delta}_{\beta} , \Theta_0 + \hat{\Delta}_{\Theta} ) - \mathcal{L}_{NT}(\beta_0, \Theta_0) \leq \frac{\varphi_{NT}}{\sqrt{NT}} (\|\Theta_0\|_{\mathrm{nuc}} - \|\Theta_0 + \hat{\Delta}_{\Theta} \|_{\mathrm{nuc}} ). \end{align}\tag{16}\] Consider the Taylor expansion of \(\mathcal{L}_{NT}(\hat{\beta}_{\mathrm{nuc}}, \hat{\Theta}_{\mathrm{nuc}})\) around \((\beta_0, \Theta_0)\): \[\label{eq:thm95consistency952} \begin{align} & \mathcal{L}_{NT}( \beta_0 + \hat{\Delta}_{\beta} , \Theta_0 + \hat{\Delta}_{\Theta} ) - \mathcal{L}_{NT}(\beta_0, \Theta_0) - \nabla_{\beta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)'\hat{\Delta} _{\beta} - \langle\nabla_{\Theta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right), \hat{\Delta}_{\Theta} \rangle\\ \stackrel{\text{(i)}}{\leq} & \frac{\varphi_{NT}}{\sqrt{NT}} \underbrace{(\|\Theta_0\|_{\mathrm{nuc}} - \| \Theta_0 + \hat{\Delta}_{\Theta} \|_{\mathrm{nuc}} ) }_{\leq \|\hat{\Delta}_{\Theta}\|_{\mathrm{nuc}}} - \underbrace{\nabla_{\beta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)'\hat{\Delta} _{\beta}}_{\leq \|\nabla_{\beta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)\| \|\hat{\Delta} _{\beta}\| } - \underbrace{\langle\nabla_{\Theta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right), \hat{\Delta}_{\Theta} \rangle}_{\leq \|\nabla_{\Theta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)\|_{\mathrm{op}}\| \hat{\Delta}_{\Theta}\|_{\mathrm{nuc}}} \\ \stackrel{\text{(ii)}}{\leq} & \frac{\varphi_{NT}}{\sqrt{NT}}\|\hat{\Delta}_{\Theta}\|_{\mathrm{nuc}} + \|\nabla_{\beta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)\| \|\hat{\Delta} _{\beta}\| + \|\nabla_{\Theta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)\|_{\mathrm{op}}\| \hat{\Delta}_{\Theta}\|_{\mathrm{nuc}} \\ \stackrel{\text{(iii)}}{\leq} & \frac{\varphi_{NT}}{\sqrt{NT}}\|\hat{\Delta}_{\Theta}\|_{\mathrm{nuc}} + \frac{ \varphi_{NT}}{1+\alpha} \|\hat{\Delta} _{\beta}\| + \frac{\varphi_{NT}}{(1+\alpha)\sqrt{NT}} \| \hat{\Delta}_{\Theta}\|_{\mathrm{nuc}} \\ \leq & \frac{2 \varphi_{NT}}{\sqrt{NT}}\left( \|\hat{\Delta}_{\Theta}\|_{\mathrm{nuc}} + \sqrt{NT} \|\hat{\Delta}_{\beta}\| \right), \end{align}\tag{17}\] where inequality (i) follows from inequality (16 ), inequality (ii) employs the triangle inequality, Cauchy-Schwarz inequality, and Hölder’s inequality (since the spectral norm is the dual norm of the nuclear norm), and inequality (iii) holds because of \(\varphi_{NT} \geq (1+\alpha) \max\{\|\nabla_{\beta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)\|, \sqrt{NT}\|\nabla_{\Theta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)\|_{\mathrm{op}}\}\), as stated in Theorem 5. In addition, we have the following inequalities for the nuclear norm: \[\begin{align} \|\hat{\Delta}_{\Theta}\|_{\mathrm{nuc}} & \stackrel{\text{(i)}}{=} \| M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0} \| _{\mathrm{nuc}} + \| \hat{\Delta}_{\Theta} - M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0} \| _{\mathrm{nuc}} \\ & \stackrel{\text{(ii)}}{\leq} \frac{2+\alpha}{\alpha}\left(\sqrt{NT}\|\hat{\Delta}_{\beta}\| + \| \hat{\Delta}_{\Theta} - M_{\Lambda_0} \hat{\Delta}_{\Theta} M_{\Gamma_0} \|_{\mathrm{nuc}} \right) + \| \hat{\Delta}_{\Theta} - M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0} \| _{\mathrm{nuc}} \\ & \stackrel{\text{(iii)}}{\leq} \frac{2+\alpha}{\alpha} \sqrt{NT}\|\hat{\Delta}_{\beta}\| + \frac{2(1+\alpha)\sqrt{2R}}{\alpha}\|\hat{\Delta}_{\Theta} - M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0}\|_{\mathrm{F}} \\ & \stackrel{\text{(iv)}}{\leq} \frac{2+\alpha}{\alpha} \sqrt{NT}\|\hat{\Delta}_{\beta}\| + \frac{2(1+\alpha)\sqrt{2R}}{\alpha}\|\hat{\Delta}_{\Theta} \|_{\mathrm{F}} \\ & \leq \frac{2(1+\alpha)\sqrt{2R}}{\alpha}( \sqrt{NT}\|\hat{\Delta}_{\beta}\| + \|\hat{\Delta}_{\Theta} \|_{\mathrm{F}}), \end{align}\] where equality (i) holds due to Lemma 4[item:local95projection951], inequality (ii) follows from Lemma 5, inequality (iii) follows from Lemma 4[item:local95projection953], and inequality (iv) is based on Lemma 4[item:local95projection952].
Therefore, combining the nuclear norm inequality with inequality 17 yields \[\label{eq:thm95consistency953} \begin{align} & \mathcal{L}_{NT}( \beta_0 + \hat{\Delta}_{\beta} , \Theta_0 + \hat{\Delta}_{\Theta} ) - \mathcal{L}_{NT}(\beta_0, \Theta_0) - \nabla_{\beta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)'\hat{\Delta} _{\beta} - \langle\nabla_{\Theta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right), \hat{\Delta}_{\Theta} \rangle\\ \leq & \frac{8(1+\alpha)\sqrt{2R}}{\alpha}\frac{\varphi_{NT}}{\sqrt{NT}} ( \sqrt{NT}\|\hat{\Delta}_{\beta}\| + \|\hat{\Delta}_{\Theta} \|_{\mathrm{F}}) \\ \leq & \frac{16(1+\alpha)\sqrt{2R}}{\alpha} \varphi_{NT} \sqrt{ \|\hat{\Delta}_{\beta}\|^2 + \frac{1}{NT}\|\hat{\Delta}_{\Theta} \|^2_\mathrm{F}}. \end{align}\tag{18}\]
By the convexity of \(\mathcal{L}_{NT}(\cdot, \cdot)\), we have \[\label{eq:thm95consistency954} \begin{align} & \mathcal{L}_{NT}(\beta_0 + \hat{\Delta}_{\beta}, \Theta_0 + \hat{\Delta}_{\Theta}) - \mathcal{L}_{NT}(\beta_0, \Theta_0) - \nabla_{\beta}\mathcal{L}_{NT}(\beta_0, \Theta_0)'\hat{\Delta}_{\beta} - \langle\nabla_{\Theta}\mathcal{L}_{NT}(\beta_0, \Theta_0), \hat{\Delta}_{\Theta} \rangle\\ \stackrel{\text{(i)}}{=} & \frac{1}{2NT}\sum_{i=1}^{N}\sum_{t=1}^{T}(-\ddot{\ell}_{it}(X_{it}'\tilde{\beta} + \tilde{\theta}_{it} ))(X_{it}' \hat{\Delta}_{\beta} + \hat{\Delta}_{\theta_{it}} )^2 \\ \stackrel{\text{(ii)}}{\geq} & \frac{{b}_{\min}}{2} \frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T}(X_{it}' \hat{\Delta}_{\beta} + \hat{\Delta}_{\theta_{it}} )^2 \\ \stackrel{\text{(iii)}}{\geq} & \frac{{b}_{\min}}{2} \left(\kappa \left(\|\hat{\Delta}_{\beta}\|^2 + \frac{1}{NT}\|\hat{\Delta}_{\Theta}\|_{\mathrm{F}}^2\right) - \eta \frac{N+T}{NT}(\log(NT))^2\right), \end{align}\tag{19}\] where equality (i) follows from the second-order Taylor expansion of \(\mathcal{L}_{NT}(\cdot, \cdot)\), inequality (ii) holds because \(-\ddot{\ell}_{it}(\cdot)\geq b_{\min}\) uniformly as stated in Assumption 4[item:smoothing95pre], and inequality (iii) follows from Lemma 5 and the RSC (Assumption 2).
By combining 18 and 19 , we obtain \[\begin{align} \frac{{b}_{\min}}{2} \left(\kappa \left(\|\hat{\Delta}_{\beta}\|^2 + \frac{1}{NT}\|\hat{\Delta}_{\Theta}\|_{\mathrm{F}}^2\right) - \eta \frac{N+T}{NT} (\log(NT))^2 \right) &\leq \frac{16(1+\alpha)\sqrt{2R}}{\alpha} \varphi_{NT} \sqrt{ \|\hat{\Delta}_{\beta}\|^2 + \frac{1}{NT}\|\hat{\Delta}_{\Theta} \|^2_{\mathrm{F}}} \\ \Rightarrow \sqrt{ \|\hat{\Delta}_{\beta}\|^2 + \frac{1}{NT}\|\hat{\Delta}_{\Theta}\|_{\mathrm{F}}^2 } & \leq \frac{a_1 \varphi_{NT} + \sqrt{a_1^2 \varphi_{NT}^2 + 4a_2^2 \frac{N+T}{NT}(\log(NT))^2}}{2} \\ & \leq a_1 \varphi_{NT} + a_2 \sqrt{\frac{N+T}{NT}}(\log(NT))^2 \\ & \leq a_1 \varphi_{NT} + \sqrt{2} a_2 \frac{\log(NT)}{\sqrt{\min\{N, T\}}}. \end{align}\] Here, \(a_1 := \frac{32(1+\alpha)\sqrt{2R}}{\alpha b_{\min}\kappa } > 0\) and \(a_2 := \sqrt{\frac{\eta}{\kappa}} >0\). Let \(c_1 := \max\{a_1, \sqrt{2} a_2\}\). It is straightforward to show that, wpa1, \[\begin{align} \| \hat{\beta}_{\mathrm{nuc}} - \beta_0\| & \leq c_1 \left(\varphi_{NT} + \log (NT)/\sqrt{\min\{N, T\}}\right), \\ \frac{1}{\sqrt{NT}}\|\hat{\Theta}_{\mathrm{nuc}} - \Theta_0\|_{\mathrm{F}} & \leq c_1 \left(\varphi_{NT} + \log (NT)/\sqrt{\min\{N, T\}}\right) . \end{align}\]
We aim to establish an estimation error bound for \(\hat{\Lambda}_{\mathrm{nuc}}\). The error bound for \(\hat{\Gamma}_{\mathrm{nuc}}\) can be directly obtained by the same method. In addition, our proof below is based on the situation where \(\Sigma_{\lambda}^{1/2}\Sigma_{\gamma}\Sigma_{\lambda}^{1/2}\) may have repeated eigenvalues. For notational simplicity, let \(\Upsilon\) be the \(R\times R\)-dimensional matrix containing the eigenvectors of \(\frac{1}{NT}(\Lambda_0'\Lambda_0)^{1/2}\Gamma_0'\Gamma_0(\Lambda_0'\Lambda_0)^{1/2}\), let \(D\) be the \(R\times R\)-dimensional diagonal matrix containing the square roots of the eigenvalues of \(\frac{1}{NT}(\Lambda_0'\Lambda_0)^{1/2}\Gamma_0'\Gamma_0(\Lambda_0'\Lambda_0)^{1/2}\), \(\Omega\) be the \(R\)-dimensional vector containing the square roots of the eigenvalues of \(\Sigma_{\lambda}^{1/2}\Sigma_{\gamma}\Sigma_{\lambda}^{1/2}\) in non-increasing order, and \(U_0\) be the matrix of left singular vectors of \(\Theta_0\). One can verify that \(U_0 = \Lambda_0(\Lambda_0'\Lambda_0)^{-1/2} \Upsilon\). In addition, let \(\hat{U}\) be the matrix containing the left singular vectors of \(\hat{\Theta}_{\mathrm{nuc}}\) such that \(\hat{U}' \hat{U} = \mathbb{I}_{R}\). Then we have \(\hat{\Lambda}_{\mathrm{nuc}} = \sqrt{N}\hat{U}\hat{D}^{1/2}_{[1:R, 1:R]}\).
We first establish the bound for the distance between two spaces spanned by \(\hat{U}\) and \(U_0\), respectively. Using Davis-Kahan Theorem (Theorem 4 in [39]), there exists an orthogonal matrix \(O^*\) such that, wpa1, \[\label{eq:thm95consistency955} \begin{align} \|\hat{U} - U_0 O^{*\prime}\|_{\mathrm{F}} \stackrel{\text{(i)}}{\leq} & \frac{2^{\frac{3}{2}}(2\|\Theta_0\|_{\mathrm{op}} + \|\hat{\Theta}_{\mathrm{nuc}} - \Theta_0 \|_{\mathrm{op}}) \|\hat{\Theta}_{\mathrm{nuc}} - \Theta_0 \|_{\mathrm{F}} }{\psi^2_{R}\left(\Theta_0\right)} \\ \stackrel{\text{(ii)}}{\leq} & \frac{8 \|\Theta_0\|_{\mathrm{op}} \|\hat{\Theta}_{\mathrm{nuc}} - \Theta_0 \|_{\mathrm{F}} }{\psi^2_{R}\left(\Theta_0\right)}, \end{align}\tag{20}\] where inequality (i) follows from the fact that \(\psi_{R+1}(\Theta_0) = 0\), inequality (ii) is based on the estimation error bound of \(\hat{\Theta}_{\mathrm{nuc}}\), implying that \(\|\hat{\Theta}_{\mathrm{nuc}} - \Theta_0 \|_{\mathrm{op}} / \|\Theta_0\|_{\mathrm{op}} \stackrel{p}{\longrightarrow}0\). The matrix \(O^*\) arises due to the possible multiplicity of eigenvalues and depends only on \((\Lambda_0, \Gamma_0)\). In addition, the strong factor assumption (Assumption 4[item:strong95factors95pre]) implies that \[\label{eq:thm95consistency956} \begin{gather} \frac{1}{\sqrt{NT}}\|\Theta_0\|_{\mathrm{op}} \stackrel{p}{\longrightarrow}\Omega_1, \quad \frac{1}{\sqrt{NT}}\psi_{R}\left(\Theta_0\right) \stackrel{p}{\longrightarrow}\Omega_{R}. \end{gather}\tag{21}\] Combining inequality (20 ) and (21 ) yields that, wpa1, \[\begin{align} \label{eq:thm95consistency957} \|\hat{U} - U_0 O^{*\prime}\|_{\mathrm{F}} & \leq \frac{16\Omega_1}{\sqrt{NT}\Omega_R^2} \|\hat{\Theta}_{\mathrm{nuc}} - \Theta_0 \|_{\mathrm{F}} \leq \frac{16 c_1 \Omega_1}{\Omega_R^2 } \left(\varphi_{NT} + \frac{\log(NT)}{\sqrt{\min\{N, T\}}}\right) . \end{align}\tag{22}\]
Now we turn to the error bound for \(\hat{\Lambda}_{\mathrm{nuc}}\). Note that \[\begin{align} \left\|\hat{\Lambda}_{\mathrm{nuc}} - \sqrt{N}\underbrace{\Lambda_0(\Lambda_0'\Lambda_0)^{-1/2} \Upsilon }_{U_0} O^{*\prime} D^{1/2} \right\|_{\mathrm{F}} = & \sqrt{N} \left\|\hat{U} \hat{D}^{1/2}_{[1:R, 1:R]} - U_0 O^{*\prime} D^{1/2} \right\|_{\mathrm{F}} \\ \leq & \sqrt{N}\left\|\hat{U} - U_0 O^{*\prime} D^{1/2}\hat{D}^{-1/2}_{[1:R, 1:R]} \right\|_{\mathrm{F}} \left\|\hat{D}^{1/2}_{[1:R, 1:R]}\right\|_{\mathrm{F}}\\ \leq & \sqrt{N} \left(\underbrace{\left\|\hat{U} - U_0 O^{*\prime} \right\|_{\mathrm{F}}}_{A_1}\left\|\hat{D}^{1/2}_{[1:R, 1:R]}\right\|_{\mathrm{F}} + \underbrace{\left\| U_0O^{*\prime} \left(\hat{D}^{1/2}_{[1:R, 1:R]} - D^{1/2} \right) \right\|_{\mathrm{F}}}_{A_2}\right) . \end{align}\] Since we have already established the bound for \(A_1\) (see 22 ), we focus on the upper bound of \(A_2\). Note that \[\begin{align} \left\|\hat{D}^{1/2}_{[1:R, 1:R]} - D^{1/2} \right\|_{\mathrm{F}} \leq & \sqrt{R} \frac{\left\|\hat{D}_{[1:R, 1:R]} - D\right\|_{\mathrm{op}}}{2\min\{\psi_{R}^{1/2} \left(\hat{D}_{[1:R, 1:R]}\right), \psi_{R}^{1/2} \left(D\right)\}} \\ \leq & \sqrt{R}\frac{\left\|\hat{D}_{[1:R, 1:R]} - D\right\|_{\mathrm{F}}}{2\min\{\psi_{R}^{1/2} \left(\hat{D}_{[1:R, 1:R]}\right), \psi_{R}^{1/2} \left(D\right)\}} \\ \leq & \frac{\sqrt{R}\|\hat{\Theta}_{\mathrm{nuc}} - \Theta_0\|_{\mathrm{F}}}{\sqrt{NT}\Omega_R^{1/2}}. \end{align}\] Then, the following inequality holds wpa1 \[\begin{align} A_2\leq \underbrace{\|U_0\|_{\mathrm{F}}}_{= \sqrt{R}}\left\|\hat{D}^{1/2}_{[1:R, 1:R]} - D^{1/2} \right\|_{\mathrm{F}} & \leq \frac{R\|\hat{\Theta}_{\mathrm{nuc}} - \Theta_0\|_{\mathrm{F}}}{\sqrt{NT}\Omega_R^{1/2}} \leq \frac{ c_1 R}{\Omega_R^{1/2} } \left(\varphi_{NT} + \frac{\log(NT)}{\sqrt{\min\{N, T\}}}\right). \end{align}\] Let \(G = D^{1/2} O^* \Upsilon'(\Lambda_0'\Lambda_0/N)^{-1/2}\). We obtain \[\begin{align} \left\|\hat{\Lambda}_{\mathrm{nuc}} - \Lambda_0 G' \right\|_{\mathrm{F}} \leq \sqrt{N} \underbrace{\left(\frac{16c_1\Omega_1^{3/2}}{\Omega_R^2} + \frac{ c_1 R}{\Omega_R^{1/2} }\right)}_{B_1} \left(\varphi_{NT} + \frac{\log(NT)}{\sqrt{\min\{N, T\}}}\right), \quad \text{wpa1}. \end{align}\] By the same method, we also show that \[\begin{align} \|\hat{\Gamma}_{\mathrm{nuc}} - \Gamma_0 G^{-1}\|_{\mathrm{F}} \leq \sqrt{T} B_1 \left(\varphi_{NT} + \frac{\log(NT)}{\sqrt{\min\{N, T\}}}\right), \quad \text{wpa1}. \end{align}\] In addition, since \[\begin{align} GG' = \underbrace{D^{1/2}}_{\stackrel{p}{\longrightarrow}\mathrm{diag}(\Omega)^{1/2}} O^* \Upsilon'\underbrace{(\Lambda_0'\Lambda_0/N)^{-1}}_{\stackrel{p}{\longrightarrow}\Sigma_{\lambda}^{-1}}\Upsilon O^{*\prime} \underbrace{D^{1/2}}_{\stackrel{p}{\longrightarrow}\mathrm{diag}(\Omega)^{1/2}} \stackrel{p}{\longrightarrow}\mathrm{diag}(\Omega)^{1/2} O^* \Upsilon' \Sigma_{\lambda}^{-1} \Upsilon O^{*\prime} \mathrm{diag}(\Omega)^{1/2}, \end{align}\] it immediately follows that, wpa1, (i) the maximum singular value of \(G\) is uniformly bounded, and (ii) the minimum singular value of \(G\) is strictly greater than zero. Therefore, this completes the proof of the theorem.
In the following text, we discuss how to construct nuisance estimators with uniformly bounded entries to ensure that \((\hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}}) \in \mathcal{B}_{\delta_{NT}}\). Although this property is not directly related to the current theorem, it will be used frequently in subsequent theoretical discussions.
As we have discussed before, one potential concern with \((\hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) as in 5 is that the entries of \((\hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) are not necessarily uniformly bounded. In the following text, we show that after truncating and normalizing the estimators in 5 , we can obtain new nuisance estimators \((\tilde{\Lambda}_{\mathrm{nuc}}, \tilde{\Gamma}_{\mathrm{nuc}})\) that satisfy the uniform boundedness condition, and consequently, \((\tilde{\Lambda}_{\mathrm{nuc}}, \tilde{\Gamma}_{\mathrm{nuc}})\in \Phi_{NT}\).
It should be noted that constructing uniformly bounded nuisance estimators is solely for the convenience of theoretical analysis. In practice, applied researchers do not need to perform this step.
Our construction is based on the following observation: Since (i) \((\Lambda_0, \Gamma_0)\) is uniformly bounded, (ii) the maximum singular value of \(G\) is uniformly bounded, and (iii) the minimum singular value of \(G\) is strictly greater than zero, each entry of \(\Lambda_0^G\) and \(\Gamma_0^G\) is uniformly bound. Thus, there exists a constant \(M>0\) that is sufficiently large but independent of \(N, T\) such that \(\|\Lambda_0^G\|_{\max}, \|\Gamma_0^G\|_{\max} \leq M\), wpa1.
We first describe how to obtain the new estimators:
For a sufficiently large constant \(M>0\) (independent of \(N, T\)), compute truncated estimators \((\bar{\Lambda}_{\mathrm{nuc}}, \bar{\Gamma}_{\mathrm{nuc}})\), defined as \[\begin{align} &\bar{\Lambda}_{\mathrm{nuc}, ir} = \left\{ \begin{array}{ll} \hat{\Lambda}_{\mathrm{nuc}, ir}, & \text{if} \quad |\hat{\Lambda}_{\mathrm{nuc}, ir}|\leq M \\ M\mathrm{sign}(\hat{\Lambda}_{\mathrm{nuc}, ir}), & \text{otherwise } \end{array} \right. \\ &\bar{\Gamma}_{\mathrm{nuc}, tr} = \left\{ \begin{array}{ll} \hat{\Gamma}_{\mathrm{nuc}, tr}, & \text{if}\quad |\hat{\Gamma}_{\mathrm{nuc}, tr}|\leq M \\ M\mathrm{sign}(\hat{\Gamma}_{\mathrm{nuc}, tr}), & \text{otherwise } \end{array} \right. \end{align}\]
Perform singular value decomposition on \(\bar{\Theta}_{\mathrm{nuc}} := \bar{\Lambda}_{\mathrm{nuc}}\bar{\Gamma}'_{\mathrm{nuc}}\), so that \(\bar{\Theta}_{\mathrm{nuc}}/\sqrt{NT} = \bar{U}\bar{D}\bar{V}\), where \(\bar{U}\in \mathbb{R}^{N\times R}\) and \(\bar{V}\in \mathbb{R}^{T\times R}\) are matrices whose columns are the left and right orthonormal singular vectors of \(\bar{\Theta}_{\mathrm{nuc}}\), respectively, and \(\bar{D}\) is a diagonal matrix whose diagonal entries are singular values of \(\bar{\Theta}_{\mathrm{nuc}}/\sqrt{NT}\) (arranged in non-increasing order). We then compute \((\tilde{\Lambda}_{\mathrm{nuc}}, \tilde{\Gamma}_{\mathrm{nuc}})\) as follows: \[\begin{gather} \tilde{\Lambda}_{\mathrm{nuc}} = \sqrt{N} \bar{U} \bar{D}^{1/2} , \quad \tilde{\Gamma}_{\mathrm{nuc}} = \sqrt{T} \bar{V} \bar{D}^{1/2}. \end{gather}\]
Since the entries of \((\bar{\Lambda}_{\mathrm{nuc}}, \bar{\Gamma}_{\mathrm{nuc}}, \Lambda_0^G, \Gamma_0^G)\) are uniformly bounded, we have the following inequalities wpa1: \[\begin{align} \|\bar{\Lambda}_{\mathrm{nuc}}-\Lambda_0^G\|_{\mathrm{F}} \leq \|\hat{\Lambda}_{\mathrm{nuc}}-\Lambda_0^G\|_{\mathrm{F}} \leq \sqrt{N} B_1 \left(\varphi_{NT} + \frac{\log(NT)}{\sqrt{\min\{N, T\}}}\right), \\ \|\bar{\Gamma}_{\mathrm{nuc}}-\Gamma_0^G\|_{\mathrm{F}} \leq \|\hat{\Gamma}_{\mathrm{nuc}}-\Gamma_0^G\|_{\mathrm{F}} \leq \sqrt{T} B_1 \left(\varphi_{NT} + \frac{\log(NT)}{\sqrt{\min\{N, T\}}}\right). \end{align}\] Thus, \[\begin{align} \|\bar{\Theta}_{\mathrm{nuc}} - \Theta_0\|_{\mathrm{F}} \leq & \underbrace{\|\bar{\Lambda}_{\mathrm{nuc}}\|_{\mathrm{F}}}_{\leq M\sqrt{NR}}\|\bar{\Gamma}_{\mathrm{nuc}} - \Gamma_0^G\|_{\mathrm{F}} + \underbrace{\|\Gamma_0^G\|_{\mathrm{F}}}_{\leq M\sqrt{TR}} \|\bar{\Lambda}_{\mathrm{nuc}} - \Lambda_0^G\|_{\mathrm{F}} \\ \leq & \underbrace{2M\sqrt{R} B_1}_{B_2} \sqrt{NT}\left(\varphi_{NT} + \frac{\log(NT)}{\sqrt{\min\{N, T\}}}\right), \quad \text{wpa1}. \end{align}\] Therefore, the estimation errors of \(\bar{\Theta}_{\mathrm{nuc}}\) and \(\hat{\Theta}_{\mathrm{nuc}}\) differ only by a constant factor. Applying the proof method in Step 3 yields that, wpa1, \[\begin{align} \frac{1}{\sqrt{N}}\|\tilde{\Lambda}_{\mathrm{nuc}} - \Lambda_0\|_{\mathrm{F}}, \frac{1}{\sqrt{T}}\|\tilde{\Gamma}_{\mathrm{nuc}} - \Gamma_0\|_{\mathrm{F}} \leq & \underbrace{\left(\frac{16B_2 \Omega_1^{3/2}}{\Omega_R^2} + \frac{ c_1 R}{\Omega_R^{1/2} }\right)}_{c_2} \left(\varphi_{NT} + \frac{\log(NT)}{\sqrt{\min\{N, T\}}}\right). \end{align}\] In addition, since \(\bar{\Theta}_{\mathrm{nuc}}\) is the product of two rank-\(R\) uniformly bounded matrices and satisfies \(\bar{\Theta}_{\mathrm{nuc}}= \tilde{\Lambda}_{\mathrm{nuc}}\tilde{\Gamma}'_{\mathrm{nuc}}\), each entry in \((\tilde{\Lambda}_{\mathrm{nuc}}, \tilde{\Gamma}_{\mathrm{nuc}})\) must be uniformly bounded wpa1. When \(\rho_{\lambda}, \rho_{\gamma}\) are sufficiently large (independent of \(N, T\)), we have \((\tilde{\Lambda}_{\mathrm{nuc}}, \tilde{\Gamma}_{\mathrm{nuc}})\in \Phi_{NT}\). Therefore, we construct new uniformly bounded nuisance estimators \((\tilde{\Lambda}_{\mathrm{nuc}}, \tilde{\Gamma}_{\mathrm{nuc}})\) and prove that they achieve the same convergence rate, differing only by a constant.
Since constructing a uniformly bounded nuisance estimator is solely for the convenience of theoretical analysis, we do not distinguish between \((\hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\) and \((\tilde{\Lambda}_{\mathrm{nuc}}, \tilde{\Gamma}_{\mathrm{nuc}})\) in the rest of the paper, with a slight abuse of notation.
Proof of Lemma 5. The proof is standard in the literature. Since \((\hat{\beta}_{\mathrm{nuc}}, \hat{\Theta}_{\mathrm{nuc}})\) solves the nuclear norm regularized optimization problem 13 , we have \[\begin{align} \label{eq:lmm95cone951} \mathcal{L}_{NT}( \beta_0 + \hat{\Delta}_{\beta} , \Theta_0 + \hat{\Delta}_{\Theta} ) - \mathcal{L}_{NT}(\beta_0, \Theta_0) \leq \frac{\varphi_{NT}}{\sqrt{NT}} (\|\Theta_0\|_{\mathrm{nuc}} - \|\Theta_0 + \hat{\Delta}_{\Theta}\|_{\mathrm{nuc}} ). \end{align}\tag{23}\] Consider the first-order Taylor expansion of \(\mathcal{L}_{NT}(\hat{\beta}_{\mathrm{nuc}}, \hat{\Theta}_{\mathrm{nuc}})\) around the true parameter. Since \(\mathcal{L}_{NT}(\cdot, \cdot)\) is convex, we have the following inequality \[\begin{align} \label{eq:lmm95cone952} \mathcal{L}_{NT}( \beta_0 + \hat{\Delta}_{\beta} , \Theta_0 + \hat{\Delta}_{\Theta} ) - \mathcal{L}_{NT}(\beta_0, \Theta_0) \geq \nabla_{\beta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)'\hat{\Delta} _{\beta} + \langle\nabla_{\Theta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right), \hat{\Delta}_{\Theta} \rangle. \end{align}\tag{24}\] Combining 23 and 24 yields \[\begin{align} \frac{\varphi_{NT}}{\sqrt{NT}} (\|\Theta_0\|_{\mathrm{nuc}} - \|\Theta_0 + \hat{\Delta}_{\Theta} \|_{\mathrm{nuc}} ) - \nabla_{\beta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)'\hat{\Delta}_{\beta} - \langle\nabla_{\Theta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right), \hat{\Delta}_{\Theta } \rangle\geq 0. \end{align}\] This implies \[\begin{align} \label{eq:lmm95cone953} \frac{\varphi_{NT}}{\sqrt{NT}} (\|\Theta_0\|_{\mathrm{nuc}} - \| \Theta_0 + \hat{\Delta}_{\Theta} \|_{\mathrm{nuc}} ) + |\nabla_{\beta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)'\hat{\Delta} _{\beta}| + | \langle\nabla_{\Theta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right), \hat{\Delta}_{\Theta } \rangle| & \geq 0. \end{align}\tag{25}\] The term \(|\nabla_{\beta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)'\hat{\Delta} _{\beta}|\) is controlled by \[\begin{align} |\nabla_{\beta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)'\hat{\Delta} _{\beta}| \stackrel{\text{(i)}}{\leq} \|\nabla_{\beta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)\| \|\hat{\Delta}_{\beta}\| \stackrel{\text{(ii)}}{\leq} \frac{1}{1 + \alpha}\varphi_{NT} \|\hat{\Delta} _{\beta}\|, \end{align}\] where inequality (i) follows from the Cauchy-Schwarz inequality, and inequality (ii) holds because of the condition \(\varphi_{NT} \geq (1+\alpha) \|\nabla_{\beta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)\|\) as stated in Theorem 5. Similarly, we can control \(| \langle\nabla_{\Theta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right), \hat{\Delta}_{\Theta } \rangle|\) as follows: \[\begin{align} | \langle\nabla_{\Theta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right) , \hat{\Delta}_{\Theta } \rangle| \stackrel{\text{(i)}}{\leq} \|\nabla_{\Theta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)\|_{\mathrm{op}} \|\hat{\Delta}_{\Theta}\|_{\mathrm{nuc}} \stackrel{\text{(ii)}}{\leq} \frac{1}{1+\alpha}\frac{\varphi_{NT}}{\sqrt{NT}}\|\hat{\Delta}_{\Theta}\|_{\mathrm{nuc}}. \end{align}\] where inequality (i) follows from the Hölder’s inequality, as the nuclear norm is the dual norm of the spectral norm, and inequality (ii) holds because of the condition \(\varphi_{NT} \geq (1+\alpha)\sqrt{NT} \|\nabla_{\Theta}\mathcal{L}_{NT}\left(\beta_0 , \Theta_0\right)\|_{\mathrm{op}}\) as stated in Theorem 5. Thus, inequality 25 can be written as \[\begin{align} \label{eq:lmm95cone954} \left(\|\Theta_0\|_{\mathrm{nuc}} - \| \Theta_0 + \hat{\Delta}_{\Theta} \|_{\mathrm{nuc}} \right) + \frac{1}{1+\alpha} \left(\sqrt{NT}\|\hat{\Delta}_{\beta}\| + \|\hat{\Delta}_{\Theta}\|_{\mathrm{nuc}}\right) \geq 0. \end{align}\tag{26}\] In addition, we have the following inequalities for the nuclear norm: \[\begin{align} \|\Theta_0 + \hat{\Delta}_{\Theta}\|_{\mathrm{nuc}} & \stackrel{\text{(i)}}{=} \| M_{\Lambda_0 } \Theta_0 M_{\Gamma_0} + M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0} \| _{\mathrm{nuc}} + \| \Theta_0 - M_{\Lambda_0 } \Theta_0 M_{\Gamma_0} + \hat{\Delta}_{\Theta} - M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0} \| _{\mathrm{nuc}} \\ & \stackrel{\text{(ii)}}{=} \| M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0} \|_{\mathrm{nuc}} + \| \Theta_0 + \hat{\Delta}_{\Theta} - M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0} \|_{\mathrm{nuc}} \\ & \stackrel{\text{(iii)}}{\geq} \| M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0} \|_{\mathrm{nuc}} + \| \Theta_0 \|_{\mathrm{nuc}} - \| \hat{\Delta}_{\Theta} - M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0} \|_{\mathrm{nuc}}. \end{align}\] Here, equality (i) holds by Lemma 4[item:local95projection951], equality (ii) follows from the fact that \(M_{\Lambda_0 } \Theta_0 M_{\Gamma_0} = 0\), and inequality (iii) follows from the triangle inequality.
Finally, we combine the nuclear norm inequality and 26 to obtain \[\begin{align} \| \hat{\Delta}_{\Theta} - M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0} \| _{\mathrm{nuc}} - \| M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0} \| _{\mathrm{nuc}} + \frac{1}{1+\alpha} \left(\sqrt{NT}\|\hat{\Delta} _{\beta}\| + \|\hat{\Delta}_{\Theta}\|_{\mathrm{nuc}}\right) \geq 0 \\ \Rightarrow \frac{2 + \alpha}{1 + \alpha}\| \hat{\Delta}_{\Theta} - M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0} \| _{\mathrm{nuc}} - \frac{\alpha}{1 + \alpha}\| M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0} \| _{\mathrm{nuc}} + \frac{1}{1+\alpha} \sqrt{NT}\|\hat{\Delta} _{\beta}\| \geq 0 \\ \frac{2+\alpha}{\alpha}\left(\| \hat{\Delta}_{\Theta} - M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0} \| _{\mathrm{nuc}} + \sqrt{NT}\|\hat{\Delta} _{\beta}\| \right) \geq \| M_{\Lambda_0 }\hat{\Delta}_{\Theta}M_{\Gamma_0} \| _{\mathrm{nuc}}. \end{align}\] This completes the proof.
Since Corollary 2 extends Corollary 1 to include predetermined covariates, we provide only the proof of Corollary 2, as the proof of Corollary 1 can be viewed as a special case.
As stated in Assumption 4[item:195pre], \(\{(Y_{it}, W_{it})\}_{1\leq t\leq T}\) is \(\phi\)-mixing with a uniformly exponential decay rate across \(i\). However, this assumption is stronger than necessary for the proof of Corollary 2. In fact, it can be relaxed to requiring \(\{(Y_{it}, W_{it})\}_{1\leq t\leq T}\) is \(\alpha\)-mixing with a uniformly sufficiently fast polynomial decay rate across \(i\).
Proof of Corollary 2. It suffices to prove that \[\max\{\|\nabla_{\beta}\mathcal{L}_{NT}(\beta_0, \Theta_0)\|, \sqrt{NT}\|\nabla_{\Theta}\mathcal{L}_{NT}(\beta_0, \Theta_0)\|_{\mathrm{op}} \} = o_p\left(\log(NT)/ \sqrt{\min\{N, T\}}\right).\]
By Assumption 4[item:sampling95pre], \(\nabla_{\beta}\mathcal{L}_{NT}(\beta_0, \Theta_0)\) is the sum of weakly dependent bounded random vectors with zero mean conditional on \((Z, \Lambda_0, \Gamma_0)\). Consequently, we apply [40] to establish that, wpa1, \[\begin{align} \label{eq:corollary95pre951} \|\nabla_{\beta}\mathcal{L}_{NT}(\beta_0, \Theta_0)\| < \log(NT) / \sqrt{NT}. \end{align}\tag{27}\] Note that the \((i, t)\) entry of the matrix \(\nabla_{\Theta}\mathcal{L}_{NT}(\beta_0, \Theta_0)\) is \(\dot{\ell}_{it}(X_{it}'\beta_0 + \theta_0)\), and it is straightforward to verify that (i) \(\{\dot{\ell}_{it}(X_{it}'\beta_0 + \theta_0)\}_{1\leq i\leq N, 1\leq t\leq T}\) is independent across \(i\) and \(\phi\)-mixing with uniformly exponential decay rate, (ii) \(\mathbb{E}_{Z, \Lambda_0, \Gamma_0}(\dot{\ell}_{it}(X_{it}'\beta_0 + \theta_0)) = 0\) by the first-order condition, and (iii) \(\dot{\ell}_{it}(X_{it}'\beta_0 + \theta_0)\) is uniformly bounded across \(i, t, N, T\) by Assumption 4[item:boundedness95pre] and Assumption 4[item:smoothing95pre]. Therefore, employing Lemma 15 gives \[\begin{align} NT \|\nabla_{\Theta}\mathcal{L}_{NT}(\beta_0, \Theta_0)\|_{\mathrm{op}} = O_p\left(\log(N + T)\sqrt{\max\{N , T\}}\right). \end{align}\] Hence, \[\begin{align} \label{eq:corollary95pre952} \sqrt{NT}\|\nabla_{\Theta}\mathcal{L}_{NT}(\beta_0, \Theta_0)\| = o_p\left(\log(NT)/\sqrt{\min\{N, T\}}\right). \end{align}\tag{28}\] We then combine 27 and 28 to complete proof.
Our proof builds on [6] (see Lemma D.3 and Lemma D.4 in their Appendix) and extends their results to accommodate serial correlation. The extension introduces additional technical complexity, particularly in deriving a high-probability upper bound for the empirical process under weak dependence. To address this issue, we apply the concentration inequality of [41] to establish the concentration bound around the expectation of the empirical process. Furthermore, we employ the block method from [42], a new sequence with independent blocks to approximate the original sequence. When the block size is sufficiently large, the dependence between separated blocks becomes negligible, thereby facilitating the theoretical analysis.
It is worth noting that our proof strategy is not only applicable to models with homogeneous slopes, but with minor modifications, can also be extended to accommodate heterogeneous slopes (e.g., [6], [7]). We believe that the flexibility of our strategy enhances the applicability of our approach to a broader class of models.
Proof of Lemma 1. Recall the definition of the constraints space: \[\begin{gather} \mathcal{C}_1 = \left\{(\Delta_{\beta}, \Delta_{\Theta})\in (\mathbb{R}^{d_X}\times \mathbb{R}^{N\times T})\mid \|M_{\Lambda_0}\Delta_{\Theta}M_{\Gamma_0}\|_{\mathrm{nuc}} \leq c_0 \left(\sqrt{NT}\|\Delta_{\beta}\| + \|\Delta_{\Theta} - M_{\Lambda_0}\Delta_{\Theta}M_{\Gamma_0}\|_{\mathrm{nuc}}\right)\right\}, \\ \mathcal{C}_2 = \left\{ (\Delta_{\beta}, \Delta_{\Theta})\in (\mathbb{R}^{d_X}\times \mathbb{R}^{N\times T})\mid \|\Delta_{\beta}\|^2 + \frac{1}{NT} \|\Delta_{\Theta}\|_{\mathrm{F}}^2 \geq \sqrt{\frac{\log (NT)}{NT}}\right\}. \end{gather}\] For notational simplicity, let \(\mathcal{C} = \mathcal{C}_1\cap \mathcal{C}_2\). Also, use \(\mathbb{E}_{\mathcal{V}}(\cdot) = \mathbb{E}(\cdot\mid \mathcal{V})\) to denote the conditional expectation, and \(\mathbb{P}_{\mathcal{V}}(\cdot) = \mathbb{P}(\cdot\mid \mathcal{V})\) to denote the conditional probability.
In this step, we aim to establish a lower bound for \(\sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}_{\mathcal{V}}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{ it}} )^2\). By Assumption 6[item:conditional95variability95RSC], there exists a constant \(\kappa_0>0\) such that \[\begin{align} \inf_{1\leq i\leq N, 1\leq t\leq T} \sigma_{\min } \left( \begin{pmatrix} \mathbb{E}_{\mathcal{V}}(X_{it}X_{it}') & \mathbb{E}_{\mathcal{V}}(X_{it}) \\ \mathbb{E}_{\mathcal{V}}(X_{it}') & 1 \end{pmatrix}\right) \geq \kappa_0. \end{align}\] It follows that \[\label{eq:lower95bound95expectation} \begin{align} \sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}_{\mathcal{V}}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 &\geq \sum_{i=1}^{N}\sum_{t=1}^{T} \begin{pmatrix} \Delta'_{\beta} & \Delta_{\theta_{it}} \end{pmatrix} \begin{pmatrix} \mathbb{E}_{\mathcal{V}}(X_{it}X_{it}') & \mathbb{E}_{\mathcal{V}}(X_{it}) \\ \mathbb{E}_{\mathcal{V}}(X_{it}') & 1 \end{pmatrix} \begin{pmatrix} \Delta_{\beta} \\ \Delta_{\theta_{it}} \end{pmatrix}\\ & \geq \kappa_0\left(NT \|\Delta_{\beta}\|^2 + \|\Delta_{\Theta}\|_{\mathrm{F}}^2\right). \end{align}\tag{29}\]
For any \(\omega>0\), define the constraint set \(\mathcal{N}(\omega)\) as \[\begin{align} \mathcal{N}(\omega) := \left\{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{C} \mid \|\Delta_{\beta}\|^2 + \frac{1}{NT}\|\Delta_{\Theta}\|_{\mathrm{F}}^2 \leq \omega, \text{ }\|\Delta_{\beta}\|_{\max}\leq 2\rho_{\beta}, \text{ }\|\Delta_{\Theta}\|_{\max} \leq 2 \rho_\theta \right\}. \end{align}\] and define the empirical process as \[\begin{align} Z(\omega) = \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left| \sum_{i=1}^{N}\sum_{t=1}^{T}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 - \mathbb{E}_{\mathcal{V}}\sum_{i=1}^{N}\sum_{t=1}^{T}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 \right|. \end{align}\] It is straightforward to verify that \[\begin{align} \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} |(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2- \mathbb{E}_{\mathcal{V}}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2| & \leq 2\sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} (X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 \\ & \leq 4\sup_{\|\Delta_{\beta}\|_{\max}\leq 2\rho_{\beta}, \|\Delta_{\theta}\|_{\max}\leq 2\rho_{\theta} } \left\{(X_{it}' \Delta_{\beta} )^2 + \Delta_{\theta_{it}}^2\right\} \\ & \leq \underbrace{16(d_X^2 \rho_{X}^2\rho_{\beta}^2 + \rho_{\theta}^2) }_{\sigma} \end{align}\] almost surely. By Assumption 6[item:conditional95weak95dependence95RSC], conditional on \(\mathcal{V}\), the sequence \(\{X_{it}\}_{1\leq t\leq T}\) is \(\phi\)-mixing with a uniform exponential decay rate. This allows us to apply [41] to obtain the following concentration inequality for \(Z(\omega)\) around its expectation: \[\begin{align} \mathbb{P}_{\mathcal{V}}\left(Z(\omega) \geq \mathbb{E}_{\mathcal{V}} Z(\omega) + \delta\right) \leq \exp\left( - L^{-1} \min\left\{\frac{\delta}{\sigma}, \frac{\delta^2}{NT\sigma^2}\right\}\right), \quad \forall \delta >0. \end{align}\] Here, \(L \in [1, \infty)\) does not depend on \(N, T\), and is only determined by the mixing properties of \(X_{it}\) conditional on \(\mathcal{V}\).
We follow the block method of [42] to derive an upper bound for \(\mathbb{E}_{\mathcal{V}}Z(\omega)\). Define \(X_i = \left(X_{i1}, X_{i2}, \ldots, X_{iT}\right)\), where \(\{X_i\}\) is independent across \(i\), and for each \(i\), \(X_i\) is \(\phi\)-mixing with mixing coefficients \(\phi(\tau)\). Let \(\tau_{NT} \leq T\) be a positive integer, and \(\mu_{NT} = \lfloor \frac{T}{2\tau_{NT}} \rfloor\) denote the largest integer less than or equal to \(\frac{T}{2\tau_{NT}}\). For each \(X_i\), we divide the sequence into \(2\mu_{NT}\) blocks, each of length \(\tau_{NT}\), with the remaining part having length at most \(2\tau_{NT}\). The subscript notation indicates that the values of \(\tau_{NT}\) and \(\mu_{NT}\) may depend on \(N\) and \(T\).
We further partition the blocks into two groups: odd-numbered blocks and even-numbered blocks. For notational simplicity, let \(\mathcal{T}^{(0)}_k\) denote the indices of elements in the \(k\)-th odd block, and \(\mathcal{T}^{(0)}\) denote the set of indices corresponding to the elements contained in odd-numbered blocks. Similarly, let \(\mathcal{T}^{(1)}_k\) denote the set of indices of elements in the \(k\)-th even block, and \(\mathcal{T}^{(1)}\) denote the set of the indices of all elements in even-numbered blocks. Specifically, \[\begin{align} \mathcal{T}^{(0)} := \bigcup_{k=1}^{\mu_{NT}} \mathcal{T}^{(0)}_k, \quad \mathcal{T}^{(0)}_k = \left\{t \mid 2(k-1)\tau_{NT} + 1\leq t\leq 2(k-1)\tau_{NT} + \tau_{NT}\right\}, \\ \mathcal{T}^{(1)} L= \bigcup_{k=1}^{\mu_{NT}} \mathcal{T}^{(1)}_k, \quad \mathcal{T}^{(1)}_k = \left\{t \mid (2k-1)\tau_{NT} + 1\leq t\leq (2k-1)\tau_{NT} + \tau_{NT}\right\}. \end{align}\] The corresponding partition of \(X_i\) can be written as \[\begin{gather} X_i^{(0)} := \left(X_i^{(0, 1)}, X_i^{(0, 2)}, \ldots, X_i^{(0, \mu_{NT})}\right), \\ X_i^{(1)} := \left( X_i^{(1, 1)} , X_i^{(1, 2)}, \ldots, X_i^{(1, \mu_{NT})}\right), \end{gather}\] where for each \(1\leq k \leq \mu_{NT}\), \[\begin{align} X_i^{(0, k)} = \left(X_{it}\mid t\in \mathcal{T}^{(0)}_k \right), \quad X_i^{(1, k)} = \left(X_{it}\mid t\in \mathcal{T}^{(1)}_k \right). \end{align}\] The remaining terms are collected into \(R_i\), where \[\begin{align} R_i = \left(X_{it}\mid t \in \mathcal{R} \right), \quad \mathcal{R} = \left\{t \mid 2\tau_{NT}\mu_{NT} + 1 \leq t \leq T \right\} \end{align}\]
In the next step, for each \(i\), we construct a new sequence with independent block structure conditional on \(\mathcal{V}\): \[\begin{align} \widetilde{X}^{(0)}_i = \left(\widetilde{X}_i^{(0, 1)}, \widetilde{X}_i^{(0, 2)}, \ldots, \widetilde{X}_i^{(0, \mu_{NT})} \right), \end{align}\] such that each block \(\widetilde{X}^{(0, k)}_i\) (of size \(\tau_{NT}\)) is independent of the others conditional on \(\mathcal{V}\). Within each block, \(\widetilde{X}^{(0, k)}_i\) follows the same conditional distribution as \(X^{(0, k)}_i\). We then construct \(\{\widetilde{X}^{(1)}_i\}\) in a similar way, i.e., \[\begin{align} \widetilde{X}^{(1)}_i := \left(\widetilde{X}_i^{(1, 1)}, \widetilde{X}_i^{(1, 2)}, \ldots, \widetilde{X}_i^{(1, \mu_{NT})}\right). \end{align}\] Here, each block is independent with the others conditional on \(\mathcal{V}\). Within each block, \(\widetilde{X}^{(1, k)}_i\) follows the same conditional distribution as \(X^{(1, k)}_i\).
Denote \(\widetilde{X} = (\widetilde{X}_i^{(0, 1)}, \widetilde{X}_i^{(1, 1)}, \ldots, \widetilde{X}_i^{(0, \mu_{NT})}, \widetilde{X}_i^{(1, \mu_{NT})})\), and for \(s \in \{0, 1\}\), define the empirical process \[\begin{align} Z^{(s)}(\omega) & := \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left| \sum_{i=1}^{N}\sum_{ t\in \mathcal{T}^{(s)}} (X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 - \mathbb{E}_{\mathcal{V}}\sum_{i=1}^{N}\sum_{ t\in \mathcal{T}^{(s)}}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 \right|, \\ \widetilde{Z}^{(s)}(\omega) & := \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left| \sum_{i=1}^{N}\sum_{ t\in \mathcal{T}^{(s)}} (\widetilde{X}_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 - \mathbb{E}_{\mathcal{V}}\sum_{i=1}^{N}\sum_{ t\in \mathcal{T}^{(s)}}(\widetilde{X}_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 \right|. \end{align}\] It is straightforward to see that \(\mathbb{E}_{\mathcal{V}}Z(\omega)\) can be bounded by \[\label{eq:lemma95RSC951} \begin{align} \mathbb{E}_{\mathcal{V}} Z(\omega) \leq & \mathbb{E}_{\mathcal{V}} Z^{(0)}(\omega) + \mathbb{E}_{\mathcal{V}} Z^{(1)}(\omega) \\ \leq & \mathbb{E}_{\mathcal{V}} \widetilde{Z}^{(0)}(\omega) + \mathbb{E}_{\mathcal{V}} \widetilde{Z}^{(1)}(\omega) + | \mathbb{E}_{\mathcal{V}} \widetilde{Z}^{(0)}(\omega) - \mathbb{E}_{\mathcal{V}}Z^{(0)}(\omega)| + | \mathbb{E}_{\mathcal{V}} \widetilde{Z}^{(1)}(\omega) - \mathbb{E}_{\mathcal{V}}Z^{(1)}(\omega)|\\ & + 2 \sigma N\tau_{NT}. \end{align}\tag{30}\] The last term on the right-hand side of the inequality, \(2\sigma N\tau_{NT}\), arises from the remaining terms when \(T\) cannot be exactly divided by \(\tau_{NT}\). The following lemma is crucial in bounding \(| \mathbb{E}_{\mathcal{V}} \widetilde{Z}^{(0)}(\omega) - \mathbb{E}_{\mathcal{V}}Z^{(0)}(\omega)|\) and \(| \mathbb{E}_{\mathcal{V}} \widetilde{Z}^{(1)}(\omega) - \mathbb{E}_{\mathcal{V}}Z^{(1)}(\omega)|\):
Lemma 6 (Based on Lemma 4.1 of [42]). For any measurable function \(h: \mathbb{R}^{N \times \tau_{NT}\mu_{NT}} \mapsto \mathbb{R}\) with bound \(M>0\), we have \[\begin{align} | \mathbb{E}_{\mathcal{V}} h(X^{(s)}_{1}, X^{(s)}_{2}, \ldots, X^{(s)}_{N}) - \mathbb{E}_{\mathcal{V}} h(\widetilde{X}^{(s)}_{1}, \widetilde{X}^{(s)}_{2}, \ldots, \widetilde{X}^{(s)}_{N}) | \leq M \left(N\mu_{NT} -1 \right) \phi(\tau_{NT}), \quad s = 0,1. \end{align}\]
To apply Lemma 6, let \[\begin{align} h^{(s)}(X^{(s)}_{1}, X^{(s)}_{2}, \ldots, X^{(s)}_{N}) : = \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left| \sum_{i=1}^{N}\sum_{ t\in \mathcal{T}^{(s)}} (X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 - \mathbb{E}_{\mathcal{V}}\sum_{i=1}^{N}\sum_{ t\in \mathcal{T}^{(s)}}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 \right| \end{align}\] It is straightforward to verify that the following inequality holds almost surely for \(s = 0,1\): \[\begin{align} |h^{(s)}(X^{(s)}_{1}, X^{(s)}_{2}, \ldots, X^{(s)}_{N})| \leq & 2 \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left| \sum_{i=1}^{N}\sum_{ t\in \mathcal{T}^{(s)}} (X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 \right| \\ \leq & 4 \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left| \sum_{i=1}^{N}\sum_{ t\in \mathcal{T}^{(s)}} (X_{it}' \Delta_{\beta})^2 + \sum_{i=1}^{N}\sum_{ t\in \mathcal{T}^{(s)}} \Delta_{\theta_{it}}^2 \right| \\ \leq & 4 \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \sum_{i=1}^{N}\sum_{ t\in \mathcal{T}^{(s)}} (X_{it}' \Delta_{\beta})^2 + 4 \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \sum_{i=1}^{N}\sum_{ t\in \mathcal{T}^{(s)}} \Delta_{\theta_{it}}^2 \\ \leq & 4 N \mu_{NT} \tau_{NT}d_X \rho_X^2 \omega + 4 NT \omega \\ \leq & 2 NT (d_X \rho_X^2 + 2)\omega. \end{align}\] Thus, applying Lemma 6 to \(| \mathbb{E}_{\mathcal{V}} \widetilde{Z}^{(s)}(\omega) - \mathbb{E}_{\mathcal{V}}Z^{(s)}(\omega)|\) (with \(M \leq 2 NT(d_X \rho_X^2 + 2)\omega\)) yields \[\label{eq:lemma95RSC952} \begin{align} | \mathbb{E}_{\mathcal{V}} \widetilde{Z}^{(s)}(\omega) - \mathbb{E}_{\mathcal{V}}Z^{(s)}(\omega)| \leq & 2 NT\left(d_X \rho_X^2 + 2\right)\omega \left(N\mu_{NT} -1 \right) \phi(\tau_{NT}) \\ \leq & \underbrace{\left(d_X \rho_X^2 + 2\right)}_{\frac{C_0}{2}} (NT)^2 \frac{\phi(\tau_{NT})}{\tau_{NT}}\omega. \end{align}\tag{31}\] Intuitively, when \(\phi(\tau_{NT})\) decays sufficiently fast as \(\tau_{NT}\rightarrow \infty\), we expect that \(\phi(\tau_{NT})/ \tau_{NT}\rightarrow 0\) sufficiently fast, so that \(| \mathbb{E}_{\mathcal{V}} \widetilde{Z}^{(s)}(\omega) - \mathbb{E}_{\mathcal{V}}Z^{(s)}(\omega)|\) can be well controlled.
We now turn to establishing a bound for \(\mathbb{E}_{\mathcal{V}} \widetilde{Z}^{(s)}(\omega)\). Since each block in \(\widetilde{X}_i\) is independent with the others conditional on \(\mathcal{V}\), it suffices to study the Rademacher process for each block. More specifically, for each \(s\), we can construct a collecion of i.i.d. Rademacher random variables \(\{\epsilon_{ik}^{(s)}\mid i = 1,2,\ldots, N, k = 1,2,\ldots, \mu_{NT}\}\), which are independent of \((\widetilde{X}^{(s)}_{1}, \widetilde{X}^{(s)}_{2}, \ldots, \widetilde{X}^{(s)}_{N})\) conditional on \(\mathcal{V}\). Using symmetrization method, we obtain \[\label{eq:lemma95RSC953} \begin{align} \mathbb{E}_{\mathcal{V}} \widetilde{Z}^{(s)}(\omega) & = \mathbb{E}_{\mathcal{V}} \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left| \sum_{i=1}^{N}\sum_{ t\in \mathcal{T}^{(s)}} (\widetilde{X}_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 - \mathbb{E}_{\mathcal{V}}\sum_{i=1}^{N}\sum_{ t\in \mathcal{T}^{(s)}}(\widetilde{X}_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 \right| \\ & \leq 2 \mathbb{E}_{\mathcal{V}, \epsilon} \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left| \sum_{i=1}^{N}\sum_{ k = 1}^{\mu_{NT}} \left(\sum_{t \in \mathcal{T}^{(s)}_k}(\widetilde{X}_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2\right) \epsilon_{ik}^{(s)}\right|. \end{align}\tag{32}\] For \(s=0\), we have \[\label{eq:lemma95RSC954} \begin{align} & 2 \mathbb{E}_{\mathcal{V}, \epsilon} \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left| \sum_{i=1}^{N}\sum_{ k = 1}^{\mu_{NT}} \left(\sum_{t \in \mathcal{T}^{(0)}_k}(\widetilde{X}_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2\right) \epsilon_{ik}^{(0)}\right| \\ \stackrel{\text{(i)}}{\leq} & 2 \mathbb{E}_{\mathcal{V}, \epsilon} \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left| \sum_{ \tau = 1}^{\tau_{NT}} \sum_{i=1}^{N} \sum_{k = 1 }^{\mu_{NT}}(\widetilde{X}_{i, 2(k-1)\tau_{NT} + \tau}' \Delta_{\beta} + \Delta_{\theta_{i, 2(k-1)\tau_{NT} + \tau}} )^2 \epsilon_{ik}^{(0)} \right| \\ \leq & 2 \sum_{ \tau = 1}^{\tau_{NT}} \mathbb{E}_{\mathcal{V}, \epsilon} \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left|\sum_{i=1}^{N} \sum_{k = 1 }^{\mu_{NT}}(\widetilde{X}_{i, 2(k-1)\tau_{NT} + \tau}' \Delta_{\beta} + \Delta_{\theta_{i, 2(k-1)\tau_{NT} + \tau}} )^2 \epsilon_{ik}^{(0)} \right| \\ \stackrel{\text{(ii)}}{\leq} & \underbrace{16 (d_X\rho_X\rho_{\beta} + \rho_{\theta})}_{\frac{C_1}{2}}\sum_{ \tau = 1}^{\tau_{NT}} \mathbb{E}_{\mathcal{V}, \epsilon} \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left|\sum_{i=1}^{N} \sum_{k = 1 }^{\mu_{NT}}(\widetilde{X}_{i, 2(k-1)\tau_{NT} + \tau}' \Delta_{\beta} + \Delta_{\theta_{i, 2(k-1)\tau_{NT} + \tau}} ) \epsilon_{ik}^{(0)} \right| \\ \leq & \frac{C_1}{2}\sum_{ \tau = 1}^{\tau_{NT}} \underbrace{\mathbb{E}_{\mathcal{V}, \epsilon} \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left|\sum_{i=1}^{N} \sum_{k = 1 }^{\mu_{NT}}\widetilde{X}_{i, 2(k-1)\tau_{NT} + \tau}' \Delta_{\beta} \epsilon_{ik}^{(0)} \right|}_{S_{1, \tau}} \\ & +\frac{C_1}{2}\sum_{ \tau = 1}^{\tau_{NT}} \underbrace{\mathbb{E}_{ \epsilon} \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left|\sum_{i=1}^{N} \sum_{k = 1 }^{\mu_{NT}} \Delta_{\theta_{i, 2(k-1)\tau_{NT} + \tau}}\epsilon_{ik}^{(0)} \right|}_{S_{2, \tau}} . \end{align}\tag{33}\] Here, we change the order of summation to obtain inequality (i). Inequality (ii) follows from the contraction property of the Rademacher process (see [43]). To derive an upper bound of \(S_{1, \tau}\), for each \(\tau\), \[\label{eq:lemma95RSC955} \begin{align} S_{1, \tau} & = \mathbb{E}_{\mathcal{V}, \epsilon} \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left|\sum_{i=1}^{N} \sum_{k = 1 }^{\mu_{NT}}\widetilde{X}_{i, 2(k-1)\tau_{NT} + \tau}' \Delta_{\beta} \epsilon_{ik}^{(0)} \right| \\ & \stackrel{\text{(i)}}{\leq} \mathbb{E}_{\mathcal{V}, \epsilon} \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left(\left\| \sum_{i=1}^{N} \sum_{k = 1 }^{\mu_{NT}}\widetilde{X}_{i, 2(k-1)\tau_{NT} + \tau}' \epsilon_{ik}^{(0)}\right\| \|\Delta_{\beta}\|\right) \\ & = \mathbb{E}_{\mathcal{V}, \epsilon} \left\| \sum_{i=1}^{N} \sum_{k = 1 }^{\mu_{NT}}\widetilde{X}_{i, 2(k-1)\tau_{NT} + \tau}' \epsilon_{ik}^{(0)}\right\| \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \|\Delta_{\beta}\| \\ & \stackrel{\text{(ii)}}{\leq} \underbrace{\sqrt{2 \pi d_X^3 \rho_X^2 }}_{\sqrt{2}C_2} \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \|\Delta_{\beta}\|\\ & \leq \sqrt{2}C_2 \sqrt{N\mu_{NT}}\sqrt{\omega} \\ & \leq C_2 \sqrt{\frac{NT\omega}{\tau_{NT}}} , \end{align}\tag{34}\] where inequality (i) follows from the Cauchy-Schwarz inequality, and inequality (ii) follows from Lemma 14 using the fact that each element in \(\widetilde{X}_{i, 2(k-1)\tau_{NT} + \tau}' \epsilon_{ik}^{(0)}\) is bounded by \(\rho_X\).
Let \(\Delta^{(0, \tau)}_{\Theta}\) be an \(N\times \mu_{NT}\) matrix such that \([\Delta^{(0, \tau)}_{\Theta}]_{ik} = \theta_{i, 2(k-1)\tau_{NT} + \tau}\) for each \(\tau =1,2,\ldots, \tau_{NT}\). Let \(E^{(0)}\) denote the \(N\times \mu_{NT}\) matrix collecting \(\epsilon^{(0)}_{ik}\). Then, \[\label{eq:lemma95RSC956} \begin{align} S_{2, \tau} = \mathbb{E}_{ \epsilon} \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left|\langle\Delta_{\Theta}^{(0, \tau)}, E^{(0)} \rangle\right| \stackrel{\text{(i)}}{\leq} \mathbb{E}_{\epsilon} \| E^{(0)} \|_{\mathrm{op}} \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)}\|\Delta_{\Theta}^{(0, \tau)} \|_{\mathrm{nuc}}, \end{align}\tag{35}\] where inequality (i) comes from the fact that nuclear norm is the duel norm of the operator norm. In addition, by [44], there exists a constant \(C_3\) that does not depend on \(N\), \(T\), or \(\tau_{NT}\) such that \(\mathbb{E}_{\epsilon}\| E^{(0)} \|_{\mathrm{op}} \leq C_3 \sqrt{N + \mu_{NT}}\), based on [44]. Furthermore, we have \[\label{eq:lemma95RSC957} \begin{align} \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)}\|\Delta_{\Theta}^{(0, \tau)} \|_{\mathrm{nuc}} \stackrel{\text{(i)}}{\leq} & \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)}\|\Delta_{\Theta}\|_{\mathrm{nuc}}\\ \stackrel{\text{(ii)}}{\leq} & \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)} \left(\|\Delta_{\Theta} - M_{\Lambda_0}\Delta_{\Theta}M_{\Gamma_0}\|_{\mathrm{nuc}} + \|M_{\Lambda_0}\Delta_{\Theta}M_{\Gamma_0}\|_{\mathrm{nuc}}\right) \\ \stackrel{\text{(iii)}}{\leq} & \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)}\left( (1 + c_0)\|\Delta_{\Theta} - M_{\Lambda_0}\Delta_{\Theta}M_{\Gamma_0}\|_{\mathrm{nuc}} + c_0 \sqrt{NT}\|\Delta_{\beta}\|\right) \\ \stackrel{\text{(iv)}}{\leq} & \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)}\left((1+c_0)\sqrt{2R}\|\Delta_{\Theta} - M_{\Lambda_0}\Delta_{\Theta}M_{\Gamma_0}\|_{\mathrm{F}} + c_0\sqrt{NT}\|\Delta_{\beta}\|\right) \\ \stackrel{\text{(v)}}{\leq} & \sup_{(\Delta_{\beta},\Delta_{\Theta})\in \mathcal{N}(\omega)}\left((1+c_0)\sqrt{2R}\|\Delta_{\Theta}\|_{\mathrm{F}} + c_0 \sqrt{NT}\|\Delta_{\beta}\|\right) \\ \stackrel{\text{(vi)}}{\leq} & \underbrace{\left((1+c_0)\sqrt{2R} + c_0 \right) }_{C_4}\sqrt{NT\omega} \\ \leq & C_4 \sqrt{NT\omega }. \end{align}\tag{36}\] Inequality (i) follows from Lemma 17, since \(\Delta_{\Theta}^{(0, \tau)}\) can be regarded as a submatrix of \(\Delta_{\Theta}\). Inequality (ii) is merely an application of triangle inequality, and inequality (iii) comes from \((\Delta_{\beta},\Delta_{\Theta}) \in \mathcal{C}\). Inequality (iv) holds because \(\Delta_{\Theta} - M_{\Lambda_0}\Delta_{\Theta}M_{\Gamma_0}\) is a matrix of rank at most \(2R\). Inequality (v) follows from the fact \(\|\Delta_{\Theta}\|_{\mathrm{F}}^2 = \|\Delta_{\Theta} - M_{\Lambda_0}\Delta_{\Theta}M_{\Gamma_0}\|_{\mathrm{F}}^2 + \|M_{\Lambda_0}\Delta_{\Theta}M_{\Gamma_0}\|_{\mathrm{F}}^2\). Finally, inequality (vi) follows directly from \((\Delta_{\beta},\Delta_{\Theta}) \in \mathcal{N}(\omega)\).
The same argument can be applied to the case \(s = 1\). Combining equations 30 —36 , we derive the following bound: \[\begin{align} \mathbb{E}_{\mathcal{V}} \widetilde{Z}^{(s)}(\omega) \leq &\frac{C_1\tau_{NT}}{2}\left(C_2 \sqrt{\frac{NT\omega}{\tau_{NT}}} + C_3 C_4 \sqrt{NT(N + \mu_{NT})\omega} \right) \\ \leq & \frac{C_1C_2}{2}\sqrt{NT\tau_{NT}\omega} + \frac{C_1C_3C_4}{2} \sqrt{NT(N+\mu_{NT})\tau_{NT}^2 \omega}, \quad s\in \{0, 1\}. \end{align}\] Therefore, \[\begin{align} \mathbb{E}_{\mathcal{V}}Z(\omega) \leq & C_1C_2 \sqrt{NT\tau_{NT}\omega} + C_1C_3C_4 \sqrt{NT(N+\mu_{NT})\tau_{NT}^2 \omega} + C_0 (NT)^2 \frac{\phi(\tau_{NT})}{\tau_{NT}}\omega + 2 \sigma N\tau_{NT} \\ \leq & (C_1C_2 + C_1C_3C_4)\sqrt{NT(N+\mu_{NT})\tau_{NT}^2 \omega} + C_0 (NT)^2 \frac{\phi(\tau_{NT})}{\tau_{NT}} \omega + 2 \sigma N\tau_{NT} \\ \leq & \frac{\kappa_0}{8} NT\omega + \left(\frac{8}{\kappa_0} (C_1C_2 + C_1C_3C_4 )^2 + 2\sigma\right) (N+\mu_{NT})\tau_{NT}^2 + C_0 (NT)^2 \frac{\phi(\tau_{NT})}{\tau_{NT}} \omega. \end{align}\] When \(\phi(\tau_{NT}) = e^{-\zeta_0 \tau_{NT}}\), let \(\tau_{NT} = \frac{2}{\zeta_0}\log (NT)\). Then we have \[\begin{align} \mathbb{E}_{\mathcal{V}}Z(\omega) \leq & \frac{\kappa_0}{8} NT\omega + \underbrace{\frac{4}{\zeta_0^2}\left( \frac{8}{\kappa_0} (C_1C_2 + C_1C_3C_4)^2 + 2\sigma + C_0\right)}_{ \eta} (N+T)(\log(NT))^2 \\ \leq & \frac{\kappa_0}{8} NT\omega + \eta (N+T)(\log(NT))^2. \end{align}\]
Substituting the upper bound for \(\mathbb{E}_{\mathcal{V}} Z(\omega)\) derived above into the concentration inequality, we obtain \[\begin{align} \mathbb{P}_{\mathcal{V}}\left(Z(\omega) \geq \frac{\kappa_0}{8} NT\omega + \eta (N+T)(\log(NT))^2 + \delta\right) \leq \exp\left( -L^{-1} \min\left\{\frac{\delta}{\sigma}, \frac{\delta^2}{NT\sigma^2}\right\}\right), \quad \forall \delta >0. \end{align}\] Let \(\delta = \frac{\kappa_0 }{8} NT\omega\). It follows that \[\label{eq:lemma95RSC958} \begin{align} \mathbb{P}_{\mathcal{V}}\left(Z(\omega) \geq \frac{\kappa_0}{4} NT\omega + \eta (N+T)(\log(NT))^2 \right) \leq \exp\left( - \min\left\{\frac{ \kappa_0 NT\omega}{8 L \sigma}, \frac{\kappa_0^2 NT \omega^2}{64 L \sigma^2}\right\}\right). \end{align}\tag{37}\]
For \(\ell = 1,2,\ldots\), define \[\begin{align} \mathcal{D}_{\ell} := \left\{(\Delta_{\beta},\Delta_{\Theta}) \in \mathcal{C} \mid 2^{\ell-1} \sqrt{\frac{\log (NT)}{NT}} \leq \|\Delta_{\beta}\|^2 + \frac{1}{NT}\|\Delta_{\Theta}\|_{\mathrm{F}}^2 \leq 2^{\ell} \sqrt{\frac{\log (NT)}{NT}} \right\} \end{align}\] It is readily verified that \(\mathcal{C} \subset \bigcup_{\ell =1}^{\infty}\mathcal{D}_{\ell}\). In addition, let \(\omega_{\ell} := 2^{\ell}\sqrt{\frac{\log (NT)}{NT}}\). Define \[\begin{align} \mathcal{E}_{\ell} := \big\{ Z(\omega_{\ell})\geq \frac{\kappa_0}{4}NT\omega_{\ell} + \eta (N+T)(\log(NT))^2 \big\}, \end{align}\] and \[\begin{align} \tilde{\mathcal{E}}_{\ell} := \bigg\{ & \left|\sum_{i=1}^{N}\sum_{t=1}^{T}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 - \sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2\right | \\ & \geq \frac{1}{2}\sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}_{\mathcal{V}}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 + \eta (N+T)(\log(NT))^2, \\ & \exists (\Delta_{\beta},\Delta_{\Theta}) \in \mathcal{D}_{\ell} \bigg\}. \end{align}\] One can verify that when \(\tilde{\mathcal{E}}_{\ell}\) happens, since \[\begin{align} \sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}_{\mathcal{V}}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 \geq NT \kappa_0\left( \|\Delta_{\beta}\|^2 + \frac{1}{NT}\|\Delta_{\Theta}\|_{\mathrm{F}}^2\right) \geq 2^{\ell-1} \kappa_0 NT \sqrt{\frac{\log (NT)}{NT}}, \end{align}\] we obtain \[\begin{align} |\sum_{i=1}^{N}\sum_{t=1}^{T}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 - \sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}_{\mathcal{V}}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2| & \geq \frac{1}{2}\sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}_{\mathcal{V}}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 + \eta (N+T)(\log(NT))^2 \\ \Rightarrow |\sum_{i=1}^{N}\sum_{t=1}^{T}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 - \sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}_{\mathcal{V}}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2| & \geq 2^{\ell-2} \kappa_0 NT \sqrt{\frac{\log NT}{NT}} + \eta (N+T)(\log(NT))^2 \\ \Rightarrow |\sum_{i=1}^{N}\sum_{t=1}^{T}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 - \sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}_{\mathcal{V}}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2| & \geq \frac{\kappa_0}{4}NT\omega_{\ell} + \eta (N+T)(\log(NT))^2 \\ \Rightarrow Z(\omega_{\ell}) & \geq \frac{\kappa_0}{4}NT\omega_{\ell} + \eta (N+T)(\log(NT))^2. \end{align}\] Therefore, \(\tilde{\mathcal{E}}_{\ell} \subset \mathcal{E}_{\ell}\). In addition, \[\label{eq:around12} \begin{align} \mathbb{P}_{\mathcal{V}}\bigg(& |\sum_{i=1}^{N}\sum_{t=1}^{T}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 - \sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}_{\mathcal{V}}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2| \\ & \geq \frac{1}{2}\sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 + \eta (N+T)(\log(NT))^2, \exists (\Delta_{\beta},\Delta_{\Theta}) \in \mathcal{C} \bigg) \\ \leq & \mathbb{P}_{\mathcal{V}}\left(\bigcup_{\ell=1}^{\infty}\tilde{\mathcal{E}}_{\ell}\right) \leq \sum_{\ell=1}^{\infty}\mathbb{P}_{\mathcal{V}}\left(\tilde{\mathcal{E}}_{\ell}\right) \leq \sum_{\ell=1}^{\infty}\mathbb{P}_{\mathcal{V}}\left(\mathcal{E}_{\ell}\right) \\ \stackrel{\text{(i)}}{\leq} & \sum_{\ell}^{\infty} \exp\left( - \frac{ \kappa_0 NT\omega_\ell}{8\Phi \sigma}\right) + \sum_{\ell}^{\infty} \exp\left( - \frac{\kappa_0^2 NT \omega_\ell^2}{64\Phi\sigma^2} \right) \\ \leq & \sum_{\ell}^{\infty} \exp\left( - \frac{ \kappa_0 2^{\ell} \sqrt{NT \log (NT)}}{8\Phi \sigma}\right) + \sum_{\ell}^{\infty} \exp\left( - \frac{\kappa_0^2 4^{\ell}\log (NT)}{64\Phi\sigma^2} \right) \\ \rightarrow & 0, \end{align}\tag{38}\] where inequality (i) follows from the probability bound 37 .
Combining the lower bound 29 and 38 , the following inequality holds for all \((\Delta_{\beta}, \Delta_{\Theta})\in \mathcal{C}\) with probability approaching \(1\): \[\begin{align} |\sum_{i=1}^{N}\sum_{t=1}^{T}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 - \sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2| \leq \frac{1}{2}\sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 + \eta (N+T)(\log(NT))^2. \end{align}\] This implies that, wpa1, \[\begin{align} \Rightarrow \sum_{i=1}^{N}\sum_{t=1}^{T}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 \geq \frac{1}{2}\sum_{i=1}^{N}\sum_{t=1}^{T}\mathbb{E}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 - \eta (N+T)(\log(NT))^2 . \end{align}\] Consequently, \[\begin{align} \Rightarrow \sum_{i=1}^{N}\sum_{t=1}^{T}(X_{it}' \Delta_{\beta} + \Delta_{\theta_{it}} )^2 \geq \frac{1}{2}\kappa_0\left(NT \|\Delta_{\beta}\|^2 + {\vert\kern-0.25ex\vert \Delta_{\Theta} \vert\kern-0.25ex\vert}_{\mathrm{F}}^2\right) - \eta (N+T)(\log(NT))^2. \end{align}\] Since the inequality above holds for any \(\mathcal{V}\), we omit the subscript \(\mathcal{V}\) and obtain the same bound under \(\mathbb{P}\). Let \(\kappa = \frac{1}{2}\kappa_0\). This completes the proof.
We begin by clarifying some related issues before proceeding the proofs.
First, we introduce the following abbreviations to simplify notations. For any \((\beta, \lambda_i, \gamma_t)\), define \(\dot{\ell}_{it} = \dot{\ell}(X_{it}'\beta + \lambda_i' \gamma_t)\) and \(\ddot{\ell}_{it} = \ddot{\ell}(X_{it}'\beta + \lambda_i' \gamma_t)\). When the log-likelihood is evaluated at the (normalized) true parameters, define \(\dot{\ell}_{it}^0 = \dot{\ell}(X_{it}'\beta_0 + \lambda_{0, i}' \gamma_{0, t})\) and \(\ddot{\ell}_{it}^0 = \ddot{\ell}(X_{it}'\beta_0 + \lambda_{0, i}' \gamma_{0, t})\). We further define the following quantities: \(\Delta_{\beta} = \beta - \beta_0\), \(\Delta_{\gamma_t} = \gamma_t - \gamma^G_{0,t}\), and \(\Delta_{\lambda_i} = \lambda_i - \lambda^G_{0,i}\), to denote the deviations of the parameters from their true values. When considering the difference between \(\ddot{\ell}_{it}\) and \(\ddot{\ell}^0_{it}\), we write \[\begin{align} \tilde{\Delta}_{Y^*_{it}} = \ddot{\ell}_{it} - \ddot{\ell}_{it}^0 = \tilde{\dddot{\ell}}_{it}(\Delta_{\beta}'X_{it} + \tilde{\Lambda}_i'\Delta_{\gamma_t} + \Delta_{\lambda_i}'\tilde{\gamma_t}), \end{align}\] which follows from the first-order Taylor expansion. Here, \(\tilde{\dddot{\ell}}_{it}\) denotes the third-order derivative of \(\ell_{it}\) evaluated at \((\tilde{\beta}, \tilde{\Lambda}, \tilde{\Gamma})\), lying on the line segment between \((\beta, \Lambda, \Gamma)\) and the normalized true parameters \((\beta_0, \Lambda_0^G, \Gamma_0^G)\). Since the parameter space of \((\beta_0, \Lambda_0, \Gamma_0)\) is bounded and all singular values of \(G\) are uniformly bounded and strictly positive wpa1, the normalized true parameters \((\beta, \Lambda_0^G, \Gamma_0^G)\) lie within a bounded space wpa1. Thus, \((\tilde{\beta}, \tilde{\Lambda}, \tilde{\Gamma})\) also lie in a bounded space.
Second, additional effort is required to address the issue that, for some \(i\) and \(t\), the nuisance parameter estimator may significantly differ from the true nuisance parameter. This challenge arises because the definition of neighborhood \(\mathcal{B}_{\delta_{NT}}\) 7 only ensures that, for any \((\Lambda, \Gamma)\) in the neighborhood, the distance between \((\Lambda, \Gamma)\) and \((\Lambda_0^G, \Gamma^G_0)\), shrinks in the Frobenius norm, However, it does not guarantee that the nuisance parameter estimates are accurate for each \(i\) and \(t\). Therefore, for each \((\beta, \Lambda, \Gamma)\in \mathcal{B}_{\delta_{NT}}\), we divide \(\Lambda\) into two parts, with subscripts for each part respectively given as \[\begin{align} \mathcal{I}_{NT} = \left\{1\leq i\leq N: \|\lambda_{i} - {\lambda}^G_{0, i} \| \leq \frac{1}{\sqrt{T} \delta_{NT} } \right\}, \quad \mathcal{I}^c_{NT} = \{1, \ldots, N\} \backslash \mathcal{I}_{NT}. \end{align}\] Similarly, we divide \(\Gamma\) into two parts, with subscripts for each part respectively given as \[\begin{align} \mathcal{T}_{NT} = \left\{1\leq t\leq T: \|\gamma_{t} - {\gamma}^G_{0, t} \| \leq \frac{1}{\sqrt{N}\delta_{NT}} \right\}, \quad \mathcal{T}^c_{NT} = \{1, \ldots, T\} \backslash \mathcal{T}_{NT}. \end{align}\] It is straightforward to observe that, when \(N\sim T\), for any \(i \in \mathcal{I}_{NT}\), the distance between \(\lambda_{i}\) and \(\lambda^G_{0, i}\) converges to zero at the rate \(T^{-1/8}\log(NT)\). Similarly, for any \(t \in \mathcal{T}_{NT}\), the distance between \(\gamma_t\) and \(\gamma^G_{0, t}\) also converges to zero as \(N, T\rightarrow\infty\) at the rate \(N^{-1/8}\log(NT)\). Note that the notations of \(\mathcal{I}_{NT}\) and \(\mathcal{T}_{NT}\) are slightly imprecise, because for different \((\beta, \Lambda, \Gamma)\in \mathcal{B}_{\delta_{NT}}\), the corresponding index sets may differ. Nevertheless, following from the Frobenius norm convergence rate (Theorem 1 and Theorem 5), we conclude that, with probability approaching \(1\), for all \((\beta, \Lambda, \Gamma)\in \mathcal{B}_{\delta_{NT}}\), the sizes of \(\mathcal{I}^c_{NT}\) and \(\mathcal{T}^c_{NT}\) are uniformly bounded by \[\begin{align} \label{eq:size95division} |\mathcal{I}^c_{NT}| \lesssim NT \delta_{NT}^4, \quad |\mathcal{T}^c_{NT}| \lesssim NT \delta_{NT}^4. \end{align}\tag{39}\] Therefore, we continue to use this imprecise notation in the following text for simpler notations. In addition, our construction 39 implies that the sizes of \(\mathcal{I}^c_{NT}\) and \(\mathcal{T}^c_{NT}\) are at most of order \(\sqrt{N}(\log(NT))^4\) when \(N\sim T\). Hence, the number of “non-convergent” nuisance estimates grows at a much slower rate relative to \(N\) and \(T\).
Third, for notational simplicity, we rescale the Hessian matrix \(\mathcal{H}_{NT}\) as \(\mathcal{H} = NT\mathcal{H}_{NT}\) and suppress its dependence on \((\beta, \Lambda, \Gamma)\) when it does not cause ambiguity. Since \(\mathcal{H}\) differs from \(\mathcal{H}_{NT}\) only by a factor \(NT\), we focus on studying the property of \(\mathcal{H}\) in the following text. Furthermore, we define \[\begin{align} H = \begin{pmatrix} H_{\beta\beta'} & H_{\beta\lambda'} & H_{\beta\gamma'} \\ H_{\lambda\beta'} & H_{\lambda\lambda'} & H_{\lambda\gamma'} \\ H_{\gamma \beta'} & H_{\gamma\lambda'} & H_{\gamma\gamma'} \end{pmatrix}, \quad F = \begin{pmatrix} 0 & 0 & 0 \\ 0 & 0 & F_{\lambda\gamma'} \\ 0 & F_{\gamma \lambda'} & 0 \end{pmatrix} \quad V = \begin{pmatrix} 0 & 0 & 0 \\ 0 & V_{\lambda\lambda'} & V_{\lambda\gamma'} \\ 0 & V_{\gamma\lambda'} & V_{\gamma\gamma'} \end{pmatrix}. \end{align}\] Here, \(H + F\) is the Hessian of the negative log-likelihood function \(\mathcal{L}_{NT}\), and \(V\) is Hessian of the penalty term. Specifically, \[\begin{align} H_{\beta\beta'} & = -\sum_{i=1}^{N}\sum_{t = 1}^{T} \ddot{\ell}_{it} X_{it} X'_{it}, \\ H_{\beta\lambda'} & = -\left[\sum_{t=1}^{T} \ddot{\ell}_{it} X_{it} \gamma'_t \right]_{i = 1,2,\ldots, N}, \\ H_{\beta\gamma'} & = -\left[\sum_{i=1}^{N} \ddot{\ell}_{it}X_{it}\lambda'_i\right]_{t = 1,2,\ldots, T}, \\ H_{\lambda\lambda'} & = -\mathrm{diag}\left\{\left[\sum_{t=1}^{T} \ddot{\ell}_{it}\gamma_t \gamma_t'\right]_{i = 1,2,\ldots, N}\right\}, \\ H_{\gamma\gamma'} & = -\mathrm{diag}\left\{\left[\sum_{i=1}^{N} \ddot{\ell}_{it}\lambda_i \lambda_i'\right]_{t= 1,2,\ldots, T}\right\}, \\ H_{\lambda\gamma'} & = \left[-\ddot{\ell}_{it}\gamma_t\lambda_i'\right]_{i = 1,2,\ldots, N, t = 1,2,\ldots, T}, \\ V_{\lambda\lambda'} & = \frac{T}{N}\left[ \lambda_i \lambda_{i'}' \right]_{i, i' = 1,2,\ldots, N}, \\ V_{\lambda\gamma'} & = \left[ -\lambda_i \gamma_{t}' \right]_{i =1,2,\ldots, N, t = 1,2,\ldots, T}, \\ V_{\gamma\gamma'} & = \frac{N}{T}\left[ \gamma_t \gamma_{t'}'\right]_{t, t' = 1,2,\ldots, T} , \\ F_{\lambda\gamma'} & = \left[-\dot{\ell}_{it} \mathbb{I}_{R}\right]_{i = 1,2,\ldots, N, t = 1,2,\ldots, T}. \end{align}\] Here, \(H_{\beta\beta'}\) is a \(d_X\times d_X\) matrix, \(H_{\beta\lambda'}\) is a \(d_X\times NR\) matrix consisting of \(N\) blocks, each of size \(d_X\times R\), and \(H_{\beta\gamma'}\) is a \(d_X\times TR\) matrix with \(T\) blocks, each of size \(d_X\times R\). \(H_{\lambda\lambda'}, V_{\lambda\lambda'}\) are \(NR\times NR\) matrices consisting of \(N^2\) blocks, each of size \(R\times R\), \(H_{\gamma\gamma'}, V_{\gamma\gamma'}\) are b \(TR\times TR\) matrices with \(T^2\) blocks, each of size \(R\times R\). In addition, \(H_{\lambda\gamma'}, V_{\lambda\gamma'}, F_{\lambda\gamma'}\) are \(NR\times TR\) matrices consisting of \(NT\) blocks, each with size \(R\times R\).
We use \(\hat{V}\) to denote the matrix \(V\) evaluated at \(\hat{\lambda}_{\mathrm{nuc},i}\) and \(\hat{\gamma}_{\mathrm{nuc},t}\). Specifically, \[\begin{align} \hat{V}_{\lambda\lambda'} & = \frac{T}{N}\left[ \hat{\lambda}_{\mathrm{nuc},i} \hat{\lambda}_{\mathrm{nuc},i'}' \right]_{i, i' = 1,2,\ldots, N\ldots}, \\ \hat{V}_{\lambda\gamma'} & = \left[ -\hat{\lambda}_{\mathrm{nuc},i} \hat{\gamma}_{\mathrm{nuc},t}' \right]_{i =1,2,\ldots, N, t = 1,2,\ldots, T}, \\ \hat{V}_{\gamma\gamma'} & = \frac{N}{T}\left[ \hat{\gamma}_{\mathrm{nuc},t} \hat{\gamma}_{\mathrm{nuc},t'}'\right]_{t, t' = 1,2,\ldots, T}. \end{align}\]
In addition, we use \(H_0, F_0, V_0\) when the matrices \(H, F, V\) are evaluated at the normalized true value \((\beta_0, \Lambda_0^G, \Gamma_0^G)\), for example, \(H_0 = H(\beta_0, \Lambda^G_0, \Gamma^G_0)\). We also use \(\mathbb{E}_0(\cdot)\) to denote the conditional expectation \(\mathbb{E}_{X, \Lambda_0, \Gamma_0}(\cdot)\) when \(X\) is strictly exogenous, or the conditional expectation \(\mathbb{E}_{Z, \Lambda_0, \Gamma_0}(\cdot)\) when we consider predetermined covariates. It is easy to verify that \[\begin{align} \label{eq:VF} \mathbb{E}_{0} F_0 = 0. \end{align}\tag{40}\]
By standard calculus, \(\mathcal{H} \in \mathbb{R}^{(d_X + R(N+T)) \times (d_X + R(N+T))}\) admits the following decomposition: \[\begin{align} \label{eq:decomposition} \mathcal{H} = H + F + \hat{V} . \end{align}\tag{41}\] We further decompose \(H\) as \(H = \tilde{H} +\tilde{H}^c\) based on partitions \(\mathcal{I}_{NT}\) and \(\mathcal{T}_{NT}\) such that \[\begin{align} \tilde{H}_{\beta\beta'} & = \sum_{i \in \mathcal{I}_{NT}, t \in \mathcal{T}_{NT}} (-\ddot{\ell}_{it}) X_{it} X'_{it}, \\ \tilde{H}_{\beta\lambda'} & = -\left[\boldsymbol{1}(i\in \mathcal{I}_{NT})\sum_{t\in \mathcal{T}_{NT}} \ddot{\ell}_{it} X_{it} \gamma'_t \right]_{i = 1,2,\ldots, N}, \\ \tilde{H}_{\beta\gamma'} & = -\left[\boldsymbol{1}(t\in \mathcal{T}_{NT})\sum_{i\in \mathcal{I}_{NT}} \ddot{\ell}_{it}X_{it}\lambda'_i\right]_{t = 1,2,\ldots, T}, \\ \tilde{H}_{\lambda\lambda'} & = -\mathrm{diag}\left\{\left[\boldsymbol{1}(i\in \mathcal{I}_{NT})\sum_{t\in \mathcal{T}_{NT}} \ddot{\ell}_{it}\gamma_t \gamma_t'\right]_{i = 1,2,\ldots, N}\right\}, \\ \tilde{H}_{\gamma\gamma'} & = -\mathrm{diag}\left\{\left[\boldsymbol{1}(t\in \mathcal{T}_{NT})\sum_{i\in \mathcal{I}_{NT}} \ddot{\ell}_{it}\lambda_i \lambda_i'\right]_{t= 1,2,\ldots, T}\right\}, \\ \tilde{H}_{\lambda\gamma'} & = -\left[\boldsymbol{1}(i \in \mathcal{I}_{NT}, t\in \mathcal{T}_{NT})\ddot{\ell}_{it}\gamma_t\lambda_i'\right]_{i = 1,2,\ldots, N, t = 1,2,\ldots, T} . \end{align}\] The matrix \(V\) can be decomposed in the same way, \(V = \tilde{V} + \tilde{V}^c\), where \[\begin{align} \tilde{V}_{\lambda\lambda'} & = \left[\frac{T}{N}\boldsymbol{1}(i, i'\in \mathcal{I}_{NT})\left(\lambda_{i} \lambda_{i'}' \right)\right]_{i, i' = 1,2,\ldots, N}, \\ \tilde{V}_{\gamma\gamma'} & = \left[\frac{N}{T}\boldsymbol{1}(t, t' \in \mathcal{T}_{NT})\left(\gamma_{ t} \gamma_{ t'}' \right)\right]_{t, t' = 1,2,\ldots, T} , \\ \tilde{V}_{\lambda\gamma'} & = \left[-\boldsymbol{1}(i\in \mathcal{I}_{NT}, t\in \mathcal{T}_{NT}) \lambda_{i} \gamma_{t}'\right]_{i =1,2,\ldots, N, t = 1,2,\ldots, T}. \end{align}\]
As Theorem 6 generalizes Theorem 2 to include predetermined covariates, we provide only the proof of the former, noting that the latter follows as a special case.
Proof of Theorem 6. It suffices to study the properties of \(\mathcal{H}\) on \(\mathcal{B}_{\delta_{NT}}\). Consider the following decomposition: \[\label{eq:Hessian95decomposition} \begin{align} \mathcal{H} = & \underbrace{ \begin{pmatrix} \mathbb{E}_{0} H_{0, \beta\beta'} & \mathbb{E}_{0} \tilde{H}_{0, \beta\lambda'} & \mathbb{E}_{0} \tilde{H}_{0, \beta\gamma'} \\ \mathbb{E}_{0} \tilde{H}_{0, \lambda\beta'} & \mathbb{E}_0 \tilde{H}_{0, \lambda\lambda'} & \mathbb{E}_0 \tilde{H}_{0, \lambda\gamma'} \\ \mathbb{E}_{0} \tilde{H}_{0, \gamma\beta'} & \mathbb{E}_0 \tilde{H}_{0, \gamma\lambda'} & \mathbb{E}_0 \tilde{H}_{0, \gamma\gamma'} \end{pmatrix} + \begin{pmatrix} 0 & 0 & 0 \\ 0 & \tilde{H}^c_{ \lambda\lambda'} & 0 \\ 0 & 0 & \tilde{H}^c_{ \gamma\gamma'} \end{pmatrix} + \mathbb{E}_0 \hat{V} }_{S_1 } \\ & + \underbrace{ \begin{pmatrix} H_{\beta\beta'} - \mathbb{E}_{0} H_{0, \beta\beta'} & H_{\beta\lambda'} - \mathbb{E}_{0} \tilde{H}_{0, \beta\lambda'} & H_{ \beta\gamma'} - \mathbb{E}_{0} \tilde{H}_{0, \beta\gamma'} \\ H_{\lambda\beta'} - \mathbb{E}_{0} \tilde{H}_{ 0, \lambda\beta'} & 0 & 0 \\ H_{\gamma\beta'} - \mathbb{E}_{0} \tilde{H}_{0, \gamma\beta'} & 0 & 0 \end{pmatrix} }_{S_2} \\ & + \underbrace{ \begin{pmatrix} 0 & 0 & 0 \\ 0 & \tilde{H}_{\lambda\lambda'} - \mathbb{E}_0 \tilde{H}_{0, \lambda\lambda'} & H_{ \lambda\gamma'} - \mathbb{E}_0 \tilde{H}_{0, \lambda\gamma'} \\ 0 & H_{ \gamma\lambda'} - \mathbb{E}_0 \tilde{H}_{0, \gamma\lambda'} & \tilde{H}_{\gamma\gamma'} - \mathbb{E}_0 \tilde{H}_{0, \gamma\gamma'} \end{pmatrix} }_{S_3} + (\hat{V} - \mathbb{E}_0 \hat{V}) + F. \end{align}\tag{42}\] Let us first examine the first term. It is straightforward to verify that \(S_1\) admits the following decomposition: \[\label{eq:thm95convexity951} \begin{align} S_1 & = \mathbb{E}_0\tilde{H}_{0} + \mathbb{E}_0 \tilde{\hat{V}} + \underbrace{ \; \begin{pmatrix} \mathbb{E}_0 \tilde{H}^c_{\beta\beta'} & 0 & 0 \\ 0 & 0& 0\\ 0 & 0 & 0 \end{pmatrix} }_{\geq 0} + \begin{pmatrix} 0 & 0 & 0 \\ 0 & \tilde{H}^c_{\lambda\lambda'} & 0\\ 0 & 0 & \tilde{H}^c_{\gamma\gamma'} \end{pmatrix} + \mathbb{E}_0 \tilde{\hat{V}}^c \\ & \geq \mathbb{E}_0\tilde{H}_{0} + \mathbb{E}_0 \tilde{\hat{V}} + \begin{pmatrix} 0 & 0 & 0 \\ 0 & \tilde{H}^c_{\lambda\lambda'} & 0\\ 0 & 0 & \tilde{H}^c_{\gamma\gamma'} \end{pmatrix} + \mathbb{E}_0 \tilde{\hat{V}}^c. \end{align}\tag{43}\] Since matrix \(\mathbb{E}_0\tilde{H}_{0} + \mathbb{E}_0 \tilde{\hat{V}}\) can be regarded as the population Hessian matrix indexed by \(i \in \mathcal{I}_{NT}\) and \(t \in \mathcal{T}_{NT}\), we obtain \[\label{eq:thm95convexity952} \begin{align} &\mathbb{E}_0\tilde{H}_{0} + \mathbb{E}_0 \tilde{\hat{V}} \\ \stackrel{\text{(i)}}{\geq} & C \begin{pmatrix} \left(NT - |\mathcal{T}^c_{NT}| N - |\mathcal{I}^c_{NT}| T\right) \mathbb{I}_{d_X} & 0 & 0 \\ 0 & \left(T - |\mathcal{T}^c_{NT}|\right)\mathrm{diag}\{D_1, \ldots, D_N\} & 0 \\ 0 & 0 & \left(N - |\mathcal{I}^c_{NT}|\right)\mathrm{diag}\{D_{N+1}, \ldots, D_{N+T}\} \end{pmatrix} \\ \stackrel{\text{(ii)}}{\geq} & \frac{1}{2} C \begin{pmatrix} NT \mathbb{I}_{d_X} & 0 & 0 \\ 0 & T\mathrm{diag}\{D_1, \ldots, D_N\} & 0 \\ 0 & 0 & N\mathrm{diag}\{D_{N+1}, \ldots, D_{N+T}\} \end{pmatrix}, \end{align}\tag{44}\] where \(D_i = 1_{\{i \in \mathcal{I}_{NT}\}} \mathbb{I}_{R}\) for any \(i=1,2,\ldots, N\), and \(D_{N+t} = 1_{\{t \in \mathcal{T}_{NT}\}} \mathbb{I}_{R}\) for any \(t=1,2,\ldots, T\). Inequality (i) follows from Assumption 5, and inequality (ii) follows from the asymptotic assumption on \(N\) and \(T\). In addition, by Lemma 7, we conclude that, there exists a constant \(B_1>0\) such that, wpa1, \[\label{eq:thm95convexity953} \begin{align} \begin{pmatrix} 0 & 0 & 0 \\ 0 & \tilde{H}^c_{\lambda\lambda'} & 0\\ 0 & 0 & \tilde{H}^c_{\gamma\gamma'} \end{pmatrix} \geq B_1 \min\{N, T\} \mathrm{diag}\left\{0_{d_X}, \mathbb{I}_R - D_1, \ldots, \mathbb{I}_R - D_N, \mathbb{I}_R - D_{N+1}, \ldots, \mathbb{I}_R -D_{N+T}\right\}. \end{align}\tag{45}\] Combining equations 43 , 44 , 45 , and \(\|\mathbb{E}_0 \tilde{\hat{V}}^c \|_{\mathrm{op}} = o_p\left(\min\{N, T\}\right)\)23, we conclude that for any \((\beta, \Lambda, \Gamma)\in \mathcal{B}_{\delta_{NT}}\), \(S_1\) is locally convex wpa1. In addition, let \(B_2 = \min\{\frac{1}{2}C, B_1\}\) irrelevant with \(N, T\), for any \((\beta, \Lambda, \Gamma)\in \mathcal{B}_{\delta_{NT}}\), \(S_1\) admits asymptotic block structure, i.e., \[\label{eq:thm95convexity954} \begin{align} S_1 \geq B_2 \begin{pmatrix} NT\mathbb{I}_{d_X} & 0 & 0 \\ 0 & T\mathbb{I}_{NR} & 0 \\ 0 & 0 & N\mathbb{I}_{TR} \end{pmatrix}, \quad \text{wpa1}. \end{align}\tag{46}\] Therefore, the optimization problem is strongly convex wpa1 uniformly within \(\mathcal{B}_{\delta_{NT}}\) if the impact (or equivalently, the maximum singular values) of the permutation terms, \(S_2\), \(S_3\), \(\hat{V}-\mathbb{E}_0\hat{V}\), and \(F\), is asymptotically negligible (within \(\mathcal{B}_{\delta_{NT}}\)) compared to \(S_1\) wpa1. Using Lemma 9, we prove that \(\frac{1}{2}S_1 + S_2\) is positive definite. In addition, by Lemma 10 and Lemma 11, we have \[\begin{align} \label{eq:thm95convexity956} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|S_3\|_{\mathrm{op}} = o_p(\min\{N, T\}) , \text{ } \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|F\|_{\mathrm{op}} = o_p(\min\{N, T\}). \end{align}\tag{47}\] Using Lemma 12, and noting that \((\hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}}) \in \mathcal{B}_{\delta_{NT}}\), we have \[\begin{align} \label{eq:thm95convexity957} \|\hat{V}-\mathbb{E}_0\hat{V}\|_{\mathrm{op}} \lesssim \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|V - V_0\|_{\mathrm{op}} = o_p(\min\{N, T\}). \end{align}\tag{48}\] (We defer the relevant lemmas and their proofs to the end of this subsection). Therefore, combining 46 —48 and applying Weyl’s theorem, we conclude that, with probability approaching \(1\), for any \((\beta, \Lambda, \Gamma)\in \mathcal{B}_{\delta_{NT}}\), \[\begin{align} \mathcal{H} \geq \frac{1}{3} B_2 \begin{pmatrix} NT\mathbb{I}_{d_X} & 0 & 0 \\ 0 & T\mathbb{I}_{NR} & 0 \\ 0 & 0 & N\mathbb{I}_{TR} \end{pmatrix}. \end{align}\] Letting \(c_5 = \frac{1}{3} B_2\) completes the proof.
Lemma 7. Under the assumptions stated in Theorem 6, there exists a constant \(B_1>0\) independent of \(N, T\) such that \[\begin{align} \inf_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \begin{pmatrix} 0 & 0 & 0 \\ 0 & H_{\lambda\lambda'} & 0\\ 0 & 0 & H_{\gamma\gamma'} \end{pmatrix} \geq B_1 \begin{pmatrix} 0_{d_X} & 0 & 0 \\ 0 & T\mathbb{I}_{NR} & 0 \\ 0 & 0 & N\mathbb{I}_{TR} \end{pmatrix}, \quad \text{wpa1}. \end{align}\]
Proof of Lemma 7. Recall \[\begin{align} H_{\lambda\lambda'} & = -\mathrm{diag}\left\{\left[\sum_{t=1}^{T} \ddot{\ell}_{it}\gamma_t \gamma_t'\right]_{i = 1,2,\ldots, N}\right\}, \\ H_{\gamma\gamma'} & = -\mathrm{diag}\left\{\left[\sum_{i=1}^{N} \ddot{\ell}_{it}\lambda_i \lambda_i'\right]_{t= 1,2,\ldots, T}\right\} . \end{align}\] For any \(i\)-th diagonal block of \(H_{\lambda\lambda'}\), with probability approaching \(1\), we have \[\begin{align} \inf_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sigma_{\min}\left(\sum_{t=1}^{T} -\ddot{\ell}_{it}\gamma_t \gamma_t' \right) & \stackrel{\text{(i)}}{\gtrsim} \inf_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sigma_{\min}\left(\sum_{t=1}^{T} \gamma_t \gamma_t' \right) \\ & \gtrsim \inf_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sigma_{\min}\left(\Gamma_0^{G\prime }\Gamma_0^G \right) - \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\|\Gamma'\Gamma - \Gamma_0^{G\prime }\Gamma_0^G \right\|_{\mathrm{op}} \\ & \stackrel{\text{(ii)}}{\gtrsim} T - \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\|\Gamma'\Gamma - \Gamma_0^{G\prime }\Gamma_0^G \right\|_{\mathrm{F}} \\ & \stackrel{\text{(iii)}}{\gtrsim} T - \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|\Gamma\|_{\mathrm{F}}\left\|\Gamma - \Gamma_0^G \right\|_{\mathrm{F}} - \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|\Gamma_0^G\|_{\mathrm{F}}\left\|\Gamma - \Gamma_0^G \right\|_{\mathrm{F}} \\ & \stackrel{\text{(iv)}}{\gtrsim} T - \sqrt{T} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\|\Gamma - \Gamma_0^G \right\|_{\mathrm{F}} \\ & \stackrel{\text{(v)}}{\gtrsim} T - T \delta_{NT} \\ & \gtrsim T, \end{align}\] where inequality (i) uses Assumption 4[item:smoothing95pre], inequality (ii) follows from Assumption 4[item:strong95factors95pre], inequality (iii) follows from the Cauchy-Schwarz inequality, inequality (iv) is based on the uniform boundedness of \(\Gamma, \Gamma_0^G\), and inequality (v) follows from Theorem 5. By the same argument, we also obtain \[\begin{align} \inf_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sigma_{\min}\left(\sum_{i=1}^{N} -\ddot{\ell}_{it}\lambda_i \lambda_i' \right) \gtrsim N, \quad \text{wpa1}. \end{align}\] Therefore, we are able to find a constant \(B_1>0\), independent of \(N, T\), and the proof is complete.
Lemma 8. Under the assumptions of Theorem 6, we have \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\|H_{\beta\beta'} - \mathbb{E}_0H_{0, \beta\beta'}\right\|_{\mathrm{op}} = O_p(NT\delta_{NT}). \end{align}\]
Proof of Lemma 8. Note that \[\label{eq:lemma:bound95Hbb951} \begin{align} H_{\beta\beta'} - \mathbb{E}_0H_{0, \beta\beta'} & = -\sum_{i=1}^{N}\sum_{t=1}^{T} (\ddot{\ell}_{it}X_{it}X_{it}') + \sum_{i=1}^{N}\sum_{t=1}^{T} \mathbb{E}_0 (\ddot{\ell}_{it}^0 X_{it}X_{it}') \\ & = -\sum_{i=1}^{N}\sum_{t=1}^{T}(\ddot{\ell}_{it}X_{it}X_{it}' - \ddot{\ell}_{it}^0 X_{it}X_{it}') -\sum_{i=1}^{N}\sum_{t=1}^{T}(\ddot{\ell}^0_{it}X_{it}X_{it}' - \mathbb{E}_0 (\ddot{\ell}_{it}^0 X_{it}X_{it}')). \end{align}\tag{49}\] We have \[\label{eq:lemma:bound95Hbb952} \begin{align} & \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\|\sum_{i=1}^{N}\sum_{t=1}^{T}(\ddot{\ell}_{it}X_{it}X_{it}' - \ddot{\ell}_{it}^0 X_{it}X_{it}') \right\|_{\mathrm{op}}\\ \stackrel{\text{(i)}}{=} & \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\| \sum_{i=1}^{N}\sum_{t=1}^{T} \underbrace{\tilde{\dddot{\ell}}_{it} \left(\Delta_{\beta}'X_{it} + \tilde{\gamma}'_{t}\Delta_{\lambda_i} + \tilde{\lambda}_{i}'\Delta_{\gamma_t} \right)}_{\tilde{\Delta}_{Y^*_{it}}} X_{it}X_{it}' \right\|_{\mathrm{op}}\\ \stackrel{\text{(ii)}}{\leq} & \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\| \sum_{i=1}^{N}\sum_{t=1}^{T}\tilde{\dddot{\ell}}_{it} (\Delta_{\beta}'X_{it}) X_{it}X_{it}' \right\|_{\mathrm{op}} \\ & + \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\| \sum_{i=1}^{N}\sum_{t=1}^{T}\tilde{\dddot{\ell}}_{it} \left( \tilde{\gamma}'_{t}\Delta_{\lambda_i} + \tilde{\lambda}_{i}'\Delta_{\gamma_t}\right) X_{it}X_{it}' \right\|_{\mathrm{op}} \\ \stackrel{\text{(iii)}}{\lesssim} & NT \delta_{NT} + \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\| \sum_{i=1}^{N}\sum_{t=1}^{T}\tilde{\dddot{\ell}}_{it} \left( \tilde{\gamma}'_{t}\Delta_{\lambda_i} + \tilde{\lambda}_{i}'\Delta_{\gamma_t}\right)X_{it}X_{it}' \right\|_{\mathrm{op}} \\ \stackrel{\text{(iv)}}{\lesssim} & \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{t=1}^{T}\left\| \sum_{i=1}^{N}\tilde{\dddot{\ell}}_{it} \tilde{\gamma}_{t}'\Delta_{\lambda_i} X_{it}X_{it}' \right\|_{\mathrm{op}} + \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{i=1}^{N}\left\| \sum_{t=1}^{T}\tilde{\dddot{\ell}}_{it} \tilde{\lambda}_{i}'\Delta_{\gamma_t} X_{it}X_{it}' \right\|_{\mathrm{op}}+ NT \delta_{NT} \\ \stackrel{\text{(v)}}{\lesssim} & \sum_{t=1}^{T} \sqrt{N}\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|\Lambda - \Lambda^G_0\|_{\mathrm{F}} + \sum_{i=1}^{N} \sqrt{T}\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|\Gamma - \Gamma^G_0\|_{\mathrm{F}} + NT\delta_{NT} \\ \stackrel{\text{(vi)}}{\lesssim} & NT \delta_{NT}. \end{align}\tag{50}\] Equation (i) follows from first-order Taylor expansion, where \(\tilde{\dddot{\ell}}_{it}\) denotes the third-order derivative of \(\ell_{it}\) evaluated at \((\tilde{\beta}, \tilde{\Lambda}, \tilde{\Gamma})\), a point on the line segment between \((\beta, \Lambda, \Gamma)\) and the true parameters. As discussed earlier, \((\tilde{\beta}, \tilde{\Lambda}, \tilde{\Gamma})\) also lies in a bounded space, as both \((\tilde{\beta}, \tilde{\Lambda}, \tilde{\Gamma})\) and true parameter lie in bounded spaces. Equation (ii) follows from the triangle inequality. Equation (iii) is derived based on the uniform boundedness of \(\tilde{\dddot{\ell}}_{it}\), \(X\), and \((\tilde{\beta}, \tilde{\Lambda}, \tilde{\Gamma})\) (Assumption 4[item:boundedness95pre] and [item:smoothing95pre]), and the bound \(\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|\beta - \beta_0\| \leq \delta_{NT}\). Inequality (iv) again follows from the triangle inequality. Inequality (v) is based on the uniform boundedness of \(\tilde{\dddot{\ell}}_{it}\), \(X\), and \((\tilde{\beta}, \tilde{\Lambda}, \tilde{\Gamma})\), along with the Cauchy-Schwarz inequality. Inequality (vi) follows from the bounds \(\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|\Lambda - \Lambda^G_0\|_{\mathrm{op}}\leq \sqrt{N}\delta_{NT}\) and \(\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|\Gamma - \Gamma^G_0\|_{\mathrm{op}}\leq \sqrt{T}\delta_{NT}\). A byproduct that will be frequently used in the subsequent analysis is \[\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{i=1}^{N}\sum_{t=1}^{T} \tilde{\Delta}_{Y^*_{it}}^2\leq NT\delta_{NT}^2,\] which can be established using a similar method.
Under Assumption 4[item:sampling95pre] and 4[item:boundedness95pre], we apply [40] to establish that for any \((\beta, \Lambda, \Gamma)\), \[\label{eq:lemma:bound95Hbb953} \begin{align} \left\|\sum_{i=1}^{N}\sum_{t=1}^{T}(\ddot{\ell}^0_{it}X_{it}X_{it}' - \mathbb{E}_0 (\ddot{\ell}_{it}^0 X_{it}X_{it}'))\right\|_{\mathrm{op}} = O_p(\sqrt{NT}). \end{align}\tag{51}\] Combining 49 , 50 , 51 , and \(\delta_{NT} \lesssim \log(NT)/\sqrt{\min\{N, T\}}\), we obtain \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\|H_{\beta\beta'} - \mathbb{E}_0H_{0, \beta\beta'}\right\|_{\mathrm{op}} = O_p(NT\delta_{NT}). \end{align}\]
Lemma 9. Under the conditions of Theorem 2, for any \((\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}\), the matrix \(\frac{1}{2}S_1 + S_2\) is positive definite wpa1.
Proof of Lemma 9. As we have demonstrated in the Proof of Theorem 6, there exists a constant \(B_2\) such that \[\begin{align} S_1 \geq B_2 \begin{pmatrix} NT\mathbb{I}_{d_X} & 0 & 0 \\ 0 & T\mathbb{I}_{NR} & 0 \\ 0 & 0 & N\mathbb{I}_{TR} \end{pmatrix} , \quad \text{wpa1. } \end{align}\] It follows that \[\begin{align} \frac{1}{2} S_1 + S_2 & \geq \frac{B_2}{2} \begin{pmatrix} NT\mathbb{I}_{d_X} & 0 & 0 \\ 0 & T\mathbb{I}_{NR} & 0 \\ 0 & 0 & N\mathbb{I}_{TR} \end{pmatrix} + \begin{pmatrix} H_{\beta\beta'} - \mathbb{E}_{0} H_{0, \beta\beta'} & H_{\beta\lambda'} - \mathbb{E}_{0} \tilde{H}_{0, \beta\lambda'} & H_{ \beta\gamma'} - \mathbb{E}_{0} \tilde{H}_{0, \beta\gamma'} \\ H_{\lambda\beta'} - \mathbb{E}_{0} \tilde{H}_{ 0, \lambda\beta'} & 0 & 0 \\ H_{\gamma\beta'} - \mathbb{E}_{0} \tilde{H}_{0, \gamma\beta'} & 0 & 0 \end{pmatrix} \\ & \stackrel{\text{(i)}}{\geq} \begin{pmatrix} \frac{B_2}{3} NT \mathbb{I}_{d_X} & H_{\beta\lambda'} - \mathbb{E}_{0} \tilde{H}_{0, \beta\lambda'} & H_{ \beta\beta'} - \mathbb{E}_{0} \tilde{H}_{0, \beta\gamma'} \\ H_{\lambda\beta'} - \mathbb{E}_{0} \tilde{H}_{ 0, \lambda\beta'} & \frac{B_2}{2} T\mathbb{I}_{NR} & 0 \\ H_{\gamma\beta'} - \mathbb{E}_{0} \tilde{H}_{0, \gamma\beta'} & 0 & \frac{B_2}{2} N\mathbb{I}_{TR} \end{pmatrix} , \quad \text{wpa1}, \end{align}\] where the inequality (i) follows from Lemma 8, which ensures that \(\left\|H_{\beta\beta'} - \mathbb{E}_0H_{0, \beta\beta'}\right\|_{\mathrm{op}}= o_p(NT)\). Therefore, the matrix \(\frac{1}{2}S_1 + S_2\) is positive definite if its Schur complement \[\begin{align} \begin{pmatrix} \frac{B_2}{2} T\mathbb{I}_{NR} & 0 \\ 0 & \frac{B_2}{2} N \mathbb{I}_{TR} \end{pmatrix} - \frac{3}{NT B_2} \begin{pmatrix} H_{\lambda\beta'} - \mathbb{E}_{0} \tilde{H}_{0, \lambda\beta'} \\ H_{\gamma\beta'} - \mathbb{E}_{0} \tilde{H}_{0, \gamma\beta'} \end{pmatrix} \begin{pmatrix} H_{\beta\lambda'} - \mathbb{E}_{0} \tilde{H}_{0, \beta\lambda'} & H_{\beta\gamma'} - \mathbb{E}_{0} \tilde{H}_{0, \beta\gamma'} \end{pmatrix} \end{align}\] is positive definite. Thus, by Weyl’s Theorem, it suffices to show that \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\| \begin{pmatrix} H_{\lambda\beta'} - \mathbb{E}_0\tilde{H}_{0, \lambda\beta'} & H_{\gamma\beta'} - \mathbb{E}_0\tilde{H}_{0, \gamma\beta'} \end{pmatrix} \right\|_{\mathrm{op}}^2 /(NT) = o_p(\min\{N, T\}). \end{align}\] Observe that \[\begin{align} \ddot{\ell}_{it}X_{it}\gamma_t' = ((\ddot{\ell}_{it}X_{it} - \ddot{\ell}^0_{it}X_{it}) + (\ddot{\ell}^0_{it}X_{it} - \mathbb{E}_0(\ddot{\ell}^0_{it}X_{it}) + \mathbb{E}_0(\ddot{\ell}^0_{it}X_{it})) (\Delta_{\gamma_t} + \gamma^{G\prime}_{0, t}) , \end{align}\] it follows that \[\begin{align} \ddot{\ell}_{it}X_{it}\gamma_t' - \mathbb{E}_0(\ddot{\ell}^0_{it}X_{it}) \gamma^{G\prime}_{0, t} = & \tilde{\dddot{\ell}}_{it}\tilde{\Delta}_{Y^*_{it}} X_{it} \gamma^{G\prime}_{0, t} + (\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0 (\ddot{\ell}^{0}_{it}X_{it}))\gamma^{G\prime}_{0, t} + \tilde{\dddot{\ell}}_{it}\tilde{\Delta}_{Y^*_{it}}X_{it}\Delta'_{\gamma_{t}} \\ & + (\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0( \ddot{\ell}^{0}_{it}X_{it}))\Delta'_{\gamma_{t}} + \mathbb{E}_0(\ddot{\ell}^{0}_{it}X_{it})\Delta'_{\gamma_{t}}, \end{align}\] where we use \(\tilde{\Delta}_{Y^*_{it}} = \Delta_{\beta}'X_{it} + \tilde{\lambda}_{i}'\Delta_{\gamma_t} + \Delta_{\lambda_i}'\tilde{\gamma}_{t}\) to simplify the expression. Here, \(\tilde{\dddot{\ell}}_{it}\) denotes the third-order derivative of \(\ell_{it}\) evaluated at \((\tilde{\beta}, \tilde{\Lambda}, \tilde{\Gamma})\), lying on the line segment between \((\beta, \Lambda, \Gamma)\) and the normalized true parameters. Since both \((\tilde{\beta}, \tilde{\Lambda}, \tilde{\Gamma})\) and true parameters lie in bounded spaces, it follows that \((\tilde{\beta}, \tilde{\Lambda}, \tilde{\Gamma})\) is also in a bounded space. Since the \(X\) and \((\tilde{\beta}, \tilde{\Lambda}, \tilde{\Gamma})\) are uniformly bounded, and \(\ell_{it}\) is four times differentiable, \(\tilde{\dddot{\ell}}_{it}\) is uniformly bounded by extreme value theorem. In the following proof, it suffices to control the Frobenius norm instead of the spectral norm because \[\begin{align} \left\| \begin{pmatrix} H_{\lambda\beta'} - \mathbb{E}_0\tilde{H}_{0, \lambda\beta'} \\ H_{\gamma\beta'} - \mathbb{E}_0\tilde{H}_{0, \gamma\beta'} \end{pmatrix} \right\|_{\mathrm{op}} \leq \left\| \begin{pmatrix} H_{ \lambda\beta'} \\ H_{ \gamma\beta'} \end{pmatrix} - \begin{pmatrix} \mathbb{E}_0 \tilde{H}_{0, \lambda\beta'} \\ \mathbb{E}_0 \tilde{H}_{0, \gamma\beta'} \end{pmatrix} \right\|_{\mathrm{F}} \leq \|H_{\lambda\beta'} - \mathbb{E}_0 \tilde{H}_{0, \lambda\beta'}\|_{\mathrm{F}} + \|H_{\gamma\beta'} - \mathbb{E}_0 \tilde{H}_{ 0, \gamma\beta'}\|_{\mathrm{F}}. \end{align}\] We focus on \(\|H_{\lambda\beta'} - \mathbb{E}_0\tilde{H}_{\lambda\beta'}\|_{\mathrm{F}}\), because the bound of \(\|H_{\gamma\beta'} - \mathbb{E}_0 \tilde{H}_{0, \gamma\beta'}\|_{\mathrm{F}}\) can be obtained using a same method. Note that \[\label{eq:bound95A} \begin{align} \|H_{\lambda\beta'} - \mathbb{E}_0\tilde{H}_{0, \lambda\beta'}\|_{\mathrm{F}}^2 = &\sum_{i\in \mathcal{I}_{NT}} \bigg\| \sum_{t\in \mathcal{T}_{NT}} \tilde{\dddot{\ell}}_{it}\Delta_{Y^*_{it}} X_{it} \gamma^{G\prime}_{0, t} + \sum_{t\in \mathcal{T}_{NT}} (\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0 (\ddot{\ell}^{0}_{it}X_{it}))\gamma^{G\prime}_{0, t} \\ & + \sum_{t\in \mathcal{T}_{NT}} \tilde{\dddot{\ell}}_{it}\tilde{\Delta}_{Y^*_{it}}X_{it}\Delta'_{\gamma_{t}} + \sum_{t\in \mathcal{T}_{NT}} (\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0( \ddot{\ell}^{0}_{it}X_{it}))\Delta'_{\gamma_{t}} \\ & + \sum_{t\in \mathcal{T}_{NT}} \mathbb{E}_0(\ddot{\ell}^{0}_{it}X_{it})\Delta'_{\gamma_{t}} + \sum_{t\in \mathcal{T}^c_{NT}}\ddot{\ell}_{it}X_{it}\gamma_t' \bigg\|_{\mathrm{F}}^2 + \sum_{i\in \mathcal{I}^c_{NT}} \left\|\sum_{t=1}^{T} \ddot{\ell}_{it}X_{it}\gamma_t' \right\|^2_{\mathrm{F}} \\ \leq & 6 \Bigg\{ \underbrace{\sum_{i\in \mathcal{I}_{NT}}\left\|\sum_{t\in \mathcal{T}_{NT}} \tilde{\dddot{\ell}}_{it}\tilde{\Delta}_{Y^*_{it}} X_{it} \gamma^{G\prime}_{0, t} \right\|_{\mathrm{F}}^2}_{A_1} + \underbrace{\sum_{i\in \mathcal{I}_{NT}} \left\| \sum_{t\in \mathcal{T}_{NT}} (\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0 (\ddot{\ell}^{0}_{it}X_{it}))\gamma^{G\prime}_{0, t} \right\|_{\mathrm{F}}^2}_{A_2} \\ & + \underbrace{\sum_{i\in \mathcal{I}_{NT}} \left\|\sum_{t\in \mathcal{T}_{NT}} \tilde{\dddot{\ell}}_{it}\tilde{\Delta}_{Y^*_{it}}X_{it}\Delta'_{\gamma_{t}} \right\|_{\mathrm{F}}^2}_{A_3} + \underbrace{\sum_{i\in \mathcal{I}_{NT}} \left\| \sum_{t\in \mathcal{T}_{NT}} (\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0( \ddot{\ell}^{0}_{it}X_{it}))\Delta'_{\gamma_{t}}\right\|_{\mathrm{F}}^2}_{A_4} \\ & + \underbrace{\sum_{i\in \mathcal{I}_{NT}}\left\|\sum_{t\in \mathcal{T}_{NT}} \mathbb{E}_0(\ddot{\ell}^{0}_{it}X_{it})\Delta'_{\gamma_{t}}\right\|_{\mathrm{F}}^2}_{A_5} + \underbrace{\sum_{i\in \mathcal{I}_{NT}}\left\|\sum_{t\in \mathcal{T}^c_{NT}} \ddot{\ell}_{it}X_{it}\gamma'_{t}\right\|_{\mathrm{F}}^2}_{A_6} \Bigg\} \\ & + \underbrace{\sum_{i\in \mathcal{I}^c_{NT}} \left\|\sum_{t=1}^{T} \ddot{\ell}_{it}X_{it}\gamma_t' \right\|^2_{\mathrm{F}}}_{A_7}. \end{align}\tag{52}\]
The first term \(A_1\) is bounded by \[\label{eq:bound95A951} \begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} A_1 & \stackrel{\text{(i)}}{\leq} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} T \sum_{i=1}^{N}\sum_{t=1}^{T} \left\| \tilde{\dddot{\ell}}_{it} X_{it} \gamma'_{0, t}\right\|_{\mathrm{F}}^2 \tilde{\Delta}_{Y^*_{it}}^2 \\ & \stackrel{\text{(ii)}}{\lesssim} T \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{i=1}^{N}\sum_{t=1}^{T}(X_{it}'\Delta_{\beta})^2 + T \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{i=1}^{N}\sum_{t=1}^{T} \left(\|\Delta_{\lambda_{i}}\|^2 + \Delta_{\gamma_{t}}\|^2\right) \\ & \lesssim NT^2 \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|\beta - \beta_0\|^2 + T^2 \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|\Lambda - \Lambda_0\|^2_{\mathrm{F}} + TN \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|\Gamma - \Gamma_0\|^2_{\mathrm{F}} \\ & \stackrel{\text{(iii)}}{\lesssim} NT^2\delta_{NT}^2, \end{align}\tag{53}\] where inequality (i) follows from the Cauchy-Schwarz inequality, inequality (ii) uses the boundedness of \(X\),\((\beta, \Lambda_0, \Gamma_0)\), \((\tilde{\beta}, \tilde{\Lambda}, \tilde{\Gamma})\), and \(\tilde{\dddot{\ell}}_{it}\), and inequality (iii) follows directly from the definition of \(\mathcal{B}_{\delta_{NT}}\).
For the second term \(A_2\), we proceed as follows: \[\label{eq:bound95A952} \begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} A_2 \leq & \sum_{i=1}^{N} \left\| \sum_{t\in \mathcal{T}_{NT}} (\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0 (\ddot{\ell}^{0}_{it}X_{it}))\gamma_{0, t}^{G\prime}\right\|_{\mathrm{F}}^2 \\ \leq & \sum_{i=1}^{N} \left\| \sum_{t=1}^{T} (\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0 (\ddot{\ell}^{0}_{it}X_{it}))\gamma_{0, t}^{G\prime}- \sum_{t\in \mathcal{T}_{NT}^c } (\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0 (\ddot{\ell}^{0}_{it}X_{it}))\gamma_{0, t}^{G\prime}\right\|_{\mathrm{F}}^2 \\ \leq & 2\sum_{i=1}^{N} \left\| \sum_{t=1}^{T} (\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0 (\ddot{\ell}^{0}_{it}X_{it}))\gamma_{0, t}^{G\prime}\right\|_{\mathrm{F}}^2 + 2 \sum_{i=1}^{N} \left\| \sum_{t\in \mathcal{T}_{NT}^c} (\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0 (\ddot{\ell}^{0}_{it}X_{it}))\gamma_{0, t}^{G\prime}\right\|_{\mathrm{F}}^2. \end{align}\tag{54}\] Since (i) \(\mathbb{E}_0 (\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0 (\ddot{\ell}^{0}_{it}X_{it})) = 0\), (ii) \((\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0 (\ddot{\ell}^{0}_{it}X_{it})) \gamma_{0, t}^{G\prime}\) is uniformly bounded, and (iii) \(\{\ddot{\ell}^{0}_{it} X_{it}\}_{1\leq t\leq T}\) (conditional on \((Z, \Lambda_0, \Gamma_0)\) or \((X, \Lambda_0, \Gamma_0)\)) satisfies the mixing condition in Assumption 4[item:sampling95pre], we apply [40] to obtain \[\begin{align} \sum_{i=1}^{N} \left\| \sum_{t=1}^{T} (\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0 (\ddot{\ell}^{0}_{it}X_{it}))\gamma_{0, t}^{G\prime}\right\|_{\mathrm{F}}^2 = O_p(NT\log(NT)). \end{align}\] Additionally, by the uniform boundedness of \((\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0 (\ddot{\ell}^{0}_{it}X_{it}))\gamma_{0, t}^{G\prime}\), we have \[\begin{align} \sum_{i=1}^{N} \left\| \sum_{t\in \mathcal{T}_{NT}^c} (\ddot{\ell}^{0}_{it}X_{it} - \mathbb{E}_0 (\ddot{\ell}^{0}_{it}X_{it}))\gamma_{0, t}^{G\prime}\right\|_{\mathrm{F}}^2 \lesssim N |\mathcal{T}_{NT}^c|^2 \lesssim N^3T^2 \delta_{NT}^8. \end{align}\] Therefore, \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} A_2 = O_p\left(N^3T^2 \delta_{NT}^8\right) . \end{align}\]
The term \(A_3\) is bounded by \[\label{eq:bound95A953} \begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} A_3 & \stackrel{\text{(i)}}{\leq} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} T \sum_{i=1}^{N} \sum_{t=1}^{T} \| \tilde{\dddot{\ell}}_{it}\tilde{\Delta}_{Y^*_{it}}X_{it} \|^2 \|\Delta_{\gamma_{t}}\|^2 \\ & \leq \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}T\sum_{t=1}^{T} \left\{\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{i=1}^{N} \|\tilde{\dddot{\ell}}_{it}\tilde{\Delta}_{Y^*_{it}}X_{it}\|^2\right\} \|\Delta_{\gamma_{t}}\|^2 \\ & \stackrel{\text{(ii)}}{\lesssim} NT^2\delta^2_{NT} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}} } \sum_{t=1}^{T} \|\Delta_{\gamma_{t}}\|^2 \\ & \stackrel{\text{(iii)}}{\lesssim} NT^3\delta^4_{NT}, \end{align}\tag{55}\] where inequality (i) uses the Cauchy-Schwarz inequality, inequality (ii) relies on the uniform boundedness of \(X\), \((\tilde{\beta}, \tilde{\Lambda}, \tilde{\Gamma})\), and \(\tilde{\dddot{\ell}}_{it}\), as well as the fact that \(\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{t=1}^{T} \tilde{\Delta}_{Y^*_{it}}^2 \leq \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{i=1}^{N}\sum_{t=1}^{T} \tilde{\Delta}_{Y^*_{it}}^2 \lesssim NT \delta_{NT}^2\) (this can be established using the same method as in the Proof of Lemma 8). Inequality (iii) follows directly from the definition of \(\mathcal{B}_{\delta_{NT}}\). Using similar arguments, we obtain that \[\begin{align} \label{eq:bound95A954} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} A_4, A_5 \lesssim NT^2\delta^2_{NT} . \end{align}\tag{56}\] We establish the upper bounds of \(A_6\) and \(A_7\) by the uniform boundedness condition: \[\label{eq:bound95A956} \begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} A_6 &\lesssim N |\mathcal{T}^c_{NT}|^2 \lesssim N^ 3 T^2 \delta_{NT}^8 \\ \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} A_7 &\lesssim T^2 |\mathcal{I}^c_{NT}| \lesssim N T^3 \delta_{NT}^4. \end{align}\tag{57}\] Thus, combining 52 , 53 , 54 , 55 , 56 , and 57 , we conclude that, wpa1, \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|H_{\lambda\beta'} - \mathbb{E}_0H_{0, \lambda\beta'}\|_{\mathrm{F}}^2 & = O_p\left(\max\{NT^2\delta_{NT}^2, NT^3\delta_{NT}^4, N^2 T \delta_{NT}^4, N^ 3 T^2 \delta_{NT}^8 \}\right) \\ & = O_p\left( \max\{N^4, T^4\} \delta_{NT}^4 \right). \end{align}\] Based on similar arguments we also show that, wpa1, \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|H_{\gamma\beta'} - \mathbb{E}_0 H_{0, \gamma\beta'}\|_{\mathrm{F}}^2 & \lesssim = O_p\left( \max\{N^4, T^4\} \delta_{NT}^4 \right). \end{align}\] Therefore, we conclude that \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\| \begin{pmatrix} H_{\lambda\beta'} - \mathbb{E}_0\tilde{H}_{0, \lambda\beta'} & H_{\gamma\beta'} - \mathbb{E}_0\tilde{H}_{0, \gamma\beta'} \end{pmatrix} \right\|_{\mathrm{op}}^2/(NT) & = o_p\left(\min\{\sqrt{N}, \sqrt{T}\}\left(\log(NT)\right)^4\right) \\ & = o_p(\min\{N, T\}). \end{align}\] This complete the proof.
Lemma 10. Under the conditions of Theorem 6, the maximum singular value of \(S_3\) in 42 satisfies \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|S_3\|_{\mathrm{op}} = o_p(\min\{N, T\}). \end{align}\]
Proof of Lemma 10. Recall that \[\begin{align} S_3 = \begin{pmatrix} \tilde{H}_{\lambda\lambda'} & H_{\lambda\gamma'} \\ H_{\gamma\lambda'} & \tilde{H}_{\gamma\gamma'} \end{pmatrix} - \begin{pmatrix} \mathbb{E}_0\tilde{H}_{0, \lambda\lambda'} & \mathbb{E}_0 \tilde{H}_{0, \lambda\gamma'} \\ \mathbb{E}_0\tilde{H}_{0, \gamma\lambda'} & \mathbb{E}_0\tilde{H}_{0, \gamma\gamma'} \end{pmatrix}. \end{align}\] The operator norm of \(S_3\) is bounded by \[\begin{align} \|S_3\|_{\mathrm{op}} \leq \underbrace{ \left\| \begin{pmatrix} \tilde{H}_{\lambda\lambda'} - \mathbb{E}_0 \tilde{H}_{0, \lambda\lambda'} & 0 \\ 0 & \tilde{H}_{\gamma\gamma'} - \mathbb{E}_0 \tilde{H}_{0, \gamma\gamma'} \end{pmatrix} \right\|_{\mathrm{op}} }_{A_1} + \underbrace{ \left\| \begin{pmatrix} 0 & H_{\lambda\gamma'} - \mathbb{E}_0 \tilde{H}_{0, \lambda\gamma'} \\ H_{\gamma\lambda'} - \mathbb{E}_0\tilde{H}_{0, \gamma\lambda'} & 0 \end{pmatrix} \right\|_{\mathrm{op}} }_{A_2}. \end{align}\]
In this step, we aim to derive an upper bound for \(A_1\). We focus on bounding the operator norm of \(\tilde{H}_{\lambda\lambda'} - \mathbb{E}_0 \tilde{H}_{0, \lambda\lambda'}\), since the same argument applies to \(\tilde{H}_{\gamma\gamma'} - \mathbb{E}_0 \tilde{H}_{0, \gamma\gamma'}\). Since \(A_1\) has a block-diagonal structure and, for any \(i\in \mathcal{I}_{NT}^c\), the \(i\)-th \(R\times R\) diagonal block of \(\tilde{H}_{\lambda\lambda'} - \mathbb{E}_0 \tilde{H}_{0, \lambda\lambda'}\) is a zero matrix, it suffices to consider the \(i\)-diagonal \(R\times R\) block for \(i\in \mathcal{I}_{NT}\).
For any \(i\in \mathcal{I}_{NT}\), we have \[\begin{align} [\tilde{H}_{\lambda\lambda'} - \mathbb{E}_0 \tilde{H}_{0, \lambda\lambda'}]_{i} = & \sum_{t\in \mathcal{T}_{NT}}\left((-\ddot{\ell}_{it})\gamma_t \gamma_{t}' - \mathbb{E}_0 (-\ddot{\ell}^0_{it})\gamma_{0, t}^{G} \gamma_{0, t}^{G\prime} \right) \\ = & -\sum_{t\in \mathcal{T}_{NT}} \ddot{\ell}_{it}\gamma^G_{0, t}\Delta_{\gamma_t}' - \sum_{t\in \mathcal{T}_{NT}}\ddot{\ell}_{it}\Delta_{\gamma_t}\gamma_{t}' - \sum_{t\in \mathcal{T}_{NT}}(\ddot{\ell}_{it} - \ddot{\ell}^0_{it} )\gamma_{0, t}^G \gamma_{0, t}^{G\prime} - \sum_{t\in \mathcal{T}_{NT}}(\ddot{\ell}^0_{it} - \mathbb{E}(\ddot{\ell}^0_{it}) )\gamma^G_{0, t} \gamma_{0, t}^{G\prime}. \end{align}\] Thus, \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|[\tilde{H}_{\lambda\lambda'} - \mathbb{E}_0 \tilde{H}_{0, \lambda\lambda'}]_{i}\|_{\mathrm{op}} \leq & \underbrace{\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\|\sum_{t\in \mathcal{T}_{NT}} \ddot{\ell}_{it}\gamma_{0, t}^G\Delta_{\gamma_t}' \right\|_{\mathrm{op}}}_{Q1} + \underbrace{\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\|\sum_{t\in \mathcal{T}_{NT}} \ddot{\ell}_{it}\Delta_{\gamma_t} \gamma_{ t}'\right\|_{\mathrm{op}}}_{Q2} \\ & \underbrace{\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\| \sum_{t\in \mathcal{T}_{NT}}(\ddot{\ell}_{it} - \ddot{\ell}^0_{it} )\gamma_{0, t}^G \gamma_{0, t}^{G\prime} \right\|_{\mathrm{op}}}_{Q3} \\ & + \underbrace{\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\| \sum_{t\in \mathcal{T}_{NT}}(\ddot{\ell}^0_{it} - \mathbb{E}(\ddot{\ell}^0_{it}) )\gamma_{0, t}^G \gamma_{0, t}^{G\prime} \right\|_{\mathrm{op}}}_{Q4}. \end{align}\]
The upper bound of \(Q_1\) is derived as follows: \[\label{eq:bound951} \begin{align} Q_1 & \leq \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\|\sum_{t\in \mathcal{T}_{NT}} \ddot{\ell}_{it}\gamma^G_{0, t}\Delta_{\gamma_t}' \right\|_{\mathrm{F}} \\ & \stackrel{\text{(i)}}{\leq} \sqrt{|\mathcal{T}_{NT}|} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left(\sum_{t\in \mathcal{T}_{NT}} \left\|\ddot{\ell}_{it}\gamma^G_{0, t}\Delta_{\gamma_{t}}'\right\|_{\mathrm{F}}^2\right)^{1/2} \\ & \leq \sqrt{T} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left(\sum_{t=1}^{T} \left\|\ddot{\ell}_{it}\gamma^G_{0, t}\Delta_{\gamma_{t}}'\right\|_{\mathrm{F}}^2\right)^{1/2} \\ & \stackrel{\text{(ii)}}{\lesssim} \sqrt{T} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|\Gamma - \Gamma_0^G\|_{\mathrm{F}} \\ & \stackrel{\text{(iii)}}{\lesssim} T \delta_{NT}, \end{align}\tag{58}\] where inequality (i) follows from the Cauchy-Schwarz inequality, inequality (ii) relies on the assumption that \({\ell}_{it}, \gamma_{0,t}\) are uniformly bounded, and inequality (iii) follows directly from the definition of \(\mathcal{B}_{\delta_{NT}}\). The upper bound of \(Q_2\) is derived via a similar method: \[\label{eq:bound952} \begin{align} Q_2 & \leq \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\|\sum_{t\in \mathcal{T}_{NT}} \ddot{\ell}_{it}\Delta_{\gamma_t}\gamma_{t}' \right\|_{\mathrm{F}} \\ & \leq \sqrt{|\mathcal{T}_{NT}|} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left(\sum_{t\in \mathcal{T}_{NT}} \left\|\ddot{\ell}_{it}\Delta_{\gamma_{t}}\gamma_{ t}'\right\|_{\mathrm{F}}^2\right)^{1/2} \\ & \leq \sqrt{T} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left(\sum_{t=1}^{T} \left\|\ddot{\ell}_{it}\Delta_{\gamma_{t}}\gamma_{ t}'\right\|_{\mathrm{F}}^2\right)^{1/2} \\ & \stackrel{\text{(i)}}{\lesssim} \sqrt{T} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|\Gamma - \Gamma_0^G \|_{\mathrm{F}} \\ & \lesssim T \delta_{NT}, \end{align}\tag{59}\] where inequality (i) uses the assumption that \({\ell}_{it}, \gamma_{t}\) are uniformly bounded. We drive the bound of \(Q_3\) by \[\label{eq:bound953} \begin{align} Q_3 \leq & \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left\|\sum_{t\in \mathcal{T}_{NT}}(\ddot{\ell}_{it} - \ddot{\ell}^0_{it} ) \gamma_{0, t}^G \gamma_{0, t}^{G\prime} \right\|_{\mathrm{F}} \\ \stackrel{\text{(i)}}{\lesssim} & \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \left|\sum_{t\in \mathcal{T}_{NT}}\tilde{\Delta}_{Y^*_{it}} \right| \\ \lesssim &\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{t\in \mathcal{T}_{NT}} |X_{it}'\Delta_{\beta} | + \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{t\in \mathcal{T}_{NT}} \|\lambda_i\| \|\gamma_t -\gamma^G_{0, t} \| + \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{t\in \mathcal{T}_{NT}} \|\gamma_{0, t}\| \|\lambda_{i} - \lambda^G_{0, i} \| \\ \lesssim &\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{t=1}^{T} |X_{it}'\Delta_{\beta} | + \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{t=1}^{T} \|\lambda_i\| \|\gamma_t -\gamma^G_{0, t} \| + \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{t= 1}^{T} \|\gamma_{0, t}\| \|\lambda_{i} - \lambda^G_{0, i} \| \\ \lesssim & T \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|\Delta_{\beta}\|_2 + \sqrt{T} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|\Gamma - \Gamma_0^G \|_{\mathrm{F}} + T \|\lambda_i - \lambda^G_{0, i}\| \\ \lesssim & T\delta_{NT} + T \delta_{NT} + \frac{\sqrt{T}}{\delta_{NT}}\\ = & o_p(\min\{N, T\}). \end{align}\tag{60}\]
Therefore, combining 58 , 59 , and 60 , we obtain \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|[\tilde{H}_{\lambda\lambda'} - \mathbb{E}_0\tilde{H}_{0, \lambda\lambda'}]_{i}\|_{\mathrm{op}} = o_p(\min\{N, T\}). \end{align}\] By the same method, we have \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|[\tilde{H}_{\gamma\gamma'} - \mathbb{E}_0\tilde{H}_{0, \gamma\gamma'}]_{t}\|_{\mathrm{op}} = o_p(\min\{N, T\}). \end{align}\] Consequently, \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|A_1\|_{\mathrm{op}} = o_p(\min\{N, T\}). \end{align}\]
In this step, we bound the influence of \(A_2\). We focus on \(H_{\lambda\gamma'} - \mathbb{E}_0 \tilde{H}_{0, \lambda\gamma'}\) and consider the expansion of its \((i, t)\)-th block, \(\ddot{\ell}_{it}\lambda_i\gamma'_t - \mathbb{E}_0 (\ddot{\ell}^0_{it})\lambda_{0,i}^G\gamma^{G\prime}_{0, t}\). When \(i\in \mathcal{I}_{NT}\) and \(t\in \mathcal{T}_{NT}\), \[\begin{align} \ddot{\ell}_{it}\lambda_i\gamma'_t - \mathbb{E}_0(\ddot{\ell}^0_{it})\lambda_{0,i}^G\gamma^{G\prime}_{0, t} & = \underbrace{\ddot{\ell}_{it} (\lambda_i \gamma'_t - \lambda_{0,i}^G\gamma^{G\prime}_{0, t}) }_{H_{1, it}} + \underbrace{(\ddot{\ell}_{it} - \ddot{\ell}^0_{it})\lambda_{0,i}^G\gamma^{G\prime}_{0, t}}_{H_{2, it}} + \underbrace{(\ddot{\ell}^0_{it} - \mathbb{E}_{\phi_0}(\ddot{\ell}^0_{it}))\lambda_{0,i}^G\gamma^{G\prime}_{0, t} }_{H_{3, it}}. \end{align}\] For notational simplicity we collect the terms above into matrices \(H_1, H_2, H_3\in \mathbb{R}^{RN\times RT}\). Note that \[\begin{align} & \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|H_{\lambda\gamma'} - \mathbb{E}_0 \tilde{H}_{0, \lambda\gamma'}\|^2_{\mathrm{op}} \\ \leq & \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} (\|H_1\|^2_{\mathrm{F}} + \|H_2\|^2_{\mathrm{F}} + \|H_3\|^2_{\mathrm{op}}) + \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\left(\sum_{i=1}^{N}\sum_{t\in \mathcal{T}_{NT}} (\ddot{\ell}_{it}\lambda_i\gamma'_t)^2 + \sum_{i\in \mathcal{I}_{NT}}\sum_{t=1}^{T} (\ddot{\ell}_{it}\lambda_i\gamma'_t)^2\right). \end{align}\] It follows that, wpa1, \[\label{eq:bound954} \begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|H_1\|^2_{\mathrm{F}} \stackrel{\text{(i)}}{\lesssim} & \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\sum_{i\in \mathcal{I}_{NT}}\|\lambda_i\|^2\sum_{t\in \mathcal{T}_{NT}} \|\gamma_t - \gamma^G_{0, t}\|^2 \\ & + \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\sum_{t\in \mathcal{T}_{NT} }\|\gamma^G_{0, t}\|^2\sum_{i\in \mathcal{I}_{NT}} \|\lambda_i - \lambda^G_{0, i}\|^2 \\ \leq & \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\sum_{i=1}^{N}\|\lambda_i\|^2\sum_{t=1}^{T} \|\gamma_t - \gamma^G_{0, t}\|^2 \\ & + \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\sum_{t=1}^{T}\|\gamma_{0, t}^G\|^2\sum_{i=1}^{N} \|\lambda_i - \lambda^G_{0, i}\|^2 \\ \stackrel{\text{(ii)}}{\lesssim} & N \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|\Gamma - \Gamma_0^G\|_{\mathrm{F}}^2 + T \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|\Lambda - \Lambda_0^G\|_{\mathrm{F}}^2 \\ \lesssim & NT\delta_{NT}^2, \end{align}\tag{61}\] Here, inequality (i) follows from the Cauchy-Schwarz inequality, and (ii) follows directly from the definition of \(\mathcal{B}_{\delta_{NT}}\). Also, wpa1, \[\label{eq:bound955} \begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|H_2\|^2_{\mathrm{F}} \lesssim & \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\sum_{i\in \mathcal{I}_{NT}} \sum_{t\in \mathcal{T}_{NT}} \tilde{\Delta}^2_{Y^*_{it}} \\ \leq & \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\sum_{i=1}^{N}\sum_{t=1}^{T} \tilde{\Delta}^2_{Y^*_{it}} \\ \stackrel{\text{(i)}}{\lesssim} & NT\delta_{NT}^2, \end{align}\tag{62}\] where inequality (i) is from the proof of Lemma 8. In addition, based on the Lemma 15, we show that, wpa1, \[\begin{align} \label{eq:bound956} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|H_3\|^2_{\mathrm{op}} \lesssim \max\{N, T\}\log((NT))^2 \lesssim NT \delta_{NT}^2 \end{align}\tag{63}\] Finally, wpa1, \[\label{eq:bound957} \begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\left(\sum_{i=1}^{N}\sum_{t\in \mathcal{T}_{NT}^c} \|\ddot{\ell}_{it}\lambda_i\gamma'_t\|_{\mathrm{F}}^2 + \sum_{i\in \mathcal{I}_{NT}^c }\sum_{t=1}^{T} \|\ddot{\ell}_{it}\lambda_i\gamma'_t\|_{\mathrm{F}}^2\right) \lesssim & N|\mathcal{T}_{NT}^c| + T|\mathcal{I}_{NT}^c| \stackrel{\text{(i)}}{\lesssim} & \max\{N^3, T^3\}\delta_{NT}^4 \end{align}\tag{64}\] where inequality (i) follows from 39 . Therefore, combining equations 61 — 64 we conclude that \(\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|A_2\|_{\mathrm{op}}= o_p(\min\{N, T\})\). Consequently, \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|S_3\|_{\mathrm{op}} = o_p(\min\{N, T\}). \end{align}\] This completes the proof.
Lemma 11. Under the conditions of Theorem 6, the operator norm of \(F\) in 42 satisfies \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|F\|_{\mathrm{op}} = o_p(\min\{N, T\}). \end{align}\]
Proof of Lemma 11. Consider the expansion of \(\dot{\ell}_{it}\) (noting that \(\mathbb{E}_0\dot{\ell}_{it}^0 = 0\)): \[\begin{align} \dot{\ell}_{it} = (\dot{\ell}_{it} - \dot{\ell}^{0}_{it})+ (\dot{\ell}^0_{it} - \mathbb{E}\dot{\ell}_{it}^0). \end{align}\] The corresponding matrix decomposition of \(F\) is given by \[\begin{align} F = \underbrace{\begin{pmatrix} 0 & 0 & 0 \\ 0 & 0 & F_{1, \lambda \gamma'} \\ 0 & F_{1, \gamma\lambda'} & 0 \end{pmatrix}}_{F_1} + \underbrace{\begin{pmatrix} 0 & 0 & 0 \\ 0 & 0 & F_{2, \lambda \gamma'} \\ 0 & F_{2, \gamma\lambda'} & 0 \end{pmatrix}}_{F_2}, \end{align}\] where \[\begin{align} F_{1, \lambda \gamma'} & = \left[(\dot{\ell}_{it} - \dot{\ell}^{0}_{it})\mathbb{I}_R\right]_{i=1,2,\ldots, N, t = 1,2,\ldots, T}, \\ F_{2, \lambda \gamma'} & = \left[(\dot{\ell}^0_{it} - \mathbb{E}_0 \dot{\ell}_{it}^0)\mathbb{I}_R\right]_{i=1,2,\ldots, N, t = 1,2,\ldots, T}. \end{align}\] The operator norm of \(F\) is therefore bounded by \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|F\|_{\mathrm{op}} \leq \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|F_1\|_{\mathrm{op}} + \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|F_2\|_{\mathrm{op}}. \end{align}\] We conclude that \(\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|F_1\|_{\mathrm{op}} = o_p(\min\{N, T\})\) using the same argument as in the proof of Lemma 9 and \(\sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|F_2\|_{\mathrm{op}} = o_p(\max \{N, T\})\) by Lemma 15. Therefore, when \(N\sim T\), we have \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|F\|_{\mathrm{op}} = o_p(\min\{N, T\}). \end{align}\] This completes the proof.
Lemma 12. Under the conditions of Theorem 6, we have \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|V - V_0\|_{\mathrm{op}} = o_p(\max\{N, T\}). \end{align}\]
Proof of Lemma 12. Recall that \[\begin{align} V_{\lambda\lambda'} & = \frac{T}{N}\left[ \lambda_i \lambda_{i'}' \right]_{i, i' = 1,2,\ldots, N }, \\ V_{\lambda\gamma'} & = \left[ -\lambda_i \gamma_{t}' \right]_{i =1,2,\ldots, N, t = 1,2,\ldots, T}, \\ V_{\gamma\gamma'} & = \frac{N}{T}\left[ \gamma_t \gamma_{t'}'\right]_{t, t' = 1,2,\ldots, T} . \end{align}\] In addition, \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|V - V_0\|_{\mathrm{op}} \leq & \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|V - V_0\|_{\mathrm{F}} \\ \leq & \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|V_{\lambda\lambda'} - V_{0, \lambda\lambda'} \|_{\mathrm{F}} + \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|V_{\gamma\gamma'} - V_{0, \gamma\gamma'} \|_{\mathrm{F}} \\ & + 2 \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|V_{\lambda\gamma'} - V_{0, \lambda\gamma'} \|_{\mathrm{F}}. \end{align}\] We first show that, wpa1, \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|V_{\lambda\lambda'} - V_{0, \lambda\lambda'} \|_{\mathrm{F}}^2 \leq & \frac{T^2}{N^2 } \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{i=1}^{N} \sum_{i'=1}^{N} \left\|\lambda_i \lambda_{i'}' - \lambda_{i ,0}^{G} \lambda_{i', 0}^{G\prime} \right\|_{\mathrm{F}}^2 \\ \stackrel{\text{(i)}}{\lesssim} & \frac{T^2}{N^2 } \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \sum_{i=1}^{N} \sum_{i'=1}^{N} \left(\|\lambda_{i'}\|^2 \|\lambda_{i} - \lambda^G_{0, i}\|^2 + \| \lambda^G_{0, i}\|^2 \|\lambda_{i'} - \lambda^G_{0, i'}\|^2 \right) \\ \stackrel{\text{(ii)}}{\lesssim} & \frac{T^2}{N } \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}}\|\Lambda - \Lambda_0^G\|_{\mathrm{F}}^2 \\ \lesssim & T^2 \delta_{NT}^2, \end{align}\] where inequality (i) employs the Cauchy-Schwarz inequality, and inequality (ii) uses the uniform boundedness of \(\Lambda\) and \(\Lambda_0^{G}\). Similarly, we have, wpa1, \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|V_{\gamma\gamma'} - V_{0, \gamma\gamma'} \|_{\mathrm{F}}^2 \lesssim N^2 \delta_{NT}^2. \end{align}\] In addition, wpa1, \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|V_{\lambda\gamma'} - V_{0, \lambda\gamma'} \|_{\mathrm{F}}^2 \stackrel{\text{(i)}}{\lesssim} & \sum_{i=1}^{N}\sum_{t=1}^{T} \left(\|\lambda_i\|^2 \|\gamma_{t} - \gamma^G_{0, t}\|^2 + \| \gamma^G_{0, t}\|^2 \|\lambda_{i} - \lambda^G_{0, i}\|^2\right) \\ \stackrel{\text{(ii)}}{\lesssim} & N \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|\Gamma - \Gamma_0^G\|_{\mathrm{F}}^2 + T \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|\Lambda - \Lambda_0^G\|_{\mathrm{F}}^2 \\ \lesssim & NT\delta^2_{NT}, \end{align}\] where inequality (i) again employs the Cauchy-Schwarz inequality, and inequality (ii) uses the uniform boundedness of \(\Gamma\) and \(\Gamma_0^{G}\). Finally, \[\begin{align} \sup_{(\beta, \Lambda, \Gamma) \in \mathcal{B}_{\delta_{NT}}} \|V - V_0\|_{\mathrm{op}} \lesssim \max\{N, T\}\delta_{NT} = o_p(\max\{N, T\}). \end{align}\] This completes the proof.
Proof of Lemma 3. Let \(\mathbb{P}_{\mathcal{U}}(\cdot): = \mathbb{P}\left(\cdot \mid \mathcal{U}\right)\) denote the probability conditional on \(\mathcal{U}\) defined in Assumption 7, and let \(\mathbb{E}_{\mathcal{U}}(\cdot): = \mathbb{E}\left(\cdot \mid \mathcal{U}\right)\) denote the expectation conditional on \(\mathcal{U}\). By 40 and 41 , we have \[\begin{align} NT \mathbb{E}_0 \mathcal{H}(\beta_0, \Lambda_0^G, \Gamma_0^G) = \mathbb{E}_{\mathcal{U}}\mathbb{E}_0 H_0 + (\mathbb{E}_0 H_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0 H_0 ) + V_0 + (\mathbb{E}_0\hat{V} - V_0). \end{align}\] In addition, by Lemma 12, we have \(\|\mathbb{E}_0\hat{V} - V_0\|_{\mathrm{op}}\lesssim \sup_{\beta, \Lambda, \Gamma\in \mathcal{B}_{\delta_{NT}}}\|V - V_0\|_{\mathrm{op}} = o_p(\max\{N, T\})\). Thus, it suffices to show that the minimum eigenvalue of \(\mathbb{E}_{\mathcal{U}}\mathbb{E} H_0 + V_0\) is strictly positive with order \(\min\{N, T\}\), and the operator norm of the perturbation \((\mathbb{E}_0 H_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0 H_0 )\) is sufficiently small relative to \(\mathbb{E}_{\mathcal{U}}\mathbb{E} H_0 + V_0\).
In this step, we show that minimum eigenvalue of \(\mathbb{E}_{\mathcal{U}}\mathbb{E}_0 H_0 + V_0\) is strictly positive and of order \(\min\{N, T\}\). For any \(i, t\), let \(\nabla^2 \ell_{it}^0\) be the \((d_X + R(N+T)) \times (d_X + R(N+T))\) Hessian matrix with respect to \((\beta, \Lambda, \Gamma)\). The matrix \(\nabla^2(-\ell_{it}^0)\) can be interpreted as the sample Hessian based on a single observation \((Y_{it}, X_{it})\). For notational simplicity, define \(L^0_{it}: = \mathbb{E}_{\mathcal{U}}\mathbb{E}_0(\ell_{it}^0)\). It is immediate that both \(\nabla^2 (-\ell_{it}^0)\) and \(\nabla^2(-L^0_{it})\) are positive semi-definite. We now consider the following rescaled version of \(\nabla^2(-L_{it}^0)\): \[\begin{align} \widecheck{ \nabla^2(-L^0_{it})} = \begin{pmatrix} \sqrt{\nu}\nabla_{\beta\beta'}^2(-L^0_{it}) & \nabla_{\beta\lambda'}^2(-L^0_{it}) & \nabla_{\beta\gamma'}^2(-L^0_{it}) \\ \nabla_{\lambda\beta'}^2(-L^0_{it}) & \sqrt{\nu}\nabla_{\lambda\lambda'}^2(-L^0_{it}) & \sqrt{\nu}\nabla_{\lambda\gamma'}^2(-L^0_{it}) \\ \nabla_{\gamma\beta'}^2 (-L^0_{it}) & \sqrt{\nu}\nabla_{\gamma\lambda'}^2(-L^0_{it}) & \sqrt{\nu} \nabla_{\gamma\gamma'}^2(-L^0_{it}) \end{pmatrix}, \end{align}\] where \(\nu\) is defined as in Assumption 7[item:conditional95variability95hessian]. Since \(\sqrt{\nu}\mathbb{E}_{\mathcal{U}}(\mathbb{E}_0(\ddot{\ell}^0_{it} X_{it}X_{it}'))\) is positive semi-definite, it follows that \(\sqrt{\nu}\nabla_{\beta\beta'}^2(-L^0_{it})\) is positive semi-definite as well. Thus, by Schur’s lemma, \(\widecheck{ \nabla^2(-L^0_{it})}\) is positive semi-definte if its Schur’s complement \(\widecheck{ \nabla^2(-L^0_{it})}\backslash \sqrt{\nu}\nabla_{\beta\beta'}^2(-L^0_{it})\) is positive semi-definite. The latter statement is true because \[\begin{align} & \begin{pmatrix} \mathbb{E}_{\mathcal{U}}(\mathbb{E}_0(-\ddot{\ell}_{it}^0)) \gamma_{t,0} \gamma_{t,0}' & \mathbb{E}_{\mathcal{U}}(\mathbb{E}_0(-\ddot{\ell}_{it}^0)) \gamma_{t,0} \lambda_{i, 0}' \\ \mathbb{E}_{\mathcal{U}}(\mathbb{E}_0(-\ddot{\ell}_{it}^0)) \lambda_{i, 0} \gamma_{t,0}' & \mathbb{E}_{\mathcal{U}}(\mathbb{E}_0(-\ddot{\ell}_{it}^0)) \lambda_{i, 0} \lambda_{i,0}' \end{pmatrix} \\ & - \mathbb{E}_{\mathcal{U}}(\mathbb{E}_0(-\ddot{\ell}_{it}^0X_{it}'))\left(\sqrt{\nu}\mathbb{E}_{\mathcal{U}}(\mathbb{E}_0(-\ddot{\ell}^0_{it} X_{it}X_{it}'))\right)^{-1}\mathbb{E}_{\mathcal{U}}(\mathbb{E}_0(-\ddot{\ell}_{it}^0X_{it})) \begin{pmatrix} \gamma_{0, t}\gamma_{0, t}' & \gamma_{0, t}\lambda_{0, i}' \\ \lambda_{0, i}\gamma_{0, t}' & \lambda_{0, i} \lambda_{0, i}' \end{pmatrix} \\ = & \left(\sqrt{\nu} \mathbb{E}_{\mathcal{U}}(\mathbb{E}_0(-\ddot{\ell}_{it}^0)) - \mathbb{E}_{\mathcal{U}}(\mathbb{E}_0(-\ddot{\ell}_{it}^0X_{it}'))\left(\sqrt{\nu}\mathbb{E}_{\mathcal{U}}(\mathbb{E}_0(-\ddot{\ell}^0_{it} X_{it}X_{it}'))\right)^{-1}\mathbb{E}_{\mathcal{U}}(\mathbb{E}_0(-\ddot{\ell}_{it}^0X_{it}))\right) \begin{pmatrix} \gamma_{0, t}\gamma_{0, t}' & \gamma_{0, t}\lambda_{0, i}' \\ \lambda_{0, i}\gamma_{0, t}' & \lambda_{0, i} \lambda_{0, i}' \end{pmatrix} \\ \geq & 0, \end{align}\] where the last inequality comes from Assumption 7[item:conditional95variability95hessian]. Thus, \[\begin{align} \mathbb{E}_{\mathcal{U}}\mathbb{E}_0H_0 = \sum_{i=1}^{N}\sum_{t=1}^{T} \nabla^2(-L^0_{it}) & \geq \sum_{i=1}^{N}\sum_{t=1}^{T} \nabla^2(-L^0_{it}) - \sum_{i=1}^{N}\sum_{t=1}^{T} \underbrace{\widecheck{ \nabla^2(-L^0_{it})}}_{\geq 0} \\ & = (1-\sqrt{\nu}) \mathbb{E}_{\mathcal{U}}\mathbb{E}_0 \begin{pmatrix} H_{0, \beta\beta'} & 0 & 0 \\ 0 & H_{0, \lambda\lambda'} & H_{0, \lambda\gamma'}\\ 0 & H_{0, \gamma\lambda'} & H_{0, \gamma\gamma'}\\ \end{pmatrix} \\ & = (1-\sqrt{\nu}) \mathbb{E}_{\mathcal{U}}\mathbb{E}_0 \begin{pmatrix} H_{0, \beta\beta'} & 0 & 0 \\ 0 & 0 & 0 \\ 0 & 0 & 0\\ \end{pmatrix} + (1-\sqrt{\nu}) \mathbb{E}_{\mathcal{U}}\mathbb{E}_0 \begin{pmatrix} 0 & 0 & 0 \\ 0 & H_{0, \lambda\lambda'} & H_{0, \lambda\gamma'}\\ 0 & H_{0, \gamma\lambda'} & H_{0, \gamma\gamma'}\\ \end{pmatrix}. \end{align}\] Hence, using the similar argument from [1], there exists a constant \(c>0\), independent of \(N, T\), such that \[\begin{align} (1-\sqrt{\nu}) \mathbb{E}_{\mathcal{U}}\mathbb{E}_0 \begin{pmatrix} 0 & 0 & 0 \\ 0 & H_{0, \lambda\lambda'} & H_{0, \lambda\gamma'}\\ 0 & H_{0, \gamma\lambda'} & H_{0, \gamma\gamma'}\\ \end{pmatrix} + V_0 & \geq c \begin{pmatrix} 0 & 0 & 0 \\ 0 & T \mathbb{I}_{R N}& 0 \\ 0 & 0& N \mathbb{I}_{R T} \end{pmatrix}, \end{align}\] and \[\begin{align} (1-\sqrt{\nu}) \mathbb{E}_{\mathcal{U}}\mathbb{E}_0 \begin{pmatrix} H_{0, \beta\beta'} & 0 & 0 \\ 0 & 0 & 0 \\ 0 & 0 & 0\\ \end{pmatrix} \geq c \begin{pmatrix} NT\mathbb{I}_{d_X} & 0 & 0 \\ 0 & 0 & 0 \\ 0 & 0 & 0\\ \end{pmatrix}. \end{align}\] Consequently, we must have \[\begin{align} \mathbb{E}_{\mathcal{U}}\mathbb{E}_0H_0 + V_0 \geq c \begin{pmatrix} NT\mathbb{I}_{d_X} & 0 & 0 \\ 0 & T \mathbb{I}_{R N}& 0 \\ 0 & 0& N \mathbb{I}_{R T} \end{pmatrix}, \end{align}\] whose minimum eigenvalue is positive with order \(\min\{N, T\}\).
In this step, we show that the perturbation \((\mathbb{E} H_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E} H_0 )\) is sufficiently small relative to \(\mathbb{E}_{\mathcal{U}}\mathbb{E}H_0 + bV_0\). Note that \[\begin{align} \mathbb{E}_0 H_0 + bV_0 = & \mathbb{E}_{\mathcal{U}}\mathbb{E}_0 H_0 + bV_0+ (E_0 H_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0 H_0) \\ \geq & c \begin{pmatrix} 0 & 0 & 0 \\ 0 & T \mathbb{I}_{R N}& 0 \\ 0 & 0& N \mathbb{I}_{R T} \end{pmatrix} + \frac{c}{2} \begin{pmatrix} NT\mathbb{I}_{d_X} & 0 & 0 \\ 0 & 0 & 0 \\ 0 & 0& 0 \end{pmatrix} + \frac{c}{2} \begin{pmatrix} NT\mathbb{I}_{d_X} & 0 & 0 \\ 0 & 0 & 0 \\ 0 & 0& 0 \end{pmatrix} \\ & + ( \mathbb{E}_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0) \begin{pmatrix} H_{0, \beta \beta'} & H_{0, \beta \lambda'} & H_{0, \beta \gamma'} \\ H_{0, \lambda \beta'} & 0 & 0 \\ H_{0, \gamma \beta'} & 0 & 0 \end{pmatrix} + (\mathbb{E}_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0 ) \begin{pmatrix} 0 & 0 & 0 \\ 0 & H_{0, \lambda \lambda'} & H_{0, \lambda \gamma'} \\ 0 & H_{0, \gamma \lambda'} & H_{0, \gamma \gamma'} \end{pmatrix} \\ = & \underbrace{ \begin{pmatrix} \frac{c}{2}NT\mathbb{I}_{d_X} + ( \mathbb{E}_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0)H_{0, \beta \beta'} & ( \mathbb{E}_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0)H_{0, \beta \lambda'} & ( \mathbb{E}_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0)H_{0, \beta \gamma'} \\ ( \mathbb{E}_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0)H_{0, \lambda\beta'} & 0& 0 \\ ( \mathbb{E}_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0)H_{0, \gamma \beta'} & 0& 0 \end{pmatrix} }_{A_1} \\ & + \underbrace{ ( \mathbb{E}_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0) \begin{pmatrix} 0 & 0 & 0 \\ 0 & H_{0, \lambda \lambda'} & 0 \\ 0 & 0 & H_{0, \gamma \gamma'} \end{pmatrix} }_{A_2} + \underbrace{ ( \mathbb{E}_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0) \begin{pmatrix} 0 & 0 & 0 \\ 0 & 0 & H_{0, \lambda \gamma'} \\ 0 & H_{0, \gamma \lambda'} & 0 \end{pmatrix} }_{A_3} \\ & + c \begin{pmatrix} \frac{NT}{2}\mathbb{I}_{d_X} & 0 & 0 \\ 0 & T \mathbb{I}_{R N}& 0 \\ 0 & 0& N \mathbb{I}_{R T} \end{pmatrix} . \end{align}\] We first study the eigenvalues of \(A_1\). For the upper-left block of \(A_1\), we establish a lower bound for its minimum eigenvalue as follows: \[\begin{align} \sigma_{\min}\left(\frac{1}{2}NT\mathbb{I}_{d_X} + ( \mathbb{E}_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0)H_{0, \beta \beta'}\right) \stackrel{\text{(i)}}{\gtrsim} NT + O_p(\sqrt{NT}) \gtrsim NT, \quad \boldsymbol{wpa1}, \end{align}\] where inequality (i) follows from the fact that \(( \mathbb{E}_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0) H_{0, \beta \beta'} = O_p(\sqrt{NT})\) by Assumption 7[item:conditional95weak95dependence95hessian] and [40]. By the same argument, each entry of \(( \mathbb{E}_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0) H_{\beta\lambda'}\) is of order \(O_p(\sqrt{T})\), and each entry of \(( \mathbb{E}_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0) H_{\beta\gamma'}\) is of order \(O_p(\sqrt{N})\). By applying the Gershgorin’s Circle Theorem, we conclude that every eigenvalue of \(A_1\) falls into one of two categories: (i) either it is positive and of order \(O_p(NT)\), or (ii) it is of order \(O_p(\sqrt{T}), O_p(\sqrt{N})\) (with an unspecified sign). Therefore, by the asymptotic assumption about \(N, T\), the minimum eigenvalue of \(A_1\) must be of order \(\min\{\sqrt{N}, \sqrt{T}\}\).
In addition, by Assumption 7[item:conditional95weak95dependence95hessian] and [40], each \(R\)-dimensional block of \((\mathbb{E}_0-\mathbb{E}_{\mathcal{U}}\mathbb{E}_0 ) H_{0, \lambda \lambda'}\) is of order \(O_p(\sqrt{T})\). By the same argument, each \(R\)-dimensional diagonal block of \((\mathbb{E}_0 - \mathbb{E}_{\mathcal{U}}\mathbb{E}_0) H_{0,\gamma\gamma'}\) is of order \(O_p(\sqrt{N})\). Since \(A_2\) has a diagonal block structure, we conclude that the minimum eigenvalue of \(A_2\) is of order \(o_p(\min\{N, T\})\) as well.
To bound the eigenvalues of \(A_3\), by Assumption 7 and Lemma 15, the operator norm of \(A_3\) is of order \(O_p(\sqrt{\max\{N, T\}}\log(NT))\), and hence the minimum eigenvalue of \(A_3\) must be of order \(o_p(\min\{N, T\})\). Combining the previous result, Weyl’s theorem, and the asymptotic assumption about \(N, T\), we conclude that the minimum eigenvalue of \(\mathbb{E}_0\mathcal{H}_{0}\) is \[\begin{align} \sigma_{\min}(\mathbb{E}_0\mathcal{H}_{0}) \gtrsim \min\{N, T\}. \end{align}\] Since \(\mathcal{H}_{0, NT} = \frac{1}{NT}\mathcal{H}_0\), it follows immediately that \(\sigma_{\min}(\mathbb{E}_0\mathcal{H}_{0, NT}) \gtrsim (\max\{N, T\})^{-1}\).
Proof of Theorem 3. Recall that \[\begin{align} (\hat{\beta}_{\mathrm{nuc}}, \hat{\Theta}_{\mathrm{nuc}}) \in \mathop{\mathrm{argmin}}_{\beta\in \mathbb{R}^{d_X}, \Theta\in \mathbb{R}^{N\times T}} \left\{\mathcal{L}_{NT}(\beta, \Theta) + \frac{\varphi_{NT}}{\sqrt{NT}}\|\Theta\|_{\mathrm{nuc}}\right\}. \end{align}\] Let \(h_{NT}(\Theta) = \frac{\varphi_{NT}}{\sqrt{NT}}\|\Theta\|_{\mathrm{nuc}}\). The proximal gradient operator is defined as \[\begin{align} \mathrm{prox}_{s_{\theta}h_{NT}}(\tilde{\Theta}) = \mathop{\mathrm{argmin}}_{\Theta } \left\{h_{NT}(\tilde{\Theta}) + \frac{1}{2}\|\tilde{\Theta} - \Theta\|_{\mathrm{F}}^2\right\}. \end{align}\] By Lemma 13, for any feasible \((\beta_1, \Theta_1), (\beta_2, \Theta_2)\) in the optimization 13 , there exist constants \(L_{\beta}, L_{\theta}\) such that \[\begin{align} \mathcal{L}_{NT}(\beta_2, \Theta_2) \leq & \mathcal{L}_{NT}(\beta_1, \Theta_1 ) + \nabla_{\beta} \mathcal{L}_{NT}(\beta_1, \Theta_1 )'(\beta_2 -\beta_1) + \langle\nabla_{\Theta} \mathcal{L}_{NT}(\beta_1, \Theta_1 ), (\Theta_2 -\Theta_1) \rangle\\ & + \frac{L_{\beta}}{2} \|\beta_2 -\beta_1 \|^2 + \frac{L_{\theta}}{2NT}\|\Theta_2 - \Theta_1 \|_{\mathrm{F}}^2. \end{align}\] We apply the standard proximal gradient argument with the quadratic upper bound established in Lemma 13. Define \(G_{s_{\theta}}(\beta,\Theta) = \frac{1}{s_{\theta}} (\Theta - \mathrm{prox}_{s_{\theta}h_{NT}}(\Theta - s_{\theta}\nabla_{\Theta}\mathcal{L}_{NT}(\beta, \Theta) ))\). The parameters update is given by \[\begin{align} \begin{pmatrix} \beta_2 \\ \Theta_2 \end{pmatrix} = \begin{pmatrix} \beta_1 -s_{\beta}\nabla_{\beta} \mathcal{L}_{NT}(\beta_1, \Theta_1) \\ \Theta_1 -s_{\theta}G_{s_{\theta}}(\beta_1,\Theta_1) \end{pmatrix}. \end{align}\] By Lemma 13, we obtain \[\begin{align} \mathcal{L}_{NT}(\beta_2, \Theta_2) \leq & \mathcal{L}_{NT}(\beta_1, \Theta_1) - s_{\beta}\| \nabla_{\beta} \mathcal{L}_{NT}(\beta_1, \Theta_1 )\|^2 - s_{\theta} \langle\nabla_{\Theta} \mathcal{L}_{NT}(\beta_1, \Theta_1 ), G_{s_{\theta}}(\beta_1,\Theta_1) \rangle\\ & + \frac{L_{\beta}s_{\beta}^2 }{2}\| \nabla_{\beta} \mathcal{L}_{NT}(\beta_1, \Theta_1 )\|^2 + \frac{L_{\theta}s_{\theta}^2 }{2NT} \|G_{s_{\theta}}(\beta_1,\Theta_1)\|_{\mathrm{F}}^2. \end{align}\] When \(s_{\beta} \leq \frac{1}{L_{\beta}}\) and \(s_{\theta} \leq \frac{NT}{L_{\theta} }\), \[\begin{align} \mathcal{L}_{NT}(\beta_2, \Theta_2) \leq & \mathcal{L}_{NT}(\beta_1, \Theta_1) - s_{\theta} \langle\nabla_{\Theta} \mathcal{L}_{NT}(\beta_1, \Theta_1 ), G_{s_{\theta}}(\beta_1,\Theta_1) \rangle\\ & - \frac{ s_{\beta} }{2}\| \nabla_{\beta} \mathcal{L}_{NT}(\beta_1, \Theta_1 )\|^2 + \frac{s_{\Theta} }{2} \|G_{s_{\theta}}(\beta_1,\Theta_1)\|_{\mathrm{F}}^2. \end{align}\] Since \(\mathcal{L}_{NT}(\cdot, \cdot)\) and \(h_{NT}(\cdot)\) are convex, we have \[\begin{align} h_{NT}(\Theta_2) & \leq h_{NT}(\Theta_0) - \langle G_{s_{\theta}}(\beta_1, \Theta_1) - \nabla_{\Theta}\mathcal{L}_{NT}(\beta_1, \Theta_1), \Theta_0 - \Theta_1 + s_{\theta}G_{s_{\theta}}(\beta_1, \Theta_1)\rangle\\ \mathcal{L}_{NT}(\beta_1, \Theta_1 ) & \leq \mathcal{L}_{NT}(\beta_0, \Theta_0) - \nabla_{\beta} \mathcal{L}_{NT}(\beta_1, \Theta_1 )'(\beta_0 -\beta_1) - \langle\nabla_{\Theta} \mathcal{L}_{NT}(\beta_1, \Theta_1 ), (\Theta_0 -\Theta_1) \rangle. \end{align}\] Adding the above inequalities yields \[\begin{align} \mathcal{L}_{NT}(\beta_2, \Theta_2) + h_{NT}(\Theta_2) \leq & \mathcal{L}_{NT}(\beta_0, \Theta_0) + h_{NT}(\Theta_0) \\ & + \nabla_{\beta}L_{NT}(\beta_1, \Theta_1)'(\beta_1 - \beta_0) + \langle G_{s_{\theta}}(\beta_1, \Theta_1), \Theta_1 - \Theta_0\rangle\\ & - \frac{ s_{\beta} }{2}\| \nabla_{\beta} \mathcal{L}_{NT}(\beta_1, \Theta_1 )\|^2 - \frac{s_{\theta} }{2} \|G_{s_{\theta}}(\beta_1,\Theta_1)\|_{\mathrm{F}}^2. \end{align}\] Let \(\mathcal{\psi}^{(k)} = \mathcal{L}_{NT}(\beta^{(k)}, \Theta^{(k)}) + h_{NT}(\Theta^{(k)})\) and \(\mathcal{\psi}_0 = \mathcal{L}_{NT}(\beta_0, \Theta_0) + h_{NT}(\Theta_0)\) for notational simplicity. Then \[\begin{align} \mathcal{\psi}^{(k+1)} - \mathcal{\psi}_0 \leq \frac{1}{2s_{\beta}}\left(\|\beta^{(k)} - \beta_0\|^2 - \|\beta^{(k+1)} - \beta_0 \|^2\right) + \frac{1}{2s_{\theta}} \left(\|\Theta^{(k)} - \Theta_0\|^2_{\mathrm{F}} - \|\Theta^{(k+1)} - \Theta_0 \|^2_{\mathrm{F}} \right). \end{align}\] Summing over \(k\) gives \[\begin{align} \mathcal{\psi}^{(k)} - \mathcal{\psi}_0 \leq \frac{1}{k}\sum_{i=0}^{k}(\mathcal{\psi}^{(i)} - \mathcal{\psi}_0 ) \leq \frac{1}{2 k s_{\beta}}\|\beta^{(0)} - \beta_0\|^2 + \frac{1}{2ks_{\theta}}\|\Theta^{(0)} - \Theta_0\|^2_{\mathrm{F}}. \end{align}\] Hence, when \(s_{\beta} \leq \frac{1}{L_{\beta}}\) and \(s_{\theta} \leq \frac{NT}{L_{\theta} }\), the proximal gradient method converges to the global minimizer with rate \(O(1/k)\).
Lemma 13 (Quadratic upper bound). For any feasible \((\beta_1, \Theta_1), (\beta_2, \Theta_2)\) in the optimization 13 , we have \[\begin{align} \mathcal{L}_{NT}(\beta_2, \Theta_2) \leq & \mathcal{L}_{NT}(\beta_1, \Theta_1 ) + \nabla_{\beta} \mathcal{L}_{NT}(\beta_1, \Theta_1 )'(\beta_2 -\beta_1) + \langle\nabla_{\Theta} \mathcal{L}_{NT}(\beta_1, \Theta_1 ), (\Theta_2 -\Theta_1) \rangle\\ & + \frac{L_{\beta}}{2} \|\beta_2 -\beta_1 \|^2 + \frac{L_{\theta}}{2NT}\|\Theta_2 - \Theta_1 \|_{\mathrm{F}}^2, \end{align}\] with \(L_{\beta} = 2 d_X b_{\max} \rho_X^2\) and \(L_{\theta} = 2b_{\max}\).
Proof of Lemma 13. Note that \[\begin{align} \mathcal{L}_{NT}(\beta_2, \Theta_2) = & \mathcal{L}_{NT}(\beta_1, \Theta_1 ) + \nabla_{\beta} \mathcal{L}_{NT}(\beta_1, \Theta_1 )'(\beta_2 -\beta_1) + \langle\nabla_{\Theta} \mathcal{L}_{NT}(\beta_1, \Theta_1 ), (\Theta_2 -\Theta_1)\rangle\\ & + \frac{1}{2} (\beta_2' -\beta'_1, \mathrm{vec}(\Theta_2)' - \mathrm{vec}(\Theta_2)')\nabla^2 \mathcal{L}_{NT}(\tilde{\beta}, \tilde{\Theta})(\beta_2' -\beta'_1, \mathrm{vec}(\Theta_2)' - \mathrm{vec}(\Theta_2)')', \end{align}\] where \[\begin{align} \nabla^2 \mathcal{L}_{NT}(\beta, \Theta ) = \begin{pmatrix} H_{\beta\beta'} & H_{\beta\theta'} \\ H_{\theta\beta'} & H_{\theta\theta'} \end{pmatrix} \geq 0, \end{align}\] and \[\begin{align} H_{\beta\beta'} & = -\frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T} X_{it}\ddot{\ell}_{it} X'_{it}, \\ H_{\beta\theta'} & = -\frac{1}{NT}\left[X_{it}\ddot{\ell}_{it}\right]_{i=1,\ldots, N, t = 1, \ldots, T}, \\ H_{\theta\theta'} & =-\frac{1}{NT}\mathrm{diag}\left\{ \ddot{\ell}_{it}\right\}_{i=1,\ldots, N, t = 1, \ldots, T}. \end{align}\] It is straightforward to verify that \[\begin{align} \begin{pmatrix} 2 H_{\beta\beta'} & 0 \\ 0 & 2 H_{\theta\theta'} \end{pmatrix} - \begin{pmatrix} H_{\beta\beta'} & H_{\beta\theta'} \\ H_{\theta\beta'} & H_{\theta\theta'} \end{pmatrix} = \begin{pmatrix} H_{\beta\beta'} & -H_{\beta\theta'} \\ -H_{\theta\beta'} & H_{\theta\theta'} \end{pmatrix} \geq 0. \end{align}\] The last inequality holds because (i) \(\nabla^2 \mathcal{L}_{NT}(\beta, \Theta )\geq 0\) and (ii) flipping the sign of off-diagonal part of the matrix does not change the eigenvalues of the matrix. Therefore, \[\begin{align} \begin{pmatrix} \beta_2 - \beta_1 \\ \mathrm{vec}(\Theta_2) - \mathrm{vec}(\Theta_1) \end{pmatrix}' \nabla^2 \mathcal{L}_{NT}(\tilde{\beta}, \tilde{\Theta} ) \begin{pmatrix} \beta_2 - \beta_1 \\ \mathrm{vec}(\Theta_2) - \mathrm{vec}(\Theta_1) \end{pmatrix} \leq 2 \sigma_{\max}(H_{\beta\beta'})\|\beta_2 - \beta_1\|^2 + 2\sigma_{\max}(H_{\theta\theta'})\|\Theta_2 - \Theta_1\|_{\mathrm{F}}^2. \end{align}\] Since \(\sigma_{\max}(H_{\beta\beta'}) \leq d_X b_{\max} \rho_X^2\) and \(\sigma_{\max}(H_{\theta\theta'}) \leq \frac{b_{\max}}{NT}\), the result follows.
Proof of Theorem 4. The proof consists of two steps. In the first step, we show that if we start from \((\beta^{(0)}, \Lambda^{(0)}, \Gamma^{(0)}) = (\hat{\beta}_{\mathrm{nuc}}, \hat{\Lambda}_{\mathrm{nuc}} ,\hat{\Gamma}_{\mathrm{nuc}}) \in \mathcal{B}_{\delta_{NT}}\), then, under properly chosen step sizes \((s_{\beta}, s_{\lambda}, s_{\gamma})\), the updated estimators \((\beta^{(1)}, \Lambda^{(1)}, \Gamma^{(1)})\) also remain in the neighborhood \(\mathcal{B}_{\delta_{NT}}\). In the second step, by establishing a local quadratic upper bound for the optimization problem 14 , we show that the sequence of updated estimators will always remain within the neighborhood \(\mathcal{B}_{\delta_{NT}}\). Additionally, after each iteration, the estimators become closer to the global minimizer than in the previous step.
First, note that when evaluated at \((\Lambda^{(0)},\Gamma^{(0)})\), we have \[\begin{align} \nabla_{\lambda} \|\hat{\Lambda}'_{\mathrm{nuc}}\Lambda /N - \Gamma' \hat{\Gamma}_{\mathrm{nuc}}/ T \|_{\mathrm{F}}^2 = 0, \quad \nabla_{\gamma}\|\hat{\Lambda}'_{\mathrm{nuc}}\Lambda / N - \Gamma' \hat{\Gamma}_{\mathrm{nuc}}/ T\|_{\mathrm{F}}^2 = 0. \end{align}\] Therefore, \[\begin{align} \beta^{(1)} = & \beta^{(0)} - s_{\beta} \nabla_{\beta}\mathcal{L}_{NT}(\beta^{(0)}, \Lambda^{(0)}, \Gamma^{(0)}), \\ \Lambda^{(1)} = & \Lambda^{(0)} - s_{\lambda} \nabla_{\lambda}\mathcal{L}_{NT}(\beta^{(0)}, \Lambda^{(0)}, \Gamma^{(0)}), \\ \Gamma^{(1)} = & \Gamma^{(0)} - s_{\gamma} \nabla_{\gamma}\mathcal{L}_{NT}(\beta^{(0)}, \Lambda^{(0)}, \Gamma^{(0)}). \end{align}\] We have the following inequality when \(s_{\beta}\lesssim 1\): \[\begin{align} \|\beta^{(1)} - \beta^{(0)}\| = & \frac{s_{\beta}}{NT}\left\| \sum_{i=1}^{N}\sum_{t=1}^{T}\dot{\ell}_{it}(\beta^{(0)\prime } X_{it} + \lambda^{(0)\prime}_i \gamma^{(0)}_t)X_{it}\right\| \\ \stackrel{\text{(i)}}{\leq} & \frac{s_{\beta}}{NT} \left\| \sum_{i=1}^{N}\sum_{t=1}^{T} \dot{\ell}^0_{it} X_{it} \right\| + \frac{s_{\beta}}{NT} \left\| \sum_{i=1}^{N}\sum_{t=1}^{T}\tilde{\ddot{\ell}}_{it} (\beta^{(0)} - \beta_0 )'X_{it}X_{it}\right\| \\ & + \frac{s_{\beta}}{NT} \left\| \sum_{i=1}^{N}\sum_{t=1}^{T}\tilde{\ddot{\ell}}_{it} \lambda^{(0)\prime}_i (\gamma_{0, t}^G - \gamma^{(0)}_t )X_{it} \right\| + \frac{s_{\beta}}{NT} \left\| \sum_{i=1}^{N}\sum_{t=1}^{T}\tilde{\ddot{\ell}}_{it} \gamma_{t}^{(0)\prime} (\lambda_{0, i}^G - \lambda^{(0)}_i )X_{it} \right\|, \end{align}\] where inequality (i) follows from a Taylor expansion and the triangle inequality. Here, \(\tilde{\ddot{\ell}}_{it}\) is an abbreviation for \(\ddot{\ell}_{it}(\tilde{\beta}'X_{it} + \tilde{\lambda}_i \tilde{\gamma}_t)\), where \((\tilde{\beta}, \tilde{\Lambda}, \tilde{\Gamma})\) lies on the line segment between \((\beta^{(0)}, \Lambda^{(0)}, \Gamma^{(0)})\) and \((\beta_0, \Lambda_0, \Gamma_0)\). Each of the four terms above can be bounded using the same argument as in the proof of Theorem 6.
Specifically, combining moment condition \(\mathbb{E}_{Z, \Lambda_0, \Gamma_0} (\dot{\ell}_{it}^{0}X_{it}) = 0\), sampling assumption (Assumption 4[item:sampling95pre]), and [40] yields \[\begin{align} \frac{s_{\beta}}{NT} \left\| \sum_{i=1}^{N}\sum_{t=1}^{T} \dot{\ell}^0_{it} X_{it} \right\| = O_p\left(\frac{\log(NT)}{\sqrt{NT}}\right). \end{align}\] In addition, since \(\{X_{it}\}_{1\leq i\leq N, 1\leq t\leq T}\), \((\tilde{\beta}, \tilde{\Lambda}, \tilde{\Gamma})\), \((\beta^{(0)}, \Lambda^{(0)}, \Gamma^{(0)})\), and \((\beta_0, \Lambda_0, \Gamma_0)\) are uniformly bounded, we obtain \[\begin{align} \frac{s_{\beta}}{NT} \left\| \sum_{i=1}^{N}\sum_{t=1}^{T}\tilde{\ddot{\ell}}_{it} (\beta^{(0)} - \beta_0 )'X_{it}X_{it}\right\| \lesssim \|\beta^{(0)} - \beta_0\| & \stackrel{\text{(i)}}{=} o_p(\delta_{NT}) , \\ \frac{s_{\beta}}{NT} \left\| \sum_{i=1}^{N}\sum_{t=1}^{T}\tilde{\ddot{\ell}}_{it} \lambda^{(0)\prime}_i (\gamma_{0, t}^G - \gamma^{(0)}_t )X_{it} \right\| \lesssim \frac{1}{\sqrt{T}} \|\Gamma^{(0)} - \Gamma_0^G\|_{\mathrm{F}} & \stackrel{\text{(ii)}}{=} o_p(\delta_{NT}), \\ \frac{s_{\beta}}{NT} \left\| \sum_{i=1}^{N}\sum_{t=1}^{T}\tilde{\ddot{\ell}}_{it} \gamma_{t}^{(0)\prime} (\lambda_{0, i}^G - \lambda^{(0)}_i )X_{it} \right\| \lesssim \frac{1}{\sqrt{N}} \|\Lambda^{(0) }- \Lambda_0^G\|_{\mathrm{F}} & \stackrel{\text{(iii)}}{=} o_p(\delta_{NT}). \end{align}\] Here, inequalities (i), (ii), and (iii) are obtained by the similar argument as in the proof of Theorem 5. Therefore, we conclude that \[\label{eq:algorithm95beta95bound} \|\beta^{(1)} - \beta^{(0)}\| = o_p(\delta_{NT}).\tag{65}\]
In addition, when \(s_{\lambda}\sim N\), \[\begin{align} \|\Lambda^{(1)} - \Lambda^{(0)}\|_{\mathrm{F}}^2 \lesssim & \frac{1}{T^2}\sum_{i=1}^{N} \left\| \sum_{t=1}^{T}\dot{\ell}_{it}(\beta^{(0) \prime} X_{it} + \lambda^{(0)\prime}_i \gamma^{(0)}_t)\gamma^{(0)}_t\right\|^2 \\ \stackrel{\text{(i)}}{\leq} & \frac{1}{T^2} \sum_{i=1}^{N} \left\| \sum_{t=1}^{T} \dot{\ell}^0_{it} \gamma_{0, t}\right\|^2 + \frac{1}{T^2} \sum_{i=1}^{N} \left\| \sum_{t=1}^{T} \dot{\ell}^0_{it} (\gamma_{0, t} - \gamma^{(0)}_{t})\right\|^2 \\ & + \frac{1}{T^2} \sum_{i=1}^{N}\left\| \sum_{t=1}^{T}\tilde{\ddot{\ell}}_{it} (\beta^{(0)} - \beta_0 )'X_{it} \gamma^{(0)}_t \right\|^2 \\ & + \frac{1}{T^2} \sum_{i=1}^{N}\left\| \sum_{t=1}^{T}\tilde{\ddot{\ell}}_{it} \lambda^{(0)\prime}_i (\gamma_{0, t}^G - \gamma^{(0)}_t ) \gamma^{(0)}_t\right\|^2 + \frac{1}{T^2} \sum_{i=1}^{N} \left\| \sum_{t=1}^{T}\tilde{\ddot{\ell}}_{it} \gamma_{t}^{(0)\prime} (\lambda_{0, i}^G - \lambda^{(0)}_i )\gamma^{(0)}_t \right\|^2, \end{align}\] where inequality (i) follows from the Taylor expansion and the triangle inequality. Similarly, we establish the following bounds: \[\begin{align} \frac{1}{T^2} \sum_{i=1}^{N} \left\| \sum_{t=1}^{T} \dot{\ell}^0_{it} \gamma_{0, t}\right\|^2 \stackrel{\text{(i)}}{=} O_p\left(\frac{N\log(T)^2}{T}\right) & = o_p(N\delta_{NT}^2 ), \\ \frac{1}{T^2} \sum_{i=1}^{N} \left\| \sum_{t=1}^{T} \dot{\ell}^0_{it} (\gamma_{0, t} - \gamma^{(0)}_{t})\right\|^2 \lesssim \frac{N}{T} \|\Gamma^{(0)} - \Gamma_0^G\|_{\mathrm{F}}^2 & \stackrel{\text{(ii)}}{=} o_p(N\delta_{NT}^2 ), \\ \frac{1}{T^2} \sum_{i=1}^{N}\left\| \sum_{t=1}^{T}\tilde{\ddot{\ell}}_{it} (\beta^{(0)} - \beta_0 )'X_{it} \gamma^{(0)}_t \right\|^2 \lesssim N \|\beta^{(0)} - \beta_0 \|_{\mathrm{F}}^2 & \stackrel{\text{(iii)}}{=} o_p(N\delta_{NT}^2 ), \\ \frac{1}{T^2} \sum_{i=1}^{N}\left\| \sum_{t=1}^{T}\tilde{\ddot{\ell}}_{it} \lambda^{(0)\prime}_i (\gamma_{0, t}^G - \gamma^{(0)}_t ) \gamma^{(0)}_t\right\|^2 \lesssim \frac{N}{T} \|\Gamma^{(0)} - \Gamma_0^G\|_{\mathrm{F}}^2 & \stackrel{\text{(iv)}}{=} o_p(N\delta_{NT}^2 ), \\ \frac{1}{T^2} \sum_{i=1}^{N} \left\| \sum_{t=1}^{T}\tilde{\ddot{\ell}}_{it} \gamma_{t}^{(0)\prime} (\lambda_{0, i}^G - \lambda^{(0)}_i )\gamma^{(0)}_t \right\|^2 \lesssim \|\Lambda^{(0)} - \Lambda_0^G\|_{\mathrm{F}}^2 & \stackrel{\text{(v)}}{=} o_p(N\delta_{NT}^2 ), \end{align}\] where inequality (i) comes from the sampling assumption (Assumption 4[item:sampling95pre]) and [40]. Inequalities (ii)—(v) are obtained by the same argument as in the proof of Theorem 5. Therefore, \[\label{eq:algorithm95lambda95bound} \frac{1}{\sqrt{N}}\|\Lambda^{(1)} - \Lambda^{(0)}\| = o_p(\delta_{NT}).\tag{66}\] Also, by a similar argument, we have \[\label{eq:algorithm95gamma95bound} \frac{1}{\sqrt{T}}\|\Gamma^{(1)} - \Gamma^{(0)}\| = o_p(\delta_{NT}).\tag{67}\] Combining equations 65 , 66 , and 67 , we show that if the optimization starts from \((\beta^{(0)}, \Lambda^{(0)}, \Gamma^{(0)}) = (\hat{\beta}_{\mathrm{nuc}}, \hat{\Lambda}_{\mathrm{nuc}} ,\hat{\Gamma}_{\mathrm{nuc}}) \in \mathcal{B}_{\delta_{NT}}\), then, under properly chosen step sizes \((s_{\beta}, s_{\lambda}, s_{\gamma})\) as in Theorem 4, the updated estimators \((\beta^{(1)}, \Lambda^{(1)}, \Gamma^{(1)})\) also remain in the neighborhood \(\mathcal{B}_{\delta_{NT}}\).
By Theorem 6, the optimization problem is strongly convex. Consequently, with properly chosen \((s_{\beta}, s_{\lambda}, s_{\gamma})\) as in Theorem 4, the first-step iterate \((\beta^{(1)}, \Lambda^{(1)}, \Gamma^{(1)})\) not only remains within the neighborhood \(\mathcal{B}_{\delta_{NT}}\) but also moves closer to the global minimizer (up to an orthogonal transformation) than \((\beta^{(0)}, \Lambda^{(0)}, \Gamma^{(0)})\).
The argument in Step 1 can be applied directly to all subsequent iterations. It is therefore straightforward to verify that each iterate remains within \(\mathcal{B}_{\delta_{NT}}\) and moves closer to the global minimizer (up to an orthogonal transformation) than the previous iterate. This completes the proof.
Lemma 14. Let \(\{x_{it} \in \mathbb{R}^{d_X}\mid i=1, \ldots,N, t = 1, \ldots, T\}\) be a collection of real vectors satisfying \(|x_{it, d}| \leq \rho_X\) for all \(i=1, \ldots,N\), \(t = 1, \ldots, T\), and \(d =1, \ldots, d_X\). Let \(\epsilon_{it}\) be Rademacher random variables independent across \(i\) and \(t\). Then, \[\begin{align} \mathbb{E}_{ \epsilon} \|\sum_{i=1}^{N}\sum_{t=1}^{T}x_{it} \epsilon_{it}\|_2 \leq D_1 \sqrt{NT}, \end{align}\] where \(D_1 = \sqrt{2\pi d_X^3\rho_X^2}\).
Proof. Since \(|x_{it,d}| \le \rho_X\), the random variable \(x_{it,d}\epsilon_{it}\) is independent, mean-zero, and sub-Gaussian with parameter \(\rho_X^2\). Therefore, for any \(d = 1,\ldots,d_X\) and any \(\delta>0\), Hoeffding’s inequality yields \[\begin{align} \mathbb{P}_{\epsilon} \left(|\sum_{i=1}^{N}\sum_{t=1}^{T}x_{it, d} \epsilon_{it}|\geq \delta\right)\leq 2 e^{-\frac{\delta^2}{2 NT \rho_X^2}}. \end{align}\] Hence, by the union bound, \[\begin{align} \left\{\|\sum_{i=1}^{N}\sum_{t=1}^{T}x_{it} \epsilon_{it}\|_2 \geq \delta\right\} &\subset \bigcup_{d=1}^{d_X}\{|\sum_{i=1}^{N}\sum_{t=1}^{T}x_{it, d} \epsilon_{it}| \geq \frac{\delta}{\sqrt{d_X}}\} \\ \Rightarrow \mathbb{P}_{\epsilon}\left(\|\sum_{i=1}^{N}\sum_{t=1}^{T}x_{it} \epsilon_{it}\|_2 \geq \delta \right) &\leq \mathbb{P}_{\epsilon}\left(\bigcup_{d=1}^{d_X}\{|\sum_{i=1}^{N}\sum_{t=1}^{T}x_{it, d} \epsilon_{it}| \geq \frac{\delta}{\sqrt{d_X}}\}\right) \\ & \leq\sum_{d=1}^{d_X} \mathbb{P}_{\epsilon}\left(|\sum_{i=1}^{N}\sum_{t=1}^{T}x_{it, d} \epsilon_{it}| \geq \frac{\delta}{\sqrt{d_X}}\right) \\ & \leq 2d_X \exp\left\{-\frac{\delta^2}{2NTd_X\rho_X^2 }\right\}. \end{align}\] Therefore, \[\begin{align} \mathbb{E}_{\epsilon} \|\sum_{i=1}^{N}\sum_{t=1}^{T}x_{it} \epsilon_{it}\|_2 = \int_{0}^{\infty} \mathbb{P}_{\epsilon} (\|\sum_{i=1}^{N}\sum_{t=1}^{T}x_{it} \epsilon_{it}\|_2 \geq \delta )\mathrm{d}\delta \leq \int_{0}^{\infty}2 d_X e^{-\frac{\delta^2}{2 NT d_X \rho_X^2}}\mathrm{d}\delta = D_1 \sqrt{NT}, \end{align}\] where \(D_1 = \sqrt{2\pi d_X^3\rho_X^2}\).
Lemma 15. Consider a random matrix \(Z\in \mathbb{R}^{N \times T}\) with uniformly bounded entries and \(\mathbb{E}Z = 0\). Assume that the rows of \(Z\) are independent. For each row \(i\), \(\{Z_{it}\}_{1\leq t\leq T}\) is \(\alpha\)-mixing with mixing coefficient \(\alpha_i(\tau)\rightarrow 0\) as \(\tau \rightarrow \infty\), where \[\begin{align} \alpha_i(\tau) = \sup_{t} \sup_{A \in \mathcal{A}^{i}_{t}, B\in \mathcal{B}^{i}_{t + \tau}}|\mathbb{P}(A\cap B) - \mathbb{P}(A)\mathbb{P}(B)|, \end{align}\] where \(\mathcal{A}^{i}_{t}\) is the sigma-field generated by \(\{\ldots, Z_{i,t-1}, Z_{i,t}, \}\), and \(\mathcal{B}^{i}_{t+\tau}\) is the sigma-field generated by \(\{Z_{i,t+\tau}, Z_{i,t+\tau+1}, \ldots\}\). Assume further that the mixing coefficients satisfy a uniform polynomial decay condition: there exist constants \(\beta > 2\) and \(C > 0\) such that \(\sup_{1\leq i\leq N}\alpha_i(\tau) \leq C\tau^{-\beta }\). Then, wpa1, \[\begin{align} \|Z\|_{\mathrm{op}} \lesssim \sqrt{\max\{N, T\}}\log (N + T) \end{align}\]
Proof of Lemma 15. We employ the rectangular matrix Bernstein inequality ([45]), stated as follows:
Theorem 7 ([45]). Consider a finite sequence \(\{Z_k\}\) of independent, random matrices with dimensions \(N\times T\), Assume that each random matrix satisfies \[\begin{align} \mathbb{E}Z_k = 0, \quad \|Z_k\|_{\mathrm{op}} \leq D. \end{align}\] Define \(\sigma^2 = \max\left\{\left\|\sum_{k} \mathbb{E}(Z_kZ_k')\right\|_{\mathrm{op}}, \left\|\sum_{k} \mathbb{E}(Z_k'Z_k)\right\|_{\mathrm{op}}\right\}\). Then for all \(\delta>0\), we have \[\begin{align} \mathbb{P}\left( \|\sum_{k}Z_k\|_{\mathrm{op}}\geq \delta \right) \leq (N + T)\exp\left\{\frac{-\delta^2}{2\sigma^2 + \frac{2}{3} D\delta}\right\}. \end{align}\].
To apply Theorem 7, let \(Z_i\) denote the \(N \times T\) matrix where the \(i\)-th row of \(Z_i\) is the \(i\)-th row of \(Z\), and all other rows are zero. Since the entries of \(Z\) are uniformly bounded, there exists a constant \(a_1>0\), irrelevant with \(N, T\), such that \[\begin{align} \label{eq:llm:independent95entry951} \|Z_{i}\|_{\mathrm{op}} \leq a_1 \sqrt{T} \leq a_1\sqrt{\max\{N, T\}}, \quad \forall i=1,2,\ldots, N. \end{align}\tag{68}\] In addition, it is straightforward to verify that \(\sum_{i=1}^{N} \mathbb{E}(Z_iZ_i')\) has a diagonal structure. Thus, there exists a constant \(a_2>0\), irrelevant with \(N, T\), such that \[\begin{align} \label{eq:llm:independent95entry952} \|\sum_{i=1}^{N} \mathbb{E}(Z_iZ_i')\|_{\mathrm{op}} \leq a_2 T . \end{align}\tag{69}\] To bound the term \(\left\|\sum_{i=1}^{N} \mathbb{E}(Z_i'Z_i)\right\|_{\mathrm{op}}\), observe that \[\begin{align} \mathbb{E}(Z_i'Z_i) = \begin{pmatrix} \gamma_i(1, 0) & \gamma_i(1, 1) & \ldots & \gamma_i(1, T-1) \\ \gamma_i(2, -1) & \gamma_i(2, 0) & \ldots & \gamma_i(1, T-2) \\ \vdots & \vdots & \ddots & \vdots \\ \gamma_i(T, 1- T ) & \gamma_i(T, 2 - T) & \ldots & \gamma_i(T, 0) \end{pmatrix}, \end{align}\] where \(\gamma_i(t, t + \tau) = \mathrm{Cov}(Z_{i, t}, Z_{i, t + \tau })\) for any \(1\leq t\leq T\). By the Gershgorin circle theorem, we show that \(\|\mathbb{E}(Z_i'Z_i)\|_{\mathrm{op}}\) is uniformly bounded by a constant \(a_3>0\). Indeed, \[\begin{align} \|\mathbb{E}(Z_i'Z_i)\|_{\mathrm{op}} & \leq \max_{1\leq t \leq T} \left\{\gamma_i(t, 0) + \sum_{\tau=1 - t}^{-1 }\left|\gamma_i(t, \tau)\right| + \sum_{\tau= 1 }^{T-t}\left|\gamma_i(t, \tau)\right| \right\} . \end{align}\] In addition, since the sequence is \(\alpha\)-mixing with a uniformly polynomial decay rate \(\sup_{1\leq i\leq N}\alpha_i(\tau) \leq C\tau^{-\beta }\), it follows from [46] that \[\begin{align} \left|\gamma_i(t, \tau)\right| \leq 4 \rho_Z^2 \alpha_i(|\tau| )^{\frac{1}{2}} = 4 \rho_Z^2 C^{\frac{1}{2}}|\tau|^{-\frac{\beta}{2}}, \end{align}\] where \(\rho_Z>0\) is a constant, independent of \(N, T\), such that \(|Z_{it}|\leq \rho_Z\) for all \(i, t, N, T\). Therefore, \[\begin{align} \label{eq:llm:independent95entry953} \|\mathbb{E}(Z_i'Z_i)\|_{\mathrm{op}} \leq \max_{1\leq t \leq T} \gamma_i(t, 0) + \sum_{\tau=1}^{T} 8 \rho_Z^2 C^{\frac{1}{2}}\tau^{-\frac{\beta}{2}} \leq \max_{1\leq t \leq T} \gamma_i(t, 0) + 8 \rho_Z^2 C^{\frac{1}{2}}\sum_{\tau=1}^{\infty} \tau^{-\frac{\beta}{2}}. \end{align}\tag{70}\] It is straightforward to verify that \(\max_{1\leq t \leq T} \gamma_i(t, 0)\leq 4\rho_Z^2\), and \(\sum_{\tau=1}^{\infty} \tau^{-\frac{\beta}{2}}<\infty\) as \(\beta>2\). Therefore, there exists a constant \(a_3 >0\), independent of \(N, T\), such that \(\|\mathbb{E}(Z_i'Z_i)\|_{\mathrm{op}}\leq a_3\), and additionally, \(\max_{1\leq i\leq N}\|\mathbb{E}(Z_i'Z_i)\|_{\mathrm{op}}\leq a_3\). Hence, we have \[\begin{align} \|\sum_{i=1}^{N}\mathbb{E}(Z_i'Z_i)\|_{\mathrm{op}} \leq \sum_{i}^{N} \| \mathbb{E}(Z_i'Z_i)\|_{\mathrm{op}}\leq a_3N. \end{align}\] Let \(a_4 = 2\max\{a_2, a_3\}\). Combining 69 and 70 , we obtain \[\begin{align} \label{eq:llm:independent95entry954} \sigma^2 = \max\left\{\|\sum_{i}^{N} \mathbb{E}(Z_iZ_i')\|_{\mathrm{op}}, \|\sum_{i}^{N} \mathbb{E}(Z_i'Z_i)\|_{\mathrm{op}} \right\} \leq a_4 \max\{N, T\}. \end{align}\tag{71}\]
Finally, since \(\{Z_i\}_{1\leq i\leq N}\) are independent and \(Z = \sum_{i=1}^{N} Z_{i}\), we apply Theorem 7 together with 68 and 71 to obtain \[\begin{align} \mathbb{P}\left(\left\|Z\right\|_{\mathrm{op}}\geq \delta \right) & \leq 2\max\{N, T\}\exp\left\{-\frac{\delta^2}{2a_4\max\{N, T\} + \frac{2}{3} a_1 \sqrt{\max\{N, T\}}\delta}\right\} . \end{align}\] Let \(\delta = \max\{4\sqrt{a_4}, 4\sqrt{\frac{2}{3}}a_1\}\sqrt{\max\{N, T\}} \log (N+T)\). Then \[\begin{align} \mathbb{P}\left(\left\|Z\right\|_{\mathrm{op}}\geq \delta \right) & \leq 2 \max\{N, T\} \exp\left\{-2\log(N + T)\right\} \leq 2 \exp\left\{-\log(N+T)\right\} \rightarrow 0. \end{align}\] Therefore, we conclude that \(\|Z\|_{\mathrm{op}}\lesssim \sqrt{\max\{N, T\}}\log (N+T)\) with probability approaching \(1\).
Lemma 16. For any block matrix \(M\in \mathbb{R}^{(n+m)\times (n+m)}\) \[\begin{align} M = \begin{pmatrix} A & B\\ B' & 0 \end{pmatrix}, \end{align}\] where \(A\in \mathbb{R}^{n\times n}\) is a positive definite matrix, \(B\in \mathbb{R}^{n\times m}\), we have \[\begin{align} \label{eq:eigenvalue95ABB0} \sigma_{\min} (M) \geq \frac{1}{2}\left(\sigma_{\min}(A) - \sqrt{\sigma^2_{\min}(A) + 4 s_{\max}^2(B)}\right). \end{align}\qquad{(3)}\] where \(\sigma(\cdot)\) denotes the eigenvalue and \(s(\cdot)\) denotes the singular value.
Proof of Lemma 16. When \(B = 0\), inequality ?? holds trivially since \(\sigma_{\min}(M) = 0\). For \(B \neq 0\), the matrix \(M\) admits negative eigenvalues; denote one of them as \(\varsigma <0\). Let \((x',y')'\) be a corresponding eigenvector. By the definition of eigenvalue, we have \[\begin{align} M(x', y')' = \varsigma (x', y')' \Rightarrow \left\{ \begin{array}{l} Ax + By = \varsigma x \\ B'x = \varsigma y \end{array} \right. . \end{align}\] Since \(\varsigma \neq 0\), substituting \(y = \varsigma^{-1} B'x\) into \(Ax + By = \varsigma x\) gives \[\begin{gather} \varsigma^2 x'x - \varsigma x'Ax - x'BB'x = 0 \quad \Rightarrow \quad \varsigma = \frac{1}{2}\left(\frac{x'Ax}{x'x } - \sqrt{\left(\frac{x'Ax}{x'x }\right)^2 + 4 \frac{x'BB'x}{x'x}}\right), \end{gather}\] where we select the negative root since \(\varsigma < 0\). It is straightforward to verify that \(\varsigma\) is increasing in \(\frac{x'Ax}{x'x }\) and decreasing in \(\frac{x'BB'x}{x'x}\). Since \[\begin{align} \frac{x'Ax}{x'x } \geq \min_{x\neq 0 }\frac{x'Ax}{x'x } = \sigma_{\min}(A), \quad\frac{x'BB'x}{x'x} \leq \max_{x} \frac{x'BB'x}{x'x} = s^2_{\max}(B), \end{align}\] we conclude that \[\begin{align} \sigma_{\min} (M) \geq \frac{1}{2}\left(\sigma_{\min}(A) - \sqrt{\sigma^2_{\min}(A) + 4 s_{\max}^2(B)}\right). \end{align}\] This completes the proof.
Lemma 17. For any block matrix \(A = [A_1, A_2]\), where \(A_1 \in \mathbb{R}^{n\times m_1}, A_2 \in \mathbb{R}^{n\times m_2}\), we have \(s_{r}(A)\geq s_{r}(A_1)\), for \(r = 1,\ldots, \min\{n, m_1 \}\). Consequently, \(\|A\|_{\mathrm{nuc}}\geq \|A_1\|_{\mathrm{nuc}}\).
Proof of Lemma 17. We use the Courant-Fischer variational characterization of singular values. Let \(u = (u_1', u_2')'\), where \(u_1\in \mathbb{R}^{m_1}\), \(u_2\in \mathbb{R}^{m_2}\). Then \[\begin{align} s_r(A) = \min_{\mathrm{dim}(V)=r-1} \max_{u\perp V, \|u\|=1} \|Au\|. \end{align}\] Since \(A = [A_1,A_2]\), for any \(u_1\in\mathbb{R}^{m_1}\),define \(\tilde{u}=(\tilde{u}_1',0)' \in \mathbb{R}^{m_1+m_2}\). Then \(A \tilde{u}= A_1 \tilde{u}_1\). Therefore, for any \((r-1)\)-dimensional subspace \(V_1\subset\mathbb{R}^{m_1}\), let \[\begin{align} \tilde{V} := \{(\tilde{v}_1',0)' : \tilde{v}_1\in V_1\}\subset \mathbb{R}^{m_1+m_2}. \end{align}\] Then \[\begin{align} \max_{u \perp \tilde{V}, \|u\|=1} \|Au\| \geq \max_{ u_1 \perp V_1, \|u_1\|=1} \|A_1 u_1\| \end{align}\] Taking the minimum over all \((r-1)\)-dimensional subspaces gives \[\begin{align} s_r(A) = \min_{\dim(V)=r-1}\max_{ u\perp V, \|u\|=1}\|Au\| \geq \min_{\dim(V_1)=r-1} \max_{u_1\perp V_1, \|u_1\|=1}\|A_1 u_1\| = s_r(A_1). \end{align}\] Finally, \[\begin{align} \|A\|_{\mathrm{nuc}}=\sum_{r=1}^{\min\{n,m_1+m_2\}} s_r(A) \geq \sum_{r=1}^{\min\{n,m_1 \}} s_r(A) \geq\sum_{r=1}^{\min\{n,m_1\}} s_r(A_1) =\|A_1\|_{\mathrm{nuc}}. \end{align}\]
This completes the proof.
University College London and CeMMAP: a.zeleneev@ucl.ac.uk.↩︎
University College London: weisheng.zhang.21@ucl.ac.uk.↩︎
We thank Iván Fernández-Val, Aureo de Paula, Martin Weidner and the participants of the 2nd UCL–CeMMAP–IFS Ph.D. Econometrics Research Day (2024) and UCL Econometrics Brownbag Seminar for their valuable comments. We thank Martin Weidner for sharing the codes and data from [1], and thank Wei Miao for careful testing of the R package and for helpful comments on the implementation.↩︎
For example, [4] proposes a method for estimating network models with (nonparametric) interactive unobserved heterogeneity that does not require solving a high-dimensional nonconvex problem. However, unlike [1], [4] focuses on identification and consistent estimation and does not provide inference tools.↩︎
Available at: https://github.com/wszhang-econ/NNRPanel.↩︎
[1] also argue that the interactive fixed effects model is sufficiently flexible to allow for homophily based on unobservables (as well as for degree heterogeneity) in network settings.↩︎
E.g., \(F(\cdot)\) can stand for the logistic or standard normal CDF in Logit and Probit models, respectively.↩︎
The representation of \(\Theta\) with \(\mathrm{rank}(\Theta)\leq R\) as \(\Theta = \Lambda \Gamma'\) is not unique. If \(\Theta = \Lambda \Gamma'\) for some \(\Lambda\) and \(\Gamma\), we also have \(\Theta = \tilde{\Lambda} \tilde{\Gamma}'\) for \(\tilde{\Lambda} = \Lambda G'\) and \(\tilde{\Gamma} = \Gamma G^{-1}\) for any invertible matrix \(G \in \mathbb{R}^{R \times R}\). This non-uniqueness manifests itself in the necessity of normalizing \(\Lambda\) and \(\Gamma\) in problem 2 in order to ensure the uniqueness of \(\hat{\Lambda}_{\mathrm{FE}}\) and \(\hat{\Gamma}_{\mathrm{FE}}\).↩︎
In principle, instead of using a gradient descent method, it is possible also employ an EM-algorithm (see, e.g., [1], [25]) initialized at \((\hat{\beta}_{\mathrm{nuc}}, \hat{\Lambda}_{\mathrm{nuc}}, \hat{\Gamma}_{\mathrm{nuc}})\).↩︎
The boundedness of \(\Lambda\) and \(\Gamma\) could, in principle, be replaced by assuming that \(\{\lambda_i\}_{1\leq i\leq N}\) and \(\{\gamma_t\}_{1\leq t\leq T}\) are sub-Gaussian sequences that are independent across \(i\) and weakly dependent over \(t\), respectively.↩︎
Developing estimation and inference methods robust to weak factors is an important but highly nontrivial problem, even in linear panels; see [17]. In this paper, we simply follow the set-up of [1] and leave the important problem of allowing for weak factors in nonlinear models for future research.↩︎
Originally introduced by [30], the RSC condition (in its various forms) has been widely applied in problems involving estimation of low-rank matrices, including matrix completion [31], reduced-rank regression [32], latent community detection [7], and econometric analysis of panel models with interactive unobserved heterogeneity [5], [6].↩︎
Notice that \(\kappa\), \(\eta\), and \(\mathcal{C}_1\) introduced in Assumption 2 all depend on \(c_0\), but we suppress this dependence for ease of notation.↩︎
The sample Hessian is a matrix-valued function of a \(d_X\)-dimensional vector \(\beta\), an \(N \times R\) parameter matrix \(\Lambda\), and a \(T \times R\) parameter matrix \(\Gamma\). These parameters are arranged as follows: \[(\beta', \text{vec}(\Lambda')', \text{vec}(\Gamma')')',\] where \(\text{vec}(\cdot)\) denotes the vectorization operator, stacking the columns of a matrix into a vector.↩︎
Notice that \(\|\nabla_{\beta}\mathcal{L}_{NT}(\beta_0, \Theta_0)\|_2\) is negligible compared to \(\sqrt{NT}\|\nabla_{\Theta}\mathcal{L}_{NT}(\beta_0, \Theta_0)\|_{\mathrm{op}}\) as \(N,T\rightarrow \infty\).↩︎
For the \(\mathrm{TS}^*\) estimator, we simply apply Algorithm 3 using the true number of factors.↩︎
See Appendix 7.2 for bias-correction implementation details.↩︎
The average number of the estimated factors is reported in the last column of the table and denoted by \(\bar{R}\).↩︎
The incorporation of additive fixed effects requires only minor modifications of the optimization algorithms presented in Section 4. These modifications and additional implementation details are provided in Appendix 7.1.↩︎
For any \(R\)-dimensional non-singular matrix \(G\), the conditional distribution of \(Y_{it}\) remains invariant under the transformations \(\lambda_i \mapsto \lambda_i G'\) and \(\gamma_t \mapsto \gamma_t G^{-1}\). This invariance allows us to freely choose different normalization methods for different purposes without affecting the inference of \(\beta_0\).↩︎
When the eigenvalues of \(\Sigma_{\lambda}\Sigma_{\gamma}\) are distinct, \(G\) is unique and does not depend on the sample \(\{(Y_{it}, X_{it})\}_{1\leq i\leq N, 1\leq t\leq T}\). However, with possible repeated eigenvalues, \(G\) is not unique and can be identified up to an orthogonal transformation.↩︎
Because by 39 , \[\begin{align} \|\tilde{\hat{V}}^c \|_{\mathrm{op}}^2 \leq \|\tilde{\hat{V}}^c\|_{\mathrm{F}}^2 = O_p\left((|\mathcal{I}_{NT}^c| + |\mathcal{I}_{NT}^c|)(N+T)\right) = O_p(NT\max\{N, T\}\delta_{NT}^4). \end{align}\] Then, we have \(\|\tilde{\hat{V}}^c \|_{\mathrm{op}} = o_p\left(\min\{N, T\}\right)\) and \(\|\mathbb{E}_0 \tilde{\hat{V}}^c \|_{\mathrm{op}} = o_p\left(\min\{N, T\}\right)\).↩︎