Semi-nonparametric estimation of spatial dynamic panel data models with nonparametric spatial weights4


Abstract

We develop a semi-nonparametric framework for spatial dynamic panel data (SDPD) models with two-way fixed effects when the spatial interaction structure is unknown beyond a distance measure. This is accomplished by modelling spatial weights in the outcome, lagged-outcome, and disturbance channels as unknown functions of underlying economic distances. These enter the SDPD system through matrix-function operators, providing a unified approach that accommodates both spatial autoregressive and matrix exponential spatial specifications. Allowing for unknown heteroskedasticity, we propose sieve GMM estimators based on a stacked set of linear and quadratic moment conditions, and derive a feasible optimal GMM estimator and a more efficient feasible best GMM estimator. As \((n,T)\to\infty\), the parametric component is \(\sqrt{n(T-1)}\)-consistent and asymptotically normal, echoing classical semi-nonparametric results. Monte Carlo experiments indicate excellent finite-sample performance. We apply the method to ‘witch’ killings as studied by [1], and find that economic-geography proximity rather than cultural-geography proximity between communities significantly amplifies spatial dependence in these economic murders.


 
JEL classification: C13, C23


Keywords: Spatial dynamic panel data models, Unknown spatial weights, Sieve-based GMM estimation, Unknown heteroskedasticity, semiparametric, nonparametric

1 Introduction↩︎

This paper develops a general framework for semi-nonparametric estimation of and inference on spatial dynamic panel data (SDPD) models with two-way fixed effects (FE), in the sense that the model contains both finite dimensional and infinite dimensional parameters of interest. SDPD models provide a workhorse framework for quantifying spillovers across units and over time, with applications in trade, public finance, housing, and regional economics. We contribute to the spatial autoregressive (SAR) literature that continues to develop extensively to accommodate various data structures and estimation challenges. Foundational studies established quasi-maximum likelihood (QML) and generalized method of moments (GMM) approaches for basic SDPD specifications [2][4]. This literature has subsequently expanded to address specific panel dynamics, such as spatial cointegration and short panels [5], [6], as well as unobserved heterogeneity, including interactive fixed effects and latent group structures [7][10]. The SDPD framework is particularly versatile because it simultaneously incorporates three channels of spatial spillovers: the outcome, the lagged outcome, and the disturbance. A leading specification with two-way FE allows for contemporaneous spatial-lag dependence, space-time-lag dependence, and spatially correlated disturbances [11].

A key maintained assumption in such models is that the spatial interaction matrices are prespecified. However, this assumption may be too strict in empirical applications. Spatial weights are typically constructed from geographic or economic proximity, and misspecification may bias estimated spillovers and distort counterfactual policy conclusions. Even within a single channel, a single pre-specified matrix can be restrictive when spillovers operate through multiple distance components or higher-order paths. Moreover, linear spatial-lag terms may fail to capture nonlinear interaction patterns generated by matrix function operators, such as matrix exponential spatial specifications (MESS); see, e.g. [12][15]. In many empirical settings, researchers only know that the interaction depends on exogenous distances, but a specific parametric link is unavailable [16].

Motivated by these considerations, we estimate the spatial interaction structure directly from the data. We allow the weight between units in each channel to be an unknown function of observable, exogenous, and time-invariant distances. These unknown distance link functions are mapped into interaction matrices through matrix functions. This formulation, which we formalize in Section 2, preserves the three-channel SDPD structure while allowing each channel to feature its own data-driven spatial interaction pattern, including SAR and MESS specifications.

Directly estimating these weights without structural assumptions or a penalized approach (see e.g. [17]) is infeasible because the number of unknown links grows quadratically with the cross-sectional dimension \(n\). To accommodate settings where \(n\) is comparable to or larger than the time dimension \(T\), we impose structure through an observed exogenous distance measure, and approximate each unknown link function by a sieve expansion. This reduces the number of unknown links from quadratic in \(n\) to a much smaller set of sieve coefficients that grows only slowly with sample size \(nT\).

There is a growing literature on estimating unknown spatial links. For cross-sectional models, [16] and [18] estimate a distance-based weights function via series approximation. For static spatial panels, [17] and [19] exploit sparsity using LASSO-type methods, though their theory requires \(T\) to be large relative to \(n\). In contrast, [20] adopts a series approximation that accommodates large \(n\) when \(T\) is finite or moderately large. In this paper, we also employ a series-based approach, but extend the scope in two significant ways. First, whereas existing research typically allows spatial interaction only through the outcome variable, our framework simultaneously incorporates all three SDPD channels: the outcome, the lagged outcome, and the disturbance. Second, based on the SDPD model, we specify the unknown spatial weights themselves as varying coefficients that depend on an exogenous distance measure. This marks a departure from the existing varying-coefficient SAR literature, which typically treats the spatial weights matrix as pre-specified and instead allows the regression coefficients to vary with covariates (see, e.g. [21][24]).

In this paper, we propose the sieve generalized method of moments estimator (GMME) for the SDPD models with unknown spatial weights matrices. Using sieve-based GMMEs in the case of unknown heteroskedasticity, we construct a feasible optimal GMME (OGMME) and a more efficient feasible best GMME (BGMME) based on quadratic and linear moments. Our study is closely related to [16], [18], and [20]. The first two studies consider cross-sectional SAR models and derive the two-stage least squares (2SLS) using linear moment conditions. The latter study analyzes a static SAR spatial panel model with individual FE and derives the efficient 2SLS estimator (2SLSE), again based on linear moments.

Our framework differs from these studies in two important aspects. First, we develop a unified SDPD framework that accommodates both SAR and MESS specifications, contrasting with [16], [18], and [25], who restrict their attention to SAR models in cross-sectional or static panel settings. This is non-trivial because the interaction between the time-lagged dependent variable and the sieve approximation error introduces complex bias terms that are absent in static settings and must be controlled. Second, our semi-nonparametric SDPD model considers unknown heteroskedastic disturbances and develops GMM estimators based on quadratic and linear moments. This setting has not been previously explored. 5 Under certain rate conditions, the estimator of the finite-dimensional parameters is \(\sqrt{n(T-1)}\)-consistent and asymptotically normal as \((n,T)\to\infty\). Monte Carlo experiments demonstrate that the proposed estimators exhibit excellent finite-sample performance.

We then apply our methods to investigate the spillover effects of so-called ‘witch’ killing. Many studies have noted that poverty and violence occur simultaneously [26][30]. There is a strong negative relationship between economic growth and crime across countries [1] . [1] empirically finds that exogenous extreme rainfall causes lower income in the region, and leads to a large increase in the murder of ‘witches’, who are nearly all elderly women. The results show that income shocks are the driver of the increase in witch killings, instead of a scapegoat culture. Based on [1], we apply our model to investigate the mechanism of spillover effects in the study of witch killings. Our results show that the economic-geography proximity rather than the cultural-geography proximity between communities significantly amplifies spatial dependence in witch killings.

The rest of this paper is organized as follows. Section 2 introduces the formal model specification, Section 3 provides the sieve-based GMM estimation methods, and Section 4 investigates the large sample properties of GMMEs. This includes a discussion of the asymptotic efficiency of the proposed estimators under unknown heteroskedasticity. Monte Carlo results for various estimators are provided in Section 5. Section 6 presents our empirical applications to the witch killings. The appendix contains the regularity assumptions, key lemmas, and proofs of the theorems, while the online supplementary material provides additional lemmas and simulation results.
Notation. For a real symmetric matrix \(H\), let \(\mu_{\max}(H)\) and \(\mu_{\min}(H)\) denote its largest and smallest eigenvalues, and let \(\rho(H)\) denote its spectral radius, which is defined as the maximum absolute value of its eigenvalues. For a generic real matrix \(H\), we write \(\|H\|\) and \(\|H\|_{\mathrm{sp}}\) for the Frobenius and spectral norms, respectively. Let \(\|H\|_{\mathrm{rc}}\equiv\max\{{\|H\|_{1},\|H\|_{\infty}}\}\), where \(\|H\|_{1}\) stands for its column sum norm and \(\|H\|_{\infty}\) stands for its row sum norm. When \(H\) is square, let \(H^{s}=H+H'\), let \(\mathtt{diagM}(H)\) denote the vector of diagonal entries of \(H\), and let \(\mathtt{Diag}_{i=1}^r(H_{i})\) denote the block diagonal matrix whose diagonal blocks are \(H_1, \cdots, H_r\).

2 Model specification↩︎

For \(i=1,\dots,n\) and \(t=1,\dots,T\), a standard parametric SDPD model typically assumes that an observable outcome \(y_{it}\) is a linear function of \(\sum_{j=1}^{n}w_{1ij}y_{jt}\) and \(\sum_{j=1}^n w_{2ij}y_{j,t-1}\), for some known ‘spatial weights’ \(w_{1ij}\) and \(w_{2ij}\). We extend this to a semi-nonparametric setting by modeling the spatial links as unknown functions of observable exogenous distances. The scalar form of the model for unit \(i\) at time \(t\) can be expressed as: \[\setlength{\abovedisplayskip}{3pt} y_{it} = \sum_{j=1}^n g_1(d_{ij})y_{jt}+\gamma y_{i,t-1} + \sum_{j=1}^n g_2(d_{ij})y_{j,t-1}+x_{it}'\beta+c_i+\alpha_{t}+ u_{it}, \setlength{\belowdisplayskip}{3pt}\] with \(u_{it}=\sum_{j=1}^n g_3(d_{ij})u_{jt}+\epsilon_{it}\), for unknown functions \(g_k(\cdot), k=1,2,3\). Here, \(d_{ij}\) represents some observed time-invariant exogenous distance measures between units \(i\) and \(j\), \(x_{it}\) is an \(\ell_x\times 1\) vector of exogenous regressors, \(c_i\) and \(\alpha_t\) are unobserved individual and time fixed effects, respectively, and \(\epsilon_{it}\) is an unobserved error term. By stacking these observations across all \(n\) units, we generalize these interactions through matrix-function operators, which allows us to provide a unified approach that accommodates both SAR and MESS specifications. Specifically, the matrix form of our model is: \[\label{eq:sdpd95np} B_{1}Y_{t}=(\gamma I_n+B_2)Y_{t-1}+X_{t}\beta+\mathbf{c}_{n}+\alpha_{t}l_{n}+U_{t}, \quad B_{3}U_{t}=E_{t},\tag{1}\] where the elements of the \(n\times n\) matrices \(B_k\) are determined by the unknown distance link functions \(g_k(d_{ij})\). These link functions form the spatial weight matrices \(G_k=\{g_k(d_{ij})\}\), which in turn characterize the operators \(B_k\) as follows: \[\begin{align} \label{eq:matrix32function} \begin{aligned} \text{SAR:} \; B_1&=I_n-G_1, \;B_2=G_2, \;B_3=I_n-G_3; \\ \text{MESS:} \; B_k&=e^{G_k} \;\text{for k=1,2,3}, \end{aligned} \end{align}\tag{2}\] with \(I_{n}\) the \(n \times n\) identity matrix. Moreover, \(Y_{t}=(y_{1t},y_{2t},...,y_{nt})'\), \(E_{t}=(\epsilon_{1t},\epsilon_{2t},...,\epsilon_{nt})'\) are \(n \times 1\) column vectors and the \(\epsilon_{it}\) are independent with zero means and homoskedastic or unknown heteroskedastic variances across \(i\) and \(t\), and \(U_{t}=(u_{1t},...,u_{nt})'\), which is a SAR/MESS process. \(X_{t}\) is an \(n\times \ell_x\) matrix of time varying regressors, \(\mathbf{c}_{n}=(c_{1},...,c_{n})'\) denotes the vector of unobserved individual FE, \(a_{t}\) is a scalar unobservable time FE, and \(l_{n}\) is the \(n\times 1\) vector of ones. The initial condition \(Y_0\) is exogenously given. We impose the normalization \(l_{n}'\mathbf{c}_{n}=0\) to avoid the loss of identification of \(c_{i}\) and \(\alpha_{t}\) because \(c_{i}+\alpha_{t}=(c_{i}-\nu)+(\alpha_{t}+\nu)\) for an arbitrary \(\nu\). For both specifications, we ensure system stability by imposing the spectral-radius condition \(\rho(A)<1\), where \(A=B_{1}^{-1}(\gamma I_n+B_2)\).

Remark 1. For the MESS specification, setting \(\gamma=\varphi-1\) allows the model in 1 to be interpreted in the intuitive form \(e^{G_1}Y_{t} = (\varphi I_n + (e^{G_2}-I_n)) Y_{t-1} + X_{t}\beta+\mathbf{c}_{n}+\alpha_{t}l_{n}+U_{t}\) with \(e^{G_3}U_{t}=E_{t}\). In this representation, \(\varphi Y_{t-1}\) captures the pure time autoregressive persistence, while the operator \((e^{G_2}-I_n)\) acting on \(Y_{t-1}\) isolates the spatio-temporal spillover effect. By the standard matrix exponential expansion \(e^{G_2}-I_n=\sum_{l=1}^{\infty}\frac{G_2^l}{l!}\), the MESS specification captures higher-order spillovers with factorially decaying weights. This stands in distinct contrast to the SAR specification, where the spatio-temporal effect is driven solely by \(G_{2}Y_{t-1}\), restricting the immediate spillover to first-order neighbors. Unlike the SAR specification which requires \(\rho(B_1)<1\) and \(\rho(B_3)<1\) for invertibility, the MESS ensures \(B_1\) and \(B_3\) are invertible for any finite \(G_1\) and \(G_3\) [14].

Throughout the paper, the subscript \(k=1, 2, 3\) refers to the contemporaneous outcome, lagged outcome, and disturbance channels, respectively. For estimation, we proceed by approximating \(g_{k}(\cdot)\) via series, i.e., \(g_{k}(d)=\sum_{p_{k}=1}^{\ell_{k}}\lambda_{p_{k}}\phi_{kp_{k}}(d)+\delta_{k}(d)\), where \(\phi_{kp_{k}}(\cdot)\) are basis functions and \(\delta_{k}(d)\) are suitably decaying approximation errors and the series lengths \(\ell_{k}\) diverge to infinity as \((n,T)\to\infty\). We then denote the sieve approximation and relevant matrices as \(\xi_{k}(d)=\sum_{p_{k}=1}^{\ell_{k}}\lambda_{p_{k}}\phi_{kp_{k}}(d)\), \(\Xi_{k}=\{\xi_{k}(d_{ij})\}\), and \(\varPhi_{kp_{k}}=\{\phi_{kp_{k}}(d_{ij})\}\). By combining 2 with these approximations and defining \(\Delta_{k}=\{\delta_{k}(d_{ij})\}\), we obtain the decomposition \(B_{k}=S_{k}+R_{k}\), where \(S_{k}\) depends on the sieve matrix \(\Xi_{k}\) and \(R_{k}\) collects the approximation error matrix. Specifically, for the SAR, \(S_{1}=I_{n}-\Xi_{1}, S_{2}=\Xi_{2}, S_{3}=I_{n}-\Xi_{3}\) with \(R_{1}=-\Delta_{1}, R_{2}=\Delta_{2}, R_{3}=-\Delta_{3}\); for the MESS, \(S_{k}=e^{\Xi_{k}}\) and \(R_{k}=S_{k}(e^{\Delta_{k}}-I_{n})\).

Substituting these into 1 , the relationship between \(Y_{t}\) and \(E_{t}\) can be reformulated as: \[\label{eq:mess95appro95Et} S_{3}S_{1}Y_{t}=S_{3}\bigl((\gamma I_{n}+S_{2})Y_{t-1}+X_{t}\beta+\mathbf{c}_{n}+\alpha_{t}l_{n}\bigr)+V_t,\tag{3}\] where the error term is \(V_{t}=r_{t}+E_{t}\) with \(r_{t}=\sum_{k=1}^3 r_{kt}\), and \(r_{1t}=-S_{3}R_{1}Y_{t}\), \(r_{2t}=S_{3}R_{2}Y_{t-1}\), and \(r_{3t}=-R_{3}U_{t}\). These components capture errors arising from contemporaneous outcome interactions, lagged outcome interactions, and disturbance feedback, respectively.

Remark 2. Our setup formally covers ‘higher-order’ SAR models, see e.g. [31]. These would simply have approximation error equal to zero. Our focus though is on the truly nonparametric case with a non-zero approximation error.

3 GMM estimation↩︎

To eliminate individual FE, we employ the forward orthogonal difference (FOD) transformation [4], [32], [33], rather than the first difference (FD) approach used in studies such as [22]. Specifically, for any \(n \times T\) matrix of variables \([K_1, \dots, K_T]\), we obtain the \(n\times (T-1)\) transformed matrix \([K_1^*, \dots, K_{T-1}^*]\) by applying an orthogonal transformation matrix \(F_{T,T-1}\), where the algebraic details are in the appendix. The merit of the FOD transformation is that it yields uncorrelated transformed idiosyncratic terms \(E_t^*\), preserving their serial uncorrelatedness, whereas FD induces an MA(1) process even when the original errors are serially uncorrelated. 6 Time FE can be eliminated by \(J_{n}=I_{n}-\frac{1}{n}l_nl_n'\), recalling that \(\alpha_{t}J_{n}l_{n}=0\). After these transformations, the reformulated model 3 for \(t=1,\dots,T-1\) becomes: 7 \[\begin{align} \label{eq:mess32nofe} J_{n}S_{3}S_{1}Y_{t}^{*}=\gamma J_{n}S_{3}Y_{t-1}^{(*,-1)}+J_{n}S_{3}S_{2}Y_{t-1}^{*}+J_{n}S_{3}X_{t}^{*}\beta +J_{n}V_{t}^{*}, \end{align}\tag{4}\] where \(Y_{t}^{*}\) and \(Y_{t-1}^{(*,-1)}\) represent the FOD-transformed contemporaneous and lagged outcome vectors, respectively. The transformed regressor matrix \(X_{t}^{*}\) and the error term \(V_{t}^{*}\) are defined analogously. These variables are constructed such that the transformation for a period \(t\) depends only on current and future observations, and the algebraic construction of these transformation matrices is shown in Section 8 of the appendix. We now introduce some useful notation for the estimates. Denote \(N=n(T-1)\) and \(\theta=(\pi',\lambda')'\), where \(\pi=(\gamma,\beta')'\) and \(\lambda=(\lambda_{1}^{\prime},\lambda_{2}',\lambda_{3}')'\) with \(\lambda_{k}=(\lambda_{k1},...,\lambda_{k\ell_{k}})'\) for \(k=1,2,3\), and \(\dim(\theta)=\ell_{\theta}\). Furthermore, \(\hat{g}_{k}(d)=\pmb{\phi}_{k}'(d)\hat{\lambda}_{k}\), where \(\pmb{\phi}_{k}(d)=[\phi_{k1}(d),...,\phi_{k\ell_{k}}(d)]\) and \(\hat{\lambda}_{k}\) is a generic estimator of \(\lambda_{k}\).

3.1 Moment conditions↩︎

For the estimation of equation 4 , we define \(\mathbf{Q}_{N}=\mathtt{Diag}_{t=1}^{T-1}(Q_{t})\) as the block-diagonal instrument variables (IV) matrix, where \(Q_{t}\) is an \(n\times\ell_{q}\) matrix of instrumental variables. A typical construction of \(Q_t\) includes \(Y_{t-1}\), \(X_{t}^*\), and their interactions with the sieve basis functions \(\varPhi_{kp_{k}}\). For example, \[\begin{align} Q_{t}=\bigl(Y_{t-1}, X_{t}^{*},[\varPhi_{21},\ldots,\varPhi_{2\ell_{2}},\varPhi_{31}\varPhi_{21},\ldots,\varPhi_{3\ell_{3}}\varPhi_{2\ell_{2}}]Y_{t-1}, [\varPhi_{11},..., \varPhi_{1\ell_{1}}, \varPhi_{31}\varPhi_{11}, \ldots,\varPhi_{3\ell_{3}}\varPhi_{1\ell_{1}}]X_{t}^{*}\bigr). \end{align}\] We that \(\ell_{q}\sim \ell_{n}\), where the symbol ‘\(\sim\)’ denotes the same rate of growth, and \(\ell_{n}=\max_{1\leq k\leq 3}\ell_k\). The linear moment condition is \(m^{\texttt{line}}_{N}(\theta)=\mathbf{Q}_{N}'\mathbf{J}_{N}\mathbf{V}_{N}^{*}(\theta),\) where \(\mathbf{J}_{N}=I_{T-1}\otimes J_{n}\), \(\mathbf{V}_{N}^{*}(\theta)=\left(\mathbf{V}_{1}^{*\prime}(\theta),\mathbf{V}_{2}^{*\prime}(\theta),...,\mathbf{V}_{T-1}^{*\prime}(\theta)\right)'\) is the vector of transformed residuals and \(\otimes\) denotes the Kronecker product. Note that \(Y_{t-1}\) serves as an asymptotically valid instrument for \(Y_{t-1}^{(*,-1)}\) since, as can be derived from Lemma 14() in the supplement, the terms involving approximation errors \(Y_{t-1}'r_t^{*}\) are asymptotically negligible. While theoretically valid, the potential for finite sample bias is explored further in our Monte Carlo simulations.

An initial estimator↩︎

Motivated by [16], [18] and [20], we consider the two-stage least square estimator (2SLSE) derived from the linear moment \(m^{\texttt{line}}_{N}(\theta)\), where \(\hat{\theta}_{ts}=\arg\min_{\theta\in\Theta}m^{\texttt{line}\prime}_{N}(\theta)\boldsymbol{M}_{Q}m^{\texttt{line}}_{N}(\theta) \;\text{with} \; \boldsymbol{M}_{Q}=\mathbf{J}_{N}\mathbf{Q}_{N}(\mathbf{Q}_{N}'\mathbf{J}_{N}\mathbf{Q}_{N})^{-1}\mathbf{Q}_{N}'\mathbf{J}_{N}.\)

3.2 Optimal GMM estimation↩︎

We now consider GMM estimation, for which we can construct additional quadratic moment conditions. We assume that the variance structure of \(E_{t}\), denoted by \(\Sigma_{t}\), has three cases: (V0) : \(\sigma^2 I_{n}\), (V1) : \(\mathtt{Diag}_{i=1}^{n}(\sigma_{i}^2)\), and (V2) : \(\sigma_{t}^2I_{n}\). There is a sense in the extant literature that combining both the linear and quadratic moments could improve GMM estimation efficiency. For related parametric SDPD models, see e.g. [4] and [8]. For related nonparametric spatial panel data models, see e.g. [22] and [34]. Specifically, \(m_{N}^{\texttt{quad}}(\theta)=\Big(\mathbf{V}_{N}^{*\prime}(\theta)\mathbf{J}_{N}\mathbf{P}_{N1}'\mathbf{J}_{N}\mathbf{V}_{N}^{*}(\theta),...,\mathbf{V}_{N}^{*\prime}(\theta)\mathbf{J}_{N}\mathbf{P}_{N\ell_{p}}'\mathbf{J}_{N}\mathbf{V}_{N}^{*}(\theta)\Big)'.\) Here, \(\mathbf{P}_{Nk}=\mathtt{Diag}_{t=1}^{T-1}(P_{kt})\), \(P_{kt}\) is some \(n\times n\) square matrix satisfying \(\mathtt{diagM}(J_{n}P_{kt}J_{n})=\mathbf{0}_{n\times 1}\). Lemma 1() in the appendix provides a way of choosing \(P_{kt}\).

Then, the stacked moment condition combining linear and quadratic moments is \[\begin{align} \label{eq:optimal95moments} m_{N}(\theta)=\frac{1}{n(T-1)}(m_{N}^{\texttt{quad}\prime}(\theta),m_{N}^{\texttt{line}\prime}(\theta))'. \end{align}\tag{5}\] According to [35], the optimal weights of the moments can be approximated by \[\begin{align} \label{eq:Var95moment} \Omega_{N}=\frac{1}{n(T-1)}\mathtt{Diag}\Big(\mathbf{\Omega}_{N1},\mathbf{\Omega}_{N2}\Big), \end{align}\tag{6}\] where \([\mathbf{\Omega}_{N1}]_{ij}=\operatorname{tr}(\mathbf{\Sigma}_{N}\mathbf{J}_{N}\mathbf{P}_{Ni}\mathbf{J}_{N}\mathbf{\Sigma}_{N}\mathbf{J}_{N}\mathbf{P}_{Nj}^{s}\mathbf{J}_{N})\), for \(i,j=1,...,\ell_{p}\), \(\mathbf{\Omega}_{N2}=\mathbf{Q}_{N}'\mathbf{J}_{N}\mathbf{\Sigma}_{N}\mathbf{J}_{N}\mathbf{Q}_{N}\), \(\mathbf{\Sigma}_{N}=\mathtt{Diag}_{t=1}^{T-1}(\Sigma_{t})\), and \(\mathbf{P}_{Nj}^{s}=\mathbf{P}_{Nj}+\mathbf{P}_{Nj}'\). The feasible optimal GMME (OGMME) is \(\hat{\theta}_{ogmm}=\arg\min_{\theta\in\Theta}m_{N}(\theta)\tilde{\Omega}_{N}^{-1}m_{N}(\theta),\) where \(\tilde{\Omega}_{N}\) is a consistent estimator of \(\Omega_{N}\).

3.3 Best GMM estimation↩︎

In this section, we consider the case in which the OGMME attains the smallest asymptotic variance, referred to as the BGMME [4], [31]. To the best of our knowledge, there is no such work on the semi-nonparametric SDPD model. For the parametric model, [4] considers only the homoskedastic case, which cannot be directly extended to the heteroskedastic setting. Furthermore, when heteroskedasticity is unknown, we must impose additional structure on \(\Sigma_{t}\) to construct a consistent estimator \(\hat{\Sigma}_{t}\), as we do in our cases [var:hetei] and [var:hetet]. 8 To eliminate time FE under unknown heteroskedasticity and to derive the best moment conditions, we introduce the following \(n \times n\) projection matrix: \[\begin{align} \label{eq:MSL} \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }=I_{n}-\Sigma_{t}^{-\frac{1}{2}}l_{n}\bigl(l_{n}'\Sigma_{t}^{-1}l_{n}\bigr)^{-1}l_{n}'\Sigma_{t}^{-\frac{1}{2}}, \end{align}\tag{7}\] which is idempotent and satisfies \(\ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\Sigma_{t}^{-\frac{1}{2}}\,l_{n} = 0\). As discussed in Section 12 of the supplement, the projection matrix \(\ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\) is well-behaved. Thus, we can employ \(\ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\) to eliminate the time FE even when \(\Sigma_{t}\) is heteroskedastic and unknown. We then pre-multiply \(\ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\Sigma_{t}^{-\frac{1}{2}}\) and employ the FOD transformation on both sides of 3 to obtain the transformed model: \[\begin{align} \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\Sigma_{t}^{-\frac{1}{2}}S_{3}S_{1}Y_{t}^{*} &=\gamma \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\Sigma_{t}^{-\frac{1}{2}}S_{3}Y_{t-1}^{(*,-1)} + \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\Sigma_{t}^{-\frac{1}{2}}S_{3}S_{2}Y_{t-1}^{*} + \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\Sigma_{t}^{-\frac{1}{2}}S_{3}X_{t}^{*}\beta \notag \\ &\quad+ \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\Sigma_{t}^{-\frac{1}{2}}V_{t}^{*}, \qquad t=1,\dots,T-1, \label{eq:trans95gmm} \end{align}\tag{8}\] with \(V_{t}^{*}=r_{t}^{*}+E_{t}^{*}\). Furthermore, denote \(\ifstrequal{t}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{t}}) }{ {J}({\Sigma_{t}}) }=\Sigma_{t}^{-\frac{1}{2}} \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\Sigma_{t}^{-\frac{1}{2}},\) and thus, the linear and quadratic moment conditions for equation 8 are now \[\begin{align} \notag &\frac{1}{n(T-1)}m_{N}^{\mathtt{line}}(\theta)=\frac{1}{n(T-1)}\mathbf{Q}_{N}^{\prime}\mathbf{\Sigma}_{N}^{-\frac{1}{2}} \ifstrequal{N}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{N}}) }{ {M}({\Sigma_{N}}) }\mathbf{\Sigma}_{N}^{-\frac{1}{2}}\mathbf{V}_{N}^{*}(\theta)=\frac{1}{n(T-1)}\mathbf{Q}_{N}^{\prime} \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{V}_{N}^{*}(\theta),\\ \label{eq:best95moments} &\frac{1}{n(T-1)}m_{Nk}^{\mathtt{quad}}(\theta)= \frac{1}{n(T-1)}\mathbf{V}_{N}^{*\prime}(\theta) \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{P}_{Nj} \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{V}_{N}^{*}(\theta), \end{align}\tag{9}\] for \(j=1,...,\ell_{p}\), where \(\ifstrequal{N}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{N}}) }{ {M}({\Sigma_{N}}) }=\mathtt{Diag}_{t=1}^{T-1} \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\) and \(\ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }=\mathtt{Diag}_{t=1}^{T-1} \ifstrequal{t}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{t}}) }{ {J}({\Sigma_{t}}) }\). Therefore, we use \(J(\Sigma_{t})\) in BGMM, and \(J_{n}\) in OGMM and 2SLS. 9

After some algebra, for \(t=1,...,T-1,\) we have \[\begin{align} \label{eq:best32Qt} Q_{t}=S_{3}\left[[\tfrac{\partial\mathbf{S}_{1N}}{\partial\lambda_{11}},\cdots,\tfrac{\partial\mathbf{S}_{1N}}{\partial\lambda_{1\ell_{1}}}]S_{1}^{-1}\mathbb{W}_{t}, [\tfrac{\partial\mathbf{S}_{2N}}{\partial\lambda_{21}},\cdots,\tfrac{\partial\mathbf{S}_{2N}}{\partial\lambda_{2\ell_{2}}}]\mathbb{Y}_{t},\mathbb{Y}_{t},X_{t}^{*}\right], \end{align}\tag{10}\] where \(\mathbb{Y}_{t}\) is defined in 23 in the appendix, and \(\mathbb{W}_{t}=(\gamma I_{n}+S_{2})\mathbb{Y}_{t}+X_{t}^{*}\beta+\alpha_{t}^{*}l_{n}\). As for the quadratic moments, \[\begin{align} \label{eq:best32Pt} \begin{aligned} & P_{t}=[P_{1t},\cdots,P_{\ell_{1}t},P_{\ell_{1}+1,t},\cdots,P_{\ell_{3}t}], \;\text{where} \; P_{j_{1},t}=\left[S_{3}\tfrac{\partial S_{1}}{\partial\lambda_{1j_{1}}}S_{1}^{-1}S_{3}^{-1}\Sigma_{t}\right]^{\diamond} \;\text{and} \; \\& P_{\ell_{1}+j_{3},t}=\left[\tfrac{\partial S_{3}}{\partial\lambda_{3j_{3}}}S_{3}^{-1}\Sigma_{t}\right]^{\diamond}, \end{aligned} \end{align}\tag{11}\] for \(j_{1}=1,...,\ell_{1}\) and \(j_{3}=1,...,\ell_{3}\), where ‘\(^{\diamond}\)’ denotes the diagonal-adjustment operator in Lemma 1() in the appendix: For any \(n\times n\) time-varying matrix \(H_t\), it modifies only the diagonal entries so that the transformed matrix has a zero diagonal, i.e., \(\mathtt{diagM}( \ifstrequal{t}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{t}}) }{ {J}({\Sigma_{t}}) }H_t^{\diamond} \ifstrequal{t}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{t}}) }{ {J}({\Sigma_{t}}) })=\mathbf{0}_{n\times 1}\) by construction. Moreover, now \([\mathbf{\Omega}_{N1}]_{ij}=\operatorname{tr}( \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{P}_{Ni} \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{P}_{Nj}^{s} \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) })\) for \(i,j=1,...,\ell_1+\ell_3\), and \(\mathbf{\Omega}_{N2}=\mathbf{Q}_{N}' \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{Q}_{N}\) by using \(\ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{\Sigma}_{N} \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }= \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\). We can then give the algorithm to obtain the proposed estimators:

Algorithm 1: Estimation Procedure


Step 1: Obtain the initial 2SLSE \(\hat{\theta}_{ts}=\min_{\theta\in\Theta}m^{\texttt{line}\prime}_{N}(\theta)\boldsymbol{M}_{Q}m^{\texttt{line}}_{N}(\theta)\) using linear moments \(m^{\texttt{line}}_{N}(\theta)=\mathbf{Q}_{N}'\mathbf{J}_{N}\mathbf{V}_{N}^{*}(\theta)\). Evaluate \(\hat{g}_{k,ts}(d)=\pmb{\phi}_{k}'(d)\hat{\lambda}_{k,ts}\) for \(k=1,2,3\). Compute the residuals \(\tilde{V}_{t}=J_{n}\bigl(V_t(\hat{\theta}_{ts})-\frac{1}{T}\sum_{h=1}^{T}V_h(\hat{\theta}_{ts})\bigr)\) and \(\tilde{\Sigma}_{t}=\Sigma_{t}(\hat{\theta}_{ts})\) based on \(\hat{\theta}_{ts}\).
Step 2: Given \(\tilde{\Sigma}_{t}\), construct \(\tilde{\Omega}_{N}\) based on equation 6 . Obtain the feasible OGMME \(\hat{\theta}_{ogmm}=\arg\min_{\theta\in\Theta}m_{N}'(\theta)\tilde{\Omega}_{N}^{-1}m_{N}(\theta)\) where \(m_{N}(\theta)\) is given in 5 . Evaluate \(\hat{g}_{k,ogmm}(d)=\pmb{\phi}_{k}'(d)\hat{\lambda}_{k,ogmm}\). Update the residuals \(\tilde{V}_{t}=J_{n}\bigl(V_t(\hat{\theta}_{ogmm})-\frac{1}{T}\sum_{h=1}^{T}V_h(\hat{\theta}_{ogmm})\bigr)\), and \(\tilde{\Sigma}_{t}=\Sigma_{t}(\hat{\theta}_{ogmm})\).
Step 3: Given \(\hat{\theta}_{ogmm}\) and \(\tilde{\Sigma}_{t}\), estimate \(\ifstrequal{t}{N}{ \pmb{M}(\tilde{\boldsymbol{\Sigma}}_{t}) }{ M(\tilde{\Sigma}_{t}) }\), \(\QLhat{t}\) and \(\PLhat{t}\) from equations 7 , 10 and 11 , respectively. Estimate \(\ifstrequal{t}{N}{ \pmb{J}({\tilde{\boldsymbol{\Sigma}}_{t}}) }{ {J}({\tilde{\Sigma}_{t}}) }=\tilde{\Sigma}_{t}^{-1/2} \ifstrequal{t}{N}{ \pmb{M}(\tilde{\boldsymbol{\Sigma}}_{t}) }{ M(\tilde{\Sigma}_{t}) }\tilde{\Sigma}_{t}^{-1/2}\). Construct \(\tilde{\Omega}_{N}\) based on equation 6 , with \(\mathbf{J}_{N}\) replaced by \(\ifstrequal{N}{N}{ \pmb{J}({\tilde{\boldsymbol{\Sigma}}_{N}}) }{ {J}({\tilde{\Sigma}_{N}}) }\). Obtain the feasible BGMME \(\hat{\theta}_{bgmm}=\arg\min_{\theta\in\Theta}m_{N}'(\theta)\tilde{\Omega}_{N}^{-1}m_{N}(\theta)\), where \(m_{N}(\theta)\) is defined as 9 . Evaluate \(\hat{g}_{k,bgmm}(d)=\pmb{\phi}_{k}'(d)\hat{\lambda}_{k,bgmm}\).


4 Asymptotic properties↩︎

We collect the full set of regularity assumptions in Section 9 in the appendix. In this section, we highlight two conditions that drive the asymptotics: the joint growth requirement on \((n,T)\) and the sieve dimension requirement on \(\ell_{n}\).

Assumption 1 (Sample size). As \((n,T)\to\infty\), \(n=o\left(T^{p/2}\right)\) for some \(p>2\).

Assumption 2 (Sieve function). () The approximation errors satisfy \(\delta_{k}(d)=O_{p}(\ell_{k}^{-\varsigma_{k}})\), \(k=1,2,3\). Let \(\ell_{n}=\max_{1\leq k\leq 3}\ell_{k}\) denote the sieve dimension, with \(\ell_{n}\to\infty\) as \((n,T)\to\infty\). Moreover, \(\sqrt{n(T-1)}\ell_{n}^{-\underline{\varsigma}}+\sqrt{\frac{\ell_{n}}{n(T-1)}}\to 0\) as \((n,T)\to\infty\), where \(\underline{\varsigma}>2\) with \(\underline{\varsigma}=\min_{1\leq k\leq 3}\varsigma_{k}\).
() The basis functions \(\{\phi_{kl}(\cdot), k=1,2,3,\;l=1,\cdots,\ell_{n}\}\) are uniformly bounded (UB) on the compact domain \(D\), and \(\pmb{\phi}_k(d)=[\phi_{k1}(d),...,\phi_{k\ell_{k}}(d)]\) satisfies \(\max\limits_{1\leq k\leq 3}\sup\limits_{d\in D}\|\pmb{\phi}_{k}(d)\|_{\mathrm{sp}}\leq C_{\ell}\sqrt{\ell_{n}}\) for some constant \(C_{\ell}\).

Assumption 1 allows a larger growth rate of \(n\) compared to \(T\). This is a mild assumption as discussed in [10]. In Assumption 2, () ensures that the sieve approximation error is asymptotically negligible, and, moreover, that the asymptotic correlation between \(\mathbf{V}_{N}^{*}\) and \(\mathbf{X}_{N}^{*}\) vanishes.10 Condition () imposes growth restrictions on the series expansion terms.

4.1 Properties of estimators of the finite-dimensional parameter↩︎

For \(k=1,2,3\), write \(\hat{\pi}_{gmm}\) for either \(\hat{\pi}_{ogmm}\) or \(\hat{\pi}_{bgmm}\), and the true estimand is \(\pi_0\). Then, we have the following theorem.

Theorem 1. Suppose that \(\tilde{\Omega}_{N}\) is evaluated by the initial consistent estimator \(\tilde{\theta}\) based on equation 6 :
() Under Assumptions 1-2, and 3-7 in the appendix, when \(\ell_{n}^{3/2-\underline{\varsigma}}+\frac{\ell_{n}^{3/2}}{n}\to 0\) as \((n,T)\to\infty\): \[\hat{\pi}_{gmm}-\pi_0=o_{p}(1).\] () Under Assumptions 1-2, and 3-7 in the appendix, when \(\sqrt{\frac{(T-1)\ell_{n}^{3}}{n}}+\sqrt{n(T-1)}\ell_{n}^{3/2-\underline{\varsigma}}\to 0\) as \((n,T)\to\infty\): \[\sqrt{n(T-1)}(\hat{\pi}_{gmm}-\pi_0)\stackrel{d}{ \to}N(0,\boldsymbol{\Sigma}_{\pi_0,gmm}),\] where \(\mathbf{\Sigma}_{\pi_0,gmm}=\lim_{n,T\to\infty}K_{\pi}\Omega_{N}K_{\pi}'=O(1)\) and \(K_{\pi}\) is defined in 33 in the appendix.

The rate condition in Theorem 1(ii) implies the one in (i). Indeed, \(\sqrt{n(T-1)}\ell_{n}^{3/2-\underline{\varsigma}} \to 0\Rightarrow \ell_{n}^{3/2-\underline{\varsigma}} \to 0,\) and \(\sqrt{\frac{(T-1)\ell_{n}^{3}}{n}}=\frac{\ell_{n}^{3/2}}{n} \sqrt{n(T-1)}\to 0\Rightarrow \frac{\ell_{n}^{3/2}}{n}\to 0\).

We next establish the consistency of the covariance estimator \(\boldsymbol{\Sigma}_{\pi_0,gmm}\) and characterize the asymptotic efficiency of the BGMME.

Theorem 2. Define \(\hat{\boldsymbol{\Sigma}}_{\pi,gmm}= \hat{K}_{\pi}\hat{\Omega}_{N}\hat{K}_{\pi}^{\prime}\), where \(\hat{K}_{\pi}\) and \(\hat{\Omega}_{N}\) are the feasible counterparts of \(K_\pi\) and \(\Omega_{N}\) evaluated at \(\hat{\pi}_{gmm}\). Under the conditions of Theorem 1(), we have:
() \(\|\hat{\boldsymbol{\Sigma}}_{\pi,gmm}-\boldsymbol{\Sigma}_{\pi_0,gmm}\|_{\mathrm{sp}}=o_{p}(1).\)
() The asymptotic variance of the BGMME attains the efficiency lower bound \(\mathbf{\Sigma}_{b_\pi}\) by the generalized Cauchy-Schwarz inequality on \(\lim_{n,T\to\infty}K_{\pi}\Omega_{N}K_{\pi}\), where \(\mathbf{\Sigma}_{b_\pi}\) is defined in equation 38 in the appendix. Accordingly, it can be consistently estimated by \(\hat{\mathbf{\Sigma}}_{b_\pi}\).

4.2 Properties of estimators of the infinite-dimensional parameter↩︎

Having established the asymptotic properties of the finite-dimensional parameters, we now turn to the inference on the unknown spatial weight functions \(g_{k}(\cdot)\). The following theorem establishes the uniform consistency and asymptotic normality of the sieve estimators \(\hat{g}_{k,gmm}(\cdot)\).

Theorem 3. Let \(\hat{g}_{k,gmm}\) be either \(\hat{g}_{k,gmm}\) or \(\hat{g}_{k,bgmm}\), and let the true \(g_{k}\) be \(g_{k0}\), \(k=1,2,3\). Let \(\hat{g}_{k,gmm}(d)=\pmb{\phi}_{k}'(d)\hat{\lambda}_{k,gmm}\). Then:
() Under Assumptions 1-2, and 3-7 in the appendix, when \(\ell_{n}^{2-\underline{\varsigma}}+\frac{\ell_{n}^{2}}{n}+\frac{\ell_{n}}{\sqrt{n(T-1)}}\to 0\) as \((n,T)\to\infty\): \[\sup_{d\in D}|\hat{g}_{k,gmm}(d)-g_{k0}(d)|=o_{p}(1).\] () Under Assumptions 1-2, and 3-7 in the appendix, when \(\sqrt{\frac{(T-1)\ell_{n}^{3}}{n}}+\sqrt{n(T-1)}\ell_{n}^{3/2-\underline{\varsigma}}\to 0\) as \((n,T)\to\infty\): \[\sqrt{\frac{n(T-1)}{\ell_{n}}}(\hat{g}_{k,gmm}(d)-g_{k0}(d))\stackrel{d}{\to}N(0,\mathbf{\Sigma}_{g_{k0},gmm}),\] where \(\mathbf{\Sigma}_{g_{k0},gmm}=O(1)\) is defined in equation 41 in the appendix.

The asymptotic results in Theorem 3 highlight two key distinctions between the nonparametric sieve estimator and the parametric case in Theorem 1. First, the rate conditions required to establish uniform consistency are slightly stronger than those for \(\hat{\pi}\). Specifically, the condition involves terms of order \(n^{-1}\ell_{n}^{2}\) rather than \(n^{-1}\ell_{n}^{3/2}\). This stricter requirement reflects the difficulty of achieving uniform convergence of the function \(\hat{g}_k(d)\) over the domain \(d\in D\), as opposed to the standard convergence of a finite parameter vector. Second, the convergence rate for asymptotic normality is slower. While the finite-dimensional estimators achieve the standard \(\sqrt{n(T-1)}\) rate, the nonparametric estimators converge at the slower rate \(\sqrt{n(T-1)/\ell_{n}}\). 11

Remark 3. () While our model assumes \((n,T) \to \infty\), the proposed estimators are also applicable in panels with large \(n\) and finite \(T\). Specifically, the 2SLSE and OGMME remain consistent under large \(n\) and fixed \(T\) for the variance structures specified in [var:homo] and [var:hetet]. However, the implementation of the BGMME requires \((n,T)\to\infty\). This is because the individual FE, which are required to construct \(\mathbb{Y}_{t}\) as defined in 23 in the appendix, cannot be consistently estimated with finite \(T\), and thus the best IVs might not be available (see [4] footnote 24).

5 Monte Carlo simulations↩︎

We conduct Monte Carlo experiments to evaluate the finite sample performance of our proposed estimators. To conserve space, we report the key baseline results for the MESS specification in the main text, while the results for the SAR specifications and additional results are provided in the supplement.

For the DGP in equation 1 , \(X_{t}\), \(\mathbf{c}_{n}\), \(\alpha_{t}\), and \(E_{t}\) are generated from independent standard normal distributions. We consider \(\pi_0=(\gamma_0,\beta_0)'=(-0.7,1)'\) for the MESS and \(\pi_0=(0.3,1)'\) for the SAR. We set \(n\in\{100, 200, 400\}\) and \(T\in\{10, 25\}\), generate the spatial panel data with \(500+T\) periods, and then use the last \(T\) periods as our sample. The initial value is generated as \(N(0, I_n)\), and we consider the case of [var:hetei] with 1,000 replications.

To generate the unknown \(G_{k}\), we randomly generate the distances \(d_{ij}\) between individuals and set \(G_{k}=G\), constructed as \(g^{*}(d_{ij})=\texttt{Normal}(-d_{ij})\) if \(i \neq j\), and \(g_{ii}=0\), where \(\texttt{Normal}(\cdot)\) is the standard normal cdf. Here, \(d_{ij}\) is the Euclidean distance between the coordinates of units \(i\) and \(j\), and we set \(d_{ij}=0\) if \(d_{ij}\) exceeds a threshold \(\bar{d}_{0}\), where \(\bar{d}_{0}\) is the 10th percentile of all \(d_{ij}\). The two-dimensional coordinates are generated from a uniform distribution \(U[0,1]\). Then, \(G=G^{*}/(1.2\|G^{*}\|_{\mathrm{sp}})\). From the construction, \(\rho(A)=0.6949\). In estimation, we use the row-normalization version of \(\varPhi_{kl}=\varPhi_{l}=\{\phi_{l}(d_{ij})\}\) for \(l=1,...,\ell_{n}\). We consider two different choices for the sieve dimension: \(\ell_{k}=\ell_{n}=2\), and \(\ell_{k}=\ell_{n}=[n^{1/5}]+2\), a choice motivated by the theoretical rate conditions derived in Section 4 and Footnote 7.

To generate the heteroskedastic variances \(\sigma_{\epsilon_i}^2\), we proceed in three steps. First, partition the \(n\) individuals into \(K\) groups and assign initial group labels \(k_i\in\{1,\dots,K\}\) accordingly; then randomly permute these labels to obtain \(\tilde{k}_i\), so that group sizes remain fixed but membership is randomized. Second, define \(u_h = u(h)=(h-1)/(K-1), h=1,\dots,K,\) and interpret \(u_{\tilde{h}_i}=u(\tilde{h}_i)\) as the value associated with the randomized group of observation \(i\). Finally, set \(\sigma_{\epsilon_i}^2 = \bigl(1 + {u_{\tilde{h}_i}}/{K}\bigr)^2\). In our design, we set \(K=3\). Thus, \(u=\{0,{1}/{2},1\}\).

For the estimation of \(\pi\), Table ¿tbl:sim:hei? shows the following: First, as \(n\), \(T\), and \(\ell_{n}\) increase, both OGMME and BGMME exhibit lower root mean squared error (RMSE) and smaller bias compared with the 2SLSE. This indicates that the GMME can improve the finite sample performance of the 2SLSE, especially when \(n\), \(T\), and \(\ell_{n}\) are large. Second, within the GMME framework, BGMME is more efficient than OGMME when \(n\) and \(T\) increase and \(\ell_{n}/(nT)\to 0\), as its empirical standard deviation (ESD) and RMSE are smaller. However, this efficiency gain is less pronounced when \(T\) and \(\ell_{n}\) are both small, which is consistent with the theoretical rate conditions. Third, as \((n, T)\) increases, the coverage probability (CP) approaches the nominal 0.95 level, suggesting that the variance approximation used for inference is reliable.

The results regarding \(\hat{G}_k\) and \(\rho(\hat{A})\) are as follows. First, as shown in Table ¿tbl:sim:rhoG95mess?, the estimates for \(\rho(\hat{A})\) are consistently less than 1 under both the baseline specification with \(\bar{d}_0=10\%\) in Panel A and the alternative threshold with \(\bar{d}_0=15\%\) in Panel B. 12 The approximation of \(\rho(\hat{A})\) improves with increasing \((n,T)\): Both bias and RMSE decline; even with \(\ell_{n}=2\), larger \(n\) and \(T\) reduce bias, while \(\ell_{n}=[n^{1/5}]+2\) produces the best performance in finite samples. This confirms that the estimated spatial dynamic system remains stable, satisfying Assumption 4 in the appendix regardless of the chosen distance cutoff. Second, from Table ¿tbl:sim:heteG?, the bias and RMSE of the components in \(\hat{\tilde{g}}_{k}\) decrease as \((n,T,\ell_{n})\) increases. Analogous to \(\pi\), BGMME yields the smallest bias and RMSE in most settings, followed by OGMME and then 2SLSE. However, this advantage may disappear when \(T\) and \(\ell_{n}\) are small. In addition, the MAE also declines as \((n,T,\ell_{n})\) increase, with the BGMME consistently yielding the smallest MAE.

Moreover, we assess robustness to functional-form misspecification by simulating data from a SAR DGP and estimating a MESS model, and vice versa. To facilitate comparability across parameterizations, Remark 1 implies a mapping between the two spatial parameters: under a SAR DGP, \(\gamma_0=0.3\) corresponds to \(\gamma_0-1\) under the MESS parameterization; under a MESS DGP, \(\gamma_0=-0.7\) corresponds to \(\gamma_0+1\) under the SAR parameterization. We set \(\beta_0=1\), keep all other design features unchanged, and maintain the stability condition \(\rho(\hat{A})<1\) in estimation. Table ¿tbl:sim:misspecification? shows that under misspecification, although the finite-dimensional \(\pi\) has smaller biases, the CPs fall below the nominal 0.95 level and deteriorate as \((n,T)\) increases.

0.9

Finite sample performance of estimators of \(\pi\) for the MESS, \(\Var(\epsilon_{it})=\sigma_i^2\).
2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\)
2-7 \((n,T,\elln)=(100,10,[n^{1/5}]+2)\)
2-7 Bias -0.0237 0.0045 -0.0255 0.0001 -0.0278 -0.0032 -0.0229 0.0042 -0.0252 0.0005 -0.0277 0.0015
ESD 0.0381 0.0344 0.0380 0.0321 0.0412 0.0349 0.0360 0.0344 0.0366 0.0317 0.0396 0.0340
RMSE 0.0449 0.0347 0.0457 0.0321 0.0497 0.0350 0.0427 0.0347 0.0444 0.0317 0.0483 0.0340
CP 0.7970 0.9070 0.7950 0.9280 0.7850 0.9250 0.8930 0.9030 0.8830 0.9100 0.8720 0.8990
2-7 \((n,T,\elln)=(100,25,[n^{1/5}]+2)\)
2-7 Bias -0.0151 0.0023 -0.0178 -0.0004 -0.0218 -0.0072 -0.0149 0.0014 -0.0170 -0.0007 -0.0180 0.0020
ESD 0.0248 0.0222 0.0247 0.0221 0.0268 0.0246 0.0242 0.0216 0.0245 0.0212 0.0266 0.0235
RMSE 0.0290 0.0223 0.0304 0.0221 0.0345 0.0256 0.0284 0.0216 0.0298 0.0212 0.0321 0.0236
CP 0.8960 0.9070 0.8840 0.9270 0.8730 0.9160 0.8980 0.9140 0.8990 0.9180 0.8880 0.9070
2-7 \((n,T,\elln)=(200,10,[n^{1/5}]+2)\)
2-7 Bias -0.0201 0.0055 -0.0230 0.0038 -0.0240 0.0003 -0.0170 0.0041 -0.0200 0.0028 -0.0198 0.0038
ESD 0.0193 0.0199 0.0193 0.0183 0.0191 0.0182 0.0193 0.0195 0.0186 0.0187 0.0184 0.0186
RMSE 0.0279 0.0206 0.0301 0.0187 0.0307 0.0182 0.0257 0.0199 0.0273 0.0189 0.0271 0.0189
CP 0.8000 0.8990 0.8060 0.9260 0.7960 0.9150 0.8430 0.8880 0.8140 0.9200 0.8040 0.9090
2-7 \((n,T,\elln)=(200,25,[n^{1/5}]+2)\)
2-7 Bias 0.0069 0.0046 0.0069 0.0046 0.0070 0.0045 -0.0066 0.0036 0.0072 0.0025 -0.0081 0.0035
ESD 0.0142 0.0131 0.0136 0.0128 0.0135 0.0126 0.0136 0.0129 0.0135 0.0124 0.0132 0.0122
RMSE 0.0158 0.0138 0.0153 0.0136 0.0152 0.0134 0.0151 0.0134 0.0153 0.0127 0.0155 0.0127
CP 0.9260 0.9400 0.9360 0.9330 0.9250 0.9220 0.9400 0.9350 0.9420 0.9420 0.9350 0.9400
2-7 \((n,T,\elln)=(400,10,[n^{1/5}]+2)\)
2-7 Bias -0.0105 0.0065 -0.0075 0.0050 -0.0074 -0.0047 -0.0101 0.0053 -0.0070 0.0043 -0.0063 0.0045
ESD 0.0169 0.0146 0.0169 0.0146 0.0167 0.0145 0.0124 0.0116 0.0125 0.0112 0.0121 0.0112
RMSE 0.0199 0.0160 0.0185 0.0154 0.0182 0.0152 0.0160 0.0127 0.0143 0.0120 0.0136 0.0120
CP 0.9120 0.9220 0.9270 0.9300 0.9190 0.9320 0.9130 0.9300 0.9320 0.9360 0.9240 0.9330
2-7 \((n,T,\elln)=(400,25,[n^{1/5}]+2)\)
2-7 Bias -0.0053 0.0016 -0.0033 -0.0005 -0.0036 -0.0044 -0.0046 0.0012 -0.0026 -0.0009 -0.0029 -0.0032
ESD 0.0094 0.0086 0.0094 0.0086 0.0093 0.0086 0.0090 0.0085 0.0090 0.0086 0.0089 0.0086
RMSE 0.0108 0.0088 0.0100 0.0086 0.0100 0.0097 0.0101 0.0086 0.0094 0.0087 0.0094 0.0092
CP 0.9250 0.9360 0.9320 0.9380 0.9250 0.9360 0.9340 0.9420 0.9400 0.9450 0.9410 0.9420

Note: The true parameters are set to \(\pi_0=(-0.7,1)'\), and \(\bard_0=10\%\). The results are based on 1,000 Monte Carlo replications. Bias denotes the mean bias of the estimates, ESD denotes the standard deviation, RMSE denotes the root mean squared error, and CP denotes the 95% coverage probability. The 2SLS refers to the two-stage least square estimator, OGMM refers to the feasible optimal GMM estimator, and BGMM refers to the feasible best GMM estimator.

0.9

0.5

Finite sample performance of estimators of \(G_{1}, G_{2}\) and \(G_{3}\) for the MESS. \(\Var(\epsilon_{it})=\sigma_i^2\).
\((n,T,\elln)=(100,10,2)\) \((n,T,\elln)=(100,10,[n^{1/5}]+2)\)
2-10 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_1\) \(\tilde{g}_2\) \(\tilde{g}_3\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
2-10 MAE 0.0346 0.0396 0.0733 0.0342 0.0393 0.0712 0.0342 0.0394 0.0693 0.0344 0.0401 0.0728 0.0340 0.0397 0.0697 0.0343 0.0395 0.0679
Bias -0.0221 -0.0326 -0.0664 -0.0206 -0.0321 -0.0648 -0.0189 -0.0315 -0.0630 -0.0219 -0.0328 -0.0661 -0.0196 -0.0320 -0.0635 -0.0182 -0.0311 -0.0618
RMSE 0.0266 0.0364 0.0774 0.0254 0.0359 0.0749 0.0246 0.0358 0.0727 0.0269 0.0374 0.0771 0.0253 0.0368 0.0734 0.0250 0.0362 0.0711
2-10 \((n,T,\elln)=(100,25,2)\) \((n,T,\elln)=(100,25,[n^{1/5}]+2)\)
2-10 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
2-10 MAE 0.0160 0.0189 0.0456 0.0159 0.0188 0.0445 0.0160 0.0189 0.0433 0.0162 0.0190 0.0483 0.0162 0.0189 0.0467 0.0167 0.0189 0.0460
Bias -0.0089 -0.0154 -0.0433 -0.0080 -0.0152 -0.0424 -0.0069 -0.0150 -0.0413 -0.0094 -0.0152 -0.0437 -0.0079 -0.0150 -0.0419 -0.0080 -0.0147 -0.0417
RMSE 0.0115 0.0177 0.0490 0.0108 0.0176 0.0476 0.0106 0.0175 0.0462 0.0111 0.0171 0.0482 0.0098 0.0170 0.0463 0.0101 0.0169 0.0458
2-10 \((n,T,\elln)=(200,10,2)\) \((n,T,\elln)=(200,10,[n^{1/5}]+2)\)
2-10 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_1\) \(\tilde{g}_2\) \(\tilde{g}_3\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
2-10 MAE 0.0333 0.0377 0.0417 0.0326 0.0376 0.0491 0.0321 0.0381 0.0466 0.0327 0.0368 0.0393 0.0319 0.0366 0.0359 0.0320 0.0363 0.0338
Bias -0.0233 -0.0325 -0.0397 -0.0208 -0.0324 -0.0371 -0.0180 -0.0328 -0.0342 -0.0228 -0.0312 -0.0371 -0.0198 -0.0308 -0.0305 -0.0177 -0.0298 -0.0312
RMSE 0.0248 0.0337 0.0461 0.0225 0.0337 0.0432 0.0203 0.0343 0.0401 0.0247 0.0327 0.0442 0.0220 0.0324 0.0378 0.0206 0.0316 0.0376
2-10 \((n,T,\elln)=(200,25,2)\) \((n,T,\elln)=(200,25,[n^{1/5}]+2)\)
2-10 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
2-10 MAE 0.0074 0.0089 0.0150 0.0074 0.0089 0.0148 0.0076 0.0089 0.0143 0.0073 0.0085 0.0149 0.0073 0.0085 0.0145 0.0072 0.0084 0.0139
Bias -0.0038 -0.0065 -0.0104 -0.0038 -0.0065 -0.0104 -0.0031 -0.0061 -0.0103 -0.0037 -0.0063 -0.0102 -0.0037 -0.0063 -0.0102 -0.0031 -0.0060 -0.0101
RMSE 0.0056 0.0089 0.0134 0.0056 0.0089 0.0134 0.0052 0.0086 0.0133 0.0052 0.0083 0.0130 0.0052 0.0084 0.0130 0.0047 0.0080 0.0129
2-10 \((n,T,\elln)=(400,10,2)\) \((n,T,\elln)=(400,10,[n^{1/5}]+2)\)
2-10 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_1\) \(\tilde{g}_2\) \(\tilde{g}_3\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
2-10 MAE 0.0086 0.0103 0.0174 0.0086 0.0103 0.0172 0.0088 0.0103 0.0166 0.0085 0.0099 0.0173 0.0085 0.0099 0.0168 0.0084 0.0097 0.0161
Bias -0.0044 -0.0075 -0.0121 -0.0044 -0.0075 -0.0121 -0.0036 -0.0071 -0.0119 -0.0043 -0.0073 -0.0118 -0.0043 -0.0073 -0.0118 -0.0036 -0.0070 -0.0117
RMSE 0.0057 0.0081 0.0128 0.0057 0.0081 0.0128 0.0058 0.0080 0.0127 0.0052 0.0077 0.0125 0.0052 0.0078 0.0125 0.0048 0.0075 0.0123
2-10 \((n,T,\elln)=(400,25,2)\) \((n,T,\elln)=(400,25,[n^{1/5}]+2)\)
2-10 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
2-10 MAE 0.0072 0.0085 0.0128 0.0072 0.0084 0.0126 0.0072 0.0085 0.0122 0.0070 0.0081 0.0124 0.0070 0.0081 0.0122 0.0069 0.0084 0.0118
Bias -0.0035 -0.0060 -0.0088 -0.0035 -0.0060 -0.0088 -0.0029 -0.0056 -0.0088 -0.0034 -0.0058 -0.0086 -0.0034 -0.0058 -0.0086 -0.0029 -0.0055 -0.0085
RMSE 0.0041 0.0065 0.0103 0.0041 0.0065 0.0103 0.0038 0.0063 0.0103 0.0038 0.0062 0.0099 0.0038 0.0062 0.0099 0.0034 0.0060 0.0098

Note: The \(\tilde{g}_{k}\) extracts the column vectors composed of non-zero elements from the upper triangular submatrix of \(G_{k}\), and \(\bard_0=10\%\). The results are based on 1,000 Monte Carlo replications. MAE denotes the mean absolute error. Bias denotes the mean bias of the estimates, and RMSE denotes the root mean squared error. The 2SLS refers to the two-stage least square estimator, OGMM refers to the feasible optimal GMM estimator, and BGMM refers to the feasible best GMM estimator.

1

Finite sample performance of estimators of \(\pi\) under misspecification, \(\Var(\epsilon_{it})=\sigma_i^2\).
DGP: SAR, Estimation: MESS
2-7 [6]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\)
2-7 \((n,T,\elln)=(100,10,2)\) \((n,T,\elln)=(100,10,[n^{1/5}]+2)\)
2-7 Bias 0.0037 -0.0024 0.0037 0.0012 0.0032 0.0012 -0.0044 0.0002 -0.0034 0.0002 -0.0034 -0.0005
RMSE 0.0442 0.0361 0.0400 0.0329 0.0399 0.0329 0.0430 0.0378 0.0400 0.0344 0.0400 0.0344
CP 0.8990 0.9000 0.8990 0.8970 0.7940 0.8610 0.9210 0.8850 0.9190 0.8820 0.7950 0.8410
2-7 \((n,T,\elln)=(100,25,2)\) \((n,T,\elln)=(100,25,[n^{1/5}]+2)\)
2-7 Bias 0.0040 0.0024 0.0029 0.0024 0.0029 -0.0020 -0.0017 0.0010 -0.0003 0.0010 -0.0003 -0.0001
RMSE 0.0340 0.0277 0.0280 0.0224 0.0280 0.0224 0.0306 0.0269 0.0265 0.0222 0.0265 0.0222
CP 0.9120 0.9100 0.9120 0.9070 0.7310 0.8230 0.9220 0.9220 0.9230 0.9220 0.7620 0.8350
2-7 \((n,T,\elln)=(200,10,2)\) \((n,T,\elln)=(200,10,[n^{1/5}]+2)\)
2-7 Bias 0.0167 0.0033 0.0163 0.0033 0.0163 -0.0008 0.0153 0.0034 0.0143 0.0034 0.0143 0.0020
RMSE 0.0332 0.0235 0.0282 0.0207 0.0282 0.0204 0.0324 0.0247 0.0275 0.0207 0.0275 0.0205
CP 0.7810 0.8880 0.7810 0.8860 0.5800 0.8340 0.7910 0.8880 0.7920 0.8880 0.6180 0.8100
2-7 \((n,T,\elln)=(200,25,2)\) \((n,T,\elln)=(200,25,[n^{1/5}]+2)\)
2-7 Bias 0.0116 0.0032 0.0116 0.0032 0.0103 0.0022 0.0098 0.0026 0.0095 0.0019 0.0095 0.0019
RMSE 0.0235 0.0153 0.0196 0.0133 0.0189 0.0131 0.0227 0.0185 0.0187 0.0131 0.0187 0.0131
CP 0.7940 0.9220 0.7960 0.9220 0.5930 0.8840 0.7980 0.9120 0.8010 0.9120 0.6050 0.7750
2-7 DGP: MESS; Estimation: SAR
2-7 [6]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\)
2-7 \((n,T,\elln)=(100,10,2)\) \((n,T,\elln)=(100,10,[n^{1/5}]+2)\)
2-7 Bias -0.0250 -0.0102 -0.0164 -0.0099 -0.0164 -0.0099 -0.0247 -0.0182 -0.0132 -0.0182 -0.0132 -0.0186
RMSE 0.0497 0.0310 0.0440 0.0308 0.0419 0.0308 0.0457 0.0384 0.0382 0.0382 0.0349 0.0384
CP 0.8880 0.9040 0.9060 0.9050 0.8070 0.8980 0.8590 0.9180 0.8940 0.9080 0.7540 0.8960
2-7 \((n,T,\elln)=(100,25,2)\) \((n,T,\elln)=(100,25,[n^{1/5}]+2)\)
2-7 Bias -0.0101 -0.0121 -0.0049 -0.0120 -0.0049 -0.0111 -0.0086 -0.0089 -0.0023 -0.0116 -0.0023 -0.0116
RMSE 0.0270 0.0243 0.0255 0.0242 0.0240 0.0238 0.0246 0.0213 0.0231 0.0226 0.0224 0.0223
CP 0.9100 0.8930 0.9310 0.8710 0.8320 0.8690 0.9120 0.8590 0.9270 0.8070 0.8070 0.7980
2-7 \((n,T,\elln)=(200,10,2)\) \((n,T,\elln)=(200,10,[n^{1/5}]+2)\)
2-7 Bias -0.0220 -0.0077 -0.0185 -0.0077 -0.0185 -0.0076 -0.0172 -0.0116 -0.0131 -0.0131 -0.0131 -0.0131
RMSE 0.0293 0.0202 0.0268 0.0191 0.0263 0.0191 0.0259 0.0204 0.0231 0.0213 0.0231 0.0211
CP 0.8220 0.9240 0.8600 0.9060 0.7070 0.9040 0.8410 0.9100 0.8660 0.8710 0.7490 0.8680
2-7 \((n,T,\elln)=(200,25,2)\) \((n,T,\elln)=(200,25,[n^{1/5}]+2)\)
2-7 Bias -0.0080 -0.0099 -0.0052 -0.0099 -0.0052 -0.0077 -0.0044 -0.0110 -0.0015 -0.0146 -0.0015 -0.0146
RMSE 0.0151 0.0163 0.0138 0.0163 0.0130 0.0149 0.0149 0.0164 0.0143 0.0190 0.0132 0.0186
CP 0.8040 0.8610 0.8110 0.8050 0.8290 0.8040 0.8290 0.8210 0.8310 0.7080 0.8390 0.7060

Note: The results are based on 1,000 Monte Carlo replications. Bias denotes the mean bias of the estimates, RMSE denotes the root mean squared error, and CP denotes the 95% coverage probability. The 2SLS refers to the two-stage least square estimator, OGMM refers to the feasible optimal GMM estimator, and BGMM refers to the feasible best GMM estimator.

6 An empirical application↩︎

[1] empirically finds that exogenous extreme rainfall causes lower income in rural Tanzania, and leads to a large increase in the murder of ‘witches’, who are nearly all elderly women. [1] further separates this income shock hypothesis from non-economic factors, namely the scapegoat culture hypothesis that families under exogenous extreme rainfall or disease epidemic shocks treat witches as scapegoats and increase their vilification and killing. The results show that income shocks are the driver of the increase in witch killings, instead of this scapegoat culture. Based on [1], we apply our proposed model to investigate the mechanism of spatial spillover effects regarding witch killings in rural Tanzania. This analysis involves a dataset from 67 villages over the period from 1992 to 2002. We plot the specific location of these villages in Figure [emp:fig95ploty] and it illustrates clear clustering within these communities.

The clustering pattern of the data motivates our study, in which we conjecture that geographic closeness leads to parallel behavioral patterns in witch killings. For example, when a village experiences severe weather or unexpected illnesses, it could lead to a rise in witch killings, an action that nearby communities might mimic. While it is natural to construct analyses using matrices such as the \(k\)-nearest neighbor matrix or inverse distance geographical matrix, these are limited by their specific forms. Therefore, we consider unknown weights \(G_{1}, G_{2}, G_{3}\), which are functions of the geographical distance \(d_{ij}\). We then approximate using a method similar to that in Section 5, where \(\ell_{k}=[n^{1/5}]+2\), \(\bar{d}_{0}\) is set to the 10%, 25%, and 50% quantiles of \(d_{ij}\), respectively. Specifically, our empirical models are \[\begin{align} \tag{12} B_{1}Y_{t}&=(\gamma I_{n}+B_2)Y_{t-1}+X_{1}\beta_{1}+FEs+U_{t},\\ \tag{13} B_{1}Y_{t}&=(\gamma I_{n}+B_2)Y_{t-1}+X_{1t}\beta_{1}+X_{2t}\beta_{2}+X_{3t}\beta_{3}+FEs+U_{t},\\ \tag{14} B_{1}Y_{t}&=(\gamma I_{n}+B_2)Y_{t-1}+X_{1t}\beta_{1}+X_{2t}\beta_{2}+X_{3t}\beta_{3}+X_{4t}\beta_{4}+FEs+U_{t}, \end{align}\] with \(B_{3}U_{t}=E_{t}\), where the description of variables is shown in Table 1.

In Table ¿tbl:emp:hetei?, we report the baseline regression results using the number of witch murders as the primary dependent variable under the case of [var:hetei]. To ensure the reliability of our estimates, we consider both the MESS and SAR specifications. The robustness checks for the MESS are presented in Table ¿tbl:emp:robust?, and the corresponding estimation results for the SAR are presented in Table ¿tbl:emp:hetei-sar?. These results are qualitatively similar to our baseline findings, confirming the robustness of our framework.13

Our baseline results reveal that as the geographic distance between communities decreases, the absolute value of estimated spatial weights increases, as shown in Figure 2. This implies that neighboring communities exert a stronger spillover effect on one another’s witch killing rates when they are located closer together, indicating stronger community spillovers in crimes. We further examine the mechanism of witch killing spillovers between communities from the economic and cultural perspectives, respectively. 14

We evaluate economic distance by measuring the similarity in annual per capita consumption expenditures, defined as \(ce_{ij}\), using the 2001 survey data. Specifically, we define the distance as \(d_{ij}=d_{ij}^{*}ce_{ij}\), where \(ce_{ij}=1+|ce_{i,2001}-ce_{j,2001}|,\) and \(d_{ij}^{*}\) is the geographical distance mentioned earlier, denoted with a superscript asterisk. We find that as the geographic distance between communities decreases, and if the similarity in annual per capita consumption expenditures increases, which indicates a smaller economic difference, the absolute value of estimated spatial weights rises, as shown in Figure 3. The mean coefficient of \(\hat{G}_{1}\) in this specification ranges from \(-0.2\) to \(-0.25\) and is significantly larger in magnitude than the baseline result of around \(-0.08\) to \(-0.1\). This finding highlights that economic similarity is a key driver of the copycat pattern in witch killings.

We then evaluate cultural distance using ethnic shares.15 Specifically, we define the cultural distance by the similarity in Sukuma ethnic shares, defined as \(st_{ij}\), over all \(t\), where \(st_{ij}=1+|st_{i}-st_{j}|\). We explicitly focus on the Sukuma population because established literature demonstrates their unique relevance to this phenomenon. Specifically, [1] notes that, according to [38], the ethnically Sukuma regions of western Tanzania historically account for roughly two-thirds of all reported witch killings in the country, even though the Sukuma ethnic group comprises only 12% of the national population. We find that as the geographic distance decreases, if the similarity in Sukuma ethnic shares increases, the absolute value of estimated spatial weights rises. However, the magnitude of the mean coefficient in this specification (around \(-0.06\) to \(-0.07\)) is not statistically larger than the magnitude of the baseline result. In other words, cultural proximity does not significantly amplify spatial dependence in witch killings. These results corroborate [1]’s conclusions on how economic and cultural factors influence crime rates, extending the framework by modeling the spatial spillovers in witch killings.

Specifically, Figure 3 illustrates that the estimated spatial weighting functions gradually vanish as the geographical economic distance increases. We note that the estimated elements in \(\hat{G}_{1}\) are predominantly negative. Since \(e^{-\hat{G}_1}=I_n+\sum_{k=1}^{\infty}{(-1)^k\hat{G}_{1}^k}/{k!}\), negative entries in \(\hat{G}_1\) result in positive off-diagonal elements in the inverse spatial filter, indicating a positive reinforcement mechanism between neighbors. We observe that if the entries of \(\hat{G}_1\) are negative, the linear term \(-\hat{G}_1\) becomes positive. This implies that the spatial multiplier \(B_1^{-1}\) contains positive off-diagonal elements, indicating a positive reinforcement mechanism between neighbors. Thus, an increase in witch killings in one village spills over to increase killings in neighboring villages, consistent with the ‘copycat’ hypothesis, even though the underlying weight function \(g_1(d)\) takes negative values. This contrasts with standard specifications that typically impose non-negative weights a priori, a flexibility emphasized by [18].

Moreover, we confirm that the stability condition \(\rho(\hat{A})<1\) holds, as shown in Table ¿tbl:emp:rhoA?, and proceed to examine the marginal effects, adopting the methodology of [39] and [40]. The short-term and long-term marginal effects with respect to the \(k\)-th independent variable are evaluated as: \[\begin{align} f_{Si}(\beta)=\mathtt{e}_{n,i}\mathtt{Diag}(S^{-1})\mathtt{e}_{n,i}'\beta \quad \text{and} \quad f_{Li}(\beta)=\mathtt{e}_{n,i}\mathtt{Diag}((S-B)^{-1})\mathtt{e}_{n,i}'\beta, \end{align}\] respectively, for \(i=1,\ldots,n\), where \(\mathtt{e}_{n,i}=(0,\ldots,\underbrace{1}_{\text{ith}},\ldots,0)'\).

Figure 5 presents the results for the case where \(\bar{d}_{0}=25\%\). The figure reveals that focusing solely on short-term marginal effects underestimates the true magnitude of the coefficients \(\beta_{1}\) and \(\beta_{4}\). The short-term estimates capture only 50% to 95% of the total long-term impact, with the average long-term effect being significantly larger. This discrepancy suggests that spatial spillovers driven by geographic-economic proximity amplify the initial shocks over time, resulting in persistent and long-lasting dynamics.

Figure 1: Scatter plot of sample villages in Tanzania

0.7

Estimation results for the MESS, \(\Var(\epsilon_{it})=\sigma_i^2\).
2SLS OGMM BGMM 2SLS OGMM BGMM 2SLS OGMM BGMM
3-5 \(\bar{d}_{0}=10\%\) \(\bar{d}_{0}=25\%\) \(\bar{d}_{0}=50\%\)
3-5 \(\hat{\beta}_{1}\) 0.0640\(^{**}\) 0.0540\(^{*}\) 0.0740\(^{**}\) 0.0540 0.0240 0.0340 0.0360 0.0660\(^{*}\) 0.0860\(^{**}\)
(0.0314) (0.0292) (0.0323) (0.0334) (0.0329) (0.0463) (0.0361) (0.0335) (0.0388)
3-5 \(\hat{\beta}_{1}\) 0.0650\(^{*}\) 0.0510\(^{*}\) 0.0710\(^{**}\) 0.0590 0.0890\(^{**}\) 0.0790\(^{*}\) 0.0400 0.0700\(^{*}\) 0.0900\(^{*}\)
(0.0382) (0.0310) (0.0357) (0.0409) (0.0425) (0.0419) (0.0424) (0.0405) (0.0463)
\(\hat{\beta}_{2}\) 0.0490 0.0520 0.0720\(^{*}\) 0.0410 0.0110 0.0310 0.0280 0.0220 0.0220
(0.0369) (0.0363) (0.0374) (0.0371) (0.0365) (0.0373) (0.0385) (0.0374) (0.0381)
\(\hat{\beta}_{3}\) -0.0120 -0.0120 -0.0080 -0.0280 -0.0580 -0.0780 -0.0070 -0.0230 -0.0430
(0.0709) (0.0650) (0.0652) (0.0658) (0.0665) (0.0984) (0.0672) (0.0658) (0.0700)
3-5 \(\hat{\beta}_{1}\) 0.0600\(^{*}\) 0.0500 0.0540 0.0560\(^{*}\) 0.0860\(^{**}\) 0.0660\(^{*}\) 0.0370 0.0670\(^{*}\) 0.0870\(^{**}\)
(0.0360) (0.0378) (0.0392) (0.0330) (0.0429) (0.0360) (0.0415) (0.0400) (0.0418)
\(\hat{\beta}_{2}\) 0.0460 0.0500 0.0570 0.0390 0.0090 0.0290 0.0270 0.0260 0.0170
(0.0362) (0.0358) (0.0372) (0.0369) (0.0364) (0.0369) (0.0382) (0.0373) (0.0382)
\(\hat{\beta}_{3}\) -0.0020 -0.0010 -0.0020 -0.0150 -0.0450 -0.0280 -0.0060 -0.0360 -0.0560
(0.0649) (0.0640) (0.0065) (0.0661) (0.0664) (0.0889) (0.0663) (0.0652) (0.0650)
\(\hat{\beta}_{4}\) 0.0870\(^{**}\) 0.0760\(^{*}\) 0.0820\(^{**}\) 0.0790\(^{*}\) 0.1090\(^{**}\) 0.0890\(^{*}\) 0.1030\(^{**}\) 0.1330\(^{***}\) 0.1530\(^{***}\)
(0.0423) (0.0402) (0.0412) (0.0436) (0.0431) (0.0477) (0.0444) (0.0434) (0.0465)
Village FEs YES
Time FEs YES

Notes: The parameter \(\hat{\beta}_1\) denotes the coefficient of extreme rainfall, \(\hat{\beta}_2\) denotes the coefficient of the previous year of extreme rainfall, \(\hat{\beta}_3\) denotes the coefficient of the current and previous year of extreme rainfall, and \(\hat{\beta}_4\) denotes the coefficient of the human disease epidemic. The term 2SLS refers to the 2SLS estimator, OGMM refers to the feasible optimal GMM estimator, and BGMM refers to the feasible best GMM estimator. The \(^{*}\) denotes \(p<0.1\), \(^{**}\) denotes \(p<0.05\) and \(^{***}\) denotes \(p<0.01\). Theoretical standard deviations are reported in parentheses.

0.8

Estimation results for the MESS robustness check.
2SLS OGMM BGMM 2SLS OGMM BGMM 2SLS OGMM BGMM
3-5 \(\bar{d}_0=10\%\) \(\bar{d}_0=25\%\) \(\bar{d}_0=50\%\)
3-5 [2]*\(\mathrm{Robust}_1\) \(\hat{\beta}_{1}\) 0.1070\(^{*}\) 0.0900\(^{*}\) 0.1100\(^{*}\) 0.0880 0.1180\(^{*}\) 0.0980\(^{*}\) 0.0420 0.0720\(^{*}\) 0.0920\(^{*}\)
(0.0617) (0.0499) (0.0610) (0.1198) (0.0709) (0.0533) (0.0634) (0.0435) (0.0558)
\(\hat{\beta}_{4}\) 0.1220 0.1020 0.1220 0.1200 0.1500\(^{**}\) 0.1300\(^{*}\) 0.1410\(^{*}\) 0.1710\(^{**}\) 0.1910\(^{**}\)
(0.0756) (0.0754) (0.0747) (0.0820) (0.0759) (0.0774) (0.0815) (0.0847) (0.0831)
3-5 [2]*\(\mathrm{Robust}_2\) \(\hat{\beta}_{1}\) 0.0470 0.0500\(^{*}\) 0.0700\(^{*}\) 0.0280 0.0200 0.0290 0.0320\(^{*}\) 0.0620\(^{*}\) 0.0420\(^{*}\)
(0.0828) (0.0297) (0.0378) (0.0893) (0.8679) (0.0232) (0.0183) (0.0365) (0.0245)
\(\hat{\beta}_{4}\) 0.1100 0.0940\(^{*}\) 0.0740\(^{*}\) 0.0900 0.0600 0.0800 0.1390 0.1090 0.0890
(0.1066) (0.0563) (0.0431) (0.1118) (0.1118) (0.1098) (0.1212) (0.1109) (0.1099)
3-5 [2]*\(\mathrm{Robust}_3\) \(\hat{\beta}_{1}\) 0.0680 0.0660 0.0860 0.0110 -0.0190 -0.0390 0.0200 0.0500 0.0700
(0.0904) (0.0865) (0.0819) (0.0977) (0.1159) (0.1441) (0.0850) (0.0845) (0.0811)
\(\hat{\beta}_{4}\) 0.1380 0.1480 0.1680 0.1940 0.1640 0.1840 0.1190 0.0830 0.0960
(0.1017) (0.1024) (0.1030) (0.1534) (0.1163) (0.1693) (0.1180) (0.1038) (0.0986)
Village FEs YES
Time FEs YES

Notes: In \(\mathrm{Robust}_1\), \(Y_{t}\) denotes witch murders per 1000 households. In \(\mathrm{Robust}_2\), \(Y_{t}\) denotes witch murders and attacks per 1000 households, and in \(\mathrm{Robust}_3\), \(Y_{t}\) denotes total murders per 1000 households. The parameter \(\hat{\beta}_1\) denotes the coefficient of extreme rainfall, and \(\hat{\beta}_4\) denotes the coefficient of the human disease epidemic. The term 2SLS refers to the two-stage least squares estimator, OGMM refers to the feasible optimal GMM estimator, and BGMM refers to the feasible best GMM estimator. The \(^{*}\) denotes \(p<0.1\), \(^{**}\) denotes \(p<0.05\) and \(^{***}\) denotes \(p<0.01\). Theoretical standard deviations are reported in parentheses. We observe that the regressors have no significant impact on total murders under the \(\mathrm{Robust}_3\) specification, and this is consistent with [1].

0.7

Estimation results for the SAR, \(\Var(\epsilon_{it})=\sigma_i^2\).
2SLS OGMM BGMM 2SLS OGMM BGMM 2SLS OGMM BGMM
3-5 \(\bar{d}_{0}=10\%\) \(\bar{d}_{0}=25\%\) \(\bar{d}_{0}=50\%\)
[emp:EP1] \(\hat\beta_1\) 0.0810\(^{*}\) 0.0880\(^{*}\) 0.0890\(^{*}\) 0.0720\(^{*}\) 0.0730\(^{*}\) 0.0760\(^{*}\) 0.0930\(^{*}\) 0.0830\(^{*}\) 0.0820\(^{*}\)
(0.0427) (0.0460) (0.0503) (0.0409) (0.0418) (0.0403) (0.0516) (0.0478) (0.0469)
3-5 [2]*[emp:EP2] \(\hat\beta_1\) 0.0770\(^{*}\) 0.0800\(^{*}\) 0.0860\(^{*}\) 0.0660\(^{*}\) 0.0690\(^{*}\) 0.0670\(^{*}\) 0.0920\(^{*}\) 0.0830\(^{*}\) 0.0840\(^{*}\)
(0.0395) (0.0439) (0.0474) (0.0363) (0.0406) (0.0344) (0.0477) (0.0426) (0.0436)
\(\hat\beta_2\) 0.1050 0.0200 0.0030 0.1510\(^{**}\) 0.0950 0.0620 0.1490\(^{*}\) 0.0840 0.0420
(0.0755) (0.0566) (0.0635) (0.0741) (0.0634) (0.0666) (0.0771) (0.0702) (0.0874)
\(\hat\beta_3\) -0.0190 0.0410 0.0270 0.0230 -0.0140 -0.0350 -0.0130 -0.0240 -0.0410
(0.1327) (0.1026) (0.1054) (0.1339) (0.1206) (0.1157) (0.1343) (0.1149) (0.1373)
3-5 [2]*[emp:EP3] \(\hat\beta_1\) 0.0900\(^{*}\) 0.0890\(^{*}\) 0.0860\(^{*}\) 0.0620\(^{*}\) 0.0620\(^{*}\) 0.0550\(^{*}\) 0.0960\(^{*}\) 0.0870\(^{**}\) 0.0870\(^{*}\)
(0.0521) (0.0510) (0.0488) (0.0350) (0.0319) (0.0333) (0.0576) (0.0445) (0.0490)
\(\hat\beta_2\) 0.1160 0.1160\(^{*}\) 0.0710 0.1380\(^{*}\) 0.1260\(^{*}\) 0.0780 0.1350\(^{*}\) 0.0730 0.0720
(0.0738) (0.0671) (0.0634) (0.0740) (0.0673) (0.0806) (0.0774) (0.0677) (0.0888)
\(\hat\beta_3\) -0.0230 -0.0230 0.0000 0.0620 -0.0230 -0.0430 0.0390 0.0050 0.0050
(0.1342) (0.1255) (0.1156) (0.1342) (0.1123) (0.1418) (0.1353) (0.1149) (0.2274)
\(\hat\beta_4\) 0.1230\(^{*}\) 0.1230\(^{*}\) 1.1930\(^{*}\) 0.1040\(^{*}\) 0.1330\(^{*}\) 0.1710\(^{*}\) 0.1170\(^{*}\) 0.0850\(^{*}\) 0.0850\(^{*}\)
(0.0739) (0.0732) (0.7025) (0.0540) (0.0746) (0.0998) (0.0608) (0.0493) (0.0485)
Village FEs YES
Time FEs YES

Notes: The parameter \(\hat{\beta}_1\) denotes the coefficient of extreme rainfall, \(\hat{\beta}_2\) denotes the coefficient of the previous year of extreme rainfall, \(\hat{\beta}_3\) denotes the coefficient of the current and previous year of extreme rainfall, and \(\hat{\beta}_4\) denotes the coefficient of the human disease epidemic. The term 2SLS refers to the two-stage least squares estimator, OGMM refers to the feasible optimal GMM estimator, and BGMM refers to the feasible best GMM estimator. The \(^{*}\) denotes \(p<0.1\), \(^{**}\) denotes \(p<0.05\) and \(^{***}\) denotes \(p<0.01\). Theoretical standard deviations are reported in parentheses.

0.8

a

b

c

Figure 2: Estimated pure-geographical \(G_{1}\), \(G_{2}\) and \(G_{3}\) based on three estimators, the case of \(\operatorname{Var}(\epsilon_{it})=\sigma_i^2\)..

a

b

Figure 3: Estimated \(G_{1}\) based on economic and cultural geographic distances for the MESS, in the case of \(\operatorname{Var}(\epsilon_{it})=\sigma_i^2\)..

a

b

Figure 4: Estimated \(G_{1}\) based on economic and cultural geographic distances for the SAR, \(\operatorname{Var}(\epsilon_{it})=\sigma_i^2\)..

a
b
c
d

Figure 5: Marginal effects with \(\bar{d}_{0}=25\%\). Top row: Short-term effects. Bottom row: Long-term effects.. a — Short-term marginal effects, rainfall, b — Short-term marginal effects, epidemic, c — Long-term marginal effects, rainfall, d — Long-term marginal effects, epidemic

7 Conclusion↩︎

We provide a method for estimating SDPD models with unknown spatial weights using series approximations, while maintaining ease of implementation. We allow the spatial weights in the outcome, lagged-outcome, and disturbance channels to be unknown functions of exogenous distance measures and cover both SAR and MESS specifications. Using sieve-based GMMEs, we construct an OGMME that is robust to unknown heteroskedasticity and a more efficient feasible BGMME via both linear and quadratic moments. This estimation approach can be directly extended to time-varying exogenous distances without any technical difficulties. Testing and selecting models between SAR and MESS matrix remain an area for future research.

0.8

8 Notations and Definitions↩︎

We first collect all notations in our paper. Although some of them are already defined in the main text we believe it will be useful to state all notations at this point.

Notation 1. () For the sieve approximation, define \(\ell_{n}=\max_{1\leq k\leq 3}\ell_{k}\) and \(\underline{\varsigma}=\min_{1\leq k\leq 3}\varsigma_{k}\).

() Define \(H=\{h_{ij}\}\) as the matrix whose \((i,j)\)-th entry is \(h_{ij}\). Then, for \(k=1,2,3\) and \(j=1,...,\ell_{k}\), we assume \(g_{k}(d)=\sum_{p_{k}=1}^{\ell_{k}}\lambda_{p_{k}}\phi_{kp_{k}}(d)+\delta_{k}(d)\), where \(\phi_{kp_{k}}(\cdot)\) are basis functions and \(\delta_{k}(d)\) are suitably decaying approximation errors. We then denote the sieve approximation and related matrices as \(\xi_{k}(d)=\sum_{p_{k}=1}^{\ell_{k}}\lambda_{p_{k}}\phi_{kp_{k}}(d)\), \(\Xi_{k}=\{\xi_{k}(d_{ij})\}\), and \(\varPhi_{kp_{k}}=\{\phi_{kp_{k}}(d_{ij})\}\). Defining \(\Delta_{k}=\{\delta_{k}(d_{ij})\}\), we obtain the decomposition \(B_{k}=S_{k}+R_{k}\), where \(S_{k}\) depends on the sieve matrix \(\Xi_{k}\) and \(R_{k}\) collects the approximation error matrix. Specifically, for the SAR, \(S_{1}=I_{n}-\Xi_{1}, S_{2}=\Xi_{2}, S_{3}=I_{n}-\Xi_{3}\) with \(R_{1}=-\Delta_{1}, R_{2}=\Delta_{2}, R_{3}=-\Delta_{3}\); for the MESS, \(S_{k}=e^{\Xi_{k}}\) and \(R_{k}=S_{k}(e^{\Delta_{k}}-I_{n})\).

() For a set of matrices \(M_{j}\), \(j=1,...,p\), the following equation represents a specific matrix operation: \[M_{1}[M_{2},M_{3},...,M_{p-1}]M_{p}=[M_{1}M_{2}M_{p},M_{1}M_{3}M_{p},...,M_{1}M_{p-1}M_{p}]\] where the column-dimension in \(M_{m}\) is equal to the row-dimension in \(M_{1}\), and the row-dimension in \(M_{m}\) is equal to the column-dimension in \(M_{p}\), for \(m=2,...,p-1\).

() For the FOD transformation, Let \([F_{T,T-1},\frac{l_T}{\sqrt{T}}]\) be the orthogonal matrix of \(J_T=I_{T}-\frac{1}{T}l_{T}l_{T}'\), where \(l_{T}\) is the \(T\times 1\) vector of ones, and then, the \(n \times T\) matrix \(\left[H_{1}, H_{2}, \cdots, H_{T}\right]\) can be transformed into the \(n(T-1)\) matrix \(\left[H_{1}^{*}, H_{2}^{*}, \cdots, H_{T-1}^{*}\right]=\left[H_{1}, H_{2}, \cdots, H_{T}\right] F_{T,T-1}\). Also, the \(n(T-1)\) matrix \([H_{0}^{(*,-1)}, H_{1}^{(*,-1)}, \cdots, H_{T-2}^{(*,-1)}]=[H_{0}, H_{1}, \cdots, H_{T-1}]F_{T,T-1}\). For example, \(Y_{t}^{*}=h_{Tt}(Y_{t}-\sum_{h=t+1}^{T}Y_{h})\) and \(Y_{t-1}^{(*,-1)}=h_{Tt}(Y_{t-1}-\sum_{h=t}^{T-1}Y_{h})\) depend on current and future variables, but not on the past ones, where \(h_{Tt}=\sqrt{\frac{T-t}{T-t+1}}\).

() For the stacked matrix form, \(\mathbf{H}_{N}^{*}=(H_{1}^{*\prime},...,H_{T-1}^{*\prime})'\) and \(\mathbf{H}_{N}^{(*,-1)}=(H_{0}^{(*,-1)\prime},...,H_{T-2}^{(*,-1)\prime})'\) for any \(n\times \ell_{h}\) matrices \(H_{t}\) and \(H_{t-1}\) with full columns. For example, \(\mathbf{Y}_{N}^{*}=(Y_{1}^{*\prime},...,Y_{T-1}^{*\prime})\), \(\mathbf{Y}_{N}^{(*,-1)}=(Y_{0}^{(*,-1)\prime},...,Y_{T-2}^{(*,-1)\prime})\) and \(\mathbf{E}_{N}^{*}=(E_{1}^{*\prime},...,E_{T-1}^{*\prime})\). For \(k=1,2,3\) and \(j=1,..,\ell_{k}\), \(\mathbf{S}_{kN}=I_{T-1}\otimes S_{k}\), \(\pmb{\Phi}_{kj,N}=I_{T-1}\otimes \varPhi_{kj}\), where \(\varPhi_{kj}\) and \(S_{k}\) are defined as Notation 1(). Thus, \(\mathbf{Z}_{N}^{*}=[\mathbf{Y}_{N}^{(*,-1)}, \mathbf{X}_{N}^{*}]\), and \(\pmb{\alpha}_{T-1}^{*}=(\alpha_{0}^{*},...,\alpha_{T-1}^{*})'\). Moreover, \(\varPhi_{k}=[\varPhi_{k1},\ldots,\varPhi_{kj}]\), \(\pmb{\Phi}_{Nk}=[\pmb{\Phi}_{k1,N},\ldots,\pmb{\Phi}_{kj,N}]\), and \(\mathbf{L}_{N}^{*}=\mathbf{S}_{3N}\left[ \mathbf{Z}_{N}^{*},\tfrac{\partial\mathbf{S}_{1N}}{\lambda_{1}}\mathbf{Y}_{N}^{*}, \tfrac{\partial\mathbf{S}_{2N}}{\lambda_{2}}\mathbf{Y}_{N}^{(*,-1)},\mathbf{0}_{N\times \ell_3}\right]\) with \(\tfrac{\partial\mathbf{S}_{kN}}{\partial\lambda_{k}}=\left[\tfrac{\partial\mathbf{S}_{kN}}{\partial\lambda_{k1}},\cdots,\tfrac{\partial\mathbf{S}_{kN}}{\partial\lambda_{kj}}\right]\).

() Denote \(N=n(T-1)\) and \(\theta=(\pi',\lambda')'\), where \(\pi=(\gamma,\beta')'\) and \(\lambda=(\lambda_{1}^{\prime},\lambda_{2}',\lambda_{3}')'\) with \(\lambda_{k}=(\lambda_{k1},...,\lambda_{kp_{k}})'\) for \(k=1,2,3\). Denote \(\hat{g}_{k}(d)=\pmb{\phi}_{k}'(d)\hat{\lambda}_{k}\), where \(\pmb{\phi}_{k}(d)=[\phi_{k1}(d),...,\phi_{k\ell_{k}}(d)]\) and \(\hat{\lambda}_{k}\) is an estimator of \(\lambda_{k}\). For the dimension, \(\dim(\theta)=\ell_{\theta}\), \(\dim(\pi)=\ell_{\pi}\), \(\dim(\lambda)=\ell_{\lambda}\), \(\dim(m_{N}^{\texttt{quad}}(\theta))=\ell_{p}\), \(\dim(m_{N}^{\texttt{line}}(\theta))=\ell_{q}\), \(\dim(m_{N}(\theta))=\ell_{m}=\ell_{p}+\ell_{q}\). .

We then introduce the following definitions.

Definition 1. () For any real time-varying \(n\times n\) random matrix \(H_t\), we say that \(H_t\) has Property UB* if \[\begin{align} \sup_{t}\|H_t\|_{\mathrm{rc}}&=O_p(1), \; \sup_{t}\|H_t\|_{\mathrm{sp}}=O_p(1), \;\text{and} \; \sup_{t}\|H_t\|=O_p(\sqrt{n}),\\ \intertext{ \text{or the non-stochastic version} } \sup_{t}\|H_t\|_{\mathrm{rc}}&=O(1), \; \sup_{t}\|H_t\|_{\mathrm{sp}}=O(1), \;\text{and} \; \sup_{t}\|H_t\|=O(\sqrt{n}) \end{align}\] () For any real time-varying \(n\times n\) random matrix \(H_t\), we say that \(H_t\) has Property \(O_p(h^j)\) if \[\begin{align} \sup\limits_{t}\|H_t\|_{\mathrm{rc}}&=O_p(h^j), \; \sup\limits_{t}\|H_t\|_{\mathrm{sp}}=O_p(h^j), \;\text{and} \; \sup\limits_{t}\|H_t\|=O_p((\sqrt{n}h)^{j}),\\ \intertext{ \text{or the non-stochastic version O(h^j)} } \sup\limits_{t}\|H_t\|_{\mathrm{rc}}&=O(h^j), \; \sup\limits_{t}\|H_t\|_{\mathrm{sp}}=O(h^j), \;\text{and} \; \sup\limits_{t}\|H_t\|=O((\sqrt{n}h)^{j}) \end{align}\] with \(h \to 0\) and \(\lim_{n \to\infty}\sqrt{n}h \to 0\) for some positive integer \(j\).*

Definition 2. For any real square matrix (or random matrix) \(H_{N}\), where \(N=n(T-1)\), we say that \(H_{N}\) has Property SP if \[\begin{align} \sup_{N}\|H_N\|_{\mathrm{sp}}&=\lim_{N\to\infty}\mu_{\max}(H_N'H_N)=O_p(1), \;\text{and} \; \bigl(\sup_{N}\mu_{\min}(H_N'H_N)\bigr)^{-1}=O_p(1),\\ \intertext{ \text{or the non-stochastic version} } \lim_{N\to\infty}\|H_N\|_{\mathrm{sp}}&=\lim_{N\to\infty}\mu_{\max}(H_N'H_N)<\infty, \;\text{and} \; \lim_{N\to\infty}\mu_{\min}(H_N'H_N)>0. \end{align}\]

Definition 1 introduces the UB and convergence to zero for certain \(n\times n\) random and nonrandom matrices in each \(t\). For example, according to Lemma 8(), \(S_{k}\)’s satisfy [propUB] and \(R_{k}\)’s satisfy [propzero] \(O_{p}(\ell_{n})\). Definition 2 captures boundedness and non-multicollinearity properties of random or fixed matrices of growing dimension.

9 Regularity Assumptions↩︎

Assumption 3. The \(\epsilon_{it}\) are independent with zero mean and finite variances that are bounded away from zero (denoted by \(\sigma^2\), \(\sigma_{i}^2\), or \(\sigma_{t}^2\)), as discussed in [var:homo], [var:hetei] and [var:hetet], respectively. Furthermore, \(\operatorname{E}|\epsilon_{it}|^{4+\delta_{\epsilon}}\leq C_{\epsilon}\) for some constant \(C_{\epsilon}\) and \(\delta_{\epsilon}>0\). 16

Assumption 4. For \(k=1,2,3\) and \(l=1,2,\cdots,\ell_{k}\):
() \(\|G_{k}\|_{\mathrm{rc}}\) and \(\|\Upsilon\|_{\mathrm{rc}}\) are uniformly bounded in \(n\), and \(\rho(A)<1\), where \(A=B_{1}^{-1}(\gamma I_{n}+B_{2})\), \(\Upsilon=\sum_{h=0}^{\infty}\mathrm{abs}(A^{h})\), \([\mathrm{abs}(A)]_{ij}=|A_{ij}|\) and \(A_{ij}\) is the \((i,j)\)-th element of \(A\). Moreover, the initial condition \(Y_{0}\) is exogenously given.
() The basis functions \(\phi_{kl}(\cdot)\) satisfy \(\sup_{i,k,l}\sum_{j=1}^{n}|\phi_{kl}(d_{ij})|<\infty\), \(\sup_{j,k,l}\sum_{i=1}^{n}|\phi_{kl}(d_{ij})|<\infty\).
() For some fixed \(\tau>1\), \(\sup_{l}|\lambda_{l}l^{\tau}|<\infty\).
(’) The \((i,j)\)th element of the approximation error matrix \(\Delta_{kl}\), defined as \(\delta_{k}(d_{ij})=\sum_{l=\ell_{k}+1}^{\infty}\lambda_{l}\phi_{kl}(d_{ij})\), satisfying \(|\delta_{k}(d_{ij})|=O(l^{-\varsigma_{k}})\) for some \(\varsigma_{k}>2\).

Assumption 3 provides regularity assumptions for the idiosyncratic term, and \(\|\Sigma_{t}\|_{\mathrm{sp}}=O(1)\) and \(\|\Sigma_{t}^{-1}\|_{\mathrm{sp}}=O(1)\) since \(\tfrac{1}{\mu_{\min}(\Sigma_{t})}=O(1)\). In particular, we require the existence of fourth moments, thereby relaxing the condition in [34] and [24], which assumes the existence of eighth moments. Assumption 4 imposes the structures of unknown spatial weights matrices \(G_{k}\)’s and approximation tail matrices \(\Delta_{k}\)’s. The condition of () is parallel to assumptions on spatial weights in parametric SDPD models, see, e.g. [2], [4].17 \(\|G_{k}\|_{\mathrm{rc}}\) requires that \(g_{k}(d_{ij})\) either has a compact support or approaches zero fast enough for large \(d_{ij}\) values, and that a growing sample increases the domain, not the density of \(d_{ij}\). Such assumptions are also discussed in [18] and [20]. Alternatively, [16] assume that \(\frac{1}{n}\sum_{i=1}^{n}\sum_{j=1}^{n}\mathbf{1}[d_{ij}\in D]<\infty\) for any fixed domain set \(D\), where \(\mathbf{1}[\cdot]\) is the indicator function. This condition is sufficient to ensure \(\|G_{k}\|_{\mathrm{rc}}<\infty\) the when \(g_{k}(\cdot)\) has bounded support \(D\). \(\rho(A)<1\) indicates that the dynamic system 1 is stable, ruling out unit root or related issues. \(\|\Upsilon\|_{\mathrm{rc}}<\infty\) combines the absolute summability condition and the UB condition of the power series of \(A\). We consider the initial condition \(Y_{0}\) is exogenously given, which does not need to be specified in the estimation or asymptotic analysis. 18 () implies that each approximation matrix \(\varPhi_{kl}=\{\phi_{kl}(d_{ij})\}\) is UB in both row and column sums. Also, it requires that there are only a finite number of neighbors for each \(i\) such that \(\phi_{kl}(d_{ij})\neq 0\) in the limit. This means \(\phi_{kl}(d_{ij})\) declines fast enough with growing \(d_{ij}\), and that observations spread out in an increasing domain. () is a ‘smoothness condition’, and the polynomial expansion can be used by this condition, as discussed in [16]. By (), we can deduce that \(|\delta_{k}(d_{ij})|=O_{p}(l^{-\varsigma_{k}})=o_{p}(1)\). (’) corresponds to a more general setting. In particular, when \(g_k(d_{ij})\) is supported on an unbounded domain, such as \([0, \infty)\) or \((-\infty, \infty)\), the basis function \(\{\phi_{kl}(\cdot)\}\) can be chosen as normalized Laguerre (or Hermite) orthonormal polynomials, following [18].

Assumption 5. \(\{x_{it}\}\), \(\{c_{i}\}\) and \(\{\alpha_{t}\}\) are uniformly bounded constants.

Assumption 5 introduces the UB condition of \(X_{t}\) and FE, which can also be found in [4] and [10]. We can also assume \(\sup_{1\leq l\leq \ell_{x}}\operatorname{E}|x_{lit}|^2\leq C_{x}\) for some constant \(C_{x}\). By Assumptions 3, 4(i) and 5, we can deduce that \(\|Y_{t}\|=O_{p}(\sqrt{n})\) and \(\|Y_{t-1}\|=O_{p}(\sqrt{n})\) by the backward substitution.

Assumption 6 (IV matrix). () The typical elements of \(Q_{t}\), denoted by \(q_{ijt}\) for \(i=1,\cdots,n\) and \(j=1,\cdots,\ell_{q}\), satisfy \(\operatorname{E}|q_{ijt}|^2\leq C_{q}\) with some constant \(C_q\). Furthermore, \(\operatorname{E}\bigl(\sum_{t=1}^{T-1}Q_{t}'E_{t}^{*}\big|\mathcal{I}_{t-1}\bigr)=0\), where \(\mathcal{I}_{t-1}\) denotes the \(\sigma\)-field spanned by \((Y_{0},\cdots,Y_{t-1})\), conditional on \((X_{1},\cdots,X_{T},\mathbf{c}_{n},\alpha_{1},\cdots,\alpha_{T}\)).
() For each \(j=1,...,\ell_{p}\), \(P_{jt}\) satisfies [propUB].
() \(\lim\limits_{n,T \to\infty}\frac{1}{n(T-1)}\mathbf{Q}_{N}'\mathbf{J}_{N}\mathbf{Q}_{N}\), \(\lim\limits_{n,T \to\infty}\frac{1}{n(T-1)}\mathbf{Q}_{N}'J_{N}\mathbf{\Sigma}_{N}J_{N}\mathbf{Q}_{N}\), and \(\lim\limits_{n,T \to\infty}\frac{1}{n(T-1)}\mathbf{Q}_{N}' \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{Q}_{N}\) satisfy [propSP].
() \(\ell_{q}\sim\ell_{n}\) and \(\ell_{p}\sim\ell_{n}\).

Assumption 6 collects the regularity conditions on the instrumental variables. Condition () imposes the standard IV requirements used in SDPD models (see, e.g. [4]). Condition () requires a UB property for the IV matrices entering the quadratic moment conditions, which in turn guarantees that \(\varPhi_{kl}\) is well-defined. Importantly, we do not impose any additional high-level restrictions on the instruments, such as \(\operatorname{tr}(J_n P_{jt})=0\) (as in [4]) or \(\mathtt{Diag}(P_{jt})=0\) (as in [44]). Condition () ensures the existence of the asymptotic covariance matrix \(\Omega_{N}\), which involves the inverse of \(\lim\limits_{n,T\to\infty}\frac{1}{n(T-1)}\mathbf{Q}_{N}'J_{N}\mathbf{Q}_{N}\) in the 2SLS approch, the inverse of \(\lim\limits_{n,T\to\infty}\frac{1}{n(T-1)}\mathbf{Q}_{N}'J_{N}\mathbf{\Sigma}_{N}J_{N}\mathbf{Q}_{N}\) in the OGMM approach and the inverse of \(\lim\limits_{n,T\to\infty}\frac{1}{n(T-1)}\mathbf{Q}_{N}' \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{\Sigma}_{N} \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{Q}_{N}=\lim\limits_{n,T\to\infty}\frac{1}{n(T-1)}\mathbf{Q}_{N}' \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{Q}_{N}\) in the BGMM approach. Finally, condition () restricts the dimension of the moment conditions, \(\ell_{m}=\ell_{p}+\ell_{q}\sim \ell_{n}\), so that the number of moments grows at the same rate as the cross-sectional dimension \(n\), in line with [24]. Unlike the setting in [4], where the many-moments problem arises from both spatial and time expansions, in our framework, it stems solely from the number of instruments, \(\ell_{m}\), which increases with the order of the spatial power series used to approximate the unknown spatial weights. To control this growth, we impose explicit rate conditions that ensure the expansion order (and hence \(\ell_{m}\)) is asymptotically dominated by both the cross-sectional dimension \(n\) and the time dimension \(T\).

Assumption 7 (Identification). () \(\frac{1}{n(T-1)}\mathbf{Q}_{N}'\mathbf{J}_{N}\mathbf{L}_{N}^{*}\) and \(\frac{1}{n(T-1)}\mathbf{L}_{N}^{*\prime} \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{L}_{N}^{*}\) satisfy [propSP], where \(\mathbf{L}_{N}^{*}\) is defined in Notation 1.
() \(\Omega_{N}\) satisfies [propSP], where \(\Omega_{N}\) is defined in equation 6 .
() \(D_{N}'\Omega_{N}^{-1}D_{N}\) satisfies [propSP], where \(D_{N}\) is defined in equation 27 .

Assumption 7 is a regularity condition to ensure the identifiability of \(\theta_{0}\). From (), we have \(\lim_{n,T\to\infty}\|D_{N}'\Omega_{N}^{-1}D_{N}\|_{\mathrm{sp}}=O(1)\).

10 Lemmas and Theorems↩︎

10.1 Key lemmas↩︎

In this section, we introduce the key lemmas, while the remaining lemmas are collected in the supplementary file. In the following part, we denote the true estimand as \(\theta_{0}=(\gamma_0',\lambda_{0}')'\).

Lemma 1. () For any time-varying \(n \times n\) matrix \(H_{t}\), denote \(H^{\diamond}_{t}\) as a transformed version of \(H_{t}\), such that \(\left[H^{\diamond}_{t}\right]_{i j}=[H_{t}]_{i j}\) if \(i \neq j\), and \[\begin{align} \label{eq:transK} \left[H^{\diamond}_{t}\right]_{ii}=\frac{1}{n-2}\sum_{j\neq i}^n\big([H_{t}]_{ij}+[H_{t}]_{ji}\big)-\frac{1}{(n-1)(n-2)}\sum_{j=1}^{n}\sum_{k\neq j}^n[H_{t}]_{jk}. \end{align}\qquad{(1)}\] Then, \(\mathtt{diagM}\left(J_n H^{\diamond}_{t} J_n\right)=\mathbf{0}\).
() In the case of [var:hetei], the vector of diagonal elements \(\mathbf{h}_{diag}^\diamond = ([H_t]_{11}^\diamond, \dots, [H_t]_{nn}^\diamond)'\) is the solution to \(C \mathbf{h}_{diag}^\diamond = d\), given by \(C^{+} d\), where: \[C = \begin{pmatrix} 1 - \frac{2\sigma_1^{-2}}{\varsigma} + \frac{\sigma_1^{-4}}{\varsigma^2} & \frac{\sigma_2^{-4}}{\varsigma^2} & \cdots & \frac{\sigma_n^{-4}}{\varsigma^2} \\ \frac{\sigma_1^{-4}}{\varsigma^2} & 1 - \frac{2\sigma_2^{-2}}{\varsigma} + \frac{\sigma_2^{-4}}{\varsigma^2} & \cdots & \frac{\sigma_n^{-4}}{\varsigma^2} \\ \vdots & \vdots & \ddots & \vdots \\ \frac{\sigma_1^{-4}}{\varsigma^2} & \frac{\sigma_2^{-4}}{\varsigma^2} & \cdots & 1 - \frac{2\sigma_n^{-2}}{\varsigma} + \frac{\sigma_n^{-4}}{\varsigma^2} \end{pmatrix},\] with \(\varsigma = \sum_{k=1}^n \sigma_k^{-2}\). The \(i\)-th element of the vector \(d\) is: \[d_i = \frac{1}{\varsigma \sigma_i^2} \left( \sum_{j \neq i} \sigma_j^{-2} [H_{t}]_{ji} + \sum_{j \neq i} \sigma_j^{-2} [H_{t}]_{ij} \right) - \frac{1}{\varsigma^2 \sigma_i^2} \sum_{j=1}^n \sum_{k \neq j} \sigma_j^{-2} \sigma_k^{-2} [H_{t}]_{jk}.\] In the cases of [var:hetet] and [var:homo], \(\left[H^{\diamond}_{t}\right]_{ii}\) is the same as equation ?? . Then, the diagonal elements of \(\ifstrequal{t}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{t}}) }{ {J}({\Sigma_{t}}) } H_t^\diamond \ifstrequal{t}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{t}}) }{ {J}({\Sigma_{t}}) }\) are equal to zero.

For (), the proof can be found in [25]. For (), the proof can be found in [25] in the case of [var:hetei]. In the cases of [var:hetet] and [var:homo], substituting \(J(\Sigma_t)=\frac{1}{\sigma_t^2} J_n\) or \(\frac{1}{\sigma^2} J_n\), we have \(\mathtt{Diag}(J_n H_t^\diamond J_n)=0\). \(\blacksquare\)

Denote \(\mathscr{G}_{N}(\theta)=m_{N}'(\theta)\Omega_{N}^{-1}m_{N}(\theta)\) and \(\operatorname{E}\mathscr{G}_{N}(\theta)=\bigl[\operatorname{E}\bigl(m_{N}(\theta)\bigl)\bigr]'\Omega_{N}^{-1}\operatorname{E}\bigl(m_{N}(\theta)\bigl)\). Then we consider the following lemma to establish the consistency, i.e., \(\|\hat{\theta}_{gmm}-\theta_{0}\|=o_{p}(1)\). The idea is borrowed from [45].

Lemma 2. Let \(\Theta_{N}=\Gamma\times\Lambda_{N}\), where \(\Gamma\subset\mathbb{R}^{\ell_{\pi}}\) is a totally UB set with fixed dimension and \(\Lambda_{N}=\{\lambda\subset\mathbb{R}^{\ell_{n}}:\|\lambda\|_{\infty}\leq C_{\lambda}\}\) for each \(l=1,\cdots,\ell_{n}\). Under Assumptions 3-7:
() For any \(\theta\in\Theta_{N}\) and \(\eta>0\), if \(\|\theta-\theta_{0}\|\geq\eta\), there exists a constant \(C_{\eta}\) such that \(\bigl|\operatorname{E}\mathscr{G}_{N}(\theta)-\operatorname{E}\mathscr{G}_{N}(\theta_{0})\bigr|\geq C_{\eta}\) when \(\ell_{n}^{1/2-\underline{\varsigma}}\to 0\) as \((n,T)\to\infty\).
() \(\sup_{\theta\in\Theta}\bigl|\mathscr{G}_{N}(\theta)-\operatorname{E}\bigl(\mathscr{G}_{N}(\theta)\bigr)\bigr|=o_{p}(1)\) when \(\ell_{n}^{2-\underline{\varsigma}}+\frac{\ell_{n}^{2}}{n}+\frac{\ell_{n}}{\sqrt{n(T-1)}}\to 0\) as \((n,T)\to\infty\).

() ensures the identification. We have \[\operatorname{E}\mathscr{G}_{N}(\theta)-\operatorname{E}\mathscr{G}_{N}(\theta_{0})= \varrho_{1N}(\theta)+2\varrho_{2N}(\theta),\] where \[\varrho_{1N}(\theta)=\left[\operatorname{E}(m_{N}(\theta))-\operatorname{E}(m_{N}(\theta_{0}))\right]'\Omega_{N}^{-1}\left[\operatorname{E}(m_{N}(\theta))-\operatorname{E}(m_{N}(\theta_{0}))\right],\] and \[\varrho_{2N}(\theta)=\left[\operatorname{E}(m_{N}(\theta))-\operatorname{E}(m_{N}(\theta_{0}))\right]'\Omega_{N}^{-1}\operatorname{E}(m_{N}(\theta_{0})).\] By the mean-value theorem, \(\operatorname{E}(m_{N}(\theta))-\operatorname{E}(m_{N}(\theta_{0}))=D_{N}(\tilde{\theta})(\theta-\theta_{0})\), where \(D_{N}\) is defined as equation 27 and \(\tilde{\theta}\), where \(\tilde{\theta}=\theta_{0}+\kappa(\theta-\theta_{0})\) with \(|\kappa|<1\). Then, \(\varrho_{1N}(\theta)=(\theta-\theta_{0})'D_{N}(\tilde{\theta})'\Omega_{N}^{-1}D_{N}(\tilde{\theta})(\theta-\theta_{0})\). From Assumption 7(), we have \(\|\Omega_{N}^{-1}\|_{\mathrm{sp}}\leq\tfrac{1}{\mu_{\min}(\Omega_{N})}=O(1)\). Combing Assumption 7(), we can find that there exists a constant \(C_{\varrho}>0\), such that for all large \((n,T)\), \(\varrho_{1N}(\theta)\geq C_{\varrho}\|\theta-\theta_{0}\|^2\). Hence, if \(\|\theta-\theta_{0}\|\geq\eta\), then \(\varrho_{1N}(\theta)\geq C_{\varrho}\eta^2\). Note that \[\begin{align} |\varrho_{2N}|&=\left|\left[\operatorname{E}(m_{N}(\theta))-\operatorname{E}(m_{N}(\theta_{0}))\right]'\Omega_{N}^{-1/2}\Omega_{N}^{-1/2}\operatorname{E}(m_{N}(\theta_{0}))\right|\leq |\varrho_{1N}|^{1/2}|\operatorname{E}'(m_{N}(\theta_{0}))\Omega_{N}^{-1}\operatorname{E}(m_{N}(\theta_{0}))|^{1/2}\\ &\leq |\varrho_{1N}|^{1/2}\cdot\|\Omega_{N}^{-1/2}\operatorname{E}(m_{N}(\theta_{0}))\| \end{align}\] by Cauchy-Schwarz inequality. From Lemma 16 in the supplement and Assumption 7(), we have \(\|\Omega_{N}^{-1/2}\operatorname{E}(m_{N}(\theta_{0}))\|\leq\|\Omega_{N}^{-1/2}\|_{\mathrm{sp}}\|\operatorname{E}(m_{N}(\theta_{0}))\|\leq C\ell_{n}^{1/2-\underline{\varsigma}}\), and thus \(|\varrho_{2N}(\theta)|\leq C\ell_{n}^{1/2-\underline{\varsigma}}|\varrho_{1N}(\theta)|^{1/2}\). For \(\|\theta-\theta_{0}\|\geq\eta\), we have \[|\operatorname{E}\mathscr{G}_{N}(\theta)-\operatorname{E}\mathscr{G}_{N}(\theta_{0})|\geq\varrho_{1N}(\theta)-|\varrho_{2N}(\theta)|\geq\varrho_{1N}(\theta)^{1/2}(\varrho_{1N}(\theta)^{1/2}-2C\ell_{n}^{1/2-\underline{\varsigma}}).\] Since \(\varrho_{1N}(\theta)^{1/2}\geq C_{\varrho}^{1/2}\eta^{1/2}\), for all large \((n,T)\), we can have \(2C\ell_{n}^{1/2-\underline{\varsigma}}\leq C_{\varrho}^{1/2}\eta^{1/2}\) as \(\ell_{n}^{1/2-\underline{\varsigma}}\) being sufficiently small, and thus \[\bigl|\operatorname{E}\mathscr{G}_{N}(\theta)-\operatorname{E}\mathscr{G}_{N}(\theta_{0})\bigr|\geq C_{\varrho}\eta=:C_{\eta}>0.\] () ensures the uniform convergence. We use the method proposed by [46]. First, note that \(\mathscr{G}_{N}(\theta)-\operatorname{E}\mathscr{G}_{N}(\theta)=H_{1N}(\theta)+2H_{2N}(\theta)\), where \[H_{1N}(\theta)=\left[m_{N}(\theta)-\operatorname{E}(m_{N}(\theta))\right]'\Omega_{N}^{-1}\left[m_{N}(\theta)-\operatorname{E}(m_{N}(\theta))\right],\] and \[H_{2N}(\theta)=\left[m_{N}(\theta)-\operatorname{E}(m_{N}(\theta))\right]'\Omega_{N}^{-1}\operatorname{E}(m_{N}(\theta)).\] From the Proof of Theorem 1, we have \(\|m_{N}(\theta)-\operatorname{E}(m_{N}(\theta))\|=O_{p}\Bigl(\ell_{n}^{3/2-\underline{\varsigma}}+\frac{\ell_{n}^{3/2}}{n}+\sqrt{\frac{\ell_{n}}{n(T-1)}}\Bigr)\) and \(\operatorname{E}\|m_{N}(\theta)\|=O(\sqrt{\ell_{n}})\). Thus, using \(\|\Omega_{N}^{-1}\|_{\mathrm{sp}}=O(1)\), we have \[\begin{align} |\mathscr{G}_{N}(\theta)-\operatorname{E}\mathscr{G}_{N}(\theta)|& \leq|H_{1N}(\theta)|+|H_{2N}(\theta)| \leq\|\Omega_{N}^{-1}\|_{\mathrm{sp}}|\left(\|m_{N}(\theta)-\operatorname{E}(m_{N}(\theta))\|^2+2\|m_{N}(\theta)-\operatorname{E}(m_{N}(\theta))\|\|m_{N}(\theta)\|\right) \\& =O_{p}\Big(\ell_{n}^{2-\underline{\varsigma}}+\frac{\ell_{n}^{2}}{n}+\frac{\ell_{n}}{\sqrt{n(T-1)}}\Big)=o_{p}(1). \end{align}\] Second, for any \(\theta, \tilde{\theta}\in \Theta_{0}\), we have \(\mathscr{G}_{N}(\theta)-\mathscr{G}_{N}(\tilde{\theta})=I_{1N}(\theta,\tilde{\theta})+2I_{2N}(\theta,\tilde{\theta})\), where \[I_{1N}(\theta,\tilde{\theta})=\left[m_{N}(\theta)-m_{N}(\tilde{\theta})\right]'\Omega_{N}^{-1}\left[m_{N}(\theta)-m_{N}(\tilde{\theta})\right],\] and \[I_{2N}(\theta,\tilde{\theta})=\left[m_{N}(\theta)-m_{N}(\tilde{\theta})\right]'\Omega_{N}^{-1}m_{N}(\tilde{\theta}).\] Then we have \(m_{N}(\theta)-m_{N}(\tilde{\theta})=D_{N}(\bar\theta)(\theta-\tilde{\theta})+o_{p}(1)\), where \(\bar\theta=\tilde{\theta}+\kappa(\theta-\tilde{\theta})\) with \(|\kappa|<1\) and \(\|m_{N}(\tilde{\theta})\|_{\mathrm{sp}}=O_{p}(\sqrt{\ell_{n}})\). Define \(B_{e}=\{\theta,\tilde{\theta}\in\Theta_{0}:\|\theta-\tilde{\theta}\|\leq e\}\) and \(h_{e}=B_{e}\sup_{u\in\Theta_{0}}\|m_{N}(u )\|=B_{e}\cdot O_{p}(\sqrt{\ell_{n}})\). Hence, \[\begin{align} \sup_{\theta,\tilde{\theta}\in B_{e}}|I_{1N}(\theta,\tilde{\theta})|& \leq \sup_{u\in\Theta_{0}}\|D_{N}(u)\|_{\mathrm{sp}}^2\|\Omega_{N}^{-1}\|_{\mathrm{sp}}B_{e}^2=O_{p}(B_{e}^2) \\ \sup_{\theta,\tilde{\theta}\in B_{e}}|I_{2N}(\theta,\tilde{\theta})|& \leq \sup_{u\in\Theta_{0}}\|D_{N}(u)\|_{\mathrm{sp}}\|\Omega_{N}^{-1}\|_{\mathrm{sp}}h_{e}=O(h_{e})=O_{p}(\sqrt{\ell_{n}}B_{e}). \end{align}\] Then, \(\sup_{\theta,\tilde{\theta}\in B_{e}}|\mathscr{G}_{N}(\theta)-\mathscr{G}_{N}(\tilde{\theta})|=O_{p}(B_e^2+h_e)\), and we can choose \(B_{e}=C\ell_{n}^{-1}\) as \((n,T)\to \infty\) such that \(B_{e}+\sqrt{\ell_{n}}B_{e}\to 0\). Combining \(|\mathscr{G}_{N}(\theta)-\operatorname{E}\mathscr{G}_{N}(\theta)|=o_{p}(1)\), the consistency follows. \(\blacksquare\)

To derive the asymptotic normality, we consider the following lemma.

Lemma 3. Let \(\mathscr{U}_{N}=\frac{1}{\sqrt{n(T-1)}}\pmb\tau'\mathbf{\Sigma}_{\pi_0,gmm}^{-1/2}K_{\pi}m_{\Delta}\), where the \(\ell_{\pi}\times\ell_{m}\) matrix \(K_{\pi}\) is defined in equation 33 , \[\begin{align} \label{eq:CLT95rv} m_{\Delta}\equiv m_{\Delta}(\theta_{0})=(\pmb{E}_{N}^{*\prime}\mathbf{J}_{N}\mathbf{P}_{N1}^{\prime}\mathbf{J}_{N}\mathbf{E}_{N}^{*},\cdots,\pmb{E}_{N}^{*\prime}\mathbf{J}_{N}\mathbf{P}_{N\ell_{p}}^{\prime}\mathbf{J}_{N}\mathbf{E}_{N}^{*},\pmb{E}_{N}^{*\prime}\mathbf{J}_{N}\mathbf{Q}_{N})', \end{align}\qquad{(2)}\] and \(\pmb\tau\) is the \(\ell_{\pi}\times 1\) vector such that \(\|\pmb\tau\|_{\mathrm{sp}}=1\). Under Assumptions 3-6, and \(\sqrt{\frac{(T-1)\ell_{n}^{3}}{n}}+\sqrt{n(T-1)}\ell_{n}^{3/2-\underline{\varsigma}}\to 0\) as \((n,T)\to\infty\), \(\mathscr{U}_{N}\stackrel{d}{\to}N(0,1)\).

By construction, we have \(\operatorname{E}(\mathscr{U}_{N})=0\) and \(\operatorname{Var}(\mathscr{U}_{N})=1\). Denote \(\mathfrak{s}\equiv\pmb\tau'\mathbf{\Sigma}_{\pi_0,gmm}^{-1/2}K_{\pi}=(\mathfrak{s}_{1},...,\mathfrak{s}_{\ell_{p}},\underbrace{\mathfrak{s}_{\ell_{p}+1},...,\mathfrak{s}_{\ell_{m}}}_{\mathfrak{s}_{\ell_{q}}'})\) be then \(1\times\ell_{m}\) vector, then we have

\(\mathscr{U}_{N}=\frac{1}{\sqrt{n(T-1)}}\left(\pmb{E}_{N}^{*\prime}(\sum_{l=1}^{\ell_{p}}\mathfrak{s}_{l}\mathbf{J}_{N}\mathbf{P}_{Nl}\mathbf{J}_{N})\mathbf{E}_{N}^{*}+\mathfrak{s}_{\ell_{q}}'\mathbf{Q}_{N}^{\prime}\mathbf{J}_{N}\mathbf{E}_{N}^{*}\right)= \frac{1}{\sqrt{n(T-1)}}\left(\pmb{E}_{N}^{*\prime}\mathscr{A}_{1N}\mathbf{E}_{N}^{*}+\mathscr{A}_{2N}\mathbf{E}_{N}^{*}\right)\), where \(\mathscr{A}_{1N}=\sum_{l=1}^{\ell_{p}}\mathfrak{s}_{l}\mathbf{J}_{N}\mathbf{P}_{Nl}\mathbf{J}_{N}\) and \(\mathscr{A}_{2N}=\mathfrak{s}_{\ell_{q}}'\mathbf{Q}_{N}'\mathbf{J}_{N}\). Thus, we have \(\|\mathscr{A}_{1N}\|_{1}=O(\sqrt{\ell_{n}})\), \(\sum_{l=1}^{N}|\mathscr{A}_{2N,l}|=O_{p}(\sqrt{\ell_{n}})\). This is because \(\|\mathscr{A}_{1N}\|_{1}\leq\sum_{l=1}^{\ell_{p}}\|\mathfrak{s}_{l}\mathbf{J}_{N}\mathbf{P}_{Nl}\mathbf{J}_{N}\|_{1}\leq\sum_{l=1}^{\ell_{p}}|\mathfrak{s}_{l}|\|\mathbf{J}_{N}\|_{1}^2\|P_{Nl}\|_{1}=O(\sqrt{\ell_{n}})\) by Assumption 6, and \(\sum_{l=1}^{N}|\mathscr{A}_{2N,l}|=O_{p}(\sqrt{\ell_{n}})\) by an analogous argument. Note that \(\operatorname{E}(\pmb{E}_{N}^{*\prime}\mathscr{A}_{1N}\mathbf{E}_{N}^{*})=\sum_{l=1}^{\ell_{p}}\mathfrak{s}_{l}\operatorname{E}\big(\pmb{E}_{N}^{*\prime}(\mathbf{J}_{N}\mathbf{P}_{Nl}\mathbf{J}_{N})\mathbf{E}_{N}^{*}\big)=0\), from Lemma 10 in the supplement, Assumption 1, 2, and \(\sqrt{\frac{(T-1)\ell_{n}^{3}}{n}}+\sqrt{n(T-1)}\ell_{n}^{3/2-\underline{\varsigma}}\to 0\) as \((n,T)\to\infty\) we have \(\mathscr{U}_{N}=\frac{1}{\sqrt{n(T-1)}}\sum_{l=1}^{N}\mathcal{X}_{l}+o_{p}(1)\), where \[\begin{align} \label{eq:calX} \mathcal{X}_{l}=\mathscr{A}_{2N,l}\epsilon_{l}+2\epsilon_{l}\sum_{j=1}^{l-1}\mathscr{A}_{1N,lj}\epsilon_{j}=\epsilon_{l}(\mathscr{A}_{2N,l}+2\sum_{j=1}^{l-1}\mathscr{A}_{1N,lj}\epsilon_{j}). \end{align}\tag{15}\] Consider the \(\sigma\)-field \(\mathcal{F}_{l}=\sigma(\epsilon_{1},...,\epsilon_{l})\), where \(\epsilon_{l}=\epsilon_{it}\), with \(l=i+n(t-1)\) and \(1\leq i\leq n\), \(1\leq t\leq T\). Then \(\mathcal{F}_{l-1}\subseteq\mathcal{F}_{l}\) and \(\mathcal{X}_{l}\) is \(\mathcal{F}_{l}\)-measurable19. Hence, \(\operatorname{E}(\mathcal{X}_{l}|\mathcal{F}_{l-1})=0 \;\text{a.s.}\), and \(\{(\mathcal{X}_{l},\mathcal{F}_{l})|1\leq l\leq N\}\) forms a martingale difference array. Then, we define \[\begin{align} \label{eq:calX95normal} \mathcal{X}_{l}^{\circ}=\frac{1}{\sqrt{N}}\mathcal{X}_{l}. \end{align}\tag{16}\] From [47], sufficient conditions 20 to verify the martingale difference array CLT are \[\begin{align} \label{eq:CLT1} \sum_{l=1}^{N}\operatorname{E}|\mathcal{X}_{l}^{\circ}|^{2+\delta}\stackrel{p}{\rightarrow} 0 \end{align}\tag{17}\] for some \(\delta>0\), and \[\begin{align} \label{eq:CLT2} \sum_{l=1}^{N}\left[\operatorname{E}(\mathcal{X}_{l}^{\circ2}|\mathcal{F}_{l-1})-\operatorname{E}(\mathcal{X}_{l}^{\circ2})\right]\stackrel{p}{\rightarrow} 0. \end{align}\tag{18}\] For 17 , denote \(u=2+\delta\), we have \(\mathcal{X}_{l}^{\circ}=N^{-u/2}\|\mathcal{X}_{l}\|_{1}^{u}\). and \(|\mathcal{X}_{l}|^{u}\leq C_u|\epsilon_{l}|^{u}|\mathscr{A}_{2N,l}|^{u}+2\|\mathscr{A}_{1N}\|_{1}^{u}\bigl(\max_{j<l}|\epsilon_{j}|^{u}\bigr)\) by Minkowski inequality and the convex inequality that \((a+b)^u\leq C_u(a^u+b^u)\), where \(u>1\) and \(C_u=2^{u-1}\). Thus, \(\operatorname{E}(\mathcal{X}_{l}|\mathcal{F}_{l-1})\leq C_{\epsilon}^2C_uC\ell_{n}^{u/2}=O(\ell_{n}^{u/2})\) and \(\sum_{l=1}^{N}\operatorname{E}|\mathcal{X}_{l}^{\circ}|^{u}=O(N)\cdot O(N^{-u/2})\cdot O(\ell_{n}^{u/2})=O\Bigl(\frac{\ell_{n}^{2+\delta}}{(n(T-1))^{\delta}}\Bigr)=o(1)\) according to Assumptions 1, 3, and the rate conditions in Theorem 1(ii). 21

For 18 , from 15 and 16 , we have \(\operatorname{E}(\mathcal{X}_{l}^{\circ 2}|\mathcal{F}_{l-1})=N^{-1}\sigma_{l}^2(\operatorname{E}\mathscr{A}_{2N,l}+2\sum_{j<l}\mathscr{A}_{1N,lj}\epsilon_{j})^2\), and \(\operatorname{E}(\mathcal{X}_{l}^{\circ 2})=N^{-1}\sigma_{l}^2(\operatorname{E}\mathscr{A}_{2N,l}^2+4\sum_{j<l}\mathscr{A}_{1N,lj}^2\sigma_{j}^2)\). Thus, \[\begin{align} \operatorname{E}(\mathcal{X}_{l}^{\circ2}|\mathcal{F}_{l-1})-\operatorname{E}(\mathcal{X}_{l}^{\circ2})& =4N^{-1}\sigma_{l}^2\Big\{\mathscr{A}_{2N,l}\sum_{j<l}\mathscr{A}_{1N,lj}\epsilon_{j}+4\big[\big(\sum_{j<l}\mathscr{A}_{1N,lj}\epsilon_{j}\big)^2-\sum_{j<l}\mathscr{A}_{1N,lj}\sigma_{j}^2\big]\Big\}\\& =4N^{-1}\sigma_{l}^2\left[\mathscr{A}_{2N,l}\sum_{j<l}\mathscr{A}_{1N,lj}\epsilon_{j}+\sum_{j<l}\mathscr{A}_{1N,lj}(\epsilon_{j}^2-\sigma_{j}^2)+2\sum_{j<l}\sum_{k<l}\mathscr{A}_{1N,lj}\mathscr{A}_{1N,lk}\epsilon_{j}\epsilon_{k}\right] \end{align}\] and \[\sum_{l=1}^{N}\left[\operatorname{E}(\mathcal{X}_{l}^{\circ2}|\mathcal{F}_{l-1})-\operatorname{E}(\mathcal{X}_{l}^{\circ2})\right] =4(\Delta_{1}+\Delta_{2}+2\Delta_{3}),\] where \[\begin{align} &\Delta_{1}=\frac{1}{N}\sum_{l=1}^{N}\sigma_{l}^2\mathscr{A}_{2N,l}\sum_{j<l}\mathscr{A}_{1N,lj}\epsilon_{j}, \Delta_{2}=\frac{1}{N}\sum_{l=1}^{N}\sigma_{l}^2\sum_{j<l}\mathscr{A}_{1N,lj}^2(\epsilon_{j}^2-\sigma_{j}^2) \;\text{and}, \;\\ &\Delta_{3}=\frac{1}{N}\sum_{l=1}^{N}\sigma_{l}^2\sum_{j<k<l}\mathscr{A}_{1N,lj}\mathscr{A}_{1N,lk}\epsilon_{j}\epsilon_{k}, \end{align}\] and thus \(\operatorname{E}(\Delta_{k})=0\) for \(k=1,2,3\). For \(\Delta_{1}\), rewrite by swapping sums: \(\Delta_{1}=\frac{1}{N}\sum_{l=1}^{N}\sigma_{l}^2\mathscr{A}_{2N,l}\sum_{j<l}\mathscr{A}_{1N,lj}\epsilon_{j}=\frac{1}{N}\sum_{j=1}^{N-1}a_{j}\epsilon_{j}\). where \(a_{j}=\sum_{l>j}\sigma_{l}^2\mathscr{A}_{2N,l}\mathscr{A}_{1N,lj}\). Thus, and \(\operatorname{E}(\Delta_{1})=\operatorname{Var}(\Delta_{1})\leq\frac{C_{\epsilon}^{2/(4+\delta_{\epsilon})}}{N^2}\sum_{j=1}^{N-1}a_{j}^2\). This is because \(\operatorname{Var}(\epsilon_{j})\leq\operatorname{E}(\epsilon_{j}^2)\leq(\operatorname{E}|\epsilon|^{4+\delta_{\epsilon}})^{2/(4+\delta_{\epsilon})}\) by the norm inequality. Using \[|a_{j}|\leq C_{\epsilon}^{2/(4+\delta_{\epsilon})}\sum_{l>j}|\mathscr{A}_{2N,l}|\sum_{l>j}|\mathscr{A}_{1N,lj}|=O_{p}(\ell_{n}),\] and \(\sum_{j=1}^{N-1}c_{j}^2=O_{p}(N\ell_{n}^{2})\), we have \(\operatorname{Var}(\Delta_{1})=O\left(\frac{\ell_{n}^{2}}{N}\right)\) and \(\|\Delta_{1}\|=O_{p}\left(\sqrt{\frac{\ell_{n}^{2}}{n(T-1)}}\right)\). For \(\Delta_{2}\), rewrite by swapping sums: \(\Delta_{2}=\frac{1}{N}\sum_{j=1}^{N-1}a_{2j}(\epsilon_{j}^2-\sigma_{j}^2)\), where \(a_{2j}=\sum_{l>j}\sigma_{l}^2\mathscr{A}_{1N,lj}^2\). Note that \[\operatorname{Var}(\epsilon_{j}^2-\sigma_{j}^2)=\operatorname{Var}(\epsilon_{j}^2)\leq\operatorname{E}(\epsilon_{j}^4)\leq C_{\epsilon}^{4/(4+\delta_{\epsilon})},\] we have \(\|\Delta_{2}\|=O_{p}\left(\sqrt{\frac{\ell_{n}^{2}}{n(T-1)}}\right)\). For \(\Delta_{3}\), rewrite by swapping sums: \(\Delta_{3}=\frac{1}{N}\sum_{1\leq j<k\leq N-1}\epsilon_{j}\epsilon_{k}a_{3jk}\), where \(a_{3jk}=\sum_{l>k}\sigma_{l}^2\mathscr{A}_{1N,lj}\mathscr{A}_{1N,lk}\). Hence, \(\Delta_{3}^2=\frac{1}{N^2}\sum_{j<k}\sum_{r<s}a_{3jk}a_{3rs}\epsilon_{j}\epsilon_{k}\epsilon_{r}\epsilon_{s} =\frac{1}{N}\sum_{j<k}a_{3jk}^2\epsilon_{j}^2\epsilon_{k}^2+2\sum_{(j<k)<(r<s)}\epsilon_{j}\epsilon_{k}\epsilon_{r}\epsilon_{s}\) where \((j<k)<(r<s)\) means sum over distinct ordered pairs, so each cross term is counted once. If \(\{j,k\}\cap\{r,s\}=\varnothing\) or \(j=r\) but \(k\neq s\), then \(\operatorname{E}\epsilon_{j}\epsilon_{k}\epsilon_{r}\epsilon_{s}=0\). The only nonzero part is the paired case \((j,k)=(r,s)\) with \(j<k\) and \(r<s\). Hence, \(\operatorname{E}(\Delta_{3}^2)=\frac{1}{N^2}\sum_{j<k}a_{3ij}^2\operatorname{E}\epsilon_{i}^2\operatorname{E}\epsilon_{j}^2\leq\frac{C_{\epsilon}^{4/(4+\delta_\epsilon)}}{N^2}\sum_{j<k}a_{3ij}^2\). By the Cauchy-Schwarz inequality, we have \[a_{3ij}^2=(\sum_{l>k}\sigma_{l}^2\mathscr{A}_{1N,lj}\mathscr{A}_{1N,lk})^2\leq C_{\epsilon}^{4/(4+\delta_\epsilon)}(\sum_{l}b_{lj}^2)(\sum_{l}b_{lk}^2)\leq C_{\epsilon}^{4/(4+\delta_\epsilon)}(\sup_{l}\sum_{l}|b_{lj}|)^2(\sup_{l}\sum_{l}|b_{lk}|)^2=O(\ell_{n}),\] \(\operatorname{E}(\Delta_{3}^2)=O(\frac{\ell_{n}^{2}}{N})\) and thus \(\|\Delta_{3}\|=O_{p}\left(\sqrt{\frac{\ell_{n}^{2}}{n(T-1)}}\right)\). Combining Assumptions 1, 3 and \(\sqrt{\frac{(T-1)\ell_{n}^{3}}{n}}\to 0\) as \((n,T)\to\infty\), we have \(\Delta_{k}\stackrel{p}{\to} 0\) for \(k=1,2,3\). The asymptotic normality of \(\mathscr{U}_{N}\) follows. \(\blacksquare\)

We then establish the consistent estimator for \(\Sigma_{t}\).

Lemma 4. Denote \(\mathscr{W}_{t}(\theta)=J_{n}(V_{t}(\theta)-\frac{1}{T}\sum_{h=1}^{T}V_{h}(\theta))\) for any \(\theta\in\Theta\). We can estimate \(\sigma_{i}^2\) by \(\hat{\sigma}_{i}^2=\frac{1}{T}\sum_{t=1}^{T}\hat{\omega}_{it}^2\) in the case of [var:hetei], and estimate \(\sigma_{t}^2\) by \(\hat{\sigma}_{t}^2=\frac{1}{n}\sum_{i=1}^{n}\hat{\omega}_{it}^2\) in the case of [var:hetet], where \(\hat{\omega}_{it}=\omega_{it}(\hat{\theta})\), and \(\omega_{it}=\omega_{it}(\theta_0)\) is the \(i\)-th entry of \(\mathscr{W}_{t}\) and \(\hat{\theta}\) is a consistent estimator of the true estimand \(\theta_0\) with \(\|\hat{\theta}-\theta_0\|=o_{p}(1)\). Under Assumptions 1, 2, and 3-7, and as \((n,T) \to \infty\), the following hold:
() In the case of [var:hetei], \(|\hat{\sigma}_i^2-\sigma_i^2|=o_{p}(1)\).
() In the case of [var:hetet], \(|\hat{\sigma}_t^2-\sigma_t^2|=o_{p}(1)\).

We define the centered components: \(\Delta\epsilon_{it}=\epsilon_{it}- \bar{\epsilon}_{i\cdot}-\bar{\epsilon}_{\cdot t}+\bar{\epsilon}_{\cdot\cdot}\), where \(\bar{\epsilon}_{i\cdot}= \frac{1}{T}\sum_{t=1}^{T}\epsilon_{it}\), \(\bar{\epsilon}_{\cdot t}= \frac{1}{n}\sum_{i=1}^{n}\epsilon_{it}\) and \(\bar{\epsilon}_{\cdot\cdot}=\frac{1}{nT}\sum_{i,t}\epsilon_{it}\). And \(\Delta r_{it}=r_{it}-\bar{r}_{i\cdot}-\bar{r}_{\cdot t}+\bar{r}_{\cdot\cdot}\) is defined by similar. The proof of () follows similarly to that of (), so we will only provide the proof for () here. Thus, \(\omega_{it}=\Delta\epsilon_{it}+\Delta r_{it}\) and \(\hat{\omega}_{it}=\omega_{it}+(\hat{\omega}_{it}-\omega_{it})\). We have \(\hat{\omega}_{it}^2=\omega_{it}^2+(\hat{\omega}_{it}-\omega_{it})^2+2\omega_{it}(\hat{\omega}_{it}-\omega_{it})\). Thus, for (), \[\hat{\sigma}_{i}^2=\dfrac{1}{T}\sum_{t=1}^{T}\omega_{it}^2+ \dfrac{1}{T}\sum_{t=1}^{T}(\hat{\omega}_{it}-\omega_{it})^2+ \dfrac{2}{T}\sum_{t=1}^{T}\omega_{it}(\hat{\omega}_{it}-\omega_{it})=a_1+a_2+a_3.\] For \(a_1\), we have \[a_1=\frac{1}{T}\sum_{t=1}^{T}(\Delta\epsilon_{it})^2+\frac{1}{T}\sum_{t=1}^{T}(\Delta r_{it})^2+\frac{2}{T}\Delta\epsilon_{it}\Delta r_{it}=a_{11}+a_{12}+a_{13}.\] Note that \(a_{11} =\frac{1}{T}\sum_{t=1}^T\epsilon_{it}^2 +(\bar{\epsilon}_{i\cdot})^2 +\frac{1}{T}\sum_{t=1}^T\bar{\epsilon}_{\cdot t}^2 +\bar{\epsilon}_{\cdot\cdot}^2 +\text{Cross Terms}\). For the single square terms, \(\operatorname{E}\left[\frac{1}{T}\sum_{t=1}^T\epsilon_{it}^2\right]=\sigma_{i}^2\), \(\operatorname{E}[\bar{\epsilon}_{i\cdot}^2]=\operatorname{E}\left[ \left( \frac{1}{T} \sum_{t=1}^T \epsilon_{it} \right)^2\right]=\frac{1}{T^2}\sum_{t=1}^T \operatorname{E}[\epsilon_{it}^2]+\frac{2}{T^2}\sum_{t<s}\operatorname{E}[\epsilon_{it}\epsilon_{is}]=\frac{\sigma_i^2}{T}=O(\frac{1}{T})\). Similarly, \(\operatorname{E}\left[\frac{1}{T}\sum_{t=1}^T\bar{\epsilon}_{\cdot t}^2\right]=O(\frac{1}{n})\) and \(\operatorname{E}\left[\bar{\epsilon}_{\cdot\cdot}^2\right]=O(\frac{1}{nT})\). By the Cauchy-Schwarz inequality, the cross terms are bounded by the product of the magnitudes. For example, the cross term with the overall mean is: \[\frac{1}{T}\sum_{t=1}^T\epsilon_{it} \bar{\epsilon}_{\cdot\cdot}= \bar{\epsilon}_{\cdot\cdot} (\frac{1}{T} \sum \epsilon_{it})=\bar{\epsilon}_{\cdot \cdot} \cdot \bar{\epsilon}_{i\cdot} =O_{p}(\frac{1}{\sqrt{nT}}) \cdot O_{p}(\frac{1}{\sqrt{T}})=O_{p}(\frac{1}{\sqrt{n}T}).\] Hence, \(a_{1}=\sigma_{i}^2+o_{p}(1)\). For \(a_{12}\) and \(a_{13}\), we can show that \(a_{12}=O_{p}(\ell_{n}^{-2\underline{\varsigma}})\) and \(|a_{13}|\leq 2\sqrt{a_{11}}\sqrt{a_{12}}=o_{p}(1)\).

Using the Mean Value Theorem, we have \(\hat{\omega}_{it}-\omega_{it}=\frac{\partial\omega_{it}}{\partial\theta'}\big|_{\theta=\bar\theta}(\hat{\theta}-\theta_0)\). Thus, \[a_{2}\leq \|\hat{\theta}-\theta_0\|^2\cdot \frac{1}{T}\sum_{t=1}^{T}\Big\Vert\frac{\partial\omega_{it}}{\partial\theta'}\big|_{\theta=\bar\theta}\Big\Vert^2=o_{p}(1)\cdot O(1)=o_{p}(1)\] by Assumption 7 and the consistency of \(\hat{\theta}\). We can show that \(|a_{3}|=o_{p}(1)\) By the Cauchy-Schwarz inequality. Combining all above, we have \(|\hat{\sigma}_i^2-\sigma_i^2|=o_{p}(1)\). \(\blacksquare\)

Lemma 5. Under Assumptions 1, 2, and 3-7 as \((n,T)\to\infty\), \(\|\hat{\Omega}_{t}-\Omega_{t}\|=o_{p}(1)\) where \(\tilde{\Omega}_{t}\) is evaluated by the consistent estimator \(\hat{\theta}\) with \(\|\hat{\theta}-\theta_0\|=o_{p}(1)\).

This result follows from the magnitudes of the order of \(\mathcal{J}_{i}\)’s with \(i=1,2\) in 28 . We can show that \(\|\tilde{\Omega}_{t}-\Omega_{t}\|=O_{p}(\|\hat{\theta}-\theta_0\|)+O_{p}(\frac{1}{n})+O_{p}(\frac{1}{T})\). The proof can be analogous to the proof of Lemma 4, and is available upon request. \(\blacksquare\)

Lemma 6. Under Assumptions 3-5, the best \(Q_{t}\) is given by equation 10 .

The intuition of the construction of \(Q_{t}\) is motivated by the first-order conditions of the GMM objective function. From 1 , we dentoe \[\begin{align} \label{eq:cA} A=B_{1}^{-1}(\gamma I_{n}+B_{2})=A_{y}+A_{r} \end{align}\tag{19}\] where \(A_{y}=S^{-1}(\gamma I_{n}+S_{2})\) and \(A_r=(B_{1}^{-1}-S_{1}^{-1})(\gamma I_{n}+S_{2})=-S_{1}^{-1}R_{1}B_{1}^{-1}(\gamma I_{n}+S_{2})\). It is straightforward to verify that \(A_x\) satisfies [propUB] and \(A_r\) satisfies [[propzero]](#propzero){reference-type=“ref” reference=“propzero”} \(O_{p}(\ell_{1}^{-\varsigma_{1}})\). For any \(h>0\), we have \[\begin{align} \label{eq:Yt43h} Y_{t+h}=A^{h+1}Y_{t-1}+\sum_{j=0}^{h}A^{j}B_{1}^{-1}(X_{t+h-j}\beta+\mathbf{c}_{n}+\alpha_{t+h-j}l_n)+\sum_{j=0}^{h}A^{j}B_{1}^{-1}\YEL{t+h-j} \end{align}\tag{20}\]

By denoting \[\begin{align} \begin{aligned} \chiL{t+h-j}&=\sum_{j=0}^{h}A^{j}B_{1}^{-1}(X_{t+h-j}\beta+\frac{1}{t-1}\sum_{h=1}^{t-1}\mathfrak{C}_{h}+\alpha_{t+h-j}l_n)\\ E_{t+h-j}&=\sum_{j=0}^{h}A^{j}B_{1}^{-1}\YEL{t+h-j}+\frac{1}{t-1}\sum_{h=0}^{t-1}\YEL{h} \end{aligned} \end{align}\] where \(\mathfrak{C}_{t}=\mathfrak{C}_{y,t}+\mathfrak{C}_{r,t}=B_{1}Y_{t}-(\gamma I_{n}+B_{2})Y_{t-1}-X_{nt}\beta-\alpha_{t}l_n\) with \(\mathfrak{C}_{y,t}=S_{1}Y_{t}-(\gamma I_{n}+S_{2})Y_{t-1}-X_{nt}\beta-\alpha_{t}l_n\) and \(\mathfrak{C}_{r,t}=R_{1}Y_{t}-R_{2}Y_{t-1}\). We then have \[\begin{align} \begin{aligned} Y_{t-1}^{(*,-1)}&=h_{Tt}(Y_{t-1}-\frac{1}{T-t}\sum_{s=t}^{T-1}Y_s)=h_{Tt}(Y_{t-1}-\frac{1}{T-t}\sum_{h=0}^{T-1-t}Y_h)\\ &=h_{Tt}\left(I_n-\frac{1}{T-t}\sum_{h=0}^{T-1-t}A^{h+1}\right)Y_{t-1}-h_{Tt}\frac{1}{T-t}\sum_{h=0}^{T-1-t}\chiL{t+h-j}-h_{Tt}\frac{1}{T-t}\sum_{h=0}^{T-1-t}E_{t+h-j} \end{aligned} \end{align}\] Since we use \(A_{y}\) to approximate \(A\) according to 19 , we then have \[\begin{align} \label{eq:ylagappro} Y_{t-1}^{(*,-1)}=h_{Tt}\left(I_n-\frac{1}{T-t}\sum_{h=0}^{T-1-t}A_{y}^{h+1}\right)Y_{t-1}-h_{Tt}\frac{1}{T-t}\sum_{h=0}^{T-1-t}\chiL{y,t+h-j}-h_{Tt}\frac{1}{T-t}\sum_{h=0}^{T-1-t}\mathscr{Q}_{t+h-j} \end{align}\tag{21}\] where \[\begin{align} \nonumber \mathscr{Q}_{t+h-j}&=&A_{r}^{h+1}Y_{t-1}+\chiL{r,t+h-j}+E_{t+h-j} \\ \nonumber \chiL{y,t+h-j}&=&\sum_{j=0}^{h}A_{y}^{j}S_{1}^{-1}(X_{t+h-j}\beta+\frac{1}{t-1}\sum_{h=1}^{t-1}\mathfrak{C}_{y,h}+\alpha_{t+h-j}l_n) \\ \nonumber \chiL{r,t+h-j}&=&\sum_{j=0}^{h}A_{r}^{j}B_{1}^{-1}(X_{t+h-j}\beta+\frac{1}{t-1}\sum_{h=1}^{t-1}\mathfrak{C}_{h}+\alpha_{t+h-j}l_n)+\sum_{j=0}^{h}A_{y}^{j}R_{1}^{-1}(X_{t+h-j}\beta+\frac{1}{t-1}\sum_{h=1}^{t-1}\mathfrak{C}_{h}+\alpha_{t+h-j}l_n) \\ \label{eq:Qt43h} &+&\sum_{j=0}^{h}A_{y}^{j}S_{1}^{-1}\frac{1}{t-1}\sum_{h=1}^{t-1}\mathfrak{C}_{r,h} \end{align}\tag{22}\] Denote \(\mathcal{I}_{t-1}\) be the sigma-algebra spanned by \((Y_{0},...,Y_{t-1})\) conditional on \((X_{1},...X_{T},\mathbf{c}_{0},\alpha_{1},...,\alpha_{T0})\) and \(\tilde{\mathscr{Q}}_{tT}=\frac{1}{T-t}\sum_{h=0}^{T-1-t}\mathscr{Q}_{t+h-j}\), from 21 , we have \[\begin{align} Y_{t-1}^{(*,-1)}=\mathbb{Y}_{t}-h_{Tt}\tilde{\mathscr{Q}}_{tT} \end{align}\] where \[\begin{align} \label{eq:bestylag} \mathbb{Y}_{t}\equiv\operatorname{E}[Y_{t-1}^{(*,-1)}|\mathcal{I}_{t-1}]=h_{Tt}\left(I_n-\frac{1}{T-t}\sum_{h=0}^{T-1-t}A_{y}^{h+1}\right)Y_{t-1}-h_{Tt}\frac{1}{T-t}\sum_{h=0}^{T-1-t}\chiL{y,t+h-j}. \end{align}\tag{23}\] We can show that \(\|\tilde{\mathscr{Q}}_{tT}\|=O_{p}(\sqrt{n}\ell_{n}^{-\underline{\varsigma}})\), which has the same asymptotic behavior as \(\|r_{t}\|\), and thus is asymptotically negligible with respect to the estimators. This is because \(\|A_{r}^{h+1}Y_{t-1}\|\leq\|A_{r}\|_{\mathrm{sp}}^{h+1}\|Y_{t-1}\|=O_{p}(\ell_{n}^{-(h+1)\underline{\varsigma}})O_{p}(\sqrt{n})=o_{p}(\sqrt{n}\ell_{n}^{-\underline{\varsigma}})\), \(\|\sum_{j=0}^{h}A_{r}^{j}\|_{\mathrm{sp}}=O_{p}(\ell_{n}^{-\underline{\varsigma}})\), \(\|\mathfrak{C}_{r,t}\|_{\mathrm{sp}}=O_{p}(\ell_{n}^{-\underline{\varsigma}})\) and \(\|X_{t}\beta+\mathbf{c}_{n}+\alpha_{t}l_n\|=O_{p}(\sqrt{n})\) by Assumptions 4 and 5.

Similarly, for some time-invariant \(n\times n\) matrix \(\mathcal{B}_{k}\) satisfying [propUB] we have \[\begin{align} \label{eq:bestBylag} \mathcal{B}_{k}\mathbb{Y}_{t}\equiv\operatorname{E}[\mathcal{B}_{k}Y_{t-1}^{(*,-1)}|\mathcal{I}_{t-1}]=h_{Tt}\mathcal{B}_{k}\bigl(I_n-\frac{1}{T-t}\sum_{h=0}^{T-1-t}A_{y}^{h+1}\bigr)Y_{t-1}-h_{Tt}\mathcal{B}_{k}\frac{1}{T-t}\sum_{h=0}^{T-1-t}\chiL{y,t+h-j} \end{align}\tag{24}\] Therefore, according to equation 25 , \(Q_{t}\) now is in 10 . \(\blacksquare\)

10.2 Proof of Theorems↩︎

Proof of Theorem 1. We establish the consistency and asymptotic normality of the OGMME in this section. The corresponding results for the BGMME follow analogously, by replacing \(\mathbf{J}_{N}\) with \(\ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\).
() From Lemma 17 in the supplement, the first-order condition at \(\theta_{0}\) times \(n(T-1)\) is approximated by22 \[\begin{align} \label{eq:FOC32of32theta} \begin{aligned} \frac{\partial m_{N}(\theta_{0})}{\partial\lambda_{1}} &\approx -\begin{pmatrix} \left(\mathbf{S}_{3N}\tfrac{\partial\mathbf{S}_{1N}}{\partial\lambda_{1}}\mathbf{Y}_{N}^{*}\right)'\mathbf{J}_{N}\mathbf{P}_{N1}^{s}\mathbf{J}_{N}\mathbf{V}_{N}^{*} \\ \vdots \\ \left(\mathbf{S}_{3N}\tfrac{\partial\mathbf{S}_{1N}}{\partial\lambda_{1}}\mathbf{Y}_{N}^{*}\right)'\mathbf{J}_{N}\mathbf{P}_{N\ell_{p}}^{s}\mathbf{J}_{N}\mathbf{V}_{N}^{*} \\ \mathbf{Q}_{N}'\mathbf{J}_{N}\mathbf{S}_{3N}\tfrac{\partial\mathbf{S}_{1N}}{\partial\lambda_{1}}\mathbf{Y}_{N}^{*} \end{pmatrix}, \frac{\partial m_{N}(\theta_{0})}{\partial\lambda_{3}} \approx \begin{pmatrix} \left(\tfrac{\partial\mathbf{S}_{3N}}{\partial\lambda_{3}}\mathbf{S}_{3N}^{-1}\mathbf{V}_{N}^{*}\right)'\mathbf{J}_{N}\mathbf{P}_{N1}^{s}\mathbf{J}_{N}\mathbf{V}_{N}^{*} \\ \vdots \\ \left(\tfrac{\partial\mathbf{S}_{3N}}{\partial\lambda_{3}}\mathbf{S}_{3N}^{-1}\mathbf{V}_{N}^{*}\right)'\mathbf{J}_{N}\mathbf{P}_{N\ell_{p}}^{s}\mathbf{J}_{N}\mathbf{V}_{N}^{*} \\ \mathbf{Q}_{N}'\mathbf{J}_{N}\tfrac{\partial\mathbf{S}_{3N}}{\partial\lambda_{3}}\mathbf{S}_{3N}^{-1}\mathbf{V}_{N}^{*} \end{pmatrix}, \\ \frac{\partial m_{N}(\theta_{0})}{\partial\lambda_{2}} &\approx -\begin{pmatrix} \left(\mathbf{S}_{3N}\tfrac{\partial\mathbf{S}_{2N}}{\partial\lambda_{2}}\mathbf{Y}_{N}^{(*,-1)}\right)'\mathbf{J}_{N}\mathbf{P}_{N1}^{s}\mathbf{J}_{N}\mathbf{V}_{N}^{*}\\ \vdots \\ \left(\mathbf{S}_{3N}\tfrac{\partial\mathbf{S}_{2N}}{\partial\lambda_{2}}\mathbf{Y}_{N}^{(*,-1)}\right)'\mathbf{J}_{N}\mathbf{P}_{N\ell_{p}}^{s}\mathbf{J}_{N}\mathbf{V}_{N}^{*}\\ \mathbf{Q}_{N}'\mathbf{J}_{N}\left(\mathbf{S}_{3N}\tfrac{\partial\mathbf{S}_{2N}}{\lambda_{2\ell_{2}}}\mathbf{Y}_{N}^{(*,-1)}\right) \end{pmatrix}, \frac{\partial m_{N}(\theta_{0})}{\partial\pi'}\approx -\begin{pmatrix} (\mathbf{S}_{3N}\mathbf{Z}_{N}^{*})'\mathbf{J}_{N}\mathbf{P}_{N1}^{s}\mathbf{J}_{N}\mathbf{V}_{N}^{*} \\ \vdots \\ (\mathbf{S}_{3N}\mathbf{Z}_{N}^{*})'\mathbf{J}_{N}\mathbf{P}_{N\ell_{p}}^{s}\mathbf{J}_{N}\mathbf{V}_{N}^{*} \\ \mathbf{Q}_{N}'\mathbf{J}_{N}\mathbf{S}_{3N}\mathbf{Z}_{N}^{*} \end{pmatrix} \end{aligned} \end{align}\tag{25}\] where \(\mathbf{Z}_{N}^{*}=[\mathbf{Y}_{N}^{(*,-1)},\mathbf{X}_{N}^{*}]\).

From the Taylor expansion, we have \[\begin{align} \label{eq:expansion32of32theta} \hat{\theta}_{gmm}-\theta_{0} = {}& -\Big[ \frac{\partial m_{N}^{\prime}(\hat{\theta}_{gmm})}{\partial \theta}\Omega_{N}^{-1} \frac{\partial m_{N}(\bar\theta^{\prime})}{\partial \theta}\Big]^{-1} \frac{\partial m_{N}^{\prime}(\hat{\theta}_{gmm})}{\partial \theta}\Omega_{N}^{-1} \frac{1}{n(T-1)}m_{N}(\theta_{0}) \end{align}\tag{26}\] and by Lemma 14 in the supplement, we have \(\frac{\partial m_{N}^{\prime}(\hat{\theta}_{gmm})}{\partial \theta}=D_{N}+O(\ell_{n}^{-\underline{\varsigma}})+o_p(1)\), where \[\begin{align} \label{eq:DN} D_{N}=-\frac{1}{n(T-1)}\operatorname{E}\begin{pmatrix} \mathbf{0}_{\ell_{p}\times\ell_{\pi}} & \mathbf{C}_{\ell_{p}\times\ell_{1}} & \mathbf{0}_{\ell_{p}\times\ell_{2}} & \mathbf{C}_{\ell_{p}\times\ell_{3}} \\ \mathbf{Q}_{N}^{\prime}\mathbf{J}_{N}\mathbf{S}_{3N}\mathbf{Z}_{N}^{*} & \mathbf{Q}_{N}^{\prime}\mathbf{J}_{N}\mathbf{S}_{3N}\tfrac{\partial\mathbf{S}_{1N}}{\lambda_{1\ell_{1}}}\mathbf{Y}_{N}^{*} & \mathbf{Q}_{N}^{\prime}\mathbf{J}_{N}\mathbf{S}_{3N}\tfrac{\partial\mathbf{S}_{2N}}{\lambda_{2\ell_{2}}}\mathbf{Y}_{N}^{(*,-1)} & \mathbf{0}_{\ell_{q}\times\ell_{3}} \end{pmatrix}, \end{align}\tag{27}\] and we denote \[ \begin{align} D_{\pi}=-\frac{1}{n(T-1)}\operatorname{E}\begin{pmatrix} \mathbf{0}_{\ell_{p}\times\ell_{\pi}}\\ \mathbf{Q}_{N}^{\prime}\mathbf{J}_{N}\mathbf{Z}_{N}^{*} \end{pmatrix} \;\text{and} \; D_{\lambda}=-\frac{1}{n(T-1)}\operatorname{E}\begin{pmatrix} \mathbf{C}_{\ell_{p}\times\ell_{1}} & \mathbf{0}_{\ell_{p}\times\ell_{2}} & \mathbf{C}_{\ell_{p}\times\ell_{3}} \\ \mathbf{Q}_{N}^{\prime}\mathbf{J}_{N}\mathbf{S}_{3N}\tfrac{\partial\mathbf{S}_{1N}}{\lambda_{1\ell_{1}}}\mathbf{Y}_{N}^{*} & \mathbf{Q}_{N}^{\prime}\mathbf{J}_{N}\mathbf{S}_{3N}\tfrac{\partial\mathbf{S}_{2N}}{\lambda_{2\ell_{2}}}\mathbf{Y}_{N}^{(*,-1)} & \mathbf{0}_{\ell_{q}\times\ell_{3}} \end{pmatrix}. \end{align}\] Furthermore, \([\mathbf{C}_{\ell_{p}\times\ell_{1}}]_{jk}= \operatorname{tr}\bigl(\mathbf{J}_{N}\mathbf{P}_{Nj}^{s}\mathbf{J}_{N}\mathbf{S}_{3N}\tfrac{\mathbf{S}_{1N}}{\partial\lambda_{1k}}\mathbf{S}_{1N}^{-1}\mathbf{S}_{3N}^{-1}\mathbf{\Sigma}_{N}\bigr)\) and \([\mathbf{C}_{\ell_{p}\times\ell_{3}}]_{jk}= \operatorname{tr}\bigl(\mathbf{J}_{N}\mathbf{P}_{Nj}^{s}\mathbf{J}_{N}\tfrac{\partial\mathbf{S}_{3N}}{\partial\lambda_{3k}}\mathbf{S}_{3N}^{-1}\mathbf{\Sigma}_{N}\bigr)\). Then we have \[\begin{align} \frac{\partial m_{N}^{\prime}(\hat{\theta}_{gmm})}{\partial \theta}\Omega_{N}^{-1}\frac{1}{n(T-1)}m_{N}(\theta_{0})= \frac{1}{n(T-1)}\mathbf{C}_{\ell_{p}\times\ell_{1}}'\mathbf{\Omega}_{1N}^{-1}\mathcal{J}_{1} +\frac{1}{n(T-1)}\mathbf{C}_{\ell_{p}\times\ell_{3}}'\mathbf{\Omega}_{1N}^{-1}\mathcal{J}_{1} +\frac{1}{n(T-1)}\mathcal{J}_{2}+o_{p}(1) \end{align}\] where \[\begin{align} \label{eq:remainder32term} \mathcal{J}_{1}=-\begin{pmatrix} \mathbf{V}_{N}^{*\prime}\mathbf{J}_{N}\mathbf{P}_{N1}^{s}\mathbf{J}_{N}\mathbf{V}_{N}^{*} \\ \vdots \\ \mathbf{V}_{N}^{*\prime}\mathbf{J}_{N}\mathbf{P}_{N\ell_{p}}^{s}\mathbf{J}_{N}\mathbf{V}_{N}^{*} \\ \mathbf{0}_{\ell_{q}\times 1} \end{pmatrix}, \;\text{and} \; \mathcal{J}_{2}=-\mathbf{L}_{N}^{*\prime}\mathbf{M}_{N}\mathbf{V}_{N}^{*}, \end{align}\tag{28}\] with \(\mathbf{L}_{N}^{*}=\big[ \mathbf{S}_{3N}\mathbf{Z}_{N}^{*}, \mathbf{S}_{3N}\tfrac{\partial\mathbf{S}_{1N}}{\partial\lambda_{1}}\mathbf{Y}_{N}^{*}, \mathbf{S}_{3N}\tfrac{\partial\mathbf{S}_{2N}}{\partial\lambda_{2}}\mathbf{Y}_{N}^{(*,-1)}, \mathbf{0}_{N\times \ell_{3}} \big]\) as defined in Notation 1(), \(\mathbf{M}_{N}=\mathtt{Diag}_{t=1}^{T-1}(M_{t})\) and \(M_{t}=J_{n}Q_{t}(Q_{t}'J_{n}\Sigma_{t}J_{n}Q_{t})^{-1}Q_{t}'J_{n}\).

Next, due to \(\mathbf{Y}_{N}^{(*,-1)'}\mathbf{M}_{N}\mathbf{E}_{N}^{*}=\sum_{t=1}^{T-1}\frac{T-t}{T-t+1}(Y_{t}-\frac{1}{T-t}\sum_{h=t}^{T-1}Y_{h})'M_{t}(E_{t}-\frac{1}{T-t}\sum_{h=t+1}^{T-1}E_{t})\), and we can deduce that the non-zero expectation part of is \(\sum_{t=1}^{T-1}\frac{1}{T-t+1}(\kappa_{1\gamma}+\kappa_{2\gamma})\), where \[\kappa_{1\gamma,t}=-E_{t}'M_{t}B_{3}(\sum_{h=1}^{T-t}A^{h-1})B_{1}^{-1}B_{3}^{-1}E_{t} \;\text{and} \; \kappa_{2\gamma,t}=\frac{1}{(T-t)}\big[\sum_{h=t+1}^{T-1}E_{h}'M_{t}(\sum_{s=1}^{T-h}A^{s-1})B_{1}^{-1}B_{3}^{-1}E_{t}\big].\] By the triangle inequality, we have \(\big\Vert\mathbf{Y}_{N}^{(*,-1)'}\mathbf{M}_{N}\mathbf{E}_{N}^{*}\big\Vert\leq\sum_{t=1}^{T-1}\frac{1}{T-t+1}\big(\|\kappa_{1\gamma}\|+\|\kappa_{2\gamma}\|\big).\) Note that \(\operatorname{E}\|\kappa_{1\gamma,t}\|=O\left(\ell_n\right) \;\text{and} \; \operatorname{E}\|\kappa_{2\gamma,t}\|=O\left(\ell_n\right)\) since \(\|M_{t}(\sum_{h=0}^{\infty}A^{h})B_{1}^{-1}B_{3}^{-1}\|_{\mathrm{rc}}=O_{p}(\ell_n)\) from Lemma 13 in the supplement. By Markov’s inequality and the fact that the magnitude order of \(\sum_{h=1}^{T-1}\frac{T-t}{T-t+1}\) is \(O(T-1)\), we have \(\Big\Vert\frac{1}{n(T-1)}\mathbf{Y}_{N}^{(*,-1)'}\mathbf{M}_{N}\mathbf{E}_{N}^{*}\Big\Vert=O_{p}\left(\ell_n\cdot\frac{\ln{T}}{n(T-1)}\right).\) Hence, we can deduce that \[\begin{align} \label{eq:order32of32Ylagr} \Big\Vert\frac{1}{n(T-1)}\mathbf{Y}_{N}^{(*,-1)\prime}\mathbf{M}_{N}\mathbf{r}_{N}^{*}\Big\Vert= O_p\left(\frac{n(T-1)\ell_{n}^{-\underline{\varsigma}}}{n(T-1)}\cdot\ell_n\right)= O_{p}\left(\ell_{n}^{1-\underline{\varsigma}}\right) \end{align}\tag{29}\] since the relationship between \(\mathbf{r}_{N}\) and \(\mathbf{E}_{N}\). Then, we have \[\begin{align} \begin{aligned} &\Big\Vert\frac{1}{n(T-1)}\left(\mathbf{S}_{3N}\tfrac{\partial\mathbf{S}_{2N}}{\partial\lambda_{2}}\mathbf{Y}_{N}^{(*,-1)}\right)^{\prime}\mathbf{M}_{N}\mathbf{E}_{N}^{*}\Big\Vert=O_{p}\left(\sqrt{\ell_{n}}\cdot\frac{\ell_n\ln{T}}{n(T-1)}\right)=O_{p}\left(\frac{\ell_{n}^{3/2}\ln{T}}{n(T-1)}\right) \;\text{and} \; \\ &\Big\Vert\frac{1}{n(T-1)}\left(\mathbf{S}_{3N}\tfrac{\partial\mathbf{S}_{2N}}{\partial\lambda_{2}}\mathbf{Y}_{N}^{(*,-1)}\right)^{\prime}\mathbf{M}_{N}\mathbf{r}_{N}^{*}\Big\Vert=O_{p}(\ell_{n})\cdot O_{p}(\sqrt{\ell_{n}})\cdot O_{p}(\ell_{n}^{-\underline{\varsigma}})=O_{p}\left(\ell_{n}^{3/2-\underline{\varsigma}}\right) \end{aligned} \end{align}\] since \(\|\tfrac{\partial\mathbf{S}_{2N}}{\partial\lambda_{2}}\|_{\mathrm{sp}}=O(\sqrt{\ell_{n}})\). For the last two terms in linear moments, we have \(\Big\Vert\frac{1}{n(T-1)}(\mathbf{S}_{3N}\tfrac{\partial\mathbf{S}_{1N}}{\partial\lambda_{1}}\mathbf{E}_{N}^{*})'\mathbf{M}_{N}\mathbf{E}_{N}^{*}\Big\Vert=O_{p}(\frac{\ell_{n}^{3/2}}{n})\), and \(\Big\Vert\frac{1}{n(T-1)}\mathbf{X}_{N}^{*\prime}\mathbf{M}_{N}\mathbf{E}_{N}^{*}\Big\Vert=O_{p}\left(\sqrt{\frac{\ell_{n}}{n(T-1)}}\right)\). Finally, from Lemma 14 in the supplement, for \(j=1,...,\ell_{p}\), we have \[\begin{align} \label{eq:order32of32VV} \Big\Vert\frac{1}{n(T-1)}\mathbf{r}_{N}^{*\prime}\JL{N}\PL{Nj}\JL{N}\mathbf{r}_{N}^{*}\Big\Vert=O_{p}(\ell_{n}^{-2\underline{\varsigma}}), \;\text{and} \; \Big\Vert\frac{1}{n(T-1)}\mathbf{r}_{N}^{*\prime}\JL{N}\PL{Nj}\JL{N}\mathbf{E}_{N}^{*}\Big\Vert=O_{p}(\ell_{n}^{-\underline{\varsigma}})=o_{p}(\ell_{n}^{1-\underline{\varsigma}}). \end{align}\tag{30}\] Thus, the bias of the quadratic moments is \(O_{p}(\ell_{n}^{-2\underline{\varsigma}}\cdot\ell_{n})=O_{p}(\ell_{n}^{1-2\underline{\varsigma}})=o_{p}(\ell_{n}^{3/2-\underline{\varsigma}}).\)

Consistency of \(\hat{\pi}_{gmm}\)↩︎

We use the \(C(\alpha)\) approach (see e.g. [25], [48]) to derive the expansion of \(\hat{\pi}_{gmm}\) and \(\hat{\lambda}_{gmm}\) from 26 . Define the projection matrix \(M_{\lambda}=I_{\ell_m}-D_{\lambda}(D_{\lambda}'\Omega_{N}^{-1}D_{\lambda})^{-1}D_{\lambda}'\Omega_{N}^{-1}\) and \(M_{\pi}=I_{\ell_m}-D_{\pi}(D_{\pi}'\Omega_{N}^{-1}D_{\pi})^{-1}D_{\pi}'\Omega_{N}^{-1}\) where \(\ell_m=\ell_{p}+\ell_{q}\). Then we have \[\begin{align} \tag{31} \hat{\pi}_{gmm}-\pi_0=-[D_{\pi}'\Omega_{N}^{-1}M_{\lambda}D_{\pi}]^{-1}D_{\pi}'\Omega_{N}^{-1}M_{\lambda}\frac{1}{n(T-1)}m_{N}(\theta_{0})+o_{p}(1) \\ \tag{32} \hat{\lambda}_{gmm}-\lambda_{0}=-[D_{\lambda}'\Omega_{N}^{-1}M_{\pi}D_{\lambda}]^{-1}D_{\lambda}'\Omega_{N}^{-1}M_{\pi}\frac{1}{n(T-1)}m_{N}(\theta_{0})+o_{p}(1) \end{align}\] Let \(\tilde{D}_{\lambda}=\Omega_{N}^{-\frac{1}{2}}D_{\lambda}\) and \(\tilde{D}_{\pi}=\Omega_{N}^{-\frac{1}{2}}D_{\pi}\), and define \(\tilde{M}_{\lambda}=\Omega_{N}^{-\frac{1}{2}}M_{\lambda}\Omega_{N}^{-\frac{1}{2}}=I_{\ell_m}-\tilde{D}_{\lambda}(\tilde{D}_{\lambda}'\tilde{D}_{\lambda})^{-1}\tilde{D}_{\lambda}'\) and \(\tilde{M}_{\pi}=\Omega_{N}^{-\frac{1}{2}}M_{\pi}\Omega_{N}^{-\frac{1}{2}}=I_{\ell_m}-\tilde{D}_{\pi}(\tilde{D}_{\pi}'\tilde{D}_{\pi})^{-1}\tilde{D}_{\pi}'\). Each \(\tilde{M}_{\bullet}\) is the orthogonal projector and thus \(\|\tilde{M}_{\bullet}\|_{\mathrm{sp}}=1\) where the subscript \(\bullet\) denotes \(\pi\) or \(\lambda\). Then, we have \(M_{\bullet}\leq \|\Omega_{N}^{1/2}\|_{\mathrm{sp}}\|\tilde{M}_{\bullet}\|_{\mathrm{sp}}\|\Omega_{N}^{1/2}\|_{\mathrm{sp}}=O(1)\). Define \[\begin{align} \label{eq:H95pi} K_{\pi}=-[D_{\pi}'\Omega_{N}^{-1}M_{\lambda}D_{\pi}]^{-1}D_{\pi}'\Omega_{N}^{-1}M_{\lambda} \;\text{and} \; K_{\lambda}=-[D_{\lambda}'\Omega_{N}^{-1}M_{\pi}D_{\lambda}]^{-1}D_{\lambda}'\Omega_{N}^{-1}M_{\pi}, \end{align}\tag{33}\] We have \(\|K_{\pi}\|_{\mathrm{sp}}\leq\Big\Vert[D_{\lambda}'\Omega_{N}^{-1}M_{\pi}D_{\lambda}]^{-1}\Big\Vert_{\mathrm{sp}}\Big\Vert D_{\lambda}'\Omega_{N}^{-1}M_{\pi}\Big\Vert_{\mathrm{sp}}=O(1)\) by Assumption 7. Thus, \[\begin{align} \label{eq:pi95og95lln} \begin{aligned} \hat{\pi}_{gmm}-\pi_0=K_{\pi}\frac{1}{n(T-1)}m_{N}(\theta_{0})=K_{\pi}\frac{1}{n(T-1)}m_{\Delta}+O_{p}(\|\mathcal{J}_{1}\|)+O_{p}(\|\mathcal{J}_{2}\|) \end{aligned} \end{align}\tag{34}\] where \(\mathcal{J}_{i}\)’s are in 28 for \(i=1,2\), and \(m_{\Delta}\) is in ?? . Therefore, combining 29 30 , we have \[\begin{align} \label{eq:rate32of32pi} \hat{\pi}_{gmm}-\pi_0=O_{p}\Bigl(\ell_{n}^{3/2-\underline{\varsigma}}+\frac{\ell_{n}^{3/2}}{n}+\frac{\ell_{n}^{3/2}\ln{T}}{n(T-1)}+\sqrt{\frac{\ell_{n}}{n(T-1)}}\Bigr)=O_{p}\Bigl(\ell_{n}^{3/2-\underline{\varsigma}}+\frac{\ell_{n}^{3/2}}{n}+\sqrt{\frac{\ell_{n}}{n(T-1)}}\Bigr). \end{align}\tag{35}\] () For the limiting distribution, define \[\begin{align} \label{eq:Var95pi} \mathbf{\Sigma}_{\pi_0,gmm}=\lim_{n,T\to\infty}K_{\pi}\Omega_{N}K_{\pi}', \end{align}\tag{36}\] and then \(\|\mathbf{\Sigma}_{\pi_0,gmm}\|_{\mathrm{sp}}\leq\|K_{\pi}\|_{\mathrm{sp}}^2\|\Omega_{N}\|_{\mathrm{sp}}=O(1)\). By 34 , we have \[\begin{align} \label{eq:piog95clt} \begin{aligned} \sqrt{n(T-1)}\mathbf{\Sigma}_{\pi_0,gmm}^{-1/2}(\hat{\pi}_{gmm}-\pi_0)&=\mathbf{\Sigma}_{\pi_0,gmm}^{-1/2}K_{\pi}\frac{1}{\sqrt{n(T-1)}}m_{\Delta}+O_{p}\left(\sqrt{n(T-1)}\Bigl(\ell_{n}^{3/2-\underline{\varsigma}}+\frac{\ell_{n}^{3/2}}{n}\Bigr)\right). \end{aligned} \end{align}\tag{37}\] It remain to show that \(\sqrt{n(T-1)}\mathbf{\Sigma}_{\pi_0,gmm}^{-1/2}(\hat{\pi}_{gmm}-\pi_0)\stackrel{d}{\to}N(0,I_{\ell_{\pi}})\). The results are directly obtained from Lemma 3 and the Cramér–Wold device.

Proof of Theorem 2. For (), define \(\Sigma_{\pi,N}=K_{\pi}\Omega_{N}K_{\pi}'\), then \(\boldsymbol{\Sigma}_{\pi_0,gmm}=\lim_{n,T\to\infty}\Sigma_{\pi,N}\). It suffices to show (a) \(\|\hat{\boldsymbol{\Sigma}}_{\pi,gmm}-\Sigma_{\pi,N}\|_{\mathrm{sp}}=o_{p}(1)\) and (b) \(\|\Sigma_{\pi,N}-\boldsymbol{\Sigma}_{\pi_0,gmm}\|_{\mathrm{sp}}=o_{p}(1)\).

For (a), note that \[\hat{\boldsymbol{\Sigma}}_{\pi,gmm}-\Sigma_{\pi,N} = (\hat{H}_{\pi}-K_\pi)\hat{\Omega}_{N}\hat{H}_{\pi}^{\prime} + K_\pi(\hat{\Omega}_{N} -\Omega_{N})\hat{H}_{\pi}^{\prime} + K_\pi\Omega_{N}(\hat{H}_{\pi}^{\prime} -K_{\pi}^{\prime}).\] Taking spectral norms and using submultiplicativity yields \[\|\hat{\boldsymbol{\Sigma}}_{\pi,gmm}-\Sigma_{\pi,N}\|_{\mathrm{sp}} \leq \|\hat{H}_{\pi}-K_\pi\|_{\mathrm{sp}} \|\hat{\Omega}_{N}\|_{\mathrm{sp}}\|\hat{H}_{\pi}\|_{\mathrm{sp}} +\|K_\pi\|_{\mathrm{sp}}\|\hat{\Omega}_{N}-\Omega_{N}\|_{\mathrm{sp}}\|\hat{H}_{\pi}\|_{\mathrm{sp}} +\|K_\pi\|_{\mathrm{sp}}\|\Omega_{N}\|_{\mathrm{sp}}\|\hat{H}_{\pi}-K_\pi\|_{\mathrm{sp}}.\] Under Assumption 7, \(K_\pi\) and \(\Omega_{N}\) are \(O(1)\). By Lemma 2\(, \|\hat{D}_{\pi}-D_{\pi}\|_{\mathrm{sp}}=o_{p}(1)\) and \(\|\hat{D}_{\lambda}-D_{\lambda}\|_{\mathrm{sp}}=o_{p}(1)\). Therefore, \(\|\hat{H}_{\pi}-K_\pi\|_{\mathrm{sp}}=o_{p}(1)\) since \(\|\hat{\Omega}_{N}-\Omega_{N}\|_{\mathrm{sp}}=o_{p}(1)\) by Lemma 5, which implies \(\|\hat{\boldsymbol{\Sigma}}_{\pi,gmm}-\Sigma_{\pi,N}\|_{\mathrm{sp}}=o_{p}(1)\). For (b), it holds by the definition \(\boldsymbol{\Sigma}_{\pi_0,gmm}=\lim_{n,T\to\infty}K_\pi\Omega_{N}K_{\pi}^{\prime}\). Combining (a) and (b) completes the proof.

For (), denote \(\mathbf{\Sigma}_{b_\pi}=\min\mathbf{\Sigma}_{\pi_0,gmm}=\min\{\lim_{n,T\to\infty}K_{\pi}'\Omega_{N}K_{\pi}\}\). The equation 36 can be equivalently written as \(\mathbf{\Sigma}_{\pi_0,gmm}=\mathtt{e}_{\pi}(\lim_{n,T\to\infty}D_{N}'\Omega_{N}^{-1}D_{N})^{-1}\mathtt{e}_{\pi}'\), where \(\mathtt{e}_{\pi}\) is the \(\ell_{\pi}\times\ell_{\theta}\) selection matrix that extracts the first \(\ell_{\pi}\) components, i.e., \(\mathtt{e}_{\pi}\theta=\pi\). In the case of BGMME, \(\mathbf{J}_{N}\) is replaced by \(\ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\), and accordingly, \(\mathbf{M}_{N}= \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{Q}_{N}(\mathbf{Q}_{N}' \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{Q}_{N})^{-1}\mathbf{Q}_{N}' \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\). For the quadratic moments, we have \(\ell_{p}=\ell_*=\ell_1+\ell_3\), and the components in \(\mathbf{C}_{\ell_*\times\ell_{1}}\) and \(\mathbf{C}_{\ell_*\times\ell_{3}}\), are \([\mathbf{C}_{\ell_*\times\ell_{1}}]_{jk}= \operatorname{tr}\bigl( \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{P}_{Nj}^{s} \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{S}_{3N}\tfrac{\mathbf{S}_{1N}}{\partial\lambda_{1k}}\mathbf{S}_{1N}^{-1}\mathbf{S}_{3N}^{-1}\mathbf{\Sigma}_{N}\bigr)\) and \([\mathbf{C}_{\ell_*\times\ell_{3}}]_{jk}= \operatorname{tr}\bigl( \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{P}_{Nj}^{s} \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\tfrac{\partial\mathbf{S}_{3N}}{\partial\lambda_{3k}}\mathbf{S}_{3N}^{-1}\mathbf{\Sigma}_{N}\bigr)\), respectively. Note that

\[\begin{align} [\mathbf{C}_{\ell_*\times\ell_{1}}]_{jk} =&\operatorname{tr}\bigl((\mathbf{\Sigma}_{N}\mathbf{S}_{3N}\tfrac{\mathbf{S}_{1N}}{\partial\lambda_{1k}}\mathbf{S}_{1N}^{-1}\mathbf{S}_{3N}^{-1})^{\prime} \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{P}_{Nj}^{s} \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\bigr)\\ =&\frac{1}{2}\operatorname{tr}\bigl([ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{\frac{1}{2}}(\mathbf{\Sigma}_{N}\mathbf{S}_{3N}\tfrac{\mathbf{S}_{1N}}{\partial\lambda_{1k}}\mathbf{S}_{1N}^{-1}\mathbf{S}_{3N}^{-1})^{s}[ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{\frac{1}{2}} [ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{\frac{1}{2}}\mathbf{P}_{Nj}^{s}[ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{\frac{1}{2}}\bigr)\\ =&\left\{\frac{1}{\sqrt{2}}\mathtt{vec}'\Bigl([ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{\frac{1}{2}}(\mathbf{\Sigma}_{N}\mathbf{S}_{3N}\tfrac{\mathbf{S}_{1N}}{\partial\lambda_{1k}}\mathbf{S}_{1N}^{-1}\mathbf{S}_{3N}^{-1})^{s}[ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{\frac{1}{2}}\Bigr)\right\} \left\{\frac{1}{\sqrt{2}}\mathtt{vec}\Bigl( [ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{\frac{1}{2}}\mathbf{P}_{Nj}^{s}[ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{\frac{1}{2}}\Bigr)\right\} \\ =&a_{1j}'a_{2k} \end{align}\] by using the symmetry of \(\mathbf{P}_{Nj}^{s}\) and \(\mathbf{\Sigma}_{N}\), together with the identity \(\operatorname{tr}(A)=1/2\operatorname{tr}(A^s)\), and \(\operatorname{tr}(AB)=\mathtt{vec}'(A)\mathtt{vec}(B)\) for any conformable matrices \(A\) and \(B\), where \(\mathtt{vec}(\cdot)\) stacks all entries of a matrix into a column vector. Similarly, \([\mathbf{C}_{\ell_*\times\ell_{3}}]_{jk} =a_{3j}'a_{2k}\), where \(a_{3j}=\frac{1}{\sqrt{2}}\mathtt{vec}\bigl([ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{\frac{1}{2}}(\mathbf{\Sigma}_{N}\tfrac{\mathbf{S}_{3N}}{\partial\lambda_{3k}}\mathbf{S}_{3N}^{-1})^{s}[ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{\frac{1}{2}}\bigr)\) and the components in \(\mathbf{\Omega}_{N1}\) is \[[\mathbf{\Omega}_{N1}]_{ij}=\operatorname{tr}( \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{P}_{Ni} \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{P}_{Nj}^{s} \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) })=1/2\operatorname{tr}( [ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{\frac{1}{2}}\mathbf{P}_{Nj}^{s}[ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{\frac{1}{2}} [ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{\frac{1}{2}}\mathbf{P}_{Nj}^{s}[ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{\frac{1}{2}})=a_{2j}'a_{2j}.\] Hence, \(D_{N}'\Omega_{N}^{-1}D_{N}=(\operatorname{E}\Sigma_{\texttt{line}}+\operatorname{E}\Sigma_{\texttt{quad}})\), where \(\Sigma_{\texttt{line}}=\lim_{n,T\to\infty}\frac{1}{n(T-1)}\mathbf{L}_{N}^{*'}\mathbf{M}_{N}\mathbf{L}_{N}^{*} =\lim_{n,T\to\infty}\frac{1}{n(T-1)}\mathbf{L}_{N}^{*'}[ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{1/2}[ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{1/2}\mathbf{Q}_{N}(\mathbf{Q}_{N}' \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{Q}_{N})^{-1}\mathbf{Q}_{N}'[ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{1/2}[ \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }]^{1/2}\mathbf{L}_{N}^{*} \leq \lim_{n,T\to\infty}\frac{1}{n(T-1)}\mathbf{L}_{N}^{*'} \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{L}_{N}^{*}\) and \(\Sigma_{\texttt{quad}}=\lim_{n,T\to\infty}\frac{1}{n(T-1)} \kappa_\lambda'\kappa_\delta(\kappa_\delta'\kappa_\delta)^{-1}\kappa_\delta'\kappa_\lambda \leq \lim_{n,T\to\infty}\frac{1}{n(T-1)}\kappa_{\lambda}'\kappa_{\lambda}\) by the Generalized Cauchy-Schwarz inequality, where \(\kappa_\lambda=(a_{11},...,a_{1\ell_1},a_{31},...,a_{3\ell_3})'\) and \(\kappa_\delta=(a_{21},...,a_{2\ell_{*}})'\). Thus, \[\begin{align} \label{eq:Var95bg95pi} \mathbf{\Sigma}_{\pi_0,bgmm}=\mathbf{\Sigma}_{b_\pi}=\mathtt{e}_{\pi}\left(\lim_{n,T\to\infty}\mathbf{\Sigma}_{\theta}\right)^{-1}\mathtt{e}_{\pi}' \end{align}\tag{38}\] where \(\mathbf{\Sigma}_{\theta}=\operatorname{E}\Bigl(\frac{1}{n(T-1)}\mathbf{L}_{N}^{*'} \ifstrequal{N}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{N}}) }{ {J}({\Sigma_{N}}) }\mathbf{L}_{N}^{*}+\mathtt{Diag}(\kappa_{\theta}'\kappa_{\theta})\Bigr)\) with \(\kappa_\theta= (\mathbf{0}_{\ell_\pi\times 1},a_{11},\dots, a_{1\ell_1}, \mathbf{0}_{\ell_{2}\times 1}, a_{31}, \dots, a_{3\ell_3})'\). Then, we can evaluate \(\hat{\mathbf{\Sigma}}_{b_\pi}\) by \(\pi_{bgmm}\). Its consistency follows by similar arguments to those presented in ().

Proof of Theorem 3. Because \(|\hat{g}_{k,gmm}(d)-g_{k0}(d)|\leq|\hat{g}_{k,gmm}(d)-\xi_{k0}(d)|+|g_{k0}(d)-\xi_{k0}(d)|\). Consequently, taking the supremum over \(d\in D\) yields \[\begin{align} \sup_{d\in D}|\hat{g}_{k,gmm}(d)-g_{k0}(d)|\leq \sup_{d\in D}\|\pmb{\phi}_{k}(d)\|_{\mathrm{sp}}\|\hat{\lambda}_{k,gmm}-\lambda_{k_0}\|+O_{p}(\ell_{n}^{-\underline{\varsigma}}). \end{align}\] According to 32 and 33 , we know that \[\begin{align} \label{eq:lambda95og95lln} \begin{aligned} \hat{\lambda}_{k,gmm}-\lambda_{k_0}=K_{\lambda}\frac{1}{n(T-1)}(\mathcal{J}_{1}+\mathcal{J}_{2}), \end{aligned} \end{align}\tag{39}\] combining Assumption 2, we have \[\begin{align} \label{eq:hatg} \begin{aligned} \sup_{d\in D}|\hat{g}_{k,gmm}(d)-g_{k0}(d)|&=O_{p}\Bigl(\ell_{n}^{1/2}\Bigl(\ell_{n}^{3/2-\underline{\varsigma}}+\frac{\ell_{n}^{3/2}}{n}+\sqrt{\frac{\ell_{n}}{n(T-1)}}\Bigr)+\ell_{n}^{-\underline{\varsigma}}\Bigr) \\&= O_{p}\Bigl(\ell_{n}^{2-\underline{\varsigma}}+\frac{\ell_{n}^{2}}{n}+\frac{\ell_{n}}{\sqrt{n(T-1)}}\Bigr). \end{aligned} \end{align}\tag{40}\] Next, let \[\begin{align} \label{eq:Var95g} \Sigma_{g_{k0},gmm} = \lim_{n,T \to \infty} \frac{1}{\ell_{n}} \phi_{k}'(d) \Sigma_{\pi_{0},gmm} \phi_{k}(d). \end{align}\tag{41}\] Using Assumption 2(), we bound the spectral norm as: \[\begin{align} \|\mathbf{\Sigma}_{g_{k0},gmm}\|_{\mathrm{sp}}\leq \dfrac{1}{\ell_{n}}\|\pmb{\phi}_{k}\|_{\mathrm{sp}}^2\|\mathbf{\Sigma}_{\pi_0,gmm}\|_{\mathrm{sp}}=\ell_{n}^{-1}\cdot O(\ell_{n})=O(1). \end{align}\] Combing equations 39 and 40 , the sufficient condition for the CLT is \[\begin{align} \sqrt{\frac{(T-1)\ell_{n}^{3}}{n}}+\sqrt{n(T-1)}\ell_{n}^{3/2-\underline{\varsigma}}\to 0 \end{align}\] as \((n,T)\to\infty\).

The sufficient rate conditions differ across estimands. Theorem 1(i) establishes consistency of the finite-dimensional parameter \(\hat{\pi}_{gmm}\) using the expansion in 31 evaluated at \(\theta_{0}\), and thus only requires \(\ell_{n}^{3/2-\underline{\varsigma}}+\ell_{n}^{3/2}/n\to0\), together with \(\sqrt{\tfrac{\ell_{n}}{n(T-1)}}\to0\) as it appears in Assumption 2(). In contrast, Lemma 2(ii) imposes the stronger condition \(\ell_{n}^{2-\underline{\varsigma}}+\tfrac{\ell_{n}^{2}}{n}+\tfrac{\ell_{n}}{\sqrt{n(T-1)}}\to0\) to guarantee uniform convergence of \(\Theta_N=\Gamma\times\Lambda_N\), which is needed for joint consistency of \(\hat{\theta}=(\hat{\pi}',\hat{\lambda}')'\). Theorem 3(i) requires the same stronger condition to obtain uniform consistency of \(\hat{g}_{k,gmm}(d)\) over \(d\in D\). The gap reflects the additional difficulty of controlling stochastic fluctuations uniformly over the growing-dimensional sieve space.

Supplementart material↩︎

A Supplement to
Semi-nonparametric estimation of spatial dynamic panel data models with nonparametric spatial weights

11 Additional lemmas↩︎

Lemma 7. Let \(A\) and \(B\) be \(n \times n\) real matrices, then \(\|AB\|\leq\|A\|_{\mathrm{sp}}\|B\|\) or \(\|AB\|\leq\|B\|_{\mathrm{sp}}\|A\|\).

This is a standard matrix norm inequality. \(\blacksquare\)

Lemma 8. () For any time-varying square matrix \(\mathcal{B}_{t}\), if \(\sup\limits_{t}\|\mathcal{B}_{t}\|_{1}=\sup\limits_{t}\|\mathcal{B}_{t}\|_{\infty}=O(1)\), then \(\sup\limits_{t}\|\mathcal{B}_{t}\|_{\mathrm{sp}}=O(1)\).
() For any time-varying square random matrix \(\mathcal{B}_{t}\), depending on an index \(n\). If \(\sup_{t}\|\mathcal{B}_{t}\|_{1}=O_{p}(1)\) and \(\sup_{t} \|\mathcal{B}_{t}\|_{\infty} = O_{p}1)\), then \(\sup_{t} \|\mathcal{B}_{t}\|_{\mathrm{sp}} = O_{p}(1)\).

() We use the standard matrix norm inequality \(\|\mathcal{B}_{t}\|_{\mathrm{sp}} \leq \sqrt{\|\mathcal{B}_{t}\|_{1} \|\mathcal{B}_{t}\|_{\infty}}\). Taking the supremum over \(t\) yields: \(\sup_{t} \|\mathcal{B}_{t}\|_{\mathrm{sp}} \leq \sup_{t} \sqrt{\|\mathcal{B}_{t}\|_{1} \|\mathcal{B}_{t}\|_{\infty}} \leq \sqrt{\left( \sup_{t} \|\mathcal{B}_{t}\|_{1} \right) \left( \sup_{t} \|\mathcal{B}_{t}\|_{\infty} \right)}=O(1).\)
() According to the proof of (), let \(X = \sup_{t} \|\mathcal{B}_{t}\|_{1}\) and \(Y = \sup_{t} \|\mathcal{B}_{t}\|_{\infty}\). Then, \(Z = XY = O_p(1)\). Since norms are non-negative, \(Z \ge 0\). Therefore, \(\sqrt{Z} = \sqrt{XY} = O_p(1)\). Hence, \(0 \leq \sup_{t} \|\mathcal{B}_{t}\|_{\mathrm{sp}} \leq \sqrt{XY}= O_p(1)\). \(\blacksquare\)

Lemma 9. Let the \(n \times n\) matrix \(A\) with its typical element \(a_{ij}\) satisfy \(\sup_{i,j}|a_{ij}|=O_{p}(d^{-\nu})\) and \(\sup_{i}\sum_{j=1}^{n}|a_{ij}|=O_{p}(d^{-\nu})\) for some \(\nu > 0\). Suppose that \(n^{1/2}d^{-\nu} \to 0\) as \(n \to\infty\), then: \[\|e^A - I\|=O_{p}(\|A\|)=O_{p}(n^{1/2}d^{-\nu}).\]

First, it is easy to conclude that \(\|A\|=O_{p}(n^{1/2}d^{-\nu})\) since \(\sum_{j=1}^{n}|a_{ij}|^2 \leq \left(\sup_{k}|a_{ik}| \right) \left(\sum_{j=1}^{n}|a_{ij}|\right)\leq O_{p}(d^{-2\nu})\). Therefore, \(\|A\|^2=O_{p}(n d^{-2\nu})\). Second, Using the triangle inequality for the Frobenius norm and the submultiplicative property, we get: \(\|e^A - I\|\leq\|A\|+\frac{\|A^2\|}{2!}+\frac{\|A^3\|}{3!}+\cdots\). Since \(\|A^k\| \leq \|A\|^k\), this becomes: \(\|e^A-I\|\leq\|A\|+\frac{\|A\|^2}{2!}+\frac{\|A\|^3}{3!}+\cdots=e^{\|A\|}-1\). Since \(\|A^k\| \leq \|A\|^k\), this becomes: \(\|e^A - I\|\leq\|A\|+\frac{\|A\|^2}{2!} + \frac{\|A\|^3}{3!}+\cdots=e^{\|A\|}-1\). Let \(X_p=\|A\|\), we have \(e^{X_p} - 1 = X_p + o_{p}(1)\). Combining above all, we conclude that \(\|e^A - I\|=O_{p}(\|A\|)\). \(\blacksquare\)

Lemma 9 shows that the MESS-type has the same asymptotic behavior as that of the SAR-type.

Lemma 10. When \(T \to\infty\), \(\sum_{t=1}^{T-1}h_{Tt}^2=O(T)\) and \(\sum_{t=1}^{T-1}(1-h_{Tt}^2)=O(\ln{T})\), where \(h_{Tt}=\sqrt{\frac{T-t}{T-t+1}}\).

Note that \(h_{Tt}^2=1-\frac{1}{T-t+1}\), and we have \(\sum_{t=1}^{T-1}h_{Tt}^2=O(T)-O(\ln{T})=O(T)\) since \(\sum_{t=1}^{T-1}\frac{1}{T-t+1}=O(\ln{T})\). \(\blacksquare\)

Lemma 11. Under Assumption 4 and \(\ell_{k}^{-\varsigma_{k}}+n^{1/2}\ell_{k}^{-\varsigma_{k}} \to 0\) as \(n \to\infty\):
() \(e^{\Delta_{k}}-I_{n}\) satisfies [[propzero]](#propzero){reference-type=“ref” reference=“propzero”} \(O_{p}(\ell_{k}^{-\varsigma_{k}})\), where \(\Delta_{k}\) is defined in Notation 1(). Also, \(R_{k}\) and \(S_{k}R_{k}\) satisfy [[propzero]](#propzero){reference-type=“ref” reference=“propzero”} \(O_{p}(\ell_{k}^{-\varsigma_{k}})\), and \(B_{k}\) and \(S_{k}\) satisfy [propUB];
() \(H_k\) satisfies [[propzero]](#propzero){reference-type=“ref” reference=“propzero”} \(O_{p}(\ell_{3}^{-\varsigma_{3}})\), where \(H_{k}=R_{k}B_{k}^{-1}\);
() \(\operatorname{E}\|r_{kt}^{*}\|=O(n^{1/2}\ell_{k}^{-\varsigma_{k}}h_{Tt})\) where \(h_{Tt}=\sqrt{\frac{T-t}{T-t+1}}\), \(r_{kt}^{*}=h_{Tt}(r_{kt}-\frac{1}{T-t}\sum_{h=t+1}^{T}r_{kt})\) and \(r_{kt}\) is defined in equation [eq:Vt].

For (), first, from Lemma 9, we can derive that \(\|e^{\Delta_{k}}-I_{n}\|=O_{p}(\|\Delta_{k}\|)=O_{p}(n^{1/2}\ell_{k}^{-\varsigma_{k}})\). Second, \(\Vert S_{k}\Vert_{\mathrm{rc}}\leq e^{\sum_{p_{k}=1}^{\ell_{k}}|\lambda_{p_{k}}|\Vert\varPhi_{kp_{k}}\Vert_{\mathrm{rc}}}=O_{p}(1)\) from Assumption 4 () and () and thus \(\|S_{k}\|_{\mathrm{sp}}=O_{p}(1)\). Third, since \(R_{k}=S_{k}(e^{\Delta_{k}}-I_{n})\), by Lemmas 7 and 9, we have \[\|R_{g_{k}}\|\leq\|S_{k}\|_{\mathrm{sp}}\|e^{\Delta_{k}}-I_{n}\|=O_{p}(1)O_{p}(\|\Delta_{k}\|)=O_{p}(n^{1/2}\ell_{k}^{-\varsigma_{k}}).\] The results of \(\|R_{k}\|_{\mathrm{rc}}=O_{p}(\ell_{k}^{-\varsigma_{k}})\) since \(\|e^{\Delta_{k}}-I_{n}\|_{\mathrm{rc}}=O_{p}(\|\Delta_{k}\|_{\mathrm{rc}})=O_{p}(\ell_{k}^{-\varsigma_{k}})\). Fourth, we have \[\|S_{k}R_{g_{k}}\|\leq\|S_{k}\|_{\mathrm{sp}}\|R_{g_{k}}\|=O_{p}(n^{1/2}\ell_{k}^{-\varsigma_{k}}).\] For (), first, combining 1 and 2 , we have \[\begin{align} \label{eq:Ut4432Et32and32RS} U_{t}=B_{3}^{-1}E_{t}=(R_{3}+S_{3})^{-1}E_{t} \end{align}\tag{42}\] When \(k=3\), we have \(r_{3t}=R_{3}U_{t}=H_{3}E_{t}\), where \(H_{3}=R_{3}B_{3}^{-1}\). Since \(B_{k}\), \(S_{k}\) and \(R_{k}\) are invertible, we have \(B_{3}^{-1}=(I_{n}+S_{3}^{-1}R_{3})^{-1}S_{3}^{-1}\). Then, we have \(\|S_{k}^{-1}R_{k}\|_{\mathrm{sp}}=O_{p}(n^{1/2}\ell_{k}^{-\varsigma_{k}})\) which implies \(\ell_{k}^{\varsigma_{k}}=o(n^{1/2})\). Thus, by the Neumann series, we have \(\|(I_{n}+S_{3}^{-1}R_{3})^{-1}\|_{\mathrm{sp}}=O(1)\), \(\|B_{3}^{-1}\|_{\mathrm{sp}}=O(1)\) and \(\|H_{3}\|_{\mathrm{sp}}=O_{p}(\ell_{k}^{-\varsigma_{3}})\). For (), consider \(k=3\), and \(r_{3t}^{*}=h_{Tt}(r_{3t}-\frac{1}{T-t}\sum_{h=t+1}^{T}r_{3t}).\) Note that \[\begin{align} \begin{aligned} \operatorname{E}\|r_{3t}\|^2& =\operatorname{E}\sum_{i=1}^{n}r_{3t,i}^2=\operatorname{E}\sum_{i=1}^{n}(\sum_{j=1}^{n}h_{3,ij}\epsilon_{jt})^2 =\operatorname{E}\sum_{i=1}^{n}\sum_{j=1}^{n}\sum_{k=1}^{n}h_{3,ij}h_{3,ik}\epsilon_{jt}\epsilon_{kt} \\ &=\operatorname{E}\sum_{i=1}^{n}\sum_{j=1}^{n}h_{3,ij}^2\epsilon_{jt}^2 \leq\sqrt{\operatorname{E}\left[\sum_{i=1}^{n}(\sum_{j=1}^{n}h_{3,ij}^2)^2\right]}\sqrt{\operatorname{E}\left[\sum_{j=1}^{n}\epsilon_{j}^4\right]}\\ &=O(\sqrt{n}\ell_{3}^{-2\varsigma_{3}})\cdot O(\sqrt{n}) =O(n\ell_{3}^{-2\varsigma_{3}}), \end{aligned} \end{align}\] by Cauchy-Schwartz inequality and the fact that \(\sum_{j=1}^{n}h_{3,ij}^2\leq(\sum_{j=1}^{n}|h_{ij}|)^2\leq \|H_{3}\|_{\infty}^2=O_{p}(\ell_{3}^{-2\varsigma_{3}})\). Thus, \(\operatorname{E}\|r_{3t}\|\leq\sqrt{\operatorname{E}\|r_{3t}\|^2}=O(n^{1/2}\ell_{k}^{-\varsigma_{3}})\). So we have \(\operatorname{E}\|\frac{1}{T-t}\sum_{h=t+1}^{T}r_{3t}\|\leq\frac{1}{T-t}\sum_{h=t+1}^{T}\operatorname{E}\|r_{3t}\|=O(n^{1/2}\ell_{k}^{-\varsigma_{3}}).\) Finally, \(\operatorname{E}\|r_{3t}^{*}\|=O(n^{1/2}\ell_{3}^{-\varsigma_{3}}h_{Tt})\). Similarly, we have \(\operatorname{E}\|r_{1t}^{*}\|=O(n^{1/2}\ell_{1}^{-\varsigma_{1}}h_{Tt})\) and \(\operatorname{E}\|r_{2t}^{*}\|=O(n^{1/2}\ell_{2}^{-\varsigma_{2}}h_{Tt})\) as \(\rho(A)<1\). \(\blacksquare\)

Lemma 12. Under Assumption 2(), the covariance of \(m_{N}^{\mathtt{line}}(\theta_{0})\) and \(m_{N}^{\mathtt{quad}}(\theta_{0})\) is asymptotically zero as \((n,T)\to\infty\).

Similar to the Proof of Theorem 1. \(\blacksquare\)

Lemma 13. Under Assumption 6, \(\|M_{t}\|_{\mathrm{rc}}=O_{p}\left(\ell_n\right)\), and \(\|M_{t}\|=O_{p}\left(\sqrt{\ell_n}\right)\).

Note that the entries of \(M_{t}\) are \[m_{ijt}=\tfrac{1}{n}\Bigl[J_{n}Q_{t}\Bigr]_{i}'\Bigl(\tfrac{1}{n}Q_{nt}'J_{n}\Sigma_{t}J_{n}Q_{nt}\Bigr)^{-1}\Bigl[J_{n}Q_{t}\Bigr]_{j}\] and thus \[|m_{ijt}|=O_{p}\Bigl(\tfrac{1}{n}\Big\Vert\bigl[J_{n}Q_{t}\bigr]_{i}\Big\Vert\Big\Vert\bigl[J_{n}Q_{t}\bigr]_{j}\Big\Vert\Bigr) =O_{p}\Bigl(\frac{\ell_{n}}{n}\Bigr)\] uniformly in \(i,j\) for each \(t\). Similarly, we also observe that \[\sum_{j=1}^{n}m_{ijt}^2=O_{p}\Bigl(\frac{\ell_{n}}{n}\Bigr).\] uniformly in \(i\) for each \(t\). So we have \(\|M_{t}\|_{\mathrm{rc}}=O_{p}(\ell_{n})\), \(\|M_{t}\|_{\mathrm{sp}}=O_{p}(\ell_{n})\) and thus \(\|M_{t}\|=O_{p}(\sqrt{\ell_{n}})\). \(\blacksquare\)

Lemma 13 indicates that the linear moments may be subject to various moment issues due to the cross-sectional dimension, which can be viewed as an extension of those presented in [4].

Lemma 14. For \(n\times n\) time-varying matrix \(\mathcal{B}_{t}\), suppose \(\mathcal{B}_{t}\) satisfies [[propzero]](#propzero){reference-type=“ref” reference=“propzero”} \(O_{p}(b)\) and \(Q_{jt}\) is the \(j\)-th column components of the IV \(Q_{t}\), under Assumptions 1, 4 and 6:
() \(\operatorname{E}\|r_{kt}^{*}\|^2=O(n^2h_{Tt}^4\ell_{k}^{-4\varsigma_{k}})\);
() \(\operatorname{E}|r_{kt}^{*\prime}{\mathcal{B}_{t}}r_{kt}^{*}|=O\left(nh_{Tt}^2\ell_{k}^{-2\varsigma_{k}}b\right)\);
() \(\operatorname{E}|r_{kt}^{*\prime}\mathcal{B}_{t}E_{t}^{*}|=O(nh_{Tt}^2\ell_{k}^{-\varsigma_{k}}b)\).
() \(\operatorname{E}|Q_{jt}'\mathcal{B}_{t}(r_{kt}^{*}+E_{t}^{*})|=O(\sqrt{n}h_{Tt}\ell_{k}^{-\varsigma_{k}}b)\).

For () \(\operatorname{E}\|r_{kt}\|^2=\operatorname{E}\left[\sum_{i}r_{kt,i}^2\right]^2=\operatorname{E}\left[\sum_{i}(\sum_{j}h_{3,ij}\epsilon_{j})^2\right]^2=\operatorname{E}(E_{t}'H_{3}'H_{3}E_{t})^2\). Note that \(H_{3}'H_{3}\) satisfies [[propzero]](#propzero){reference-type=“ref” reference=“propzero”} \(O_{p}(h^2)\) with \(h=\ell_{3}^{-\varsigma_{3}}\), and we have \(\operatorname{E}\|r_{kt}\|^2=O(n^2h_{Tt}^4\ell_{k}^{-4\varsigma_{k}})\). For (), taking \(k=3\) as an example, according to Lemmas 7 and 11, we know that the dominant term is \(h_{Tt}\operatorname{E}|r_{3t}^{\prime}{\mathcal{B}_{t}}r_{3t}|\). According to the Cauchy-Schwarz inequality, \(\operatorname{E}|r_{kt}^{\prime}{\mathcal{B}_{t}}r_{3t}|\leq\sqrt{\operatorname{E}\|r_{kt}^{\prime}{\mathcal{B}_{t}}r_{3t}\|^2}=\sqrt{\operatorname{E}\|E_{t}^{\prime}H_{3}'\mathcal{B}_{t}H_{3}E_{t}\|^2}\), also, \(H_{3}'\mathcal{B}_{t}H_{3}\) satisfies [[propzero]](#propzero){reference-type=“ref” reference=“propzero”} \(O_{p}(b\ell_{3}^{-2\varsigma_{3}})\). Thus, \(\operatorname{E}\|r_{3t}^{\prime}{\mathcal{B}_{t}}r_{3t}\|^2=O\left(n^2\ell_{3}^{-4\varsigma_{3}}b^2\right)\). The result of () can be proved similarly to (). For (), we know that \(Q_{t}\) can be expressed as a linear combination of \(Y_{t-1}^{(*,-1)}\) and \(X_{t}^{*}\). Since that \(Y_{t-1}^{(*,-1)}\) and \(X_{t}^{*}\) are both uncorrelated with \(E_{t}^{*}\), the desired result follows immediately by arguments analogous to those used for () and (). \(\blacksquare\)

Lemma 15. Under Assumption 1, \(\operatorname{Var}( \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\Sigma_{t}^{-1/2}E_{t}^{*})= \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }+o(1)\).

For the case of [var:homo] and [var:hetei], we have \(\operatorname{Var}( \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\Sigma_{t}^{-1/2}E_{t}^{*})= \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\). Hence, we only need to investigate the case of [var:hetet]. Denote \(H_{t}= \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\Sigma_{t}^{-1/2}E_{t}^{*}\). To prove that \(\operatorname{Var}(H_t)= \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }+o(1)\) under the case of [var:hetet] where \(\Sigma_t=\sigma_t^2 I_n\), we first observe that the scaling matrix simplifies to \(\Sigma_t^{-1/2}=\sigma_t^{-1} I_n\) and the projection matrix reduces to the standard centering matrix \(M(\Sigma_t)=I_n- (\sigma_t^{-1}l_n)(n \sigma_t^{-2})^{-1}(\sigma_t^{-1}l_n)'=J_n\). Since the variance of the forward orthogonal deviation error \(E_t^*\) is dominated by the current period variance \(\sigma_t^2 I_n\) for large \(T\), specifically, \(\operatorname{Var}(E_t^*) = \sigma_t^2 I_n + O((T-t)^{-1})\), the variance of the transformed error becomes \(\operatorname{Var}(H_t)=J_n(\sigma_t^{-1} I_n)[\sigma_t^2 I_n + o(1)](\sigma_t^{-1} I_n)J_n =J_n(I_n+o(1))J_n=J_n+o(1)\). Given that \(M(\Sigma_t)=J_n\) in this specification, it follows that \(\operatorname{Var}(H_t)= \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }+o_{p}(1)\) as \(T \to \infty\). \(\blacksquare\)

Lemma 16. Under Assumptions 3-6, \(\operatorname{E}\|m_{N}(\theta_{0})\|=o(1)\).

Observe that \[\begin{align} \begin{aligned} \operatorname{E}\|m_{N}(\theta_{0})\|& \leq\frac{1}{n(T-1)}\Bigl(\operatorname{E}\|\sum_{j=1}^{\ell_{p}}\mathbf{V}_{N}^{*\prime}\mathbf{J}_{N}\mathbf{P}_{Nj}\mathbf{J}_{N}\mathbf{V}_{N}^{*}\|+\operatorname{E}\|\sum_{j=1}^{\ell_{q}}\mathbf{Q}_{Nj}'\mathbf{J}_{N}\mathbf{V}_{N}^{*}\|\Bigr) \\& =O\left((\ell_{n}^{-2\underline{\varsigma}}+\ell_{n}^{-\underline{\varsigma}})\cdot\ell_{n}^{1/2}\right)=o(1) \end{aligned} \end{align}\] where the last two relations follow by the same argument used to bound \(\mathscr{A}_{1N}\) and \(\mathscr{A}_{2N}\) in the proof of Lemma 3. \(\blacksquare\)

Lemma 17. () For the MESS-type, when \(j=1,...,\ell_{k}\) and \(k=1,2,3\), under Assumption 2 and \(S_{k}\)’s satisfy [propUB], we have \[\begin{align} \|\tfrac{\partial S_{k}}{\partial\lambda_{kj_0}}-\varPhi_{kj}S_{k}\|_{\mathrm{sp}}=O(\|\varPhi_{kj}\|_{\mathrm{sp}})=O(1). \end{align}\] () \(\|\frac{\partial S_{k}}{\partial\lambda_{k_0}}\|_{\mathrm{sp}}=O(\sqrt{\ell_{n}})\), where \(\frac{\partial S_{k}}{\partial\lambda_{k_0}}=\left[\frac{\partial S_{k}}{\partial\lambda_{k_{10}}},...,\frac{\partial S_{k}}{\partial\lambda_{{k\ell_{k}}_0}}\right]\) for the true estimand \(\lambda_{kj_0}\).

Since \(\varPhi_{kj}\) does not commute with \(S_{k}\), the Fréchet derivative of the matrix exponential gives \[\begin{align} \tfrac{\partial S_{k}}{\partial\lambda_{kj_0}}=\int_0^1 e^{(1-t) \Xi_k}\varPhi_{kj} e^{t\Xi_k} d s \end{align}\] Because the norm of an integral is less than the integral of the norm, we have \[\begin{align} \begin{aligned} \|\tfrac{\partial S_{k}}{\partial\lambda_{kj_0}}\|_{\mathrm{sp}}&\leq \left(\int_0^1\|e^{(1-t) \Xi_k}\varPhi_{kj} e^{t\Xi_k}\|_{\mathrm{sp}}dt\right)\leq \left(\int_0^1\|e^{(1-t)\Xi_k}\|_{\mathrm{sp}} \|\varPhi_{kj}\|_{\mathrm{sp}}\|e^{t\Xi_k}\|_{\mathrm{sp}}dt\right) \\ &\leq \|\varPhi_{kj}\|_{\mathrm{sp}}\left(\int_0^1e^{(1-t)\|\Xi_k\|_{\mathrm{sp}}} e^{t\|\Xi_k\|_{\mathrm{sp}}}dt\right)\leq \|\varPhi_{kj}\|_{\mathrm{sp}}\cdot e^{\|\Xi_k\|_{\mathrm{sp}}}=O(\|\varPhi_{kj}\|_{\mathrm{sp}})=O(1). \end{aligned} \end{align}\] from Assumption 2. Thus, we have \[\begin{align} \|\tfrac{\partial S_{k}}{\partial\lambda_{kj_0}}-\varPhi_{kj}S_{k}\|_{\mathrm{sp}}\leq\|\tfrac{\partial S_{k}}{\partial\lambda_{kj_0}}\|_{\mathrm{sp}}+\|\varPhi_{kj}S_{k}\|_{\mathrm{sp}}=O(\|\varPhi_{kj}\|_{\mathrm{sp}}). \end{align}\] \(\blacksquare\)

12 Estimation of heteroskedastic variances↩︎

First, we explain the transformation operator \(\ifstrequal{t}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{t}}) }{ {J}({\Sigma_{t}}) }\) derived in BGMME, which can be motivated from the approximated log-likelihood function \[\begin{align} L_{nT}(\theta,\mathbf{c}_{n},\boldsymbol{\alpha}_{T},\Sigma_{t})=-\frac{nT}{2}\ln(2\pi)-\frac{nT}{2}\ln|\Sigma_{t}(\theta)|+T\bigl(\ln|(S_{1}(\lambda)|+\ln|S_{3}(\lambda)|\bigr)-\sum_{t=1}^{T}V_{t}^{c\prime}(\theta)\Sigma_{t}^{-1}(\theta)V_{t}^{c}(\theta). \end{align}\] The use of the approximate likelihood relies on the negligibility of \(r_{t}\), which in turn permits the replacement of the true \(g_{k0}\) with asymptotically negligible cost. Concentrating out \(\boldsymbol{\alpha}_{T}\) by the first-order condition, we have \[\begin{align} L_{nT}(\theta,\mathbf{c}_{n},\Sigma_{t})=-\frac{nT}{2}\ln(2\pi)-\frac{nT}{2}\ln|\Sigma_{t}(\theta)|+T\bigl(\ln|(S_{1}(\lambda)|+\ln|S_{3}(\lambda)|\bigr)-\sum_{t=1}^{T}V_{t}^{c\prime}(\theta) \ifstrequal{t}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{t}}) }{ {J}({\Sigma_{t}}) }V_{t}^{c}(\theta), \end{align}\] where \(V_{t}^{c}(\theta)=S_{3}(\lambda)S_{1}(\lambda)Y_{t}-S_{3}\bigl((\gamma I_{n}+S_{2})Y_{t-1}+X_{t}\beta+\mathbf{c}_{n}\bigr)\) and \(\ifstrequal{t}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{t}}) }{ {J}({\Sigma_{t}}) }=\Sigma_{t}^{-1}-\Sigma_{t}^{-1}l_{n}(l_{n}'\Sigma_{t}^{-1}l_{n})^{-1}l_{n}'\Sigma_{t}^{-1}\). Thus, \(\ifstrequal{t}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{t}}) }{ {J}({\Sigma_{t}}) }\) enables the best moment conditions to mimic the score of the likelihood function, and thus it can provide the moment conditions more efficiently than the counterparts relying on the operator \(J_{n}\). The BGMME has two main advantages over MLE. First, when the model is SAR-type, it avoids evaluating the Jacobian determinant. Second, the BGMME is subject only to approximation bias from sieves, whereas MLE can suffer additional bias from the incidental-parameters problem, as documented in the dynamic panel data literature.

13 Additional results↩︎

In this section, we present comprehensive results from our Monte Carlo experiments. While the main text focused on MESS-type specifications, here we provide the corresponding finite-sample performance for SAR-type models. Tables are presented on the following pages.

0.8

1

Finite sample performance of estimators of \(\pi\) for the MESS-type, \(\Var(\epsilon_{it})=\sigma_i^2\). Additional results.
\((n,T,\elln)=(100,10,2)\) \((n,T,\elln)=(100,10,[n^{1/5}]+2)\)
2-7 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\)
2-7 Bias -0.0539 0.0062 -0.0544 0.0049 -0.0562 0.0044 -0.0508 0.0047 -0.0517 0.0036 -0.0524 -0.0016
ESD 0.0393 0.0381 0.0388 0.0345 0.0435 0.0386 0.0381 0.0350 0.0374 0.0339 0.0398 0.0375
RMSE 0.0667 0.0386 0.0668 0.0348 0.0711 0.0388 0.0635 0.0353 0.0638 0.0341 0.0658 0.0375
CP 0.7800 0.9190 0.7770 0.9240 0.7260 0.8750 0.7890 0.8950 0.7940 0.9180 0.7220 0.8840
2-7 \((n,T,\elln)=(100,25,[n^{1/5}]+2)\)
2-7 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\)
2-7 Bias -0.0445 0.0095 -0.0434 0.0080 -0.0429 0.0077 -0.0441 0.0074 -0.0433 0.0060 -0.0428 0.0008
ESD 0.0246 0.0243 0.0223 0.0240 0.0211 0.0214 0.0231 0.0237 0.0204 0.0219 0.0201 0.0205
RMSE 0.0508 0.0261 0.0488 0.0253 0.0478 0.0227 0.0498 0.0248 0.0479 0.0227 0.0473 0.0205
CP 0.8970 0.9090 0.9000 0.9270 0.9070 0.8340 0.8960 0.9010 0.9010 0.9140 0.8990 0.9060
2-7 \((n,T,\elln)=(200,10,[n^{1/5}]+2)\)
2-7 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\)
2-7 Bias -0.0396 0.0071 -0.0372 0.0067 -0.0369 0.0066 -0.0349 0.0050 -0.0343 0.0035 -0.0339 -0.0050
ESD 0.0320 0.0282 0.0260 0.0246 0.0256 0.0231 0.0299 0.0278 0.0249 0.0232 0.0249 0.0222
RMSE 0.0509 0.0291 0.0454 0.0255 0.0449 0.0240 0.0460 0.0282 0.0424 0.0235 0.0421 0.0228
CP 0.8850 0.8820 0.8540 0.9020 0.8770 0.8520 0.8930 0.8950 0.8840 0.9060 0.8820 0.9030
2-7 \((n,T,\elln)=(200,25,[n^{1/4}])\)
2-7 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\)
2-7 Bias -0.0199 0.0079 -0.0124 0.0087 -0.0123 0.0063 -0.0155 0.0062 -0.0109 0.0055 -0.0103 -0.0048
ESD 0.0138 0.0145 0.0135 0.0130 0.0132 0.0128 0.0135 0.0135 0.0133 0.0128 0.0130 0.0126
RMSE 0.0242 0.0165 0.0183 0.0156 0.0180 0.0143 0.0206 0.0149 0.0172 0.0139 0.0166 0.0135
CP 0.9140 0.8990 0.9160 0.9210 0.9250 0.9320 0.9310 0.9610 0.9680 0.9210 0.9700 0.9440
2-7 \((n,T,\elln)=(400,10,[n^{1/5}]+2)\)
2-7 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\)
2-7 Bias -0.0173 -0.0055 -0.0161 0.0031 -0.0153 0.0010 -0.0151 0.0023 -0.0131 0.0002 -0.0126 -0.0010
ESD 0.0241 0.0233 0.0177 0.0158 0.0177 0.0157 0.0241 0.0226 0.0170 0.0155 0.0170 0.0154
RMSE 0.0297 0.0239 0.0239 0.0161 0.0234 0.0157 0.0284 0.0227 0.0215 0.0155 0.0212 0.0154
CP 0.9470 0.9140 0.9610 0.9180 0.9520 0.9230 0.9260 0.9200 0.9450 0.9240 0.9180 0.9420
2-7 \((n,T,\elln)=(400,25,[n^{1/5}]+2)\)
2-7 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\)
2-7 Bias -0.0142 -0.0095 -0.0125 0.0031 -0.0122 0.0010 -0.0133 0.0029 -0.0113 0.0010 -0.0071 0.0008
ESD 0.0152 0.0155 0.0095 0.0100 0.0095 0.0099 0.0140 0.0121 0.0091 0.0091 0.0091 0.0090
RMSE 0.0208 0.0182 0.0157 0.0105 0.0155 0.0100 0.0193 0.0124 0.0145 0.0092 0.0115 0.0090
CP 0.9430 0.9510 0.9530 0.9180 0.9430 0.9640 0.9570 0.9490 0.9310 0.9520 0.9420 0.9400

Note: The true parameters are set to \(\pi_0=(0.3,1)'\), and \(\bard_0=15\%\). The results are based on 1,000 Monte Carlo replications. Bias denotes the mean bias of the estimates, ESD denotes the standard deviation, RMSE denotes the root mean squared error, and CP denotes the 95% coverage probability. The 2SLS refers to the two-stage least square estimator, OGMM refers to the feasible optimal GMM estimator, and BGMM refers to the feasible best GMM estimator.

0.5

Finite sample performance of estimators of \(G_{1}, G_{2}\) and \(G_{3}\) for the MESS-type, \(\Var(\epsilon_{it})=\sigma_i^2\). Additional results.
\((n,T,\elln)=(100,10,2)\) \((n,T,\elln)=(100,10,[n^{1/5}]+2)\)
2-20 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_1\) \(\tilde{g}_2\) \(\tilde{g}_3\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
MAE 0.0339 0.0352 0.0530 0.0338 0.0351 0.0514 0.0340 0.0352 0.0505 MAE 0.0323 0.0342 0.0509 0.0322 0.0341 0.0499 0.0327 0.0343 0.0489
Bias -0.0204 -0.0257 -0.0464 -0.0201 -0.0259 -0.0457 -0.0197 -0.0258 -0.0451 Bias -0.0187 -0.0255 -0.0450 -0.0185 -0.0257 -0.0446 -0.0180 -0.0255 -0.0441
RMSE 0.0228 0.0288 0.0531 0.0225 0.0289 0.0517 0.0224 0.0289 0.0506 RMSE 0.0209 0.0284 0.0509 0.0207 0.0285 0.0500 0.0204 0.0284 0.0491
\((n,T,\elln)=(100,25,2)\) \((n,T,\elln)=(100,25,[n^{1/5}]+2)\)
2-20 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
MAE 0.0199 0.0209 0.0375 0.0199 0.0209 0.0368 0.0198 0.0209 0.0361 MAE 0.0194 0.0209 0.0322 0.0192 0.0208 0.0317 0.0192 0.0208 0.0311
Bias -0.0159 -0.0180 -0.0352 -0.0156 -0.0181 -0.0346 -0.0153 -0.0180 -0.0342 Bias -0.0128 -0.0162 -0.0305 -0.0125 -0.0163 -0.0302 -0.0121 -0.0162 -0.0298
RMSE 0.0164 0.0191 0.0370 0.0160 0.0191 0.0362 0.0158 0.0191 0.0357 RMSE 0.0137 0.0175 0.0326 0.0133 0.0175 0.0321 0.0130 0.0175 0.0316
\((n,T,\elln)=(200,10,2)\) \((n,T,\elln)=(200,10,[n^{1/5}]+2)\)
2-20 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_1\) \(\tilde{g}_2\) \(\tilde{g}_3\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
MAE 0.0329 0.0342 0.0509 0.0329 0.0341 0.0499 0.0331 0.0343 0.0489 MAE 0.0323 0.0329 0.0488 0.0322 0.0328 0.0478 0.0327 0.0330 0.0467
Bias -0.0190 -0.0255 -0.0475 -0.0185 -0.0257 -0.0467 -0.0180 -0.0255 -0.0457 Bias -0.0187 -0.0240 -0.0450 -0.0185 -0.0243 -0.0446 -0.0175 -0.0247 -0.0441
RMSE 0.0216 0.0287 0.0541 0.0211 0.0288 0.0526 0.0209 0.0287 0.0512 RMSE 0.0196 0.0250 0.0472 0.0194 0.0252 0.0466 0.0186 0.0257 0.0459
\((n,T,\elln)=(200,25,2)\) \((n,T,\elln)=(200,25,[n^{1/5}]+2)\)
2-20 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
MAE 0.0094 0.0099 0.0198 0.0094 0.0099 0.0198 0.0093 0.0099 0.0194 MAE 0.0092 0.0098 0.0166 0.0091 0.0098 0.0166 0.0092 0.0099 0.0164
Bias -0.0075 -0.0084 -0.0181 -0.0074 -0.0083 -0.0181 -0.0074 -0.0083 -0.0180 Bias -0.0062 -0.0076 -0.0159 -0.0062 -0.0075 -0.0159 -0.0061 -0.0075 -0.0158
RMSE 0.0077 0.0090 0.0190 0.0076 0.0089 0.0190 0.0076 0.0089 0.0189 RMSE 0.0068 0.0083 0.0173 0.0068 0.0082 0.0173 0.0067 0.0083 0.0170
\((n,T,\elln)=(400,10,2)\) \((n,T,\elln)=(400,10,[n^{1/5}]+2)\)
2-20 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_1\) \(\tilde{g}_2\) \(\tilde{g}_3\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
MAE 0.0195 0.0200 0.0367 0.0195 0.0200 0.0360 0.0195 0.0200 0.0354 MAE 0.0186 0.0196 0.0314 0.0184 0.0196 0.0309 0.0183 0.0197 0.0302
Bias -0.0158 -0.0180 -0.0357 -0.0153 -0.0181 -0.0350 -0.0148 -0.0182 -0.0344 Bias -0.0126 -0.0159 -0.0313 -0.0122 -0.0160 -0.0308 -0.0115 -0.0163 -0.0301
RMSE 0.0160 0.0183 0.0363 0.0154 0.0184 0.0355 0.0149 0.0186 0.0349 RMSE 0.0129 0.0163 0.0320 0.0125 0.0164 0.0314 0.0119 0.0167 0.0306
\((n,T,\elln)=(400,25,2)\) \((n,T,\elln)=(400,25,[n^{1/5}]+2)\)
2-20 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
MAE 0.0092 0.0095 0.0188 0.0092 0.0095 0.0187 0.0092 0.0096 0.0184 MAE 0.0088 0.0095 0.0169 0.0088 0.0095 0.0169 0.0088 0.0096 0.0164
Bias -0.0075 -0.0087 -0.0181 -0.0074 -0.0086 -0.0180 -0.0073 -0.0087 -0.0178 Bias -0.0064 -0.0078 -0.0169 -0.0064 -0.0078 -0.0168 -0.0060 -0.0079 -0.0164
RMSE 0.0076 0.0089 0.0183 0.0075 0.0088 0.0182 0.0074 0.0089 0.0180 RMSE 0.0066 0.0081 0.0173 0.0066 0.0081 0.0172 0.0062 0.0082 0.0167

Note: The \(\tilde{g}_{k}\) extracts the column vectors composed of non-zero elements from the upper triangular submatrix of \(G_{k}\), and \(\bard_0=15\%\). The results are based on 1,000 Monte Carlo replications. MAE denotes the mean absolute error. Bias denotes the mean bias of the estimates, and RMSE denotes the root mean squared error. The 2SLS refers to the two-stage least square estimator, OGMM refers to the feasible optimal GMM estimator, and BGMM refers to the feasible best GMM estimator.

1

Finite sample performance of estimators of \(\pi\) for the SAR-type, \(\Var(\epsilon_{it})=\sigma_i^2\).
\((n,T,\elln)=(100,10,2)\) \((n,T,\elln)=(100,10,[n^{1/5}]+2)\)
2-7 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\)
2-7 Bias 0.0021 -0.0024 0.0108 -0.0015 0.0110 -0.0018 -0.0020 -0.0041 0.0089 -0.0029 0.0087 -0.0027
ESD 0.0361 0.0292 0.0406 0.0287 0.0416 0.0294 0.0378 0.0288 0.0424 0.0283 0.0423 0.0292
RMSE 0.0362 0.0293 0.0420 0.0288 0.0430 0.0295 0.0378 0.0291 0.0433 0.0285 0.0432 0.0294
CP 0.7970 0.9070 0.7950 0.9280 0.7850 0.9250 0.8930 0.9030 0.8830 0.9100 0.8720 0.8990
2-7 \((n,T,\elln)=(100,25,[n^{1/5}]+2)\)
2-7 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\)
2-7 Bias 0.0163 -0.0002 0.0163 -0.0002 0.0149 0.0000 0.0152 0.0012 0.0152 0.0012 0.0125 0.0002
ESD 0.0221 0.0175 0.0221 0.0163 0.0212 0.0163 0.0202 0.0173 0.0202 0.0165 0.0196 0.0164
RMSE 0.0275 0.0175 0.0275 0.0163 0.0259 0.0163 0.0253 0.0173 0.0253 0.0165 0.0232 0.0164
CP 0.8960 0.9070 0.8840 0.9270 0.8730 0.9160 0.8980 0.9140 0.8990 0.9180 0.8880 0.9070
2-7 \((n,T,\elln)=(200,10,[n^{1/5}]+2)\)
2-7 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\)
2-7 Bias 0.0088 -0.0017 0.0086 -0.0007 0.0035 -0.0006 0.0057 -0.0002 0.0054 -0.0003 -0.0017 -0.0023
ESD 0.0286 0.0206 0.0283 0.0203 0.0270 0.0200 0.0270 0.0211 0.0266 0.0208 0.0252 0.0206
RMSE 0.0299 0.0207 0.0296 0.0203 0.0272 0.0200 0.0276 0.0211 0.0271 0.0208 0.0253 0.0207
CP 0.8000 0.8990 0.8060 0.9260 0.7960 0.9150 0.8430 0.8880 0.8140 0.9200 0.8040 0.9090
2-7 \((n,T,\elln)=(200,25,[n^{1/5}]+2)\)
2-7 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\)
2-7 Bias 0.0123 0.0019 0.0123 0.0019 0.0109 -0.0002 0.0089 0.0017 0.0089 0.0017 0.0063 -0.0006
ESD 0.0150 0.0118 0.0150 0.0115 0.0147 0.0115 0.0141 0.0120 0.0139 0.0115 0.0139 0.0115
RMSE 0.0194 0.0120 0.0194 0.0117 0.0183 0.0115 0.0167 0.0121 0.0165 0.0116 0.0153 0.0115
CP 0.9260 0.9400 0.9360 0.9330 0.9250 0.9220 0.9400 0.9350 0.9420 0.9420 0.9350 0.9400
2-7 \((n,T,\elln)=(400,10,[n^{1/5}]+2)\)
2-7 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\)
2-7 Bias 0.0063 0.0011 0.0062 0.0011 0.0033 -0.0004 0.0050 0.0018 0.0049 0.0015 0.0009 0.0000
ESD 0.0198 0.0145 0.0198 0.0144 0.0193 0.0144 0.0189 0.0146 0.0185 0.0145 0.0181 0.0144
RMSE 0.0208 0.0145 0.0207 0.0144 0.0196 0.0144 0.0196 0.0147 0.0191 0.0146 0.0181 0.0144
CP 0.9120 0.9220 0.9270 0.9300 0.9190 0.9320 0.9130 0.9300 0.9320 0.9360 0.9240 0.9330
2-7 \((n,T,\elln)=(400,25,[n^{1/5}]+2)\)
2-7 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\) \(\gamma\) \(\beta\)
2-7 Bias 0.0073 0.0019 0.0073 0.0019 0.0070 0.0001 0.0060 0.0018 0.0060 0.0018 0.0054 -0.0007
ESD 0.0102 0.0081 0.0102 0.0081 0.0100 0.0081 0.0095 0.0084 0.0095 0.0084 0.0095 0.0083
RMSE 0.0125 0.0083 0.0125 0.0083 0.0122 0.0081 0.0112 0.0086 0.0112 0.0086 0.0109 0.0083
CP 0.9250 0.9360 0.9320 0.9380 0.9250 0.9360 0.9340 0.9420 0.9400 0.9450 0.9410 0.9420

Note: The true parameters are set to \(\pi_0=(0.3,1)'\), and \(\bard_0=10\%\). The results are based on 1,000 Monte Carlo replications. Bias denotes the mean bias of the estimates, ESD denotes the standard deviation, RMSE denotes the root mean squared error, and CP denotes the 95% coverage probability. The 2SLS refers to the two-stage least square estimator, OGMM refers to the feasible optimal GMM estimator, and BGMM refers to the feasible best GMM estimator.

1

Estimated \(\rho(\hat{A})\) for the SAR-type, \(\Var(\epsilon_{it})=\sigma_i^2\), \(\rho(A)=0.7333\).
\((n,T)=(100,10)\) \((n,T,\elln)=(100,25)\)
2-7 [4]* \(\elln=2\) \(\elln=[n^{1/5}]+2\) \(\elln=2\) \(\elln=[n^{1/5}]+2\)
2-7 2SLS OGMM BGMM 2SLS OGMM BGMM 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 Mean 0.6065 0.6044 0.6096 0.6096 0.6118 0.6121 0.6485 0.6274 0.6282 0.6566 0.6464 0.6454
ESD 0.1834 0.1786 0.1785 0.1724 0.1742 0.1735 0.1353 0.1420 0.1408 0.1327 0.1365 0.1355
RMSE 0.2230 0.2202 0.2172 0.2122 0.2124 0.2117 0.1597 0.1772 0.1757 0.1533 0.1618 0.1615
2-7 \((n,T,\elln)=(200,25)\)
2-7 [4]* \(\elln=2\) \(\elln=[n^{1/5}]+2\) \(\elln=2\) \(\elln=[n^{1/5}]+2\)
2-7 2SLS OGMM BGMM 2SLS OGMM BGMM 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 Mean 0.6319 0.6355 0.6614 0.6368 0.6349 0.6355 0.6885 0.6740 0.6741 0.7007 0.6926 0.6924
ESD 0.1520 0.1542 0.1535 0.1508 0.1557 0.1535 0.0918 0.0827 0.0827 0.0898 0.0859 0.0862
RMSE 0.1827 0.1826 0.1695 0.1791 0.1842 0.1820 0.1022 0.1018 0.1017 0.0955 0.0951 0.0954
2-7 \((n,T,\elln)=(400,25)\)
2-7 [4]* \(\elln=2\) \(\elln=[n^{1/5}]+2\) \(\elln=2\) \(\elln=[n^{1/5}]+2\)
2-7 2SLS OGMM BGMM 2SLS OGMM BGMM 2SLS OGMM BGMM 2SLS OGMM BGMM
2-7 Mean 0.6818 0.6535 0.6535 0.6901 0.6619 0.6606 0.7489 0.7131 0.7131 0.7533 0.7534 0.7545
ESD 0.1237 0.1162 0.1162 0.1205 0.1160 0.1161 0.0855 0.0835 0.0835 0.0803 0.0806 0.0812
RMSE 0.1340 0.1410 0.1410 0.1280 0.1363 0.1370 0.0870 0.0859 0.0859 0.0828 0.0830 0.0839

Note: The \(\bard_0=10\%\), and the results are based on 1,000 Monte Carlo replications. Mean denotes the mean of the estimates. Bias denotes the mean bias of the estimates, ESD denotes the standard deviation, and RMSE denotes the root mean squared error. The 2SLS refers to the two-stage least square estimator, OGMM refers to the feasible optimal GMM estimator, and BGMM refers to the feasible best GMM estimator.

0.5

Finite sample performance of estimators of \(G_{1}, G_{2}\) and \(G_{3}\) for the SAR-type, \(\Var(\epsilon_{it})=\sigma_i^2\).
\((n,T,\elln)=(100,10,2)\) \((n,T,\elln)=(100,10,[n^{1/5}]+2)\)
2-20 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_1\) \(\tilde{g}_2\) \(\tilde{g}_3\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
MAE 0.0158 0.0178 - 0.0158 0.0190 0.0193 0.0159 0.0190 0.0195 MAE 0.0119 0.0123 - 0.0117 0.0124 0.0122 0.0116 0.0124 0.0123
Bias 0.0017 -0.0039 - -0.0060 -0.0022 0.0016 -0.0056 -0.0023 0.0017 Bias 0.0013 -0.0012 - -0.0070 -0.0005 0.0012 -0.0069 -0.0005 0.0013
RMSE 0.0157 0.0193 - 0.0165 0.0213 0.0206 0.0167 0.0212 0.0208 RMSE 0.0081 0.0098 - 0.0102 0.0099 0.0084 0.0101 0.0099 0.0084
\((n,T,\elln)=(100,25,2)\) \((n,T,\elln)=(100,25,[n^{1/5}]+2)\)
2-20 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
MAE 0.0124 0.0124 - 0.0117 0.0129 0.0122 0.0117 0.0129 0.0122 MAE 0.0119 0.0123 - 0.0117 0.0124 0.0122 0.0116 0.0124 0.0123
Bias 0.0026 -0.0014 - -0.0056 -0.0001 0.0018 -0.0054 -0.0001 0.0019 Bias 0.0013 -0.0012 - -0.0070 -0.0005 0.0012 -0.0069 -0.0005 0.0013
RMSE 0.0099 0.0110 - 0.0103 0.0115 0.0097 0.0103 0.0115 0.0097 RMSE 0.0081 0.0098 - 0.0102 0.0099 0.0084 0.0101 0.0099 0.0084
\((n,T,\elln)=(200,10,2)\) \((n,T,\elln)=(200,10,[n^{1/5}]+2)\)
2-20 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_1\) \(\tilde{g}_2\) \(\tilde{g}_3\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
MAE 0.0082 0.0090 - 0.0077 0.0094 0.0093 0.0078 0.0094 0.0092 MAE 0.0074 0.0079 - 0.0069 0.0083 0.0071 0.0069 0.0083 0.0071
Bias 0.0022 -0.0026 - -0.0015 -0.0013 0.0001 -0.0015 -0.0013 0.0000 Bias 0.0011 -0.0025 - -0.0030 -0.0013 -0.0004 -0.0029 -0.0013 -0.0004
RMSE 0.0083 0.0100 - 0.0081 0.0105 0.0100 0.0083 0.0105 0.0099 RMSE 0.0061 0.0079 - 0.0065 0.0084 0.0061 0.0065 0.0083 0.0061
\((n,T,\elln)=(200,25,2)\) \((n,T,\elln)=(200,25,[n^{1/5}]+2)\)
2-20 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
MAE 0.0064 0.0060 - 0.0057 0.0059 0.0058 0.0057 0.0059 0.0058 MAE 0.0063 0.0058 - 0.0056 0.0059 0.0056 0.0056 0.0059 0.0056
Bias 0.0029 -0.0018 - -0.0004 -0.0009 0.0000 -0.0005 -0.0009 -0.0001 Bias 0.0017 -0.0017 - -0.0011 -0.0010 -0.0008 -0.0011 -0.0010 -0.0008
RMSE 0.0054 0.0055 - 0.0047 0.0053 0.0048 0.0048 0.0053 0.0047 RMSE 0.0039 0.0045 - 0.0036 0.0044 0.0037 0.0036 0.0044 0.0037
\((n,T,\elln)=(400,10,2)\) \((n,T,\elln)=(400,10,[n^{1/5}]+2)\)
2-20 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_1\) \(\tilde{g}_2\) \(\tilde{g}_3\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
MAE 0.0042 0.0046 - 0.0039 0.0048 0.0040 0.0039 0.0048 0.0040 MAE 0.0039 0.0043 - 0.0035 0.0044 0.0037 0.0035 0.0044 0.0037
Bias 0.0016 -0.0004 - 0.0003 0.0002 0.0007 0.0003 0.0002 0.0008 Bias 0.0013 -0.0005 - -0.0001 0.0001 0.0004 -0.0001 0.0001 0.0005
RMSE 0.0044 0.0052 - 0.0040 0.0054 0.0045 0.0041 0.0054 0.0044 RMSE 0.0035 0.0045 - 0.0031 0.0046 0.0033 0.0031 0.0046 0.0033
\((n,T,\elln)=(400,25,2)\) \((n,T,\elln)=(400,25,[n^{1/5}]+2)\)
2-20 [4]* 2SLS OGMM BGMM 2SLS OGMM BGMM
2-10 \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\) \(\tilde{g}_{1}\) \(\tilde{g}_{2}\) \(\tilde{g}_{3}\)
MAE 0.0035 0.0031 - 0.0032 0.0031 0.0027 0.0032 0.0031 0.0027 MAE 0.0034 0.0032 - 0.0029 0.0031 0.0025 0.0029 0.0031 0.0025
Bias 0.0022 0.0001 - 0.0016 0.0002 0.0005 0.0016 0.0002 0.0005 Bias 0.0017 0.0003 - 0.0008 0.0003 0.0000 0.0008 0.0003 0.0000
RMSE 0.0033 0.0030 - 0.0028 0.0030 0.0022 0.0028 0.0030 0.0022 RMSE 0.0026 0.0026 - 0.0017 0.0026 0.0015 0.0017 0.0026 0.0015

Note: The \(\tilde{g}_{k}\) extracts the column vectors composed of non-zero elements from the upper triangular submatrix of \(G_{k}\), and \(\bard_0=10\%\). The results are based on 1,000 Monte Carlo replications. MAE denotes the mean absolute error. Bias denotes the mean bias of the estimates, and RMSE denotes the root mean squared error. The 2SLS refers to the two-stage least square estimator, OGMM refers to the feasible optimal GMM estimator, and BGMM refers to the feasible best GMM estimator.

References↩︎

[1]
Miguel, E. (2005). Poverty and witch killing. The Review of Economic Studies 72, 1153–1172.
[2]
Yu, J., R. M. De Jong, and L.-f. Lee (2008). Quasi-maximum likelihood estimators for spatial dynamic panel data with fixed effects when both \(n\) and \({T}\) are large. Journal of Econometrics 146, 118–134.
[3]
Lee, L.-f. and J. Yu (2010). A spatial dynamic panel data model with both time and individual fixed effects. Econometric Theory 26, 564–597.
[4]
Lee, L.-f. and J. Yu (2014). Efficient GMM estimation of spatial dynamic panel data models with fixed effects. Journal of Econometrics 180, 174–197.
[5]
Yu, J., R. de Jong, and L.-f. Lee (2012). Estimation for spatial dynamic panel data with fixed effects: The case of spatial cointegration. Journal of Econometrics 167, 16–37.
[6]
Su, L. and Z. Yang (2015). estimation of dynamic panel data models with spatial errors. Journal of Econometrics 185, 230–258.
[7]
Shi, W. and L.-f. Lee (2017). Spatial dynamic panel data models with interactive fixed effects. Journal of Econometrics 197, 323–347.
[8]
Kuersteiner, G. M. and I. R. Prucha (2020). Dynamic spatial panel models: Networks, common shocks, and sequential exogeneity. Econometrica 88, 2109–2146.
[9]
Bai, J. and K. Li (2021). Dynamic spatial panel data models with common shocks. Journal of Econometrics.
[10]
Su, L., W. Wang, and X. Xu (2023). Identifying latent group structures in spatial dynamic panels. Journal of Econometrics 235, 1955–1980.
[11]
Yang, Z. (2018). Unified M-estimation of fixed-effects spatial dynamic models with short panels. Journal of Econometrics 205, 423–447.
[12]
LeSage, J. P. and R. K. Pace (2007). A matrix exponential spatial specification. Journal of Econometrics 140, 190–214.
[13]
Han, X. and L.-f. Lee (2013). Model selection using J-test for the spatial autoregressive model vs. the matrix exponential spatial model. Regional Science and Urban Economics 43, 250–271.
[14]
Debarsy, N., F. Jin, and L.-F. Lee (2015). Large sample properties of the matrix exponential spatial specification with an application to FDI. Journal of Econometrics 188, 1–21.
[15]
Yang, Y. (2022). Unified M-estimation of matrix exponential spatial dynamic panel specification. Econometric Reviews 41, 729–748.
[16]
Pinkse, J., M. E. Slade, and C. Brett (2002). Spatial price competition: A semiparametric approach. Econometrica 70, 1111–1153.
[17]
Lam, C. and P. C. Souza (2020). Estimation and selection of spatial weight matrix in a spatial lag model. Journal of Business & Economic Statistics 38, 693–710.
[18]
Sun, Y. (2016). Functional-coefficient spatial autoregressive models with nonparametric spatial weights. Journal of Econometrics 195, 134–153.
[19]
De Paula, A., I. Rasul, and P. C. Souza (2025). Identifying network ties from panel data: Theory and an application to tax competition. Review of Economic Studies 92, 2691–2729.
[20]
Chen, S., X. Song, and J. Yu (2025). Spatial panel data models with nonparametric spatial weights. Working paper.
[21]
Su, L. and S. Jin (2012). Sieve estimation of panel data models with cross section dependence. Journal of Econometrics 169, 34–47.
[22]
Sun, Y. and E. Malikov (2018). Estimation and inference in functional-coefficient spatial autoregressive panel data models with fixed effects. Journal of Econometrics 203, 359–378.
[23]
Hoshino, T. (2022). Sieve IV estimation of cross-sectional interaction models with nonparametric endogenous effect. Journal of Econometrics 229, 263–275.
[24]
Gupta, A., X. Qu, S. Srisuma, and J. Zhang (2025). Wald inference on varying coefficients. https://arxiv.org/abs/2502.03084.
[25]
Chen, Y., X. Han, F. Jin, and J. Zhang (2025). Efficient and sequential estimation of high-order spatial dynamic panels with time-varying dominant units. Available at SSRN 5167495.
[26]
Huang, C.-C., D. Laing, and P. Wang (2004). Crime and poverty: A search-theoretic approach. International Economic Review 45, 909–938.
[27]
Mehlum, H., E. Miguel, and R. Torvik (2006). Poverty and crime in 19th century Germany. Journal of Urban Economics 59, 370–388.
[28]
Khanna, G., C. Medina, A. Nyshadham, C. Posso, and J. Tamayo (2021, March). Job loss, credit, and crime in Colombia. American Economic Review: Insights 3, 97–114.
[29]
McGuirk, E. and M. Burke (2020). The economic origins of conflict in Africa. Journal of Political Economy 128, 3940–3997.
[30]
Heilmann, K., M. E. Kahn, and C. K. Tang (2021). The urban crime and heat gradient in high and low poverty areas. Journal of Public Economics 197, 104408.
[31]
Lee, L.-f. and X. Liu (2010). Efficient GMM estimation of high order spatial autoregressive models with autoregressive disturbances. Econometric Theory 26, 187–230.
[32]
Arellano, M. and O. Bover (1995). Another look at the instrumental variable estimation of error-components models. Journal of Econometrics 68, 29–51.
[33]
Alvarez, J. and M. Arellano (2003). The time series and cross-section asymptotics of dynamic panel data estimators. Econometrica 71, 1121–1159.
[34]
Yang, Z., X. Song, and J. Yu (2025). Estimation of spatial autoregressive panel data models with nonparametric endogenous effect. Journal of Econometrics 252, 106112.
[35]
Hansen, L. P. (1982). Large sample properties of generalized method of moments estimators. Econometrica, 1029–1054.
[36]
Hall, P. and L.-S. Huang (2001). Nonparametric kernel regression subject to monotonicity constraints. The Annals of Statistics 29, 624–647.
[37]
Malikov, E. and Y. Sun (2017). Semiparametric estimation and testing of smooth coefficient spatial autoregressive models. Journal of Econometrics 199, 12–34.
[38]
Mesaki, S. (1994). . In R. Abrahams (Ed.), Witchcraft in Contemporary Tanzania. African Studies Center, University of Cambridge.
[39]
LeSage, J. and R. K. Pace (2009). Introduction to Spatial Econometrics. Chapman and Hall/CRC.
[40]
Elhorst, J. P. (2014). Spatial panel models. Handbook of Regional Science 3, 705.
[41]
Pesaran, M. H. and C. F. Yang (2021). Estimation and inference in spatial models with dominant units. Journal of Econometrics 221, 591–615.
[42]
Lee, L.-f., C. Yang, and J. Yu (2022). and efficient GMM estimation of spatial autoregressive models with dominant (popular) units. Journal of Business & Economic Statistics, 1–38.
[43]
Zhang, J., C. Zhao, and X. Qu (2025). Nonlinear spatial dynamic panel data models with endogenous dominant units: An application to share data. Journal of Business & Economic Statistics 43, 150–163.
[44]
Lin, X. and L.-f. Lee (2010). estimation of spatial autoregressive models with unknown heteroskedasticity. Journal of Econometrics 157, 34–52.
[45]
Su, L. and T. Hoshino (2016). Sieve instrumental variable quantile regression estimation of functional coefficient models. Journal of Econometrics 191, 231–254.
[46]
Andrews, D. W. (1992). Generic uniform convergence. Econometric Theory 8, 241–257.
[47]
Hall, P. and C. C. Heyde (2014). Martingale Limit Theory and its Application. Academic press.
[48]
Jin, F. and L.-f. Lee (2020). Asymptotically efficient root estimators for spatial autoregressive models with spatial autoregressive disturbances. Economics Letters 194, 109397.

  1. Department of Economics, Queen’s University, Dunning Hall, 94 University Avenue, Kingston, Ontario K7L 3N6, Canada, and Department of Economics, University of Essex, Wivenhoe Park, Colchester CO4 3SQ, UK. Email: abhimanyu.g@queensu.ca.↩︎

  2. Department of Economics, Antai College of Economics and Management, Shanghai Jiao Tong University, 1954 Huashan Road, Shanghai, 200030, China PRC. Email: xiqu@sjtu.edu.cn.↩︎

  3. International Business School, Shanghai University of International Business and Economics, 201620, China PRC. Email: jiajun30@suibe.edu.cn.↩︎

  4. Gupta’s research was supported in part by the Social Science and Humanities Research Council of Canada. Qu’s research was supported by the National Natural Science Foundation of China via grant 72595872. Zhang’s research was supported by the National Natural Science Foundation of China via grant 72503134.↩︎

  5. Our model can be viewed as a semi-nonparametric counterpart of [4], delivering a more flexible specification of the spatial interaction structure.↩︎

  6. In our semi-nonparametric setting, the transformed error term is \(V_t^{*} = E_t^{*} + r_t^{*}\), where \(r_t^{*}\) represents the transformed sieve approximation error. As shown in Lemma 11 in the supplement, under certain rate conditions, \(r_t^{*}\) is asymptotically negligible.↩︎

  7. We approximate the unknown weight functions by a sieve space of dimension \(\ell_{n}=\max_{1\leq k\leq 3}\ell_{k}\), where \(\ell_{k}\to\infty\) but grows sufficiently slowly relative to \(n,T\). This contrasts with [22], who nonparametrically estimate coefficient functions of a \(T\)-dimensional vector of covariates. Hence, the first issue in [22] does not arise in our setting.↩︎

  8. It is difficult to treat the fully two-dimensional unknown heteroskedastic case \(\sigma_{it}^2\) because we cannot obtain feasible best moment conditions to construct the BGMME. A common way to restore feasibility is to impose a separable structure, for example \(\sigma_{it}^2 = \sigma_{i}^2 \sigma_{t}^2\) or \(\sigma_{it}^2 = \sigma_{i}^2 + \sigma_{t}^2\), which reduces the number of variance parameters and makes consistent estimation possible. Similar discussions can be found in [25].↩︎

  9. When the variance of \(\epsilon_{it}\) is cross-sectionally homogeneous, i.e., \(\sigma_{it}^2=\sigma_{t}^2\) or \(\sigma^2\), we have \(\ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }=J_{n}\) and \(\ifstrequal{t}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{t}}) }{ {J}({\Sigma_{t}}) }=\sigma_{t}^{-1}J_{n}\) or \(\ifstrequal{t}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{t}}) }{ {J}({\Sigma_{t}}) }=\sigma^{-1}J_{n}\). Also, note that \(\ifstrequal{t}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{t}}) }{ {J}({\Sigma_{t}}) }\Sigma_{t} \ifstrequal{t}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{t}}) }{ {J}({\Sigma_{t}}) }=\Sigma_{t}^{-1/2} \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\Sigma_{t}^{-1/2}\Sigma_{t}\Sigma_{t}^{-1/2} \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\Sigma_{t}^{-1/2}=\Sigma_{t}^{-1/2} \ifstrequal{t}{N}{ \pmb{M}({\boldsymbol{\Sigma}_{t}}) }{ {M}({\Sigma_{t}}) }\Sigma_{t}^{-1/2}= \ifstrequal{t}{N}{ \pmb{J}({\boldsymbol{\Sigma}_{t}}) }{ {J}({\Sigma_{t}}) }\).↩︎

  10. In the cross-sectional setting for each \(t\), we also need \(\sqrt{n}\ell_{n}^{-\underline{\varsigma}}\to 0\), which is implied by \(\sqrt{n(T-1)}\ell_{n}^{-\underline{\varsigma}}\to 0\) as \((n,T)\to\infty\).↩︎

  11. For illustration, let \(T\to\infty\) and set \(n=O(T^{13/5})\) and \(\ell_{n}=O(n^{1/5})\). With \(\underline{\varsigma}=p=6\), \(\frac{n}{T^{p/2}}=O(T^{-2/5})\) so Assumption 1 holds. Note that \(\ell_{n}^{2-\underline{\varsigma}}\to 0\) as \(\underline{\varsigma}>2\), \(\frac{\ell_{n}^{2}}{n}=O(n^{-3/5})=O(T^{-39/25})\), \(\frac{\ell_{n}}{\sqrt{n(T-1)}}=O(T^{-32/25})\), \(\sqrt{\frac{(T-1)\ell_{n}^{3}}{n}}=O(T^{-1/50})\) and \(\sqrt{n(T-1)}\ell_{n}^{3/2-\underline{\varsigma}}=O(T^{-27/50})\). Thus, this choice of \((n,T,\ell_{n})\) satisfies all necessary rate conditions.↩︎

  12. If the estimated \(\rho(\hat{A})\geq 1\), following the approach of [36], we can employ an Euclidean minimum distance estimator: \(\hat{\theta}_{c}=\arg\min_{\theta} \sum_{k=1}^{\ell_{\theta}} w_k(\hat{\theta}_{unc, k}-\theta_k)^2\) subject to \(\rho(A(\theta))\leq 1-\epsilon\), where \(\hat{\theta}_{unc, k}\) is the unconstrained estimator, and the weights satisfy \(\sum_{k}w_k=1\). For example, \(w_{k}=\operatorname{Var}(\hat{\theta}_{unc, k})/\sum_{l=1}^{\ell_{\theta}}\operatorname{Var}(\hat{\theta}_{unc, l})\). This approach is also discussed in nonparametric SAR models [22], [37].↩︎

  13. The Monte Carlo results indicate that the estimator performs well in the [var:hetei] case with \(T=10\). This validates our use of the [var:hetei] specification in the empirical application, where the sample size is comparable (\(T=11\)).↩︎

  14. Because the household survey in [1] was conducted as a single cross-section, consumption data are only available for the year 2001. We construct our economic distance matrix using these 2001 values. While using end-of-period data introduces potential endogeneity, we follow the methodology of [1], who utilizes this same 2001 cross-section to proxy for village characteristics across the entire panel. This approach relies on the reasonable assumption that the relative wealth distances between rural villages are structurally persistent over a single decade, even in the presence of transient weather shocks.↩︎

  15. Although the demographic data were collected during the 2001 survey described in [1], village-level ethnic composition is highly time-invariant. Therefore, we use these ethnic shares to construct the cultural distance matrix for the entire sample period, thereby avoiding the endogeneity concerns associated with time-varying economic variables.↩︎

  16. If \(0<\delta_{\epsilon}<1\), where \(\delta_{\epsilon}\) is in Assumption 3, then we need \(\delta>\frac{p-2}{p+4}\) to guarantee the CLT. Discussions are in footnote . Moreover, Assumption 1 is required for the joint limit theory where \((n,T) \to \infty\). For the fixed \(T\) case in Remark 3, this rate condition is not required.↩︎

  17. One may also consider settings in which \(G_k\) contains “stars”, so that the column sums grow with \(n\), see, e.g. [41][43]. We leave this extension for future research.↩︎

  18. For the finite \(T\) case, one may specify the initial condition that is endogenous. Such a specification is beyond the scope of this paper and deserves further study.↩︎

  19. We use the trivial \(\sigma\)-field to define \(\mathcal{F}_{0}\).↩︎

  20. We need to verify \(\sum_{l=1}^{N}\operatorname{E}\left(\mathcal{X}_{l}^{\circ2}\mathbf{1}[|\mathcal{X}_{l}^{\circ}|>e]\right)\stackrel{p}{\rightarrow} 0\) for some constant \(e\). By the elementary truncation bound \(\mathcal{X}_{l}^{\circ2}\mathbf{1}[|\mathcal{X}_{l}^{\circ}|>e]\leq e^{-\delta}|\mathcal{X}_{l}^{\circ}||^{2+\delta}\), and use the Markov inequality, we can verify the sufficient condition 17 .↩︎

  21. In Theorem 1(ii), we need \(\sqrt{\frac{(T-1)\ell_{n}^{3}}{n}}\to 0\) which implies \(\ell_{n}=o\big((\frac{n}{T})^{1/3}\big)\). Then, \(O\bigl(\frac{\ell_{n}^{2+\delta}}{(n(T-1))^{\delta}}\bigr)=o\bigl(n^{(1-\delta)/3}T^{-(1+2\delta)/3}\bigr)\). Using \(n=o(T^{p/2})\), \(n^{(1-\delta)/3}T^{-(1+2\delta)/3}\) goes to zero provided \((p/2)(1-\delta)<1+2\delta\). Therefore, if \(0<\delta<1\), we need \(\delta>\frac{p-2}{p+4}\).↩︎

  22. Note that \(r_{t}^{*}\) also contains \(\lambda\), however, it is easy to verify that the order of \(\|\frac{\partial r_{t}^{*}}{\partial\lambda_{j}'}\|_{\mathrm{sp}}=O_{p}(\ell_{n}^{-\underline{\varsigma}})\) by Assumption 4, which is negligible. Also, we need to divide \(n(T-1)\) to investigate the consistency and divide \(\sqrt{n(T-1)}\) to investigate the asymptotical normality. As shown in Lemma 17, the approximation terms will be negligible.↩︎