Precise Asymptotics for Spectral Methods in Mixed Generalized Linear Models4


Abstract

In a mixed generalized linear model, the goal is to learn multiple signals from unlabeled observations: each sample comes from exactly one signal, but it is not known which one. We consider the prototypical problem of estimating two statistically independent signals in a mixed generalized linear model with Gaussian covariates. Spectral methods are a popular class of estimators which output the top two eigenvectors of a suitable data-dependent matrix. However, despite the wide applicability, their design is still obtained via heuristic considerations, and the number of samples \(n\) needed to guarantee recovery is super-linear in the signal dimension \(d\). In this paper, we develop exact asymptotics on spectral methods in the challenging proportional regime in which \(n, d\) grow large and their ratio converges to a finite constant. This allows us optimize the design of the spectral method, and combine it with a simple linear estimator, to minimize the estimation error. Our characterization exploits a mix of tools from random matrices, free probability and the theory of approximate message passing algorithms. Numerical simulations for mixed linear regression and phase retrieval demonstrate the advantage enabled by our analysis over existing designs of spectral methods.

Spectral estimator, generalized linear models, mixed regression, high dimensional asymptotics, random matrix theory, Approximate Message Passing (AMP).

62E20, 62J05, 62J12.

1 Introduction↩︎

We consider the problem of learning multiple \(d\)-dimensional vectors from \(n\) unlabeled observations coming from a mixed generalized linear model (GLM): \[\begin{align} y_i &= q\paren{\inprod{a_i}{x_{\upsilon_i}^*} ,\eps_i}, \qquad i\in [n]=\{1, \ldots, n\}. \label{eqn:def-y-glmm-intro} \end{align}\tag{1}\] Here, \(x_1^*, \ldots, x_\ell^*\in\mathbb{R}^d\) are the \(\ell\) signals (regression vectors) to be recovered from the observation vector \(y=(y_1, \ldots, y_n)\in \mathbb{R}^n\) and the known design matrix \(A=[a_1, \ldots, a_n]^\top\in \mathbb{R}^{n\times d}\). For \(i \in [n]\), \(\eps_i\) is a noise variable, and \(\upsilon_i\) is an \([\ell]\)-valued latent variable, i.e., it indicates which signal each observation comes from, and is unknown to the statistician. The notation \(\inprod{\cdot}{\cdot}\) denotes the Euclidean inner product, and \(q:\mathbb{R}^2\to\mathbb{R}\) is a known link function. For \(\ell=1\), 1 reduces to a generalized linear model [1], which covers many widely studied problems in statistical estimation including linear regression, logistic regression, phase retrieval [2], [3], and 1-bit compressed sensing [4]. The regression model with \(\ell=1\) implicitly assumes a homogeneous population, in which a single regression vector suffices to capture the features of the entire sample. In practice, it is often the case that the observations may come from multiple sub-populations. Mixed GLMs offer a flexible solution in settings with unlabeled heterogeneous data, and have found applications in a variety of fields including biology, physics, and economics [5][8]. When \(q(g, \eps)=g+\eps\), 1 reduces to the widely studied mixture of linear regressions [9][17].

A natural approach to estimate the vectors \(x_1^*, \ldots, x_\ell^*\) from \(y\) and \(A\) is via the maximum-likelihood estimator (assuming a statistical model for \((\eps_i)_{i \in [n]}\) is available). However, the corresponding optimization problem is non-convex and NP-hard [13]. Thus, various low-complexity alternatives — mostly focusing on mixed linear regression — have been proposed: examples include expectation-maximization (EM) [10], [11], [18], alternating minimization [13], [15], [17], convex relaxation [19], moment descent methods [20], [21], and the use of tractable non-convex objectives [14], [22]. Many of these methods are iterative in nature and require a “warm start” with an initial guess correlated with the ground truth. Spectral methods are a popular way to provide such initialization [13]. A variety of estimators based on the spectral decomposition of data-dependent matrices or tensors have been proposed for mixed GLMs [12], [13], [23]. In this paper, we focus on a spectral method that estimates the \(\ell\) signals via the top-\(\ell\) principal eigenvectors of the following data-dependent matrix: \[\begin{align} D = \frac{1}{n} \sum_{i = 1}^n \cT(y_i) a_i a_i^\top \in\bbR^{d\times d} ,\label{eqn:def-mtx-t-and-d-intro} \end{align}\tag{2}\] where \(\mathcal{T}:\mathbb{R}\to\mathbb{R}\) is a suitably chosen preprocessing function. This spectral estimator with the preprocessing function \(\cT(y) = y^2\) was studied for mixed linear regression by Yi et al.[13], who showed that the signals can be accurately recovered when the number of observations \(n\) is of order \(d\log d\). Furthermore, existing theoretical results for all estimators (including spectral, alternating minimization and EM) require \(n\) to be of order at least \(d\log d\) to guarantee accurate recovery [12], [13], [20], [21], [23]. This leads to the following natural questions:

What is the optimal sample complexity of a spectral estimator based on 2 ?
Can we carry out a principled optimization of the preprocessing function \(\cT\)?

A simpler alternative to obtain an initial estimate is to use the linear estimator \[\begin{align} \frac{1}{n} \sum_{i = 1}^n \cL(y_i) a_i \;\in\bbR^d, \label{eqn:def-lin-estimator-intro} \end{align}\tag{3}\] where \(\cL: \reals \to \reals\) is a suitable preprocessing function. The performance analysis of this linear estimator for the mixed GLM can be carried out similarly to that for the non-mixed case (\(\ell=1\)); the analysis for the latter is given in [24] and in [25]. Thus, another natural question is:

What is the optimal way to combine a spectral estimator based on 2 and the linear estimator in 3 ?

1.1 Main contributions↩︎

In this paper, we resolve the questions above for the recovery of two independent signals \(x_1^*, x_2^*\) with a Gaussian design matrix \(A\). This is achieved by characterizing the high-dimensional limit of the joint empirical distribution of (i) the signals \(x_1^*, x_2^*\), (ii) the linear estimator in 3 , and (iii) spectral estimators based on the matrix in 2 . Our analysis holds in the proportional setting where \(n,d\to\infty\) with \(n/d\to\delta\in(0,\infty)\). That is, we consider the regime where the ratio between sample size and signal dimension tends to a constant, as opposed to most analyses of mixed GLMs in the literature which assume \(n = \Omega (d \log d)\). Our major findings are summarized as follows.

  • Our master theorem (1) characterizes the joint distribution of the linear estimator, the spectral estimator, and the signals in the high-dimensional limit. This joint distribution characterization holds for arbitrary preprocessing functions \(\cL,\cT\colon\bbR\to\bbR\) in [eqn:def-mtx-t-and-d-intro,eqn:def-lin-estimator-intro] (subject to mild regularity conditions). The limiting joint distribution is expressed as the law of a set of jointly Gaussian random variables whose covariance structure is explicitly derived in terms of the model and the preprocessing functions.

  • As an immediate consequence of the distributional characterization, we derive the normalized correlations (or ‘overlaps’) between the linear/spectral estimator and the signals (2/3). The linear estimator achieves a strictly positive overlap with each signal for any \(\delta > 0\), provided a strictly positive overlap can be attained for some \(\delta>0\). In contrast, for the spectral estimator, we identify a threshold (depending on the preprocessing function \(\cT\)) such that strictly positive overlap is attained as soon as \(\delta\) exceeds this threshold. In general, there is no clear winner between the spectral and the linear estimator, and which one performs better depends on the setting.

  • In fact, it is best to combine the linear and spectral estimators: our master theorem also allows us to compute the limiting overlap of a class of such combinations. In particular, the Bayes-optimal combination can be derived, which turns out to be linear in the two estimators due to the Gaussianity of their high-dimensional limits (1).

  • We determine the optimal preprocessing functions \(\cL^*, \cT_1^*,\cT_2^*\colon\bbR\to\bbR\) for the linear and spectral estimators that maximize the overlap between the estimator and each signal ([lem:linear-optimal-overlap,thm:opt-spec]). The optimal overlaps of linear and spectral estimators reveal intriguing behaviors of mixed models. In particular, there is a single function \(\cL^*\) that simultaneously maximizes the overlap between the linear estimator and each signal. In contrast, for the spectral method, one needs to employ two different functions \(\cT_1^*,\cT_2^*\) in order to achieve the maximal overlaps with \(x_1^*\), \(x_2^*\), respectively. Furthermore, the optimal overlap of the spectral estimator with each signal approaches \(1\) — the best possible value — as the aspect ratio \(\delta\) grows. We remark that the same is not true for the linear estimator: the optimal overlap with each signal remains strictly less than 1 even as \(\delta \to \infty\), as long as there is a strictly positive fraction of observations corresponding to each signal.

Our precise asymptotic analysis leads to a significant improvement over previous designs of spectral methods, as showcased in 1 for noiseless mixed linear regression. The continuous lines correspond to our theoretical predictions (“pred.”), which closely match the points coming from the simulations (“sim.”). The following methods are compared: (i) optimal spectral method (black), obtained from [thm:opt-spec]; (ii) optimal linear method (blue), obtained from [lem:linear-optimal-overlap]; (iii) combined estimator (“combo”) (red), obtained from 1; (iv) spectral estimator for mixed linear regression proposed in [13] (yellow); (v) spectral estimator which optimizes the overlap in the non-mixed setting (green), proposed in [26]. The spectral methods resulting from our sharp analysis (red, black) significantly outperform existing methods (green, yellow), especially for low values of \(\delta\). More details on the experimental setup and additional simulation results can be found in 4.

a
b

Figure 1: Noiseless mixed linear regression with mixing parameter (i.e., probability that a sample corresponds to \(x_1^*\)) \(\alpha = 0.6\). Overlaps with the first signal \(x_1^*\) (left) and the second signal \(x_2^*\) (right), computed via simulation (“sim.”) and the theoretical prediction (“pred.”), are plotted as a function of the aspect ratio \(\delta =n/d\). The signal dimension is \(d=2000\). We note that our optimal spectral estimator enables weak recovery of both signals at a smaller \(\delta\) than existing spectral estimators designed for non-mixed data. E.g., in the right panel, our optimal spectral estimator starts weakly recovering \(x_2^*\) when \(\delta > 3.1\) while other spectral estimators require at least \(\delta > 5.3\).. a — Recovery of \(x_1^*\), b — Recovery of \(x_2^*\)

1.1.0.1 Proof techniques

We exploit a combination of tools from free probability, random matrices and the theory of approximate message passing (AMP). Generalized approximate message passing (GAMP) refers to a family of iterative algorithms [27] with the following key feature: the joint distribution of the iterates is accurately tracked by a simple deterministic recursion, called state evolution. Our strategy to obtain the joint distribution of the linear/spectral estimators and the signals in the master theorem (1) is to design a GAMP that (i) outputs the linear estimator as the first iterate and (ii) then implements a power method, so that its fixed point corresponds to the spectral estimator. One challenge in the implementation of this strategy is that the state evolution of GAMP, in its original form for vanilla (non-mixed) GLMs, only records the correlation of its iterates with a single signal. To circumvent this issue, we equip GAMP with a state evolution recursion involving both signals, and run a pair of GAMP iterations converging to the first and second top eigenvector of the spectral matrix \(D\) in 2 , respectively. A second, even more fundamental, challenge is that, for the power method to converge to the desired eigenvector, a spectral gap between the corresponding eigenvalue and the rest of the spectrum is required. For non-mixed GLMs, the spectral analysis was carried out in earlier work [28], [29], which characterized the limiting eigenvalues of \(D\) as well as the overlaps using tools from random matrix theory. Here, the difficulty comes from the mixed effect of the model, leading to additional matrix terms which appear challenging to control. Our approach is to decompose \(D\) into the sum of two matrices, \(D_1\) and \(D_2\), each consisting of components only pertaining to the first and second signal, respectively. Now, \(D_1,D_2\) can be individually viewed as generated from a non-mixed GLM, hence their limiting spectra are well understood. The key observation is then that, by assuming both signals to be independent and uniformly distributed on the sphere, \(D_1\) and \(D_2\) become asymptotically free5. Thus, we are able to characterize the sum of these two spiked matrices by using techniques from free probability.

1.2 Related work↩︎

Mixtures of generalized linear models have been studied in machine learning as ‘hierarchical mixtures of experts’ [30]. Bayesian methods for this problem were investigated in [9], [31], [32]. Khalili and Chen [18] proposed a penalized likelihood approach for variable selection in mixed GLMs, showing consistency in the low-dimensional setting (where the dimension \(d\) is fixed as \(n\) grows). Städler et al.[11] analyzed \(\ell_1\)-penalized estimators for high-dimensional mixed linear regression (MLR). Zhang et al.[16] considered MLR with two sparse components, when the mixing proportion and the covariance structure of the covariates are unknown. The works [11], [16], [18] all use variants of the EM algorithm for optimizing a suitable penalized likelihood function. Balakrishnan et al.[33] and Klusowski et al.[34] obtained statistical guarantees on the performance of EM for a class of problems, including symmetric MLR with \(x_1^* = - x_2^*\). Variants of EM for symmetric MLR were also analyzed in [35][37]. Minimax lower bounds were obtained in [38], and statistical-computational gaps were recently studied in [39]. Kong et al.[40] studied MLR as a canonical example of meta-learning: in the setting where the number of signals (\(\ell\)) is large, they derived conditions under which a large number of signals with a few observations can compensate for the lack of signals with abundantly many observations. The prediction error of MLR in the non-realizable setting, where no generative model is assumed for the data, was studied in [41]. Chandrasekher et al.[42] analyzed the performance of iterative algorithms (not including AMP) for mixtures of GLMs. They provide a sharp characterization of the per-iteration error with sample-splitting in the regime \(n \asymp d\, \mathrm{polylog}(d)\), assuming a Gaussian design and a random initialization. An AMP estimator for mixed GLMs was recently studied in [43]. We emphasize that the focus of the current paper is not on using the AMP algorithm as an estimator for mixed GLMs. Rather, we use AMP as a proof technique to obtain a precise distributional characterization of the spectral estimator, and use this characterization to optimize its accuracy.

Spectral methods based on 2 were introduced in [44] for standard GLMs (non-mixed, with \(\ell=1\)). For the special case of phase retrieval, a series of works has provided increasingly refined bounds on the number of samples needed to guarantee signal recovery via the spectral method [45][47]. This type of analysis is based on matrix concentration inequalities, a technique that typically does not return exact values for the overlap between the signal and the estimate. More recently, an exact high-dimensional analysis for generalized linear models was carried out in [28], [29]. These works focus on the regime of interest in this paper: \(n\) and \(d\) growing at a proportional rate \(\delta\). This sharp analysis allows for the optimization of the preprocessing function: the choice of \(\cT\) minimizing the value of \(\delta\) (and, hence, the amount of data) needed to achieve a strictly positive overlap was provided in [29]; furthermore, the choice of \(\cT\) maximizing the overlap was provided in [26]. Going beyond the proportional regime in which \(n\) is linear in \(d\), bounds on the sample complexity required for moment methods (including spectral) to achieve non-vanishing overlap were recently obtained in [48]. The aforementioned analyses assume a Gaussian design matrix. Beyond this assumption, [49] provides precise asymptotics for design matrices sampled from the Haar distribution, and [50] studies rotationally invariant designs. Moving to the mixed regression setting (\(\ell>1\)), Yi et al.[13] proposed a spectral estimator based on 2 with \(\cT(y) = y^2\). The analysis is based on concentration inequalities and requires the number of samples \(n\) to be of order \(d\log d\) for accurate recovery. Estimators based on spectral decomposition of data-dependent tensors were proposed for MLR in [12] and for mixed GLMs in [23]. However, these methods require \(n\) to be of order at least \(d^3\) for accurate recovery. Our work is the first to establish exact asymptotics for a mixed GLM in the linear sample-size regime: \(n, d\to\infty\) with \(n/d\to\delta\in (0, \infty)\). To achieve this goal, our strategy differs from analyses of spectral methods in the non-mixed setting [28], [29] which reduce the study of the spectrum of \(D\) to that of a rank-1 perturbation. In contrast, our analysis is based on a combination of techniques from free probability and approximate message passing (AMP).

AMP is a family of iterative algorithms that has been applied to several problems in high-dimensional statistics, including estimation in linear models [51][53], generalized linear models [27], [54], [55], and low-rank matrix estimation [56][58], see also the review [59]. A key feature of AMP algorithms is that under suitable model assumptions, the empirical joint distribution of their iterates can be exactly characterized in the high-dimensional limit, in terms of a simple scalar recursion called state evolution. By taking advantage of this characterization, AMP methods have been used to derive exact high-dimensional asymptotics for convex penalized estimators such as LASSO [60], M-estimators [61], logistic regression [55], and SLOPE [62]. AMP algorithms have been initialized via spectral methods in the context of low-rank matrix estimation [63] and generalized linear models [64]. Furthermore, they have been used – in a non-mixed setting – to combine linear and spectral estimators [25]. A finite-sample analysis which allows the number of iterations to grow roughly as \(\log n\) (\(n\) being the ambient dimension) was put forward in [65], and the recent paper [66] improves this guarantee to a linear (in \(n\)) number of iterations. This could potentially allow to study settings in which \(\delta=n/d\) approaches the spectral threshold. The works on AMP discussed above all assume i.i.d. Gaussian matrices. A number of recent papers have proposed generalizations of AMP for the much broader class of rotationally invariant matrices, e.g., [67][74].

Finally, we mention the recent paper [75] that derived precise asymptotics of spectral estimators for multi-index models by generalizing the techniques in [28], [29]. However, [75] did not derive the joint distribution of spectral and linear estimators, or the (optimal) combination of the two. To the best of our knowledge, such results do not follow immediately from the pure random matrix theory-type results in [75], but require additional work to handle the correlation between spectral and linear estimators. These are achieved in the present work using a mix of tools from free probability and Approximate Message Passing (AMP).

2 Preliminaries↩︎

The \(i\)-th element in \(a\in\bbR^p\) is denoted by \(a_i\). If a vector has multiple subscripts, the component index is the last one. For a symmetric \(M\in\bbR^{p\times p}\), we denote by \(\mu_M\) its empirical spectral distribution. The (real) eigenvalues of \(M\) are \(\lambda_1(M)\ge\lambda_2(M)\ge\cdots\ge\lambda_p(M)\), and the corresponding eigenvectors are \(v_1(M),v_2(M),\cdots,v_p(M)\). The \((i,j)\)-th entry of \(M\) is denoted by \(M_{i,j}\). For a random variable \(X\), \(\supp(X)\) denotes the support of its density function. The orthogonal group in dimension \(p\) is \(\bbO(p)\mathrel{\vcenter{:}}=\curbrkt{O\in\bbR^{p\times p} : OO^\top = O^\top O = I_p}\). The unit sphere in dimension \(p\) is \(\bbS^{p-1} \mathrel{\vcenter{:}}= \curbrkt{x\in\bbR^p : \normtwo{x} = 1}\). For two distributions \(P\) and \(Q\), \(P\ot Q\) is their product distribution, and \(P^{\ot k}\) is the \(k\)-fold product distribution of \(P\).

2.0.0.1 Model

We consider a two-component mixed GLM with signal vectors \(x_1^*,x_2^*\in\bbS^{d-1}\), covariate vectors \(a_1,\cdots,a_n\in\bbR^d\) , and a known link function \(q\colon\bbR^2\to\bbR\). Let \(P_\eps\) be a noise distribution over \(\bbR\). The \(n\) observations \(y_1,\cdots,y_n\in\bbR\) are generated as: \[\begin{align} y_i &= q\paren{\inprod{a_i}{\eta_ix_1^* + (1-\eta_i) x_2^*} ,\, \eps_i}, \qquad i \in [n] . \label{eqn:def-y-glmm} \end{align}\tag{4}\] Here, the vector of latent variables \(\veta \mathrel{\vcenter{:}}= (\eta_1,\cdots,\eta_n) \sim \bern(\alpha)^\tn\) indicates which signal is selected by each observation, and is unobserved. The latent variable vector \(\veta\), the signals \(x_1^*, x_2^*\), the covariate vectors \(a_1, \ldots, a_n\), and the noise vector \(\veps \mathrel{\vcenter{:}}= (\eps_1,\cdots,\eps_n)\sim P_\eps^\tn\) are mutually independent. Then, 4 is equivalent to \[y_i\mid \inprod{a_i}{\eta_ix_1^* + (1-\eta_i) x_2^*} \sim p(\cdot \mid \inprod{a_i}{\eta_ix_1^* + (1-\eta_i) x_2^*}),\label{eq:condlaw}\tag{5}\] where \(p(\cdot |g)\) denotes the distribution of \(q(g, \eps)\) for a fixed \(g\in\mathbb{R}\) and \(\eps\sim P_\eps\) independent of \(g\). The design matrix is \(A=[a_1^\top, \ldots, a_n^\top]^\top\in\bbR^{n\times d}\). Given \(A\), upon observing \(y=( y_1,\cdots,y_n)\in\bbR^n\), our goal is to estimate \(x_1^*\) and \(x_2^*\). Given a pair of estimators \(\wh{x}_1 = \wh{x}_1(y,A), \, \wh{x}_2 = \wh{x}_2(y,A)\in\bbR^d\), we measure the performance via their overlap with the respective signals: \[\begin{align} \lim_{d\to\infty} \frac{\abs{\inprod{\wh{x}_1}{x_1^*}}}{\normtwo{\wh{x}_1}\normtwo{x_1^*}} , \quad \lim_{d\to\infty} \frac{\abs{\inprod{\wh{x}_2}{x_2^*}}}{\normtwo{\wh{x}_2}\normtwo{x_2^*}} . \notag \end{align}\] Throughout the paper, the following assumptions are imposed.

As for [itm:assump-signal-distr], choosing signals uniform on the sphere corresponds to having no structural information about them. This requirement is natural, since spectral methods are typically unable to exploit prior information about the signal. We expect that all results of the present paper hold for the more relaxed setting where the signals \(x_1^*, x_2^*\) are independent of the design matrix \(A\) and the noise vector \(\veps\), and satisfy \[\begin{align} && \lim_{d\to\infty} \normtwo{x_1^*} = \lim_{d\to\infty} \normtwo{x_2^*} &= 1 , & \lim_{d\to\infty} \inprod{x_1^*}{x_2^*} &= 0 . & & \notag \end{align}\] In particular, we believe that the assumption that the signals are uniformly distributed on the sphere is not required. In that case, 35 in the argument of the reduction to the free sum of two random matrices is no longer an exact equality (in distribution). However, we still expect asymptotic freeness to hold in the proportional limit. We leave its formal justification to future work. Understanding the effect of correlation on the performance of spectral estimators and the design of the optimal preprocessing function is an exciting future direction. [itm:assump-alpha] is without loss of generality: if \(0<\alpha<1/2\), one can simply interchange the roles of \(x_1^*\) and \(x_2^*\). When \(\alpha = 1/2\), the top two eigenvectors given by the spectral method correspond to the same limiting eigenvalue as \(n\to \infty\). These eigenvectors provide an estimate on the space spanned by \(x_1^*, x_2^*\), and in order to estimate the individual signals, an additional \(1\)-dimensional grid search is required. Provided this extra step is carried out, our results still apply, see [rk:alpha-half-master,rk:alpha-half-rmt,rk:alpha-half-amp]. [itm:assump-gaussian-design] is common in the related literature [13], [26], [28], [29], and the potential universality beyond Gaussian design matrices is discussed in 6.

2.0.0.2 Linear estimator

Given the preprocessing function \(\cL\colon\bbR\to\bbR\), the linear estimator is \[\begin{align} \xlin &\mathrel{\vcenter{:}}= \frac{1}{n} A^\top \cL(y) = \frac{1}{n} \sum_{i = 1}^n \cL(y_i) a_i \in\bbR^d , \label{eqn:def-lin-estimator} \end{align}\tag{6}\] where \(\cL\) is applied component-wise, i.e., \(\cL(y) = (\cL(y_1),\cdots,\cL(y_n))\). Let \(Y\) be defined as \[\begin{align} Y=q(G,\eps), \quad \text{ where } (G,\eps)\sim\cN(0,1)\otimes P_\eps . \label{eqn:def-rv-y} \end{align}\tag{7}\] We make the following assumption on \(\cL\).

The first condition guarantees that the linear method w.r.t.\(\cL\) attains positive overlaps with both signals, and the second condition is rather mild and purely technical.

2.0.0.3 Spectral estimator

Let \(\cT\colon\bbR\to\bbR\) be a preprocessing function, and consider \[\begin{align} T &\mathrel{\vcenter{:}}= \mathop{\mathrm{diag}}(\cT(y)) \in\bbR^{n\times n} , \quad D \mathrel{\vcenter{:}}= \frac{1}{n} A^\top TA = \frac{1}{n} \sum_{i = 1}^n \cT(y_i) a_ia_i^\top \in\bbR^{d\times d} , \label{eqn:def-mtx-t-and-d} \end{align}\tag{8}\] where \(\cT(y) = (\cT(y_1),\cdots,\cT(y_n))\). Then, the spectral method computes the top two eigenvectors \(v_1(D),v_2(D)\) of \(D\) as estimates of \(x_1^*,x_2^*\). We make the following assumption on \(\cT\).

In words, we require \(\cT\) to be bounded, with strictly positive upper edge of its range. A bounded preprocessing function is also required in the non-mixed setting [28], [29]. The requirement on the \(\sup\) to be strictly positive is purely technical, and it simply rules out the trivial cases in which the spectral matrix \(D\) is all-zero with high probability. [itm:assump-preproc-spec] is satisfied by the preprocessing function that maximizes the overlap (cf.[thm:opt-spec]).

3 Main results↩︎

We start by defining a few auxiliary quantities. Let \(\delta_1 = \alpha\delta,\delta_2 = (1-\alpha)\delta\), and \(Z = \cT(Y)\), with \(Y\) as defined in 7 . Define \(\phi\colon(\sup\supp(Z),\infty)\to\bbR\) and \(\psi\colon(\sup\supp(Z),\infty)\times(0,\infty)\to\bbR\) as \[\begin{align} \phi(\lambda) &\mathrel{\vcenter{:}}= \lambda\,\expt{\frac{ZG^2}{\lambda - Z}} , \tag{9} \\ \psi(\lambda;\Delta) &\mathrel{\vcenter{:}}= \lambda\paren{\frac{1}{\Delta} + \expt{\frac{Z}{\lambda - Z}}} . \tag{10} \end{align}\] In what follows, we will set the second argument \(\Delta\) of \(\psi\) to \(\delta, \delta_1\) and \(\delta_2\). For \(\Delta\in\{\delta, \delta_1, \delta_2\}\), let \(\ol\lambda(\Delta)>\sup\supp(Z)\) be the minimum point of \(\psi(\cdot;\Delta)\), i.e., \[\begin{align} \ol\lambda(\Delta) &\mathrel{\vcenter{:}}= \argmin_{\lambda>\sup\supp(Z)} \psi(\lambda;\Delta) . \label{eqn:def-lam-bar} \end{align}\tag{11}\] Since \(\psi\) is convex in its first argument (see 4), this minimum point is obtained by setting the derivative to \(0\). Furthermore, define \(\zeta\colon(\sup\supp(Z),\infty)\times(0,\infty)\to\bbR\) as \[\begin{align} \zeta(\lambda; \Delta) &\mathrel{\vcenter{:}}= \psi(\max\{\lambda,\ol\lambda(\Delta)\}; \Delta) . \label{eqn:def-zeta-i} \end{align}\tag{12}\] Finally, for \(i\in\{1, 2\}\), by [29], the equation \(\zeta(\lambda;\delta_i) = \phi(\lambda)\) admits a unique solution in \(\lambda\in(\sup\supp(Z),\infty)\) which we call \(\lambda^*(\delta_i)\). The functions \(\psi(\lambda;\Delta),\phi(\lambda),\zeta(\lambda;\Delta)\) together with the parameters \(\lambda^*(\Delta),\ol\lambda(\Delta)\) are plotted in 2 for \(\Delta\in\{\delta,\delta_1,\delta_2\}\). Some convexity and monotonicity properties of these functions can be found in 4.

Figure 2: Plot of \psi(\lambda;\Delta),\phi(\lambda),\zeta(\lambda;\Delta) as functions of \lambda with \Delta\in\{\delta,\delta_1,\delta_2\}.

The empirical distribution of a vector \(u \in \reals^d\) is given by \(\frac{1}{d}\sum_{i=1}^d \delta_{u_i}\), where \(\delta_{u_i}\) denotes a Dirac delta mass on \(u_i\). Similarly, the joint empirical distribution of the rows of a matrix \((u^1, u^2, \ldots, u^t) \in \reals^{d \times t}\) is \(\frac{1}{d} \sum_{i=1}^d \delta_{(u^1_i, \ldots, u^t_i)}\). Our master theorem is an exact characterization in the high-dimensional limit of the joint empirical distribution of the rows of the signals, the linear estimator, and the spectral estimators. In particular, we show that this joint empirical distribution converges to the law of a Gaussian random vector with a specified covariance matrix. The result is stated in terms of the following parameters: the asymptotic correlations \(\rho_1^{\lin}, \rho_2^{\lin}\) between the linear estimator and the two signals; the asymptotic normalized Euclidean norm \(n^{\lin}\) of the linear estimator; and the asymptotic correlations \(\rho_1^{\spec}, \rho_2^{\spec}\) between the spectral estimators and the two signals. The formulas for these quantities are: \[\begin{align} & n^{\lin} \mathrel{\vcenter{:}}= \bigg( (\alpha^2+(1-\alpha)^2) \expt{G\cL(Y)}^2 + \frac{\expt{\cL(Y)^2}}{\delta} \bigg)^{\frac{1}{2}}, \\ & \rho_1^{\lin} \mathrel{\vcenter{:}}= \frac{\alpha \expt{G\cL(Y)}}{n^{\lin}}, \quad \rho_2^{\lin} \mathrel{\vcenter{:}}= \frac{ (1-\alpha) \expt{G\cL(Y)}}{n^{\lin}}, \tag{13} \\ \begin{aligned} & \rho_1^{\spec} \mathrel{\vcenter{:}}= \paren{ \frac{\frac{1}{\delta} - \expt{\paren{\frac{Z}{\lambda^*(\delta_1) - Z}}^2}}{\frac{1}{\delta} + \alpha\expt{\paren{\frac{Z}{\lambda^*(\delta_1) - Z}}^2(G^2 - 1)}} }^{\frac{1}{2}} , \\ &\rho_2^{\spec} \mathrel{\vcenter{:}}= \paren{ \frac{\frac{1}{\delta} - \expt{\paren{\frac{Z}{\lambda^*(\delta_2) - Z}}^2}}{\frac{1}{\delta} + (1-\alpha)\expt{\paren{\frac{Z}{\lambda^*(\delta_2) - Z}}^2(G^2 - 1)}} }^{\frac{1}{2}} . \end{aligned} \tag{14} \end{align} \tag{15}\] 1 is stated in terms of pseudo-Lipschitz test functions. A function \(\Psi: \reals^m \to \reals\) is pseudo-Lipschitz of order \(k \geq 1\), denoted \(\Psi \in \PL(k)\), if there is a constant \(C > 0\) such that \[\normtwo{\Psi(x)-\Psi(y)} \le C (1 + \| x \|_2^{k-1} + \| y \|_2^{k-1} )\normtwo{x - y}, \label{eq:PL295prop}\tag{16}\] for all \(x,y \in \mathbb{R}^m\). Examples of pseudo-Lipschitz functions of order two are: \(\Psi(u)=u^2\) and \(\Psi(u,v) = \abs{uv}\), for \(u, v \in \reals\). We consider pseudo-Lipschitz test functions of order two, as those suffice to compute the asymptotic overlaps between the signals and the various estimators. One could extend 1 to test functions in \(\PL(k)\) for \(k>2\), at the cost of a more involved argument and an additional assumption on the finiteness of the moments of \(P_\eps\).

Theorem 1 (Master theorem on joint distribution). Consider the setting of 2, and let [itm:assump-signal-distr,itm:assump-alpha,itm:assump-noise-distr,itm:assump-gaussian-design,itm:assump-proportional,itm:assump-preproc-lin,itm:assump-preproc-spec] hold. Define the following rescaled vectors of Euclidean norm \(\sqrt{d}\): \(x^{\lin} = \sqrt{d} \, \xlin/\normtwo{\xlin}\), and for \(i\in \{1, 2\}\), \(\ol{x}^*_i = \sqrt{d} x_i^*\), \(x^{\spec}_i=s_i\sqrt{d} \, v_i(D)\), where the sign \(s_i\in\{-1,1\}\) is chosen such that \(\inprod{s_iv_i(D)}{x_i^*}\ge0\). Then, the following holds almost surely for any \(\PL(2)\) function \(\Psi: \reals^3 \to \reals\). If \(\lambda^*(\delta_1) > \ol\lambda(\delta)\), then \[\begin{align} & \lim_{d \to \infty} \, \frac{1}{d} \sum_{i=1}^d \Psi(\ol{x}^*_{1,i}, x^{\lin}_i, x^{\spec}_{1,i}) = \expt{\Psi(X_1, \, \rho_1^{\lin} X_1 + \rho_2^{\lin} X_2 + W^{\lin} ,\, \rho_1^{\spec} X_1 + W_1^{\spec}) }. \label{eq:psiX1joint} \end{align}\qquad{(1)}\] Similarly, if \(\lambda^*(\delta_2) > \ol\lambda(\delta)\), then \[\begin{align} & \lim_{d \to \infty} \, \frac{1}{d} \sum_{i=1}^d \Psi(\ol{x}^*_{2,i}, x^{\lin}_i, x^{\spec}_{2,i}) = \expt{ \Psi(X_2, \, \rho_1^{\lin} X_1 + \rho_2^{\lin} X_2 + W^{\lin} ,\, \rho_2^{\spec} X_2 + W_2^{\spec}) }. \label{eq:psiX2joint} \end{align}\qquad{(2)}\] Here \((X_1, X_2) \sim \cN(0,1)^{\ot2}\), the pairs \((W^{\lin}, W_1^{\spec})\) and \((W^{\lin}, W_2^{\spec})\) are independent of \((X_1, X_2)\) and each pair is jointly Gaussian with zero mean and covariance given by \[\begin{align} & \E[(W^{\lin})^2]=1-(\rho_1^{\lin})^2-(\rho_2^{\lin})^2, \quad \E[(W_1^{\spec})^2]=1-(\rho_1^{\spec})^2, \quad \E[(W_2^{\spec})^2]=1-(\rho_2^{\spec})^2 , \\ & \E[W^{\lin} W_1^{\spec}] = \frac{\alpha \rho_1^{\spec}}{n^{\lin}} \expt{ \frac{ G\cL(Y) Z}{\lambda^*(\delta_1) -Z}}, \quad \E[W^{\lin} W_2^{\spec}] = \frac{(1-\alpha) \rho_2^{\spec}}{n^{\lin}} \expt{ \frac{G\cL(Y) Z}{\lambda^*(\delta_2) -Z}}. \end{align} \notag\]

The outline of the argument is presented in 5. The full proof is given in 10, and it relies on the characterization of the eigenvalues of \(D\) carried out in 2, which is stated and proved in 9.

?? is equivalent to the statement that the joint empirical distribution of \((\ol{x}^*_{1}, x^{\lin}, x^{\spec}_{1})\) converges in Wasserstein-2 distance to the joint law of \((X_2, \, \rho_1^{\lin} X_1 + \rho_2^{\lin} X_2 + W^{\lin},\, \rho_1^{\spec} X_1 + W_1^{\spec})\). The equivalence between convergence of empirical distributions in Wasserstein distance and convergence of empirical averages of pseudo-Lipschitz functions is proved in [59].

The validity of the description of the joint law of the first signal and the linear/spectral estimators in ?? relies on two assumptions: \(\expt{G\cL(Y)}\ne0\) for the linear estimator, and \(\lambda^*(\delta_1)>\ol\lambda(\delta)\) for the spectral one. They guarantee that both estimators achieve non-zero asymptotic overlaps with \(x_1^*\), i.e., \(\rho_1^{\lin}\ne0\) and \(\rho_1^{\spec}>0\). If either condition is not satisfied, a conclusion similar to ?? still holds with \(\Psi\colon\bbR\times\bbR\to\bbR\) only taking \(x_1^*\) and the non-trivial estimator as inputs. Specifically, if only the linear estimator is effective, then we terminate GAMP in 39 after one step \(t=0\) and obtain the distributional characterization; if only the spectral estimator is effective, then the initializer in 44 ensures that the same proof goes through without modifications, again leading to the desired conclusion. An analogous argument holds for the second signal.

We have been using \(v_i(D)\) to estimate \(x_i^*\), for \(i\in\{1,2\}\). In fact, \(v_1(D)\) is asymptotically uncorrelated with \(x_2^*\) (and \(v_2(D)\) asymptotically uncorrelated with \(x_1^*\)), according to the characterization of the asymptotic distribution of the top two eigenvectors in 10.1. Intuitively, this phenomenon arises due to the orthogonality between the two signals, and fails to hold otherwise. For instance, if \(\lim_{d\to\infty} \inprod{x_1^*}{x_2^*} = \rho \ne 0\), then both the first and second eigenvectors are asymptotically correlated with both signals. This can be formally justified by specializing [75] to mixtures of GLMs and is numerically corroborated in [75].

As the eigenvectors of a matrix are insensitive to sign flip, the spectral estimators \(x^{\spec}_1,x^{\spec}_2\) are defined up to a change of sign. In 1, we pick the signs so that the resulting overlaps \(\rho_1^{\spec}, \rho_2^{\spec}\) are positive. In practice, there is a simple way to resolve the sign ambiguity: one can match the sign of \(\expt{(\rho_1^{\lin} X_1 + \rho_2^{\lin} X_2 + W^{\lin})\, (\rho_i^{\spec} X_i + W_i^{\spec})}\) with that of the scalar product \(\inprod{x^{\lin}}{x^{\spec}_i}\), as the latter can be computed empirically (without knowing \(x_1^*,x_2^*\)).

Even though we assume \(\alpha\in(1/2,1)\) (see [itm:assump-alpha]), the conclusion of 1 still holds for \(\alpha = 1/2\) with a slight modification in the definition of the spectral estimators. In this case, as \(n\to \infty\) the top two eigenvectors given by the spectral method correspond to the same limiting eigenvalue. These eigenvectors, \(v_1(D)\) and \(v_2(D)\), estimate the subspace spanned by \(x_1^*, x_2^*\). To estimate each individual signal, we search for a vector in \(\spn\curbrkt{v_1(D),v_2(D)}\) whose correlation with \(x^{\lin}\) is closest to the theoretical prediction from 1. Indeed, let \(x_1^{\spec},x_2^{\spec}\) be defined as \[\begin{align} x_i^{\spec} &\mathrel{\vcenter{:}}= \argmin_{v\,\in\,\spn\curbrkt{v_1(D),v_2(D)}\cap\sqrt{d}\,\bbS^{d-1}} \abs{\frac{\inprod{v}{x^{\lin}}}{\sqrt{d}} - \paren{\rho_i^{\lin} \rho_i^{\spec} + \expt{W^{\lin} W_i^{\spec}}}} , \quad \text{for } i\in\{1,2\} . \label{eqn:xspec-alpha-half} \end{align}\tag{17}\] Then, [eq:psiX1joint,eq:psiX2joint] hold, provided \(\expt{G\cL(Y)} \ne 0\) (which guarantees that the linear estimator attains nonzero overlaps; see [itm:assump-preproc-lin] and 13 ). We stress that 17 is computable in practice since it only involves \(x^{\lin}\) and theoretical predictions. If \(x^{\lin}\) is ineffective (which is the case, for example, in mixed phase retrieval, as mentioned in 7), a similar grid search can still be performed if the statistician is given as side information a vector with known correlation with a signal. The reader is referred to [rk:alpha-half-rmt,rk:alpha-half-amp] for the adaptation of our proofs to the case \(\alpha=1/2\).

Equipped with 1, we can combine the linear and spectral estimators to improve the performance in the recovery of \(x_1^*\) and \(x_2^*\). Formally, consider the (rescaled) linear and spectral estimators \(x^{\lin}\in\sqrt{d}\,\bbS^{d-1}\) and \(x_1^{\spec},x_2^{\spec}\in\sqrt{d}\,\bbS^{d-1}\). Define \[\begin{align} X^{\lin}\mathrel{\vcenter{:}}=\rho_1^{\lin} X_1 + \rho_2^{\lin} X_2 + W^{\lin} , \quad X_1^{\spec}\mathrel{\vcenter{:}}=\rho_1^{\spec} X_1 + W_1^{\spec} , \quad X_2^{\spec}\mathrel{\vcenter{:}}=\rho_2^{\spec} X_2 + W_2^{\spec} . \label{eqn:def-rv-xlin-xspec} \end{align}\tag{18}\] 1 states that the joint empirical distribution of the estimators \((x^{\lin},x_1^{\spec},x_2^{\spec})\) converges to the law of \((X^{\lin}, X_1^{\spec}, X_2^{\spec} )\). For \(i\in\{1,2\}\), define the set of functions \[\begin{align} \cC_i &\mathrel{\vcenter{:}}= \curbrkt{C_i\colon\bbR\times\bbR\to\bbRs.t.\expt{C_i(X^{\lin}, X_i^{\spec})^2}\in(0,\infty)} . \label{eqn:set-comb-fun-1} \end{align}\tag{19}\] Then, for any \(C_i\in\cC_i\), the combined estimator \(x_i^{\comb}\) is defined as \[\begin{align} x^{\comb}_i \mathrel{\vcenter{:}}= C_i(x^{\lin}, x_i^{\spec}), \label{eqn:def-combo-estimator} \end{align}\tag{20}\] where \(C_i\) acts on its inputs component-wise, i.e., \(x_{i,j}^{\comb} = C_i(x_j^{\lin}, x_{i,j}^{\spec})\) for any \(j\in[d]\). Now, ?? reduces the vector problem of estimating \(x_i^*\) given \((x^{\lin}, x^{\spec}_i)\) to the scalar problem of estimating \(X_i\) from \(X^{\lin}\) and \(X_i^{\spec}\). The Bayes-optimal combined estimator that minimizes the expected squared error for this scalar problem is \(\expt{X_i \mid X^{\lin}, X_i^{\spec}}\). Recalling from 1 that \((X_i, X^{\lin}, X_i^{\spec})\) are jointly Gaussian, the Bayes-optimal combined estimator is a linear combination of \((X^{\lin}, X_i^{\spec})\). The performance of this combined estimator is formalized in the following corollary, whose proof is given in 11.

Corollary 1 (Bayes-optimal linear-spectral combination). Consider the setting of 1. For \(i\in\{1,2\}\), define \(C_i^*\colon\bbR\times\bbR\to\bbR\) as follows: \[\begin{align} C^*_i(X^{\lin}, X_i^{\spec}) &\mathrel{\vcenter{:}}= \expt{X_i \mid X^{\lin}, \, X_i^{\spec}} = \frac{1}{1-\nu_i^2} \, \paren{\xi_i X^{\lin} \, + \, \zeta_i X_i^{\spec}}, \label{eqn:opt-combinator-1} \end{align}\qquad{(3)}\] where \[\begin{align} \nu_i \mathrel{\vcenter{:}}= \rho_i^{\lin} \rho_i^{\spec} + \expt{W^{\lin} W_i^{\spec}} , \quad \xi_i \mathrel{\vcenter{:}}= \rho_i^{\lin} - \rho_i^{\spec}\nu_i , \quad \zeta_i \mathrel{\vcenter{:}}= \rho_i^{\spec} - \rho_i^{\lin}\nu_i . \notag \end{align}\] For \(i\in\{1,2\}\), let \(x_i^{\comb}\) be the combined estimators defined in 20 w.r.t.\(C_i^*\), respectively. Then, almost surely we have \[\begin{align} \lim_{d\to\infty} \frac{\abs{\inprod{x_i^{\comb}}{x_i^*}}}{\normtwo{x_i^{\comb}}\normtwo{x_i^*}} &= \frac{1}{1 - \nu_i^2} \paren{\xi_i^2 + \zeta_i^2 + 2\xi_i\zeta_i\paren{\rho_i^{\lin}\rho_i^{\spec} + \expt{W^{\lin}W_i^{\spec}}}}^{1/2} \eqqcolon \mathsf{OL}_i^{\comb} . \notag \end{align}\] Furthermore, for any \((C_1,C_2)\in\cC_1\times\cC_2\), the corresponding combined estimators \(\wt{x}_1^{\comb},\wt{x}_2^{\comb}\) defined w.r.t.\(C_1,C_2\) through 20 , satisfy \[\begin{align} \lim_{d\to\infty} \frac{\abs{\inprod{\wt{x}_i^{\comb}}{x_i^*}}}{\normtwo{\wt{x}_i^{\comb}}\normtwo{x_i^*}} &= \frac{\abs{\expt{X_i C_i(X^{\lin}, X_i^{\spec})}}}{\sqrt{\expt{C_i(X^{\lin}, X_i^{\spec})^2}}} \le \mathsf{OL}_i^{\comb} , \qquad i\in\{1,2\}. \notag \end{align}\]

3.1 Linear estimator↩︎

1 allows us to derive the asymptotic overlap of each signal with the linear estimator in 6 .

Corollary 2 (Overlaps, linear). Consider the setting of 2, and let [itm:assump-signal-distr,itm:assump-alpha,itm:assump-noise-distr,itm:assump-gaussian-design,itm:assump-proportional,itm:assump-preproc-lin] hold. Then, almost surely, \[\begin{align} \lim_{d \to \infty} \, \frac{\inprod{\xlin}{x_i^*}}{\normtwo{\xlin}\normtwo{x_i^*}} & = \rho_i^{\lin} \, , \quad i\in\{1,2\}. \label{eqn:overlap-lin-1-main} \end{align}\qquad{(4)}\]

Proof. Choose \(\Psi(a,b,c) = ab\), and note that \(\Psi\in\PL(2)\). Then, as \(\normtwo{\xlin} = \normtwo{\ol{x}^*_i}= \sqrt{d}\), the left side of [eq:psiX1joint,eq:psiX2joint] recovers the overlaps in ?? for \(i=1,2\), and the right sides of [eq:psiX1joint,eq:psiX2joint] become \(\rho_1^{\lin},\rho_2^{\lin}\) (defined in 13 ). ◻

From ?? and the definitions of \(\rho_1^{\lin}, \rho_2^{\lin}\) in 13 , we have that the linear estimator achieves positive overlap with each signal for any positive \(\delta\), as long as \(\expt{G\cL(Y)}>0\). As \(\delta\to\infty\), the limiting overlaps approach \(\sqrt{\frac{\alpha^2}{\alpha^2 + (1-\alpha)^2}}\) and \(\sqrt{\frac{(1-\alpha)^2}{\alpha^2 + (1-\alpha)^2}}\), and they are strictly less than \(1\) for any \(\alpha\in (1/2,1)\). In contrast, the overlap of the spectral estimator becomes positive only when \(\delta\) exceeds a certain threshold (see [rk:univ-lb-spec-thr]). However, once this threshold is exceeded, the spectral estimator yields overlaps approaching \(1\) as \(\delta\) grows (see [rk:overlap-approach-1]). We also note that beyond the spectral threshold, the Bayes-optimal combination of the linear and spectral estimators has a larger overlap than either of the individual estimators (see Figure 1).

Using the limiting overlap of a linear estimator in 2, we can optimize the performance over the choice of \(\cL\) (subject to [itm:assump-preproc-lin]). Let \[\begin{align} \cI &\mathrel{\vcenter{:}}= \curbrkt{\cL\colon\bbR\to\bbR \, \text{ Lipschitz} { s.t. } \expt{G\cL(Y)}\ne0, \;\expt{\abs{G\cL(Y)}}<\infty} \label{eqn:def95I} \end{align}\tag{21}\] be the set of functions \(\cL\) satisfying [itm:assump-preproc-lin]. For \(i\in\{1,2\}\) and \(\delta\in(0,\infty)\), define the optimal overlaps among linear estimators as \[\begin{align} \mathsf{OL}_i^{\lin} &\mathrel{\vcenter{:}}= \sup_{\cL\in\cI} \rho_i^{\lin} . \notag \end{align}\] Furthermore, if \(\cI= \emptyset\), we set \(\mathsf{OL}_1^{\lin}=\mathsf{OL}_2^{\lin}=0\). In words, \(\mathsf{OL}_i^{\lin}\) (\(i\in\{1,2\}\)) is the largest overlap with the \(i\)-th signal that can be achieved by a linear estimator. Then, we have the following characterization of the optimal overlaps. The proof is contained in 12.

Consider the setting of 2, and let [itm:assump-signal-distr,itm:assump-alpha,itm:assump-noise-distr,itm:assump-gaussian-design,itm:assump-proportional] hold. Assume further that \[\begin{align} \int_{\supp(Y)} \frac{\expt{Gp(y|G)}^2}{\expt{p(y|G)}} \diff y &\in (0,\infty) , \label{eqn:cond-lin-eff} \end{align}\tag{22}\] where \(p(y|g)\) is the conditional law in 5 and the expectation is taken w.r.t.\(G\sim\cN(0,1)\). Then, for any \(\delta\in(0,\infty)\), writing \(\alpha_1:=\alpha\) and \(\alpha_2:=(1-\alpha)\), we have \[\begin{align} \mathsf{OL}_i^{\lin} &= \paren{\frac{\alpha_1^2+ \alpha_2^2}{\alpha_i^2} + \frac{1}{\alpha_i^2\delta} \cdot \frac{1}{\int_{\supp(Y)} \frac{\expt{Gp(y|G)}^2}{\expt{p(y|G)}} \diff y}}^{-1/2}, \qquad i \in \{1,2\}. \label{eq:RHSinsp} \end{align}\tag{23}\] Moreover, define \(\cL^*\colon\bbR\to\bbR\) as \[\begin{align} \cL^*(y) &= \frac{\expt{Gp(y|G)}}{\expt{p(y|G)}} . \notag \end{align}\] Then, \(\cL^*\in\cI\) and for any \(\delta\in(0,\infty)\), both \(\mathsf{OL}_1^{\lin},\mathsf{OL}_2^{\lin}\) are simultaneously achieved by \(\cL^*\).

22 ensures that the linear estimator asymptotically achieves strictly positive overlap with the signals. In fact, if \[\begin{align} \int_{\supp(Y)} \frac{\expt{Gp(y|G)}^2}{\expt{p(y|G)}} \diff y &= 0 , \notag \end{align}\] then, from the RHS of 23 , we obtain that \(\mathsf{OL}_1^{\lin}=\mathsf{OL}_2^{\lin}=0\) for any \(\delta\in(0,\infty)\). This is the case for mixed phase retrieval, as mentioned in 7. We note that the condition in 22 also appears in the non-mixed setting (see Appendix C.1 of [25]).

3.2 Spectral estimator↩︎

The limiting value of the overlaps for the spectral estimator can be obtained similarly to 2.

Corollary 3 (Overlaps, spectral). Consider the setting of 2, and let [itm:assump-signal-distr,itm:assump-alpha,itm:assump-noise-distr,itm:assump-gaussian-design,itm:assump-proportional,itm:assump-preproc-spec] hold. Then, for \(i\in\{1,2\}\), if \(\lambda^*(\delta_i) > \ol\lambda(\delta)\), we have that, almost surely, \[\begin{align} \lim_{d\to\infty} \frac{\abs{\inprod{v_i(D)}{x_i^*}}}{\normtwo{v_i(D)}\normtwo{x_i^*}} &= \rho_i^{\spec} . \label{eqn:overlap-spectral-1-main} \end{align}\qquad{(5)}\]

We focus here on the recovery of the first signal, and an analogous discussion is valid for the second one. As \(\lambda^*(\delta_1)\) approaches \(\ol\lambda(\delta)\) from above, the RHS of ?? tends to \(0\). Indeed, as \(\lambda^*(\delta_1)\searrow \ol\lambda(\delta)\), one can readily verify that \(\expt{\paren{\frac{Z}{\lambda^*(\delta_1) - Z}}^2}\nearrow \frac{1}{\delta}\) and consequently the numerator of \(\rho_1^{\spec}\) (cf.@eq:eqn:def-rho-spec-12 ) decreases to \(0\). Furthermore, in the non-mixed setting (\(\alpha=1\)), the analysis of [28], [29] gives that, when \(\lambda^*(\delta) < \ol\lambda(\delta)\), the corresponding overlap vanishes. While we do not formally prove that the condition \(\lambda^*(\delta_1) > \ol\lambda(\delta)\) is necessary for the spectral method to have non-vanishing overlap, these two observations point strongly in that direction. A third piece of supporting evidence is provided in [rk:pseig].

Equipped with 3, we can optimize both (i) the spectral threshold, namely, the minimum value of \(\delta\) needed to satisfy the condition \(\lambda^*(\delta_1) > \ol\lambda(\delta)\) which gives a strictly positive overlap, and (ii) the limiting overlap given by the right side of ?? . Formally, for \(i\in\{1,2\}\) and \(\delta\in(0,\infty)\), let \[\begin{align} \cH_i &\mathrel{\vcenter{:}}= \curbrkt{\cT\colon\bbR\to\bbR \text{ Lipschitz}s.t.\begin{array}{c} \displaystyle \inf_{y\in\supp(Y)} \cT(y) > -\infty , \; \displaystyle 0 < \sup_{y\in\supp(Y)} \cT(y) < \infty , \\ \prob{\cT(Y) = 0} < 1 , \; \lambda^*(\delta_i) > \ol\lambda(\delta) \end{array} } \label{eqn:def95Hi} \end{align}\tag{24}\] be the set of functions \(\cT\) satisfying [itm:assump-preproc-spec] such that \(\lambda^*(\delta_i)>\ol\lambda(\delta)\) holds. We recall that \(\delta_1 = \alpha\delta, \delta_2 = (1-\alpha)\delta\) and \(\lambda^*(\cdot), \ol\lambda(\cdot)\) depend on the choice of the preprocessing function. Noting that \(\cH_i\) depends on \(\delta\), we can define the spectral threshold for the \(i\)-th signal as \[\begin{align} \delta_i^{\spec} &\mathrel{\vcenter{:}}= \inf\curbrkt{\delta\in(0,\infty) : \cH_i \ne \emptyset} \, , \quad i\in\{1,2\}. \notag \end{align}\] In words, this is the smallest \(\delta\) such that there exists a preprocessing function satisfying \(\lambda^*(\delta_i) > \ol\lambda(\delta)\) (and, hence, leading to non-vanishing limiting overlap). Furthermore, for \(i\in\{1,2\}\) and \(\delta>\delta_i^{\spec}\), define the optimal overlap as \[\begin{align} \mathsf{OL}_i^{\spec} &\mathrel{\vcenter{:}}= \sup_{\cT\in\cH_i} \rho_i^{\spec}. \notag \end{align}\] In words, for a given \(\delta>\delta_i^{\spec}\), \(\mathsf{OL}_i^{\spec}\) is the largest overlap with preprocessing functions that satisfy \(\lambda^*(\delta_i) > \ol\lambda(\delta)\). We note that the supremum is guaranteed to be over a nonempty set as \(\delta>\delta_i^{\spec}\). At this point, we can state the following result whose proof is given in 13.

Consider the setting of 2, and let [itm:assump-signal-distr,itm:assump-alpha,itm:assump-noise-distr,itm:assump-gaussian-design,itm:assump-proportional] hold. Let \(\alpha_1:=\alpha\) and \(\alpha_2:=(1-\alpha)\). Then, for \(i \in \{1,2\}\) we have \[\begin{align} \delta_i^{\spec} &= \frac{1}{\alpha_i^2\int_{\supp(Y)} \frac{\expt{p(y|G)(G^2 - 1)}^2}{\expt{p(y|G)}} \diff y} \, , \label{eqn:spec-bound-1} \end{align}\tag{25}\] and for \(\delta>\delta_i^{\spec}\), \[\begin{align} \mathsf{OL}_i^{\spec} &= \frac{1}{\sqrt{\beta_i^*(\delta,\alpha) + \alpha_i}} \, , \label{eqn:opt-overlap-1} \end{align}\tag{26}\] where \(\beta_i^*(\delta,\alpha)\in(1-\alpha_i,\infty)\) are the unique solutions to the following pair of fixed point equations: \[\begin{align} (\beta_i^*(\delta,\alpha) - (1-\alpha_i)) \int_{\supp(Y)} \frac{\expt{p(y|G)(G^2-1)}^2}{\alpha_i \expt{p(y|G)G^2} + \beta_i^*(\delta,\alpha) \expt{p(y|G)}} \diff y &= \frac{1}{\alpha_i^2 \delta} , \quad i \in\{1, 2\}.\label{eqn:beta1-fp} \end{align}\tag{27}\] Finally, for \(i \in \{1,2\}\), define \(\cT_i^*\colon\bbR\to\bbR\) as \[\begin{align} \cT_i^*(y) &= 1 - \frac{1}{\alpha_i\cdot \frac{\expt{p(y|G)G^2}}{\expt{p(y|G)}} + (1-\alpha_i)} , \quad \text{ where } G\sim\cN(0,1). \label{eqn:opt-preprocessor} \end{align}\tag{28}\] Then, for \(\delta>\delta_i^{\spec}\), we have: (i) \(\cT_i^*\in\cH_i\), and (ii) the value of \(\mathsf{OL}_i^{\spec}\) is achieved by \(\cT_i^*\).

Finding the spectral and linear estimators that jointly maximize the overlap between the optimal combination of the two and the \(i\)-th signal (where \(i\in\{1,2\}\)) amounts to solving the following constrained optimization problem over a pair of functions \(\cT, \cL\): \[\begin{align} & \sup_{(\cT, \cL) \in \cH_i \times \cI} \mathsf{OL}_i^{\mathrm{comb}} . \label{eqn:opt} \end{align}\tag{29}\] In the above display, \(\cH_i\) (defined in 24 ) is the set of spectral preprocessing functions that satisfy [itm:assump-preproc-spec] and are effective for estimating \(x_i^*\) (i.e., \(\lambda^*(\delta_i) > \ol{\lambda}(\delta)\)); \(\cI\) (defined in 21 ) is the set of linear preprocessing functions that satisfy [itm:assump-preproc-lin] (and therefore are effective for estimating both signals); \(\mathsf{OL}_i^{\mathrm{comb}}\) (defined in 1) is the asymptotic overlap between the optimally combined estimator (with respect to fixed \(\cT, \cL\)) and \(x_i^*\). 29 is an explicit yet challenging functional optimization problem that remains open. Note that in the special cases of \(\alpha =0\) or \(\alpha=1\), 29 reduces to an analogous optimization problem for (non-mixed) GLMs whose resolution was left open in [25].

In 14, we show that the spectral thresholds \(\delta_1^{\spec}\) and \(\delta_2^{\spec}\) are always at least \(\delta_1^*\mathrel{\vcenter{:}}= \frac{1}{2\alpha^2}\) and \(\delta_2^*\mathrel{\vcenter{:}}= \frac{1}{2(1-\alpha)^2}\), for any conditional law \(p(\cdot \mid g)\) in 5 (i.e., regardless of the model). These lower bounds coincide with the spectral thresholds for both noiseless linear regression and noiseless phase retrieval, see [rk:lin-regr-phase-retrieval-coincide]. Thus, unlike the linear estimator, the spectral estimator (even the optimal one) does not achieve weak recovery for all \(\delta>0\); it gives positive overlaps only when the aspect ratio \(\delta\) exceeds a certain value. We highlight that the threshold associated to our proposed optimal spectral estimator is significantly lower than that corresponding to spectral estimators proposed earlier in the literature [13], [26], see 3.

a
b

Figure 3: Smallest \(\delta\) required by different spectral estimators to weakly recover signals for noiseless mixed phase retrieval. The spectral threshold is plotted as a function of a varying mixing parameter \(\alpha \in[0.6, 0.8]\). Our optimal spectral estimator always attains the lowest threshold. We note that these thresholds remain the same for noiseless mixed linear regression, due to the design of the corresponding estimators.. a — Recovery of \(x_1^*\), b — Recovery of \(x_2^*\)

The optimal limiting overlaps in 26 approach \(1\) as \(\delta\to\infty\) provided \[\begin{align} \int_{\supp(Y)} \frac{\expt{p(y|G)(G^2-1)}^2}{\alpha \expt{p(y|G)G^2} + (1-\alpha) \expt{p(y|G)}} \diff y& \, \in (0,\infty) . \label{eqn:overlap-approach-1-cond} \end{align}\tag{30}\] To show this, consider the optimal limiting overlap between the spectral estimator and the first signal, which by 26 equals \(\frac{1}{\sqrt{\beta_1^*(\delta,\alpha) + \alpha}}\). To show the claim, it suffices to show \(\beta_1^*(\delta,\alpha)\xrightarrow{\delta\to\infty}1-\alpha\). From 27 , the fixed point equation defining \(\beta_1^*(\infty,\alpha)\) becomes \[\begin{align} (\beta_1^*(\infty,\alpha) - (1-\alpha)) \int_{\supp(Y)} \frac{\expt{p(y|G)(G^2-1)}^2}{\alpha \expt{p(y|G)G^2} + \beta_1^*(\infty,\alpha) \expt{p(y|G)}} \diff y &= 0 , \label{eqn:beta1-fp-delta-infty} \end{align}\tag{31}\] as \(\delta\to\infty\). Since 30 holds, the unique solution to 31 has to be \(\beta_1^*(\infty,\alpha) = 1-\alpha\). This proves the claim. We note that the condition 30 is satisfied by the mixed linear regression model.

4 Numerical experiments↩︎

The experimental results in [fig:all,fig:noiseless-lin-regr,fig:lin-regr-vs-phase-retr] show that the performance of the various estimators (linear, spectral and combined) closely match the asymptotic predictions in various settings. Furthermore, 6 shows that our estimators exhibit improvements over existing spectral estimators designed for non-mixed data even when the signals have mild correlation. In all plots, the signal dimension is \(d=2000\), and the vertical and horizontal axes represent the overlap and the aspect ratio \(\delta\). The solid curves correspond to the theoretical predictions whose analytic expressions are in 7. Discrete points (little squares, triangles, asterisks, etc.) are computed using synthetic data. Each of these points is the mean of \(10\) i.i.d.trials together with error bars at \(1\) standard deviation. Additional comments on experimental setup and results are deferred to 8.

Figure 4: Spectral estimators for noiseless mixed linear regression, with mixing parameter \alpha \in \{ 0.6, 0.8\}. Optimal spectral estimators given by 50 are used. Overlaps with both signals x_1^*,x_2^*, computed from simulation (“sim.”) and prediction (“pred.”), are plotted as a function of the aspect ratio \delta. Same numerics apply to noiseless phase retrieval (see [rk:lin-regr-phase-retrieval-coincide]).
a
b

Figure 5: Spectral estimators for mixed linear regression and mixed phase retrieval. Optimal spectral estimators ([eqn:opt-prec-spec-lin-regr-main,eqn:opt-prec-spec-phase-retr-main]) are used. Overlaps with the first signal \(x_1^*\) (left plot) and with both signals \(x_1^*,x_2^*\) (right plot), computed from simulation (“sim.”) and prediction (“pred.”), are plotted as a function of the aspect ratio \(\delta\).. a — \(\alpha = 0.8\) and \(\sigma \in \{0.8,1.5\}\)., b — \(\alpha = 0.6\) and \(\sigma = 1.5\).

a
b

Figure 6: Performance comparison for correlated signals. The setting is noiseless mixed linear regression with mixing parameter \(\alpha = 0.6\) and signal correlation \(\inprod{x_1^*}{x_2^*} = \rho\) where \(\rho = 0.1\). Overlaps with \(x_1^*\) (left) and \(x_2^*\) (right) are plotted as a function of the aspect ratio \(\delta\). The signal dimension is \(d=2000\).. a — Recovery of \(x_1^*\), b — Recovery of \(x_2^*\)

5 Proof outline↩︎

The proof of 1 combines AMP with random matrix theory (RMT) tools. We now outline the high-level ideas in the analysis.

5.0.0.1 Eigenvalues via random matrix theory

The first step is to understand the spectrum of \(D\), in particular, the right edge of the bulk and the outlier(s). This involves the following challenges:

  • The matrix \(D\) in 8 can be thought of as an instance of spiked matrix model. Its structure is, however, more sophisticated than the canonical “signal plus noise” model. Indeed, the potential spikes of \(D\) result from two signals through the composition of the link function \(q\) and the spectral preprocessing function \(\cT\).

  • The analysis of the limiting spectrum for non-mixed GLMs is provided in [28], [29]. In our mixed setting, applying the strategy of [28], [29] to analyze the spectrum of \(D\) results in additional matrix terms which are hard to bound.

The key idea is to decompose \(D\) into the sum of two asymptotically free random matrices, consisting of the observations corresponding to the first and second signal. To be more specific, let us condition on \(\eta_1, \cdots, \eta_n\) and assume for notational convenience that \(\eta_i = 1\) for \(1\le i\le n_1\) and \(\eta_i = 0\) for \(n_1+1\le i\le n\), for some \(0\le n_1\le n\). Let \(n_2 = n - n_1\). Note that, almost surely, \(n_1/d\to\delta_1, n_2/d\to\delta_2\). Now, we can write the matrices of interest in block form: \[\begin{align} A &= \begin{bmatrix} A_1 \\ A_2 \end{bmatrix} , \quad T = \begin{bmatrix} T_1 & 0_{n_1\times n_2} \\ 0_{ n_2\times n_1} & T_2 \end{bmatrix} , \label{eqn:decomp-a-t} \end{align}\tag{32}\] where \(A_1\in\bbR^{n_1\times d}, A_2\in\bbR^{ n_2\times d}\) and \(T_1\in\bbR^{n_1\times n_1}, T_2\in\bbR^{ n_2\times n_2}\). We also let \(\veps_1 = (\eps_1,\cdots,\eps_{n_1})\) and \(\veps_2 = (\eps_{n_1+1},\cdots,\eps_n)\). Then, \[\begin{align} A^\top TA &= \begin{bmatrix} A_1^\top & A_2^\top \end{bmatrix} \begin{bmatrix} T_1 & 0_{n_1\times n_2} \\ 0_{ n_2\times n_1} & T_2 \end{bmatrix} \begin{bmatrix} A_1 \\ A_2 \end{bmatrix} = A_1^\top T_1 A_1 + A_2^\top T_2 A_2 . \label{eqn:decomp-d} \end{align}\tag{33}\] Note that, for \(i\in \{1,2\}\), \(A_i^\top T_i A_i = A_i^\top \mathop{\mathrm{diag}}(\cT(q(A_i x_i^*, \veps_i))) A_i\). Since \(A_1,x_1^*,\veps_1\) and \(A_2,x_2^*,\veps_2\) are mutually independent, \(A_1^\top T_1 A_1\) is independent of \(A_2^\top T_2 A_2\). However, \(A_1\) and \(T_1\) are not independent, neither are \(A_2\) and \(T_2\). When considered in isolation, \(A_1^\top T_1 A_1\) and \(A_2^\top T_2 A_2\) are obtained from a non-mixed GLM with aspect ratio discounted by \(\alpha\) and \(1-\alpha\), respectively. Thanks to [28], [29], their limiting spectra are well understood. Now, the crucial observation is that \(A_1^\top T_1 A_1\) and \(A_2^\top T_2 A_2\) are asymptotically free. Indeed, let \(O\sim \haar(\bbO(d))\) be a matrix sampled uniformly from the orthogonal group \(\bbO(d)\) and independent of everything else. Then, \[\begin{align} A_1^\top T_1 A_1 &+ A_2^\top T_2 A_2 = A_1^\top \mathop{\mathrm{diag}}(\cT(q(A_1 x_1^*, \veps_1))) A_1 + A_2^\top \mathop{\mathrm{diag}}(\cT(q(A_2 x_2^*, \veps_2))) A_2 \notag \\ &\eqqlaw A_1^\top \mathop{\mathrm{diag}}(\cT(q(A_1 x_1^*, \veps_1))) A_1 + (A_2O)^\top \mathop{\mathrm{diag}}(\cT(q((A_2O)x_2^*, \veps_2))) (A_2O) \tag{34} \\ &= A_1^\top \mathop{\mathrm{diag}}(\cT(q(A_1 x_1^*, \veps_1))) A_1 + O^\top A_2^\top \mathop{\mathrm{diag}}(\cT(q(A_2 (Ox_2^*), \veps_2))) A_2O \notag \\ &\eqqlaw A_1^\top \mathop{\mathrm{diag}}(\cT(q(A_1 x_1^*, \veps_1))) A_1 + O^\top A_2^\top \mathop{\mathrm{diag}}(\cT(q(A_2x_2^*, \veps_2))) A_2O \tag{35} \\ &= A_1^\top T_1 A_1 + O^\top A_2^\top T_2 A_2O . \tag{36} \end{align}\] 34 follows from the independence of \(A_1, A_2\), and from the rotational invariance of isotropic Gaussians. 35 follows since \(O\) and \(Ox_2^*\) are independent if \(O\sim \haar(\bbO(d))\) and \(x_2^*\sim\unif(\bbS^{d-1})\). In this step, we crucially use the assumption that \(x_1^*\) and \(x_2^*\) are independent and each uniformly distributed over \(\bbS^{d-1}\).

The asymptotic freeness shown in 36 allows us to study the (free) sum of \(A_1^\top T_1 A_1\) and \(A_2^\top T_2 A_2\) using the tools developed in [76]. Indeed, the analysis carried out in [sec:bulk,sec:outliers] implies the following characterization of the top three limiting eigenvalues of \(D\) (see 2): \[\begin{align} \lim_{d\to\infty} \lambda_1(D) &= \zeta(\lambda^*(\delta_1); \delta) , \; \lim_{d\to\infty} \lambda_2(D) = \zeta(\lambda^*(\delta_2); \delta) , \; \lim_{d\to\infty} \lambda_3(D) = \zeta(\ol\lambda(\delta); \delta) . \label{eqn:eigval-tmp} \end{align}\tag{37}\] Here, it is helpful to recall the definitions of \(\zeta(\cdot;\cdot)\) (see 12 ), \(\lambda^*(\cdot)\) (see page ) and \(\ol{\lambda}(\cdot)\) (see 11 ). Moreover, using the convexity of the function \(\zeta(\cdot; \delta)\), it can be shown that for \(i\in\{1,2\}\), \(\lambda_i(D)\) is strictly larger than \(\lambda_3(D)\) in the high-dimensional limit if \(\lambda^*(\delta_i) > \ol{\lambda}(\delta)\), meaning that \(\lambda_i(D)\) is detached from the bulk spectrum of \(D\) and becomes an outlier eigenvalue. Therefore, \(D\) exhibits a spectral gap between the \(i\)-th eigenvalue and the right edge of the bulk. In that case, the limiting eigenvalues admit the more explicit expressions reported in [rk:explicit-formula-eigval]. The existence of a spectral gap will be crucially used in proving the convergence of GAMP iterates to spectral estimators, as discussed below.

5.0.0.2 Joint distribution via GAMP

The convergence results in [eq:psiX1joint,eq:psiX2joint] are obtained using a generalized approximate message passing (GAMP) algorithm [27]. In a mixed GLM, since the observations \((y_i)_{i \in [n]}\) are unlabeled (i.e., it is unknown to the estimator whether each \(y_i\) is generated from the first or the second signal), estimating both signals is more challenging than estimating each one from an individual non-mixed GLM. However, the existing state evolution result for GAMP [27], [59] is derived for a non-mixed model, and only keeps track of the effect of a single signal. We generalize the GAMP state evolution result to mixed GLMs (see [prop:GAMP95SE]), so that the state evolution recursion tracks the effect of both signals. For convenience, let us work with the following rescalings: \[\begin{align} \Abar &\mathrel{\vcenter{:}}= \frac{1}{\sqrt{d}}\, A , \quad \xone \mathrel{\vcenter{:}}= \sqrt{d}\, x_1^* , \quad \xtwo \mathrel{\vcenter{:}}= \sqrt{d}\, x_2^* , \quad \Dbar \mathrel{\vcenter{:}}= \Abar^\top \, T\, \Abar = \frac{n}{d} A^\top TA . \label{eqn:rescaled-tmp} \end{align}\tag{38}\] Given two sequences of denoising functions \(f_{t+1}\colon\bbR^3\to\bbR, g_t\colon\bbR^2\to\bbR\) (for each iteration \(t\ge0\)), GAMP maintains a pair of iterates \(u^t\in\bbR^n, v^{t+1}\in\bbR^d\) according to \[\begin{align} \begin{aligned} u^t &= \frac{1}{\sqrt{\delta}} \Abar \tv^t - \sfb_t \tu^{t-1}, \quad \tu^t = g_{t}(u^{t};y) , \\ v^{t+1} &= \frac{1}{\sqrt{\delta}} \Abar^\top \tu^t - \sfc_t \tv^t, \quad \tv^{t+1}=f_{t+1}(v^{t+1}; \, \xone, \xtwo) , \end{aligned} \label{eq:gamp-eqn-general-tmp} \end{align}\tag{39}\] where \(f_{t+1},g_t\) are applied component-wise, i.e., \(f_{t+1}(v^{t+1}; \, \xone, \xtwo)=(f_{t+1}(v^{t+1}_1; \, \ol{x}_{1,1}^*, \ol{x}_{2,1}^*)\), \(\ldots, f_{t+1}(v^{t+1}_d; \, \ol{x}_{1,d}^*, \ol{x}_{2,d}^*)), g_t(u^t; y)=(g_t(u^t_1; y_1), \ldots, g_t(u^t_n; y_n))\). The scalars \(\sfb_t, \sfc_t\) are defined as \[\sfb_t =\frac{1}{n}\sum_{i=1}^d f_t'(v_i^t; \, \ol{x}_{1,i}^*, \ol{x}_{2,i}^*), \qquad \sfc_t = \frac{1}{n}\sum_{i=1}^n g_t'(u_i^t; y_i), \notag\] where \(f_t'\) and \(g_t'\) each denote the derivative with respect to the first argument. The iteration is initialized with a given \(\tv^0 \in \bbR^d\) and \(\tu^{-1}=0_n\). Under the assumption that the design matrix is Gaussian (indeed \(\Abar _{i,j} \iid \cN(0,1/d)\) according to [itm:assump-gaussian-design]), the joint empirical distribution of \(u^t,v^{t+1}\) converges (as \(n,d \to \infty\) with \(n/d \to \delta\)) to the law of a pair of jointly Gaussian random variables \(U_t,V_{t+1}\): \[U_t \mathrel{\vcenter{:}}= \mu_{1,t} G_1 + \mu_{2,t} G_2 + W_{U,t} , \quad V_{t+1} \mathrel{\vcenter{:}}= \chi_{1,t+1} X_1 + \chi_{2,t+1} X_2 + W_{V,t+1} , \notag\] where \((G_1, G_2, W_{U,t}) \sim \normal(0,1) \otimes \normal(0,1) \otimes \normal(0, \sigma_{U,t}^2)\), and \((X_1, X_2, W_{V,t+1}) \sim \normal(0,1) \otimes \normal(0,1) \otimes \normal(0,\sigma_{V,t+1}^2)\). The covariance structure of these jointly Gaussian random variables is described by a set of recursions called state evolution: \[\begin{gather} \mu_{1,t} = \frac{1}{\sqrt{\delta}} \E [ X_1 f_t(V_t; \, X_1, X_2) ], \quad \mu_{2,t} = \frac{1}{\sqrt{\delta}} \E[ X_2 f_t(V_t; \, X_1, X_2) ], \nonumber \\ \sigma_{U,t}^2 = \frac{1}{\delta}\E[ f_t(V_t; \, X_1, X_2)^2 ] - \mu_{1,t}^2 - \mu_{2,t}^2 \, , \nonumber \\ \chi_{1,t+1} = \sqrt{\delta} \left( \E[G_1 g_t(U_t; \tY) ] - \E[ g_t'(U_t; \tY)] \mu_{1,t} \right), \nonumber \\ \chi_{2,t+1} = \sqrt{\delta} \left( \E [ G_2 g_t(U_t; \tY) ] - \E [g_t'(U_t; \tY)] \mu_{2,t} \right), \nonumber \\ \qquad \sigma_{V,t+1}^2 = \E[g_t(U_t; \tY)^2 ] , \nonumber \end{gather}\] where the random variable \(\tY\) is given by \[\tY = q(\eta G_1 + (1-\eta)G_2, \, \eps), \text{ with } (G_1, G_2, \eta, \eps) \sim \normal(0,1) \ot \normal(0,1) \ot \text{Bern}(\alpha) \ot P_{\eps}. \notag\] The recursion is initialized as \[\begin{align} &\mu_{1,0} =\frac{1}{\sqrt{\delta}} \lim_{d \to \infty} \frac{\langle \xone \, , \, \tv^0 \rangle}{d} , \;\; \mu_{2,0} =\frac{1}{\sqrt{\delta}} \lim_{d \to \infty} \frac{\langle \xtwo \, , \, \tv^0 \rangle}{d} , \;\; \sigma_{U,0}^2= \frac{1}{\delta} \lim_{d \to \infty} \frac{\normtwo{\tv^0}^2}{d} - \mu_{1,0}^2 - \mu_{2,0}^2 . \notag \end{align}\] The proof of convergence of the empirical distributions of \(u^t,v^{t+1}\) to the laws of \(U_t,V_{t+1}\), given in 15, uses a reduction to an abstract AMP recursion with matrix-valued iterates for which a state evolution result was established in [59], [77]. For details, see the formal statements in [prop:GAMP95SE] which track the joint empirical distribution of all iterates.

At this point, the linear estimator is readily obtained via the iterate of GAMP run for one step (\(t=0\)). For \(t\ge1\), we tailor the denoisers \((f_{t+1}, g_t)_{t\ge1}\) so that the iterates of GAMP implement a power method, which for large enough \(t\), gives the first and second eigenvector of the spectral matrix \(D\) (defined in 8 ). Specifically, consider the GAMP iteration in 39 with the initializer \(\tv^0 = 0_d\), and the following choice of denoisers: \[\begin{align} & g_0(u^0; y) = \sqrt{\delta} \cL(y), \qquad f_1(v; \, \xone, \xtwo)= f(\xone, \xtwo), \\ & g_t(u; \, y) = \sqrt{\delta} \, u \, \cF(y), \quad f_{t+1}(v; \, \xone, \xtwo) = \frac{v}{\beta_{t+1}}, \quad t \ge 1, \end{align} \label{eq:ft95gt95choice-tmp}\tag{40}\] where \(\cF: \reals \to \reals\) is bounded and Lipschitz, \(f: \reals^2 \to \reals\) is Lipschitz, and \(\beta_{t+1} \mathrel{\vcenter{:}}= \sqrt{\chi_{1,t+1}^2 + \chi_{2,t+1}^2 + \sigma_{V,t+1}^2}\). To prove 1, we select two pairs of functions \((f, \cF)\), in terms of the spectral preprocessing function \(\cT\) (see [eq:F1x1-tmp,eq:F2x2-tmp]). With the above choice of \((f_{t+1})_{t\ge1}, (g_t)_{t\ge1}\), the GAMP iteration becomes \[\begin{align} & u^0 =0_n, \quad v^1= \Abar^\top \cL(y), \\ & u^{t} = \frac{1}{\sqrt{\delta} \, \beta_{t}} \paren{\Abar v^{t} \, - \, F u^{t-1}}, \quad v^{t+1} = \Abar^\top F u^t - \frac{\sqrt{\delta}}{\beta_t} \, \E[ \cF(\tY) ] \, v^t, \qquad t \ge 2, \end{align} \label{eq:newGAMP-tmp}\tag{41}\] where \(F = \mathop{\mathrm{diag}}(\cF(y_1), \ldots, \cF(y_n))\). First, note that the iterate \(v^1\) coincides with the linear estimator \(\xlin\) in 6 . Furthermore, we show that in the high-dimensional limit, as \(t\to\infty\), the iterate \(v^t\) is aligned with an eigenvector of the matrix \[\label{eq:newmat} \Abar^\top F(\sqrt{\delta} \beta_{\infty}I_n + F)^{-1} \Abar,\tag{42}\] where \(\beta_{\infty} = \lim\limits_{t \to \infty} \beta_t\). To justify the claim, assume the iterates \(u^t, v^{t+1}\) converge to the limits \(u^\infty, v^\infty\) in the sense that \(\lim\limits_{t \to \infty} \lim\limits_{d \to \infty} \frac{1}{d} \normtwo{u^t - u^{\infty}}^2 =0\) and \(\lim\limits_{t \to \infty} \lim\limits_{d \to \infty} \frac{1}{d} \normtwo{v^t - v^{\infty}}^2 =0\). Then, from 41 we can derive \[v^{\infty}\left( 1 + \frac{\sqrt{\delta}}{\beta_\infty} \E[ \cF(\tY)] \right) = \Abar^{\top}F (\sqrt{\delta}\beta_{\infty}I_n + F)^{-1}\Abar v^{\infty}. \label{eq:evector95eqn-tmp}\tag{43}\] Therefore, \(v^{\infty}\) is an eigenvector of the matrix in 42 , and the GAMP iteration of 41 is effectively a power method.

Recall that our goal is to obtain via GAMP the two leading eigenvectors of \(\Abar^{\top}T\Abar\). Hence, we pick \(\cF\) so that \(F (\sqrt{\delta}\beta_{\infty}I_n + F)^{-1} = c\, T\), for some constant \(c\). To this end, we analyze the iteration in 41 with two choices for the function \(\cF(y)\) and initialization \(\tv^0\). \[\begin{align} &\text{Choice 1}: \quad \cF_1(y) \mathrel{\vcenter{:}}= \frac{\cT(y)}{\lambda^*(\delta_1) - \cT(y)}, \quad f(\xone, \xtwo)=\xone, \tag{44} \\ &\text{Choice 2}: \quad \cF_2(y) \mathrel{\vcenter{:}}= \frac{\cT(y)}{\lambda^*(\delta_2) - \cT(y)}, \quad f(\xone, \xtwo)=\xtwo, \tag{45} \end{align}\] where for \(i\in\{1, 2\}\), \(\lambda^*(\delta_i)\) is the unique solution of \(\zeta(\lambda;\delta_i) = \phi(\lambda)\) (see page ). The above two choices are motivated by the characterization of the limiting eigenvalues in 37 and [rk:explicit-formula-eigval]. As outlined below, choice 1 (resp.choice 2) ensures that 43 becomes an eigen-equation for the first (resp.second) eigenvalue of \(\Dbar\).

3 shows that the state evolution parameters \((\chi_{1,t},\chi_{2,t},\sigma_{V,t}^2)\) for choice 1 satisfy \[\begin{align} \lim_{t\to\infty} \chi_{1,t} &= \frac{\rho_1^{\spec}}{\sqrt{\delta}} , \quad \lim_{t\to\infty} \sigma_{V,t}^2 = \frac{1- (\rho_1^{\spec})^2}{\delta} , \quad \text{ and } \; \chi_{2,t} = 0 , \;\forall t\ge2, \notag \end{align}\] where the quantity \(\rho_1^{\spec}\) was defined in 14 . Hence, \[\begin{align} \beta_\infty &= \lim_{t\to\infty} \sqrt{\chi_{1,t}^2 + \chi_{2,t}^2 + \sigma_{V,t}^2} = \frac{1}{\sqrt{\delta}} , \notag \end{align}\] and 43 becomes: \[v^{\infty}\left( 1 + \delta \E\left[ \frac{\cT(Y)}{\lambda^*(\delta_1) - \cT(Y)} \right] \right) = \frac{1}{\lambda^*(\delta_1)} \Abar^{\top} T \Abar v^\infty . \label{eq:evector95eqn951-tmp}\tag{46}\] With choice 1, 46 gives that the GAMP iterate converges to an eigenvector of \(\Dbar = \Abar^\top T \Abar\) corresponding to the eigenvalue \(\lambda^*(\delta_1)\left( 1 + \delta \E\left[ \frac{\cT(Y)}{\lambda^*(\delta_1) - \cT(Y)} \right] \right)\). Similarly, with choice 2, the GAMP iterate converges to an eigenvector of \(\Dbar\) corresponding to the eigenvalue \(\lambda^*(\delta_2)\left( 1 + \delta \E\left[ \frac{\cT(Y)}{\lambda^*(\delta_2) - \cT(Y)} \right] \right)\). These claims match the rigorous eigenvalue characterization in 37 and [rk:explicit-formula-eigval]. At this point, note that power methods (and therefore our GAMP iterations in 41 ) crucially require a spectral gap to converge to the desired eigenvector. This spectral gap is guaranteed precisely by 37 provided \(\lambda^*(\delta_1) > \ol{\lambda}(\delta)\) (resp.\(\lambda^*(\delta_2) > \ol{\lambda}(\delta)\)), which gives that \(\lambda_1(\Dbar)\) (resp.\(\lambda_2(\Dbar)\)) is asymptotically an outlier in the spectrum of \(\Dbar\). As a consequence, we can rigorously prove the convergence of the GAMP iterates under choice 1 (resp.choice 2) to \(v_1(\Dbar)\) (resp.\(v_2(\Dbar)\)). To conclude, the iterate \(v^1\) in the GAMP iteration in 41 equals the linear estimator, and \(v^{t+1}\) asymptotically aligns with the spectral estimator. Since the state evolution tracks the limiting joint distribution of all iterates, the characterization in [eq:psiX1joint,eq:psiX2joint] follows. We stress that GAMP in our argument is used only as a tool for analysis and is not part of the estimators. The actual estimators (spectral and linear) can be computed by a combination of the following simple operations: (i) applying a component-wise nonlinearity, (ii) matrix-vector/-matrix multiplication, (iii) computation of eigenvectors.

5.0.0.3 Optimal linear and spectral estimators

The master theorem (1) holds for arbitrary linear and spectral preprocessing functions \(\cL,\cT\) satisfying the stated assumptions. Specializing 1 to linear and spectral estimators alone and using the explicit formulas for their limiting overlaps (given in [lem:linear-overlap,thm:overlap-spectral]), we find the optimal preprocessing functions \(\cL^*,\cT_1^*,\cT_2^*\) that maximize the limiting overlaps. This is done in [lem:linear-optimal-overlap,thm:opt-spec] by casting the optimization problem as a variational problem and solving it explicitly.

6 Discussion↩︎

6.0.0.1 Universality beyond the Gaussian design matrix

A natural question is whether the predictions obtained under an i.i.d.Gaussian design are valid more generally. This topic has been investigated in the random matrix theory literature [78][82], and a recent line of research has focused on AMP [83][88]. We note that none of these results is directly applicable to our setting, and the problem also remains open in the non-mixed (i.e., \(\alpha=1\)) setup. However, the aforementioned body of work suggests that the Gaussian predictions may hold for much more general – even “almost deterministic” – design matrices.

6.0.0.2 Mixed GLM with multiple components

We focus on the mixed GLM with two components, but our approach is well suited to handle mixed GLMs with multiple components. We now briefly sketch how to generalize our main 1. The other results (overlaps for linear, spectral and combined estimators, and their optimization) are generalized in a similar fashion.

Let \(x_1^*,\cdots,x_\ell^*\in\bbR^d\) be \(\ell\) signal vectors, and let the observation \(y = (y_1,\cdots,y_n)\in\bbR^n\) be generated as \(y_i = q\paren{\inprod{a_i}{x_{\upsilon_i}^*}, \eps_i}\). The latent vector \(\underline{\upsilon} = (\upsilon_1,\cdots,\upsilon_n)\) is a sequence of i.i.d.mixing random variables s.t. \(\prob{\upsilon_i = j} = \alpha_j\) for all \(i\in [n]\) and \(j\in [\ell]\). We assume that \(x_1^*,\cdots,x_\ell^*\) are i.i.d.and uniform on the unit sphere, and \(1>\alpha_1>\alpha_2>\cdots>\alpha_\ell>0\) (corresponding to [itm:assump-signal-distr,itm:assump-alpha]). We also impose our previous [itm:assump-noise-distr,itm:assump-gaussian-design,itm:assump-proportional], and assume that \(\ell\) is a constant (independent of \(n,d\)). For \(i\in[\ell]\), let \(\delta_i \mathrel{\vcenter{:}}= \alpha_i\delta\) and \[\begin{align} n^{\lin} &\mathrel{\vcenter{:}}= \paren{\paren{\sum_{k = 1}^\ell \alpha_k^2} \expt{G\cL(Y)}^2 + \frac{\expt{\cL(Y)^2}}{\delta}}^{1/2} , \notag \\ \rho_i^{\lin} &\mathrel{\vcenter{:}}= \frac{\alpha_i \expt{G\cL(Y)}}{n^{\lin}} , \qquad \rho_i^{\spec} \mathrel{\vcenter{:}}= \paren{\frac{\frac{1}{\delta} - \expt{\paren{\frac{Z}{\lambda^*(\delta_i) - Z}}^2}}{\frac{1}{\delta} + \alpha_i \expt{\paren{\frac{Z}{\lambda^*(\delta_i) - Z}}^2 (G^2 - 1)}}}^{1/2} . \notag \end{align}\] Here \(Z = \cT(Y)\), \(Y= q(G, \eps)\), and \(G \sim \cN(0,1)\), as before. Then, under the same setting of 1 with \(\ol{x}_i^*\) and \(x_i^{\spec}\) defined similarly for \(i\in [\ell]\), we have that, if \(\lambda^*(\delta_i) > \ol\lambda(\delta)\), \[\begin{align} \lim_{d\to\infty} \frac{1}{d} \sum_{j = 1}^d \Psi(\ol{x}_{i,j}^*, x_j^{\lin}, x_{i,j}^{\spec}) &= \expt{\Psi\paren{X_i, \, \sum_{k = 1}^\ell \rho_k^{\lin} X_k + W^{\lin}, \, \rho_i^{\spec} X_i + W_i^{\spec}}} , \notag \end{align}\] where \((X_1,\cdots,X_\ell) \sim \cN(0,1)^{\ot\ell}\), \((W^{\lin}, W_i^{\spec})\) is independent of \((X_1,\cdots,X_\ell)\) and is jointly Gaussian with zero mean and covariance given by \[\begin{align} \expt{(W^{\lin})^2} &= 1 - \sum_{k = 1}^\ell (\rho_k^{\lin})^2 , \; \expt{(W_i^{\spec})^2} = 1 - (\rho_i^{\spec})^2 , \notag \\ \expt{W^{\lin} W_i^{\spec}} &= \frac{\alpha_i \rho_i^{\spec}}{n^{\lin}} \expt{\frac{G\cL(Y)Z}{\lambda^*(\delta_i) - Z}} . \notag \end{align}\]

The result on the eigenvalues of the spectral matrix \(D\) can be obtained by following the strategy detailed in 9 (and sketched in 5). Indeed, \(D\) can be decomposed into the (asymptotically) free sum of the \(\ell\) components associated to each of the signals, and [76] is well equipped to characterize its top \(\ell+1\) eigenvalues. To derive the limiting joint empirical law of the \(i\)-th signal and the linear and spectral estimators, we can then run a GAMP algorithm similar to 40 with denoisers tailored for the \(i\)-th signal. The condition \(\lambda^*(\delta_i) > \ol\lambda(\delta)\) guarantees the existence of a spectral gap between the \(i\)-th largest eigenvalue of \(D\) and the rest of its spectrum, which in turn is leveraged to argue the convergence of GAMP to the desired eigenvector. This yields results analogous to 1 with \(\alpha\) underlying ?? therein replaced with \(\alpha_i\), for \(1\le i \le \ell\).

6.0.0.3 Lower bounds for inference in mixed GLMs

In the non-mixed setting, [29] derives an information-theoretic threshold \(\delta^{\infthr}\) such that, for \(\delta < \delta^{\infthr}\), no estimation method gives a non-trivial estimate of the signal.6 Furthermore, for noiseless phase retrieval, \(\delta^{\infthr}=1/2\), which matches the threshold achieved by a spectral method. In the mixed setting, denoting by \(\delta^{\infthr}_1,\delta^{\infthr}_2\) the information-theoretic thresholds corresponding to the two signals, we have \[\begin{align} \delta^{\infthr}_1 \ge \frac{1}{\alpha} \delta^{\infthr}, \qquad \delta^{\infthr}_2 \ge \frac{1}{1-\alpha} \delta^{\infthr} . \label{eqn:it-thr-trivial} \end{align}\tag{47}\] To see this, note that, if a genie reveals the values of the mixing variables \((\eta_1,\cdots,\eta_n)\), then the estimation problem given mixed data with aspect ratio \(\delta\) can be decoupled into two non-mixed ones with aspect ratios \(\alpha\delta\) and \((1-\alpha)\delta\). We also remark that adapting the second moment method of [29] to our mixed setting does not improve the bound in 47 (hence, this derivation is omitted). Following the strategy of [89] – which establishes the exact asymptotics of the minimum mean squared error and, thus, gives a tight bound in the non-mixed case – requires additional ideas beyond the scope of this paper, so it is left for future research.

As a final remark, let us contrast 47 with the spectral bounds mentioned in [rk:univ-lb-spec-thr], which are universal in the sense that they hold for any mixed GLM. In particular, we note that the former scales as \((1/\alpha, 1/(1-\alpha))\), while the latter as \((1/\alpha^2, 1/(1-\alpha)^2)\), which suggests a gap between what is achievable information-theoretically and algorithmically. The possibility of a statistical-computational trade-off is also suggested by the fact that, for \(\alpha=1/2\) and antipodal signals (\(x_1^*=-x_2^*\)), mixed linear regression reduces to phase retrieval, which is widely believed to have such a gap, see e.g.[39], [90][92]. Closing the gap or understanding its fundamental nature remains an intriguing open question for future investigation.

6.0.0.4 Organization of the supplementary material

The supplementary material is organized as follows. Two illustrative examples of our main results in 3 are given in 7. Discussions on numerical simulations in 4 are provided in 8. The proof of the master theorem (1) is divided across two sections. 9 contains a characterization of the top three eigenvalues of the matrix \(D\) which is used in the analysis of GAMP in the following section. The limiting joint law of the signal, the linear and the spectral estimators in 1 is then proved in 10 using a GAMP algorithm and its characterization via state evolution. The proof of the state evolution characterization is deferred to 15. [app:bayes-opt-comb,sec:linear-estimator-pf,sec:spec-estimator-pf,sec:lower-bound-spec-thr] contain the proofs of various consequences of the master theorem. Several auxiliary lemmas are in 16.

7 Two illustrative examples↩︎

We specialize the results in [sec:results-spec,sec:results-lin] to two prototypical examples of mixed GLMs: the mixed linear regression model where \[\begin{align} q(g,\eps) &= g + \eps , \quad \veps\sim \cN(0,\sigma^2 I_n) , \label{eqn:noisy-lin-regr-model} \end{align}\tag{48}\] and the mixed phase retrieval model where \[\begin{align} q(g,\eps) &= |g| + \eps , \quad \veps \sim\cN(0,\sigma^2 I_n) . \label{eqn:noisy-phase-retrieval-model} \end{align}\tag{49}\] The explicit formulas for the optimal preprocessing functions, the optimal overlaps, and the thresholds (for spectral estimators) are collected in the following corollaries. Throughout this section, for brevity we write \(\alpha_1=\alpha\) and \(\alpha_2=(1-\alpha)\). Let us first consider linear estimators.

Corollary 4 (Mixed linear regression, linear estimator). Consider the mixed linear regression model in 48 , and let [itm:assump-signal-distr,itm:assump-alpha,itm:assump-gaussian-design,itm:assump-proportional] hold. Then, the optimal preprocessing function \(\cL^*\) defined in [lem:linear-optimal-overlap] is given by \[\begin{align} \cL^*(y) &= \frac{y}{1+\sigma^2}. \label{eqn:opt-lin-for-lin-regr-main} \end{align}\qquad{(6)}\] Recalling that \(\xlin \mathrel{\vcenter{:}}= \frac{1}{n} A^\top \cL^*(y)\), we almost surely have: \[\begin{align} \lim_{d \to \infty} \, \frac{\inprod{\xlin}{x_i^*}}{\normtwo{\xlin}\normtwo{x_1^*}} &= \paren{\frac{\alpha_1^2 + \alpha_2^2}{\alpha_i^2} + \frac{1+\sigma^2}{\alpha_i^2\delta}}^{-1} , \qquad i \in \{1,2\}. \notag \end{align}\]

For mixed phase retrieval, one can readily check that the overlap of the linear estimator with each signal is always vanishing regardless of the choice of the preprocessing function. Next, we consider spectral estimators.

Corollary 5 (Mixed linear regression, spectral estimator). Consider the mixed linear regression model and let [itm:assump-signal-distr,itm:assump-alpha,itm:assump-gaussian-design,itm:assump-proportional] hold. Then, for \(i \in \{1,2 \}\), the optimal preprocessing function \(\cT_i^*\) defined in [thm:opt-spec] is: \[\begin{align} \cT_i^*(y) &= 1 - \frac{1}{\alpha_i \cdot \frac{y^2 + \sigma^2 + \sigma^4}{(1+\sigma^2)^2} + (1-\alpha_i)}. \label{eqn:opt-prec-spec-lin-regr-main} \end{align}\qquad{(7)}\] Let \(T_i^* = \mathop{\mathrm{diag}}(\cT_i^*(y))\) and \(D_i^* = \frac{1}{n}A^\top T_i^* A\), for \(i \in \{1,2\}\). Denote by \(v_1(D_i^*),v_2(D_i^*)\) the eigenvectors of \(D_i^*\) corresponding to the two largest eigenvalues. Then for \(\delta > \frac{(1+\sigma^2)^2}{2\alpha_i^2}\), we almost surely have \[\begin{align} \lim_{d\to\infty} \frac{\abs{\inprod{v_i(D_i^*)}{x_i^*}}}{\normtwo{v_i(D_i^*)}\normtwo{x_i^*}} &= \frac{1}{\sqrt{\beta_i^*(\delta,\alpha,\sigma) + \alpha_i}} , \notag \end{align}\] where \(\beta_i^*(\delta,\alpha,\sigma)\) is the unique solution in \((1-\alpha_i,\infty)\) to the following fixed point equation: \[\begin{gather} (\beta_i^*(\delta,\alpha,\sigma) - (1-\alpha_i)) \left[ -\frac{\alpha_i + \beta_i^*(\delta,\alpha,\sigma)}{\alpha_i^2} + \paren{\frac{\alpha_i+\beta_i^*(\delta,\alpha,\sigma)}{\alpha_i}}^2 \right.\\ \left. \times \sqrt{\frac{ \pi(1+\sigma^2)^2}{2\alpha_i(\sigma^2\alpha_i +(1+\sigma^2)\beta_i^*(\delta,\alpha,\sigma))}} \right. \left. \exp\paren{\frac{\sigma^2\alpha_i + (1+\sigma^2)\beta_i^*(\delta,\alpha,\sigma)}{2\alpha_i}} \right. \\ \left. \times \erfc\paren{\sqrt{\frac{\sigma^2\alpha_i + (1+\sigma^2)\beta_i^*(\delta,\alpha,\sigma)}{2\alpha_i}}} \right] = \frac{1}{\alpha_i^2\delta} . \notag \end{gather}\]

Corollary 6 (Mixed phase retrieval, spectral). Consider the mixed phase retrieval model in 49 , and let [itm:assump-signal-distr,itm:assump-alpha,itm:assump-gaussian-design,itm:assump-proportional] hold. Then, for \(i \in \{1,2 \}\), the optimal preprocessing function \(\cT_i^*\) defined in [thm:opt-spec] is: \[\begin{align} \cT_i^*(y) &= 1 - \frac{1}{\alpha_i \Delta(y) + (1-\alpha_i)} , \label{eqn:opt-prec-spec-phase-retr-main} \end{align}\qquad{(8)}\] where the auxiliary function \(\Delta\colon\bbR\to\bbR\) is defined as \[\begin{align} \Delta(y) &\mathrel{\vcenter{:}}= \frac{y^2 + \sigma^2 + \sigma^4}{(1+\sigma^2)^2} + \sqrt{\frac{2}{\pi}} \cdot \frac{\sigma y \exp\paren{-\frac{y^2}{2\sigma^2(1+\sigma^2)}}}{(1+\sigma^2)^{3/2}}\sqrbrkt{1 + \erf\paren{\frac{y}{\sqrt{2\sigma^2(1+\sigma^2)}}}}^{-1} . \notag \end{align}\] Let \(T_i^* = \mathop{\mathrm{diag}}(\cT_i^*(y))\in\bbR^{n\times n}\) and \(D_i^* = \frac{1}{n}A^\top T_i^* A\in\bbR^{d\times d}\), for \(i\in\{1,2\}\). Denote by \(v_1(D_i^*),v_2(D_i^*)\) the eigenvectors of \(D_i^*\) corresponding to the two largest eigenvalues Let \[\begin{align} \delta^*_i &= \frac{1}{\alpha_i^2} \paren{ \frac{2}{(1+\sigma^2)^2} + \frac{4\sigma^5 h(\sigma^2)}{\pi^{3/2}(1+\sigma^2)^2} }^{-1} , \quad \text{ where } h(\sigma^2) \mathrel{\vcenter{:}}= \int_\bbR \frac{\exp\paren{-(2+\sigma^2)z^2} z^2}{1+\erf(z)} \diff z. \end{align}\] Define the functions \(m_0,m_1\colon\bbR\to\bbR\) and \(I\colon[1/2,1]\times(0,\infty)\to\bbR\) as \[\begin{align} m_0(y) &\mathrel{\vcenter{:}}= \frac{1}{\sqrt{2\pi(1+\sigma^2)}} \exp\paren{-\frac{y^2}{2(1+\sigma^2)}} \sqrbrkt{1 + \erf\paren{\frac{y}{\sqrt{2\sigma^2(1+\sigma^2)}}}} , \notag \\ m_1(y) &\mathrel{\vcenter{:}}= m_0(y) \frac{y^2 + \sigma^2 + \sigma^4}{(1+\sigma^2)^2} + \frac{\sigma y}{\pi(1+\sigma^2)^2} \exp\paren{-\frac{y^2}{2\sigma^2}} , \notag \\ I(\alpha,\beta) &\mathrel{\vcenter{:}}= \int_{\supp(Y)} \frac{(m_2(y) - m_0(y))^2}{\alpha m_2(y) + \beta m_0(y)} \diff y . \notag \end{align}\] Then for \(i\in\{1,2\}\) if \(\delta > \delta^*_i\), we almost surely have: \[\begin{align} \lim_{d\to\infty} \frac{\abs{\inprod{v_i(D_i^*)}{x_i^*}}}{\normtwo{v_i(D_i^*)}\normtwo{x_i^*}} &= \frac{1}{\sqrt{\beta_i^*(\delta,\alpha,\sigma) + \alpha_i}} \, , \notag \end{align}\] where \(\beta_i^*(\delta,\alpha,\sigma)\) is the unique solution in \((1-\alpha_i,\infty)\) to the fixed point equation: \[\begin{align} (\beta_i^*(\delta,\alpha,\sigma) - (1-\alpha_i)) I(\alpha_i, \beta_i^*(\delta,\alpha,\sigma)) &= \frac{1}{\alpha_i^2\delta} . \notag \end{align}\]

We note that the performance of the optimal spectral estimators given in [cor:noisy-mixed-linear-regr-spec,cor:noisy-mixed-phase-retrieval-spec] coincides for mixed noiseless linear regression and mixed noiseless phase retrieval. Specifically, for both models, when \(\sigma^2=0\), the spectral thresholds and the optimal preprocessing functions are: \[\begin{align} \delta_i^* &= \frac{1}{2\alpha_i^2} , \quad \cT_i^*(y) = 1-\frac{1}{\alpha_i y^2 + (1-\alpha_i)} \, , \label{eqn:rk-spec-preproc-same} \end{align}\tag{50}\] and the corresponding overlaps are \(\frac{1}{\sqrt{\beta_i^*(\delta,\alpha,0) + \alpha_i}}\), for \(i \in \{1,2\}\). Here \(\beta_i^*(\delta,\alpha,0)\) is the solution to the fixed point equation in 5 with \(\sigma=0\). In fact, one can verify that even the first-order dependence of the spectral thresholds on the noise variance \(\sigma\) coincides for the noisy versions of these two problems: \(\delta^*_i = \frac{1 + 2\sigma^2}{2\alpha_i^2} + \cO(\sigma^4)\) for \(i \in \{1,2\}\). This phenomenon is because the optimal preprocessing functions \(\cT_i^*(y)\) in both models depend only on \(y^2\), and are therefore invariant to the signs of the observations \((y_1, \ldots, y_n)\).

8 Discussion on numerical experiments↩︎

We make a few remarks on the numerical results in 4.

  • 1 shows numerical results for the recovery of the first and second signal, respectively, from a noiseless linear regression model (i.e., the model in 48 with \(\sigma = 0\)) with mixing parameter \(\alpha = 0.6\). We plot overlaps obtained via (i) the optimal spectral estimator in 50 , (ii) the optimal linear estimator in ?? , (iii) the Bayes-optimal linear combination of the estimators in (i) and (ii) (as per 1), (iv) the spectral estimator proposed in [13] whose preprocessing function \(\cT^{\mathrm{YCS}}\) is: \[\begin{align} \cT^{\mathrm{YCS}}(y) = \min\curbrkt{y^2, 10} , \label{eqn:ycs-truncated} \end{align}\tag{51}\] and (v) the spectral estimator proposed in [26] whose preprocessing function \(\cT^{\mathrm{LAL}}\) is: \[\begin{align} \cT^{\mathrm{LAL}}(y) = \max\curbrkt{1 - \frac{1}{y^2}, -10} . \label{eqn:lal-truncated} \end{align}\tag{52}\] The estimators in [eqn:ycs-truncated,eqn:lal-truncated] are truncated at \(+10\) and \(-10\), respectively, in order to compute our theoretical predictions. Choosing a larger value in magnitude for the truncation does not lead to improved empirical performance. This is because these choices are not optimal for estimation from mixed models. Our combined estimator and our optimal design of the spectral method yield substantially larger overlaps compared to existing heuristic choices, such as those in [eqn:ycs-truncated,eqn:lal-truncated].

  • In 4, we consider the recovery of both signals for noiseless mixed linear regression (link function given by 48 with \(\sigma=0\)), using the spectral estimator with optimal preprocessing functions given by 50 . Overlaps are plotted for two values of the mixing parameter \(\alpha\in\{0.6, 0.8\}\). The results for noiseless phase retrieval are identical (for both simulations with \(d=2000\) and the asymptotic prediction), as noted in [rk:lin-regr-phase-retrieval-coincide].

  • In 5 (a), we compare mixed linear regression and mixed phase retrieval ([eqn:noisy-lin-regr-model,eqn:noisy-phase-retrieval-model]), under their respective optimal spectral estimators ([eqn:opt-prec-spec-lin-regr-main,eqn:opt-prec-spec-phase-retr-main]). For each model, we plot the overlap with the first signal for two different values of the noise standard deviation \(\sigma \in \{0.8, 1.5\}\). In all cases, the mixing parameter is fixed to be \(\alpha = 0.8\). Though for \(\sigma = 0\) the curves for both models coincide, the gap between phase retrieval and linear regression grows with \(\sigma\), with increasingly better performance for phase retrieval.

  • In 5 (b), we consider the recovery of both signals for mixed linear regression and mixed phase retrieval with mixing parameter \(\alpha = 0.6\) and noise standard deviation \(\sigma = 1.5\). The overlaps for linear regression are noticeably lower compared to phase retrieval, showing how the model noise makes the latter problem easier for spectral estimation.

  • In 6, we test our estimators against those in [eqn:ycs-truncated,eqn:lal-truncated] under the same setting of 1 with a change in the prior: each signal is uniform on \(\bbS^{d-1}\) with their correlation being \(\inprod{x_1^*}{x_2^*} = \rho = 0.1\). The results show that our estimators retain their superiority. Similar improvements are observed for \(\rho = 0.3\), verifying the robustness of our estimators to mild signal correlation.

9 Eigenvalues via random matrix theory↩︎

The characterization of the limiting joint law of spectral and linear estimators in 1 is based on the analysis of a Generalized Approximate Message Passing (GAMP) algorithm. The proof of convergence of the GAMP iteration to the desired high-dimensional limit, whenever the conditions \(\lambda^*(\delta_1) > \ol\lambda(\delta)\) and/or \(\lambda^*(\delta_2) > \ol\lambda(\delta)\) are satisfied, crucially relies on the existence of an eigengap in the matrix \(D\) (defined in 8 ). In this section, we derive the limits of the top three eigenvalues of \(D\). This result, stated as 2 below, is then used in 10 to prove 1.

Theorem 2 (Eigenvalues). Consider the setting of 2 and let [itm:assump-signal-distr,itm:assump-alpha,itm:assump-noise-distr,itm:assump-gaussian-design,itm:assump-proportional,itm:assump-preproc-spec] hold. Then we have \[\begin{align} \lim_{d\to\infty} \lambda_1(D) &= \zeta(\lambda^*(\delta_1); \delta) , \quad \lim_{d\to\infty} \lambda_2(D) = \zeta(\lambda^*(\delta_2); \delta) , \quad \lim_{d\to\infty} \lambda_3(D) = \zeta(\ol\lambda(\delta); \delta) , \notag \end{align}\] almost surely. Furthermore,

  1. If \(\lambda^*(\delta_1) > \lambda^*(\delta_2) > \ol\lambda(\delta)\), then \[\begin{align} \zeta(\lambda^*(\delta_1); \delta) > \zeta(\lambda^*(\delta_2); \delta) > \zeta(\ol\lambda(\delta); \delta) ; \notag \end{align}\]

  2. If \(\lambda^*(\delta_1) > \ol\lambda(\delta) \ge \lambda^*(\delta_2)\), then \[\begin{align} \zeta(\lambda^*(\delta_1); \delta) > \zeta(\lambda^*(\delta_2); \delta) = \zeta(\ol\lambda(\delta); \delta) ; \notag \end{align}\]

  3. If \(\ol\lambda(\delta) \ge \lambda^*(\delta_1) > \lambda^*(\delta_2)\), then \[\begin{align} \zeta(\lambda^*(\delta_1); \delta) = \zeta(\lambda^*(\delta_2); \delta) = \zeta(\ol\lambda(\delta); \delta) . \notag \end{align}\]

2 shows a phase transition phenomenon for the top three eigenvalues of \(D\): (i) the top two eigenvalues escape from the bulk of \(D\) if \(\lambda^*(\delta_1) > \lambda^*(\delta_2) > \ol\lambda(\delta)\); (ii) only the largest eigenvalue escapes from the bulk if \(\lambda^*(\delta_1) > \ol\lambda(\delta) \ge \lambda^*(\delta_2)\); (iii) no outlier eigenvalue exists if \(\ol\lambda(\delta)\ge\lambda^*(\delta_1) > \lambda^*(\delta_2)\). See 2 on page . In words, the condition \(\lambda^*(\delta_i)> \ol\lambda(\delta)\) is necessary and sufficient for the \(i\)-th eigenvalue to escape the bulk of the spectrum. This provides an additional piece of evidence (see also [rk:cnvan]) suggesting that such condition is also necessary and sufficient for the corresponding eigenvector to have non-vanishing overlap with the signal. In fact, phase transitions in the behavior of eigenvalues typically correspond to phase transitions in the behavior of the related eigenvectors, see e.g.[28], [29], [93], [94].

Similar results to 2 hold for \(\alpha = 1/2\). In this case, the limits of the first and second eigenvalues of \(D\) coincide and equal \(\zeta(\lambda^*(\delta/2); \delta)\). The limit of the right edge of the bulk of \(D\) does not depend on \(\alpha\) and remains the same (\(\ol\lambda(\delta)\)) as in 2. Therefore, we get two cases: (i) if \(\lambda^*(\delta/2) > \ol\lambda(\delta)\), the top two eigenvalues of \(D\) are repeated outliers; otherwise, (ii) \(\ol\lambda(\delta) \ge \lambda^*(\delta/2)\) and there is no outlier eigenvalue in the limiting spectrum of \(D\).

By the definition of \(\zeta(\lambda;\delta)\) (cf.@eq:eqn:def-zeta-i ), we can write the limits of the eigenvalues in the following more explicit form, which will be convenient in 10: \[\begin{align} \zeta(\lambda^*(\delta_1); \delta) &= \begin{cases} \lambda^*(\delta_1)\paren{\frac{1}{\delta} + \expt{\frac{Z}{\lambda^*(\delta_1) - Z}}} , & \lambda^*(\delta_1) > \ol\lambda(\delta) \\ \ol\lambda(\delta)\paren{\frac{1}{\delta} + \expt{\frac{Z}{\ol\lambda(\delta) - Z}}} , & \lambda^*(\delta_1) \le \ol\lambda(\delta) \end{cases}, \notag \\ \zeta(\lambda^*(\delta_2); \delta) &= \begin{cases} \lambda^*(\delta_2)\paren{\frac{1}{\delta} + \expt{\frac{Z}{\lambda^*(\delta_2) - Z}}} , & \lambda^*(\delta_2) > \ol\lambda(\delta) \\ \ol\lambda(\delta)\paren{\frac{1}{\delta} + \expt{\frac{Z}{\ol\lambda(\delta) - Z}}} , & \lambda^*(\delta_2) \le \ol\lambda(\delta) \end{cases}, \notag \\ \zeta(\ol\lambda(\delta); \delta) &= \ol\lambda(\delta)\paren{\frac{1}{\delta} + \expt{\frac{Z}{\ol\lambda(\delta) - Z}}} . \notag \end{align}\]

Proof of 2. The proof is divided into three steps. Specifically, we first condition on \(\eta_1, \cdots, \eta_n\) and write \(D\) as the sum of two asymptotically free spiked random matrices as on page of the main text. Then, the limit of \(\lambda_3(D)\) is determined in 1. Finally, the limits of \(\lambda_1(D),\lambda_2(D)\) and the monotonicity properties of the limiting eigenvalues in [itm:mono-2,itm:mono-1,itm:mono-0] of the theorem are given by 2. ◻

9.1 Right edge of the bulk of \(D\)↩︎

Before proceeding to the analysis, let us introduce some more notation. Let \[\begin{align} D_1 = \frac{1}{n} A_1^\top T_1A_1, \quad D_2=\frac{1}{n}A_2^\top T_2A_2 . \label{eqn:d1-d2} \end{align}\tag{53}\] Therefore \(D=D_1+D_2\) according to 33 . We first calculate the limiting value of the right edge of the bulk of the spectrum of \(D\).

Lemma 1. Consider the setting of 2. Let [itm:assump-signal-distr,itm:assump-alpha,itm:assump-noise-distr,itm:assump-gaussian-design,itm:assump-proportional,itm:assump-preproc-spec] hold. Denote by \(\mu_D\) the empirical spectral distribution of \(D\). Then \[\begin{align} \lim_{d\to\infty} \sup\supp(\mu_{D}) &= \frac{1}{\delta} \cdot s^{-1}_{\mu_1\boxplus\mu_2}(-1/\ol\lambda(\delta)) , \label{eqn:lim-bulk} \end{align}\qquad{(9)}\] almost surely, where \(\ol\lambda(\delta)\) is the solution to \[\begin{align} \expt{\paren{\frac{Z}{\ol\lambda(\delta) - Z}}^2} &= \frac{1}{\delta} \label{eqn:lam-bar-rmt} \end{align}\qquad{(10)}\] and the function \(s_{\mu_1\boxplus\mu_2}^{-1}\) is defined as \[\begin{align} s_{\mu_1\boxplus\mu_2}^{-1}(z) &= -\frac{1}{z} + \delta\, \expt{\frac{Z}{1+zZ}} . \label{eqn:s-sum-inv-def} \end{align}\qquad{(11)}\]

The function \(s_{\mu_1\boxplus\mu_2}^{-1}\) is the inverse Stieltjes transform of the free additive convolution of the limiting spectral distributions \(\mu_1\) of \(\frac{n}{d}D_1\) and \(\mu_2\) of \(\frac{n}{d}D_2\). Furthermore, \(s_{\mu_1\boxplus\mu_2}^{-1}(\lambda)\) is precisely \(\delta \psi(\lambda;\delta)\), where \(\psi(\lambda;\delta)\) defined in 10 . We note also that the parameter \(\ol\lambda(\delta)\) defined in ?? is the same as that defined through 11 . (See 5.) The connection shall become more transparent in the proof below.

Proof of 1. First note that the scaling factor \(\frac{1}{d}\) in [29] is different from our scaling \(\frac{1}{n}\) in the definition of \(D\) (cf.@eq:eqn:def-mtx-t-and-d ). We therefore consider \(\wt{D} = \frac{1}{d}A^\top TA\) for the convenience of applying Lemma 3 in [29]. All results regarding the matrix \(\wt{D}\) can be translated to \(D\) by inserting a factor \(\frac{d}{n}\to\frac{1}{\delta}\) at proper places, since \(D = \frac{d}{n}\wt{D}\).

Let \[\begin{align} \wt{D}_1 = \frac{1}{d} A_1^\top T_1 A_1 , \quad \wt{D}_2 = \frac{1}{d} A_2^\top T_2 A_2 . \label{eqn:def-di-tilde} \end{align}\tag{54}\] By 33 , \(\wt{D}=\wt{D}_1+\wt{D}_2\). Let \(\mu_1\) and \(\mu_2\) be the limiting spectral distributions of \(\wt{D}_1\) and \(\wt{D}_2\), respectively, as \(n_1,n_2,d\to\infty\) with \(n_1/d\to\delta_1 = \alpha\delta\) and \(n_2/d\to\delta_2 = (1-\alpha)\delta\). As argued on page , \(\wt{D}_1\) and \(\wt{D}_2\) are asymptotically free. Hence, the limiting spectral distribution of \(\wt{D}\) is given by the free additive convolution of \(\mu_1,\mu_2\), denoted by \(\mu_1\boxplus\mu_2\) [95], [96]. It remains to compute \(\sup\supp(\mu_1\boxplus\mu_2)\).

A careful inspection of the proof of [29] shows that the bulk of the spectrum of \(\wt{D}_i\), i.e., \(\lambda_2(\wt{D}_i)\ge\cdots\ge\lambda_d(\wt{D}_i)\), interlaces the spectrum of \(E_i \mathrel{\vcenter{:}}= \frac{1}{d} \wt{A}_i^\top T_i \wt{A}_i\in\bbR^{(d-1)\times(d-1)}\) for \(i\in\{1,2\}\), respectively. Specifically, \[\begin{align} \lambda_1(E_i) \ge \lambda_2(\wt{D}_i) \ge \lambda_2(E_i) \ge \lambda_3(\wt{D}_i) \ge \cdots \ge \lambda_{d-2}(E_i) \ge \lambda_{d-1}(\wt{D}_i) \ge \lambda_{d-1}(E_i) \ge \lambda_d(\wt{D}_i) . \label{eqn:interlace} \end{align}\tag{55}\] Here, \(T_i = \mathop{\mathrm{diag}}(\cT(q(A_i x_i^*,\veps_i)))\) (recall 32 ) and \(\wt{A}_i\in\bbR^{n_i\times(d-1)}\) is an independent matrix with i.i.d.\(\cN(0,1)\) entries. In particular, \(T_i\) and \(\wt{A}_i\) are independent. We also define, for \(i\in \{1,2\}\), \(\wt{E}_i \mathrel{\vcenter{:}}= \frac{1}{d-1} \wt{A}_i^\top T_i \wt{A}_i\). Note that \(E_i = \frac{d-1}{d} \wt{E}_i\), \(n_1/(d-1)\to\delta_1\) and \(n_2/(d-1)\to\delta_2\).

Since each \(y_i\) (for \(1\le i\le n\)) is i.i.d., the limiting spectral distributions of \(T_1\) and \(T_2\) are in fact the same and both equal the law of \(Z\). Thus, Lemma 3 in [29] provides us with a characterization of the limiting spectral distribution of \(\wt{E}_i\): \[\begin{align} \mu_{\wt{E}_1} &\to \wt{\mu}_1 , \quad \mu_{\wt{E}_2} \to \wt{\mu}_2 , \notag \end{align}\] weakly as \(n_1,n_2,d\to\infty\) with \(n_1/(d-1)\to\delta_1,n_2/(d-1)\to\delta_2\). Furthermore, the limiting spectral distributions admit the following explicit description through the inverse Stieltjes transform: \[\begin{align} s_{\wt{\mu}_1}^{-1}(z) &= -\frac{1}{z} + \delta_1 \expt{\frac{Z}{1+zZ}} , \quad s_{\wt{\mu}_2}^{-1}(z) = -\frac{1}{z} + \delta_2 \expt{\frac{Z}{1+zZ}} . \label{eqn:s-inv-mui} \end{align}\tag{56}\] In view of the scaling factor \(\frac{d-1}{d}\to1\), the limiting spectral distributions of \(E_1,E_2\) are also given by \(\wt{\mu}_1,\wt{\mu}_2\), respectively. Recall that the bulks of the spectra of \(\wt{D}_1,\wt{D}_2\) interlace the spectra of \(E_1,E_2\), respectively (cf.@eq:eqn:interlace ). Since \(\wt{D}_1\) and \(\wt{D}_2\) can each have at most one outlier eigenvalue by Lemma 2 in [29], the limiting spectral distributions \(\mu_1,\mu_2\) of \(\wt{D}_1,\wt{D}_2\), respectively, are the same as \(\wt{\mu}_1,\wt{\mu}_2\) whose inverse Stieltjes transforms are shown in 56 .

Let \(R_\mu(z) \mathrel{\vcenter{:}}= s_\mu^{-1}(-z) - \frac{1}{z}\) denote the \(R\)-transform [97] of \(\mu\). Then, a well-known fact in free probability theory is that \(R_{\mu_1\boxplus\mu_2}(z) = R_{\mu_1}(z) + R_{\mu_2}(z)\). Thus, \[\begin{align} s_{\mu_1\boxplus\mu_2}^{-1}(z) &= s_{\mu_1}^{-1}(z) + s_{\mu_2}^{-1}(z) + \frac{1}{z} = s_{\wt{\mu}_1}^{-1}(z) + s_{\wt{\mu}_2}^{-1}(z) + \frac{1}{z} = -\frac{1}{z} + \delta \expt{\frac{Z}{1+zZ}} . \label{eqn:s-inv-sum} \end{align}\tag{57}\] Given \(s_{\mu_1\boxplus\mu_2}^{-1}\), one can calculate \(\sup\supp(\mu_1\boxplus\mu_2)\) which is in turn the limiting value of \(\sup\supp(\mu_{\wt{D}})\), where \(\mu_{\wt{D}}\) denotes the empirical spectral distribution of \(\wt{D}\). This can be accomplished thanks to the results in [98] (see also [99]): \[\begin{align} \lim_{d\to\infty} \sup\supp(\mu_{\wt{D}}) &= \sup\supp(\mu_1\boxplus\mu_2) \tag{58} \\ &= \min_{\lambda>\sup\supp(Z)} s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda) \notag \\ &= \min_{\lambda>\sup\supp(Z)} \lambda + \delta\expt{\frac{Z}{1 - Z/\lambda}} . \tag{59} \end{align}\] The convergence in 58 holds almost surely since \[\begin{align} \mu_{\wt{D}} &= \mu_{\wt{D}_1 + \wt{D}_2} \to \mu_1 \boxplus \mu_2 \notag \end{align}\] weakly [95], [96]. To solve the minimization problem in 59 , we observe that the function \(s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda)\) can be written in terms of \(\psi(\lambda;\delta)\) defined in 10 : \[\begin{align} s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda) &= \lambda + \delta\expt{\frac{Z}{1 - Z/\lambda}} = \delta \lambda \paren{\frac{1}{\delta} + \expt{\frac{Z}{\lambda - Z}}} = \delta \cdot \psi(\lambda;\delta) . \notag \end{align}\] Since \(\psi(\lambda;\delta)\) is convex in the first argument (cf.4), so is \(s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda)\) as a function of \(\lambda\). As a result, the minimizer \(\ol\lambda(\delta)\) in 59 is the critical point of \(s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda)\). That is, \[\begin{align} \left. \frac{\diff}{\diff\lambda} s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda)\right|_{\lambda = \ol\lambda(\delta)} &= 1 - \delta\expt{\paren{\frac{Z}{\ol\lambda(\delta) - Z}}^2} = 0 , \notag \end{align}\] i.e., \(\ol\lambda(\delta)\) is the solution to the following equation \[\begin{align} \expt{\paren{\frac{Z}{\ol\lambda(\delta) - Z}}^2} &= \frac{1}{\delta} . \label{eqn:fp-taustar} \end{align}\tag{60}\] The minimum value in 59 is therefore \[\begin{align} s_{\mu_1\boxplus\mu_2}^{-1}(-1/\ol\lambda(\delta)) &= \ol\lambda(\delta)\paren{1 + \delta\expt{\frac{Z}{\ol\lambda(\delta) - Z}}} . \label{eqn:bulk-of-free-sum-formula} \end{align}\tag{61}\]

At this point, we have successfully computed the limiting value of \(\sup\supp(\mu_{\wt{D}})\). However, recall that the original matrix we are interested in is \(D = \frac{d}{n} (\wt{D}_1 + \wt{D}_2)\). Therefore, \[\begin{align} \lim_{d\to\infty} \sup\supp(\mu_D) &= \lim_{d\to\infty} \frac{d}{n} \sup\supp(\mu_{\wt{D}}) = \ol\lambda(\delta)\paren{\frac{1}{\delta} + \expt{\frac{Z}{\ol\lambda(\delta) - Z}}} , \notag \end{align}\] where \(\ol\lambda(\delta)\) satisfies 60 . This concludes the proof. ◻

9.2 Outlier eigenvalues of \(D\)↩︎

Finally, we need to understand the outliers in the spectrum of \(D\).

Lemma 2. Consider the setting of 2. Let [itm:assump-signal-distr,itm:assump-alpha,itm:assump-noise-distr,itm:assump-gaussian-design,itm:assump-proportional,itm:assump-preproc-spec] hold. Let the function \(s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda)\) be given by ?? . Then, the following statements hold.

  1. \(\lambda^*(\delta_1) > \lambda^*(\delta_2)\);

  2. If \(\lambda^*(\delta_1) > \lambda^*(\delta_2) > \ol\lambda(\delta)\), then \(s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda^*(\delta_1)) > s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda^*(\delta_2))\);

  3. For \(i\in \{1,2\}\), if \(\lambda^*(\delta_i) > \ol\lambda(\delta)\), then \(s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda^*(\delta_i)) > s_{\mu_1\boxplus\mu_2}^{-1}(-1/\ol\lambda(\delta))\);

  4. We have that, almost surely, \[\begin{align} \lim_{d\to\infty} \lambda_1(D) &= \frac{1}{\delta} \cdot s_{\mu_1\boxplus\mu_2}^{-1}(-1/\max\{\lambda^*(\delta_1),\ol\lambda(\delta)\}) , \label{eqn:lim-lam1-d-stieltjes} \\ \lim_{d\to\infty} \lambda_2(D) &= \frac{1}{\delta} \cdot s_{\mu_1\boxplus\mu_2}^{-1}(-1/\max\{\lambda^*(\delta_2),\ol\lambda(\delta)\}) , \label{eqn:lim-lam2-d-stieltjes} \\ \lim_{d\to\infty} \lambda_3(D) &= \frac{1}{\delta} \cdot \sup\supp(\mu_1\boxplus\mu_2) = \frac{1}{\delta} \cdot s_{\mu_1\boxplus\mu_2}^{-1}(-1/\ol\lambda(\delta)) . \label{eqn:lim-lam3-d-stieltjes} \end{align}\] {#eq: sublabel=eq:eqn:lim-lam1-d-stieltjes,eq:eqn:lim-lam2-d-stieltjes,eq:eqn:lim-lam3-d-stieltjes}

Recalling the definition \(\zeta(\lambda;\delta) = \psi(\max\{\lambda,\ol\lambda(\delta)\};\delta)\) (cf.@eq:eqn:def-zeta-i ) and the relation \(\frac{1}{\delta}\cdot s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda) = \psi(\lambda;\delta)\) (cf.[rk:stieltjes-equals-psi]), we can write the limiting values of \(\lambda_1(D),\lambda_2(D),\lambda_3(D)\) in [eqn:lim-lam1-d-stieltjes,eqn:lim-lam2-d-stieltjes,eqn:lim-lam3-d-stieltjes] as \[\begin{align} \zeta(\lambda^*(\delta_1);\delta) \ge \zeta(\lambda^*(\delta_2);\delta) \ge \zeta(\ol\lambda(\delta);\delta), \label{eqn:order} \end{align}\tag{62}\] respectively. To see why the above chain of inequalities holds, note that by [itm:outlier-concl-3] of 2, \(\zeta(\lambda^*(\delta_i); \delta)>\zeta(\ol\lambda(\delta);\delta)\) if \(\lambda^*(\delta_i)>\ol\lambda(\delta)\) and \(\zeta(\lambda^*(\delta_i); \delta)=\zeta(\ol\lambda(\delta);\delta)\) otherwise. So \[\begin{align} \zeta(\lambda^*(\delta_i); \delta)\ge\zeta(\ol\lambda(\delta);\delta) \label{eqn:order-1} \end{align}\tag{63}\] is always true for \(i\in \{1,2\}\). Also, by [itm:outlier-concl-2] of 2, \(\zeta(\lambda^*(\delta_1);\delta)\ge\zeta(\lambda^*(\delta_2);\delta)>\zeta(\ol\lambda(\delta);\delta)\) if \(\lambda^*(\delta_1)\ge\lambda^*(\delta_2)>\ol\lambda(\delta)\). If \(\lambda^*(\delta_1)\ge\ol\lambda(\delta)\ge\lambda^*(\delta_2)\), \(\zeta(\lambda^*(\delta_2);\delta) = \zeta(\ol\lambda(\delta); \delta) \le \zeta(\lambda^*(\delta_1);\delta)\) by 63 . If \(\ol\lambda(\delta)\ge\lambda^*(\delta_1)\ge\lambda^*(\delta_2)\), \(\zeta(\lambda^*(\delta_1);\delta) = \zeta(\lambda^*(\delta_2);\delta) = \zeta(\ol\lambda(\delta); \delta)\). All cases have been exhausted in light of [itm:outlier-concl-1] of 2. In any case, \[\begin{align} \zeta(\lambda^*(\delta_1);\delta)\ge\zeta(\lambda^*(\delta_2);\delta) \label{eqn:order-2} \end{align}\tag{64}\] holds. 62 then follows from [eqn:order-1,eqn:order-2].

Proof of 2. The proof is divided into three parts. We first explicitly evaluate the theoretical predictions of the limiting values of the top three eigenvalues of \(D\). The convergence of the outlier eigenvalues and the right edge of the bulk to the respective predictions is then formally justified in the second part. Finally, several properties concerning the spectral threshold and the limiting eigenvalues are proved in the third part.

9.2.0.1 Limiting eigenvalues

To understand the outlier eigenvalues of \(D = D_1 + D_2\), we need to first understand the outlier eigenvalues of \(D_1\) and \(D_2\) individually. To calibrate the scaling, let us define \[\begin{align} D_1' \mathrel{\vcenter{:}}= \frac{1}{n_1}A_1^\top T_1 A_1 ,\quad D_2' \mathrel{\vcenter{:}}= \frac{1}{n_2}A_2^\top T_2 A_2 .\notag \end{align}\] Lemma 2 in [29] applies to the above matrices \(D_1',D_2'\) and implies that each of \(D_1'\) and \(D_2'\) has a potential outlier eigenvalue \(\lambda_1(D_1')\) and \(\lambda_1(D_2')\), respectively. As \(n_1,n_2,d\to\infty\) with \(n_1/d\to\delta_1\) and \(n_2/d\to\delta_2\), they converge almost surely to the following limiting values: \[\begin{align} \lim_{d\to\infty} \lambda_1(D_1') &= \zeta(\lambda^*(\delta_1); \delta_1) , \quad \lim_{d\to\infty} \lambda_1(D_2') = \zeta(\lambda^*(\delta_2); \delta_2) , \notag \end{align}\] where \(\lambda^*(\delta_1)\) and \(\lambda^*(\delta_2)\) are the solutions to \[\begin{align} \zeta(\lambda^*(\delta_1); \delta_1) &= \phi(\lambda^*(\delta_1)) , \quad \zeta(\lambda^*(\delta_2); \delta_2) = \phi(\lambda^*(\delta_2)) , \notag \end{align}\] respectively. For \(i\in\{1,2\}\), let us assume that \(\lambda_1(D_i')\) is indeed an outlier eigenvalue of \(D_i'\), that is, its limiting value \(\zeta(\lambda^*(\delta_i);\delta_i)\) lies outside the bulk of the limiting spectrum of \(D_i'\). According to Lemma 2 in [29], this happens if and only if \(\lambda^*(\delta_i) > \ol\lambda(\delta_i)\). In this case, the limiting value of the outlier eigenvalue can be written more explicitly as \[\begin{align} \zeta(\lambda^*(\delta_i);\delta_i) = \psi(\lambda^*(\delta_i);\delta_i) = \lambda^*(\delta_i) \paren{\frac{1}{\delta_i} + \expt{\frac{Z}{\lambda^*(\delta_i) - Z}}} , \label{eqn:outlier-of-summand} \end{align}\tag{65}\] where \(\lambda^*(\delta_i)\) is the solution to \[\begin{align} \expt{\frac{Z(G^2 - 1)}{\lambda^*(\delta_i) - Z}} &= \frac{1}{\delta_i} . \label{eqn:def-lam-star-explicit} \end{align}\tag{66}\]

Let us first translate the above result (i.e., [eqn:outlier-of-summand,eqn:def-lam-star-explicit]) regarding \(D_1',D_2'\) to \(\wt{D}_1,\wt{D}_2\) defined in 54 . Since \(n_1/d\to\delta_1,n_2/d\to\delta_2\) and \(\wt{D}_1 = \frac{n_1}{d}D_1',\wt{D}_2 = \frac{n_2}{d}D_2'\), we have that, almost surely, \[\begin{align} \lim_{d\to\infty} \lambda_1(\wt D_i) &= \delta_i\lambda^*(\delta_i)\paren{\frac{1}{\delta_i} + \expt{\frac{Z}{\lambda^*(\delta_i) - Z}}} = \lambda^*(\delta_i)\paren{1 + \delta_i\expt{\frac{Z}{\lambda^*(\delta_i) - Z}}} \eqqcolon\theta_i , \notag \end{align}\] where we have denoted the limiting value of \(\lambda_1(\wt D_i)\) by \(\theta_i\). In view of the definition of \(s_{\mu_i}^{-1}\) in 56 , we recognize that \[\begin{align} \theta_i = s_{\mu_i}^{-1}(-1/\lambda^*(\delta_i)) . \label{eqn:thetai-equals-si-inv} \end{align}\tag{67}\]

Provided with the individual outlier of \(\wt D_i\) (cf.@eq:eqn:thetai-equals-si-inv ), we now invoke [76] to determine how an outlier of \(\wt D_i\) is mapped to the spectrum of \(\wt D\) by the free additive convolution. Specifically, the limiting value, denoted by \(\rho_i\), of the potential outlier of \(\wt{D} = \wt{D}_1+\wt{D}_2\) resulting from \(\theta_i\) is given by \[\begin{align} \rho_i \mathrel{\vcenter{:}}= w_i^{-1}(\theta_i) , \label{eqn:outliers-sum} \end{align}\tag{68}\] where \(w_1,w_2\) are the pair of subordination functions associated with the pair of distributions \(\mu_1,\mu_2\).

As the name suggests, \(w_1,w_2\) enjoy the following subordination property (cf.[76]): \[\begin{align} s_{\mu_1\boxplus\mu_2}(z) &= s_{\mu_1}(w_1(z)) = s_{\mu_2}(w_2(z)) = \frac{1}{z - (w_1(z) + w_2(z))} . \label{eqn:subord-prop} \end{align}\tag{69}\] To understand the value of \(\rho_i = w_i^{-1}(\theta_i)\) (cf.@eq:eqn:outliers-sum ), let us compute \[\begin{align} s_{\mu_1\boxplus\mu_2}(w_i^{-1}(\theta_i)) &= s_{\mu_i}(\theta_i) = -1/\lambda^*(\delta_i). \label{eqn:implies-wi-inv} \end{align}\tag{70}\] The first equality is by the subordination property (69 ) and the second one by the observation in 67 . 70 then gives \[\begin{align} \rho_i &= w_i^{-1}(\theta_i) = s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda^*(\delta_i)) . \label{eqn:rhoi} \end{align}\tag{71}\]

To translate the result in 71 regarding \(\wt{D}\) to \(D\) in 53 , we simply note that \(D = \frac{d}{n}\wt{D}\) and \(d/n\to1/\delta\). Therefore, the limiting eigenvalue of \(D\) resulting from the outlier eigenvalue of \(D_i\) is given by \[\begin{align} \frac{1}{\delta} \cdot \rho_i = \frac{1}{\delta} \cdot s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda^*(\delta_i)) , \label{eqn:outlier-d-ref} \end{align}\tag{72}\] almost surely.

9.2.0.2 Convergence of eigenvalues

We then formally justify that the right edge of the bulk and the outlier eigenvalues of \(D\) indeed converge to the theoretical predictions in [eqn:lim-bulk,eqn:outlier-d-ref], respectively, as \(d\to\infty\), therefore confirming the validity of the latter formulas. Let \(\cK_0 \mathrel{\vcenter{:}}=\supp(\mu_1\boxplus\mu_2)\). For \(i\in\{1,2\}\), let \(\cK_i\) be the singleton set \(\{\rho_i\}\) if \(\theta_i\notin\supp(\mu_i)\) and \(\emptyset\) otherwise. Let \(\cK\mathrel{\vcenter{:}}=\cK_0\cup\cK_1\cup\cK_2\). Then the first statement of [76] guarantees that for any \(\eps>0\), \[\begin{align} \prob{\exists d_0,\,\forall d>d_0,\,\{\lambda_i(\wt D)\}_{i=1}^d\subset \cK_\eps} &= 1 , \label{eqn:use-belinschi-1} \end{align}\tag{73}\] where \(\cK_\eps\) denotes the \(\eps\)-enlargement of \(\cK\), i.e., \[\begin{align} \cK_\eps \mathrel{\vcenter{:}}= \curbrkt{\rho\in\bbR : \inf_{\rho'\in\cK} |\rho - \rho'| \le \eps} . \notag \end{align}\] In words, 73 says that almost surely for every sufficiently large dimension \(d\), the spectrum of \(\wt D\) is contained in an arbitrarily small neighbourhood of \(\cK\). Furthermore, suppose \(\rho\in\cK_1\cup\cK_2\) and \(\rho\notin\cK_0\), that is, \(\rho\) is an outlier in the limiting spectrum of \(\wt D\). Assume also that \(\eps>0\) is sufficiently small so that \((\rho - 2\eps,\rho+2\eps) \cap \cK = \{\rho\}\). Then \[\begin{align} \prob{\exists d_0,\,\forall d>d_0,\,\abs{\{\lambda_i(\wt D)\}_{i=1}^d \cap (\rho-2\eps,\rho+2\eps)} = \indicator{w_1(\rho) = \theta_1} + \indicator{w_2(\rho) = \theta_2}} &= 1 . \label{eqn:use-belinschi-2} \end{align}\tag{74}\] In words, 74 says that almost surely for every sufficiently large dimension \(d\), the outlier \(\theta_1\) (resp.\(\theta_2\)) in the limiting spectrum of \(\wt D_1\) (resp.\(\wt D_2\)) is mapped to \(w_1^{-1}(\theta_1)\) (resp.\(w_2^{-1}(\theta_2)\)) in the limiting spectrum of \(\wt D\). Since \(D,D_1,D_2\) and \(\wt D,\wt D_1,\wt D_2\) only differ by a \(\delta\) factor, similar statements hold true for \(D,D_1,D_2\) as well.

Combining [eqn:outlier-d-ref,eqn:use-belinschi-1,eqn:use-belinschi-2] yields [eqn:lim-lam1-d-stieltjes,eqn:lim-lam2-d-stieltjes,eqn:lim-lam3-d-stieltjes] in [itm:outlier-concl-4] of 2.

9.2.0.3 Properties of spectral threshold and limiting eigenvalues

We identify under what condition \(\rho_i = w_i^{-1}(\theta_i)\) is an outlier in the limiting spectrum of \(\wt D\). For this to be the case, \(\theta_i\) is necessarily an outlier in the limiting spectrum of \(\wt D_i\), which is assumed in the preceding derivations. As Lemma 2 in [29] guaranteed, a sufficient and necessary condition for this event is \(\lambda^*(\delta_i) > \ol\lambda(\delta_i)\). Under the free additive convolution, the outlier \(\theta_i\) of \(\wt D_i\) is then mapped to \(w_i^{-1}(\theta_i)\). Let us compare \(w_i^{-1}(\theta_i)\) with \(\sup\supp(\mu_1\boxplus\mu_2)\), i.e., the right edge of the bulk of the limiting spectral distribution of \(\wt D = \wt D_1 + \wt D_2\). The former quantity equals \(s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda^*(\delta_i))\) (as derived in 71 ) and the latter one equals \(s_{\mu_1\boxplus\mu_2}^{-1}(-1/\ol\lambda(\delta))\) (see 61 in the proof of 1). Recall the following two facts:

  1. \(s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda) = \delta \cdot \psi(\lambda;\delta)\) (as observed in [rk:stieltjes-equals-psi]);

  2. \(\psi(\lambda;\delta)\) is convex in \(\lambda\) and increasing for \(\lambda\in[\ol\lambda(\delta),\infty)\) (proved in 4).

We therefore conclude that \(s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda^*(\delta_i)) > s_{\mu_1\boxplus\mu_2}^{-1}(-1/\ol\lambda(\delta))\) if \(\lambda^*(\delta_i) > \ol\lambda(\delta)\). This establishes [itm:outlier-concl-3] of 2. This condition is more stringent than the previous one \(\lambda^*(\delta_i) > \ol\lambda(\delta_i)\). This can be seen by inspecting the definitions (see, e.g., ?? in 5) of \(\ol\lambda(\delta_i)\) and \(\ol\lambda(\delta)\): \[\begin{align} \expt{\paren{\frac{Z}{\ol\lambda(\delta) - Z}}^2} &= \frac{1}{\delta} , \quad \expt{\paren{\frac{Z}{\ol\lambda(\delta_i) - Z}}^2} = \frac{1}{\delta_i} , \label{eqn:lambda-bar-order} \end{align}\tag{75}\] respectively, and realizing that \(\ol\lambda(\delta) > \ol\lambda(\delta_i)\) since \(\delta > \delta_i\).

We pause and make the following remark regarding the effect of the free additive convolution on the outliers in the spectra of the addends. Comparing 71 with the limiting value of the right edge of the bulk (cf.@eq:eqn:bulk-of-free-sum-formula ), we note the following: \(\lambda_1(\wt{D}_i)\) being an outlier eigenvalue of \(\wt{D}_i\) does not imply that its image \(\rho_i\) under the free additive convolution is also an outlier eigenvalue of \(\wt{D} = \wt{D}_1+\wt{D}_2\). In fact, it can be buried strictly inside the bulk, which happens if \(\ol\lambda(\delta_i) < \lambda^*(\delta_i) < \ol\lambda(\delta)\).

We then show \(\lambda^*(\delta_1) > \lambda^*(\delta_2)\) in [itm:outlier-concl-1] of 2. Recall that \(\lambda^*(\delta_1)\) and \(\lambda^*(\delta_2)\) are the unique solutions to \(\zeta(\lambda^*(\delta_1); \delta_1) = \phi(\lambda^*(\delta_1))\) and \(\zeta(\lambda^*(\delta_2); \delta_2) = \phi(\lambda^*(\delta_2))\), respectively. Since \(\zeta(\cdot;\delta_1),\zeta(\cdot;\delta_2)\) are non-decreasing and \(\phi(\cdot)\) is strictly decreasing, it suffices to show \[\begin{align} \zeta(\lambda;\delta_1) < \zeta(\lambda;\delta_2) \label{eqn:zeta-toshow} \end{align}\tag{76}\] for any \(\lambda>\sup\supp(Z)\). We do so in four steps. (The following arguments are best understood with 2 in mind.)

  1. First we claim that \(\ol\lambda(\delta_1) > \ol\lambda(\delta_2)\). This follows from a similar observation as in 75 and the assumption \(\alpha>1/2\) (cf.[itm:assump-alpha]) which implies \(\delta_1>\delta_2\).

  2. Second we claim that \(\psi(\ol\lambda(\delta_1); \delta_1) < \psi(\ol\lambda(\delta_2); \delta_2)\). Indeed, \[\begin{align} \psi(\ol\lambda(\delta_1); \delta_1) &< \psi(\ol\lambda(\delta_2); \delta_1) < \psi(\ol\lambda(\delta_2); \delta_2) . \notag \end{align}\] The first inequality follows since \(\psi(\lambda;\delta_1)\) is strictly decreasing for \(\lambda\le\ol\lambda(\delta_1)\) (see [itm:property-2] of 4) and \(\ol\lambda(\delta_1) > \ol\lambda(\delta_2)\) as shown in [itm:step1] above. The second inequality follows since \[\begin{align} \psi(\cdot;\delta_1)<\psi(\cdot;\delta_2) \label{eqn:psi-order} \end{align}\tag{77}\] for any \(\lambda>\sup\supp(Z)\) (see the definition of \(\psi\) in 10 and also [itm:property-3] of 4). Note that in this step we use \(\sup\supp(Z)>0\) in [itm:assump-preproc-spec]. This shows that 76 holds for any \(\lambda\le\ol\lambda(\delta_2)\).

  3. We then claim that 76 holds for any \(\lambda\ge\ol\lambda(\delta_1)\). This is because, in this regime, we have \[\begin{align} \zeta(\lambda; \delta_1) = \psi(\lambda; \delta_1) < \psi(\lambda; \delta_2) = \zeta(\lambda; \delta_2) \notag \end{align}\] using the definition of \(\zeta(\cdot;\delta_i)\) (cf.@eq:eqn:def-zeta-i ) and 77 .

  4. Finally, it remains to verify that 76 holds for \(\ol\lambda(\delta_2)\le\lambda\le\ol\lambda(\delta_1)\). Indeed, we have \[\begin{align} \zeta(\lambda; \delta_1) &= \psi(\ol\lambda(\delta_1); \delta_1) <\psi(\ol\lambda(\delta_2); \delta_2) <\psi(\lambda; \delta_2) . \notag \end{align}\] The equality is by definition of \(\zeta(\cdot;\delta_1)\). The first inequality is by [itm:step2] above. The second inequality follows since \(\psi(\cdot;\delta_2)\) is strictly increasing for \(\lambda\ge\ol\lambda(\delta_2)\) (see [itm:property-2] of 4).

Combining [itm:step1,itm:step2,itm:step3,itm:step4] above then proves 76 which implies [itm:outlier-concl-1] of 2.

Since \(\ol\lambda(\delta)\) is the (unique) critical point of \(s_{\mu_1\boxplus\mu_2}^{-1}(-1/\lambda)\) which is increasing for \(\lambda\ge\ol\lambda(\delta)\), [itm:outlier-concl-2] of 2 then follows. This concludes the argument. ◻

10 Joint distribution via Approximate Message Passing↩︎

The limiting joint distribution in 1 is obtained via a generalized approximate message passing (GAMP) algorithm whose iterates converge to the top two eigenvectors of \(D = A^\top T A\). Within this section, we adopt the following rescaling for the convenience of applying the GAMP machinery: \[\begin{align} \Abar &\mathrel{\vcenter{:}}= \frac{1}{\sqrt{d}}\, A , \quad \xone \mathrel{\vcenter{:}}= \sqrt{d}\, x_1^* , \quad \xtwo \mathrel{\vcenter{:}}= \sqrt{d}\, x_2^* , \quad \Dbar \mathrel{\vcenter{:}}= \Abar^\top \, T\, \Abar = \frac{n}{d} A^\top TA . \label{eqn:rescaled} \end{align}\tag{78}\] Due to [itm:assump-signal-distr,itm:assump-gaussian-design], we have \(\Abar _{i,j} \iid \cN(0,1/d)\) and \(\xone,\xtwo \iid \unif(\sqrt{d}\,\bbS^{d-1})\). Let \(\ol{a}_i^\top\in\bbR^d\) denote the \(i\)-th row of \(\Abar\). Then, we have \[\begin{align} y_i &= q\paren{\inprod{a_i}{\eta_i x_1^* + (1-\eta_i) x_2^*}, \eps_i} = q\paren{\inprod{\ol{a}_i}{\eta_i \xone + (1-\eta_i) \xtwo}, \eps_i} . \notag \end{align}\] Therefore, \(y\in\bbR^n\) and related quantities such as \(T\in\bbR^{n\times n}\) (defined in 8 ) do not have to be rescaled. The overlaps are invariant under rescaling of \(D\). Furthermore, since \(n/d\to\delta\), the limiting eigenvalues of \(\Dbar\) are equal to those of \(D\) multiplied by \(\delta\) in view of 78 .

We first extend the GAMP algorithm for the non-mixed GLM [27] and its associated state evolution analysis to the mixed GLM model. The GAMP algorithm is defined in terms of a sequence of Lipschitz functions \(g_t:\mathbb{R}^2 \to \mathbb{R}\) and \(f_{t+1}:\mathbb{R}^3 \to \mathbb{R}\), for \(t \geq 0\). For \(t\ge 0\), the algorithm iteratively computes \(u^t, \tu^t \in \bbR^n\) and \(v^{t+1}, \tv^{t+1} \in \bbR^d\) as follows: \[\begin{align} \begin{aligned} u^t &= \frac{1}{\sqrt{\delta}} \Abar \tv^t - \sfb_t \tu^{t-1}, \quad \tu^t = g_{t}(u^{t};y) , \\ v^{t+1} &= \frac{1}{\sqrt{\delta}} \Abar^\top \tu^t - \sfc_t \tv^t, \quad \tv^{t+1}=f_{t+1}(v^{t+1}; \, \xone, \xtwo). \end{aligned} \label{eq:gamp-eqn-general} \end{align}\tag{79}\] The iteration is initialized with a given \(\tv^0 \in \bbR^d\) and \(\tu^{-1}=0_n\). The functions \(f_t\) and \(g_t\) are applied component-wise, i.e., \(f_t(v^t; \, \xone, \xtwo)=(f_t(v^t_1; \, \ol{x}_{1,1}^*, \ol{x}_{2,1}^*)\), \(\ldots, f_t(v^t_d; \, \ol{x}_{1,d}^*, \ol{x}_{2,d}^*))\) and \(g_t(u^t; y)=(g_t(u^t_1; y_1), \ldots, g_t(u^t_n; y_n))\). The scalars \(\sfb_t, \sfc_t\) are defined as \[\sfb_t =\frac{1}{n}\sum_{i=1}^d f_t'(v_i^t; \, \ol{x}_{1,i}^*, \ol{x}_{2,i}^*), \qquad \sfc_t = \frac{1}{n}\sum_{i=1}^n g_t'(u_i^t; y_i), \label{eq:GAMP95onsager}\tag{80}\] where \(f_t'\) and \(g_t'\) each denote the derivative with respect to the first argument.

An important feature of the GAMP algorithm is that as \(d \to \infty\), the empirical distributions of the iterates \(u^t\) and \(v^{t+1}\) converge to the laws of well-defined scalar random variables \(U_t\) and \(V_{t+1}\), respectively. Specifically, for \(t \ge 0\), let \[U_t \mathrel{\vcenter{:}}= \mu_{1,t} G_1 + \mu_{2,t} G_2 + W_{U,t} , \quad V_{t+1} \mathrel{\vcenter{:}}= \chi_{1,t+1} X_1 + \chi_{2,t+1} X_2 + W_{V,t+1}, \label{eq:UtVt95def}\tag{81}\] where \((G_1, G_2, W_{U,t}) \sim \normal(0,1) \otimes \normal(0,1) \otimes \normal(0, \sigma_{U,t}^2)\), and \((X_1, X_2, W_{V,t+1}) \sim \normal(0,1) \otimes \normal(0,1) \otimes \normal(0,\sigma_{V,t+1}^2)\). The random variables \(X_1, X_2\) are distributed according to limiting laws of the signals \(\xone, \xtwo\), and \(G_1, G_2\) according to the limiting laws of \(\Abar \xone, \Abar \xtwo\). Since \(\xone, \xtwo\) are independent and uniformly distributed on the sphere, we have \(X_1, X_2\iid\normal(0, 1)\). The deterministic coefficients \((\mu_{1,t}, \mu_{2,t}, \sigma_{U,t}, \chi_{1,t+1}, \chi_{2, t+1}, \sigma_{V, t+1})\) are computed using the following state evolution recursion: \[\begin{align} & \mu_{1,t} = \frac{1}{\sqrt{\delta}} \E [ X_1 f_t(V_t; \, X_1, X_2) ], \quad \mu_{2,t} = \frac{1}{\sqrt{\delta}} \E[ X_2 f_t(V_t; \, X_1, X_2) ], \label{eq:muU95update} \\ & \sigma_{U,t}^2 = \frac{1}{\delta}\E[ f_t(V_t; \, X_1, X_2)^2 ] - \mu_{1,t}^2 - \mu_{2,t}^2 \, , \nonumber \\ & \chi_{1,t+1} = \sqrt{\delta} \left( \E[G_1 g_t(U_t; \tY) ] - \E[ g_t'(U_t; \tY)] \mu_{1,t} \right), \notag \\ & \chi_{2,t+1} = \sqrt{\delta} \left( \E [ G_2 g_t(U_t; \tY) ] - \E [g_t'(U_t; \tY)] \mu_{2,t} \right), \nonumber \\ & \sigma_{V,t+1}^2 = \E[g_t(U_t; \tY)^2 ]. \nonumber \end{align}\tag{82}\] Here the random variable \(\tY\) is given by \[\tY= q(\eta G_1 + (1-\eta)G_2, \, \eps), \text{ where } (G_1, G_2, \eta, \eps) \sim \normal(0,1) \ot \normal(0,1) \ot \text{Bern}(\alpha) \ot P_{\eps}. \label{eq:G1G2Y95joint}\tag{83}\] The state evolution recursion is initialized in terms of the limiting correlation of the initializer \(\tv^0\) with each of the signals \(\xone\) and \(\xtwo\). The existence of these limiting correlations is guaranteed by imposing the following condition on \(\tv^0\):

This assumption is typical in AMP algorithms [59], and our initializer for proving 1 will be \(\tv^0=0_d\), which trivially satisfies [ass:gamp-init]. [ass:gamp-init] allows us to initialize the state evolution recursion as: \[\begin{align} \begin{aligned} &\mu_{1,0} =\frac{1}{\sqrt{\delta}} \lim_{d \to \infty} \frac{\langle \xone \, , \, \tv^0 \rangle}{d} = \frac{1}{\sqrt{\delta}} \E[ F_0(X_1, X_2) X_1 ], \\ &\mu_{2,0} =\frac{1}{\sqrt{\delta}} \lim_{d \to \infty} \frac{\langle \xtwo \, , \, \tv^0 \rangle}{d} = \frac{1}{\sqrt{\delta}} \E[ F_0(X_1, X_2) X_2 ], \\ &\sigma_{U,0}^2= \frac{1}{\delta} \lim_{d \to \infty} \frac{\normtwo{\tv^0}^2}{d} - \mu_{1,0}^2 - \mu_{2,0}^2 = \frac{1}{\delta} \E[F_0(X_1, X_2)^2] - \mu_{1,0}^2 - \mu_{2,0}^2. \end{aligned} \label{eq:SE95gen95init} \end{align}\tag{84}\]

The sequences of random variables \((W_{U,t})_{t \geq 0}\) and \((W_{V,t+1})_{t \geq 0}\) in 81 are each jointly Gaussian with zero mean and the following covariance structure: \[\E[W_{U,0} W_{U,t}] = \frac{1}{\delta} \E[F_0(X_1, X_2) \, f_t(V_t; X_1, X_2) ] - \mu_{1,0} \mu_{1,t} - \mu_{1,0} \mu_{2,t}, \qquad t \geq 1, \label{eq:WV95corr95init}\tag{85}\] and for \(r, t \geq 1\): \[\begin{align} \tag{86} & \E [ W_{V,r} W_{V,t} ] = \E\left[g_{r-1}(U_{r-1}; \tY) \, g_{t-1}(U_{t-1}; \tY) \right], \\ & \E[W_{U,r} W_{U,t}] = \frac{1}{\delta}\E [ f_r(V_r; \, X_1, X_2) f_t(V_t; \, X_1, X_2) ] - \mu_{1,r} \mu_{1,t} - \mu_{1,r} \mu_{2,t}. \tag{87} \end{align}\] Note that for \(r=t\) we have \(\E [ W_{U,t}^2 ] = \sigma_{U,t}^2\) and \(\E [W_{V,t+1}^2 ]=\sigma_{V,t}^2\).

The state evolution result for the GAMP is stated in terms of pseudo-Lipschitz test functions (see 16 ).

Consider the setup of 1 and the GAMP iteration in 79 , with initialization \(\tv^0\) that satisfies [ass:gamp-init]. Assume that for \(t \ge 0\), the functions \(g_t: \reals^2 \to \reals\) and \(f_{t+1}: \reals^3 \to \reals\) are Lipschitz. Let \(g_1 \mathrel{\vcenter{:}}= \Abar \xone, g_2\mathrel{\vcenter{:}}= \Abar \xtwo\). Then, the following holds almost surely for any \(\PL(2)\) function \(\Psi: \reals^{t+3} \to \reals\), for \(t \geq 0\): \[\begin{align} & \lim_{n \to \infty}\frac{1}{n} \sum_{i=1}^n \Psi( g_{1, i}, g_{2, i}, u^t_i, u^{t-1}_i, \ldots, u^0_{i} ) = \E[ \Psi( G_1, G_2, \, U_t, \, U_{t-1}, \, \ldots, U_0) ], \tag{88} \\ & \lim_{d \to \infty} \frac{1}{d} \sum_{i=1}^d \Psi(\ol{x}_{1,i}^*, \ol{x}_{2,i}^*, v^{t+1}_i, v^{t}_i, \ldots, v^1_i) = \E[ \Psi(X_1, \, X_2, \, V_{t+1}, \, V_{t}, \, \ldots, V_1) ], \tag{89} \end{align}\] where the distributions of the random vectors \(( G_1, G_2, U_t, \ldots, U_{0})\) and \((X_1, X_2, V_{t+1}, \ldots, V_1)\) are given by the state evolution recursion in 81 to 87 .

The proof of the proposition, given in 15, uses a reduction to an abstract AMP recursion with matrix-valued iterates for which a state evolution result was established in [77].

The result in 89 is equivalent to the statement that the joint empirical distribution of the rows of \((\xone, \xtwo, v^t, \ldots, v^1)\) converges in Wasserstein-\(2\) distance to the joint law of \((X_1, X_2, V_t, \ldots, V_1)\) (see [59]). A similar equivalence holds for the result in 88 .

The result in [prop:GAMP95SE] also applies to the GAMP algorithm in which the memory coefficients \((\sfb_t, \sfc_t)\) in 80 are replaced with their deterministic limits \(\bar{\sfb}_t, \bar{\sfc}_t\) computed via state evolution: \[\bar{\sfb}_t = \frac{1}{\delta} \E[ f_t'(V_t; \, X_1, X_2) ], \qquad \bar{\sfc}_t = \E[ g_t'(U_t; \tY) ]. \label{eq:detOnsager}\tag{90}\] This equivalence follows from an argument similar to [59].

10.1 GAMP as a method to compute the linear and spectral estimators↩︎

Consider the GAMP iteration in 79 with the initializer \(\tv^0=0\), and the following choice of functions: \[\begin{align} & g_0(u^0; y) = \sqrt{\delta} \cL(y), \qquad f_1(v; \, \xone, \xtwo)= f(\xone, \xtwo), \\ & g_t(u; \, y) = \sqrt{\delta} \, u \, \cF(y), \quad f_{t+1}(v; \, \xone, \xtwo) = \frac{v}{\beta_{t+1}}, \quad t \ge 1, \label{eq:ft95gt95choice} \end{align}\tag{91}\] where \(\cF: \reals \to \reals\) is bounded and Lipschitz, \(f: \reals \to \reals\) is Lipschitz, and \(\beta_{t+1}\) is a constant, defined iteratively for \(t \ge 0\) via the state evolution equations below (96 ). To prove 1, we will consider two different choices for the pair of functions \((f, \cF)\), in terms of the spectral preprocessing function \(\cT\) (see [eq:F1x1,eq:F2x2]).

With the above choice of \(f_t, g_t\), the memory coefficients in 80 are given by \[\label{eq:coeffbct} \sfc_0= \sfb_1=0, \qquad \sfc_t = \sqrt{\delta} \cdot \frac{1}{n}\sum_{i=1}^n \cF(y_i), \qquad \sfb_{t+1} =\frac{1}{\delta \beta_{t+1}}.\tag{92}\] Replacing the parameter \(\sfc_t\) with its almost sure limit \(\bar{\sfc}_t = \sqrt{\delta} \, \E[\cF(\tY)]\), the GAMP iteration becomes \[\label{eq:newGAMP} \begin{align} & u^0 =0, \quad v^1= \Abar^\top \cL(y), \\ & u^1= \frac{1}{\sqrt{\delta}} \Abar f(\xone, \xtwo), \quad v^2 = \frac{1}{\sqrt{\delta}} \Abar^\top F u^1 - \sqrt{\delta} \E[ \cF(\tY) ] f(\xone,\xtwo), \\ & u^{t} = \frac{1}{\sqrt{\delta} \, \beta_{t}} \paren{\Abar v^{t} \, - \, F u^{t-1}}, \quad v^{t+1} = \Abar^\top F u^t - \frac{\sqrt{\delta}}{\beta_t} \, \E[ \cF(\tY) ] \, v^t, \qquad t \ge 2, \end{align}\tag{93}\] where \(F = \mathop{\mathrm{diag}}(\cF(y_1), \ldots, \cF(y_n))\). With \(f_t, g_t\) given by 91 , the initialization for the state evolution in 82 to 84 is: \[\begin{align} & \mu_{1,0} =\mu_{2,0}= \sigma_{U,0}^2=0, \nonumber \\ &\chi_{1,1} = \delta \E[ G_1 \cL(\tY)], \quad \chi_{2,1} = \delta \E[ G_2 \cL(\tY)], \quad \sigma_{V,1}^2 = \delta \E[\cL(\tY)^2 ], \nonumber \\ & \mu_{1,1} = \frac{1}{\sqrt{\delta}}\expt{X_1 f(X_1, X_2)}, \quad \mu_{2,1} = \frac{1}{\sqrt{\delta}}\expt{X_2 f(X_1, X_2)}, \nonumber \\ & \sigma_{U,1}^2 = \frac{1}{\delta} \expt{f(X_1, X_2)^2} - \mu_{1,1}^2 - \mu_{2,1}^2, \label{eq:SE95lin95init} \end{align}\tag{94}\] where the joint distribution of \((G_1, G_2, \tY)\) is given by 83 . Furthermore, for \(t \ge 1\): \[\begin{align} & \chi_{1,t+1} = \delta \mu_{1,t} \, \E [ \cF(\tY)(G_1^2-1) ], \qquad \chi_{2,t+1} = \delta \mu_{2,t} \, \E [ \cF(\tY)(G_2^2-1) ], \nonumber \\ & \sigma_{V,t+1}^2 = \delta \paren{ \mu_{1,t}^2 \E [\cF(\tY)^2 G_1^2] + \mu_{2,t}^2 \E[\cF(\tY)^2 G_2^2] + \sigma_{U,t}^2 \E[ \cF(\tY)^2] } , \tag{95} \\ & \beta_{t+1} \mathrel{\vcenter{:}}= \sqrt{\chi_{1,t+1}^2 + \chi_{2,t+1}^2 + \sigma_{V,t+1}^2} \, , \tag{96} \\ & \mu_{1,t+1} = \frac{\chi_{1,t+1}}{\sqrt{\delta} \beta_{t+1}}, \qquad \mu_{2,{t+1}} = \frac{\chi_{2,t+1}}{\sqrt{\delta} \beta_{t+1}}, \qquad \sigma_{U,t+1}^2= \frac{\sigma_{V,t+1}^2}{\delta \beta_{t+1}^2}. \tag{97} \end{align}\]

First note that the iterate \(v^1\) coincides with the linear estimator \(\xlin\) in 6 . We will show that in the high-dimensional limit the iterate \(v^t\) is aligned with an eigenvector of the matrix \(M \mathrel{\vcenter{:}}= \Abar^\top F(\sqrt{\delta} \beta_{\infty}I_n + F)^{-1} \Abar\), as \(t \to \infty\). (3 shows that \(\beta_{\infty} = \lim\limits_{t \to \infty} \beta_t\) is well-defined for our choices of \(\cF\) and initializations.) For a heuristic justification of this claim, assume the iterates \(u^t, v^t\) converge to the limits \(u^\infty, v^\infty\) in the sense that \(\lim\limits_{t \to \infty} \lim\limits_{d \to \infty} \frac{1}{d} \normtwo{u^t - u^{\infty}}^2 =0\) and \(\lim\limits_{t \to \infty} \lim\limits_{d \to \infty} \frac{1}{d} \normtwo{v^t - v^{\infty}}^2 =0\). Then, from 93 these limits satisfy \[u^\infty = \frac{1}{\sqrt{\delta} \, \beta_\infty} \paren{ \Abar v^\infty \, - \, F u^{\infty} }, \qquad v^{\infty} = \Abar^\top F u^\infty - \frac{\sqrt{\delta}}{\beta_\infty} \, \E[ \cF(\tY) ] \, v^\infty,\] which after simplification, can be written as: \[v^{\infty}\left( 1 + \frac{\sqrt{\delta}}{\beta_\infty} \E[ \cF(\tY)] \right) = \Abar^{\top}F (\sqrt{\delta}\beta_{\infty}I_n + F)^{-1}\Abar v^{\infty}. \label{eq:evector95eqn}\tag{98}\] Therefore, \(v^{\infty}\) is an eigenvector of the matrix \(\Abar^{\top}F (\sqrt{\delta}\beta_{\infty}I_n + F)^{-1}\Abar\), and the GAMP iteration 93 is effectively a power method.

We wish to obtain via GAMP the two leading eigenvectors of the matrix \(\Abar^{\top}T\Abar\), so the heuristic above indicates that we should choose \[\begin{align} \cF(y) &= \frac{c \sqrt{\delta} \beta_\infty \cT(y)}{1 - c \cT(y)} , \notag \end{align}\] so that \(F (\sqrt{\delta}\beta_{\infty}I_n + F)^{-1} = c\, T\), for some constant \(c\). For estimating the \(i\)-th signal, we fix the values of \(\beta_\infty\) and \(c\) by enforcing the following two constraints: \[\begin{align} 1 &= \lim_{d\to\infty} \frac{\normtwo{\wt{v}^\infty}}{\sqrt{d}} = \lim_{d\to\infty} \frac{\normtwo{v^\infty}}{\beta_\infty \sqrt{d}} , \notag \\ \frac{1}{\sqrt{\delta}} &= \lim_{d\to\infty} \frac{\normtwo{v^\infty}}{\sqrt{d}} = \sqrt{\delta} \expt{ \frac{c \sqrt{\delta} \beta_\infty \cT(\wt{Y})}{1 - c \cT(\wt{Y})} (G_i^2 - 1) } , \notag \end{align}\] where the last equality in the second line is by state evolution (formally shown in 3). Upon simplifications, the above two conditions are equivalent to \[\begin{align} && \beta_\infty &= \frac{1}{\sqrt{\delta}} , & c &= \frac{1}{\lambda^*(\delta_i)} , & & \notag \end{align}\] which in turn motivates the choice of \(\cF\): \[\begin{align} \cF(y) &= \frac{\cT(y)}{\lambda^*(\delta_i) - \cT(y)} . \notag \end{align}\]

Formally, we analyze the iteration in 93 with two choices for the function \(\cF(y)\) and initialization \(\tv^0\): \[\begin{align} &\text{ Choice 1}: \quad \cF_1(y) \mathrel{\vcenter{:}}= \frac{\cT(y)}{\lambda^*(\delta_1) - \cT(y)}, \quad f(\xone, \xtwo)=\xone, \tag{99} \\ &\text{ Choice 2}: \quad \cF_2(y) \mathrel{\vcenter{:}}= \frac{\cT(y)}{\lambda^*(\delta_2) - \cT(y)}, \quad f(\xone, \xtwo)=\xtwo. \tag{100} \end{align}\] Here, we recall that, for \(i\in\{1, 2\}\), \(\lambda^*(\delta_i)\) is the unique solution of \(\zeta(\lambda;\delta_i) = \phi(\lambda)\) (see page ). The initializations in [eq:F1x1,eq:F2x2] are not feasible in practice since they depend on the unknown signals \(\xone\) and \(\xtwo\), but this is not an issue as we use the GAMP in 91 only as a proof technique.

We now examine the state evolution recursion in [eq:chi_sigmaV_lin,eq:beta_t_def,eq:mu_sigmaU_lin] under each of these choices.

10.1.0.1 Choice 1

From 94 , this corresponds to the initialization \[\begin{align} &\chi_{1,1} = \delta \E[ G_1 \cL(\tY)], \; \chi_{2,1} = \delta \E[ G_2 \cL(\tY)], \; \sigma_{V,1}^2 = \delta \E[\cL(\tY)^2 ], \; \mu_{1,1} = \frac{1}{\sqrt{\delta}}, \; \mu_{2,1} = \sigma_{U,1}^2 = 0. \label{eq:SElin95init95x1} \end{align}\tag{101}\] For \(t \ge 1\), the state evolution equations in [eq:chi_sigmaV_lin,eq:beta_t_def,eq:mu_sigmaU_lin] reduce to: \[\begin{align} & \chi_{1,t+1} = \delta \mu_{1,t} \, \E [ \cF_1(\tY)(G_1^2-1) ], \quad \sigma_{V,t+1}^2 = \delta \left( \mu_{1,t}^2 \E[ \cF_1(\tY)^2 G_1^2 ] \, + \, \sigma_{U,t}^2 \E[ \cF_1(\tY)^2 ] \right), \\ & \beta_{t+1} = \sqrt{\chi_{1,t+1}^2 + \sigma_{V,t+1}^2} \, , \qquad \mu_{1,t+1} = \frac{\chi_{1,t+1}}{\sqrt{\delta} \beta_{t+1}} \, , \qquad \sigma_{U,t+1}^2= \frac{\sigma_{V,t+1}^2}{\delta \beta_{t+1}^2} \, , \\ \end{align} \label{eq:mu195chi195lin}\tag{102}\] and \(\mu_{2,t+1} = \chi_{2,t+1} =0\) for \(t \ge 1\). Using this in [prop:GAMP95SE], we obtain that: \[\lim_{d \to \infty} \frac{\la \xone, v^{1} \ra}{d} = \chi_{1,1}, \; \lim_{d \to \infty} \frac{\la \xtwo, v^{1} \ra}{d} = \chi_{2,1}, \; \lim_{d \to \infty} \frac{\la \xone, v^{t+1} \ra}{d} = \chi_{1,t+1}, \; \lim_{d \to \infty} \frac{\la \xtwo, v^{t+1} \ra}{d} = 0,\] for \(t \ge 1\). Thus, when initialized with \(f(\xone, \xtwo) =\xone\), the GAMP iterates \(\{ v^{t+1} \}_{t \ge 1}\) are asymptotically uncorrelated with the signal \(\xtwo\).

10.1.0.2 Choice 2

This corresponds to the initialization \[\begin{align} &\chi_{1,1} = \delta \E[ G_1 \cL(\tY)], \; \chi_{2,1} = \delta \E[ G_2 \cL(\tY)], \; \sigma_{V,1}^2 = \delta \E[\cL(\tY)^2 ], \; \mu_{2,1} = \frac{1}{\sqrt{\delta}}, \; \mu_{1,1}=\sigma_{U,1}^2 = 0. \label{eq:SElin95init95x2} \end{align}\tag{103}\] The state evolution equations are: \(\mu_{1,t+1} = \chi_{1,t+1} =0\) for \(t \ge 1\), and \[\begin{align} & \chi_{2,t+1} = \delta \mu_{2,t} \, \E [ \cF_2(\tY)(G_2^2-1) ], \quad \sigma_{V,t+1}^2 = \delta \left( \mu_{2,t}^2 \E[ \cF_2(\tY)^2 G_2^2 ] \, + \, \sigma_{U,t}^2 \E[ \cF_2(\tY)^2 ] \right), \\ & \beta_{t+1} = \sqrt{\chi_{2,t+1}^2 + \sigma_{V,t+1}^2} \, , \qquad \mu_{2,t+1} = \frac{\chi_{2,t+1}}{\sqrt{\delta} \beta_{t+1}}, \qquad \sigma_{U,t+1}^2= \frac{\sigma_{V,t+1}^2}{\delta \beta_{t+1}^2} \, . \end{align} \label{eq:mu295chi295lin}\tag{104}\] Using this in [prop:GAMP95SE], we obtain that for \(t \ge 1\), \[\lim_{d \to \infty} \frac{\la \xone, v^{1} \ra}{d} = \chi_{1,1}, \; \lim_{d \to \infty} \frac{\la \xtwo, v^{1} \ra}{d} = \chi_{2,1}, \; \lim_{d \to \infty} \frac{\la \xone, v^{t+1} \ra}{d} = 0, \; \lim_{d \to \infty} \frac{\la \xtwo, v^{t+1} \ra}{d} = \chi_{2,t+1}.\] The following lemma gives the fixed point of state evolution under choices 1 and 2.

Lemma 3 (Limiting values of state evolution parameters). Consider the state evolution recursion under choice \(i \in \{1,2\}\). Assume assume that \(\E [ \cF_i(\tY)(G_i^2-1) ]>0\) and \(\delta > \frac{\E [ \cF_i(\tY)^2 ]}{ (\E[ \cF_i(\tY)(G_i^2-1)])^2}\). Then, as \(t\to \infty\) the state evolution parameters \((\chi_{i,t}, \sigma_{V,t}^2)\) converge to the fixed point \((\tilde{\chi}_i, \tilde{\sigma}_i^2 )\), where \[\tilde{\chi}_i = \sqrt{\frac{\tilde{\beta}_i^2(\tilde{\beta}_i^2 - \E[\cF_i(\tY)^2])}{\tilde{\beta}_i^2 + \E[\cF_i(\tY)^2G_i^2] - \E[\cF_i(\tY)^2] }}, \qquad \tilde{\sigma}_i^2 = \frac{\tilde{\beta}_i^2 \E[\cF_i(\tY)^2 G_i^2]}{ \tilde{\beta}_i^2 + \E[\cF_i(\tY)^2G_i^2] - \E[\cF_i(\tY)^2]}, \label{eq:FP1}\qquad{(12)}\] and \[\tilde{\beta}_i^2 = \tilde{\chi}_i^2 + \tilde{\sigma}_i^2 = \delta\, ( \E[\cF_i(\tY)(G_i^2-1)])^2. \label{eq:beta195def}\qquad{(13)}\]

Proof. The proof is identical to that of Lemma 5.2 in [25], which analyzes GAMP for a non-mixed GLM with \(f_t, g_t\) given by 91 . The state evolution recursion under choice 1 in [eq:SElin_init_x1,eq:mu1_chi1_lin] has the same form for all values of \(\alpha \in [1/2,1)\). The value of \(\alpha\) affects the recursion only through the joint distribution of \((\wt{Y}, G_1) = (q(\eta G_1 + (1-\eta)G_2,\eps), G_1)\), where \(\eta \sim \bern(\alpha)\). The proof of Lemma 5.2 in [25] does not depend on this joint distribution and applies for any \(\alpha\) such that the lower bound on \(\delta\) in the statement of the first part is satisfied. The argument for choice 2, where the joint distribution determining the state evolution in [eq:SElin_init_x2,eq:mu2_chi2_lin] is \((\wt{Y}, G_2) = (q(\eta G_1 + (1-\eta)G_2,\eps), G_2)\), is identical. ◻

It is convenient to express the state evolution fixed points in 3 in terms of the joint law of \((G, Y)\), where \(Y = q(G, \eps)\), with \(G \sim \normal(0,1)\) and \(\eps \sim P_{\eps}\) are independent. Recalling the joint law of \((\tY, G_1, G_2)\) given in 83 and the definitions of \(\cF_1,\cF_2\) in [eq:F1x1,eq:F2x2], we have \[\begin{align} \E[\cF_1(\tY)] &= \E[\cF_1(Y)] = \E\left[ \frac{\cT(Y)}{\lambda^*(\delta_1) - \cT(Y)} \right], \notag \\ \E[\cF_1(\tY)^2] &= \E[\cF_1(Y)^2] = \E\left[ \frac{\cT(Y)^2}{ (\lambda^*(\delta_1) - \cT(Y))^2} \right ], \notag \\ \E[\cF_1(\tY)G_1^2] &= \alpha \E[\cF_1(q(G_1, \eps)) G_1^2 ] + (1-\alpha) \E[\cF_1(q(G_2, \eps)) G_1^2 ] \notag \\ &= \alpha \E \left[ \frac{\cT(Y) G^2}{\lambda^*(\delta_1) - \cT(Y)} \right] + (1-\alpha) \E \left[ \frac{\cT(Y)}{\lambda^*(\delta_1) - \cT(Y)} \right] \notag \\ &= \frac{1}{\delta} + \E \left[ \frac{\cT(Y)}{\lambda^*(\delta_1) - \cT(Y)} \right], \label{eq:SE95GY1} \end{align}\tag{105}\] where the last equality holds because \(\E \left[ \frac{\cT(Y) (G^2-1)}{\lambda^*(\delta_1) - \cT(Y)} \right] = \frac{1}{\delta_1}\) from ?? , and \(\delta_1 = \alpha \delta\). Similarly, we obtain \[\begin{align} & \E[\cF_2(\tY)] = \E[\cF_2(Y)] = \E\left[ \frac{\cT(Y)}{\lambda^*(\delta_2) - \cT(Y)} \right], \quad \E[\cF_2(\tY)^2] = \E\left[ \frac{\cT(Y)^2}{ (\lambda^*(\delta_2) - \cT(Y))^2} \right ], \nonumber \\ & \E[\cF_2(\tY)G_2^2] = \frac{1}{\delta} + \E \left[ \frac{\cT(Y)}{\lambda^*(\delta_2) - \cT(Y)} \right]. \label{eq:SE95GY21} \end{align}\tag{106}\] Using [eq:SE_GY1,eq:SE_GY21], the formula for \(\tbeta_i^2\) in ?? becomes: \[\begin{align} \tbeta_i^2 = \frac{1}{\delta}, \qquad i\in\{1,2\}. \label{eq:tbeta195def} \end{align}\tag{107}\] We similarly obtain \[\begin{align} \tilde{\chi}_i = \frac{\rho_i^{\spec}}{\sqrt{\delta}}, \quad \tilde{\sigma}_1^2 = \frac{1- (\rho_i^{\spec})^2}{\delta}, \qquad i\in\{1,2\}. \label{eq:tchi95tsigma95formulas} \end{align}\tag{108}\] where \(\rho_1^{\spec}, \rho_2^{\spec}\) are defined in 15 .

10.1.0.3 Proof heuristic

Let us revisit the heuristic sanity-check in 98 . For \(i \in \{1,2\}\), under choice \(i\) with \(\cF= \cF_i\), \(F= F_i\mathrel{\vcenter{:}}=\mathop{\mathrm{diag}}(\cF_i(y_1), \ldots, \cF_i(y_n))\), and \(\beta_{\infty} = \tilde{\beta}_i\), by using the formulas above for \(\tilde{\beta}_i\) and \(\E[\cF_i(\tY)]\), 98 becomes: \[v^{\infty}\left( 1 + \delta \E\left[ \frac{\cT(Y)}{\lambda^*(\delta_i) - \cT(Y)} \right] \right) = \Abar^{\top}F_i ( I_n + F_i)^{-1}\Abar v^{\infty} = \frac{1}{\lambda^*(\delta_i)} \Abar^{\top} T \Abar v^\infty, \label{eq:evector95eqn951}\tag{109}\] where we recall that \(T= \mathop{\mathrm{diag}}(\cT(y_1), \ldots, \cT(y_n))\). Therefore, with choice \(i\), 109 suggests that the GAMP iterate converges to an eigenvector of \(\Dbar = \Abar^\top T \Abar\) corresponding to the eigenvalue \(\lambda^*(\delta_i)\left( 1 + \delta \E\left[ \frac{\cT(Y)}{\lambda^*(\delta_i) - \cT(Y)} \right] \right)\). Moreover, when \(\lambda^*(\delta_i) > \ol{\lambda}(\delta)\), 2 and [rk:explicit-formula-eigval] tell us that the leading eigenvalue of \(\Dbar\) converges to: \[\label{eq:limeig1} \lambda_i( \Dbar ) \stackrel{d \to \infty}{\longrightarrow} \lambda^*(\delta_i)\left( 1 + \delta \E\left[ \frac{\cT(Y)}{\lambda^*(\delta_i) - \cT(Y)} \right] \right).\tag{110}\] Therefore, 109 indicates that the GAMP iterates under each choices \(1\) and \(2\) converge to the eigenvectors corresponding to the two largest eigenvalues of \(\Dbar\), when \(\lambda^*(\delta_1) > \lambda^*(\delta_2) > \ol{\lambda}(\delta)\). We now make this claim rigorous.

10.2 Proof of 1↩︎

Consider the GAMP iteration in 93 for \(t \ge 2\). By substituting the expression for \(u^t\) in the \(v^{t+1}\) update, the iteration can be rewritten as follows: \[\begin{align} u^t &= \frac{1}{\sqrt{\delta} \, \beta_t} \paren{ \Abar v^t \, - \, F u^{t-1} }, \qquad v^{t+1} = \frac{1}{\sqrt{\delta} \beta_t}\sqrbrkt{ \paren{\Abar^\top F \Abar - \delta \E[\cF(\tY)] \, I_d} v^t \, - \, \Abar^\top F^2 u^{t-1} }. \label{eq:GAMPrewrite1} \end{align}\tag{111}\] In the remainder of the proof, we will assume that \(t \geq 2\). Define \[\begin{align} e_1^t &= u^{t}-u^{t-1},\tag{112}\\ e_2^t &= v^{t+1}-v^{t}.\tag{113} \end{align}\] By combining 112 with 111 , we have \[\label{eq:ut1new} u^{t-1}=(F+\sqrt{\delta}\beta_tI_n)^{-1}\Abar v^t-\sqrt{\delta}\beta_t(F+\sqrt{\delta}\beta_tI_n)^{-1}e_1^t.\tag{114}\] Substituting the expression for \(u^{t-1}\) in 114 into 111 and recalling from [eq:SE_GY1,eq:SE_GY21] that \(\expt{\cF(\wt{Y})} = \expt{\cF(Y)}\), we obtain: \[\begin{align} v^{t+1} &= \left(\Abar^\top F(F+\sqrt{\delta}\beta_t I_n)^{-1} \Abar - \frac{\sqrt{\delta} \E[\cF(Y)]}{\beta_t} \, I_d \right) v^t +\Abar^\top F^2(F+\sqrt{\delta}\beta_tI_n)^{-1}e_1^t \nonumber \\ &= \left(\Abar^\top F(F+I_n)^{-1} \Abar -\delta \E[\cF(Y)] \, I_d \right) v^t \notag \\ &\quad +(1-\sqrt{\delta}\beta_t)\Abar^\top F(F+I_n)^{-1}(F+\sqrt{\delta}\beta_tI_n)^{-1} \Abar v^t \nonumber \\ & \quad + \delta\E[\cF(Y)]\left(1-\frac{1}{\sqrt{\delta}\beta_t}\right) v^t+\Abar^\top F^2(F+\sqrt{\delta}\beta_t I_n)^{-1}e_1^t. \label{eq:GAMPrewrite3} \end{align}\tag{115}\] Let \[\label{eq:GAMPrewrite4} e_3^t = \left(\Abar^\top F(F+I_n)^{-1} \Abar -(\delta \E[\cF(Y)]+1) \, I_d \right) v^t.\tag{116}\] Using this in 115 along with 113 , we obtain \[\label{eq:deferr3} \begin{align} e_3^t&=e_2^t-(1-\sqrt{\delta}\beta_t)\Abar^\top F(F+I_n)^{-1}(F+\sqrt{\delta}\beta_tI_n)^{-1} \Abar v^t \\ & \quad - \delta\E[\cF(Y)]\left(1-\frac{1}{\sqrt{\delta}\beta_t}\right) v^t-\Abar^\top F^2(F+\sqrt{\delta}\beta_tI_n)^{-1}e_1^t. \end{align}\tag{117}\]

We now prove the two claims of 1 via choices 1 and 2, respectively. All the limits in the remainder of the proof hold almost surely, so we won’t specify this explicitly.

10.2.1 Proof of ??↩︎

Consider the GAMP algorithm with choice 1, as defined in 99 . With \(\cF(y) = \cF_1(y)\), we have: \[F(F + I_n)^{-1} = \frac{1}{\lambda^*(\delta_1)} T , \qquad \E[\cF(Y)] = \E\left[ \frac{\cT(Y)}{ \lambda^*(\delta_1) - \cT(Y)} \right]. \label{eq:choice95195params}\tag{118}\] Recalling the notation \(\Dbar = \Abar^{\top} T \Abar\), let us decompose \(v^t\) into a component in the direction of \(\eigvone\) plus an orthogonal component \(r_1^t\): \[\label{eq:projvt1} v^t = \xi_{1,t}\,\eigvone+r_1^t,\tag{119}\] where \(\xi_{1,t}=\la v^t, \eigvone \ra\). Substituting 119 in the definition of \(e_3^t\) in 116 and using 118 , we obtain \[\begin{gather} \left( \frac{\Dbar}{\lambda^*(\delta_1)} - \E \left[ \frac{ \delta \cT(Y)}{ \lambda^*(\delta_1) - \cT(Y)} + 1 \right] I_d \right) r_1^t \\ = e_3^t \, + \, \xi_{1,t} \left( \delta \E \left[ \frac{\cT(Y)}{ \lambda^*(\delta_1) - \cT(Y)} \right] + 1 - \frac{\lambda_1(\Dbar)}{\lambda^*(\delta_1)}\right) \eigvone . \label{eq:r1t95exp} \end{gather}\tag{120}\]

The idea of the proof is to prove that \(\lim\limits_{t \to \infty}\lim\limits_{d \to \infty} \| r_1^t \|_2^2/d = 0\), which from 119 implies that the GAMP iterate is aligned with \(\eigvone\) in the limit. To show this, we first claim that for all sufficiently large \(n\):\[\begin{align} \normtwo{ \left( \frac{\Dbar}{\lambda^*(\delta_1)} - \E \left[ \frac{ \delta \cT(Y)}{ \lambda^*(\delta_1) - \cT(Y)} + 1 \right] I_d \right) r_1^t} \ge C \normtwo{r_1^t}, \label{eq:r1t95LB} \end{align}\tag{121}\] for some constant \(C >0\) that does not depend on \(n\). We then consider the right side of 120 and show that under choice 1: \[\begin{align} \lim_{t \to \infty} \lim_{d \to \infty} \, \frac{1}{\sqrt{d}} \normtwo{e_3^t \, + \, \xi_{1,t} \left( \delta \E \left[ \frac{\cT(Y)}{ \lambda^*(\delta_1) - \cT(Y)} \right] + 1 - \frac{\lambda_1(\Dbar)}{\lambda^*(\delta_1)}\right) \eigvone} =0. \label{eq:xi195e395lim} \end{align}\tag{122}\] We now derive the result in ?? using [eq:r1t_LB,eq:xi1_e3_lim], deferring the proofs of these claims to the end of the section. Using [eq:r1t_LB,eq:xi1_e3_lim] in 120 , we have that \[\lim_{t \to \infty}\lim_{d \to \infty} \frac{ \normtwo{ r_1^t }^2}{d} =0. \label{eq:r1t95lim}\tag{123}\] From the decomposition of \(v^t\) in 119 , we have \[\normtwo{v^t}^2 = \xi_{1,t}^2 \, + \, \normtwo{r_{1}^t}^2 \, , \label{eq:vtnorm95decomp}\tag{124}\] since \(r_{1}^t\) is orthogonal to \(\eigvone\) and \(\normtwo{\eigvone}=1\). From [prop:GAMP95SE], we have \[\begin{align} \lim_{d \to \infty} \frac{ \normtwo{v^t}^2}{d} = \E[V_t^2]=\beta_t^2, \qquad t\ge1. \end{align}\] Moreover, from 3 and 107 , under choice 1, \(\lim\limits_{t \to \infty} \beta_t^2 = \tbeta_1^2 = \frac{1}{\delta}\). Therefore, \[\lim_{t \to \infty} \lim_{d \to \infty} \frac{ \normtwo{v^t}^2}{d} = \frac{1}{\delta}. \label{eq:vt95def}\tag{125}\] Combining this with [eq:vtnorm_decomp,eq:r1t_lim] yields \[\lim_{t \to \infty} \lim_{d \to \infty} \frac{\xi_{1,t}^2}{d} = \frac{1}{\delta}. \label{eq:xi1lim}\tag{126}\] Using [eq:xi1lim,eq:r1t_lim] in 119 , and recalling the definition of \(x^{\spec}_1\) from the statement of 1, we have \[\begin{align} \lim_{t\to \infty} \lim_{d \to \infty} \frac{\| \sqrt{\delta} \, v^t - x_1^{\spec} \|_2}{\sqrt{d}} =0. \label{eq:vt95x1spec95diff} \end{align}\tag{127}\] For any \(\PL(2)\) function \(\Psi: \reals^3 \to \reals\), by an application of Cauchy-Schwarz inequality, we have that [59] \[\begin{align} & \abs{\frac{1}{d} \sum_{i=1}^d \Psi( \ol{x}^*_{1,i}, x^{\lin}_i, x^{\spec}_{1,i}) - \frac{1}{d} \sum_{i=1}^d \Psi( \ol{x}^*_{1,i}, x^{\lin}_i, \sqrt{\delta} v^{t}_{i})} \nonumber \\ & \le C \frac{\|\sqrt{\delta} v^t - x_1^{\spec} \|_2}{\sqrt{d}} \left(1 + \frac{\| \xone \|_2}{\sqrt{d}} + \frac{\| x^{\lin} \|_2}{\sqrt{d}} + \frac{\| x_1^{\spec} \|_2}{\sqrt{d}} + \frac{\| v^t \|_2}{\sqrt{d}} \right). \label{eq:Psi95diff1} \end{align}\tag{128}\] We have that \(\| \xone \|_2 = \| x^{\lin} \|_2 = \| x_1^{\spec} \|_2 =\sqrt{d}\), by the definitions in the theorem statement. Therefore, using [eq:vt_def,eq:vt_x1spec_diff] in 128 , we obtain: \[\begin{align} \lim_{d \to \infty} \, \abs{\frac{1}{d} \sum_{i=1}^d \Psi( \ol{x}^*_{1,i}, x^{\lin}_i, x^{\spec}_{1,i}) - \frac{1}{d} \sum_{i=1}^d \Psi( \ol{x}^*_{1,i}, x^{\lin}_i, \sqrt{\delta} v^{t}_{i})} = 0. \end{align}\] Recall from 93 that the GAMP iterate \(v^1 = \xlin\), and \(x^{\lin} = \sqrt{d} \, \xlin/\normtwo{\xlin}\). From [prop:GAMP95SE], we have that \(\lim\limits_{d \to \infty} \frac{\normtwo{\xlin}}{d} = \sqrt{\E[V_1^2]}\). Using [prop:GAMP95SE] again, we have that \[\begin{align} \lim_{d \to \infty} \frac{1}{d} \sum_{i=1}^d \Psi\left( \ol{x}^*_{1,i}, x^{\lin}_i, \sqrt{\delta} v^{t}_{i} \right) = \expt{\Psi\left(X_1, \, \frac{V_1}{\sqrt{\E[V_1^2]}} ,\, \sqrt{\delta} V_t \right) } . \label{eq:Psi95conv1} \end{align}\tag{129}\] From the definitions of \(V_1, V_t\) in 81 , and the state evolution equations for choice 1 in [eq:SElin_init_x1,eq:mu1_chi1_lin], we have \[\begin{align} & V_1 = \chi_{1,1} X_1 + \chi_{2,1} X_2 + W_{V,1}, \quad V_t= \chi_{1,t} X_1 \, + \, W_{V,t}, \quad t \ge 2. \end{align}\] Here \(W_{V,1} \sim \normal(0, \, \delta \E[\cL(\tY)^2 ] )\) and \(W_{V,t} \sim \normal(0, \sigma_{V,t}^2)\) are independent of \((X_1, X_2)\), and from [eq:WV_corr,eq:ft_gt_choice], their covariance is given by \[\begin{align} \expt{W_{V,1}W_{V,t}}= \delta \expt{ U_{t-1} \cL(\tY) \cF(\tY)} = \delta \alpha \mu_{1, t-1} \expt{ G \cL(Y) \cF(Y)}, \end{align}\] where in the last line we have used that \(\mu_{2,t-1} =0\) under choice 1. Hence, for \(t \ge 2\), 129 becomes \[\begin{align} & \lim_{d \to \infty} \frac{1}{d} \sum_{i=1}^d \Psi\left( \ol{x}^*_{1,i}, x^{\lin}_i, \sqrt{\delta} v^{t}_{i} \right) \notag \\ &= \expt{\Psi\left(X_1, \, \frac{\chi_{1,1} X_1 + \chi_{2,1} X_2 + W_{V,1}}{\sqrt{\chi_{1,1}^2 + \chi_{2,1}^2 + \delta \E[\cL(Y)^2]}} ,\, \sqrt{\delta} (\chi_{1,t} X_1 + W_{V,t}) \right) }. \label{eq:Psi95conv2} \end{align}\tag{130}\] To obtain the result in ?? , we take \(t \to \infty\) on both sides above and show that \[\begin{align} & \lim_{t \to \infty} \expt{\Psi\left(X_1, \, \frac{\chi_{1,1} X_1 + \chi_{2,1} X_2 + W_{V,1}}{\sqrt{\chi_{1,1}^2 + \chi_{2,1}^2 + \delta \E[\cL(Y)^2 ]}} ,\, \sqrt{\delta}(\chi_{1,t} X_1 + W_{V,t}) \right) } \\ & = \expt{\Psi\left(X_1, \, \frac{\chi_{1,1} X_1 + \chi_{2,1} X_2 + W_{V,1}}{\sqrt{\chi_{1,1}^2 + \chi_{2,1}^2 + \delta \E[\cL(Y)^2 ]}} ,\, \sqrt{\delta}(\tilde{\chi}_{1} X_1 + W_{V, \infty}) \right) }, \end{align} \label{eq:limExpfinal}\tag{131}\] where \((W_{V,1}, W_{V, \infty})\) are jointly Gaussian with \[W_{V, \infty} \sim \normal( 0, \tilde{\sigma}_{1}^2), \qquad \expt{W_{V,1} W_{V, \infty} } = \tilde{\chi}_1 \delta \alpha \expt{ G \cL(Y) \cF(Y)}.\] Here \(\tilde{\chi}_{1}\) and \(\tilde{\sigma}_{1}^2\) are given by 108 . Using 3, we have \[\begin{align} & \lim_{t \to \infty} \expt{W_{V,t}^2} = \lim_{t \to \infty} \sigma_{V,t}^2 = \tilde{\sigma}_1^2 = \expt{W_{V, \infty}^2}, \\ & \lim_{t \to \infty} \expt{W_{V,1}W_{V,t}} = \delta \alpha \expt{ G \cL(Y) \cF(Y)} \lim_{t \to \infty} \mu_{1, t-1} = \tilde{\chi}_1 \delta \alpha \expt{ G \cL(Y) \cF(Y)} = \expt{W_{V,1} W_{V, \infty} } , \end{align} \label{eq:rv95conv}\tag{132}\] where in the second line, we have used the formula for \(\mu_{1, t-1}\) from 102 and that \(\lim\limits_{t \to \infty} \beta_t =1/\sqrt{\delta}\) (from 107 ). 132 implies that the sequence of zero mean jointly Gaussian pairs \((W_{V,1}, W_{V,t})_{t \geq 1}\) converges in distribution to the jointly Gaussian pair \((W_{V,1}, W_{V, \infty})\). To show 131 , we use Lemma 4.5 in [100]. We apply this result taking \(Q_t\) to be the distribution of \[\left(X_1, \, \frac{\chi_{1,1} X_1 + \chi_{2,1} X_2 + W_{V,1}}{\sqrt{\chi_{1,1}^2 + \chi_{2,1}^2 + \delta \E[\cL(Y)^2 ]}} ,\, \sqrt{\delta}( \chi_{1,t} X_1 + W_{V,t}) \right), \;\] and \(Q\) to be the distribution of \[\left(X_1, \, \frac{\chi_{1,1} X_1 + \chi_{2,1} X_2 + W_{V,1}}{\sqrt{\chi_{1,1}^2 + \chi_{2,1}^2 + \delta \E[\cL(Y)^2 ]}} ,\, \sqrt{\delta}(\tilde{\chi}_{1} X_1 + W_{V, \infty}) \right).\] Since \(\chi_{1,t} \to \tilde{\chi}_1\) and the limits in 132 hold, the sequence \((Q_t)_{t \geq 2}\) converges weakly to \(Q\). In our case, \(\Psi: \reals^3 \to \reals\) is PL(\(2\)), and therefore \(\Psi(a,b,c) \leq C'(1 + \abs{a}^2 + \abs{b}^2+\abs{c}^2)\), for all \((a,b,c) \in \reals^3\) for some constant \(C'\). Choosing \(h(a,b,c) = \abs{a}^2 + \abs{b}^2 + \abs{c}^2\), we have \(\frac{\abs{\Psi}}{1 + h} \leq C'\). Furthermore, \(\int h \, \de Q_t\) is a linear combination of \(\{ \chi_{1,t}^2, \sigma_{V,t}^2\}\), with coefficients that do not depend on \(t\). The integral \(\int h \, \de Q\) has the same form, except that \(\chi_{1,t}, \sigma_{V,t}\) are replaced by \(\tilde{\chi}_{1}, \tilde{\sigma}_{1}\), respectively. Since \(\chi_{1,t} \to \tilde{\chi}_{1}\), \(\sigma_{V,t} \to \tilde{\sigma}_{1}\), we have that \(\lim_{t \to \infty} \int h \, \de Q_t = \int h \, \de Q\). Therefore, by applying Lemma 4.5 in [100], we have that \[\lim_{t \to \infty} \int \Psi \, \de Q_t = \int \Psi \, \de Q, \label{eq:Psi95Qint95conv}\tag{133}\] which is equivalent to 131 . From 108 , we recall that \(\tilde{\chi}_1 = \frac{\rho_1^{\spec}}{\sqrt{\delta}}\), \(\tilde{\sigma}^2_1 = \frac{1 - (\rho_1^{\spec})^2 }{\delta}\). Using these and the formulas for \(\chi_{1,1}, \chi_{2,1}\) from 101 in 131 , and taking \(t \to \infty\) in 130 yields the result in ?? .

It remains to prove [eq:r1t_LB,eq:xi1_e3_lim].

10.2.1.1 Proof of 121

We recall from 2 and [rk:explicit-formula-eigval] that when \(\lambda^*(\delta_1) > \ol{\lambda}(\delta)\), the top eigenvalue of \(\Dbar \mathrel{\vcenter{:}}= \Abar^\top T \Abar\) converges almost surely to \[\begin{align} & \lim_{d \to \infty} \lambda_1( \Dbar ) = \lambda^*(\delta_1)\left( 1 + \delta \E\left[ \frac{\cT(Y)}{\lambda^*(\delta_1) - \cT(Y)} \right] \right). \\ \end{align} \label{eq:eigD95lims}\tag{134}\] Moreover, when \(\lambda^*(\delta_1) > \ol{\lambda}(\delta)\), 2 also guarantees a strict separation between the first and second eigenvalues, i.e., \[\lim_{d \to \infty} \lambda_1( \Dbar ) > \lim_{d \to \infty} \lambda_2( \Dbar) = \zeta(\lambda^*(\delta_2); \delta) \ge \lim_{d \to \infty} \lambda_3( \Dbar ) = \zeta(\ol{\lambda}(\delta); \delta). \label{eq:eigD95lims2}\tag{135}\] Let \[\begin{align} M_1 \mathrel{\vcenter{:}}= \frac{\Dbar}{\lambda^*(\delta_1)} - \E \left[ \frac{ \delta \cT(Y)}{ \lambda^*(\delta_1) - \cT(Y)} + 1 \right] I_d . \label{eq:M1def} \end{align}\tag{136}\] As \(M_1\) is symmetric, it can be written as \(M_1 = Q \Lambda Q^{\top}\), with \(Q\) an orthogonal matrix consisting of the eigenvectors of \(M_1\) and \(\Lambda\) a diagonal matrix with the eigenvalues. Note that the eigenvectors of \(M_1\) are the same as those of \(\Dbar\) and its eigenvalues are: \[\lambda_i(M_1) = \frac{\lambda_i(\Dbar)}{\lambda^*(\delta_1)} - \E \left[ \frac{ \delta \cT(Y)}{ \lambda^*(\delta_1) - \cT(Y)} + 1 \right], \quad i=1, \ldots,d. \label{eq:M195eigs}\tag{137}\] Since \(r_1^t\) is orthogonal to \(\eigvone = v_1(M_1)\), we have \(M_1 r_1^t = Q \Lambda' Q^{\top}r_1^t\), where \(\Lambda'\) is obtained from \(\Lambda\) by replacing \(\lambda_1(M_1)\) with any other value. Here we replace \(\lambda_1(M_1)\) by \(\lambda_2(M_1)\). We therefore have \[\begin{align} \normtwo{M_1 r_1^t}^2 & = \normtwo{ Q \Lambda' Q^{\top} r_1^t}^ 2 \ge \normtwo{r_1^t}^2 \, \min_{ s\in\bbS^{d-1} } \normtwo{ Q \Lambda' Q^{\top} s}^2 \notag\\ &= \normtwo{r_1^t}^2 \, \min_{ s\in\bbS^{d-1} }\langle s, Q (\Lambda')^2 Q^{\top} s \rangle = \normtwo{r_1^t}^2 \, \lambda_d( Q (\Lambda')^2 Q^{\top}), \label{eq:norm95M1r1} \end{align}\tag{138}\] where the last equality follows from the variational characterization of the smallest eigenvalue of a symmetric matrix (Courant–Fischer theorem). Note that \[\lambda_d( Q (\Lambda')^2 Q^{\top}) =\lambda_d( (\Lambda')^2 ) = \min_{i \in \{2, \ldots, d\}} \lambda_i(M_1)^2. \label{eq:min95lambdai95M1}\tag{139}\] From the formula for \(\lambda_i(M_1)\) in 137 and the limiting eigenvalues of \(\Dbar\) in [eq:eigD_lims,eq:eigD_lims2], we have \[\lim_{d \to \infty} \lambda_1(M_1) =0, \qquad \lim_{d \to \infty} \, \min_{i\in \{2, \ldots, d\}} \lambda_i(M_1)^2 = C >0, \label{eq:lim95lambdai95M1}\tag{140}\] for a universal constant \(C\). Combining [eq:norm_M1r1,eq:min_lambdai_M1,eq:lim_lambdai_M1] shows that the lower bound in 121 holds for all sufficiently large \(n\).

10.2.1.2 Proof of 122

Since \(\normtwo{\eigvone}=1\), by the triangle inequality we have \[\begin{align} & \normtwo{e_3^t \, + \, \xi_{1,t} \left( \delta \E \left[ \frac{\cT(Y)}{ \lambda^*(\delta_1) - \cT(Y)} \right] + 1 - \frac{\lambda_1(\Dbar)}{\lambda^*(\delta_1)}\right) \eigvone} \nonumber \\ & \le \normtwo{e_3^t} \, + \, \abs{\xi_{1,t}} \cdot \abs{ \delta \E \left[ \frac{\cT(Y)}{ \lambda^*(\delta_1) - \cT(Y)} \right] + 1 - \frac{\lambda_1(\Dbar)}{\lambda^*(\delta_1)}}. \label{eq:trineq1} \end{align}\tag{141}\] From [eq:xi1lim,eq:eigD_lims] we have: \[\begin{align} \lim_{t \to \infty} \lim_{d \to \infty} \, \frac{\abs{\xi_{1,t}}}{\sqrt{d}} = \frac{1}{\sqrt{\delta}}, \qquad \lim_{d \to \infty} \, \abs{ \delta \E \left[ \frac{\cT(Y)}{ \lambda^*(\delta_1) - \cT(Y)} \right] + 1 - \frac{\lambda_1(\Dbar)}{\lambda^*(\delta_1)}}=0. \label{eq:xit95lim} \end{align}\tag{142}\] Therefore, the second term in 141 converges to \(0\). For the term \(\normtwo{e_3^t}\), using the triangle inequality in the expression in 117 we obtain: \[\label{eq:e395trianeq} \begin{align} \normtwo{e_3^t} \le & \, \normtwo{e_2^t} + \abs{1-\sqrt{\delta}\beta_t} \normop{\Abar } \normop{F(F+I_n)^{-1}(F+\sqrt{\delta}\beta_tI_n)^{-1}} \normop{\Abar } \normtwo{v^t} \\ & + \delta \abs{\E[\cF(Y)]\left(1-\frac{1}{\sqrt{\delta}\beta_t}\right)} \normtwo{v^t} + \normop{\Abar } \normop{F^2(F+\sqrt{\delta}\beta_tI_n)^{-1}} \normtwo{e_1^t}. \end{align}\tag{143}\] Here we have used the fact that \(\normtwo{Mv} \le \normop{M} \normtwo{v}\) for any matrix \(M\) and vector \(v\), and that the operator norm is sub-multiplicative.

Since \(\Abar\) is i.i.d.Gaussian, its operator norm is bounded almost surely as \(n\) grows. With \(\cF(y)\) given by choice 1 in 99 , the diagonal matrices in 143 become: \[\begin{align} & F(F+I_n)^{-1}(F+\sqrt{\delta}\beta_t I_n)^{-1}= \frac{1}{ \lambda^*(\delta_1)}\, T[\lambda^*(\delta_1) I_n - T][\lambda^*(\delta_1)\sqrt{\delta} \beta_t I_n + (1- \sqrt{\delta}\beta_t)T]^{-1}, \nonumber \\ & F^2(F+\sqrt{\delta}\beta_tI_n)^{-1} = T^2[\lambda^*(\delta_1) I_n - T]^{-1}[\lambda^*(\delta_1)\sqrt{\delta} \beta_t I_n + (1- \sqrt{\delta}\beta_t)T]^{-1}. \label{eq:Fdiag95matrices} \end{align}\tag{144}\] Recalling that \(\beta_t >0\) for \(t > 0\), \(\lim\limits_{t \to \infty} \sqrt{\delta}\beta_t =1\) and that \(\cT(\cdot)\) is bounded, the operator norms of the diagonal matrices in 144 are both bounded.

[prop:GAMP95SE] and 3 together imply that \[\begin{align} \label{eq:vt95lim} \lim_{t \to \infty} \lim_{d \to \infty} \frac{\normtwo{v^t}^2}{d} = \lim_{t \to \infty} \beta_t^2 = \frac{1}{\delta}. \end{align}\tag{145}\] Recalling that \(e_1^t= u^t-u^{t-1}\) and \(e_2^t=v^{t+1}-v^t\), using [prop:GAMP95SE] and 3 we can also show [25] that \[\begin{align} \lim_{t \to \infty} \lim_{n \to \infty} \frac{\normtwo{e_1^t}^2}{n} = 0, \qquad \lim_{t \to \infty} \lim_{d \to \infty} \frac{\normtwo{e_2^t}^2}{d} = 0. \label{eq:e1t95e2t95lims} \end{align}\tag{146}\] Using [eq:Fdiag_matrices,eq:vt_lim,eq:e1t_e2t_lims] in 143 shows that \(\lim\limits_{t \to \infty} \lim\limits_{d \to \infty} \normtwo{e_3^t}/\sqrt{d}=0\) which, together with [eq:trineq1,eq:xit_lim], completes the proof of 122 .

10.2.2 Proof of ??↩︎

With choice 2, as defined in 100 , we have \(\cF(y) = \cF_2(y)\), which yields \[F(F + I_n)^{-1} = \frac{1}{\lambda^*(\delta_2)} T, \qquad \E[\cF(Y)] = \E\left[ \frac{\cT(Y)}{ \lambda^*(\delta_2) - \cT(Y)} \right]. \label{eq:choice95295params}\tag{147}\] We decompose the GAMP iterate \(v^t\) into a component in the direction of \(\eigvtwo\) plus an orthogonal component \(r_2^t\): \[\label{eq:projvt2} v^t = \xi_{2,t} \eigvtwo+r_2^t,\tag{148}\] where \(\xi_{2,t}=\la v^t, \eigvtwo \ra\). Substituting 148 in the definition of \(e_3^t\) in 116 and using 147 , we obtain \[\begin{gather} \left( \frac{\Dbar}{\lambda^*(\delta_2)} - \E \left[ \frac{ \delta \cT(Y)}{ \lambda^*(\delta_2) - \cT(Y)} + 1 \right] I_d \right) r_2^t \\ = e_3^t \, + \, \xi_{2,t} \left( 1+ \delta \E \left[ \frac{\cT(Y)}{ \lambda^*(\delta_2) - \cT(Y)} \right] - \frac{\lambda_2(\Dbar)}{\lambda^*(\delta_2)} \right) \eigvtwo . \label{eq:r2t95exp} \end{gather}\tag{149}\] We show that \(\lim\limits_{t \to \infty}\lim\limits_{d \to \infty} \| r_2^t \|^2/d = 0\), which from 148 implies that the GAMP iterate is aligned with \(\eigvtwo\) in the limit. To show this, we first claim that for all sufficiently large \(n\):\[\begin{align} \normtwo{ \left( \frac{\Dbar}{\lambda^*(\delta_2)} - \E \left[ \frac{ \delta \cT(Y)}{ \lambda^*(\delta_2) - \cT(Y)} + 1 \right] I_d \right) r_2^t} \ge C \normtwo{r_2^t}, \label{eq:r2t95LB} \end{align}\tag{150}\] for some constant \(C >0\). We then show that under choice 2: \[\begin{align} \lim_{t \to \infty} \lim_{d \to \infty} \, \frac{1}{\sqrt{d}} \normtwo{e_3^t \, + \, \xi_{2,t} \left( \delta \E \left[ \frac{\cT(Y)}{ \lambda^*(\delta_2) - \cT(Y)} \right] + 1 - \frac{\lambda_2(\Dbar)}{\lambda^*(\delta_2)}\right) \eigvtwo} =0. \label{eq:xi295e395lim} \end{align}\tag{151}\] Given the claims in [eq:r2t_LB,eq:xi2_e3_lim], the result in ?? is obtained using the same steps as 123 to 133 , by replacing \(\xone, x_1^{\spec}, r_{1}^t, \alpha, \xi_{1,t}, \tilde{\chi}_1, \tilde{\sigma}_1\) with \(\xtwo, x_2^{\spec}, r_{2}^t, (1-\alpha), \xi_{2,t}, \tilde{\chi}_2, \tilde{\sigma}_2\), respectively.

The proof of 150 is along the same lines as that of 121 . Other than replacing notation as above, the only change is in the argument from 136 to 140 . Here we define the matrix \(M_2 \mathrel{\vcenter{:}}= \frac{\Dbar}{\lambda^*(\delta_2)} - \E \left[ \frac{ \delta \cT(Y)}{ \lambda^*(\delta_2) - \cT(Y)} + 1 \right]I_d\), which can be written as \(M_2 = Q \bar{\Lambda} Q^{\top}\), where \(\bar{\Lambda}\) is a diagonal matrix containing the eigenvalues of \(M_2\), and \(Q\) is an orthogonal matrix with the eigenvectors. The eigenvectors of \(M_2\) are the same as those of \(\Dbar\) and its eigenvalues are: \[\lambda_i(M_2) = \frac{\lambda_i(\Dbar)}{\lambda^*(\delta_2)} - \E \left[ \frac{ \delta \cT(Y)}{ \lambda^*(\delta_2) - \cT(Y)} + 1 \right], \quad i=1, \ldots,d. \label{eq:M295eigs}\tag{152}\] 2 and [rk:explicit-formula-eigval] guarantee that when \(\lambda^*(\delta_1) > \lambda^*(\delta_2) > \ol{\lambda}(\delta)\), there is strict separation between the top three three eigenvalues of \(\Dbar \mathrel{\vcenter{:}}= \Abar^\top T \Abar\). The limits of these eigenvalues are: \[\begin{align} \lim_{d \to \infty} \lambda_1( \Dbar ) &= \lambda^*(\delta_1)\left( 1 + \delta \E\left[ \frac{\cT(Y)}{\lambda^*(\delta_1) - \cT(Y)} \right] \right) \\ & > \lim_{d \to \infty} \lambda_2( \Dbar ) = \lambda^*(\delta_2)\left( 1 + \delta \E\left[ \frac{\cT(Y)}{\lambda^*(\delta_2) - \cT(Y)} \right] \right) \\ & > \lim_{d \to \infty} \lambda_3( \Dbar ) = \bar{\lambda}(\delta)\left( 1 + \delta \E\left[ \frac{\cT(Y)}{\bar{\lambda}(\delta) - \cT(Y)} \right] \right). \end{align} \label{eq:eigD95lims3}\tag{153}\] Since \(r_2^t\) is orthogonal to \(\eigvtwo = v_2(M_2)\), we have \(M_2 r_2^t = Q \bar{\Lambda}' Q^{\top}r_1^t\), where \(\bar{\Lambda}'\) is obtained from \(\bar{\Lambda}\) by replacing the second eigenvalue \(\lambda_2(M_2)\) with any other value. Here we replace \(\lambda_2(M_2)\) by \(\lambda_1(M_2)\). Then, using the eigenvalue limits in 153 together with arguments analogous to [eq:norm_M1r1,eq:min_lambdai_M1,eq:lim_lambdai_M1], we obtain: \[\lim_{d \to \infty} \lambda_2(M_2) =0, \qquad \lim_{d \to \infty} \, \min_{i \ne 2} \lambda_i(M_2)^2 = C >0, \label{eq:lim95lambdai95M2}\tag{154}\] for a universal constant \(C\).

The proof of 151 is essentially identical to that of 122 , and is omitted. This completes the proof of 1.

To obtain the result mentioned in [rk:alpha-half-master], we analyze a pair of GAMP algorithms with the same design of denoisers and initializers as in choices 1 and 2. In particular, one can show that \(v^t\) converges to a pair of linearly independent vectors in the span of \(v_1(\Dbar)\) and \(v_2(\Dbar)\) under choices 1 and 2. To prove the claim, under choice 1, we decompose \(v^t\) into the projection onto \(\spn\curbrkt{v_1(\Dbar), v_2(\Dbar)}\) and the orthogonal component \(r_1^t\). Via a similar analysis, 123 can be shown to hold, provided that \(\lambda^*(\delta/2) > \ol\lambda(\delta)\), since this condition ensures the existence of a spectral gap (cf.[rk:alpha-half-rmt]). Hence, \(v^t\) converges to a vector in \(\spn\curbrkt{v_1(\Dbar), v_2(\Dbar)}\). Furthermore, 129 continues to hold by state evolution ([prop:GAMP95SE]). As a result, \(v^t\) converges to a vector \(\tv_1\) whose limiting empirical distribution has the law of \(\tilde{\chi}_1 X_1 + W_{V,\infty}\), with \(W_{V,\infty}\) independent of \((X_1,X_2)\). Similarly, under choice 2, \(v^t\) converges to another vector \(\tv_2\) in \(\spn\curbrkt{v_1(\Dbar), v_2(\Dbar)}\) whose limiting empirical distribution has the law of \(\tilde{\chi}_2 X_2 + W_{V,\infty}'\), with \(W_{V,\infty}'\) independent of \((X_1,X_2)\). By recognizing that the inner products of \(\tv_1 , \tv_2\) with \(x_1^*\) differ, one readily obtains that \(\tv_1,\tv_2\) are linearly independent. Therefore, in the high-dimensional limit, the GAMP iterates recover \(\spn\curbrkt{v_1(\Dbar),v_2(\Dbar)}\). At this point, we can find the vector in \(\spn\curbrkt{v_1(\Dbar),v_2(\Dbar)}\) with the desired limiting joint law as in [eq:psiX1joint,eq:psiX2joint] by matching its correlation with the linear estimator \(x^{\lin}\) via 17 . This last step of grid search can be effectively carried out when the overlap attained by \(x^{\lin}\) is non-zero, i.e., \(\expt{G\cL(Y)}\ne0\).

11 Bayes-optimal combination (proof of 1)↩︎

In the proof, we only consider combined estimators for the estimation of \(x_1^*\). The arguments for the estimation of \(x_2^*\) are similar, and therefore omitted. For any \(C_1\in\cC_1\) (the latter set is defined in 19 ), consider the combined estimator \(x_1^{\comb} \mathrel{\vcenter{:}}= C_1(x^{\lin}, x_1^{\spec})\). By 1, we have \[\begin{align} \lim_{d\to\infty} \frac{\abs{\inprod{x_1^{\comb}}{x_1^*}}}{\normtwo{x_1^{\comb}}\normtwo{x_1^*}} &= \frac{\abs{\expt{X_1 C_1(X^{\lin}, X_1^{\spec})}}}{\sqrt{\expt{C_1(X^{\lin}, X_1^{\spec})^2}}} . \notag \end{align}\] almost surely. Here we use the fact that \(\expt{X_1^2} = 1\) which follows from \(\xone\in\sqrt{d}\,\bbS^{d-1}\). Now, the optimality of the conditional expectation function \(C_1^*\) follows from the Cauchy–Schwarz inequality: \[\begin{align} \frac{\abs{\expt{X_1 C_1(X^{\lin}, X_1^{\spec})}}}{\sqrt{\expt{C_1(X^{\lin}, X_1^{\spec})^2}}} &= \frac{\abs{\expt{\expt{X_1 \condon X^{\lin}, X_1^{\spec}} C_1(X^{\lin}, X_1^{\spec})}}}{\sqrt{\expt{C_1(X^{\lin}, X_1^{\spec})^2}}} \notag \\ &\le \sqrt{\expt{\expt{X_1 \condon X^{\lin}, X_1^{\spec}}^2}}, \label{eqn:bayes-opt-bound} \end{align}\tag{155}\] with equality in 155 if \(C_1=C_1^*\).

We then compute \(\expt{X_1 \condon X^{\lin}, X_1^{\spec}}\) from the joint distribution of \((X_1, X^{\lin}, X^{\spec})\) given by [eqn:def-asymp-corr,eqn:def-rv-xlin-xspec]. Under [itm:assump-signal-distr], we have \((X_1,X_2)\sim\cN(0,1)^{\ot2}\). Using this it can be verified that \((X_1, X^{\lin}, X^{\spec})\) are jointly Gaussian with zero mean and the following covariance matrix: \[\begin{align} \begin{bmatrix} 1 & \rho_1^{\lin} & \rho_1^{\spec} \\ \rho_1^{\lin} & 1 & \rho_1^{\lin}\rho_1^{\spec} + \expt{W^{\lin}W_1^{\spec}} \\ \rho_1^{\spec} & \rho_1^{\lin}\rho_1^{\spec} + \expt{W^{\lin}W_1^{\spec}} & 1 \end{bmatrix} . \notag \end{align}\] Let \(\nu_1 = \rho_1^{\lin} \rho_1^{\spec} + \expt{W^{\lin} W_1^{\spec}}\). Using the covariance structure above, we obtain that \(X_1\) conditioned on \((X^{\lin}, X_1^{\spec})\) is a Gaussian random variable with mean \[\begin{align} \wt{\mu} &\mathrel{\vcenter{:}}= \begin{bmatrix} \rho_1^{\lin} & \rho_1^{\spec} \end{bmatrix} \begin{bmatrix} 1 & \nu_1 \\ \nu_1 & 1 \end{bmatrix}^{-1} \begin{bmatrix} X^{\lin} \\ X_1^{\spec} \end{bmatrix} \notag \end{align}\] and variance \[\begin{align} \wt{\sigma}^2 &\mathrel{\vcenter{:}}= 1 - \begin{bmatrix} \rho_1^{\lin} & \rho_1^{\spec} \end{bmatrix} \begin{bmatrix} 1 & \nu_1 \\ \nu_1 & 1 \end{bmatrix}^{-1} \begin{bmatrix} \rho_1^{\lin} \\ \rho_1^{\spec} \end{bmatrix} . \notag \end{align}\] Therefore, the Bayes-optimal combined estimator is given by \[\begin{align} \expt{X_1\condon X^{\lin}, X_1^{\spec}} &= \wt{\mu} = \frac{1}{1 - \nu_1^2} \sqrbrkt{\paren{\rho_1^{\lin} - \rho_1^{\spec}\nu_1} X^{\lin} + \paren{\rho_1^{\spec} - \rho_1^{\lin}\nu_1} X_1^{\spec}} , \notag \end{align}\] which agrees with the expression in ?? . Finally, the explicit formulas of the overlaps given by the Bayes-optimal combined estimators can be obtained from [eqn:bayes-opt-bound,eqn:opt-combinator-1] via elementary algebraic manipulations.

12 Additional proofs for linear estimator (proof of [lem:linear-optimal-overlap])↩︎

With the characterization of the limiting overlaps of the linear estimator in 2, we can maximize them over the choice of the preprocessing function \(\cL\colon\bbR\to\bbR\). For \(i\ge0\), let \[\begin{align} m_i(y) &\mathrel{\vcenter{:}}= \expt{G^i p(y|G)} , \label{eqn:def-mi} \end{align}\tag{156}\] where \(G\sim\cN(0,1)\) and \(p(y|g)\) is defined in 5 . Using \(m_0,m_1\), the squared limiting overlap between \(\xlin\) and \(x_1^*\) in ?? can be expressed in the following way: \[\begin{gather} \frac{\alpha^2 \expt{G\cL(Y)}^2}{(\alpha^2+(1-\alpha)^2) \expt{G\cL(Y)}^2 + \expt{\cL(Y)^2}/\delta} = \paren{\frac{\alpha^2+(1-\alpha)^2}{\alpha^2} + \frac{1}{\alpha^2\delta} \cdot \frac{\expt{\cL(Y)^2}}{\expt{G\cL(Y)}^2}}^{-1} \\ = \paren{\frac{\alpha^2+(1-\alpha)^2}{\alpha^2} + \frac{1}{\alpha^2\delta} \cdot \frac{\int_{\supp(Y)} m_0(y) \cL(y)^2 \diff y}{\paren{\int_{\supp(Y)} m_1(y) \cL(y) \diff y}^2}}^{-1} , \label{eqn:overlap-using-m0m1} \end{gather}\tag{157}\] provided \(\int_{\supp(Y)} m_1(y) \cL(y) \diff y \ne0\) and \(\expt{|G\cL(Y)|}<\infty\). The optimization of overlap can be formalized as the following maximization problem: \[\begin{align} \paren{\ollinone}^2 &\mathrel{\vcenter{:}}= \sup_{\cL\colon\bbR\to\bbR} \paren{\frac{\alpha^2+(1-\alpha)^2}{\alpha^2} + \frac{1}{\alpha^2\delta} \cdot \frac{\int_{\supp(Y)} m_0(y) \cL(y)^2 \diff y}{\paren{\int_{\supp(Y)} m_1(y) \cL(y) \diff y}^2}}^{-1} \notag \\ &\phantom{\mathrel{\vcenter{:}}=} \suchthat \int_{\supp(Y)} m_1(y) \cL(y) \diff y \ne0 \tag{158} \\ &\phantom{\mathrel{\vcenter{:}}= \suchthat} \expt{|G\cL(Y)|}<\infty . \tag{159} \end{align}\] Therefore, maximizing \(\ollinone\) is equivalent to solving the following minimization problem: \[\begin{align} \optlin &\mathrel{\vcenter{:}}= \inf_{\cL\colon\bbR\to\bbR} \frac{\int_{\supp(Y)} m_0(y) \cL(y)^2 \diff y}{\paren{\int_{\supp(Y)} m_1(y) \cL(y) \diff y}^2} \quad \suchthat \quad \text{\Cref{eqn:ol-lin-constraint-1,eqn:ol-lin-constraint-2}} . \notag \end{align}\] This optimization problem has been studied in [25]. In particular, under the condition \[\begin{align} \int_{\supp(Y)} \frac{m_1(y)^2}{m_0(y)} \diff y \in(0,\infty) \label{eqn:lin-regularity-condition} \end{align}\tag{160}\] (which is equivalent to 22 ), we have \(\optlin = \paren{\int_{\supp(Y)} \frac{m_1(y)^2}{m_0(y)} \diff y}^{-1}\), attained by \(\cL^*\colon\bbR\to\bbR\) defined as (cf.[25]) \[\begin{align} \cL^*(y) = \frac{m_1(y)}{m_0(y)} , \label{eqn:opt-lin-preprocessor} \end{align}\tag{161}\] which satisfies \(\int_{\supp(Y)} m_1(y) \cL^*(y) \diff y > 0\) and \(\expt{|G\cL^*(Y)|}<\infty\).

Therefore, the value of the original problem \(\ollinone\) we are interested in is given by \[\begin{align} \paren{\ollinone}^2 &= \paren{\frac{\alpha^2+(1-\alpha)^2}{\alpha^2} + \frac{1}{\alpha^2\delta} \cdot \frac{1}{\int_{\supp(Y)} \frac{m_1(y)^2}{m_0(y)} \diff y}}^{-1} . \label{eqn:opt-overlp-lin-1} \end{align}\tag{162}\] Analogously, the optimal (over the choice of \(\cL\colon\bbR\to\bbR\)) limiting overlap between \(\xlin\) and \(x_2^*\) equals \[\begin{align} \paren{\ollintwo}^2 &= \paren{\frac{\alpha^2+(1-\alpha)^2}{(1-\alpha)^2} + \frac{1}{(1-\alpha)^2\delta} \cdot \frac{1}{\int_{\supp(Y)} \frac{m_1(y)^2}{m_0(y)} \diff y}}^{-1} , \label{eqn:opt-overlp-lin-2} \end{align}\tag{163}\] which is also achieved by \(\cL^*\colon\bbR\to\bbR\) defined in 161 .

13 Additional proofs for spectral estimator ([thm:opt-spec])↩︎

13.1 Optimization of spectral threshold↩︎

Let us consider weak recovery of \(x_1^*\). (The analysis for the recovery of \(x_2^*\) is completely analogous and is therefore omitted.) For a given preprocessing function \(\cT\colon\bbR\to\bbR\), we know from 3 and [rk:cnvan] that weak recovery of \(x_1^*\) is possible (i.e., \(\rho_1^{\spec} >0\)) when \(\lambda^*(\delta_1) > \ol\lambda(\delta)\). This condition is equivalent to \(\phi(\ol\lambda(\delta)) > \zeta(\ol\lambda(\delta); \delta_1) = \psi(\ol\lambda(\delta); \delta_1)\), or more explicitly, \[\begin{align} \expt{\frac{Z}{\ol\lambda(\delta) - Z}(G^2 - 1)} &> \frac{1}{\delta_1} = \frac{1}{\alpha\delta}. \label{eqn:optthr-cond1} \end{align}\tag{164}\] Here we recall that \(Z=\cT(Y)\), and \(\ol\lambda(\delta)\) satisfies \(\psi'(\ol\lambda(\delta);\delta) = 0\), or equivalently (see 5), it is the solution to \[\begin{align} \expt{\paren{\frac{Z}{\ol\lambda(\delta) - Z}}^2} &= \frac{1}{\delta} . \label{eqn:optthr-cond2} \end{align}\tag{165}\] 164 assumes that \(\ol\lambda(\delta)>0\), which is satisfied as \(\ol\lambda(\delta)>\sup\supp(\cT(Y))\) by definition (cf.@eq:eqn:def-lam-bar ) and the RHS is strictly positive by [itm:assump-preproc-spec]. Due to scaling invariance, we claim that \(\ol\lambda(\delta)\) can be assumed to be \(1\). Indeed, both [eqn:optthr-cond1,eqn:optthr-cond2] depend on \(\cT\colon\bbR\to\bbR\) only through the ratio \(\frac{\cT(Y)}{\ol\lambda(\delta) - \cT(Y)}\), therefore any given \(\cT\) and the corresponding \(\ol\lambda(\delta)\) derived via 165 can be replaced with \(\cT/\ol\lambda(\delta)\) and \(1\), respectively. [eqn:optthr-cond1,eqn:optthr-cond2] then become \[\begin{align} \expt{\frac{Z}{1 - Z}(G^2 - 1)} > \frac{1}{\alpha\delta} , \qquad \expt{\paren{\frac{Z}{1 - Z}}^2} = \frac{1}{\delta} .\label{eqn:optthr-cond1-2} \end{align}\tag{166}\] For convenience, let \(\Gamma(y) \mathrel{\vcenter{:}}= \frac{\cT(y)}{1 - \cT(y)}\). The expectations in 166 can then be written as \[\begin{align} \expt{\paren{\frac{Z}{1 - Z}}^2} = \expt{\paren{\frac{\cT(Y)}{1 - \cT(Y)}}^2} = \expt{\Gamma(Y)^2} &= \int_{\supp(Y)} \expt{p(y|G)}\cdot\Gamma(y)^2 \diff y , \label{eqn:expand1} \end{align}\tag{167}\] where \(G \sim \cN(0,1)\). Similarly \[\begin{align} \expt{\frac{Z}{1 - Z}(G^2 - 1)} &= \int_{\supp(Y)} \expt{p(y|G)(G^2 - 1)}\cdot\Gamma(y) \diff y . \label{eqn:expand2} \end{align}\tag{168}\] So the conditions in 166 can be further written as \[\begin{align} \int_{\supp(Y)} \expt{p(y|G)}\cdot\Gamma(y)^2 \diff y &= \frac{1}{\delta} , \label{eqn:optthr-cond1-3} \end{align}\tag{169}\] and \[\begin{align} \int_{\supp(Y)} \expt{p(y|G)(G^2 - 1)}\cdot\Gamma(y) \diff y &> \frac{1}{\alpha\delta} . \label{eqn:optthr-cond2-3} \end{align}\tag{170}\] In the above form, one can apply Cauchy–Schwarz inequality to 170 given the equality condition in 169 . Hence, we can upper bound the LHS of 170 as \[\begin{align} & \int_{\supp(Y)} \expt{p(y|G)(G^2 - 1)}\cdot \Gamma(y) \diff y \notag \\ &= \int_{\supp(Y)} \sqrt{\expt{p(y|G)}}\cdot\Gamma(y) \cdot \frac{\expt{p(y|G)(G^2 - 1)}}{\sqrt{\expt{p(y|G)}}} \diff y \notag \\ & \le \sqrt{\int_{\supp(Y)} \expt{p(y|G)}\cdot\Gamma(y)^2 \diff y} \cdot\sqrt{\int_{\supp(Y)} \frac{\expt{p(y|G)(G^2 - 1)}^2}{\expt{p(y|G)}} \diff y} \notag \\ & = \frac{1}{\sqrt{\delta}} \cdot \frac{1}{\alpha\sqrt{\delta^*_1}}, \label{eqn:optthr-cond3-3} \end{align}\tag{171}\] where the last equality is obtained using 169 and by defining \[\begin{align} \delta^*_1 &\mathrel{\vcenter{:}}= \frac{1}{\alpha^2\int_{\supp(Y)} \frac{\expt{p(y|G)(G^2 - 1)}^2}{\expt{p(y|G)}} \diff y} \label{eqn:def-delta-1-star} \end{align}\tag{172}\] Combining [eqn:optthr-cond2-3,eqn:optthr-cond3-3], we obtain the condition \(\delta > \delta^*_1\).

So far we have shown that if \(\lambda^*(\delta_1) > \ol\lambda(\delta)\) under a preprocessing function \(\cT\colon\bbR\to\bbR\), then it must be the case that \(\delta > \delta^*_1\). In what follows, we show that whenever \(\delta>\delta^*_1\), one can find a preprocessing function \(\wt\cT_1^*\colon\bbR\to\bbR\) that achieves equality in both [eqn:optthr-cond1-3,eqn:optthr-cond3-3], guaranteeing that \(\lambda^*(\delta_1) > \ol\lambda(\delta)\). Consider any \(\delta > \delta^*_1\). Our goal is to find a function \(\wt\cT^*_1\colon\bbR\to\bbR\) satisfying \[\begin{align} \expt{\paren{\frac{\wt\cT_1^*(Y)}{1 - \wt\cT_1^*(Y)}}^2} &= \frac{1}{\delta} , \label{eqn:find-optthr-1} \end{align}\tag{173}\] and \[\begin{align} \expt{\frac{\wt\cT_1^*(Y)}{1 - \wt\cT^*_1(Y)}(G^2 - 1)} = \frac{1}{\alpha\sqrt{\delta}\sqrt{\delta^*_1}} > \frac{1}{\alpha\delta}. \label{eqn:find-optthr-2-strong} \end{align}\tag{174}\] Note that according to 173 , the corresponding \(\ol\lambda(\delta)\) associated with \(\wt\cT^*_1\) (to be constructed) is equal to \(1\). Constructing such a \(\wt\cT_1^*\) is equivalent to constructing a function \(\cT_1^*\), which we define via \[\begin{align} \frac{\wt\cT_1^*(y)}{1 - \wt\cT_1^*(y)} &= \sqrt{\frac{\delta^*_1}{\delta}}\cdot\frac{\cT_1^*(y)}{1 - \cT_1^*(y)} \label{eqn:tilde-t1star} \end{align}\tag{175}\] for every \(y\in\bbR\). Now [eqn:find-optthr-1,eqn:find-optthr-2-strong] are equivalent to \[\begin{align} \expt{\paren{\frac{\cT_1^*(Y)}{1 - \cT_1^*(Y)}}^2} &= \frac{1}{\delta^*_1} , \qquad \expt{\frac{\cT_1^*(Y)}{1 - \cT_1^*(Y)}(G^2 - 1)} = \frac{1}{\alpha\delta^*_1} .\label{eqn:find-optthr-1-2} \end{align}\tag{176}\] Before proceeding, let us define the following functions for convenience: \[\begin{align} \Gamma_1^*(y) &\mathrel{\vcenter{:}}= \frac{\cT_1^*(y)}{1 - \cT_1^*(y)} , \qquad m_0(y) \mathrel{\vcenter{:}}= \expt{p(y|G)}, \qquad m_2(y) \mathrel{\vcenter{:}}= \expt{p(y|G)G^2}. \label{eqn:gamma1star} \end{align}\tag{177}\] With the above notation, \(\delta^*_1\) in 172 can be written as \[\begin{align} \delta^*_1 &= \frac{1}{\alpha^2\int_{\supp(Y)} \frac{(m_2(y) - m_0(y))^2}{m_0(y)} \diff y} , \label{eqn:def-delta1-star-pf} \end{align}\tag{178}\] and the LHSs of 176 can be written as \[\begin{align} \expt{\paren{\frac{\cT_1^*(Y)}{1 - \cT_1^*(Y)}}^2} &= \int_{\supp(Y)} \Gamma_1^*(y)^2 m_0(y) \diff y \notag \end{align}\] and \[\begin{align} \expt{\frac{\cT_1^*(Y)}{1 - \cT_1^*(Y)}(G^2 - 1)} &= \int_{\supp(Y)} \Gamma_1^*(y) (m_2(y) - m_0(y)) \diff y \notag \end{align}\] respectively, using [eqn:expand1,eqn:expand2]. Then, 176 becomes \[\begin{align} \int_{\supp(Y)} \Gamma_1^*(y)^2 m_0(y) \diff y &= \alpha^2\int_{\supp(Y)} \frac{(m_2(y) - m_0(y))^2}{m_0(y)} \diff y \label{eqn:find-optthr-1-3} \end{align}\tag{179}\] and \[\begin{align} \int_{\supp(Y)} \Gamma_1^*(y) (m_2(y) - m_0(y)) \diff y &= \alpha\int_{\supp(Y)} \frac{(m_2(y) - m_0(y))^2}{m_0(y)} \diff y . \label{eqn:find-optthr-2-3} \end{align}\tag{180}\] By inspecting [eqn:find-optthr-1-3,eqn:find-optthr-2-3], we conclude that the following choice of \(\Gamma_1^*\colon\bbR\to\bbR\) meets both conditions: \[\begin{align} \Gamma_1^*(y) &\mathrel{\vcenter{:}}= \alpha\cdot\frac{m_2(y) - m_0(y)}{m_0(y)} . \notag \end{align}\] By 177 , this gives a choice of \(\cT_1^*\): \[\begin{align} \cT_1^*(y) &= \frac{\Gamma_1^*(y)}{1 + \Gamma_1^*(y)} = \frac{\alpha\cdot\frac{m_2(y) - m_0(y)}{m_0(y)}}{1 + \alpha\cdot\frac{m_2(y) - m_0(y)}{m_0(y)}} = 1 - \frac{1}{\alpha\cdot\frac{m_2(y)}{m_0(y)} + (1-\alpha)} . \label{eqn:opt-spec-T-thr} \end{align}\tag{181}\] Using the relation in 175 , we can then determine \(\wt\cT_1^*(y)\): \[\begin{align} \wt\cT_1^*(y) &= 1 - \frac{1}{1 + \sqrt{\frac{\delta^*_1}{\delta}}\cdot\frac{\cT_1^*(y)}{1 - \cT_1^*(y)}} = 1 - \frac{1}{\sqrt{\frac{\delta^*_1}{\delta}} \paren{\alpha\cdot\frac{m_2(y)}{m_0(y)} + (1-\alpha)} + \paren{1 - \sqrt{\frac{\delta^*_1}{\delta}}\cdot}} . \label{eqn:opt-spec-tilde-T} \end{align}\tag{182}\]

We have found a candidate function \(\wt\cT_1^*\colon\bbR\to\bbR\) that satisfies both [eqn:find-optthr-1,eqn:find-optthr-2-strong]. It remains to verify that this function meets [itm:assump-preproc-spec]. In 13.2 (see [eqn:inf-lb,eqn:sup-lb,eqn:sup-ub]), we will show that \(\cT_1^*\) satisfies \[\begin{align} 0 < \sup_{y\in\supp(Y)} \cT_1^*(y) < 1 , \quad \inf_{y\in\supp(Y)} \cT_1^*(y) > -\infty . \label{eqn:T-prop-to-prove} \end{align}\tag{183}\] Since \(\cT_1^*(y) < 1\) for every \(y\in\supp(Y)\), writing \[\begin{align} \wt{\cT}_1^*(y) &= 1 - \frac{1}{\frac{\sqrt{\delta_1^*/\delta}}{1 - \cT_1^*(y)} + 1 - \sqrt{\delta_1^*/\delta}} , \notag \end{align}\] we see that \(\wt{\cT}_1^*(y)\) increases as \(\cT_1^*(y)\) increases. Therefore, 183 implies \[\begin{align} 0 < \sup_{y\in\supp(Y)} \wt{\cT}_1^*(y) < 1 , \quad \inf_{y\in\supp(Y)} \wt{\cT}_1^*(y) > 1 - \frac{1}{1 - \sqrt{\delta_1^*/\delta}} , \end{align}\] which certifies that \(\wt{\cT}_1^*\) satisfies [itm:assump-preproc-spec].

13.2 Optimization of spectral overlap↩︎

We will now identify the optimal preprocessing function that achieves the largest asymptotic overlap with \(x_1^*\) whenever \(\lambda^*(\delta_1) > \ol\lambda(\delta)\). We again only focus on \(x_1^*\), as the argument for \(x_2^*\) is analogous. Recall from 3 that the squared overlap induced by a generic preprocessing function converges almost surely to \[\begin{align} \lim_{d\to\infty} \frac{\inprod{v_1(D)}{x_1^*}^2}{\normtwo{v_1(D)^2}^2 \normtwo{x_1^*}^2} & = \frac{1}{\alpha}\cdot\frac{\frac{1}{\delta} - \expt{\paren{\frac{Z}{\lambda^*(\delta_1) - Z}}^2}}{\frac{1}{\delta_1} + \expt{\paren{\frac{Z}{\lambda^*(\delta_1) - Z}}^2(G^2 - 1)}}= \frac{1}{\alpha}\cdot\frac{\psi'(\lambda^*(\delta_1);\delta)}{\psi'(\lambda^*(\delta_1);\delta_1) - \phi'(\lambda^*(\delta_1))}, \notag \end{align}\] where the second equality readily follows after some manipulations. Let \(\cF\) denote the set of feasible preprocessing functions: \[\begin{align} \cF &\mathrel{\vcenter{:}}= \curbrkt{\cT\colon\bbR\to\bbR : \inf_{y\in\supp(Y)} \cT(y) > -\infty , \; 0 < \sup_{y\in\supp(Y)} \cT(y) < \infty , \; \prob{\cT(Y) = 0} < 1 } , \label{eqn:feasible-t} \end{align}\tag{184}\] where \(Y = q(G,\eps)\) and \((G,\eps)\sim\cN(0,1)\otimes P_\eps\). The variational problem we would like to solve is \[\begin{align} \mathsf{OL}_1^2 &= \sup_{\cT_1\in\cF} \, \frac{1}{\alpha}\cdot\frac{\psi'(\lambda^*(\delta_1);\delta)}{\psi'(\lambda^*(\delta_1);\delta_1) - \phi'(\lambda^*(\delta_1))} \notag \\ & \quad\suchthat\quad \zeta(\lambda^*(\delta_1); \delta_1) = \phi(\lambda^*(\delta_1)) \notag \\ &= \frac{1}{\alpha} \sup_{\cT_1\in\cF} \frac{\psi'(\lambda^*(\delta_1);\delta)}{\psi'(\lambda^*(\delta_1);\delta_1) - \phi'(\lambda^*(\delta_1))} \notag \\ & \phantom{=}\suchthat\quad \psi(\lambda^*(\delta_1);\delta_1) = \phi(\lambda^*(\delta_1)), \quad \psi'(\lambda^*(\delta_1);\delta) > 0 \label{eqn:cond-opt-overlap-for-rk} \end{align}\tag{185}\] where \(\mathsf{OL}_1^2\) is the squared overlap for the first signal. The first constraint in 185 is by the definition of \(\lambda^*(\delta_1)\) and the second one is an equivalent characterization of the spectral threshold. We claim that \(\lambda^*(\delta_1)\) can be assumed without loss of generality to be \(1\). To see this, according to the explicit formulas for \(\phi'(\cdot)\), \(\psi'(\cdot; \delta_1)\), and \(\lambda^*(\delta_1)\) (see [eqn:phi-deriv,eqn:derivative-psi,eqn:def-lami-star-supcrit-explicit]), we note that both the objective and constraints of the optimization problem \(\mathsf{OL}_1^2\) depend on \(\cT_1\colon\bbR\to\bbR\) only through the ratio \(\frac{\cT_1(Y)}{\lambda^*(\delta_1) - \cT_1(Y)}\). Therefore, any \(\cT_1\) and the corresponding \(\lambda^*(\delta_1)\) computed via 222 can be replaced without affecting anything with \(\cT_1/\lambda^*(\delta_1)\) and \(1\), respectively. Recall that \(\Gamma_1\colon\bbR\to\bbR\) is defined as \[\begin{align} \Gamma_1(y) &\mathrel{\vcenter{:}}= \frac{\cT_1(y)}{1 - \cT_1(y)} = \frac{1}{1 - \cT_1(y)} - 1 . \label{eqn:gamma-t} \end{align}\tag{186}\] Since \(\lambda^*(\delta_1) = 1\) and \(\lambda^*(\delta_1)\) is the solution to ?? in \((\sup\supp(Z),\infty)\), we need an additional constraint on \(\cT_1\): \[\begin{align} 1 > \sup\supp(\cT_1(Y)) = \sup_{y\in\supp(Y)} \cT_1(y) , \label{eqn:sup-t1-ub} \end{align}\tag{187}\] which, in light of 186 , translates to the following constraint on \(\Gamma_1\): \[\begin{align} \sup_{y\in\supp(Y)} \Gamma_1(y) &> -1 . \label{eqn:gamma-cstr-3} \end{align}\tag{188}\] Let \(\cG\) be the feasible set of \(\Gamma_1\colon\bbR\to\bbR\). To give a precise definition of \(\cG\), we now translate the constraints on \(\cT_1\) to constraints on \(\Gamma_1\). According to 187 and the second equality in 186 , \(\Gamma_1(y)\) increases as \(\cT_1(y)\) increases in \((-\infty,1)\). Therefore, the constraints \(\sup\limits_{y\in\supp(Y)} \cT(y) \in (0,\infty)\) and \(\inf\limits_{y\in\supp(Y)} \cT_1(y) > -\infty\) translate to \[\begin{align} \sup_{y\in\supp(Y)} \Gamma_1(y) &\in (0,\infty), \qquad \inf_{y\in\supp(Y)} \Gamma_1(y) > -1 .\label{eqn:gamma-cstr-1} \end{align}\tag{189}\] [eqn:gamma-cstr-1,eqn:gamma-cstr-3] yield the following definition of \(\cG\): \[\begin{align} \cG &\mathrel{\vcenter{:}}= \curbrkt{\Gamma\colon\bbR\to\bbR : \inf_{y\in\supp(Y)} \Gamma(y) > -1 , \; 0 < \sup_{y\in\supp(Y)} \Gamma(y) < \infty , \; \prob{\Gamma(Y) = 0} < 1 } . \label{eqn:def-feasible-set-gamma} \end{align}\tag{190}\]

Using the explicit representations of various functionals and variables, we can then write the optimization problem \(\mathsf{OL}_1^2\) as \[\begin{align} \mathsf{OL}_1^2 &= \sup_{\cT_1\in\cF} \frac{\frac{1}{\delta} - \expt{\paren{\frac{Z}{\lambda^*(\delta_1) - Z}}^2}}{\frac{1}{\delta} + \alpha \expt{\paren{\frac{Z}{\lambda^*(\delta_1) - Z}}^2(G^2 - 1)}} \notag \\ &= \sup_{\Gamma_1\in\cG} \paren{\frac{\frac{1}{\delta} + \alpha\int_{\supp(Y)} \Gamma_1(y)^2 m_2(y) \diff y - \alpha\int_{\supp(Y)} \Gamma_1(y)^2 m_0(y) \diff y}{\frac{1}{\delta} - \int_{\supp(Y)}\Gamma_1(y)^2 m_0(y) \diff y}}^{-1} \notag \\ &= \sup_{\Gamma_1\in\cG} \paren{\frac{\frac{1-\alpha}{\delta} + \alpha \int_{\supp(Y)} \Gamma_1(y)^2 m_2(y) \diff y}{\frac{1}{\delta} - \int_{\supp(Y)} \Gamma_1(y)^2 m_0(y) \diff y} + \alpha}^{-1} , \notag \end{align}\] subject to the conditions \[\begin{align} \expt{\frac{Z}{\lambda^*(\delta_1) - Z}(G^2 - 1)} &= \frac{1}{\alpha\delta} , \qquad \frac{1}{\delta} > \expt{\paren{\frac{Z}{\lambda^*(\delta_1) - Z}}^2} , \label{eqn:two-cond} \end{align}\tag{191}\] which can be alternatively written as follows using the notation in 13.1: \[\begin{align} \frac{1}{\alpha\delta} &= \int_{\supp(Y)} \Gamma_1(y) (m_2(y)-m_0(y)) \diff y , \qquad \frac{1}{\delta} > \int_{\supp(Y)} \Gamma_1(y)^2 m_0(y) \diff y . \label{eqn:cond-opt-overlap} \end{align}\tag{192}\] Note that in the first identity in 191 , we use \(\lambda^*(\delta_1)\ne0\) which holds since \(\lambda^*(\delta_1)>\sup\supp(Z)>0\) by [itm:assump-preproc-spec]. We observe that to solve \(\mathsf{OL}_1^2\), it suffices to solve the following minimization problem: \[\begin{align} \mathsf{OPT}_1 &\mathrel{\vcenter{:}}= \inf_{\Gamma_1\in\cG} \, \paren{\frac{\frac{1-\alpha}{\delta} + \alpha \int_{\supp(Y)} \Gamma_1(y)^2 m_2(y) \diff y}{\frac{1}{\delta} - \int_{\supp(Y)} \Gamma_1(y)^2 m_0(y) \diff y}} \notag \\ &\phantom{\mathrel{\vcenter{:}}=} \suchthat\quad \text{\Cref{eqn:cond-opt-overlap}} . \notag \end{align}\] Recall the standard fact that the minimum value of a function (subject to constraints) is the smallest level \(\beta\) such that the \(\beta\)-level set (after taking the intersection with the constraint set) is non-empty. For any given \(\beta>0\), define \(\cL_\beta\) as the set of \(\Gamma_1\in\cG\) which induces an objective value at most \(\beta\) and satisfies all conditions of \(\mathsf{OPT}_1\): \[\begin{align} \cL_\beta &\mathrel{\vcenter{:}}= \curbrkt{\Gamma_1\in\cG : \begin{array}{l} \frac{\frac{1-\alpha}{\delta} + \alpha \int_{\supp(Y)} \Gamma_1(y)^2 m_2(y) \diff y}{\frac{1}{\delta} - \int_{\supp(Y)} \Gamma_1(y)^2 m_0(y) \diff y} \le \beta \\ \frac{1}{\alpha\delta} = \int_{\supp(Y)} \Gamma_1(y) (m_2(y)-m_0(y)) \diff y \\ \frac{1}{\delta} > \int_{\supp(Y)} \Gamma_1(y)^2 m_0(y) \diff y \end{array}} . \notag \end{align}\] The first constraint describes the \(\beta\)-level set of the original objective of \(\mathsf{OPT}_1\). Then \(\mathsf{OPT}_1\) can be further written as \[\begin{align} \mathsf{OPT}_1 &\mathrel{\vcenter{:}}= \inf\curbrkt{\beta>0 : \cL_\beta \ne\emptyset} . \label{eqn:beta-opt1} \end{align}\tag{193}\]

The first constraint in \(\cL_\beta\) is equivalent to: \[\begin{align} \frac{1-\alpha}{\delta} + \alpha\int_{\supp(Y)} \Gamma_1(y)^2 m_2(y) \diff y &\le \frac{\beta}{\delta} - \beta \int_{\supp(Y)} \Gamma_1(y)^2 m_0(y) \diff y, \notag \end{align}\] or \[\begin{align} \int_{\supp(Y)} \Gamma_1(y)^2 (\alpha m_2(y) + \beta m_0(y)) \diff y &\le \frac{\beta - (1 - \alpha)}{\delta}, \label{eqn:before-divide-beta} \end{align}\tag{194}\] since the third condition guarantees that the denominator of the LHS of the first inequality is positive. We also claim that \(\beta-(1-\alpha)>0\) and divide both sides by it to obtain \[\begin{align} \int_{\supp(Y)} \Gamma_1(y)^2 \frac{\alpha m_2(y) + \beta m_0(y)}{\beta - (1 - \alpha)} \diff y &\le \frac{1}{\delta} . \label{eqn:levelset-cond1-equiv} \end{align}\tag{195}\] To see why \(\beta-(1-\alpha)>0\), recall that \(\beta\) is a possible value of \(\mathsf{OPT}_1\) which in turn has the following relation to \(\mathsf{OL}_1^2\): \[\begin{align} \mathsf{OL}_1^2 &= \frac{1}{\mathsf{OPT}_1 + \alpha} . \label{eqn:ol-opt} \end{align}\tag{196}\] Since \(0\le \mathsf{OL}_1^2\le 1\), this implies \(\beta\ge1-\alpha\). Furthermore, if \(\beta = 1-\alpha\), 194 implies that \(\Gamma_1(y) = 0\) for almost every \(y\in\supp(Y)\). However, since \(\Gamma_1\in\cG\), this cannot happen according to the third condition in the definition of \(\cG\) (cf.@eq:eqn:def-feasible-set-gamma ). Therefore \(\beta>1 - \alpha\).

The following observation can further simplify the description of the set \(\cL_\beta\). The third (inequality) constraint can be dropped since \(m_2(y) , m_0(y)\ge0\) and \[\begin{align} \int_{\supp(Y)} \Gamma_1(y)^2 \frac{\alpha m_2(y) + \beta m_0(y)}{\beta - (1 - \alpha)} \diff y \ge \int_{\supp(Y)} \Gamma_1(y)^2 m_0(y) \diff y . \notag \end{align}\] Therefore the third constraint has already been guaranteed to be true given the first (inequality) constraint which is, as we just argued, equivalent to 195 .

Now, with the above observations, we arrive at the following equivalent description of the \(\beta\)-level set \(\cL_\beta\): \[\begin{align} \cL_\beta &= \curbrkt{\Gamma_1\in\cG : \begin{array}{l} \int_{\supp(Y)} \Gamma_1(y)^2 \frac{\alpha m_2(y) + \beta m_0(y)}{\beta - (1 - \alpha)} \diff y \le \frac{1}{\delta} \\ \frac{1}{\alpha\delta} = \int_{\supp(Y)} \Gamma_1(y) (m_2(y)-m_0(y)) \diff y \end{array}} . \notag \end{align}\] We observe that for any fixed \(\beta>0\), \(\cL_\beta\ne\emptyset\) if and only if \(\mathsf{INF}_1^{(\beta)} \le \frac{1}{\delta}\) where \(\mathsf{INF}_1^{(\beta)}\) is the value of the following constrained minimization problem: \[\begin{align} \mathsf{INF}_1^{(\beta)} &\mathrel{\vcenter{:}}= \inf_{\Gamma_1\in\cG} \int_{\supp(Y)} \Gamma_1(y)^2 \frac{\alpha m_2(y) + \beta m_0(y)}{\beta - (1 - \alpha)} \diff y \notag \\ &\phantom{\mathrel{\vcenter{:}}=} \suchthat\quad \frac{1}{\alpha\delta} = \int_{\supp(Y)} \Gamma_1(y) (m_2(y)-m_0(y)) \diff y . \notag \end{align}\]

We turn to compute the value of \(\mathsf{INF}_1^{(\beta)}\). Though \(\mathsf{INF}_1^{(\beta)}\) appears as a constrained variational problem over function space, one can obtain its optimum (and the corresponding optimizer) from a different perspective by casting it as a linear program over a certain Hilbert space. Specifically, motivated by the form of the objective of \(\mathsf{INF}_1^{(\beta)}\), we define the following inner product on the function space: \[\begin{align} \inprod{f}{g}_{\beta, \alpha} &\mathrel{\vcenter{:}}= \int_{\supp(Y)} f(y)g(y) \frac{\alpha m_2(y) + \beta m_0(y)}{\beta - (1 - \alpha)} \diff y . \notag \end{align}\] This induces the norm \(\norm{\beta,\alpha}{f} = \sqrt{\inprod{f}{f}_{\beta,\alpha}}\) With the above notation, \(\mathsf{INF}_1^{(\beta)}\) can be written as \[\begin{align} \mathsf{INF}_1^{(\beta)} &= \inf_{\Gamma_1\in\cG} \norm{\beta,\alpha}{\Gamma_1}^2 \notag \\ &\phantom{=} \suchthat \quad \inprod{\Gamma_1}{\frac{m_2 - m_0}{\frac{\alpha m_2 + \beta m_0}{\beta - (1-\alpha)}}}_{\beta,\alpha} = \frac{1}{\alpha\delta} . \notag \end{align}\] In words, \(\mathsf{INF}_1^{(\beta)}\) outputs the smallest norm of \(\Gamma_1\) whose linear projection onto a given vector (viewed as an element in the defined Hilbert space) \(\frac{m_2 - m_0}{\frac{\alpha m_2 + \beta m_0}{\beta - (1-\alpha)}}\) is fixed. It is now geometrically clear that the minimizer \(\Gamma_1^{(\beta)}\) must be aligned with the vector \(\frac{m_2 - m_0}{\frac{\alpha m_2 + \beta m_0}{\beta - (1-\alpha)}}\). Therefore \(\Gamma_1^{(\beta)} = a^* \cdot \frac{m_2 - m_0}{\frac{\alpha m_2 + \beta m_0}{\beta - (1-\alpha)}}\), where the scalar \(a^*\in\bbR\) is uniquely determined from the equality condition: \[\begin{align} \int_{\supp(Y)} a^* \cdot \frac{m_2(y) - m_0(y)}{\frac{\alpha m_2(y) + \beta m_0(y)}{\beta - (1-\alpha)}} \cdot (m_2(y)-m_0(y)) \diff y = \frac{1}{\alpha\delta} , \notag \end{align}\] i.e., \[\begin{align} a^* &= \paren{\alpha\delta(\beta - (1-\alpha)) \int_{\supp(Y)} \frac{(m_2(y) - m_0(y))^2}{\alpha m_2(y) + \beta m_0(y)} \diff y}^{-1} . \notag \end{align}\] Let \[\begin{align} f_\alpha(\beta) &\mathrel{\vcenter{:}}= (\beta - (1-\alpha)) \int_{\supp(Y)} \frac{(m_2(y) - m_0(y))^2}{\alpha m_2(y) + \beta m_0(y)} \diff y . \notag \end{align}\] We have found the minimizer of \(\mathsf{INF}_1^{(\beta)}\): \[\begin{align} \Gamma_1^{(\beta)} &= \frac{1}{\alpha\delta f_\alpha(\beta)} (\beta - (1-\alpha)) \frac{m_2 - m_0}{\alpha m_2 + \beta m_0} , \notag \end{align}\] and the resulting minimum value is given by \[\begin{align} \mathsf{INF}_1^{(\beta)} &= \int_{\supp(Y)} \Gamma_1^{(\beta)}(y)^2 \frac{\alpha m_2(y) + \beta m_0(y)}{\beta - (1 - \alpha)} \diff y \notag \\ &= \int_{\supp(Y)} \paren{\frac{1}{\alpha\delta f_\alpha(\beta)} (\beta - (1-\alpha)) \frac{m_2(y) - m_0(y)}{\alpha m_2(y) + \beta m_0(y)}}^2 \frac{\alpha m_2(y) + \beta m_0(y)}{\beta - (1 - \alpha)} \diff y \notag \\ &= \frac{1}{(\alpha\delta f_\alpha(\beta))^2} \int_{\supp(Y)} (\beta - (1-\alpha)) \frac{(m_2(y) - m_0(y))^2}{\alpha m_2(y) + \beta m_0(y)} \diff y \, = \, \frac{1}{\alpha^2\delta^2f_\alpha(\beta)} . \notag \end{align}\] It follows that \[\begin{align} \cL_\beta\ne\emptyset & \quad \iff \quad \mathsf{INF}_1^{(\beta)} \le\frac{1}{\delta} \quad \iff \quad \frac{1}{\alpha^2\delta^2f_\alpha(\beta)} \le \frac{1}{\delta} \quad \iff \quad f_\alpha(\beta) \ge \frac{1}{\alpha^2\delta}. \notag \end{align}\] Recalling 193 , the value of \(\mathsf{OPT}_1\) is therefore equal to \[\begin{align} \mathsf{OPT}_1 &\mathrel{\vcenter{:}}= \inf\curbrkt{\beta>0 : f_\alpha(\beta) \ge \frac{1}{\alpha^2\delta}} . \notag \end{align}\] Writing \(f_\alpha\) as \[\begin{align} f_\alpha(\beta) &= (\beta - (1-\alpha)) \int_{\supp(Y)} \frac{(m_2(y) - m_0(y))^2}{\alpha m_2(y) + \beta m_0(y)} \diff y = \int_{\supp(Y)} \frac{(m_2(y) - m_0(y))^2}{\frac{\alpha m_2(y)}{\beta - (1-\alpha)} + \frac{m_0(y)}{1 - \frac{1-\alpha}{\beta}}} \diff y , \notag \end{align}\] we see that \(f_\alpha\) is increasing in \(\beta\). This implies that \(\mathsf{OPT}_1\) is equal to the critical \(\beta_1^*(\delta,\alpha)\) that solves the following equation \[\begin{align} f_\alpha(\beta_1^*(\delta,\alpha)) &= \frac{1}{\alpha^2 \delta} . \label{eqn:beta-fixed} \end{align}\tag{197}\]

Putting our findings together, we have shown that the value of \(\mathsf{OPT}_1\) equals \(\beta_1^*(\delta,\alpha)\) satisfying 197 , or more explicitly, \[\begin{align} (\beta_1^*(\delta,\alpha) - (1-\alpha)) \int_{\supp(Y)} \frac{(m_2(y) - m_0(y))^2}{\alpha m_2(y) + \beta_1^*(\delta,\alpha) m_0(y)} \diff y = \frac{1}{\alpha^2 \delta} , \label{eqn:opt-val-beta-star-fp} \end{align}\tag{198}\] and the corresponding optimizer is given by \[\begin{align} \Gamma_1^* = \Gamma_1^{(\beta_1^*(\delta,\alpha))} &= \frac{1}{\alpha\delta f_\alpha(\beta_1^*(\delta,\alpha))} (\beta_1^*(\delta,\alpha) - (1-\alpha)) \frac{m_2 - m_0}{\alpha m_2 + \beta_1^*(\delta,\alpha) m_0} \notag \\ &= \alpha (\beta_1^*(\delta,\alpha) - (1-\alpha)) \frac{m_2 - m_0}{\alpha m_2 + \beta_1^*(\delta,\alpha) m_0} , \label{eqn:opt-gamma} \end{align}\tag{199}\] where the second equality follows from 198 .

In light of the relation between \(\mathsf{OL}_1^2\) and \(\mathsf{OPT}_1\) (cf.@eq:eqn:ol-opt ) and the relation between \(\Gamma_1\) and \(\cT_1\) (cf.@eq:eqn:gamma-t ), it is straightforward to translate results in [eqn:opt-val-beta-star-fp,eqn:opt-gamma] to the original problem \(\mathsf{OL}_1^2\) we are interested in. Indeed, the value of \(\mathsf{OL}_1^2\) equals \[\begin{align} \mathsf{OL}_1^2 &= \frac{1}{\beta_1^*(\delta,\alpha) + \alpha} \label{eqn:opt-overlap-formula} \end{align}\tag{200}\] and is achieved by \[\begin{align} \cT_1^* &= \frac{\Gamma_1^*}{1 + \Gamma_1^*} = \frac{\alpha (\beta_1^*(\delta,\alpha) - (1-\alpha)) \frac{m_2 - m_0}{\alpha m_2 + \beta_1^*(\delta,\alpha) m_0}}{1 + \alpha (\beta_1^*(\delta,\alpha) - (1-\alpha)) \frac{m_2 - m_0}{\alpha m_2 + \beta_1^*(\delta,\alpha) m_0}} \notag \\ &= \frac{\alpha(\beta_1^*(\delta,\alpha) - (1-\alpha)) (m_2 - m_0)}{\alpha (\beta_1^*(\delta,\alpha) + \alpha) m_2 + (1-\alpha)(\beta_1^*(\delta,\alpha) + \alpha) m_0} . \notag \end{align}\] Recall that a multiplicative scaling of \(\cT_1^*\) does not change its performance in terms of spectral threshold and overlap. Therefore, we multiply the above expression by \(\frac{\beta_1^*(\delta,\alpha) + \alpha}{\beta_1^*(\delta,\alpha) - (1-\alpha)}\) and redefine \(\cT_1^*\) as \[\begin{align} \cT_1^* &= \frac{m_2 - m_0}{m_2 + \frac{1-\alpha}{\alpha} m_0} = 1 - \frac{1}{\alpha\cdot\frac{m_2}{m_0} + (1-\alpha)} . \label{eqn:opt-preproc-formula} \end{align}\tag{201}\] We observe that \(\cT_1^*\) in 201 above is the same as that in 181 obtained in 13.1. This is discussed in [rk:tilting] below.

Finally, we claim that the supremum in \(\mathsf{OL}_1^2\) can be achieved, by verifying that \(\cT_1^*\) meets [itm:assump-preproc-spec]. Indeed, letting \((G,\eps)\sim\cN(0,1)\otimes P_\eps\), we have \[\begin{align} \inf_{y\in\supp(\cT_1^*(q(G,\eps)))} \cT_1^*(y) &\ge 1 - \frac{1}{1-\alpha} = -\frac{\alpha}{1-\alpha} > -\infty , \label{eqn:inf-lb} \end{align}\tag{202}\] provided \(\alpha<1\). Also, it trivially holds that \[\begin{align} \sup_{y\in\supp(\cT_1^*(q(G,\eps)))} \cT_1^*(y) &< 1 . \label{eqn:sup-ub} \end{align}\tag{203}\] We then verify \(\prob{\cT_1^*(q(G,\eps)) = 0} < 1\). To this end, observe that \(m_2\) cannot be identically equal to \(m_0\), otherwise \(\delta_1^* = \delta_2^* = \infty\) (cf.@eq:eqn:def-delta1-star-pf ). Therefore \(m_2/m_0\) is not constantly \(1\) and \(\cT_1^*\) is not constantly \(0\). Finally, we verify that \[\begin{align} \sup_{y\in\supp(\cT_1^*(q(G,\eps)))} \frac{m_2(y)}{m_0(y)} &> 1 , \label{eqn:m2-larger-than95m0} \end{align}\tag{204}\] which implies \[\begin{align} \sup_{y\in\supp(\cT_1^*(q(G,\eps)))} \cT_1^*(y) &> 1 - \frac{1}{\alpha + (1-\alpha)} = 0 . \label{eqn:sup-lb} \end{align}\tag{205}\] 204 follows since \(m_2\not\equiv m_0\) and by 207 , \[\begin{align} \int_{\sup(\cT_1^*(q(G,\eps)))} (m_2(y) - m_0(y) )\diff y = 0 , \notag \end{align}\] hence there must exist \(y\in\sup(\cT_1^*(q(G,\eps)))\) such that \(m_2(y)>m_0(y)\).

As noted in the proofs, [eqn:opt-spec-T-thr,eqn:opt-preproc-formula] coincide. The first part of 13.1 exhibits a lower bound \(\delta_1^{\spec} \ge \delta_1^*\) (defined in 172 ), whereas the second part (from 173 onward) shows \(\delta_1^{\spec} \le \delta_1^*\). The upper bound can be alternatively obtained in 13.2 by substituting in the condition \(\psi'(\lambda^*(\delta_1);\delta)>0\) (cf.@eq:eqn:cond-opt-overlap-for-rk ) the function \(\cT_1^*\) and recognizing that the condition is equivalent to \(\delta > \delta_1^*\). This recognition is, however, not entirely obvious and we find it more transparent to directly derive the upper bound in the second part of 13.1. The price is that a slightly tilted version of \(\cT_1^*\) (see \(\wt\cT_1^*\) in 182 ) is constructed. The tilting is an artifact of the proof technique and we expect that \(\cT_1^*\) simultaneously minimizes the spectral threshold and maximizes the limiting overlap above the threshold.

14 Universal lower bound on spectral threshold (missing proof in [rk:univ-lb-spec-thr])↩︎

We show that for \(i \in \{1,2\}\), we have \(\delta^{\spec}_i \ge \frac{1}{2\alpha_i^2}\) for any mixed GLM. This follows from the upper bound \[\begin{align} \int_{\supp(Y)} \frac{(m_2(y) - m_0(y))^2}{m_0(y)} \diff y &\le 2,\label{eqn:upperbound-to-prove} \end{align}\tag{206}\] in view of 25 . Here the functions \(m_2\) and \(m_0\) are defined in 177 .

A similar argument has been made in [28] for the non-mixed complex generalized linear model where the output \(y_i\) depends on the linear measurement \(g_i = \inprod{a_i}{x^*}\) through its modulus: \(y_i \sim p(\cdot\mid |g_i|)\). Here we provide a similar argument for the real case without requiring the modulus. Recalling the definition of \(m_2\) and \(m_0\) (cf.@eq:eqn:gamma1star ), we have \[\begin{align} \int_{\supp(Y)} \frac{(m_2(y) - m_0(y))^2}{m_0(y)} \diff y &= \int_{\supp(Y)} \frac{m_2(y)^2}{m_0(y)} \diff y - 2 \int_{\supp(Y)} m_2(y) \diff y + \int_{\supp(Y)} m_0(y) \diff y \notag \\ &= \int_{\supp(Y)} \frac{m_2(y)^2}{m_0(y)} \diff y - 1 , \notag \end{align}\] since \[\begin{align} \begin{aligned} \int_{\supp(Y)} m_2(y) \diff y &= \int_{\supp(Y)} \expt{p(y|G)G^2} \diff y = \expt{G^2} = 1 , \\ \int_{\supp(Y)} m_0(y) \diff y &= \int_{\supp(Y)} \expt{p(y|G)} \diff y = 1 . \end{aligned} \label{eqn:m2-m0-int-1} \end{align}\tag{207}\] To show 206 , it then suffices to show \(\int \frac{m_2^2}{m_0}\le3\). Using the Cauchy–Schwarz inequality, we can bound \(m_2^2\) as follows: \[\begin{align} m_2(y)^2 = \paren{\int_\bbR f(g)p(y|g)g^2 \diff g}^2 & = \paren{\int_\bbR \sqrt{f(g)p(y|g)}g^2 \cdot \sqrt{f(g)p(y|g)} \diff g}^2 \notag \\ &\le \paren{\int_\bbR f(g)p(y|g) g^4 \diff g} \cdot \paren{\int_\bbR f(g)p(y|g) \diff g} \notag \\ &= \paren{\int_\bbR f(g)p(y|g) g^4 \diff g} \cdot m_0(y) , \notag \end{align}\] where \(f(g)\) denotes the standard Gaussian density function. The desired bound then follows: \[\begin{align} \int_{\supp(Y)} \frac{m_2(y)^2}{m_0(y)} \diff y &\le \int_{\supp(Y)}\int_\bbR f(g)p(y|g) g^4 \diff g\diff y \notag \\ &= \int_\bbR \sqrbrkt{f(g)\paren{\int_{\supp(Y)} p(y|g) \diff y}g^4} \diff g = \expt{G^4} = 3 . \notag \end{align}\]

15 State evolution of GAMP for mixed GLMs (proof of [prop:GAMP95SE])↩︎

Recall the GAMP iteration in 79 : \[\begin{align} \begin{aligned} u^t &= \frac{1}{\sqrt{\delta}} \Abar f_t(v^t; \, \xone, \xtwo) - \sfb_t \, g_{t-1}(u^{t-1};y), \qquad v^{t+1} = \frac{1}{\sqrt{\delta}} \Abar^\top g_{t}(u^t;y) - \sfc_t \, f_t(v^t;\, \xone, \xtwo), \end{aligned} \label{eq:gamp-eqn-det} \end{align}\tag{208}\] where the memory coefficients are \(\sfb_t =\frac{1}{n}\sum_{i=1}^d f_t'(v_i^t;\, \ol{x}_{1,i}^*, \ol{x}_{2,i}^*)\) and \(\sfc_t = \frac{1}{n}\sum_{i=1}^n g_t'(u_i^t; y_i)\). Using the initialization \(\tv^0\), the iteration starts with \(u^0= \frac{1}{\sqrt{\delta}}\Abar \tv^0\).

To prove the proposition, we rewrite 208 as an AMP iteration with matrix-valued iterates. This matrix-valued AMP is designed to be a special case of an abstract AMP iteration for which a state evolution result has been established in [59], [77]. This state evolution result is then translated to obtain the results in [eq:psiG,eq:psiX]. (See also [101] for an analysis of a general graph-based AMP iteration that includes AMP with matrix-valued iterates.) Given the iteration in 208 , for \(t \ge 1\), let \[\begin{align} e^t \mathrel{\vcenter{:}}= \begin{bmatrix} g_1 & g_2 & u^t \end{bmatrix} , \qquad h^{t+1} \mathrel{\vcenter{:}}= \, & v^{t+1} - \chi_{1,t+1}\xone - \chi_{2,t+1}\xtwo \, , \label{eq:etht195def} \end{align}\tag{209}\] where we recall that \(g_1 = \Abar \xone\), \(g_2 = \Abar \xtwo\), and \(\chi_{1,t}, \chi_{2,t}\) are the state evolution parameters computed via the recursion in [eq:UtVt_def,eq:muU_update,eq:G1G2Y_joint]. We also introduce the functions \(\brf_t: \reals^3 \to \reals^3\) and \(\brg_t:\reals^5 \to \reals\), defined as: \[\begin{align} & \brf_t(h^t; \, \xone , \xtwo) = \begin{bmatrix} \sqrt{\delta} \xone, & \sqrt{\delta} \xtwo, & f_t(h^t+ \chi_{1,t}\xone + \chi_{2,t}\xtwo; \xone,\xtwo ) \end{bmatrix}, \\ & \brg_t(e^t; \, \veta, \veps) = g_t(e^t_3; \, q(\veta \odot e^t_1 + (1-\veta) \odot e^t_2, \veps)). \end{align}\] Here, \(\odot\) denotes element-wise multiplication, \(\brf_t\) and \(\brg_t\) act row-wise on their matrix-valued inputs and \(e^t_{j}\) denotes the \(j\)-th column of \(e^t \in \reals^{n \times 3}\). We also recall the notation \(\veta=(\eta_1, \ldots, \eta_n)\), \(\veps=(\eps_1, \ldots, \eps_n)\) and that \(y = q( \veta \odot g_{1} + (1-\veta) \odot g_2, \, \veps) = q(\veta \odot e^t_1 + (1-\veta) \odot e^t_2, \, \veps)\). With these definitions, we claim that the AMP iteration in 208 is equivalent to the following one: \[\begin{align} & e^t = \frac{1}{\sqrt{\delta}} \Abar \brf_t(h^t; \, \xone , \xtwo) \, - \, \brg_{t-1}(e^{t-1}; \, \veta, \veps) \sfB_t^{\top}, \\ & h^{t+1} = \frac{1}{\sqrt{\delta}} \Abar^{\top} \brg_t(e^t; \, \veta, \veps) \, - \, \brf_t(h^t; \, \xone , \xtwo) \sfC_t^{\top}, \end{align} \label{eq:et95ht195def}\tag{210}\] where \(\sfB_t \in \reals^{3 \times 1}\) and \(\sfC_t \in \reals^{1 \times 3}\) are given by: \[\begin{align} & \sfB_t = \begin{bmatrix} 0 & 0 & \frac{1}{n}\sum_{i=1}^d f_t'(h_i^t + \chi_{1,t}\ol{x}^*_{1,i} + \chi_{2,t}\ol{x}^*_{2,i}; \ol{x}_{1,i}^*,\ol{x}_{2,i}^*) \end{bmatrix} ^{\top}, \nonumber \\ & \sfC_t = \begin{bmatrix} \E[\partial_{1} g_t(U_t; \, q(\eta G_1 + (1-\eta) G_2, \eps))] \\ \E[\partial_{2} g_t(U_t; \, q(\eta G_1 + (1-\eta) G_2, \eps))] \\ \frac{1}{n}\sum_{i=1}^n g_t'(u_i^t; \, q(\eta_i g^{1}_i + (1-\eta_i) g^{2}_i, \epsilon_i) )] \end{bmatrix}^\top . \label{eq:sBsC95def} \end{align}\tag{211}\] Here \(\partial_{1} g_t\) and \(\partial_{2} g_t\) refer to the partial derivatives of \(g_t(u; \, q( \eta g_1 + \eta g_2, \eps) )\) with respect to \(g_1\) and \(g_2\), respectively, \((G_1, G_2, \eta, \eps) \sim \normal(0,1) \otimes \normal(0,1) \otimes \bern(\alpha) \otimes P_{\epsilon}\), and \(U_t\) is defined as in 81 . The iteration in 210 is initialized with \(e^0= [g_1 \quad g_2 \quad \frac{1}{\sqrt{\delta}}\Abar\tv^0 ]\) .

The equivalence between [eq:gamp-eqn-det,eq:et_ht1_def] can be seen by substituting in 210 the definitions of \(e^t\) and \(h^{t+1}\) from 209 , and the fact that by Stein’s lemma, \(\chi_{1,t+1}\) and \(\chi_{2,t+1}\) defined in [eq:UtVt_def,eq:muU_update,eq:G1G2Y_joint] can be expressed as [59]: \[\chi_{1,t+1} = \sqrt{\delta}\, \E[\partial_{1} g_t(U_t; q(\eta G_1 + (1-\eta) G_2, \eps))] , \; \chi_{2,t+1} = \sqrt{\delta}\, \E[\partial_{2} g_t(U_t; q(\eta G_1 + (1-\eta) G_2, \eps))].\] The recursion in 210 is a special case of the abstract AMP recursion with matrix-valued iterates for which a state evolution result has been established [77], [59]. The standard form of the abstract AMP recursion uses empirical estimates (instead of expected values) for the first two entries of \(\sfC_t\) in 211 . However, the state evolution result remains valid for the recursion in 210 (see Remark 4.3 of [59]). This result states that the empirical distributions of the rows of \(e^t\) and \(h^{t+1}\) converge to the Gaussian distributions \(\normal(0, \Sigma_t)\) and \(\normal(0, \Omega_{t+1})\), respectively. The covariances \(\Sigma_t \in \reals^{3 \times 3}\) and \(\Omega_{t+1} \in \reals\) are defined by the following state evolution recursion: \[\begin{align} &\Sigma_{t} = \frac{1}{\delta} \E[\brf_t(G^{\omega}_t; \, X_1, X_2) \brf_t(G^{\omega}_t; \, X_1, X_2)^{\top}] \tag{212} \\ &= { \begin{bmatrix} 1 & 0 & \frac{\E[ X_1 f_t( G^{\omega}_t + \chi_{1,t}X_1 + \chi_{2,t}X_2; X_1,X_2) ]}{\sqrt{\delta}} \\ 0 & 1 & \frac{\E[ X_2 f_t( G^{\omega}_t + \chi_{1,t}X_1 + \chi_{2,t}X_2; X_1,X_2) ]}{\sqrt{\delta}} \\ \frac{\E[ X_1 f_t( G^{\omega}_t + \chi_{1,t}X_1 + \chi_{2,t}X_2; X_1,X_2) ]}{\sqrt{\delta}} & \frac{\E[ X_2 f_t( G^{\omega}_t + \chi_{1,t}X_1 + \chi_{2,t}X_2; X_1,X_2) ]}{\sqrt{\delta}} & \frac{\E[(f_t( G^{\omega}_t + \chi_{1,t}X_1 + \chi_{2,t}X_2; X_1,X_2))^2]}{\delta} \end{bmatrix} } , \nonumber \\ &\Omega_{t+1} = \E[ (\brg_t(G^{\sigma}_t; \, \eta, \epsilon))^2 ] = \E[(g_t(G^{\sigma}_{t,3}; \, q(G^{\sigma}_{t,1}, G^{\sigma}_{t,2}, \eta, \epsilon)))^2], \tag{213} \end{align}\] where \(G^{\sigma}_t \equiv (G^{\sigma}_{t,1}, G^{\sigma}_{t,2}, G^{\sigma}_{t,3} ) \sim \normal(0, \Sigma_t)\) is independent of \((\eta, \eps) \sim \bern(\alpha) \otimes P_{\epsilon}\), and \((G^{\omega}_t, X_1, X_2) \sim \normal(0, \Omega_t) \otimes \normal(0,1) \otimes \normal(0,1)\). The recursion is initialized with \[\Sigma_0 = \begin{bmatrix} 1 & 0 & \frac{1}{\sqrt{\delta}}\E[ F_0(X_1, X_2) X_1 ] \\ 0 & 1 & \frac{1}{\sqrt{\delta}}\E[ F_0(X_1, X_2) X_2 ] \\ \frac{1}{\sqrt{\delta}}\E[ F_0(X_1, X_2) X_1 ] & \frac{1}{\sqrt{\delta}}\E[ F_0(X_1, X_2) X_2 ] & \frac{1}{\delta} \E[ (F_0(X_1, X_2))^2 ] \end{bmatrix}. \label{eq:Sigma0}\tag{214}\]

The sequences \((G^{\sigma}_t)_{t \geq 0}\) and \((G^{\omega}_{t+1})_{t \geq 0}\) are each jointly Gaussian with the following covariance structure: \[\begin{align} &G^{\sigma}_{t,1} = G_1, \qquad G^{\sigma}_{t,2} = G_2, \quad \forall \, t \ge 0 \quad \text{ where } (G_1, G_2) \sim \normal(0,1) \otimes \normal(0,1), \\ &\expt{G^{\sigma}_{0,3}G^{\sigma}_{t,3}} = \frac{1}{\delta}\expt{F_0(X_1, X_2) f_t(G^{\omega}_t + \chi_{1,t}X_1 + \chi_{2,t}X_2; X_1,X_2)}, \quad t \ge 1, \end{align} \label{eq:Gsigma95Gomega}\tag{215}\] and for \(r, t \ge 1\): \[\begin{multline} \expt{G^{\omega}_{r} G^{\omega}_t } = \E[g_{r-1}(G^{\sigma}_{r-1,3}; \, q(\eta G^{\sigma}_{r-1,1} + (1-\eta) G^{\sigma}_{r-1,2}, \epsilon) \\ \times g_{t-1}(G^{\sigma}_{t-1,3}; \, q(\eta G^{\sigma}_{t-1,1} + (1 - \eta) G^{\sigma}_{t-1,2}, \epsilon)], \end{multline} \begin{multline} \expt{G^{\sigma}_{r,3}G^{\sigma}_{t,3}} =\frac{1}{\delta} \expt{f_r(G^{\omega}_r + \chi_{1,r}X_1 + \chi_{2,r}X_2; X_1,X_2) \, f_t(G^{\omega}_t + \chi_{1,t}X_1 + \chi_{2,t}X_2; X_1,X_2)}. \end{multline} \label{eq:Grt95corr}\tag{216}\]

The following proposition follows from the state evolution result in [77], [59] for AMP with matrix-valued iterates.

With setup and assumptions of [prop:GAMP95SE], consider the AMP recursion in 210 . The following holds almost surely for any \(\PL(2)\) functions \(\Psi: \reals^{t+1} \to \reals\), \(\Phi: \reals^{t+3} \to \reals\), for \(t \geq 0\): \[\begin{align} & \lim_{n \to \infty}\frac{1}{n} \sum_{i=1}^n \Psi( e^t_i, e^{t-1}_i, \ldots, e^0_{i} ) = \E[ \Psi( G^{\sigma}_t, \, G^{\sigma}_{t-1}, \, \ldots, G^{\sigma}_0) ], \tag{217} \\ & \lim_{d \to \infty} \frac{1}{d} \sum_{i=1}^d \Phi( \ol{x}^*_{1,i}, \ol{x}^*_{2,i}, h^{t+1}_i, h^{t}_i, \ldots, h^1_i) = \E[ \Phi(X_1, X_2, G^{\omega}_{t+1}, \, G^{\omega}_{t}, \, \ldots, G^{\omega}_1) ], \tag{218} \end{align}\] where the joint distributions of \((G^{\sigma}_t, \, G^{\sigma}_{t-1}, \, \ldots, G^{\sigma}_0)\) and \((G^{\omega}_{t+1}, \, G^{\omega}_{t}, \, \ldots, G^{\omega}_1))\) are as given in [eq:Sigmat_rec,eq:Omegat_rec,eq:Sigma0,eq:Gsigma_Gomega,eq:Grt_corr].

Recall the definitions of \(e^t, h^{t+1}\) from 209 . Then, the results in [eq:psiGsigma,eq:psiGomega] imply that for any \(\PL(2)\) function \(\Phi: \reals^{t+3} \to \reals\), we have for \(t \ge 0\): \[\begin{align} & \lim_{d \to \infty} \frac{1}{d} \sum_{i=1}^d \Phi( \ol{x}^*_{1,i}, \ol{x}^*_{2,i}, v^{t+1}_i, \ldots, v^1_i) \nonumber \\ & \qquad = \E[ \Phi(X_1, X_2, \chi_{1,t+1}X_1 + \chi_{2,t+1}X_2 + G^{\omega}_{t+1}, \, \ldots, \, \chi_{1,1}X_1 + \chi_{2,1}X_2 +G^{\omega}_1) ] \tag{219} \\ & \lim_{n \to \infty}\frac{1}{n} \sum_{i=1}^n \Phi( g_{1,i}, g_{2,i}, u^t_{i}, \ldots, u^0_{i} ) =\expt{\Phi(G_1, G_2, G^{\sigma}_{t,3}, \ldots, G^{\sigma}_{0,3} )}. \tag{220} \end{align}\] Recalling \(G^{\sigma}_{t,1}=G_1\) and \(G^{\sigma}_{t,2}=G_2\) for \(t \ge 0\), using 212 we can write \[\begin{gather} G^{\sigma}_{t,3} = \expt{G^{\sigma}_{t,3} \mid G^{\sigma}_{t,1}, G^{\sigma}_{t,2}} \, + \, W^{\sigma}_t = \frac{1}{\sqrt{\delta}} \expt{X_1 f_t( G^{\omega}_t + \chi_{1,t}X_1 + \chi_{2,t}X_2; X_1,X_2)} G_1 \notag \\ + \, \frac{1}{\sqrt{\delta}} \expt{ X_2 f_t( G^{\omega}_t + \chi_{1,t}X_1 + \chi_{2,t}X_2; X_1,X_2) } G_2 \, + \, W^{\sigma}_t, \notag \end{gather}\] where \(W^{\sigma}_t\) is a zero-mean Gaussian independent of \((G_1, G_2)\) with variance \[\begin{gather} \expt{(W^{\sigma}_t)^2} = \frac{1}{\delta} \E[(f_t( G^{\omega}_t + \chi_{1,t}X_1 + \chi_{2,t}X_2; X_1,X_2))^2] \\ - \frac{1}{\delta} (\expt{X_1 f_t( G^{\omega}_t + \chi_{1,t}X_1 + \chi_{2,t}X_2; X_1,X_2)} )^2 - \frac{1}{\delta} ( \expt{ X_2 f_t( G^{\omega}_t + \chi_{1,t}X_1 + \chi_{2,t}X_2; X_1,X_2) })^2. \notag \end{gather}\] To complete the proof, for \(t \ge 0\), we define: \[\begin{align} & W_{U,t} \mathrel{\vcenter{:}}= W^{\sigma}_t, \quad W_{V,t+1}\mathrel{\vcenter{:}}= G^{\omega}_{t+1}, \\ & \mu_{1,t} \mathrel{\vcenter{:}}= \frac{1}{\sqrt{\delta}} \E[ X_1 f_t( G^{\omega}_t + \chi_{1,t}X_1 + \chi_{2,t}X_2; X_1,X_2) ], \\ & \mu_{2,t} \mathrel{\vcenter{:}}= \frac{1}{\sqrt{\delta}} \E[ X_2 f_t( G^{\omega}_t + \chi_{1,t}X_1 + \chi_{2,t}X_2; X_1,X_2) ], \\ & U_t\mathrel{\vcenter{:}}= G^{\sigma}_{3,t} = \mu_{1,t} G_1 + \mu_{2,t} G_2 + W_{U,t}, \quad V_{t+1} = \chi_{1,t}X_1 + \chi_{2,t}X_2 + W_{V, t+1}. \end{align} \notag\] Using these in the convergence statements in [eq:thm_result_vi,eq:thm_result_ui] gives the result of [prop:GAMP95SE].

16 Auxiliary lemmas↩︎

Lemma 4. Consider the setting of 2 and let [itm:assump-preproc-spec] hold. Then the following properties of \(\phi(\lambda),\psi(\lambda;\Delta)\) hold.

  1. \(\phi(\cdot)\) is strictly decreasing;

  2. For any \(\Delta>0\), \(\psi(\cdot;\Delta)\) is strictly convex in the first argument;

  3. For any \(\lambda>\sup\supp(Z)\), \(\psi(\lambda;\cdot)\) is strictly decreasing in the second argument.

Proof. The proof follows by checking the derivatives. For \(\phi\), we have \[\begin{align} \frac{\diff}{\diff\lambda}\phi(\lambda) &= \expt{\frac{ZG^2}{\lambda-Z}} - \lambda\expt{\frac{ZG^2}{(\lambda-Z)^2}} = - \expt{\paren{\frac{ZG}{\lambda-Z}}^2} < 0 , \label{eqn:phi-deriv} \end{align}\tag{221}\] The last strict inequality holds since \(Z\) is not almost surely zero, i.e., \(\prob{Z = 0}<1\) in [itm:assump-preproc-spec]. [itm:property-1] of 4 then follows. For \(\psi\), we have \[\begin{align} \frac{\partial}{\partial\lambda}\psi(\lambda;\Delta) &= \frac{1}{\Delta} + \expt{\frac{Z}{\lambda - Z}} - \lambda \expt{\frac{Z}{(\lambda-Z)^2}} = \frac{1}{\Delta} - \expt{\frac{Z^2}{(\lambda-Z)^2}} , \label{eqn:derivative-psi} \end{align}\tag{222}\] and \[\begin{align} \frac{\partial^2}{\partial\lambda^2}\psi(\lambda;\Delta) &= 2\expt{\frac{Z^2}{(\lambda-Z)^3}} . \notag \end{align}\] Since \(\lambda>\sup\supp(Z)\) and \(Z\) is not almost surely zero, \(\psi''(\cdot;\Delta)>0\) and therefore [itm:property-2] of 4 holds. Finally, [itm:property-3] of 4 is obvious, provided \(\lambda>\sup\supp(Z)>0\) where the second inequality is guaranteed by [itm:assump-preproc-spec]. ◻

Lemma 5 (Explicit formulas). Fix any \(\Delta>0\). The parameter \(\ol\lambda(\Delta)\) satisfies \[\begin{align} \expt{\paren{\frac{Z}{\ol\lambda(\Delta) - Z}}^2} &= \frac{1}{\Delta} . \label{eqn:def-lam-bar-explicit} \end{align}\qquad{(14)}\] If \(\lambda^*(\Delta) > \ol\lambda(\Delta)\), the parameter \(\lambda^*(\Delta)\) satisfies \[\begin{align} \expt{\frac{Z(G^2 - 1)}{\lambda^*(\Delta) - Z}} &= \frac{1}{\Delta} . \label{eqn:def-lami-star-supcrit-explicit} \end{align}\qquad{(15)}\]

Proof. Since \(\ol\lambda(\Delta)\) (cf.@eq:eqn:def-lam-bar ) is the minimum point of \(\psi(\cdot;\Delta)\), it satisfies \[\begin{align} \left.\frac{\partial}{\partial\lambda}\psi(\lambda;\Delta)\right|_{\lambda = \ol\lambda(\Delta)} &= 0 . \notag \end{align}\] This gives ?? according to 222 .

Under the condition \(\lambda^*(\Delta) > \ol\lambda(\Delta)\), we have \(\zeta(\lambda^*(\Delta);\Delta) = \psi(\lambda^*(\Delta);\Delta)\) and \(\lambda^*(\Delta)\) satisfies the fixed point equation \(\psi(\lambda^*(\Delta);\Delta) = \phi(\lambda^*(\Delta))\) which in turn can be written more explicitly as ?? . ◻

?? is often used with \(\Delta = \delta_i\) (\(i\in\{1,2\}\)) under the condition \(\lambda^*(\delta_i) > \ol\lambda(\delta)\). This is legitimate since \(\ol\lambda(\delta) > \ol\lambda(\delta_i)\) (see, e.g., 75 ) and the latter condition is stronger than the one in 5.

References↩︎

[1]
P. McCullagh and J. Nedler, Generalized linear models. Chapman; Hall/CRC, 1989.
[2]
Y. Shechtman, Y. C. Eldar, O. Cohen, H. N. Chapman, J. Miao, and M. Segev, “Phase retrieval with application to optical imaging: A contemporary overview,” IEEE Signal Processing Magazine, vol. 32, no. 3, pp. 87–109, 2015.
[3]
A. Fannjiang and T. Strohmer, “The numerics of phase retrieval,” Acta Numerica, vol. 29, p. 125?228, 2020.
[4]
P. T. Boufounos and booktitle=Conference. on I. S. and S. (CISS). Baraniuk Richard G, “1-bit compressive sensing,” 2008, pp. 16–21.
[5]
G. J. McLachlan and D. Peel, John Wiley & Sons, 2004.
[6]
B. Grün and F. Leisch, “Applications of finite mixtures of regression models,” 2007.
[7]
Q. Li, R. Shi, and F. Liang, PloS one, pp. 1–18, 2019.
[8]
E. Devijver, Y. Goude, and J.-M. Poggi, “Clustering electricity consumers using high-dimensional regression mixture models,” Applied Stochastic Models in Business and Industry, pp. 159–177, 2020.
[9]
K. Viele and B. Tong, “Modeling with mixtures of linear regressions,” Statistics and Computing, vol. 12, pp. 315–330, 2002.
[10]
S. Faria and G. Soromenho, “Fitting mixtures of linear regressions,” Journal of Statistical Computation and Simulation, vol. 80, no. 2, pp. 201–225, 2010.
[11]
N. Städler, P. Bühlmann, and S. van de Geer, \(\ell_1\)-penalization for mixture regression models,” TEST, vol. 19, no. 2, p. 209?256, 2010.
[12]
A. T. Chaganty and P. Liang, “Spectral experts for estimating mixtures of linear regressions,” International Conference on Machine Learning, p. 1040?1048, 2013.
[13]
X. Yi, C. Caramanis, and booktitle=International. C. on M. L. Sanghavi Sujay, “Alternating minimization for mixed linear regression,” 2014 , organization={PMLR}, pp. 613–621.
[14]
K. Zhong, P. Jain, and booktitle=Advances. in N. I. P. S. Inderjit S. Dhillon, “Mixed linear regression with multiple components,” 2016, pp. 2190–2198.
[15]
Y. Shen and book Sanghavi Sujay, “Advances in neural information processing systems,” 2019, vol. 32.
[16]
L. Zhang, R. Ma, T. T. Cai, and note=arXiv:2011. 03598. Hongzhe Li, “Estimation, confidence intervals, and large-scale hypotheses testing for high-dimensional mixed linear regression,” 2020.
[17]
A. Ghosh and booktitle =. I. C. on A. I. and S. Kannan Ramchandran, “Alternating minimization converges super-linearly for mixed linear regression,” 2020, pp. 1093–1103.
[18]
“Variable selection in finite mixture of regression models,” Journal of the American Statistical Association, vol. 102, no. 479, pp. 1025–1038, 2007.
[19]
Y. Chen, X. Yi, and booktitle=Conference. on L. T. Constantine Caramanis, “A convex formulation for mixed regression with two components: Minimax optimal rates,” 2014, pp. 560–604.
[20]
Y. Li and Y. Liang, Conference On Learning Theory, pp. 1125–1144, 2018.
[21]
S. Chen, J. Li, and Z. Song, “Learning mixtures of linear regressions in subexponential time via Fourier moments,” Proceedings of the 52nd Annual ACM Symposium on Theory of Computing, pp. 587–600, 2020.
[22]
A. Barik and booktitle=International. C. on M. L. Jean Honorio, “Sparse mixed linear regression with guarantees: Taming an intractable problem with invex relaxation,” 2022, vol. 162, pp. 1627–1646.
[23]
H. Sedghi, M. Janzamin, and booktitle=Artificial. I. and S. Anandkumar Anima, “Provable tensor methods for learning mixtures of generalized linear models,” 2016 , organization={PMLR}, pp. 1223–1231.
[24]
Y. Plan, R. Vershynin, and E. Yudovina, “High-dimensional estimation with geometric constraints,” Information and Inference, vol. 6, no. 1, pp. 1–40, 2017.
[25]
M. Mondelli, C. Thrampoulidis, and R. Venkataramanan, “Optimal combination of linear and spectral estimators for generalized linear models,” Foundations of Computational Mathematics, vol. 22, no. 5, pp. 1513–1566, 2022.
[26]
W. Luo, W. Alghamdi, and Y. M. Lu, “Optimal spectral initialization for signal recovery with applications to phase retrieval,” IEEE Transactions on Signal Processing, vol. 67, no. 9, pp. 2347–2356, 2019.
[27]
book Rangan Sundeep, “2011 IEEE international symposium on information theory proceedings , title=Generalized approximate message passing for estimation with random linear mixing,” 2011, pp. 2168–2172, keywords=Equations;Estimation;Approximation algorithms;Algorithm design and analysis;Mathematical model;Belief propagation;Approximation methods;Optimization;random matrices;estimation;belief propagation;compressed sensing.
[28]
Y. M. Lu and G. Li, “Phase transitions of spectral initialization for high-dimensional non-convex estimation,” Information and Inference: A Journal of the IMA, vol. 9, no. 3, pp. 507–541, 2020.
[29]
M. Mondelli and A. Montanari, “Fundamental limits of weak recovery with applications to phase retrieval,” Foundations of Computational Mathematics, vol. 19, no. 3, pp. 703–773, 2019.
[30]
M. I. Jordan and R. A. Jacobs, “Hierarchical mixtures of experts and the EM algorithm,” Neural Computation, vol. 6, no. 2, pp. 181–214, Mar. 1994.
[31]
“Bayesian inference in mixtures-of-experts and hierarchical mixtures-of-experts models with an application to speech recognition,” Journal of the American Statistical Association, vol. 91, no. 435, pp. 953–960, 1996.
[32]
S. Waterhouse, D. MacKay, and book Robinson Anthony, “Advances in neural information processing systems,” 1995, vol. 8.
[33]
S. Balakrishnan, M. J. Wainwright, and B. Yu, “Statistical guarantees for the EM algorithm: From population to sample-based analysis,” The Annals of Statistics, vol. 45, no. 1, pp. 77–120, 2017.
[34]
J. M. Klusowski, D. Yang, and W. D. Brinda, “Estimating the coefficients of a mixture of two linear regressions by expectation maximization,” IEEE Transactions on Information Theory, vol. 65, pp. 3515–3524, 2019.
[35]
Z. Wang, Q. Gu, Y. Ning, and H. Liu, “High dimensional EM algorithm: Statistical optimization and asymptotic normality,” Advances in neural information processing systems, pp. 2521–2529, 2015.
[36]
X. Yi and booktitle=Advances. in N. I. P. S. Constantine Caramanis, “Regularized EM algorithms: A unified framework and statistical guarantees,” 2015, pp. 1567–1575.
[37]
R. Zhu, L. Wang, C. Zhai, and booktitle=International. C. on M. L. Quanquan Gu, “High-dimensional variance-reduced stochastic gradient expectation-maximization algorithm,” 2017, vol. 70, pp. 4180–4188.
[38]
J. Fan, H. Liu, Z. Wang, and note=arXiv:1808. 06996. Zhuoran Yang, “Curse of heterogeneity: Computational barriers in sparse mixture models and phase retrieval,” 2018.
[39]
G. Arpino and booktitle =. P. of T. S. C. on L. T. Venkataramanan Ramji, “Statistical-computational tradeoffs in mixed sparse linear regression,” 2023, vol. 195 , series = Proceedings of Machine Learning Research, pp. 921–986.
[40]
W. Kong, R. Somani, Z. Song, S. Kakade, and booktitle=International. C. on M. L. Sewoong Oh, “Meta-learning for mixed linear regression,” 2020, vol. 119, pp. 5394–5404.
[41]
S. Pal, A. Mazumdar, R. Sen, and booktitle=International. C. on M. L. Ghosh Avishek, “On learning mixture of linear regressions in the non-realizable setting,” 2022, pp. 17202–17220.
[42]
K. A. Chandrasekher, A. Pananjady, and C. Thrampoulidis, “Sharp global convergence guarantees for iterative nonconvex optimization with random data,” Ann. Statist. , FJOURNAL = The Annals of Statistics, vol. 51, no. 1, pp. 179–210, 2023.
[43]
N. Tan and R. Venkataramanan, “Mixed regression via approximate message passing,” Journal of Machine Learning Research, vol. 24, no. 317, pp. 1–44, 2023.
[44]
K.-C. Li, “On principal Hessian directions for data visualization and dimension reduction: Another application of Stein’s lemma,” Journal of the American Statistical Association, vol. 87, no. 420, pp. 1025–1039, 1992.
[45]
P. Netrapalli, P. Jain, and booktitle=Advances. in N. I. P. S. (NIPS). Sanghavi Sujay, “Phase retrieval using alternating minimization,” 2013, pp. 2796–2804.
[46]
E. J. Candès, T. Strohmer, and V. Voroninski, “Phaselift: Exact and stable signal recovery from magnitude measurements via convex programming,” Communications on Pure and Applied Mathematics, vol. 66, no. 8, pp. 1241–1274, 2013.
[47]
Y. Chen and booktitle=Advances. in N. I. P. S. (NIPS). Candès Emmanuel J., “Solving random quadratic systems of equations is nearly as easy as solving linear systems,” 2015, pp. 739–747.
[48]
A. Damian, L. Pillaud-Vivien, J. Lee, and booktitle =. P. of T. S. C. on L. T. Bruna Joan, “Computational-statistical gaps in gaussian single-index models (extended abstract),” 2024, vol. 247 , series = Proceedings of Machine Learning Research, pp. 1262–1262.
[49]
R. Dudeja, M. Bakhshizadeh, J. Ma, and A. Maleki, “Analysis of spectral methods for phase retrieval with random orthogonal matrices,” IEEE Transactions on Information Theory, vol. 66, no. 8, pp. 5182–5203, 2020.
[50]
A. Maillard, F. Krzakala, Y. M. Lu, and booktitle=Mathematical. and S. M. L. Zdeborová Lenka, “Construction of optimal spectral methods in phase retrieval,” 2022 , organization={PMLR}, pp. 693–720.
[51]
D. L. Donoho, A. Maleki, and A. Montanari, “Message passing algorithms for compressed sensing,” Proceedings of the National Academy of Sciences, vol. 106, pp. 18914–18919, 2009.
[52]
M. Bayati and A. Montanari, “The dynamics of message passing on dense graphs, with applications to compressed sensing,” IEEE Transactions on Information Theory, vol. 57, pp. 764–785, 2011.
[53]
F. Krzakala, M. Mézard, F. Sausset, Y. Sun, and L. Zdeborová, “Probabilistic reconstruction in compressed sensing: Algorithms, phase diagrams, and threshold achieving matrices,” Journal of Statistical Mechanics: Theory and Experiment, vol. 2012, no. 8, p. P08009, 2012.
[54]
P. Schniter and D.-A. =. 2019. 12:15:37. +0100. Rangan Sundeep, “Compressive phase retrieval via generalized approximate message passing,” IEEE Transactions on Signal Processing, vol. 63, no. 4, pp. 1043–1055, 2014.
[55]
P. Sur and D.-A. =. 2019. 10:46:06. +0100. Candès Emmanuel J, “A modern maximum-likelihood theory for high-dimensional logistic regression,” Proceedings of the National Academy of Sciences, vol. 116, no. 29, pp. 14516–14525, 2019.
[56]
Y. Deshpande and B. Montanari Andrea, “IEEE international symposium on information theory (ISIT) , date-modified = 2017-12-23 08:45:04 +0000,” 2014, pp. 2197–2201, Title = Information–theoretically optimal sparse PCA.
[57]
S. Rangan, V. Goyal, and book Fletcher Alyson K, “Advances in neural information processing systems,” 2009, vol. 22.
[58]
T. Lesieur, F. Krzakala, and L. Zdeborová, “Constrained low-rank matrix estimation: Phase transitions, approximate message passing and applications,” Journal of Statistical Mechanics: Theory and Experiment, vol. 2017, no. 7, p. 073403, 2017.
[59]
O. Y. Feng, R. Venkataramanan, C. Rush, and R. J. Samworth, “A unifying tutorial on approximate message passing,” Foundations and Trends in Machine Learning, vol. 15, no. 4, pp. 335–536, 2022.
[60]
M. Bayati and A. Montanari, “The LASSO risk for gaussian matrices,” IEEE Transactions on Information Theory, vol. 58, pp. 1997–2017, 2012.
[61]
D. Donoho and A. Montanari, “High dimensional robust m-estimation: Asymptotic variance via approximate message passing,” Probability Theory and Related Fields, vol. 166, no. 3, pp. 935–969, 2016.
[62]
Z. Bu, J. M. Klusowski, C. Rush, and W. J. Su, “Algorithmic analysis and statistical estimation of SLOPE via approximate message passing,” IEEE Transactions on Information Theory, vol. 67, no. 1, pp. 506–537, 2020.
[63]
A. Montanari and R. Venkataramanan, “Estimation of low-rank matrices via approximate message passing,” Annals of Statistics, vol. 45, no. 1, pp. 321–345, 2021.
[64]
M. Mondelli and booktitle=International. C. on A. I. and S. Venkataramanan Ramji, “Approximate message passing with spectral initialization for generalized linear models,” 2021 , organization={PMLR}, pp. 397–405.
[65]
C. Rush and R. Venkataramanan, “Finite-sample analysis of approximate message passing,” IEEE Transactions on Information Theory, vol. 64, no. 11, pp. 7264–7286, 2018.
[66]
G. Li and Y. Wei, “A non-asymptotic framework for approximate message passing in spiked models,” arXiv preprint arXiv:2208.03313, 2022.
[67]
M. Opper, B. Cakmak, and O. Winther, “A theory of solving TAP equations for Ising models with general invariant random matrices,” Journal of Physics A: Mathematical and Theoretical, vol. 49, no. 11, p. 114002, 2016.
[68]
J. Ma and L. Ping, “Orthogonal AMP,” IEEE Access, vol. 5, pp. 2020–2033, 2017.
[69]
S. Rangan, P. Schniter, and A. K. Fletcher, “Vector approximate message passing,” IEEE Transactions on Information Theory, vol. 65, no. 10, pp. 6664–6684, 2019.
[70]
K. Takeuchi, “Rigorous dynamics of expectation-propagation-based signal recovery from unitarily invariant measurements,” IEEE Transactions on Information Theory, vol. 66, no. 1, pp. 368–386, 2020.
[71]
X. Zhong, T. Wang, and Z. Fan, “Approximate message passing for orthogonally invariant ensembles: Multivariate non-linearities and spectral initialization,” Inf. Inference , FJOURNAL = Information and Inference. A Journal of the IMA, vol. 13, no. 3, pp. Paper No. iaae024, 2024.
[72]
Z. Fan, “Approximate message passing algorithms for rotationally invariant matrices,” Annals of Statistics, vol. 50, no. 1, pp. 197–224, 2022.
[73]
M. Mondelli and book Venkataramanan Ramji, “Advances in neural information processing systems,” 2021, vol. 34, pp. 29616–29629.
[74]
R. Venkataramanan, K. Kögler, and booktitle =. P. of the 39th. I. C. on M. L. Mondelli Marco, “Estimation in rotationally invariant generalized linear models via approximate message passing,” 2022, vol. 162 , series = Proceedings of Machine Learning Research, pp. 22120–22144.
[75]
F. Kovacevic, Y. Zhang, and M. Mondelli, “Spectral estimators for multi-index models: Precise asymptotics and optimal weak recovery , booktitle=Conference on Learning Theory (COLT),” 2025.
[76]
S. T. Belinschi, H. Bercovici, M. Capitaine, and M. Février, “Outliers in the spectrum of large deformed unitarily invariant models,” Ann. Probab. , FJOURNAL = The Annals of Probability, vol. 45, no. 6A, pp. 3571–3625, 2017.
[77]
A. Javanmard and A. Montanari, “State evolution for general approximate message passing algorithms, with applications to spatial coupling,” Information and Inference, pp. 115–144, 2013.
[78]
T. Tao and V. Vu, “Random matrices: Localization of the eigenvalues and the necessity of four moments,” Acta Math. Vietnam. , FJOURNAL = Acta Mathematica Vietnamica, vol. 36, no. 2, pp. 431–449, 2011.
[79]
L. Erdős, H.-T. Yau, and J. Yin, “Bulk universality for generalized Wigner matrices,” Probab. Theory Related Fields , FJOURNAL = Probability Theory and Related Fields, vol. 154, no. 1–2, pp. 341–407, 2012.
[80]
A. M. Tulino, G. Caire, S. Shamai, and S. Verdu, “Capacity of channels with frequency-selective and time-selective fading,” IEEE Transactions on Information Theory, vol. 56, no. 3, pp. 1187–1215, 2010.
[81]
B. Farrell, “Limiting empirical singular value distribution of restrictions of discrete Fourier transform matrices,” J. Fourier Anal. Appl. , FJOURNAL = The Journal of Fourier Analysis and Applications, vol. 17, no. 4, pp. 733–753, 2011.
[82]
G. W. Anderson and keywords =. F. probability,. A. liberation,. R. matrices,. U. matrices,. H. matrices Brendan Farrell, “Asymptotically liberating sequences of random unitary matrices,” Advances in Mathematics, vol. 255, pp. 381–413, 2014.
[83]
M. Bayati, M. Lelarge, and A. Montanari, “Universality in polytope phase transitions and message passing algorithms,” The Annals of Applied Probability, vol. 25, no. 2, pp. 753–822, keywords = compressed sensing, message passing, polytope neighborliness, random matrices, Universality, 2015.
[84]
W.-K. Chen and W.-K. Lam, “Universality of approximate message passing algorithms,” Electron. J. Probab. , FJOURNAL = Electronic Journal of Probability, vol. 26, no. 4235487, pp. Paper No. 36, 44, MRCLASS = 60F05 (60B20 62H12 94A12), MR, 2021.
[85]
R. Dudeja and M. Bakhshizadeh, “Universality of linearized message passing for phase retrieval with structured sensing matrices,” IEEE Trans. Inform. Theory , FJOURNAL = Institute of Electrical and Electronics Engineers. Transactions on Information Theory, vol. 68, no. 11, pp. 7545–7574, 2022.
[86]
T. Wang, X. Zhong, and Z. Fan, “Universality of approximate message passing algorithms and tensor networks,” Ann. Appl. Probab. , FJOURNAL = The Annals of Applied Probability, vol. 34, no. 4, pp. 3943–3994, 2024.
[87]
R. Dudeja, Y. M. Lu, and S. Sen, “Universality of approximate message passing with semirandom matrices,” The Annals of Probability, vol. 51, no. 5, pp. 1616–1683, keywords = message passing, random matrices, Spin glasses, Universality, 2023.
[88]
R. Dudeja, S. Sen, and Y. M. Lu, “Spectral universality in regularized linear regression with nearly deterministic sensing matrices,” IEEE Transactions on Information Theory, pp., keywords=Sensors;Vectors;Noise;Symmetric matrices;Linear regression;Discrete cosine transforms;Optimization;universality;approximate message passing;compressed sensing;regularized linear regression, 2024.
[89]
J. Barbier, F. Krzakala, N. Macris, L. Miolane, and L. Zdeborová, “Optimal errors and phase transitions in high-dimensional generalized linear models,” Proc. Natl. Acad. Sci. USA , FJOURNAL = Proceedings of the National Academy of Sciences of the United States of America, vol. 116, no. 12, pp. 5451–5460, 2019.
[90]
A. Maillard, B. Loureiro, F. Krzakala, and booktitle=Advances. in N. I. P. S. Zdeborová Lenka, “Phase retrieval in high dimensions: Statistical and computational phase transitions,” 2020, vol. 33, pp. 11071–11082.
[91]
M. Brennan and booktitle =. P. of T. T. C. on L. T. Bresler Guy, “Reducibility and statistical-computational gaps from secret leakage,” 2020, vol. 125 , series = Proceedings of Machine Learning Research, pp. 648–847.
[92]
M. Celentano, A. Montanari, and booktitle=Conference. on L. T. Wu Yuchen, “The estimation error of general first order methods,” 2020 , organization={PMLR}, pp. 1078–1141.
[93]
F. Benaych-Georges and R. R. Nadakuditi, “The eigenvalues and eigenvectors of finite, low rank perturbations of large random matrices,” Advances in Mathematics, vol. 227, no. 1, pp. 494–521, 2011.
[94]
F. Benaych-Georges and R. R. Nadakuditi, “The singular values and vectors of low rank perturbations of large rectangular random matrices,” Journal of Multivariate Analysis, vol. 111, pp. 120–135, 2012.
[95]
D. Voiculescu, “Limit laws for random matrices and free products,” Invent. Math. , FJOURNAL = Inventiones Mathematicae, vol. 104, no. 1, pp. 201–220, 1991.
[96]
R. Speicher, “Free convolution and the random sum of matrices,” Publ. Res. Inst. Math. Sci. , FJOURNAL = Kyoto University. Research Institute for Mathematical Sciences. Publications, vol. 29, no. 5, pp. 731–744, 1993.
[97]
D. Voiculescu, “Addition of certain noncommuting random variables,” J. Funct. Anal. , FJOURNAL = Journal of Functional Analysis, vol. 66, no. 3, pp. 323–346, 1986.
[98]
Z. Bai and J. Yao, “On sample eigenvalues in a generalized spiked population model,” Journal of Multivariate Analysis, vol. 106, pp. 167–177, 2012.
[99]
J. W. Silverstein and S.-I. Choi, “Analysis of the limiting spectral distribution of large-dimensional random matrices,” J. Multivariate Anal. , FJOURNAL = Journal of Multivariate Analysis, vol. 54, no. 2, pp. 295–309, 1995.
[100]
L. D?mbgen, R. Samworth, and f. Schuhmacher Dominic", “Approximation by log-concave distributions, with applications to regression,” Annals of Statistics", journal = "Annals of Statistics, vol. 39, no. 2, pp. 702–730, Apr. 2011.
[101]
C. Gerbelot and R. Berthier, “Graph-based approximate message passing iterations,” Inf. Inference , FJOURNAL = Information and Inference. A Journal of the IMA, vol. 12, no. 4, pp. Paper No. iaad020, 67, 2023.

  1. School of Mathematics, University of Bristol ().↩︎

  2. Institute of Science and Technology Austria ().↩︎

  3. Department of Engineering, University of Cambridge ().↩︎

  4. Submitted to the editors 2026-07-12.↩︎

  5. Asymptotic freeness can be thought of as the random matrix analogue of independence of random variables.↩︎

  6. More formally, it is proved that the minimum mean squared error achieved by the Bayes-optimal estimator coincides with the error of a trivial estimator which always outputs the all-0 vector.↩︎