Breaking the One-Dimensional Expressibility–Trainability Tradeoff


Abstract

Expressive parameterized quantum circuits (PQCs) are often designed under a dilemma: the growth of expressibility and entangling power (EP) that improves Hilbert-space coverage is also expected to randomize an ansatz and activate barren-plateau (BP) conditions. We show that this dilemma is not a one-dimensional tradeoff. The usual picture collapses three inequivalent objects—parameter-ensemble coverage, fixed-circuit entangling response, and local gradient moments—into one scalar narrative. For a fixed circuit probed by Haar-product inputs, EP is a global two-copy mean of the output-entanglement distribution, whereas entangling-power deviation (EPD) is a global four-copy fluctuation descriptor. Gradient variance, however, is a local two-copy contraction selected by a parameter light cone and a cost observable. This moment hierarchy yields an analytic separation: equal EP need not imply equal trainability, as witnessed by equal-EP circuits with different EPDs and different gradient variances. These separations turn EP and EPD into a two-dial design rule for PQC ansatzes: EP measures how far the circuit has moved along the coverage dial, while EPD monitors whether input-dependent variability remains. We find that ansatz routes can reach high, Haar-like coverage before EPD and gradient variance collapse, showing that coverage and BP activation are distinct crossover events. The EP/EPD framework thus breaks the apparent one-dimensional expressibility–trainability tradeoff into a practical design rule: search for highly expressive PQCs in the window where coverage is high but BP-like homogenization has not yet erased trainable structure.

↩︎

Introduction.—Parameterized quantum circuits (PQCs) are judged by two requirements that pull in opposite directions. PQCs must generate sufficiently rich state families for variational quantum algorithms and/or quantum machine learning [1][5]. PQCs must also retain gradients large enough to be estimated and used by a classical optimizer. The first requirement is commonly quantified through expressibility, namely the closeness of parameter-sampled states to Haar coverage, together with the average Meyer-Wallach entanglement generated by the circuit, which we call entangling power (EP) [6], [7]. The second requirement is trainability: the circuit must not become so randomized that the cost landscape loses usable gradient information. Barren-plateau theory made this requirement sharp by proving that sufficiently random circuits can have exponentially small gradients [8][10]. These facts are often compressed into the slogan that increasing expressibility together with EP drives a circuit toward Haar-random structure and hence toward barren-plateau (BP) behavior.

This slogan is useful as a warning against blindly random circuits, but it is too coarse to serve as a design principle. A BP criterion tests whether the gradient-relevant moment has already randomized; it does not say what kind of coverage was gained, nor whether input-dependent structure remains before that randomization. The reason is threefold. First, expressibility is a property of a parameter ensemble, whereas EP is naturally a property of one fixed unitary tested on many product inputs. Second, EP is a global scalar, whereas the gradient of a particular parameter is a local object selected by a cost observable and an effective light cone. Third, a mean does not determine a fluctuation profile. These distinctions help explain why the descriptor-performance relations in classification learning studies can be nuanced: expressibility may correlate with learning accuracy, while the average EP can be much less predictive [11].

In this Letter, we formulate the missing structure as a moment hierarchy. For that, we extend the entangling-power deviation (EPD) viewpoint, previously developed for bipartite gates [12], to \(n\)-qubit PQCs. The result is a two-dial design principle: First, a coverage dial records how far the ansatz has moved toward high mean entangling strength or Haar-like coverage. Second, a variability dial records whether the output entanglement generated from product inputs remains input-dependent. We show analytically that these dials are distinct, prove that equal mean EP does not imply equal trainability, and show that useful ansatz routes can reach high coverage before EPD and gradient variance collapse.

Our central claim is that the aforementioned slogan, i.e., the expressibility–vs–trainability tradeoff, is not a hard one-dimensional obstruction: high coverage can be reached before the BP condition is locally activated. In this window, EP records that the ansatz has become sufficiently nonlocal, while EPD records that product-input responses have not homogenized. Thus, our EP/EPD framework identifies high-coverage, BP-free candidates by separating mean coverage from the collapse of input-dependent variability. It provides a constructive screening principle for PQC design: maximize coverage without crossing into the low-EPD regime where gradient variance is expected to collapse.

Fixed-circuit descriptors.—We start by separating the two ensemble levels that are often conflated.

Definition 1 (Ensemble levels). Let \(\hat{U}_L(\boldsymbol{\theta})\) be a depth-\(L\) PQC on \(\mathcal{H}_n=(\mathbb{C}^2)^{\otimes n}\). The parameter ensemble \(\nu_L\) samples circuit instances \(\hat{U}_L(\boldsymbol{\theta})\) and underlies expressibility. The product-input ensemble \[\begin{align} \mu_{\mathrm{prod}} = \bigotimes_{j=1}^n\mu_{\mathrm{Haar}}^{(1)}, \quad \left|\psi\right\rangle = \bigotimes_{j=1}^n\left|\psi_j\right\rangle, \label{eq:product95ensemble95main} \end{align}\tag{1}\] with independent one-qubit Haar states \(\left|\psi_j\right\rangle\), probes a single fixed circuit. These two ensembles are not interchangeable.

For a pure state \(\left|\phi\right\rangle\), we use the Meyer-Wallach entanglement [13] \[\begin{align} Q(\left|\phi\right\rangle) = \frac{2}{n}\sum_{j=1}^n\left(1-\operatorname{Tr}\hat{\rho}_j^2\right), \quad \hat{\rho}_j = \operatorname{Tr}_{\bar j}\left|\phi\right\rangle\!\!\left\langle \phi\right|. \label{eq:Q95def95main} \end{align}\tag{2}\] It satisfies \(0\le Q\le1\) and vanishes on fully product states.

We now define the two fixed-circuit quantities used throughout this Letter.

Definition 2 (Fixed-circuit EP and EPD). For a fixed unitary \(\hat{U}\in U(2^n)\), \[\begin{align} \mathrm{EP}(\hat{U}) &=& \mathbb{E}_{\psi\sim\mu_{\mathrm{prod}}}\bigl[Q(\hat{U}\left|\psi\right\rangle)\bigr], \nonumber \\ \mathrm{EPD}(\hat{U}) &=& \sqrt{\operatorname{Var}_{\psi\sim\mu_{\mathrm{prod}}}\bigl[Q(\hat{U}\left|\psi\right\rangle)\bigr]} . \label{eq:EP95EPD95def95main} \end{align}\tag{3}\] EP is the mean entangling strength of a circuit instance. EPD is the standard deviation of the same output-entanglement distribution, and hence measures the input-state sensitivity [12]. Both are invariant under local pre- and post-unitaries.

The fixed-circuit viewpoint has two practical advantages. First, it separates the nonlocal action of a circuit block from the arbitrary local basis in which the block is written. Second, it lets us ask a question: does the same circuit entangle all product inputs in nearly the same way, or does it preserve strong input dependence? This second question is invisible to the mean but is directly relevant to the trainability, because the gradients are also response functions of a circuit with respect to perturbations of an input, a parameter, or an observable.

The distinction between the two descriptors becomes exact in copy space. Let \(S_j\) swap \(j\)-qubit between two copies and let us define \(\hat{K}_Q = \frac{2}{n} \sum_{j=1}^n (\hat{\mathbb{1}} - \hat{S}_j)\). Then, \(Q(\left|\phi\right\rangle)=\operatorname{Tr}[\hat{K}_Q \hat{\rho}^{\otimes2}]\). On four copies, we define \(\hat{T}_Q = \hat{K}_Q^{(12)}\hat{K}_Q^{(34)}\), where the superscripts denote copy pairs. For the fixed-circuit moment \[\begin{align} \hat{M}_{\hat{U}}^{(t)}=\mathbb{E}_{\psi \sim \mu_{\mathrm{prod}}} \left[\left(\hat{U}\left|\psi\right\rangle\!\!\left\langle \psi\right|\hat{U}^\dagger\right)^{\otimes t}\right],\label{eq:MU95t95main} \end{align}\tag{4}\] we have the following structural identity.

Theorem 1 (Moment hierarchy). For every fixed circuit \(\hat{U}\), \[\begin{align} \mathrm{EP}(\hat{U}) &=& \operatorname{Tr}\bigl[\hat{K}_Q \hat{M}_{\hat{U}}^{(2)}\bigr], \nonumber \\ \mathrm{EPD}(\hat{U})^2 &=& \operatorname{Tr}\bigl[\hat{T}_Q \hat{M}_{\hat{U}}^{(4)}\bigr] - \mathrm{EP}(\hat{U})^2. \label{eq:main95hierarchy95theorem} \end{align}\tag{5}\] Thus, EP is a global two-copy scalar, while EPD is a global four-copy fluctuation descriptor.

Proof sketch. The local swap identity \(\operatorname{Tr}[\hat{S}_j (\hat{A} \otimes \hat{B})] = \operatorname{Tr}[(\operatorname{Tr}_{\bar j} \hat{A}) (\operatorname{Tr}_{\bar j} \hat{B})]\) gives Eq. (2 ) as \(\operatorname{Tr}[\hat{K}_Q \hat{\rho}^{\otimes 2}]\). Squaring the same two-copy expectation on the two disjoint copy pairs \((1,2)\) and \((3,4)\) gives \(Q(\left|\phi\right\rangle)^2 = \operatorname{Tr}[\hat{T}_Q \hat{\rho}^{\otimes4}]\). Averaging these two identities over Haar-product inputs after applying \(\hat{U}\) yields Eq. (5 ). The full swap derivation and norm bounds are given in Sec. II of the Supplemental Material. ◻

Proposition 1 (Spectral obstruction). For every \(t \ge 1\), \(\hat{M}_{\hat{U}}^{(t)} = \hat{U}^{\otimes t} \hat{M}_{\mathrm{prod}}^{(t)} \hat{U}^{\dagger\otimes t}\). Hence, the spectrum of \(\hat{M}_{\hat{U}}^{(t)}\) is independent of \(\hat{U}\). For \(n > 1\) and \(t \ge 2\), it differs from the spectrum of the full Haar pure-state moment.

Proof sketch. Linearity pulls \(\hat{U}^{\otimes t}\) outside the product-input expectation. The product-input moment is proportional to the projector onto the locally symmetric subspace \(\bigotimes_j\mathrm{Sym}^t(\mathbb{C}^2_j)\), of dimension \((t+1)^n\). The full Haar moment is proportional to the projector onto \(\mathrm{Sym}^t(\mathcal{H}_n)\), of dimension \(\binom{2^n+t-1}{t}\). These subspaces are not equal for \(n > 1, t \ge 2\); the detailed support comparison is given in Sec. I of the Supplemental Material. Therefore, Haar is used below as a scalar entanglement benchmark and as a local light-cone benchmark, not as a global operator limit of fixed-circuit product-input moments. ◻

This obstruction is important for interpretation. When a deep parameterized circuit is called “Haar-like”, the statement usually concerns a particular observable, fidelity distribution, or low-order local contraction. It cannot mean that one fixed unitary acting on Haar-product inputs literally produces the full Haar moment operator on the global Hilbert space.

Gradient variance is local.—Consider a cost \[\begin{align} C(\boldsymbol{\theta},\psi) = \operatorname{Tr}[\hat{O}\hat{\rho}_\psi(\boldsymbol{\theta})], \end{align}\] where \(\hat{\rho}_\psi(\boldsymbol{\theta})=\hat{U}(\boldsymbol{\theta})\left|\psi\right\rangle\!\!\left\langle \psi\right|\hat{U}(\boldsymbol{\theta})^\dagger\) and \(\hat{O} = \hat{O}^\dagger\) is the bounded cost observable whose expectation value defines the optimization objective. Here and below, \(\boldsymbol{\theta}\) denotes the full parameter vector, whereas \(\theta_i\) denotes its \(i\)th component. Suppose that \(\theta_i\) occurs in the elementary gate \(\exp(-i \theta_i \hat{G}_i)\) with Hermitian generator \(\hat{G}_i\). We decompose the circuit as \[\begin{align} \hat{U}(\boldsymbol{\theta})=\hat{U}_i^{\rm post}\exp(-i \theta_i \hat{G}_i)\hat{U}_i^{\rm pre}, \label{eq:prepost95decomposition95main} \end{align}\tag{6}\] where \(\hat{U}_i^{\rm pre}\) collects all gates before the parameterized gate and \(\hat{U}_i^{\rm post}\) collects all gates after it. The output-picture generator is therefore \(\hat{H}_i^{\rm out} = \hat{U}_i^{\rm post} \hat{G}_i \hat{U}_i^{{\rm post},\dagger}\), so \[\begin{align} \partial_{\theta_i}\hat{\rho}_\psi = - i \bigl[ \hat{H}_i^{\rm out}, \hat{\rho}_\psi \bigr]. \label{eq:state95derivative95main} \end{align}\tag{7}\] Only gates after the parameter, contained in \(\hat{U}_i^{\rm post}\), dress the generator; the cost observable then selects the part of that dressed operator that can contribute to the gradient.

Definition 3 (Effective gradient observable). For parameter \(\theta_i\) and cost observable \(\hat{O}\), define \[\begin{align} \hat{B}_i = i \bigl[ \hat{H}_i^{\rm out}, \hat{O} \bigr], \quad \mathrm{LC}(i) = \mathrm{supp}(\hat{B}_i), \label{eq:Bi95def95main} \end{align}\tag{8}\] where \(\hat{B}_i\) is the effective gradient observable; it is Hermitian and traceless because it is \(i\) times a commutator of Hermitian operators. The set \(\mathrm{LC}(i)\) is its effective light cone. Let \(r_i=\left|\mathrm{LC}(i)\right|\), \(d_i=2^{r_i}\), and let \(\hat{\sigma}_\psi^{(i)}=\operatorname{Tr}_{\overline{\mathrm{LC}(i)}}[\hat{\rho}_\psi]\) be the reduced output state on \(\mathrm{LC}(i)\). Here, \(\mathrm{supp}(\hat{X})\) denotes the smallest set of qubits outside which \(\hat{X}\) acts trivially, i.e., \(\hat{X}=\hat{X}_{\mathrm{supp}}\otimes\hat{\mathbb{1}}_{\overline{\mathrm{supp}}}\) up to identity factors.

Proposition 2 (Local two-copy gradient criterion). For Haar-product inputs, \[\begin{align} g_i(\psi) = \partial_{\theta_i}C = \operatorname{Tr}[\hat{B}_i \hat{\sigma}_\psi^{(i)}], \quad \mathbb{E}_\psi[g_i(\psi)]=0, \end{align}\] and \[\begin{align} \operatorname{Var}_\psi [g_i(\psi)] &=&\operatorname{Tr}\left[ \left( \hat{B}_i \otimes \hat{B}_i \right) \hat{M}_{i,\mathrm{LC}}^{(2)} \right], \nonumber \\ M_{i,\mathrm{LC}}^{(2)} &=&\mathbb{E}_\psi \left[ \bigl( \hat{\sigma}_\psi^{(i)} \bigr)^{\otimes 2} \right]. \label{eq:grad95var95local95main} \end{align}\qquad{(1)}\] The Haar benchmark on the light cone is \[\begin{align} \operatorname{Var}_{\mathrm{Haar},r_i}[g_i] = \frac{\operatorname{Tr}\bigl( \hat{B}_i^2 \bigr)}{d_i(d_i+1)} \le \frac{4\|{\hat{G}_i}\|_\infty^2 \|{\hat{O}}\|_\infty^2}{2^{r_i}+1}. \label{eq:haar95scale95main} \end{align}\qquad{(2)}\]

Proof sketch. Eq. (7 ) gives \(g_i=\operatorname{Tr}[\hat{B}_i \hat{\rho}_\psi]\). Since \(\hat{B}_i\) is supported only on \(\mathrm{LC}(i)\), this is the reduced-state contraction above. Squaring the linear functional gives \(g_i^2=\operatorname{Tr}\bigl[(\hat{B}_i \otimes \hat{B}_i)\bigl(\hat{\sigma}_\psi^{(i)}\bigr)^{\otimes2}\bigr]\). Averaging gives Eq. (?? ). The mean vanishes because the product-input first moment is maximally mixed and \(\operatorname{Tr}\hat{B}_i=0\). The Haar expression follows from the standard second moment \((\hat{\mathbb{1}}+\hat{S})/[d_i(d_i+1)]\) for a traceless observable. Details, including the dressed-generator construction, light-cone reduction, and local moment-gap form, are given in Secs. III A–III C of the Supplemental Material. ◻

Proposition 1 and Theorem 1 explain why a one-number global mean descriptor cannot determine trainability. EP is a single global contraction of \(\hat{M}^{(2)}_{\hat{U}}\). A gradient variance is a different contraction, local to \(\mathrm{LC}(i)\) and depending on \(i\) and \(\hat{O}\). EPD is not the trainability invariant either, but because it probes fourth-order fluctuation structure, it contains information discarded by EP.

Operationally, Eq. (?? ) says that the plateau question must be asked parameter by parameter. If \(\hat{O}\) is local and the forward cone of \(\hat{G}_i\) does not reach it, then \(\hat{B}_i=0\) and the gradient vanishes for a purely causal reason. If the cone is small, the Haar scale in Eq. (?? ) need not be exponentially small in the total system size. If the cone grows to \(\Theta(n)\) and the local two-copy contraction becomes Haar-like, the variance is exponentially suppressed. Thus, the relevant randomization is not global in the first instance; it is the randomization visible on the effective subsystem selected by the parameter and cost.

Equal EP does not imply equal trainability.—The hierarchy leads to a non-identifiability statement. Let \(\mathcal{W}_{\boldsymbol{\theta}}(\hat{E})\) denote a fixed trainable wrapper into which a fixed entangling block \(\hat{E}\) is inserted. The parameter locations, generators, input ensemble, and cost observable are held fixed; only the entangling block \(\hat{E}\) is changed. For such a comparison, define the parameter-resolved gradient-variance profile \[\begin{align} \Gamma_{\hat{E}}(\boldsymbol{\theta}) = \left( \operatorname{Var}_\psi[g_1^{(\hat{E})}(\psi,\boldsymbol{\theta})],\ldots, \operatorname{Var}_\psi[g_p^{(\hat{E})}(\psi,\boldsymbol{\theta})] \right), \label{eq:gamma95profile95main} \end{align}\tag{9}\] where \(g_k^{(\hat{E})}\) is the gradient of the common cost with respect to the \(k\)th parameter after inserting \(\hat{E}\).

Theorem 2 (Non-identifiability from mean EP). There exists a finite-dimensional PQC architecture and a fixed cost observable for which the trainability profile does not factor through the scalar mean descriptor EP. Equivalently, there are two entangling blocks \(\hat{E}\) and \(\hat{E}'\) in the same wrapper such that \[\begin{align} \mathrm{EP}(\mathcal{W}_{\boldsymbol{\theta}}(\hat{E}))=\mathrm{EP}(\mathcal{W}_{\boldsymbol{\theta}}(\hat{E}')) \quad \text{for all}~\boldsymbol{\theta}, \label{eq:equal95ep95formal95main} \end{align}\tag{10}\] whereas, for the same parameter vector (indeed for all \(\boldsymbol{\theta}\) in the constructive witness), \[\begin{align} \Gamma_{\hat{E}}(\boldsymbol{\theta}) \ne \Gamma_{\hat{E}'}(\boldsymbol{\theta}). \label{eq:unequal95gamma95formal95main} \end{align}\tag{11}\] Consequently, no function \(F:\mathbb{R}\to\mathbb{R}^p\) can satisfy \(\Gamma_{\hat{E}}(\boldsymbol{\theta}) = F(\mathrm{EP}(\mathcal{W}_{\boldsymbol{\theta}}(\hat{E})))\) on this architecture class. The witness can also be chosen so that the two circuits have different EPDs.

Proof sketch. Equality of EP fixes only the single global contraction \(\operatorname{Tr}[\hat{K}_Q \hat{M}^{(2)}_{\hat{U}}]\). In contrast, Eq. (?? ) shows that each gradient variance is a parameter- and cost-selected local contraction \(\operatorname{Tr}[(\hat{B}_i \otimes \hat{B}_i) \hat{M}_{i,\mathrm{LC}}^{(2)}]\). These contractions need not be determined by the EP contraction. The statement is made constructive in Sec. IV B of the Supplemental Material, where two locally inequivalent entangling blocks with identical Meyer-Wallach EP but different EPD are inserted into the same one-parameter wrapper and yield different gradient variances for the same cost. The resulting non-factorization through EP is formulated abstractly in Sec. IV C of the Supplemental Material. ◻

Theorem 2 rules out the strongest possible mean-descriptor picture. Equal EP means equal average entangling strength over product inputs; it does not fix the width of the output-entanglement distribution, nor does it fix how the local two-copy moment is oriented relative to a chosen generator and cost observable. Thus, two circuits can be indistinguishable by the coverage while lying in different trainability regimes.

This is also the role of EPD. EPD is not itself a barren-plateau invariant, because trainability remains local and cost-dependent. However, EPD is a strictly richer fixed-circuit descriptor than EP in the sense relevant here: it detects fluctuation structure that the mean discards and that can accompany a change in an actual gradient variance.

Figure 1: Ansatz trajectories in the EP-EPD two-dial plane for n=5,10,15,20. Curves denote ansatz families; markers trace L=1,\ldots,6, with labels at the first and final depths and arrows in the increasing-L direction. The axes are the coverage and variability dials in Eq. (12 ). Gray, green, and red mark the underexpressive, sweet-spot, and near-plateau regimes in Eq. (14 ), with thresholds fixed by Eq. (13 ).

Two-dial routes for PQC design.—We now use this hierarchy as a design protocol. For an ansatz family \(A\), increasing depth \(L\) traces a route whose two coordinates measure coverage and residual input-dependent variability. The objective is not to maximize depth, expressibility, or mean EP alone, but to locate a high-coverage, nonhomogenized window before the gradient-controlling local moments approach their Haar scale.

Fig. 1 implements this EP/EPD scan for five PQC ansatz families: ring-CZ HEA, line-CNOT HEA, brickwork RXX, QAOA-style ring-ZZ, and ring-CRX HEA. For \(n=5,10,15,20\) and \(L=1,\ldots,6\), we estimate \[\begin{align} \overline{\mathrm{EP}}_{A,L}^{(n)} &=& \mathbb{E}_{\boldsymbol{\theta} \sim \nu_{A,L}^{(n)}} \bigl[ \mathrm{EP}(\hat{U}_{A,L}^{(n)}(\boldsymbol{\theta})) \bigr], \nonumber \\ \overline{\mathrm{EPD}}_{A,L}^{(n)} &=&\mathbb{E}_{\boldsymbol{\theta} \sim \nu_{A,L}^{(n)}} \bigl[ \mathrm{EPD}(\hat{U}_{A,L}^{(n)}(\boldsymbol{\theta})) \bigr]. \end{align}\] Here, \(\nu_{A,L}^{(n)}\) denotes the parameter-sampling measure on the depth-\(L\) parameter space \(\Theta_{A,L}^{(n)}\) for ansatz family \(A\) and system size \(n\); in the route scan it is the product distribution over the circuit angles used to sample fixed circuit instances, i.e., the family-specific version of the ensemble \(\nu_L\) introduced above. The two plotted coordinates are \[\begin{align} x_{A,L}^{(n)}=\frac{\overline{\mathrm{EP}}_{A,L}^{(n)}}{Q_{\mathrm{Haar}}^{(n)}}, \qquad y_{A,L}^{(n)}=\frac{\overline{\mathrm{EPD}}_{A,L}^{(n)}}{\max_{A,L}\overline{\mathrm{EPD}}_{A,L}^{(n)}}, \label{eq:two95dial95coords95main} \end{align}\tag{12}\] where \(Q_{\mathrm{Haar}}^{(n)}=(2^n-2)/(2^n+1)\). Thus \(x\) is a coverage dial and \(y\) is the residual variability dial within the same \(n\) panel. Projecting onto \(x\) alone recovers the usual EP-vs-BP reading—larger mean entanglement as motion toward Haar-like behavior [8][10]—but cannot separate a useful high-coverage circuit from a homogenized one. The EPD dial resolves this ambiguity.

For reproducible region boundaries, we use the route data rather than hand-tuned constants. For a finite set \(Z\subset[0,1]\), sort the distinct values and consider midpoint cuts \(\tau\) between adjacent values. Each cut partitions \(Z\) into \(Z_{<\tau}\) and \(Z_{\ge\tau}\), with means \(\bar z_{<}\) and \(\bar z_{\ge}\), and assigns \[\begin{align} W_Z(\tau) = \sum_{z \in Z_{<\tau}} \left( z - \bar{z}_{<} \right)^2 + \sum_{z \in Z_{\ge\tau}}\left( z - \bar{z}_{\ge} \right)^2. \end{align}\] We define \(\tau_*(Z)=\arg\min_\tau W_Z(\tau)\), equivalently the Otsu/Fisher maximum-separation split [14], and set \[\begin{align} x_c^{(n)} = \tau_*\big(\{x_{A,L}^{(n)}\}_{A,L}\big), \quad y_c^{(n)} = \tau_*\big(Y_+^{(n)}\big), \label{eq:threshold95rule95main} \end{align}\tag{13}\] where \(Y_+^{(n)}=\{y_{A,L}^{(n)}:x_{A,L}^{(n)}\ge x_c^{(n)}\}\). The \(y\) cut is conditional on coverage so that low variability is interpreted as plateau-like homogenization only after the circuit has passed the coverage cut. Details and thresholds are in Sec. V-C of the Supplemental Material. We then define \[\begin{align} \mathcal{S}^{(n)}=\{(A,L):x_{A,L}^{(n)}\ge x_c^{(n)}, \;y_{A,L}^{(n)}\ge y_c^{(n)}\}. \label{eq:sweet95region95main} \end{align}\tag{14}\] The green sweet spot in Fig. 1 is \(\mathcal{S}^{(n)}\): coverage has passed the cut while EPD remains open; gray is underexpressive, and red is high-coverage but homogenized. We rank candidates by \[\begin{align} s_{A,L}^{(n)}=x_{A,L}^{(n)}y_{A,L}^{(n)}, \label{eq:sweet95score95main} \end{align}\tag{15}\] which rewards simultaneous coverage and noncollapsed variability. Thresholds and best-score depths are in Table I of Sec. V C of the Supplemental Material. Only the EPD dial separates the green high-coverage/nonhomogenized window from the red high-coverage/homogenized sector.

Definition 4 (Two-dial trainability window). For a chosen ansatz family and cost class, a practical trainability window is a depth interval in which \((A,L)\in\mathcal{S}^{(n)}\), or more generally in which the coverage is already high while \(\overline{\mathrm{EPD}}_{A,L}^{(n)}\) has not yet collapsed relative to its early-depth scale.

This definition is a screening principle, not a replacement for the local gradient criterion: the actual variance is still selected by the parameter light cone and cost observable in Eq. (?? ). The scan says where task-specific optimization is most worthwhile—away from the underexpressive gray and already homogenized red regions, and toward high-coverage points whose EPD dial remains open.

A consistency check in Sec. V-C of the Supplemental Material compares normalized EPD with representative gradient variance. The relation is not universal because Eq. (?? ) is local and cost-dependent, but low EPD broadly coincides with smaller gradients, especially at larger \(n\).

Figure 2: Task-level validation on the n=5 route plane. Points are ansatz-depth pairs, color is empirical success probability, and dashed lines are x_c,y_c. High coverage with low EPD is less favorable for training.
Figure 3: Region-wise diagnostics: success probability, first-success iteration (lower is faster), and initial SPSA gradient scale. The sweet region is the fastest high-coverage region and outperforms near-plateau in success probability.

Task-level validation.—We finally test the descriptor-defined window in a descriptor-blind \(n=5\) teacher–student QML benchmark. The EP/EPD coordinates, region labels, and sweet scores are fixed before training; all \(30\) ansatz-depth candidates use the same embedding, readout, loss, and Adam-SPSA budget for two sweet-region teacher tasks.

Figs. 2 and 3 show that, within the high-coverage subset, the sweet region outperforms the low-EPD near-plateau sector in accuracy, loss, success probability, and first-success time, reaching the threshold after a median of \(5\) iterations versus \(20\) near the plateau. Thus, the sweet score can be a first-pass design rule. Full protocol, result tables, and analysis are in Sec. VI of the Supplemental Material.

Discussion.—We have introduced an EP/EPD framework that separates three objects often compressed into a single expressibility–trainability narrative: EP is a global two-copy mean, EPD is a global four-copy fluctuation descriptor, and gradient variance is a local two-copy contraction selected by a parameter light cone and cost observable. This hierarchy proves that equal EP does not imply equal EPD or equal gradient variance, and Fig. 1 shows that Haar-like coverage and BP-like homogenization are distinct finite-depth crossovers.

The implication is that expressibility-vs-trainability is not a hard one-dimensional tradeoff. EP reports coverage, but only EPD tells whether high-coverage circuits still retain input-dependent variability. The EP/EPD scan therefore turns the usual EP-vs-BP tension into an ansatz-design rule: choose blocks, connectivities, and depth schedules that move far enough along the coverage dial while keeping the variability dial open. This matters for variational algorithms, hybrid heuristics, and quantum machine learning, where useful coverage must coexist with finite-shot gradients.

Acknowledgments.—This work was supported by the Ministry of Science and ICT/National Research Foundation of Korea (RS-2023-NR119924, RS-2024-00432214, RS-2025-03532992, and RS-2025-18362970), the IITP grant funded by the Korean government (RS-2019-II190003, RS-2026-25519864), and the Korean ARPA-H Project through KHIDI, funded by the Ministry of Health & Welfare, Republic of Korea (RS-2025-25456722). We thank the Yonsei University Quantum Computing Project Group for support and access to Quantum System One.

Supplemental Material for “Breaking the One-Dimensional Expressibility–Trainability Tradeoff”↩︎

1 Ensembles, fixed-circuit descriptors, and the spectral obstruction↩︎

This Supplemental Material follows the notation and proof of the original manuscript. Its purpose is threefold. First, it gives the detailed derivations behind the \(2\)-copy/\(4\)-copy hierarchy for fixed-circuit entangling power (EP) SM-Zanardi2000EP? and entangling-power deviation (EPD) SM-ChoBang2026EPD?. Second, it gives the local light-cone proof of the gradient-variance criterion used in the main manuscript. Third, it records the numerical protocols, supplementary diagnostics, and numerical analysis.

A useful way to read this Supplemental Material is to keep two orderings distinct. The first is the copy order of the moment: EP is determined by a two-copy contraction, EPD by a four-copy contraction, and the gradient variance by a local two-copy contraction. The second is the ensemble order: one may first average over product inputs for a fixed circuit and only then average over circuit parameters, or one may form Sim-style descriptors directly from parameter-sampled output states. These operations do not define the same object. The main manuscript uses this distinction to argue that a one-dimensional expressibility narrative is incomplete, while the present document supplies the algebra behind that statement.

Throughout, the \(n\)-qubit Hilbert space is \[\begin{align} \mathcal{H}_n=(\mathbb{C}^2)^{\otimes n}, \label{eq:sm95Hn} \end{align}\tag{16}\] and a depth-\(L\) parameterized circuit is denoted by \[\begin{align} \hat{U}_L(\boldsymbol{\theta})\in U(2^n), \qquad \boldsymbol{\theta}\in\Theta_L\subset\mathbb{R}^{p_L}. \label{eq:sm95pqc} \end{align}\tag{17}\] Here, \(\boldsymbol{\theta}=(\theta_1,\ldots,\theta_{p_L})\) is a single parameter vector whose \(p_L\) scalar components are the independent trainable parameters of the depth-\(L\) ansatz. The number \(p_L\) may depend on the depth, and \(\Theta_L\) denotes the allowed joint domain of these parameters within the parameter space \(\mathbb{R}^{p_L}\). Two randomness levels are used. The parameter ensemble \(\nu_L\) samples a circuit instance and underlies Sim-style expressibility SM-Sim2019expressibility?. The product-input ensemble used for fixed-circuit EP and EPD is instead \[\begin{align} \mu_{\mathrm{prod}}=\bigotimes_{j=1}^n\mu_{\mathrm H}^{(1)}, \qquad \left|\psi\right\rangle=\bigotimes_{j=1}^n\left|\psi_j\right\rangle, \qquad \left|\psi_j\right\rangle\sim\mu_{\mathrm H}^{(1)}. \label{eq:sm95product95ensemble} \end{align}\tag{18}\] Thus, expressibility samples many circuit instances from a fixed reference input, whereas EP/EPD probe one fixed circuit with many product inputs.

For a pure state \(\left|\phi\right\rangle\), the Meyer-Wallach global entanglement measure is SM-MeyerWallach2002? \[\begin{align} Q(\left|\phi\right\rangle)=\frac{2}{n}\sum_{j=1}^{n}\left(1-\operatorname{Tr}\hat{\rho}_j^2\right), \qquad \hat{\rho}_j=\operatorname{Tr}_{\bar j}\left|\phi\right\rangle\!\!\left\langle \phi\right|. \label{eq:sm95Q95def} \end{align}\tag{19}\] Here, \(\bar j:=\{1,\ldots,n\}\setminus\{j\}\) denotes the complement of qubit \(j\). Thus, \(\operatorname{Tr}_{\bar j}\) traces out all qubits except \(j\) and leaves the one-qubit reduced density operator \(\hat{\rho}_j\). Since \(1/2\le \operatorname{Tr}\hat{\rho}_j^2\le1\), one has \(0\le Q\le1\). Also, \(Q(\left|\phi\right\rangle)=0\) if and only if \(\left|\phi\right\rangle\) is fully separable: if \(Q=0\), every one-qubit reduction is pure; purity of a subsystem of a global pure state implies factorization across the corresponding bipartition; iterating over all sites gives full product structure.

For a fixed unitary \(\hat{U}\), we define \[\begin{align} \mathrm{EP}(\hat{U}) := \mathbb{E}_{\psi\sim\mu_{\mathrm{prod}}}\left[Q(\hat{U}\left|\psi\right\rangle)\right], \qquad \mathrm{EPD}(\hat{U}) := \sqrt{\operatorname{Var}_{\psi\sim\mu_{\mathrm{prod}}}\left[Q(\hat{U}\left|\psi\right\rangle)\right]}. \label{eq:sm95EP95EPD95def} \end{align}\tag{20}\] If \(\hat{V}_{\rm in}, \hat{V}_{\rm out}\in U(2)^{\otimes n}\) are local unitaries, then \[\begin{align} \mathrm{EP}(\hat{V}_{\rm out} \hat{U}\hat{V}_{\rm in})=\mathrm{EP}(\hat{U}), \qquad \mathrm{EPD}(\hat{V}_{\rm out} \hat{U}\hat{V}_{\rm in})=\mathrm{EPD}(\hat{U}). \label{eq:sm95local95invariance} \end{align}\tag{21}\] The post-multiplication by \(\hat{V}_{\rm out}\) preserves \(Q\), while the pre-multiplication by \(\hat{V}_{\rm in}\) preserves the product-Haar measure.

For \(t\ge1\), we define the fixed-circuit output moment \[\begin{align} \hat{M}^{(t)}_{\hat{U}}:= \mathbb{E}_{\psi\sim\mu_{\mathrm{prod}}} \left[\left(\hat{U}\left|\psi\right\rangle\!\!\left\langle \psi\right|\hat{U}^\dagger\right)^{\otimes t}\right]. \label{eq:sm95MUt} \end{align}\tag{22}\] For one qubit, \[\begin{align} \int d\mu_{\mathrm H}^{(1)}(\psi)(\left|\psi\right\rangle\!\!\left\langle \psi\right|)^{\otimes t} = \frac{\hat{\Pi}^{(t,1)}_{\rm sym}}{t+1}, \label{eq:sm95one95qubit95moment} \end{align}\tag{23}\] where \(\hat{\Pi}^{(t,1)}_{\rm sym}\) projects onto \(\mathrm{Sym}^t(\mathbb{C}^2)\). Grouping together, the \(t\) copies of each physical qubit gives \[\begin{align} \hat{M}^{(t)}_{\rm prod} = \frac{1}{(t+1)^n}\hat{\Pi}^{(t)}_{\rm loc}, \qquad \hat{\Pi}^{(t)}_{\rm loc}=\bigotimes_{j=1}^n \hat{\Pi}^{(t)}_{{\rm sym},j}. \label{eq:sm95prod95moment} \end{align}\tag{24}\] The full Haar pure-state moment on \(d=2^n\) dimensions is \[\begin{align} \hat{M}^{(t)}_{\rm Haar} =\frac{1}{D_{n,t}}\hat{\Pi}^{(t)}_{\rm glob}, \qquad D_{n,t}=\binom{2^n+t-1}{t}, \label{eq:sm95haar95moment} \end{align}\tag{25}\] where \(\hat{\Pi}^{(t)}_{\rm glob}\) projects onto \(\mathrm{Sym}^t(\mathcal{H}_n)\).

Proposition 3 (Spectral obstruction). For every fixed unitary \(\hat{U}\) and every \(t\ge1\), \[\begin{align} \hat{M}^{(t)}_{\hat{U}}=\hat{U}^{\otimes t}\hat{M}^{(t)}_{\rm prod}\hat{U}^{\dagger\otimes t}. \label{eq:sm95spectral95orbit} \end{align}\qquad{(3)}\] Therefore, the spectrum of \(\hat{M}^{(t)}_{\hat{U}}\) is independent of \(\hat{U}\). If \(n>1\) and \(t\ge2\), then \(\hat{M}^{(t)}_{\rm prod}\) and \(\hat{M}^{(t)}_{\rm Haar}\) do not have the same spectrum.

Proof. —Eq. (?? ) follows by linearity: \[\begin{align} \hat{M}^{(t)}_{\hat{U}} &=&\mathbb{E}_\psi\left[\hat{U}^{\otimes t}(\left|\psi\right\rangle\!\!\left\langle \psi\right|)^{\otimes t}\hat{U}^{\dagger\otimes t}\right] =\hat{U}^{\otimes t}\hat{M}^{(t)}_{\rm prod}\hat{U}^{\dagger\otimes t}. \end{align}\] Thus \(\hat{M}^{(t)}_{\hat{U}}\) is unitarily conjugate to \(\hat{M}^{(t)}_{\rm prod}\).

By Eqs. (24 ) and (25 ), the two reference moments are proportional to projectors with support dimensions \[\begin{align} \dim \left( \mathop{\mathrm{supp}}(\hat{M}^{(t)}_{\rm prod}) \right) = (t+1)^n, \quad \dim \left( \mathop{\mathrm{supp}}(\hat{M}^{(t)}_{\rm Haar}) \right) = \binom{2^n+t-1}{t}. \label{eq:sm95rank95compare} \end{align}\tag{26}\] The locally symmetric subspace is strictly contained in the globally symmetric subspace when \(n>1\) and \(t\ge2\). To see strictness, take the vector obtained by symmetrizing one copy of \(\left|110\cdots 0\right\rangle\) among \(t\) copies while all other copies are \(\left|0^n\right\rangle\). This vector is globally copy-symmetric, but it is not invariant under swapping only the first-qubit factors of two copies. Hence, the supports, ranks, and spectra differ. ◻

This proposition is the reason that Haar is used in the main manuscript as a scalar entanglement reference and as a local-moment benchmark, not as an operator-level limit of fixed-circuit product-input moments. More generally, let \(\mathcal{H}=\mathcal{H}_A\otimes\mathcal{H}_B\), with \(d_A=\dim\mathcal{H}_A\) and \(d_B=\dim\mathcal{H}_B\), and let \(\hat{\rho}_A=\operatorname{Tr}_B\left|\phi\right\rangle\!\!\left\langle \phi\right|\) for a Haar-random \(\left|\phi\right\rangle\in\mathcal{H}\). Lubkin’s formula states \(\mathbb{E}_{\phi\sim\mathrm{Haar}}[\operatorname{Tr}(\hat{\rho}_A^2)]=(d_A+d_B)/(d_Ad_B+1)\) SM-Lubkin1978?. Indeed, the Haar two-copy moment \(\mathbb{E}[\left|\phi\right\rangle\!\!\left\langle \phi\right|^{\otimes2}]=(\hat{\mathbb{1}}+\hat{S}_A\hat{S}_B)/[d_Ad_B(d_Ad_B+1)]\), together with \(\operatorname{Tr}(\hat{\rho}_A^2)=\operatorname{Tr}[\hat{S}_A\left|\phi\right\rangle\!\!\left\langle \phi\right|^{\otimes2}]\), gives this result by a single swap contraction. Taking \(d_A=2\) and \(d_B=2^{n-1}\) yields the one-qubit-versus-rest value \[\begin{align} \mathbb{E}_{\phi\sim\mathrm{Haar}}\left[\operatorname{Tr}\hat{\rho}_j^2\right] =\frac{2+2^{n-1}}{2^n+1}. \label{eq:sm95lubkin} \end{align}\tag{27}\] Substituting this into Eq. (19 ) gives \[\begin{align} Q_{\mathrm{Haar}} = \mathbb{E}_{\phi\sim\mathrm{Haar}}[Q(\left|\phi\right\rangle)] = \frac{2^n-2}{2^n+1}. \label{eq:sm95Q95Haar} \end{align}\tag{28}\]

2 Exact two-copy and four-copy formulas for EP and EPD↩︎

Let \(\hat{S}_j\) denote the unitary that swaps qubit \(j\) between two copies and acts trivially elsewhere. The following identity is repeatedly used to convert local purities into two-copy expectation values. If \(\hat{A},\hat{B}\) are operators on \(\mathcal{H}_n\), then \[\begin{align} \operatorname{Tr}[\hat{S}_j(\hat{A}\otimes \hat{B})] = \operatorname{Tr}\left[(\operatorname{Tr}_{\bar j}\hat{A})(\operatorname{Tr}_{\bar j}\hat{B})\right]. \label{eq:sm95swap95identity} \end{align}\tag{29}\] To verify Eq. (29 ) explicitly, write basis states as \(\left|a,\mu\right\rangle=\left|a\right\rangle_j\left|\mu\right\rangle_{\bar j}\) (or \(\left|b,\nu\right\rangle=\left|b\right\rangle_j\left|\nu\right\rangle_{\bar j}\)), where \(a,b\in\{0,1\}\) are single-qubit labels and \(\mu,\nu\) label an orthonormal basis of the complement. The swap acts as \[\begin{align} \hat{S}_j(\left|a,\mu\right\rangle\otimes\left|b,\nu\right\rangle)=\left|b,\mu\right\rangle\otimes\left|a,\nu\right\rangle. \end{align}\] Therefore, \[\begin{align} \operatorname{Tr}[\hat{S}_j(\hat{A}\otimes \hat{B})] &=&\sum_{a,b,\mu,\nu} \left(\left\langle a,\mu\right|\otimes\left\langle b,\nu\right|\right) \hat{S}_j (\hat{A}\otimes \hat{B}) \left(\left|a,\mu\right\rangle\otimes\left|b,\nu\right\rangle\right) \nonumber\\ &=&\sum_{a,b,\mu,\nu} \left\langle b,\mu\right|\hat{A}\left|a,\mu\right\rangle \left\langle a,\nu\right|\hat{B}\left|b,\nu\right\rangle \nonumber\\ &=&\sum_{a,b}\left\langle b\right|(\operatorname{Tr}_{\bar j}\hat{A})\left|a\right\rangle\left\langle a\right|(\operatorname{Tr}_{\bar j}\hat{B})\left|b\right\rangle \nonumber\\ &=&\operatorname{Tr}[(\operatorname{Tr}_{\bar j}\hat{A})(\operatorname{Tr}_{\bar j}\hat{B})]. \end{align}\] For \(\hat{\rho}=\left|\phi\right\rangle\!\!\left\langle \phi\right|\), Eq. (29 ) yields \[\begin{align} \operatorname{Tr}[\hat{S}_j\hat{\rho}^{\otimes2}]=\operatorname{Tr}\hat{\rho}_j^2. \end{align}\] Therefore, with \[\begin{align} \hat{K}_Q=\frac{2}{n}\sum_{j=1}^n(\hat{\mathbb{1}}-\hat{S}_j), \label{eq:sm95KQ} \end{align}\tag{30}\] one has the exact two-copy expression \[\begin{align} Q(\left|\phi\right\rangle)=\operatorname{Tr}[\hat{K}_Q\hat{\rho}^{\otimes2}]. \label{eq:sm95Q95two95copy} \end{align}\tag{31}\] On four copies, let \(\hat{K}_Q^{(12)}\) and \(\hat{K}_Q^{(34)}\) act on copy pairs \((1,2)\) and \((3,4)\), respectively. Define \[\begin{align} \hat{T}_Q=\hat{K}_Q^{(12)}\hat{K}_Q^{(34)} =\frac{4}{n^2}\sum_{j,k=1}^n(\hat{\mathbb{1}}-\hat{S}_j^{(12)})(\hat{\mathbb{1}}-\hat{S}_k^{(34)}). \label{eq:sm95TQ} \end{align}\tag{32}\] Since \(\hat{\rho}^{\otimes4}=\hat{\rho}^{\otimes2}_{12}\otimes\hat{\rho}^{\otimes2}_{34}\), \[\begin{align} Q(\left|\phi\right\rangle)^2=\operatorname{Tr}[\hat{T}_Q\hat{\rho}^{\otimes4}]. \label{eq:sm95Q95four95copy} \end{align}\tag{33}\]

Theorem 3 (Exact moment representation). For a fixed circuit \(\hat{U}\), \[\begin{align} \mathrm{EP}(\hat{U})=\operatorname{Tr}[\hat{K}_Q\hat{M}^{(2)}_{\hat{U}}], \label{eq:sm95EP95two95copy} \end{align}\tag{34}\] and \[\begin{align} \mathrm{EPD}(\hat{U})^2 =\operatorname{Tr}[\hat{T}_Q\hat{M}^{(4)}_{\hat{U}}]-\left(\operatorname{Tr}[\hat{K}_Q\hat{M}^{(2)}_{\hat{U}}]\right)^2. \label{eq:sm95EPD95four95copy} \end{align}\tag{35}\]

Proof. —For \(\hat{\rho}_\psi(\hat{U})=\hat{U}\left|\psi\right\rangle\!\!\left\langle \psi\right|\hat{U}^\dagger\), Eq. (31 ) gives \[\begin{align} Q(\hat{U}\left|\psi\right\rangle)=\operatorname{Tr}[\hat{K}_Q\hat{\rho}_\psi(\hat{U})^{\otimes2}]. \end{align}\] Averaging and interchanging trace and expectation gives Eq. (34 ). Similarly, Eq. (33 ) gives \[\begin{align} \mathbb{E}_\psi[Q(\hat{U}\left|\psi\right\rangle)^2]=\operatorname{Tr}[\hat{T}_Q\hat{M}^{(4)}_{\hat{U}}]. \end{align}\] Subtracting the square of the mean gives Eq. (35 ). ◻

The theorem is the precise sense in which EP and EPD probe different moment orders. Equality of EP values constrains only one scalar contraction of a two-copy moment and does not determine the four-copy fluctuation profile.

For completeness, we record the stability bounds used implicitly in the main manuscript. For two circuits \(\hat{U},\hat{V}\), \[\begin{align} \left|\mathrm{EP}(\hat{U})-\mathrm{EP}(\hat{V})\right| \le \left\|\hat{K}_Q\right\|_2\left\|\hat{M}^{(2)}_{\hat{U}}-\hat{M}^{(2)}_{\hat{V}}\right\|_2, \label{eq:sm95EP95stability} \end{align}\tag{36}\] and \[\begin{align} \left|\mathrm{EPD}(\hat{U})^2-\mathrm{EPD}(\hat{V})^2\right| \le \left\|\hat{T}_Q\right\|_2\left\|\hat{M}^{(4)}_{\hat{U}}-\hat{M}^{(4)}_{\hat{V}}\right\|_2 +2\left\|\hat{K}_Q\right\|_2\left\|\hat{M}^{(2)}_{\hat{U}}-\hat{M}^{(2)}_{\hat{V}}\right\|_2. \label{eq:sm95EPD95stability} \end{align}\tag{37}\] The first inequality is Hilbert-Schmidt Cauchy-Schwarz. The second follows by applying the same inequality to the four-copy term and bounding the change in the squared mean by \(2\left|\mathrm{EP}(\hat{U})-\mathrm{EP}(\hat{V})\right|\), since \(0\le\mathrm{EP}\le1\).

Expanding \(\operatorname{Tr}\hat{K}_Q^2\) gives \[\begin{align} \left\|\hat{K}_Q\right\|_2^2=4^n\frac{n+3}{n}, \qquad \left\|\hat{T}_Q\right\|_2=4^n\frac{n+3}{n}. \label{eq:sm95norms} \end{align}\tag{38}\] Indeed, \[\begin{align} \left\|\hat{K}_Q\right\|_2^2 =\operatorname{Tr}\hat{K}_Q^2 =\frac{4}{n^2}\sum_{j,k=1}^n\operatorname{Tr}[(\hat{\mathbb{1}}-\hat{S}_j)(\hat{\mathbb{1}}-\hat{S}_k)]. \end{align}\] If \(j=k\), then \((\hat{\mathbb{1}}-\hat{S}_j)^2=2(\hat{\mathbb{1}}-\hat{S}_j)\) and \[\begin{align} \operatorname{Tr}[(\hat{\mathbb{1}}-\hat{S}_j)^2]=2\operatorname{Tr}(\hat{\mathbb{1}}-\hat{S}_j)=2(4^n-2\cdot4^{n-1})=4^n. \end{align}\] If \(j\ne k\), the trace factorizes over the two nontrivial site pairs and the remaining \(n-2\) sites: \[\begin{align} \operatorname{Tr}[(\hat{\mathbb{1}}-\hat{S}_j)(\hat{\mathbb{1}}-\hat{S}_k)] &=& \operatorname{Tr}[\hat{\mathbb{1}}-\hat{S}_j-\hat{S}_k+\hat{S}_j\hat{S}_k] \nonumber\\ &=& 4^n-2\cdot4^{n-1}-2\cdot4^{n-1}+2\cdot2\cdot4^{n-2} \nonumber\\ &=& 2\cdot2\cdot4^{n-2}=4^{n-1}. \end{align}\] Therefore, \[\begin{align} \left\|\hat{K}_Q\right\|_2^2=\frac{4}{n^2}\left(n4^n+n(n-1)4^{n-1}\right)=4^n\frac{n+3}{n}. \end{align}\] Because \(\hat{T}_Q=\hat{K}_Q\otimes \hat{K}_Q\) under the copy grouping \((12)|(34)\), the Hilbert-Schmidt norm is multiplicative and \(\left\|\hat{T}_Q\right\|_2=\left\|\hat{K}_Q\right\|_2^2\).

3 Gradient Variance and the Local 2-Copy Criterion for BP↩︎

We now turn to trainability. The main purpose of this section is to isolate, in exact algebraic form, the object that controls the second moment of gradients. This is where the usual expressibility–vs–barren plateau (BP) narrative must be sharpened. The quantity \(\mathrm{EP}(\hat{U})\) derived in Sec. 2 is a global two-copy scalar functional of a fixed circuit, while \(\mathrm{EPD}(\hat{U})\) is a global four-copy scalar functional. By contrast, the variance of the gradient with respect to a specific parameter is controlled not by any global scalar descriptor, but by a local two-copy moment associated with the parameter’s effective light cone.

This is the precise sense in which the barren plateaus are fundamentally a local second-moment phenomenon. In the language of McClean-type results, one often says that the gradients vanish when the relevant circuit becomes sufficiently random or design-like SM-McClean2018barrenplateaus?, SM-Cerezo2022challenges?. What the present section makes explicit is that the operative notion is not a global one-number summary such as expressibility or average entangling power, but the second-copy structure seen by each parameter through its causal cone.

Throughout this section, we consider a differentiable \(n\)-qubit PQC \[\begin{align} \hat{U}(\boldsymbol{\theta}) = \hat{U}_m(\boldsymbol{\theta})\cdots \hat{U}_2(\boldsymbol{\theta})\hat{U}_1(\boldsymbol{\theta}), \label{eq:BP95circuit95seq} \end{align}\tag{39}\] built from geometrically local gates on a fixed connectivity graph. For notational clarity, we assume that each parameter \(\theta_i\) appears in a single elementary gate. If the same parameter controls several gates, the total gradient is the sum of the corresponding single-gate contributions, and all arguments below apply termwise.

We study a cost function of the form \[\begin{align} C(\boldsymbol{\theta},\psi) = \operatorname{Tr}\left[ \hat{O}\,\hat{\rho}_{\psi}(\boldsymbol{\theta}) \right], \qquad \hat{\rho}_{\psi}(\boldsymbol{\theta}) := \hat{U}(\boldsymbol{\theta})\left|\psi\right\rangle\!\!\left\langle \psi\right|\hat{U}(\boldsymbol{\theta})^{\dagger}, \label{eq:BP95cost} \end{align}\tag{40}\] where \(\hat{O}=\hat{O}^{\dagger}\) is a bounded observable, which we normalize as \[\begin{align} \left\|\hat{O}\right\|_{\infty}\le 1. \end{align}\] The input state is again sampled from the Haar-product ensemble introduced in Sec. 1, \[\begin{align} \left|\psi\right\rangle = \bigotimes_{j=1}^{n}\left|\psi_j\right\rangle, \qquad \left|\psi_j\right\rangle\sim \mu_{\mathrm{Haar}}^{(1)} \quad \text{independently}. \label{eq:BP95input} \end{align}\tag{41}\] For a fixed parameter index \(i\), we denote the corresponding gradient by \[\begin{align} g_i(\psi) := \partial_{\theta_i} C(\boldsymbol{\theta},\psi). \label{eq:def95gi} \end{align}\tag{42}\]

3.1 Dressed generators and effective light cones↩︎

Let the \(i\)th parameter enter the circuit through a gate of the form \[\begin{align} \hat{U}_i(\theta_i)=e^{-i\theta_i \hat{G}_i}, \label{eq:param95gate} \end{align}\tag{43}\] where \(\hat{G}_i=\hat{G}_i^{\dagger}\) is a Hermitian generator supported on a bounded set of qubits \(\Lambda_i\). We decompose the total circuit around this gate as \[\begin{align} \hat{U}(\boldsymbol{\theta}) = \hat{U}_i^{\rm post}(\boldsymbol{\theta})\,e^{-i\theta_i \hat{G}_i}\,\hat{U}_i^{\rm pre}(\boldsymbol{\theta}), \label{eq:circuit95split} \end{align}\tag{44}\] where \(\hat{U}_i^{\rm pre}\) collects all gates before the parameterized gate and \(\hat{U}_i^{\rm post}\) collects all gates after it, matching the pre/post notation used in the main manuscript.

The key output-picture object is the dressed generator \[\begin{align} \hat{H}_i^{\mathrm{out}} := i\, \left( \partial_{\theta_i}\hat{U}(\boldsymbol{\theta}) \right) \hat{U}(\boldsymbol{\theta})^{\dagger}. \label{eq:def95Hout} \end{align}\tag{45}\] Using Eq. (44 ), we obtain the exact identity \[\begin{align} \hat{H}_i^{\mathrm{out}} = \hat{U}_i^{\rm post}(\boldsymbol{\theta})\,\hat{G}_i\,\hat{U}_i^{\rm post}(\boldsymbol{\theta})^{\dagger}. \label{eq:Hout95explicit} \end{align}\tag{46}\] In particular, \(\hat{H}_i^{\mathrm{out}}\) is Hermitian and \[\begin{align} \left\|\hat{H}_i^{\mathrm{out}}\right\|_{\infty} = \left\|\hat{G}_i\right\|_{\infty}. \label{eq:Hout95norm} \end{align}\tag{47}\]

Eq. (46 ) shows that only the gates after the parameter dress the generator in the output picture. Because the circuit is local, the support of \(\hat{H}_i^{\mathrm{out}}\) can be obtained recursively by propagating the support \(\Lambda_i\) through later gates: conjugation by a gate disjoint from the current support leaves the support unchanged, whereas conjugation by a gate overlapping the current support enlarges it at most to the union of the two supports. The resulting output subset is the usual forward light cone of the parameter. We denote it by \[\begin{align} \mathcal{F}(i) := \operatorname{supp}\left(\hat{H}_i^{\mathrm{out}}\right). \label{eq:def95forward95cone} \end{align}\tag{48}\]

Taking the derivative of the output state gives \[\begin{align} \partial_{\theta_i}\hat{\rho}_{\psi}(\boldsymbol{\theta}) &=& \left(\partial_{\theta_i}\hat{U}\right)\left|\psi\right\rangle\!\!\left\langle \psi\right|\hat{U}^{\dagger} + \hat{U}\left|\psi\right\rangle\!\!\left\langle \psi\right|\left(\partial_{\theta_i}\hat{U}^{\dagger}\right) \nonumber\\ &=& -i\hat{H}_i^{\mathrm{out}}\hat{\rho}_{\psi}(\boldsymbol{\theta}) + i\hat{\rho}_{\psi}(\boldsymbol{\theta})\hat{H}_i^{\mathrm{out}} \nonumber\\ &=& -i\left[ \hat{H}_i^{\mathrm{out}}, \hat{\rho}_{\psi}(\boldsymbol{\theta}) \right]. \label{eq:rho95derivative95comm} \end{align}\tag{49}\] Hence, the gradient can be written as \[\begin{align} g_i(\psi) &=& \operatorname{Tr}\left[ \hat{O}\,\partial_{\theta_i}\hat{\rho}_{\psi}(\boldsymbol{\theta}) \right] = -i\operatorname{Tr}\left[ \hat{O}\left(\hat{H}_i^{\mathrm{out}}\hat{\rho}_{\psi}-\hat{\rho}_{\psi}\hat{H}_i^{\mathrm{out}}\right) \right] \nonumber\\ &=& \operatorname{Tr}\left[ i\left[\hat{H}_i^{\mathrm{out}},\hat{O}\right]\hat{\rho}_{\psi}(\boldsymbol{\theta}) \right]. \label{eq:gradient95commutator95output} \end{align}\tag{50}\]

This motivates the following definition.

Definition 5 (Effective gradient observable and effective light cone). For the parameter \(\theta_i\), we define the effective gradient observable \[\begin{align} \hat{B}_i := i\left[ \hat{H}_i^{\mathrm{out}},\hat{O} \right]. \label{eq:def95Bi} \end{align}\tag{51}\] Its support \[\begin{align} \mathrm{LC}(i) := \operatorname{supp}(\hat{B}_i) \label{eq:def95LCi} \end{align}\tag{52}\] is called the effective light cone of the parameter with respect to the chosen cost operator. We denote its size by \[\begin{align} r_i := |\mathrm{LC}(i)|, \qquad d_i := 2^{r_i}. \label{eq:def95ri95di} \end{align}\tag{53}\]

By construction, \[\begin{align} \mathrm{LC}(i) \subseteq \mathcal{F}(i)\cup \operatorname{supp}(\hat{O}). \label{eq:LC95subset} \end{align}\tag{54}\] In particular, if \(\hat{O}\) is local and its support is disjoint from the forward light cone \(\mathcal{F}(i)\), then \([\hat{H}_i^{\mathrm{out}},\hat{O}]=0\) and the corresponding gradient vanishes identically. Thus, the relevant subsystem for trainability is not the entire circuit, but the causal overlap between the parameter and the cost operator.

The observable \(\hat{B}_i\) inherits several basic structural properties from Eq. (51 ).

Lemma 1 (Basic properties of \(\hat{B}_i\)). For each parameter index \(i\), the operator \(\hat{B}_i\) defined in Eq. (51 ) is Hermitian, traceless, and satisfies \[\begin{align} \left\|\hat{B}_i\right\|_{\infty} \le 2\left\|\hat{G}_i\right\|_{\infty}\left\|\hat{O}\right\|_{\infty}. \label{eq:Bi95norm95bound} \end{align}\tag{55}\]

Proof. —Hermiticity follows from \[\begin{align} \hat{B}_i^{\dagger} = \left( i[\hat{H}_i^{\mathrm{out}},\hat{O}] \right)^{\dagger} = -i[\hat{O},\hat{H}_i^{\mathrm{out}}] = i[\hat{H}_i^{\mathrm{out}},\hat{O}] = \hat{B}_i, \end{align}\] since both \(\hat{H}_i^{\mathrm{out}}\) and \(\hat{O}\) are Hermitian. Tracelessness follows from the cyclicity of trace: \[\begin{align} \operatorname{Tr}(\hat{B}_i) = i\operatorname{Tr}\left(\hat{H}_i^{\mathrm{out}}\hat{O}-\hat{O}\hat{H}_i^{\mathrm{out}}\right) = 0. \end{align}\] Finally, we have \[\begin{align} \left\|\hat{B}_i\right\|_{\infty} = \left\|i[\hat{H}_i^{\mathrm{out}},\hat{O}]\right\|_{\infty} \le \left\|\hat{H}_i^{\mathrm{out}}\hat{O}\right\|_{\infty} + \left\|\hat{O}\hat{H}_i^{\mathrm{out}}\right\|_{\infty} \le 2\left\|\hat{H}_i^{\mathrm{out}}\right\|_{\infty}\left\|\hat{O}\right\|_{\infty} = 2\left\|\hat{G}_i\right\|_{\infty}\left\|\hat{O}\right\|_{\infty}, \end{align}\] where Eq. (47 ) was used in the last step. ◻

3.2 Exact local 2-copy formula for gradient moments↩︎

We now prove the central structural statement of this section. The gradient associated with a given parameter is a linear functional of the reduced output state on its effective light cone, and its second moment is determined exactly by the corresponding local two-copy moment.

Proposition 4 (Light-cone reduction of gradient moments). Let \(g_i(\psi)=\partial_{\theta_i}C(\boldsymbol{\theta},\psi)\) be the gradient with respect to the parameter \(\theta_i\), and let \(\hat{B}_i\) and \(\mathrm{LC}(i)\) be as in Definition 5. Define the reduced output state on the effective light cone by \[\begin{align} \hat{\sigma}_{\psi}^{(i)} := \operatorname{Tr}_{\overline{\mathrm{LC}(i)}} \left[ \hat{\rho}_{\psi}(\boldsymbol{\theta}) \right]. \label{eq:def95sigma95i} \end{align}\qquad{(4)}\] Then, \[\begin{align} g_i(\psi) = \operatorname{Tr}\left[ \hat{B}_i\,\hat{\sigma}_{\psi}^{(i)} \right]. \label{eq:gradient95local95linear} \end{align}\qquad{(5)}\]

Moreover, if we define the local two-copy moment \[\begin{align} \hat{M}_{i,\mathrm{LC}}^{(2)} := \mathbb{E}_{\psi\sim\mu_{\mathrm{prod}}} \left[ \left( \hat{\sigma}_{\psi}^{(i)} \right)^{\otimes 2} \right], \label{eq:def95local95moment2} \end{align}\qquad{(6)}\] then \[\begin{align} \mathbb{E}_{\psi}\left[g_i(\psi)^2\right] = \operatorname{Tr}\left[ (\hat{B}_i\otimes \hat{B}_i)\, \hat{M}_{i,\mathrm{LC}}^{(2)} \right]. \label{eq:gradient95second95moment95local} \end{align}\qquad{(7)}\] In fact, \(\mathbb{E}_{\psi}[g_i(\psi)]=0\), and therefore \[\begin{align} \operatorname{Var}_{\psi}[g_i(\psi)] = \operatorname{Tr}\left[ (\hat{B}_i\otimes \hat{B}_i)\, \hat{M}_{i,\mathrm{LC}}^{(2)} \right]. \label{eq:gradient95variance95local} \end{align}\qquad{(8)}\]

Proof. —Eq. (?? ) follows directly from Eq. (50 ) and the fact that \(\hat{B}_i\) is supported on \(\mathrm{LC}(i)\). Indeed, \[\begin{align} g_i(\psi) = \operatorname{Tr}\left[ \hat{B}_i\,\hat{\rho}_{\psi}(\boldsymbol{\theta}) \right] = \operatorname{Tr}\left[ \hat{B}_i\, \operatorname{Tr}_{\overline{\mathrm{LC}(i)}} \left( \hat{\rho}_{\psi}(\boldsymbol{\theta}) \right) \right] = \operatorname{Tr}\left[ \hat{B}_i\,\hat{\sigma}_{\psi}^{(i)} \right]. \end{align}\]

To compute the second moment, square Eq. (?? ) and use the standard identity \[\begin{align} \operatorname{Tr}[\hat{X}\hat{\omega}]\operatorname{Tr}[\hat{Y}\hat{\omega}] = \operatorname{Tr}\left[ (\hat{X} \otimes \hat{Y})\, \hat{\omega}^{\otimes 2} \right] \label{eq:trace95square95identity} \end{align}\tag{56}\] for any operators \(\hat{X}, \hat{Y}, \hat{\omega}\) on the same Hilbert space. With \(\hat{X} = \hat{Y} = \hat{B}_i\) and \(\hat{\omega}=\hat{\sigma}_{\psi}^{(i)}\), this gives \[\begin{align} g_i(\psi)^2 = \operatorname{Tr}\left[ (\hat{B}_i\otimes \hat{B}_i)\, \left(\hat{\sigma}_{\psi}^{(i)}\right)^{\otimes 2} \right]. \end{align}\] Averaging over \(\psi\) and using linearity of trace and expectation, \[\begin{align} \mathbb{E}_{\psi}[g_i(\psi)^2] &=& \operatorname{Tr}\left[ (\hat{B}_i\otimes \hat{B}_i)\, \mathbb{E}_{\psi} \left[ \left(\hat{\sigma}_{\psi}^{(i)}\right)^{\otimes 2} \right] \right] \nonumber\\ &=& \operatorname{Tr}\left[ (\hat{B}_i\otimes \hat{B}_i)\, \hat{M}_{i,\mathrm{LC}}^{(2)} \right], \end{align}\] which proves Eq. (?? ).

It remains to show that the mean gradient vanishes under the Haar-product input ensemble. First note that \[\begin{align} \mathbb{E}_{\psi\sim\mu_{\mathrm{prod}}} \left[ \left|\psi\right\rangle\!\!\left\langle \psi\right| \right] = \left( \frac{\hat{\mathbb{1}}}{2} \right)^{\otimes n} = \frac{\hat{\mathbb{1}}}{2^n}. \label{eq:avg95input95maxmixed} \end{align}\tag{57}\] Therefore, \[\begin{align} \mathbb{E}_{\psi} \left[ \hat{\rho}_{\psi}(\boldsymbol{\theta}) \right] = \hat{U}(\boldsymbol{\theta}) \frac{\hat{\mathbb{1}}}{2^n} \hat{U}(\boldsymbol{\theta})^{\dagger} = \frac{\hat{\mathbb{1}}}{2^n}. \label{eq:avg95output95maxmixed} \end{align}\tag{58}\] Taking the partial trace over the complement of \(\mathrm{LC}(i)\), \[\begin{align} \mathbb{E}_{\psi} \left[ \hat{\sigma}_{\psi}^{(i)} \right] = \frac{\hat{\mathbb{1}}_{\mathrm{LC}(i)}}{2^{r_i}} = \frac{\hat{\mathbb{1}}_{\mathrm{LC}(i)}}{d_i}. \label{eq:avg95sigma95maxmixed} \end{align}\tag{59}\] Hence, \[\begin{align} \mathbb{E}_{\psi}[g_i(\psi)] &=& \operatorname{Tr}\left[ \hat{B}_i\, \mathbb{E}_{\psi}[\hat{\sigma}_{\psi}^{(i)}] \right] \nonumber\\ &=& \frac{1}{d_i}\operatorname{Tr}(\hat{B}_i) = 0, \end{align}\] where Lemma 1 was used in the last step. Consequently, \[\begin{align} \operatorname{Var}_{\psi}[g_i(\psi)] = \mathbb{E}_{\psi}[g_i(\psi)^2], \end{align}\] and Eq. (?? ) follows from Eq. (?? ). ◻

Proposition 4 is the exact local statement underlying this paper’s trainability narrative. The relevant object is neither the global expressibility of the full ansatz nor the global EP of the full circuit. For each parameter, the gradient variance is determined by a specific local two-copy moment, namely \(\hat{M}_{i,\mathrm{LC}}^{(2)}\), contracted against the parameter-dependent observable \(\hat{B}_i\otimes \hat{B}_i\).

In this sense, the variance of the gradient is a local analogue of the global two-copy formula for EP. The crucial difference is that the object depends on the parameter index \(i\) and on the effective light cone induced by the cost operator. This is exactly why one should not expect a single global scalar descriptor to reconstruct the full trainability profile \(\{\operatorname{Var}(g_i)\}_i\).

3.3 Local Haar benchmark and barren-plateau scaling↩︎

We next compare the exact local second moment to a Haar benchmark on the \(r_i\)-qubit effective light-cone Hilbert space. As emphasized already in Sec. 1, this benchmark should not be interpreted as an operator-level convergence claim for the fixed-circuit/product-input ensemble. Rather, it is a scalar reference value for the contraction relevant to gradient variance.

Let \(\mu_{\mathrm{Haar}}^{(r_i)}\) denote the Haar measure on pure states of the \(r_i\)-qubit Hilbert space associated with \(\mathrm{LC}(i)\), and let \[\begin{align} \hat{M}_{\mathrm{Haar},r_i}^{(2)} := \mathbb{E}_{\phi\sim\mu_{\mathrm{Haar}}^{(r_i)}} \left[ \left|\phi\right\rangle\!\!\left\langle \phi\right|^{\otimes 2} \right] = \frac{\hat{\mathbb{1}}+\hat{S}_i}{d_i(d_i+1)}, \label{eq:local95Haar95moment} \end{align}\tag{60}\] where \(\hat{S}_i\) is the full swap operator on two copies of the \(r_i\)-qubit light-cone Hilbert space.

We first compute the Haar benchmark exactly for any traceless local observable.

Lemma 2 (Haar benchmark for linear observables). Let \(\hat{B}=\hat{B}^{\dagger}\) be an operator on a \(d\)-dimensional Hilbert space with \(\operatorname{Tr}(\hat{B})=0\). For a Haar-random pure state \(\left|\phi\right\rangle\), define \[\begin{align} h_{\hat{B}}(\phi) := \operatorname{Tr}\left[ \hat{B}\,\left|\phi\right\rangle\!\!\left\langle \phi\right| \right]. \label{eq:def95hB} \end{align}\tag{61}\] Then, we have \[\begin{align} \mathbb{E}_{\phi}[h_{\hat{B}}(\phi)] = 0, \qquad \operatorname{Var}_{\phi}[h_{\hat{B}}(\phi)] = \frac{\operatorname{Tr}(\hat{B}^2)}{d(d+1)}. \label{eq:hB95var95exact} \end{align}\tag{62}\] In particular, \[\begin{align} \operatorname{Var}_{\phi}[h_{\hat{B}}(\phi)] \le \frac{\left\|\hat{B}\right\|_{\infty}^{2}}{d+1}. \label{eq:hB95var95bound} \end{align}\tag{63}\]

Proof. —The mean is immediate from the Haar first moment: \[\begin{align} \mathbb{E}_{\phi}\left[\left|\phi\right\rangle\!\!\left\langle \phi\right|\right] = \frac{\hat{\mathbb{1}}}{d}, \end{align}\] hence \[\begin{align} \mathbb{E}_{\phi}[h_{\hat{B}}(\phi)] = \frac{\operatorname{Tr}(\hat{B})}{d} = 0. \end{align}\]

For the second moment, use Eq. (56 ) with \(\hat{\omega}=\left|\phi\right\rangle\!\!\left\langle \phi\right|\) and \(\hat{X}=\hat{Y}=\hat{B}\): \[\begin{align} h_{\hat{B}}(\phi)^2 = \operatorname{Tr}\left[ (\hat{B}\otimes \hat{B}) \left|\phi\right\rangle\!\!\left\langle \phi\right|^{\otimes 2} \right]. \end{align}\] Averaging over \(\phi\) gives \[\begin{align} \mathbb{E}_{\phi}[h_{\hat{B}}(\phi)^2] &=& \operatorname{Tr}\left[ (\hat{B}\otimes \hat{B})\, \hat{M}_{\mathrm{Haar}}^{(2)} \right] = \frac{1}{d(d+1)} \operatorname{Tr}\left[ (\hat{B}\otimes \hat{B})(\hat{\mathbb{1}}+\hat{S}) \right] \nonumber\\ &=& \frac{1}{d(d+1)} \left( \operatorname{Tr}(\hat{B})^2+\operatorname{Tr}(\hat{B}^2) \right) = \frac{\operatorname{Tr}(\hat{B}^2)}{d(d+1)}. \end{align}\] Since the mean is zero, this is the variance, proving Eq. (62 ). Finally, \[\begin{align} \operatorname{Tr}(\hat{B}^2)\le d\left\|\hat{B}\right\|_{\infty}^{2}, \end{align}\] which immediately yields Eq. (63 ). ◻

Applying the lemma to the operator \(\hat{B}_i\) from Proposition 4 yields the local Haar benchmark for gradients.

Corollary 1 (Local 2-copy criterion for barren-plateau scaling). Let \(g_i(\psi)\) be the gradient associated with the parameter \(\theta_i\), and let \(\hat{B}_i\), \(r_i\), \(d_i\), and \(\hat{M}_{i,\mathrm{LC}}^{(2)}\) be as in Proposition 4. Define the Haar benchmark variance on the \(r_i\)-qubit light-cone Hilbert space by \[\begin{align} \operatorname{Var}_{\mathrm{Haar},r_i}[g_i] := \operatorname{Var}_{\phi\sim\mu_{\mathrm{Haar}}^{(r_i)}} \left[ \operatorname{Tr}\left(\hat{B}_i\left|\phi\right\rangle\!\!\left\langle \phi\right|\right) \right] = \operatorname{Tr}\left[(\hat{B}_i\otimes\hat{B}_i)\,\hat{M}_{\mathrm{Haar},r_i}^{(2)}\right]. \label{eq:def95local95Haar95var} \end{align}\tag{64}\] Then, \[\begin{align} \operatorname{Var}_{\psi}[g_i(\psi)] = \operatorname{Var}_{\mathrm{Haar},r_i}[g_i] + \operatorname{Tr}\left[ (\hat{B}_i\otimes \hat{B}_i) \left( \hat{M}_{i,\mathrm{LC}}^{(2)}-\hat{M}_{\mathrm{Haar},r_i}^{(2)} \right) \right]. \label{eq:var95difference95exact} \end{align}\tag{65}\] Consequently, \[\begin{align} \left| \operatorname{Var}_{\psi}[g_i(\psi)] - \operatorname{Var}_{\mathrm{Haar},r_i}[g_i] \right| \le \left\|\hat{B}_i\right\|_{2}^{2}\, \left\| \hat{M}_{i,\mathrm{LC}}^{(2)}-\hat{M}_{\mathrm{Haar},r_i}^{(2)} \right\|_{2}. \label{eq:var95difference95bound} \end{align}\tag{66}\] Moreover, \[\begin{align} \operatorname{Var}_{\mathrm{Haar},r_i}[g_i] = \frac{\operatorname{Tr}(\hat{B}_i^2)}{d_i(d_i+1)} \le \frac{\left\|\hat{B}_i\right\|_{\infty}^{2}}{d_i+1} \le \frac{4\left\|\hat{G}_i\right\|_{\infty}^{2}\left\|\hat{O}\right\|_{\infty}^{2}}{2^{r_i}+1}. \label{eq:local95Haar95BP95scale} \end{align}\tag{67}\] Therefore, whenever the correction term in Eq. (65 ) is at most of order \(2^{-r_i}\), one has \[\begin{align} \operatorname{Var}_{\psi}[g_i(\psi)] = O(2^{-r_i}). \label{eq:BP95scale95result} \end{align}\tag{68}\] In particular, if \(r_i=\Theta(n)\), this is exponentially small in the total system size and thus exhibits barren-plateau scaling.

Proof. —The contraction form in Eq. (64 ) is Lemma 2 with \(\hat{B}=\hat{B}_i\) and \(d=d_i\). To make the decomposition in Eq. (65 ) explicit, write \[\begin{align} \hat{M}_{i,\mathrm{LC}}^{(2)} = \hat{M}_{\mathrm{Haar},r_i}^{(2)} + \left( \hat{M}_{i,\mathrm{LC}}^{(2)}- \hat{M}_{\mathrm{Haar},r_i}^{(2)} \right). \label{eq:local95moment95Haar95decomposition} \end{align}\tag{69}\] Substituting this identity into Eq. (?? ) and using Eq. (64 ) yields Eq. (65 ).

For Eq. (66 ), apply the Hilbert–Schmidt Cauchy–Schwarz inequality: \[\begin{align} \left|\operatorname{Tr}\left[ (\hat{B}_i\otimes \hat{B}_i) \left( \hat{M}_{i,\mathrm{LC}}^{(2)}-\hat{M}_{\mathrm{Haar},r_i}^{(2)} \right) \right]\right| &\le& \left\|\hat{B}_i\otimes \hat{B}_i\right\|_{2}\, \left\| \hat{M}_{i,\mathrm{LC}}^{(2)} \hat{M}_{\mathrm{Haar},r_i}^{(2)} \right\|_{2} \nonumber\\ &=& \left\|\hat{B}_i\right\|_{2}^{2}\, \left\| \hat{M}_{i,\mathrm{LC}}^{(2)}-\hat{M}_{\mathrm{Haar},r_i}^{(2)} \right\|_{2}. \end{align}\]

Finally, Eq. (67 ) is Lemma 2 together with Lemma 1: \[\begin{align} \operatorname{Var}_{\mathrm{Haar},r_i}[g_i] \le \frac{\left\|\hat{B}_i\right\|_{\infty}^{2}}{d_i+1} \le \frac{4\left\|\hat{G}_i\right\|_{\infty}^{2}\left\|\hat{O}\right\|_{\infty}^{2}}{2^{r_i}+1}. \end{align}\] The final statement then follows directly from Eq. (65 ). ◻

Corollary 1 is the precise local statement behind the informal slogan that “2-design behavior causes the barren plateaus.” What actually enters the variance of the gradient is the contraction of the local two-copy moment against \(\hat{B}_i\otimes \hat{B}_i\). If that contraction is already Haar-like on an effective subsystem of size \(r_i\), then the gradient variance is forced down to the Haar scale \(O(2^{-r_i})\).

A comment is needed for consistency with Sec. 1. In the strict fixed-circuit/product-input model, Proposition 3 prevents literal operator convergence of the full moment orbit to a pure-state Haar moment operator. The same caveat applies here if one interprets \(\hat{M}_{i,\mathrm{LC}}^{(2)}\) as arising from a fixed-circuit orbit of local product inputs. Accordingly, the rigorous content of Corollary 1 should be read at the level of the relevant scalar contraction in Eq. (65 ), not as an unrestricted operator identity claim. In the usual random-parameter or random-circuit setting of McClean-type theorems, local approximate 2-designity is precisely a sufficient mechanism that makes this contraction Haar-like. In the present fixed-circuit formulation, the same mechanism is isolated as an exact local second-moment criterion.

It is also worth stressing how the light-cone size \(r_i\) depends on the architecture and on the choice of cost function. If \(\hat{O}\) is a sum of local observables, \[\begin{align} \hat{O}=\sum_{\alpha} c_{\alpha} \hat{O}_{\alpha}, \label{eq:O95local95sum} \end{align}\tag{70}\] then \[\begin{align} \hat{B}_i = \sum_{\alpha} c_{\alpha}\, i[\hat{H}_i^{\mathrm{out}},\hat{O}_{\alpha}], \label{eq:Bi95local95sum} \end{align}\tag{71}\] and only those terms whose supports overlap the forward cone \(\mathcal{F}(i)\) contribute. For bounded-depth circuits on bounded-degree lattices and local costs, the effective light cone often remains small, so the Haar benchmark in Eq. (67 ) is not exponentially small in \(n\). By contrast, for global costs or sufficiently deep circuits, the effective light cone can grow to \(r_i=\Theta(n)\), in which case the same formula yields the familiar exponentially vanishing gradient scale. This recovers, in a single unified expression, the well-known distinction between local-cost and global-cost barren plateaus.

3.4 Interpretation of global descriptors and local trainability↩︎

We can now articulate the structural relation among the quantities studied so far.

Remark 1 (Moment hierarchy of coverage, fluctuation, and trainability). The results of Secs. 2 and 3 can be condensed as \[\begin{align} \text{BP is local-2-copy},\qquad \text{EP is global-2-copy-scalar},\qquad \text{EPD is global-4-copy-scalar}. \label{eq:BP95EP95EPD95slogan} \end{align}\tag{72}\] Accordingly, \(\mathrm{EP}(\hat{U})\) alone cannot determine barren-plateau behavior. It is a global mean-entangling descriptor of one fixed circuit, whereas \(\operatorname{Var}(g_i)\) is controlled parameter-by-parameter by a local second moment on the effective light cone. The quantity \(\mathrm{EPD}(\hat{U})\) does not determine \(\operatorname{Var}(g_i)\) either, but because it probes higher-order fluctuation structure invisible to EP, it can serve as a more informative diagnostic of the randomization process associated with plateau formation.

Remark 1 should be read as the conceptual punchline of the present section. The gradient variance is not governed by a one-dimensional trade-off between “more entanglement” and “less trainability.” Rather, trainability is a local causal property. The role of global descriptors such as expressibility, EP, and EPD is indirect: they provide increasingly refined information about how broadly and how unevenly the circuit explores the state space, but the actual gradient suppression occurs when the local two-copy contraction seen by each parameter has already reached its Haar-like scale.

This is precisely the reason why one can, in principle, have an intermediate regime in which a circuit already has large EP or near-saturated coverage, while the local two-copy moments governing the gradients have not yet fully collapsed. It is also the reason why EPD, being a fluctuation-sensitive descriptor rather than a pure mean descriptor, can empirically align more closely with plateau onset than EP alone. The next sections exploit these observations to develop the structural EP–EPD picture and the associated numerical diagnostics of trainability.

4 Why Mean Descriptors Are Not Enough↩︎

The preceding sections already suggest a strong obstruction to any one-dimensional account of trainability. On the one hand, the fixed-circuit entangling power \(\mathrm{EP}(\hat{U})\) is a single scalar functional of the global two-copy output moment: \[\begin{align} \mathrm{EP}(\hat{U})=\operatorname{Tr}\left[\hat{K}_Q \hat{M}_{\hat{U}}^{(2)}\right]. \label{eq:mean95not95enough95EP95recall} \end{align}\tag{73}\] On the other hand, the gradient-variance profile of a parametrized circuit, \[\begin{align} \Gamma(\hat{U}) := \left( \operatorname{Var}_{\psi}[g_1(\psi)], \dots, \operatorname{Var}_{\psi}[g_p(\psi)] \right), \label{eq:def95trainability95profile} \end{align}\tag{74}\] is determined parameter-by-parameter by a family of local two-copy contractions on effective light cones: \[\begin{align} \operatorname{Var}_{\psi}[g_i(\psi)] = \operatorname{Tr}\left[ (\hat{B}_i\otimes \hat{B}_i)\, \hat{M}_{i,\mathrm{LC}}^{(2)} \right]. \label{eq:mean95not95enough95var95recall} \end{align}\tag{75}\] The structural mismatch is immediate: a single global scalar cannot, in general, encode an entire parameter-indexed collection of local second moments.

The aim of the present section is to make this obstruction explicit. We do so in the strongest possible form, namely by constructing two entangling blocks that have the same mean entangling power but induce different gradient variances when inserted into the same trainable architecture. In this sense, the failure of mean descriptors is not merely a matter of asymptotic scaling or numerical correlation. It is already present at the level of an explicit, finite-dimensional, analytically tractable counterexample.

At the same time, the example will show why EPD is a useful refinement. Although EPD is still not the trainability object itself, it distinguishes the two blocks that EP cannot distinguish. This is exactly the kind of separation one expects from the hierarchy established earlier: \[\begin{align} \text{EP: global mean,} \qquad \text{EPD: global fluctuation,} \qquad \text{BP: local second-moment trainability.} \label{eq:hierarchy95recall95mean95not95enough} \end{align}\tag{76}\]

4.1 A one-dimensional descriptor cannot parametrize a trainability profile↩︎

Before giving the explicit construction, it is useful to state precisely what is meant by “mean descriptors are not enough.” For a fixed architecture and fixed cost class, the trainability data of interest is not a single number but the vector \(\Gamma(\hat{U})\) in Eq. (74 ). Even if one focuses on a single designated parameter, the associated gradient variance remains a local object that depends on the interaction between the parameter generator, the cost observable, and the effective light cone. This already suggests non-identifiability from any one-number global mean descriptor.

To prepare the constructive theorem below, we record two elementary lemmas.

Lemma 3 (Conjugation of anticommuting Pauli strings). Let \(\hat{P}\) and \(\hat{Q}\) be Hermitian Pauli strings satisfying \[\begin{align} \hat{P}^2=\hat{Q}^2=\hat{\mathbb{1}}, \qquad \{\hat{P},\hat{Q}\}=0. \end{align}\] Then, for all real \(\theta\), \[\begin{align} e^{i\theta \hat{P}} \hat{Q} e^{-i\theta \hat{P}} = \hat{Q}\cos(2\theta) + i\hat{P}\hat{Q} \sin(2\theta). \label{eq:anticommuting95pauli95conjugation} \end{align}\tag{77}\]

Proof. —Define \[\begin{align} F(\theta):=e^{i\theta \hat{P}}\hat{Q}e^{-i\theta \hat{P}}. \end{align}\] Then, \[\begin{align} F'(\theta) = ie^{i\theta \hat{P}}[\hat{P},\hat{Q}]e^{-i\theta \hat{P}} = 2i\,e^{i\theta \hat{P}}\hat{P}\hat{Q}\,e^{-i\theta \hat{P}}, \label{eq:Fprime95anticomm} \end{align}\tag{78}\] because \(\{\hat{P},\hat{Q}\}=0\) implies \([\hat{P},\hat{Q}]=2\hat{P}\hat{Q}\). Likewise, with \(R(\theta):=e^{i\theta \hat{P}}(i\hat{P}\hat{Q})e^{-i\theta \hat{P}}\), one finds \[\begin{align} R'(\theta) &=& ie^{i\theta \hat{P}}[\hat{P},i\hat{P}\hat{Q}]e^{-i\theta \hat{P}}. \end{align}\] Since \(\hat{P}(i\hat{P}\hat{Q})=i\hat{Q}\) and \((i\hat{P}\hat{Q})\hat{P}=-i\hat{Q}\), we have \[\begin{align} [\hat{P},i\hat{P}\hat{Q}]=2i\hat{Q}, \end{align}\] and hence, \[\begin{align} R'(\theta)=-2F(\theta). \label{eq:Rprime95anticomm} \end{align}\tag{79}\] The pair \((F,R)\) therefore satisfies the linear system \[\begin{align} F'(\theta)=2R(\theta), \qquad R'(\theta)=-2F(\theta), \end{align}\] with initial conditions \[\begin{align} F(0)=\hat{Q}, \qquad R(0)=i\hat{P}\hat{Q}. \end{align}\] The unique solution is exactly Eq. (77 ). ◻

Lemma 4 (Single-qubit Haar moments of Bloch coordinates). Let \(\left|\psi\right\rangle\sim \mu_{\mathrm{Haar}}^{(1)}\) be Haar-random on one qubit, and define \[\begin{align} x:=\left\langle \psi\right|\hat{X}\left|\psi\right\rangle, \qquad y:=\left\langle \psi\right|\hat{Y}\left|\psi\right\rangle, \qquad z:=\left\langle \psi\right|\hat{Z}\left|\psi\right\rangle. \end{align}\] Then, \[\begin{align} \mathbb{E}[x]=\mathbb{E}[y]=\mathbb{E}[z]=0, \quad\text{and}\quad \mathbb{E}[x^2]=\mathbb{E}[y^2]=\mathbb{E}[z^2]=\frac{1}{3}. \label{eq:bloch95second95moments} \end{align}\tag{80}\] More generally, \[\begin{align} \mathbb{E}[x_ax_b]=\frac{\delta_{ab}}{3}, \qquad a,b\in\{x,y,z\}. \label{eq:bloch95covariance} \end{align}\tag{81}\]

Proof. —The Haar measure on one qubit induces the uniform distribution on the Bloch sphere. Hence the random vector \[\begin{align} \vec{r} := (x,y,z) \end{align}\] is rotationally invariant and satisfies \[\begin{align} x^2+y^2+z^2=1 \end{align}\] almost surely. Rotational invariance implies \[\begin{align} \mathbb{E}[x]=\mathbb{E}[y]=\mathbb{E}[z]=0 \end{align}\] and \[\begin{align} \mathbb{E}[x^2]=\mathbb{E}[y^2]=\mathbb{E}[z^2]. \end{align}\] Summing these equal quantities and using \(\mathbb{E}[x^2+y^2+z^2]=1\) gives Eq. (80 ). The mixed moments in Eq. (81 ) vanish by rotational symmetry. ◻

4.2 Constructive equal-EP, unequal-trainability theorem↩︎

We now present the promised explicit separation. The two entanglers are chosen from the canonical two-qubit KAK family analyzed in the preceding EPD work SM-ChoBang2026EPD?. Concretely, define \[\begin{align} \hat{E}_{\mathrm{C}} := \exp\left( -i\frac{\pi}{4}\hat{X}_1\hat{X}_2 \right), \label{eq:def95EC} \end{align}\tag{82}\] and \[\begin{align} \hat{E}_{\mathrm{B}} := \exp\left( -i\frac{\pi}{4}\hat{X}_1\hat{X}_2 -i\frac{\pi}{8}\hat{Y}_1\hat{Y}_2 \right). \label{eq:def95EB} \end{align}\tag{83}\] Because \(\hat{X}_1\hat{X}_2\) and \(\hat{Y}_1\hat{Y}_2\) commute, one may also write \[\begin{align} \hat{E}_{\mathrm{B}} = \exp\left( -i\frac{\pi}{4}\hat{X}_1\hat{X}_2 \right) \exp\left( -i\frac{\pi}{8}\hat{Y}_1\hat{Y}_2 \right). \label{eq:def95EB95factorized} \end{align}\tag{84}\] These are the canonical Cartan representatives with parameters \[\begin{align} \chi_{\mathrm{C}}=\left(\frac{\pi}{4},0,0\right), \qquad \chi_{\mathrm{B}}=\left(\frac{\pi}{4},\frac{\pi}{8},0\right), \end{align}\] which are locally equivalent to the CNOT and \(B\)-gate classes, respectively SM-ChoBang2026EPD?.

We embed each entangler into the same one-parameter hardware-efficient block, \[\begin{align} \hat{U}_{\alpha}(\theta) := \hat{E}_{\alpha} \left( e^{-i\theta \hat{Z}_1/2}\otimes \hat{\mathbb{1}}_2 \right), \qquad \alpha\in\{\mathrm{C},\mathrm{B}\}, \label{eq:def95UEalpha} \end{align}\tag{85}\] and consider the same cost observable for both: \[\begin{align} \hat{O}:= \hat{\mathbb{1}}_1\otimes \hat{Z}_2. \label{eq:def95O95separation} \end{align}\tag{86}\] For an input product state \(\left|\psi\right\rangle=\left|\psi_1\right\rangle\otimes \left|\psi_2\right\rangle\) sampled from \(\mu_{\mathrm{prod}}\), define \[\begin{align} C_{\alpha}(\theta,\psi) := \left\langle \psi\right| \hat{U}_{\alpha}(\theta)^{\dagger} \hat{O} \hat{U}_{\alpha}(\theta) \left|\psi\right\rangle, \label{eq:def95Calpha} \end{align}\tag{87}\] and the corresponding gradient \[\begin{align} g_{\alpha}(\psi,\theta) := \partial_{\theta} C_{\alpha}(\theta,\psi). \label{eq:def95galpha} \end{align}\tag{88}\]

We can now state the central result of this section.

Theorem 4 (Constructive equal-EP, unequal-trainability theorem). The two one-parameter blocks \(\hat{U}_{\mathrm{C}}(\theta)\) and \(\hat{U}_{\mathrm{B}}(\theta)\) defined in Eq. (85 ) satisfy the following properties.

(i) Equal mean entangling power but unequal entangling-power deviation SM-ChoBang2026EPD?. For every \(\theta\), \[\begin{align} \mathrm{EP}(\hat{U}_{\mathrm{C}}(\theta)) = \mathrm{EP}(\hat{U}_{\mathrm{B}}(\theta)) = \frac{4}{9}, \label{eq:equal95EP95theta95all} \end{align}\tag{89}\] whereas \[\begin{align} \mathrm{EPD}(\hat{U}_{\mathrm{C}}(\theta)) = \frac{4\sqrt{11}}{45}, \qquad \mathrm{EPD}(\hat{U}_{\mathrm{B}}(\theta)) = \frac{2\sqrt{35}}{45}. \label{eq:unequal95EPD95theta95all} \end{align}\tag{90}\] In particular, \[\begin{align} \mathrm{EP}(\hat{U}_{\mathrm{C}}(\theta)) = \mathrm{EP}(\hat{U}_{\mathrm{B}}(\theta)), \qquad \mathrm{EPD}(\hat{U}_{\mathrm{C}}(\theta)) \neq \mathrm{EPD}(\hat{U}_{\mathrm{B}}(\theta)). \label{eq:sameEP95diffEPD} \end{align}\tag{91}\]

(ii) Unequal gradient variance in the same trainable architecture. For every \(\theta\), \[\begin{align} \operatorname{Var}_{\psi}\left[g_{\mathrm{C}}(\psi,\theta)\right] = \frac{1}{9}, \qquad \operatorname{Var}_{\psi}\left[g_{\mathrm{B}}(\psi,\theta)\right] = \frac{1}{18}. \label{eq:grad95var95explicit95difference} \end{align}\tag{92}\] Hence, the two blocks have identical mean entangling power but different trainability, even though the trainable layer, parameter placement, and cost observable are otherwise the same.

Proof. —We divide the proof into two parts.

Part I: same EP, different EPD.

By Remark [rem:LU95invariance], EP and EPD are invariant under local pre- and post-unitaries. Therefore, the fixed-circuit descriptors of \(\hat{E}_{\mathrm{C}}\) and \(\hat{E}_{\mathrm{B}}\) agree with those of the CNOT and \(B\) local-equivalence classes analyzed in Ref. SM-ChoBang2026EPD?. That work uses the two-qubit linear entropy \[\begin{align} E(\left|\phi\right\rangle) = 1-\operatorname{Tr}\hat{\rho}_1^2, \end{align}\] whereas our present two-qubit descriptor is the Meyer–Wallach quantity \[\begin{align} Q(\left|\phi\right\rangle) = \frac{2}{2}\sum_{j=1}^{2}(1-\operatorname{Tr}\hat{\rho}_j^2) = 2E(\left|\phi\right\rangle). \label{eq:Q95equals952E95twobits} \end{align}\tag{93}\] Hence, the EP and EPD values from Ref. SM-ChoBang2026EPD? are multiplied by a factor of \(2\) in our normalization. Since the CNOT and \(B\) classes have \[\begin{align} \mathrm{ep}_{E}(\mathrm{CNOT}) = \mathrm{ep}_{E}(B) = \frac{2}{9}, \qquad \Delta_{E}(\mathrm{CNOT}) = \frac{2\sqrt{11}}{45}, \qquad \Delta_{E}(B) = \frac{\sqrt{35}}{45}, \end{align}\] we obtain \[\begin{align} \mathrm{EP}(\hat{E}_{\mathrm{C}}) = \mathrm{EP}(\hat{E}_{\mathrm{B}}) = \frac{4}{9}, \label{eq:EP95EC95EB} \end{align}\tag{94}\] and \[\begin{align} \mathrm{EPD}(\hat{E}_{\mathrm{C}}) = \frac{4\sqrt{11}}{45}, \qquad \mathrm{EPD}(\hat{E}_{\mathrm{B}}) = \frac{2\sqrt{35}}{45}. \label{eq:EPD95EC95EB} \end{align}\tag{95}\] Now \(\hat{U}_{\alpha}(\theta)\) differs from \(\hat{E}_{\alpha}\) only by right multiplication with the local unitary \(e^{-i\theta \hat{Z}_1/2}\otimes \hat{\mathbb{1}}_2\), so Remark [rem:LU95invariance] gives \[\begin{align} \mathrm{EP}(\hat{U}_{\alpha}(\theta))=\mathrm{EP}(\hat{E}_{\alpha}), \qquad \mathrm{EPD}(\hat{U}_{\alpha}(\theta))=\mathrm{EPD}(\hat{E}_{\alpha}) \end{align}\] for all \(\theta\) and both \(\alpha\in\{\mathrm{C},\mathrm{B}\}\). This proves Eqs. (89 )–(91 ).

Part II: explicit gradient variance calculation.

We now compute the gradient with respect to the trainable \(\hat{Z}_1\)-rotation. Write \[\begin{align} \hat{R}_Z(\theta):=e^{-i\theta \hat{Z}_1/2}\otimes \hat{\mathbb{1}}_2. \end{align}\] Then, \[\begin{align} \hat{U}_{\alpha}(\theta)=\hat{E}_{\alpha}\hat{R}_Z(\theta), \end{align}\] so Eq. (87 ) becomes \[\begin{align} C_{\alpha}(\theta,\psi) = \left\langle \psi\right| \hat{R}_Z(\theta)^{\dagger} \hat{E}_{\alpha}^{\dagger}\hat{O}\hat{E}_{\alpha} \hat{R}_Z(\theta) \left|\psi\right\rangle. \label{eq:Calpha95input95picture} \end{align}\tag{96}\] Differentiating, \[\begin{align} g_{\alpha}(\psi,\theta) &=& \left\langle \psi\right| \hat{R}_Z(\theta)^{\dagger} \hat{A}_{\alpha} \hat{R}_Z(\theta) \left|\psi\right\rangle, \label{eq:galpha95input95picture} \end{align}\tag{97}\] where \[\begin{align} \hat{A}_{\alpha} := \frac{i}{2} \left[ \hat{Z}_1\otimes \hat{\mathbb{1}}_2, \hat{E}_{\alpha}^{\dagger}\hat{O}\hat{E}_{\alpha} \right]. \label{eq:def95Aalpha} \end{align}\tag{98}\] Thus, the problem reduces to computing the dressed observable \(\hat{E}_{\alpha}^{\dagger}\hat{O}\hat{E}_{\alpha}\).

We first treat the \(\mathrm{C}\) block. Since \(\hat{X}_1\hat{X}_2\) anticommutes with \(\hat{O}=\hat{\mathbb{1}}_1\otimes \hat{Z}_2\), Lemma 3 with \[\begin{align} \hat{P}=\hat{X}_1\hat{X}_2, \qquad \hat{Q}=\hat{\mathbb{1}}_1\otimes \hat{Z}_2, \qquad \theta=\frac{\pi}{4} \end{align}\] yields \[\begin{align} \hat{E}_{\mathrm{C}}^{\dagger}\hat{O}\hat{E}_{\mathrm{C}} = e^{i\frac{\pi}{4}\hat{X}_1\hat{X}_2} (\hat{\mathbb{1}}_1\otimes \hat{Z}_2) e^{-i\frac{\pi}{4}\hat{X}_1\hat{X}_2} = i(\hat{X}_1\hat{X}_2)(\hat{\mathbb{1}}_1\otimes \hat{Z}_2). \label{eq:EC95conjugates95O} \end{align}\tag{99}\] Using \(\hat{X}\hat{Z}=-i\hat{Y}\), we obtain \[\begin{align} i(\hat{X}_1\hat{X}_2)(\hat{\mathbb{1}}_1\otimes \hat{Z}_2) = i\left(\hat{X}_1\otimes \hat{X}_2\hat{Z}_2\right) = i\left(\hat{X}_1\otimes (-i\hat{Y}_2)\right) = \hat{X}_1\hat{Y}_2. \label{eq:EC95conjugates95O95XY} \end{align}\tag{100}\] Therefore, \[\begin{align} \hat{A}_{\mathrm{C}} &=& \frac{i}{2} \left[ \hat{Z}_1\otimes \hat{\mathbb{1}}_2, \hat{X}_1\hat{Y}_2 \right]. \end{align}\] Now \[\begin{align} (\hat{Z}_1\otimes \hat{\mathbb{1}}_2)(\hat{X}_1\hat{Y}_2)=i\hat{Y}_1\hat{Y}_2, \qquad (\hat{X}_1\hat{Y}_2)(\hat{Z}_1\otimes \hat{\mathbb{1}}_2)=-i\hat{Y}_1\hat{Y}_2, \end{align}\] so \[\begin{align} \hat{A}_{\mathrm{C}} = \frac{i}{2}(2i\hat{Y}_1\hat{Y}_2) = -\,\hat{Y}_1\hat{Y}_2. \label{eq:AC95equals95YY} \end{align}\tag{101}\]

Next consider the \(\mathrm{B}\) block. Since \(\hat{X}_1\hat{X}_2\) and \(\hat{Y}_1\hat{Y}_2\) commute, Eq. (84 ) gives \[\begin{align} \hat{E}_{\mathrm{B}}^{\dagger}\hat{O}\hat{E}_{\mathrm{B}} &=& e^{i\frac{\pi}{8}\hat{Y}_1\hat{Y}_2} \left( e^{i\frac{\pi}{4}\hat{X}_1\hat{X}_2} (\hat{\mathbb{1}}_1\otimes \hat{Z}_2) e^{-i\frac{\pi}{4}\hat{X}_1\hat{X}_2} \right) e^{-i\frac{\pi}{8}\hat{Y}_1\hat{Y}_2} \nonumber\\ &=& e^{i\frac{\pi}{8}\hat{Y}_1\hat{Y}_2} (\hat{X}_1\hat{Y}_2) e^{-i\frac{\pi}{8}\hat{Y}_1\hat{Y}_2}, \label{eq:EB95conjugation95step1} \end{align}\tag{102}\] where Eq. (100 ) was used. Since \(\hat{Y}_1\hat{Y}_2\) anticommutes with \(\hat{X}_1\hat{Y}_2\), Lemma 3 now gives \[\begin{align} \hat{E}_{\mathrm{B}}^{\dagger}\hat{O}\hat{E}_{\mathrm{B}} &=& (\hat{X}_1\hat{Y}_2)\cos\frac{\pi}{4} + i(\hat{Y}_1\hat{Y}_2)(\hat{X}_1\hat{Y}_2)\sin\frac{\pi}{4}. \end{align}\] Because \(\hat{Y}_1\hat{X}_1=-i\hat{Z}_1\) and \(\hat{Y}_2\hat{Y}_2=\hat{\mathbb{1}}_2\), we have \[\begin{align} i(\hat{Y}_1\hat{Y}_2)(\hat{X}_1\hat{Y}_2) = i(\hat{Y}_1\hat{X}_1\otimes \hat{\mathbb{1}}_2) = i(-i\hat{Z}_1\otimes \hat{\mathbb{1}}_2) = \hat{Z}_1\otimes \hat{\mathbb{1}}_2. \end{align}\] Hence \[\begin{align} \hat{E}_{\mathrm{B}}^{\dagger}\hat{O}\hat{E}_{\mathrm{B}} = \frac{1}{\sqrt{2}} \left( \hat{X}_1\hat{Y}_2 + \hat{Z}_1\otimes \hat{\mathbb{1}}_2 \right). \label{eq:EB95conjugates95O} \end{align}\tag{103}\] Substituting into Eq. (98 ), \[\begin{align} \hat{A}_{\mathrm{B}} = \frac{i}{2} \left[ \hat{Z}_1\otimes \hat{\mathbb{1}}_2, \frac{1}{\sqrt{2}} \left( \hat{X}_1\hat{Y}_2+\hat{Z}_1\otimes \hat{\mathbb{1}}_2 \right) \right] = \frac{i}{2\sqrt{2}} \left[ \hat{Z}_1\otimes \hat{\mathbb{1}}_2, \hat{X}_1\hat{Y}_2 \right] = -\frac{1}{\sqrt{2}}\,\hat{Y}_1\hat{Y}_2, \label{eq:AB95equals95YY} \end{align}\tag{104}\] where Eq. (101 ) was used in the last step.

Eqs. (101 ) and (104 ) show that the two architectures reduce to the same input observable \(\hat{Y}_1\hat{Y}_2\), but with different amplitudes. For an input product state \(\left|\psi_1\right\rangle\otimes \left|\psi_2\right\rangle\), define \[\begin{align} y_j := \left\langle \psi_j\right|\hat{Y}_j\left|\psi_j\right\rangle, \qquad j=1,2. \end{align}\] Then Eq. (97 ) gives \[\begin{align} g_{\mathrm{C}}(\psi,\theta) = -\, \left\langle \psi\right| \hat{R}_Z(\theta)^{\dagger}(\hat{Y}_1\hat{Y}_2)\hat{R}_Z(\theta) \left|\psi\right\rangle, \label{eq:gC95theta} \end{align}\tag{105}\] and \[\begin{align} g_{\mathrm{B}}(\psi,\theta) = -\frac{1}{\sqrt{2}}\, \left\langle \psi\right| \hat{R}_Z(\theta)^{\dagger}(\hat{Y}_1\hat{Y}_2)\hat{R}_Z(\theta) \left|\psi\right\rangle. \label{eq:gB95theta} \end{align}\tag{106}\] Now \(\hat{R}_Z(\theta)\) is a local unitary and \(\mu_{\mathrm{prod}}\) is invariant under independent single-qubit Haar rotations. Therefore the distribution of \[\begin{align} \left\langle \psi\right| \hat{R}_Z(\theta)^{\dagger}(\hat{Y}_1\hat{Y}_2)\hat{R}_Z(\theta) \left|\psi\right\rangle \end{align}\] is the same as that of \[\begin{align} \left\langle \psi\right|\hat{Y}_1\hat{Y}_2\left|\psi\right\rangle = y_1y_2. \end{align}\] Hence the gradient distribution is independent of \(\theta\), and it suffices to compute at \(\theta=0\): \[\begin{align} g_{\mathrm{C}}(\psi,0)=-y_1y_2, \qquad g_{\mathrm{B}}(\psi,0)=-\frac{1}{\sqrt{2}}y_1y_2. \label{eq:gC95gB95y1y2} \end{align}\tag{107}\]

By Lemma 4, the independent random variables \(y_1\) and \(y_2\) satisfy \[\begin{align} \mathbb{E}[y_j]=0, \qquad \mathbb{E}[y_j^2]=\frac{1}{3}. \end{align}\] Therefore, \[\begin{align} \mathbb{E}[y_1y_2]=\mathbb{E}[y_1]\mathbb{E}[y_2]=0, \end{align}\] and \[\begin{align} \mathbb{E}[y_1^2y_2^2] = \mathbb{E}[y_1^2]\mathbb{E}[y_2^2] = \frac{1}{9}. \end{align}\] Thus \[\begin{align} \operatorname{Var}_{\psi}[g_{\mathrm{C}}(\psi,\theta)] &=& \operatorname{Var}_{\psi}[y_1y_2] = \frac{1}{9}, \\ \operatorname{Var}_{\psi}[g_{\mathrm{B}}(\psi,\theta)] &=& \frac{1}{2}\operatorname{Var}_{\psi}[y_1y_2] = \frac{1}{18}, \end{align}\] for all \(\theta\). This proves Eq. (92 ). ◻

Theorem 4 is the key structural separation result of the paper. It exhibits, in a fully analytic manner, two circuit blocks that are indistinguishable by the scalar mean descriptor EP but distinguishable both by EPD and by an actual trainability metric, namely the gradient variance of the same parameter in the same trainable architecture.

Several important consequences follow immediately.

4.3 Non-identifiability of barren plateaus from scalar mean descriptors↩︎

We can now state the general structural conclusion.

Proposition 5 (Non-identifiability of trainability from scalar mean descriptors). Fix an architecture class and a cost observable, and let \[\begin{align} \Gamma(\hat{U}) = \left( \operatorname{Var}_{\psi}[g_1(\psi)], \dots, \operatorname{Var}_{\psi}[g_p(\psi)] \right) \end{align}\] denote the gradient-variance profile with respect to the Haar-product input ensemble. Then the map \(\hat{U}\mapsto \Gamma(\hat{U})\) does not* factor through the scalar mean descriptor \(\mathrm{EP}(\hat{U})\). Equivalently, there does not exist a function \[\begin{align} F:\mathbb{R}\to \mathbb{R}^{p} \end{align}\] such that \[\begin{align} \Gamma(\hat{U})=F(\mathrm{EP}(\hat{U})) \qquad \text{for all }\hat{U} \end{align}\] in the architecture class.*

Proof. Theorem 4 provides two one-parameter blocks \(\hat{U}_{\mathrm{C}}(\theta)\) and \(\hat{U}_{\mathrm{B}}(\theta)\) such that \[\begin{align} \mathrm{EP}(\hat{U}_{\mathrm{C}}(\theta)) = \mathrm{EP}(\hat{U}_{\mathrm{B}}(\theta)) \end{align}\] for all \(\theta\), while \[\begin{align} \operatorname{Var}_{\psi}[g_{\mathrm{C}}(\psi,\theta)] \neq \operatorname{Var}_{\psi}[g_{\mathrm{B}}(\psi,\theta)]. \end{align}\] Hence even the variance of a single designated parameter cannot be recovered as a function of \(\mathrm{EP}(\hat{U})\) alone. A fortiori, the full profile \(\Gamma(\hat{U})\) cannot factor through \(\mathrm{EP}(\hat{U})\). ◻

Proposition 5 is stronger than a mere empirical warning. It says that the insufficiency of mean descriptors is a matter of principle, not of imperfect regression or finite-sample error. The map \[\begin{align} \hat{U}\longmapsto \mathrm{EP}(\hat{U}) \end{align}\] forgets information that is genuinely relevant to trainability.

This conclusion is entirely consistent with the moment-theoretic analysis of Secs. 2 and 3. The quantity \(\mathrm{EP}(\hat{U})\) is a single contraction of the global two-copy output moment against the observable \(\hat{K}_Q\). By contrast, each gradient variance is a different contraction of a parameter-dependent local two-copy moment against a parameter-dependent local observable \(\hat{B}_i\otimes \hat{B}_i\). There is no reason, in general, for all of that local information to be recoverable from one global scalar.

The explicit example above also clarifies the role of EPD. Since \[\begin{align} \mathrm{EP}(\hat{U}_{\mathrm{C}}(\theta)) = \mathrm{EP}(\hat{U}_{\mathrm{B}}(\theta)), \qquad \mathrm{EPD}(\hat{U}_{\mathrm{C}}(\theta)) \neq \mathrm{EPD}(\hat{U}_{\mathrm{B}}(\theta)), \end{align}\] the higher-order fluctuation descriptor already distinguishes the two blocks before any gradient is computed. In this sense, EPD is strictly more informative than EP for this pair. However, it is important not to overstate the conclusion.

Remark 2 (EPD is more informative than EP, but still not the trainability invariant). Theorem 4 does not imply that EPD alone determines barren-plateau behavior. EPD remains a global scalar four-copy quantity, whereas gradient variance is a local two-copy quantity. What the theorem does show is more precise: EPD can distinguish circuits that are invisible to EP and that, once embedded into the same trainable block, exhibit different gradient variances. Thus EPD is not the trainability object itself, but it is a strictly richer diagnostic than EP.

The non-identifiability statement can be rephrased in the language of information loss. A scalar mean descriptor compresses the circuit to one number, the trainability profile is, in general, a many-parameter object. Theorem 4 shows that this compression is already too coarse at the two-qubit block level. Consequently, when such blocks are used as building units of hardware-efficient ansätze, the inability of a mean descriptor to resolve local trainability is not an exceptional pathology but a generic structural limitation.

4.4 Interpretation for expressibility and barren-plateau phenomenology↩︎

Although the explicit construction above is formulated in terms of fixed-circuit EP and EPD, its interpretive consequence extends directly to the broader expressibility literature. Indeed, scalar expressibility indices are likewise one-number summaries of a circuit family. Such descriptors can correlate with broad architectural trends, and in many random-circuit settings they do. But the present analysis shows why no such scalar should be expected to provide a complete criterion for plateau onset.

The correct hierarchy is now clear:

  1. A global mean descriptor such as EP measures average entangling strength.

  2. A global fluctuation descriptor such as EPD measures the input dependence of that strength.

  3. The actual trainability object is the parameter-resolved local two-copy contraction on the effective light cone.

Accordingly, the common informal slogan \[\begin{align} \text{``more expressibility''} \Longrightarrow \text{``more barren plateau''} \end{align}\] must be interpreted with care. What can be true, under additional assumptions, is that stronger randomization tends to push several descriptors in the same direction. But the descriptors live at different levels of the moment hierarchy and encode different information. A mean descriptor can saturate before the local second moments relevant to trainability have fully collapsed.

This is exactly the regime that motivates the practical two-dial picture advocated in this paper. One dial measures coverage or average entangling strength, the other measures fluctuation or variability. Theorem 4 shows that fixing the first dial does not fix the second, nor does it fix trainability. Therefore any meaningful account of barren plateaus must go beyond a single mean descriptor.

We summarize the lesson of this section as follows:

Equal mean entangling power does not imply equal trainability.

The obstruction is not numerical but structural, and it arises already at the smallest nontrivial entangling scale. This is why the transition from mean descriptors to fluctuation-sensitive and local-moment-sensitive diagnostics is not merely a refinement of language, but a necessary step in understanding PQC trainability.

5 Numerical protocols and supplementary diagnostics↩︎

This section gives the numerical details supporting the depth-onset and route-scan claims in the main manuscript. The guiding question is the following: as the depth \(L\) of a PQC family is increased, can one locate a regime in which the circuit has already acquired substantial coverage while its product-input response has not yet homogenized?

All simulations use product-Haar input probes for fixed-circuit EP and EPD. Family-level quantities are obtained in fixed-circuit order: for each sampled parameter vector, the product-input distribution of \(Q(\hat{U}_L(\boldsymbol{\theta})\left|\psi\right\rangle)\) is first estimated; the corresponding fixed-circuit mean and standard deviation give EP and EPD; only after this step are the resulting descriptors averaged over parameter samples. This ordering is essential. It keeps the present EP/EPD descriptors distinct from Sim-style parameter-ensemble expressibility, in which one samples many circuit instances from a fixed reference input.

The numerical section should therefore be read with the same hierarchy as the analytical sections: \[\begin{align} \text{coverage dial} &:& \text{how far the family has moved toward high mean entanglement},\nonumber\\ \text{variability dial} &:& \text{how much product-input response width remains},\nonumber\\ \text{gradient probe} &:& \text{a cost-dependent consistency check of local trainability}. \label{eq:sm95numeric95hierarchy} \end{align}\tag{108}\] The first two entries define the EP/EPD scan; the third is not used to define the scan and is included only to compare the diagnostic with an explicit gradient scale.

5.1 Main example: Ring-CZ depth scan↩︎

The ring-CZ hardware-efficient ansatz used in the onset diagnostic is \[\begin{align} \hat{U}_L(\boldsymbol{\theta})=\prod_{\ell=1}^{L} \left[ \left(\prod_{j=1}^{n}CZ_{j,j+1}\right) \left(\prod_{j=1}^{n}\hat{R}_z(\phi_{\ell j})\hat{R}_y(\vartheta_{\ell j})\right) \right], \label{eq:sm95ringcz} \end{align}\tag{109}\] with periodic boundary conditions. All parameters are sampled independently from \([0,2\pi)\). This family is useful as a first analysis example because it exhibits the separation between three depth-dependent events: coverage saturation, mean-EP growth, and variability/gradient collapse.

The Sim-style coverage proxy is SM-Sim2019expressibility? \[\begin{align} X_L=\exp[-D_{\rm KL}(P_L(F)\Vert P_{\mathrm{Haar}}(F))], \label{eq:sm95expr95proxy} \end{align}\tag{110}\] where \(P_L(F)\) is the empirical pairwise-fidelity distribution obtained from parameter sampling and the reference input \(\left|0\right\rangle^{\otimes n}\). A larger \(X_L\) means that the parameter-sampled fidelity distribution is closer to the Haar fidelity distribution. This quantity is intentionally included as a contrast: it is a parameter-ensemble coverage proxy, whereas EP and EPD are fixed-circuit product-input descriptors.

For the onset comparison, we normalize increasing quantities as \[\begin{align} \widetilde{X}_L=\frac{X_L}{\max_{L'}X_{L'}}, \qquad \widetilde{\mathrm{EP}}_L=\frac{\mathrm{EP}_L-\min_{L'}\mathrm{EP}_{L'}}{\max_{L'}\mathrm{EP}_{L'}-\min_{L'}\mathrm{EP}_{L'}}, \label{eq:sm95normalized95increasing} \end{align}\tag{111}\] and define collapse ratios \[\begin{align} R_{\mathrm{EPD}}(L)=\frac{\mathrm{EPD}_L}{\mathrm{EPD}_1}, \qquad R_g(L)=\frac{\overline{V}_L}{\overline{V}_1}. \label{eq:sm95collapse95ratios} \end{align}\tag{112}\] Here, \(\overline{V}_L\) denotes the input-induced variance of a representative finite-difference gradient at depth \(L\). Thus, \(R_{\mathrm{EPD}}\) and \(R_g\) decrease when the output-entanglement distribution and the representative gradient scale collapse relative to their \(L=1\) values.

Figure 4: Depth scan for the ring-CZ hardware-efficient ansatz. The four panels show the Sim-style coverage proxy, family-averaged EP, family-averaged EPD, and the input-induced variance of a representative gradient. The key feature is the mismatch between the early saturation of the coverage proxy and the later collapse of EPD and the gradient scale.

Fig. 4 should be read as follows: (a) The coverage proxy \(X_L\) rises rapidly with depth. For the finite-size data shown here, the \(n=6\) curve is already close to its scan maximum at small depth. This means that the parameter ensemble has quickly learned to reproduce a Haar-like fidelity distribution. (b) The mean fixed-circuit EP generally increases with depth, but it does not carry the same information as the coverage proxy. It measures the average Meyer-Wallach entanglement generated from product inputs by a fixed circuit and then averaged over circuit instances. (c) EPD decreases after an intermediate depth. This is the narrowing of the product-input output-entanglement distribution: different product inputs begin to produce nearly the same amount of entanglement. (d) The representative gradient variance also decreases, on a logarithmic scale, in the same depth range where EPD collapses. This is consistent with the local two-copy criterion of Sec. III, although it is not a universal functional relation because the gradient variance depends on the chosen parameter and cost. Thus, Fig. 4 separates two statements that are often conflated: a circuit family can look highly covering according to a fidelity-distribution proxy before the fixed-circuit response has fully homogenized.

Figure 5: Normalized onset comparison for the n=6 ring-CZ ansatz. The coverage proxy reaches a near-ceiling value at L=2, while EPD and the representative gradient variance both pass their half-collapse thresholds at L=4 in this finite-sample run.

Fig. 5 restates the same information in threshold language. For the finite-sample \(n=6\) scan, the ordering is \[\begin{align} L^{95\%}_{\rm expr}=2, \qquad L^{90\%}_{\rm EP}=5, \qquad L^{50\%}_{\rm EPD}=4, \qquad L^{50\%}_{g}=4. \label{eq:sm95thresholds95expanded} \end{align}\tag{113}\] The meaning of these four depths is summarized in Table 1. The thresholds are not universal constants; they are a compact way to compare crossover events in one reproducible finite-size scan. The robust observation is the ordering: coverage can become large before EPD and the representative gradient variance have collapsed.

Table 1: Interpretation of the onset depths in Eq. ([eq:sm95thresholds95expanded]). The values are finite-sample diagnostics for the \(n=6\) ring-CZ scan, not universal constants.
quantity depth interpretation
\(L^{95\%}_{\rm expr}\) 2 the Sim-style coverage proxy has reached \(95\%\) of its maximum value in the depth scan; coverage is already near its finite-sample ceiling.
\(L^{90\%}_{\rm EP}\) 5 the normalized mean fixed-circuit EP reaches \(90\%\) of its scan range; average entangling strength continues to grow after the coverage proxy has saturated.
\(L^{50\%}_{\rm EPD}\) 4 EPD has fallen to one half of its \(L=1\) value; product-input response variability has significantly narrowed.
\(L^{50\%}_{g}\) 4 the representative gradient variance has fallen to one half of its \(L=1\) value; this finite-difference probe aligns with the EPD-collapse depth in this scan.

The interval between early coverage saturation and later variability collapse is the numerical prototype of the design window emphasized in the main manuscript. In this window, the ansatz is no longer trivially shallow, but it has not yet reached the low-EPD regime associated with BP-like homogenization.

Fig. 6 shows histograms of \(Q(\hat{U}_L\left|\psi\right\rangle)\) over Haar-random product inputs for \(n=6\) ring-CZ circuits. This plot is the most direct way to understand what EP and EPD measure. EP tracks the horizontal drift of the distribution; EPD tracks its width.

Figure 6: Collapse of the output-entanglement distribution with depth. Histograms of Q(\hat{U}_L\left|\psi\right\rangle) are shown for representative n=6 ring-CZ circuits at L=1,3,7. Increasing depth shifts the distribution to larger Q and then narrows it.

The three histograms should be read as follows. At \(L=1\), the circuit is shallow. The distribution is relatively broad because different product inputs experience noticeably different entangling responses. At the intermediate depth, the distribution has moved to the right while still retaining a visible width; this is the desired qualitative signature of high coverage without full homogenization. At larger depth, the histogram is concentrated near a high-entanglement value. The mean is high, but the width is small. This last regime is precisely why EP alone is insufficient: a high-EP circuit can be either useful and still input-sensitive, or high-EP and already homogenized.

A small EPD therefore has two different meanings depending on coverage. A shallow circuit can have small EPD simply because it barely entangles any product input. A deep circuit can have small EPD because it entangles almost all product inputs in the same way. For this reason, EPD should always be read together with a coverage coordinate.

5.2 Ansatz trajectory scan↩︎

The depth scan above studies one ansatz family. The route scan generalizes the same idea to five families: ring-CZ HEA, line-CNOT HEA, brickwork RXX, QAOA-style ring-ZZ, and ring-CRX HEA. Fig. 7 shows the representative \(n=5\) and \(L=1\) circuit blocks used in the route scan. The drawings use five wires only as a schematic stencil; for larger \(n\), the same local pattern is extended along the line or ring.

Figure 7: Representative L=1 circuit blocks for the five ansatz families used in the route scan SM-Sim2019expressibility?, SM-Kandala2017HardwareEfficient?, SM-Farhi2014QAOA?. Top row: (a) ring-CZ HEA, (b) line-CNOT HEA, and (c) brickwork RXX. Bottom row: (d) QAOA-style ring-ZZ and (e) ring-CRX HEA. The figure is drawn for five qubits for readability; the simulations extend the same connectivity pattern to n=5,10,15,20.

The route-scan simulations use the following elementary gates: \[\begin{align} \hat{R}_X(\theta)=e^{-i\theta \hat{X}/2}, \quad \hat{R}_Y(\theta)=e^{-i\theta \hat{Y}/2}, \quad \hat{R}_Z(\theta)=e^{-i\theta \hat{Z}/2}, \quad \hat{R}_{XX}(\theta)=e^{-i\theta \hat{X}\otimes \hat{X}/2}, \quad \hat{R}_{ZZ}(\theta)=e^{-i\theta \hat{Z}\otimes \hat{Z}/2}, \end{align}\] with standard \(CZ\), \(CNOT\), and controlled-\(\hat{R}_X\) gates. The five ansatz families are:

  • RingCZ-HEA: each layer applies \(\hat{R}_Y\) and \(\hat{R}_Z\) rotations to every qubit, followed by a periodic ring of \(CZ\) gates.

  • LineCNOT-HEA: each layer applies \(\hat{R}_Y\) and \(\hat{R}_Z\) rotations to every qubit, followed by a nearest-neighbor CNOT chain. The control direction is alternated in the implementation across layers to avoid a fixed orientation bias.

  • Brickwork-RXX: each layer applies \(\hat{R}_Y\) and \(\hat{R}_Z\) rotations to every qubit, followed by \(\hat{R}_{XX}\) gates on alternating even and odd nearest-neighbor bonds.

  • QAOA-RingZZ: each layer applies \(\hat{R}_{ZZ}(\gamma_\ell)\) gates on the ring followed by single-qubit \(\hat{R}_X(\beta_\ell)\) mixers.

  • RingCRX-HEA: each layer applies \(\hat{R}_Y\) and \(\hat{R}_Z\) rotations to every qubit, followed by controlled-\(\hat{R}_X\) gates on a ring, with the directed pattern alternated across layers.

These families were chosen to produce visibly different route geometries. Line-like entanglers can move rapidly along the coverage direction but may also lose EPD quickly. Brickwork interactions can preserve variability longer while accumulating coverage more slowly. Ring-based and controlled-rotation hardware-efficient families often create intermediate-depth windows. QAOA-style ring-ZZ uses a more constrained parameter structure and can enter the sweet region at a different depth scale.

The representative gradient-variance probe uses the global projector cost \(\hat{O}_0=\left|0^n\right\rangle\!\!\left\langle 0^n\right|\) and the first trainable rotation parameter of each ansatz family. For hardware-efficient ansatze this is the first-layer single-qubit \(\hat{R}_Y\) parameter on qubit \(0\); for the QAOA-style ansatz it is the first mixer angle. In the reproduction package this diagnostic is evaluated with a small-budget central finite difference. It is included only to compare the EPD collapse with a concrete cost-dependent gradient scale.

The simulation budgets are intentionally modest and are stored in the metadata file. They are summarized in Table 2 below. Increasing these budgets improves Monte Carlo precision without altering the pipeline.

Table 2: Monte Carlo budgets used in the route-scan script. \(N_\theta\) is the number of parameter samples at each ansatz-depth-width point, \(N_\psi\) is the number of product-input probes per circuit sample, and \(N_{\rm targ}\) is the number of one-qubit reductions sampled when target-qubit subsampling is used for larger systems.
qubits \(n\) \(N_\theta\) \(N_\psi\) \(N_{\rm targ}\)
5 16 40 5
10 8 24 6
15 4 12 4
20 2 4 2

At each \(n\in\{5,10,15,20\}\) and \(L=1,\ldots,6\), the code estimates \[\begin{align} \overline{\mathrm{EP}}_{A,L}^{(n)}=\mathbb{E}_{\boldsymbol{\theta}\sim\nu_{A,L}^{(n)}}[\mathrm{EP}(\hat{U}_{A,L}^{(n)}(\boldsymbol{\theta}))], \qquad \overline{\mathrm{EPD}}_{A,L}^{(n)}=\mathbb{E}_{\boldsymbol{\theta}\sim\nu_{A,L}^{(n)}}[\mathrm{EPD}(\hat{U}_{A,L}^{(n)}(\boldsymbol{\theta}))], \label{eq:sm95family95averages} \end{align}\tag{114}\] where \(A\) labels the ansatz family and \(\nu_{A,L}^{(n)}\) is the ansatz-, depth-, and size-dependent product measure over trainable angles. Eq. (114 ) again encodes the fixed-circuit ordering: for each sampled \(\boldsymbol{\theta}\), EP and EPD are first estimated over product-Haar inputs, and only then averaged over the parameter ensemble.

5.2.1 Algorithmic pipeline↩︎

The numerical pipeline is deliberately written in fixed-circuit order. For each sampled parameter vector, the code first estimates the product-input distribution of \(Q(\hat{U}_L(\boldsymbol{\theta})\left|\psi\right\rangle)\) and records its mean and standard deviation. Only after this fixed-circuit step are the resulting EP and EPD values averaged over parameter samples. This order is essential: averaging \(Q\) over parameter samples from a fixed reference input would reproduce the Sim-style entangling-capability convention rather than the fixed-circuit EP/EPD convention used here.

The numerical steps are summarized in Algorithm 8.

Figure 8: Fixed-circuit EP/EPD route scan and gradient diagnostic

For the probe, we use \[\begin{align} g_i(\psi)\simeq \frac{\left|\left\langle 0^n\right|\hat{U}(\boldsymbol{\theta}_{+i})\left|\psi\right\rangle\right|^2- \left|\left\langle 0^n\right|\hat{U}(\boldsymbol{\theta}_{-i})\left|\psi\right\rangle\right|^2}{2\delta}, \qquad \delta=10^{-6}, \label{eq:sm95fd95gradient} \end{align}\tag{115}\] where \(\boldsymbol{\theta}_{\pm i}\) denotes the same parameter vector as \(\boldsymbol{\theta}\) except that the \(i\)th component is shifted to \(\theta_i\pm\delta\). The analytic results in the main manuscript do not depend on this numerical probe. It is a cost-dependent consistency check showing that low-EPD regions often coincide with smaller representative gradient scales, as expected from the local two-copy criterion, but it is not a substitute for a task-specific trainability analysis.

5.2.2 Operational definition of the sweet spot↩︎

For each qubit number \(n\), we convert the family-averaged descriptors into two panel-normalized coordinates, \[\begin{align} x_{A,L}^{(n)}=\frac{\overline{\mathrm{EP}}_{A,L}^{(n)}}{Q_{\mathrm{Haar}}^{(n)}}, \qquad y_{A,L}^{(n)}=\frac{\overline{\mathrm{EPD}}_{A,L}^{(n)}}{\max_{A,L}\overline{\mathrm{EPD}}_{A,L}^{(n)}}, \label{eq:sm95route95xy} \end{align}\tag{116}\] where \[\begin{align} Q_{\mathrm{Haar}}^{(n)}=\frac{2^n-2}{2^n+1} \label{eq:sm95qhaar95route} \end{align}\tag{117}\] is the Haar mean of the Meyer-Wallach entanglement. The horizontal coordinate \(x\) is the coverage dial: it measures how close the family-averaged mean entanglement is to the Haar benchmark at the same \(n\). The vertical coordinate \(y\) is the variability dial: it measures how much input-dependent output-entanglement spread remains relative to the largest EPD value observed in the same \(n\) panel.

The shaded regions in the main Fig. 1 are fixed by the route data, not by hand-tuned constants. For a finite set \(Z=\{z_m\}_{m=1}^M\subset[0,1]\), let \(\mathcal{T}(Z)\) be the midpoint cuts between adjacent sorted distinct values. For \(\tau\in\mathcal{T}(Z)\) define \[\begin{align} W_Z(\tau)=\sum_{z_m<\tau}(z_m-\bar z_<)^2+ \sum_{z_m\ge\tau}(z_m-\bar z_\ge)^2, \label{eq:sm95within95cluster} \end{align}\tag{118}\] where \(\bar z_<\) and \(\bar z_\ge\) are the means of the two groups. We set \[\begin{align} \tau_*(Z)=\arg\min_{\tau\in\mathcal{T}(Z)} W_Z(\tau). \label{eq:sm95tau95star} \end{align}\tag{119}\] Equivalently, \(\tau_*\) is the one-dimensional Otsu/Fisher split that maximizes the separation between the two classes while keeping each class internally compact SM-Otsu1979Threshold?, SM-Fisher1936Taxonomic?. This is a useful choice because it introduces no external threshold, percentile, or gradient information.

The coverage boundary is the split of all coverage coordinates in a panel, \[\begin{align} x_c^{(n)}=\tau_*\!\left(\{x_{A,L}^{(n)}\}_{A,L}\right). \label{eq:sm95xc95rule} \end{align}\tag{120}\] The variability boundary is then computed only inside the high-coverage subset, \[\begin{align} y_c^{(n)}=\tau_*\!\left(\{y_{A,L}^{(n)}:x_{A,L}^{(n)}\ge x_c^{(n)}\}_{A,L}\right). \label{eq:sm95yc95rule} \end{align}\tag{121}\] The conditional second step is essential. Low variability at low coverage can simply mean that the circuit is too weak to produce a broad entanglement distribution. It is not evidence of BP-like homogenization. The EPD-collapse cut is therefore learned only from points that already passed the coverage cut.

The three regions used in the route scan are \[\begin{align} \mathcal{U}^{(n)} &=& \{(A,L):x_{A,L}^{(n)}<x_c^{(n)}\},\nonumber\\ \mathcal{S}^{(n)} &=& \{(A,L):x_{A,L}^{(n)}\ge x_c^{(n)},\;y_{A,L}^{(n)}\ge y_c^{(n)}\},\nonumber\\ \mathcal{P}^{(n)} &=& \{(A,L):x_{A,L}^{(n)}\ge x_c^{(n)},\;y_{A,L}^{(n)}<y_c^{(n)}\}. \label{eq:sm95regions} \end{align}\tag{122}\] Here, \(\mathcal{U}\) is the underexpressive region: the coverage dial has not crossed the data-driven coverage cut. High variability here is not enough, because the family has not yet reached substantial coverage. \(\mathcal{S}\) is the sweet spot: the coverage dial is high and the variability dial remains open. This is the desired high-coverage/nonhomogenized design window. \(\mathcal{P}\) is the near-plateau or homogenized high-coverage region

The reproduced route scan gives \[\begin{align} (x_c^{(5)},y_c^{(5)})&=&(0.731,0.472),\nonumber\\ (x_c^{(10)},y_c^{(10)})&=&(0.700,0.352),\nonumber\\ (x_c^{(15)},y_c^{(15)})&=&(0.698,0.321),\nonumber\\ (x_c^{(20)},y_c^{(20)})&=&(0.712,0.263). \label{eq:sm95threshold95values} \end{align}\tag{123}\] The absolute threshold values should not be interpreted as universal physical constants. What is meaningful is the route geometry: a point in \(\mathcal{S}^{(n)}\) has already moved far along the coverage direction but still retains a broad output-entanglement distribution over product inputs, whereas a point in \(\mathcal{P}^{(n)}\) has comparable coverage but has lost that variability.

For a compact one-number ranking inside each panel, we use the descriptive sweet score \[\begin{align} s_{A,L}^{(n)}=x_{A,L}^{(n)}y_{A,L}^{(n)}. \label{eq:sm95sweet95score} \end{align}\tag{124}\] The score rewards points that are simultaneously far to the right and high in the two-dial plane. It is not a new objective function, and it is not used as a theorem hypothesis. In particular, an underexpressive point with very large \(y\) can still have a sizable \(s\). Therefore Eq. (124 ) must be read together with the region classification in Eq. (122 ).

Table 3: Meaning of the factors in the sweet score \(s_{A,L}^{(n)}=x_{A,L}^{(n)}y_{A,L}^{(n)}\).
symbol role interpretation
\(x_{A,L}^{(n)}\) coverage factor large when the family-averaged EP is close to the Haar entanglement benchmark. It prevents a shallow low-coverage circuit from being selected only because it has high variability.
\(y_{A,L}^{(n)}\) variability factor large when the fixed-circuit product-input response remains broad. It distinguishes high-coverage/nonhomogenized circuits from high-coverage/homogenized circuits.
\(s_{A,L}^{(n)}\) descriptive rank useful for ranking candidate depths within a panel, but not a universal trainability metric. It should be interpreted only after checking whether the point lies in \(\mathcal U\), \(\mathcal S\), or \(\mathcal P\).

5.2.3 Detailed reading of the route geometry↩︎

Fig. 1 in the main manuscript plots the trajectories of the five ansatz families in the \((x,y)\) plane as \(L\) increases from \(1\) to \(6\). The route geometry carries the physical information. Moving to the right means that the family has gained mean entangling coverage. Moving downward means that the product-input response has homogenized. A useful route moves sufficiently far right before it moves too far down.

For \(n=5\), two types of behavior are already visible. LineCNOT-HEA starts in the sweet region and then rapidly moves toward high coverage with reduced variability. RingCZ-HEA and RingCRX-HEA pass through an intermediate region before descending toward lower EPD. Brickwork-RXX retains appreciable variability but remains mostly under the coverage cut. QAOA-RingZZ reaches the sweet region only at later depth. Thus even for small \(n\), EP alone would merge late high-coverage points with genuinely useful sweet-spot points, while EPD separates them. For \(n=10\), the intermediate-depth windows become clearer. RingCZ-HEA and RingCRX-HEA develop visible sweet-spot portions before falling into the near-plateau sector. LineCNOT-HEA reaches high coverage early and then loses variability quickly. Brickwork-RXX remains mostly underexpressive, despite retaining width in the output-entanglement distribution. QAOA-RingZZ approaches the useful region more slowly. The lesson is that faster coverage is not automatically better: a route that arrives at high \(x\) only after \(y\) has collapsed is less attractive than one that enters \(\mathcal{S}^{(n)}\). For \(n=15\), the distinction between coverage and homogenization is sharper. LineCNOT-HEA and the deepest RingCZ-HEA points lie at high \(x\) but low \(y\), illustrating why a one-dimensional EP-vs-BP reading is incomplete. RingCZ-HEA and RingCRX-HEA still show middle-depth sweet windows, whereas QAOA-RingZZ reaches a useful region at a later depth. Brickwork-RXX again demonstrates the opposite failure mode: noncollapsed variability without enough coverage. For \(n=20\), the newly added panel confirms that the two-dial interpretation is not restricted to the smaller scans. The value of \(y_c^{(20)}\) is lower than in the smaller panels because it is a panel-wise split of the normalized diagnostic cloud, not a universal BP threshold. RingCZ-HEA retains a middle-depth high-coverage/nonhomogenized window before descending. LineCNOT-HEA largely occupies the high-coverage/low-variability side, with only a narrow sweet-spot passage. Brickwork-RXX and QAOA-RingZZ reach the useful region later. RingCRX-HEA also illustrates the importance of depth scheduling: too shallow is underexpressive, whereas too deep tends toward homogenization.

The route-scan conclusion is therefore structural rather than family-specific. The different entanglers and connectivities trace different paths through the two-dial plane. The design task is to choose circuit blocks, connectivities, and depth schedules that move far enough along the coverage dial while keeping the variability dial open.

Table 4 below also explains why the route plot should be read panel by panel. The coverage split is relatively stable near \(0.7\) for this diagnostic cloud, while the variability split decreases with \(n\). This does not mean that \(y_c\) is a system-size-independent trainability threshold. It means that the normalized EPD cloud changes with system size and must be separated within each panel.

Table 4: Panel-wise thresholds and region counts for the reproduced route scan. The counts are over the \(5\) ansatz families and \(6\) depths in each \(n\) panel.
\(n\) \(x_c^{(n)}\) \(y_c^{(n)}\) \(\abs{\mathcal U^{(n)}}\) \(\abs{\mathcal S^{(n)}}\) \(\abs{\mathcal P^{(n)}}\)
5 0.731 0.472 14 6 10
10 0.700 0.352 15 8 7
15 0.698 0.321 15 6 9
20 0.712 0.263 15 7 8

Table ¿tbl:tab:sm95sweet95summary? below lists the data-driven thresholds and, for each ansatz family and qubit number, the depth that maximizes \(s_{A,L}^{(n)}\). The additional region column is important: the largest sweet score for a family is not automatically a sweet-spot point if the coverage cut is not passed.

max width=0.98

The score table reinforces the same caution. In several rows, the maximum score occurs in \(\mathcal{U}^{(n)}\) because \(y\) is large while \(x\) is still below the coverage cut. These points should not be read as final sweet-spot candidates. The practical screening rule is: first identify whether the route reaches \(\mathcal{S}^{(n)}\), then use the score and task-specific gradient checks to rank nearby depths.

6 Task-level validation on QML benchmarks↩︎

The route scan in the previous Sec. V is task-independent. It uses only fixed-circuit EP/EPD descriptors and therefore cannot, by itself, certify high accuracy for every possible dataset, data encoding, optimizer, readout, and loss. A natural question is therefore whether the high-coverage/nonhomogenized region selected by the two-dial scan is truly useful in an actual learning loop. In this section, we address this question by adding a small descriptor-blind QML validation. The purpose is not to turn the sweet score into a universal accuracy theorem. Rather, it is to test whether the scan identifies ansatz-depth points at which task-specific training is most worthwhile.

The validation uses the \(n=5\) route cloud from Fig. 1 of the main text. All \(30\) ansatz-depth points, namely the five ansatz families at \(L=1,\ldots,6\), are trained under the same supervised-learning protocol. The EP/EPD coordinates, region assignments, and sweet scores are fixed before training and are not adjusted using the task labels. Thus, the comparison is descriptor-blind: the learning task only tests the candidates selected by the scan. The two summary figures for this validation are included in the main text; the present section keeps the full protocol, tables, and interpretation.

6.1 Benchmark construction and training protocol↩︎

We use a simple teacher-student binary classification benchmark. For each input \(\boldsymbol{x}=(x_1,x_2)\in[-1,1]^2\), the data embedding prepares \[\begin{align} \left|\eta(\boldsymbol{x})\right\rangle=\bigotimes_{j=1}^{5}\hat{R}_z\!\left(b_j(\boldsymbol{x})\right)\hat{R}_y\!\left(a_j(\boldsymbol{x})\right)\left|0\right\rangle, \label{eq:sm95task95embedding} \end{align}\tag{125}\] where the angles \(a_j,b_j\) are fixed linear combinations of \(x_1\) and \(x_2\). The same embedding is used for every ansatz-depth point. We then generate labels using two fixed sweet-region teacher circuits, \(\hat{T}_1=({\rm RingCZ\text{-}HEA},L=3)\) and \(\hat{T}_2=({\rm RingCRX\text{-}HEA},L=3)\), both of which lie in the green region of the \(n=5\) route scan. For teacher \(\hat{T}\), we define \[\begin{align} z_T(\boldsymbol{x})=\left\langle \eta(\boldsymbol{x})\right|\hat{U}_T^\dagger \hat{Z}_0 \hat{U}_T\left|\eta(\boldsymbol{x})\right\rangle. \label{eq:sm95teacher95score} \end{align}\tag{126}\] The binary label is obtained by thresholding \(z_T(\boldsymbol{x})\) at the median value on the training set, so that the two classes are approximately balanced: \[\begin{align} y_T(\boldsymbol{x})=\begin{cases} +1, & z_T(\boldsymbol{x})\ge m_T,\\ -1, & z_T(\boldsymbol{x})<m_T, \end{cases} \qquad m_T={\rm median}_{\boldsymbol{x}\in\mathcal{D}_{\rm train}}z_T(\boldsymbol{x}). \label{eq:sm95teacher95label} \end{align}\tag{127}\] This teacher-student construction is useful because it probes an actual QML learning loop while keeping the target function in the representational regime that the two-dial scan is designed to find: high coverage is useful, but the teacher itself is not taken from the low-EPD near-plateau region.

For a student ansatz \((A,L)\), the trainable prediction is \[\begin{align} f_{A,L}(\boldsymbol{x};\boldsymbol{\theta})=\left\langle \eta(\boldsymbol{x})\right|\hat{U}_{A,L}(\boldsymbol{\theta})^\dagger \hat{Z}_0 \hat{U}_{A,L}(\boldsymbol{\theta})\left|\eta(\boldsymbol{x})\right\rangle, \label{eq:sm95student95readout} \end{align}\tag{128}\] and the loss is the mean squared error \[\begin{align} \mathcal{L}_{A,L}(\boldsymbol{\theta})=\frac{1}{|\mathcal{D}_{\rm train}|}\sum_{(\boldsymbol{x},y)\in\mathcal{D}_{\rm train}}\left(f_{A,L}(\boldsymbol{x};\boldsymbol{\theta})-y\right)^2. \label{eq:sm95task95loss} \end{align}\tag{129}\] All circuits are initialized from the same type of random angle distribution used in the route scan, namely independent uniform angles in \([0,2\pi)\). We train with full-batch Adam-SPSA for \(50\) iterations. SPSA is used only to keep the benchmark architecture-neutral: each update requires two loss evaluations independent of the number of parameters. For each ansatz-depth point and each teacher task, we use two independent optimizer seeds. The training set has \(40\) samples and the test set has \(96\) samples.

The validation metrics are listed in Table 5. A run is called successful when its final test accuracy is at least \(75\%\). The initial gradient scale is the per-parameter squared norm of the first SPSA gradient estimate, \[\begin{align} G_0=\frac{1}{p_{A,L}}\left\|\widehat{\nabla}_{\rm SPSA}\mathcal{L}_{A,L}(\boldsymbol{\theta}_0)\right\|_2^2, \label{eq:sm95initial95spsa95scale} \end{align}\tag{130}\] which should be read as an optimizer-level proxy for the available initial training signal, not as the exact local two-copy variance of Sec. III.

Table 5: Task-level validation metrics. The benchmark is not used to define EP, EPD, or the sweet spot; it is used only after the descriptor scan to test whether the selected ansatz-depth regions behave as useful QML candidates.
metric symbol interpretation
final test accuracy \(A_{\rm test}\) classification accuracy after the fixed training budget
final training loss \(\mathcal L_{\rm final}\) residual optimization error on the training set
success probability \(P_{\rm succ}\) fraction of runs with \(A_{\rm test}\ge0.75\)
first success iteration \(T_{\rm succ}\) first logged iteration at which the success threshold is reached
initial gradient scale \(G_0\) per-parameter squared norm of the first SPSA gradient estimate

6.2 Results: sweet region as a useful screening window, not a guarantee↩︎

Table 6 groups the task-level results by the three EP/EPD regions. The most relevant comparison for the two-dial claim is the high-coverage comparison between the green and red regions. Both have passed the coverage cut, but they differ in whether the variability dial remains open. In this comparison, the sweet region has higher median test accuracy, lower median final loss, larger success probability, and a much faster first-success time than the near-plateau region. In particular, the median first-success iteration is \(5\) in the sweet region but \(20\) in the near-plateau region. This supports the intended interpretation of the green region as the depth window in which high-coverage training is most worthwhile.

Table 6: Region-wise task-level validation for the \(n=5\) route cloud. The table reports medians over all ansatz-depth points, the two teacher tasks, and two optimizer seeds. The success threshold is \(A_{\rm test}\ge0.75\).
region runs median \(A_{\rm test}\) median \(\mathcal L_{\rm final}\) \(P_{\rm succ}\) median \(T_{\rm succ}\) median \(G_0\)
underexpressive 56 0.682 0.793 0.357 12.5 \(3.51\times10^{-1}\)
sweet spot 24 0.656 0.804 0.333 5.0 \(1.72\times10^{-1}\)
near plateau 40 0.609 0.865 0.200 20.0 \(2.91\times10^{-1}\)

The underexpressive region deserves a careful interpretation. Some gray-region circuits perform competitively on these small \(n=5\) teacher tasks. This is not a contradiction of the two-dial framework. It means that a small finite task may not require the full coverage that the descriptor scan regards as desirable. The gray region can therefore contain trainable circuits for simple tasks, but it is not the high-coverage design window targeted by the main Letter. The operational question addressed by the sweet spot is different: once high coverage has been reached, should one continue increasing depth into the homogenized red sector? The answer from Table 6 is no; the high-coverage/noncollapsed region gives better task-level behavior than the high-coverage/low-EPD region under the same training protocol.

The two task-level validation plots have been moved to the main text. The EP/EPD scatter plot shown there gives the same result directly in the route plane: the color indicates the empirical success probability of each ansatz-depth point, and the high-coverage/low-EPD sector contains several low-success points. Successful points are more often found away from the collapsed lower-right corner. This is exactly the design information lost by projecting the route cloud onto the coverage coordinate alone.

The region-wise diagnostic plot in the main text isolates three aggregate quantities. The first panel shows that sweet-region success is slightly below the gray region but above the red near-plateau region. The second panel is the most diagnostic for trainability: among successful runs, the sweet region reaches the success threshold earlier than both gray and red regions. The third panel reports the initial SPSA gradient scale. These quantities are not expected to be monotonic functions of EPD, because the exact gradient object is local and cost-dependent. Nevertheless, the red sector does not provide an advantage despite having the largest coverage; this is the task-level counterpart of the main two-dial message.

We also compare simple screening rules in Table 7 below. The rule based on the product \(s=xy\) selects points that balance coverage and variability and gives a larger success probability than selecting by coverage alone. Selecting by variability alone can also perform well in this small task, but it includes low-coverage gray points; this is why the main text defines the sweet spot by applying the variability cut only after the coverage cut. Depth alone and parameter count alone are not reliable substitutes for the two-dial geometry.

Table 7: Comparison of simple descriptor-blind screening rules. Each rule selects six ansatz-depth points from the \(n=5\) route cloud before any task training. The table then reports the median test accuracy, mean success probability, and median initial gradient scale of the selected points.
screening rule median \(A_{\rm test}\) mean \(P_{\rm succ}\) median \(G_0\)
sweet score \(s=xy\) 0.648 0.375 \(3.51\times10^{-1}\)
coverage \(x\) only 0.570 0.125 \(4.11\times10^{-1}\)
variability \(y\) only 0.703 0.417 \(5.26\times10^{-1}\)
deepest \(L\) only 0.698 0.333 \(2.72\times10^{-1}\)
fewest parameters 0.589 0.208 \(3.22\times10^{-1}\)
all points 0.641 0.300 \(3.46\times10^{-1}\)

Finally, Table 8 reports Spearman rank correlations between scalar predictors and task success. The coverage coordinate alone is negatively correlated with success in this benchmark, because many of the highest-coverage points lie in the low-EPD near-plateau sector. By contrast, the variability coordinate and the sweet score have positive correlations with success. These numbers should not be overinterpreted as universal constants. Their role is to verify the central logic: once one moves beyond very small tasks, EP-only or coverage-only screening is incomplete, and retaining noncollapsed variability is useful information for ansatz selection.

Table 8: Spearman rank correlation between descriptor-level predictors and success probability in the task-level validation. Positive values indicate that larger predictor values tend to rank ansatz-depth points with higher success probability.
predictor Spearman \(\rho\) with \(P_{\rm succ}\)
coverage dial \(x\) -0.251
variability dial \(y\) 0.374
sweet score \(s=xy\) 0.364
mean EP -0.251
EPD 0.374
depth \(L\) -0.061
parameter count -0.102

The validation therefore answers the practical question in a deliberately limited sense. A high sweet score is not a universal guarantee of high QML accuracy, and the present paper does not claim such a no-free-lunch reversal. Instead, the EP/EPD scan is a pre-training filter. It identifies the regime in which coverage has been gained without obvious homogenization, and the task-level check shows that this regime is a better high-coverage target than the low-EPD near-plateau sector.

7 Interpretive remarks and limitations of the diagnostic↩︎

The central theoretical hierarchy is \[\begin{align} \mathrm{coverage} &:& \operatorname{Tr}[\hat{K}_Q\hat{M}^{(2)}_{\hat{U}}],\nonumber\\ \mathrm{variability} &:& \operatorname{Tr}[\hat{T}_Q\hat{M}^{(4)}_{\hat{U}}]-\mathrm{EP}^2,\nonumber\\ \mathrm{trainability} &:& \operatorname{Tr}[(\hat{B}_i\otimes \hat{B}_i)\hat{M}^{(2)}_{i,\mathrm{LC}}]. \label{eq:sm95final95hierarchy} \end{align}\tag{131}\] The first term is global and mean-like, the second is global but fluctuation-sensitive, and the third is local and parameter-resolved. This is why the Letter does not claim that EPD is the barren-plateau invariant. Rather, EPD is a compact fourth-moment diagnostic of the same homogenization process that, when seen locally by a parameter light cone, suppresses gradient variance.

This distinction also clarifies how to interpret small EPD. A shallow underexpressive circuit may have small EPD because it generates little entanglement at all. A deep randomizing circuit may have small EPD because it generates almost the same high entanglement for most product inputs. The diagnostic is therefore most informative in combination with a coverage coordinate: high normalized EP together with low EPD indicates high-coverage/low-variability homogenization. The route plot in the Letter is designed precisely to show this joint behavior.

Several extensions are natural. One may replace the Meyer-Wallach functional by other entanglement or correlation functionals; one may replace the Haar-product input distribution by a data-conditioned distribution; and one may adapt the moment formulas to noisy channels. These generalizations change the operational meaning of the descriptors but not the main lesson: mean coverage, fluctuation width, and local trainability are distinct objects and should not be collapsed into a single scalar trade-off.

References↩︎

[1]
A. Peruzzo, J. McClean, P. Shadbolt, M.-H. Yung, X.-Q. Zhou, P. J. Love, A. Aspuru-Guzik, and J. L. O’Brien, https://doi.org/10.1038/ncomms5213.
[2]
E. Farhi, J. Goldstone, and S. Gutmann, A quantum approximate optimization algorithm(2014), https://arxiv.org/abs/1411.4028.
[3]
A. Kandala, A. Mezzacapo, K. Temme, M. Takita, M. Brink, J. M. Chow, and J. M. Gambetta, https://doi.org/10.1038/nature23879.
[4]
V. Havlíček, A. D. Córcoles, K. Temme, A. W. Harrow, A. Kandala, J. M. Chow, and J. M. Gambetta, https://doi.org/10.1038/s41586-019-0980-2.
[5]
M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio, and P. J. Coles, https://doi.org/10.1038/s42254-021-00348-9.
[6]
S. Sim, P. D. Johnson, and A. Aspuru-Guzik, https://doi.org/10.1002/qute.201900070.
[7]
P. Zanardi, C. Zalka, and L. Faoro, https://doi.org/10.1103/PhysRevA.62.030301.
[8]
J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, and H. Neven, https://doi.org/10.1038/s41467-018-07090-4.
[9]
M. Cerezo, A. Sone, T. Volkoff, L. Cincio, and P. J. Coles, https://doi.org/10.1038/s41467-021-21728-w.
[10]
Z. Holmes, K. Sharma, M. Cerezo, and P. J. Coles, https://doi.org/10.1103/PRXQuantum.3.010313.
[11]
T. Hubregtsen, J. Pichlmeier, P. Stecher, and K. Bertels, https://doi.org/10.1007/s42484-021-00038-w.
[12]
K. Cho and J. Bang, Physical Review A 113, 012442 (2026), https://arxiv.org/abs/2508.00301.
[13]
D. A. Meyer and N. R. Wallach, https://doi.org/10.1063/1.1497700.
[14]
N. Otsu, https://doi.org/10.1109/TSMC.1979.4310076.