Optimal Shadow Estimation with Minimal Measurement Settings


Abstract

Shadow estimation is a powerful framework for predicting quantum properties from randomized measurements. While \(3\)-design protocols achieve optimal worst-case performance, the minimal number of measurement bases required for such optimality has remained open. Here we prove that \(\Theta(d^2)\) measurement bases are both necessary and sufficient for worst-case optimal shadow estimation and construct an explicit basis family. In stark contrast, any state \(2\)-design already suffices for average-case optimality: the mean squared shadow norm of normalized observables is bounded by a universal constant, and we prove strong concentration for Haar-random states, yielding constant sample complexity for generic pure-state fidelity estimation. Easily implementable \(2\)-designs—from mutually unbiased bases, cyclic measurements, or shallow \(\mathcal{O}(\log n)\)-depth circuits—enable optimal average-case protocols with remarkably simple measurement strategies. Our results establish a fundamental complexity separation: worst-case estimation requires \(\Theta(d^2)\) bases, whereas average-case performance requires only \(\Theta(d)\) bases, with broad implications for quantum information theory and near-term experiments.

Introduction—Characterizing quantum systems efficiently is a central challenge in quantum science and technology. Full quantum state tomography demands resources that scale exponentially with system size, motivating the development of more targeted approaches. The classical shadow framework [1] provides a powerful alternative: randomized measurements paired with classical post-processing can predict key properties of an unknown quantum state—including expectation values, fidelities, and entanglement witnesses—without full state reconstruction. Shadow estimation has since found broad applications in fidelity estimation [2][5], entanglement detection [6][11], and Hamiltonian learning [12][15], with experimental demonstrations on photonic [16], [17], superconducting [18], [19], and trapped-ion [6], [20] platforms. In practice, the number of distinct measurement bases directly determines the calibration and implementation cost of a shadow protocol, making the identification of minimal measurement schemes a pressing concern.

A shadow estimation protocol is specified by a measurement ensemble \(\mathcal{E}\) of weighted pure states forming a rank-one positive operator-valued measure (POVM) [1], [21], [22]. Its efficiency is governed by the shadow norm \(\|O\|_\mathcal{E}\) of the target observable \(O\), which characterizes the single-shot estimation variance: to estimate an expectation value within additive error \(\varepsilon\), \(\mathcal{O}(\|O\|_\mathcal{E}^2/\varepsilon^2)\) samples suffice regardless of the unknown state. Protocols based on unitary 3-designs—such as the multiqubit Clifford group [23], [24]—achieve the optimal worst-case shadow norm \(\|O\|_\mathcal{E}^2\leq 3\|O\|_2^2\) [1], but require large, highly structured ensembles, which are challenging to implement on near-term hardware. Simpler schemes based on state 2-designs—such as complete sets of mutually unbiased bases (MUBs) [25][27] and cyclic measurements [28]—are far more practical, yet their worst-case squared shadow norm grows linearly in \(d\), rendering them suboptimal.

Despite rapid recent progress on shadow estimation based on various measurement ensembles [1], [29][39], existing approaches either fail to achieve uniformly efficient estimation for general observables or rely on measurements with superpolynomially many outcomes. A fundamental question therefore remains open: What is the minimal measurement complexity required for optimal shadow estimation, and can simple measurements still be efficient in typical scenarios?

Here we provide a complete answer, revealing a striking complexity separation. First, we prove that \(\Theta(d^2)\) measurement bases are both necessary and sufficient for worst-case optimal shadow estimation, and construct an explicit basis family from phase 3-designs and MUBs. Second, we show that any state 2-design already achieves average-case optimal performance with only \(\Theta(d)\) bases: the mean squared shadow norm over arbitrary target observables is bounded by \(\left(2\sqrt{2}+1\right)\|O\|_2^2\) (Theorem 3). For normalized observables, this bound evaluates to a constant independent of \(d\), close to the 3-design optimum. The worst-case suboptimality of 2-designs thus vanishes in typical scenarios. Third, for fidelity estimation of Haar-random target states, we establish strong concentration of the shadow norm, yielding constant sample complexity for generic pure states even with simple 2-design measurements.

Our results reveal a fundamental \(d^2\)-versus-\(d\) gap between worst-case and average-case measurement complexities, demonstrating that easily implementable ensembles—including complete sets of MUBs, cyclic measurements, and shallow \(\mathcal{O}(\log n)\)-depth circuits [19], [40][42]—already realize optimal protocols for typical tasks. Meanwhile, our results highlight the foundational significance of MUBs and symmetric informationally complete (SIC) measurements [43][45] in quantum learning.

Preliminaries—Let \(\mathcal{H}\) be a \(d\)-dimensional complex Hilbert space (\(d\geq 2\)). Denote by \(\mathcal{L}(\mathcal{H})\) and \(\mathcal{L}^{\mathrm{H}}_0(\mathcal{H})\) the spaces of linear operators and traceless Hermitian operators on \(\mathcal{H}\), respectively; denote by \(\mathcal{D}(\mathcal{H})\) the set of density operators, and by \(\mathcal{P}(\mathcal{H})\subset\mathcal{D}(\mathcal{H})\) the subset of pure states. For a general operator \(O\in \mathcal{L}(\mathcal{H})\), we denote by \(O_0=O-\mathop{\mathrm{tr}}(O)\mathbb{1}/d\) the traceless part of \(O\), where \(\mathbb{1}\) is the identity operator; for a pure state \(\ket{\phi}\in\mathcal{H}\), we write \(\phi=\ket{\phi}\mkern-2mu\bra{\phi}\) and \(\phi_0=\phi-\mathbb{1}/d\). The notation \(\|\cdot\|_p\) is used for the Schatten \(p\)-norm and \(\|\cdot\|\) for the operator norm.

Given a positive integer \(t\), a state \(t\)-design is a weighted ensemble of pure states whose \(t\)-th moment reproduces the Haar average; it mimics Haar-random states up to \(t\)-th order statistics [44][47]. State 2-designs can be realized by SIC ensembles [43][45], complete sets of MUBs [25][27], or cyclic constructions [28]; the set of multiqubit stabilizer states forms a state 3-design [23], [24], [48]. Unitary \(t\)-designs are defined analogously for ensembles of unitary operators.

A shadow estimation protocol is specified by a state ensemble \(\mathcal{E}=\{\ket{\phi_i},w_i\}_i\) with \(w_i>0\) and \({\sum_i w_i \phi_i=\mathbb{1}/d}\), which ensures that \(\{d\,w_i\phi_i\}_i\) forms a valid rank-one POVM on \(\mathcal{H}\). The associated measurement channel is \[\begin{align} \label{eq:MeasCh} \mathcal{M}_\mathcal{E}(O) = d\sum_i w_i \mathop{\mathrm{tr}}(\phi_i O)\,\phi_i. \end{align}\tag{1}\] When \(\mathcal{E}\) is informationally complete (IC), i.e., \(\{\phi_i\}_i\) spans \(\mathcal{L}(\mathcal{H})\), the channel \(\mathcal{M}_\mathcal{E}\) is invertible. The inverse \(\mathcal{M}_\mathcal{E}^{-1}\) acts as the reconstruction map: each snapshot \(\hat{\rho}_i \mathrel{\vcenter{:}}= \mathcal{M}_\mathcal{E}^{-1}(\phi_i)\) satisfies \(\mathbb{E}[\hat{\rho}_i] = \rho\), making it an unbiased estimator of \(\rho\). Consequently, \(\mathop{\mathrm{tr}}(O \hat{\rho}_i)\) is an unbiased estimator of \(\mathop{\mathrm{tr}}(O\rho)\) for any observable \(O\).

The sample complexity of estimating \(\mathop{\mathrm{tr}}(O\rho)\) to accuracy \(\varepsilon\) with constant success probability scales as \(\|O_0\|_\mathcal{E}^2/\varepsilon^2\) [1], where the squared shadow norm is defined as follows. For a traceless observable \(O_0\in\mathcal{L}^{\mathrm{H}}_0(\mathcal{H})\) (the traceless part of \(O\)), the state-dependent variant is \[\begin{align} \label{eq:normSD} \|O_0\|_{\mathcal{E},\rho}^2 \mathrel{\vcenter{:}}= d\sum_i w_i \mathop{\mathrm{tr}}(\phi_i\rho)\bigl[\mathop{\mathrm{tr}}\bigl(O_0\,\hat{\rho}_i\bigr)\bigr]^2, \end{align}\tag{2}\] and the state-independent (worst-case) one is \(\|O_0\|_\mathcal{E}^2 \mathrel{\vcenter{:}}= \max_{\rho\in\mathcal{D}(\mathcal{H})} \|O_0\|_{\mathcal{E},\rho}^2\). The worst-case norm provides a universal guarantee for all unknown states, while its \(\rho\)-dependent counterpart quantifies performance for a given target state and can be substantially smaller. Henceforth, observables are assumed traceless except in the context of fidelity estimation or otherwise stated.

A key quantity in our analysis is the normalized \(t\)-th frame potential \[\label{eq:FPt} \bar{\Phi}_t(\mathcal{E}) \mathrel{\vcenter{:}}= D_{[t]}\,\Phi_t(\mathcal{E}),\;\; \Phi_t(\mathcal{E}) \mathrel{\vcenter{:}}= \sum_{i,j} w_i w_j \bigl[\mathop{\mathrm{tr}}(\phi_i\phi_j)\bigr]^t,\tag{3}\] where \(D_{[t]}=\binom{d+t-1}{t}\) is the dimension of the symmetric subspace of \(\mathcal{H}^{\otimes t}\). It is well known that \(\bar{\Phi}_t(\mathcal{E})\geq 1\), with equality if and only if \(\mathcal{E}\) forms a state \(t\)-design [45][47]. In this work, the case \(t=3\) plays the central role: \(\bar{\Phi}_3(\mathcal{E})\) quantifies how closely a 2-design approximates a 3-design and governs the performance gap in shadow estimation. For any state 2-design, its magnitude is controlled by:

Proposition 1. If \(\mathcal{E}\) is a state 2-design in dimension \(d\), then \[\begin{align} \label{eq:FP3UB} \bar{\Phi}_3(\mathcal{E}) \leq \frac{(d+2)(d^2+2d-1)}{6d^2}. \end{align}\qquad{(1)}\] In particular, \(\bar{\Phi}_3(\mathcal{E})=\mathcal{O}(d)\) for any state 2-design.

Proposition 1 and other results are proved in Supplemental Material (SM) [49], which includes Refs. [50], [51].

Minimal optimal protocols in the worst-case setting—We first establish a fundamental lower bound on the worst-case shadow norm achievable with a given number of measurement outcomes.

Proposition 2. Suppose the ensemble \(\mathcal{E}\) consists of \(K\) pure states in \(\mathcal{H}\) and induces an IC-POVM. Then \[\begin{align} \label{eq:ShNormLUB} \max_{\psi\in\mathcal{P}(\mathcal{H})} \|\psi_0\|_\mathcal{E}^2 \geq\frac{(d^2-1)^2}{Kd}. \end{align}\qquad{(2)}\]

This result encodes a fundamental trade-off: fewer measurement outcomes necessarily imply larger worst-case shadow norms. Achieving optimal worst-case performance—meaning \(\|O\|_\mathcal{E}^2\leq C\|O\|_2^2\) for a universal constant \(C\)—requires at least \(\Omega(d^3)\) outcomes, or equivalently \(\Omega(d^2)\) orthonormal bases, given that each basis corresponds to \(d\) outcomes. Incidentally, when \(\mathcal{E}\) forms a state 2-design, the worst-case shadow norm satisfies \(\|O\|_\mathcal{E}^2\leq (d+1)\|O\|_2^2\) [21], [34], [35]. For typical 2-designs consisting of \(\mathcal{O}(d^2)\) states—such as SICs and complete sets of MUBs—the lower and upper bounds match up to constant factors, demonstrating that standard 2-design measurements are suboptimal for worst-case performance. We now show that the \(\Omega(d^2)\)-basis lower bound is tight.

Theorem 1. In dimension \(d\), \(\Theta(d^2)\) measurement bases are both necessary and sufficient for worst-case optimal shadow estimation.

Our explicit construction achieving this scaling rests on two ingredients: phase 3-designs and MUBs. A pure state \(\ket{\phi}\in \mathcal{H}\) is a phase state with respect to the computational basis \(\{\ket{k}\}_k\) if all its entries have the same magnitude. A phase \(t\)-design is a finite ensemble of phase states whose \(t\)-th moment matches that of the uniform random-phase (URP) ensemble \[\label{eq:URPmain} \mathcal{T}= \left\{\frac{1}{\sqrt{d}}\sum_{k=0}^{d-1}\mathrm{e}^{\mathrm{i}\varphi_k}\ket{k}\colon \mathrm{e}^{\mathrm{i}\varphi_k}\sim \mathrm{U}(1)\right\},\tag{4}\] which geometrically forms a \(d\)-dimensional torus. Intuitively, a phase \(t\)-design captures the same statistical correlations as fully random phases up to order \(t\), using only finitely many states. Mathematically, it is equivalent to a projective toric \(t\)-design studied in Ref. [52]. As shown in the End Matter, a phase 2-design can be constructed from \(3p\) bases and a phase 3-design from \(7p^2\) bases, where \(p\) is any prime satisfying \(p \geq \max\{d,5\}\) and can be chosen such that \(p \leq \max\{2d,5\}\). When \(d\) is not divisible by \(3\), the number of bases required for a phase 3-design can be reduced to \(3p^2\). Since any phase \(t\)-design (with \(t\geq 1\)) is also a state 1-design, it gives rise to a valid POVM.

The construction proceeds as follows. Consider \(N\geq 2\) MUBs \(\mathcal{B}_1,\mathcal{B}_2,\ldots,\mathcal{B}_N\). For each basis \(\mathcal{B}_j\), we construct a phase 3-design \(\mathcal{E}_{\mathcal{B}_j}\) with respect to \(\mathcal{B}_j\). The combined ensemble \(\mathcal{E}_N=\bigsqcup_{j=1}^N \mathcal{E}_{\mathcal{B}_j}/N\) assigns uniform probability across all sub-ensembles and comprises at most \(7Np^2\) bases in total. MUBs ensure that the combined ensemble probes complementary degrees of freedom: while a single phase 3-design can only access coherence relative to its reference basis, MUBs guarantee that no observable direction is left unexplored.

Theorem 2. Suppose \(\mathcal{E}_N\) is the ensemble constructed from \(N\) phase 3-design ensembles based on MUBs with \(2\leq N\leq d+1\), and \(O\in\mathcal{L}^{\mathrm{H}}_0(\mathcal{H})\). Then \[\begin{align} \label{eq:MMSoptWCMUB} \|O\|_{\mathcal{E}_N}^2 \leq \frac{3N^2+N}{(N-1)^2}\|O\|_2^2 \leq 14\|O\|_2^2. \end{align}\qquad{(3)}\]

Already for \(N=2\) MUBs, we have \(\|O\|_{\mathcal{E}_2}^2 \leq 14\|O\|_2^2\), achieving constant worst-case shadow norm with \(\Theta(d^2)\) bases—matching the lower bound from Proposition 2 up to a constant factor. The bound improves monotonically with \(N\): for \(N=d+1\) (a complete set of MUBs), the prefactor reads \(3+7/d+4/d^2\), which is quite close to the optimal value of \(3\) achieved by state 3-designs. In practice, \(N=2\) or \(3\) MUBs strike a favorable balance between measurement overhead and shadow norm, as adding MUBs rapidly decreases the prefactor at first but yields diminishing returns thereafter. The ensemble \(\mathcal{E}_N\) is not a state 2-design unless \(N=d+1\); nevertheless, the measurement channel \(\mathcal{M}_{\mathcal{E}_N}\) and reconstruction map \(\mathcal{M}_{\mathcal{E}_N}^{-1}\) both admit simple closed forms (see the End Matter), enabling straightforward construction of unbiased estimators.

Figure 1: Construction of the optimal measurement ensemble \mathcal{E}_N for worst-case shadow estimation from mutually unbiased bases (MUBs) and phase 3-designs. Each phase 3-design is constructed from \Theta(d^2) bases. The ensemble \mathcal{E}_N is composed of \Theta(d^2) bases whenever N=\mathcal{O}(1) but can achieve optimal scaling in the worst-case shadow norm as long as N\geq 2.

Minimal optimal protocols for average performance—We now turn to average performance and show that far fewer measurement bases suffice. We call a protocol average-case optimal if the mean squared shadow norm over a unitary orbit—observables sharing the spectrum of \(O\) but with arbitrary eigenbases—is bounded by \(c\|O\|_2^2\) for a dimension-independent constant \(c\).

For \(O\in\mathcal{L}(\mathcal{H})\) and a unitary ensemble \(\mathcal{U}\) on \(\mathrm{U}(\mathcal{H})\), define the orbit \[\begin{align} \Xi(O,\mathcal{U})\mathrel{\vcenter{:}}= \left\{UOU^\dagger \mid U\sim\mathcal{U}\right\}. \end{align}\] Let \(\|\Xi(O,\mathcal{U})\|_\mathcal{E}^2\) denote the mean squared shadow norm averaged over \(\mathcal{U}\), and \(\|\Xi(O,\mathcal{U})\|_{\mathcal{E},\rho}^2\) its state-dependent counterpart. Note that \(\mathcal{U}\) defines only the observable class over which performance is averaged; it plays no role in the measurement protocol. For a unitary 2-design \(\mathcal{U}\) and a state 2-design \(\mathcal{E}\) [21], [53], we have \[\begin{align} \label{eq:SDopt} \|\Xi(O,\mathcal{U})\|_{\mathcal{E},\rho}^2 = \frac{d+1}{d}\|O\|_2^2, \end{align}\tag{5}\] so any state 2-design is already optimal in the state-dependent setting. However, this is of limited practical interest because \(\rho\) is usually unknown. The state-independent norm \(\|\Xi(O,\mathcal{U})\|_\mathcal{E}^2\), providing a universal guarantee regardless of \(\rho\), can be much larger; the following theorem bounds this gap.

Theorem 3. Suppose \(\mathcal{E}\) is a state \(2\)-design, \(O\in\mathcal{L}^{\mathrm{H}}_0(\mathcal{H})\), and \(\mathcal{U}\) is a unitary \(4\)-design. Then \[\begin{align} \|\Xi(O,\mathcal{U})\|_\mathcal{E}^2 &\le \left(1 + \sqrt{\frac{24}{d}\big[\bar{\Phi}_3(\mathcal{E}){-}1\big] + 4}\,\right)\|O\|_2^2 \nonumber\\ &\le \left(2\sqrt{2}+1\right)\|O\|_2^2. \label{eq:AvgShNormUB} \end{align}\qquad{(4)}\]

The 4-design assumption on \(\mathcal{U}\) (satisfied, e.g., by the Haar measure) ensures sufficient moment control over the orbit. The parameter \(\bar{\Phi}_3\) quantifies deviation from a 3-design [\(\bar{\Phi}_3 = \mathcal{O}(d)\) for any state 2-design by Proposition 1]: at \(\bar{\Phi}_3=1\) the bound recovers the 3-design optimum \(3\|O\|_2^2\), while even the worst 2-design incurs at most a constant-factor overhead.

State 2-designs can be constructed from \(d^2+\mathcal{O}\left(d^{1.525}\right)\) states [54] or \(\mathcal{O}(d)\) orthonormal bases [55], [56]. Explicitly, mixing a fixed basis \(\mathcal{B}\) [with probability \(1/(d+1)\)] with a phase 2-design over \(\mathcal{B}\) [with probability \(d/(d+1)\)] yields a valid state 2-design. Meanwhile, a phase 2-design requires at most \(6d\) bases (Proposition 4 in the End Matter). Conversely, at least \(d^2\) states or \(d+1\) bases are necessary to construct an IC ensemble and guarantee finite shadow norms for all observables in \(\Xi(O,\mathcal{U})\), assuming that the orbit spans \(\mathcal{L}^{\mathrm{H}}_0(\mathcal{H})\).

Theorem 4. In dimension \(d\), \(\Theta(d)\) measurement bases—equivalently \(\Theta(d^2)\) measurement outcomes—are necessary and sufficient for average-case optimal shadow estimation.

The contrast is striking: average-case optimal shadow estimation requires only \(\Theta(d)\) bases, versus \(\Theta(d^2)\) in the worst case—a fundamental quadratic gap. Notably, ensembles as simple as complete sets of MUBs [25][27], cyclic measurements [28], or shallow \(\mathcal{O}(\log n)\)-depth circuits [40] already form state 2-designs and thus achieve average-case optimal performance, enabling practical protocols with unexpectedly economical measurement strategies. Moreover, \(d+1\) bases suffice to realize average-case optimal shadow estimation whenever \(d\) is a prime power, since a complete set of MUBs exists in every prime-power dimension [25][27]. Furthermore, if a SIC exists in every dimension—as conjectured and supported by strong evidence [43][45], [57]—then \(d^2\) outcomes suffice to realize a minimal optimal protocol.

Fidelity estimation—Predicting the fidelity \(\mathop{\mathrm{tr}}(\rho\psi)\) between an unknown state \(\rho\) and a known target state \(\psi\) is among the most important applications of shadow estimation, central to state verification, benchmarking, and certification [4], [5]. Here the observable is the projector \(\psi\), so sample complexity is governed by \(\|\psi_0\|_\mathcal{E}^2\) with \(\psi_0=\psi-\mathbb{1}/d\). Theorem 3 immediately bounds the mean squared shadow norm for Haar-random target states. The following proposition further establishes a large-deviation guarantee, whose proof requires an independent and considerably more involved argument (see SM Sec. 5).

Proposition 3. Suppose \(\mathcal{E}\) is a state \(2\)-design, \(\psi\) is a Haar-random pure state, and \(k>0\). Then \[\begin{gather} \underset{\psi\sim\mathrm{Haar}}{\mathbb{E}}\left\|\psi_0\right\|_\mathcal{E}^2 \leq\left(2\sqrt{2}+1\right), \label{eq:ShNormFidUB} \\ \Pr\left\{\|\psi_0\|_\mathcal{E}^2 \ge\sqrt{48[1+k\xi(\mathcal{E})]}-1\right\} \le\frac{1}{(d+6)^3k^2}, \label{eq:ShNormFidLD} \end{gather}\] {#eq: sublabel=eq:eq:ShNormFidUB,eq:eq:ShNormFidLD} where \[\begin{align} \xi(\mathcal{E})\mathrel{\vcenter{:}}=\sqrt{102[\bar{\Phi}_3(\mathcal{E})-1]+6}<\sqrt{17d}. \label{eq:Xi} \end{align}\qquad{(5)}\]

For SIC ensembles and complete sets of MUBs, the normalized third frame potentials read \[\begin{align} \label{eq:bPhi3SICMUB} \bar{\Phi}_3(\mathcal{E}_\mathrm{SIC})=\frac{(d+2)(d+3)}{6(d+1)},\;\, \bar{\Phi}_3(\mathcal{E}_\mathrm{MUB})=\frac{(d+1)(d+2)}{6d}, \end{align}\tag{6}\] each scaling as \(d/6\) in the large-\(d\) limit and attaining near-maximal values among 2-designs (cf.Proposition 1). Nevertheless, Fig. 2 demonstrates that these canonical 2-designs already closely match the 3-design benchmark for average-case performance.

Figure 2: Mean squared shadow norm for fidelity estimation of Haar-random pure states versus dimension d. SIC and MUB results are compared with upper bounds from Eq. (?? ), the exact 3-design result, and the worst-case 2-design bound [based on Proposition 1 and Eq. (?? )]. Notably, the average performance across all 2-designs, including SICs and MUBs, nearly matches the 3-design benchmark.

Since \(\bar{\Phi}_3(\mathcal{E})=\mathcal{O}(d)\) for any 2-design (Proposition 1), choosing \(k=\mathcal{O}\left(1/\sqrt{d}\mkern 2mu\right)\) in Eq. (?? ) yields \(\|\psi_0\|_\mathcal{E}^2=\mathcal{O}(1)\) except with probability \(\mathcal{O}(1/d^2)\). Hence, the shadow norm concentrates tightly around its mean for Haar-random \(\psi\), with a universally fast rate for all state 2-designs; dependence on \(\bar{\Phi}_3\) only marginally slows convergence for larger \(\bar{\Phi}_3\). States with anomalously large shadow norms thus form a vanishingly small subset of the state space: generic pure-state fidelity estimation achieves constant sample complexity with any 2-design measurement, constructible from only \(\mathcal{O}(d)\) bases. This stands in stark contrast to the \(\Theta(d^2)\) bases required for worst-case optimality.

For stabilizer measurements, by means of involved analysis, a universal upper bound on the shadow norm was established in Refs. [58], [59], scaling with local dimension but not qudit number. However, that bound cannot imply constant sample complexity for generic fidelity estimation. Our concentration result is stronger in this regard: the shadow norm remains \(\mathcal{O}(1)\) with high probability for any state 2-design, independent of both the system dimension and local dimension. This prediction is corroborated by numerical calculation on \(n\)-qudit stabilizer measurements, as illustrated in Fig. 3 in the End Matter.

Beyond Haar-random states, the mean squared shadow norm also admits constant upper bounds for other physically motivated ensembles of target states: URP ensembles, relevant to Hamiltonian simulation, and Clifford orbits, relevant to randomized benchmarking and quantum error correction (Proposition 6 in the End Matter).

Summary and outlook—We have determined the fundamental measurement complexity of shadow estimation. In the worst case, \(\Theta(d^2)\) bases are both necessary and sufficient for optimal performance; our explicit construction from phase 3-designs and MUBs achieves this without requiring exact state 3-design structure. In the average case, only \(\Theta(d)\) bases suffice: any state 2-design yields a constant mean squared shadow norm governed by \(\bar{\Phi}_3\). For fidelity estimation, the shadow norm concentrates tightly for Haar-random targets, implying constant sample complexity for generic instances. Together, these results reveal a sharp quadratic gap between worst-case and average-case complexities: \(d^2\) versus \(d\) bases. For a \(10\)-qubit system, this reduces the required number of bases from \({\sim}10^6\) to \({\sim}10^3\), demonstrating that the practical cost can be far lower than the worst-case bound suggests.

Looking forward, extending these results to multipartite settings and structured observable classes—such as local Hamiltonians, where problem-specific structure may yield further reductions—remains an important open direction. On the practical side, easily implementable 2-designs—realized via complete sets of MUBs [34], [35], cyclic measurements [28], or shallow \(\mathcal{O}(\log n)\)-depth circuits [19], [40][42]—already achieve optimal average-case performance. This is especially relevant for decoherence-limited devices, where measurement overhead is a primary bottleneck. Finally, connections with Hamiltonian-driven shadow protocols [32], [33] offer a promising route: the system’s intrinsic dynamics could naturally generate structured measurements, enabling optimal shadow estimation without external randomization.

Acknowledgments—We thank Changhao Yi, Xiaodi Li, and Xinyang Shu for inspiring discussions. This work is supported by Shanghai Science and Technology Innovation Action Plan (Grant No.24LZ1400200), National Natural Science Foundation of China (Grant No.), Quantum Science and Technology-National Science and Technology Major Project (Grant No.2024ZD0300101), National Key Research and Development Program of China (Grant No.2022YFA1404204), and Shanghai Municipal Science and Technology Major Project (Grant No.2019SHZDZX01).



End Matter


Appendix A: Reconstruction map for the combined phase-design ensemble \(\mathcal{E}_N\)—Consider \(N\geq 2\) MUBs \(\mathcal{B}_1,\mathcal{B}_2,\ldots,\mathcal{B}_N\) and the ensemble \(\mathcal{E}_N\) used in Theorem 2. Any operator \(O\in\mathcal{L}(\mathcal{H})\) can be decomposed as \[\begin{align} O = \frac{\mathop{\mathrm{tr}}(O)\mathbb{1}}{d} + O_\perp + \sum_{j=1}^{N} O_{\mathcal{B}_j}, \end{align}\] where \(O_{\mathcal{B}_j}\) denotes the traceless component diagonal in basis \(\mathcal{B}_j\), and \(O_\perp\) is the component orthogonal to all \(\mathcal{B}_j\)-diagonal operators, which vanishes when \(N=d+1\). With respect to this decomposition, the reconstruction map takes the simple form \[\begin{align} \label{eq:ReconMap} \mathcal{M}_{\mathcal{E}_N}^{-1}(O) = \frac{\mathop{\mathrm{tr}}(O)\mathbb{1}}{d} + dO_\perp + \sum_{j=1}^{N}\frac{Nd}{N-1}O_{\mathcal{B}_j}. \end{align}\tag{7}\] For a complete set of MUBs (\(N=d+1\)), the ensemble \(\mathcal{E}_N\) forms a 2-design, and the reconstruction map reduces to the familiar form \(\mathcal{M}_{\mathcal{E}_{d+1}}^{-1}(O) = (d+1)O - \mathop{\mathrm{tr}}(O)\mathbb{1}\) [1].

Appendix B: Construction of nearly tight phase 2- and 3-designs from orthonormal bases—To construct a phase 2-design (3-design), at least \(\Theta(d^2)\) (\(\Theta(d^3)\)) states are required. Here we provide nearly tight constructions from orthonormal bases. In the case of phase 2-designs, our construction is a reformulation of the construction in Ref. [56], but in a much simpler language. This reformulation is crucial to generalization to phase 3-designs.

Given any positive integer \(t\), define the function \[f_t(x)\mathrel{\vcenter{:}}= \left\lfloor \frac{tx}{d}\right\rfloor,\quad x\in \mathscr{I}:= \{0,1,\ldots,d-1\}.\] Let \(p\) be any prime that satisfies \(p\geq\max\{d,3\}\). Then a phase 2-design can be constructed as follows: \[\label{eq:caT2} \begin{gather} \mathcal{T}_2\mathrel{\vcenter{:}}= \left\{\ket{\psi_2(a,b,c)} \colon a\in \mathbb{Z}_d,\, b\in \mathbb{Z}_3,\, c\in \mathbb{Z}_p\right\},\\ \ket{\psi_2(a,b,c)}\mathrel{\vcenter{:}}= \frac{1}{\sqrt{d}} \sum_{x\in \mathscr{I}} \omega_d^{ax}\omega_3^{bf_2(x)}\omega_p^{cx^2}\ket{x}, \end{gather}\tag{8}\] where \(\omega_k=\mathrm{e}^{2\pi\mathrm{i}/k}\) for a positive integer \(k\) is a primitive \(k\)-th root of unity; note that \(\mathcal{T}_2\) is composed of \(3p\) bases. When \(d\) is odd, we have an alternative construction using only \(2p\) bases: \[\label{eq:tcaT2} \begin{gather} \tilde{\mathcal{T}}_2\mathrel{\vcenter{:}}=\left\{\ket{\tilde{\psi}_2(a,b,c)} \colon a\in \mathbb{Z}_d,\, b\in \mathbb{Z}_2,\, c\in \mathbb{Z}_p\right\},\\ \ket{\tilde{\psi}_2(a,b,c)}\mathrel{\vcenter{:}}=\frac{1}{\sqrt{d}} \sum_{x\in \mathscr{I}} \omega_d^{ax}\omega_2^{bx}\omega_p^{cx^2}\ket{x}. \end{gather}\tag{9}\]

Proposition 4. The ensembles \(\mathcal{T}_2\) and \(\tilde{\mathcal{T}}_2\) defined in Eqs. (8 ) and (9 ) form phase 2-designs.

Next, we turn to phase 3-designs. Let \(p\) be any prime that satisfies \(p\geq\max\{d,5\}\). Then a phase 3-design can be constructed as follows: \[\label{eq:caT3} \begin{gather} \mathcal{T}_3\mathrel{\vcenter{:}}=\left\{\ket{\psi_3(a,b,c)} \colon a\in \mathbb{Z}_d,\, b\in \mathbb{Z}_7,\, c_2, c_3\in \mathbb{Z}_p\right\},\\ \ket{\psi_3(a,b,c)}\mathrel{\vcenter{:}}=\frac{1}{\sqrt{d}} \sum_{x\in \mathscr{I}} \omega_d^{ax}\omega_7^{bf_3(x)}\omega_p^{c_2x^2+c_3x^3}\ket{x}; \end{gather}\tag{10}\] note that \(\mathcal{T}_3\) is composed of \(7p^2\) bases. When \(d\) is not divisible by \(3\), we have an alternative construction using only \(3p^2\) bases: \[\label{eq:tcaT3} \begin{gather} \tilde{\mathcal{T}}_3\mathrel{\vcenter{:}}=\left\{\ket{\tilde{\psi}_3(a,b,c)} \colon a\in \mathbb{Z}_d,\, b\in \mathbb{Z}_3,\, c_2, c_3\in \mathbb{Z}_p\right\},\\ \ket{\tilde{\psi}_3(a,b,c)}\mathrel{\vcenter{:}}=\frac{1}{\sqrt{d}} \sum_{x\in \mathscr{I}} \omega_d^{ax}\omega_3^{bx}\omega_p^{c_2x^2+c_3x^3}\ket{x}. \end{gather}\tag{11}\]

Proposition 5. The ensembles \(\mathcal{T}_3\) and \(\tilde{\mathcal{T}}_3\) defined in Eqs. (10 ) and (11 ) form phase 3-designs.

Proposition 5 is a simple corollary of Lemma 2 in SM Sec. 1; Proposition 4 follows from similar but simpler reasoning. It is well known that, for any positive integer \(d\), there exists a prime \(p\) with \(p\leq 2d\) [60]. Therefore, for any dimension \(d\), a phase 2-design can be constructed from no more than \(6d\) bases, while a phase 3-design can be constructed from no more than \(28d^2\) bases. In the large-\(d\) limit, it is possible to tighten these bounds.

Appendix C: Numerical evidence for concentration in stabilizer shadow estimation—Figure 3 plots the probability upper bound from Proposition 3 for the event \(\|\psi_0\|_{\mathcal{E}}^2 \geq 9\). Here, \(\psi\) denotes a Haar-random pure state of \(n\) qudits with local dimension \(p\), and fidelity estimation is implemented via stabilizer measurements. These measurements form a Clifford orbit which, for odd prime \(p\), is a state 2-design but not a 3-design [23], [24], [48]. The worst-case shadow norm satisfies \(\|\psi_0\|_\mathcal{E}^2\leq 2p-1\) [58], with the bound approximately saturated by stabilizer states when \(n\geq 2\). This worst-case scaling with \(p\) might suggest suboptimality for fidelity estimation. However, Fig. 3 demonstrates that the large-deviation probability vanishes rapidly with increasing \(n\) for various choices of \(p\). The normalized third frame potential \(\bar{\Phi}_3(\mathcal{E})=(p+1)(d+2)/[3(d+p)]\) [59] approaches \((p+1)/3\) for large \(n\); together with Proposition 3, this yields an \(\mathcal{O}\left(1/d^2\right)\) tail bound that explains the observed concentration and confirms constant sample complexity for typical fidelity estimation. Related numerical observations were reported in Ref. [58]; our Theorem 3 and Proposition 3 provide the rigorous theoretical underpinning that has been missing in the prior work.

Figure 3: Probability that \|\psi_0\|_\mathcal{E}^2\geq 9 for a Haar-random pure state \psi in fidelity estimation based on n-qudit stabilizer measurements with local dimension p. The rapid decay confirms constant sample complexity for generic target states, in agreement with Theorem 3 and Proposition 3.

Appendix D: Fidelity estimation beyond Haar-random pure states—The main text establishes concentration for Haar-random target states. Here we show that constant mean squared shadow norms also hold for other physically motivated ensembles.

Proposition 6. Suppose \(\mathcal{E}\) is a state 2-design. Then: (i)* For the uniform random-phase (URP) ensemble with respect to any orthonormal basis, \[\begin{align} \underset{\psi\sim\mathrm{URP}}{\mathbb{E}}\left\|\psi_0\right\|_\mathcal{E}^2 \le\sqrt{\frac{48(d+1)(d+2)(d+3)}{d^3}}-1. \label{eq:AShNormRP} \end{align}\tag{12}\] (ii) For \(d=2^n\) and the \(n\)-qubit Clifford group \(\mathrm{Cl}(n)\), \[\begin{align} \underset{U\sim\mathrm{Cl}(n)}{\mathbb{E}} \left\|U\psi_0 U^\dagger\right\|_\mathcal{E}^2 \le 2\sqrt{15}-1 \quad \forall\,\psi\in\mathcal{P}(\mathcal{H}). \label{eq:AShNormFidUBCli} \end{align}\tag{13}\] *

URP ensembles—The URP ensemble models long-time outputs of Hamiltonian evolution with random energy eigenvalues under a no-resonance condition [61], a standard setting in quantum thermalization. The bound Eq. (12 ) holds for the URP ensemble with respect to any orthonormal basis and converges to \(\sqrt{48}-1\approx 5.93\) as \(d\to\infty\), guaranteeing efficient fidelity estimation with any 2-design measurement regardless of system size.

Clifford orbits—The bound Eq. (13 ) is uniform over all initial states \(\psi\) and equals \(2\sqrt{15}-1\approx 6.75\). It applies to any Clifford orbit, including orbits of stabilizer states, central to quantum error correction, and magic states, relevant to magic state distillation. Since Clifford orbits arise naturally in verification and benchmarking, efficient fidelity estimation for these targets is particularly valuable in practice.

Comparison—Both bounds are slightly larger than the Haar-random mean of \(2\sqrt{2}+1\approx 3.83\), reflecting less randomness in these ensembles. Nevertheless, all three bounds are \(\mathcal{O}(1)\), confirming that constant sample complexity for fidelity estimation is a robust structural feature of 2-design measurements, not an artifact of Haar randomness.

Optimal Shadow Estimation with Minimal Measurement Settings: Supplemental Material

In this Supplemental Material (SM), we prove the results presented in the main text and End Matter, including Propositions 16 as well as Theorems 2 and 3. In addition, we provide some auxiliary results on frame potentials, the combined phase-design ensemble based on MUBs (optimal for worst-case shadow estimation), and shadow norms.

As in the main text, let \(\mathcal{H}\) be a complex Hilbert space of dimension \(d\ge2\). Denote by \(\mathcal{L}(\mathcal{H})\), \(\mathcal{L}^{\mathrm{H}}(\mathcal{H})\), \(\mathcal{L}_0(\mathcal{H})\), and \(\mathcal{L}^{\mathrm{H}}_0(\mathcal{H})\) the spaces of linear operators, Hermitian operators, traceless operators, and traceless Hermitian operators on \(\mathcal{H}\), respectively. Denote by \(\mathcal{D}(\mathcal{H})\) the set of density operators on \(\mathcal{H}\) and by \(\mathcal{P}(\mathcal{H})\) the subset of pure state projectors. For any operator \(O \in \mathcal{L}(\mathcal{H})\), we write \(O_0=O-\mathop{\mathrm{tr}}(O)\mathbb{1}/d\) for the traceless part of \(O\); for any pure state \(\ket{\phi} \in \mathcal{H}\), we write \(\phi = \ket{\phi}\mkern-2mu\bra{\phi}\) for the corresponding projector and \(\phi_0=\phi-\mathbb{1}/d\). Given a positive integer \(t\), denote by \(P_{[t]}\) the projector onto the \(t\)-partite symmetric subspace of \(\mathcal{H}^{\otimes t}\) and by \(D_{[t]}=\mathop{\mathrm{tr}}(P_{[t]})\) the dimension of the symmetric subspace.

We also introduce the vectorization map between operators on \(\mathcal{H}\) and vectors in \(\mathcal{H}\otimes \mathcal{H}\). For any \(O \in \mathcal{L}(\mathcal{H})\), its vectorization is defined in the computational basis \(\{\ket{k}\}_{k=0}^{d-1}\) as \[\begin{align} \dket{O} \mathrel{\vcenter{:}}= \sum_{k,l} O_{kl}\ket{k}\otimes \ket{l}, \end{align}\] where \(O_{kl}=\bra{k}O\ket{l}\).

1 State \(t\)-designs and phase \(t\)-designs↩︎

Here we discuss state \(t\)-designs and phase \(t\)-designs in more detail and derive auxiliary results relevant to this work.

Let \(\mathcal{E}= \{\ket{\phi_i}, w_i\}_{i}\) be an ensemble of pure states in \(\mathcal{H}\), where \(w_i>0\) and \(\sum_i w_i=1\). The \(t\)-th moment operator and normalized moment operator of \(\mathcal{E}\) are defined as \[Q_t(\mathcal{E})\mathrel{\vcenter{:}}=\sum_i w_i\phi_i^{\otimes t}, \quad \bar{Q}_t(\mathcal{E})\mathrel{\vcenter{:}}= D_{[t]}Q_t(\mathcal{E}).\] The \(t\)-th frame potential and normalized \(t\)-th frame potential of \(\mathcal{E}\) are defined as [see Eq. (3 )] \[\label{eq:FPtSupp} \Phi_t(\mathcal{E}) \mathrel{\vcenter{:}}= \sum_{i,j} w_i w_j \bigl[\mathop{\mathrm{tr}}(\phi_i\phi_j)\bigr]^t,\quad \bar{\Phi}_t(\mathcal{E}) \mathrel{\vcenter{:}}= D_{[t]}\,\Phi_t(\mathcal{E}),\tag{14}\] which can also be expressed as \[\Phi_t(\mathcal{E})=\mathop{\mathrm{tr}}\left\{[Q_t(\mathcal{E})]^2\right\},\quad \bar{\Phi}_t(\mathcal{E})=D_{[t]}\mathop{\mathrm{tr}}\left\{[Q_t(\mathcal{E})]^2\right\}=\frac{\mathop{\mathrm{tr}}\left\{[\bar{Q}_t(\mathcal{E})]^2\right\}}{D_{[t]}}.\] By definition, we have \[\label{eq:Phittp1} \Phi_{t+1}(\mathcal{E}) \leq \Phi_t(\mathcal{E}),\quad \bar{\Phi}_{t+1}(\mathcal{E})\leq \frac{d+t}{t+1} \bar{\Phi}_t(\mathcal{E}),\tag{15}\] given that \(D_{[t+1]}=(d+t)D_{[t]}/(t+1)\).

1.1 Haar-random states and state \(t\)-designs↩︎

State \(t\)-designs, also known as complex projective \(t\)-designs, are configurations of pure states whose \(t\)-th moments match those of Haar-random states [44][47]. They can be viewed as the complex analog of spherical designs on the real unit sphere and have found applications in areas such as approximation theory, combinatorics, and quantum information. The ensemble \(\mathcal{E}\) is a state \(t\)-design if \(\bar{Q}_t(\mathcal{E})=P_{[t]}\). It is known that \(\bar{\Phi}_t(\mathcal{E})\geq 1\), and the lower bound is saturated if and only if \(\mathcal{E}\) is a \(t\)-design. In addition, any state \(t\)-design in dimension \(d\) has at least \[\begin{align} \binom{d+\lceil t/2\rceil-1}{\lceil t/2\rceil} \binom{d+\lfloor t/2\rfloor-1}{\lfloor t/2\rfloor} \end{align}\] elements [46], [50]; the lower bound simplifies to \(d^2\) and \(d^2(d+1)/2\) for \(t=2\) and \(t=3\), respectively. When \(t=2\), the lower bound is saturated if and only if the ensemble \(\mathcal{E}\) induces a SIC-POVM. It is conjectured and supported by strong evidence that a SIC exists in every finite dimension [43][45], [57]. In sharp contrast, it remains open whether a 3-design can always be constructed using only \(\mathcal{O}(d^3)\) states. As a prominent example, the set of \(n\)-qubit stabilizer states forms a 3-design [23], [24], [48], yet the number of such states grows superpolynomially with dimension \(d\).

1.2 Uniform random phase states and phase \(t\)-designs↩︎

Given a positive integer \(t\), two sequences \(\boldsymbol{k},\boldsymbol{l}\in \{0,1,\dots,d-1\}^t\) are said to be equivalent if one can be obtained from the other by a permutation of entries. The equivalence class of \(\boldsymbol{k}\) is denoted by \([\boldsymbol{k}]\), and its cardinality by \(|[\boldsymbol{k}]|\). Let \(\mathcal{B}= \{\ket{k}_\mathcal{B}\}_{k=0}^{d-1}\) be an orthonormal basis of \(\mathcal{H}\). For each equivalence class \([\boldsymbol{k}]\), we define the corresponding symmetric ket in \(\mathcal{H}^{\otimes t}\) by \[\begin{align} \ket{[\boldsymbol{k}]}_\mathcal{B}\mathrel{\vcenter{:}}= \frac{1}{\sqrt{|[\boldsymbol{k}]|}} \sum_{\boldsymbol{l}\in [\boldsymbol{k}]} \ket{l_1}_\mathcal{B}\otimes \cdots \otimes \ket{l_t}_\mathcal{B}. \end{align}\] Recall that the uniform random phase (URP) ensemble \(\mathcal{T}_\mathcal{B}\) with respect to \(\mathcal{B}=\{\ket{k}_\mathcal{B}\}\) has the form \[\mathcal{T}_\mathcal{B}=\left\{\frac{1}{\sqrt{d}} \sum_{k=0}^{d-1} \mathrm{e}^{\mathrm{i}\varphi_k} \ket{k}_\mathcal{B}\colon \mathrm{e}^{\mathrm{i}\varphi_k}\sim \mathrm{U}(1)\right\}, \label{eq:URPsupp}\tag{16}\] where \(\varphi = (\varphi_0, \varphi_1, \dots, \varphi_{d-1})\) is uniformly distributed over \([0,2\pi)^d\). When \(\mathcal{B}\) is the computational basis, we omit the subscript \(\mathcal{B}\) and simply write \(\mathcal{T}\).

Lemma 1. For the URP ensemble \(\mathcal{T}_\mathcal{B}\), we have \[\begin{align} Q_t(\mathcal{T}_\mathcal{B}) = \frac{1}{d^t} \sum_{[\boldsymbol{k}]} |[\boldsymbol{k}]| \ket{[\boldsymbol{k}]}\mkern-2mu\bra{[\boldsymbol{k}]}_\mathcal{B}\le \frac{t!}{d^t}P_{[t]},\quad t=1,2,\ldots \label{eq:URP95AVG} \end{align}\qquad{(6)}\]

We now define phase designs. Let \(\mathcal{E}_\mathcal{B}\) be an ensemble of phase states (equal-modulus states) of the form \[\begin{align} \label{eq:PhaseState} \ket{\phi} = \frac{1}{\sqrt d} \sum_{k=0}^{d-1}\mathrm{e}^{\mathrm{i}\varphi_k}\ket{k}_\mathcal{B}. \end{align}\tag{17}\] The ensemble \(\mathcal{E}_\mathcal{B}\) is called a phase \(t\)-design with respect to \(\mathcal{B}\) if \[\begin{align} Q_t(\mathcal{E}_\mathcal{B}) = Q_t(\mathcal{T}_\mathcal{B}). \end{align}\] By definition, any phase \(t\)-design with \(t\geq 1\) is also a state 1-design and thus defines a POVM. In addition, a phase \(t\)-design determines a projective toric \(t\)-design [52], and vice versa.

Proof of Lemma 1. By definition, a state drawn from the URP ensemble has the form in Eq. (17 ) with \(\mathrm{e}^{\mathrm{i}\varphi_k}\sim \mathrm{U}(1)\). The \(t\)-th tensor power of the corresponding projector reads \[\begin{align} \phi^{\otimes t} = \left(\ket{\phi}\mkern-2mu\bra{\phi}\right)^{\otimes t} = \frac{1}{d^t} \sum_{k_1,\dots,k_t} \sum_{l_1,\dots,l_t} \mathrm{e}^{\mathrm{i}(\varphi_{k_1}+\dots+\varphi_{k_t}-\varphi_{l_1}-\dots-\varphi_{l_t})} \ket{k_1\cdots k_t}\bra{l_1\cdots l_t}_\mathcal{B}. \end{align}\] The expectation of the phase factor over \(\mathcal{T}_\mathcal{B}\) reads \[\begin{align} \underset{\phi\sim\mathcal{T}_\mathcal{B}}{\mathbb{E}} \left[\mathrm{e}^{\mathrm{i}(\sum_{s=1}^t \varphi_{k_s} - \sum_{s=1}^t \varphi_{l_s})}\right] = \begin{cases} 1 & \text{if } [\boldsymbol{k}]=[\boldsymbol{l}], \\ 0 & \text{otherwise}. \end{cases} \end{align}\] The above two equations together imply the equality in Eq. (?? ). The inequality holds because \(\left\{\ket{[\boldsymbol{k}]}_\mathcal{B}\right\}_{[\boldsymbol{k}]}\) forms an orthonormal basis of the symmetric subspace of \(\mathcal{H}^{\otimes t}\) and \(|[\boldsymbol{k}]|\leq t!\). ◻

1.3 Proof of Proposition 1↩︎

Proof of Proposition 1. Suppose the state ensemble \(\mathcal{E}\) has the form \(\mathcal{E}=\{\ket{\phi_i},w_i\}_i\) and let \(\phi_{i,0}=\phi_i-\mathbb{1}/{d}\) for each index \(i\). Since \(0\leq \mathop{\mathrm{tr}}(\phi_i\phi_j)\leq 1\) for all \(i,j\), we have \[\begin{align} \sum_{i,j}w_iw_j\mathop{\mathrm{tr}}(\phi_i\phi_j)\left[\mathop{\mathrm{tr}}\left(\phi_{i,0}\phi_{j,0}\right)\right]^2\le \sum_{i,j}w_iw_j\left[\mathop{\mathrm{tr}}\left(\phi_{i,0}\phi_{j,0}\right)\right]^2. \label{eq:phi95395eq1} \end{align}\tag{18}\] Direct calculation yields \[\begin{align} \sum_{i,j}w_iw_j\mathop{\mathrm{tr}}(\phi_i\phi_j)\left[\mathop{\mathrm{tr}}\left(\phi_{i,0}\phi_{j,0}\right)\right]^2&=\Phi_3(\mathcal{E})-\frac{2}{d}\Phi_2(\mathcal{E})+\frac{1}{d^2}\Phi_1(\mathcal{E})=\Phi_3(\mathcal{E})-\frac{3d-1}{d^3(d+1)},\\ \sum_{i,j}w_iw_j\left[\mathop{\mathrm{tr}}\left(\phi_{i,0}\phi_{j,0}\right)\right]^2&=\Phi_2(\mathcal{E})-\frac{2}{d}\Phi_1(\mathcal{E})+\frac{1}{d^2}=\frac{d-1}{d^2(d+1)}, \end{align}\] where the second equality in each line holds because \(\Phi_1(\mathcal{E})=1/d\) and \(\Phi_2(\mathcal{E})=2/[d(d+1)]\), given that \(\mathcal{E}\) is a 2-design. Combining the above three equations yields \[\Phi_3(\mathcal{E})\leq \frac{d^2+2d-1}{d^3(d+1)},\quad \bar{\Phi}_3(\mathcal{E}) \leq \frac{(d+2)(d^2+2d-1)}{6d^2}.\] which confirms Proposition 1. ◻

1.4 Proofs of Propositions 4 and 5↩︎

Proposition 5 is a simple corollary of the following lemma, which guarantees that the third moment operators of \(\mathcal{T}_3\) and \(\tilde{\mathcal{T}}_3\) have the form in Lemma 1. Proposition 4 follows from similar but simpler reasoning.

Lemma 2. Suppose \(x_1, x_2, x_3, y_1, y_2, y_3\in\{0,1,2,\ldots, d-1\}\). Then the sequence \(y_1, y_2, y_3\) is a permutation of \(x_1, x_2, x_3\) if and only if \[\label{eq:ModularCondition} \begin{align} x_1+x_2+x_3&=y_1+y_2+y_3\!\mod d,\\ f_3(x_1)+f_3(x_2)+f_3(x_3)&=f_3(y_1)+f_3(y_2)+f_3(y_3)\!\mod 7,\\ x_1^2+x_2^2+x_3^2&=y_1^2+y_2^2+y_3^2\!\mod p,\\ x_1^3+x_2^3+x_3^3&=y_1^3+y_2^3+y_3^3\!\mod p. \end{align}\qquad{(7)}\] If \(d\) is not divisible by 3, then the sequence \(y_1, y_2, y_3\) is a permutation of \(x_1, x_2, x_3\) if and only if \[\label{eq:ModularCondition2} \begin{align} x_1+x_2+x_3&=y_1+y_2+y_3\!\mod d,\\ x_1+x_2+x_3&=y_1+y_2+y_3\!\mod 3,\\ x_1^2+x_2^2+x_3^2&=y_1^2+y_2^2+y_3^2\!\mod p,\\ x_1^3+x_2^3+x_3^3&=y_1^3+y_2^3+y_3^3\!\mod p. \end{align}\qquad{(8)}\]

Proof of Lemma 2. If the sequence \(y_1, y_2, y_3\) is a permutation of \(x_1, x_2, x_3\), then Eq. (?? ) holds automatically. Conversely, suppose Eq. (?? ) holds. Then the second line implies that \[f_3(x_1)+f_3(x_2)+f_3(x_3)=f_3(y_1)+f_3(y_2)+f_3(y_3),\quad |y_1+y_2+y_3-x_1-x_2-x_3|<d,\] given that \(0\leq f_3(x)\leq 2\) for \(x\in \{0,1,2,\ldots, d-1\}\). The first two lines in Eq. (?? ) together yield \(x_1+x_2+x_3=y_1+y_2+y_3\) and \[\label{eq:ElementaryPoly1} x_1+x_2+x_3=y_1+y_2+y_3\!\mod p.\tag{19}\] In conjunction with the last two lines in Eq. (?? ), we can deduce that \[\label{eq:ElementaryPoly23} \begin{align} x_1 x_2 +x_2 x_3+x_3x_1 &= y_1 y_2 +y_2 y_3+y_3y_1 \!\mod p, \\ x_1 x_2 x_3 &= y_1 y_2 y_3 \!\mod p. \end{align}\tag{20}\] Equations (19 ) and (20 ) together imply that the sequence \(y_1, y_2, y_3\) is a permutation of \(x_1, x_2, x_3\).

Next, suppose \(d\) is not divisible by 3; then \(d\) and 3 are coprime. If the sequence \(y_1, y_2, y_3\) is a permutation of \(x_1, x_2, x_3\), then Eq. (?? ) holds automatically. Conversely, suppose Eq. (?? ) holds. The first two lines together give \(x_1+x_2+x_3=y_1+y_2+y_3 \!\mod 3d\), which in turn implies that \(x_1+x_2+x_3=y_1+y_2+y_3\). So Eqs. (19 ) and (20 ) hold as before, and the sequence \(y_1, y_2, y_3\) is a permutation of \(x_1, x_2, x_3\). ◻

1.5 Auxiliary results on frame potentials↩︎

Here we introduce auxiliary results on (normalized) frame potentials and their variants, which are useful for studying shadow norms in shadow estimation.

Suppose \(\mathcal{E}=\{\ket{\phi_{i}}, w_i\}_{i}\) is a pure state ensemble on \(\mathcal{H}\) and \(\psi \in \mathcal{P}(\mathcal{H})\). We define the \(t\)-th frame potential of \(\mathcal{E}\) relative to \(\psi\) and its normalized version as \[\begin{align} \Phi_t(\mathcal{E},\psi)\mathrel{\vcenter{:}}=\sum_iw_i\left[\mathop{\mathrm{tr}}(\psi \phi_i)\right]^t,\quad \bar{\Phi}_t(\mathcal{E},\psi)\mathrel{\vcenter{:}}= D_{[t]}\Phi_t(\mathcal{E},\psi). \end{align}\] where \(D_{[t]}=\binom{d + t - 1}{t}\) is the dimension of the symmetric subspace of \(\mathcal{H}^{\otimes t}\). These definitions are also applicable when \(t=0\), in which case \(\bar{\Phi}_t(\mathcal{E},\psi)=\Phi_t(\mathcal{E},\psi)=\bar{\Phi}_t(\mathcal{E})=\Phi_t(\mathcal{E})=D_{[t]}=1\). By definition, \(\Phi_t(\mathcal{E})=\sum_j w_j\Phi_t(\mathcal{E},\phi_j)\), which implies \[\min_{\psi\in \mathcal{P}(\mathcal{H})} \Phi_t(\mathcal{E},\psi) \le \Phi_t(\mathcal{E}) \le \max_{\psi\in \mathcal{P}(\mathcal{H})} \Phi_t(\mathcal{E},\psi).\]

Lemma 3. Suppose \(\mathcal{E}\) is a state ensemble on \(\mathcal{H}\), \(\psi\in\mathcal{P}(\mathcal{H})\), and \(t\) is a positive integer. Then \[\begin{align} \Phi_t^2(\mathcal{E})&\leq \Phi_{t-1}(\mathcal{E}) \Phi_{t+1}(\mathcal{E}), & \Phi_t^2(\mathcal{E},\psi)&\leq \Phi_{t-1}(\mathcal{E},\psi) \Phi_{t+1}(\mathcal{E},\psi), \label{eq:FPrelation}\\ \frac{\bar{\Phi}_t^2(\mathcal{E})}{D_{[t]}^2}&\leq \frac{\bar{\Phi}_{t-1}(\mathcal{E}) \bar{\Phi}_{t+1}(\mathcal{E})}{D_{[t-1]}D_{[t+1]}},& \frac{\bar{\Phi}_t^2(\mathcal{E},\psi)}{D_{[t]}^2}&\leq \frac{\bar{\Phi}_{t-1}(\mathcal{E},\psi) \bar{\Phi}_{t+1}(\mathcal{E},\psi)}{D_{[t-1]}D_{[t+1]}}. \label{eq:NFPrelation} \end{align}\] {#eq: sublabel=eq:eq:FPrelation,eq:eq:NFPrelation} If \(\mathcal{E}\) forms a state 2-design, then \[\begin{align} \bar{\Phi}_3(\mathcal{E},\psi)\ge\dfrac{2(d+2)}{3(d+1)}> \dfrac{2}{3}, \quad \bar{\Phi}_4(\mathcal{E},\psi)\ge\dfrac{(d+2)(d+3)}{3(d+1)^2}> \dfrac{1}{3}, \label{eq:NFP34LB} \end{align}\qquad{(9)}\] and the same bounds hold with \(\bar{\Phi}_t(\mathcal{E},\psi)\) replaced by \(\bar{\Phi}_t(\mathcal{E})\) for \(t=3,4\).

Lemma 4. Suppose \(\psi,\phi\in\mathcal{P}(\mathcal{H})\). Then \[\begin{align} \mathop{\mathrm{tr}}\left[P_{[8]}\left(\psi^{\otimes 4}\otimes \phi^{\otimes 4}\right)\right]=\frac{1+16\mathop{\mathrm{tr}}(\psi\phi)+36\left[\mathop{\mathrm{tr}}(\psi\phi)\right]^2+16\left[\mathop{\mathrm{tr}}(\psi\phi)\right]^3+\left[\mathop{\mathrm{tr}}(\psi\phi)\right]^4}{70}. \end{align}\]

Lemma 5. Suppose \(\mathcal{E}\) is a state ensemble on \(\mathcal{H}\) and \(\psi \in \mathcal{P}(\mathcal{H})\) is a Haar-random pure state. Then \[\begin{align} \underset{\psi\sim \mathrm{Haar}}{\mathbb{E}} \bar{\Phi}_4(\mathcal{E},\psi)&=1, \label{eq:4th95moment}\\ \underset{\psi\sim \mathrm{Haar}}{\mathbb{E}} \left[\bar{\Phi}_4(\mathcal{E},\psi)\right]^2&=\frac{D_1}{D_2}\left(1+\dfrac{16\bar{\Phi}_1(\mathcal{E})}{D_{[1]}}+\dfrac{36\bar{\Phi}_2(\mathcal{E})}{D_{[2]}}+\dfrac{16\bar{\Phi}_3(\mathcal{E})}{D_{[3]}}+\dfrac{\bar{\Phi}_4(\mathcal{E})}{D_{[4]}}\right), \label{eq:4th95squared} \end{align}\] {#eq: sublabel=eq:eq:4th95moment,eq:eq:4th95squared} where \(D_1=d(d+1)(d+2)(d+3)\) and \(D_2=(d+4)(d+5)(d+6)(d+7)\). If \(\mathcal{E}\) forms a state 2-design, then \[\begin{align} \mathrm{Var}(\bar{\Phi}_4(\mathcal{E},\psi))&\le\frac{6(17d+51)[\bar{\Phi}_3(\mathcal{E})-1]+6(d-1)}{D_2}<\frac{[\xi(\mathcal{E})]^2}{(d+6)^3}, \label{eq:4th95mo95var} \end{align}\qquad{(10)}\] where \(\xi(\mathcal{E})=\sqrt{102[\bar{\Phi}_3(\mathcal{E})-1]+6}\), as defined in Eq. (?? ).

Lemma 6. Suppose \(\mathcal{E}\) is a state ensemble on \(\mathcal{H}\) and \(\mathcal{B}\) is an orthonormal basis of \(\mathcal{H}\). Then \[\begin{align} \underset{\psi\sim \mathcal{T}_\mathcal{B}}{\mathbb{E}}\bar{\Phi}_t(\mathcal{E}, \psi)\le\frac{t!D_{[t]}}{d^t},\label{eq:RP95URP} \end{align}\qquad{(11)}\] where \(\mathcal{T}_\mathcal{B}\) is the uniform random phase ensemble with respect to \(\mathcal{B}\), as defined in Eq. (16 ).

1.6 Proofs of auxiliary results Lemmas 3 and 6↩︎

Proof of Lemma 3. Equation (?? ) follows from the Cauchy–Schwarz inequality applied to the definitions of frame potentials and relative frame potentials. Equation (?? ) is a simple corollary of Eq. (?? ) given that \(\bar{\Phi}_t(\mathcal{E})=D_{[t]}\Phi_t(\mathcal{E})\) and \(\bar{\Phi}_t(\mathcal{E},\psi)=D_{[t]}\Phi_t(\mathcal{E},\psi)\) for any positive integer \(t\).

Next, suppose \(\mathcal{E}\) is a 2-design; then \(\bar{\Phi}_1(\mathcal{E},\psi)=\bar{\Phi}_1(\mathcal{E})=1\) and \(\bar{\Phi}_2(\mathcal{E},\psi)=\bar{\Phi}_2(\mathcal{E})=1\). By virtue of Eq. (?? ) with \(t=2,3\), we deduce that \[\begin{align} \bar{\Phi}_3(\mathcal{E},\psi)&\ge \frac{D_{[1]}D_{[3]}\bar{\Phi}_2^2(\mathcal{E},\psi)}{D_{[2]}^2 \bar{\Phi}_1(\mathcal{E},\psi)}=\frac{D_{[1]}D_{[3]}}{D_{[2]}^2}=\dfrac{2(d+2)}{3(d+1)}> \dfrac{2}{3},\\ \bar{\Phi}_4(\mathcal{E},\psi)&\ge \frac{D_{[2]}D_{[4]}\bar{\Phi}_3^2(\mathcal{E},\psi)}{D_{[3]}^2 \bar{\Phi}_2(\mathcal{E},\psi)} =\frac{3(d+3)\bar{\Phi}_3^2(\mathcal{E},\psi)}{4(d+2)} \ge\dfrac{(d+2)(d+3)}{3(d+1)^2}> \dfrac{1}{3}, \end{align}\] which confirm Eq. (?? ). The same reasoning applies with \(\bar{\Phi}_t(\mathcal{E},\psi)\) replaced by \(\bar{\Phi}_t(\mathcal{E})\) for \(t=1,2,3,4\). This completes the proof of Lemma 3. ◻

Proof of Lemma 4. Note that \(P_{[8]}=\sum_{\sigma \in S_8} R(\sigma)/8!\), where \(S_t\) for a positive integer \(t\) denotes the symmetric group of \(t\) numbers, and \(R(\sigma)\) denotes the unitary representation of the permutation \(\sigma\). Accordingly, \[\begin{align} \mathop{\mathrm{tr}}\left[P_{[8]}\left(\psi^{\otimes 4}\otimes \phi^{\otimes 4}\right)\right] = \frac{1}{8!} \sum_{\sigma \in S_8} \mathop{\mathrm{tr}}\left[R(\sigma)\left(\psi^{\otimes 4}\otimes \phi^{\otimes 4} \right)\right]= \frac{1}{8!} \sum_{\sigma \in S_8} [\mathop{\mathrm{tr}}(\psi\phi)]^{\gamma(\sigma)}, \end{align}\] where \(\gamma(\sigma)=|\sigma(\{1,2,3,4\})\cap\{5,6,7,8\}|\). A straightforward counting argument shows that \[\begin{align} |\{\sigma \in S_8 \mid \gamma(\sigma)=m \}|=\begin{cases} 576 & m=0,\, 4,\\ 9216 & m=1,\, 3,\\ 20736 & m=2. \end{cases} \end{align}\] Combining the above two equations, we obtain \[\begin{align} \mathop{\mathrm{tr}}\left[P_{[8]}\left(\psi^{\otimes 4}\otimes \phi^{\otimes 4}\right)\right]&= \frac{576+9216\mathop{\mathrm{tr}}(\psi\phi)+20736\left[\mathop{\mathrm{tr}}(\psi\phi)\right]^2+9216\left[\mathop{\mathrm{tr}}(\psi\phi)\right]^3+576\left[\mathop{\mathrm{tr}}(\psi\phi)\right]^4}{8!}\nonumber\\ &=\frac{1+16\mathop{\mathrm{tr}}(\psi\phi)+36\left[\mathop{\mathrm{tr}}(\psi\phi)\right]^2+16\left[\mathop{\mathrm{tr}}(\psi\phi)\right]^3+\left[\mathop{\mathrm{tr}}(\psi\phi)\right]^4}{70}, \label{eq:8th95moment} \end{align}\tag{21}\] which completes the proof of Lemma 4. ◻

Proof of Lemma 5. For any positive integer \(t\), we have \[\begin{align} \underset{\psi\sim \mathrm{Haar}}{\mathbb{E}}\psi^{\otimes t}=\dfrac{P_{[t]}}{D_{[t]}}, \end{align}\] where \(P_{[t]}\) denotes the projector onto the \(t\)-partite symmetric subspace of \(\mathcal{H}^{\otimes t}\) and \(D_{[t]}=\mathop{\mathrm{tr}}(P_{[t]})=\binom{d + t - 1}{t}\). Based on this observation, Eq. (?? ) can be proved as follows: \[\begin{align} \underset{\psi\sim \mathrm{Haar}}{\mathbb{E}} \bar{\Phi}_4(\mathcal{E},\psi)=D_{[4]}\underset{\psi\sim \mathrm{Haar}}{\mathbb{E}}\sum_iw_i\left[\mathop{\mathrm{tr}}(\psi\phi_i)\right]^4=D_{[4]}\sum_iw_i\mathop{\mathrm{tr}}\left[\left(\underset{\psi\sim \mathrm{Haar}}{\mathbb{E}} \psi^{\otimes 4}\right)\phi_i^{\otimes 4}\right]=\sum_iw_i\mathop{\mathrm{tr}}\left(P_{[4]}\phi_i^{\otimes 4}\right)=1. \end{align}\] In conjunction with Lemma 4, Eq. (?? ) can be verified as follows: \[\begin{align} &\underset{\psi\sim \mathrm{Haar}}{\mathbb{E}} \left[\bar{\Phi}_4(\mathcal{E},\psi)\right]^2=D_{[4]}^2\sum_{i,j}w_iw_j\underset{\psi\sim \mathrm{Haar}}{\mathbb{E}}\left[\left[\mathop{\mathrm{tr}}(\psi \phi_{i})\right]^4\left[\mathop{\mathrm{tr}}(\psi \phi_{j})\right]^4\right] = \dfrac{D_{[4]}^2}{D_{[8]}}\sum_{i,j}w_iw_j\mathop{\mathrm{tr}}\left[P_{[8]}\left(\phi_{i}^{\otimes 4}\otimes \phi_{j}^{\otimes 4}\right)\right]\nonumber\\ &=\dfrac{D_{[4]}^2}{D_{[8]}}\sum_{i,j}\frac{w_iw_j\left\{1+16\mathop{\mathrm{tr}}(\phi_i\phi_j)+36\left[\mathop{\mathrm{tr}}(\phi_i\phi_j)\right]^2+16\left[\mathop{\mathrm{tr}}(\phi_i\phi_j)\right]^3+\left[\mathop{\mathrm{tr}}(\phi_i\phi_j)\right]^4\right\}}{70}\nonumber\\ &=\frac{D_{[4]}^2}{70D_{[8]}}\left[1+16\Phi_1(\mathcal{E})+36\Phi_2(\mathcal{E})+16\Phi_3(\mathcal{E})+\Phi_4(\mathcal{E})\right]=\frac{D_1}{D_2}\left(1+\dfrac{16\bar{\Phi}_1(\mathcal{E})}{D_{[1]}}+\dfrac{36\bar{\Phi}_2(\mathcal{E})}{D_{[2]}}+\dfrac{16\bar{\Phi}_3(\mathcal{E})}{D_{[3]}}+\dfrac{\bar{\Phi}_4(\mathcal{E})}{D_{[4]}}\right), \end{align}\] where \(D_1=d(d+1)(d+2)(d+3)\) and \(D_2=(d+4)(d+5)(d+6)(d+7)\).

Next, suppose \(\mathcal{E}\) is a state 2-design; then \(\bar{\Phi}_2(\mathcal{E})=\bar{\Phi}_1(\mathcal{E})=1\). The variance of \(\bar{\Phi}_4(\mathcal{E},\psi)\) can be upper bounded as follows: \[\begin{align} \mathrm{Var}(\bar{\Phi}_4(\mathcal{E},\psi)) &= \underset{\psi\sim \mathrm{Haar}}{\mathbb{E}} \left[\bar{\Phi}_4(\mathcal{E},\psi)\right]^2-\left[\underset{\psi\sim \mathrm{Haar}}{\mathbb{E}}\bar{\Phi}_4(\mathcal{E},\psi)\right]^2=\frac{D_1}{D_2}\left(1+\dfrac{16\bar{\Phi}_1(\mathcal{E})}{D_{[1]}}+\dfrac{36\bar{\Phi}_2(\mathcal{E})}{D_{[2]}}+\dfrac{16\bar{\Phi}_3(\mathcal{E})}{D_{[3]}}+\dfrac{\bar{\Phi}_4(\mathcal{E})}{D_{[4]}}\right)-1\nonumber\\ &=\frac{24[(4d+12)\bar{\Phi}_3(\mathcal{E})+\bar{\Phi}_4(\mathcal{E})-4d-13]}{D_2}\le \frac{6(17d+51)[\bar{\Phi}_3(\mathcal{E})-1]+6(d-1)}{D_2}\nonumber\\ &<\frac{(d+3)\{102[\bar{\Phi}_3(\mathcal{E})-1]+6\}}{D_2} <\frac{[\xi(\mathcal{E})]^2}{(d+6)^3}, \end{align}\] which confirms Eq. (?? ). Here the first inequality holds because \(\bar{\Phi}_4(\mathcal{E})\le {(d+3)\bar{\Phi}_3(\mathcal{E})}/{4}\) by Eq. (15 ), the second inequality holds because \[\bar{\Phi}_3(\mathcal{E})\ge 1,\quad \frac{(17d+51)[\bar{\Phi}_3(\mathcal{E})-1]+d-1}{d+3}<17[\bar{\Phi}_3(\mathcal{E})-1]+1,\] and the last inequality holds because \((d+3)/D_2<1/(d+6)^3\). This completes the proof of Lemma 5. ◻

Proof of Lemma 6. Suppose the ensemble \(\mathcal{E}\) has the form \(\mathcal{E}=\{\ket{\phi_i},w_i\}_i\). Then Eq. (?? ) can be proved as follows: \[\begin{align} \underset{\psi\sim \mathcal{T}_\mathcal{B}}{\mathbb{E}}\bar{\Phi}_t(\mathcal{E}, \psi)&=D_{[t]}\sum_iw_i\underset{\psi\sim \mathcal{T}_\mathcal{B}}{\mathbb{E}}\left[\mathop{\mathrm{tr}}(\phi_i\psi)\right]^t=D_{[t]}\sum_iw_i\mathop{\mathrm{tr}}\left[\phi_i^{\otimes t}\left( \underset{\psi\sim \mathcal{T}_\mathcal{B}}{\mathbb{E}}\psi^{\otimes t}\right)\right]\nonumber\\&\le \frac{t! D_{[t]}}{d^t}\sum_iw_i\mathop{\mathrm{tr}}\left(\mkern 2mu\phi_i^{\otimes t}P_{[t]}\right)=\frac{t! D_{[t]}}{d^t}, \end{align}\] where the inequality follows from Lemma 1. ◻

2 Shadow estimation within the framework of rank-1 IC-POVMs↩︎

In this section we briefly introduce shadow estimation within the framework of rank-1 IC POVMs, following the presentation in the main text. Let \(\mathcal{E}= \{\ket{\phi_i}, w_i\}_i\) be a state ensemble that forms a \(1\)-design, which induces the POVM \(\{ d w_i \phi_i \}_i\). If \(\mathcal{E}\) is informationally complete (IC), any quantum state \(\rho \in \mathcal{D}(\mathcal{H})\) can be reconstructed as \(\rho = d \sum_i w_i \mathop{\mathrm{tr}}(\phi_i \rho)\, \hat{\rho}_i\), where \(\hat{\rho}_i\) is the estimator associated with outcome \(i\). When \(\mathcal{E}\) is informationally overcomplete, the choice of \(\hat{\rho}_i\) is not unique. In the shadow-estimation framework, one usually employs the canonical reconstruction, which is the standard choice in the absence of prior information about \(\rho\).

The canonical measurement channel can be represented in two equivalent ways. The first directly defines the channel action: \[\begin{align} \mathcal{M}_\mathcal{E}(\rho) &= d \sum_i w_i \mathop{\mathrm{tr}}(\phi_i \rho)\, \phi_i. \label{eq:FD1} \end{align}\tag{22}\] The second uses the vectorization (double-ket) notation. For any operator \(A \in \mathcal{L}^{\mathrm{H}}(\mathcal{H})\), we define its vectorization \(\dket{A} \in \mathcal{H}\otimes \mathcal{H}\) by expanding \(\dket{A} = \sum_{k,l} A_{kl}\, \ket{k}\otimes\ket{l}\) in a fixed orthonormal basis, so that the Hilbert–Schmidt inner product becomes \(\dbraket{A}{B}=\mathop{\mathrm{tr}}\left(A^\dagger B\right)\). The outer product \(\dketbra{A}{A} = \dket{A}\dbra{A}\) then acts as a linear operator on this space. With this notation the measurement channel becomes \[\begin{align} \mathcal{M}_\mathcal{E}= d \sum_i w_i \dketbra{\phi_i}{\phi_i}, \label{eq:FD2} \end{align}\tag{23}\] where \(\mathcal{M}_\mathcal{E}\) is viewed as an operator acting on vectorized operators via \(\mathcal{M}_\mathcal{E}\dket{A} = \dket{\mathcal{M}_\mathcal{E}(A)}\). This representation makes the positive semidefiniteness of \(\mathcal{M}_\mathcal{E}\) manifest and facilitates spectral analysis and inversion of the channel. By definition, \(\mathcal{M}_\mathcal{E}(\mathbb{1}) = \mathbb{1}\), so \(\mathbb{1}\) is an eigenoperator of \(\mathcal{M}_\mathcal{E}\) with eigenvalue \(1\). With a slight abuse of notation, we use \(\mathcal{M}_\mathcal{E}\) to denote both representations of the measurement channel when no confusion arises. Let \(\bar{\mathbf{I}}\) denote the projector onto the space of traceless Hermitian operators. The restricted channel is \[\begin{align} \bar{\mathcal{M}}_\mathcal{E}\mathrel{\vcenter{:}}= \bar{\mathbf{I}}\, \mathcal{M}_\mathcal{E}\, \bar{\mathbf{I}}= d \sum_i w_i \dketbra{\phi_{i,0}}{\phi_{i,0}}, \end{align}\] where \(\phi_{i,0} = \phi_i - \mathbb{1}/d\) is the traceless part of \(\phi_i\). Because \(\mathcal{E}\) forms a \(1\)-design, this simplifies to \[\begin{align} \bar{\mathcal{M}}_\mathcal{E}= \mathcal{M}_\mathcal{E}- \frac{1}{d}\dketbra{\mathbb{1}}{\mathbb{1}}. \end{align}\]

For any observable \(O \in \mathcal{L}^{\mathrm{H}}(\mathcal{H})\) and an IC-POVM \(\mathcal{M}= \{E_i\}_i\), the expectation value \(\mathop{\mathrm{tr}}(\rho O)\) can be estimated as \[\begin{align} \mathop{\mathrm{tr}}(\rho O) = \sum_i \mathop{\mathrm{tr}}(E_i \rho)\, \hat{o}_i, \end{align}\] where \(\hat{o}_i = \mathop{\mathrm{tr}}(\hat{\rho}_i O)\) and \(\hat{\rho}_i\) is the canonical reconstruction operator for outcome \(i\). The single-shot estimation variance reads \[\begin{align} \operatorname{Var}(\hat{o}) = \sum_i \mathop{\mathrm{tr}}(E_i \rho)\, \hat{o}_i^2 - [\mathop{\mathrm{tr}}(\rho O)]^2 = \sum_i \mathop{\mathrm{tr}}(E_i \rho)\, \hat{o}_{0,i}^2 - [\mathop{\mathrm{tr}}(\rho O_0)]^2, \end{align}\] where \(O_0 = O - \mathop{\mathrm{tr}}(O)\mathbb{1}/d\) is the traceless part of \(O\) and \(\hat{o}_{0,i} = \mathop{\mathrm{tr}}(\hat{\rho}_i O_0)\). The second equality shows that the variance depends solely on \(O_0\). The dominant term defines the state-dependent squared shadow norm: \[\begin{align} \|O_0\|_{\mathcal{E},\rho}^2 \mathrel{\vcenter{:}}= \sum_i \mathop{\mathrm{tr}}(E_i \rho)\, \hat{o}_{0,i}^2 = d \sum_i w_i \mathop{\mathrm{tr}}(\rho \phi_i) \left\{ \mathop{\mathrm{tr}}\left[ \phi_i\, \mathcal{M}_\mathcal{E}^{-1}(O_0) \right] \right\}^2=d\mathop{\mathrm{tr}}\left\{Q_3(\mathcal{E}) \left[\rho \otimes \mathcal{M}_{\mathcal{E}}^{-1}(O_0)\otimes \mathcal{M}_{\mathcal{E}}^{-1}(O_0)\right]\right\}, \label{eq:ShadowNormSD} \end{align}\tag{24}\] where \(\mathcal{M}_\mathcal{E}^{-1}\) denotes the inverse of the measurement channel. The (state-independent) squared shadow norm is obtained by maximizing \(\|O\|_{\mathcal{E},\rho}^2\) over all states: \[\begin{align} \|O_0\|_\mathcal{E}^2 \mathrel{\vcenter{:}}= \max_{\rho\in\mathcal{D}(\mathcal{H})} \|O_0\|_{\mathcal{E},\rho}^2 = \left\| d \sum_i w_i \phi_i \left\{ \mathop{\mathrm{tr}}\left[ \phi_i\, \mathcal{M}_\mathcal{E}^{-1}(O_0) \right] \right\}^2 \right\|=\left\|d\mathop{\mathrm{tr}}_{2,3}\left\{Q_3(\mathcal{E}) \left[\mathbb{1}\otimes \mathcal{M}_{\mathcal{E}}^{-1}(O_0)\otimes \mathcal{M}_{\mathcal{E}}^{-1}(O_0)\right]\right\}\right\|, \label{eq:norm2} \end{align}\tag{25}\] which applies to any traceless operator \(O_0 \in \mathcal{L}^{\mathrm{H}}_0(\mathcal{H})\).

If \(\mathcal{E}\) forms a state \(2\)-design, the measurement channel takes the simple form \[\begin{align} \mathcal{M}_\mathcal{E}= \frac{\mathbf{I}+ \dketbra{\mathbb{1}}{\mathbb{1}}}{d+1}, \end{align}\] where \(\mathbf{I}\) is the identity superoperator on \(\mathcal{L}^{\mathrm{H}}(\mathcal{H})\). Restricted to traceless operators, its inverse gives \[\begin{align} \mathcal{M}_\mathcal{E}^{-1}(O_0) = \bar{\mathcal{M}}_\mathcal{E}^{-1}(O_0) = (d+1)\, O_0. \end{align}\] Consequently, the squared shadow norm reduces to \[\begin{align} \|O_0\|_\mathcal{E}^2 = \left\| d(d+1)^2 \sum_i w_i \phi_i \left[ \mathop{\mathrm{tr}}(\phi_i O_0) \right]^2 \right\|=\left\|d(d+1)^2\mathop{\mathrm{tr}}_{2,3}\left[Q_3(\mathcal{E}) \left(\mathbb{1}\otimes O_0\otimes O_0\right)\right]\right\|. \end{align}\] If \(\mathcal{E}\) further forms a state \(3\)-design, then \(Q_3(\mathcal{E})=6P_{[3]}/[d(d+1)(d+2)]\), and we obtain the closed-form expression \[\begin{align} \|O_0\|_\mathcal{E}^2 = \frac{d+1}{d+2}\left( 2\|O_0\|^2 + \|O_0\|_2^2 \right) \le 3\|O_0\|_2^2. \end{align}\] For fidelity estimation, where the observable is a pure-state projector \(\psi\) with traceless part \(\psi_0 = \psi - \mathbb{1}/d\), the shadow norm evaluates to \[\begin{align} \|\psi_0\|_\mathcal{E}^2 = \frac{(3d-2)(d+1)(d-1)}{d^2(d+2)}. \label{eq:3design95pure} \end{align}\tag{26}\]

3 Proofs of results on worst-case optimal shadow estimation↩︎

3.1 Proof of Proposition 2↩︎

Proof. Suppose the ensemble has the form \(\mathcal{E}=\{\ket{\phi_i},w_i\}_{i=1}^{K}\). For any collection of states \(\{\rho_j\}_{j=1}^K \subset \mathcal{D}(\mathcal{H})\), any collection of rank-1 projectors \(\{\psi_j\}_{j=1}^K \subset \mathcal{P}(\mathcal{H})\), and any probability distribution \(\{v_j\}_{j=1}^K\) (i.e., \(v_j \ge 0\) and \(\sum_j v_j = 1\)), we have \[\begin{align} \max_{\psi\in \mathcal{P}(\mathcal{H})} \|\psi_0\|_\mathcal{E}^2 &\mathrel{\vcenter{:}}= \max_{\rho\in \mathcal{D}(\mathcal{H}),\,\psi\in \mathcal{P}(\mathcal{H})} d\sum_i w_i\mathop{\mathrm{tr}}(\rho \phi_i) \left\{\mathop{\mathrm{tr}}\left[\phi_i\mathcal{M}_\mathcal{E}^{-1}(\psi_0) \right]\right\}^2 \nonumber\\ &\,\ge d\sum_{i,j} w_i v_j{\mathop{\mathrm{tr}}(\rho_j \phi_i) \left\{\mathop{\mathrm{tr}}\left[\phi_i\mathcal{M}_\mathcal{E}^{-1}(\psi_{j,0}) \right]\right\}^2}, \label{eq:LBK} \end{align}\tag{27}\] where the inequality follows from the fact that the maximum is lower bounded by any convex combination. To establish the inequality in Proposition 2, we now choose \[\begin{align} \rho_j=\phi_j, \qquad \psi_{j}=\phi_j, \qquad v_j=\frac{1}{K}. \end{align}\] Then we obtain \[\begin{align} \max_{\psi\in \mathcal{P}(\mathcal{H})} \|\psi_0\|_\mathcal{E}^2 &\ge d\sum_{i,j} w_i v_j \mathop{\mathrm{tr}}(\rho_j \phi_i) \left\{\mathop{\mathrm{tr}}\left[\phi_i\mathcal{M}_\mathcal{E}^{-1}(\phi_{j,0}) \right]\right\}^2 \ge \frac{d}{K}\sum_i w_i \mathop{\mathrm{tr}}(\rho_i \phi_i) \left\{\mathop{\mathrm{tr}}\left[\phi_i\mathcal{M}_\mathcal{E}^{-1}(\phi_{i,0}) \right]\right\}^2 \nonumber\\ &= \frac{d}{K}\sum_i w_i \left\{\mathop{\mathrm{tr}}\left[\phi_{i,0}\bar{\mathcal{M}}_\mathcal{E}^{-1}(\phi_{i,0}) \right]\right\}^2.\label{eq:LBIC} \end{align}\tag{28}\] The second inequality holds because \(\mathop{\mathrm{tr}}(\rho_j\phi_i)\) and \(\left\{\mathop{\mathrm{tr}}\left[\phi_i\mathcal{M}_\mathcal{E}^{-1}(\phi_{j,0}) \right]\right\}^2\) are nonnegative for all \(i,j\); the equality follows from the facts that \(\phi_{i,0}\) is traceless and that \(\mathcal{M}_\mathcal{E}\) restricted to \(\mathcal{L}^{\mathrm{H}}_0(\mathcal{H})\) coincides with \(\bar{\mathcal{M}}_\mathcal{E}\).

Since \(f(x)=x^t\) is convex for \(t>1\), Jensen’s inequality implies that \[\begin{align} \sum_i w_i \left\{\mathop{\mathrm{tr}}\left[\phi_{i,0}\bar{\mathcal{M}}_\mathcal{E}^{-1}(\phi_{i,0}) \right]\right\}^2 \ge \left\{\sum_i w_i \mathop{\mathrm{tr}}\left[\phi_{i,0}\bar{\mathcal{M}}_\mathcal{E}^{-1}(\phi_{i,0}) \right]\right\}^2.\label{eq:Jensen} \end{align}\tag{29}\] To evaluate the right-hand side, note that \[\begin{align} \sum_i w_i \mathop{\mathrm{tr}}\left[\phi_{i,0}\bar{\mathcal{M}}_\mathcal{E}^{-1}(\phi_{i,0}) \right] &= \mathop{\mathrm{Tr}}\left[\sum_i w_i \dketbra{\phi_{i,0}}{\phi_{i,0}}\bar{\mathcal{M}}_\mathcal{E}^{-1}\right] = \frac{1}{d}\mathop{\mathrm{Tr}}\left(\bar{\mathcal{M}}_\mathcal{E}\bar{\mathcal{M}}_\mathcal{E}^{-1}\right) = \frac{1}{d}\mathop{\mathrm{Tr}}(\bar{\mathbf{I}})=\frac{d^2-1}{d}, \label{eq:TrBId} \end{align}\tag{30}\] where \(\mathop{\mathrm{Tr}}\) denotes the trace on \(\mathcal{L}^{\mathrm{H}}(\mathcal{H})\) (viewing \(\mathcal{M}_\mathcal{E}\) as a linear operator on the operator space), and \(\bar{\mathbf{I}}\) denotes the identity on \(\mathcal{L}^{\mathrm{H}}_0(\mathcal{H})\). Combining Eqs. 2830 yields \[\begin{align} \max_{\psi\in \mathcal{P}(\mathcal{H})} \|\psi_0\|_\mathcal{E}^2 \ge \frac{d}{K}\left(\frac{d^2-1}{d}\right)^2 = \frac{(d^2-1)^2}{Kd}, \end{align}\] which establishes the desired inequality in Proposition 2. ◻

3.2 Auxiliary results on the combined phase-design ensemble based on MUBs↩︎

In this section, we clarify the basic properties of the combined phase 3-design ensemble \(\mathcal{E}_N\) with \(2\le N\le d+1\), as introduced in the main text, and provide a proof of Theorem 2. This ensemble is a uniform mixture of \(N\) phase 3-design ensembles \(\mathcal{E}_{\mathcal{B}_j}\) associated with \(N\) mutually unbiased bases (MUBs), denoted by \(\mathcal{B}_j=\{\ket{k}_{\mathcal{B}_j}\}_{k=0}^{d-1}\) for \(j=1,2,\ldots,N\). More precisely, \(\mathcal{E}_N\) can be expressed as follows: \[\begin{align} \mathcal{E}_N = \bigsqcup_{j=1}^{N} \frac{1}{N}\mathcal{E}_{\mathcal{B}_j}. \end{align}\]

Proposition 7. The second and third moment operators of \(\mathcal{E}_N\) read \[\begin{align} Q_2(\mathcal{E}_N) &= \frac{2}{d^2}P_{[2]} -\frac{1}{Nd^2}\sum_{j=1}^{N}\sum_{k=0}^{d-1}(\ket{kk}\mkern-2mu\bra{kk})_{\mathcal{B}_j} \label{eq:2ndMA}\\ Q_3(\mathcal{E}_N) &= \frac{6}{d^3}P_{[3]} -\frac{5}{Nd^3}\sum_{j=1}^{N}\sum_{k=0}^{d-1}(\ket{kkk}\mkern-2mu\bra{kkk})_{\mathcal{B}_j} -\frac{3}{Nd^3}\sum_{j=1}^{N}\sum_{[\boldsymbol{k}]\colon |[\boldsymbol{k}]|=3} (\ket{[\boldsymbol{k}]}\mkern-2mu\bra{[\boldsymbol{k}]})_{\mathcal{B}_j}. \label{eq:3ndMA} \end{align}\] {#eq: sublabel=eq:eq:2ndMA,eq:eq:3ndMA}

Proposition 8. The second frame potential of \(\mathcal{E}_N\) reads \[\Phi_2(\mathcal{E}_N)=\frac{2d^2-2d+1}{d^4}+\frac{d-1}{Nd^4}.\] The ensemble \(\mathcal{E}_N\) is a state \(2\)-design if and only if \(N=d+1\).

Proposition 7 is a direct corollary of Lemma 1, and Proposition 8 follows from Proposition 7, given that the bases \(\mathcal{B}_1, \mathcal{B}_2, \ldots, \mathcal{B}_N\) are mutually unbiased. When \(N = d+1\), the \(N\) bases form a complete set of MUBs and thus constitute a 2-design, which implies that \[\frac{1}{Nd}\sum_{j=1}^{N}\sum_{k=0}^{d-1}(\ket{kk}\mkern-2mu\bra{kk})_{\mathcal{B}_j} = \frac{2P_{[2]}}{d(d+1)}.\] As a corollary, Proposition 7 yields \(Q_2(\mathcal{E}_N) = 2P_{[2]}/[d(d+1)]\), consistent with \(\mathcal{E}_N\) forming a 2-design.

Next, we decompose any operator \(O \in \mathcal{L}^{\mathrm{H}}(\mathcal{H})\) into orthogonal parts as \[\begin{align} O =\frac{\mathop{\mathrm{tr}}(O)}{d}\mathbb{1}+O_{\perp}+ \sum_{j=1}^{N} O_{\mathcal{B}_j} , \end{align}\] where \(O_{\mathcal{B}_j}\) denotes the traceless diagonal part of \(O\) with respect to basis \(\mathcal{B}_j\), and \(O_{\perp}\) is the remaining orthogonal component. When \(N=d+1\), we have \(O_{\perp}=0\).

Proposition 9. Suppose \(O\in \mathcal{L}^{\mathrm{H}}(\mathcal{H})\). Then \[\begin{align} \mathcal{M}_{\mathcal{E}_N}(O) &= \frac{\mathop{\mathrm{tr}}(O)}{d}\mathbb{1}+\frac{1}{d}O_{\perp} + \sum_{j=1}^{N}\frac{N-1}{Nd}O_{\mathcal{B}_j}, \label{eq:caMcaE95N}\\ \mathcal{M}^{-1}_{\mathcal{E}_N}(O) &= \frac{\mathop{\mathrm{tr}}(O)}{d}\mathbb{1}+dO_{\perp} +\sum_{j=1}^{N}\frac{Nd}{N-1}O_{\mathcal{B}_j}, \label{eq:caMcaE95Ninv} \end{align}\] {#eq: sublabel=eq:eq:caMcaE95N,eq:eq:caMcaE95Ninv} where \(\mathcal{M}_{\mathcal{E}_N}\) is defined in Eqs. (22 ) and (23 ).

Lemma 7. Suppose \(\rho\in \mathcal{D}(\mathcal{H})\), \(O\in \mathcal{L}^{\mathrm{H}}_0(\mathcal{H})\), and \(\mathcal{B}= \{\ket{k}_\mathcal{B}\}_{k=0}^{d-1}\) is an orthonormal basis of \(\mathbb{C}^d\). Define \[\begin{align} \Pi_\mathcal{B}(\rho, O) \mathrel{\vcenter{:}}= \mathop{\mathrm{tr}}\left[\sum_{[\boldsymbol{k}]\colon|[\boldsymbol{k}]|=3} |[\boldsymbol{k}]|(\ket{[\boldsymbol{k}]}\mkern-2mu\bra{[\boldsymbol{k}]})_\mathcal{B}\left(\rho \otimes O\otimes O\right)\right]. \end{align}\] Then \[\begin{align} \label{eq:PI3A} -\Pi_\mathcal{B}(\rho, O) \le \mathop{\mathrm{tr}}\left(\rho O_\mathcal{B}^2\right) + \left\|O_\mathcal{B}\right\|_2^2. \end{align}\qquad{(12)}\]

Lemma 8. Suppose \(\rho\in \mathcal{D}(\mathcal{H})\) and \(O\in \mathcal{L}^{\mathrm{H}}_0(\mathcal{H})\). Then \[\begin{align} \label{eq:7N} d^3\mathop{\mathrm{tr}}\left[Q_3(\mathcal{E}_N)\left(\rho\otimes O\otimes O\right)\right] \le \left(3+\frac{1}{N}\right)\|O\|_2^2. \end{align}\qquad{(13)}\]

Proof of Proposition 9. Suppose the ensemble \(\mathcal{E}_N\) has the form \(\mathcal{E}_N=\{\ket{\phi_i}, w_i\}_i\). By Eq. (22 ), for any \(O\in \mathcal{L}^{\mathrm{H}}(\mathcal{H})\), the measurement channel acts as \[\begin{align} \mathcal{M}_{\mathcal{E}_N}(O) = d \sum_i w_i\mathop{\mathrm{tr}}(\phi_i O)\,\phi_i = d\mathop{\mathrm{tr}}_1\left[Q_2(\mathcal{E}_N)\left(O\otimes \mathbb{1}\right)\right]. \end{align}\] Direct calculation based on Eq. (?? ) in Proposition 7 yields \[\begin{align} \mathcal{M}_{\mathcal{E}_N}(\mathbb{1})=\mathbb{1},\qquad \mathcal{M}_{\mathcal{E}_N}(O_{\mathcal{B}_j})=\frac{N-1}{Nd}\,O_{\mathcal{B}_j},\qquad \mathcal{M}_{\mathcal{E}_N}(O_\perp)=\frac{1}{d}\,O_\perp. \end{align}\] By linearity, we have \[\begin{align} \mathcal{M}_{\mathcal{E}_N}(O) &= \frac{\mathop{\mathrm{tr}}(O)}{d}\mathbb{1}+\frac{1}{d}O_{\perp} + \sum_{j=1}^{N}\frac{N-1}{Nd}O_{\mathcal{B}_j}, \end{align}\] which confirms Eq. (?? ). Since \(\mathcal{M}_{\mathcal{E}_N}(O)\) preserves the orthogonal decomposition, by inverting it on each eigenspace, we readily obtain the expression for \(\mathcal{M}_{\mathcal{E}_N}^{-1}(O)\) shown in Eq. (?? ). ◻

Proof of Lemma 7. Without loss of generality, assume that \(\mathcal{B}\) is the computational basis and \(\ket{k}_\mathcal{B}=\ket{k}\) for \(k=0,1,\ldots, d-1\). Then \[\begin{align} \sum_{[\boldsymbol{k}]\colon|[\boldsymbol{k}]|=3} |[\boldsymbol{k}]|\ket{[\boldsymbol{k}]}\mkern-2mu\bra{[\boldsymbol{k}]} &= \sum_{\substack{k,l \\ k \neq l}}\ket{kll}\mkern-2mu\bra{kll} + \sum_{\substack{k,l \\ k \neq l}}\ket{kll}\mkern-2mu\bra{lkl} + \sum_{\substack{k,l \\ k \neq l}}\ket{kll}\mkern-2mu\bra{llk} + \sum_{\substack{k,l \\ k \neq l}}\ket{lkl}\mkern-2mu\bra{kll} + \sum_{\substack{k,l \\ k \neq l}}\ket{lkl}\mkern-2mu\bra{lkl} + \sum_{\substack{k,l \\ k \neq l}}\ket{lkl}\mkern-2mu\bra{llk}\nonumber\\ &\,\hphantom{=}\,+ \sum_{\substack{k,l \\ k \neq l}}\ket{llk}\mkern-2mu\bra{kll} + \sum_{\substack{k,l \\ k \neq l}}\ket{llk}\mkern-2mu\bra{lkl} + \sum_{\substack{k,l \\ k \neq l}}\ket{llk}\mkern-2mu\bra{llk}. \end{align}\] Taking the trace against \(\rho \otimes O \otimes O\) gives \[\begin{align} \Pi_\mathcal{B}(\rho, O) &= \sum_{\substack{k,l \\ k \neq l}}\rho_{kk}O_{ll}^2 + 2\sum_{\substack{k,l \\ k \neq l}}\rho_{ll}\left|O_{kl}\right|^2 + 4\sum_{\substack{k,l \\ k \neq l}}\Re\left(\rho_{kl}O_{lk}\right)O_{ll} + 2\sum_{\substack{k,l \\ k \neq l}}\rho_{ll}O_{ll}O_{kk}\nonumber\\ &= 2\sum_{\substack{k,l \\ k \neq l}}\rho_{kk}O_{ll}^2 + 2\sum_{\substack{k,l \\ k \neq l}}\rho_{ll}\left|O_{lk}\right|^2 + 4\sum_{\substack{k,l \\ k \neq l}}\Re\left(\rho_{kl}O_{lk}\right)O_{ll} - \sum_{l}\rho_{ll}O_{ll}^2 - \sum_l O_{ll}^2, \label{eq:PirhoX0} \end{align}\tag{31}\] where the second equality uses the facts \(|O_{kl}|=|O_{lk}|\), \(\sum_k\rho_{kk}=\mathop{\mathrm{tr}}(\rho)=1\), and \(\sum_k O_{kk}=\mathop{\mathrm{tr}}(O)=0\). The sum of the first three terms is nonnegative because \[\rho_{kk}O_{ll}^2 + \rho_{ll}\left|O_{lk}\right|^2 \geq 2\sqrt{\rho_{kk}\rho_{ll}}\,|O_{lk}O_{ll}| \geq 2|\rho_{kl}O_{lk}O_{ll}| \geq -2\Re(\rho_{kl}O_{lk})O_{ll}\quad \forall k,l,\] where the second inequality holds because every \(2\times 2\) principal submatrix of \(\rho\) with respect to \(\mathcal{B}\) is positive semidefinite. Therefore, \[\begin{align} \Pi_\mathcal{B}(\rho, O) \ge -\mathop{\mathrm{tr}}\left(\rho O_\mathcal{B}^2\right) - \left\|O_\mathcal{B}\right\|_2^2, \end{align}\] which establishes Eq. (?? ). ◻

Proof of Lemma 8. From Proposition 7, we have \[\begin{align} Q_3(\mathcal{E}_N) = \frac{6}{d^3}P_{[3]} - \frac{5}{Nd^3}\sum_{j=1}^{N}\sum_{k=0}^{d-1}(\ket{kkk}\mkern-2mu\bra{kkk})_{\mathcal{B}_j} - \frac{3}{Nd^3}\sum_{j=1}^{N}\sum_{[\boldsymbol{k}]\colon|[\boldsymbol{k}]|=3} (\ket{[\boldsymbol{k}]}\mkern-2mu\bra{[\boldsymbol{k}]})_{\mathcal{B}_j}. \end{align}\] Since \(O\) is traceless by assumption, we can apply the following two key bounds: \[\begin{align} 6\mathop{\mathrm{tr}}\left[P_{[3]}\left(\rho\otimes O\otimes O\right)\right] &= 2\mathop{\mathrm{tr}}(\rho O^2) + \|O\|_2^2 \le 3\|O\|_2^2, \quad -\Pi_{\mathcal{B}_j}(\rho, O) \le \mathop{\mathrm{tr}}\left(\rho O_{\mathcal{B}_j}^2\right) + \left\|O_{\mathcal{B}_j}\right\|_2^2, \end{align}\] where the latter follows from Lemma 7. Combining the above two equations, we obtain \[\begin{align} d^3\mathop{\mathrm{tr}}\left[Q_3(\mathcal{E}_N)\left(\rho\otimes O\otimes O\right)\right] &= 6\mathop{\mathrm{tr}}\left[P_{[3]}\left(\rho\otimes O\otimes O\right)\right] - \frac{5}{N}\sum_{j=1}^N\mathop{\mathrm{tr}}\left(\rho O_{\mathcal{B}_j}^2\right) - \frac{1}{N}\sum_{j=1}^N\Pi_{\mathcal{B}_j}(\rho,O)\nonumber\\ &\le 3\|O\|_2^2 + \frac{1}{N}\sum_{j=1}^N\left\|O_{\mathcal{B}_j}\right\|_2^2 - \frac{4}{N}\sum_{j=1}^N\mathop{\mathrm{tr}}\left(\rho O_{\mathcal{B}_j}^2\right)\le \left(3+\frac{1}{N}\right)\|O\|_2^2, \label{eq:P4rhoX} \end{align}\tag{32}\] where the last step holds because \(\sum_{j=1}^N\left\|O_{\mathcal{B}_j}\right\|_2^2 \le \|O\|_2^2\) and \(\mathop{\mathrm{tr}}\bigl(\rho O_{\mathcal{B}_j}^2\bigr)\geq0\) for each \(j\). ◻

3.3 Proof of Theorem 2↩︎

Proof. By Eq. (24 ) and Lemma 8, we bound the squared shadow norm as follows: \[\begin{align} \left\|O\right\|_{\mathcal{E}_N}^2=d\max_{\rho\in \mathcal{D}(\mathcal{H})}\mathop{\mathrm{tr}}\left\{Q_3(\mathcal{E}) \left[\rho \otimes \mathcal{M}_{\mathcal{E}_N}^{-1}(O)\otimes \mathcal{M}_{\mathcal{E}_N}^{-1}(O)\right]\right\}\le \frac{3N+1}{Nd^2}\left\|\mathcal{M}_{\mathcal{E}_N}^{-1}(O)\right\|_2^2. \label{eq:N1} \end{align}\tag{33}\] According to Proposition 9, \(\mathcal{M}_{\mathcal{E}_N}^{-1}(O)\) has the form \[\begin{align} \mathcal{M}_{\mathcal{E}_N}^{-1}(O) = dO_{\perp}+\sum_{j=1}^N\frac{Nd}{N-1}\,O_{\mathcal{B}_j} , \end{align}\] where the identity component vanishes because \(\mathop{\mathrm{tr}}(O)=0\) by assumption. These components are mutually orthogonal with respect to the Hilbert–Schmidt inner product, so \[\begin{align} \left\|\mathcal{M}_{\mathcal{E}_N}^{-1}(O)\right\|_2^2 = d^2\left[\left\|O_{\perp}\right\|_2^2+\frac{N^2}{(N-1)^2}\sum_{j=1}^N\left\|O_{\mathcal{B}_j}\right\|_2^2\right]\le \frac{N^2 d^2}{(N-1)^2}\,\|O\|_2^2, \label{eq:N295expand} \end{align}\tag{34}\] where the inequality holds because \(\sum_{j=1}^N\left\|O_{\mathcal{B}_j}\right\|_2^2 + \left\|O_{\perp}\right\|_2^2 = \|O\|_2^2\). Together with Eq. (33 ), this equation means \[\begin{align} \|O\|_{\mathcal{E}_N}^2 \le \frac{3N+1}{Nd^2} \cdot \frac{N^2 d^2}{(N-1)^2}\,\|O\|_2^2 = \frac{N(3N+1)}{(N-1)^2}\,\|O\|_2^2, \end{align}\] which completes the proof of Theorem 2. ◻

4 Fourth moments of Haar-random observables and Clifford orbits↩︎

4.1 Fourth moments of Haar-random observables↩︎

In preparation for studying the average shadow norm achieved by 2-design POVMs, we recall some basic results about Schur–Weyl duality for the unitary group.

Let \(\mathrm{U}(\mathcal{H})\) be the unitary group on \(\mathcal{H}\) and \(S_t\) the symmetric group of \(t\) numbers. According to Schur–Weyl duality, the \(t\)-th tensor power \(\mathcal{H}^{\otimes t}\) decomposes into a multiplicity-free direct sum of irreducible representations of \(\mathrm{U}(\mathcal{H}) \times S_t\): \[\begin{align} \mathcal{H}^{\otimes t}=\bigoplus_\lambda \mathcal{H}_\lambda=\bigoplus_\lambda \mathcal{W}_\lambda\otimes \mathcal{S}_\lambda, \end{align}\] where each \(\lambda\) labels a nonincreasing partition of \(t\) into at most \(d\) parts, \(\mathcal{W}_\lambda\) is the Weyl module carrying an irreducible representation of \(\mathrm{U}(\mathcal{H})\), and \(\mathcal{S}_\lambda\) is the Specht module carrying an irreducible representation of \(S_t\). The dimensions of \(\mathcal{W}_\lambda\) and \(\mathcal{S}_\lambda\) are denoted by \(D_\lambda\) and \(d_\lambda\), respectively. The projector \(P_{\lambda}\) onto \(\mathcal{H}_\lambda\) can be expressed as \[\begin{align} P_{\lambda}=\frac{d_\lambda}{24}\sum_{\sigma\in S_4}\chi_\lambda(\sigma)R(\sigma), \label{eq:Plambda} \end{align}\tag{35}\] where \(\chi_\lambda(\sigma)\) is the character of \(\sigma\) in the representation \(\lambda\), and \(R(\sigma)\) is the unitary operator representing the permutation \(\sigma\). By definition, \(\mathop{\mathrm{tr}}(P_\lambda)=\mathop{\mathrm{rank}}(P_\lambda)=d_\lambda D_\lambda\).

Lemma 9. Suppose \(\mathcal{U}\) is a unitary 4-design and \(O \in \mathcal{L}(\mathcal{H})\). Then \[\begin{align} \underset{U\sim\mathcal{U}}{\mathbb{E}}\, U^{\otimes 4} O^{\otimes 4} U^{\dagger \otimes 4}=\sum_{\lambda}\mathop{\mathrm{tr}}\left( P_{\lambda} O^{\otimes 4}\right)\frac{P_\lambda}{d_\lambda D_\lambda}, \end{align}\] where \(\underset{U\sim\mathcal{U}}{\mathbb{E}}\) denotes the average over \(\mathcal{U}\).

Proof. By assumption, \(\mathcal{U}\) is a unitary 4-design, which implies that \[\begin{align} \underset{U\sim\mathcal{U}}{\mathbb{E}}\, U^{\otimes 4} O^{\otimes 4} U^{\dagger \otimes 4}=\underset{U\sim \mathrm{Haar}}{\mathbb{E}}\, U^{\otimes 4} O^{\otimes 4} U^{\dagger \otimes 4}. \end{align}\] Thus both sides are invariant under permutations and diagonal unitary transformations of the form \(U^{\otimes 4}\) for \(U\in \mathrm{U}(\mathcal{H})\). The lemma then follows directly from Schur–Weyl duality. ◻

Table 1: Dimensions of the Specht module \(\caS_\lambda\) and Weyl module \(\caW_\lambda\).
\(\lambda\) \(d_\lambda\) \(D_\lambda\)
\([4]\) \(1\) \(\dfrac{d(d+1)(d+2)(d+3)}{24}\)
\([1,1,1,1]\) \(1\) \(\dfrac{d(d-1)(d-2)(d-3)}{24}\)
\([2,2]\) \(2\) \(\dfrac{d^2(d^2-1)}{12}\)
\([2,1,1]\) \(3\) \(\dfrac{d(d-2)(d^2-1)}{8}\)
\([3,1]\) \(3\) \(\dfrac{d(d+2)(d^2-1)}{8}\)
Table 2: Characters of the symmetric group \(S_4\).
Cycle type \((1^4)\) \((2^2)\) \((2,1^2)\) \((3,1)\) \((4)\)
Order \(1\) \(2\) \(2\) \(3\) \(4\)
No.of elements \(1\) \(3\) \(6\) \(8\) \(6\)
\(\chi_1 = [4]\) \(1\) \(1\) \(1\) \(1\) \(1\)
\(\chi_2 = [1,1,1,1]\) \(1\) \(1\) \(-1\) \(1\) \(-1\)
\(\chi_3 = [2,2]\) \(2\) \(2\) \(0\) \(-1\) \(0\)
\(\chi_4 = [2,1,1]\) \(3\) \(-1\) \(-1\) \(0\) \(1\)
\(\chi_5 = [3,1]\) \(3\) \(-1\) \(1\) \(0\) \(-1\)

In the rest of this section, we focus on the case \(t = 4\). There are five potential partitions: \([4]\), \([2,2]\), \([2,1,1]\), \([3,1]\), and \([1,1,1,1]\), where \([2,1,1]\) is relevant when \(d\geq 3\), and \([1,1,1,1]\) is relevant when \(d\geq 4\). The dimensions of the Specht and Weyl modules for these partitions are presented in Table 1, and the corresponding characters in Table 2. Using Table 2 and Eq. (35 ), we obtain the following lemma.

Lemma 10. Suppose \(O \in \mathcal{L}^{\mathrm{H}}_0(\mathcal{H})\) and \(\phi_i, \phi_j \in \mathcal{P}(\mathcal{H})\). Then \[\begin{align} {2} \mathop{\mathrm{tr}}\left( P_{[4]} O^{\otimes 4}\right)&=\frac{\|O\|_2^4+2\|O\|_4^4}{8},&\quad \mathop{\mathrm{tr}}\left[P_{[4]} \left(\phi_i^{\otimes 2}\otimes\phi_j^{\otimes 2}\right)\right]&=\frac{\left[\mathop{\mathrm{tr}}(\phi_i\phi_j)\right]^2+4\mathop{\mathrm{tr}}(\phi_i\phi_j)+1}{6}, \nonumber\\ \mathop{\mathrm{tr}}\left( P_{[2,2]} O^{\otimes 4}\right)&=\frac{\|O\|_2^4}{2},&\quad \mathop{\mathrm{tr}}\left[ P_{[2,2]} \left(\phi_i^{\otimes 2}\otimes\phi_j^{\otimes 2}\right)\right]&=\frac{\left[\mathop{\mathrm{tr}}(\phi_i\phi_j)\right]^2-2\mathop{\mathrm{tr}}(\phi_i\phi_j)+1}{3}, \nonumber\\ \mathop{\mathrm{tr}}\left( P_{[1,1,1,1]} O^{\otimes 4}\right)&=\frac{\|O\|_2^4-2\|O\|_4^4}{8},&\quad\mathop{\mathrm{tr}}\left[ P_{[1,1,1,1]} \left(\phi_i^{\otimes 2}\otimes\phi_j^{\otimes 2}\right)\right]&=0,\\ \mathop{\mathrm{tr}}\left( P_{[2,1,1]} O^{\otimes 4}\right)&=\frac{-3\|O\|_2^4+6\|O\|_4^4}{8},&\quad \mathop{\mathrm{tr}}\left[ P_{[2,1,1]}\left(\phi_i^{\otimes 2}\otimes\phi_j^{\otimes 2}\right)\right]&=0, \nonumber\\ \mathop{\mathrm{tr}}\left( P_{[3,1]} O^{\otimes 4}\right)&=\frac{-3\|O\|_2^4-6\|O\|_4^4}{8},&\quad\mathop{\mathrm{tr}}\left[ P_{[3,1]}\left(\phi_i^{\otimes 2}\otimes\phi_j^{\otimes 2}\right)\right]&=\frac{-\left[\mathop{\mathrm{tr}}(\phi_i\phi_j)\right]^2+1}{2}.\nonumber \end{align}\]

4.2 Fourth moments of Clifford orbits↩︎

In this appendix, we focus on an \(n\)-qubit system, whose Hilbert space \(\mathcal{H}\) has dimension \(d = 2^n\).

Let \(\{I,X,Y,Z\}^{\otimes n}\) be the set of \(n\)-qubit Pauli operators. Then \(\{W^{\otimes 4}\mid W\in \{I,X,Y,Z\}^{\otimes n}\}\) forms a stabilizer group with stabilizer projector \[P_n=\frac{1}{d^2}\sum_{W\in \{I,X,Y,Z\}^{\otimes n}} W^{\otimes 4}.\] For a pure state \(\psi \in \mathcal{P}(\mathcal{H})\), the stabilizer 2-Rényi entropy [62] is defined as \[\begin{align} M_2(\psi)\mathrel{\vcenter{:}}=-\log_2\sum_{W\in\{I,X,Y,Z\}^{\otimes n}}\left[\mathop{\mathrm{tr}}(W\psi)\right]^4+\log_2 d = -\log_2\mathop{\mathrm{tr}}\left(P_n\psi^{\otimes 4}\right)-\log_2 d. \end{align}\] By definition, \(\mathop{\mathrm{tr}}\left(P_n\psi^{\otimes 4}\right)=2^{-M_2(\psi)}/d\). As established in Ref. [51], the stabilizer 2-Rényi entropy satisfies \[\begin{align} 0\le M_2(\psi)\le \log_2(d+1)-1 \quad \forall \psi\in \mathcal{P}(\mathcal{H}), \label{eq:stab95bound} \end{align}\tag{36}\] where the lower bound is saturated if and only if \(\psi\) is a stabilizer state. Under the action of the Clifford group, the symmetric subspace of \(\mathcal{H}^{\otimes 4}\) decomposes into two inequivalent irreducible components with projectors \[\begin{align} P_+=P_{[4]}P_n=P_nP_{[4]},\quad P_-=P_{[4]}(\mathbb{1}-P_n)=(\mathbb{1}-P_n)P_{[4]}, \end{align}\] and dimensions \[\begin{align} D^{+}_{[4]}=\mathop{\mathrm{tr}}(P_+)=\frac{(d+1)(d+2)}{6},\quad D^{-}_{[4]}=\mathop{\mathrm{tr}}(P_-)=\frac{(d^2-1)(d+2)(d+4)}{24}. \label{eq:D4pm} \end{align}\tag{37}\] By construction, we have \[\begin{align} \mathop{\mathrm{tr}}\left(P_+\psi^{\otimes 4}\right)=\mathop{\mathrm{tr}}\left(P_n\psi^{\otimes 4}\right)=\frac{2^{-M_2(\psi)}}{d},\quad \mathop{\mathrm{tr}}\left(P_-\psi^{\otimes 4}\right)=1-\frac{2^{-M_2(\psi)}}{d}\quad \forall \psi\in \mathcal{P}(\mathcal{H}). \label{eq:PpPmProb} \end{align}\tag{38}\]

Lemma 11. Suppose \(\psi, \phi \in \mathcal{P}(\mathcal{H})\). Then \[\begin{align} \underset{U\sim \mathrm{Cl}(n)}{\mathbb{E}}\left(U \psi U^{\dagger}\right)^{\otimes 4} &=\frac{2^{-M_{2}(\psi)}}{d\,D^{+}_{[4]}}\,P_++\frac{1}{D^{-}_{[4]}}\left(1-\frac{2^{-M_{2}(\psi)}}{d}\right)P_-, \label{eq:4thTwirlCliff}\\ \underset{U\sim \mathrm{Cl}(n)}{\mathbb{E}}\left[\mathop{\mathrm{tr}}\left(\phi\, U \psi U^{\dagger}\right)\right]^4 &=\frac{2^{-M_{2}(\psi)-M_{2}(\phi)}}{d^2\,D^{+}_{[4]}}+\frac{1}{D^{-}_{[4]}}\left(1-\frac{2^{-M_{2}(\phi)}}{d}\right)\left(1-\frac{2^{-M_{2}(\psi)}}{d}\right)\leq \frac{5(d+3)}{4(d+4)D_{[4]}}\leq \frac{5}{4D_{[4]}}, \label{eq:4thTwirlClifftr} \end{align}\] {#eq: sublabel=eq:eq:4thTwirlCliff,eq:eq:4thTwirlClifftr} where \(D^{+}_{[4]}\) and \(D^{-}_{[4]}\) are given in Eq. (37 ), and \(D_{[4]}\) is given in Table 1.

Proof of Lemma 11. Equation (?? ) follows from Schur’s lemma and Eq. (38 ); it was essentially proved in Ref. [51], though without the concept of stabilizer 2-Rényi entropy. The equality in Eq. (?? ) is a direct corollary of Eqs. (38 ) and (?? ). The first inequality in Eq. (?? ) follows from Eq. (36 ), and the second is trivial. ◻

The following lemma is a simple corollary of Lemma 11.

Lemma 12. Suppose \(\mathcal{E}=\{\ket{\phi_{i}}, w_i\}_i\) is a state ensemble on \(\mathcal{H}\), and \(\psi \in \mathcal{P}(\mathcal{H})\) is a pure state. Then \[\begin{align} \underset{U\sim \mathrm{Cl}(n)}{\mathbb{E}}\Phi_4\left(\mathcal{E},U\psi U^{\dagger}\right)& =\frac{\sum_i 2^{-M_{2}(\phi_i)-M_{2}(\psi)}w_i}{d^2\,D^{+}_{[4]}}+\frac{1}{D^{-}_{[4]}}\left(1-\frac{\sum_i 2^{-M_{2}(\phi_i)}w_i}{d}\right)\left(1-\frac{2^{-M_{2}(\psi)}}{d}\right)\leq \frac{5(d+3)}{4(d+4)D_{[4]}}\leq \frac{5}{4D_{[4]}}. \end{align}\]

5 Proofs of results on average-case optimal shadow estimation↩︎

In this section, we prove our main results on optimal shadow estimation in the average-case setting, including Theorem 3 and Proposition 3 in the main text as well as Proposition 6 in the End Matter.

5.1 Auxiliary results on shadow norms↩︎

Lemma 13. Suppose \(A\) is a positive semidefinite operator acting on \(\mathcal{H}\) with \(\|A\|_1 = a\) and \(\|A\|_2^2 = b\), where \(a\) and \(b\) are positive numbers satisfying \({a^2}/{d}\le b \le a^2\). Then \[\begin{align} \|A\|\leq\frac{a + \sqrt{(d-1)(db - a^2)}}{d}, \label{eq:MaxEigUB} \end{align}\qquad{(14)}\] and the inequality is saturated when all eigenvalues of \(A\), except for the largest one, are equal.

Proof of Lemma 13. Let \(\mu_1 \ge \mu_2 \ge \dots \ge \mu_d \ge 0\) be the eigenvalues of \(A\) in nonincreasing order; then \(\|A\|= \mu_1\). By definition, \[\begin{align} \sum_{i=1}^d \mu_i =\|A\|_1= a,\quad \quad \sum_{i=1}^d \mu_i^2 =\|A\|_2^2= b. \end{align}\] Applying the Cauchy–Schwarz inequality, we deduce that \[\begin{align} (a - \mu_1)^2=\left( \sum_{i=2}^d \mu_i \right)^2 \le (d-1) \sum_{i=2}^d \mu_i^2=(d-1)\left(b - \mu_1^2\right), \end{align}\] which simplifies to \[\begin{align} d\mu_1^2 - 2a\mu_1 + \left[a^2 - (d-1)b\right] &\le 0. \end{align}\] This quadratic inequality implies that \[\begin{align} \|A\|=\mu_1\le \frac{a + \sqrt{(d-1)(db - a^2)}}{d}, \end{align}\] which confirms Eq. (?? ). If \(\mu_2 = \mu_3 = \dots = \mu_d\), then all three inequalities above are saturated, which completes the proof of Lemma 13. ◻

Lemma 14. Suppose \(O\in \mathcal{L}^{\mathrm{H}}_0(\mathcal{H})\). Then \[\frac{1}{d}\|O\|_2^4 \leq \|O\|_4^4\leq \frac{d-1}{d}\|O\|_2^4. \label{eq:4NormO2Norm}\qquad{(15)}\]

Proof. The first inequality in Eq. (?? ) follows from the Cauchy–Schwarz inequality. To prove the second inequality, denote by \(O_+\) and \(O_-\) the positive and negative parts of \(O\), respectively. Then, by assumption, we have \[\mathop{\mathrm{tr}}(O_+)=\mathop{\mathrm{tr}}(O_-),\quad \frac{1}{d-1} \|O_+\|_2^2\leq \|O_-\|_2^2\leq (d-1)\|O_+\|_2^2.\] Without loss of generality, we can assume that \(\|O_-\|_2^2\leq \|O_+\|_2^2=1\); then \(\|O\|_2^2=\|O_+\|_2^2+\|O_-\|_2^2\geq d/(d-1)\). In addition, the eigenvalues of \(O\) are bounded in absolute value by 1, so that \[\|O\|_4^4=\|O_+\|_4^4+\|O_-\|_4^4\leq \|O_+\|_2^2+\|O_-\|_2^2= \|O\|_2^2 \leq \frac{d-1}{d}\|O\|_2^4,\] which confirms the second inequality in Eq. (?? ) and completes the proof of Lemma 14. ◻

Proposition 10. Suppose \(\mathcal{E}\) is a state \(2\)-design, \(O \in \mathcal{L}^{\mathrm{H}}_0(\mathcal{H})\), and \(\mathcal{U}\) is a unitary \(4\)-design. Then \[\begin{align} \left\|\Xi(O,\mathcal{U})\right\|_\mathcal{E}^2 \le \left(\frac{d+1}{d}+\sqrt{f(d,r)\bar{\Phi}_3(\mathcal{E})+g(d,r)}\mkern 2mu\right)\|O\|_2^2, \label{eq:AvgShNormTUB} \end{align}\qquad{(16)}\] where \[r=\frac{\|O\|_4^4}{\|O\|_2^4}, \quad f(d,r)=\frac{12(d+1)^2\left[(d^2+3d+3)+(d^2+d)r\right]}{d^2(d+2)^2(d+3)},\quad g(d,r)=\frac{4(d+1)^2\left[(d^2-3d)r-(4d+3)\right]}{d^2(d+2)(d+3)}. \label{eq:func}\qquad{(17)}\]

By definition in Eq. (?? ) and Lemma 14, it is straightforward to verify that \[\label{eq:func2} \frac{1}{d}\le r\le \frac{d-1}{d},\quad f(d,r) \le \frac{24}{d} - \frac{26}{d^2},\quad g(d,r) \le 4-\frac{g_d}{d},\quad g_d\mathrel{\vcenter{:}}=\begin{cases} 18 & d=2,\\ 22 & d=3,\\ 25 & d=4,5,\\ 29 &d\geq 6,\\ 30 &d\geq 7. \end{cases}\tag{39}\]

Proof of Proposition 10. Suppose the ensemble \(\mathcal{E}\) has the form \(\mathcal{E}=\{\ket{\phi_i},w_i\}_i\) and let \(A\mathrel{\vcenter{:}}=\sum_i w_i\left[\mathop{\mathrm{tr}}\left(\phi_i O\right)\right]^2 \phi_i\). Then \[\begin{align} \|A\|_1=\frac{1}{d(d+1)}\|O\|_2^2,\quad \|A\|_2^2=\sum_{i,j}w_iw_j\left[\mathop{\mathrm{tr}}(O\phi_{i}) \right]^2\left[\mathop{\mathrm{tr}}(O\phi_{j})\right]^2\mathop{\mathrm{tr}}(\phi_{i}\phi_{j}). \end{align}\] Using Lemma 13, we deduce that \[\begin{align} &\|O\|_\mathcal{E}^2\le (d+1)^2\sqrt{d(d-1) \sum_{i,j}w_iw_j\left[\mathop{\mathrm{tr}}(O\phi_{i}) \right]^2\left[\mathop{\mathrm{tr}}(O\phi_{j})\right]^2\mathop{\mathrm{tr}}(\phi_{i}\phi_{j})-\frac{d-1}{d^2(d+1)^2}\|O\|_2^4}+\frac{d+1}{d}\|O\|_2^2. \end{align}\] Applying Jensen’s inequality, we further derive an upper bound for \(\underset{U\sim\mathcal{U}}{\mathbb{E}}\left\|UOU^{\dagger}\right\|_\mathcal{E}^2\): \[\begin{align} \underset{U\sim\mathcal{U}}{\mathbb{E}}\left\|UOU^{\dagger}\right\|_\mathcal{E}^2&\le (d+1)^2\sqrt{d(d-1) \underset{U\sim\mathcal{U}}{\mathbb{E}}\sum_{i,j}w_iw_j\left[\mathop{\mathrm{tr}}(UOU^{\dagger}\phi_{i}) \right]^2\left[\mathop{\mathrm{tr}}(UOU^{\dagger}\phi_{j})\right]^2\mathop{\mathrm{tr}}(\phi_{i}\phi_{j})-\frac{d-1}{d^2(d+1)^2}\|O\|_2^4} \nonumber\\ &\,\hphantom{=}\,+\frac{d+1}{d}\|O\|_2^2. \label{eq:AvgShNormTUBProof1} \end{align}\tag{40}\] The expectation under the square root can be evaluated as follows: \[\begin{align} &\underset{U\sim\mathcal{U}}{\mathbb{E}} \sum_{i,j}w_iw_j\left[\mathop{\mathrm{tr}}(UOU^{\dagger}\phi_{i}) \right]^2\left[\mathop{\mathrm{tr}}(UOU^{\dagger}\phi_{j})\right]^2\mathop{\mathrm{tr}}(\phi_{i}\phi_{j})= \sum_{i,j}w_iw_j\mathop{\mathrm{tr}}\left[\underset{U\sim\mathcal{U}}{\mathbb{E}}U^{\otimes 4}O^{\otimes 4}U^{\dagger \otimes 4}\left( \phi_i^{\otimes 2}\otimes\phi_j^{\otimes 2}\right)\right]\mathop{\mathrm{tr}}(\phi_{i}\phi_{j})\nonumber\\ &=\sum_{i,j}\sum_{\lambda}\frac{w_iw_j}{d_\lambda D_\lambda}\mathop{\mathrm{tr}}\left( P_{\lambda} O^{\otimes 4}\right)\mathop{\mathrm{tr}}\left(P_\lambda\phi_i^{\otimes 2}\otimes\phi_j^{\otimes 2}\right)\mathop{\mathrm{tr}}(\phi_{i}\phi_{j})=h_1(d,O)\Phi_1+h_2(d,O)\Phi_2+h_3(d,O)\Phi_3, \label{eq:AvgNormTUBProof2} \end{align}\tag{41}\] where \(\lambda\) represents a nonincreasing partition of 4 into at most \(d\) parts (see SM Sec. 4) and \[\begin{align} h_1(d,O)&=\frac{(d^2+3d+6)\left[\mathop{\mathrm{tr}}(O^2)\right]^2-4d\mathop{\mathrm{tr}}(O^4)}{d^2(d-1)(d+1)(d+2)(d+3)},\\ h_2(d,O)&=\frac{-(12d+12)\left[\mathop{\mathrm{tr}}(O^2)\right]^2+(4d^2-4d)\mathop{\mathrm{tr}}(O^4)}{d^2(d-1)(d+1)(d+2)(d+3)},\\ h_3(d,O)&=\frac{(2d^2+6d+6)\left[\mathop{\mathrm{tr}}(O^2)\right]^2+(2d^2+2d)\mathop{\mathrm{tr}}(O^4)}{d^2(d-1)(d+1)(d+2)(d+3)}. \end{align}\] The second equality in Eq. (41 ) follows from Lemma 9, and the third equality is derived using Lemma 10.

Combining Eqs. (40 ) and (41 ), we deduce that \[\begin{align} \left\|\Xi(O,\mathcal{U})\right\|_\mathcal{E}^2 &\le\frac{d+1}{d}\|O\|_2^2 +\sqrt{d(d+1)^4(d-1)\left[h_1(d,O)\Phi_1+h_2(d,O)\Phi_2+h_3(d,O)\Phi_3\right]-\frac{(d+1)^2(d-1)}{d^2}\|O\|_2^4}\nonumber\\ &=\left(\frac{d+1}{d}+\sqrt{f(d,r)\bar{\Phi}_3(\mathcal{E})+g(d,r)}\mkern 2mu\right)\|O\|_2^2, \label{eq:avg95G95up} \end{align}\tag{42}\] where \(\Phi_j\) is shorthand for the frame potential \(\Phi_j(\mathcal{E})\) for \(j=1, 2, 3\), and the equality holds because \(\Phi_1=1/d\) and \(\Phi_2=2/[d(d+1)]\) (since \(\mathcal{E}\) is a state 2-design) together with the definition \(r=\|O\|_4^4/\|O\|_2^4\). This completes the proof of Proposition 10. ◻

5.2 Proof of Theorem 3↩︎

Proof of Theorem 3. By virtue of Proposition 10 and the inequalities in Eq. (39 ), we deduce that \[\begin{align} \left\|\Xi(O,\mathcal{U})\right\|_\mathcal{E}^2 &\le \left(\frac{d+1}{d} + \sqrt{\left(\frac{24}{d}-\frac{26}{d^2}\right)\bar{\Phi}_3(\mathcal{E}) + 4 - \frac{g_d}{d}}\mkern 2mu\right)\|O\|_2^2\le \left(\frac{d+1}{d} + \sqrt{\frac{24}{d}\bar{\Phi}_3(\mathcal{E}) + 4 - \frac{26}{d^2} -\frac{g_d}{d}}\mkern 2mu\right)\|O\|_2^2\nonumber\\ &\le \left(\frac{d+1}{d} + \sqrt{\frac{24}{d}\bar{\Phi}_3(\mathcal{E}) + 4 - \frac{30}{d}}\mkern 2mu\right)\|O\|_2^2\leq \left(1 + \sqrt{\frac{24}{d}[\bar{\Phi}_3(\mathcal{E})-1] + 4 }\mkern 2mu\right)\|O\|_2^2\nonumber\\ &\le \left(2\sqrt{2}+1\right)\|O\|_2^2. \end{align}\] Here the second inequality holds because \(\bar{\Phi}_3(\mathcal{E})\ge 1\), the third holds because \((26/d)+g_d\geq 30\) by Eq. (39 ), the last holds because \(\bar{\Phi}_3(\mathcal{E}) \le (d + 2)(d^2 + 2d - 1)/(6d^2)\) by Proposition 1, which implies that \(24[\bar{\Phi}_3(\mathcal{E})-1]/d\leq 4\), and the fourth inequality follows from the concavity of the square root: \[\begin{align} \sqrt{\frac{24}{d}\bar{\Phi}_3(\mathcal{E}) + 4 - \frac{30}{d}}\leq \sqrt{\frac{24}{d}[\bar{\Phi}_3(\mathcal{E})-1] + 4 }-\frac{3}{d\sqrt{\frac{24}{d}[\bar{\Phi}_3(\mathcal{E})-1] + 4 }} \leq \sqrt{\frac{24}{d}[\bar{\Phi}_3(\mathcal{E})-1] + 4 }-\frac{1}{d}. \end{align}\] This completes the proof of Theorem 3. ◻

5.3 Auxiliary results on shadow norms for fidelity estimation↩︎

Suppose \(\mathcal{E}=\{\ket{\phi_i},w_i\}_i\) is a state 2-design on \(\mathcal{H}\), \(\psi \in \mathcal{P}(\mathcal{H})\), and \(\psi_0=\psi-{\mathbb{1}}/{d}\). To better understand the properties of the squared shadow norm \(\|\psi_0\|_\mathcal{E}^2\) tied to fidelity estimation, we introduce some auxiliary results. Define \[J(\mathcal{E},\psi)\mathrel{\vcenter{:}}= d(d + 1)^2\sum_iw_i\left[\mathop{\mathrm{tr}}\left(\phi_i\psi\right)\right]^2\phi_{i},\quad \eta(\mathcal{E},\psi)\mathrel{\vcenter{:}}=\left\|J(\mathcal{E},\psi)\right\|. \label{eq:JN}\tag{43}\]

Proposition 11. Suppose \(\mathcal{E}\) forms a state 2-design on \(\mathcal{H}\) and \(\psi\in\mathcal{P}(\mathcal{H})\). Then \[\begin{align} \eta(\mathcal{E},\psi)-3-\frac{2}{d}\le\left\|\psi_0\right\|_\mathcal{E}^2=\left\|J(\mathcal{E},\psi)-\frac{2(d+1)}{d}\psi\right\|-1+\frac{1}{d^2}\le \eta(\mathcal{E},\psi)-1+\frac{1}{d^2}.\label{eq:w} \end{align}\qquad{(18)}\]

Proposition 12. Suppose \(\mathcal{E}\) forms a state 2-design on \(\mathcal{H}\) and \(\psi \in \mathcal{P}(\mathcal{H})\). Then \[\begin{align} \dfrac{6(d+1)}{d+2}\bar{\Phi}_3(\mathcal{E},\psi)\leq \eta(\mathcal{E},\psi)&\le \mkern 2mu\dfrac{d+1}{\sqrt{(d+2)(d+3)}}\sqrt{48\bar{\Phi}_4(\mathcal{E},\psi)}\leq \sqrt{48\bar{\Phi}_4(\mathcal{E},\psi)},\label{eq:QnormLUB}\\ \dfrac{6(d+1)}{d+2}\bar{\Phi}_3(\mathcal{E},\psi)-3-\frac{2}{d}\leq \left\|\psi_0\right\|_\mathcal{E}^2&\le \dfrac{d+1}{\sqrt{(d+2)(d+3)}}\sqrt{48\bar{\Phi}_4(\mathcal{E},\psi)}-1+\frac{1}{d^2}\leq \sqrt{48\bar{\Phi}_4(\mathcal{E},\psi)}-1.\label{eq:EpsiShNormLUB} \end{align}\] {#eq: sublabel=eq:eq:QnormLUB,eq:eq:EpsiShNormLUB}

Proof of Proposition 11. By assumption, \(\mathcal{E}\) is a state 2-design, so \(\mathcal{M}_\mathcal{E}^{-1}(O)=(d+1)O\) whenever \(O\) is traceless. By virtue of 25 with \(O=\psi_0=\psi-\mathbb{1}/d\), we deduce that \[\begin{align} \left\|\psi_0\right\|_\mathcal{E}^2&=\left\|d(d+1)^2\sum_iw_i\left[\mathop{\mathrm{tr}}\left ( \phi_{i}\psi_0\right )\right]^2\phi_{i}\right\| \nonumber\\&=\left\|d(d+1)^2\sum_iw_i\left[\mathop{\mathrm{tr}}\left ( \phi_i\psi\right )\right]^2\phi_i-2(d+1)^2\sum_iw_i\mathop{\mathrm{tr}}\left ( \phi_i\psi\right )\phi_i+\dfrac{(d+1)^2}{d}\sum_iw_i\phi_i\right\| \nonumber\\ &=\left\|J(\mathcal{E},\psi)-\frac{2(d+1)}{d}\psi-\left(1-\frac{1}{d^2}\right)\mathbb{1}\right\|=\left\|J(\mathcal{E},\psi)-\frac{2(d+1)}{d}\psi\right\|-1+\frac{1}{d^2}, \label{eq:CRFE95norm95fidelity} \end{align}\tag{44}\] which confirms the equality in Eq. (?? ). Here the third equality holds because \(\sum_iw_i\phi_i={\mathbb{1}}/{d}\) and \(\sum_iw_i\mathop{\mathrm{tr}}\left(\phi_i\psi\right)\phi_i={(\mathbb{1}+\psi)}/{[d(d+1)]}\), given that \(\mathcal{E}\) is a state 2-design. In addition, the operator \(J(\mathcal{E},\psi)-2(d+1)\psi/d\) is positive semidefinite, which implies that \[\eta(\mathcal{E},\psi)- \frac{2(d+1)}{d} \leq \left\|J(\mathcal{E},\psi)-\frac{2(d+1)}{d}\psi\right\|\leq \eta(\mathcal{E},\psi).\] The above two equations together confirm Eq. (?? ) and complete the proof of Proposition 11. ◻

Proof of Proposition 12. The first inequality in Eq. (?? ) can be proved as follows: \[\begin{align} \eta(\mathcal{E},\psi)&=\max_{\rho}{d(d+1)^2}\sum_iw_i\mathop{\mathrm{tr}}(\phi_i\rho)\left[\mathop{\mathrm{tr}}(\phi_i \psi)\right]^2\ge {d(d+1)^2}\sum_iw_i\left[\mathop{\mathrm{tr}}(\phi_i \psi)\right]^3\nonumber\\ &=d(d+1)^2\Phi_3(\mathcal{E},\psi)=\dfrac{6(d+1)}{d+2}\bar{\Phi}_3(\mathcal{E},\psi), \label{eq:w95lower} \end{align}\tag{45}\] where the three equalities hold by definition. By the Cauchy–Schwarz inequality, we further deduce that \[\begin{align} \eta(\mathcal{E},\psi)&=\max_{\rho}{d(d+1)^2}\sum_iw_i\mathop{\mathrm{tr}}(\phi_i\rho)\left[\mathop{\mathrm{tr}}(\phi_i \psi)\right]^2 \le \max_{\rho}d(d+1)^2\sqrt{\Phi_2(\mathcal{E},\rho)}\sqrt{\Phi_4(\mathcal{E},\psi)}\nonumber\\ &=\sqrt{2}\mkern 2mud^{1/2}(d+1)^{3/2}\sqrt{\Phi_4(\mathcal{E},\psi)}=\dfrac{(d+1)}{\sqrt{(d+2)(d+3)}}\sqrt{48\bar{\Phi}_4(\mathcal{E},\psi)}\le\sqrt{48\bar{\Phi}_4(\mathcal{E},\psi)}, \label{eq:w95upper} \end{align}\tag{46}\] where the second equality holds because \(\max_\rho \Phi_2(\mathcal{E},\rho)= 2/[d(d+1)]\), given that \(\mathcal{E}\) forms a state 2-design by assumption, and the third equality holds by definition. The above two equations together confirm Eq. (?? ).

Next, the first two inequalities in Eq. (?? ) follow from Proposition 11 and Eq. (?? ). The last inequality in Eq. (?? ) is a simple corollary of the following inequalities: \[\begin{align} &\left(1- \dfrac{d+1}{\sqrt{(d+2)(d+3)}}\right)\sqrt{48\bar{\Phi}_4(\mathcal{E},\psi)}\ge \dfrac{1}{d+2}\sqrt{48\bar{\Phi}_4(\mathcal{E},\psi)}\ge\dfrac{4}{d+2}> \dfrac{1}{d^2}, \end{align}\] where the second inequality holds because \(\bar{\Phi}_4(\mathcal{E},\psi)\geq 1/3\) by Lemma 3. This completes the proof of Proposition 12. ◻

5.4 Proofs of Proposition 3 and Proposition 6↩︎

Proof of Proposition 3. Equation (?? ) is a simple corollary of Theorem 3 given that \(\|\psi_0\|_2^2=(d-1)/d\). To prove Eq. (?? ), note that \(\left\|\psi_0\right\|_\mathcal{E}^2\le\sqrt{48\bar{\Phi}_4(\mathcal{E},\psi)}-1\) for all \(\psi\in \mathcal{P}(\mathcal{H})\) by Proposition 12. If \(\left\|\psi_0 \right\|_\mathcal{E}^2\ge \sqrt{48[1+k\xi(\mathcal{E})]}-1\) for some \(k>0\), then \[\begin{align} \bar{\Phi}_4(\mathcal{E},\psi)-\mathbb{E}\bar{\Phi}_4(\mathcal{E},\psi) = \bar{\Phi}_4(\mathcal{E},\psi)-1\geq k \xi(\mathcal{E})>k \sqrt{(d+6)^3\mathrm{Var}(\bar{\Phi}_4(\mathcal{E},\psi))}, \end{align}\] where the last inequality follows from Lemma 5. Therefore, by Chebyshev’s inequality, we have \[\begin{align} \Pr \left\{ \left\|\psi_0 \right\|_\mathcal{E}^2 \ge \sqrt{48[1+k\xi(\mathcal{E})]}-1\right\} \le \Pr \left\{\bar{\Phi}_4(\mathcal{E},\psi)-1\ge k \xi(\mathcal{E})\right\} \le\frac{1}{(d+6)^3k^2}\le \frac{1}{d^3k^2}\quad \forall k>0, \end{align}\] which confirms Eq. (?? ). Finally, the inequality in Eq. (?? ) holds because \(\bar{\Phi}_3(\mathcal{E}) \le (d + 2)(d^2 + 2d - 1)/(6d^2)\) by Proposition 1, which completes the proof of Proposition 3. ◻

Proof of Proposition 6. To prove Eq. (12 ) in Proposition 6, without loss of generality, we can assume that \(\psi\sim\mathcal{T}\), where \(\mathcal{T}\) is the URP ensemble with respect to the computational basis. By Proposition 12 and Lemma 6, we can deduce that \[\begin{align} \underset{\psi\sim\mathcal{T}}{\mathbb{E}}\left\|\psi_0\right\|_\mathcal{E}^2\le \underset{\psi\sim\mathcal{T}}{\mathbb{E}}\sqrt{48\bar{\Phi}_4(\mathcal{E},\psi)}-1\le \sqrt{\underset{\psi\sim\mathcal{T}}{\mathbb{E}}48\bar{\Phi}_4(\mathcal{E},\psi)}-1\le24\sqrt{\frac{2D_{[4]}}{d^4}}=\sqrt{\frac{48(d+1)(d+2)(d+3)}{d^3}}, \end{align}\] which confirms Eq. (12 ).

Next, we turn to Eq. (13 ), assuming that \(d=2^n\). By Proposition 12 and Lemma 12, we can deduce that \[\begin{align} \underset{U\sim \mathrm{Cl}(n)}{ \mathbb{E}}\left\|U\psi_0U^\dagger\right\|_\mathcal{E}^2&\le \sqrt{2}\mkern 2mud^{1/2}(d+1)^{3/2}\underset{U\sim \mathrm{Cl}(n)}{\mathbb{E}}\sqrt{\Phi_4(\mathcal{E},U\psi U^{\dagger})}-1+\frac{1}{d^2}\nonumber\\&\le \sqrt{2}\mkern 2mud^{1/2}(d+1)^{3/2}\sqrt{\underset{U\sim \mathrm{Cl}(n)}{\mathbb{E}}\Phi_4(\mathcal{E},U\psi U^{\dagger})}-1+\frac{1}{d^2} \nonumber\\ &\leq \sqrt{2}\mkern 2mud^{1/2}(d+1)^{3/2}\sqrt{\frac{5(d+3)}{4(d+4)D_{[4]}}}-1+\frac{1}{d^2}\nonumber\\ &=\sqrt{\frac{60(d+1)^2}{(d+2)(d+4)}}-1+\frac{1}{d^2}\le \frac{2\sqrt{15}\mkern 2mu(d+1)}{d+2}-1+\frac{1}{d^2}\le 2\sqrt{15}-1, \end{align}\] which confirms Eq. (13 ) and completes the proof of Proposition 3. ◻

6 Further numerical results on mean squared shadow norms↩︎

In this appendix, we present additional numerical results on the mean squared shadow norms for several families of observables and state ensembles.

6.1 Shadow norms for fidelity estimation of Haar-random states↩︎

Figure 4: Mean squared shadow norms \mathbb{E}_{\psi\sim\mathrm{Haar}}\|\psi_0\|_{\mathcal{E}}^2 for fidelity estimation of Haar-random pure states with SIC-POVMs (blue) and MUBs (red). Error bars denote the standard deviation over 20 000 random pure states.

Here we provide additional results on shadow norms for fidelity estimation based on SIC-POVMs and complete sets of MUBs (available for prime-power dimensions). As in the main text, SIC-POVMs are generated by the Heisenberg–Weyl group from fiducial states listed in Ref. [57]; complete sets of MUBs are constructed according to Refs. [25], [27]. To complement Fig. 2 in the main text, Fig. 4 provides a refined comparison between SIC-POVMs and MUBs for Haar-random pure states, with error bars quantifying statistical fluctuations. Both measurement schemes exhibit convergence of the squared shadow norms toward their ensemble averages with increasing \(d\), indicating suppressed variability in higher dimensions. MUBs yield slightly narrower error bars at small \(d\), while the dispersion becomes comparable to that of SIC-POVMs at large \(d\), implying asymptotically similar performance.

6.2 Shadow norms for fidelity estimation of states in Clifford orbits↩︎

Next, we consider fidelity estimation of states in orbits of the \(n\)-qubit Clifford group using SIC-POVMs. Figure 5 (a) displays the ensemble averages \(\mathbb{E}_{U\sim\mathrm{Cl}(n)}\bigl\|U\psi_0 U^{\dagger}\bigr\|_{\mathcal{E}}^2\) for the stabilizer orbit and a random Clifford orbit. The mean values for the two orbits nearly coincide, especially for \(n\geq 3\). SIC-POVMs are generated by the Heisenberg–Weyl group from fiducial states listed in Ref. [57]; for the special case \(n=7\) (\(d=128\)), in which no fiducial is tabulated in that reference, we adopt the online database linked in Ref. [43].

a

b

Figure 5: Mean squared shadow norm \(\mathbb{E}_{\psi\sim\mathcal{T}}\bigl\|\psi_0\bigr\|_{\mathcal{E}}^2\) over the URP ensemble for fidelity estimation with SIC-POVMs and MUBs. The result achieved by an exact 3-design is also shown as a benchmark.. a — Mean squared shadow norm \(\mathbb{E}_{U\sim\mathrm{Cl}(n)}\bigl\|U\psi_0 U^{\dagger}\bigr\|_{\mathcal{E}}^2\) for fidelity estimation with SIC-POVMs. Results on the stabilizer orbit (blue circles) and a random Clifford orbit (red squares) nearly coincide, especially for \(n\geq 5\). For each \(n\), 2000 random Clifford unitaries are sampled.

6.3 Shadow norms for fidelity estimation of uniform random phase states↩︎

Next, we consider the ensemble \(\mathcal{T}\) of uniform random phase (URP) states defined over the computational basis. Figure 5 displays the mean squared shadow norms \(\mathbb{E}_{\psi\sim\mathcal{T}}\bigl\|\psi_0\bigr\|_{\mathcal{E}}^2\) for both SIC-POVMs and MUBs (one of the bases coincides with the computational basis). Both schemes exhibit similar scaling behavior, with SIC-POVMs maintaining a slight advantage across all dimensions shown except for \(d=2\).

6.4 Shadow norms of observables from the Gaussian unitary ensemble↩︎

a

b

Figure 6: Mean squared shadow norms over normalized traceless GUE observables versus dimension \(d\). Scatter points show numerical results over 2000 GUE samples for SIC-POVMs and MUBs (the latter restricted to prime-power dimensions). Lines with markers show analytical upper bounds from Proposition 13 with \(r=2/d\) for the worst-case 2-design (Proposition 1), SIC-POVMs, and MUBs.. a — Comparison of \(\mathbb{E}\bigl[\|O_0\|_4^4/\|O_0\|_2^4\bigr]\), \(\mathbb{E}\mkern 2mu\|O_0\|_4^4\big/\bigl(\mathbb{E}\mkern 2mu\|O_0\|_2^2\bigr)^2\), and \(\mathbb{E}\mkern 2mu\|O_0\|_4^4\big/\mathbb{E}\mkern 2mu\|O_0\|_2^4\) for GUE observables. For each dimension, 2000 observables are sampled. The curve \(2/d\) is shown as a benchmark.

Finally, we extend the analysis beyond pure-state observables to Hermitian operators sampled from the Gaussian unitary ensemble (GUE). For a state 2-design \(\mathcal{E}\), we introduce the normalized mean squared shadow norm \(\mathbb{E}_{O\sim\mathrm{GUE}}\bigl[\|O_0\|_{\mathcal{E}}^2/\|O_0\|_2^2\bigr]\), where \(O_0 = O - \mathop{\mathrm{tr}}(O)\,\mathbb{1}/d\) and the expectation is taken over the GUE. The following result parallels Proposition 10.

Proposition 13. Suppose \(\mathcal{E}\) is a state \(2\)-design on \(\mathcal{H}\). Then \[\begin{align} \label{eq:avg95norm95Gauss95bound} \mathbb{E}_{O\sim\mathrm{GUE}}\frac{\|O_0\|_{\mathcal{E}}^2}{\|O_0\|_2^2} \le \frac{d+1}{d} + \sqrt{f(d,r)\,\bar{\Phi}_3(\mathcal{E}) + g(d,r)}\,, \end{align}\qquad{(19)}\] where the expectation is well defined since \(\|O_0\|_2 > 0\) almost surely under the GUE, \[\begin{align} \label{eq:r95def} r \mathrel{\vcenter{:}}= \mathbb{E}_{O\sim\mathrm{GUE}}\frac{\|O_0\|_4^{4}}{\|O_0\|_2^{4}}\,, \end{align}\qquad{(20)}\] and \(f(d,r)\), \(g(d,r)\) are defined in Eq. (?? ) of Proposition 10.

Numerical results in Fig. 6 (a) suggest that the mean moment ratio \(r\) converges to \(2/d\) as \(d\) increases, with negligible deviation for \(d\geq 10\). Figure 6 presents mean squared shadow norms over normalized traceless GUE observables for SIC-POVMs and MUBs. As benchmarks, the figure also displays the analytical upper bounds from Proposition 13 evaluated at \(r=2/d\) for the worse 2-design (Proposition 1) as well as SIC-POVMs and MUBs.

References↩︎

[1]
H.-Y. Huang, R. Kueng, and J. Preskill, titlePredicting many properties of a quantum system from very few measurements, https://doi.org/10.1038/s41567-020-0932-7 journal Nat. Phys. 16, 1050 (2020).
[2]
A.  Elben, B. Vermersch, R. van Bijnen, author C. Kokail, T. Brydges, C. Maier, M. K. Joshi, R.  Blatt, C. F. Roos, and P. Zoller, Cross-platform verification of intermediate scale quantum devices, https://doi.org/10.1103/PhysRevLett.124.010504 Phys. Rev. Lett. 124, 010504 (2020)NoStop.
[3]
Z.  Yang, D. Chen, author Z. Li, and H. Zhu, High-precision fidelity estimation with common randomized measurements(2025), https://arxiv.org/abs/2511.22509arXiv:2511.22509 [quant-ph].
[4]
J.  Eisert, D. Hangleiter, N. Walk, I. Roth, D. Markham, R. Parekh, U. Chabaud, and E. Kashefi, Quantum certification and benchmarking, https://doi.org/10.1038/s42254-020-0186-4 journal Nat. Rev. Phys. 2, pages 382 (2020).
[5]
M.  Kliesch and I. Roth, Theory of quantum system certification, https://doi.org/10.1103/PRXQuantum.2.010201PRX Quantum volume 2, 010201 ( 2021).
[6]
T.  Brydges, A. Elben, P. Jurcevic, author B. Vermersch, C. Maier, B. P. Lanyon, P. Zoller, R. Blatt, and C. F. Roos, Probing Rényi entanglement entropy via randomized measurements, https://doi.org/10.1126/science.aau4963 journal Science 364, 260 (2019).
[7]
A.  Elben, R. Kueng, H.-Y. R. Huang, author R. van Bijnen, C. Kokail, M. Dalmonte, P. Calabrese, B. Kraus, J. Preskill, P. Zoller, and B. Vermersch, Mixed-state entanglement from local randomized measurements, https://doi.org/10.1103/PhysRevLett.125.200501 Phys. Rev. Lett. 125, 200501 (2020)NoStop.
[8]
Y.  Zhou, P. Zeng, and Z. Liu, titleSingle-copies estimation of entanglement negativity, https://doi.org/10.1103/PhysRevLett.125.200502Phys. Rev. Lett. 125, 200502 ( 2020).
[9]
A.  Neven, J. Carrasco, V. Vitale, author C. Kokail, A. Elben, M. Dalmonte, P. Calabrese, P. Zoller, B. Vermersch, R. Kueng, and B. Kraus, Symmetry-resolved entanglement detection using partial transpose moments, https://doi.org/10.1038/s41534-021-00487-y.
[10]
S.  Liu, Q. He, author M. Huber, O. Gühne, and G. Vitagliano, Characterizing entanglement dimensionality from randomized measurements, https://doi.org/10.1103/PRXQuantum.4.020324PRX Quantum volume 4, 020324 ( 2023).
[11]
C.  Yi, X. Li, and H. Zhu, titleCertifying entanglement dimensionality by \(k\)-reduction moments, https://doi.org/10.1103/cc1n-gmj1PRX Quantum volume 7, 010356 ( 2026).
[12]
E.  Bairey, I. Arad, and N. H. Lindner, Learning a local Hamiltonian from local measurements, https://doi.org/10.1103/PhysRevLett.122.020504Phys. Rev. Lett. 122, 020504 ( 2019).
[13]
A.  Anshu, S. Arunachalam, T. Kuwahara, and M. Soleimanifar, Sample-efficient learning of quantum many-body systems, in https://doi.org/10.1109/FOCS46700.2020.000692020 IEEE 61st Annual Symposium on Foundations of Computer Science (FOCS)(2020) pp. 685–691.
[14]
C.  Hadfield, S. Bravyi, R. Raymond, and author A. Mezzacapo, Measurements of quantum Hamiltonians with locally-biased classical shadows, https://doi.org/10.1007/s00220-022-04343-8Commun. Math. Phys. 391, 951 ( 2022).
[15]
H.-Y. Huang, Y. Tong, D. Fang, and author Y. Su, Learning many-body Hamiltonians with Heisenberg-limited scaling, https://doi.org/10.1103/PhysRevLett.130.200403 Phys. Rev. Lett. 130, 200403 (2023).
[16]
G. I. Struchalin, Y. A. Zagorovskii, E. V. Kovlakov, S. S. Straupe, and S. P. Kulik, Experimental estimation of quantum state properties from classical shadows, https://doi.org/10.1103/PRXQuantum.2.010307 journal PRX Quantum 2, 010307 (2021).
[17]
T.  Zhang, J. Sun, author X.-X. Fang, X.-M. Zhang, X. Yuan, and H. Lu, title Experimental quantum state measurement with classical shadows, https://doi.org/10.1103/PhysRevLett.127.200501.
[18]
H.-Y. Huang, M. Broughton, J. Cotler, author S. Chen, J. Li, M. Mohseni, H. Neven, R. Babbush, R. Kueng, J. Preskill, and J. R. McClean, Quantum advantage in learning from experiments, https://doi.org/10.1126/science.abn7293 journal Science 376, 1182 (2022).
[19]
H.-Y. Hu, A. Gu, author S. Majumder, H. Ren, Y. Zhang, D. S. Wang, Y.-Z. You, Z. Minev, author S. F. Yelin, and author A. Seif, Demonstration of robust and efficient quantum property learning with shallow shadows, https://doi.org/10.1038/s41467-025-57349-w journal Nat. Commun. 16, pages 2943 (2025).
[20]
R.  Stricker, M. Meth, L. Postler, author C. Edmunds, C. Ferrie, R. Blatt, P. Schindler, T. Monz, R.  Kueng, and M.  Ringbauer, Experimental single-setting quantum state tomography, https://doi.org/10.1103/PRXQuantum.3.040310 journal PRX Quantum 3, 040310 (2022).
[21]
L.  Innocenti, S. Lorenzo, I. Palmisano, author F. Albarelli, A. Ferraro, M. Paternostro, and G. M. Palma, Shadow tomography on general measurement frames, https://doi.org/10.1103/PRXQuantum.4.040328 journal PRX Quantum 4, 040328 (2023).
[22]
H. C. Nguyen, J. L. Bönsel, J.  Steinberg, and O.  Gühne, Optimizing shadow tomography with generalized measurements, https://doi.org/10.1103/PhysRevLett.129.220502 Phys. Rev. Lett. 129, 220502 (2022).
[23]
H.  Zhu, Multiqubit Clifford groups are unitary 3-designs, https://doi.org/10.1103/PhysRevA.96.062336Phys. Rev. A volume 96, 062336 ( 2017).
[24]
Z.  Webb, The Clifford group forms a unitary 3-design, https://doi.org/10.5555/3179439.3179447Quantum Inf. Comput. 16, 1379 ( 2016).
[25]
T.  Durt, B.-G. Englert, I. Bengtsson, and K. Życzkowski, On mutually unbiased bases, https://doi.org/10.1142/S0219749910006502 journal Int. J. Quantum Inf. 8, pages 535 (2010).
[26]
I. D. Ivanovic, Geometrical description of quantal state determination, https://doi.org/10.1088/0305-4470/14/12/019 journal J. Phys. A 14, 3241 (1981).
[27]
W. K. Wootters and B. D. Fields, Optimal state-determination by mutually unbiased measurements, Annals of Physics 191, 363 ( 1989).
[28]
V.  Gonzalez Avella, J.  Czartowski, D.  Goyeneche, and K.  Życzkowski, Cyclic measurements and simplified quantum state tomography, https://doi.org/10.22331/q-2025-06-04-1763 journal Quantum 9, 1763 (2025).
[29]
A.  Elben, S. T. Flammia, H.-Y. Huang, author R. Kueng, J. Preskill, B. Vermersch, and P. Zoller, title The randomized measurement toolbox, https://doi.org/10.1038/s42254-022-00535-2 journal Nat. Rev. Phys. 5, pages 9 (2023).
[30]
H.-Y. Hu, S. Choi, and Y.-Z. You, titleClassical shadow tomography with locally scrambled quantum dynamics, https://doi.org/10.1103/PhysRevResearch.5.023027Phys. Rev. Res. 5, 023027 ( 2023).
[31]
K.  Bu, D. E. Koh, R. J. Garcia, and A. Jaffe, titleClassical shadows with Pauli-invariant unitary ensembles, https://doi.org/10.1038/s41534-023-00801-wNoStop.
[32]
H.-Y. Hu and Y.-Z. You, Hamiltonian-driven shadow tomography of quantum states, https://doi.org/10.1103/PhysRevResearch.4.013054 Phys. Rev. Res. 4, 013054 (2022).
[33]
Z.  Liu, Z. Hao, and H.-Y. Hu, titlePredicting arbitrary state properties from single Hamiltonian quench dynamics, https://doi.org/10.1103/PhysRevResearch.6.043118 Phys. Rev. Res. 6, 043118 (2024).
[34]
Q.  Zhang, Q. Liu, and Y. Zhou, titleMinimal-Clifford shadow estimation by mutually unbiased bases, https://doi.org/10.1103/PhysRevApplied.21.064001Phys. Rev. Appl. 21, 064001 ( 2024).
[35]
Y.  Wang and W. Cui, Classical shadow tomography with mutually unbiased bases, https://doi.org/10.1103/PhysRevA.109.062406 journal Phys. Rev. A 109, pages 062406 (2024).
[36]
G.  Park, Y. S. Teo, and H. Jeong, titleResource-efficient shadow tomography using equatorial stabilizer measurements, https://doi.org/10.1103/9pbp-jzr9 journal Phys. Rev. Res. 7, pages 033097 (2025).
[37]
Y.  Wu, C. Wang, author J. Yao, H. Zhai, Y.-Z. You, and P. Zhang, Contractive unitary and classical shadow tomography, https://doi.org/10.1038/s41534-026-01227-w journal npj Quantum Inf. 12, pages 86 (2026).
[38]
M.  Ippoliti, Classical shadows based on locally-entangled measurements, https://doi.org/10.22331/q-2024-03-21-1293 journal Quantum 8, 1293 (2024).
[39]
M.  West, F. Sauvage, A. Sen, R. Forestano, D. Wierichs, N. Killoran, D. Grinko, M. Cerezo, and M. Larocca, Classical shadows with arbitrary group representations(2026), https://arxiv.org/abs/2604.01429.
[40]
R.  Cleve, D. Leung, L. Liu, and author C. Wang, Near-linear constructions of exact unitary 2-designs, https://doi.org/10.26421/QIC16.9-10-1NoStop.
[41]
C.  Bertoni, J. Haferkamp, M. Hinsche, author M. Ioannou, J. Eisert, and H. Pashayan, title Shallow shadows: Expectation estimation using low-depth random Clifford circuits, https://doi.org/10.1103/PhysRevLett.133.020602 Phys. Rev. Lett. 133, 020602 (2024).
[42]
T.  Schuster, J.  Haferkamp, and H.-Y. Huang, Random unitaries in extremely low depth, https://doi.org/10.1126/science.adv8590Science volume 389, 92 (2025)NoStop.
[43]
C. A. Fuchs, M. C. Hoang, and B. C. Stacey, The SIC question: History and state of play, https://doi.org/10.3390/axioms6030021.
[44]
G.  Zauner, Quantum designs: Foundations of a noncommutative design theory, https://doi.org/10.1142/S0219749911006776 journal Int. J. Quantum Inf. 9, pages 445 (2011).
[45]
J. M. Renes, R. Blume-Kohout, A. J. Scott, and C. M. Caves, titleSymmetric informationally complete quantum measurements, https://doi.org/10.1063/1.1737053NoStop.
[46]
A. J. Scott, Tight informationally complete quantum measurements, https://doi.org/10.1088/0305-4470/39/43/009 journal J. Phys. A 39, 13507 (2006).
[47]
A.  Ambainis and J.  Emerson, Quantum \(t\)-designs: \(t\)-wise independence in the quantum world, in https://doi.org/10.1109/CCC.2007.26Twenty-Second Annual IEEE Conference on Computational Complexity (CCC’07)(2007) pp. 129–140.
[48]
R.  Kueng and D. Gross, Qubit stabilizer states are complex projective 3-designs(2015), https://arxiv.org/abs/1510.02767arXiv:1510.02767 [quant-ph]NoStop.
[49]
See Supplemental Material for proofs and additional results.
[50]
S. G. Hoggar, \(t\)-designs in projective spaces, https://doi.org/10.1016/S0195-6698(82)80035-8Eur. J. Comb. 3, 233 ( 1982).
[51]
H.  Zhu, R. Kueng, author M. Grassl, and D. Gross, The Clifford group fails gracefully to be a unitary 4-design(year2016), https://arxiv.org/abs/1609.08172.
[52]
J. T. Iosue, T. C. Mooney, A. Ehrenberg, and A. V. Gorshkov, Projective toric designs, quantum state designs, and mutually unbiased bases, https://doi.org/10.22331/q-2024-12-03-1546 journal Quantum 8, 1546 (2024).
[53]
D.  Chen and H. Zhu, Nonstabilizerness enhances thrifty shadow estimation(2024), https://arxiv.org/abs/2410.23977arXiv:2410.23977 [quant-ph]NoStop.
[54]
J.  Jasper and D. G. Mixon, Nearly tight weighted 2-designs in complex projective spaces of every dimension, in https://doi.org/10.1109/SampTA64769.2025.11133524 booktitle 2025 International Conference on Sampling Theory and Applications (SampTA)(2025) pp. 1–5.
[55]
A.  Roy and A. Scott, Weighted complex projective 2-designs from bases: optimal state determination by orthogonal measurements, https://doi.org/10.1063/1.2748617 journal J. Math. Phys. 48, pages 072110 (2007).
[56]
G.  McConnell and D.  Gross, Efficient 2-designs from bases exist, https://doi.org/10.26421/QIC8.8-9-4NoStop.
[57]
A. J. Scott and M. Grassl, Symmetric informationally complete positive-operator-valued measures: A new computer study, https://doi.org/10.1063/1.3374022 journal J. Math. Phys. 51, pages 042203 (2010).
[58]
C.  Mao, C. Yi, and H. Zhu, titleQudit shadow estimation based on the Clifford group and the power of a single magic gate, https://doi.org/10.1103/PhysRevLett.134.160801 Phys. Rev. Lett. 134, 160801 (2025).
[59]
H.  Zhu, C. Mao, and C. Yi, Third moments of qudit Clifford orbits and 3-designs based on magic orbits(2024), https://arxiv.org/abs/2410.13575arXiv:2410.13575 [quant-ph]NoStop.
[60]
G. H. Hardy and E. M. Wright, An Introduction to the Theory of Numbers, 6th ed. (Oxford University Press, Oxford, 2008).
[61]
D. K. Mark, F. Surace, A. Elben, author A. L. Shaw, J. Choi, G. Refael, M. Endres, and S. Choi, Maximum entropy principle in deep thermalization and in hilbert-space ergodicity, https://doi.org/10.1103/PhysRevX.14.041051.
[62]
L.  Leone, S. F. E. Oliviero, and A.  Hamma, Stabilizer Rényi entropy, https://doi.org/10.1103/PhysRevLett.128.050402.