The limits of erasure-based postselection for quantum error mitigation


Abstract

In both classical and quantum error correction, heralded erasures are known to be easier to tolerate than unheralded general stochastic errors. Whilst an established benefit of loss-dominant quantum architectures such as photonic qubits, this fact has received renewed interest, with a pivot towards reconstructing other architectures to be erasure-dominant, such as dual-rail transmons. This work investigates exploiting these ‘erasure qubits’ in the near term by using postselection as a technique for error mitigation, wherein circuit shots detecting any erased qubits are discarded from the computational ensemble and repeated. Firstly, we outline a numerical framework for representing circuit-level erasure noise and present ‘erado’, an open-source library capable of simulating erasure noise and postselection. Secondly, we investigate the effects of both erasure noise and noise in the erasure checks themselves on the quantum Fourier transform (QFT), in the additional presence of gate depolarising noise. A worked example is provided of postselection fully mitigating against the erasure channel for erasure check error rates less than 3.0%. We also show how a postselected dual-rail system can surpass a fundamental noise floor at the kiloquop scale where a comparable single-rail system cannot, justifying this approach in the NISQ regime before (and, perhaps, combined with) the practical arrival of QEC.

1 Introduction↩︎

A central challenge in many quantum computing architectures – and especially in superconducting qubits – is the presence of energy relaxation processes, commonly characterised by the \(T_1\) lifetime [1]. In conventional architectures, this decay mechanism manifests as an unheralded amplitude damping channel, introducing stochastic errors that directly contribute to logical failure rates [2]; the error mechanism is unheralded in that we cannot directly know if or where discrete stochastic errors are induced, requiring algorithmic inference via error-correcting codes [3][5]. This places stringent requirements on coherence, calibration, and ultimately the overhead required for quantum error correction (QEC).

In this work, we consider an approach based on erasure-aware qubit encodings, in which dominant relaxation events are heralded (that is, can be detected) and conditionally removed from the computational ensemble, i.e.postselection [6]. In a dual-rail encoding, two physical modes are used to represent a single logical qubit, enabling the detection of prevailing loss mechanisms through a single end-of-line measurement [7], [8]. This provides a practical route to mitigating the dominant source of noise in such systems, not only in their fault-tolerant future when combined with QEC, but also in the noisy intermediate-scale (NISQ) regime, wherein active QEC remains challenging for current hardware capabilities.

In Section 2, technical background is introduced: we formally define the dual-rail codespace, describe an erasure noise model appropriate for relaxation-dominant systems, explain the propagation of erasures through circuits, and analyse the role of imperfect erasure detection in postselection.

In Section 3, we overview our simulational techniques developed for benchmarking the effects of noise parameters upon the performance and cost of erasure-based postselection. We have published this codebase as erado, an open-source Qiskit-based [9] library providing tools for the simulation of erasure noise and postselection with arbitrary quantum circuits.4

In Section 4, we demonstrate these techniques in a series of results using the quantum Fourier transform (QFT) as a case study; in particular, we explore how the overhead of postselection scales with the size of the circuit, as well as how the mean fidelity with and without postselection is affected by erasure noise, erasure check noise and gate depolarising noise. Finally, we also show how postselected dual-rail systems can outperform conventional single-rail systems, accounting for qubit idling at the kiloquop scale.

2 Background↩︎

2.1 Erasure noise↩︎

In classical information theory, an error most conventionally takes the form of a bit flip at an unknown (or ‘unheralded’) location, where the binary symmetric channel is a noise model which flips each bit to the opposite state with uniform probability. A linear code with distance \(d\) can detect up to \(d-1\) flip errors, or correct up to \(\lfloor(d-1)/2\rfloor\) flip errors [3].5 In contrast, an erasure is the loss of state (i.e.neither 0 nor 1) at a known (or ‘heralded’) location, where the equivalent binary erasure channel erases each bit with uniform probability [10]. As the location is known, an appropriate code/decoder can correspondingly correct up to \(d-1\) erasures.

In quantum information theory, unheralded stochastic noise generalises to a statistical ensemble of Pauli operators acting on the logical subspace: \[\begin{align} \mathcal{P}(\rho) = &(1-p_X-p_Y-p_Z)\rho \\ &+ p_X X\rho X + p_Y Y\rho Y + p_Z Z\rho Z \;. \notag \end{align}\] A special case is the depolarising channel, wherein the probability associated with each Pauli operator is equal, such that the state \(\rho\) is mapped towards the maximally-mixed state [2]: \[\begin{align} \mathcal{D}(\rho) &= (1-p)\rho + \frac{p}{3}(X\rho X + Y\rho Y + Z\rho Z) \\ &= (1-p)\rho + p\frac{I}{2} \;. \end{align}\]

Correspondingly, an erasure is distinguished from Pauli noise by the fact that the location (and, often, the time) of the error is known; rather than requiring inference by decoding stabiliser measurements, erasure events provide classical information indicating that a qubit has left the computational subspace. Erasure-dominance can thus be a desirable system property, with highly-efficient codes [11], [12] and decoders [13], [14] designed for this noise channel. This has been considered an advantage of quantum architectures with natural erasure-dominance, such as dual-rail optical qubits, wherein photon loss is a major error mechanism [7].

2.2 Modelling dual-rail qubits↩︎

In conventional superconducting devices such as transmons, the primary error mechanisms are time-varying amplitude damping and phase damping, characterised by time constants \(T_1\) and \(T_2\) respectively, such that depolarisation is a dominant noise channel [1], [2]. Contrastingly, leakage into non-computational states is less statistically significant. Erasure qubits are instead engineered such that leakage events are both dominant and heralded [15]. For example, transmons can be used to construct the dual-rail encoded dimon qubit (DDQ) [8]; as a direct analogue of a dual-rail photonic qubit, it encodes one logical qubit into two physical modes \(\@ifnextchar\bra{{|{ab}\rangle}\!}{{|{ab}\rangle}}\), such that the computational basis can be defined as \[\@ifnextchar\bra{{|{0_L}\rangle}\!}{{|{0_L}\rangle}} = \@ifnextchar\bra{{|{10}\rangle}\!}{{|{10}\rangle}} \;, \; \@ifnextchar\bra{{|{1_L}\rangle}\!}{{|{1_L}\rangle}} = \@ifnextchar\bra{{|{01}\rangle}\!}{{|{01}\rangle}} \;,\] where \(\@ifnextchar\bra{{|{10}\rangle}\!}{{|{10}\rangle}}\) denotes a single excitation in mode \(a\) with none in mode \(b\), and vice versa. The logical codespace is therefore the single-excitation manifold \[C = \mathop{\mathrm{span}}\{ \@ifnextchar\bra{{|{10}\rangle}\!}{{|{10}\rangle}}, \@ifnextchar\bra{{|{01}\rangle}\!}{{|{01}\rangle}}\} \;.\]

A key feature of this encoding is that thermal (i.e.spin–lattice \(T_1\)) relaxation corresponds to loss of excitation, such that the energy transitions \[\@ifnextchar\bra{{|{10}\rangle}\!}{{|{10}\rangle}}\mapsto \@ifnextchar\bra{{|{00}\rangle}\!}{{|{00}\rangle}} \;, \; \@ifnextchar\bra{{|{01}\rangle}\!}{{|{01}\rangle}}\mapsto \@ifnextchar\bra{{|{00}\rangle}\!}{{|{00}\rangle}} \;,\] take the system out of the logical subspace \(C\). Thus, any decay into the noncomputational ground state \(\@ifnextchar\bra{{|{e}\rangle}\!}{{|{e}\rangle}}= \@ifnextchar\bra{{|{00}\rangle}\!}{{|{00}\rangle}}\) is an erasure event. This can be interpreted as a physical error-detecting code for the erasure channel.6

We can therefore define the erasure subspace as \[E = \mathop{\mathrm{span}}\{ \@ifnextchar\bra{{|{e}\rangle}\!}{{|{e}\rangle}}\} = \mathop{\mathrm{span}}\{ \@ifnextchar\bra{{|{00}\rangle}\!}{{|{00}\rangle}}\} \;,\] and note that thermal relaxation becomes a detectable leakage process \(C \to E\).

This erasure noise channel may similarly be written explicitly in an operator–sum form as \[\begin{align} \mathcal{E}(\rho) &= (1-p_e) \mathcal{E}_0 \rho \mathcal{E}_0^\dagger + p_e \mathcal{E}_1 \rho \mathcal{E}_1^\dagger \;, \\ \mathcal{E}_0 &= I \;, \\ \mathcal{E}_1 &= \@ifnextchar\bra{{|{e}\rangle}\!}{{|{e}\rangle}}\bra{10}+ \@ifnextchar\bra{{|{e}\rangle}\!}{{|{e}\rangle}}\bra{01} \;, \end{align}\] where the erasure rate \(p_e\) is the uniform probability with which each dual-rail qubit transitions into the erased state \(\@ifnextchar\bra{{|{e}\rangle}\!}{{|{e}\rangle}}\). Conditioned on no erasure, residual noise may still act as an effective Pauli channel within the logical codespace \(C\), such that a convenient overall phenomenological noise model takes the form \[\mathcal{N}(\rho) = (1-p_e) \mathcal{P}(\rho) + p_e \mathcal{E}_1 \rho \mathcal{E}_1^\dagger \;,\] capturing the physically-relevant regime in which erasures dominate but a background of Pauli-type errors remain.

A crucial distinction between erasure noise and Pauli noise is that erasures induce conditional circuit evolution; once a qubit has left the computational subspace, subsequent gates no longer correspond to meaningful logical operations. Take \(U_L\) to be any unitary matrix representing a coherent operation within the logical subspace \(C\) (i.e.a quantum circuit gate) such that \[\begin{align} &U_L : C \to C \;, \\ \text{i.e.} \;&U_L \@ifnextchar\bra{{|{\psi}\rangle}\!}{{|{\psi}\rangle}} = \@ifnextchar\bra{{|{\phi}\rangle}\!}{{|{\phi}\rangle}} \;, \; \@ifnextchar\bra{{|{\phi}\rangle}\!}{{|{\phi}\rangle}} \in C \quad \forall \@ifnextchar\bra{{|{\psi}\rangle}\!}{{|{\psi}\rangle}} \in C \;. \end{align}\] Then, in our model, we treat the dynamics of the system after an erasure out of \(C\) as \[U_L \@ifnextchar\bra{{|{e}\rangle}\!}{{|{e}\rangle}} = \@ifnextchar\bra{{|{e}\rangle}\!}{{|{e}\rangle}}\] such that the erased state \(\@ifnextchar\bra{{|{e}\rangle}\!}{{|{e}\rangle}}\) is stabilised by all logical operators \(U_L\); that is, all subsequent operations act trivially on the erased state. Operationally, this corresponds to replacing all circuit gates on an erased qubit with no-ops.

More generally, for an erasure qubit with arbitrarily-sized \(E\), \[U_L : E \to E \;,\] such that the orthogonal subspaces \(C\) and \(E\) are disjoint under logical operators.

This model captures the physical intuition that once the encoded excitation is lost, the logical information is irrecoverable – it cannot be restored by the remainder of the circuit. Importantly, this conditional structure must be handled carefully when performing circuit-level simulations or constructing detector–error models (DEMs) [16].

2.3 Postselection↩︎

Postselection is the reduction of a probability distribution conditioned on the outcome of some event(s); as a hypothetical computational feature, it has been shown to significantly improve the complexity space of nondeterministic algorithms – both classical and quantum – by conditioning the outcome ensemble upon successful or otherwise-desired states. More formally, the classical bounded-error probabilistic polynomial-time complexity class \(\mathsf{BPP}\) is improved to \(\mathsf{BPP_{path}}\) with postselection [17], and the equivalent quantum class \(\mathsf{BQP}\) improved to \(\mathsf{PostBQP}\) (which is equal to the large \(\mathsf{PP}\)) [6].

In this work, we investigate the practical application of this conceptual technique as error mitigation for the erasure channel, in which we simply reject circuit shots for which any erasure events were detected.

Let us consider the experimentally-motivated setting where erasure detection is performed exactly once at the end of circuit execution, a.k.a.end-of-line (EOL). In the ideal limit of perfect erasure detection, postselection removes all \(T_1\) events, yielding behaviour approaching an infinite-\(T_1\) device, up to re-excitations. A primary objective of this paper is to quantify how closely practical dual-rail systems can approach this limit.

In realistic devices, erasure detection is imperfect. We model errors on the end-of-line erasure checks (one bit per logical qubit) as a binary asymmetric channel with the following parameters:

  1. false positive (\(0 \mapsto 1\)) rate \(p_{\mathrm{fp}}\);

  2. false negative (\(1 \mapsto 0\)) rate \(p_{\mathrm{fn}}\).

False positives cause successful runs to be rejected, reducing sampling efficiency, whilst false negatives cause unsuccessful runs to be allowed into the accepted ensemble, limiting the overall improvement in logical fidelity. These effects therefore define a trade-off between improved error suppression and sampling overhead, with the utility of postselection depending critically on both the physical erasure rate and the fidelity of erasure checks.

The remainder of this paper develops quantitative models of this improvement, explores the role of imperfect erasure detection, and compares the resulting performance against the infinite-\(T_1\) limit.

3 Numerical techniques↩︎

a
b

Figure 1: Illustrations of the two equivalent erasure noise models implemented in erado with a toy circuit example.. a — Example shot with the erasure circuit sampler. An erasure is stochastically induced on \(q_0\) and \(q_1\) by the first \(\mathrm{CX}\) gate, so all subsequent gates touching these qubits are deleted from the circuit representation., b — The same base circuit (cropped for clarity) instead transformed by the erasure transpiler pass. An additional classical register stores an erasure flag for each data qubit. An auxiliary qubit \(q_\textrm{eraser}\) populates erasures independently onto these wires via an auxiliary bit-flip channel with rate \(p_e\) on its measurement operations. All gates on data qubits are then conditioned on all qubit arguments having unset erasure flags.

3.1 Circuit sampling/transpilation↩︎

In this section, we outline the techniques used in our open-source erasure–postselection codebase erado, which implements a circuit-level noise simulation capable of applying heralded erasure dynamics to arbitrary quantum circuits.

Two equivalent implementations of this model are provided, both of which are illustrated in Figure 1. The first of these (the ErasureCircuitSampler) enforces erasure noise by deleting gates affected by erasure events from the circuit (Figure 1 (a)). This approach directly models the conditional circuit evolution induced by loss: once a qubit is erased, subsequent operations acting on that qubit no longer implement meaningful dynamics within the computational subspace and are therefore treated as no-ops. Rather than inserting additional operations or markers into the circuit, the sampler generates a modified circuit for each shot by removing the gates that would act on erased degrees of freedom. The principal benefit is that no additional overhead is introduced into the circuit representation itself; the principal drawback is that circuits are sampled and simulated on a per-shot basis.

The second of these (the ErasurePass) instead implements a circuit transpilation pass which adds a classical register containing the Boolean erasure state of each qubit, and wraps all gates in classical logic such that they only execute conditionally on none of the qubit arguments being erased (Figure 1 (b)). The principal benefit is that this yields a single precomputed circuit which fully implements the noise model across all shots; the principal drawback is that the addition of classical logic significantly slows down the simulation of quantum circuits.

With either implementation, this model occupies a complementary regime to simulational techniques based on stabilisers and detector–error models (DEMs) [18]. It captures conditional dynamics intrinsic to erasure processes and supports arbitrary Qiskit circuits (and backends, in the case of the circuit sampler), at the cost of either per-shot circuit sampling or the introduction of classical control flow. This makes it well-suited to near-term studies of erasure-aware primitives, postselection strategies, and the impact of imperfect erasure checks; meanwhile, scalable DEM-based pipelines remain essential for large-distance QEC studies and broad architectural exploration.

In the remainder of this paper, of these two implementations, we refer to and use the ErasureCircuitSampler for numerically characterising erasure-induced behaviour in circuits (empirically, it outperformed the erasure transpiler pass in overall runtime on our high-performance computing platforms).

3.2 Sampling semantics↩︎

We assume a uniform erasure rate \(p_e\) applied across the circuit execution. The circuit sampler supports two closely-related semantics controlled by a Boolean option erasure_before_gates:

  1. post-gate erasure, where erasure is inflicted immediately after a gate acts;

  2. pre-gate erasure, where erasure is inflicted immediately before a gate acts.

These two conventions represent common physical interpretations (e.g.erasure associated with idle evolution, versus erasure associated with driven operations) and allow sensitivity studies to bracket hardware-dependent timing assumptions. In both cases, once an erasure event occurs on a qubit, all subsequent gates that touch that qubit are removed from the shot circuit. For multi-qubit gates, this deletion rule naturally removes any operation for which one participant has been erased.

Take a circuit represented as an ordered collection of gates \(G\) which, as a tuple, can be defined as a set of index–operator pairs \((t,U)\) in the form \[G = \{ (1,U_1) , (2,U_2) , \dots , (\abs{G},U_{\abs{G}}) \} \;.\] We denote the gate index with \(t\) as it is hereinafter used interchangeably with the more general concept of time, under our chosen circuit timing convention. Each element in \(G\) is therefore a candidate location for an erasure event. For each shot \(\sigma\), the sampler stochastically draws a vector of independent erasure flags \[\vec{e}_\sigma = \left( {e_\sigma}_t \mid t \in [1 \mathinner{\ldotp \ldotp}\abs{G}], {e_\sigma}_t \sim \mathrm{Bernoulli}(p_e) \right) \;,\] yielding a set of erasure locations for each shot \[\Omega_\sigma = \left\{ t \mid {e_\sigma}_t = 1 \right\} \;.\]

3.3 Deletion lookup table (LUT)↩︎

To avoid computing deletion sets repeatedly over many shots, the ErasureCircuitSampler precomputes a lookup table (LUT) at construction time. Given an input circuit, we enumerate the ordered gate representation \(G\) and, treating each element as a candidate erasure location, cache the set of downstream gates which would be deleted.

Formally, each erasure event at location \(t\) maps to a deletion set \[\tau(t) \subseteq [t \mathinner{\ldotp \ldotp}\abs{G}] \;,\] consisting of all gate indices \(t' \ge t\) where the support of \(U_{t'}\) includes any of the qubits in the support of the deleted gate \(U_t\) (if using post-gate erasure, the above subset and inequality become strict).

The lookup table therefore enables \(O(1)\) retrieval of the relevant deletion set for each sampled erasure event, reducing the problem of shot generation to set union and filtering. This design is particularly effective for parameter sweeps, since the lookup table depends only on circuit structure and erasure timing convention, and thus can be reused across many runs with different erasure rates, backend configurations, or analysis modes.

3.4 Shot generation and execution↩︎

For each shot \(\sigma\), an erasure-sampled circuit \(G_\sigma\) is constructed by deleting from \(G\) the union of deletion sets implied by the erasure locations \(\Omega_\sigma\); that is, the tuple of all deleted gate locations \(D_\sigma\) is \[D_\sigma = \left\{ (t', U_{t'}) \;\middle|\;t' \in \bigcup_{t\in\Omega_\sigma} \tau(t) \right\} \;,\] such that \[G_\sigma = G \setminus D_\sigma \;.\]

The ErasureCircuitSampler implementing erasure noise as gate deletion means that the resulting shot circuit \(G_\sigma\) contains only a subset of the original circuit primitives, introducing no additional operations, measurements or classical control flow (in contrast to the ErasurePass). In this sense, the method enforces conditional ‘no-op’ dynamics without modifying the gate set.

3.5 Randomness and concurrency↩︎

The ErasureCircuitSampler implements reproducible stochastic sampling through a multiprocessing-safe random number generator (RNG) interface (MultiprocessingRNG). Importantly, the seed method controls the entropy used to generate erasure patterns, and also correctly overrides the seed behaviour of the backend used for simulating circuit shots. This ensures that both erasure sampling and circuit execution are reproducible and stable under parallel execution, enabling deterministic reruns and statistically-valid comparisons across parameter sweeps.

Because shots are independent, this method is embarrassingly parallel; in practice, we distribute shots across worker processes. By aggregating desired results7 via multiprocessing-safe optional callbacks, we also avoid the need to store full shot-level traces (unless required for debugging or other detailed analysis).

Finally, since erasure-aware simulation can occasionally produce pathological cases for certain backends or compilation settings (e.g.extreme deletion patterns, backend-specific slowdowns or system-level multiprocessing limits), the ErasureCircuitSampler supports a per-shot timeout parameter, which upper-bounds the allowed execution time for a single shot before termination. This provides robust behaviour in large Monte Carlo simulations and prevents a small number of outliers from dominating total runtime.

4 Results↩︎

4.1 Model and metrics↩︎

a

b

Figure 2: Cost of postselection for QFT with 5000 target shots. The trends show erasure noise with ideal checks (green), erasure noise with check noise (orange), and erasure noise with check noise and gate depolarising noise (blue). Note how false positives dominate to strictly increase the rejection rate over ideal checks. Error bars for rejection rate show the \(95\%\) confidence interval, calculated via the Clopper-Pearson method for binomial proportions..

In these results, we study the effects of noise channels and postselection on the performance of a linear-connectivity quantum Fourier transform (QFT) construction, as given in [19]. Arbitrarily but consistently, all circuits are transpiled to the universal basis gate set \(\{\mathrm{R_Z}, \mathrm{\sqrt{X}}, \mathrm{CX}\}\). For a number of qubits \(n\in[2 \mathinner{\ldotp \ldotp}16]\), the number of gates \(\abs{G}\) in this QFT construction scales quadratically as \[\label{eq:qft-gates} \abs{G} = 3n^2-2n+2 \;.\tag{1}\]

The erasure rate \(p_e\) is the uniform probability with which any of the \(\abs{G}\) gates in the circuit inflicts an erasure upon any of its qubit arguments. This is equivalent to an amplitude damping channel with emission rate \(p_e\), where \(p_e\) models thermal relaxation by assuming the time-varying form \(1-e^{-t/T_1}\) [2]. For example, a gate time \(\require{physics} t=\qty{500}{\nano\second}\) and a physical relaxation time \(\require{physics} T_1=\qty{50}{\micro\second}\) equates to an erasure rate \(p_e \approx 0.010\).

Erasure check noise is modelled with a single error rate \(q\) for simplicity, encompassing both false positives and false negatives as \(q=p_\mathrm{fp}=p_\mathrm{fn}\). Unheralded Pauli noise is modelled as a depolarising channel with 1-qubit gate error rate \(p_\mathrm{1Q}=0.001\) and 2-qubit rate \(p_\mathrm{2Q}=0.010\).

When postselection is enabled, the simulation runs until the number of accepted shots \(s_a\) equals some target (e.g. shots). This leads to a number of rejected shots \(s_r\), with total shots \(s=s_a+s_r\). The rejection rate \(R\) is thus defined as the proportion of rejected shots: \[\label{eq:rejection-rate} R = \frac{s_r}{s} = 1-\frac{s_a}{s} \;.\tag{2}\] A shot is rejected if any qubits are erased, such that the theoretical rejection rate scales binomially as \[\label{eq:rejection-rate-trend} R = 1-(1-p_e)^{\abs{G}} \;.\tag{3}\] Correspondingly, we define the overhead of postselection as the proportion of total shots to accepted shots \(s/s_a\). The theoretical trend for the overhead is derived by combining Equations 2 and 3 : \[\begin{align} \label{eq:overhead} \frac{s}{s_a} &= (1-p_e)^{-\abs{G}} \\ &\in (1-p_e)^{-O(n^2)} \;. \end{align}\tag{4}\] Therefore, we can say that this postselection problem scales exponentially, in that its overhead scales with a polynomial exponent in \(n\) (as per the complexity class \(\mathsf{EXPTIME}\)). We verify these theoretical trends for the rejection rate and overhead empirically in Figure 2.

Finally, the per-shot circuit fidelity \(F\) is defined conventionally as \[F( \@ifnextchar\bra{{|{\psi}\rangle}\!}{{|{\psi}\rangle}}, \@ifnextchar\bra{{|{\Psi}\rangle}\!}{{|{\Psi}\rangle}}) = \abs{\braket{\psi | \Psi}}^2 \;,\] where \(\@ifnextchar\bra{{|{\psi}\rangle}\!}{{|{\psi}\rangle}}\) is the end-of-line statevector snapshot of the simulation (before final measurements) and \(\@ifnextchar\bra{{|{\Psi}\rangle}\!}{{|{\Psi}\rangle}}\) is the ideal statevector representation calculated directly from the circuit definition for each \(n\).

a
b
c

Figure 3: Mean and minimum circuit fidelity for QFT with 5000 target shots, with and without postselection, for \(p_e=0.005\) and \(q=0.010\). The \(T_1\to\infty\) reference is equivalent to a standalone depolarising model (i.e. \(p_e=q=0\)). Error bars for mean fidelity (here and hereinafter) show the \(95\%\) confidence interval, calculated via Student’s \(t\) distribution of the standard error of the mean (SEM).. a — Ideal checks and ideal circuit. Note how postselection definitionally yields \(F=1.0\)., b — Noisy checks and ideal circuit. Note the small-but-nonzero effect of false negatives upon mean fidelity., c — Noisy checks and noisy (depolarising) circuit. Note how postselection approximately achieves \(T_1\to\infty\), excepting false negatives.

4.2 Interaction of noise models↩︎

The impact of postselection on the fidelity \(F\) of the QFT algorithm is presented in Figure 3; in particular, we observe how erasure and postselection impacts the fidelity with respect to the limit of \(T_1\to\infty\), i.e.the regime where \(p_e=0\) and the remaining unheralded Pauli errors (i.e.our gate depolarising channel) alone constitute the noise model. The fidelity for this regime is included across all three subfigures – note that this is a deterministic reference series, calculable directly via a density-matrix representation of the circuit: \[\begin{align} \@ifnextchar\bra{{|{\Psi}\rangle}\!}{{|{\Psi}\rangle}} &= (U_{\abs{G}} \dots U_2U_1) \@ifnextchar\bra{{|{0}\rangle}\!}{{|{0}\rangle}}^{\otimes n} \;, \\ \rho &= \@ifnextchar\bra{{|{\Psi}\rangle}\!}{{|{\Psi}\rangle}}\bra{\Psi} \;, \\ F(\rho,\mathcal{D}(\rho)) &= \left( \mathop{\mathrm{tr}}{\sqrt{\sqrt{\rho} \mathcal{D}(\rho) \sqrt{\rho}}} \right)^2 \;. \label{eq:fidelity-dm} \end{align}\tag{5}\]

Firstly, Figure 3 (a) shows the impact of postselection on QFT with \(p_e=0.005\) perfect erasure checks (\(q=0\)) and no depolarising noise. With perfect postselection, the mean (and minimum) fidelity is by definition improved to exactly 1.0, because erasures constitute the sole noise channel, which is perfectly mitigated.

Secondly, Figure 3 (b) shows the same simulation but with imperfect erasure checks, where \(q=p_\mathrm{fp}=p_\mathrm{fn}=0.010\). Here, we can see how false negatives have a distinct but small impact on the expected \(F=1.0\) result with postselection, by erroneously allowing shots experiencing erasure into the accepted ensemble. Note again that in contrast, false positives do not affect the fidelity, but instead drive up the rejection rate/overhead (as is demonstrated starkly in Figure 2).

Thirdly, Figure 3 (c) shows the same simulation but with imperfect erasure checks and unheralded gate depolarising noise (\(p_\mathrm{1Q}=0.001,p_\mathrm{2Q}=0.010\)). Two key points are immediately apparent from this data:

  1. without postselection, the erasure and depolarising channels have stacking negative effects on fidelity;

  2. with postselection, the erasure channel is mitigated such that the mean fidelity accurately matches \(\lim_{T_1\to\infty}F\), excepting false negatives.

a

b

Figure 4: Gain in mean circuit fidelity for QFT achieved with postselection. For the absolute fidelity gain \(\Delta F\) (top), the curves are fitted via nonlinear least-squares regression on the difference between two reverse logistic (i.e.sigmoid) functions, as in Figure 3 (c) (with error bars propagated accordingly), with their approximated maxima shown. The relative gain (bottom) shows the quotient rather than the difference..

To summarise, this data shows the dynamics of adding in three different ‘layers’ of noise channel:

  1. gate-deleting erasure noise;

  2. classical bit-flip noise on the erasure checks;

  3. gate depolarising noise as unheralded Pauli error.

It is clear that postselection fully mitigates the erasure channel, such that \(F=1.0\) in the absence of other channels and \(F=F(\rho,\mathcal{D}(\rho))\) in the presence of a depolarising channel \(\mathcal{D}\), with an apparently minor impact by false negatives when \(q=0.010\).

In the next two subsections, we consider the following emergent questions: by how much does postselection improve fidelity with the full noise model (Section 4.3), and how close to \(\lim_{T_1\to\infty}F\) can we get with imperfect-check postselection (Section 4.4)?

4.3 Fidelity gain with postselection↩︎

Table 1: Drop in mean circuit fidelity for \(p_e=0.010\) and \(n=7\) for QFT with 10,000 target shots (as seen in Figure [fig:q-sweep]). Note that for \(q < 0.030\), \(\Delta F_{T_1\to\infty}\) is within error of zero.
\(q\) \(\Delta F_{T_1\to\infty}\)
0.005 \(\;\: \, 0.0022 \pm 0.0095\)
0.010 \(-0.0010 \pm 0.0095\)
0.015 \(-0.0073 \pm 0.0095\)
0.050 \(-0.0219 \pm 0.0095\)

Figure 3 shows how – for both the erasure and depolarising channels – the mean fidelity decays with a reverse sigmoid (i.e.S-shaped) curve as \(n\) increases. Of particular interest is how postselection under gate depolarising noise shallows this curve compared to no postselection (Figure 3 (c)), appearing to yield a gap of maximum magnitude at exactly one point. This gap therefore corresponds to the maximum absolute gain in circuit fidelity \(\Delta F\) achieved with postselection.

To study this further, we repeat the simulation for four different combinations of \(p_e\) and \(q\) and plot \(\Delta F\) as a function of \(n\) in Figure 4. As we may intuit from the sigmoids by inspection, \(\Delta F(n)\) takes the form of a unimodal bell-shaped curve.8 For each series, we fit the curve as a difference of two reversed logistic functions and maximise this function, yielding \[\begin{align} \mathop{\mathrm{arg\,max}}_n{\Delta F} &\approx 7 \;, \quad p_e=0.010 \;, \\ \mathop{\mathrm{arg\,max}}_n{\Delta F} &\approx 8 \;, \quad p_e=0.005 \;. \end{align}\] In other words, as \(p_e\) increases, the maximum fidelity gain from postselection also increases, but occurs for lower \(n\). Equivalently, the relative fidelity gain increases monotonically with \(n\).

Figure 5: Mean circuit fidelity for QFT with 10,000 target shots, for p_e=0.010 and n=7. As the check error rate q tends to zero, postselection perfectly mitigates the erasure channel, leaving only the gate depolarising channel. The curve fit is a ‘quadratic decay’ of the form F=d - (aq^2+bq+c), approximating the polynomial structure in the binomial expansion of false-negative events.

4.4 Approaching the limit of \(T_1\to\infty\)↩︎

Taking \(p_e=0.010\) and \(n=7\) as an example ‘sweet spot’ of maximum fidelity gain from the previous section (Figure 4), let us now consider the question of how effectively postselection mitigates against the erasure channel given imperfect erasure checks. For these values, due to gate depolarising noise, the expected fidelity with perfect postselection (from Equation 5 ) is \[\lim_{T_1\to\infty}{F} \approx 0.6021836859864419 \;.\]

In Figure 5, we plot the mean circuit fidelity and sweep the full entropy space of the erasure check error rate \(0 \leq q \leq 0.5\), observing that \[\lim_{q \to 0}{F} = \lim_{T_1\to\infty}{F} \;.\] As \(q\) increases, \(F\) decays from \(\lim_{T_1\to\infty}{F}\) by a polynomial factor arising from the binomial expansion of false-negative events.

From this data, we indicate in Table 1 the drop in mean circuit fidelity (compared to perfect postselection) caused by some realistic NISQ-era benchmarks for \(q\), noting that for \(q < 0.030\), \(\Delta F_{T_1\to\infty}\) is within error of zero.9 These values are consistent with the small influence of imperfect checks under the erasure channel seen between Figures 3 (a) and Figures 3 (b).

4.5 Quantum operation benchmarks↩︎

Table 2: Depolarising error budgets for QFT in the \(\qty{100}{\quop}\), \(\qty{500}{\quop}\) and \(\qty{1}{\kilo\quop}\) regimes. The improvement in mean fidelity from single-rail to postselected dual-rail models (as seen in Figure [fig:quop-benchmarks]) is listed as \(\Delta F\).
\(n\) \(\abs{G}\) \(p_\mathrm{1Q}\) \(p_\mathrm{2Q}\) \(\Delta F\)
5 149 0.016 0.050 0.03330360
10 514 0.004 0.013 0.07587702
15 1029 0.002 0.005 0.14007602

The number of gates \(\abs{G}\) is a key metric in describing the scale of quantum computation in practice, popularly given the unit quop as a parallel to the classical flop (i.e.quantum operations versus floating-point operations) [20]. We conclude our results with datasets representative of practical \(\require{physics} \qty{100}{\quop}\), \(\require{physics} \qty{500}{\quop}\) and \(\require{physics} \qty{1}{\kilo\quop}\) regimes.

To compare the performance of postselected dual-rail qubits against conventional systems of comparable specifications, we replace the erasure channel with an amplitude damping channel (whilst keeping the gate depolarising channel constant). As this amplitude damping channel takes an emission rate equal to the erasure rate \(p_e\), it models the same physical relaxation time \(T_1\) as a single-rail qubit.

Additionally, we account for decoherence caused by qubit idling; this is a significant source of noise, which can be mitigated by techniques such as dynamical decoupling, wherein circuit-preserving operations are performed on qubits during idle periods [21]. To model idling error in our simulations, noisy identity gates representing discrete idle operations are populated in the circuit using a modified dynamical decoupling scheduler (see Appendix 6).

Figure 6: Mean circuit fidelity for QFT with s_a=1000 target shots for both single-rail and postselected dual-rail qubits, assuming p_e=0.008 and q=0.005. The dashed horizontal line marks a 1/\sqrt{1000} \approx 3\% fidelity benchmark [22]. Table 2 lists the depolarising error budgets. Idling error is included, such that the datapoints correspond to \require{physics}
\qty{100}{\quop}, \require{physics}
\qty{500}{\quop} and \require{physics}
\qty{1}{\kilo\quop}.

Without idle operations, for some select \(n\), the number of gates \(\abs{G}(n)\) in the QFT construction as per Equation 1 is \[\begin{align} \abs{G}(5) &= 67 \;, \\ \abs{G}(10) &= 282 \;, \\ \abs{G}(15) &= 647 \;. \end{align}\] With idle operations added using our representation, these benchmarks increase to \[\begin{align} \abs{G}(5) &= 149 \;, \\ \abs{G}(10) &= 514 \;, \\ \abs{G}(15) &= 1029 \;, \end{align}\] thus corresponding to the \(\require{physics} \qty{100}{\quop}\), \(\require{physics} \qty{500}{\quop}\) and \(\require{physics} \qty{1}{\kilo\quop}\) regimes, respectively.

Figure 6 compares the mean fidelity for QFT between these single-rail and postselected dual-rail models. Across both models, we assume an emission rate \(p_e=0.008\) (approximately corresponding to a gate time \(\require{physics} t=\qty{500}{\nano\second}\) and physical relaxation time \(\require{physics} T_1=\qty{60}{\micro\second}\)) with accepted shots \(s_a=1000\). In the dual-rail case, we assume an erasure check error rate \(q=0.005\).

The depolarising error rates \(p_\mathrm{1Q}\) and \(p_\mathrm{2Q}\), listed in Table 2, are chosen for different \(n\) as rough estimates of practical error budgets, such that the factor in overall fidelity is no worse than 50% due to 1Q gate error, and 20% due to 2Q gate error.

We can see in Figure 6 that transitioning from the single-rail to the postselected dual-rail model consistently increases the mean fidelity, with this increase \(\Delta F\) listed in Table 2. In particular, we take a fidelity benchmark of \(1/\sqrt{s_a} \approx 0.03\), as an estimate of the noise floor resulting from shot noise as a Poisson process [22]. For \(n=15\), the single-rail model falls short of this minimum-required fidelity at \(F\approx0.023\), but the postselected dual-rail model comfortably surpasses it at \(F\approx0.163\).

5 Conclusion↩︎

In this work, we have presented a theoretical and simulational framework for the use of postselection with erasure-dominant qubits, e.g.dual-rail transmons. Acknowledging that the overhead of postselection scales exponentially due to the rejection rate (Equation 4 ; Figure 2), we have demonstrated empirically how postselection can accurately mitigate the erasure channel in terms of mean circuit fidelity for QFT, even in the presence of noisy erasure checks (Figure 3).

We have shown how the gain in fidelity achieved by postselection on such dual-rail systems increases with the erasure rate \(p_e\) (or, equivalently, decreases against the physical relaxation time \(T_1\)), with the maximum gain emerging at smaller circuit sizes as \(p_e\) increases (Figure 4). We have also shown how the fidelity converges polynomially to the limit of \(T_1\to\infty\) (in which only the gate depolarising channel remains), with the difference from this limit arriving within error of zero for erasure check error rate \(q<0.030\) (Figure 5; Table 1).

Finally, we have shown that a conventional single-rail system does not surpass an estimate of the fundamental noise floor for QFT in the kiloquop regime, whereas a comparable postselected dual-rail system does (Figure 6; Table 2). We believe these results may justify the use of erasure-based postselection as a near-term strategy for error mitigation, before (and, perhaps, combined with) the practical arrival of scalable quantum fault-tolerance.

Acknowledgements↩︎

This work was supported by the Innovate UK Quantum Missions pilot competition 10148061 DECIDE: Dimon error correction integrated into a data-centre environment.

We also thank Bryn Bell and Ailsa Keyser for their contributions and reviews.

6 Modelling idling error↩︎

Dynamical decoupling (DD) is a technique in which a sequence of quantum operations (e.g.\(\mathrm{X}\mathrm{X})\) that composes to the identity – and thus does not alter the overall circuit – is inserted into idle periods and acts to mitigate decoherence caused by qubit idling [21]. A DD scheduler is provided as a circuit transpiler pass by the qiskit-ibm-runtime package [9].

To simulate idling error in our simulations in Section 4.5, we use this transpiler pass to insert identity gates on qubits during idle periods; these identity gates are subject to all of the same noise channels we consider as the circuit gates.

This circuit transformation is implemented in erado by first performing an as-late-as-possible (ALAP) scheduling pass, and then performing a DD scheduling pass. We set every circuit gate to be exactly 1 unit of time in duration, and every identity gate inserted by the DD scheduler to be 0.8 units of time, with sequence_min_length_ratios equal to 1.0. With a DD sequence of \(\mathrm{I}\mathrm{I}\), the scheduler would therefore insert nothing into an idle period of 1 unit, and an \(\mathrm{I}\mathrm{I}\) sequence into an idle period of 2 units or greater.

In order to fill arbitrarily-long idle periods with arbitrarily-many identity operators, we populate the list of possible DD sequences with even-weight sequences of repeated \(\mathrm{I}\) gates in descending priority. Given a maximum idle sequence length \(k\), where \(k\) is necessarily even, the ordered list of possible idle sequences is \[( \mathrm{I}^k , \mathrm{I}^{k-2} , \dots , \mathrm{I}^2 ) \;.\] Resultingly, the number of identity gates inserted into an idle period of length \(t\) units of time is equal to \[\min\left\{ t\left\lfloor\frac{t}{2}\right\rfloor , k \right\} \;.\]

In a relatively sparse circuit such as QFT, increasing the number of qubits \(n\) notably increases the total amount of time spent idling by each qubit. We therefore use \(k\) as a parameter to tune the density of idling gates populated into the circuit representation for increasingly large \(n\). For example, a value of \(k=14\) is used in Section 4.5 to most closely map \(\abs{G}\) to the three operational regimes discussed, which also broadly agrees with ParityQC QFT compilations [23]. This is justified from the perspective of DD as a technique, wherein the mitigation achieved by aggressively populating DD sequences forms a trade-off with infidelity in the sequence gates themselves.

References↩︎

[1]
J. J. Burnett et al., “Decoherence benchmarking of superconducting qubits,” npj Quantum Inf, vol. 5, no. 1, p. 54, Jun. 2019, doi: 10.1038/s41534-019-0168-5.
[2]
M. A. Nielsen and I. L. Chuang, Quantum computation and quantum information, 10th Anniversary Edition. Cambridge University Press, 2010.
[3]
F. J. MacWilliams and N. J. A. Sloane, The theory of error-correcting codes. Elsevier, 1977.
[4]
A. R. Calderbank and P. W. Shor, “Good quantum error-correcting codes exist,” Phys. Rev. A, vol. 54, no. 2, pp. 1098–1105, Aug. 1996, doi: 10.1103/PhysRevA.54.1098.
[5]
E. Dennis, A. Kitaev, A. Landahl, and J. Preskill, “Topological quantum memory,” Journal of Mathematical Physics, vol. 43, no. 9, pp. 4452–4505, Sep. 2002, doi: 10.1063/1.1499754.
[6]
S. Aaronson, “Quantum computing, postselection, and probabilistic polynomial-time,” Proc. A, vol. 461, no. 2063, pp. 3473–3482, Sep. 2005, doi: 10.1098/rspa.2005.1546.
[7]
P. Kok, W. J. Munro, K. Nemoto, T. C. Ralph, J. P. Dowling, and G. J. Milburn, “Linear optical quantum computing with photonic qubits,” Rev. Mod. Phys., vol. 79, no. 1, pp. 135–174, Jan. 2007, doi: 10.1103/RevModPhys.79.135.
[8]
J. Wills, M. T. Haque, and B. Vlastakis, “Error-detected coherence metrology of a dual-rail encoded fixed-frequency multimode superconducting qubit,” Jun. 2025, Pre-published, doi: 10.48550/arXiv.2506.15420.
[9]
A. Javadi-Abhari et al., “Quantum computing with Qiskit,” Jun. 2024, Pre-published, doi: 10.48550/arXiv.2405.08810.
[10]
D. J. C. MacKay, Information theory, inference and learning algorithms. Cambridge University Press, 2003.
[11]
I. S. Reed and G. Solomon, “Polynomial codes over certain finite fields,” Soc. Indust. Appl. Math., vol. 8, no. 2, pp. 300–304, Jun. 1960, doi: 10.1137/0108018.
[12]
M. Grassl, T. Beth, and T. Pellizzari, “Codes for the quantum erasure channel,” Phys. Rev. A, vol. 56, no. 1, pp. 33–38, Jul. 1997, doi: 10.1103/PhysRevA.56.33.
[13]
N. Delfosse and G. Zémor, “Linear-time maximum likelihood decoding of surface codes over the quantum erasure channel,” Phys. Rev. Res., vol. 2, no. 3, p. 033042, Jul. 2020, doi: 10.1103/PhysRevResearch.2.033042.
[14]
S. J. Griffiths and D. E. Browne, “Union-find quantum decoding without union-find,” Phys. Rev. Res., vol. 6, no. 1, p. 013154, Feb. 2024, doi: 10.1103/PhysRevResearch.6.013154.
[15]
A. Kubica, A. Haim, Y. Vaknin, H. Levine, F. Brandão, and A. Retzker, “Erasure qubits: Overcoming the \(T_1\) limit in superconducting circuits,” Phys. Rev. X, vol. 13, no. 4, p. 041022, Nov. 2023, doi: 10.1103/PhysRevX.13.041022.
[16]
P.-J. H. S. Derks, A. Townsend-Teague, A. G. Burchards, and J. Eisert, “Designing fault-tolerant circuits using detector error models,” Quantum, vol. 9, p. 1905, Nov. 2025, doi: 10.22331/q-2025-11-06-1905.
[17]
Y. Han, L. A. Hemaspaandra, and T. Thierauf, “Threshold computation and cryptographic security,” SIAM J. Comput., vol. 26, no. 1, pp. 59–78, Feb. 1997, doi: 10.1137/S0097539792240467.
[18]
C. Gidney, “Stim: A fast stabilizer circuit simulator,” Quantum, vol. 5, p. 497, Jul. 2021, doi: 10.22331/q-2021-07-06-497.
[19]
A. G. Fowler, S. J. Devitt, and L. C. L. Hollenberg, “Implementation of Shor’s algorithm on a linear nearest neighbour qubit array,” Quantum Information & Computation, vol. 4, no. 4, pp. 237–251, Jul. 2004, [Online]. Available: https://arxiv.org/abs/quant-ph/0402196.
[20]
J. Preskill, “Beyond NISQ: The megaquop machine,” ACM Transactions on Quantum Computing, vol. 6, no. 3, pp. 18:1–18:7, Apr. 2025, doi: 10.1145/3723153.
[21]
P. Krantz, M. Kjaergaard, F. Yan, T. P. Orlando, S. Gustavsson, and W. D. Oliver, “A quantum engineer’s guide to superconducting qubits,” Applied Physics Reviews, vol. 6, no. 2, p. 021318, Jun. 2019, doi: 10.1063/1.5089550.
[22]
I. G. Hughes and T. P. A. Hase, Measurements and their uncertainties: A practical guide to modern error analysis. Oxford University Press, 2010.
[23]
B. Klaver et al., SWAP-less implementation of quantum algorithms,” Phys. Rev. A, vol. 113, no. 1, p. 012443, Jan. 2026, doi: 10.1103/2wzk-fnhx.

  1. sgriffiths@oqc.tech↩︎

  2. jfriel@oqc.tech↩︎

  3. bvlastakis@oqc.tech↩︎

  4. https://github.com/oqc-community/erado↩︎

  5. A code’s distance is the minimum Hamming distance between any two codewords, which is equivalent to the weight of the smallest undetectable error.↩︎

  6. Note that this is not a linear code, as \(C\) is not closed under addition (readily apparent from the fact that the all-zero state is not a codeword).↩︎

  7. (e.g.accepted/rejected proportions, logical state measurements and circuit fidelities)↩︎

  8. In fact, the integral of any unimodal bell-shaped curve is sigmoidal by definition.↩︎

  9. See Figure 3 for error calculation.↩︎