July 10, 2026
Accurate noise classification is essential for operating near-term quantum processors, yet existing approaches, such as quantum process tomography, scale exponentially with system size, limiting their practicality for routine calibration. We propose a scalable noise fingerprinting pipeline that combines structured classical shadow tomography with physics-informed feature engineering to identify noise channels from a fixed set of 3-qubit probe circuits. Each sample is represented by a 279-dimensional feature vector constructed from randomized Pauli measurements and derived observables, designed to resolve physically similar noise channels that produce overlapping signatures under generic measurement sets. We evaluate three classifiers, i.e., random forest, extra trees, and a multilayer perceptron, on a dataset of 14,000 labeled samples spanning 10 noise types. The random forest classifier achieves the highest test accuracy of 0.8426 with a macro F1 score of 0.8437, outperforming both baselines. Confusion analysis reveals that many noise types are classified with high reliability, with the remaining confusions occurring between channels sharing similar physical decay mechanisms, motivating future work on richer probe states and noise parameter estimation.
quantum noise classification, classical shadow tomography, noise fingerprinting, machine learning
Quantum processors based on superconducting qubits and trapped ions are rapidly advancing toward practical applications, yet their performance remains fundamentally constrained by hardware noise [1]. Noise in these systems arises from a variety of physical mechanisms, including energy relaxation, dephasing, and measurement errors, each of which degrades quantum information in distinct ways. Understanding the dominant noise sources in a given device is therefore a prerequisite for effective error mitigation and the design of noise-aware quantum algorithms.
Existing approaches such as quantum process tomography [2] provide complete noise classification but require resources that scale exponentially with system size, making them impractical for routine calibration of large-scale devices. Direct fidelity estimation and randomized benchmarking offer partial improvements in scalability, yet remain limited to specific noise models or rely on hardcoded decision rules that do not generalize across diverse error channels. As processors scale in qubit count and circuit depth, efficient and generalizable noise identification becomes increasingly critical [1].
In this work, we address this gap by proposing a noise fingerprinting pipeline1 that exploits the efficiency of classical shadow tomography [3] to extract structured measurement data from a small set of 3-qubit probe circuits [4]. Rather than relying on generic observables, we construct a physics-informed feature representation designed to amplify distinctions between physically similar noise channels. We benchmark three machine learning classifiers on ten candidate noise models and demonstrate that ensemble methods substantially outperform a neural baseline, achieving over 0.84 test accuracy on a dataset of 14,000 labeled samples.
Our contributions are twofold. First, the paper proposes a shadow-based noise fingerprinting pipeline using fixed 3-qubit probe circuits with a 279-dimensional physics-informed feature representation combining Pauli-shadow observables and derived coherence/population/asymmetry features. Second, we propose an empirical evaluation across ten Qiskit noise models and three classifiers, including confusion and scaling analysis. Our artifacts are publicly available at https://github.com/AdaHgrace/Noise_Fingerprinting.
Quantum noise in noisy intermediate-scale quantum (NISQ) devices [1] arises from distinct physical mechanisms modeled as quantum channels acting on the system density matrix [5]. We consider ten noise models spanning a broad range of error types: depolarizing, phase flip, bit flip, readout error, phase damping, thermal relaxation, amplitude damping, phase-amplitude damping, Pauli-asymmetric, and reset noise [5]. These range from symmetric decoherence channels to combined energy-and-coherence loss models, providing a diverse and physically realistic benchmark for noise fingerprinting.
Classical shadow tomography [3] provides an efficient protocol for estimating many properties of a quantum state from a relatively small number of randomized measurements. In this framework, a quantum state is repeatedly prepared and measured in randomly chosen Pauli bases, and the resulting measurement outcomes are used to construct a classical representation of the state, known as a classical shadow. This approach achieves a significant reduction in measurement overhead compared to full state tomography, making it particularly well-suited for extracting noise-relevant observables from near-term quantum devices. Extensions of the classical shadow framework to noisy settings [6] have further demonstrated its robustness under realistic device conditions.
Quantum process tomography [2] and direct fidelity estimation provide complete characterizations of noise channels but require measurement resources that scale exponentially with the number of qubits, making them impractical for routine calibration of devices beyond a few qubits. Randomized benchmarking offers a more scalable alternative by estimating average gate fidelities through sequences of random Clifford gates, yet it does not resolve individual noise types or provide the per-channel discrimination needed for automated noise fingerprinting. These limitations motivate the development of methods that combine efficient measurement protocols with data-driven classification to identify noise models at scale.
Recent work has demonstrated that machine learning classifiers can identify noise signatures from measurement data [7], though most approaches are limited to a small number of noise types or rely on idealized probe states that may not reflect realistic device conditions. Neural network approaches have been used to classify Markovian and non-Markovian noise processes from time-series measurement data [8], while large-scale studies have benchmarked machine learning models including random forests and multilayer perceptrons (MLPs) for error mitigation on hardware with up to 100 qubits [9]. In this work, we hypothesize that ensemble methods such as random forests may offer interpretability advantages over purely neural approaches for structured feature spaces derived from quantum measurements, a question we investigate empirically in Section 4. Our work builds on these foundations by combining structured shadow measurements with physically motivated feature engineering to enable classification across a broader and more realistic set of ten noise models.

Figure 1: Shadow-based noise fingerprinting pipeline consisting of four stages: probe circuit preparation, noisy device execution, randomized shadow measurements, and machine learning classification..
Our noise fingerprinting pipeline takes as input the measurement outcomes from a fixed set of probe circuits executed on the noisy device and outputs a predicted noise type from a set of ten candidate models. The pipeline consists of four stages: 1) probe state preparation, 2) randomized shadow measurements, 3) feature extraction, and 4) machine learning classification, as illustrated in Fig. 1. At each stage, the design choices are motivated by the need to extract maximally discriminative information from a minimal number of measurements, reflecting the practical constraints of near-term quantum hardware. The full pipeline is implemented in simulation using Qiskit [10] and evaluated on a dataset of 14,000 labeled samples distributed equally across the ten noise models.
We use 3-qubit QAOA circuits [4] as complex probe states, owing to their ability to generate entangled states with rich coherence structure that is sensitive to a broad range of noise mechanisms. QAOA circuits are parameterized by alternating cost and mixer layers, and in this work we use a single-layer ansatz with randomly sampled parameters to ensure diversity across probe instances. These complex probe states are supplemented with simple structured states, including computational basis states \(|0\rangle^{\otimes 3}\), \(|1\rangle^{\otimes 3}\), and uniform superpositions \(|+\rangle^{\otimes 3}\), to capture both population-level and coherence-level noise signatures respectively. The combination of complex and simple probe states ensures that the feature extraction stage receives complementary information, improving the discriminability of physically similar noise channels such as phase damping and thermal relaxation.
Each sample is represented by a 279-dimensional feature vector constructed from 9 probe circuits, contributing 31 features per probe circuit. The probe set consists of 5 QAOA probe circuits with randomly sampled parameters and 4 simple structured states. The first 18 features per probe are raw Pauli expectation values estimated via the classical shadow protocol [3], in which each probe circuit is followed by a randomly sampled Pauli basis rotation prior to measurement, and expectation values are reconstructed from the resulting bit strings over 200 shots. We estimate 9 single-qubit observables covering all three Pauli axes on each of the 3 qubits, and 9 same-axis two-qubit correlation observables covering all qubit pairs, e.g. \(X_1X_2I_3\) and \(Z_1I_2Z_3\).
The remaining 13 features per probe are physically motivated derived quantities. Let \(\bar{X}\), \(\bar{Y}\), \(\bar{Z}\) denote the mean absolute response along each Pauli axis. These include: axis means \(\bar{X}\), \(\bar{Y}\), \(\bar{Z}\); coherence and population strengths \(\bar{X}+\bar{Y}\) and \(\bar{Z}\); pairwise differences \(\bar{X}-\bar{Y}\), \(\bar{Z}-\bar{X}\), \(\bar{Z}-\bar{Y}\); and coherence-to-population ratios \(\frac{\bar{X}+\bar{Y}}{\bar{Z}+\varepsilon}\) and \(\frac{\bar{Z}}{\bar{X}+\bar{Y}+\varepsilon}\). This structured design is motivated by the observation that physically similar channels such as depolarizing and Pauli-asymmetric noise produce nearly indistinguishable signatures under generic observable sets, necessitating physics-informed features to resolve fine-grained distinctions between noise types.
We benchmark three machine learning classifiers on the extracted feature vectors to assess whether model complexity is a limiting factor in noise type identification. The first two classifiers are ensemble tree methods: random forest and extra trees. Both methods construct a large number of decision trees over random subsets of the training data and features, aggregating their predictions through majority voting. Extra trees introduces additional randomness by selecting split thresholds randomly rather than optimally, which can improve generalization on high-dimensional feature spaces. The third classifier is an MLP with three hidden layers of 256, 128, and 64 units respectively, using ReLU activations and the Adam optimizer, included as a neural baseline to evaluate whether the added expressivity of a neural network provides any advantage over ensemble methods on this task.
All three classifiers are implemented using scikit-learn [11]. Hyperparameters are left at their default values to provide a controlled baseline evaluation, isolating the effect of the feature representation rather than classifier tuning.
We generated a dataset of 14,000 labeled samples distributed equally across 10 noise types, yielding 1,400 samples per class. Each sample corresponds to a single noise configuration applied to a 3-qubit probe circuit, with noise strength sampled uniformly at random from \([0.01, 0.15]\) for all noise types. Feature extraction is performed, producing a 279-dimensional feature vector per sample from structured shadow measurements under 200 shots per probe circuit. We use a 20% held-out test set of 2,800 samples; the remainder is split 85:15 into 9,520 training and 1,680 validation samples.
| Classifier | Accuracy | Macro F1 |
|---|---|---|
| Random Forest | \(0.8426 \pm 0.0036\) | \(0.8437 \pm 0.0039\) |
| Extra Trees | \(0.8406 \pm 0.0019\) | \(0.8416 \pm 0.0024\) |
| MLP | \(0.7925 \pm 0.0042\) | \(0.7924 \pm 0.0046\) |
| Values reported as mean \(\pm\) standard deviation over 3 random seeds. | ||
Table ¿tbl:tab:results? summarizes the classification performance of all three models evaluated on the held-out test set. The random forest classifier achieved the highest mean test accuracy of 0.8426 with a macro F1 score of 0.8437, outperforming both extra trees (0.8406, F1: 0.8416) and the MLP baseline (0.7925, F1: 0.7924). The low standard deviations across all three classifiers indicate that these results are stable across different random seeds and train/test splits. The strong performance of both ensemble methods relative to the MLP suggests that the structured feature representation is well-suited to tree-based classifiers.

Figure 2: Confusion matrix for the random forest classifier evaluated on the held-out test set across ten noise types..
Figure 2 presents the confusion matrix for the random forest classifier evaluated on the held-out test set. Readout error, phase flip, bit flip, and thermal relaxation achieve high per-class accuracy, with the classifier correctly identifying the majority of samples in each of these categories. The primary misclassifications occur between phase damping and thermal relaxation, and between phase-amplitude damping and reset noise, suggesting that these channel pairs produce similar feature signatures under the current measurement protocol. The remaining noise types, including amplitude damping, phase-amplitude damping, and Pauli-asymmetric noise, are classified with moderate accuracy, with their errors spread across several physically related channels rather than concentrated in a single confusion.

Figure 3: Test accuracy as a function of samples per class for random forest, extra trees, and MLP classifiers..
Figure 3 presents classification accuracy as a function of samples per class for all three classifiers. At 300 samples per class, all three models perform near chance level, achieving accuracies in the range of 0.46-0.49. Accuracy increases consistently with dataset size across all classifiers, with random forest and extra trees reaching 0.8426 and 0.8406, respectively, at 1,400 samples per class, while the MLP achieves 0.7925 at the same scale. The ensemble methods maintain comparable performance throughout the scaling curve, with both consistently outperforming the MLP baseline at larger dataset sizes. The steep initial improvement followed by a gradual plateau suggests that the feature representation contains sufficient discriminative information for most noise classes, but that a minimum sample threshold is required to resolve decision boundaries between physically similar channels.
The high accuracy of the ensemble classifiers demonstrates that structured shadow measurements with physics-informed features provide a discriminative representation for most noise models.The confusion between phase damping and thermal relaxation is consistent with their shared physical mechanism of coherence decay, which produces similar expectation value profiles under the current probe and observable set. Similarly, the confusion between phase-amplitude damping and reset noise reflects the difficulty of distinguishing channels that produce overlapping effects on both population and coherence observables. The rapid overfitting of the MLP without corresponding accuracy gains suggests that the primary bottleneck is not model expressivity but rather the intrinsic separability of the feature representation for these overlapping noise types. Further gains are therefore more likely to come from richer probes or observables than from more expressive classifiers.
The current framework has several limitations that should be acknowledged. First, the pipeline is limited to 3-qubit circuits, and it remains unclear whether the feature representation scales effectively to larger qubit counts without a corresponding increase in measurement overhead. Second, the pipeline does not yet support noise parameter estimation, providing only a discrete classification of noise type rather than a continuous characterization of noise strength. Third, three derived features associated with an unpopulated mixed-axis observable group were found to be constant across samples, indicating a small amount of redundancy in the 279-dimensional feature space that future work could address with genuine mixed-Pauli observables. Fourth, all experiments are conducted in simulation, and the impact of finite shot noise, state preparation and measurement errors, and hardware-specific noise correlations on classification performance has not yet been evaluated. Finally, the current sprobe circuit set may not provide sufficient coverage of the feature space for noise channels that are only weakly excited by QAOA-structured states.
This paper presented a shadow-based noise fingerprinting pipeline for classifying quantum noise channels from a fixed set of 3-qubit probe circuits. By combining randomized Pauli measurements with physics-informed feature engineering, the proposed method achieved 0.8426 accuracy and 0.8437 macro F1 across ten simulated noise models, with tree-based ensemble classifiers outperforming the neural baseline.
The results suggest that structured shadow-derived features can capture discriminative signatures for many common quantum noise channels while requiring substantially less information than full process tomography. The confusion analysis further shows that remaining errors are concentrated among physically similar channels, indicating that the primary limitation lies in the separability of the current probe-feature design rather than classifier capacity.
Future work will extend the framework to larger systems, incorporate noise-parameter estimation, evaluate robustness under real hardware noise and finite-shot variation, and explore probe designs that better separate physically overlapping channels.
Here, we use the term “noise fingerprinting” to refer to supervised classification of predefined noise-model families from shadow-derived measurement features. Our study focuses on simulated Qiskit noise channels and does not attempt full noise characterization or continuous noise-parameter estimation.↩︎