July 01, 2026
Artificial intelligence (AI) and quantum information (QI) are rapidly co-evolving. AI is becoming a practical tool for learning, designing, controlling, and verifying quantum systems, while QI offers new computational models, representational structures, and learning-theoretic questions for AI. This survey reviews the interface from both directions. In the AI for QI direction, we organize recent progress around the central tasks of extracting information from limited measurements, training and discovering quantum algorithms, stabilizing noisy hardware, automating experimental and programming workflows, and extending learning-based methods to sensing and networking. In the QI for AI direction, we examine how quantum computation and quantum-inspired structures affect learning through algorithmic speedups, expressivity, trainability, generalization, neural-network design, and tensor-network representations. We close by identifying cross-cutting challenges in reproducibility, scalability, hardware realism, and co-design, arguing that progress will depend on tighter integration of theory, experiment, and hybrid quantum–classical systems.
Artificial intelligence (AI) and quantum information (QI) science and technology are two rapidly advancing fields whose development is increasingly intertwined, though in an asymmetric way. AI already provides powerful tools for modeling complex systems and supporting inference, optimization, design, and discovery. Quantum computing seeks computational advantages from superposition, interference, and entanglement, with the long-term goal of solving problems that remain intractable for classical hardware. Related QI technologies, including quantum sensing and quantum networking, exploit quantum resources to enhance measurement, discrimination, communication, and distributed information processing. The interaction between these areas is becoming central to future progress: AI is emerging as an essential tool for building, controlling, and operating quantum devices, optimizing quantum sensors, and designing quantum networking protocols, while quantum computing and quantum-inspired methods are being explored as new computational and representational resources for machine learning and AI.
This interaction is driven by a common technical reality: both AI and QI must reason about complex systems through limited and noisy observations. A quantum device or protocol rarely exposes the object of interest directly. The relevant state, signal, or noise process has to be inferred from measurements and then acted on under tight experimental constraints. This is why AI methods now appear throughout the quantum-technology stack, supporting device characterization, variational training, hardware control, sensing, networking, and increasingly autonomous workflows. The reverse direction starts from a different but related question: whether quantum mechanics offers useful computational or representational structure for AI. Quantum computers naturally manipulate high-dimensional linear-algebraic objects, while quantum models force one to think carefully about encoding, entanglement, and what information can actually be read out from measurement. Even when this does not lead to an immediate speedup, it can suggest new ways to analyze model capacity, generalization, and efficient classical architectures.
This overlap has become especially timely. On the AI side, large language models have moved beyond prompt-conditioned text generation toward tool use, long-horizon planning, and autonomous scientific workflows, to the point that recent work treats agentic science as an emerging research paradigm [1]–[3]. AlphaEvolve [4] and Aurora [5] illustrate the same shift from another angle: AI is beginning to participate directly in algorithm design and scientific prediction workflows, with a growing role across the full scientific pipeline. On the quantum side, some of the newest signals appear directly at the AI–quantum interface. Quantum oracle sketching [6] revisits a long-standing concern in quantum machine learning by showing that small quantum processors can give exponential space advantages for learning tasks on massive classical data without assuming full quantum random access memory (QRAM). Recent calibration benchmarks [7] and AI-based surface-code pre-decoders [8] point to a growing role for AI in quantum calibration and error-correction workflows. Recent demonstrations in networked quantum interferometry [9] and integrated photonic quantum optics [10] also show that QI technologies are developing across sensing, networking, hardware platforms, and quantum processors. This convergence motivates a more comprehensive review of the interface between AI and QI.
Position of this work relative to prior surveys. Recent reviews cover important parts of this literature. Broad overviews document AI across the quantum-computing workflow [11] and the wider bidirectional notion of quantum AI [12], [13]. Other recent surveys focus on specific subproblems, including quantum machine learning [14], [15], the complexity of learning quantum states [16], AI methods for representing and characterizing quantum systems [17], machine learning for estimation and control of quantum systems [18], barren plateaus in variational quantum models [19], [20], and tensor-network or low-rank methods for machine learning [21], [22].
This review connects the two directions, AI for QI and QI for AI, through recurring questions about learning, representation, optimization, control, and resource constraints. On the AI for QI side, we focus on how learning-theoretic, optimization-theoretic, and control-theoretic ideas help characterize quantum systems, train quantum models, stabilize noisy hardware, and design quantum sensing and networking protocols. On the QI for AI side, we examine algorithmic speedups together with how quantum models and quantum-inspired constructions alter representation power, generalization behavior, and architectural design. We also highlight tensor networks as an important bridge because they connect quantum many-body structure, efficient representation, and practical machine learning through explicit multilinear models.
Organization. The review proceeds as follows. Section 2 collects the quantum-mechanical, machine-learning, and computational preliminaries used throughout the paper. Section 3 surveys artificial intelligence for quantum information, covering statistical learning of quantum systems, AI-theoretic perspectives on quantum algorithms, AI-assisted quantum algorithm discovery, learning, correction, and control of noisy quantum systems, autonomous quantum workflow orchestration, and AI-enabled quantum sensing and networking. Section 4 turns to quantum computing and quantum-inspired methods for artificial intelligence. There we discuss quantum algorithmic speedups for classical machine learning, conditions and limits of quantum learning benefits, quantum neural networks, quantum-inspired analyses of classical neural networks, and tensor-network methods for machine learning. Section 5 identifies cross-cutting challenges, and Section 6 concludes.
Across all sections, our central argument is that the interface between AI and QI is an emerging theory and practice of learning, representation, and control under quantum constraints. For convenience, Table 1 summarizes the globally recurring notation used throughout this review.
This section collects background on how quantum states and measurements are represented, how dynamics and noise are modeled, and how parameterized circuits become learnable objects under finite-shot data.
We begin with the state space. An \(n\)-qubit system is described on the Hilbert space \(\mathcal{H}=(\mathbb{C}^{2})^{\otimes n}\). Its physical state is represented by a density operator \(\rho\) satisfying \(\rho\succeq 0\) and \(\require{physics} \Tr(\rho)=1\). Pure states are the special case \(\rho=\ket{\psi}\bra{\psi}\), whereas mixed states describe either classical uncertainty or entanglement with degrees of freedom that have been ignored. For a bipartite system \(AB\), the state of subsystem \(A\) is the reduced density operator \(\require{physics} \rho_A=\Tr_B(\rho_{AB})\). This notion of reduction is important throughout the section, because later discussions of locality, data embeddings, and barren plateaus often depend on whether reduced states remain structured or become close to maximally mixed.
Observables are represented by Hermitian operators \(O\), and their expectation values are given by \[\require{physics} \langle O\rangle_{\rho}=\Tr(O\rho).\] Measurements are more generally described by a positive-operator-valued measure (POVM) \(\{M_s\}\), where each \(M_s\succeq 0\) and \(\sum_s M_s=\mathbb{I}\). The probability of observing outcome \(s\) is given by the Born rule \[\require{physics} p_{\rho}(s)=\Tr(M_s\rho).\] The key practical point is that quantum information is not accessed directly. What one obtains experimentally is a collection of classical samples distributed according to \(p_\rho(s)\). Thus, in essentially all AI-for-quantum settings, the learning algorithm infers latent quantum objects from partial and noisy measurement data, without direct access to \(\rho\) itself [23], [24].
Dynamics and noise enter through maps on quantum states. In closed systems, evolution is generated by a Hamiltonian \(H\) and acts unitarily as \(\rho\mapsto U\rho U^\dagger\) with \(U=e^{-itH}\). In open systems and on real hardware, evolution is more generally modeled by a completely positive trace-preserving map, or quantum channel, denoted \(\mathcal{E}\). A quantum channel is a linear map that sends density operators to density operators, and it admits the Kraus form \[\mathcal{E}(\rho)=\sum_a K_a \rho K_a^\dagger, \qquad \sum_a K_a^\dagger K_a=\mathbb{I}.\] In this review, the word process is used in essentially the same sense, namely an input-output map acting on quantum states. In continuous-time descriptions one often writes \(\dot{\rho}=\mathcal{L}(\rho)\), where \(\mathcal{L}\) is a Liouvillian generator. This distinction matters because the tasks surveyed below range from learning static quantum properties to learning unknown dynamics, noise models, and control responses on NISQ devices [25].
In this language, quantum noise refers to departures from ideal unitary evolution, such as decoherence, dissipation, and control imperfections, and these effects are often absorbed into an effective channel or Liouvillian description. A particularly common source of experimental error is state preparation and measurement (SPAM) error, meaning that the prepared input state and the recorded measurement outcome may both differ systematically from the intended idealized model. This matters throughout the review because realistic learning and control problems often require simultaneous inference of target dynamics and nuisance noise processes.
Another recurring object is the parameterized quantum circuit (PQC), or ansatz. A typical PQC is a unitary family \(U(\boldsymbol{\theta})\) built from elementary gates whose angles are collected in a parameter vector \(\boldsymbol{\theta}=(\theta_1,\ldots,\theta_p)\). In many applications one starts from an input state \(\rho_{\mathrm{in}}\), applies \(U(\boldsymbol{\theta})\), and reads out an observable \(O\). This produces the expectation-valued model \[\require{physics} f(\boldsymbol{\theta})=\Tr\!\left[O\,U(\boldsymbol{\theta})\rho_{\mathrm{in}}U^\dagger(\boldsymbol{\theta})\right].\] In quantum machine learning, one often inserts a data-encoding circuit before or inside the ansatz, so the model may also depend on an input \(x\) through a feature state \(\ket{\phi(x)}\) or an encoded density operator \(\rho(x)\). In variational quantum algorithms, the same structure appears with a task-specific cost observable in place of a supervised-learning label.
Once the expectation value above is inserted into a classical loss function \(\mathscr{L}(\boldsymbol{\theta})\), training proceeds by a hybrid quantum–classical loop. The classical optimizer updates \(\boldsymbol{\theta}\), while the quantum device is repeatedly queried to estimate \(f(\boldsymbol{\theta})\) and, when needed, its gradients. This is where the statistical character of the problem becomes unavoidable: these quantities are estimated through finitely many circuit repetitions. As a result, optimization quality depends on how many shots are available and how noisy the hardware is.
This setup explains several recurring themes in Section 3. Some subsections focus on inference from measurement data, some on the optimization of variational models, and others on robustness under finite-shot and noisy conditions. In all three cases, the quantum object of interest is only indirectly accessible, and the available data, feedback, and optimization signals are statistical and resource-limited. One-off local indices in theorem statements, asymptotic bounds, or subsection-specific examples are defined in place. In particular, we reserve \(\mathcal{L}\) for Liouvillian generators and use \(\mathscr{L}\) for loss or optimization objectives.
This review uses a small set of standard AI and ML ideas repeatedly. In supervised learning, one observes a training set \(\mathcal{S}_{\mathrm{tr}}=\{(x_i,y_i)\}_{i=1}^{N}\) drawn from an underlying data distribution \(\mathcal{D}\), chooses a hypothesis or model \(f_{\boldsymbol{\theta}}\) from a parameterized family, and fits the parameters by minimizing an empirical objective \[\widehat R(\boldsymbol{\theta}) = \frac{1}{N}\sum_{i=1}^{N} \ell\!\left(f_{\boldsymbol{\theta}}(x_i),y_i\right),\] where \(\ell\) is a per-example loss. The corresponding population risk is \[R(\boldsymbol{\theta}) = \mathbb{E}_{(x,y)\sim\mathcal{D}} \!\left[\ell\!\left(f_{\boldsymbol{\theta}}(x),y\right)\right].\] The gap \[\mathcal{G}(\boldsymbol{\theta}) = R(\boldsymbol{\theta})-\widehat R(\boldsymbol{\theta})\] is the basic object in generalization analysis [26], [27]. In practice, one often optimizes a regularized objective \[\widehat{\boldsymbol{\theta}} \in \arg\min_{\boldsymbol{\theta}} \left[ \widehat R(\boldsymbol{\theta}) + \lambda_{\mathrm{reg}}\operatorname{Reg}(\boldsymbol{\theta}) \right],\] where \(\operatorname{Reg}\) encodes a structural preference such as sparsity, smoothness, locality, or small norm. Learning guarantees often bound the probability that the generalization gap exceeds a tolerance, \[\Pr\!\left[ \mathcal{G}(\widehat{\boldsymbol{\theta}})>\varepsilon \right] \le \delta,\] with sample complexity controlled by three factors. The first is the target accuracy \(\varepsilon\), and the second is the allowed failure probability \(\delta\). The third factor is the part that changes from setting to setting: it measures how hard the model class or prediction task is. In classical PAC (Probably Approximately Correct) learning this role is played by the VC (Vapnik-Chervonenkis) dimension \(d_{\mathrm{VC}}\), which is a combinatorial measure that quantifies the capacity, or complexity, of a hypothesis class. In classical-shadow prediction it is often the shadow norm, or the locality \(k\) of the observables. In variational-model generalization it can be the number of active gates or trainable parameters. In quantum machine learning (QML) generalization bounds it can be a margin or an effective dimension computed from the trained model and the data. For unsupervised, generative, or system-identification tasks, labels may be absent and the objective may be a likelihood, reconstruction error, divergence, or prediction loss for measured trajectories. This empirical-risk language underlies our discussions of quantum tomography, Hamiltonian and channel learning, QML benchmarks, and tensor-network learning models.
Training is the process of finding useful parameters under finite computation, data, and measurement budgets. Gradient-based methods use derivatives of \(\widehat R\) or of a task-specific cost, while derivative-free methods, Bayesian optimization, evolutionary search, and reinforcement learning can be preferable when gradients are unavailable, expensive, noisy, or delayed. In this review, trainability refers to whether useful optimization signal can be obtained and exploited with feasible resources; sample efficiency refers to how much experimental or training data are needed; and inductive bias refers to structural assumptions built into the model, such as locality, symmetry, sparsity, temporal memory, or graph structure. These notions recur in the analysis of barren plateaus, quantum kernels, neural decoders, autonomous calibration, sensing, and networking.
Several model families appear throughout the review. A kernel method represents data through pairwise similarities \(k(x,x')\), producing a Gram matrix \[\mathbf{K}_{ij}=k(x_i,x_j)\] on the training set and reducing many supervised tasks to regularized optimization in the induced feature space [28], [29]. A typical kernel predictor has the form \[f_{\boldsymbol{\alpha}}(x) = \sum_{i=1}^{N} \alpha_i k(x,x_i), \qquad \boldsymbol{\alpha} =(\mathbf{K}+\lambda_{\mathrm{reg}}I_N)^{-1}\mathbf{y}\] in kernel ridge regression. In quantum kernel methods, the similarity is often estimated from overlaps of quantum feature states, for example \(k(x,x')=|\langle\phi(x)|\phi(x')\rangle|^2\).
A neural network is a parameterized composition of linear maps and nonlinear or attention-based transformations. A feed-forward example can be written as \[\begin{align} h_0 &= x,\\ h_\ell &= \sigma_\ell(W_\ell h_{\ell-1}+b_\ell),\\ f_{\boldsymbol{\theta}}(x) &= W_L h_{L-1}+b_L, \end{align}\] where the trainable parameters \(\boldsymbol{\theta}\) collect the weights and biases. Different neural architectures encode different biases: CNNs (Convolutional Neural Networks) emphasize locality, RNNs (Recurrent Neural Networks) and sequence models emphasize temporal structure, Transformers use attention to model long-range dependencies, and graph neural networks propagate information along graph edges. In a Transformer block, the basic attention operation is \[\operatorname{Attn}(Q_{\mathrm{att}},K_{\mathrm{att}},V_{\mathrm{att}}) = \operatorname{softmax}\!\left( \frac{Q_{\mathrm{att}}K_{\mathrm{att}}^{T}}{\sqrt{d_{\mathrm{att}}}} \right)V_{\mathrm{att}},\] where \(K_{\mathrm{att}}\) denotes attention keys and should not be confused with the Gram matrix \(\mathbf{K}\). Tangent-kernel analyses study training by linearizing \(f_{\boldsymbol{\theta}}\) around its initialization. The classical neural tangent kernel (NTK) [30], \[\Theta^{\mathrm{NTK}}_{ij}(\boldsymbol{\theta}) = \nabla_{\boldsymbol{\theta}} f_{\boldsymbol{\theta}}(x_i)^{T} \nabla_{\boldsymbol{\theta}} f_{\boldsymbol{\theta}}(x_j),\] motivates the quantum neural tangent kernel (QNTK) discussion later in the review. For graph-structured inputs, a typical message-passing layer updates a node representation by aggregating information from its neighbors, \[h_v^{(\ell+1)} = \psi_\ell\!\left( h_v^{(\ell)}, \bigoplus_{u\in\mathcal{N}(v)} \phi_\ell(h_v^{(\ell)},h_u^{(\ell)},e_{uv}) \right),\] where \(\mathcal{N}(v)\) is the neighborhood of node \(v\), \(e_{uv}\) is an edge feature, \(\oplus\) is a permutation-invariant aggregation operation, and \(\phi_\ell,\psi_\ell\) are learnable maps.
Generative models learn a distribution \(p_{\boldsymbol{\theta}}(x)\) or a sampler whose outputs resemble the data distribution. For explicit density models, maximum-likelihood training minimizes \[\widehat R_{\mathrm{NLL}}(\boldsymbol{\theta}) = -\frac{1}{N}\sum_{i=1}^{N} \log p_{\boldsymbol{\theta}}(x_i).\] The NLL is the negative average log-likelihood of the dataset, so minimizing \(\widehat{R}_{\text{NLL }}\) encourages the model to assign high probability density to the observed samples. For sequence models, including most large language models, the learned distribution is often factorized autoregressively as \[p_{\boldsymbol{\theta}}(x_{1:T}) = \prod_{t=1}^{T} p_{\boldsymbol{\theta}}(x_t\mid x_{<t}).\] This category includes explicit density models, implicit samplers, Born-machine-style models, and diffusion models that generate samples through learned denoising dynamics [31]. A simplified diffusion-style denoising objective has the form \[\min_{\boldsymbol{\theta}} \mathbb{E}_{t,x_0,\boldsymbol{\xi}} \!\left[ \left\| \boldsymbol{\xi} - \boldsymbol{\xi}_{\boldsymbol{\theta}}(x_t,t) \right\|_2^2 \right],\] where \(x_t\) is a noised version of a data point and \(\boldsymbol{\xi}_{\boldsymbol{\theta}}\) is a learned denoiser. Foundation models and large language models are high-capacity generative models trained on broad corpora and adapted through prompting, fine-tuning, retrieval, or tool use. These ideas appear below in discussions of quantum generative models, tensor-network generative models, scientific foundation models, autonomous programming, and self-driving laboratory workflows.
For losses derived from negative log-likelihoods, an important curvature object is the Fisher information matrix. Given a dataset \(\mathcal{D}=\left\{x_i\right\}_{i=1}^M\), the empirical Fisher matrix is defined as \[\tilde{F}_{\mu \nu}(\boldsymbol{\theta})=\frac{1}{M} \sum_{i=1}^M \partial_\mu \log p_{\boldsymbol{\theta}}\left(x_i\right) \partial_\nu \log p_{\boldsymbol{\theta}}\left(x_i\right)~,\] equivalently as the sample average of outer products of the score function. It captures how sensitively the model distribution changes under infinitesimal parameter variations and is used in natural-gradient methods to precondition updates according to the local information geometry of the model. In quantum variational settings, a related object is the quantum Fisher information matrix (QFIM), which measures the distinguishability of nearby parameterized quantum states. For a normalized pure-state family \(\{\ket{\psi(\boldsymbol{\theta})}\}\), writing \(\ket{\partial_\mu\psi}:=\partial_\mu\ket{\psi(\boldsymbol{\theta})}\), one common convention is \[\begin{align} F^{\mathrm{Q}}_{\mu\nu}(\boldsymbol{\theta}) &=4\,\mathrm{Re}\!\Big[ \braket{\partial_\mu\psi|\partial_\nu\psi} \\ &\qquad -\braket{\partial_\mu\psi|\psi}\braket{\psi|\partial_\nu\psi} \Big]. \end{align}\] which is four times the Fubini–Study metric, or equivalently four times the real part of the quantum geometric tensor [32]. Thus, the empirical Fisher \(\tilde{F}\) used in supervised-learning losses and the QFIM used in quantum natural-gradient methods are distinct but analogous metric-like objects.
Several AI-for-quantum settings are sequential decision problems. Reinforcement learning (RL) models an agent interacting with an environment through states, actions, transitions, and rewards, often formalized as a Markov decision process with transition law \(P_{\mathrm{RL}}(z_{t+1}\mid z_t,a_t)\), reward \(r_t\), and discount factor \(\gamma_{\mathrm{RL}}\). A policy \(\pi(a\mid z)\) maps an observed state or belief state \(z\) to an action distribution, and training seeks a policy with high expected return \[J(\pi) = \mathbb{E}_{\pi}\!\left[ \sum_{t\ge 0}\gamma_{\mathrm{RL}}^{t} r_t \right].\] For a fixed policy, the value function satisfies the Bellman relation \[V^{\pi}(z) = \mathbb{E}_{a\sim\pi(\cdot\mid z),\,z'\sim P_{\mathrm{RL}}(\cdot\mid z,a)} \!\left[ r(z,a)+\gamma_{\mathrm{RL}}V^{\pi}(z') \right].\] When the agent observes only partial information, the problem is naturally partially observable; this is common in quantum networking, online calibration, and feedback control. Agentic AI systems extend this closed-loop view to language-model-driven workflows: a model or collection of models plans, calls tools, observes outputs, updates memory, and revises actions over multiple steps [1], [2]. This terminology is used below for autonomous quantum programming, self-driving laboratories, and formal-verification workflows.
| Symbol | Meaning | Remarks |
|---|---|---|
| \(n\) | number of qubits | The Hilbert-space dimension typically scales as \(2^n\). |
| \(\mathcal{H}\) | Hilbert space | For \(n\) qubits, \(\mathcal{H}=(\mathbb{C}^{2})^{\otimes n}\). |
| \(\rho\) | density operator / quantum state | The central latent object in state-learning and variational settings. |
| \(\rho_A\) | reduced state of subsystem \(A\) | Obtained from \(\rho_{AB}\) by partial trace over subsystem \(B\). |
| \(O\) | Hermitian observable | Expectation values are written as \(\langle O\rangle_\rho=\Tr(O\rho)\). |
| \(\{M_s\}\) | POVM | The measurement operators associated with classical outcomes \(s\). |
| \(s\) | classical measurement outcome | Drawn according to the Born-rule distribution \(p_\rho(s)\). |
| \(p_\rho(s)\) | measurement-outcome distribution | Defined by \(p_\rho(s)=\Tr(M_s\rho)\). |
| \(\Tr\) | trace | Used for expectations, Born probabilities, and reduced states. |
| \(H\) | Hamiltonian | Generator of closed-system unitary dynamics. |
| \(\mathcal{L}\) | Liouvillian / Lindbladian generator | Reserved throughout the paper for open-system continuous-time dynamics. |
| \(\mathcal{E}\) | quantum channel | A completely positive trace-preserving map. |
| \(U\) | unitary evolution or circuit | Includes both dynamical evolution and parameterized circuits. |
| \(t\) | evolution time | Used in dynamical prediction and simulation formulas. |
| \(\boldsymbol{\theta}\) | trainable parameter vector | Used for PQCs, Hamiltonian models, and other parameterized hypotheses. |
| \(p\) | number of trainable parameters | The length of \(\boldsymbol{\theta}\) when that quantity is used explicitly. |
| \(x\) | classical input / data point | Used in quantum feature maps, kernels, and supervised learning models. |
| \(y_i\), \(\mathbf{y}\) | target label and label vector | \(\mathbf{y}=(y_1,\ldots,y_N)^\top\) in supervised settings. |
| \(\mathcal{D}\) | data distribution | Used for population-risk and generalization statements. |
| \(\mathcal{S}_{\mathrm{tr}}\) | training set | Usually written as \(\{(x_i,y_i)\}_{i=1}^{N}\) in supervised settings. |
| \(\ell\) | per-example loss | Used to define empirical and population risks. |
| \(\widehat R\), \(R\), \(\mathcal{G}\) | empirical risk, population risk, and generalization gap | The gap is written as \(\mathcal{G}=R-\widehat R\) when it is useful to name it. |
| \(d_{\mathrm{VC}}\) | VC dimension | Used for PAC-style sample-complexity statements; other sections use quantities such as margins, shadow norms, or active-parameter counts. |
| \(\lambda_{\mathrm{reg}}\), \(\operatorname{Reg}\) | regularization strength and penalty | Used in classical and quantum learning objectives when structural bias is imposed explicitly. |
| \(\tilde y_j\) | observed measurement-derived datum | Used for experimentally observed classical values in system-identification losses. |
| \(y_j(\theta)\) | model prediction for datum \(j\) | The prediction generated by a parameterized Hamiltonian or Liouvillian model. |
| \(f(\boldsymbol{\theta})\), \(f_{\boldsymbol{\theta}}(x)\) | model output or predictor | May denote an expectation-valued model or a supervised predictor, depending on context. |
| \(p_{\boldsymbol{\theta}}(x)\) | learned data distribution | Used for generative models, including Born-machine and diffusion-style settings. |
| \(\mathscr{L}\) | loss or optimization objective | Reserved throughout the paper for training losses, empirical risks, and related objectives. |
| \(\pi(a\mid z)\) | policy in sequential decision-making | Maps an observed state or belief state \(z\) to an action distribution. |
| \(z_t\), \(a_t\), \(r_t\) | state or observation, action, and reward | Used in RL, feedback control, and quantum-network decision problems. |
| \(P_{\mathrm{RL}}\), \(J(\pi)\), \(V^\pi\) | RL transition law, return, and value function | Used for MDP and POMDP formulations of control and networking. |
| \(\gamma_{\mathrm{RL}}\) | RL discount factor | Weights future rewards; the subscript avoids conflict with other local uses of \(\gamma\). |
| \(\mathbf{F}\) | quantum Fisher information matrix (QFIM) | Appears in quantum natural-gradient methods. |
| \(\mathcal{Q}_{\mu\nu}\) | quantum geometric tensor (QGT) | Its real part gives the Fubini–Study metric tensor. |
| \(g_{\mu\nu}\) | Fubini–Study metric tensor | Defines the local geometry of the variational state manifold. |
| \(k(x,x')\) | kernel function | Usually estimated from quantum feature states or overlaps. |
| \(\mathbf{K}\), \(\boldsymbol{\alpha}\) | Gram matrix and kernel coefficients | \(\mathbf{K}_{ij}=k(x_i,x_j)\); \(\boldsymbol{\alpha}\) denotes the coefficients in a kernel predictor. |
| \(\Theta^{\mathrm{NTK}}_{ij}\), \(\Theta_{ij}(\boldsymbol{\theta})\) | NTK and QNTK entries | Measure parameter-space sensitivity correlations between inputs \(x_i\) and \(x_j\). |
| \(|\psi\rangle\) | pure quantum state vector | Used for pure states; mixed states are represented by \(\rho\). |
| \(N\) | system size or dataset size | Denotes matrix dimension in QLSA contexts or sample count in learning contexts. |
| \(\mathcal{O}(\cdot)\), \(\Omega(\cdot)\) | asymptotic upper and lower bounds | Standard Bachmann–Landau notation for complexity scaling. |
| \(\lambda_i\) | eigenvalue | Eigenvalues of a matrix or Hamiltonian; central to QLSA discussions. |
| \(D\) | problem or Hilbert-space dimension | Used for the ambient dimension when that quantity is global to the problem. |
| \(\varepsilon\) | target accuracy / approximation error | Used in complexity and approximation guarantees. |
| \(\delta\) | failure probability / confidence parameter | Used in sample-complexity and concentration bounds. |
| \(|b\rangle\), \(|x\rangle\) | encoded input and solution states | Standard notation in quantum linear-system algorithms. |
| \(U_A\) | block-encoding of matrix \(A\) | Satisfies \((\langle 0| \otimes I)\, U_A\, (|0\rangle \otimes I)=A/\alpha\). |
| \(\kappa\) | condition number | Governs the complexity of quantum linear-system algorithms. |
| \(S(\rho_A)\) | von Neumann entanglement entropy | Used in the discussion of entanglement structure and scaling laws. |
| \(\mathcal{T}\), \(A^{(k)}\), \(r_k\), \(\chi\) | tensor-network object, local cores, bond dimensions, and truncation rank | Used in the tensor-network section; these symbols are fixed there with their standard meanings. |
4pt
Section 2.1 emphasized physical states, measurements, noise, and variational learning on quantum hardware. For quantum methods in AI, quantum mechanics plays a different role. In the algorithmic part of the review, QC is treated as a computational model for AI-relevant linear algebra, learning, and optimization tasks; later subsections also consider quantum-inspired models and analyses. The key questions are how classical data are encoded into quantum states, how matrices and operators are accessed, and what information can actually be extracted from the quantum output.
A common primitive is state preparation. A vector \(\mathbf{b}\in\mathbb{C}^N\) may be encoded as the normalized quantum state \[|b\rangle=\frac{1}{\|\mathbf{b}\|}\sum_{i=1}^{N} b_i |i\rangle .\] This amplitude-encoding viewpoint allows an \(n=\log_2 N\) qubit register to represent an \(N\)-dimensional vector compactly. At the same time, it makes clear why data loading is such a central issue in QC-for-AI: a speedup based on manipulating \(|b\rangle\) is meaningful only if the state can itself be prepared efficiently, or if one is given an input model that grants comparably efficient oracle access [33].
A second recurring primitive is matrix access. Quantum algorithms typically access a matrix \(A\) through sparse-access queries or a block-encoding \(U_A\) satisfying \[(\langle 0| \otimes I)\, U_A\, (|0\rangle \otimes I) = A/\alpha\] for some normalization factor \(\alpha\) [34], [35]. Once such access is available, techniques such as Hamiltonian simulation, linear combination of unitaries (LCU), and quantum singular value transform (QSVT) can implement matrix functions \(f(A)\), turning inversion, spectral filtering, and related operations into reusable quantum subroutines.
Finally, quantum outputs are usually not full classical objects. What one often obtains is a quantum state, a sample distribution, or an expectation value such as \(\langle O\rangle_{\psi}=\langle \psi|O|\psi\rangle\). Recovering all \(N\) entries of a vector or all entries of a matrix would typically destroy a polylogarithmic speedup. For that reason, the complexity claims in Section 4 must always be read together with their access and readout assumptions, not in isolation.
These computational preliminaries explain the organization of Section 4. Some subsections study quantum algorithmic speedups for linear algebra and optimization, some ask when quantum data or quantum feature maps yield genuine learning advantages, and others analyze variational or tensor-network models inspired by quantum structure. What unifies them is that QC enters as a structured linear-algebraic resource whose power depends critically on state preparation, operator access, and measurement-limited readout.
A central topic in AI for QI is learning quantum systems from measurement data. In realistic experiments, the state \(\rho\) and its governing dynamics, such as a Hamiltonian \(H\) for closed systems or a Liouvillian generator \(\mathcal{L}\) for open systems, are not observed directly. They are inferred from finite, noisy, and often incomplete classical outcomes produced by quantum measurements. Using the measurement formalism introduced in Sec. 2.1, a POVM \(\{M_s\}\) generates an outcome \(s\) according to \[\require{physics} s \sim p_{\rho}(s), \qquad p_{\rho}(s)=\Tr(M_s\rho),\] where \(p_{\rho}(s)\) is the probability of observing outcome \(s\) when the system is in state \(\rho\), and \(\require{physics} \Tr\) denotes the trace. This is the Born rule written in sampling form. For projective measurements \(M_s=\Pi_s\) and a pure state \(\rho=\ket{\psi}\bra{\psi}\), it becomes \[\require{physics} p_{\rho}(s)=\Tr(\Pi_s\rho)=\bra{\psi}\Pi_s\ket{\psi},\] and, for \(\Pi_s=\ket{s}\bra{s}\), further reduces to \(p_{\rho}(s)=|\langle s \mid \psi\rangle|^2\).
This sampling view turns quantum system identification into a statistical-inference problem: the learner must infer latent quantum structure from finite data under measurement constraints. It also connects state and process inference to quantum learning theory, which studies what quantum objects can be learned, under which access models, and with what resource costs [16]. Therefore the organizing principle is statistical learning from measurement data. The analogy with classical statistical learning [26], [27] is therefore helpful but limited, because quantum measurements impose constraints that ordinary classical datasets do not. Table 2 summarizes this correspondence.
We use this lens to separate three related goals. First, state reconstruction aims to learn a representation of \(\rho\) itself. Second, property prediction estimates selected observables without reconstructing the full density matrix. Third, dynamical learning infers Hamiltonians, Liouvillians, or channels from input-output or time-resolved data. The discussion below first treats state reconstruction and property prediction, using quantum state tomography, structured ansätze, and classical-shadow methods as representative examples. It then turns to dynamical learning, including Hamiltonian, Liouvillian, and channel learning.
| Statistical learning | Quantum system inference | Example |
|---|---|---|
| Unknown target | Quantum state, dynamical generator, or channel | \(\rho\), \(H\), \(\mathcal{L}\), \(\mathcal{E}\) |
| Data distribution | Distribution from measurement | \(p_{\rho}(s)=\Tr(M_s\rho)\) |
| Sample | Single-shot measurement outcome | \(s \sim p_{\rho}(s)\) |
| Sufficient statistic / summary | Classical representation derived from measurement outcomes | Pauli expectation values, classical shadows |
| Hypothesis class | Structured family of candidate quantum objects | Rank-\(r\) states, \(k\)-local Hamiltonians, tensor-network states, sparse Lindbladians, Pauli channels |
| Prediction target | Expectation value or dynamical prediction from the learned object | \(\Tr(O\rho)\), \(\Tr\!\big[O\,e^{-iHt}\rho\, e^{iHt}\big]\), \(\Tr\!\big[O\,e^{t\mathcal{L}}(\rho)\big]\) |
| Loss / risk | Discrepancy between predicted and observed measurement outcomes | Squared error, negative log-likelihood, trace distance |
| Sample complexity | Number of state copies or measurement shots needed for a target accuracy | \(\Theta(4^n/\epsilon^2)\) for full tomography; \(O(\max_i\|O_{i,0}\|_{\mathrm{shadow}}^2\log(M/\delta)/\epsilon^2)\) for fixed-observable prediction, with \(O_{i,0}=O_i-\Tr(O_i)\mathbb{I}/2^n\) |
| Inductive bias | Physical or structural constraint used to regularize inference | Locality, sparsity, positivity, low rank, bounded interaction range |
8pt
Quantum state tomography, classical shadows, and prediction of physical properties. The most direct learning task is full quantum state tomography: reconstruct an unknown state from measurement data. This task quickly becomes infeasible for many-body systems. An \(n\)-qubit state is described by a \(2^n \times 2^n\) density matrix, so a generic state has on the order of \(4^n\) real parameters. Correspondingly, a standard upper-bound scaling for the copy complexity of full reconstruction is \[N_{\mathrm{full}} = O\!\left(\frac{4^n}{\epsilon_{\mathrm{rec}}^2}\right),\] where \(\epsilon_{\mathrm{rec}}\) is the target reconstruction error in trace distance, i.e., one seeks an estimate \(\hat{\rho}\) such that \(\tfrac{1}{2}\|\rho-\hat{\rho}\|_1 \le \epsilon_{\mathrm{rec}}\) [36].
This scaling motivates tomography methods that exploit structure. A standard example is compressed sensing. For rank-\(r\) states, the number of randomly chosen measurement settings or Pauli expectation values can scale as [37], [38] \[m_{\mathrm{CS}} = O\!\big(r D\, \mathrm{polylog}(D)\big),\] where \(m_{\mathrm{CS}}\) is the compressed-sensing measurement complexity, \(r\) is the rank of \(\rho\), \(D=2^n\) is the Hilbert-space dimension, and \(\mathrm{polylog}(D)\) denotes a polynomial in \(\log D\). Here \(m_{\mathrm{CS}}\) counts settings or expectation values rather than total experimental shots; finite-sample implementations also include accuracy and confidence factors. The important point is that, under a low-rank assumption, the setting count can be reduced relative to generic \(O(D^2)\)-parameter tomography when \(r \ll D\). Other structured reconstructions, including tensor-network and locally constrained tomography, use entanglement or locality assumptions instead of low rank in the full Hilbert space [39], [40]. These methods are examples of inductive bias.
Many experiments do not require a full density matrix. They ask instead for selected properties of \(\rho\), such as an energy \(\require{physics} E=\Tr(H\rho)\), a magnetization or local order parameter \(\require{physics} m=\Tr(M\rho)\), or a two-point correlation function \(\require{physics} C_{ij}=\Tr(O_{ij}\rho)\). Shadow tomography formalizes this property-prediction task [41]: the goal is to estimate many specified observables without reconstructing the entire state. The classical-shadow framework of [42] gives a practical randomized-measurement procedure for this task. Each randomized measurement produces a single classical snapshot, and the collection of snapshots can be reused to estimate many observables with provable guarantees.
Concretely, a classical-shadow snapshot \(\hat{\rho}\) is designed to be an unbiased estimate of the unknown state: \(\mathbb{E}[\hat{\rho}]=\rho\). Therefore, for any observable \(O\), the quantity \(\require{physics} \hat{o}:=\Tr(O\hat{\rho})\) is a noisy one-shot estimate of the desired value \(\require{physics} \Tr(O\rho)\). This leads to the central question: how many such snapshots are needed before the averaged estimates are accurate for all target observables? In the resulting sample bound, only the traceless part of each observable matters. The identity component is already known from \(\require{physics} \Tr(\rho)=1\): if \(\require{physics} O_{i,0}:=O_i-\Tr(O_i)\mathbb{I}/2^n\), then \(\require{physics} \Tr(O_i\rho)=\Tr(O_{i,0}\rho)+\Tr(O_i)/2^n\). Thus only \(O_{i,0}\) contributes to the statistical difficulty. For a target set \(\{O_i\}_{i=1}^M\), the classical-shadow theorem gives the sample bound [42] \[N_{\mathrm{shadow}} = O\!\left( \frac{\log(M/\delta)}{\epsilon_{\mathrm{obs}}^2} \max_i \|O_{i,0}\|_{\mathrm{shadow}}^2 \right),\] where \(N_{\mathrm{shadow}}\) is the number of copies or randomized measurements needed to estimate all \(\require{physics} \Tr(O_i\rho)\) to additive error \(\epsilon_{\mathrm{obs}}\) with success probability at least \(1-\delta\).
The task-dependent quantity in this bound is the shadow norm. It measures how much a one-shot estimate can fluctuate under the chosen randomized-measurement ensemble. A convenient expression is \[\require{physics} \|O\|_{\mathrm{shadow}}^2 := \max_{\sigma}\, \mathbb{E}_{\sigma}\!\left[\hat{o}^{\,2}\right], \qquad \hat{o}:=\Tr(O\hat{\rho}),\] where \(\mathbb{E}_{\sigma}\) means that the snapshot is generated from input state \(\sigma\), including the randomness in the unitary, the measurement outcome, and the construction of \(\hat{\rho}\) [42]. The maximization over \(\sigma\) makes the quantity a worst-case guarantee, independent of the unknown state. In words, a small shadow norm means that a single snapshot gives a relatively stable estimate. A large shadow norm means that the one-shot estimates have larger worst-case fluctuations, so more independent snapshots must be averaged to reach the same target error \(\epsilon_{\mathrm{obs}}\) and failure probability \(\delta\).
A useful way to interpret this abstract norm is to look at the measurement ensemble most often used in near-term experiments: random single-qubit Pauli measurements. In this protocol, each qubit is independently measured in a randomly chosen \(X\), \(Y\), or \(Z\) basis. A Pauli string is a tensor product of single-qubit Pauli operators and identities, such as \(Z_1X_3Y_7\). Its weight \(k\) is the number of non-identity factors, equivalently the number of qubits on which the operator acts nontrivially. This set of qubits is the support of the Pauli string. For this measurement ensemble, a weight-\(k\) Pauli string \(O\) with eigenvalues \(\pm1\) has \(\|O\|_{\mathrm{shadow}}^2 = 3^{k}\) [42]. Combining this with the sample bound above shows why the measurement cost depends on \(k\) rather than directly on the total number of qubits \(n\): one pays for the number of qubits involved in the observable, not for all qubits in the device. Thus a two-qubit correlation can remain cheap to estimate even in a large system, whereas a Pauli string acting nontrivially on \(O(n)\) qubits can still require exponentially many snapshots. For observables that are sums of many Pauli terms, the coefficients and number of terms enter through the shadow norm. Thus classical shadows give dimension-efficient prediction for structured observable families, not a dimension-free solution to arbitrary state reconstruction.
The same shadow representation can also be used as input to classical machine learning models [43], [44]. Instead of reconstructing one particular state, one converts measurement data from many related quantum states into compact classical features and trains a supervised model to predict physical properties of unseen states drawn from the same family. Under the distributional and property-class assumptions in these works, the sample complexity is controlled by the structure of the target property and need not scale with the full Hilbert-space dimension [42], [43]. This should be read as a structured property-learning guarantee, not as a dimension-free guarantee for arbitrary properties.
Neural-network quantum state tomography takes a complementary route. Rather than fixing a linear shadow estimator, it parametrizes the state itself with a flexible model. In neural-network quantum state tomography [45], for example, a restricted Boltzmann machine or related neural ansatz is trained directly on measurement outcomes to approximate the target state. Such models can capture states that are not well matched to simple low-rank or sparse ansätze, but they typically do not provide the same general closed-form statistical guarantees as compressed-sensing or shadow-based frameworks. Neural-shadow methods [46] aim to combine these advantages by using shadow-derived losses together with neural state models, retaining some measurement-efficiency benefits of shadows under controlled assumptions while increasing representational flexibility.
Learning dynamical laws from measurement data. Static inference is only part of the problem. Many experiments aim to learn the law that generates the observed dynamics. In the simplest closed-system setting, one seeks an unknown Hamiltonian \[H(\boldsymbol{\theta}) = \sum_{\mu=1}^{p} \theta_\mu\, P_\mu,\] where \(\{P_\mu\}\) is a known operator basis and the coefficients \(\theta_\mu\) are inferred from time-resolved measurement data. This expansion also shows why structure is essential. If the basis is unrestricted, then \(p\) can be \(\Theta(4^n)\), so full identification generically has exponential parameter and sample requirements. Physically relevant Hamiltonians typically have locality, sparsity, bounded interaction range, or related structure that reduces the effective complexity. Local-Hamiltonian learning exploits such assumptions [47]. For certain geometrically local Hamiltonians, protocols [48], [49] achieve sample complexity polynomial in the system size \(n\) and inverse target accuracy, under specified locality, control, and measurement-access assumptions, using experimentally accessible product-state preparations and local measurements.
Practical Hamiltonian-learning methods differ mainly in how data are chosen and how candidate models are fit. Bayesian and adaptive protocols [50] choose input states, evolution times, and measurements to distinguish competing Hamiltonian hypotheses, then update a posterior distribution after each round of data. Simulator-based fitting protocols compare measured data with predictions from parameterized models, using energies, variances, equilibrium states, or dynamical responses; these are best viewed as model calibration or validation unless the identifiability assumptions are specified [51], [52].
A common supervised formulation casts Hamiltonian learning as empirical risk minimization over dynamical data. Given input states \(\rho_j\), evolution times \(t_j\), measured observables \(O_j\), and measured expectation values \(\tilde{y}_j\), a candidate Hamiltonian \(H(\theta)\) predicts \[\require{physics} y_j(\theta)= \Tr\!\left[ O_j\, e^{-iH(\theta)t_j}\rho_j e^{iH(\theta)t_j} \right],\] for datum \(j\). One then minimizes \[\mathscr{L}_{\mathrm{Ham}}(\theta) = \frac{1}{N}\sum_{j=1}^{N} \left( \tilde{y}_j-y_j(\theta) \right)^2 .\] Here the structured family \(\{H(\theta)\}\) is the hypothesis class, and \(y_j(\theta)\) is the model prediction for the \(j\)th experiment.
Open-system learning follows the same template, but the physical constraints are stronger. For Markovian dynamics, the candidate Hamiltonian is replaced by a parameterized Liouvillian \(\mathcal{L}_{\theta}\). To ensure that \(e^{t\mathcal{L}_{\theta}}\) is a physical channel, the hypothesis class is usually restricted to a GKLS/Lindblad form with positive rates, or to another parametrization that guarantees complete positivity and trace preservation. The prediction becomes \[\require{physics} y_j(\theta)= \Tr\!\left[ O_j\, e^{t_j\mathcal{L}_{\theta}}(\rho_j) \right],\] with empirical risk \[\mathscr{L}_{\mathrm{open}}(\theta) = \frac{1}{N}\sum_{j=1}^{N} \left( \tilde{y}_j-y_j(\theta) \right)^2 .\] For time-ordered trajectories or repeated-shot outcome counts, the squared-error loss is usually replaced by a negative log-likelihood derived from the corresponding probabilistic model. This formulation also makes robustness part of the learning problem. Dissipation, decoherence, and state preparation and measurement (SPAM) errors can be statistically confounded with the effective dynamics unless they are modeled jointly or calibrated independently. Self-consistent characterization methods such as gate-set tomography address related SPAM and gauge-freedom issues [53], [54]. Recent large-scale examples include learning sparse Pauli-Lindblad models [55], weakly dissipative Liouvillian learning [56], and Lindblad learning from time-series data [57].
An even broader view treats the unknown object as a quantum channel \(\mathcal{E}\), i.e., a completely positive trace-preserving input-output map. Hamiltonian and Liouvillian dynamics are then structured special cases. As recalled in sSec. 2.1, closed-system evolution generated by \(H\) defines the channel \(\rho\mapsto e^{-itH}\rho e^{itH}\), while a continuous-time open-system model generated by \(\mathcal{L}\) induces the channel family \(\mathcal{E}_t=e^{t\mathcal{L}}\). General channel learning also covers effective noise, control imperfections, and black-box processes that may not admit a simple generator description. The prediction rule is \[\require{physics} y_j(\mathcal{E})=\Tr\!\left[O_j\, \mathcal{E}(\rho_j)\right],\] so the learner selects a channel hypothesis from finite input-output data. This channel-learning viewpoint connects quantum process tomography [58]–[60], shadow-based channel learning [61], and process learning without input control [62]. It is complementary to gate-set tomography and randomized benchmarking, which handle SPAM/gauge freedoms or report operational error rates rather than reconstructing a single CPTP map [54], [63]. Across these settings, the central object is an unknown input-output map that can extend beyond a generator of time evolution.
Across state, property, and dynamics learning, three issues recur. The first is the fundamental limit question: what can be learned, under which measurement model, and with how many samples? For full state tomography, near-optimal worst-case bounds [36] clarify the copy complexity needed to reconstruct an entire unknown state to a target accuracy. Shadow tomography [41], classical shadows [42], and learning from noisy quantum experiments [64] give complementary limits for more restricted prediction tasks. More broadly, learnability can depend sharply on the allowed data-collection operations. Exponential separations can appear between protocols that allow joint measurements or quantum memory and protocols that measure each copy separately and process only classical outcomes [64], [65].
The second issue is data acquisition. In many classical learning problems, the learner receives a fixed dataset. Quantum experiments often allow partial control over how data are gathered. Adaptive and active learning protocols [50], [66], [67] exploit this control by choosing measurement bases, probe states, or evolution times online to maximize information gain or reduce posterior uncertainty. Such protocols are especially natural in tomography and Hamiltonian learning, where measurement design can strongly affect statistical efficiency.
The third issue is robustness. Noise and SPAM errors are not merely experimental nuisances; they are part of the learnability question. Under specified noise and measurement-depth assumptions, work on learning from noisy quantum experiments and on robust ultra-shallow or shallow-shadow protocols [64], [68], [69] shows how shadow-based methods can preserve useful bias, variance, and sample-complexity control. These results also clarify how noise changes estimator bias, variance, and feasible measurement depth. The common message is that learnability depends jointly on the target task, the structure of the quantum object, the measurement model, and the available computational resources.
A central challenge in QC is optimization over high-dimensional design spaces, especially variational circuits such as VQEs [70] and QAOA [71]. These problems are intrinsically hard for two reasons: First, their loss landscapes are often nonconvex and can exhibit barren plateaus [72], [73]. Second, even when useful directions exist, finite-shot estimation and hardware noise make them difficult to resolve [74]. This subsection uses AI theory to explain the learning dynamics of these variational quantum algorithms (VQAs). We therefore treat a VQA as a learnable system. The four themes below develop this viewpoint by examining whether usable optimization signal exists, how quantum geometry should shape updates, what effective function class the model induces, and how finite shots and hardware noise change the picture [19], [23], [24], [32], [74]–[79].
In classical machine learning, trainability usually refers to whether an optimizer can find parameters with low training loss using feasible computational resources The same idea applies to VQAs and QML, but quantum training has additional constraints. The ansatz and cost construction shape the loss landscape, expectation values and gradients must be estimated from finitely many measurement shots, and the data encoding can change gradient scaling or even create dataset-dependent optimization pathologies [80]. In the gradient-based setting considered here, a model is trainable when useful optimization signals remain resolvable and the total resources needed to find good parameters scale at most polynomially with problem size and target precision.
Barren Plateau. The central threat to trainability in quantum settings is the barren plateau (BP) phenomenon [73]. The BP phenomenon can be made precise by looking at gradients at random initializations. If, for every trainable parameter, these gradients have zero mean and an exponentially small variance as the number of qubits grows, then the cost function is said to exhibit a BP [73], [81]. Formally, for a parameterized cost function \(C(\boldsymbol{\theta})\), this means that the following two conditions hold for all variational parameters \(\theta_\nu \in \boldsymbol{\theta}\), with parameters \(\boldsymbol{\theta}\sim p(\boldsymbol{\theta})\) drawn from a specified initialization distribution: \[\begin{align} &\mathbb{E}_{\boldsymbol{\theta}\sim p(\boldsymbol{\theta})}[\partial_\nu C(\boldsymbol{\theta})] = 0, \tag{1}\\ &\mathrm{Var}_{\boldsymbol{\theta}\sim p(\boldsymbol{\theta})}[\partial_\nu C(\boldsymbol{\theta})] \in \mathcal{O}(1/\alpha^n), \quad \alpha > 1, \tag{2} \end{align}\] where \(\partial_\nu C(\boldsymbol{\theta}) = \partial C(\boldsymbol{\theta})/\partial \theta_\nu\), \(n\) is the number of qubits, and \(\alpha>1\) is a constant independent of \(n\). In many standard random-initialization settings, including common Pauli-rotation parameterizations with symmetric sampling, the mean gradient is zero [73], [81]. Together with the exponentially vanishing variance, this implies that for typical initializations the gradients are exponentially concentrated near zero as the system size grows [73], [80], [81]. This exponential suppression renders gradient-based optimization ineffective, as an exponential number of measurement shots is needed to resolve the vanishingly small gradient signal from statistical noise [80], [82].
[80] show that BP results also apply to supervised QML models. Under suitable assumptions, if the underlying linear expectation values exhibit a BP, then common QML losses such as mean squared error and negative log-likelihood inherit the same exponential gradient suppression. Familiar BP mechanisms, including global cost functions and deep unstructured circuits, therefore remain problematic in supervised learning settings. The same work identifies a second failure mode termed the dataset-induced BP: highly entangling embeddings can make local reduced states nearly maximally mixed, causing gradients to concentrate even when the measurement is local and the QNN is shallow. The same work also identifies a related issue: the curvature information used by natural-gradient methods can suffer from the same problem. Under the negative-log-likelihood assumptions of [80], the empirical Fisher entries obey \[\mathbb{E}_{\boldsymbol{\theta}\sim p(\boldsymbol{\theta})} \!\left[ |\tilde{F}_{\mu\nu}(\boldsymbol{\theta})| \right] \in \mathcal{O}(1/\alpha^n),\] where \(\tilde{F}_{\mu\nu}(\boldsymbol{\theta})\) is the \((\mu,\nu)\) entry of the empirical Fisher information matrix, \(p(\boldsymbol{\theta})\) is the initialization distribution, \(n\) is the number of qubits, and \(\alpha>1\) is independent of \(n\). Exponentially small Fisher entries are exponentially hard to estimate reliably, which limits the usefulness of natural-gradient preconditioning in this regime.
Later work extends the above conclusions beyond gradient-based QML. [74] prove analogous exponential concentration in quantum kernel methods, and [83] report predictive collapse in deep data re-uploading models on high-dimensional data. Symmetry or covariance can avoid or mitigate such concentration in structured settings [84], [85]. Overall, BPs are a family of mechanisms controlled by the ansatz, observable or cost construction, initialization, locality, and noise processes. A recent survey [20] synthesizes these factors together with the main mitigation strategies.
Leveraging Dynamical Lie Algebras and Adjoint Representations to Characterize BPs. Recent works highlight a shift from empirical diagnosis toward criteria for predicting when BPs arise and how their severity scales with system size [19], [20]. Two representative directions are as follows. First, the dynamical Lie algebra (DLA) viewpoint [75] characterizes an ansatz by the smallest Lie algebra generated by the circuit’s elementary gate generators under commutators. Using this object, [75] derive variance expressions for sufficiently deep circuits and show that BP onset can be inferred from the algebraic growth of the ansatz. Second, the adjoint-representation viewpoint [76] analyzes the same algebraic structure in the Heisenberg picture, where the circuit transforms observables by conjugation, \(O \mapsto U^\dagger O U\). For the class they call Lie-algebra-supported ansätze, in which the relevant observable is supported on the dynamical Lie algebra \(\mathfrak{g}\), [76] express the gradient variance in terms of the decomposition \(\mathfrak{g}=\bigoplus_\alpha \mathfrak{g}_\alpha\oplus\mathfrak{c}\) into simple ideals and a center. In the Haar or 2-design setting, their result has the form \[\mathrm{Var}[\partial_\nu \langle O\rangle_{\rho_{\mathrm{in}}}] = \sum_{\alpha} \frac{ \|G_{\nu,\alpha}\|_{\mathrm{K}}^2 \|O_{\alpha}\|_{\mathrm{F}}^2 \|\rho_{\mathrm{in},\alpha}\|_{\mathrm{F}}^2 }{d_{\mathfrak{g}_\alpha}^2},\] where \(G_{\nu,\alpha}\), \(O_\alpha\), and \(\rho_{\mathrm{in},\alpha}\) denote the components of the parameter generator, observable, and input state associated with the ideal \(\mathfrak{g}_\alpha\), and \(d_{\mathfrak{g}_\alpha}=\dim(\mathfrak{g}_\alpha)\). This formula makes the intuition concrete: gradients become small when the relevant support is spread over exponentially large algebraic components, or when the generator, state, and observable have little overlapping support on the same components. When locality, symmetry, or problem structure keeps these components small and aligned, useful gradient signal can persist.
Constructive BP-free regimes and mitigation strategies. Several recent works identify constructive BP-free regimes by adding architectural or initialization structure to expressive ansätze. [86] show that a Hamiltonian variational ansatz (HVA) can avoid BPs while staying close to the underlying Hamiltonian structure. [87] demonstrate BP-free behavior in finite local-depth circuits (FLDCs) that can still support long-range entanglement, reinforcing the idea that locality-structured architectures can remain trainable at scale. [88] propose a related entanglement-based strategy that steers the effective ensemble away from overly mixing regimes.
Beyond such BP-free regimes, practical mitigation methods target initialization, sparsity, and optimizer design. BEINIT [89] initializes gate parameters from a data-dependent Beta distribution and adds perturbations during gradient descent, empirically reducing the chance that a QNN becomes trapped in a BP. QAdaPrune [90] adaptively prunes redundant or weakly contributing variational parameters. The resulting sparse parameter sets can match the unpruned circuit and sometimes improve trainability when the original circuit stalls.
Gradient-free training provides another optimizer-level mitigation route. Although BP-induced concentration can also limit gradient-free methods [82], the learned meta-optimizer of [91] proposes QNN parameters directly without estimating gradients on the quantum device and reports improved minima with fewer circuit evaluations in the studied settings. More generally, [92] construct a modified parameterized quantum circuit by inserting a trainable gadget layer into an existing circuit. Each gadget couples a system qubit to an ancilla initialized in \(\ket{0}\) through a single-qubit operation and three trainable two-qubit rotations, \(R_{XX}\), \(R_{YY}\), and \(R_{ZZ}\), after which the ancilla is discarded; the resulting ansatz is therefore a parameterized quantum channel rather than a purely unitary circuit. The construction preserves expressivity because setting the gadget parameters to zero recovers the original PQC, while suitable placement of the gadget layer yields inverse-polynomial lower bounds on the loss variance and on the gradient variance of parameters after the gadget layer for local observables. They also introduce an activation step that adds extra gadgets around parameters before the gadget layer when those parameters remain difficult to train.
Together, these works show that BP avoidance can come from architectural structure, initialization, sparsification, controlled perturbations, or optimizer design that limits unnecessary mixing while preserving task-relevant correlations.
Traps and Trainability Tradeoffs. BPs are only one aspect of optimization hardness. Another is the presence of traps, i.e., suboptimal local minima or other stationary regions where local optimization can stall even if gradients do not vanish exponentially. [93] show that BP landscapes can be “swamped with traps,” so avoiding vanishing gradients alone is not sufficient for trainability. At a higher level, [94] identify a tension between BP-free design and quantum advantage: under certain notions of provable BP absence, circuits may become easier to simulate classically. This suggests that trainability and possible quantum speedups should be analyzed together. [95] connect the discussion to modern AI theory by showing that broad classes of deep, random QNNs converge in function space to Gaussian processes, reinforcing a concentration-based view of trainability at large scale.
Taken together, the recent literature places BPs inside a broader picture of trainability. Besides, they also include traps and poor local minima, and the practical difficulty of resolving weak signals under finite measurement budgets [19], [20], [74], [93], [96]. From this perspective, an important goal for future research is to predict trainability before large-scale optimization begins.
This subsection focuses on one prominent class of optimization: geometry-aware optimization. We ask and survey how the geometry of the quantum state manifold can inform parameter updates. Natural-gradient preconditioning is the main example discussed below, while the broader geometric viewpoint also involves QGT/QFIM-based diagnostics, metric estimation, and geometry-driven variational dynamics. A parameterized quantum circuit (PQC) generates a smooth family of quantum states \(\ket{\psi(\boldsymbol{\theta})}\), so optimization can be viewed in parameter space and on the underlying manifold of quantum states [32], [97], [98]. For a normalized pure-state family \(\{\ket{\psi(\boldsymbol{\theta})}\}_{\boldsymbol{\theta}}\), with \(\boldsymbol{\theta}=(\theta_1,\ldots,\theta_p)\), the quantum geometric tensor (QGT) is defined as \[\mathcal{Q}_{\mu\nu}(\boldsymbol{\theta}) = \bra{\partial_\mu \psi(\boldsymbol{\theta})} \Bigl(I-\ket{\psi(\boldsymbol{\theta})}\bra{\psi(\boldsymbol{\theta})}\Bigr) \ket{\partial_\nu \psi(\boldsymbol{\theta})},\] where \(\ket{\partial_\mu \psi(\boldsymbol{\theta})}:=\partial \ket{\psi(\boldsymbol{\theta})}/\partial \theta_\mu\) [97], [98]. Its real part, \[g_{\mu\nu}(\boldsymbol{\theta}) := \mathrm{Re}\,\mathcal{Q}_{\mu\nu}(\boldsymbol{\theta}),\] is the Fubini–Study metric tensor, which measures the infinitesimal distance between neighboring quantum states on the state manifold. Equivalently, it defines the line element \(ds_{\mathrm{FS}}^2=\sum_{\mu,\nu} g_{\mu\nu} d\theta_\mu d\theta_\nu\). For pure states, this metric is directly related to the quantum Fisher information matrix (QFIM) [32] up to a constant factor. This geometric viewpoint leads to the quantum natural gradient (QNG) [32]: QNG rescales the gradient by the inverse QFIM so that the update follows steepest descent with respect to the state-space metric, schematically \(\Delta\boldsymbol{\theta}\propto -\mathbf{F}^{-1}\nabla_{\boldsymbol{\theta}}\mathscr{L}\), where \(\mathbf{F}\) denotes the QFIM and \(\mathscr{L}\) denotes the training loss, or more generally the optimization objective. The main intuition is that apparent plateaus in parameter space can partly reflect poor conditioning of the map \(\boldsymbol{\theta}\mapsto\ket{\psi(\boldsymbol{\theta})}\). Geometry-aware updates aim to correct this mismatch. The main drawback is cost: in both classical and quantum settings, natural-gradient methods require estimating, storing, and inverting a Fisher-type metric, which can become computationally expensive, measurement-expensive, or numerically unstable in high dimensions [32], [99], [100].
Recent work develops this geometric viewpoint along several complementary directions. On the optimizer-design side, qBang [78] is a representative recent example that combines QFIM or QGT-based preconditioning with momentum to accelerate navigation of flat energy landscapes in VQA objectives. On the estimation side, because forming a full QFIM can be measurement-expensive, structure-exploiting protocols are being developed. For example, [100] study commuting-block circuits, in which the relevant generators can be grouped into mutually commuting blocks, and propose a QFIM-estimation protocol that exploits this structure to reduce the number of required state preparations. Complementarily, robust hardware-facing methods for estimating Fisher-type geometric quantities have also advanced: [101] demonstrate robust estimation of the quantum Fisher information (QFI, i.e., the single-parameter or directional version of the QFIM) on a quantum processor, emphasizing robustness to experimental imperfections. More broadly, the same geometric ideas also appear in variational time-evolution methods. The connection to QITE is that exact imaginary-time evolution is generally nonunitary and leaves the variational ansatz class, so variational QITE projects this flow back onto the tangent space of the parameterized state manifold. This projection is measured with the Fubini–Study/QFIM geometry, leading to metric-dependent parameter updates closely related to QNG. [102] formulate variational quantum time evolution without computing the QGT explicitly, directly targeting a key bottleneck in geometry-driven updates. Along related lines, [103] develop an analytic theory of quantum imaginary time evolution (QITE) by making this connection explicit: they relate QITE, in the continuous-time variational limit, to quantum natural-gradient dynamics and use that geometric formulation to analyze its update dynamics. A separate but practically important question is whether geometry-aware optimization remains useful on noisy hardware. In that direction, [104] study QNG on noisy platforms using the quantum approximate optimization algorithm (QAOA) as a case study, providing evidence and diagnostics for the robustness of metric-preconditioned optimization under realistic noise.
The discussion above suggests several concrete future directions for geometry-aware optimization. First, optimizer design should move beyond applying QNG as a standalone preconditioner and study how geometric information can be combined with momentum, adaptive stepsizes, shot allocation, and other noise-aware optimization heuristics [78], [104]. Second, metric estimation remains a central practical challenge: future methods need cheaper ways to estimate, approximate, or avoid explicitly forming QGT/QFIM objects by exploiting circuit structure, locality, commuting blocks, or low-rank approximations [100]–[102]. Third, the link between geometry-aware optimization and variational time evolution deserves further clarification, especially for QITE and related methods where the update rule can be viewed as projecting nonunitary dynamics onto a parameterized state manifold [102], [103]. Finally, trainability claims for geometry-aware updates should be tested under finite shots, hardware noise, and numerical ill-conditioning, since the usefulness of geometric preconditioning ultimately depends on whether the relevant metric information is resolvable on realistic devices [101], [104].
Kernel methods provide another way to analyze quantum learning [28], [29]. For fixed kernels, standard models such as kernel ridge regression and support vector machines reduce training to convex optimization. In quantum kernel methods (QKMs), one defines a quantum feature map \(x \mapsto \ket{\phi(x)}\) via a state-preparation circuit. For normalized pure-state embeddings, a common choice is the fidelity kernel \[k(x,x') \;=\; \big|\langle \phi(x)\,|\,\phi(x')\rangle\big|^2 .\] This setup makes a uniquely quantum constraint explicit. Once a kernel \(k(\cdot,\cdot)\) is fixed, the downstream classical learning problem is standard and typically reduces to regularized convex optimization [28], [29]. In quantum kernel methods, the training kernel matrix or Gram matrix must be estimated from finitely many quantum measurements. As a result, statistical performance depends jointly on the inductive bias of the kernel and the shot complexity required to resolve pairwise similarities to useful precision [74].
Exponential concentration. A key recent result is that quantum kernels can exhibit exponential concentration with the number of qubits [74]. Let \(k(x,x')\) denote the kernel value for a pair of inputs \(x,x'\sim\mathcal{D}\) drawn from a data distribution \(\mathcal{D}\), and let \(n\) be the number of qubits in the corresponding encoded feature states. The kernel is said to be probabilistically exponentially concentrated [74] toward a data-independent value \(\mu\) if, for every fixed resolution threshold \(\delta>0\), \[\begin{align} \Pr_{x,x'\sim\mathcal{D}}\!\left[\,|k(x,x')-\mu|\ge \delta\,\right] &\le \frac{\beta}{\delta^2}, \\ \beta &\in \mathcal{O}(1/b^n), \qquad b>1, \end{align}\] where \(\mu\) is the concentration value, \(\beta\) controls the exponentially small tail, and the probability is taken over input pairs. As noted, an equivalent variance-based criterion is \[\mathrm{Var}_{x,x'\sim\mathcal{D}}[k(x,x')] \in \mathcal{O}(1/b^n),\] where the variance is again over input pairs [74]. For fidelity-type kernels, where \(0\le k(x,x')\le 1\), this means that kernel values over different input pairs become increasingly indistinguishable as \(n\) grows; in many relevant cases they concentrate around a constant or even an exponentially small value. If the intrinsic spread of \(k(x,x')\) across input pairs is exponentially small, then a polynomial number of measurement shots cannot resolve kernel values to a precision finer than that spread. Measurement noise then overwhelms the remaining data dependence, and the off-diagonal entries of the estimated Gram matrix become effectively indistinguishable unless one uses exponentially many measurements [74]. This loss of variation is especially damaging because kernel methods are completely determined by the Gram matrix \(\mathbf{K}\) on the training set. Once \(\mathbf{K}\) is fixed, training reduces to solving a convex problem (e.g., kernel ridge regression or support vector machines), and predictions for a new input depend only on its similarities to the training points [28], [29]. If the off-diagonal entries of \(\mathbf{K}\) become effectively data-independent under feasible shot budgets, the learned predictor becomes insensitive to input variation, yielding a “trivial” classifier or regressor even though the optimization itself is easy [74]. This highlights a uniquely quantum bottleneck: the estimability of the kernel matrix under finite-shot measurement constraints can limit performance even when the downstream optimization problem is convex [74], [84], [85].
Structured quantum kernels: symmetry, generalization, and implementability. Recent work treats concentration as a kernel-design problem. [105] argue that useful quantum kernels need problem-specific inductive bias together with large implicit feature spaces. Symmetry is one important source of such bias. If the data distribution or target function is invariant, or equivariant, under a group action, then a symmetry-aligned kernel can treat symmetry-related inputs consistently and suppress task-irrelevant variation. A covariant kernel [84], [85] implements this idea by making the feature map transform compatibly with the group action. Informally, if a group element \(g\) maps \(x\) to \(g\cdot x\), then the encoded features transform in a matched way and the kernel respects the symmetry, often through a relation such as \(k(g\!\cdot\! x,g\!\cdot\! x')=k(x,x')\). [85] and [84] show that such symmetry-structured kernels can mitigate or avoid exponential concentration while retaining trainability on structured data.
Once a kernel remains informative and estimable, the next question is generalization. A fixed quantum kernel leads to an ordinary kernel-regression or kernel-classification problem on its Gram matrix, so phenomena from classical kernel learning can reappear in quantum feature spaces. For a training set \(\{(x_i,y_i)\}_{i=1}^{N}\), a kernel predictor \(f(x)=\sum_{i=1}^{N}\alpha_i k(x,x_i)\) interpolates the data when there exists \(\boldsymbol{\alpha}=(\alpha_1,\ldots,\alpha_N)^\top\) such that \(f(x_i)=y_i\) for every training point. Equivalently, \(\mathbf{K}\boldsymbol{\alpha}=\mathbf{y}\), where \(\mathbf{y}=(y_1,\ldots,y_N)^\top\).
The regime at or beyond this interpolation threshold is often called overparameterized [106], [107]. In this regime, double descent [106], [107] means that test error can vary nonmonotonically with model complexity: it decreases at first, rises near the interpolation threshold, and then decreases again deeper in the interpolating regime. Figure 1 gives a schematic example. [106] demonstrate this behavior in quantum kernel methods, while [107] study benign overfitting with quantum kernels and provide numerical evidence that interpolating solutions can generalize well in suitable regimes.
A related theoretical question is representability [108]: can a desired kernel function be realized by a quantum feature map? The usual setup encodes each input \(x\) into a quantum state \(\ket{\phi(x)}\) and defines the kernel from state overlaps, for example \(k(x,x') = |\langle\phi(x)\mid\phi(x')\rangle|^2\). In this setting, [108] show that any kernel has an embedding-based quantum realization in principle, and they identify broad kernel classes with efficient embedding constructions.
Recent work extends this discussion in several directions. [109] introduce entangled tensor kernels, showing that embedding quantum kernels belong to a broader structural class that helps clarify their inductive bias and possible routes to dequantization. [110] study operator-valued quantum kernels for structured outputs and richer learning tasks. [111] propose neural quantum kernels, where a QNN is trained first and the learned representation is then turned into a problem-informed kernel.
These results make it useful to distinguish representability in principle from efficient implementability. The former asks whether some quantum embedding reproduces the desired kernel. The latter asks whether the embedding can be realized with feasible circuit depth, width, and measurement cost. A kernel can therefore be representable in theory while remaining impractical to implement or estimate. Quantum kernels should be evaluated along all three axes: representability, estimability under feasible shot budgets, and statistical behavior in the induced overparameterized regime.
Connection to tangent kernel theory: the Quantum Neural Tangent Kernel (QNTK). The kernel viewpoint also connects naturally to tangent-kernel analyses of VQAs, where one studies training dynamics through the local sensitivity of the model output with respect to parameters. In this literature, the Quantum Neural Tangent Kernel (QNTK) [112] is a central diagnostic that generalizes the neural tangent kernel philosophy [30], [113] from classical deep learning to PQCs. Following [112], [114], for a training set \(\{x_i\}_{i=1}^{N}\) of size \(N\), a parameter vector \(\boldsymbol{\theta}=(\theta_1,\ldots,\theta_p)\) with \(p\) variational parameters, and a model output \(f_{\boldsymbol{\theta}}(x)\), the QNTK is defined as the data-dependent kernel matrix \(\mathbf{\Theta}(\boldsymbol{\theta})\in\mathbb{R}^{N\times N}\) with entries \[\Theta_{ij}(\boldsymbol{\theta}) = \sum_{\ell=1}^{p} \frac{\partial f_{\boldsymbol{\theta}}(x_i)}{\partial \theta_\ell} \frac{\partial f_{\boldsymbol{\theta}}(x_j)}{\partial \theta_\ell},\] so each entry measures the similarity of the parameter-space sensitivities induced by inputs \(x_i\) and \(x_j\). It is sometimes useful to consider a scalar proxy for the overall gradient signal strength during training. For the same parameter vector \(\boldsymbol{\theta}\) and an error functional \(e(\boldsymbol{\theta})=\langle O\rangle - O_0\), where \(O\) is the measured observable and \(O_0\) is the target value, one such proxy at training step \(t\) is \[K(t) = \sum_{\ell=1}^{p} \left(\frac{\partial e(\boldsymbol{\theta})}{\partial \theta_\ell}\right)^2\bigg|_{\boldsymbol{\theta}=\boldsymbol{\theta}(t)},\] which is the squared \(\ell_2\)-norm of the gradient of the training error. Although this scalar quantity is not the full QNTK, it quantifies how strongly the error responds to joint perturbations across all parameters and therefore serves as a proxy for the rate at which gradient descent can reduce the error in early training [112], [114]. Figure 2 summarizes the two dynamical pictures that recur in the QNTK literature: an early-time frozen-kernel regime, where training is governed by an almost constant tangent kernel, and an evolving-kernel regime, where the kernel changes with the parameters and representation learning becomes possible.
QNTK theory began by adapting classical neural tangent kernel theory to VQAs. [112] introduced QNTK as a tool for analyzing training dynamics. [114] then developed an analytic theory for wide quantum neural networks in an overparametrized regime. In their setting, the number of parameters is large enough that the convergence rate satisfies \(\gamma=\mathcal{O}(1)\), which permits exact solutions for the training dynamics and a detailed description of the optimization trajectory. Related work uses the same tangent-space view to study parameter redundancy. [115] show that circuit symmetries can create redundant directions in parameter space and that symmetry-aware pruning can improve parameter efficiency without substantially degrading performance. [116] extends QNTK to graph-structured data through GraphQNTK.
Recent QNTK work studies concentration, noise, and training dynamics beyond the earliest regime. [117] use QNTK to relate quantum laziness, meaning exponential suppression in the number of qubits, to barren plateaus, meaning flat loss landscapes with vanishing gradients. They clarify the distinction between these phenomena and demonstrate noise resilience in overparametrized regimes. [118] show that circuit expressivity, the ability of a circuit family to represent diverse quantum states, can itself lead to QNTK concentration and thereby connect expressibility measures to training dynamics.
These results also make a practical point. Even when the tangent kernel predicts a useful optimization signal, that signal is built from gradients and Jacobian entries that must be estimated from finitely many measurements. If those quantities cannot be estimated accurately, the optimizer cannot reliably use the tangent-kernel information. QNTK-based learning therefore faces the same estimability constraint as overlap-based quantum kernels [74], [118].
Other work extends QNTK beyond the initial frozen-kernel picture. [119] identify a dynamical transition in deep QNNs when the target value crosses the minimum achievable value. Their analysis reveals a duality between QNTK and total-error dynamics described by generalized Lotka–Volterra equations. This transition separates frozen-kernel and frozen-error phases and gives a more detailed picture of late-time VQA training. [120] find similar transitions for quantum-data-driven learning, while [121] derive limits on learning random quantum data within the QNTK framework. Beyond standard VQA training, [103] develop a QNTK-based analytic theory for quantum imaginary time evolution. On the practical side, [122] use QNTK for QNN diagnostics, and [123] extend QNTK ideas to quantum-enhanced neural contextual bandits. These works make QNTK a useful tool for connecting circuit design, optimization dynamics, and trainability.
QNTK theory has important limitations. It applies primarily to the early phase of training, when the circuit remains close to its initialization. The dynamical phase transition found by [119] shows that late-time behavior can be more complex, which matters for applications that require the full optimization trajectory. Most QNTK analyses also assume noiseless circuits, while real quantum hardware is noisy. [117] show noise resilience in overparametrized regimes, but the general relationship between noise, QNTK, and trainability remains open.
Across these developments, stronger representational power can also amplify concentration and measurement burden. Claims about trainable or generalizable quantum kernels should therefore specify whether the kernel is estimable: how many measurements are needed to distinguish inputs reliably. They should also identify the inductive biases, such as symmetry or covariance, that keep kernels both expressive and estimable [74], [84], [85], [108].
A final AI-theoretic question is the most operational one for near-term QC: does variational training remain scalable once finite-shot fluctuations and hardware noise are unavoidable? In practice, VQAs optimize expectation values estimated from a finite number of circuit repetitions (“shots”) and, on real devices, under noisy gates and decoherence. From a learning-theoretic viewpoint, scalability then depends on three concrete questions: can the cost and its gradients still be estimated accurately enough with a reasonable number of shots? Can optimization remain stable when those estimates are noisy? How strongly does hardware noise reshape the loss landscape itself?
How noise limits scalability. Finite-shot estimation makes both the loss and its gradients random variables. When gradients are small, estimator variance can dominate the effective training dynamics. Recent analyses show that these fluctuations can determine the practical scaling of the quantum–classical training loop, linking convergence to shot budgets and step-size choices [124]. Systematic studies of variational optimization under finite shots and hardware noise also identify regimes in which noise swamps the gains from deeper circuits or more parameters unless the measurement budget grows accordingly [79].
Hardware noise also changes the objective being optimized by altering the prepared state and the induced cost function. More realistic noise models can steer the dynamics in systematic ways. [96] show that noise can induce barren-plateau-like suppression and drive training toward noise-dependent fixed points, while [93] emphasize that plateau regimes can be “swamped with traps,” meaning that many local structures may stall optimization even when gradients are not exactly zero. Practical failure can therefore come from both weak signals and trap-rich landscapes.
Recent algorithms respond by making the hybrid pipeline more measurement-efficient and noise-aware. [125] propose random coordinate descent for parameterized quantum circuits, reducing gradient-estimation overhead by updating only a subset of parameters. [126] argue that optimizers should be evaluated under latency and measurement-execution constraints, turning iteration complexity into realistic wall-clock complexity for hybrid loops. [127] introduce adaptive shot allocation, treating the shot budget as a learnable resource-allocation problem, and [128] develop a few-shot QAOA protocol for obtaining high-quality parameters with limited measurements.
Another direction is geometry-aware preconditioning. [104] study quantum natural gradient (QNG) for QAOA and report improved robustness relative to vanilla gradient descent across hardware-relevant noise models. These results suggest that trainability claims in the noiseless setting should be paired with an explicit account of shot complexity and noise-induced landscape deformation. Noise is not always detrimental: [129] show that controlled stochastic noise can help escape barren plateaus and flat regions, acting as a regularizer in certain regimes. Noise-aware algorithm design therefore includes both mitigation and beneficial stochasticity.
Taken together, recent progress makes “scalability under noise” a question about signal resolution. As the system grows, the cost and gradient signals must remain resolvable with feasible measurement resources. Optimizer design and measurement allocation must also be coordinated so that the training loop preserves signal-to-noise ratio and avoids noise-induced plateaus or trap-dominated regimes [79], [93], [96], [104], [124], [125], [127].
AI for quantum algorithm discovery can happen at various hierarchical levels, including circuit, architecture, abstract mathematical algorithms, or co-design among them. Early work by Williams and Gray demonstrated that automated search could be used to design quantum circuits, framing circuit construction as an optimization problem over possible gate sequences rather than as a purely manual task [130]. Although the circuits studied in such early work were necessarily small, the key conceptual contribution remains important: quantum circuit design can be treated as a search problem over a combinatorial design space. This perspective provides a foundation for later approaches based on genetic programming, reinforcement learning, differentiable architecture search, and symbolic rewriting.
Evolutionary methods continue to be relevant in modern quantum circuit generation. Stein and Färber revisit genetic programming for quantum circuit discovery and explicitly incorporate measures related to quantum advantage into the generation process [131]. This is an important shift from simply searching for circuits that reproduce a desired input-output behavior or minimize gate count. By including quantum advantage as part of the objective, the search process is encouraged to identify circuits whose structure is not merely classically reproducible or trivially simulable. This suggests that AI-based circuit generation should not only optimize syntactic circuit properties, but also include higher-level criteria tied to computational usefulness, expressivity, or potential advantage.
Reinforcement learning for circuit discovery. Reinforcement learning has recently emerged as one of the most active approaches for automated quantum circuit discovery. In this setting, an agent sequentially constructs or modifies a circuit by selecting actions such as adding gates, applying transformations, or choosing higher-level circuit components. Zen et al. apply this idea to the discovery of fault-tolerant logical state-preparation circuits, showing that reinforcement learning can be used to search for circuits that satisfy nontrivial constraints imposed by quantum error correction [132]. This is especially relevant for fault-tolerant quantum computing, where useful circuits must obey logical encoding constraints, avoid propagating errors catastrophically, and often optimize depth or resource overhead. In contrast to generic circuit synthesis, this work demonstrates AI-driven discovery in a setting where the target circuits are directly connected to error-corrected quantum computation.
A major challenge for reinforcement learning-based circuit discovery is scalability. Searching gate by gate becomes increasingly difficult as the number of qubits, circuit depth, and constraint complexity grow. Olle et al. address this issue by introducing “gadgets” as higher-level action primitives for reinforcement learning [133]. Rather than forcing the agent to rediscover useful subcircuits repeatedly from elementary gates, the search can operate over reusable circuit motifs. This hierarchical approach reduces the effective search depth and makes it possible to discover more complex circuits. A related direction is the automated discovery of such gadgets themselves, where repeated or useful substructures are identified and promoted to new building blocks for subsequent learning [134]. Together, these works point toward a form of compositional circuit discovery, in which AI systems gradually build libraries of reusable quantum subroutines.
Co-designing circuit and architecture. Beyond direct circuit construction, reinforcement learning has also been applied to circuit synthesis, transpilation, and optimization. Kremer et al. study practical and efficient quantum circuit synthesis and transpiling with reinforcement learning, targeting tasks where the goal is to produce circuits that satisfy architectural or compilation constraints while remaining resource efficient [135]. This line of work is important because algorithm discovery cannot be separated from implementability: a theoretically attractive circuit may become impractical after compilation to a real hardware topology or fault-tolerant gate set. Similarly, Riu et al. combine reinforcement learning with ZX-calculus, allowing agents to optimize quantum circuits through graph-rewriting rules rather than only through gate-level transformations [136]. This symbolic perspective is promising because ZX-calculus provides a mathematically structured representation in which non-obvious circuit equivalences can be exploited during search.
Another direction is differentiable quantum architecture search. Zhang et al. formulate quantum circuit architecture design as a differentiable optimization problem, enabling gradient-based search over parameterized circuit structures [137]. This approach is particularly relevant for VQAs, where the choice of ansatz strongly affects trainability, expressivity, and hardware efficiency. Compared with reinforcement learning or genetic programming, differentiable architecture search can exploit continuous relaxations of discrete design decisions, potentially improving search efficiency. However, it also raises questions about whether the relaxed optimization landscape faithfully represents the discrete circuit-design problem.
Outlook. These progress from small-scale automated circuit design to increasingly structured and scalable AI-assisted discovery frameworks suggest that the field is moving beyond mere circuit optimization toward the automated discovery of reusable, hardware-aware, and fault-tolerance-compatible quantum algorithmic components. One important open question is to use AI to verify the correctness of large-scale circuits or quantum algorithms. Another promising future direction is to leverage AI and formal proof tools to discover algorithms based on entirely new mathematical structures. One example of such a link between algorithms and mathematics is the equivalence between quantum signal processing algorithms [138]–[140] and nonlinear Fourier transform [141], [142]. Another example is group and representation theory and the corresponding quantum algorithms such as Schur transform and generalized phase estimation [143]. Other mathematical structures will inspire novel quantum algorithms that are unknown before. In addition, there are more design spaces across the entire quantum computing stack where AI can be helpful to explore, including co-designing QEC codes with algorithms and architectures.
Correction and control are needed because realistic QC hardware has several coupled sources of error. Physical qubits are unavoidably noisy [25], [144]. These imperfections can be grouped into four broad categories: (i) decoherence, which causes stored quantum information to relax or lose phase coherence over time; (ii) control imperfections, which produce gate errors and can push the state out of the intended computational subspace; (iii) readout infidelity, meaning that the measured bit may not match the underlying qubit state; and (iv) cross-talk and temporal drift, meaning that neighboring qubits can influence one another and calibrated device parameters can slowly change over time.
Quantum error correction (QEC) addresses these limitations by encoding logical information redundantly across many physical qubits and repeatedly measuring stabilizer operators to infer and correct faults without disturbing the logical state [145], [146]. At the abstract level, QEC is defined by a code subspace \(\mathcal{C}\), a noise channel \(\mathcal{N}\), and a recovery channel \(\mathcal{R}\) such that \((\mathcal{R}\circ\mathcal{N})(\rho)=\rho\) for all logical states \(\rho\) encoded in \(\mathcal{C}\). Stabilizer-based QEC [147], [148] is a practically important realization of this framework.
Each syndrome measurement cycle produces classical data, and a classical decoder uses those data to infer the most likely fault pattern and apply a recovery operation. QEC is therefore a sequential inference and feedback-control problem: syndrome data arrive round by round, and the system must update its fault estimate and recovery action in real time. AI methods support this broader workflow through decoding, noise characterization, adaptive control, and code discovery.
| Task | Typical AI methods | Role in noisy quantum systems |
|---|---|---|
| Syndrome decoding | CNNs, recurrent models, Transformers, GNNs, differentiable message passing | Infer likely fault patterns from repeated syndrome measurements under strict latency constraints. |
| Noise-model learning and adaptation | Supervised learning, likelihood-based inference, adaptive decoders, surrogate models | Estimate drift, correlations, leakage, biased noise, and readout errors, then adapt decoders or calibration policies. |
| Adaptive QEC and code discovery | Reinforcement learning, search, differentiable optimization, generative design | Choose recovery actions, measurement schedules, code layouts, or protocol structures matched to hardware constraints. |
| Control and calibration | Bayesian optimization, RL, neural controllers, autonomous experimental agents | Tune device parameters, stabilize operation, and close the loop between characterization, correction, and hardware control. |
4pt
Table 3 organizes the discussion in this subsection by task and AI method family. Here, decoding belongs most directly to the standard inner loop of QEC, whereas noise adaptation, adaptive control, and code discovery are better viewed as supporting tasks in the broader workflow of making QEC effective on realistic hardware.
The decoder must process syndrome data in real time, handle faulty measurements, and generalize across noise models and code families, all within the strict latency budgets imposed by qubit coherence times. A broad range of neural architectures are applied to this problem, each exploiting different structural properties of the syndrome data.
Convolutional and feedforward neural decoders. CNNs are among the earliest neural architectures applied to decoding, exploiting the planar lattice structure of the surface code to assign local error likelihoods [149]–[151]. A recent systems-oriented example is the AI-based pre-decoding framework of [8], which uses local, parallel neural pre-decoders for surface-code syndrome volumes before passing residual syndromes to a downstream global decoder such as PyMatching, achieving microsecond-scale decoding runtimes on NVIDIA GPUs while reducing logical error rates relative to global decoding alone. A parallel line of work uses feedforward, or fully connected, neural networks as another decoder family. In this family, [152] design scalable syndrome decoders for surface codes of arbitrary shape and size, incorporating noise-model flexibility for biased and spatially inhomogeneous noise, and related feedforward decoders are also benchmarked directly on IBM quantum processors [153].
Transformer and recurrent decoders. A major direction replaces spatial convolutions with sequence models that process repeated syndrome rounds as a time series, capturing temporal correlations across measurement cycles. A recurrent transformer decoder demonstrates performance advantages over minimum-weight perfect matching (MWPM) on both simulated and experimental data from Google’s Sycamore processor [154]. AlphaQubit [155] extends this approach with a two-stage training strategy: large-scale simulation pretraining followed by finetuning on scarce experimental data. The resulting decoder achieves improved logical error rates on distance-3 and distance-5 surface-code experiments, as well as on simulated distances up to 11 under realistic noise including cross-talk, leakage, and analog readout. This pretraining-then-finetuning paradigm is now widely adopted: [156] apply a similar strategy and explicitly incorporate analog readout information to improve decoding accuracy. Transfer learning across code distances is further explored using Transformer decoders, reporting over an order-of-magnitude reduction in training cost [157]. End-to-end approaches that optimize directly against decoding performance metrics extend naturally to the practically important setting of faulty syndrome measurements [158].
Graph neural network decoders. Graph neural networks (GNNs) offer a complementary approach well suited to codes with irregular sparse structure, where MWPM is inapplicable and belief propagation (BP) is impaired by short cycles in the Tanner graph. By learning message-passing schedules directly from syndrome data on the Tanner or detector graph, GNN decoders generalize across code families. [159] show that sufficiently expressive GNN decoders trained on simulated and experimental data can match matching-based approaches for stabilizer codes. The Astra decoder [160] operates on Tanner graphs and reports improved thresholds and logical error rates over BP+OSD for qLDPC codes, while demonstrating extrapolation to larger code distances using models trained on smaller instances. Differentiable iterative decoders interleaving BP stages with GNN layers are proposed to suppress trapping sets arising from short Tanner-graph cycles [161], and GNN-based message-passing decoders are benchmarked across several qLDPC designs [162]. GraphQEC [163] pushes further toward universality, proposing a code-agnostic neural decoder with linear-time complexity that generalizes across surface codes, color codes, and qLDPC codes within a single model.
Reinforcement learning decoders. RL frames decoding as a sequential decision-making problem in which an agent is trained by assigning rewards for successful corrections [164]. Deep Q-learning agents are applied to toric and surface codes under faulty syndrome measurements, achieving performance comparable to MWPM at low error rates [165].
Future directions. Three decoder trends now stand out. First, moving to larger code distances and more general code families will require substantially larger and more representative training sets, since direct supervised scaling quickly becomes expensive even when transfer across code distances is possible [157], [160], [163]. Second, higher decoding accuracy will increasingly depend on fine-tuning pretrained models on experimental QPU data, because hardware-specific effects such as leakage, analog readout, drift, and correlated faults are difficult to capture faithfully with stylized simulated noise alone [155], [156]. Third, practical impact will hinge on real-time deployment: strong offline accuracy is not sufficient unless neural decoders can be integrated into low-latency classical control stacks whose inference time remains compatible with syndrome-cycle deadlines and feedback bandwidth constraints [145], [166].
Most decoders assume a known, stationary noise model, yet hardware noise drifts over time, exhibits spatial correlations, and frequently deviates from the idealized depolarizing channel used during training. The Decoding Graph Re-weighting (DGR) method [167] estimates up-to-date edge probabilities and correlations from matching statistics accumulated across syndrome rounds, and uses these estimates to reweight the MWPM decoding graph, reporting substantial logical error rate improvements under noise mismatch and correlated noise. Bayesian inference methods are used to estimate parameters of general noise models, including time-varying channels via sequential Monte Carlo [168]. Recent work extends syndrome-based learning to circuit-level noise models via Fourier analysis and compressed sensing, reporting sample-complexity savings over direct logical benchmarking [169].
RL demonstrates broad utility across the QEC pipeline beyond decoding. At the hardware level, QEC detection events can serve directly as a learning signal for steering physical control parameters under drift: [166] demonstrate this paradigm experimentally, showing improved stability of the surface-code logical error rate against injected drift, with scaling simulations extending to distance 15. Model-free RL is credited as a key component in achieving coherence gains greater than unity in a 3D superconducting cavity memory [145], and in optimizing GKP-encoded qudit memories [146]. Multi-agent RL is also proposed to simultaneously discover QEC cycle components and adapt to time-varying noise channels in a model-free manner [170].
RL also drives the automated discovery of hardware-tailored QEC codes. [171] train a noise-aware RL agent to simultaneously discover stabilizer codes and encoding circuits under specified gate sets, connectivity graphs, and noise models, scaling to approximately 20–25 physical qubits and distance-5 codes. [172] combine RL with the Quantum Lego tensor-network code framework to optimize code distance and logical error metrics under biased noise, discovering constructions including an optimal \([[17,1,3]]\)-type code for asymmetric noise settings. RL optimization of tensor network code geometries has similarly been shown to identify optimal stabilizer codes more efficiently than random search [173].
Recent progress in AI has made agentic systems [1], [2], [174] a major direction for large language models (LLMs). These systems pursue long-horizon goals through repeated cycles of planning, acting, observing, and revision. Foundational work such as ReAct [175] and Toolformer [176] shows that language models can be embedded in closed loops that combine reasoning, tool use, and environmental interaction. More recent surveys [1], [2], [174] treat autonomous and multi-agent LLM systems as a systems layer above the base model, built from memory, planning, tool invocation, coordination, and self-refinement. These developments create opportunities across quantum information science and technology: agentic systems can orchestrate research and engineering pipelines that include task decomposition, experimental execution, and iterative refinement [177], [178]. This section surveys three main thrusts at the intersection of autonomous AI and quantum workflows: autonomous quantum programming, self-driving quantum laboratories, and formal frameworks for discovery and verification.
A central application is the automated generation, debugging, and optimization of quantum circuits from natural language descriptions. QAgent [179] is a representative LLM-powered multi-agent system that fully automates OpenQASM programming for NISQ devices. It integrates task planning, retrieval-augmented generation (RAG) for long-term context, few-shot in-context learning, chain-of-thought (CoT) reasoning, and reflection mechanisms. QAgent decomposes quantum problems into sub-tasks, routes them to either a Dynamic-few-shot Coder or a Tools-augmented Coder, and employs a Reflection Agent for iterative debugging, achieving a 71.6% improvement in code generation correctness over prior static LLM approaches across multiple base models.
Agent-Q [180] takes a complementary approach by constructing a large-scale training dataset of over 14,000 quantum circuits (covering QAOA, VQE, and adaptive VQE) and fine-tuning LLMs to generate syntactically correct parameterized quantum circuits in OpenQASM 3.0. Agent-Q is designed to be integrated into autonomous workflows where generated circuits serve as starting points for further optimization, functioning as templates in quantum machine learning and as benchmarks for compilers and hardware. [181] further develops a multi-agent framework that uses semantic analysis and multi-pass inference alongside an error correction module to enhance LLM-based quantum code generation with fault-tolerant considerations.
Benchmarks and Dedicated Quantum Code Models. The rapid development of LLM-based quantum programming tools has created a need for standardized evaluation. Qiskit HumanEval [182] introduces the first benchmark specifically designed for quantum code generation, with more than 100 hand-curated tasks across eight categories of
quantum programming functionality. Alongside the benchmark, IBM releases granite-8b-qiskit, a domain-adapted code LLM that significantly outperforms general-purpose models on quantum coding tasks.
Subsequent benchmarks address complementary dimensions. QCoder [183] incorporates hardware-simulator-based evaluation and human-written reference solutions from programming contests, enabling assessment of whether generated circuits execute correctly on simulated quantum devices. QuanBench [184] evaluates both syntactic correctness and quantum semantic fidelity via process-fidelity metrics, testing LLMs on algorithmic reasoning beyond API compliance. These benchmarks show that current LLMs perform reasonably on basic quantum circuit construction but still struggle with deeper algorithmic reasoning and hardware-aware optimization. This motivates further work on domain-specific fine-tuning, such as the reinforcement-learning-based QSpark [185], which fine-tunes Qwen2.5-Coder on Qiskit HumanEval, and on multi-agent architectures that decompose complex quantum programming tasks into manageable subtasks.
Beyond code generation, recent work asks whether LLMs can reason about quantum algorithms at a deeper symbolic level. [186] train GroverGPT, an 8-billion-parameter language model designed for quantum search, and show that large-scale pretraining on quantum-structured data can help LLMs internalize algorithmic logic together with circuit syntax. Complementing this direction, [187] introduce a framework for symbolic analysis of Grover’s search algorithm that combines chain-of-thought (CoT) reasoning with quantum-native tokenization, enabling step-by-step symbolic execution of quantum algorithms.
The recent integration of artificial intelligence into experimental quantum science has shifted the role of automation from running batch sequences to closing the loop on design, fabrication, control, and characterization of quantum hardware. Several recent reviews give a cross-stack picture of this shift [18], [188]–[192]. We use “self-driving” broadly to describe AI systems that close at least one experimental feedback loop—proposing an action, executing or recommending it, observing the result, and updating the next action—and we use “AI systems” equally broadly, encompassing a wide spectrum of tools that range from classical nonlinear-optimization algorithms in specific settings, to reinforcement-learning (RL) and neural-network (NN) methods, to frontier large language models (LLMs). Depending on the experimental platform, a typical quantum laboratory progresses through some or all of the following steps to complete an experiment: (i) design of the experiment and devices, (ii) fabrication of the devices, and (iii) control and measurement—including device tune-up or state preparation, device / Hamiltonian characterization, control-parameter optimization, and readout-parameter optimization. Although a fully self-driving quantum laboratory remains out of reach today, workflows with varying levels of AI assistance and automation have been demonstrated for each of these individual steps. In what follows, we organize the discussion according to these steps.
Experiment design. Quantum-optics experiments built from a finite vocabulary of building blocks, such as single-photon sources, beam splitters, mirrors, shift-parameterized holograms, and Dove prisms, are a natural target for AI design. Each block has a well-defined unitary action, so the design problem becomes a structured search over how to stack the blocks. An early demonstration was MELVIN [193], which autonomously composed standard quantum-optics elements into experimental setups generating complex multi-photon states. It identified previously unknown layouts for high-dimensional GHZ states and asymmetrically entangled states, and learned from simpler instances to accelerate discovery on more complex ones. Soon after, projective-simulation-based active learning [194] showed that a reinforcement-learning-style (RL) agent could generate high-dimensional entangled multi-photon states more efficiently than random search and rediscover useful quantum-optics design strategies.
The next wave of these tools emphasized interpretability. Theseus [195] introduced a graph-based representation of quantum-optical experiments and an inverse-design algorithm orders of magnitude faster than MELVIN, with solutions that scientists can read and reason about. Klaus [196] reformulated photonic experiment design as a Boolean satisfiability problem hybridized with continuous optimization, returning exact minimal configurations and bringing logic AI into quantum experimental design for the first time. The open-source successor PyTheus [197] discovered one hundred diverse experimental designs spanning highly entangled states, optimized measurement schemes, communication protocols, and multi-particle gates. In 2026, language-model meta-design [198] took the loop a step further by training a transformer to generate Python programs that solve entire families of quantum-optical design problems in a single forward pass. Because the generated code is human-readable, researchers can inspect the underlying design principles directly, and the trained model produced novel experimental generalizations of important condensed-matter quantum states.
Device design. Superconducting-qubit hardware is a useful example because it offers substantial design flexibility and relatively tight agreement between designed and fabricated devices. The design problem has two stages. The first maps a target Hamiltonian or quantum function to a lumped-element circuit model and its parameters. The second maps that circuit model to a manufacturable layout through geometry optimization and electromagnetic simulation.
The first stage has seen the most direct AI uptake. SCILLA [199] is an automated closed-loop design tool that proposes superconducting circuit diagrams whose Hamiltonian spectrum and noise sensitivities match a user-specified target; it is demonstrated on 4-local couplers, and the same group’s experimental realization of tunable three-body interactions [200] provides hardware context for the interaction designs SCILLA targets. A differentiable framework due to Ni et al. [201] jointly optimizes device parameters and control parameters in a single gradient-based pass, extending GRAPE-style optimization to the device layer. A more recent automatic-differentiation framework on top of the SQcircuit package [202] computes eigensystem gradients of large sparse circuit Hamiltonians and applies them to qubit discovery, including metrics for decoherence and fabrication-error robustness. Cárdenas-López et al. [203] use genetic algorithms to design superconducting-circuit elements with target spectra, selection rules, and built-in resilience to fabrication variability.
Layout-level work connects circuit design to manufacturable geometries. Li and Jin [204] derive an analytic zero-coupling condition for the qubit-coupler-qubit (QCQ) architecture, together with a geometry-only upper bound on the effective qubit–qubit coupling. They propose an automated layout procedure that approaches that bound and validate it by electromagnetic simulation, including a millimeter-scale long-range QCQ layout. Predictive machine learning on Qiskit Metal and ANSYS simulation data [205] forward-predicts transmon engineering parameters such as frequencies, anharmonicities, and couplings, enabling fast design-space exploration.
Two complementary tools support this layout-level workflow. SQuADDS [206] provides an open-source database of experimentally validated superconducting device designs, each generable through Qiskit Metal and simulated via finite-element solvers, together with a front-end that returns a “best-guess” design from desired circuit parameters. QDesignOptimizer [207] closes the optimization loop with an open-source Python package that combines Ansys HFSS electromagnetic simulation, pyEPR Energy-Participation-Ratio analysis, and Qiskit-Metal under a physics-guided nonlinear-optimization strategy that targets user-specified mode frequencies, decay rates, and coupling strengths across qubits, resonators, and couplers. At the chip and system level, graph neural networks (GNNs) have begun to replace heuristic layout tools. Ai and Liu [208] introduce a GNN-based parameter-design algorithm whose “three-stair scaling” mechanism mitigates quantum crosstalk on circuits with up to roughly 870 qubits, reducing errors to about 51% of the previous state of the art and cutting runtime from 90 minutes to 27 seconds. Lu et al. [209] develop a closed-loop neural-network surrogate plus Bayesian optimization (BO) for superconducting-chip frequency configuration, accounting for nonlinear error mechanisms such as crosstalk and decoherence that linear models cannot capture, and validate the method by randomized benchmarking and cross-entropy benchmarking.
For integrated quantum photonics, the inverse-design pipeline spans emitter selection, device-element design, and deterministic on-chip emitter assembly, as surveyed by Kudyshev et al. [210].
Fabrication and device preparation. Once a design exists, AI plays two complementary fabrication roles. The first is to identify patterns in empirical or in-line measurements, either to fine-tune the recipe or to select optimized device regions. The second is to drive robotic actuators that physically execute the recipe. A direct demonstration of the latter is robotic chip-scale nanofabrication of Josephson-junction devices [211]. Automating the wet-chemical resist-development step with a robotic arm reduces chip-to-chip junction-resistance spread from roughly 7% under human operation to roughly 2% with the robot, pointing toward more consistent device fabrication.
For the analytics-driven role, CNN analysis of in-line scanning-electron-microscopy (SEM) micrographs has been used to optimize quantum-dot qubit nanofabrication, including lithographic proximity effects and process quality control [212]. Atomic-precision silicon-qubit fabrication has been a particularly fruitful target. Real-time machine-learning prediction of donor number during scanning-tunneling-microscope (STM) patterning enables atomic-precision manufacturing of donor qubits in silicon [213], while million-atom STM simulations combined with machine learning provide post-fabrication metrology that infers donor number and positions for atomic-scale devices [214]. Two earlier modules from the Wolkow group make this pipeline autonomous: deep-learning identification of defects on hydrogen-terminated silicon routes patterning into defect-free regions before hydrogen lithography [215], and an earlier CNN automatically detects degraded STM tips and triggers in-situ reconditioning [216]. Machine-learning-enhanced in-situ electron-beam lithography (EBL) has been used to deterministically integrate single InGaAs quantum dots into photonic nanostructures [217]. At the materials-growth front, supervised regression and BO over diamond synthesis and post-processing parameters improve nitrogen-vacancy (NV) magnetometry sensitivity by 300% over an average sample and 55% over the previous champion [218], with Shapley-importance rankings revealing the dominant growth parameters.
Control: atomic, molecular, and optical (AMO) experiments. The control layer is where AI in quantum experiments has accumulated the deepest experimental track record. On atomic, molecular, and optical (AMO) platforms, Wigley et al. [219] established the canonical example of self-optimizing quantum experiments by reducing the number of optimization iterations needed to produce a Bose–Einstein condensate (BEC) by roughly an order of magnitude. BO has since been pushed to genuinely high-dimensional control with up to 55 simultaneous controls in BEC preparation [220], and reinforcement learning has extended these gains: Milson et al. [221] demonstrate experimental RL preparation of ultracold quantum gases over many simultaneous control parameters, while Reinschmidt et al. [222] use RL to control a magneto-optical trap, discover robust operating modes, and successfully transfer in-silico learning to a real experiment. BO has also been used to improve robust many-body state-preparation protocols relevant to few-body fractional quantum Hall physics [223], bridging AI experiment control and many-body quantum-state preparation.
Control: quantum-dot device tune-up. Quantum-dot devices were an early adopter of machine-learning autotuning, as reviewed by Zwolak and Taylor [224]. Kalantre et al. [225] used CNNs for state recognition and automated tuning, in-situ machine-learning autotuning on hardware was demonstrated for double-dot devices [226], and Moon et al. [227] achieved completely automatic tuning of a quantum device across gate-voltage spaces of dimension up to eight, with median tuning times below 70 minutes and decisive speedups over random search. Schuff et al. [228] have demonstrated fully autonomous tuning of a spin qubit from a grounded device through to coherent Rabi oscillations, combining deep learning, BO, and computer vision; RL has also been used to decide what to measure, reducing measurement burden on quantum devices including automatic identification of double-dot bias triangles [229].
Control: device or Hamiltonian characterization. Closely related to control is the question of what model the controller is acting on. This has led to a parallel body of work on Hamiltonian learning and device characterization. Béjanin et al. [230] formulate resonant-coupling parameter estimation using offline data acquisition and modeling combined with online Bayesian learning. Genois et al. [231] combine domain knowledge with a recurrent neural network (RNN) to learn the dynamics of a superconducting qubit from experimental data; the same collaboration later extends this RNN-based characterization to detailed numerical quantum optimal control of single qubits [232]. BO calibration with a state-parameter estimator for non-Markovian environments has been studied numerically [233], learning-based calibration of flux crosstalk has been demonstrated in transmon arrays [234], and real-time binary-search Hamiltonian tracking achieves exponential scaling of calibration precision in measurement count [235].
Hardware-aware characterization is also moving toward real-time and multimodal settings. FPGA-enabled real-time \(T_1\) tracking resolves fast fluctuations of \(T_1\) that previous nonadaptive protocols averaged out, improving temporal resolution by two orders of magnitude [236]. Deep transfer learning maps fluxonium spectra to Hamiltonian parameters, addressing the spectrum-inversion problem that is harder for fluxonium and 0-\(\pi\) qubits than for transmons [237]. A recent multimodal turn treats calibration plots themselves as inputs to vision-language models: [7] introduce QCalEval, a benchmark for quantum-calibration plot understanding with 243 samples across 87 scenario types and 22 experiment families, and release NVIDIA Ising Calibration 1 as an open-weight reference model for this setting.
Beyond superconducting hardware, autonomous characterization agents have operated directly on real quantum hardware. An autonomous protocol combining unsupervised machine learning with a genetic algorithm reverse-engineers Hamiltonian models from nitrogen-vacancy-center experiments and retrieves plausible Hamiltonian models in 74% of experimental-data instances [238]. The Quantum Model Learning Agent (QMLA), built around exploration-tree strategies and Elo-style objective functions, identifies the correct interaction family in 72% of simulated Ising / Heisenberg / Hubbard cases [239]. On a reconfigurable integrated photonic device, an experimental gray-box agent simultaneously controls the device and reconstructs internal unitaries / Hamiltonians, accommodating fabrication imperfections [240].
Control: gate and operation-parameter optimization. Model-free RL has become a workhorse for non-classical state preparation and gate design. Sivak et al. [241] proposed and numerically demonstrated model-free RL with adaptive measurement-based feedback for non-classical-state preparation, an approach that has subsequently been adapted in Gottesman–Kitaev–Preskill (GKP) control work. Porotti et al. [242] study coherent transport and feedback control with deep RL, including Fock-state preparation under weak nonlinear measurements. Li et al. [243] introduce a sample-efficient RL-from-demonstration approach for quantum gate calibration, targeting workflows where individual experiments are expensive.
Machine learning has also been used to find pulse parameters for unconventional gates, including three-qubit Toffoli and parity-check operations in transmon systems [244]. On real IBM hardware, Baum et al. [245] show that experimental deep RL can produce error-robust gate sets. Single-qubit gate design has been demonstrated with model-free RL operating on real-time readout signals [246]. In numerical simulation, model-free RL has designed two-qubit transmon entangling gates such as CNOT and cross-resonance [247], and an RL-parameterized ansatz has been combined with classical optimal control to design fast two-qubit gates [248]. An embedded / FPGA implementation pipeline for arbitrary single-qubit rotations [249] and a multi-objective deep-RL formulation for superconducting quantum optimal control [250] extend this thread toward deployed hardware and multi-criterion optimization. Across gate-design work, a clear separation emerges between numerical pulse design and a growing set of hardware-validated demonstrations, with the latter increasingly using real-time feedback.
Control: readout or state-discrimination parameter optimization. Readout is the layer where machine learning has had a particularly cleanly measured impact. Magesan et al. [251] introduced machine-learning-based discrimination of quantum-measurement trajectories using IBM superconducting-qubit data, and trapped-ion readout was likewise improved by neural networks trained on experimental data [252]. Hidden Markov models capture temporal structure in readout signals [253]. Recurrent networks reconstruct full quantum dynamics from physical observations in superconducting circuits [254]; an LSTM estimator tracks fast superconducting-qubit dynamics [255]; and deep networks improve frequency-multiplexed five-qubit superconducting readout [256].
Other work improves the signal-processing and deployment layers. Excited-state-promoted readout has been combined with feedforward classifiers [257], hardware-efficient machine-learning architectures have been engineered for readout scaling [258], and autoencoders [259] and Bayesian learning [260] each enhance discrimination at the signal level. Time-resolved modulated networks [261] move beyond binary discrimination to full state tomography, while path-signature features of the time-domain readout [262] further improve assignment fidelity across multiplexed setups.
A growing body of work pushes inference onto field-programmable gate arrays (FPGAs) at nanosecond latency, enabling mid-circuit measurement and feedback. Machine-learning-powered FPGA-based real-time discrimination has demonstrated mid-circuit measurements for superconducting qubits [263], and a low-latency neural-network accelerator has been designed for multi-qubit-state discrimination [264]. RL has been used to co-optimize readout pulses and classifiers [265]. End-to-end deployment workflows now combine the RFSoC FPGA framework with the hls4ml (high-level-synthesis-for-machine-learning) deployment toolchain [266], knowledge-distilled lightweight networks tailored for FPGA readout [267], and dedicated architectures for multi-level (qutrit / leakage-aware) readout at scale [268]. Bayesian-learning measurement-error mitigation has been extended to multi-qubit experiments on near-term superconducting devices [269], and reservoir computing provides a computationally lighter alternative to deep networks for crosstalk-aware multiplexed readout [270]. Closing the feedback loop, Reuer et al. [271] realized a deep-RL agent in real-time FPGA hardware demonstrated on a qubit-reset task.
Agentic quantum control. Agentic AI systems are beginning to close the full control-and-measurement loop by translating human instructions directly into laboratory experiments. An agent-based framework [272] translates natural language into executable laboratory scripts and reads figures, such as plots and scope traces, to choose the next measurement step. An LLM-assisted superconducting-qubit experimentation framework [273] demonstrates two end-to-end tasks: finding readout resonators from scratch, and reading a research paper to reproduce the quantum non-demolition (QND) measurement protocol it describes.
On the conceptual side, [274] provides the first formal definition of quantum agents and proposes architectures that integrate quantum workflows with autonomous agent-based systems. Their analysis covers both how QI may enhance AI capabilities and how autonomous AI can support quantum workflows. [275] presents Ax-Prover, a multi-agent theorem-proving system in Lean that combines LLM reasoning with formal verification tools. Ax-Prover outperforms frontier LLMs on the authors’ QuantumTheorems benchmark and further demonstrates assistant capabilities in quantum cryptography.
Going beyond individual theorem proving, MerLean [276] introduces a fully automated agentic framework for end-to-end autoformalization of quantum-computing research papers. MerLean extracts mathematical statements directly from LaTeX source files, formalizes them into verified Lean 4 code built on Mathlib, and translates the results back into human-readable LaTeX for semantic review. Evaluated on three theoretical quantum-computing papers, MerLean produces 2,050 Lean declarations from 114 statements, reducing the verification burden to only newly introduced definitions and axioms. This pipeline offers both a practical tool for machine-verified peer review of quantum-computing research and a scalable engine for mining high-quality synthetic data to train future reasoning models. A more domain-specific formalization effort is the end-to-end Lean 4 formalization of quantum error correction by [277]. This highlights a complementary role for proof assistants: not only autoformalizing research papers, but also closing trust gaps in concrete fault-tolerant quantum-computing primitives such as QEC definitions, code properties, and correctness certificates.
[278] examines the governance and ethical implications of combining autonomous AI with QI. Beyond orchestration and verification, autonomous AI is also emerging as a creative partner in physics discovery. [279] reports a collaboration between GPT-5.2 and physicists in which the model autonomously conjectures a closed-form formula showing that single-minus tree-level gluon scattering amplitudes are nonzero in the half-collinear regime, overturning a decades-old assumption in quantum field theory, and derives a formal proof subsequently verified by the human collaborators.
Despite rapid progress, significant challenges remain: the reliability and verifiability of autonomously generated quantum protocols, the limited context windows of current LLMs for representing complex quantum systems [280], and the need for benchmarks that go beyond syntactic correctness to evaluate quantum semantic fidelity and hardware-aware optimization [179], [180], [182], [184]. The emerging benchmarking infrastructure (Qiskit HumanEval, QCoder, QuanBench) and autoformalization pipelines (Ax-Prover, MerLean) provide initial foundations, but substantial gaps remain in evaluating agentic systems on end-to-end quantum research workflows that span algorithm design, implementation, execution, and interpretation.
Artificial intelligence is emerging as a useful design tool for advanced quantum sensing protocols. In conventional quantum metrology, the goal is often to estimate an unknown parameter with minimum variance. Many sensing applications require a downstream decision, such as classification, hypothesis testing, anomaly detection, or extraction of a nonlinear feature. This motivates a task-oriented formulation in which the probe state, sensing interaction, quantum control, measurement, and classical decision rule are optimized end to end. This view naturally connects quantum sensing with VQAs and quantum machine learning.
A key example is supervised learning assisted by an entangled sensor network (SLAEN) [281]. In SLAEN, distributed quantum sensors coordinate entanglement and measurement strategy to perform a classification task directly at the physical layer. The entangled network learns collective observables matched to the decision boundary. The original theory considered a support-vector-machine-inspired classification problem and trained a squeezing–linear-optics circuit to minimize classification error. An experimental demonstration later showed quantum-enhanced data classification using a variational entangled sensor network for multidimensional radio-frequency signals [282]. These works established that entanglement can be shaped to improve parameter precision and reduce task-specific decision error.
Gaussian SLAEN architectures are naturally suited to linear classification tasks. This limitation was addressed by controllable bosonic variational sensor networks, where universal bosonic control enables nonlinear classification by training non-Gaussian sensing and measurement operations [283]. In this setting, the sensor acts as a trainable quantum feature processor: the physical channel embeds data into a quantum state, while the variational circuit maps nonlinear task-relevant features onto measurement outcomes. Related ideas appear in quantum computational sensing, where quantum processing is integrated with the sensing interaction so that the device directly outputs a property or class label [284]. Recent experimental work on quantum computational displacement sensing used a superconducting qubit–oscillator platform and trained parameterized circuits to classify displacement signals, reporting a task-specific sensing advantage over estimate-then-classify baselines [285].
The classification formulation is also closely related to quantum state hypothesis testing. Variational circuits can be trained to minimize state-discrimination error, as in quantum convolutional neural networks for many-body phase classification [286], universal discriminative quantum neural networks [287], and noisy quantum neural-network state discrimination [288]. Analyses of circuit depth and classification-error decay further clarify how variational expressivity, noise, and trainability determine sensing performance [289].
Adaptive quantum sensing has roots in optical coherent-state discrimination, where the receiver is optimized to distinguish nonorthogonal quantum states with minimum error. The Kennedy receiver introduced a displacement-and-photon-counting strategy for binary coherent states [290], and the Dolinar receiver showed that real-time feedback can attain the Helstrom bound for binary coherent-state discrimination [291]. An early experimental milestone was the closed-loop coherent-state measurement, which used real-time quantum feedback to emulate an optimal measurement for optical coherent-state discrimination [292]. Conditional-nulling ideas were later extended to optical codeword demodulation, where optimized coherent-pulse nulling, photon counting, and quantum feedforward achieved error rates below direct detection [293].
This receiver viewpoint naturally connects to adaptive and learning-based sensing. Becerra and co-workers demonstrated adaptive displacement and photon-counting receivers beating the standard quantum limit for multiple nonorthogonal coherent states [294], followed by photon-number-resolving receivers robust to realistic imperfections [295], multi-state discrimination below the quantum noise limit [296], and robust binary coherent-state discrimination [297]. Machine-learning approaches then generalized adaptivity from hand-designed feedback to optimized policies: Hentschel and Sanders used machine learning to design adaptive phase-estimation strategies [298]; Fiderer, Schuff, and Braun introduced neural-network heuristics for adaptive Bayesian quantum estimation [299]; QREAL used adaptive learning to optimize a feed-forward quantum receiver under different coherent-state discrimination tasks and operating conditions [300]; and Cimini et al. applied deep reinforcement learning to quantum multiparameter estimation [301]. These works show a chronological progression from analytically designed adaptive receivers to AI-optimized sensing policies.
A parallel literature on variational quantum sensing optimizes probes and measurements for metrological objectives in ions and atoms. Examples include variational spin-squeezing algorithms on programmable sensors [302], optimal metrology with trapped-ion quantum sensors [303], and variational multiparameter quantum metrology for vector-field sensing [304]. These works often minimize estimation error and share the same design principle: use trainable, hardware-native quantum circuits to discover entangled probes and measurements adapted to the task.
Overall, AI-enabled quantum sensing reframes the sensor as a trainable quantum information processor optimized for an operational task. The central opportunity is end-to-end design: entanglement, squeezing, nonlinear control, measurement, and classical postprocessing can be co-optimized under realistic hardware constraints. Open challenges include rigorous advantage criteria, fair resource accounting, sample-efficient training, robustness to noise and model mismatch, and scalable benchmarks. Recent theory and experiments suggest that AI-designed quantum sensors may be especially powerful when the desired output is a decision, feature, or hypothesis.
Quantum networks aim to distribute entanglement between distant nodes so that the resulting shared resource can support tasks such as quantum key distribution (QKD), distributed sensing, and distributed quantum computation [305]. The central obstacle is that entanglement cannot be amplified or copied the way classical signals can, so photon loss and decoherence accumulate exponentially with fiber length. The standard solution is the quantum repeater, which divides a long channel into shorter segments, generates entangled pairs across each segment, and connects them using entanglement swapping at intermediate nodes. Connection and local operations are themselves imperfect, so each swap lowers the fidelity of the resulting pair, and repeaters therefore interleave swapping with entanglement purification, which probabilistically distills higher-fidelity pairs from lower-fidelity ones at the cost of consuming additional pairs and additional classical communication. The resulting nested protocols deliver high-fidelity end-to-end entanglement using resources that grow only polynomially with distance [306], [307].
Once repeater nodes are connected into networks rather than linear chains, additional control problems appear. Switches must route entanglement among many users, and grids of repeaters must support multiple simultaneous flows. The network must decide which users to serve, which paths to use, and how to share limited memory across competing requests, so the operation of every node becomes a sequence of decisions: when to attempt generation, when to swap, when to purify, when to discard a stored pair whose fidelity has decayed, and when to consume a pair for the application. These decisions are naturally modeled as a Markov decision process (MDP), with states capturing the occupancy and quality of each memory and actions corresponding to the local operations available at each node.
Several features of realistic networks make the resulting MDPs hard for classical methods. The state space grows rapidly with the number of memories, links, and pending classical messages, so exact dynamic programming becomes intractable beyond small systems; even the bipartite–tripartite switch with two memories per link is analyzed only numerically [308], [309], and routing in 2D grids is solved through carefully designed local heuristics rather than optimal policies [310]. Operationally relevant objectives such as the secret-key rate (SKR) of a QKD protocol, or the joint rate region across competing entanglement flows, are non-linear in fidelity and generation rate, breaking the additive-reward assumption that standard reinforcement learning (RL) relies on. And classical-communication delays between nodes mean that any node’s view of the network is necessarily stale, so the underlying decision problem is partially observable rather than fully Markovian. Recent work therefore turns to RL, graph neural networks, and partially observable MDP (POMDP) methods to discover network control policies, alongside machine-learning approaches to characterize the links and nodes those policies must operate on.
RL for entanglement distribution and repeater control. The starting point for learning-based repeater control is the finite MDP formulation of Iñesta et al. [311], which models a homogeneous repeater chain with per-memory cutoffs that discard stored entangled links once their fidelity has decayed too far. The state encodes the ages of stored entangled links, and the action at each time step is the set of nodes that perform a swap. Value iteration and policy iteration yield globally optimal policies for end-to-end entanglement delivery time, outperforming the standard swap-as-soon-as-possible (swap-asap) heuristic. The framework assumes global instantaneous knowledge of the MDP state, an idealization that subsequent work systematically relaxes.
Reiß and van Loock [312] apply deep RL to a four-segment repeater chain targeting the BB84 SKR. The generalized non-additive objective they adopt is one whose equivalence to the true SKR they cannot rigorously establish. Haldar et al. [313] use Q-learning on inhomogeneous chains to discover dynamic, state-adaptive cutoff and node-collaboration policies that improve on swap-asap in both waiting time and fidelity, identifying global knowledge of the chain state as the structural source of advantage. Classical-communication costs, however, remain outside their model.
Subsequent work targets that cost directly. Haldar et al. [314] introduce quasi-local policies for multiplexed repeater chains, in which each node acts on information from a bounded sub-region of the network; this sharply reduces classical-communication overhead while remaining competitive with fully global strategies. Li et al. [315] take the complementary centralized route, folding classical-communication delays into the MDP state via an action–result history; their RL policy outperforms the natural wait-for-broadcast generalization of swap-asap in the high-success-probability regime, where partial-information decision-making outweighs the cost of acting on stale information. Mobayenjarihani et al. [316] characterize the same wait-versus-act trade-off analytically through an optimistic purification scheme in which nodes proceed without waiting for heralding messages, trading classical-communication latency against the risk of acting on failed pairs. Casado et al. [317] extend the RL framing to entanglement-distribution decisions on general network topologies.
Yau et al. [318] return to the question of the objective itself, developing an RL framework that directly optimizes non-linear, differentiable application-driven objectives such as the BB84 SKR. Their MDP models two-node entanglement distribution with multi-memory configurations, multiplexing, and distillation, and includes classical-communication-induced uncertainty about distillation outcomes as part of the state. To handle the non-additive objective, they extend standard policy-gradient methods to optimize functions of multiple expected discounted returns, capturing the joint dependence on fidelity and generation time. The resulting policies improve the SKR over threshold-based baselines by up to 23% in the parameter regimes considered.
Routing, resource allocation, and protocol primitives. Beyond a single chain, networks of repeaters serving multiple users introduce decisions about which path to use for each request, how to schedule competing requests against finite memory, and how to revise routing when links fluctuate. Early deep RL approaches address these directly through deep Q-networks for request scheduling on qubit-limited grids [319], proximal policy optimization (PPO) agents that exploit the non-additive composition of quantum errors to find higher-fidelity routes than classical shortest-path algorithms [320], and Q-learning policies that cache unused entangled links and proactively swap commonly used segments [321]. More recent work confronts larger topologies and time-varying links: graph neural networks with local message passing yield policies competitive with global-information baselines on both terrestrial and satellite topologies [322], [323], and belief-state methods cast routing as a POMDP with formal convergence guarantees under time-varying decoherence [324].
These controllers depend on resource primitives that are themselves under active study. Vardoyan and Wehner [325] extend classical network utility maximization to quantum networks, providing the formal framework for jointly allocating entanglement rate and quality across competing users with fairness encoded through log-composition of per-route utilities. Related work develops the protocol- and design-level primitives these controllers use: entanglement buffering with closed-form availability and consumed-fidelity bounds for any purification policy [326], and machine-learning surrogates of expensive network simulators for tuning memory allocations and protocol parameters at a fraction of the simulation cost [327].
Network characterization. The controllers and routing agents above presuppose channel parameters that, in practice, must be estimated from measurements without direct access to the internal links. Quantum network tomography (QNT) addresses this problem by inferring per-link channel parameters from end-to-end measurements at peripheral nodes, in analogy to its classical counterpart but contending with quantum-specific obstacles such as the multiplicative composition of channel parameters and the sign ambiguities it induces. Guedes de Andrade et al. [328] formalize QNT for star topologies with single-parameter Pauli channels and characterize the regimes in which their multicast-based estimators achieve identifiability. Wang et al. [329] extend the framework to arbitrary topologies and full Pauli channels via a Mergecast protocol that entangles peripheral qubits at an intermediate node, paired with a progressive etching procedure that identifies internal channels working from the periphery inward; they further give estimators for state-preparation and measurement errors and validate the combined workflow under photon loss and memory decoherence.
Machine learning offers a complementary route that does not assume a fixed channel class. Mukherjee et al. [330] demonstrate a multi-stage neural-network pipeline on a Rydberg-array black box that classifies the number of network nodes, regresses their positions, and recovers the system–environment Hamiltonian and Lindblad operators from a single-time snapshot of excitation transport, with reconstruction degrading once decoherence becomes comparable to the mean dipolar coupling. At the device layer beneath such a network, Canonici et al. [331] apply supervised learning to estimate physical noise parameters of a neutral-atom processor from occupation-probability measurements. Together these directions cover targets ranging from device-level noise to network-level open-system models, supplying the channel parameters and cutoff times the controllers above treat as inputs.
Physical-layer networking optimization. Beyond network-level tasks, AI also has significant potential to enhance physical-layer protocols for quantum networking and transduction. Early work by Krastanov et al. [332] used genetic algorithms to optimize entanglement purification protocols for qubit systems. More recently, Zhao et al. [333] introduced LOCCNet, a machine-learning framework for the design and optimization of distributed quantum information-processing protocols, including entanglement purification. This approach was later extended to dynamic LOCCNet [334], which improves scalability by constructing larger LOCC protocols from smaller adaptive modules. In a hardware-oriented setting, Zhang et al. [335] developed a qubit–qumode hybrid variational quantum circuit to optimize the physical-layer distillation of entangled qubits from noisy entangled qumodes, enabling improved performance over conventional time-bin approaches. Later on, similar technique are applied to variational quantum transduction [336]. Another related direction is variational quantum network optimization: Doolittle et al. [337] employed variational quantum circuits to optimize network-level performance measures such as nonlocality, using noisy qubit channels as the underlying network model.
Outlook. The picture above is largely one of single-agent learning over a fixed substrate. Several directions push beyond that frame. Distributed and multi-agent learning over quantum networks is one: DeRieux and Saad [338] propose entangled quantum multi-agent reinforcement learning, in which agents share an entangled critic over a quantum channel and coordinate without exchanging local observations. Quantum-secured federated learning over QKD networks goes the other way, with the network’s QKD service underwriting classical aggregation rather than itself becoming the learning target. Both directions sharpen, rather than resolve, the open challenges that already cut across the section. Chief among them is scaling RL controllers from elementary links to long multi-hop chains under realistic partial observability, and ensuring that the operating regime of a deployed controller does not silently invalidate the channel statistics on which it was trained — characterization and control must ultimately be co-designed rather than staged. Benchmarking AI-discovered policies against analytic baselines on standardized quantum-network simulators, and bridging the gap between elementary-link RL and continuous multi-hop network operation remain practical priorities.
Many machine learning workloads rely on a small set of linear algebra primitives. Two examples are solving linear systems and extracting spectral structure through eigendecomposition or singular value decomposition. Quantum algorithms can sometimes accelerate these primitives by exploiting superposition and entanglement in exponentially large Hilbert spaces. The usual strategy is to place a quantum subroutine inside a larger ML pipeline. For instance, a quantum linear system algorithm (QLSA) can solve the normal equation in least-squares regression, or serve as a subroutine inside a larger optimization routine. Whether the speedup survives in practice depends on how the data are encoded, what access model is assumed, and how much information must be read out at the end. This subsection traces the development from the foundational HHL algorithm to modern QLSAs and then surveys applications in optimization, machine learning, and scientific computing. For a concise primer on this algorithmic family, see [33].
HHL Algorithm: A Foundational Quantum Linear Systems Solver. The Harrow-Hassidim-Lloyd (HHL) algorithm [339] is the foundational quantum algorithm for the linear systems problem. Formally, it addresses a state-preparation version of this problem: given a matrix \(A\) and an efficiently preparable input state \(|b\rangle\), the goal is to prepare a state \(|x\rangle\) satisfying \[\left\|\,|x\rangle-\frac{A^{-1}|b\rangle}{\|A^{-1}|b\rangle\|}\right\| \le \varepsilon ,\] where \(\varepsilon\) is the target error in the output state [33], [339]. This definition makes the key point explicit: HHL returns a quantum state that approximates the normalized solution. In the usual HHL input model, \(A\) is an \(s\)-sparse Hermitian \(N \times N\) matrix. This means that each row has at most \(s\) nonzero entries and that the row can be queried efficiently: given a row index, one can identify the nonzero column indices and obtain their values in time \(\mathcal{O}(s)\) [339]. The primer of [33] formalizes this access model and explains why it is needed for efficient Hamiltonian simulation.
The original HHL algorithm performs matrix inversion by working in the eigenbasis of \(A\). It encodes the input vector \(\mathbf{b}\) as a quantum state \(|b\rangle\), treats \(A\) as a Hamiltonian whose evolution can be efficiently simulated, and uses quantum phase estimation to coherently estimate the eigenvalues of \(A\) into an ancillary register. Conditioned on these eigenvalue estimates, the algorithm applies a controlled rotation to another ancilla state whose successful post-selection weights each eigencomponent by the inverse of its eigenvalue. Applying inverse phase estimation then uncomputes the eigenvalue register, leaving a normalized quantum state proportional to \(A^{-1}|b\rangle\). Thus HHL implements matrix inversion by leaving the eigenvectors unchanged while inverting the corresponding nonzero eigenvalues, producing a quantum state from which selected properties of the solution can be estimated by measurement.
If the input state \(|b\rangle\) can also be prepared efficiently, then the original HHL runtime scales as \(\widetilde{\mathcal{O}}(\log(N)\, s^2 \kappa^2/\varepsilon)\), where \(\kappa\) is the condition number of \(A\) [339]. This can yield an exponential improvement in the system size \(N\) when the quantum state output is itself a useful final representation. Subsequent work by [340] improves the precision dependence, reducing the runtime to \(\mathcal{O}(\log(N)\kappa \log(1/\varepsilon))\) up to the same model-dependent overheads.
This improvement uses the linear combination of unitaries (LCU) technique. In LCU, the inverse transformation is approximated by a weighted sum of efficiently implementable unitaries, often built from short-time evolutions under \(A\). By combining these unitaries coherently, the algorithm applies an approximation to \(A^{-1}|b\rangle\) directly, without estimating every eigenvalue one by one. This idea foreshadows the modern QLSA methods discussed below.
HHL and QLSAs for Machine Learning Applications. The HHL algorithm and its successors apply to a broad range of machine learning tasks in which the computational bottleneck is a linear algebra primitive. Some of the earliest examples come from regression and classification, but recent work continues to expand these ideas to broader regression families, multi-class classification, and modern generative models. We organize these applications by task type:
Regression. In least-squares regression, one seeks a parameter vector \(\boldsymbol{\theta}\) that minimizes the squared residual norm \(\|X\boldsymbol{\theta}-\mathbf{y}\|_2^2\). The corresponding first-order optimality condition is the normal equation \((X^T X)\boldsymbol{\theta} = X^T \mathbf{y}\), a linear system that HHL can address when the data matrix \(X\) and target vector \(\mathbf{y}\) admit efficient quantum state preparation. Building on this idea, [341] gives a quantum algorithm for least-squares fitting that efficiently estimates fit quality and, in many cases, recovers a concise fitting function. More recently, [342] extends the discussion beyond standard linear fitting to a broader family of regression tasks, including ridge-, Lasso-, Huber-, and \(\ell_p\)-type objectives, and shows that quantum techniques can provide up to a quadratic improvement in the sample parameter under their structured access assumptions.
Classification. [343] introduces a quantum least-squares support vector machine (LS-SVM) that applies HHL-style matrix inversion to an augmented system built from the training-data kernel matrix. Under strong data-access and low-rank assumptions, the method achieves logarithmic scaling in both feature dimension and training-set size. Subsequent work refines this approach. [344] studies optimized HHL circuits for LS-SVM classifiers, with an emphasis on reducing circuit size and execution cost. The multi-class study of [345] compares quantum-kernel QSVMs with HHL-based LS-SVM schemes using two reductions to binary classification. The first is a one-vs-rest (OvR) construction, in which one binary classifier is trained for each class against the union of all remaining classes, and the predicted label is chosen from the classifier with the largest decision score. The second is a two-step hierarchical construction. In the three-class astronomical-object task built from a reduced Sloan Digital Sky Survey (SDSS) dataset, the first classifier isolates one chosen class from the other two, and a second classifier distinguishes between the two remaining classes. That comparison highlights a practical tradeoff: the HHL-based route retains attractive asymptotic scaling, but it is more sensitive to noise and tends to underperform the kernel-based alternative on the reduced SDSS benchmark considered in that work.
Limitations and Practical Constraints. The applications above share two bottlenecks. First, classical data must be encoded efficiently into quantum states, often through QRAM-style assumptions. Second, the quantum output is a state, so extracting a full classical solution may require enough measurements to erase the speedup. QLSA-based advantages are therefore most relevant when expectation values or other partial functions of the solution are sufficient.
HHL and related QLSAs also face practical constraints emphasized early by [346] and summarized in later QLSA surveys such as [33]. The original HHL algorithm requires fault-tolerant quantum computation with long coherence times to implement quantum phase estimation and its inverse reliably [339], which makes it unsuitable for current NISQ devices [14]. Its runtime depends quadratically on the condition number \(\kappa\) of the matrix \(A\), scaling as \(\mathcal{O}(\kappa^2/\varepsilon)\), and later QLSAs still retain condition-number dependence even when they improve this factor to linear. Data encoding remains another bottleneck: preparing classical data as quantum states using QRAM or related access models may require \(\mathcal{O}(N)\) operations and can negate the speedup [346], [347]. [347] further argue that most asymptotic quantum linear-algebra advantages disappear with active QRAM architectures, where every query requires external intervention and control. Scalable passive QRAM would need memory routing to proceed after the query is initiated without further external input or energy, which rests on demanding physical assumptions.
Readout creates a second practical barrier. Measurements of the output state \(|x\rangle\) reveal only partial information, and reconstructing the full solution may require \(\mathcal{O}(N)\) runs of the algorithm [33], [346]. Applicability is further limited by the need for an efficient quantum representation of \(A\), such as sparse access or block-encoding; arbitrary matrices do not admit such representations without prohibitive overhead. Later dequantization results sharpen this issue: for low-rank linear systems and related QML pipelines, quantum-inspired classical algorithms can sometimes recover polylogarithmic dimension dependence under sampling-access assumptions analogous to state-preparation assumptions [348], [349]. These limitations mean that HHL is best suited for structured problems with favorable condition numbers, efficient data encoding, and applications that require only partial information about the solution. The algorithmic developments below are best understood against these constraints.
Beyond linear algebra primitives. Recent work extends QLSAs to more complex ML workloads. [350] develops fault-tolerant quantum algorithms for sparse, sufficiently dissipative large-scale models with small learning rates, showing that certain stochastic-gradient-descent dynamics can be simulated in time polylogarithmic in the model size. [351] gives a particularly current example on the generative-model side by formulating quantum algorithms for representative diffusion probability models [31], [352]–[354], and by recasting higher-order ODE solvers such as DPM-solver-\(k\) [355] and UniPC [356] into quantum linear-solver and Hamiltonian-simulation primitives. These extensions illustrate ongoing efforts to carry quantum algorithmic advantages beyond core linear algebra, while maintaining rigorous specification of the access models and problem structures under which speedups hold.
Beyond HHL: Modern Quantum Linear System Algorithms. Modern QLSAs are built around two shifts. First, as recalled in the computational preliminaries in Sec. 2.3, block-encoding [34], [35] gives a flexible matrix access model. A matrix \(A\) is embedded as a subblock of a unitary \(U_A\) such that \((\langle 0| \otimes I)\, U_A\, (|0\rangle \otimes I) = A/\alpha\) for some normalization factor \(\alpha\). This replaces the sparse Hamiltonian-simulation oracle assumed in HHL with a more composable primitive.
Second, modern algorithms can approximate the matrix function \(f(A)=A^{-1}\) directly as a polynomial or rational transformation [340], [357], thereby bypassing QPE. The LCU-based algorithm of [340] is an early concrete example: it represents \(A^{-1}\) as a linear combination of unitary operators derived from Hamiltonian simulation and achieves the first exponential improvement in \(\varepsilon\) over HHL. The QSVT framework [357] later gives a general mechanism. Given a block-encoding of \(A\), QSVT applies a bounded polynomial \(p(x)\) to the singular values of \(A\) using \(\mathcal{O}(\mathrm{deg}(p))\) applications of \(U_A\) and its inverse, interleaved with single-qubit signal-processing rotations. To solve \(Ax=b\), one chooses \(p(x)\approx 1/x\) on the spectral interval \([1/\kappa,1]\), which yields the optimal \(\mathcal{O}(\kappa\log(1/\varepsilon))\) query complexity. Building on this perspective, [358] introduces eigenstate filtering, and [359] further develops fast inversion methods and preconditioned quantum linear system solvers within the block-encoding framework.
A different line of work uses integral and adiabatic ideas. The Laplace transform identity \(A^{-1} = \int_0^\infty e^{-At}\, dt\) expresses the inverse as a continuous superposition of matrix exponentials. [360] discretizes such an integral via quadrature rules and approximates \(A^{-1}|b\rangle\) as a linear combination of time-evolved states \(e^{-At_j}|b\rangle\) at quadrature nodes \(t_j\). Each term can be implemented through Hamiltonian simulation. [361] achieves the same optimal \(\mathcal{O}(\kappa \log(1/\varepsilon))\) complexity through a combination of the discrete adiabatic theorem and qubitized quantum walks. [362] shows that this approach is especially natural for structured systems such as low-rank tensor-sum linear systems arising from discretized partial differential equations.
Most recently, [363] develops kernel reflection, a conceptually simple quantum linear-system solver that extends the eigenstate-filtering idea of [358]. When the solution norm \(\|x\|\) is known, a single reflection about the kernel of an augmented linear operator suffices. When \(\|x\|\) is unknown, it can be estimated with \(\mathcal{O}(\log\log\kappa)\) kernel projections. These modern QLSAs achieve the same information-theoretically optimal \(\mathcal{O}(\kappa \log(1/\varepsilon))\) scaling [339], but they use different strategies. LCU and QSVT treat inversion as a directly implementable matrix function, adiabatic methods follow continuous paths in solution space, and kernel reflection projects onto the solution subspace through a small number of reflections. Together, these developments turn quantum linear system solving from a single algorithm into a broader toolkit.
Modern QLSAs also serve as computational engines inside larger machine-learning and scientific-computing pipelines. In optimization, many ML training procedures reduce to structured programs whose Newton steps require solving linear systems. Quantum interior point methods use QLSAs inside each iteration: [364] initiate this direction for linear and semidefinite programs, and [365] develop a predictor-corrector variant.
Later work handles the fact that QLSAs return approximate linear solves. Inexact-feasible formulations preserve primal feasibility despite QLSA error [366], [367]. Related ideas extend to quadratic [368] and semidefinite [369] optimization. Iterative refinement controls growing condition numbers near optimality [370], and preconditioning reduces condition-number dependence in the duality gap [371]. This line culminates in an optimally scaling QIPM framework for dense large-scale linear optimization [372]. A different route, due to [373], uses Grover-accelerated leverage-score sampling to approximate the Hessian and gradient of the barrier function. Quantum SDP solvers based on matrix multiplicative weight updates form a complementary direction [374]–[376].
QLSAs also appear as subroutines in recommendation systems [377], quantum gradient descent for least squares [378], and scientific-computing tasks such as discretized differential equations [379]. Across these examples, the lesson is consistent: QLSAs are most compelling when the linear system has favorable structure, the oracle model is realistic, and the end-to-end pipeline still has an advantage after data loading and readout are included.
Other Quantum Linear-Algebra Primitives for ML. Not all quantum speedups for ML linear-algebra tasks are rooted in solving linear systems. Two notable examples rely on fundamentally different quantum primitives.
For dimensionality reduction, [380] proposes quantum principal component analysis (qPCA) for unknown low-rank density matrices. qPCA uses density matrix exponentiation to extract dominant eigenvectors and eigenvalues in quantum form with runtime polylogarithmic in the dimension under the paper’s access assumptions.
For topological data analysis, [381] gives quantum algorithms that estimate Betti numbers across persistent homology scales and diagonalize the combinatorial Laplacian via quantum simulation—again without reducing the problem to a linear system. Under QRAM-style access assumptions, these routines can offer exponential improvements over the best classical methods known at the time. [382] provides complexity-theoretic evidence that this TDA formulation is likely resistant to dequantization, making it a comparatively strong candidate for quantum speedup.
Quantum learning advantages are most plausible when the learning task matches the quantum resources used for data access, encoding, and processing. Evidence for this structure dependence appears in quantum sample complexity, data access models, and the geometry of quantum feature maps. A growing body of results indicates that the clearest learning advantages arise when the data themselves come from quantum physical systems.
Learning Complexity and Sample Scaling Laws The classical PAC framework [383], [384] gives sample complexity \(\Theta(d_{\mathrm{VC}}/\varepsilon)\) in the realizable setting and \(\Theta(d_{\mathrm{VC}}/\varepsilon^2)\) in the agnostic setting, suppressing logarithmic dependence on the confidence parameter \(\delta\). Here \(d_{\mathrm{VC}}\) is the VC dimension, which measures the richness of the hypothesis class. Arunachalam and de Wolf [385], [386] prove that quantum and classical sample complexity coincide up to constant factors in both settings; exponential gaps appear only in query complexity. Huang, Kueng, and Preskill [387] further show that potential quantum gains are restricted in the average-case prediction setting.
Quantum state prediction gives a related lesson. The classical shadow framework [42] predicts \(M\) linear functions with \(O(\log M \cdot \max_i\|O_{i,0}\|^2_{\mathrm{shadow}}/\varepsilon^2)\) samples, where \(\require{physics} O_{i,0}=O_i-\Tr(O_i)\mathbb{I}/2^n\). For weight-\(k\) Pauli observables under local Pauli measurements, the shadow norm scales as \(3^k\), so local observables can be estimated efficiently while global operators can remain exponentially costly. Generalization analyses point in the same direction. Caro et al. [388] bound generalization error by \(\sqrt{T/N}\), tightening to \(\sqrt{K/N}\) when only \(K\!\ll\!T\) gates are active, while Gil-Fuster et al. [389] show that these uniform bounds can substantially overestimate true error because QNNs can fit random labels. Consistent with these observations, the expressivity–trainability relationship is nonmonotonic [390]: increasing expressive capacity does not necessarily improve learning performance. These results indicate that quantum models rarely achieve advantage through reduced sample complexity alone.
Provable Quantum Benefits in Physical Learning. When the learning target is a quantum system, the measurement strategy often determines the access complexity. Aaronson [391] establishes shadow tomography: \(M\) binary observables require only \(\tilde{O}(\varepsilon^{-4}\log^4\!M\log D)\) copies, improved from \(\varepsilon^{-5}\) using the online framework of [392]. Huang et al. [43] prove an information-theoretic exponential gap between separable and entangled measurements, verified experimentally on 40-qubit superconducting processors. Chen et al. [65] show that access complexity interpolates continuously: \(k\) memory qubits move the learner between classical and fully quantum sample complexity.
Hamiltonian learning shows similar structure. Haah–Kothari–Tang [393] achieves \(O(\log N/(\beta\varepsilon)^2)\) from high-temperature Gibbs states, while Huang et al. [394] attain Heisenberg-limited \(1/\varepsilon\) scaling with entangled sensors. Tang’s dequantization program [395], [396] shows that quantum speedups on low-rank classical data are often classically emulable; robust advantage persists mainly for quantum-structured or high-rank data.
Very recently, [6] establish a provable exponential advantage on classical data by showing that a polylogarithmic-size quantum computer can perform classification and dimension reduction on massive datasets that any sub-exponential-size classical machine cannot match. Their primitive, quantum oracle sketching, accesses classical data in superposition using only random samples, bypasses QRAM entirely, and is validated on single-cell RNA-seq and sentiment classification with fewer than 60 logical qubits.
Expressivity, Trainability, and the Limits of High-Dimensional Hilbert Spaces. High-dimensional Hilbert spaces alone do not guarantee learning benefits. As discussed in detail in Sec. 3.2 (see in particular the analyses of quantum kernels and barren plateaus), quantum feature maps can embed data into exponentially large spaces, yet the resulting models may suffer from kernel concentration [74], barren plateaus [73], [75], [81], and noise-induced trainability loss [397]. The main tension is that highly expressive circuits can be hard to train, while circuits that avoid barren plateaus often have restricted algebraic structure and may be classically simulable [94]. Structured architectures such as equivariant circuits [398], [399] offer one route forward by constraining the dynamical Lie algebra while preserving task-relevant expressivity. In infinite-dimensional continuous-variable systems, linear optical models are efficiently trainable [400], whereas universal discrete-continuous hybrid models exhibit a distinctive energy-dependent barren plateau phenomenon [401].
Outlook The structure-dependence principle emerges consistently across all three axes: quantum gains in learning arise from structural alignment, not raw dimensionality. Quantum sensing for learning is supported by rigorous information-theoretic results and hardware verification [43]. Quantum processing for learning faces the trainability simulability dilemma [75], [94]. Key open problems: (i) existence of a simultaneously non-simulable and barren-plateau-free circuit family; (ii) whether quantum data gains compound iteratively; (iii) beyond-worst-case regimes for variational training [402]; (iv) quantum sample complexity gains under structural distributional assumptions [385].
Building on recent tutorials and surveys of quantum neural network (QNN) architectures and models, such as [15], this subsection benchmarks QNNs against classical neural networks (NNs) by addressing three questions: (i) What provable quantum benefits of QNNs over classical NNs are established theoretically? (ii) Do these provable benefits translate into practical gains on common, real-world workloads? (iii) What gaps remain between theoretical analysis and practical demonstrations? We survey existing benchmarking results comparing classical and quantum models on representative tasks including generative modeling, classification, and time-series prediction to assess whether, and under what conditions, such gains manifest in practice.
Theoretical Analysis of Provable Quantum Benefits. For theoretical frameworks of provable quantum benefits in QNNs, we present four main directions.
One line of theory links quantum benefits in QNNs to quantum contextuality [403]–[405]. Contextuality is a foundational feature of quantum mechanics: the outcome of a measurement can depend on which other compatible measurements are performed jointly, and no classical hidden-variable model can reproduce this dependence [406], [407]. [405] show that a quantum-enhanced hidden Markov model (HMM) with a \(D\)-dimensional latent space requires a classical HMM with at least \(D^{\Omega(\log D)}\) hidden states to simulate, where \(\Omega(\cdot)\) denotes an asymptotic lower bound. Equivalently, an \(n\)-qubit quantum model can require \(2^{\Omega(n^2)}\) classical states. The gap arises from quantum contextuality, demonstrated via the Mermin–Peres magic square construction, and is further validated through numerical experiments.
Building on this idea, [403] introduce contextual recurrent neural networks (CRNNs). They prove unconditionally that CRNNs with \(\mathcal{O}(n)\) qumodes can express distributions that require \(\Omega(n^2)\)-dimensional latent spaces in any reasonable classical sequence model. This gives a quadratic memory gap and the first unconditional expressivity benefit of a quantum neural network over a classical neural network on classical data. Considering trainability barriers, [404] propose \(k\)-hypergraph recurrent neural networks (\(k\)-HRNNs). On the \((l,n,k)\)-hypergraph stabilizer measurement task, a \(k\)-HRNN with \(\mathcal{O}(n)\) qumodes solves the task perfectly, while any classical network satisfying smoothness assumptions requires a latent space of dimension at least \(\binom{n}{k}-1\). For constant \(k\), this yields an \(\Omega(n^k)\)-versus-\(\mathcal{O}(n)\) memory gap, with the polynomial degree controlled by the choice of \(k\).
Another line of work emphasizes entanglement as a computational resource [408]. [408] show that a quantum model with \(\mathcal{O}(1)\) parameters and \(2n\) Bell pairs achieves a perfect score on a magic-square translation task, whereas any communication-bounded classical model requires \(\Omega(n)\) parameters to achieve a score above \(2^{-o(n)}\). This constant-versus-linear gap is robust to depolarizing noise up to a threshold strength \(p^{*}\approx 0.0064\).
A third approach realizes benefits by using QNNs to instantiate quantum algorithms that already admit provable speedups. [409] show that any concept class efficiently learnable in the quantum statistical query (QSQ) model [410] can also be efficiently learned by parameterized quantum circuits. Because QSQ can solve parity learning, which requires exponentially many classical statistical queries, QNNs inherit this exponential benefit over classical SQ learners for such hard concept classes.
Finally, QNNs can be studied with classical learning-theoretic tools. [390] introduce the effective dimension, a capacity measure derived from the Fisher information matrix [411], and prove a generalization bound in terms of this quantity. They show empirically that certain QNN architectures achieve higher effective dimension than comparable classical feedforward networks with the same parameter count, and train faster. This framework is largely task-agnostic.
Overall, existing theoretical studies of quantum benefits in QNNs make impressive progress in delivering clean, provable results with clear physical interpretations. They also exhibit a notable characteristic: most results are inherently task-specific, and they establish benefits on carefully constructed problem families often based on contextuality, entanglement-assisted nonlocal correlations, or translation-style formulations.
Practical Benchmarking between QNNs and NNs. A fair comparison between QNNs and classical NNs requires careful benchmark design. As noted in [412], QML benchmarks can be distorted by unfair baselines, inappropriate evaluation metrics, or metrics that miss quantum-specific gains. The authors therefore emphasize careful experimental design and principled evaluation protocols. To support standardized benchmarking, [413] develop LazyQML, a Python library for systematic comparison of quantum and classical learning models. These methodological contributions provide the basis for rigorous empirical evaluation of QNNs.
The theoretical results above establish quantum benefits for specific, often abstract, task families. Practical advantage is less settled. Current benchmarking studies show a more nuanced picture: quantum models can be competitive in some regimes, but the benefits depend strongly on the task, data structure, hardware assumptions, and classical baseline. We focus on three common workload categories: generative modeling, classification, and time-series prediction.
Generative modeling is one of the most actively benchmarked areas. [414] compare quantum and classical generative models, define evaluation protocols, and identify conditions under which quantum generative models show gains. Building on this framework, [415] characterize performance across architectures and datasets, revealing scalability patterns and tradeoffs. [416] study the generalization behavior of quantum circuit Born machines and find regimes in which they generalize beyond the training data. [417] demonstrate quantum generative adversarial networks for image generation on quantum hardware, providing an early hardware implementation of quantum GANs. These works show that quantum generative models can be competitive in specific regimes, while scalability and generalization remain open challenges. Recently, [418] and others [419] propose quantum diffusion models for generative learning, with a focus on pure-state ensemble learning. This direction has stimulated studies in machine-learning applications [420]–[423] and theory [424], [425]. The possible quantum advantage of quantum diffusion models remains largely unexplored.
For classification tasks, benchmark studies report mixed results. [426] study quantum-kernel training for classification and find advantages for certain data characteristics, with benefits that remain dataset-dependent. [427] conduct a reproducibility study of quantum machine learning methods in fundus analysis, revealing challenges in reproducing quantum ML results and highlighting the need for standardized evaluation protocols.
For time-series prediction, [428] compare quantum and classical approaches using variational QML. They find that quantum models can match or exceed classical performance in some scenarios, especially when the time series has favorable structure, although quantum-circuit overhead often limits practical gains. Beyond standard accuracy metrics, [429] benchmark adversarially robust QML at scale. They find that quantum models can have different robustness properties from classical models, but the advantage is not universal and depends on the model architecture and attack strategy.
Together, these benchmarks show that QNN gains are highly context-dependent. They appear in specific tasks, data regimes, and problem sizes, especially when quantum feature spaces, entanglement, or interference align with the problem structure. Standardized tools such as LazyQML and carefully designed evaluation protocols are therefore essential for fair comparison between quantum and classical approaches.
Gaps between Theory and Application. The theory and benchmarks above point to the same conclusion: QNN performance is highly context-dependent, and a gap remains between provable benefits and real-world applications. Theoretical results often rely on abstract learning settings, idealized assumptions, or restricted models such as SQ/QSQ frameworks. These results identify possible mechanisms for quantum benefit, but they do not directly guarantee gains on common practical workloads. An important open question is how to use the theory to guide the design of efficient and practically useful QNNs.
| Reference | Quantum Property | Model |
|---|---|---|
| [430] | Entanglement (entanglement entropy) | RBM (Restricted Boltzmann Machine) |
| [431] | Quantum entanglement (entanglement measures) | Deep neural networks (various architectures) |
| [432] | Quantum field theory | Wide neural networks (infinite-width limit) |
| [433] | Entanglement (entanglement transition) | MPS (Matrix Product State) tensor network |
| [434] | Transfer entropy, O-information | MPS (Matrix Product State) tensor network |
| [435] | Artificial entanglement (entanglement entropy) | LLM (Large Language Models) |
| [436] | Quantum-inspired categorical frameworks | Neural networks (general architectures) |
Direct quantum algorithms for machine learning face significant practical challenges [346], [351], [396], [437]–[439]. A separate line of work uses quantum concepts to analyze classical machine learning architectures. We focus on how ideas from quantum information theory, especially entanglement, can be used to study classical neural networks. We call this approach quantum-inspired analysis of classical neural networks.
Quantum Entanglement. Quantum entanglement is a fundamental concept in quantum information theory that describes non-classical correlations between quantum systems. It explains why a subsystem can have nonzero entropy even when the full quantum system is in a pure state [440]. One standard way to quantify this effect is entanglement entropy [441], [442].
Entanglement Measures and Quantification. Entanglement in quantum systems is quantified using measures derived from quantum information theory. For a bipartite quantum system described by a density matrix \(\rho_{AB}\), the entanglement entropy is defined as the von Neumann entropy of the reduced density matrix: \[\require{physics} S(\rho_A) = -\Tr(\rho_A \log \rho_A),\] where \(\require{physics} \rho_A = \Tr_B(\rho_{AB})\) is the reduced density matrix obtained by tracing out subsystem \(B\). For a bipartite pure state, the Schmidt decomposition gives \[|\psi\rangle_{AB} = \sum_i \lambda_i |i\rangle_A \otimes |i\rangle_B .\] The Schmidt coefficients \(\lambda_i\) determine the entanglement, and the corresponding entanglement entropy is \[S = -\sum_i \lambda_i^2 \log \lambda_i^2 .\] Higher entanglement entropy indicates stronger non-classical correlation between the two subsystems. For mixed states, additional measures such as logarithmic negativity or entanglement of formation are often used.
Area Law and Volume Law Scaling. The scaling behavior of entanglement entropy with subsystem size plays a crucial role in understanding quantum many-body systems. For a subsystem \(A\) with linear size \(\ell_A\) in spatial dimension \(d_{\mathrm{sp}}\), the area law states that the entanglement entropy scales as: \[S(\rho_A) \sim \ell_A^{d_{\mathrm{sp}}-1},\] where \(d_{\mathrm{sp}}\) is the spatial dimension. This means that in one dimension (\(d_{\mathrm{sp}}=1\)), the entanglement entropy is constant (scales as \(\ell_A^0\)), while in two dimensions (\(d_{\mathrm{sp}}=2\)), it scales linearly with the boundary size. The area law holds for ground states of gapped local Hamiltonians and reflects the local nature of correlations in such systems. In contrast, the volume law describes systems where entanglement entropy scales with the volume of the subsystem: \[S(\rho_A) \sim |A|,\] where \(|A|\) denotes the size of subsystem \(A\). Volume-law scaling typically appears in highly entangled states, such as random states or thermal states at finite temperature. Critical systems at quantum phase transitions can exhibit logarithmic violations of the area law. In some systems, transitions between area-law and volume-law behavior signal changes such as many-body-localization transitions or measurement-induced phase transitions. Other quantum phase transitions involve a change from area-law scaling to logarithmic area-law violation.
Early work connects entanglement concepts from quantum information theory to machine learning. [430] interpret the restricted Boltzmann machine (RBM) as a variational ansatz for a quantum many-body wavefunction and ask what entanglement structures this ansatz can represent. They show that short-range RBMs exhibit area-law scaling, while long-range RBMs can exhibit volume-law scaling under certain constructions. In this setting, entanglement measures describe how efficiently the neural network represents a quantum state. [443] show an equivalence between the function realized by a deep convolutional arithmetic circuit (ConvAC) and a quantum many-body wavefunction. Building on these works, [431] apply quantum entanglement measures to deep convolutional and recurrent neural networks. Beyond entanglement, [432] connect neural networks with quantum field theory and use field-theoretic tools to analyze training dynamics, generalization, and expressivity, especially in the infinite-width limit.
Recent work applies these ideas to specific learning phenomena in AI. Grokking [444] refers to cases where a neural network generalizes only after an extended period of training. [433] analyze grokking in tensor-network machine learning and interpret it as an entanglement transition: the model moves from a less entangled phase to a more entangled phase that captures the structure of the problem. Complementing this work, [434] use transfer entropy and O-information to detect grokking in tensor-network multi-class classification. These information-theoretic measures track how information flows and is shared across different parts of the network during training.
More recently, [435] introduce artificial entanglement, defined as the entanglement entropy of neural-network parameters. They use this measure to study fine-tuning dynamics in large language models and numerically identify area-law and volume-law behavior. They also derive, under suitable limits, an attention Cardy formula that connects artificial entanglement entropy to the structure of attention patterns. The same work draws an analogy to the No-Hair Theorem in black hole physics to explain why low-rank updates can achieve strong performance with far fewer trainable parameters.
Beyond entanglement measures, [436] develop a compositional approach to explainable AI using quantum-inspired categorical frameworks, especially category theory and diagrammatic reasoning. Their goal is to analyze how neural networks compose simpler operations into complex behavior. We also treat tensor networks as a quantum-inspired framework for analyzing neural-network learning dynamics, and we review this line of work in Section 4.5.
To sum up, we organize these works in Table 4. These quantum-inspired perspectives offer new ways to analyze classical AI and form one thread in the broader physics-of-AI literature [445]–[447]. They also require care. Classical neural networks do not exhibit true quantum entanglement, so analogies between entanglement and classical correlations must be interpreted precisely. It also remains open when quantum-inspired measures lead to practical improvements in neural-network design or training. A useful next step is to connect these measures more directly to observable changes in optimization, generalization, or architecture design.
Tensor network methods are quantum-inspired machine-learning tools built from mathematical structures that originated in quantum many-body physics [448]–[451]. They can be implemented efficiently on classical computers and are useful for representing high-dimensional data with controlled correlation structure. Table 5 provides an overview of representative tensor network methods discussed in this section.
| Reference | TN Structure | Target Model / Task |
|---|---|---|
| Tensorizing Neural Networks for Compression and Efficiency | ||
| [452] | TT (MPS) | FC layers (VGG) |
| [453] | TT | Exponential machines |
| [454] | TT (MPS) | LLM fine-tuning (LoRA) |
| [455] | TT (MPS) | LLM (MoE) |
| [456] | Tensor factorization | Transformer attention |
| [457] | Tensor decomposition | Transformer \(Q/K/V\) |
| Tensor Networks as Standalone Learning Models | ||
| [458] | MPS | Supervised classification |
| [459] | MPS | Generative classification |
| [460] | PEPS | Image classification |
| [461], [462] | MERA + MPS | Classification and regression |
| [463] | TTN | ML with tree structure |
| [464] | Deep TTN | Image recognition |
| [465] | TN | Medical image classification |
| [466] | MPS | Time-series models |
| [467] | MPS | Unsupervised generative models |
| [468] | TTN | Generative models |
| [469] | TN | Unsupervised ML |
| [470] | PEPS | 2D generative models |
| [471] | TN | Continuous generative models |
| [472] | TN (MPS) | Privacy-preserving ML |
| Theoretical Equivalence between TNs and NNs | ||
| [473], [474] | Hierarchical TN | Deep CNNs |
| [431], [443] | TN | Deep NNs |
| [435] | MPS | LLM (PEFT) |
| [475] | TNS | RBM |
| [476] | Generalized TN | Probabilistic graphical models |
| [477] | TN/MPS | RNNs |
| [478] | Kronecker-factorized TN | Higher-order Transformers |
| [479] | Tensor attention | Transformers |
| [480] | High-order attention tensors | Transformer stack |
| [481] | TN | TN-ML models |
Important tensor network structures include matrix product states (MPS) [450], [482], tree tensor networks (TTN) [483], projected entangled pair states (PEPS) [451], Multiscale Entanglement Renormalization Ansatz (MERA) [484], and Branching MERA [485]. These structures represent high-dimensional tensors compactly by factorizing them into lower-dimensional tensors connected through virtual bonds; see [486] for a detailed introduction.
Their connection to QC comes from a shared mathematical origin. Tensor networks were developed to represent quantum many-body wavefunctions efficiently [448]–[451], where entanglement structure controls how large the representation must be. They are now used across QC applications, including simulation of quantum computation, quantum circuit synthesis, quantum error correction and mitigation, and quantum machine learning; see [487] for a detailed introduction. We therefore treat tensor network methods as quantum-inspired approaches.
Several surveys have examined tensor networks in the context of machine learning. [488] provide a comprehensive overview of tensorial neural networks (combinations of TNs and NNs) and cover how tensor decompositions are integrated into CNNs, RNNs, Transformers, GNNs, and LLMs for model compression, as well as their use in data fusion and multi-task learning. [489] focus on tensor networks as quantum-inspired machine learning models, highlighting their intrinsic interpretability grounded in quantum information theory and their potential for deployment on quantum hardware. [21] organize the field through an AI4Science and Science4AI dual perspective, reviewing both TN-assisted scientific simulation and TN-based neural network tensorization. [487] survey tensor networks specifically within QC and quantum machine learning. [22] focuses on low-rank tensor methods for the theoretical analysis of neural networks and connects tensor decompositions to neural-network theory. Our review places tensor network methods within the broader study of quantum information and artificial intelligence. We organize the discussion along three axes: tensorizing neural networks for compression and efficiency, tensor networks as standalone learning models, and theoretical equivalences between tensor networks and neural networks. We emphasize how quantum information-theoretic concepts, especially entanglement structure and bond-dimension constraints, inform both the design and analysis of machine learning architectures.
We first introduce the fundamental concepts, notation, and operations of multilinear algebra and tensor network theory.
Tensor Notation. We follow the notational conventions of [473]. Vectors are denoted by bold lowercase letters, e.g., \(\mathbf{v} \in \mathbb{R}^{d}\), with coordinates in regular typeface, e.g., \(v_i \in \mathbb{R}\). Matrices are denoted by uppercase letters, e.g., \(M \in \mathbb{R}^{d_1 \times d_2}\). Tensors (multi-dimensional arrays) of order \(n \geq 3\) are denoted by calligraphic letters, e.g., \(\mathcal{T} \in \mathbb{R}^{d_1 \times d_2 \times \cdots \times d_n}\), where \(d_i\) is the dimension of the \(i\)-th mode and specific entries are referenced as \(\mathcal{T}_{i_1, i_2, \ldots, i_n}\) with \(i_k \in [d_k] \mathrel{\vcenter{:}}= \{1, \ldots, d_k\}\). The number of modes \(n\) is the order of the tensor, generalizing scalars (\(n=0\)), vectors (\(n=1\)), and matrices (\(n=2\)). Superscripts index elements within a collection, e.g., \(A^{(k)}\) denotes the \(k\)-th tensor in a sequence.
Tensor Operations. Fundamental tensor operations form the computational building blocks of these networks. (i) Tensor contraction generalizes matrix multiplication to higher-order tensors by summing over shared indices: given \(A_{i\beta}\) and \(B_{\beta j}\), their contraction yields \(C_{ij} = \sum_{\beta} A_{i\beta} B_{\beta j}\). (ii) Reshaping (unfolding/matricization) reorganizes a tensor into a matrix by grouping modes: the mode-\(k\) unfolding of \(\mathcal{T} \in \mathbb{R}^{d_1 \times \cdots \times d_n}\) creates a matrix \(T_{(k)}\) by arranging the mode-\(k\) fibers as columns, enabling the application of standard linear algebra techniques such as SVD. (iii) The tensor product \(\mathcal{C} = \mathcal{A} \otimes \mathcal{B}\) combines two tensors into a higher-order tensor whose order is the sum of the orders of \(\mathcal{A}\) and \(\mathcal{B}\), without contraction over any shared index. Figure 3 summarizes these basic tensor objects and operations in the graphical language commonly used throughout tensor network theory.
SVD and Low-Rank Approximation. The Singular Value Decomposition (SVD) is the basic tool behind tensor network compression. Any real matrix \(M\) can be decomposed as \(M = U \Sigma V^T\), where \(U\) and \(V\) are orthogonal matrices and \(\Sigma\) is a diagonal matrix of non-negative singular values \(\sigma_i\). In tensor networks, SVD has two main roles. First, it enables truncation: retaining only the largest \(\chi\) singular values gives the optimal rank-\(\chi\) approximation of the matrix in Frobenius norm by the Eckart–Young theorem [490]. Applied to tensor matricizations at each bond, this truncation reduces the bond dimension by discarding weak correlations. Second, SVD produces canonical forms: it orthogonalizes local tensors and transforms the network into left- or right-canonical form, improving numerical stability and making the representation easier to manipulate.
Tensor Decomposition and Networks. A tensor network factorizes a high-dimensional tensor into a set of low-order component tensors connected by virtual bonds. The most common structure is MPS, also known as the Tensor-Train (TT) decomposition [491]. An order-\(n\) tensor \(\mathcal{T}\) is decomposed as: \[\mathcal{T}_{i_1, i_2, \ldots, i_n} = \sum_{\alpha_1, \dots, \alpha_{n-1}} A^{(1)}_{i_1, \alpha_1} A^{(2)}_{\alpha_1, i_2, \alpha_2} \cdots A^{(n)}_{\alpha_{n-1}, i_n},\] where \(A^{(k)}\) are the local core tensors and \(\alpha_k\) are the auxiliary virtual bond indices of dimension \(r_k\) (the bond dimension). The boundary tensors \(A^{(1)} \in \mathbb{R}^{d_1 \times r_1}\) and \(A^{(n)} \in \mathbb{R}^{r_{n-1} \times d_n}\) are matrices, while the interior cores \(A^{(k)} \in \mathbb{R}^{r_{k-1} \times d_k \times r_k}\) (\(2 \leq k \leq n-1\)) are order-3 tensors. The total number of parameters scales as \(O(n d r^2)\), compared to the exponential scaling \(O(d^n)\) of the full tensor. The bond dimension \(r = \max_k r_k\) controls the expressivity: a larger \(r\) captures more correlations (entanglement) between variables, while a smaller \(r\) imposes a tighter information bottleneck. Figure 4 illustrates how SVD truncation induces low-rank tensor-network factorizations such as MPS, with the retained singular values directly controlling the bond dimension.
Tensor network methods have traditionally been applied to calculate ground- and low-lying excited states [492]–[494], thermal states [495]–[497], non-equilibrium steady states [498]–[500], and time evolution [501], of low-dimensional many body quantum systems with short-range interactions. Examples include strongly-correlated electron systems in condensed matter [502]–[507], chemical and molecular systems [508], [509], and cold atoms in optical lattices [510]–[513]. The general formalism of tensor networks, the different types of networks, their connection to entanglement in many-body quantum systems, and examples of applications are covered in several reviews [486], [514]–[519].
Many recent applications of tensor networks successfully demonstrate the efficient simulation of strongly correlated quantum systems beyond the low-dimensional and short-range interaction limits [520]. First, entanglement remains upper-bounded by a constant value in non-critical one-dimensional systems with long-range interactions with locally-bounded Hamiltonians [521], such as chains with dipolar interactions [522]–[524]. In addition, systems that feature effective long-range interactions when represented as one-dimensional lattices have been investigated with MPS representations of their states. Examples include impurity models [525]–[529], nonequilibrium systems connected to discretized reservoirs [530]–[532], and lattices embedded in optical cavities [533]–[539]. Crucially, some of the most remarkable demonstrations of the power of tensor networks have been achieved in simulations of high-dimensional interacting quantum systems. Notable examples are bilayer systems [540]–[542], as well as three-dimensional lattices when combined with mean-field approximations [543], [544]. Furthermore, efficient simulations of two- and three-dimensional systems incorporating contractions based on belief propagation have challenged the most recent claims of quantum advantage [545], [546].
A novel multidisciplinary avenue of tensor network applications, mostly based on MPS structures, corresponds to the simulation of high-dimensional partial differential equations (PDEs). Accurate numerical analysis of multiscale dynamical systems, such as turbulent flows [547], suffers from a complexity similar to that of exact diagonalization of many-body quantum systems. Namely, direct numerical simulation (DNS) of turbulence relies on resolving all relevant length and time scales of the dynamics, leading to an exponential complexity with the Reynolds number (the ratio of inertial and viscous forces). This complexity motivates the application of tensor network structures to compress the high-order tensors that encode the transport variables of turbulent flows. In particular, representing such variables as MPS [548] (also known as quantics MPS or tensor trains [491], [549]–[553]), and performing operations such as derivatives, element-wise multiplication [554]–[556], and Fourier transform within the MPS manifold [557]–[560], the nonlinear governing PDEs such as the Navier-Stokes equations can be efficiently solved. This quantum-inspired application was first demonstrated for incompressible turbulent flows in Ref. [561], where a 2D temporally-developing jet and a 3D Taylor-Green vortex were accurately simulated while requiring one order of magnitude fewer parameters compared to DNS. Such a significant compression was enabled by the notion of locality among length scales resulting from the energy cascade mechanism in turbulent flows. Further applications of MPS for PDEs and continuous functions include simulations of various types of classical fluid dynamics systems [562]–[573], quantum turbulence [574], [575], plasma physics [576]–[578], phonon transport in crystals [579], wave equations [580], and calculation of Green’s functions in quantum field theory [581], [582]. Efforts in these directions have also been made using tree tensor networks, also referred to as Tucker tensors [583], [584]. Whether other types of network geometries enable an efficient simulation of high-dimensional multiscale PDEs remains an open question.
To take advantage of emerging quantum computing resources, a natural strategy consists of performing classical tensor network simulations first and continuing them on quantum hardware via shallow circuits [585]. Recent works based on efficient MPS representation of flows at high Reynolds numbers [586]–[588] suggest that such states can be efficiently uploaded to quantum hardware via diverse strategies [589]–[597]. Moreover, the recent demonstration that matrix product operators can be well represented as quantum circuits[586], [598] indicates that tensor-network time evolution algorithms can be performed on the uploaded quantum MPS. Similarly to classical simulations [554]–[556], the main bottleneck of the quantum simulation of PDEs is the implementation of nonlinearities. However, dealing with nonlinear terms is much more complex in the quantum computing regime [599], [600]. This complexity comes mainly from the no cloning theorem and the exponential cost of measuring variables encoded in a large number of qubits. A proposal to overcome these limitations, built on MPS encoding, is based on variational algorithms using quantum nonlinear processing units [601]–[604]. Alternatively, variational quantum time evolution algorithms [605]–[607] could be used to solve PDEs [608], [609] based on an MPS Ansatz [610].
A fundamental application of tensor networks in machine learning is neural-network compression. Early work uses tensor-train (TT) decompositions to compress fully connected layers and to model higher-order feature interactions with far fewer parameters [452], [453], [611]. The same basic idea later extends to convolutional, recurrent, transformer, and graph models [612]–[626].
More recently, TT-style factorizations move into LLMs, where they are used for both model compression and parameter-efficient fine-tuning (PEFT) [627], [628]. Representative examples include TT-LoRA, which compresses LoRA-style updates [454], and TT-LoRA MoE, which combines TT structure with mixture-of-experts design [455]. This direction now sits within a broader family of tensor-network-inspired PEFT methods [629]–[639].
Tensor decompositions also help at inference time, for example through KV-cache compression [640] and production-oriented TN compression pipelines [641]. A closely related line of work tensorizes attention itself in addition to the weights: recent methods factorize long-context attention, compress the \(Q/K/V\) maps, or redesign attention kernels around tensor structure [456], [457], [642]. For a more comprehensive survey, see [488].
The predominance of MPS (or equivalently, TT) in neural-network compression is mainly due to its favorable balance between expressivity and computational tractability. Other topologies, such as tensor rings (TR), PEPS, and TTN, encode different structural biases and may be better matched to some architectures. Their use in large-scale language models nevertheless remains limited. The main obstacles are more expensive contractions, weaker alignment with common network geometries, and the lack of optimization pipelines as mature as those available for TT. Even so, these alternative geometries remain promising for targeted components such as hierarchical feedforward blocks or attention modules with structured correlation patterns.
Beyond compression, tensor networks can also serve as standalone learning models. Existing work spans supervised classification, generative modeling, and privacy-preserving learning [458]–[461], [463]–[465], [467]–[469], [471], [472].
[458] show that MPS can act as effective classifiers, especially when the data have sequential or locally correlated structure. This basic supervised-learning idea now extends in several directions. [459] combine generative and discriminative objectives in a tensor-network classifier. [460] use PEPS to handle two-dimensional spatial correlations, while [461]–[464] develop hierarchical and tree-based architectures that better capture multiscale structure. Domain-specific applications now include medical imaging [465] and time-series learning [466], where the network geometry aligns naturally with the data.
Tensor networks also form a useful family of generative models. MPS-based approaches learn high-dimensional distributions efficiently [467], [469], TTN variants capture multiscale correlations [468], and PEPS-based models extend the same logic to higher-dimensional data [470]. More recent work pushes these models to continuous distributions [471]. Beyond prediction and generation, tensor factorizations are also used to support privacy-preserving learning frameworks [472].
Challenges. Despite these advances, standalone tensor network models face several persistent challenges that limit their broader adoption. A fundamental limitation shared across nearly all tensor network architectures is the sensitivity to input ordering: MPS-based models require imposing a one-dimensional ordering on inherently multi-dimensional data such as images, and the choice of this ordering can significantly impact performance [458], [467]. While higher-dimensional tensor network structures such as PEPS [460], [470] and MERA-based architectures [461] alleviate this issue by natively accommodating spatial correlations, they introduce a different bottleneck: their contraction costs scale exponentially on classical hardware, severely limiting scalability [460]. The bond dimension remains a critical hyperparameter governing the expressiveness–efficiency tradeoff across all architectures; increasing it improves model capacity but incurs rapidly growing computational and memory costs [458], [466]. In generative modeling, early MPS-based Born machines suffer from exponentially decaying correlations that restrict their ability to model complex, long-range dependencies [467], [468], and the overall performance of tensor network generative models has historically lagged behind state-of-the-art neural network approaches, with prior work acknowledging that existing tensor network models “only work as a proof of principle” [469]. Although recent developments such as autoregressive MPS [469], deep tree tensor networks [464], and continuous Born machines [471] have substantially narrowed this gap, [464] still note that “current tensor networks are predominantly suited for simpler tasks and face limitations in computational efficiency and expressive power.” For time-series analysis, the MPS ansatz naturally aligns with sequential data structures [466], but its applicability to very long or high-dimensional time series remains constrained by bond dimension scaling. In privacy-preserving settings, the canonical form-based protection proposed by [472] currently applies rigorously only to MPS architectures, and extending such guarantees to more general tensor network families remains an open challenge. More broadly, tensor network models are inherently multilinear, lacking the nonlinear activation functions that endow deep neural networks with their expressive power, which may fundamentally limit their capacity for certain learning tasks, though this same property provides advantages in interpretability and theoretical analysis.
Collectively, these limitations suggest that the path forward for standalone tensor network models lies in exploiting their unique structural advantages: exact tractability of probabilistic quantities (e.g., partition functions and marginals), principled interpretability through entanglement measures, intrinsic privacy-preserving properties via gauge symmetry, and natural compatibility with quantum hardware. We argue that future research might focus on (i) developing hybrid architectures that combine the interpretability and structural guarantees of tensor networks with the nonlinear expressiveness of neural networks, (ii) designing adaptive tensor network topologies that automatically discover data-dependent optimal structures, and (iii) extending the theoretical privacy and generalization guarantees currently established for MPS to broader tensor network families, thereby enabling principled deployment in safety-critical and privacy-sensitive applications.
Recent theory connects TNs and NNs by treating both as parameterizations of high-order multilinear objects. TNs represent functions or distributions through explicit tensor contraction graphs, where topology and bond dimension impose structured low-rank constraints. In quantum language, these are also low-entanglement constraints. Many classical ML models can also be recast as generalized tensor factorizations once their computations are expressed over suitable feature maps. This gives a direct way to compare architecture, expressivity, and inductive bias [21], [22], [488].
Expressivity and depth efficiency via tensor decompositions. A seminal line of work establishes that the expressive power of deep architectures can be analyzed using tensor decomposition theory. In particular, deep convolutional-style computations correspond to hierarchical tensor decompositions, while shallower counterparts correspond to flatter decompositions, yielding formal depth-efficiency statements: for broad families of functions, hierarchical factorizations can achieve compact representations that would require exponentially many parameters in shallow forms [473], [474]. This “tensorization” perspective makes architectural motifs such as locality, pooling, and weight sharing amenable to rigorous rank-based analysis.
Entanglement-inspired measures of representational capacity. Another influential bridge interprets neural architectures as representations of many-body wavefunctions. This makes it possible to use entanglement-related quantities to characterize the complexity of functions they can represent. In this formulation, network design choices determine entanglement capacity, often through graph cuts or effective bond dimensions in the corresponding TN picture. This gives a principled account of how architectural inductive bias arises [431], [443]. Such ideas have recently resurfaced in the context of parameter-efficient fine-tuning, where effective rank/entanglement measures help explain why low-dimensional updates can still induce rich task-specific behavior [435].
Constructive correspondences: RBMs, graphical models, and sequence models. Beyond expressivity analysis, several works provide constructive translations between model classes. Restricted Boltzmann machines (RBMs) admit explicit equivalences with tensor network states (TNS), clarifying when and how RBMs can emulate particular TN topologies and how entanglement/rank constraints bound their representational power [475]. Probabilistic graphical models can likewise be embedded into generalized TN formalisms (often using copy/replication tensors), yielding a shared language for inference and supervised learning [476]. For sequential data, TN representations of quantum states can be mapped to recurrent-style computations, resulting in tensorial RNNs that inherit TN training intuitions and regularization mechanisms [477].
Transformers through tensor-network methods. Recent work has further enriched the theoretical toolkit connecting TNs and Transformers. Higher-order Transformers generalize attention to tensor-structured inputs using Kronecker-factorized interactions, offering favorable scaling for multi-axis data [478]. Learning-theoretic advances show that certain higher-order/tensor attention mechanisms admit provably efficient training, including near-linear-time gradient computation under suitable assumptions [479]. End-to-end analysis tools reformulate the Transformer stack via high-order attention tensors to unify attention, MLPs, normalization, and residual pathways into a single tensor-centric description [480].
Limits, assumptions, and when equivalence helps. Equivalence should not be interpreted as universality. Recent no-free-lunch results characterize regimes in which TN-based ML models cannot guarantee generalization without assumptions on data structure and encoding [481], and analyses of classical data distributions highlight cases where strict low-entanglement TN biases may be insufficient for efficient description [643]. Taken together, these developments suggest a nuanced conclusion for this survey: TNs provide a mathematically transparent language for understanding and designing neural architectures, but their benefits depend critically on matching TN topology and rank constraints to the compositional structure of the target data and task [22], [488].
Recent surveys in this area typically close with the challenges and outlooks that cut across subfields [11], [12], [644]. However we identify that AI for QI and QI for AI share several hard problem structures: (1) Both rely on indirect or noisy access to the object of interest; (2) Both operate under severe resource constraints; (3) Both still face a large gap between elegant theory and robust large-scale deployment. Concretely, three cross-cutting challenges appear as detailed below:
Benchmarking and Baselines. A first challenge is evaluation. In AI for QI, many methods are validated on simulators, simplified noise models, or narrowly defined hardware tasks. This makes it difficult to compare performance across platforms, code distances, control settings, sensing tasks, networking protocols, and experimental pipelines. In QI for AI, performance claims are often benchmarked against weak classical baselines, incomplete end-to-end resource accounting, or problem instances that already encode favorable structure. Reported improvements can therefore reflect a mismatch in evaluation protocol and may obscure whether a genuine algorithmic or scientific advantage has been demonstrated. Therefore, more standardized datasets, careful ablations, transparent classical baselines, and clearer resource reporting are needed. Just as importantly, the field would benefit from more negative results and stress tests. Failures under noise drift, distribution shift, limited training data, or larger system size are scientifically valuable and should become part of the evidence base.
Scalability, Data, and Hardware Realism. A second challenge is the transition from proof-of-principle settings to realistic scales. Across both directions, many appealing results are still established on small systems, synthetic datasets, idealized noise models, or asymptotic regimes far from current hardware. In AI for QI, the main bottlenecks are larger training requirements, online adaptation to drift, and real-time deployment limits set by classical feedback latency and experimental overhead. In QI for AI, the bottlenecks are data loading, fault-tolerance overhead, trainability, and the question of whether asymptotic speedups survive realistic constants and error rates. Bridging this gap will require tighter interaction between theory, simulation, and experiment. We expect increasing importance of methods fine-tuned on experimental QPU data, benchmark suites derived from real hardware traces, and analyses that report both asymptotic complexity and the regime in which a proposed advantage could plausibly appear. Future surveys will be most useful when they clearly separate mathematically interesting results, near-term deployable techniques, and genuinely fault-tolerant long-term opportunities.
Toward Co-Designed Quantum–AI Systems. The third challenge is co-designing quantum–AI system. The strongest results in this area emerge when the learning model, the physical constraints, the optimization loop, and the computational architecture are designed together. Examples already appear throughout this survey: geometry-aware training for variational circuits, adaptive decoders tuned to hardware noise, AI-assisted experimental control, AI-designed sensing and networking protocols, tensor-network methods that exploit problem structure, and quantum-inspired models that feed back into classical architecture design.
Looking ahead, we expect the most productive direction to be hybrid and hierarchical. In the near term, AI will likely continue to deliver practical gains in calibration, control, error mitigation, decoding, compilation, sensing, networking, and autonomous experimentation. In parallel, quantum computing may contribute more selectively to AI through specialized subroutines, new representational principles, and sharper theoretical insights into learning and expressivity. Over a longer horizon, the field may evolve toward AI-native quantum stacks in which learning algorithms, compilers, controllers, simulators, sensors, networked devices, and hardware are coupled in closed loop. At the same time, responsible progress will require restraint in claims of quantum speedup and greater clarity about uncertainty, deployment assumptions, and energy or infrastructure costs.
This survey reviews the interface between AI and QI from both directions. On one side, AI is increasingly central to the practical development of QI, from learning and characterizing quantum systems to optimizing variational algorithms, designing sensing and networking protocols, stabilizing noisy hardware, and automating laboratory and programming workflows. On the other side, quantum computing and quantum-inspired methods offer new algorithmic primitives, representational tools, and theoretical perspectives for machine learning and artificial intelligence.
Our main message is that these two directions are best understood together. They are linked by recurring questions about what can be learned from limited data, how high-dimensional models can be represented efficiently, which optimization signals remain usable at scale, how robustness is maintained under noise and finite resources, and when a claimed advantage is scientifically meaningful. In that sense, the interface between AI and QI is an emerging common language for learning, representation, and control under quantum constraints. We expect future progress to require rigorous benchmarking, hardware-aware modeling, and the co-design of hybrid quantum–classical systems.
MC, YG, XJ, YL, JW, ZW, and JL are supported in part by the University of Pittsburgh, School of Computing and Information, Department of Computer Science, Pitt Cyber, Pitt Momentum fund, PQI Community Collaboration Awards, John C. Mascaro Faculty Scholar in Sustainability, Switzerland NSF award 2000-1-243053, NSF award 2535915, Thinking Machines Lab and Cisco Research. This research used resources of the Oak Ridge Leadership Computing Facility, which is a DOE Office of Science User Facility supported under Contract DE-AC05-00OR22725. BZ and QZ acknowledges support from NSF (CCF-2240641, 2350153, OMA-2326746), AFOSR MURI FA9550-24-1-0349, ONR MURI N000142612102 and DOE ARPA-E Grant No. DE-AR0002067. YLiu acknowledges support from DOE Advanced Scientific Computing Research under contract number DE-SC0025384, and NSF OSI-2531350 (with a sub-contract from Duke University). XZ acknowledges the support from the AWS Center for Quantum Computing, Samsung Global Research Outreach program, and Columbia University. KPS thanks the U.S. Department of Energy, Office of Science, Advanced Scientific Computing Research (ASCR) program, for support under Award Number DE-SC0026264, and PQI Community Collaboration Awards.