SoK: Critical Evaluation of Quantum Machine Learning for Adversarial Robustness

1University of Texas at El Paso 2University of Dhaka 3University of Massachusetts Amherst
📌 Accepted at IEEE S&P 2026

Abstract

Quantum Machine Learning (QML) integrates quantum computational principles into learning algorithms, offering the potential for improved representational capacity and computational efficiency. Nevertheless, the security and robustness of QML systems remain largely underexplored, particularly under adversarial conditions. We present the first comprehensive systematization of adversarial robustness in QML, integrating conceptual organization with empirical evaluation across black-, gray-, and white-box threat models. We implement five representative attacks across all three threat models: a label-flipping data poisoning attack under black-box; an encoder-level indiscriminate data poisoning attack and a proxy-model-based clean-label backdoor attack under gray-box; and a circuit-level backdoor attack (QTrojan) and gradient-based evasion attacks (FGSM and PGD) under white-box. We evaluate the attacks using a Quantum Multilayer Perceptron (QMLP) trained on MNIST and AZ-Class across circuit depths of 2, 5, 10 and 50 layers and two encoding schemes (angle and amplitude).

Our extensive evaluations reveal a fundamental accuracy–robustness trade-off. In particular, amplitude encoding yields the highest clean accuracy (92.6% on MNIST, 67% on AZ-Class); however, it collapses under adversarial perturbations and depolarizing noise, while shallow angle-encoded models remain substantially more stable. In addition, QUID is highly effective under noiseless conditions but is weakened by noise, whereas the proxy-model backdoor persists unless the circuit itself is overwhelmed, highlighting that noise is an asymmetric and unreliable passive defense. Furthermore, the circuit-level backdoor fails in the multi-class setting, indicating a scalability constraint. Finally, QMLP models are more robust than Classical Multi-Layer Perceptron (CMLP) models under label-flipping attacks but are substantially more vulnerable to gradient-based evasion, motivating the need for quantum-specific defenses. We conclude by proposing a threat-aware, noise-resilient design framework for secure and robust QML deployment. The dataset and code are available at https://github.com/IQSeC-Lab/SoK-QML.

QML Architecture

Architecture of Parameterized Quantum Circuit (PQC) as a QML model
Figure 1: Architecture of Parameterized Quantum Circuit (PQC) as a QML model. The circuit comprises three key components: (1) data encoding layers that map classical inputs into quantum states, (2) parameterized quantum gates that define the model's behavior and are iteratively optimized, and (3) measurement operations that extract classical information for loss evaluation.

Most practical QML systems use a hybrid quantum-classical approach. The quantum processor executes Parameterized Quantum Circuits (PQCs) that prepare quantum states, while a classical optimizer updates trainable parameters to minimize a loss function. A typical QML pipeline consists of three stages: data embedding, parameterized quantum gates, and measurement followed by classical postprocessing.

We evaluate two data embedding strategies:

  • Angle Encoding: Each classical feature is mapped to a rotation angle of a single-qubit gate. This method is simple and hardware-efficient but requires at least one qubit per input feature.
  • Amplitude Encoding: The entire input vector is encoded into the amplitudes of a quantum state. This method is qubit-efficient but harder to implement on NISQ hardware due to complex state preparation and noise sensitivity.

Taxonomy of Attack Models

Schematic of adversarial attack surfaces in QML pipeline
Figure 2: Schematic of adversarial attack surfaces in QML pipeline, illustrating potential threat vectors and their points of insertion. Each attack is annotated as (T): Training-time, (I): Inference-time, (D): Dual-phase, (B): Black-Box, (G): Gray-Box, (W): White-Box.

We present a taxonomy of adversarial threats in QML based on the adversary's level of access—black-box, gray-box, and white-box—and the quantum-specific mechanisms exploited during the attack.

  • Black-Box Attacks: The adversary lacks visibility into the QML model's architecture, parameters, and hardware. Attacks include model extraction, crosstalk-induced side-channel attacks, membership inference, data poisoning (label-flipping), and adversarial evasion.
  • Gray-Box Attacks: Adversaries have partial visibility into the QML system, including access to intermediate representations such as encoded quantum states and transpiled circuits. Attacks include backdoor attacks, Quantum Indiscriminate Data Poisoning (QUID), and pulse-level attacks.
  • White-Box Attacks: Adversaries have full visibility and control over the QML system. Attacks include circuit-level backdoors (QTrojan), reverse engineering, hardware/compiler trojans, input inference, and gradient-based evasion attacks (FGSM and PGD).

Experimental Setup

Model Architecture: We implement QMLP models following a hybrid quantum-classical design. The classical component is implemented in PyTorch and the quantum circuit is simulated using PennyLane's default.qubit backend. To emulate realistic hardware behavior, depolarizing noise with probability p = 0.01 is applied via the Qiskit Aer backend. The QMLP architecture varies in circuit depth (2, 5, 10, and 50 layers) and encoding strategy (angle and amplitude).

Datasets: We use two multiclass datasets from distinct domains:

  • MNIST: The most widely used dataset in prior QML works, used for 10-class classification of handwritten digits.
  • AZ-Class: A twenty-three-class Android malware dataset based on behavioral features, enabling cross-domain evaluation of QML robustness.

Our QMLP employs a 9-qubit circuit. For angle encoding, inputs are reduced to 9 principal components (one per qubit) via PCA, whereas for amplitude encoding, inputs are compressed to 512 dimensions to match the circuit's Hilbert space.

Attacks Evaluated: We implement five representative attacks spanning all three threat models:

  • Black-Box: Label-flipping data poisoning with label smoothing defense
  • Gray-Box: QUID (Quantum Indiscriminate Data Poisoning) with Q-Detection defense; Huang-Zhang backdoor attack
  • White-Box: QTrojan circuit-level backdoor; FGSM and PGD gradient-based evasion attacks

Baseline Results

Baseline Accuracies (%) of QMLP across encoding schemes and circuit depths under noiseless and depolarized noise (p = 0.01) conditions
Encoding Layers AZ-Class (NL) AZ-Class (DN) MNIST (NL) MNIST (DN)
Angle 2 49.8 47.0 66.0 59.2
5 52.3 22.4 68.0 45.0
10 54.8 4.9 83.6 23.6
50 32.4 5.2 76.9 9.9
Amplitude 2 40.1 5.1 50.7 9.9
5 48.1 5.0 70.3 9.5
10 55.0 4.78 79.8 10.1
50 67.0 4.81 92.6 9.7
CMLP – 95.89 – 96.63 –

NL = Noiseless, DN = Depolarized Noise. Angle encoding favors shallow circuits (highest accuracy at 10-layer PQC), while amplitude encoding benefits from deeper circuits, achieving 92.6% on MNIST at 50 layers. Noise severely limits accuracy, with shallow angle-encoded circuits showing relatively higher resilience.

AZ Angle MNIST Angle AZ Amplitude MNIST Amplitude
Figure 3: Performance of the QMLP model with varying circuit layers and encoding schemes on the AZ-Class and MNIST datasets under noiseless condition.

Attack Evaluation Results

Black-Box: Label-Flipping Attack

QMLP models exhibit greater robustness than CMLP models under label-flipping attacks, retaining relative accuracies close to 90% under a 50% label-flipping attack, while CMLP loses nearly 50% of its clean accuracy. Label smoothing provides no meaningful benefit for QMLP and only modest gains for CMLP.

AZ Angle LF MNIST Angle LF AZ Amplitude LF MNIST Amplitude LF
Figure 4: Performance of QMLP under label-flipping (LF) and label-flipping with label-smoothing (LFLS) across datasets (AZ-Class, MNIST) and encodings (Angle, Amplitude) in noiseless conditions.

Gray-Box: QUID Poisoning Attack

Amplitude-encoded models are more vulnerable to QUID, with ASR exceeding 90% at poison ratio 0.5. Angle encoding is more resilient at low poison ratios but degrades at higher ratios. Depolarizing noise weakens QUID on amplitude encoding by disrupting Hilbert-space structure, acting as a natural defense. Q-Detection is more effective with angle encoding at moderate poison ratios.

Noiseless Angle QUID Noiseless Amplitude QUID Noisy Angle QUID Noisy Amplitude QUID
Figure 5: Attack Success Rate (ASR) of the QUID attack without Q-Detection across Amplitude and Angle encodings under Noiseless and Noisy quantum circuit settings with varying layers.

Gray-Box: Huang-Zhang Backdoor Attack

The attack remains highly stealthy across encodings, depths, and poison ratios. Under noiseless conditions, amplitude encoding is markedly more vulnerable because its single-shot embedding gives the optimized trigger a stable entry point. Unlike QUID, the Huang trigger is embedded in the trained weights and persists as long as the model remains functional.

Noiseless Angle Huang Noiseless Amplitude Huang Noisy Angle Huang Noisy Amplitude Huang
Figure 6: Attack Success Rate (ASR) of the Huang backdoor attack with amplitude and angle encodings under noiseless and noisy settings with varying layers.

White-Box: QTrojan Attack

QTrojan preserves clean data accuracy (CDA) across all depths and datasets, making it a stealthy attack. However, ASR is consistently low across all configurations, contrasting sharply with the 100% ASR reported in the original paper, due to the increased number of classes in our evaluation. This reveals that QTrojan's effectiveness degrades significantly with the number of output classes.

ASR of QTrojan across circuit depths
Figure 7: ASR of QTrojan across circuit depths under noiseless and noisy conditions using angle encoding.

White-Box: FGSM and PGD Evasion Attacks

Shallow angle-encoded QMLPs retain moderate robustness, while deeper circuits collapse due to over-entanglement. Amplitude-encoded models collapse uniformly across all depths under small perturbations. At higher perturbation strengths (ε≥0.10), shallow angle-encoded QMLPs outperform CMLP on AZ-Class under both FGSM and PGD, revealing a fundamental accuracy–robustness trade-off.

AZ Angle PGD AZ Amplitude PGD MNIST Angle PGD MNIST Amplitude PGD
Figure 8: Performance of QMLP with varying perturbation strengths of the PGD attack in a noiseless environment.

Secure and Robust QML Pipeline

Proposed Secure and Robust QML Pipeline
Figure 9: Proposed Secure and Robust QML Pipeline.

Based on our empirical findings, we present a modular security framework for QML systems. Our results highlight three key lessons: classical defenses do not transfer directly to QML, noise is an unreliable and asymmetric passive defense, and encoding choice is the most consequential architectural security decision. The pipeline covers:

  • Defining the Adversarial Threat Model: Identifying which components can be observed or modified, aligned with our taxonomy of black-box, gray-box, and white-box threats.
  • Encoder-Level Security: The encoding stage is the most security-critical interface. Encoding should include quantum-aware validation, randomized or obfuscated schemes, and classical-quantum consistency checks.
  • Securing Quantum Circuit Architecture: Quantum logic locking (QLL) and E-LoQ enforce key-dependent circuit behavior, while obfuscation through gate reordering and dummy-gate insertion makes reverse engineering harder.
  • Hardware-Aware Security Mechanisms: Active hardware-level defenses including randomized qubit mapping, instruction reordering, dynamical decoupling, and hardware noise fingerprinting for tamper detection.
  • Secure Partitioning and Model Distribution: Partitioning strategies that divide functionality across multiple quantum backends via isolated sub-circuits communicating through secure classical channels.

Key Contributions

  • We present a comprehensive systematization of adversarial attack models in QML, categorized by the adversary's level of access (black-box, gray-box, white-box).
  • We empirically analyze how data encoding schemes (angle vs. amplitude) and circuit depth (2, 5, 10, and 50 layers) influence QML performance and robustness under NISQ constraints.
  • We perform a comparative study with classical machine learning (CMLP) to highlight distinctive vulnerability patterns and motivate the need for quantum-specific defenses.
  • We implement representative attacks and defenses across all three threat models and find that QMLP models are more robust against classical label-flipping attacks but substantially more vulnerable to gradient-based perturbation attacks.
  • We propose a structured design pipeline for developing secure and robust QML models integrating threat modeling, encoding strategy selection, and robustness evaluation under quantum noise.

Key Findings

  • Encoding choice is a key security decision: Amplitude encoding achieves the highest clean accuracy but collapses under adversarial perturbations and noise. Shallow angle-encoded models offer a better accuracy–robustness trade-off.
  • Noise is an unreliable passive defense: It weakens QUID on amplitude encoding by disrupting encoded-state structure, but does not stop the Huang backdoor, whose trigger is embedded in model weights.
  • Classical defenses do not transfer directly: Label smoothing provides no meaningful protection for QMLP, underscoring that reducing the attack surface at deployment is necessary when algorithmic defenses alone are insufficient.
  • Circuit-level backdoors face scalability limits: QTrojan's effectiveness degrades significantly with the number of output classes, failing in multi-class settings.
  • QMLP vs. CMLP trade-offs: QMLP models are more robust than CMLP under label-flipping attacks but substantially more vulnerable to gradient-based evasion (FGSM, PGD).

BibTeX

@inproceedings{nowmi2026sok,
  title     = {SoK: Critical Evaluation of Quantum Machine Learning for Adversarial Robustness},
  author    = {Nowmi, Saeefa Rubaiyet and Lopez, Jesus and Imon, Md Mahmudul Alam
               and Pouryousef, Shahrooz and Rahman, Mohammad Saidur},
  booktitle = {2026 IEEE Symposium on Security and Privacy (S\&P)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2511.14989}
}