Quantum Machine Learning (QML) integrates quantum computational principles into learning algorithms, offering the potential for improved representational capacity and computational efficiency. Nevertheless, the security and robustness of QML systems remain largely underexplored, particularly under adversarial conditions. We present the first comprehensive systematization of adversarial robustness in QML, integrating conceptual organization with empirical evaluation across black-, gray-, and white-box threat models. We implement five representative attacks across all three threat models: a label-flipping data poisoning attack under black-box; an encoder-level indiscriminate data poisoning attack and a proxy-model-based clean-label backdoor attack under gray-box; and a circuit-level backdoor attack (QTrojan) and gradient-based evasion attacks (FGSM and PGD) under white-box. We evaluate the attacks using a Quantum Multilayer Perceptron (QMLP) trained on MNIST and AZ-Class across circuit depths of 2, 5, 10 and 50 layers and two encoding schemes (angle and amplitude).
Our extensive evaluations reveal a fundamental accuracy–robustness trade-off. In particular, amplitude encoding yields the highest clean accuracy (92.6% on MNIST, 67% on AZ-Class); however, it collapses under adversarial perturbations and depolarizing noise, while shallow angle-encoded models remain substantially more stable. In addition, QUID is highly effective under noiseless conditions but is weakened by noise, whereas the proxy-model backdoor persists unless the circuit itself is overwhelmed, highlighting that noise is an asymmetric and unreliable passive defense. Furthermore, the circuit-level backdoor fails in the multi-class setting, indicating a scalability constraint. Finally, QMLP models are more robust than Classical Multi-Layer Perceptron (CMLP) models under label-flipping attacks but are substantially more vulnerable to gradient-based evasion, motivating the need for quantum-specific defenses. We conclude by proposing a threat-aware, noise-resilient design framework for secure and robust QML deployment. The dataset and code are available at https://github.com/IQSeC-Lab/SoK-QML.
Most practical QML systems use a hybrid quantum-classical approach. The quantum processor executes Parameterized Quantum Circuits (PQCs) that prepare quantum states, while a classical optimizer updates trainable parameters to minimize a loss function. A typical QML pipeline consists of three stages: data embedding, parameterized quantum gates, and measurement followed by classical postprocessing.
We evaluate two data embedding strategies:
We present a taxonomy of adversarial threats in QML based on the adversary's level of access—black-box, gray-box, and white-box—and the quantum-specific mechanisms exploited during the attack.
Model Architecture: We implement QMLP models following a hybrid quantum-classical design. The classical component is implemented in PyTorch and the quantum circuit is simulated using PennyLane's default.qubit backend. To emulate realistic hardware behavior, depolarizing noise with probability p = 0.01 is applied via the Qiskit Aer backend. The QMLP architecture varies in circuit depth (2, 5, 10, and 50 layers) and encoding strategy (angle and amplitude).
Datasets: We use two multiclass datasets from distinct domains:
Our QMLP employs a 9-qubit circuit. For angle encoding, inputs are reduced to 9 principal components (one per qubit) via PCA, whereas for amplitude encoding, inputs are compressed to 512 dimensions to match the circuit's Hilbert space.
Attacks Evaluated: We implement five representative attacks spanning all three threat models:
| Encoding | Layers | AZ-Class (NL) | AZ-Class (DN) | MNIST (NL) | MNIST (DN) |
|---|---|---|---|---|---|
| Angle | 2 | 49.8 | 47.0 | 66.0 | 59.2 |
| 5 | 52.3 | 22.4 | 68.0 | 45.0 | |
| 10 | 54.8 | 4.9 | 83.6 | 23.6 | |
| 50 | 32.4 | 5.2 | 76.9 | 9.9 | |
| Amplitude | 2 | 40.1 | 5.1 | 50.7 | 9.9 |
| 5 | 48.1 | 5.0 | 70.3 | 9.5 | |
| 10 | 55.0 | 4.78 | 79.8 | 10.1 | |
| 50 | 67.0 | 4.81 | 92.6 | 9.7 | |
| CMLP | – | 95.89 | – | 96.63 | – |
NL = Noiseless, DN = Depolarized Noise. Angle encoding favors shallow circuits (highest accuracy at 10-layer PQC), while amplitude encoding benefits from deeper circuits, achieving 92.6% on MNIST at 50 layers. Noise severely limits accuracy, with shallow angle-encoded circuits showing relatively higher resilience.
QMLP models exhibit greater robustness than CMLP models under label-flipping attacks, retaining relative accuracies close to 90% under a 50% label-flipping attack, while CMLP loses nearly 50% of its clean accuracy. Label smoothing provides no meaningful benefit for QMLP and only modest gains for CMLP.
Amplitude-encoded models are more vulnerable to QUID, with ASR exceeding 90% at poison ratio 0.5. Angle encoding is more resilient at low poison ratios but degrades at higher ratios. Depolarizing noise weakens QUID on amplitude encoding by disrupting Hilbert-space structure, acting as a natural defense. Q-Detection is more effective with angle encoding at moderate poison ratios.
The attack remains highly stealthy across encodings, depths, and poison ratios. Under noiseless conditions, amplitude encoding is markedly more vulnerable because its single-shot embedding gives the optimized trigger a stable entry point. Unlike QUID, the Huang trigger is embedded in the trained weights and persists as long as the model remains functional.
QTrojan preserves clean data accuracy (CDA) across all depths and datasets, making it a stealthy attack. However, ASR is consistently low across all configurations, contrasting sharply with the 100% ASR reported in the original paper, due to the increased number of classes in our evaluation. This reveals that QTrojan's effectiveness degrades significantly with the number of output classes.
Shallow angle-encoded QMLPs retain moderate robustness, while deeper circuits collapse due to over-entanglement. Amplitude-encoded models collapse uniformly across all depths under small perturbations. At higher perturbation strengths (ε≥0.10), shallow angle-encoded QMLPs outperform CMLP on AZ-Class under both FGSM and PGD, revealing a fundamental accuracy–robustness trade-off.
Based on our empirical findings, we present a modular security framework for QML systems. Our results highlight three key lessons: classical defenses do not transfer directly to QML, noise is an unreliable and asymmetric passive defense, and encoding choice is the most consequential architectural security decision. The pipeline covers:
@inproceedings{nowmi2026sok,
title = {SoK: Critical Evaluation of Quantum Machine Learning for Adversarial Robustness},
author = {Nowmi, Saeefa Rubaiyet and Lopez, Jesus and Imon, Md Mahmudul Alam
and Pouryousef, Shahrooz and Rahman, Mohammad Saidur},
booktitle = {2026 IEEE Symposium on Security and Privacy (S\&P)},
year = {2026},
url = {https://arxiv.org/abs/2511.14989}
}