QiT: Quantum-Inspired Transformer for Visual Recognition Task
quant-ph, cs.AI, cs.CV, physics.app-ph
Submitted: 2026-09-15
Updated: 2026-09-15
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Quantum machine learning offers a compelling representational perspective: angle-encoded states inhabit Hilbert spaces in which periodic similarities and interactions can be expressed naturally.
Terminology
Abstract
Quantum machine learning offers a compelling representational perspective: angle-encoded states inhabit Hilbert spaces in which periodic similarities and interactions can be expressed naturally. Realizing this perspective for visual recognition remains difficult, however, because present quantum neural networks are constrained by limited qubit counts, costly circuit simulation and measurement, noise, and unstable optimization on noisy intermediate-scale quantum devices. We investigate whether useful structural ideas from quantum models can instead be realized as scalable classical Transformer operations. We introduce QiT, a Quantum-inspired Transformer for vision tasks with three components: (i) angle-inspired encoding that maps image tokens to learned trigonometric Hilbert-space features analogous to quantum rotation-based state encoding; (ii) self-attention over these periodic features, inducing a classical cosine kernel approximated to quantum fidelity kernels; and (iii) gated multiplicative emulation, a trainable classical surrogate for interaction terms found in variational circuits. All components are differentiable tensor operations, so QiT claims neither quantum computation nor quantum speedup and retains the O(N 2D) attention complexity of a standard Vision Transformer. Across image-classification benchmarks, QiT is competitive with a matched classical Transformer while avoiding the severe runtime cost observed for a small simulated quantum Transformer. QiT-B reaches 78.3% ImageNet-1K top-1 accuracy with 45.7M parameters and 11.5 GFLOPs. These results position QiT as a scalable baseline for isolating and evaluating quantum-motivated inductive biases in visual recognition.
Sources
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- LocalMamba: Visual State Space Model with Windowed Selective Scan
- FNet: Mixing Tokens with Fourier Transforms
- VMamba: Visual State Space Model
- SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time series
- Solar Extreme UV radiation and quark nugget dark matter model
- PlainMamba: Improving Non-Hierarchical Mamba in Visual Recognition
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
Related papers
- Reconquering Bell sampling on qudits: stabilizer learning and testing, quantum pseudorandomness bounds, and more
- Encrypted clones can leak: Classification of informative subsets in Quantum Encrypted Cloning
- Polynomial-time classical and quantum simulation of quantum impurity models
- Theory of quantum-enhanced interferometry with general Markovian light sources
- A convergent hierarchy of spectral gap certificates for qubit Hamiltonians
- Universal Bound and Phase Transition in Many-Body Fermionic Non-Gaussianity