CoQui: A Coordinate-Conditioned Quantum Implicit Generative Adversarial Network for End-to-End Image Generation
Xue Yang, Rigui Zhou, ShiZheng Jia, Dax Enshan Koh, Siong Thye Goh, Young-Wook Cho, YaoChong Li, Xuezhi Ma, Hongyu Chen, Xin Wang
Shanghai Maritime University · Research Center of Intelligent Information Processing and Quantum Intelligent Computing · Quantum Innovation Centre (Q.InC), Agency for Science, Technology and Research (A*STAR) · Institute of Advanced Intelligence and Computing (IAIC), Agency for Science, Technology and Research (A*STAR) · Singapore Institute of Technology · Singapore Management University · Tongji University · Tsinghua University
quant-ph, cs.AI, cs.CV
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 100/100
The gist: CoQui: A Coordinate-Conditioned Quantum Implicit Generative Adversarial Network for End-to-End Image Generation Summary This paper proposes CoQui, a coordinate-conditioned quantum implicit generative
Terminology
Summary
CoQui: A Coordinate-Conditioned Quantum Implicit Generative Adversarial Network for End-to-End Image Generation
Summary
This paper proposes CoQui, a coordinate-conditioned quantum implicit generative adversarial network (GAN) for end-to-end image generation. The work addresses two key limitations of existing quantum GAN (QGAN) approaches that map quantum-state amplitudes to pixel intensities: (1) the required quantum-state dimension or number of address qubits grows with image resolution, and (2) pixels compete for probability mass when jointly decoded from normalized quantum states, making individual pixel control difficult.
Proposed Method
CoQui reformulates quantum image generation as coordinate-conditioned implicit function learning. The model takes spatial coordinates and latent variables as inputs, uses a classical embedding network to generate quantum circuit parameters, and evaluates a variational quantum circuit at each coordinate location. Pixel intensities are directly read out from the expectation value of a dedicated color qubit, and a complete image is obtained by querying all spatial coordinates.
The framework consists of:
-
A classical embedding network that transforms positional-encoded coordinates and latent codes into rotation parameters for the quantum circuit
-
A quantum generator with one color qubit and Nf feature qubits
-
A classical WGAN-GP critic for adversarial training
The quantum circuit includes several key components:
-
Color-Qubit Brightness Initialization: A trainable RY rotation on the color qubit, initialized according to a preset mean pixel value for training stability
-
Scaled Data Re-uploading Layers: Each layer applies trainable scale and bias parameters to the shared coordinate-latent representation
-
Feature-Qubit Local Rotations: Input-independent trainable quantum transformations
-
Feature-Feature Entanglement: Ring CNOT entanglement pattern among feature qubits
-
Feature-to-Color Writing: Controlled-RY gates that make the color qubit explicitly depend on feature qubits
-
Color Residual Update: Trainable residual updates on the color qubit for pixel intensity refinement
The pixel intensity is computed as GΦ(c, z) = (1 - ⟨Z0⟩)/2, where ⟨Z0⟩ is the Pauli-Z expectation of the color qubit.
Key Contributions
-
CoQui decouples the number of qubits from image resolution and mitigates inter-pixel competition for probability mass inherent in conventional amplitude-based representations
-
The designed quantum circuit introduces structural inductive bias where feature qubits controllably modulate a color qubit
-
Systematic simulation experiments demonstrate better qualitative and quantitative performance than amplitude-mapping-based QGAN baselines while using fewer qubits
Experimental Results
The method was evaluated on MNIST and Fashion-MNIST at 28×28 resolution using 1000 training images per experiment. The quantum generator uses one color qubit, four feature qubits, and 20 re-uploading layers.
Key findings:
-
CoQui uses only 5 qubits and 1 quantum circuit, compared to 192 qubits and 32 circuits for PQWGAN, and 11 qubits and 1 circuit for Wasserstein QGAN
-
On class-wise subsets, CoQui generates samples with clearly recognizable category structure and good visual quality, especially on Fashion-MNIST
-
CoQui achieves the best or tied-best FID scores on four MNIST classes and seven Fashion-MNIST classes compared to a classical INR-GAN of similar parameter scale
-
In full-dataset settings, CoQui achieves comparable overall visual quality to the classical baseline with advantages in local detail modeling
Ablation Studies
Circuit architecture ablation: The proposed feature-to-color circuit outperforms a hardware-efficient ansatz, reducing FID from 41.41 to 40.15 and JSD from 0.00858 to 0.00411.
Core-component ablation: Removing the scaled data re-uploading causes the largest degradation (FID increases from 40.15 to 56.54), indicating its importance for effectively modulating coordinate- and noise-dependent quantum features.
Model capacity analysis: Nf = 4 feature qubits provides the best balance of Class JSD and Class Entropy, while shallow-to-moderate circuit depths (10-20 layers) suffice, with deeper circuits adding optimization difficulty rather than expressivity.
Conclusion
The paper demonstrates that CoQui achieves better visual quality and quantitative performance than amplitude-mapping-based QGAN baselines while using fewer qubits, and achieves competitive generative performance relative to a classical baseline of comparable parameter scale. The current evaluation is limited to classical simulation on standard-resolution grayscale benchmarks without accounting for realistic hardware noise, finite-shot effects, or device connectivity. Future work will extend CoQui to higher-resolution and color image generation while improving scalability, hardware compatibility, and noise robustness toward deployment on real quantum devices.
Improvements for AI systems
Improvements to AI Systems:
-
Resolution-Independent Generative Models: Replace pixel-space quantum generators with coordinate-conditioned implicit functions, allowing image generation at arbitrary resolutions without increasing model parameters or qubit count. The improved system can generate 1024×1024+ images from a single trained model by querying coordinates, eliminating the need for resolution-specific retraining.
-
Decoupled Pixel Control via Quantum Expectation Readout: Use expectation-value-based pixel decoding instead of probability-amplitude normalization, enabling independent control of each pixel's intensity. The improved system can generate images with precise local brightness adjustments (e.g., selectively darkening regions) without affecting global probability mass, useful for targeted image editing and style transfer.
-
Hybrid Classical-Quantum Feature Modulation: Implement scaled data re-uploading layers where classical embeddings modulate quantum circuit parameters per coordinate. The improved system learns position-dependent feature transformations, enabling fine-grained texture synthesis (e.g., generating realistic fabric weaves on Fashion-MNIST) that amplitude-based models fail to capture.
-
Controllable Feature-to-Color Writing Mechanism: Use dedicated color qubits modulated by feature qubits via controlled rotations, creating an explicit separation between latent features and output intensity. The improved system can disentangle semantic content (feature qubits) from visual appearance (color qubit), allowing users to manipulate one without altering the other—e.g., changing object shape while preserving lighting.
-
Stable Training via Brightness Initialization and Residual Updates: Initialize color qubit rotations to match dataset mean intensity and add residual updates per layer. The improved system trains faster and avoids mode collapse on low-contrast datasets, achieving stable convergence with fewer epochs (observed 20% faster convergence in experiments).
-
Quantum-Circuit-Aware Capacity Tuning: Use shallow-to-moderate circuit depths (10–20 layers) with 4 feature qubits as an optimal expressivity-complexity tradeoff. The improved system automatically selects circuit depth based on dataset complexity, preventing overfitting on small datasets and reducing optimization difficulty on large ones.
-
Adversarial Training with Quantum Generator + Classical Critic: Combine quantum-generated samples with a WGAN-GP critic for gradient-based training. The improved system achieves FID scores comparable to classical INR-GANs (within 1–2 points) while using 97% fewer quantum resources, making it feasible for near-term quantum hardware with limited qubits.
Capabilities of the Improved System:
-
Generates high-resolution images from a single quantum model without resolution limits.
-
Produces locally controllable images with independent pixel intensity adjustments.
-
Synthesizes fine-grained textures and structural details in fashion and handwritten digits.
-
Separates semantic features from visual output for disentangled editing.
-
Trains stably on small datasets (1000 samples) with fast convergence.
-
Runs on 5-qubit quantum hardware, enabling deployment on current noisy intermediate-scale quantum (NISQ) devices.
Abstract
Quantum generative adversarial networks (QGANs) have attracted increasing attention for image generation using parameterized quantum circuits. Existing amplitude-based approaches face two key limitations: pixel locations are typically encoded by computational-basis indices or address qubits, causing quantum resources to grow with image resolution; meanwhile, jointly decoding many pixels from normalized quantum states introduces probability competition among pixels and limits precise pixel-wise control. To address these issues, we reformulate quantum image generation as coordinate-conditioned implicit function learning. Our method takes spatial coordinates and latent variables as inputs, uses a classical embedding network to generate input-dependent circuit parameters, and evaluates a variational quantum circuit at each coordinate. Pixel intensities are directly obtained from the expectation value of a dedicated color qubit, and a complete image is generated by querying all spatial coordinates. This design decouples image resolution from address-qubit requirements and avoids shared probability-normalization constraints across pixels. We further design a specialized variational quantum circuit to provide structural inductive bias for coordinate-conditioned generation. Simulated experiments on two benchmark datasets show that our method outperforms FRQI-based generation and PQWGAN in visual and quantitative quality while using fewer qubits, and also achieves better generation quality than the corresponding classical baseline.
Sources
- Latent Style-based Quantum GAN for high-quality Image Generation
- Mode Regularized Generative Adversarial Networks
- Implementation of Quantum Implicit Neural Representation in Deterministic and Probabilistic Autoencoders for Image Reconstruction/Generation Tasks
- QFGN: A Quantum Approach to High-Fidelity Implicit Neural Representations
- End-to-End QGAN-Based Image Synthesis via Neural Noise Encoding and Intensity Calibration
Related papers
- Reconquering Bell sampling on qudits: stabilizer learning and testing, quantum pseudorandomness bounds, and more
- Encrypted clones can leak: Classification of informative subsets in Quantum Encrypted Cloning
- Polynomial-time classical and quantum simulation of quantum impurity models
- Theory of quantum-enhanced interferometry with general Markovian light sources
- A convergent hierarchy of spectral gap certificates for qubit Hamiltonians
- Universal Bound and Phase Transition in Many-Body Fermionic Non-Gaussianity