Quantum Large Language Models via Tensor Network Disentanglers
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Quantum Large Language Models via Tensor Network Disentanglers".
Jane: The paper was written by Borja Aizpurua, Fernando Loren, Saeed S. Jahromi, Sukhbinder Singh and Roman Orus from Multiverse Computing and University of Navarra and Donostia International Physics Center and Ikerbasque Foundation for Science.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that just hit arXiv — it’s called “Quantum Large Language Models via Tensor Network Disentanglers.” Jane, I have to say, that title alone got me excited before I even read a word.
Jane: Same here, Tom. And the author list is a who’s who in quantum tensor networks — Borja Aizpurua, Saeed Jahromi, Sukhbinder Singh, and Román Orús. Orús has been pushing tensor network methods into AI for years, so seeing this group tackle LLMs is a big deal.
Tom: For our listeners who might not live in the quantum world — what’s the big idea? What are they actually proposing?
Jane: So, think of a large language model like a giant machine with millions of knobs — those are the weight matrices. They’re what the model learns during training. The problem is they take up tons of memory and energy. This paper says: what if we replace those weight matrices with a combination of quantum circuits and tensor networks?
Tom: And tensor networks are like a way of compressing huge amounts of information by focusing on the most important correlations, right?
Jane: Exactly. It’s like taking a massive photo and compressing it to a JPEG — you keep the parts your eye notices most. Tensor networks do that for mathematical data. The twist here is they’re adding quantum circuits on top, which can capture even more complex patterns than a classical compression can.
Tom: So they’re not just compressing — they’re upgrading. That’s the part that got me. They’re saying the quantum circuits can find correlations that the original weight matrix simply couldn’t represent.
Jane: And that’s the bold claim. They’re not just making the model smaller — they’re making it potentially smarter. The paper even says the accuracy could go beyond the classical model. That’s a strong statement, and we’ll get into how they pull that off in a minute.
Tom: Before we do — Meng, you’re the engineer here. When you hear “replace weight matrices with quantum circuits,” what’s your first reaction?
Meng: Honestly, my first thought is: how do you even train that? Quantum circuits are unitary — they’re reversible, they preserve information. A weight matrix isn’t. So you’re adding a constraint that might hurt performance. But the paper addresses that head-on, and I’m curious to see how.
Jane: Good instinct, Meng. And that’s exactly the bridge we’re about to cross. Stick with us.
Summary: Tom: Alright, we’re back with “Quantum Large Language Models via Tensor Network Disentanglers.” Jane, let’s get into the meat of the summary. What’s the actual method here?
Jane: So the core idea is pretty elegant. You take a weight matrix from a layer in an LLM — say, in the self-attention or the feed-forward part. First, you compress it into a tensor network, specifically a Matrix Product Operator, or MPO. That’s the step from earlier work — they call it CompactifAI — and it already gives you huge memory savings.
Tom: Right, and that part’s been shown to work. But then they go further.
Jane: They do. They then apply two quantum circuits — one on the input side, one on the output side — to “disentangle” the MPO. Think of it like untangling a knot: the circuits remove as much complexity as possible from the tensor network, leaving behind a simpler MPO in the middle.
Tom: And the key equation is MPO old equals U times MPO new times V-dagger, right? So the quantum circuits squeeze the complexity out of the original operator.
Jane: Exactly. And because the circuits are unitary, they don’t lose information — they just rearrange it. The remaining MPO has a smaller bond dimension, meaning it’s cheaper to store. So you get a shallow quantum circuit plus a slim tensor network, and together they reproduce the original weight matrix.
Meng: But here’s the thing — you said “reproduce.” So this is just a compression scheme? Where’s the improvement?
Jane: Great question. The improvement comes after you’ve encoded the classical model. You then extend the quantum circuits — make them deeper — and increase the bond dimension of the middle MPO. That adds new parameters, new correlations that the original model never had. So you’re not just copying the LLM; you’re upgrading it.
Tom: And the paper claims this can push accuracy beyond the classical model. That’s the punchline.
Jane: It is. And they’re careful to say the memory overhead grows polynomially with the circuit depth, so it stays manageable. They also mention you can combine this with other compression tricks like quantization and pruning.
Lu: If I can jump in — what excites me is that the disentanglers are a well-established tool from quantum many-body physics. They’ve been used for years in entanglement renormalization. So they’re not inventing a new algorithm from scratch; they’re applying a proven technique to a new domain. That makes the proposal much more credible.
Jane: Exactly, Lu. And that’s why the summary feels solid — it’s built on known tools, not magic.
Tom: So we’ve got the recipe: compress, disentangle, then expand. But how do you actually train those quantum circuits? That’s the question I want to dig into next.
Improvements: Tom: Welcome back. We’re still on “Quantum Large Language Models via Tensor Network Disentanglers.” Jane, we left off asking about training. How do you actually optimize those quantum circuits?
Jane: So the paper suggests a variational approach — similar to what’s used in quantum chemistry, like the VQE algorithm. You start with the encoded model, then you tweak the parameters in the quantum circuits and the tensors in the MPO to minimize some loss function. It’s a self-consistent loop.
Meng: But on real hardware, you can’t just feed in a vector and read out a vector. You have to encode the input as a quantum state and then sample the output. That introduces noise and approximation errors.
Jane: Right, and the paper acknowledges that. They say you can encode the input using techniques like quantum GANs or tensor network methods. And after running the circuit, you sample to reconstruct the state. The errors can be controlled by making the circuits deeper and the MPO wider.
Lu: That’s the key trade-off. Deeper circuits capture more correlations but also amplify hardware noise. The paper’s bet is that on near-term devices — hundreds of qubits — you can still get a net benefit. And that matches IBM’s roadmap for quantum-centric supercomputing.
Tom: So they’re aiming at NISQ devices, not fault-tolerant quantum computers. That’s a big deal because it means this could be tested soon, not in twenty years.
Jane: Exactly. And here’s another improvement they mention — you don’t have to train from scratch. You can take an existing model like LlaMA or BERT, encode it into this format, and then fine-tune. That’s much more practical than training a quantum model from zero.
Meng: But what about the actual speedup? Even if the model is more accurate, if it takes ten times longer to run, is it worth it?
Jane: That’s the million-dollar question. The paper doesn’t give benchmark numbers yet — they say quantitative results are coming in a future version. But the memory savings are already proven from the earlier CompactifAI work. The accuracy boost is the new claim.
Lu: And I’d add — the improvement isn’t just about speed. It’s about representational power. A quantum circuit can encode correlations that are exponentially hard to represent classically. So even if the hardware is slower, the model might solve problems that classical LLMs simply can’t.
Tom: So the improvement is two-fold: you get a smaller model, and you get a potentially more capable model. That’s the dream combination.
Jane: It is. And they’re also saying you can stack this with classical compression — quantization, pruning, low-rank factorization. So it’s not either/or. It’s a toolkit.
Meng: Okay, I’m starting to see the practical angle. But I still want to know — how many qubits are we talking about for a real model like LlaMA?
Jane: The paper says hundreds of qubits and layers should suffice for some current open-source models. That’s ambitious but not crazy. It’s within reach of current hardware roadmaps.
Tom: So we’re not talking about science fiction here. We’re talking about something that could actually run in the next few years.
Jane: That’s the hope. And that’s what makes this paper exciting — it’s a concrete bridge between today’s LLMs and tomorrow’s quantum hardware.
Conclusion: Tom: Alright, we’re wrapping up our time with “Quantum Large Language Models via Tensor Network Disentanglers.” Jane, give us the final take — what did we learn?
Jane: We learned that you can take a classical LLM, compress its weight matrices into tensor networks, then use quantum circuits to disentangle them — leaving a compact core that can be expanded and retrained to beat the original model. The method is practical, builds on proven tools, and targets near-term quantum hardware.
Tom: And the authors are clear that this is just the beginning — they’re already evaluating it on LlaMA models and generalizations. So we might see real numbers soon.
Lu: I think the biggest impact is conceptual. This shows that quantum computing isn’t just for niche physics problems — it can plug directly into the most important AI systems we have. That’s a shift in mindset.
Meng: And from an engineering standpoint, the fact that you can start from an existing model and fine-tune — not train from scratch — makes this actually feasible. That’s what gives me hope it’ll work.
Jane: And for the rest of us, it means the next generation of language models might not just be bigger — they might be fundamentally different. Quantum-enhanced, with correlations we can’t even describe classically.
Tom: Well said, Jane. That’s a good note to end on. Thanks to Lu, Meng, and all our listeners for joining us. We’ll be back next time with another paper from the arXiv. Until then, keep asking big questions.
Jane: And keep an eye on those quantum roadmaps. Bye, everyone.
Borja Aizpurua, Fernando Loren, Saeed S. Jahromi, Sukhbinder Singh, Roman Orus
Multiverse Computing · University of Navarra · Donostia International Physics Center · Ikerbasque Foundation for Science
quant-ph, cs.AI, cs.LG
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: 4 pages, 2 figures
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 49/100
The gist: OpenAI's ChatGPT, built on the transformer architecture, has led to many LLMs, but the main challenge is "their enormous energy consumption." The paper highlights that training ChatGPT-3 alone
Key concepts
- Weight Matrices
- These are the parameters that a large language model learns during training. They take up significant memory and energy.
- Tensor Networks
- These are a method used to compress huge amounts of mathematical data by focusing on the most important correlations, similar to compressing an image into a JPEG.
- Quantum Circuits
- These are unitary operations that can be applied to the tensor network. They are used in this paper to 'disentangle' the network, removing complexity and potentially capturing more patterns than classical methods.
Terminology
Summary
Summary
The paper proposes a method to enhance the performance of Large Language Models (LLMs) by integrating quantum computing and quantum-inspired techniques. Specifically, the approach involves replacing the weight matrices in the Self-Attention and Multi-layer Perceptron layers with a combination of two variational quantum circuits and a quantum-inspired tensor network, such as a Matrix Product Operator (MPO).
This substitution enables the reproduction of classical LLM functionality by decomposing weight matrices through the application of tensor network disentanglers and MPOs, leveraging well-established tensor network techniques.
By incorporating more complex and deeper quantum circuits, along with increasing the bond dimensions of the MPOs, the method captures additional correlations within the quantum-enhanced LLM, leading to improved accuracy beyond classical models while maintaining low memory overhead.
The paper begins by noting the context: OpenAI's ChatGPT, built on the transformer architecture, has led to many LLMs, but the main challenge is their enormous energy consumption.
The paper highlights that training ChatGPT-3 alone incurred an estimated 100 million in electricity costs, with expenses expected to double every ten months. This has prompted research into compression techniques, with a particularly promising approach to LLM compression is the use of quantum-inspired tensor networks, as originally proposed in Ref. [7].
In parallel, quantum computing is gaining traction, with NISQ (Noisy Intermediate-Scale Quantum) devices enabling complex variational quantum circuits
for applications such as quantum optimization and quantum machine learning.
The main idea is to replace the weight matrices in the deep layers of LLMs with (unitary) quantum circuits combined with (arbitrary) Tensor Networks.
The paper notes that previous work demonstrated that substituting weight matrices with TNs in LLMs can achieve over 90% memory compression while preserving model accuracy.
The new step is to replace the weight matrix with a variational quantum circuit, which would enable the model to capture a significantly higher degree of correlations, far beyond what the weight matrix alone can represent.
However, the unitarity of the quantum circuit would introduce constraints in optimization, potentially lowering accuracy. To address this, the paper proposes combining the quantum circuits with a non-unitary TN.
The resulting model is a generalization of the approach in Ref. [7], where weight matrices in deep layers are replaced by (i) a Variational Quantum Circuit (VQC), followed by (ii) an arbitrary TN, such as a Matrix Product Operator (MPO), and then followed again by (iii) a second Variational Quantum Circuit.
Omitting (i) and (iii) recovers the tensorized and compressed LLM models from Ref. [7]. By incorporating quantum circuits, the hybrid quantum LLM architecture captures a vast amount of quantum correlations, in addition to the correlations present in the classical (potentially tensorized) LLM.
This increase in correlations implies "at worst, enhanced accuracy compared to classical models, while the associated memory overhead scales proportionally with the depth of the quantum circuit, which in the worst case grows polynomially with the number of input bits to the layer."
The paper provides a concrete example in Fig. 1, illustrating how to implement this approach within the transformer architecture used in LLMs such as LlaMA. The Self-Attention (SA) and Multi-layer Perceptron (MP) layers contain large weight matrices, which are particularly suitable for our method.
The approach replaces these matrices with the combination of VQC+TN. The methodology can also be extended to more complex architectures like Mixtral8x7B.
The paper outlines a procedure to convert an existing classical LLM into the VQC+TN format. For a deep layer characterized by a weight matrix W, the algorithm is:
-
Apply an MPO decomposition to W and truncate the bond dimension χ to preserve accuracy, implementing the CompactifAI algorithm from Ref. [7]. This reduces memory usage by discarding irrelevant parameters.
-
Compute two quantum circuits of disentanglers, one for the input and one for the output, using two-body unitary gates to remove as much entanglement as possible from the MPO. This is a well-established procedure in tensor networks, forming the core of Entanglement Renormalization (ER) and the Multiscale Entanglement Renormalization Ansatz (MERA). The disentanglers can be computed efficiently using iterative methods from Ref. [23].
-
Compute a new MPO representing the remaining part that cannot be disentangled, described by the equation MPOold ≈ U × MPOnew × V†, where U and V† are the unitary quantum circuits of disentanglers. Since MPOold is not necessarily Hermitian, U ≠ V in general. Since U and V† remove entanglement, it follows that χnew ≤ χold. The new MPO can be computed using MPOnew ≈ U† × MPOold × V, efficiently carried out using standard tensor network approximation techniques such as the Time-Evolving Block Decimation (TEBD) algorithm for operators.
The paper notes that this procedure is advantageous because it ensures a shallow quantum circuit combined with a low-dimensional MPO.
Alternative approaches, such as using the polar decomposition W = U × P, result in a non-shallow quantum circuit for U and a P with a very large bond dimension, due to the positivity constraint.
The disentangler approach offers the most compact possible representation of W.
To enhance the original classical LLM, once expressed in this format, the process is to extend the quantum circuits for U and V† and increase the bond dimensions for MPOnew.
The model parameters can be optimized through various techniques: the unitaries in the quantum circuits can be optimized variationally using a self-consistent method similar to the Variational Quantum Eigensolver (VQE) algorithm, and the tensors in the MPO can be retrained using distributed tensor network retraining techniques as implemented in Ref. [7].
When deploying on actual quantum hardware, the input to the quantum circuits must be encoded as a quantum state, and the output is obtained through sampling via qubit measurements. Techniques for encoding include quantum GANs and tensor network methods. After executing the optimized quantum circuit for U, sampling is performed to reconstruct the input state for MPOnew, which can also be encoded as a tensor network. Similarly, the input to V† is encoded as a quantum state and the output is estimated via sampling. These steps introduce additional approximations, but they can be controlled and compensated by increasing the depth of U, V† and the bond dimensions of MPOnew.
The paper concludes that the performance of this method is currently under evaluation for LlaMA models and generalizations thereof, and quantitative results will be reported in a future version of this manuscript.
The method can also be combined with standard compression techniques such as quantization, distillation, pruning, and low-rank approximations. The paper states that our first estimations indicate that quantum circuits with hundreds of qubits and layers should suffice to improve some of the current open-source LLMs,
matching the roadmap of quantum hardware providers such as IBM. The vision is that Quantum LLMs, like the ones described in this paper, may become in the mid-term one of the first practical applications of noisy quantum computers.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system:
Implementation: Replace all weight matrices in Self-Attention and Multi-layer Perceptron layers with the VQC+TN architecture (two variational quantum circuits sandwiching a Matrix Product Operator).
What the improved system can do:
-
Capture correlations beyond what classical weight matrices can represent, since the combination of unitaries and tensor networks spans a larger hypothesis space
-
Maintain accuracy at over 90% memory compression compared to classical LLMs, as demonstrated in the prior CompactifAI work
-
Scale memory overhead polynomially with circuit depth rather than quadratically with matrix dimensions
Implementation: For any existing trained LLM (e.g., LlaMA, BERT), apply the three-step conversion:
-
MPO decomposition with bond dimension truncation
-
Compute input/output disentangler circuits U and V† using entanglement renormalization techniques
-
Compute the residual MPOnew via TEBD, satisfying MPOold ≈ U × MPOnew × V†
Implementation: After conversion, extend the depth of U and V† circuits and increase MPOnew bond dimensions, then optimize:
-
Unitaries via self-consistent VQE-style variational optimization
-
MPO tensors via distributed tensor network retraining (as in CompactifAI)
Implementation: During inference, encode input activations as quantum states (via quantum GANs or tensor network encoding), pass through U, then MPOnew (classically), then V†, and sample outputs via qubit measurements.
Implementation: Apply the VQC+TN replacement on top of, or in combination with, quantization, distillation, pruning, and low-rank factorization.
Net capability of the improved AI system: It is a hybrid LLM that (a) matches or exceeds classical accuracy, (b) reduces memory by >90%, (c) runs on near-term quantum hardware, and (d) can be derived from any existing open-source LLM without full retraining.
Sources
- LLaMA: Open and Efficient Foundation Language Models
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- CompactifAI: Extreme Compression of Large Language Models using Quantum-Inspired Tensor Networks
- Improving Gradient Methods via Coordinate Transformations: Applications to Quantum Machine Learning
- A Practical Introduction to Tensor Networks: Matrix Product States and Projected Entangled Pair States
- Regular language quantum states
- Tensorizing Neural Networks
- Quantum-Inspired Tensor Neural Networks for Partial Differential Equations
- Variational Tensor Neural Networks for Deep Learning
- Tensor Networks Meet Neural Networks: A Survey and Future Perspectives
- Mixtral of Experts
- Distilling the Knowledge in a Neural Network
- Speeding up Convolutional Neural Networks with Low Rank Expansions
Related papers
- Reconquering Bell sampling on qudits: stabilizer learning and testing, quantum pseudorandomness bounds, and more
- Encrypted clones can leak: Classification of informative subsets in Quantum Encrypted Cloning
- Polynomial-time classical and quantum simulation of quantum impurity models
- Theory of quantum-enhanced interferometry with general Markovian light sources
- A convergent hierarchy of spectral gap certificates for qubit Hamiltonians
- Universal Bound and Phase Transition in Many-Body Fermionic Non-Gaussianity