The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
summary
The gist
The paper "The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures" by Alexander M.
In short
The episode analyzes 'The End of Transformers?', a paper surveying alternative AI architectures to address the Transformer's quadratic complexity bottleneck. Hosts discuss how while pure Transformers remain best for massive frontier models, sub-quadratic methods are rapidly improving efficiency for smaller, edge devices.
Key concepts
- Transformer
- The core architecture behind major language models like GPT and Llama. Its 'attention' mechanism is powerful for understanding context but suffers from quadratic complexity, meaning computation cost increases exponentially with input length.
- Quadratic Complexity
- A computational bottleneck where the required processing power grows according to the square of the input size. For example, doubling the text length quadruples the necessary computation.
- Sub-Quadratic Architectures
- Alternative models (like State Space Models and linear RNNs) designed to process text more efficiently than Transformers. They achieve speed by avoiding full attention calculations, making them ideal for memory-constrained devices.
- Hybrid Models
- Architectures that combine the best parts of different systems, such as mixing fast State Space Models with targeted attention layers. This aims to gain efficiency while patching weaknesses like poor recall ability.
Terminology used across episodes
This episode discusses
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures · Paper Radio
- Titans: Learning to Memorize at Test Time
- Longformer: The Long-Document Transformer
- RecurrentGemma: Moving Past Transformers for Efficient Open Language Models
- Evaluating Large Language Models Trained on Code
- Generating Long Sequences with Sparse Transformers
- Rethinking Attention with Performers
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
- The Llama 3 Herd of Models · Paper Radio
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Log-Linear Attention
- Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
- Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers
- Mixtral of Experts
- Scaling Laws for Neural Language Models
- MiniMax-01: Scaling Foundation Models with Lightning Attention
The paper
The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures · Read on arXiv
Technical University of Munich
Transformers have dominated sequence processing tasks for the past seven years -- most notably language modeling. However, the inherent quadratic complexity of their attention mechanism remains a significant bottleneck as context length increases. This paper surveys recent efforts to overcome this bottleneck, including advances in (sub-quadratic) attention variants, recurrent neural networks, state space models, and hybrid architectures. We critically analyze these approaches in terms of compute and memory complexity, benchmark results, and fundamental limitations to assess whether the dominance of pure-attention transformers may soon be challenged.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures".
Jane: The paper was written by Alexander M. Fichtl, Jeremias Bohn, Josefin Kelber, Edoardo Mosca and Georg Groh from Technical University of Munich.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making waves with a pretty bold title: "The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures." Jane, I have to say, just reading that title gave me a little shiver.
Jane: Same here, Tom. It's a question a lot of people in the field have been whispering about for a while. The paper, from the team at the Technical University of Munich, is essentially asking whether the architecture that powers basically every major language model is about to get dethroned.
Tom: Right, and for our listeners who might not be deep in the weeds, can you break down what a Transformer actually is and why this question matters so much?
Jane: Absolutely. So, the Transformer is the backbone of models like GPT and Llama. Its superpower is something called "attention," which lets the model look at every single word in a sentence and figure out how it relates to every other word. That's incredibly powerful for understanding context.
Tom: But there's a catch, right? That power comes at a cost.
Jane: A huge cost. The paper calls it quadratic complexity. If you double the length of your input text, the amount of computation needed doesn't just double—it quadruples. So, processing a one hundred thousand-word document is exponentially more expensive than a ten thousand-word one.
Tom: And that's the bottleneck this paper is all about. The authors from TUM are surveying all the attempts to break through that wall. They're looking at things like state space models, linear RNNs, and hybrids that mix different approaches.
Jane: Exactly. The title is a bit provocative, but the paper itself is a really careful survey. They're not just declaring the Transformer dead. They're laying out the contenders and asking if any of them can actually take the crown.
Tom: So, what's the verdict? Is the end actually near?
Jane: Well, that's what we're going to dig into. The short answer is: it's complicated. The paper's conclusion suggests that for the biggest, most powerful "frontier" models, the Transformer is still king. But for smaller, edge devices, these new architectures are already making a real difference.
Tom: I love that. It's not a simple yes or no. It's a nuanced battlefield. We've got a lot to unpack here, so let's get into the details of what these alternative architectures actually are.
Jane: Sounds good. Let's do it.
Summary: Tom: So, we're back with "The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures." Jane, we've established that attention is expensive. What are the paper's main categories of alternatives trying to do about it?
Jane: The paper organizes the field into a few big buckets. First, you have sub-quadratic attention variants. These are attempts to make attention itself faster, either by approximating the calculations or by making the attention pattern sparse, so the model only looks at a few relevant tokens instead of everything.
Tom: And then there are the linear RNNs and State Space Models, right? Those are the ones that seem to get the most hype.
Jane: Right. Instead of looking at the whole sequence at once, these models process it step-by-step, compressing what they've seen into a fixed-size "state" or memory. Think of it like reading a book and taking notes on a single index card. You can't write down everything, but you can capture the gist. That makes them incredibly fast and memory-efficient.
Tom: But the paper points out a serious weakness there. If you only have that one index card, you can't recall a specific fact from page three hundred if you didn't write it down. The paper mentions this as a limitation in "lookup table" tasks, like associative recall.
Jane: Exactly. That's the trade-off. You trade the ability to instantly recall any detail for speed and efficiency. That's why the third category, hybrids, is so interesting. These models try to get the best of both worlds by combining attention with these faster mechanisms.
Tom: The paper calls them "striped" and "fusion" hybrids. Can you explain that?
Jane: Sure. A striped hybrid is like a sandwich—you have alternating layers. One layer might be a fast state space model, and the next layer is a full attention layer. A fusion hybrid is more like a smoothie—both mechanisms are computed in parallel and their outputs are combined.
Tom: And what are the results? Does this actually work?
Jane: The paper has a great benchmark table. In the smaller model size range, around one billion parameters, some of these alternatives like Samba and RWKV-seven actually outperform full-attention models like Llama three point two on several reasoning benchmarks. That's a big deal.
Tom: But then you look at the bigger models, the fourteen to seventy billion parameter range, and the picture changes completely.
Jane: It does. In that range, the top performers are still the pure Transformer-based models like Qwen and Llama. The hybrids like Griffin and Jamba are competitive, but they're not dominating. And when you look at the absolute frontier models, the ones with hundreds of billions of parameters, there are no sub-quadratic models in the top ten.
Tom: So, the bigger you get, the more the Transformer's power seems to win out, despite the cost.
Jane: Precisely. It suggests that when you have massive compute, the expressiveness of full attention is hard to beat. But when compute is scarce, these alternatives are becoming very compelling.
Tom: That's a fantastic summary. Now, I want to get into the "why" behind those results. What are the fundamental limits that keep these alternatives from taking over?
Jane: That's the perfect setup for our next segment.
Improvements: Tom: Welcome back. We're still digging into "The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures." Jane, we saw the benchmark results. Now, the paper gets into the theoretical weeds about why these architectures are limited. What did you find most interesting there?
Jane: The most fascinating part for me was the section on fundamental limitations. The paper points out that both Transformers and State Space Models belong to the same theoretical complexity class, called TC0. That's a mathematical way of saying they have similar theoretical limits on what they can compute.
Tom: So, even though SSMs are faster, they're not more powerful in theory?
Jane: Exactly. They're both limited in their ability to handle tasks that require true step-by-step logic, like tracking state or simulating a finite automaton. But the paper also highlights a unique weakness of the recurrent models: their state size is fixed.
Tom: Right, the index card problem we talked about earlier.
Jane: Exactly. The paper argues that this fixed state creates a fundamental problem for tasks that require recalling arbitrary information from the input. It even cites proofs showing that these models need a certain amount of memory just to perform simple recall tasks reliably, whereas a Transformer can theoretically attend to any token directly.
Tom: So, the improvements these new architectures offer—the speed, the efficiency—come at the cost of a hard ceiling on their recall ability.
Jane: That's the core tension. The paper calls it a "lookup table" limitation. If the answer is in the input, a Transformer can find it. A state space model might have already forgotten it.
Tom: But hold on, the paper also talks about how hybrids are trying to fix this. They add a little bit of attention back in to handle those recall tasks, right?
Jane: Yes. That's why hybrids are so promising. They use the fast, efficient layers for the bulk of the processing, and then sprinkle in attention layers specifically to handle the tasks where the recurrent models fail. It's a way to get most of the speed benefit while patching the biggest weakness.
Tom: And the paper also mentions some really novel ideas, like the Titans architecture with its "neural long-term memory." That's a different way of thinking about the problem altogether.
Jane: It is. Instead of just compressing everything into a state, it's about explicitly deciding what's surprising and worth remembering. It's a much more sophisticated memory system.
Tom: So, the "improvements" aren't just about making things faster. They're about redesigning how these models remember and process information.
Jane: Right. The paper suggests the future isn't a single architecture, but a toolbox of different primitives that we can mix and match based on the task.
Tom: I think we're ready to wrap this up. Let's bring it home with our final thoughts.
Conclusion: Tom: Alright, we're at the end of our discussion on "The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures." Jane, if you had to give our listeners the one-sentence takeaway, what would it be?
Jane: I'd say this: the Transformer isn't dead, but its monopoly is over. This paper from the Technical University of Munich shows us that we're entering an era of architectural diversity, where the best tool depends on the job.
Tom: And that's a really important point. For the massive, frontier models, the pure Transformer is still the champion. Its power and expressiveness are unmatched when you have the compute to feed it.
Jane: But for the edge—your phone, a smart speaker, a car—where memory and battery are precious, the sub-quadratic architectures like Mamba, RWKV, and the various hybrids are already winning. They offer a way to get capable AI without the massive overhead.
Tom: The paper's analysis of the trade-offs was so clear. It's not about which architecture is "better." It's about understanding the fundamental limits of each approach.
Jane: And that's what I appreciated most. They didn't just list the models. They explained *why* they work and *why* they fail. The theoretical limitations, like the recall problem with fixed-size states, are crucial for anyone trying to build these systems.
Tom: So, what's the future look like?
Jane: The paper hints at a future of "mixture of architectures," where a single system might route different types of queries to different specialized components. It's a much more flexible and efficient vision than just scaling up one giant Transformer.
Tom: I love that. It's a shift from "one model to rule them all" to a more modular, specialized approach. We're saying goodbye to this paper, but we're definitely going to be following the work it surveys.
Jane: Absolutely. It's an exciting time. The field is opening up, and the next big breakthrough might not look like a Transformer at all.
Tom: Well said. Thanks for joining us, everyone. We'll be back soon with another paper. Until then, keep questioning the status quo.
Jane: See you next time.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization