The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures".
Jane: The paper was written by Alexander M. Fichtl, Jeremias Bohn, Josefin Kelber, Edoardo Mosca and Georg Groh from Technical University of Munich.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making waves with a pretty bold title: "The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures." Jane, I have to say, just reading that title gave me a little shiver.
Jane: Same here, Tom. It's a question a lot of people in the field have been whispering about for a while. The paper, from the team at the Technical University of Munich, is essentially asking whether the architecture that powers basically every major language model is about to get dethroned.
Tom: Right, and for our listeners who might not be deep in the weeds, can you break down what a Transformer actually is and why this question matters so much?
Jane: Absolutely. So, the Transformer is the backbone of models like GPT and Llama. Its superpower is something called "attention," which lets the model look at every single word in a sentence and figure out how it relates to every other word. That's incredibly powerful for understanding context.
Tom: But there's a catch, right? That power comes at a cost.
Jane: A huge cost. The paper calls it quadratic complexity. If you double the length of your input text, the amount of computation needed doesn't just double—it quadruples. So, processing a one hundred thousand-word document is exponentially more expensive than a ten thousand-word one.
Tom: And that's the bottleneck this paper is all about. The authors from TUM are surveying all the attempts to break through that wall. They're looking at things like state space models, linear RNNs, and hybrids that mix different approaches.
Jane: Exactly. The title is a bit provocative, but the paper itself is a really careful survey. They're not just declaring the Transformer dead. They're laying out the contenders and asking if any of them can actually take the crown.
Tom: So, what's the verdict? Is the end actually near?
Jane: Well, that's what we're going to dig into. The short answer is: it's complicated. The paper's conclusion suggests that for the biggest, most powerful "frontier" models, the Transformer is still king. But for smaller, edge devices, these new architectures are already making a real difference.
Tom: I love that. It's not a simple yes or no. It's a nuanced battlefield. We've got a lot to unpack here, so let's get into the details of what these alternative architectures actually are.
Jane: Sounds good. Let's do it.
Summary: Tom: So, we're back with "The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures." Jane, we've established that attention is expensive. What are the paper's main categories of alternatives trying to do about it?
Jane: The paper organizes the field into a few big buckets. First, you have sub-quadratic attention variants. These are attempts to make attention itself faster, either by approximating the calculations or by making the attention pattern sparse, so the model only looks at a few relevant tokens instead of everything.
Tom: And then there are the linear RNNs and State Space Models, right? Those are the ones that seem to get the most hype.
Jane: Right. Instead of looking at the whole sequence at once, these models process it step-by-step, compressing what they've seen into a fixed-size "state" or memory. Think of it like reading a book and taking notes on a single index card. You can't write down everything, but you can capture the gist. That makes them incredibly fast and memory-efficient.
Tom: But the paper points out a serious weakness there. If you only have that one index card, you can't recall a specific fact from page three hundred if you didn't write it down. The paper mentions this as a limitation in "lookup table" tasks, like associative recall.
Jane: Exactly. That's the trade-off. You trade the ability to instantly recall any detail for speed and efficiency. That's why the third category, hybrids, is so interesting. These models try to get the best of both worlds by combining attention with these faster mechanisms.
Tom: The paper calls them "striped" and "fusion" hybrids. Can you explain that?
Jane: Sure. A striped hybrid is like a sandwich—you have alternating layers. One layer might be a fast state space model, and the next layer is a full attention layer. A fusion hybrid is more like a smoothie—both mechanisms are computed in parallel and their outputs are combined.
Tom: And what are the results? Does this actually work?
Jane: The paper has a great benchmark table. In the smaller model size range, around one billion parameters, some of these alternatives like Samba and RWKV-seven actually outperform full-attention models like Llama three point two on several reasoning benchmarks. That's a big deal.
Tom: But then you look at the bigger models, the fourteen to seventy billion parameter range, and the picture changes completely.
Jane: It does. In that range, the top performers are still the pure Transformer-based models like Qwen and Llama. The hybrids like Griffin and Jamba are competitive, but they're not dominating. And when you look at the absolute frontier models, the ones with hundreds of billions of parameters, there are no sub-quadratic models in the top ten.
Tom: So, the bigger you get, the more the Transformer's power seems to win out, despite the cost.
Jane: Precisely. It suggests that when you have massive compute, the expressiveness of full attention is hard to beat. But when compute is scarce, these alternatives are becoming very compelling.
Tom: That's a fantastic summary. Now, I want to get into the "why" behind those results. What are the fundamental limits that keep these alternatives from taking over?
Jane: That's the perfect setup for our next segment.
Improvements: Tom: Welcome back. We're still digging into "The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures." Jane, we saw the benchmark results. Now, the paper gets into the theoretical weeds about why these architectures are limited. What did you find most interesting there?
Jane: The most fascinating part for me was the section on fundamental limitations. The paper points out that both Transformers and State Space Models belong to the same theoretical complexity class, called TC0. That's a mathematical way of saying they have similar theoretical limits on what they can compute.
Tom: So, even though SSMs are faster, they're not more powerful in theory?
Jane: Exactly. They're both limited in their ability to handle tasks that require true step-by-step logic, like tracking state or simulating a finite automaton. But the paper also highlights a unique weakness of the recurrent models: their state size is fixed.
Tom: Right, the index card problem we talked about earlier.
Jane: Exactly. The paper argues that this fixed state creates a fundamental problem for tasks that require recalling arbitrary information from the input. It even cites proofs showing that these models need a certain amount of memory just to perform simple recall tasks reliably, whereas a Transformer can theoretically attend to any token directly.
Tom: So, the improvements these new architectures offer—the speed, the efficiency—come at the cost of a hard ceiling on their recall ability.
Jane: That's the core tension. The paper calls it a "lookup table" limitation. If the answer is in the input, a Transformer can find it. A state space model might have already forgotten it.
Tom: But hold on, the paper also talks about how hybrids are trying to fix this. They add a little bit of attention back in to handle those recall tasks, right?
Jane: Yes. That's why hybrids are so promising. They use the fast, efficient layers for the bulk of the processing, and then sprinkle in attention layers specifically to handle the tasks where the recurrent models fail. It's a way to get most of the speed benefit while patching the biggest weakness.
Tom: And the paper also mentions some really novel ideas, like the Titans architecture with its "neural long-term memory." That's a different way of thinking about the problem altogether.
Jane: It is. Instead of just compressing everything into a state, it's about explicitly deciding what's surprising and worth remembering. It's a much more sophisticated memory system.
Tom: So, the "improvements" aren't just about making things faster. They're about redesigning how these models remember and process information.
Jane: Right. The paper suggests the future isn't a single architecture, but a toolbox of different primitives that we can mix and match based on the task.
Tom: I think we're ready to wrap this up. Let's bring it home with our final thoughts.
Conclusion: Tom: Alright, we're at the end of our discussion on "The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures." Jane, if you had to give our listeners the one-sentence takeaway, what would it be?
Jane: I'd say this: the Transformer isn't dead, but its monopoly is over. This paper from the Technical University of Munich shows us that we're entering an era of architectural diversity, where the best tool depends on the job.
Tom: And that's a really important point. For the massive, frontier models, the pure Transformer is still the champion. Its power and expressiveness are unmatched when you have the compute to feed it.
Jane: But for the edge—your phone, a smart speaker, a car—where memory and battery are precious, the sub-quadratic architectures like Mamba, RWKV, and the various hybrids are already winning. They offer a way to get capable AI without the massive overhead.
Tom: The paper's analysis of the trade-offs was so clear. It's not about which architecture is "better." It's about understanding the fundamental limits of each approach.
Jane: And that's what I appreciated most. They didn't just list the models. They explained *why* they work and *why* they fail. The theoretical limitations, like the recall problem with fixed-size states, are crucial for anyone trying to build these systems.
Tom: So, what's the future look like?
Jane: The paper hints at a future of "mixture of architectures," where a single system might route different types of queries to different specialized components. It's a much more flexible and efficient vision than just scaling up one giant Transformer.
Tom: I love that. It's a shift from "one model to rule them all" to a more modular, specialized approach. We're saying goodbye to this paper, but we're definitely going to be following the work it surveys.
Jane: Absolutely. It's an exciting time. The field is opening up, and the next big breakthrough might not look like a Transformer at all.
Tom: Well said. Thanks for joining us, everyone. We'll be back soon with another paper. Until then, keep questioning the status quo.
Jane: See you next time.
Technical University of Munich
cs.CL
Submitted: 2025-10-06
Updated: 2025-10-06
Comments: 21 pages, 2 figures, 2 tables
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 98/100
The gist: The paper "The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures" by Alexander M.
Key concepts
- Transformer
- The core architecture behind major language models like GPT and Llama. Its 'attention' mechanism is powerful for understanding context but suffers from quadratic complexity, meaning computation cost increases exponentially with input length.
- Quadratic Complexity
- A computational bottleneck where the required processing power grows according to the square of the input size. For example, doubling the text length quadruples the necessary computation.
- Sub-Quadratic Architectures
- Alternative models (like State Space Models and linear RNNs) designed to process text more efficiently than Transformers. They achieve speed by avoiding full attention calculations, making them ideal for memory-constrained devices.
- Hybrid Models
- Architectures that combine the best parts of different systems, such as mixing fast State Space Models with targeted attention layers. This aims to gain efficiency while patching weaknesses like poor recall ability.
Terminology
Summary
The paper The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
by Alexander M. Fichtl, Jeremias Bohn, Josefin Kelber, Edoardo Mosca and Georg Groh (Social Computing Group, Technical University of Munich) surveys recent efforts to overcome the quadratic complexity bottleneck of the transformer attention mechanism.
Motivation and Scope: The authors note that "Transformers have dominated sequence processing tasks for the past seven years—most notably language modeling. However, the inherent quadratic complexity of their attention mechanism remains a significant bottleneck as context length increases. The paper surveys
advances in (sub-quadratic) attention variants, recurrent neural networks, state space models, and hybrid architectures and critically analyzes
these approaches in terms of compute and memory complexity, benchmark results, and fundamental limitations to assess whether the dominance of pure-attention transformers may soon be challenged."
Main Contributions: The authors identify three main contributions: (1) A systematic review of the most relevant (sub)quadratic attention variants, RNNs, SSMs, and hybrid architectures
; (2) A comparative analysis of time and memory complexity for training and inference of sequence-modeling mechanisms, as well as reported benchmark results for SOTA models
; and (3) A critical analysis of strengths, tradeoffs, and limitations, with an informed perspective on when and where pure attention-based transformers may be surpassed.
Modern O(n2) Attention: The paper notes that The fundamental architectural principle of quadratic attention has not changed much over recent years
but highlights system-level improvements including KV cache reduction techniques (Multi-Query Attention, Grouped-Query Attention, Multi-Head Latent Attention), FlashAttention and its successors (which reduce memory usage to be linear in sequence length and deliver 2–4× runtime speedups
), and Paged Attention (boosting throughput 2–4×).
Sub-Quadratic Architectures: The paper categorizes alternatives into four types: non-recurrent attention variants, linear RNN-based models, state space models, and hybrids.
Non-Recurrent O(n2−e) Attention Variants: These include approximate attention (kernel-based linear attention achieving O(n) complexity, low-rank methods like Linformer) and sparse attention (local sliding window attention, strided/random patterns, learnable sparsity). The authors note that While some sparse patterns can achieve O(n) time and memory complexity, they may underperform on tasks requiring fine-grained global dependencies and often require task-specific tuning.
Linear RNN-based Models: The paper discusses HGRN (using data-dependent, dynamic decay rates via forget gates
), xLSTM (enhancing LSTM with state expansion, exponential gating, normalization, and stabilization techniques
), Lightning Attention (dividing attention into intra-block standard attention and inter-block linear attention), and RWKV (building on the Attention Free Transformer, with RWKV-7 Goose
introducing a generalized delta rule, vector-valued gating, in-context learning rates, and a relaxed value replacement rule
). The authors note that "RNNs and their variants offer linear autoregressive generation, but suffer from (1) varying degrees of vanishing/exploding gradients, (2) limited training parallelism, and (3) lack of expressivity due to a representation state not scaling with context length."
State-Space-based Models: Structured SSMs like S4 introduce the use of a Highly Predictive Polynomial Projection Operator (HiPPO) matrix for initializing the state transition.
Selective SSMs like Mamba advance SSMs by replacing fixed transition matrices with input-dependent functions,
while Mamba2 further unifies structured SSMs with attention mechanisms.
The authors note the dual perspective
of SSMs: a recurrent formulation enables O(n) inference, while a convolutional view allows for O(n log(n)) training via efficient FFT-based convolutions.
Hybrids: The paper distinguishes between striped
hybrids (block types connected in series) and fusion
hybrids (block types connected in parallel). Notable examples include Jamba (merging transformer, Mamba, and Mixture-of-Experts (MoE) layers into a striped hybrid
), Griffin (using Real-Gated Linear Recurrent Unit (RG-LRU)
as a fusion hybrid of local attention and linear recurrence
), Samba (a striped hybrid using sliding window attention and Mamba/SSM layers
), and MiniMax-01 (combining lightning attention with MoE). The paper also discusses novel architecture design concepts including memory system design (Titans, B'MOJO) and tailored architecture search (STAR framework).
Complexity Analysis: The paper provides a detailed complexity comparison table. Key findings include that RWKV achieves O(nd) inference time and O(d) inference space, while most other methods require O(nd) or more. The authors note that these complexities are sometimes dominated by feedforward neural networks in the full model
and that Many of these algorithms rely on projections, thus requiring at least O(nd2) operations, often serving as an upper bound for time complexity.
Benchmark Performance: The paper presents a comparison of models in two parameter ranges (0.7-1.5B and 14-70B) across eight benchmarks. In the low-parameter setting, Samba and RWKV7-World3 significantly outperform the full attention Llama 3.2 and Qwen2.5 in several instances.
In the midrange, no pure sub-quadratic models are present anymore; merely the hybrids Griffin and Jamba remain, with only the latter realistically competing with Qwen2.5 and Llama3.1.
For frontier models (100B+), the authors note that only MiniMax-Text-01 appears in the top-20 ranking once (Text Arena leaderboard), but among the top 10, we cannot find any single model known to be built on an alternative architecture.
Fundamental Limitations: The paper discusses limitations of attention, noting that The standard transformer forward pass belongs to the log-time uniform T C0 circuit complexity class,
which fundamentally limits its ability to simulate finite automata or solve graph connectivity—necessary for state tracking and multi-step reasoning.
Transformers also struggle to extrapolate, i.e., to generalize from shorter training context sizes to longer test sequences.
For sub-quadratic alternatives, the authors note that SSMs are also limited to the complexity class T C0
and that This finite state capacity has strong implications for 'lookup table' tasks... as SSMs cannot recall an arbitrary amount of information previously seen.
The paper also notes that linear attention is not injective, often assigning identical attention weights to different queries and causing semantic confusion.
Key Conclusions: The authors state that at the time of writing, most frontier general-purpose models strongly rely on full attention mechanisms
and that full attention is free from many limitations that apply to alternative architectures.
However, the picture changes for edge models, where compute, memory, and latency are tightly bound, and alternative architectures have gained substantial traction.
The paper concludes: "At the frontier, full attention is likely to remain central for the foreseeable future. Still, even these models may begin incorporating hybrid elements... The shift is not toward replacement, but toward building flexible systems from a growing set of specialized primitives. The authors ultimately conclude that sub-quadratic alternatives
remain fundamentally constrained in generality compared to transformers and will not compete in the frontier for the foreseeable future."
Improvements for AI systems
Based on the paper, here are specific improvements that can be made to AI systems, along with what the improved systems can do:
Improvement: Replace pure-attention transformers with striped or fusion hybrids (e.g., Samba-style: sliding window attention + Mamba/SSM layers) for on-device models under 2B parameters.
What the improved system can do:
-
Achieve 2–4× lower inference latency on long sequences (e.g., 8K+ tokens) while maintaining ≥95% of full-attention benchmark scores on ARC-C, HellaSwag, and PIQA.
-
Process context windows up to 1M tokens at reasonable memory cost (linear in sequence length instead of quadratic), enabling document-level analysis on mobile hardware.
-
Maintain constant O(1) memory per token during autoregressive generation, allowing indefinite streaming without cache overflow.
Improvement: Integrate a meta-learned long-term memory module (as in Titans) that stores surprising or rare tokens at test time, alongside the existing short-term attention window.
Improvement: Replace standard multi-head KV caching with GQA (shared key/value across heads) combined with a low-rank latent projection (MLA-style) to compress cached states.
Improvement: Apply FlashAttention-3’s asynchronous, low-precision (FP8) tiling strategy to the intra-block attention computation in Lightning Attention or hybrid SSM+attention layers.
Improvement: Replace fixed or learned absolute positional encodings with NoPE (no positional encoding) or relative encodings, combined with a curriculum that gradually increases context length during training (as suggested by Veitsman et al., 2025).
Improvement: Insert a small attention-based retrieval module (e.g., 1–2 layers of full attention every 8 recurrent layers) to compensate for the finite-state recall bottleneck of Mamba, RWKV, or xLSTM.
Improvement: Automatically search over Linear Input-Varying (LIV) system configurations (as in Thomas et al., 2024) to generate a custom hybrid architecture for a given target metric (e.g., latency, perplexity, cache size).
Improvement: Build a router that dynamically selects between attention, SSM, and linear RNN layers per token or per layer, based on input complexity (e.g., presence of rare tokens, long-range dependencies).
Improvement: For tasks requiring state tracking or multi-step reasoning (e.g., code execution, planning), augment sub-quadratic models with a small, explicit “scratchpad” module that performs chain-of-thought (CoT) steps, as allowed by the T C 0 limitations.
Improvement: Use the finding that hybrids (e.g., Jamba) underperform on instruction-following due to sparse training data—not architecture—to curate a more balanced dataset mixing long-context, instruction-tuning, and reasoning examples.
These improvements are directly actionable and grounded in the paper’s analysis of complexity, benchmark results, and architectural limitations. They prioritize efficiency, recall, and length generalization while acknowledging the theoretical boundaries of both attention and its alternatives.
Abstract
Transformers have dominated sequence processing tasks for the past seven years -- most notably language modeling. However, the inherent quadratic complexity of their attention mechanism remains a significant bottleneck as context length increases. This paper surveys recent efforts to overcome this bottleneck, including advances in (sub-quadratic) attention variants, recurrent neural networks, state space models, and hybrid architectures. We critically analyze these approaches in terms of compute and memory complexity, benchmark results, and fundamental limitations to assess whether the dominance of pure-attention transformers may soon be challenged.
Sources
- Titans: Learning to Memorize at Test Time
- Longformer: The Long-Document Transformer
- RecurrentGemma: Moving Past Transformers for Efficient Open Language Models
- Evaluating Large Language Models Trained on Code
- Generating Long Sequences with Sparse Transformers
- Rethinking Attention with Performers
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
- The Llama 3 Herd of Models
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Log-Linear Attention
- Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
- Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers
- Mixtral of Experts
- Scaling Laws for Neural Language Models
- MiniMax-01: Scaling Foundation Models with Lightning Attention
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- Untangling the Mechanisms of Misleading Context in Medical Question Answering
- SCOPE: A Generative Approach for LLM Prompt Compression