Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores
summary
The gist
Vision Transformers (ViTs) are challenged by their quadratic computational cost due to dense token-to-token self-attention, which limits their scalability in high-resolution domains.
In short
Vision Transformers face quadratic costs from dense self-attention. VECA proposes replacing this with an efficient core-periphery structure using a small set of learned 'core tokens.' This architecture allows patches to communicate indirectly through these cores, achieving linear computational scaling while maintaining high visual representation quality.
Key concepts
- Core-Periphery Structure
- This is the attention mechanism where a small group of learned 'core tokens' (the core) interacts with all image patches (the periphery). Instead of every patch talking to every other patch, communication flows through these central cores, creating a sparse and efficient network structure.
- Active Core Budget
- VECA is trained to use different subsets of the learned core tokens at each step. This allows the model to learn a 'coarse-to-fine' representation: early cores capture broad information, while later cores focus on more specific, local details.
- Core Coordinate State
- Each core token has a continuous spatial coordinate that changes as the network learns. This state is updated layer by layer using a learned residual mechanism. This allows each core to maintain both its semantic meaning and an evolving position in the image plane.
- Budget-Adaptive Representation Learning
- The model samples different active core budgets during training. This forces the cores to specialize; some cores become 'object-centric' and attend to diverse objects, while others evolve into semantically aligned groups across different network layers.
Terminology used across episodes
This episode discusses
- Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores · Paper Radio
- How Much Position Information Do Convolutional Neural Networks Encode?
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- How Do Vision Transformers Work?
- Contrastive Learning Rivals Masked Image Modeling in Fine-tuning via Feature Distillation
- Theia: Distilling Diverse Vision Foundation Models for Robot Learning
- DINOv3
- Triangular Dropout: Variable Network Width without Retraining
- Distributional Principal Autoencoders
- A PCA-like Autoencoder
- Slimmable Neural Networks
- Once-for-All: Train One Network and Specialize it for Efficient Deployment
- SortedNet: A Scalable and Generalized Framework for Training Modular Deep Neural Networks
- Flextron: Many-in-One Flexible Large Language Model
- Adaptive Computation Time for Recurrent Neural Networks
- Universal Transformers
- Accelerating Optimization via Differentiable Stopping Time
- LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation
- Token Merging: Your ViT But Faster
- ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inference
- Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
The paper
Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores · Read on arXiv
Carnegie Mellon University · University of Hong Kong (HKU) · Columbia University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores".
Tom: Vision Transformers (ViTs) are challenged by their quadratic computational cost due to dense token-to-token self-attention, which limits their scalability in high-resolution domains.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, we've been diving into the paper "Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores," and it seems like the main idea is challenging the idea that all those dense patch-to-patch interactions are absolutely necessary for a vision transformer to learn good visual representations.
Jane: Exactly, Tom; this paper proposes VECA, which suggests we can achieve effective learning without needing every single patch to talk directly to every other patch.
Lu: What's really compelling here is the core-periphery structure they introduce, where you have a small set of learned "core tokens" acting as a clique that mediates the communication between all the image patches, which is pretty creative thinking for this kind of architecture.
Meng: From an engineering standpoint, it sounds like reducing that attention cost from quadratic to linear when C is fixed is something we can actually see in practice.
Lalam: I'm seeing a really interesting cultural implication here; if we can learn rich visual data with less brute-force computation, it means models could become much more accessible and deployable across a wider variety of devices.
Tom: That makes sense, Jane; the abstract claims that this architecture allows for learning without direct patch-to-patch interaction to achieve strong data-driven scaling.
Jane: It really does matter because the original ViTs scale strongly by using all-to-all self-attention, but that flexibility comes with a computational cost that grows quadratically with image resolution, which severely limits them in high-resolution tasks.
Meng: I'm looking at the FLOPs data they show; for instance, at the one thousand twenty-four by one thousand twenty-four resolution, the FLOPs for VECA-Small drop significantly from about 30 point 67G down to around 5 point 72G, which is a reduction of over five times.
Lu: And it’s not just about the raw speed; they also show that the method produces stable embeddings across different resolutions when compared against their teacher model, DINOv3. That stability is a big deal for deploying models in varied visual environments.
Lalam: And looking at the results, they achieved competitive performance with the DINOv3 teacher on classification and dense spatial tasks like VOC Context, showing a gap of less than zero point two eight mIoU difference on segmentation. This suggests that we can maintain high accuracy while using this more efficient structure.
Paper summary: Tom: So, to put it simply for our listeners, the paper "Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores" is proposing a new way for vision transformers to learn complex visual data by replacing the heavy, quadratic patch-to-patch attention with a more efficient core-periphery structure that uses only a small set of learned tokens to connect everything.
Jane: That's the essence of it, Tom; they are showing that effective visual representations can be learned without needing every single patch to interact directly with every other patch, which is a significant departure from the standard approach.
Meng: From an engineering perspective, the concept of maintaining spatially aligned dense representations while introducing this core-periphery structure seems like a clever way to balance computational efficiency and feature quality. I wonder how robust this coordination between the patch tokens and the cores holds up during deep processing layers.
Lu: The paper details how each core token gets an evolving spatial coordinate, updated through a learned residual mechanism, which allows them to maintain both a semantic representation and a position within the image plane. That dynamic positioning capability is something I find very exciting for future generative applications.
Lalam: And the idea of budget-adaptive learning, where they sample from an ordered core bank RM to create coarse-to-fine representations, suggests we can tailor the model's focus based on what we need in a specific task. This hints at a future where models could dynamically adjust their computational complexity on the fly.
Tom: It really puts things into perspective for us; this paper shows that the necessity of all-to-all attention isn't absolute, and by using these learned cores, we can trade compute for accuracy in a way that scales much better as image resolution increases.
Jane: And the implications are big because it opens up a new building block for vision models; if this elastic core-periphery attention structure is adopted, we could see much more efficient and scalable vision transformers in deployment across many applications.
Paper summary: Meng: I'm thinking about how this impacts real-world inference; having linear scaling with image resolution means we could run these kinds of complex visual tasks on edge devices that currently struggle with the quadratic scaling of standard ViTs. That practical efficiency is what gets my attention.
Lu: The emergent behaviors they observe in the core tokens, specifically how they evolve from being isotropic to becoming semantically aligned groups across layers under different core budgets, points toward a rich internal representation learning process. This suggests the cores are learning meaningful organizational structures within the visual data itself.
Lalam: If we look at this from a broader cultural view, it means that the underlying AI architecture can become inherently more flexible and less constrained by strict computational requirements for every single task, allowing for more diverse and specialized applications in vision systems.
Tom: So, to summarize for our listeners: the paper "Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores" argues that we don't always need dense patch-to-patch interaction, proposing VECA as a way to use efficient linear communication through learned core tokens to achieve strong visual learning, which is very relevant for scaling vision models.
Jane: It really highlights that the architecture of self-attention can be adapted based on the problem at hand, not just sticking to the standard quadratic approach.
Meng: And we should keep an eye on how this core-periphery structure plays out when we integrate it into larger, multimodal models; practical integration is always the next hurdle for these kinds of novel architectural ideas.
Lu: The way they handle the coordinate updates for those cores across layers through that learned residual mechanism is a sophisticated piece of design that shows how spatial awareness can be integrated into the attention mechanism itself. That level of detail in maintaining geometric context is impressive.
Lalam: For me, the implication for our culture is that this pushes us toward developing AI systems that are inherently more adaptive to resource constraints while still delivering high performance, which feels like a very necessary direction for responsible innovation.
Conclusion: Tom: So, we’ve been looking at how this paper tackles the big question of whether vision transformers really need that heavy, dense patch-to-patch attention mechanism and what it means for the authors and their work overall.
Jane: Right, Tom; essentially, they are exploring if we can learn excellent visual data representations without every single image patch having to directly talk to every other patch.
Lu: The authors of this paper are really pushing the boundaries of how we structure attention in these models, moving away from the standard quadratic scaling issue that plagues high-resolution vision tasks.
Meng: From a practical standpoint, I'm curious about the real-world impact; does this architecture offer a tangible way to make these massive models run more efficiently on actual hardware?
Lalam: The core idea is that by using just a small set of learned tokens, we can achieve global communication in an organized way that is much more manageable computationally.
Tom: Exactly; they’re showing that you don't always need all-to-all attention to get high-quality visual understanding, and the paper lays out how this core structure works in detail.
Jane: It’s about finding a smarter way for the AI to communicate across an image without incurring that massive computational price every time it processes a picture.
Lu: The methodology is clever because they build this core-periphery structure where those learned cores mediate the interaction between all the patches in a linear time fashion.
Meng: That linear scaling aspect is what keeps my interest piqued; if we can get better performance without that quadratic hit, it changes how we think about deploying these systems.
Lalam: I think this work has huge cultural implications because it suggests vision AI can become more resource-aware and scalable across different devices and applications.
Tom: It really is a fascinating look at the architecture itself, showing us an alternative pathway to building powerful vision models that handle high resolution better.
Jane: So, the main point for us as listeners is that this paper introduces a new way to structure attention that prioritizes efficiency while maintaining strong visual representation quality.
Lu: The authors’ approach with budget-adaptive learning also makes sense because it suggests we can tailor the model's focus based on what we need in a specific visual task.
Meng: That adaptive nature is where the real engineering challenge lies, figuring out how to implement that flexibility reliably without losing accuracy during training.
Lalam: I think this pushes us toward developing AI systems that are inherently more flexible and less constrained by strict computational requirements for every single task we throw at them.
Tom: We’ve seen the results show competitive performance with established models, which is a strong sign of the architecture's promise in practice.
Jane: It shows that this core-periphery idea isn't just theoretical; it actually delivers results on both classification and dense prediction tasks.
Lu: The way they handle those evolving spatial coordinates for the cores across different layers is a sophisticated design choice that really shows how deep geometric context can be maintained.
Meng: That level of detail in keeping track of spatial information during the network's processing is impressive, and I want to see how robust that stays under heavy workloads.
Lalam: Ultimately, this paper opens up a new building block for vision models, suggesting that we can trade compute for accuracy in a way that scales much better as image resolution increases.
Tom: It really is an important piece of research because it challenges the standard assumption about what self-attention requires to achieve high performance in vision AI.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck