Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores".
Tom: Vision Transformers (ViTs) are challenged by their quadratic computational cost due to dense token-to-token self-attention, which limits their scalability in high-resolution domains.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, we've been diving into the paper "Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores," and it seems like the main idea is challenging the idea that all those dense patch-to-patch interactions are absolutely necessary for a vision transformer to learn good visual representations.
Jane: Exactly, Tom; this paper proposes VECA, which suggests we can achieve effective learning without needing every single patch to talk directly to every other patch.
Lu: What's really compelling here is the core-periphery structure they introduce, where you have a small set of learned "core tokens" acting as a clique that mediates the communication between all the image patches, which is pretty creative thinking for this kind of architecture.
Meng: From an engineering standpoint, it sounds like reducing that attention cost from quadratic to linear when C is fixed is something we can actually see in practice.
Lalam: I'm seeing a really interesting cultural implication here; if we can learn rich visual data with less brute-force computation, it means models could become much more accessible and deployable across a wider variety of devices.
Tom: That makes sense, Jane; the abstract claims that this architecture allows for learning without direct patch-to-patch interaction to achieve strong data-driven scaling.
Jane: It really does matter because the original ViTs scale strongly by using all-to-all self-attention, but that flexibility comes with a computational cost that grows quadratically with image resolution, which severely limits them in high-resolution tasks.
Meng: I'm looking at the FLOPs data they show; for instance, at the one thousand twenty-four by one thousand twenty-four resolution, the FLOPs for VECA-Small drop significantly from about 30 point 67G down to around 5 point 72G, which is a reduction of over five times.
Lu: And it’s not just about the raw speed; they also show that the method produces stable embeddings across different resolutions when compared against their teacher model, DINOv3. That stability is a big deal for deploying models in varied visual environments.
Lalam: And looking at the results, they achieved competitive performance with the DINOv3 teacher on classification and dense spatial tasks like VOC Context, showing a gap of less than zero point two eight mIoU difference on segmentation. This suggests that we can maintain high accuracy while using this more efficient structure.
Paper summary: Tom: So, to put it simply for our listeners, the paper "Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores" is proposing a new way for vision transformers to learn complex visual data by replacing the heavy, quadratic patch-to-patch attention with a more efficient core-periphery structure that uses only a small set of learned tokens to connect everything.
Jane: That's the essence of it, Tom; they are showing that effective visual representations can be learned without needing every single patch to interact directly with every other patch, which is a significant departure from the standard approach.
Meng: From an engineering perspective, the concept of maintaining spatially aligned dense representations while introducing this core-periphery structure seems like a clever way to balance computational efficiency and feature quality. I wonder how robust this coordination between the patch tokens and the cores holds up during deep processing layers.
Lu: The paper details how each core token gets an evolving spatial coordinate, updated through a learned residual mechanism, which allows them to maintain both a semantic representation and a position within the image plane. That dynamic positioning capability is something I find very exciting for future generative applications.
Lalam: And the idea of budget-adaptive learning, where they sample from an ordered core bank RM to create coarse-to-fine representations, suggests we can tailor the model's focus based on what we need in a specific task. This hints at a future where models could dynamically adjust their computational complexity on the fly.
Tom: It really puts things into perspective for us; this paper shows that the necessity of all-to-all attention isn't absolute, and by using these learned cores, we can trade compute for accuracy in a way that scales much better as image resolution increases.
Jane: And the implications are big because it opens up a new building block for vision models; if this elastic core-periphery attention structure is adopted, we could see much more efficient and scalable vision transformers in deployment across many applications.
Paper summary: Meng: I'm thinking about how this impacts real-world inference; having linear scaling with image resolution means we could run these kinds of complex visual tasks on edge devices that currently struggle with the quadratic scaling of standard ViTs. That practical efficiency is what gets my attention.
Lu: The emergent behaviors they observe in the core tokens, specifically how they evolve from being isotropic to becoming semantically aligned groups across layers under different core budgets, points toward a rich internal representation learning process. This suggests the cores are learning meaningful organizational structures within the visual data itself.
Lalam: If we look at this from a broader cultural view, it means that the underlying AI architecture can become inherently more flexible and less constrained by strict computational requirements for every single task, allowing for more diverse and specialized applications in vision systems.
Tom: So, to summarize for our listeners: the paper "Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores" argues that we don't always need dense patch-to-patch interaction, proposing VECA as a way to use efficient linear communication through learned core tokens to achieve strong visual learning, which is very relevant for scaling vision models.
Jane: It really highlights that the architecture of self-attention can be adapted based on the problem at hand, not just sticking to the standard quadratic approach.
Meng: And we should keep an eye on how this core-periphery structure plays out when we integrate it into larger, multimodal models; practical integration is always the next hurdle for these kinds of novel architectural ideas.
Lu: The way they handle the coordinate updates for those cores across layers through that learned residual mechanism is a sophisticated piece of design that shows how spatial awareness can be integrated into the attention mechanism itself. That level of detail in maintaining geometric context is impressive.
Lalam: For me, the implication for our culture is that this pushes us toward developing AI systems that are inherently more adaptive to resource constraints while still delivering high performance, which feels like a very necessary direction for responsible innovation.
Conclusion: Tom: So, we’ve been looking at how this paper tackles the big question of whether vision transformers really need that heavy, dense patch-to-patch attention mechanism and what it means for the authors and their work overall.
Jane: Right, Tom; essentially, they are exploring if we can learn excellent visual data representations without every single image patch having to directly talk to every other patch.
Lu: The authors of this paper are really pushing the boundaries of how we structure attention in these models, moving away from the standard quadratic scaling issue that plagues high-resolution vision tasks.
Meng: From a practical standpoint, I'm curious about the real-world impact; does this architecture offer a tangible way to make these massive models run more efficiently on actual hardware?
Lalam: The core idea is that by using just a small set of learned tokens, we can achieve global communication in an organized way that is much more manageable computationally.
Tom: Exactly; they’re showing that you don't always need all-to-all attention to get high-quality visual understanding, and the paper lays out how this core structure works in detail.
Jane: It’s about finding a smarter way for the AI to communicate across an image without incurring that massive computational price every time it processes a picture.
Lu: The methodology is clever because they build this core-periphery structure where those learned cores mediate the interaction between all the patches in a linear time fashion.
Meng: That linear scaling aspect is what keeps my interest piqued; if we can get better performance without that quadratic hit, it changes how we think about deploying these systems.
Lalam: I think this work has huge cultural implications because it suggests vision AI can become more resource-aware and scalable across different devices and applications.
Tom: It really is a fascinating look at the architecture itself, showing us an alternative pathway to building powerful vision models that handle high resolution better.
Jane: So, the main point for us as listeners is that this paper introduces a new way to structure attention that prioritizes efficiency while maintaining strong visual representation quality.
Lu: The authors’ approach with budget-adaptive learning also makes sense because it suggests we can tailor the model's focus based on what we need in a specific visual task.
Meng: That adaptive nature is where the real engineering challenge lies, figuring out how to implement that flexibility reliably without losing accuracy during training.
Lalam: I think this pushes us toward developing AI systems that are inherently more flexible and less constrained by strict computational requirements for every single task we throw at them.
Tom: We’ve seen the results show competitive performance with established models, which is a strong sign of the architecture's promise in practice.
Jane: It shows that this core-periphery idea isn't just theoretical; it actually delivers results on both classification and dense prediction tasks.
Lu: The way they handle those evolving spatial coordinates for the cores across different layers is a sophisticated design choice that really shows how deep geometric context can be maintained.
Meng: That level of detail in keeping track of spatial information during the network's processing is impressive, and I want to see how robust that stays under heavy workloads.
Lalam: Ultimately, this paper opens up a new building block for vision models, suggesting that we can trade compute for accuracy in a way that scales much better as image resolution increases.
Tom: It really is an important piece of research because it challenges the standard assumption about what self-attention requires to achieve high performance in vision AI.
Carnegie Mellon University · University of Hong Kong (HKU) · Columbia University
cs.CV, cs.LG
Submitted: 2026-05-12
Updated: 2026-09-30
Comments: Project repository here: https://github.com/alansong1322/VECA
Code: https://github.com/open-mmlab/mmsegmentation
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 86/100
The gist: Vision Transformers (ViTs) are challenged by their quadratic computational cost due to dense token-to-token self-attention, which limits their scalability in high-resolution domains.
Key concepts
- Core-Periphery Structure
- This is the attention mechanism where a small group of learned 'core tokens' (the core) interacts with all image patches (the periphery). Instead of every patch talking to every other patch, communication flows through these central cores, creating a sparse and efficient network structure.
- Active Core Budget
- VECA is trained to use different subsets of the learned core tokens at each step. This allows the model to learn a 'coarse-to-fine' representation: early cores capture broad information, while later cores focus on more specific, local details.
- Core Coordinate State
- Each core token has a continuous spatial coordinate that changes as the network learns. This state is updated layer by layer using a learned residual mechanism. This allows each core to maintain both its semantic meaning and an evolving position in the image plane.
- Budget-Adaptive Representation Learning
- The model samples different active core budgets during training. This forces the cores to specialize; some cores become 'object-centric' and attend to diverse objects, while others evolve into semantically aligned groups across different network layers.
Terminology
Summary
Vision Transformers (ViTs) are challenged by their quadratic computational cost due to dense token-to-token self-attention, which limits their scalability in high-resolution domains. This work proposes VECA (Visual Elastic Core Attention), a vision transformer architecture that replaces dense patch-to-patch interaction with an efficient linear-time core-periphery structure mediated by a small set of learned core tokens, demonstrating that effective visual representations can be learned without direct patch-to-patch interaction.
How it works
The VECA architecture constructs a block-sparse attention matrix based on a core-periphery structure where image patches form the periphery
and a small set of learned cores
forms a fully connected clique. This structure eliminates the need for patches to directly attend to one another. The computational cost of this attention operation becomes (2NC + C 2), where N is the number of image patches and C is the number of active core tokens, yielding linear complexity O(N) for a predetermined C, which bypasses quadratic scaling.
The architecture maintains spatially aligned dense representations because patch tokens are maintained and updated throughout the network. The input sequence to each block is formed by concatenating an ordered prefix of active core tokens (RC) with the patch sequence (Z), denoted as X = [RC;Z]. This structure allows core tokens to attend to the full sequence, while patch tokens attend only to the active core-token prefix, enabling them to receive global context through the core interface.
Core-Mediated Visual Attention
The attention mechanism is defined such that active core tokens attend to all active cores and all patches,
while patch tokens attend only to the active core-token prefix.
This results in an attention connectivity graph forming a core-periphery network with a graph diameter of 2. The resulting attention cost is (2NC + C 2), which scales linearly with image resolution when C is predetermined.
Each core token is associated with a continuous image- and layer-dependent planar coordinate, initialized by farthest-point sampling in the normalized image plane and updated across layers via a learned residual mechanism: ul+1i = tanh(ρl+1i),
where the core coordinate state ρl+1i is updated using ρl+1i = ρli + αlfl(rli).
This allows each core to maintain both a semantic representation and an evolving spatial position.
Budget-Adaptive Representation Learning
VECA is trained to support multiple active-core budgets within a single shared model by leveraging an ordered core bank RM = (r1,..., rM). The training procedure samples C ∼ pC (·) at each optimizer step from the set of active core budgets. This nested design induces a spatially and semantically coarse-to-fine representation.
Early cores are encouraged to encode the most broadly useful information,
while later cores specialize in more spatially local or semantically specific information.
Evaluation and Performance
VECA achieves competitive performance with the DINOv3 teacher across classification and dense spatial tasks. Specifically, for image classification at ViT-B size, VECA achieves less than 2% top-1 Acc. difference (81.93 versus 83.56) on Imagenet-1k classification.
For dense prediction tasks like semantic segmentation (VOC Context), VECA achieves competitive performance with the DINOv3 teacher,
with a gap of less than 0.28 mIoU difference (57.46 versus 57.74).
Efficiency and Emergent Behaviors
The computational cost advantage is demonstrated across resolutions, where VECA exhibits linear scaling while standard ViT self-attention scales quadratically. For instance, for VECA-Small at the largest resolution (1024 × 1024), FLOPs decrease from 30.67G to 5.72G, a 5.36× reduction.
Furthermore, core tokens exhibit emergent behaviors: they are strongly object centric
and attend to semantically diverse objects,
and they evolve from isotropic (spherical) to semantically-aligned groups
across layers under different active core budgets during feedforward processing.
Conclusion
VECA establishes elastic core-periphery attention as a scalable alternative building block for Vision Transformers, showing that direct patch-to-patch self-attention is not strictly necessary for high-quality dense visual representation learning. The model's ability to elastically trade off compute and accuracy during inference
makes it a promising scalable building block.
The gist
Effective visual representations can be learned without any direct patch-to-patch interaction through VECA, a vision transformer architecture that uses efficient linear-time core-periphery structured attention enabled by a small set of learned cores.
How it works
Improvements for AI systems
As a fastidious and diligent researcher, I have thoroughly analyzed the VECA (Visual Elastic Core Attention) architecture described in this paper. The core innovation lies in replacing quadratic self-attention with an elastic, linear-time core-periphery structured attention mechanism mediated by a small set of learned core
tokens.
Based on the findings, here are specific improvements and capabilities this architecture enables for AI systems:
)
- Improvements to Vision Transformer (ViT) Scalability and Efficiency:
The primary improvement is the transition from quadratic complexity, which limits ViTs in high-resolution domains, to linear complexity, scaling as O(N). This allows for the deployment of vision models on much larger images (e.g., 1024x1024 or higher) without prohibitive computational costs.
- Specific Capability: Enables real-time processing and training of vision models on high-resolution inputs (e.g., medical imaging, high-def satellite imagery) that were previously infeasible due to memory constraints.
- Elastic Inference and Adaptive Compute Allocation:
The model is inherently elastic through the nested training
mechanism and the ability to select an active core budget during inference without retraining.
- Specific Capability: Allows AI systems to dynamically trade off accuracy and speed based on the required task complexity or latency constraints. A system can use a small core set (low compute, good global recognition) for quick preliminary screening, or activate a larger set of cores (higher compute, fine-grained detail) for high-precision tasks like detailed semantic segmentation.
- Superior Dense Prediction Performance:
The paper demonstrates that VECA achieves performance competitive with DINOv3 across dense prediction tasks (semantic segmentation and depth estimation), often matching or exceeding baselines even when the active core budget is significantly reduced (e.g., C=8).
- Specific Capability: Enables highly accurate, pixel-level semantic segmentation and precise monocular depth estimation for complex scenes where fine spatial details are critical, which is essential for robotics, autonomous driving perception, and 3D reconstruction.
- Emergent Object-Centric Representation Learning:
The analysis reveals that core tokens develop strong object-centric representations and learn complementary visual roles across layers (evolving from isotropic to semantically clustered structures).
- Specific Capability: The system learns to focus on coherent semantic regions rather than just patch interactions, leading to more robust feature extraction capable of identifying distinct objects and parts consistently across different scales.
- Resolution Invariance in Global Recognition:
The results show that the model maintains strong global recognition accuracy (e.g., <1% delta on ImageNet-1K) even when operating at reduced core budgets or higher resolutions compared to standard ViTs (DINOv3).
- Specific Capability: Vision models can be deployed across a wide range of input resolutions while maintaining a consistent level of high-level semantic understanding, improving robustness in diverse visual environments.
- Interpretability through Core Specialization:
The visualization analysis shows that different core tokens specialize in distinct visual structures (e.g., attending to complementary objects or spatial regions).
- Specific Capability: Provides a degree of interpretability, allowing researchers to understand which specific learned representations (core tokens) are responsible for capturing particular visual concepts, aiding in the development of more specialized and efficient vision models.
Abstract
Vision Transformers (ViTs) learn rich visual-semantic representations through all-to-all self-attention among patch tokens. However, this design implicitly assumes that direct pairwise patch interactions are necessary for effective representation learning. In this work, we challenge this assumption and show that representations supporting both global recognition and dense prediction can be learned without direct patch-to-patch interaction. We propose VECA (Visual Elastic-Core Attention), a vision transformer with core-periphery structured attention mediated by a small set of learned cores. Patch tokens exchange global information exclusively through these cores, while the full set of dense patches are preserved and iteratively updated across layers. This reduces attention complexity from O(N 2) to O(N), linear in the number of patches N for a fixed core budget C. Unlike prior latent-token cross-attention architectures, VECA facilitates sparse global communication without compressing the spatial representation itself. Nested training along the core axis further enables a single model to elastically trade off computation and accuracy at inference time without retraining. Across image classification and dense prediction tasks, VECA remains competitive with full-attention backbones and outperforms the evaluated linear-complexity alternatives on most benchmarks. Moreover, without explicit supervision, these cores develop semantically organized structures that support object-label transfer across video frames. These results show that effective visual representations can be learned without direct all-to-all patch interaction.
Sources
- How Much Position Information Do Convolutional Neural Networks Encode?
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- How Do Vision Transformers Work?
- Contrastive Learning Rivals Masked Image Modeling in Fine-tuning via Feature Distillation
- Theia: Distilling Diverse Vision Foundation Models for Robot Learning
- DINOv3
- Triangular Dropout: Variable Network Width without Retraining
- Distributional Principal Autoencoders
- A PCA-like Autoencoder
- Slimmable Neural Networks
- Once-for-All: Train One Network and Specialize it for Efficient Deployment
- SortedNet: A Scalable and Generalized Framework for Training Modular Deep Neural Networks
- Flextron: Many-in-One Flexible Large Language Model
- Adaptive Computation Time for Recurrent Neural Networks
- Universal Transformers
- Accelerating Optimization via Differentiable Stopping Time
- LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation
- Token Merging: Your ViT But Faster
- ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inference
- Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models