HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers
summary
In short
The episode discusses HiAP, a framework for automatically pruning Vision Transformers to make them smaller and faster without losing accuracy. The hosts explain how HiAP uses learned gates and cost models to simultaneously learn an efficient architecture during training. They conclude that this method simplifies compression by making it part of the learning process.
Key concepts
- HiAP
- Hierarchical Auto-Pruning, a framework that automatically prunes Vision Transformers. It works at multiple levels—like blocks, attention heads, and dimensions—simultaneously to create a compact model.
- Multi-Granular Pruning
- A technique where pruning is applied across different levels of the model structure at once. This allows the system to decide which parts of the network to remove based on their role, such as removing entire blocks or individual attention heads.
- Gumbel-Sigmoid Gates
- These are learnable switches used in HiAP that decide whether a part of the model should be kept or pruned. They start as 'soft' during training and harden into firm on-off decisions, allowing the model to learn which structures are important for the task.
- MACs (Multiply-Accumulate Operations)
- A measure used in HiAP to calculate the computational cost of model structures. The framework adds this cost to the training loss, forcing the model to learn both accuracy and computational efficiency simultaneously.
Terminology used across episodes
This episode discusses
- HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers · Paper Radio
- Conditional Computation in Neural Networks for faster models
- Token Merging: Your ViT But Faster
- ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware
- Learning Sparse Neural Networks through L 0 Regularization
- MDP: Multidimensional Vision Model Pruning with Latency Constraint
- DBP: Discrimination Based Block-Level Pruning for Deep Model Acceleration
- Layer Pruning via Fusible Residual Convolutional Block for Deep Neural Networks
- AutoSlim: Towards One-Shot Architecture Search for Channel Numbers
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Reducing Transformer Depth on Demand with Structured Dropout
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations
- Unified Visual Transformer Compression
- Vision Transformer Pruning
The paper
HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers · Read on arXiv
Andy Li, Aiden Durrant, Milan Markovic, Georgios Leontidis
University of Aberdeen · University of East Anglia · UiT The Arctic University of Norway
Vision Transformers require significant computational resources and memory bandwidth, severely limiting their deployment on resource-constraint hardware. Most structured pruning methods reduce theoretical cost effectively, yet they typically operate at a single structural granularity and depend on multi-stage pipelines with importance ranking, auxiliary solvers or post-hoc magnitude thresholding, followed by a separate fine-tuning phase to recover accuracy. We propose Hierarchical Auto-Pruning (HiAP), which casts ViT pruning as a single budget-aware learning problem and jointly allocates sparsity across four granularities in one end-to-end phase. HiAP introduces stochastic Gumbel-Sigmoid gates at macro level (attention heads and FFN blocks) and micro level (intra-head dimensions and FFN neurons), and optimizes them against the task loss together with an analytical MAC cost term. The budget coefficient steers the network to a target compute level while the gates gradually harden into a dense, smaller sub-network at convergence. It does not require importance heuristics, ranking metrics, auxillary solvers or secondary fine-tuning. On ImageNet with DeiT small, HiAP automatically discovers hetergenous architectures, pruning depths, heads, and width by different amount across layers, and reaches competitive accuracies against substantially more complex pruning pipelines at comparable compute from a single training run.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers".
Jane: The paper was written by Andy Li, Aiden Durrant, Milan Markovic and Georgios Leontidis from University of Aberdeen and University of East Anglia and UiT The Arctic University of Norway.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone! Today we're diving into a paper that's got a mouthful of a title — "HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers." Jane, what's your first read on this one?
Jane: Tom, I'm genuinely excited about this. Vision transformers are those massive AI models that power image recognition, but they're so heavy they barely fit on phones or edge devices. This paper is basically about making them lean without losing their brains.
Tom: And "HiAP" — that's a clever acronym, right? Hierarchical Auto-Pruning. The "hierarchical" part is what caught my eye, because most pruning methods work at one level only. This one works at multiple levels at once.
Jane: Exactly. Think of it like renovating a house. Some methods only remove furniture — that's fine-grained pruning. Others knock down entire walls — that's coarse pruning. HiAP does both simultaneously, deciding which walls to remove and which furniture to keep, all in one go.
Tom: And the "auto" part is huge. It's not using some hand-crafted rule to decide what to prune. The model learns it during training. That's a big philosophical shift from older approaches.
Jane: Right, and the authors are from Aberdeen, East Anglia, and UiT in Norway — a solid European collaboration. They're tackling a real deployment problem, not just chasing benchmarks.
Tom: So the title tells us it's automatic, it's multi-granular, and it's stochastic — meaning there's randomness baked into the learning process. That randomness is actually a feature, not a bug. It helps the model explore different architectures.
Jane: And the "framework" part means it's not just a trick for one model — it's a general approach you can apply to different vision transformers. That's what makes it exciting for the field.
Tom: I want to get into the weeds on how it actually works, but first — Jane, why should a regular person care about pruning vision transformers?
Jane: Because these models are everywhere. Your phone's camera, medical imaging, self-driving cars. If we can make them smaller and faster while keeping accuracy, that means better AI on devices that don't have a supercomputer behind them.
Tom: And that's the promise here. Alright, let's dig into the actual method next — how does HiAP decide what to cut?
Summary: Jane: So we're back with "HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers," and Tom, I want to unpack the core idea because it's genuinely clever.
Tom: Please do, because the abstract threw a lot at us — Gumbel-Sigmoid gates, MACs, budget-aware learning. Break it down for our listeners.
Jane: Okay, imagine you're packing a suitcase for a trip. You have a weight limit — that's the compute budget. Old methods would say "remove thirty percent of everything" uniformly. HiAP instead learns which items to leave behind by trying different combinations during packing, and it gets feedback on whether the suitcase still works.
Tom: And the "Gumbel-Sigmoid gates" — those are the little switches that decide keep or prune, right?
Jane: Exactly. Each attention head, each block, each neuron gets a learnable switch. During training, these switches are "soft" — they can be partially on, partially off. As training progresses, they harden into firm on-off decisions. It's like slowly turning a dimmer switch into a regular light switch.
Tom: And the beauty is that the model learns which switches to flip based on the actual task — not some arbitrary importance score. That's the "auto" in auto-pruning.
Jane: Right. And they've got this clever cost model. They calculate exactly how many multiply-accumulate operations — MACs — each structure costs. Then they add that cost to the training loss. So the model is simultaneously learning to be accurate and learning to be cheap.
Tom: They tested it on ImageNet with DeiT-Small, which is a standard vision transformer. They got it down from four point six Giga-MACs to about three point one — that's a third less compute — while keeping accuracy within half a percent of the original. That's impressive.
Jane: And here's the kicker — they did it in a single training phase. No separate search step, no fine-tuning afterwards. The model discovers its own compact architecture while learning the task. That's a huge simplification over methods that need multiple stages.
Tom: So it's not just about the final numbers — it's about how you get there. The process itself is simpler and more elegant.
Jane: And that matters for reproducibility. Multi-stage pipelines are notoriously hard to get right. This is one loss function, one training run, done.
Tom: I'm curious about what they actually discovered — like, which parts of the network got pruned. That's our next segment.
Improvements: Tom: Welcome back. We're still on "HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers," and I want to talk about what the model actually learned to prune. Jane, this is where it gets fascinating.
Jane: Tom, the most striking finding is that the model didn't prune uniformly. It learned a heterogeneous strategy. For example, it completely removed the final feed-forward block in the last transformer layer. That's depth pruning — it just bypasses that entire block.
Tom: And it also varied the number of attention heads per layer. Some layers kept more heads, some fewer. It's not a one-size-fits-all approach.
Jane: Right. And within the heads that survived, it also pruned individual dimensions. So you have this three-level hierarchy — blocks, heads, and dimensions — all being pruned simultaneously, each at different rates across layers.
Tom: That's the "multi-granular" part of the title really showing off. But here's what I find interesting — they also studied what happens if you only prune at one level. Like, what if you only prune macro-structures?
Jane: They ran those experiments, and the results are telling. If you only prune heads and blocks, you get an unbalanced network — the MLP neurons stay bloated. If you only prune micro-structures, you keep all heads but starve the MLPs. Neither is optimal.
Tom: So the magic is in the balance. They found a two:one ratio between macro and micro penalties worked best. That's a practical insight — if you're building a pruning system, you need both levels working together.
Jane: And they validated it on CIFAR-ten too, comparing against standard heuristics like L1-norm ranking. HiAP beat those baselines at both moderate and aggressive compression levels.
Tom: But the real proof is on hardware, right? Did they actually measure speedups?
Jane: They did. On a single GPU, the pruned model ran about one point four four times faster — from five point five seven milliseconds to three point eight six milliseconds per inference. That's a real, measurable improvement, not just theoretical FLOPs savings.
Tom: So it's not just a paper exercise. These pruned models actually run faster on real hardware. That's the kind of result that gets engineers excited.
Jane: And it's because they physically remove the structures — they don't leave soft masks that need special sparse kernels. The output is a dense, smaller network that runs natively.
Tom: That's a huge practical advantage. Alright, let's wrap up with our final thoughts on the impact of this work.
Conclusion: Tom: And we're back for our final segment on "HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers." Jane, give us the big-picture takeaway.
Jane: Tom, this paper represents a shift in how we think about model compression. Instead of pruning as a separate post-processing step, HiAP makes it part of the learning itself. The model learns its own compact architecture while learning the task — one phase, one loss function, done.
Tom: And the results speak for themselves. Competitive accuracy at a third less compute on ImageNet, real speedups on hardware, and a method that's simpler than the multi-stage pipelines that dominate the field.
Jane: The authors acknowledge a limitation, though — they optimize for MACs, not actual latency or energy. Those don't always align perfectly, depending on the hardware. But that's a natural next step.
Tom: What excites me is the potential for combination. They mention composing HiAP with token pruning, quantization, and compiler optimizations. If you stack all those, you could get enormous compression.
Jane: And that's the path to running these powerful models on phones, cameras, and medical devices. The environmental impact matters too — smaller models use less energy, which is better for the planet.
Tom: So as we say goodbye to this paper, what's the one thing you want listeners to remember?
Jane: That pruning doesn't have to be a complicated, multi-stage engineering ordeal. With the right formulation, you can let the model figure out its own efficient architecture — and it does a remarkably good job.
Tom: And that's a beautiful idea — giving the model agency over its own size. Thanks for joining us, everyone. Next time, we'll be looking at another exciting paper from the arXiv. Until then, keep learning!
Jane: And keep pruning! See you next time.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language