HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers".
Jane: The paper was written by Andy Li, Aiden Durrant, Milan Markovic and Georgios Leontidis from University of Aberdeen and University of East Anglia and UiT The Arctic University of Norway.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone! Today we're diving into a paper that's got a mouthful of a title — "HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers." Jane, what's your first read on this one?
Jane: Tom, I'm genuinely excited about this. Vision transformers are those massive AI models that power image recognition, but they're so heavy they barely fit on phones or edge devices. This paper is basically about making them lean without losing their brains.
Tom: And "HiAP" — that's a clever acronym, right? Hierarchical Auto-Pruning. The "hierarchical" part is what caught my eye, because most pruning methods work at one level only. This one works at multiple levels at once.
Jane: Exactly. Think of it like renovating a house. Some methods only remove furniture — that's fine-grained pruning. Others knock down entire walls — that's coarse pruning. HiAP does both simultaneously, deciding which walls to remove and which furniture to keep, all in one go.
Tom: And the "auto" part is huge. It's not using some hand-crafted rule to decide what to prune. The model learns it during training. That's a big philosophical shift from older approaches.
Jane: Right, and the authors are from Aberdeen, East Anglia, and UiT in Norway — a solid European collaboration. They're tackling a real deployment problem, not just chasing benchmarks.
Tom: So the title tells us it's automatic, it's multi-granular, and it's stochastic — meaning there's randomness baked into the learning process. That randomness is actually a feature, not a bug. It helps the model explore different architectures.
Jane: And the "framework" part means it's not just a trick for one model — it's a general approach you can apply to different vision transformers. That's what makes it exciting for the field.
Tom: I want to get into the weeds on how it actually works, but first — Jane, why should a regular person care about pruning vision transformers?
Jane: Because these models are everywhere. Your phone's camera, medical imaging, self-driving cars. If we can make them smaller and faster while keeping accuracy, that means better AI on devices that don't have a supercomputer behind them.
Tom: And that's the promise here. Alright, let's dig into the actual method next — how does HiAP decide what to cut?
Summary: Jane: So we're back with "HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers," and Tom, I want to unpack the core idea because it's genuinely clever.
Tom: Please do, because the abstract threw a lot at us — Gumbel-Sigmoid gates, MACs, budget-aware learning. Break it down for our listeners.
Jane: Okay, imagine you're packing a suitcase for a trip. You have a weight limit — that's the compute budget. Old methods would say "remove thirty percent of everything" uniformly. HiAP instead learns which items to leave behind by trying different combinations during packing, and it gets feedback on whether the suitcase still works.
Tom: And the "Gumbel-Sigmoid gates" — those are the little switches that decide keep or prune, right?
Jane: Exactly. Each attention head, each block, each neuron gets a learnable switch. During training, these switches are "soft" — they can be partially on, partially off. As training progresses, they harden into firm on-off decisions. It's like slowly turning a dimmer switch into a regular light switch.
Tom: And the beauty is that the model learns which switches to flip based on the actual task — not some arbitrary importance score. That's the "auto" in auto-pruning.
Jane: Right. And they've got this clever cost model. They calculate exactly how many multiply-accumulate operations — MACs — each structure costs. Then they add that cost to the training loss. So the model is simultaneously learning to be accurate and learning to be cheap.
Tom: They tested it on ImageNet with DeiT-Small, which is a standard vision transformer. They got it down from four point six Giga-MACs to about three point one — that's a third less compute — while keeping accuracy within half a percent of the original. That's impressive.
Jane: And here's the kicker — they did it in a single training phase. No separate search step, no fine-tuning afterwards. The model discovers its own compact architecture while learning the task. That's a huge simplification over methods that need multiple stages.
Tom: So it's not just about the final numbers — it's about how you get there. The process itself is simpler and more elegant.
Jane: And that matters for reproducibility. Multi-stage pipelines are notoriously hard to get right. This is one loss function, one training run, done.
Tom: I'm curious about what they actually discovered — like, which parts of the network got pruned. That's our next segment.
Improvements: Tom: Welcome back. We're still on "HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers," and I want to talk about what the model actually learned to prune. Jane, this is where it gets fascinating.
Jane: Tom, the most striking finding is that the model didn't prune uniformly. It learned a heterogeneous strategy. For example, it completely removed the final feed-forward block in the last transformer layer. That's depth pruning — it just bypasses that entire block.
Tom: And it also varied the number of attention heads per layer. Some layers kept more heads, some fewer. It's not a one-size-fits-all approach.
Jane: Right. And within the heads that survived, it also pruned individual dimensions. So you have this three-level hierarchy — blocks, heads, and dimensions — all being pruned simultaneously, each at different rates across layers.
Tom: That's the "multi-granular" part of the title really showing off. But here's what I find interesting — they also studied what happens if you only prune at one level. Like, what if you only prune macro-structures?
Jane: They ran those experiments, and the results are telling. If you only prune heads and blocks, you get an unbalanced network — the MLP neurons stay bloated. If you only prune micro-structures, you keep all heads but starve the MLPs. Neither is optimal.
Tom: So the magic is in the balance. They found a two:one ratio between macro and micro penalties worked best. That's a practical insight — if you're building a pruning system, you need both levels working together.
Jane: And they validated it on CIFAR-ten too, comparing against standard heuristics like L1-norm ranking. HiAP beat those baselines at both moderate and aggressive compression levels.
Tom: But the real proof is on hardware, right? Did they actually measure speedups?
Jane: They did. On a single GPU, the pruned model ran about one point four four times faster — from five point five seven milliseconds to three point eight six milliseconds per inference. That's a real, measurable improvement, not just theoretical FLOPs savings.
Tom: So it's not just a paper exercise. These pruned models actually run faster on real hardware. That's the kind of result that gets engineers excited.
Jane: And it's because they physically remove the structures — they don't leave soft masks that need special sparse kernels. The output is a dense, smaller network that runs natively.
Tom: That's a huge practical advantage. Alright, let's wrap up with our final thoughts on the impact of this work.
Conclusion: Tom: And we're back for our final segment on "HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers." Jane, give us the big-picture takeaway.
Jane: Tom, this paper represents a shift in how we think about model compression. Instead of pruning as a separate post-processing step, HiAP makes it part of the learning itself. The model learns its own compact architecture while learning the task — one phase, one loss function, done.
Tom: And the results speak for themselves. Competitive accuracy at a third less compute on ImageNet, real speedups on hardware, and a method that's simpler than the multi-stage pipelines that dominate the field.
Jane: The authors acknowledge a limitation, though — they optimize for MACs, not actual latency or energy. Those don't always align perfectly, depending on the hardware. But that's a natural next step.
Tom: What excites me is the potential for combination. They mention composing HiAP with token pruning, quantization, and compiler optimizations. If you stack all those, you could get enormous compression.
Jane: And that's the path to running these powerful models on phones, cameras, and medical devices. The environmental impact matters too — smaller models use less energy, which is better for the planet.
Tom: So as we say goodbye to this paper, what's the one thing you want listeners to remember?
Jane: That pruning doesn't have to be a complicated, multi-stage engineering ordeal. With the right formulation, you can let the model figure out its own efficient architecture — and it does a remarkably good job.
Tom: And that's a beautiful idea — giving the model agency over its own size. Thanks for joining us, everyone. Next time, we'll be looking at another exciting paper from the arXiv. Until then, keep learning!
Jane: And keep pruning! See you next time.
Andy Li, Aiden Durrant, Milan Markovic, Georgios Leontidis
University of Aberdeen · University of East Anglia · UiT The Arctic University of Norway
cs.CV, cs.LG
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: V1:14 pages, 9 figures, 3 Tables V2:different layout, more ablations
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 76/100
Key concepts
- HiAP
- Hierarchical Auto-Pruning, a framework that automatically prunes Vision Transformers. It works at multiple levels—like blocks, attention heads, and dimensions—simultaneously to create a compact model.
- Multi-Granular Pruning
- A technique where pruning is applied across different levels of the model structure at once. This allows the system to decide which parts of the network to remove based on their role, such as removing entire blocks or individual attention heads.
- Gumbel-Sigmoid Gates
- These are learnable switches used in HiAP that decide whether a part of the model should be kept or pruned. They start as 'soft' during training and harden into firm on-off decisions, allowing the model to learn which structures are important for the task.
- MACs (Multiply-Accumulate Operations)
- A measure used in HiAP to calculate the computational cost of model structures. The framework adds this cost to the training loss, forcing the model to learn both accuracy and computational efficiency simultaneously.
Terminology
Summary
Affiliations: University of Aberdeen, University of East Anglia, UiT The Arctic University of Norway
Publication: arXiv:2603.12222v2 [cs.CV], 28 Jun 2026
Vision Transformers (ViTs) require significant computational resources and memory bandwidth, severely limiting their deployment on resource-constrained hardware. The paper identifies two recurring limitations in existing structured pruning methods for ViTs:
First, single-granularity pruning leaves one bottleneck unaddressed. Methods that prune only micro-structures (e.g., intra-head dimensions) reduce FLOPs but preserve the network's overall depth and the number of attention heads. The paper notes that latency and energy on modern hardware are often dominated not by compute but by memory bandwidth: transferring data between DRAM and SRAM can cost orders of magnitude more energy than the corresponding computation.
ViTs are particularly exposed to this cost since preserving every layer and materializing each N × N attention map requires repeated weight loads and large intermediate transfers.
Conversely, methods that exclusively prune macro structures (e.g., entire heads, blocks) often risk more performance degradation as they lose the network's representation capacity to a greater extent.
Second, multi-stage pipelines are complex and hard to reproduce. Many strong methods decompose pruning into separate searching, importance ranking, post-hoc thresholding, and retraining stages, each adding its own design choices, hyperparameters, and auxiliary machinery (graph evaluations, solvers, scheduled mask updates etc).
HiAP casts pruning as a single-shot, budget-aware learning problem solved in one end-to-end phase.
Rather than hand-designing per-layer sparsity ratios or running a separate search-and-retrain pipeline, HiAP specifies a global compute budget and learns the model weights together with a per-layer, per-granularity sparsity allocation that meets it.
HiAP introduces two distinct levels of gating inside each Transformer block:
Macro-Level Pruning: Macro-gates control coarse units including entire attention heads (denoted as g l,h) and FFN blocks (denoted as b l). When g l,h = 0, the h-th attention head in layer l is completely bypassed. When b l = 0, the entire FFN block in layer l is removed.
Micro-Level Pruning: Micro-gates operate within active macro-structures, dynamically pruning internal matrix dimensions. For an active attention head, gates d l,h ∈ 0,1 D h prune the value-path dimensions. For an active FFN block, gates c l ∈ 0,1 D ffn prune the intermediate hidden neurons.
The paper states: "This hierarchy is central to our method, and it allows the network to independently decide whether to drop a coarse structure entirely to save substantial computational overhead, or merely narrow its width to preserve representation."
HiAP formulates an analytical, differentiable accounting of the network's Multiply-Accumulate operations (MACs).
The computational cost is decomposed into a static overhead C const (e.g., patch embedding, layer normalization) and dynamic costs governed by the gates. Three specific marginal cost constants are isolated:
-
C1 = 2ND(3D h) + 2N2D h: The macro-overhead for computing the dense Q, K, V projections and the attention map QK T for a single head.
-
C2 = 2ND + 2N2: The micro-cost of computing the attention output for a single surviving value dimension j.
-
C3 = 4ND: The micro-cost of computing a single surviving intermediate neuron in the FFN block.
The total prunable computational cost is linearly decomposed as:
E[C(G)] = Σ l=1 L Σ h=1 H (C1 · E[g l,h] + C2 Σ j=1 D h E[g l,h · d l,h,j]) + Σ l=1 L Σ k=1 D ffn C3 · E[b l · c l,k]
The paper emphasizes: "This linear decomposition is crucial - it allows the optimization process to cleanly attribute hardware penalties to individual structures, explicitly penalizing the network for keeping empty attention heads open, while allowing FFN blocks to scale costs purely by their active neuron count."
Each gate is parameterized by a learnable gate logit α. During the forward pass, a continuous gate value ẑ ∈ (0,1) is sampled by adding logistic noise ε and applying a sigmoid with temperature τ:
ẑ = σ((α + ε)/τ)
The total optimization objective combines the primary task loss, cost-aware structural regularization, and architectural feasibility penalties:
L total = L task + λ macro L macro + λ micro L micro + L feasibility
For the task loss, the paper uses a combination of standard cross-entropy and Knowledge Distillation (KD)
with a pre-trained, dense teacher model provid[ing] soft targets, which are crucial for maintaining the representations of the student network as its capacity is dynamically pruned.
The macro and micro computational costs are penalized separately, controlled by independent hyperparameters λ macro and λ micro. This decoupling allows us to explicitly manage the trade-off between coarse and fine-grained sparsity.
A feasibility penalty L feasibility prevents structural collapse by enforcing minimum retention quotas using a ReLU threshold: L f,head = Σ l ReLU(k min − Σ h g l,h)2 penalizes layers whose active head count falls below a minimum k min. Analogous terms apply minimum-ratio thresholds to surviving attention dimensions (γ attn) within each active head and to FFN neurons (γ ffn) within each active block.
HiAP eliminates the two-phase pipeline by unifying search and training into a single continuous phase. The dense ViT is trained end-to-end alongside the gate logits. The Gumbel-Sigmoid temperature τ is gradually annealed from an initial value τ0 to a near-zero minimum τ min.
The paper describes the dynamic co-adaptation process: "In the early epochs (high τ), the gates behave as stochastic dropout mechanisms, continuously forcing the surviving network weights to learn robust, distributed representations. As training progresses and τ decays, the gate distributions sharpen, naturally converging toward deterministic binary decisions. Because the weights are co-adapted to this gradually hardening topology, the network seamlessly transitions into its final sparse state without the catastrophic 'gradient shock' typically associated with sudden structural pruning."
Upon completion of training, gates are deterministically hardened using a simple probability threshold (ẑ > 0.5). "The dimensions of head value matrices are physically truncated according to the micro-gates, and entirely pruned heads and FFN blocks are deleted. This yields a natively fast, physically compressed Vision Transformer ready for immediate inference without the need for secondary fine-tuning."
The paper presents three theoretical propositions:
Lemma 1 (Expressivity: strict superset): Let A head be the set of architectures reachable by head/block (macro) gates only, and let A hiap be those additionally using micro-gates over attention dimensions or FFN neurons. If any layer has D h > 1 or D ffn > 1, then A head ⊊ A hiap.
Proposition 1 (Budget linearity): Under the accounting convention, there exist nonnegative weights w i for each prunable unit such that the expected prunable cost decomposes linearly as E[C] = Σ i w i E[z i], where z i ∈ g l,h, g l,hd l,h,j, b l c l,k. No independence assumptions are required
because by linearity of expectation, E[C] = Σ i w i E[z i] holds regardless of correlations among gates.
Proposition 2 (Soft-to-hard budget alignment): As τ → 0 anneals the Gumbel-Sigmoid gates to near-binary samples, "the differentiable expected cost E[C] converges to the discrete cost as variance collapses, using a fixed threshold θ = 0.5 natively finds a sub-network such that f(0.5) − C target ≤ ε for a small tolerance ε > 0, subject to the discrete step size of widths."
Datasets: Primary evaluations on ImageNet using DeiT-Small (dense baseline 4.6G MACs). Controlled ablations and latency profiling on CIFAR-10 with a custom 6-layer ViT-Tiny variant (embedding dimension 192, 3 heads, 174M MACs).
Implementation Details: For ImageNet, training runs 300 epochs using AdamW optimizer with learning rate 5 × 10−5 and global batch size 256. The Gumbel-Sigmoid temperature τ is annealed exponentially from 2.0 to 0.5. A dense, pre-trained RegNetY-160 serves as the teacher model (α KD = 0.7, T = 4.0).
At 3.1G MACs (approximately 33% reduction), HiAP achieves 79.33% Top-1 accuracy (baseline 79.75%), a drop of only −0.42%. At 2.9G MACs, it achieves 78.64% (−1.11% vs. dense). The paper notes: HiAP achieves these competitive trade-offs through an elegantly simple, unified mechanism
compared to methods like GOHSP and ViT-Slim that rely on highly complex, heavily engineered pipelines
requiring auxiliary graph evaluations, iterative importance ranking, and multi-stage masking.
The paper analyzes the evolution of structural gates during training on ImageNet:
Early Macro-Level Reductions: HiAP prioritizes the removal of coarse macro-structures early in the training process, as these provide the largest immediate reductions in the computational penalty L macro.
Within the first 10 epochs, the network reduces active attention heads from 6 to an average of 2-4 per layer. Most notably, the algorithm consistently identifies the FFN block in the final transformer layer (L = 12) as entirely redundant, permanently closing its macro-gate (b12 = 0).
Micro-Level Fine-Tuning: As macro-gates stabilize, the network transitions to exploiting micro-sparsity. Earlier layers maintain nearly full capacity (1400 active neurons out of 1536), while deeper layers are compressed much more aggressively (1200 active neurons).
Intra-head dimensions are often reduced from 64 down to 32 or fewer dimensions per head.
Validation of the Decoupled Loss: Because empty structural overhead (C1) is heavily penalized, the network learns to drop entire heads rather than keeping all 6 heads open with tiny, inefficient dimensions.
The paper evaluates penalty ratios spanning macro-only, 5:1, 2:1, 1.5:1, and micro-only configurations. The empirical data suggests that a balanced 2:1 ratio offers the most favorable accuracy-cost balance for DeiT-Small, enforcing a stable, hierarchical pruning trajectory between the macro and micro structures.
With aggressive λ macro (macro-only), the network shows a very high reduction in the remaining heads
and bypasses several blocks entirely,
but the MLP blocks, however, retain nearly 100% of their neurons, producing an inefficient and unbalanced distribution of the remaining MAC budget.
With aggressive λ micro (micro-only), the network retains all heads in every layer, and many MLP blocks are effectively pruned, although some recover later in training.
The paper concludes: Macro-only and micro-only each yield unbalanced allocations (bloated MLPs, or unstable indirect block elimination), whereas only the coupled penalty distributes sparsity cleanly across both granularities.
HiAP is compared against l1-norm importance ranking (targeting FFN neurons) and a Uniform-Ratio baseline. At the 33% budget, HiAP achieves 87.56% accuracy versus 87.15% for l1-structured and 86.63% for uniform-ratio. At approximately 50% reduction, HiAP achieves 87.25% versus 86.80% for l1-structured. The paper notes: because the network retains critical feature pathways, yielding a +0.93% accuracy improvement over the uniform baseline at the 33% budget.
Throughput and Hardware Efficiency: "When profiling the 33.1% pruned model (batch size 1, 50 inference runs), measured latency improves from 5.57 ms to 3.86 ms on a single GPU. This represents a ≈1.44× inference speedup, confirming that the discovered sub-networks do not rely on sparse convolution engines to realize their efficiency."
The paper lists three main contributions:
-
Unifying macro-level (heads, FFN blocks) and micro-level (intra-head dimensions, FFN neurons) structured pruning into a single differentiable phase, where weights and architecture co-adapt without a separate search or retraining stage.
-
Introducing a linear, differentiable decomposition of expected MACs into pre-structure marginal costs, enabling a single budget-aware objective to allocate sparsities jointly across all granularities without per-layer ratios, proxy ranking metrics, or offline solvers.
-
On ImageNet-1K with DeiT-small, achieving accuracy competitive with substantially more complex multi-stage pipelines at comparable compute, with physically extracted sub-networks delivering measurable speedups on real hardware.
The paper acknowledges: our method's objective has been to optimize expected MACs, not calibrated latency/energy, so realized speedups can vary with hardware/kernels.
Future directions include closing the MACs-to-latency gap with platform-calibrated latency/energy signals, and composing HiAP with token pruning, quantization, and compiler-level optimizations to expand beyond classification.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and the resulting capabilities:
1. Unified Multi-Granularity Pruning in a Single Phase
-
Replace multi-stage pipelines (search → rank → threshold → fine-tune) with a single end-to-end training loop that jointly optimizes weights and architecture gates
-
Implement hierarchical Gumbel-Sigmoid gates at both macro (heads, FFN blocks) and micro (intra-head dimensions, FFN neurons) levels simultaneously
-
Use temperature annealing (τ: 2.0 → 0.5) to gradually harden gates from stochastic to deterministic without gradient shock
2. Differentiable Analytical MAC Cost Modeling
-
Replace heuristic importance rankings or proxy metrics with an exact, linear decomposition of expected MACs into per-structure marginal costs (C1, C2, C3)
-
Decouple macro and micro cost penalties (λ macro, λ micro) to independently control coarse vs. fine-grained sparsity
-
Add feasibility penalties (ReLU-thresholded retention quotas) to prevent layer collapse during optimization
3. Autonomous Heterogeneous Architecture Discovery
-
Eliminate manual per-layer sparsity ratios; let the network learn its own heterogeneous allocation (e.g., pruning final FFN block entirely while varying head counts and widths across depth)
-
Enable emergent behaviors like dropping entire heads rather than keeping them with tiny, inefficient dimensions
-
Co-adapt weights and structure throughout training, avoiding the need for separate fine-tuning
4. Physical Sub-Network Extraction Without Post-Processing
-
Harden gates via simple probability threshold (ẑ > 0.5) at convergence
-
Physically delete pruned heads, blocks, dimensions, and neurons—no sparse convolution engines required
-
Deliver immediate hardware speedups (measured 1.44× on GPU) without secondary fine-tuning
For Vision Transformer Deployment:
-
Automatically compress a DeiT-Small from 4.6G to 3.1G MACs (33% reduction) while maintaining competitive accuracy (79.33% Top-1 on ImageNet) in a single training run
-
Discover architectures that reduce both memory bandwidth (by pruning blocks/heads) and compute FLOPs (by pruning micro-structures) simultaneously
-
Produce deployable, dense sub-networks that run natively on standard hardware without specialized sparse kernels
For Resource-Constrained Environments:
-
Meet strict computational budgets (e.g., 3.1G, 2.9G MACs) with a single hyperparameter (budget coefficient) rather than complex per-layer tuning
-
Achieve measurable latency improvements (5.57ms → 3.86ms on GPU) from physically extracted models
-
Maintain accuracy better than uniform-ratio or l1-norm heuristics at both moderate (33%) and aggressive (50%) compression levels
For Research and Development:
-
Eliminate the need for auxiliary solvers, graph evaluations, or iterative importance ranking—reducing pipeline complexity and reproducibility barriers
-
Provide a clean, theoretically grounded framework (with proofs for expressivity, budget linearity, and soft-to-hard alignment) that can be extended to other architectures
-
Enable controlled ablation of macro:micro penalty ratios (2:1 optimal for DeiT-Small) to understand sparsity allocation trade-offs
For Future Extensions:
-
The framework is complementary to token pruning, quantization, and compiler optimizations, allowing stacking of efficiency techniques
-
The linear cost decomposition can be adapted to platform-calibrated latency/energy models instead of raw MACs
-
The gating mechanism can be applied to other transformer variants (BERT, GPT, etc.) with minimal modification
Abstract
Vision Transformers require significant computational resources and memory bandwidth, severely limiting their deployment on resource-constraint hardware. Most structured pruning methods reduce theoretical cost effectively, yet they typically operate at a single structural granularity and depend on multi-stage pipelines with importance ranking, auxiliary solvers or post-hoc magnitude thresholding, followed by a separate fine-tuning phase to recover accuracy. We propose Hierarchical Auto-Pruning (HiAP), which casts ViT pruning as a single budget-aware learning problem and jointly allocates sparsity across four granularities in one end-to-end phase. HiAP introduces stochastic Gumbel-Sigmoid gates at macro level (attention heads and FFN blocks) and micro level (intra-head dimensions and FFN neurons), and optimizes them against the task loss together with an analytical MAC cost term. The budget coefficient steers the network to a target compute level while the gates gradually harden into a dense, smaller sub-network at convergence. It does not require importance heuristics, ranking metrics, auxillary solvers or secondary fine-tuning. On ImageNet with DeiT small, HiAP automatically discovers hetergenous architectures, pruning depths, heads, and width by different amount across layers, and reaches competitive accuracies against substantially more complex pruning pipelines at comparable compute from a single training run.
Sources
- Conditional Computation in Neural Networks for faster models
- Token Merging: Your ViT But Faster
- ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware
- Learning Sparse Neural Networks through $L_0$ Regularization
- MDP: Multidimensional Vision Model Pruning with Latency Constraint
- DBP: Discrimination Based Block-Level Pruning for Deep Model Acceleration
- Layer Pruning via Fusible Residual Convolutional Block for Deep Neural Networks
- AutoSlim: Towards One-Shot Architecture Search for Channel Numbers
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Reducing Transformer Depth on Demand with Structured Dropout
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations
- Unified Visual Transformer Compression
- Vision Transformer Pruning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models