Exploring Sparsity for Parameter Efficient Fine Tuning Using Wavelets for Vision

arXiv:2505.12532 · cs.CV, cs.AI, cs.LG, eess.IV, eess.SP · Submitted 2026-08-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Exploring Sparsity for Parameter Efficient Fine Tuning Using Wavelets for Vision".

Jane: The paper was written by Ahmet Bilican, M. Akın Yılmaz, A. Murat Tekalp and R. Gökberk Cinbiş from ETH Zürich and Codeway AI Research and Koç University and Middle East Technical University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everybody. Today we are digging into a paper that has been making waves in the AI community, and it's called "Exploring Sparsity for Parameter Efficient Fine Tuning Using Wavelets for Vision." Jane, I have to say, just reading that title got me excited.

Jane: Tom, I'm glad you're excited, because this one is genuinely clever. The basic problem is that these huge models like Stable Diffusion are amazing, but fine-tuning them for a specific task usually costs a fortune in compute and memory. This paper is about doing that fine-tuning with way fewer trainable parameters.

Tom: Right, and the trick is in the title. "Sparsity" means you're only updating a tiny fraction of the weights, and "wavelets" are the mathematical tool they use to decide which parts to update. It's like instead of repainting the whole house, you're just touching up the spots that actually need it.

Jane: Exactly. And the authors are from ETH Zurich, Koç University, and Middle East Technical University. They've got a nice mix of academic rigor and practical engineering. The paper is called WaveFT for short, and it's already been added to the Hugging Face PEFT library, which is a big deal.

Tom: Hugging Face is basically the standard library for this kind of work, so getting included there means people can actually use it today. Jane, for our listeners who aren't deep in the weeds, why does this matter for the real world?

Jane: Think about a company that wants to personalize an image generation model for their product catalog. With full fine-tuning, you'd need a massive GPU cluster. With WaveFT, you can do it with a fraction of the resources, and the results are actually better than the current standard method, LoRA.

Tom: Better results with less compute. That's the dream. And the paper shows it works across image generation, image classification, and even language tasks. We're going to spend the rest of the show unpacking how they pulled that off.

Jane: And trust me, the "how" involves some beautiful math about wavelets and gradient coverage. We'll get into that in the next segment.

Tom: Stick around, folks. This one's a goodie.

Summary: Tom: So we've set the stage. Now let's talk about what this paper actually does. The core idea in "Exploring Sparsity for Parameter Efficient Fine Tuning Using Wavelets for Vision" is to take the weight update matrix, transform it into the wavelet domain, and then only train a sparse set of coefficients there.

Jane: Right, and I want to make sure everyone gets why wavelets are special. Imagine you have a picture, and you want to compress it. A Fourier transform looks at the whole image globally, like a single lens. A wavelet transform looks at it in small local patches, like a mosaic of tiles. That locality is the secret sauce here.

Tom: And why does locality matter for fine-tuning? Because when you're adapting a pretrained model, the gradients—the signals telling you what to change—tend to be sparse and localized. You're not changing everything; you're tweaking specific features.

Jane: Exactly. And the paper has this beautiful theoretical framework. They prove that sparse methods like WaveFT produce high-rank updates, which means they can represent much more complex changes than LoRA's low-rank updates. LoRA is stuck in a low-dimensional subspace, like a car that can only turn left.

Tom: And that's why the results are so striking. On the DreamBooth benchmark for personalized image generation, WaveFT gets a DINO similarity of zero point four nine five, compared to LoRA's zero point four six three. That's a significant jump in subject fidelity—the generated images actually look more like the subject you're trying to personalize.

Jane: And they do it with fewer parameters. WaveFT at 72K parameters beats LoRA at 581K parameters on image classification. That's an eight times reduction in trainable parameters while getting better accuracy. Tom, that's not an incremental improvement; that's a paradigm shift.

Tom: It really is. And they also show that WaveFT produces more diverse outputs, which is huge for creative applications. The LPIPS diversity score is zero point three four eight for WaveFT versus zero point three zero nine for LoRA. So not only are the images more faithful, they're also more varied.

Jane: And that diversity comes from the high-rank updates. LoRA's low-rank constraint limits the variety of changes it can make. WaveFT escapes that bottleneck, so the model can explore a much richer space of possible adaptations.

Tom: So the summary is: wavelets give you local awareness, sparsity gives you efficiency, and high-rank updates give you expressiveness. It's a triple threat.

Jane: And we haven't even talked about the gradient coverage framework yet, which explains why wavelets beat both direct sparsity and Fourier methods. That's coming up.

Improvements: Tom: Welcome back. We've covered the basics, but now I want to dig into the improvements this paper suggests over existing methods. Jane, what's the key differentiator here?

Jane: It's all about what they call "gradient coverage." When you're fine-tuning, only a fraction of the weight positions have useful gradients. The paper defines this as gradient sparsity, and it varies by task. In vision tasks, gradients are often very sparse, like ten percent of positions.

Tom: And that's where wavelets shine. Because a wavelet coefficient aggregates information from a local neighborhood, it has a much better chance of catching a useful gradient signal. The paper shows that at ten percent gradient sparsity, WaveFT has three point four times better coverage than direct weight sparsity like SHiRA.

Jane: Right. SHiRA is the baseline that just picks random individual weights to update. Each weight only sees its own gradient. Wavelets see a whole patch. So if a gradient is nearby, the wavelet coefficient catches it, even if it's not exactly at that position.

Tom: But wait, why not just use Fourier transforms? They have global coverage, so they'd catch everything, right?

Jane: That's the clever part. Fourier transforms have global receptive fields, which means each coefficient aggregates gradients from the entire matrix. When gradients are sparse and localized, those global signals interfere destructively. They cancel each other out. The paper shows FourierFT absolutely fails on personalized image generation, with a DINO score of zero point two one five versus WaveFT's zero point four nine five.

Tom: So it's not just about coverage; it's about the right kind of coverage. Wavelets are semi-local, which means they aggregate coherent gradients from a neighborhood without the global interference problem. It's like the Goldilocks zone of gradient aggregation.

Jane: And the paper validates this with a really clever experiment. They randomly permute the input tokens to the attention layers, destroying any spatial structure. WaveFT still outperforms SHiRA. That proves the advantage comes from gradient coverage, not from exploiting spatial locality.

Tom: That's a beautiful ablation. It isolates the mechanism. And they also show that as you increase the parameter budget, the gap between WaveFT and SHiRA narrows. At higher budgets, SHiRA's point-wise updates eventually achieve enough coverage, so the wavelet advantage diminishes.

Jane: Which tells us something practical: WaveFT is most valuable in the parameter-constrained regime, which is exactly where you need it. When you're trying to squeeze every bit of efficiency out of a small budget, wavelets give you the best return.

Tom: And they also show that on NLP tasks, the advantage shrinks because gradients are denser. So this isn't a universal win; it's a targeted improvement for the right conditions.

Jane: Right, and knowing when a method works is just as important as knowing that it works. That's what makes this paper so solid.

Conclusion: Tom: Alright, we've covered a lot of ground on "Exploring Sparsity for Parameter Efficient Fine Tuning Using Wavelets for Vision." Let's wrap this up. Jane, what's the one-sentence takeaway for our listeners?

Jane: WaveFT shows that by learning sparse updates in the wavelet domain, you can fine-tune massive models with far fewer parameters and get better results than the current standard methods, especially for vision tasks.

Tom: And it's not just theory. The paper has rigorous experiments across text-to-image generation with SDXL, image classification with ViT-Base, and language understanding with GLUE. The results are consistent and the theoretical framework explains why.

Jane: I also love that they've open-sourced it and got it into the Hugging Face PEFT library. That means it's not just a paper; it's a tool that practitioners can use today.

Tom: And the implications go beyond just saving compute. Better parameter efficiency means smaller organizations can fine-tune state-of-the-art models. It democratizes access to AI customization.

Jane: And the diversity improvement is huge for creative industries. If you're generating images for a brand, you want variety, not just the same image with different prompts. WaveFT delivers that.

Tom: Alright, let's give a round of applause to the authors for this one. It's a clever idea, rigorously validated, and practically useful. That's a rare combination.

Jane: Agreed. We'll be watching to see how this gets adopted in the community. Thanks for tuning in, and we'll see you for the next paper.

Tom: Take care, everyone.

Ahmet Bilican, M. Akın Yılmaz, A. Murat Tekalp, R. Gökberk Cinbiş

ETH Zürich · Codeway AI Research · Koç University · Middle East Technical University

cs.CV, cs.AI, cs.LG, eess.IV, eess.SP

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/Lightning-AI/torchmetrics

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 43/100

The gist: "We propose Wavelet Fine-Tuning (WaveFT), which learns a sparse set of p parameters in the wavelet domain representation of the weight update matrix ∆W.

Key concepts

Sparsity
In this context, sparsity means only updating a tiny fraction of the model's weights during fine-tuning instead of changing every single weight. This makes the process much more parameter-efficient.
Wavelets
Wavelets are mathematical tools that transform data into small local patches, similar to a mosaic. They are used here to decide which specific parts of the model's weights need updating based on localized information.
Gradient Coverage
This refers to how well the chosen method captures useful gradient signals during fine-tuning. Wavelets are effective because they aggregate information from a local neighborhood, allowing them to catch sparse and localized gradient signals better than direct sparsity methods like SHiRA.

Terminology

Summary

Summary

The paper introduces Wavelet Fine-Tuning (WaveFT), a novel Parameter-Efficient Fine-Tuning (PEFT) method that learns sparse updates in the wavelet domain of weight matrices. The authors state: "We propose Wavelet Fine-Tuning (WaveFT), which learns a sparse set of p parameters in the wavelet domain representation of the weight update matrix ∆W. These learned coefficients are transformed back to the weight domain via the Inverse Discrete Wavelet Transform (IDWT)."

The method addresses limitations of Low-Rank Adaptation (LoRA), which relies on an integer rank r ≥ 1 imposes two limitations: the minimum rank forces allocating more parameters than necessary for simple adaptations, and discrete rank increments prevent fine-grained control across layers. WaveFT enables parameter budgets well below LoRA’s minimum at r = 1 through its sparse parameterization governed by p.

The paper provides two complementary theoretical justifications for why the wavelet domain is preferable to direct weight-space sparsity (SHiRA). First, both WaveFT and SHiRA produce high-rank updates, escaping LoRA’s subspace bottleneck that confines modifications to an r-dimensional subspace. Second, wavelet coefficients have semi-local receptive fields that aggregate gradient information from spatially coherent neighborhoods, yielding better gradient coverage than weight-domain sparsity, without the destructive interference that plagues global Fourier bases.

The theoretical framework includes several key lemmas. Lemma 1 establishes that a sufficiently sparse random matrix ∆W (where sufficiently sparse implies the number of non-zero elements p is at least n(ln n + cn)) is highly likely to be full rank as matrix dimensions grow. Lemma 2 formalizes LoRA's Subspace Bottleneck: For ∆W = BA⊤ with B ∈ Rm×r, A ∈ Rn×r: 1. im(∆W) ⊆ span(cols of B); 2. ker(∆W) ⊇ ker(A⊤). Lemma 3 proves Block-Sparse Interpolation Capacity, showing that "if a sparse support pattern S is suitably structured relative to a set of k linearly independent inputs x(l) and desired outputs y (l), then a ∆W confined to this support S can perfectly interpolate these target transformations."

The paper introduces a gradient coverage framework with Definition 4 (Gradient Receptive Field): For parameter θi in coefficient matrix C, the gradient receptive field Ri ⊆ [m] × [n] is the set of weight positions whose gradients influence θi: Ri = (u, v): ∂Wuv /∂θi ̸= 0. The receptive field sizes differ by method: SHiRA (identity): Ri = 1; WaveFT (wavelet, filter size κ): Ri = O(κ2); FourierFT (Fourier): Ri = mn. Proposition 5 (Gradient Coverage) states: "When ρ is small and gradient positions are approximately uniformly distributed across the weight matrix, the probability that a parameter with receptive field size R receives gradient signal is approximately: P (receives gradient) ≈ 1 − (1 − ρ)R."

The framework explains task-dependent performance: "Vision personalization (sparse gradients): WaveFT > SHiRA > LoRA; NLP tasks (denser gradients): FourierFT > SHiRA ≈ WaveFT. The authors note that FourierFT achieves maximal coverage but global receptive fields create two critical problems": destructive interference from gradients from distant positions, mapped through complex phase factors, can cancel each other, and an information bottleneck because sparse FourierFT uses only p ≪ mn coefficients, so each trainable coefficient must therefore encode information aggregated from all positions, without sufficient degrees of freedom.

Experiments span three domains. For personalized text-to-image generation using SDXL on the DreamBooth benchmark, WaveFT achieves 0.495 DINO similarity versus LoRA’s 0.463 at equivalent parameter counts, while also improving output diversity (LPIPS: 0.348 vs 0.309). WaveFT also outperforms SHiRA (DINO: 0.495 vs 0.467, CLIP-I: 0.655 vs 0.645). FourierFT exhibits substantially degraded subject fidelity (DINO: 0.215), consistent with our theoretical prediction that global basis functions suffer from destructive interference.

For image classification on ViT-Base across eight datasets, WaveFT attains 78.29% average accuracy with only 72K parameters, outperforming LoRA (77.58% at 581K parameters) and FourierFT (77.75%). WaveFT shows particular strength on fine-grained classification tasks (StanfordCars: 48.12%, FGVC: 31.53%).

For language understanding on GLUE with RoBERTa-base, WaveFT and SHiRA perform comparably (83.9 vs 84.2 average), both slightly below FourierFT (85.0), consistent with the theoretical prediction that NLP fine-tuning involves denser, more distributed gradient patterns, reducing the locality advantage of wavelets.

Key ablation findings include: (1) Robustness to Input Permutation: Performance remained largely unchanged... WaveFT still outperforming SHiRA, demonstrating that WaveFT’s advantage arises from improved gradient coverage under sparsity, not from exploiting spatial locality; (2) WaveFT yields more stable and better results than SHiRA due to it being less sensitive to the random selection of p trainable locations across different seeds; (3) Various wavelet families (Coiflets, Daubechies, Symlets) yielded strong, comparable performance; (4) Zero-initialization of the p trainable parameters in the coefficient matrix C proved robust. Gaussian initialization performed drastically worse; (5) Allocating a fixed p to each layer outperformed allocating parameters proportionally to layer size; (6) The output scaling factor λ provides a controllable fidelity-alignment trade-off.

The paper concludes: WaveFT achieves state-of-the-art results among PEFT methods on vision tasks, with particular strength on fine-grained classification and personalized generation where gradients are sparse and localized. The authors also note that WaveFT has officially been included in the Hugging Face PEFT library.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:

Implementation: Add a WaveFT layer to the PEFT library that:

  • Applies 2D Discrete Wavelet Transform (DWT) to weight matrices

  • Selects p coefficients uniformly at random for training

  • Uses zero initialization (Gaussian initialization performs 38× worse)

  • Applies Inverse DWT (IDWT) to reconstruct the update matrix

  • Uses Daubechies 1 (Haar) wavelet as default (comparable performance to more complex wavelets at lower compute)

Resulting capability: Fine-tune models with 8× fewer parameters than LoRA r=1 (72K vs 581K for ViT-Base) while achieving higher accuracy (78.29% vs 77.58% average across 8 vision datasets).

Abstract

Efficiently adapting large pretrained models is critical under tight compute and memory budgets. While Parameter-Efficient Fine-Tuning (PEFT) methods like LoRA achieve efficiency through low-rank updates, their discrete rank constraint limits fine-grained parameter control and confines adaptations to low-dimensional subspaces. We propose Wavelet Fine-Tuning (WaveFT), which learns sparse updates in the wavelet domain of weight matrices, enabling fine-grained control over trainable parameters well below LoRA's minimum rank. Wavelet bases provide semi-local receptive fields that aggregate spatially coherent gradients, offering better coverage than direct weight sparsity (SHiRA) without the destructive interference of global Fourier bases (FourierFT). We provide theoretical analysis showing: (i) sparse methods achieve high-rank updates, avoiding LoRA's subspace bottleneck and enabling higher representational capacity, and (ii) a gradient coverage framework explaining when WaveFT is preferable. We perform experiments across text-to-image generation, image classification, and language understanding. WaveFT demonstrates state-of-the-art results among PEFT methods for vision tasks, where wavelets effectively capture sparse gradient structure through improved coverage, while performing comparably on NLP tasks. WaveFT has officially been included in the Hugging Face PEFT library (huggingface.co/docs/peft/en/package reference/waveft).

Sources

Related papers