Fisher-Guided Progressive Parameter Selection for Adaptive Fine-Tuning

arXiv:2606.10196 · cs.CV, cs.AI · Submitted 2026-06-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Fisher-Guided Progressive Parameter Selection for Adaptive Fine-Tuning".

Tom: FisherAdapTune is a Fisher-guided Adaptive Fine-Tuning framework designed to dynamically select which parameter groups should remain trainable during fine-tuning by tracking the temporal drift of their Fisher geometry.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Now we're moving on to the title and authors of this paper, "Fisher-Guided Progressive Parameter Selection for Adaptive Fine-Tuning." The title itself really tells you the essence of what they’re tackling here: guiding parameter selection using Fisher information.

Jane: It sounds quite technical, Tom; I hope you can explain what that actually means in plain English for our listeners who might not be deep into differential geometry.

Lu: Think of it this way: instead of picking a set of layers based on their position in the architecture, they are using the temporal evolution of the Fisher geometry to dynamically select which parameter groups should remain trainable.

Meng: So, it’s like having an intelligent system that watches how much each part of the model is still actively shaping its understanding as it learns something new.

Lalam: That sounds incredibly sophisticated, Lu; like giving the AI a way to self-assess its own learning progress on a granular level.

Tom: Precisely, Lu; they are using the Fisher Information Matrix to estimate parameter importance based on how sensitive the loss is to changes in those parameters during training.

Jane: That sensitivity is what we usually try to measure, but this framework takes that sensitivity and applies it over time across different parameter groups, which is what makes it unique.

Lu: They start from a PAC-Bayesian view of fine-tuning and decompose the generalization error bound into Fisher-weighted update costs to show that parameters whose curvature contribution has stabilized can be frozen.

Tom: That's a very specific theoretical underpinning; they aren't just guessing what to freeze; they are mathematically proving why freezing those specific components reduces the generalization error bound according to Equation (four) in the paper.

Jane: It grounds the selection in learning theory, which gives it a strong theoretical foundation that goes beyond just looking at empirical results on specific tasks.

Meng: From an engineering perspective, that theoretical grounding is reassuring because it suggests we aren't just tinkering with hyperparameters; we have a mathematical reason for our choices.

Lalam: It’s exciting because it moves us toward a more principled approach to AI development where the adaptation strategy is driven by the inherent geometry of the learning process itself.

Tom: So, to summarize, this paper proposes using Fisher-Guided Progressive Parameter Selection for Adaptive Fine-Tuning as a framework that dynamically selects parameter groups based on tracking temporal drift in Fisher geometry.

Jane: And it’s about showing that parameters whose influence has stabilized can be frozen to reduce the error bound without disrupting the remaining adaptation dynamics.

The paper's summary: Tom: Now we need to get into what they actually propose, so the paper summarizes FisherAdapTune as a task-aware fine-tuning framework that dynamically selects which parameter groups to update instead of fixing a subset before training.

Jane: So, in simple terms, this means the system assesses the learning progress iteration by iteration and decides on the fly whether to freeze or update specific parts of the model based on that assessment.

Lu: Exactly, Jane; it’s about moving away from those fixed architectural heuristics where you might freeze "top layers" just because that's a common practice in other papers.

Meng: That sounds like a significant operational improvement because we avoid those rigid rules and instead let the data dictate the required update budget dynamically.

Lalam: If this is true, it means our fine-tuning process becomes much more fluid, adapting its resources according to what the task demands at any given moment during training.

Tom: The core mechanism involves computing per-layer Fisher curvature estimates and then using those estimates to build probability distributions over parameter groups.

Jane: Then they compute the consecutive Fisher drift between these distributions, which is the key signal that tells them if a group is still changing significantly or has stabilized.

Lu: Under assumption three point seven about smooth parameter updates, this drift is approximated by delta t, F t, delta t, (six), which directly links the contribution change to the Fisher matrix evolution.

Tom: That equation is a powerful piece of math because it shows that the change in parameter contribution is directly tied to how much its local curvature is changing over time.

Jane: When the Fisher stabilizes, meaning F theta t, about F theta t-one that drift becomes negligible, which is exactly when they decide to freeze those parameters.

Meng: I’m still curious about the practical implementation details; how do we manage the computation of those per-layer Fisher estimates efficiently without it becoming a bottleneck during the actual training process?

Lalam: If we can manage that estimation cost effectively, then this framework could become a standard component in our production fine-tuning pipeline, making adaptation much more scalable and efficient.

Tom: The paper shows that to make this criterion actionable, they map the Fisher tensor at iteration t and layer into a probability distribution via log-normalization to capture structural drift in a scale-invariant way.

Jane: That mapping step is clever because it allows them to use the Jensen-Shannon divergence between these distributions, which gives them a symmetric and bounded metric for comparison.

Lu: Proposition three point eight establishes that this Jensen-Shannon divergence satisfies two key properties: one shows structural sensitivity, meaning it's zero only if the log-normalized Fisher distributions have not changed in shape.

Tom: So, to summarize the summary, FisherAdapTune uses these structural drift measures to progressively select parameter groups by freezing those whose contribution has stabilized according to a data-driven criterion.

Jane: This framework offers a principled mechanism for identifying parameters that actively reshape the model's local loss landscape while allowing stabilized groups to be frozen.

The paper's improvements: Tom: Let’s talk about the specific improvements this paper suggests, which are really centered around moving from static heuristics to dynamic, data-driven criteria for parameter selection.

Jane: So it moves beyond just picking layers based on fixed rules and instead uses the temporal evolution of the Fisher geometry to make that decision dynamically during training.

Lu: The primary improvement is that they use a scale-invariant Jensen-Shannon distance between consecutive Fisher distributions as the core metric to quantify structural drift, which is much more robust than previous methods.

Meng: That robustness is important because it means we aren't getting fooled by irrelevant changes in parameter magnitudes that don't actually reflect meaningful shifts in the model’s adaptation strategy.

Lalam: This seems like a huge step toward developing AI systems that are less susceptible to overfitting because they won't waste resources on redundant parameters that have already found their place.

Tom: Another improvement is the ability to differentiate between parameters that are actively reshaping the model’s local loss landscape—those with high Fisher drift—and those whose influence has stabilized, which is a direct measure of adaptation relevance.

Jane: So, this allows for a very targeted approach where we only keep updating what is currently driving the task adaptation forward, effectively controlling the computational budget on a per-parameter basis.

Lu: By tracking these curvature shifts as proxies for adaptation relevance, they establish a direct link between the Fisher dynamics and which parameters are meaningful to keep active during training.

Tom: This also leads to better in-distribution performance and robustness because by preventing over-specialization, the model retains essential task-dependent components while discarding redundant ones.

Jane: So, the improvement is that we can maintain high performance on unseen data because we’re not forcing every parameter to change just for the sake of being updated.

Conclusion: Tom: Alright team, let's wrap up with the conclusion of "Fisher-Guided Progressive Parameter Selection for Adaptive Fine-Tuning." Essentially, this paper confirms that using Fisher dynamics provides a principled way to select parameter groups during fine-tuning based on their temporal curvature contribution.

Jane: The main takeaway is that we can now stop relying on fixed rules and start using dynamic, task-aware criteria derived from the Fisher Information Matrix.

Lu: It’s a significant theoretical advancement because it provides a mathematically rigorous basis for identifying task-relevant parameters through the analysis of their loss sensitivity.

Meng: Operationally, this means we can potentially achieve better performance with fewer trainable parameters and lower computational overhead during adaptation phases, which is always something we want.

Lalam: I think the real impact is that it pushes our entire AI culture toward a more principled way of thinking about how models adapt and manage their resources dynamically.

Tom: It’s a solid piece of research that shows how to use the Fisher dynamics to guide parameter selection progressively during fine-tuning, which we can definitely start exploring.

Jane: It really gives us a strong tool for controlling our training process with more mathematical insight into what's actually happening inside the model.

Ghodsiyeh Rostami, Po-Han Chen, Mahdi S. Hosseini

Concordia University

cs.CV, cs.AI

Submitted: 2026-06-08

Updated: 2026-09-28

Importance score: 90/100

The gist: FisherAdapTune is a Fisher-guided Adaptive Fine-Tuning framework designed to dynamically select which parameter groups should remain trainable during fine-tuning by tracking the temporal drift of

Key concepts

FisherAdapTune
A task-aware fine-tuning framework that dynamically selects which parameter groups to update based on learning progress. It moves away from fixed architectural rules by assessing layer updates iteration by iteration.
Fisher geometry
The geometric structure related to the Fisher Information Matrix, which estimates parameter importance based on how sensitive the loss function is to changes in those parameters during training. Tracking its temporal drift helps determine if a parameter group has stabilized.
Jensen-Shannon divergence
A metric used to measure structural drift between consecutive Fisher distributions. It provides a symmetric and bounded way to compare parameter groups, indicating whether their shape or structural sensitivity has changed significantly over time.

Terminology

Summary

FisherAdapTune is a Fisher-guided Adaptive Fine-Tuning framework designed to dynamically select which parameter groups should remain trainable during fine-tuning by tracking the temporal drift of their Fisher geometry. This method addresses the limitation of existing Parameter-Efficient Fine-Tuning (PEFT) methods, which often rely on fixed architectural heuristics rather than dynamic, task-aware criteria. By grounding its selection criterion in the temporal evolution of the Fisher Information Matrix (FIM), FisherAdapTune provides a principled mechanism to identify parameters that actively contribute to adaptation while allowing stabilized groups to be frozen, thereby improving generalization and parameter efficiency.

Theoretical Foundation for Generalization Bounds

The framework is built upon Probably Approximately Correct (PAC-Bayesian) learning theory. The paper shows that the generalization error bound is governed by the divergence between the posterior distribution after fine-tuning and the pretrained prior distribution, DKL(QT P). By relating this divergence to parameter updates using a second-order expansion, it establishes that these updates are controlled by curvature-weighted increments. Specifically, Equation (4) shows that the total divergence decomposes into a sum of curvature-weighted updates along the optimization trajectory:

DKL(QT P) ≲ T X−1 t=0 1/2 δ⊤ t Fθt δt, (4). This formulation shifts the perspective from parameter magnitude to curvature-aligned change, providing a theoretical foundation for using Fisher dynamics to identify which parameters meaningfully contribute to task adaptation.

Curvature Shift as an Adaptation Signal

To make the theoretical criterion actionable, the paper decomposes the Fisher-governed KL cost across layers and uses resulting curvature shifts as a proxy for adaptation relevance. The core signal is defined by tracking the per-layer contribution drift: ct,l − ct−1,l. Under Assumption 3.7 (Smooth parameter updates), this drift is approximated as:

ct,l − ct−1,l ≈ δ⊤ t,l∆Ft,lδt,l (6). When the Fisher stabilizes (Fθt,l ≈ Fθt-1,l), the contribution drift becomes negligible. Conversely, layers with significant Fisher drift continue to reshape their local curvature and contribute meaningfully to task adaptation.

Scale-Invariant Measure of Structural Drift

To capture structural changes in Fisher information while remaining invariant to scale, the paper maps the Fisher tensor at iteration t and layer l to a probability distribution via log-normalization: pt,l(i) = log(1+F t l (i)) / sum j log(1+F t l (j)). The evolution of this distribution across training iterations is quantified using the Jensen-Shannon (JS) divergence, dJS(p, q), which yields a symmetric and bounded metric. Proposition 3.8 establishes that dJS satisfies two key properties:

  1. Structural sensitivity: dJS(pt,l, pt−1,l) = 0 if and only if the log-normalized Fisher distributions coincide (i.e., the normalized shape tensor Pt l has not changed).

  2. Scale invariance: dJS is invariant to uniform scaling F t l → αF t l for any α > 0, since the log-normalization step removes the overall magnitude.

FisherAdapTune Algorithm and Parameter Selection Dynamics

The FisherAdapTune framework implements this criterion by progressively freezing stabilized groups. The algorithm proceeds as follows:

  1. Compute loss L(D; θ) and gradients ∇θL.

  2. If t mod m = 0, estimate per-layer Fisher curvature F˜t Di via Eq. (9).

  3. Partition F˜t Di column-wise into k blocks and build PMFs p(F˜t wik) via Eqs. (11)-(14).

  4. Compute consecutive Fisher drift: ˜d t JS,ik = dJS(p t−1 ik, pt ik) (9).

  5. Smooth and accumulate the drift: ˆd t JS,ik ← β ˆd t−1 JS,ik + (1 − β) ˜d t JS,ik.

  6. If t mod n = 0, compute global statistics of JS distances to set an adaptive threshold τ ← µ + λσ (12).

  7. Freeze stabilized parameter groups: S ← S: ¯dJS,ik ≥ τ (14).

  8. Update only active parameter groups: θS ← θS − η∇θS L(D; θ) (15).

Empirical Validation and Results

Experiments on crack segmentation across SAM2 and SegFormer architectures demonstrate the effectiveness of FisherAdapTune. The results show that FisherAdapTune recovers most of the benefit of full fine-tuning with fewer effective trainable parameters.

Improvements for AI systems

Here are the specific improvements to AI systems based on FisherAdapTune, along with what those improved systems can achieve:


The core improvement lies in moving from static or heuristic parameter selection for fine-tuning to a dynamic, data-driven criterion based on the temporal evolution of local curvature (Fisher Information Matrix).

  1. [Dynamic Parameter Sculpting and Budget Optimization]:

The system will no longer use fixed architectural rules (e.g., freezing top layers) or arbitrary layer schedules. Instead, it will dynamically adjust the trainable parameter set during training by monitoring the Jensen-Shannon distance between consecutive Fisher distributions of different parameter groups.

  1. [Curvature-Aware Adaptation Strategy]:

The system can differentiate between parameters that are actively reshaping the model's local loss landscape (high Fisher drift) and those whose influence has stabilized (low Fisher drift). It will automatically freeze or deactivate parameters whose contribution to the generalization error bound has saturated, ensuring computational resources are focused only on parameters that continue to drive task adaptation.

  1. [Improved In-Distribution Performance and Robustness]:

By preventing over-specialization to the training set—a common pitfall of full fine-tuning—the system will achieve competitive or superior in-distribution performance while maintaining better generalization bounds under distribution shifts (zero-shot transfer). The adaptive thresholding mechanism ensures the model retains essential task-dependent components while discarding redundant ones.

  1. [Enhanced Parameter Efficiency and Reduced Computational Cost]:

The system will consistently achieve the performance benefits of full fine-tuning (or even exceed them in zero-shot settings) while using a significantly smaller, dynamically determined active parameter set compared to fixed PEFT methods like LoRA or BitFit. This translates directly into reduced training time, lower memory footprint, and more sustainable deployment costs for large models.

  1. [Task-Aware Transfer Learning Hierarchy Recovery]:

The framework automatically recovers a standard transfer learning hierarchy (e.g., freezing generic feature-extraction components early while retaining deeper semantic aggregation and decoder modules) based on the empirical observation that generic representations stabilize quickly, while task-dependent modules require longer adaptation time.

This improved AI system can:

  1. [Perform High-Accuracy, Efficient Fine-Tuning]: It can fine-tune massive foundation models (like SAM2 or SegFormer variants) for complex vision tasks (e.g., crack segmentation) with significantly fewer trainable parameters while maintaining state-of-the-art accuracy on both the target domain and in unseen scenarios.

  2. [Optimize Training Efficiency]: It will reduce the effective number of trainable parameters, leading to faster convergence and lower computational overhead during the adaptation phase of a downstream task.

  3. [Increase Model Robustness]: The curvature-based selection mechanism ensures that the resulting model is not over-specialized, leading to superior zero-shot transfer capabilities and better handling of out-of-distribution data shifts compared to models trained with static PEFT methods.

  4. [Provide Principled Adaptation]: It moves beyond trial and error heuristic tuning by grounding parameter selection in the underlying geometry of the loss landscape (Fisher dynamics), providing a principled, mathematically rigorous basis for adaptive fine-tuning strategies across various model architectures.

Sources

Related papers