EfficientXpert: Efficient Domain Adaptation for Large Language Models via Propagation-Aware Pruning

arXiv:2511.19935 · cs.LG, cs.CL · Submitted 2025-11-25 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "EfficientXpert: Efficient Domain Adaptation for Large Language Models via Propagation-Aware Pruning".

Jane: The gist: EfficientXpert proposes a lightweight domain pruning framework that combines propagation-aware pruning with an efficient adapter update algorithm to create sparse,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Now, let's look closer at what this paper is actually proposing with "EfficientXpert: Efficient Domain Adaptation for Large Language Models via Propagation-Aware Pruning." They are presenting a lightweight framework that blends propagation-aware pruning with an efficient adapter update algorithm.

Jane: The core thesis is that this combination allows you to create sparse, domain-adapted language models that maintain performance comparable to dense models even when sparsity is quite high.

Lu: It introduces two main innovations: the ForeSight Mask, which is a domain-aware, dynamic pruning method incorporating forward error propagation into its scoring mechanism.

Meng: So it doesn't just prune weights randomly; it estimates the impact of zeroing out each weight on downstream representations before pruning anything.

Tom: Right. It derives a per-weight importance score by approximating the loss increase incurred by zeroing a single entry while holding other parameters fixed, coupling the weight magnitude and input energy with downstream amplification.

Jane: Then you have the Partial Brain Surgeon, which is an efficient adapter realignment step that solves a ridge regression to perform post-pruning recovery.

Lu: This part updates the low-rank factors by solving a closed-form weighted ridge problem derived from a diagonal activation-norm approximation, which reallocates the limited rank budget effectively.

Meng: It sounds like they are using this recovery step to ensure that the low-rank updates stay aligned with the evolving sparse structure, controlling drift on retained weights.

Tom: And their results show that across health and legal tasks, EfficientXpert keeps up to ninety-eight percent and ninety-eight point four three percent of dense model performance at a forty percent sparsity level on LLaMA 7B <ref:2511.19935#pg2,and 98.43% of dense model performance at>.

Jane: Plus, they manage to keep the end-to-end wall clock runtime within about eighteen percent of plain LoRA when using the same training configuration <ref:2511.19935#pg2,end-to-end wall clock runtime within>.

Lu: This is significant because it shows that you can get this level of performance retention with minimal training time and comparable memory usage to dense models.

Meng: It confirms that this isn't just a theoretical exercise; it’s a method that actually translates into practical efficiency gains for deployment.

Tom: Exactly, and the whole point is to enable a one-step transformation of general pretrained models into sparse, domain-adapted experts using this framework.

Conclusion: Jane: So wrapping up on "EfficientXpert: Efficient Domain Adaptation for Large Language Models via Propagation-Aware Pruning," the paper proposes a unified way to prune and adapt LLMs that is much more tailored to specific domains.

Tom: The authors, Songlin Zhao, Michael Pitts, and Zhuwei Qin from UC Berkeley and San Francisco State University, have put forward this framework by combining the ForeSight Mask with the Partial Brain Surgeon update.

Lu: What this means for the future is that we can move toward creating models that are not just general-purpose but are specialized experts ready to be deployed efficiently in niche areas like law or healthcare.

Meng: It suggests a path where we don't have to choose between using a massive model and achieving necessary specialization, because we can do both at once with this type of method.

Jane: In simple terms, they show that you can achieve high domain specialization while keeping the overhead low enough for real-world deployment in resource-constrained settings.

Tom: The title itself tells us they are aiming for efficient adaptation specifically within a domain context, which is key to addressing the deployment challenges of these large language models.

Lu: It’s about creating these specialized experts that run efficiently, not just general compressed versions that lose capability everywhere.

Meng: This work gives practitioners a concrete method for how to adapt pretraining effectively without ballooning the training costs or memory requirements excessively.

Songlin Zhao, Michael Pitts, Zhuwei Qin

University of California, Berkeley · San Francisco State University

cs.LG, cs.CL

Submitted: 2025-11-25

Updated: 2026-10-03

Importance score: 92/100

The gist: The gist: EfficientXpert proposes a lightweight domain pruning framework that combines propagation-aware pruning with an efficient adapter update algorithm to create sparse, domain-adapted LLMs with

Key concepts

ForeSight Mask
This is a dynamic, gradientfree pruning method that assesses weight importance by estimating how much zeroing a single weight affects downstream results. It considers the weight's magnitude, input energy along its row, and amplification along its column to create an accurate importance score.
Partial Brain Surgeon
This component handles post-pruning recovery by updating low-rank factors after pruning. It solves a closed-form weighted ridge problem to reallocate the limited rank budget, compensating for the loss while ensuring updates align with the sparse pattern and controlling drift on retained weights.
Grassmann Distance
This metric measures the similarity between two LoRA adapters based on cosine angles derived from their weight matrices. The analysis shows that domain identity, rather than task type, strongly dictates these distances, helping to separate domain-specific subspace directions.
Domain Sensitivity Analysis
This analysis uses Grassmann distance and training entropy to show how pruning affects model behavior. It reveals that FORESIGHT and FORESIGHT+PBS maintain higher entropy during finetuning compared to dense LoRA, suggesting less over-confident outputs while preserving strong in-domain performance.

Terminology

Summary

The gist: EfficientXpert proposes a lightweight domain pruning framework that combines propagation-aware pruning with an efficient adapter update algorithm to create sparse, domain-adapted LLMs with performance comparable to dense models at low sparsity levels. This work matters because it addresses the deployment barrier of large language models by reducing inference memory and computation while preserving domain performance.

EfficientXpert Framework

EfficientXpert introduces a lightweight framework that integrates pruning with LoRA fine-tuning to produce sparse, domain-adapted LLMs with minimal overhead It consists of two tightly coupled components: the ForeSight Mask and the Partial Brain Surgeon The ForeSight Mask is a domain-aware, dynamic, gradientfree pruning method that incorporates forward error propagation into its scoring mechanism by estimating the impact of weights on downstream representations It derives a per-weight importance score by approximating the loss increase incurred by zeroing a single entry while holding other parameters fixed This score couples (i) the magnitude of the pruned weight, (ii) the input-energy along the affected row via (X⊤X)ii, and (iii) the downstream amplification along the affected column via (U2U⊤2)jj

Partial Brain Surgeon

Complementing ForeSight Mask, Partial Brain Surgeon performs a post-pruning adapter recovery step Given the fixed mask, it updates the low-rank factors (e.g., updating B with A fixed) by solving a closed-form weighted ridge problem derived from a diagonal activation-norm approximation This recovery reallocates the limited rank budget to compensate for pruning-induced loss while explicitly controlling drift on retained weights The closed-form solution is given by ∆bi = − U1,i,Si DS A⊤S + λIr−1(4) This step suppresses pruned weights with minimal adapter perturbation, ensuring its updates remain aligned with the sparse pattern and enabling more efficient learning within the ForeSight Mask’s support

Performance and Efficiency

Across a comprehensive suite of health and legal tasks, EfficientXpert retains up to 98% and 98.43% of dense model performance at 40% sparsity on LLaMA 7B It achieves this while keeping end to end wall clock runtime within 18% of plain LoRA under the same training configuration and maintaining a comparable peak GPU memory footprint The time complexity for ForeSight costs O(mnr), whereas Partial Brain Surgeon (PBS) can be executed in parallel across rows, achieving an effective constant time runtime in practice under sufficient compute resources

Domain Sensitivity Analysis

The Grassmann distance between two LoRA adapters of the same weight is given by d(U1, U2) = Xr i=1 cos−1(σi) 2 The analysis shows that task similar but cross domain pairs concentrate at noticeably larger distances, while task diverse but intra domain pairs cluster at smaller distances This separation indicates that domain identity, rather than task type, more strongly determines the subspace directions emphasized by finetuning Furthermore, training entropy analysis reveals that FORESIGHT and FORESIGHT+PBS maintain higher entropy throughout finetuning compared to dense LoRA (and LoRA+Wanda), indicating a less over-confident output distribution while preserving strong in-domain perplexity

Conclusion

EfficientXpert introduces a domain-aware pruning framework that transforms general-purpose LLMs into sparse, domain-specialized experts with minimal overhead By integrating the propagation-aware ForeSight Mask and the lightweight Partial Brain Surgeon update into LoRA finetuning, EfficientXpert enables a unified, end-to-end pruning and adaptation process Our experiments in the health and legal domains show that EfficientXpert consistently outperforms existing pruning baselines across all sparsity levels Moreover, EfficientXpert incurs only modest additional finetuning time relative to dense LoRA while maintaining essentially the same peak memory footprint All evaluations were conducted on 4 NVIDIA A100 GPUs with 80GB of memory with model loaded at float 16 precision The license for the models and datasets used in this paper is as follows: Llama-2-7b-hf: The model is licensed under the LLaMA 2 community license, Llama-3-8b: The model is licensed under META LLama 3 community license, PubMedQA: The dataset is licensed under MIT license, MedNLI: The dataset is licensed under the Physionet Credentialed Health Data License Version 1.5.0, HQS: The dataset is licensed under Apache License 2.0, CaseHold: The dataset is licensed under Apache License 2.0, ContractNLI: The dataset is licensed under the CC BY 4.0 license.

--- Page 7 ---

Table 4: Qwen 8B Legal EfficientXpert Runtime and Memory. Prune% Total Runtime Peak Memory Legal ppl –– –– 106 mins 45728.68 MiB 5.1969

--- Page 3 ---

Figure 2: (a) Comparison of EfficientXpert with existing domain pruning methods. (b) Overview of EfficientXpert framework, including ForeSight Mask and Partial Brain Surgeon

--- Page 4 ---

Algorithm 1 EfficientXpert 1: Input: f(X M, W, B, A), D, Dcal, s, r, T Output: f⋆

--- Page 6 ---Table 2: Evaluation in the Legal domain. Best pruned scores are bold—–– ––

--- Page 8 ---Table 3: Qwen 8B Health EfficientXpert Runtime and Memory. Prune% Total Runtime Peak Memory Health ppl –– ––

--- Page 10 ---Figure 5: Projection energy by full name across domains. L0 L1 L2 L3 L4 L5 L6 L7 L8 L9L10L11L12L13L14L15L16L17L20

--- Page 9 ---Figure 6: Grassmann distance between LoRA adapters and weight matrices across layers. Layerwise Average Grassmann Distance Across Domains Domain CaseHold Health Legal PubMedQA

--- Page 5 ---Table 8: Dense Health Domain Training Hyperparameters Hyperparameter Value model name meta-llama/Llama-2-7b-hf learning rate 1 × 10−4 per device batch size 2 gradient accumulation steps 4 num train epochs 3 weight decay 0.01 fp16 True save steps 1000 max seq length 2048 lora r 8 lora alpha 16 lora dropout none lora bias none lora target modules q proj, o proj, v proj, k proj, gate proj, up proj, down proj

--- Page 13 ---Table 13: Performance metrics for different methods at various ranks r at 50% sparsity on Llama-3.2-1B Method Rank Med PPL MedNLI Acc MedNLI F1 PubMedQA Acc PubMedQA F1 HQS R1 HQS R2 HQS RL

--- Page 9 ---Figure 8: Grassmann distance between LoRA adapters and weight matrices by weight type Domain CaseHold Health Legal PubMedQA

--- Page 4 ---Figure 7: Projection energy across layers. L0 L1 L2 L3L4L5L6L7

--- Page 14 ---Table 12: Overview of legal datasets used for evaluation. ContractNLI CaseHOLD BillSum

--- Page 15 ---Table 9: Dense Legal Domain Training Hyperparameters Hyperparameter Value model name meta-llama/Llama-3-8b learning rate 1 × 10−4 per device batch size 1 gradient accumulation steps 8 num train epochs 3 weight decay 0.

Improvements for AI systems

  1. textbfPropagation-Aware Pruning for Domain Adaptation (ForeSight Mask): This mechanism addresses representation errors that propagate through subsequent layers by scoring a pruning mask based on the induced change in the two-layer output, which allows pruning decisions to be informed by downstream impact rather than just local weight magnitude.

  2. textbfEnd-to-End Sparse Domain Adaptation (EfficientXpert): This framework enables a one-step transformation of general pretrained models into sparse, domain-adapted experts by combining the ForeSight Mask with the Partial Brain Surgeon to ensure that low-rank updates stay aligned with the evolving sparse structure.

  3. textbfLightweight Adapter Realignment (Partial Brain Surgeon): This step solves a ridge regression to suppress pruned coordinates in constant time complexity, ensuring that low-rank updates stay aligned with the evolving sparse structure and reallocates the limited rank budget to compensate for pruning-induced loss.

  4. textbfResource-Efficient Training Pipeline: By utilizing these methods, EfficientXpert achieves a fine-tuning cost comparable to plain LoRA and operates within 1% of LoRA’s peak GPU memory footprint, making domain adaptation feasible in resource-constrained environments.

  5. textbfDomain Sensitivity Analysis: The analysis reveals that domain identity, rather than task type, more strongly determines the subspace directions emphasized by finetuning, suggesting that pruning strategies should be organized by domain for optimal performance retention.

Sources

Related papers