Least but not Last: Fine-tuning Intermediate Principal Components for Better Performance-Forgetting Trade-Offs
cs.LG
Submitted: 2026-02-03
Updated: 2026-08-31
License: http://creativecommons.org/licenses/by/4.0/
The gist: Low-Rank Adaptation (LoRA) methods have emerged as crucial techniques for adapting large pre-trained models to downstream tasks under computational and memory constraints.
Terminology
Abstract
Low-Rank Adaptation (LoRA) methods have emerged as crucial techniques for adapting large pre-trained models to downstream tasks under computational and memory constraints. However, they face a fundamental challenge in balancing task-specific performance gains against catastrophic forgetting of pre-trained knowledge, where existing methods provide inconsistent recommendations. This paper presents a comprehensive analysis of the performance-forgetting trade-offs inherent in low-rank adaptation using principal components of weight matrices as initialization. Our investigation reveals that fine-tuning intermediate components leads to better balance and robustness to high learning rates than first (PiSSA) and last (MiLoRA) components in existing work. Building on these findings, we provide practical guidelines for initialization of LoRA methods to balance the performance-forgetting trade-off. In a thorough empirical study on a variety of computer vision and NLP tasks we confirm that these guidelines achieve high accuracy and reduced forgetting.
Sources
- LoRA Learns Less and Forgets Less
- VeRA: Vector-based Random Matrix Adaptation
- DoRA: Weight-Decomposed Low-Rank Adaptation
- Periodicity of power Fibonacci sequences modulus a Fibonacci number
- Plasmonic detection of the parity anomaly in a two-dimensional Chern insulator
- 1LoRA: Summation Compression for Very Low-Rank Adaptation
- LoRA: Low-Rank Adaptation of Large Language Models
- Scaling Laws for Forgetting When Fine-Tuning Large Language Models
- Scaling Laws for Neural Language Models
- Continual Learning: Applications and the Road Forward
- MiLoRA: Harnessing Minor Singular Components for Parameter-Efficient LLM Finetuning
- DiffFit: Unlocking Transferability of Large Diffusion Models via Simple Parameter-Efficient Fine-Tuning
- BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
- AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks