AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention
Juncheng Liao, Jinfan Lv, Guoming Wang, Jupeng Zheng, Ling Xiao, Siliang Tang
Zhejiang University · Fudan University · Sun Yat-sen University · Hokkaido University
cs.CL
Submitted: 2026-08-12
Updated: 2026-08-13
Code: https://github.com/kaln27/AWARe
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention introduces a fine-tuning method for Multimodal Large Language Models (MLLMs) that mitigates catastrophic
Terminology
Summary
AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention introduces a fine-tuning method for Multimodal Large Language Models (MLLMs) that mitigates catastrophic forgetting by dynamically controlling parameter updates based on activation patterns. The paper states: "AWARe assigns activation-based importance scores to parameters, selectively freezing those essential for preserving prior capabilities while allowing less important parameters to adapt to new tasks. Importantly, AWARe operates without modifying model architectures, ensuring compatibility with existing inference engines."
The method has two phases: Knowledge Profiling, which estimates neuron saliency from a small calibration set, and Constrained Fine-tuning, which freezes the most salient parameters during downstream training.
The saliency estimation involves computing L2-norm of activations along the sequence length dimension, applying per-sample L2-normalization across the hidden dimension, and averaging across the batch to obtain a saliency score for each neuron. A retention ratio ρ controls the fraction of parameters frozen, with high-saliency neurons frozen via a binary gradient mask during optimization.
AWARe is applied specifically to the linear projection layers within the self-attention mechanism (i.e., q proj, k proj, v proj) and the linear layers within the multimodal projector (mm projector),
while MLP blocks and output projections remain entirely frozen. The paper notes this selection is informed by research showing updating self-attention projections tends to cause significantly less catastrophic forgetting of pre-existing knowledge compared to the multilayer perceptron (MLP) blocks.
Experiments use LLaVA-v1.5-7b as the base model, with upstream tasks including OKVQA, OCRVQA, GQA, and TextVQA, and downstream tasks including IconQA and COCO-Caption. Results show AWARe achieves superior harmonic balance between plasticity and stability compared to baselines like LoRA, DoRA, DARE, Orth-Reg, Model Tailor, LoRASculpt, and SPIDER. For IconQA, AWARe achieves H=103.2 compared to the sub-optimal 98.6, and for COCO-Caption, H=108.0 compared to 104.2. When upstream data is restricted, using MMMU as a general-purpose calibration set remains highly effective.
On the MLLM-DCL continual learning benchmark, AWARe achieves the highest average per-task performance (Avg) and best final performance (Last), improving overall averages by 3.18 and 1.77 points over the strongest baseline. The paper reports: AWARe achieves the highest average per-task performance (Avg) and the best final performance after the last task (Last), improving the overall averages by 3.18 and 1.77 points over the strongest baseline, respectively.
Ablation studies demonstrate that activation-based selection outperforms random selection, weight norm-based selection, and hybrid approaches. The Global-Highest selection strategy at a 30% retention ratio achieves optimal balance, with the paper noting freezing the top 30% of self-attention parameters is often sufficient to preserve upstream knowledge while keeping downstream performance competitive.
Calibration dataset composition experiments show that using all upstream sources yields the most balanced performance, though individual upstream tasks also preserve knowledge effectively. Sensitivity analysis shows minimal standard deviation across random runs, confirming statistical stability.
The paper also validates AWARe on Qwen2.5-VL, showing it consistently preserves the strongest upstream performance among fine-tuned methods while retaining competitive downstream adaptation.
Parameter efficiency analysis shows AWARe updates approximately 17.5% of model parameters. The main contributions are summarized as: introducing the AWARe framework, demonstrating efficiency and simplicity without architectural changes, and providing comprehensive validation through ablations and benchmarks.
Improvements for AI systems
Improvements to AI Systems:
-
Adaptive Parameter Freezing for Continual Learning: Implement activation-weighted importance scoring to dynamically freeze high-saliency parameters during fine-tuning, enabling models to retain pre-existing knowledge while adapting to new tasks without architectural changes.
-
Calibration-Efficient Knowledge Preservation: Use a small, general-purpose calibration set (e.g., MMMU) to estimate neuron saliency via L2-norm activations, allowing rapid deployment in data-scarce scenarios while maintaining robust upstream performance.
-
Selective Layer-Specific Fine-Tuning: Restrict parameter updates to self-attention projections (q/k/v) and multimodal projector layers, while freezing MLP blocks and output projections, reducing catastrophic forgetting by 3.18 points (Avg) and 1.77 points (Last) on continual learning benchmarks.
-
Optimal Retention Ratio Control: Freeze the top 30% of salient parameters (Global-Highest strategy) to achieve the best harmonic balance between plasticity and stability, improving harmonic scores to 103.2 (IconQA) and 108.0 (COCO-Caption) over sub-optimal baselines.
-
Activation-Based Saliency Scoring: Replace random or weight-norm-based selection with activation-derived importance (per-sample L2-normalization, batch averaging), yielding statistically stable performance with minimal variance across runs.
What the Improved AI System Can Do:
-
Continual Multimodal Learning: A vision-language model that sequentially learns new tasks (e.g., visual question answering, captioning) without forgetting prior knowledge, maintaining high accuracy on both old and new benchmarks.
-
Efficient Domain Adaptation: Rapidly fine-tune on new domains with limited data, preserving general-purpose capabilities while achieving competitive downstream performance, using only 17.5% of parameters.
-
Deployment-Ready Compatibility: Operate seamlessly with existing inference engines, as no architectural modifications are required, enabling drop-in integration into production systems.
-
Robust Knowledge Retention: Maintain superior performance on upstream tasks (e.g., OKVQA, TextVQA) even after extensive downstream training, with a harmonic balance exceeding current state-of-the-art methods.
-
Scalable Calibration: Leverage a single, reusable calibration dataset (e.g., MMMU) to profile neuron saliency across different base models (e.g., LLaVA, Qwen2.5-VL), reducing setup overhead for new deployments.
Abstract
Multimodal Large Language Models (MLLMs) exhibit strong generalization and reasoning abilities due to large-scale multimodal pre-training. However, fine-tuning these models on downstream tasks often leads to catastrophic forgetting, where newly learned task-specific knowledge degrades previously acquired capabilities. This issue arises because gradient updates for new tasks overwrite parameters critical to prior knowledge, limiting the practical deployment of MLLMs. To address this challenge, we propose Activation-Weighted Adaptive REtention (AWARe), a fine-tuning method that mitigates catastrophic forgetting by dynamically controlling parameter updates based on activation patterns. AWARe assigns activation-based importance scores to parameters, selectively freezing those essential for preserving prior capabilities while allowing less important parameters to adapt to new tasks. Importantly, AWARe operates without modifying model architectures, ensuring compatibility with existing inference engines. Extensive experiments demonstrate that AWARe effectively preserves upstream capabilities while achieving superior downstream performance compared to existing methods. Code is available at https://github.com/kaln27/AWARe.
Sources
- On Tiny Episodic Memories in Continual Learning
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- LoRA Learns Less and Forgets Less
- The Power of Scale for Parameter-Efficient Prompt Tuning
- LLaVA-OneVision: Easy Visual Task Transfer
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- Revisiting Catastrophic Forgetting in Large Language Model Tuning
- Learning to Learn without Forgetting by Maximizing Transfer and Minimizing Interference
- Multimodal Instruction Tuning with Conditional Mixture of LoRA
- An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
- ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Forgetting before Learning: Utilizing Parametric Arithmetic for Knowledge Updating in Large Language Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering