Continual Learning for Sequential Personalization of Small Language Models: A Stability Monitoring Analysis
cs.LG
Submitted: 2026-06-26
Updated: 2026-09-12
Comments: Corrected an implementation error in next-token selection for left-padded evaluation batches. Recomputed KL divergence, entropy, and margin for all runs, and updated the related tables, figures, discussion, and conclusions. Training, task accuracy, and continual learning results are unchanged
Code: https://github.com/tspthomas/slm_stability_cl
License: http://creativecommons.org/licenses/by/4.0/
The gist: Small Language Models (SLMs) are increasingly being considered for deployment on edge devices such as laptops, enabling private, low-latency, and locally personalized applications.
Terminology
Abstract
Small Language Models (SLMs) are increasingly being considered for deployment on edge devices such as laptops, enabling private, low-latency, and locally personalized applications. However, personalization requires models to adapt over time to evolving user- or task-specific data, placing them in a continual learning setting. This creates the risk of catastrophic forgetting, where learning new information degrades performance on previously learned tasks or broader model capabilities. Recent benchmarks such as TRACE have shown that continual fine-tuning can significantly degrade the general abilities of aligned large language models. In this work, we present a study for sequential LoRA personalization of SLMs. We save model checkpoints after each adaptation stage and evaluate them on current tasks, previously seen tasks, and a fixed reference set. This checkpoint-level protocol enables us to monitor task performance, forgetting, and reference set drift over time. We show that lightweight reference set distributional diagnostics can reveal model-specific instability patterns during sequential LoRA personalization of SLMs, including cases where task-level metrics alone hide harmful adaptation. We hope this can highlight new research avenues for monitoring stability of SLMs in a continual learning setting.
Sources
- Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages
- The Llama 3 Herd of Models
- DATA: Decomposed Attention-based Task Adaptation for Rehearsal-Free Continual Learning
- Decoupled Weight Decay Regularization
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection
- Gemma 3 Technical Report
- TRACE: A Comprehensive Benchmark for Continual Learning in Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks