BPG: Balancing Plasticity and Generalization for Domain Incremental Learning
Qiang Wang, Songlin Dong, Shaokun Wang, Jizhou Han, Xiang Song, Chenhao Ding, Yuhang He, Yihong Gong
Xi'an Jiaotong University · Shenzhen University of Advanced Technology · Harbin Institute of Technology
cs.CV, cs.AI, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 75/100
The gist: BPG: Balancing Plasticity and Generalization for Domain Incremental Learning Summary This paper introduces BPG, a unified framework for Domain Incremental Learning (DIL) that addresses two key
Terminology
Summary
BPG: Balancing Plasticity and Generalization for Domain Incremental Learning
Summary
This paper introduces BPG, a unified framework for Domain Incremental Learning (DIL) that addresses two key limitations of existing parameter-isolation methods: (1) the plasticity challenge
arising from uniform adapter capacity allocation across domains of varying difficulty, and (2) the generalization challenge
caused by brittle hard domain selection at inference time.
Problem Statement
The authors identify that most existing methods assign fixed-capacity modules (e.g., uniform adapters or prompts) to every domain, overlooking the substantial variation in domain difficulty.
They demonstrate that easier domains attain peak performance at small adapter dimensions and may even degrade when granted excessive capacity through overfitting, whereas harder domains continue to benefit from larger capacity.
Additionally, they note that hard selection is inherently fragile
because samples near domain boundaries are prone to misassignment, causing a catastrophic mismatch between the test sample and the applied domain-specific model,
with domain ID errors degrading early-domain accuracy by over seven points.
Proposed Method
BPG consists of two complementary components:
- BPG-Adapter (Domain-wise Capacity Allocation): This component dynamically determines each domain's adapter hidden dimension based on a feature separability score. The separability score is computed as:
-
Between-class scatter: Rtbcs = (1/C) Σ µt,c − µt2
-
Within-class scatter: Rtwcs = (1/C) Σ (1/Nt,c) Σ xt,c,i − µt,c2
-
Separability score: st = Rtbcs / (Rtwcs + ε)
The capacity allocation rule follows an inverse-proportionality relationship: rt = r0 · (s0/st), where s0 and r0 are the reference domain's (ImageNet) separability and hidden dimension. This means hard domains (low st) should receive more parameters to improve plasticity, while easy domains (high st) can be modeled with smaller adapters to avoid unnecessary capacity.
- BPG-Inference (Soft Mixture Strategy): This replaces hard domain selection with a confidence-weighted aggregation over multiple domain-specific models. For each test sample, it:
-
Computes distances to k-means prototypes (k=5) per domain: at = min z − νt,j2
-
Converts distances to normalized confidences: wt = exp(−at) / Σ exp(−ai)
-
Applies confidence-guided pruning (domains below uniform threshold are set to zero)
-
Fuses surviving domain logits: yfinal = Σ wt′ yt
Theoretical Justification
The paper provides a qualitative risk decomposition showing that expected risk Rt(rt) = Rapprox t(rt) + Rgen t(rt), where approximation error decreases exponentially with rt and estimation error increases with √(rt/Nt). Through Lagrangian optimization, they derive that optimal allocation rt⋆ ∝ 1/κt, and since κt ∝ st (monotonically increasing map), this yields rt⋆ ∝ 1/st, consistent with their allocation rule.
Experimental Results
The method is evaluated on three benchmarks:
-
DomainNet: BPG achieves AT=72.19% and FT=0.22% on ViT, surpassing ICON (67.95%) by 4.24% in accuracy and SOYO (1.26%) by 1.04% in forgetting. On CLIP, BPG attains 75.72% and 0.59% in AT and FT, exceeding CP-Prompt (73.35%) by 2.37%.
-
CDDB: BPG achieves the best AT (88.55%) and FT (0.68%) on ViT, outperforming all rehearsal-based methods including DyTox (86.21%). On CLIP, BPG reaches 93.91% with 0.12% forgetting, approaching the Upper Bound (94.40%) within 0.49%.
-
CORe50: BPG attains 91.87% on ViT (surpassing PC's 91.35%) and 92.46% on CLIP (surpassing MoP-CLIP's 92.29%).
Key Ablation Findings
-
Both BPG-Adapter and BPG-Inference independently improve accuracy while reducing forgetting; their combination achieves the best results.
-
BPG-Inference is a plug-and-play module that consistently improves existing methods (e.g., +5.42% for S-iPrompts on DomainNet).
-
BPG-Adapter consistently outperforms uniform adapter counterparts across comparable hidden dimensions.
-
The adaptive capacity allocation transfers well to both Adapter and LoRA fine-tuning paradigms.
-
BPG-Inference mitigates domain-ID misselection: Domain 1 accuracy improves from 71.06% (hard selection) to 76.83% (soft mixture) after training on all six domains.
Efficiency Analysis
BPG adds only 0.38 hours for separability scoring (one-time cost) and finishes in 23.02 hours total on DomainNet, just +4.8% over uniform r=64 baseline. Inference latency increases from 11.5 ms to 74.2 ms per image due to logit fusion across all domains, but this requires no learnable router.
Conclusion
The authors state: "We address two key bottlenecks in domain incremental learning: insufficient plasticity from uniform adapter capacity and brittle generalization from hard domain selection at inference. We propose BPG, a unified framework comprising BPG-Adapter, which adaptively allocates adapter hidden dimensions from feature separability, and BPG-Inference, which soft-mixes logits from multiple domain-specific models. Experiments on DomainNet, CDDB, and CORe50 show that BPG achieves state-of-the-art average accuracy with near-zero forgetting, validating the need to jointly adapt plasticity and generalization in lifelong learning."
Improvements for AI systems
Based on this paper, I can improve AI systems in the following ways:
1. Adaptive Parameter Allocation for Continual Learning Systems
- Implement a feature-separability scoring mechanism (between-class vs. within-class scatter) to dynamically allocate model capacity per task/domain. This allows the system to automatically give more parameters to hard domains and fewer to easy ones, preventing overfitting on simple tasks while ensuring sufficient plasticity for complex ones. The improved system can maintain high accuracy across heterogeneous domains without manual hyperparameter tuning per domain.
2. Confidence-Weighted Inference for Domain-Agnostic Deployment
- Replace brittle hard domain selection with a soft mixture strategy that computes k-means prototype distances, converts them to normalized confidences, and fuses logits from multiple domain-specific models. This makes the system robust to ambiguous or boundary samples, reducing catastrophic misclassification when domain identity is uncertain. The improved system can operate reliably in real-world scenarios where input domain is unknown or mixed.
3. Plug-and-Play Robustness Module for Existing Incremental Learning Models
- Add BPG-Inference as a drop-in inference-time module to any existing domain-incremental model (e.g., prompt-based or adapter-based). This module improves average accuracy by up to +5.42% without retraining, by softly aggregating predictions from all domain-specific components rather than committing to a single one. The improved system can be retrofitted onto deployed models to enhance generalization with zero additional training cost.
4. Risk-Aware Capacity Optimization via Theoretical Guarantees
- Use the derived risk decomposition (approximation error decreases exponentially with capacity, estimation error increases with √(capacity/N)) to guide capacity allocation in any lifelong learning system. This enables principled trade-off between plasticity and generalization, ensuring the system allocates resources optimally even when domain difficulty is unknown a priori. The improved system can self-tune its architecture complexity per task, achieving near-zero forgetting while maximizing accuracy.
5. Efficient Domain-Specific Fine-Tuning with Adaptive LoRA/Adapter Ranks
- Apply the inverse-proportionality rule (rank ∝ 1/separability) to set LoRA or adapter ranks per domain during fine-tuning. This reduces total parameter count and training time (e.g., only +4.8% overhead vs. uniform capacity) while improving accuracy, making the system more scalable for large-scale multi-domain deployment with limited compute. The improved system can fine-tune on many domains with lower memory footprint and faster convergence.
Abstract
Deep neural networks excel in various tasks but struggle to generalize across evolving data distributions, leading to significant performance degradation under domain shifts. Domain incremental learning (DIL) addresses this challenge by enabling models to continuously adapt while retaining prior knowledge. Among existing DIL approaches, the parameter-isolation paradigm achieves state-of-the-art performance. However, these methods often adopt a one-size-fits-all approach to adapt to new domains, resulting in either insufficient learning capacity or redundant parameters. In this work, we propose BPG, a unified framework that addresses both challenges through two complementary components: BPG-Adapter, which dynamically determines each domain's adapter hidden dimension based on domain-specific feature separability, and BPG-Inference, a soft domain mixture strategy that integrates multiple domain-specific models at test time, mitigating domain ID misselection. Experimental results on DomainNet, CDDB, and CORe50 demonstrate that BPG consistently outperforms uniform adapter-based approaches and hard domain selection strategies, achieving state-of-the-art average accuracy while reducing forgetting to as low as 0.22% on DomainNet.
Sources
- Dynamic Integration of Task-Specific Adapters for Class Incremental Learning
- Class-Independent Increment: An Efficient Approach for Multi-label Class-Incremental Learning
- GFPL: Generative Federated Prototype Learning for Resource-Constrained and Data-Imbalanced Vision Task
- LPT: Less-overfitting Prompt Tuning for Vision-Language Model
- Trajectory-Diversity-Driven Robust Vision-and-Language Navigation
- Consistent Supervised-Unsupervised Alignment for Generalized Category Discovery
- Beyond CLIP Generalization: Against Forward&Backward Forgetting Adapter for Continual Learning of Vision-Language Models
- VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
- Subspace Regularizers for Few-Shot Class Incremental Learning
- Space Rotation with Basis Transformation for Training-free Test-Time Adaptation
- Variational Prototype Replays for Continual Learning
- Is Parameter Isolation Better for Prompt-Based Continual Learning?
- Continual Learning and Catastrophic Forgetting
- AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning
- DyLoRA: Parameter Efficient Tuning of Pre-trained Models using Dynamic Search-Free Low-Rank Adaptation
- GeLoRA: Geometric Adaptive Ranks For Efficient LoRA Fine-tuning
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Continual Knowledge Consolidation LORA for Domain Incremental Learning
- On Tiny Episodic Memories in Continual Learning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models