Bilinear Optimization Divergence: Diagnosing Factor-Constrained LoRA Continual Learning
cs.LG
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: 22 pages, 3 figures. Code is available at [https://github.com/legend91019/My_first](https://github.com/legend91019/My_first)
Code: https://github.com/legend91019/My_first
License: http://creativecommons.org/licenses/by/4.0/
The gist: Orthogonality in a LoRA factor does not by itself specify what the composed update protects: the answer depends on the task-start state, the parameterization, and the realized optimizer displacement.
Terminology
Abstract
Orthogonality in a LoRA factor does not by itself specify what the composed update protects: the answer depends on the task-start state, the parameterization, and the realized optimizer displacement. We formalize this question through Bilinear Optimization Divergence (BOD), an anchor-relative diagnostic of effective-update response on selected historical features. The finite-step analysis distinguishes two cases. In a shared adapter, protecting the routing displacement leaves a learned-anchor residual through the changing companion factor. In a fresh zero-output block, a feasible routing state can protect the composed update while both current factors remain trainable. These conditions yield Semi-Frozen Orthogonal Routing (SFOR) for shared adapters and current-block hard protection for cumulative O-LoRA; Weight Residual Projection (WRP) enforces the required displacement after the optimizer step. Controlled two-task traces verify the predicted residual paths, reducing normalized historical response from 19.12% to 0.005% in the shared family and from 7.72% to 0.002% in the cumulative family. Four-task experiments on Qwen3-8B characterize the resulting trade-offs: SFOR improves backward transfer (BWT) from-2.47 to-0.86 with nearly unchanged average accuracy (AA), while O-LoRA hard protection improves three-order mean AA from 80.27% to 81.30% and forgetting measure (FM) from 2.20 to 0.43. Component controls also show that stricter feasibility need not improve final task performance. Together, the analysis and evidence provide an architecture-conditioned account of which constraint to enforce, how to enforce it, and how to interpret its empirical value.
Sources
- Janus-LoRA: A Balanced Low-Rank Adaptation for Continual Learning
- Spectral Imbalance Causes Forgetting in Low-Rank Continual Adaptation
- LoRA+: Efficient Low Rank Adaptation of Large Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Adam: A Method for Stochastic Optimization
- TRGP: Trust Region Gradient Projection for Continual Learning
- Decoupled Weight Decay Regularization
- LFPT5: A Unified Framework for Lifelong Few-shot Language Learning Based on Prompt Tuning of T5
- Progressive Prompts: Continual Learning for Language Models
- Gradient Projection Memory for Continual Learning
- Qwen3 Technical Report
- Asymmetry in Low-Rank Adapters of Foundation Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks