Understanding and Enforcing Weight Disentanglement in Task Arithmetic
cs.AI
Submitted: 2026-04-18
Updated: 2026-08-27
Comments: CVPR 2026 Oral
Code: https://github.com/RL-MIND/OrthoReg
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Layer Normalization
- Scaling Laws for Neural Language Models
- Pre-trained Large Language Models for Financial Sentiment Analysis
- An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
- LLaMA: Open and Efficient Foundation Language Models
- Multi-Task Model Merging via Adaptive Weight Disentanglement
- Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
- Efficient Adaptation of Pre-trained Vision Transformer underpinned by Approximately Orthogonal Fine-Tuning Strategy
- Task Vector Bases: A Unified and Scalable Framework for Compressed Task Arithmetic
- A Survey of Controllable Text Generation using Transformer-based Pre-trained Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection