On Emergent Capabilities and Model Merging
cs.LG, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: main paper has 8 pages, 5 figures, and 4 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: Fine-tuned checkpoints and adapters now fill public repositories, and the most common operation applied to these artifacts is model merging: arithmetic on their weights that assembles capabilities
Terminology
Abstract
Fine-tuned checkpoints and adapters now fill public repositories, and the most common operation applied to these artifacts is model merging: arithmetic on their weights that assembles capabilities cheaply. We ask what this operation does to emergent capabilities: behaviors an artifact carries that were never an explicit training target. Studying two independent testbeds (activation oracles and emergent-misaligned models) across three model families, we find that the answer is threefold. First, merging preserves an emergent capability that both parents carry: merging two misaligned checkpoints retains most of their broad misalignment across the whole mixing range. Second, merging cannot create an emergent capability that is superadditive in its parents: no weighted merge of two single-task oracles reaches the jointly-trained oracle's auditing ability. Third, when only one parent carries the capability, merging dilutes it faster than the trained capability that accompanies it: the gap is significant in most settings. In short, emergent behaviors of an artifact do not compose the way its trained capability does.
Sources
- Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation
- Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language
- Gemma 2: Improving Open Language Models at a Practical Size
- The Llama 3 Herd of Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
- Qwen2.5 Technical Report
- Qwen3 Technical Report
- Convergent Linear Representations of Emergent Misalignment
- Model Organisms for Emergent Misalignment
- Predicting Neural Network Accuracy from Weights
- Emergent Abilities of Large Language Models
- Demystifying Mergeability: Interpretable Properties to Predict Model Merging Success
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks