From Parameters to Behaviors: A Survey of Model Fusion for Large Language Models
cs.CL
Submitted: 2026-09-17
Updated: 2026-09-17
Comments: 25 pages, 4 figures. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026
Code: https://github.com/Baicaihaochi/Awesome-Model-Fusion-Survey
License: http://creativecommons.org/licenses/by/4.0/
The gist: Model fusion integrates the capabilities from source models into a single target model.
Terminology
Abstract
Model fusion integrates the capabilities from source models into a single target model. As of June 2026, Hugging Face hosts more than 2M models. This growing pool provides a rich base for model reuse and capability integration. Yet existing surveys often cover only separate parts of this space, and they do not provide a unified definition or a systematic taxonomy. This survey defines model fusion and organizes prior work into three levels: parameter-level, representation-level, and behavior-level fusion. We also review related metrics, benchmarks, and applications, summarize current challenges, and identify future directions. Our goal is to provide a clear map of this area and support future work on model fusion. A comprehensive list of papers about model fusion is available at https://github.com/Baicaihaochi/Awesome-Model-Fusion-Survey.
Sources
- X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs
- An Empirical Study and Theoretical Explanation on Task-Level Model-Merging Collapse
- Harnessing Multiple Large Language Models: A Survey on LLM Ensemble
- Transport and Merge: Cross-Architecture Merging for Large Language Models
- Model Merging via Multi-Teacher Knowledge Distillation
- DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling
- SODA: Semi On-Policy Black-Box Distillation for Large Language Models
- SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging
- GLM-5: from Vibe Coding to Agentic Engineering
- FeatCal: Feature Calibration for Post-Merging Models
- Distilling the Knowledge in a Neural Network
- A Systematic Study of In-the-Wild Model Merging for Large Language Models
- CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation
- DP-OPD: Differentially Private On-Policy Distillation for Language Models
- If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs
- Draft-OPD: On-Policy Distillation for Speculative Draft Models
- Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
- A Unified Generalization Framework for Model Merging: Trade-offs, Non-Linearity, and Scaling Laws
- MAGIC: Achieving Superior Model Merging via Magnitude Calibration
- Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering