Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
cs.CL
Submitted: 2025-11-17
Updated: 2026-09-18
Code: https://github.com/facebookresearch/llm_souping
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) have displayed remarkable capabilities across diverse domains, but their training remains resource- and time-intensive, requiring massive computational resources and
Terminology
Abstract
Large Language Models (LLMs) have displayed remarkable capabilities across diverse domains, but their training remains resource- and time-intensive, requiring massive computational resources and careful orchestration of training procedures. Model souping-the practice of averaging weights from multiple models of the same architecture-has emerged as a promising pre- and post-training technique that can enhance performance without expensive retraining. We observe that previous souping approaches can lead to collapse in precision-sensitive LLM capabilities. In this paper, we introduce SoCE, a principled approach for model souping to overcome this shortcoming. The proposed method utilizes benchmark composition to identify optimal model candidates and applies non-uniform weighted averaging to maximize performance. Contrary to previous approaches, our method leverages the observation that different clusters (or categories) of points within a benchmark often exhibit low inter-correlations in model performance. SoCE identifies "expert" models for each weakly-correlated category cluster and combines them using optimized weighted averaging rather than uniform weights. We demonstrate that SoCE improves performance and robustness across multiple domains and achieves state-of-the-art results on the Berkeley Function Calling Leaderboard.
Sources
- GPT-4 Technical Report
- Can a Single Model Master Both Multi-turn Conversations and Tool Use? CoALM: A Unified Conversational Agentic Language Model
- Breaking Language Barriers in Multilingual Mathematical Reasoning: Insights and Observations
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- The FLORES-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation
- Training Compute-Optimal Large Language Models
- Model Stock: All we need is just a few fine-tuned models
- Soup to go: mitigating forgetting during continual learning with model averaging
- Model Merging in Pre-training of Large Language Models
- ToolACE: Winning the Points of LLM Function Calling
- APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
- Code Llama: Open Foundation Models for Code
- MAPO: Advancing Multilingual Reasoning through Multilingual Alignment-as-Preference Optimization
- Language Models are Multilingual Chain-of-Thought Reasoners
- Massive Activations in Large Language Models
- Gemini: A Family of Highly Capable Multimodal Models
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- AdaMerging: Adaptive Model Merging for Multi-Task Learning
- The Super Weight in Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering