CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation

arXiv:2608.12842 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Yuchen Liu, Zongzhen Yang, Binhang Qi, Hailong Sun, Xiang Gao

Beihang University · Hangzhou Innovation Institute of Beihang University · National University of Singapore

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/raids-lab/crater

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 100/100

The gist: CABS+ is an enhanced model merging framework that extends the Conflict-Aware and Balanced Sparsification (CABS) method to address its limitations in time complexity, GPU memory consumption, and

Terminology

Summary

CABS+ is an enhanced model merging framework that extends the Conflict-Aware and Balanced Sparsification (CABS) method to address its limitations in time complexity, GPU memory consumption, and optimization bias. The paper states: "To address these limitations, we extend CABS and propose the enhanced method, termed CABS+. Specifically, the Adaptive Weight Allocation (AWA) strategy optimizes merging coefficients via a gradient-free search scheme to reduce time complexity, while the asymmetric fitness function promotes more comprehensive performance gains across tasks."

The core methodological contributions are twofold. First, CABS+ introduces the Adaptive Weight Allocation (AWA) strategy, which replaces the grid search used in CABS with a gradient-free optimization mechanism based on the Covariance Matrix Adaptation Evolution Strategy (CMA-ES). The paper explains: "Since this method does not require gradient computation or backpropagation, the GPU memory usage is strictly limited to the level of the inference stage, making it possible to merge billion parameter scale models on consumer-level GPUs such as the V100." AWA incorporates boundary constraints (with lower bound l=0.1 and upper bound u=2) and an asymmetric fitness evaluation function that applies a large penalty (α=100) to tasks whose loss has increased and a smaller penalty (β=1) to tasks whose loss has decreased, thereby mitigating the tendency of optimization to be dominated by tasks with larger loss scales. The paper notes: The final fitness function is obtained as the sum of the asymmetric scores across all tasks. AWA also benefits from the CABS pruning stage: Since task vectors have already undergone conflict reduction during the CABS pruning stage, the resulting optimization landscape for merging coefficients becomes smoother with reduced non-convexity.

Second, the paper conducts a systematic empirical investigation of model mergeability and proposes a new metric, the Relative Synergy Score (RSS), defined as: RSS = (Scoremerged − Scoreideal) / Scoreideal × 100%. This metric quantifies whether merging produces synergistic gains (positive RSS) or destructive interference (negative RSS). The empirical study identifies six critical factors affecting mergeability: learning rate, finetuning epoch, task heterogeneity, data distribution heterogeneity, model architecture, and model scale. Key findings include: lower learning rates lead to higher parameter overlap but non-monotonic merge performance; finetuning epochs exhibit an inverted-U pattern with optimal performance at 3 epochs; task heterogeneity shows a strong negative correlation with synergistic merge performance (RSS ranges from +0.73% for low heterogeneity to −14.06% for high heterogeneity); data distribution differences significantly reduce mergeability (RSS drops from +0.38% to −7.32% when merging models trained on similar vs. dissimilar datasets); encoder-only and decoder-only architectures respond differently to sparsification strategies; and larger models exhibit better mergeability (RSS improves from −1.06% at 7B scale to +1.28% at 70B scale).

The experimental evaluation covers 27 datasets and 5 models, including large language models (Mistral-7B-v0.1, Qwen-2.5-7B-Instruct), small-scale language models (RoBERTa, GPT-2), and vision models (ViT-B/32). Results show that CABS+ achieves significant performance improvements: Compared with AdaMerging and WUDIMerging, CABS+ achieves overall performance improvements of 16.97% and 12.93%, respectively. In efficiency comparisons, on Mistral models, CABS+ requires less than 25% of the GPU memory used by AdaMerging and achieves a nearly 4× speedup in merging time compared with WUDIMerging. Specifically, AdaMerging uses 66.70GB GPU memory on Mistral while CABS+ uses only 15.95GB, and WUDIMerging takes 4 hours while CABS+ takes 1 hour. On RoBERTa with 4 tasks, CABS+ achieves 82.42 average performance compared to 80.52 for AdaMerging and 81.79 for WUDIMerging, while using less GPU memory (3.77GB vs. 6.69GB and 5.10GB).

The paper also demonstrates robustness across varying task numbers and model architectures. On RoBERTa, CABS+ consistently outperforms baselines across 2, 4, and 6 task merging settings. On GPT-2 with 6 tasks, CABS+ achieves 68.71 average performance, improving over CABS by 1.56%. On the Open LLM Leaderboard 2 benchmark with Qwen-2.5-7B, CABS+ achieves 45.09 and 45.10 average performance under different merging orders, improving over CABS by 2.37% and 2.86% respectively. On the LLM Leaderboard with Mistral-7B, CABS+ achieves 76.71 average performance, slightly below WUDIMerging's 76.82 but above CABS's 76.50. On vision tasks with ViT-B/32, CABS+ achieves 82.50 average performance, improving over CABS's 81.08 and remaining competitive with WUDIMerging's 82.30.

The paper concludes: "This paper proposed CABS+, an enhanced model merging framework that extends CABS with the Adaptive Weight Allocation strategy. By replacing grid search with gradient-free coefficient optimization, CABS+ improves merging efficiency, reduces GPU memory consumption, and promotes more balanced performance improvements across tasks. Future work will explore multimodal and heterogeneous model merging to further broaden its applicability."

Improvements for AI systems

Improvements to AI Systems:

  1. Gradient-Free Model Merging with Consumer-Grade GPU Feasibility
  • Implement CABS+’s Adaptive Weight Allocation (AWA) using CMA-ES to merge billion-parameter models (e.g., 7B–70B) without backpropagation, reducing GPU memory to inference-level (e.g., 15.95GB vs. 66.70GB for AdaMerging on Mistral-7B).

  • Enable real-time merging of large language models on single V100 GPUs, allowing smaller labs or edge devices to combine specialized models (e.g., domain-tuned chatbots) without high-end hardware.

  1. Balanced Multi-Task Performance via Asymmetric Fitness
  • Adopt the asymmetric fitness function (penalty α=100 for regressed tasks, β=1 for improved tasks) to prevent optimization from being skewed by high-loss tasks.

  • Result: Merged models achieve more uniform performance across all tasks, avoiding catastrophic forgetting of easy tasks while still improving hard ones—critical for multi-task assistants (e.g., a model handling code, math, and dialogue without sacrificing any).

  1. Predictive Mergeability Assessment with Relative Synergy Score (RSS)
  • Use RSS = (Score merged − Score ideal)/Score ideal × 100% to pre-screen model pairs before merging, identifying synergistic (positive RSS) vs. destructive (negative RSS) combinations.

  • Automatically reject merges with high task heterogeneity (RSS drops to −14.06%) or dissimilar data distributions (−7.32%), saving compute and preventing performance degradation in production pipelines.

  1. Hyperparameter-Aware Merging for Optimal Configurations
  • Leverage the empirical findings (e.g., inverted-U for finetuning epochs, optimal at 3; lower learning rates improve parameter overlap) to automatically set finetuning and merging hyperparameters.

  • Improved system can self-tune: given a target task set, it recommends learning rate, epoch count, and sparsification ratio to maximize mergeability before training begins.

  1. Architecture-Specific Sparsification Adaptation
  • Apply CABS+’s pruning stage differently for encoder-only (e.g., RoBERTa) vs. decoder-only (e.g., GPT-2, Mistral) models, based on the paper’s finding that they respond differently to sparsification.

  • Result: Higher merged accuracy across heterogeneous model families (e.g., 82.42 on RoBERTa-4 tasks, 68.71 on GPT-2-6 tasks) without manual architecture tuning.

  1. Scalable Multi-Task Merging with Order Invariance
  • Use CABS+’s robustness to merging order (e.g., Qwen-2.5-7B achieves 45.09 and 45.10 under different orders) to build modular AI systems where new task models can be added incrementally without re-optimizing the entire merge.

  • Enables continuous learning: a system can merge a newly finetuned model into an existing ensemble in 1 hour (vs. 4 hours for WUDIMerging) with stable performance.

  1. Cross-Domain Synergy Detection for Vision-Language Models
  • Apply RSS and AWA to vision models (ViT-B/32) to merge CLIP-like encoders trained on different visual domains, achieving 82.50 average performance (vs. 81.08 for CABS) while maintaining low memory.

  • Improved system can combine a visual encoder for medical images with another for natural scenes, producing a unified model that excels in both domains without retraining.

  1. Efficient Benchmark-Driven Model Selection
  • Use CABS+’s 27-dataset evaluation to build a recommendation engine that predicts which base models and finetuning recipes yield the highest RSS for a given application.

  • The system can automatically select the best candidate models (e.g., Mistral-7B vs. Qwen-2.5-7B) for merging based on task overlap and heterogeneity, reducing trial-and-error.

Sources

Related papers