ACE: Adapter Consolidation across Experts for Parameter-Efficient Fine-Tuning of MoE LLMs

arXiv:2609.06072 · cs.LG, cs.AI, cs.CL · Submitted 2026-09-05 · Read on arXiv

cs.LG, cs.AI, cs.CL

Submitted: 2026-09-05

Updated: 2026-09-05

Comments: 23 pages, 13 figures. Accepted to EMNLP 2026

Code: https://github.com/UbiquitousAILab/ACE

License: http://creativecommons.org/licenses/by/4.0/

The gist: Parameter-efficient fine-tuning (PEFT) of mixture-of-experts (MoE) models commonly attaches a separate low-rank adapter to each expert.

Terminology

Abstract

Parameter-efficient fine-tuning (PEFT) of mixture-of-experts (MoE) models commonly attaches a separate low-rank adapter to each expert. This expert-wise design fragments adaptation in three ways: capacity is split across narrow low-rank updates, gradient supervision becomes sparse and imbalanced under sparse routing, and execution is decomposed into many small GEMMs. We find that such expert-wise separation is often unnecessary, as subsets of LoRA adapters become functionally similar during fine-tuning, revealing redundancy among expert-specific adapters. Based on this redundancy, we propose ACE (Adapter Consolidation across Experts), which groups redundant experts and replaces their expert-specific adapters with group-shared higher-rank LoRA modules under the same PEFT budget. ACE further introduces grouped adapter execution, which consolidates fragmented expert-wise adapter computations into fewer, larger group-level GEMMs. Across evaluations covering 12 datasets and four MoE backbones, ACE achieves the highest observed mean accuracy among the parameter-matched PEFT methods on the three backbones with complete baseline coverage, while providing 1.31 times to 1.48 times wall-clock training speedup over expert-wise LoRA without increasing peak memory. Our code is available at https://github.com/UbiquitousAILab/ACE.

Related papers