AdaFuse: Adaptive Ensemble Decoding with Test-Time Scaling for LLMs
cs.CL, cs.AI
Submitted: 2026-01-09
Updated: 2026-08-30
Comments: ACL 2026 Main
Code: https://github.com/CCM0111/AdaFuse
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) exhibit complementary strengths arising from differences in pretraining data, model architectures, and decoding behaviors.
Terminology
Abstract
Large language models (LLMs) exhibit complementary strengths arising from differences in pretraining data, model architectures, and decoding behaviors. Inference-time ensembling provides a practical way to combine these capabilities without retraining. However, existing ensemble approaches suffer from fundamental limitations. Most rely on fixed fusion granularity, which lacks the flexibility required for mid-generation adaptation and fails to adapt to different generation characteristics across tasks. To address these challenges, we propose AdaFuse, an adaptive ensemble decoding framework that dynamically selects semantically appropriate fusion units during generation. Rather than committing to a fixed granularity, AdaFuse adjusts fusion behavior on the fly based on the decoding context, with words serving as basic building blocks for alignment. To be specific, we introduce an uncertainty-based criterion to decide whether to apply ensembling at each decoding step. Under confident decoding states, the model continues generation directly. In less certain states, AdaFuse invokes a diversity-aware scaling strategy to explore alternative candidate continuations and inform ensemble decisions. This design establishes a synergistic interaction between adaptive ensembling and test-time scaling, where ensemble decisions guide targeted exploration, and the resulting diversity in turn strengthens ensemble quality. Experiments on open-domain question answering, arithmetic reasoning, and machine translation demonstrate that AdaFuse consistently outperforms strong ensemble baselines, achieving an average relative improvement of 6.88%. The code is available at https://github.com/CCM0111/AdaFuse.
Sources
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Measuring Massive Multitask Language Understanding
- ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration
- RouterBench: A Benchmark for Multi-LLM Routing System
- InternLM2 Technical Report
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- WAPITI: A Watermark for Finetuned Open-Source LLMs
- Training Verifiers to Solve Math Word Problems
- RLAE: Reinforcement Learning-Assisted Ensemble for LLMs
- Mistral 7B
- LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
- Fusing Models with Complementary Expertise
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Life-Cycle Routing Vulnerabilities of LLM Router
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- Cool-Fusion: Fuse Large Language Models without Training
- Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models
- SpecFuse: Ensembling Large Language Models via Next-Segment Prediction
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering