Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning
cs.AI
Submitted: 2025-04-14
Updated: 2026-08-25
Code: https://github.com/jincan333/MAS-TTS
License: http://creativecommons.org/licenses/by/4.0/
The gist: Test-Time Scaling has emerged as a powerful method to extend the reasoning capabilities of Large Language Models.
Terminology
Abstract
Test-Time Scaling has emerged as a powerful method to extend the reasoning capabilities of Large Language Models. However, single-agent TTS faces significant scalability bottlenecks, as excessively long reasoning traces lead to increased inference costs and stability issues caused by context management failures. To address these limitations, we propose leveraging Multi-Agent Systems as a structural upgrade to standard TTS. By decomposing monolithic reasoning chains into distinct, manageable contexts across multiple agents, MAS offers a more robust framework for scaling reasoning. We validate this approach by introducing M500, a dataset comprising 500 high-quality multi-agent, multi-turn collaborative reasoning traces generated via DeepSeek-R1. Through Supervised Fine-Tuning on M500, we enable open-source models to internalize collaborative reasoning patterns and show improved TTS performance in MAS. Furthermore, we propose an adaptive scaling strategy incorporating a ``CEO'' agent to dynamically guide the reasoning process and optimize collaboration depth. Extensive experiments within the AgentVerse framework demonstrate that our fine-tuned models, Qwen2.5-32B-MAS and Phi4-14B-MAS, significantly outperform their base counterparts. Codes are available at https://github.com/jincan333/MAS-TTS.
Sources
- GPT-4 Technical Report
- Critique-out-Loud Reward Models
- Program Synthesis with Large Language Models
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- The Role of Deductive and Inductive Reasoning in Large Language Models
- Evaluating Large Language Models Trained on Code
- AI-GAs: AI-generating algorithms, an alternate paradigm for producing general artificial intelligence
- When One LLM Drools, Multi-LLM Collaboration Rules
- Interpretable Contrastive Monte Carlo Tree Search Reasoning
- Think before you speak: Training Language Models With Pause Tokens
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Node-as-Agent: Graph Agentic Network
- SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning
- You Only Fine-tune Once: Many-Shot In-Context Fine-Tuning for Large Language Models
- T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
- Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model
- Qwen2.5-Coder Technical Report
- GPT-4o System Card
- Rewarding Chatbots for Real-World Engagement with Millions of Users
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection