SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks
cs.CL
Submitted: 2026-06-14
Updated: 2026-08-28
Code: https://github.com/llexieguo/SciOrch
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Frontier scientific reasoning remains a major challenge for large language models (LLMs), where even the strongest commercial systems fall short of expert-level performance.
Terminology
Abstract
Frontier scientific reasoning remains a major challenge for large language models (LLMs), where even the strongest commercial systems fall short of expert-level performance. A closer look at model behavior reveals substantial complementarity that single-model evaluation hides: different frontier models excel on different question types, and no single model captures the full picture. We present SciOrch, a framework that trains a lightweight 8B model to orchestrate frontier LLMs for scientific reasoning. The orchestrator decomposes each question, delegates sub-problems to selected commercial models through API calls, and synthesizes a final answer. Training such an orchestrator is fundamentally harder than conventional agentic RL: each action triggers an API call that is expensive in both dollar cost and latency, making standard online rollouts infeasible. We address this with MCTS-based approach, producing diverse orchestration trajectories, extracting per-node single-turn samples, and optimizing the orchestrator with GRPO-style training. On a 240-question test set spanning SGI-Reasoning and Scientists' First Exam, SciOrch reaches 56.66% average accuracy, outperforming the strongest single commercial model by 3.74% and the strongest multi-agent baseline by 3.33%. It also attains the best accuracy on both SGI and SFE with less than half the API cost of typical multi-agent methods.
Sources
- The Impact of Large Language Models on Scientific Discovery: a Preliminary Study using GPT-4
- Qwen3-VL Technical Report
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models
- Prompt-to-Leaderboard
- Accelerating scientific discovery with Co-Scientist
- Large Language Model based Multi-Agents: A Survey of Progress and Challenges
- REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
- Automated Design of Agentic Systems
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
- Are LLMs Ready for Real-World Materials Discovery?
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration
- AgentSquare: Automatic LLM Agent Search in Modular Design Space
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering