Learning to Configure Agentic AI Systems
cs.AI
Submitted: 2026-02-12
Updated: 2026-09-05
Comments: 22 pages, 12 figures
Code: https://github.com/langchain-ai/langchain
License: http://creativecommons.org/licenses/by/4.0/
The gist: Configuring LLM-based agent systems involves choosing workflows, tools, token budgets, and prompts from a large combinatorial design space, and is typically handled today by fixed templates or
Terminology
Abstract
Configuring LLM-based agent systems involves choosing workflows, tools, token budgets, and prompts from a large combinatorial design space, and is typically handled today by fixed templates or hand-tuned heuristics that apply the same configuration regardless of query difficulty, leading to brittle behavior and wasted compute. To address this, we formulate agent configuration as a semi-Markov decision process (SMDP) where each configuration acts as a temporally extended option that determines how an agent system processes a query, and introduce introduce ARC (Agentic Resource & Configuration learner), a lightweight hierarchical policy that dynamically selects query-specific agent configurations. Across reasoning, tool-use, and agentic benchmarks, ARC consistently improves over budget-matched tool-augmented LLMs, increasing average reasoning accuracy by 31.3%, tool-use accuracy by 13.95%, and doubling τ-Bench (Airline) Pass 1 success from 9.0% to 18.0%. These results demonstrate that learning per-query agent configurations is a powerful alternative to "one size fits all" designs.
Sources
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- Training Verifiers to Solve Math Word Problems
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs
- LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- AgentBench: Evaluating LLMs as Agents
- Iterative Reasoning Preference Optimization
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- Proximal Policy Optimization Algorithms
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- Qwen2.5 Technical Report
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- WebArena: A Realistic Web Environment for Building Autonomous Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection