Small Language Models are the Future of Agentic AI
cs.AI
Submitted: 2025-06-02
Updated: 2026-09-22
Code: https://github.com/ai-dynamo/dynamo
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) are often praised for exhibiting near-human performance on a wide range of tasks and valued for their ability to hold a general conversation.
Terminology
Abstract
Large language models (LLMs) are often praised for exhibiting near-human performance on a wide range of tasks and valued for their ability to hold a general conversation. The rise of agentic AI systems is, however, ushering in a mass of applications in which language models perform a small number of specialized tasks repetitively and with little variation. Here we lay out the position that small language models (SLMs) are sufficiently powerful, inherently more suitable, and necessarily more economical for many invocations in agentic systems, and are therefore the future of agentic AI. Our argumentation is grounded in the current level of capabilities exhibited by SLMs, the common architectures of agentic systems, and the economy of LM deployment. We further argue that in situations where general-purpose conversational abilities are essential, heterogeneous agentic systems (i.e., agents invoking multiple different models) are the natural choice. We discuss the potential barriers for the adoption of SLMs in agentic systems and outline a general LLM-to-SLM agent conversion algorithm. Our position, formulated as a value statement, highlights the significance of the operational and economic impact even a partial shift from LLMs to SLMs is to have on the AI agent industry. We aim to stimulate the discussion on the effective use of AI resources and hope to advance the efforts to lower the costs of AI of the present day. Calling for both contributions to and critique of our position, we commit to publishing all such correspondence at https://research.nvidia.com/labs/lpr/slm-agents.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- DELIFT: Data Efficient Language model Instruction Fine Tuning
- Minifinetuning: Low-Data Generation Domain Adaptation through Corrective Self-Distillation
- Tiny Transformers Excel at Sentence Compression
- Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
- Improving language models by retrieving from trillions of tokens
- Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training
- Hymba: A Hybrid-head Architecture for Small Language Models
- Text Compression for Efficient Language Generation
- Scaling Laws for Transfer
- Training Compute-Optimal Large Language Models
- LoRA: Low-Rank Adaptation of Large Language Models
- MatFormer: Nested Transformer for Elastic Inference
- DoRA: Weight-Decomposed Low-Rank Adaptation
- Large Model Strategic Thinking, Small Model Efficiency: Transferring Theory of Mind in Large Language Models
- Small Language Models: Survey, Measurements, and Insights
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey
- Agentic AI Needs a Systems Theory
- A Comprehensive Overview of Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection