SFT or RL for Tool-Calling Agents? A Controlled Study Across Data, Method, and Scale
cs.CL
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: Accepted to the REALM Workshop at EMNLP 2026
Code: https://github.com/talkiq/dialpad-ai-research
License: http://creativecommons.org/licenses/by/4.0/
The gist: Limited controlled evidence exists on how training data, adaptation method, and model scale jointly affect tool-calling performance in language-model agents.
Terminology
Abstract
Limited controlled evidence exists on how training data, adaptation method, and model scale jointly affect tool-calling performance in language-model agents. We evaluate supervised fine-tuning (SFT) with LoRA, reinforcement learning (RL) via Group Relative Policy Optimization (GRPO), and SFT followed by GRPO across six Qwen3 models from 0.6B to 32B parameters, covering both in-distribution performance and cross-dataset transfer. SFT with LoRA is the strongest in-distribution method throughout the 0.6B-32B range and best in 15 out of 18 experimental settings. On cross-dataset transfer, the methods are closer: GRPO wins 29 out of 54 settings where training and test datasets differ, but its margin over SFT averages under one point, and SFT->GRPO is rarely strongest in either comparison. Dataset mixing gives consistently strong transfer while staying close to specialized in-distribution training, regardless of method. Additional analysis further confirms that LoRA outperforms full-parameter fine-tuning, demonstrating that LoRA better preserves pretrained agentic behavior.
Sources
- Small Language Models are the Future of Agentic AI
- LoRA Learns Less and Forgets Less
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
- APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
- ToolRL: Reward is All Tool Learning Needs
- LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- Qwen3 Technical Report
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- ReAct: Synergizing Reasoning and Acting in Language Models
- Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering