DualSQL: Text-to-SQL with Multi-Agent Reinforcement Learning
cs.CL, cs.AI, cs.DB
Submitted: 2026-09-16
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: State-of-the-art Text-to-SQL systems are typically multi-agent pipelines centered around two fundamental tasks: schema linking and SQL generation.
Terminology
Abstract
State-of-the-art Text-to-SQL systems are typically multi-agent pipelines centered around two fundamental tasks: schema linking and SQL generation. However, existing work trains separate models for each task, failing to leverage the synergy between these interrelated tasks. In this work, we propose DualSQL, a new Text-to-SQL system consisting of two agents powered by a single model backbone. The agents share the same model weights and agentic scaffold, enabling joint optimization through a robust multi-agent reinforcement learning (RL) framework. We design three database access tools to facilitate effective multi-step reasoning grounded to interactions with the databases. To improve training and avoid model collapse, we introduce a set of rollout guardrail mechanisms that stabilizes multi-agent RL training, supporting DualSQL to keep improving during training. We also introduce a new SQL correctness metric, robust execution match (REX), to more accurately judge SQL correctness and assign reward signals. Being trained on only 3755 examples, DualSQL-4B achieves an impressive 68.0% execution accuracy on the BIRD development set, matching previous 7B models. DualSQL-8B further improves to 71.1%, outperforming previous state-of-the-art single-model solutions with 32B parameters. These results demonstrate the strength of joint multi-agent reinforcement learning for building high performance Text-to-SQL pipelines.
Sources
- Evaluating Long-Context Reasoning in LLM-Based WebAgents
- Agentic Reinforced Policy Optimization
- Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
- Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- The Dawn of Natural Language to SQL: Are We Fully Ready?
- AgentSM: Semantic Memory for Agentic Text-to-SQL
- LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
- Automatic Metadata Extraction for Text-to-SQL
- XiYan-SQL: A Novel Multi-Generator Framework For Text-to-SQL
- TableGPT2: A Large Multimodal Model with Tabular Data Integration
- CHESS: Contextual Harnessing for Efficient SQL Synthesis
- DB-Explore: Automated Database Exploration and Instruction Synthesis for Text-to-SQL
- SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement Learning
- SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
- Agentar-Scale-SQL: Advancing Text-to-SQL through Orchestrated Test-Time Scaling
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
- Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering