HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments
cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/YangHaolin0526/HarnessSQL
Terminology
Sources
- FireAct: Toward Language Agent Fine-tuning
- ReFoRCE: A Text-to-SQL Agent with Self-Refinement, Consensus Enforcement, and Column Exploration
- BIRD-INTERACT: Re-imagining Text-to-SQL Evaluation for Large Language Models via Lens of Dynamic Interactions
- SWE-SQL: Illuminating LLM Pathways to Solve User SQL Issues in Real-World Applications
- ToolACE: Winning the Points of LLM Function Calling
- AgentBench: Evaluating LLMs as Agents
- APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
- AgentInstruct: Toward Generative Teaching with Agentic Flows
- FlexSQL: Flexible Exploration and Execution Make Better Text-to-SQL Agents
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- CHESS: Contextual Harnessing for Efficient SQL Synthesis
- ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
- Executable Code Actions Elicit Better LLM Agents
- Does Synthetic Layered Design Data Benefit Layered Design Decomposition?
- MTIR-SQL: Multi-turn Tool-Integrated Reasoning Reinforcement Learning for Text-to-SQL
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
- AgentTuning: Enabling Generalized Agent Abilities for LLMs
- Group Sequence Policy Optimization
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering