UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents
cs.AI
Submitted: 2026-04-13
Updated: 2026-09-02
Comments: 25 pages, 10 figures, 17 tables. Code and datasets are publicly available at: https://github.com/EIT-NLP/UniToolCall
Code: https://github.com/EIT-NLP/UniToolCall
License: http://creativecommons.org/licenses/by/4.0/
The gist: Tool-use capability is a fundamental component of LLM agents, enabling them to interact with external systems through structured function calls.
Terminology
Abstract
Tool-use capability is a fundamental component of LLM agents, enabling them to interact with external systems through structured function calls. However, existing research exhibits inconsistent interaction representations, largely overlooks the structural distribution of tool-use trajectories, and relies on incompatible evaluation benchmarks. We present UniToolCall, a unified framework for tool learning that standardizes the entire pipeline from toolset construction and dataset generation to evaluation. The framework curates a large tool pool of 22k+ tools and constructs a hybrid training corpus of 390k+ instances by combining 10 standardized public datasets with structurally controlled synthetic trajectories. It explicitly models diverse interaction patterns, including single-hop vs. multi-hop and single-turn vs. multi-turn, while capturing both serial and parallel execution structures. To support coherent multi-turn reasoning, we further introduce an Anchor Linkage mechanism that enforces cross-turn dependencies. Furthermore, we convert 7 public benchmarks into a unified Query--Action--Observation--Answer (QAOA) representation with fine-grained evaluation at the function-call, turn, and conversation levels. Experiments show that fine-tuning Qwen3-8B on our dataset substantially improves tool-use performance. Under the distractor-heavy Hybrid-20 setting, achieves 93.0% single-turn Strict Precision, outperforming commercial models including GPT, Gemini, and Claude.
Sources
- ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
- ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Agent AI: Surveying the Horizons of Multimodal Interaction
- MCP-RADAR: A Multi-Dimensional Benchmark for Evaluating Tool Use Capabilities in Large Language Models
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- Efficient and Scalable Estimation of Tool Representations in Vector Space
- Advancing SLM Tool-Use Capability using Reinforcement Learning
- ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
- ASTRA: Automated Synthesis of agentic Trajectories and Reinforcement Arenas
- HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Device Scenarios
- Qwen3 Technical Report
- LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
- ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection