MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-17
Updated: 2026-09-17
License: http://creativecommons.org/licenses/by/4.0/
The gist: Tool learning enables large language models (LLMs) to use external tools for tasks beyond parametric knowledge.
Terminology
Abstract
Tool learning enables large language models (LLMs) to use external tools for tasks beyond parametric knowledge. Reinforcement learning can optimize tool-call behavior from feedback, but current methods still face two problems: fixed-threshold curricula can become misaligned with the policy's evolving capability boundary, and additive rewards can leak argument-level credit when the predicted tool is wrong. To address these problems, we propose MATCH, a closed-loop framework for model-aware tool learning with curriculum scheduling and hierarchically gated rewards. Model-Aware Curriculum Learning (MACL) maintains reward-derived sample difficulty that co-evolves with the policy, and each epoch selects samples near the current capability boundary together with a top-k pool of harder cases. Hierarchical Tool-call Gated Reward (HTGR) scores tool name, argument key, and argument value as a gated chain, granting credit at each level only when prerequisites hold. The same HTGR rewards drive both GRPO updates and MACL's difficulty refresh, closing the loop between policy optimization and sample scheduling. On API-Bank and BFCL V3, MATCH reaches 72.19% and 62.87% overall accuracy, outperforming the main supervised and RL-based baselines. Backbone experiments further show consistent improvements across four backbones from two model families.
Sources
- ToolSample: Dual Dynamic Sampling Methods with Curriculum Learning for RL-based Tool Learning
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- ToRL: Scaling Tool-Integrated RL
- Let's Verify Step by Step
- Understanding Tool-Integrated Reasoning
- ToolACE: Winning the Points of LLM Function Calling
- The Llama 3 Herd of Models
- GPT-4 Technical Report
- Gorilla: Large Language Model Connected with Massive APIs
- ToolRL: Reward is All Tool Learning Needs
- Online Difficulty Filtering for Reasoning Oriented Reinforcement Learning
- FireAct: Toward Language Agent Fine-tuning
- Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
- Tool Zero: Training Tool-Augmented LLMs via Pure RL from Scratch
- iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Use
- xLAM: A Family of Large Action Models to Empower AI Agent Systems
- Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
- Tool Learning with Foundation Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Acting Less is Reasoning More! Teaching Model to Act Efficiently
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks