Toolcompass: Guiding Tool Trialing, Not Suppressing It
cs.AI, cs.LG
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/StonyBrookNLP/appworld
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model (LLM) agents must generalize from tools seen during training to unseen tools at deployment.
Terminology
Abstract
Large language model (LLM) agents must generalize from tools seen during training to unseen tools at deployment. A key challenge is tool trialing, i.e., excessive trials waste the interaction budget, whereas selective trials enable exploration of unfamiliar tools. Existing outcome-based post-training leaves wasteful trials unguided, while turn-level supervision may suppress necessary exploration. We introduce ToolCompass, a post-training framework that guides tool trialing by organizing tool-call representations according to shared functions. Specifically, ToolCompass models each function class as a von Mises--Fisher distribution and jointly reduces intra-function variation across domains and increases inter-function separation. This structure transfers experience from seen tools to functionally similar unseen tools, directing exploration away from unrelated alternatives. ToolCompass requires no ground-truth call traces or unseen-tool access and incurs no inference overhead. Experiments on AppWorld and FTRL show consistent gains across GRPO, RFT, and DMPO. improves AppWorld OOD task success by up to 10.71 percentage points over vanilla post-training and performs best among competitive baselines on both benchmarks.
Sources
- Reinforcement Learning for Long-Horizon Interactive LLM Agents
- Teaching LLMs to Learn Tool Trialing and Execution through Environment Interaction
- Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Cyclical Entropy Eruption: Entropy Dynamics in Agent Reinforcement Learning
- Budget-Aware Tool-Use Enables Effective Agent Scaling
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use
- Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
- TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents
- Chain-of-Tools: Utilizing Massive Unseen Tools in the CoT Reasoning of Frozen Language Models
- Can RL Improve Generalization of LLM Agents? An Empirical Study
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
- GLM-5: from Vibe Coding to Agentic Engineering
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection