Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives
cs.SE, cs.AI, cs.CL, cs.LG, cs.MA
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/tatsu-lab/stanford_alpaca
Terminology
Sources
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- Hammer: Robust Function-Calling for On-Device Language Models via Function Masking
- ToolACE: Winning the Points of LLM Function Calling
- APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- LongFuncEval: Measuring the effectiveness of long context models for function calling
- Agent Primitives: Reusable Latent Building Blocks for Multi-Agent Systems
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- GPT-4 Technical Report
- The Llama 3 Herd of Models
- Mixtral of Experts
- DeepSeek-V3 Technical Report
- GPT-4o System Card
- Qwen2.5 Technical Report
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Qwen3 Technical Report
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Prompt Injection Attack to Tool Selection in LLM Agents
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties