Agent-First Tool API: A Semantic Interface Paradigm for Enterprise AI Agent Systems

arXiv:2605.10555 · cs.AI · Submitted 2026-08-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Agent-First Tool API: A Semantic Interface Paradigm for Enterprise AI Agent Systems".

Jane: The paper was written by Kai Pan and Rong Hou from A2A Lab.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title and Authors: Tom: We're looking at this massive paper by Kai Pan and Rong Hou titled "Agent-First Tool API: Rethinking Enterprise Service Interfaces for LLM-Native Execution," and it feels like we're witnessing a fundamental shift in how software is designed. Before, everything was built for a human clicking buttons, but this paper argues that the AI agents are becoming the primary users.

Jane: That’s absolutely right, Tom; they aren't just tools anymore. The authors are suggesting that our current APIs—the ones we call CRUD interfaces—are inherently mismatched for LLM planners because those UIs expect precise inputs and return raw data, not a guided planning experience.

Lu: I find the title itself fascinating because it implies that the AI isn't just plugging into existing tools; it's demanding a complete redesign of the interface to support its own reasoning process. It’s a demand for semantic intelligence in the API structure.

Meng: From a practical standpoint, this is huge because many companies are struggling with "brittle prompt engineering" where they have to constantly patch LLMs to handle simple tasks; this paper is proposing a solution that makes those ad-hoc workarounds unnecessary.

Lalam: When I process the idea of "LLM-Native Execution," I see a future where AI agents are treated not as mere execution engines, but as sophisticated decision-makers who need the right context and justification for every action they take, which is exactly what this paper seems to provide.

Tom: It's clear that Pan and Hou aren't just patching an old system; they're building a whole new framework where the tools are designed specifically to support the core reasoning loop of LLMs, which is a massive undertaking.

Jane: And it doesn’s just about the authors too; their work shows how this paradigm can be applied across eighty-five different tools in a complex multi-tenant system, proving that it' isn't just a theoretical concept but something scalable.

Lu: I think the paper suggests that if we want trustworthy AI, we need to stop designing systems for human eyes and start designing them for LLM brains. It forces us to rethink what "the interface" really means in modern architecture.

Meng: That’s the core message: moving beyond just how a tool is called, and focusing on how the entire interaction loop—from intention to verification—is structured and supported by making the the service itself is highly important for a scalable solution.

Lalam: I think this work opens up possibilities where AI can perform complex enterprise tasks not just with competence, but with demonstrable accountability right at the semantic level.

Summary and Contributions: Tom: So, we've established the need for this paradigm; now let's look at what Pan and Hou actually propose as their three core contributions to solve those original architectural mismatches. The paper "Agent-First Tool API" is packed with specific technical ideas that are really innovative.

Jane: They’re introducing a Six-Verb Semantic Protocol, which they are defining as a structured reasoning loop. This means instead of just one call, the every task follows a sequence: semantic search, resolve candidates, preview action, execute action, verify result, and recover from error.

Tom: And that’s not just some abstract list of verbs; it's a formal Finite State Machine that ensures the agent can successfully navigate its workflow or handle failure in a specific way. It's designed to be robust.

Lu: I love the idea of "resolve candidates" being a step, too. It acknowledges that an LLM agent doesn’t just know exactly what it means; it needs to disambiguate real-world entities based on the relevance scores provided by the tool, making its planning more intelligent.

Meng: The second big piece is the Normalized Tool Contract, or NTC. This is a complete paradigm shift because it mandates that every single response from an agency tool must include confidence scores and evidential provenance.

Jane: The NTC truly augments the response to provide decision support, which means the LLM can see not just *what* happened, but *how sure* the system is about what happened before making its next choice.

Lalam: From a cultural perspective, this creates a mechanism for trust; it compels AI to explain its reasoning and justify its actions with verifiable evidence, which is crucial when we' are delegating high-stakes enterprise tasks to automated systems.

Tom: And finally, there’s the third contribution: this dual-layer governance pipeline. This handles permissions not just based on who is using the tool, but also based on what the tool is even allowed to do—capability versus object scope.

Lu: That separation of capacity from context is a critical architectural insight, ensuring that even if someone has high privileges, they can't operate outside their specific organizational boundaries.

Meng: This governance is highly practical because it means we don't have to bolt on extra layers; the security and workflow logic are built right into the tool itself.

Jane: It’s all working together in this paper, making sure that every step—from deciding which tool to use to verifying the result—is guided by structured data and strong safety checks.

Improvements and Mechanisms: Tom: We’ve covered the framework; now let's dig into the mechanics. How does this "Agent-First Tool API" actually improve upon a traditional setup, especially regarding how it handles inputs and risk? The paper offers some very concrete improvements.

Jane: One huge improvement is that the tool accepts descriptive natural language input instead of demanding exact IDs. This shifts the resolution burden to deterministic backend logic, which is far more reliable than relying on an LLM’s context window to figure out what a human meant.

Tom: And if it's not sure, it doesn't just fail silently; the tool returns a disambiguation prompt, which is way smarter than just throwing a four hundred four error. It guides the agent toward the correct choice instead of making it give up.

Lu: The governance part is also deeply improved through this dual-layer model; we aren't just checking if someone has permission, we are checking if they have permission *and* if that operation respects the boundaries of their assigned organizational scope.

Meng: And on top of that, the system uses dynamic risk escalation. If an operation is low risk but becomes a high-risk batch operation, the required approval level changes automatically—it scales with the action.

Jane: That's a huge leap forward in safety; we’re moving past static permissions to a system that reacts to specific operational conditions before taking any action.

Tom: And if this dynamic risk gets too high, the built-in approval gate triggers an automatic suspension and workflow, which is a powerful way to prevent catastrophic mistakes without needing external orchestration.

Lu: I'm particularly impressed by the formal verification of the Six-Verb Protocol; that gives us a mathematical guarantee that no task can ever get stuck in an unrecoverable state, ensuring system integrity.

Meng: The design also mandates idempotency for write operations, which is essential because when AI retries a failed call, it might otherwise create duplicate records—this prevents data corruption automatically.

Lalam: Seeing the integration of these safety mechanisms—from semantic search to dynamic risk checks—makes me feel that AI is being given the tools not just to perform tasks, but to perform them responsibly and with demonstrable internal oversight.

Conclusion and Wrap-up: Tom: We've seen how this new paradigm solves the fundamental mismatch between human-driven APIs and LLM planners, especially in a complex enterprise environment. The paper "Agent-First Tool API" is essentially proposing a new contract for trust between the incredibly powerful AI and the reliable backend systems.

Jane: The experimental results were quite striking; they demonstrated that by moving from this traditional CRUD approach, we saw a significant increase in task success rate and much lower rates of hallucination errors.

Tom: But Meng mentioned earlier that this requires a dual API approach, which has its own necessary complexity and overhead. It’s not just a quick fix or an easy patch; it demands careful infrastructure investment.

Meng: It is an investment in infrastructure, absolutely; but the cost of that maintenance is offset by the massive reduction in human intervention required to fix AI mistakes. The reliability gained justifies the engineering effort for us at this scale.

Lu: We should also look at how this relates to future work—specifically how we can formally verify tool composition safety, ensuring that when we chain many tools together, the overall system invariants are preserved.

Lalam: This is about establishing a new standard of trust in the AI; that this level of accountability is what's needed to move forward with truly autonomous decision-making at scale.

Tom: That’s a powerful thought, Lalam; it moves us away from just trusting the LLM and toward trusting the system design itself as a guarantee of correctness.

Jane: So, as we wrap up our discussion on "Agent-First Tool API: A Semantic Interface Paradigm for Enterprise AI Agent Systems," we have a very clear picture of how to bridge that gap between human-driven APIs and autonomous agents.

Lu: I can't wait to see how this translates into more complex interactions across different systems; the possibilities for chaining these are vast.

Meng: Hopefully, we can implement this at scale while managing that infrastructure overhead efficiently across different corporate stacks.

Lalam: I'm excited to see a culture where the AI is not seen as a black box, but as a rigorously audited and accountable agent in production environments.

Kai Pan, Rong Hou

A2A Lab

cs.AI

Submitted: 2026-08-20

Updated: 2026-08-21

Code: https://github.com/langchain-ai/langchain

Importance score: 85/100

The gist: * Summary of Agent-First Tool API The paper addresses a critical design mismatch in enterprise software architecture where Large Language Model (LLM) agents are increasingly acting as primary

Key concepts

Six-Verb Semantic Protocol
A structured reasoning loop defining how an agent executes tasks. It mandates a sequence of steps: semantic search, resolve candidates, preview action, execute action, verify result, and recover from error.
Normalized Tool Contract (NTC)
A requirement that every agency tool response must include confidence scores and evidential provenance. This allows the LLM to understand not only what happened but also how certain the system is about the outcome.
Dual-Layer Governance Pipeline
A security mechanism that governs access based on two factors: who is using the tool (capability) and what the tool is allowed to do within specific boundaries (object scope).
LLM-Native Execution
The concept of designing systems where AI agents are treated as sophisticated decision-makers. This requires APIs that support complex reasoning and context, rather than just simple data input/output.

Terminology

Summary

Summary of Agent-First Tool API

The paper addresses a critical design mismatch in enterprise software architecture where Large Language Model (LLM) agents are increasingly acting as primary operators, yet the APIs they invoke remain designed for human-driven User Interfaces (UIs). Traditional CRUD interfaces are fundamentally incompatible with autonomous agentic execution, operating under five assumptions that fail when the caller is an LLM planner:

  1. Assumption 1: The caller knows exact parameters.

  2. Assumption 2: Responses are renderable data (requiring confidence scores and provenance).

  3. Assumption 3: One request suffices (agents operate in multi-turn reasoning loops).

  4. Assumption 4: Permissions equal user permissions (a second layer is needed for tool capability).

  5. Assumption 5: Errors are HTTP status codes (agents require structured failure descriptions and recovery strategies).

To resolve this paradigm gap, the authors propose the Agent-First Tool API design paradigm, which reconceives tool interfaces as goal-achievement protocols rather than form-submission endpoints. This paradigm introduces three core contributions:

** 1. The Six-Verb Semantic Protocol**

The paper defines a structured reasoning loop for every tool interaction, formalized in a Finite State Machine (FSM) and a sequence of six semantic verbs. This protocol structures the interaction as:

(T, q) = S, R, P, E, V, C

Where the phases are defined as follows:

  • Phase 1: Semantic Search (S): Given natural-language intent q, the tool returns a ranked candidate set C = (c i, s i) where s i is a relevance score.

  • Phase 2: Resolve Candidates (R): If multiple candidates exist, the tool provides a disambiguation prompt.

  • Phase 3: Preview Action (P): Before state mutation, the tool produces a dry-run summary showing what would change (preview).

  • Phase 4: Execute Action (E):: The tool applies the mutation within a transactional boundary. Write-mode tools enforce idempotency via caller-supplied idempotency keys.

  • Phase 5: Verify Result (V): Post-execution, the tool rereads the affected state and compares it against the agent’s original intent, returning a match: bool and evidence.

  • Phase 6: Recover from Error (C): error. On failure at any phase, the tool returns structured recovery guidance.

** 2. Normalized Tool Contract (NTC)**

The NTC standardizes every tool response into a decision-support structure that an LLM planner can reason over. The NTC augments output with:

  • Confidence: A float in [0, 1] representing the tool author's prior knowledge (c static). This is combined with empirical success rate (w) using a sliding-window calibration mechanism: c calibrated = alpha times c static + (1 - alpha) times w.

  • Evidential Provenance: An array of evidence detailing the justification for the result.

  • Action Continuity: A next actions array suggesting logical follow-up tools, which helps reducing the planner’s search space and preventing dead-end tool sequences.

** 3. Dual-Layer Governance Pipeline**

To manage autonomous execution, a six-layer validation pipeline is implemented:

  1. Schema Validation (Input parameters checked against JSON Schema).

  2. Capability Permission (RBAC check).

  3. Object Scope (Enforcing a four-level hierarchy: Tenant Brand Store User).

  4. Dynamic Risk Assessment.

  5. Approval Gate (If the assessed risk exceeds the autoexecute threshold, execution is suspended and requires human approval).

  6. Execution within a transactional boundary.

The dual-layer permission model separates what can I do (capability) from what can I see (object scope), ensuring that tool discovery is tenant-scoped while execution is filtered by the full organizational hierarchy.

** Technical Implementation and Evaluation**

The paradigm was validated in a production SaaS system spanning 85 tools across 6 business domains. The evaluation involved a controlled experiment comparing the Agent-First approach against a baseline (CRUD+ReAct).

Key findings from this comparison include:

  • A 37.5% improvement in task success rate for the Agent-First paradigm.

  • A 5.8× improvement in successful error recovery.

  • A significant reduction in human intervention, dropping from 22% (CRUD) to 6% (Agent-First).

The system architecture utilizes a dual API surface: the traditional CRUD API serves the human dashboard, while the Agent Tool API serves the LLM planner exclusively. The tool names follow a verb.noun convention (e.g., ticket.search) which is used for both provide semantic clarity to the LLM and enabling wildcard capability rules for governance.

The paper concludes that while the Agent-First Tool API introduces measurable overhead (e.g, 300% increase in token count per response), this cost is offset by a reduction in exploratory retry loops, leading to a net decrease in task-level latency and significantly enhanced reliability—making it a prerequisite for trustworthy agentic AI in production enterprise environments.

Improvements for AI systems

The Agent-First Tool API paradigm fundamentally shifts the design focus from CRUD interaction (human-centric form submission) to Goal-Achievement Protocol (AI-native reasoning). The following improvements represent the necessary architectural changes:

Instead of treating tools as atomic functions, they are exposed as sequences of semantic verbs that guide the LLM planner through a structured, multi-turn reasoning process.

What the Improved AI System Can Do:

  • Handle Ambiguity Gracefully (Semantic Search & Resolve Candidates): The system no longer fails when encountering vague or descriptive input (e.g., the downtown branch). It executes semantic search to retrieve a ranked candidate set, and if ambiguity persists, it invokes resolve candidates, providing the LLM with a structured prompt for disambiguation.

  • Execute Non-Atomic Tasks (Six-Verb Protocol): Complex operations like ticket creation are broken down into sequential phases: (S) Semantic Search to (R) Resolve Candidates to (P) Preview Action to (E) Execute Action to (V) Verify Result to (C) Recover from Error. The AI system autonomously flows through these phases, ensuring the task is not stalled or hallucinated.

The response structure is standardized to include decision-support metadata, allowing the LLM planner to make informed decisions based on evidence, rather than merely parsing raw JSON data.

Authorization is separated into capability checks (what is allowed) and object scoping (what can be acted upon), ensuring fine-grained control over autonomous operations.

The design mandates that complex, state-mutating tools are treated as transactional processes with built-in safety checks.


The adoption of this paradigm allows the AI system to achieve:

  • 37.5% Increase in Task Success Rate: By eliminating ambiguity-induced failures (e.g., incorrect ID resolution).

  • Drastic Reduction in Human Intervention: Moving from 22% required human intervention (CRUD) to only 6% (Agent-First), as the system handles errors and ambiguities autonomously.

  • Structured Error Recovery: Instead of opaque HTTP status codes, the system provides actionable, structured guidance (cause, candidates, suggestion) when a tool fails, enabling self-healing behavior.

Sources

Related papers