Beyond Static Endpoints: Tool Programs as an Interface for Flexible Agentic Web Services
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Beyond Static Endpoints".
Tom: Tool programs are introduced as an executable representation of tool intent to address the representational bottleneck posed by static API endpoints in agentic web services.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, we’re talking about "Beyond Static Endpoints: Tool Programs as an Interface for Flexible Agentic Web Services," and the authors are Liu, Li, Zhang, Ma—a pretty solid group coming from different AI research backgrounds. The core idea they’re pushing is that the standard way of calling web services through static endpoints just doesn't handle complex tasks well when those tasks involve loops or conditional logic.
Jane: It seems like the authors are arguing that this current static endpoint setup creates a brittle interface because it forces the agent to do all the heavy lifting of sequencing and decision-making on its own side for every single step, which just adds latency.
Lu: That's right, and they lay out a very clear comparison in Figure one showing how stepwise endpoints lead to a sequence of one plus N plus one requests for N reasoning rounds, while the tool program approach aims for just one request and one reasoning round.
Meng: If we can consolidate that into a single execution unit, it drastically simplifies the orchestration layer on our end and reduces the overhead associated with managing those many sequential calls.
Lalam: I think this shift from a sequence of local decisions to an executable program interface is really exciting because it suggests a more structured way for agents to express complex procedures rather than just reactive prompts.
The paper's summary: Tom: Let's look at the summary of "Beyond Static Endpoints: Tool Programs as an Interface for Flexible Agentic Web Services." They basically introduce TOOLPRO, which is this executable representation of tool intent that compactly encodes those multi-step service interactions along with explicit effect types.
Jane: The main point there is that TOOLPRO combines a few key mechanisms: constraint-guided program construction to make sure the code actually runs, effect-aware replay to handle state changes safely during repairs, and a profile-driven policy to decide whether running the program or just calling things step by step makes more sense.
Lu: The effect-aware replay is particularly clever because it enforces exactly once semantics for any write operations across iterative repair attempts, which addresses the issue of state corruption that happens when things fail mid-execution.
Meng: So, they’re not just about making it run once; they are specifically designed to ensure that if a tool program fails and gets re-run, it doesn't accidentally trigger the same database update twice. That’s a critical practical detail for any service we build.
Lalam: It sounds like the combination of these three features—making it executable, keeping state safe during fixes, and intelligently choosing between execution modes—is what makes this tool program concept so powerful for complex agentic tasks.
The paper's improvements: Tom: The paper points out some specific improvements they’ve made to make TOOLPRO practical. First, they use constraint-guided program construction with lightweight formatting constraints and compiler feedback to fix common failures on the service side rather than just failing entirely client-side.
Jane: That means if the agent writes a slightly malformed tool program, the service runtime can attempt a repair using those constraints, which cuts down on those frustrating client-server back-and-forth interactions we see today.
Lu: They also introduced effect-aware replay to handle side effects under repair by maintaining a per-intent instance log of committed write outcomes and checking against it before re-executing a dynamic write call.
Meng: That mechanism for exactly once semantics on WRITE operations is key because it directly tackles the state corruption problem we discussed earlier when we think about long-horizon workflows that involve many modifications.
Lalam: And finally, they have this profile-driven consolidation rule that uses moving averages of things like round trip time and decision overhead to decide when to execute the tool program versus falling back to simple stepwise calling.
Conclusion: Tom: So, wrapping up on "Beyond Static Endpoints: Tool Programs as an Interface for Flexible Agentic Web Services," the authors show that this approach can reduce end-to-end latency by up to fifty-three point four percent and client-side traffic by up to ninety-six point one percent, with those gains getting bigger when network latency or workflow complexity is higher.
Jane: It really boils down to moving the heavy lifting of sequencing into the service, giving agents a single object to send instead of a long chain of dependent calls, which makes things much more efficient overall.
Lu: The implications for creating more robust agentic systems are significant because it provides a structured, executable way to define intent that respects the underlying service's capabilities and failure modes.
Meng: From an engineering standpoint, this gives us a concrete blueprint for building tools that can reliably handle multi-step logic with built-in safety mechanisms like effect replay.
Lalam: For our culture, I see this as a step toward allowing agents to perform much more sophisticated, reliable tasks autonomously because the interaction layer is no longer just brittle static endpoints but something truly executable and safe.
Mugeng Liu, Shuoqi Li, Yixuan Zhang, Yun Ma
cs.AI, cs.CL, cs.SE
Submitted: 2026-06-18
Updated: 2026-06-18
Comments: Accepted by ICML 2026
Code: https://github.com/morgen52/toolpro_icml26
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 87/100
The gist: Tool programs are introduced as an executable representation of tool intent to address the representational bottleneck posed by static API endpoints in agentic web services.
Key concepts
- Tool Programs
- Tool Programs are executable representations of tool intent that package multi-step interactions into one object. They include explicit effect types, allowing the agent to delegate complex workflows to a service runtime instead of relying on fragile sequences of static API calls.
- Effect-Aware Replay
- This mechanism ensures safety during iterative repairs by tracking committed WRITE operations. When re-executing code, it checks if a WRITE operation has already been successfully committed, returning the cached result instantly to prevent duplicate state changes and maintain exactly-once semantics.
- Profile-Driven Consolidation Rule
- This rule decides whether to run a complex Tool Program or use simple stepwise calls. It uses metrics like network latency (TRTT) and decision overhead (TDEC) to predict the net benefit of execution versus calling, switching modes adaptively based on predicted costs.
- WebAssembly Substrate
- TOOLPRO uses WebAssembly as a secure execution environment. This provides a strong sandbox for untrusted code generated by LLMs while allowing the system to mediate all side effects through a unified interface stub.
Terminology
Summary
Tool programs are introduced as an executable representation of tool intent to address the representational bottleneck posed by static API endpoints in agentic web services. This approach allows agents to delegate multi-step interactions to a service-side runtime, enabling features like turn reduction and effect-aware execution, which significantly improves efficiency in long-horizon workflows.
The core problem addressed is the inefficiency of stepwise endpoint sequences.
Static endpoints force an agent into a brittle sequence of calls interleaved with multiround reasoning. This interface fragments a coherent procedure into local decisions conditioned on intermediate responses, leading to inflated latency via repeated round trips and decision overhead, over-/under-fetching of data, and brittle recovery mechanisms where partial failures trigger cascading retries that can duplicate state-modifying operations and corrupt service state. The key insight is that endpoint sequences are a weak interface for expressing tool intent because they fragment the plan into local decisions.
TOOLPRO operationalizes tool programs as an interface by addressing three core challenges.
The paper proposes TOOLPRO, which packages a multi-step interaction as a single executable object with explicit effect types. It tackles: (1) Executability through constraint-guided program construction,
which uses lightweight formatting constraints and compiler/runtime feedback to resolve common failures via service-side repair. (2) Side effects under repair using effect-aware replay
that enforces exactly-once semantics for WRITE operations across iterative repair and re-execution. (3) When to consolidate by applying a profile-driven consolidation rule
that adaptively selects between program execution and stepwise calling based on predicted net benefit.
TOOLPRO's mechanism for effect safety is effect-aware replay.
To ensure retry safety under repair-driven re-executions, TOOLPRO intercepts every dynamic call at the interface boundary. READ calls are always forwarded to the service since they do not introduce new side effects. WRITE calls are replay-protected by maintaining a per-intent instance log of committed WRITE outcomes (History log H and Working log W). When a re-execution reaches a dynamic WRITE, it matches it against H; if matched, the cached outcome is returned without an external call. If no match exists, the call is emitted once and recorded in W. A conservative discipline dictates that if a repaired program changes the arguments or relative order of any committed WRITE prefix in H, replay is disabled and stepwise calling falls back to surface diagnostics.
The system employs a synthesis–project–compile–execute pipeline for execution.
TOOLPRO follows this pipeline: (1) Client-side synthesis with lightweight interface checks to reject obviously misaligned programs early. (2) Server-side projection into a constrained interface-program surface
via the deterministic projection Π(·), which rewrites interactions to the unified stub CALL(e, a) and rejects unsupported features. (3) Compilation and execution in a sandbox utilize feedback-driven, bounded in-place repair using compiler diagnostics for compilation failures and lightweight runtime traces for runtime failures. (4) A safe fallback
mechanism reverts to stepwise endpoint calling if the attempt budget is exceeded or constraints are violated, surfacing auditable diagnostics.
The system utilizes a profile-driven consolidation policy to decide execution mode.
TOOLPRO uses online profiling and decision rules based on moving averages of TRTT (client–service RTT), TDEC (per-step client-side decision overhead), and TBUILD (program construction time). It predicts the net benefit of program execution over stepwise calling using the model: ∆T = (N − 1) · (TRTT + TDEC) − TBUILD. If ∆T > 0, it executes the tool program; otherwise, it selects stepwise calling. This policy bootstraps with initial runs and enables tool program execution only when the synthesized structure is clearly multi-step, making a switch explicit based on network latency and workflow complexity.
Experimental results demonstrate substantial efficiency gains.
TOOLPRO reduces end-to-end latency by up to 53.4% and client-side traffic by up to 96.1% across diverse workflows, with gains increasing under higher network latency and workflow complexity. For complex benchmarks (cbench1–cbench3), TOOLPRO improved task accuracy from 0.60 to 0.93, and for cross-service branching (cbench4), it reduced latency by 53.4% and cut client-side traffic by 96.1%. The policy successfully switches between stepwise mode in low-latency conditions and program execution when RTT dominates the cost.
The implementation relies on WebAssembly as a secure substrate.
TOOLPRO uses WebAssembly (Wasm) as the execution substrate because it provides a strong sandbox boundary for untrusted, LLM-generated code with no ambient authority,
and a capability-style host interface that allows mediation of all side effects through the unified CALL(e, a, eff) stub.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems by implementing the TOOLPRO framework, along with a description of what these improved systems can achieve:
The implementation of TOOLPRO fundamentally shifts agentic interaction from a brittle, reactive sequence of static API calls to a robust, executable program interface. This enables the following specific capabilities for improved AI systems:
-
Consolidation of Multi-Step Workflows into Single Executable Units:
TOOLPRO allows an LLM agent to represent complex, procedural tasks (involving loops, conditionals, and intermediate data bindings) not as a sequence of N sequential decisions requiring N separate client-server round trips, but as a single tool program.
The improved AI system can perform long-horizon tasks—such as Find the best hotel matching criteria
(as shown in Figure 1)—in drastically reduced latency by executing the entire logic server-side in one submission cycle, instead of interleaving N client-side reasoning rounds.
-
Guaranteed Exactly-Once State Modification:
By enforcing effect-aware replay semantics for WRITE operations across repair and re-execution cycles, TOOLPRO ensures that state changes on the underlying service (e.g., updating a record in a database) occur at most once, even if the agent's tool program fails mid-execution and requires repair or retry.
The improved AI system can reliably perform transactional or multi-step updates in dynamic environments without causing data corruption due to duplicated side effects during recovery attempts.
-
Automated, Bounded Self-Repair of LLM-Generated Code:
The constraint-guided construction pipeline allows the server runtime to automatically catch and fix common failures (type mismatches, missing imports, incorrect control flow) in LLM-generated tool programs using compiler diagnostics and runtime traces within a fixed attempt budget.
The improved AI system can execute tools generated by an LLM with much higher reliability than current methods; instead of failing entirely or requiring a complete human re-prompt, the system self-correcting fixes minor logical errors in the program structure before attempting execution again.
-
Adaptive Efficiency Switching (Program vs. Stepwise Mode):
The profile-driven consolidation policy dynamically decides whether to use the highly efficient tool program execution or fall back to stepwise calling based on real-time cost modeling (balancing build time, network RTT, and decision overhead).
The improved AI system can achieve optimal performance across diverse conditions: in low-latency environments, it chooses the faster stepwise mode; under high latency or complex workflows where consolidation pays off, it switches to program execution to maximize speed.
-
Optimized Client-Side Traffic Management:
By transmitting a compact tool program once instead of repeatedly sending tool specifications and intermediate context across N rounds, TOOLPRO significantly reduces client-to-LLM traffic (up to 96.1% reduction in experiments).
The improved AI system can operate more efficiently under strict network constraints or high inference costs by minimizing the number of back-and-forth communication exchanges required between the agent and the LLM.
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection