Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility

arXiv:2607.14108 · cs.CL, cs.AI, cs.SE · Submitted 2026-05-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Eta Given Delta".

Tom: Tool efficiency and marginal tool utility are introduced as new quantitative metrics to evaluate how useful tools are in an LLM agent trajectory,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, to wrap up our discussion on "Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility," this paper introduces these new quantitative metrics to evaluate tool usefulness in an LLM agent trajectory, focusing specifically on marginal tool utility and aggregate tool efficiency.

Jane: That’s right, Tom. The core idea is that by defining marginal tool utility as the difference in solution likelihood when a specific tool call is included versus excluded, they can determine if a tool call actually contributes positively to the task's correctness.

Lu: What this means for us conceptually is that we have moved from guessing which tools are useful to having a mathematical framework that tells us precisely when a tool provides value, even in complex agent workflows.

Meng: The implications are that application-layer engineering teams can design leaner and more maintainable agent tool suites by quantifying the rate of useful tool calls based on whether removing a specific tool affects accuracy.

Lalam: Ultimately, the authors conclude that these well-defined quantitative metrics enable the construction of the leanest possible tool suite for an LLM agent given a particular task, which is a significant contribution to how we think about agent development.

Tom: That’s what it boils down to: using these metrics gives us a rigorous way to evaluate and optimize the set of tools an AI agent uses, moving evaluation beyond simple accuracy checks.

Jane: It suggests that empirical results using accuracy as a proxy align with the conclusions drawn from analyzing marginal tool utility and tool efficiency values, providing a new dimension for evaluating agent performance.

Lu: I think this work is going to be a springboard for future benchmark designs and agent harness engineering, pushing us toward building systems where every component has been rigorously justified by its utility.

Meng: For me, the impact will be seen in more efficient deployment of AI agents because we can systematically prune unnecessary components from the toolset before they get deployed widely.

Lalam: I hope this research inspires a culture of rigorous evaluation where we don't just build tools, but rigorously test and quantify their actual impact on the agent's success.

Conclusion: Tom: So we've looked at all this technical stuff about marginal tool utility and efficiency metrics in agent trajectories, but now we need to wrap up what it actually means for us on air.

Jane: Exactly, Tom; the paper "Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility" is basically giving us a new way to measure if an AI's tools are actually helping it solve the problem.

Lu: The authors, I think they really nailed how you can distinguish between just using a tool and using that tool in the right way for a specific task.

Meng: From my side, it’s interesting because this gives us a concrete number to look at when we're building those complex AI workflows in the real world.

Lalam: I see it as something that lets us build systems that are inherently more efficient and less wasteful when they interact with external tools.

Tom: That's the core idea, Jane; they show how you can get a direct measure of tool usefulness separate from just looking at how often an AI tries to use a tool.

Jane: It helps us understand that not every attempted tool call is equally valuable for getting the right answer, and this metric quantifies that difference.

Lu: It opens up possibilities for designing agents where we can systematically prune tools that don't contribute positively to the final result.

Meng: Imagine being able to build tool suites where you know exactly which ones are worth keeping based on this utility score.

Lalam: For me, this pushes us toward a culture of building incredibly lean and effective AI systems, focusing our development energy where it truly matters for the outcome.

Tom: It’s about moving beyond just accuracy scores to have a real understanding of how useful each part of an agent's toolset is.

Jane: And that understanding helps us see that we can create much more maintainable AI agents by focusing on those high-utility interactions.

Lu: This could lead to some really cool new ways to structure agent architectures based on these utility calculations.

Meng: It moves the conversation from "does it work?" to "how well is it working, and what parts are actually driving that success."

Lalam: And in the long run, this kind of rigorous evaluation could improve how we design and deploy AI across different applications.

cs.CL, cs.AI, cs.SE

Submitted: 2026-05-07

Updated: 2026-09-27

Importance score: 92/100

The gist: Tool efficiency and marginal tool utility are introduced as new quantitative metrics to evaluate how useful tools are in an LLM agent trajectory, aiming to provide a direct measure of tool usefulness

Key concepts

Marginal Tool Utility (∆α(αi))
This metric checks if executing a specific tool call actually improves the likelihood of solving the task correctly. It compares the success probability before and after that single tool call is made. A positive utility means the tool was useful.
Tool Efficiency (ηα(τ))
This metric measures how effective your entire set of tools is. It is calculated by dividing the number of 'useful' tool calls (those with positive marginal utility) by the total number of tool calls in a sequence. Higher efficiency means fewer unnecessary tools are being used.
Agent Trajectory (τ)
This represents the ordered sequence of actions an LLM agent takes while working on a task. It is a timeline showing alternating steps between reasoning tokens and actual tool calls, which forms the basis for measuring tool performance.
LLM-as-a-Judge
This is the method used to classify each individual tool call as having positive or non-positive marginal utility. An LLM is prompted to judge whether a specific action helped or hurt the final task accuracy.

Terminology

Summary

Tool efficiency and marginal tool utility are introduced as new quantitative metrics to evaluate how useful tools are in an LLM agent trajectory, aiming to provide a direct measure of tool usefulness distinct from accuracy proxies. This work matters because it offers a way for application-layer engineering teams to design leaner and more maintainable agent tool suites by quantifying the rate of useful tool calls based on whether removing a specific tool affects accuracy.

Defining Marginal Tool Utility

The paper introduces marginal tool utility, denoted as marginal tool utility or ∆α(αi), which is defined per individual tool call instance within an agent trajectory. Formally, it is defined as the difference between the likelihood that the task is solved correctly given tool calls up to and including αi and the likelihood that the task is solved correctly given tool calls up to and excluding αi. The core idea supporting this metric is that a tool call is useful if and only if "upon its execution the likelihood of correctness increases (∆α(αi) > 0). To determine this sign, the authors opt for using LLM-as-a-Judge to classify each tool call as either having a positive or non-positive marginal utility, yielding a label of positive or non positive."

Defining Tool Efficiency

Tool efficiency, denoted as ηα(τ), is defined based on the signs of the marginal tool utility. It is formally calculated as the ratio of the number of useful tool calls to the total number of tool calls in a trajectory, where a useful call is one with "∆α(αi) > 0. This metric allows for an aggregate measure, such as mean tool efficiency, across multiple agent trajectories, which can be used to assess the overall health of a tool suite. The authors hypothesize and empirically show that tool efficiency increases when a non-useful tool is removed from the tool suite."

Methodology and Validation

The methodology formalizes an agent trajectory τ as an ordered sequence of token sequences, alternating between reasoning tokens (t i) and tool calls (r i), resulting in a sequence of tool calls αi ≜ (ri, oi). The expected correctness of the final answer y is defined by the indicator function I(f1(y) · … · fM(y) = 1), where f j is the unit test result. To validate the definitions, experiments were conducted on the APEX-SWE Observability benchmark. Tool ablations were performed by comparing three variants: default: Full suite of MCP tools, grafana: Only Grafana/Loki MCP is available, and no-mcp: Absolutely no MCP tools are available.

Empirical Results and Findings

The empirical results demonstrate that task accuracy increases when a tool has positive aggregate tool utility, while efficiency increases with the removal of non-useful tools. For instance, in the default harness for GPT-5.3-Codex, the aggregate utility for Grafana/Loki was +25, while Mattermost and Plane had non-positive utilities (-35 and-30 respectively). This supports the hypothesis that Grafana/Loki is a useful tool, aligning with the observation that agents benefit more from reading logs (Grafana/Loki) than from secondary sources like discussions (Mattermost) or specifications (Plane). Tool efficiency showed significant increases when Mattermost and Plane were removed, suggesting they are unnecessarily included in the agent’s tool suite.

Future Directions

The paper suggests several directions for future work leveraging these metrics. These include Minimizing sub-optimal middle calls, potentially using marginal tool utility as a component of reward functions for reinforcement learning, or designing agents that can backtrack if consecutive calls show non-positive marginal utility. Furthermore, the tool efficiency metric can be used to assess the overall health of a tool suite to ensure it remains lean and maintainable. The authors also posit that harness engineering must be informed by each agent’s particular backbone model, as different models can yield the same task accuracy with different tool efficiencies.

Limitations

A limitation noted is that confidence scores returned by an LLM-as-a-Judge are not the most rigorous measure, although they indicate a general trend regarding non-positive classifications. Additionally, the tool ablations were only performed on read-only tools (MCPs), and future work seeks to assess the marginal tool utility of write tool calls.

Conclusion

The authors conclude that their new definitions enable the construction of the leanest possible tool suite for an LLM agent given a particular task, suggesting that empirical results using accuracy as a proxy align with conclusions drawn from analyzing marginal tool utility and tool efficiency values. These well-defined quantitative metrics are intended to accelerate progress in LLM research and engineering by providing new dimensions for evaluation.

Improvements for AI systems

Here are the specific improvements to AI systems that can be derived from this research paper:

  1. Improve Tool Suite Design via Predictive Utility Assessment:

  2. Enable Dynamic, Lean Harness Engineering:

  3. Optimize Agent Reasoning by Prioritizing High-Utility Tool Calls:


AI Systems Improvement Details:

  1. A system can be implemented that performs an offline or online analysis of agent trajectories to calculate the aggregate marginal utility for every tool in a given suite. This system would use the LLM-as-a-Judge methodology described (Section 3.2) to assign a positive (+) or non-positive (-) utility label to each tool call based on whether its inclusion increases the probability of solving the task correctly.

  2. The system can then automatically generate a Tool Utility Map for any given agent harness configuration, identifying which tools are consistently useful (positive aggregate utility) and which are redundant or detrimental (non-positive aggregate utility).

  3. This map would directly inform lean tool suite engineering by providing actionable data to remove non-useful MCP tools (like Mattermost or Plane in the APEX-SWE example) without sacrificing accuracy, as empirically shown in Section 4.3.

  4. For inference-time optimization, the system can be designed to monitor the sequence of tool calls and dynamically intervene if consecutive tool calls show non-positive marginal utility, prompting the agent to backtrack or pivot to an alternative reasoning path (as suggested in Section 5).

  5. The system enables More Informed Harness Engineering by allowing developers to compare different backbone models using the same task accuracy while assessing the efficiency of their respective tool suites, guiding them toward model-specific harness designs rather than a one-size-fits-all approach.

  6. For future self-improving agents, the system provides a rich reward signal during execution that is not just based on final accuracy but also on the instantaneous utility of intermediate tool calls, allowing agents to optimize their context and tool usage mid-execution.

Sources

Related papers